Glossary
crawl budget
In short: Crawl budget is the number of pages a crawler is willing and able to fetch from a site in a given period, set by server capacity and perceived value.
Crawl budget is the practical limit on how much of a site a crawler will fetch. It combines a crawl rate limit (how fast the crawler can request pages without hurting the server) and crawl demand (how much the crawler wants those pages, based on popularity and freshness). Googlebot has one; so do GPTBot, ClaudeBot, PerplexityBot and OAI-SearchBot, each with its own rules.
For small sites with a few hundred pages, budget is rarely a problem. It becomes one when a site generates many low-value URLs: faceted navigation, calendar pages, session parameters, infinite scroll, duplicate paths without canonicals, or redirect chains. The crawler spends its allowance on those and may never reach or refresh the service pages you actually want cited.
AI crawlers can be more aggressive than Googlebot, which tempts site owners to block them outright. Throttling via a CDN or a Crawl-delay directive, which some bots respect, keeps you reachable without the load.
Ways to protect budget:
- Fix redirect chains and 4xx/5xx pages.
- Add canonicals and block parameterised duplicates in
robots.txt. - Keep an accurate sitemap and remove orphan pages.
- Improve server response time.
Seoptist's audit (100 pages on Starter, 500 on Growth and per Agency client) surfaces each of these as checklist items.