In-depth guide · Updated July 29, 2026

Crawl Budget for Large Sites

This guide is for large or rapidly changing sites. If you only need the definition and a triage check, read the crawl budget quick answer. Here we cover URL inventory, host capacity, status codes, logs, and measurement.

1. The Two Components of Crawl Budget

Google describes crawl budget through two main elements that together shape actual crawl behavior:

ComponentQuestion It AnswersInfluenced By
Crawl Capacity Limit"How hard can Google crawl without breaking the site?"Server response time, error rates, sustained responsiveness
Crawl Demand"How much does Google want to crawl?"Perceived inventory, popularity, staleness, site events

Critical distinction: Crawling is retrieval; indexing is a later evaluation and consolidation step. Being crawled does not guarantee being indexed, and an exclusion status should be investigated rather than assigned to crawl budget by default.

Google defines crawl budget at the hostname levelwww.example.com and code.example.com have separate budgets. Architecture choices (subdomains vs subfolders) can change how crawl resources are partitioned.

2. When Crawl Budget Actually Matters

Google provides rough estimates to help site owners decide whether the advanced guide is relevant. They are not exact thresholds:

  • 1M+ unique pages with content that changes moderately often (weekly)
  • 10K+ unique pages with rapidly changing content (daily)
  • Large % of URLs classified as "Discovered – currently not indexed"

A smaller site whose new pages are discovered promptly usually needs only a current sitemap and routine Page Indexing checks. Page count alone does not prove a crawl-budget problem.

Evidence Worth Investigating

  • A large share of URLs reported as "Discovered – currently not indexed"; the status alone does not identify the cause
  • High crawl volume on low-value templates (filters, sort permutations) while important pages have low crawl frequency
  • Excessive redirect chains inflating crawl requests (each hop counts separately)
  • Many soft-404s and thin/duplicate URLs being re-crawled

3. Crawl Budget Optimization Techniques

TechniqueImpactEffort
Consolidate duplicate inventoryFocuses requests on canonical, useful URLsDepends on templates
Fix redirect chains (A→C not A→B→C)Avoids unnecessary hopsUsually moderate
Improve host stability and response timeCan improve crawl capacity when host load is limitingEngineering work
Use crawlable internal linksHelps discovery and site understandingTemplate and editorial work
Keep XML sitemaps currentSupplies intended URLs and accurate lastmod valuesOften automated
Block unimportant crawl spacesPrevents requests to selected URL patternsRequires careful testing
Return 404/410 for removed URLsProvides a clear removal signalTemplate dependent

Prioritized Workflow

High Impact, First Sprint

  • 1. Build URL inventory segmented by template & value — identify largest crawl sinks
  • 2. Fix systemic redirect chains and internal links pointing to redirected URLs
  • 3. Resolve soft-404 patterns (return real 404/410s or add proper content)
  • 4. Choose faceted-navigation controls from the URL's purpose; consolidate duplicates and block only spaces that should not be crawled

Parallel Engineering Track

  • 5. Improve server stability and response times — reduce DNS/network errors and 5xx
  • 6. Ensure internal linking uses crawlable <a href> patterns

Discovery and Measurement

  • 7. Keep sitemaps focused on canonical, indexable URLs with accurate <lastmod>; optional template splits can make monitoring easier
  • 8. Audit pagination — unique URLs, self-canonical per page, no fragments

4. Noindex vs Robots.txt vs 404: Decision Guide

MethodUse WhenCrawl Budget Impact
robots.txt"Don't crawl at all" — URLs never needed for searchSaves crawl requests
noindexPage must exist for users but shouldn't appear in searchNo savings — Google must crawl to see noindex
404/410Content truly removed — want URL to stop being crawledStrong signal not to re-crawl

5. Measurement: KPIs and Data Sources

Choose a comparison window long enough to cover normal crawl variation for your site, and record deployments, outages, migrations, and major publishing events:

  • Crawl requests/day — overall and by template (from Search Console Crawl Stats)
  • Average response time + host status — verify server improvements reflect in crawling
  • Discovery vs refresh mix — investigate changes alongside releases, migrations, sitemap changes, and new URL patterns
  • Crawl waste share — redirects, 4xx, soft-404 as % of total crawl
  • "Discovered – not indexed" trend — investigate changes alongside quality, duplication, and technical signals
  • Time-to-first-crawl — track with URL Inspection for new/updated pages

Data Sources

  • Search Console Crawl Stats: Macro trends, bots, response types, response-time trends
  • CDN and server logs: Direct request evidence at the layers you log; account for caching and retention gaps
  • Page Indexing report: Diagnostic indexing categories that need interpretation with other evidence
  • URL Inspection: Per-URL verification (rendered HTML, indexing status, canonical)

What This Means for You

If your site fits Google's large-site profiles, start with a URL inventory and evidence from Crawl Stats and server logs. Our technical SEO checklist can help organize related checks, but it cannot diagnose crawl allocation without your site's actual data.

Primary Source

Google Crawling Infrastructure: Optimize your crawl budget ↗

Related Guides

Frequently Asked Questions

Crawl budget is the practical limit on how many URLs a search engine crawler will (a) be able to fetch without harming your servers (crawl capacity limit) and (b) want to fetch because the content appears valuable and fresh (crawl demand). It's defined per hostname.
Google aims this advanced guide mainly at very large sites with roughly 1M+ pages that change moderately often, sites with roughly 10K+ pages that change daily, and sites with many URLs reported as Discovered – currently not indexed. Google says those figures are rough classifications, not exact thresholds.
No. Google explicitly separates crawling (retrieval) from indexing (evaluation and consolidation). Being crawled means Google fetched your page, but it may decide not to index it based on quality, duplication, or canonicalization signals.
Robots.txt blocking prevents crawling of specified URLs, which reduces crawl requests. However, Google notes this doesn't automatically 'shift' the freed capacity to other pages unless Google is already hitting your site's serving limit.
No — noindex doesn't save crawl budget because Google must crawl the page to see the noindex directive. Use robots.txt to prevent crawling entirely, or use 404/410 for truly removed content. Reserve noindex for pages that must be accessible to users but shouldn't appear in search.
No. Googlebot does not process the non-standard crawl-delay directive in robots.txt. If crawling is causing an emergency, follow Google's current guidance for reducing Googlebot's crawl rate; for ordinary optimization, improve stability and URL inventory.

Ready to Scale Your SEO?

Generate optimized content, review it with SEO checks, and publish to WordPress from one workflow.

Start 3-Day Free Trial