Crawl Efficiency: Why Pages Stay Out of the Index
Crawl efficiency is how cheaply a search engine can fetch, read and re-read a site's documents — measured most directly by the HTML crawl rate. The engine budgets crawling like any other expense: every request costs computation, and pages that cost more to process than their quality justifies are fetched, then shelved.
On this page — 5 sections
What Is the HTML Crawl Rate?
Quick answer
The share of crawl requests that successfully fetch a page's HTML. It is a direct measure of how effectively an engine spends resources on a site: the higher the crawl rate, the more the engine is described as caring about the site.
Every fetch of a page costs the engine computation, so crawling is budgeted like any other expense. The HTML crawl rate — the share of crawl requests that return real HTML — is the direct measure of how effectively a site converts that budget into understanding. The documented reading is candid: "the higher the crawl rate, the more Google cares about the site." Crawl attention is allocated by a URL scheduler that weighs how often a URL changes and how important it is, populating different crawl layers from daily to near real-time.
That makes the crawl rate a feedback signal rather than a setting. A site that is cheap to fetch, consistent across mobile and desktop, and quick to respond is revisited more; a site that wastes fetches trains the scheduler to come back less. Server response times belong to the same account: fast responses "increase crawl efficiency," and a 499 error — the client closing the connection before the server answers — means the crawler received nothing and there is no retry.
Why Do Pages Stay "Crawled – Currently Not Indexed"?
Quick answer
Because indexing is an investment decision. Pages that cost more to process than their quality justifies — thin, duplicated, temporary or unnecessary pages — are crawled but held back, and the site's overall crawl profile records the verdict.
Search Console states Crawled – currently not indexed and Discovered – currently not indexed as statuses; the doctrine reads them as verdicts. They mark pages the engine fetched (or found) and declined to index because the value did not justify the cost: thin or duplicated content, unnecessary pages, low-quality subdomains, mobile versions missing what desktop shows. The governing rule is the one that runs the whole cost account — the cost of ranking a site cannot exceed the cost of not ranking it.
"If ranking you is costlier than not ranking you, deindexing begins."
— Cost-of-retrieval doctrine
Two consequences follow. First, the status is not a penalty to appeal; it is an accounting entry to correct, by making the page worth processing or removing it. Second, the verdict is portfolio-level: a crawl profile full of declined pages shades the engine's judgment of the whole site, so pruning dead weight is not just tidiness — it changes how the next page is evaluated.
Which Numbers Measure Crawl Efficiency?
Quick answer
Three thresholds: an HTML crawl rate of at least 99%, a combined 200-and-304 status rate of 99% or higher, and every HTML crawl landing on an indexable URL that returns 200 and appears in both the sitemap and internal links.
The documented targets are unambiguous, and every one of them is checkable in server logs:
| Indicator | Target | What it tells the engine |
|---|---|---|
| HTML crawl rate | At least 99% | Whether crawl requests return real HTML — failures mean the engine learns nothing and, for 499 responses, never retries |
| Combined 200/304 status rate | 99% or higher | That the server answers correctly and that unchanged pages validate cleanly instead of being refetched |
| Crawl destination ratio | 100% of HTML crawls on indexable URLs | That crawls land on pages returning 200 that appear in both the sitemap and internal links |
| Server response times | Fast, consistently | Slow responses shrink crawl efficiency; a 499 error means the crawler received nothing and there is no retry |
One measurement habit outranks the rest: raw log files. Search Console is documented as exposing only around 30% of actual crawl data, so the log file is the honest record of what the crawler asked for, what it received, and where the budget went.
What Wastes a Crawl Budget?
Quick answer
Work the engine does for nothing: query-parameterized filter URLs that force canonicalization effort, pages deleted within two weeks, low-quality subdomains, near-duplicates, and pages with no query to trigger an index. Every unnecessary page increases the cost of retrieval.
The waste list is short and specific:
- Query-parameterized filter URLs: faceted URLs that "force Google into a huge level of computational consumption for canonicalization and URL consolidation" — filtering systems should be non-crawlable.
- Temporary pages: publishing "too many pages that will be deleted or removed within two weeks" makes crawling the site "less worthwhile"; ephemeral content should declare itself with an
unavailable_afterrobots meta. - Low-quality subdomains: the engine keeps crawling them, "whose content quality also impacts Google's judgment of your website."
- Near-duplicates: duplicate detection exists to skip redundant fetching — it "speeds up the crawling and saves bandwidth" when the site does not fight it.
- Pages with no query behind them: documents with no search demand to trigger an index entry are processed and then shelved — pure cost, no return.
How Do You Keep a Site Worth Crawling?
Quick answer
Treat the crawl report as a quality ledger: prune what stays Crawled – currently not indexed, consolidate duplicates, clean parameterized URLs, keep crawl paths short, read raw log files rather than sampling, and keep publishing natural, original content.
The discipline that keeps a site on the investment side of the ledger:
- Prune the declined pages: review
Crawled – currently not indexedandDiscovered – currently not indexedreports, then remove or consolidate what does not earn its URL — documented cases report significant click gains after pruning. - Consolidate duplicates: one strong page beats three weak reflections; overlap beyond the threshold turns documents into near-duplicates the engine must choose between.
- Close the parameter holes: make filtered views non-crawlable and keep URLs short, since short URLs are "easier to parse, resolve and request."
- Shorten the crawl paths: consistent internal links from strong pages let the crawler reach every important document in few hops.
- Watch the real logs: 499 errors, redirect chains and dead parameters show up in raw logs long before any dashboard.
- Keep the content worth fetching: natural, original, semantically precise pages are the only yield that justifies the spend — gibberish and scaled output are the fastest way to lose the budget.
This article is part of the Search Engine Understanding & SEO series — How search engines read queries, pages, layout and user behavior, explained in plain terms with service-business examples.
About the author
Mohamed Youns
Semantic SEO Engineer · Author & system developer
Mohamed Youns writes about how search engines understand content — the same standards he applies when building semantic systems at Nut Hub. nut-hub.org