Search Console’s 165 unindexed pages point beyond a simple crawl shortage

Search Console’s 165 unindexed pages point beyond a simple crawl shortage

The Search Console snapshot—two alternate canonical URLs and 165 pages marked “Crawled – currently not indexed”—does not prove that requests to recrawl old pages are blocking new pages. It does expose a more important risk: Google can spend finite crawl attention on low-value or duplicate URLs while still declining to index pages it has already fetched.

Google Search Console - Page indexing - Why pages aren’t indexed
Google Search Console – Page indexing – Why pages aren’t indexed

Google Search Console can make a technical SEO problem look more mechanical than it is. A site owner sees old URLs being validated or recrawled, new URLs failing to appear in search, and a large “Crawled – currently not indexed” count. The intuitive conclusion is that the old pages are occupying a queue that should belong to the new ones. Google’s own documentation supports part of that concern but not the absolute claim. Google says it allocates crawling algorithmically, has finite crawl capacity, and warns that time spent crawling URLs it should not crawl can reduce exploration elsewhere on a site.

Yet a recrawl request does not create a documented one-for-one block on a new URL. Google also says crawl requests do not guarantee indexing, and the status “Crawled – currently not indexed” means the crawler has already reached the page but the indexing system did not store it. That distinction is decisive for the supplied snapshot: two alternate canonical URLs are usually normal, while 165 crawled-but-not-indexed URLs point first to selection, duplication, content value, rendering or site-level indexing signals—not merely a shortage of crawl slots.

The Search Console numbers describe two different problems

The two rows in the snapshot should not be added together as if they represented one indexing failure. Google defines “Alternate page with proper canonical tag” as a non-canonical version of a page for which Google recognizes another URL as the representative version. Search Console explicitly says duplicate or alternate pages generally should not be indexed and that seeing them reported can be a good sign when the intended canonical is indexed. The two alternate URLs are therefore not evidence of a crawl crisis by themselves.

The 165 URLs in “Crawled – currently not indexed” carry a different meaning. Google has already fetched those URLs. The failure, at least at the point represented by that status, happens after discovery and crawl. Google’s documentation says indexing is not guaranteed and that page content, metadata, technical accessibility and site design can affect whether a processed page enters the index. Independent reporting has also documented Google representatives linking this status to cases where pages are technically crawlable but not compelling enough for the index, although technical failures can still produce the same symptom.

That makes the 165 count a diagnostic population, not a diagnosis. Some URLs may be near duplicates, outdated pages, thin variants, soft-404-like pages, rendered pages whose main content is missing to Googlebot, or genuinely useful pages that Google has not yet selected. Search Console itself cautions that a URL in this state “may or may not” be indexed later; the status is not a permanent verdict. The practical question is therefore not “How do we force 165 URLs into Google?” It is “Which of these 165 pages deserve to be indexed, and what evidence does the site give Google that they are distinct, important and technically complete?”

Validation is a recrawl workflow, not a priority switch

The “Started” label is easy to misread. In the Page indexing report, it belongs to Search Console’s validation process after a site owner clicks “Validate Fix.” Google says it first checks a sample; if the sample passes, validation enters the Started state and Search Console works through known URLs affected by that issue. Importantly, Google says only URLs with known instances of the issue are queued for recrawling, not the whole site.

That workflow can generate recrawling, but Google does not document it as a mechanism that freezes discovery or indexing of newly published pages. Normal crawling continues alongside validation. Google even notes that it can detect fixed instances during ordinary crawling without a site owner ever starting validation. This is why the statement that validation of old pages “blocks” new indexing is too strong.

There is still a resource argument worth taking seriously. Google’s July 2026 crawl-budget documentation says site crawl capacity is finite and advises site owners to manage URL inventory because excessive crawling of URLs that should not be crawled can mean Google’s crawlers do not explore the rest of the site as effectively. The defensible thesis is not that the validation button steals a fixed slot from every new page; it is that a noisy inventory can make Google spend scarce crawl effort on URLs with little search value. On a large site, or one with limited crawl demand or server constraints, that can lengthen the path between publishing, discovery, crawling and eventual indexing.

Crawl capacity and crawl demand decide where Googlebot spends time

Google describes crawling as an algorithmic scheduling problem. Googlebot decides which sites to crawl, how often to crawl them and how many pages to fetch. Its current crawl-budget documentation separates the system into capacity—the amount the site and Google’s infrastructure can sustain—and demand, which reflects Google’s interest in recrawling or discovering particular URLs. This matters because a site does not own a fixed daily allowance that can be manually reassigned URL by URL.

Request Indexing in URL Inspection is therefore a signal, not a command. Google says a recrawl request can take days to weeks and does not guarantee inclusion in search results; it also warns that repeated requests for the same URL will not make crawling happen faster. For many URLs, Google recommends using a sitemap rather than submitting URLs one at a time. The purpose of the sitemap is to help discovery and tell Google which URLs the publisher considers important, while the <lastmod> value can help scheduling when it accurately reflects a substantial page change.

The architecture of the site supplies another set of priorities. Google discovers pages through links, and its documentation states that crawlable internal links help it find new URLs. A newly published page linked prominently from the home page, category hub and relevant high-value pages sends a very different discovery signal from an orphan URL that exists only in a sitemap. If old URLs dominate internal links, parameter spaces, archives or sitemap inventories, they can keep presenting themselves as crawl candidates even when the business wants Google to focus elsewhere.

That is the hidden mechanism behind many “crawl budget” complaints. The issue is often not a single manual request. It is the cumulative URL inventory that the site exposes to the crawler.

Canonicals reduce duplication but do not erase crawl cost immediately

Canonicalization is designed to consolidate duplicate or very similar URLs around one representative version. Google treats rel="canonical" as a strong signal, and it also uses redirects and sitemap inclusion as canonicalization signals. If the two “Alternate page with proper canonical tag” URLs in the snapshot point to the intended canonical pages, their non-indexed status is expected behavior rather than something to validate into disappearance.

The subtlety is that canonicalization happens after Google learns enough about the URLs to cluster them. Duplicate URLs can still consume discovery, fetch and processing resources before Google consolidates them. Google’s crawl-budget documentation now explicitly recommends consolidating duplicate content so crawling focuses on unique content rather than merely unique URLs. That is especially relevant to faceted navigation, tracking parameters, print versions, protocol or hostname variants, and CMS-generated archives that expose many paths to substantially the same content.

A canonical tag is therefore not a crawl-blocking directive. If a URL remains internally linked, listed in sitemaps or otherwise discoverable, Google can revisit it even when it canonicalizes elsewhere. The publisher’s job is to make the preferred URL structurally obvious: use consistent internal links, include preferred canonicals in sitemaps, redirect obsolete equivalents where a permanent move exists, and avoid producing unnecessary duplicate paths. Google’s canonical guidance says sitemap URLs should be the preferred canonical versions, while internal consistency strengthens the signal.

This distinction explains why aggressively “fixing” the two alternate URLs may be wasted work. If they are legitimate alternates, there is nothing to index. The higher-value audit is the 165-page group: determine whether those are canonical URLs that the site genuinely expects to rank, then investigate why Google fetched them without retaining them.

The 165 crawled pages shift attention from crawling to selection

“Crawled – currently not indexed” is often treated as a crawl-budget warning because it appears in the indexing report. But the words themselves place the URL beyond the first bottleneck: Googlebot reached the page. Google’s Search documentation says the indexing stage processes text, key content signals and canonical relationships, then may store the canonical page in the index; not every processed page is stored.

For the supplied 165 URLs, a useful audit starts with segmentation rather than mass resubmission. Group them by template and purpose: product or service pages, articles, tag pages, pagination, location pages, old campaign pages, parameter URLs and any autogenerated variants. Then inspect representative URLs from each group with URL Inspection. Google’s tool shows the indexed version, whether the page is indexable, the last crawl information and the canonical selected by Google; its live test can reveal whether Googlebot can currently obtain the intended content.

If many URLs share a template and the rendered content available to Google is mostly boilerplate, the problem can look site-wide even though the root cause is templated. Search Engine Journal reported a 2026 example in which a live inspection exposed a migration-related rendering problem: Google could crawl the URLs, but the rendered content available in testing was largely missing. The same report noted Google discussion around “commodity” pages that add little beyond material already available elsewhere. Those are different causes with the same Search Console label, which is why clicking Request Indexing on all 165 URLs can create activity without resolving the reason they were excluded.

The strongest candidates for intervention are canonical pages tied to real search demand, unique business information or original editorial value. Pages whose only purpose is to duplicate another route should instead be consolidated or removed from the indexable inventory.

Old URLs become expensive when the site keeps advertising them

An old page is not automatically a problem. Evergreen pages, historic documentation and older articles can remain useful and deserve regular crawling when they still satisfy user demand. The problem begins when obsolete or duplicate URLs continue to look important to the crawler because the site keeps linking to them, listing them in XML sitemaps, changing superficial timestamps or generating new parameter combinations.

Google says a sitemap communicates which pages a publisher considers important and that accurate <lastmod> data can be used when it reliably reflects substantive updates. Artificially refreshing dates on unchanged content therefore creates the wrong signal. Google’s crawl documentation is equally clear that reducing unnecessary URL inventory can help crawling concentrate on unique content.

For URLs that are permanently gone, meaningful HTTP status codes matter. Google recommends 404 or 410 for content that no longer exists and 301-style permanent redirects when a page has moved to a relevant replacement. Google may continue revisiting known 4xx URLs for a period because disappearance can be temporary, so cleanup does not instantly eliminate all crawl traffic. Search Console’s own indexing documentation explains that Google keeps trying known unavailable URLs for a while.

Blocking is also easy to misuse. A robots.txt disallow stops crawling, but it is not the same as noindex; if Google cannot crawl a URL, it cannot see a page-level noindex directive. Removing low-value URLs from crawl paths requires choosing the right control for the intent: redirect true replacements, return 404/410 for genuinely removed content, use noindex when a reachable page should stay out of search, and use robots rules when the objective is specifically to prevent crawling of spaces that do not need Googlebot access.

Crawl Stats can prove whether old URLs are consuming attention

The Page indexing report alone cannot establish that old URLs are crowding out new ones. To test that thesis, use Search Console’s Crawl Stats report. Google says the report shows crawling history, including request volume, download size, response time, host status and breakdowns of crawl requests. The evidence you need is a pattern of crawler behavior, not merely a count of excluded URLs.

Start by looking at crawl requests by response code and file type, then inspect sample URLs from server logs if available. A high share of requests to obsolete parameters, repeated redirects, duplicate archives or low-value filters would support the argument that crawl effort is being spent inefficiently. Conversely, if Googlebot is already crawling the new URLs quickly and those URLs later land in “Crawled – currently not indexed,” the bottleneck is not discovery or crawl capacity; it is selection after crawling.

Host health matters as a competing explanation. Google says availability problems can prevent crawling at the rate its systems otherwise want, while improving server availability does not automatically create more crawl demand. The Crawl Stats report therefore needs to be read alongside response times, 5xx rates and host status. A slow or unstable server can suppress effective crawling even when URL inventory is clean.

This is also where the user’s proposed causal chain can be tested rather than assumed. Compare publication dates of new pages with their first crawl dates, then compare that period with spikes in validation-related or legacy-URL crawling. If new pages wait to be crawled while old low-value URLs dominate requests, crawl allocation is a credible contributor. If new pages are crawled promptly but remain unindexed, the 165-page problem sits downstream. Google’s pipeline separates discovery, crawling and indexing, and the remedy should follow the stage that is actually failing.

The strongest fix is to shrink ambiguity, not submit more requests

For a site showing the supplied pattern, the first operational decision should be to stop treating every excluded URL as a defect. The two proper-canonical alternates should be checked for intended canonical targets, then left alone if the canonicalization is correct. Google explicitly says the objective is to index the canonical version of important pages, not achieve 100% indexing of every known URL. Success is a cleaner indexable set, not a zero in every exclusion category.

The 165 crawled-but-not-indexed URLs deserve triage. Select a sample across templates and commercial importance. For each one, verify an HTTP 200 response, absence of accidental noindex, rendered primary content, self-canonical or intended canonical behavior, crawlable internal links, sitemap inclusion only when the page should be indexed, and meaningful differentiation from existing indexed pages. Google’s technical requirements state that indexable pages need a successful response and indexable content; its Search documentation adds that indexing can still fail because of page content, metadata or site design.

Next, remove contradictory signals. Do not list obsolete, redirected, duplicate or non-indexable URLs as if they were priority pages. Keep sitemaps focused on canonical URLs that matter, and use accurate modification dates. Strengthen internal links to new high-priority pages from crawlable hubs rather than relying on repeated manual submission. Reserve Request Indexing for important URLs after a genuine fix or meaningful content change; Google itself says the tool is quota-limited in some workflows and should be reserved for important pages.

Finally, separate two scorecards: crawl latency and index acceptance. A new URL that is discovered and crawled within a reasonable interval but not indexed needs a different intervention from a new URL that Google has not crawled at all. That separation prevents teams from spending days “optimizing crawl budget” when the actual problem is content selection—or rewriting content when server health and URL explosion are holding the crawler back.

The evidence does not support a simple queue-blocking story

The strongest version of the warning—“requesting crawls of old pages will block indexing of new pages”—goes beyond what Google documents. Search Console validation does queue affected URLs for recrawling, but Google says only the known affected URLs are included in that validation workflow, and normal crawling can independently detect fixes. Manual Request Indexing is also not presented as a fixed queue where one old URL necessarily displaces one new URL. Google repeatedly describes crawling as algorithmic and driven by capacity and demand.

The weaker claim is both more accurate and more useful: unnecessary old, duplicate or low-value URLs can consume crawl resources and dilute the signals that tell Google what deserves attention. Google’s 2026 crawl-budget guidance says that if crawlers spend too much time on URLs they should not crawl, they may explore the rest of the site less effectively. That creates a plausible mechanism for delayed discovery or recrawling on sufficiently large, noisy or constrained sites.

But the supplied 165 count demands a second conclusion. Those URLs have already crossed the crawl barrier. If they are important canonical pages, the next investigation belongs to index selection: content uniqueness, usefulness, rendering, canonical clustering and site architecture. Recent specialist reporting on Google’s explanations of “Crawled – currently not indexed” reinforces that quality and technical rendering can both be involved, so no single cause should be inferred from the status alone.

The forward judgment is conditional. If Crawl Stats and server logs show old URL families absorbing a large share of requests while new URLs wait unvisited, reduce that inventory aggressively. If new pages are fetched quickly and then join the 165, stop chasing the crawler and improve the reasons Google has to keep those pages in the index. That distinction turns an alarming Search Console number into a testable SEO diagnosis.

Questions site owners should ask about these indexing statuses

Does “Crawled – currently not indexed” mean Google could not crawl the page?

No. It means Google crawled the URL but did not index it at that time. Google says the URL may or may not be indexed later.

Does requesting indexing for an old URL block a new URL?

Google does not document a one-for-one blocking relationship. Crawl capacity is finite, however, and unnecessary crawling can reduce the attention available for other URLs on a site.

Is “Alternate page with proper canonical tag” an error?

Usually not. It generally means Google recognized the URL as an alternate and selected another URL as canonical. The intended canonical page, not every duplicate, is the page that should be indexed.

What does “Started” mean in Search Console validation?

It means the initial validation sample passed and Search Console has begun checking known affected URLs. Google says those known issue URLs are queued for recrawling during validation, not the entire site.

Should all 165 crawled-but-not-indexed URLs be resubmitted?

No. First determine which URLs genuinely deserve indexing. Repeated crawl requests do not guarantee indexing, and Google recommends prioritizing useful, canonical pages.

Where can I see whether Googlebot is wasting requests on old URLs?

Use the Crawl Stats report in Search Console and, where available, server logs. Crawl Stats shows request history, response patterns, host status and other crawl activity that can reveal which URL classes Googlebot is fetching.

Should old pages be blocked in robots.txt?

Only when the objective is to prevent crawling. robots.txt is not a reliable substitute for noindex, because Google must crawl a page to see a page-level noindex directive.

Do sitemaps force Google to crawl and index new pages?

No. Sitemaps help Google discover URLs and understand which pages the publisher considers important, but they do not guarantee crawling or indexing. Accurate values can help Google schedule recrawls.

What should be checked first on an important page that was crawled but not indexed?

Check the live rendered page, HTTP status, indexability directives, canonical selection, internal linking and whether the content is sufficiently distinct and useful. URL Inspection is the primary Google tool for page-level diagnosi

Author:
Jan Bielik
CEO & Founder of Webiano Digital & Marketing Agency

Google Search Console - Page indexing - Why pages aren’t indexed
Google Search Console – Page indexing – Why pages aren’t indexed
Search Console’s 165 unindexed pages point beyond a simple crawl shortage
Search Console’s 165 unindexed pages point beyond a simple crawl shortage

This article is an original analysis supported by the sources cited below

Crawl Budget Management

Google’s current crawl-budget documentation supplied the central evidence on finite crawl capacity, crawl demand, duplicate URL consolidation and the risk of spending crawl activity on URLs that do not merit it.

Page indexing report

Google Search Console documentation defined the supplied indexing statuses, explained why alternate canonical URLs are normally excluded, and documented the validation workflow and “Started” state.

Ask Google to recrawl your URLs

Google’s recrawl guidance established that crawl requests do not guarantee indexing, may take time, and should not be treated as commands that force immediate inclusion.

In-Depth Guide to How Google Search Works

Google’s search-process documentation supported the distinction between URL discovery, crawling, processing and indexing, including the fact that not every processed page is indexed.

What is URL canonicalization

Google’s canonicalization documentation defined canonical URLs and explained why duplicate URL variants are consolidated around a representative page.

How to specify a canonical URL with rel=”canonical” and other methods

This guidance supported the discussion of canonical signals, sitemap consistency and the use of preferred URLs across site architecture.

Learn about sitemaps

Google’s sitemap documentation established that sitemaps improve discovery and communicate important URLs without guaranteeing either crawling or indexing.

Build and submit a sitemap

This source supplied Google’s guidance on accurate <lastmod> values and their potential use in scheduling recrawls after meaningful changes.

URL Inspection tool

Google’s URL Inspection documentation supported the recommended page-level checks for indexed versions, indexability and live testing.

Crawl Stats report

This Search Console reference established what crawl-history and host-level evidence can be examined when testing whether old URL families are consuming crawler attention.

Google Search technical requirements

Google’s technical requirements supported the checks for successful HTTP responses, accessible pages and indexable content.

SEO link best practices for Google

Google’s link guidance supported the role of crawlable internal links in URL discovery and the recommendation to expose new priority pages through site architecture.

Block Search indexing with noindex

Google’s noindex documentation supplied the distinction between preventing crawling with robots.txt and preventing indexing with a page-level or header directive.

Troubleshoot Google Search crawling errors

Google’s troubleshooting guide supported the distinction between server availability constraints and crawl demand when diagnosing insufficient crawling.

Why your pages are stuck in crawled-currently not indexed

Search Engine Journal’s July 2026 reporting provided independent context on technical rendering failures and Google event discussion around low-differentiation or commodity content in this status.

Google explains reasons for crawled not indexed

Search Engine Journal’s reporting on Google explanations supplied independent context for why similarity and content value can be involved after a URL has already been crawled.

Citing this article? Brief excerpts are welcome. Please credit Webiano.digital, name the author where stated, and include a link to https://webiano.digital and to this original article. Full or substantial republication requires prior written permission. Read our Copyright and Content Use Policy.

This article was prepared with the assistance of artificial intelligence tools. The content underwent expert human review, and Webiano Digital & Marketing Agency assumes editorial responsibility for its final version and publication.