Google’s own guidance on generative AI features opens with a claim that most GEO vendors quietly skip past: there are no additional requirements to appear in AI Overviews or AI Mode, and no special optimizations are necessary. The company states plainly that its AI features draw on the same ranking and quality systems as classic Search, that structured data is not required for generative AI search, that no special schema.org markup exists for it, and that llms.txt files are not used by Google Search and neither help nor harm visibility.
Table of Contents
GEO in 2026 means earning retrieval, not gaming a ranking
That is an uncomfortable starting point for an industry that has spent two years selling AI-search packages. It is also the most useful starting point available, because it tells you where the leverage actually sits. If the retrieval layer feeding a generative answer is largely the same retrieval layer feeding the blue links, then the practical GEO checklist is not a separate discipline bolted onto SEO. It is a reordering of SEO priorities around a different consumption pattern, plus a genuinely new set of access-control, measurement and licensing decisions that classic SEO never had to make.
The consumption pattern is what changed. Pew Research Center’s April 2025 field study, built on 68,879 unique Google searches from 900 U.S. adults, found that users clicked a traditional search result on 8% of visits to result pages carrying an AI summary, against 15% on pages without one. Clicks on the sources cited inside the summary happened on 1% of visits. Sessions ended entirely on 26% of pages with an AI summary versus 16% without. Ahrefs’ later analysis of AI Overview citations found that 38% of them come from pages already ranking in the top 10 for the query — meaning the citation pool overlaps heavily with the classic ranking pool, but not completely, and the difference is where GEO lives.
By mid-2026 the surface itself had grown past the point where treating it as an edge case makes sense. Google told developers at I/O 2026 that AI Mode had passed one billion monthly users, with query volume doubling quarter over quarter since launch, and that Gemini 3.5 Flash had become the default model behind it globally. Gemini 3 became the default for AI Overviews in January 2026. Search Console gained dedicated generative AI performance reports in June 2026, which means the measurement gap that made GEO arguments unfalsifiable for two years has partially closed.
What a practical universal checklist has to do is separate three categories of work that get sold as one. The first category is table stakes that were already SEO and remain SEO: crawlability, indexability, rendering, clear page structure, factual accuracy, entity clarity. The second is genuinely new operational work: crawler-level access decisions across a dozen user agents, prompt-level visibility sampling, log-based retrieval monitoring, licensing posture, agentic-commerce data readiness. The third is folklore: llms.txt files, keyword-shaped rewriting for LLMs, schema deployed as a persuasion device, purchased “brand mentions”, content chopped into artificially tiny chunks.
The folklore category matters more than it should, because it consumes budget that would produce measurable results elsewhere. A 2026 critical survey of 45 GEO studies published between November 2023 and July 2026 concluded that keyword stuffing shows null or negative effects, that generic formatting recipes generalize poorly across engines, and that fixed heuristics mostly fail — the C-SEO Bench evaluation it cites found only three of 54 method–domain combinations produced statistically positive results. The same survey found that the strongest and most repeatable signal remains plain query–document relevance.
So the checklist that follows is deliberately conservative about mechanism and aggressive about execution. It assumes you cannot control which model retrieves you, in what order, or how it paraphrases you. It assumes you can control whether you are fetchable, whether your claims are extractable, whether your identity is unambiguous, whether third parties describe you accurately, and whether you can prove any of it changed. That is a smaller ambition than “rank first in ChatGPT”, and it is the only version that holds up when someone audits the work.
Anatomy of a generative answer from query to citation
Understanding the pipeline is not academic. Every item on a defensible checklist maps to a specific stage, and the stages have different failure modes.
A generative answer in Google’s AI Mode or an AI Overview begins with query fan-out: the system decomposes the user’s prompt into multiple related sub-queries and issues them against the index in parallel. Google describes this openly in its documentation on AI features, saying fan-out lets the system “show a wider and more diverse set of helpful links” than a single query would. Practitioner analyses put typical fan-out at somewhere between a handful and a couple of dozen sub-queries depending on prompt complexity, though the exact count is not disclosed and varies by model version.
The consequence is immediate and often missed. You are not competing for one query. You are competing for a set of sub-queries you never see, most of which are more specific than the prompt the user typed. A page that ranks eleventh for the head term but first for a narrow sub-question can enter the candidate pool through the side door. A page that ranks third for the head term and has nothing to say about any of the sub-questions may not enter at all.
The second stage is retrieval and candidate assembly. Documents that match sub-queries get pulled into a working set, usually at passage rather than whole-page granularity. The retrieval index for Google’s AI features is the Search index; for ChatGPT search it is OpenAI’s own index built by OAI-SearchBot plus live fetches; for Perplexity it is a mix of its own crawl and third-party search APIs; for Copilot it is Bing’s index. These indexes do not agree with each other. The 2026 survey found URL-level Jaccard similarity between Google’s classic results, AI Overviews and Gemini answers of only 0.11 to 0.18, and found that across Bing Chat and Perplexity only 26% of domains appeared in both. A checklist that treats “AI search” as one destination will produce visibility in one engine and blindness in another.
The third stage is context construction. The retrieved passages are ordered and packed into the model’s context window, and this ordering matters more than most content advice acknowledges. The survey identifies document position within the context window as one of only two consistently predictive factors, alongside query–document relevance. You cannot control position directly. You can influence it through the relevance signals that determine it, and you can make sure that when your passage is included, it contains a self-contained, attributable claim rather than a sentence that only makes sense after three paragraphs of setup.
The fourth stage is generation with citation. The model writes an answer and attaches source links. This is where the widest gap between publisher expectation and system behaviour appears. Citation is not proportional to contribution. A page can be retrieved, read, and used to shape an answer without being cited, and a page can be cited without having contributed much. Research the survey collects found only 51.5% of generated sentences fully supported by their cited sources in one early study, with 74.5% of citations correctly supporting the proposition they were attached to, and a 2026 analysis putting roughly 11% of claims as insufficiently supported. Citation is a lossy, partly stochastic layer on top of retrieval.
The fifth stage is display, and it varies by surface. Google renders links as chips, side panels, and inline anchors depending on device and answer type. ChatGPT renders inline numbered citations plus a source list. Perplexity leads with numbered sources. The click-through economics differ enormously by placement, which is why the Pew figure of 1% clicks on summary sources and the industry’s more optimistic per-engine numbers can both be true of different surfaces.
The sixth stage — the one that actually pays — is behaviour: the click, the branded search afterwards, the direct visit next week, the conversion. This is the stage with the weakest evidence base in the entire field. The survey rates behavioural outcomes as having “very low” support, noting a single suggestive quasi-experiment estimating a 1.82× traffic multiplier with confidence intervals wide enough to include almost nothing.
Every stage is a separate checklist item with a separate diagnostic. Not fetchable is stage zero. Not retrieved is a relevance and coverage problem. Retrieved but never cited is an extractability and authority problem. Cited but no clicks is a placement and answer-completeness problem. Clicks but no conversions is a landing-page problem that has nothing to do with GEO at all. Teams that skip the diagnosis and jump to tactics usually fix the wrong stage.
Evidence behind the famous 40 percent visibility claim
Almost every GEO pitch deck traces back, directly or through three layers of citation, to one paper: GEO: Generative Engine Optimization by Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande, presented at KDD 2024. The paper introduced the term, built GEO-bench as a benchmark of user queries across domains, and reported that its optimization framework “can boost visibility by up to 40% in generative engine responses.”
That sentence has been doing an enormous amount of commercial work for two years, and it does not mean what most people repeat. The 2026 critical survey reconstructs the underlying figure: the headline number is a relative gain in Position-Adjusted Word Count from 19.3% to 27.2%, measured inside a fixed five-document context window. Position-Adjusted Word Count is a proxy for how much of the generated answer your document accounts for, weighted by where it appears. It is a reasonable research metric. It is not clicks, not citations, not customers.
Three limitations follow, and a practical checklist has to be built around them rather than in spite of them.
First, the experiment holds retrieval constant. The five candidate documents are given. The paper measures what happens to your share of the answer once you are already in the context window. It says nothing about whether you get there. The survey calls this the field’s central scope confusion: most studies measure conditional effects given retrieval, while practitioners want unconditional effects on discoverability. Optimizing content for extraction when you are not being retrieved is polishing a document nobody opens.
Second, the gains are domain-dependent. The original paper says so explicitly — efficacy varies across domains, and domain-specific methods are needed. The C-SEO Bench follow-up sharpened this into an uncomfortable finding: across 54 method–domain combinations, only three showed statistically positive results. Techniques that work in one vertical routinely produce nothing in another, which is why cross-industry GEO playbooks underperform their case studies.
Third, and most consequential for anyone selling GEO as a durable advantage, the gains erode with adoption. C-SEO Bench documents congestion effects approaching a zero-sum game: when every competitor adds statistics, quotations and authoritative framing, the relative advantage of doing so collapses toward zero. The absolute quality of the corpus rises, which is arguably good for readers, but the differential visibility gain that justified the invoice disappears.
What survives all three critiques? The survey’s own summary is worth reading precisely: “already-retrieved content can causally alter its citation or use, but no reviewed technique shows a stable, longitudinal, cross-platform causal effect on organic discoverability or downstream behavior.”
That is the honest evidentiary floor for GEO in 2026. Two things have strong support: query–document relevance predicts retrieval, and position within the context window influences citation. Several things have moderate support: extractable evidence in the form of statistics, definitions and quotations can make a passage easier to use; structured data and recent dates show context-dependent benefits; learned optimization outperforms fixed heuristics in controlled settings. Several things have poor or negative support: keyword stuffing, universal formatting recipes, and any claim of a stable cross-engine visibility lever.
There is a methodological reason the evidence stays thin, and it is not laziness. The survey reports that repeating identical queries within 24 hours produces Jaccard similarity of only 0.34 to 0.42 across engines — meaning roughly half to two-thirds of the cited sources change between two runs of the same prompt on the same day. To detect a real effect against that noise, the authors recommend seven to eight repetitions per prompt. Almost no commercial GEO study does this. Almost every “we increased AI visibility 300%” case study is measuring engine variance and calling it causation.
Two further biases inflate published results. Studies routinely discard responses that contain no citations, which removes the denominator and makes citation look more attainable than it is. And LLM judges used to score visibility often come from the same model family as the generator, introducing circularity — the model rewards text that looks like its own output.
None of this means GEO is fake. It means the honest version is narrower and more operational than the marketed version. You can reliably improve whether you are fetchable, whether your claims are extractable, whether your entity is unambiguous, whether third parties describe you accurately, and whether you can measure any of it. You cannot reliably promise a citation rate. A checklist built on the first set of promises is defensible. One built on the second is a liability the first time a client asks for a controlled test.
Seven distinct things people mean by AI visibility
The word “visibility” is where most GEO conversations go wrong, because five people in a meeting use it to mean five different things. The 2026 survey proposes a decomposition that is worth adopting as working vocabulary, because it turns arguments into measurable questions.
The survey’s visibility vector separates seven components. Discoverability is the probability that your document is retrieved at all for a relevant query. Context exposure is where it lands in the model’s working set and how much room it gets. Citation probability is whether the model attaches a visible link to you. Prominence is how much of the answer you account for and how early you appear in it. Absorption is whether your facts made it into the answer regardless of attribution. Fidelity is whether the claims attributed to you are accurate. Behaviour is whether anything happened afterwards — a click, a signup, a purchase.
Read that list against how AI visibility tools are sold and the mismatch is obvious. Most tools measure citation probability and brand mention frequency, which are two of seven components, and the two most sensitive to engine variance. Absorption — the case where a model uses your comparison table to build its answer and cites nobody — is invisible to almost every commercial tracker and is probably the most common outcome for well-structured commercial content.
The survey pairs the vector with a nine-level evidence hierarchy running from easiest to hardest to establish: activation (whether the AI answer triggered at all), retrieval, mention, citation, prominence, coverage, absorption, fidelity, behaviour. Confidence is rated high at the lower levels and very low at behaviour. That gradient should determine what you promise and what you report.
Activation deserves more attention than it gets. The survey reports that Google AI Overviews activate on 13.7% of queries overall but 64.7% of question-formatted queries. If your category is dominated by navigational and transactional phrasing, the AI answer may simply not appear for most of your demand, and the correct GEO investment is close to zero. If your category is dominated by “how do I”, “which is better”, “is it safe to” phrasing, activation is near-universal and the investment case is strong. Measuring activation rate across your own keyword set is the cheapest, highest-value GEO diagnostic available, and it takes an afternoon.
Fidelity is the component most likely to become a legal and brand-safety concern rather than a marketing one. When an assistant misstates your pricing, invents a product limitation, or attributes a competitor’s failing to you, the damage is real and the remedy is unclear. A 2026 analysis found credible-source shares between 71.4% and 86.3% depending on the assistant and topic, which leaves a substantial residue of answers grounded in weak sources. Monitoring fidelity means checking what engines say about you, not just whether they link to you.
Behaviour is where most reporting quietly cheats. Because referrer data from AI surfaces is incomplete — some engines pass a referrer, some strip it, some route through intermediary domains, and Google’s AI features report inside the ordinary Search performance data rather than as a separate channel — attributing revenue to AI visibility requires modelling rather than counting. Anyone presenting a clean AI-attributed revenue figure without explaining the model behind it is presenting an estimate as a measurement.
The practical instruction is to pick one component per objective and measure that. Brand teams should track mention and fidelity. Demand-generation teams should track activation and citation on commercial-intent prompts. Technical teams should track retrieval proxies from server logs. Executive reporting should track behaviour with the confidence intervals attached. Collapsing all seven into a single “AI visibility score” produces a number that moves for reasons nobody can explain, which is worse than no number.
Crawl access as the first and least glamorous requirement
Every sophisticated GEO tactic is worthless if the fetch fails. This is the least discussed and highest-frequency failure in real audits, and it has become more common rather than less as infrastructure teams have tightened bot controls in response to AI crawling volumes.
The scale of that crawling explains the tightening. Cloudflare’s network data shows 57.4% of traffic to HTML across its network coming from bots, with training-related crawlers accounting for 50.6% of total traffic. More than half of crawler requests re-fetch pages that have not changed. The crawl-to-referral ratios Cloudflare published are the numbers that changed publisher sentiment: Anthropic’s crawler at roughly 38,000 pages fetched per referral sent back, OpenAI’s at about 1,091 crawls per referral. Compared with the rough parity of the classic Googlebot bargain, that is not an exchange, and publishers have responded accordingly.
The result is an environment where blocking is often the default rather than a decision. WAF rule sets ship with AI crawler categories pre-enabled. CDN bot-management products classify unfamiliar user agents as suspicious. Cloudflare moved to blocking AI crawlers by default for new domains. Rate limiters trip on legitimate crawlers that fetch in bursts. A marketing team can spend a quarter on GEO content while the security team has been returning 403 to OAI-SearchBot the entire time, and nothing in either team’s dashboard will surface the conflict.
The audit is mechanical and should be run before anything else. Fetch a representative sample of your important URLs while presenting each of the AI user agents you care about, from an IP that is not on an allowlist, and record the status code and the response body length. Do this for the home page, a category page, a product or service page, a long-form guide, a pricing page, and a page behind any consent or geo logic. Compare against the same fetch presenting Googlebot and presenting an ordinary browser string.
Four distinct failure patterns turn up repeatedly. Hard blocks return 403 or 429 to specific user agents. Soft blocks return 200 with an interstitial, a challenge page, or a near-empty shell — the worst kind, because monitoring that only checks status codes reports success. Geo-conditional serving returns different content depending on the requesting IP’s country, and most AI crawlers fetch from a narrow set of data-centre regions, so a site that serves a country-selector page to unrecognised regions serves that page to the crawler forever. Consent-wall serving returns a cookie banner with the article body absent, common in the EU and directly self-defeating for European publishers.
Verification is the piece teams skip. Both OpenAI and Anthropic publish IP ranges as JSON files for each of their crawlers, and Google publishes ranges for Googlebot and its other fetchers. Blocking by user-agent string alone is unreliable in both directions: bad actors spoof crawler strings to bypass rate limits, and legitimate crawlers get blocked by rules aimed at spoofers. Allow by verified IP range plus user agent, not by user agent alone, and revisit the ranges quarterly because they change.
Robots.txt hygiene deserves its own pass. Common defects: wildcard disallow rules inherited from a staging configuration; disallowing /wp-content/ or /assets/ in a way that blocks the CSS and JS needed to render the page; disallowing internal search or faceted paths with patterns broad enough to catch real content; conflicting group precedence where a specific user-agent group unintentionally overrides a more permissive wildcard group. Because robots.txt group matching selects the most specific matching group and ignores the others, adding User-agent: GPTBot with a single Disallow: line does not inherit your general Allow rules — a subtlety that has produced real accidental blocks.
Crawl budget is the other half of access. Google’s guidance calls out crawl budget management specifically for large sites and sites that update frequently. If an AI system’s index of you is eighteen months stale, your carefully restructured pages do not exist from its point of view. Sitemap accuracy, lastmod honesty, internal linking depth, and server response time all determine refresh rate. A page that takes four seconds to respond gets crawled less often than one that takes 400 milliseconds, and in a retrieval-driven world that translates directly into a staler representation of your business.
Two practical additions for 2026. Many AI systems now perform live, user-initiated fetches at query time — OpenAI’s ChatGPT-User, Anthropic’s Claude-User, Perplexity’s user-triggered fetcher. These are documented as user-initiated actions that may not follow robots.txt rules, and they fail differently: they time out. A page that renders in eight seconds may simply not return in time for the assistant to use it, regardless of your robots.txt posture. Time-to-first-meaningful-content is now a retrieval variable, not just a user-experience one.
The second addition is agentic access. Google’s optimization guide points site owners toward agent-friendly website practices and flags emerging protocols for agent-mediated commerce. Agents driving headless browsers behave like impatient users with no tolerance for interstitials, and they will become a measurable share of non-human traffic. Whatever your posture on training crawlers, deciding it deliberately for agents is now part of the same access review.
A crawler control map for search, training and user fetches
The single most useful thing a technical team can do for GEO in an afternoon is replace their all-or-nothing bot policy with a deliberate map. The vendors have made this possible by splitting their crawlers by purpose, and most sites have not caught up.
OpenAI documents four agents. OAI-SearchBot builds the index that surfaces sites in ChatGPT’s search features; OpenAI states directly that sites opted out of it will not appear in ChatGPT search answers, though they may still appear as navigational links. GPTBot collects content for training foundation models, and OpenAI states that disallowing it indicates content should not be used in training. OAI-AdsBot checks the safety and relevance of pages submitted as ad landing pages, and OpenAI states that the data it collects is not used for model training. ChatGPT-User handles user-initiated fetches, is not used for automatic crawling or search eligibility, and may not follow robots.txt.
Anthropic split its crawlers along the same lines in a February 2026 documentation update, retiring the older Claude-Web and anthropic-ai agents. ClaudeBot gathers training data. Claude-SearchBot indexes content for Claude’s search results, and Anthropic warns that blocking it “prevents our system from indexing your content for search optimization, which may reduce your site’s visibility and accuracy in user search results.” Claude-User fetches pages when a Claude user asks it to browse.
Google’s arrangement predates both and works differently, which is the detail that trips people up. Googlebot is the single crawler behind Search, AI Overviews and AI Mode. There is no separate AI-features crawler to allow or block. Google-Extended is not a crawler at all — it is a token in robots.txt that controls whether already-crawled content may be used for training Gemini models and for grounding in Google’s other AI products, without affecting Search inclusion. Google-CloudVertexBot handles Vertex AI grounding for customers who request site crawls. GoogleOther covers research and product use.
Microsoft continues to use Bingbot for both Bing and Copilot, with nocache and noarchive directives available as partial controls over how content is used in generative answers. Perplexity operates PerplexityBot for indexing and a separate user-triggered fetcher, with its crawler documentation and IP ranges published; the company has also been the subject of disputes about fetches that bypassed publisher blocks, which is a reminder that documented behaviour and observed behaviour are not always identical.
AI crawler control matrix by purpose
| Purpose | OpenAI | Anthropic | Microsoft | Perplexity | |
|---|---|---|---|---|---|
| Classic search index | Googlebot | — | — | Bingbot | PerplexityBot |
| AI answer citations | Googlebot (same crawler) | OAI-SearchBot | Claude-SearchBot | Bingbot | PerplexityBot |
| Model training | Google-Extended token | GPTBot | ClaudeBot | Bingbot with nocache/noarchive | Not separately documented |
| Live user-initiated fetch | Google-CloudVertexBot (grounding) | ChatGPT-User | Claude-User | Copilot fetches | Perplexity-User |
| Advertising checks | — | OAI-AdsBot | — | — | — |
The matrix is the artefact to hand your infrastructure team, because it converts an ideological argument about AI into four separate technical decisions with different business consequences.
The decision framework that follows is straightforward once the agents are separated. If you want AI-answer citations, you must allow the search crawlers — OAI-SearchBot, Claude-SearchBot, Bingbot, PerplexityBot — and you must keep Googlebot fully allowed, because there is no way to appear in AI Overviews or AI Mode while excluding Googlebot. If you object to training use, block the training crawlers — GPTBot, ClaudeBot — and set Google-Extended. These are independent choices, and the common publisher position of “cite me, don’t train on me” is now technically expressible for the major vendors.
Three caveats keep this from being a solved problem. Blocking training crawlers does nothing about content already ingested, and nothing about third-party datasets like Common Crawl unless you block CCBot too. User-initiated fetchers are documented as not necessarily honouring robots.txt, so a robots.txt-only posture leaves that channel open by design. And smaller or less scrupulous crawlers ignore robots.txt entirely, which is the argument for enforcing preferences at the network edge rather than in a text file.
Review cadence matters more than the initial configuration. Vendors add, rename and retire agents several times a year — Anthropic’s February 2026 split is one example, OpenAI’s addition of an ads crawler another. A robots.txt written in 2024 is now describing a world that no longer exists. Put the crawler map on a quarterly review with a named owner, and log the diff each time.
Rendering, JavaScript and the content AI systems never see
Google’s optimization guide tells site owners to follow JavaScript SEO best practices and states that AI features rely on publicly accessible, crawlable content. That sentence is doing a lot of quiet work, because the rendering behaviour of AI retrieval systems is more varied and less forgiving than Googlebot’s.
Googlebot renders JavaScript. It has for years, with a queue and a delay, and Google’s AI features draw on the rendered index. That means a client-side React application can appear in AI Overviews, provided the rendering succeeds and the delay is tolerable. The other engines do not offer the same guarantee, and several do not render at all.
The practical test is unglamorous: fetch your page with JavaScript disabled and read what comes back. If the main content is absent from the raw HTML, you are betting your visibility in every non-Google AI surface on whichever engines happen to run a browser. Some do for some request types. Many fetch raw HTML, strip tags, and hand the result to the model. If your raw HTML contains a <div id=”root”></div> and a bundle reference, the model receives nothing.
Four rendering patterns produce recurring problems.
Content behind interaction. Accordions, tabs, “read more” toggles and modal dialogs are fine if the content exists in the DOM on load and is hidden with CSS. They are invisible if the content is fetched on click. The distinction is not visible to a human reviewer and is trivially detectable in the raw HTML, which is why it belongs on a checklist rather than in a designer’s judgment.
Infinite scroll and virtualised lists. Category pages, documentation indexes and resource libraries that load items as the user scrolls present a handful of items to a non-rendering fetcher. Paginated equivalents with real links solve this and cost almost nothing.
Client-side routing without server responses. Applications that handle routes in the browser and return the same shell HTML for every URL are, from a retrieval perspective, a single page. Each answerable question needs its own server-rendered URL with its own content.
Hydration mismatch and partial rendering. Server-side rendering that emits a skeleton and fills it after hydration behaves like client-side rendering for any fetcher that does not wait. Time-to-content is the variable, not the framework name.
The fix hierarchy is well established and unchanged by AI: server-side rendering or static generation for content pages, with client-side interactivity layered on top. Prerendering for crawlers is acceptable when the prerendered output matches what users see; when it diverges, it becomes cloaking, which carries real ranking risk and is trivially detectable by comparing user-agent-conditional responses.
Two AI-specific additions belong here.
The first is speed as a retrieval gate rather than a ranking factor. Live user-initiated fetchers operate inside a conversation turn. If a page needs six seconds of round trips before the article body exists, the assistant either abandons it or uses a partial version. Time-to-first-meaningful-content for a cold, uncached, non-rendering fetch is now worth measuring as its own metric, separate from Core Web Vitals, which measure a warm browser experience.
The second is markup noise. Google states that perfectly semantic HTML is not required and that its systems handle imperfect structure. That is true for Google. It is less true for pipelines that convert HTML to text or Markdown before chunking, where a page with fourteen nested wrapper divs, inline SVG icons in every heading, and navigation duplicated three times for responsive breakpoints produces a text extraction where the article body is a minority of the tokens. Clean, shallow markup with one navigation block, one main content region, and headings that are actually heading elements produces a cleaner extraction, and cleaner extractions are easier to cite accurately.
A related practice worth adopting: use <main>, <article>, and <nav> landmarks, and mark repeated boilerplate with data-nosnippet where you genuinely do not want it quoted. Google honours data-nosnippet for snippet generation including in AI features, which makes it a precision tool for keeping cookie notices, disclaimers and promotional interstitials out of quoted material. Used carelessly on real content, it removes you from AI answers entirely — which is exactly what some publishers want, and exactly what most marketing teams do not.
Finally, test what the models actually receive rather than what your CMS thinks it published. Convert your own pages to plain text with a standard extraction library, read the output, and ask whether a reader with only that text could answer the questions your page is meant to answer. That five-minute exercise finds more real GEO problems than most paid audits.
Indexability and snippet permissions as a single decision
Google’s documentation is explicit that content must be indexable with snippets enabled to appear in AI Overviews and AI Mode. This turns a set of directives that most teams treat as unrelated legacy settings into a single coherent permission layer, and misconfiguration here produces total invisibility with no error message anywhere.
The directives that matter, and what each does in an AI context:
noindex removes the page from the index entirely. No index entry, no retrieval, no citation. Obvious, and still found on pages nobody intended — staging templates promoted to production, filtered category variants, gated resource pages, and entire subdirectories where a header rule applied more broadly than intended.
nosnippet prevents any text snippet from being shown. Because AI features generate summaries from indexable snippet-eligible content, nosnippet withdraws the page from that generation. The page can still rank as a link. It will not be quoted or summarised. Publishers who deployed nosnippet during the 2024–2025 wave of AI anxiety and never revisited it are frequently the same publishers now asking why they are absent from AI answers.
max-snippet:[n] caps snippet length in characters. Setting a low value — a tactic once used to force Google to show titles rather than descriptions — constrains what AI features can use. A max-snippet:0 is functionally nosnippet.
data-nosnippet excludes specific page regions at the element level. This is the precision instrument: keep your paywall notice, author bio boilerplate, related-articles module and legal disclaimer out of quotable material while leaving the article body fully available.
Google-Extended in robots.txt governs training and grounding in Google’s other AI products, not Search inclusion. Setting it does not remove you from AI Overviews or AI Mode. Teams routinely conflate the two, either setting Google-Extended and expecting to disappear from AI Overviews, or refusing to set it out of fear that they will.
nocache and noarchive for Bing restrict caching and archiving and have been described by Microsoft as controls affecting generative use in Copilot. Their effect is partial and less precisely documented than Google’s equivalents.
The audit is a crawl of your own site checking, per URL: HTTP status, robots meta directives in both the HTML head and the X-Robots-Tag response header, canonical target, and presence of data-nosnippet on content regions. The response-header path is the one that hides: an X-Robots-Tag: noindex applied at the CDN or load-balancer level does not appear in your CMS, your page source as viewed in a browser, or most SEO plugins’ reporting. It appears only in the response headers, and it has quietly de-indexed entire sections of large sites.
Canonicalisation deserves attention because retrieval operates on passages but attribution operates on URLs. If the same content exists at four URLs — with and without a trailing slash, with and without tracking parameters, on a print variant, on an AMP-era legacy path — the retrieval system may pull from one and cite another, or may split signals across all four. Consolidate first, then measure. Otherwise your citation tracking will undercount you by attributing your own citations to URLs you are not monitoring.
Two further items belong on this checklist.
Paywalls and gating. Content behind a hard paywall is not publicly accessible and will not be retrieved unless you serve it to crawlers, which without proper structured-data signalling is cloaking. The supported path is isAccessibleForFree structured data with the paywalled section marked, which lets crawlers understand the arrangement. The strategic question underneath it — how much of your paid content to expose so that AI systems know it exists — is a business decision, not a technical one, and it is one of the genuinely hard calls of 2026.
Consent management. In the EU, a consent management platform that blocks content rendering until a choice is made will serve the banner and nothing else to any fetcher that cannot consent. Because AI crawlers overwhelmingly fetch from a limited set of data-centre regions, this failure is often invisible to teams testing from their own browsers. Serve content, not consent walls, to non-human agents, and handle tracking consent through script gating rather than content gating. This is also the more defensible reading of the underlying data-protection obligations, since the content itself is not personal data processing.
Page architecture that survives chunking and passage retrieval
Retrieval systems do not read pages. They read passages. Understanding that one fact changes how you build a page more than any other piece of GEO advice, and it also explains why one popular piece of GEO advice is wrong.
The wrong advice is to break content into tiny pieces. Google addresses this directly, stating there is no requirement to break content into small chunks for AI and that its systems understand the nuance of multiple topics on a page. Artificially fragmenting content into 150-word micro-pages produces thin pages that struggle to rank, which removes them from the candidate pool that AI features draw on. The pattern is self-defeating.
The right version of the same instinct is different: write in self-contained units inside a well-organised long page. A passage is a self-contained unit when a reader who sees only that passage can understand the claim, knows what it refers to, and can identify who is making it. Most web writing fails this test because it relies on the preceding paragraph for its subject, its qualifier, or its numbers.
Consider the difference in practice. “This makes it the fastest option in the category” is not self-contained — “this” and “the category” are both unresolved. “The 2026 model completes the cycle in 42 minutes, faster than any other machine in the compact class” is self-contained, quotable, and attributable. Nothing about that rewrite is AI-specific; it is what good technical writing has always looked like. The difference is that in a passage-retrieval world, the failure mode is not “slightly harder to read” but “unusable as evidence”.
Google’s guidance supports the underlying structural point: organise content clearly with paragraphs, sections and descriptive headings. Descriptive headings do real work in a chunked pipeline because many extraction approaches carry the nearest heading forward as context for the passage beneath it. A heading that says “Pricing” gives the passage almost no context. A heading that says “Pricing for teams under 50 seats” makes every sentence beneath it interpretable in isolation.
Six structural practices follow, and each is testable.
One question per section. A section that answers a single identifiable question produces passages that answer that question. A section that wanders across three topics produces passages that answer none of them cleanly.
Answer position early in the section. Put the direct answer in the first or second sentence under the heading, then develop it. The convention journalists call the inverted pyramid maps almost exactly onto how passages get selected and quoted. Burying the answer under 200 words of context means the retrieved passage contains the context and not the answer.
Explicit subjects instead of pronouns at paragraph boundaries. Restate the entity at the start of each paragraph rather than carrying “it” or “they” across a boundary. This reads slightly more repetitively to a human and dramatically better to a chunker.
Numbers, units and dates inline. “Roughly a third” is unusable; “34% as of March 2026” is citable. Attach the date to the claim rather than relying on the page’s publication date, because retrieval frequently separates the passage from the page metadata.
Tables for genuinely tabular comparisons, prose for reasoning. Tables extract cleanly and are heavily used by retrieval systems for comparison queries. Reasoning stuffed into a table loses its logic. Comparisons stuffed into prose lose their structure.
Real heading elements in real hierarchy. Styled <div> elements that look like headings are not headings. Skipped levels confuse structural parsers. This is basic, and it is broken on a majority of the marketing sites I have audited.
There is a length question underneath all of this, and the honest answer is that it depends on the query type rather than on a word count. Comprehensive pages covering a topic and its adjacent sub-questions perform well on fan-out because a single page can match several sub-queries. Narrow pages perform well when the query is narrow and the competition is comprehensive but shallow. The failure mode to avoid is the middle: a 2,000-word page that covers six topics at 300 words each, deep enough to look substantial and shallow enough to lose every sub-query to a specialist page.
One more architectural point, aimed at documentation and support content specifically. Anchor links to individual sections are worth implementing properly, with stable IDs and a visible table of contents. Some AI surfaces deep-link to page fragments, and a page with meaningful fragment targets can be cited at the section level. This is low-cost, and it also improves the human experience of long pages, which is the test any GEO tactic should pass before you spend money on it.
Extractable claims and the answer-first paragraph
The strongest moderately-supported finding in the GEO literature is that extractable evidence makes a passage easier for a model to use. The 2026 survey places statistics, definitions and quotations in its moderate-evidence tier — not proven to lift discoverability, but plausibly influencing whether an already-retrieved passage gets used and cited. That is a narrow claim with wide practical consequences, because making claims extractable costs almost nothing and improves the content for human readers regardless.
An extractable claim has five properties. It is specific rather than general. It is attributable — the source of the number or the assertion is named in or adjacent to the sentence. It is self-contained, resolvable without the surrounding paragraph. It is falsifiable, stated precisely enough that it could be checked. And it is dated, so a model deciding between conflicting claims can prefer the current one.
Run any commercial page through those five tests and the results are usually bleak. “Our platform helps teams work faster” fails all five. “Teams using the shared inbox closed support tickets 18% faster in our 2026 customer study of 240 accounts” passes all five, and it took the same number of words.
Definitions are the highest-yield instance of this, and the most neglected. Generative answers frequently need a one-sentence definition of the entity, product category, or process at hand. A page that contains a clean, standalone definition of its subject early in the body is a natural candidate for that slot. The pattern is simple: the term, a copula, the category, the distinguishing property. “Generative engine optimization is the practice of increasing a source’s presence in AI-generated answers, distinguished from classic SEO by its focus on citation and inclusion in synthesised responses rather than link position.” That sentence can be lifted whole. Most product pages do not contain one anywhere.
Quotations work for a related reason. A quoted expert statement carries an attributable speaker, which gives a model something safe to reproduce. Original quotes from named people with stated credentials are worth more than paraphrased consensus, because the paraphrase is available everywhere and the quote is available only from you. This is one of the few places where GEO advantage is genuinely defensible rather than competitive-erosion-prone, since a competitor can copy your formatting but cannot copy your interview.
There are three traps in this area, and all three are actively being sold as best practice.
The first is the FAQ block deployed as a citation-farming device. Question-and-answer formatting does make answers extractable, and it does align with the question-formatted queries that trigger AI answers at high rates. It stops working when the questions are invented to host keywords rather than because readers ask them, and when the answers restate the page rather than adding anything. Google’s spam policies cover scaled content abuse, and a page carrying twenty synthetic questions with two-sentence answers is squarely in that territory. Write the questions your support team actually receives, answer them properly, and cap the block at the number of real questions you have.
The second is statistic manufacturing. The survey’s white-hat test set includes evidentiary authenticity: statistics must be verifiable. The GEO industry’s own literature is a case study in the failure, with widely repeated figures like “36% more likely to appear in AI summaries” that trace to no traceable source. Citing an unverifiable number does not just risk correction; it risks being the source a model cites for a false claim, which is a reputational exposure with no upside.
The third is over-formatting. Bullet lists, bolded fragments and short paragraphs everywhere produce a page that looks machine-optimised and reads badly. The survey found generic formatting recipes generalize poorly across engines. Formatting should follow the content’s structure, not a template applied uniformly, and a page where everything is emphasised has emphasised nothing.
A closing practice that ties this section to measurement: keep a list of the twenty claims you most want attributed to your brand, written in extractable form, and check quarterly whether AI assistants reproduce them and whether they attribute them to you. That list is simultaneously a content brief, a PR brief, and a fidelity monitoring baseline. It is the most useful single artefact I know of for running GEO as an ongoing programme rather than a project.
Structured data reconsidered as description, not persuasion
Structured data is the most contested item on every GEO checklist, and the disagreement persists because both camps are arguing about different things.
The evidence, stated fairly. Google says structured data is not required for generative AI search and that no special schema.org markup exists for it, while recommending it as part of overall SEO for rich-results eligibility. Microsoft has been more positive, with representatives stating that schema markup helps Bing’s language models understand content for Copilot — the only first-party confirmation from a major AI platform that markup feeds a generative pipeline. OpenAI, Anthropic and Perplexity have made no public statements either way. On the independent side, a Search Atlas study found no correlation between schema coverage and citation rates across OpenAI, Gemini and Perplexity. The 2026 academic survey places structured data in its moderate tier with context-dependent benefits. No peer-reviewed study establishes a causal link between markup and citation.
The industry’s numbers, meanwhile, do not survive scrutiny. Claims like “36% more likely to appear in AI summaries” and “60% visibility loss without schema” circulate widely and trace back to nothing verifiable. A checklist that justifies structured data with those figures is building on sand, and clients increasingly know it.
So why does structured data still belong on a practical checklist? Because the argument for it does not depend on a citation-rate lift.
The first reason is eligibility for surfaces that are unambiguously fed by markup. Product, Review, Recipe, Event, JobPosting, Video and Organization markup drive rich results and feature eligibility in Google Search, and those surfaces continue to exist inside and alongside AI experiences. Merchant listings, product prices in shopping surfaces, and event details in Google’s answer panels all draw on structured feeds and markup. Losing those is a concrete, measurable cost, whether or not markup influences a generative citation.
The second reason is entity resolution, which matters enormously in an AI context. Organization markup with sameAs links to your Wikidata entry, your LinkedIn page, your Crunchbase profile and your registry listings does the work of telling machines that these identifiers refer to the same company. Person markup on author pages with sameAs links to a professional profile does the same for people. Knowledge graphs are built from exactly these correspondences, and knowledge graphs feed grounding. This is a description problem, and structured data is a description format.
The third reason is disambiguation of things models otherwise guess at. Prices with currency and validity dates. Availability. Service areas. Opening hours. Qualification requirements. Membership of professional bodies. Whether content is free or paywalled. Whether a review is first-party or aggregated. When a model has to state a fact about you and the only available signal is prose, it infers. When there is markup, it has less room to infer.
What structured data cannot do is argue. This is the trap, and it is the reason “schema for GEO” packages disappoint. Markup that asserts your product is the best, stuffs keywords into description fields, or fabricates AggregateRating values on pages with no reviews is a spam vector, not a visibility tactic. Google’s structured data guidelines treat markup that does not represent visible page content as a violation, and enforcement includes manual actions. Structured data describes what is on the page. It does not add claims to it.
A defensible structured-data checklist, ordered by expected return:
Organization on the home page, with legal name, alternate names, logo, contact points, address, founding date, identifiers such as VAT or company registration number where public, and sameAs links to every authoritative external profile you control. This is the single highest-value markup for entity clarity and it is missing or minimal on most sites.
WebSite with SearchAction if you have internal search, plus a stable canonical url.
Article or NewsArticle on editorial content, with author as a linked Person entity rather than a string, datePublished and dateModified that reflect reality, publisher linked to the Organization node, and isAccessibleForFree where relevant.
Person on author pages, with credentials, affiliations, and sameAs links. This is where E-E-A-T signalling becomes machine-readable rather than rhetorical.
Product with Offer on commercial pages, with price, currency, availability and priceValidUntil, matching what the page displays and what your feeds say.
BreadcrumbList for hierarchy, FAQPage only where the visible page genuinely contains a Q&A block, and HowTo only for genuine procedures.
Two implementation notes that determine whether any of this works. Use a connected graph rather than isolated blocks: give nodes @id values and reference them, so the Organization referenced as publisher on an article is the same node as the Organization on the home page. And validate continuously, not once, because CMS template changes silently break markup and nobody notices until a rich result disappears. Google’s Rich Results Test and the Schema Markup Validator cover different things — the former checks Google feature eligibility, the latter checks schema.org validity — and you want both in a scheduled check.
The honest positioning: structured data is infrastructure with low cost, no downside when implemented truthfully, clear benefits in surfaces that definitely use it, and unproven benefits in generative citation. That is enough to justify doing it and not enough to justify selling it as the core of a GEO programme.
Entity clarity, brand disambiguation and the identity layer
If one area of GEO is undersold relative to its actual importance, it is entity clarity. Generative systems answer questions about things. If a system cannot reliably determine what your organisation is, what it sells, where it operates and how it differs from a similarly named entity, then every downstream tactic operates on a foundation that keeps shifting.
The problem is more common than teams expect. Companies share names across jurisdictions and industries. Products share names with unrelated software. Rebrands leave two identities coexisting in the corpus for years. Subsidiaries, trading names and legal entities differ. Local franchises carry the parent name with different service areas. Every one of those situations produces answers that blend two entities, and the blend is usually invisible until someone asks the assistant a question and reads the response carefully.
Diagnosing entity confusion takes an hour. Ask five assistants a set of plain questions: what is [brand], what does [brand] do, where is [brand] based, who founded [brand], who are [brand]’s competitors, is [brand] the same as [similar name]. Record the answers verbatim across several runs, because engine variance means one run tells you nothing. Then classify the errors: wrong industry, wrong geography, wrong founder, merged with another entity, outdated after a rebrand, confused with a competitor, hallucinated product that does not exist.
The remedies are structural rather than editorial, and they work on the sources these systems actually ground on.
A canonical about page that reads like a reference entry, not a brand manifesto. Legal name, trading names, founding year, founders with full names, headquarters with a real address, markets served, what the company sells stated in category terms a stranger would use, size indicators, and any parent or subsidiary relationships. Most about pages contain none of this and a great deal of positioning language. The reference version is what gets quoted.
Consistent naming across every surface you control. One canonical spelling, one canonical legal form, one canonical product name per product. Sites that alternate between three variants of their own product name train the corpus to treat them as three things.
Wikidata, where the entity meets notability criteria. Wikidata is openly licensed, machine-readable, and used as a grounding source across the industry. An accurate Wikidata item with correct industry classification, location, founding date, official website and external identifiers does more for entity clarity than a quarter of content production. It is also community-governed, so editing it promotionally will get reverted and can damage your standing; the correct approach is factual, sourced and modest.
Authoritative third-party profiles kept current. LinkedIn company page, Crunchbase, industry association directories, national business registries, app stores, review platforms, and — for regulated sectors — the regulator’s own public register. These are cited disproportionately by AI systems because they are structured, stable and trusted. A stale Crunchbase entry describing your business as it existed in 2021 will keep resurfacing in answers about your business as it exists now.
sameAs linkage from your Organization markup to every one of those profiles, so the correspondence is explicit rather than inferred.
Explicit disambiguation content where a real collision exists. If another company shares your name, a short, factual page stating who you are and who you are not is legitimate and useful, and it gives systems something to retrieve when the ambiguity surfaces. This is not a competitor-attack page; it is a clarification.
Personal entities deserve the same treatment when your strategy depends on named expertise. Author pages with credentials, institutional affiliations, publication history and external profile links let systems connect a byline to a real person with a real record. When they cannot, the byline carries no weight in whatever internal assessment of source quality the system performs. Building author entities is slow, cumulative and one of the few GEO investments that appreciates rather than erodes.
One caution against a tactic being actively sold. “Entity building” packages that generate dozens of low-quality profile listings, syndicated press releases and directory entries are the AI-era version of link farms. Google’s guidance explicitly says that seeking inauthentic mentions across the web is ineffective and contradicts its spam policies. The mechanism that makes entity clarity work is consistency across sources that are already trusted, not volume across sources that are not.
Freshness signals, dates and the maintenance discipline
Recency shows up in the GEO literature as a context-dependent benefit rather than a universal ranking factor, which matches what practitioners observe: for some query types recency dominates, for others it is nearly irrelevant, and treating it as a blanket priority wastes effort while treating it as unimportant produces answers built on your outdated pages.
The query types where recency matters are identifiable in advance. Anything with a year in it. Anything where the correct answer changes — pricing, availability, regulation, compatibility, version numbers, best-of comparisons, tax and compliance thresholds. Anything about an ongoing situation. Where the answer is stable — definitions, mechanisms, history, physics — a well-written page from 2019 competes fine, and republishing it with a new date signals nothing except that you changed a date.
That last practice deserves direct criticism because it is widespread. Bulk-updating dateModified across a site without changing content is a manipulation attempt with a poor risk-to-reward ratio. It corrupts your own ability to tell which pages are genuinely current, it produces a mismatch between the stated date and the content that any evaluation of quality will eventually catch, and it does not change the underlying text that retrieval operates on. If nothing changed, the date should not change.
The productive version of freshness work is a maintenance system, and it has five components.
An inventory with decay rates. Every content asset gets a review interval based on how fast its subject changes: pricing pages monthly, regulatory guides quarterly, integration documentation on release cadence, foundational explainers annually. This turns “we should update our content” into a scheduled workload with owners.
Dates on claims, not just on pages. Write “as of March 2026” next to the number inside the sentence. Retrieval separates passages from page metadata, so a passage carrying its own date survives the separation. This single habit does more for how your content behaves in generative answers than most technical work, and it also makes your own audits trivial.
Both datePublished and dateModified, honestly. Visible on the page and in structured data, matching each other, with a short note on what changed for substantive revisions. Publications that show “updated 12 March 2026 — revised pricing table and added the new compliance deadline” are giving both readers and machines something to work with.
Deprecation rather than deletion. When information becomes wrong, the worst outcome is leaving it live; the second worst is deleting the URL and losing the accumulated signals. The better pattern is updating in place with an explicit note that the previous guidance changed and when. Models trained or grounded on your old content will reproduce it; a live page that explicitly corrects the old claim gives the retrieval layer something to override it with.
A monitoring loop for stale claims resurfacing. Ask assistants the questions where you have changed your answer, and check whether they are still giving the old one. This is the only way to discover that a 2023 price point is still being quoted about you.
There is a harder version of this problem that deserves its own note: content already absorbed into model weights cannot be updated. Training-time knowledge is fixed until the next training run. If a model learned an incorrect fact about your business in 2024, no amount of publishing corrects the weights. What publishing does is give retrieval-augmented systems a current source that contradicts the stale parameter, which is why grounded answers correct faster than ungrounded ones. For businesses with a factual error embedded in the corpus, the practical response is to make the corrected version abundant, well-sourced and present on the third-party profiles that grounding systems reach for — and to accept that ungrounded answers will lag.
The operational point that ties this together: freshness in a GEO context is a data-integrity discipline, not a publishing tactic. The goal is that every factual claim about your business is correct, dated and consistent across every source a machine might read. That is unglamorous, it does not produce a monthly deliverable that looks impressive, and it is where a large share of the real damage sits.
Off-site presence and the third-party substrate AI reads
The most commonly cited sources in AI answers are not brand websites. Peec AI’s March 2026 analysis of 30 million sources across AI platforms found Reddit ranked first, YouTube second, with LinkedIn, Wikipedia and Forbes in the top five, and Yelp and G2 appearing heavily in recommendation queries. Platform preferences diverged: ChatGPT favoured Wikipedia, Reddit and editorial sites like Forbes; Google leaned toward Facebook and Yelp; Perplexity emphasised Reddit, LinkedIn and G2 for B2B queries.
Read that finding for what it implies rather than as a curiosity. For a large share of commercial queries, the sources shaping the answer about your company are pages you do not own. A buyer asking “best project management tool for construction firms” gets an answer assembled largely from review platforms, community threads, listicles and comparison sites. Your own product page may be cited as a supporting link. It is rarely the substrate.
That reframes the checklist. On-site work determines whether you can be cited when the system reaches for a primary source. Off-site work determines what the system believes about you when it reaches for anything else — which is most of the time.
The practical off-site layer breaks into five workstreams, in rough order of return.
Review and comparison platforms relevant to your category. G2 and Capterra for software, Trustpilot for consumer services, Yelp and Google for local, sector-specific registers for regulated services. The task is not “get more reviews” as a vanity metric but making sure the structured fields are complete and accurate: category classification, pricing model, feature list, integrations, target company size, geography. These platforms are heavily cited because their data is structured and comparable, and an incomplete profile produces an incomplete or wrong description of you in answers.
Wikipedia and Wikidata, where notability genuinely applies. Wikipedia’s role in AI grounding is out of proportion to its share of the web, and Wikidata’s structured claims are used directly. Both are community-governed with strict conflict-of-interest norms; the correct engagement is factual correction with reliable sources, disclosed. Attempting promotional editing is both against the rules and counterproductive.
Independent editorial coverage in outlets your category’s audience and the models both trust. This is public relations, and the AI era has raised its return rather than lowered it. A substantive article in a respected trade publication, with your named expert quoted and your data cited, becomes a retrievable third-party source that says what you want said in a voice that carries more weight than your own site. The measurable version of this is tracking whether your claims appear in AI answers attributed to third-party coverage rather than to you — which counts as a win, not a loss.
Structured industry data sources. Standards bodies, professional associations, trade registries, patent and trademark databases, funding databases, regulatory registers. These are stable, authoritative and machine-readable, and for regulated industries they are frequently the decisive source. A law firm’s entry in a bar association directory or a clinic’s entry in a health regulator’s register carries more grounding weight than anything on the firm’s own site.
Video, specifically YouTube. Its second-place position in the citation index is not an accident: transcripts are available, the platform is well-indexed, and Google has obvious reasons to surface it. For categories where demonstration matters — software walkthroughs, physical products, procedures — a well-titled, well-described video with an accurate transcript is a retrievable asset with a different competitive field than text.
What does not work, and is being sold hard: bulk directory submission, syndicated press-release networks, paid “brand mention” packages, and AI-generated guest posts placed at volume. Google’s guidance states that seeking inauthentic mentions across the web is ineffective and contradicts its spam policies. The mechanism these packages imagine — that models count mentions — is not how grounding works, and the sources they place on are precisely the low-authority pages that grounding systems discount.
The uncomfortable strategic conclusion: if your category’s answers are built from Reddit threads and G2 profiles, then your GEO budget belongs substantially outside your own website. Most GEO proposals are on-site content plans because on-site content is what agencies sell. The distribution of cited sources says the money should be split differently.
Community platforms and the Reddit problem
Reddit’s position at the top of the AI citation index creates a strategic problem that no honest GEO checklist can resolve cleanly, and pretending otherwise is how brands get themselves into trouble.
The mechanics are clear enough. Reddit threads are long, conversational, contain multiple viewpoints, use natural question phrasing, and carry visible signals of agreement through voting. That structure maps unusually well onto the questions people ask assistants — “is X worth it”, “what do people actually use for Y”, “did anyone else have this problem”. Reddit also has commercial data arrangements with major AI companies, which puts its content in a different position than an ordinary forum. And by mid-2026 the relationship had become contentious enough that reports of Reddit and large publishers weighing restrictions on Google over AI-driven referral declines were circulating widely.
For a brand, three routes exist, and only two are defensible.
The indefensible route is astroturfing: creating accounts to recommend your product, seeding threads, paying users for positive comments. Beyond the platform rules, the advertising-standards exposure in most jurisdictions is real — undisclosed paid endorsement is regulated as deceptive practice in the EU under the Unfair Commercial Practices Directive and by the FTC in the United States. The detection risk is also rising rather than falling, and a documented astroturfing campaign is a bigger brand event than any visibility gain it produced.
The first defensible route is genuine participation with disclosure. Employees who are experts in their field, identified as working for the company, answering questions in relevant communities without pitching. Most subreddits tolerate and some welcome this when the disclosure is clear and the contribution is substantive. It is slow, it does not scale, it cannot be delegated to an agency, and it produces the exact artefact that gets cited: an informed answer in a community thread from a named person with disclosed affiliation.
The second defensible route is making yourself easy to recommend accurately. You do not control what a community says about you, but you strongly influence it. Clear pricing pages that do not require a sales call. Honest limitation documentation. Public changelogs. Responsive support with visible resolution. Migration guides that acknowledge competitors by name. Communities recommend products they can describe confidently, and they can only describe confidently what the vendor has documented plainly. A recurring pattern in threads about opaque vendors is “nobody knows what it costs” — which becomes the answer an assistant gives about your pricing.
A related tactic that is legitimate and underused: monitor and correct. Community threads containing factual errors about your product — a limitation that was removed two versions ago, a price that changed, a compatibility claim that was never true — are worth a disclosed, polite, sourced correction from an identified employee. This is not reputation manipulation; it is accuracy maintenance on a source that grounding systems read. It is also usually welcomed, because the community would rather be right.
The same logic applies across the other community-shaped platforms. Stack Overflow and its network for technical answers, where a well-received answer from a company engineer is a durable asset. Q&A sites in specific verticals. Professional groups on LinkedIn. Discord and Slack communities are largely invisible to crawlers, which makes them poor GEO targets and fine community-building targets — a distinction worth keeping straight so nobody reports Discord engagement as AI visibility work.
Two structural cautions.
Community sentiment is a lagging, sticky variable. A bad period of service quality produces threads that keep getting cited for years. There is no content tactic that overrides a body of negative first-hand accounts, and attempting one produces the worst version of the problem, where the assistant contrasts your marketing claims with user reports. The remedy for negative community sentiment is fixing the product, and any GEO consultant who offers a different remedy is selling something else.
Citation of a thread is not endorsement of your position within it. When an assistant cites a Reddit thread about your category, it may be synthesising a consensus that goes against you. Tracking “we appeared in the answer” as a success metric here is misleading. Track sentiment and accuracy of the mention, not the mention itself — which is why the fidelity component of the visibility vector matters as much as the citation component for anyone with active community discussion.
Query fan-out and deliberate topical coverage
Query fan-out is the mechanism that makes topical coverage a retrieval strategy rather than a content-marketing aspiration. Google describes fan-out as issuing multiple related searches to assemble a broader set of supporting links. The practical implication is that a single user prompt generates a set of sub-queries, and your presence in the answer depends on matching some of them.
That changes how to plan content. Classic keyword planning starts from search volume and picks head terms. Fan-out planning starts from a user situation and enumerates the questions that situation generates, most of which have negligible individual search volume and all of which may appear as sub-queries.
The method is straightforward and does not require tooling. Take a realistic prompt a buyer would type into an assistant — not a keyword, a full sentence with context. “I run a 30-person accounting practice in Slovakia and need to replace our document management system before the new e-invoicing rules take effect.” Then write out every question a competent adviser would need answered to respond: what the rules require and when, what document retention periods apply, which systems support the required formats, what migration from the incumbent involves, what it costs at that headcount, what the implementation timeline looks like, what happens to historical records, whether local hosting is required, who provides support in the local language.
That list is the fan-out surface for that prompt. Nine sub-questions, most with low individual search volume, several of which no competitor has answered properly. A site that answers all nine — properly, in self-contained passages, with dates and specifics — enters the candidate pool through multiple doors for that prompt and for dozens of adjacent ones. A site that has a page titled “document management software” enters through none.
Three implementation patterns follow.
Cluster architecture with real depth at the leaves. A hub page establishing the topic and defining terms, with substantial pages beneath it answering specific sub-questions in full. The failure mode is a hub with fifteen thin children; the working version has fewer children with more substance each. Interlinking should be descriptive rather than decorative — link text that states what the target answers.
Coverage auditing against enumerated questions rather than keyword lists. Build the question inventory for your top ten buyer situations, then map existing content against it and mark each question as answered well, answered thinly, answered incorrectly, or absent. The gaps are your content plan, prioritised by commercial proximity rather than volume. This audit routinely finds that a company has forty pages targeting the same six head terms and nothing addressing the questions that actually precede a purchase.
Adjacent-domain coverage where the buyer’s question crosses a boundary. The accounting-practice prompt above involves regulatory content, software comparison, migration logistics and local-market specifics. Most vendors write only the software comparison. The regulatory and migration content is where the sub-queries are least contested and where being the source that explains the situation correctly builds the entity association you want.
Two cautions.
The first is against automating the enumeration into content production. Generating a page per sub-question at scale produces exactly the pattern Google’s spam policies describe as scaled content abuse, and it produces thin pages that fail to rank, which removes them from the candidate pool. Fan-out planning tells you what to write. It does not license writing it badly at volume.
The second is against confusing fan-out simulation tools with fan-out reality. Several tools now claim to reveal the sub-queries a given prompt generates. Google does not publish them. These tools are producing plausible reconstructions using their own language models, which is useful as an ideation aid and is not observation. Treat the output as a hypothesis about buyer questions, validate it against your sales calls and support tickets, and do not report it as data about Google’s behaviour.
Comparison pages and the recommendation query
Recommendation queries are where AI search most directly intercepts commercial demand, and they are the query class where the citation index findings bite hardest. When a buyer asks an assistant which product to choose, the answer draws heavily on review platforms, community threads and third-party comparison content — and on the small number of vendor pages that behave like reference material rather than sales collateral.
The dominant vendor mistake is writing comparison pages that no neutral system would use. A page titled “us vs competitor” that presents a feature matrix where the vendor has every checkmark and the competitor has almost none is not usable as evidence. It is also frequently inaccurate, because it reflects the competitor’s product as it existed when the page was written. Systems assembling a comparison need a source that states verifiable facts about both options; a page that states favourable facts about one is a low-quality candidate.
The version that works is uncomfortable for marketing teams because it requires conceding things.
State the honest fit boundaries. “Our platform suits teams above roughly 20 users with a dedicated administrator; below that, the setup overhead is hard to justify and a simpler tool is usually the better choice.” That sentence is a gift to a retrieval system: specific, self-contained, falsifiable, and clearly not marketing copy. It is also the sentence that stops unqualified prospects from entering your funnel, which most sales leaders eventually recognise as an improvement.
Compare on dimensions that are checkable. Pricing model and actual numbers. Deployment options. Data residency. Integration list. Compliance certifications with names and dates. Support hours and channels. Migration path and effort. Contract terms. These are facts that a system can verify against other sources, and consistency across sources is what builds source reliability.
Keep competitor claims current and sourced. Reference the competitor’s own documentation with a date. When their pricing changes, update yours. An out-of-date competitor claim is worse than no claim: it is a demonstrable inaccuracy that undermines the credibility of everything else on the page, and in some jurisdictions comparative advertising rules make it a legal exposure as well. In the EU, comparative advertising is permitted under the Misleading and Comparative Advertising Directive only where the comparison is objective, verifiable and not misleading — a standard that most vendor comparison pages would fail.
Publish the category overview, not just your position in it. A page that maps the category honestly — what the main approaches are, which types of buyer each suits, what the trade-offs are, and where you sit — is the kind of source a system reaches for when a user asks an open question. It also earns links and third-party citation in a way that a versus page never does.
Alternatives pages deserve a specific mention because they are the highest-intent version of this content and the most frequently botched. A page titled “alternatives to [competitor]” that lists five options with your own product first and four straw men beneath it is transparently self-serving. A page that lists the genuine alternatives, describes each accurately including the ones that beat you on specific dimensions, and explains which buyer profile each suits, is a reference document. The commercial argument for the honest version is that it converts better in a world where buyers arrive already informed by an assistant that read both pages.
Pricing transparency deserves its own line on the checklist because it has become a hard constraint rather than a preference. When a buyer asks an assistant what a product costs and the vendor does not publish a price, the assistant either says the price is not disclosed, quotes a third-party estimate that may be wrong, or quotes a competitor’s price by mistake. None of those outcomes serve the vendor. The “contact sales” pattern was already losing ground in B2B software; in an assistant-mediated purchase process it removes you from the comparison entirely for any buyer who does not want a call. Publishing a starting price, a pricing model, and the variables that move it is now a retrieval requirement, not just a conversion-rate question.
A final note on the review substrate underneath recommendation queries. Because review platforms are cited so heavily, the accuracy of your structured profile on them functions as pricing and feature documentation whether you treat it that way or not. A G2 profile with a stale feature list and no pricing information will be quoted as your feature list and your pricing opacity. Auditing third-party profiles for accuracy belongs on the same quarterly cycle as auditing your own pages, and in categories dominated by recommendation queries it is arguably the higher priority.
Product data, feeds and agentic commerce readiness
For anyone selling physical or digital products, the highest-leverage GEO work in 2026 is not content. It is data quality in the feeds and structured sources that shopping and agentic surfaces consume.
Google’s own optimization guidance says this directly: use Merchant Center feeds and Google Business Profiles to enable visibility in AI responses, and consider its business agent capabilities for conversational experiences. That is a different instruction from anything in the content-optimization literature, and it is the one with the clearest mechanism — these are structured pipelines with documented fields, not a probabilistic retrieval layer.
The layer above feeds is agentic commerce, and 2026 is the year it stopped being speculative. Google published the Universal Commerce Protocol as an open specification for agent-mediated commerce, with developer documentation and merchant tooling, and announced agentic shopping and booking capabilities across its surfaces at I/O 2026, including agents that handle local services bookings and, in selected categories, place calls to businesses on the user’s behalf. OpenAI has its own agentic commerce work with instant checkout mechanics and product feed ingestion. Perplexity has run merchant programmes. Multiple protocols are competing, none has won, and merchants are being asked to support several.
The practical checklist here is concrete in a way most GEO advice is not.
Feed completeness and accuracy first. GTIN or MPN where they exist, brand, precise product titles that a human would recognise, category taxonomy mapped correctly, availability that reflects real stock, price including currency and any validity window, shipping and return terms, and images that meet platform requirements. Incomplete feeds produce products that are absent from agent-visible inventory, and the absence is silent.
Attribute depth beyond the minimum. Size, material, colour, compatibility, dimensions, weight, power requirements, certifications, age suitability. Agents filter on attributes. A product with fifteen populated attributes is findable for fifteen kinds of query; a product with four is findable for four. This is the single most under-invested area in most retail catalogues, and it is measurable, unglamorous data work rather than strategy.
Consistency between feed, page markup and page content. A price of €249 in the feed, €279 in the Product markup and “from €199” in the page copy is a trust failure that produces suppression on some surfaces and wrong answers on others. Automated consistency checking between these three sources belongs in your build pipeline.
Returns, warranty and delivery terms as structured, retrievable facts. These are decisive in purchase decisions and are frequently buried in a PDF or a legal page. Stating them plainly with numbers — “30-day returns, buyer pays return shipping, refund within 14 days of receipt” — makes them quotable. Agents comparing two products on total cost of ownership need exactly this.
Availability signalling that is honest about lead times. “In stock” that means “ships in three weeks” is the kind of mismatch that agentic purchasing surfaces will punish faster than human shoppers did, because the agent’s job is to satisfy a stated constraint.
Protocol readiness assessed rather than assumed. Determine which agentic commerce protocols your platform supports natively, which your payment provider supports, and what a transaction initiated by an agent looks like in your order flow — including how you handle fraud checks, address verification and consent when the buyer is not present in a browser session. Most merchants have not run this exercise. The gap between “our products appear in an agent’s recommendation” and “an agent can complete a purchase” is where the revenue sits, and it is an engineering and operations problem, not a marketing one.
Two risks worth naming. Agentic purchasing compresses the moment of decision, which reduces the influence of on-site persuasion and increases the influence of structured comparison data — a shift in where marketing spend produces returns. And delegating checkout to third-party agents raises questions about customer data ownership, consent records, and who holds the relationship afterwards. Merchants who treat agentic commerce purely as a visibility channel and not as a channel-conflict question will find the terms set for them.
Local, service-area and physical-world visibility
Local queries are the most assistant-friendly query class in existence — a stated need, a stated location, an expected short list — and they run on a data layer that has almost nothing to do with content optimization.
Google’s guidance points local businesses at Google Business Profile as the enabling asset for AI visibility. That is the mechanism: the profile is a structured record with verified attributes, and it feeds both classic local results and the generative answers that draw on them. Everything else is secondary.
The checklist for local is therefore short and specific.
Profile completeness on every field that exists. Primary category chosen precisely rather than broadly, secondary categories where genuinely applicable, service list with descriptions, attributes including accessibility and payment methods, hours including holiday exceptions, service area defined accurately for mobile businesses, and photographs that are current. Empty fields become absent facts, and absent facts become answers that say the information is unavailable.
Name, address and phone consistency across every source that carries them. The website, the profile, the industry directories, the regulator’s register, the mapping providers, the review platforms. Inconsistency here produces entity fragmentation, and entity fragmentation in local search produces a business that appears twice with contradictory hours.
Reviews as a substrate, with response. Volume and recency both matter, and responses matter because they are text that describes what happened. A pattern of specific, non-templated responses to reviews produces a body of content about your service quality that grounding systems read. Review gating and incentivised reviews violate platform policies and, in the EU, consumer-protection rules on fake reviews that were tightened under the Omnibus Directive.
Location pages that contain real local information rather than the same 400 words with the city name swapped. Templated location pages at scale are a long-standing spam pattern and are now also useless for retrieval, because they contain no locally specific fact worth quoting. A location page worth having contains the actual address, the actual staff, the actual hours, the actual service area boundaries, parking and transit specifics, and any locally relevant regulatory or seasonal information. Ten real ones beat four hundred generated ones.
Mapping and voice ecosystem coverage beyond Google: Apple Maps via Business Connect, Bing Places, OpenStreetMap where accuracy matters, and the aggregators that feed navigation systems. Assistants embedded in cars and phones draw on these.
Booking and availability data where the category supports it. Google’s I/O 2026 announcements included agentic booking for local services and, in selected categories, agents calling businesses directly on a user’s behalf. A business with real-time availability exposed through a supported booking integration is actionable by an agent. One that requires a phone call during business hours is, at best, going to receive an automated call it may not recognise. Deciding how you want to handle agent-initiated contact — and whether your staff can tell the difference between an agent and a customer — is a 2026 operational question for service businesses.
Two observations about local that are worth stating plainly because they cut against the general GEO narrative.
First, local visibility in AI answers is more tractable and more measurable than almost any other GEO objective. The inputs are structured, the platforms are documented, the results appear in Business Profile insights and call tracking, and the competitive set is small enough that a well-maintained profile in a market of mediocre profiles produces a visible difference. For a plumbing company, a dental practice or a regional law firm, this is where the entire budget should go before anyone writes a blog post.
Second, the physical-world layer is where AI answer errors do the most immediate damage. Wrong hours send a customer to a closed door. A wrong address sends them to the wrong building. A wrong service area produces a booking you cannot fulfil. These are not visibility problems; they are operational failures mediated by an information system. Monitoring what assistants say about your practical details — hours, location, whether you take a given insurance, whether you serve a given postcode — is a customer-service obligation before it is a marketing task.
A measurement stack that survives audit
Two years of GEO reporting has been built on numbers that do not survive questioning, and the correction is arriving because the data situation improved. Google introduced generative AI performance reports in Search Console in June 2026, which for the first time gives site owners first-party visibility into performance in Google’s generative surfaces rather than inference from third-party sampling. Bing Webmaster Tools has added its own AI performance reporting. The era where “we can’t measure AI visibility” excused unfalsifiable claims is closing, and reporting practices need to change with it.
A measurement stack that holds up has four layers, and they answer different questions.
Layer one: first-party platform reporting. Search Console’s generative AI performance data for Google surfaces, Bing Webmaster Tools for Microsoft surfaces. These are the only sources with actual platform data rather than sampled reconstruction. Google’s own guidance warns explicitly against third-party tools claiming access to internal Google metrics, because no such access exists. Treat platform reporting as the baseline and anything else as supplementary.
Layer two: server-side retrieval evidence. Log analysis of verified AI crawler activity — which pages are fetched, how often, by which agents, with what response codes. This measures the input to retrieval rather than the output, and it is the layer nobody can fake.
Layer three: prompt-level sampling. A defined prompt set queried repeatedly across engines on a schedule, with results recorded. This measures citation, mention, sentiment and fidelity. It is legitimate and it is noisy, and the noise level determines the methodology, which the next section covers in detail.
Layer four: outcome data. Referral traffic where the referrer survives, assisted conversions, branded search volume, direct traffic patterns, and self-reported attribution from forms and sales conversations. This is the layer executives care about and the layer with the weakest measurement, which means it should carry the most explicit uncertainty.
AI visibility measurement layers and what each can honestly claim
| Layer | Data source | Answers | Reliability | Common misuse |
|---|---|---|---|---|
| Platform reporting | Search Console generative AI reports, Bing Webmaster Tools | Impressions and clicks in the platform’s own AI surfaces | High, but scoped to that platform | Presented as total AI visibility across all engines |
| Server logs | Verified crawler fetches by user agent and URL | Whether AI systems can and do fetch your content | High and unspoofable if IP-verified | Fetch counts reported as citation counts |
| Prompt sampling | Scheduled queries across engines | Citation, mention, sentiment, factual accuracy | Moderate, needs repetition to beat variance | Single-run results reported as trends |
| Outcome data | GA4, CRM, self-reported attribution, branded search | Whether visibility produced business results | Low to moderate, requires modelling | Modelled estimates presented as measured revenue |
Each layer answers a question the others cannot, and each fails differently — which is why a single composite “AI visibility score” hides more than it reveals.
Three reporting disciplines make the difference between a stack and a dashboard.
Separate leading from lagging indicators. Crawler access, coverage of enumerated questions, entity consistency and profile completeness are leading indicators you control. Citation rate and referral traffic are lagging indicators you influence. Reporting only the lagging ones produces a programme that looks like it is failing during the months when the work is actually being done.
Attach uncertainty to every number that has any. A citation rate from a 30-prompt sample run once has a confidence interval wide enough to be useless. Saying so is not weakness; it is the thing that makes the numbers you do report credible.
Baseline before you start. Almost no GEO engagement has a proper pre-intervention baseline, which makes every subsequent claim of improvement unfalsifiable. Record activation rate, citation rate, entity accuracy and crawler access before changing anything, and record it with enough repetitions to be a real baseline rather than a single reading.
One structural note on Google specifically. Google’s documentation states that traffic from AI features appears in Search Console’s Performance report under the Web search type, and that there is no separate tracking mechanism for it beyond the generative AI reports. That means AI Overview clicks have been mixed into your organic numbers all along. Any analysis that compared “organic traffic” before and after an AI rollout was comparing a category to a modified version of itself, which explains a great deal of the confusion in published traffic studies.
Sampling protocol for prompt-level tracking
Prompt-level tracking is the most common form of AI visibility measurement and the most frequently done wrong. The core problem is variance, and the size of it is now documented: repeating identical queries within 24 hours produces Jaccard similarity of only 0.34 to 0.42 across engines, meaning a majority of cited sources can change between two runs of the same prompt on the same day. The 2026 survey’s recommendation, derived from that variance, is seven to eight repetitions per prompt to detect a real effect.
Almost no commercial tool does this, because it multiplies cost by seven. Which means almost every AI visibility chart in circulation is plotting noise with a trend line through it.
A protocol that produces defensible numbers looks like this.
Define the prompt set from real demand, not from keywords. Prompts should be full sentences a real buyer would type, drawn from sales call recordings, support tickets, and the questions your team actually receives. Thirty to eighty prompts is a workable range for a single business. Stratify them: informational, comparison, recommendation, troubleshooting, local, and brand-specific. Report by stratum, because performance differs enormously between them and a blended average hides it.
Fix the prompt text and version it. Changing prompt wording between measurement periods breaks comparability. When you need to change a prompt, treat it as a new prompt and start its history fresh.
Repeat each prompt at least five times per measurement period, ideally seven or eight, spread across the period rather than fired in a burst. Record every run, not a summary.
Control what you can and record what you cannot. Log the engine, the model version where visible, the date and time, the geographic origin of the request, the account state (logged in or not, memory or personalisation enabled or not), and whether any search-grounding toggle was active. Personalisation and memory features materially change results, and a measurement taken from a logged-in account that has been researching your category for months is measuring that account, not the engine.
Record the full response text, not just whether you appeared. The response text is what lets you score mention, sentiment, factual accuracy and competitor presence later, including for questions you did not think to ask at collection time. Storage is cheap; re-collection is not, because the engine will have changed.
Score on multiple dimensions. Whether an AI answer appeared at all (activation). Whether your brand was mentioned. Whether you were cited with a link. Which URL was cited. Your share of the answer. Whether the claims attributed to you were accurate. Which competitors appeared. Which third-party sources were cited. The competitor and third-party columns are the most useful and the most commonly omitted, because they tell you what the answer is actually made of.
Report as rates with intervals, across periods, by stratum. “Cited in 34% of runs (95% CI 26–43%) across 240 runs of 30 recommendation prompts, up from 21% (CI 15–29%) in the previous quarter” is a defensible sentence. “AI visibility up 62%” is not.
Two methodological traps worth naming explicitly, because both appear in vendor reporting.
Selection on the outcome. Discarding runs where no AI answer appeared, or where no citations were present, removes the denominator and inflates every rate you compute. The survey identifies this as a pervasive flaw in published GEO research. Activation and no-citation runs are data, not failures of collection.
Model-family circularity. Using an LLM to judge whether an answer favours your brand, where the judge comes from the same family as the generator, introduces a bias toward text that resembles the model’s own output. If you use automated scoring — and at volume you must — validate it against a human-scored sample, blind to condition, and report the agreement rate.
A final practical note on cost. A proper protocol at 50 prompts, 7 repetitions, 5 engines, monthly, is 1,750 queries a month with full response capture. That is entirely feasible via APIs at modest cost, and it is a fraction of what most brands pay for tools that run one query per prompt per week. The constraint on rigorous AI visibility measurement is not budget. It is that rigorous measurement produces less flattering charts.
Server logs as the most honest AI visibility dataset
Log analysis is the least fashionable item on this checklist and the one I would keep if I could keep only one. Everything else in AI visibility measurement is inference from outputs. Logs are direct observation of inputs, they cannot be spoofed if you verify properly, and they answer questions no other source can.
What logs tell you, in order of usefulness.
Whether AI crawlers reach your content at all. This is the stage-zero diagnostic, and it fails more often than anyone expects. A log query for the documented user agents — OAI-SearchBot, GPTBot, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Bingbot, Googlebot, Google-Extended-affected fetches, plus the user-initiated fetchers — grouped by response code, immediately reveals blocks, challenges and rate limiting. A site returning 403 to OAI-SearchBot on 40% of requests has an infrastructure problem no content investment will fix.
Which content they prioritise. Fetch frequency by URL and by section tells you what each system considers worth revisiting. Sections that are never fetched are effectively invisible to that engine. This is a coverage map you cannot get any other way.
How stale their picture of you is. Time since last fetch per URL, per crawler, is your staleness map. A pricing page last fetched by a given crawler four months ago will be described using four-month-old prices no matter what you publish today.
Whether user-initiated fetches are happening. ChatGPT-User, Claude-User and equivalent agents fetch during a conversation. Their presence in your logs is evidence that assistants are sending real users’ questions at your pages in real time — the closest thing to a leading indicator of assistant-mediated interest that exists. Correlating those fetches with the URLs involved tells you which pages assistants reach for when a user pushes for detail.
Whether crawl volume is proportionate. Cloudflare’s data — more than half of crawler requests re-fetching unchanged pages, and crawl-to-referral ratios in the thousands or tens of thousands — is the industry aggregate. Your own ratio is computable from your logs and your referral data, and it is the number that should inform your licensing and blocking posture rather than someone else’s average.
Implementation requirements are modest but non-negotiable on one point. Verify crawler identity by IP, not by user-agent string. OpenAI, Anthropic and Google publish IP ranges for their crawlers as machine-readable files; unverified user-agent strings in logs include a substantial volume of spoofed traffic, and counting it produces both inflated crawl numbers and misdirected blocking decisions. Automate the verification against the published ranges and refresh them on a schedule.
Practical setup: retain raw access logs for at least 12 months if storage allows, parse into a queryable store, tag requests with verified crawler identity, and build four standing reports — response-code distribution by crawler, fetch frequency by site section, staleness by URL, and user-initiated fetch volume over time. This is a week of engineering work and it produces the only AI visibility data in your organisation that nobody can argue with.
Two caveats. Logs are increasingly incomplete at the edge: CDN-cached responses may not reach origin logs, so edge logs are usually the right source rather than application logs. And crawler activity is a necessary but not sufficient condition for citation — a page fetched daily may never be cited, because retrieval and citation are separate stages. Report fetches as fetches. The most common misuse of log data in GEO reporting is presenting crawler visits as evidence of AI visibility, which conflates stage one with stage four and is exactly the kind of category error this whole field needs less of.
Attribution beyond the referrer header
The attribution problem in AI search is structural, not a tooling gap, and understanding why prevents a great deal of wasted effort chasing a number that does not exist.
Three mechanisms break clean attribution. Google’s AI features report inside ordinary Search data — Google states that traffic from AI features appears in Search Console’s Performance report under the Web search type, meaning AI Overview and AI Mode clicks arrive in analytics as ordinary organic Google traffic. Assistant referrers are inconsistent: some pass a recognisable referring domain, some strip it, some route through intermediary domains, and behaviour varies by client, platform and whether the link opened in an in-app browser. And the largest share of AI influence produces no click at all — the Pew data showing 1% of visits producing a click on a cited source is the clearest available statement of that.
What follows is a practical attribution approach that accepts the limits.
Capture what the referrer does give you. Build a channel grouping in GA4 or your analytics platform for known assistant referrers — the ChatGPT domains, Perplexity, Copilot, Gemini, Claude, and the smaller engines — and maintain it, because new hosts appear regularly. This undercounts systematically. It is still worth having, because the traffic it does capture is measurable, comparable over time, and behaviourally distinct.
Segment and study the behaviour of that traffic rather than just counting it. Assistant-referred sessions frequently show different patterns from organic search sessions: fewer pages per session, higher engagement on the landing page, different device mix, and in several published analyses better conversion rates on commercial intent. Whether the multiplier is the widely quoted “four times more valuable” or something more modest is not settled — those figures come from vendor analyses with varying rigour, and the academic literature offers only one quasi-experimental estimate of a 1.82× traffic multiplier with wide confidence intervals. Treat the direction as plausible and the magnitude as unknown.
Add a self-reported attribution field to your primary conversion form. “How did you hear about us?” with an option for AI assistant, as a free-text or option-list field. This is the single most informative attribution instrument available for assistant-mediated discovery, because it captures the influence that produced no click and no referrer. It is noisy, subject to recall bias, and vastly better than nothing. Sales teams should also be asking the question in discovery calls and logging the answer in the CRM.
Watch branded search and direct traffic as proxies. Assistant-mediated discovery frequently ends with the user searching your brand name or typing your domain, because that is how people behave when they have been told a name. Rising branded search volume with flat non-branded volume, in a period when your assistant citation rate rose, is corroborating evidence. It is not proof, and it should be presented as corroboration.
Model rather than count, and say that you are modelling. A defensible executive number looks like: “Assistant-referred sessions, measurably attributed: 4,200. Estimated assistant-influenced sessions including branded-search and direct effects: 11,000 to 19,000, based on self-reported attribution rates applied to branded and direct segments. The wide range reflects the attribution gap and should not be narrowed without better instrumentation.” That sentence has survived every board conversation I have watched it enter, because it is honest about what it is.
Two things to stop doing.
Stop reporting “AI traffic” as a channel with a precise number. The number is a floor, not a measurement, and presenting it as a measurement means that when the floor rises for technical reasons — an engine starting to pass referrers — you will report a growth figure that reflects nothing but instrumentation.
Stop attributing organic declines to AI without controlling for anything else. Traffic fell for many sites in 2025 and 2026, and AI answers are clearly one cause: Chartbeat data across more than 2,500 news sites showed Google search traffic to those sites down 33% year over year between November 2024 and November 2025, with Google Discover down 21%, and Reuters Institute’s January 2026 survey of 280 media leaders across 51 countries found an expectation of a further 43% decline in search referrals over three years. Those are real and large. They are also concurrent with core algorithm updates, spam-policy enforcement, changes in Discover behaviour, and shifting user habits. A site-level decline needs a query-level and page-level diagnosis before it gets an AI explanation, because the remedy differs completely depending on the cause.
Sector-by-sector reading of the same checklist
The checklist is universal. Its priority order is not. The same twelve items produce entirely different sequencing depending on how your category’s demand is phrased, how heavily AI answers activate on it, and what the consequences of an inaccurate answer are.
Publishing and media. This sector faces the sharpest version of the problem: high activation, heavy summarisation, and answers that frequently satisfy the user’s need completely. Reuters Institute’s 2026 survey found media leaders expecting a 43% decline in search referrals over three years, with roughly one in five expecting substantial revenue from AI platform deals and two-thirds reporting no job savings from AI efficiency gains yet. The strategic response documented in that survey is not GEO — it is direct-audience investment, with net positive sentiment toward YouTube, TikTok and Instagram, and 76% of managers shifting staff toward creator-like roles. For publishers, the checklist priorities are licensing posture, crawler control, entity and author authority, and original reporting that cannot be synthesised from other sources. Optimising a news article for citation is a marginal exercise when the underlying economics of referral have moved.
B2B software. High activation on comparison and recommendation queries, and a citation substrate dominated by third parties — G2, Capterra, Reddit, LinkedIn. Peec AI’s data showing Perplexity emphasising Reddit, LinkedIn and G2 for B2B queries is directly actionable. Priorities: third-party profile accuracy and completeness, published pricing, honest comparison and alternatives content, fit-boundary statements, integration and compliance documentation as structured facts, and community participation with disclosure. On-site blog content ranks low on this list, which contradicts most B2B content strategies.
E-commerce and retail. The mechanism is feed quality and agentic commerce readiness, not content. Priorities: attribute depth, feed-to-page-to-markup consistency, availability honesty, returns and shipping as structured facts, protocol readiness for agent-initiated purchase, and review substrate health. A retailer with 4,000 SKUs and an average of five populated attributes each has a larger addressable improvement in agent visibility than any content programme could deliver.
Financial services. High activation on informational queries, severe consequences for inaccurate answers, and heavy regulatory constraint on what may be published and how. Priorities: factual accuracy and fidelity monitoring above visibility, regulator-register presence, clear jurisdictional scoping of every claim, dated rates and thresholds, author credentials as machine-readable entities, and disclaimers that survive extraction. The specific risk is a passage lifted without its regulatory qualifier — which is a compliance event, not a marketing outcome. data-nosnippet on qualifying language is exactly the wrong move; the qualifier needs to be inside the same sentence as the claim so it cannot be separated.
Healthcare. The highest-stakes fidelity environment. AI systems apply additional care to health topics and draw disproportionately on institutional sources. Priorities: institutional and regulator-register presence, clinician credentials as entities, alignment with recognised clinical guidance with citation, explicit scope-of-practice and jurisdiction statements, and monitoring of what assistants say about your specific services. Visibility work that increases citation while allowing an ambiguous clinical claim to circulate is a net negative.
Professional services — legal, accounting, consulting. Demand is heavily situation-shaped, which makes fan-out coverage unusually productive: the enumerated-question method described earlier maps almost perfectly onto how clients describe problems. Priorities: named-expert entity building, jurisdiction-specific substantive content, professional-register presence, local and service-area data, and honest scope statements about what you do and do not handle. This is the sector where the individual expert’s entity matters more than the firm’s.
Local and home services. Almost entirely a structured-data exercise. Priorities: Business Profile completeness, NAP consistency, review volume and response, real location pages, booking and availability integration, and agent-contact readiness. Content investment beyond a handful of genuinely local pages has poor returns here relative to profile and review work.
Manufacturing and industrial B2B. Long, technical, specification-driven queries with low volume and high value, and a competitive field that has mostly not done the work. Priorities: specification data as structured, extractable facts; compatibility and standards compliance stated precisely; datasheets in HTML rather than only PDF; distributor and catalogue-platform accuracy; and technical documentation depth. Converting a PDF-only specification library into indexable HTML is frequently the highest-return single action available in this sector.
The pattern across all eight: the further your category sits from content-shaped demand and the closer it sits to structured, verifiable facts, the less your GEO programme should look like content marketing and the more it should look like data governance. Most agencies sell the first shape to every sector, which is why results are inconsistent in ways that have nothing to do with the engines.
Licensing, compensation and the economics of being crawled
Access control and licensing have become part of the GEO checklist rather than a separate legal matter, because the decision about who may fetch your content now determines where you can appear.
The economic argument that drove this is quantitative. Cloudflare’s published network data shows 57.4% of traffic to HTML across its network coming from bots, with training-related crawlers at 50.6% of total traffic and more than half of crawler requests re-fetching pages that have not changed. The crawl-to-referral ratios are the figures that ended the assumption of reciprocity: roughly 38,000 pages fetched per referral for Anthropic’s crawler and about 1,091 crawls per referral for OpenAI’s. Whatever one thinks about the ethics, those ratios are not the bargain publishers accepted with search engines.
The market response has moved fast through three phases. Cloudflare launched pay-per-crawl in private beta on 1 July 2025, letting publishers return HTTP 402 Payment Required with a price per crawler and acting as merchant of record; it expanded the pricing options in August 2025; it moved to blocking AI crawlers by default for new domains; and on 1 July 2026 it announced a shift from charging per crawl to compensating per citation, with named partners including Ceramic.ai and You.com. The stated rationale is that crawling is a poor proxy for value, since a page may be crawled once and cited thousands of times or crawled constantly and never cited.
In parallel, Really Simple Licensing launched on 10 September 2025 as an open, RSS-derived protocol for expressing machine-readable licensing terms via robots.txt and an accompanying licensing document, supporting free, attribution, subscription, pay-per-crawl and pay-per-inference models. Its initial backers included Reddit, Yahoo, People Inc., Internet Brands, Ziff Davis, O’Reilly Media, Medium, The Daily Beast, wikiHow, Raptive, Ranker, Quora, Fastly and Evolve Media, alongside a nonprofit RSL Collective offering collective negotiation on the model of music rights organisations. RSL AI Licensing reached 1.0 status as an industry standard with expanded capabilities during 2026.
The standards track is running underneath all of it. The IETF AI Preferences working group, chartered in early 2025 and co-chaired by Suresh Krishnan and Mark Nottingham, is producing two deliverables: a common vocabulary for expressing preferences about how content may be collected and processed for AI, and mechanisms for attaching that vocabulary to content either by embedding it or through robots.txt-like formats. The working group’s premise is stated bluntly in its own announcement — AI vendors currently use a confusing array of non-standard signals. Until AIPREF completes, expressing preferences means maintaining a per-vendor patchwork, which is precisely the operational burden the standard exists to remove.
What this means for a practical checklist.
Decide your posture per purpose, not per company. Search citation, training, user-initiated fetch and agent access are four separate questions with different business answers. The most common defensible position for a commercial site is: allow all search crawlers, allow user fetchers, block training crawlers, decide agent access deliberately. The most common defensible position for a publisher with licensable archives is the same, plus an explicit licensing offer.
Publish machine-readable terms. Whether via RSL, a licensing statement in robots.txt, or terms-of-use language that specifically addresses automated access and AI training, having explicit terms matters for any future negotiation or dispute. Silence has historically been read as permission.
Enforce at the edge, not only in a text file. Robots.txt is a request. Bot management at the CDN with IP verification is enforcement. Crawlers that ignore robots.txt — and user-initiated fetchers that documentedly may not follow it — are only addressable at the network layer.
Quantify your own ratio before deciding. Compute crawls per referral per crawler from your own logs and referral data. A site receiving meaningful referral volume from an engine is in a different position from one receiving none, and the industry average tells you nothing about your case.
Understand what blocking does not achieve. It does not remove content already ingested. It does not affect third-party datasets unless those are blocked separately. It does not prevent your content reaching a model via a syndication partner, an aggregator, or a scraper. And it does remove you from that engine’s citations, which for a business dependent on discovery is a real cost that should be priced rather than assumed away.
The honest summary of the economics: for most commercial businesses, the referral and brand value of appearing in AI answers exceeds the licensing value of their content, and blocking search crawlers is self-harm. For publishers whose product is the content itself, the calculus inverts, and the emerging pay-per-citation and collective-licensing structures are the first mechanisms that make the inverted position economically coherent rather than merely defensive.
Regulatory and compliance boundaries
GEO sits inside a regulatory perimeter that most practitioners have not read, and several standard tactics sit on the wrong side of it.
The EU AI Act’s general-purpose AI obligations took effect on 2 August 2025, with the accompanying General-Purpose AI Code of Practice providing the compliance route. The provisions relevant to content owners are the copyright-policy requirement and the training-data transparency requirement: providers must maintain a policy for complying with EU copyright law, including honouring machine-readable reservations of rights, and must publish a sufficiently detailed summary of training content. The practical consequence for publishers is that a machine-readable rights reservation is no longer purely advisory — it is the mechanism the Act’s copyright provisions reference, which is a substantial part of why the RSL and AIPREF efforts exist. Enforcement of the GPAI provisions phases in through 2026 and 2027, with obligations for models placed on the market before the cut-off following a longer timeline.
Article 50 transparency duties matter to anyone deploying AI in their own content production or customer interactions: AI-generated or manipulated content in certain categories must be disclosed, chatbots must be identifiable as such, and synthetic media requires marking. A content operation using generative tools at scale should have a documented position on disclosure.
Consumer protection law constrains several popular tactics. Undisclosed paid endorsement — the mechanism behind community astroturfing and paid “brand mention” services — is a prohibited practice under the EU’s Unfair Commercial Practices Directive as amended, and is enforced by the FTC in the United States under its endorsement guides. Fake and incentivised reviews were addressed directly in the EU’s Omnibus Directive, which requires traders to disclose whether and how they verify reviews and prohibits submitting or commissioning false reviews. Any GEO tactic that manufactures third-party endorsement is a consumer-protection exposure before it is an SEO risk.
Comparative advertising rules apply to the comparison and alternatives content that recommendation queries reward. Under the EU’s Misleading and Comparative Advertising Directive, comparative advertising is lawful where it compares goods meeting the same needs, compares verifiable and representative features objectively, and does not mislead or denigrate. A feature matrix built to flatter would fail that test in a formal complaint, which is one more reason the honest version of comparison content is the commercially sound one.
Sector-specific regimes override general practice. Financial promotion rules in most jurisdictions require specific risk warnings attached to specific claims. Medical device and pharmaceutical advertising rules restrict claims and require approved language. Legal services advertising is constrained by professional bodies. The AI-specific wrinkle is extraction: a compliant page can produce a non-compliant quotation when a system lifts the claim and leaves the warning behind. The mitigation is structural — put the qualifier in the same sentence as the claim, not in a footer or a separate paragraph — and it belongs in content templates rather than in a reviewer’s memory.
Data protection intersects at two points. Consent-management implementations that block content rendering to non-consenting agents produce the invisibility problem described earlier, and the better implementation gates scripts rather than content. And any prompt-level monitoring programme that captures assistant responses containing personal data — reviews naming individuals, complaints, employee mentions — is processing personal data and needs a lawful basis and a retention policy like any other dataset.
Accessibility deserves a mention because it correlates with the structural work GEO requires. Semantic headings, real heading hierarchy, text alternatives, transcripts for video, and content that exists without JavaScript are accessibility requirements under the European Accessibility Act’s phased application and equivalent regimes elsewhere. They are also, almost item for item, the technical foundation for machine extractability. The overlap is close enough that framing the work as accessibility compliance often unlocks budget that GEO framing does not.
The disclosure question that has no settled answer yet: should content optimised for machine consumption be labelled as such, and should AI-assisted content be labelled? Google’s position is that AI-assisted content is acceptable when it meets quality and spam standards, with no labelling requirement for Search. The EU AI Act’s Article 50 duties bite on specific categories rather than on marketing copy generally. The prudent position for a business with a brand to protect is to disclose AI involvement where a reader would reasonably want to know, and to treat the absence of a legal requirement as a temporary condition rather than a permanent licence.
Manipulation, prompt injection and the line GEO cannot cross
The GEO field has a black-hat wing, and the reason to discuss it on a practical checklist is not curiosity. It is that several techniques being sold as clever are either ineffective, actively dangerous to the client, or both — and that the boundary between aggressive optimization and manipulation has been drawn more clearly than most practitioners realise.
The 2026 survey proposes four cumulative tests for distinguishing white-hat optimization from manipulation, and they are the most usable formulation I have seen. Semantic preservation: after the rewrite, the facts remain true. Evidentiary authenticity: statistics and quotations are verifiable. Content–instruction separation: the page contains no text intended to instruct a model rather than inform a reader. Disclosure and fairness: commercial intent is transparent and competitors are not misrepresented. Reorganising paragraphs, adding a definition, citing a real source, restructuring for extractability — all pass. Hidden model-directed text, fabricated testimonials, invented statistics, and misrepresenting a competitor’s product — all fail.
The third test is where the interesting failures live. Prompt injection via web content — embedding instructions in a page intended to be read by a model rather than a human, in white-on-white text, zero-size fonts, aria-hidden containers, HTML comments or CSS-clipped elements — became a documented tactic and then became a security category. It is now ranked as the top LLM security risk in OWASP’s classification, and Google has publicly described prompt injection as having moved from theoretical concern into observed abuse.
Three reasons it does not work as a marketing tactic, in increasing order of severity.
It is increasingly detected and neutralised. Retrieval pipelines sanitise inputs, strip invisible text, and separate retrieved content from instruction context. The technique’s window was narrow and has largely closed for the major engines.
It is cloaking under existing search policy. Hidden text serving a different message to machines than to humans is a long-established violation with manual-action consequences, and there is no reason to expect a carve-out because the machine is a language model.
It is an attack, not an optimization. Instructions embedded in a page that alter an assistant’s behaviour toward a user are a form of unauthorised interference with someone else’s system and someone else’s session. In a jurisdiction with computer-misuse legislation, the analysis is not obviously favourable to the publisher. A brand caught doing this does not have an SEO problem. It has a disclosure event.
Two adjacent tactics deserve the same treatment. Synthetic authority signals — AI-generated author personas with fabricated credentials, invented institutional affiliations, fake certifications — fail the disclosure test and create a directly falsifiable lie about the business. Manufactured consensus — coordinated posting across communities and review platforms to produce apparent agreement — fails both disclosure and fairness, and carries the consumer-protection exposure described earlier.
There is a defensive side to prompt injection that belongs on the checklist for a completely different reason. If your site accepts user-generated content — reviews, comments, forum posts, profile fields, support tickets — that content can carry injection payloads aimed at any AI system that later reads your pages, including your own. A review containing hidden instructions is a vector into your own AI-powered support tool, your summarisation features, and any assistant that grounds on your site. Sanitising user-generated content for invisible text, unusual Unicode, and instruction-shaped patterns is now a security control, not a spam filter.
The same applies to any AI feature you operate. A site-search assistant, a documentation chatbot or a product-recommendation agent grounded on your own content inherits every injection payload in that content. Treating retrieved content as untrusted input, with the same suspicion you apply to form submissions, is the baseline mitigation.
The commercial argument against black-hat GEO is stronger than the ethical one, which is why it works better in a client conversation. These techniques have short half-lives, are detectable, produce no durable asset, and convert a marketing budget into a legal and reputational liability. The white-hat programme — access, extractability, entity clarity, third-party accuracy, measurement — produces compounding assets that survive engine changes. That is not a moral argument. It is an asset-quality argument, and it is the one that survives contact with a CFO.
Failure modes, misattribution and reputational exposure
Visibility in AI answers is not uniformly good, and a checklist that only pushes toward more citation without monitoring what is being said is incomplete in a way that eventually produces an incident.
The fidelity data is the starting point. Research collected in the 2026 survey found only 51.5% of generated sentences fully supported by their cited sources in one early evaluation, with 74.5% of citations correctly supporting the attached proposition, and a 2026 analysis putting roughly 11% of claims as insufficiently supported. Credible-source shares ranged from 71.4% to 86.3% depending on the assistant and the topic. Those numbers describe a system that is usually approximately right and regularly specifically wrong.
The failure modes that produce business consequences, with their diagnostics.
Wrong facts attributed to you. An assistant states your price, your policy, your capability or your limitation incorrectly and cites your site. The user acts on it. This is the most common damaging failure and it is detectable only by asking the questions and reading the answers. Remedy: make the correct fact unambiguous, dated and prominent on the page the system is citing, and check whether an outdated cached version is the actual source.
Your content used to support a claim you did not make. A passage lifted without its qualifier, a statistic separated from its scope, a recommendation separated from the conditions under which it applies. This is the extraction risk that matters most in regulated sectors. Remedy: qualifiers inside sentences, scope stated inline, and no reliance on adjacent paragraphs to carry conditions.
Competitor conflation. The answer describes your competitor’s feature as yours or vice versa. Common where naming is similar or where a third-party comparison source is inaccurate. Remedy: entity disambiguation work and correction of the third-party source.
Stale information persisting. Discontinued products, former pricing, retired policies, staff who left. Remedy: the maintenance discipline described earlier, plus explicit correction pages where the old claim was widely reproduced.
Negative synthesis. The assistant accurately summarises a body of genuine criticism. This is not a visibility failure and there is no content remedy. Recognising the difference between an information problem and a product problem is the useful skill here.
Absence from your own category. You do not appear in answers where competitors do. Diagnose by stage: not activated (no AI answer for these queries at all), not fetchable (logs), not retrieved (no coverage of the sub-questions), retrieved but not cited (extractability), or cited but not mentioned by name (entity clarity).
Building the monitoring loop is straightforward once the failure taxonomy exists. Take your prompt set, add a specific block of factual-accuracy prompts — what does [brand] cost, does [brand] support [thing], is [brand] available in [country], what are [brand]’s limitations, is [brand] compliant with [standard] — run them across engines on the sampling protocol, and score the responses for factual accuracy rather than presence. Route material errors to whoever owns the underlying fact, not to the marketing team, because the fix is usually a documentation or data change rather than a content change.
One organisational recommendation that has proven itself: give someone explicit ownership of “what AI systems say about us”, with a monthly report that leads on accuracy and sentiment and treats citation rate as secondary. The teams that run GEO as an accuracy programme with a visibility side-effect consistently produce better outcomes than the teams running it as a visibility programme, partly because accuracy work fixes the underlying data that drives visibility, and partly because it produces findings that other departments act on.
Original research as the one durable citation advantage
Almost every GEO tactic erodes under adoption. The C-SEO Bench finding of congestion approaching a zero-sum game applies to formatting, to structure, to definitions, to schema — anything a competitor can replicate by reading your page and copying the pattern. One category does not erode: facts that exist only because you produced them.
Original data is the closest thing to a defensible position in AI search, for a mechanical reason rather than a rhetorical one. When an assistant needs a number and only one source has it, the retrieval problem has a single answer. When ten sources have the same number, citation distributes across them and drifts toward whichever is most authoritative or most recently fetched. Proprietary data collapses the competitive set to one.
What counts as producible original data, in rough order of cost.
Aggregated operational data you already hold. Any business with volume has a dataset nobody else can compute. Average resolution times. Failure rates by category. Seasonal demand patterns. Regional price variation. Typical implementation duration by company size. Migration timelines. Most of this can be published in aggregate without disclosing anything sensitive, and most of it answers questions buyers actually ask. This is the highest-return content asset available to most businesses and almost nobody produces it, because it requires cooperation between marketing and whoever owns the data warehouse.
Structured customer surveys with published methodology. Sample size, sampling frame, dates, question wording, and the raw distributions. A survey of 400 practitioners in your sector, published properly, becomes the citable source for its subject for as long as it stays current. Publish the methodology visibly, because that is what separates a citable study from a marketing statistic.
Benchmark and comparison testing. Measured, reproducible tests of things in your category — performance, durability, accuracy, cost over time — with the method described in enough detail that someone could repeat it. This is expensive and it produces the kind of source that third parties cite, which is how you get into the substrate rather than just onto the citation list.
Longitudinal tracking. The same measurement, repeated on a schedule, published each period. The value compounds: a price index in its fourth year is a reference series, and reference series get cited reflexively.
Expert judgment attributed to a named person. Not data, but similarly non-replicable. A named specialist with a documented record stating a view, with reasoning, is a quotable source a competitor cannot copy. This is why author entity work and original research reinforce each other.
Three implementation requirements determine whether any of this gets cited.
Publish the number in extractable form. The finding in a sentence with its number, its unit, its scope and its date, near the top of the page. A study whose headline finding is only available in a chart image or a downloadable PDF is invisible to text retrieval. Charts should have text equivalents, and PDFs should have HTML versions — the PDF-only research report is one of the most common self-inflicted GEO wounds in industrial and professional sectors.
Make the methodology visible and the data downloadable. Both increase third-party citation, which is what moves you into the substrate that recommendation queries draw on.
Give it a stable, canonical URL and update it in place. A study that moves URLs each year fragments its own citation history. One URL, updated annually, with a visible revision note, accumulates authority.
The strategic point, stated plainly: if your GEO budget is spent entirely on rewriting existing pages, you are competing in the category where advantage erodes fastest. Moving a meaningful share of it into producing facts that did not previously exist is the only reliable way to build a position that a competitor cannot copy by reading your site. It is slower, it is harder to sell as a monthly retainer, and it is the part of the work that still matters in three years.
The checklist assembled in execution order
Everything above collapses into a sequence. The order matters more than the contents, because most failing GEO programmes are doing items eight through twelve while items one through three are broken.
Phase one — access and eligibility. Do this before anything else, and do not proceed until it passes.
Fetch a sample of important URLs as each AI user agent from a non-allowlisted IP and record status codes and body length. Fix hard blocks, soft blocks, challenge pages and rate limits. Verify crawlers by published IP range rather than user-agent string. Audit robots.txt for inherited disallows, group-precedence errors and blocked render resources. Check for X-Robots-Tag headers applied at the CDN. Confirm no unintended noindex, nosnippet or low max-snippet on content pages. Confirm content exists in raw HTML without JavaScript. Confirm consent and geo logic serve content rather than walls to non-human agents. Measure cold, uncached time-to-content. Set your crawler posture deliberately per purpose — search, training, user fetch, agent — and document it with an owner and a quarterly review.
Phase two — identity. Do this second, because everything downstream describes an entity that must be unambiguous.
Publish a reference-quality about page with legal name, trading names, founders, founding date, headquarters, markets, category description and corporate relationships. Standardise naming across every surface. Implement connected Organization and Person markup with @id references and sameAs links to every authoritative external profile. Create or correct the Wikidata item where notability applies. Audit and correct every third-party profile: LinkedIn, Crunchbase, industry associations, registries, regulator registers, app stores, review platforms. Test entity accuracy by asking five assistants who you are, several times, and log the errors.
Phase three — coverage. Now content becomes worth doing.
Enumerate the questions generated by your top buyer situations rather than picking keywords. Map existing content against them and mark each as answered well, thin, wrong or absent. Prioritise gaps by commercial proximity. Build clusters with real depth at the leaves. Publish pricing. Publish honest fit boundaries and limitations. Publish category overviews and accurate alternatives content. Convert PDF-only material to HTML.
Phase four — extractability. Applied to what phase three produced.
One question per section. Answer in the first two sentences under the heading. Descriptive headings that make passages self-interpreting. Explicit subjects at paragraph boundaries. Numbers with units, scope and inline dates. Standalone definitions of your core terms. Qualifiers inside the sentences they qualify. Tables for tabular comparisons. Real heading elements in real hierarchy. Then read your own pages as plain-text extractions and check whether the claims survive.
Phase five — substrate. The part most programmes skip and most answers depend on.
Complete and correct your profiles on the review and comparison platforms your category’s answers are actually built from. Earn independent editorial coverage with named experts and cited data. Participate in relevant communities with disclosure. Correct factual errors about you in community threads. Publish video with accurate transcripts where demonstration matters. Ensure presence in the structured industry sources your sector’s answers rely on.
Phase six — structured operational data, weighted by sector.
For retail: feed completeness, attribute depth, feed-page-markup consistency, availability honesty, returns and shipping as structured facts, agentic protocol readiness. For local and services: Business Profile completeness, NAP consistency, review volume and response, real location pages, booking integration, agent-contact readiness. For industrial: specifications as extractable HTML facts, standards compliance stated precisely, distributor accuracy.
Phase seven — original research.
Identify aggregated operational data you can publish. Run one properly documented survey or benchmark. Publish findings in extractable form with visible methodology at a stable URL. Plan the second edition before the first is published.
Phase eight — measurement, which should have started in phase one.
Baseline before intervention. Connect Search Console generative AI reports and Bing Webmaster Tools AI reporting. Build verified-crawler log analysis with four standing reports. Define a stratified prompt set from real demand. Run five to eight repetitions per prompt per period with full response capture. Score activation, mention, citation, share, accuracy, competitors and third-party sources. Report rates with intervals by stratum. Add self-reported attribution to conversion forms and sales discovery. Model influenced volume with explicit ranges. Never report fetches as citations or modelled estimates as measurements.
Phase nine — governance, permanently.
Quarterly crawler-map review with a named owner. Monthly accuracy report on what assistants say about you, routed to whoever owns the underlying facts. Content inventory with per-asset decay rates and review dates. Third-party profile audit on the same cycle as on-site audits. Licensing posture reviewed against your own crawl-to-referral ratios. Compliance review of extraction risk in regulated claims. User-generated content sanitised against injection payloads. A standing prohibition on hidden text, fabricated statistics, synthetic authority and manufactured consensus, written down so nobody has to relitigate it.
Two things to keep in view while executing all of it. The first is that the honest evidentiary position — already-retrieved content can causally alter its own citation, but no technique has shown a stable cross-platform effect on discoverability or behaviour — means every promise should be scoped to a stage. The second is that most of this list is good engineering, good documentation, good data governance and good public relations. The reason it works is not that it games a generative model. It is that a business whose facts are accurate, complete, consistent, structured and reachable is easier for any system to describe correctly — and being described correctly, at scale, by machines that increasingly stand between you and your market, is the actual objective.
Questions readers ask about GEO in practice
Not structurally. Google states that its AI features rely on the same ranking and quality systems as classic Search, and that no special optimizations are required to appear in AI Overviews or AI Mode. GEO is a reordering of SEO priorities around passage retrieval and citation, plus genuinely new work in crawler access control, prompt-level measurement, licensing posture and structured operational data.
Not for Google Search. Google’s documentation states plainly that it does not use llms.txt files and that they neither help nor harm visibility. No major AI vendor has documented using it as a retrieval signal. It costs almost nothing to publish, and it should not be presented to a client as a visibility measure.
No. Google says structured data is not required for generative AI search and that no special schema.org markup exists for it. Microsoft has said schema helps Bing’s models understand content for Copilot. Independent testing has found no correlation between schema coverage and citation rates across OpenAI, Gemini and Perplexity. Implement it for feature eligibility and entity clarity, not on a promise of citation lift.
For Google’s AI Overviews and AI Mode, Googlebot — there is no separate AI crawler. For ChatGPT search, OAI-SearchBot. For Claude’s search, Claude-SearchBot. For Copilot, Bingbot. For Perplexity, PerplexityBot. Blocking any of these removes you from that engine’s answers.
Yes, for the major vendors. Block GPTBot and ClaudeBot for training while allowing OAI-SearchBot and Claude-SearchBot for search, and set Google-Extended, which controls training and grounding in Google’s other AI products without affecting Search inclusion. Blocking does not remove content already ingested, and user-initiated fetchers are documented as possibly not following robots.txt.
Not in the way it is usually quoted. The figure is a relative gain in Position-Adjusted Word Count from 19.3% to 27.2% inside a fixed five-document context window. It measures share of a generated answer once a document is already retrieved. It is not a click, a citation or a customer, and it says nothing about whether you get retrieved in the first place.
Pew Research Center’s study of 68,879 Google searches found users clicked a traditional result on 8% of visits to pages with an AI summary versus 15% without, with clicks on sources cited inside the summary occurring on 1% of visits. Sessions ended entirely on 26% of pages with a summary versus 16% without.
Peec AI’s analysis of 30 million sources found Reddit first and YouTube second, with LinkedIn, Wikipedia and Forbes in the top five, and Yelp and G2 prominent in recommendation queries. Preferences differ by engine: ChatGPT favoured Wikipedia, Reddit and Forbes; Google leaned toward Facebook and Yelp; Perplexity emphasised Reddit, LinkedIn and G2 for B2B.
Largely no. URL-level Jaccard similarity between Google’s classic results, AI Overviews and Gemini answers has been measured at 0.11 to 0.18, and only 26% of domains appeared in both Bing Chat and Perplexity responses in one comparison. Cross-engine visibility has to be measured per engine.
Seven to eight repetitions per prompt is the published recommendation, because repeating identical queries within 24 hours produces Jaccard similarity of only 0.34 to 0.42 across engines. Single-run measurements are indistinguishable from engine variance.
Search Console gained generative AI performance reports in June 2026 for Google’s AI surfaces, and Bing Webmaster Tools offers AI performance reporting for Microsoft’s. Google’s guidance warns against third-party tools claiming access to internal Google metrics, because no such access exists. Traffic from Google’s AI features also appears in the ordinary Search performance data under the Web search type.
Three reasons. Google’s AI features report as ordinary organic Search traffic rather than as a separate channel. Assistant referrers are inconsistent, with some engines stripping the referring domain. And most AI influence produces no click at all. Analytics figures for AI referrals are a floor, not a measurement.
Question-and-answer formatting makes answers extractable and aligns with the question-formatted queries that trigger AI answers most often. It stops working when questions are invented to host keywords and answers restate the page, which falls under scaled content abuse in Google’s spam policies. Use real questions from support and sales, and stop when you run out of them.
No. Google states there is no requirement to break content into tiny pieces and that its systems understand multiple topics on a page. Fragmenting content produces thin pages that struggle to rank, which removes them from the pool AI features draw on. Write self-contained passages inside well-organised long pages instead.
Substantially, for commercial queries. When a vendor does not publish a price, assistants either state that pricing is undisclosed, quote a third-party estimate that may be wrong, or attribute a competitor’s price. Publishing a starting price, a pricing model and the variables that move it is now a retrieval requirement rather than a conversion preference.
It is a documented tactic that has largely stopped working, is treated as cloaking under existing search policy, and is classified as the top LLM security risk by OWASP. Retrieval pipelines sanitise invisible text and separate retrieved content from instruction context. A brand caught doing it has a disclosure problem, not a ranking problem.
Undisclosed paid endorsement and fake reviews are prohibited under EU consumer-protection rules and enforced by the FTC in the United States. Comparative advertising must be objective and verifiable under EU rules. The EU AI Act’s general-purpose AI obligations, effective from August 2025, require providers to honour machine-readable rights reservations and publish training-content summaries, which is why machine-readable licensing signals now carry legal weight.
Mechanisms are emerging. Cloudflare ran pay-per-crawl from July 2025 and announced a shift to compensating per citation on 1 July 2026. Really Simple Licensing launched in September 2025 with pay-per-crawl and pay-per-inference models and a collective licensing organisation, backed by Reddit, Yahoo, Ziff Davis, O’Reilly, Medium, Quora and others. The IETF’s AI Preferences working group is standardising how preferences are expressed.
Verifying that AI crawlers can actually fetch your content, and fixing it where they cannot. Blocks, challenge pages, consent walls and client-side rendering failures make every other item on the list irrelevant, and they are common enough that the check should precede any content or markup work.
No. Google AI Overviews activate on roughly 13.7% of queries overall but 64.7% of question-formatted ones. If your demand is dominated by navigational and transactional phrasing, activation is low and the investment case is weak. Measuring activation rate across your own keyword set takes an afternoon and should precede any budget decision.
Author:
Jan Bielik
CEO & Founder of Webiano Digital & Marketing Agency

This article is an original analysis supported by the sources cited below
Google’s guide to optimizing for generative AI features on Google Search Google’s primary guidance for site owners, stating that no special optimizations are required for AI features, that llms.txt is not used, and that structured data is not required for generative AI search.
AI features and your website Google Search Central documentation on how content appears in AI Overviews and AI Mode, query fan-out, the nosnippet and max-snippet controls, Google-Extended, and where AI traffic appears in Search Console.
Introducing Search Generative AI performance reports in Search Console Google’s June 2026 announcement of first-party reporting for performance in its generative AI surfaces.
Top ways to ensure your content performs well in Google’s AI experiences on Search Google’s earlier statement of the same position, useful for tracking how the official guidance has and has not changed.
Google Search’s I/O 2026 updates Google’s announcement of AI Mode passing one billion monthly users, the Gemini 3.5 Flash default, the redesigned search box, and agentic information, booking and coding features.
Google users are less likely to click on links when an AI summary appears in the results Pew Research Center’s field study of 68,879 searches, the source of the 8% versus 15% click rates and the 1% click rate on cited sources.
GEO: Generative Engine Optimization The Aggarwal et al. paper that introduced the term and the widely quoted 40% visibility figure, with the GEO-bench benchmark.
GEO: Generative Engine Optimization, KDD 2024 proceedings The peer-reviewed publication record for the foundational GEO paper.
Optimizing visibility in generative engines: a critical survey of generative engine optimization A 2026 critical review of 45 GEO studies, the source of the visibility vector, the reframing of the 40% claim, the cross-engine agreement figures, the repetition requirement, and the white-hat test set.
38% of AI Overview citations pull from the top 10 Ahrefs’ analysis of the overlap between AI Overview citations and classic organic rankings.
Overview of OpenAI crawlers OpenAI’s documentation of OAI-SearchBot, GPTBot, OAI-AdsBot and ChatGPT-User, including the statement that opting out of OAI-SearchBot removes a site from ChatGPT search answers.
Anthropic’s Claude bots make robots.txt decisions more granular Coverage of Anthropic’s February 2026 split into ClaudeBot, Claude-SearchBot and Claude-User, and the practical implications for allowing citations while blocking training.
Google confirms llms.txt has no current implementation Reporting on Google’s position that llms.txt is not used by Search.
IETF setting standards for AI preferences The IETF’s announcement of the AIPREF working group, its two deliverables and its premise that vendors currently use non-standard signals.
draft-ietf-aipref-vocab The working group’s draft vocabulary for expressing content usage preferences for AI.
New RSL web standard and collective rights organization automate content licensing The Really Simple Licensing launch announcement, its licence models including pay-per-crawl and pay-per-inference, and its initial publisher backers.
RSL: Really Simple Licensing The standard’s specification and implementation documentation.
How Really Simple Licensing may change online content licensing Legal analysis of RSL’s mechanics and enforceability from a commercial law perspective.
Cloudflare stops charging AI per crawl and starts paying per answer Reporting on Cloudflare’s July 2026 shift to citation-based compensation, with the crawl-to-referral ratios and the bot traffic share figures.
AI search engines cite Reddit, YouTube and LinkedIn most Coverage of Peec AI’s analysis of 30 million sources and the platform-by-platform citation preferences.
News publishers expect search traffic to fall by more than 40% in the next three years The Reuters Institute’s January 2026 survey of 280 media leaders across 51 countries, with the Chartbeat traffic decline data.
ChatGPT traffic analysis: insights from 17 months of clickstream data Semrush’s clickstream study of ChatGPT query and referral behaviour.
Google AI Overviews get Gemini 3 upgrade Reporting on Gemini 3 becoming the default model behind AI Overviews in January 2026.
Under the hood: Universal Commerce Protocol Google’s technical documentation of its open protocol for agent-mediated commerce.
New tech and tools for retailers to succeed in an agentic shopping era Google’s announcement of agentic commerce capabilities and the merchant tooling that supports them.
Does schema markup help LLMs? What the evidence actually shows A review of the first-party statements and independent testing on structured data and LLM citation, including the absence of peer-reviewed evidence and the unsourced statistics circulating in the industry.
Hidden prompt injection: the black hat trick AI outgrew Analysis of hidden-instruction tactics in web content, why they stopped working, and how they are treated under existing search policy.
Overview of the General-Purpose AI Code of Practice Reference material on the EU AI Act’s general-purpose AI obligations, the copyright policy requirement and the training-content summary requirement.
The EU AI Act’s transparency rules: a practical guide to Article 50 Reference material on disclosure duties for AI-generated content and AI-mediated interactions.
| Citing this article? Brief excerpts are welcome. Please credit Webiano.digital, name the author where stated, and include a link to https://webiano.digital and to this original article. Full or substantial republication requires prior written permission. Read our Copyright and Content Use Policy. |
This article was prepared with the assistance of artificial intelligence tools. The content underwent expert human review, and Webiano Digital & Marketing Agency assumes editorial responsibility for its final version and publication.















