Why GEO and ‘AI SEO’ Don’t Work as a Separate Discipline
Generative engine optimization (GEO) is the idea that you can optimize content for AI answer systems like ChatGPT, Google AI Overviews, and Perplexity using tactics that reach beyond classical SEO. After reading the GEO papers and the industry citation data, I think the idea falls apart. Once you subtract classical SEO, content quality, PR, YouTube, and brand-building, almost nothing is left for a GEO vendor to sell you.
The labels stack up fast. GEO, AEO (answer engine optimization), and “AI SEO” all point at the same promise: a separate optimization layer for generative answers. This piece walks through how AI answers actually get built, what Google says about optimizing for them, what the research shows, and the few inputs that move the needle.
Download the report: PDF · LaTeX source
How Do AI Search Answers Actually Work?
AI answer surfaces work differently from a ranked results page. With Google’s classic list, you adjust a page and watch it move up or down, so study and iteration pay off. A generative answer comes out of a different pipeline. Retrieval runs first: the system fans the query out into sub-queries, runs them against a search index, and feeds the retrieved sources to the model as context. Then the model generates text, conditioned on those sources plus its system instructions and safety filters, and stochastic (partly random) decoding produces the actual tokens.
Google’s documentation on AI features says AI Overviews and AI Mode “may use different models and techniques” and that responses and links “will vary.” OpenAI’s API docs describe chat completions as non-deterministic by default. The global swap of AI Overviews to Gemini 3 on January 27, 2026 made the stakes concrete. SE Ranking measured roughly 42% of previously cited domains replaced after the swap, with 32% more sources cited per response. SE Ranking, “Gemini 3 impact on AI Overviews: Nearly half of cited domains changed, 32% more sources per answer and sourceless bug fixed,” February 27, 2026. 100,000-keyword study; the 42.4% domain-replacement figure is from the post-bug-fix dataset. The churn concentrated in the long tail, where the most-cited domains held their spots: among the top 500, only one dropped out. Even so, a page-level change can do nothing about an event at the model level. And the swaps keep coming. In May 2026, Gemini 3.5 Flash became the default model in AI Mode worldwide, and by the end of that month Gemini powered the default backend across Google Search. Google, “Google AI announcements from May 2026” (blog.google), and Google I/O 2026 Search updates. Gemini 3.5 Flash reached general availability and became the default model in AI Mode, which passed one billion monthly active users in May 2026; by the end of the month, Gemini models were the default backend across Google Search. The May 21, 2026 broad core update rolled out separately in the same window. Google re-rolls the selection logic on its own release cadence.
The Gemini 3 Swap, January 27, 2026
One model swap replaced 42% of cited domains overnight, almost all of them in the long tail. A page-level change can do nothing about an event at the model level.
Where Do AI Models Get Their Information?
Two different things feed an AI answer, and people mix them up. Training data shaped the model’s weights months before it deployed. At inference, those weights are all that runs. The system has no step where it pulls your page out of the training data, even if it crawled your page at the right moment. When a chatbot answer feels out of date, retrieval usually skipped that query, so the model answered from weights with a training cutoff. Forcing retrieval fixes that: use Perplexity, turn on AI Mode, switch browsing on, or ask the model to search. Training updates can substitute for none of this, because training happens in batches and inference queries the weights, never the training set.
What Does Google Say About GEO and AEO?
Google has been clear about what works. Its Search Central guidance on AI features still says there are “no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary.” The only changes since that doc first appeared are measurement disclosures: a February 2026 note confirming AI Overview clicks and impressions count in Search Console, and a dedicated Search Generative AI performance report that began rolling out in Search Console on June 3, 2026. Both describe how Google reports the data. Google asks for no special schema, markup, or AI-specific text file. The advice is ordinary SEO.
On May 15, 2026, Google went further and published a dedicated guide, Optimizing your website for generative AI features on Google Search. Its governing line: “From Google Search’s perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.” The guide names AEO and GEO directly, describes the same retrieval and query fan-out architecture covered here, and tells site owners to skip llms.txt files, content chunking, copy rewritten for AI systems, structured-data overfocus, and inauthentic mentions. Every tactic the GEO vendors sell turns up on Google’s list of things to ignore.
On June 5, 2026, Google added guidance on third-party SEO tools, services, and advice and revised its hiring guide. It states that third-party tools “don’t have access to our internal ranking data” and “can’t guarantee performance,” and it lists tools that promise “improvements for AI experiences and search formats (also known as ‘AEO’ or ‘GEO’ tools)” among the claims to scrutinize. Google moved from calling this work SEO to naming the vendor category and asking site owners to check whether the pitch matches its published guidance.
One area does point past classical SEO, and Google flagged it too: agentic experiences and agent-readable site structure. At Google I/O 2026 and Marketing Live, the Universal Commerce Protocol shipped a Universal Cart with Google Pay checkout across Search, Gemini, and Maps, backed by more than 20 partners including Shopify, Etsy, Target, and Walmart, with expansion underway to YouTube and to verticals like hotels and food delivery. This is the real frontier: a new agent-readable surface with its own protocol. GEO vendors aren’t selling it, and the work is young enough that I’d watch it rather than bet on it.
Does the GEO Research Actually Hold Up?
The academic case for GEO is thinner than its popularity suggests. The founding paper, Aggarwal et al. at KDD 2024, built what the authors call a “generative engine” that retrieved the top five Google results and generated an answer with GPT-3.5-turbo. A page outside those five results stayed invisible to everything the paper tested. Deployed systems now use frontier models, embedding-based retrieval, and query fan-out, so the distance between that experiment and production runs wide.
The follow-up literature mostly concedes the gap. AutoGEO (October 2025) describes earlier work as “manually designed heuristics.” AutoGEO, arXiv:2510.11438, October 2025. The “manually designed heuristics” phrase appears in the related-work section characterizing prior GEO approaches. E-GEO, from researchers at Columbia and MIT (November 2025), calls the GEO literature “nascent, with many open questions regarding effective strategies and their practical impacts.” An April 2026 arXiv preprint, “Beyond Retrieval,” argues the whole retrieval-centric GEO approach is architecturally fragile, though three of its authors work for the AI vendor whose product serves as the case study, so I’d read it as a sharp critique with a sales angle attached. Zhao, Li, Meng, Zhang, Liu, “Beyond Retrieval: Modeling Confidence Decay and Deterministic Agentic Platforms in Generative Engine Optimization,” arXiv:2604.03656, April 4, 2026. It’s an arXiv preprint; KDD hasn’t published it. Three of the five authors are affiliated with Yishu Research, an industrial AI vendor, and the empirical case study is Yishu’s own product (EasyNote). Useful directional critique of retrieval-centric GEO, and also a vendor pitch.
The 2026 work that does report wins stays in the same lane. A structural-feature-engineering paper, GEO-SFE, claims a 17.3% citation lift from reorganizing document structure, chunking, and visual emphasis. That describes writing and formatting content well. It’s good practice, and it’s also just editing.
Can You Game an AI Model?
The short answer is no, and the reason sits in how these models work. Most GEO intuition rests on the wrong mental model. A classic ranker is a scoring function. It sorts URLs by signals, so when you push the signals, you climb the list. That lever is mechanical, and it’s real. A generative answer comes from something else. After retrieval pulls candidate sources, the model reads them and judges, passage by passage, whether each one answers the question. It behaves more like a careful researcher reading the pages Google returned and deciding which ones are responsive. The unit of selection is “does this source help me answer the user,” judged semantically. That difference is why the discipline fails to carry over.
It explains the missing lever, too. You can no more manipulate a sharp reader’s judgment of whether your page answered the question than you can fool a good editor. The durable way to be judged useful is to be useful. A reasoning evaluator is built to see through copy that performs authority without earning it, and the models get better at this with each release. Treating that judgment as a ranking formula with a knob you can turn is a category error.
Model makers also train directly against this kind of influence. OpenAI’s instruction hierarchy gives untrusted quoted and tool content “no authority” when instructions conflict. Anthropic scans untrusted external content with classifiers and uses reinforcement learning to make models resist injection. Anthropic, public security documentation on prompt-injection mitigation, including Constitutional Classifiers (Sharma et al., arXiv:2501.18837). NVIDIA’s Augmented Intermediate Representations (May 2025) reports a 1.6x to 9.2x drop in attack success against gradient-based attacks, with full robustness to the static attacks it held out. Kariyappa, Suh (NVIDIA), “Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations,” arXiv:2505.18907, May 2025. The 1.6x to 9.2x figure is specifically against gradient-based (GCG with momentum) attacks; against static attacks (Naive, Ignore, Completion, Escape Separation) AIR reaches 0% attack success, including on the two held-out attack types. The arms race runs both ways: backdoor-powered injection work (October 2025) shows defenses can fall to fine-tuning data poisoning, though that attack assumes control of the training data, which a GEO vendor lacks. Chen, Li, Sui, Song, Hooi, “Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods,” arXiv:2510.03705, October 4, 2025. The threat model assumes the attacker can poison supervised fine-tuning data, a real concern for open-weight or downstream-finetuned deployments and a remote one for closed frontier models served behind APIs. The more your optimization tries to influence the model rather than inform the reader, the better the odds it gets filtered.
How Often Do AI Models Hallucinate?
Often enough that “share of AI voice” is a shaky thing to sell. GPT-5.5 launched on April 23, 2026 with coverage touting a “60% reduction in hallucinations,” a figure that never appears in OpenAI’s own system card. The card reports gains that are far smaller: claims that run 23% more likely to be correct, and responses that carry a factual error 3% less often. Independent testing on Artificial Analysis’s AA-Omniscience benchmark put GPT-5.5’s hallucination rate at 86%, the highest of any flagship. Artificial Analysis, “OpenAI’s GPT-5.5 is the new leading AI model,” April 23, 2026, and the OpenAI GPT-5.5 system card (deploymentsafety.openai.com/gpt-5-5). The “60%” figure that circulated on launch day appears nowhere in the system card; the card reports 23% higher claim-level accuracy and a 3% lower response-level error rate, measured on de-identified ChatGPT conversations users had flagged as containing factual errors. The 60% number looks like a loose conflation of older GPT-5-series reductions (roughly 83% for GPT-5 thinking vs o3, 45% for GPT-5 vs GPT-4o, and 33% for GPT-5.4 vs GPT-5.2). AA-Omniscience defines hallucination rate as the share of non-correct responses where the model confabulated rather than abstaining, which differs from the share of all answers that are wrong. The Claude Opus 4.8 figure (35.9%, essentially flat against Opus 4.7’s 36%) is from Artificial Analysis, “Claude Opus 4.8 analysis and benchmarks,” May 28, 2026. Claude Opus 4.8 sits at 36% and Gemini 3.1 Pro at 50%. The Opus figure is the telling one: Opus 4.8 held its rate flat at 35.9% across a full generation, which shows a model maker can keep calibration steady when it chooses to. A separate cross-model audit (February 2026) found reference-fabrication rates from 11.4% to 56.8%, depending on the model and the query. Naser, “How LLMs Cite and Why It Matters: A Cross-Model Audit of Reference Fabrication,” arXiv:2603.03299, February 2026. 10 commercially deployed LLMs, 69,557 citation instances across four academic domains, verified against CrossRef, OpenAlex, and Semantic Scholar. Pick a benchmark and the frontier looks great; pick another and it looks shaky.
Hallucination Rates Across Flagship Models
The “60% reduction” headline for GPT-5.5 didn’t survive independent testing. AA-Omniscience scores the share of non-correct responses where the model confabulated instead of abstaining.
Do AI Citations Track Google Rankings?
Ahrefs found the share of AI Overview citations drawn from top-10 organic positions fell from 76% in July 2025 to 38% by early 2026. BrightEdge, tracking the same question over 16 months, reports a headline that looks opposite: overlap with organic rankings climbing from 32% to 54%. Its own data reconciles the two. Only about 17% of citations come from the top 10, and BrightEdge calls positions 21 through 100 “the sweet spot.” Put the studies together and they agree. Citations increasingly come from pages that rank organically somewhere, with the top 10 doing less of the work. Ranking is still the gate. The promise that a top-10 position buys you the citation is the part that’s breaking.
Profound’s analysis of 240 million ChatGPT citations found 40% to 60% of cited domains change month to month for the same query, and 70% to 90% turn over completely across six months. Profound longitudinal analysis of 240 million ChatGPT citations, 2026. The 70% to 90% six-month turnover figure is from Profound’s own reporting, consolidated in third-party writeups including Machine Relations’ “Citation Drift” analysis. Semrush’s 13-week study watched Reddit’s share of ChatGPT citations fall from about 60% to about 10% in six weeks during August and September 2025. That timing lined up with Google removing its num=100 parameter on September 11, though Semrush’s own analyst credits a ChatGPT-side change instead, a deliberate down-weighting of its most over-cited domains. Semrush 13-week analysis of ~230,000 prompts (weekly snapshots, July to October 2025) across ChatGPT, Google AI Mode, and Perplexity. Reddit’s ChatGPT citation share fell from ~60% to ~10% in six weeks during August and September 2025. The drop lined up with Google removing its num=100 parameter on September 11, but Semrush analyst Sergei Rogulin disputes that as the cause: only ~34% of Reddit’s rankings sit in positions 21 through 100, too few to explain a 50-point fall, so he credits ChatGPT deliberately reducing over-citation of its highest-frequency domains. Perplexity and AI Mode showed no matching drop, consistent with a ChatGPT-specific retrieval change. Corroborated in 5W’s AI Platform Citation Source Index 2026 (May 1, 2026) and ALM Corp’s March 2026 cross-platform analysis.
DeepTRACE (September 2025), the closest successor to Stanford’s 2023 verifiability work, found that current generative search systems “frequently produce one-sided, highly confident responses” with weak attribution. Liu, Zhang, Liang, “Evaluating Verifiability in Generative Search Engines,” EMNLP Findings 2023 (arXiv:2304.09848). 51.5% citation recall and 74.5% citation precision in the deployed systems studied at the time. DeepTRACE (Venkit et al., arXiv:2509.04499, September 2, 2025) extends that work into an eight-dimensional audit and supplies the “one-sided, highly confident” language. Building a KPI on a surface whose selection logic gets rewritten in a single week asks clients to chase a target that keeps moving.
Reddit’s Share of ChatGPT Citations, August to October 2025
Reddit’s share of ChatGPT citations fell roughly 50 points. The two solid points are reported figures; the dashed path between them is interpolated.
What Actually Drives AI Search Visibility?
Two things move with AI visibility, and both sit outside the GEO playbook. The first is YouTube. Ahrefs’s study of 75,000 brands found that mentions on YouTube, in titles, transcripts, and descriptions, correlate with AI Overview visibility more strongly than any other signal it tested. Ahrefs (Linehan), “An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied),” 2025. Branded web mentions correlate with AIO visibility at 0.664 (about three times backlinks at 0.218); YouTube mentions were the single strongest signal among all factors tested. The German health-query finding is from SE Ranking’s January 2026 AIO source analysis, which placed YouTube above official medical bodies on health-information queries. YouTube is now the most-cited domain in AI Overviews overall, up 34% over six months, and SE Ranking found YouTube outranking official medical bodies as a source on German health queries. AI Overviews will cite a YouTube video ahead of a state licensing board.
What Correlates With AI Overview Visibility
Every off-site brand signal outranks every on-site link metric. Branded web mentions correlate with AI Overview visibility about three times as strongly as backlinks.
YouTube mentions, in titles, transcripts, and descriptions, were the single strongest signal among every factor Ahrefs tested. YouTube is now the most-cited domain in AI Overviews overall.
The second is click economics. Pew Research Center analyzed 68,879 queries and found an 8% click rate when an AI summary appeared, against 15% when none did, a 46.7% relative drop. Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results,” July 22, 2025. Browsing data from about 900 U.S. adults, March 2025; 68,879 queries analyzed. 8% click rate with an AI summary present, 15% without, and 1% on links inside the summary. Ahrefs’s December 2025 CTR update separately estimated a 58% drop in position-1 organic click-through with AIO present. Just 1% of users clicked a link inside the summary. Ahrefs put the hit to position-1 organic click-through at 58%. Strong classical SEO keeps you eligible to be cited, and the click loss still lands once an AI summary triggers. Say that to clients plainly before anyone conflates a ranking with traffic.
Where the Clicks Go When an AI Summary Appears
Click-throughs fall 46.7% when an AI summary triggers. The 1% click rate on links inside the summary is the number clients should see before they conflate ranking with traffic.
What Actually Works for GEO?
Do the SEO well, and add the few inputs that help. Crawlability, clear and well-sourced content, and real authority decide whether you win retrieval, because winning retrieval means being the best answer for the query in ordinary SEO terms. Add YouTube where a client can support it. Add earned media: Muck Rack’s “What Is AI Reading?” report, re-run in May 2026, found that 84% of AI citations come from earned media, journalism alone accounts for about 27%, and paid or advertorial content draws 0.3%. That covers most of the playbook.
Here’s the test for any GEO pitch. Ask which specific product surface it optimizes for, which retrieval mechanism that surface uses, and how long the tactic survives the next model update. A vendor who can answer none of those is repackaging work you should already do. Subtract classical SEO, content quality, PR, YouTube, and brand-building from “AI SEO,” and almost nothing remains. Profound raised $96 million at a $1 billion valuation in February 2026, which shows how much capital is chasing measurement for a surface whose logic Google swapped overnight in January and rerolled again in May. Profound press release, Globe Newswire, and Fortune coverage, February 24, 2026. $96 million Series C led by Lightspeed, with participation from Sequoia, Kleiner Perkins, Saga, South Park Commons, and Evantic. Total raised: $155 million. The company is 18 months old.
Frequently Asked Questions About GEO and AI SEO
Is GEO the Same as SEO?
For the tactics that survive contact with a real system, yes. Google’s own guidance says optimizing for generative AI search “is optimizing for the search experience, and thus still SEO.” Crawlability, strong content, and authority win retrieval, and those are classical SEO. The GEO-specific add-ons, from llms.txt to AI-only copy rewrites, are the parts Google tells you to skip.
Does GEO Actually Work?
The evidence is weak. The founding research tested a toy system on the top five Google results with an old model, later papers call the field “nascent,” and the surface that AI answers clearly favor, organic ranking, is something you earn through SEO. Formatting content clearly helps a little, mostly because it helps human readers too.
What’s the Difference Between GEO, AEO, and SEO?
GEO (generative engine optimization) and AEO (answer engine optimization) both describe optimizing for AI-generated answers. SEO covers ranking in search results broadly. In practice the inputs overlap almost completely, and Google treats AEO and GEO as labels for AI-focused SEO rather than separate disciplines.
Should I Buy a GEO Tool or Service?
Ask three questions first: which product surface it optimizes for, which retrieval mechanism that surface uses, and how long the tactic lasts past the next model update. Google’s June 2026 guidance also warns that third-party tools “don’t have access to our internal ranking data” and “can’t guarantee performance.” A measurement dashboard can earn its keep. A tactic that promises to game the model is selling you the wrong thing.
Does llms.txt Help With AI Search?
Probably not. No evidence shows it helps, and Google lists llms.txt among the tactics to skip in its generative-AI guidance. No major AI search product has committed to reading it as a ranking or citation input. Spend that time on content and crawlability.