ora research · september 2026 · abstract
In 2023, AI answered from its training data, and hallucinated when it didn't. An industry grew around seeding that knowledge: AEO, with a playbook of Reddit threads, listicles, and citations on other sites. In August 2026 that playbook broke. Reddit's share of ChatGPT citations fell 86% in a week as ChatGPT started querying official sites directly. The buyer's agent now opens your pages, reads them, and verifies before it recommends anything.
We ran ~38,000 of those journeys across multiple harnesses running on claude and gpt models, against 1,000+ real business sites, each answering a real buyer question, split by one axis: is the site ready for agents? Much of what we found was unexpected, starting with what agents do when a site shuts them out.
00 / the shift, in public data
We put our buyer questions to earlier OpenAI generations and measured how much of each answer came from memory rather than retrieval. The last non-reasoning flagship, gpt-4.1, still built about half its answer from memory. Then reasoning and tool use became the default, and it falls off a cliff: down to 7–10% in the models agents run today.
Even the sources it cites are a moving target. Reddit's share of ChatGPT citations fell 86% in a week once the model began running site: searches. Any off-site channel can vanish in one update, so none is worth optimizing for. The page it keeps returning to is your own.
Note: models don't search on everything. Stable facts still come from memory. But pricing, setup, and comparison questions need current information, and in agent harnesses fetching is the default.
the experiment
We ran the study across 1,000+ live business sites. Real buyer questions went to four agent harnesses spanning Claude and GPT models: claude-agent-sdk, claude-code, openclaw, and eve. That comes to ~38,000 agent journeys and 140,000+ page fetches from those sites, each journey run in its own isolated environment. Beforehand we scored, ranked, and tagged every site for agent-readiness and controlled for confounds like brand fame and industry. Full experimental design is in the preprint.
01 / trace replay · one ask, two journeys
Two real traces for the same buyer question, "what can I try for free?", on the same harness. twilio.com answers from its own docs. hashicorp.com blocks the agent twice and the journey ends on two competitors' blogs.
two real traces, claude-agent-sdk · claude-sonnet-4-6 · click a panel to expand
02 / recommendation
We measure how strongly an agent recommends each business. An agent-ready site earns a clear, confident recommendation nearly twice as often as a blocked one.
The lift is steepest exactly where buying decisions are researched: 2.5× in IT infrastructure, 2.3× in sales & marketing, and up to 2.6× depending on which stack the buyer's agent runs on. The flip side holds too: when a site is not agent-ready, answers that both judges rate too weak to recommend at all are 2.5× more common.
| category | ready | not ready | lift |
|---|---|---|---|
| IT Infrastructure | 24% | 9% | 2.5× |
| Sales & Marketing | 21% | 9% | 2.3× |
| Artificial Intelligence | 23% | 10% | 2.2× |
| Development | 32% | 14% | 2.2× |
| stack | ready | not ready | lift |
|---|---|---|---|
| claude-codeclaude-haiku-4-5 | 5% | 2% | 2.6× |
| claude-agent-sdkclaude-sonnet-4-6 | 14% | 5% | 2.6× |
| openclawgpt-5.4-mini | 25% | 14% | 1.8× |
| evegpt-5.4 | 36% | 20% | 1.8× |
The four ways an answer undermines its own recommendation, agent-ready vs not:
03 / grounding
Across the finished answers, training knowledge stays at 7–10% either way, so agents are not falling back on training data. Whatever the site does not provide, web search fills in: on sites that are not agent-ready, almost half the answer (42%) comes from outside them.
The less an agent can read your site, the more it searches the web: 2.6× from the most readable sites to the least. Every search is another chance the agent gets fed information you never published, or stopped updating: other people's pages, stale mirrors, whatever the ranker serves.
Web search is provided by Tavily on openclaw and eve; Claude stacks use built-in search. The ambiguous middle band of scores (0.50–0.65) is excluded by design.
Search reliance is a stack personality. openclaw navigates by search and claude-code barely searches, but losing readability pushes every stack, and every category, the same way. Web searches per run, agent-ready → not:
| stack | ready | not ready | lift |
|---|---|---|---|
| claude-agent-sdkclaude-sonnet-4-6 | 0.4 | 0.9 | 2.4× |
| claude-codeclaude-haiku-4-5 | 0.1 | 0.3 | 3.5× |
| openclawgpt-5.4-mini | 6.9 | 10.4 | 1.5× |
| evegpt-5.4 | 1.2 | 1.9 | 1.6× |
| category | ready | not ready | lift |
|---|---|---|---|
| Consumer | 1.9 | 3.7 | 2.0× |
| IT Infrastructure | 2.0 | 3.9 | 1.9× |
| Development | 1.7 | 3.2 | 1.8× |
| Artificial Intelligence | 1.8 | 3.1 | 1.7× |
| Sales & Marketing | 2.2 | 3.6 | 1.6× |
04 / cost & performance
On sites that are not agent-ready, runs take 23% more turns (5.5 → 6.8), get blocked by the site 2.1× as often, and take 15% longer to reach an answer (48s → 55s). The extra turns and retries are billed to whoever runs the agent.
Cost per run understates it. A 403 is cheap; the expensive part is that only 56% of runs on not-agent-ready sites end grounded in the business, against 78% on agent-ready sites. Spread the cost of the failed runs over the answers that did use your site and the price is +64% per grounded answer, averaged across all harnesses using claude and gpt.
05 / accuracy
Every answer is graded fact by fact against ground truth captured from each business's own site: 30,633 asked facts across 2,426 answers and a sampled 131 of the 1,056 businesses. Wrong facts are rare in both conditions. The failure that grows when the agent can't read the site is an answer that contains none of the facts the user asked for. It is fluent and confident, so hallucination metrics and AEO dashboards do not flag it.
Answers built from the business's own pages are 1.4× as accurate as answers built from web search and write-ups on other sites, and the gap is widest (+64%) on pricing questions, where secondhand information ages fastest. Grouping answers by how much of their evidence came from the company's own pages, accuracy climbs from 41% to 56%.
Prices have one fresh source, your own pages, so reading them pays. Setup is harder: two-thirds of its facts never surface, and when stated they are equally right from either source. What moves setup is findability: sites where agents could ground setup answers score 31% vs 18%, from both sources alike.
06 / what to do about it
AEO optimized what models knew about you from training data. AX means making your site readable to the agent that is already on it. Serve the answer in the initial HTML, don't block the agents your buyers send, and put pricing, trials, and setup on canonical pages an agent can fetch. Everything this study measured moves together once agents can read you: recommendation, grounding, cost, and accuracy.
We think the arbitrage is closing. Tactics that worked on the model's training knowledge (seeded threads, listicles on other sites, citation placement) lose leverage each time retrieval moves closer to the source, and the Reddit data in section 00 is what that looks like in practice. What is left is the basics: SEO so agents can find you, AX so agents can read you.
07 / conclusions
Five findings hold across the four Claude and ChatGPT harnesses and every business category we tested. Together they describe a market where training data no longer decides who gets recommended.
The full report publishes the traces behind every number.
08 / sources
ora research (2026). AX is the new AEO. ora.ai/research, september 2026. arXiv preprint forthcoming.