ora research · september 2026 · abstract

AX is the new AEO

In 2023, AI answered from its training data, and hallucinated when it didn't. An industry grew around seeding that knowledge: AEO, with a playbook of Reddit threads, listicles, and citations on other sites. In August 2026 that playbook broke. Reddit's share of ChatGPT citations fell 86% in a week as ChatGPT started querying official sites directly. The buyer's agent now opens your pages, reads them, and verifies before it recommends anything.

We ran ~38,000 of those journeys across multiple harnesses running on claude and gpt models, against 1,000+ real business sites, each answering a real buyer question, split by one axis: is the site ready for agents? Much of what we found was unexpected, starting with what agents do when a site shuts them out.

Agent-ready businesses get recommended up to 2.6× more often

share of answers both blind judges rate a clear endorsement · by agent stack
agent-readynot agent-ready
20%11% 1.9× 5%2% 2.6× 14%5% 2.6× 25%14% 1.8× 36%20% 1.8×
all stacks37,927 runs pooled
claude-codeclaude-haiku-4-5
claude-agent-sdkclaude-sonnet-4-6
openclawgpt-5.4-mini
evegpt-5.4
37,927 agent runs · 1,056 businesses · september 2026ora.ai/research

00 / the shift, in public data

Agents rely far less on training data

We put our buyer questions to earlier OpenAI generations and measured how much of each answer came from memory rather than retrieval. The last non-reasoning flagship, gpt-4.1, still built about half its answer from memory. Then reasoning and tool use became the default, and it falls off a cliff: down to 7–10% in the models agents run today.

How much of the answer comes from the model's training knowledge

same buyer questions, each OpenAI generation from gpt-4.1 onward
last non-reasoning model 52% 37% 27% 14% reasoning models latest generation this study 7–10% 2025 Q2 Q3 Q4 2026 Q2 Q3 Q4 52% → 7% from gpt-4.1 to today's agents
630 runs, 90 buyer questions per OpenAI generation, web tools optional, every claim scored by an independent judge, points are release half-year averagesora.ai/research

Three public signals show the same turn toward grounding

claude searches when the answer needs fresh data
9 in 10
recency-dependent questions trigger a live web search
"Look It Up", arXiv 2026 · source
chatgpt searches anchored to the current year
6%
GPT-5.2
53%
GPT-5.3
87%
GPT-5.5
TUM & maestra.ai · source
reddit citations fall as chatgpt runs site: searches
3.8% 0.5% jul 7 aug 17
Promptwatch & Qwairy · source

Even the sources it cites are a moving target. Reddit's share of ChatGPT citations fell 86% in a week once the model began running site: searches. Any off-site channel can vanish in one update, so none is worth optimizing for. The page it keeps returning to is your own.

Note: models don't search on everything. Stable facts still come from memory. But pricing, setup, and comparison questions need current information, and in agent harnesses fetching is the default.

the experiment

How we designed the research

We ran the study across 1,000+ live business sites. Real buyer questions went to four agent harnesses spanning Claude and GPT models: claude-agent-sdk, claude-code, openclaw, and eve. That comes to ~38,000 agent journeys and 140,000+ page fetches from those sites, each journey run in its own isolated environment. Beforehand we scored, ranked, and tagged every site for agent-readiness and controlled for confounds like brand fame and industry. Full experimental design is in the preprint.

01 / trace replay · one ask, two journeys

Where the agent actually goes

Two real traces for the same buyer question, "what can I try for free?", on the same harness. twilio.com answers from its own docs. hashicorp.com blocks the agent twice and the journey ends on two competitors' blogs.

the site's own page errored / dead end web search another site

two real traces, claude-agent-sdk · claude-sonnet-4-6 · click a panel to expand

02 / recommendation

Agents recommend you twice as often when they can read you

We measure how strongly an agent recommends each business. An agent-ready site earns a clear, confident recommendation nearly twice as often as a blocked one.

One in five answers clearly recommends an agent-ready business. One in nine when it isn't.

both blind judges rate the answer a clear endorsement · all stacks pooled
1.9×
agent-ready
20%
not agent-ready
11%
independent recommendation: how strong a case the agent could build · two blind judges (claude and gpt) must both give the top grade · they agree within a point on 88%ora.ai/research

The lift is steepest exactly where buying decisions are researched: 2.5× in IT infrastructure, 2.3× in sales & marketing, and up to 2.6× depending on which stack the buyer's agent runs on. The flip side holds too: when a site is not agent-ready, answers that both judges rate too weak to recommend at all are 2.5× more common.

breakdown · by category, by stack, and how the hedge looks
steepest categories · endorsement rate, both judges
categoryreadynot readylift
IT Infrastructure24%9%2.5×
Sales & Marketing21%9%2.3×
Artificial Intelligence23%10%2.2×
Development32%14%2.2×
by harness
stackreadynot readylift
claude-codeclaude-haiku-4-55%2%2.6×
claude-agent-sdkclaude-sonnet-4-614%5%2.6×
openclawgpt-5.4-mini25%14%1.8×
evegpt-5.436%20%1.8×

What the hedge looks like

The four ways an answer undermines its own recommendation, agent-ready vs not:

admits it could not find or access the information"the site returned 403, so…"
4%
16%
4.4×
vouches from secondhand sourcescites aggregators instead of the business
5%
15%
3.0×
vague, noncommittal endorsementno concrete prices, steps, or links
7%
12%
1.8×
punts the user to go check themselves"contact sales for a quote"
15%
22%
1.4×
agent-readynot agent-ready

03 / grounding

When agents can't read you, the answer is built without you

Across the finished answers, training knowledge stays at 7–10% either way, so agents are not falling back on training data. Whatever the site does not provide, web search fills in: on sites that are not agent-ready, almost half the answer (42%) comes from outside them.

When the agent can't read you, a quarter of the answer comes from web search

what the finished answer is built from · training knowledge stays at 7–10% either way
your pagesexternalweb searchtraining knowledge
agent-ready78% your pages · 3% external · 12% web search · 7% training knowledge
not agent-ready58% your pages · 7% external · 25% web search · 10% training knowledge
share of the finished answer, by evidence sourceora.ai/research
In ~99% of the journeys that hit a dead end on a site, the agent answered anyway, built from whatever it found elsewhere.

Web search rises as readability falls

The less an agent can read your site, the more it searches the web: 2.6× from the most readable sites to the least. Every search is another chance the agent gets fed information you never published, or stopped updating: other people's pages, stale mirrors, whatever the ranker serves.

Web search is provided by Tavily on openclaw and eve; Claude stacks use built-in search. The ambiguous middle band of scores (0.50–0.65) is excluded by design.

Web searches per run, by readability

average per agent run · 1,056 domains
0.20.40.60.8 4.54.03.13.02.72.32.12.01.8 2.6× more web searches
← not agent-readyora accessibility scoreagent-ready →
ora.ai/research
breakdown · web searches by harness and domain category

Search reliance is a stack personality. openclaw navigates by search and claude-code barely searches, but losing readability pushes every stack, and every category, the same way. Web searches per run, agent-ready → not:

by harness
stackreadynot readylift
claude-agent-sdkclaude-sonnet-4-60.40.92.4×
claude-codeclaude-haiku-4-50.10.33.5×
openclawgpt-5.4-mini6.910.41.5×
evegpt-5.41.21.91.6×
by domain category · steepest 5
categoryreadynot readylift
Consumer1.93.72.0×
IT Infrastructure2.03.91.9×
Development1.73.21.8×
Artificial Intelligence1.83.11.7×
Sales & Marketing2.23.61.6×

04 / cost & performance

Blocked sites cost the agent's operator more per answer

On sites that are not agent-ready, runs take 23% more turns (5.5 → 6.8), get blocked by the site 2.1× as often, and take 15% longer to reach an answer (48s → 55s). The extra turns and retries are billed to whoever runs the agent.

Cost per run understates it. A 403 is cheap; the expensive part is that only 56% of runs on not-agent-ready sites end grounded in the business, against 78% on agent-ready sites. Spread the cost of the failed runs over the answers that did use your site and the price is +64% per grounded answer, averaged across all harnesses using claude and gpt.

Every grounded answer costs the agent's operator up to 93% more when the site isn't agent-ready

cost per answer grounded in the business · agent-ready → not agent-ready · per stack
openclawgpt-5.4-mini
+93%
$0.135 → $0.260
claude-agent-sdkclaude-sonnet-4-6
+93%
$0.112 → $0.216
claude-codeclaude-haiku-4-5
+57%
$0.020 → $0.031
evegpt-5.4
+11%
$0.091 → $0.101
+64% averaged across stacksora.ai/research

05 / accuracy

The common failure is a missing fact

Every answer is graded fact by fact against ground truth captured from each business's own site: 30,633 asked facts across 2,426 answers and a sampled 131 of the 1,056 businesses. Wrong facts are rare in both conditions. The failure that grows when the agent can't read the site is an answer that contains none of the facts the user asked for. It is fluent and confident, so hallucination metrics and AEO dashboards do not flag it.

Answers built off-site are 3.7× more likely to contain none of the facts the buyer asked for

read from the site7%
from the web / other sites25%
share of answers with none of the facts the buyer asked for · 3.7× higher off-site · paired within domain, harness, and question type
2,426 graded answers · 131 businessesora.ai/research

Wrong facts stay rare either way. Missing facts grow from 29% to 45%.

the fate of every asked fact · by where the evidence came from
stated righthalf rightstated wrongnever mentioned
read from the site40% stated right · 27% half right · 4% wrong · 29% never mentioned
from the web / other sites28% stated right · 21% half right · 6% wrong · 45% never mentioned
30,633 asked facts graded against each business's own siteora.ai/research

The source predicts the accuracy

Answers built from the business's own pages are 1.4× as accurate as answers built from web search and write-ups on other sites, and the gap is widest (+64%) on pricing questions, where secondhand information ages fastest. Grouping answers by how much of their evidence came from the company's own pages, accuracy climbs from 41% to 56%.

Accuracy climbs +37% as the evidence comes from your own site

by share of evidence read from your site
20% 40% 60% 80% 41%43%47%50%56%
← noneevidence read from your siteall →
2,426 answersora.ai/research
breakdown · accuracy by question type, empty answers by stack
accuracy by question type · built from the site vs from the web
pricing
60%
37%
+64%
features
46%
39%
+18%
setup
23.9%
23.5%
+2%
empty answers by stack · built from the site vs from the web
claude-codeclaude-haiku-4-5
15%
70%
4.8×
claude-agent-sdkclaude-sonnet-4-6
11%
21%
1.9×
openclawgpt-5.4-mini
7%
9%
1.2×
evegpt-5.4
7%
7%
1.0×
read from the sitefrom the web or other sites

Pricing pays. Setup never surfaces.

Prices have one fresh source, your own pages, so reading them pays. Setup is harder: two-thirds of its facts never surface, and when stated they are equally right from either source. What moves setup is findability: sites where agents could ground setup answers score 31% vs 18%, from both sources alike.

06 / what to do about it

The fix is a site agents can read.

AEO optimized what models knew about you from training data. AX means making your site readable to the agent that is already on it. Serve the answer in the initial HTML, don't block the agents your buyers send, and put pricing, trials, and setup on canonical pages an agent can fetch. Everything this study measured moves together once agents can read you: recommendation, grounding, cost, and accuracy.

We think the arbitrage is closing. Tactics that worked on the model's training knowledge (seeded threads, listicles on other sites, citation placement) lose leverage each time retrieval moves closer to the source, and the Reddit data in section 00 is what that looks like in practice. What is left is the basics: SEO so agents can find you, AX so agents can read you.

07 / conclusions

What we conclude

Five findings hold across the four Claude and ChatGPT harnesses and every business category we tested. Together they describe a market where training data no longer decides who gets recommended.

  1. 01
    Training knowledge is no longer where the answer comes from.
    Agents now ground answers in the live web: frontier models search about nine times out of ten when a fact is fresh, and in our traces training knowledge supplies just 7 to 10% of the finished answer, readable site or not. Optimizing what the model knows about you optimizes a shrinking slice.
  2. 02
    Readability is the largest single lever we measured.
    Agent-ready businesses are recommended 1.9× more often on average and up to 2.6× depending on the harness. When the agent can read the site, 78% of the finished answer is built from the business's own pages, against 58% when it cannot.
  3. 03
    When the agent can't read you, it builds the answer without you.
    Blocked agents do not give up. They go to web search, whose share of the answer doubles from 12% to 25%. The answer built off-site is 3.7× more likely to contain none of the facts the buyer asked for, yet it reads fluent and confident, so nothing flags it. And accuracy follows the source: an answer grounded in your own pages is 1.4× as accurate.
  4. 04
    Blocking agents is paid for by the agent's operator, then by you.
    On not-agent-ready sites, runs take 23% more turns, hit a block 2.1× as often, and each grounded answer costs the operator up to 93% more. Every harness is under pressure to route around sites that behave this way.
  5. 05
    The AEO arbitrage is closing.
    Tactics aimed at the model's training knowledge (seeded threads, listicles on other sites, citation placement) lose leverage each time retrieval moves closer to the source. In August 2026, Reddit's share of ChatGPT citations fell 86% in a week. What remains is the basics: SEO so agents can find you, AX so agents can read you.

The full report publishes the traces behind every number.

08 / sources

01
ora research, september 2026 · arXiv preprint
02
open specification and data · agentready.org · the practices ranked by measured agent behavior
03
the dataset behind this report: a stratified 10% domain sample of the 37,927 journeys (106 domains, 3,816 journeys) as flat CSVs, with the script that recomputes the aggregate numbers and reports the delta against the published ones
04
the agent-readiness ranker that splits the 1,056 businesses in this study · 100+ checks, per-layer breakdown
05
TUM and maestra.ai, 2026 · 1,650 runs, 370k search results · share of Business-tier runs whose searches are anchored to the current year: GPT-5.2 ~6%, 5.3 53%, 5.5 87% (Plus-tier runs 60% / 79%)
06
Sahil Kale, arXiv:2511.18931, 2026 · on recency-dependent queries, search-invocation rates of 91.0% (Claude Haiku 4.5) and 93.9% (Claude Sonnet 4.6), vs 87.5–92.1% for GPT-5-mini / GPT-5; end-to-end accuracy still caps near 66–71%
07
Promptwatch, august 2026 · Reddit's share of ChatGPT Search citations, 3.83% (jul 18 to aug 7) to 0.52% (aug 14 to 17) · labeled provisional by the authors; trailing share has since recovered to about 1.5%
08
Qwairy, august 2026 · site: queries rose from near 0% to 23–24% of ChatGPT's background searches from august 8, redirecting citations from discussion platforms to official sites
cite this abstract

ora research (2026). AX is the new AEO. ora.ai/research, september 2026. arXiv preprint forthcoming.