HomeBlog

Two AEO Red Flags With Real Evidence Behind Them

Two AEO proposal claims fail against real published data: guaranteed citation counts, and schema sold as a citation lever. Here is what the studies found.

Last reviewed:
September 23, 2026
· Reviewed quarterly for accuracy
Two AEO Red Flags With Real Evidence Behind Them
Key Facts

Two AEO red flags have real evidence behind them, not opinion: guaranteed citation counts, and schema sold as the core method. SparkToro and Gumshoe.ai found under a 1-in-100 chance of the same AI brand list twice across 2,961 runs, which rules out any fixed guarantee. A 2026 Ahrefs study of 1,885 pages found schema alone produced no real citation lift.

TL;DR
  • The variability is measured, not assumed. SparkToro and Gumshoe.ai found under a 1% chance of an identical brand list across repeat runs of the same prompt.
  • Schema's lift is measured too, and it is zero. Ahrefs tested 1,885 pages against 4,000 matched controls and found no real citation gain.
  • A guarantee is not confidence, it is a tell. Both studies agree: fixed promises do not match how these systems actually behave.
  • Everything else sold as a red flag is judgment, not measurement, and belongs in a different piece.
Decision Matrix
Claim in a proposalWhat the evidence actually showsWhat a credible answer sounds like
"We guarantee X citations by Y date"The same brand list showed up less than once per 100 repeat runs of one prompt (SparkToro and Gumshoe.ai)"We track citation share as a trend across repeat runs, because the output changes run to run"
"Schema and llms.txt work is our core AEO method"Adding schema to 1,885 pages produced no measurable citation gain against 4,000 matched controls (Ahrefs)"Schema is a basic requirement, alongside content structure and clear entities, not a lever on its own"
Will not commit to a fixed number or date (Steelman)The honest reading of both studies, not evasion"We will give you a range and a re-check cadence, and explain why one number would mislead you"
The Verdict

Two specific AEO claims have real, controlled evidence behind them: that no agency can guarantee a fixed citation count, and that schema alone does not move citations. Everything else marketed as a red flag is judgment, not measurement. When a proposal contradicts either of these two studies, that is not a stylistic quibble, it is a factual error.

Why Can't Any Agency Guarantee a Citation Count?

The odds of the same prompt returning an identical brand list, and the 3 things that varied every run
No fixed list means nothing fixed to guarantee.

Because the system behind answer engine optimisation (AEO) does not produce a fixed answer to guarantee. SparkToro and Gumshoe.ai's research had 600 volunteers run 12 different prompts through 3 tools a combined 2,961 times, over November and December 2025. Rand Fishkin and Patrick O'Donnell tested prompts requesting brand recommendations, each run 60 to 100 times, across ChatGPT, Claude and Google's AI Overview, with AI Mode standing in when AI Overviews did not show.

Three things varied on every single run:

  • Which brands appeared in the answer at all
  • How many brands were listed
  • The order they came in

The result: under a 1-in-100 chance that ChatGPT or Google's AI would return the identical brand list twice for the same prompt. The same list in the same order showed up closer to 1 in 1,000 times.

An AI answer is generated fresh each time from a probability distribution, not retrieved from a stored ranking. There is no fixed list for an agency to guarantee.

What Did the SparkToro Study Actually Measure?

DetailWhat it was
Run byRand Fishkin (SparkToro) and Patrick O'Donnell (Gumshoe.ai)
Sample600 volunteers, 12 prompts, a combined 2,961 runs across November and December 2025
Platforms testedChatGPT, Claude, Google's AI Overview, with AI Mode when Overviews did not show
MethodPrompts requesting brand recommendations, each run 60 to 100 times, responses normalised into ordered brand lists
Headline resultUnder a 1-in-100 chance of an identical brand list across repeat runs of the same prompt
Evidence strengthField study, methodology and data published openly. Discloses a commercial link to an AI-tracking vendor. Not peer-reviewed

Ranking position is close to meaningless here. Running a prompt repeatedly does tell you something a single run cannot: which recommendations are more or less likely to appear at all. The authors are careful about how far that goes. They note that how often a brand shows up in a topic may have less to do with its prominence than with how many options the engine had to choose from, and that prompt wording is a further problem, since across 142 volunteer-written prompts barely 2 looked alike.

So appearance rate over many runs is the more useful number, not a clean one. Our own generative engine optimisation work is built around that distinction, tracking share of voice over repeat runs rather than promising a single number.

Does Schema Actually Earn AI Citations?

The raw schema correlation beside the controlled result, with the citation effect on each platform
The correlation vanishes once the test controls for it.

The correlation looks real until someone controls for it. In an analysis of 6 million URLs, pages that got cited were almost 3 times more likely to carry schema in the JavaScript Object Notation for Linked Data (JSON-LD) format, which is the number most pitches use first. Ahrefs ran the controlled version instead: schema added to 1,885 pages, compared against 4,000 matched pages that did not get it.

PlatformCitation effect after adding schema
ChatGPT+2.2%, within noise
Google AI Mode+2.4%, within noise
Google AI Overviews-4.6%, a small but real decline

Four separate tests on the same dataset all found the same result, which rules out a one-off. What the raw correlation was actually measuring was the kind of team that adds schema in the first place: one that already writes clearer, better-structured content. Schema was a marker of that team, not the cause of the citation.

Google's own guidance on AI features supports the same reading. It states plainly that you do not need to create new machine-readable files, AI text files or markup to appear in these features, and that there is no special schema.org structured data you need to add.

What Did the Ahrefs Study Actually Measure?

DetailWhat it was
Run byAhrefs
Sample1,885 pages that added JSON-LD schema, August 2025 to March 2026
Control group4,000 matched pages that never added schema, 3 per treated page, from different domains with similar prior citation levels
MethodMatched difference-in-differences analysis, run as 4 separate tests
Headline resultNo meaningful lift on ChatGPT or Google AI Mode; a small decline on Google AI Overviews
Evidence strengthControlled test with matched controls, run as 4 separate checks that agreed. Independent of any vendor with a stake in the result

None of this makes schema worthless. It still helps machines parse a page correctly, but that does not mean it earns more citations on its own. The claim it disproves is narrower and more specific: that schema by itself is a citation lever an agency can sell as its core method. If a proposal relies mostly on schema, check what else it commits to in writing before you sign, covered in our AEO agency contracts checklist.

Where Do Both Studies Fall Short?

What each study rules out and where each one stops, side by side
Both findings hold. Neither is the final word.

Neither study is the final word, and a credible proposal should be able to say so openly.

  • SparkToro and Gumshoe.ai. The authors themselves called for larger follow-up work, disclosed a commercial link to an AI-tracking vendor, and note this is not peer-reviewed research. They also flag that prompt wording varies so widely between real buyers that any tracked prompt set is only a sample of what people actually ask.
  • Ahrefs. The sample was pages that already had a meaningful citation baseline before schema was added. It does not test whether schema helps a page that is not yet cited at all, which is a different question.

Neither gap changes the headline finding. It changes how much you can conclude from it. Use the results to rule out guarantees and oversold schema pitches, not to claim either study proves what it does not.

How Do You Check a Proposal Against This Live?

Three claims you might hear in a proposal and the question to put back each time
Swap the pitch for the question, on the call.
Phrase you hearWhat it is dodgingAsk this instead
"We guarantee X citations by [date]"The SparkToro and Gumshoe.ai finding that no fixed list is possible"What is the range across your last 10 client reports?"
"Our schema strategy is the core of our AEO method"The Ahrefs finding that schema alone shows no lift"What else are you doing besides markup?"
"We have cracked the algorithm"There is no single algorithm to crack"Which engine, specifically, and how does your method differ per engine?"

These 2 are not the whole evaluation. How to choose an AEO agency covers the fuller checklist, including reporting, diagnosis and references. How to vet an AEO partner covers what a dodge on those broader questions actually sounds like on a live call.

What this piece adds is narrower and more important than a longer list: 2 claims that are not a matter of opinion, because the data to check them already exists and is public. A proposal that contradicts either one is not showing bad judgment. It is stating something the evidence has already ruled out.

Content Creation

Check a proposal against the evidence, not the pitch

Get a baseline for citation rate, mention rate and share of voice, so you can judge what any proposal is actually promising to move.

Frequently Asked Questions

FAQs

Does the SparkToro finding mean AI visibility cannot be tracked at all?

No, but it is messier than most tools admit. Appearance rate across many runs tells you more than rank or an exact brand list does. The authors also warn that appearance rate is shaped by how many options the engine had, and that real buyer prompts vary far more than any tracked set.

Does the Ahrefs finding mean schema should be skipped entirely?

No. Schema still helps a machine parse a page correctly. It only disproves the claim that schema alone drives citation, which is a different thing from being worthless.

Why did the schema correlation look real in the first place?

Because teams that add schema tend to also write clearer, better-structured content. The controlled test isolates schema from those other habits, and the citation effect disappears.

Is a study of 2,961 runs really enough to call this settled?

Strong enough to rule out fixed guarantees, not strong enough to close the question completely. Treat it as the best evidence available, not the final word.

What should a credible AEO proposal say instead of a guarantee?

A range, a re-check cadence, and a plain explanation of why a single fixed number would mislead you. That answer is the honest one, not the evasive one.

SHARE