Why Can't Any Agency Guarantee a Citation Count?

Because the system behind answer engine optimisation (AEO) does not produce a fixed answer to guarantee. SparkToro and Gumshoe.ai's research had 600 volunteers run 12 different prompts through 3 tools a combined 2,961 times, over November and December 2025. Rand Fishkin and Patrick O'Donnell tested prompts requesting brand recommendations, each run 60 to 100 times, across ChatGPT, Claude and Google's AI Overview, with AI Mode standing in when AI Overviews did not show.
Three things varied on every single run:
- Which brands appeared in the answer at all
- How many brands were listed
- The order they came in
The result: under a 1-in-100 chance that ChatGPT or Google's AI would return the identical brand list twice for the same prompt. The same list in the same order showed up closer to 1 in 1,000 times.
An AI answer is generated fresh each time from a probability distribution, not retrieved from a stored ranking. There is no fixed list for an agency to guarantee.
What Did the SparkToro Study Actually Measure?
| Detail | What it was |
|---|---|
| Run by | Rand Fishkin (SparkToro) and Patrick O'Donnell (Gumshoe.ai) |
| Sample | 600 volunteers, 12 prompts, a combined 2,961 runs across November and December 2025 |
| Platforms tested | ChatGPT, Claude, Google's AI Overview, with AI Mode when Overviews did not show |
| Method | Prompts requesting brand recommendations, each run 60 to 100 times, responses normalised into ordered brand lists |
| Headline result | Under a 1-in-100 chance of an identical brand list across repeat runs of the same prompt |
| Evidence strength | Field study, methodology and data published openly. Discloses a commercial link to an AI-tracking vendor. Not peer-reviewed |
Ranking position is close to meaningless here. Running a prompt repeatedly does tell you something a single run cannot: which recommendations are more or less likely to appear at all. The authors are careful about how far that goes. They note that how often a brand shows up in a topic may have less to do with its prominence than with how many options the engine had to choose from, and that prompt wording is a further problem, since across 142 volunteer-written prompts barely 2 looked alike.
So appearance rate over many runs is the more useful number, not a clean one. Our own generative engine optimisation work is built around that distinction, tracking share of voice over repeat runs rather than promising a single number.
Does Schema Actually Earn AI Citations?

The correlation looks real until someone controls for it. In an analysis of 6 million URLs, pages that got cited were almost 3 times more likely to carry schema in the JavaScript Object Notation for Linked Data (JSON-LD) format, which is the number most pitches use first. Ahrefs ran the controlled version instead: schema added to 1,885 pages, compared against 4,000 matched pages that did not get it.
| Platform | Citation effect after adding schema |
|---|---|
| ChatGPT | +2.2%, within noise |
| Google AI Mode | +2.4%, within noise |
| Google AI Overviews | -4.6%, a small but real decline |
Four separate tests on the same dataset all found the same result, which rules out a one-off. What the raw correlation was actually measuring was the kind of team that adds schema in the first place: one that already writes clearer, better-structured content. Schema was a marker of that team, not the cause of the citation.
Google's own guidance on AI features supports the same reading. It states plainly that you do not need to create new machine-readable files, AI text files or markup to appear in these features, and that there is no special schema.org structured data you need to add.
What Did the Ahrefs Study Actually Measure?
| Detail | What it was |
|---|---|
| Run by | Ahrefs |
| Sample | 1,885 pages that added JSON-LD schema, August 2025 to March 2026 |
| Control group | 4,000 matched pages that never added schema, 3 per treated page, from different domains with similar prior citation levels |
| Method | Matched difference-in-differences analysis, run as 4 separate tests |
| Headline result | No meaningful lift on ChatGPT or Google AI Mode; a small decline on Google AI Overviews |
| Evidence strength | Controlled test with matched controls, run as 4 separate checks that agreed. Independent of any vendor with a stake in the result |
None of this makes schema worthless. It still helps machines parse a page correctly, but that does not mean it earns more citations on its own. The claim it disproves is narrower and more specific: that schema by itself is a citation lever an agency can sell as its core method. If a proposal relies mostly on schema, check what else it commits to in writing before you sign, covered in our AEO agency contracts checklist.
Where Do Both Studies Fall Short?

Neither study is the final word, and a credible proposal should be able to say so openly.
- SparkToro and Gumshoe.ai. The authors themselves called for larger follow-up work, disclosed a commercial link to an AI-tracking vendor, and note this is not peer-reviewed research. They also flag that prompt wording varies so widely between real buyers that any tracked prompt set is only a sample of what people actually ask.
- Ahrefs. The sample was pages that already had a meaningful citation baseline before schema was added. It does not test whether schema helps a page that is not yet cited at all, which is a different question.
Neither gap changes the headline finding. It changes how much you can conclude from it. Use the results to rule out guarantees and oversold schema pitches, not to claim either study proves what it does not.
How Do You Check a Proposal Against This Live?

| Phrase you hear | What it is dodging | Ask this instead |
|---|---|---|
| "We guarantee X citations by [date]" | The SparkToro and Gumshoe.ai finding that no fixed list is possible | "What is the range across your last 10 client reports?" |
| "Our schema strategy is the core of our AEO method" | The Ahrefs finding that schema alone shows no lift | "What else are you doing besides markup?" |
| "We have cracked the algorithm" | There is no single algorithm to crack | "Which engine, specifically, and how does your method differ per engine?" |
These 2 are not the whole evaluation. How to choose an AEO agency covers the fuller checklist, including reporting, diagnosis and references. How to vet an AEO partner covers what a dodge on those broader questions actually sounds like on a live call.
What this piece adds is narrower and more important than a longer list: 2 claims that are not a matter of opinion, because the data to check them already exists and is public. A proposal that contradicts either one is not showing bad judgment. It is stating something the evidence has already ruled out.
Content Creation
Get a baseline for citation rate, mention rate and share of voice, so you can judge what any proposal is actually promising to move.
FAQs
Does the SparkToro finding mean AI visibility cannot be tracked at all?
No, but it is messier than most tools admit. Appearance rate across many runs tells you more than rank or an exact brand list does. The authors also warn that appearance rate is shaped by how many options the engine had, and that real buyer prompts vary far more than any tracked set.
Does the Ahrefs finding mean schema should be skipped entirely?
No. Schema still helps a machine parse a page correctly. It only disproves the claim that schema alone drives citation, which is a different thing from being worthless.
Why did the schema correlation look real in the first place?
Because teams that add schema tend to also write clearer, better-structured content. The controlled test isolates schema from those other habits, and the citation effect disappears.
Is a study of 2,961 runs really enough to call this settled?
Strong enough to rule out fixed guarantees, not strong enough to close the question completely. Treat it as the best evidence available, not the final word.
What should a credible AEO proposal say instead of a guarantee?
A range, a re-check cadence, and a plain explanation of why a single fixed number would mislead you. That answer is the honest one, not the evasive one.

