What Proof Should You Demand Before Trusting an Agency's Past Results?

Answer engine optimisation (AEO) and generative engine optimisation (GEO) have become the phrase every agency reaches for right now, whether or not they can actually deliver it. That crowding is exactly why the burden of proof matters here more than in an established discipline: before trusting any AEO result, demand 3 things. A live query dashboard you can re-run yourself, a case study naming the AI engine tested, and a reference client willing to verify the figures. TrustRadius's 2026 report found that 94% of B2B buyers independently verify AI-generated research before acting on it. Apply the same standard to the agency's own claims.
A live query means the agency runs the search prompt in front of you, on the day, in real time. A result that cannot be reproduced today is a claim about what once happened, not proof of what the agency built.
Ask for the specific AI engine. ChatGPT, Perplexity, Gemini, and Google AI Overviews each have different citation behaviours and content source preferences. A case study that does not name the engine cannot be compared across platforms.
Ask for the reference client's name, not a vague vertical description. If the agency will not provide a name until after you sign, ask why. The strength of that evidence should also influence how you choose an AEO agency before adding it to your shortlist.
How Do You Verify an AEO Case Study Is Real, Not Cherry-Picked?

Check whether the proof comes from sources an AI engine actually cites. 84% of AI citations come from earned media, not brand-owned pages, per Muck Rack's Generative Pulse report (May 2026). A case study built only from on-site content misses most of where citation growth should show up. Ask the agency which third-party publications are driving their client's citations.
The on-site trap: many AEO case studies show traffic growth from the client's own website. That proves search engine performance. It does not prove AI engines cite the brand in response to buyer prompts. These are different results measured differently.
Ask specifically: "Which publications did the AI engine cite when mentioning your client?" If the agency cannot name them, the evidence is on-site traffic data only.
A static screenshot proves a result existed at one moment. A live query proves it exists today. That distinction matters when comparing AEO agencies in Australia, particularly where providers rely on case studies to demonstrate citation growth.
Which Metrics Actually Prove Citation Growth, Not Vanity Growth?

A result proves citation growth only when it separates 3 figures: citation rate, mention rate, and share of voice. They move independently. Profound's 11.84-billion-citation study found brand-site citation share ranging from 47% on ChatGPT to 69% on Google Gemini. A blended percentage that combines all 3 hides which engine is performing and which is not.
Citation Rate
Citation rate is the percentage of AI-generated answers that name the brand as a direct source. A high citation rate means the AI engine pulls the brand explicitly when responding to relevant prompts. It is the clearest signal of active recommendation.
Mention Rate
Mention rate is the percentage of AI answers where the brand appears anywhere in the response, including indirect references that do not attribute it as the primary source. High mention rate with low citation rate means the AI recognises the brand but does not lead with it.
Share of Voice
Share of voice measures the brand's citation volume against named competitors on the same prompt set. Without a competitor baseline, a 20% citation rate is not interpretable. Against a competitor at 8%, it is a 2.4x advantage.
The Kynection AEO case study shows what separated tracking looks like in practice: 22.88% AI share of voice, first of 40 tracked competitors in its category and about 3.6x the nearest global rival (measured by our own tracker, 26 August 2026, 8,891 runs across 393 prompts). That separation is only visible when citation rate, mention rate, and share of voice are tracked independently, across a content programme that now runs to 71 deep-dive pieces.
What Should a Reference Call Cover Before You Sign?
A reference call should confirm 3 things: whether the client's dashboard matches what was reported, whether the result held after the campaign ended, and whether the timeline matched. 73% of B2B decision-makers trust peer recommendations above all other sources. A 15-minute call is the cheapest verification available.
According to UserEvidence's 2026 research, 26% of deals fail for customer-evidence reasons: return on investment (ROI) not proven, a lack of references, or no clear differentiation. A reference call addresses all 3 before you commit a budget.
Ask 3 questions on the call:
- "Do the numbers on your own dashboard match what the agency reported in the case study?" If the client cannot access a tracking dashboard, ask how the result was measured.
- "Did the result hold after the campaign push ended?" A result that peaked during heavy content production and then dropped is not a structural citation gain.
- "Did the work complete on the quoted timeline?" Timeline accuracy is a proxy for how well the agency scopes and manages its content programme.
Request 2 reference clients from different verticals where possible. 1 reference does not tell you whether results transfer across industries.
What Red Flags Mean a Case Study Will Not Hold Up?

With so many agencies now claiming GEO or AEO capability overnight, the same handful of red flags tend to separate a real result from a repackaged one. A case study fails scrutiny when it cannot name the AI engine tested, cannot be reproduced on a live query, or credits growth to "AI optimisation" with no detail on what changed. Each pattern points to a number reported once and never verified.
Watch for these 5 patterns:
- No AI engine named: The case study says "AI visibility improved" with no mention of ChatGPT, Perplexity, Gemini, or AI Overviews. Visibility on which platform? Against which prompts?
- Result cannot be reproduced live: If the agency cannot run the search prompt today, the result no longer exists or never existed at scale.
- Growth credited to "AI optimisation" with no mechanism: What content changed? Which URLs were targeted? Which prompts moved? A claim without a mechanism is not evidence.
- On-site traffic data only: Session counts and pageview growth are search engine metrics. They do not measure AI citation rate, mention rate, or share of voice.
- No reference client willing to speak: An agency with 3 or more successful campaigns has at least 1 client willing to verify results on a call. "Clients prefer anonymity" is not a sufficient answer.
Two or more of these patterns together means the case study was assembled after the fact, not tracked in real time. It is the same discipline behind our Content Strategy work: every proof point is tracked from campaign start, not constructed at pitch time.
How Do You Confirm a Result Holds After You Sign?
The verification question does not end at signing. Half of all content cited by AI engines is less than 13 weeks old, according to Profound's research on citation decay (August 2026). A result confirmed at pitch can disappear before renewal. Ask the agency how they track citation rate, mention rate, and share of voice after work begins.
Ask for monthly snapshots from a previous engagement: citation rate by engine, mention rate trend, and share of voice against 3 to 5 named competitors. If the agency cannot produce these from a past client, they have not been tracking results continuously.
A result that grows and holds over 6 months is a structural gain. A result that peaks during heavy content production and then decays is a campaign peak. The difference only becomes clear when the agency shows month-by-month data, not a single end-point snapshot.
Content Creation
Before committing to an AEO programme, establish where your brand stands today. A free AEO audit benchmarks your citation rate, mention rate and share of voice against your top three competitors.
FAQs
How do you tell if an AEO case study is real?
Ask for a live query you can watch in real time, not a screenshot or a PDF. A real result can be reproduced today on the same AI engine. If the agency can only provide static images, ask why the result cannot be shown live. The ability to run the query today is the clearest signal the work is active and the result holds.
What is the difference between citation rate and mention rate?
Citation rate is the percentage of AI-generated answers that name your brand as a direct source. Mention rate is the percentage of answers where your brand appears anywhere in the response, including indirect references. Both figures matter and move at different speeds. An agency reporting 1 blended number has chosen the easiest figure to present, not the most complete picture.
How many reference clients should an AEO agency provide?
Request 2 reference clients from different industries. 1 reference does not tell you whether results transfer across verticals. An agency that has run 3 or more successful campaigns has at least 1 client in a different vertical who will take a call to verify results.
Which AI engines should an AEO case study cover?
At minimum, ChatGPT and Google AI Overviews, which account for the largest share of B2B AI search traffic. A case study covering 1 engine says nothing about performance on the others. Perplexity and Google Gemini are relevant for B2B buyers who use research-mode AI tools.
How long does it take to see measurable AEO results?
Citation rate improvements appear within the first monthly reporting cycle in any active content programme. Share of voice movement against named competitors takes 3 to 6 months and depends on how active the competitor content programme is. An agency without monthly tracking data from a past client has not been measuring results continuously.

