HomeBlog

How to Evaluate an AEO Agency's Past Results

What to check before trusting an AEO agency’s case studies: live dashboards, the real metrics, and the reference questions that expose inflated claims.

Last reviewed:
September 21, 2026
· Reviewed quarterly for accuracy
How to Evaluate an AEO Agency's Past Results
Key Facts

Answer engine optimisation (AEO) is a new enough discipline that almost every agency now claims it can do generative engine optimisation (GEO) or AEO, with no shared standard for what "results" actually means. That makes it more important, not less, to check whether an AEO agency's past results can survive scrutiny. Here are 3 ways to confirm they can deliver. Ask for a live dashboard you can re-run yourself, not a static screenshot. Confirm the case study names the specific AI engines tested. Ask a reference client to verify the numbers by phone. A result that passes all 3 checks is evidence, anything short of that is just marketing copy.

TL;DR
  • Ask for a live dashboard, not a screenshot: A query dashboard you can re-run yourself is proof. A static image is only a claim.
  • Check which AI engines were tested: ChatGPT, Perplexity, Gemini, and Google AI Overviews behave differently. A result on one engine says nothing about the others.
  • Separate citation rate from mention rate and share of voice: A case study that reports 1 blended number picked the easiest metric, not the full picture. Intelligent Resourcing reports all 3 separately on every client dashboard, for this reason.
  • Call the reference client: 26% of business-to-business (B2B) deals fail for customer-evidence reasons (UserEvidence, 2026), and 73% of B2B decision-makers trust peer recommendations above vendor claims (Reddit and SurveyMonkey, 2026). A 15-minute call is the cheapest verification in any agency selection process.
  • Watch for on-site-only proof: 84% of AI citations come from earned media, not brand-owned pages. A case study built from brand content alone does not prove citation growth elsewhere.
Decision Matrix
Proof typeWhat it confirmsWhat it does not confirmWhen this is acceptable
Live query dashboard (re-runnable)Current citation rate by engine, reproducible todayHistorical trend; whether the result held over timeEngine named; result reproduced live during the meeting
Named AI engine breakdownEngine-specific performance on the tested prompt setCross-engine performance; a ChatGPT result does not equal a Gemini resultBrief covers a single AI platform only
Reference client verified by phoneThird-party confirmation of the reported numbersScale beyond 1 client in 1 verticalClient's vertical matches yours exactly
3 separated metrics: citation rate, mention rate, share of voiceFull visibility picture across distinct tracking dimensionsAbsolute market position without a competitor baselineAll 3 figures present alongside a named competitor set
Self-reported trafficOn-site performance targets; direct-response goalsAI citation growth in third-party sourcesBrief explicitly scoped to direct-response only; AI citation visibility not part of the deliverable
The Verdict

Not every agency with real results has a clean dashboard ready to share: pilots, early campaigns, and rapid-growth periods produce proof that is harder to package cleanly. But in a category this new, where almost anyone can put "GEO/AEO" on a services page, the burden sits with the agency to prove it, not with you to take it on trust. Any agency that has produced reproducible citation growth can name the AI engine, show you the prompt set, and point you to a reference client willing to confirm the numbers directly. If it cannot do all 3, the result is not yet verified.

What Proof Should You Demand Before Trusting an Agency's Past Results?

The three proofs to demand from an AEO agency, and what each one fails on
Pass all three, or the result is not verified.

Answer engine optimisation (AEO) and generative engine optimisation (GEO) have become the phrase every agency reaches for right now, whether or not they can actually deliver it. That crowding is exactly why the burden of proof matters here more than in an established discipline: before trusting any AEO result, demand 3 things. A live query dashboard you can re-run yourself, a case study naming the AI engine tested, and a reference client willing to verify the figures. TrustRadius's 2026 report found that 94% of B2B buyers independently verify AI-generated research before acting on it. Apply the same standard to the agency's own claims.

A live query means the agency runs the search prompt in front of you, on the day, in real time. A result that cannot be reproduced today is a claim about what once happened, not proof of what the agency built.

Ask for the specific AI engine. ChatGPT, Perplexity, Gemini, and Google AI Overviews each have different citation behaviours and content source preferences. A case study that does not name the engine cannot be compared across platforms.

Ask for the reference client's name, not a vague vertical description. If the agency will not provide a name until after you sign, ask why. The strength of that evidence should also influence how you choose an AEO agency before adding it to your shortlist.

How Do You Verify an AEO Case Study Is Real, Not Cherry-Picked?

Earned media against brand-owned pages as a share of AI citations, plus the split by engine
On-site proof misses most of where citations come from.

Check whether the proof comes from sources an AI engine actually cites. 84% of AI citations come from earned media, not brand-owned pages, per Muck Rack's Generative Pulse report (May 2026). A case study built only from on-site content misses most of where citation growth should show up. Ask the agency which third-party publications are driving their client's citations.

The on-site trap: many AEO case studies show traffic growth from the client's own website. That proves search engine performance. It does not prove AI engines cite the brand in response to buyer prompts. These are different results measured differently.

Ask specifically: "Which publications did the AI engine cite when mentioning your client?" If the agency cannot name them, the evidence is on-site traffic data only.

A static screenshot proves a result existed at one moment. A live query proves it exists today. That distinction matters when comparing AEO agencies in Australia, particularly where providers rely on case studies to demonstrate citation growth.

Which Metrics Actually Prove Citation Growth, Not Vanity Growth?

Citation rate, mention rate and share of voice compared, with what a weak result looks like
One blended number hides which engine is working.

A result proves citation growth only when it separates 3 figures: citation rate, mention rate, and share of voice. They move independently. Profound's 11.84-billion-citation study found brand-site citation share ranging from 47% on ChatGPT to 69% on Google Gemini. A blended percentage that combines all 3 hides which engine is performing and which is not.

Citation Rate

Citation rate is the percentage of AI-generated answers that name the brand as a direct source. A high citation rate means the AI engine pulls the brand explicitly when responding to relevant prompts. It is the clearest signal of active recommendation.

Mention Rate

Mention rate is the percentage of AI answers where the brand appears anywhere in the response, including indirect references that do not attribute it as the primary source. High mention rate with low citation rate means the AI recognises the brand but does not lead with it.

Share of Voice

Share of voice measures the brand's citation volume against named competitors on the same prompt set. Without a competitor baseline, a 20% citation rate is not interpretable. Against a competitor at 8%, it is a 2.4x advantage.

The Kynection AEO case study shows what separated tracking looks like in practice: 22.88% AI share of voice, first of 40 tracked competitors in its category and about 3.6x the nearest global rival (measured by our own tracker, 26 August 2026, 8,891 runs across 393 prompts). That separation is only visible when citation rate, mention rate, and share of voice are tracked independently, across a content programme that now runs to 71 deep-dive pieces.

What Should a Reference Call Cover Before You Sign?

A reference call should confirm 3 things: whether the client's dashboard matches what was reported, whether the result held after the campaign ended, and whether the timeline matched. 73% of B2B decision-makers trust peer recommendations above all other sources. A 15-minute call is the cheapest verification available.

According to UserEvidence's 2026 research, 26% of deals fail for customer-evidence reasons: return on investment (ROI) not proven, a lack of references, or no clear differentiation. A reference call addresses all 3 before you commit a budget.

Ask 3 questions on the call:

  • "Do the numbers on your own dashboard match what the agency reported in the case study?" If the client cannot access a tracking dashboard, ask how the result was measured.
  • "Did the result hold after the campaign push ended?" A result that peaked during heavy content production and then dropped is not a structural citation gain.
  • "Did the work complete on the quoted timeline?" Timeline accuracy is a proxy for how well the agency scopes and manages its content programme.

Request 2 reference clients from different verticals where possible. 1 reference does not tell you whether results transfer across industries.

What Red Flags Mean a Case Study Will Not Hold Up?

Five red flags that mean an AEO case study was assembled after the fact
Two or more together is the tell.

With so many agencies now claiming GEO or AEO capability overnight, the same handful of red flags tend to separate a real result from a repackaged one. A case study fails scrutiny when it cannot name the AI engine tested, cannot be reproduced on a live query, or credits growth to "AI optimisation" with no detail on what changed. Each pattern points to a number reported once and never verified.

Watch for these 5 patterns:

  • No AI engine named: The case study says "AI visibility improved" with no mention of ChatGPT, Perplexity, Gemini, or AI Overviews. Visibility on which platform? Against which prompts?
  • Result cannot be reproduced live: If the agency cannot run the search prompt today, the result no longer exists or never existed at scale.
  • Growth credited to "AI optimisation" with no mechanism: What content changed? Which URLs were targeted? Which prompts moved? A claim without a mechanism is not evidence.
  • On-site traffic data only: Session counts and pageview growth are search engine metrics. They do not measure AI citation rate, mention rate, or share of voice.
  • No reference client willing to speak: An agency with 3 or more successful campaigns has at least 1 client willing to verify results on a call. "Clients prefer anonymity" is not a sufficient answer.

Two or more of these patterns together means the case study was assembled after the fact, not tracked in real time. It is the same discipline behind our Content Strategy work: every proof point is tracked from campaign start, not constructed at pitch time.

How Do You Confirm a Result Holds After You Sign?

The verification question does not end at signing. Half of all content cited by AI engines is less than 13 weeks old, according to Profound's research on citation decay (August 2026). A result confirmed at pitch can disappear before renewal. Ask the agency how they track citation rate, mention rate, and share of voice after work begins.

Ask for monthly snapshots from a previous engagement: citation rate by engine, mention rate trend, and share of voice against 3 to 5 named competitors. If the agency cannot produce these from a past client, they have not been tracking results continuously.

A result that grows and holds over 6 months is a structural gain. A result that peaks during heavy content production and then decays is a campaign peak. The difference only becomes clear when the agency shows month-by-month data, not a single end-point snapshot.

Content Creation

See what your AI visibility actually looks like

Before committing to an AEO programme, establish where your brand stands today. A free AEO audit benchmarks your citation rate, mention rate and share of voice against your top three competitors.

Frequently Asked Questions

FAQs

How do you tell if an AEO case study is real?

Ask for a live query you can watch in real time, not a screenshot or a PDF. A real result can be reproduced today on the same AI engine. If the agency can only provide static images, ask why the result cannot be shown live. The ability to run the query today is the clearest signal the work is active and the result holds.

What is the difference between citation rate and mention rate?

Citation rate is the percentage of AI-generated answers that name your brand as a direct source. Mention rate is the percentage of answers where your brand appears anywhere in the response, including indirect references. Both figures matter and move at different speeds. An agency reporting 1 blended number has chosen the easiest figure to present, not the most complete picture.

How many reference clients should an AEO agency provide?

Request 2 reference clients from different industries. 1 reference does not tell you whether results transfer across verticals. An agency that has run 3 or more successful campaigns has at least 1 client in a different vertical who will take a call to verify results.

Which AI engines should an AEO case study cover?

At minimum, ChatGPT and Google AI Overviews, which account for the largest share of B2B AI search traffic. A case study covering 1 engine says nothing about performance on the others. Perplexity and Google Gemini are relevant for B2B buyers who use research-mode AI tools.

How long does it take to see measurable AEO results?

Citation rate improvements appear within the first monthly reporting cycle in any active content programme. Share of voice movement against named competitors takes 3 to 6 months and depends on how active the competitor content programme is. An agency without monthly tracking data from a past client has not been measuring results continuously.

SHARE