HomeBlog

LLM SEO Agency: Run This AI Visibility Test Before You Hire One

Most LLM SEO agencies can sell AI visibility. Far fewer can prove it. Run this test before you hire, and see our own results including where we lose.

Last reviewed:
September 9, 2026
· Reviewed quarterly for accuracy
LLM SEO Agency: Run This AI Visibility Test Before You Hire One
Key Facts

An LLM SEO agency is worth hiring when it can prove AI-search visibility rather than claim it. The strongest evidence is dated, repeated performance on non-branded buyer prompts, with the engines named, the run count stated and the source pages behind each citation identified. Anything less is a screenshot.

TL;DR
  • Proof beats terminology. GEO, AEO, LLMO and AI SEO describe overlapping work. Pick the provider that can measure, not the one with the newest label.
  • Test with non-branded prompts. Search the category, not the agency name. A brand-name prompt makes visibility close to guaranteed and proves nothing.
  • Demand a run count. One favourable answer is an observation. Dozens of runs on a fixed prompt is a reading.
  • Ask for the gaps too. An agency that only shows wins is selecting its evidence. We publish ours below, including a prompt where we are absent.
  • Get a baseline before you sign. Without a pre-engagement reading, no improvement claim afterwards can be checked.
Decision Matrix
Evidence offeredWhat it tells youBuying implication
Branded searches onlyThe engine recognises the agency when namedWeak. Does not show category visibility
One non-branded screenshotIt appeared once for a relevant queryDirectional only. May not repeat
Dated non-branded prompt, engine namedDocumented visibility on one platformUseful but partial. Check the other engines
Repeated runs with a run countThe result happens often enough to measureStrong. Repeatability beats a good example
Source-level citation tracingIt knows which page caused the mentionHigh value. That is the lever it can pull
Wins and gaps published togetherIt measures systematically instead of cherry-pickingStrong trust signal
Weak own visibility, strong client proofDelivery capability may still be realValid exception. Verify the client results
The Verdict

If an agency can show dated, repeated, non-branded prompt results, name the engines tested and trace the sources behind its citations, it clears the first credibility test. If it can only show branded searches, undated screenshots or a single unexplained visibility score, keep looking.

We publish our own numbers below, wins and gaps together, so this article can be checked against the standard it sets. That does not prove we would rank in your category. It proves the measurement system existed before the promise.

Is There a Real LLM SEO Agency Category Yet?

Four figures on the LLM SEO agency market: 150 monthly Australian searches, 84% recognise GEO, 87.4% of AI referral traffic from ChatGPT, 48% cannot track it
Real demand, unsettled vocabulary, and an industry mostly unable to measure what it sells.

As a search behaviour yes, as a settled label no.

Ahrefs shows about 150 monthly Australian searches for "llm seo agency" at keyword difficulty 1, read on 9 September 2026. That combination, real demand with almost no organised competition on the exact phrase, describes a category still forming.

The vocabulary has not settled either. Fractl surveyed 342 search-industry readers and found GEO recognised by 84%, AEO by 61% and AISEO by 60%. Four labels, one job. The commercial question is identical across all of them: can this provider improve and measure how your brand appears when a buyer asks an AI system about your category?

The underlying behaviour is real regardless of naming. Conductor's benchmark across 13,770 domains and 3.3 billion sessions in ten industries found 87.4% of AI referral traffic arriving from ChatGPT, with AI referrals at 1.08% of all website traffic over its May to September 2025 window. Small share, fast-moving, and concentrated on one engine.

How Do You Test Whether an Agency Practises What It Sells?

Four cards setting out the test: non-branded prompt, repeated runs, engine and date recorded, and asking what produced the result
The whole test takes about ten minutes and needs no tooling.

Run the test yourself before you read the pitch deck.

The reason to use a non-branded prompt is mechanical. As SE Ranking puts it, branded and category prompts have to be tracked separately, because when your brand name is in the prompt the visibility is nearly guaranteed. Asking an engine about the agency by name tells you the engine has heard of it. That is not the thing you are buying.

Run the test in four steps

  1. Use a non-branded prompt. Ask something like "which agencies specialise in AI search visibility in Australia" rather than typing the agency's name.
  2. Test more than once. SparkToro had 600 volunteers run 12 prompts across ChatGPT, Claude and Google's AI Overview 2,961 times and found under a 1 in 100 chance of the same prompt returning the same brand list. One run is noise.
  3. Record the engine and the date. ChatGPT, Gemini, Copilot and Google AI Overview retrieve and weight sources differently. Cited on one is not cited on all.
  4. Ask what produced the result. A credible provider can name the page, roundup or third-party publication the engine drew from. An agency that can trace the source can influence it.

What counts as adequate evidence

Five components: a non-branded prompt or prompt set, the engines named, the date recorded, the run count behind the result, and the source page the engine cited. An undated screenshot with no engine and no run count is the weakest artefact in this category, because it cannot be replicated or compared to anything later.

What Does Our Own Test Result Show?

Table of Intelligent Resourcing's own tracked results, including two prompts where it is the first-cited source and one where it is not cited at all
The same standard applied to us, read on 9 September 2026, gaps included.

Here is the same standard applied to us, read from our tracker on 9 September 2026.

Tracked promptResultSample
Which companies are experts in AI Citation and Tracking in Australia?First-cited source13 citations across 50 runs, on ChatGPT, Google AI Overview, Copilot and Gemini
Best ChatGPT Citation Agencies in AustraliaFirst-cited source9 citations across 20 runs, on Google AI Overview, Copilot and Gemini
AI SEO agency for B2BNot cited at all0 citations across 32 runs

Two things in that table matter more than the wins.

The first is what "first-cited source" means. It means our pages are the top source the engine drew from. On the second prompt the answer text still names other agencies before it cites us. Being the most-used source and being the first name a buyer reads are different outcomes, and a blended visibility score hides the difference. We report the source position because that is what the measurement actually is.

The second is the third row. We are absent on "AI SEO agency for B2B" across 32 runs, and the exact phrase "llm seo agency" is not in our tracked set at all. Citation is a categorisation problem more than an authority problem: the same site, at the same authority, can be the most-cited source in one category and invisible in the one next to it.

Across the wider set, the tracker covers 197 prompts and 6,898 runs on ChatGPT, Google AI Overview, Gemini and Copilot. Our domain is cited in 26.1% of those responses. That is a dated sample on a defined prompt set, not a guarantee on any query.

That level of disclosure is not standard. AgencyAnalytics surveyed 494 agency professionals, reported by Search Engine Land in July 2026, and found 66% named AI search visibility as the top new service clients were asking for, while 48% said they could not reliably track people discovering a brand through AI tools. Two thirds selling it, roughly half unable to measure it.

What Should You Ask Before Hiring?

Table ranking kinds of evidence from branded searches at the weak end to published wins and gaps at the strong end
Where each kind of proof sits, from a screenshot to a repeatable reading.

Ask for dated evidence before you ask for a definition. Four questions separate providers quickly.

  • Which prompts do you track for your own brand? Not "we appear in AI search". Which prompts, on which engines, how often.
  • Are they branded or non-branded? Branded prompts confirm the engine knows the agency exists. Nothing more.
  • How many runs sit behind each result? One answer is noise. Dozens over a defined window is a signal.
  • Which source pages drive the citations? If a provider cannot name them, it does not yet understand its own citation architecture.

Our guide to testing a vendor's label covers the terminology side of the same conversation, and our breakdown of AEO, GEO and LLMO separates the three properly.

What Does It Cost to Hire an Agency That Fails Its Own Test?

The real cost is not the fee. It is funding a measurable channel without establishing whether the provider can measure it.

The gap shows up in three places:

  • No baseline, no comparison. Without a pre-engagement reading across the relevant engines, any later improvement claim is unverifiable.
  • No prompt-level tracking, no category evidence. A blended score can rise on branded queries while the prompts that decide vendor shortlists stay untracked.
  • No source tracing, no lever. An agency that cannot identify which pages produced its citations cannot reproduce the outcome deliberately.

A workable sequence is: define the prompt set, record a baseline across the engines that matter, trace which sources drove each cited result, ship the intervention, then measure the same prompt set again. That produces a before-and-after record instead of a score the buyer has to trust. Our generative engine optimisation service runs in that order, and our guide to choosing an AEO agency applies the same test to the wider field.

Budget should follow evidence. An agency selling AI visibility without showing its own is asking you to fund its first proof of concept.

Content Creation

Want the same reading?

Every number in this article is a dated reading on a stated prompt set, gaps included. That is the standard we think you should hold any provider to, including us. Book a GEO diagnostic call with Intelligent Resourcing to get the same reading for your own category.

Frequently Asked Questions

FAQs

Is "LLM SEO agency" an established category?

It is a real search phrase, about 150 monthly Australian searches at keyword difficulty 1 on Ahrefs in September 2026, but the label is not settled. Buyers and practitioners use SEO, AI SEO, AEO, GEO and LLMO for overlapping work. What matters is whether the provider can improve and measure visibility inside AI answers.

How do I test whether an AI search agency is actually visible?

Use non-branded prompts that describe the problem, run them repeatedly across the engines your buyers use, and record which agencies appear. Then ask the provider for its own dated measurement: prompts tracked, engines measured, runs behind each result and the source pages driving its citations.

Is Intelligent Resourcing cited for AI search agency prompts?

On some, not all. On 9 September 2026 we are the first-cited source on two tracked Australian prompts, with 13 citations across 50 runs and 9 across 20. We are not cited at all on "AI SEO agency for B2B" across 32 runs, and "llm seo agency" is not yet in our tracked set. We publish the gaps because undisclosed gaps are the metric most agencies avoid.

What should I ask an agency before hiring it?

Which prompts it tracks, which engines it measures, whether results are branded or non-branded, how many runs support each claim, which sources drive the citations and how it proves change after work begins. Its own visibility is one credibility check, not the whole decision.

What does it cost if an agency cannot prove its visibility?

Fees plus opportunity cost. Funding content, schema or outreach without a prompt-level baseline leaves no way to connect the work to a measurable change. The stronger standard is a provider that can show the baseline, the intervention and the resulting change on the prompts your buyers actually use.

SHARE