HomeBlog

How Do AI Engines Evaluate Technical and Spec-Heavy Content?

Spec sheet SEO fails when AI engines can’t read your specs. Live tests show schema-only and tab-loaded values get skipped. Run the 5-step access test first.

Last reviewed:
October 5, 2026
· Reviewed quarterly for accuracy
How Do AI Engines Evaluate Technical and Spec-Heavy Content?
Key Facts

Manufacturers miss out on artificial intelligence (AI) answers when engines can’t read their specs. Specs often hide in tabs that load late, scanned datasheets or schema code. AI engines read visible text, split each question into several searches and quote short passages. In a 2025 searchVIU test, no engine read a price kept only in schema.

TL;DR
  • Visible text is the spec sheet engines see. Values kept only in hidden code were missed by every system tested, and a value loaded by script was read by just one.
  • Schema isn’t a proven citation lever. A controlled Ahrefs test of 1,885 pages found no meaningful lift on any platform.
  • One engineering question becomes many searches. Each sub-question needs its own short, labelled answer on the page.
  • Put the operating envelope first. Language models use the start and the end of a long input far better than the middle.
  • Ask for the test behind every tactic. Many popular fixes for spec pages have never been properly tested.
Decision Matrix
Where a spec value livesRead in searchVIU’s live test?Controlled evidence of a citation liftWhere it earns its place
Visible HyperText Markup Language (HTML) table with a plain summaryYes, by ChatGPT and GeminiNo test has measured it on its own yetThe main home for any number a buyer asks about
Table loaded by JavaScript (tabs, configurators)Only by GeminiNone foundInteractive tools, with key values also written as plain text
Product schema only, in JavaScript Object Notation for Linked Data (JSON-LD)No system read itTested on 1,885 pages: no meaningful liftGoogle Merchant Center listings and rich results
Download-only Portable Document Format (PDF) datasheetNot covered by the testNone foundCertified drawings and compliance records engineers file
llms.txt file pointing at datasheetsNot covered by the testNone found. 97% of files got no requests in May 2026Setting rules for AI tools
Steelman: keep schema and PDFsSchema feeds Google Merchant Center and rich results. Ahrefs only tested pages engines already citedUnknown for pages no engine has found yetA signed PDF stays the right home for certified drawings and compliance documents
The Verdict

Put every number a buyer asks about into visible HTML on the product page, with a short plain summary near the top. Treat schema and PDFs as extras, since they aren’t how engines judge your specs. If you sell through Google Merchant Center, or the PDF is the certified record, keep both, but still repeat the key values as plain text.

What Can an AI Engine Actually Read on a Spec Page?

When an AI engine opens a page directly, it reads the page’s HTML text and little else. searchVIU put 8 prices on one test page in October 2025, some visible on screen and some hidden in code, then asked 5 AI systems to read them. No system read a price kept only in JavaScript Object Notation for Linked Data (JSON-LD) code, and only Gemini read a price loaded by JavaScript.

Manufacturers often go missing from AI answers before a buyer ever sees their site, as our guide to why manufacturers miss AI answers explains. This guide looks at the cause: what an engine does with a datasheet, a spec table or a product tab.

searchVIU’s 8-test experiment asked each system the same price question. The results split by how each system reached the page:

AI systemHow it reached the pageWhat it found
ChatGPTOpened the page directlyThe 3 prices shown as visible text, and none of the other 5
GeminiOpened the page directlyThe visible prices plus the JavaScript-loaded price
ClaudeOpened the page directlyNo prices at all, even the visible ones
Google AI ModeSearched its own index2 prices, only after the page was indexed
PerplexitySearched its own indexOnly the JavaScript-loaded price, after indexing

Manufacturer sites hide specs from engines in 3 common places:

  • Script-loaded tabs and configurators. The pressure rating appears on screen, but only after JavaScript runs.
  • Schema-only values. The Product markup holds the dimensions, but the page body doesn’t repeat them.
  • Image-only tables and scanned datasheets. The numbers exist only as pixels, so there’s nothing to quote.

If a value matters to a buyer, write it as plain text in the HTML.

The six places a spec value can sit, with whether each was read in the searchVIU live test
Visible text is the only placement that more than one system read.

Does Schema Markup Change How Engines Judge Spec Data?

On current evidence, schema doesn’t change how often engines cite a page. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, each matched against 3 control pages that never added it. ChatGPT and AI Mode citations barely moved, and AI Overviews dipped slightly.

The test has limits. Every page was already cited often before the schema went on, so nobody knows yet whether schema helps a page that engines haven’t found. It also grouped all schema types together, so Product schema wasn’t tested on its own.

A 2026 retrieval experiment points the same way. Adding JSON-LD to web pages gave only modest gains in answer accuracy. The big gains, close to 30%, came from richer entity pages with clear navigation.

So keep Product schema for 2 jobs: Google Merchant Center listings and rich results, the extra details Google shows under a result. Make sure every value in it also appears in the visible text.

Prices found out of eight by Gemini, ChatGPT, Google AI Mode, Perplexity and Claude on the same test page
Not one of the five found the price that lived only in schema.

How Does One Engineering Question Become Several Searches?

AI search engines break one question into several smaller searches, a process called query fan-out. Seer Interactive ran 501 prompts through Google’s Gemini 3 model and found 10.7 searches per prompt on average, and up to 28. Each sub-search pulls its own passage, so each sub-question needs an answer.

Seer Interactive’s fan-out research adds a detail that matters for industrial search. 95% of those searches had no monthly search volume, so keyword tools never show them. You have to work them out from how engineers actually specify a part.

Take an example question: “Which ball valve suits steam lines in a food plant?” An engine is likely to split it into searches like these:

  1. Pressure and temperature rating for steam service.
  2. Body and seat materials suited to steam.
  3. Hygiene or food-contact standards the valve meets.
  4. End connections and sizes available.
  5. Local stock and lead time.

How the page is laid out decides how many of those searches it can answer:

Page layoutWhat the engine gets
One labelled block per sub-questionA passage to quote for each of the 5 searches
One dense table with no labelsNumbers without context, so it quotes a page that explains them

Mapping these sub-questions across a whole product range is the core of the content we build for AI search.

One ball valve question split into five sub-searches, with the fan-out averages from 501 Gemini 3 prompts
Each sub-search pulls its own passage, so each needs its own labelled block.

Where Should the Operating Envelope Sit on a Long Spec Page?

Put it in the first screen, not buried mid-page. Research on how language models read long inputs found they use what sits at the start and the end of a context far better than what sits in the middle. On a spec page with a dozen tables, the middle is where a value is most likely to be passed over.

A long-context study published in Transactions of the Association for Computational Linguistics tested models on multi-document question answering and key-value retrieval. Accuracy was highest when the passage that mattered sat at the beginning or the end, and fell sharply when it sat in the middle, even for models built for long contexts.

One caveat travels with this finding. It measured a model reading a long input, not an engine choosing which page to cite, so treat it as a reason to front-load rather than proof that front-loading earns citations.

Seer’s fan-out data points the same way on naming. 26.4% of the fan-out queries included a brand name, and the longest query in the set put 4 named product lines against each other. Searches like that match a page that writes its model numbers and standards out in full.

So the first screen of a product page should hold 3 things:

  • A one-paragraph summary of where it works. What the product is, its key ratings, its material and the standard it meets.
  • Identifiers written in full. The model number and each certification written as text, beyond any logo strip.
  • One limit of use. Where the product should not be used, stated plainly.

Moving specs out of download-only PDF files follows the same logic, as our PDF-to-HTML migration workflow shows.

Start, middle and end of a long input marked for how well models use each, beside the three first-screen items
Front-load the envelope, because the middle of a long page is the weak spot.

Which Spec-Content Tactics Are Tested, and Which Are Assumed?

Only a few spec page tactics have been tested directly. Tests show what engines can read when they open a page, and schema was tested and showed no lift. Most other popular fixes rest on patterns or reasoning.

TacticEvidence behind itWhat the evidence supports
Plain-text specs in HTMLDirect test, 5 AI systems, October 2025ChatGPT and Gemini read them on a live fetch
JSON-LD schemaControlled test, 1,885 pages, 2026No meaningful citation lift on pages already cited
Answer-first layoutLong-context retrieval tests, published 2023Models use the start and the end of an input best. Not tested on web pages
PDF-to-HTML migrationReasoning from the tests aboveLikely helps image-only PDFs. Not tested on its own
llms.txtServer logs, 137,210 domainsAlmost no engine requests the file

The llms.txt figure comes from Ahrefs’ server-log study. About 38,000 of those domains had published the file, and 97% of the files got no requests at all in May 2026.

Intelligent Resourcing doesn’t sell llms.txt as a way to get cited. It takes 5 minutes and does no harm, but it only tells AI tools what they can use. It doesn’t help you get found.

Ask one question of every line in any proposal: what proof shows this gets you cited more? A pattern is fine as long as it says it’s a pattern. Our guide to search engine optimisation (SEO) for large language models (LLMs) rates the wider set of tactics the same way.

What Should a Manufacturer Test First?

Start by checking what an engine can reach and read before you change any copy. A clean robots.txt file doesn’t prove engines can get in. Your firewall or content delivery network (CDN) can still block AI crawlers, or let them through.

A study of AI crawler blocking, accepted at the Association for Computing Machinery’s (ACM) 2025 Internet Measurement Conference, found blockers built into network services stop AI crawlers more firmly than robots.txt. Few sites used them yet. Either way, your network settings decide what actually gets through.

A 5-step test shows where your specs stand:

  1. Fetch 5 key product pages as plain HTML. Search the source for the 3 numbers buyers ask about most.
  2. Check your CDN and firewall bot settings. Robots.txt says what you want. Your firewall decides what gets through.
  3. Ask ChatGPT, Gemini and Perplexity your 10 most common engineering questions. Record which source each engine quotes.
  4. Rewrite the first screen of the 5 weakest pages. Add the operating limits, part numbers and standards as plain text.
  5. Re-run the same questions after a recrawl. Compare against 5 pages you left untouched, so a platform-wide shift doesn’t look like your win.

Content Creation

Which spec questions do you already win?

Know which spec questions you already win in AI answers, and which ones a competitor’s page is taking, before you touch a datasheet.

Frequently Asked Questions

FAQs

Do AI engines read schema markup on product pages?

Not when they open a page directly, based on current tests. In searchVIU’s October 2025 test, no AI system read a price kept only in JSON-LD. Ahrefs tracked 1,885 cited pages that added schema and saw no real rise in citations.

Can ChatGPT read specifications inside a PDF datasheet?

As of October 2026, the public tests covered here didn’t compare PDF and HTML specs. Scanned or image-only datasheets hold numbers as pixels, so there’s no text to quote. Keep the PDF and repeat the key values as plain HTML text.

Why does ChatGPT miss specs shown in product tabs?

Many product tabs load their values with JavaScript after the page opens. In searchVIU’s test, ChatGPT missed the JavaScript-loaded price while Gemini found it. Writing key values as plain text on the page removes the problem for every engine.

Does adding more technical detail help AI citation?

Detail helps when an engine can read it and reach it early. Long-context tests show models use the start and the end of an input better than the middle, so an operating envelope buried mid-page is the easiest thing to miss. Seer Interactive also found 26.4% of Gemini 3 fan-out queries named a brand, so model numbers and standards written out in full are what those searches match.

What is a query fan-out in AI search?

Query fan-out is how an AI engine splits one question into several smaller searches. Seer Interactive found Google’s Gemini 3 model runs 10.7 searches per question on average. Each sub-search pulls its own passage, so each sub-question needs a clear answer.

SHARE