What Can an AI Engine Actually Read on a Spec Page?
When an AI engine opens a page directly, it reads the page’s HTML text and little else. searchVIU put 8 prices on one test page in October 2025, some visible on screen and some hidden in code, then asked 5 AI systems to read them. No system read a price kept only in JavaScript Object Notation for Linked Data (JSON-LD) code, and only Gemini read a price loaded by JavaScript.
Manufacturers often go missing from AI answers before a buyer ever sees their site, as our guide to why manufacturers miss AI answers explains. This guide looks at the cause: what an engine does with a datasheet, a spec table or a product tab.
searchVIU’s 8-test experiment asked each system the same price question. The results split by how each system reached the page:
| AI system | How it reached the page | What it found |
|---|---|---|
| ChatGPT | Opened the page directly | The 3 prices shown as visible text, and none of the other 5 |
| Gemini | Opened the page directly | The visible prices plus the JavaScript-loaded price |
| Claude | Opened the page directly | No prices at all, even the visible ones |
| Google AI Mode | Searched its own index | 2 prices, only after the page was indexed |
| Perplexity | Searched its own index | Only the JavaScript-loaded price, after indexing |
Manufacturer sites hide specs from engines in 3 common places:
- Script-loaded tabs and configurators. The pressure rating appears on screen, but only after JavaScript runs.
- Schema-only values. The Product markup holds the dimensions, but the page body doesn’t repeat them.
- Image-only tables and scanned datasheets. The numbers exist only as pixels, so there’s nothing to quote.
If a value matters to a buyer, write it as plain text in the HTML.

Does Schema Markup Change How Engines Judge Spec Data?
On current evidence, schema doesn’t change how often engines cite a page. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, each matched against 3 control pages that never added it. ChatGPT and AI Mode citations barely moved, and AI Overviews dipped slightly.
The test has limits. Every page was already cited often before the schema went on, so nobody knows yet whether schema helps a page that engines haven’t found. It also grouped all schema types together, so Product schema wasn’t tested on its own.
A 2026 retrieval experiment points the same way. Adding JSON-LD to web pages gave only modest gains in answer accuracy. The big gains, close to 30%, came from richer entity pages with clear navigation.
So keep Product schema for 2 jobs: Google Merchant Center listings and rich results, the extra details Google shows under a result. Make sure every value in it also appears in the visible text.

How Does One Engineering Question Become Several Searches?
AI search engines break one question into several smaller searches, a process called query fan-out. Seer Interactive ran 501 prompts through Google’s Gemini 3 model and found 10.7 searches per prompt on average, and up to 28. Each sub-search pulls its own passage, so each sub-question needs an answer.
Seer Interactive’s fan-out research adds a detail that matters for industrial search. 95% of those searches had no monthly search volume, so keyword tools never show them. You have to work them out from how engineers actually specify a part.
Take an example question: “Which ball valve suits steam lines in a food plant?” An engine is likely to split it into searches like these:
- Pressure and temperature rating for steam service.
- Body and seat materials suited to steam.
- Hygiene or food-contact standards the valve meets.
- End connections and sizes available.
- Local stock and lead time.
How the page is laid out decides how many of those searches it can answer:
| Page layout | What the engine gets |
|---|---|
| One labelled block per sub-question | A passage to quote for each of the 5 searches |
| One dense table with no labels | Numbers without context, so it quotes a page that explains them |
Mapping these sub-questions across a whole product range is the core of the content we build for AI search.

Where Should the Operating Envelope Sit on a Long Spec Page?
Put it in the first screen, not buried mid-page. Research on how language models read long inputs found they use what sits at the start and the end of a context far better than what sits in the middle. On a spec page with a dozen tables, the middle is where a value is most likely to be passed over.
A long-context study published in Transactions of the Association for Computational Linguistics tested models on multi-document question answering and key-value retrieval. Accuracy was highest when the passage that mattered sat at the beginning or the end, and fell sharply when it sat in the middle, even for models built for long contexts.
One caveat travels with this finding. It measured a model reading a long input, not an engine choosing which page to cite, so treat it as a reason to front-load rather than proof that front-loading earns citations.
Seer’s fan-out data points the same way on naming. 26.4% of the fan-out queries included a brand name, and the longest query in the set put 4 named product lines against each other. Searches like that match a page that writes its model numbers and standards out in full.
So the first screen of a product page should hold 3 things:
- A one-paragraph summary of where it works. What the product is, its key ratings, its material and the standard it meets.
- Identifiers written in full. The model number and each certification written as text, beyond any logo strip.
- One limit of use. Where the product should not be used, stated plainly.
Moving specs out of download-only PDF files follows the same logic, as our PDF-to-HTML migration workflow shows.

Which Spec-Content Tactics Are Tested, and Which Are Assumed?
Only a few spec page tactics have been tested directly. Tests show what engines can read when they open a page, and schema was tested and showed no lift. Most other popular fixes rest on patterns or reasoning.
| Tactic | Evidence behind it | What the evidence supports |
|---|---|---|
| Plain-text specs in HTML | Direct test, 5 AI systems, October 2025 | ChatGPT and Gemini read them on a live fetch |
| JSON-LD schema | Controlled test, 1,885 pages, 2026 | No meaningful citation lift on pages already cited |
| Answer-first layout | Long-context retrieval tests, published 2023 | Models use the start and the end of an input best. Not tested on web pages |
| PDF-to-HTML migration | Reasoning from the tests above | Likely helps image-only PDFs. Not tested on its own |
| llms.txt | Server logs, 137,210 domains | Almost no engine requests the file |
The llms.txt figure comes from Ahrefs’ server-log study. About 38,000 of those domains had published the file, and 97% of the files got no requests at all in May 2026.
Intelligent Resourcing doesn’t sell llms.txt as a way to get cited. It takes 5 minutes and does no harm, but it only tells AI tools what they can use. It doesn’t help you get found.
Ask one question of every line in any proposal: what proof shows this gets you cited more? A pattern is fine as long as it says it’s a pattern. Our guide to search engine optimisation (SEO) for large language models (LLMs) rates the wider set of tactics the same way.
What Should a Manufacturer Test First?
Start by checking what an engine can reach and read before you change any copy. A clean robots.txt file doesn’t prove engines can get in. Your firewall or content delivery network (CDN) can still block AI crawlers, or let them through.
A study of AI crawler blocking, accepted at the Association for Computing Machinery’s (ACM) 2025 Internet Measurement Conference, found blockers built into network services stop AI crawlers more firmly than robots.txt. Few sites used them yet. Either way, your network settings decide what actually gets through.
A 5-step test shows where your specs stand:
- Fetch 5 key product pages as plain HTML. Search the source for the 3 numbers buyers ask about most.
- Check your CDN and firewall bot settings. Robots.txt says what you want. Your firewall decides what gets through.
- Ask ChatGPT, Gemini and Perplexity your 10 most common engineering questions. Record which source each engine quotes.
- Rewrite the first screen of the 5 weakest pages. Add the operating limits, part numbers and standards as plain text.
- Re-run the same questions after a recrawl. Compare against 5 pages you left untouched, so a platform-wide shift doesn’t look like your win.
Content Creation
Know which spec questions you already win in AI answers, and which ones a competitor’s page is taking, before you touch a datasheet.
FAQs
Do AI engines read schema markup on product pages?
Not when they open a page directly, based on current tests. In searchVIU’s October 2025 test, no AI system read a price kept only in JSON-LD. Ahrefs tracked 1,885 cited pages that added schema and saw no real rise in citations.
Can ChatGPT read specifications inside a PDF datasheet?
As of October 2026, the public tests covered here didn’t compare PDF and HTML specs. Scanned or image-only datasheets hold numbers as pixels, so there’s no text to quote. Keep the PDF and repeat the key values as plain HTML text.
Why does ChatGPT miss specs shown in product tabs?
Many product tabs load their values with JavaScript after the page opens. In searchVIU’s test, ChatGPT missed the JavaScript-loaded price while Gemini found it. Writing key values as plain text on the page removes the problem for every engine.
Does adding more technical detail help AI citation?
Detail helps when an engine can read it and reach it early. Long-context tests show models use the start and the end of an input better than the middle, so an operating envelope buried mid-page is the easiest thing to miss. Seer Interactive also found 26.4% of Gemini 3 fan-out queries named a brand, so model numbers and standards written out in full are what those searches match.
What is a query fan-out in AI search?
Query fan-out is how an AI engine splits one question into several smaller searches. Seer Interactive found Google’s Gemini 3 model runs 10.7 searches per question on average. Each sub-search pulls its own passage, so each sub-question needs a clear answer.

