GEO · SEO · AI Search · Structured Data

llms.txt is a proposal, not a ranking signal: what AI search actually documents

Published · Updated

Publishing an llms.txt file is not what gets you cited. Here is what the specification actually requires, what Google says you can skip, and which GEO lever has a measured number behind it.

A lot of GEO advice starts by telling you to publish an llms.txt file and stops there. That is a proposal with a specification, not a documented ranking input, and the two get conflated constantly. Here is what the published sources actually say you can configure for AI search, what is merely proposed, and which lever has a measured number attached to it.

What llms.txt actually specifies

The llmstxt.org proposal is for a Markdown file at /llms.txt that provides LLM-friendly content. The specification is small. An H1 with the name of the project or site is the only required section. An Optional section exists by convention for secondary information, holding links an agent can skip when a shorter context is needed. That is close to the whole contract.

The proposal goes further than the index file, and this is the part most rollouts drop. It also suggests that pages with information agents might need provide a clean Markdown version of those pages at the same URL as the original page. An index pointing at pages no agent can parse cleanly adds a file, not clarity. The second half is the work; the first half is the table of contents.

What Google documents, and what it says you can skip

Google's guidance on AI features is blunt about eligibility. To be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. The same page states plainly that you do not need to create new machine-readable files, AI text files, or markup to appear in these features.

That documentation also describes a query fan-out technique: AI Overviews and AI Mode may issue multiple related searches across subtopics and data sources to develop a response. The practical consequence is that one page tuned for one head query is competing against a spread of subtopic queries. Breadth of genuinely indexed, answerable pages matters more than any single file at the root of the domain.

Measurement follows the same logic. Google reports these appearances in the Performance report, within the Web search type. There is no separate AI dashboard to wait for, which also means no isolated attribution: those appearances sit inside the same numbers as everything else. Any GEO claim built on a clean split between AI traffic and classic search traffic is drawing a line the reporting does not give you.

The binary switch: OpenAI's crawler tags

OpenAI publishes two robots.txt tags, OAI-SearchBot and GPTBot, to let webmasters manage how their sites and content work with AI. The distinction matters. Sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though they can still appear as navigational links. Disallowing GPTBot indicates that a site's content should not be used in training generative AI foundation models. Visibility and training are separate decisions.

Two operational notes. These are opt-outs, not opt-ins: there is nothing to publish in order to be included, only something to publish in order to be excluded. And changes are not instant. For search results, it can take ~24 hours from a robots.txt update for OpenAI's systems to adjust. Audit the robots file you already have before you add a new file beside it.

Structured data is the lever with a number on it

Of everything in this pillar, structured data is the only element in the published guidance carrying a measured outcome. Rotten Tomatoes added structured data to 100,000 unique pages and measured a 25% higher click-through rate for pages enhanced with structured data, compared to pages without it. That is a click-through effect in Google Search, not an AI citation rate, and it should not be resold as one.

The quality rules are specific. You must include all the required properties for an object to be eligible for appearance in Google Search with enhanced display. Beyond that, it is more important to supply fewer but complete and accurate recommended properties than to provide every possible recommended property with less complete, badly formed, or inaccurate data. Do not create blank pages to hold markup, or describe information that is not visible to the user.

What this looked like on a build

On Tatano Energy, a biomass boiler manufacturer, the starting point was one generic site with no structured data and no localised content, against competitors established in each of four European markets. The rebuild put 4 country domains into separate indexation, served 7 languages and published 8 SEO articles every day with 0 manual intervention. Against query fan-out, that breadth is the point: many indexed pages, each able to answer a different subtopic search.

None of that depended on a proposal. It depended on requirements that are already documented: pages that are indexed, snippet-eligible, structurally marked up where the markup reflects visible content, and numerous enough to cover the subtopics a fan-out will generate. The daily autoblog keeps widening that surface without anyone touching it, which is what makes the approach survive contact with a small team.

If you want a defensible order of work: get pages indexed and snippet-eligible first, then complete the structured data you can fill accurately rather than every type you could technically emit, then decide deliberately what OAI-SearchBot and GPTBot are allowed to do. Add llms.txt if you will also ship clean Markdown versions of the pages it points to. Treat it as a convenience for agents, not as the thing that gets you cited.

Sources

llmstxt.org — The /llms.txt file, v2 — https://llmstxt.org/

Google Search Central — AI features and your website — https://developers.google.com/search/docs/appearance/ai-features

Google Search Central — Introduction to structured data markup in Google Search — https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data

OpenAI — Overview of OpenAI Crawlers — https://developers.openai.com/api/docs/bots

Neurolinks case study — Four markets, one codebase — https://neurolinks.be/work/tatano-energy

Working on a project where these methods apply?