llms.txt: What It Is, How to Create One and Whether It Works in 2026
llms.txt promises to tell AI which pages matter. We break down the format, the Ahrefs and SE Ranking data and Google's position, then show step by step how to create one.

Table of contents
- What is llms.txt and what does it look like?
- How is llms.txt different from robots.txt and sitemap.xml?
- Do AI systems read llms.txt?
- Does llms.txt help you get AI citations?
- Who is llms.txt actually useful for?
- How do you create llms.txt in five steps?
- What mistakes do people make with llms.txt?
- How do you check that llms.txt works?
- What should you do for GEO instead of chasing llms.txt?
- FAQ
- Practical summary: what to do next
llms.txt is a markdown file in the root of a site that tells language models which pages matter. Google does not support it, and the large studies found no link between the file and AI citations. Below: how the file is built, who it helps, how to make one in an hour, and what actually works for visibility in AI search.

What is llms.txt and what does it look like?
llms.txt is a markdown file at /llms.txt that gives language models a shortlist of a site's most important pages with short descriptions. The format was proposed by Jeremy Howard of Answer.AI in September 2024. It is a community proposal, not a ratified standard, and the file does not control crawler access.
The structure is minimal. The file contains an H1 with the project name, a blockquote with a short summary, optional free text without headings, and zero or more H2 sections with lists of links. Each list item is a link, optionally followed by a colon and a note. The only mandatory element is the H1, but a file without link sections is pointless.
An example for a fictional agency site:
# Acme Digital
> An SEO and GEO agency for B2B companies in Ukraine.
> We work on technical SEO, content and visibility in AI search.
## Services
- [SEO promotion](https://example.com/services/seo): audit, technical SEO, content strategy
- [GEO promotion](https://example.com/services/geo): work on citations in AI answers
## Knowledge base
- [What is a canonical URL](https://example.com/kb/canonical): definition, code example, common mistakes
- [What is E-E-A-T](https://example.com/kb/eeat): how Google assesses expertise
## Optional
- [Case studies](https://example.com/cases): results of client projectsThe section named Optional has a special role: in the specification it holds secondary material that a model may skip if it is short on context.
How is llms.txt different from robots.txt and sitemap.xml?
robots.txt sets rules for crawlers, sitemap.xml lists all pages for indexing, and llms.txt is an editorial selection of what models should read first. The file neither allows nor forbids anything. It is a hint, not an instruction.
| File | What it does | Who uses it |
|---|---|---|
| robots.txt | Allows or blocks crawling | Search and AI crawlers (voluntarily, but widely) |
| sitemap.xml | Lists all URLs for indexing | Search engines |
| llms.txt | Picks key pages with descriptions | Some AI agents and developer tools |
Hence a practical rule: if you need to close a section off from AI bots, you need robots.txt. llms.txt is no use for that.

Do AI systems read llms.txt?
Google says outright that AI Overviews and AI Mode do not need the file. The other major model developers have not documented that their crawlers read llms.txt on third-party sites. The file is fetched mostly by coding agents, while the search-type AI bots that citations depend on almost never request it.
Google's position is recorded in the official Search Central guide of 15 May 2026. Among the myths about optimising for generative search, the guide lists the need for llms.txt, markdown pages, "chunking" content into short paragraphs and a special "AI-friendly" writing style. At the same time, the guide describes only Google Search: it does not explain how ChatGPT, Claude or Perplexity choose their sources.
Other companies present a different picture. Anthropic, OpenAI, Perplexity, Meta and Mistral publish llms.txt for their own documentation, but have not committed to reading the format on other sites. That is publisher behaviour, not proof that their crawlers consume such files across the internet.
Does llms.txt help you get AI citations?
According to the available research, no. SE Ranking's analysis of 300 thousand domains found no correlation between having the file and citation frequency. Ahrefs' logs from 137 thousand domains showed that the vast majority of files get no requests at all. The data describes the state of things in 2026 and may change, but for now it is unambiguous.
97% of llms.txt files received no requests during a month of observation. — Ahrefs, via Search Engine Journal, 2026
The details of this data matter. Among the requests that did arrive, search-type AI bots accounted for only 1.1%, and 12% came from tools that audit the files rather than use them. Ahrefs also measured requests, not whether a model acted on what it read.
The SE Ranking study gave a similar result: the file was found on only 10.13% of domains, and a citation-prediction model with the llms.txt feature removed became more accurate. In other words, for predicting citations the presence of the file worked as noise.
Another check is citations themselves. Of 18,000 citations in Search Engine Land's sample, only six led to llms.txt files, and "optimised" versions stuffed with keywords got none.
6 of 18,000 citations led to llms.txt files (0.03%). — Search Engine Land, 2026
A caveat on this data: Ahrefs works with a sample of technically advanced clients, so the overall picture may differ. But none of the known independent studies has shown a positive effect on citations.
Who is llms.txt actually useful for?
The file helps most for documentation and API products, where developers work through AI assistants. There it removes a concrete problem: the agent gets a clean map of sections instead of parsing HTML full of navigation and scripts. For service sites and blogs the effect is most likely minimal.
You can see this in who uses it. Stripe, Vercel, Cloudflare, Anthropic, Coinbase, Pinecone and Cursor publish llms.txt because their users build products with AI coding assistants. Documentation platforms such as Mintlify generate these files automatically, so many sites got them "by default".
For an agency site the calculation is simple: the file costs an hour or two of work and breaks nothing, but you should not expect traffic from it. It makes sense as cheap insurance in case some AI tool starts to take it into account, and as a discipline: to assemble the 15-20 most important pages you have to choose and describe them.
How do you create llms.txt in five steps?
Create a text file in the domain root, open it with an H1 holding the brand name and a one- or two-sentence blockquote, group the links under H2 headings and add a description to each. Keep the file short: it is a selection, not a copy of the sitemap.

- Choose the pages. Keep 15-30 URLs: key services, your best reference materials, the about and contact pages. Drop everything that is not meant for people.
- Write the H1 and the summary. The brand name, then one or two sentences on what the company does and for whom. A model that reads only the top of the file should already understand who you are.
- Group the links under H2. Sections by area: services, knowledge base, case studies. Move secondary material into
Optional. - Describe every link. The format is
[Title](URL): what is on the page. The description should match the page's content, with no promotional epithets. - Publish at
/llms.txt. The address must return a 200 code with no redirects. Subdirectories such as/api/llms.txtor/.well-known/llms.txtare not in the specification, so keep the file in the root.
If you want a model to be able to take the full text in one request, also add llms-full.txt: this companion file embeds the content of the pages in a single document so the model can process a whole section without crawling page by page. For most corporate sites that is overkill.
What mistakes do people make with llms.txt?
The most common mistakes are turning the file into a copy of the sitemap, stuffing it with keywords and expecting it to replace the rest of the optimisation. Each makes the file less useful and adds no visibility.
- A file with 400 links. It becomes a sitemap with extra steps. The point of the file is selection, not completeness.
- Keywords in descriptions. Search Engine Land's data shows that "search" versions of the files got no citations at all. A model has no use for semantic noise.
- Broken links and redirects. The agent follows a link and runs into a 404 or a redirect chain.
- A separate version "for bots". If the file's content differs substantially from what people see, it comes close to cloaking. A public index of pages is fine; a hidden version is not.
- Contradicting robots.txt. If you block GPTBot, ClaudeBot or OAI-SearchBot, the file simply will not be read by those bots. Decide first which bots you let in.
- A forgotten file. After the site structure changes, the list goes stale. Add a review of the file to your release checklist.
How do you check that llms.txt works?
Check three things: that the file returns a 200 code, that all links are live, and that bots actually request it. The last one is visible only in server logs: look for requests to /llms.txt and see which user agents appear.
A typical check goes like this:
- open
yourdomain/llms.txtin a browser and check the display and encoding; - run the list of links through any crawler to catch 404s and redirects;
- filter server logs by the
/llms.txtpath and group by user agent.
There is also an automatic check: according to the format's author, Lighthouse in Chrome checks for llms.txt as part of its "agentic browsing" audits. But remember that requests from auditors do not mean models use the file: according to Ahrefs, a sizeable share of the requests are exactly that.
What should you do for GEO instead of chasing llms.txt?
Google's guide offers the same principles as for classic SEO: original, useful content, a clear structure and semantic HTML markup for people. Giving the answer right under the heading and backing it with data works well. That is what produces material AI systems can cite.
From Google's point of view, optimising for generative search is optimising the search experience, which is to say still SEO. For other AI systems such as ChatGPT or Perplexity the rules are not documented as clearly, but the practical base is the same: an indexed page, clear text, access permitted in robots.txt and brand mentions on third-party sites.
So the order of priorities is this: indexing and content quality first, then markup and authority, and only after that small experiments like llms.txt. For tools that help track citations and optimise text, read our review of AI tools for content.

FAQ
Can llms.txt hurt SEO?
No, as long as the file is public and contains only useful links. It does not affect indexing and does not change crawling rules. The risk appears when the file contains content that differs from the site (which looks like cloaking), or when it has broken links.
How is llms.txt different from llms-full.txt?
llms.txt is a short catalogue of links and descriptions. llms-full.txt embeds the full text of the pages in a single document so the model gets the context in one request. For large documentation that is convenient; for a small service site it is usually unnecessary.
Can I block AI bots from crawling my site through llms.txt?
No. The file has no access directives; it is just a list of recommended pages. To restrict crawling use robots.txt, where you set rules for GPTBot, ClaudeBot, OAI-SearchBot, Google-Extended and other bots.
What language should llms.txt be written in?
Write the descriptions in the language of the pages themselves. The titles and summary are best given in the language of your main audience. The specification does not regulate language, so this is a matter of consistency with the site's content, not a technical requirement.
How often should I update the file?
Whenever the structure of the key sections changes: a new service, a moved page, a deleted article. For a small site, a review once a quarter and after every URL change is enough.
Practical summary: what to do next
llms.txt is a cheap experiment, not a strategy. Make it in an hour if you have a clear list of key pages, and do not expect more citations: the available data does not support that. The real impact on visibility in AI answers comes from indexing, the quality and structure of content, and brand authority.
If you want to work on your presence in AI search systematically, start with an audit of these basics. The LuchanLabs team can help with AI search visibility and SEO promotion: from the technical base to content that gets cited.
Sources
- Howard J. /llms.txt — a proposal to provide information to help LLMs use websites, Answer.AI, 2024
- The llms.txt specification (v2), llmstxt.org
- Practical Ecommerce. Google says AI optimization is just SEO (review of Google's guide of 15.05.2026)
- Search Engine Journal. 97% of llms.txt files got no requests, Ahrefs data shows
- Search Engine Journal. LLMs.txt shows no clear effect on AI citations, based on 300k domains
- SE Ranking. Does LLMs.txt impact your AI visibility and citations?
- Search Engine Land. Why LLM-only pages aren't the answer to AI search, 2026
- Limy. LLMs.txt in 2026: The Full Guide
- NoHacks. The AI User-Agent Landscape in 2026
About the author
The LuchanLabs Team
SEO and AIO strategists
We rank websites at the top of Google and make brands visible inside ChatGPT, Gemini and Perplexity answers. We write about what we test on client projects every day.
All blog articlesRelated articles
Content marketingTop 5 AI Tools for Writing Content in 2026
The AI writing assistant market is heading for $4.88bn, so the question is no longer whether to use one, but which. Here are five proven tools with real pricing, genuine strengths and the limits their landing pages never mention.
The LuchanLabs Team9 min read