Discoverability & AEOOctober 6, 202610 min read

What Is llms.txt—and Should Your Website Have One?

llms.txt is a plain-text Markdown file at a website's root that gives an AI a curated map of its most useful pages. It was proposed to help an AI work with a site it has already reached, not to win citations: Ahrefs found 97% of published files received zero requests in May 2026.

An llms.txt file was never built to get your business cited by ChatGPT. Jeremy Howard proposed it in 2024 with the idea that an AI already working with a site could find the right pages quickly. Developer documentation sites adopted it first, and other sites followed, though no major AI company has said it reads the file. Ahrefs watched roughly 38,000 published files for a month in 2026, and 97% were never requested once.

What is llms.txt?

llms.txt is a plain-text file, written in Markdown (plain text with simple formatting marks), placed at a website's root (yoursite.com/llms.txt). Jeremy Howard, co-founder of Answer.AI, published the proposed standard on 3 September 2024. The file is a short, curated map of a site: what it is, and which pages are worth reading. The problem it addressed was practical. A page wraps its information in navigation, ads and JavaScript, converting that back into clean text is difficult and imprecise, and a model's context window (how much text it can read at once) is still too small for a whole site (llmstxt.org). The only required element in an llms.txt file is a single H1 with the site's name. A short blockquote summary and sections of links are optional.

It spread through documentation first. Mintlify switched it on for every docs site it hosts in November 2024, and Anthropic's own documentation publishes one, alongside a full-text companion, llms-full.txt, that Mintlify developed with Anthropic.

What llms.txt does and doesn't do

It was written for one job: helping an AI that has already reached a site, or a tool working through documentation, find the right pages without wading through clutter. It was never designed to decide which site an AI sends someone to. That mismatch is a source of the confusion. Ask "Will it get me cited?" and you're setting it a different job from the one it was built for.

On the question people do ask, the evidence points one way.

Google Search doesn't use it. Google's AI search optimisation guide, published in May 2026, lists llms.txt among the tactics you don't need, and Gary Illyes has confirmed Google isn't pursuing it (Search Engine Journal, May 2026).

The limit is built in. On Google's Search Off the Record podcast (episode 111, 15 June 2026), John Mueller said an AI system can't use llms.txt to choose between sites: the file is a site describing itself, and every site can describe itself as the best. He compared it to the keywords meta tag, which stopped working for the same reason (Search Engine Journal, June 2026).

AI systems aren't fetching it. Ahrefs analysed server logs for 137,210 domains in May 2026. Of the roughly 38,000 with a valid file, 97% received zero requests. Named AI bots made 19.5% of the requests that did arrive; AI retrieval bots, the ones that fetch pages to answer a question, about 1%. Slackbot fetched the files more often than PerplexityBot. No AI bot requested /llms.txt on sites that don't have one (Ahrefs, June 2026). Ahrefs measured requests, not whether anything acted on what it fetched.

It doesn't track with citations. SE Ranking compared about 300,000 domains against how often each was cited in AI answers and found no relationship with having the file. Removing it from their prediction model made the model more accurate (SE Ranking, November 2025, reported by Search Engine Journal).

One narrow use survives. Mueller said the file might help an agent that's already on your site finish a task faster. That's real, and small.

Who is paying attention to llms.txt

The evidence looks contradictory until you separate the audiences.

AI search and answer crawlers: barely reading it, per the logs above.

Coding assistants and documentation: real use. Documentation platforms generate the file by default, and the fetch data is in the llms.txt vs llms-full.txt section below.

Browser agents: experimental. Chrome's Lighthouse 13.3 added an Agentic Browsing category in May 2026 that checks whether a site provides the file. It's framed around whether a browser agent can navigate efficiently, not around search, and Lighthouse treats the file as optional. In Ahrefs' data the audit generated about 22 requests, roughly 1 in 1,000.

Google disagrees with itself here. Search says skip it. Chrome's Lighthouse, which shipped days earlier, checks for it. And in December 2025 an llms.txt file briefly appeared on Google's own Search Central developer documentation (Search Engine Journal, May 2026).

How many sites have one

Adoption depends on who you ask.

Ahrefs, May 2026

Sample 137,210 domains using Ahrefs Web Analytics

Sites with llms.txt 28% (a technical-leaning sample)

SE Ranking, November 2025

Sample about 300,000 domains

Sites with llms.txt 10.13%

ProGEO.ai, March 2026

Sample Fortune 500

Sites with llms.txt 7.4% (37 companies)

Originality.ai, May 2026

Sample 3M+ sites tracked

Sites with llms.txt 36,120 files, up 8.8x in 12 months

The range runs from 7.4% of the Fortune 500 to 28% of Ahrefs' analytics customers because the samples differ. Adoption has grown. In Ahrefs' log data, reading barely registered. No major AI company has publicly committed to reading the file in production (Ahrefs, June 2026).

llms.txt vs robots.txt vs sitemap.xml

These three do different jobs.

robots.txt is a gatekeeper that controls access: which crawlers (GPTBot, ClaudeBot and others) are allowed to fetch which pages. (For the full crawler-by-crawler rules, which bots to allow and which to block, see AI Training vs AI Citation, not repeated here.)

sitemap.xml is an index: every URL on the site, so a crawler knows what exists.

llms.txt is neither. It doesn't block or allow anything, and it isn't a complete list. It's a curated summary, closer to a table of contents than a map. The spec describes it as a complement to robots.txt, providing context for allowed content (llmstxt.org).

None of the three guarantees an AI system reads or uses it. And all three are voluntary conventions a well-behaved crawler chooses to respect.

llms.txt vs llms-full.txt

llms.txt is the index, while llms-full.txt carries the content as well. It's one Markdown file holding the full text of the pages the index links to, so a tool can load a whole documentation set in a single fetch instead of following links. The Next.js and Cloudflare documentation each publish one.

It's also the file with fetch data behind it. Mintlify analysed seven days of traffic across 25 companies' documentation and found a median of 14 visits to llms.txt and 79 to llms-full.txt, with ChatGPT driving most of the larger file's traffic (Mintlify, July 2026). That's one vendor's data, on documentation sites only, so don't read that as an extensive study. Ahrefs' study measured the index file and nothing else, so its 97% says nothing about llms-full.txt either way.

For a site that isn't documentation, the same rule applies with a cost on top: llms-full.txt duplicates your whole site into one file that has to be regenerated every time a page changes.

Will llms.txt take off?

As a way of getting chosen, I don't expect it to. The limit Mueller describes isn't about how many sites adopt it. A file a site writes about itself can't help a system pick between sites, and every extra site writing the same kind of file makes the claims look more alike. Where it can last is the job it was built for: an agent that's already on a site, or working through documentation, finding the right pages, faster. Chrome checking for it and documentation platforms generating it both point that way. If you want one because you want to be cited, the file isn't the lever.

How to create an llms.txt file

The llms.txt format is plain CommonMark Markdown:

  • An H1 with your site or project name (the only required line)
  • An optional blockquote with a one- or two-line summary
  • Optional H2 sections grouping links by category, each a Markdown link with an optional short description

A minimal llms.txt example:

# Aimee Q Devlin

> Systems and infrastructure architect. AEO, technical SEO, and infrastructure builds for founder-led businesses.

## Services
- [Infrastructure Audit](https://aimeeqdevlin.com/services/infrastructure-audit): technical diagnosis and fix list
- [Infrastructure Architecture](https://aimeeqdevlin.com/services/infrastructure-architecture): full design and specification

## Insights
- [What Is AEO](https://aimeeqdevlin.com/insights/what-is-aeo): the discoverability layer explained

Keep it curated, not comprehensive. Every link needs to resolve (a 200 status, not a redirect or a 404), and none should point at a page that only renders via JavaScript, since the point is giving a system something it can read without executing a script. The spec goes one step further: the links should point at LLM-friendly content, such as Markdown versions of your pages (llmstxt.org). If you can serve those, do.

If you'd rather not start from a blank file, free llms.txt generators will crawl your site and draft one. Edit the draft down, because a generator lists everything and the point of the file is curation.

Should your website have an llms.txt file?

Have one. It's quick, it's useful if you run documentation, an API or developer tooling, and if the convention takes hold you'll already be there. Don't build one expecting citations, rankings or traffic. Put that effort into schema markup, an entity graph (how your site's structured data connects into one consistent description of the business), crawl access for the AI crawlers that do request pages, and answer-first content structure. See what moves the needle → What Is AEO

How to check yours

Run the AI Crawler Checker to test AI crawler access: whether AI crawlers can reach and read your site. For the file itself, a free third-party llms.txt validator will check it against the spec in seconds: H1 placement, valid Markdown, working links. And to see whether anything is fetching it, look for /llms.txt in your hosting or server request logs. A month of logs will tell you more than any guide, including this one.


Aimee Q Devlin is a Systems and Infrastructure Architect based in San Miguel de Allende, Mexico. She works with founders and operators of established businesses who are ready to rebuild their systems properly—including the infrastructure that makes those systems discoverable. The PRISM Diagnostic is where engagements begin.

What is llms.txt?

llms.txt is a plain-text Markdown file at a website's root that gives AI systems a curated map of the site's most useful pages. It was proposed in 2024 and remains an optional, unofficial convention, not a web standard.

Who created llms.txt?

Jeremy Howard, co-founder of Answer.AI, proposed it on 3 September 2024. It's a community proposal, and no major AI company has publicly committed to reading it in production.

Does llms.txt work?

Not for discovery, ranking or citations, based on current evidence. Ahrefs' May 2026 log study of 137,210 domains found 97% of published files received zero requests. It may help an agent already on your site, or a coding tool working through documentation.

Is llms.txt a ranking signal?

No. Google's AI search guidance lists it among the tactics you don't need, and SE Ranking's analysis of about 300,000 domains found no relationship between having the file and how often a domain is cited in AI answers.

What's the difference between llms.txt and robots.txt?

robots.txt controls which crawlers may access which pages. llms.txt is a curated content summary with no access-control function at all: it doesn't block or allow anything.

What's the difference between llms.txt and a sitemap?

A sitemap.xml lists every URL on a site, built for completeness. llms.txt is a short, curated summary of only the most useful pages, built for a quick read, not a full index.

What is llms-full.txt?

llms-full.txt is the companion to llms.txt that holds the full text of the linked pages in one Markdown file, so a tool can load a documentation set in a single fetch. Documentation platforms generate it, and Ahrefs' large log study measured the index file only.

How do I create an llms.txt file?

Write a plain Markdown file named llms.txt with an H1 for the site name, an optional one-line summary, and sections of links to your most useful pages. Upload it to your site's root so it loads at yoursite.com/llms.txt, and check that every link returns a 200.

Do I need an llms.txt file?

Yes, have one. It's quick, it's useful if you run documentation, an API or developer tooling, and if the convention takes hold you'll already be there. Don't expect it to help your citations, rankings or traffic on its own: schema markup, crawl access and answer-first content come first.

→ The PRISM Diagnostic

→ What Is AEO

→ How to Get Your Business Cited by ChatGPT, Perplexity, and Claude

→ AI Training vs AI Citation

›Sources
Aimee Q Devlin—Systems Architect and infrastructure builder based in San Miguel de Allende, Mexico

Aimee Q Devlin

Aimee Q Devlin is a Systems and Infrastructure Architect based in San Miguel de Allende, Mexico. She works with founders and operators of established businesses whose sites aren't ranking, converting, or being cited by AI—and builds the infrastructure that fixes it properly. She developed the PRISM Framework, an AEO framework for making founder-led businesses visible to ChatGPT, Perplexity, and the AI engines shaping discovery in 2026. The PRISM Diagnostic is where engagements begin.

About Aimee →

Something isn't working. Let's find out what.

You don't need to have it diagnosed before we talk. That's what the PRISM Diagnostic is for. 60 minutes, USD $350.

See the PRISM Diagnostic