The Complete Guide to llms.txt

llms.txt is a plain-markdown file at your site's root that gives AI engines a curated, human-written map of your most important content — a summary of what your site is, plus links to key pages, organized by section. It's an emerging convention, not a ratified standard, and not every AI engine honours it yet, but it costs almost nothing to publish and there's no downside to shipping one.
6 min read · Updated

What is llms.txt?

llms.txt is a plain-markdown file, published at your domain's root, that gives AI engines a short, curated, human-written summary of your site — what it is, who it's for, and links to the pages that matter most, organized by section.

Think of it as the difference between a sitemap.xml (a complete, automated list of every URL) and a table of contents someone wrote by hand to help a new reader find the right starting point. AI engines can use it to quickly understand a site's scope before deciding what to retrieve.

What should actually be in it?

A minimal, well-structured llms.txt has:

# Your Site Name

> One or two sentences describing what the site is and who it's for.

## Core pages
- [Page title](url): one-line description of what it covers
- [Page title](url): one-line description

## Guides
- [Guide title](url): one-line description

## Reference / Glossary
- [Glossary](url): one-line description

The one-line descriptions matter — they're what helps an engine decide whether a page is relevant to a given question without having to fetch and parse it first.

What should you leave out?

Everything that isn't genuinely a top-level entry point: don't list every blog post, every tag page, every paginated archive. That's what sitemap.xml and normal crawling are for. A curated file with 30–80 genuinely important links is more useful than an automated dump of thousands — the curation is the entire value of the file.

What are its real limitations?

Be honest about this, because overselling it undermines the rest of your AEO work:

  • It's an emerging convention, not a ratified standard. There's no governing body enforcing it the way there is for robots.txt.
  • Not every AI engine reads it yet. Adoption is growing but inconsistent across providers.
  • It doesn't replace good on-page structure. An engine that does read your llms.txt still has to retrieve and parse the actual pages — llms.txt helps it decide what to fetch, not what to cite.

Given the cost is close to zero and the downside is none, publishing one is a reasonable, low-effort part of an AEO strategy — just don't treat it as a substitute for the actual content-structure work.

Frequently asked questions

Is llms.txt the same as robots.txt?

No — robots.txt controls crawler access (what's allowed or disallowed); llms.txt is a curated content map, more like a hand-written sitemap aimed at AI engines specifically, describing what your site is and pointing to the pages that matter most.

Where does llms.txt need to live?

At the root of your domain — yoursite.com/llms.txt — the same convention as robots.txt and sitemap.xml, so engines that look for it know exactly where to check.

Does every page need to be listed in llms.txt?

No — that defeats the purpose. It should be curated: your most important, most representative pages per section, not an automated dump of every URL (that's what sitemap.xml is for). A curated 50-line file is more useful to an engine than an automated 5,000-line one.

Keep reading

Or let Omnitopical do all of this.

Everything in this guide, run automatically on one domain — mapped, written, interlinked, and published with 25 machine files. $99/mo.

Map my site's topical authority