The Complete Guide to llms.txt
What is llms.txt?
llms.txt is a plain-markdown file, published at your domain's root, that gives AI engines a short, curated, human-written summary of your site — what it is, who it's for, and links to the pages that matter most, organized by section.
Think of it as the difference between a sitemap.xml (a complete, automated list of every URL) and a table of contents someone wrote by hand to help a new reader find the right starting point. AI engines can use it to quickly understand a site's scope before deciding what to retrieve.
What should actually be in it?
A minimal, well-structured llms.txt has:
# Your Site Name
> One or two sentences describing what the site is and who it's for.
## Core pages
- [Page title](url): one-line description of what it covers
- [Page title](url): one-line description
## Guides
- [Guide title](url): one-line description
## Reference / Glossary
- [Glossary](url): one-line description
The one-line descriptions matter — they're what helps an engine decide whether a page is relevant to a given question without having to fetch and parse it first.
What should you leave out?
Everything that isn't genuinely a top-level entry point: don't list every blog post, every tag page, every paginated archive. That's what sitemap.xml and normal crawling are for. A curated file with 30–80 genuinely important links is more useful than an automated dump of thousands — the curation is the entire value of the file.
What are its real limitations?
Be honest about this, because overselling it undermines the rest of your AEO work:
- It's an emerging convention, not a ratified standard. There's no governing body enforcing it the way there is for robots.txt.
- Not every AI engine reads it yet. Adoption is growing but inconsistent across providers.
- It doesn't replace good on-page structure. An engine that does read your llms.txt still has to retrieve and parse the actual pages — llms.txt helps it decide what to fetch, not what to cite.
Given the cost is close to zero and the downside is none, publishing one is a reasonable, low-effort part of an AEO strategy — just don't treat it as a substitute for the actual content-structure work.
Frequently asked questions
Is llms.txt the same as robots.txt?
No — robots.txt controls crawler access (what's allowed or disallowed); llms.txt is a curated content map, more like a hand-written sitemap aimed at AI engines specifically, describing what your site is and pointing to the pages that matter most.
Where does llms.txt need to live?
At the root of your domain — yoursite.com/llms.txt — the same convention as robots.txt and sitemap.xml, so engines that look for it know exactly where to check.
Does every page need to be listed in llms.txt?
No — that defeats the purpose. It should be curated: your most important, most representative pages per section, not an automated dump of every URL (that's what sitemap.xml is for). A curated 50-line file is more useful to an engine than an automated 5,000-line one.
Keep reading
Or let Omnitopical do all of this.
Everything in this guide, run automatically on one domain — mapped, written, interlinked, and published with 25 machine files. $99/mo.
Map my site's topical authority