llms.txt: the standard for being cited by ChatGPT and Perplexity
The llms.txt file is a simple convention: a Markdown file placed at the root of your site that presents its content in a format an AI can read. The goal is to help models such as ChatGPT, Perplexity, or Claude understand your site and, ideally, cite it correctly. This guide covers the exact specification, the difference from llms-full.txt, the step-by-step setup, and what it really changes for visibility in AI responses (GEO)—without overselling it.
#Why llms.txt exists
A modern web page is a nightmare for a machine to parse: menus, cookie banners, scripts, ad blocks, and deeply nested HTML. Humans filter all that out without thinking; a language model has to guess where the useful content is, and its context window is limited. As a result, when an AI reads your page, a large part of its token budget goes to noise.
llms.txt addresses this problem. It's a Markdown-formatted text file served at the root of your domain (at /llms.txt) that gives an AI a clean, concise map of your site: what it's about and which pages to visit for details. The idea is directly inspired by robots.txt (which tells crawlers where they can go) and sitemap.xml (which lists URLs), but designed for an LLM to read rather than a conventional indexing bot.
#The llms.txt specification
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The format is intentionally minimal and relies entirely on standard Markdown, in a precise order so that a program can parse it:
- An H1 heading
- The name of the site or project. This is the only required element of the specification.
- A blockquote
- A summary sentence (preceded by >) describing the site. Strongly recommended: it is often the first thing the AI reads.
- Free-form paragraphs
- Zero or more untitled text blocks to provide context: who the site is for, how to read it, and any conventions.
- H2 sections
- Each H2 groups a list of Markdown links to your important pages, in the format [Title](url): an optional note describing the page.
- An “Optional” section
- An H2 named exactly Optional, with its links marked as secondary: a rushed AI can ignore them to save context.
#llms.txt vs llms-full.txt
The two files serve different purposes, and there is no requirement to have both.
- llms.txt
- An index. Short and structured, it lists your pages with brief descriptions. The AI reads it to understand the site architecture, then fetches the pages it’s interested in. This is the file to create first.
- llms-full.txt
- The site's (or documentation's) complete content concatenated into one large Markdown file. The AI can ingest it in one block without making additional requests. Useful for technical documentation you want to load in its entirety into a long context.
In practice: llms.txt for an editorial site such as a blog or guide catalog; llms-full.txt as a supplement if your content is documentation that should be possible to “paste” in full into an LLM. Note that llms-full.txt can quickly grow to several megabytes—generate it only if it makes sense for your use case.
#Create your file step by step
- 01List the pages that really matterDo not include everything. Select the 10 to 50 pages that carry your value: cornerstone guides, product pages, and key documentation. The goal is a signal-dense index, not a copy of the sitemap.
- 02Write the headerAn H1 with the site name, followed by a one-sentence blockquote that precisely summarizes what visitors can find on your site. Be factual: AI will often reuse this sentence verbatim to introduce your site.
- 03Group links by H2 sectionOne section per topic (for example, “Installation,” “Hardware,” and “RAG”). For each link, add a short note after the colon explaining what the page provides. These notes help the AI choose the right page.
- 04Add an Optional sectionMove service pages there (legal notices, contact, about) so the AI knows they are secondary.
- 05Validate the MarkdownMake sure it is valid Markdown and that all links are absolute and work. A syntax error defeats the purpose of the file.
#Deploy on a static site or WordPress
The file must be accessible at the exact address https://votre-domaine.fr/llms.txt, served as plain text (Content-Type text/plain or text/markdown). The method depends on your hosting provider.
On a static site (Hugo, Astro, Eleventy, or a simple HTML directory), it's trivial: place the llms.txt file at the root of the published directory (often public/ or static/), alongside robots.txt. It will be served automatically.
On WordPress, there is no dedicated field. Three options: upload the file to the root via FTP/SFTP (next to wp-config.php); use a recent SEO plugin (several, including versions of Yoast and Rank Math, can generate an llms.txt); or add a rule that serves the file from an endpoint. FTP remains the simplest and most predictable approach.
#GEO: getting cited by AI answer engines
GEO (Generative Engine Optimization) is to ChatGPT, Perplexity, or Google AI Overviews what SEO is to traditional search: the set of practices for appearing — and being cited — in AI-generated answers. The fundamental change is simple: on an answer engine, users read a summary and click much less often. What matters is no longer just ranking on the first page, but being the source the AI uses and attributes.
llms.txt follows the same logic: giving models clean access to well-structured, clearly described content. But let's be precise about the chain of causality, because that's where many articles exaggerate.
- What actually helps GEO
- Factual content structured into clear sections, with direct answers at the top of the page and dated, sourced data. Answer engines reward that, llms.txt or not.
- The role of llms.txt
- Make your site’s structure easier to understand and provide descriptions that AI can reuse. It’s a machine-readability accelerator, not a magical ranking signal.
#What llms.txt does not do (let’s be honest)
Two structural limitations to keep in mind. First, llms.txt blocks no one and forces no one: a crawler that does not know about it will continue reading your HTML normally—for controlling AI bot access, configure robots.txt (and AI crawler user-agent directives), not llms.txt. Second, a file can never make up for weak content: if your pages are vague or outdated, a polished index will not make them citable.
So why deploy it? Because the cost is nearly zero, the exercise forces you to clarify your site’s architecture and descriptions—which also helps human readers and traditional SEO—and if adoption takes off, you’ll already be ready. It’s low-cost insurance, not a magic wand.
#Field report: quelllm.fr deploys it
We’ve been publishing an llms.txt file at the root of quelllm.fr for several months. In practice, it’s generated at build time from our guide catalog: each guide becomes a line, grouped by category (Installation, Hardware, RAG…), with its description taken from the meta field. This ensures the file never becomes outdated as long as a guide exists.
- What we observed
- It’s impossible to cleanly attribute an increase in citations to the file alone: too many variables are changing at the same time (new content, visibility, search engine evolution). So we won’t claim a quantified effect.
- The real measurable benefit
- The generation exercise forced us to write short, accurate, jargon-free descriptions for every guide. This cleanup helped everywhere: meta tags, search previews, and editorial consistency.
- Cost
- Once the generation script is written, there is zero maintenance. This benefit-to-cost ratio is what justifies keeping it, regardless of its exact GEO impact.
Our position: deploy it for cleanliness and anticipation, not because you're promised citations. Put your main effort where it really matters—the quality and structure of the content itself.
#Go further
These guides expand on the concepts covered here, focusing on how AI systems retrieve and cite information:
- Understand RAG
- “What is RAG and how does it work?” explains how AI retrieves external content to answer — the very mechanism llms.txt aims to facilitate.
- Reliability and citations
- “Hallucinations: why your local LLM makes things up and how to limit them” shows why sourcing and structure matter so much for reliable answers.
- Source notebooks
- “NotebookLM locally: open-source alternatives” concretely illustrates the source-based questions and answers mentioned above.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.