Intermediate 10 minGEO

llms.txt: the standard for being cited by ChatGPT and Perplexity

The llms.txt file is a simple convention: a Markdown file placed at the root of your site that presents its content in a format an AI can read. The goal is to help models such as ChatGPT, Perplexity, or Claude understand your site and, ideally, cite it correctly. This guide covers the exact specification, the difference from llms-full.txt, the step-by-step setup, and what it really changes for visibility in AI responses (GEO)—without overselling it.

By Mohamed Meguedmi·Update 2026-09-03·Tested on Windows, macOS, and Linux

#Why llms.txt exists

A modern web page is a nightmare for a machine to parse: menus, cookie banners, scripts, ad blocks, and deeply nested HTML. Humans filter all that out without thinking; a language model has to guess where the useful content is, and its context window is limited. As a result, when an AI reads your page, a large part of its token budget goes to noise.

llms.txt addresses this problem. It's a Markdown-formatted text file served at the root of your domain (at /llms.txt) that gives an AI a clean, concise map of your site: what it's about and which pages to visit for details. The idea is directly inspired by robots.txt (which tells crawlers where they can go) and sitemap.xml (which lists URLs), but designed for an LLM to read rather than a conventional indexing bot.

i
Where the standard comes from
llms.txt was proposed in September 2024 by Jeremy Howard (co-founder of Answer.AI and fast.ai). The reference specification is published at llmstxt.org. It is an open community convention, not a standard imposed by OpenAI or Google.

#The llms.txt specification

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The format is intentionally minimal and relies entirely on standard Markdown, in a precise order so that a program can parse it:

An H1 heading
The name of the site or project. This is the only required element of the specification.
A blockquote
A summary sentence (preceded by >) describing the site. Strongly recommended: it is often the first thing the AI reads.
Free-form paragraphs
Zero or more untitled text blocks to provide context: who the site is for, how to read it, and any conventions.
H2 sections
Each H2 groups a list of Markdown links to your important pages, in the format [Title](url): an optional note describing the page.
An “Optional” section
An H2 named exactly Optional, with its links marked as secondary: a rushed AI can ignore them to save context.
/llms.txt
# QuelLLM.fr

> Guides francophones pour installer et faire tourner des LLM open-weight en local (Ollama, LM Studio, llama.cpp), du choix du matériel au RAG.

Site en français destiné aux utilisateurs qui veulent auto-héberger une IA sans cloud. Chaque guide est autonome et daté.

## Guides essentiels

- [Installer Ollama en 5 minutes](https://quelllm.fr/guides/installer-ollama.html): prérequis RAM/GPU et premier modèle
- [Choisir sa quantification (Q4, Q5, Q8, FP16)](https://quelllm.fr/guides/choisir-quantification-q4-q5-q8.html): compromis qualité/mémoire
- [Le RAG expliqué simplement](https://quelllm.fr/guides/rag-debutant-c-est-quoi-comment-ca-fonctionne.html): discuter avec ses documents

## Matériel

- [Quel LLM sur RTX 4090 ?](https://quelllm.fr/guides/llm-rtx-4090.html): la référence grand public 24 Go

## Optional

- [Mentions légales](https://quelllm.fr/legal.html): informations éditeur
→
Absolute URLs and links that point to clean targets
Always use absolute URLs (https://...): the file may be read out of context. And if your pages have a raw Markdown version (many documentation generators expose /page.md), point to it instead of the HTML: the AI gets more context and you use fewer tokens.

#llms.txt vs llms-full.txt

The two files serve different purposes, and there is no requirement to have both.

llms.txt
An index. Short and structured, it lists your pages with brief descriptions. The AI reads it to understand the site architecture, then fetches the pages it’s interested in. This is the file to create first.
llms-full.txt
The site's (or documentation's) complete content concatenated into one large Markdown file. The AI can ingest it in one block without making additional requests. Useful for technical documentation you want to load in its entirety into a long context.

In practice: llms.txt for an editorial site such as a blog or guide catalog; llms-full.txt as a supplement if your content is documentation that should be possible to “paste” in full into an LLM. Note that llms-full.txt can quickly grow to several megabytes—generate it only if it makes sense for your use case.

#Create your file step by step

  1. 01
    List the pages that really matter
    Do not include everything. Select the 10 to 50 pages that carry your value: cornerstone guides, product pages, and key documentation. The goal is a signal-dense index, not a copy of the sitemap.
  2. 02
    Write the header
    An H1 with the site name, followed by a one-sentence blockquote that precisely summarizes what visitors can find on your site. Be factual: AI will often reuse this sentence verbatim to introduce your site.
  3. 03
    Group links by H2 section
    One section per topic (for example, “Installation,” “Hardware,” and “RAG”). For each link, add a short note after the colon explaining what the page provides. These notes help the AI choose the right page.
  4. 04
    Add an Optional section
    Move service pages there (legal notices, contact, about) so the AI knows they are secondary.
  5. 05
    Validate the Markdown
    Make sure it is valid Markdown and that all links are absolute and work. A syntax error defeats the purpose of the file.
!
Don’t let it become outdated
A frozen llms.txt file that points to dead URLs or describes a site that has changed is worse than no file at all: it gives AI a false picture of you. Regenerate it after every major redesign, ideally through a build script.

#Deploy on a static site or WordPress

The file must be accessible at the exact address https://votre-domaine.fr/llms.txt, served as plain text (Content-Type text/plain or text/markdown). The method depends on your hosting provider.

On a static site (Hugo, Astro, Eleventy, or a simple HTML directory), it's trivial: place the llms.txt file at the root of the published directory (often public/ or static/), alongside robots.txt. It will be served automatically.

Terminal
# Déposer le fichier à la racine du site publié
cp llms.txt ./public/llms.txt

# Vérifier qu'il est bien servi une fois en ligne
curl -I https://votre-domaine.fr/llms.txt
# Attendu : HTTP/2 200 et un Content-Type text/plain ou text/markdown

On WordPress, there is no dedicated field. Three options: upload the file to the root via FTP/SFTP (next to wp-config.php); use a recent SEO plugin (several, including versions of Yoast and Rank Math, can generate an llms.txt); or add a rule that serves the file from an endpoint. FTP remains the simplest and most predictable approach.

→
Report it, like a sitemap
Nothing requires you to declare it, but you can add a line to robots.txt to make it discoverable, following the model of the Sitemap directive. Some teams also add a link in the <head>. These conventions are not yet standardized: placement at the root remains the key point.

#GEO: getting cited by AI answer engines

GEO (Generative Engine Optimization) is to ChatGPT, Perplexity, or Google AI Overviews what SEO is to traditional search: the set of practices for appearing — and being cited — in AI-generated answers. The fundamental change is simple: on an answer engine, users read a summary and click much less often. What matters is no longer just ranking on the first page, but being the source the AI uses and attributes.

llms.txt follows the same logic: giving models clean access to well-structured, clearly described content. But let's be precise about the chain of causality, because that's where many articles exaggerate.

What actually helps GEO
Factual content structured into clear sections, with direct answers at the top of the page and dated, sourced data. Answer engines reward that, llms.txt or not.
The role of llms.txt
Make your site’s structure easier to understand and provide descriptions that AI can reuse. It’s a machine-readability accelerator, not a magical ranking signal.

#What llms.txt does not do (let’s be honest)

!
No major crawler has confirmed using it as a signal
As of today, neither OpenAI, Perplexity, nor Anthropic has officially announced that it reads llms.txt to rank or cite sites, and Google has publicly stated that it does not use it for search. Be wary of any article promising “X% more citations thanks to llms.txt”: this cannot be verified today.

Two structural limitations to keep in mind. First, llms.txt blocks no one and forces no one: a crawler that does not know about it will continue reading your HTML normally—for controlling AI bot access, configure robots.txt (and AI crawler user-agent directives), not llms.txt. Second, a file can never make up for weak content: if your pages are vague or outdated, a polished index will not make them citable.

So why deploy it? Because the cost is nearly zero, the exercise forces you to clarify your site’s architecture and descriptions—which also helps human readers and traditional SEO—and if adoption takes off, you’ll already be ready. It’s low-cost insurance, not a magic wand.

#Field report: quelllm.fr deploys it

We’ve been publishing an llms.txt file at the root of quelllm.fr for several months. In practice, it’s generated at build time from our guide catalog: each guide becomes a line, grouped by category (Installation, Hardware, RAG…), with its description taken from the meta field. This ensures the file never becomes outdated as long as a guide exists.

What we observed
It’s impossible to cleanly attribute an increase in citations to the file alone: too many variables are changing at the same time (new content, visibility, search engine evolution). So we won’t claim a quantified effect.
The real measurable benefit
The generation exercise forced us to write short, accurate, jargon-free descriptions for every guide. This cleanup helped everywhere: meta tags, search previews, and editorial consistency.
Cost
Once the generation script is written, there is zero maintenance. This benefit-to-cost ratio is what justifies keeping it, regardless of its exact GEO impact.

Our position: deploy it for cleanliness and anticipation, not because you're promised citations. Put your main effort where it really matters—the quality and structure of the content itself.


#Go further

These guides expand on the concepts covered here, focusing on how AI systems retrieve and cite information:

Understand RAG
“What is RAG and how does it work?” explains how AI retrieves external content to answer — the very mechanism llms.txt aims to facilitate.
Reliability and citations
“Hallucinations: why your local LLM makes things up and how to limit them” shows why sourcing and structure matter so much for reliable answers.
Source notebooks
“NotebookLM locally: open-source alternatives” concretely illustrates the source-based questions and answers mentioned above.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.