What is llms.txt and how does it help AI search engines?
llms.txt is a plain-text file served from the root of your domain (https://tudominio.com/llms.txt). It holds a structured description of your business, products, pricing and policies, written specifically for generative AI engines (ChatGPT, Claude, Perplexity) to read. Think of it as the modern equivalent of robots.txt, but for LLMs.
The standard was proposed by Jeremy Howard (Answer.AI) in 2024 and is becoming established as a GEO best practice. The typical structure is hierarchical markdown: an H1 with the company name, then sections covering services, pricing, contact details, legal identity and links to more detailed pages. The idea is to give the AI a condensed view of "who you are and what you offer" so it doesn't have to crawl your whole website.
Unlike a sitemap (which is for search engine crawlers), llms.txt is designed to be interpreted by language models during fine-tuning, indexing or real-time answering (when the LLM's bot crawls a site to respond to a user).
An applied example
STAKKER SYSTEMS serves its llms.txt at https://stakker.es/llms.txt, with its products, reference prices, ideal customer profile, declared tech stack, contact details, tax information and key URLs. When ChatGPT gets a question about "AI receptionists in Málaga", the bot can read that llms.txt as a concise source and cite STAKKER with the correct information.
When it is worth it
- You have a brand or product that you want LLMs to describe correctly.
- Your pricing and services change rarely, so you can keep them up to date in a single file.
- You compete in sectors where people research with AI before buying.
- Your site is static or easy to rebuild, which makes updating llms.txt trivial.
Common mistakes
- Filling the file with marketing copy instead of verifiable facts. LLMs mark down hype and favour concrete data.
- Letting llms.txt and your public website contradict each other. If the figures don't match, the AI loses trust in you.
- Blocking AI bots in robots.txt while publishing llms.txt. It's a contradiction that reduces how much of your site gets indexed.
- Forgetting to include your legal entity (registered name, tax ID, registered address). It builds trust with both LLMs and users.
- Not updating llms.txt when your pricing changes. LLMs crawl periodically and will quote outdated data.
How we do it at STAKKER
STAKKER SYSTEMS publishes its own llms.txt as part of the GEO Foundation add-on. It includes a consistency check against your public website, a complementary ai-instructions.md file and a robots.txt optimised for the main AI bots (GPTBot, ClaudeBot, PerplexityBot, etc.). We also set this up for clients who sign up for GEO.
Frequently asked questions
Is llms.txt an official standard?
It's a proposal from Jeremy Howard (2024) and there's no formal body behind it yet. In practice, most serious LLMs respect it and plenty of technical websites already publish one.
Do I have to update llms.txt manually?
Yes, unless you have a build script that regenerates it from a single source. STAKKER regenerates it automatically from data/pricing.ts and its services database.
Is there a difference between llms.txt and ai-instructions.md?
Yes. llms.txt is the specific file in your root that follows the standard's convention. ai-instructions.md is an optional add-on with more technical instructions aimed at LLMs (how they should describe you, what to avoid). STAKKER publishes both.
If I publish llms.txt, will I automatically appear in ChatGPT?
No, but it raises your chances considerably. ChatGPT needs to crawl your site (which depends on SEO, authority and inbound links), understand your niche and consider you reliable. llms.txt is just one of several signals, but an important one.
Related terms
- GEOGEO (Generative Engine Optimization) is the practice of optimising your content, schema and entity signals so that generative AI engines such as ChatGPT, Claude and Perplexity cite you when answering questions in your sector.
- RAGRAG (Retrieval-Augmented Generation) is a technique that combines a language model (LLM) with a search system over your own knowledge base. Before answering, the system looks for relevant documents in your content (FAQs, products, pricing, contracts) and hands them to the LLM as context, so it answers from verifiable data rather than whatever the model remembers from its training.