An agent forwarded a sales email last week: $900 to add an llms.txt file to their website so AI tools would start recommending them. The file takes minutes to create, Google's documentation says no such file is needed, and a study of 137,210 sites found almost none of them were ever fetched.
What llms.txt was proposed to be
Jeremy Howard of Answer.AI published the proposal on September 3, 2024. The idea was reasonable and modest.
AI models work with a limited amount of text at once. A website's pages are cluttered with navigation, banners, and scripts that waste that space. So the proposal suggested a plain text file at the root of a site, at yoursite.com/llms.txt, listing the pages a model should read and what each one covers, written in a simple format.
Think of it as a short reading list for a machine. The published specification asks for the site name as a heading, a one-line summary, and sections of links with brief descriptions. Some sites also publish a larger file containing the full text of their documentation.
Two things about this deserve emphasis. It was offered as a proposal, not as a standard any company had agreed to support. And it was aimed mainly at technical documentation, where an AI coding assistant needs to read a software library's docs, which is a different situation from a home buyer asking which agent to call.
What Google says
Google's documentation on its AI features is unusually direct. "You don't need to create new machine readable files, AI text files, or markup to appear in these features," it states, and adds that "There's also no special schema.org structured data that you need to add," per Google Search Central.
The same page says there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."
What Google does list is ordinary. Your pages need to be indexed and allowed to show snippets. Crawling should be permitted in your robots.txt file. Internal links should let a crawler move around the site. Important content should be in text rather than locked inside images or video. Structured data should match what a visitor actually sees. Business Profile and product information should be current.
That is the same list that has applied to normal search results for years. Google Search does not use llms.txt as a ranking signal or as an input to its AI features.
What the measured data shows
Opinions about llms.txt are plentiful. Measurements are rarer, which makes the Ahrefs study the most useful evidence available.
Ahrefs examined 137,210 domains using their web analytics data for May 2026. Of those, 28 percent, about 38,360 sites, published a valid llms.txt file. Then they looked at whether anything actually requested those files.
97 percent of them received zero requests. Nothing fetched them at all during the month.
Of the 3 percent that were fetched, about 1,100 domains, 96 percent of requests came from bots rather than people. Breaking those bots down, AI systems of all kinds accounted for 19.5 percent of requests. The remainder were SEO audit tools at 21.7 percent, unidentified requests at 14.9 percent, general web crawlers at 13.1 percent, technology profiling tools at 11.6 percent, and tools studying llms.txt itself at 12.1 percent. Search Engine Journal covered the same findings.
Put plainly: most of these files are read by nobody, and a large share of the reading that does happen is audit software checking whether the file exists.
As for the platforms themselves, no major answer engine has published a commitment to use llms.txt when deciding which sources to cite in consumer answers. Some companies reference the format in developer documentation settings, where an AI coding tool reads technical docs. That is a genuine use, and it is not the same as an AI assistant choosing which real estate agent to name.
Why it is being sold to agents anyway
Three reasons, and none of them are about the file working.
It is fast to deliver. A vendor writes a text file, uploads it, and invoices. There is no content to research and nothing to maintain.
It cannot be quickly disproved. AI citations appear slowly, vary between users, and change with how a question is phrased. If an agent asks in three months whether the file helped, the honest answer is hard to establish either way, which suits a seller.
It sounds technical. "We optimised your site for large language models" lands better than "we wrote four pages about local rental rules," even though the second one is the work that matters.
The pattern is familiar. Every time search behaviour shifts, a simple technical artifact gets sold as the key to it. Ask any vendor offering this what they will measure, when they will measure it, and whether the price includes changes to your actual pages. Our post on free AI visibility tools covers similar claims in the measurement category.
What genuinely affects AI citation
The mechanism is less mysterious than the marketing suggests, and it comes down to four things.
The crawler has to reach your pages. If your robots.txt blocks AI crawlers, or your content only appears after a script runs, or your neighbourhood information lives inside an image of a chart, none of it can be used. This is the most common real problem on agent sites and the least discussed.
The answer has to be stated directly, in text. An assistant quoting you needs a passage it can lift. A page that says "monthly fees in these buildings typically include water, building insurance, and landscaping" gives it that. A page that gestures at the topic across five paragraphs does not. Our post on answer engine optimization covers how to write passages that can be quoted.
Structured data should match what a visitor sees. Google's guidance is that structured data must reflect the visible content. It helps a machine understand what a page is, and it does not create eligibility on its own.
The facts have to be ones nobody else published. This is the part agents control and undervalue. An assistant does not need your page to define a mortgage. It needs a source when someone asks what a specific city requires this year, what recent sales on a named street closed at, or what a particular building's documents say. General content is already inside the model. Local current facts are not. Our post on leads from AI search walks through how those citations turn into contacts.
The honest verdict
llms.txt is harmless. It sits at your site's root, costs nothing to host, and exposes nothing a crawler could not already find, since it lists pages that are already public.
It is also low priority. The measured evidence says almost nothing reads it, no major answer engine has committed to using it for citations, and Google states plainly that no AI text file is required.
So the practical position: if you or your web person want to add one, spend the twenty minutes and move on. If you already have one, leave it. If somebody is charging you for it as an AI visibility service, that price is buying the wrong thing.
The risk worth naming is substitution. An agent who adds the file and believes the AI visibility problem is handled has stopped at the step that does nothing. The pages that answer real local questions are still unwritten, and those are what an assistant needs.
It is possible this changes. If a major platform announces it reads these files and uses them to pick sources, the calculation shifts and it takes minutes to act on. Until that announcement exists, plan around what is documented today.
What to do with the budget instead
Same money, better return, in rough order.
Write three pages that answer local questions nobody in your market has answered in plain text. Rules that changed this year. What fees cover in the building types you sell. What a specific street's recent sales looked like, with dates.
Check that AI crawlers can reach your site at all, and that your key content is text rather than an image or a video with no transcript.
Make sure your Google Business Profile is complete and your reviews are current, since local answers lean on that information.
Then measure it the only way that works reliably. Write down the ten questions your clients actually ask, run them through the assistants your clients use, and record whether you appear. Repeat monthly. It is manual and it beats any dashboard, because these answers vary by person and phrasing.
Our post on the AI search visibility gap covers why most agent sites are invisible to these tools in the first place, and it has nothing to do with a missing file.
The takeaway
llms.txt is a reasonable proposal from September 2024 that the major platforms have not adopted for deciding citations. Google's documentation says no AI text file is needed for its AI features, and Ahrefs found 97 percent of published files received zero requests across 137,210 domains in May 2026. Add one if you like, since it is free and harmless. Do not pay for it, and do not let it stand in for the work that decides whether an assistant can quote you: pages a crawler can reach, answers stated plainly in text, and local facts that exist nowhere else.



