AI Search

    Does an llms.txt File Help a Real Estate Website?

    Saige Team·October 6, 2026·12 min read
    Does an llms.txt File Help a Real Estate Website?

    An agent forwarded a sales email last week: $900 to add an llms.txt file to their website so AI tools would start recommending them. The file takes minutes to create, Google's documentation says no such file is needed, and a study of 137,210 sites found almost none of them were ever fetched.

    What llms.txt was proposed to be

    Jeremy Howard of Answer.AI published the proposal on September 3, 2024. The idea was reasonable and modest.

    AI models work with a limited amount of text at once. A website's pages are cluttered with navigation, banners, and scripts that waste that space. So the proposal suggested a plain text file at the root of a site, at yoursite.com/llms.txt, listing the pages a model should read and what each one covers, written in a simple format.

    Think of it as a short reading list for a machine. The published specification asks for the site name as a heading, a one-line summary, and sections of links with brief descriptions. Some sites also publish a larger file containing the full text of their documentation.

    Two things about this deserve emphasis. It was offered as a proposal, not as a standard any company had agreed to support. And it was aimed mainly at technical documentation, where an AI coding assistant needs to read a software library's docs, which is a different situation from a home buyer asking which agent to call.

    What Google says

    Google's documentation on its AI features is unusually direct. "You don't need to create new machine readable files, AI text files, or markup to appear in these features," it states, and adds that "There's also no special schema.org structured data that you need to add," per Google Search Central.

    The same page says there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary."

    What Google does list is ordinary. Your pages need to be indexed and allowed to show snippets. Crawling should be permitted in your robots.txt file. Internal links should let a crawler move around the site. Important content should be in text rather than locked inside images or video. Structured data should match what a visitor actually sees. Business Profile and product information should be current.

    That is the same list that has applied to normal search results for years. Google Search does not use llms.txt as a ranking signal or as an input to its AI features.

    What the measured data shows

    Opinions about llms.txt are plentiful. Measurements are rarer, which makes the Ahrefs study the most useful evidence available.

    Ahrefs examined 137,210 domains using their web analytics data for May 2026. Of those, 28 percent, about 38,360 sites, published a valid llms.txt file. Then they looked at whether anything actually requested those files.

    97 percent of them received zero requests. Nothing fetched them at all during the month.

    Of the 3 percent that were fetched, about 1,100 domains, 96 percent of requests came from bots rather than people. Breaking those bots down, AI systems of all kinds accounted for 19.5 percent of requests. The remainder were SEO audit tools at 21.7 percent, unidentified requests at 14.9 percent, general web crawlers at 13.1 percent, technology profiling tools at 11.6 percent, and tools studying llms.txt itself at 12.1 percent. Search Engine Journal covered the same findings.

    Put plainly: most of these files are read by nobody, and a large share of the reading that does happen is audit software checking whether the file exists.

    As for the platforms themselves, no major answer engine has published a commitment to use llms.txt when deciding which sources to cite in consumer answers. Some companies reference the format in developer documentation settings, where an AI coding tool reads technical docs. That is a genuine use, and it is not the same as an AI assistant choosing which real estate agent to name.

    Why it is being sold to agents anyway

    Three reasons, and none of them are about the file working.

    It is fast to deliver. A vendor writes a text file, uploads it, and invoices. There is no content to research and nothing to maintain.

    It cannot be quickly disproved. AI citations appear slowly, vary between users, and change with how a question is phrased. If an agent asks in three months whether the file helped, the honest answer is hard to establish either way, which suits a seller.

    It sounds technical. "We optimised your site for large language models" lands better than "we wrote four pages about local rental rules," even though the second one is the work that matters.

    The pattern is familiar. Every time search behaviour shifts, a simple technical artifact gets sold as the key to it. Ask any vendor offering this what they will measure, when they will measure it, and whether the price includes changes to your actual pages. Our post on free AI visibility tools covers similar claims in the measurement category.

    What genuinely affects AI citation

    The mechanism is less mysterious than the marketing suggests, and it comes down to four things.

    The crawler has to reach your pages. If your robots.txt blocks AI crawlers, or your content only appears after a script runs, or your neighbourhood information lives inside an image of a chart, none of it can be used. This is the most common real problem on agent sites and the least discussed.

    The answer has to be stated directly, in text. An assistant quoting you needs a passage it can lift. A page that says "monthly fees in these buildings typically include water, building insurance, and landscaping" gives it that. A page that gestures at the topic across five paragraphs does not. Our post on answer engine optimization covers how to write passages that can be quoted.

    Structured data should match what a visitor sees. Google's guidance is that structured data must reflect the visible content. It helps a machine understand what a page is, and it does not create eligibility on its own.

    The facts have to be ones nobody else published. This is the part agents control and undervalue. An assistant does not need your page to define a mortgage. It needs a source when someone asks what a specific city requires this year, what recent sales on a named street closed at, or what a particular building's documents say. General content is already inside the model. Local current facts are not. Our post on leads from AI search walks through how those citations turn into contacts.

    The honest verdict

    llms.txt is harmless. It sits at your site's root, costs nothing to host, and exposes nothing a crawler could not already find, since it lists pages that are already public.

    It is also low priority. The measured evidence says almost nothing reads it, no major answer engine has committed to using it for citations, and Google states plainly that no AI text file is required.

    So the practical position: if you or your web person want to add one, spend the twenty minutes and move on. If you already have one, leave it. If somebody is charging you for it as an AI visibility service, that price is buying the wrong thing.

    The risk worth naming is substitution. An agent who adds the file and believes the AI visibility problem is handled has stopped at the step that does nothing. The pages that answer real local questions are still unwritten, and those are what an assistant needs.

    It is possible this changes. If a major platform announces it reads these files and uses them to pick sources, the calculation shifts and it takes minutes to act on. Until that announcement exists, plan around what is documented today.

    What to do with the budget instead

    Same money, better return, in rough order.

    Write three pages that answer local questions nobody in your market has answered in plain text. Rules that changed this year. What fees cover in the building types you sell. What a specific street's recent sales looked like, with dates.

    Check that AI crawlers can reach your site at all, and that your key content is text rather than an image or a video with no transcript.

    Make sure your Google Business Profile is complete and your reviews are current, since local answers lean on that information.

    Then measure it the only way that works reliably. Write down the ten questions your clients actually ask, run them through the assistants your clients use, and record whether you appear. Repeat monthly. It is manual and it beats any dashboard, because these answers vary by person and phrasing.

    Our post on the AI search visibility gap covers why most agent sites are invisible to these tools in the first place, and it has nothing to do with a missing file.

    The takeaway

    llms.txt is a reasonable proposal from September 2024 that the major platforms have not adopted for deciding citations. Google's documentation says no AI text file is needed for its AI features, and Ahrefs found 97 percent of published files received zero requests across 137,210 domains in May 2026. Add one if you like, since it is free and harmless. Do not pay for it, and do not let it stand in for the work that decides whether an assistant can quote you: pages a crawler can reach, answers stated plainly in text, and local facts that exist nowhere else.

    Plot shares general guidance for real estate agents and brokers. It is not individualized business, financial, or legal advice for your specific situation.

    Frequently asked questions

    A marketing company quoted me $900 to add an llms.txt file so ChatGPT recommends me. Should I pay it?

    No. The file is a plain text list of your pages that you or your web person can add in a few minutes at no cost. More importantly, no major AI platform has publicly committed to using it to decide who gets cited in answers, and Google's documentation says no AI text file is needed for its AI features. If the quote is for AI visibility, ask what else is included, because the file itself is not the work.

    I already added llms.txt to my site. Should I remove it now?

    There is no need to remove it. The file sits at your site's root, costs nothing to host, and does no harm. Treat it the way you would treat a tidy folder structure: reasonable housekeeping with no traffic attached. What matters is that adding it does not replace the work that affects citations, which is the content on your actual pages.

    If Google ignores llms.txt, does that mean AI Overviews ignore my site's structure entirely?

    Structure still matters, just through normal pages rather than a special file. Google's documentation lists what helps: allowing crawling, good internal linking, keeping important content in text form, and structured data that matches what visitors see. Those are the same fundamentals that affect ordinary search results. The file is the part that does nothing, while the page structure is doing real work.

    Has any AI company actually said they read llms.txt?

    Not for deciding what to cite in consumer answers. Some companies reference the format in developer documentation contexts, where an AI coding tool is reading technical docs, and that is a different use case from a home buyer asking for an agent recommendation. As of this writing, no major answer engine has published a statement that llms.txt influences which sources it quotes. Treat any vendor claim otherwise as something to ask for evidence on.

    What does the research actually show about whether these files get read?

    Ahrefs analysed 137,210 domains and found that 97 percent of published llms.txt files received zero requests during May 2026, meaning nothing fetched them at all. Among the 3 percent that were fetched, AI bots accounted for 19.5 percent of requests, with the rest coming from audit tools, general crawlers, and tools studying the file itself. That is measured behaviour rather than opinion, which is why it is the most useful evidence available.

    Then why are so many agencies selling llms.txt as an AI visibility service?

    Because it takes minutes to deliver and is hard for a client to check. AI citations are slow to appear, difficult to measure, and vary by question, so nobody can quickly prove the file did nothing. It also sounds technical enough to feel like real work. Ask any vendor selling it what they will measure and when, and whether the price includes content changes.

    What actually decides whether ChatGPT or Perplexity mentions my name for a local question?

    Having a page that answers the specific question, in text, that their crawler can reach. General questions get answered from what the model already knows, so your page is never needed. Local and current questions need a source, and that is where an agent can win. Our post on leads from AI search covers the mechanism in more detail.

    I run a small agent site with 12 pages. Is there any version of this file worth making?

    The honest answer for a 12-page site is that the file would list pages any crawler already finds through your navigation and sitemap. The effort is small enough that you can do it if you want the tidiness. The same twenty minutes spent making sure each of those 12 pages answers one specific local question in plain text will do more for citations.

    Does having llms.txt risk anything, like giving away content to AI companies?

    It does not expose anything a crawler could not already reach, since it just lists pages that are already public. If your concern is AI tools using your content at all, the control for that is your robots.txt file and the crawler permissions in it, which is a separate decision. Blocking AI crawlers also removes you from AI answers, which is usually the opposite of what an agent wants.

    A competitor added llms.txt and now appears in AI answers. Did the file do that?

    Almost certainly not by itself, and the timing is worth examining. Most sites that add the file also update content, improve pages, or start publishing more, and any of those would be the likelier cause. The Ahrefs data showing 97 percent of these files never get fetched makes a direct causal link hard to support. Look at what else changed on their site in the same period.

    If I should not spend money on llms.txt, what should that budget go toward for AI visibility?

    Pages that answer real local questions nobody else has answered. What the rental rules in your city say this year, what a monthly fee typically covers in your buildings, what recent sales on a specific street closed at. Those are facts an AI assistant cannot get from its training data and has to source. Add a complete Google Business Profile and reviews, which feed the local picture.

    How would I even tell if an AI tool is citing my site?

    Ask the questions your clients ask, in the tools they use, and write down what comes back. Run the same set of five or ten questions once a month and note whether your name or your pages appear. It is manual and it is the most reliable check available, because these answers vary by user and by phrasing. Our post on free AI visibility tools covers what the automated options can and cannot tell you.

    Free guide for agents

    Find out why Google and ChatGPT skip your website

    The Real Estate Ranking Guide walks through what both of them read, what they ignore, and a 90-day plan you can start this month. Free, and it opens as soon as you submit.

    Call 604.401.4849