Get started
Features Overview Testimonial Faq Contact

97% of llms.txt Files Are Never Read: What the Data Shows (and What to Do Instead)

June 19, 2026

There’s a recurring pattern in SEO history: a new optimization tactic gets announced, practitioners rush to implement it, and then the data quietly reveals it doesn’t work. In 2026, llms.txt is following that pattern almost exactly.

A new large-scale study from Ahrefs analyzed 137,000 websites that had published llms.txt files — the proposed standard for telling AI crawlers what content to use and avoid. Their finding: 97% of those files were never fetched by any known AI crawler. Of the 3% that were read, there’s no evidence that reading the file changed how any AI system cited or used the content.

Meanwhile, Google added an llms.txt check to Chrome Lighthouse — the developer audit tool used by millions — while simultaneously confirming that llms.txt files have no effect on Search or AI Overviews rankings. The message is mixed, the enthusiasm is real, and the actual impact is nearly zero.

Here’s what the data actually shows about AI visibility — and where the real leverage is.

What llms.txt Was Supposed to Do

The llms.txt proposal, put forward by Jeremy Howard (fast.ai) in late 2024, was designed as a “robots.txt for AI.” The idea: place a structured markdown file at yourdomain.com/llms.txt that lists your key pages, explains your content, and tells AI systems how to use it. A companion format, llms-full.txt, would contain the actual content in AI-friendly form.

The theory was sensible. As AI systems increasingly browse the web to generate answers, having a clean, machine-readable summary of your site’s content could make you easier to cite. It’s the same logic that made sitemaps useful — give the machine a map, and it will find you faster.

The problem is that no major AI provider committed to using it. Not OpenAI, not Google, not Anthropic, not Perplexity. The file format exists; the readership doesn’t.

The Ahrefs Study: 137,000 Sites, 97% Ignored

Ahrefs crawled 137,000 domains that had published llms.txt files and cross-referenced that against known AI crawler logs. The result: only 3% of those files showed any evidence of being fetched by AI systems.

That 3% figure likely overstates the impact. Crawling a file and using it are different things. A crawler might fetch llms.txt as part of a general crawl sweep — the same way Googlebot fetches robots.txt on every domain — without any downstream effect on how content is cited or ranked.

No AI provider has published documentation explaining how llms.txt influences their citation decisions. The specification remains a proposal, not a standard. Implementing it today is the equivalent of writing a letter in a language no one reads, then mailing it to an address where no one collects the post.

Google’s Lighthouse Signal: What It Actually Means

In June 2026, Google added llms.txt detection to Chrome Lighthouse. This created a wave of coverage suggesting Google was endorsing the format. The reality is more specific.

Google framed the Lighthouse check under “agentic browsing” — the scenario where AI assistants navigate the web on behalf of users. The suggestion is that if AI agents browse your site, having structured content helps them understand it. That’s plausible. But Google was explicit on the ranking question: llms.txt does not affect Search rankings or AI Overviews inclusion.

Lighthouse is a diagnostic tool. Adding a check for llms.txt is similar to Lighthouse checking whether your page has a meta description — it surfaces the option, it doesn’t validate that using it will improve your results. Google has added Lighthouse checks for dozens of signals over the years that had marginal or zero impact on actual rankings. This appears to be another one.

What Actually Predicts AI Citation

While the llms.txt data is negative, other data tells a more useful story about what AI systems do respond to.

A 2026 SSRN study by Kurt Fischman analyzed thousands of queries across major AI systems and found that pages with comprehensive structured data were 3.2x more likely to appear in AI Overviews and AI-generated answers. The effect was strongest for schema types that include specific, measurable attributes — pricing, ratings, availability, specifications. Attribute-rich schema gave AI systems something concrete to cite.

OtterlyAI ran a controlled experiment adding schema markup to pages that previously had none, then measured AI citation rates before and after. Pages with schema were cited more frequently, with the effect concentrated in queries where the AI needed to surface a specific fact — a price, a date, a specification — rather than summarize general information.

Digital Applied found a 2.3x citation rate difference between pages with schema and those without, across a sample of e-commerce and service pages. The pattern held across ChatGPT, Perplexity, and Google’s AI Overview.

The common thread: AI systems cite sources that make facts easy to extract. Structured data is a machine-readable claim. llms.txt is a file that most AI systems don’t read. The delta in impact between the two is not subtle.

Schema Markup as AI Visibility Infrastructure

The practical implication is that schema markup — which has existed since 2011 and is fully supported by every major search engine — is the primary lever for AI citation visibility.

The schema types with the strongest AI citation signal are those tied to specific, verifiable claims:

  • Product with offers, price, aggregateRating — AI systems cite prices and ratings directly in answers
  • FAQPage — directly matches the question-answer format AI systems prefer
  • HowTo — structured steps are frequently extracted verbatim
  • Article with datePublished, author, publisher — signals freshness and authority
  • LocalBusiness with hours, location, and contact data — extracted constantly for local queries

The Fischman study found that schema completeness mattered as much as schema presence. A Product page with only name and description showed weak citation lift. The same page with price, availability, aggregateRating, and brand showed strong citation lift. AI systems reward completeness because complete structured data gives them more to work with.

Entity Signals and Earned Mentions

Structured data alone doesn’t fully explain AI citation patterns. The Fischman study also identified entity signals — whether a brand, person, or organization is referenced across multiple credible sources — as a secondary predictor.

Pages that were cited by AI systems tended to belong to entities that also appeared in Wikipedia, Wikidata, or were referenced in high-authority publications. This tracks with how AI systems build their training data and knowledge graphs: entities with cross-source validation are treated as more reliable.

For most websites, building entity presence means two things: creating a complete, consistent brand identity across structured data (use the same legal name, address, and URL consistently across all schema), and earning genuine third-party mentions in publications that AI training sets include.

Neither of these is a quick fix. Both are legitimate — which is partly why they work. AI systems are better at detecting citation-gaming than search engines were in 2005, because they process the semantic content of pages, not just keyword signals.

Should You Publish an llms.txt File?

The cost of publishing llms.txt is low: 30 minutes, a markdown file, a deployment. If you have a technical team and want to future-proof against the scenario where AI crawlers do start reading these files, implementing it is reasonable.

But if your goal is to appear more frequently in AI-generated answers in 2026, llms.txt is not where your time should go. The data is clear on that.

The higher-ROI path:

  1. Audit your schema coverage — most sites have incomplete or missing structured data on the pages most likely to be cited
  2. Add attribute-rich schema to your highest-value content — product pages, FAQ pages, how-to guides
  3. Build entity consistency across your web presence
  4. Target earned mentions in publications with AI citation authority

Schema markup isn’t exciting. It doesn’t generate the kind of launch-week enthusiasm that a new file format does. But the Ahrefs, Fischman, OtterlyAI, and Digital Applied studies are all pointing in the same direction: structured data is the signal AI systems currently respond to.

Conclusion

llms.txt is a well-intentioned proposal that arrived before anyone committed to reading it. The Ahrefs study is the clearest signal we have: 97% of those files sit unread while the sites that invested in comprehensive schema markup are getting cited 2-3x more often by AI systems.

The framework for AI visibility in 2026 isn’t meaningfully different from what it was for structured data in 2015 — give AI systems machine-readable, attribute-rich claims about your content, build consistent entity signals, and earn third-party validation. The delivery mechanism is newer; the underlying logic isn’t.

If you want to see how your site currently performs on AI visibility signals — structured data coverage, entity consistency, citation potential — run a free audit at ai-visibility.llmagnet.com. It covers the signals that the current data shows actually matter.

Liked it? Share on social media

More articles:

Your Content Has a 13-Week Window for AI Citations. Here’s the Data and What To Do About It.
One GEO Strategy Won’t Work Across All AI Platforms. Here’s the Data That Proves It.
Schema Markup Gets You 2.5x More AI Visibility. Here’s Exactly What to Implement
YouTube Is Now a GEO Channel. Here’s What Actually Gets Cited.