Research: how AI engines find and recommend businesses
Predictions written down first. Misses published next to the wins.
Lefty Media Co measures how AI answer engines like ChatGPT, Claude, and Perplexity find, read, and recommend small businesses. We measure it on websites we run ourselves, from server logs and stored answers, and we publish what we find. Every number on this page has a date, a denominator, and a way to be wrong. Papers are deposited on Zenodo with a DOI under Nathan Hall's ORCID 0009-0005-3040-9995.
Who is crawling us right now
Our clients watch this panel in their Locker. This is the same panel for every site we run or monitor, in public. Verified crawlers only: the crawler name and its network address both have to check out. A Google search advocate has said no AI system uses llms.txt, and a large industry study found 97% of the files it checked were never read; both looked at the file at the site root. The pair in the middle is the paper’s measurement, still running.
The snapshot is rewritten every 15 minutes so a crawler reading this page gets real numbers. In a browser the tiles refresh live.
The paper measured this pair over 165 site-days: 1,549 announced, 2 root. This is the same count, still running.
Pages = content pages. AI files = the files we deploy for engines (/ai/*, llms.txt). Plumbing = robots.txt, sitemaps, feeds, what every crawler reads first. Refused = came, and the server answered 429.
/ai/llms.txt/llms.txt/ai/ai-schema.jsonLog scale, so small crawlers stay visible beside large ones. Totals are the last 30 days.
Verified crawlers read at least one content page on 30 of the last 30 days. Hover the chart for the pages read on any day.
These arrived and named themselves. Their operators publish no address list, so there is nothing to check the claim against. Not counted above, and not counted is not the same as did not happen.
760 · 4d ago
96 · 2h ago
55 · 15d ago
23 · 1d ago
6 · 18d ago
This page: 0 verified AI crawler reads in the last 30 days.
Verified = the crawler name and the source address both match the operator's published ranges, checked against a list refreshed weekly (ClaudeBot: the network block registered to Anthropic, network-attributed, as the paper marks it). Pages = content pages; AI files = /ai/*, llms.txt, .well-known; plumbing = robots.txt, sitemaps, feeds. Nine sites we run or monitor. Googlebot is a search crawler and is logged on only two of them, so it is left out. No client domain leaves this route. Paper: doi.org/10.5281/zenodo.22814843. If the count is unavailable, this panel says so. It never shows a number we made up.
Published
Discovery, not adoption: crawler retrieval of an llms.txt resource announced through robots.txt, across seven sites on one server, and what a server log can and cannot show.
Nathan Hall. Preprint, September 17, 2026. doi.org/10.5281/zenodo.22814844
Finding: across six sites in 165 site-days, verified crawlers fetched an llms.txt file announced in robots.txt 1,549 times, 9.4 per site-day. The same file left un-announced at the root was fetched twice. Discovery is real. Use of the content in an answer is not shown, and the paper says so. Data, scripts, and a manifest are deposited with the paper; client sites are lettered Site A to F.
Version 1.1, September 21, 2026. doi.org/10.5281/zenodo.22880411. Adds a section answering the directory-link confound raised by Google’s John Mueller. No counts changed. All versions: doi.org/10.5281/zenodo.22814843.
Detecting undisclosed model changes from local session logs: a method, and a worked example in which the operator was one day off.
Nathan Hall. Preprint, August 26, 2026. doi.org/10.5281/zenodo.22117618
Finding: a working method for detecting that the model behind an AI tool changed when nobody announced it, from the tool’s own local session logs. The worked example includes the operator’s own miss: his estimate of the change date was one day off.
Both are on Zenodo under ORCID 0009-0005-3040-9995. Cite the DOI, not this page. This page changes.
Running now
Four registered studies with the prediction locked and the read date set. The result goes here on the read date, whichever way it lands.
1. Filename control, leftymediaco.com. Registered September 16, 2026, before applying.
Question: does the crawler follow the robots.txt Sitemap line, or does it want the name llms.txt? Treatment: a byte-identical copy of the file under a different name, announced the same way, nothing else changed.
Locked prediction: “within 7 days of applying, ClaudeBot (network-attributed) fetches [the copy] at the same rate as llms.txt on this site, within the Poisson 95% interval of llms.txt’s rate over the same days. Bingbot: same, within 14 days.”
Reads: September 24, 2026 (ClaudeBot) and October 1, 2026 (Bingbot). Result posted here either way.
2. Announcement control, Site F. Registered September 16, 2026, before applying.
Question: does one Sitemap line in robots.txt start fetches on a site whose root llms.txt had never been fetched by a verified crawler? Treatment: that one line. No link tag, no other change.
Locked prediction: “ClaudeBot (network-attributed) llms.txt fetches on this site go from 0 to about 6/day within 7 days of its next robots.txt read; Bingbot from 0 to about 2/day within 14.” And the kill condition, in the same document: “If fetches stay at 0 while robots.txt is being read, the robots.txt line is not the mechanism and paper 3.3 is wrong.”
Reads: September 24 and October 1, 2026.
3. Content, not fetches. Registered September 17, 2026, amended once the same day, both before placement.
Question: does anything in llms.txt ever reach an answer? Three true facts about Lefty Media Co, unknown to the web, committed by SHA-256. One goes only in llms.txt. One goes only on one page of this site. One goes nowhere. Then we ask the engines every day for 30 days.
Locked prediction: the llms.txt-only fact scores 0 hits from ungrounded ChatGPT and Claude for 30 days, with a stated 15% chance that any row hits without the file in the retrieved context. “If it happens, it is the first evidence in this fleet that llms.txt CONTENT reached an answer, and it is reported as such.”
Placement: September 23, 2026 or later. Window: 30 days.
4. Known but never recommended. Registered September 2, 2026. Starts October 1, 2026.
Question: can a business be well known to the engines and still never be recommended? Data: our daily sweep of nine sites on three engines, two question types: recognition (does AI know who you are) and recommendation (does AI pick you when someone asks who is best).
Locked prediction: “Across the fleet, sites scoring above 75% on geo questions will show no correlated floor on ai_discoverability: at least two will sit below 5% while above 75% geo.”
Kill condition: “If every site above 75% geo also sits above 25% discovery, the two are not independent and this is wrong.”
Closes when eight sites have 30 or more days each under one unchanged question set. Interim, on mixed question sets and not the result: four of nine sites sit above 75% recognition and below 2% recommendation, one of them at 0 of 2,583.
Falsified
A prediction that cannot fail is not a measurement. These failed.
Small business sites are faked more than the industry benchmark.
Locked prediction: “Every Locker exceeds the reported 5.7% industry spoof rate, and the smallest sites exceed it by the widest margin.”
Result: two of six sites came in below 5.7%, and the smallest site was the lowest. The sensor on those sites was a PHP hook that could not see the traffic that does the spoofing. The instrument failed, not the world. A follow-up study on the sensor is registered and running.
A name that resolves to a different profession makes the engines answer about that profession.
Locked prediction: from the names alone, before reading any holdout answer, at least four of five holdout sites would match the per-site call.
Result: two of five. Falsified at the 0.5 null. The mechanism is real where it fires (see the Casebook) and does not generalize from a name alone.
Casebook
Single observations with the numbers attached. A case is a thing that happened once, on record. It becomes a study when it gets a prediction.
ClaudeBot came back the day after the site changed, and read the changed pages first.
leftymediaco.com, verified Anthropic addresses only. September 1 to 9, 2026: a visit every day, three files each time (robots.txt, the sitemap, llms.txt), zero content pages. September 10 to 19: no visits. September 19: eighteen logged changes to this site. September 20: 19 hits, 13 paths, 11 of them content pages, the changed pages first. One site, one event, no causal claim. Open question: how it knew the new pages existed before its first sitemap read of the day. Next read: September 26, 2026.
The wrong answer is not random. It is retrieved from the name’s neighborhood.
A Kansas City mobile dog gym. Asked with no web access what the business does, GPT-4o answered GPS fleet tracking, fuel monitoring, and keyless entry. Asked with web access, Perplexity returned auto transport companies and four federal motor-carrier records that share the name. The hallucination was not invented. It was retrieved from the wrong neighborhood. We then registered the general version on five holdout sites and it landed on two, so the case stands and the law does not. The follow-up separates name collision from improvisation.
How we count
Where the numbers come from. A JSON sensor in the web server on every site we run, all on one server we control. Every request is kept, not sampled.
Who counts as a crawler. A request counts as a verified AI crawler only when its name and its network address both match the operator’s published ranges. The user-agent string alone never counts. Requests that fail the check are reported as a number and never shown as reads.
Our own traffic. Every test request we send carries a tag and is excluded. In one read our own office address wearing a crawler’s name was caught by the same check, which is the check working.
Answers. Each engine is asked the same questions on a schedule. Every answer is stored with the engine, the model, the date, and the sources it opened, so a number can be recomputed later from the stored rows instead of a fresh run.
Predictions first. The prediction and what would prove it wrong are written and dated before the data comes in. Where a fact has to stay hidden, it is committed by SHA-256 and revealed with the result.
Data. Deposited with each paper on Zenodo: the crawler rows, the scripts, the outputs, and a manifest. Client sites are lettered. Server logs that contain human visitors are never published.
One instrument. The number we report to clients, Pick Rate, is computed by the same code as the number we publish here.
Questions people ask about this research
Does anyone read llms.txt?
Yes, when the site tells crawlers where it is. Across six sites we run, verified crawlers fetched an llms.txt file announced in robots.txt 1,549 times in 165 site-days. The same file at the site root, not announced, was fetched twice. That is discovery. Whether the content is used in an answer is a separate question, and we are testing it now.
What is a verified AI crawler?
A request whose crawler name and network address both match the operator's published ranges. Anyone can type "GPTBot" into a user-agent string. The address is what we check, and requests that fail are counted separately and never shown as reads.
What is a preregistered study?
The prediction, and what would prove it wrong, are written down and dated before the data comes in. Ours are in the deposited papers and in the Running section of this page, with the read dates.
Why publish the misses?
A page that only shows wins is an ad. A prediction that failed is what makes the ones that held worth believing.
Where is the data?
On Zenodo, with each paper: the crawler rows, the scripts, the outputs, and a manifest. Client sites are lettered Site A to F. Server logs with human visitors are never published.
Can I cite this research?
Yes. Cite the DOI of the paper, not this page. Papers are on Zenodo under ORCID 0009-0005-3040-9995 and do not change after deposit. New versions get their own version DOI.
Who does this research?
Nathan Hall, founder of Lefty Media Co in Kansas City, Missouri. Lefty Media Co measures AI discoverability for small businesses and publishes what it learns.
See what AI says about you.
Drop in your website and watch what ChatGPT, Claude, and Perplexity say about your business right now. Thirty seconds, no email, no pitch.
Free. Thirty seconds. No email.