Readiness study · 2026-08-31
Your website may already be answering AI agents. You didn’t write what it says.
We read 6,284 small business websites across 10 industries. 26% of them publish a file that exists for one purpose — telling an AI agent what the business is. Almost none of their owners wrote it, chose it, or know it is there. Presence is arriving as a platform default. Control is not arriving at all.
What was measured
Public HTTP GETs of the documents an AI agent looks for — /llms.txt, three/.well-known/ documents, robots.txt, a sitemap, and the organisation data on the homepage — following each site’s own robots.txt. Nothing was submitted. No page was altered. No business was contacted.
Businesses were found in OpenStreetMap’s open data, filtered to those publishing a website of their own, across 10 areas. Two rules govern every number here. A check that could not complete is recorded as unknown, never as absent — a server that answered 503 is a server that failed to answer, not a business that publishes nothing. And 1,887 businesses whose core documents could not be read are excluded from every percentage below rather than counted as failing.
1. Presence arrived by default
A year ago /llms.txt was a proposal. Today 1,242 of the 4,856 businesses where we could decide the question publish one.
| Industry | Checked | Publish one | Share |
|---|---|---|---|
| Events & bridal | 7 | 5 | 71% |
| Antique dealers | 55 | 24 | 44% |
| Galleries | 54 | 23 | 43% |
| Self-care | 814 | 240 | 29% |
| Building trades | 292 | 81 | 28% |
| Trade suppliers | 258 | 72 | 28% |
| Professional services | 712 | 195 | 27% |
| Clinical | 786 | 180 | 23% |
| Hospitality | 1,495 | 336 | 22% |
| Wellness | 383 | 86 | 22% |
That looks like an adoption curve. It is a release note.
Of the 1,238 files we could re-read, 734 name their own author — and the authors are five pieces of software.
| Wrote the file | Files |
|---|---|
| Wix | 338 |
| Shopify | 157 |
| Yoast SEO | 104 |
| All in One SEO | 87 |
| Rank Math | 37 |
| other self-declared generator | 4 |
| Spillover | 4 |
| WordPress (unspecified plugin) | 3 |
The other 504 name no author. We count that as no author named — not as “written by the business”. Plenty of software does not sign its work.
And it goes past a text file. 498 of these sites declare a working agent interface, in as many words: “This site is powered by Wix and supports the Model Context Protocol (MCP) for agentic AI access.” A garden centre in Alabama and an ice cream maker in Tennessee are both queryable by an assistant today, over a live protocol, because their storefront platform shipped it. Nobody at either business decided that.
2. Presence is not control
Here is what none of the 6,284 businesses publish.
| Document | Publish it | Of |
|---|---|---|
| A machine-readable record of what this business is | 0 | 4,407 |
| A catalogue of what it does, with evidence | 0 | 4,405 |
| Its own terms for what an agent may ask or commit it to | 0 | 4,403 |
A platform-generated file says what the platform decided to say. It is assembled from the pages and products already on the site. It does not say what the business will quote a price on and what it will not. It does not say which services it actually performs versus which it lists for search traffic. It does not say the hours it will honour, the area it will travel to, or which questions it wants a human on.
An agent reading a platform-generated file learns the site’s page structure. An agent reading nothing at all learns the same thing from the HTML. In both cases the business is the subject of the description rather than its author. That is the gap. Presence is now free and automatic; accuracy, capability and answering on your own terms are not.
3. Three quarters are outside even the automatic layer
3,614 of 4,856 businesses — 74% — publish no /llms.txt at all. These are disproportionately the businesses not on a platform that ships one: independent sites, older builds, custom work, and small firms whose site was built once and left alone. No plugin update is coming for them.
A further 213 block AI crawlers outright.
| Agent disallowed in robots.txt | Businesses |
|---|---|
| CCBot | 196 |
| GPTBot | 193 |
| ClaudeBot | 187 |
| Google-Extended | 186 |
| Applebot-Extended | 180 |
| Claude-Web | 11 |
| PerplexityBot | 11 |
We make no claim about whether blocking is right or wrong; some businesses have good reasons. What we observe is that it is almost always all of them at once, which is what a pasted line looks like rather than five separate decisions.
4. What we got wrong while measuring this
Two findings were caught before publication, and both would have been ours to answer for.
We built a scoring band only reachable by buying our own product. Our assessor bands a site “agent-ready” at 70 out of 100. The highest score any business achieved was 57. Three of the seven checks are the agent-presence documents this firm publishes and sells; they were present zero times, so the top band was arithmetically unreachable. “No small business is agent-ready” would have been true against our own definition and false to every reader. No business is given a score anywhere on this page or in the dataset.
We reported 498 businesses as having no agent interface, because we only looked for our own filename. The check reads /.well-known/agent-interaction.json and nothing else, while Wix was writing MCP support into files we had already downloaded. A measure that can only see its author’s spec reports the rest of the world as empty. The figure above is a floor, not a total: MCP was only looked for inside llms.txt files.
We publish both errors because a study that reports only its successes is asking to be trusted rather than checked.
Look up any business in the study
Every one of the 6,284 businesses is listed with the facts observed about it: which documents were present, which absent, and which could not be read. No business is scored, graded or ranked. We report what a URL returned. What that is worth is your judgment, and theirs.
Download the complete dataset — 6,284 businesses, JSON.
What this study does not claim
- Nothing about outcomes. We measured what is published, not whether it changes how often a business is surfaced, cited or recommended by any assistant. That is a different study and we have not done it.
- Nothing about the 504 unattributed files. They may have been authored, or generated by software that does not sign its work. We do not know which.
- Nothing about the 1,887 businesses whose documents could not be read, beyond the individual checks that did complete.
- Nothing about businesses without a website in OpenStreetMap — a different and almost certainly larger population than this one.
- Nothing about whether blocking AI crawlers is a mistake. We counted it.
Check it yourself
Every figure on this page is generated from the dataset rather than typed, and a verification gate refuses the build if any of them stops reproducing. Any business listed can rerun the same checks on its own domain in a browser.
This is a point-in-time reading. One business in our sample was recorded as having no /llms.txt and returned a 44KB one when re-checked an hour later; we do not know whether the site changed or our earlier read differed, and we have not named it as either. Any business that wants to disagree with its line can change the answer today.
If we have something wrong, we would like to be told, and we will correct it in place with a note saying what changed. Tell us, or check your own site free.
Study by River Cade Concepts. The scheme behind the check is published in full.