BlogState of the agentic web
State of the agentic web
We fetched the agent-facing files and endpoints of 150,074 hosts drawn from the Tranco list. This is what the sites that published them got wrong.
Rules for AI are eleven times more common than interfaces for agents
Published something a machine can read
88,484sites out of the 150,074 we asked published at least one machine-readable file. Most of that is robots.txt and sitemaps, which are decades older than any of this.
Wrote rules for AI
34,160sites name specific AI crawlers in robots.txt, or publish a signal saying how their content may be used. What those rules say is a separate question: this counts sites that decided to have a policy at all.
Published an interface
3,007sites published something an agent can call: 368 an MCP card, 73 an agent card, 8 NLWeb. For every one of them, eleven have written a rule about AI instead.
Every industry writes rules faster than it builds interfaces
Policy is commonest on online shop sites, at 62%, and rarest on nonprofit or association ones, at 22%. Shops are the exception on the other side: 24% of them publish something an agent can act through, against 3% or less everywhere else.
- Online shop62%
- News or publisher61%
- Marketplace or classifieds60%
- Community or social49%
- Reference or directory46%
- Gambling42%
- Entertainment or media32%
- Finance29%
- Software or SaaS29%
- Education27%
- Company or brand24%
- Government or public body24%
- Nonprofit or association22%
Everything we measured, by industry
Almost every site here publishes robots.txt, and around seven in ten publish a sitemap. Nearly four in ten write rules naming individual AI crawlers. After that the grid is nearly empty: everything built specifically for agents is still rare in every industry.
Which AI crawlers sites name in robots.txt
- gptbot31%reads to train
- claudebot27%reads to train
- ccbot24%reads to train
- google-extended22%reads to train
- chatgpt-user20%fetches for a person
- bytespider20%reads to train
- amazonbot19%reads to train
- perplexitybot18%reads to train
- anthropic-ai16%reads to train
- applebot-extended16%reads to train
- meta-externalagent15%reads to train
- oai-searchbot13%fetches for a person
- perplexity-user5.1%fetches for a person
- duckassistbot3.9%fetches for a person
- novaact0.1%fetches for a person
- operator0.1%fetches for a person
35% of sitemaps break the rules
Sites with a sitemap
61,587sites publish one. The format was agreed in 2005, so it has had 21 years to settle down.
Sitemaps with a fault
21,448of them break a rule of the sitemap format somewhere: 35% of everyone who publishes one.
Faults Google overlooks
11,780of those break only one rule: the sitemap sits in a subfolder and lists pages outside it. Google reads those pages anyway if you submitted the file in Search Console.
Faults no submission fixes
9,668are wrong however you submit them: one in every six sites that publish a sitemap.
Bigger sitemaps break more often
62%of sitemaps with over 50,000 URLs have a fault, against 29% of those with under 100 URLs. We could not count the URLs in 1,510 of them, so those are left out of this split.
Which mistakes actually happen
The single commonest mistake is in the sitemap: the file sits in a subfolder and lists pages outside that folder. 14,176 sites do it, 50% of every fault we counted. The sitemap format says those pages do not count, though the big search engines read them anyway if you submitted the file in Search Console.
XML sitemap61,587 sites publish one
- lists pages above its own folder14,17623%
- lists pages on another domain6,1069.9%
- missing the XML namespace2,0773.4%
- is not a sitemap at all1,5622.5%
- uses relative URLs4920.8%
- is over the size limit3050.5%
- is an empty file2150.3%
- will not parse2020.3%
llms.txt13,730 sites publish one
- starts with text before its title1,0227.4%
- has no title at all2171.6%
- hides a heading inside a collapsed block720.5%
robots.txt85,379 sites publish one
- puts a rule before any User-agent line8131.0%
- has a line missing its colon2540.3%
OpenAPI description497 sites publish one
- describes a value without saying what type it is17635%
- leaves an operation with no name to call it by11223%
- gives two operations the same name6613%
Per-crawler rules in robots.txt33,626 sites publish one
- allows and disallows the same path1340.4%
- misspells a crawler's name800.2%
UCP profile2,202 sites publish one
- leaves out a registry the profile requires1074.9%
Agent card73 sites publish one
- leaves out fields the card is required to carry5474%
Every technology, and how often it is wrong
Of the sites that published each one, how many got it wrong. Open a row for its faults.
- Untyped schema35% · 176 of 497
- Missing operation ID23% · 112 of 497
- Duplicate operation ID13% · 66 of 497
- Invalid version5.4% · 27 of 497
- Invalid document2.4% · 12 of 497
- Out of directory URLs23% · 14,176 of 61,587
- Cross host URLs9.9% · 6,106 of 61,587
- Missing namespace3.4% · 2,077 of 61,587
- Wrong document type2.5% · 1,562 of 61,587
- Relative URLs0.8% · 492 of 61,587
- Too large0.5% · 305 of 61,587
- Empty document0.3% · 215 of 61,587
- Malformed XML0.3% · 202 of 61,587
- Content before h17.4% · 1,022 of 13,730
- Missing h11.6% · 217 of 13,730
- Heading in details0.5% · 72 of 13,730
- Missing required registries4.9% · 107 of 2,202
- Missing transports1.3% · 28 of 2,202
- Invalid document0.3% · 7 of 2,202
- Invalid version0.1% · 2 of 2,202
- Content Signals4.3% · 64 of 1,495
- Directive before group1.0% · 813 of 85,379
- Missing colon0.3% · 254 of 85,379
- Empty user agent0.0% · 15 of 85,379
- Oversized0.0% · 11 of 85,379
- Contradictory allow disallow0.4% · 134 of 33,626
- Misspelled token0.2% · 80 of 33,626
A bar counts every site with any fault in that file, so it is not a count of the worst fault. 35% of sitemaps have something wrong does not mean 35% are broken XML files. That is 0.3% of them. The usual problem is a sitemap in a subfolder that lists pages outside it, which the format says do not count even though the big search engines read them anyway. Open a row to see which faults actually fire.
Who refuses an identified agent
14,227 of 150,074 sites turned us away with a 401, 403 or 429. We know because we said who we were. A crawler dressed as a browser cannot measure this at all.
How this was measured
We checked 150,074 hosts, taken from the top 150,451 names in the Tranco research ranking, as it stood on 14 September 2026, between 17 September 2026 and 19 September 2026. Tranco combines several popularity sources over 30 days, some counting how often a name is looked up and some how often people visit, so a good part of the list is infrastructure rather than websites. The last 377 names in that range have no row: the run stops once it has enough hosts and its workers take the list in turns, so they do not stop on the same rank. Nothing was left out for anything to do with the sites themselves.
One question, asked of every site that tried. When a site publishes a file or an endpoint meant for AI agents, how often is it implemented wrongly, and in what way. Every figure on this page is counted over the sites that published the thing being measured, never over the corpus.
What this study does not claim. It does not say what share of the web has adopted anything. The list behind it ranks domain names by popularity, not by whether they are websites: 27,992 of these names never replied and another 14,227 turned us away, and much of the rest is not a website at all but nameservers, content delivery, API hosts and parked names. An adoption rate needs a denominator of sites, and this frame gives a denominator of hosts that answered a request.
We used the same checks Lumar runs on a customer's own site, not a special version written for this study. Every fault counted here is one the product would report. Files are fetched; no page is rendered and no JavaScript is run.
Breaking a standard and being awkward for an agent are not the same thing, and we count both. Of the 116 faults we can report, 48 break a rule in a published specification. The other 68 are things a specification allows but an agent cannot work with. A missing operation ID in an OpenAPI description is the clear case: the specification does not require one, so the document is valid, but an agent choosing a call has nothing stable to name it by. Shapes a specification explicitly permits are recorded and never counted as faults, which is why a robots.txt naming one crawler in two groups is absent from these charts: RFC 9309 tells crawlers to merge them, so every rule still applies.
We told every site who we were. Lumar's crawler normally identifies itself as Googlebot-compatible, and sites check that claim by looking up who owns the visiting address. Ours is not Google, so the check fails and the visit looks like something pretending to be Google. On a trial of 36 sites, the Googlebot-style name was turned away by 8 of them and this one by 2. Being honest got us more data, not less.
Every dot is a hundred hosts
150,074names from the Tranco ranking, tried in rank order. A popularity ranking of domain names, which is not the same as a list of websites.
Not all of them answer
107,855of the 150,074 answered our knock. 27,992 never replied and 14,227 turned us away, and nothing is claimed about any of them.
Answering is not the same as letting us read
100,008let us read something. The rest answered a home page and then refused every file we asked for, which proves nothing about what they publish.
This is the population every rate is counted over
88,484published at least one file or endpoint meant for agents. Not a share of the web, but a count of the sites that tried.
And this is what the study is about
24,119of them got something wrong. 27% of everyone who tried.
107,855 of 150,074 sites answered our knock.
What became of the rest. Answering the home page is not the same as letting us read the files, and this chart counts only the first.
- Answered normally69% · 103,869
- Answered, with an error2.7% · 3,986
- Turned us away9.5% · 14,227
- Never answered19% · 27,992
A technology is only given a percentage once at least 400 sites publish it, and is shown as a plain count below that. It is a reporting threshold, not a significance test: this is not a random sample of the web, so the point is simply not to let a reader over-read a handful of sites. Turning us away is us failing to look rather than a site publishing nothing, so no absence is ever counted against a host that refused us. What a host did serve is counted like anyone else's, even when it refused the first knock at its home page.
How the industries were assigned. Each host was put in exactly one category by a language model reading its home page, falling back to the domain name where no page could be read. It placed 84,149 of the 88,484 publishing hosts; the 4,335 it could not place are left off the industry charts and counted everywhere else. A category appears on the industry charts only once 900 of its hosts publish something we can check. The categories are ours, not a standard taxonomy, and no accuracy audit has been run on them, so treat an industry comparison as indicative rather than measured. Read the ends of that chart, not the middle: a gap of a point or two between neighbouring rows is well inside what a mislabelled host or two could move.
A rule about AI is not the same as a rule against it. What this study can see is that a site named a crawler in robots.txt, or published a content signal, an ai.txt or a mining reservation. It cannot see which way those rules point: a group headed by a crawler's name is just as valid with an Allow in it as with a Disallow, and Content-Signals, ai.txt and TDMRep all state preferences that may permit as easily as refuse. So the figures here count sites that have decided to have a policy, not sites that have shut the door. Restricting a crawler and publishing an interface are not opposites either: doing both is a coherent position, and arguably the considered one.
Using these figures. The findings on this page are published under CC BY 4.0: quote them, chart them and build on them, including commercially, as long as you credit Lumar and link back. We publish the findings rather than the underlying rows, so there is no data file to download.