agentic‑readiness docs
/
Build one › Run a scan ›

BlogState of the agentic web

State of the agentic web

We fetched the agent-facing files and endpoints of 150,074 hosts drawn from the Tranco list. This is what the sites that published them got wrong.

Rules for AI are eleven times more common than interfaces for agents

34,160 sites set a policy for AIper-crawler rules in robots.txt 33,626 · Content-Signals 1,495 · TDMRep 213 · ai.txt 1263,007 sites published an interfaceUCP 2,202 · OpenAPI 497 · MCP server card 368 · MCP authentication 191 · agent card 73 · NLWeb 8

Published something a machine can read

88,484

sites out of the 150,074 we asked published at least one machine-readable file. Most of that is robots.txt and sitemaps, which are decades older than any of this.

Wrote rules for AI

34,160

sites name specific AI crawlers in robots.txt, or publish a signal saying how their content may be used. What those rules say is a separate question: this counts sites that decided to have a policy at all.

Published an interface

3,007

sites published something an agent can call: 368 an MCP card, 73 an agent card, 8 NLWeb. For every one of them, eleven have written a rule about AI instead.

Every industry writes rules faster than it builds interfaces

Policy is commonest on online shop sites, at 62%, and rarest on nonprofit or association ones, at 22%. Shops are the exception on the other side: 24% of them publish something an agent can act through, against 3% or less everywhere else.

has a policy for AIhas an interface agents can call
  1. Online shop62%
  2. News or publisher61%
  3. Marketplace or classifieds60%
  4. Community or social49%
  5. Reference or directory46%
  6. Gambling42%
  7. Entertainment or media32%
  8. Finance29%
  9. Software or SaaS29%
  10. Education27%
  11. Company or brand24%
  12. Government or public body24%
  13. Nonprofit or association22%
Tap or hover a row for what each side counts. An industry is shown once 900 of its sites publish something we can check; adult sites are counted in every figure on this page and left off this chart.

Everything we measured, by industry

Almost every site here publishes robots.txt, and around seven in ten publish a sitemap. Nearly four in ten write rules naming individual AI crawlers. After that the grid is nearly empty: everything built specifically for agents is still rare in every industry.

robots.txtsitemapAI-bot rulesllms.txtContent-SignalsUCPOpenAPIMCP cardMCP authTDMRepai.txtagent cardNLWebSoftware or SaaS97762731News or publisher98836010Online shop9880623324Company or brand95742416Reference or directory9662459Entertainment or media967532Marketplace or classifieds98686016Education9668279Gambling947641Nonprofit or association9669229Community or social975748Government or public body945824Finance97782925Local or physical business94773013Parked or placeholder963236Health96773215Infrastructure964023
Tap or hover any cell for what that technology is and why an agent would care. The grid scrolls sideways on a narrow screen. Shading is square-root scaled, because on a linear scale every column past the third would be indistinguishable from empty. Numbers appear above 8%.

Which AI crawlers sites name in robots.txt

  1. gptbot31%reads to train
  2. claudebot27%reads to train
  3. ccbot24%reads to train
  4. google-extended22%reads to train
  5. chatgpt-user20%fetches for a person
  6. bytespider20%reads to train
  7. amazonbot19%reads to train
  8. perplexitybot18%reads to train
  9. anthropic-ai16%reads to train
  10. applebot-extended16%reads to train
  11. meta-externalagent15%reads to train
  12. oai-searchbot13%fetches for a person
  13. perplexity-user5.1%fetches for a person
  14. duckassistbot3.9%fetches for a person
  15. novaact0.1%fetches for a person
  16. operator0.1%fetches for a person
Share of the 33,626 sites that name any crawler. Tap or hover a row for its counts.

35% of sitemaps break the rules

under 100 URLs29%100 to 1,000 URLs25%1,000 to 10,000 URLs34%10,000 to 50,000 URLs45%over 50,000 URLs62%

Sites with a sitemap

61,587

sites publish one. The format was agreed in 2005, so it has had 21 years to settle down.

Sitemaps with a fault

21,448

of them break a rule of the sitemap format somewhere: 35% of everyone who publishes one.

Faults Google overlooks

11,780

of those break only one rule: the sitemap sits in a subfolder and lists pages outside it. Google reads those pages anyway if you submitted the file in Search Console.

Faults no submission fixes

9,668

are wrong however you submit them: one in every six sites that publish a sitemap.

Bigger sitemaps break more often

62%

of sitemaps with over 50,000 URLs have a fault, against 29% of those with under 100 URLs. We could not count the URLs in 1,510 of them, so those are left out of this split.

Which mistakes actually happen

The single commonest mistake is in the sitemap: the file sits in a subfolder and lists pages outside that folder. 14,176 sites do it, 50% of every fault we counted. The sitemap format says those pages do not count, though the big search engines read them anyway if you submitted the file in Search Console.

XML sitemap61,587 sites publish one

  1. lists pages above its own folder14,17623%
  2. lists pages on another domain6,1069.9%
  3. missing the XML namespace2,0773.4%
  4. is not a sitemap at all1,5622.5%
  5. uses relative URLs4920.8%
  6. is over the size limit3050.5%
  7. is an empty file2150.3%
  8. will not parse2020.3%

llms.txt13,730 sites publish one

  1. starts with text before its title1,0227.4%
  2. has no title at all2171.6%
  3. hides a heading inside a collapsed block720.5%

robots.txt85,379 sites publish one

  1. puts a rule before any User-agent line8131.0%
  2. has a line missing its colon2540.3%

OpenAPI description497 sites publish one

  1. describes a value without saying what type it is17635%
  2. leaves an operation with no name to call it by11223%
  3. gives two operations the same name6613%

Per-crawler rules in robots.txt33,626 sites publish one

  1. allows and disallows the same path1340.4%
  2. misspells a crawler's name800.2%

UCP profile2,202 sites publish one

  1. leaves out a registry the profile requires1074.9%

Agent card73 sites publish one

  1. leaves out fields the card is required to carry5474%
Faults firing on 50+ sites, grouped by the file they break. The share is of that file's own publishers. Tap or hover a row for its numbers.

Every technology, and how often it is wrong

Of the sites that published each one, how many got it wrong. Open a row for its faults.

  1. Content Signals4.3% · 64 of 1,495
Of the sites that publish each one, the share with at least one fault.

A bar counts every site with any fault in that file, so it is not a count of the worst fault. 35% of sitemaps have something wrong does not mean 35% are broken XML files. That is 0.3% of them. The usual problem is a sitemap in a subfolder that lists pages outside it, which the format says do not count even though the big search engines read them anyway. Open a row to see which faults actually fire.

Who refuses an identified agent

14,227 of 150,074 sites turned us away with a 401, 403 or 429. We know because we said who we were. A crawler dressed as a browser cannot measure this at all.

How this was measured

We checked 150,074 hosts, taken from the top 150,451 names in the Tranco research ranking, as it stood on 14 September 2026, between 17 September 2026 and 19 September 2026. Tranco combines several popularity sources over 30 days, some counting how often a name is looked up and some how often people visit, so a good part of the list is infrastructure rather than websites. The last 377 names in that range have no row: the run stops once it has enough hosts and its workers take the list in turns, so they do not stop on the same rank. Nothing was left out for anything to do with the sites themselves.

One question, asked of every site that tried. When a site publishes a file or an endpoint meant for AI agents, how often is it implemented wrongly, and in what way. Every figure on this page is counted over the sites that published the thing being measured, never over the corpus.

What this study does not claim. It does not say what share of the web has adopted anything. The list behind it ranks domain names by popularity, not by whether they are websites: 27,992 of these names never replied and another 14,227 turned us away, and much of the rest is not a website at all but nameservers, content delivery, API hosts and parked names. An adoption rate needs a denominator of sites, and this frame gives a denominator of hosts that answered a request.

We used the same checks Lumar runs on a customer's own site, not a special version written for this study. Every fault counted here is one the product would report. Files are fetched; no page is rendered and no JavaScript is run.

Breaking a standard and being awkward for an agent are not the same thing, and we count both. Of the 116 faults we can report, 48 break a rule in a published specification. The other 68 are things a specification allows but an agent cannot work with. A missing operation ID in an OpenAPI description is the clear case: the specification does not require one, so the document is valid, but an agent choosing a call has nothing stable to name it by. Shapes a specification explicitly permits are recorded and never counted as faults, which is why a robots.txt naming one crawler in two groups is absent from these charts: RFC 9309 tells crawlers to merge them, so every rule still applies.

We told every site who we were. Lumar's crawler normally identifies itself as Googlebot-compatible, and sites check that claim by looking up who owns the visiting address. Ours is not Google, so the check fails and the visit looks like something pretending to be Google. On a trial of 36 sites, the Googlebot-style name was turned away by 8 of them and this one by 2. Being honest got us more data, not less.

Every dot is a hundred hosts

150,074

names from the Tranco ranking, tried in rank order. A popularity ranking of domain names, which is not the same as a list of websites.

Not all of them answer

107,855

of the 150,074 answered our knock. 27,992 never replied and 14,227 turned us away, and nothing is claimed about any of them.

Answering is not the same as letting us read

100,008

let us read something. The rest answered a home page and then refused every file we asked for, which proves nothing about what they publish.

This is the population every rate is counted over

88,484

published at least one file or endpoint meant for agents. Not a share of the web, but a count of the sites that tried.

And this is what the study is about

24,119

of them got something wrong. 27% of everyone who tried.

107,855 of 150,074 sites answered our knock.

What became of the rest. Answering the home page is not the same as letting us read the files, and this chart counts only the first.

  1. Answered normally69% · 103,869
  2. Answered, with an error2.7% · 3,986
  3. Turned us away9.5% · 14,227
  4. Never answered19% · 27,992

A technology is only given a percentage once at least 400 sites publish it, and is shown as a plain count below that. It is a reporting threshold, not a significance test: this is not a random sample of the web, so the point is simply not to let a reader over-read a handful of sites. Turning us away is us failing to look rather than a site publishing nothing, so no absence is ever counted against a host that refused us. What a host did serve is counted like anyone else's, even when it refused the first knock at its home page.

How the industries were assigned. Each host was put in exactly one category by a language model reading its home page, falling back to the domain name where no page could be read. It placed 84,149 of the 88,484 publishing hosts; the 4,335 it could not place are left off the industry charts and counted everywhere else. A category appears on the industry charts only once 900 of its hosts publish something we can check. The categories are ours, not a standard taxonomy, and no accuracy audit has been run on them, so treat an industry comparison as indicative rather than measured. Read the ends of that chart, not the middle: a gap of a point or two between neighbouring rows is well inside what a mislabelled host or two could move.

A rule about AI is not the same as a rule against it. What this study can see is that a site named a crawler in robots.txt, or published a content signal, an ai.txt or a mining reservation. It cannot see which way those rules point: a group headed by a crawler's name is just as valid with an Allow in it as with a Disallow, and Content-Signals, ai.txt and TDMRep all state preferences that may permit as easily as refuse. So the figures here count sites that have decided to have a policy, not sites that have shut the door. Restricting a crawler and publishing an interface are not opposites either: doing both is a coherent position, and arguably the considered one.

Using these figures. The findings on this page are published under CC BY 4.0: quote them, chart them and build on them, including commercially, as long as you credit Lumar and link back. We publish the findings rather than the underlying rows, so there is no data file to download.