Scan at scale
Scan at scale
The scanner on this site checks one URL at a time. The same checks run as a Lumar custom metric container across an entire estate, and arrive as 34 predefined reports you can filter, trend and assign.
One container, every page and every host
Everything the single-URL scanner does is a Lumar custom metric container (id 2246) that runs inside a normal crawl. Link it to a project and every crawl scores agent readiness alongside the SEO, accessibility and site-speed metrics already being collected. There is nothing separate to schedule and no second crawl to pay for.
The checks land in two places, because agent-facing technologies live in two places. Page signals vary per URL — a canonical, Product schema, an in-page tool declaration — so they attach to the page row, and you filter them like any other page metric. Site-wide signals are one fact per host — robots.txt, an llms.txt, a sitemap, an MCP server card, a live probe of the endpoint it advertises — so they are fetched once per host and arrive as their own rows rather than being copied onto every URL.
That split is what makes this workable on a large site. A million-page crawl does not fetch robots.txt a million times, and a host-level fault is reported once rather than as a million identical page issues.
It reports what is broken, never what is absent
No report on this page will tell you that you have not adopted something. These are emerging technologies and most sites have not adopted most of them, which is not a finding. Every report is gated on the technology actually being present, so a page enters one only when the site tried to implement something and got it wrong.
That is what makes the output a work queue rather than a scorecard. A row is a thing a team published, that does not do what it was published to do, with the specific defect named — a relative canonical, a sitemap listing URLs outside its own folder, an agent card missing a required field, an MCP endpoint that does not answer the protocol its card advertises.
Where we could not check something, the row says so rather than guessing. A timeout or a blocked connection is us failing to look, and it is reported as unassessed, never as a fault on the site.
The 34 reports
Grouped the way they appear in Lumar. Each one opens a page explaining what it returns, which columns it shows, and which technology it is about.
Overview
Discoverability
- Pages with a broken canonicalper-URL
- Pages with a relative canonical URLper-URL
- Pages with broken Link headersper-URL
- Sites with a broken sitemapsite-wide
- Sites with a broken llms.txtsite-wide
Content
- Pages with invalid JSON-LDper-URL
- Pages with unparseable JSON-LDper-URL
- Pages whose content only appears after JavaScriptper-URL
- Pages with a broken Markdown twinper-URL
- Pages with missing or invalid language/encoding declarationsper-URL
Access
- Sites with a malformed robots.txtsite-wide
- Sites with broken AI bot rulessite-wide
- Sites with a misspelled AI bot tokensite-wide
- Sites with invalid Content Signalssite-wide
- Sites with an invalid ai.txtsite-wide
- Sites with an invalid TDM reservation (TDMRep)site-wide
Capabilities
- Pages with broken WebMCP toolsper-URL
- Pages with undescribed WebMCP toolsper-URL
- Pages with a broken schema.org Action (potentialAction)per-URL
- Sites with a broken MCP server cardsite-wide
- Sites with an invalid OpenAPI contractsite-wide
- Sites with a broken NLWeb /ask endpointsite-wide
- Sites with a broken NLWeb /mcp serversite-wide
- Sites with broken MCP authorization discoverysite-wide
- Sites whose MCP authorization a conforming agent must refusesite-wide
Commerce
- Pages with broken Product schemaper-URL
- Products missing core Offer fieldsper-URL
- Sites with a broken UCP profilesite-wide
Trust
Running it on your own site
The container is available to Lumar accounts and runs on crawls you are already scheduling. If you want to see what it finds on one URL first, the scanner on this site is the same code and needs no account.
You can also read what each technology is and how it is checked, or the measurements we have published from running these checks across 150,074 domains.