SectorsReference or directory
Agent technologies for reference sites
For reference sites, 11 agent technologies are standard, 4 are worth building and 8 are worth watching. 45% of the reference sites we scanned publish AI-bot rules, and 37% of their sitemaps are broken. We answered the questions below for a typical site in this sector; change any that are wrong for yours and the list updates.
Where to start
- Fix what you already publish. Scan your site: agents read these files today, and a broken one tells them something false. In our study, 37% of reference sites with a sitemap had a broken one.
- Cover what is standard. Anything under "Should have already" you do not have yet, easiest first.
- Then choose by cost. Pick from "Worth building" by what each one takes to build and to run. Leave the rest on watch.
Should have already 11
Standard equipment for a site like yours. If you have it, scan your site to check it is right.
AI-bot rulesNamed rules for GPTBot, ClaudeBot and friends, rather than one blanket policy.lownone45%
What it takes
- A decision on which AI crawlers may read the site
- A rule for each of them in robots.txt
Where it stands
Platform default, rising. 45% of reference sites have it · 0.6% of all published ones are broken.
Canonical URLTells an agent which address is the real one for this page.lownone
What it takes
- A canonical link tag in each page template, pointing at the page's preferred address
Where it stands
Web baseline, flat.
Content SignalsCloudflare's robots.txt vocabulary for search, AI input and AI training.lownone2.1%
What it takes
- One line in robots.txt saying how your content may be used
- Or Cloudflare's managed robots.txt, which adds it
Where it stands
Platform default, rising. 2.1% of reference sites have it · 4.3% of all published ones are broken.
Editorial metadataAuthor, dates and provenance behind the content.lownone
What it takes
- Author, publisher and dates in the structured data of each article
Where it stands
Web baseline, flat.
HTTP Link headerRelations declared in headers, readable without parsing the body.lownone
What it takes
- A server or CDN rule that adds Link headers for the page's alternates
Where it stands
Web baseline, flat.
JSON-LDStructured data describing what this page is about.lownone
What it takes
- Structured data in each page template, describing what the page is about
Where it stands
Web baseline, flat.
Language and encodingDeclared language and charset, so text is interpreted correctly.lownone
What it takes
- A lang attribute and a character set declared on every page
Where it stands
Web baseline, flat.
robots.txtThe crawl-policy file every agent reads first.lownone96%
What it takes
- A robots.txt file at the root of the site
Where it stands
Web baseline, flat. 96% of reference sites have it · 1.3% of all published ones are broken.
schema:ActionDeclares actions an agent could take from this page.lownone
What it takes
- Markup describing the actions a page supports, such as search or booking
Where it stands
Web baseline, flat.
XML sitemapThe URL inventory agents and crawlers use for discovery.lownone62%
What it takes
- A sitemap your CMS or a build step generates
- A job that keeps it current as pages change
Where it stands
Web baseline, flat. 62% of reference sites have it · 37% of theirs are broken.
Server-side renderingWhether the content exists before JavaScript runs. Most agents never run it.highlow
What it takes
- Page content in the HTML the server sends, not only after scripts run
- For a site that renders in the browser today, a framework change
Where it stands
Web baseline, flat.
Worth building 4
In production at other organisations and gaining ground. Weigh each one by what it takes.
TDMRepThe W3C text and data mining reservation, used mainly by publishers.lownone1%
What it takes
- A decision on whether you reserve text and data mining rights
- A JSON file or header saying so, with a link to your licensing terms
Where it stands
Early production, rising. 1% of reference sites have it · 19% of all published ones are broken.
llms.txtA hand-written Markdown index of a site's most useful pages, written for language models.lowlow9.5%
What it takes
- A Markdown file listing your most useful pages
- Someone to keep it current
Where it stands
Early production, rising. 9.5% of reference sites have it · 9.4% of all published ones are broken.
Markdown negotiationMarkdown served to agents that ask for it with Accept: text/markdown.mediumlow
What it takes
- A server or CDN that answers a request for Markdown with Markdown
- A Vary: Accept header so caches keep the two versions apart
Where it stands
Early production, rising.
Markdown twinA clean low-token markdown version of the page for agents.mediumlow
What it takes
- A build step that turns each page into Markdown
- The Markdown served at the page's address with .md added
Where it stands
Early production, rising.
Keep an eye on 8
It fits your site, but it is early or still settling. Nothing to build yet; worth knowing about.
A2A agent cardNeeds an A2A agentDeclares an agent this site operates, for agent-to-agent work.lownone0.2%
What it takes
- A JSON file at /.well-known/agent-card.json describing your agent's skills and address
It describes an A2A agent you already run. Building and hosting that agent is a separate project, and the bigger one.
Where it stands
Early pilots, rising. 11 of 5,642 reference sites have it · 79% of all published ones are broken.
ai.txtAttribution preferences for AI systems that use the content.lownone0.2%
What it takes
- A text file at the root of the site stating your AI usage terms
Where it stands
Specification only, flat. 11 of 5,642 reference sites have it.
AIPREF Content-UsageThe IETF's standards-track vocabulary for stating how content may be used by AI systems.lownone
What it takes
- Usage preferences declared in robots.txt or an HTTP header, once the standard settles
Where it stands
Specification only, rising.
auth.mdTelling an agent how to register for an account and get its own credentials, without a signup form or a consent screen.lownone
What it takes
- A Markdown file explaining how agents should sign in
- Only worth writing once the convention settles
Where it stands
Early pilots, rising.
MCP server cardNeeds an MCP serverAdvertises a Model Context Protocol endpoint agents can call.lownone0.5%
What it takes
- A JSON file at /.well-known/mcp.json giving your server's address, transports and tools
It describes an MCP server you already run. Building, hosting and securing that server is a separate project, and the bigger one.
Where it stands
Early pilots, rising. 0.5% of reference sites have it · 22% of all published ones are broken.
WebMCPIn-page tools a browsing agent can call directly.mediumlow
What it takes
- Front-end code that registers your page's actions as tools
- Tool design: names, descriptions and inputs an agent can rely on
Where it stands
Early pilots, rising.
MCP authorizationNeeds an MCP serverThe OAuth discovery chain protecting that MCP endpoint.highmedium<0.1%
What it takes
- An OAuth authorization server, or a provider that runs one
- Protected-resource metadata so agents can find it
- Token checks on every MCP request
It protects an MCP server you already run, which is a separate project. The authorization itself is the effort rated here.
Where it stands
Early production, rising. 3 of 5,642 reference sites have it · 37% of all published ones are broken.
NLWeb endpointMicrosoft's conversational /ask route returning schema.org answers.highhigh0%
What it takes
- A vector database
- An embedding model
- A pipeline that keeps the embeddings in sync with your content
- A language model to write the answers, paid per question
- An /ask endpoint to host and scale
Where it stands
Early pilots, receding. None of the 5,642 reference sites we scanned have it yet.
Not for you yet 9
These do not fit your answers above. Change an answer and they move.
- Only if you sell online. Product / Offer schema, UCP profile, ACP, AP2
- Only if you offer an API. API catalog, OpenAPI description, MPP (Machine Payments Protocol), x402
- Only if you run bots or agents of your own. Web Bot Auth directory
How we decide
- Fit comes from your answers. Each technology is for one kind of site. We answered the questions for a typical site in your sector; change any that are wrong for you.
- The group comes from how far along it is. Web baseline and platform default are standard; early production is worth building; earlier than that, or receding, is one to watch.
- Few peers is not a reason to skip. The share of your sector that publishes something is shown for context and never moves it down: early is not the same as irrelevant.
- Effort and running cost are our estimates for a typical site on a mainstream platform. Yours may differ.
- Peer numbers come from our study of 150,074 domains, scanned in September 2026. Cautions for regulated sectors are our judgment, not legal advice.