SectorsNews or publisher
Agent technologies for news sites
For news sites, 10 agent technologies are standard, 4 are worth building and 3 are worth watching. 60% of the news sites we scanned publish AI-bot rules, and 45% of their sitemaps are broken. We answered the questions below for a typical site in this sector; change any that are wrong for yours and the list updates.
Where to start
- Fix what you already publish. Scan your site: agents read these files today, and a broken one tells them something false. In our study, 45% of news sites with a sitemap had a broken one.
- Cover what is standard. Anything under "Should have already" you do not have yet, easiest first.
- Then choose by cost. Pick from "Worth building" by what each one takes to build and to run. Leave the rest on watch.
Should have already 10
Standard equipment for a site like yours. If you have it, scan your site to check it is right.
AI-bot rulesNamed rules for GPTBot, ClaudeBot and friends, rather than one blanket policy.lownone60%
What it takes
- A decision on which AI crawlers may read the site
- A rule for each of them in robots.txt
Where it stands
Platform default, rising. 60% of news sites have it · 0.6% of all published ones are broken.
Canonical URLTells an agent which address is the real one for this page.lownone
What it takes
- A canonical link tag in each page template, pointing at the page's preferred address
Where it stands
Web baseline, flat.
Content SignalsCloudflare's robots.txt vocabulary for search, AI input and AI training.lownone2%
What it takes
- One line in robots.txt saying how your content may be used
- Or Cloudflare's managed robots.txt, which adds it
Where it stands
Platform default, rising. 2% of news sites have it · 4.3% of all published ones are broken.
Editorial metadataAuthor, dates and provenance behind the content.lownone
What it takes
- Author, publisher and dates in the structured data of each article
Where it stands
Web baseline, flat.
HTTP Link headerRelations declared in headers, readable without parsing the body.lownone
What it takes
- A server or CDN rule that adds Link headers for the page's alternates
Where it stands
Web baseline, flat.
JSON-LDStructured data describing what this page is about.lownone
What it takes
- Structured data in each page template, describing what the page is about
Where it stands
Web baseline, flat.
Language and encodingDeclared language and charset, so text is interpreted correctly.lownone
What it takes
- A lang attribute and a character set declared on every page
Where it stands
Web baseline, flat.
robots.txtThe crawl-policy file every agent reads first.lownone98%
What it takes
- A robots.txt file at the root of the site
Where it stands
Web baseline, flat. 98% of news sites have it · 1.3% of all published ones are broken.
XML sitemapThe URL inventory agents and crawlers use for discovery.lownone83%
What it takes
- A sitemap your CMS or a build step generates
- A job that keeps it current as pages change
Where it stands
Web baseline, flat. 83% of news sites have it · 45% of theirs are broken.
Server-side renderingWhether the content exists before JavaScript runs. Most agents never run it.highlow
What it takes
- Page content in the HTML the server sends, not only after scripts run
- For a site that renders in the browser today, a framework change
Where it stands
Web baseline, flat.
Worth building 4
In production at other organisations and gaining ground. Weigh each one by what it takes.
TDMRepThe W3C text and data mining reservation, used mainly by publishers.lownone0.9%
What it takes
- A decision on whether you reserve text and data mining rights
- A JSON file or header saying so, with a link to your licensing terms
Where it stands
Early production, rising. 0.9% of news sites have it · 19% of all published ones are broken.
llms.txtA hand-written Markdown index of a site's most useful pages, written for language models.lowlow9.8%
What it takes
- A Markdown file listing your most useful pages
- Someone to keep it current
Where it stands
Early production, rising. 9.8% of news sites have it · 9.4% of all published ones are broken.
Markdown negotiationMarkdown served to agents that ask for it with Accept: text/markdown.mediumlow
What it takes
- A server or CDN that answers a request for Markdown with Markdown
- A Vary: Accept header so caches keep the two versions apart
Where it stands
Early production, rising.
Markdown twinA clean low-token markdown version of the page for agents.mediumlow
What it takes
- A build step that turns each page into Markdown
- The Markdown served at the page's address with .md added
Where it stands
Early production, rising.
Keep an eye on 3
It fits your site, but it is early or still settling. Nothing to build yet; worth knowing about.
ai.txtAttribution preferences for AI systems that use the content.lownone0.3%
What it takes
- A text file at the root of the site stating your AI usage terms
Where it stands
Specification only, flat. 0.3% of news sites have it.
AIPREF Content-UsageThe IETF's standards-track vocabulary for stating how content may be used by AI systems.lownone
What it takes
- Usage preferences declared in robots.txt or an HTTP header, once the standard settles
Where it stands
Specification only, rising.
NLWeb endpointMicrosoft's conversational /ask route returning schema.org answers.highhigh<0.1%
What it takes
- A vector database
- An embedding model
- A pipeline that keeps the embeddings in sync with your content
- A language model to write the answers, paid per question
- An /ask endpoint to host and scale
Where it stands
Early pilots, receding. 2 of 9,364 news sites have it.
Not for you yet 15
These do not fit your answers above. Change an answer and they move.
- Only if you sell online. Product / Offer schema, UCP profile, ACP, AP2
- Only if visitors can do things on the site an agent could do for them. A2A agent card, auth.md, MCP server card, schema:Action, WebMCP, MCP authorization
- Only if you offer an API. API catalog, OpenAPI description, MPP (Machine Payments Protocol), x402
- Only if you run bots or agents of your own. Web Bot Auth directory
How we decide
- Fit comes from your answers. Each technology is for one kind of site. We answered the questions for a typical site in your sector; change any that are wrong for you.
- The group comes from how far along it is. Web baseline and platform default are standard; early production is worth building; earlier than that, or receding, is one to watch.
- Few peers is not a reason to skip. The share of your sector that publishes something is shown for context and never moves it down: early is not the same as irrelevant.
- Effort and running cost are our estimates for a typical site on a mainstream platform. Yours may differ.
- Peer numbers come from our study of 150,074 domains, scanned in September 2026. Cautions for regulated sectors are our judgment, not legal advice.