Measured behaviour, not opinion.
A run fetches a bounded set of a site’s public URLs once: the homepage, the pages its sitemap and links lead to, robots.txt, and the well-known documents an agent looks for. Every probe is a pure function over that one set of responses. A probe never fetches, never asks a model, and never states a cause. What it records is a status code, a header, a parsed field: a fact you can reproduce with curl.
A probe applies to a site because of what the site is or because of a capability it publicly advertises. A site that publishes no API is not measured on API probes. A site that publishes one badly is measured on them and fails them, because a half-built public surface is a new liability for an agent, not a neutral act.
A probe for a young standard ships as emerging: reported in full, with its evidence and its fix, and scoring nothing. It joins the figure only at a dated catalogue version, and only once Heading’s own citation data shows AI answers depend on it. That is the one instrument this measurement has that a checklist does not.
Four layers, one figure, a grade.
Each probe belongs to one of four layers. Within a layer, the points a site earned are divided by the points available to it and multiplied by the layer’s weight. The figure is the sum over the layers that applied, normalised to 100 over their weights. A layer nothing applied to is left out rather than scored zero.
The figure, 0 to 100
Σ (layer points earned ÷ layer points available × layer weight) ÷ Σ applicable layer weights × 100| Grade | Score | Reads as |
|---|---|---|
| A+ | 95 and above | Exemplary |
| A | 86 and above | Strong |
| B | 70 and above | Competitive |
| C | 48 and above | Developing |
| D | 28 and above | Limited |
| F | below 28 | Not ready |
A bonus probe is upside only: earned, it adds to both sides of its layer; unearned, it leaves both, so a site is never marked down for a convention almost nobody serves yet. A probe that applied and could not be read stays in the figure and earns nothing, so a site that breaks our fetches cannot score well on the probes that happened to work. Each failing probe carries the points a full pass adds to the figure, and fixes are ranked by that number.
Can an agent or an answer engine find this brand at all?
Discovery, weight 207 probes, 1 scored
AI crawler access stated
Required2 pointsEvery websiterobots-ai-policy
- Why an agent needs it
- robots.txt is where a site states which AI crawlers may read it. A crawler disallowed at the root does not fetch the site, and an answer engine built on that crawler has nothing of the site to work from.
- How it is measured
- robots.txt was read and evaluated at the root path for each AI crawler user agent. It passes when no retrieval crawler is disallowed, and a disallow is reported with the file's own lines quoted. Training-control tokens are reported and do not decide the outcome.
- When it fails
- Remove the root-path Disallow rule for each AI crawler named in the evidence, or keep the rule as a deliberate access policy. This probe reports the file and does not obey it.
- Specification
- https://www.rfc-editor.org/rfc/rfc9309
Resource catalogue published
Required1 pointEvery surface the site exposesard-catalog
- Why an agent needs it
- A resource catalogue at a well-known path tells an agent, in one document, which documents the site publishes for machines: its API description, its llms.txt, its MCP server. Without one an agent tries the conventional paths one by one.
- How it is measured
- /.well-known/ard.json and then /.well-known/ai-catalog.json were requested. It passes when one of them is served as parseable JSON carrying at least one entry.
- When it fails
- Publish an Agentic Resource Discovery manifest at /.well-known/ard.json listing each machine-facing document the site serves, with its URL and its type.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
AI catalogue published
Recommended1 pointEvery surface the site exposesai-catalog-published
- Why an agent needs it
- The AI catalogue was the first well-known address for a site's machine-readable inventory, and agents still try it. A site that serves one is found by the clients that look there.
- How it is measured
- /.well-known/ai-catalog.json was requested. It passes when the document is served as parseable JSON.
- When it fails
- Serve /.well-known/ai-catalog.json as parseable JSON, or publish the newer /.well-known/ard.json in its place.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Catalogue entries are usable
Recommended2 pointsEvery surface the site exposesard-entries-valid
- Why an agent needs it
- A catalogue entry with a relative path or no type leaves an agent guessing where the document is and what it is for. A usable entry is one it can fetch and file without reading anything else.
- How it is measured
- Each entry of the served catalogue was read for a URL and a type. It passes when every entry carries an absolute http(s) URL and a non-empty type.
- Not applicable when
- No resource catalogue was served.
- When it fails
- Give every catalogue entry an absolute URL and a type naming what the document is.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Catalogue names a trust manifest
Recommended2 pointsEvery surface the site exposesard-trust-manifest
- Why an agent needs it
- A trust manifest is how an agent tells a site's own catalogue from one somebody else published about it. A catalogue that names one gives the agent something to verify rather than something to take on faith.
- How it is measured
- The served catalogue was read for a trust, trustManifest or attestation field. It passes when the field names an absolute URL.
- Not applicable when
- No resource catalogue was served.
- When it fails
- Add a trust reference to the catalogue naming the URL of a signed manifest that attests to the site's identity, or leave it out until the site operates a signing key.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Own site ranks for the brand's name
Required3 pointsEvery websitebrand-search-accuracy
- Why an agent needs it
- An answer engine that is asked about a brand searches for its name first. If the brand's own site is not among the results, the engine describes whatever is: a marketplace listing, a same-name business, a review site.
- How it is measured
- One search-engine read was made for the brand's name, and the top ten organic results were recorded. It passes when the brand's own domain is among them. The source is read once per run and the result is stored; no search engine is asked again on read.
- Not applicable when
- No brand name was available to search for, so no search snapshot was taken.
- When it fails
- Make the brand's own site the first answer to a search for the brand's name: a homepage title naming the brand, Organization structured data, and the same name on the site's own profiles.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Public reference article names the site
Recommended4 pointsEvery websitewikipedia-presence
- Why an agent needs it
- A public reference article is the record an answer engine trusts most when it describes a business, and the knowledge base behind it is where the engine learns which website is the business's own. A brand with neither is described from whatever else the engine finds.
- How it is measured
- The encyclopaedia's search was read for the brand's name, up to five articles. Each article's knowledge-base record was read for its official-website claim and compared with the brand's domain. It passes when one article's claim names the domain. A same-name article naming another site does not pass.
- Not applicable when
- No brand name was available to search for, so no reference read was made.
- When it fails
- Where the business meets the encyclopaedia's notability rules, see that its article exists and that its knowledge-base record names the site as the official website. Nothing on the site itself changes this.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Once found, can a machine read the site and its documents?
Access, weight 3037 probes, 11 scored
Sitemap published
Required2 pointsEvery websitesitemap-exists
- Why an agent needs it
- A sitemap is the one list of a site's URLs an agent can read in a single request. Without it, an agent finds pages only by following links, and misses the ones nothing links to.
- How it is measured
- A sitemap was requested from every Sitemap: directive in robots.txt and from /sitemap.xml. It passes when one of them returned URL entries.
- When it fails
- Publish a sitemap at /sitemap.xml and name it in a Sitemap: directive in robots.txt.
- Specification
- https://www.sitemaps.org/protocol.html
Sitemap entries dated
Recommended1 pointEvery websitesitemap-freshness
- Why an agent needs it
- A dated sitemap entry tells an agent which pages changed and which it can trust from a previous read. Without dates, every page has to be fetched again to find out.
- How it is measured
- The sitemap's own lastmod elements were counted across its entries. It passes when at least half of them publish one.
- Not applicable when
- No sitemap with URL entries was found.
- When it fails
- Serve a Last-Modified header on the pages listed in the sitemap, and give each sitemap entry a lastmod date.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.sitemaps.org/protocol.html
llms.txt published
Required1 pointEvery websitellms-txt-exists
- Why an agent needs it
- An llms.txt is a map of the site written for a model: which pages describe the business and what each one says. An agent that finds it reads the right pages first instead of guessing.
- How it is measured
- /llms.txt was requested. It passes when the file was served with content.
- When it fails
- Publish a plain-text /llms.txt listing the pages that describe what the business does, each with a one-line summary.
- Specification
- https://llmstxt.org/
llms.txt is a usable map
Recommended2 pointsEvery websitellms-txt-quality
- Why an agent needs it
- An llms.txt that is empty or unstructured points an agent nowhere. A title and annotated links are what make the file a map rather than a placeholder.
- How it is measured
- The /llms.txt that was served was read as text. It passes when the file carries an H1 title and at least three markdown links, which is what separates a curated map from an empty file or a stub.
- Not applicable when
- No /llms.txt was served.
- When it fails
- Give /llms.txt an H1 title naming the business and a list of markdown links to the pages that describe it, each with a one-line summary.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://llmstxt.org/
llms.txt links resolve
Recommended2 pointsEvery websitellms-txt-links-resolve
- Why an agent needs it
- A map is only useful if the addresses on it answer. A link in llms.txt that returns an error sends an agent to a page that is not there.
- How it is measured
- Links advertised in /llms.txt were matched against the pages this run already fetched. It passes when every matched link answered. A link no page in this run covers is reported as unknown and is not counted either way.
- Not applicable when
- No link in /llms.txt matched a page this Run fetched.
- When it fails
- Point every link in /llms.txt at a URL that answers with the page, or remove the link from the file.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://llmstxt.org/
Agent discovery document published
Required2 pointsEvery websiteagent-discovery-file
- Why an agent needs it
- An agent card or server descriptor at a well-known path tells an agent what the site offers it and where to send requests, with no documentation page in between. It is the one document on a site addressed to an agent rather than to a browser.
- How it is measured
- /.well-known/agent.json, /.well-known/mcp.json and /.well-known/ai-plugin.json were requested. It passes when one of them was served with content.
- When it fails
- Publish an agent card at /.well-known/agent.json naming what the site offers an agent, where to send requests and how to authenticate.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://a2a-protocol.org/latest/specification/
Link header advertises documents
Required1 pointEvery websitelink-header-discovery
- Why an agent needs it
- A Link header arrives with the response an agent already asked for, so it costs nothing to read. Naming the sitemap or a description there is the cheapest way to point an agent at the site's machine-readable documents.
- How it is measured
- The homepage response headers were read for a Link header. It passes when one advertises a sitemap, a description, a service document, an API catalogue, or an alternate representation in a machine-readable media type.
- Not applicable when
- The homepage returned no response to read headers from.
- When it fails
- Send a Link header on the homepage naming the sitemap (rel="sitemap") and any machine-readable description of the site (rel="describedby").
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.rfc-editor.org/rfc/rfc8288
A2A agent card published
Recommended2 pointsEvery websitea2a-agent-card
- Why an agent needs it
- An agent card is how one agent learns what another offers and where to send it work, in the shape the A2A protocol fixes. A site that serves one is reachable by that protocol's clients without a person in between.
- How it is measured
- /.well-known/agent-card.json was requested. It passes when the document is served as parseable JSON carrying a name and an absolute URL, on the card itself or on its first supported interface.
- When it fails
- Publish an A2A agent card at /.well-known/agent-card.json naming the agent and the URL of at least one interface an agent can reach it on.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://a2a-protocol.org/latest/specification/
Agent documents say when to use the site
Required3 pointsEvery websiteagent-instruction
- Why an agent needs it
- A document an agent reads about a site is only useful if it says what to do with the site and when. A list of links with no guidance leaves the agent to infer the site's purpose from its pages.
- How it is measured
- The llms.txt, agent card and MCP descriptor the site serves were read for guidance addressed to an agent: a line naming when to use the site, what an agent can do with it, or its capabilities. It passes when at least one document carries such a line.
- Not applicable when
- The site serves no document addressed to an agent: no llms.txt, agent card or MCP descriptor.
- When it fails
- Add a short section to llms.txt, the agent card or the MCP descriptor saying what an agent can do with the site and when to use it.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
User-triggered fetchers allowed
Required2 pointsEvery websiterobots-agent-user-policy
- Why an agent needs it
- When a person asks an assistant to read a page, the assistant fetches it under one of these user agents. A root-path disallow for them means the assistant cannot read the site even when a person asks it to.
- How it is measured
- robots.txt was evaluated at the root path for ChatGPT-User, Claude-User and Perplexity-User, the user agents an assistant sends when a person asks it to read a page. It passes when none is disallowed, and a disallow is reported with the file's own lines quoted.
- When it fails
- Remove the root-path Disallow rule for each user-triggered fetcher named in the evidence, or keep it as a deliberate policy. This probe reports the file and does not obey it.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.rfc-editor.org/rfc/rfc9309
Schema feeds named in robots.txt
Emerging1 pointEvery websitenlweb-schema-feeds
- Why an agent needs it
- A schema feed is a bulk, structured copy of what the site sells or publishes, and the Schemamap directive is how a site tells a client where it is. A client that finds one reads the site's items without crawling its pages.
- How it is measured
- robots.txt was read for Schemamap: directives. It passes when at least one names a URL.
- When it fails
- Add a Schemamap: directive to robots.txt naming the URL of a schema.org feed of the site's items.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
llms.txt is split by section
Emerging1 pointEvery websitemodular-llms-txt
- Why an agent needs it
- One llms.txt for a large site is a document too long for an agent to read in full. Section files let it read the part about pricing, or the API, and skip the rest.
- How it is measured
- The served llms.txt was read for links whose target is itself an llms.txt or llms-full.txt. It passes when at least one is linked.
- Not applicable when
- No /llms.txt was served.
- When it fails
- Split a large llms.txt into one file per section and link each from the root file, so an agent reads the part it needs.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://llmstxt.org/
Readable without JavaScript
Required3 pointsEvery websitecontent-without-javascript
- Why an agent needs it
- An agent reads the HTML the server returns and runs no JavaScript. A page whose text is rendered by a script is an empty page to it.
- How it is measured
- Every crawled page's raw HTML was measured for the text a recipient running no JavaScript would read. It passes when at least half the crawled pages carry 600 characters or more of that text, and the homepage does too.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Render each page's main text into the HTML the server returns, so the text is in the response body before any script runs.
Text against markup
Recommended1 pointEvery websitecontent-efficiency
- Why an agent needs it
- An agent reads a page within a fixed budget of text. A response that is mostly markup, styles and payload spends that budget before the page says anything.
- How it is measured
- The crawled pages' visible text was divided by the bytes the origin sent. It passes when text is at least 5% of the response across the crawl.
- Not applicable when
- No crawled page returned a body to measure.
- When it fails
- Move inlined data, styles and scripts out of the document, and keep the page's text in the markup the origin returns.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Heading structure
Recommended2 pointsEvery websiteheading-structure
- Why an agent needs it
- One h1 per page tells an agent what the page is about before it reads the body. Headings are the outline it uses to find the section that answers its question.
- How it is measured
- Each crawled page's <h1> elements were counted. It passes when at least 80% of the crawled pages carry exactly one.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Give each page one <h1> naming what that page is, and mark the sections below it with <h2> and <h3>.
Title and description
Required2 pointsEvery websitemetadata-completeness
- Why an agent needs it
- The title and the meta description are the first thing an agent reads about a page, and often the only thing it reads before deciding whether to open it. A one-word title or a missing description leaves that decision to chance.
- How it is measured
- Each crawled page's <title> and meta description were read. It passes when at least 80% of the crawled pages carry a title of 10 characters or more and a description of 50 characters or more.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Give each page a <title> naming the page and the business, and a meta description of one or two sentences stating what is on it.
Markdown on request
Emerging1 pointEvery websitemarkdown-negotiation
- Why an agent needs it
- Markdown carries a page's text with almost none of its markup. A site that serves it on request hands an agent the content in the form it reads best.
- How it is measured
- The homepage was re-requested with Accept: text/markdown. It passes when the origin answers with a markdown content type. Whether the response also varies on Accept is the next probe's question.
- Not applicable when
- The homepage answered none of the requests naming a machine-readable form.
- When it fails
- Serve a markdown form of each page to a request whose Accept header names text/markdown.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Redirect chains
Recommended1 pointEvery websiteredirect-hygiene
- Why an agent needs it
- Every redirect hop is a request an agent makes before it reads anything. A long chain spends its budget, and a loop ends the visit.
- How it is measured
- The redirect chain of the homepage and every crawled page was measured. It passes when no URL takes more than two hops and none of them looped.
- Not applicable when
- No fetched URL returned a response to read a redirect chain from.
- When it fails
- Point each redirect at its final URL in one hop, and remove the rules that send a URL back to one already in its chain.
Agent view on request
Emerging2 pointsEvery websiteagent-mode-view
- Why an agent needs it
- A page built for a browser carries navigation, scripts and styling an agent has no use for. A site that answers a mode=agent request with the page's substance alone hands the agent less to read and nothing to strip.
- How it is measured
- The homepage was re-requested with ?mode=agent. It passes when the answer is served in a non-HTML type, or as HTML at least 200 characters long and under 90% the size of the plain page.
- Not applicable when
- The homepage or its ?mode=agent request returned no response to compare.
- When it fails
- Answer a request carrying ?mode=agent with a reduced view of the page: its text and its controls, without the chrome, as markdown or as a smaller HTML document.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Pages readable without an account
Recommended2 pointsEvery websitedocs-auth-gate
- Why an agent needs it
- An agent has no account and cannot sign in. A page the sitemap advertises that answers with a login wall is a page the agent was sent to and cannot read.
- How it is measured
- Every crawled page's status and final URL were read. It fails when one answered 401 or 403, or redirected to a sign-in path.
- Not applicable when
- No page was fetched in this Run.
- When it fails
- Serve the pages the sitemap and the homepage link to without a sign-in, or remove the gated URLs from the sitemap so an agent is not sent to them.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Markdown at .md URLs
Emerging2 pointsEvery websitemarkdown-url-fallback
- Why an agent needs it
- A markdown twin at a predictable URL is the simplest way to hand an agent a page's text: no negotiation, no headers, one request it can guess. An agent that finds /about.md reads the page in a tenth of the bytes.
- How it is measured
- /index.md was requested, and the .md twin of up to three readable crawled pages. It passes when any of them is served with a markdown content type.
- When it fails
- Serve a markdown copy of each page at the page's path with .md appended, and the homepage's at /index.md.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Pricing as markdown
Emerging2 pointsEvery websitepricing-md
- Why an agent needs it
- Pricing is the page an agent comparing options reads first, and a pricing page built from tabs and toggles is the hardest kind to read from HTML. A markdown copy at a fixed address hands the agent the plans as a list.
- How it is measured
- /pricing.md was requested. It passes when the document is served with a markdown content type.
- Not applicable when
- No crawled page is named for pricing.
- When it fails
- Serve the pricing page's content as markdown at /pricing.md: each plan, what it includes and what it costs.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Markdown code fences balanced
Emerging1 pointEvery websitecode-fence-validity
- Why an agent needs it
- An unclosed code fence turns the rest of a markdown document into one code block. An agent reading it sees a page of source where the text was.
- How it is measured
- Every markdown document the collection holds was scanned for fence lines (three or more backticks or tildes at the start of a line). It passes when each document carries an even number of them.
- Not applicable when
- No markdown document was served.
- When it fails
- Close every fenced code block in the served markdown with a matching fence line.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Markdown pages carry frontmatter
Emerging1 pointEvery websitemarkdown-frontmatter
- Why an agent needs it
- Frontmatter is where a markdown page states its title, its address and its date in a form a parser reads before the body. Without it an agent has to infer what page it is holding from the prose.
- How it is measured
- Every markdown document the collection holds was read for an opening frontmatter block, a --- line closed by another. It passes when each document carries one.
- Not applicable when
- No markdown document was served.
- When it fails
- Open each served markdown page with a frontmatter block stating at least its title and its canonical URL.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Markdown twin advertised in the page
Emerging1 pointEvery websitemarkdown-link-alternate
- Why an agent needs it
- A link element naming the markdown twin tells an agent where the readable copy is from the page it already fetched, with no guessing at URLs. The target has to answer as markdown, or the link points at nothing.
- How it is measured
- The homepage HTML was read for <link rel="alternate" type="text/markdown">, and the URL it names was requested. It passes when the link exists and its target is served with a markdown content type.
- Not applicable when
- The homepage was not readable.
- When it fails
- Add <link rel="alternate" type="text/markdown" href="..."> to each page naming its markdown twin, and serve that URL with a markdown content type.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Negotiated markdown varies on Accept
Recommended1 pointEvery websitemarkdown-negotiation-vary
- Why an agent needs it
- A cache between the agent and the origin keys on the URL. Without Vary: Accept it hands the HTML to the client that asked for markdown, or the markdown to the browser, and the negotiation stops working the moment a cache is involved.
- How it is measured
- The homepage was re-requested with Accept: text/markdown. It passes when the origin answers with a markdown content type and a Vary header naming Accept.
- Not applicable when
- The homepage answered none of the requests naming a machine-readable form.
- When it fails
- Send Vary: Accept on every response whose body depends on the Accept header, so a cache keeps the HTML and the markdown apart.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.rfc-editor.org/rfc/rfc9110#name-vary
Pages fit one read
Recommended1 pointEvery websitepage-token-budget
- Why an agent needs it
- An agent reads a page within a fixed budget of text, and a page that overruns it is read in part or not at all. The part it reads is the top, which is rarely the part that answers the question.
- How it is measured
- Each crawled page's extracted text was measured. It passes when no page carries more than 100,000 characters of text outside its markup.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Split a page carrying more than about 100,000 characters of text into several pages, each on one subject.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Structured data on the homepage
Required4 pointsEvery websitestructured-data-present
- Why an agent needs it
- JSON-LD is how a site states what it is in a form a machine parses rather than infers. Without it, an agent has to work out from prose whether this is a shop, a clinic or a software company.
- How it is measured
- The homepage HTML was read for <script type="application/ld+json"> blocks. It passes when at least one block parsed as JSON and declared a schema.org @type.
- Not applicable when
- The homepage was not readable.
- When it fails
- Add a <script type="application/ld+json"> block to the homepage HTML describing the business with a schema.org @type.
- Specification
- https://www.w3.org/TR/json-ld11/
Organization details in structured data
Recommended3 pointsEvery websiteorganization-schema-completeness
- Why an agent needs it
- An Organization block names the business and says how to reach it. An agent asked to contact or locate the business reads these fields, and a block carrying a name alone leaves it with nothing to act on.
- How it is measured
- The homepage's JSON-LD was read for an Organization node. It passes when one carries a name, a contact point (a contactPoint node, a telephone or an email) and a postal address. Each of the three earns a point on its own.
- Not applicable when
- The homepage published no structured data.
- When it fails
- Give the homepage's Organization block a name, a contactPoint with a telephone or an email, and a postal address.
- Specification
- https://schema.org/Organization
Breadth of structured data
Recommended2 pointsEvery websiteschema-type-breadth
- Why an agent needs it
- WebSite and WebPage describe the wrapper. A third type is the first one that describes the business itself: what it sells, what its pages answer, who it is for.
- How it is measured
- Distinct schema.org types were counted across the homepage's JSON-LD and every crawled page. It passes when the site publishes at least three.
- Not applicable when
- The site published no schema.org type.
- When it fails
- Describe more of the site in JSON-LD than the page wrapper: the business itself, the things it sells, and the questions its pages answer.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://schema.org/docs/full.html
Links to external entities
Recommended2 pointsEvery websiteentity-linking
- Why an agent needs it
- A sameAs link ties the business to a profile an agent may already hold a record for, so it can tell this business from another with the same name.
- How it is measured
- The Organization nodes in the homepage's JSON-LD were read for sameAs links. It passes when at least one names an absolute URL on a host other than the site's own.
- Not applicable when
- The homepage published no Organization node.
- When it fails
- Add a sameAs array to the homepage's Organization block listing the business's own profiles on other sites.
- Specification
- https://schema.org/sameAs
About, contact and privacy pages
Required3 pointsEvery websitetrust-anchor-pages
- Why an agent needs it
- An about page, a contact page and a privacy policy are what a reader checks before trusting a business, and an agent checks the same three. A site missing one of them answers the question of who this is with a gap.
- How it is measured
- The crawled pages were read for an about, a contact and a privacy page. Each earns a point when a page answered and rendered at least 500 characters of text outside its markup. A page that returned an error is recorded as absent; one that answered with almost no text is recorded as thin.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Serve an about, a contact and a privacy page at stable URLs, each rendering its text in the HTML response rather than after JavaScript runs.
Pricing published
Required3 pointsEvery websitepricing-discoverable
- Why an agent needs it
- An agent comparing options reads the prices it can find and nothing else. A price quoted only on enquiry is absent from the comparison.
- How it is measured
- The crawled pages were read for one named for pricing, and the site's structured data for an Offer or a PriceSpecification. It passes when a pricing page rendered at least 200 characters of text, or when a price type appears in structured data.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Publish prices on a page an agent can read without running JavaScript, or describe them in JSON-LD as an Offer or a PriceSpecification.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://schema.org/Offer
OpenAPI document usable
Required7 pointsSites that advertise an HTTP APIopenapi-document-usable
- Why an agent needs it
- An OpenAPI document is how an agent learns what an API does and how to call it. A document that is served but empty or malformed claims a capability and then delivers none of it.
- How it is measured
- The OpenAPI document served at /openapi.json or /.well-known/openapi.json was parsed. It passes when the document is valid JSON, declares its OpenAPI or Swagger version, and describes at least one path.
- Not applicable when
- The site serves no OpenAPI document at a conventional path.
- When it fails
- Serve a parseable OpenAPI document that declares its version and describes at least one path, or remove the document from its public path.
- Specification
- https://spec.openapis.org/oas/latest.html
API catalogue published
Required2 pointsSites that advertise an HTTP APIapi-catalog
- Why an agent needs it
- An API catalogue lists every API a site publishes from one well-known address. An agent that finds it does not have to discover the APIs one by one.
- How it is measured
- RFC 9727 API catalogue discovery. It passes when a catalogue document is served at /.well-known/api-catalog, or when a response carries an api-catalog link relation naming one.
- Not applicable when
- The site advertises no HTTP API: no OpenAPI document, API catalogue or plugin manifest was served.
- When it fails
- Serve an RFC 9727 catalogue at /.well-known/api-catalog listing each published API, or carry an api-catalog link relation on a response that already names one.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.rfc-editor.org/rfc/rfc9727
MCP server discoverable
Recommended2 pointsSites that advertise an MCP servermcp-server-discoverable
- Why an agent needs it
- An MCP server an agent can find without reading prose is one it can connect to on its own. A descriptor at the well-known path is that address; a documentation page is not.
- How it is measured
- Where the site's MCP server is advertised. It passes when a descriptor is served at /.well-known/mcp.json as parseable JSON, and fails when the only advertisement is prose an agent has to read.
- Not applicable when
- The site advertises no MCP server: no descriptor, document or page named one.
- When it fails
- Serve a parseable MCP descriptor at /.well-known/mcp.json naming the server's endpoint, or withdraw the endpoint from the pages that publish it.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://modelcontextprotocol.io/specification/latest
API documentation linked from the homepage
Required3 pointsSites that advertise an HTTP APIpublic-api-docs
- Why an agent needs it
- An agent that has found a site's API needs its documentation, and the homepage is where it looks first. Documentation that is published but linked from nowhere an agent lands is documentation it does not find.
- How it is measured
- The homepage's links were read for one to the site's own developer documentation: a docs, developers, api or dev subdomain of the site, or a /docs, /developers, /api or /reference path on it. It passes when at least one is linked.
- Not applicable when
- The site advertises no HTTP API: no OpenAPI document, API catalogue or plugin manifest was served.
- When it fails
- Link the API documentation from the homepage, on a docs or developer subdomain or under a /docs path on the site itself.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Can an agent operate what it reads?
Usability, weight 4013 probes, 2 scored
Bot signing keys published
Emerging2 pointsEvery websiteweb-bot-auth-directory
- Why an agent needs it
- An agent that signs its requests can prove which operator sent them, and an origin that publishes its keys can be checked the same way. The directory is where that proof is published.
- How it is measured
- /.well-known/http-message-signatures-directory was requested. It passes when the document is served as parseable JSON carrying a non-empty keys array.
- When it fails
- Serve a JSON Web Key Set at /.well-known/http-message-signatures-directory carrying the public keys the site signs its own bot requests with.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Missing paths return 404
Recommended2 pointsEvery websiteagent-friendly-404
- Why an agent needs it
- A real 404 tells an agent a page does not exist. A site that answers every unknown path with a 200 and a page makes a miss look like a hit, so the agent reads the wrong page as the answer.
- How it is measured
- The convention paths the collector requests were read for how the origin answers a path it does not serve. It fails when two or more of them answered 200 with something other than the JSON document the path is for.
- Not applicable when
- None of the convention paths the collector requests returned a response to read.
- When it fails
- Answer a path the site does not serve with a 404 status, and keep the catch-all route off paths that carry no page.
Content is marked out
Recommended3 pointsEvery websitepage-landmarks
- Why an agent needs it
- A main landmark tells an agent where the page's own content starts. Without one, it reads the header, the navigation and the cookie banner before it reaches the article.
- How it is measured
- Every page read was checked for a main landmark, as an element or an explicit role. It passes when at least half of them carry one.
- Not applicable when
- No page's markup could be read for its landmarks.
- When it fails
- Wrap each page's own content in a main element, leaving the header, navigation and footer outside it.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.w3.org/TR/wai-aria-1.2/#landmark_roles
Controls are real controls
Recommended3 pointsEvery websiteinteractive-controls-native
- Why an agent needs it
- An agent driving a browser finds controls through the accessibility tree. A div with a click handler is not in it: nothing says the element can be activated, so the agent cannot use it.
- How it is measured
- Elements carrying an explicit button role or an inline click handler were counted against real buttons and links in the served HTML. It passes when at least half the pages keep the improvised share low.
- Not applicable when
- No page carried a clickable element to measure.
- When it fails
- Use a button element for anything that acts, and an anchor with an href for anything that navigates, rather than attaching a click handler to a div or span.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.w3.org/TR/html-aria/
Controls say what they do
Recommended2 pointsEvery websitecontrols-have-names
- Why an agent needs it
- An agent chooses a control by its name. A button with no text and no label announces itself as a button and nothing else, so the agent cannot tell what it does.
- How it is measured
- Buttons and links were checked for text, an aria-label, an aria-labelledby or a title. It passes when at least half the pages name almost all of theirs.
- Not applicable when
- No page carried a button or a link to name.
- When it fails
- Give every button and link readable text, or an aria-label when it carries only an icon.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.w3.org/TR/accname-1.2/
Form fields are labelled
Recommended2 pointsEvery websiteform-fields-labelled
- Why an agent needs it
- A form field's label is how an agent knows what the field asks for. An unlabelled field is a box it fills in without knowing what it is answering.
- How it is measured
- Form fields were checked for a label, an aria-label, an aria-labelledby or a title. It passes when at least half the pages with a form label every field.
- Not applicable when
- No page carried a form field.
- When it fails
- Give every form field a label element bound to it, or an aria-label when no visible label is wanted.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Hidden text describes, not instructs
Recommended2 pointsEvery websiteassistive-text-injection-safe
- Why an agent needs it
- aria-label and alt text are read by an agent and by almost nobody else looking at the page. Text there that reads as an instruction is addressed to the agent rather than describing the element, and it is the place prompt injection hides.
- How it is measured
- Text that only an assistive recipient reads, from aria-label and alt attributes, was scanned for phrasing shaped like an instruction to a model. It passes when none was found.
- Not applicable when
- No page carried text only an assistive client reads.
- When it fails
- Keep aria-label and alt text to a description of the element, and remove any text addressed to a reader other than the person using the page.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
WebMCP tools declared
Required5 pointsEvery websitewebmcp
- Why an agent needs it
- WebMCP lets a page tell an agent driving a browser which actions it offers and how to call them, instead of leaving the agent to guess from buttons. A page that declares its tools can be operated; one that does not has to be worked out.
- How it is measured
- Every readable crawled page's HTML was scanned for a toolname attribute and for a navigator.modelContext reference in inline script. It passes when at least one page carries either.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Declare the actions an agent may take on the page as WebMCP tools: a toolname attribute on each form or control, or a navigator.modelContext.registerTool call describing it.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://webmachinelearning.github.io/webmcp/
A2UI markers present
Emerging2 pointsEvery websitea2ui-support
- Why an agent needs it
- A2UI is how a site hands an agent a piece of its own interface to render inside the agent's answer, rather than a link. The markers are the only trace of it a fetch can see.
- How it is measured
- Every readable crawled page's HTML was scanned for the A2UI media type, a data-a2ui attribute or an a2ui renderer reference. It passes when at least one page carries one.
- Not applicable when
- No crawled page was readable in this Run.
- When it fails
- Serve A2UI surfaces where an agent should render the site's own components: the application/a2ui+json media type or data-a2ui attributes on the elements that carry them.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
JSON error responses
Required4 pointsSites that advertise an HTTP APIjson-error-responses
- Why an agent needs it
- An agent calling an API reads the error body to decide what to do next. An HTML error page gives it a web page where it needs a machine-readable reason.
- How it is measured
- The content type the origin declared on the error responses this Run received. It passes when every typed error response is JSON, and fails when one is an HTML page.
- Not applicable when
- The site advertises no HTTP API: no OpenAPI document, API catalogue or plugin manifest was served.
- When it fails
- Answer a request the origin cannot serve with a JSON body and a JSON content type, rather than a rendered HTML page.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://www.rfc-editor.org/rfc/rfc9457
MCP server card published
Recommended2 pointsSites that advertise an MCP servermcp-server-card
- Why an agent needs it
- A descriptor that names the server and says where to connect lets an agent start with no person in the loop. An empty object at the right path tells it nothing.
- How it is measured
- The MCP descriptor's own contents. It passes when the descriptor names the server and either states its version or names the endpoint to connect to.
- Not applicable when
- The site advertises no MCP server: no descriptor, document or page named one.
- When it fails
- Publish a descriptor that names the MCP server and states either its version or the endpoint an agent connects to.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
- Specification
- https://modelcontextprotocol.io/specification/latest
Protected-resource metadata
Required2 pointsSites that advertise an HTTP APIoauth-protected-resource-metadata
- Why an agent needs it
- An agent that has to authenticate needs to know what the resource is and which authorization server issues tokens for it before its first request. Protected-resource metadata is where a site states both.
- How it is measured
- RFC 9728 protected-resource metadata at /.well-known/oauth-protected-resource. It passes when the document parses, names the resource, and either lists its authorization servers or is served beside an authorization-server document.
- Not applicable when
- The site advertises no authenticated surface: no OpenAPI security scheme names OAuth and no plugin manifest does.
- When it fails
- Serve RFC 9728 metadata at /.well-known/oauth-protected-resource naming the resource identifier and the authorization servers that issue tokens for it.
- Bonus
- Upside only. A pass adds this probe to the figure on both sides; a site that does not serve it is not marked down.
- Specification
- https://www.rfc-editor.org/rfc/rfc9728
NLWeb /ask answers
Emerging1 pointEvery surface the site exposesnlweb-ask
- Why an agent needs it
- NLWeb gives an agent a question-answering endpoint on the site's own data. A site that advertises it and does not answer has told the agent about a door that does not open.
- How it is measured
- When robots.txt carried a Schemamap directive, /ask was requested once with a generic question and streaming off. It passes when the endpoint answers with a JSON content type and a parseable JSON body.
- Not applicable when
- The site does not advertise NLWeb: robots.txt carries no Schemamap directive.
- When it fails
- Answer GET /ask?query=... with a JSON document, as the NLWeb protocol specifies, or remove the Schemamap directive that advertises it.
- Scoring
- Reported in full and not yet scored. It joins the figure at a dated catalogue version, once the lab shows agents or AI answers depend on it.
Can an agent complete a purchase?
Payments, weight 100 probes, 0 scored
No probe scores into this layer yet. It is listed so the weights are stated in full, and it is left out of every figure until one does.
Six statuses, and what each one does to the figure.
| Status | Meaning |
|---|---|
| Passed | The site was read, and answered as the probe requires. The probe earns its full weight. |
| Partly met | The site was read, and answered part of what the probe requires. The probe earns the part it met. |
| Failed | The site was read, and did not answer as the probe requires. The probe earns nothing and stays in the figure. |
| Not evaluated | The probe applied and could not be read. It stays in the figure and earns nothing, so a fault of ours is never rounded away. |
| Pending | The probe applied and its evidence had not landed when the run was read. It counts toward nothing until it does. |
| Not applicable | The probe is outside this site's measurement, for a reason the result states. It counts neither for nor against the figure. |
A probe is added and retired, never renamed.
A probe’s id is written into every stored result, so two runs months apart compare row for row. A probe that stops running keeps its id and its definition, and a past run keeps the result it recorded. These probes no longer run:
AI crawlers reach the page (retired)
Emerging0 pointsEvery websiteai-crawler-reachability
- Why an agent needs it
- Retired. Whether an origin serves a genuine AI crawler cannot be observed from Heading's own address, so nothing on the site depends on this result.
- How it is measured
- Retired. This probe re-requested the homepage as each of four AI crawlers from Heading's own address and read a refusal as a block. An origin that verifies a crawler by its source address refuses that request and serves the genuine crawler, so the observation could not tell the two apart.
- When it fails
- Nothing to change on the site: this result came from a request Heading sent from its own address while carrying a crawler's user agent, and an origin that verifies crawlers by their source address refuses that request correctly. The probe has been retired. Read what robots.txt states in the AI crawler access probe instead.
- Scoring
- Retired. It no longer runs; a past run keeps the result it recorded.
AI crawlers get the whole page (retired)
Emerging0 pointsEvery websiteai-crawler-content-parity
- Why an agent needs it
- Retired. Whether an origin serves a genuine AI crawler cannot be observed from Heading's own address, so nothing on the site depends on this result.
- How it is measured
- Retired. This probe compared the page each AI crawler user agent received with the page Heading received, from the same address. A challenge or a smaller page served to an unverified request is the same correct behaviour a refusal is, so the comparison could not tell a wall from a check.
- When it fails
- Nothing to change on the site: this result came from a request Heading sent from its own address while carrying a crawler's user agent, and an origin that verifies crawlers by their source address refuses that request correctly. The probe has been retired. Read what robots.txt states in the AI crawler access probe instead.
- Scoring
- Retired. It no longer runs; a past run keeps the result it recorded.
Read a grade with its probes attached.
- Two sites are measured over different probes. A site that advertises an API is measured on more than a brochure site. The grade is a share of what applied, and each run states what applied and why the rest did not.
- A probe reports from outside. It sees what a public request from Heading’s own address sees. It cannot see a firewall rule, a geographic restriction or what an origin serves to a verified crawler, and it does not pretend to.
- A pass is a fact, not a promise. Passing a probe means the site answered as the probe requires. It says nothing about how often the site is named or cited in AI answers; that is a separate measurement with its own evidence.
- Separate observation from cause. A grade that moves after a change does not prove the change moved it. Each run is stamped with the catalogue version it was measured against, so a step caused by a catalogue change has a date to point at.