ax-audit 3.6.0 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (201) hide show
  1. package/CHANGELOG.md +138 -0
  2. package/LICENSE +1 -1
  3. package/README.md +58 -36
  4. package/dist/baseline.d.ts +2 -0
  5. package/dist/baseline.d.ts.map +1 -1
  6. package/dist/baseline.js +42 -4
  7. package/dist/baseline.js.map +1 -1
  8. package/dist/check-ids.d.ts +19 -0
  9. package/dist/check-ids.d.ts.map +1 -0
  10. package/dist/check-ids.js +53 -0
  11. package/dist/check-ids.js.map +1 -0
  12. package/dist/checks/agent-access.d.ts +23 -6
  13. package/dist/checks/agent-access.d.ts.map +1 -1
  14. package/dist/checks/agent-access.js +200 -54
  15. package/dist/checks/agent-access.js.map +1 -1
  16. package/dist/checks/agent-card.d.ts +37 -0
  17. package/dist/checks/agent-card.d.ts.map +1 -0
  18. package/dist/checks/agent-card.js +352 -0
  19. package/dist/checks/agent-card.js.map +1 -0
  20. package/dist/checks/agent-operability.d.ts +66 -0
  21. package/dist/checks/agent-operability.d.ts.map +1 -0
  22. package/dist/checks/agent-operability.js +383 -0
  23. package/dist/checks/agent-operability.js.map +1 -0
  24. package/dist/checks/agent-skills.d.ts +24 -0
  25. package/dist/checks/agent-skills.d.ts.map +1 -0
  26. package/dist/checks/agent-skills.js +316 -0
  27. package/dist/checks/agent-skills.js.map +1 -0
  28. package/dist/checks/ai-catalog.d.ts +28 -0
  29. package/dist/checks/ai-catalog.d.ts.map +1 -0
  30. package/dist/checks/ai-catalog.js +254 -0
  31. package/dist/checks/ai-catalog.js.map +1 -0
  32. package/dist/checks/ai-directives.d.ts +57 -0
  33. package/dist/checks/ai-directives.d.ts.map +1 -0
  34. package/dist/checks/ai-directives.js +263 -0
  35. package/dist/checks/ai-directives.js.map +1 -0
  36. package/dist/checks/api-discovery.d.ts +26 -0
  37. package/dist/checks/api-discovery.d.ts.map +1 -0
  38. package/dist/checks/api-discovery.js +432 -0
  39. package/dist/checks/api-discovery.js.map +1 -0
  40. package/dist/checks/auth-discovery.d.ts +28 -0
  41. package/dist/checks/auth-discovery.d.ts.map +1 -0
  42. package/dist/checks/auth-discovery.js +213 -0
  43. package/dist/checks/auth-discovery.js.map +1 -0
  44. package/dist/checks/commerce-discovery.d.ts +40 -0
  45. package/dist/checks/commerce-discovery.d.ts.map +1 -0
  46. package/dist/checks/commerce-discovery.js +295 -0
  47. package/dist/checks/commerce-discovery.js.map +1 -0
  48. package/dist/checks/content-negotiation.d.ts.map +1 -1
  49. package/dist/checks/content-negotiation.js +135 -20
  50. package/dist/checks/content-negotiation.js.map +1 -1
  51. package/dist/checks/crawl-efficiency.d.ts +13 -1
  52. package/dist/checks/crawl-efficiency.d.ts.map +1 -1
  53. package/dist/checks/crawl-efficiency.js +65 -1
  54. package/dist/checks/crawl-efficiency.js.map +1 -1
  55. package/dist/checks/frontmatter.d.ts +34 -0
  56. package/dist/checks/frontmatter.d.ts.map +1 -0
  57. package/dist/checks/frontmatter.js +100 -0
  58. package/dist/checks/frontmatter.js.map +1 -0
  59. package/dist/checks/html-rendering.d.ts.map +1 -1
  60. package/dist/checks/html-rendering.js +0 -1
  61. package/dist/checks/html-rendering.js.map +1 -1
  62. package/dist/checks/html-utils.d.ts +10 -0
  63. package/dist/checks/html-utils.d.ts.map +1 -1
  64. package/dist/checks/html-utils.js +19 -0
  65. package/dist/checks/html-utils.js.map +1 -1
  66. package/dist/checks/http-headers.d.ts.map +1 -1
  67. package/dist/checks/http-headers.js +82 -10
  68. package/dist/checks/http-headers.js.map +1 -1
  69. package/dist/checks/http-hygiene.d.ts +26 -0
  70. package/dist/checks/http-hygiene.d.ts.map +1 -0
  71. package/dist/checks/http-hygiene.js +257 -0
  72. package/dist/checks/http-hygiene.js.map +1 -0
  73. package/dist/checks/index.d.ts.map +1 -1
  74. package/dist/checks/index.js +24 -8
  75. package/dist/checks/index.js.map +1 -1
  76. package/dist/checks/llms-txt.d.ts +15 -0
  77. package/dist/checks/llms-txt.d.ts.map +1 -1
  78. package/dist/checks/llms-txt.js +162 -2
  79. package/dist/checks/llms-txt.js.map +1 -1
  80. package/dist/checks/mcp-discovery.d.ts +30 -0
  81. package/dist/checks/mcp-discovery.d.ts.map +1 -0
  82. package/dist/checks/mcp-discovery.js +523 -0
  83. package/dist/checks/mcp-discovery.js.map +1 -0
  84. package/dist/checks/meta-tags.d.ts.map +1 -1
  85. package/dist/checks/meta-tags.js +6 -5
  86. package/dist/checks/meta-tags.js.map +1 -1
  87. package/dist/checks/robots-parser.d.ts +110 -0
  88. package/dist/checks/robots-parser.d.ts.map +1 -0
  89. package/dist/checks/robots-parser.js +277 -0
  90. package/dist/checks/robots-parser.js.map +1 -0
  91. package/dist/checks/robots-txt.d.ts +2 -20
  92. package/dist/checks/robots-txt.d.ts.map +1 -1
  93. package/dist/checks/robots-txt.js +219 -120
  94. package/dist/checks/robots-txt.js.map +1 -1
  95. package/dist/checks/rsl.d.ts +0 -2
  96. package/dist/checks/rsl.d.ts.map +1 -1
  97. package/dist/checks/rsl.js +1 -11
  98. package/dist/checks/rsl.js.map +1 -1
  99. package/dist/checks/security-txt.d.ts.map +1 -1
  100. package/dist/checks/security-txt.js +0 -1
  101. package/dist/checks/security-txt.js.map +1 -1
  102. package/dist/checks/seo-basics.d.ts.map +1 -1
  103. package/dist/checks/seo-basics.js +0 -1
  104. package/dist/checks/seo-basics.js.map +1 -1
  105. package/dist/checks/sitemap.d.ts.map +1 -1
  106. package/dist/checks/sitemap.js +0 -1
  107. package/dist/checks/sitemap.js.map +1 -1
  108. package/dist/checks/structured-data.d.ts.map +1 -1
  109. package/dist/checks/structured-data.js +215 -4
  110. package/dist/checks/structured-data.js.map +1 -1
  111. package/dist/checks/structured-fields.d.ts +46 -0
  112. package/dist/checks/structured-fields.d.ts.map +1 -0
  113. package/dist/checks/structured-fields.js +112 -0
  114. package/dist/checks/structured-fields.js.map +1 -0
  115. package/dist/checks/surface.d.ts +59 -0
  116. package/dist/checks/surface.d.ts.map +1 -0
  117. package/dist/checks/surface.js +106 -0
  118. package/dist/checks/surface.js.map +1 -0
  119. package/dist/checks/tls-https.d.ts.map +1 -1
  120. package/dist/checks/tls-https.js +0 -1
  121. package/dist/checks/tls-https.js.map +1 -1
  122. package/dist/checks/usage-policy.d.ts +53 -0
  123. package/dist/checks/usage-policy.d.ts.map +1 -0
  124. package/dist/checks/usage-policy.js +339 -0
  125. package/dist/checks/usage-policy.js.map +1 -0
  126. package/dist/checks/utils.d.ts +25 -1
  127. package/dist/checks/utils.d.ts.map +1 -1
  128. package/dist/checks/utils.js +33 -1
  129. package/dist/checks/utils.js.map +1 -1
  130. package/dist/checks/waf.d.ts +75 -0
  131. package/dist/checks/waf.d.ts.map +1 -0
  132. package/dist/checks/waf.js +203 -0
  133. package/dist/checks/waf.js.map +1 -0
  134. package/dist/checks/webmcp.d.ts +55 -0
  135. package/dist/checks/webmcp.d.ts.map +1 -0
  136. package/dist/checks/webmcp.js +209 -0
  137. package/dist/checks/webmcp.js.map +1 -0
  138. package/dist/checks/well-known.d.ts +38 -0
  139. package/dist/checks/well-known.d.ts.map +1 -0
  140. package/dist/checks/well-known.js +202 -0
  141. package/dist/checks/well-known.js.map +1 -0
  142. package/dist/cli.d.ts +14 -0
  143. package/dist/cli.d.ts.map +1 -1
  144. package/dist/cli.js +124 -4
  145. package/dist/cli.js.map +1 -1
  146. package/dist/constants.d.ts +195 -14
  147. package/dist/constants.d.ts.map +1 -1
  148. package/dist/constants.js +598 -70
  149. package/dist/constants.js.map +1 -1
  150. package/dist/fetcher.d.ts.map +1 -1
  151. package/dist/fetcher.js +44 -14
  152. package/dist/fetcher.js.map +1 -1
  153. package/dist/guide-urls.js +1 -1
  154. package/dist/guide-urls.js.map +1 -1
  155. package/dist/index.d.ts +2 -1
  156. package/dist/index.d.ts.map +1 -1
  157. package/dist/index.js +1 -0
  158. package/dist/index.js.map +1 -1
  159. package/dist/orchestrator.d.ts.map +1 -1
  160. package/dist/orchestrator.js +5 -1
  161. package/dist/orchestrator.js.map +1 -1
  162. package/dist/reporter/html.d.ts +9 -0
  163. package/dist/reporter/html.d.ts.map +1 -1
  164. package/dist/reporter/html.js +49 -11
  165. package/dist/reporter/html.js.map +1 -1
  166. package/dist/reporter/markdown.d.ts.map +1 -1
  167. package/dist/reporter/markdown.js +36 -6
  168. package/dist/reporter/markdown.js.map +1 -1
  169. package/dist/reporter/terminal.d.ts.map +1 -1
  170. package/dist/reporter/terminal.js +36 -1
  171. package/dist/reporter/terminal.js.map +1 -1
  172. package/dist/scorer.d.ts +10 -0
  173. package/dist/scorer.d.ts.map +1 -1
  174. package/dist/scorer.js +22 -5
  175. package/dist/scorer.js.map +1 -1
  176. package/dist/types.d.ts +90 -3
  177. package/dist/types.d.ts.map +1 -1
  178. package/docs/architecture.md +27 -11
  179. package/docs/checks.md +285 -52
  180. package/docs/cli.md +36 -0
  181. package/docs/concepts.md +27 -13
  182. package/docs/faq.md +18 -6
  183. package/docs/getting-started.md +20 -13
  184. package/docs/roadmap.md +367 -0
  185. package/package.json +13 -5
  186. package/dist/checks/agent-json.d.ts +0 -14
  187. package/dist/checks/agent-json.d.ts.map +0 -1
  188. package/dist/checks/agent-json.js +0 -167
  189. package/dist/checks/agent-json.js.map +0 -1
  190. package/dist/checks/mcp.d.ts +0 -4
  191. package/dist/checks/mcp.d.ts.map +0 -1
  192. package/dist/checks/mcp.js +0 -162
  193. package/dist/checks/mcp.js.map +0 -1
  194. package/dist/checks/openapi.d.ts +0 -4
  195. package/dist/checks/openapi.d.ts.map +0 -1
  196. package/dist/checks/openapi.js +0 -121
  197. package/dist/checks/openapi.js.map +0 -1
  198. package/dist/checks/well-known-ai.d.ts +0 -17
  199. package/dist/checks/well-known-ai.d.ts.map +0 -1
  200. package/dist/checks/well-known-ai.js +0 -123
  201. package/dist/checks/well-known-ai.js.map +0 -1
package/docs/checks.md CHANGED
@@ -1,8 +1,29 @@
1
1
  # Checks Reference
2
2
 
3
- ax-audit runs 18 checks. Fourteen are **weighted** (summing to 100% of the overall score); four are **informational** in 3.x — they run and report findings but carry weight 0 until v4.0, because score-affecting changes are treated as breaking (see [CHANGELOG 3.0.0](../CHANGELOG.md)).
3
+ ax-audit runs 26 checks. Twenty-three are **weighted**, summing to 100. Three rest on draft specifications and stay at weight 0: scoring a site against a specification that may be renamed next quarter would make the number less trustworthy, not more.
4
4
 
5
- This page documents the **exact scoring** of every check: each deduction, bonus, and formula, extracted from the source. Every finding links to a step-by-step remediation guide at `lucioduran.com/projects/ax-audit/guides/<check-id>`.
5
+ Checks are grouped into five areas, and reports are ordered by them: **content** (is there substance an agent can read?), **discovery** (can an agent find your machine-readable files?), **access** (can it actually retrieve them?), **policy** (what usage rights do you declare?), and **protocols** (what can an agent call?).
6
+
7
+ Some checks are **conditional**. A blog has no API to describe, no MCP server to advertise and nothing to authorize, so those checks report **n/a** and are excluded from the score rather than counted as failures. Everything counted against a site is something the site could have done. `--profile api|mcp|agent|docs|commerce|all` forces them applicable, for auditing against what a site intends to become.
8
+
9
+ ## Weights
10
+
11
+ | Area | Total | Checks |
12
+ | --- | --- | --- |
13
+ | Content | 33 | html-rendering 11 · agent-operability 7 · structured-data 6 · seo-basics 5 · content-negotiation 4 |
14
+ | Access | 24 | agent-access 9 · ai-directives 6 · http-hygiene 4 · tls-https 3 · crawl-efficiency 2 |
15
+ | Discovery | 21 | robots-txt 9 · llms-txt 5 · http-headers 4 · sitemap 2 · meta-tags 1 |
16
+ | Protocols | 13 | api-discovery 4 · agent-card 3 · mcp-discovery 3 · agent-skills 2 · auth-discovery 1 |
17
+ | Policy | 9 | usage-policy 4 · security-txt 3 · rsl 2 |
18
+ | Draft specifications | 0 | ai-catalog · webmcp · commerce-discovery |
19
+
20
+ Content leads because the failure that breaks the most agents is a page with nothing in its HTML: most crawlers do not run JavaScript, and a site whose content appears only after hydration is invisible no matter how many discovery files it publishes. Access is second because a firewall rule or a `nosnippet` directive silently undoes everything else, and those are the failures operators are least likely to know about. llms.txt sits at 5 rather than the 11 it carried in 3.x, because most published files are never fetched by an AI search crawler.
21
+
22
+ Weights live in `CHECK_WEIGHTS` in `src/constants.ts`, and only there. Checks used to declare their own alongside it, and the two drifted.
23
+
24
+ Every probed path is labelled by standing — **IANA-registered**, **vendor convention**, **draft**, or **legacy** — because the agent web mixes registered URIs with drafts that get renamed. A missing draft file is not the same kind of finding as a missing registered one, and reports say which is which.
25
+
26
+ This page documents the **exact scoring** of every check: each deduction, bonus, and formula, extracted from the source. Every finding links to a step-by-step remediation guide at `axrush.com/guides/<check-id>`.
6
27
 
7
28
  **Reading the tables:** each check starts at 100 unless noted. Deductions stack additively; `buildResult` clamps the final score to [0, 100]. "Hard fail" rows short-circuit the check.
8
29
 
@@ -24,10 +45,19 @@ This page documents the **exact scoring** of every check: each deduction, bonus,
24
45
  | No Markdown links | −10 |
25
46
  | Content under 100 characters | −10 |
26
47
  | `/llms-full.txt` also available | **+10** (capped at 100) |
48
+ | Broken links among a sample of 15 | informational, 0 in 3.x |
49
+ | Redirecting or duplicated links | informational, 0 |
50
+ | No `rel="describedby"` pointing at the file (llms.txt v2) | informational, 0 |
51
+ | No per-page Markdown mirror (`/index.md`, `/index.html.md`) | informational, 0 |
52
+ | File over 50 KB | informational, 0 |
53
+
54
+ Every report states plainly that Google says Search ignores llms.txt and that most published files are never fetched by an AI search crawler, while Claude Code, Cursor and OpenCode do read it. It is a developer-tooling signal, not a search-visibility one.
27
55
 
28
56
  ### `robots-txt` — 11%
29
57
 
30
- AI-crawler configuration. Core crawlers: GPTBot, ClaudeBot, ChatGPT-User, Claude-SearchBot, Google-Extended, PerplexityBot, OAI-SearchBot, CCBot.
58
+ AI-crawler configuration. Scoring runs against the frozen 3.x core set (GPTBot, ClaudeBot, ChatGPT-User, Claude-SearchBot, Google-Extended, PerplexityBot, OAI-SearchBot, CCBot); the wider September-2026 core set adds Meta-ExternalAgent, Applebot-Extended, Amazonbot and Bytespider, reported but not scored until 4.0.
59
+
60
+ Findings are tiered by what a client does with a page, because that determines the cost of blocking it. Blocking a **training** crawler is a policy choice and is reported as such. Blocking a **search** crawler removes the site from that assistant's answers. Blocking a **user-triggered fetcher** often does nothing, because most vendors document that robots.txt may not apply to them.
31
61
 
32
62
  | Condition | Points |
33
63
  | --- | --- |
@@ -38,7 +68,10 @@ AI-crawler configuration. Core crawlers: GPTBot, ClaudeBot, ChatGPT-User, Claude
38
68
  | Known AI crawler(s) explicitly blocked (`Disallow: /`) | −3 per crawler |
39
69
  | No `Sitemap:` directive | −5 |
40
70
  | Partial path restrictions on AI crawlers | warn only, 0 |
41
- | [Content Signals](https://contentsignals.org) findings (declared / malformed / unknown / missing) | informational, 0 in 3.x |
71
+ | Blocking a crawler token added in 3.7 (`meta-webindexer`, `Amzn-SearchBot`, …) | informational, 0 in 3.x |
72
+ | Rules targeting a retired or fictional token (`GeminiBot`, `Claude-Web`, `NeevaBot`, …) | informational, 0 |
73
+ | [Content Signals](https://contentsignals.org) findings, including the `use=immediate\|reference\|full` field | informational, 0 in 3.x |
74
+ | [IETF AIPREF](https://datatracker.ietf.org/wg/aipref/documents/) `Content-Usage:` findings, including vocabulary mix-ups | informational, 0 in 3.x |
42
75
 
43
76
  ### `html-rendering` — 9%
44
77
 
@@ -72,54 +105,84 @@ JSON-LD on the homepage. Key entity types: Person, Organization, WebSite, WebPag
72
105
  | Only one key entity type | −10 |
73
106
  | No `@graph` array | −5 |
74
107
  | No `BreadcrumbList` | −5 |
108
+ | No `author`, no `sameAs`, or an author with no `publisher` | informational, 0 in 3.x |
109
+ | No `dateModified` / `datePublished`, a future date, or content over 730 days old | informational, 0 |
110
+ | `headline` or `name` not present in the visible text | informational, 0 |
111
+
112
+ The visible-text comparison is Google's one explicit requirement for structured data and AI features. It reads static HTML, so text rendered by script reads as missing here too, and the finding says so.
75
113
 
76
114
  ### `http-headers` — 9%
77
115
 
78
- Security headers, AI discovery `Link` headers (RFC 5988-parsed), CORS on `.well-known`.
116
+ Security headers, AI discovery `Link` headers (RFC 5988-parsed), CORS on `.well-known`. Either Agent Card path satisfies the discovery-link requirement.
79
117
 
80
118
  | Condition | Points |
81
119
  | --- | --- |
82
120
  | No headers retrievable | **hard fail → 0** |
83
121
  | Missing critical security header (HSTS, X-Content-Type-Options) | −10 each |
84
122
  | Only 1–3 of the 7 tracked security headers present | −5 |
85
- | `Link` header missing both llms.txt and agent.json references | −15 |
123
+ | `Link` header missing both llms.txt and the Agent Card | −15 |
86
124
  | `Link` header missing one of the two | −5 |
87
- | No CORS on `/.well-known/agent.json` | −10 |
125
+ | No CORS on the Agent Card | −10 |
126
+ | Additional discovery relations (`describedby`, `api-catalog`, `service-desc`, `service-doc`, `ai-catalog`, `c2pa-manifest`, `license`, markdown `alternate`) and the `X-Llms-Txt` header | informational, 0 in 3.x |
88
127
 
89
- ### `agent-json` — 7%
128
+ ### `agent-card` — 7%
90
129
 
91
- `/.well-known/agent.json` [A2A Agent Card](https://a2a-protocol.org). Required fields: `name`, `description`, `url`, `skills`.
130
+ The [A2A Agent Card](https://a2a-protocol.org), probed at `/.well-known/agent-card.json` (IANA-registered since A2A v0.3.0, 2025-07-30) and then at the pre-0.3 path `/.well-known/agent.json`. *Former id: `agent-json`, still accepted in `--checks` and in saved baselines.*
131
+
132
+ Two spec generations are in the wild, and the check detects which one a card follows from its own structure rather than from a version field:
133
+
134
+ - **A2A 1.0** (2026-03-12) declares every endpoint inside `supportedInterfaces[]`. Required: `name`, `description`, `version`, `capabilities`, `supportedInterfaces`, `defaultInputModes`, `defaultOutputModes`, `skills`.
135
+ - **A2A 0.3** declares a top-level `url` and `protocolVersion`. Required: those two plus `name`, `description`, `version`, `capabilities`, `defaultInputModes`, `defaultOutputModes`, `skills`.
92
136
 
93
137
  | Condition | Points |
94
138
  | --- | --- |
95
- | Not found | **hard fail → 0** |
139
+ | Not found at either path | **hard fail → 0** |
96
140
  | Invalid JSON | **→ 10** |
97
- | Wrong Content-Type (expected `application/json`) | −5 |
98
- | Missing required field | −15 per field |
99
- | `url` on a different origin | −5 |
100
- | `url` not an absolute URL | −5 |
141
+ | Served only from the pre-0.3 `agent.json` path | warn only, 0 |
142
+ | Wrong Content-Type (expected `application/json` or `application/a2a+json`) | −5 |
143
+ | Card shape matches neither generation | −30 |
144
+ | Missing required field for the detected generation | −15 per field |
145
+ | `supportedInterfaces[]` empty (1.0) | −15 |
146
+ | Interface missing `url`, `protocolBinding` or `protocolVersion` (1.0) | −10 |
147
+ | Unrecognised `protocolBinding` (not JSONRPC / GRPC / HTTP+JSON) | −5 |
148
+ | Interface or `url` on a different origin | −5 |
149
+ | `url` not an absolute URL (0.3) | −5 |
101
150
  | `skills` empty | −10 |
102
151
  | `skills` entries missing `id` or `description` | −5 |
103
- | No `protocolVersion` | −5 |
104
- | No optional fields (`capabilities`, `authentication`, `documentationUrl`) | −5 |
152
+ | Uses `authentication`, removed from the spec in 0.2.x | −5 |
153
+ | No optional descriptive fields (`provider`, `documentationUrl`, `iconUrl`) | −5 |
154
+
155
+ ### `mcp-discovery` — 7%
156
+
157
+ How an agent finds this site's [Model Context Protocol](https://modelcontextprotocol.io) server. *Former id: `mcp`.*
105
158
 
106
- ### `mcp` — 7%
159
+ `/.well-known/mcp.json` was never part of the MCP specification. What emerged instead is the **server card**, which deliberately carries no `tools[]` — tool lists come from a live `tools/list` call, and a static copy drifts the day it is written. Discovery is probed in this order:
107
160
 
108
- `/.well-known/mcp.json` [Model Context Protocol](https://modelcontextprotocol.io) server configuration.
161
+ 1. `/.well-known/ai-catalog.json` entries of type `application/mcp-server-card+json` *(draft)*
162
+ 2. `/.well-known/mcp/server-card.json`, `/.well-known/mcp/server-cards.json` *(vendor convention: Cloudflare, Mintlify)*
163
+ 3. `<endpoint>/server-card` for `/mcp`, `/api/mcp`, `/sse` *(the MCP extension's own recommendation)*
164
+ 4. `/.well-known/mcp.json` *(legacy)*
165
+
166
+ An HTML response counts as absence, not a malformed card: SPA catch-alls answer every unknown path with the index shell.
167
+
168
+ **When a server card is found:**
109
169
 
110
170
  | Condition | Points |
111
171
  | --- | --- |
112
- | Not found | **hard fail → 0** |
172
+ | Wrong Content-Type (expected `application/json` or `application/mcp-server-card+json`) | −5 |
113
173
  | Invalid JSON | **→ 10** |
114
- | Wrong Content-Type | −5 |
115
- | Missing `name` | −10 |
116
- | Missing `description` | −5 |
117
- | No `tools` array, or empty | −15 |
118
- | No tool has a description | −10 |
119
- | Some tools missing descriptions | −5 |
120
- | No `resources` | −5 |
121
- | No protocol version | −5 |
174
+ | Missing `$schema`, `name`, `version` or `description` | −15 each |
175
+ | `name` not in reverse-DNS form | −5 |
176
+ | No `remotes[]` | −15 |
177
+ | Remote with an unrecognised transport (not `streamable-http` / `sse`) | −5 |
178
+ | Remote missing a `url` | −10 |
179
+ | No `supportedProtocolVersions` | −10 |
180
+ | Only pre-2025-06 protocol revisions | −10 |
181
+ | Unrecognised protocol version | −5 |
122
182
  | No CORS headers | −10 |
183
+ | Declares `tools[]`, which the schema omits by design | warn only, 0 |
184
+
185
+ **When only `/.well-known/mcp.json` is found**, the pre-3.7 rules are applied unchanged so the score is exactly what 3.6 produced (missing `name` −10, missing `description` −5, no `tools` −15, no tool descriptions −10 / −5, no `resources` −5, no version −5, no CORS −10, wrong Content-Type −5). The path itself is reported, not penalised.
123
186
 
124
187
  ### `seo-basics` — 7%
125
188
 
@@ -161,7 +224,7 @@ AI meta tags (`ai:summary`, `ai:content_type`, `ai:author`, `ai:api`, `ai:agent_
161
224
  | 0 AI meta tags | −18 |
162
225
  | Only 1–2 AI meta tags | −12 |
163
226
  | No `rel="alternate"` → llms.txt | −12 |
164
- | No `rel="alternate"` → agent.json | −8 |
227
+ | No `rel="alternate"` → the Agent Card | −8 |
165
228
  | No `rel="me"` identity links | −8 |
166
229
  | No Open Graph tags at all | −12 |
167
230
  | OG required incomplete (`og:title`, `og:description`, `og:url`, `og:type`) | −8 |
@@ -170,21 +233,33 @@ AI meta tags (`ai:summary`, `ai:content_type`, `ai:author`, `ai:api`, `ai:agent_
170
233
  | Twitter required incomplete (`twitter:card`, `twitter:title`, `twitter:description`) | −5 |
171
234
  | Twitter recommended incomplete (`twitter:image`) | −2 |
172
235
 
173
- ### `openapi` — 6%
236
+ ### `api-discovery` — 6%
237
+
238
+ Whether an agent can find, and read, a machine-readable API description. *Former id: `openapi`.*
174
239
 
175
- `/.well-known/openapi.json`.
240
+ `/.well-known/openapi.json` is a folk convention — unregistered, and not prescribed by the OpenAPI specification, which recommends the file name `openapi.json` without a location. Discovery is probed in order of authority:
241
+
242
+ 1. `/.well-known/api-catalog` *(RFC 9727, IANA-registered)* → its `service-desc` links
243
+ 2. `Link: rel="service-desc"` on the homepage *(RFC 8631)*
244
+ 3. `<link rel="service-desc">` in the HTML head
245
+ 4. Conventional paths: `/.well-known/openapi.json`, `/openapi.json`, `/openapi.yaml`, `/.well-known/openapi.yaml`, `/api/openapi.json`, `/v1/openapi.json`, `/swagger.json`, `/api-docs`, `/asyncapi.json`, `/arazzo.json`
176
246
 
177
247
  | Condition | Points |
178
248
  | --- | --- |
179
- | Not found | **hard fail → 0** |
249
+ | No description found by any mechanism | **hard fail → 0** |
180
250
  | Invalid JSON | **→ 10** |
181
- | Wrong Content-Type | −5 |
251
+ | Wrong Content-Type on a JSON document | −5 |
182
252
  | No `openapi`/`swagger` version field | −20 |
183
253
  | Swagger 2.x instead of OpenAPI 3.x | −10 |
184
254
  | Missing `info.title` | −10 |
185
255
  | Missing `info.description` | −5 |
186
256
  | No `paths` documented | −15 |
187
257
  | No `servers` | −5 |
258
+ | Found only by guessing a path (nothing links to it) | warn only, 0 |
259
+ | `operationId` coverage below 100% | informational, 0 in 3.x |
260
+ | API catalog present but empty, or entries missing `anchor` / `service-doc` | warn only, 0 |
261
+
262
+ YAML descriptions are recognised and reported, but only surface-validated: ax-audit ships no YAML parser, and the finding says so rather than pretending otherwise.
188
263
 
189
264
  ### `tls-https` — 5%
190
265
 
@@ -220,21 +295,13 @@ Located via robots.txt `Sitemap:` or `/sitemap.xml`. Limits: 50,000 URLs / 50 MB
220
295
  | `<lastmod>` coverage < 50% | −5 |
221
296
  | Newest `<lastmod>` older than 365 days | −5 |
222
297
 
223
- ### `well-known-ai` — 3%
224
-
225
- Emerging AI discovery files. **Purely proportional** — no deductions:
226
-
227
- ```
228
- score = round(present / 5 × 100)
229
- ```
230
-
231
- over `/.well-known/ai.txt` (Spawning), `/.well-known/genai.txt`, `/ai-plugin.json`, `/agents.json`, `/.well-known/nlweb.json`. Files with invalid content produce warnings without counting as present.
232
-
233
298
  ---
234
299
 
235
- ## Informational checks (weight 0 in 3.x)
300
+ ## Checks on draft specifications (weight 0)
301
+
302
+ `ai-catalog`, `webmcp` and `commerce-discovery` run on every audit and report full findings, but never affect the score. Each rests on a specification that is still a draft and may be renamed; scoring a site against one would make the number less trustworthy, not more.
236
303
 
237
- These run on every audit and report full findings, but do not affect the overall score or baselines. They gain weight in v4.0.
304
+ `well-known-ai` was removed in 4.0. Re-verification found three of its five scored files had no consumer: `/.well-known/nlweb.json` appears in no NLWeb release, `genai.txt` has no specification, and `/ai-plugin.json` described a product shut down in 2024. Its live probes moved into the checks that own them — TDMRep into `usage-policy`, the Web Bot Auth key directory and `AGENTS.md` into the reporting they belong to.
238
305
 
239
306
  ### `content-negotiation` — Markdown for Agents
240
307
 
@@ -250,6 +317,11 @@ Probes the homepage with `Accept: text/markdown` — the pattern served by Cloud
250
317
  | Body is a relabeled HTML document | −25 |
251
318
  | `Vary` does not include `Accept` | −15 |
252
319
  | Markdown not smaller than HTML | warn only, 0 |
320
+ | Origin-reported token counts (`x-markdown-tokens` / `x-original-tokens`) | informational, 0 |
321
+ | No frontmatter, or frontmatter with no title / canonical URL / date | informational, 0 |
322
+ | User-agent negotiation or a `.md` suffix URL, when Accept negotiation fails | informational, 0 |
323
+
324
+ The probe sends the Accept header a real agent sends (`text/markdown, text/html;q=0.9, */*;q=0.1`). A bare `text/markdown` would pass against an implementation that fails every real request.
253
325
 
254
326
  ### `rsl` — Really Simple Licensing
255
327
 
@@ -270,23 +342,27 @@ Probes the homepage with `Accept: text/markdown` — the pattern served by Cloud
270
342
  | Tokens outside the RSL 1.0 vocabulary (incl. pre-1.0 draft tokens) | −5 |
271
343
  | Invalid `payment` type | −5 |
272
344
 
273
- ### `agent-access` — Cloaking detection
345
+ ### `agent-access` — blocking and cloaking detection
274
346
 
275
- Probes the homepage with realistic UAs for the 8 core AI crawlers and compares status + visible text against the default-UA baseline. **Credit-ratio formula:**
347
+ Probes the homepage with realistic user agents for the 10 core AI crawlers that actually issue requests (`Google-Extended` and `Applebot-Extended` are robots.txt control tokens, so probing with them tests nothing) and compares each response against the default-UA baseline. **Credit-ratio formula:**
276
348
 
277
349
  ```
278
- score = round(credit / 8 × 100)
350
+ score = round(credit / 10 × 100)
279
351
  ```
280
352
 
353
+ Responses are classified by *how* a request was turned away, because the remedies differ completely:
354
+
281
355
  | Outcome per crawler | Credit |
282
356
  | --- | --- |
283
- | Equivalent response | 1 |
284
- | Blocked, consistent with robots.txt `Disallow` (explicit or wildcard) | 1 |
285
- | 200 but < 50% of baseline visible text (baseline ≥ 200 chars) | 0.5 |
286
- | Blocked while robots.txt allows (or doesn't restrict) it | 0 |
357
+ | Same page as a regular client | 1 |
358
+ | Refused, consistent with an explicit robots.txt `Disallow` | 1 |
359
+ | Priced access (`402` + `crawler-price`) or an RSL licence challenge | 1 |
360
+ | JavaScript challenge (`cf-mitigated: challenge`, `x-vercel-mitigated`, AWS WAF's `202`), Web Bot Auth demand, rate limit, or a refusal from a bot-verifying CDN | 0.75, **inconclusive** |
361
+ | Different page than the baseline: less text, or a changed title / h1 / JSON-LD block count | 0.5 |
362
+ | Refused by a plain origin while robots.txt permits it | 0 |
287
363
  | Baseline request itself fails | **hard fail → 0** |
288
364
 
289
- Caveat: WAFs using Web Bot Auth / IP verification may pass the real crawler while rejecting this unverified probe — confirm against WAF logs before changing rules.
365
+ The probe is unsigned and comes from the auditor's own network, so an edge that verifies crawlers by IP range or Web Bot Auth signature will reject it while admitting the genuine crawler. Those outcomes are reported as inconclusive with the exact header observed, never as "blocks AI crawlers". Confirm against WAF logs before changing a rule.
290
366
 
291
367
  ### `crawl-efficiency`
292
368
 
@@ -299,6 +375,161 @@ Caveat: WAFs using Web Bot Auth / IP verification may pass the real crawler whil
299
375
  | Validator present but conditional request not answered with `304` | −15 |
300
376
  | Page > 2 MB decompressed | −10 |
301
377
  | Page > 500 KB decompressed | −5 |
378
+ | Content tokens, wire tokens and markup share (estimated at 4 chars per token) | informational, 0 |
379
+ | Response over 2s | informational, 0 |
380
+
381
+ ---
382
+
383
+ ### `ai-directives` — page-level AI controls
384
+
385
+ The controls Google and Microsoft document that they honor, read from robots meta tags and `X-Robots-Tag` (including the user-agent-scoped header form).
386
+
387
+ | Condition | Points |
388
+ | --- | --- |
389
+ | Homepage HTML unavailable | **hard fail → 0** |
390
+ | `noindex` or `none` | **hard fail → 0** — invisible to every search-grounded assistant |
391
+ | `nosnippet`, or `max-snippet:0` | −30 — excluded as a direct input to Google AI Overviews and AI Mode |
392
+ | `noarchive` | −30 — excluded from Microsoft Copilot grounding |
393
+ | `nocache` | −10 — Copilot may use only the URL, title and snippet |
394
+ | `data-nosnippet` wrapping `<main>`, `<article>` or `<body>` | −20 |
395
+ | `noimageindex` | −5 |
396
+ | `max-snippet:[n]` under 160 | warn only, 0 |
397
+ | `noai` / `noimageai` | reported, 0 — no major operator documents honoring them |
398
+ | robots.txt disallows `Google-Extended` with no snippet directive set | warn only, 0 |
399
+
400
+ That last row is the finding this check exists for. `Google-Extended` governs Gemini training and grounding in Gemini Apps and Vertex AI, **not** AI Overviews, which follow Googlebot and the snippet directives. A site that disallows it expecting to leave AI Overviews has opted out of the thing it probably did not mind.
401
+
402
+ ### `usage-policy` — do your usage signals agree?
403
+
404
+ Normalises every machine-readable usage declaration onto three questions — may you train on it, ground an answer in it, index it — and reports where they disagree.
405
+
406
+ | Mechanism | Training | Grounding | Search |
407
+ | --- | --- | --- | --- |
408
+ | Content Signals (robots.txt or header) | `ai-train=yes\|no` | `ai-input=yes\|no` | `search=yes\|no` |
409
+ | IETF AIPREF (robots.txt or header) | `train-ai=y\|n` | *(no category yet)* | `search=y\|n` |
410
+ | RSL licence | `ai-train` | `ai-input` | `ai-index`, `search` |
411
+ | TDMRep (meta > header > well-known) | `tdm-reservation: 0\|1` | — | — |
412
+ | robots meta | `noai` | — | — |
413
+
414
+ | Condition | Points |
415
+ | --- | --- |
416
+ | No declaration of any kind | **→ 40** |
417
+ | Two mechanisms give opposite answers on one dimension | −25 per dimension |
418
+ | `Content-Usage` header outside the AIPREF vocabulary | warn only, 0 |
419
+ | A dimension no declaration covers | warn only, 0 |
420
+
421
+ Every report states that only robots.txt access rules are documented as honored by major AI operators. The rest are declarations whose weight is legal rather than technical.
422
+
423
+ ### `http-hygiene` — status-code honesty
424
+
425
+ | Condition | Points |
426
+ | --- | --- |
427
+ | A nonexistent path returns 200 | −30 |
428
+ | A nonexistent path redirects | −20 |
429
+ | A 429 with no `Retry-After` | −20 |
430
+ | A 429 with `Retry-After` on the second request | −10 |
431
+ | `HEAD` refused (405/501) | −10 |
432
+ | Over one redirect hop to the homepage | −10 |
433
+ | No `Content-Type` header | −10 |
434
+ | No charset in the header or the document | −10 |
435
+ | `<html lang>` disagrees with `Content-Language` | −5 |
436
+ | A nonexistent path returns 403/401 | −5 |
437
+ | Empty 404 body | −5 |
438
+ | The 404 probe was challenged by bot management | warn only, 0 |
439
+
440
+ ### `ai-catalog` — the index of everything callable
441
+
442
+ Discovery, in the order Lighthouse's `ard-schema` audit uses: robots.txt `Agentmap:`, `Link: rel="ai-catalog"`, `<link rel="ai-catalog">`, then `/.well-known/ai-catalog.json` and `/.well-known/ard.json`. Both specifications are drafts, so absence warns and scores nothing.
443
+
444
+ | Condition | Points |
445
+ | --- | --- |
446
+ | Catalog is not valid JSON | **→ 10** |
447
+ | An entry points at a document that cannot be fetched | −15 each |
448
+ | No entries | −20 |
449
+ | Entry missing identifier, type, or url/data | −10 |
450
+ | No `specVersion` / no `host` | −5 each |
451
+ | Entry served with a different media type than declared | −5 |
452
+
453
+ ### `agent-skills` — installable procedures
454
+
455
+ Conditional: **n/a** unless the site has a developer-facing surface (documentation links, llms.txt, or an API description). Probes `/.well-known/agent-skills/index.json`, `/.well-known/skills/index.json`, then `/skill.md`.
456
+
457
+ | Condition | Points |
458
+ | --- | --- |
459
+ | Index is not valid JSON | **→ 10** |
460
+ | A sampled skill is unreachable, or its frontmatter name disagrees with the index | −10 per problem |
461
+ | Index lists no skills | −30 |
462
+ | A skill has no description | −15 |
463
+ | No entry carries a url | −15 |
464
+ | A skill name is outside `[a-z0-9-]{1,64}` | −10 |
465
+ | Malformed digest, unknown type, over-long description, or no `$schema` | −5 each |
466
+ | No skill declares a digest | −5 |
467
+ | A single `/skill.md` with no index | −20 |
468
+
469
+ ### `webmcp` — forms as callable tools
470
+
471
+ Conditional: **n/a** on a page with no forms and no WebMCP code. Never asks for WebMCP — it is a Community Group draft in a Chrome origin trial.
472
+
473
+ | Condition | Points |
474
+ | --- | --- |
475
+ | `toolname` with no `tooldescription`, or the reverse | −30 |
476
+ | Parameters with no `toolparamdescription` | −15 |
477
+ | Tool name is not a usable identifier | −10 |
478
+ | Deprecated `navigator.modelContext` namespace | −10 |
479
+ | Forms present, none annotated | warn only, 0 |
480
+
481
+ ### `agent-operability` — can a browser agent work this page?
482
+
483
+ Browser agents read the accessibility tree, not the pixels. A `<div onclick>` styled as a button has no role and no name, so it does not appear in the tree at all: the agent does not see a button it cannot press, it sees nothing.
484
+
485
+ | Condition | Points |
486
+ | --- | --- |
487
+ | Homepage HTML unavailable | **hard fail → 0** |
488
+ | Under 90% of buttons and links have an accessible name | −20 |
489
+ | Under 90% of form controls are labelled | −20 |
490
+ | Clickable elements that are not buttons or links (no `role` + `tabindex`) | −15 |
491
+ | A CAPTCHA or `<meta http-equiv="refresh">` on the entry page | −15 |
492
+ | Links with no `href` or a `javascript:` one | −10 |
493
+ | Tables with no `<th>` | −10 |
494
+ | Untitled iframes, `<time>` without `datetime`, heading-level skips, unsized media, no `<html lang>` | −5 each |
495
+
496
+ Names are read from visible text, `aria-label`, `aria-labelledby`, `title`, image `alt`, and SVG `<title>`. Labels from `<label for>`, a wrapping label, or ARIA.
497
+
498
+ Every run ends with a method note: this reads markup, not a rendered accessibility tree, so labels attached by script and roles computed at runtime are invisible to it. A low proportion is a prompt to check the real tree, not a count to act on blindly. Every finding is also a plain accessibility defect.
499
+
500
+ ### `commerce-discovery` — Universal Commerce Protocol
501
+
502
+ Conditional: **n/a** unless the page shows storefront signals. A lone `Offer` is a price statement, not a catalog, so it counts only alongside a cart link.
503
+
504
+ | Condition | Points |
505
+ | --- | --- |
506
+ | Profile requires authentication | **hard fail → 0** |
507
+ | Profile is not valid JSON | **→ 10** |
508
+ | No `ucp` object | **→ 20** |
509
+ | No services declared | −25 |
510
+ | A declared schema URL cannot be fetched | −20 |
511
+ | No version | −20 |
512
+ | No payment handlers | −15 |
513
+ | Version is not a specification date | −10 |
514
+ | No signing keys | −10 |
515
+ | No schema URL declared | −10 |
516
+ | Service name not in reverse-DNS form, unnamed handler | −5 each |
517
+
518
+ The OpenAI and Stripe Agentic Commerce Protocol defines no manifest, and AP2 advertises through an A2A card extension, so neither is probed.
519
+
520
+ ### `auth-discovery` — can an agent get credentials?
521
+
522
+ Conditional: **n/a** unless the site exposes an API description, API catalog, MCP server card or commerce profile. Follows the RFC 9728 chain from `WWW-Authenticate` or `/.well-known/oauth-protected-resource` to the authorization server's RFC 8414 or OpenID metadata.
523
+
524
+ | Condition | Points |
525
+ | --- | --- |
526
+ | Metadata names no authorization server | **→ 40** |
527
+ | Authorization server publishes no discovery metadata | −30 |
528
+ | Invalid issuer URL | −25 |
529
+ | Missing `issuer`, `authorization_endpoint` or `token_endpoint` | −15 each |
530
+ | No PKCE with `S256` | −15 |
531
+ | No dynamic registration and no Client ID Metadata Documents | −10 |
532
+ | No `resource` identifier | −10 |
302
533
 
303
534
  ---
304
535
 
@@ -312,6 +543,8 @@ overall = round( Σ (score_i / 100 × weight_i) / Σ weight_i × 100 )
312
543
 
313
544
  When every selected check has weight 0 (e.g. `--checks rsl`), the overall falls back to a plain average of check scores.
314
545
 
546
+ Checks reporting `applicable: false` are excluded from both the numerator and the denominator. A check whose meta exists but produced no result — because it crashed — still counts at full weight, so a broken check cannot inflate a score by shrinking the denominator.
547
+
315
548
  | Grade | Score | Exit code |
316
549
  | --- | --- | --- |
317
550
  | Excellent | 90–100 | 0 |
package/docs/cli.md CHANGED
@@ -65,3 +65,39 @@ npx ax-audit $(cat urls.txt) --concurrency 4 --json > batch.json
65
65
  # Show me only what is broken
66
66
  npx ax-audit https://your-site.com --only-failures
67
67
  ```
68
+
69
+ ## Agent-surface profiles
70
+
71
+ Protocol checks (`api-discovery`, `agent-card`, `mcp-discovery`, `agent-skills`, `auth-discovery`, `commerce-discovery`) apply only where the site has the corresponding surface. A blog has no API to describe, so scoring it zero would say something false; those checks report **n/a** and leave the denominator.
72
+
73
+ `--profile` overrides that, for auditing against what a site intends to build rather than what it already has.
74
+
75
+ ```bash
76
+ ax-audit https://example.com --profile api # audit as though it offered an API
77
+ ax-audit https://example.com --profile all # every protocol check applies
78
+ ```
79
+
80
+ Values: `auto` (default, detect from the site), `api`, `mcp`, `agent`, `docs`, `commerce`, `all`.
81
+
82
+ ## Focusing on one area
83
+
84
+ ```bash
85
+ ax-audit https://example.com --category access # only access checks
86
+ ax-audit https://example.com --category content,discovery # two areas
87
+ ```
88
+
89
+ `--category` narrows an explicit `--checks` selection rather than replacing it, and errors if the intersection is empty.
90
+
91
+ ## Per-area CI gates
92
+
93
+ An overall score hides an area that is entirely broken: a site can score 80 while every access check fails, because the other four carry it.
94
+
95
+ ```bash
96
+ ax-audit https://example.com --fail-on-category access:70,content:80
97
+ ```
98
+
99
+ Each area is reported against its threshold on stderr, so the output survives `--output json` being piped elsewhere. An area with no applicable checks is reported as not evaluated rather than failed: a build must not break because a site has no commerce surface.
100
+
101
+ ## Baselines across a scoring change
102
+
103
+ Baselines record which scoring model produced them. Comparing a baseline written by an older version shows the deltas but suspends regression gating, because a rescore is not something the site did. Re-save with `--save-baseline` to resume gating. Checks that changed applicability are excluded from regressions and improvements for the same reason.
package/docs/concepts.md CHANGED
@@ -7,7 +7,7 @@
7
7
  Agents are not browsers. Three differences drive every check:
8
8
 
9
9
  1. **They mostly don't run JavaScript.** GPTBot, ClaudeBot, CCBot and most crawlers fetch raw HTML. A client-rendered SPA that returns an empty `<div id="root">` is, to them, a blank page. (`html-rendering`, `content-negotiation`)
10
- 2. **They look for declared structure, not visual layout.** An agent would rather read a `/llms.txt` summary or a JSON-LD graph than infer meaning from your CSS grid. (`llms-txt`, `structured-data`, `meta-tags`, `agent-json`, `mcp`, `openapi`)
10
+ 2. **They look for declared structure, not visual layout.** An agent would rather read a `/llms.txt` summary or a JSON-LD graph than infer meaning from your CSS grid. (`llms-txt`, `structured-data`, `meta-tags`, `agent-card`, `mcp-discovery`, `api-discovery`)
11
11
  3. **Their access is a policy and economic question, not just a technical one.** Who may crawl, for what use, at what price, under what license — these now have machine-readable answers. (`robots-txt`, Content Signals, `rsl`, `agent-access`)
12
12
 
13
13
  Bot traffic is projected to exceed human traffic by 2029. AX is the interface layer for that shift.
@@ -30,11 +30,15 @@ These answer: *can an agent find your content and actually read it?*
30
30
 
31
31
  | Standard | What it is | Check |
32
32
  | --- | --- | --- |
33
- | **[A2A — Agent2Agent](https://a2a-protocol.org)** | An "Agent Card" at `/.well-known/agent.json` advertising your agent's identity and skills, so other agents can interoperate. | `agent-json` |
34
- | **[MCP — Model Context Protocol](https://modelcontextprotocol.io)** | A manifest at `/.well-known/mcp.json` describing tools and resources an agent can call. The emerging standard for exposing capabilities to LLMs. | `mcp` |
35
- | **[OpenAPI](https://www.openapis.org)** | The long-standing machine-readable API description; agents use it to call your endpoints. | `openapi` |
36
- | **Emerging discovery files** | `ai.txt`, `genai.txt`, `ai-plugin.json`, `agents.json`, `nlweb.json` — competing/early conventions, scored as coverage bonus. | `well-known-ai` |
37
- | **AI meta tags & discovery links** | `ai:*` meta tags and `rel="alternate"` links pointing agents to your llms.txt / agent.json. | `meta-tags` |
33
+ | **[A2A — Agent2Agent](https://a2a-protocol.org)** | An Agent Card at `/.well-known/agent-card.json` advertising your identity and skills, so other agents can interoperate. The path moved there in v0.3.0 and is IANA-registered; v1.0 changed the card's shape. | `agent-card` |
34
+ | **[MCP — Model Context Protocol](https://modelcontextprotocol.io)** | A **server card** identifying your MCP server and its remote endpoints. Note it carries no tool list: tools come from a live `tools/list` call, and a static copy drifts. `/.well-known/mcp.json` was never part of the specification. | `mcp-discovery` |
35
+ | **[OpenAPI](https://www.openapis.org)** and **[RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html)** | The machine-readable API description, plus the registered `/.well-known/api-catalog` that points at it. Agents use both to call your endpoints without a human reading your docs. | `api-discovery` |
36
+ | **[Agent Skills](https://agentskills.io)** | SKILL.md documents an agent installs and follows — setup steps, argument shapes, mistakes to avoid. Answers "how do I do the thing this site is for", where llms.txt answers "what is here". | `agent-skills` |
37
+ | **[RFC 9728 / RFC 8414](https://www.rfc-editor.org/rfc/rfc9728.html)** | OAuth metadata. A human hitting a 401 reads your docs; an agent cannot, so the answer has to be in the response. | `auth-discovery` |
38
+ | **[UCP](https://developers.google.com/merchant/ucp/guides/ucp-profile)** | The Universal Commerce Protocol profile at `/.well-known/ucp` — the one agentic-commerce specification with published site-side discovery. | `commerce-discovery` |
39
+ | **[WebMCP](https://webmachinelearning.github.io/webmcp/)** | Declaring a form as a callable tool, so an agent invokes it rather than driving it pixel by pixel. A Community Group draft in a Chrome origin trial. | `webmcp` |
40
+ | **AI catalogs** | `ai-catalog.json` and `ard.json`, two competing drafts for one index listing everything above, so a client stops probing four conventions. | `ai-catalog` |
41
+ | **AI meta tags & discovery links** | `ai:*` meta tags and the `rel` relations — `alternate`, `describedby`, `service-desc`, `api-catalog` — that point agents at your files instead of making them guess paths. | `meta-tags`, `http-headers` |
38
42
 
39
43
  These answer: *once an agent arrives, can it understand what you offer and act on it?*
40
44
 
@@ -44,14 +48,20 @@ This is the newest and fastest-moving family — the response to "AI scraped my
44
48
 
45
49
  | Standard | What it is | Check |
46
50
  | --- | --- | --- |
47
- | **[Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)** | The original robots.txt — *who* may crawl *what*. ax-audit grades coverage of 48 known AI crawlers. | `robots-txt` |
48
- | **[Content Signals](https://contentsignals.org)** | A robots.txt extension (Cloudflare, CC0) declaring *how* content may be used after access: `search`, `ai-input`, `ai-train`. Served by default on 3.8M+ Cloudflare domains. | `robots-txt` (findings) |
49
- | **[RSL — Really Simple Licensing](https://rslstandard.org)** | A full machine-readable licensing layer (license.xml): permits/prohibits vocabularies, payment models (free, attribution, pay-per-crawl, pay-per-inference). Endorsed by 1,500+ publishers. | `rsl` |
50
- | **Cloaking integrity** | Not a standard but a failure mode: your stated policy (robots.txt allows GPTBot) contradicting enforcement (WAF returns 403). | `agent-access` |
51
+ | **[Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)** | The original robots.txt — *who* may crawl *what*. ax-audit knows 57 AI clients, grouped by what they do with a page, because blocking a search crawler costs citations while blocking a training crawler is a policy choice. | `robots-txt` |
52
+ | **[Content Signals](https://contentsignals.org)** | A robots.txt extension (Cloudflare, CC0) declaring *how* content may be used after access: `search`, `ai-input`, `ai-train`, plus the `use` field added in 2026. Served by default on millions of Cloudflare domains. | `robots-txt`, `usage-policy` |
53
+ | **[IETF AIPREF](https://datatracker.ietf.org/wg/aipref/documents/)** | The standards-track answer to the same question: a `Content-Usage` directive or header over `train-ai` and `search`. Still pre-last-call. Note the token inversion against Content Signals — `train-ai` here, `ai-train` there. | `robots-txt`, `usage-policy` |
54
+ | **[RSL — Really Simple Licensing](https://rslstandard.org)** | A full machine-readable licensing layer: permits/prohibits vocabularies, payment models, and an Open Licensing Protocol for negotiating access. | `rsl`, `usage-policy` |
55
+ | **[TDMRep](https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240510/)** | A text-and-data-mining reservation, named in the EU GPAI Code of Practice. Its weight is legal rather than technical. | `usage-policy` |
56
+ | **Page-level AI directives** | `nosnippet` and `max-snippet` for Google AI Overviews, `noarchive` and `nocache` for Copilot grounding. Unlike everything above, the vendors document that they honor these. | `ai-directives` |
57
+ | **Cloaking and blocking integrity** | Not a standard but a failure mode: robots.txt allows GPTBot and the firewall returns 403. | `agent-access` |
58
+ | **Signal consistency** | Also not a standard: whether your Content Signals, AIPREF, RSL and TDMRep declarations agree with each other. Five documents maintained by hand do not stay in sync, and the terms that apply then depend on which file a crawler read. | `usage-policy` |
51
59
 
52
60
  These answer: *have you expressed your access and usage policy in a form agents can honor — and does your infrastructure actually match it?*
53
61
 
54
- The progression is one of increasing expressiveness: robots.txt says **who/where**, Content Signals adds **how it may be used**, RSL adds **under what license and price**.
62
+ The progression is one of increasing expressiveness: robots.txt says **who and where**, Content Signals and AIPREF add **how it may be used**, RSL adds **under what licence and price**.
63
+
64
+ One caveat worth stating plainly: only robots.txt access rules are documented as honored by Google, OpenAI, Anthropic and Microsoft. Content Signals, AIPREF, RSL and TDMRep are declarations. Their weight is legal rather than technical, and ax-audit says so on every run rather than implying a crawler will obey.
55
65
 
56
66
  ### 4. Transport, efficiency & hygiene
57
67
 
@@ -78,10 +88,14 @@ A fully AX-ready site tells a coherent story across all four:
78
88
 
79
89
  > "Here's my content in a form you can read **(family 1)**, here's the interface to interact with me **(family 2)**, here's exactly who may use it and how, for what license **(family 3)**, over a fast and trustworthy connection **(family 4)**."
80
90
 
81
- ax-audit's weighting reflects today's leverage: discovery and readability (`llms-txt`, `robots-txt`, `html-rendering`, `structured-data`, `http-headers`) carry the most weight because they're the highest-impact, most-adopted signals. The governance and efficiency standards are informational in 3.x — real and worth adopting, but still stabilizing — and gain weight in v4.0.
91
+ ax-audit's weighting follows what actually stops an agent. **Content** carries the most weight, led by `html-rendering`, because a page whose content only exists after hydration is invisible to the crawlers that do not run JavaScript, no matter how many discovery files it publishes. **Access** is next, because a firewall rule or a `nosnippet` directive silently undoes everything else.
92
+
93
+ `llms-txt` sits at 5 rather than the 11 it carried in 3.x. Adoption studies find most published files are never fetched by an AI search crawler, and Google has stated that Search ignores them; the vendors that do read them are coding agents like Claude Code and Cursor. It is worth publishing, as developer tooling, and it is not worth twice what having content is worth.
94
+
95
+ Checks resting on draft specifications — `ai-catalog`, `webmcp`, `commerce-discovery` — are reported but never scored. A specification that may be renamed next quarter should not move your number.
82
96
 
83
97
  ## See also
84
98
 
85
99
  - [getting-started.md](./getting-started.md) — run your first audit
86
100
  - [checks.md](./checks.md) — exact scoring per standard
87
- - The [remediation guides](https://lucioduran.com/projects/ax-audit/guides) — how to implement each one
101
+ - The [remediation guides](https://axrush.com/guides) — how to implement each one