ax-audit 3.6.0 → 4.0.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +138 -0
- package/LICENSE +1 -1
- package/README.md +58 -36
- package/dist/baseline.d.ts +2 -0
- package/dist/baseline.d.ts.map +1 -1
- package/dist/baseline.js +42 -4
- package/dist/baseline.js.map +1 -1
- package/dist/check-ids.d.ts +19 -0
- package/dist/check-ids.d.ts.map +1 -0
- package/dist/check-ids.js +53 -0
- package/dist/check-ids.js.map +1 -0
- package/dist/checks/agent-access.d.ts +23 -6
- package/dist/checks/agent-access.d.ts.map +1 -1
- package/dist/checks/agent-access.js +200 -54
- package/dist/checks/agent-access.js.map +1 -1
- package/dist/checks/agent-card.d.ts +37 -0
- package/dist/checks/agent-card.d.ts.map +1 -0
- package/dist/checks/agent-card.js +352 -0
- package/dist/checks/agent-card.js.map +1 -0
- package/dist/checks/agent-operability.d.ts +66 -0
- package/dist/checks/agent-operability.d.ts.map +1 -0
- package/dist/checks/agent-operability.js +383 -0
- package/dist/checks/agent-operability.js.map +1 -0
- package/dist/checks/agent-skills.d.ts +24 -0
- package/dist/checks/agent-skills.d.ts.map +1 -0
- package/dist/checks/agent-skills.js +316 -0
- package/dist/checks/agent-skills.js.map +1 -0
- package/dist/checks/ai-catalog.d.ts +28 -0
- package/dist/checks/ai-catalog.d.ts.map +1 -0
- package/dist/checks/ai-catalog.js +254 -0
- package/dist/checks/ai-catalog.js.map +1 -0
- package/dist/checks/ai-directives.d.ts +57 -0
- package/dist/checks/ai-directives.d.ts.map +1 -0
- package/dist/checks/ai-directives.js +263 -0
- package/dist/checks/ai-directives.js.map +1 -0
- package/dist/checks/api-discovery.d.ts +26 -0
- package/dist/checks/api-discovery.d.ts.map +1 -0
- package/dist/checks/api-discovery.js +432 -0
- package/dist/checks/api-discovery.js.map +1 -0
- package/dist/checks/auth-discovery.d.ts +28 -0
- package/dist/checks/auth-discovery.d.ts.map +1 -0
- package/dist/checks/auth-discovery.js +213 -0
- package/dist/checks/auth-discovery.js.map +1 -0
- package/dist/checks/commerce-discovery.d.ts +40 -0
- package/dist/checks/commerce-discovery.d.ts.map +1 -0
- package/dist/checks/commerce-discovery.js +295 -0
- package/dist/checks/commerce-discovery.js.map +1 -0
- package/dist/checks/content-negotiation.d.ts.map +1 -1
- package/dist/checks/content-negotiation.js +135 -20
- package/dist/checks/content-negotiation.js.map +1 -1
- package/dist/checks/crawl-efficiency.d.ts +13 -1
- package/dist/checks/crawl-efficiency.d.ts.map +1 -1
- package/dist/checks/crawl-efficiency.js +65 -1
- package/dist/checks/crawl-efficiency.js.map +1 -1
- package/dist/checks/frontmatter.d.ts +34 -0
- package/dist/checks/frontmatter.d.ts.map +1 -0
- package/dist/checks/frontmatter.js +100 -0
- package/dist/checks/frontmatter.js.map +1 -0
- package/dist/checks/html-rendering.d.ts.map +1 -1
- package/dist/checks/html-rendering.js +0 -1
- package/dist/checks/html-rendering.js.map +1 -1
- package/dist/checks/html-utils.d.ts +10 -0
- package/dist/checks/html-utils.d.ts.map +1 -1
- package/dist/checks/html-utils.js +19 -0
- package/dist/checks/html-utils.js.map +1 -1
- package/dist/checks/http-headers.d.ts.map +1 -1
- package/dist/checks/http-headers.js +82 -10
- package/dist/checks/http-headers.js.map +1 -1
- package/dist/checks/http-hygiene.d.ts +26 -0
- package/dist/checks/http-hygiene.d.ts.map +1 -0
- package/dist/checks/http-hygiene.js +257 -0
- package/dist/checks/http-hygiene.js.map +1 -0
- package/dist/checks/index.d.ts.map +1 -1
- package/dist/checks/index.js +24 -8
- package/dist/checks/index.js.map +1 -1
- package/dist/checks/llms-txt.d.ts +15 -0
- package/dist/checks/llms-txt.d.ts.map +1 -1
- package/dist/checks/llms-txt.js +162 -2
- package/dist/checks/llms-txt.js.map +1 -1
- package/dist/checks/mcp-discovery.d.ts +30 -0
- package/dist/checks/mcp-discovery.d.ts.map +1 -0
- package/dist/checks/mcp-discovery.js +523 -0
- package/dist/checks/mcp-discovery.js.map +1 -0
- package/dist/checks/meta-tags.d.ts.map +1 -1
- package/dist/checks/meta-tags.js +6 -5
- package/dist/checks/meta-tags.js.map +1 -1
- package/dist/checks/robots-parser.d.ts +110 -0
- package/dist/checks/robots-parser.d.ts.map +1 -0
- package/dist/checks/robots-parser.js +277 -0
- package/dist/checks/robots-parser.js.map +1 -0
- package/dist/checks/robots-txt.d.ts +2 -20
- package/dist/checks/robots-txt.d.ts.map +1 -1
- package/dist/checks/robots-txt.js +219 -120
- package/dist/checks/robots-txt.js.map +1 -1
- package/dist/checks/rsl.d.ts +0 -2
- package/dist/checks/rsl.d.ts.map +1 -1
- package/dist/checks/rsl.js +1 -11
- package/dist/checks/rsl.js.map +1 -1
- package/dist/checks/security-txt.d.ts.map +1 -1
- package/dist/checks/security-txt.js +0 -1
- package/dist/checks/security-txt.js.map +1 -1
- package/dist/checks/seo-basics.d.ts.map +1 -1
- package/dist/checks/seo-basics.js +0 -1
- package/dist/checks/seo-basics.js.map +1 -1
- package/dist/checks/sitemap.d.ts.map +1 -1
- package/dist/checks/sitemap.js +0 -1
- package/dist/checks/sitemap.js.map +1 -1
- package/dist/checks/structured-data.d.ts.map +1 -1
- package/dist/checks/structured-data.js +215 -4
- package/dist/checks/structured-data.js.map +1 -1
- package/dist/checks/structured-fields.d.ts +46 -0
- package/dist/checks/structured-fields.d.ts.map +1 -0
- package/dist/checks/structured-fields.js +112 -0
- package/dist/checks/structured-fields.js.map +1 -0
- package/dist/checks/surface.d.ts +59 -0
- package/dist/checks/surface.d.ts.map +1 -0
- package/dist/checks/surface.js +106 -0
- package/dist/checks/surface.js.map +1 -0
- package/dist/checks/tls-https.d.ts.map +1 -1
- package/dist/checks/tls-https.js +0 -1
- package/dist/checks/tls-https.js.map +1 -1
- package/dist/checks/usage-policy.d.ts +53 -0
- package/dist/checks/usage-policy.d.ts.map +1 -0
- package/dist/checks/usage-policy.js +339 -0
- package/dist/checks/usage-policy.js.map +1 -0
- package/dist/checks/utils.d.ts +25 -1
- package/dist/checks/utils.d.ts.map +1 -1
- package/dist/checks/utils.js +33 -1
- package/dist/checks/utils.js.map +1 -1
- package/dist/checks/waf.d.ts +75 -0
- package/dist/checks/waf.d.ts.map +1 -0
- package/dist/checks/waf.js +203 -0
- package/dist/checks/waf.js.map +1 -0
- package/dist/checks/webmcp.d.ts +55 -0
- package/dist/checks/webmcp.d.ts.map +1 -0
- package/dist/checks/webmcp.js +209 -0
- package/dist/checks/webmcp.js.map +1 -0
- package/dist/checks/well-known.d.ts +38 -0
- package/dist/checks/well-known.d.ts.map +1 -0
- package/dist/checks/well-known.js +202 -0
- package/dist/checks/well-known.js.map +1 -0
- package/dist/cli.d.ts +14 -0
- package/dist/cli.d.ts.map +1 -1
- package/dist/cli.js +124 -4
- package/dist/cli.js.map +1 -1
- package/dist/constants.d.ts +195 -14
- package/dist/constants.d.ts.map +1 -1
- package/dist/constants.js +598 -70
- package/dist/constants.js.map +1 -1
- package/dist/fetcher.d.ts.map +1 -1
- package/dist/fetcher.js +44 -14
- package/dist/fetcher.js.map +1 -1
- package/dist/guide-urls.js +1 -1
- package/dist/guide-urls.js.map +1 -1
- package/dist/index.d.ts +2 -1
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +1 -0
- package/dist/index.js.map +1 -1
- package/dist/orchestrator.d.ts.map +1 -1
- package/dist/orchestrator.js +5 -1
- package/dist/orchestrator.js.map +1 -1
- package/dist/reporter/html.d.ts +9 -0
- package/dist/reporter/html.d.ts.map +1 -1
- package/dist/reporter/html.js +49 -11
- package/dist/reporter/html.js.map +1 -1
- package/dist/reporter/markdown.d.ts.map +1 -1
- package/dist/reporter/markdown.js +36 -6
- package/dist/reporter/markdown.js.map +1 -1
- package/dist/reporter/terminal.d.ts.map +1 -1
- package/dist/reporter/terminal.js +36 -1
- package/dist/reporter/terminal.js.map +1 -1
- package/dist/scorer.d.ts +10 -0
- package/dist/scorer.d.ts.map +1 -1
- package/dist/scorer.js +22 -5
- package/dist/scorer.js.map +1 -1
- package/dist/types.d.ts +90 -3
- package/dist/types.d.ts.map +1 -1
- package/docs/architecture.md +27 -11
- package/docs/checks.md +285 -52
- package/docs/cli.md +36 -0
- package/docs/concepts.md +27 -13
- package/docs/faq.md +18 -6
- package/docs/getting-started.md +20 -13
- package/docs/roadmap.md +367 -0
- package/package.json +13 -5
- package/dist/checks/agent-json.d.ts +0 -14
- package/dist/checks/agent-json.d.ts.map +0 -1
- package/dist/checks/agent-json.js +0 -167
- package/dist/checks/agent-json.js.map +0 -1
- package/dist/checks/mcp.d.ts +0 -4
- package/dist/checks/mcp.d.ts.map +0 -1
- package/dist/checks/mcp.js +0 -162
- package/dist/checks/mcp.js.map +0 -1
- package/dist/checks/openapi.d.ts +0 -4
- package/dist/checks/openapi.d.ts.map +0 -1
- package/dist/checks/openapi.js +0 -121
- package/dist/checks/openapi.js.map +0 -1
- package/dist/checks/well-known-ai.d.ts +0 -17
- package/dist/checks/well-known-ai.d.ts.map +0 -1
- package/dist/checks/well-known-ai.js +0 -123
- package/dist/checks/well-known-ai.js.map +0 -1
package/docs/checks.md
CHANGED
|
@@ -1,8 +1,29 @@
|
|
|
1
1
|
# Checks Reference
|
|
2
2
|
|
|
3
|
-
ax-audit runs
|
|
3
|
+
ax-audit runs 26 checks. Twenty-three are **weighted**, summing to 100. Three rest on draft specifications and stay at weight 0: scoring a site against a specification that may be renamed next quarter would make the number less trustworthy, not more.
|
|
4
4
|
|
|
5
|
-
|
|
5
|
+
Checks are grouped into five areas, and reports are ordered by them: **content** (is there substance an agent can read?), **discovery** (can an agent find your machine-readable files?), **access** (can it actually retrieve them?), **policy** (what usage rights do you declare?), and **protocols** (what can an agent call?).
|
|
6
|
+
|
|
7
|
+
Some checks are **conditional**. A blog has no API to describe, no MCP server to advertise and nothing to authorize, so those checks report **n/a** and are excluded from the score rather than counted as failures. Everything counted against a site is something the site could have done. `--profile api|mcp|agent|docs|commerce|all` forces them applicable, for auditing against what a site intends to become.
|
|
8
|
+
|
|
9
|
+
## Weights
|
|
10
|
+
|
|
11
|
+
| Area | Total | Checks |
|
|
12
|
+
| --- | --- | --- |
|
|
13
|
+
| Content | 33 | html-rendering 11 · agent-operability 7 · structured-data 6 · seo-basics 5 · content-negotiation 4 |
|
|
14
|
+
| Access | 24 | agent-access 9 · ai-directives 6 · http-hygiene 4 · tls-https 3 · crawl-efficiency 2 |
|
|
15
|
+
| Discovery | 21 | robots-txt 9 · llms-txt 5 · http-headers 4 · sitemap 2 · meta-tags 1 |
|
|
16
|
+
| Protocols | 13 | api-discovery 4 · agent-card 3 · mcp-discovery 3 · agent-skills 2 · auth-discovery 1 |
|
|
17
|
+
| Policy | 9 | usage-policy 4 · security-txt 3 · rsl 2 |
|
|
18
|
+
| Draft specifications | 0 | ai-catalog · webmcp · commerce-discovery |
|
|
19
|
+
|
|
20
|
+
Content leads because the failure that breaks the most agents is a page with nothing in its HTML: most crawlers do not run JavaScript, and a site whose content appears only after hydration is invisible no matter how many discovery files it publishes. Access is second because a firewall rule or a `nosnippet` directive silently undoes everything else, and those are the failures operators are least likely to know about. llms.txt sits at 5 rather than the 11 it carried in 3.x, because most published files are never fetched by an AI search crawler.
|
|
21
|
+
|
|
22
|
+
Weights live in `CHECK_WEIGHTS` in `src/constants.ts`, and only there. Checks used to declare their own alongside it, and the two drifted.
|
|
23
|
+
|
|
24
|
+
Every probed path is labelled by standing — **IANA-registered**, **vendor convention**, **draft**, or **legacy** — because the agent web mixes registered URIs with drafts that get renamed. A missing draft file is not the same kind of finding as a missing registered one, and reports say which is which.
|
|
25
|
+
|
|
26
|
+
This page documents the **exact scoring** of every check: each deduction, bonus, and formula, extracted from the source. Every finding links to a step-by-step remediation guide at `axrush.com/guides/<check-id>`.
|
|
6
27
|
|
|
7
28
|
**Reading the tables:** each check starts at 100 unless noted. Deductions stack additively; `buildResult` clamps the final score to [0, 100]. "Hard fail" rows short-circuit the check.
|
|
8
29
|
|
|
@@ -24,10 +45,19 @@ This page documents the **exact scoring** of every check: each deduction, bonus,
|
|
|
24
45
|
| No Markdown links | −10 |
|
|
25
46
|
| Content under 100 characters | −10 |
|
|
26
47
|
| `/llms-full.txt` also available | **+10** (capped at 100) |
|
|
48
|
+
| Broken links among a sample of 15 | informational, 0 in 3.x |
|
|
49
|
+
| Redirecting or duplicated links | informational, 0 |
|
|
50
|
+
| No `rel="describedby"` pointing at the file (llms.txt v2) | informational, 0 |
|
|
51
|
+
| No per-page Markdown mirror (`/index.md`, `/index.html.md`) | informational, 0 |
|
|
52
|
+
| File over 50 KB | informational, 0 |
|
|
53
|
+
|
|
54
|
+
Every report states plainly that Google says Search ignores llms.txt and that most published files are never fetched by an AI search crawler, while Claude Code, Cursor and OpenCode do read it. It is a developer-tooling signal, not a search-visibility one.
|
|
27
55
|
|
|
28
56
|
### `robots-txt` — 11%
|
|
29
57
|
|
|
30
|
-
AI-crawler configuration.
|
|
58
|
+
AI-crawler configuration. Scoring runs against the frozen 3.x core set (GPTBot, ClaudeBot, ChatGPT-User, Claude-SearchBot, Google-Extended, PerplexityBot, OAI-SearchBot, CCBot); the wider September-2026 core set adds Meta-ExternalAgent, Applebot-Extended, Amazonbot and Bytespider, reported but not scored until 4.0.
|
|
59
|
+
|
|
60
|
+
Findings are tiered by what a client does with a page, because that determines the cost of blocking it. Blocking a **training** crawler is a policy choice and is reported as such. Blocking a **search** crawler removes the site from that assistant's answers. Blocking a **user-triggered fetcher** often does nothing, because most vendors document that robots.txt may not apply to them.
|
|
31
61
|
|
|
32
62
|
| Condition | Points |
|
|
33
63
|
| --- | --- |
|
|
@@ -38,7 +68,10 @@ AI-crawler configuration. Core crawlers: GPTBot, ClaudeBot, ChatGPT-User, Claude
|
|
|
38
68
|
| Known AI crawler(s) explicitly blocked (`Disallow: /`) | −3 per crawler |
|
|
39
69
|
| No `Sitemap:` directive | −5 |
|
|
40
70
|
| Partial path restrictions on AI crawlers | warn only, 0 |
|
|
41
|
-
|
|
|
71
|
+
| Blocking a crawler token added in 3.7 (`meta-webindexer`, `Amzn-SearchBot`, …) | informational, 0 in 3.x |
|
|
72
|
+
| Rules targeting a retired or fictional token (`GeminiBot`, `Claude-Web`, `NeevaBot`, …) | informational, 0 |
|
|
73
|
+
| [Content Signals](https://contentsignals.org) findings, including the `use=immediate\|reference\|full` field | informational, 0 in 3.x |
|
|
74
|
+
| [IETF AIPREF](https://datatracker.ietf.org/wg/aipref/documents/) `Content-Usage:` findings, including vocabulary mix-ups | informational, 0 in 3.x |
|
|
42
75
|
|
|
43
76
|
### `html-rendering` — 9%
|
|
44
77
|
|
|
@@ -72,54 +105,84 @@ JSON-LD on the homepage. Key entity types: Person, Organization, WebSite, WebPag
|
|
|
72
105
|
| Only one key entity type | −10 |
|
|
73
106
|
| No `@graph` array | −5 |
|
|
74
107
|
| No `BreadcrumbList` | −5 |
|
|
108
|
+
| No `author`, no `sameAs`, or an author with no `publisher` | informational, 0 in 3.x |
|
|
109
|
+
| No `dateModified` / `datePublished`, a future date, or content over 730 days old | informational, 0 |
|
|
110
|
+
| `headline` or `name` not present in the visible text | informational, 0 |
|
|
111
|
+
|
|
112
|
+
The visible-text comparison is Google's one explicit requirement for structured data and AI features. It reads static HTML, so text rendered by script reads as missing here too, and the finding says so.
|
|
75
113
|
|
|
76
114
|
### `http-headers` — 9%
|
|
77
115
|
|
|
78
|
-
Security headers, AI discovery `Link` headers (RFC 5988-parsed), CORS on `.well-known`.
|
|
116
|
+
Security headers, AI discovery `Link` headers (RFC 5988-parsed), CORS on `.well-known`. Either Agent Card path satisfies the discovery-link requirement.
|
|
79
117
|
|
|
80
118
|
| Condition | Points |
|
|
81
119
|
| --- | --- |
|
|
82
120
|
| No headers retrievable | **hard fail → 0** |
|
|
83
121
|
| Missing critical security header (HSTS, X-Content-Type-Options) | −10 each |
|
|
84
122
|
| Only 1–3 of the 7 tracked security headers present | −5 |
|
|
85
|
-
| `Link` header missing both llms.txt and
|
|
123
|
+
| `Link` header missing both llms.txt and the Agent Card | −15 |
|
|
86
124
|
| `Link` header missing one of the two | −5 |
|
|
87
|
-
| No CORS on
|
|
125
|
+
| No CORS on the Agent Card | −10 |
|
|
126
|
+
| Additional discovery relations (`describedby`, `api-catalog`, `service-desc`, `service-doc`, `ai-catalog`, `c2pa-manifest`, `license`, markdown `alternate`) and the `X-Llms-Txt` header | informational, 0 in 3.x |
|
|
88
127
|
|
|
89
|
-
### `agent-
|
|
128
|
+
### `agent-card` — 7%
|
|
90
129
|
|
|
91
|
-
|
|
130
|
+
The [A2A Agent Card](https://a2a-protocol.org), probed at `/.well-known/agent-card.json` (IANA-registered since A2A v0.3.0, 2025-07-30) and then at the pre-0.3 path `/.well-known/agent.json`. *Former id: `agent-json`, still accepted in `--checks` and in saved baselines.*
|
|
131
|
+
|
|
132
|
+
Two spec generations are in the wild, and the check detects which one a card follows from its own structure rather than from a version field:
|
|
133
|
+
|
|
134
|
+
- **A2A 1.0** (2026-03-12) declares every endpoint inside `supportedInterfaces[]`. Required: `name`, `description`, `version`, `capabilities`, `supportedInterfaces`, `defaultInputModes`, `defaultOutputModes`, `skills`.
|
|
135
|
+
- **A2A 0.3** declares a top-level `url` and `protocolVersion`. Required: those two plus `name`, `description`, `version`, `capabilities`, `defaultInputModes`, `defaultOutputModes`, `skills`.
|
|
92
136
|
|
|
93
137
|
| Condition | Points |
|
|
94
138
|
| --- | --- |
|
|
95
|
-
| Not found | **hard fail → 0** |
|
|
139
|
+
| Not found at either path | **hard fail → 0** |
|
|
96
140
|
| Invalid JSON | **→ 10** |
|
|
97
|
-
|
|
|
98
|
-
|
|
|
99
|
-
|
|
|
100
|
-
|
|
|
141
|
+
| Served only from the pre-0.3 `agent.json` path | warn only, 0 |
|
|
142
|
+
| Wrong Content-Type (expected `application/json` or `application/a2a+json`) | −5 |
|
|
143
|
+
| Card shape matches neither generation | −30 |
|
|
144
|
+
| Missing required field for the detected generation | −15 per field |
|
|
145
|
+
| `supportedInterfaces[]` empty (1.0) | −15 |
|
|
146
|
+
| Interface missing `url`, `protocolBinding` or `protocolVersion` (1.0) | −10 |
|
|
147
|
+
| Unrecognised `protocolBinding` (not JSONRPC / GRPC / HTTP+JSON) | −5 |
|
|
148
|
+
| Interface or `url` on a different origin | −5 |
|
|
149
|
+
| `url` not an absolute URL (0.3) | −5 |
|
|
101
150
|
| `skills` empty | −10 |
|
|
102
151
|
| `skills` entries missing `id` or `description` | −5 |
|
|
103
|
-
|
|
|
104
|
-
| No optional fields (`
|
|
152
|
+
| Uses `authentication`, removed from the spec in 0.2.x | −5 |
|
|
153
|
+
| No optional descriptive fields (`provider`, `documentationUrl`, `iconUrl`) | −5 |
|
|
154
|
+
|
|
155
|
+
### `mcp-discovery` — 7%
|
|
156
|
+
|
|
157
|
+
How an agent finds this site's [Model Context Protocol](https://modelcontextprotocol.io) server. *Former id: `mcp`.*
|
|
105
158
|
|
|
106
|
-
|
|
159
|
+
`/.well-known/mcp.json` was never part of the MCP specification. What emerged instead is the **server card**, which deliberately carries no `tools[]` — tool lists come from a live `tools/list` call, and a static copy drifts the day it is written. Discovery is probed in this order:
|
|
107
160
|
|
|
108
|
-
`/.well-known/
|
|
161
|
+
1. `/.well-known/ai-catalog.json` entries of type `application/mcp-server-card+json` *(draft)*
|
|
162
|
+
2. `/.well-known/mcp/server-card.json`, `/.well-known/mcp/server-cards.json` *(vendor convention: Cloudflare, Mintlify)*
|
|
163
|
+
3. `<endpoint>/server-card` for `/mcp`, `/api/mcp`, `/sse` *(the MCP extension's own recommendation)*
|
|
164
|
+
4. `/.well-known/mcp.json` *(legacy)*
|
|
165
|
+
|
|
166
|
+
An HTML response counts as absence, not a malformed card: SPA catch-alls answer every unknown path with the index shell.
|
|
167
|
+
|
|
168
|
+
**When a server card is found:**
|
|
109
169
|
|
|
110
170
|
| Condition | Points |
|
|
111
171
|
| --- | --- |
|
|
112
|
-
|
|
|
172
|
+
| Wrong Content-Type (expected `application/json` or `application/mcp-server-card+json`) | −5 |
|
|
113
173
|
| Invalid JSON | **→ 10** |
|
|
114
|
-
|
|
|
115
|
-
|
|
|
116
|
-
|
|
|
117
|
-
|
|
|
118
|
-
|
|
|
119
|
-
|
|
|
120
|
-
|
|
|
121
|
-
|
|
|
174
|
+
| Missing `$schema`, `name`, `version` or `description` | −15 each |
|
|
175
|
+
| `name` not in reverse-DNS form | −5 |
|
|
176
|
+
| No `remotes[]` | −15 |
|
|
177
|
+
| Remote with an unrecognised transport (not `streamable-http` / `sse`) | −5 |
|
|
178
|
+
| Remote missing a `url` | −10 |
|
|
179
|
+
| No `supportedProtocolVersions` | −10 |
|
|
180
|
+
| Only pre-2025-06 protocol revisions | −10 |
|
|
181
|
+
| Unrecognised protocol version | −5 |
|
|
122
182
|
| No CORS headers | −10 |
|
|
183
|
+
| Declares `tools[]`, which the schema omits by design | warn only, 0 |
|
|
184
|
+
|
|
185
|
+
**When only `/.well-known/mcp.json` is found**, the pre-3.7 rules are applied unchanged so the score is exactly what 3.6 produced (missing `name` −10, missing `description` −5, no `tools` −15, no tool descriptions −10 / −5, no `resources` −5, no version −5, no CORS −10, wrong Content-Type −5). The path itself is reported, not penalised.
|
|
123
186
|
|
|
124
187
|
### `seo-basics` — 7%
|
|
125
188
|
|
|
@@ -161,7 +224,7 @@ AI meta tags (`ai:summary`, `ai:content_type`, `ai:author`, `ai:api`, `ai:agent_
|
|
|
161
224
|
| 0 AI meta tags | −18 |
|
|
162
225
|
| Only 1–2 AI meta tags | −12 |
|
|
163
226
|
| No `rel="alternate"` → llms.txt | −12 |
|
|
164
|
-
| No `rel="alternate"` →
|
|
227
|
+
| No `rel="alternate"` → the Agent Card | −8 |
|
|
165
228
|
| No `rel="me"` identity links | −8 |
|
|
166
229
|
| No Open Graph tags at all | −12 |
|
|
167
230
|
| OG required incomplete (`og:title`, `og:description`, `og:url`, `og:type`) | −8 |
|
|
@@ -170,21 +233,33 @@ AI meta tags (`ai:summary`, `ai:content_type`, `ai:author`, `ai:api`, `ai:agent_
|
|
|
170
233
|
| Twitter required incomplete (`twitter:card`, `twitter:title`, `twitter:description`) | −5 |
|
|
171
234
|
| Twitter recommended incomplete (`twitter:image`) | −2 |
|
|
172
235
|
|
|
173
|
-
### `
|
|
236
|
+
### `api-discovery` — 6%
|
|
237
|
+
|
|
238
|
+
Whether an agent can find, and read, a machine-readable API description. *Former id: `openapi`.*
|
|
174
239
|
|
|
175
|
-
`/.well-known/openapi.json
|
|
240
|
+
`/.well-known/openapi.json` is a folk convention — unregistered, and not prescribed by the OpenAPI specification, which recommends the file name `openapi.json` without a location. Discovery is probed in order of authority:
|
|
241
|
+
|
|
242
|
+
1. `/.well-known/api-catalog` *(RFC 9727, IANA-registered)* → its `service-desc` links
|
|
243
|
+
2. `Link: rel="service-desc"` on the homepage *(RFC 8631)*
|
|
244
|
+
3. `<link rel="service-desc">` in the HTML head
|
|
245
|
+
4. Conventional paths: `/.well-known/openapi.json`, `/openapi.json`, `/openapi.yaml`, `/.well-known/openapi.yaml`, `/api/openapi.json`, `/v1/openapi.json`, `/swagger.json`, `/api-docs`, `/asyncapi.json`, `/arazzo.json`
|
|
176
246
|
|
|
177
247
|
| Condition | Points |
|
|
178
248
|
| --- | --- |
|
|
179
|
-
|
|
|
249
|
+
| No description found by any mechanism | **hard fail → 0** |
|
|
180
250
|
| Invalid JSON | **→ 10** |
|
|
181
|
-
| Wrong Content-Type | −5 |
|
|
251
|
+
| Wrong Content-Type on a JSON document | −5 |
|
|
182
252
|
| No `openapi`/`swagger` version field | −20 |
|
|
183
253
|
| Swagger 2.x instead of OpenAPI 3.x | −10 |
|
|
184
254
|
| Missing `info.title` | −10 |
|
|
185
255
|
| Missing `info.description` | −5 |
|
|
186
256
|
| No `paths` documented | −15 |
|
|
187
257
|
| No `servers` | −5 |
|
|
258
|
+
| Found only by guessing a path (nothing links to it) | warn only, 0 |
|
|
259
|
+
| `operationId` coverage below 100% | informational, 0 in 3.x |
|
|
260
|
+
| API catalog present but empty, or entries missing `anchor` / `service-doc` | warn only, 0 |
|
|
261
|
+
|
|
262
|
+
YAML descriptions are recognised and reported, but only surface-validated: ax-audit ships no YAML parser, and the finding says so rather than pretending otherwise.
|
|
188
263
|
|
|
189
264
|
### `tls-https` — 5%
|
|
190
265
|
|
|
@@ -220,21 +295,13 @@ Located via robots.txt `Sitemap:` or `/sitemap.xml`. Limits: 50,000 URLs / 50 MB
|
|
|
220
295
|
| `<lastmod>` coverage < 50% | −5 |
|
|
221
296
|
| Newest `<lastmod>` older than 365 days | −5 |
|
|
222
297
|
|
|
223
|
-
### `well-known-ai` — 3%
|
|
224
|
-
|
|
225
|
-
Emerging AI discovery files. **Purely proportional** — no deductions:
|
|
226
|
-
|
|
227
|
-
```
|
|
228
|
-
score = round(present / 5 × 100)
|
|
229
|
-
```
|
|
230
|
-
|
|
231
|
-
over `/.well-known/ai.txt` (Spawning), `/.well-known/genai.txt`, `/ai-plugin.json`, `/agents.json`, `/.well-known/nlweb.json`. Files with invalid content produce warnings without counting as present.
|
|
232
|
-
|
|
233
298
|
---
|
|
234
299
|
|
|
235
|
-
##
|
|
300
|
+
## Checks on draft specifications (weight 0)
|
|
301
|
+
|
|
302
|
+
`ai-catalog`, `webmcp` and `commerce-discovery` run on every audit and report full findings, but never affect the score. Each rests on a specification that is still a draft and may be renamed; scoring a site against one would make the number less trustworthy, not more.
|
|
236
303
|
|
|
237
|
-
|
|
304
|
+
`well-known-ai` was removed in 4.0. Re-verification found three of its five scored files had no consumer: `/.well-known/nlweb.json` appears in no NLWeb release, `genai.txt` has no specification, and `/ai-plugin.json` described a product shut down in 2024. Its live probes moved into the checks that own them — TDMRep into `usage-policy`, the Web Bot Auth key directory and `AGENTS.md` into the reporting they belong to.
|
|
238
305
|
|
|
239
306
|
### `content-negotiation` — Markdown for Agents
|
|
240
307
|
|
|
@@ -250,6 +317,11 @@ Probes the homepage with `Accept: text/markdown` — the pattern served by Cloud
|
|
|
250
317
|
| Body is a relabeled HTML document | −25 |
|
|
251
318
|
| `Vary` does not include `Accept` | −15 |
|
|
252
319
|
| Markdown not smaller than HTML | warn only, 0 |
|
|
320
|
+
| Origin-reported token counts (`x-markdown-tokens` / `x-original-tokens`) | informational, 0 |
|
|
321
|
+
| No frontmatter, or frontmatter with no title / canonical URL / date | informational, 0 |
|
|
322
|
+
| User-agent negotiation or a `.md` suffix URL, when Accept negotiation fails | informational, 0 |
|
|
323
|
+
|
|
324
|
+
The probe sends the Accept header a real agent sends (`text/markdown, text/html;q=0.9, */*;q=0.1`). A bare `text/markdown` would pass against an implementation that fails every real request.
|
|
253
325
|
|
|
254
326
|
### `rsl` — Really Simple Licensing
|
|
255
327
|
|
|
@@ -270,23 +342,27 @@ Probes the homepage with `Accept: text/markdown` — the pattern served by Cloud
|
|
|
270
342
|
| Tokens outside the RSL 1.0 vocabulary (incl. pre-1.0 draft tokens) | −5 |
|
|
271
343
|
| Invalid `payment` type | −5 |
|
|
272
344
|
|
|
273
|
-
### `agent-access` —
|
|
345
|
+
### `agent-access` — blocking and cloaking detection
|
|
274
346
|
|
|
275
|
-
Probes the homepage with realistic
|
|
347
|
+
Probes the homepage with realistic user agents for the 10 core AI crawlers that actually issue requests (`Google-Extended` and `Applebot-Extended` are robots.txt control tokens, so probing with them tests nothing) and compares each response against the default-UA baseline. **Credit-ratio formula:**
|
|
276
348
|
|
|
277
349
|
```
|
|
278
|
-
score = round(credit /
|
|
350
|
+
score = round(credit / 10 × 100)
|
|
279
351
|
```
|
|
280
352
|
|
|
353
|
+
Responses are classified by *how* a request was turned away, because the remedies differ completely:
|
|
354
|
+
|
|
281
355
|
| Outcome per crawler | Credit |
|
|
282
356
|
| --- | --- |
|
|
283
|
-
|
|
|
284
|
-
|
|
|
285
|
-
|
|
|
286
|
-
|
|
|
357
|
+
| Same page as a regular client | 1 |
|
|
358
|
+
| Refused, consistent with an explicit robots.txt `Disallow` | 1 |
|
|
359
|
+
| Priced access (`402` + `crawler-price`) or an RSL licence challenge | 1 |
|
|
360
|
+
| JavaScript challenge (`cf-mitigated: challenge`, `x-vercel-mitigated`, AWS WAF's `202`), Web Bot Auth demand, rate limit, or a refusal from a bot-verifying CDN | 0.75, **inconclusive** |
|
|
361
|
+
| Different page than the baseline: less text, or a changed title / h1 / JSON-LD block count | 0.5 |
|
|
362
|
+
| Refused by a plain origin while robots.txt permits it | 0 |
|
|
287
363
|
| Baseline request itself fails | **hard fail → 0** |
|
|
288
364
|
|
|
289
|
-
|
|
365
|
+
The probe is unsigned and comes from the auditor's own network, so an edge that verifies crawlers by IP range or Web Bot Auth signature will reject it while admitting the genuine crawler. Those outcomes are reported as inconclusive with the exact header observed, never as "blocks AI crawlers". Confirm against WAF logs before changing a rule.
|
|
290
366
|
|
|
291
367
|
### `crawl-efficiency`
|
|
292
368
|
|
|
@@ -299,6 +375,161 @@ Caveat: WAFs using Web Bot Auth / IP verification may pass the real crawler whil
|
|
|
299
375
|
| Validator present but conditional request not answered with `304` | −15 |
|
|
300
376
|
| Page > 2 MB decompressed | −10 |
|
|
301
377
|
| Page > 500 KB decompressed | −5 |
|
|
378
|
+
| Content tokens, wire tokens and markup share (estimated at 4 chars per token) | informational, 0 |
|
|
379
|
+
| Response over 2s | informational, 0 |
|
|
380
|
+
|
|
381
|
+
---
|
|
382
|
+
|
|
383
|
+
### `ai-directives` — page-level AI controls
|
|
384
|
+
|
|
385
|
+
The controls Google and Microsoft document that they honor, read from robots meta tags and `X-Robots-Tag` (including the user-agent-scoped header form).
|
|
386
|
+
|
|
387
|
+
| Condition | Points |
|
|
388
|
+
| --- | --- |
|
|
389
|
+
| Homepage HTML unavailable | **hard fail → 0** |
|
|
390
|
+
| `noindex` or `none` | **hard fail → 0** — invisible to every search-grounded assistant |
|
|
391
|
+
| `nosnippet`, or `max-snippet:0` | −30 — excluded as a direct input to Google AI Overviews and AI Mode |
|
|
392
|
+
| `noarchive` | −30 — excluded from Microsoft Copilot grounding |
|
|
393
|
+
| `nocache` | −10 — Copilot may use only the URL, title and snippet |
|
|
394
|
+
| `data-nosnippet` wrapping `<main>`, `<article>` or `<body>` | −20 |
|
|
395
|
+
| `noimageindex` | −5 |
|
|
396
|
+
| `max-snippet:[n]` under 160 | warn only, 0 |
|
|
397
|
+
| `noai` / `noimageai` | reported, 0 — no major operator documents honoring them |
|
|
398
|
+
| robots.txt disallows `Google-Extended` with no snippet directive set | warn only, 0 |
|
|
399
|
+
|
|
400
|
+
That last row is the finding this check exists for. `Google-Extended` governs Gemini training and grounding in Gemini Apps and Vertex AI, **not** AI Overviews, which follow Googlebot and the snippet directives. A site that disallows it expecting to leave AI Overviews has opted out of the thing it probably did not mind.
|
|
401
|
+
|
|
402
|
+
### `usage-policy` — do your usage signals agree?
|
|
403
|
+
|
|
404
|
+
Normalises every machine-readable usage declaration onto three questions — may you train on it, ground an answer in it, index it — and reports where they disagree.
|
|
405
|
+
|
|
406
|
+
| Mechanism | Training | Grounding | Search |
|
|
407
|
+
| --- | --- | --- | --- |
|
|
408
|
+
| Content Signals (robots.txt or header) | `ai-train=yes\|no` | `ai-input=yes\|no` | `search=yes\|no` |
|
|
409
|
+
| IETF AIPREF (robots.txt or header) | `train-ai=y\|n` | *(no category yet)* | `search=y\|n` |
|
|
410
|
+
| RSL licence | `ai-train` | `ai-input` | `ai-index`, `search` |
|
|
411
|
+
| TDMRep (meta > header > well-known) | `tdm-reservation: 0\|1` | — | — |
|
|
412
|
+
| robots meta | `noai` | — | — |
|
|
413
|
+
|
|
414
|
+
| Condition | Points |
|
|
415
|
+
| --- | --- |
|
|
416
|
+
| No declaration of any kind | **→ 40** |
|
|
417
|
+
| Two mechanisms give opposite answers on one dimension | −25 per dimension |
|
|
418
|
+
| `Content-Usage` header outside the AIPREF vocabulary | warn only, 0 |
|
|
419
|
+
| A dimension no declaration covers | warn only, 0 |
|
|
420
|
+
|
|
421
|
+
Every report states that only robots.txt access rules are documented as honored by major AI operators. The rest are declarations whose weight is legal rather than technical.
|
|
422
|
+
|
|
423
|
+
### `http-hygiene` — status-code honesty
|
|
424
|
+
|
|
425
|
+
| Condition | Points |
|
|
426
|
+
| --- | --- |
|
|
427
|
+
| A nonexistent path returns 200 | −30 |
|
|
428
|
+
| A nonexistent path redirects | −20 |
|
|
429
|
+
| A 429 with no `Retry-After` | −20 |
|
|
430
|
+
| A 429 with `Retry-After` on the second request | −10 |
|
|
431
|
+
| `HEAD` refused (405/501) | −10 |
|
|
432
|
+
| Over one redirect hop to the homepage | −10 |
|
|
433
|
+
| No `Content-Type` header | −10 |
|
|
434
|
+
| No charset in the header or the document | −10 |
|
|
435
|
+
| `<html lang>` disagrees with `Content-Language` | −5 |
|
|
436
|
+
| A nonexistent path returns 403/401 | −5 |
|
|
437
|
+
| Empty 404 body | −5 |
|
|
438
|
+
| The 404 probe was challenged by bot management | warn only, 0 |
|
|
439
|
+
|
|
440
|
+
### `ai-catalog` — the index of everything callable
|
|
441
|
+
|
|
442
|
+
Discovery, in the order Lighthouse's `ard-schema` audit uses: robots.txt `Agentmap:`, `Link: rel="ai-catalog"`, `<link rel="ai-catalog">`, then `/.well-known/ai-catalog.json` and `/.well-known/ard.json`. Both specifications are drafts, so absence warns and scores nothing.
|
|
443
|
+
|
|
444
|
+
| Condition | Points |
|
|
445
|
+
| --- | --- |
|
|
446
|
+
| Catalog is not valid JSON | **→ 10** |
|
|
447
|
+
| An entry points at a document that cannot be fetched | −15 each |
|
|
448
|
+
| No entries | −20 |
|
|
449
|
+
| Entry missing identifier, type, or url/data | −10 |
|
|
450
|
+
| No `specVersion` / no `host` | −5 each |
|
|
451
|
+
| Entry served with a different media type than declared | −5 |
|
|
452
|
+
|
|
453
|
+
### `agent-skills` — installable procedures
|
|
454
|
+
|
|
455
|
+
Conditional: **n/a** unless the site has a developer-facing surface (documentation links, llms.txt, or an API description). Probes `/.well-known/agent-skills/index.json`, `/.well-known/skills/index.json`, then `/skill.md`.
|
|
456
|
+
|
|
457
|
+
| Condition | Points |
|
|
458
|
+
| --- | --- |
|
|
459
|
+
| Index is not valid JSON | **→ 10** |
|
|
460
|
+
| A sampled skill is unreachable, or its frontmatter name disagrees with the index | −10 per problem |
|
|
461
|
+
| Index lists no skills | −30 |
|
|
462
|
+
| A skill has no description | −15 |
|
|
463
|
+
| No entry carries a url | −15 |
|
|
464
|
+
| A skill name is outside `[a-z0-9-]{1,64}` | −10 |
|
|
465
|
+
| Malformed digest, unknown type, over-long description, or no `$schema` | −5 each |
|
|
466
|
+
| No skill declares a digest | −5 |
|
|
467
|
+
| A single `/skill.md` with no index | −20 |
|
|
468
|
+
|
|
469
|
+
### `webmcp` — forms as callable tools
|
|
470
|
+
|
|
471
|
+
Conditional: **n/a** on a page with no forms and no WebMCP code. Never asks for WebMCP — it is a Community Group draft in a Chrome origin trial.
|
|
472
|
+
|
|
473
|
+
| Condition | Points |
|
|
474
|
+
| --- | --- |
|
|
475
|
+
| `toolname` with no `tooldescription`, or the reverse | −30 |
|
|
476
|
+
| Parameters with no `toolparamdescription` | −15 |
|
|
477
|
+
| Tool name is not a usable identifier | −10 |
|
|
478
|
+
| Deprecated `navigator.modelContext` namespace | −10 |
|
|
479
|
+
| Forms present, none annotated | warn only, 0 |
|
|
480
|
+
|
|
481
|
+
### `agent-operability` — can a browser agent work this page?
|
|
482
|
+
|
|
483
|
+
Browser agents read the accessibility tree, not the pixels. A `<div onclick>` styled as a button has no role and no name, so it does not appear in the tree at all: the agent does not see a button it cannot press, it sees nothing.
|
|
484
|
+
|
|
485
|
+
| Condition | Points |
|
|
486
|
+
| --- | --- |
|
|
487
|
+
| Homepage HTML unavailable | **hard fail → 0** |
|
|
488
|
+
| Under 90% of buttons and links have an accessible name | −20 |
|
|
489
|
+
| Under 90% of form controls are labelled | −20 |
|
|
490
|
+
| Clickable elements that are not buttons or links (no `role` + `tabindex`) | −15 |
|
|
491
|
+
| A CAPTCHA or `<meta http-equiv="refresh">` on the entry page | −15 |
|
|
492
|
+
| Links with no `href` or a `javascript:` one | −10 |
|
|
493
|
+
| Tables with no `<th>` | −10 |
|
|
494
|
+
| Untitled iframes, `<time>` without `datetime`, heading-level skips, unsized media, no `<html lang>` | −5 each |
|
|
495
|
+
|
|
496
|
+
Names are read from visible text, `aria-label`, `aria-labelledby`, `title`, image `alt`, and SVG `<title>`. Labels from `<label for>`, a wrapping label, or ARIA.
|
|
497
|
+
|
|
498
|
+
Every run ends with a method note: this reads markup, not a rendered accessibility tree, so labels attached by script and roles computed at runtime are invisible to it. A low proportion is a prompt to check the real tree, not a count to act on blindly. Every finding is also a plain accessibility defect.
|
|
499
|
+
|
|
500
|
+
### `commerce-discovery` — Universal Commerce Protocol
|
|
501
|
+
|
|
502
|
+
Conditional: **n/a** unless the page shows storefront signals. A lone `Offer` is a price statement, not a catalog, so it counts only alongside a cart link.
|
|
503
|
+
|
|
504
|
+
| Condition | Points |
|
|
505
|
+
| --- | --- |
|
|
506
|
+
| Profile requires authentication | **hard fail → 0** |
|
|
507
|
+
| Profile is not valid JSON | **→ 10** |
|
|
508
|
+
| No `ucp` object | **→ 20** |
|
|
509
|
+
| No services declared | −25 |
|
|
510
|
+
| A declared schema URL cannot be fetched | −20 |
|
|
511
|
+
| No version | −20 |
|
|
512
|
+
| No payment handlers | −15 |
|
|
513
|
+
| Version is not a specification date | −10 |
|
|
514
|
+
| No signing keys | −10 |
|
|
515
|
+
| No schema URL declared | −10 |
|
|
516
|
+
| Service name not in reverse-DNS form, unnamed handler | −5 each |
|
|
517
|
+
|
|
518
|
+
The OpenAI and Stripe Agentic Commerce Protocol defines no manifest, and AP2 advertises through an A2A card extension, so neither is probed.
|
|
519
|
+
|
|
520
|
+
### `auth-discovery` — can an agent get credentials?
|
|
521
|
+
|
|
522
|
+
Conditional: **n/a** unless the site exposes an API description, API catalog, MCP server card or commerce profile. Follows the RFC 9728 chain from `WWW-Authenticate` or `/.well-known/oauth-protected-resource` to the authorization server's RFC 8414 or OpenID metadata.
|
|
523
|
+
|
|
524
|
+
| Condition | Points |
|
|
525
|
+
| --- | --- |
|
|
526
|
+
| Metadata names no authorization server | **→ 40** |
|
|
527
|
+
| Authorization server publishes no discovery metadata | −30 |
|
|
528
|
+
| Invalid issuer URL | −25 |
|
|
529
|
+
| Missing `issuer`, `authorization_endpoint` or `token_endpoint` | −15 each |
|
|
530
|
+
| No PKCE with `S256` | −15 |
|
|
531
|
+
| No dynamic registration and no Client ID Metadata Documents | −10 |
|
|
532
|
+
| No `resource` identifier | −10 |
|
|
302
533
|
|
|
303
534
|
---
|
|
304
535
|
|
|
@@ -312,6 +543,8 @@ overall = round( Σ (score_i / 100 × weight_i) / Σ weight_i × 100 )
|
|
|
312
543
|
|
|
313
544
|
When every selected check has weight 0 (e.g. `--checks rsl`), the overall falls back to a plain average of check scores.
|
|
314
545
|
|
|
546
|
+
Checks reporting `applicable: false` are excluded from both the numerator and the denominator. A check whose meta exists but produced no result — because it crashed — still counts at full weight, so a broken check cannot inflate a score by shrinking the denominator.
|
|
547
|
+
|
|
315
548
|
| Grade | Score | Exit code |
|
|
316
549
|
| --- | --- | --- |
|
|
317
550
|
| Excellent | 90–100 | 0 |
|
package/docs/cli.md
CHANGED
|
@@ -65,3 +65,39 @@ npx ax-audit $(cat urls.txt) --concurrency 4 --json > batch.json
|
|
|
65
65
|
# Show me only what is broken
|
|
66
66
|
npx ax-audit https://your-site.com --only-failures
|
|
67
67
|
```
|
|
68
|
+
|
|
69
|
+
## Agent-surface profiles
|
|
70
|
+
|
|
71
|
+
Protocol checks (`api-discovery`, `agent-card`, `mcp-discovery`, `agent-skills`, `auth-discovery`, `commerce-discovery`) apply only where the site has the corresponding surface. A blog has no API to describe, so scoring it zero would say something false; those checks report **n/a** and leave the denominator.
|
|
72
|
+
|
|
73
|
+
`--profile` overrides that, for auditing against what a site intends to build rather than what it already has.
|
|
74
|
+
|
|
75
|
+
```bash
|
|
76
|
+
ax-audit https://example.com --profile api # audit as though it offered an API
|
|
77
|
+
ax-audit https://example.com --profile all # every protocol check applies
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Values: `auto` (default, detect from the site), `api`, `mcp`, `agent`, `docs`, `commerce`, `all`.
|
|
81
|
+
|
|
82
|
+
## Focusing on one area
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
ax-audit https://example.com --category access # only access checks
|
|
86
|
+
ax-audit https://example.com --category content,discovery # two areas
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
`--category` narrows an explicit `--checks` selection rather than replacing it, and errors if the intersection is empty.
|
|
90
|
+
|
|
91
|
+
## Per-area CI gates
|
|
92
|
+
|
|
93
|
+
An overall score hides an area that is entirely broken: a site can score 80 while every access check fails, because the other four carry it.
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
ax-audit https://example.com --fail-on-category access:70,content:80
|
|
97
|
+
```
|
|
98
|
+
|
|
99
|
+
Each area is reported against its threshold on stderr, so the output survives `--output json` being piped elsewhere. An area with no applicable checks is reported as not evaluated rather than failed: a build must not break because a site has no commerce surface.
|
|
100
|
+
|
|
101
|
+
## Baselines across a scoring change
|
|
102
|
+
|
|
103
|
+
Baselines record which scoring model produced them. Comparing a baseline written by an older version shows the deltas but suspends regression gating, because a rescore is not something the site did. Re-save with `--save-baseline` to resume gating. Checks that changed applicability are excluded from regressions and improvements for the same reason.
|
package/docs/concepts.md
CHANGED
|
@@ -7,7 +7,7 @@
|
|
|
7
7
|
Agents are not browsers. Three differences drive every check:
|
|
8
8
|
|
|
9
9
|
1. **They mostly don't run JavaScript.** GPTBot, ClaudeBot, CCBot and most crawlers fetch raw HTML. A client-rendered SPA that returns an empty `<div id="root">` is, to them, a blank page. (`html-rendering`, `content-negotiation`)
|
|
10
|
-
2. **They look for declared structure, not visual layout.** An agent would rather read a `/llms.txt` summary or a JSON-LD graph than infer meaning from your CSS grid. (`llms-txt`, `structured-data`, `meta-tags`, `agent-
|
|
10
|
+
2. **They look for declared structure, not visual layout.** An agent would rather read a `/llms.txt` summary or a JSON-LD graph than infer meaning from your CSS grid. (`llms-txt`, `structured-data`, `meta-tags`, `agent-card`, `mcp-discovery`, `api-discovery`)
|
|
11
11
|
3. **Their access is a policy and economic question, not just a technical one.** Who may crawl, for what use, at what price, under what license — these now have machine-readable answers. (`robots-txt`, Content Signals, `rsl`, `agent-access`)
|
|
12
12
|
|
|
13
13
|
Bot traffic is projected to exceed human traffic by 2029. AX is the interface layer for that shift.
|
|
@@ -30,11 +30,15 @@ These answer: *can an agent find your content and actually read it?*
|
|
|
30
30
|
|
|
31
31
|
| Standard | What it is | Check |
|
|
32
32
|
| --- | --- | --- |
|
|
33
|
-
| **[A2A — Agent2Agent](https://a2a-protocol.org)** | An
|
|
34
|
-
| **[MCP — Model Context Protocol](https://modelcontextprotocol.io)** | A
|
|
35
|
-
| **[OpenAPI](https://www.openapis.org)** | The
|
|
36
|
-
| **
|
|
37
|
-
| **
|
|
33
|
+
| **[A2A — Agent2Agent](https://a2a-protocol.org)** | An Agent Card at `/.well-known/agent-card.json` advertising your identity and skills, so other agents can interoperate. The path moved there in v0.3.0 and is IANA-registered; v1.0 changed the card's shape. | `agent-card` |
|
|
34
|
+
| **[MCP — Model Context Protocol](https://modelcontextprotocol.io)** | A **server card** identifying your MCP server and its remote endpoints. Note it carries no tool list: tools come from a live `tools/list` call, and a static copy drifts. `/.well-known/mcp.json` was never part of the specification. | `mcp-discovery` |
|
|
35
|
+
| **[OpenAPI](https://www.openapis.org)** and **[RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html)** | The machine-readable API description, plus the registered `/.well-known/api-catalog` that points at it. Agents use both to call your endpoints without a human reading your docs. | `api-discovery` |
|
|
36
|
+
| **[Agent Skills](https://agentskills.io)** | SKILL.md documents an agent installs and follows — setup steps, argument shapes, mistakes to avoid. Answers "how do I do the thing this site is for", where llms.txt answers "what is here". | `agent-skills` |
|
|
37
|
+
| **[RFC 9728 / RFC 8414](https://www.rfc-editor.org/rfc/rfc9728.html)** | OAuth metadata. A human hitting a 401 reads your docs; an agent cannot, so the answer has to be in the response. | `auth-discovery` |
|
|
38
|
+
| **[UCP](https://developers.google.com/merchant/ucp/guides/ucp-profile)** | The Universal Commerce Protocol profile at `/.well-known/ucp` — the one agentic-commerce specification with published site-side discovery. | `commerce-discovery` |
|
|
39
|
+
| **[WebMCP](https://webmachinelearning.github.io/webmcp/)** | Declaring a form as a callable tool, so an agent invokes it rather than driving it pixel by pixel. A Community Group draft in a Chrome origin trial. | `webmcp` |
|
|
40
|
+
| **AI catalogs** | `ai-catalog.json` and `ard.json`, two competing drafts for one index listing everything above, so a client stops probing four conventions. | `ai-catalog` |
|
|
41
|
+
| **AI meta tags & discovery links** | `ai:*` meta tags and the `rel` relations — `alternate`, `describedby`, `service-desc`, `api-catalog` — that point agents at your files instead of making them guess paths. | `meta-tags`, `http-headers` |
|
|
38
42
|
|
|
39
43
|
These answer: *once an agent arrives, can it understand what you offer and act on it?*
|
|
40
44
|
|
|
@@ -44,14 +48,20 @@ This is the newest and fastest-moving family — the response to "AI scraped my
|
|
|
44
48
|
|
|
45
49
|
| Standard | What it is | Check |
|
|
46
50
|
| --- | --- | --- |
|
|
47
|
-
| **[Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)** | The original robots.txt — *who* may crawl *what*. ax-audit
|
|
48
|
-
| **[Content Signals](https://contentsignals.org)** | A robots.txt extension (Cloudflare, CC0) declaring *how* content may be used after access: `search`, `ai-input`, `ai-train
|
|
49
|
-
| **[
|
|
50
|
-
| **
|
|
51
|
+
| **[Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)** | The original robots.txt — *who* may crawl *what*. ax-audit knows 57 AI clients, grouped by what they do with a page, because blocking a search crawler costs citations while blocking a training crawler is a policy choice. | `robots-txt` |
|
|
52
|
+
| **[Content Signals](https://contentsignals.org)** | A robots.txt extension (Cloudflare, CC0) declaring *how* content may be used after access: `search`, `ai-input`, `ai-train`, plus the `use` field added in 2026. Served by default on millions of Cloudflare domains. | `robots-txt`, `usage-policy` |
|
|
53
|
+
| **[IETF AIPREF](https://datatracker.ietf.org/wg/aipref/documents/)** | The standards-track answer to the same question: a `Content-Usage` directive or header over `train-ai` and `search`. Still pre-last-call. Note the token inversion against Content Signals — `train-ai` here, `ai-train` there. | `robots-txt`, `usage-policy` |
|
|
54
|
+
| **[RSL — Really Simple Licensing](https://rslstandard.org)** | A full machine-readable licensing layer: permits/prohibits vocabularies, payment models, and an Open Licensing Protocol for negotiating access. | `rsl`, `usage-policy` |
|
|
55
|
+
| **[TDMRep](https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240510/)** | A text-and-data-mining reservation, named in the EU GPAI Code of Practice. Its weight is legal rather than technical. | `usage-policy` |
|
|
56
|
+
| **Page-level AI directives** | `nosnippet` and `max-snippet` for Google AI Overviews, `noarchive` and `nocache` for Copilot grounding. Unlike everything above, the vendors document that they honor these. | `ai-directives` |
|
|
57
|
+
| **Cloaking and blocking integrity** | Not a standard but a failure mode: robots.txt allows GPTBot and the firewall returns 403. | `agent-access` |
|
|
58
|
+
| **Signal consistency** | Also not a standard: whether your Content Signals, AIPREF, RSL and TDMRep declarations agree with each other. Five documents maintained by hand do not stay in sync, and the terms that apply then depend on which file a crawler read. | `usage-policy` |
|
|
51
59
|
|
|
52
60
|
These answer: *have you expressed your access and usage policy in a form agents can honor — and does your infrastructure actually match it?*
|
|
53
61
|
|
|
54
|
-
The progression is one of increasing expressiveness: robots.txt says **who
|
|
62
|
+
The progression is one of increasing expressiveness: robots.txt says **who and where**, Content Signals and AIPREF add **how it may be used**, RSL adds **under what licence and price**.
|
|
63
|
+
|
|
64
|
+
One caveat worth stating plainly: only robots.txt access rules are documented as honored by Google, OpenAI, Anthropic and Microsoft. Content Signals, AIPREF, RSL and TDMRep are declarations. Their weight is legal rather than technical, and ax-audit says so on every run rather than implying a crawler will obey.
|
|
55
65
|
|
|
56
66
|
### 4. Transport, efficiency & hygiene
|
|
57
67
|
|
|
@@ -78,10 +88,14 @@ A fully AX-ready site tells a coherent story across all four:
|
|
|
78
88
|
|
|
79
89
|
> "Here's my content in a form you can read **(family 1)**, here's the interface to interact with me **(family 2)**, here's exactly who may use it and how, for what license **(family 3)**, over a fast and trustworthy connection **(family 4)**."
|
|
80
90
|
|
|
81
|
-
ax-audit's weighting
|
|
91
|
+
ax-audit's weighting follows what actually stops an agent. **Content** carries the most weight, led by `html-rendering`, because a page whose content only exists after hydration is invisible to the crawlers that do not run JavaScript, no matter how many discovery files it publishes. **Access** is next, because a firewall rule or a `nosnippet` directive silently undoes everything else.
|
|
92
|
+
|
|
93
|
+
`llms-txt` sits at 5 rather than the 11 it carried in 3.x. Adoption studies find most published files are never fetched by an AI search crawler, and Google has stated that Search ignores them; the vendors that do read them are coding agents like Claude Code and Cursor. It is worth publishing, as developer tooling, and it is not worth twice what having content is worth.
|
|
94
|
+
|
|
95
|
+
Checks resting on draft specifications — `ai-catalog`, `webmcp`, `commerce-discovery` — are reported but never scored. A specification that may be renamed next quarter should not move your number.
|
|
82
96
|
|
|
83
97
|
## See also
|
|
84
98
|
|
|
85
99
|
- [getting-started.md](./getting-started.md) — run your first audit
|
|
86
100
|
- [checks.md](./checks.md) — exact scoring per standard
|
|
87
|
-
- The [remediation guides](https://
|
|
101
|
+
- The [remediation guides](https://axrush.com/guides) — how to implement each one
|