ax-audit 3.1.0 → 4.0.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (208) hide show
  1. package/CHANGELOG.md +198 -0
  2. package/LICENSE +1 -1
  3. package/README.md +89 -231
  4. package/dist/baseline.d.ts +2 -0
  5. package/dist/baseline.d.ts.map +1 -1
  6. package/dist/baseline.js +42 -4
  7. package/dist/baseline.js.map +1 -1
  8. package/dist/check-ids.d.ts +19 -0
  9. package/dist/check-ids.d.ts.map +1 -0
  10. package/dist/check-ids.js +53 -0
  11. package/dist/check-ids.js.map +1 -0
  12. package/dist/checks/agent-access.d.ts +33 -0
  13. package/dist/checks/agent-access.d.ts.map +1 -0
  14. package/dist/checks/agent-access.js +256 -0
  15. package/dist/checks/agent-access.js.map +1 -0
  16. package/dist/checks/agent-card.d.ts +37 -0
  17. package/dist/checks/agent-card.d.ts.map +1 -0
  18. package/dist/checks/agent-card.js +352 -0
  19. package/dist/checks/agent-card.js.map +1 -0
  20. package/dist/checks/agent-operability.d.ts +66 -0
  21. package/dist/checks/agent-operability.d.ts.map +1 -0
  22. package/dist/checks/agent-operability.js +383 -0
  23. package/dist/checks/agent-operability.js.map +1 -0
  24. package/dist/checks/agent-skills.d.ts +24 -0
  25. package/dist/checks/agent-skills.d.ts.map +1 -0
  26. package/dist/checks/agent-skills.js +316 -0
  27. package/dist/checks/agent-skills.js.map +1 -0
  28. package/dist/checks/ai-catalog.d.ts +28 -0
  29. package/dist/checks/ai-catalog.d.ts.map +1 -0
  30. package/dist/checks/ai-catalog.js +254 -0
  31. package/dist/checks/ai-catalog.js.map +1 -0
  32. package/dist/checks/ai-directives.d.ts +57 -0
  33. package/dist/checks/ai-directives.d.ts.map +1 -0
  34. package/dist/checks/ai-directives.js +263 -0
  35. package/dist/checks/ai-directives.js.map +1 -0
  36. package/dist/checks/api-discovery.d.ts +26 -0
  37. package/dist/checks/api-discovery.d.ts.map +1 -0
  38. package/dist/checks/api-discovery.js +432 -0
  39. package/dist/checks/api-discovery.js.map +1 -0
  40. package/dist/checks/auth-discovery.d.ts +28 -0
  41. package/dist/checks/auth-discovery.d.ts.map +1 -0
  42. package/dist/checks/auth-discovery.js +213 -0
  43. package/dist/checks/auth-discovery.js.map +1 -0
  44. package/dist/checks/commerce-discovery.d.ts +40 -0
  45. package/dist/checks/commerce-discovery.d.ts.map +1 -0
  46. package/dist/checks/commerce-discovery.js +295 -0
  47. package/dist/checks/commerce-discovery.js.map +1 -0
  48. package/dist/checks/content-negotiation.d.ts.map +1 -1
  49. package/dist/checks/content-negotiation.js +135 -20
  50. package/dist/checks/content-negotiation.js.map +1 -1
  51. package/dist/checks/crawl-efficiency.d.ts +16 -0
  52. package/dist/checks/crawl-efficiency.d.ts.map +1 -0
  53. package/dist/checks/crawl-efficiency.js +186 -0
  54. package/dist/checks/crawl-efficiency.js.map +1 -0
  55. package/dist/checks/frontmatter.d.ts +34 -0
  56. package/dist/checks/frontmatter.d.ts.map +1 -0
  57. package/dist/checks/frontmatter.js +100 -0
  58. package/dist/checks/frontmatter.js.map +1 -0
  59. package/dist/checks/html-rendering.d.ts.map +1 -1
  60. package/dist/checks/html-rendering.js +0 -1
  61. package/dist/checks/html-rendering.js.map +1 -1
  62. package/dist/checks/html-utils.d.ts +10 -0
  63. package/dist/checks/html-utils.d.ts.map +1 -1
  64. package/dist/checks/html-utils.js +19 -0
  65. package/dist/checks/html-utils.js.map +1 -1
  66. package/dist/checks/http-headers.d.ts.map +1 -1
  67. package/dist/checks/http-headers.js +82 -10
  68. package/dist/checks/http-headers.js.map +1 -1
  69. package/dist/checks/http-hygiene.d.ts +26 -0
  70. package/dist/checks/http-hygiene.d.ts.map +1 -0
  71. package/dist/checks/http-hygiene.js +257 -0
  72. package/dist/checks/http-hygiene.js.map +1 -0
  73. package/dist/checks/index.d.ts.map +1 -1
  74. package/dist/checks/index.js +30 -8
  75. package/dist/checks/index.js.map +1 -1
  76. package/dist/checks/llms-txt.d.ts +15 -0
  77. package/dist/checks/llms-txt.d.ts.map +1 -1
  78. package/dist/checks/llms-txt.js +162 -2
  79. package/dist/checks/llms-txt.js.map +1 -1
  80. package/dist/checks/mcp-discovery.d.ts +30 -0
  81. package/dist/checks/mcp-discovery.d.ts.map +1 -0
  82. package/dist/checks/mcp-discovery.js +523 -0
  83. package/dist/checks/mcp-discovery.js.map +1 -0
  84. package/dist/checks/meta-tags.d.ts.map +1 -1
  85. package/dist/checks/meta-tags.js +6 -5
  86. package/dist/checks/meta-tags.js.map +1 -1
  87. package/dist/checks/robots-parser.d.ts +110 -0
  88. package/dist/checks/robots-parser.d.ts.map +1 -0
  89. package/dist/checks/robots-parser.js +277 -0
  90. package/dist/checks/robots-parser.js.map +1 -0
  91. package/dist/checks/robots-txt.d.ts +2 -0
  92. package/dist/checks/robots-txt.d.ts.map +1 -1
  93. package/dist/checks/robots-txt.js +252 -45
  94. package/dist/checks/robots-txt.js.map +1 -1
  95. package/dist/checks/{mcp.d.ts → rsl.d.ts} +1 -1
  96. package/dist/checks/rsl.d.ts.map +1 -0
  97. package/dist/checks/rsl.js +242 -0
  98. package/dist/checks/rsl.js.map +1 -0
  99. package/dist/checks/security-txt.d.ts.map +1 -1
  100. package/dist/checks/security-txt.js +0 -1
  101. package/dist/checks/security-txt.js.map +1 -1
  102. package/dist/checks/seo-basics.d.ts.map +1 -1
  103. package/dist/checks/seo-basics.js +0 -1
  104. package/dist/checks/seo-basics.js.map +1 -1
  105. package/dist/checks/sitemap.d.ts.map +1 -1
  106. package/dist/checks/sitemap.js +0 -1
  107. package/dist/checks/sitemap.js.map +1 -1
  108. package/dist/checks/structured-data.d.ts.map +1 -1
  109. package/dist/checks/structured-data.js +215 -4
  110. package/dist/checks/structured-data.js.map +1 -1
  111. package/dist/checks/structured-fields.d.ts +46 -0
  112. package/dist/checks/structured-fields.d.ts.map +1 -0
  113. package/dist/checks/structured-fields.js +112 -0
  114. package/dist/checks/structured-fields.js.map +1 -0
  115. package/dist/checks/surface.d.ts +59 -0
  116. package/dist/checks/surface.d.ts.map +1 -0
  117. package/dist/checks/surface.js +106 -0
  118. package/dist/checks/surface.js.map +1 -0
  119. package/dist/checks/tls-https.d.ts.map +1 -1
  120. package/dist/checks/tls-https.js +0 -1
  121. package/dist/checks/tls-https.js.map +1 -1
  122. package/dist/checks/usage-policy.d.ts +53 -0
  123. package/dist/checks/usage-policy.d.ts.map +1 -0
  124. package/dist/checks/usage-policy.js +339 -0
  125. package/dist/checks/usage-policy.js.map +1 -0
  126. package/dist/checks/utils.d.ts +25 -1
  127. package/dist/checks/utils.d.ts.map +1 -1
  128. package/dist/checks/utils.js +33 -1
  129. package/dist/checks/utils.js.map +1 -1
  130. package/dist/checks/waf.d.ts +75 -0
  131. package/dist/checks/waf.d.ts.map +1 -0
  132. package/dist/checks/waf.js +203 -0
  133. package/dist/checks/waf.js.map +1 -0
  134. package/dist/checks/webmcp.d.ts +55 -0
  135. package/dist/checks/webmcp.d.ts.map +1 -0
  136. package/dist/checks/webmcp.js +209 -0
  137. package/dist/checks/webmcp.js.map +1 -0
  138. package/dist/checks/well-known.d.ts +38 -0
  139. package/dist/checks/well-known.d.ts.map +1 -0
  140. package/dist/checks/well-known.js +202 -0
  141. package/dist/checks/well-known.js.map +1 -0
  142. package/dist/cli.d.ts +14 -0
  143. package/dist/cli.d.ts.map +1 -1
  144. package/dist/cli.js +144 -6
  145. package/dist/cli.js.map +1 -1
  146. package/dist/constants.d.ts +212 -14
  147. package/dist/constants.d.ts.map +1 -1
  148. package/dist/constants.js +625 -59
  149. package/dist/constants.js.map +1 -1
  150. package/dist/fetcher.d.ts +5 -1
  151. package/dist/fetcher.d.ts.map +1 -1
  152. package/dist/fetcher.js +62 -27
  153. package/dist/fetcher.js.map +1 -1
  154. package/dist/guide-urls.js +1 -1
  155. package/dist/guide-urls.js.map +1 -1
  156. package/dist/index.d.ts +3 -1
  157. package/dist/index.d.ts.map +1 -1
  158. package/dist/index.js +2 -0
  159. package/dist/index.js.map +1 -1
  160. package/dist/orchestrator.d.ts +2 -2
  161. package/dist/orchestrator.d.ts.map +1 -1
  162. package/dist/orchestrator.js +18 -7
  163. package/dist/orchestrator.js.map +1 -1
  164. package/dist/reporter/html.d.ts +9 -0
  165. package/dist/reporter/html.d.ts.map +1 -1
  166. package/dist/reporter/html.js +49 -11
  167. package/dist/reporter/html.js.map +1 -1
  168. package/dist/reporter/index.d.ts.map +1 -1
  169. package/dist/reporter/index.js +7 -0
  170. package/dist/reporter/index.js.map +1 -1
  171. package/dist/reporter/markdown.d.ts +8 -0
  172. package/dist/reporter/markdown.d.ts.map +1 -0
  173. package/dist/reporter/markdown.js +106 -0
  174. package/dist/reporter/markdown.js.map +1 -0
  175. package/dist/reporter/terminal.d.ts.map +1 -1
  176. package/dist/reporter/terminal.js +36 -1
  177. package/dist/reporter/terminal.js.map +1 -1
  178. package/dist/scorer.d.ts +10 -0
  179. package/dist/scorer.d.ts.map +1 -1
  180. package/dist/scorer.js +22 -5
  181. package/dist/scorer.js.map +1 -1
  182. package/dist/types.d.ts +96 -3
  183. package/dist/types.d.ts.map +1 -1
  184. package/docs/api.md +200 -0
  185. package/docs/architecture.md +104 -0
  186. package/docs/checks.md +555 -0
  187. package/docs/ci.md +89 -0
  188. package/docs/cli.md +103 -0
  189. package/docs/concepts.md +101 -0
  190. package/docs/faq.md +89 -0
  191. package/docs/getting-started.md +108 -0
  192. package/docs/roadmap.md +367 -0
  193. package/package.json +14 -5
  194. package/dist/checks/agent-json.d.ts +0 -14
  195. package/dist/checks/agent-json.d.ts.map +0 -1
  196. package/dist/checks/agent-json.js +0 -167
  197. package/dist/checks/agent-json.js.map +0 -1
  198. package/dist/checks/mcp.d.ts.map +0 -1
  199. package/dist/checks/mcp.js +0 -162
  200. package/dist/checks/mcp.js.map +0 -1
  201. package/dist/checks/openapi.d.ts +0 -4
  202. package/dist/checks/openapi.d.ts.map +0 -1
  203. package/dist/checks/openapi.js +0 -121
  204. package/dist/checks/openapi.js.map +0 -1
  205. package/dist/checks/well-known-ai.d.ts +0 -17
  206. package/dist/checks/well-known-ai.d.ts.map +0 -1
  207. package/dist/checks/well-known-ai.js +0 -123
  208. package/dist/checks/well-known-ai.js.map +0 -1
package/docs/cli.md ADDED
@@ -0,0 +1,103 @@
1
+ # CLI Reference
2
+
3
+ ```bash
4
+ ax-audit <urls...> [options]
5
+ ```
6
+
7
+ One or more fully qualified URLs (scheme required). A single URL produces a full report; multiple URLs run in batch mode with a summary table.
8
+
9
+ ## Options
10
+
11
+ | Flag | Default | Description |
12
+ | --- | --- | --- |
13
+ | `--output <format>` | `terminal` | Output format: `terminal`, `json`, `html`, `markdown`. Invalid values error out. |
14
+ | `--json` | — | Shorthand for `--output json`. |
15
+ | `--checks <list>` | all | Comma-separated check IDs to run (see [checks.md](./checks.md)). Unknown IDs error with the list of valid ones. |
16
+ | `--timeout <ms>` | `10000` | Per-request timeout in milliseconds. |
17
+ | `--retries <n>` | `2` | Retry attempts for transient fetch failures (network errors, timeouts, 408/425/429/5xx) with exponential backoff from 250ms. `0` disables retries. |
18
+ | `--concurrency <n>` | `1` | Batch mode only: maximum URLs audited in parallel. Output order always matches input order. |
19
+ | `--verbose` | — | Log every HTTP request, cache hit, retry, and per-check score to stderr. |
20
+ | `--only-failures` | — | Hide passing findings; checks with only passes are omitted entirely. |
21
+ | `--save-baseline <path>` | — | Save this audit as a baseline JSON file. |
22
+ | `--baseline <path>` | — | Compare against a saved baseline; shows per-check deltas (▲/▼). Single-URL mode only. |
23
+ | `--fail-on-regression <points>` | — | Exit 1 if any check regresses more than N points vs the baseline. Requires `--baseline`. |
24
+ | `-v, --version` | — | Print version. |
25
+
26
+ ## Output formats
27
+
28
+ - **terminal** — colored report with score bar, per-check sections, and PASS/WARN/FAIL findings.
29
+ - **json** — the full `AuditReport` (plus `baselineDiff` when `--baseline` is used). Stable shape for CI pipelines.
30
+ - **html** — self-contained page (score gauge, dark/light mode, collapsible sections). Pipe to a file: `ax-audit <url> --output html > report.html`.
31
+ - **markdown** — summary table + per-check findings with status emoji. Built for CI logs and PR comments: `ax-audit <url> --output markdown > report.md`.
32
+
33
+ ## Exit codes
34
+
35
+ | Code | Meaning |
36
+ | --- | --- |
37
+ | `0` | Score ≥ 70 (single), or all URLs ≥ 70 (batch), and no regression beyond the `--fail-on-regression` threshold. |
38
+ | `1` | Score < 70, any batch URL < 70, invalid arguments, or regression beyond threshold. |
39
+ | `2` | Fatal error (network failure on the audit itself, unreadable baseline file). |
40
+
41
+ ## Baseline workflow
42
+
43
+ ```bash
44
+ # First run — record the baseline
45
+ ax-audit https://your-site.com --save-baseline .ax-baseline.json
46
+
47
+ # Subsequent runs — compare and gate
48
+ ax-audit https://your-site.com --baseline .ax-baseline.json --fail-on-regression 5
49
+ ```
50
+
51
+ The baseline stores the overall score and per-check scores. Checks added after the baseline was saved appear as new (no delta); removed checks are ignored.
52
+
53
+ ## Examples
54
+
55
+ ```bash
56
+ # Quick audit
57
+ npx ax-audit https://your-site.com
58
+
59
+ # Only the AI-licensing surface
60
+ npx ax-audit https://your-site.com --checks robots-txt,rsl,content-negotiation
61
+
62
+ # Batch, 4 at a time, machine-readable
63
+ npx ax-audit $(cat urls.txt) --concurrency 4 --json > batch.json
64
+
65
+ # Show me only what is broken
66
+ npx ax-audit https://your-site.com --only-failures
67
+ ```
68
+
69
+ ## Agent-surface profiles
70
+
71
+ Protocol checks (`api-discovery`, `agent-card`, `mcp-discovery`, `agent-skills`, `auth-discovery`, `commerce-discovery`) apply only where the site has the corresponding surface. A blog has no API to describe, so scoring it zero would say something false; those checks report **n/a** and leave the denominator.
72
+
73
+ `--profile` overrides that, for auditing against what a site intends to build rather than what it already has.
74
+
75
+ ```bash
76
+ ax-audit https://example.com --profile api # audit as though it offered an API
77
+ ax-audit https://example.com --profile all # every protocol check applies
78
+ ```
79
+
80
+ Values: `auto` (default, detect from the site), `api`, `mcp`, `agent`, `docs`, `commerce`, `all`.
81
+
82
+ ## Focusing on one area
83
+
84
+ ```bash
85
+ ax-audit https://example.com --category access # only access checks
86
+ ax-audit https://example.com --category content,discovery # two areas
87
+ ```
88
+
89
+ `--category` narrows an explicit `--checks` selection rather than replacing it, and errors if the intersection is empty.
90
+
91
+ ## Per-area CI gates
92
+
93
+ An overall score hides an area that is entirely broken: a site can score 80 while every access check fails, because the other four carry it.
94
+
95
+ ```bash
96
+ ax-audit https://example.com --fail-on-category access:70,content:80
97
+ ```
98
+
99
+ Each area is reported against its threshold on stderr, so the output survives `--output json` being piped elsewhere. An area with no applicable checks is reported as not evaluated rather than failed: a build must not break because a site has no commerce surface.
100
+
101
+ ## Baselines across a scoring change
102
+
103
+ Baselines record which scoring model produced them. Comparing a baseline written by an older version shows the deltas but suspends regression gating, because a rescore is not something the site did. Re-save with `--save-baseline` to resume gating. Checks that changed applicability are excluded from regressions and improvements for the same reason.
@@ -0,0 +1,101 @@
1
+ # Concepts: the AX standards landscape
2
+
3
+ "AI Agent Experience" (AX) is the sum of the conventions a site uses to be discovered, read, governed, and transacted with by autonomous AI agents and crawlers — the way "web accessibility" is the sum of conventions for assistive technology. This page maps the standards ax-audit checks against, why each exists, and how they relate. It's the conceptual companion to the mechanical detail in [checks.md](./checks.md).
4
+
5
+ ## Why AX is its own discipline
6
+
7
+ Agents are not browsers. Three differences drive every check:
8
+
9
+ 1. **They mostly don't run JavaScript.** GPTBot, ClaudeBot, CCBot and most crawlers fetch raw HTML. A client-rendered SPA that returns an empty `<div id="root">` is, to them, a blank page. (`html-rendering`, `content-negotiation`)
10
+ 2. **They look for declared structure, not visual layout.** An agent would rather read a `/llms.txt` summary or a JSON-LD graph than infer meaning from your CSS grid. (`llms-txt`, `structured-data`, `meta-tags`, `agent-card`, `mcp-discovery`, `api-discovery`)
11
+ 3. **Their access is a policy and economic question, not just a technical one.** Who may crawl, for what use, at what price, under what license — these now have machine-readable answers. (`robots-txt`, Content Signals, `rsl`, `agent-access`)
12
+
13
+ Bot traffic is projected to exceed human traffic by 2029. AX is the interface layer for that shift.
14
+
15
+ ## The four families of standards
16
+
17
+ ### 1. Content discovery & readability
18
+
19
+ | Standard | What it is | Check |
20
+ | --- | --- | --- |
21
+ | **[llms.txt](https://llmstxt.org)** | A Markdown file at your root summarizing your site for LLMs, with curated links. The "sitemap for AI." | `llms-txt` |
22
+ | **Server-side rendering** | Delivering real content in the HTML response, not assembling it client-side. | `html-rendering` |
23
+ | **[Markdown for Agents](https://developers.cloudflare.com/fundamentals/reference/markdown-for-agents/)** | Content negotiation: serve clean Markdown when a client sends `Accept: text/markdown`. ~80% fewer tokens than HTML. | `content-negotiation` |
24
+ | **schema.org / JSON-LD** | Structured data describing entities (Person, Organization, Product) in a graph agents can parse. | `structured-data` |
25
+ | **Sitemaps** | The classic XML index, still how crawlers enumerate your URLs. | `sitemap` |
26
+
27
+ These answer: *can an agent find your content and actually read it?*
28
+
29
+ ### 2. Agent interaction surface
30
+
31
+ | Standard | What it is | Check |
32
+ | --- | --- | --- |
33
+ | **[A2A — Agent2Agent](https://a2a-protocol.org)** | An Agent Card at `/.well-known/agent-card.json` advertising your identity and skills, so other agents can interoperate. The path moved there in v0.3.0 and is IANA-registered; v1.0 changed the card's shape. | `agent-card` |
34
+ | **[MCP — Model Context Protocol](https://modelcontextprotocol.io)** | A **server card** identifying your MCP server and its remote endpoints. Note it carries no tool list: tools come from a live `tools/list` call, and a static copy drifts. `/.well-known/mcp.json` was never part of the specification. | `mcp-discovery` |
35
+ | **[OpenAPI](https://www.openapis.org)** and **[RFC 9727](https://www.rfc-editor.org/rfc/rfc9727.html)** | The machine-readable API description, plus the registered `/.well-known/api-catalog` that points at it. Agents use both to call your endpoints without a human reading your docs. | `api-discovery` |
36
+ | **[Agent Skills](https://agentskills.io)** | SKILL.md documents an agent installs and follows — setup steps, argument shapes, mistakes to avoid. Answers "how do I do the thing this site is for", where llms.txt answers "what is here". | `agent-skills` |
37
+ | **[RFC 9728 / RFC 8414](https://www.rfc-editor.org/rfc/rfc9728.html)** | OAuth metadata. A human hitting a 401 reads your docs; an agent cannot, so the answer has to be in the response. | `auth-discovery` |
38
+ | **[UCP](https://developers.google.com/merchant/ucp/guides/ucp-profile)** | The Universal Commerce Protocol profile at `/.well-known/ucp` — the one agentic-commerce specification with published site-side discovery. | `commerce-discovery` |
39
+ | **[WebMCP](https://webmachinelearning.github.io/webmcp/)** | Declaring a form as a callable tool, so an agent invokes it rather than driving it pixel by pixel. A Community Group draft in a Chrome origin trial. | `webmcp` |
40
+ | **AI catalogs** | `ai-catalog.json` and `ard.json`, two competing drafts for one index listing everything above, so a client stops probing four conventions. | `ai-catalog` |
41
+ | **AI meta tags & discovery links** | `ai:*` meta tags and the `rel` relations — `alternate`, `describedby`, `service-desc`, `api-catalog` — that point agents at your files instead of making them guess paths. | `meta-tags`, `http-headers` |
42
+
43
+ These answer: *once an agent arrives, can it understand what you offer and act on it?*
44
+
45
+ ### 3. Access governance & licensing
46
+
47
+ This is the newest and fastest-moving family — the response to "AI scraped my content and now competes with me."
48
+
49
+ | Standard | What it is | Check |
50
+ | --- | --- | --- |
51
+ | **[Robots Exclusion Protocol](https://www.rfc-editor.org/rfc/rfc9309.html)** | The original robots.txt — *who* may crawl *what*. ax-audit knows 57 AI clients, grouped by what they do with a page, because blocking a search crawler costs citations while blocking a training crawler is a policy choice. | `robots-txt` |
52
+ | **[Content Signals](https://contentsignals.org)** | A robots.txt extension (Cloudflare, CC0) declaring *how* content may be used after access: `search`, `ai-input`, `ai-train`, plus the `use` field added in 2026. Served by default on millions of Cloudflare domains. | `robots-txt`, `usage-policy` |
53
+ | **[IETF AIPREF](https://datatracker.ietf.org/wg/aipref/documents/)** | The standards-track answer to the same question: a `Content-Usage` directive or header over `train-ai` and `search`. Still pre-last-call. Note the token inversion against Content Signals — `train-ai` here, `ai-train` there. | `robots-txt`, `usage-policy` |
54
+ | **[RSL — Really Simple Licensing](https://rslstandard.org)** | A full machine-readable licensing layer: permits/prohibits vocabularies, payment models, and an Open Licensing Protocol for negotiating access. | `rsl`, `usage-policy` |
55
+ | **[TDMRep](https://www.w3.org/community/reports/tdmrep/CG-FINAL-tdmrep-20240510/)** | A text-and-data-mining reservation, named in the EU GPAI Code of Practice. Its weight is legal rather than technical. | `usage-policy` |
56
+ | **Page-level AI directives** | `nosnippet` and `max-snippet` for Google AI Overviews, `noarchive` and `nocache` for Copilot grounding. Unlike everything above, the vendors document that they honor these. | `ai-directives` |
57
+ | **Cloaking and blocking integrity** | Not a standard but a failure mode: robots.txt allows GPTBot and the firewall returns 403. | `agent-access` |
58
+ | **Signal consistency** | Also not a standard: whether your Content Signals, AIPREF, RSL and TDMRep declarations agree with each other. Five documents maintained by hand do not stay in sync, and the terms that apply then depend on which file a crawler read. | `usage-policy` |
59
+
60
+ These answer: *have you expressed your access and usage policy in a form agents can honor — and does your infrastructure actually match it?*
61
+
62
+ The progression is one of increasing expressiveness: robots.txt says **who and where**, Content Signals and AIPREF add **how it may be used**, RSL adds **under what licence and price**.
63
+
64
+ One caveat worth stating plainly: only robots.txt access rules are documented as honored by Google, OpenAI, Anthropic and Microsoft. Content Signals, AIPREF, RSL and TDMRep are declarations. Their weight is legal rather than technical, and ax-audit says so on every run rather than implying a crawler will obey.
65
+
66
+ ### 4. Transport, efficiency & hygiene
67
+
68
+ | Standard | What it is | Check |
69
+ | --- | --- | --- |
70
+ | **TLS / HSTS** | HTTPS everywhere; many agents refuse plaintext origins. | `tls-https` |
71
+ | **HTTP security & discovery headers** | Security headers plus `Link` headers advertising your AI files. | `http-headers` |
72
+ | **Compression & conditional GET** | Brotli/gzip and `ETag`/`304` — crawl cost matters when bots dominate traffic. | `crawl-efficiency` |
73
+ | **[RFC 9116 security.txt](https://www.rfc-editor.org/rfc/rfc9116)** | A machine-readable security contact. | `security-txt` |
74
+ | **SEO basics** | Title, description, canonical, lang, hreflang — agents use the same head-tag fundamentals search engines do. | `seo-basics` |
75
+
76
+ These answer: *is the connection trustworthy, cheap, and well-formed?*
77
+
78
+ ## On the horizon (not yet scored)
79
+
80
+ Two standards are maturing and worth watching:
81
+
82
+ - **[Web Bot Auth](https://datatracker.ietf.org/doc/draft-meunier-web-bot-auth-architecture/)** — cryptographic crawler verification via HTTP Message Signatures (RFC 9421). Bots sign requests with a key published at `/.well-known/http-message-signatures-directory`; sites verify identity instead of guessing from user-agent strings. Already implemented by Cloudflare and Google (`agent.bot.goog`). It directly affects the `agent-access` check: a WAF using Web Bot Auth may pass a real, signed crawler while rejecting ax-audit's unsigned probe — which is why that check's findings carry an explicit verified-bots caveat.
83
+ - **Pay-per-crawl / HTTP 402** — Cloudflare and the RSL payment vocabulary point toward metered, paid agent access. RSL already encodes the terms; enforcement protocols (Open License Protocol, x402) are emerging.
84
+
85
+ ## How the families compose
86
+
87
+ A fully AX-ready site tells a coherent story across all four:
88
+
89
+ > "Here's my content in a form you can read **(family 1)**, here's the interface to interact with me **(family 2)**, here's exactly who may use it and how, for what license **(family 3)**, over a fast and trustworthy connection **(family 4)**."
90
+
91
+ ax-audit's weighting follows what actually stops an agent. **Content** carries the most weight, led by `html-rendering`, because a page whose content only exists after hydration is invisible to the crawlers that do not run JavaScript, no matter how many discovery files it publishes. **Access** is next, because a firewall rule or a `nosnippet` directive silently undoes everything else.
92
+
93
+ `llms-txt` sits at 5 rather than the 11 it carried in 3.x. Adoption studies find most published files are never fetched by an AI search crawler, and Google has stated that Search ignores them; the vendors that do read them are coding agents like Claude Code and Cursor. It is worth publishing, as developer tooling, and it is not worth twice what having content is worth.
94
+
95
+ Checks resting on draft specifications — `ai-catalog`, `webmcp`, `commerce-discovery` — are reported but never scored. A specification that may be renamed next quarter should not move your number.
96
+
97
+ ## See also
98
+
99
+ - [getting-started.md](./getting-started.md) — run your first audit
100
+ - [checks.md](./checks.md) — exact scoring per standard
101
+ - The [remediation guides](https://axrush.com/guides) — how to implement each one
package/docs/faq.md ADDED
@@ -0,0 +1,89 @@
1
+ # FAQ & Troubleshooting
2
+
3
+ ## Scores & results
4
+
5
+ ### Why did my score change after upgrading ax-audit?
6
+
7
+ In 3.x, score changes on the same site are treated as **breaking** and only happen in major or minor releases that explicitly say so. The 3.0.0 release redistributed weights across 14 checks and added Content-Type penalties — see its CHANGELOG entry. Every check added since (3.1.0–3.6.0) ships at **weight 0** precisely so your score and baselines don't move. To track changes deliberately, use `--baseline` (see [cli.md](./cli.md)).
8
+
9
+ ### Why is my score lower than my Lighthouse / SEO score?
10
+
11
+ ax-audit measures the *AI-agent* surface, not performance, accessibility, or human SEO. A fast, beautiful site can still score poorly if it has no `llms.txt`, ships an empty SPA shell to non-JS crawlers, and exposes no structured data. That gap is the reason the tool exists.
12
+
13
+ ### A check shows 0 but the file exists — why?
14
+
15
+ Most "not found" hard-fails mean the request didn't return a 2xx. Common causes: the file is served with a redirect chain that breaks, a non-2xx status, or — most often — a WAF/bot-rule blocking ax-audit's request (see below). Re-run with `--verbose` to see the exact status per request.
16
+
17
+ ### What's the difference between a weighted and an informational check?
18
+
19
+ Weighted checks (14) sum to 100% and determine your overall score. Informational checks (4: `content-negotiation`, `rsl`, `agent-access`, `crawl-efficiency`) run and report full findings but contribute 0 to the score in 3.x. They gain weight in v4.0. The Content Signals findings inside `robots-txt` are likewise informational.
20
+
21
+ ## False positives & caveats
22
+
23
+ ### `agent-access` flags crawlers as blocked, but my real crawlers work fine
24
+
25
+ This is the most important caveat in the tool. ax-audit's probe sends a user-agent *containing* the crawler token (e.g. `...GPTBot/1.0`) but it is **not** the real, verified crawler. If your WAF verifies bots cryptographically ([Web Bot Auth](./concepts.md)) or by IP range, it will correctly pass the genuine GPTBot while rejecting ax-audit's unverified probe. **Before changing any WAF rule, confirm against your WAF logs** whether real crawler traffic is actually being served. If it is, this finding is a false positive for your setup.
26
+
27
+ ### Where did `well-known-ai` go?
28
+
29
+ Removed in 4.0. Three of the five files it scored turned out to have no consumer at all: `/.well-known/nlweb.json` appears in no NLWeb release, document or commit; `genai.txt` has no specification; and `/ai-plugin.json` described ChatGPT plugins, shut down in April 2024. Marking a site down for omitting a file nobody specified is worse than not checking it.
30
+
31
+ ### A check says `n/a` — is that bad?
32
+
33
+ No, it is the point. A blog has no API to describe and no MCP server to advertise, so those checks report `n/a` and leave the score entirely rather than counting as failures. Everything counted against your site is something your site could have done.
34
+
35
+ If you are *planning* to build one of those surfaces and want it audited anyway, use `--profile api`, `--profile mcp`, or `--profile all`.
36
+
37
+ ### My score changed after upgrading to 4.0
38
+
39
+ Expected. 4.0 redistributed the weights and made protocol checks conditional. Sites with no API or MCP surface generally rise, because they are no longer marked down for lacking things they do not have.
40
+
41
+ Your saved baseline still works, but regression gating is suspended until you re-save it: comparing across a scoring change would report regressions your site did not cause. Run `ax-audit <url> --save-baseline .ax-baseline.json` to resume gating.
42
+
43
+ ### `crawl-efficiency` says no compression, but my CDN compresses
44
+
45
+ The check reads the `Content-Encoding` header on the response it received. If a proxy between ax-audit and your origin strips or fails to negotiate compression, you'll see this. Verify directly: `curl -sI -H 'Accept-Encoding: br, gzip' https://your-site.com | grep -i content-encoding`.
46
+
47
+ ### `content-negotiation` fails but I don't serve Markdown
48
+
49
+ That's expected — most sites don't yet. It's informational (weight 0). Adopt it when you're ready; the [guide](https://axrush.com/guides/content-negotiation) covers Cloudflare/Vercel zero-code options.
50
+
51
+ ## Running the tool
52
+
53
+ ### My WAF is blocking ax-audit itself
54
+
55
+ ax-audit sends a `User-Agent` of `ax-audit/<version> (+https://github.com/duranitech/ax-audit)`. If your firewall challenges unknown agents, allowlist that UA (or the IP you run from) for the duration of the audit. Note that several checks deliberately send *other* user-agents (`agent-access`) and unusual `Accept` headers (`content-negotiation`) — a WAF rejecting those is itself a finding, not a tool bug.
56
+
57
+ ### How do I audit a staging site behind auth?
58
+
59
+ ax-audit has no auth support today. Options: run it from inside the network perimeter, temporarily allowlist its UA/IP, or audit a public preview deployment (the typical CI pattern — see [ci.md](./ci.md)).
60
+
61
+ ### Audits are slow / flaky on cold deployments
62
+
63
+ Transient failures (timeouts, 5xx) retry automatically with backoff — raise `--retries` (default 2) for very cold preview environments and `--timeout` (default 10000ms) for slow origins. In batch mode, `--concurrency` speeds up multi-URL runs.
64
+
65
+ ### Can I run only some checks?
66
+
67
+ Yes: `--checks llms-txt,robots-txt,rsl`. Note the overall score then averages *only* those checks, so a subset run isn't comparable to a full-audit score. Unknown IDs error out with the valid list.
68
+
69
+ ### Is there rate limiting I should know about?
70
+
71
+ The tool itself doesn't rate-limit, but it makes several requests per audit (one per check, plus follow-ups for conditional GET, content negotiation, and the 8 `agent-access` probes). All responses are cached per run, so repeated checks of the same URL don't re-fetch. Be considerate auditing sites you don't own.
72
+
73
+ ## Integration
74
+
75
+ ### Does it work in CI?
76
+
77
+ Yes — exit codes gate the build (`0` = Good/Excellent, `1` = Fair/Poor). See [ci.md](./ci.md) for GitHub Actions recipes including PR comments via `--output markdown` and regression gates via `--baseline`.
78
+
79
+ ### Can I consume results programmatically?
80
+
81
+ Yes — `import { audit } from 'ax-audit'` returns a typed `AuditReport`. See [api.md](./api.md).
82
+
83
+ ### How do I generate the files ax-audit checks for?
84
+
85
+ Use [ax-init](https://github.com/duranitech/ax-init) — it generates `llms.txt`, `robots.txt`, an Agent Card, `security.txt`, structured data, and header snippets, then you verify with `npx ax-audit`. Check its output against this tool: the standards moved in 2026, and a generator written against the older paths will produce files agents no longer look for.
86
+
87
+ ## Still stuck?
88
+
89
+ Open an issue at [github.com/duranitech/ax-audit/issues](https://github.com/duranitech/ax-audit/issues) with the output of `npx ax-audit <url> --verbose`.
@@ -0,0 +1,108 @@
1
+ # Getting Started
2
+
3
+ This walkthrough takes you from zero to a passing AX score: run your first audit, learn to read the report, and fix findings in the order that moves your score most.
4
+
5
+ ## 1. Run your first audit
6
+
7
+ No install needed:
8
+
9
+ ```bash
10
+ npx ax-audit https://your-site.com
11
+ ```
12
+
13
+ You get a report like:
14
+
15
+ ```
16
+ AX Audit Report
17
+ https://your-site.com
18
+
19
+ ██████████████████████░░░░░░░░░░░░░░░░░░ 56/100 Fair
20
+
21
+ LLMs.txt (0/100)
22
+ FAIL /llms.txt not found
23
+ ...
24
+ ```
25
+
26
+ Three things to locate immediately:
27
+
28
+ - **The overall score and grade.** 0–100, weighted across 14 checks. Grades: Excellent (≥90), Good (≥70), Fair (≥50), Poor (<50). The CLI exits `0` at Good or better — that is the CI gate.
29
+ - **Per-check scores.** Each check is independent and scored 0–100. The weight of each check is in [checks.md](./checks.md).
30
+ - **Findings.** Every `WARN`/`FAIL` line carries a hint and a `learnMoreUrl` to a remediation guide with copy-pasteable fixes.
31
+
32
+ To see only what needs fixing:
33
+
34
+ ```bash
35
+ npx ax-audit https://your-site.com --only-failures
36
+ ```
37
+
38
+ ## 2. Understand what you're optimizing
39
+
40
+ AI agents interact with your site differently than browsers: most don't execute JavaScript, they look for machine-readable discovery files, and they respect (or at least read) your declared crawler policy. The audit measures three layers — if you're new to the standards involved (llms.txt, A2A, MCP, RSL, Content Signals), read [concepts.md](./concepts.md) first:
41
+
42
+ 1. **Content** — is there substance an agent can read, and can a browser agent act on it? (`html-rendering`, `agent-operability`, `structured-data`, `seo-basics`, `content-negotiation`)
43
+ 2. **Access** — can an agent actually retrieve it? (`agent-access`, `ai-directives`, `http-hygiene`, `tls-https`, `crawl-efficiency`)
44
+ 3. **Discovery** — can an agent find your machine-readable files? (`robots-txt`, `llms-txt`, `http-headers`, `sitemap`, `meta-tags`)
45
+ 4. **Policy** — what usage rights do you declare, and do they agree? (`usage-policy`, `security-txt`, `rsl`)
46
+ 5. **Protocols** — what can an agent call? (`api-discovery`, `agent-card`, `mcp-discovery`, `agent-skills`, `auth-discovery`)
47
+
48
+ Protocol checks are conditional: if your site has no API to describe and no MCP server, they report **n/a** and leave the score alone rather than counting as failures.
49
+
50
+ ## 3. Fix in impact order
51
+
52
+ The fastest path from Fair to Good, by weight and typical effort:
53
+
54
+ | Step | Check | Weight | Typical effort |
55
+ | --- | --- | --- | --- |
56
+ | 1 | Verify server-rendered content | 11% | Free if you already render on the server; significant if you ship an SPA shell. Nothing else matters if an agent sees an empty page. |
57
+ | 2 | Confirm nothing blocks AI crawlers | 9% | 15 minutes in your WAF. Check `agent-access` first: your robots.txt may say one thing and your firewall another. |
58
+ | 3 | Configure `robots.txt` for the 12 core AI crawlers | 9% | 15 minutes; `npx ax-init` generates it |
59
+ | 4 | Name your buttons and label your inputs | 7% | Hours to days, and it is accessibility work you owed anyway |
60
+ | 5 | Check your page-level AI directives | 6% | 10 minutes. A stray `nosnippet` removes you from AI Overviews. |
61
+ | 6 | Add JSON-LD structured data | 6% | 1–2 hours |
62
+ | 7 | Create `/llms.txt` | 5% | 30 minutes — it is a Markdown file. Read the check's note on who actually fetches it first. |
63
+
64
+ Steps 1 and 2 are the ones worth doing today. They are also the two most likely to be silently broken: a hydration-only page and a firewall rule are both invisible from inside the site.
65
+
66
+ The remaining checks (`seo-basics`, `content-negotiation`, `usage-policy`, `http-hygiene`, `http-headers`, `security-txt`, `tls-https`, `sitemap`, `crawl-efficiency`, `rsl`, `meta-tags`) are mostly configuration; the remediation guides give exact snippets for Nginx, Vercel, Netlify, and Express.
67
+
68
+ Re-run after each fix — all requests are cached per run, so audits are fast and cheap.
69
+
70
+ ## 4. Lock in your progress with a baseline
71
+
72
+ Once you reach a score you're happy with, freeze it:
73
+
74
+ ```bash
75
+ npx ax-audit https://your-site.com --save-baseline .ax-baseline.json
76
+ git add .ax-baseline.json && git commit -m "chore: AX baseline"
77
+ ```
78
+
79
+ From then on, compare every run against it:
80
+
81
+ ```bash
82
+ npx ax-audit https://your-site.com --baseline .ax-baseline.json --fail-on-regression 5
83
+ ```
84
+
85
+ This catches drift you didn't cause — a CDN toggle, a WAF rule, a header dropped in a refactor. Wire it into CI with the recipes in [ci.md](./ci.md).
86
+
87
+ ## 5. Look at the informational checks
88
+
89
+ Four checks report findings without affecting your score yet (they will in v4.0): `content-negotiation`, `rsl`, `agent-access`, `crawl-efficiency`. Treat them as the early-warning lane — they cover the newest standards, and fixing them now means v4.0 changes nothing for you.
90
+
91
+ The one to check first is `agent-access`: it detects the failure mode you cannot see — your robots.txt allows GPTBot while your WAF returns it a 403:
92
+
93
+ ```bash
94
+ npx ax-audit https://your-site.com --checks agent-access
95
+ ```
96
+
97
+ ## Common first-run questions
98
+
99
+ - **"My score seems harsh."** The audit measures the AI-agent surface, not site quality. A beautiful SPA with no llms.txt, no structured data, and an empty `#root` div is genuinely poor AX — that's the point of the tool.
100
+ - **"A check crashed / network error."** Transient failures retry automatically (`--retries`, default 2). For slow staging environments raise `--timeout`.
101
+ - **"Which findings are safe to ignore?"** See the [FAQ](./faq.md) — notably the `agent-access` verified-bots caveat, and the checks resting on draft specifications, which never affect your score.
102
+
103
+ ## Next steps
104
+
105
+ - [checks.md](./checks.md) — exact scoring of all 26 checks, with the weight table
106
+ - [concepts.md](./concepts.md) — the AX standards landscape explained
107
+ - [cli.md](./cli.md) — every flag · [ci.md](./ci.md) — CI recipes · [api.md](./api.md) — programmatic use
108
+ - [ax-init](https://github.com/duranitech/ax-init) — generates most of the files this tool audits