@surea11y/core 1.6.0 → 1.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (59) hide show
  1. package/CHANGELOG.md +47 -0
  2. package/README.md +24 -38
  3. package/docs/ACT_RULE_MAPPING.md +8 -6
  4. package/docs/API_STABILITY.md +51 -3
  5. package/docs/BINDING_AUTHORS_GUIDE.md +104 -2
  6. package/docs/DESIGN_CHALLENGES.md +66 -0
  7. package/docs/EARL.md +100 -0
  8. package/docs/ENGINE_OPTIONS.md +28 -2
  9. package/docs/INTEGRATION.md +4 -2
  10. package/docs/LIMITATIONS.md +3 -1
  11. package/docs/OUTPUT_SCHEMA.md +44 -6
  12. package/docs/POLICY.md +1 -1
  13. package/docs/RULE_AUTHORING.md +11 -12
  14. package/docs/RULE_CATALOG.md +76 -26
  15. package/docs/RULE_HELPERS.md +333 -0
  16. package/docs/RULE_TAXONOMY.md +25 -4
  17. package/docs/SARIF.md +21 -2
  18. package/docs/WCAG_CONFORMANCE.md +9 -1
  19. package/package.json +9 -3
  20. package/src/checks/automatic/aria-allowed-attr.js +6 -0
  21. package/src/checks/automatic/aria-allowed-role.js +32 -23
  22. package/src/checks/automatic/aria-braille-equivalent.js +18 -10
  23. package/src/checks/automatic/aria-conditional-attr.js +17 -10
  24. package/src/checks/automatic/aria-deprecated-role.js +12 -0
  25. package/src/checks/automatic/aria-hidden-body.js +1 -1
  26. package/src/checks/automatic/aria-prohibited-attr.js +5 -0
  27. package/src/checks/automatic/aria-prohibited-children.js +6 -6
  28. package/src/checks/automatic/aria-required-attr.js +59 -12
  29. package/src/checks/automatic/aria-required-children.js +33 -16
  30. package/src/checks/automatic/aria-required-parent.js +32 -6
  31. package/src/checks/automatic/aria-role-name-present.js +1 -1
  32. package/src/checks/automatic/aria-roles-valid.js +52 -21
  33. package/src/checks/automatic/aria-valid-attr-value.js +74 -21
  34. package/src/checks/automatic/aria-valid-attr.js +14 -9
  35. package/src/checks/automatic/avoid-inline-spacing.js +133 -6
  36. package/src/checks/automatic/contrast-computable.js +10 -0
  37. package/src/checks/automatic/contrast-enhanced.js +12 -0
  38. package/src/checks/automatic/contrast-minimum.js +12 -0
  39. package/src/checks/automatic/css-orientation-lock.js +42 -5
  40. package/src/checks/automatic/duplicate-id-aria.js +5 -0
  41. package/src/checks/automatic/duplicate-id.js +13 -8
  42. package/src/checks/automatic/form-control-single-label.js +9 -0
  43. package/src/checks/automatic/identical-iframes-same-purpose.js +229 -0
  44. package/src/checks/automatic/iframe-focusable-content.js +5 -0
  45. package/src/checks/automatic/label-in-name.js +38 -56
  46. package/src/checks/automatic/link-in-text-block.js +279 -23
  47. package/src/checks/automatic/target-size-minimum.js +84 -5
  48. package/src/checks/automatic/td-has-header.js +19 -18
  49. package/src/checks/manual/form-control-label-quality-manual.js +134 -24
  50. package/src/checks/manual/landmark-complementary-is-top-level-manual.js +231 -0
  51. package/src/checks/manual/password-paste-enabled-manual.js +255 -0
  52. package/src/core.js +3863 -44184
  53. package/src/earl.js +144 -0
  54. package/src/sarif.js +22 -2
  55. package/surea11y.browser.js +10 -41039
  56. package/surea11y.i18n.de.js +2 -21
  57. package/surea11y.i18n.es.js +2 -21
  58. package/surea11y.i18n.fr.js +2 -21
  59. /package/src/checks/automatic/{role-img-alt-present.js → role-img-text-alternative-present.js} +0 -0
package/docs/EARL.md ADDED
@@ -0,0 +1,100 @@
1
+ # EARL report
2
+
3
+ `@surea11y/core/earl` renders scan results as [EARL 1.0](https://www.w3.org/TR/EARL10-Schema/) in JSON-LD — the vocabulary the W3C publishes for stating "this tool tested this thing and got this result", and the format the [ACT Rules](https://act-rules.github.io/) community group accepts as an implementation report.
4
+
5
+ ```js
6
+ const { renderEarlReport } = require('@surea11y/core/earl');
7
+
8
+ const report = renderEarlReport(result, {
9
+ assertor: { name: 'surea11y', version: '1.7.0' },
10
+ mode: 'earl:automatic'
11
+ });
12
+
13
+ fs.writeFileSync('earl.jsonld', JSON.stringify(report, null, 2));
14
+ ```
15
+
16
+ ## What it is for
17
+
18
+ Two audiences, and they want the same document for different reasons.
19
+
20
+ An **implementation report** tells the ACT Rules community group how this engine behaves against their test cases, which is what gets an engine listed alongside the other implementations. Listing is not endorsement and the W3C does not verify the data — the accurate phrasing is *"listed as an ACT implementation"*, never *"W3C certified"*.
21
+
22
+ A **consumer** gets an interchange format. EARL is what accessibility tooling reads when it has to combine results from more than one source — an automated scan and a manual audit, say, or several engines — because every assertion carries who asserted it and how. That is worth having whether or not anything is ever submitted anywhere.
23
+
24
+ ## How it differs from the other reporters
25
+
26
+ [`SARIF.md`](./SARIF.md) and [`REPORT.md`](./REPORT.md) both carry **violations only**: they report `fail` and `cantTell` occurrences and drop everything else. EARL is the opposite. Every rule that ran becomes an assertion, `pass` and `inapplicable` included, because an implementation report is a claim about what the engine decided *everywhere*. A rule that stayed silent because it found nothing applicable is evidence, not noise — it is how a reader distinguishes "this engine checked and found nothing to check" from "this engine does not implement that rule at all".
27
+
28
+ ## Shape
29
+
30
+ The graph groups by subject rather than being a flat list of assertions:
31
+
32
+ ```json
33
+ {
34
+ "@context": "https://www.w3.org/WAI/content-assets/wcag-act-rules/earl-context.json",
35
+ "@graph": [
36
+ {
37
+ "@type": "TestSubject",
38
+ "source": "https://example.test/",
39
+ "assertions": [
40
+ {
41
+ "@type": "Assertion",
42
+ "test": { "title": "img-alt-present", "isPartOf": ["WCAG2:non-text-content"] },
43
+ "result": { "outcome": "earl:failed" },
44
+ "assertedBy": {
45
+ "@type": "Assertor",
46
+ "name": "surea11y",
47
+ "release": { "@type": "Version", "revision": "1.7.0" }
48
+ },
49
+ "mode": "earl:automatic"
50
+ }
51
+ ]
52
+ }
53
+ ]
54
+ }
55
+ ```
56
+
57
+ - **`source`** is the scanned URL, or `about:blank` when a result carries none.
58
+ - **`test.title`** is the engine's own rule id. In ACT terms a rule is the *procedure* the implementation ran, which is exactly what a rule id names.
59
+ - **`test.isPartOf`** lists the Success Criteria that rule maps to, as `WCAG2:<criterion-id>`. Omitted entirely for a rule claiming no criterion — `aria-allowed-role` is the engine's one automatic rule in that position, and asserting an empty list would read as "maps to nothing we could find" rather than "deliberately maps to none".
60
+ - **`assertedBy`** and **`mode`** appear only when you supply them.
61
+
62
+ Criterion ids are derived from the criterion's own title (`Non-text Content` → `non-text-content`). `normativeMappings` also carries Understanding-document references and non-WCAG standards, which share `standard: "WCAG"` and a `requirement` with the real thing; a Success Criterion is the entry that states a conformance level and claims no other document type, and only those are read.
63
+
64
+ ## Outcomes
65
+
66
+ | Engine | EARL |
67
+ |---|---|
68
+ | `pass` | `earl:passed` |
69
+ | `fail` | `earl:failed` |
70
+ | `cantTell` | `earl:cantTell` |
71
+ | `notApplicable` | `earl:inapplicable` |
72
+
73
+ `earl:untested` has no counterpart: a rule that did not run produces no result to assert on, so it contributes no assertion rather than an untested one.
74
+
75
+ **`cantTell` does not cost conformance credit.** ACT's own consistency rules allow an automated implementation to report "cannot tell" on some — though not all — examples and still count as consistent. What a partially consistent implementation may *not* do is produce a false positive: failing an example the rule says should pass, or that is inapplicable. That is the gate worth watching, and it is a property of the rules rather than of this reporter. As of the last full run the engine produces **zero false positives** across the 798 ACT examples covering its 58 matched rules — see [`ACT_RULE_MAPPING.md`](./ACT_RULE_MAPPING.md), and re-run `scripts/act-testcase-check.js` for the live figure rather than trusting this one indefinitely.
76
+
77
+ ## Several results, one report
78
+
79
+ `renderEarlReport` takes an array as readily as a single result, because a report covering many pages is the normal case:
80
+
81
+ ```js
82
+ renderEarlReport([homeResult, checkoutResult, searchResult], { assertor });
83
+ ```
84
+
85
+ Results sharing a URL merge into one subject — a caller scanning the same page under different `engineOptions` is still describing one resource, and the context has no way to express two subjects with the same source. Where two results assert on the same rule for the same URL, the last one wins.
86
+
87
+ Output is deterministic: subjects sort by source, assertions by rule id, and the same inputs produce byte-identical output in any order. That is what makes a diff between two engine versions meaningful.
88
+
89
+ ## Options
90
+
91
+ | Option | Meaning |
92
+ |---|---|
93
+ | `assertor` | `{ name, version }`. Defaults the name to `surea11y`; pass `null` to omit `assertedBy` entirely. |
94
+ | `mode` | An EARL test mode such as `'earl:automatic'`. Omitted when not supplied. |
95
+
96
+ ## See also
97
+
98
+ - [`OUTPUT_SCHEMA.md`](./OUTPUT_SCHEMA.md) — the result this reads
99
+ - [`ACT_RULE_MAPPING.md`](./ACT_RULE_MAPPING.md) — which ACT rules the engine's rules correspond to
100
+ - [`API_STABILITY.md`](./API_STABILITY.md) — what is covered by semver
@@ -55,10 +55,34 @@ Since versions are cumulative (2.1 = 2.0 + new; 2.2 = 2.0 + 2.1 + new), select a
55
55
  { tags: ['wcag22a', 'wcag22aa', 'wcag22aaa'] }
56
56
  ```
57
57
 
58
- **One SC goes the other way.** WCAG 2.2 removed SC 4.1.1 Parsing — the only criterion ever dropped rather than added. A rule mapped to it carries its 2.0-origin tag (`wcag2a`) like any other baseline rule, plus `wcag22-removed`, and the version tag sets above therefore include it under a 2.2 target, where it does not belong. Exclude it explicitly:
58
+ **One SC goes the other way.** WCAG 2.2 removed SC 4.1.1 Parsing — the only criterion ever dropped rather than added. A rule mapped to it carries its 2.0-origin tag (`wcag2a`) like any other baseline rule, plus `wcag22-removed`, and the version tag sets above therefore include it under a 2.2 target, where it does not belong.
59
+
60
+ You do not have to do anything about that. The engine resolves a **target WCAG version** for every run and, when that target is 2.2, a `wcag22-removed` rule cannot report `fail`: it still runs, still reports every occurrence it found, but its outcome is coerced to `cantTell` and the result carries a `wcagVersionScope` field saying why (see [`OUTPUT_SCHEMA.md`](./OUTPUT_SCHEMA.md#a-check-result-checksresultsi)). Nothing is silently dropped, and a 2.2 run is not gated by a criterion 2.2 does not contain.
61
+
62
+ The target version is resolved in this order:
63
+
64
+ 1. `engineOptions.wcagVersion` — `'2.0'`, `'2.1'` or `'2.2'`, if you set it.
65
+ 2. The version-origin tags in your own filter: a set topping out at `wcag21a`/`wcag21aa` reads as a 2.1 target, one containing any `wcag22*` tag as 2.2, one with only `wcag2*` tags as 2.0. Only those nine tags count — an SC tag (`wcag411`) or `best-practice` says nothing about a version.
66
+ 3. Otherwise `'2.2'`, this engine's default target.
67
+
68
+ ```js
69
+ // Nothing to declare: a plain run already targets 2.2, so a duplicate id
70
+ // comes back cantTell rather than fail.
71
+ runDomRulesInPage(url, null, {}, null);
72
+
73
+ // Conformance-testing against 2.1, where SC 4.1.1 still exists:
74
+ runDomRulesInPage(url, null, { wcagVersion: '2.1' }, null);
75
+
76
+ // Same thing, implied by the tag set — no extra option needed:
77
+ runDomRulesInPage(url, null, {}, { tags: ['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa'] });
78
+ ```
79
+
80
+ The resolved target is reported back on every result as `engine.wcagVersion`, so you can confirm which one a run actually used.
81
+
82
+ If you would rather not see the rule at all under 2.2, exclude it outright — the tag is still there for exactly that:
59
83
 
60
84
  ```js
61
- // WCAG 2.2 AA conformance, without the criterion 2.2 removed:
85
+ // WCAG 2.2 AA conformance, with the removed criterion left out entirely:
62
86
  {
63
87
  tags: ['wcag2a', 'wcag2aa', 'wcag21a', 'wcag21aa', 'wcag22a', 'wcag22aa'],
64
88
  excludeTags: ['wcag22-removed']
@@ -86,6 +110,7 @@ runDomRulesInPage(url, null, {
86
110
  ```js
87
111
  const engineOptions = {
88
112
  locale: 'en', // default 'en'; de-DE falls back to de, then to en per string
113
+ wcagVersion: '2.2', // default '2.2' — the conformance target, see "Filtering by WCAG version" above
89
114
  messages: { de: { /* key: text */ } }, // optional caller-supplied dictionaries; win over built-in ones
90
115
  includeHiddenElements: false, // default false — set true to evaluate hidden/collapsed subtrees too
91
116
  includeShadowDom: true, // default true — opt OUT with `false` to skip open shadow roots
@@ -131,6 +156,7 @@ const engineOptions = {
131
156
  | Option | Meaning |
132
157
  |---|---|
133
158
  | `locale` | Any string. A code with a subtag falls back to its base language first, so `de-DE` uses `de`; failing that, English. Individual strings then fall back the same way (chosen locale → `en` → the rule's literal English text), so a partly-translated locale never produces missing text. All of that is silent in the strings themselves, so the result reports what actually happened in `engine.locale` — check it if you need to know whether you got the language you asked for. See [`I18N.md`](./I18N.md). |
159
+ | `wcagVersion` | `'2.0'`, `'2.1'` or `'2.2'` — which version of WCAG the run is conformance-testing against. Defaults to whatever your version-origin tags imply, and to `'2.2'` when they imply nothing. The only thing it currently changes is SC 4.1.1 Parsing, removed in 2.2: under a 2.2 target a rule tagged `wcag22-removed` still runs and still reports its occurrences, but cannot `fail` — see ["Filtering by WCAG version"](#filtering-by-wcag-version-21-vs-22) above. Any other value is ignored and the default applies. |
134
160
  | `messages` | Optional `{ [locale]: { key: text } }`. Checked before the engine's own tables, so it can override individual strings or supply a language the build does not carry. Keys you omit fall back normally, so a partial override is fine. This is how the standalone browser bundle receives a locale side file, and it is the only way to get a dictionary into a page context, since the in-page runner is serialized and cannot read files. See [`I18N.md`](./I18N.md). |
135
161
  | `includeHiddenElements` | Default `false`: helper queries exclude elements hidden by structural/CSS mechanisms such as `display:none`, `[hidden]`, closed `<details>`, and hidden rendering-only host elements (with descendants excluded too). Set `true` to include those hidden/collapsed subtrees in evaluation (legacy/static-markup behavior). |
136
162
  | `includeShadowDom` | Default `true`: rules using `helpers.queryAllSmart` traverse into open shadow roots. Set `false` to scan only the light DOM. Closed shadow roots are never reachable either way (no DOM API exposes them). |
@@ -167,6 +167,8 @@ Notes for CI specifically:
167
167
 
168
168
  **How it works**: a parent frame's `runa11yCoreAcrossFrames()` call pings each direct child `<iframe>`/`<frame>` via `postMessage`; if — and only if — that child has *also* called `a11yCoreEnableFrameResponder()` (its own opt-in to being scannable from above), it runs its own scan and replies with the result, which the parent includes. **A non-cooperating frame (the common case for most third-party embeds you don't control) is simply unreachable** — the same-origin policy allows no way around it from inside the page.
169
169
 
170
+ **Who a responder answers**: the frame that embeds it, and nothing else. Enabling the responder is consent to be scanned *from above*, not by anything that can reach you — a sibling frame can obtain a reference through `parent.frames[i]` and `postMessage` to you across origins, and a scan result carries `occurrences[].html`, which is DOM content the same-origin policy otherwise makes unreadable to it. A `run` command whose sender is not the direct parent is ignored, as is one arriving at a window nothing embeds. Replies are matched the same way: only the frame a request was addressed to can answer it, so another window cannot settle a scan in flight by naming its id. The relay is hop-by-hop — a grandchild is reached through its own parent — so a legitimate request always arrives from the direct parent.
171
+
170
172
  ```js
171
173
  // Inside the embedded/child page (e.g. a widget's own bundle), once, at load:
172
174
  const { a11yCoreEnableFrameResponder } = require('@surea11y/core');
@@ -191,6 +193,6 @@ A few things worth knowing:
191
193
  - **Async, unlike the other two runners** — `postMessage` round-trips can't be synchronous, so this is a separate, Promise-returning pair rather than an `engineOptions` flag on `runa11yCoreInPage` (which stays synchronous, unchanged, for every existing caller).
192
194
  - **`engineOptions.pingWaitTime`** (default `500`ms) and **`engineOptions.frameWaitTime`** (default `60000`ms) control how long a child frame gets to answer a ping and a full run request respectively.
193
195
  - **No jsdom/Node equivalent** — this is browser-only. jsdom's window/frame model doesn't meaningfully represent independent-realm cross-origin `postMessage`, and the feature has no purpose in Node anyway.
194
- - **Bundler-free, like `runa11yCoreInPage`** — both functions are fully self-contained (their own private copy of the rule catalog and every helper they need), so raw-source injection (a bookmarklet, a content script with no build step) works with zero bundler needed, exactly like `runa11yCoreInPage` already does. If you *do* use a normal bundler/`require`/`import`, that works too, unchanged.
195
- - **Cost of that self-containment**: `src/core.js` grew from ~1.86MB to ~3.1MB when these two functions landed, since each needed its own complete private copy of the rule catalog and shared helpers rather than sharing the outer `RULE_IMPLS`; it is ~4.3MB now, and grows with every rule added, three times over. If this file's size ever becomes a real problem, the fix would be to drop the bundler-free requirement for just these two functions (accepting that cross-frame scanning in "plain script injection" mode needs a real bundler, unlike `runa11yCoreInPage` alone) rather than tripling the embedded catalog again for some future feature.
196
+ - **Bundler-free, like `runa11yCoreInPage`** — raw-source injection (a bookmarklet, a content script with no build step) still works with no bundler, but the slice you inject has to start at the `// SELF-CONTAINED in-page runner` marker rather than at the cross-frame block. `runa11yCoreAcrossFrames` scans its own frame by calling `runa11yCoreInPage`, so the two travel together. Everything from that marker to `module.exports` is one contiguous chunk with no `require()` in it. If you *do* use a normal bundler/`require`/`import`, that works too, unchanged.
197
+ - **Cost**: these two functions used to carry their own private copy of the rule catalog and helpers, which put the catalog in `src/core.js` three times over and took the file to ~4.3MB. They share `runa11yCoreInPage`'s copy now, which brings it to ~2.67MB and leaves one copy to grow as rules are added.
196
198
  - **No origin/identity check on the sender** beyond the message's own namespaced envelope. Running a read-only scan and replying with DOM-derived results isn't a privileged operation; the content involved is no more sensitive than what's already rendered on the page.
@@ -13,13 +13,15 @@ surea11y is a **static DOM scan**: it reads the DOM tree and computed styles at
13
13
  ## Environment-dependent — depends on how you run it
14
14
 
15
15
  - **jsdom (Node, no real browser) has no CSS layout engine.** Rules needing real geometry — most notably `target-size-minimum` (WCAG 2.5.8, needs real `getBoundingClientRect()`) — report `notApplicable` under plain jsdom rather than guess. Run under a real browser (Puppeteer/Playwright — see [`INTEGRATION.md`](./INTEGRATION.md) Pattern 2) to get real findings from these rules.
16
+ - **Whether text wraps needs layout, so text-spacing findings on non-wrapping text are reported for review.** WCAG 1.4.12 and the ACT rules behind it apply only to text containing a soft wrap break, which a static scan cannot establish. `avoid-inline-spacing` treats text as wrapping by default, so an ordinary forced value below the metric still fails; where no wrap is possible for a reason that *is* visible without layout — text not allowed to wrap, or a fixed-width element inside a horizontally scrolling ancestor — it reports `cantTell` instead. Text that never wraps for some other reason is still reported as a failure.
16
17
  - **`<dialog>` and other elements hidden by the default UA stylesheet** (no `open` attribute, `display: none` by spec), along with any other subtree hidden via `display:none`, `visibility:hidden`, `[hidden]`, or closed `<details>`, are **excluded from rule evaluation by default** — matching the visibility-aware behavior of other established engines. This is a deliberate default (`engineOptions.includeHiddenElements: false`), not an oversight: hidden content isn't reachable by assistive technology or keyboard until it's shown, so flagging a markup defect inside it by default would often be noise. Set `engineOptions.includeHiddenElements: true` to evaluate hidden/collapsed subtrees anyway — e.g. to catch a markup defect (like a broken ARIA ID reference) before a dialog ever opens. See [`ENGINE_OPTIONS.md`](./ENGINE_OPTIONS.md#engineoptions--the-rest) for the option and exactly which hiding mechanisms it covers.
17
18
  - **jsdom's computed `text-shadow` is unreliable on a second read of the same element.** Confirmed in jsdom 29.1.1: reading a computed `text-shadow` value a second time on the same element — through any accessor, from any freshly-requested `CSSStyleDeclaration` for that element, regardless of caching — silently returns a different, wrong "no shadow" value instead of the real declared one. The first read is always correct. This engine works around it internally by reading each element's `text-shadow` exactly once per run and caching the result (see `__textShadowInfoEl` in `src/core/contrast-helpers.js`), so a single scan is unaffected. It only surfaces if you read `getComputedStyle(el).textShadow` yourself, more than once, against the same jsdom-parsed element — a real browser has no such bug.
19
+ - **Under jsdom, scan time grows with the square of DOM depth, not with element count.** jsdom resolves inherited CSS by walking an element's ancestor chain on every `getComputedStyle` call, so one call per element costs O(elements x depth). Measured on jsdom 29.1.1 with no engine code involved: 4000 elements in one chain take 8.3s of `getComputedStyle` alone, against 0.25s for the same 4000 as siblings. The engine's own ancestor walks are capped and stay linear, so this is jsdom's cost rather than the rules'. It affects `runDomRulesInPage` and anything built on it, including `@surea11y/test-matchers`; a real browser computes inherited style natively and does not have this shape. Component frameworks routinely nest 100-300 deep, which is comfortably fast — it becomes noticeable past roughly 1000.
18
20
  - **Static markup vs. live/post-hydration DOM state.** The rule logic itself is DOM-source-agnostic — it evaluates whatever DOM it's handed, whether that's jsdom-parsed static HTML (Pattern 1) or an already-loaded, already-hydrated real browser tab (Pattern 2, see [`INTEGRATION.md`](./INTEGRATION.md)). But the CLI (`npx @surea11y/cli scan <url>`) specifically fetches static HTML only, with no JS execution — see [the CLI docs](https://github.com/SureA11y/cli/blob/main/docs/CLI.md). For a JS-framework-hydrated widget whose server-rendered markup intentionally ships one state before client JS syncs it (e.g. `<input type="checkbox" aria-checked="true">` shipped before client JS sets the native `checked` property to match on hydration — an extremely common, entirely legitimate pattern), a CLI scan only sees the pre-hydration markup. A scan running inside an actual loaded browser tab sees the post-hydration state instead, so the two can disagree on exactly this class of element for reasons that have nothing to do with rule correctness. That's why `aria-checked-state-mismatch` is capped at `manual`/`cantTell` rather than a hard `fail`. If you need live-DOM accuracy for hydration-sensitive checks, run the library directly against an already-loaded page via Pattern 2, not the static-fetch CLI.
19
21
 
20
22
  ## Not attempted: judgment calls that aren't automatable safely
21
23
 
22
- These have no comparably safe heuristic at this engine's confidence bar (`fail` must stay reserved for deterministic, high-confidence violations, full stop). Building them anyway would either catch almost nothing (too narrow to be useful) or risk real false positives (too broad to trust):
24
+ These have no comparably safe heuristic at this engine's bar (`fail` must stay reserved for deterministic violations, full stop). Building them anyway would either catch almost nothing (too narrow to be useful) or risk real false positives (too broad to trust):
23
25
 
24
26
  - **"Is this heading/label text meaningful?"** — real headings and labels are enormously varied and legitimately short ("FAQ," "Name," "Overview" are all fine), so nothing decides from markup whether a heading describes the section under it or a label describes the field beside it. What *is* decidable is that some strings cannot describe anything: `heading-quality` and `form-control-label-quality` flag leftover placeholders, numbered template slots, filenames and URLs against curated exact-match lists, the same precision-over-recall trade-off `link-name-quality` makes. Both are `manual` rules capped at `cantTell` — they raise a candidate for review, they never assert the text is wrong.
25
27
  - **"Does this error message describe the problem?"** — what triggers a validation error and its content are almost always JS/validation-library-driven, invisible to a static scan in the first place; not just a heuristic-design problem.
@@ -18,7 +18,8 @@ This is the exact shape of the object returned by `runDomRulesInPage(...)` / `ru
18
18
  engine: {
19
19
  tag: string,
20
20
  schemaVersion: string,
21
- locale: { requested: string, resolved: string, reason: string }
21
+ locale: { requested: string, resolved: string, reason: string },
22
+ wcagVersion: "2.0" | "2.1" | "2.2"
22
23
  },
23
24
  url: string | null,
24
25
  title: string | null,
@@ -37,6 +38,7 @@ This is the exact shape of the object returned by `runDomRulesInPage(...)` / `ru
37
38
  | `engine.schemaVersion` | The result-schema version (`"1.0.0"`). Bump-worthy if this document's shape ever changes incompatibly — pin to it if you're parsing output programmatically. See [`API_STABILITY.md`](./API_STABILITY.md) for the full stable/unstable field list and version-bump policy. |
38
39
  | `engine.locale` | Which dictionary the run actually used. `requested` is your `engineOptions.locale` after trimming (`"en"` if you passed nothing or a non-string); `resolved` is the locale whose dictionary was used; `reason` explains the pairing. Because locale fallback is graceful and per-string, asking for a language the build does not carry produces English text rather than an error — this field is how you find that out without reading the strings. Reported once per result: a run uses one dictionary throughout. |
39
40
  | `engine.locale.reason` | `"ok"` — you got the dictionary you asked for, and it carries every key. `"primary-subtag"` — your code had a subtag with no dictionary of its own, so its base language was used: `"de-DE"` resolves to `"de"`. `"dictionary-not-loaded"` — the project ships that language, but this build does not carry it and none was supplied (the standalone browser bundle, without its locale side file). `"unknown-locale"` — the project has no such translation at all. `"partial-dictionary"` — the dictionary was used but is missing keys English has, so those strings fell back to English. Treat the set as open; later releases can add to it. |
41
+ | `engine.wcagVersion` | Which version of WCAG this run was conformance-tested against: your `engineOptions.wcagVersion`, or what your version-origin tags implied, or the default `"2.2"`. It affects one thing today — a rule mapped only to SC 4.1.1 Parsing cannot `fail` under a 2.2 target (see `checksResults[i].wcagVersionScope` below). Reported once per result: a run has one target throughout. |
40
42
  | `url` | The `pageUrl` argument you passed in, or `document.location.href` if you passed `null`/omitted it, or `null` if neither is available. |
41
43
  | `title` | `document.title` at scan time, or `null`. |
42
44
  | `timestamp` | **Not auto-generated.** Only set if you pass `engineOptions.timestamp` as a non-empty string — the engine has no built-in clock (deterministic-by-design). If you want a scan timestamp in the result, supply it yourself. |
@@ -95,6 +97,11 @@ This is the exact shape of the object returned by `runDomRulesInPage(...)` / `ru
95
97
  },
96
98
  engineOptions: object, // the resolved engineOptions this rule actually ran under
97
99
  schemaVersion: string,
100
+ wcagVersionScope?: { // present only when the target WCAG version changed this outcome
101
+ target: "2.0" | "2.1" | "2.2",
102
+ removedSc: string[],
103
+ coercedFrom: "fail"
104
+ },
98
105
  error?: string // present only if the rule threw — see below
99
106
  }
100
107
  ```
@@ -102,14 +109,17 @@ This is the exact shape of the object returned by `runDomRulesInPage(...)` / `ru
102
109
  Notes:
103
110
 
104
111
  - **`outcome` vs `outcomeNormalized`**: identical except `notApplicable` becomes `"inapplicable"` in `outcomeNormalized`. Both are provided so you can match either your own vocabulary or the engine's internal one.
105
- - **`type: "manual"` rules can never report `outcome: "fail"`.** If a manual rule's own logic would have said `fail`, the engine coerces it to `cantTell` and appends an explanatory note to `error` — this is enforced centrally (`policy.coerceManualFailToCantTell`, on by default under the `a11y` policy contract; see [`POLICY.md`](./POLICY.md)), not something each rule has to remember. `fail` is reserved for deterministic, high-confidence, `type: "automatic"` findings only.
112
+ - **`type: "manual"` rules can never report `outcome: "fail"`.** If a manual rule's own logic would have said `fail`, the engine coerces it to `cantTell` and appends an explanatory note to `error` — this is enforced centrally (`policy.coerceManualFailToCantTell`, on by default under the `a11y` policy contract; see [`POLICY.md`](./POLICY.md)), not something each rule has to remember. `fail` is reserved for deterministic, `type: "automatic"` findings only.
106
113
  - **`meta.normativeMappings`** is how a check result ties back to a WCAG Success Criterion — `[]` for rules with no formal WCAG mapping (this engine calls them advisory `type: "manual"` rules). See [`WCAG_CONFORMANCE.md`](./WCAG_CONFORMANCE.md) for how these roll up.
114
+ - **`wcagVersionScope`**: only present when the run's target WCAG version turned this rule's `fail` into a `cantTell` — today that means a rule mapped to SC 4.1.1 Parsing (`duplicate-id`) under the default 2.2 target, since 2.2 removed that criterion. `removedSc` lists the criteria that stopped existing, `target` is the version that removed them, and `coercedFrom` is the outcome the rule itself reported. The occurrences are the rule's own, unchanged — nothing was dropped, only the conformance verdict was. Absent on every other result, and **never** reported through `error`: nothing went wrong. See [`ENGINE_OPTIONS.md`](./ENGINE_OPTIONS.md#filtering-by-wcag-version-21-vs-22).
107
115
  - **`error`**: only present if the rule implementation threw an uncaught exception, or if the manual-fail coercion above fired. A thrown rule always surfaces as `outcome: "cantTell"` with `occurrences: []` and `error` set to the exception message — the engine never lets one broken rule crash the whole scan.
108
116
  - **`engineOptions`** on each result is the *resolved* options object (after locale/contrast defaults were applied), not literally what you passed in — useful for confirming what a given rule actually saw, especially the resolved `locale` and `contrast.mode`/`contrast.rootCanvasFallback`.
109
117
 
110
118
  ## An occurrence (`occurrences[i]`)
111
119
 
112
- Only present when `outcome` is `fail` or `cantTell` (a `pass`/`notApplicable` result has `occurrences: []` — this engine does not enumerate the elements it passed, only the ones it flagged).
120
+ Normally present only when `outcome` is `fail` or `cantTell`: a `pass` result has `occurrences: []`, since this engine does not enumerate the elements it passed, only the ones it flagged.
121
+
122
+ `notApplicable` is the one exception. A rule that had nothing to judge may attach a single occurrence saying why, and the contrast rules do exactly that when no text had a computable background — the difference between "checked, nothing to flag" and "could not check" is one this engine reports rather than hides. Such an occurrence describes the scan, not an element, so its `selector` is empty. Do not read `occurrences.length` as a violation count without checking `outcome` first.
113
123
 
114
124
  ```ts
115
125
  {
@@ -119,6 +129,13 @@ Only present when `outcome` is `fail` or `cantTell` (a `pass`/`notApplicable` re
119
129
  summary: string,
120
130
  hint: string,
121
131
  i18n: { summaryKey: string, hintKey: string, params: object } | null,
132
+ occurrenceOutcome?: "fail" | "cantTell", // present when the rule graded its findings into tiers
133
+ uncertainty?: { // present only on a cantTell-tier occurrence
134
+ code: "not-computable" | "runtime-dependent" | "spec-only"
135
+ | "equivalence-unknown" | "judgement-required" | "out-of-scope",
136
+ needed?: string, // what would settle the question
137
+ evidence?: object // what the rule did establish, rule-specific
138
+ },
122
139
  data: {
123
140
  visibilityFilter?: { eligible: boolean, targetSet: string, accEligible: boolean | null, reasons: string[] },
124
141
  details?: object // rule-specific, non-normative — see below
@@ -128,14 +145,33 @@ Only present when `outcome` is `fail` or `cantTell` (a `pass`/`notApplicable` re
128
145
 
129
146
  | Field | Meaning |
130
147
  |---|---|
131
- | `selector` | A best-effort CSS selector built to resolve back to the flagged element (see `helpers.buildSelector` in `RULE_AUTHORING.md`). Not guaranteed unique in adversarial DOM shapes, but the engine actively verifies it resolves to the reported element before using it. |
148
+ | `selector` | A best-effort CSS selector built to resolve back to the flagged element (see `helpers.buildSelector` in `RULE_AUTHORING.md`). Not guaranteed unique in adversarial DOM shapes, but the engine actively verifies it resolves to the reported element before using it. The exception is a rule whose finding *is* an absent element: `page-title-present` reports `head > title` with an `html` of `<title>(missing)</title>`, neither of which is on the page. Both are constants, so the fingerprint they feed stays stable, but do not treat `selector` as resolvable or `html` as real markup without checking the rule reported something that exists. |
132
149
  | `html` | An outer-HTML snippet of the flagged element — use this as your primary "which element" signal when `includeShadowDom: true` (selectors don't pierce shadow boundaries). |
133
150
  | `structuralPath` | The flagged element's sibling-index path from `documentElement` down to it (e.g. `[1, 0, 2]`) — `[]` if the element *is* `documentElement`, `null` if it couldn't be determined. A more robust element-identity mechanism than `selector` alone: it survives DOM changes a selector string wouldn't (an id/class rename, for instance), at the cost of not being usable as an actual CSS selector. Computed from the element reference when the rule kept one, otherwise by re-resolving `selector` against the document (same caveat as `selector` itself: a non-unique selector could resolve to a different element than intended). |
134
151
  | `summary` | Human-readable, already localized ("This button has no accessible name."). |
135
152
  | `hint` | Human-readable remediation guidance, already localized. |
136
153
  | `i18n` | The raw translation keys behind `summary`/`hint`, if you want to re-render them in a different locale yourself without re-running the scan. `null` if the occurrence didn't use key-based i18n. |
137
154
  | `data.visibilityFilter` | Present on most occurrences: why the engine considered this element eligible (or not) under whichever eligibility model the rule used. `eligible` is that result; `targetSet` says which model produced it (`'dom'`: raw DOM/CSS visibility — most rules; `'acc'`: accessibility-tree eligibility). `accEligible` mirrors `eligible` only when `targetSet` is `'acc'`, otherwise `null` — don't read it as a second, independent signal. `reasons` is a list of machine-readable exclusion codes when `eligible: false`. |
138
- | `data.details` | Rule-specific structured data (e.g. `reasonCode`, computed metrics, resolved references) — **non-normative**: useful for building richer UI or debugging, but never changes what `outcome`/`severity` mean. Shape varies per rule; treat as best-effort extra context, not a stable contract. |
155
+ | `data.details` | Rule-specific structured data (computed metrics, resolved references) — **non-normative**: useful for building richer UI or debugging, but never changes what `outcome`/`severity` mean. Shape varies per rule; treat as best-effort extra context, not a stable contract. The one exception is `data.details.reasonCode`, which **is** stable: it identifies *which* of a rule's findings this is, and together with `ruleId` and `html` forms the fingerprint baselines and SARIF are keyed on. A rule may gain a new reason code in a minor release; a shipped one does not change. See [`API_STABILITY.md`](./API_STABILITY.md#finding-identity). |
156
+ | `occurrenceOutcome` | Which tier this occurrence belongs to, on a rule that graded its findings into a confident `fail` tier and a needs-review `cantTell` tier. A rule reporting one tier only omits it, in which case the result's own `outcome` is the occurrence's tier. This is why a `fail` result can carry `cantTell`-tier occurrences: the aggregate outcome stays singular so CI can still gate on it, without discarding the findings that only warranted review. |
157
+ | `uncertainty` | Why this finding could not be decided — see [Uncertainty codes](#uncertainty-codes) below. |
158
+
159
+ ### Uncertainty codes
160
+
161
+ A `cantTell` says the engine did not decide. `uncertainty` says **why**, from a closed vocabulary, so a consumer can branch on the reason rather than parse a summary string. It is present only on a `cantTell`-tier occurrence: a `fail`-tier one would be claiming the rule both decided and did not, so the engine drops it.
162
+
163
+ | `code` | Meaning | Typical shape |
164
+ |---|---|---|
165
+ | `not-computable` | The evidence the rule needed could not be read in this environment. | A cross-origin stylesheet, a background colour that resolves to no value, an `src` that will not resolve. |
166
+ | `runtime-dependent` | The markup cannot settle it because script decides at runtime. | An `aria-controls` naming an element the widget builds when it opens. |
167
+ | `spec-only` | A real specification violation, but the exposed name, role and value survive it, so no Success Criterion is established as failed. | An ARIA attribute whose absence the specification supplies a default for. |
168
+ | `equivalence-unknown` | Two things may or may not serve the same purpose, and neither the markup nor the content settles it. | Two frames sharing an accessible name but embedding different resources. |
169
+ | `judgement-required` | The question is inherently a human call. | Whether an undersized target is essential; every `type: "manual"` rule. |
170
+ | `out-of-scope` | The finding is real but falls outside the standard this run targets. | A rule mapped only to a criterion the target WCAG version removed. |
171
+
172
+ `needed` states, in one sentence, what would settle the question — the thing a reviewer has to go and check. `evidence` carries what the rule *did* establish, so the reviewer starts from the engine's work rather than repeating it; its shape is rule-specific and, like `data.details`, not a stable contract. The `code` is: new codes may be added in a minor release, but an existing one does not change meaning, so branch on the codes you know and treat an unrecognised one as "needs review" rather than an error.
173
+
174
+ Every automatic rule that can report `cantTell` carries this, and a test holds that line so a new one cannot arrive without it. The `out-of-scope` code is attached by the engine rather than by a rule, on the same occurrences that produce a result-level `wcagVersionScope`. Manual rules do not carry it: `judgement-required` is what `type: "manual"` already means, so repeating it per occurrence would say nothing the result does not.
139
175
 
140
176
  ## A composite result (`rulesResults[i]`)
141
177
 
@@ -164,7 +200,7 @@ Rollup precedence (deterministic, in this order): **any contributor `fail` → c
164
200
 
165
201
  | Outcome | Meaning | Can appear on `type: "manual"`? |
166
202
  |---|---|---|
167
- | `fail` | Deterministic, high-confidence, normative violation — no heuristics, no guessing. | No (coerced to `cantTell`) |
203
+ | `fail` | Deterministic, normative violation — the decision procedure guesses at nothing. | No (coerced to `cantTell`) |
168
204
  | `pass` | The rule's applicable target(s) exist and none were flagged. | Yes |
169
205
  | `cantTell` | Requires human judgment — either genuinely ambiguous, or a `manual` rule's advisory finding. | Yes |
170
206
  | `notApplicable` | The rule found no elements it applies to on this page/scope. | Yes |
@@ -176,6 +212,8 @@ Rollup precedence (deterministic, in this order): **any contributor `fail` → c
176
212
  - `severity`: `minor` < `moderate` < `serious` < `critical` — the rule author's assessment of user impact, independent of `confidence`.
177
213
  - `confidence`: `low` < `medium` < `high` — how certain the engine is that a `fail`/`cantTell` verdict is correct. Both are informational metadata for prioritization; neither changes `outcome`'s meaning.
178
214
 
215
+ A `fail` is not always `confidence: "high"`, and that is not a contradiction. The outcome describes the decision procedure — it resolved the question without guessing — while `confidence` describes the model that decision was made against. A handful of automatic rules decide deterministically against something that is itself an approximation (the curated WAI-ARIA role tables, the native-role mappings, an accessibility tree inferred from static markup) and report `medium`: `aria-required-children`, `aria-prohibited-children`, `aria-required-parent`, `aria-allowed-attr`, `form-control-programmatic-label-present`, `svg-image-text-alternative-present`, `video-poster-text-alternative-present` and `target-size-minimum`. `confidence` is on every result, so a consumer that wants only the most certain failures can gate on it directly; `policy.allowedConfidence` will not do it for you, since a disallowed value is replaced with the rule's own `defaultConfidence` rather than changing the outcome (see [`POLICY.md`](./POLICY.md)).
216
+
179
217
  ## Worked example
180
218
 
181
219
  Scanning `<img src="logo.png">` (no `alt`) and `<button></button>` (no accessible name), scoped to just those two rules via `runOnly: { includeRuleIds: [...] }` (see [`ENGINE_OPTIONS.md`](./ENGINE_OPTIONS.md) — this is **not** a bare array):
package/docs/POLICY.md CHANGED
@@ -68,4 +68,4 @@ Neither of these ever throws — policy resolution is designed to always produce
68
68
 
69
69
  ## Why this exists as a separate layer
70
70
 
71
- Keeping outcome-integrity rules (like "manual rules can't fail") in a policy layer — rather than hard-coded into every rule, or worse, left to each rule author's discretion — means the guarantee holds even if a rule's own logic has a bug, and means different consumers can have different appetites for risk (a CI gate vs. an internal audit dashboard) without forking the rule set itself. This protects the engine's core guarantee: `fail` must always mean "deterministic, high-confidence, normative violation," full stop.
71
+ Keeping outcome-integrity rules (like "manual rules can't fail") in a policy layer — rather than hard-coded into every rule, or worse, left to each rule author's discretion — means the guarantee holds even if a rule's own logic has a bug, and means different consumers can have different appetites for risk (a CI gate vs. an internal audit dashboard) without forking the rule set itself. This protects the engine's core guarantee: `fail` must always mean "deterministic, normative violation," full stop — see [`OUTPUT_SCHEMA.md`](./OUTPUT_SCHEMA.md#severity-and-confidence-values) for why that is not the same as "high-confidence."
@@ -254,17 +254,15 @@ that one is on you.
254
254
 
255
255
  ## 6) Helpers contract used by rules (ctx.helpers)
256
256
 
257
- Rules use helpers returned by `createDomHelpers()`.
258
-
259
- Helpers observed in this repo include:
260
- - `queryAll`, `queryAllDeep`, `queryAllSmart`
261
- - `getOuterHtmlSnippet`
262
- - `buildSimpleSelector`, `buildSelector`
263
- - `isAccTreeEligible`, `getEligibilityInfo`
264
- - `resolveIdRefs`, `getTextFromIdRefs`
265
- - `getAccessibleNameInfo`, `getAccessibleDescriptionInfo`
266
- - `getTextAlternativeInfo`
267
- - `getRoleInfo`, `getFocusableInfo`
257
+ Rules use helpers returned by `createDomHelpers()`. The most load-bearing ones —
258
+ `queryAllSmart` (query with shadow/hidden/exclude handling built in),
259
+ `getAccessibleNameInfo`/`getAccessibleDescriptionInfo`/`getTextAlternativeInfo` (naming),
260
+ `isAccTreeEligible`/`getEligibilityInfo` (visibility), `getRoleInfo`/`getFocusableInfo`
261
+ (role/focus) — cover most rules.
262
+
263
+ **See [`RULE_HELPERS.md`](./RULE_HELPERS.md) for the full reference** (~35 helpers plus
264
+ the `contrast.*`/`aria.*` namespaces), with what each one does and when to reach for it
265
+ instead of reimplementing the logic in a new rule.
268
266
 
269
267
  ### 6.1 Shadow DOM scanning
270
268
 
@@ -484,4 +482,5 @@ This writes `tests/fixtures/INDEX.md` (human-readable), `tests/fixtures/index.js
484
482
  counts, for external tooling to enumerate and load fixtures directly) and
485
483
  `tests/fixtures/index.html` (the same listing as a browsable page). Commit all three
486
484
  alongside the fixture and test changes. A rule shipped without its fixture is treated
487
- the same as a rule shipped without tests — not done.
485
+ the same as a rule shipped without tests — not done. `npm run fixtures:check` reports
486
+ a stale index without rewriting it, and CI fails on one.