@redocly/recheck 0.4.1 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (69) hide show
  1. package/README.md +365 -127
  2. package/dist/cli.js +59 -0
  3. package/dist/cli.js.map +1 -1
  4. package/dist/commands/baseline.d.ts +10 -0
  5. package/dist/commands/baseline.d.ts.map +1 -0
  6. package/dist/commands/baseline.js +59 -0
  7. package/dist/commands/baseline.js.map +1 -0
  8. package/dist/commands/readability.d.ts +10 -0
  9. package/dist/commands/readability.d.ts.map +1 -0
  10. package/dist/commands/readability.js +90 -0
  11. package/dist/commands/readability.js.map +1 -0
  12. package/dist/commands/run.d.ts.map +1 -1
  13. package/dist/commands/run.js +52 -8
  14. package/dist/commands/run.js.map +1 -1
  15. package/dist/config/load.d.ts +2 -0
  16. package/dist/config/load.d.ts.map +1 -1
  17. package/dist/config/load.js +3 -0
  18. package/dist/config/load.js.map +1 -1
  19. package/dist/config/schema.d.ts +4 -0
  20. package/dist/config/schema.d.ts.map +1 -1
  21. package/dist/config/schema.js +3 -0
  22. package/dist/config/schema.js.map +1 -1
  23. package/dist/config/validate.d.ts +2 -0
  24. package/dist/config/validate.d.ts.map +1 -1
  25. package/dist/config/validate.js +3 -1
  26. package/dist/config/validate.js.map +1 -1
  27. package/dist/core/baseline.d.ts +49 -0
  28. package/dist/core/baseline.d.ts.map +1 -0
  29. package/dist/core/baseline.js +0 -0
  30. package/dist/core/baseline.js.map +1 -0
  31. package/dist/core/prose-extract.d.ts +6 -0
  32. package/dist/core/prose-extract.d.ts.map +1 -0
  33. package/dist/core/prose-extract.js +79 -0
  34. package/dist/core/prose-extract.js.map +1 -0
  35. package/dist/core/readability.d.ts +17 -0
  36. package/dist/core/readability.d.ts.map +1 -0
  37. package/dist/core/readability.js +37 -0
  38. package/dist/core/readability.js.map +1 -0
  39. package/dist/core/runner.d.ts.map +1 -1
  40. package/dist/core/runner.js +1 -0
  41. package/dist/core/runner.js.map +1 -1
  42. package/dist/data/markdoc-realm-schema.d.ts.map +1 -1
  43. package/dist/data/markdoc-realm-schema.js +13 -0
  44. package/dist/data/markdoc-realm-schema.js.map +1 -1
  45. package/dist/reporter/formats/json.d.ts +5 -1
  46. package/dist/reporter/formats/json.d.ts.map +1 -1
  47. package/dist/reporter/formats/json.js +2 -1
  48. package/dist/reporter/formats/json.js.map +1 -1
  49. package/dist/reporter/index.js +2 -2
  50. package/dist/reporter/index.js.map +1 -1
  51. package/dist/reporter/summary.d.ts +1 -1
  52. package/dist/reporter/summary.d.ts.map +1 -1
  53. package/dist/reporter/summary.js +2 -1
  54. package/dist/reporter/summary.js.map +1 -1
  55. package/dist/rules/scope/metric.d.ts +2 -14
  56. package/dist/rules/scope/metric.d.ts.map +1 -1
  57. package/dist/rules/scope/metric.js +11 -187
  58. package/dist/rules/scope/metric.js.map +1 -1
  59. package/dist/rules/scope/semantic-line-breaks.d.ts.map +1 -1
  60. package/dist/rules/scope/semantic-line-breaks.js +58 -35
  61. package/dist/rules/scope/semantic-line-breaks.js.map +1 -1
  62. package/dist/rules/token/link-fragments.d.ts.map +1 -1
  63. package/dist/rules/token/link-fragments.js +251 -91
  64. package/dist/rules/token/link-fragments.js.map +1 -1
  65. package/dist/rules/types.d.ts +2 -0
  66. package/dist/rules/types.d.ts.map +1 -1
  67. package/dist/types/reporting.d.ts +9 -0
  68. package/dist/types/reporting.d.ts.map +1 -1
  69. package/package.json +2 -2
package/README.md CHANGED
@@ -133,7 +133,8 @@ git diff --name-only origin/main... | node dist/cli.js run . --changed-only
133
133
 
134
134
  ## Library API
135
135
 
136
- The CLI is a thin wrapper around a public library API, published from `packages/recheck`'s `dist/index.js`. This is the intended integration point for embedding Recheck in another tool (a build step, an editor extension, or another CLI like Redocly CLI's `lint`) rather than shelling out:
136
+ The CLI is a thin wrapper around a public library API, published from `packages/recheck`'s `dist/index.js`.
137
+ This is the intended integration point for embedding Recheck in another tool (a build step, an editor extension, or another CLI like Redocly CLI's `lint`) rather than shelling out:
137
138
 
138
139
  ```ts
139
140
  import { lintContent, lintFiles } from '@redocly/recheck';
@@ -162,17 +163,26 @@ const { problems: fileProblems, fixedFiles } = await lintFiles(['README.md'], co
162
163
 
163
164
  Key exports:
164
165
 
165
- - **`parseMarkdown(content, options?)`** — parses a markdown string into a micromark-based token tree, once. Every other API in this list builds on this tree rather than re-parsing. `options.markdoc` is a boolean here: `true` also tokenizes `{% ... %}` Markdoc tag spans into `markdocTag` tokens, and `false` or omitted gives you the same tree as passing no options at all. The [object form](#markdoc-aware-linting-markdoc-true) (`{ schema, extend }`) is a config-file concept only — it resolves down to this boolean before any file is parsed, and `ParseOptions.markdoc` does not accept it.
166
+ - **`parseMarkdown(content, options?)`** — parses a markdown string into a micromark-based token tree, once.
167
+ Every other API in this list builds on this tree rather than re-parsing.
168
+ `options.markdoc` is a boolean here: `true` also tokenizes `{% ... %}` Markdoc tag spans into `markdocTag` tokens, and `false` or omitted gives you the same tree as passing no options at all.
169
+ The [object form](#markdoc-aware-linting-markdoc-true) (`{ schema, extend }`) is a config-file concept only — it resolves down to this boolean before any file is parsed, and `ParseOptions.markdoc` does not accept it.
166
170
  - **`extractScopes(tree, content)`** — segments a parsed token tree into Vale-style scopes (`sentence`, `paragraph`, `heading`, `list-item`, `blockquote`, `table.cell`, etc.) for prose/style rules to run against.
167
- - **`lintContent(content, config, opts?)`** — lints a single in-memory markdown string against a config; no disk access. Rules that need on-disk facts (e.g. `max-image-size`) require `opts.metadata` to be supplied by the caller.
171
+ - **`lintContent(content, config, opts?)`** — lints a single in-memory markdown string against a config; no disk access.
172
+ Rules that need on-disk facts (e.g. `max-image-size`) require `opts.metadata` to be supplied by the caller.
168
173
  - **`lintFiles(paths, config, opts?)`** — lints markdown files from disk; pass `{ fix: true }` to also write auto-fixes back, looping lint → fix → re-lint until the file converges.
169
174
  Files that can't be read are skipped (with a console warning) and reported in the returned `skippedFiles` (`{ path, reason }[]`), so callers can detect incomplete coverage programmatically.
170
175
  `opts.root` sets the lint root that image-metadata loading is confined to (default `process.cwd()`) — image refs resolving outside it are treated as missing without touching the disk.
171
176
  `opts.maxProblems` caps the total problems collected: once a file's lint pushes the run to the cap, later files aren't linted at all and the returned `truncated` flag is set.
172
- - **`runRules(files, rules)`** — the lower-level engine entry point for callers that already have a `NormalizedRule[]` (e.g. from `loadConfig`) and want to run against an explicit in-memory file list, bypassing `lintFiles`'s own config loading/validation. Under `{ fix: true }` its `RunResult` separates the fixes that genuinely landed (`fixes`) from proposals dropped by overlap resolution (`skippedFixes`).
173
- - **`applyFixesToContent(content, fixes)`** — applies `Fix` edits to a string, preserving the file's own line endings (CRLF files stay CRLF). Returns `{ content, applied, skipped }`: every input fix is classified as genuinely applied or skipped (overlapping edits, out-of-range lines), so callers can report what actually changed rather than every proposal.
174
- - **`computeTextStatistics(prose)`** — computes word/sentence/syllable/character/complex-word counts for a plain prose string (not markdown — extract prose from a scope first). Sentence counting reuses `splitSentences` internally, so it agrees with the rest of the engine on sentence boundaries. Tokenization is ASCII-only by design (accented or non-Latin letters don't count as word characters), so readability scores are meaningful for English prose.
175
- - **`computeReadability(formula, stats)`** — scores a `TextStatistics` object with one of six standard readability formulas: `flesch-reading-ease`, `flesch-kincaid-grade`, `gunning-fog`, `smog`, `coleman-liau`, `automated-readability`. Returns `0` (rather than `NaN`/`Infinity`) when `stats.words` or `stats.sentences` is `0`.
177
+ - **`runRules(files, rules)`** — the lower-level engine entry point for callers that already have a `NormalizedRule[]` (e.g. from `loadConfig`) and want to run against an explicit in-memory file list, bypassing `lintFiles`'s own config loading/validation.
178
+ Under `{ fix: true }` its `RunResult` separates the fixes that genuinely landed (`fixes`) from proposals dropped by overlap resolution (`skippedFixes`).
179
+ - **`applyFixesToContent(content, fixes)`** — applies `Fix` edits to a string, preserving the file's own line endings (CRLF files stay CRLF).
180
+ Returns `{ content, applied, skipped }`: every input fix is classified as genuinely applied or skipped (overlapping edits, out-of-range lines), so callers can report what actually changed rather than every proposal.
181
+ - **`computeTextStatistics(prose)`** — computes word/sentence/syllable/character/complex-word counts for a plain prose string (not markdown — extract prose from a scope first).
182
+ Sentence counting reuses `splitSentences` internally, so it agrees with the rest of the engine on sentence boundaries.
183
+ Tokenization is ASCII-only by design (accented or non-Latin letters don't count as word characters), so readability scores are meaningful for English prose.
184
+ - **`computeReadability(formula, stats)`** — scores a `TextStatistics` object with one of six standard readability formulas: `flesch-reading-ease`, `flesch-kincaid-grade`, `gunning-fog`, `smog`, `coleman-liau`, `automated-readability`.
185
+ Returns `0` (rather than `NaN`/`Infinity`) when `stats.words` or `stats.sentences` is `0`.
176
186
  - **`TECHNICAL_PROPER_NOUNS`** — the [built-in technical proper-noun vocabulary](#built-in-technical-proper-noun-vocabulary) `capitalization`/`spelling` consume by default; re-exported so you can read it or build your own tooling around the same list.
177
187
 
178
188
  This is exactly the surface a host tool needs to add both markdown-structure linting and prose/style linting to content it already has in memory — for example, linting the markdown inside an OpenAPI `description` field without writing it to a temp file first.
@@ -246,6 +256,58 @@ recheck/ul-style-dash:
246
256
  style: dash
247
257
  ```
248
258
 
259
+ ## Baseline
260
+
261
+ A baseline lets a team adopt recheck on a large document set with no cleanup project first: record the findings that exist today, then fail only on new ones.
262
+
263
+ ```bash
264
+ recheck baseline # writes recheck-baseline.yaml next to your config
265
+ ```
266
+
267
+ Activate it with one config line:
268
+
269
+ ```yaml
270
+ baseline: ./recheck-baseline.yaml
271
+ ```
272
+
273
+ The file stores one count per file per rule, errors only, sorted for stable diffs:
274
+
275
+ ```yaml
276
+ version: 1
277
+ files:
278
+ docs/index.md:
279
+ recheck/semantic-line-breaks: 3
280
+ ```
281
+
282
+ With the baseline active, `recheck run`:
283
+
284
+ - **suppresses** findings whose (file, rule) count matches the baseline, and reports how many matched;
285
+ - **fails** when a count rises — the group's findings are printed with `(baseline 3, found 5)` context;
286
+ - **fails** when a count falls, because the baseline is stale — the message says to run `recheck baseline` and commit the result.
287
+ Counts only step down, so the file equals reality at every green commit.
288
+
289
+ Warnings are never baselined; they do not affect exit codes.
290
+ Partial runs (`--rule`, `--changed-only`, a narrower path) compare only the files they scanned and the rules they ran, so they never false-alarm about what they did not see.
291
+ Line numbers are deliberately not stored: counts survive unrelated edits, and a baseline diff in review reads as "this PR pays down 4 findings."
292
+ A renamed file is a new path with no budget, so its pre-existing findings report as new until you regenerate — the baseline diff then shows the counts moving from the old path to the new one.
293
+
294
+ ## Readability
295
+
296
+ `recheck readability` reports scores per file: Flesch reading ease, Flesch-Kincaid grade, Automated Readability Index (ARI), words, and sentences, plus medians.
297
+ ARI is a grade level computed from exact character counts, with no syllable heuristic, which makes it steadier on technical vocabulary.
298
+ It is score-shaped, not rule-shaped: it never gates and always exits 0 when it ran.
299
+ To gate on a bound, use the `metric` assertion — both read the same prose and the same formulas, so they can never disagree.
300
+
301
+ ```bash
302
+ recheck readability docs
303
+ recheck readability docs --output json
304
+ recheck readability docs --changed-only < changed.txt # score only listed files
305
+ ```
306
+
307
+ The score reads flowing prose the way standard readability tools do: headings, code, and Markdoc tags are excluded, and every block ends a sentence.
308
+ A file with no prose reports `—` (null in JSON) rather than a fake zero.
309
+ In CI, run it twice — once on the PR head and once on the merge-base worktree — and join on file to show each changed page's score change.
310
+
249
311
  ## Exceptions
250
312
 
251
313
  Rules can be configured with exceptions to skip specific files or lines:
@@ -307,9 +369,11 @@ recheck/no-trailing-spaces:
307
369
  ## Inline Directives
308
370
 
309
371
  Beyond config-level `exceptions`, individual Markdown files can silence rules
310
- inline with HTML comments — the same mechanism ESLint/Vale users expect. A
372
+ inline with HTML comments — the same mechanism ESLint/Vale users expect.
373
+ A
311
374
  directive names rules by their **short name** (`oxford-comma`) or **full
312
- name** (`recheck/oxford-comma`) — both work. A directive is inert inside a
375
+ name** (`recheck/oxford-comma`) — both work.
376
+ A directive is inert inside a
313
377
  fenced code block (it has to be real, parsed HTML, not just matching text).
314
378
 
315
379
  ```markdown
@@ -342,7 +406,8 @@ The five forms:
342
406
 
343
407
  Rule naming: list one or more rules space-separated, by short name
344
408
  (`oxford-comma`) or full name (`recheck/oxford-comma`) — both work on every
345
- form that accepts names; omitting names targets every rule. Naming a rule
409
+ form that accepts names; omitting names targets every rule.
410
+ Naming a rule
346
411
  that isn't configured produces a warning (`recheck-directive`, severity
347
412
  `warn`) pointing at the directive's line — useful for catching a typo in
348
413
  the disabled rule name — but disables nothing.
@@ -355,7 +420,9 @@ Rules are defined using `assertions` that specify their behavior:
355
420
 
356
421
  #### Swap Assertions (`swap`)
357
422
  Text replacement with configurable options.
358
- **Fixable**: each match is replaced with its pair's value, with the matched text's own casing applied to the replacement -- an all-lowercase match inserts the replacement as configured, a Capitalized match capitalizes just the replacement's first word, and an ALL-CAPS match (2+ letters) uppercases the whole replacement; any other (mixed-case) casing is left as configured, since it carries no reliable intent to infer. This matters most with `ignoreCase: true`: without it, a sentence-initial `'Behaviour'` would be fixed to literal `'behavior'`, silently lowercasing the start of the sentence -- with it, it fixes to `'Behavior'`. (With `keysAreRegex: true`, casing is inferred from the MATCHED text, not the regex key, so this applies uniformly to regex keys too.)
423
+ **Fixable**: each match is replaced with its pair's value, with the matched text's own casing applied to the replacement -- an all-lowercase match inserts the replacement as configured, a Capitalized match capitalizes just the replacement's first word, and an ALL-CAPS match (2+ letters) uppercases the whole replacement; any other (mixed-case) casing is left as configured, since it carries no reliable intent to infer.
424
+ This matters most with `ignoreCase: true`: without it, a sentence-initial `'Behaviour'` would be fixed to literal `'behavior'`, silently lowercasing the start of the sentence -- with it, it fixes to `'Behavior'`.
425
+ (With `keysAreRegex: true`, casing is inferred from the MATCHED text, not the regex key, so this applies uniformly to regex keys too.)
359
426
  When two pairs' matches overlap in the source (a compound key together with the shorter keys it contains), the longest match wins and is reported and fixed as one span.
360
427
 
361
428
  ```yaml
@@ -397,7 +464,8 @@ assertions:
397
464
  | `includeCode` | `boolean` | No | Matches inside inline code spans (`` `like this` ``) are skipped by default, so a token like `master` doesn't fire inside `` `git checkout master` ``. Set `true` to scan inline code too. Default `false`. |
398
465
 
399
466
  #### Occurrence Assertions (`occurrence`)
400
- Vale-parity `occurrence` check: counts regex matches within each scoped segment and flags the segment when the count falls outside `[min, max]`. `min: 1` with no `max` acts as an existence check — it flags a segment where the pattern is missing entirely.
467
+ Vale-parity `occurrence` check: counts regex matches within each scoped segment and flags the segment when the count falls outside `[min, max]`.
468
+ `min: 1` with no `max` acts as an existence check — it flags a segment where the pattern is missing entirely.
401
469
 
402
470
  ```yaml
403
471
  assertions:
@@ -413,9 +481,11 @@ assertions:
413
481
  | `max` | `number` | At least one of `min`/`max` | Maximum allowed match count; more matches is a violation. |
414
482
  | `ignoreCase` | `boolean` | No | Matches `pattern` case-insensitively. Default `false`. |
415
483
 
416
- Omitting both `min` and `max` is a validation error — an occurrence assertion with no bound can never report anything. An unknown option key under `occurrence` is likewise a validation error.
484
+ Omitting both `min` and `max` is a validation error — an occurrence assertion with no bound can never report anything.
485
+ An unknown option key under `occurrence` is likewise a validation error.
417
486
 
418
- The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the actual match count, 2nd = the bound that was violated** (`min` or `max`, whichever applied), e.g. `'Too many sentences (%s found, max %s).'` → `'Too many sentences (4 found, max 3).'`. Not fixable (detection-only): a count-based violation has no single match position to anchor an edit to.
487
+ The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the actual match count, 2nd = the bound that was violated** (`min` or `max`, whichever applied), e.g. `'Too many sentences (%s found, max %s).'` → `'Too many sentences (4 found, max 3).'`.
488
+ Not fixable (detection-only): a count-based violation has no single match position to anchor an edit to.
419
489
 
420
490
  #### Repetition Assertions (`repetition`)
421
491
  Vale-parity `repetition` check: flags an adjacent repeated word — two tokens matching `pattern`, separated only by whitespace (which may include a single hard-wrap newline), so `'the theory'` is not flagged (different words) but `'the the'` and a hard-wrapped `'the\nthe rest'` are. **Fixable**: collapses the pair back to one occurrence, keeping the FIRST token's casing/text (so `'The the'` fixes to `'The'`, not `'the'`).
@@ -433,10 +503,12 @@ assertions:
433
503
 
434
504
  An unknown option key under `repetition` is a validation error, as is a non-string `pattern` or a non-boolean `ignoreCase`; both options are optional, so an empty `repetition: {}` is valid.
435
505
 
436
- The rule's `message` gets one positional `%s` substitution: the repeated word itself, e.g. `'Repeated word "%s".'` → `'Repeated word "the".'`. Fix idempotency holds under repeated `--fix` passes: `'the the the'` converges to `'the'`.
506
+ The rule's `message` gets one positional `%s` substitution: the repeated word itself, e.g. `'Repeated word "%s".'` → `'Repeated word "the".'`.
507
+ Fix idempotency holds under repeated `--fix` passes: `'the the the'` converges to `'the'`.
437
508
 
438
509
  #### Consistency Assertions (`consistency`)
439
- Vale-parity `consistency` check: each `either` entry declares one alternative group — the key and the value are the two variants (both matched as literals with word boundaries, like `swap` keys). Whichever variant appears **first in the file (by source order)** wins file-wide; every later occurrence of the other variant is flagged. **Fixable**: each later occurrence is replaced with the winning variant **literally as written in `either`** — unlike `swap`, the losing match's own casing is not preserved here (with `ignoreCase: true`, a later `'Behaviour'` in a `behavior`-first document fixes to `'behavior'`).
510
+ Vale-parity `consistency` check: each `either` entry declares one alternative group — the key and the value are the two variants (both matched as literals with word boundaries, like `swap` keys).
511
+ Whichever variant appears **first in the file (by source order)** wins file-wide; every later occurrence of the other variant is flagged. **Fixable**: each later occurrence is replaced with the winning variant **literally as written in `either`** — unlike `swap`, the losing match's own casing is not preserved here (with `ignoreCase: true`, a later `'Behaviour'` in a `behavior`-first document fixes to `'behavior'`).
440
512
 
441
513
  ```yaml
442
514
  assertions:
@@ -451,14 +523,16 @@ assertions:
451
523
  | `either` | `object` | Yes | Map of variant pairs; key and value are the two alternatives of one group. Each pair gets its own independent first-seen winner. Must be non-empty. |
452
524
  | `ignoreCase` | `boolean` | No | Matches variants case-insensitively (so `'Behaviour'` counts as an occurrence of `behaviour`). Default `false`. |
453
525
 
454
- Omitting `either`, leaving it empty, or giving it non-string or empty-string keys or values is a validation error — a consistency assertion with no variant pairs can never report anything, and an empty-string key would otherwise reach the scan loop as a zero-width regex that never terminates. An unknown option key under `consistency` is likewise a validation error.
526
+ Omitting `either`, leaving it empty, or giving it non-string or empty-string keys or values is a validation error — a consistency assertion with no variant pairs can never report anything, and an empty-string key would otherwise reach the scan loop as a zero-width regex that never terminates.
527
+ An unknown option key under `consistency` is likewise a validation error.
455
528
 
456
529
  Matches from overlapping scopes (e.g. `scope: [paragraph, sentence]`, where every sentence segment sits inside its paragraph segment) are deduplicated by source position before the winner is decided, so each occurrence is counted — and fixed — exactly once.
457
530
 
458
531
  The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the offending (later) match, 2nd = the first-seen winner**, e.g. `'Inconsistent spelling: "%s" conflicts with first-seen "%s".'` → `'Inconsistent spelling: "behaviour" conflicts with first-seen "behavior".'`.
459
532
 
460
533
  #### Conditional Assertions (`conditional`)
461
- Vale-parity `conditional` check: if `first` (a regex pattern) matches anywhere within the rule's scoped segments, `second` (a regex pattern) must exist **somewhere in the whole file** — checked against the full raw file content, not just the rule's own scope, so a `second` match sitting inside a code block still satisfies a rule scoped to `paragraph`. When `second` is absent file-wide, every `first` match becomes its own problem, at its exact source position. **Detection-only** (not fixable) — there is no single well-defined edit that would "introduce" `second`.
534
+ Vale-parity `conditional` check: if `first` (a regex pattern) matches anywhere within the rule's scoped segments, `second` (a regex pattern) must exist **somewhere in the whole file** — checked against the full raw file content, not just the rule's own scope, so a `second` match sitting inside a code block still satisfies a rule scoped to `paragraph`.
535
+ When `second` is absent file-wide, every `first` match becomes its own problem, at its exact source position. **Detection-only** (not fixable) — there is no single well-defined edit that would "introduce" `second`.
462
536
 
463
537
  ```yaml
464
538
  assertions:
@@ -473,14 +547,16 @@ assertions:
473
547
  | `second` | `string` | Yes | Regex; must match somewhere in the whole file content once `first` has matched. Non-empty. |
474
548
  | `ignoreCase` | `boolean` | No | Matches both `first` and `second` case-insensitively. Default `false`. |
475
549
 
476
- Unlike `swap`/`consistency`'s escaped-literal variants, `first` and `second` are raw user regex patterns (like `pattern`'s `tokens`). Missing, empty, or non-string `first`/`second` is a validation error, as is an unknown option key or a non-boolean `ignoreCase` — but `first`/`second` are **not** validated as compilable regexes at config-load time; an invalid regex in either one silently produces zero problems at runtime instead (same convention as `pattern`).
550
+ Unlike `swap`/`consistency`'s escaped-literal variants, `first` and `second` are raw user regex patterns (like `pattern`'s `tokens`).
551
+ Missing, empty, or non-string `first`/`second` is a validation error, as is an unknown option key or a non-boolean `ignoreCase` — but `first`/`second` are **not** validated as compilable regexes at config-load time; an invalid regex in either one silently produces zero problems at runtime instead (same convention as `pattern`).
477
552
 
478
553
  Matches from overlapping scopes (e.g. `scope: [paragraph, sentence]`) are deduplicated by source position, so each occurrence of `first` is reported exactly once.
479
554
 
480
555
  The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the offending `first` match, 2nd = the `second` pattern that was never introduced**, e.g. `'"%s" appears but "%s" was never introduced.'` → `'"TODO" appears but "DONE" was never introduced.'`.
481
556
 
482
557
  #### Capitalization Assertions (`capitalization`)
483
- Vale-parity `capitalization` check: flags (and — for four of its `match` values — fixes) a scoped segment whose text doesn't already match the required casing. `match` is one of `$title`, `$sentence`, `$lower`, `$upper`, or else a **custom regex** the whole segment text must satisfy.
558
+ Vale-parity `capitalization` check: flags (and — for four of its `match` values — fixes) a scoped segment whose text doesn't already match the required casing.
559
+ `match` is one of `$title`, `$sentence`, `$lower`, `$upper`, or else a **custom regex** the whole segment text must satisfy.
484
560
 
485
561
  ```yaml
486
562
  assertions:
@@ -503,24 +579,31 @@ Unknown option keys, a missing/empty `match`, an invalid `style`, a non-string-a
503
579
 
504
580
  **`$title`** — AP or Chicago title case, implemented in `rules/scope/title-case.ts`'s `apTitleCase`/`chicagoTitleCase`:
505
581
  - The first and last word are **always** capitalized, regardless of any stopword list.
506
- - A hyphenated compound (e.g. `well-known`) runs **each hyphen part** through the same stopword test a standalone word gets for the active style — `well-known` → `Well-Known`, but `editor-in-chief` → `Editor-in-Chief` (`in` is a stopword in both styles). The compound's first part always capitalizes when the compound opens the title, and its last part always capitalizes when the compound closes the title — e.g. `the new state-of-the-art` → `The New State-of-the-Art` (changed from the original simplification by product decision during execution, 2026-07-27).
582
+ - A hyphenated compound (e.g. `well-known`) runs **each hyphen part** through the same stopword test a standalone word gets for the active style — `well-known` → `Well-Known`, but `editor-in-chief` → `Editor-in-Chief` (`in` is a stopword in both styles).
583
+ The compound's first part always capitalizes when the compound opens the title, and its last part always capitalizes when the compound closes the title — e.g. `the new state-of-the-art` → `The New State-of-the-Art` (changed from the original simplification by product decision during execution, 2026-07-27).
507
584
  - A word already in ALL-CAPS (2+ letters, e.g. an acronym like `API`) is left exactly as written.
508
585
  - **AP** (default) lowercases articles (`a`, `an`, `the`), coordinating conjunctions (`and`, `but`, `or`, `nor`, `for`, `so`, `yet`), and prepositions of **3 letters or fewer** (`at`, `by`, `in`, `of`, `off`, `on`, `out`, `to`, `up`, `via`).
509
586
  - **Chicago** lowercases the same articles/conjunctions, plus **every** preposition regardless of length (the short ones above, plus `about`, `above`, `across`, `after`, `against`, `along`, `among`, `around`, `before`, `behind`, `below`, `between`, `during`, `through`, `toward`, `under`, `until`, `with`, `within`, `without`) — e.g. Chicago lowercases `'...walking through the park'` → `'...walking through the Park'`, where AP capitalizes `Through`.
510
587
 
511
588
  **`$sentence`** — only the first word is capitalized; every other word is lowercased unless it's an `exceptions` entry (as-written) or already ALL-CAPS (left alone).
512
589
 
513
- **Word position counts a phrase exception as one word.** A *phrase* exception (one containing whitespace or a dot, like `Node.js` or `VS Code` — see the phrase-matching note above) is a single atomic token in the word sequence the `$`-styles case: it's emitted in its exact as-written form, and it **occupies a position**, so it never changes which word counts as first or last. With `exceptions: [VS Code]`, the already-correctly-cased heading `## VS Code actions for teams` produces no finding under `$sentence` (`actions` is the second word, not the first), and `## a guide to Node.js` becomes `## A Guide to Node.js` under `$title`/AP (`Node.js` is the last word, so `to` is a mid-title stopword and stays lowercase). Single-word exceptions (e.g. `GitHub`) behave as they always have — resolved by lookup rather than position.
590
+ **Word position counts a phrase exception as one word.** A *phrase* exception (one containing whitespace or a dot, like `Node.js` or `VS Code` — see the phrase-matching note above) is a single atomic token in the word sequence the `$`-styles case: it's emitted in its exact as-written form, and it **occupies a position**, so it never changes which word counts as first or last.
591
+ With `exceptions: [VS Code]`, the already-correctly-cased heading `## VS Code actions for teams` produces no finding under `$sentence` (`actions` is the second word, not the first), and `## a guide to Node.js` becomes `## A Guide to Node.js` under `$title`/AP (`Node.js` is the last word, so `to` is a mid-title stopword and stays lowercase).
592
+ Single-word exceptions (e.g. `GitHub`) behave as they always have — resolved by lookup rather than position.
514
593
 
515
- This used to be a bug, tracked as [Redocly/redocly#25610](https://github.com/Redocly/redocly/issues/25610) and fixed since: phrase exceptions were previously *masked out* of the text before word position was computed, which made a leading phrase promote the next word to sentence-initial under `$sentence` (`## VS Code actions for teams` was flagged, and under `fix: true` rewritten to `## VS Code Actions for teams`) and made a trailing phrase promote the preceding word to last-word position under `$title` (`a guide to Node.js` → `A Guide To Node.js`). If you had worked around it by rephrasing headings or by swapping in a custom regex `match`, neither is needed any more. See `rules/scope/title-case.ts`'s `recaseWords` for the tokenization that replaced the masking.
594
+ This used to be a bug, tracked as [Redocly/redocly#25610](https://github.com/Redocly/redocly/issues/25610) and fixed since: phrase exceptions were previously *masked out* of the text before word position was computed, which made a leading phrase promote the next word to sentence-initial under `$sentence` (`## VS Code actions for teams` was flagged, and under `fix: true` rewritten to `## VS Code Actions for teams`) and made a trailing phrase promote the preceding word to last-word position under `$title` (`a guide to Node.js` → `A Guide To Node.js`).
595
+ If you had worked around it by rephrasing headings or by swapping in a custom regex `match`, neither is needed any more.
596
+ See `rules/scope/title-case.ts`'s `recaseWords` for the tokenization that replaced the masking.
516
597
 
517
598
  **`$lower`** / **`$upper`** — the whole segment must be all-lowercase / all-uppercase respectively; no exceptions/ALL-CAPS carve-out (unconditional, matching Vale's own `$lower`/`$upper`).
518
599
 
519
- **Custom regex** — the whole segment text must satisfy the pattern. **Detection-only**: unlike the four `$`-styles, a failing regex is flagged but never auto-fixed, even though the rule itself is registered fixable. Like `pattern`'s `tokens`, an invalid regex is caught and silently produces zero problems rather than crashing the run.
600
+ **Custom regex** — the whole segment text must satisfy the pattern. **Detection-only**: unlike the four `$`-styles, a failing regex is flagged but never auto-fixed, even though the rule itself is registered fixable.
601
+ Like `pattern`'s `tokens`, an invalid regex is caught and silently produces zero problems rather than crashing the run.
520
602
 
521
- **Inline code is frozen.** A backtick-delimited span in the segment text (e.g. a heading like `'the `configFile` option'`) is treated like an exception: its content is never flagged or rewritten by any of the four `$`-styles, even if it would otherwise land on the first/last word.
603
+ **Inline code is frozen.** A backtick-delimited span in the segment text (e.g. a heading like ``'the `configFile` option'``) is treated like an exception: its content is never flagged or rewritten by any of the four `$`-styles, even if it would otherwise land on the first/last word.
522
604
 
523
- **Fixable** for `$title`/`$sentence`/`$lower`/`$upper` only, one segment-wide edit per flagged segment. A **multi-line** segment (e.g. a soft-wrapped paragraph) is skipped entirely under these four styles — neither a problem nor a fix — since a `Fix` can only rewrite a single line; a custom regex `match` has no such restriction and still checks (and reports) multi-line segments, since it never produces a fix regardless of segment span.
605
+ **Fixable** for `$title`/`$sentence`/`$lower`/`$upper` only, one segment-wide edit per flagged segment.
606
+ A **multi-line** segment (e.g. a soft-wrapped paragraph) is skipped entirely under these four styles — neither a problem nor a fix — since a `Fix` can only rewrite a single line; a custom regex `match` has no such restriction and still checks (and reports) multi-line segments, since it never produces a fix regardless of segment span.
524
607
 
525
608
  The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the segment's own text (first line only), 2nd = the `match` value itself** (e.g. `'$title'`, or the literal regex source for custom-regex mode), e.g. `'"%s" should use %s capitalization.'` → `'"the great escape" should use $title capitalization.'`.
526
609
 
@@ -542,15 +625,25 @@ assertions:
542
625
 
543
626
  Omitting both `min` and `max`, an unrecognized `formula`, or an unknown option key are all validation errors — a metric assertion with no bound can never report anything, and an unrecognized formula would otherwise reach the scoring engine's own exhaustive-switch failure at lint time instead of at config validation.
544
627
 
545
- **Always summary-scoped.** Unlike every other assertion above, `metric` does not honor a configurable `scope:` — readability is a property of the WHOLE document's prose, not something a selector could sensibly narrow (a readability score isn't meaningful for one paragraph in isolation the way an `occurrence` count is). Config validation forces every `metric` rule to `scope: summary`. Omit `scope` on a `metric` rule (or write `scope: summary` explicitly); configuring any other scope prints a warning (`metric is always summary-scoped; ignoring configured scope ...`) and applies `summary` behavior anyway. Text from overlapping segments (e.g. a list nested inside a blockquote) is deduplicated by source position, same as `consistency`/`conditional` above.
628
+ **Always summary-scoped.** Unlike every other assertion above, `metric` does not honor a configurable `scope:` — readability is a property of the WHOLE document's prose, not something a selector could sensibly narrow (a readability score isn't meaningful for one paragraph in isolation the way an `occurrence` count is).
629
+ Config validation forces every `metric` rule to `scope: summary`.
630
+ Omit `scope` on a `metric` rule (or write `scope: summary` explicitly); configuring any other scope prints a warning (`metric is always summary-scoped; ignoring configured scope ...`) and applies `summary` behavior anyway.
631
+ Text from overlapping segments (e.g. a list nested inside a blockquote) is deduplicated by source position, same as `consistency`/`conditional` above.
546
632
 
547
- **What the score reads.** The metric scores flowing prose the way standard readability tools do: `paragraph`, `list-item`, `blockquote`, `table.cell`, and `table.header` text counts; **headings are excluded**, and `code`, `frontmatter`, `html`, `comment`, `alt`, and `link` content is never counted. Every block that does not end in terminal punctuation ends a sentence — an unpunctuated list item is one sentence, not a fragment fused into its neighbors. Without that rule, a run of bullets scored as one enormous "sentence" and pushed Flesch reading ease far below zero; with it, scores line up with other readability tools within syllable-heuristic differences.
633
+ **What the score reads.** The metric scores flowing prose the way standard readability tools do: `paragraph`, `list-item`, `blockquote`, `table.cell`, and `table.header` text counts; **headings are excluded**, and `code`, `frontmatter`, `html`, `comment`, `alt`, and `link` content is never counted.
634
+ Every block that does not end in terminal punctuation ends a sentence — an unpunctuated list item is one sentence, not a fragment fused into its neighbors.
635
+ Without that rule, a run of bullets scored as one enormous "sentence" and pushed Flesch reading ease far below zero; with it, scores line up with other readability tools within syllable-heuristic differences.
548
636
 
549
- **Non-prose stripping.** Before scoring, each segment's text also has Markdoc tag-marker spans (`{% tag attr="x" %}`, `{% /tag %}`, and the `{%- ... -%}` trim variant) and backtick-delimited inline code spans stripped out — neither is readable prose, and both otherwise skew word/syllable counts. Prose between two block-tag markers still counts (only the marker spans themselves are removed); a paragraph consisting only of tag markers contributes nothing. Multi-backtick delimiters (`` ``like this`` ``) are handled conservatively as a simple open-run/close-run pair match, not a full CommonMark-correct implementation.
637
+ **Non-prose stripping.** Before scoring, each segment's text also has Markdoc tag-marker spans (`{% tag attr="x" %}`, `{% /tag %}`, and the `{%- ... -%}` trim variant) and backtick-delimited inline code spans stripped out — neither is readable prose, and both otherwise skew word/syllable counts.
638
+ Prose between two block-tag markers still counts (only the marker spans themselves are removed); a paragraph consisting only of tag markers contributes nothing.
639
+ Multi-backtick delimiters (`` ``like this`` ``) are handled conservatively as a simple open-run/close-run pair match, not a full CommonMark-correct implementation.
550
640
 
551
- **Detection-only** (not fixable) — there is no single edit that would "fix" a readability score. Reports at most **one** problem per file, always at `line: 1, column: 1` (there is no single source position a whole-document score belongs to) — never divided by zero: a file with no prose at all (empty, or only code/frontmatter) is never flagged, regardless of `min`/`max`.
641
+ **Detection-only** (not fixable) — there is no single edit that would "fix" a readability score.
642
+ Reports at most **one** problem per file, always at `line: 1, column: 1` (there is no single source position a whole-document score belongs to) — never divided by zero: a file with no prose at all (empty, or only code/frontmatter) is never flagged, regardless of `min`/`max`.
552
643
 
553
- The rule's `message` is substituted against up to **four** values, in this order: **1st = the formula name, 2nd = the computed score, 3rd = `min` (or `-∞` if unset), 4th = `max` (or `∞` if unset)** — e.g. the internal fallback `'Readability (%s) is %s; expected between %s and %s.'` → `'Readability (flesch-reading-ease) is 42.1; expected between 60 and ∞.'`. The `message` validation cap is per-assertion: a `metric` rule's `message` may use up to **4** `%s` placeholders (one per value above), while every other assertion stays capped at 2. Fewer placeholders than values is fine — substitution is positional, so a 2-slot message receives the leading values (formula name, then score).
644
+ The rule's `message` is substituted against up to **four** values, in this order: **1st = the formula name, 2nd = the computed score, 3rd = `min` (or `-∞` if unset), 4th = `max` (or `∞` if unset)** — e.g. the internal fallback `'Readability (%s) is %s; expected between %s and %s.'` → `'Readability (flesch-reading-ease) is 42.1; expected between 60 and ∞.'`.
645
+ The `message` validation cap is per-assertion: a `metric` rule's `message` may use up to **4** `%s` placeholders (one per value above), while every other assertion stays capped at 2.
646
+ Fewer placeholders than values is fine — substitution is positional, so a 2-slot message receives the leading values (formula name, then score).
554
647
 
555
648
  <!-- recheck-disable-next-line no-gerund-headings -->
556
649
  #### Spelling Assertions (`spelling`)
@@ -570,9 +663,11 @@ assertions:
570
663
  | `ignore` | `string[]` | No | Regex patterns; a token matching ANY of them is never flagged, e.g. `['\bAcme\w*']` to allow every inflection of a brand name. An invalid pattern is silently ignored, same convention as `pattern`'s `tokens`. |
571
664
  | `builtinVocabulary` | `boolean` | No | Default `true`. Whether [`TECHNICAL_PROPER_NOUNS`](#built-in-technical-proper-noun-vocabulary) is unioned into the accepted-word set alongside `vocab`. A multi-token entry (`Node.js`, `VS Code`) is split into its individual words, each accepted separately — correct for a per-word spell check, unlike `capitalization`'s whole-phrase matching. Set `false` for a closed vocabulary of only this rule's own `vocab`. |
572
665
 
573
- All options are optional — an empty `spelling: {}` is valid (default dictionary, no extra vocabulary, no ignore patterns, built-in vocabulary on). Unknown option keys, a non-string/empty-string `dictionary`, a `vocab`/`ignore` entry that isn't a non-empty string, or a non-boolean `builtinVocabulary` are all validation errors.
666
+ All options are optional — an empty `spelling: {}` is valid (default dictionary, no extra vocabulary, no ignore patterns, built-in vocabulary on).
667
+ Unknown option keys, a non-string/empty-string `dictionary`, a `vocab`/`ignore` entry that isn't a non-empty string, or a non-boolean `builtinVocabulary` are all validation errors.
574
668
 
575
- **Optional peer dependencies — install to enable.** `nspell` and its default dictionary (`dictionary-en`) are **optional peer dependencies**: installing `@redocly/recheck` itself pulls in **neither**. Enable `spelling` with:
669
+ **Optional peer dependencies — install to enable.** `nspell` and its default dictionary (`dictionary-en`) are **optional peer dependencies**: installing `@redocly/recheck` itself pulls in **neither**.
670
+ Enable `spelling` with:
576
671
 
577
672
  ```bash
578
673
  npm i nspell dictionary-en
@@ -586,34 +681,52 @@ npm i nspell
586
681
 
587
682
  If a config enables `spelling` without the required peer(s) installed, `recheck validate` fails with an actionable error naming the exact command above — never a bare `Cannot find module 'nspell'` surfacing for the first time at lint time.
588
683
 
589
- **Dictionaries load lazily.** Neither `nspell` nor `dictionary-en` is imported unless some rule in your config actually has a `spelling` assertion — a config without one never touches either package, at either `validate` or lint time. The loaded speller (including the ~500KB parsed dictionary) is cached per dictionary source for the process's lifetime, so every file/rule sharing the same `dictionary` (or the shared default) reuses one instance rather than reloading it per call.
684
+ **Dictionaries load lazily.** Neither `nspell` nor `dictionary-en` is imported unless some rule in your config actually has a `spelling` assertion — a config without one never touches either package, at either `validate` or lint time.
685
+ The loaded speller (including the ~500KB parsed dictionary) is cached per dictionary source for the process's lifetime, so every file/rule sharing the same `dictionary` (or the shared default) reuses one instance rather than reloading it per call.
590
686
 
591
- **Word tokenization.** Words are matched with `/\p{L}+(?:['’]\p{L}+)?/gu` — Unicode letter runs, with an optional apostrophe-joined suffix so contractions (`don't`, `it's`) tokenize as one word. A token is skipped (never checked) when it's in `vocab` (case-insensitively), matches any `ignore` pattern, is ALL-CAPS (2+ letters, e.g. an acronym) — matching the same ALL-CAPS carve-out `$title`/`$sentence` capitalization use — or is digit-adjacent (see below). Because `\p{L}` can never match a digit, a token touching one is never captured WHOLE by the tokenizer in the first place: a digit-adjacent identifier like `config2` still splits into a letter-only fragment (`config`) as its own regex match. Rather than checking that fragment like any other word, a digit-adjacency guard looks at the character immediately before and after each match and skips it when either neighbor is a digit — so common digit-bearing identifiers (`sha256` → `sha`, `utf8` → `utf`, `oauth2` → `oauth`, `es6` → `es`, `log4j` → both `log` and `j`, `2fast` → `fast`) are no longer flagged as false-positive misspellings. This mitigates, but doesn't eliminate, every false positive from the tokenizer's inability to capture digits at all — a token entirely surrounded by non-digit characters is still checked normally, so a genuine misspelling elsewhere in the same sentence is still flagged.
687
+ **Word tokenization.** Words are matched with `/\p{L}+(?:['’]\p{L}+)?/gu` — Unicode letter runs, with an optional apostrophe-joined suffix so contractions (`don't`, `it's`) tokenize as one word.
688
+ A token is skipped (never checked) when it's in `vocab` (case-insensitively), matches any `ignore` pattern, is ALL-CAPS (2+ letters, e.g. an acronym) — matching the same ALL-CAPS carve-out `$title`/`$sentence` capitalization use — or is digit-adjacent (see below).
689
+ Because `\p{L}` can never match a digit, a token touching one is never captured WHOLE by the tokenizer in the first place: a digit-adjacent identifier like `config2` still splits into a letter-only fragment (`config`) as its own regex match.
690
+ Rather than checking that fragment like any other word, a digit-adjacency guard looks at the character immediately before and after each match and skips it when either neighbor is a digit — so common digit-bearing identifiers (`sha256` → `sha`, `utf8` → `utf`, `oauth2` → `oauth`, `es6` → `es`, `log4j` → both `log` and `j`, `2fast` → `fast`) are no longer flagged as false-positive misspellings.
691
+ This mitigates, but doesn't eliminate, every false positive from the tokenizer's inability to capture digits at all — a token entirely surrounded by non-digit characters is still checked normally, so a genuine misspelling elsewhere in the same sentence is still flagged.
592
692
 
593
- **Code is never spell-checked, by construction of scope segmentation — not something this assertion special-cases.** A fenced or indented code block is its own `scope: 'code'` segment, entirely distinct from `paragraph`/`heading`/etc.; scoping `spelling` to prose (the common case, e.g. `scope: paragraph` or an array of prose scopes) means `ctx.segments` never contains one. A backtick-delimited **inline** code span, though, remains embedded as raw text inside a prose segment's own content (verified directly against the extractor) — those spans are masked out before tokenizing, the same length-preserving technique `capitalization`'s backtick-span freezing uses, so positions of any remaining flagged word stay exact. Scoping `spelling` to `all`/`raw` (or leaving `scope` at its default) checks the whole raw file, literal code included — same default-scope behavior every other native assertion (`swap`, `pattern`, ...) has.
693
+ **Code is never spell-checked, by construction of scope segmentation — not something this assertion special-cases.** A fenced or indented code block is its own `scope: 'code'` segment, entirely distinct from `paragraph`/`heading`/etc.; scoping `spelling` to prose (the common case, e.g. `scope: paragraph` or an array of prose scopes) means `ctx.segments` never contains one.
694
+ A backtick-delimited **inline** code span, though, remains embedded as raw text inside a prose segment's own content (verified directly against the extractor) — those spans are masked out before tokenizing, the same length-preserving technique `capitalization`'s backtick-span freezing uses, so positions of any remaining flagged word stay exact.
695
+ Scoping `spelling` to `all`/`raw` (or leaving `scope` at its default) checks the whole raw file, literal code included — same default-scope behavior every other native assertion (`swap`, `pattern`, ...) has.
594
696
 
595
- **Detection-only** — no `fix`. The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the unrecognized word, 2nd = a suggestion suffix** — either `''` (zero suggestions) or `' — did you mean: a, b, c?'` (one to three, comma-joined) — e.g. the internal fallback `'Unknown word "%s"%s'` → `'Unknown word "wrold" — did you mean: wold, world?'`.
697
+ **Detection-only** — no `fix`.
698
+ The rule's `message` gets two positional `%s` substitutions, in this order: **1st = the unrecognized word, 2nd = a suggestion suffix** — either `''` (zero suggestions) or `' — did you mean: a, b, c?'` (one to three, comma-joined) — e.g. the internal fallback `'Unknown word "%s"%s'` → `'Unknown word "wrold" — did you mean: wold, world?'`.
596
699
 
597
700
  #### Built-in technical proper-noun vocabulary
598
701
 
599
- `capitalization` and `spelling` both ship a built-in list of common technical/product proper nouns — `TECHNICAL_PROPER_NOUNS`, exported from `@redocly/recheck`'s public API (`import { TECHNICAL_PROPER_NOUNS } from '@redocly/recheck'`) so you can read or extend it yourself. It exists so a config that turns on sentence-case headings or spelling doesn't immediately need to hand-list the same 15+ mixed-case technology names every project already has to deal with (`OpenAPI`, `npm`, `Node.js`, `VS Code`, ...).
702
+ `capitalization` and `spelling` both ship a built-in list of common technical/product proper nouns — `TECHNICAL_PROPER_NOUNS`, exported from `@redocly/recheck`'s public API (`import { TECHNICAL_PROPER_NOUNS } from '@redocly/recheck'`) so you can read or extend it yourself.
703
+ It exists so a config that turns on sentence-case headings or spelling doesn't immediately need to hand-list the same 15+ mixed-case technology names every project already has to deal with (`OpenAPI`, `npm`, `Node.js`, `VS Code`, ...).
600
704
 
601
705
  **On by default**, per rule:
602
706
  - `capitalization` unions it into `exceptions` (so a listed name keeps its as-written casing under every `$`-style, including `$sentence`).
603
707
  - `spelling` unions it into `vocab` (so those words are never reported as misspellings), splitting any multi-token entry into its individual words first — a per-word spell check has no way to accept a whole phrase atomically the way `capitalization`'s phrase matching does.
604
708
  - Either union is opted out of independently with that rule's own `builtinVocabulary: false`, restoring strict pre-built-in behavior (a closed vocabulary of only what you list yourself).
605
- - Your own `exceptions`/`vocab` on the same rule **compose** with the built-ins rather than replacing them — unlike a preset-shipped list on the same rule key, which a same-key override *would* replace entirely (see [`extends` presets](#extends-presets) above). This is exactly how [`recheck/prose`](#extends-presets)'s `capitalization` rule gets its protection for common technical nouns without shipping any `exceptions` of its own.
709
+ - Your own `exceptions`/`vocab` on the same rule **compose** with the built-ins rather than replacing them — unlike a preset-shipped list on the same rule key, which a same-key override *would* replace entirely (see [`extends` presets](#extends-presets) above).
710
+ This is exactly how [`recheck/prose`](#extends-presets)'s `capitalization` rule gets its protection for common technical nouns without shipping any `exceptions` of its own.
606
711
 
607
712
  **Multi-token entries work.** An entry containing a dot or whitespace (`Node.js`, `VS Code`, `Visual Studio Code`, `GitHub Actions`, `Google Cloud`, `Azure DevOps`) is matched by `capitalization` as a whole phrase against the segment text (longest-match-first, case-insensitive but otherwise literal) and preserved verbatim — not looked up per word, which is what a single-token entry like `GitHub` still gets.
608
713
 
609
- **Inclusion bar** (why an entry is — or isn't — in the list, and the bar to clear before proposing one): an entry qualifies if it's an unambiguous technology, product, or company name whose exception listing wouldn't *weaken* capitalization/spelling checks — concretely, its lowercase form must not be a legitimate English word in its own right. That covers ordinary Title-Case brand names (`Android`, `Kubernetes`, `Redocly`) just as much as entries with an internal capital (`OpenAPI`, `GraphQL`), a dot (`Node.js`), or forced lowercase (`npm`) — `$sentence` lowercases every non-first word regardless of how "ordinary" its casing looks, so plain Title-Case names need protection too. Excluded, deliberately:
610
- - **Pure ALL-CAPS acronyms** (`JWT`, `YAML`) — already handled structurally by the ALL-CAPS carve-out both `capitalization` and `spelling` apply, so listing them adds maintenance for no behavior change. Note this is narrower than "looks like an acronym": `OAuth` and `AsyncAPI` are mixed-case, not pure ALL-CAPS, and are in the list.
611
- - **Terms with legitimate lowercase prose usage** — generic English (`cloud`, `apps`), words that are ALSO ordinary English words even though they're Redocly product names too (`Realm`, `Replay`, `Respect` — listing them would force-capitalize ordinary usage like "we respect your privacy"; `Node` — the common technical noun, superseded by the `Node.js` phrase entry for the platform specifically), and — caught by a later audit, not the original pass — ordinary brand-shaped words with a real dictionary meaning (`Chrome`, `Markdown`, `Postman`, `Prettier`, `Safari`, `Swagger`, `Windows`; see `src/data/proper-nouns.ts`'s header for each one's disqualifying lowercase usage). A few real dictionary words (`Android`, `Docker`, `TypeScript`) were judged rare enough in ordinary lowercase usage to keep anyway — a documented, deliberate risk-acceptance, not an oversight. List your own such names in your rule's own `exceptions`/`vocab`, which compose with this list as described above.
714
+ **Inclusion bar** (why an entry is — or isn't — in the list, and the bar to clear before proposing one): an entry qualifies if it's an unambiguous technology, product, or company name whose exception listing wouldn't *weaken* capitalization/spelling checks — concretely, its lowercase form must not be a legitimate English word in its own right.
715
+ That covers ordinary Title-Case brand names (`Android`, `Kubernetes`, `Redocly`) just as much as entries with an internal capital (`OpenAPI`, `GraphQL`), a dot (`Node.js`), or forced lowercase (`npm`) — `$sentence` lowercases every non-first word regardless of how "ordinary" its casing looks, so plain Title-Case names need protection too.
716
+ Excluded, deliberately:
717
+ - **Pure ALL-CAPS acronyms** (`JWT`, `YAML`) — already handled structurally by the ALL-CAPS carve-out both `capitalization` and `spelling` apply, so listing them adds maintenance for no behavior change.
718
+ Note this is narrower than "looks like an acronym": `OAuth` and `AsyncAPI` are mixed-case, not pure ALL-CAPS, and are in the list.
719
+ - **Terms with legitimate lowercase prose usage** — generic English (`cloud`, `apps`), words that are ALSO ordinary English words even though they're Redocly product names too (`Realm`, `Replay`, `Respect` — listing them would force-capitalize ordinary usage like "we respect your privacy"; `Node` — the common technical noun, superseded by the `Node.js` phrase entry for the platform specifically), and — caught by a later audit, not the original pass — ordinary brand-shaped words with a real dictionary meaning (`Chrome`, `Markdown`, `Postman`, `Prettier`, `Safari`, `Swagger`, `Windows`; see `src/data/proper-nouns.ts`'s header for each one's disqualifying lowercase usage).
720
+ A few real dictionary words (`Android`, `Docker`, `TypeScript`) were judged rare enough in ordinary lowercase usage to keep anyway — a documented, deliberate risk-acceptance, not an oversight.
721
+ List your own such names in your rule's own `exceptions`/`vocab`, which compose with this list as described above.
612
722
 
613
- Two automated tests in `src/data/__tests__/proper-nouns.test.ts` enforce this: one checks every entry's shape against the bar above — no pure ALL-CAPS, and, mechanically, no single-token entry whose lowercase form the REAL spelling dictionary (`dictionary-en`/`nspell`, the same pair `spelling` loads at runtime) accepts as a legitimate English word, unless it's named in an explicit accepted-risk allowlist — plus alphabetization and no duplicates. A round-trip guard separately drives every entry through the real `capitalization` and `spelling` rules and fails the suite if any entry can't actually be protected — the list can't silently regress into decoration.
723
+ Two automated tests in `src/data/__tests__/proper-nouns.test.ts` enforce this: one checks every entry's shape against the bar above — no pure ALL-CAPS, and, mechanically, no single-token entry whose lowercase form the REAL spelling dictionary (`dictionary-en`/`nspell`, the same pair `spelling` loads at runtime) accepts as a legitimate English word, unless it's named in an explicit accepted-risk allowlist — plus alphabetization and no duplicates.
724
+ A round-trip guard separately drives every entry through the real `capitalization` and `spelling` rules and fails the suite if any entry can't actually be protected — the list can't silently regress into decoration.
614
725
 
615
726
  #### Length Assertions (`length`)
616
- Recheck-original, detection-only check: measures each scoped segment's size — in characters, words, or sentences — and flags a segment whose measurement falls outside `[min, max]`. Unlike `metric` (always whole-document), `length` honors whatever `scope` the rule configures — e.g. `scope: alt` to cap image alt text, or `scope: sentence` to cap sentence length in words. [`recheck/google`](#extends-presets) ships this for the guide's stated "fewer than 26 words per sentence" limit (`google/sentence-length`); Microsoft's 150-character alt-text cap is the other published example of this shape.
727
+ Recheck-original, detection-only check: measures each scoped segment's size — in characters, words, or sentences — and flags a segment whose measurement falls outside `[min, max]`.
728
+ Unlike `metric` (always whole-document), `length` honors whatever `scope` the rule configures — e.g. `scope: alt` to cap image alt text, or `scope: sentence` to cap sentence length in words.
729
+ [`recheck/google`](#extends-presets) ships this for the guide's stated "fewer than 26 words per sentence" limit (`google/sentence-length`); Microsoft's 150-character alt-text cap is the other published example of this shape.
617
730
 
618
731
  ```yaml
619
732
  assertions:
@@ -628,11 +741,14 @@ assertions:
628
741
  | `min` | `number` | At least one of `min`/`max` | Minimum allowed size; a smaller segment is a violation. |
629
742
  | `max` | `number` | At least one of `min`/`max` | Maximum allowed size; a larger segment is a violation. |
630
743
 
631
- Omitting both `min` and `max`, a missing/unrecognized `unit`, or an unknown option key are all validation errors — same reasoning as `occurrence`/`metric` above. An inverted range (`min` > `max`) is also an error.
744
+ Omitting both `min` and `max`, a missing/unrecognized `unit`, or an unknown option key are all validation errors — same reasoning as `occurrence`/`metric` above.
745
+ An inverted range (`min` > `max`) is also an error.
632
746
 
633
- **Detection-only** (not fixable) — there is no single edit that would resize a segment to fit. Reports at most one problem per flagged segment, at the segment's own `startLine`/`startColumn`.
747
+ **Detection-only** (not fixable) — there is no single edit that would resize a segment to fit.
748
+ Reports at most one problem per flagged segment, at the segment's own `startLine`/`startColumn`.
634
749
 
635
- The rule's `message` gets three positional `%s` substitutions, in this order: **1st = the segment's measured size, 2nd = the unit name, 3rd = the bound that was violated** (`min` or `max`, whichever applied), e.g. the internal fallback `'Segment is %s %s; at most %s allowed'` → `'Segment is 151 characters; at most 150 allowed'`. The message validation cap for `length` is **3** placeholders (one per value above), same reasoning as `metric`'s 4-cap.
750
+ The rule's `message` gets three positional `%s` substitutions, in this order: **1st = the segment's measured size, 2nd = the unit name, 3rd = the bound that was violated** (`min` or `max`, whichever applied), e.g. the internal fallback `'Segment is %s %s; at most %s allowed'` → `'Segment is 151 characters; at most 150 allowed'`.
751
+ The message validation cap for `length` is **3** placeholders (one per value above), same reasoning as `metric`'s 4-cap.
636
752
 
637
753
  #### Built-in Prose Assertions
638
754
  Beyond `swap`, `pattern`, `occurrence`, `repetition`, `consistency`, `conditional`, `capitalization`, `metric`, `spelling`, and `length` above, Recheck ships a small set of native prose/format checks:
@@ -644,16 +760,19 @@ Three of the assertions above (`repetition`, `consistency`, `capitalization`) ar
644
760
  ### Recheck-original structural rules
645
761
 
646
762
  Seven rules have no markdownlint counterpart, so they sit outside the 53-rule parity set
647
- (and outside the parity comparison). All seven are **detection-only** (`fix: false`). The
763
+ (and outside the parity comparison).
764
+ All seven are **detection-only** (`fix: false`).
765
+ The
648
766
  canonical list is `RECHECK_ORIGINAL_TOKEN_RULE_NAMES` in `src/rules/token/index.ts`.
649
767
 
650
- The table below covers five of them. The other two — `markdoc-unknown-tag` and
768
+ The table below covers five of them.
769
+ The other two — `markdoc-unknown-tag` and
651
770
  `markdoc-attributes` — need a tag schema to check anything, so they are documented with
652
771
  the [`recheck/markdoc`](#extends-presets) preset instead.
653
772
 
654
773
  | Rule | Flags | Why |
655
774
  |---|---|---|
656
- | `no-empty-headings` | A heading whose text content is empty (`# `, or markup that renders to nothing such as `## <span></span>`) | An empty heading still lands in the document outline and in screen-reader heading navigation. Inline code counts as content, so `` # `config.yaml` `` is fine. |
775
+ | `no-empty-headings` | A heading whose text content is empty (a bare `#`, or markup that renders to nothing such as `## <span></span>`) | An empty heading still lands in the document outline and in screen-reader heading navigation. Inline code counts as content, so `` # `config.yaml` `` is fine. |
657
776
  | `no-duplicate-link-destinations` | The second and later links to one destination when the link **text** differs from the first occurrence's | Screen-reader users listing a page's links hear one target described inconsistently; the texts also drift apart over time. Repeating the *same* text for the same destination is ordinary prose and is not flagged. Resolves reference links through their definition. |
658
777
  | `list-length` | A list (ordered or unordered) with fewer than `min` items (default 2) or more than `max` items (no default — unbounded unless set) | A single-item list usually reads better as a plain sentence, and a very long list asks readers to hold too many parallel items in mind. Every list is evaluated independently, including nested sublists — a short sublist is flagged even when its parent list is long enough. |
659
778
  | `markdoc-syntax` | A grammar-level Markdoc tag error — a malformed span, an unquoted "bareword" attribute/primary value, or a close tag carrying attributes | These are invalid under real Markdoc's own grammar regardless of any tag schema, so the rule fires on custom/unknown tags and under `schema: false` alike. See the [`recheck/markdoc`](#extends-presets) preset bullet below for the full behavior and a config example. |
@@ -664,10 +783,12 @@ individually as shown below.
664
783
 
665
784
  `markdoc-syntax` and `markdoc-pairing` work the other way around: they ship only inside
666
785
  the [`recheck/markdoc`](#extends-presets) preset, and both need `markdoc: true` (or the
667
- object form) to ever see a Markdoc tag token. Naming either rule key on its own, without
786
+ object form) to ever see a Markdoc tag token.
787
+ Naming either rule key on its own, without
668
788
  the flag, validates but can never report anything — and you get no warning about it,
669
789
  because the stale-config warning fires on `extends` containing `"recheck/markdoc"`
670
- (`warnStaleMarkdocPreset` in `config/validate.ts`), not on individual rule keys. The
790
+ (`warnStaleMarkdocPreset` in `config/validate.ts`), not on individual rule keys.
791
+ The
671
792
  preset bullet below covers both rules' full behavior with the flag on, plus a config
672
793
  example.
673
794
 
@@ -691,7 +812,8 @@ recheck/list-length:
691
812
  list-length: { min: 2, max: 10 }
692
813
  ```
693
814
 
694
- For markdown structure/format rules (headings, lists, links, tables, whitespace, and 49 more), see [Markdownlint parity](#markdownlint-parity) below — `no-trailing-spaces`, `no-hard-tabs`, `line-length`, `ul-style` (bullet style), `no-duplicate-heading`, and `link-fragments` are all part of that 53-rule set, not this native list. (`max-line-length`, `bullet-style`, `no-duplicate-headings`, and `no-broken-fragment-links` were pre-parity native ids for those same rules; they were removed rather than kept as aliases — see [Migrating from markdownlint](#migrating-from-markdownlint).)
815
+ For markdown structure/format rules (headings, lists, links, tables, whitespace, and 49 more), see [Markdownlint parity](#markdownlint-parity) below — `no-trailing-spaces`, `no-hard-tabs`, `line-length`, `ul-style` (bullet style), `no-duplicate-heading`, and `link-fragments` are all part of that 53-rule set, not this native list.
816
+ (`max-line-length`, `bullet-style`, `no-duplicate-headings`, and `no-broken-fragment-links` were pre-parity native ids for those same rules; they were removed rather than kept as aliases — see [Migrate from markdownlint](#migrate-from-markdownlint).)
695
817
 
696
818
  ### Enhanced Scope Support
697
819
 
@@ -731,11 +853,14 @@ Opt-in — off by default, since Liquid/Jinja templates use the same `{% %}` del
731
853
  markdoc: true # shorthand for `{ schema: 'realm' }`
732
854
  ```
733
855
 
734
- Writing *about* Markdoc syntax rather than using it (docs like this one, a tutorial, a changelog entry)? Wrap the literal `{% ... %}` in a code span — `` `{% partial /%}` `` — instead of leaving it bare in prose. Code spans never tokenize as Markdoc tags whether the flag is on or off, so that's the escape hatch.
856
+ Writing *about* Markdoc syntax rather than using it (docs like this one, a tutorial, a changelog entry)?
857
+ Wrap the literal `{% ... %}` in a code span — `` `{% partial /%}` `` — instead of leaving it bare in prose.
858
+ Code spans never tokenize as Markdoc tags whether the flag is on or off, so that's the escape hatch.
735
859
 
736
860
  #### Object form: choosing or extending the tag schema
737
861
 
738
- `markdoc: true` is shorthand for the common case. The object form adds two things the
862
+ `markdoc: true` is shorthand for the common case.
863
+ The object form adds two things the
739
864
  boolean can't express: turning the schema-aware checks off while keeping tag tokenization,
740
865
  and layering a project's own custom tags over the built-in schema.
741
866
 
@@ -754,27 +879,35 @@ markdoc:
754
879
  ```
755
880
 
756
881
  - **`schema: realm`** — the same built-in schema `markdoc: true` uses: `@markdoc/markdoc`'s
757
- own built-in tags composed with `@redocly/theme`'s tag definitions. It's generated from a
882
+ own built-in tags composed with `@redocly/theme`'s tag definitions.
883
+ It's generated from a
758
884
  theme build rather than hand-written, and a test fails if it drifts out of sync (see
759
- [CONTRIBUTING.md](CONTRIBUTING.md) for the regeneration command). This is what most
885
+ [CONTRIBUTING.md](CONTRIBUTING.md) for the regeneration command).
886
+ This is what most
760
887
  projects want, and what the four [`recheck/markdoc`](#extends-presets) rules validate
761
888
  against by default.
762
889
  - **`schema: false`** — tokenization and tag **pairing** still run, so `markdoc.tag` scope,
763
890
  prose-scope exclusion, fix protection, and `markdoc-syntax`/`markdoc-pairing`'s
764
- grammar-level checks all still work. Only the two schema-dependent rules
891
+ grammar-level checks all still work.
892
+ Only the two schema-dependent rules
765
893
  (`markdoc-unknown-tag`, `markdoc-attributes`) go inert, since there's no schema left for
766
- "unknown tag" or "missing required attribute" to mean anything against. Use this if you
894
+ "unknown tag" or "missing required attribute" to mean anything against.
895
+ Use this if you
767
896
  write Markdoc tags but don't have (or don't want) a schema to validate them against.
768
- - **`extend.tags`** — merges your own tag definitions over the base schema. On a name
897
+ - **`extend.tags`** — merges your own tag definitions over the base schema.
898
+ On a name
769
899
  collision the merge is a **whole-tag replace**, matching how Markdoc's own config
770
- composition works, not a per-attribute deep merge. Declare your project's custom tags
900
+ composition works, not a per-attribute deep merge.
901
+ Declare your project's custom tags
771
902
  here (for example, a docs site's own `@theme/markdoc/schema.ts` overrides) so
772
903
  `markdoc-unknown-tag` and `markdoc-attributes` validate against your real tag surface
773
- instead of flagging every custom tag as unknown. Under `schema: false` there is no base
904
+ instead of flagging every custom tag as unknown.
905
+ Under `schema: false` there is no base
774
906
  to merge over, so `extend` does nothing.
775
907
  - **`extend.tagsFile`** — the same tag-definition surface as `extend.tags`, but sourced from
776
- a separate YAML file instead of written inline into `recheck.yaml`. This is the shape
777
- [`recheck markdoc-schema`](#generating-a-tagsfile-recheck-markdoc-schema) below generates,
908
+ a separate YAML file instead of written inline into `recheck.yaml`.
909
+ This is the shape
910
+ [`recheck markdoc-schema`](#generate-a-tagsfile-recheck-markdoc-schema) below generates,
778
911
  so a project with tags defined in TypeScript (a `@theme/markdoc/schema.ts` module, say)
779
912
  never hand-transcribes them into YAML.
780
913
  ```yaml
@@ -789,7 +922,8 @@ markdoc:
789
922
  wherever `recheck` is invoked from.
790
923
  - **Precedence**: tags merge in the order built-in schema → `tagsFile` → inline
791
924
  `extend.tags`, each layer a whole-tag replace on a name collision (same rule as
792
- `extend.tags` above). `tags` and `tagsFile` can both be set on the same `extend` block;
925
+ `extend.tags` above).
926
+ `tags` and `tagsFile` can both be set on the same `extend` block;
793
927
  `extend` with neither key is rejected by config validation as a likely no-op.
794
928
  - **Errors are fatal to the whole run, not a silent markdoc downgrade.** A `tagsFile` that
795
929
  doesn't exist, isn't valid YAML, isn't a YAML map, or contains a tag entry with an
@@ -800,15 +934,21 @@ markdoc:
800
934
 
801
935
  Turning `markdoc` on (either form) changes how every prose rule sees a Markdoc tag, not just `markdoc.tag` (above):
802
936
 
803
- - **Prose scopes exclude the tag itself.** `paragraph`, `heading`, `list-item`, `blockquote`, and `table.header`/`table.cell` all blank a tag's own `{% ... %}` span out of their content before any rule runs — a `swap`/`pattern`/`capitalization` match can't fire on the tag's syntax, and a `length`/`metric` count doesn't include it. The blanking is position-preserving (same-width spaces, never a deletion), so real text on either side of a tag keeps its exact line and column.
937
+ - **Prose scopes exclude the tag itself.** `paragraph`, `heading`, `list-item`, `blockquote`, and `table.header`/`table.cell` all blank a tag's own `{% ... %}` span out of their content before any rule runs — a `swap`/`pattern`/`capitalization` match can't fire on the tag's syntax, and a `length`/`metric` count doesn't include it.
938
+ The blanking is position-preserving (same-width spaces, never a deletion), so real text on either side of a tag keeps its exact line and column.
804
939
  - **A segment with no prose left isn't emitted at all.** A heading or table cell whose entire text IS a tag (`# {% #anchor %}`) produces no `heading.h1`/`table.cell` segment — there's nothing for a heading or cell rule to check, so none fires on it.
805
940
  - **`--fix` never rewrites a Markdoc tag's bytes.** Every proposed fix is checked against the document's tag spans before it's applied: one that doesn't touch a tag goes through untouched, one that fully covers a tag with a same-length replacement gets the tag spliced back in, and anything that would change a tag's length or split it in half is withheld instead — a withheld fix is reported (`skippedFixes` in the [Library API](#library-api)), not silently swallowed.
806
941
  - **Two CommonMark constructs Markdoc doesn't have stop being recognized.** Markdoc's own tokenizer disables indented code blocks and setext headings (the `Title\n===\n` underline form) unconditionally, which is how Realm renders, so `markdoc: true` disables them too — and only while the flag is on:
807
- - A 4+-space-indented block that would otherwise be an indented code block parses as ordinary content instead: a paragraph, list, or fence, whichever the un-indented text would have produced. This shows up most with a block-positioned tag followed immediately by more indented lines, and with genuinely indented example text. Realm renders both as prose, so matching that is the intent.
808
- - A text line immediately followed by a `---`/`===` line, with no blank line between, no longer forms a heading. This also fixes a common false positive: a tag on its own line (`{% table %}`, say) directly followed by a `---` line is ordinary Markdoc table-row syntax, but without the flag it reads as a setext heading whose text is the tag itself. With the flag on, the tag is its own token and can't merge into a paragraph that `---` would complete.
809
- - Practical effect: on documents using either construct, expect `heading-style`, `blanks-around-headings`, `capitalization`, and `code-block-style` findings to shift when you first turn the flag on. They are moving to match how Markdoc actually renders, not regressing.
942
+ - A 4+-space-indented block that would otherwise be an indented code block parses as ordinary content instead: a paragraph, list, or fence, whichever the un-indented text would have produced.
943
+ This shows up most with a block-positioned tag followed immediately by more indented lines, and with genuinely indented example text.
944
+ Realm renders both as prose, so matching that is the intent.
945
+ - A text line immediately followed by a `---`/`===` line, with no blank line between, no longer forms a heading.
946
+ This also fixes a common false positive: a tag on its own line (`{% table %}`, say) directly followed by a `---` line is ordinary Markdoc table-row syntax, but without the flag it reads as a setext heading whose text is the tag itself.
947
+ With the flag on, the tag is its own token and can't merge into a paragraph that `---` would complete.
948
+ - Practical effect: on documents using either construct, expect `heading-style`, `blanks-around-headings`, `capitalization`, and `code-block-style` findings to shift when you first turn the flag on.
949
+ They are moving to match how Markdoc actually renders, not regressing.
810
950
 
811
- ### Generating a tagsFile: `recheck markdoc-schema`
951
+ ### Generate a tagsFile: `recheck markdoc-schema`
812
952
 
813
953
  Projects that define their own Markdoc tags in TypeScript — a `@theme/markdoc/schema.ts`
814
954
  module exporting a `tags` map, the shape both `docs/realm` and `docs/intranet` use in this
@@ -820,20 +960,24 @@ recheck markdoc-schema --from path/to/schema.ts --out recheck-markdoc-tags.yaml
820
960
  ```
821
961
 
822
962
  - **`--from <path>`** (repeatable) — a project schema module to extract tags from, resolved
823
- relative to the current working directory. The module must export `tags` (named or on a
963
+ relative to the current working directory.
964
+ The module must export `tags` (named or on a
824
965
  `default` object) mapping tag name to a Markdoc tag config; only the statically-checkable
825
966
  facets (`selfClosing`, and each attribute's `type`/`required`/`default`/`enum`) are
826
967
  extracted — anything richer (a custom attribute class, a `validate()` function) is written
827
- out as `dynamic: true`, the same reduction the built-in `realm` schema goes through. Pass
968
+ out as `dynamic: true`, the same reduction the built-in `realm` schema goes through.
969
+ Pass
828
970
  `--from` more than once to merge several modules; an identical tag definition repeated
829
971
  across modules is fine, but two modules disagreeing about the same tag's shape fails the
830
972
  command rather than letting flag order silently pick one.
831
973
  - **`--out <path>`** — where to write the generated YAML, resolved relative to the current
832
- working directory. The file opens with a generated-file header naming its source module(s)
974
+ working directory.
975
+ The file opens with a generated-file header naming its source module(s)
833
976
  and the exact command to regenerate it.
834
977
  - **`--check`** — verifies the output file matches what a fresh generation would produce,
835
978
  without writing it: exits `0` and prints `<path> is up to date.` when it matches, exits `1`
836
- and prints a one-line diagnosis (file missing, or stale) otherwise. This is what a CI drift
979
+ and prints a one-line diagnosis (file missing, or stale) otherwise.
980
+ This is what a CI drift
837
981
  check should call — see this repo's own wiring below.
838
982
 
839
983
  **TypeScript sources need a loader.** `recheck markdoc-schema` dynamic-`import()`s each
@@ -846,7 +990,8 @@ node_modules/.bin/recheck markdoc-schema … (Cannot find module '<a module your
846
990
  imports>' imported from '<path to your schema.ts>')
847
991
  ```
848
992
 
849
- (wrapped above for line length; the real message is one line. The parenthetical is Node's
993
+ (wrapped above for line length; the real message is one line.
994
+ The parenthetical is Node's
850
995
  own error and its shape varies: for a schema whose extensionless internal imports plain
851
996
  `node` cannot resolve — the common case — the "imported from" path is your schema file
852
997
  itself; for a `--from` path that doesn't exist at all it is recheck's own command module.)
@@ -859,7 +1004,8 @@ works under either.
859
1004
  long-term source of truth: [issue #25666](https://github.com/Redocly/redocly/issues/25666)
860
1005
  tracks Realm itself emitting one canonical, statics-only Markdoc tag/schema manifest, which
861
1006
  would let this generator (and its drift check) retire in favor of reading that manifest
862
- directly. Until then, `recheck markdoc-schema` is the supported way to keep a project's
1007
+ directly.
1008
+ Until then, `recheck markdoc-schema` is the supported way to keep a project's
863
1009
  `tagsFile` in sync with its real tag schema modules.
864
1010
 
865
1011
  #### Worked example: this repo's own setup
@@ -896,7 +1042,8 @@ equivalent but isn't: the script itself already ends in `pnpm --filter @redocly/
896
1042
  tsx dist/cli.js markdoc-schema …`, and pnpm's own `--` forwarding through that nested `exec`
897
1043
  makes yargs read `--check` as a positional argument instead of the `--check` flag — the
898
1044
  command then silently regenerates the file and always exits `0`, defeating the whole point
899
- of a drift check. Always call it as `pnpm run recheck:markdoc-tags --check`, with no extra
1045
+ of a drift check.
1046
+ Always call it as `pnpm run recheck:markdoc-tags --check`, with no extra
900
1047
  `--`.
901
1048
 
902
1049
  CI runs exactly that check on every PR, as its own step in
@@ -921,10 +1068,12 @@ Rules can be configured with different severity levels:
921
1068
  ### Auto-Fix Safety
922
1069
 
923
1070
  Fixability is declared by each assertion, not by config — a rule can be automatically corrected if and only if its assertion implements a `fix()`.
924
- The `autoFixable` config key was removed — rules declare fixability; use `fix: false` to opt out. A config that still sets `autoFixable` now fails validation with an unknown-property error.
1071
+ The `autoFixable` config key was removed — rules declare fixability; use `fix: false` to opt out.
1072
+ A config that still sets `autoFixable` now fails validation with an unknown-property error.
925
1073
  To opt a rule out of auto-fixing, set `fix: false` on it instead.
926
1074
 
927
- The `enabled` config key was likewise removed — it was schema-legal but never actually consulted by the engine (use `severity: off` to disable a rule). A config that still sets `enabled` now fails validation with an unknown-property error.
1075
+ The `enabled` config key was likewise removed — it was schema-legal but never actually consulted by the engine (use `severity: off` to disable a rule).
1076
+ A config that still sets `enabled` now fails validation with an unknown-property error.
928
1077
 
929
1078
  - ✅ **Fixable native assertions**: `swap`, `semantic-line-breaks`, `repetition`, `consistency`, `capitalization` (except its custom-regex `match` mode, which is always detection-only)
930
1079
  - ❌ **Not fixable native assertions**: `pattern`, `max-image-size`, `occurrence`, `conditional`, `metric`, `spelling`
@@ -932,7 +1081,8 @@ The `enabled` config key was likewise removed — it was schema-legal but never
932
1081
 
933
1082
  ## Markdownlint parity
934
1083
 
935
- Recheck ports all 53 of [markdownlint](https://github.com/DavidAnson/markdownlint)'s built-in rules (MD001-MD060, minus retired ids) as native `assertions`, verified against upstream by a differential parity harness (see [Parity with markdownlint](#parity-with-markdownlint) below). Enable the full set with one line:
1084
+ Recheck ports all 53 of [markdownlint](https://github.com/DavidAnson/markdownlint)'s built-in rules (MD001-MD060, minus retired ids) as native `assertions`, verified against upstream by a differential parity harness (see [Parity with markdownlint](#parity-with-markdownlint) below).
1085
+ Enable the full set with one line:
936
1086
 
937
1087
  ```yaml
938
1088
  extends: [recheck/markdown]
@@ -1009,25 +1159,55 @@ Generated from the built rule registry — `dist/rules/token/index.js`'s `allTok
1009
1159
 
1010
1160
  ### `extends` presets
1011
1161
 
1012
- Recheck ships nine built-in presets, referenced by id under `extends:`. Presets are applied in listed order, then your own rule keys are merged on top — **your config always wins**: a rule key you define overrides the same key from a preset, and per-assertion options you set override just that assertion's preset options (other preset options for the same rule are preserved).
1162
+ Recheck ships nine built-in presets, referenced by id under `extends:`.
1163
+ Presets are applied in listed order, then your own rule keys are merged on top — **your config always wins**: a rule key you define overrides the same key from a preset, and per-assertion options you set override just that assertion's preset options (other preset options for the same rule are preserved).
1013
1164
 
1014
- There is one exception. A rule may attach a milder severity to some of its own reports, and your config can't escalate those: `recheck/markdoc-attributes` reports unknown attributes at `warn` no matter what severity you give the rule (see the `recheck/markdoc` bullet below). `severity: 'off'` still works as expected — it disables the rule entirely, so no reports of any severity.
1165
+ There is one exception.
1166
+ A rule may attach a milder severity to some of its own reports, and your config can't escalate those: `recheck/markdoc-attributes` reports unknown attributes at `warn` no matter what severity you give the rule (see the `recheck/markdoc` bullet below).
1167
+ `severity: 'off'` still works as expected — it disables the rule entirely, so no reports of any severity.
1015
1168
 
1016
- - **`recheck/markdown`** — the full 53-rule set from the table above, all at `severity: error` with upstream-faithful default options. Equivalent to markdownlint's `{ default: true }`.
1169
+ - **`recheck/markdown`** — the full 53-rule set from the table above, all at `severity: error` with upstream-faithful default options.
1170
+ Equivalent to markdownlint's `{ default: true }`.
1017
1171
  - **`recheck/markdown-relaxed`** — mirrors markdownlint's own `style/relaxed.json`: the same 53 rules, with `no-trailing-spaces`, `no-hard-tabs`, `no-multiple-blanks`, `no-multiple-space-blockquote`, `no-blanks-blockquote`, `line-length`, `ul-indent`, `no-inline-html`, `no-bare-urls`, `fenced-code-language`, and `first-line-h1` turned off.
1018
1172
  - **`recheck/minimal`** — a small, high-signal set: `no-trailing-spaces`, `no-hard-tabs`, `single-trailing-newline`, `no-reversed-links`, `no-empty-links`.
1019
- - **`recheck/prose`** — Recheck's Vale-parity starter set, all at `severity: warn`: `repetition` (default options), `consistency` (one US spelling enforced file-wide for `behavior`/`color`/`license`/`organize` vs. their British spellings, matched with `ignoreCase: true` so a capitalized, sentence-initial variant like `Colour` still counts), and `capitalization` (`$sentence`, `scope: heading` only, `fix: false`, no preset-level `exceptions` — see below). All three are scoped to prose segments — `repetition` and `consistency` to `summary` (the document's prose: paragraph, heading, list-item, blockquote, and table-cell text), `capitalization` to headings — so the preset never flags (and `--fix` never rewrites) code samples or frontmatter. `extends: [recheck/markdown, recheck/prose]` is the one-liner that replaces a markdownlint + Vale combo. See [Opt-in prose assertions](#opt-in-prose-assertions) below for three more prose assertions that exist but are deliberately **not** in this preset.
1020
- - **`recheck/markdoc`** — four Recheck-original rules that check Markdoc tag syntax itself (`{% tag attr="value" %}`) rather than prose or markdownlint parity. All four are `fix: false`.
1173
+ - **`recheck/prose`** — Recheck's Vale-parity starter set, all at `severity: warn`: `repetition` (default options), `consistency` (one US spelling enforced file-wide for `behavior`/`color`/`license`/`organize` vs. their British spellings, matched with `ignoreCase: true` so a capitalized, sentence-initial variant like `Colour` still counts), and `capitalization` (`$sentence`, `scope: heading` only, `fix: false`, no preset-level `exceptions` — see below).
1174
+ All three are scoped to prose segments — `repetition` and `consistency` to `summary` (the document's prose: paragraph, heading, list-item, blockquote, and table-cell text), `capitalization` to headings — so the preset never flags (and `--fix` never rewrites) code samples or frontmatter.
1175
+ `extends: [recheck/markdown, recheck/prose]` is the one-liner that replaces a markdownlint + Vale combo.
1176
+ See [Opt-in prose assertions](#opt-in-prose-assertions) below for three more prose assertions that exist but are deliberately **not** in this preset.
1177
+ - **`recheck/markdoc`** — four Recheck-original rules that check Markdoc tag syntax itself (`{% tag attr="value" %}`) rather than prose or markdownlint parity.
1178
+ All four are `fix: false`.
1021
1179
  - `recheck/markdoc-syntax` (`error`) — malformed spans, unquoted "bareword" values, and close tags carrying attributes.
1022
1180
  - `recheck/markdoc-pairing` (`error`) — unclosed, orphaned, or crossed tag pairs, and a self-closing tag written without a slash or given a close tag it shouldn't have.
1023
1181
  - `recheck/markdoc-unknown-tag` (`warn`, because custom tags are common) — a tag name the schema doesn't declare.
1024
- - `recheck/markdoc-attributes` (mixed) — a missing required attribute, an enum or type violation, and a duplicate attribute report at `error`; an unknown attribute name, whether named or a stray positional value, always reports at `warn`. That `warn` is set per report by the rule itself, so it wins over the rule's configured severity: setting the rule to `severity: error` does not escalate those reports. Only `severity: 'off'` removes them, by disabling the rule.
1025
-
1026
- **These rules only fire when Markdoc tokenization is also on.** Set `markdoc: true` (or the object form, see below) alongside `extends: [recheck/markdoc]`. Extending the preset without the flag validates, but prints a console warning that the four rules can never report. The flag stays an explicit opt-in because Liquid and Jinja templates use the same `{% %}` delimiters for unrelated syntax, so Recheck never assumes it.
1027
- - **`recheck/google`** — Google's developer documentation style guide (https://developers.google.com/style), CC BY 4.0, synced 2026-07-29. 99 rules covering heading/list/table/link structure, sentence-case headings, sentence length, voice and contractions, plain language, product naming, compound word forms, and inclusive/precise-language terminology — all derived from the *live* guide (see `packages/recheck/presets/google/PROVENANCE.md` for the rule -> source page -> quote -> verdict table, including everything considered and NOT shipped, and why). `extends: [recheck/google]` is a one-line adoption of Google's style; combine with `recheck/markdown` for full structural linting too. Rule ids are namespaced `google/<rule>` (not `recheck/<rule>`) so they never collide with the markdownlint-parity or other style-guide presets. Structural/mechanical rules (heading hierarchy, list mechanics, alt-text presence, sentence length) are `severity: error`; every word-choice, terminology, and punctuation-convention rule is `severity: warn`. See `packages/recheck/presets/google/sources.json` for the fetched-page hashes. **Adopting this preset has a real, measured performance cost — roughly 2.7× the standard `recheck/markdown`-only profile's lint time on a docs-sized document set** — see [Performance](#performance) below (Phase 4) before turning it on in CI.
1028
- - **`recheck/microsoft`** — the Microsoft Writing Style Guide (https://learn.microsoft.com/en-us/style-guide/welcome/), CC BY 4.0 (via the guide's backing GitHub repository's LICENSE file — no `learn.microsoft.com` page states the licence itself, see `packages/recheck/presets/microsoft/PROVENANCE.md`), synced 2026-07-30. 93 rules covering heading/list/table/alt-text structure, the guide's own numeric thresholds (paragraph length, list length, comma density, alt-text length), its signature "use contractions" rule, US spelling, bias-free and people-first terminology, and a large A-Z terminology word list — all derived from the *live* guide and checked against four independent verification passes (~490 rules/entries across ~340 page fetches), with every Tier-1 pair anchored or demoted to detection-only wherever it was found capable of rewriting correct prose. Rule ids are namespaced `microsoft/<rule>`. Structural rules and the A-Z word list's three unconditional tiers are `severity: error`; voice, punctuation-convention, and UI-terminology rules are `severity: warn`. Audience-conditional and UI-conditional entries (Microsoft's own "Tier 4") are never enforced, and developer-audience carve-outs relevant to API documentation (`header`, `context menu`, `disk`, `directory`) are excluded rather than misfiring on Redocly's own docs — see `packages/recheck/presets/microsoft/PROVENANCE.md` for the full table, every excluded candidate, and why. Unlike `recheck/google` (which allows `click`), this preset bans all input-specific UI verbs (`click`, `press`, `hit`) in favor of `select` — the sharpest divergence between the two guides. See `packages/recheck/presets/microsoft/sources.json` for the fetched-page hashes.
1029
- - **`recheck/inclusive-language`** — composable, guide-agnostic: the *intersection* of `recheck/google` and `recheck/microsoft`'s inclusive/bias-free/ableist/accessibility content — terminology both flagship guides independently state should be avoided (`slave`, `master/slave`, `blacklist`/`whitelist`, `DMZ`, `grayed-out`, `he/she`, `normal person`/`healthy person`, `suffering from`/`victim of`, `differently abled`, `crippled`, `nuke`). All `warn` severity, all detection-only. Needed no new web fetch — every term was already confirmed against a live page by five existing verification reports; see `packages/recheck/presets/inclusive-language/PROVENANCE.md` for the report → row → term table and every single-guide term left out on purpose. Layer it onto either flagship or onto `recheck/prose`: `extends: [recheck/google, recheck/inclusive-language]`. **Because it's built as an intersection, every one of its 11 rules is already shipped by at least one flagship's own preset** (measured: 7 of 11 duplicate a `google/*` finding on the same span when stacked onto `recheck/google` alone, 6 of 11 duplicate a `microsoft/*` finding when stacked onto `recheck/microsoft` alone — see `packages/recheck/presets/inclusive-language/PROVENANCE.md`'s "Duplicate-finding audit"). Its full, zero-duplicate value is realized standalone, with `recheck/prose`, or on a project using neither flagship; stacked onto exactly one flagship it still fills that flagship's own gaps, but expect a majority of its findings to be reported twice.
1030
- - **`recheck/plain-language`** — composable, derived from the *live* US federal plain-language guidance (`digital.gov/guides/plain-language`; public domain, no attribution constraint). Smaller than a first read of the old `plainlanguage.gov` site would suggest: that site is now dead and redirects to a much thinner overview, so there's no sentence-length or readability-`metric` rule (`metric` stays a documented [opt-in](#opt-in-prose-assertions), unchanged) — only paragraph length (the one family with real, quotable numbers), filler/wordy phrases, complex-word substitutes, redundant pairs, double negatives, and jargon-to-plain examples. **`shall` is never flagged** — it's a defined RFC 2119 normative keyword used throughout specs and API docs, exactly what Recheck lints; `implement` and `command` carry the identical technical-sense collision and are excluded the same way. All `warn`/`error` (paragraph-length ceiling only) severity, all detection-only. `in order to` and `utilize`/`utilization` are deliberately NOT shipped despite being live, verbatim guide content — both flagships already ship the identical pair, so keeping them here would only ever produce a duplicate finding, never new coverage (measured: this cut duplicate findings on the same fixture from 6 to 3 against `recheck/google`, and from 5 to 3 against `recheck/microsoft`). The 3 that remain are an accepted paragraph-length overlap with `recheck/microsoft` (two independently-sourced numbers, not the same fact restated) and a coincidental substring collision with `use-contractions`, not content duplication. See `packages/recheck/presets/plain-language/PROVENANCE.md` for every rule's source quote, every family considered and left out, and the full duplicate-finding audit.
1182
+ - `recheck/markdoc-attributes` (mixed) — a missing required attribute, an enum or type violation, and a duplicate attribute report at `error`; an unknown attribute name, whether named or a stray positional value, always reports at `warn`.
1183
+ That `warn` is set per report by the rule itself, so it wins over the rule's configured severity: setting the rule to `severity: error` does not escalate those reports.
1184
+ Only `severity: 'off'` removes them, by disabling the rule.
1185
+
1186
+ **These rules only fire when Markdoc tokenization is also on.** Set `markdoc: true` (or the object form, see below) alongside `extends: [recheck/markdoc]`.
1187
+ Extending the preset without the flag validates, but prints a console warning that the four rules can never report.
1188
+ The flag stays an explicit opt-in because Liquid and Jinja templates use the same `{% %}` delimiters for unrelated syntax, so Recheck never assumes it.
1189
+ - **`recheck/google`** — Google's developer documentation style guide (https://developers.google.com/style), CC BY 4.0, synced 2026-07-29. 99 rules covering heading/list/table/link structure, sentence-case headings, sentence length, voice and contractions, plain language, product naming, compound word forms, and inclusive/precise-language terminology — all derived from the *live* guide (see `packages/recheck/presets/google/PROVENANCE.md` for the rule -> source page -> quote -> verdict table, including everything considered and NOT shipped, and why).
1190
+ `extends: [recheck/google]` is a one-line adoption of Google's style; combine with `recheck/markdown` for full structural linting too.
1191
+ Rule ids are namespaced `google/<rule>` (not `recheck/<rule>`) so they never collide with the markdownlint-parity or other style-guide presets.
1192
+ Structural/mechanical rules (heading hierarchy, list mechanics, alt-text presence, sentence length) are `severity: error`; every word-choice, terminology, and punctuation-convention rule is `severity: warn`.
1193
+ See `packages/recheck/presets/google/sources.json` for the fetched-page hashes. **Adopting this preset has a real, measured performance cost — roughly 2.7× the standard `recheck/markdown`-only profile's lint time on a docs-sized document set** — see [Performance](#performance) below (Phase 4) before turning it on in CI.
1194
+ - **`recheck/microsoft`** — the Microsoft Writing Style Guide (https://learn.microsoft.com/en-us/style-guide/welcome/), CC BY 4.0 (via the guide's backing GitHub repository's LICENSE file — no `learn.microsoft.com` page states the licence itself, see `packages/recheck/presets/microsoft/PROVENANCE.md`), synced 2026-07-30. 93 rules covering heading/list/table/alt-text structure, the guide's own numeric thresholds (paragraph length, list length, comma density, alt-text length), its signature "use contractions" rule, US spelling, bias-free and people-first terminology, and a large A-Z terminology word list — all derived from the *live* guide and checked against four independent verification passes (~490 rules/entries across ~340 page fetches), with every Tier-1 pair anchored or demoted to detection-only wherever it was found capable of rewriting correct prose.
1195
+ Rule ids are namespaced `microsoft/<rule>`.
1196
+ Structural rules and the A-Z word list's three unconditional tiers are `severity: error`; voice, punctuation-convention, and UI-terminology rules are `severity: warn`.
1197
+ Audience-conditional and UI-conditional entries (Microsoft's own "Tier 4") are never enforced, and developer-audience carve-outs relevant to API documentation (`header`, `context menu`, `disk`, `directory`) are excluded rather than misfiring on Redocly's own docs — see `packages/recheck/presets/microsoft/PROVENANCE.md` for the full table, every excluded candidate, and why.
1198
+ Unlike `recheck/google` (which allows `click`), this preset bans all input-specific UI verbs (`click`, `press`, `hit`) in favor of `select` — the sharpest divergence between the two guides.
1199
+ See `packages/recheck/presets/microsoft/sources.json` for the fetched-page hashes.
1200
+ - **`recheck/inclusive-language`** — composable, guide-agnostic: the *intersection* of `recheck/google` and `recheck/microsoft`'s inclusive/bias-free/ableist/accessibility content — terminology both flagship guides independently state should be avoided (`slave`, `master/slave`, `blacklist`/`whitelist`, `DMZ`, `grayed-out`, `he/she`, `normal person`/`healthy person`, `suffering from`/`victim of`, `differently abled`, `crippled`, `nuke`).
1201
+ All `warn` severity, all detection-only.
1202
+ Needed no new web fetch — every term was already confirmed against a live page by five existing verification reports; see `packages/recheck/presets/inclusive-language/PROVENANCE.md` for the report → row → term table and every single-guide term left out on purpose.
1203
+ Layer it onto either flagship or onto `recheck/prose`: `extends: [recheck/google, recheck/inclusive-language]`. **Because it's built as an intersection, every one of its 11 rules is already shipped by at least one flagship's own preset** (measured: 7 of 11 duplicate a `google/*` finding on the same span when stacked onto `recheck/google` alone, 6 of 11 duplicate a `microsoft/*` finding when stacked onto `recheck/microsoft` alone — see `packages/recheck/presets/inclusive-language/PROVENANCE.md`'s "Duplicate-finding audit").
1204
+ Its full, zero-duplicate value is realized standalone, with `recheck/prose`, or on a project using neither flagship; stacked onto exactly one flagship it still fills that flagship's own gaps, but expect a majority of its findings to be reported twice.
1205
+ - **`recheck/plain-language`** — composable, derived from the *live* US federal plain-language guidance (`digital.gov/guides/plain-language`; public domain, no attribution constraint).
1206
+ Smaller than a first read of the old `plainlanguage.gov` site would suggest: that site is now dead and redirects to a much thinner overview, so there's no sentence-length or readability-`metric` rule (`metric` stays a documented [opt-in](#opt-in-prose-assertions), unchanged) — only paragraph length (the one family with real, quotable numbers), filler/wordy phrases, complex-word substitutes, redundant pairs, double negatives, and jargon-to-plain examples. **`shall` is never flagged** — it's a defined RFC 2119 normative keyword used throughout specs and API docs, exactly what Recheck lints; `implement` and `command` carry the identical technical-sense collision and are excluded the same way.
1207
+ All `warn`/`error` (paragraph-length ceiling only) severity, all detection-only.
1208
+ `in order to` and `utilize`/`utilization` are deliberately NOT shipped despite being live, verbatim guide content — both flagships already ship the identical pair, so keeping them here would only ever produce a duplicate finding, never new coverage (measured: this cut duplicate findings on the same fixture from 6 to 3 against `recheck/google`, and from 5 to 3 against `recheck/microsoft`).
1209
+ The 3 that remain are an accepted paragraph-length overlap with `recheck/microsoft` (two independently-sourced numbers, not the same fact restated) and a coincidental substring collision with `use-contractions`, not content duplication.
1210
+ See `packages/recheck/presets/plain-language/PROVENANCE.md` for every rule's source quote, every family considered and left out, and the full duplicate-finding audit.
1031
1211
 
1032
1212
  **All four of the presets above — `recheck/google`, `recheck/microsoft`,
1033
1213
  `recheck/inclusive-language`, and `recheck/plain-language` — are
@@ -1039,28 +1219,35 @@ There is one exception. A rule may attach a milder severity to some of its own r
1039
1219
  makes one fixable again — see `preset-google.test.ts`'s and
1040
1220
  `preset-microsoft.test.ts`'s "is detection-only" describe blocks, and
1041
1221
  `preset-composition.test.ts`'s list-driven version covering all four.
1042
- `recheck/google` and `recheck/microsoft` once had fixable rules. Every
1222
+ `recheck/google` and `recheck/microsoft` once had fixable rules.
1223
+ Every
1043
1224
  attempt to define a safe subset of them found the fixes corrupting
1044
1225
  genuinely correct prose, in every category previously believed safe:
1045
1226
  spelling (Hemingway's correctly spelled *A Moveable Feast* → "A Movable
1046
1227
  Feast"), hyphenation ("read only the introduction" → "read-only the
1047
1228
  introduction"), and at least one outright inversion of meaning ("No SQL is
1048
- used here" → "NoSQL is used here"). A rule's *category* does not predict
1229
+ used here" → "NoSQL is used here").
1230
+ A rule's *category* does not predict
1049
1231
  fix safety at this scale: a style guide states intent while
1050
1232
  `swap`/`consistency`/`pattern` match tokens, and narrowing which
1051
- categories count as "safe" doesn't close that gap. Detection is
1233
+ categories count as "safe" doesn't close that gap.
1234
+ Detection is
1052
1235
  unaffected — every rule still runs and reports, and you apply the fix
1053
- yourself with the judgment style guidance has always required. This is the
1236
+ yourself with the judgment style guidance has always required.
1237
+ This is the
1054
1238
  same reason Vale, the tool these presets replace, never shipped this class
1055
- of bug. See the "Detection-only" sections of
1239
+ of bug.
1240
+ See the "Detection-only" sections of
1056
1241
  `packages/recheck/presets/google/PROVENANCE.md` and
1057
1242
  `packages/recheck/presets/microsoft/PROVENANCE.md` for the full history.
1058
1243
 
1059
1244
  The heading rule uses **sentence case** *(changed from AP title case by product decision 2026-07-29: Redocly's own guide, Google, and Microsoft all mandate sentence case)*.
1060
1245
  Two details make that default safe out of the box:
1061
1246
 
1062
- - **The [built-in technical proper-noun vocabulary](#built-in-technical-proper-noun-vocabulary)** — `TECHNICAL_PROPER_NOUNS` — is unioned into `exceptions` by `capitalization` itself (default `builtinVocabulary: true`), so this preset doesn't ship its own copy: `$sentence` still won't flag `OpenAPI`, `GitHub`, `macOS`, and the rest of that list out of the box. It's a common-vocabulary floor, not a full brand list — extend it with your own product/company names via this rule's own `exceptions`, which **compose** with the built-ins rather than replacing them (unlike a preset-shipped list, which a same-key override would have replaced entirely).
1063
- - **`fix: false`** — a sentence-case auto-fix would lowercase any proper noun the built-ins and your own `exceptions` don't cover, silently damaging content. Set `fix: true` on your own `recheck/capitalization` key (or drop the key) once your exceptions list covers your vocabulary.
1247
+ - **The [built-in technical proper-noun vocabulary](#built-in-technical-proper-noun-vocabulary)** — `TECHNICAL_PROPER_NOUNS` — is unioned into `exceptions` by `capitalization` itself (default `builtinVocabulary: true`), so this preset doesn't ship its own copy: `$sentence` still won't flag `OpenAPI`, `GitHub`, `macOS`, and the rest of that list out of the box.
1248
+ It's a common-vocabulary floor, not a full brand list — extend it with your own product/company names via this rule's own `exceptions`, which **compose** with the built-ins rather than replacing them (unlike a preset-shipped list, which a same-key override would have replaced entirely).
1249
+ - **`fix: false`** — a sentence-case auto-fix would lowercase any proper noun the built-ins and your own `exceptions` don't cover, silently damaging content.
1250
+ Set `fix: true` on your own `recheck/capitalization` key (or drop the key) once your exceptions list covers your vocabulary.
1064
1251
 
1065
1252
  ```yaml
1066
1253
  extends:
@@ -1076,9 +1263,10 @@ recheck/line-length:
1076
1263
 
1077
1264
  Multiple presets can be listed; later presets in the list override earlier ones for the same rule key, before your own top-level rule keys are merged in last.
1078
1265
 
1079
- ### Tuning a preset
1266
+ ### Tune a preset
1080
1267
 
1081
- Adopting a whole style-guide preset doesn't mean accepting every rule at its shipped severity. Because your own config's rule keys always win over a preset's (see above), you can turn individual rules off, downgrade them, or silence single occurrences — all verified against a live build, not just read from source:
1268
+ Adopting a whole style-guide preset doesn't mean accepting every rule at its shipped severity.
1269
+ Because your own config's rule keys always win over a preset's (see above), you can turn individual rules off, downgrade them, or silence single occurrences — all verified against a live build, not just read from source:
1082
1270
 
1083
1271
  ```yaml
1084
1272
  extends: [recheck/markdown, recheck/microsoft]
@@ -1103,7 +1291,9 @@ Click the hot link to continue.
1103
1291
  <!-- recheck-disable-file -->
1104
1292
  ```
1105
1293
 
1106
- **The sharp edge:** merging a user override on top of a preset rule happens per *assertion id*, not per option inside it. Setting a partial override on a bundled `swap` or `pattern` rule doesn't just change the one option you named — it **replaces that assertion object entirely**, silently discarding everything else it carried. For example:
1294
+ **The sharp edge:** merging a user override on top of a preset rule happens per *assertion id*, not per option inside it.
1295
+ Setting a partial override on a bundled `swap` or `pattern` rule doesn't just change the one option you named — it **replaces that assertion object entirely**, silently discarding everything else it carried.
1296
+ For example:
1107
1297
 
1108
1298
  ```yaml
1109
1299
  microsoft/spelling-hyphenation:
@@ -1118,7 +1308,9 @@ drops the preset's whole `pairs` map along with it, and the config then fails va
1118
1308
  Rule "microsoft/spelling-hyphenation": swap requires a "pairs" object mapping find -> replace strings
1119
1309
  ```
1120
1310
 
1121
- So today, to reject just one term out of a bundled `swap`/`pattern` rule, your options are: turn the whole rule off, restate its entire `pairs`/`tokens` yourself, or inline-disable each occurrence as shown above. Two assertion types already have a real per-term escape hatch that doesn't hit this edge: `capitalization`'s `exceptions` (an array of allowed terms that **composes** with the built-in technical-proper-noun vocabulary and anything else you add, rather than replacing it) and `spelling`'s `ignore`. A per-term opt-out for `swap`/`pattern` is a known follow-up, not shipped yet.
1311
+ So today, to reject just one term out of a bundled `swap`/`pattern` rule, your options are: turn the whole rule off, restate its entire `pairs`/`tokens` yourself, or inline-disable each occurrence as shown above.
1312
+ Two assertion types already have a real per-term escape hatch that doesn't hit this edge: `capitalization`'s `exceptions` (an array of allowed terms that **composes** with the built-in technical-proper-noun vocabulary and anything else you add, rather than replacing it) and `spelling`'s `ignore`.
1313
+ A per-term opt-out for `swap`/`pattern` is a known follow-up, not shipped yet.
1122
1314
 
1123
1315
  ### Example configs
1124
1316
 
@@ -1127,15 +1319,20 @@ So today, to reject just one term out of a bundled `swap`/`pattern` rule, your o
1127
1319
  Each file has four parts, in this order:
1128
1320
 
1129
1321
  1. **An attribution header** — source, license, and sync date, as YAML comments (mirrors that preset's `PROVENANCE.md`).
1130
- 2. **`# What to paste`** — the actual adoption cost: a two-to-four-line `extends` block. This is the only part most readers need; everything below it is supporting material, not something to copy.
1131
- 3. **`# How to tune it`** — override patterns verified to work today (turn a rule off, downgrade its severity, inline-disable one occurrence with an HTML comment), plus a documented sharp edge: overriding one option on a bundled `swap`/`pattern` rule's `assertions` **replaces that assertion entirely**, silently discarding options like a `pairs` map you didn't restate (merging is per *assertion id*, not per option) — restate the whole map, turn the rule off, or use an inline directive instead. `capitalization`'s `exceptions` and `spelling`'s `ignore` are the two assertion types that already have a real per-term escape hatch; an equivalent for `swap`/`pattern` is a known follow-up, not shipped yet.
1132
- 4. **`# Full expansion (reference)`** — the preset's entire resolved rule set (alphabetized), so a reader can see exactly what they're adopting without running the tool. Every value here is identical to what the `extends` block above already resolves to, so copying this section too is redundant, not broken — it's for reading, not pasting.
1322
+ 2. **`# What to paste`** — the actual adoption cost: a two-to-four-line `extends` block.
1323
+ This is the only part most readers need; everything below it is supporting material, not something to copy.
1324
+ 3. **`# How to tune it`** — override patterns verified to work today (turn a rule off, downgrade its severity, inline-disable one occurrence with an HTML comment), plus a documented sharp edge: overriding one option on a bundled `swap`/`pattern` rule's `assertions` **replaces that assertion entirely**, silently discarding options like a `pairs` map you didn't restate (merging is per *assertion id*, not per option) — restate the whole map, turn the rule off, or use an inline directive instead.
1325
+ `capitalization`'s `exceptions` and `spelling`'s `ignore` are the two assertion types that already have a real per-term escape hatch; an equivalent for `swap`/`pattern` is a known follow-up, not shipped yet.
1326
+ 4. **`# Full expansion (reference)`** — the preset's entire resolved rule set (alphabetized), so a reader can see exactly what they're adopting without running the tool.
1327
+ Every value here is identical to what the `extends` block above already resolves to, so copying this section too is redundant, not broken — it's for reading, not pasting.
1133
1328
 
1134
1329
  A hand-maintained appendix is appended verbatim after part 4: NOISY candidates the guide states but the preset doesn't enforce (shown as the rule they'd be if shipped, commented out, with a one-line false-positive note each) and a checklist of guide content that needs a human, not a linter (NOT-ENFORCEABLE — active voice, missing-Oxford-comma detection, and similar).
1135
1330
 
1136
1331
  ### Opt-in prose assertions
1137
1332
 
1138
- `recheck/prose` (above) intentionally ships only `repetition`, `consistency`, and `capitalization` — a small, broadly-applicable default. Three more Vale-parity/native assertions exist (see [Assertion Types](#assertion-types) above for full per-option tables) but are **not shipped in any preset**, because their thresholds, patterns, or dictionaries are inherently project-specific rather than having one right-for-everyone default: `conditional`, `metric`, `spelling`. (`length` and `occurrence` used to be entries here; neither is an opt-in any more — [`recheck/google`](#extends-presets) ships `length` directly for the guide's sentence-length limit, and [`recheck/microsoft`](#extends-presets) ships `occurrence` directly for the guide's comma-density rule, so neither one's default bounds are "no one right answer" any more.) Add any of the three by copying its rule below into your own config, alongside `extends: [recheck/prose]`:
1333
+ `recheck/prose` (above) intentionally ships only `repetition`, `consistency`, and `capitalization` — a small, broadly-applicable default.
1334
+ Three more Vale-parity/native assertions exist (see [Assertion Types](#assertion-types) above for full per-option tables) but are **not shipped in any preset**, because their thresholds, patterns, or dictionaries are inherently project-specific rather than having one right-for-everyone default: `conditional`, `metric`, `spelling`.
1335
+ (`length` and `occurrence` used to be entries here; neither is an opt-in any more — [`recheck/google`](#extends-presets) ships `length` directly for the guide's sentence-length limit, and [`recheck/microsoft`](#extends-presets) ships `occurrence` directly for the guide's comma-density rule, so neither one's default bounds are "no one right answer" any more.) Add any of the three by copying its rule below into your own config, alongside `extends: [recheck/prose]`:
1139
1336
 
1140
1337
  ```yaml
1141
1338
  extends: [recheck/markdown, recheck/prose]
@@ -1171,7 +1368,7 @@ recheck/us-spelling-check:
1171
1368
 
1172
1369
  Each snippet's `severity`, `message`, `scope`, and `exceptions` are yours to adjust — see [Rule Types and Assertions](#rule-types-and-assertions) for every option each assertion accepts, and [Inline Directives](#inline-directives) to silence any one of them on a specific line or file with an HTML comment instead of turning it off entirely.
1173
1370
 
1174
- ### Migrating from markdownlint
1371
+ ### Migrate from markdownlint
1175
1372
 
1176
1373
  A markdownlint config maps onto Recheck almost 1:1 — `extends` a preset, then override individual rules by their Recheck name (same short name markdownlint uses, e.g. `line-length` for MD013) under `assertions`:
1177
1374
 
@@ -1210,10 +1407,27 @@ Two other ids are **upstream markdownlint's own alternate rule names**, not a Re
1210
1407
 
1211
1408
  **Two intentional behavior changes** vs. plain markdownlint defaults, both on rules that predate the parity port and kept their exact ids:
1212
1409
 
1213
- - **`no-trailing-spaces` (MD009)** now exempts lines with *exactly 2* trailing spaces by default (a markdown hard line break), instead of flagging all trailing whitespace. Set `strict: true` to restore the old flag-everything behavior (matches markdownlint's default).
1214
- - **`no-hard-tabs` (MD010)**'s `spacesPerTab` option now defaults to `1` (matching markdownlint's own upstream default) — Recheck's earlier, pre-parity native rule had defaulted this to `2`. If you were relying on that old default, set `spacesPerTab: 2` explicitly.
1410
+ - **`no-trailing-spaces` (MD009)** now exempts lines with *exactly 2* trailing spaces by default (a markdown hard line break), instead of flagging all trailing whitespace.
1411
+ Set `strict: true` to restore the old flag-everything behavior (matches markdownlint's default).
1412
+ - **`no-hard-tabs` (MD010)**'s `spacesPerTab` option now defaults to `1` (matching markdownlint's own upstream default) — Recheck's earlier, pre-parity native rule had defaulted this to `2`.
1413
+ If you were relying on that old default, set `spacesPerTab: 2` explicitly.
1414
+
1415
+ Token-rule options aren't individually schema-validated, so an option name a rule doesn't recognize is silently ignored — it has no effect, and produces no warning or error.
1416
+
1417
+ ### Cross-file link validation
1215
1418
 
1216
- Token-rule options aren't individually schema-validated, so an option name a rule doesn't recognize (e.g. a leftover `checkExternalFiles` on `link-fragments`, from the old `no-broken-fragment-links` rule, which never implemented it either) is silently ignored — it has no effect, and produces no warning or error.
1419
+ `link-fragments` accepts `crossFile: true` to validate links across files, replacing external link checkers such as `mlc` for repo-internal links:
1420
+
1421
+ - A relative link or image target must exist on disk (`[x](./missing.md)` flags).
1422
+ - A `file.md#anchor` fragment must exist in the target file's headings and anchors.
1423
+ - Extensionless links resolve the way the Realm router does: `./page` tries `page.md`, and a directory link reads its `index.md`.
1424
+ - Site-root absolute paths (`/x/y`) resolve against the `rootDir` option when set, and are skipped without it.
1425
+ In a monorepo with several docs projects, give `rootDir` a map from source-directory prefix to that directory's site root; the longest matching prefix wins, and files under no prefix keep the skip.
1426
+ Paths are relative to the working directory.
1427
+ - External URLs and `mailto:` are skipped.
1428
+
1429
+ Each target file is read once per run and cached by modification time.
1430
+ The option is off by default; in-document fragment checking is unchanged.
1217
1431
 
1218
1432
  ### Known differences from markdownlint
1219
1433
 
@@ -1223,19 +1437,26 @@ Token-rule options aren't individually schema-validated, so an option name a rul
1223
1437
  <!-- markdownlint&#45;disable -->
1224
1438
  ```
1225
1439
 
1226
- (dash HTML-escaped above so this very README doesn't trip markdownlint's own directive scanner — markdownlint recognizes these directives even inside fenced code, so the literal syntax can't appear here unescaped). Markdownlint's HTML-comment-based per-line/per-region rule toggles (`markdownlint-disable`, `markdownlint-disable-next-line`, `markdownlint-enable`, etc.) are a distinct engine feature, not a rule port, and Recheck doesn't parse them today. Use config-level `exceptions` (file/line patterns) or `excludes`/`appliesTo` to achieve the same effect. Native support is a possible future addition; no decision has been made yet.
1440
+ (dash HTML-escaped above so this very README doesn't trip markdownlint's own directive scanner — markdownlint recognizes these directives even inside fenced code, so the literal syntax can't appear here unescaped).
1441
+ Markdownlint's HTML-comment-based per-line/per-region rule toggles (`markdownlint-disable`, `markdownlint-disable-next-line`, `markdownlint-enable`, etc.) are a distinct engine feature, not a rule port, and Recheck doesn't parse them today.
1442
+ Use config-level `exceptions` (file/line patterns) or `excludes`/`appliesTo` to achieve the same effect.
1443
+ Native support is a possible future addition; no decision has been made yet.
1227
1444
 
1228
1445
  ### Parity with markdownlint
1229
1446
 
1230
- The 53 ported rules are checked against upstream markdownlint by a differential harness (`pnpm parity`, `benchmarks/parity/run-parity.mjs`): the harness lints the same real-world document set with both Recheck (via a config translated from markdownlint's option surface) and markdownlint itself, then set-diffs the findings. As of this writing it reports **zero unexplained differences** across:
1447
+ The 53 ported rules are checked against upstream markdownlint by a differential harness (`pnpm parity`, `benchmarks/parity/run-parity.mjs`): the harness lints the same real-world document set with both Recheck (via a config translated from markdownlint's option surface) and markdownlint itself, then set-diffs the findings.
1448
+ As of this writing it reports **zero unexplained differences** across:
1231
1449
 
1232
1450
  - `mdn-content` (MDN Web Docs, ~14.5k files)
1233
1451
  - `electron` (Electron's docs + repo markdown)
1234
1452
  - `monorepo-docs` (this monorepo's own `docs/` tree)
1235
1453
 
1236
- on both the `default` profile (full `recheck/markdown` preset vs. markdownlint `{ default: true }`) and a `rebilly` profile (a real third-party `.markdownlint.yaml` translated to Recheck config). A small, explicitly documented allowlist (`benchmarks/parity/allowlist.json`) covers the one known permanent engine-surface gap — inline `markdownlint-disable` directives (see [Known differences](#known-differences-from-markdownlint) above) — scoped to the exact rules it can affect (MD010, MD011, MD033, MD059). Everything else matches exactly.
1454
+ on both the `default` profile (full `recheck/markdown` preset vs. markdownlint `{ default: true }`) and a `rebilly` profile (a real third-party `.markdownlint.yaml` translated to Recheck config).
1455
+ A small, explicitly documented allowlist (`benchmarks/parity/allowlist.json`) covers the one known permanent engine-surface gap — inline `markdownlint-disable` directives (see [Known differences](#known-differences-from-markdownlint) above) — scoped to the exact rules it can affect (MD010, MD011, MD033, MD059).
1456
+ Everything else matches exactly.
1237
1457
 
1238
- **`pnpm parity` requires `--corpus`** — running it bare exits `2` with a usage error (`Usage: node benchmarks/parity/run-parity.mjs --corpus <name> [--profile default|rebilly] [--rules MD001,MD013]`) rather than running against a default document set. Always pass a document-set name, e.g. `pnpm parity --corpus monorepo-docs --profile default`.
1458
+ **`pnpm parity` requires `--corpus`** — running it bare exits `2` with a usage error (`Usage: node benchmarks/parity/run-parity.mjs --corpus <name> [--profile default|rebilly] [--rules MD001,MD013]`) rather than running against a default document set.
1459
+ Always pass a document-set name, e.g. `pnpm parity --corpus monorepo-docs --profile default`.
1239
1460
 
1240
1461
  ## GitHub Actions Integration
1241
1462
 
@@ -1303,19 +1524,35 @@ Recheck trades a bit of the Phase 1 constant-factor lead for full rule-count par
1303
1524
 
1304
1525
  - **Phase 3** (prose profile — the standard Phase 2 rule set vs. that same set plus the Vale-parity prose additions, `monorepo-docs` document set, the same set on both sides):
1305
1526
  - Standard profile (`recheck-mdl-preset.yaml`, `extends: [recheck/markdown]`, 53 rules, `run-recheck-mdl-preset.mjs`, 965 files): **4072ms** median — **-8.9% vs the Phase 2 recording** (`phase2-parity-recheck`, 4468ms, 954 files), i.e. no regression from the Phase 3 rule-registry additions, since the standard profile doesn't exercise any of them (the small speedup is run-to-run variance plus the document-set size difference, not an optimization claim).
1306
- - Prose profile (`recheck-prose-bench.yaml`, `extends: [recheck/markdown, recheck/prose]` plus one opt-in `occurrence` rule and one opt-in `conditional` rule, 58 rules, `run-recheck-prose.mjs`, same 965 files): **4547ms** median — **+11.7%** vs. the standard profile above, for the five added prose/scope rules (`repetition`, `consistency`, `capitalization`, and the two opt-ins). Comfortably under the "investigate if >2x standard" threshold; no pathological per-rule cost found. There is no hard pass/fail gate for this profile (new profile, first recording) — these numbers establish its baseline.
1307
- - Measured 2026-07-27 (local time; the result JSONs record the UTC date `2026-07-28`, so the file dates and this measurement date differ by design, not by error) at commit `27a8b6feb10` (immediately prior to the commit that added this benchmark profile), on an Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine; single 3-run session, not a statistically rigorous multi-session average — treat the deltas as directional, not precise. `pnpm bench --subject <script> --corpus monorepo-docs --runs 3 --record <label>` (median of 3 timed runs after 1 warm-up); recorded to `benchmarks/results/phase3-prose-standard.json` / `phase3-prose-profile.json`. A `--corpus self` (2-file) smoke pair recorded to `phase3-prose-standard-self.json` / `phase3-prose-profile-self.json` proves the harness/subject-script mechanics end-to-end but isn't large enough to be a meaningful timing signal on its own.
1527
+ - Prose profile (`recheck-prose-bench.yaml`, `extends: [recheck/markdown, recheck/prose]` plus one opt-in `occurrence` rule and one opt-in `conditional` rule, 58 rules, `run-recheck-prose.mjs`, same 965 files): **4547ms** median — **+11.7%** vs. the standard profile above, for the five added prose/scope rules (`repetition`, `consistency`, `capitalization`, and the two opt-ins).
1528
+ Comfortably under the "investigate if >2x standard" threshold; no pathological per-rule cost found.
1529
+ There is no hard pass/fail gate for this profile (new profile, first recording) — these numbers establish its baseline.
1530
+ - Measured 2026-07-27 (local time; the result JSONs record the UTC date `2026-07-28`, so the file dates and this measurement date differ by design, not by error) at commit `27a8b6feb10` (immediately prior to the commit that added this benchmark profile), on an Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine; single 3-run session, not a statistically rigorous multi-session average — treat the deltas as directional, not precise.
1531
+ `pnpm bench --subject <script> --corpus monorepo-docs --runs 3 --record <label>` (median of 3 timed runs after 1 warm-up); recorded to `benchmarks/results/phase3-prose-standard.json` / `phase3-prose-profile.json`.
1532
+ A `--corpus self` (2-file) smoke pair recorded to `phase3-prose-standard-self.json` / `phase3-prose-profile-self.json` proves the harness/subject-script mechanics end-to-end but isn't large enough to be a meaningful timing signal on its own.
1308
1533
 
1309
1534
  - **Phase 4** (refreshed standard/prose figures plus the new `recheck/google` profile, `monorepo-docs` document set, the same set across all three, 968 files — the document set grew by 3 files since the Phase 3 recording):
1310
- - Standard profile (same config as Phase 2/3, `recheck-mdl-preset.yaml`, 53 rules): **4350ms** median (runs 4341/4350/4369ms — 28ms spread, 0.6% of median: a tight, trustworthy measurement). This refreshes — and for current comparisons supersedes — Phase 3's `4072ms`/965-file recording; the ~7% difference is within normal session-to-session noise (different process/cache/scheduler state), not a regression, and there is still no rule-registry change that would affect this profile.
1311
- - Prose profile (same config as Phase 3, `recheck-prose-bench.yaml`, 58 rules): **4801ms** median (runs 4599/4801/4840ms — 241ms spread, 5.0% of median) — **+10.4%** vs. the refreshed standard profile above. This is the requested refresh of the Phase 3 prose figures, which the `capitalization` default's AP-title-case → sentence-case change (2026-07-29, see [`extends` presets](#extends-presets) below) made marginally stale, since this profile's `capitalization` rule is exactly what that change touched. The new delta (+10.4%) is close to Phase 3's original (+11.7%); given the 5.0% run-to-run spread observed here, treat both numbers as directionally consistent, not as proof of a precise change in cost — same posture Phase 3 itself took.
1535
+ - Standard profile (same config as Phase 2/3, `recheck-mdl-preset.yaml`, 53 rules): **4350ms** median (runs 4341/4350/4369ms — 28ms spread, 0.6% of median: a tight, trustworthy measurement).
1536
+ This refreshes — and for current comparisons supersedes — Phase 3's `4072ms`/965-file recording; the ~7% difference is within normal session-to-session noise (different process/cache/scheduler state), not a regression, and there is still no rule-registry change that would affect this profile.
1537
+ - Prose profile (same config as Phase 3, `recheck-prose-bench.yaml`, 58 rules): **4801ms** median (runs 4599/4801/4840ms — 241ms spread, 5.0% of median) — **+10.4%** vs. the refreshed standard profile above.
1538
+ This is the requested refresh of the Phase 3 prose figures, which the `capitalization` default's AP-title-case → sentence-case change (2026-07-29, see [`extends` presets](#extends-presets) below) made marginally stale, since this profile's `capitalization` rule is exactly what that change touched.
1539
+ The new delta (+10.4%) is close to Phase 3's original (+11.7%); given the 5.0% run-to-run spread observed here, treat both numbers as directionally consistent, not as proof of a precise change in cost — same posture Phase 3 itself took.
1312
1540
  - **`recheck/google` profile** (new — `recheck-google-bench.yaml`, `extends: [recheck/markdown, recheck/google]`, 152 rules total: the same 53-rule structural set plus all 99 rules `recheck/google` ships, `run-recheck-google.mjs`, same 968 files): **11792ms** median (runs 11628/11792/12164ms — 536ms spread, 4.5% of median).
1313
- - **This profile is substantially more expensive than either profile above: roughly 2.7× (~+171%) the standard profile's median time**, a materially different result from the ~1.1–1.2× range Phase 2/3 established for the markdownlint-parity and Vale-parity workloads. This is recorded as a **new baseline on its own terms, not a regression against the standard-profile's ±40% parity gate** — that gate applies only to `recheck/markdown` vs. markdownlint's equivalent rule set (Phase 2, unaffected by this preset's addition) and was never meant to bound a ~99-rule prose-preset workload layered on top of it. The number is reported as measured, without tuning the preset to improve it.
1314
- - The ~171% delta is far larger than the ≤5% run-to-run spread measured on all three profiles this session, so the *direction and rough magnitude* of "the Google preset costs several times more than the structural set alone" is trustworthy; the precise "171%" is not — this is a single 3-run-median session on one developer machine, not a statistically rigorous benchmark. Read it as "roughly 2.5–3×," not as a number with two decimal digits of meaning.
1315
- - **What this means for adopters**: turning on `extends: [recheck/google]` roughly triples per-run lint time on a docs-sized document set (968 files: ~4.3s → ~11.8s). That is a real, user-facing cost, stated here so it is visible before adoption rather than discovered later in CI — projects sensitive to CI duration should budget for it explicitly (e.g., a separate, non-blocking job, or a scheduled run) rather than assuming `recheck/google` is free to layer on top of `recheck/markdown`.
1316
- - Measured 2026-07-30 at commit `7c50caaaf05` (immediately prior to the commit that adds this benchmark profile), on the same Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine as Phase 1–3; single 3-run session per profile (1 warm-up + 3 timed runs), same caveats as Phase 3 — treat deltas as directional, not precise. `pnpm bench --subject <script> --corpus monorepo-docs --runs 3 --record <label>`; recorded to `benchmarks/results/phase4-google-standard.json` / `phase4-google-prose-refresh.json` / `phase4-google-profile.json`.
1317
-
1318
- - **Markdoc parse cost** (`parseMarkdown` in isolation — no rules, no config validation — flag off vs. flag on, `monorepo-docs` document set, 981 files, `run-recheck-parse.mjs` / `run-recheck-parse-markdoc.mjs`): flag off **2825ms** median (runs 2784/2825/2966ms) vs. flag on **2896ms** median (runs 2864/2896/3101ms) — **+2.5%**. That gap is smaller than the ~240–280ms run-to-run spread on either side, so read it as "no measurable parse-time cost from turning `markdoc: true` on" rather than a precise 2.5% overhead; the number is reported as measured. This is its own baseline row — there is no earlier "parse only, flag off" recording to compare against — and it is deliberately not gated against the ±40% markdownlint-parity budget above, which covers the 53-rule structural comparison and never sets this flag. Measured 2026-08-02 at commit `db8df4378ff`, on the same Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine as the phases above; single 3-run session per side. `pnpm bench --subject benchmarks/run-recheck-parse.mjs --corpus monorepo-docs --runs 3 --record <label>` (and the `-markdoc` sibling script for the flag-on side); recorded to `benchmarks/results/markdoc-parse-cost-flag-off.json` / `markdoc-parse-cost-flag-on.json`.
1541
+ - **This profile is substantially more expensive than either profile above: roughly 2.7× (~+171%) the standard profile's median time**, a materially different result from the ~1.1–1.2× range Phase 2/3 established for the markdownlint-parity and Vale-parity workloads.
1542
+ This is recorded as a **new baseline on its own terms, not a regression against the standard-profile's ±40% parity gate** — that gate applies only to `recheck/markdown` vs. markdownlint's equivalent rule set (Phase 2, unaffected by this preset's addition) and was never meant to bound a ~99-rule prose-preset workload layered on top of it.
1543
+ The number is reported as measured, without tuning the preset to improve it.
1544
+ - The ~171% delta is far larger than the ≤5% run-to-run spread measured on all three profiles this session, so the *direction and rough magnitude* of "the Google preset costs several times more than the structural set alone" is trustworthy; the precise "171%" is not — this is a single 3-run-median session on one developer machine, not a statistically rigorous benchmark.
1545
+ Read it as "roughly 2.5–3×," not as a number with two decimal digits of meaning.
1546
+ - **What this means for adopters**: turning on `extends: [recheck/google]` roughly triples per-run lint time on a docs-sized document set (968 files: ~4.3s → ~11.8s).
1547
+ That is a real, user-facing cost, stated here so it is visible before adoption rather than discovered later in CI — projects sensitive to CI duration should budget for it explicitly (e.g., a separate, non-blocking job, or a scheduled run) rather than assuming `recheck/google` is free to layer on top of `recheck/markdown`.
1548
+ - Measured 2026-07-30 at commit `7c50caaaf05` (immediately prior to the commit that adds this benchmark profile), on the same Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine as Phase 1–3; single 3-run session per profile (1 warm-up + 3 timed runs), same caveats as Phase 3 — treat deltas as directional, not precise.
1549
+ `pnpm bench --subject <script> --corpus monorepo-docs --runs 3 --record <label>`; recorded to `benchmarks/results/phase4-google-standard.json` / `phase4-google-prose-refresh.json` / `phase4-google-profile.json`.
1550
+
1551
+ - **Markdoc parse cost** (`parseMarkdown` in isolation — no rules, no config validation — flag off vs. flag on, `monorepo-docs` document set, 981 files, `run-recheck-parse.mjs` / `run-recheck-parse-markdoc.mjs`): flag off **2825ms** median (runs 2784/2825/2966ms) vs. flag on **2896ms** median (runs 2864/2896/3101ms) — **+2.5%**.
1552
+ That gap is smaller than the ~240–280ms run-to-run spread on either side, so read it as "no measurable parse-time cost from turning `markdoc: true` on" rather than a precise 2.5% overhead; the number is reported as measured.
1553
+ This is its own baseline row — there is no earlier "parse only, flag off" recording to compare against — and it is deliberately not gated against the ±40% markdownlint-parity budget above, which covers the 53-rule structural comparison and never sets this flag.
1554
+ Measured 2026-08-02 at commit `db8df4378ff`, on the same Apple M2 Max / Darwin 24.6.0 / Node v23.7.0 machine as the phases above; single 3-run session per side.
1555
+ `pnpm bench --subject benchmarks/run-recheck-parse.mjs --corpus monorepo-docs --runs 3 --record <label>` (and the `-markdoc` sibling script for the flag-on side); recorded to `benchmarks/results/markdoc-parse-cost-flag-off.json` / `markdoc-parse-cost-flag-on.json`.
1319
1556
  - ✅ **Scalable**: File-first architecture with rule indexing optimizes for large repositories; the same micromark AST backs both markdown-structure rules and prose-scope rules, so combining both rule families costs one parse, not two.
1320
1557
 
1321
1558
  ## Dependencies
@@ -1330,7 +1567,8 @@ Recheck trades a bit of the Phase 1 constant-factor lead for full rule-count par
1330
1567
  - `string-width` - measures the display width of strings containing wide/ambiguous-width or ANSI-styled characters, for CLI table output alignment
1331
1568
 
1332
1569
  ### Optional Peer Dependencies
1333
- - `nspell` + `dictionary-en` - Hunspell-compatible spell checker (and its bundled English dictionary) backing the `spelling` assertion (see [Spelling Assertions](#spelling-assertions-spelling) above). **Not installed by installing `@redocly/recheck`** — both are declared `optional: true` in `peerDependenciesMeta`, loaded lazily via dynamic `import()` only when a config actually enables `spelling`. Run `npm i nspell dictionary-en` to enable it (or just `npm i nspell` if every `spelling` rule supplies its own `dictionary` path).
1570
+ - `nspell` + `dictionary-en` - Hunspell-compatible spell checker (and its bundled English dictionary) backing the `spelling` assertion (see [Spelling Assertions](#spelling-assertions-spelling) above). **Not installed by installing `@redocly/recheck`** — both are declared `optional: true` in `peerDependenciesMeta`, loaded lazily via dynamic `import()` only when a config actually enables `spelling`.
1571
+ Run `npm i nspell dictionary-en` to enable it (or just `npm i nspell` if every `spelling` rule supplies its own `dictionary` path).
1334
1572
 
1335
1573
  ### Development Dependencies
1336
1574
  - `vitest` - Modern testing framework
@@ -1499,7 +1737,7 @@ The enhanced pattern matching checks patterns against:
1499
1737
 
1500
1738
  This allows flexible targeting while maintaining backward compatibility.
1501
1739
 
1502
- ## Contributing
1740
+ ## Contribute
1503
1741
 
1504
1742
  Want to add new assertions or improve existing ones?
1505
1743
  Check out our **[Contributing Guide](src/rules/CONTRIBUTING.md)** for: