polytypo 1.1.0 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (66) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +9 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +227 -6
  7. data/lib/polytypo/data/fixtures/el.json +9 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +28 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +690 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +9 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +50 -1
  13. data/lib/polytypo/data/fixtures/fr.json +282 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +47 -1
  21. data/lib/polytypo/data/fixtures/sv.json +9 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/apostrophe.rb +10 -1
  55. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  56. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  57. data/lib/polytypo/engine/rules/nbsp.rb +53 -20
  58. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  59. data/lib/polytypo/engine/rules/spaces.rb +8 -1
  60. data/lib/polytypo/errors.rb +3 -0
  61. data/lib/polytypo/modes/runner.rb +17 -0
  62. data/lib/polytypo/modes/spans.rb +30 -2
  63. data/lib/polytypo/modes/yaml.rb +312 -0
  64. data/lib/polytypo/version.rb +1 -1
  65. data/lib/polytypo.rb +126 -15
  66. metadata +31 -1
@@ -0,0 +1,1281 @@
1
+ # Modes — the L2 contract
2
+
3
+ **Not a rule.** No entry in `spec/rules/order.json`, no locale data, no edits of its own. This
4
+ document specifies how a `html`, `markdown` or `yaml` document is decomposed into processable
5
+ text, how the rule pipeline is applied to it, and how the result is reassembled. It is normative
6
+ for all five runtimes and is parser-agnostic by construction: `parse5`, `nokogiri`, `lxml`,
7
+ `golang.org/x/net/html` and PHP's DOM disagree about almost everything this document does not
8
+ forbid them from doing. `yaml` mode is parser-**free** rather than parser-agnostic, for the
9
+ reason §3.8.1 measures.
10
+ **Spec version:** 1.3.0 (0.1.0 for everything except §3.3's class-membership table rows for
11
+ `nbsp` and `apostrophe`, split in 1.2.0, and §3.8, added in 1.3.0).
12
+
13
+ ---
14
+
15
+ ## 1. Purpose
16
+
17
+ `text` mode hands the pipeline one string. `html`, `markdown` and `yaml` hand it a document in
18
+ which _most_ of the characters must not be touched at all — markup, code, URLs, keys — and the
19
+ processable text is scattered across dozens of disconnected fragments. Two things then have to
20
+ be decided and neither is obvious: **what the rules see**, and **what comes out**.
21
+
22
+ The second is the easier one and this document answers it absolutely: **the output is the input
23
+ with a set of disjoint substring replacements applied, and nothing else.** The source is
24
+ _located_, never re-emitted. That is the only formulation under which five parsers can agree,
25
+ because it removes serialisation from the contract entirely — and in `yaml` mode it is also what
26
+ makes an entire class of reported corruption unreachable (§3.8.1).
27
+
28
+ The first is the interesting one, it is argued rather than asserted in §3.2, and everything
29
+ else in this document follows from it.
30
+
31
+ ---
32
+
33
+ ## 2. Locale data consumed
34
+
35
+ **None.** The mode layer is locale-independent. It decides _which characters_ the pipeline
36
+ sees; the pipeline decides what to do with them.
37
+
38
+ ---
39
+
40
+ ## 3. Algorithm
41
+
42
+ ### 3.1 Definitions
43
+
44
+ - A **skipped region** is a maximal span of the source document that the pipeline must never
45
+ see and must never modify.
46
+ - A **processable span** is a maximal span of the source that is not skipped. Each span is
47
+ identified by its offsets in the **original source**, and those offsets are the only handle
48
+ the mode layer keeps.
49
+ - The **span sequence** `S₁ … Sₘ` is the processable spans in document order.
50
+ - `text` mode is the degenerate case: one span covering the whole input, no skipped regions.
51
+ Every statement below holds for it trivially.
52
+
53
+ ### 3.2 The span model — the decision everything follows from
54
+
55
+ Three models are available. The choice is not a matter of taste; two of them produce wrong
56
+ output on ordinary documents.
57
+
58
+ **Model A — per-span, independent.** Run the whole pipeline on each span separately. Simple,
59
+ parallelisable, and **wrong**. Consider the entirely ordinary
60
+
61
+ ```html
62
+ "He said <em>'hi'</em> loudly"
63
+ ```
64
+
65
+ The spans are `"He said `, `'hi'` and ` loudly"`. Processed independently, the two double
66
+ quotes are unmatched in their own spans and stay straight, while `'hi'` pairs _in isolation_ at
67
+ **depth 1** — so it takes the locale's **primary** glyphs and comes out `“hi”` in `en-US`, where
68
+ the correct answer is the secondary pair `‘hi’` nested inside a converted outer quotation. That
69
+ is not a miss, it is visibly wrong output on a construction that appears in every second
70
+ paragraph of edited prose. Model A is rejected.
71
+
72
+ **Model B — naive concatenation.** Concatenate the spans, run the pipeline, redistribute by
73
+ offset. This fixes nesting, and introduces two defects of its own, both from **adjacencies that
74
+ do not exist in the document**:
75
+
76
+ - `"a"<code>x</code>"b"` concatenates to `"a""b"`. The `quotes` pass 1 same-kind adjacency veto
77
+ (`quotes.md` §3.2) sees two identical adjacent marks that are _not_ adjacent in the document,
78
+ vetoes both, and the surviving outer pair then quotes across the whole construction:
79
+ `“a""b”`. Wrong output.
80
+ - `He said <code>x</code> "hi"` concatenates to `He said "hi"` with two spaces that render as
81
+ one and are not adjacent in the source. `spaces` collapses them, deleting a character from a
82
+ text node that had a single space in it.
83
+
84
+ Model B is rejected. Note that both defects are invisible to every per-rule argument, because
85
+ each rule is behaving exactly as specified on the input it was given.
86
+
87
+ **Model C — concatenation with an explicit boundary marker. Adopted.** The spans are
88
+ concatenated with a **boundary marker** between each adjacent pair. The pipeline runs once, over
89
+ the whole marker-separated array. Edits are then redistributed to spans by offset.
90
+
91
+ The marker gives the rules what Model A denies them — the knowledge that `'hi'` sits inside a
92
+ larger quotation — while denying them what Model B wrongly grants: the belief that the last
93
+ character of one span touches the first character of the next.
94
+
95
+ **Markers are negative integers in the code-point array**, not Unicode code points. Every
96
+ runtime's array-of-code-points representation (`number[]`, `[]rune`, `list[int]`, `int[]`) holds
97
+ them without collision, and no input text can contain them. Implementations must not substitute
98
+ real code points — a private-use character or a noncharacter — because that reintroduces the
99
+ possibility of collision with author content and makes the classification table below a lie.
100
+
101
+ **There are two markers, and which one is used is decided by the source text of the gap:**
102
+
103
+ | Marker | Used when | Classified as |
104
+ | ------------------------ | ------------------------------------------------------------------------ | --------------------------------------------------- |
105
+ | **−1** — inline boundary | the skipped region between the two spans contains **no** line terminator | §3.3 |
106
+ | **−2** — line boundary | the skipped region contains **at least one** line terminator | a member of **`BREAK`**, for every rule, everywhere |
107
+
108
+ The test is on the raw source bytes of the gap, so it is decidable without asking the parser
109
+ anything and is identical in five runtimes. `a<em>x</em>b` gives −1; `foo\nbar`, `foo\n\nbar`
110
+ and `<p>a</p>\n<p>b</p>` all give −2.
111
+
112
+ **Why a line boundary is a `BREAK` and not just another opaque marker.** Every rule already has
113
+ correct, specified, fixture-covered behaviour at a line terminator, because `text` mode has real
114
+ ones: `spaces` refuses to touch a run that borders a `BREAK` (`spaces.md` §3.2 step 4), `dashes`
115
+ declines a token with a `BREAK` at `cp[L]`/`cp[R]` (§3.2 step 5), `quotes` counts `BREAK` inside
116
+ `SPACELIKE`, and `nbsp` never inserts against one. Reusing that behaviour is free and makes a
117
+ mode's output on hard-wrapped prose **identical to `text` mode's on the same characters**, which
118
+ is the property a mode that silently converts less than `text` cannot claim. See §7.7 for the
119
+ alternative that was rejected.
120
+
121
+ It also removes a parser dependency that would otherwise have been load-bearing. A Markdown hard
122
+ break is two spaces before a line ending; whether those spaces land inside a span or in a
123
+ structural token is a decision each parser makes differently. With a −2 marker the question does
124
+ not arise: the spaces sit next to a `BREAK`, and `spaces` protects a run bordering a `BREAK` in
125
+ every mode and every runtime.
126
+
127
+ ### 3.3 How the marker is classified
128
+
129
+ Every rule already decides from character classes. The **−1 inline marker's** membership is
130
+ fixed here, once, and is normative for all rules present and future.
131
+
132
+ > **The class names below are per-rule, not spec-wide.** Each rule document defines its own
133
+ > `SPACELIKE`, `OPENISH` and `CLOSEISH`, and they differ: `nbsp`'s `SPACELIKE` is the largest
134
+ > (it includes the fixed-width spaces), and its `OPENISH`/`CLOSEISH` are **locale-data-driven**,
135
+ > since they contain the declared quote glyphs. Read each row as "the marker is a member of that
136
+ > rule's class of this name, wherever the rule defines one" — not as a claim that one class of
137
+ > that name exists. Conflating two same-named classes across documents is precisely what
138
+ > produced the `STRIP-BEFORE` divergence recorded in `dashes.md` §3.2 step 9. (The **−2 line marker** is
139
+ simply a member of `BREAK` and of nothing else, so it needs no table: every rule's existing
140
+ `BREAK` handling applies to it unchanged.)
141
+
142
+ | Class family | Marker is a member? |
143
+ | ------------------------------------------------------------------------------------------- | ------------------------------------------- |
144
+ | `SPACELIKE`, `SP`, `NOBREAK`, `OTHER-SPACE`, `BREAK` | **no** |
145
+ | `LETTER`, `DIGIT`, `ALNUM`, `UPPER`, `WORDISH`, `ROMAN` | **no** |
146
+ | `DASH`, `INERT-DASH`, `DASHISH`, `SENTENCE-DASH`, `HY`, `NBHY` | **no** |
147
+ | `DOTLIKE`, `STRIP-BEFORE`, `TERMINAL` | **no** |
148
+ | `STRAIGHT`, `SQ`, `DQ` | **no** |
149
+ | `OPEN-BRACKET`, `CLOSE-BRACKET` | **no** |
150
+ | any literal matching list (`abbreviations`, `beforeUnits`, `hyphen.*`, the trademark table) | **no** — the marker never matches a literal |
151
+ | `OPENISH` **and** `CLOSEISH` (`quotes`, `apostrophe`) | **yes, both** |
152
+ | `CLOSEISH` (`nbsp`) | **yes** |
153
+ | `OPENISH` (`nbsp`), `OPENQUOTE` (`apostrophe`) | **no** |
154
+
155
+ and one exemption:
156
+
157
+ > In `quotes` §3.2's `canOpen` right-test, which rejects a `right` in `CLOSEISH`, the marker
158
+ > receives **the same exemption `STRAIGHT` receives** and does not disqualify.
159
+
160
+ Everywhere else the marker is opaque content: it is "a content character" and nothing more —
161
+ **with one exception, which is normative and which resolves a contradiction between this
162
+ document and `spaces.md`.**
163
+
164
+ > ### Edge tests — the marker behaves as `NONE`
165
+ >
166
+ > Where a rule asks not _"what character is here"_ but _"am I at the edge of the text I am
167
+ > allowed to modify"_, the marker behaves as **`NONE`**, exactly as the end of the array does.
168
+ >
169
+ > There is currently **exactly one such test in the whole spec**: the boundary guard in
170
+ > `spaces.md` §3.2 step 4. Every other test in every other rule asks the first question, and the
171
+ > marker remains opaque content for all of them.
172
+
173
+ **Why the exception exists, and why it is not a licence to add more.** Read without it, this
174
+ document said a span-final run of spaces has content on both sides, so `spaces` may collapse or
175
+ delete it. `spaces.md` step 4 said the opposite in the same breath, and named the reason — a
176
+ trailing run "in `html` mode [is] frequently the only separator between two inline elements".
177
+ Both cannot be true. The contradiction is resolved in favour of `spaces.md` for three reasons:
178
+
179
+ - **`spaces` is the only rule that deletes.** Its boundary guard exists because deletion at an
180
+ edge is irreversible and the rule cannot see past the edge to know whether the character
181
+ mattered. At a span edge it genuinely cannot: there is markup there. Every other rule either
182
+ replaces one-for-one or is already constrained by §3.4, so none of them needs the exception and
183
+ none of them may claim it.
184
+ - **The literal reading is destructive, not merely different.** `<em>mot</em> !` in `fr` has its
185
+ space deleted by `spaces` (`!` is in `STRIP-BEFORE`), after which `nbsp`'s U+202F insertion is
186
+ discarded by the edge-growth rule because the insertion point is now a span edge. Output:
187
+ `<em>mot</em>!` — a character removed and nothing put back — where `text` mode on the same
188
+ prose gives `mot⍹!`. §7.4 calls a mode that diverges from `text` on identical characters
189
+ indefensible, and this would have been the sharpest instance of it. Under the exception the
190
+ space survives, `nbsp` converts it in place (1 → 1, permitted at an edge), and the two modes
191
+ agree.
192
+ - **It makes the round-trip guarantee stronger, never weaker.** The exception only ever causes
193
+ `spaces` to do less: `a <!--x--> b` is returned untouched instead of collapsed. Fewer edits
194
+ to bytes the rule cannot fully see is the direction §4 already points in.
195
+
196
+ **The generalisation, for future rules:** a rule that _deletes_ must treat a span edge as the end
197
+ of the text. A rule that replaces or inserts must not — §3.4 governs it instead. If a rule is
198
+ ever added that deletes, it inherits this clause and must say so here.
199
+
200
+ **Why `nbsp` differs (spec 1.2.0).** `nbsp` reads `CLOSEISH` in one place only, the right-context
201
+ guard of N1/N2 (`nbsp.md` §3.3 step 2). Membership there is what lets the no-break space come back
202
+ after `spaces` deleted the typed one before a mark that ends an inline element:
203
+ `<strong>Label :</strong>` in French. It reads `OPENISH` at the quote-glyph guard (`nbsp.md` §3.3
204
+ step 3), in the left-boundary tests of N3, N7, N9 and N10, and in N3's following-token guard. At
205
+ the quote-glyph guard, membership would lose the narrow space in `<em>non</em> !`. At the other
206
+ tests nobody has measured its effect.
207
+ Up to spec 1.1.0 this row read "yes, both" for `nbsp` too, while every runtime implemented
208
+ "neither". `nbsp.md` §7 item 12 records how the split was decided. `apostrophe`'s `OPENQUOTE`
209
+ (`apostrophe.md` §3.1, spec 1.2.0) excludes the marker for a simpler reason: case 3 already
210
+ accepts it through `CLOSEISH`, so membership would change nothing.
211
+
212
+ **Why dual `OPENISH`/`CLOSEISH` membership plus the exemption.** These three settings are what
213
+ make quotation marks pair correctly across an inline element, and they were derived by working
214
+ the cases, not by analogy:
215
+
216
+ | Input | Concatenation | Result |
217
+ | -------------------------------- | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
218
+ | `"<em>hello</em>"` | `"⟦hello⟧"` | the opening mark's `right` is the marker; it is exempted from the `CLOSEISH` rejection, so `canOpen` holds. The closing mark's `left` is the marker — not `NONE`, not `SPACELIKE` — so `canClose` holds. **They pair.** |
219
+ | `"a"<code>x</code>"b"` | `"a"⟦"b"` | the second mark's `right` is the marker, in `CLOSEISH`, so `canClose` holds and `"a"` pairs. The third mark's `left` is the marker, in `OPENISH`, so `canOpen` holds and `"b"` pairs. **Two pairs, correctly.** |
220
+ | `"He said <em>'hi'</em> loudly"` | `"He said ⟦'hi'⟧ loudly"` | outer pair at depth 1 → primary; inner pair enclosed by it at depth 2 → **secondary**. The Model A defect is gone. |
221
+
222
+ Treating the marker as `NONE` instead — the intuitive choice, "a span edge is like the edge of
223
+ the text" — fails the first row: both marks are dropped and nothing converts. Treating it as
224
+ plain opaque content fails the second: the closing mark of `"a"` is not `canClose` because its
225
+ `right` is in none of the accepted classes. Only the dual membership plus the exemption passes
226
+ all three.
227
+
228
+ ### 3.4 Every other rule gets the right behaviour for free
229
+
230
+ Because the marker is opaque content and is in none of their classes (apart from the memberships
231
+ in the table of §3.3), the remaining rules decline to work across a boundary **without any
232
+ special-casing**:
233
+
234
+ | Situation | Concatenation | Outcome |
235
+ | ----------------------------- | ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
236
+ | `<em>foo </em>- bar` | `foo ⟦- bar` | `dashes`: the code point left of the dash is the marker, not a space, so `lsp = 0` while `rsp = 1` — the symmetry guard declines. Miss, not damage |
237
+ | `He said <code>x</code> "hi"` | `He said ⟦ "hi"` | `spaces`: two runs of length 1 separated by the marker, not one run of 2. No collapse |
238
+ | `5 <em>km</em>` | `5 ⟦km` | `nbsp` N5 requires `cp[a-1]` to be `SP` or `NBSP`; it is the marker. Declines |
239
+ | `из<em>-под</em>` | `из⟦-под` | `hyphen`: no listed form matches across the marker. Declines |
240
+ | `(c<em>)</em>` | `(c⟦)` | `symbols`: the literal `(c)` does not match. Declines |
241
+
242
+ **No rule can produce an edit whose span contains a marker.** Every rule's edit span is a
243
+ contiguous run of characters drawn from classes the marker does not belong to, or a literal
244
+ match the marker cannot participate in. This is what makes redistribution unambiguous, and it
245
+ is a **proof obligation on every future rule**: if a new rule could match across a marker, it
246
+ must state how its edits are redistributed. As a safety net, an implementation that computes an
247
+ edit whose span contains a marker **must discard that edit** and should report it as a bug.
248
+
249
+ **The edge-growth rule.** An edit is described by the span it replaces, `cp[p … q]` in the
250
+ concatenation, and a replacement sequence of length `r`. For an insertion, `q = p - 1` and the
251
+ replaced length is zero. Every edit lies wholly within one processable span `S = cp[s₀ … s₁]`.
252
+
253
+ > **An edit is discarded if it would place code points at an extremity of its span that were
254
+ > not there before.** Formally, with `d = q - p + 1` the replaced length and `w` the replacement:
255
+ >
256
+ > - if `p = s₀` and (`r > d`, **or** `r > 0` and `w[0]` is U+0020 and `cp[p]` is not U+0020)
257
+ > → **discard**;
258
+ > - if `q = s₁` and (`r > d`, **or** `r > 0` and `w[r−1]` is U+0020 and `cp[q]` is not U+0020)
259
+ > → **discard**.
260
+ >
261
+ > An insertion has `d = 0`, so `r > d` always holds and an insertion is discarded exactly when
262
+ > its position coincides with a span edge. That is what this section said before; the rule below
263
+ > generalises it.
264
+
265
+ **The second clause, and why the length test alone was not the rule it claimed to be (spec
266
+ 1.3.0).** The sentence above is the rule; `r > d` was an incorrect formalisation of it, and the
267
+ gap is **`r = d`**. `dashes` P3 admits a run of two _or three_ dashes, so a tight `---` is a
268
+ 3 → 3 replacement in every `-spaced` locale: the length test sees nothing, and U+0020 lands on
269
+ both extremities of the span anyway. Measured, and both cases are ones this document already
270
+ claims to have closed:
271
+
272
+ | Mode | Input | Locale | Before the second clause | What breaks |
273
+ | ---------- | ----------------- | ------- | ------------------------ | --------------------------------------------------------- |
274
+ | `html` | `a<em>---</em>b` | `de-DE` | `a<em> – </em>b` | the element begins and ends with a space it never held — the harm §7.3 exists to prevent |
275
+ | `markdown` | `x *---* y` | `de-DE` | `x * – * y` | the asterisks are de-flanked, so the span partition of §5 item 2 is not stable |
276
+ | `yaml` | `description: a:---b` | `de-DE` | `description: a: – b` | **the document no longer parses** — `: ` is now a mapping indicator |
277
+
278
+ The second clause tests the **character**, not the length, and it tests only U+0020 because
279
+ U+0020 is the only code point any rule emits whose meaning comes from its position rather than
280
+ from itself (§5 item 2). It leaves every case in the table below unchanged: a one-for-one
281
+ `"` → `“` still applies at an edge, a contraction still applies, and a `-spaced` edit that
282
+ already spans the surrounding spaces still applies, because there a U+0020 replaces a U+0020.
283
+
284
+ **This is a behaviour change to `html` and `markdown`, which shipped at v1.0.0**, not only to the
285
+ mode added in 1.3.0: three dashes at a span extremity are now left alone in a `-spaced` locale,
286
+ exactly as two already were. It needs its own changelog entry, in the terms the rest of the
287
+ public copy uses.
288
+
289
+ **Why it had to be generalised, and what it costs.** The previous formulation discarded
290
+ _insertions_ only. But a rule can grow a span with a **replacement**: `dashes` emits a `-spaced`
291
+ form by replacing the dash token with `U+0020 – U+0020`, which is one code point becoming three.
292
+ When the dash is the whole span, the two new spaces land on the span's edges, and the previous
293
+ filter did not see them because no insertion occurred. Three consequences, all observed:
294
+
295
+ - **html** — `a<em>--</em>b` became `a<em> – </em>b`. The element now begins and ends with a
296
+ space it never contained: the precise harm §7.3 exists to prevent, arriving by a route §7.3
297
+ did not cover.
298
+ - **markdown, span partition** — `*–*` became `* – *`, which de-flanks the asterisks so that on
299
+ the next run they are literal content rather than emphasis delimiters. §5 item 2 requires the
300
+ span partition to be stable; it was not.
301
+ - **markdown, unbounded growth** — `a\n\n–\n\n'` became `a\n\n – \n\n'` and then
302
+ `a\n\n – \n\n'`. The emitted space migrates into the line prefix, which is _outside every
303
+ span_, so the next run starts from a fresh span and emits another one. The document grows
304
+ without bound on every re-processing — which for content re-processed on each save is not a
305
+ cosmetic defect.
306
+
307
+ Both clauses are mechanically checkable from `(p, q, r, s₀, s₁)` and the replacement's own first
308
+ and last code points, with no rule cooperation and no knowledge of Markdown, HTML or YAML syntax.
309
+ Together they distinguish exactly the cases that matter:
310
+
311
+ | Edit | At an edge? | `d` → `r` | Verdict |
312
+ | -------------------------------------- | ----------- | --------- | ----------------------------------------------------------- |
313
+ | `"` → `“` (`quotes`) | yes | 1 → 1 | **applied** — a one-for-one replacement never grows an edge |
314
+ | `--` → `␣–␣` (`dashes`, `-spaced`) | yes | 2 → 3 | **discarded** by the length clause |
315
+ | `---` → `␣–␣` (`dashes` P3, `-spaced`) | yes | 3 → 3 | **discarded** by the character clause — the length clause misses it entirely |
316
+ | `a--b` → `a␣–␣b` (same edit, interior) | no | 2 → 3 | **applied** — growth is only a problem at an extremity |
317
+ | `␣--␣` → `␣–␣` (the edit spans its own spaces) | yes | 4 → 3 | **applied** — a U+0020 replaces a U+0020, so no edge changed |
318
+ | `(c)` → `©` (`symbols`) | yes | 3 → 1 | **applied** — shrinking is always safe |
319
+ | U+0020 → U+00A0 (`nbsp`) | yes | 1 → 1 | **applied** — U+00A0 is not U+0020 |
320
+ | insertion of U+202F (`nbsp` N1/N2) | yes | 0 → 1 | **discarded**, as before |
321
+
322
+ **Deletion at an edge is not restricted**, and does not need to be. A rule can only delete
323
+ U+0020 (`spaces`), and a deletion cannot bring into existence a structural token that was not
324
+ already there — it can only remove a character from inside a span. The whitespace that _is_
325
+ structural in Markdown — a line prefix, a list marker's indentation, the two spaces of a hard
326
+ break — is protected by the −2 line marker (§3.2): it sits beside a `BREAK`, and `spaces`
327
+ refuses to touch a run bordering one, in every mode.
328
+
329
+ `fr` `mot<em>!</em>` therefore still keeps no narrow no-break space, for the reason §7.3 gives,
330
+ and `a<em>--</em>b` is now left alone as well in every `-spaced` locale (§7.9).
331
+
332
+ (The original witness for this defect was `a<em>–</em>b`, with an authored en dash. It is still a
333
+ fixed point, but it no longer demonstrates anything: `dashes` §3.2 step 2a declines every token
334
+ containing an authored U+2013 or U+2014, so that input is now refused a step earlier and never
335
+ reaches the edge-growth rule. `--` is the live carrier.)
336
+
337
+ ### 3.5 Applying the result
338
+
339
+ 1. Extract the span sequence from the source. Record each span's **source offsets**.
340
+ 2. Build the code-point array `S₁ ⌢ [m₁] ⌢ S₂ ⌢ [m₂] ⌢ … ⌢ Sₘ`, where each `mₖ` is **−1 or −2
341
+ by §3.2's test on the raw source of that gap** — not always −1. In `yaml` the −2 is the common
342
+ case, since a block scalar's spans are separated by a line terminator.
343
+ 3. Run the pipeline **once**, in `order.json` order, over that array.
344
+ 4. Each edit lies wholly within one span (§3.4). Map it back to source offsets.
345
+ 5. Emit the **original source bytes**, with those replacements applied and nothing else changed.
346
+
347
+ Step 5 is the round-trip guarantee and is stated in full in §4.
348
+
349
+ ### 3.6 Skip list — `html`
350
+
351
+ Skipped, exhaustively and exactly:
352
+
353
+ - the **entire subtree** of `code`, `pre`, `kbd`, `samp`, `var`, `script`, `style`, `textarea`,
354
+ `svg`, `math` — including any nested elements, whatever they are. A `<pre><em>x</em></pre>`
355
+ is skipped whole;
356
+ - **every attribute**, name and value, of every element, without exception;
357
+ - **every well-formed character reference** — `&name;`, `&#1234;`, `&#x2014;` — as an opaque
358
+ unit. A text node containing one is split into spans around it, so the reference's spelling is
359
+ preserved exactly, which is the only way `&nbsp;` does not become a literal U+00A0 and back.
360
+ **A bare `&` that begins no well-formed reference stays inside its span.** No rule emits or
361
+ deletes `&`, so it is in no danger; lifting it out would create a boundary that suppresses
362
+ conversions on both sides of it for nothing — `Tom & Jerry's "book"` would lose the pairing of
363
+ its quotation marks. Well-formedness, not the presence of an ampersand, is what makes a
364
+ skipped region;
365
+ - comments, CDATA, doctype, processing instructions, and any prologue.
366
+
367
+ **Element names compare case-insensitively**, per HTML: `<CODE>`, `<Code>` and `<code>` are all
368
+ skipped. (In MDX, JSX names compare case-sensitively — §3.7.3.)
369
+
370
+ `svg` and `math` are skipped although neither appears in PLAN.md §3.2's list. Neither is prose:
371
+ in MathML a quotation mark, a hyphen and a prime are **operators and identifiers**, where
372
+ substituting a curly glyph or an en dash changes what the expression means, not how it looks.
373
+ The accepted cost is that `<svg><text>` and `<svg><title>` do hold real prose and are now left
374
+ untypeset. That asymmetry is deliberate — widening a skip list later is additive, while
375
+ narrowing one after release breaks every document that depended on the wider behaviour.
376
+
377
+ Everything else is processable, **including unknown and custom elements** (`<my-callout>`,
378
+ `<Foo>`). An unknown element is far more likely to be a wrapper than a code container, and
379
+ guessing from its name is exactly the kind of heuristic that behaves differently in five
380
+ runtimes. **The skip list is closed**: extending it is a spec change, not an implementation
381
+ decision.
382
+
383
+ ### 3.7 Skip list — `markdown`
384
+
385
+ #### 3.7.1 The dialect is chosen by the caller, never detected
386
+
387
+ `markdown` is not one language. CommonMark and MDX disagree on ordinary documents: **MDX has no
388
+ indented code blocks and no `<https://…>` autolinks; CommonMark has no `{…}` expressions and no
389
+ JSX.** A document is frequently valid in both and means different things in each.
390
+
391
+ > **`markdown` mode takes a required `dialect` option**, one of:
392
+ >
393
+ > | value | language |
394
+ > | -------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
395
+ > | `"commonmark"` | CommonMark 0.31 **plus GFM** — tables, strikethrough, task lists, autolink literals |
396
+ > | `"mdx"` | MDX 3 — CommonMark plus GFM, **minus** indented code blocks and `<…>` autolinks, **plus** JSX elements and `{…}` expression containers |
397
+ >
398
+ > It has **no default**, and omitting it when `mode` is `"markdown"` throws, exactly as an
399
+ > omitted `locale` does (PLAN.md §5.1). `dialect` is ignored in `text` and `html` modes.
400
+
401
+ **Detection is forbidden**, and that is the important half. A heuristic is available and looks
402
+ reasonable — parse as MDX, and if it succeeds _and_ finds an MDX construct, call it MDX — but it
403
+ is silently wrong in exactly the case that matters. **One `<https://example.com>` autolink makes
404
+ a file invalid MDX**, so the heuristic falls back to CommonMark, and
405
+ `export const meta = {slug: "une-note"}` is then no longer an ESM statement but a **paragraph**:
406
+ it becomes processable text, and `fr` puts a narrow no-break space before its colon. A silent
407
+ false positive inside a machine-read field, in the author's own content format, produced by a
408
+ dialect the caller never chose.
409
+
410
+ The caller always knows which dialect they have — it is the file extension — and the library
411
+ never can. Requiring the option moves an unanswerable question to the party that holds the
412
+ answer, which is the reasoning that already makes `locale` required. A separate mode id for MDX
413
+ was rejected because MDX **is** Markdown with two constructs swapped, so an option on the mode
414
+ it varies says what is true; spec 1.3.0 later added `yaml` as a fourth mode id (§3.8), and the
415
+ two decisions do not conflict — YAML has no mode to hang off, and a `dialect` on nothing is not
416
+ a shape this contract has. **Ratified by the operator as public contract**: the throw carries
417
+ its own code, `POLYTYPO_INVALID_DIALECT`. See §7.8.
418
+
419
+ #### 3.7.2 A document that does not parse in its declared dialect
420
+
421
+ > **`transform` throws, carrying the stable code `POLYTYPO_MALFORMED_INPUT`.** The parser's own
422
+ > error type must never escape. **Ratified by the operator as public contract** — this code and
423
+ > `POLYTYPO_INVALID_DIALECT` are part of the taxonomy, not proposals.
424
+
425
+ **This can only happen in one place, and knowing that is what decides it.** Neither of the other
426
+ two languages can fail: HTML parsing is specified with total error recovery, and **every byte
427
+ sequence is valid CommonMark** — there is no such thing as a CommonMark syntax error. Only
428
+ `dialect: "mdx"` can reject a document, and only because MDX embeds JavaScript: an unterminated
429
+ JSX element, a malformed `{…}` expression, a broken `export`. A document that fails there is one
430
+ the author's own build already refuses.
431
+
432
+ Given that, returning the input untouched is the worse option. It would mean polytypo silently
433
+ succeeding on a file that is broken, hiding a build error behind a typography pass, and — the
434
+ common case — hiding the caller's own mistake of naming the wrong dialect. §3.7.1 forbids
435
+ detection precisely so that a wrong dialect is the caller's error rather than the library's
436
+ guess; swallowing the consequence would give that decision back with none of the information.
437
+ PLAN.md §5.1 fails fast on an unknown locale for the same reason, and this is the same shape.
438
+
439
+ Two constraints on the throw:
440
+
441
+ - **The runtime's parser error must be wrapped, never propagated.** A `VFileMessage`, a
442
+ `Nokogiri::SyntaxError` or a Python exception on the public surface puts a dependency's type in
443
+ the contract and is unreproducible in the other four runtimes. Only the code is contractual
444
+ (ARCHITECTURE.md §4.6); a message and a source position are useful and are not part of it.
445
+ - **This is the only way input can make `transform` throw.** Every other throw is caused by
446
+ options or by locale data. `transform` was previously total on its input, and it no longer is
447
+ — that is a real change in the shape of the public contract, and it is why the new code needs
448
+ sign-off (§7.8) rather than being an implementation detail.
449
+
450
+ #### 3.7.3 What is skipped
451
+
452
+ Skipped, exhaustively:
453
+
454
+ - **frontmatter** — a metadata block at the very start of the document, delimited by `---`
455
+ (YAML) or `+++` (TOML), skipped whole including its delimiters. Without it the closing `---`
456
+ reads as a setext underline, `title: Une note` becomes a paragraph, and `fr` inserts a narrow
457
+ no-break space before the colon of a machine-read field. That is a guaranteed false positive
458
+ on the M4 corpus (PLAN.md §8), where every file opens with frontmatter;
459
+ - fenced code blocks, including the info string and the fences;
460
+ - indented code blocks — **`commonmark` only**; MDX has none;
461
+ - inline code spans, including the backticks;
462
+ - autolinks `<https://…>` — **`commonmark` only**;
463
+ - **link and image destinations and titles** — the `(…)` of `[text](url "title")`, and the
464
+ definition line of a reference link. The **link text** is processable;
465
+ - **GFM constructs**: table pipes, delimiter and alignment rows, task-list checkboxes,
466
+ strikethrough delimiters, and footnote definition labels. GFM is enabled in **both** dialects,
467
+ and saying so is not decoration: §4's promise that table alignment rows survive is empty
468
+ unless tables are recognised at all, and §6's `[P: html, markdown]` marking on `dashes`' URL
469
+ bullet holds **only** because GFM autolink literals make a bare `https://…` a link — in pure
470
+ CommonMark a bare URL is ordinary text and would be typeset;
471
+ - HTML blocks and inline raw HTML, which are handed to the `html` skip list of §3.6 rather than
472
+ processed as markdown. **Inline HTML arrives as isolated tags with Markdown between them**,
473
+ not as a tree, so §3.6's subtree rule must be implemented as a **stack of open skipped
474
+ elements**: push on a start tag whose name is in the skip list, pop on its matching end tag,
475
+ and treat everything as skipped while the stack is non-empty. An unclosed skipped start tag
476
+ skips to the end of the block. Without the stack, `<code>a *b* c</code>` leaks its Markdown
477
+ middle back into a span;
478
+ - **MDX only**: JSX expression containers `{…}` in full, and every JSX attribute. **JSX element
479
+ children are processable** — `<Callout>Some text here</Callout>` is prose the author wrote and
480
+ wants typeset, and MDX is the author's own content format (PLAN.md §8, M4).
481
+
482
+ **Element names compare case-sensitively in JSX and case-insensitively in HTML.** `<Code>` is an
483
+ MDX component and is **not** in the skip list; `<code>`, `<CODE>` and `<Code>` in raw HTML all
484
+ are. This follows the two languages rather than any choice of ours, and an implementation that
485
+ lower-cases JSX names before matching will skip a component's children.
486
+
487
+ Nesting follows the same rule as `html`: a skipped construct is skipped whole, including
488
+ anything that looks processable inside it.
489
+
490
+ ### 3.8 Skip list — `yaml`
491
+
492
+ **Spec 1.3.0.** `yaml` is the fourth mode id, and PLAN.md §3.2's "exactly three modes" is amended
493
+ by that decision. Two things about it decide everything below, and neither is true of the other
494
+ two modes: **no parser is used** (§3.8.1), and **the caller names the keys whose values hold
495
+ prose** (§3.8.2). The first is forced by what five ecosystems can report; the second by what YAML
496
+ is.
497
+
498
+ #### 3.8.1 Why a scanner, and why the issue's stated blocker is not one
499
+
500
+ The request (#18) arrived with a data-corruption report attached. A caller with an `openapi.yaml`
501
+ wrote their own splicer over a YAML library, and a block scalar's decoded value carries its
502
+ trailing line terminator only for some chomping indicators; they wrote the value back with the
503
+ wrong assumption and the following key was absorbed into the string. The file stayed
504
+ syntactically valid and became semantically wrong.
505
+
506
+ **That bug class is unreachable here, and it is unreachable by construction rather than by
507
+ care.** It is a property of decode → mutate → **re-encode**, and §1 already forbids the third
508
+ step: the source is located, never serialised. A chomping indicator is a byte of the source that
509
+ no span contains, so nothing this document permits can misread it, rewrite it, or lose it. The
510
+ same sentence disposes of indentation, anchors, tag handles and quoting style.
511
+
512
+ What the mode needs, then, is not a parser but a **locator**: something that reports the source
513
+ offsets of scalar content. The obvious move is to take those offsets from each ecosystem's YAML
514
+ library, and it does not survive contact with the five runtimes. Measured, 2026-09-20:
515
+
516
+ | Runtime | Library | Scalar positions |
517
+ | ------- | -------------------------- | -------------------------------------------------------------------- |
518
+ | JS | `yaml` | start **and** end, exact to the token |
519
+ | Python | PyYAML | start **and** end, exact to the token |
520
+ | Ruby | Psych | start **and** end, as line/column |
521
+ | Go | `gopkg.in/yaml.v3` | **start only** — `Node` carries `Line`/`Column` and no end position |
522
+ | PHP | `symfony/yaml` | **none** — the public surface decodes to values and reports no nodes |
523
+
524
+ Two of five cannot supply what the contract needs, and PHP has no alternative: `ext-yaml` binds
525
+ libyaml's value API and reports no positions either. A parser-backed `yaml` mode is therefore not
526
+ portable, and the §4 cost paragraph — "a runtime whose parser cannot report source offsets cannot
527
+ implement `html` mode conformantly" — would have excluded two runtimes outright rather than
528
+ describing a constraint they could meet.
529
+
530
+ So the locator is **specified here and hand-written per runtime**, like every core rule and like
531
+ locale resolution: a single left-to-right scan over the code-point array, no regular expressions,
532
+ no lookbehind. That is a cost — it is the first format knowledge polytypo owns rather than
533
+ delegates — and it buys the one thing delegation could not: the same spans in five runtimes.
534
+
535
+ #### 3.8.2 Why the caller names the keys, and why no heuristic can
536
+
537
+ > **`yaml` mode takes a required `keys` option: the list of mapping keys whose scalar values are
538
+ > processable.** It has **no default**, and omitting it when `mode` is `"yaml"` throws
539
+ > `POLYTYPO_INVALID_OPTION`, exactly as an omitted `dialect` throws in `markdown` mode. A value
540
+ > that is not a list of strings throws the same code. An **empty list is legal** and yields no
541
+ > spans — "process nothing" is a choice a caller may make, not an error.
542
+
543
+ The first draft of this section had no such option. It processed every string scalar and decided
544
+ prose by a content test — a scalar had to contain a space and a letter. Both the option and this
545
+ paragraph exist because that draft was **measured against this repository's own
546
+ `.github/workflows/`, and it corrupted them**: 37 of 106 spans changed, in `en-US`, `de-DE` and
547
+ `fr` alike. Two of the results:
548
+
549
+ ```
550
+ run: |
551
+ if ! git cat-file -e "$SHA"; then → if! git cat-file -e "$SHA"; then
552
+ name: Node ${{ matrix.node }} → name: Node ${{matrix.node}}
553
+ ```
554
+
555
+ The first is a shell script that no longer parses. Both passed the content test comfortably —
556
+ they have spaces and letters, because shell and template expressions are written in words.
557
+
558
+ **The reason no better content test exists is structural, and it is the whole argument for the
559
+ option.** `html` and `markdown` are prose formats with islands of code in them, so a closed skip
560
+ list works: the islands are marked by the syntax itself — a `<code>` element, a fence — and
561
+ naming them is a finite job this document can do once. **YAML is the inverse: a data format with
562
+ islands of prose in it**, and nothing in YAML's syntax distinguishes them. `description`,
563
+ `summary` and `title` hold sentences; `run`, `command`, `if`, `image`, `pattern` and `name` hold
564
+ things a machine reads. They are the same construct, spelled the same way, and they differ only
565
+ in what the schema above the YAML means by them — which the caller knows and the library cannot.
566
+
567
+ That is the same shape as the `dialect` decision (§3.7.1) and it is settled the same way:
568
+ **requiring the option moves an unanswerable question to the party that holds the answer.** The
569
+ request itself predicted it — "probably an option for which keys to process, since typesetting a
570
+ `url:` or an `id:` value is not wanted" — and the measurement above is what turned that from a
571
+ reasonable expectation into a requirement.
572
+
573
+ **Matching is exact and deliberately dumb**, so that five hand-written scanners cannot disagree:
574
+ a key matches iff its source text, with trailing U+0020 removed, equals a member of `keys` code
575
+ point for code point. No case folding — ARCHITECTURE.md §4.4 forbids locale-dependent case
576
+ operations and Turkish dotless ı is the standing reason. No paths, no globs, no wildcards, no
577
+ nesting: a key named `description` is processable wherever it occurs and at any depth. A **quoted
578
+ key never matches**, because §3.8.4 step 5 has already declined the line; that is an accepted
579
+ cost, recorded in §7.11.
580
+
581
+ **The option also removes a defect the content test carried**, which is why nothing of it
582
+ survives. A content test is a predicate on the text inside a span, and `spaces` can delete the
583
+ very U+0020 the predicate reads: `ref: ${{ steps.pin.outputs.sha }}` yielded one span on the
584
+ first run, `ref: ${{steps.pin.outputs.sha}}` on the second, and no span at all — §5's obligation
585
+ that the span partition be stable, **falsified on a one-line document**. `keys` is a predicate on
586
+ the key, which no rule can reach, so the partition is fixed by the source and the obligation
587
+ holds again. §5's yaml paragraph records this as the reason the predicate must stay structural.
588
+
589
+ #### 3.8.3 Skip by default — the rule the scan is built around
590
+
591
+ > **A construct the scan does not recognise with certainty yields no spans.** The worst outcome
592
+ > of a gap in the scan is prose left untypeset. It is never a changed byte.
593
+
594
+ This is the inverse of the posture `html` takes, where everything not in a closed skip list is
595
+ processable (§3.6). The asymmetry is deliberate and follows from §3.8.1: an HTML adapter has a
596
+ conforming parser telling it what every construct is, so "process what is not skipped" is a claim
597
+ it can back. The YAML scan has no such oracle, so it may only claim what it has proved. **`yaml`
598
+ mode names what is processable and skips the rest**, and the list below is closed: extending it
599
+ is a spec change.
600
+
601
+ A consequence worth stating: **`transform` never throws on input in `yaml` mode.** There is no
602
+ declared grammar to violate, so §3.7.2's `POLYTYPO_MALFORMED_INPUT` has no `yaml` counterpart. A
603
+ file that is not YAML at all yields few spans or none and comes back byte for byte. The only
604
+ throw this mode adds is the option throw of §3.8.2, which is about the call, not the input.
605
+
606
+ #### 3.8.4 The scan
607
+
608
+ The source is the code-point array of §3.1. A **line** is a maximal run containing no U+000A,
609
+ **and a U+000D immediately before that U+000A is not part of the line** — it is a terminator like
610
+ the U+000A itself, so it lies outside every span and comes back untouched. Line terminators are
611
+ never inside a span, so no rule can create or destroy one. `indent` is the number of leading
612
+ U+0020 on the line and `i` the index of the first code point after them.
613
+
614
+ The carriage-return clause is normative rather than obvious, and it is here because five runtimes
615
+ would otherwise inherit five answers from five standard-library calls. Without it a CRLF file
616
+ diverges from the same bytes with LF: the block header of §3.8.5 reads as `|` followed by U+000D
617
+ and is unrecognised, so the block yields no spans, while a plain scalar carries the carriage
618
+ return **inside** its span and hands it to the rules as content.
619
+
620
+ > **Normative, and the repair of a defect the first draft shipped: a line consumed by a construct
621
+ > is never scanned again.** Step 7 defines the one case — an inline value's continuation lines —
622
+ > and without it a multi-line quoted scalar, a multi-line flow collection and a folded plain
623
+ > scalar all leak their continuation lines back into the scan as if they were mappings.
624
+
625
+ Lines are processed in order.
626
+
627
+ 1. The line yields no spans if it is empty or all U+0020, or if it contains U+0009 anywhere. A
628
+ tab makes indentation undecidable, which is what every step below depends on.
629
+ 2. The line yields no spans if `cp[i]` is `#` (comment) or `%` (directive).
630
+ 3. The line yields no spans if it begins at `i` with `---` or with `...` and the next code point
631
+ is U+0020 or the line ends there. The trailing-content form is declined as well as the bare
632
+ one: `--- key: value` is a node introduced by a document marker, and reading the marker as part
633
+ of a key is how the first draft produced a key of `--- key`.
634
+ 4. **Block sequence entries are consumed, not skipped.** While `cp[i]` is `-` and `cp[i+1]` is
635
+ U+0020, advance `i` past the marker and past any U+0020 after it. Repeat — `- - key: v` is
636
+ legal and nests. If `i` reaches the end of the line, it yields no spans.
637
+ 5. **Find the key.** Scan `j` upward from `i`.
638
+ - If `cp[j]` is `"`, `'`, `{`, `[`, `&`, `*`, `!` or `#`, the line yields no spans. This test
639
+ applies at **every** `j`, not only at `j = i`: it declines a quoted, flow, anchored, aliased
640
+ or tagged key, and equally a key that merely contains one of those characters anywhere —
641
+ `a!b: value` yields nothing. The wider reading is the normative one, and saying so is what
642
+ stops a second implementer from testing only the first position and getting a different span
643
+ set. The cost is recorded in §7.11.
644
+ - If `cp[j]` is `:` **and** `j` is the last code point of the line or `cp[j+1]` is U+0020, the
645
+ key is `cp[i … j−1]` with trailing U+0020 removed, and the scan continues at step 6.
646
+ - **Otherwise `j` advances by one and the scan continues**, including when `cp[j]` is a colon
647
+ that is not followed by U+0020 or the line end — such a colon is an ordinary character of the
648
+ key, and `a:b: v` has the key `a:b`. Stating the else-branch is not pedantry: without it the
649
+ first draft admitted two readings of that line, which is the exact divergence §3.8.1 says
650
+ the scan exists to prevent.
651
+ - If no such `j` exists, or the key is empty, the line yields no spans.
652
+ 6. Let `v` be the index of the first code point after the colon that is not U+0020. **If there is
653
+ none, the value is empty**: the following, more-indented lines are a nested node and are
654
+ scanned on their own. The line yields no spans; go to the next line.
655
+ 7. Otherwise the line carries an **inline value**, and the **value run** — every following line
656
+ that is blank or indented more than `indent` — belongs to that value. Those lines are never
657
+ scanned as lines of their own. The only spans they can contribute are the ones step 9 takes
658
+ from a recognised block scalar; every other step below yields no spans for the whole run.
659
+ 8. **The key must be listed.** If the key of step 5 is not a member of `keys` (§3.8.2), the line
660
+ and its value run yield no spans. This is the step that makes `run:`, `if:` and `image:`
661
+ unreachable, and it is checked before the value's form so that an unlisted key costs nothing
662
+ to decline.
663
+ 9. `cp[v]` selects the form:
664
+ - `&`, `*` or `!` — anchor, alias or tag — and `{` or `[` — a flow collection: no spans.
665
+ - `#`: the value is absent and only a comment follows: no spans.
666
+ - `|` or `>`: §3.8.5, block scalar.
667
+ - `"` or `'`: §3.8.6, quoted scalar.
668
+ - anything else: §3.8.6, plain scalar.
669
+
670
+ #### 3.8.5 Block scalars — one span per content line, and the chomping indicator never enters
671
+
672
+ The header is `cp[v]` followed by **at most one** chomping indicator (`-` or `+`) and **at most
673
+ one** indentation indicator (`1`–`9`), in either order, then optional U+0020 and an optional `#`
674
+ comment, then the end of the line. Any other header is unrecognised and the block yields no
675
+ spans. **The header is never inside a span.**
676
+
677
+ The block's content is the value run of §3.8.4 step 7. Let `bi` be the indentation of its first
678
+ non-blank line, or `indent + n` when the header carried an explicit indentation indicator `n`.
679
+ **One definition, and three conditions that make the whole block yield no spans** — the block is
680
+ consumed either way, never rescanned:
681
+
682
+ - any line of the run contains U+0009;
683
+ - the run's first non-blank line is indented less than `bi`, which can only happen when an
684
+ explicit indicator disagrees with the block as written;
685
+ - any later non-blank line of the run is indented less than `bi`.
686
+
687
+ The last one is the case the first draft left with two readings — `k: |` followed by a line at
688
+ four spaces and then one at two — and it is settled by bailing rather than by choosing, because
689
+ either choice would be a guess about which line the author meant to be content.
690
+
691
+ > **Each non-blank content line contributes exactly one span, `[lineStart + bi, lineEnd)`.** Blank
692
+ > lines contribute none. The indentation is outside every span; so is every line terminator, and
693
+ > so is the run of line terminators at the end of the block that the chomping indicator governs.
694
+
695
+ That last clause is the direct answer to the report in §3.8.1: **the trailing newlines are not in
696
+ any span, so no `yaml`-mode output can add or remove one.** `|`, `|-`, `|+`, `>`, `>-` and `>+`
697
+ are handled identically here because the difference between them lives entirely in bytes the mode
698
+ never touches.
699
+
700
+ Extra indentation beyond `bi` on a content line stays **inside** that line's span, because in a
701
+ literal block it is content. It is safe there because each content line is its own span and the
702
+ gap between two of them contains a line terminator: the marker is a **−2**, a member of `BREAK`,
703
+ and `spaces` refuses to touch a run bordering one (`spaces.md` §3.2 step 4). The −1 marker's
704
+ `NONE` exception of §3.3 is a different mechanism and is not what protects it.
705
+
706
+ #### 3.8.6 Quoted and plain scalars
707
+
708
+ **Quoted.** The scalar must open and close **on the same line**; if the matching quote is not on
709
+ that line the value yields no spans, and its continuation lines are consumed by §3.8.4 step 7
710
+ rather than rescanned. After the closing quote only U+0020 and an optional `#` comment may
711
+ follow. The span is the content between the quotes, exclusive of both.
712
+
713
+ Two bails make source characters and content characters the same thing, which §3.1's offset model
714
+ requires:
715
+
716
+ - a double-quoted scalar whose content contains U+005C yields no spans — `\n`, `\"` and `é`
717
+ are content the source spells with more characters than it has;
718
+ - a single-quoted scalar whose content contains two consecutive U+0027 yields no spans, for the
719
+ same reason.
720
+
721
+ This is the treatment `html` gives a character reference (§3.6), reached from the same constraint
722
+ rather than by analogy. **No colon test applies to a quoted scalar** — quoting neutralises the
723
+ colon, and applying the plain-scalar test here is what made the first draft decline
724
+ `title: "Chapter 1: the beginning"`.
725
+
726
+ **Plain.** The scalar runs to the end of the line, minus a trailing comment — the first `#`
727
+ preceded by U+0020, and that U+0020 with it — and minus any remaining trailing U+0020. Then:
728
+
729
+ - **the continuation bail**: if the value run of §3.8.4 step 7 is not empty, the scalar is a
730
+ multi-line plain scalar and yields no spans. Its line folding is a construct this scan does not
731
+ claim;
732
+ - **compact nesting is not a value**: the scalar yields no spans if it begins with `-` followed by
733
+ U+0020 or the line end (`key: - item` opens a sequence), or if it contains a colon followed by
734
+ U+0020 or the line end (`key: a .:` is a mapping whose key is `a .`). The second test is exact
735
+ rather than conservative — **a plain scalar can never contain a colon in that position** — and
736
+ it runs here, after the comment has been removed, so that `key: prose # note: here` is not
737
+ declined for a colon that is not in the scalar at all.
738
+
739
+ > **`:` and `#` are opaque one-code-point skipped units inside a plain scalar.** The scalar's
740
+ > range is split at each of them, exactly as a character reference splits an HTML text node.
741
+
742
+ Without the split, `yaml` mode corrupts documents, and the two witnesses are ordinary:
743
+
744
+ | Source | Locale | Without the split | Reparsed as |
745
+ | ---------- | ------- | ----------------- | -------------------------------------------------- |
746
+ | `k: a:--b` | `de-DE` | `k: a: – b` | **a parse error** — `: ` is now a mapping indicator |
747
+ | `k: a--#b` | `de-DE` | `k: a – #b` | `k` is `a –` — the rest became a comment |
748
+
749
+ Both come from the same source: `dashes` emits U+0020 in every `-spaced` locale, and **U+0020 is
750
+ the only thing that makes either character structural** — a `:` is a mapping indicator only when
751
+ a U+0020 follows it, a `#` a comment introducer only when a U+0020 precedes it. With the split,
752
+ `:` and `#` lie outside every span, so the dash token sits at a span **extremity**, and §3.4's
753
+ edge-growth rule discards a replacement there that is longer than what it replaced. The `-spaced`
754
+ form is 2 → 3. It is discarded, in `de-DE`, `fr` and `ru` alike.
755
+
756
+ **The em-tight form is 2 → 1, and it is applied** — `k: a:--b` becomes `k: a:—b` in `en-US`,
757
+ which reparses as the same one plain scalar it was. That is not a hole in the argument, it is the
758
+ argument: a contraction cannot emit U+0020, and U+0020 is the whole of what the split protects.
759
+ The length test of §3.4 separates exactly the two, with no knowledge of YAML — which is what
760
+ makes it bind rules not yet written, while a per-rule observation would not.
761
+
762
+ The accepted cost is that a conversion whose replacement would **grow** against a colon or a hash
763
+ is missed: `fr` inserts no narrow no-break space before a colon inside a YAML scalar, and
764
+ `one--two` takes its spaced dash only where the split leaves it interior to a span. That is the
765
+ same trade §7.3 already made for `mot<em>!</em>`, and in the same direction: a miss is visible to
766
+ the author and fixable in the source; a corrupted document is neither.
767
+
768
+ #### 3.8.7 What this was tested against
769
+
770
+ Three bodies of evidence, and the order matters: the first two were run against the **keyless
771
+ draft** and are reported here for what they failed to catch, which is as much a part of this
772
+ section as what they established.
773
+
774
+ **A real corpus, span placement only.** 70 YAML files — GitHub Actions workflows,
775
+ `docker-compose`, Dependabot, issue-form templates, and two OpenAPI documents including the one
776
+ the request came from. The scan's spans were checked against the `yaml` package's own scalar
777
+ ranges: **3135 spans, none overlapping a key, none outside a string value scalar, none whose text
778
+ was not a literal substring of that scalar's decoded value.** That establishes the **geometry** of
779
+ the scan and nothing else. It never ran the pipeline, so every span it approved could still have
780
+ been a shell script — and 37 of them were. Read alone it is exactly the kind of measurement that
781
+ looks like proof and is not.
782
+
783
+ **A bounded exhaustive sweep**, in the form `pipeline-idempotency.md` §6 requires of the rules.
784
+ Payloads of length 1–4 over the alphabet `a`, U+0020, `-`, **`---`**, `:`, `#`, `'`, `"`, `\` and
785
+ `.` in eleven document templates — all three scalar forms, sequence entries, compact nesting,
786
+ nested mappings, an interior block-scalar line, and a quoted scalar left open across a line — each
787
+ transformed in eight locales, reparsed, and compared against the original's **structure and
788
+ non-string leaves**. **977 680 cases, 99 274 of them changed by the transform: no parse failure,
789
+ no structural change, no idempotency failure.**
790
+
791
+ **The sweep is discriminating, and that was checked rather than assumed.** With §3.4's character
792
+ clause reverted, the same run reports **480 structural changes** — `k: a:---a` becoming
793
+ `k: a: – a`, a string turning into a nested mapping — in every locale that emits a spaced dash. A
794
+ measurement that passes both with and without the thing it is supposed to protect proves nothing,
795
+ and two earlier versions of this sweep were in exactly that position.
796
+
797
+ **Three things it could not find**, each worth naming because each cost something.
798
+
799
+ - It did not find the workflow damage, because a shell script inside a block scalar is a string
800
+ before the transform and a string after it. **Structural equality is not integrity**, and a
801
+ sweep shaped like this one cannot be the only guard on a data format.
802
+ - Its first alphabet omitted U+0022, so it could not reach a quoted scalar's own delimiters. The
803
+ multi-line-quoted template and the two characters it needs were added afterwards, and §5 now
804
+ requires them.
805
+ - Its alphabet then held `-` but not `---`, and at a payload length of four no combination of
806
+ single dashes reaches a three-dash run flanked by content. That is why the `r = d` hole in §3.4
807
+ survived two full runs of this sweep at 649 440 cases each, reported clean both times. **`---`
808
+ is in the alphabet as one token for that reason**, and §5's sweep obligation now names it.
809
+
810
+ **A real corpus, end to end.** The scan plus the full pipeline over the same 70 files in eight
811
+ locales, comparing the result against the original byte for byte. This is the measurement that
812
+ produced §3.8.2 and the only one of the three that can see a value damaged without being
813
+ destroyed:
814
+
815
+ | `keys` | Result |
816
+ | ------ | ------- |
817
+ | none — every string scalar processable (the keyless draft) | **37 of 106 spans changed in this repository's own workflows**, including `if !` → `if!` and `"$SHA"` → `“$SHA”` inside a `run:` block |
818
+ | every key that occurs anywhere in the corpus, 575 of them | 560 file/locale cases: no parse failure, no structural change, no idempotency failure — the scan's own safety envelope, at its most exposed |
819
+ | `description`, `summary`, `title` — what a caller would actually pass | **no workflow file changes at all**; 8 files change, and every one of them is prose-bearing: the two OpenAPI documents, four issue-form templates, and a package manifest |
820
+
821
+ The third row is the mode working as specified, and the first is why the option is not optional.
822
+
823
+ ---
824
+
825
+ ## 4. The round-trip guarantee
826
+
827
+ > **`transform` in `html`, `markdown` and `yaml` mode returns the input source with a set of
828
+ > disjoint substring replacements applied at recorded offsets. No other byte of the input
829
+ > changes, ever.**
830
+
831
+ This is stronger than "byte-identical when no changes are needed" (PLAN.md §3.2) and it implies
832
+ it. It is stated as a prohibition because that is the only form five parsers cannot drift on:
833
+
834
+ - **[P] The document is never serialised.** The parser is used to locate spans and is then
835
+ discarded. An implementation that reconstructs output from a DOM is non-conforming even if
836
+ its output happens to match.
837
+ - **[P] Attribute quoting is preserved** — `class=foo`, `class='foo'`, `class="foo"` all
838
+ survive as written. Attributes are skipped _and_ never re-emitted.
839
+ - **[P] Self-closing and void-element forms are preserved** — `<br>`, `<br/>`, `<br />`.
840
+ - **[P] Character-reference spelling is preserved** — `&amp;` does not become `&#38;`, `&nbsp;`
841
+ does not become U+00A0, and a bare `&` that a parser would repair stays a bare `&`.
842
+ - **[P] Tag name case, attribute order, duplicate attributes, and whitespace inside tags are
843
+ preserved.**
844
+ - **[P] The prologue, doctype, comments, and trailing whitespace are preserved.**
845
+ - **[P] Markdown is never reformatted** — list markers, emphasis delimiters, heading style,
846
+ table alignment, hard-break spaces and line endings are all outside every span.
847
+ - **[P] Malformed input is not repaired.** Unclosed tags, stray `<`, mismatched nesting: the
848
+ parser's error recovery affects only where spans are found, never what is emitted.
849
+
850
+ - **[P] YAML quoting style, indentation, anchors, tag handles and chomping indicators are
851
+ preserved** — `yaml` mode uses no parser and no emitter, so none of them is ever decoded and
852
+ rewritten (§3.8.1). A block scalar's trailing line terminators lie outside every span and are
853
+ returned exactly as written, whichever chomping indicator governs them.
854
+
855
+ The cost is real and is accepted: a runtime whose parser cannot report source offsets cannot
856
+ implement `html` mode conformantly. That constraint belongs to the implementation, not to this
857
+ contract. It does not bind `yaml` mode, which has no parser to ask — §3.8.1 records why asking
858
+ one was not an option in the first place.
859
+
860
+ ---
861
+
862
+ ## 5. Idempotency under composition
863
+
864
+ `pipeline-idempotency.md` proves `T(T(x)) = T(x)` for a single text run, on the composition
865
+ obligation **CO**: no rule creates work for an earlier-ordered rule. Modes do not weaken CO, and
866
+ they do not carry it over for free either. Write `M` for the whole mode transform.
867
+
868
+ **`M(M(x)) = M(x)` requires three things**, of which only the first is already proved:
869
+
870
+ 1. **The pipeline is idempotent on the concatenated array.** This is exactly
871
+ `pipeline-idempotency.md`, applied to an array that happens to contain markers. The markers
872
+ are inert — no rule matches them, no rule edits them, and they are in none of the classes any
873
+ guard tests except `OPENISH`/`CLOSEISH` for `quotes`, where they behave like a fixed piece of
874
+ punctuation that no rule can change. So every CO discharge in every rule document holds
875
+ verbatim.
876
+
877
+ 2. **The span partition is stable.** `M` must produce the same sequence of spans for `M(x)` as
878
+ for `x`. This is a **new obligation, invisible to every per-rule argument**, and it is the
879
+ one a mode adapter can break on its own: if applying edits changed how the document parses,
880
+ the second run would see different spans, the concatenation would differ, and the rules would
881
+ legitimately reach different conclusions.
882
+
883
+ An earlier revision discharged this by claiming that **every code point the rules emit is
884
+ syntactically inert in `html` and `markdown`**. That claim is **false**, and the
885
+ counterexample is the most ordinary character in the set. **U+0020 is not a markup character
886
+ and is not inert**: beside `*` or `_` it decides whether a delimiter run is left- or
887
+ right-flanking, and at the start of a line it is a line prefix, a list continuation or
888
+ indented code. `dashes` emits U+0020 in every `-spaced` locale. The claim was wrong when it
889
+ was written, and it is what hid the defect now repaired in §3.4.
890
+
891
+ The honest discharge has four parts, and the third is **positional rather than alphabetic** —
892
+ the same move as CO-S in `pipeline-idempotency.md` §5.1a, applied to _where_ a character may
893
+ be placed rather than to _which_ characters exist:
894
+
895
+ - **§4 forbids reserialisation**, so markup bytes are untouched by construction. Nothing a
896
+ rule does can move, requote or respell a tag, an entity or a fence.
897
+ - **No rule emits a markup character.** Of the code points the rules can emit —
898
+ U+2018/U+2019/U+201A–U+201F, U+00AB/U+00BB, U+2039/U+203A, U+2013/U+2014, U+2026, U+00A0,
899
+ U+202F, U+2011, U+00D7, U+00B1, U+00A9, U+00AE, U+2122, U+002E, U+0020 — none is `<`, `&`,
900
+ `>`, `*`, `_`, `` ` ``, `[`, `]`, `(`, `)`, `#`, `|` or a line terminator. This part of the
901
+ old claim survives and covers every emitted character **except U+0020**.
902
+ - **U+0020 is made safe by position, not by inertness.** The edge-growth rule (§3.4) means a
903
+ rule can introduce a code point only at a position **strictly interior** to a span. An
904
+ interior position has span content on both sides, so an emitted U+0020 can never become
905
+ adjacent to a delimiter — delimiters are structural tokens, outside every span — and can
906
+ never begin a line, because a line terminator is always a −2 marker and a line start is
907
+ therefore a span edge or outside spans entirely. **The characters that decide Markdown
908
+ structure and the positions a rule may write to are disjoint sets.** That is what makes the
909
+ partition stable; no property of U+0020 itself is relied upon.
910
+ - **Deletion** is only ever of a U+0020, and structural whitespace borders a `BREAK` — a real
911
+ one in `text` mode, a −2 marker in the others — where `spaces` refuses to act
912
+ (`spaces.md` §3.2 step 4). A deletion therefore cannot turn `- item` into `-item`, collapse
913
+ an indented code block, or destroy a hard break.
914
+
915
+ > **Normative, and binding on future rules:** a rule may emit a code point that is
916
+ > syntactically significant in a mode **only** where the edge-growth rule confines it to a
917
+ > span interior. A rule that needs to place such a character at a span edge must specify its
918
+ > own span-stability argument here first, and must not assume a character is inert merely
919
+ > because it is not markup. U+0020 is the standing counterexample.
920
+
921
+ **In `yaml` the same four parts hold, the third needs one addition, and the obligation needs
922
+ a fifth part this mode is the first to require.**
923
+
924
+ The addition to the third: YAML's structural characters are a larger set than Markdown's, but
925
+ only two of them are made structural by an adjacent U+0020 — `:` becomes a mapping indicator
926
+ when a U+0020 follows it, and `#` becomes a comment introducer when a U+0020 precedes it — and
927
+ those are exactly the two that §3.8.6 lifts out of every plain scalar as opaque units. They are
928
+ therefore outside every span, an emitted U+0020 can never become adjacent to one, and the dash
929
+ token that would have emitted it sits at a span extremity, where §3.4's length test discards
930
+ the 2 → 3 replacement that emits U+0020 and admits the 2 → 1 replacement that cannot. In a quoted scalar both are neutralised by the
931
+ quoting; in a block scalar both are literal content. Of the remaining emitted code points, only
932
+ U+002E is structural in YAML **syntax** at all, and only as `...` at the start of a line — a
933
+ position no span contains, since a plain scalar's span begins after its key and its colon, and
934
+ a block scalar's begins after an indentation of at least one U+0020. Deletion is unchanged:
935
+ `spaces` may delete only a U+0020, and a block scalar's indentation borders a −2 marker.
936
+
937
+ > **What that argument does not cover, and what §7.11 accepts.** A plain scalar's **type** is
938
+ > resolved from its whole text, not from any one character, so an edit anywhere inside one can
939
+ > change a value's tag without touching anything syntactic. Measured: `description: 1 .`
940
+ > becomes `description: 1.` in every locale, because `spaces` strips the space before a
941
+ > sentence-final dot — and the value goes from the string `"1 ."` to the number `1`. No
942
+ > enumeration of emitted code points can close this, because the mechanism is not a character
943
+ > but a whole-token resolution; the span-stability obligation of item 2 is unaffected, since
944
+ > the predicate that selects the span is the key. This is a cost of processing plain scalars at
945
+ > all, it applies only to a value a caller has explicitly listed as prose, and it is accepted
946
+ > rather than closed.
947
+
948
+ > **The fifth part, normative and binding on future modes: the predicate that decides whether
949
+ > a span exists must not read anything a rule can change.** In `html` and `markdown` the
950
+ > question never arises, because span selection reads only the document's structure. `yaml`'s
951
+ > first draft decided a plain scalar's processability from **its own text** — it had to contain
952
+ > a space and a letter — and `spaces` can delete that space. One line was enough to falsify the
953
+ > obligation: `ref: ${{ steps.pin.outputs.sha }}` yields one span, comes back as
954
+ > `ref: ${{steps.pin.outputs.sha}}`, and yields none on the next run. `M(M(x)) = M(x)` survived
955
+ > that only by accident — a lost span means fewer edits, and the first run's edits happened to
956
+ > be fixed points — but the obligation itself was false, so the conclusion rested on nothing.
957
+ > §3.8.2's `keys` option is the repair: the predicate reads the **key**, which lies outside
958
+ > every span and which no rule can reach, so the partition is a function of the source alone.
959
+
960
+ A content-dependent predicate is not merely risky here, it is unarguable: item 2's whole
961
+ method is to show that the characters rules may write and the positions they may write to are
962
+ disjoint from what decides structure. A predicate over span content puts the rules' own output
963
+ back inside that decision, and there is no version of the argument that survives it.
964
+
965
+ 3. **Redistribution is deterministic.** Guaranteed by §3.4: no edit spans a marker, and any edit
966
+ that would place code points at a span extremity which were not there before — an insertion at
967
+ a boundary, or a replacement longer than what it replaced — is discarded rather than assigned
968
+ to a side. The discard is a pure function of `(p, q, r, s₀, s₁)`, so it happens identically on
969
+ every run: a rule that was declined at an edge once is declined there always, and the output
970
+ contains nothing for the next run to reconsider.
971
+
972
+ Given 1–3, `M(M(x)) = M(x)`.
973
+
974
+ **The testing obligation of `pipeline-idempotency.md` §6 extends to modes**: the bounded
975
+ exhaustive sweep must be run in `html`, `markdown` and `yaml` mode over documents that place
976
+ span boundaries inside the swept string — at minimum a template like `A<em>B</em>C` with the
977
+ sweep alphabet distributed across `A`, `B` and `C`. A sweep that only ever produces one span
978
+ tests none of this document. In `yaml` the sweep alphabet must include `:`, `#`, `-` and **`---`
979
+ as one token**, since those are the characters whose adjacency the argument above turns on and a
980
+ three-dash run is not reachable from single dashes at a bounded payload length; and it must
981
+ include U+0022 and U+005C, without which the sweep cannot reach a quoted scalar's delimiters at
982
+ all. §3.8.7 records what each run did and did not establish — including that a sweep comparing
983
+ structure and types is blind to a string whose content is damaged, and that an alphabet without
984
+ `---` reported clean twice while §3.4's `r = d` hole was open.
985
+
986
+ ---
987
+
988
+ ## 6. Which §4 claims change
989
+
990
+ Per `pipeline-idempotency.md` §5.2, a **[P]** claim is a promise about `transform`. Several were
991
+ written for `text` mode and are mode-dependent. The notation extends: **[P: html, markdown]**
992
+ means the claim holds in those modes only.
993
+
994
+ - **"Anything inside a skipped region"**, asserted in every rule's §4, is
995
+ **[P: html, markdown, yaml]** and is now _defined_ rather than assumed — §3.6, §3.7 and §3.8
996
+ are what those bullets refer to. In `text` mode there are no skipped regions and the bullet is
997
+ vacuous. In `yaml` the definition runs the other way round (§3.8.2): the skipped region is
998
+ everything the scan did not claim.
999
+ - **`dashes` §4 "URLs, code spans, fenced code, HTML attributes"** — **[P: html, markdown]**.
1000
+ In `text` mode a URL is ordinary text. The rule is nevertheless safe there, but by its own
1001
+ guards (P1 declines the tight hyphens in `a-b`, the cluster guard declines `2026-08-15`), not
1002
+ because anything skipped it. The distinction matters to a reader deciding whether to run
1003
+ `text` mode over a document containing URLs. **`yaml` sits with `text` here, not with the
1004
+ other two**: a URL written inside a listed key's scalar is ordinary text and nothing skips it.
1005
+ A URL that is the whole value of a `url:` key is a different matter and is skipped, but by
1006
+ §3.8.2's `keys` option rather than by any rule of this bullet — the caller did not list that
1007
+ key.
1008
+ - **`symbols` §4 "`(c)` and `(r)` in a call or index position"**, and the whole `0x1F` family —
1009
+ **[P] in all modes**, since they rest on guards S1/M3/M4 rather than on skipping. §7.2 of that
1010
+ document already records that the `(tm)` exemption's exposure is `text`-mode-only, and this is
1011
+ why.
1012
+ - **`ellipsis` §4 "`../`, `./..`"** and **`spaces` §4's path protection** — **[P] in all modes**.
1013
+ They rest on the lone-dot condition (`spaces.md` §3.4), not on skipping.
1014
+ - **`quotes` §4 "Any unbalanced mark"** — **[P] in all modes**, and _strengthened_ by §3.3: a
1015
+ mark that would have been unbalanced within its own span may now pair across an element, which
1016
+ converts more, never less.
1017
+ - **`nbsp` §4 "The start or end of a text unit"** — **[P] in all modes, and strengthened.** A
1018
+ "text unit" is the concatenation, and §3.4 additionally refuses any insertion at a span
1019
+ boundary, so no span ever begins or ends with a character `nbsp` put there. The guarantee is
1020
+ now about element boundaries as well as document boundaries.
1021
+
1022
+ - **`dashes` §4, the `-spaced` forms** — a new **[R]** consequence in `html`/`markdown`, not a
1023
+ change to any existing bullet. A `-spaced` locale converts a dash only where the replacement
1024
+ fits inside a span without touching either edge, so `a<em>--</em>b` is left alone in `de-DE`
1025
+ while `en-US` converts it — only a contraction survives at an edge (§7.9). Nothing in `dashes.md` is false; the rule simply gets fewer chances
1026
+ to fire. §7.9 records it.
1027
+
1028
+ No rule's §4 becomes _false_ in a mode. Two become narrower ([P: html, markdown]); none becomes
1029
+ rule-local.
1030
+
1031
+ ---
1032
+
1033
+ ## 7. Open questions
1034
+
1035
+ 1. _(Settled — both skipped.)_ `svg` and `math` are in §3.6's skip list although neither is in
1036
+ PLAN.md §3.2's. The decisive argument was MathML: a quotation mark, a hyphen and a prime are
1037
+ **operators and identifiers** there, so substitution changes meaning rather than appearance.
1038
+ The accepted cost is `<svg><text>` and `<svg><title>`, which do hold prose and are now left
1039
+ untypeset. Direction matters here — widening a skip list is additive, narrowing one after
1040
+ release breaks documents — so the conservative side was taken deliberately rather than
1041
+ pending evidence.
1042
+ 2. **A skipped element between two halves of a word.** `un<code>x</code>believable` yields two
1043
+ spans that no rule joins, which is correct, but it also means `hyphen` and the `nbsp` literal
1044
+ lists silently fail on any word interrupted by markup. Correct and conservative, but worth a
1045
+ fixture so it is a decision rather than a surprise.
1046
+ 3. **An insertion at a span boundary is discarded, and the miss is permanent under this
1047
+ contract.** (Since the edge-growth rule of §3.4, this is one case of a general prohibition on
1048
+ writing a code point onto a span edge; the argument below is what established the principle,
1049
+ and it applies unchanged to the replacement case that generalised it.) `fr` `mot<em>!</em>` keeps no narrow no-break space. This is an argued
1050
+ limitation, not an oversight, and the argument is the mirror pair — neither side is
1051
+ universally safe:
1052
+
1053
+ | Input | Spans | Assign to left span | Assign to right span |
1054
+ | --------------- | ---------- | ------------------- | -------------------- |
1055
+ | `mot<em>!</em>` | `mot`, `!` | `mot⍹<em>!</em>` ✓ | `mot<em>⍹!</em>` ✗ |
1056
+ | `<em>mot</em>!` | `mot`, `!` | `<em>mot⍹</em>!` ✗ | `<em>mot</em>⍹!` ✓ |
1057
+
1058
+ Whichever rule is adopted, one of these two ordinary documents ends up with an element that
1059
+ begins or ends with a space that was never in it. That is invisible in a plain rendering and
1060
+ immediately visible the moment the element carries an underline, a border or a background —
1061
+ and the package must never put a character inside an element it did not come from. A miss is
1062
+ visible to the author and fixable in the source; corrupted markup is neither.
1063
+
1064
+ Note that only _insertion_ is affected. `<em>mot</em> !` already has a U+0020 inside the
1065
+ right-hand span, so N2 converts it in place and French spacing works normally; the loss is
1066
+ confined to documents where the space does not exist at all and the punctuation is wrapped.
1067
+
1068
+ **A real fix needs something this contract does not have:** the ability to represent
1069
+ "between two nodes" as a position, so that an inserted character could be emitted as a
1070
+ sibling of both elements rather than a child of either. That is a different document model —
1071
+ the adapter would have to be able to _add_ to the document rather than only replace inside
1072
+ spans, which §4's round-trip prohibition rules out by design. Any future attempt starts by
1073
+ reopening §4, not by revisiting this paragraph.
1074
+
1075
+ 4. _(Settled — soft breaks are `BREAK`, not boundaries.)_ A soft line break inside a paragraph
1076
+ is a **−2 line marker** (§3.2), classified as a member of `BREAK` for every rule. The
1077
+ alternative readings were both worse. Treating it as an ordinary **−1 inline boundary** — what
1078
+ an implementation gets by default, since a line ending is a structural token like any other —
1079
+ makes a mode diverge from `text` mode on the same characters: a `"` at the start of a wrapped
1080
+ line has a `BREAK` on its left in `text` (so it can only open) and an opaque content marker in
1081
+ `html`/`markdown` (so it could also close), and `"foo\n"bar` then pairs in one and not the
1082
+ other. A mode that converts _differently_ from `text` on identical prose is indefensible; a
1083
+ mode that converts _less_ is nearly as bad, and hard-wrapped prose is precisely the M4 corpus.
1084
+ Putting the line terminator physically **inside** the span was the other candidate and it
1085
+ fails on structure: a continuation line's block prefix — a blockquote `>`, a list item's
1086
+ indentation — must stay outside, which makes the span non-contiguous and breaks the
1087
+ source-offset model §4 depends on.
1088
+ The −2 marker gets the benefit of both: spans stay contiguous, and every rule's existing,
1089
+ fixture-covered `BREAK` behaviour applies unchanged. It also removes a parser dependency that
1090
+ would otherwise have been load-bearing — whether the two spaces of a Markdown hard break land
1091
+ inside a span or in a structural token is a per-parser decision, and with a −2 marker beside
1092
+ them `spaces` protects them either way.
1093
+ 5. **Adjacent text nodes.** Some parsers report `a<!-- c -->b` or a long text run as two adjacent
1094
+ text nodes with no element between them. This document treats every gap between processable
1095
+ spans as a boundary, so such a split would suppress conversions that ought to happen.
1096
+ Implementations **should** coalesce adjacent processable spans separated by nothing in the
1097
+ source; whether that is required, and how it interacts with comments, is not settled here.
1098
+ 6. **`markdown` emphasis delimiters inside a span.** `*foo*` is markup, but a naive adapter that
1099
+ walks an `mdast` tree gets `foo` as a text node and the asterisks as structure — so they are
1100
+ outside every span, which is right. An adapter that instead processes raw source lines would
1101
+ have them inside a span and could let a rule edit next to them. The tree-walking approach is
1102
+ assumed throughout; it is not stated as a requirement because it names an implementation
1103
+ strategy, and it probably should be.
1104
+ 7. **No mode covers plain-text email or reStructuredText**, and the `rules` option cannot
1105
+ express "skip this region" for a caller with their own format. That is out of scope for v1
1106
+ (PLAN.md §4) and is recorded so the span model is not mistaken for an extension point. _(YAML
1107
+ was in this item until 1.3.0 and is now §3.8; the two formats named here are not, and the
1108
+ argument that moved YAML does not carry over — it turned on a measured corpus and a
1109
+ corruption report, neither of which exists for these.)_
1110
+
1111
+ 8. **`dialect` and `POLYTYPO_MALFORMED_INPUT` are new public surface and need operator
1112
+ sign-off.** Both are specified — I did not leave them open — but both change the contract.
1113
+
1114
+ The first is `dialect` (§3.7.1): a required option on `markdown` mode rather than a fourth
1115
+ mode id, because mode ids are public API and PLAN.md §3.2 names exactly three. Two things
1116
+ there are mine to flag rather than to decide: whether the operator prefers the option or a
1117
+ `markdown-mdx` mode id after all; and which code the missing or invalid value raises. I
1118
+ specified a throw with no default, on the fail-fast reasoning that already governs `locale`,
1119
+ but ARCHITECTURE.md §4.6 has no code for it — `POLYTYPO_INVALID_MODE` is the closest fit and
1120
+ is arguably a lie, since the mode is valid and only its dialect is not. A
1121
+ `POLYTYPO_INVALID_DIALECT` code would be honest, and is a contract addition.
1122
+
1123
+ The second is `POLYTYPO_MALFORMED_INPUT` (§3.7.2). It adds a code to the taxonomy **and**
1124
+ makes `transform` non-total on its input for the first time. I judged that right on
1125
+ fail-fast grounds and because the only reachable case is a genuinely broken MDX file, but a
1126
+ caller running polytypo over a large corpus may reasonably want a document-level failure not
1127
+ to abort a batch. If so, the answer is a wrapper in _their_ code, not an option here — this
1128
+ spec has no error-suppression switch and should not acquire one.
1129
+
1130
+ 9. **The edge-growth rule costs conversions that `text` mode makes, and only a contraction
1131
+ survives at a span edge.** That sentence is the whole rule, and it follows from §3.4's length
1132
+ test rather than sitting beside it: an edit at an extremity is applied iff `r ≤ d`. So a
1133
+ `dashes` promotion at a span edge converts **only where the locale's parenthetical form is
1134
+ shorter than the input that triggered it**.
1135
+
1136
+ Measured across all nine locales, on `a<em>--</em>b`:
1137
+
1138
+ | Locale | `dash.parenthetical` | At a span edge | Interior (`a<em>x--y</em>b`) |
1139
+ | --- | --- | --- | --- |
1140
+ | `en-US` | `em-tight` | **converts** — `a<em>—</em>b`, since 2 → 1 is a contraction | converts |
1141
+ | `de-CH`, `de-DE`, `en-GB`, `fi`, `sv` | `en-spaced` | unchanged — 2 → 3 grows both edges | converts, `a<em>x – y</em>b` |
1142
+ | `fr`, `ru` | `em-spaced` | unchanged — same reason | converts |
1143
+ | `el` | `none` | unchanged — emits nothing anywhere | unchanged |
1144
+
1145
+ `en-US` is the only locale whose parenthetical form is shorter than the `--` that triggers it,
1146
+ and it is therefore the only one that converts at an edge. The asymmetry is real — the same
1147
+ document is typeset differently in `de-DE` and `en-US` for a reason that has nothing to do
1148
+ with German or American typography — but it applies to a **narrower class** than an earlier
1149
+ revision of this item claimed, and it now has a mechanical explanation rather than a
1150
+ coincidental one.
1151
+
1152
+ That earlier revision was also **factually wrong** in a way worth recording, because the error
1153
+ outlived the thing that caused it. It said `a<em>–</em>b` "keeps its en dash unspaced in
1154
+ `de-DE` … while `en-US` (`em-tight`) converts it happily". Neither half is true any more:
1155
+ `dashes` §3.2 step 2a declines **every** token containing an authored U+2013 or U+2014, so
1156
+ nothing converts that input in any locale. The claim was correct when written and was
1157
+ invalidated by a change in another document; the carrier died and the sentence did not notice.
1158
+ §3.4's worked table had the identical fault and was rebuilt on `--` for the same reason.
1159
+
1160
+ 10. **Whether a −2 marker should also end a `quotes` pairing scope** is open. Today a quotation
1161
+ opened in one paragraph and never closed can still pair with a mark in the _next_ paragraph,
1162
+ because the stack is not reset at a `BREAK` — which is exactly the behaviour `text` mode has,
1163
+ so the two agree. It may nonetheless be wrong in both, and `quotes.md` §7.5 already records
1164
+ the multi-paragraph question. If that is ever resolved, it must be resolved for `text` and
1165
+ the modes together, not here.
1166
+
1167
+ 11. **`yaml` mode's accepted misses, and one accepted cost.** Every entry but the last is a
1168
+ construct the scan declines to claim, so prose inside it is returned untouched; each could be
1169
+ admitted later without breaking a document that relies on today's behaviour, and none could
1170
+ be withdrawn later without breaking one. They are recorded together because the direction is
1171
+ the argument — and the last entry is here precisely because it is the one that does not share
1172
+ it:
1173
+
1174
+ - **a bare sequence item** — `- Some prose here` yields no spans, because §3.8.4 step 5
1175
+ requires a processable scalar to be the value of a key, and a bare item has none for
1176
+ `keys` to match. `- key: value` is unaffected;
1177
+ - **a key containing `"`, `'`, `{`, `[`, `&`, `*`, `!` or `#` anywhere** — `"description"`,
1178
+ but also `a!b` — declined at step 5 before the key is compared, so a document that quotes
1179
+ its keys gets nothing from this mode;
1180
+ - **a plain scalar whose type is resolved from its whole text** — a listed key's value can
1181
+ change tag without any syntactic character being touched (§5 item 2's note). `1 .` becomes
1182
+ `1.` and stops being a string. This one is not a miss but an accepted cost, and it is the
1183
+ only entry here that is not purely in the widening direction;
1184
+ - **a multi-line plain scalar**, and a quoted scalar or flow collection that does not close
1185
+ on its opening line: all three are consumed by step 7 and yield no spans;
1186
+ - **a flow collection** — `tags: [one, two]` and `{a: b}` — even on one line;
1187
+ - **a quoted scalar containing an escape** — `"say \"hi\""` — where the source spells the
1188
+ content with more characters than it has;
1189
+ - **a block scalar whose indentation is ambiguous**, per §3.8.5's three bail conditions.
1190
+
1191
+ 12. **`keys` matches a bare name at any depth, and nothing narrower.** `description` is
1192
+ processable wherever it occurs, which is what makes the option portable — no paths, no
1193
+ globs, no schema. The cost is that a caller who wants `components.schemas.*.description`
1194
+ but not `info.description` cannot say so, and a caller whose document has a machine-read
1195
+ `description` somewhere in it must choose between typesetting that one too and typesetting
1196
+ none of them. A path syntax is the obvious extension and is deliberately not specified here:
1197
+ it is a small language, five runtimes would have to agree on it exactly, and no measured
1198
+ document has yet needed it. If one does, this is where it reopens, and the extension is
1199
+ additive — a caller passing bare names keeps today's behaviour.
1200
+
1201
+ ## 8. Fixture coverage strategy (non-normative)
1202
+
1203
+ This section records why `spec/fixtures/` does not carry the full every-locale × four-modes
1204
+ Cartesian product, and what a smaller set must still prove instead. It has **no normative
1205
+ force** — it does not change §§1–7 — and if it drifts out of sync with the actual fixtures, that
1206
+ is a defect in this section, not license to distrust the fixtures.
1207
+
1208
+ **Why the full product is not required.** §3.2–§3.5 establish that a mode adapter's only job is
1209
+ to identify which spans of the source are processable text and which are structural markup:
1210
+ locating skip-list boundaries, computing marker gaps, and reassembling by offset. Once a span is
1211
+ handed off, `runOverSpans` (`src/engine/span-runner.ts`, used by `html` and `markdown`) calls
1212
+ exactly the same per-rule `apply()` functions, against the same locale data, that `text` mode's
1213
+ `runRules` (`src/engine/text-pipeline.ts`) calls — there is one implementation per rule id,
1214
+ shared by every mode, not one per mode. A locale's quote glyphs, dash conventions, or `nbsp`
1215
+ targets are therefore already fully exercised by `text`-mode fixtures; an `html`/`markdown`
1216
+ fixture in the same locale mostly re-tests that same rule logic through an extra layer of span
1217
+ bookkeeping, not something the locale itself changes about it.
1218
+
1219
+ What *does* vary by mode is the span-selection and reassembly machinery, and that machinery is
1220
+ locale-agnostic: the HTML skip list and the CommonMark/MDX skip lists never consult `LocaleData`.
1221
+ A representative sample proves the machinery correct; the full product would mostly multiply
1222
+ proof of the same machinery by locale count, without adding coverage of anything the locale
1223
+ changes.
1224
+
1225
+ **What representative coverage requires instead**, and where it currently lives:
1226
+
1227
+ 1. **HTML span selection** — the skip list, character-reference handling, and round-trip
1228
+ guarantee, exercised directly (`tests/modes/html.test.ts`) and via fixtures.
1229
+ 2. **CommonMark span selection** — fenced/indented/inline code, autolinks, link destinations and
1230
+ titles, reference-link definitions, and the round-trip guarantee
1231
+ (`tests/modes/markdown.test.ts`).
1232
+ 3. **MDX span selection** — expression containers, JSX attributes and children, ESM export
1233
+ blocks, and the JSX-vs-skipped-element case distinction (`tests/modes/markdown.test.ts`).
1234
+ 3a. **YAML span selection** — the three scalar forms, the `keys` option (a listed key, an
1235
+ unlisted one, a quoted one, the same name at two depths), block-scalar indentation and
1236
+ chomping, the `:`/`#` split of §3.8.6, step 7's consumption of an inline value's continuation
1237
+ lines, and every bail of §3.8.4. `yaml` carries a heavier fixture burden than the other two
1238
+ for a reason §8's opening argument does not cover: its span selection is **specified rather
1239
+ than delegated** (§3.8.1), so a fixture is the only thing standing between five hand-written
1240
+ scanners and five different answers. The bounded sweep of §3.8.7 is part of this item, not an
1241
+ extra — and so is the end-to-end corpus run, for the reason §3.8.7 gives: a sweep that
1242
+ compares structure cannot see a string being damaged without being destroyed.
1243
+ 4. **At least two materially different locale outputs per mode/dialect**, so a fixture is not
1244
+ merely "the same English output with a different `locale` field": `spec/fixtures/fr.json` and
1245
+ `fr-CA.json` carry `html`/`markdown`/`mdx` cases whose guillemets-plus-U+00A0 output is
1246
+ structurally different from `en-US`'s curly quotes, not just a different glyph in the same
1247
+ shape. `yaml` carries `en-US`, `de-DE` and `fr`, and for that mode **one of them must be a
1248
+ `-spaced` locale**: §3.8.6's `:`/`#` split only discriminates where the replacement grows, so
1249
+ an em-tight-only fixture set passes with the split removed.
1250
+ 5. **Non-ASCII text and code-point/offset boundaries** — an astral-character (surrogate-pair)
1251
+ preservation case in HTML mode, and genuinely accented non-ASCII prose exercised through the
1252
+ French MDX and CommonMark fixtures and round-trip tests.
1253
+ 6. **Byte-identical skipped regions** — the round-trip guarantee ("returns a document/article
1254
+ that needs no changes byte for byte") asserted directly for HTML, CommonMark, MDX and YAML.
1255
+ In `yaml` the guarantee is also what proves §3.8.2: a document the scan does not understand
1256
+ must come back unchanged, so an unrecognised construct is a fixture, not a hope.
1257
+ 7. **Every locale-specific transformation independently covered through `text`-mode fixtures** —
1258
+ the conformance suite's own per-locale, per-rule coverage, not this file.
1259
+
1260
+ **Enforcement**, split across the two places that actually check each item — neither file alone
1261
+ covers all eight:
1262
+
1263
+ - `tests/conformance/mode-fixture-strategy.test.ts` reads `spec/fixtures/*.json` directly and
1264
+ protects item 4 (at least two distinct locales carry `html`, `markdown`/`commonmark`,
1265
+ `markdown`/`mdx` and `yaml` fixtures, and at least one `yaml` locale is `-spaced`), the
1266
+ code-point half of item 5 (at least one fixture per mode contains a non-ASCII code point,
1267
+ `yaml` included), and item 7 (every locale has a `text`-mode fixture for
1268
+ *every* canonical rule id in `spec/rules/order.json`, checked per locale/rule id pair, not
1269
+ merely "the locale has a fixture for some rule"). It does not inspect span selection or
1270
+ round-trip behaviour at all.
1271
+ - **In each runtime repository**, `tests/modes/html.test.ts`, `tests/modes/markdown.test.ts` and
1272
+ `tests/modes/yaml.test.ts` protect items 1–3a (HTML, CommonMark, MDX and YAML span selection
1273
+ respectively), the parser-boundary half of item 5 (an astral-character/surrogate-pair
1274
+ preservation case in HTML mode), and item 6 (the round-trip guarantee — "returns a
1275
+ document/article that needs no changes byte for byte" — asserted directly for HTML, CommonMark,
1276
+ MDX and YAML). Those files are not in this repository, which has no engine to run them through,
1277
+ so this bullet is an obligation on every runtime: one that ships `yaml` without them is not
1278
+ conformant for the mode however green its fixtures are.
1279
+
1280
+ If a future edit narrows fixture coverage, or removes a span-selection or round-trip assertion,
1281
+ the corresponding test above fails before this section's claim goes silently stale.