polytypo 1.2.0 → 1.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +648 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,605 @@
1
+ # Pipeline idempotency
2
+
3
+ **Not a rule.** No entry in `spec/rules/order.json`, no locale data, no edits. This document
4
+ states the invariant that `transform` as a whole must satisfy, proves that per-rule
5
+ idempotency does not imply it, and defines the obligation each rule must discharge so that it
6
+ does.
7
+ **Spec version:** 1.2.0 (0.1.0 for everything except S-b's word-start clause, added in 1.2.0).
8
+
9
+ ---
10
+
11
+ ## 1. Purpose
12
+
13
+ PLAN.md §3.4 makes `transform(transform(x)) == transform(x)` a hard invariant **on the public
14
+ function**, not on individual rules. Every rule document carries a §5 arguing that _that rule_
15
+ is a fixed point on its own output. Those arguments are necessary and they are not sufficient:
16
+ the composition of eight individually idempotent functions is not in general idempotent, and
17
+ in this pipeline it demonstrably was not. Two defect families were found by exhaustive search
18
+ over short strings, both of them invisible to every per-rule argument because no per-rule
19
+ argument is allowed to mention another rule.
20
+
21
+ This document supplies the missing layer. It is short, and the obligation it imposes is
22
+ mechanical enough to be checked by reading a rule's "Must not touch" section against a table.
23
+
24
+ ---
25
+
26
+ ## 2. The invariant, and why per-rule idempotency does not give it
27
+
28
+ Write the pipeline as `T = R₈ ∘ R₇ ∘ … ∘ R₂ₐ ∘ R₂ ∘ R₁`, the rules in `order.json` order:
29
+
30
+ | # | rule | order |
31
+ | --- | ------------ | ----- |
32
+ | R₁ | `spaces` | 10 |
33
+ | R₂ | `ellipsis` | 20 |
34
+ | R₂ₐ | `ranges` | 25 |
35
+ | R₃ | `dashes` | 30 |
36
+ | R₄ | `hyphen` | 35 |
37
+ | R₅ | `quotes` | 40 |
38
+ | R₆ | `apostrophe` | 50 |
39
+ | R₇ | `symbols` | 60 |
40
+ | R₈ | `nbsp` | 70 |
41
+
42
+ (A disabled rule is removed from the sequence and never reorders the rest, so every statement
43
+ below holds for any subset, in the same relative order.)
44
+
45
+ **Why `ranges` is `R₂ₐ` and not `R₃`.** It was split out of `dashes` in spec 0.5.0, after the
46
+ subscripts in this document had been cited by name from eight other rule documents. Renumbering
47
+ `R₃ … R₈` to make the sequence contiguous would rewrite every one of those references for no
48
+ gain in meaning, so the rule takes the letter and the composition `T = R₈ ∘ R₇ ∘ … ∘ R₂ₐ ∘ … ∘ R₁`
49
+ reads with it in place at order 25. `ranges` is also the one rule that is **off by default**, so
50
+ for most callers the sequence is literally the eight numbered ones; every proof below is stated
51
+ over "the rules that run", which is the subset the caller's `rules` option selected.
52
+
53
+ For each rule define the predicate
54
+
55
+ > **`Iᵢ(y)` ⟺ `Rᵢ(y) = y`** — "rule `i` is a no-op on `y`".
56
+
57
+ **Lemma.** If `y = T(x)` satisfies `I₁ ∧ I₂ ∧ I₂ₐ ∧ … ∧ I₈`, then `T(y) = y`.
58
+ _Proof._ `R₁(y) = y` by `I₁`; then `R₂(R₁(y)) = R₂(y) = y` by `I₂`; then `R₂ₐ(y) = y` by `I₂ₐ`;
59
+ and so on through `R₈`. ∎
60
+
61
+ So the whole problem reduces to: **make the final output a fixed point of every rule, not just
62
+ of the last one that touched it.**
63
+
64
+ Per-rule idempotency gives `Iᵢ` immediately after `Rᵢ` runs. What it does not give is that
65
+ `Iᵢ` still holds at the _end_ of the pipeline. A later rule may undo it. That is exactly the
66
+ gap, and it yields the obligation:
67
+
68
+ > ### The composition obligation (CO)
69
+ >
70
+ > **For every pair `i < j`: if `Iᵢ(y)` holds, then `Iᵢ(Rⱼ(y))` holds.**
71
+ >
72
+ > In words: **a rule must never create work for an earlier-ordered rule.**
73
+
74
+ With CO, an induction over `j` gives the Lemma's premise: after `Rⱼ` has run, `I₁ … Iⱼ` all
75
+ hold — `Iⱼ` by `Rⱼ`'s own idempotency, and `I₁ … Iⱼ₋₁` because `Rⱼ` preserved them. After the
76
+ last rule all of them hold, and `T(T(x)) = T(x)`. (`R₂ₐ` takes its place in the induction between
77
+ `R₂` and `R₃`; the letter is a numbering artefact of the spec 0.5.0 split, not a gap in the
78
+ sequence.)
79
+
80
+ Note what CO does **not** require: nothing about `j < i`. A rule may freely create work for a
81
+ _later_ rule, because the later rule has not run yet and will clean it up in the same pass.
82
+ That asymmetry is the whole content of "the rules run in a fixed order".
83
+
84
+ ### 2.1 Why not simply iterate to a fixed point
85
+
86
+ Because it would make the invariant true by construction and hide precisely the defects this
87
+ document exists to find. It would also weaken the per-rule guarantee the package sells (a rule
88
+ that is a fixed point in isolation is a testable, portable claim; "the loop converges
89
+ eventually" is not), and any difference in iteration bound between two runtimes — or any input
90
+ where one runtime converges in two passes and another in three — becomes a conformance
91
+ divergence rather than a bug. **The rules must compose correctly.** Iteration is forbidden.
92
+
93
+ ---
94
+
95
+ ## 3. The output invariants
96
+
97
+ To discharge CO a rule author needs to know what `Iᵢ` actually says for each earlier rule, in
98
+ terms concrete enough to check. These are the four that constrain anything.
99
+
100
+ ### I₁ — `spaces`
101
+
102
+ `spaces` deletes U+0020 in three situations and collapses runs. `I₁(y)` therefore requires
103
+ that `y` contains **no U+0020 in any of these positions**:
104
+
105
+ - **S-a** a run of two or more U+0020 with a `CONTENT` code point on both sides;
106
+ - **S-b** a U+0020 whose right neighbour is in `STRIP-BEFORE` = { U+002C, U+002E, U+003B,
107
+ U+003A, U+0021, U+003F } and whose left neighbour is `CONTENT` — **and**, when that right
108
+ neighbour is U+002E, the maximal run of { U+002E, U+2026 } beginning there has length exactly
109
+ 1 and the code point after that dot is neither a `LETTER` nor an ASCII digit (`spaces.md` §3.4;
110
+ the second clause since spec 1.2.0) — **and** unless the emoticon guard's eye side fires, i.e. that right
111
+ neighbour is the eye of a recognised emoticon (`spaces.md` §3.6);
112
+ U+2026 is not a member of `STRIP-BEFORE`;
113
+ - **S-c** a U+0020 whose left neighbour is in { U+0028, U+005B, U+007B } — unless the
114
+ empty-bracket guard applies, and unless the emoticon guard's mouth side fires, i.e. that left
115
+ neighbour is the mouth of a recognised emoticon (`spaces.md` §3.6);
116
+ - **S-d** a U+0020 whose right neighbour is in { U+0029, U+005D, U+007D } — same
117
+ empty-bracket proviso. The emoticon guard does not apply here: it suppresses the
118
+ `STRIP-BEFORE` and `OPEN-BRACKET` clauses only, never this one.
119
+
120
+ See `spaces.md` §3.1–§3.3 for the exact definitions of `CONTENT` and the guard, and §3.6 for the
121
+ emoticon guard's two sides. Both emoticon provisos, and the lone-dot condition's word-start clause
122
+ added in spec 1.2.0, narrow the forbidden set, so every discharge
123
+ already written against S-b, S-c or S-d stays valid without re-derivation — a rule that emits no
124
+ U+0020 in the wider set emits none in the narrower one.
125
+
126
+ **Who can violate it.** Only a rule that emits U+0020. In the current pipeline that is
127
+ `dashes` alone (`-spaced` forms). `nbsp` emits U+00A0 and U+202F, which are `CONTENT` to
128
+ `spaces` and are never touched. `quotes` (0.3.0) deletes a code point inside a quote pair when
129
+ `innerSpace = "none"`, but never a U+0020 that survives the deletion — the maximal-run deletion
130
+ and the landing guard together ensure the two surviving neighbours are always `CONTENT`
131
+ (`quotes.md` 3.7, 5). Every other rule replaces code points one-for-one or shortens a run, and
132
+ none of them can create a U+0020 or bring two apart-standing ones together.
133
+
134
+ ### I₂ — `ellipsis`
135
+
136
+ `I₂(y)` requires no maximal run over { U+002E, U+2026 } that is either three or more code
137
+ points long, or of length ≥ 2 containing a U+2026; plus the locale's terminal form.
138
+
139
+ **Who can violate it.** Only a rule that emits U+002E or U+2026, or that deletes a code point
140
+ standing between two such runs. `quotes` (0.3.0) deletes code points, but only a maximal
141
+ `INLINE-SPACE` run landing on `ALNUM ∪ QUOTEMARK` — never a code point standing between two dot
142
+ runs, since a quote glyph is never itself U+002E/U+2026 and the deletion's landing class contains
143
+ neither (`quotes.md` 5, composition obligation against `I₂`). `I₂` is otherwise unconditionally
144
+ preserved and needs no attention from rule authors, but it is listed so that a future rule
145
+ emitting a full stop knows it has an obligation.
146
+
147
+ ### I₂ₐ — `ranges`
148
+
149
+ `I₂ₐ(y)` requires that no **range candidate** in `y` is one `ranges` would edit — `ranges.md`
150
+ §3.2 and §3.2a for candidacy, §3.2's G1-G5 for the guards, §3.3 for the replacement. A candidate
151
+ is a dash token whose two flanks are `DIGIT`, or are `DIGIT` once a `CLOSED-SYMBOL` matched on
152
+ the opposite member has been walked over.
153
+
154
+ **Who can violate it.** `dashes` (R₃), and only `dashes` — which is why the obligation is
155
+ discharged in that rule's own §5.3 rather than here. `ranges`' guards G1, G2 and G3 read
156
+ `before`/`after`, code points outside the token, and a `dashes` edit that turns a tight token
157
+ into a spaced one replaces a dash at exactly such a position with a U+0020. Two of `dashes`'
158
+ shared guards exist for this and no other reason: the **cluster guard** (`dashes.md` §3.2
159
+ step 7) makes the whole neighbourhood inert when it holds two dash runs, and **T1** (§3.2
160
+ step 8) declines the tight-to-spaced transition at the one-space distance the cluster guard
161
+ cannot see, because a cluster ends at a space.
162
+
163
+ **Spec 1.3.0 widened `I₂ₐ`'s domain, and T1 with it.** `ranges.md` §3.2a made a `CLOSED-SYMBOL`
164
+ written closed up to a digit run part of a range member, so T1's reach became transparent to one
165
+ such symbol at each of two positions per side. `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
166
+ `$1 - $1--a` are the witnesses, one per transparency position and side, each of which produced a
167
+ `transform` idempotency defect before the amendment in the locales whose `dash.parenthetical` is
168
+ spaced **and** whose `dash.range` is not `"none"`, and each of which now behaves exactly as its
169
+ all-digit analogue always did. The remedy is unconditional on `dash.range`, so its cost is wider
170
+ than the defect was — see `ranges.md` §3.2a. The cluster guard's
171
+ alphabet was deliberately **not** widened — see `dashes.md` §3.2 steps 7-8 for the comparison.
172
+
173
+ **Nothing ordered after `ranges` can violate `I₂ₐ`, and the argument is positional, not
174
+ alphabetic.** `hyphen` (R₄) emits U+2011, which *is* in `INERT-DASH` and therefore *is* in G2's
175
+ exclusion set — the alphabetic argument would be false. What holds instead is that `hyphen`
176
+ replaces a U+002D **in place**, inside a word, so any position where its output could satisfy G2
177
+ already held a `DASH` and already failed G2. `quotes` (R₅) and `apostrophe` (R₆) replace code
178
+ points one for one with quote glyphs, and `quotes`' one deletion removes an `INLINE-SPACE` run
179
+ whose two sides are a quote glyph and a `DELETE-LANDING` member — neither is a digit or a
180
+ `CLOSED-SYMBOL`, so no range member's `before`/`after` moves. `symbols` (R₇) emits only U+00A9,
181
+ U+00AE and U+2122 and deletes only a `(`…`)` span. `nbsp` (R₈) converts a space it found and
182
+ never inserts one where a guard reads (`nbsp.md` §3.3), and U+00A0/U+202F satisfy none of G1, G2
183
+ or G3.
184
+
185
+ ### I₃ — `dashes`
186
+
187
+ `I₃(y)` requires that no **dash token** in `y` is one `dashes` would edit — see `dashes.md`
188
+ §3.2 for the admissibility gate and §3.3/§3.4 for the branches. In practice a later rule
189
+ violates `I₃` when it changes the **spacing** around a U+002D/U+2013/U+2014, because spacing
190
+ is what `dashes` reads to decide whether a stroke is a parenthetical dash at all.
191
+
192
+ **Who can violate it.** Nobody, since spec 0.1.0 — and it is worth recording that this line
193
+ previously read _"`nbsp`, which inserts and converts space-like code points"_, which was true
194
+ and was the source of two defects. `dashes` now treats both U+00A0 and U+202F as making an
195
+ adjacent token inert (`dashes.md` §3.2 step 3), so `E(nbsp) = { U+00A0, U+202F }` is wholly
196
+ inert for it and **CO-S** discharges the pair structurally. `apostrophe` and `symbols`
197
+ replace code points one-for-one with characters in none of `dashes`' classes; `hyphen` emits
198
+ U+2011, which `dashes` treats identically to U+002D everywhere it matters. `quotes` (0.3.0)
199
+ replaces code points one-for-one with quote glyphs (also outside every `dashes` class) and, when
200
+ `innerSpace = "none"`, deletes a maximal `INLINE-SPACE` run whose two sides are a quote glyph and
201
+ a `DELETE-LANDING` member — neither in `DASH ∪ INERT-DASH`, so no dash run is ever adjacent to a
202
+ deleted run and no token's `lsp`/`rsp` changes (`quotes.md` 3.7, 5, composition obligation
203
+ against `I₃`).
204
+
205
+ ### I₄ — `hyphen`
206
+
207
+ `I₄(y)` requires no occurrence of a listed `hyphen` form whose hyphen is still U+002D. A later
208
+ rule violates it only by changing a form's word boundaries, i.e. by inserting or removing a
209
+ code point immediately beside a listed form. `nbsp` inserts only next to listed punctuation
210
+ or beside a quote glyph, so the inserted space never lands between two letters; `symbols`
211
+ deletes only `(`…`)` spans; `quotes` (0.3.0) deletes only a run whose landing is in
212
+ `ALNUM ∪ QUOTEMARK`, but the deletion always removes an `INLINE-SPACE` run that is *not itself*
213
+ inside a word — `hyphen`'s `WORDISH = ALNUM ∪ HYPHENISH` excludes both a U+0020 and a quote
214
+ glyph, so the deletion never lands inside a listed form (`quotes.md` 5). `I₄` is preserved.
215
+
216
+ ### I₅ … I₈
217
+
218
+ `quotes`, `apostrophe`, `symbols` and `nbsp` are the last four rules; only `nbsp` has anything
219
+ after it, and nothing runs after `nbsp` at all. `I₅` must be preserved by `apostrophe`,
220
+ `symbols` and `nbsp`; `I₆` by `symbols` and `nbsp`; `I₇` by `nbsp`.
221
+
222
+ `I₆` and `I₇` are discharged in the respective documents by the "later rule changes only code
223
+ points whose class membership is unchanged, or changed only in the direction that removes
224
+ candidacy" shape. `I₅` (0.3.0) is discharged differently, because `quotes`' own idempotency no
225
+ longer rests on that shape at all — `quotes.md` 0.1.0's Claim 3 ("capabilities on the second run
226
+ are a subset of the first run's") **no longer exists**; under mandate 1 a converted mark is a
227
+ candidate again on the next run, so capabilities are not monotone and no subset argument is
228
+ available. `quotes.md` §5's replacement is two lemmas plus a certification gate that checks its
229
+ own output rather than relying on an argument about it: **Lemma A** (glyph-blindness — no
230
+ candidate's verdict depends on *which* quote glyph a neighbour is) makes `apostrophe`'s emission
231
+ structurally inert as a neighbour (Corollary A1); **Lemma B** (space-inertness at every position
232
+ `nbsp` can reach) makes `nbsp`'s three insertion sites structurally inert. Both are CO-S
233
+ discharges in the sense of §5.1a below, over a *reachable-position* alphabet rather than a
234
+ whole-emission alphabet for Lemma B specifically — see `quotes.md` §5 for why the alphabet alone
235
+ is not sufficient there.
236
+
237
+ ---
238
+
239
+ ## 4. The two defect families
240
+
241
+ Both were found by an exhaustive sweep over every string of length 0–4 drawn from
242
+ `{ " ' - SPACE . 1 a }` across all nine locales. Both are CO violations. Neither is visible
243
+ to any per-rule idempotency argument, and both were pinned as failing tests rather than hidden
244
+ behind a precondition.
245
+
246
+ ### Family 1 — a spaced dash emitted before a full stop
247
+
248
+ `dashes` (R₃) violates `I₁` (`spaces`, R₁).
249
+
250
+ ```
251
+ de-DE: .--. → . – . → . –.
252
+ ```
253
+
254
+ Minimal shapes: `"--.`, `'--.`, `.--.`, `1--.`, `a--.`. Reproduces in all seven locales whose
255
+ `dash.parenthetical` is `em-spaced` or `en-spaced`; `en-US` is `em-tight` and is clean.
256
+
257
+ `dashes` emits `U+0020 – U+0020` for a `-spaced` locale. The trailing U+0020 now sits directly
258
+ before a full stop, which is violation **S-b**: on the next pass `spaces` deletes it, the
259
+ token becomes asymmetrically spaced, and `dashes` then declines it. Each rule is a fixed point
260
+ on its own output; the composition is not.
261
+
262
+ The same shape occurs with brackets — `(--a` → `( – a` → `(– a` (violation **S-c**) and
263
+ `a--)` → `a – )` → `a –)` (violation **S-d**).
264
+
265
+ **Repaired in `dashes`**, §3.2 step 9: a token may not be given a `-spaced` form when the
266
+ space that form would emit stands in a position `spaces` would delete. The alternative —
267
+ teaching `spaces` not to strip a space that follows a dash — was rejected: stripping a space
268
+ before punctuation is `spaces`' core job, the exception would have to fire for hyphens too
269
+ (`a - .`), and the inputs it would protect (`word – .`, `word – ,`) do not occur in real copy.
270
+ Declining leaves the input byte-identical, which is the conservative side of the ship
271
+ criterion.
272
+
273
+ ### Family 2 — a quoted hyphen acquires no-break spacing
274
+
275
+ `nbsp` (R₈) violates `I₃` (`dashes`, R₃), by way of a shape `quotes` (R₅) created.
276
+
277
+ ```
278
+ fr: "-" → «-» → «⍽-⍽» → «⍽–⍽»
279
+ ```
280
+
281
+ 37 shapes, minimal `"-"`. `fr` only, because it is the only v1 locale with a non-`none`
282
+ `quotes.*.innerSpace`.
283
+
284
+ In the first pass `dashes` sees `"-"`: a bare U+002D with no spacing, which guard P1 declines.
285
+ `quotes` then produces `«-»` and `nbsp` inserts the guillemet inner spaces. On the second pass
286
+ `dashes` sees a U+002D flanked by space-like code points on both sides — indistinguishable,
287
+ by its own classes, from a spaced parenthetical dash — and converts it.
288
+
289
+ **Repaired in `dashes`**, §3.2 step 3: a `NOBREAK-SPACE` counts as the token's spacing on the
290
+ **left only**. The fix has to live in `dashes` and not in `nbsp` for a concrete reason:
291
+ `order.json` gives `dashes` `"localeData": ["dash"]`, so it cannot read `quotes.primary.open`
292
+ and literally cannot recognise a guillemet. `nbsp` could be taught not to insert beside a
293
+ hyphen, but that would require `nbsp` to encode `dashes`' admissibility rules, which is a
294
+ layering violation and a second copy of a subtle guard. The asymmetry is justified on its own
295
+ terms: `nbsp` promotes the space **before** a dash (Russian binds an em dash to the preceding
296
+ word) and never the space after one, so a no-break space to the right of a dash was never put
297
+ there for that dash's sake.
298
+
299
+ ---
300
+
301
+ ## 5. The proof obligation on every rule
302
+
303
+ Every `spec/rules/<id>.md` §5 must contain, in addition to its own idempotency argument, a
304
+ subsection discharging CO. It answers exactly two questions:
305
+
306
+ 1. **What does this rule emit?** Enumerate the code points it can insert, and the positions in
307
+ which it can delete or replace one. This is usually three lines.
308
+ 2. **For each earlier-ordered rule, can that emission violate its invariant?** Walk §3's table
309
+ for every rule with a lower `order`. State the answer and the reason, even when the answer
310
+ is trivially no.
311
+
312
+ A rule with `order` 10 has an empty obligation and says so.
313
+
314
+ When a new rule is added, `docs/ARCHITECTURE.md` §5's five-step checklist gains an implicit
315
+ sixth item: discharge CO against every rule ordered before it, **and** re-check every rule
316
+ ordered after it, since the new rule's invariant becomes something they must now preserve.
317
+
318
+ ### 5.1a CO-S — the structural discharge, and why case analysis is not enough
319
+
320
+ CO is sound. It was nevertheless violated four times, always in the same direction — a later
321
+ rule changing an earlier rule's verdict — and three of those four were _repairs of each other_.
322
+ The pattern is worth naming, because the defect was never in CO and always in the **discharge**.
323
+
324
+ A discharge of the form _"rule `Rⱼ` changes spacing, and I checked the cases I could think of"_
325
+ is not a proof. It is a survey, it terminates when the author runs out of imagination, and each
326
+ of the three failed `dashes`/`nbsp` repairs was exactly that. The alternative is available and
327
+ is usually cheaper to write:
328
+
329
+ > ### CO-S — sufficient condition for a structural discharge
330
+ >
331
+ > Let `E(Rⱼ)` be the set of code points `Rⱼ` can emit or insert. If **every member of `E(Rⱼ)`
332
+ > is inert for `Rᵢ`** — meaning `Rᵢ` declines to form or edit any token adjacent to it — then
333
+ > `Rⱼ` cannot create work for `Rᵢ` **on any input whatsoever**, and CO is discharged for the
334
+ > pair `(i, j)` without enumerating a single case.
335
+
336
+ CO-S is not always achievable, but where it is, it is the discharge to write. The `dashes` /
337
+ `nbsp` pair is the worked example. `E(nbsp) = { U+00A0, U+202F }` — an **upper bound** as of spec
338
+ 1.3.0, not an exact set: under `narrowNbsp: "nbsp"` (nbsp.md §3.1a) it shrinks to `{ U+00A0 }`.
339
+ Every discharge against it survives that unchanged, and by construction rather than by luck: each
340
+ one names **both** code points together, through a class that holds both (`spaces`'
341
+ `PROTECTED-SPACE`, `dashes`' and `ranges`' `NOBREAK-SPACE`, `quotes`' `INLINE-SPACE`,
342
+ `apostrophe`'s `SPACELIKE`), so a discharge that holds for the pair holds for either subset. A
343
+ future option that made `E(nbsp)` **larger** would not be free this way, and would have to be
344
+ re-argued here. Three successive attempts
345
+ tried to specify _which_ of those counted as dash spacing and _on which side_:
346
+
347
+ | Formulation | Fixed | Exposed |
348
+ | ------------------------------------------------ | ----------------------- | -------------------------------------------------------------- |
349
+ | a `NOBREAK-SPACE` counts as spacing | the original tight case | `«⍽-⍽»` — family 2 |
350
+ | …counts on the **left** only | family 2 | `«⍽–␣x` — defect (d), reachable from the pipeline's own output |
351
+ | **neither counts; a dash touching one is inert** | both, and the class | — |
352
+
353
+ Only the third satisfies CO-S, and it is also the shortest to state and the only one whose
354
+ correctness does not depend on which witnesses were in front of the author. A discharge that
355
+ has to be re-derived every time the other rule changes is a defect waiting for a new locale.
356
+
357
+ **When writing a CO discharge, try CO-S first.** If the earlier rule cannot be made inert to
358
+ the later rule's whole emission alphabet, say so explicitly and enumerate — but treat the
359
+ enumeration as a known-weak argument and make the sweep (§6) the real control.
360
+
361
+ ### 5.2 Rule-local and pipeline-level claims
362
+
363
+ Every rule document has a §4 "Must not touch". Those lists were written rule-by-rule, and they
364
+ read — to anyone who has not memorised `order.json` — as promises about `transform`. Some of
365
+ them are not. A §4 bullet is really one of two different assertions:
366
+
367
+ - **[R] rule-local.** _"Given its input, this rule does not edit X."_ A statement about one
368
+ function. Always checkable from that rule's §3 alone.
369
+ - **[P] pipeline-level.** _"`transform` does not edit X."_ A statement about the composition,
370
+ and it is true only if **X survives every rule R₁…R₈** — not merely the one whose §4 it
371
+ appears in.
372
+
373
+ The gap between them is not academic. `ellipsis` §4 claimed `e.g. ..` was untouched; it is
374
+ untouched _by `ellipsis`_, but `spaces` deleted the U+0020 first, the run merged with the
375
+ abbreviation's dot, and the pipeline returned `e.g…`. Both rules were behaving exactly as
376
+ specified. The false statement was the scope of the claim, not the behaviour of either rule.
377
+
378
+ > **Derivation rule for a [P] claim.** X is protected end-to-end iff, for every rule `Rᵢ`, no
379
+ > sequence of edits by `R₁ … Rᵢ₋₁` can transform the input containing X into something `Rᵢ`
380
+ > would edit. In practice this reduces to one question per earlier rule: _can it change a code
381
+ > point adjacent to X, or delete one inside it?_ — because every rule in this pipeline decides
382
+ > from a bounded neighbourhood.
383
+
384
+ **Marking is mandatory.** Every §4 bullet in every rule document carries **[P]** or **[R]**.
385
+ A bullet describing a user-visible protection — paths, code-like text, identifiers, URLs,
386
+ abbreviations, measurements — must be **[P]**, because that is what a README reader will assume
387
+ it means; if the pipeline cannot deliver it, the defect is fixed rather than the claim
388
+ downgraded. **[R]** is reserved for statements that are genuinely about the rule's own
389
+ mechanics, and each one names the rule that can falsify it.
390
+
391
+ **Reading is not sufficient to find these.** Both instances above, and four more, were found by
392
+ _writing conformance fixtures_, not by reviewing prose — including by an agent that had read
393
+ all of these documents. A §4 bullet is a hypothesis until a fixture exercises it through the
394
+ whole pipeline. §6's sweep obligation is the mechanical half of this; per-bullet fixtures are
395
+ the other half, and they are what actually caught the class.
396
+
397
+ ### 5.3 What this is worth, honestly
398
+
399
+ CO is a sufficient condition, not a necessary one. A pipeline could violate CO and still be
400
+ idempotent, if the work one rule creates for an earlier one happens to be undone again. Such a
401
+ pipeline would be idempotent by luck and would break on the next locale, so the spec requires
402
+ CO rather than the weaker property.
403
+
404
+ CO is also only as good as the invariant statements in §3. `I₁` is stated precisely because
405
+ `spaces` is simple; `I₃` is stated loosely ("no token `dashes` would edit") because `dashes`
406
+ is not. A loose invariant means the obligation is discharged by argument rather than by
407
+ mechanical check, and arguments about `dashes` have now been wrong three times. The
408
+ compensating control is the exhaustive sweep, not the prose — which is why §6 makes it a
409
+ release gate.
410
+
411
+ ---
412
+
413
+ ## 6. Testing obligation
414
+
415
+ The idempotency property test must include, in every runtime:
416
+
417
+ 1. **A bounded exhaustive sweep, in two tiers.** Both are run for every locale in
418
+ `spec/locales/registry.json`, asserting `transform(transform(x)) == transform(x)`.
419
+
420
+ | tier | alphabet | length |
421
+ |---|---|---|
422
+ | **wide** | the consumed **and** emitted sets below | 0–5 |
423
+ | **deep** | a core of at least `1`, `-`, U+0020, `.`, U+2060 | 0–8 |
424
+ | **quotes** (spec 0.3.0) | `{ U+0022, U+0027, U+002D, U+0020, a }` ∪ every distinct quote glyph in `spec/locales/registry.json`'s locales | 0–6, per locale |
425
+
426
+ **Why the quotes tier exists, separately from wide/deep.** `quotes` (0.3.0)'s worst witness —
427
+ `‘""-"-"` in `en-GB` — is seven code points over `{ ‘ " - }`, and the committed deep tier's
428
+ core alphabet cannot reach it: it has neither the apostrophe-shaped opener nor a second
429
+ dash. Under mandate 1, `quotes.md` §6's **consumed** and **emitted** columns coincide for the
430
+ first time — every glyph the rule can produce is also now something it reads — so a tier keyed
431
+ to that single, self-consistent alphabet is the natural unit rather than folding it into wide
432
+ or deep, which are keyed to different rules' emission alphabets.
433
+
434
+ **Why two tiers, and why the bound moved.** A single length-4 bound was normative until
435
+ `dashes` defect (e) — witness `1-1 - 1`, **seven characters** — shipped underneath it. Nothing
436
+ was excluded and no precondition hid it; the bound simply could not reach it, and the suite
437
+ was green throughout. **A bound that cannot reach a known witness will hide the next one.**
438
+ The wide tier catches interactions between many characters, and grows combinatorially, so it
439
+ stays shallow; the deep tier is narrow enough (5 symbols to length 8 is 390 625 strings per
440
+ locale) to run to a length where multi-token shapes actually appear. U+2060 is in the deep
441
+ alphabet because `dashes` emits it and the guards must be inert to it (§5.1a).
442
+
443
+ **Standing obligation:** when a defect is found whose witness the committed bound cannot
444
+ reach, the bound is wrong. Raise it, or add the witness's alphabet to the deep tier, in the
445
+ same change that fixes the defect. **A biased or exhaustive generator is not
446
+ optional**: uniform random strings over a large alphabet essentially never produce `.--.`,
447
+ and the defect it hides may be the important one.
448
+
449
+ **The sweep alphabet must contain the code points the rules _emit_, not only those they
450
+ consume.** At minimum:
451
+
452
+ | | |
453
+ | ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
454
+ | **consumed** | `"` `'` `-` U+0020 `.` `1` `a`, plus U+2010 (hyphen) and U+2212 (minus sign) — spec 0.2.0 added both to `dashes`' `DASH` class (`dashes.md` §3.1) |
455
+ | **emitted** | every locale `quotes.*.open`/`close` glyph in the registry — at minimum `«` `»` `“` `”` `„` `‘` `’` — plus U+2013, U+2014, U+2026, U+00A0, U+202F, U+2011 |
456
+
457
+ This is a normative requirement and it was learned the expensive way. **Idempotency is a
458
+ statement about re-processing output**, so an alphabet drawn only from input characters
459
+ tests the wrong language: it can only reach the subset of the output space that happens to
460
+ be spelled with input characters. Defect (d) of `dashes.md` §5.2 — `«⍽–␣"`, a shape built
461
+ entirely from characters the pipeline itself emits — was unreachable until the emitted
462
+ alphabet was added, and before that it surfaced on roughly **one property-test seed in
463
+ three**, where it read as flaky infrastructure rather than as a defect. A test that fails
464
+ one run in three and passes the rest is worse than no test, because it trains everyone
465
+ looking at it to re-run.
466
+
467
+ Two consequences for whoever maintains the sweep: the alphabet **grows when a locale is
468
+ added**, since a new locale's quote glyphs are new output characters; and the sweep must be
469
+ **deterministic and exhaustive** rather than sampled, so that a failure is reproducible from
470
+ the seed-free description alone.
471
+
472
+ **Named witnesses pinned as individual conformance fixtures, not only as sweep coverage**
473
+ (spec 0.3.0): `‘""-"-"` (`en-GB`), `'"‘` (`ru`), `«␣**"` (`fr`), `"<p class="x">«[t](u)`
474
+ (`fr-CA`, `html`), `" --x"` (`en-US`), `" "`, and `«»`. Each exercises the `quotes` gate or the
475
+ `I₃` landing guard on a shape the general sweep would eventually reach but that a reviewer
476
+ should be able to find by name — see `spec/rules/quotes.md` §6.
477
+
478
+ Two more, found by running the implementation rather than specified by either design
479
+ document: `«"<p class="x">"` (`fr`, `html`) and `««”` (`fr`, `text`), pinned in
480
+ `spec/fixtures/fr.json` as `fr-quotes-030-cobug-html-tag-boundary` and
481
+ `fr-quotes-030-cobug-double-open-then-closer`. Both exercise `quotes.md` §3.2's V1
482
+ `gapInsertable` clause — without it, `nbsp` inserting a space at a position V1's literal
483
+ adjacency check depended on lets a pairing the gate declined on one pipeline pass certify on
484
+ the next, which is a CO violation (§2) rather than an ordinary idempotency defect.
485
+
486
+ 2. **A per-rule sweep.** The same assertion with a single rule enabled, which localises a
487
+ failure to a rule rather than to the composition.
488
+ 3. **A composition sweep.** The same assertion with the full pipeline. A failure here that
489
+ passes (2) is a CO violation, and the report should say so — that is the signal this
490
+ document exists to make legible.
491
+
492
+ 4. **A test whose carrier dies keeps passing, and proves nothing.** When a rule narrows, some
493
+ assertions elsewhere stop exercising what they were written for while continuing to go green
494
+ — they pass *for a different reason than when they were written*. That is strictly worse than
495
+ a failing test, because nothing draws attention to it.
496
+
497
+ This is not hypothetical: `dashes` §3.2 step 2a stopped the rule restyling an authored dash,
498
+ and every test and worked example carried on `–` → `␣–␣` became inert overnight. One of them
499
+ was in `modes.md` §3.4's normative table, where it had been the illustration of the whole
500
+ edge-growth rule. It was caught only because a *sibling* assertion in the same test failed and
501
+ someone asked why the other two still passed.
502
+
503
+ **Obligation when a rule narrows:** audit every assertion that uses the narrowed construct as
504
+ its carrier, and for each one ask whether it still fails when the behaviour it names is broken.
505
+ Rebuild it on a live carrier or delete it. The formulation worth keeping is the implementer's:
506
+ **the carrier is dead, the assertion is alive, the meaning is lost.**
507
+
508
+ Two related checks belong to the same discipline:
509
+
510
+ - **A round-trip fixture must have teeth.** A document asserted to come back byte-identical
511
+ proves something about skip lists only if it *would* change otherwise. If it is also
512
+ unchanged in `text` mode, it proves nothing at all — the assertion holds for the trivial
513
+ reason. Every such fixture should be verified to change materially in `text` mode.
514
+ - **A guard that has become unreachable is deleted, not left to rot** (§5.1a) — the same
515
+ principle one layer down, applied to the rule instead of the test.
516
+
517
+ **Preconditions (`fc.pre`, `assume`, and equivalents) that exclude a failing shape are not
518
+ permitted in the committed suite** except as a temporary, named, individually-pinned
519
+ containment for a defect that is already reported. Every such precondition must name the
520
+ document section that will remove it.
521
+
522
+ ---
523
+
524
+ ## 6a. The Unicode pin, and what it can portably require
525
+
526
+ `spec/UNICODE` contains `17.0`. Until this section it had **no reader**: no rule document cited
527
+ it, so an implementation could derive its character tables from any UCD version and still pass
528
+ conformance. That is the defect — not the file's absence, but its inertness.
529
+
530
+ ### 6a.1 What the pin means
531
+
532
+ Two claims get conflated, and only one of them is portable.
533
+
534
+ 1. **"The embedded tables are those derived from UCD 17.0."** Enforceable in every runtime,
535
+ determines output, and is what the pin means. Each port generates its `LETTER` (L ∪ M),
536
+ `UPPER` (Lu ∪ Lt) and simple-uppercase tables from the pinned UCD and checks them in.
537
+ 2. **"The host runtime's UCD is 17.0."** **Not portable, and therefore not required.** JS exposes
538
+ `process.versions.unicode`, Python `unicodedata.unidata_version`, Go `unicode.Version` — but
539
+ PHP only through `intl`/ICU, which is not always installed, and Ruby has no dependable public
540
+ accessor. A normative requirement of this form would be unimplementable in at least one target
541
+ runtime, which is what ARCHITECTURE.md §4 exists to prevent.
542
+
543
+ > **`spec/UNICODE` is normative for the derived tables, not for the host runtime.** Every rule
544
+ > document that uses `LETTER`, `UPPER` or a case mapping must cite it. A port whose host UCD
545
+ > version *is* readable should additionally assert it against the pin as a cheap extra gate;
546
+ > where it is not readable that gate does not exist, and the spec must not pretend otherwise.
547
+
548
+ A host-side drift detector proves that the table matches *some* UCD and that this host agrees
549
+ with it. **Only fixtures can prove that five runtimes agree with each other.**
550
+
551
+ ### 6a.2 Canary fixtures
552
+
553
+ Ordinary conformance cases whose expected output depends on the tables being right, so that table
554
+ drift surfaces in the conformance run rather than in one runtime's unit tests.
555
+
556
+ **On version canaries, honestly.** I cannot name a code point whose general category or simple
557
+ uppercase mapping demonstrably changed between UCD 16.0 and 17.0 from anything I can read here.
558
+ Naming one on recollection would be worse than naming none: **a canary that cannot fire advertises
559
+ a guarantee nobody has.** The version-drift canary must therefore be **generated, not hand-picked**
560
+ — a build step diffs the pinned UCD against the previous major version, takes the code points whose
561
+ `General_Category` or `Simple_Uppercase_Mapping` differs, and emits fixture rows asserting the
562
+ pinned behaviour. Generated rows fire by construction, and the set is re-derived whenever the pin
563
+ moves. If the diff is empty for a given version pair, the suite should say so rather than silently
564
+ contain nothing.
565
+
566
+ **Derivation canaries, which I can specify with confidence and which catch the likelier failure.**
567
+ The realistic defect is not "a port used UCD 16.0"; it is **"a port called the host's letter
568
+ predicate instead of the embedded table"**. Three cases, each chosen because a naive host call
569
+ gives a *different answer* from this spec's definition, in every Unicode version:
570
+
571
+ | Canary | Code point | Fires when |
572
+ |---|---|---|
573
+ | **`LETTER` includes marks** | U+0301 combining acute, category `Mn` | the port used `\p{L}`, `unicode.IsLetter`, `str.isalpha` or `ctype_alpha`, all of which exclude `Mn`. This spec's `LETTER` is **L ∪ M**, so a decomposed `é` (`e` + U+0301) must behave as one letter — `apostrophe.md` §6 case 2 already depends on it |
574
+ | **`UPPER` includes titlecase** | U+01C5 `Dž`, category `Lt` | the port used an `isUpper` that tests `Lu` only. This spec's `UPPER` is **Lu ∪ Lt** (`nbsp.md` §3.1) |
575
+ | **Case mapping is simple and locale-independent** | U+0069 `i` / U+0049 `I` | the port used a locale-sensitive uppercase. Under a Turkish host locale `i` maps to `İ` (U+0130), so `nbsp`'s first-character leniency (§3.5) stops matching a capitalised short word. This is ARCHITECTURE.md §4.4's dotless-ı hazard, made testable |
576
+
577
+ These prove the tables were **derived correctly**; the generated rows prove they were derived from
578
+ the **pinned version**. Both are needed, and only the second depends on data I cannot supply here.
579
+
580
+ ---
581
+
582
+ ## 7. Open questions
583
+
584
+ 1. **`I₃` is not stated precisely enough to check mechanically.** "No token `dashes` would
585
+ edit" is a restatement of the rule, not an invariant. A checkable form would enumerate the
586
+ shapes — and the fact that I cannot write it in half a page is itself evidence that
587
+ `dashes` is the most complex rule in the spec and the one most likely to break again.
588
+ 2. **CO has not been verified for rule subsets.** The `rules` option lets a caller disable any
589
+ rule. Disabling `spaces` removes `I₁` from the obligation set, which is harmless; but
590
+ disabling a rule can also _expose_ text to a later rule that the disabled rule would have
591
+ normalised, and no sweep currently runs over subsets. The combinatorics are mild (2⁸ = 256
592
+ configurations × the length-4 sweep) and this should probably be a CI job.
593
+ 3. _(Settled.)_ Modes are now covered by [modes.md](modes.md), which was written to answer
594
+ this item. The premise it was raised under turned out to be the wrong one: the pipeline does
595
+ **not** run per span. Spans are concatenated with an explicit boundary marker and the
596
+ pipeline runs once over the whole array, so a space at a span boundary and a dash in the
597
+ neighbouring span are seen together — and are declined together, because the marker breaks
598
+ the adjacency that would make them a token. `modes.md` §5 adds one obligation this document
599
+ cannot see: **the span partition must be stable between runs**, which constrains what code
600
+ points a rule may emit.
601
+ 4. _(Settled, spec 0.3.0.)_ **The `nbsp`-before-`quotes` direction.** `quotes.md` §2.1's
602
+ constraint **Q-P** — every `nbsp.beforePunctuation`/`nbsp.narrowBeforePunctuation` entry must
603
+ be a member of `quotes`' `CLOSEISH` — is enforced by `scripts/validate-spec.mjs` and is what
604
+ makes `quotes.md` §5's Lemma B cover `nbsp`'s N1/N2 insertion sites as well as N8, by proof
605
+ rather than by case analysis over the shipped data alone. All ten shipped locales satisfy it.