polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,605 @@
|
|
|
1
|
+
# Pipeline idempotency
|
|
2
|
+
|
|
3
|
+
**Not a rule.** No entry in `spec/rules/order.json`, no locale data, no edits. This document
|
|
4
|
+
states the invariant that `transform` as a whole must satisfy, proves that per-rule
|
|
5
|
+
idempotency does not imply it, and defines the obligation each rule must discharge so that it
|
|
6
|
+
does.
|
|
7
|
+
**Spec version:** 1.2.0 (0.1.0 for everything except S-b's word-start clause, added in 1.2.0).
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## 1. Purpose
|
|
12
|
+
|
|
13
|
+
PLAN.md §3.4 makes `transform(transform(x)) == transform(x)` a hard invariant **on the public
|
|
14
|
+
function**, not on individual rules. Every rule document carries a §5 arguing that _that rule_
|
|
15
|
+
is a fixed point on its own output. Those arguments are necessary and they are not sufficient:
|
|
16
|
+
the composition of eight individually idempotent functions is not in general idempotent, and
|
|
17
|
+
in this pipeline it demonstrably was not. Two defect families were found by exhaustive search
|
|
18
|
+
over short strings, both of them invisible to every per-rule argument because no per-rule
|
|
19
|
+
argument is allowed to mention another rule.
|
|
20
|
+
|
|
21
|
+
This document supplies the missing layer. It is short, and the obligation it imposes is
|
|
22
|
+
mechanical enough to be checked by reading a rule's "Must not touch" section against a table.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## 2. The invariant, and why per-rule idempotency does not give it
|
|
27
|
+
|
|
28
|
+
Write the pipeline as `T = R₈ ∘ R₇ ∘ … ∘ R₂ₐ ∘ R₂ ∘ R₁`, the rules in `order.json` order:
|
|
29
|
+
|
|
30
|
+
| # | rule | order |
|
|
31
|
+
| --- | ------------ | ----- |
|
|
32
|
+
| R₁ | `spaces` | 10 |
|
|
33
|
+
| R₂ | `ellipsis` | 20 |
|
|
34
|
+
| R₂ₐ | `ranges` | 25 |
|
|
35
|
+
| R₃ | `dashes` | 30 |
|
|
36
|
+
| R₄ | `hyphen` | 35 |
|
|
37
|
+
| R₅ | `quotes` | 40 |
|
|
38
|
+
| R₆ | `apostrophe` | 50 |
|
|
39
|
+
| R₇ | `symbols` | 60 |
|
|
40
|
+
| R₈ | `nbsp` | 70 |
|
|
41
|
+
|
|
42
|
+
(A disabled rule is removed from the sequence and never reorders the rest, so every statement
|
|
43
|
+
below holds for any subset, in the same relative order.)
|
|
44
|
+
|
|
45
|
+
**Why `ranges` is `R₂ₐ` and not `R₃`.** It was split out of `dashes` in spec 0.5.0, after the
|
|
46
|
+
subscripts in this document had been cited by name from eight other rule documents. Renumbering
|
|
47
|
+
`R₃ … R₈` to make the sequence contiguous would rewrite every one of those references for no
|
|
48
|
+
gain in meaning, so the rule takes the letter and the composition `T = R₈ ∘ R₇ ∘ … ∘ R₂ₐ ∘ … ∘ R₁`
|
|
49
|
+
reads with it in place at order 25. `ranges` is also the one rule that is **off by default**, so
|
|
50
|
+
for most callers the sequence is literally the eight numbered ones; every proof below is stated
|
|
51
|
+
over "the rules that run", which is the subset the caller's `rules` option selected.
|
|
52
|
+
|
|
53
|
+
For each rule define the predicate
|
|
54
|
+
|
|
55
|
+
> **`Iᵢ(y)` ⟺ `Rᵢ(y) = y`** — "rule `i` is a no-op on `y`".
|
|
56
|
+
|
|
57
|
+
**Lemma.** If `y = T(x)` satisfies `I₁ ∧ I₂ ∧ I₂ₐ ∧ … ∧ I₈`, then `T(y) = y`.
|
|
58
|
+
_Proof._ `R₁(y) = y` by `I₁`; then `R₂(R₁(y)) = R₂(y) = y` by `I₂`; then `R₂ₐ(y) = y` by `I₂ₐ`;
|
|
59
|
+
and so on through `R₈`. ∎
|
|
60
|
+
|
|
61
|
+
So the whole problem reduces to: **make the final output a fixed point of every rule, not just
|
|
62
|
+
of the last one that touched it.**
|
|
63
|
+
|
|
64
|
+
Per-rule idempotency gives `Iᵢ` immediately after `Rᵢ` runs. What it does not give is that
|
|
65
|
+
`Iᵢ` still holds at the _end_ of the pipeline. A later rule may undo it. That is exactly the
|
|
66
|
+
gap, and it yields the obligation:
|
|
67
|
+
|
|
68
|
+
> ### The composition obligation (CO)
|
|
69
|
+
>
|
|
70
|
+
> **For every pair `i < j`: if `Iᵢ(y)` holds, then `Iᵢ(Rⱼ(y))` holds.**
|
|
71
|
+
>
|
|
72
|
+
> In words: **a rule must never create work for an earlier-ordered rule.**
|
|
73
|
+
|
|
74
|
+
With CO, an induction over `j` gives the Lemma's premise: after `Rⱼ` has run, `I₁ … Iⱼ` all
|
|
75
|
+
hold — `Iⱼ` by `Rⱼ`'s own idempotency, and `I₁ … Iⱼ₋₁` because `Rⱼ` preserved them. After the
|
|
76
|
+
last rule all of them hold, and `T(T(x)) = T(x)`. (`R₂ₐ` takes its place in the induction between
|
|
77
|
+
`R₂` and `R₃`; the letter is a numbering artefact of the spec 0.5.0 split, not a gap in the
|
|
78
|
+
sequence.)
|
|
79
|
+
|
|
80
|
+
Note what CO does **not** require: nothing about `j < i`. A rule may freely create work for a
|
|
81
|
+
_later_ rule, because the later rule has not run yet and will clean it up in the same pass.
|
|
82
|
+
That asymmetry is the whole content of "the rules run in a fixed order".
|
|
83
|
+
|
|
84
|
+
### 2.1 Why not simply iterate to a fixed point
|
|
85
|
+
|
|
86
|
+
Because it would make the invariant true by construction and hide precisely the defects this
|
|
87
|
+
document exists to find. It would also weaken the per-rule guarantee the package sells (a rule
|
|
88
|
+
that is a fixed point in isolation is a testable, portable claim; "the loop converges
|
|
89
|
+
eventually" is not), and any difference in iteration bound between two runtimes — or any input
|
|
90
|
+
where one runtime converges in two passes and another in three — becomes a conformance
|
|
91
|
+
divergence rather than a bug. **The rules must compose correctly.** Iteration is forbidden.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## 3. The output invariants
|
|
96
|
+
|
|
97
|
+
To discharge CO a rule author needs to know what `Iᵢ` actually says for each earlier rule, in
|
|
98
|
+
terms concrete enough to check. These are the four that constrain anything.
|
|
99
|
+
|
|
100
|
+
### I₁ — `spaces`
|
|
101
|
+
|
|
102
|
+
`spaces` deletes U+0020 in three situations and collapses runs. `I₁(y)` therefore requires
|
|
103
|
+
that `y` contains **no U+0020 in any of these positions**:
|
|
104
|
+
|
|
105
|
+
- **S-a** a run of two or more U+0020 with a `CONTENT` code point on both sides;
|
|
106
|
+
- **S-b** a U+0020 whose right neighbour is in `STRIP-BEFORE` = { U+002C, U+002E, U+003B,
|
|
107
|
+
U+003A, U+0021, U+003F } and whose left neighbour is `CONTENT` — **and**, when that right
|
|
108
|
+
neighbour is U+002E, the maximal run of { U+002E, U+2026 } beginning there has length exactly
|
|
109
|
+
1 and the code point after that dot is neither a `LETTER` nor an ASCII digit (`spaces.md` §3.4;
|
|
110
|
+
the second clause since spec 1.2.0) — **and** unless the emoticon guard's eye side fires, i.e. that right
|
|
111
|
+
neighbour is the eye of a recognised emoticon (`spaces.md` §3.6);
|
|
112
|
+
U+2026 is not a member of `STRIP-BEFORE`;
|
|
113
|
+
- **S-c** a U+0020 whose left neighbour is in { U+0028, U+005B, U+007B } — unless the
|
|
114
|
+
empty-bracket guard applies, and unless the emoticon guard's mouth side fires, i.e. that left
|
|
115
|
+
neighbour is the mouth of a recognised emoticon (`spaces.md` §3.6);
|
|
116
|
+
- **S-d** a U+0020 whose right neighbour is in { U+0029, U+005D, U+007D } — same
|
|
117
|
+
empty-bracket proviso. The emoticon guard does not apply here: it suppresses the
|
|
118
|
+
`STRIP-BEFORE` and `OPEN-BRACKET` clauses only, never this one.
|
|
119
|
+
|
|
120
|
+
See `spaces.md` §3.1–§3.3 for the exact definitions of `CONTENT` and the guard, and §3.6 for the
|
|
121
|
+
emoticon guard's two sides. Both emoticon provisos, and the lone-dot condition's word-start clause
|
|
122
|
+
added in spec 1.2.0, narrow the forbidden set, so every discharge
|
|
123
|
+
already written against S-b, S-c or S-d stays valid without re-derivation — a rule that emits no
|
|
124
|
+
U+0020 in the wider set emits none in the narrower one.
|
|
125
|
+
|
|
126
|
+
**Who can violate it.** Only a rule that emits U+0020. In the current pipeline that is
|
|
127
|
+
`dashes` alone (`-spaced` forms). `nbsp` emits U+00A0 and U+202F, which are `CONTENT` to
|
|
128
|
+
`spaces` and are never touched. `quotes` (0.3.0) deletes a code point inside a quote pair when
|
|
129
|
+
`innerSpace = "none"`, but never a U+0020 that survives the deletion — the maximal-run deletion
|
|
130
|
+
and the landing guard together ensure the two surviving neighbours are always `CONTENT`
|
|
131
|
+
(`quotes.md` 3.7, 5). Every other rule replaces code points one-for-one or shortens a run, and
|
|
132
|
+
none of them can create a U+0020 or bring two apart-standing ones together.
|
|
133
|
+
|
|
134
|
+
### I₂ — `ellipsis`
|
|
135
|
+
|
|
136
|
+
`I₂(y)` requires no maximal run over { U+002E, U+2026 } that is either three or more code
|
|
137
|
+
points long, or of length ≥ 2 containing a U+2026; plus the locale's terminal form.
|
|
138
|
+
|
|
139
|
+
**Who can violate it.** Only a rule that emits U+002E or U+2026, or that deletes a code point
|
|
140
|
+
standing between two such runs. `quotes` (0.3.0) deletes code points, but only a maximal
|
|
141
|
+
`INLINE-SPACE` run landing on `ALNUM ∪ QUOTEMARK` — never a code point standing between two dot
|
|
142
|
+
runs, since a quote glyph is never itself U+002E/U+2026 and the deletion's landing class contains
|
|
143
|
+
neither (`quotes.md` 5, composition obligation against `I₂`). `I₂` is otherwise unconditionally
|
|
144
|
+
preserved and needs no attention from rule authors, but it is listed so that a future rule
|
|
145
|
+
emitting a full stop knows it has an obligation.
|
|
146
|
+
|
|
147
|
+
### I₂ₐ — `ranges`
|
|
148
|
+
|
|
149
|
+
`I₂ₐ(y)` requires that no **range candidate** in `y` is one `ranges` would edit — `ranges.md`
|
|
150
|
+
§3.2 and §3.2a for candidacy, §3.2's G1-G5 for the guards, §3.3 for the replacement. A candidate
|
|
151
|
+
is a dash token whose two flanks are `DIGIT`, or are `DIGIT` once a `CLOSED-SYMBOL` matched on
|
|
152
|
+
the opposite member has been walked over.
|
|
153
|
+
|
|
154
|
+
**Who can violate it.** `dashes` (R₃), and only `dashes` — which is why the obligation is
|
|
155
|
+
discharged in that rule's own §5.3 rather than here. `ranges`' guards G1, G2 and G3 read
|
|
156
|
+
`before`/`after`, code points outside the token, and a `dashes` edit that turns a tight token
|
|
157
|
+
into a spaced one replaces a dash at exactly such a position with a U+0020. Two of `dashes`'
|
|
158
|
+
shared guards exist for this and no other reason: the **cluster guard** (`dashes.md` §3.2
|
|
159
|
+
step 7) makes the whole neighbourhood inert when it holds two dash runs, and **T1** (§3.2
|
|
160
|
+
step 8) declines the tight-to-spaced transition at the one-space distance the cluster guard
|
|
161
|
+
cannot see, because a cluster ends at a space.
|
|
162
|
+
|
|
163
|
+
**Spec 1.3.0 widened `I₂ₐ`'s domain, and T1 with it.** `ranges.md` §3.2a made a `CLOSED-SYMBOL`
|
|
164
|
+
written closed up to a digit run part of a range member, so T1's reach became transparent to one
|
|
165
|
+
such symbol at each of two positions per side. `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
|
|
166
|
+
`$1 - $1--a` are the witnesses, one per transparency position and side, each of which produced a
|
|
167
|
+
`transform` idempotency defect before the amendment in the locales whose `dash.parenthetical` is
|
|
168
|
+
spaced **and** whose `dash.range` is not `"none"`, and each of which now behaves exactly as its
|
|
169
|
+
all-digit analogue always did. The remedy is unconditional on `dash.range`, so its cost is wider
|
|
170
|
+
than the defect was — see `ranges.md` §3.2a. The cluster guard's
|
|
171
|
+
alphabet was deliberately **not** widened — see `dashes.md` §3.2 steps 7-8 for the comparison.
|
|
172
|
+
|
|
173
|
+
**Nothing ordered after `ranges` can violate `I₂ₐ`, and the argument is positional, not
|
|
174
|
+
alphabetic.** `hyphen` (R₄) emits U+2011, which *is* in `INERT-DASH` and therefore *is* in G2's
|
|
175
|
+
exclusion set — the alphabetic argument would be false. What holds instead is that `hyphen`
|
|
176
|
+
replaces a U+002D **in place**, inside a word, so any position where its output could satisfy G2
|
|
177
|
+
already held a `DASH` and already failed G2. `quotes` (R₅) and `apostrophe` (R₆) replace code
|
|
178
|
+
points one for one with quote glyphs, and `quotes`' one deletion removes an `INLINE-SPACE` run
|
|
179
|
+
whose two sides are a quote glyph and a `DELETE-LANDING` member — neither is a digit or a
|
|
180
|
+
`CLOSED-SYMBOL`, so no range member's `before`/`after` moves. `symbols` (R₇) emits only U+00A9,
|
|
181
|
+
U+00AE and U+2122 and deletes only a `(`…`)` span. `nbsp` (R₈) converts a space it found and
|
|
182
|
+
never inserts one where a guard reads (`nbsp.md` §3.3), and U+00A0/U+202F satisfy none of G1, G2
|
|
183
|
+
or G3.
|
|
184
|
+
|
|
185
|
+
### I₃ — `dashes`
|
|
186
|
+
|
|
187
|
+
`I₃(y)` requires that no **dash token** in `y` is one `dashes` would edit — see `dashes.md`
|
|
188
|
+
§3.2 for the admissibility gate and §3.3/§3.4 for the branches. In practice a later rule
|
|
189
|
+
violates `I₃` when it changes the **spacing** around a U+002D/U+2013/U+2014, because spacing
|
|
190
|
+
is what `dashes` reads to decide whether a stroke is a parenthetical dash at all.
|
|
191
|
+
|
|
192
|
+
**Who can violate it.** Nobody, since spec 0.1.0 — and it is worth recording that this line
|
|
193
|
+
previously read _"`nbsp`, which inserts and converts space-like code points"_, which was true
|
|
194
|
+
and was the source of two defects. `dashes` now treats both U+00A0 and U+202F as making an
|
|
195
|
+
adjacent token inert (`dashes.md` §3.2 step 3), so `E(nbsp) = { U+00A0, U+202F }` is wholly
|
|
196
|
+
inert for it and **CO-S** discharges the pair structurally. `apostrophe` and `symbols`
|
|
197
|
+
replace code points one-for-one with characters in none of `dashes`' classes; `hyphen` emits
|
|
198
|
+
U+2011, which `dashes` treats identically to U+002D everywhere it matters. `quotes` (0.3.0)
|
|
199
|
+
replaces code points one-for-one with quote glyphs (also outside every `dashes` class) and, when
|
|
200
|
+
`innerSpace = "none"`, deletes a maximal `INLINE-SPACE` run whose two sides are a quote glyph and
|
|
201
|
+
a `DELETE-LANDING` member — neither in `DASH ∪ INERT-DASH`, so no dash run is ever adjacent to a
|
|
202
|
+
deleted run and no token's `lsp`/`rsp` changes (`quotes.md` 3.7, 5, composition obligation
|
|
203
|
+
against `I₃`).
|
|
204
|
+
|
|
205
|
+
### I₄ — `hyphen`
|
|
206
|
+
|
|
207
|
+
`I₄(y)` requires no occurrence of a listed `hyphen` form whose hyphen is still U+002D. A later
|
|
208
|
+
rule violates it only by changing a form's word boundaries, i.e. by inserting or removing a
|
|
209
|
+
code point immediately beside a listed form. `nbsp` inserts only next to listed punctuation
|
|
210
|
+
or beside a quote glyph, so the inserted space never lands between two letters; `symbols`
|
|
211
|
+
deletes only `(`…`)` spans; `quotes` (0.3.0) deletes only a run whose landing is in
|
|
212
|
+
`ALNUM ∪ QUOTEMARK`, but the deletion always removes an `INLINE-SPACE` run that is *not itself*
|
|
213
|
+
inside a word — `hyphen`'s `WORDISH = ALNUM ∪ HYPHENISH` excludes both a U+0020 and a quote
|
|
214
|
+
glyph, so the deletion never lands inside a listed form (`quotes.md` 5). `I₄` is preserved.
|
|
215
|
+
|
|
216
|
+
### I₅ … I₈
|
|
217
|
+
|
|
218
|
+
`quotes`, `apostrophe`, `symbols` and `nbsp` are the last four rules; only `nbsp` has anything
|
|
219
|
+
after it, and nothing runs after `nbsp` at all. `I₅` must be preserved by `apostrophe`,
|
|
220
|
+
`symbols` and `nbsp`; `I₆` by `symbols` and `nbsp`; `I₇` by `nbsp`.
|
|
221
|
+
|
|
222
|
+
`I₆` and `I₇` are discharged in the respective documents by the "later rule changes only code
|
|
223
|
+
points whose class membership is unchanged, or changed only in the direction that removes
|
|
224
|
+
candidacy" shape. `I₅` (0.3.0) is discharged differently, because `quotes`' own idempotency no
|
|
225
|
+
longer rests on that shape at all — `quotes.md` 0.1.0's Claim 3 ("capabilities on the second run
|
|
226
|
+
are a subset of the first run's") **no longer exists**; under mandate 1 a converted mark is a
|
|
227
|
+
candidate again on the next run, so capabilities are not monotone and no subset argument is
|
|
228
|
+
available. `quotes.md` §5's replacement is two lemmas plus a certification gate that checks its
|
|
229
|
+
own output rather than relying on an argument about it: **Lemma A** (glyph-blindness — no
|
|
230
|
+
candidate's verdict depends on *which* quote glyph a neighbour is) makes `apostrophe`'s emission
|
|
231
|
+
structurally inert as a neighbour (Corollary A1); **Lemma B** (space-inertness at every position
|
|
232
|
+
`nbsp` can reach) makes `nbsp`'s three insertion sites structurally inert. Both are CO-S
|
|
233
|
+
discharges in the sense of §5.1a below, over a *reachable-position* alphabet rather than a
|
|
234
|
+
whole-emission alphabet for Lemma B specifically — see `quotes.md` §5 for why the alphabet alone
|
|
235
|
+
is not sufficient there.
|
|
236
|
+
|
|
237
|
+
---
|
|
238
|
+
|
|
239
|
+
## 4. The two defect families
|
|
240
|
+
|
|
241
|
+
Both were found by an exhaustive sweep over every string of length 0–4 drawn from
|
|
242
|
+
`{ " ' - SPACE . 1 a }` across all nine locales. Both are CO violations. Neither is visible
|
|
243
|
+
to any per-rule idempotency argument, and both were pinned as failing tests rather than hidden
|
|
244
|
+
behind a precondition.
|
|
245
|
+
|
|
246
|
+
### Family 1 — a spaced dash emitted before a full stop
|
|
247
|
+
|
|
248
|
+
`dashes` (R₃) violates `I₁` (`spaces`, R₁).
|
|
249
|
+
|
|
250
|
+
```
|
|
251
|
+
de-DE: .--. → . – . → . –.
|
|
252
|
+
```
|
|
253
|
+
|
|
254
|
+
Minimal shapes: `"--.`, `'--.`, `.--.`, `1--.`, `a--.`. Reproduces in all seven locales whose
|
|
255
|
+
`dash.parenthetical` is `em-spaced` or `en-spaced`; `en-US` is `em-tight` and is clean.
|
|
256
|
+
|
|
257
|
+
`dashes` emits `U+0020 – U+0020` for a `-spaced` locale. The trailing U+0020 now sits directly
|
|
258
|
+
before a full stop, which is violation **S-b**: on the next pass `spaces` deletes it, the
|
|
259
|
+
token becomes asymmetrically spaced, and `dashes` then declines it. Each rule is a fixed point
|
|
260
|
+
on its own output; the composition is not.
|
|
261
|
+
|
|
262
|
+
The same shape occurs with brackets — `(--a` → `( – a` → `(– a` (violation **S-c**) and
|
|
263
|
+
`a--)` → `a – )` → `a –)` (violation **S-d**).
|
|
264
|
+
|
|
265
|
+
**Repaired in `dashes`**, §3.2 step 9: a token may not be given a `-spaced` form when the
|
|
266
|
+
space that form would emit stands in a position `spaces` would delete. The alternative —
|
|
267
|
+
teaching `spaces` not to strip a space that follows a dash — was rejected: stripping a space
|
|
268
|
+
before punctuation is `spaces`' core job, the exception would have to fire for hyphens too
|
|
269
|
+
(`a - .`), and the inputs it would protect (`word – .`, `word – ,`) do not occur in real copy.
|
|
270
|
+
Declining leaves the input byte-identical, which is the conservative side of the ship
|
|
271
|
+
criterion.
|
|
272
|
+
|
|
273
|
+
### Family 2 — a quoted hyphen acquires no-break spacing
|
|
274
|
+
|
|
275
|
+
`nbsp` (R₈) violates `I₃` (`dashes`, R₃), by way of a shape `quotes` (R₅) created.
|
|
276
|
+
|
|
277
|
+
```
|
|
278
|
+
fr: "-" → «-» → «⍽-⍽» → «⍽–⍽»
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
37 shapes, minimal `"-"`. `fr` only, because it is the only v1 locale with a non-`none`
|
|
282
|
+
`quotes.*.innerSpace`.
|
|
283
|
+
|
|
284
|
+
In the first pass `dashes` sees `"-"`: a bare U+002D with no spacing, which guard P1 declines.
|
|
285
|
+
`quotes` then produces `«-»` and `nbsp` inserts the guillemet inner spaces. On the second pass
|
|
286
|
+
`dashes` sees a U+002D flanked by space-like code points on both sides — indistinguishable,
|
|
287
|
+
by its own classes, from a spaced parenthetical dash — and converts it.
|
|
288
|
+
|
|
289
|
+
**Repaired in `dashes`**, §3.2 step 3: a `NOBREAK-SPACE` counts as the token's spacing on the
|
|
290
|
+
**left only**. The fix has to live in `dashes` and not in `nbsp` for a concrete reason:
|
|
291
|
+
`order.json` gives `dashes` `"localeData": ["dash"]`, so it cannot read `quotes.primary.open`
|
|
292
|
+
and literally cannot recognise a guillemet. `nbsp` could be taught not to insert beside a
|
|
293
|
+
hyphen, but that would require `nbsp` to encode `dashes`' admissibility rules, which is a
|
|
294
|
+
layering violation and a second copy of a subtle guard. The asymmetry is justified on its own
|
|
295
|
+
terms: `nbsp` promotes the space **before** a dash (Russian binds an em dash to the preceding
|
|
296
|
+
word) and never the space after one, so a no-break space to the right of a dash was never put
|
|
297
|
+
there for that dash's sake.
|
|
298
|
+
|
|
299
|
+
---
|
|
300
|
+
|
|
301
|
+
## 5. The proof obligation on every rule
|
|
302
|
+
|
|
303
|
+
Every `spec/rules/<id>.md` §5 must contain, in addition to its own idempotency argument, a
|
|
304
|
+
subsection discharging CO. It answers exactly two questions:
|
|
305
|
+
|
|
306
|
+
1. **What does this rule emit?** Enumerate the code points it can insert, and the positions in
|
|
307
|
+
which it can delete or replace one. This is usually three lines.
|
|
308
|
+
2. **For each earlier-ordered rule, can that emission violate its invariant?** Walk §3's table
|
|
309
|
+
for every rule with a lower `order`. State the answer and the reason, even when the answer
|
|
310
|
+
is trivially no.
|
|
311
|
+
|
|
312
|
+
A rule with `order` 10 has an empty obligation and says so.
|
|
313
|
+
|
|
314
|
+
When a new rule is added, `docs/ARCHITECTURE.md` §5's five-step checklist gains an implicit
|
|
315
|
+
sixth item: discharge CO against every rule ordered before it, **and** re-check every rule
|
|
316
|
+
ordered after it, since the new rule's invariant becomes something they must now preserve.
|
|
317
|
+
|
|
318
|
+
### 5.1a CO-S — the structural discharge, and why case analysis is not enough
|
|
319
|
+
|
|
320
|
+
CO is sound. It was nevertheless violated four times, always in the same direction — a later
|
|
321
|
+
rule changing an earlier rule's verdict — and three of those four were _repairs of each other_.
|
|
322
|
+
The pattern is worth naming, because the defect was never in CO and always in the **discharge**.
|
|
323
|
+
|
|
324
|
+
A discharge of the form _"rule `Rⱼ` changes spacing, and I checked the cases I could think of"_
|
|
325
|
+
is not a proof. It is a survey, it terminates when the author runs out of imagination, and each
|
|
326
|
+
of the three failed `dashes`/`nbsp` repairs was exactly that. The alternative is available and
|
|
327
|
+
is usually cheaper to write:
|
|
328
|
+
|
|
329
|
+
> ### CO-S — sufficient condition for a structural discharge
|
|
330
|
+
>
|
|
331
|
+
> Let `E(Rⱼ)` be the set of code points `Rⱼ` can emit or insert. If **every member of `E(Rⱼ)`
|
|
332
|
+
> is inert for `Rᵢ`** — meaning `Rᵢ` declines to form or edit any token adjacent to it — then
|
|
333
|
+
> `Rⱼ` cannot create work for `Rᵢ` **on any input whatsoever**, and CO is discharged for the
|
|
334
|
+
> pair `(i, j)` without enumerating a single case.
|
|
335
|
+
|
|
336
|
+
CO-S is not always achievable, but where it is, it is the discharge to write. The `dashes` /
|
|
337
|
+
`nbsp` pair is the worked example. `E(nbsp) = { U+00A0, U+202F }` — an **upper bound** as of spec
|
|
338
|
+
1.3.0, not an exact set: under `narrowNbsp: "nbsp"` (nbsp.md §3.1a) it shrinks to `{ U+00A0 }`.
|
|
339
|
+
Every discharge against it survives that unchanged, and by construction rather than by luck: each
|
|
340
|
+
one names **both** code points together, through a class that holds both (`spaces`'
|
|
341
|
+
`PROTECTED-SPACE`, `dashes`' and `ranges`' `NOBREAK-SPACE`, `quotes`' `INLINE-SPACE`,
|
|
342
|
+
`apostrophe`'s `SPACELIKE`), so a discharge that holds for the pair holds for either subset. A
|
|
343
|
+
future option that made `E(nbsp)` **larger** would not be free this way, and would have to be
|
|
344
|
+
re-argued here. Three successive attempts
|
|
345
|
+
tried to specify _which_ of those counted as dash spacing and _on which side_:
|
|
346
|
+
|
|
347
|
+
| Formulation | Fixed | Exposed |
|
|
348
|
+
| ------------------------------------------------ | ----------------------- | -------------------------------------------------------------- |
|
|
349
|
+
| a `NOBREAK-SPACE` counts as spacing | the original tight case | `«⍽-⍽»` — family 2 |
|
|
350
|
+
| …counts on the **left** only | family 2 | `«⍽–␣x` — defect (d), reachable from the pipeline's own output |
|
|
351
|
+
| **neither counts; a dash touching one is inert** | both, and the class | — |
|
|
352
|
+
|
|
353
|
+
Only the third satisfies CO-S, and it is also the shortest to state and the only one whose
|
|
354
|
+
correctness does not depend on which witnesses were in front of the author. A discharge that
|
|
355
|
+
has to be re-derived every time the other rule changes is a defect waiting for a new locale.
|
|
356
|
+
|
|
357
|
+
**When writing a CO discharge, try CO-S first.** If the earlier rule cannot be made inert to
|
|
358
|
+
the later rule's whole emission alphabet, say so explicitly and enumerate — but treat the
|
|
359
|
+
enumeration as a known-weak argument and make the sweep (§6) the real control.
|
|
360
|
+
|
|
361
|
+
### 5.2 Rule-local and pipeline-level claims
|
|
362
|
+
|
|
363
|
+
Every rule document has a §4 "Must not touch". Those lists were written rule-by-rule, and they
|
|
364
|
+
read — to anyone who has not memorised `order.json` — as promises about `transform`. Some of
|
|
365
|
+
them are not. A §4 bullet is really one of two different assertions:
|
|
366
|
+
|
|
367
|
+
- **[R] rule-local.** _"Given its input, this rule does not edit X."_ A statement about one
|
|
368
|
+
function. Always checkable from that rule's §3 alone.
|
|
369
|
+
- **[P] pipeline-level.** _"`transform` does not edit X."_ A statement about the composition,
|
|
370
|
+
and it is true only if **X survives every rule R₁…R₈** — not merely the one whose §4 it
|
|
371
|
+
appears in.
|
|
372
|
+
|
|
373
|
+
The gap between them is not academic. `ellipsis` §4 claimed `e.g. ..` was untouched; it is
|
|
374
|
+
untouched _by `ellipsis`_, but `spaces` deleted the U+0020 first, the run merged with the
|
|
375
|
+
abbreviation's dot, and the pipeline returned `e.g…`. Both rules were behaving exactly as
|
|
376
|
+
specified. The false statement was the scope of the claim, not the behaviour of either rule.
|
|
377
|
+
|
|
378
|
+
> **Derivation rule for a [P] claim.** X is protected end-to-end iff, for every rule `Rᵢ`, no
|
|
379
|
+
> sequence of edits by `R₁ … Rᵢ₋₁` can transform the input containing X into something `Rᵢ`
|
|
380
|
+
> would edit. In practice this reduces to one question per earlier rule: _can it change a code
|
|
381
|
+
> point adjacent to X, or delete one inside it?_ — because every rule in this pipeline decides
|
|
382
|
+
> from a bounded neighbourhood.
|
|
383
|
+
|
|
384
|
+
**Marking is mandatory.** Every §4 bullet in every rule document carries **[P]** or **[R]**.
|
|
385
|
+
A bullet describing a user-visible protection — paths, code-like text, identifiers, URLs,
|
|
386
|
+
abbreviations, measurements — must be **[P]**, because that is what a README reader will assume
|
|
387
|
+
it means; if the pipeline cannot deliver it, the defect is fixed rather than the claim
|
|
388
|
+
downgraded. **[R]** is reserved for statements that are genuinely about the rule's own
|
|
389
|
+
mechanics, and each one names the rule that can falsify it.
|
|
390
|
+
|
|
391
|
+
**Reading is not sufficient to find these.** Both instances above, and four more, were found by
|
|
392
|
+
_writing conformance fixtures_, not by reviewing prose — including by an agent that had read
|
|
393
|
+
all of these documents. A §4 bullet is a hypothesis until a fixture exercises it through the
|
|
394
|
+
whole pipeline. §6's sweep obligation is the mechanical half of this; per-bullet fixtures are
|
|
395
|
+
the other half, and they are what actually caught the class.
|
|
396
|
+
|
|
397
|
+
### 5.3 What this is worth, honestly
|
|
398
|
+
|
|
399
|
+
CO is a sufficient condition, not a necessary one. A pipeline could violate CO and still be
|
|
400
|
+
idempotent, if the work one rule creates for an earlier one happens to be undone again. Such a
|
|
401
|
+
pipeline would be idempotent by luck and would break on the next locale, so the spec requires
|
|
402
|
+
CO rather than the weaker property.
|
|
403
|
+
|
|
404
|
+
CO is also only as good as the invariant statements in §3. `I₁` is stated precisely because
|
|
405
|
+
`spaces` is simple; `I₃` is stated loosely ("no token `dashes` would edit") because `dashes`
|
|
406
|
+
is not. A loose invariant means the obligation is discharged by argument rather than by
|
|
407
|
+
mechanical check, and arguments about `dashes` have now been wrong three times. The
|
|
408
|
+
compensating control is the exhaustive sweep, not the prose — which is why §6 makes it a
|
|
409
|
+
release gate.
|
|
410
|
+
|
|
411
|
+
---
|
|
412
|
+
|
|
413
|
+
## 6. Testing obligation
|
|
414
|
+
|
|
415
|
+
The idempotency property test must include, in every runtime:
|
|
416
|
+
|
|
417
|
+
1. **A bounded exhaustive sweep, in two tiers.** Both are run for every locale in
|
|
418
|
+
`spec/locales/registry.json`, asserting `transform(transform(x)) == transform(x)`.
|
|
419
|
+
|
|
420
|
+
| tier | alphabet | length |
|
|
421
|
+
|---|---|---|
|
|
422
|
+
| **wide** | the consumed **and** emitted sets below | 0–5 |
|
|
423
|
+
| **deep** | a core of at least `1`, `-`, U+0020, `.`, U+2060 | 0–8 |
|
|
424
|
+
| **quotes** (spec 0.3.0) | `{ U+0022, U+0027, U+002D, U+0020, a }` ∪ every distinct quote glyph in `spec/locales/registry.json`'s locales | 0–6, per locale |
|
|
425
|
+
|
|
426
|
+
**Why the quotes tier exists, separately from wide/deep.** `quotes` (0.3.0)'s worst witness —
|
|
427
|
+
`‘""-"-"` in `en-GB` — is seven code points over `{ ‘ " - }`, and the committed deep tier's
|
|
428
|
+
core alphabet cannot reach it: it has neither the apostrophe-shaped opener nor a second
|
|
429
|
+
dash. Under mandate 1, `quotes.md` §6's **consumed** and **emitted** columns coincide for the
|
|
430
|
+
first time — every glyph the rule can produce is also now something it reads — so a tier keyed
|
|
431
|
+
to that single, self-consistent alphabet is the natural unit rather than folding it into wide
|
|
432
|
+
or deep, which are keyed to different rules' emission alphabets.
|
|
433
|
+
|
|
434
|
+
**Why two tiers, and why the bound moved.** A single length-4 bound was normative until
|
|
435
|
+
`dashes` defect (e) — witness `1-1 - 1`, **seven characters** — shipped underneath it. Nothing
|
|
436
|
+
was excluded and no precondition hid it; the bound simply could not reach it, and the suite
|
|
437
|
+
was green throughout. **A bound that cannot reach a known witness will hide the next one.**
|
|
438
|
+
The wide tier catches interactions between many characters, and grows combinatorially, so it
|
|
439
|
+
stays shallow; the deep tier is narrow enough (5 symbols to length 8 is 390 625 strings per
|
|
440
|
+
locale) to run to a length where multi-token shapes actually appear. U+2060 is in the deep
|
|
441
|
+
alphabet because `dashes` emits it and the guards must be inert to it (§5.1a).
|
|
442
|
+
|
|
443
|
+
**Standing obligation:** when a defect is found whose witness the committed bound cannot
|
|
444
|
+
reach, the bound is wrong. Raise it, or add the witness's alphabet to the deep tier, in the
|
|
445
|
+
same change that fixes the defect. **A biased or exhaustive generator is not
|
|
446
|
+
optional**: uniform random strings over a large alphabet essentially never produce `.--.`,
|
|
447
|
+
and the defect it hides may be the important one.
|
|
448
|
+
|
|
449
|
+
**The sweep alphabet must contain the code points the rules _emit_, not only those they
|
|
450
|
+
consume.** At minimum:
|
|
451
|
+
|
|
452
|
+
| | |
|
|
453
|
+
| ------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
454
|
+
| **consumed** | `"` `'` `-` U+0020 `.` `1` `a`, plus U+2010 (hyphen) and U+2212 (minus sign) — spec 0.2.0 added both to `dashes`' `DASH` class (`dashes.md` §3.1) |
|
|
455
|
+
| **emitted** | every locale `quotes.*.open`/`close` glyph in the registry — at minimum `«` `»` `“` `”` `„` `‘` `’` — plus U+2013, U+2014, U+2026, U+00A0, U+202F, U+2011 |
|
|
456
|
+
|
|
457
|
+
This is a normative requirement and it was learned the expensive way. **Idempotency is a
|
|
458
|
+
statement about re-processing output**, so an alphabet drawn only from input characters
|
|
459
|
+
tests the wrong language: it can only reach the subset of the output space that happens to
|
|
460
|
+
be spelled with input characters. Defect (d) of `dashes.md` §5.2 — `«⍽–␣"`, a shape built
|
|
461
|
+
entirely from characters the pipeline itself emits — was unreachable until the emitted
|
|
462
|
+
alphabet was added, and before that it surfaced on roughly **one property-test seed in
|
|
463
|
+
three**, where it read as flaky infrastructure rather than as a defect. A test that fails
|
|
464
|
+
one run in three and passes the rest is worse than no test, because it trains everyone
|
|
465
|
+
looking at it to re-run.
|
|
466
|
+
|
|
467
|
+
Two consequences for whoever maintains the sweep: the alphabet **grows when a locale is
|
|
468
|
+
added**, since a new locale's quote glyphs are new output characters; and the sweep must be
|
|
469
|
+
**deterministic and exhaustive** rather than sampled, so that a failure is reproducible from
|
|
470
|
+
the seed-free description alone.
|
|
471
|
+
|
|
472
|
+
**Named witnesses pinned as individual conformance fixtures, not only as sweep coverage**
|
|
473
|
+
(spec 0.3.0): `‘""-"-"` (`en-GB`), `'"‘` (`ru`), `«␣**"` (`fr`), `"<p class="x">«[t](u)`
|
|
474
|
+
(`fr-CA`, `html`), `" --x"` (`en-US`), `" "`, and `«»`. Each exercises the `quotes` gate or the
|
|
475
|
+
`I₃` landing guard on a shape the general sweep would eventually reach but that a reviewer
|
|
476
|
+
should be able to find by name — see `spec/rules/quotes.md` §6.
|
|
477
|
+
|
|
478
|
+
Two more, found by running the implementation rather than specified by either design
|
|
479
|
+
document: `«"<p class="x">"` (`fr`, `html`) and `««”` (`fr`, `text`), pinned in
|
|
480
|
+
`spec/fixtures/fr.json` as `fr-quotes-030-cobug-html-tag-boundary` and
|
|
481
|
+
`fr-quotes-030-cobug-double-open-then-closer`. Both exercise `quotes.md` §3.2's V1
|
|
482
|
+
`gapInsertable` clause — without it, `nbsp` inserting a space at a position V1's literal
|
|
483
|
+
adjacency check depended on lets a pairing the gate declined on one pipeline pass certify on
|
|
484
|
+
the next, which is a CO violation (§2) rather than an ordinary idempotency defect.
|
|
485
|
+
|
|
486
|
+
2. **A per-rule sweep.** The same assertion with a single rule enabled, which localises a
|
|
487
|
+
failure to a rule rather than to the composition.
|
|
488
|
+
3. **A composition sweep.** The same assertion with the full pipeline. A failure here that
|
|
489
|
+
passes (2) is a CO violation, and the report should say so — that is the signal this
|
|
490
|
+
document exists to make legible.
|
|
491
|
+
|
|
492
|
+
4. **A test whose carrier dies keeps passing, and proves nothing.** When a rule narrows, some
|
|
493
|
+
assertions elsewhere stop exercising what they were written for while continuing to go green
|
|
494
|
+
— they pass *for a different reason than when they were written*. That is strictly worse than
|
|
495
|
+
a failing test, because nothing draws attention to it.
|
|
496
|
+
|
|
497
|
+
This is not hypothetical: `dashes` §3.2 step 2a stopped the rule restyling an authored dash,
|
|
498
|
+
and every test and worked example carried on `–` → `␣–␣` became inert overnight. One of them
|
|
499
|
+
was in `modes.md` §3.4's normative table, where it had been the illustration of the whole
|
|
500
|
+
edge-growth rule. It was caught only because a *sibling* assertion in the same test failed and
|
|
501
|
+
someone asked why the other two still passed.
|
|
502
|
+
|
|
503
|
+
**Obligation when a rule narrows:** audit every assertion that uses the narrowed construct as
|
|
504
|
+
its carrier, and for each one ask whether it still fails when the behaviour it names is broken.
|
|
505
|
+
Rebuild it on a live carrier or delete it. The formulation worth keeping is the implementer's:
|
|
506
|
+
**the carrier is dead, the assertion is alive, the meaning is lost.**
|
|
507
|
+
|
|
508
|
+
Two related checks belong to the same discipline:
|
|
509
|
+
|
|
510
|
+
- **A round-trip fixture must have teeth.** A document asserted to come back byte-identical
|
|
511
|
+
proves something about skip lists only if it *would* change otherwise. If it is also
|
|
512
|
+
unchanged in `text` mode, it proves nothing at all — the assertion holds for the trivial
|
|
513
|
+
reason. Every such fixture should be verified to change materially in `text` mode.
|
|
514
|
+
- **A guard that has become unreachable is deleted, not left to rot** (§5.1a) — the same
|
|
515
|
+
principle one layer down, applied to the rule instead of the test.
|
|
516
|
+
|
|
517
|
+
**Preconditions (`fc.pre`, `assume`, and equivalents) that exclude a failing shape are not
|
|
518
|
+
permitted in the committed suite** except as a temporary, named, individually-pinned
|
|
519
|
+
containment for a defect that is already reported. Every such precondition must name the
|
|
520
|
+
document section that will remove it.
|
|
521
|
+
|
|
522
|
+
---
|
|
523
|
+
|
|
524
|
+
## 6a. The Unicode pin, and what it can portably require
|
|
525
|
+
|
|
526
|
+
`spec/UNICODE` contains `17.0`. Until this section it had **no reader**: no rule document cited
|
|
527
|
+
it, so an implementation could derive its character tables from any UCD version and still pass
|
|
528
|
+
conformance. That is the defect — not the file's absence, but its inertness.
|
|
529
|
+
|
|
530
|
+
### 6a.1 What the pin means
|
|
531
|
+
|
|
532
|
+
Two claims get conflated, and only one of them is portable.
|
|
533
|
+
|
|
534
|
+
1. **"The embedded tables are those derived from UCD 17.0."** Enforceable in every runtime,
|
|
535
|
+
determines output, and is what the pin means. Each port generates its `LETTER` (L ∪ M),
|
|
536
|
+
`UPPER` (Lu ∪ Lt) and simple-uppercase tables from the pinned UCD and checks them in.
|
|
537
|
+
2. **"The host runtime's UCD is 17.0."** **Not portable, and therefore not required.** JS exposes
|
|
538
|
+
`process.versions.unicode`, Python `unicodedata.unidata_version`, Go `unicode.Version` — but
|
|
539
|
+
PHP only through `intl`/ICU, which is not always installed, and Ruby has no dependable public
|
|
540
|
+
accessor. A normative requirement of this form would be unimplementable in at least one target
|
|
541
|
+
runtime, which is what ARCHITECTURE.md §4 exists to prevent.
|
|
542
|
+
|
|
543
|
+
> **`spec/UNICODE` is normative for the derived tables, not for the host runtime.** Every rule
|
|
544
|
+
> document that uses `LETTER`, `UPPER` or a case mapping must cite it. A port whose host UCD
|
|
545
|
+
> version *is* readable should additionally assert it against the pin as a cheap extra gate;
|
|
546
|
+
> where it is not readable that gate does not exist, and the spec must not pretend otherwise.
|
|
547
|
+
|
|
548
|
+
A host-side drift detector proves that the table matches *some* UCD and that this host agrees
|
|
549
|
+
with it. **Only fixtures can prove that five runtimes agree with each other.**
|
|
550
|
+
|
|
551
|
+
### 6a.2 Canary fixtures
|
|
552
|
+
|
|
553
|
+
Ordinary conformance cases whose expected output depends on the tables being right, so that table
|
|
554
|
+
drift surfaces in the conformance run rather than in one runtime's unit tests.
|
|
555
|
+
|
|
556
|
+
**On version canaries, honestly.** I cannot name a code point whose general category or simple
|
|
557
|
+
uppercase mapping demonstrably changed between UCD 16.0 and 17.0 from anything I can read here.
|
|
558
|
+
Naming one on recollection would be worse than naming none: **a canary that cannot fire advertises
|
|
559
|
+
a guarantee nobody has.** The version-drift canary must therefore be **generated, not hand-picked**
|
|
560
|
+
— a build step diffs the pinned UCD against the previous major version, takes the code points whose
|
|
561
|
+
`General_Category` or `Simple_Uppercase_Mapping` differs, and emits fixture rows asserting the
|
|
562
|
+
pinned behaviour. Generated rows fire by construction, and the set is re-derived whenever the pin
|
|
563
|
+
moves. If the diff is empty for a given version pair, the suite should say so rather than silently
|
|
564
|
+
contain nothing.
|
|
565
|
+
|
|
566
|
+
**Derivation canaries, which I can specify with confidence and which catch the likelier failure.**
|
|
567
|
+
The realistic defect is not "a port used UCD 16.0"; it is **"a port called the host's letter
|
|
568
|
+
predicate instead of the embedded table"**. Three cases, each chosen because a naive host call
|
|
569
|
+
gives a *different answer* from this spec's definition, in every Unicode version:
|
|
570
|
+
|
|
571
|
+
| Canary | Code point | Fires when |
|
|
572
|
+
|---|---|---|
|
|
573
|
+
| **`LETTER` includes marks** | U+0301 combining acute, category `Mn` | the port used `\p{L}`, `unicode.IsLetter`, `str.isalpha` or `ctype_alpha`, all of which exclude `Mn`. This spec's `LETTER` is **L ∪ M**, so a decomposed `é` (`e` + U+0301) must behave as one letter — `apostrophe.md` §6 case 2 already depends on it |
|
|
574
|
+
| **`UPPER` includes titlecase** | U+01C5 `Dž`, category `Lt` | the port used an `isUpper` that tests `Lu` only. This spec's `UPPER` is **Lu ∪ Lt** (`nbsp.md` §3.1) |
|
|
575
|
+
| **Case mapping is simple and locale-independent** | U+0069 `i` / U+0049 `I` | the port used a locale-sensitive uppercase. Under a Turkish host locale `i` maps to `İ` (U+0130), so `nbsp`'s first-character leniency (§3.5) stops matching a capitalised short word. This is ARCHITECTURE.md §4.4's dotless-ı hazard, made testable |
|
|
576
|
+
|
|
577
|
+
These prove the tables were **derived correctly**; the generated rows prove they were derived from
|
|
578
|
+
the **pinned version**. Both are needed, and only the second depends on data I cannot supply here.
|
|
579
|
+
|
|
580
|
+
---
|
|
581
|
+
|
|
582
|
+
## 7. Open questions
|
|
583
|
+
|
|
584
|
+
1. **`I₃` is not stated precisely enough to check mechanically.** "No token `dashes` would
|
|
585
|
+
edit" is a restatement of the rule, not an invariant. A checkable form would enumerate the
|
|
586
|
+
shapes — and the fact that I cannot write it in half a page is itself evidence that
|
|
587
|
+
`dashes` is the most complex rule in the spec and the one most likely to break again.
|
|
588
|
+
2. **CO has not been verified for rule subsets.** The `rules` option lets a caller disable any
|
|
589
|
+
rule. Disabling `spaces` removes `I₁` from the obligation set, which is harmless; but
|
|
590
|
+
disabling a rule can also _expose_ text to a later rule that the disabled rule would have
|
|
591
|
+
normalised, and no sweep currently runs over subsets. The combinatorics are mild (2⁸ = 256
|
|
592
|
+
configurations × the length-4 sweep) and this should probably be a CI job.
|
|
593
|
+
3. _(Settled.)_ Modes are now covered by [modes.md](modes.md), which was written to answer
|
|
594
|
+
this item. The premise it was raised under turned out to be the wrong one: the pipeline does
|
|
595
|
+
**not** run per span. Spans are concatenated with an explicit boundary marker and the
|
|
596
|
+
pipeline runs once over the whole array, so a space at a span boundary and a dash in the
|
|
597
|
+
neighbouring span are seen together — and are declined together, because the marker breaks
|
|
598
|
+
the adjacency that would make them a token. `modes.md` §5 adds one obligation this document
|
|
599
|
+
cannot see: **the span partition must be stable between runs**, which constrains what code
|
|
600
|
+
points a rule may emit.
|
|
601
|
+
4. _(Settled, spec 0.3.0.)_ **The `nbsp`-before-`quotes` direction.** `quotes.md` §2.1's
|
|
602
|
+
constraint **Q-P** — every `nbsp.beforePunctuation`/`nbsp.narrowBeforePunctuation` entry must
|
|
603
|
+
be a member of `quotes`' `CLOSEISH` — is enforced by `scripts/validate-spec.mjs` and is what
|
|
604
|
+
makes `quotes.md` §5's Lemma B cover `nbsp`'s N1/N2 insertion sites as well as N8, by proof
|
|
605
|
+
rather than by case analysis over the shipped data alone. All ten shipped locales satisfy it.
|