polytypo 1.2.0 → 1.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +648 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,1324 @@
1
+ # Rule: `quotes`
2
+
3
+ **Order:** 40. **Default:** on. **Modes:** text, html, markdown, yaml.
4
+ **Spec version:** 1.1.0 (0.4.1 for everything except the universal medial-`n` elision veto
5
+ described in §3.2 and the History section below).
6
+
7
+ ---
8
+
9
+ ## 0. What changed from 0.1.0, and why it is a rewrite rather than a patch
10
+
11
+ 0.1.0's whole architecture rested on one sentence: *the rule reasons about the straight input
12
+ marks and never about the curly output glyphs.* Two operator decisions withdraw that sentence:
13
+
14
+ 1. **Any existing quote glyph is a re-typesetting candidate**, this locale's own or a foreign
15
+ one. `«Wort»` in German prose becomes `„Wort“`. Same principle as `dashes` §3.2 step 2a
16
+ applies to dash length: a quote glyph's identity in ordinary prose is at least as often a
17
+ copy-paste artefact or unfamiliarity as a deliberate choice, and an author who wants their own
18
+ typography untouched has always had the option of not running the pipeline. **0.1.0 §3.8 is
19
+ withdrawn.**
20
+ 2. **A space touching a quote mark is sloppiness, not evidence.** `" hello"` is `“hello”`; the
21
+ space is deleted. `« bonjour »` is a formed pair on the first application, not only after
22
+ `nbsp` has touched it.
23
+
24
+ Once every quote glyph is a candidate, **the rule's input alphabet equals its output alphabet**,
25
+ and 0.1.0's idempotency proof — whose first line is "a converted position is not a candidate in
26
+ pass 1 of the second run" — is not weakened, it is **false**. Three ideas replace it, and each
27
+ closes one class of the four defects that motivated this revision:
28
+
29
+ | Idea | Closes |
30
+ | --- | --- |
31
+ | **Glyph-blindness (Lemma A, §5).** Every quote mark belongs to `QUOTEMARK`, which is a member of *both* `OPENISH` and `CLOSEISH` and exempt from `canOpen`'s closeish rejection. A candidate's verdict never depends on *which* quote glyph its neighbour is. | neighbour drift; `apostrophe` (R₆) becomes structurally invisible as a neighbour |
32
+ | **Directional space-skipping (Lemma B, §5).** `canOpen` skips a space on its inner (right) side, `canClose` on its inner (left) side, unconditionally; each reads its *outer* side literally, except at the two positions `nbsp` can insert — derived from locale data, not hand-picked. | mandate 2; a run-1/run-2 asymmetry across an `nbsp`-inserted space; an HTML span boundary the same shape exposed |
33
+ | **Certification (the gate, §3.5).** The accepted pairing is *checked*, not proved: render the hypothetical output, re-run passes 1–2 on it, and decline pairs until the re-run reproduces the accepted set exactly. | width drift interacting with an unmatched candidate — the one remaining non-invariance, and provably non-local |
34
+
35
+ The division of labour is the point. **Lemmas A and B cover what the gate cannot see** — edits
36
+ made by *other* rules after `quotes` has run. **The gate covers everything `quotes` itself
37
+ does** — instabilities in the vetoes, the width partition, the depth assignment — without
38
+ anyone having to enumerate them by hand. A future change to the capability tests can break
39
+ quality; it cannot break idempotency, because the gate checks the actual re-derivation rather
40
+ than trusting an argument about it.
41
+
42
+ ---
43
+
44
+ ## 1. Purpose
45
+
46
+ `quotes` resolves every quotation mark in the text — typewriter (U+0022, U+0027) or typographic
47
+ (`«`, `“`, `„`, `‘`, `”`, …), the locale's own or a foreign locale's — into the pair of glyphs
48
+ the locale prescribes, alternating between the primary and secondary pairs by nesting depth, and
49
+ normalising the spacing immediately inside each pair.
50
+
51
+ ---
52
+
53
+ ## 2. Locale data consumed
54
+
55
+ - `quotes.primary.open` / `.close` / `.innerSpace`
56
+ - `quotes.secondary.open` / `.close` / `.innerSpace`
57
+ - `quotes.elisionIdioms` (spec 0.4.0) — an array of `{ left, elided, right }` triples, each a
58
+ literal string, consulted only by the **listed elision veto** (§3.2) to **decline** pairing
59
+ two `NARROW` marks as a quotation when the full context matches: the elided content between
60
+ the marks, the word immediately before the opening mark, and the word immediately after the
61
+ closing mark (`rock 'n' roll`'s `{ left: "rock", elided: "n", right: "roll" }`). May be
62
+ empty, and empty is the default for a locale with no verified idiom of this shape. A
63
+ decline-only list can never widen what this rule pairs, only narrow it — the same principle
64
+ stated normatively in `nbsp.md` §2.1 and governing every list-valued field in
65
+ `locale.schema.json`.
66
+
67
+ **All three fields are required, and matching `elided` alone is deliberately not a supported
68
+ shape.** A bare word list cannot distinguish an idiom from an arbitrary quotation of the same
69
+ word — `The letter 'n' is common.` and `He said "press 'n' now".` are both genuine
70
+ quotations of the letter *n*, and a list keyed on `elided` alone would falsely elide both.
71
+ The surrounding context is the evidence that makes `rock 'n' roll` an idiom and these two
72
+ sentences ordinary prose; §3.2 states exactly how much of it must match.
73
+
74
+ **`left` and `right` must consist entirely of `LETTER` code points** — §3.2's Word definition
75
+ is a maximal `LETTER` run, and a `left`/`right` entry is normatively *authored as* that word.
76
+ This is not a limit the matcher itself enforces at match time: the comparison in §3.2 is a
77
+ literal code-point equality test, so a configured entry containing a digit or punctuation
78
+ code point would still match those exact bytes wherever they occur, not fail to match at all.
79
+ The constraint exists precisely because the matcher does not check it — an unvalidated entry
80
+ could make the veto fire on text no citation ever attested, silently exceeding what the
81
+ normative Word definition authorises. `scripts/validate-spec.mjs` and `scripts/gen-locales.mjs`
82
+ therefore reject such an entry at the data boundary, before it can be authored into a locale
83
+ file at all. `elided` carries no such constraint: its match is a literal equality test over
84
+ the raw code points between the two marks, not a `LETTER`-run walk, so it may contain any
85
+ non-empty literal. Both validators read the identical pinned table `src/engine/unicode.ts`
86
+ imports (`scripts/lib/is-letter.mjs`), so validation and the runtime engine cannot classify a
87
+ code point differently.
88
+
89
+ `innerSpace` is read and acted on in **one direction only**: `quotes` *deletes* an inner space
90
+ when the assigned pair's `innerSpace` is `"none"`, and never inserts or converts one. Insertion
91
+ and conversion stay with `nbsp` (order 70), exactly as 0.1.0 always said. See §3.7.
92
+
93
+ ### 2.1 Two new normative locale constraints
94
+
95
+ Both are needed by §5's proof. **All ten shipped locales already satisfy both**, verified against
96
+ `spec/locales/*.json`:
97
+
98
+ - **Q-W — a pair is width-homogeneous.** For each of `primary` and `secondary`, `open` and
99
+ `close` must have the same quote *width* (§3.1). Without it a formed pair straddles both
100
+ stacks after rendering and the gate would decline every pair in a same-glyph document.
101
+ - **Q-A — a spaced pair must not use the apostrophe glyph.** If `innerSpace ≠ "none"`, neither
102
+ `open` nor `close` may be U+2019. Without it `apostrophe`'s output would land in a
103
+ `nbsp`-insertion-adjacent skip set and Lemma A's corollary would break.
104
+
105
+ One constraint belongs to `nbsp`'s data and is stated normatively in `nbsp.md` §2:
106
+
107
+ - **Q-P — every code point in `nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must
108
+ be a member of this rule's `CLOSEISH`.** This is what makes Lemma B cover N1/N2 as well as N8.
109
+ The shipped lists (`:`, `;`, `!`, `?` in `fr`/`fr-CA`, empty elsewhere) satisfy it.
110
+
111
+ **Evidence for locale data.** A locale attests **membership**; the rule owns the **mechanism**.
112
+ Stated normatively in `nbsp.md` §2.1; it governs every locale field this rule reads.
113
+
114
+ ---
115
+
116
+ ## 3. Algorithm
117
+
118
+ Input is a code-point array `cp[0 … n-1]`. Five passes plus an emit; no backtracking inside a
119
+ pass, no regular expression, no native-string indexing.
120
+
121
+ ### 3.1 Character classes
122
+
123
+ | Class | Members |
124
+ | --- | --- |
125
+ | `DQ` | U+0022 |
126
+ | `SQ` | U+0027 |
127
+ | `STRAIGHT` | `DQ` ∪ `SQ` |
128
+ | `WIDE` | U+0022, U+00AB, U+00BB, U+201C, U+201D, U+201E, U+201F, U+301D, U+301E, U+301F |
129
+ | `NARROW` | U+0027, U+2018, U+2019, U+201A, U+201B, U+2039, U+203A |
130
+ | `QUOTEMARK` | `WIDE` ∪ `NARROW` (disjoint) |
131
+ | `DIGIT` | U+0030–U+0039 |
132
+ | `LETTER` | general categories `Lu Ll Lt Lm Lo Mn Mc Me` |
133
+ | `ALNUM` | `LETTER` ∪ `DIGIT` |
134
+ | `BREAK` | U+000A, U+000D, U+000B, U+000C, U+0085, U+2028, U+2029, plus `LINE_MARKER` |
135
+ | `INLINE-SPACE` | U+0020, U+0009, U+00A0, U+202F, U+2007, U+2009, U+200A |
136
+ | `SPACELIKE` | `INLINE-SPACE` ∪ `BREAK` |
137
+ | `OPENISH` | U+0028 `(` U+005B `[` U+007B `{`, `MARKER`, ∪ `QUOTEMARK` |
138
+ | `CLOSEISH` | U+0029, U+005D, U+007D, U+002C, U+002E, U+003B, U+003A, U+0021, U+003F, U+2026, U+2013, U+2014, `MARKER`, ∪ `QUOTEMARK` |
139
+ | `DASHISH` | U+002D, U+2011, U+2013, U+2014 |
140
+ | `DELETE-LANDING` | `ALNUM` ∪ `QUOTEMARK` |
141
+ | `NONE` | the pseudo-class "index out of range" |
142
+
143
+ **`QUOTEMARK` is a member of both `OPENISH` and `CLOSEISH`, and is exempt from `canOpen`'s
144
+ closeish rejection** — the same dual membership `STRAIGHT` had in 0.1.0, widened to the whole
145
+ class. This is Lemma A's entire mechanism and must not be simplified back to per-glyph lists.
146
+ `MARKER` keeps its 0.1.0 treatment (`modes.md` §3.3): in both classes, exempt from the
147
+ rejection, and **not** in `SPACELIKE`, so no skip walk ever crosses a span boundary.
148
+
149
+ **Deliberately excluded from `QUOTEMARK`:** U+2032/U+2033 (primes — a separate rule with its own
150
+ false-positive profile, §7), U+02BC (a `Lm` letter), U+0060 and U+00B4. Neither candidates nor
151
+ produced.
152
+
153
+ **Unicode version.** As 0.1.0: the pinned UCD in `spec/UNICODE` is normative for the derived
154
+ tables, not for the host runtime (`pipeline-idempotency.md` §6a).
155
+
156
+ **A declared quote glyph must not be in `SPACELIKE`, `ALNUM` or `STRAIGHT`**, and by Q-W must
157
+ share its pair's `WIDE`/`NARROW` bucket with its partner glyph. `singleChar` in
158
+ `locale.schema.json` enforces the `STRAIGHT` exclusion.
159
+
160
+ ### 3.1a Locale-derived skip sets
161
+
162
+ Computed once per call, from locale data alone:
163
+
164
+ ```
165
+ SPACE-RIGHT = { p.open : p ∈ {primary, secondary}, p.innerSpace ≠ "none", p.open ≠ p.close }
166
+ SPACE-LEFT = { p.close : p ∈ {primary, secondary}, p.innerSpace ≠ "none", p.open ≠ p.close }
167
+ ```
168
+
169
+ These are **exactly the positions at which `nbsp` N8 can insert a space** (`nbsp.md` §3.10;
170
+ `nbsp.ts` `quotesSubRule` skips a pair whose `open` equals its `close`, and skips
171
+ `innerSpace: "none"`). For the ten shipped locales: `SPACE-RIGHT = {«}` and `SPACE-LEFT = {»}`
172
+ in `fr` and `fr-CA`; both empty everywhere else. The sets are derived from the *reason*, not
173
+ hand-written — if a locale gains a spaced pair, they follow automatically and Lemma B keeps
174
+ holding.
175
+
176
+ ### 3.2 Pass 1 — collect and classify candidates
177
+
178
+ Walk `i` from `0` to `n-1`. Skip unless `cp[i] ∈ QUOTEMARK`. Write `g = cp[i]` and compute four
179
+ neighbour reads:
180
+
181
+ - `Llit = cp[i-1]`, or `NONE`; `Rlit = cp[i+1]`, or `NONE`;
182
+ - `Lskip` = the first `cp[j]`, `j < i` descending, with `cp[j] ∉ INLINE-SPACE`; `NONE` if the
183
+ walk leaves the array;
184
+ - `Rskip` = symmetric to the right.
185
+
186
+ `MARKER` and every `BREAK` stop a skip walk, because neither is in `INLINE-SPACE` — a skip never
187
+ crosses a span boundary or a line terminator.
188
+
189
+ Then two *directed* reads:
190
+
191
+ - `openLeft` = `Lskip` if `g ∈ SPACE-LEFT`, else `Llit`;
192
+ - `closeRight` = `Rskip` if `g ∈ SPACE-RIGHT`, else `Rlit`.
193
+
194
+ The two capabilities:
195
+
196
+ ```
197
+ canOpen ⟺ ( openLeft = NONE ∨ openLeft ∈ SPACELIKE ∨ openLeft ∈ OPENISH ∨ openLeft ∈ DASHISH )
198
+ ∧ ( Rskip ≠ NONE ∧ Rskip ∉ SPACELIKE
199
+ ∧ ( Rskip ∉ CLOSEISH ∨ Rskip ∈ QUOTEMARK ∨ Rskip = MARKER ) )
200
+
201
+ canClose ⟺ ( Lskip ≠ NONE ∧ Lskip ∉ SPACELIKE )
202
+ ∧ ( closeRight = NONE ∨ closeRight ∈ SPACELIKE ∨ closeRight ∈ CLOSEISH ∨ closeRight ∈ DASHISH )
203
+ ```
204
+
205
+ `Rskip`/`Lskip` can still be a `BREAK` (a skip walk stops there), which is why both
206
+ `∉ SPACELIKE` tests are retained: a mark at a line end does not open across the break, a mark at
207
+ a line start does not close back over it.
208
+
209
+ **Why this exact asymmetry.**
210
+
211
+ - `canOpen` skips right and `canClose` skips left: each capability treats its *inner* side — the
212
+ side facing the quoted text under that hypothesis — as noise. That is mandate 2, stated once
213
+ and applied uniformly, unconditionally, on every locale.
214
+ - Each capability reads its *outer* side literally. The outer space is not noise, it is the
215
+ evidence. `"hi" to me` is a closing mark **because** a space follows it; skipping that space
216
+ (reading `t`) would destroy the commonest closing shape in every language. `He said "hi. She
217
+ said "bye."` is preserved by exactly this: the first mark's `closeRight` is the letter `h`, so
218
+ it is `canOpen` only, and the stray mark cannot swallow the real quotation.
219
+ - The two exceptions to "outer side is literal" are *derived*, not chosen: `nbsp` can insert on
220
+ the right of a `SPACE-RIGHT` glyph and the left of a `SPACE-LEFT` glyph, and those are the only
221
+ two outer reads it can reach. Making exactly those two reads skip is what makes every verdict
222
+ inert to `nbsp` (Lemma B).
223
+
224
+ **Medial-elision veto** (`NARROW` marks only, literal reads):
225
+
226
+ > if `g ∈ NARROW` and `Llit ∈ ALNUM` and `Rlit ∈ ALNUM`, set both capabilities false.
227
+
228
+ `don't`, `l'été`, `O'Brien`, `1990's` — and, on a second pipeline pass, `don't` with U+2019,
229
+ because `apostrophe` has converted the mark and U+2019 is also `NARROW`. Widening from `SQ` to
230
+ `NARROW` is what makes `apostrophe`'s output inert here.
231
+
232
+ **Listed elision veto** (spec 0.4.0, `NARROW` marks only, locale data `quotes.elisionIdioms`,
233
+ literal reads). **Matching the elided content alone is not this veto's shape.** A bare word
234
+ list cannot distinguish `rock 'n' roll` from an arbitrary quotation of the same word —
235
+ `The letter 'n' is common.` and `He said "press 'n' now".` are both genuine quotations that a
236
+ content-only match would falsely elide (this was tried and rejected in an earlier revision of
237
+ this field; both sentences are pinned as negative fixtures, `spec/fixtures/en-US.json`). The
238
+ veto therefore requires the **full context**: the elided content, and a literal word on each
239
+ outer side of the two marks.
240
+
241
+ > **Word.** A maximal run of `LETTER` code points. Its outer boundary — the code point
242
+ > immediately before its first code point, or immediately after its last — is either `NONE`
243
+ > (document or span start/end) or **not** `LETTER` and **not** `DIGIT`. This is a whole-word
244
+ > test: `"MyRock 'n' roll"` does not match a `left` of `"Rock"`, because the code point before
245
+ > `R` is `y`, a `LETTER`.
246
+ >
247
+ > **The permitted space.** Exactly one code point of `INLINE-SPACE` (`spec 3.1`) separates a
248
+ > context word from its adjacent mark — not `NONE`, not `OPENISH`, not `DASHISH`, unlike
249
+ > `canOpen`/`canClose`'s own outer tests above. The veto needs an actual word to anchor to, so
250
+ > a mark at document or span start, or immediately after an opening bracket, has no `left` word
251
+ > and cannot be vetoed on that side — §3.2's `MARKER` (`modes.md` §3.3) is not `INLINE-SPACE`
252
+ > either, so a context word split from a mark by an element or span boundary does not match
253
+ > (`<em>rock</em> 'n' roll` is unaffected; §6 rows H1–H3).
254
+ >
255
+ > Computed once, before capabilities are assigned, over every index `i` with `cp[i] ∈ NARROW`:
256
+ > for each entry `{ left, elided, right } ∈ quotes.elisionIdioms`, of `k` code points in
257
+ > `elided`: if `cp[i+1 … i+k]` equals `elided` **exactly**, code point for code point — no case
258
+ > leniency, on the elided content — and `cp[i+k+1] ∈ NARROW` (call this index `j`), and:
259
+ >
260
+ > - `cp[i-1]` is a single `INLINE-SPACE` code point, and the word ending there equals `left`;
261
+ > - `cp[j+1]` is a single `INLINE-SPACE` code point, and the word starting after it equals `right`;
262
+ >
263
+ > then both `i` and `j` are **vetoed**: both capabilities of `i` and both capabilities of `j`
264
+ > are set to `false`, overriding every other test in this section — including capabilities
265
+ > computed for `j` when the main walk reaches it.
266
+ >
267
+ > **Exact matching, with one narrow exception.** `elided` is always matched exactly, code point
268
+ > for code point — no leniency, ever, on the elided content itself. `left` and `right` are
269
+ > matched exactly **except their first code point**, which is compared case-insensitively —
270
+ > implemented directly by code point (ASCII `A`–`Z` folds to `a`–`z`; every other code point,
271
+ > including every non-ASCII letter, must match exactly), never through a platform locale
272
+ > function (`ARCHITECTURE.md` §4.4). This is the same leniency `nbsp.md` §3.5's
273
+ > `afterShortWords` already uses, for the same reason: `Rock 'n' roll is great.` (sentence-
274
+ > initial capital) is exactly as much the idiom as `rock 'n' roll` is, and `rock 'N' roll` —
275
+ > case varying anywhere past the first code point — is not: it fails the exact `elided` test in
276
+ > this specific idiom (`elided = "n"`, and `N ≠ n`) and is correctly left alone (§6 row N5).
277
+
278
+ `rock 'n' roll` with `elisionIdioms = [{ left: "rock", elided: "n", right: "roll" }]`: the
279
+ leading mark's `left` word is `rock`, the elided run spells exactly `n`, the trailing mark's
280
+ `right` word is `roll` — all three match, both marks are vetoed, survive pass 1 as U+0027, and
281
+ `apostrophe` (order 50) renders each independently: leading elision (`apostrophe.md` §3.3 case
282
+ 4) on the first, trailing elision (case 3) on the second, giving `rock ’n’ roll`. Without the
283
+ veto both marks are `canOpen`/`canClose` candidates in their own right — the leading mark has
284
+ no `ALNUM` on its left so the medial veto above does not apply to it, and likewise for the
285
+ trailing mark's right — and pass 2 pairs them as an ordinary quotation.
286
+
287
+ **This is a bounded, literal check, not a heuristic.** The lookahead from `i` and `j` is
288
+ bounded by the total length of `left` + `elided` + `right` for each configured idiom — no
289
+ regex, no priorities, no open-ended word class; the same shape `nbsp.md`'s
290
+ `beforeNumber`/`beforeWord` literal matching already uses on one side of a mark
291
+ (`nbsp.md` §3.11), applied here to both sides of a pair of marks at once. It can only ever
292
+ **prevent** a pairing pass 2 would otherwise have formed — an empty `quotes.elisionIdioms`
293
+ makes it a total no-op, and every `NARROW` mark is classified exactly as it was before spec
294
+ 0.4.0.
295
+
296
+ **Behaviour at punctuation and document boundaries.** Trailing punctuation after `right` (a
297
+ period, a comma, a closing quotation mark) does not block the match: the word-boundary test
298
+ only requires the code point after `right`'s last letter to be non-`LETTER`/non-`DIGIT`, and
299
+ punctuation satisfies that trivially (`I love rock 'n' roll.` still vetoes; §6 row P3).
300
+
301
+ **A context word may itself begin at document/span start or end at document/span end — the
302
+ Word definition above already says so (`NONE` is an accepted outer boundary).** `Rock 'n' roll
303
+ is great.`, where `Rock` is the very first thing in the document, still vetoes (§6 row P2): the
304
+ mark's own left neighbour `cp[i-1]` is a real `INLINE-SPACE` code point (the space after
305
+ `Rock`), and `wordEndsAt` finds `Rock` ending there with a legal `NONE` boundary one step
306
+ further left. **What cannot happen is the mark itself sitting at that extremity.** A `NARROW`
307
+ mark at the very start or end of the text unit has `Llit`/`Rlit` = `NONE`, which is not
308
+ `INLINE-SPACE`, so its own outer test fails before a word is even sought — the veto needs an
309
+ actual space *and* an actual word beyond it, and a mark with nothing to its outer side has
310
+ neither. The constraint is on the mark's adjacency, not on where the word may fall.
311
+
312
+ **Membership evidence, not mechanism**, stated normatively in `nbsp.md` §2.1 and governing every
313
+ list-valued field in `locale.schema.json`: a locale attests that a `{ left, elided, right }`
314
+ triple belongs in this list; this rule owns the mechanism. Each triple is independently
315
+ evidenced — `{ left: "rock", elided: "n", right: "roll" }` is cited (§6, `en-US.json`
316
+ `sources`), and that citation supports no other triple: it does not, for instance, attest
317
+ `{ left: "fish", elided: "n", right: "chips" }`, which is not listed here and would need its
318
+ own citation before being added. This is deliberately **not** a general elision-word list, and
319
+ it must not grow into one merely because the mechanism could technically encode it. A locale
320
+ with no citable evidence for a triple of this exact shape ships an empty list rather than a
321
+ guessed one — the same discipline `PLAN.md` §6.1 states for every locale field, and the reason
322
+ this veto is populated for only one locale at launch (§6).
323
+
324
+ **Known, accepted residual ambiguity — not eliminated by the full-context requirement.** The
325
+ veto has no notion of authorial intent: it matches code points, not meaning, so a document that
326
+ explicitly states its own intent to write three separate tokens — the word `rock`, a genuinely
327
+ quoted letter `n`, and the word `roll` — and then, for illustration, reproduces the identical
328
+ surface sequence `rock 'n' roll`, still gets that final occurrence vetoed:
329
+
330
+ ```
331
+ The sequence is the word rock, the quoted letter 'n', and the word roll: rock 'n' roll.
332
+ ```
333
+
334
+ transforms to
335
+
336
+ ```
337
+ The sequence is the word rock, the quoted letter “n”, and the word roll: rock ’n’ roll.
338
+ ```
339
+
340
+ The **first** `'n'` is correctly left alone: its `left` context is `letter`, not `rock`, and it
341
+ is immediately followed by a comma rather than a space, so no configured idiom matches and it
342
+ pairs as an ordinary depth-1 quotation — exactly the mechanism §6 rows N1/N2 already prove. The
343
+ **second** occurrence — `rock 'n' roll` at the end — matches `left`/`elided`/`right` exactly,
344
+ regardless of the fact that the same sentence has just told the reader it is meant as a literal
345
+ concatenation of the three tokens just described, not a fresh invocation of the idiom. **Under
346
+ the authorial intent the sentence itself states, this conversion is a real semantic false
347
+ positive** — the final `'n'` is meant as a repetition of the same quoted letter named earlier in
348
+ the sentence, not the idiom, and eliding it changes what the author wrote to mean. It is
349
+ accepted as an **unavoidable** false positive for a bounded, surface-form contract, not a
350
+ tolerated one: `left`, `elided`, `right` and the single permitted `INLINE-SPACE` on each side
351
+ are the entire alphabet this veto is defined over, and none of them can encode "this is a
352
+ demonstration, not an utterance of the idiom" — no finite extension of that alphabet could,
353
+ short of the veto reasoning about meaning, which is out of scope for a literal, bounded
354
+ contract (§3.2). The implementation still **conforms exactly** to its own normative matcher on
355
+ this input: the two occurrences are given the treatment the matcher's rules specify byte for
356
+ byte, and the defect is that the matcher's contract cannot distinguish these bytes from the
357
+ idiom it exists to recognise — because, written this way, they are the same bytes. Recorded
358
+ here in the same spirit as `dashes.md` §7.9, and §6 row N6 pins it as a fixture, so the exposure
359
+ is visible in the conformance suite rather
360
+ than discovered in someone's content.
361
+
362
+ **Universal medial-`n` elision veto (spec 1.1.0).** The listed elision veto above closes
363
+ `rock 'n' roll` for `en-US`, where a citation exists. It closes nothing for `en-GB`, `de-DE`,
364
+ `de-CH`, `fr`, `fr-CA`, `ru`, `fi`, `sv` or `el`, whose `elisionIdioms` lists are empty. Without a
365
+ further mechanism the marks in `rock 'n' roll` fall through to ordinary pairing in all nine and
366
+ are typeset as a quotation.
367
+
368
+ **Operator decision, 2026-09-09: `rock 'n' roll` and `rock'n'roll` are international, and both
369
+ take U+2019 in every locale.** The idiom is not a locale fact — it is a fixed borrowed string that
370
+ appears untranslated in prose in all ten — so the mechanism is locale-blind and lives in this rule
371
+ rather than in locale data. Locale data could not carry it in any case: the `sources` discipline
372
+ (§2) would block an entry in a locale whose own orthography writes the borrowing out instead
373
+ (Russian writes рок-н-ролл), and an entry withheld for that reason would leave uncovered exactly
374
+ the locales this decision is about.
375
+
376
+ The predicate (`src/rules/quote-ambiguity.ts`, `computeAmbiguousShapeIndices`) is **this rule's
377
+ alone** as of spec 1.1.0. Under 0.5.0 `apostrophe` consulted the same module to decide what to skip;
378
+ it no longer skips anything, so the module has one caller and there is nothing left for the two
379
+ rules to drift apart on. The shape: a pair of `NARROW` marks enclosing **exactly one code point**,
380
+ either U+006E `n` or U+004E `N`, with **at least one** `INLINE-SPACE` code point immediately
381
+ outside each mark. `rock 'n' roll`, `fish 'n' chips` and `rock 'N' roll` all match;
382
+ `She chose 'A' today`, `They said 'no' yesterday` and `say 'yes' now` do not, and pair as ordinary
383
+ quotations.
384
+
385
+ The closed-up form `rock'n'roll` is not this veto's business at all — the medial-elision veto
386
+ above already declines both marks on `ALNUM` neighbours, and `apostrophe` converts them. It needs
387
+ no special case and gets none.
388
+
389
+ **Why `NARROW` and not straight ASCII only — an idempotency obligation, not a preference.** Spec
390
+ 0.5.0's predicate matched U+0027 alone. That was sound *there*, because its outcome was to
391
+ preserve the author's straight marks byte-identically, so a second pipeline pass saw the same
392
+ bytes and re-vetoed. This veto instead lets `apostrophe` **convert** both marks to U+2019, and a
393
+ straight-only predicate therefore would not recognise its own output: pass 2 would pair
394
+ `rock ’n’ roll` as an ordinary `NARROW` quotation on the next run. Measured, not argued — with a
395
+ straight-only predicate, `rock ’n’ roll` becomes `rock «n» roll` in `ru` and `rock ”n” roll` in
396
+ `fi`, an idempotency violation and so a release blocker (`pipeline-idempotency.md` §1). Matching
397
+ the whole `NARROW` class makes the converted form a fixed point in all ten locales.
398
+
399
+ **Why exactly one code point `n`, and not a letter-count shape.** 0.5.0's predicate matched 1 to 3
400
+ `LETTER` code points, which cannot tell the idiom from an ordinary short nested quotation and
401
+ declined both — `«это 'моё' дело»` in `ru`, `“He said 'no' to me,”` in `en-US`. Counting letters
402
+ is wrong here twice over. It is too broad: a genuine short quotation is far commoner in prose than
403
+ the idiom, so declining it is the larger error. And it is not normalization-stable — `LETTER`
404
+ includes `Mn` (§3.1) and this rule never normalizes its input (`ARCHITECTURE.md` §4.3), so `'моё'`
405
+ is three `LETTER` code points composed and four decomposed, and the same visible word took two
406
+ different branches depending on which form the author's editor happened to save. A comparison
407
+ against a fixed pair of code points has no such dependency.
408
+
409
+ **"At least one", deliberately, not "exactly one".** Only the single code point immediately
410
+ adjacent to each mark is tested; a longer run of inline spaces further out — `rock 'n' roll`,
411
+ doubled — does not invalidate the match. Narrowing this to "exactly one" was considered and
412
+ rejected: it would make a doubled-space variant of `rock 'n' roll` fall through to ordinary
413
+ quote-pairing, reintroducing a false-positive quotation conversion the veto exists to prevent, in
414
+ exchange for no compensating benefit. The existing listed-idiom matcher (above) may remain
415
+ stricter on this point without contradiction — with extra spaces its own literal `left`/`right`
416
+ word-boundary test can fail to match the cited en-US `rock`/`n`/`roll` tuple, and when it does,
417
+ this veto still wins and the marks are still converted.
418
+
419
+ **Every position this shape matches is added to the same veto set the listed-idiom check
420
+ populates**, and both mechanisms have the identical outcome: `quotes` declines the pairing, both
421
+ marks survive pass 2 unmatched, and `apostrophe` (order 50) renders each independently by its own
422
+ structural case ladder — leading elision (`apostrophe.md` §3.3 case 4) on the first, trailing
423
+ elision (case 3) on the second, giving `rock ’n’ roll`. `en-US`'s cited `elisionIdioms` entry now
424
+ vetoes a strict subset of what this veto vetoes; it is retained because the listed mechanism
425
+ matches multi-code-point `elided` content this one does not, and because the two can never
426
+ disagree about a position both cover. **Spec 0.5.0's preserve set is withdrawn**
427
+ (`computePreserveIndices`, `apostrophe.md` §3.4): it existed to stop `apostrophe`'s case ladder
428
+ from converting marks 0.5.0 wanted preserved, and conversion is now the specified outcome for
429
+ every position this veto matches, so `apostrophe` needs no knowledge of the veto at all and its
430
+ ordinary ladder does the right thing unaided.
431
+
432
+ **Already-curly input is matched, and preserved as typed.** A pair typed directly with U+2019 is
433
+ in `NARROW`, so it is vetoed like any other — and since `apostrophe` emits U+2019 only in place of
434
+ U+0027 (`apostrophe.md` §1), it edits nothing there: `rock ’n’ roll` is a fixed point, and
435
+ `rock ‘n’ roll` keeps the author's own U+2018 rather than being re-typeset as that locale's
436
+ quotation. This is the direct opposite of 0.5.0's rule for curly input, and the reason is the
437
+ idempotency obligation two paragraphs above, not a change of view about authorial intent.
438
+
439
+ **Accepted false positive, recorded rather than tolerated.** `The letter 'n' is common.` becomes
440
+ `The letter ’n’ is common.` — a genuine quotation of the letter *n*, rendered as an elision. So
441
+ does the same shape one level deeper, `He said "press 'n' now".` Both are pinned as fixtures
442
+ (`spec/fixtures/en-US.json`) so the exposure sits in the conformance suite rather than in
443
+ someone's content. The veto's alphabet is one code point and the spaces around it; nothing in that
444
+ alphabet can encode "this is a quoted letter, not an elided one", and no bounded literal contract
445
+ over these bytes could. The trade is deliberate and was made with the numbers in view: 0.5.0 paid
446
+ for this same shape by declining **every** short quotation in **every** locale, a far larger and
447
+ far commoner class of error than quoting the single letter *n*.
448
+
449
+ **V1 — same-V1-identity adjacency veto** (both widths):
450
+
451
+ > Define `V1ID(c) = U+2019 if c = U+0027, else c` — the **V1 identity** of a code point. `V1ID`
452
+ > is the identity function everywhere except at U+0027; it is *not* a claim that a given U+0027
453
+ > will become U+2019 (see below). Let `gapInsertable = g ∈ SPACE-RIGHT ∨ g ∈ SPACE-LEFT`. If
454
+ > `V1ID(Llit) = V1ID(g)`, or (`Llit ∈ INLINE-SPACE` and `V1ID(Lskip) = V1ID(g)` and
455
+ > `gapInsertable`), set both capabilities false. Symmetrically for `Rlit`/`Rskip`.
456
+
457
+ `""`, `''`, `««`, `””` and any longer run of one glyph, by the first (literal) clause. This is
458
+ 0.1.0's veto, generalised, and it keeps the `"""a""` witness dead and `"a "" b"` intact.
459
+ **Granularity is deliberately "same V1 identity", not "same width"**: the narrower test
460
+ reproduces both witnesses, and a width-based test would veto the ordinary mixed boundary `"'`.
461
+
462
+ **`V1ID` and why V1 cannot compare raw code points (spec 0.4.1).** `apostrophe` (R₆) is the one
463
+ rule ordered after `quotes` that can change a candidate's own code point, and every edit it makes
464
+ replaces a U+0027 with U+2019 — the only code point it ever emits for that position
465
+ (`apostrophe.md` §1). **This is not an unconditional mapping.** `apostrophe`'s own case ladder
466
+ (`apostrophe.md` §3.3) leaves a U+0027 unedited under two of its six cases — the **prime guard**
467
+ (case 1: `6' 2"`, a foot/inch mark) and **case 5** (`a ' b`, `''`: nothing inferable) — and
468
+ `apostrophe.md` §5 names both explicitly as surviving, unedited, to the rule's own second-run
469
+ argument. So only a *subset* of the U+0027s `quotes` leaves unmatched are ever actually converted.
470
+
471
+ Every other capability test in this section reads *class* membership, which is unaffected by
472
+ that substitution regardless of whether it happens (Lemma A, §5) — but V1, uniquely, compares
473
+ *code points*, because that is its job: telling `""` from `"'`. A literal comparison is therefore
474
+ the one place a U+0027 neighbour's *eventual* fate matters, and `quotes` cannot know that fate
475
+ from inside its own pass: whether a given U+0027 reaches `apostrophe` at all — as opposed to
476
+ being claimed by `quotes` itself and rendered as a quote glyph — and what its *own* neighbours
477
+ will read once `quotes` has finished rendering the whole accepted set, are exactly the outputs
478
+ this rule is still computing when V1 runs. Deciding V1 exactly would mean re-deriving
479
+ `apostrophe`'s verdict against `quotes`' own not-yet-final output — a circular second pipeline
480
+ inside this rule, not a fix. (A per-position "is this specific U+0027 definitely going to survive
481
+ unmatched with these exact neighbours" analysis was considered and rejected for the same reason:
482
+ neighbouring accepted pairs can rewrite the very code points that analysis would need to read
483
+ first.)
484
+
485
+ `V1ID` is therefore a **conservative over-approximation**, not an exact rule: it treats *every*
486
+ U+0027 as if it might already be, or might become, U+2019, and every U+2019 as if it might be an
487
+ unconverted U+0027 — regardless of which of `apostrophe`'s six cases (1, 2, 3, 3a, 4, 5) will actually apply.
488
+
489
+ **What "one-directional" bounds, precisely — a local claim, not a pipeline-level one.** At a
490
+ single V1 comparison, `V1ID` can only ever *add* a veto: it merges two identities that raw
491
+ equality would have kept apart, and it never grants `canOpen`/`canClose` to a candidate that
492
+ would not otherwise have had one — `V1ID` only ever makes the equality test *stricter*, never
493
+ looser. **That local fact does not extend to the rule's global output.** Pass 2 (§3.3, the stack
494
+ pairing) and pass 4 (§3.5, the certification gate) both operate over the *whole* candidate list at
495
+ once, not candidate-by-candidate: declining one candidate can remove a crossing pair, a nesting
496
+ conflict, or a certification instability that was previously the reason a *different, unrelated*
497
+ pair failed to certify. An extra local veto can therefore **indirectly** let another pair certify
498
+ that previously could not, reassign which glyphs/depth a pair receives, or change what `nbsp`
499
+ (order 70) inserts downstream — this is not a hypothetical: **the reported counterexample is
500
+ exactly this shape.** On pass 1, `V1ID` declining the crossing `NARROW` pair (`U+2019`
501
+ adjacent to the unmatched `U+0027`) is precisely what lets the `WIDE` pair certify instead of
502
+ being caught in the same gate rejection that, pre-fix, declined both. Vetoing one candidate
503
+ *enabled* a different conversion; "extra decline only" was never literally true of the rule's
504
+ output, only of the single comparison `V1ID` changes.
505
+
506
+ The actual safety argument therefore does **not** rest on a monotonicity claim. It rests on: (1)
507
+ a missed pairing is the failure mode this whole spec already treats as safe — "missing a possible
508
+ improvement is safer than damaging text" (`PLAN.md` §6.1,
509
+ `docs/AUDIT_REMEDIATION_AND_RELEASE_PLAN.md` §3.1) — against the defect `V1ID` closes, a hard
510
+ release blocker (an idempotency violation); and (2) **inspection of every output `V1ID` globally
511
+ changes within a reproducible bounded sweep**, not an a-priori argument about direction. The "V1ID
512
+ empirical audit" section below is that inspection: every changed output in the sweep — declined,
513
+ newly enabled, or reassigned — is individually classified, and none is a false positive against
514
+ plausible prose. `tests/rules/quotes.test.ts` pins one permanent regression case per distinct
515
+ mechanism found: a pure conservative veto (prime-guard and case-5 U+0027 preserved exactly), the
516
+ indirect newly-enabled-outer-pair mechanism (the reported counterexample's own shape), a pairing
517
+ reassignment, and the `fr`/`fr-CA` downstream `nbsp` consequence — plus a direct test that the
518
+ plain U+0027 vs. U+2019 distinction still governs `apostrophe` and the medial veto exactly as
519
+ before, unaffected by `V1ID`, which is scoped to V1's own comparison alone.
520
+
521
+ **Reachability, precisely.** `QUOTEMARK`, V1, and `apostrophe`'s case ladder are defined
522
+ identically for every locale, so the *algorithmic* shape (an unmatched U+0027 adjacent to a
523
+ `QUOTEMARK` candidate whose glyph is U+2019) is constructible in every locale with a synthetic
524
+ string, and `tests/rules/quotes.test.ts` sweeps it across every entry in `KNOWN_LOCALES` for
525
+ exactly that reason — that is a claim about the algorithm, not about naturally occurring prose.
526
+ Separately, and more narrowly: `en-GB`'s primary `close` and `fi`/`sv`'s secondary pair (both
527
+ sides) are U+2019 *as that locale's own prescribed glyph* — not merely a candidate a foreign or
528
+ adversarial input could contain — so those three locales are where an author's ordinary,
529
+ correctly-typed document (one already-closed quotation immediately followed by, say, a
530
+ possessive apostrophe with no separating space) is the most plausible route to this shape without
531
+ any copy-pasted or malformed input at all. `de-CH`, where the defect was first found, reaches it
532
+ only via a mixed straight/curly input (`spec/fixtures/de-CH.json`,
533
+ `de-ch-quotes-041-cobug-apostrophe-v1`) — de-CH's own glyphs are `« »`/`‹ ›`, neither of which is
534
+ U+2019, so the witness there is adversarial, not a natural-content route.
535
+
536
+ **V1ID empirical audit.** `quotes.ts`'s V1 block was twice temporarily reverted to the pre-0.4.1
537
+ raw comparison, run against a bounded, deterministic alphabet across every locale in
538
+ `spec/locales/registry.json`, and restored — a controlled experiment, not a committed second
539
+ engine, and not itself part of this repository (the exact reversion is exactly the diff `git
540
+ diff` shows against `tests/rules/quotes.test.ts`'s permanent assertions, which is what actually
541
+ pins the result). Exact counts from that run are deliberately **not** recorded here: an
542
+ uncommitted, one-off harness is not something a future maintainer can re-run from this document
543
+ alone, and this spec's normative claims must not depend on an experiment nobody can reproduce.
544
+ What is recorded, and *is* reproducible — by running `npx vitest run tests/rules/quotes.test.ts`
545
+ against `main` — is the durable conclusion the audit reached:
546
+
547
+ - Every output `V1ID` changes relative to the raw comparison, across the swept alphabet and every
548
+ locale, falls into exactly one of four mechanisms: a pure extra conservative veto (the mark
549
+ stays exactly as authored); an *indirect* newly-enabled conversion (vetoing one candidate
550
+ removes a crossing/certification conflict and lets a different, unrelated pair certify — the
551
+ reported counterexample's own shape, run in general); a pairing reassignment (the same input
552
+ pairs differently, not more or less); and the `fr`/`fr-CA` downstream `nbsp` consequence of a
553
+ differently accepted or declined pair. No difference in the swept range was unexplained.
554
+ - Every newly-enabled or reassigned output in the swept range was inspected against ordinary-prose
555
+ plausibility, not merely checked for a digit adjacency. None resembles content a human author
556
+ would write and reject: the newly-enabled class never contains a letter or digit at all (its
557
+ witnesses are runs of quote-class punctuation with no word between them, not prose); the
558
+ reassignment class only ever chooses between two already-plausible readings of a short ambiguous
559
+ fragment (e.g. an ordinary quotation vs. two independent elisions), never invents a reading with
560
+ no textual basis.
561
+ - A separate, narrower sweep confirmed the raw-comparison variant is genuinely non-idempotent
562
+ within a bound that reaches the reported defect's minimal witness shape, and that `V1ID` is not,
563
+ across every locale — the direct measurement of the closure this fix provides, not merely of its
564
+ conservative cost.
565
+
566
+ `tests/rules/quotes.test.ts`'s "V1ID" describe block pins one permanent, exact-output regression
567
+ case for each of the four mechanisms above (plus the U+0027/U+2019-still-matters-outside-V1
568
+ control) — that is the reproducible evidence this section rests on.
569
+
570
+ **The second clause is not in either design document that fed this revision, and closes a
571
+ composition-obligation violation found only by running the implementation.** A literal-only V1
572
+ lets the certification gate correctly decline a pairing when rendering it would place two
573
+ identical glyphs strictly adjacent (e.g. `«` immediately left of a mark the gate would render to
574
+ `«`) — but if `nbsp` (a *later* rule) subsequently inserts a space at exactly that gap, on the
575
+ *next* full-pipeline pass `quotes` sees the two occurrences with a space between them, the
576
+ literal adjacency is gone, and the *same* pairing certifies. That is `nbsp` creating work for an
577
+ earlier-ordered rule, forbidden by `pipeline-idempotency.md` §2's composition obligation, and it
578
+ reproduces with two witnesses, pinned in `spec/fixtures/fr.json` as `fr-quotes-030-cobug-html-tag-boundary`
579
+ (`«"<p class="x">"` in `html` mode) and `fr-quotes-030-cobug-double-open-then-closer` (`««”` in
580
+ `text` mode). The fix reads the
581
+ skip in exactly the two positions Lemma B (§5) already names as `nbsp`'s whole insertion
582
+ surface next to a quote mark — a candidate's own `g` value is shared identically with the
583
+ same-glyph neighbour a skip reaches, so checking `g`'s own `SPACE-RIGHT`/`SPACE-LEFT` membership
584
+ covers both of Lemma B's insertion sites at once, regardless of which of the two occurrences is
585
+ on which side of the gap. **A plain skip-based V1** (comparing `Lskip`/`Rskip` to `g`
586
+ unconditionally, with no `gapInsertable` guard) **was tried and rejected**: it also vetoes two
587
+ genuinely distinct, space-separated quotations (`'a' 'b'`, ordinary prose, not a typewriter
588
+ artefact) in every locale, including ones with no spaced pair at all — the guard is what keeps
589
+ the veto scoped to positions `nbsp` can actually reach.
590
+
591
+ Under mandate 1 the veto is **no longer permanent** — a converted mark can acquire or lose an
592
+ identical neighbour between runs — and that is stated rather than papered over: **the gate
593
+ (§3.5), not this veto, is the guarantor of stability against `quotes`' own re-derivation.** The
594
+ `gapInsertable` clause is what extends that guarantee to survive `nbsp` acting *between*
595
+ pipeline passes, which the gate alone — being a property of a single call — cannot see.
596
+
597
+ Record `{ index, width ∈ {WIDE, NARROW}, canOpen, canClose }` for every `i` with at least one
598
+ capability true. Candidates with both false are dropped and never touched.
599
+
600
+ Pass 1 still does **not** decide that a `NARROW` mark with a letter to the right and a space to
601
+ the left is a leading elision. It is `canOpen` and enters the pairing; if it finds no partner it
602
+ survives and `apostrophe` renders it U+2019. The pairing settles the ambiguity, not a heuristic.
603
+
604
+ ### 3.3 Pass 2 — pair the candidates, one stack per width
605
+
606
+ **There are two stacks, one for `WIDE` and one for `NARROW`, and a candidate only ever touches
607
+ the stack for its own width.** 0.1.0's per-kind split generalised from "the literal code point
608
+ `DQ`/`SQ`" to "the glyph's visual width" — the property that actually carried the argument: no
609
+ reader opens a quotation with a double mark and closes it with a single one, in any language.
610
+ This is what keeps a Greek elision (`σ'`, `NARROW`) from closing a `WIDE` opener.
611
+
612
+ Walk the candidate list in index order. For candidate `c`, let `S` be the stack for `c.width`:
613
+
614
+ 1. If `c.canClose`, `S` is non-empty, and **`¬vacuous(top(S).index, c.index)`** → pop `o`, record
615
+ the pair `(o.index, c.index)`.
616
+ 2. Else if `c.canOpen` → push `c` onto `S`.
617
+ 3. Else record `c` as unmatched.
618
+
619
+ When the walk ends, everything still on either stack is unmatched.
620
+
621
+ > **`vacuous(a, b)`** ⟺ every `cp[k]` with `a < k < b` is in `INLINE-SPACE` (vacuously true when
622
+ > `b = a + 1`).
623
+
624
+ **The vacuity condition is new and is forced by mandate 1.** In 0.1.0 an empty pair was
625
+ impossible: two adjacent marks of the same kind were vetoed by V1, and two of different kinds
626
+ were on different stacks. Now `«»` is two *different* code points of the *same* width, and `" "`
627
+ is two identical marks V1 cannot see past a space; both would enclose no content. The condition
628
+ sits in step 1 rather than as a fourth outcome: a closer that fails step 1 **falls through to
629
+ step 2 and then step 3**, so every candidate still reaches exactly one of three outcomes and the
630
+ accepted/unmatched partition stays exhaustive — that exhaustiveness is what makes the "declined"
631
+ set well-defined for the gate. **Do not add a fourth outcome to this list.**
632
+
633
+ Step 1 before step 2 — **closing takes precedence** — is retained unchanged. An `«` whose
634
+ right-hand space `nbsp` inserted reads `closeRight = Rskip`, exactly as it did before `nbsp`
635
+ touched it, so the tie-break's ambiguity does not shift under mandate 1.
636
+
637
+ ### 3.4 Pass 3 — assign glyphs by depth
638
+
639
+ For a set of pairs `A`, the **depth** of `p ∈ A` is
640
+
641
+ > `depth_A(p) = 1 + |{ q ∈ A : q.open < p.open ∧ p.close < q.close }|`
642
+
643
+ — one plus the number of *accepted* pairs that strictly enclose it. Odd depth → `quotes.primary`;
644
+ even depth → `quotes.secondary`. Depths beyond 2 alternate by the same parity.
645
+
646
+ **Depth is computed over the accepted set `A`, not over the raw pass-2 output.** On a second run
647
+ the accepted set *is* the raw set, so a depth taken over the raw set on run 1 and over the
648
+ accepted set on run 2 would disagree whenever the gate declined anything. Depth over `A` agrees
649
+ on both runs by construction.
650
+
651
+ **Depth, not the mark's original width, selects the pair.** `'He said "no" to me,' she noted.`
652
+ promotes the outer `NARROW` marks to `en-US`'s `WIDE` primary pair and demotes the inner ones —
653
+ the reason width drift exists at all, and the case the certification gate exists to re-verify.
654
+
655
+ ### 3.5 Pass 4 — the certification gate
656
+
657
+ This pass replaces 0.1.0's Claims 1–3. It is a *check the algorithm performs*, not a property a
658
+ reader must verify by argument alone.
659
+
660
+ **Rendering.** `render(cp, A)` is the array obtained by applying, for every `p ∈ A` with assigned
661
+ pair `P = pairFor(depth_A(p))`:
662
+
663
+ - replace `cp[p.open]` with `P.open` and `cp[p.close]` with `P.close`;
664
+ - **if and only if `P.innerSpace = "none"`**, delete the pair's two *inner runs* where the
665
+ landing guard permits:
666
+ - the **open-side run** = the maximal `INLINE-SPACE` run starting at `p.open + 1` (possibly
667
+ empty); its **landing** is the first code point past it;
668
+ - the **close-side run** = the maximal `INLINE-SPACE` run ending at `p.close - 1`; its landing
669
+ is the first code point before it;
670
+ - a run is deleted iff it is non-empty, its landing is in `DELETE-LANDING`, and it is not
671
+ simultaneously both of the pair's runs (a pair enclosing nothing but spaces deletes
672
+ neither — unreachable for an accepted pair given §3.3's vacuity condition, but stated because
673
+ the construction is otherwise silent about it).
674
+
675
+ `render` also yields the order-preserving index map `π` from surviving input indices to output
676
+ indices.
677
+
678
+ **Certification loop.**
679
+
680
+ ```
681
+ A := the pair set from pass 2
682
+ loop:
683
+ if A = ∅: accept ∅ and stop
684
+ (y, π) := render(cp, A)
685
+ B := passes 1–2 applied to y (raw, ungated)
686
+ if { (π(p.open), π(p.close)) : p ∈ A } = B: accept A and stop
687
+ A' := { p ∈ A : (π(p.open), π(p.close)) ∈ B }
688
+ if A' = A: A := A \ { the pair of A with the greatest `open` index }
689
+ else: A := A'
690
+ ```
691
+
692
+ **Every clause is normative, including the tie-break.** When the intersection fails to shrink
693
+ `A` — exactly when `B ⊋ π(A)`, i.e. the rendered array admits a pairing the input did not — one
694
+ pair is removed, and *which* pair is specified so two ports cannot disagree.
695
+
696
+ **Termination.** Each iteration either accepts or strictly shrinks `A`. `A` is finite and `∅`
697
+ accepts unconditionally, so the loop runs at most `|A₀| + 1` times. Cost is `O(k·n)` with `k` the
698
+ number of pairs; in the overwhelmingly common case the first round certifies and the extra cost
699
+ is one `O(n)` pass. An implementation may short-circuit on that.
700
+
701
+ **Why the exit condition is `B = π(A)` and not `π(A) ⊆ B`.** Exiting with extra pairs in `B` is
702
+ unsound: the second run *starts* from `B`, and if `B` certifies, the second run renders `B`'s
703
+ extra pairs and the output differs. The intersection alone does not terminate at a sound state,
704
+ which is why the forced-removal clause exists.
705
+
706
+ **Empirical status of the forced-removal clause.** An exhaustive sweep of ~850,000 inputs across
707
+ all ten locales (the alphabet of §9's fixtures, to length 4–5, plus the named witnesses) did not
708
+ observe the clause firing even once — every case that reached the gate either certified in round
709
+ 1 or was resolved by the plain intersection shrinking `A`. The clause is implemented as specified
710
+ (it is strictly more conservative than omitting it, never less sound, so it costs nothing even if
711
+ unreachable in practice) but its necessity is **not** demonstrated by a concrete witness as of
712
+ this writing. If a future sweep finds one, it belongs in `spec/fixtures/` as a named case.
713
+
714
+ ### 3.6 Pass 5 — emit
715
+
716
+ For each `p` in the accepted set:
717
+
718
+ - an edit replacing `cp[p.open]` with the assigned `open` glyph, **omitted if the code point is
719
+ already that glyph**;
720
+ - likewise at `p.close`;
721
+ - a deletion edit for each inner run the landing guard permitted, **omitted if the run is
722
+ empty**.
723
+
724
+ An edit whose replacement is identical to the span it replaces is never emitted — the
725
+ invisible-edit principle `dashes.md` §3.3.1 states for the joiner, applied so that
726
+ already-correct text produces a byte-identical no-op rather than a diff of self-replacements.
727
+ Suppression is per **mark**, not per pair: a pair where only one side needed a glyph change
728
+ emits exactly one glyph edit.
729
+
730
+ Emitted spans are pairwise disjoint: replacements sit at distinct mark indices; a deletion span
731
+ lies strictly between a mark and its landing; and since accepted pairs are properly nested or
732
+ disjoint (stack discipline) and a pair enclosing only spaces deletes neither run, no two spans
733
+ can claim the same index. Edits are sorted ascending before returning.
734
+
735
+ Unmatched candidates and declined pairs produce no edit and keep their code point exactly.
736
+
737
+ ### 3.7 `innerSpace`: `quotes` deletes, `nbsp` inserts
738
+
739
+ The split is asymmetric on purpose.
740
+
741
+ - `innerSpace = "none"` — a space inside the pair is *wrong* and nobody else can remove it.
742
+ `nbsp` has no deletion capability at all. `quotes` deletes it, subject to the landing guard.
743
+ - `innerSpace ≠ "none"` — a space inside the pair is *required*, and `nbsp` (order 70) already
744
+ inserts and converts it, correctly and idempotently, at `nbsp.md` §3.10. `quotes` emits
745
+ nothing and deletes nothing there.
746
+
747
+ So `quotes` still never inserts a space, `nbsp` keeps ownership of all no-break spacing, and
748
+ there is no second copy of anyone's rules. `"mot"` → `«mot»` → (via `nbsp`) `«` U+00A0 `mot`
749
+ U+00A0 `»` — U+00A0 because `fr`'s primary pair sets `innerSpace: "nbsp"`, and **no shipped
750
+ locale sets `"narrow-nbsp"`** (nbsp.md §3.1a); this read U+202F until spec 1.3.0, which is a
751
+ claim about a locale file that the file never supported. `« mot »` → (no edit from `quotes`) →
752
+ the same output via `nbsp`. Both paths
753
+ converge, which is the property the pipeline needs.
754
+
755
+ **The landing guard is not a heuristic; it is a structural CO discharge.** The deleted run's near
756
+ side is a quote glyph and its far side is in `ALNUM ∪ QUOTEMARK`. Every class the four
757
+ earlier-ordered rules read as evidence — `spaces`' `CONTENT`/`STRIP-BEFORE`/bracket sets,
758
+ `ellipsis`' dot runs, `dashes`' `DASH`, `INERT-DASH`, `DIGIT`, `SPACE ∪ NOBREAK-SPACE`, `LETTER`,
759
+ `ROMAN`, `hyphen`'s `WORDISH` — classifies a quote glyph and a U+0020 **identically** (both
760
+ outside all of them, except `SPACE`, which the guard's own construction keeps away from any dash
761
+ run). `ALNUM ∪ QUOTEMARK` is the largest landing class for which that holds; letting the
762
+ deletion land on a dash is what would produce `" --x"` → `“--x”` → `“—x”`, a violation of `I₃`.
763
+ Two absolutes fall out of the same construction: **a `BREAK` is never deleted** (not in
764
+ `INLINE-SPACE`, so it terminates the run and, being outside `DELETE-LANDING`, declines the
765
+ deletion) and **`MARKER` is never crossed** (same mechanism), so the span partition is stable
766
+ between runs as `modes.md` §5 requires.
767
+
768
+ ---
769
+
770
+ ## 4. Must not touch
771
+
772
+ Each bullet is **[P]** (a guarantee of `transform`) or **[R]** (true of this rule alone, with the
773
+ rule that can falsify it named), per `pipeline-idempotency.md` §5.2.
774
+
775
+ - **[P] A medial apostrophe:** `don't`, `l'été`, `O'Brien`, `Hawai'i`, `1990's` — and their
776
+ U+2019 forms on a second pass. Vetoed in §3.2.
777
+ - **[R] A leading elision that finds no partner** (`'90s`, `'tis`) and **[R] a trailing
778
+ possessive that finds no partner** (`the dogs' bowls`). Handed to `apostrophe`.
779
+ - **[R] A foot or inch mark:** `6' 2"`. Both marks are `canClose` only with an empty stack. The
780
+ 0.1.0 caveat still applies verbatim: in `"6' 2"` the foot mark pairs, and the protection is not
781
+ pipeline-level.
782
+ - **[P] Any unbalanced mark.** Convert what is unambiguous, leave the rest — §3.6 of 0.1.0's
783
+ policy is unchanged.
784
+ - **[P] Any run of two or more identical quote marks:** `""`, `'''`, `««`, `””`. V1.
785
+ - **[P] A pair enclosing nothing but spaces:** `«»`, `" "`, `” ”`. §3.3's vacuity condition.
786
+ - **[P] U+2032, U+2033, U+02BC, U+0060, U+00B4** and the LaTeX idioms `` ` `` and `''`. Not in
787
+ `QUOTEMARK`.
788
+ - **[P] Anything inside a skipped region.** Attribute values, code spans, fenced code, `<pre>`,
789
+ URLs — removed by the mode adapter.
790
+ - **[P] Line terminators.** Never deleted, never crossed by a skip walk.
791
+ - **[R] Outer spacing.** This rule deletes only *inner* spacing and inserts nothing.
792
+
793
+ **No longer on this list, by operator decision:** every non-straight quotation glyph. `“ ” « »
794
+ „ ‘ ’ ‚ ‹ ›` and the rest are now candidates. `en-GB` text containing a correct `“…”` pair will
795
+ be rewritten to `‘…’`; `de-DE` text containing `«Wort»` becomes `„Wort“`. That is mandate 1, and
796
+ it is the largest behavioural change in this revision.
797
+
798
+ **Known, accepted residual risk — not new under this revision.** A leading elision that has
799
+ already been curled to U+2019 (`’90s were fun,’ he said`) can, if a later unrelated closing mark
800
+ exists on the same stack, be paired as an opener and lose its elided-digit reading. This was
801
+ already possible for straight input under 0.1.0 (`'90s were fun,' he said` is corrupted
802
+ identically) — it is §7 item 3's family, "no local way to distinguish an elision from a real
803
+ opening quote", now also reachable via curly input because mandate 1 makes curly input a
804
+ candidate at all. **§7 item 8's `quotes.elisionIdioms` (spec 0.4.0) does not cover this case**:
805
+ that veto fires only on a listed idiom's full left/elided/right context found together, not on
806
+ a lone elision separated from any partner by arbitrary distance. It is **not** a new defect this
807
+ revision introduces; §5's Corollary A1 explains why no U+2019-specific restriction is applied, and
808
+ `spec/fixtures/` pins the verified behaviour across `en-GB`, `en-US`, `fi` and `sv` (§6, rows
809
+ E1–E4) so the tradeoff is a citable fact rather than an assumption.
810
+
811
+ ---
812
+
813
+ ## 5. Idempotency argument
814
+
815
+ 0.1.0's Claims 1 and 2 are **withdrawn**, not weakened. Claim 1 ("a converted position is not a
816
+ candidate on the second run") is false by construction under mandate 1, and with it goes the
817
+ permanence of V1: two marks that were distinct can both be assigned the same glyph, so the
818
+ veto's verdict is no longer immutable. Claim 2's monotonicity is likewise gone: a converted mark
819
+ can *gain* a capability. Claim 3 rested on both. Three replacements follow.
820
+
821
+ ### Lemma A — glyph-blindness
822
+
823
+ > Replacing any `QUOTEMARK` at index `j` with any other `QUOTEMARK` changes no capability of any
824
+ > candidate at any index `i ≠ j`.
825
+
826
+ *Proof.* A candidate's four neighbour reads are compared against `NONE`, `SPACELIKE`, `OPENISH`,
827
+ `CLOSEISH`, `DASHISH`, `QUOTEMARK`, `MARKER` and `ALNUM` (the medial veto), and V1 compares
828
+ `V1ID` of the relevant code points against `V1ID(g)` (spec 0.4.1). Every `QUOTEMARK` is in
829
+ `OPENISH`, in `CLOSEISH`, in `QUOTEMARK`, in none of `SPACELIKE`, `DASHISH`, `ALNUM`, and is
830
+ neither `MARKER` nor `NONE`. All seven class tests are therefore **invariant**, not merely
831
+ monotone. The skip walks are invariant because `QUOTEMARK ∩ INLINE-SPACE = ∅`.
832
+
833
+ **V1 is invariant too, as of spec 0.4.1, and this required changing V1 itself, not just arguing
834
+ about it.** Before 0.4.1, V1 compared raw code points, and replacing a neighbour's `U+0027` with
835
+ `U+2019` — the one substitution this lemma's hypothesis is actually exercised against by
836
+ `apostrophe`, via Corollary A1 below — could change whether that neighbour's code point equalled
837
+ a candidate's own `g`, which is exactly the comparison V1 makes. That was a real gap, not a
838
+ theorem this document previously proved: the claim once made here — that "V1's movement is
839
+ precisely what the gate certifies" — conflated two different sources of movement. The
840
+ certification gate (§3.5) re-derives candidates from `quotes`' *own* hypothetical render, within
841
+ one call; it has no visibility into a *different* rule editing the input between two separate
842
+ pipeline invocations, which is exactly what `apostrophe` does. §3.2's `V1ID` closes the actual
843
+ gap: `V1ID(U+0027) = V1ID(U+2019) = U+2019`, and `V1ID` is the identity elsewhere, so replacing a
844
+ `U+0027` neighbour with `U+2019` (or vice versa) changes neither side of any V1 comparison in
845
+ `V1ID`-space. V1's verdict is now invariant under Lemma A's hypothesis by the same construction
846
+ as the other seven tests, not by appeal to a mechanism (the gate) that cannot see the edit in
847
+ question. ∎
848
+
849
+ **The listed elision veto (spec 0.4.0) is covered by the same argument, not a new proof
850
+ obligation.** Every input it reads is one Lemma A already accounts for or one nothing in this
851
+ pipeline ever touches:
852
+
853
+ - `NARROW` membership of `cp[i]` and `cp[j]` — invariant, same as every test above.
854
+ - The elided content `cp[i+1 … i+k]` — a literal `LETTER` run (in the shipped `en-US` entry;
855
+ the schema permits any non-empty string, but an idiom's `elided` field is authored as
856
+ letters). `apostrophe` (R₆) never touches a `LETTER`; it replaces exactly one `SQ` with
857
+ exactly one U+2019 and nothing else (`apostrophe.md` §1). Unaffected by any rule ordered at
858
+ or before `quotes` for the same reason.
859
+ - The single `INLINE-SPACE` code point on each outer side, and the `left`/`right` `LETTER` runs
860
+ beyond it — `apostrophe` touches neither an `INLINE-SPACE` code point nor a `LETTER`; nothing
861
+ ordered before `quotes` does either (each earlier rule's own §4 "Must not touch" already
862
+ states this for ordinary prose letters and spacing it did not itself just emit, and none of
863
+ them emits a `LETTER`).
864
+
865
+ So substituting `g`'s own U+0027 for U+2019 at `i` or `j` (the only edit `apostrophe` can make
866
+ to this construction, and the only one reachable in this direction) changes none of the four
867
+ inputs above, and the veto's verdict on a given pair of indices is invariant under Lemma A's
868
+ hypothesis exactly as every other capability test is. Concretely: on the pipeline's second
869
+ pass, `rock ’n’ roll`'s two marks are U+2019 rather than U+0027, still `NARROW`, still bounded
870
+ by the identical literal `rock`/`n`/`roll` context — the veto fires identically, both marks are
871
+ vetoed again, `quotes` makes no pairing, and `apostrophe` does not act on U+2019 at all (§4).
872
+ The construction is a fixed point (§6 row P4 pins it).
873
+
874
+ **Modes.** `text` is the base case above. In `html` and `markdown`, `modes.md` §3.2's
875
+ concatenation-with-marker model means the veto's bounded lookaround can land on a `MARKER` (a
876
+ negative integer, in none of `LETTER`, `INLINE-SPACE`, `NARROW`) exactly where a real code
877
+ point would otherwise be — a context word split from its mark by an element or span boundary
878
+ fails the `INLINE-SPACE` test at that position and the veto does not fire (§6 rows H1–H3). This
879
+ is not a special case for modes: it is the same "word not found" outcome an ordinary document
880
+ boundary produces (`Llit`/`Rlit` = `NONE`), and `modes.md` §5's stability argument — no rule may
881
+ emit a code point at a position that changes how the next pass partitions spans — is
882
+ unaffected, because this veto neither emits anything nor changes any span boundary; it only
883
+ narrows which pairs `quotes` itself forms, which `modes.md` §5 part 3 already covers for the
884
+ rest of this rule's determinism.
885
+
886
+ > **Corollary A1 — `apostrophe` (R₆) is structurally invisible (spec 0.4.1: was claimed, not yet
887
+ > true, before `V1ID`).** `apostrophe` replaces U+0027 with U+2019 and nothing else. As a
888
+ > *neighbour*, Lemma A applies — **including its V1 clause**, now that V1 itself compares `V1ID`
889
+ > rather than raw code points. As the mark *itself*: both `U+0027` and `U+2019` are in `NARROW`,
890
+ > so the stack partition is unchanged; the medial veto is stated over `NARROW`, so its verdict is
891
+ > unchanged; neither is in `SPACE-RIGHT`/`SPACE-LEFT`, by constraint **Q-A**; and `V1ID` maps
892
+ > both to the same value, so V1's verdict on the mark's own candidacy is unchanged too. Every
893
+ > verdict is identical. This is a genuine **CO-S** discharge in the sense of
894
+ > `pipeline-idempotency.md` §5.1a — `E(apostrophe) = {U+2019}` is wholly inert for this rule,
895
+ > because `V1ID` makes it indistinguishable from the one code point it can only ever have
896
+ > replaced.
897
+ >
898
+ > **This corollary was wrong for one release** (spec 0.3.0–0.4.0): it asserted glyph-blindness
899
+ > covered V1 by deferring to Lemma A's text, and Lemma A's own proof at the time explicitly
900
+ > carved V1 out ("the only test that can move") and waved at the certification gate as the
901
+ > guarantor — a claim about a *different* mechanism (single-call self-re-derivation) standing in
902
+ > for a proof about *this* one (cross-call stability against a later rule). The gap was real: a
903
+ > `NARROW` candidate `g = U+2019` adjacent to an unmatched `U+0027` had a literal-neighbour
904
+ > comparison that read differently before and after `apostrophe` converted that neighbour on the
905
+ > *previous* pipeline pass, which is precisely `apostrophe` creating work for `quotes` —
906
+ > forbidden by `pipeline-idempotency.md` §2. Found by `fast-check` in `de-CH`
907
+ > (`spec/fixtures/de-CH.json`, `de-ch-quotes-041-cobug-apostrophe-v1`).
908
+ >
909
+ > **Reachability, stated precisely rather than as a blanket "every locale."** The *algorithm* —
910
+ > `QUOTEMARK`, V1, and `apostrophe`'s case ladder — is locale-independent, so the same synthetic
911
+ > witness can be constructed and passed through every locale's data, and
912
+ > `tests/rules/quotes.test.ts` does exactly that as a regression sweep. That is a claim about
913
+ > algorithmic constructibility, not evidence that the *pre-fix* defect actually reproduced in
914
+ > every locale — it was verified directly only in `de-CH`, where `« »`/`‹ ›` (neither U+2019) make
915
+ > the witness adversarial rather than naturally occurring. Separately: `en-GB`'s primary `close`
916
+ > and `fi`/`sv`'s secondary pair (both sides) *are* U+2019 as those locales' own prescribed
917
+ > glyphs, which is what makes this shape especially plausible from ordinary, correctly-typed
918
+ > content there — not merely reachable by construction — even though no canonical fixture is
919
+ > pinned in those locales for it (§3.2's `V1ID` discussion states the same reachability
920
+ > distinction and gives the empirical delta of the fix, "V1ID empirical audit" below).
921
+ >
922
+ > **On the elision-vs-closer tradeoff.** No U+2019-specific "can never open" restriction is
923
+ > applied. Such a restriction was considered and rejected: `fi`/`sv`'s secondary pair is U+2019
924
+ > on *both* sides, and forcing `canOpen = false` for every U+2019 would mean an already-correct
925
+ > `fi`/`sv` secondary pair could never be *recognised* as a pair at all — only ever declined —
926
+ > which directly contradicts mandate 1's purpose for that locale family. The cost of not
927
+ > restricting is stated in §4 and verified in §6 (rows E1–E4): it is the same pre-existing
928
+ > `rock 'n' roll`-family ambiguity 0.1.0 already documented for straight input, now also
929
+ > reachable via curly input, not a new class of damage.
930
+
931
+ ### Lemma B — space-inertness at every position `nbsp` can reach
932
+
933
+ > Inserting a code point of `INLINE-SPACE`, or converting one `INLINE-SPACE` code point to
934
+ > another, at any position `nbsp` can act on, changes no capability of any candidate.
935
+
936
+ *Proof.* Conversion is trivial: every test reads class membership and U+0020, U+00A0, U+202F are
937
+ all in `INLINE-SPACE`. For insertion, `nbsp` has exactly three insertion sites (`nbsp.ts`
938
+ `claimInsertion`):
939
+
940
+ 1. **N8, immediately right of an occurrence of `g ∈ SPACE-RIGHT`.**
941
+ - *`g` itself*: `canOpen` reads `Rskip` (skips) and `openLeft` (untouched); `canClose` reads
942
+ `Lskip` (untouched) and `closeRight`, which is `Rskip` because `g ∈ SPACE-RIGHT` (skips).
943
+ All four invariant.
944
+ - *the candidate `c` whose `Llit` was `g`*: `c.canOpen` reads `openLeft`, which is `Llit`
945
+ (unless `c ∈ SPACE-LEFT`, in which case it skips): `g ∈ OPENISH` is accepted,
946
+ `INLINE-SPACE ⊂ SPACELIKE` is accepted — same verdict. `c.canClose` reads `Lskip`, which
947
+ skips the insertion back to `g` — same. Right-hand reads untouched.
948
+ - no other index has a changed neighbourhood.
949
+ 2. **N8, immediately left of `h ∈ SPACE-LEFT`.** Mirror image: `h.canOpen` reads
950
+ `openLeft = Lskip` (skips, because `h ∈ SPACE-LEFT`); `h.canClose` reads `Lskip`. The
951
+ candidate whose `Rlit` was `h` reads `closeRight = Rlit`, moving from `h ∈ CLOSEISH`
952
+ (accepted) to `INLINE-SPACE` (accepted) — same verdict; and `Rskip`, which skips back to `h`.
953
+ 3. **N1/N2, immediately left of a mark `m ∈ nbsp.beforePunctuation ∪
954
+ nbsp.narrowBeforePunctuation`.** The only quote mark whose neighbourhood changes is a `g`
955
+ with `Rlit = m`. Its `closeRight` moves from `m` to `INLINE-SPACE`; by constraint **Q-P**,
956
+ `m ∈ CLOSEISH`, so both are accepted — same verdict. Its `Rskip` skips back to `m`. (`nbsp`'s
957
+ own step 3 already declines to insert after an opening glyph.)
958
+
959
+ No other sub-rule inserts; N3–N7, N9, N10 only convert. ∎
960
+
961
+ **V1's `gapInsertable` clause (§3.2) is what extends this proof to V1 itself.** The four
962
+ capability tests above are the ones Lemma B's statement covers directly, but V1 is also a test a
963
+ candidate's own capabilities depend on, and it compares *code points*, not class membership —
964
+ Lemma B's "accepts `SPACELIKE` either way" argument does not apply to it. `gapInsertable` is
965
+ defined to hold exactly when `g` is a member of `SPACE-RIGHT` or `SPACE-LEFT` — precisely the
966
+ condition under which `nbsp`'s insertion sites 1 and 2 above can place a code point at the one
967
+ gap V1's second clause reads across — so V1's verdict is invariant to that insertion by the same
968
+ argument, case by case: at insertion site 1, the candidate immediately right of `g' ∈ SPACE-RIGHT`
969
+ has `gapInsertable = true` when its own glyph equals `g'` (the only case V1's second clause can
970
+ ever fire for it), and the literal-vs-skipped reads agree on whether that neighbour is `g'`,
971
+ exactly as case 1 above already shows for `canOpen`/`canClose`. Insertion site 2 is the mirror
972
+ image for `SPACE-LEFT`. Insertion site 3 (N1/N2) never inserts adjacent to a `QUOTEMARK` glyph on
973
+ the side V1 reads (it inserts beside listed punctuation, never beside a quote mark's own
974
+ same-glyph neighbour), so it cannot affect `gapInsertable`'s condition at all.
975
+
976
+ > **Corollary B1.** A run-1/run-2 asymmetry that used to arise from `quotes` reading a
977
+ > `SPACE-RIGHT` glyph's right side literally on one run and skipping it on the other cannot arise
978
+ > under §3.2: the read is `Rskip` on **both** runs, so `« **"` pairs on the first application and
979
+ > re-pairs identically on the second.
980
+
981
+ ### Certification — the theorem
982
+
983
+ > **Theorem.** For every input `x` and every locale, `Q(Q(x)) = Q(x)`.
984
+
985
+ *Proof.* Let `A` be the accepted set and `y = Q(x) = render(x, A)`.
986
+
987
+ **Case `A = ∅`.** Then `y = x`. `Q` is a deterministic function of its input and the locale, with
988
+ no state (`ARCHITECTURE.md` §7), so `Q(y) = Q(x) = x = y`. ∎
989
+
990
+ **Case `A ≠ ∅`.** The loop exited on the equality branch, so `G₀(y) = π(A)`, where `G₀` denotes
991
+ passes 1–2. Run `Q` on `y`. Passes 1–2 give `A₀' = G₀(y) = π(A)`. The gate's first round computes
992
+ `render(y, π(A))`, and this equals `y`:
993
+
994
+ - **Glyphs.** `π` is order-preserving, so `depth_{π(A)}(π(p)) = depth_A(p)` for every `p`; the
995
+ assigned pair is the same, and `y` already carries those glyphs at those indices by
996
+ construction of `render`.
997
+ - **Inner spacing.** For a pair whose `innerSpace ≠ "none"`, `render` never touches spacing, on
998
+ either run. For `innerSpace = "none"`: a run that was deleted is gone, so there is nothing to
999
+ delete; a run that was declined is still present and its landing is unchanged — `render` maps
1000
+ `QUOTEMARK` to `QUOTEMARK` and leaves `ALNUM` alone, so a landing outside `DELETE-LANDING`
1001
+ stays outside it, and the same run is declined again.
1002
+
1003
+ Therefore `G₀(render(y, π(A))) = G₀(y) = π(A)`, the gate certifies in round 1, and
1004
+ `Q(y) = render(y, π(A)) = y`. ∎
1005
+
1006
+ ### Composition obligation
1007
+
1008
+ This rule is **R₅**; the obligation runs against `spaces`, `ellipsis`, `dashes` and `hyphen`, and
1009
+ its own `I₅` must survive `apostrophe`, `symbols` and `nbsp`.
1010
+
1011
+ **What this rule emits.** `E(quotes)` = the four locale quote glyphs, each replacing exactly one
1012
+ `QUOTEMARK` at the same index. Plus deletions of a maximal `INLINE-SPACE` run whose near side is
1013
+ a quote glyph and whose far side is in `ALNUM ∪ QUOTEMARK`. **Nothing is inserted.**
1014
+
1015
+ **The listed elision veto (spec 0.4.0) changes none of the below.** It adds no emission — a
1016
+ vetoed pair of marks is simply never paired, so `E(quotes)` is exactly what it was — and can
1017
+ only shrink the set of pairs this rule forms, never grow it. Every discharge that follows was
1018
+ argued against a `quotes` whose positive behaviour is a superset of the current one, so each
1019
+ still holds without re-verification.
1020
+
1021
+ - **Against `I₁` (`spaces`).** Discharged. No U+0020 is emitted, so no S-b/S-c/S-d position can
1022
+ be created. Deletion removes a *maximal* run, so it cannot bring two U+0020 into contact
1023
+ (S-a); its neighbours after deletion are a quote glyph and an `ALNUM`/`QUOTEMARK`, both
1024
+ `CONTENT`.
1025
+ - **Against `I₂` (`ellipsis`).** Discharged. No U+002E or U+2026 is emitted; the deletion's
1026
+ landing is in `ALNUM ∪ QUOTEMARK`, which contains no dot, so no two dot runs are joined.
1027
+ - **Against `I₃` (`dashes`).** Discharged **structurally**. A quote glyph is in none of `DASH`,
1028
+ `INERT-DASH`, `DIGIT`, `SPACE`, `NOBREAK-SPACE`, `JOINER`, `ROMAN` or `LETTER`, so a
1029
+ `cp[L]`/`cp[R]`/`before`/`after` moving from one quote mark to another moves nowhere. The
1030
+ deletion's two sides are a quote glyph and a `DELETE-LANDING` member, **neither in
1031
+ `DASH ∪ INERT-DASH`** — so no dash run is ever adjacent to a deleted run, no token's
1032
+ `lsp`/`rsp` changes, and the landing separates the deletion from any dash run by at least one
1033
+ code point. This is the discharge the naive "delete the inner space unconditionally"
1034
+ formulation failed, with witness `" --x"` → `“--x”` → `“—x”` (row D1, §6).
1035
+ - **Against `I₄` (`hyphen`).** Discharged. `hyphen`'s boundary test is `¬WORDISH`, and both a
1036
+ U+0020 and a quote glyph are outside `WORDISH = ALNUM ∪ HYPHENISH`. The deletion never lands
1037
+ inside a word, only against its first or last code point.
1038
+ - **`I₅` against `apostrophe` (R₆).** Discharged by Corollary A1 (CO-S).
1039
+ - **`I₅` against `nbsp` (R₈).** Discharged by Lemma B (CO-S-shaped, over `nbsp`'s insertion
1040
+ *sites* rather than its emission alphabet alone — the alphabet by itself is not enough, because
1041
+ a `SPACELIKE` insertion is *not* inert at an arbitrary position; it is inert at exactly the
1042
+ positions `nbsp` can reach).
1043
+ - **`I₅` against `symbols` (R₇).** **Known-weak: a case analysis, not CO-S.** `symbols`
1044
+ collapses `(c)`/`(r)`/`(tm)` to a single sign, deleting code points, and converts `x` to `×`
1045
+ between numerals. A quote mark adjacent to a collapsed span sees `(` on its right (moving to
1046
+ `©`, which is neither `OPENISH` nor `CLOSEISH`) or `)` on its left (moving to `©`). Checked
1047
+ over all four capability tests, every reachable verdict coincides: `canOpen`'s right test
1048
+ accepts `(` and accepts `©`; `canClose`'s right test rejects both; `canOpen`'s left test
1049
+ rejects `)` and rejects `©`; `canClose`'s left test accepts both as non-`NONE`. This is not
1050
+ made structural — the sweep is the control, per §5.1a of `pipeline-idempotency.md`.
1051
+
1052
+ ---
1053
+
1054
+ ## 6. Worked examples
1055
+
1056
+ `␣` = U+0020, `⍽` = U+00A0, `⟶` = no change. Every row is verified by running
1057
+ `src/rules/quotes.ts` against the fixture of the same name in `spec/fixtures/`; none is
1058
+ hand-traced-only.
1059
+
1060
+ ### `en-US` — primary `“ ”`, secondary `‘ ’`, `innerSpace: none`
1061
+
1062
+ | # | Input | Output | Why |
1063
+ | --- | --- | --- | --- |
1064
+ | 1 | `She said "hello" twice.` | `She said “hello” twice.` | depth 1 → primary |
1065
+ | 3 | `'He said "no" to me,' she noted.` | `“He said ‘no’ to me,” she noted.` | depth, not width, selects the pair; the gate certifies despite both streams swapping width |
1066
+ | 4 | `He said "hi. She said "bye."` | `He said "hi. She said “bye.”` | the first mark's `closeRight` is the letter `h`, so it is `canOpen` only and cannot swallow the real quotation |
1067
+ | 5 | `Don't touch it — it's the '90s.` | ⟶ | two medial vetoes; `'90s` is `canOpen`, unmatched, handed to `apostrophe` |
1068
+ | 6 | `He is 6' 2" tall.` | ⟶ | both marks `canClose` only, both stacks empty |
1069
+ | 7b | `"""a""` | ⟶ | every mark has an identical literal neighbour, V1 drops all five |
1070
+ | 7c | `"a "" b"` | `“a "" b”` | the inner `""` is vetoed; the outer pair converts; the gate certifies |
1071
+ | 8 | `“Already curly,” he said, and "this too."` | `“Already curly,” he said, and “this too.”` | the curly pair is now a real depth-1 pair, not an ignored one |
1072
+ | A | `“He said ‘no’ twice.”` | ⟶ | already-correct nesting is a no-op |
1073
+ | M2a | `" hello"` | `“hello”` | mandate 2: `canOpen` reads `Rskip = h`; the inner run's landing is `h ∈ ALNUM`, deleted |
1074
+ | M2b | `"hello "` | `“hello”` | mirror case |
1075
+ | D1 | `" --x"` | `“ --x”` | the inner run's landing is `-`, outside `DELETE-LANDING` → deletion declined, glyphs still convert |
1076
+ | E1 | `’90s were fun,’ he said` | `“90s were fun,” he said` | **elision-vs-closer, curly input** — see §4's residual-risk note |
1077
+ | E1s | `'90s were fun,' he said` | `“90s were fun,” he said` | the straight-input control: the same corruption, pre-existing since 0.1.0 |
1078
+
1079
+ #### Listed elision veto — `en-US`, `elisionIdioms = [{ left: "rock", elided: "n", right: "roll" }]`
1080
+
1081
+ Positive — the full context matches and both marks are vetoed, so `apostrophe` (order 50) renders
1082
+ each independently:
1083
+
1084
+ | # | Input | Output | Why |
1085
+ | --- | --- | --- | --- |
1086
+ | P1 | `rock 'n' roll` | `rock ’n’ roll` | canonical form |
1087
+ | P2 | `Rock 'n' roll is great.` | `Rock ’n’ roll is great.` | sentence-initial capital — first-code-point leniency on `left` |
1088
+ | P3 | `I love rock 'n' roll.` | `I love rock ’n’ roll.` | trailing period after `right` does not block the word-boundary test |
1089
+ | P4 | `rock ’n’ roll` | ⟶ | already-curly: U+2019 is `NARROW`, the veto fires identically on the already-correct form, `apostrophe` does not act on U+2019 at all — a fixed point (§5) |
1090
+
1091
+ Negative **for the listed mechanism only** — the elided content matches but the full
1092
+ `left`/`elided`/`right` context does not, so no configured idiom fires. Since spec 1.1.0 every row
1093
+ below is nevertheless matched by the universal medial-`n` veto (§3.2), so every one converts, and
1094
+ this table's job is to show where the two mechanisms differ rather than where the marks fall
1095
+ through. N1 and N2 remain the rows that matter most: they were reported as falsely elided by the
1096
+ earlier, context-free `elisionForms` design, and 1.1.0 accepts that specific cost knowingly and
1097
+ narrowly (§3.2, "Accepted false positive") — which is not the same as readmitting the unbounded
1098
+ word list that was rejected.
1099
+
1100
+ **These five outputs were last accurate before spec 0.5.0.** The table was not updated when 0.5.0's
1101
+ veto landed, so as written it described neither 0.5.0's behaviour nor the fixtures'; the outputs
1102
+ below are 1.1.0's, verified against `spec/fixtures/en-US.json`.
1103
+
1104
+ | # | Input | Output | Why |
1105
+ | --- | --- | --- | --- |
1106
+ | N1 | `The letter 'n' is common.` | `The letter ’n’ is common.` | no `left`/`right` context at all, so no idiom fires — but the universal veto matches the bare `'n'` and `apostrophe` converts both marks. The accepted false positive of §3.2, stated there in full |
1107
+ | N2 | `He said "press 'n' now".` | `He said “press ’n’ now”.` | the same shape one level deeper; §3.2's predicate reads a mark's immediate neighbours and knows nothing of enclosing pairs, so nesting depth never reaches it |
1108
+ | N3 | `rock 'n' pop` | `rock ’n’ pop` | `right` is `pop`, not `roll` — no configured idiom matches, and since 1.1.0 none is needed |
1109
+ | N4 | `fish 'n' chips` | `fish ’n’ chips` | not sourced or listed: no `{ left: "fish", … }` entry exists in `en-US.json`, and none is required for the marks to convert |
1110
+ | N5 | `rock 'N' roll` | `rock ’N’ roll` | `elided` is matched **exactly** by the listed mechanism, so `N ≠ n` still fails *it*; the universal veto matches U+004E as well as U+006E, and converts both marks |
1111
+
1112
+ **Known, accepted residual ambiguity (not a negative-fixture guarantee).** The full-context
1113
+ requirement narrows the false-positive surface to genuine surface-form coincidence; it cannot
1114
+ see authorial intent, only code points. A document that explicitly states it means three
1115
+ separate tokens — a word, a genuinely quoted single letter, another word — and then reproduces
1116
+ the identical surface sequence for illustration still gets that final occurrence vetoed:
1117
+
1118
+ | # | Input | Output | Why |
1119
+ | --- | --- | --- | --- |
1120
+ | N6 | `The sequence is the word rock, the quoted letter 'n', and the word roll: rock 'n' roll.` | `The sequence is the word rock, the quoted letter “n”, and the word roll: rock ’n’ roll.` | **the first `'n'` is correctly NOT vetoed**, and by both mechanisms independently: the mark closing it is followed by a comma rather than an `INLINE-SPACE`, which fails the universal veto's outer test (§3.2), and `left` is `letter` rather than `rock`, which fails the listed idiom's. It pairs as an ordinary quotation. **The second `rock 'n' roll` IS vetoed and pinned as known, accepted risk, not fixed.** Under the intent the sentence itself states, this is a genuine semantic false positive — accepted as an unavoidable one for this bounded, surface-form contract, since `left`/`elided`/`right` and the single permitted `INLINE-SPACE` are all it is defined over, and none of them can encode "this is a demonstration, not an utterance of the idiom." The implementation still conforms exactly to its own matcher here: the bytes are the idiom's bytes, byte for byte, and the matcher cannot see past that. Recorded so the exposure is visible in the conformance suite rather than discovered in someone's content (§3.2's residual-ambiguity note) |
1121
+
1122
+ #### Elision vetoes across span boundaries — `html`, `markdown`
1123
+
1124
+ `modes.md` §3.2's boundary marker is not `INLINE-SPACE`, so a **context word** separated from its
1125
+ mark by an element or span boundary does not satisfy the listed idiom's outer test, and that idiom
1126
+ does not fire. The universal medial-`n` veto (§3.2) reads no context word at all — only the single
1127
+ enclosed code point and the space immediately outside each mark — so a boundary further out is
1128
+ invisible to it and it fires on all four rows. Since spec 1.1.0 the split cases therefore converge
1129
+ with the unsplit one instead of diverging from it.
1130
+
1131
+ **H1–H3's outputs were last accurate before spec 0.5.0** and are corrected here for the same
1132
+ reason as N1–N5 above; they are verified against `spec/fixtures/en-US.json`.
1133
+
1134
+ | # | Mode | Input | Output | Why |
1135
+ | --- | --- | --- | --- | --- |
1136
+ | H0 | `html` | `<p>rock 'n' roll</p>` | `<p>rock ’n’ roll</p>` | idiom whole within one text node — positive control, both mechanisms match |
1137
+ | H1 | `html` | `<p><em>rock</em> 'n' roll</p>` | `<p><em>rock</em> ’n’ roll</p>` | `left` word is inside a different span, so the *listed* idiom does not match; the universal veto never looks for a `left` word, and the opening mark's own left neighbour is still a real `INLINE-SPACE` |
1138
+ | H2 | `html` | `<p>rock 'n' <em>roll</em></p>` | `<p>rock ’n’ <em>roll</em></p>` | mirror case on `right` |
1139
+ | H3 | `markdown` (`commonmark`) | `*rock* 'n' roll\n` | `*rock* ’n’ roll\n` | `left` word is inside an emphasis span; same reasoning as H1, and the result matches `text` mode on the same characters split the same way |
1140
+
1141
+ ### `en-GB` — primary `‘ ’`, secondary `“ ”`
1142
+
1143
+ | # | Input | Output | Why |
1144
+ | --- | --- | --- | --- |
1145
+ | B2 | `‘""-"-"` | ⟶ | width drift interacting with the same-glyph adjacency veto; the gate declines to ∅ |
1146
+ | E2 | `’90s were fun,’ he said` | `‘90s were fun,’ he said` | elision-vs-closer, curly input |
1147
+ | E2s | `'90s were fun,' he said` | `‘90s were fun,’ he said` | straight-input control |
1148
+
1149
+ ### `fi` — primary `” ”` (U+201D both sides), secondary `’ ’` (U+2019 both sides)
1150
+
1151
+ | # | Input | Output | Why |
1152
+ | --- | --- | --- | --- |
1153
+ | 9 | `Hän sanoi "moi" ja lähti.` | `Hän sanoi ”moi” ja lähti.` | the pairing, not the glyph, knows which is which |
1154
+ | B | `”Hän sanoi ’moi’”, totesin.` | ⟶ | same-glyph already-correct nesting is a no-op |
1155
+ | 11 | `"Hän sanoi 'moi'", totesin.` | `”Hän sanoi ’moi’”, totesin.` | row B's input reached from straight marks |
1156
+ | U1 | `”Hän sanoi ”moi” ja lähti.` | ⟶ | unbalanced same-glyph: three `”`, the stray leading one is left alone |
1157
+ | E3 | `’90s were fun,’ he said` | `”90s were fun,” he said` | elision-vs-closer, curly input |
1158
+ | E3s | `'90s were fun,' he said` | `”90s were fun,” he said` | straight-input control |
1159
+
1160
+ ### `sv` — same shape as `fi`
1161
+
1162
+ | # | Input | Output | Why |
1163
+ | --- | --- | --- | --- |
1164
+ | E4 | `’90s were fun,’ he said` | `”90s were fun,” he said` | elision-vs-closer, curly input |
1165
+ | E4s | `'90s were fun,' he said` | `”90s were fun,” he said` | straight-input control |
1166
+
1167
+ ### `fr` — primary `« »`, `innerSpace: "nbsp"`; secondary `“ ”`, `innerSpace: "none"`
1168
+
1169
+ | # | Input | Output *(after `nbsp`)* | Why |
1170
+ | --- | --- | --- | --- |
1171
+ | 12 | `Il a dit "bonjour".` | `Il a dit «⍽bonjour⍽».` | `quotes` emits the glyphs; `nbsp` the U+00A0 |
1172
+ | 13 | `Il a dit «␣bonjour␣».` | `Il a dit «⍽bonjour⍽».` | mandate 2's second half: the pair *forms on the first application* and produces no edit from `quotes` because the glyphs are already right |
1173
+ | 3a | `«␣**"` | `«⍽**⍽»` *(after `nbsp`)* | width drift interacting with a run-1 `nbsp` insertion; `quotes` alone certifies the pair with one edit |
1174
+
1175
+ ### `fr-CA`, `html` mode
1176
+
1177
+ | # | Input | Output | Why |
1178
+ | --- | --- | --- | --- |
1179
+ | 3b | `"<p class="x">«[t](u)` | ⟶ | the mode adapter splices a `MARKER` between the two text spans; `«`'s `closeRight` reads `Rskip` (because `« ∈ SPACE-RIGHT`) and finds `[`, so it is never `canClose` and no pair forms, on either application |
1180
+
1181
+ ### `ru` — primary `« »`, secondary `„ “`
1182
+
1183
+ | # | Input | Output | Why |
1184
+ | --- | --- | --- | --- |
1185
+ | 14 | `Он сказал: "это 'моё' дело".` | `Он сказал: «это „моё“ дело».` | depths 1 and 2 |
1186
+ | B1 | `'"‘` | ⟶ | width drift with no candidate left unmatched to be swallowed; the gate declines to ∅ in one round |
1187
+
1188
+ ### `el` — primary `« »`, secondary `“ ”` (both `WIDE`)
1189
+
1190
+ | # | Input | Output | Why |
1191
+ | --- | --- | --- | --- |
1192
+ | 16 | `"Είπε 'όχι' σε μένα", σημείωσε.` | `«Είπε “όχι” σε μένα», σημείωσε.` | width drift with no gate intervention — evidence the gate's cost is near zero |
1193
+ | 18 | `Είπε "σ' αυτό το βιβλίο" χθες.` | `Είπε «σ' αυτό το βιβλίο» χθες.` | the elision is `NARROW`, the quotation `WIDE`, different stacks — the elision survives untouched for `apostrophe` to convert on the next rule |
1194
+
1195
+ ### `de-DE` — primary `„ “`, secondary `‚ ‘`
1196
+
1197
+ | # | Input | Output | Why |
1198
+ | --- | --- | --- | --- |
1199
+ | C | `Er sagte «Wort» leise.` | `Er sagte „Wort“ leise.` | foreign guillemets convert |
1200
+ | C2 | `Er sagte « Wort » leise.` | `Er sagte „Wort“ leise.` | same, plus both inner runs deleted (landings `W` and `t`) |
1201
+
1202
+ ### Adversarial sweep — `ru`/`el`/`fr`/`fr-CA` (§7 item 9)
1203
+
1204
+ These four locales' primary and secondary pairs are **both `WIDE`**, so the width-keyed stack
1205
+ split does not independently separate a re-scanned primary stream from a re-scanned secondary
1206
+ stream the way it does for `en-US`/`en-GB`/`fi`/`sv`. An exhaustive sweep over
1207
+ `{«, », „, “, ", ', a, ␣}` (the locale's own glyphs plus the shared control alphabet) to length 5
1208
+ — 149,796 inputs across the four locales — found **zero** idempotency failures; the medial and
1209
+ same-glyph vetoes plus the certification gate together protect already-curly interleaved text in
1210
+ this family, even without an independent width partition. Recorded as measured, not asserted.
1211
+
1212
+ ---
1213
+
1214
+ ## 7. Open questions, and what could not be closed
1215
+
1216
+ Ordered by how much this matters.
1217
+
1218
+ 1. **Mandate 1's false-positive surface is genuinely larger, and M4 is where it will show.** A
1219
+ lone `«` can now pair with a distant `"` (row 3a converts an input 0.1.0 left alone). Any
1220
+ document with an odd stray typographic quote — a `»` used as a bullet, a `’` used as a prime,
1221
+ a `«` in a citation — can recruit a partner from arbitrarily far away, and unlike a straight
1222
+ mark the result is not visibly "unfinished" to a proofreader. This is a **product** risk, not
1223
+ an idempotency one; run the M4 corpus against this change specifically before anything else.
1224
+ 2. **The `symbols` (R₇) discharge is a case analysis**, not structural. Frequency: low (needs a
1225
+ quote mark literally adjacent to a `(c)`/`(r)`/`(tm)` span), but not closed structurally.
1226
+ 3. **The elision-vs-closer tradeoff (§4, §5 Corollary A1, §6 rows E1–E4) is verified, not
1227
+ eliminated.** A leading elision that has already been rendered U+2019 and finds no partner of
1228
+ its own can still be recruited as an opener by an unrelated closer elsewhere in the text —
1229
+ that is a different mechanism from item 8's closed-idiom case below (`quotes.elisionIdioms`,
1230
+ spec 0.4.0) and is **not** closed by it: the veto only fires on a *listed idiom's full
1231
+ left/elided/right context* found together, not on a lone leading elision separated from any
1232
+ partner by arbitrary distance. No fix is proposed here.
1233
+ 4. **The gate's forced-removal clause's necessity is asserted, not demonstrated with a concrete
1234
+ witness.** ~850,000 swept inputs did not trigger it. Implemented anyway (free insurance); if a
1235
+ future witness fires it, record it as a fixture.
1236
+ 5. **Declined deletions leave visibly sloppy output** (`" --x"` → `“ --x”`, row D1). This is the
1237
+ correct side to be wrong on (the alternative is a confirmed `I₃` violation) but is a real,
1238
+ acknowledged gap in mandate 2's coverage.
1239
+ 6. **`" "` and `«»` are prevented by an explicit vacuity guard, not by construction**, unlike
1240
+ 0.1.0 where an empty pair was structurally impossible. Weaker footing than 0.1.0 had, though
1241
+ the gate catches any resulting instability.
1242
+ 7. **Worst case is `O(n²)`** (gate rounds × pass cost). A pathological span of alternating quote
1243
+ marks is quadratic; cap rounds if this matters in practice.
1244
+ 8. **`rock 'n' roll` — closed for every locale in spec 1.1.0 by the universal medial-`n` veto
1245
+ (§3.2), by operator decision that the idiom is international rather than a locale fact.** It
1246
+ was closed for `en-US` alone in spec 0.4.0 via `quotes.elisionIdioms =
1247
+ [{ left: "rock", elided: "n", right: "roll" }]` (Chicago Manual of Style Online on the
1248
+ mark's function, American Heritage Dictionary on the spaced variant's existence — see
1249
+ `en-US.json` `sources`), and that entry is retained and still cited; it now vetoes a strict
1250
+ subset of what the universal veto vetoes.
1251
+ **What 1.1.0 knowingly accepts in exchange** is the cost that sank a first design during the
1252
+ development of 0.4.0 (a bare `elisionForms = ["n"]` word list, matching only the elided
1253
+ content, prototyped and rejected before any release or commit — it never shipped and no user
1254
+ ever ran it): ordinary quotations of the letter *n* are elided — `The letter 'n' is common.`,
1255
+ `He said "press 'n' now".`, §6 rows N1/N2. The reasoning that rejected it in 0.4.0 was not
1256
+ wrong; what changed is the alternative it was measured against. In 0.4.0 the alternative was
1257
+ `elisionIdioms`' three-part contract, which is strictly better for `en-US`. From 0.5.0 the
1258
+ de-facto alternative for the other nine locales was the general 1–3-`LETTER` shape veto, which
1259
+ bought the same protection by declining **every** short quotation in **every** locale —
1260
+ a much larger and commoner error class than quoting a single letter, and one this repository's
1261
+ own promo examples tripped over in `ru`, `fr`, `fr-CA`, `fi` and `sv`. 1.1.0 takes the smaller
1262
+ error knowingly, and N1/N2 stay pinned so it is visible rather than forgotten.
1263
+ The evidentiary notes that governed per-locale entries are retained because they still govern
1264
+ `elisionIdioms` itself, which is unchanged: `en-GB` has no independent citation and carries an
1265
+ extra risk the other locales do not, since its primary pair *is* the single quote, so `'n'` is
1266
+ more plausible as genuine quoted dialogue there than in a locale whose primary pair is double;
1267
+ and `fi`/`sv` write the idiom **closed up** (`rock'n'roll`) in their own dictionaries, so a
1268
+ spaced-form entry there would be evidenced-inert. Neither observation blocks the universal
1269
+ veto, which rests on the operator decision rather than on per-locale citation. **Still open,
1270
+ all locales:** `en-GB` single-first (now more consequential under mandate 1); quotation across
1271
+ a paragraph boundary; nesting deeper than 2; primes; surviving-mark invisibility to the
1272
+ conformance matrix.
1273
+ 9. **`ru`/`el`/`fr`/`fr-CA` already-curly protection is narrower than 0.1.0's** for these
1274
+ same-width-primary/secondary locales specifically — the two-stack split no longer
1275
+ independently protects already-curly text there the way it protects freshly-typed straight
1276
+ text. §6's dedicated adversarial sweep found no failures, but the protection that remains is
1277
+ the vetoes and the gate, not an independent structural guarantee.
1278
+
1279
+ ---
1280
+
1281
+ ## History
1282
+
1283
+ 0.1.0 shipped with the four-pass, `STRAIGHT`-only design whose idempotency proof is quoted and
1284
+ withdrawn in §0/§5 above. 0.2.0 was `dashes`' version bump and did not touch this file. 0.3.0 is
1285
+ the revision described through most of this document: mandates 1 and 2, the five-pass
1286
+ algorithm, the certification gate, and the locale-schema constraints in §2.1.
1287
+
1288
+ 0.4.0 adds the **listed elision veto** (§2, §3.2): `quotes.elisionIdioms`, a locale-data list of
1289
+ `{ left, elided, right }` triples that declines to pair two `NARROW` marks as a quotation when
1290
+ the full surrounding context matches a configured idiom, closing the `rock 'n' roll` defect
1291
+ (§7 item 8) for `en-US`. A first design considered during the same development — a bare
1292
+ `elisionForms` word list matching only the elided content — was reviewed and rejected before
1293
+ any commit or release once it was found to falsely elide ordinary quotations of a single letter
1294
+ (`The letter 'n' is common.`); it never shipped. §6's `N1`/`N2` rows and this history entry
1295
+ exist so that design is not silently reintroduced.
1296
+
1297
+ 0.5.0 adds the **general ambiguous-medial-span veto** (§3.2), closing the same class of defect for
1298
+ every locale without a cited `elisionIdioms` entry — not by inferring an idiom, but by preserving
1299
+ the author's straight ASCII marks unconverted. The shared predicate this and `apostrophe`
1300
+ (order 50) both consult lives in one module, `src/rules/quote-ambiguity.ts` (JS reference
1301
+ implementation), specifically so `apostrophe`'s own structural case ladder cannot independently
1302
+ curl a mark this rule has deliberately left alone — see `apostrophe.md` §3.4 for why that
1303
+ composition risk was real, not hypothetical.
1304
+
1305
+ 1.1.0 **replaces 0.5.0's veto with the universal medial-`n` elision veto** (§3.2), on the operator
1306
+ decision of 2026-09-09 that `rock 'n' roll` and `rock'n'roll` are international and take U+2019 in
1307
+ every locale. Three things change together, and none of them works without the other two:
1308
+
1309
+ - **The shape narrows** from 1–3 `LETTER` code points to exactly one code point, U+006E or U+004E.
1310
+ 0.5.0's shape could not distinguish the idiom from an ordinary short nested quotation and
1311
+ declined both, which is what made `«это 'моё' дело»` and `“He said 'no' to me,”` come out with
1312
+ the inner marks unconverted. It was also normalization-dependent, since `LETTER` includes `Mn`.
1313
+ - **The outcome inverts** from preserve to convert: matched marks are no longer held back from
1314
+ `apostrophe`, which renders each by its ordinary case ladder. 0.5.0's preserve set
1315
+ (`computePreserveIndices`, `apostrophe.md` §3.4) is withdrawn along with the reason it existed.
1316
+ - **The predicate widens** from straight ASCII to the whole `NARROW` class. This is forced by the
1317
+ second change, not chosen: a veto that produces U+2019 must recognise U+2019, or its own output
1318
+ pairs as a quotation on the next pipeline pass. A straight-only predicate was implemented first
1319
+ and measured to do exactly that — `rock ’n’ roll` → `rock «n» roll` in `ru`, `rock ”n” roll` in
1320
+ `fi` — an idempotency violation and so a release blocker.
1321
+
1322
+ The cost accepted, knowingly and with §6's N1/N2 rows kept as its permanent witnesses, is that a
1323
+ genuine quotation of the letter *n* is elided. §7 item 8 records why that is the smaller error
1324
+ than the class 0.5.0 traded it for.