polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,1324 @@
|
|
|
1
|
+
# Rule: `quotes`
|
|
2
|
+
|
|
3
|
+
**Order:** 40. **Default:** on. **Modes:** text, html, markdown, yaml.
|
|
4
|
+
**Spec version:** 1.1.0 (0.4.1 for everything except the universal medial-`n` elision veto
|
|
5
|
+
described in §3.2 and the History section below).
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 0. What changed from 0.1.0, and why it is a rewrite rather than a patch
|
|
10
|
+
|
|
11
|
+
0.1.0's whole architecture rested on one sentence: *the rule reasons about the straight input
|
|
12
|
+
marks and never about the curly output glyphs.* Two operator decisions withdraw that sentence:
|
|
13
|
+
|
|
14
|
+
1. **Any existing quote glyph is a re-typesetting candidate**, this locale's own or a foreign
|
|
15
|
+
one. `«Wort»` in German prose becomes `„Wort“`. Same principle as `dashes` §3.2 step 2a
|
|
16
|
+
applies to dash length: a quote glyph's identity in ordinary prose is at least as often a
|
|
17
|
+
copy-paste artefact or unfamiliarity as a deliberate choice, and an author who wants their own
|
|
18
|
+
typography untouched has always had the option of not running the pipeline. **0.1.0 §3.8 is
|
|
19
|
+
withdrawn.**
|
|
20
|
+
2. **A space touching a quote mark is sloppiness, not evidence.** `" hello"` is `“hello”`; the
|
|
21
|
+
space is deleted. `« bonjour »` is a formed pair on the first application, not only after
|
|
22
|
+
`nbsp` has touched it.
|
|
23
|
+
|
|
24
|
+
Once every quote glyph is a candidate, **the rule's input alphabet equals its output alphabet**,
|
|
25
|
+
and 0.1.0's idempotency proof — whose first line is "a converted position is not a candidate in
|
|
26
|
+
pass 1 of the second run" — is not weakened, it is **false**. Three ideas replace it, and each
|
|
27
|
+
closes one class of the four defects that motivated this revision:
|
|
28
|
+
|
|
29
|
+
| Idea | Closes |
|
|
30
|
+
| --- | --- |
|
|
31
|
+
| **Glyph-blindness (Lemma A, §5).** Every quote mark belongs to `QUOTEMARK`, which is a member of *both* `OPENISH` and `CLOSEISH` and exempt from `canOpen`'s closeish rejection. A candidate's verdict never depends on *which* quote glyph its neighbour is. | neighbour drift; `apostrophe` (R₆) becomes structurally invisible as a neighbour |
|
|
32
|
+
| **Directional space-skipping (Lemma B, §5).** `canOpen` skips a space on its inner (right) side, `canClose` on its inner (left) side, unconditionally; each reads its *outer* side literally, except at the two positions `nbsp` can insert — derived from locale data, not hand-picked. | mandate 2; a run-1/run-2 asymmetry across an `nbsp`-inserted space; an HTML span boundary the same shape exposed |
|
|
33
|
+
| **Certification (the gate, §3.5).** The accepted pairing is *checked*, not proved: render the hypothetical output, re-run passes 1–2 on it, and decline pairs until the re-run reproduces the accepted set exactly. | width drift interacting with an unmatched candidate — the one remaining non-invariance, and provably non-local |
|
|
34
|
+
|
|
35
|
+
The division of labour is the point. **Lemmas A and B cover what the gate cannot see** — edits
|
|
36
|
+
made by *other* rules after `quotes` has run. **The gate covers everything `quotes` itself
|
|
37
|
+
does** — instabilities in the vetoes, the width partition, the depth assignment — without
|
|
38
|
+
anyone having to enumerate them by hand. A future change to the capability tests can break
|
|
39
|
+
quality; it cannot break idempotency, because the gate checks the actual re-derivation rather
|
|
40
|
+
than trusting an argument about it.
|
|
41
|
+
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
## 1. Purpose
|
|
45
|
+
|
|
46
|
+
`quotes` resolves every quotation mark in the text — typewriter (U+0022, U+0027) or typographic
|
|
47
|
+
(`«`, `“`, `„`, `‘`, `”`, …), the locale's own or a foreign locale's — into the pair of glyphs
|
|
48
|
+
the locale prescribes, alternating between the primary and secondary pairs by nesting depth, and
|
|
49
|
+
normalising the spacing immediately inside each pair.
|
|
50
|
+
|
|
51
|
+
---
|
|
52
|
+
|
|
53
|
+
## 2. Locale data consumed
|
|
54
|
+
|
|
55
|
+
- `quotes.primary.open` / `.close` / `.innerSpace`
|
|
56
|
+
- `quotes.secondary.open` / `.close` / `.innerSpace`
|
|
57
|
+
- `quotes.elisionIdioms` (spec 0.4.0) — an array of `{ left, elided, right }` triples, each a
|
|
58
|
+
literal string, consulted only by the **listed elision veto** (§3.2) to **decline** pairing
|
|
59
|
+
two `NARROW` marks as a quotation when the full context matches: the elided content between
|
|
60
|
+
the marks, the word immediately before the opening mark, and the word immediately after the
|
|
61
|
+
closing mark (`rock 'n' roll`'s `{ left: "rock", elided: "n", right: "roll" }`). May be
|
|
62
|
+
empty, and empty is the default for a locale with no verified idiom of this shape. A
|
|
63
|
+
decline-only list can never widen what this rule pairs, only narrow it — the same principle
|
|
64
|
+
stated normatively in `nbsp.md` §2.1 and governing every list-valued field in
|
|
65
|
+
`locale.schema.json`.
|
|
66
|
+
|
|
67
|
+
**All three fields are required, and matching `elided` alone is deliberately not a supported
|
|
68
|
+
shape.** A bare word list cannot distinguish an idiom from an arbitrary quotation of the same
|
|
69
|
+
word — `The letter 'n' is common.` and `He said "press 'n' now".` are both genuine
|
|
70
|
+
quotations of the letter *n*, and a list keyed on `elided` alone would falsely elide both.
|
|
71
|
+
The surrounding context is the evidence that makes `rock 'n' roll` an idiom and these two
|
|
72
|
+
sentences ordinary prose; §3.2 states exactly how much of it must match.
|
|
73
|
+
|
|
74
|
+
**`left` and `right` must consist entirely of `LETTER` code points** — §3.2's Word definition
|
|
75
|
+
is a maximal `LETTER` run, and a `left`/`right` entry is normatively *authored as* that word.
|
|
76
|
+
This is not a limit the matcher itself enforces at match time: the comparison in §3.2 is a
|
|
77
|
+
literal code-point equality test, so a configured entry containing a digit or punctuation
|
|
78
|
+
code point would still match those exact bytes wherever they occur, not fail to match at all.
|
|
79
|
+
The constraint exists precisely because the matcher does not check it — an unvalidated entry
|
|
80
|
+
could make the veto fire on text no citation ever attested, silently exceeding what the
|
|
81
|
+
normative Word definition authorises. `scripts/validate-spec.mjs` and `scripts/gen-locales.mjs`
|
|
82
|
+
therefore reject such an entry at the data boundary, before it can be authored into a locale
|
|
83
|
+
file at all. `elided` carries no such constraint: its match is a literal equality test over
|
|
84
|
+
the raw code points between the two marks, not a `LETTER`-run walk, so it may contain any
|
|
85
|
+
non-empty literal. Both validators read the identical pinned table `src/engine/unicode.ts`
|
|
86
|
+
imports (`scripts/lib/is-letter.mjs`), so validation and the runtime engine cannot classify a
|
|
87
|
+
code point differently.
|
|
88
|
+
|
|
89
|
+
`innerSpace` is read and acted on in **one direction only**: `quotes` *deletes* an inner space
|
|
90
|
+
when the assigned pair's `innerSpace` is `"none"`, and never inserts or converts one. Insertion
|
|
91
|
+
and conversion stay with `nbsp` (order 70), exactly as 0.1.0 always said. See §3.7.
|
|
92
|
+
|
|
93
|
+
### 2.1 Two new normative locale constraints
|
|
94
|
+
|
|
95
|
+
Both are needed by §5's proof. **All ten shipped locales already satisfy both**, verified against
|
|
96
|
+
`spec/locales/*.json`:
|
|
97
|
+
|
|
98
|
+
- **Q-W — a pair is width-homogeneous.** For each of `primary` and `secondary`, `open` and
|
|
99
|
+
`close` must have the same quote *width* (§3.1). Without it a formed pair straddles both
|
|
100
|
+
stacks after rendering and the gate would decline every pair in a same-glyph document.
|
|
101
|
+
- **Q-A — a spaced pair must not use the apostrophe glyph.** If `innerSpace ≠ "none"`, neither
|
|
102
|
+
`open` nor `close` may be U+2019. Without it `apostrophe`'s output would land in a
|
|
103
|
+
`nbsp`-insertion-adjacent skip set and Lemma A's corollary would break.
|
|
104
|
+
|
|
105
|
+
One constraint belongs to `nbsp`'s data and is stated normatively in `nbsp.md` §2:
|
|
106
|
+
|
|
107
|
+
- **Q-P — every code point in `nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must
|
|
108
|
+
be a member of this rule's `CLOSEISH`.** This is what makes Lemma B cover N1/N2 as well as N8.
|
|
109
|
+
The shipped lists (`:`, `;`, `!`, `?` in `fr`/`fr-CA`, empty elsewhere) satisfy it.
|
|
110
|
+
|
|
111
|
+
**Evidence for locale data.** A locale attests **membership**; the rule owns the **mechanism**.
|
|
112
|
+
Stated normatively in `nbsp.md` §2.1; it governs every locale field this rule reads.
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## 3. Algorithm
|
|
117
|
+
|
|
118
|
+
Input is a code-point array `cp[0 … n-1]`. Five passes plus an emit; no backtracking inside a
|
|
119
|
+
pass, no regular expression, no native-string indexing.
|
|
120
|
+
|
|
121
|
+
### 3.1 Character classes
|
|
122
|
+
|
|
123
|
+
| Class | Members |
|
|
124
|
+
| --- | --- |
|
|
125
|
+
| `DQ` | U+0022 |
|
|
126
|
+
| `SQ` | U+0027 |
|
|
127
|
+
| `STRAIGHT` | `DQ` ∪ `SQ` |
|
|
128
|
+
| `WIDE` | U+0022, U+00AB, U+00BB, U+201C, U+201D, U+201E, U+201F, U+301D, U+301E, U+301F |
|
|
129
|
+
| `NARROW` | U+0027, U+2018, U+2019, U+201A, U+201B, U+2039, U+203A |
|
|
130
|
+
| `QUOTEMARK` | `WIDE` ∪ `NARROW` (disjoint) |
|
|
131
|
+
| `DIGIT` | U+0030–U+0039 |
|
|
132
|
+
| `LETTER` | general categories `Lu Ll Lt Lm Lo Mn Mc Me` |
|
|
133
|
+
| `ALNUM` | `LETTER` ∪ `DIGIT` |
|
|
134
|
+
| `BREAK` | U+000A, U+000D, U+000B, U+000C, U+0085, U+2028, U+2029, plus `LINE_MARKER` |
|
|
135
|
+
| `INLINE-SPACE` | U+0020, U+0009, U+00A0, U+202F, U+2007, U+2009, U+200A |
|
|
136
|
+
| `SPACELIKE` | `INLINE-SPACE` ∪ `BREAK` |
|
|
137
|
+
| `OPENISH` | U+0028 `(` U+005B `[` U+007B `{`, `MARKER`, ∪ `QUOTEMARK` |
|
|
138
|
+
| `CLOSEISH` | U+0029, U+005D, U+007D, U+002C, U+002E, U+003B, U+003A, U+0021, U+003F, U+2026, U+2013, U+2014, `MARKER`, ∪ `QUOTEMARK` |
|
|
139
|
+
| `DASHISH` | U+002D, U+2011, U+2013, U+2014 |
|
|
140
|
+
| `DELETE-LANDING` | `ALNUM` ∪ `QUOTEMARK` |
|
|
141
|
+
| `NONE` | the pseudo-class "index out of range" |
|
|
142
|
+
|
|
143
|
+
**`QUOTEMARK` is a member of both `OPENISH` and `CLOSEISH`, and is exempt from `canOpen`'s
|
|
144
|
+
closeish rejection** — the same dual membership `STRAIGHT` had in 0.1.0, widened to the whole
|
|
145
|
+
class. This is Lemma A's entire mechanism and must not be simplified back to per-glyph lists.
|
|
146
|
+
`MARKER` keeps its 0.1.0 treatment (`modes.md` §3.3): in both classes, exempt from the
|
|
147
|
+
rejection, and **not** in `SPACELIKE`, so no skip walk ever crosses a span boundary.
|
|
148
|
+
|
|
149
|
+
**Deliberately excluded from `QUOTEMARK`:** U+2032/U+2033 (primes — a separate rule with its own
|
|
150
|
+
false-positive profile, §7), U+02BC (a `Lm` letter), U+0060 and U+00B4. Neither candidates nor
|
|
151
|
+
produced.
|
|
152
|
+
|
|
153
|
+
**Unicode version.** As 0.1.0: the pinned UCD in `spec/UNICODE` is normative for the derived
|
|
154
|
+
tables, not for the host runtime (`pipeline-idempotency.md` §6a).
|
|
155
|
+
|
|
156
|
+
**A declared quote glyph must not be in `SPACELIKE`, `ALNUM` or `STRAIGHT`**, and by Q-W must
|
|
157
|
+
share its pair's `WIDE`/`NARROW` bucket with its partner glyph. `singleChar` in
|
|
158
|
+
`locale.schema.json` enforces the `STRAIGHT` exclusion.
|
|
159
|
+
|
|
160
|
+
### 3.1a Locale-derived skip sets
|
|
161
|
+
|
|
162
|
+
Computed once per call, from locale data alone:
|
|
163
|
+
|
|
164
|
+
```
|
|
165
|
+
SPACE-RIGHT = { p.open : p ∈ {primary, secondary}, p.innerSpace ≠ "none", p.open ≠ p.close }
|
|
166
|
+
SPACE-LEFT = { p.close : p ∈ {primary, secondary}, p.innerSpace ≠ "none", p.open ≠ p.close }
|
|
167
|
+
```
|
|
168
|
+
|
|
169
|
+
These are **exactly the positions at which `nbsp` N8 can insert a space** (`nbsp.md` §3.10;
|
|
170
|
+
`nbsp.ts` `quotesSubRule` skips a pair whose `open` equals its `close`, and skips
|
|
171
|
+
`innerSpace: "none"`). For the ten shipped locales: `SPACE-RIGHT = {«}` and `SPACE-LEFT = {»}`
|
|
172
|
+
in `fr` and `fr-CA`; both empty everywhere else. The sets are derived from the *reason*, not
|
|
173
|
+
hand-written — if a locale gains a spaced pair, they follow automatically and Lemma B keeps
|
|
174
|
+
holding.
|
|
175
|
+
|
|
176
|
+
### 3.2 Pass 1 — collect and classify candidates
|
|
177
|
+
|
|
178
|
+
Walk `i` from `0` to `n-1`. Skip unless `cp[i] ∈ QUOTEMARK`. Write `g = cp[i]` and compute four
|
|
179
|
+
neighbour reads:
|
|
180
|
+
|
|
181
|
+
- `Llit = cp[i-1]`, or `NONE`; `Rlit = cp[i+1]`, or `NONE`;
|
|
182
|
+
- `Lskip` = the first `cp[j]`, `j < i` descending, with `cp[j] ∉ INLINE-SPACE`; `NONE` if the
|
|
183
|
+
walk leaves the array;
|
|
184
|
+
- `Rskip` = symmetric to the right.
|
|
185
|
+
|
|
186
|
+
`MARKER` and every `BREAK` stop a skip walk, because neither is in `INLINE-SPACE` — a skip never
|
|
187
|
+
crosses a span boundary or a line terminator.
|
|
188
|
+
|
|
189
|
+
Then two *directed* reads:
|
|
190
|
+
|
|
191
|
+
- `openLeft` = `Lskip` if `g ∈ SPACE-LEFT`, else `Llit`;
|
|
192
|
+
- `closeRight` = `Rskip` if `g ∈ SPACE-RIGHT`, else `Rlit`.
|
|
193
|
+
|
|
194
|
+
The two capabilities:
|
|
195
|
+
|
|
196
|
+
```
|
|
197
|
+
canOpen ⟺ ( openLeft = NONE ∨ openLeft ∈ SPACELIKE ∨ openLeft ∈ OPENISH ∨ openLeft ∈ DASHISH )
|
|
198
|
+
∧ ( Rskip ≠ NONE ∧ Rskip ∉ SPACELIKE
|
|
199
|
+
∧ ( Rskip ∉ CLOSEISH ∨ Rskip ∈ QUOTEMARK ∨ Rskip = MARKER ) )
|
|
200
|
+
|
|
201
|
+
canClose ⟺ ( Lskip ≠ NONE ∧ Lskip ∉ SPACELIKE )
|
|
202
|
+
∧ ( closeRight = NONE ∨ closeRight ∈ SPACELIKE ∨ closeRight ∈ CLOSEISH ∨ closeRight ∈ DASHISH )
|
|
203
|
+
```
|
|
204
|
+
|
|
205
|
+
`Rskip`/`Lskip` can still be a `BREAK` (a skip walk stops there), which is why both
|
|
206
|
+
`∉ SPACELIKE` tests are retained: a mark at a line end does not open across the break, a mark at
|
|
207
|
+
a line start does not close back over it.
|
|
208
|
+
|
|
209
|
+
**Why this exact asymmetry.**
|
|
210
|
+
|
|
211
|
+
- `canOpen` skips right and `canClose` skips left: each capability treats its *inner* side — the
|
|
212
|
+
side facing the quoted text under that hypothesis — as noise. That is mandate 2, stated once
|
|
213
|
+
and applied uniformly, unconditionally, on every locale.
|
|
214
|
+
- Each capability reads its *outer* side literally. The outer space is not noise, it is the
|
|
215
|
+
evidence. `"hi" to me` is a closing mark **because** a space follows it; skipping that space
|
|
216
|
+
(reading `t`) would destroy the commonest closing shape in every language. `He said "hi. She
|
|
217
|
+
said "bye."` is preserved by exactly this: the first mark's `closeRight` is the letter `h`, so
|
|
218
|
+
it is `canOpen` only, and the stray mark cannot swallow the real quotation.
|
|
219
|
+
- The two exceptions to "outer side is literal" are *derived*, not chosen: `nbsp` can insert on
|
|
220
|
+
the right of a `SPACE-RIGHT` glyph and the left of a `SPACE-LEFT` glyph, and those are the only
|
|
221
|
+
two outer reads it can reach. Making exactly those two reads skip is what makes every verdict
|
|
222
|
+
inert to `nbsp` (Lemma B).
|
|
223
|
+
|
|
224
|
+
**Medial-elision veto** (`NARROW` marks only, literal reads):
|
|
225
|
+
|
|
226
|
+
> if `g ∈ NARROW` and `Llit ∈ ALNUM` and `Rlit ∈ ALNUM`, set both capabilities false.
|
|
227
|
+
|
|
228
|
+
`don't`, `l'été`, `O'Brien`, `1990's` — and, on a second pipeline pass, `don't` with U+2019,
|
|
229
|
+
because `apostrophe` has converted the mark and U+2019 is also `NARROW`. Widening from `SQ` to
|
|
230
|
+
`NARROW` is what makes `apostrophe`'s output inert here.
|
|
231
|
+
|
|
232
|
+
**Listed elision veto** (spec 0.4.0, `NARROW` marks only, locale data `quotes.elisionIdioms`,
|
|
233
|
+
literal reads). **Matching the elided content alone is not this veto's shape.** A bare word
|
|
234
|
+
list cannot distinguish `rock 'n' roll` from an arbitrary quotation of the same word —
|
|
235
|
+
`The letter 'n' is common.` and `He said "press 'n' now".` are both genuine quotations that a
|
|
236
|
+
content-only match would falsely elide (this was tried and rejected in an earlier revision of
|
|
237
|
+
this field; both sentences are pinned as negative fixtures, `spec/fixtures/en-US.json`). The
|
|
238
|
+
veto therefore requires the **full context**: the elided content, and a literal word on each
|
|
239
|
+
outer side of the two marks.
|
|
240
|
+
|
|
241
|
+
> **Word.** A maximal run of `LETTER` code points. Its outer boundary — the code point
|
|
242
|
+
> immediately before its first code point, or immediately after its last — is either `NONE`
|
|
243
|
+
> (document or span start/end) or **not** `LETTER` and **not** `DIGIT`. This is a whole-word
|
|
244
|
+
> test: `"MyRock 'n' roll"` does not match a `left` of `"Rock"`, because the code point before
|
|
245
|
+
> `R` is `y`, a `LETTER`.
|
|
246
|
+
>
|
|
247
|
+
> **The permitted space.** Exactly one code point of `INLINE-SPACE` (`spec 3.1`) separates a
|
|
248
|
+
> context word from its adjacent mark — not `NONE`, not `OPENISH`, not `DASHISH`, unlike
|
|
249
|
+
> `canOpen`/`canClose`'s own outer tests above. The veto needs an actual word to anchor to, so
|
|
250
|
+
> a mark at document or span start, or immediately after an opening bracket, has no `left` word
|
|
251
|
+
> and cannot be vetoed on that side — §3.2's `MARKER` (`modes.md` §3.3) is not `INLINE-SPACE`
|
|
252
|
+
> either, so a context word split from a mark by an element or span boundary does not match
|
|
253
|
+
> (`<em>rock</em> 'n' roll` is unaffected; §6 rows H1–H3).
|
|
254
|
+
>
|
|
255
|
+
> Computed once, before capabilities are assigned, over every index `i` with `cp[i] ∈ NARROW`:
|
|
256
|
+
> for each entry `{ left, elided, right } ∈ quotes.elisionIdioms`, of `k` code points in
|
|
257
|
+
> `elided`: if `cp[i+1 … i+k]` equals `elided` **exactly**, code point for code point — no case
|
|
258
|
+
> leniency, on the elided content — and `cp[i+k+1] ∈ NARROW` (call this index `j`), and:
|
|
259
|
+
>
|
|
260
|
+
> - `cp[i-1]` is a single `INLINE-SPACE` code point, and the word ending there equals `left`;
|
|
261
|
+
> - `cp[j+1]` is a single `INLINE-SPACE` code point, and the word starting after it equals `right`;
|
|
262
|
+
>
|
|
263
|
+
> then both `i` and `j` are **vetoed**: both capabilities of `i` and both capabilities of `j`
|
|
264
|
+
> are set to `false`, overriding every other test in this section — including capabilities
|
|
265
|
+
> computed for `j` when the main walk reaches it.
|
|
266
|
+
>
|
|
267
|
+
> **Exact matching, with one narrow exception.** `elided` is always matched exactly, code point
|
|
268
|
+
> for code point — no leniency, ever, on the elided content itself. `left` and `right` are
|
|
269
|
+
> matched exactly **except their first code point**, which is compared case-insensitively —
|
|
270
|
+
> implemented directly by code point (ASCII `A`–`Z` folds to `a`–`z`; every other code point,
|
|
271
|
+
> including every non-ASCII letter, must match exactly), never through a platform locale
|
|
272
|
+
> function (`ARCHITECTURE.md` §4.4). This is the same leniency `nbsp.md` §3.5's
|
|
273
|
+
> `afterShortWords` already uses, for the same reason: `Rock 'n' roll is great.` (sentence-
|
|
274
|
+
> initial capital) is exactly as much the idiom as `rock 'n' roll` is, and `rock 'N' roll` —
|
|
275
|
+
> case varying anywhere past the first code point — is not: it fails the exact `elided` test in
|
|
276
|
+
> this specific idiom (`elided = "n"`, and `N ≠ n`) and is correctly left alone (§6 row N5).
|
|
277
|
+
|
|
278
|
+
`rock 'n' roll` with `elisionIdioms = [{ left: "rock", elided: "n", right: "roll" }]`: the
|
|
279
|
+
leading mark's `left` word is `rock`, the elided run spells exactly `n`, the trailing mark's
|
|
280
|
+
`right` word is `roll` — all three match, both marks are vetoed, survive pass 1 as U+0027, and
|
|
281
|
+
`apostrophe` (order 50) renders each independently: leading elision (`apostrophe.md` §3.3 case
|
|
282
|
+
4) on the first, trailing elision (case 3) on the second, giving `rock ’n’ roll`. Without the
|
|
283
|
+
veto both marks are `canOpen`/`canClose` candidates in their own right — the leading mark has
|
|
284
|
+
no `ALNUM` on its left so the medial veto above does not apply to it, and likewise for the
|
|
285
|
+
trailing mark's right — and pass 2 pairs them as an ordinary quotation.
|
|
286
|
+
|
|
287
|
+
**This is a bounded, literal check, not a heuristic.** The lookahead from `i` and `j` is
|
|
288
|
+
bounded by the total length of `left` + `elided` + `right` for each configured idiom — no
|
|
289
|
+
regex, no priorities, no open-ended word class; the same shape `nbsp.md`'s
|
|
290
|
+
`beforeNumber`/`beforeWord` literal matching already uses on one side of a mark
|
|
291
|
+
(`nbsp.md` §3.11), applied here to both sides of a pair of marks at once. It can only ever
|
|
292
|
+
**prevent** a pairing pass 2 would otherwise have formed — an empty `quotes.elisionIdioms`
|
|
293
|
+
makes it a total no-op, and every `NARROW` mark is classified exactly as it was before spec
|
|
294
|
+
0.4.0.
|
|
295
|
+
|
|
296
|
+
**Behaviour at punctuation and document boundaries.** Trailing punctuation after `right` (a
|
|
297
|
+
period, a comma, a closing quotation mark) does not block the match: the word-boundary test
|
|
298
|
+
only requires the code point after `right`'s last letter to be non-`LETTER`/non-`DIGIT`, and
|
|
299
|
+
punctuation satisfies that trivially (`I love rock 'n' roll.` still vetoes; §6 row P3).
|
|
300
|
+
|
|
301
|
+
**A context word may itself begin at document/span start or end at document/span end — the
|
|
302
|
+
Word definition above already says so (`NONE` is an accepted outer boundary).** `Rock 'n' roll
|
|
303
|
+
is great.`, where `Rock` is the very first thing in the document, still vetoes (§6 row P2): the
|
|
304
|
+
mark's own left neighbour `cp[i-1]` is a real `INLINE-SPACE` code point (the space after
|
|
305
|
+
`Rock`), and `wordEndsAt` finds `Rock` ending there with a legal `NONE` boundary one step
|
|
306
|
+
further left. **What cannot happen is the mark itself sitting at that extremity.** A `NARROW`
|
|
307
|
+
mark at the very start or end of the text unit has `Llit`/`Rlit` = `NONE`, which is not
|
|
308
|
+
`INLINE-SPACE`, so its own outer test fails before a word is even sought — the veto needs an
|
|
309
|
+
actual space *and* an actual word beyond it, and a mark with nothing to its outer side has
|
|
310
|
+
neither. The constraint is on the mark's adjacency, not on where the word may fall.
|
|
311
|
+
|
|
312
|
+
**Membership evidence, not mechanism**, stated normatively in `nbsp.md` §2.1 and governing every
|
|
313
|
+
list-valued field in `locale.schema.json`: a locale attests that a `{ left, elided, right }`
|
|
314
|
+
triple belongs in this list; this rule owns the mechanism. Each triple is independently
|
|
315
|
+
evidenced — `{ left: "rock", elided: "n", right: "roll" }` is cited (§6, `en-US.json`
|
|
316
|
+
`sources`), and that citation supports no other triple: it does not, for instance, attest
|
|
317
|
+
`{ left: "fish", elided: "n", right: "chips" }`, which is not listed here and would need its
|
|
318
|
+
own citation before being added. This is deliberately **not** a general elision-word list, and
|
|
319
|
+
it must not grow into one merely because the mechanism could technically encode it. A locale
|
|
320
|
+
with no citable evidence for a triple of this exact shape ships an empty list rather than a
|
|
321
|
+
guessed one — the same discipline `PLAN.md` §6.1 states for every locale field, and the reason
|
|
322
|
+
this veto is populated for only one locale at launch (§6).
|
|
323
|
+
|
|
324
|
+
**Known, accepted residual ambiguity — not eliminated by the full-context requirement.** The
|
|
325
|
+
veto has no notion of authorial intent: it matches code points, not meaning, so a document that
|
|
326
|
+
explicitly states its own intent to write three separate tokens — the word `rock`, a genuinely
|
|
327
|
+
quoted letter `n`, and the word `roll` — and then, for illustration, reproduces the identical
|
|
328
|
+
surface sequence `rock 'n' roll`, still gets that final occurrence vetoed:
|
|
329
|
+
|
|
330
|
+
```
|
|
331
|
+
The sequence is the word rock, the quoted letter 'n', and the word roll: rock 'n' roll.
|
|
332
|
+
```
|
|
333
|
+
|
|
334
|
+
transforms to
|
|
335
|
+
|
|
336
|
+
```
|
|
337
|
+
The sequence is the word rock, the quoted letter “n”, and the word roll: rock ’n’ roll.
|
|
338
|
+
```
|
|
339
|
+
|
|
340
|
+
The **first** `'n'` is correctly left alone: its `left` context is `letter`, not `rock`, and it
|
|
341
|
+
is immediately followed by a comma rather than a space, so no configured idiom matches and it
|
|
342
|
+
pairs as an ordinary depth-1 quotation — exactly the mechanism §6 rows N1/N2 already prove. The
|
|
343
|
+
**second** occurrence — `rock 'n' roll` at the end — matches `left`/`elided`/`right` exactly,
|
|
344
|
+
regardless of the fact that the same sentence has just told the reader it is meant as a literal
|
|
345
|
+
concatenation of the three tokens just described, not a fresh invocation of the idiom. **Under
|
|
346
|
+
the authorial intent the sentence itself states, this conversion is a real semantic false
|
|
347
|
+
positive** — the final `'n'` is meant as a repetition of the same quoted letter named earlier in
|
|
348
|
+
the sentence, not the idiom, and eliding it changes what the author wrote to mean. It is
|
|
349
|
+
accepted as an **unavoidable** false positive for a bounded, surface-form contract, not a
|
|
350
|
+
tolerated one: `left`, `elided`, `right` and the single permitted `INLINE-SPACE` on each side
|
|
351
|
+
are the entire alphabet this veto is defined over, and none of them can encode "this is a
|
|
352
|
+
demonstration, not an utterance of the idiom" — no finite extension of that alphabet could,
|
|
353
|
+
short of the veto reasoning about meaning, which is out of scope for a literal, bounded
|
|
354
|
+
contract (§3.2). The implementation still **conforms exactly** to its own normative matcher on
|
|
355
|
+
this input: the two occurrences are given the treatment the matcher's rules specify byte for
|
|
356
|
+
byte, and the defect is that the matcher's contract cannot distinguish these bytes from the
|
|
357
|
+
idiom it exists to recognise — because, written this way, they are the same bytes. Recorded
|
|
358
|
+
here in the same spirit as `dashes.md` §7.9, and §6 row N6 pins it as a fixture, so the exposure
|
|
359
|
+
is visible in the conformance suite rather
|
|
360
|
+
than discovered in someone's content.
|
|
361
|
+
|
|
362
|
+
**Universal medial-`n` elision veto (spec 1.1.0).** The listed elision veto above closes
|
|
363
|
+
`rock 'n' roll` for `en-US`, where a citation exists. It closes nothing for `en-GB`, `de-DE`,
|
|
364
|
+
`de-CH`, `fr`, `fr-CA`, `ru`, `fi`, `sv` or `el`, whose `elisionIdioms` lists are empty. Without a
|
|
365
|
+
further mechanism the marks in `rock 'n' roll` fall through to ordinary pairing in all nine and
|
|
366
|
+
are typeset as a quotation.
|
|
367
|
+
|
|
368
|
+
**Operator decision, 2026-09-09: `rock 'n' roll` and `rock'n'roll` are international, and both
|
|
369
|
+
take U+2019 in every locale.** The idiom is not a locale fact — it is a fixed borrowed string that
|
|
370
|
+
appears untranslated in prose in all ten — so the mechanism is locale-blind and lives in this rule
|
|
371
|
+
rather than in locale data. Locale data could not carry it in any case: the `sources` discipline
|
|
372
|
+
(§2) would block an entry in a locale whose own orthography writes the borrowing out instead
|
|
373
|
+
(Russian writes рок-н-ролл), and an entry withheld for that reason would leave uncovered exactly
|
|
374
|
+
the locales this decision is about.
|
|
375
|
+
|
|
376
|
+
The predicate (`src/rules/quote-ambiguity.ts`, `computeAmbiguousShapeIndices`) is **this rule's
|
|
377
|
+
alone** as of spec 1.1.0. Under 0.5.0 `apostrophe` consulted the same module to decide what to skip;
|
|
378
|
+
it no longer skips anything, so the module has one caller and there is nothing left for the two
|
|
379
|
+
rules to drift apart on. The shape: a pair of `NARROW` marks enclosing **exactly one code point**,
|
|
380
|
+
either U+006E `n` or U+004E `N`, with **at least one** `INLINE-SPACE` code point immediately
|
|
381
|
+
outside each mark. `rock 'n' roll`, `fish 'n' chips` and `rock 'N' roll` all match;
|
|
382
|
+
`She chose 'A' today`, `They said 'no' yesterday` and `say 'yes' now` do not, and pair as ordinary
|
|
383
|
+
quotations.
|
|
384
|
+
|
|
385
|
+
The closed-up form `rock'n'roll` is not this veto's business at all — the medial-elision veto
|
|
386
|
+
above already declines both marks on `ALNUM` neighbours, and `apostrophe` converts them. It needs
|
|
387
|
+
no special case and gets none.
|
|
388
|
+
|
|
389
|
+
**Why `NARROW` and not straight ASCII only — an idempotency obligation, not a preference.** Spec
|
|
390
|
+
0.5.0's predicate matched U+0027 alone. That was sound *there*, because its outcome was to
|
|
391
|
+
preserve the author's straight marks byte-identically, so a second pipeline pass saw the same
|
|
392
|
+
bytes and re-vetoed. This veto instead lets `apostrophe` **convert** both marks to U+2019, and a
|
|
393
|
+
straight-only predicate therefore would not recognise its own output: pass 2 would pair
|
|
394
|
+
`rock ’n’ roll` as an ordinary `NARROW` quotation on the next run. Measured, not argued — with a
|
|
395
|
+
straight-only predicate, `rock ’n’ roll` becomes `rock «n» roll` in `ru` and `rock ”n” roll` in
|
|
396
|
+
`fi`, an idempotency violation and so a release blocker (`pipeline-idempotency.md` §1). Matching
|
|
397
|
+
the whole `NARROW` class makes the converted form a fixed point in all ten locales.
|
|
398
|
+
|
|
399
|
+
**Why exactly one code point `n`, and not a letter-count shape.** 0.5.0's predicate matched 1 to 3
|
|
400
|
+
`LETTER` code points, which cannot tell the idiom from an ordinary short nested quotation and
|
|
401
|
+
declined both — `«это 'моё' дело»` in `ru`, `“He said 'no' to me,”` in `en-US`. Counting letters
|
|
402
|
+
is wrong here twice over. It is too broad: a genuine short quotation is far commoner in prose than
|
|
403
|
+
the idiom, so declining it is the larger error. And it is not normalization-stable — `LETTER`
|
|
404
|
+
includes `Mn` (§3.1) and this rule never normalizes its input (`ARCHITECTURE.md` §4.3), so `'моё'`
|
|
405
|
+
is three `LETTER` code points composed and four decomposed, and the same visible word took two
|
|
406
|
+
different branches depending on which form the author's editor happened to save. A comparison
|
|
407
|
+
against a fixed pair of code points has no such dependency.
|
|
408
|
+
|
|
409
|
+
**"At least one", deliberately, not "exactly one".** Only the single code point immediately
|
|
410
|
+
adjacent to each mark is tested; a longer run of inline spaces further out — `rock 'n' roll`,
|
|
411
|
+
doubled — does not invalidate the match. Narrowing this to "exactly one" was considered and
|
|
412
|
+
rejected: it would make a doubled-space variant of `rock 'n' roll` fall through to ordinary
|
|
413
|
+
quote-pairing, reintroducing a false-positive quotation conversion the veto exists to prevent, in
|
|
414
|
+
exchange for no compensating benefit. The existing listed-idiom matcher (above) may remain
|
|
415
|
+
stricter on this point without contradiction — with extra spaces its own literal `left`/`right`
|
|
416
|
+
word-boundary test can fail to match the cited en-US `rock`/`n`/`roll` tuple, and when it does,
|
|
417
|
+
this veto still wins and the marks are still converted.
|
|
418
|
+
|
|
419
|
+
**Every position this shape matches is added to the same veto set the listed-idiom check
|
|
420
|
+
populates**, and both mechanisms have the identical outcome: `quotes` declines the pairing, both
|
|
421
|
+
marks survive pass 2 unmatched, and `apostrophe` (order 50) renders each independently by its own
|
|
422
|
+
structural case ladder — leading elision (`apostrophe.md` §3.3 case 4) on the first, trailing
|
|
423
|
+
elision (case 3) on the second, giving `rock ’n’ roll`. `en-US`'s cited `elisionIdioms` entry now
|
|
424
|
+
vetoes a strict subset of what this veto vetoes; it is retained because the listed mechanism
|
|
425
|
+
matches multi-code-point `elided` content this one does not, and because the two can never
|
|
426
|
+
disagree about a position both cover. **Spec 0.5.0's preserve set is withdrawn**
|
|
427
|
+
(`computePreserveIndices`, `apostrophe.md` §3.4): it existed to stop `apostrophe`'s case ladder
|
|
428
|
+
from converting marks 0.5.0 wanted preserved, and conversion is now the specified outcome for
|
|
429
|
+
every position this veto matches, so `apostrophe` needs no knowledge of the veto at all and its
|
|
430
|
+
ordinary ladder does the right thing unaided.
|
|
431
|
+
|
|
432
|
+
**Already-curly input is matched, and preserved as typed.** A pair typed directly with U+2019 is
|
|
433
|
+
in `NARROW`, so it is vetoed like any other — and since `apostrophe` emits U+2019 only in place of
|
|
434
|
+
U+0027 (`apostrophe.md` §1), it edits nothing there: `rock ’n’ roll` is a fixed point, and
|
|
435
|
+
`rock ‘n’ roll` keeps the author's own U+2018 rather than being re-typeset as that locale's
|
|
436
|
+
quotation. This is the direct opposite of 0.5.0's rule for curly input, and the reason is the
|
|
437
|
+
idempotency obligation two paragraphs above, not a change of view about authorial intent.
|
|
438
|
+
|
|
439
|
+
**Accepted false positive, recorded rather than tolerated.** `The letter 'n' is common.` becomes
|
|
440
|
+
`The letter ’n’ is common.` — a genuine quotation of the letter *n*, rendered as an elision. So
|
|
441
|
+
does the same shape one level deeper, `He said "press 'n' now".` Both are pinned as fixtures
|
|
442
|
+
(`spec/fixtures/en-US.json`) so the exposure sits in the conformance suite rather than in
|
|
443
|
+
someone's content. The veto's alphabet is one code point and the spaces around it; nothing in that
|
|
444
|
+
alphabet can encode "this is a quoted letter, not an elided one", and no bounded literal contract
|
|
445
|
+
over these bytes could. The trade is deliberate and was made with the numbers in view: 0.5.0 paid
|
|
446
|
+
for this same shape by declining **every** short quotation in **every** locale, a far larger and
|
|
447
|
+
far commoner class of error than quoting the single letter *n*.
|
|
448
|
+
|
|
449
|
+
**V1 — same-V1-identity adjacency veto** (both widths):
|
|
450
|
+
|
|
451
|
+
> Define `V1ID(c) = U+2019 if c = U+0027, else c` — the **V1 identity** of a code point. `V1ID`
|
|
452
|
+
> is the identity function everywhere except at U+0027; it is *not* a claim that a given U+0027
|
|
453
|
+
> will become U+2019 (see below). Let `gapInsertable = g ∈ SPACE-RIGHT ∨ g ∈ SPACE-LEFT`. If
|
|
454
|
+
> `V1ID(Llit) = V1ID(g)`, or (`Llit ∈ INLINE-SPACE` and `V1ID(Lskip) = V1ID(g)` and
|
|
455
|
+
> `gapInsertable`), set both capabilities false. Symmetrically for `Rlit`/`Rskip`.
|
|
456
|
+
|
|
457
|
+
`""`, `''`, `««`, `””` and any longer run of one glyph, by the first (literal) clause. This is
|
|
458
|
+
0.1.0's veto, generalised, and it keeps the `"""a""` witness dead and `"a "" b"` intact.
|
|
459
|
+
**Granularity is deliberately "same V1 identity", not "same width"**: the narrower test
|
|
460
|
+
reproduces both witnesses, and a width-based test would veto the ordinary mixed boundary `"'`.
|
|
461
|
+
|
|
462
|
+
**`V1ID` and why V1 cannot compare raw code points (spec 0.4.1).** `apostrophe` (R₆) is the one
|
|
463
|
+
rule ordered after `quotes` that can change a candidate's own code point, and every edit it makes
|
|
464
|
+
replaces a U+0027 with U+2019 — the only code point it ever emits for that position
|
|
465
|
+
(`apostrophe.md` §1). **This is not an unconditional mapping.** `apostrophe`'s own case ladder
|
|
466
|
+
(`apostrophe.md` §3.3) leaves a U+0027 unedited under two of its six cases — the **prime guard**
|
|
467
|
+
(case 1: `6' 2"`, a foot/inch mark) and **case 5** (`a ' b`, `''`: nothing inferable) — and
|
|
468
|
+
`apostrophe.md` §5 names both explicitly as surviving, unedited, to the rule's own second-run
|
|
469
|
+
argument. So only a *subset* of the U+0027s `quotes` leaves unmatched are ever actually converted.
|
|
470
|
+
|
|
471
|
+
Every other capability test in this section reads *class* membership, which is unaffected by
|
|
472
|
+
that substitution regardless of whether it happens (Lemma A, §5) — but V1, uniquely, compares
|
|
473
|
+
*code points*, because that is its job: telling `""` from `"'`. A literal comparison is therefore
|
|
474
|
+
the one place a U+0027 neighbour's *eventual* fate matters, and `quotes` cannot know that fate
|
|
475
|
+
from inside its own pass: whether a given U+0027 reaches `apostrophe` at all — as opposed to
|
|
476
|
+
being claimed by `quotes` itself and rendered as a quote glyph — and what its *own* neighbours
|
|
477
|
+
will read once `quotes` has finished rendering the whole accepted set, are exactly the outputs
|
|
478
|
+
this rule is still computing when V1 runs. Deciding V1 exactly would mean re-deriving
|
|
479
|
+
`apostrophe`'s verdict against `quotes`' own not-yet-final output — a circular second pipeline
|
|
480
|
+
inside this rule, not a fix. (A per-position "is this specific U+0027 definitely going to survive
|
|
481
|
+
unmatched with these exact neighbours" analysis was considered and rejected for the same reason:
|
|
482
|
+
neighbouring accepted pairs can rewrite the very code points that analysis would need to read
|
|
483
|
+
first.)
|
|
484
|
+
|
|
485
|
+
`V1ID` is therefore a **conservative over-approximation**, not an exact rule: it treats *every*
|
|
486
|
+
U+0027 as if it might already be, or might become, U+2019, and every U+2019 as if it might be an
|
|
487
|
+
unconverted U+0027 — regardless of which of `apostrophe`'s six cases (1, 2, 3, 3a, 4, 5) will actually apply.
|
|
488
|
+
|
|
489
|
+
**What "one-directional" bounds, precisely — a local claim, not a pipeline-level one.** At a
|
|
490
|
+
single V1 comparison, `V1ID` can only ever *add* a veto: it merges two identities that raw
|
|
491
|
+
equality would have kept apart, and it never grants `canOpen`/`canClose` to a candidate that
|
|
492
|
+
would not otherwise have had one — `V1ID` only ever makes the equality test *stricter*, never
|
|
493
|
+
looser. **That local fact does not extend to the rule's global output.** Pass 2 (§3.3, the stack
|
|
494
|
+
pairing) and pass 4 (§3.5, the certification gate) both operate over the *whole* candidate list at
|
|
495
|
+
once, not candidate-by-candidate: declining one candidate can remove a crossing pair, a nesting
|
|
496
|
+
conflict, or a certification instability that was previously the reason a *different, unrelated*
|
|
497
|
+
pair failed to certify. An extra local veto can therefore **indirectly** let another pair certify
|
|
498
|
+
that previously could not, reassign which glyphs/depth a pair receives, or change what `nbsp`
|
|
499
|
+
(order 70) inserts downstream — this is not a hypothetical: **the reported counterexample is
|
|
500
|
+
exactly this shape.** On pass 1, `V1ID` declining the crossing `NARROW` pair (`U+2019`
|
|
501
|
+
adjacent to the unmatched `U+0027`) is precisely what lets the `WIDE` pair certify instead of
|
|
502
|
+
being caught in the same gate rejection that, pre-fix, declined both. Vetoing one candidate
|
|
503
|
+
*enabled* a different conversion; "extra decline only" was never literally true of the rule's
|
|
504
|
+
output, only of the single comparison `V1ID` changes.
|
|
505
|
+
|
|
506
|
+
The actual safety argument therefore does **not** rest on a monotonicity claim. It rests on: (1)
|
|
507
|
+
a missed pairing is the failure mode this whole spec already treats as safe — "missing a possible
|
|
508
|
+
improvement is safer than damaging text" (`PLAN.md` §6.1,
|
|
509
|
+
`docs/AUDIT_REMEDIATION_AND_RELEASE_PLAN.md` §3.1) — against the defect `V1ID` closes, a hard
|
|
510
|
+
release blocker (an idempotency violation); and (2) **inspection of every output `V1ID` globally
|
|
511
|
+
changes within a reproducible bounded sweep**, not an a-priori argument about direction. The "V1ID
|
|
512
|
+
empirical audit" section below is that inspection: every changed output in the sweep — declined,
|
|
513
|
+
newly enabled, or reassigned — is individually classified, and none is a false positive against
|
|
514
|
+
plausible prose. `tests/rules/quotes.test.ts` pins one permanent regression case per distinct
|
|
515
|
+
mechanism found: a pure conservative veto (prime-guard and case-5 U+0027 preserved exactly), the
|
|
516
|
+
indirect newly-enabled-outer-pair mechanism (the reported counterexample's own shape), a pairing
|
|
517
|
+
reassignment, and the `fr`/`fr-CA` downstream `nbsp` consequence — plus a direct test that the
|
|
518
|
+
plain U+0027 vs. U+2019 distinction still governs `apostrophe` and the medial veto exactly as
|
|
519
|
+
before, unaffected by `V1ID`, which is scoped to V1's own comparison alone.
|
|
520
|
+
|
|
521
|
+
**Reachability, precisely.** `QUOTEMARK`, V1, and `apostrophe`'s case ladder are defined
|
|
522
|
+
identically for every locale, so the *algorithmic* shape (an unmatched U+0027 adjacent to a
|
|
523
|
+
`QUOTEMARK` candidate whose glyph is U+2019) is constructible in every locale with a synthetic
|
|
524
|
+
string, and `tests/rules/quotes.test.ts` sweeps it across every entry in `KNOWN_LOCALES` for
|
|
525
|
+
exactly that reason — that is a claim about the algorithm, not about naturally occurring prose.
|
|
526
|
+
Separately, and more narrowly: `en-GB`'s primary `close` and `fi`/`sv`'s secondary pair (both
|
|
527
|
+
sides) are U+2019 *as that locale's own prescribed glyph* — not merely a candidate a foreign or
|
|
528
|
+
adversarial input could contain — so those three locales are where an author's ordinary,
|
|
529
|
+
correctly-typed document (one already-closed quotation immediately followed by, say, a
|
|
530
|
+
possessive apostrophe with no separating space) is the most plausible route to this shape without
|
|
531
|
+
any copy-pasted or malformed input at all. `de-CH`, where the defect was first found, reaches it
|
|
532
|
+
only via a mixed straight/curly input (`spec/fixtures/de-CH.json`,
|
|
533
|
+
`de-ch-quotes-041-cobug-apostrophe-v1`) — de-CH's own glyphs are `« »`/`‹ ›`, neither of which is
|
|
534
|
+
U+2019, so the witness there is adversarial, not a natural-content route.
|
|
535
|
+
|
|
536
|
+
**V1ID empirical audit.** `quotes.ts`'s V1 block was twice temporarily reverted to the pre-0.4.1
|
|
537
|
+
raw comparison, run against a bounded, deterministic alphabet across every locale in
|
|
538
|
+
`spec/locales/registry.json`, and restored — a controlled experiment, not a committed second
|
|
539
|
+
engine, and not itself part of this repository (the exact reversion is exactly the diff `git
|
|
540
|
+
diff` shows against `tests/rules/quotes.test.ts`'s permanent assertions, which is what actually
|
|
541
|
+
pins the result). Exact counts from that run are deliberately **not** recorded here: an
|
|
542
|
+
uncommitted, one-off harness is not something a future maintainer can re-run from this document
|
|
543
|
+
alone, and this spec's normative claims must not depend on an experiment nobody can reproduce.
|
|
544
|
+
What is recorded, and *is* reproducible — by running `npx vitest run tests/rules/quotes.test.ts`
|
|
545
|
+
against `main` — is the durable conclusion the audit reached:
|
|
546
|
+
|
|
547
|
+
- Every output `V1ID` changes relative to the raw comparison, across the swept alphabet and every
|
|
548
|
+
locale, falls into exactly one of four mechanisms: a pure extra conservative veto (the mark
|
|
549
|
+
stays exactly as authored); an *indirect* newly-enabled conversion (vetoing one candidate
|
|
550
|
+
removes a crossing/certification conflict and lets a different, unrelated pair certify — the
|
|
551
|
+
reported counterexample's own shape, run in general); a pairing reassignment (the same input
|
|
552
|
+
pairs differently, not more or less); and the `fr`/`fr-CA` downstream `nbsp` consequence of a
|
|
553
|
+
differently accepted or declined pair. No difference in the swept range was unexplained.
|
|
554
|
+
- Every newly-enabled or reassigned output in the swept range was inspected against ordinary-prose
|
|
555
|
+
plausibility, not merely checked for a digit adjacency. None resembles content a human author
|
|
556
|
+
would write and reject: the newly-enabled class never contains a letter or digit at all (its
|
|
557
|
+
witnesses are runs of quote-class punctuation with no word between them, not prose); the
|
|
558
|
+
reassignment class only ever chooses between two already-plausible readings of a short ambiguous
|
|
559
|
+
fragment (e.g. an ordinary quotation vs. two independent elisions), never invents a reading with
|
|
560
|
+
no textual basis.
|
|
561
|
+
- A separate, narrower sweep confirmed the raw-comparison variant is genuinely non-idempotent
|
|
562
|
+
within a bound that reaches the reported defect's minimal witness shape, and that `V1ID` is not,
|
|
563
|
+
across every locale — the direct measurement of the closure this fix provides, not merely of its
|
|
564
|
+
conservative cost.
|
|
565
|
+
|
|
566
|
+
`tests/rules/quotes.test.ts`'s "V1ID" describe block pins one permanent, exact-output regression
|
|
567
|
+
case for each of the four mechanisms above (plus the U+0027/U+2019-still-matters-outside-V1
|
|
568
|
+
control) — that is the reproducible evidence this section rests on.
|
|
569
|
+
|
|
570
|
+
**The second clause is not in either design document that fed this revision, and closes a
|
|
571
|
+
composition-obligation violation found only by running the implementation.** A literal-only V1
|
|
572
|
+
lets the certification gate correctly decline a pairing when rendering it would place two
|
|
573
|
+
identical glyphs strictly adjacent (e.g. `«` immediately left of a mark the gate would render to
|
|
574
|
+
`«`) — but if `nbsp` (a *later* rule) subsequently inserts a space at exactly that gap, on the
|
|
575
|
+
*next* full-pipeline pass `quotes` sees the two occurrences with a space between them, the
|
|
576
|
+
literal adjacency is gone, and the *same* pairing certifies. That is `nbsp` creating work for an
|
|
577
|
+
earlier-ordered rule, forbidden by `pipeline-idempotency.md` §2's composition obligation, and it
|
|
578
|
+
reproduces with two witnesses, pinned in `spec/fixtures/fr.json` as `fr-quotes-030-cobug-html-tag-boundary`
|
|
579
|
+
(`«"<p class="x">"` in `html` mode) and `fr-quotes-030-cobug-double-open-then-closer` (`««”` in
|
|
580
|
+
`text` mode). The fix reads the
|
|
581
|
+
skip in exactly the two positions Lemma B (§5) already names as `nbsp`'s whole insertion
|
|
582
|
+
surface next to a quote mark — a candidate's own `g` value is shared identically with the
|
|
583
|
+
same-glyph neighbour a skip reaches, so checking `g`'s own `SPACE-RIGHT`/`SPACE-LEFT` membership
|
|
584
|
+
covers both of Lemma B's insertion sites at once, regardless of which of the two occurrences is
|
|
585
|
+
on which side of the gap. **A plain skip-based V1** (comparing `Lskip`/`Rskip` to `g`
|
|
586
|
+
unconditionally, with no `gapInsertable` guard) **was tried and rejected**: it also vetoes two
|
|
587
|
+
genuinely distinct, space-separated quotations (`'a' 'b'`, ordinary prose, not a typewriter
|
|
588
|
+
artefact) in every locale, including ones with no spaced pair at all — the guard is what keeps
|
|
589
|
+
the veto scoped to positions `nbsp` can actually reach.
|
|
590
|
+
|
|
591
|
+
Under mandate 1 the veto is **no longer permanent** — a converted mark can acquire or lose an
|
|
592
|
+
identical neighbour between runs — and that is stated rather than papered over: **the gate
|
|
593
|
+
(§3.5), not this veto, is the guarantor of stability against `quotes`' own re-derivation.** The
|
|
594
|
+
`gapInsertable` clause is what extends that guarantee to survive `nbsp` acting *between*
|
|
595
|
+
pipeline passes, which the gate alone — being a property of a single call — cannot see.
|
|
596
|
+
|
|
597
|
+
Record `{ index, width ∈ {WIDE, NARROW}, canOpen, canClose }` for every `i` with at least one
|
|
598
|
+
capability true. Candidates with both false are dropped and never touched.
|
|
599
|
+
|
|
600
|
+
Pass 1 still does **not** decide that a `NARROW` mark with a letter to the right and a space to
|
|
601
|
+
the left is a leading elision. It is `canOpen` and enters the pairing; if it finds no partner it
|
|
602
|
+
survives and `apostrophe` renders it U+2019. The pairing settles the ambiguity, not a heuristic.
|
|
603
|
+
|
|
604
|
+
### 3.3 Pass 2 — pair the candidates, one stack per width
|
|
605
|
+
|
|
606
|
+
**There are two stacks, one for `WIDE` and one for `NARROW`, and a candidate only ever touches
|
|
607
|
+
the stack for its own width.** 0.1.0's per-kind split generalised from "the literal code point
|
|
608
|
+
`DQ`/`SQ`" to "the glyph's visual width" — the property that actually carried the argument: no
|
|
609
|
+
reader opens a quotation with a double mark and closes it with a single one, in any language.
|
|
610
|
+
This is what keeps a Greek elision (`σ'`, `NARROW`) from closing a `WIDE` opener.
|
|
611
|
+
|
|
612
|
+
Walk the candidate list in index order. For candidate `c`, let `S` be the stack for `c.width`:
|
|
613
|
+
|
|
614
|
+
1. If `c.canClose`, `S` is non-empty, and **`¬vacuous(top(S).index, c.index)`** → pop `o`, record
|
|
615
|
+
the pair `(o.index, c.index)`.
|
|
616
|
+
2. Else if `c.canOpen` → push `c` onto `S`.
|
|
617
|
+
3. Else record `c` as unmatched.
|
|
618
|
+
|
|
619
|
+
When the walk ends, everything still on either stack is unmatched.
|
|
620
|
+
|
|
621
|
+
> **`vacuous(a, b)`** ⟺ every `cp[k]` with `a < k < b` is in `INLINE-SPACE` (vacuously true when
|
|
622
|
+
> `b = a + 1`).
|
|
623
|
+
|
|
624
|
+
**The vacuity condition is new and is forced by mandate 1.** In 0.1.0 an empty pair was
|
|
625
|
+
impossible: two adjacent marks of the same kind were vetoed by V1, and two of different kinds
|
|
626
|
+
were on different stacks. Now `«»` is two *different* code points of the *same* width, and `" "`
|
|
627
|
+
is two identical marks V1 cannot see past a space; both would enclose no content. The condition
|
|
628
|
+
sits in step 1 rather than as a fourth outcome: a closer that fails step 1 **falls through to
|
|
629
|
+
step 2 and then step 3**, so every candidate still reaches exactly one of three outcomes and the
|
|
630
|
+
accepted/unmatched partition stays exhaustive — that exhaustiveness is what makes the "declined"
|
|
631
|
+
set well-defined for the gate. **Do not add a fourth outcome to this list.**
|
|
632
|
+
|
|
633
|
+
Step 1 before step 2 — **closing takes precedence** — is retained unchanged. An `«` whose
|
|
634
|
+
right-hand space `nbsp` inserted reads `closeRight = Rskip`, exactly as it did before `nbsp`
|
|
635
|
+
touched it, so the tie-break's ambiguity does not shift under mandate 1.
|
|
636
|
+
|
|
637
|
+
### 3.4 Pass 3 — assign glyphs by depth
|
|
638
|
+
|
|
639
|
+
For a set of pairs `A`, the **depth** of `p ∈ A` is
|
|
640
|
+
|
|
641
|
+
> `depth_A(p) = 1 + |{ q ∈ A : q.open < p.open ∧ p.close < q.close }|`
|
|
642
|
+
|
|
643
|
+
— one plus the number of *accepted* pairs that strictly enclose it. Odd depth → `quotes.primary`;
|
|
644
|
+
even depth → `quotes.secondary`. Depths beyond 2 alternate by the same parity.
|
|
645
|
+
|
|
646
|
+
**Depth is computed over the accepted set `A`, not over the raw pass-2 output.** On a second run
|
|
647
|
+
the accepted set *is* the raw set, so a depth taken over the raw set on run 1 and over the
|
|
648
|
+
accepted set on run 2 would disagree whenever the gate declined anything. Depth over `A` agrees
|
|
649
|
+
on both runs by construction.
|
|
650
|
+
|
|
651
|
+
**Depth, not the mark's original width, selects the pair.** `'He said "no" to me,' she noted.`
|
|
652
|
+
promotes the outer `NARROW` marks to `en-US`'s `WIDE` primary pair and demotes the inner ones —
|
|
653
|
+
the reason width drift exists at all, and the case the certification gate exists to re-verify.
|
|
654
|
+
|
|
655
|
+
### 3.5 Pass 4 — the certification gate
|
|
656
|
+
|
|
657
|
+
This pass replaces 0.1.0's Claims 1–3. It is a *check the algorithm performs*, not a property a
|
|
658
|
+
reader must verify by argument alone.
|
|
659
|
+
|
|
660
|
+
**Rendering.** `render(cp, A)` is the array obtained by applying, for every `p ∈ A` with assigned
|
|
661
|
+
pair `P = pairFor(depth_A(p))`:
|
|
662
|
+
|
|
663
|
+
- replace `cp[p.open]` with `P.open` and `cp[p.close]` with `P.close`;
|
|
664
|
+
- **if and only if `P.innerSpace = "none"`**, delete the pair's two *inner runs* where the
|
|
665
|
+
landing guard permits:
|
|
666
|
+
- the **open-side run** = the maximal `INLINE-SPACE` run starting at `p.open + 1` (possibly
|
|
667
|
+
empty); its **landing** is the first code point past it;
|
|
668
|
+
- the **close-side run** = the maximal `INLINE-SPACE` run ending at `p.close - 1`; its landing
|
|
669
|
+
is the first code point before it;
|
|
670
|
+
- a run is deleted iff it is non-empty, its landing is in `DELETE-LANDING`, and it is not
|
|
671
|
+
simultaneously both of the pair's runs (a pair enclosing nothing but spaces deletes
|
|
672
|
+
neither — unreachable for an accepted pair given §3.3's vacuity condition, but stated because
|
|
673
|
+
the construction is otherwise silent about it).
|
|
674
|
+
|
|
675
|
+
`render` also yields the order-preserving index map `π` from surviving input indices to output
|
|
676
|
+
indices.
|
|
677
|
+
|
|
678
|
+
**Certification loop.**
|
|
679
|
+
|
|
680
|
+
```
|
|
681
|
+
A := the pair set from pass 2
|
|
682
|
+
loop:
|
|
683
|
+
if A = ∅: accept ∅ and stop
|
|
684
|
+
(y, π) := render(cp, A)
|
|
685
|
+
B := passes 1–2 applied to y (raw, ungated)
|
|
686
|
+
if { (π(p.open), π(p.close)) : p ∈ A } = B: accept A and stop
|
|
687
|
+
A' := { p ∈ A : (π(p.open), π(p.close)) ∈ B }
|
|
688
|
+
if A' = A: A := A \ { the pair of A with the greatest `open` index }
|
|
689
|
+
else: A := A'
|
|
690
|
+
```
|
|
691
|
+
|
|
692
|
+
**Every clause is normative, including the tie-break.** When the intersection fails to shrink
|
|
693
|
+
`A` — exactly when `B ⊋ π(A)`, i.e. the rendered array admits a pairing the input did not — one
|
|
694
|
+
pair is removed, and *which* pair is specified so two ports cannot disagree.
|
|
695
|
+
|
|
696
|
+
**Termination.** Each iteration either accepts or strictly shrinks `A`. `A` is finite and `∅`
|
|
697
|
+
accepts unconditionally, so the loop runs at most `|A₀| + 1` times. Cost is `O(k·n)` with `k` the
|
|
698
|
+
number of pairs; in the overwhelmingly common case the first round certifies and the extra cost
|
|
699
|
+
is one `O(n)` pass. An implementation may short-circuit on that.
|
|
700
|
+
|
|
701
|
+
**Why the exit condition is `B = π(A)` and not `π(A) ⊆ B`.** Exiting with extra pairs in `B` is
|
|
702
|
+
unsound: the second run *starts* from `B`, and if `B` certifies, the second run renders `B`'s
|
|
703
|
+
extra pairs and the output differs. The intersection alone does not terminate at a sound state,
|
|
704
|
+
which is why the forced-removal clause exists.
|
|
705
|
+
|
|
706
|
+
**Empirical status of the forced-removal clause.** An exhaustive sweep of ~850,000 inputs across
|
|
707
|
+
all ten locales (the alphabet of §9's fixtures, to length 4–5, plus the named witnesses) did not
|
|
708
|
+
observe the clause firing even once — every case that reached the gate either certified in round
|
|
709
|
+
1 or was resolved by the plain intersection shrinking `A`. The clause is implemented as specified
|
|
710
|
+
(it is strictly more conservative than omitting it, never less sound, so it costs nothing even if
|
|
711
|
+
unreachable in practice) but its necessity is **not** demonstrated by a concrete witness as of
|
|
712
|
+
this writing. If a future sweep finds one, it belongs in `spec/fixtures/` as a named case.
|
|
713
|
+
|
|
714
|
+
### 3.6 Pass 5 — emit
|
|
715
|
+
|
|
716
|
+
For each `p` in the accepted set:
|
|
717
|
+
|
|
718
|
+
- an edit replacing `cp[p.open]` with the assigned `open` glyph, **omitted if the code point is
|
|
719
|
+
already that glyph**;
|
|
720
|
+
- likewise at `p.close`;
|
|
721
|
+
- a deletion edit for each inner run the landing guard permitted, **omitted if the run is
|
|
722
|
+
empty**.
|
|
723
|
+
|
|
724
|
+
An edit whose replacement is identical to the span it replaces is never emitted — the
|
|
725
|
+
invisible-edit principle `dashes.md` §3.3.1 states for the joiner, applied so that
|
|
726
|
+
already-correct text produces a byte-identical no-op rather than a diff of self-replacements.
|
|
727
|
+
Suppression is per **mark**, not per pair: a pair where only one side needed a glyph change
|
|
728
|
+
emits exactly one glyph edit.
|
|
729
|
+
|
|
730
|
+
Emitted spans are pairwise disjoint: replacements sit at distinct mark indices; a deletion span
|
|
731
|
+
lies strictly between a mark and its landing; and since accepted pairs are properly nested or
|
|
732
|
+
disjoint (stack discipline) and a pair enclosing only spaces deletes neither run, no two spans
|
|
733
|
+
can claim the same index. Edits are sorted ascending before returning.
|
|
734
|
+
|
|
735
|
+
Unmatched candidates and declined pairs produce no edit and keep their code point exactly.
|
|
736
|
+
|
|
737
|
+
### 3.7 `innerSpace`: `quotes` deletes, `nbsp` inserts
|
|
738
|
+
|
|
739
|
+
The split is asymmetric on purpose.
|
|
740
|
+
|
|
741
|
+
- `innerSpace = "none"` — a space inside the pair is *wrong* and nobody else can remove it.
|
|
742
|
+
`nbsp` has no deletion capability at all. `quotes` deletes it, subject to the landing guard.
|
|
743
|
+
- `innerSpace ≠ "none"` — a space inside the pair is *required*, and `nbsp` (order 70) already
|
|
744
|
+
inserts and converts it, correctly and idempotently, at `nbsp.md` §3.10. `quotes` emits
|
|
745
|
+
nothing and deletes nothing there.
|
|
746
|
+
|
|
747
|
+
So `quotes` still never inserts a space, `nbsp` keeps ownership of all no-break spacing, and
|
|
748
|
+
there is no second copy of anyone's rules. `"mot"` → `«mot»` → (via `nbsp`) `«` U+00A0 `mot`
|
|
749
|
+
U+00A0 `»` — U+00A0 because `fr`'s primary pair sets `innerSpace: "nbsp"`, and **no shipped
|
|
750
|
+
locale sets `"narrow-nbsp"`** (nbsp.md §3.1a); this read U+202F until spec 1.3.0, which is a
|
|
751
|
+
claim about a locale file that the file never supported. `« mot »` → (no edit from `quotes`) →
|
|
752
|
+
the same output via `nbsp`. Both paths
|
|
753
|
+
converge, which is the property the pipeline needs.
|
|
754
|
+
|
|
755
|
+
**The landing guard is not a heuristic; it is a structural CO discharge.** The deleted run's near
|
|
756
|
+
side is a quote glyph and its far side is in `ALNUM ∪ QUOTEMARK`. Every class the four
|
|
757
|
+
earlier-ordered rules read as evidence — `spaces`' `CONTENT`/`STRIP-BEFORE`/bracket sets,
|
|
758
|
+
`ellipsis`' dot runs, `dashes`' `DASH`, `INERT-DASH`, `DIGIT`, `SPACE ∪ NOBREAK-SPACE`, `LETTER`,
|
|
759
|
+
`ROMAN`, `hyphen`'s `WORDISH` — classifies a quote glyph and a U+0020 **identically** (both
|
|
760
|
+
outside all of them, except `SPACE`, which the guard's own construction keeps away from any dash
|
|
761
|
+
run). `ALNUM ∪ QUOTEMARK` is the largest landing class for which that holds; letting the
|
|
762
|
+
deletion land on a dash is what would produce `" --x"` → `“--x”` → `“—x”`, a violation of `I₃`.
|
|
763
|
+
Two absolutes fall out of the same construction: **a `BREAK` is never deleted** (not in
|
|
764
|
+
`INLINE-SPACE`, so it terminates the run and, being outside `DELETE-LANDING`, declines the
|
|
765
|
+
deletion) and **`MARKER` is never crossed** (same mechanism), so the span partition is stable
|
|
766
|
+
between runs as `modes.md` §5 requires.
|
|
767
|
+
|
|
768
|
+
---
|
|
769
|
+
|
|
770
|
+
## 4. Must not touch
|
|
771
|
+
|
|
772
|
+
Each bullet is **[P]** (a guarantee of `transform`) or **[R]** (true of this rule alone, with the
|
|
773
|
+
rule that can falsify it named), per `pipeline-idempotency.md` §5.2.
|
|
774
|
+
|
|
775
|
+
- **[P] A medial apostrophe:** `don't`, `l'été`, `O'Brien`, `Hawai'i`, `1990's` — and their
|
|
776
|
+
U+2019 forms on a second pass. Vetoed in §3.2.
|
|
777
|
+
- **[R] A leading elision that finds no partner** (`'90s`, `'tis`) and **[R] a trailing
|
|
778
|
+
possessive that finds no partner** (`the dogs' bowls`). Handed to `apostrophe`.
|
|
779
|
+
- **[R] A foot or inch mark:** `6' 2"`. Both marks are `canClose` only with an empty stack. The
|
|
780
|
+
0.1.0 caveat still applies verbatim: in `"6' 2"` the foot mark pairs, and the protection is not
|
|
781
|
+
pipeline-level.
|
|
782
|
+
- **[P] Any unbalanced mark.** Convert what is unambiguous, leave the rest — §3.6 of 0.1.0's
|
|
783
|
+
policy is unchanged.
|
|
784
|
+
- **[P] Any run of two or more identical quote marks:** `""`, `'''`, `««`, `””`. V1.
|
|
785
|
+
- **[P] A pair enclosing nothing but spaces:** `«»`, `" "`, `” ”`. §3.3's vacuity condition.
|
|
786
|
+
- **[P] U+2032, U+2033, U+02BC, U+0060, U+00B4** and the LaTeX idioms `` ` `` and `''`. Not in
|
|
787
|
+
`QUOTEMARK`.
|
|
788
|
+
- **[P] Anything inside a skipped region.** Attribute values, code spans, fenced code, `<pre>`,
|
|
789
|
+
URLs — removed by the mode adapter.
|
|
790
|
+
- **[P] Line terminators.** Never deleted, never crossed by a skip walk.
|
|
791
|
+
- **[R] Outer spacing.** This rule deletes only *inner* spacing and inserts nothing.
|
|
792
|
+
|
|
793
|
+
**No longer on this list, by operator decision:** every non-straight quotation glyph. `“ ” « »
|
|
794
|
+
„ ‘ ’ ‚ ‹ ›` and the rest are now candidates. `en-GB` text containing a correct `“…”` pair will
|
|
795
|
+
be rewritten to `‘…’`; `de-DE` text containing `«Wort»` becomes `„Wort“`. That is mandate 1, and
|
|
796
|
+
it is the largest behavioural change in this revision.
|
|
797
|
+
|
|
798
|
+
**Known, accepted residual risk — not new under this revision.** A leading elision that has
|
|
799
|
+
already been curled to U+2019 (`’90s were fun,’ he said`) can, if a later unrelated closing mark
|
|
800
|
+
exists on the same stack, be paired as an opener and lose its elided-digit reading. This was
|
|
801
|
+
already possible for straight input under 0.1.0 (`'90s were fun,' he said` is corrupted
|
|
802
|
+
identically) — it is §7 item 3's family, "no local way to distinguish an elision from a real
|
|
803
|
+
opening quote", now also reachable via curly input because mandate 1 makes curly input a
|
|
804
|
+
candidate at all. **§7 item 8's `quotes.elisionIdioms` (spec 0.4.0) does not cover this case**:
|
|
805
|
+
that veto fires only on a listed idiom's full left/elided/right context found together, not on
|
|
806
|
+
a lone elision separated from any partner by arbitrary distance. It is **not** a new defect this
|
|
807
|
+
revision introduces; §5's Corollary A1 explains why no U+2019-specific restriction is applied, and
|
|
808
|
+
`spec/fixtures/` pins the verified behaviour across `en-GB`, `en-US`, `fi` and `sv` (§6, rows
|
|
809
|
+
E1–E4) so the tradeoff is a citable fact rather than an assumption.
|
|
810
|
+
|
|
811
|
+
---
|
|
812
|
+
|
|
813
|
+
## 5. Idempotency argument
|
|
814
|
+
|
|
815
|
+
0.1.0's Claims 1 and 2 are **withdrawn**, not weakened. Claim 1 ("a converted position is not a
|
|
816
|
+
candidate on the second run") is false by construction under mandate 1, and with it goes the
|
|
817
|
+
permanence of V1: two marks that were distinct can both be assigned the same glyph, so the
|
|
818
|
+
veto's verdict is no longer immutable. Claim 2's monotonicity is likewise gone: a converted mark
|
|
819
|
+
can *gain* a capability. Claim 3 rested on both. Three replacements follow.
|
|
820
|
+
|
|
821
|
+
### Lemma A — glyph-blindness
|
|
822
|
+
|
|
823
|
+
> Replacing any `QUOTEMARK` at index `j` with any other `QUOTEMARK` changes no capability of any
|
|
824
|
+
> candidate at any index `i ≠ j`.
|
|
825
|
+
|
|
826
|
+
*Proof.* A candidate's four neighbour reads are compared against `NONE`, `SPACELIKE`, `OPENISH`,
|
|
827
|
+
`CLOSEISH`, `DASHISH`, `QUOTEMARK`, `MARKER` and `ALNUM` (the medial veto), and V1 compares
|
|
828
|
+
`V1ID` of the relevant code points against `V1ID(g)` (spec 0.4.1). Every `QUOTEMARK` is in
|
|
829
|
+
`OPENISH`, in `CLOSEISH`, in `QUOTEMARK`, in none of `SPACELIKE`, `DASHISH`, `ALNUM`, and is
|
|
830
|
+
neither `MARKER` nor `NONE`. All seven class tests are therefore **invariant**, not merely
|
|
831
|
+
monotone. The skip walks are invariant because `QUOTEMARK ∩ INLINE-SPACE = ∅`.
|
|
832
|
+
|
|
833
|
+
**V1 is invariant too, as of spec 0.4.1, and this required changing V1 itself, not just arguing
|
|
834
|
+
about it.** Before 0.4.1, V1 compared raw code points, and replacing a neighbour's `U+0027` with
|
|
835
|
+
`U+2019` — the one substitution this lemma's hypothesis is actually exercised against by
|
|
836
|
+
`apostrophe`, via Corollary A1 below — could change whether that neighbour's code point equalled
|
|
837
|
+
a candidate's own `g`, which is exactly the comparison V1 makes. That was a real gap, not a
|
|
838
|
+
theorem this document previously proved: the claim once made here — that "V1's movement is
|
|
839
|
+
precisely what the gate certifies" — conflated two different sources of movement. The
|
|
840
|
+
certification gate (§3.5) re-derives candidates from `quotes`' *own* hypothetical render, within
|
|
841
|
+
one call; it has no visibility into a *different* rule editing the input between two separate
|
|
842
|
+
pipeline invocations, which is exactly what `apostrophe` does. §3.2's `V1ID` closes the actual
|
|
843
|
+
gap: `V1ID(U+0027) = V1ID(U+2019) = U+2019`, and `V1ID` is the identity elsewhere, so replacing a
|
|
844
|
+
`U+0027` neighbour with `U+2019` (or vice versa) changes neither side of any V1 comparison in
|
|
845
|
+
`V1ID`-space. V1's verdict is now invariant under Lemma A's hypothesis by the same construction
|
|
846
|
+
as the other seven tests, not by appeal to a mechanism (the gate) that cannot see the edit in
|
|
847
|
+
question. ∎
|
|
848
|
+
|
|
849
|
+
**The listed elision veto (spec 0.4.0) is covered by the same argument, not a new proof
|
|
850
|
+
obligation.** Every input it reads is one Lemma A already accounts for or one nothing in this
|
|
851
|
+
pipeline ever touches:
|
|
852
|
+
|
|
853
|
+
- `NARROW` membership of `cp[i]` and `cp[j]` — invariant, same as every test above.
|
|
854
|
+
- The elided content `cp[i+1 … i+k]` — a literal `LETTER` run (in the shipped `en-US` entry;
|
|
855
|
+
the schema permits any non-empty string, but an idiom's `elided` field is authored as
|
|
856
|
+
letters). `apostrophe` (R₆) never touches a `LETTER`; it replaces exactly one `SQ` with
|
|
857
|
+
exactly one U+2019 and nothing else (`apostrophe.md` §1). Unaffected by any rule ordered at
|
|
858
|
+
or before `quotes` for the same reason.
|
|
859
|
+
- The single `INLINE-SPACE` code point on each outer side, and the `left`/`right` `LETTER` runs
|
|
860
|
+
beyond it — `apostrophe` touches neither an `INLINE-SPACE` code point nor a `LETTER`; nothing
|
|
861
|
+
ordered before `quotes` does either (each earlier rule's own §4 "Must not touch" already
|
|
862
|
+
states this for ordinary prose letters and spacing it did not itself just emit, and none of
|
|
863
|
+
them emits a `LETTER`).
|
|
864
|
+
|
|
865
|
+
So substituting `g`'s own U+0027 for U+2019 at `i` or `j` (the only edit `apostrophe` can make
|
|
866
|
+
to this construction, and the only one reachable in this direction) changes none of the four
|
|
867
|
+
inputs above, and the veto's verdict on a given pair of indices is invariant under Lemma A's
|
|
868
|
+
hypothesis exactly as every other capability test is. Concretely: on the pipeline's second
|
|
869
|
+
pass, `rock ’n’ roll`'s two marks are U+2019 rather than U+0027, still `NARROW`, still bounded
|
|
870
|
+
by the identical literal `rock`/`n`/`roll` context — the veto fires identically, both marks are
|
|
871
|
+
vetoed again, `quotes` makes no pairing, and `apostrophe` does not act on U+2019 at all (§4).
|
|
872
|
+
The construction is a fixed point (§6 row P4 pins it).
|
|
873
|
+
|
|
874
|
+
**Modes.** `text` is the base case above. In `html` and `markdown`, `modes.md` §3.2's
|
|
875
|
+
concatenation-with-marker model means the veto's bounded lookaround can land on a `MARKER` (a
|
|
876
|
+
negative integer, in none of `LETTER`, `INLINE-SPACE`, `NARROW`) exactly where a real code
|
|
877
|
+
point would otherwise be — a context word split from its mark by an element or span boundary
|
|
878
|
+
fails the `INLINE-SPACE` test at that position and the veto does not fire (§6 rows H1–H3). This
|
|
879
|
+
is not a special case for modes: it is the same "word not found" outcome an ordinary document
|
|
880
|
+
boundary produces (`Llit`/`Rlit` = `NONE`), and `modes.md` §5's stability argument — no rule may
|
|
881
|
+
emit a code point at a position that changes how the next pass partitions spans — is
|
|
882
|
+
unaffected, because this veto neither emits anything nor changes any span boundary; it only
|
|
883
|
+
narrows which pairs `quotes` itself forms, which `modes.md` §5 part 3 already covers for the
|
|
884
|
+
rest of this rule's determinism.
|
|
885
|
+
|
|
886
|
+
> **Corollary A1 — `apostrophe` (R₆) is structurally invisible (spec 0.4.1: was claimed, not yet
|
|
887
|
+
> true, before `V1ID`).** `apostrophe` replaces U+0027 with U+2019 and nothing else. As a
|
|
888
|
+
> *neighbour*, Lemma A applies — **including its V1 clause**, now that V1 itself compares `V1ID`
|
|
889
|
+
> rather than raw code points. As the mark *itself*: both `U+0027` and `U+2019` are in `NARROW`,
|
|
890
|
+
> so the stack partition is unchanged; the medial veto is stated over `NARROW`, so its verdict is
|
|
891
|
+
> unchanged; neither is in `SPACE-RIGHT`/`SPACE-LEFT`, by constraint **Q-A**; and `V1ID` maps
|
|
892
|
+
> both to the same value, so V1's verdict on the mark's own candidacy is unchanged too. Every
|
|
893
|
+
> verdict is identical. This is a genuine **CO-S** discharge in the sense of
|
|
894
|
+
> `pipeline-idempotency.md` §5.1a — `E(apostrophe) = {U+2019}` is wholly inert for this rule,
|
|
895
|
+
> because `V1ID` makes it indistinguishable from the one code point it can only ever have
|
|
896
|
+
> replaced.
|
|
897
|
+
>
|
|
898
|
+
> **This corollary was wrong for one release** (spec 0.3.0–0.4.0): it asserted glyph-blindness
|
|
899
|
+
> covered V1 by deferring to Lemma A's text, and Lemma A's own proof at the time explicitly
|
|
900
|
+
> carved V1 out ("the only test that can move") and waved at the certification gate as the
|
|
901
|
+
> guarantor — a claim about a *different* mechanism (single-call self-re-derivation) standing in
|
|
902
|
+
> for a proof about *this* one (cross-call stability against a later rule). The gap was real: a
|
|
903
|
+
> `NARROW` candidate `g = U+2019` adjacent to an unmatched `U+0027` had a literal-neighbour
|
|
904
|
+
> comparison that read differently before and after `apostrophe` converted that neighbour on the
|
|
905
|
+
> *previous* pipeline pass, which is precisely `apostrophe` creating work for `quotes` —
|
|
906
|
+
> forbidden by `pipeline-idempotency.md` §2. Found by `fast-check` in `de-CH`
|
|
907
|
+
> (`spec/fixtures/de-CH.json`, `de-ch-quotes-041-cobug-apostrophe-v1`).
|
|
908
|
+
>
|
|
909
|
+
> **Reachability, stated precisely rather than as a blanket "every locale."** The *algorithm* —
|
|
910
|
+
> `QUOTEMARK`, V1, and `apostrophe`'s case ladder — is locale-independent, so the same synthetic
|
|
911
|
+
> witness can be constructed and passed through every locale's data, and
|
|
912
|
+
> `tests/rules/quotes.test.ts` does exactly that as a regression sweep. That is a claim about
|
|
913
|
+
> algorithmic constructibility, not evidence that the *pre-fix* defect actually reproduced in
|
|
914
|
+
> every locale — it was verified directly only in `de-CH`, where `« »`/`‹ ›` (neither U+2019) make
|
|
915
|
+
> the witness adversarial rather than naturally occurring. Separately: `en-GB`'s primary `close`
|
|
916
|
+
> and `fi`/`sv`'s secondary pair (both sides) *are* U+2019 as those locales' own prescribed
|
|
917
|
+
> glyphs, which is what makes this shape especially plausible from ordinary, correctly-typed
|
|
918
|
+
> content there — not merely reachable by construction — even though no canonical fixture is
|
|
919
|
+
> pinned in those locales for it (§3.2's `V1ID` discussion states the same reachability
|
|
920
|
+
> distinction and gives the empirical delta of the fix, "V1ID empirical audit" below).
|
|
921
|
+
>
|
|
922
|
+
> **On the elision-vs-closer tradeoff.** No U+2019-specific "can never open" restriction is
|
|
923
|
+
> applied. Such a restriction was considered and rejected: `fi`/`sv`'s secondary pair is U+2019
|
|
924
|
+
> on *both* sides, and forcing `canOpen = false` for every U+2019 would mean an already-correct
|
|
925
|
+
> `fi`/`sv` secondary pair could never be *recognised* as a pair at all — only ever declined —
|
|
926
|
+
> which directly contradicts mandate 1's purpose for that locale family. The cost of not
|
|
927
|
+
> restricting is stated in §4 and verified in §6 (rows E1–E4): it is the same pre-existing
|
|
928
|
+
> `rock 'n' roll`-family ambiguity 0.1.0 already documented for straight input, now also
|
|
929
|
+
> reachable via curly input, not a new class of damage.
|
|
930
|
+
|
|
931
|
+
### Lemma B — space-inertness at every position `nbsp` can reach
|
|
932
|
+
|
|
933
|
+
> Inserting a code point of `INLINE-SPACE`, or converting one `INLINE-SPACE` code point to
|
|
934
|
+
> another, at any position `nbsp` can act on, changes no capability of any candidate.
|
|
935
|
+
|
|
936
|
+
*Proof.* Conversion is trivial: every test reads class membership and U+0020, U+00A0, U+202F are
|
|
937
|
+
all in `INLINE-SPACE`. For insertion, `nbsp` has exactly three insertion sites (`nbsp.ts`
|
|
938
|
+
`claimInsertion`):
|
|
939
|
+
|
|
940
|
+
1. **N8, immediately right of an occurrence of `g ∈ SPACE-RIGHT`.**
|
|
941
|
+
- *`g` itself*: `canOpen` reads `Rskip` (skips) and `openLeft` (untouched); `canClose` reads
|
|
942
|
+
`Lskip` (untouched) and `closeRight`, which is `Rskip` because `g ∈ SPACE-RIGHT` (skips).
|
|
943
|
+
All four invariant.
|
|
944
|
+
- *the candidate `c` whose `Llit` was `g`*: `c.canOpen` reads `openLeft`, which is `Llit`
|
|
945
|
+
(unless `c ∈ SPACE-LEFT`, in which case it skips): `g ∈ OPENISH` is accepted,
|
|
946
|
+
`INLINE-SPACE ⊂ SPACELIKE` is accepted — same verdict. `c.canClose` reads `Lskip`, which
|
|
947
|
+
skips the insertion back to `g` — same. Right-hand reads untouched.
|
|
948
|
+
- no other index has a changed neighbourhood.
|
|
949
|
+
2. **N8, immediately left of `h ∈ SPACE-LEFT`.** Mirror image: `h.canOpen` reads
|
|
950
|
+
`openLeft = Lskip` (skips, because `h ∈ SPACE-LEFT`); `h.canClose` reads `Lskip`. The
|
|
951
|
+
candidate whose `Rlit` was `h` reads `closeRight = Rlit`, moving from `h ∈ CLOSEISH`
|
|
952
|
+
(accepted) to `INLINE-SPACE` (accepted) — same verdict; and `Rskip`, which skips back to `h`.
|
|
953
|
+
3. **N1/N2, immediately left of a mark `m ∈ nbsp.beforePunctuation ∪
|
|
954
|
+
nbsp.narrowBeforePunctuation`.** The only quote mark whose neighbourhood changes is a `g`
|
|
955
|
+
with `Rlit = m`. Its `closeRight` moves from `m` to `INLINE-SPACE`; by constraint **Q-P**,
|
|
956
|
+
`m ∈ CLOSEISH`, so both are accepted — same verdict. Its `Rskip` skips back to `m`. (`nbsp`'s
|
|
957
|
+
own step 3 already declines to insert after an opening glyph.)
|
|
958
|
+
|
|
959
|
+
No other sub-rule inserts; N3–N7, N9, N10 only convert. ∎
|
|
960
|
+
|
|
961
|
+
**V1's `gapInsertable` clause (§3.2) is what extends this proof to V1 itself.** The four
|
|
962
|
+
capability tests above are the ones Lemma B's statement covers directly, but V1 is also a test a
|
|
963
|
+
candidate's own capabilities depend on, and it compares *code points*, not class membership —
|
|
964
|
+
Lemma B's "accepts `SPACELIKE` either way" argument does not apply to it. `gapInsertable` is
|
|
965
|
+
defined to hold exactly when `g` is a member of `SPACE-RIGHT` or `SPACE-LEFT` — precisely the
|
|
966
|
+
condition under which `nbsp`'s insertion sites 1 and 2 above can place a code point at the one
|
|
967
|
+
gap V1's second clause reads across — so V1's verdict is invariant to that insertion by the same
|
|
968
|
+
argument, case by case: at insertion site 1, the candidate immediately right of `g' ∈ SPACE-RIGHT`
|
|
969
|
+
has `gapInsertable = true` when its own glyph equals `g'` (the only case V1's second clause can
|
|
970
|
+
ever fire for it), and the literal-vs-skipped reads agree on whether that neighbour is `g'`,
|
|
971
|
+
exactly as case 1 above already shows for `canOpen`/`canClose`. Insertion site 2 is the mirror
|
|
972
|
+
image for `SPACE-LEFT`. Insertion site 3 (N1/N2) never inserts adjacent to a `QUOTEMARK` glyph on
|
|
973
|
+
the side V1 reads (it inserts beside listed punctuation, never beside a quote mark's own
|
|
974
|
+
same-glyph neighbour), so it cannot affect `gapInsertable`'s condition at all.
|
|
975
|
+
|
|
976
|
+
> **Corollary B1.** A run-1/run-2 asymmetry that used to arise from `quotes` reading a
|
|
977
|
+
> `SPACE-RIGHT` glyph's right side literally on one run and skipping it on the other cannot arise
|
|
978
|
+
> under §3.2: the read is `Rskip` on **both** runs, so `« **"` pairs on the first application and
|
|
979
|
+
> re-pairs identically on the second.
|
|
980
|
+
|
|
981
|
+
### Certification — the theorem
|
|
982
|
+
|
|
983
|
+
> **Theorem.** For every input `x` and every locale, `Q(Q(x)) = Q(x)`.
|
|
984
|
+
|
|
985
|
+
*Proof.* Let `A` be the accepted set and `y = Q(x) = render(x, A)`.
|
|
986
|
+
|
|
987
|
+
**Case `A = ∅`.** Then `y = x`. `Q` is a deterministic function of its input and the locale, with
|
|
988
|
+
no state (`ARCHITECTURE.md` §7), so `Q(y) = Q(x) = x = y`. ∎
|
|
989
|
+
|
|
990
|
+
**Case `A ≠ ∅`.** The loop exited on the equality branch, so `G₀(y) = π(A)`, where `G₀` denotes
|
|
991
|
+
passes 1–2. Run `Q` on `y`. Passes 1–2 give `A₀' = G₀(y) = π(A)`. The gate's first round computes
|
|
992
|
+
`render(y, π(A))`, and this equals `y`:
|
|
993
|
+
|
|
994
|
+
- **Glyphs.** `π` is order-preserving, so `depth_{π(A)}(π(p)) = depth_A(p)` for every `p`; the
|
|
995
|
+
assigned pair is the same, and `y` already carries those glyphs at those indices by
|
|
996
|
+
construction of `render`.
|
|
997
|
+
- **Inner spacing.** For a pair whose `innerSpace ≠ "none"`, `render` never touches spacing, on
|
|
998
|
+
either run. For `innerSpace = "none"`: a run that was deleted is gone, so there is nothing to
|
|
999
|
+
delete; a run that was declined is still present and its landing is unchanged — `render` maps
|
|
1000
|
+
`QUOTEMARK` to `QUOTEMARK` and leaves `ALNUM` alone, so a landing outside `DELETE-LANDING`
|
|
1001
|
+
stays outside it, and the same run is declined again.
|
|
1002
|
+
|
|
1003
|
+
Therefore `G₀(render(y, π(A))) = G₀(y) = π(A)`, the gate certifies in round 1, and
|
|
1004
|
+
`Q(y) = render(y, π(A)) = y`. ∎
|
|
1005
|
+
|
|
1006
|
+
### Composition obligation
|
|
1007
|
+
|
|
1008
|
+
This rule is **R₅**; the obligation runs against `spaces`, `ellipsis`, `dashes` and `hyphen`, and
|
|
1009
|
+
its own `I₅` must survive `apostrophe`, `symbols` and `nbsp`.
|
|
1010
|
+
|
|
1011
|
+
**What this rule emits.** `E(quotes)` = the four locale quote glyphs, each replacing exactly one
|
|
1012
|
+
`QUOTEMARK` at the same index. Plus deletions of a maximal `INLINE-SPACE` run whose near side is
|
|
1013
|
+
a quote glyph and whose far side is in `ALNUM ∪ QUOTEMARK`. **Nothing is inserted.**
|
|
1014
|
+
|
|
1015
|
+
**The listed elision veto (spec 0.4.0) changes none of the below.** It adds no emission — a
|
|
1016
|
+
vetoed pair of marks is simply never paired, so `E(quotes)` is exactly what it was — and can
|
|
1017
|
+
only shrink the set of pairs this rule forms, never grow it. Every discharge that follows was
|
|
1018
|
+
argued against a `quotes` whose positive behaviour is a superset of the current one, so each
|
|
1019
|
+
still holds without re-verification.
|
|
1020
|
+
|
|
1021
|
+
- **Against `I₁` (`spaces`).** Discharged. No U+0020 is emitted, so no S-b/S-c/S-d position can
|
|
1022
|
+
be created. Deletion removes a *maximal* run, so it cannot bring two U+0020 into contact
|
|
1023
|
+
(S-a); its neighbours after deletion are a quote glyph and an `ALNUM`/`QUOTEMARK`, both
|
|
1024
|
+
`CONTENT`.
|
|
1025
|
+
- **Against `I₂` (`ellipsis`).** Discharged. No U+002E or U+2026 is emitted; the deletion's
|
|
1026
|
+
landing is in `ALNUM ∪ QUOTEMARK`, which contains no dot, so no two dot runs are joined.
|
|
1027
|
+
- **Against `I₃` (`dashes`).** Discharged **structurally**. A quote glyph is in none of `DASH`,
|
|
1028
|
+
`INERT-DASH`, `DIGIT`, `SPACE`, `NOBREAK-SPACE`, `JOINER`, `ROMAN` or `LETTER`, so a
|
|
1029
|
+
`cp[L]`/`cp[R]`/`before`/`after` moving from one quote mark to another moves nowhere. The
|
|
1030
|
+
deletion's two sides are a quote glyph and a `DELETE-LANDING` member, **neither in
|
|
1031
|
+
`DASH ∪ INERT-DASH`** — so no dash run is ever adjacent to a deleted run, no token's
|
|
1032
|
+
`lsp`/`rsp` changes, and the landing separates the deletion from any dash run by at least one
|
|
1033
|
+
code point. This is the discharge the naive "delete the inner space unconditionally"
|
|
1034
|
+
formulation failed, with witness `" --x"` → `“--x”` → `“—x”` (row D1, §6).
|
|
1035
|
+
- **Against `I₄` (`hyphen`).** Discharged. `hyphen`'s boundary test is `¬WORDISH`, and both a
|
|
1036
|
+
U+0020 and a quote glyph are outside `WORDISH = ALNUM ∪ HYPHENISH`. The deletion never lands
|
|
1037
|
+
inside a word, only against its first or last code point.
|
|
1038
|
+
- **`I₅` against `apostrophe` (R₆).** Discharged by Corollary A1 (CO-S).
|
|
1039
|
+
- **`I₅` against `nbsp` (R₈).** Discharged by Lemma B (CO-S-shaped, over `nbsp`'s insertion
|
|
1040
|
+
*sites* rather than its emission alphabet alone — the alphabet by itself is not enough, because
|
|
1041
|
+
a `SPACELIKE` insertion is *not* inert at an arbitrary position; it is inert at exactly the
|
|
1042
|
+
positions `nbsp` can reach).
|
|
1043
|
+
- **`I₅` against `symbols` (R₇).** **Known-weak: a case analysis, not CO-S.** `symbols`
|
|
1044
|
+
collapses `(c)`/`(r)`/`(tm)` to a single sign, deleting code points, and converts `x` to `×`
|
|
1045
|
+
between numerals. A quote mark adjacent to a collapsed span sees `(` on its right (moving to
|
|
1046
|
+
`©`, which is neither `OPENISH` nor `CLOSEISH`) or `)` on its left (moving to `©`). Checked
|
|
1047
|
+
over all four capability tests, every reachable verdict coincides: `canOpen`'s right test
|
|
1048
|
+
accepts `(` and accepts `©`; `canClose`'s right test rejects both; `canOpen`'s left test
|
|
1049
|
+
rejects `)` and rejects `©`; `canClose`'s left test accepts both as non-`NONE`. This is not
|
|
1050
|
+
made structural — the sweep is the control, per §5.1a of `pipeline-idempotency.md`.
|
|
1051
|
+
|
|
1052
|
+
---
|
|
1053
|
+
|
|
1054
|
+
## 6. Worked examples
|
|
1055
|
+
|
|
1056
|
+
`␣` = U+0020, `⍽` = U+00A0, `⟶` = no change. Every row is verified by running
|
|
1057
|
+
`src/rules/quotes.ts` against the fixture of the same name in `spec/fixtures/`; none is
|
|
1058
|
+
hand-traced-only.
|
|
1059
|
+
|
|
1060
|
+
### `en-US` — primary `“ ”`, secondary `‘ ’`, `innerSpace: none`
|
|
1061
|
+
|
|
1062
|
+
| # | Input | Output | Why |
|
|
1063
|
+
| --- | --- | --- | --- |
|
|
1064
|
+
| 1 | `She said "hello" twice.` | `She said “hello” twice.` | depth 1 → primary |
|
|
1065
|
+
| 3 | `'He said "no" to me,' she noted.` | `“He said ‘no’ to me,” she noted.` | depth, not width, selects the pair; the gate certifies despite both streams swapping width |
|
|
1066
|
+
| 4 | `He said "hi. She said "bye."` | `He said "hi. She said “bye.”` | the first mark's `closeRight` is the letter `h`, so it is `canOpen` only and cannot swallow the real quotation |
|
|
1067
|
+
| 5 | `Don't touch it — it's the '90s.` | ⟶ | two medial vetoes; `'90s` is `canOpen`, unmatched, handed to `apostrophe` |
|
|
1068
|
+
| 6 | `He is 6' 2" tall.` | ⟶ | both marks `canClose` only, both stacks empty |
|
|
1069
|
+
| 7b | `"""a""` | ⟶ | every mark has an identical literal neighbour, V1 drops all five |
|
|
1070
|
+
| 7c | `"a "" b"` | `“a "" b”` | the inner `""` is vetoed; the outer pair converts; the gate certifies |
|
|
1071
|
+
| 8 | `“Already curly,” he said, and "this too."` | `“Already curly,” he said, and “this too.”` | the curly pair is now a real depth-1 pair, not an ignored one |
|
|
1072
|
+
| A | `“He said ‘no’ twice.”` | ⟶ | already-correct nesting is a no-op |
|
|
1073
|
+
| M2a | `" hello"` | `“hello”` | mandate 2: `canOpen` reads `Rskip = h`; the inner run's landing is `h ∈ ALNUM`, deleted |
|
|
1074
|
+
| M2b | `"hello "` | `“hello”` | mirror case |
|
|
1075
|
+
| D1 | `" --x"` | `“ --x”` | the inner run's landing is `-`, outside `DELETE-LANDING` → deletion declined, glyphs still convert |
|
|
1076
|
+
| E1 | `’90s were fun,’ he said` | `“90s were fun,” he said` | **elision-vs-closer, curly input** — see §4's residual-risk note |
|
|
1077
|
+
| E1s | `'90s were fun,' he said` | `“90s were fun,” he said` | the straight-input control: the same corruption, pre-existing since 0.1.0 |
|
|
1078
|
+
|
|
1079
|
+
#### Listed elision veto — `en-US`, `elisionIdioms = [{ left: "rock", elided: "n", right: "roll" }]`
|
|
1080
|
+
|
|
1081
|
+
Positive — the full context matches and both marks are vetoed, so `apostrophe` (order 50) renders
|
|
1082
|
+
each independently:
|
|
1083
|
+
|
|
1084
|
+
| # | Input | Output | Why |
|
|
1085
|
+
| --- | --- | --- | --- |
|
|
1086
|
+
| P1 | `rock 'n' roll` | `rock ’n’ roll` | canonical form |
|
|
1087
|
+
| P2 | `Rock 'n' roll is great.` | `Rock ’n’ roll is great.` | sentence-initial capital — first-code-point leniency on `left` |
|
|
1088
|
+
| P3 | `I love rock 'n' roll.` | `I love rock ’n’ roll.` | trailing period after `right` does not block the word-boundary test |
|
|
1089
|
+
| P4 | `rock ’n’ roll` | ⟶ | already-curly: U+2019 is `NARROW`, the veto fires identically on the already-correct form, `apostrophe` does not act on U+2019 at all — a fixed point (§5) |
|
|
1090
|
+
|
|
1091
|
+
Negative **for the listed mechanism only** — the elided content matches but the full
|
|
1092
|
+
`left`/`elided`/`right` context does not, so no configured idiom fires. Since spec 1.1.0 every row
|
|
1093
|
+
below is nevertheless matched by the universal medial-`n` veto (§3.2), so every one converts, and
|
|
1094
|
+
this table's job is to show where the two mechanisms differ rather than where the marks fall
|
|
1095
|
+
through. N1 and N2 remain the rows that matter most: they were reported as falsely elided by the
|
|
1096
|
+
earlier, context-free `elisionForms` design, and 1.1.0 accepts that specific cost knowingly and
|
|
1097
|
+
narrowly (§3.2, "Accepted false positive") — which is not the same as readmitting the unbounded
|
|
1098
|
+
word list that was rejected.
|
|
1099
|
+
|
|
1100
|
+
**These five outputs were last accurate before spec 0.5.0.** The table was not updated when 0.5.0's
|
|
1101
|
+
veto landed, so as written it described neither 0.5.0's behaviour nor the fixtures'; the outputs
|
|
1102
|
+
below are 1.1.0's, verified against `spec/fixtures/en-US.json`.
|
|
1103
|
+
|
|
1104
|
+
| # | Input | Output | Why |
|
|
1105
|
+
| --- | --- | --- | --- |
|
|
1106
|
+
| N1 | `The letter 'n' is common.` | `The letter ’n’ is common.` | no `left`/`right` context at all, so no idiom fires — but the universal veto matches the bare `'n'` and `apostrophe` converts both marks. The accepted false positive of §3.2, stated there in full |
|
|
1107
|
+
| N2 | `He said "press 'n' now".` | `He said “press ’n’ now”.` | the same shape one level deeper; §3.2's predicate reads a mark's immediate neighbours and knows nothing of enclosing pairs, so nesting depth never reaches it |
|
|
1108
|
+
| N3 | `rock 'n' pop` | `rock ’n’ pop` | `right` is `pop`, not `roll` — no configured idiom matches, and since 1.1.0 none is needed |
|
|
1109
|
+
| N4 | `fish 'n' chips` | `fish ’n’ chips` | not sourced or listed: no `{ left: "fish", … }` entry exists in `en-US.json`, and none is required for the marks to convert |
|
|
1110
|
+
| N5 | `rock 'N' roll` | `rock ’N’ roll` | `elided` is matched **exactly** by the listed mechanism, so `N ≠ n` still fails *it*; the universal veto matches U+004E as well as U+006E, and converts both marks |
|
|
1111
|
+
|
|
1112
|
+
**Known, accepted residual ambiguity (not a negative-fixture guarantee).** The full-context
|
|
1113
|
+
requirement narrows the false-positive surface to genuine surface-form coincidence; it cannot
|
|
1114
|
+
see authorial intent, only code points. A document that explicitly states it means three
|
|
1115
|
+
separate tokens — a word, a genuinely quoted single letter, another word — and then reproduces
|
|
1116
|
+
the identical surface sequence for illustration still gets that final occurrence vetoed:
|
|
1117
|
+
|
|
1118
|
+
| # | Input | Output | Why |
|
|
1119
|
+
| --- | --- | --- | --- |
|
|
1120
|
+
| N6 | `The sequence is the word rock, the quoted letter 'n', and the word roll: rock 'n' roll.` | `The sequence is the word rock, the quoted letter “n”, and the word roll: rock ’n’ roll.` | **the first `'n'` is correctly NOT vetoed**, and by both mechanisms independently: the mark closing it is followed by a comma rather than an `INLINE-SPACE`, which fails the universal veto's outer test (§3.2), and `left` is `letter` rather than `rock`, which fails the listed idiom's. It pairs as an ordinary quotation. **The second `rock 'n' roll` IS vetoed and pinned as known, accepted risk, not fixed.** Under the intent the sentence itself states, this is a genuine semantic false positive — accepted as an unavoidable one for this bounded, surface-form contract, since `left`/`elided`/`right` and the single permitted `INLINE-SPACE` are all it is defined over, and none of them can encode "this is a demonstration, not an utterance of the idiom." The implementation still conforms exactly to its own matcher here: the bytes are the idiom's bytes, byte for byte, and the matcher cannot see past that. Recorded so the exposure is visible in the conformance suite rather than discovered in someone's content (§3.2's residual-ambiguity note) |
|
|
1121
|
+
|
|
1122
|
+
#### Elision vetoes across span boundaries — `html`, `markdown`
|
|
1123
|
+
|
|
1124
|
+
`modes.md` §3.2's boundary marker is not `INLINE-SPACE`, so a **context word** separated from its
|
|
1125
|
+
mark by an element or span boundary does not satisfy the listed idiom's outer test, and that idiom
|
|
1126
|
+
does not fire. The universal medial-`n` veto (§3.2) reads no context word at all — only the single
|
|
1127
|
+
enclosed code point and the space immediately outside each mark — so a boundary further out is
|
|
1128
|
+
invisible to it and it fires on all four rows. Since spec 1.1.0 the split cases therefore converge
|
|
1129
|
+
with the unsplit one instead of diverging from it.
|
|
1130
|
+
|
|
1131
|
+
**H1–H3's outputs were last accurate before spec 0.5.0** and are corrected here for the same
|
|
1132
|
+
reason as N1–N5 above; they are verified against `spec/fixtures/en-US.json`.
|
|
1133
|
+
|
|
1134
|
+
| # | Mode | Input | Output | Why |
|
|
1135
|
+
| --- | --- | --- | --- | --- |
|
|
1136
|
+
| H0 | `html` | `<p>rock 'n' roll</p>` | `<p>rock ’n’ roll</p>` | idiom whole within one text node — positive control, both mechanisms match |
|
|
1137
|
+
| H1 | `html` | `<p><em>rock</em> 'n' roll</p>` | `<p><em>rock</em> ’n’ roll</p>` | `left` word is inside a different span, so the *listed* idiom does not match; the universal veto never looks for a `left` word, and the opening mark's own left neighbour is still a real `INLINE-SPACE` |
|
|
1138
|
+
| H2 | `html` | `<p>rock 'n' <em>roll</em></p>` | `<p>rock ’n’ <em>roll</em></p>` | mirror case on `right` |
|
|
1139
|
+
| H3 | `markdown` (`commonmark`) | `*rock* 'n' roll\n` | `*rock* ’n’ roll\n` | `left` word is inside an emphasis span; same reasoning as H1, and the result matches `text` mode on the same characters split the same way |
|
|
1140
|
+
|
|
1141
|
+
### `en-GB` — primary `‘ ’`, secondary `“ ”`
|
|
1142
|
+
|
|
1143
|
+
| # | Input | Output | Why |
|
|
1144
|
+
| --- | --- | --- | --- |
|
|
1145
|
+
| B2 | `‘""-"-"` | ⟶ | width drift interacting with the same-glyph adjacency veto; the gate declines to ∅ |
|
|
1146
|
+
| E2 | `’90s were fun,’ he said` | `‘90s were fun,’ he said` | elision-vs-closer, curly input |
|
|
1147
|
+
| E2s | `'90s were fun,' he said` | `‘90s were fun,’ he said` | straight-input control |
|
|
1148
|
+
|
|
1149
|
+
### `fi` — primary `” ”` (U+201D both sides), secondary `’ ’` (U+2019 both sides)
|
|
1150
|
+
|
|
1151
|
+
| # | Input | Output | Why |
|
|
1152
|
+
| --- | --- | --- | --- |
|
|
1153
|
+
| 9 | `Hän sanoi "moi" ja lähti.` | `Hän sanoi ”moi” ja lähti.` | the pairing, not the glyph, knows which is which |
|
|
1154
|
+
| B | `”Hän sanoi ’moi’”, totesin.` | ⟶ | same-glyph already-correct nesting is a no-op |
|
|
1155
|
+
| 11 | `"Hän sanoi 'moi'", totesin.` | `”Hän sanoi ’moi’”, totesin.` | row B's input reached from straight marks |
|
|
1156
|
+
| U1 | `”Hän sanoi ”moi” ja lähti.` | ⟶ | unbalanced same-glyph: three `”`, the stray leading one is left alone |
|
|
1157
|
+
| E3 | `’90s were fun,’ he said` | `”90s were fun,” he said` | elision-vs-closer, curly input |
|
|
1158
|
+
| E3s | `'90s were fun,' he said` | `”90s were fun,” he said` | straight-input control |
|
|
1159
|
+
|
|
1160
|
+
### `sv` — same shape as `fi`
|
|
1161
|
+
|
|
1162
|
+
| # | Input | Output | Why |
|
|
1163
|
+
| --- | --- | --- | --- |
|
|
1164
|
+
| E4 | `’90s were fun,’ he said` | `”90s were fun,” he said` | elision-vs-closer, curly input |
|
|
1165
|
+
| E4s | `'90s were fun,' he said` | `”90s were fun,” he said` | straight-input control |
|
|
1166
|
+
|
|
1167
|
+
### `fr` — primary `« »`, `innerSpace: "nbsp"`; secondary `“ ”`, `innerSpace: "none"`
|
|
1168
|
+
|
|
1169
|
+
| # | Input | Output *(after `nbsp`)* | Why |
|
|
1170
|
+
| --- | --- | --- | --- |
|
|
1171
|
+
| 12 | `Il a dit "bonjour".` | `Il a dit «⍽bonjour⍽».` | `quotes` emits the glyphs; `nbsp` the U+00A0 |
|
|
1172
|
+
| 13 | `Il a dit «␣bonjour␣».` | `Il a dit «⍽bonjour⍽».` | mandate 2's second half: the pair *forms on the first application* and produces no edit from `quotes` because the glyphs are already right |
|
|
1173
|
+
| 3a | `«␣**"` | `«⍽**⍽»` *(after `nbsp`)* | width drift interacting with a run-1 `nbsp` insertion; `quotes` alone certifies the pair with one edit |
|
|
1174
|
+
|
|
1175
|
+
### `fr-CA`, `html` mode
|
|
1176
|
+
|
|
1177
|
+
| # | Input | Output | Why |
|
|
1178
|
+
| --- | --- | --- | --- |
|
|
1179
|
+
| 3b | `"<p class="x">«[t](u)` | ⟶ | the mode adapter splices a `MARKER` between the two text spans; `«`'s `closeRight` reads `Rskip` (because `« ∈ SPACE-RIGHT`) and finds `[`, so it is never `canClose` and no pair forms, on either application |
|
|
1180
|
+
|
|
1181
|
+
### `ru` — primary `« »`, secondary `„ “`
|
|
1182
|
+
|
|
1183
|
+
| # | Input | Output | Why |
|
|
1184
|
+
| --- | --- | --- | --- |
|
|
1185
|
+
| 14 | `Он сказал: "это 'моё' дело".` | `Он сказал: «это „моё“ дело».` | depths 1 and 2 |
|
|
1186
|
+
| B1 | `'"‘` | ⟶ | width drift with no candidate left unmatched to be swallowed; the gate declines to ∅ in one round |
|
|
1187
|
+
|
|
1188
|
+
### `el` — primary `« »`, secondary `“ ”` (both `WIDE`)
|
|
1189
|
+
|
|
1190
|
+
| # | Input | Output | Why |
|
|
1191
|
+
| --- | --- | --- | --- |
|
|
1192
|
+
| 16 | `"Είπε 'όχι' σε μένα", σημείωσε.` | `«Είπε “όχι” σε μένα», σημείωσε.` | width drift with no gate intervention — evidence the gate's cost is near zero |
|
|
1193
|
+
| 18 | `Είπε "σ' αυτό το βιβλίο" χθες.` | `Είπε «σ' αυτό το βιβλίο» χθες.` | the elision is `NARROW`, the quotation `WIDE`, different stacks — the elision survives untouched for `apostrophe` to convert on the next rule |
|
|
1194
|
+
|
|
1195
|
+
### `de-DE` — primary `„ “`, secondary `‚ ‘`
|
|
1196
|
+
|
|
1197
|
+
| # | Input | Output | Why |
|
|
1198
|
+
| --- | --- | --- | --- |
|
|
1199
|
+
| C | `Er sagte «Wort» leise.` | `Er sagte „Wort“ leise.` | foreign guillemets convert |
|
|
1200
|
+
| C2 | `Er sagte « Wort » leise.` | `Er sagte „Wort“ leise.` | same, plus both inner runs deleted (landings `W` and `t`) |
|
|
1201
|
+
|
|
1202
|
+
### Adversarial sweep — `ru`/`el`/`fr`/`fr-CA` (§7 item 9)
|
|
1203
|
+
|
|
1204
|
+
These four locales' primary and secondary pairs are **both `WIDE`**, so the width-keyed stack
|
|
1205
|
+
split does not independently separate a re-scanned primary stream from a re-scanned secondary
|
|
1206
|
+
stream the way it does for `en-US`/`en-GB`/`fi`/`sv`. An exhaustive sweep over
|
|
1207
|
+
`{«, », „, “, ", ', a, ␣}` (the locale's own glyphs plus the shared control alphabet) to length 5
|
|
1208
|
+
— 149,796 inputs across the four locales — found **zero** idempotency failures; the medial and
|
|
1209
|
+
same-glyph vetoes plus the certification gate together protect already-curly interleaved text in
|
|
1210
|
+
this family, even without an independent width partition. Recorded as measured, not asserted.
|
|
1211
|
+
|
|
1212
|
+
---
|
|
1213
|
+
|
|
1214
|
+
## 7. Open questions, and what could not be closed
|
|
1215
|
+
|
|
1216
|
+
Ordered by how much this matters.
|
|
1217
|
+
|
|
1218
|
+
1. **Mandate 1's false-positive surface is genuinely larger, and M4 is where it will show.** A
|
|
1219
|
+
lone `«` can now pair with a distant `"` (row 3a converts an input 0.1.0 left alone). Any
|
|
1220
|
+
document with an odd stray typographic quote — a `»` used as a bullet, a `’` used as a prime,
|
|
1221
|
+
a `«` in a citation — can recruit a partner from arbitrarily far away, and unlike a straight
|
|
1222
|
+
mark the result is not visibly "unfinished" to a proofreader. This is a **product** risk, not
|
|
1223
|
+
an idempotency one; run the M4 corpus against this change specifically before anything else.
|
|
1224
|
+
2. **The `symbols` (R₇) discharge is a case analysis**, not structural. Frequency: low (needs a
|
|
1225
|
+
quote mark literally adjacent to a `(c)`/`(r)`/`(tm)` span), but not closed structurally.
|
|
1226
|
+
3. **The elision-vs-closer tradeoff (§4, §5 Corollary A1, §6 rows E1–E4) is verified, not
|
|
1227
|
+
eliminated.** A leading elision that has already been rendered U+2019 and finds no partner of
|
|
1228
|
+
its own can still be recruited as an opener by an unrelated closer elsewhere in the text —
|
|
1229
|
+
that is a different mechanism from item 8's closed-idiom case below (`quotes.elisionIdioms`,
|
|
1230
|
+
spec 0.4.0) and is **not** closed by it: the veto only fires on a *listed idiom's full
|
|
1231
|
+
left/elided/right context* found together, not on a lone leading elision separated from any
|
|
1232
|
+
partner by arbitrary distance. No fix is proposed here.
|
|
1233
|
+
4. **The gate's forced-removal clause's necessity is asserted, not demonstrated with a concrete
|
|
1234
|
+
witness.** ~850,000 swept inputs did not trigger it. Implemented anyway (free insurance); if a
|
|
1235
|
+
future witness fires it, record it as a fixture.
|
|
1236
|
+
5. **Declined deletions leave visibly sloppy output** (`" --x"` → `“ --x”`, row D1). This is the
|
|
1237
|
+
correct side to be wrong on (the alternative is a confirmed `I₃` violation) but is a real,
|
|
1238
|
+
acknowledged gap in mandate 2's coverage.
|
|
1239
|
+
6. **`" "` and `«»` are prevented by an explicit vacuity guard, not by construction**, unlike
|
|
1240
|
+
0.1.0 where an empty pair was structurally impossible. Weaker footing than 0.1.0 had, though
|
|
1241
|
+
the gate catches any resulting instability.
|
|
1242
|
+
7. **Worst case is `O(n²)`** (gate rounds × pass cost). A pathological span of alternating quote
|
|
1243
|
+
marks is quadratic; cap rounds if this matters in practice.
|
|
1244
|
+
8. **`rock 'n' roll` — closed for every locale in spec 1.1.0 by the universal medial-`n` veto
|
|
1245
|
+
(§3.2), by operator decision that the idiom is international rather than a locale fact.** It
|
|
1246
|
+
was closed for `en-US` alone in spec 0.4.0 via `quotes.elisionIdioms =
|
|
1247
|
+
[{ left: "rock", elided: "n", right: "roll" }]` (Chicago Manual of Style Online on the
|
|
1248
|
+
mark's function, American Heritage Dictionary on the spaced variant's existence — see
|
|
1249
|
+
`en-US.json` `sources`), and that entry is retained and still cited; it now vetoes a strict
|
|
1250
|
+
subset of what the universal veto vetoes.
|
|
1251
|
+
**What 1.1.0 knowingly accepts in exchange** is the cost that sank a first design during the
|
|
1252
|
+
development of 0.4.0 (a bare `elisionForms = ["n"]` word list, matching only the elided
|
|
1253
|
+
content, prototyped and rejected before any release or commit — it never shipped and no user
|
|
1254
|
+
ever ran it): ordinary quotations of the letter *n* are elided — `The letter 'n' is common.`,
|
|
1255
|
+
`He said "press 'n' now".`, §6 rows N1/N2. The reasoning that rejected it in 0.4.0 was not
|
|
1256
|
+
wrong; what changed is the alternative it was measured against. In 0.4.0 the alternative was
|
|
1257
|
+
`elisionIdioms`' three-part contract, which is strictly better for `en-US`. From 0.5.0 the
|
|
1258
|
+
de-facto alternative for the other nine locales was the general 1–3-`LETTER` shape veto, which
|
|
1259
|
+
bought the same protection by declining **every** short quotation in **every** locale —
|
|
1260
|
+
a much larger and commoner error class than quoting a single letter, and one this repository's
|
|
1261
|
+
own promo examples tripped over in `ru`, `fr`, `fr-CA`, `fi` and `sv`. 1.1.0 takes the smaller
|
|
1262
|
+
error knowingly, and N1/N2 stay pinned so it is visible rather than forgotten.
|
|
1263
|
+
The evidentiary notes that governed per-locale entries are retained because they still govern
|
|
1264
|
+
`elisionIdioms` itself, which is unchanged: `en-GB` has no independent citation and carries an
|
|
1265
|
+
extra risk the other locales do not, since its primary pair *is* the single quote, so `'n'` is
|
|
1266
|
+
more plausible as genuine quoted dialogue there than in a locale whose primary pair is double;
|
|
1267
|
+
and `fi`/`sv` write the idiom **closed up** (`rock'n'roll`) in their own dictionaries, so a
|
|
1268
|
+
spaced-form entry there would be evidenced-inert. Neither observation blocks the universal
|
|
1269
|
+
veto, which rests on the operator decision rather than on per-locale citation. **Still open,
|
|
1270
|
+
all locales:** `en-GB` single-first (now more consequential under mandate 1); quotation across
|
|
1271
|
+
a paragraph boundary; nesting deeper than 2; primes; surviving-mark invisibility to the
|
|
1272
|
+
conformance matrix.
|
|
1273
|
+
9. **`ru`/`el`/`fr`/`fr-CA` already-curly protection is narrower than 0.1.0's** for these
|
|
1274
|
+
same-width-primary/secondary locales specifically — the two-stack split no longer
|
|
1275
|
+
independently protects already-curly text there the way it protects freshly-typed straight
|
|
1276
|
+
text. §6's dedicated adversarial sweep found no failures, but the protection that remains is
|
|
1277
|
+
the vetoes and the gate, not an independent structural guarantee.
|
|
1278
|
+
|
|
1279
|
+
---
|
|
1280
|
+
|
|
1281
|
+
## History
|
|
1282
|
+
|
|
1283
|
+
0.1.0 shipped with the four-pass, `STRAIGHT`-only design whose idempotency proof is quoted and
|
|
1284
|
+
withdrawn in §0/§5 above. 0.2.0 was `dashes`' version bump and did not touch this file. 0.3.0 is
|
|
1285
|
+
the revision described through most of this document: mandates 1 and 2, the five-pass
|
|
1286
|
+
algorithm, the certification gate, and the locale-schema constraints in §2.1.
|
|
1287
|
+
|
|
1288
|
+
0.4.0 adds the **listed elision veto** (§2, §3.2): `quotes.elisionIdioms`, a locale-data list of
|
|
1289
|
+
`{ left, elided, right }` triples that declines to pair two `NARROW` marks as a quotation when
|
|
1290
|
+
the full surrounding context matches a configured idiom, closing the `rock 'n' roll` defect
|
|
1291
|
+
(§7 item 8) for `en-US`. A first design considered during the same development — a bare
|
|
1292
|
+
`elisionForms` word list matching only the elided content — was reviewed and rejected before
|
|
1293
|
+
any commit or release once it was found to falsely elide ordinary quotations of a single letter
|
|
1294
|
+
(`The letter 'n' is common.`); it never shipped. §6's `N1`/`N2` rows and this history entry
|
|
1295
|
+
exist so that design is not silently reintroduced.
|
|
1296
|
+
|
|
1297
|
+
0.5.0 adds the **general ambiguous-medial-span veto** (§3.2), closing the same class of defect for
|
|
1298
|
+
every locale without a cited `elisionIdioms` entry — not by inferring an idiom, but by preserving
|
|
1299
|
+
the author's straight ASCII marks unconverted. The shared predicate this and `apostrophe`
|
|
1300
|
+
(order 50) both consult lives in one module, `src/rules/quote-ambiguity.ts` (JS reference
|
|
1301
|
+
implementation), specifically so `apostrophe`'s own structural case ladder cannot independently
|
|
1302
|
+
curl a mark this rule has deliberately left alone — see `apostrophe.md` §3.4 for why that
|
|
1303
|
+
composition risk was real, not hypothetical.
|
|
1304
|
+
|
|
1305
|
+
1.1.0 **replaces 0.5.0's veto with the universal medial-`n` elision veto** (§3.2), on the operator
|
|
1306
|
+
decision of 2026-09-09 that `rock 'n' roll` and `rock'n'roll` are international and take U+2019 in
|
|
1307
|
+
every locale. Three things change together, and none of them works without the other two:
|
|
1308
|
+
|
|
1309
|
+
- **The shape narrows** from 1–3 `LETTER` code points to exactly one code point, U+006E or U+004E.
|
|
1310
|
+
0.5.0's shape could not distinguish the idiom from an ordinary short nested quotation and
|
|
1311
|
+
declined both, which is what made `«это 'моё' дело»` and `“He said 'no' to me,”` come out with
|
|
1312
|
+
the inner marks unconverted. It was also normalization-dependent, since `LETTER` includes `Mn`.
|
|
1313
|
+
- **The outcome inverts** from preserve to convert: matched marks are no longer held back from
|
|
1314
|
+
`apostrophe`, which renders each by its ordinary case ladder. 0.5.0's preserve set
|
|
1315
|
+
(`computePreserveIndices`, `apostrophe.md` §3.4) is withdrawn along with the reason it existed.
|
|
1316
|
+
- **The predicate widens** from straight ASCII to the whole `NARROW` class. This is forced by the
|
|
1317
|
+
second change, not chosen: a veto that produces U+2019 must recognise U+2019, or its own output
|
|
1318
|
+
pairs as a quotation on the next pipeline pass. A straight-only predicate was implemented first
|
|
1319
|
+
and measured to do exactly that — `rock ’n’ roll` → `rock «n» roll` in `ru`, `rock ”n” roll` in
|
|
1320
|
+
`fi` — an idempotency violation and so a release blocker.
|
|
1321
|
+
|
|
1322
|
+
The cost accepted, knowingly and with §6's N1/N2 rows kept as its permanent witnesses, is that a
|
|
1323
|
+
genuine quotation of the letter *n* is elided. §7 item 8 records why that is the smaller error
|
|
1324
|
+
than the class 0.5.0 traded it for.
|