polytypo 1.2.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +626 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,1157 @@
|
|
|
1
|
+
# Rule: `nbsp`
|
|
2
|
+
|
|
3
|
+
**Order:** 70 (last). **Default:** on. **Modes:** text, html, markdown, yaml.
|
|
4
|
+
**Spec version:** 1.3.0 (0.1.0 for everything except §3.9's `initialBinding` change (0.6.0),
|
|
5
|
+
§3.3's span-boundary paragraphs (1.2.0) and §3.3's character-reference guard (1.3.0), noted
|
|
6
|
+
inline).
|
|
7
|
+
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
## 1. Purpose
|
|
11
|
+
|
|
12
|
+
`nbsp` is the rule that makes a line break fall where the language allows it. It converts
|
|
13
|
+
ordinary spaces to U+00A0 (no-break space) or U+202F (narrow no-break space), and inserts one
|
|
14
|
+
of those where a locale requires a space that the author did not type — before French
|
|
15
|
+
`? ! ; :`, inside French guillemets, between a number and its unit, between a Russian
|
|
16
|
+
preposition and the word it governs, between initials and a surname. It runs **last**, after
|
|
17
|
+
every other rule has settled the characters around each candidate space, which is what makes
|
|
18
|
+
its decisions simple: by the time it runs, `spaces` has normalised every ordinary space run
|
|
19
|
+
to length one, `dashes` has fixed the dash forms, and `quotes` has produced the final quote
|
|
20
|
+
glyphs whose inner spacing this rule then owns. It is also the rule with the largest surface
|
|
21
|
+
for oscillation, so almost all of its design effort goes into being a strict no-op on text
|
|
22
|
+
that already carries the right no-break spaces.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## 2. Locale data consumed
|
|
27
|
+
|
|
28
|
+
- `nbsp.beforePunctuation` — array of single characters that take U+00A0 before them
|
|
29
|
+
- `nbsp.narrowBeforePunctuation` — array of single characters that take U+202F before them
|
|
30
|
+
- `nbsp.afterShortWords` — array of lowercase words
|
|
31
|
+
- `nbsp.abbreviations` — array of literal multi-token abbreviations written with U+0020
|
|
32
|
+
- `nbsp.beforeUnits` — array of units that bind to a preceding number
|
|
33
|
+
- `nbsp.beforeNumber` — array of abbreviations that bind to a following **number**
|
|
34
|
+
- `nbsp.beforeWord` — array of abbreviations that bind to a following **word**
|
|
35
|
+
- `nbsp.afterSymbols` — array of symbols that bind to a following number
|
|
36
|
+
- `nbsp.initialBinding` — enum: `"none" | "chain" | "single"` (spec 0.6.0; replaces the boolean
|
|
37
|
+
`bindInitials`). See §3.9.
|
|
38
|
+
- `quotes.primary.open`, `quotes.primary.close`, `quotes.primary.innerSpace`
|
|
39
|
+
- `quotes.secondary.open`, `quotes.secondary.close`, `quotes.secondary.innerSpace`
|
|
40
|
+
|
|
41
|
+
**Constraint Q-P (`quotes.md` §2.1), normative on this rule's own data.** Every code point in
|
|
42
|
+
`nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must be a member of `quotes`'
|
|
43
|
+
`CLOSEISH` (`quotes.md` §3.1). This is what makes `quotes`' Lemma B cover N1/N2 as well as N8:
|
|
44
|
+
Lemma B's case 3 needs `m ∈ CLOSEISH` to know that a mark whose `closeRight` moves from `m` to a
|
|
45
|
+
`SPACELIKE` insertion is accepted either way. The shipped lists (`:`, `;`, `!`, `?` in
|
|
46
|
+
`fr`/`fr-CA`, empty elsewhere) satisfy it; `pipeline-idempotency.md` §7 item 4's schema gap is
|
|
47
|
+
closed by `locale.schema.json`'s `nbspClosePunctuation` check. A locale adding a new
|
|
48
|
+
`beforePunctuation`/`narrowBeforePunctuation` entry outside `CLOSEISH` reopens Lemma B and needs
|
|
49
|
+
a `quotes.md` §5 re-derivation, not just a data change.
|
|
50
|
+
|
|
51
|
+
### 2.1 A locale attests membership; the rule owns the mechanism
|
|
52
|
+
|
|
53
|
+
**Normative, and general — it governs every list-valued field in `locale.schema.json`, not only
|
|
54
|
+
this rule's.**
|
|
55
|
+
|
|
56
|
+
> A `sources` citation must support exactly what the locale file **can express**: which tokens
|
|
57
|
+
> are in which list, which enum value is chosen, which boolean is set. It is **not** required to
|
|
58
|
+
> support the behaviour the rule attaches to that membership.
|
|
59
|
+
|
|
60
|
+
So a locale populating `beforeUnits` needs a source establishing that **these tokens are units
|
|
61
|
+
and are written with a space after the number**. It does **not** need a source saying that space
|
|
62
|
+
is non-breaking, or that it is U+00A0 rather than U+202F, or that the rule converts an existing
|
|
63
|
+
space rather than inserting a missing one. Those are decisions this document has already made,
|
|
64
|
+
once, for every locale.
|
|
65
|
+
|
|
66
|
+
**The decisive argument is that evidentiary burden must track expressive power.** There is no
|
|
67
|
+
field in which a locale can say "bind these units" or "do not bind them" — the schema offers a
|
|
68
|
+
list and nothing else. A locale therefore cannot be *wrong* about the mechanism, cannot vary it,
|
|
69
|
+
and cannot be asked to cite it. Requiring a citation for a claim the file has no way to make is
|
|
70
|
+
not rigour; it is a category error, and its effect is to empty fields that are correctly
|
|
71
|
+
populated.
|
|
72
|
+
|
|
73
|
+
Four supporting points:
|
|
74
|
+
|
|
75
|
+
- **The schema already says so.** `beforeUnits` is defined as "Units and signs that bind to a
|
|
76
|
+
preceding number with U+00A0". The binding is in the *field's* definition, which is where a
|
|
77
|
+
uniform mechanism belongs.
|
|
78
|
+
- **Every other field works this way, and must.** `afterShortWords` enumerates words and N3
|
|
79
|
+
decides what happens to the following space; nobody asks Мильчин to name U+00A0 by code point.
|
|
80
|
+
`abbreviations` lists strings and N4 converts their internal spaces. `initialBinding` (spec
|
|
81
|
+
0.6.0; the enum that replaced the boolean `bindInitials`) is still mechanism, not a code-point
|
|
82
|
+
claim — no typographic authority names U+00A0 by code point for it, and this document still
|
|
83
|
+
owns exactly what C1/C2 do — but the enum is more citable than the boolean it replaced: Chicago
|
|
84
|
+
states "two or more initials" (`"chain"`, en-US/de-DE/de-CH/ru) and André states a single
|
|
85
|
+
abbreviated first name (`"single"`, fr/fr-CA) as two *distinguishable* claims a locale file can
|
|
86
|
+
now actually choose between, where the boolean could only say on or off. `"none"` remains the
|
|
87
|
+
uncitable negative it always was.
|
|
88
|
+
- **PLAN.md §6 and ARCHITECTURE.md §5.1 already draw this line.** Declarative facts that vary by
|
|
89
|
+
locale go in the file; algorithms that do not vary go in the rule. *Which tokens are units*
|
|
90
|
+
varies by language. *A measurement does not break between its number and its unit* does not —
|
|
91
|
+
it is the same universal principle that justifies the range binding in `dashes.md` §3.3.1, and
|
|
92
|
+
that one is argued from UAX #14 and Chicago/Мильчин centrally rather than per locale.
|
|
93
|
+
- **Authorities describe how text should look and behave when set, not how it is encoded.** This
|
|
94
|
+
spec has drawn that distinction twice already and in the same direction: for Greek
|
|
95
|
+
(`spaces.md` §3.5 — the guide instructs a typist, it does not ask a tool to delete) and for
|
|
96
|
+
vulgar fractions (`symbols.md` §7.6 — the sources describe how a fraction should *look*). A
|
|
97
|
+
source stating "a space separates the number from the unit", plus a rule stating "that space
|
|
98
|
+
must not break", is not the source being over-read.
|
|
99
|
+
|
|
100
|
+
**What this does not license.** The membership claim still needs a source, and the principle
|
|
101
|
+
makes that burden *sharper*, not looser:
|
|
102
|
+
|
|
103
|
+
- Listing a token the source does not attest is unsupported however the mechanism is owned.
|
|
104
|
+
- `fi`'s `abbreviations: ["fil. maist."]` is right on exactly this test: Kotus attests the
|
|
105
|
+
internal shape of **one** abbreviation, which is a membership claim about a single token that
|
|
106
|
+
happens to contain a space. Joining two tokens a source merely shows side by side would be a
|
|
107
|
+
membership claim the source does not make, and would stay unsupported under either reading.
|
|
108
|
+
- A field whose *shape* the source contradicts is still wrong. This principle relocates the
|
|
109
|
+
burden; it does not lower it.
|
|
110
|
+
|
|
111
|
+
**Consequence, recorded so it is not re-litigated per locale.** A `beforeUnits` list resting on a
|
|
112
|
+
source that states the space but not its non-breaking character is **correctly populated**.
|
|
113
|
+
`fi.beforeUnits` is restored on that footing, and `en-GB`'s and `sv`'s measurement units stand on
|
|
114
|
+
the same one. The reverse reading — requiring each locale to attest the mechanism — would leave
|
|
115
|
+
N5 effectively dead outside `en-US` and `ru`, which is a large behavioural consequence to arrive
|
|
116
|
+
at one file at a time; if it is ever taken, it must be taken deliberately and written here.
|
|
117
|
+
|
|
118
|
+
---
|
|
119
|
+
|
|
120
|
+
**Precondition.** `nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must be
|
|
121
|
+
disjoint. `locale.schema.json` does not enforce this (reported); an implementation that finds
|
|
122
|
+
a character in both must raise `POLYTYPO_MALFORMED_LOCALE_DATA` rather than pick one.
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## 3. Algorithm
|
|
127
|
+
|
|
128
|
+
Input is a code-point array `cp[0 … n-1]`.
|
|
129
|
+
|
|
130
|
+
### 3.1 Character classes
|
|
131
|
+
|
|
132
|
+
| Class | Members |
|
|
133
|
+
| --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
134
|
+
| `SP` | U+0020 (space) |
|
|
135
|
+
| `NBSP` | U+00A0 |
|
|
136
|
+
| `NNBSP` | U+202F |
|
|
137
|
+
| `NOBREAK` | `NBSP` ∪ `NNBSP` |
|
|
138
|
+
| `SPACELIKE` | `SP` ∪ `NOBREAK` ∪ U+0009 ∪ U+2007 ∪ U+2008 ∪ U+2009 ∪ U+200A ∪ U+2000–U+2006 ∪ U+205F ∪ U+3000 ∪ every member of `BREAK` |
|
|
139
|
+
| `OTHER-SPACE` | the fixed-width spaces above: U+2000–U+200A, U+205F, U+3000. In `SPACELIKE`, but **never** converted and never a valid "already correct" state |
|
|
140
|
+
| `BREAK` | U+000A, U+000D, U+000B, U+000C, U+0085, U+2028, U+2029 |
|
|
141
|
+
| `DIGIT` | U+0030–U+0039 |
|
|
142
|
+
| `LETTER` | general category `Lu`, `Ll`, `Lt`, `Lm`, `Lo`, `Mn`, `Mc`, `Me` |
|
|
143
|
+
| `UPPER` | general category `Lu` or `Lt` |
|
|
144
|
+
| `ALNUM` | `LETTER` ∪ `DIGIT` |
|
|
145
|
+
| `OPENISH` | U+0028 `(` U+005B `[` U+007B `{` and every locale `open` glyph |
|
|
146
|
+
| `SENTENCE-DASH` | U+2013, U+2014 only. **Not** U+002D and **not** U+2011 — see §3.5 |
|
|
147
|
+
| `DASHISH` | U+002D, U+2011, U+2013, U+2014. U+2011 is included because `hyphen` (order 35) produces it; a class holding only U+002D would make a word-boundary verdict depend on whether `hyphen` had run (`hyphen.md` §3.2) |
|
|
148
|
+
| `CLOSEISH` | U+0029 `)` U+005D `]` U+007D `}` and every locale `close` glyph |
|
|
149
|
+
| `NONE` | index out of range |
|
|
150
|
+
|
|
151
|
+
**Unicode version.** The general categories and case mappings this rule reads are those of the UCD version pinned in `spec/UNICODE` (`17.0`). The pin is normative for the **derived tables**, not for the host runtime — see [pipeline-idempotency.md](pipeline-idempotency.md) §6a, which also specifies the canary fixtures that make the pin detectable.
|
|
152
|
+
|
|
153
|
+
**The inline span boundary marker is in `CLOSEISH` and not in `OPENISH` (spec 1.2.0).** In `html`
|
|
154
|
+
and `markdown` mode the pipeline runs over spans joined by a −1 marker ([modes.md](modes.md) §3.2).
|
|
155
|
+
For this rule the marker is a member of `CLOSEISH` only; [modes.md](modes.md) §3.3 records the
|
|
156
|
+
split and §3.3 steps 2 and 3 below give the reason for each half. The −2 line marker is a member of
|
|
157
|
+
`BREAK`, as it is for every rule.
|
|
158
|
+
|
|
159
|
+
**`NOBREAK` is a member of `SPACELIKE`.** Every word/number/token boundary test in this rule
|
|
160
|
+
uses `SPACELIKE`, never `SP`. This single decision is what makes the rule idempotent: after a
|
|
161
|
+
conversion the boundary that justified it still reads as a boundary.
|
|
162
|
+
|
|
163
|
+
### 3.1a The narrow target, and the `narrowNbsp` option (spec 1.3.0)
|
|
164
|
+
|
|
165
|
+
**This rule is the only thing in the spec that emits U+202F**, through exactly two sub-rules: N2
|
|
166
|
+
(§3.4), always, and N8 (§3.10), when the pair's `innerSpace` is `"narrow-nbsp"`. **No shipped
|
|
167
|
+
locale sets `"narrow-nbsp"`** — `fr`, the only locale with a non-empty `narrowBeforePunctuation`,
|
|
168
|
+
sets its primary pair to `"nbsp"` — so today the option's whole visible effect is N2's. The N8
|
|
169
|
+
clause is written anyway, because the two must move together the day a locale does set it, and a
|
|
170
|
+
substitution that covered one emitter and not the other would be a defect nobody could see until
|
|
171
|
+
that locale landed. `quotes` places
|
|
172
|
+
the glyphs and leaves the inner space to N8 (quotes.md §5); no other rule writes a no-break space
|
|
173
|
+
of any width.
|
|
174
|
+
|
|
175
|
+
That makes one substitution expressible without touching anything else:
|
|
176
|
+
|
|
177
|
+
> **`NARROW-TARGET`** is U+202F, **unless** the caller passed `narrowNbsp: "nbsp"`, in which case
|
|
178
|
+
> it is U+00A0.
|
|
179
|
+
|
|
180
|
+
**N2 and N8 read `NARROW-TARGET` everywhere they name U+202F** — the state their "already
|
|
181
|
+
correct" branch recognises, the code point they convert a space to, and the code point they
|
|
182
|
+
insert. Nothing else in this document changes.
|
|
183
|
+
|
|
184
|
+
**Why a caller would ask.** U+202F is missing from many common text faces (Manrope and EB
|
|
185
|
+
Garamond among them), so a browser falls back to another face for that one character and the
|
|
186
|
+
advance width changes mid-line; a PDF renderer with no fallback drops the glyph entirely. That
|
|
187
|
+
is a **rendering** fact about the reader's fonts, not a typographic one about the language, which
|
|
188
|
+
is why it is a caller option and not locale data: `fr` still sets a narrow space before `?`, and
|
|
189
|
+
the locale file still says so.
|
|
190
|
+
|
|
191
|
+
**Why it rewrites the target rather than post-processing the output.** A caller can already write
|
|
192
|
+
`transform(...).replaceAll("\u202f", "\u00a0")`, and that is stable **only as long as it always
|
|
193
|
+
runs**: feed its output back through `transform` and N2 sees U+00A0 where its target is U+202F,
|
|
194
|
+
converts it back, and the two steps disagree for ever. Rewriting the target makes the result a
|
|
195
|
+
fixed point by construction — §5's argument is stated per sub-rule **target**, so substituting the
|
|
196
|
+
target carries it over verbatim, with no new case to check.
|
|
197
|
+
|
|
198
|
+
**Validation, and where it happens.** `narrowNbsp` is `"narrow"` or `"nbsp"`; absent means
|
|
199
|
+
`"narrow"`. Any other value raises `POLYTYPO_INVALID_OPTION` (`ARCHITECTURE.md` §4.6, new in spec
|
|
200
|
+
1.3.0 and general to every option added from 1.3.0 onward). It is checked **immediately after
|
|
201
|
+
`mode` and before `rules`**, so the full order is `mode` → `narrowNbsp` → `rules` → `locale` →
|
|
202
|
+
`dialect`: the two checks that read nothing but the call itself come first, then rule ids, then
|
|
203
|
+
locale data, then the dialect and the parse. The existing pairs are untouched — an unknown rule
|
|
204
|
+
still wins over an unknown locale — and this order is public, tested behaviour like the rest.
|
|
205
|
+
|
|
206
|
+
The check belongs to the **call**, not to this rule: it runs whether or not `nbsp` is enabled, so
|
|
207
|
+
`rules: { nbsp: false }` with a misspelled `narrowNbsp` still raises rather than silently
|
|
208
|
+
accepting a value that would have mattered.
|
|
209
|
+
|
|
210
|
+
**Conformance.** `spec/fixtures/*.json` carries the option as a case-level `narrowNbsp` field,
|
|
211
|
+
passed through to `transform` exactly as `rules` is; omitting it means `"narrow"`, so every case
|
|
212
|
+
written before spec 1.3.0 keeps its meaning. Note what a fixture **cannot** express: the schema's
|
|
213
|
+
own enum admits only the two valid values, so `POLYTYPO_INVALID_OPTION` has no fixture, the same
|
|
214
|
+
way a missing or misspelled `dialect` has none — a fixture with `mode: "markdown"` is required to
|
|
215
|
+
name a valid dialect. Option-validation errors are each runtime's own test, and the order in the
|
|
216
|
+
paragraph above is what those tests assert.
|
|
217
|
+
|
|
218
|
+
**What it does not do.**
|
|
219
|
+
|
|
220
|
+
- It does not change **which** positions take a no-break space. That is locale data (§2) and is
|
|
221
|
+
untouched: the same indices are claimed, by the same sub-rules, under the same guards.
|
|
222
|
+
- It does not change what the rule **reads**. U+202F stays in `NOBREAK` and `SPACELIKE` (§3.1),
|
|
223
|
+
so every guard still treats an authored narrow space as the no-break space it is. The one
|
|
224
|
+
visible consequence is that an authored U+202F at an index N2 or N8 claims is now converted to
|
|
225
|
+
U+00A0, because the rule normalises a claimed index to its target and the target has moved —
|
|
226
|
+
the same normalisation N2 already performs on an authored U+00A0 in the default configuration.
|
|
227
|
+
An authored U+202F anywhere else is left alone (§4).
|
|
228
|
+
- It does not touch **U+2011**, which `hyphen` (order 35) emits. Binding a hyphen is that rule's
|
|
229
|
+
entire job, so `rules: { hyphen: false }` already asks for exactly this and no new option is
|
|
230
|
+
warranted.
|
|
231
|
+
- It does not touch **U+2060**, which `ranges` (order 25) emits when it binds a tight range
|
|
232
|
+
(ranges.md §3.3.1). `ranges` is off by default, so a caller who has not enabled it never sees
|
|
233
|
+
one; a caller who has can turn it off again. See §7.8.
|
|
234
|
+
|
|
235
|
+
### 3.2 Structure: ten sub-rules, first claim wins
|
|
236
|
+
|
|
237
|
+
The rule is one left-to-right scan that evaluates **ten** independent sub-rules, N1 through N10. (It was eight before `beforeNumber` and `beforeWord` were added as N9 and N10; a reader who stops at N8 loses both, and with them the abbreviation binding for four locales.) Two sub-rules
|
|
238
|
+
may target the same index with different replacements (French `«` followed by a word that
|
|
239
|
+
starts with `«`… or, more realistically, a unit that is also a listed short word). The
|
|
240
|
+
resolution is fixed and total:
|
|
241
|
+
|
|
242
|
+
> Sub-rules are evaluated in the order **N1, N2, N3, N4, N5, N6, N7, N8, N9, N10** as written
|
|
243
|
+
> below.
|
|
244
|
+
> Each produces candidate edits keyed by the index of the space (or insertion point) it
|
|
245
|
+
> claims. **The first sub-rule to claim an index wins; every later edit at that index is
|
|
246
|
+
> discarded.** No index is ever edited twice.
|
|
247
|
+
|
|
248
|
+
This is stated because a map- or set-based implementation that resolves conflicts by
|
|
249
|
+
iteration order will behave differently in Go (ARCHITECTURE.md §4.5).
|
|
250
|
+
|
|
251
|
+
> **First-claim-wins is not sufficient on its own, and relying on it as though it were is a
|
|
252
|
+
> bug.** It resolves a conflict only when both sub-rules actually _emit an edit_ at the index.
|
|
253
|
+
> A sub-rule whose "already correct" branch fires emits nothing and therefore **claims
|
|
254
|
+
> nothing**, silently yielding the index to a lower-priority sub-rule — which then edits it,
|
|
255
|
+
> after which the higher-priority sub-rule is no longer satisfied, and the two alternate on
|
|
256
|
+
> successive runs. That is not a hypothetical: it is exactly how N2 and N8 oscillated on the
|
|
257
|
+
> French input `«?` (§3.10.1). **Sub-rules that target different code points must therefore be
|
|
258
|
+
> made disjoint by construction, not merely ordered.** Where two sub-rules could want
|
|
259
|
+
> different characters at one index, one of them must be given a guard that makes it decline
|
|
260
|
+
> the index outright, so that its verdict does not depend on what the other did.
|
|
261
|
+
|
|
262
|
+
**Why N9 and N10 are last, and in that order.** `nbsp.abbreviations` (N4) holds multi-token
|
|
263
|
+
forms such as `т. д.`; if a locale also lists `т.` in `beforeWord`, the space inside the
|
|
264
|
+
longer, more specific form must be claimed by N4, so N4 must precede N10. `beforeNumber`
|
|
265
|
+
precedes `beforeWord` so that an abbreviation listed in both — `S.` before `12` and before
|
|
266
|
+
`Petersburg` — resolves to the number reading first; the two are then disjoint by their own
|
|
267
|
+
guards (N9 requires a following digit, N10 a following letter) and the ordering is belt and
|
|
268
|
+
braces rather than a real tie-break. `initialBinding` (N7, spec 0.6.0) precedes both because an initial is a
|
|
269
|
+
narrower, structurally-recognised form than a listed abbreviation, and the two agree on the
|
|
270
|
+
replacement (U+00A0) wherever they overlap.
|
|
271
|
+
|
|
272
|
+
### 3.3 N1 — `beforePunctuation` (U+00A0)
|
|
273
|
+
|
|
274
|
+
For each index `i` such that `cp[i]` is a member of `nbsp.beforePunctuation`:
|
|
275
|
+
|
|
276
|
+
1. **Run guard.** If `cp[i-1]` is a member of `beforePunctuation` ∪ `narrowBeforePunctuation`,
|
|
277
|
+
skip. Only the first mark of `?!` or `!!!` takes the space.
|
|
278
|
+
2. **Right-context guard.** Let `after = cp[i+1]` or `NONE`. If `after` is **not** `NONE`,
|
|
279
|
+
not in `SPACELIKE`, not in `CLOSEISH`, **not U+2026**, and not a member of
|
|
280
|
+
`beforePunctuation` ∪ `narrowBeforePunctuation`, skip.
|
|
281
|
+
|
|
282
|
+
U+2026 is in the accepted set because the guard exists to detect a punctuation mark that is
|
|
283
|
+
_inside a token_ — `http://`, `12:30` — and an ellipsis after a question mark is nothing of
|
|
284
|
+
the kind. Without it, French `Vraiment?…` got no narrow space while `Vraiment ?` did, purely
|
|
285
|
+
because `ellipsis` (order 20) had converted the dots before `nbsp` looked. The space belongs
|
|
286
|
+
before the `?` whatever follows it. This is the guard that protects
|
|
287
|
+
`http://example.org` and `12:30` in a locale that lists U+003A — after the colon comes
|
|
288
|
+
`/` or a digit, so nothing happens.
|
|
289
|
+
|
|
290
|
+
**A span boundary marker after the mark is accepted (spec 1.2.0)**, because the marker is in
|
|
291
|
+
`CLOSEISH` (§3.1). This is what makes the `spaces` round trip (`spaces.md` §1) work when an
|
|
292
|
+
inline element closes or a `<br>` follows the mark. `spaces` deletes the U+0020 in
|
|
293
|
+
`<strong>Résistant au gel :</strong> il` and `Et la réglementation ?<br>Oui`, because in both
|
|
294
|
+
the run has content on each side and `right` is in `STRIP-BEFORE`. This sub-rule must then put
|
|
295
|
+
the no-break form back. Before 1.2.0 every runtime left the marker out of `CLOSEISH`, this
|
|
296
|
+
guard refused, and the author's space was lost for good: `gel:</strong>`, `réglementation?<br>`.
|
|
297
|
+
The insertion is at the mark's own index, inside the span, so the edge-growth rule
|
|
298
|
+
([modes.md](modes.md) §3.4) does not discard it. When the mark is alone in its span
|
|
299
|
+
(`mot<em>!</em>`) the insertion point is the span edge and is still discarded, as before.
|
|
300
|
+
|
|
301
|
+
The cost is stated here so it is not rediscovered as a bug. The guard cannot see past the
|
|
302
|
+
marker, so a colon that ends one span while a digit or a `/` starts the next is accepted as
|
|
303
|
+
sentence punctuation: `fr` `12:<b>30</b>` becomes `12⍽:<b>30</b>`. That needs a time, a URL
|
|
304
|
+
or a ratio split by markup exactly at the colon. The shape this repairs, `**Label :**` and
|
|
305
|
+
`<strong>Label :</strong>`, is the standard French definition-list and FAQ pattern, and the
|
|
306
|
+
failure it repairs deletes a character.
|
|
307
|
+
|
|
308
|
+
3. **Quote-glyph guard.** If `cp[i-1]` is in `SPACELIKE` and `cp[i-2]` is in `OPENISH`, or if
|
|
309
|
+
`cp[i-1]` is in `OPENISH`, **skip**. The space immediately after an opening quotation glyph
|
|
310
|
+
belongs to `quotes.innerSpace` and is owned by N8; two sub-rules must not both have an
|
|
311
|
+
opinion about it. Without this guard N2 and N8 alternate for ever on the French input `«?`
|
|
312
|
+
— see §3.10.1, which is the defect this guard repairs.
|
|
313
|
+
|
|
314
|
+
**A span boundary marker never triggers this guard**, because it is not in this rule's
|
|
315
|
+
`OPENISH` (§3.1). N8 matches the literal `P.open` glyph, so it never owns the space beside a
|
|
316
|
+
marker, and nothing is left for the guard to protect. Counting the marker as `OPENISH` would
|
|
317
|
+
make this guard decline ordinary French prose that follows an inline element or a link:
|
|
318
|
+
`Il dit <em>non</em> ! Oui` would keep a plain U+0020 before `!`, and the Markdown
|
|
319
|
+
`Voir [ceci](http://x.org) : oui` a plain U+0020 before `:`. The same membership would also
|
|
320
|
+
widen the left-boundary tests of N3, N7, N9 and N10, and N3's following-token guard, at span
|
|
321
|
+
edges. Spec 1.2.0 does not make that change (§7 item 12).
|
|
322
|
+
4. **Character-reference guard (spec 1.3.0).** If `cp[i]` is U+003B, and the code points to its
|
|
323
|
+
left have the shape of a character reference, **skip**. Concretely: let `j = i-1` and walk
|
|
324
|
+
left while `cp[j]` is an ASCII letter or an ASCII digit, giving a run of `len` code points;
|
|
325
|
+
if `len` is at least 1 and at most 32, and `cp[j]` is U+0023, decrement `j` once more; then
|
|
326
|
+
if `cp[j]` is U+0026, this `;` terminates a reference and the sub-rule emits nothing.
|
|
327
|
+
|
|
328
|
+
The reason is that in `text` mode there is no markup concept at all (`modes.md` §3.1), so a
|
|
329
|
+
French locale, whose `narrowBeforePunctuation` contains `;`, inserted U+202F before the `;`
|
|
330
|
+
that **ends** the reference: `Bonjour : oui` became `Bonjour ·: oui` and
|
|
331
|
+
`Tom & Jerry` became `Tom &· Jerry`. The input stopped being what it was — this is the
|
|
332
|
+
one case found where a rule corrupts the input's own syntax rather than merely typesetting
|
|
333
|
+
something it should not have. In `html` mode the same strings were already safe, because
|
|
334
|
+
§3.6 makes a well-formed reference opaque; the guard makes `text` mode stop destroying them
|
|
335
|
+
too, which is the mode people reach for when they have "just a string".
|
|
336
|
+
|
|
337
|
+
The test is the **shape** of a reference, not membership of the HTML named-reference table.
|
|
338
|
+
That table is thousands of entries and would have to be identical in five runtimes, for no
|
|
339
|
+
gain: the guard's job is to decline, and declining on `¬aname;` costs nothing. The bound
|
|
340
|
+
of 32 code points is the length of the longest named reference (31) plus one, and it keeps
|
|
341
|
+
the left walk bounded rather than open-ended.
|
|
342
|
+
|
|
343
|
+
Accepted cost, pinned: a French semicolon that directly follows a token containing `&` with
|
|
344
|
+
no space, `R&D;`, keeps a plain U+0020 or no space at all rather than gaining U+202F. Such a
|
|
345
|
+
token is not French prose, and the alternative is breaking every character reference in the
|
|
346
|
+
language's own documents.
|
|
347
|
+
5. Let `left = cp[i-1]` or `NONE`.
|
|
348
|
+
- If `left` is `NBSP` → **already correct**, emit nothing.
|
|
349
|
+
- If `left` is `SP` or `NNBSP` → emit an edit replacing `cp[i-1]` with U+00A0.
|
|
350
|
+
- If `left` is in `OTHER-SPACE` → **skip**. A thin or figure space was placed deliberately;
|
|
351
|
+
§4 promises it is left alone, and treating it as content would insert a _second_ space
|
|
352
|
+
beside it.
|
|
353
|
+
- If `left` is `NONE`, in `BREAK`, or U+0009 → skip. A line must not begin with a no-break
|
|
354
|
+
space.
|
|
355
|
+
- Otherwise (`left` is a content character) → emit an edit **inserting** U+00A0 at
|
|
356
|
+
index `i`.
|
|
357
|
+
|
|
358
|
+
### 3.4 N2 — `narrowBeforePunctuation` (`NARROW-TARGET`)
|
|
359
|
+
|
|
360
|
+
Identical to N1 with `NARROW-TARGET` (§3.1a — U+202F unless the caller substituted it) in place
|
|
361
|
+
of U+00A0: already-correct means `left` **is** `NARROW-TARGET`; a `SP`, or a `NOBREAK` member
|
|
362
|
+
that is not `NARROW-TARGET`, is converted to it; a content character causes an insertion of it.
|
|
363
|
+
In the default configuration that reads exactly as it always has — already-correct is `NNBSP`,
|
|
364
|
+
and `SP` or `NBSP` converts to U+202F. The
|
|
365
|
+
quote-glyph guard, the `OTHER-SPACE` guard and the character-reference guard apply unchanged —
|
|
366
|
+
the last one matters most here, since `;` is in `narrowBeforePunctuation` for `fr` and nowhere in
|
|
367
|
+
`beforePunctuation` for any shipped locale, so N2 is the sub-rule that was destroying references.
|
|
368
|
+
|
|
369
|
+
Because N1 runs first, a character listed in both arrays would be handled by N1 — which is
|
|
370
|
+
why §2 requires the arrays to be disjoint and requires an implementation to fail loudly
|
|
371
|
+
rather than rely on that precedence.
|
|
372
|
+
|
|
373
|
+
### 3.5 N3 — `afterShortWords` (U+00A0)
|
|
374
|
+
|
|
375
|
+
For each entry `w` in `nbsp.afterShortWords` (a lowercase word of `k` code points):
|
|
376
|
+
|
|
377
|
+
1. Find each index `a` where `cp[a … a+k-1]` matches `w` under the **first-character-lenient**
|
|
378
|
+
comparison: `cp[a+j] = w[j]` exactly for `j ≥ 1`, and for `j = 0` either `cp[a] = w[0]` or
|
|
379
|
+
`cp[a]` is the Unicode **simple uppercase mapping** of `w[0]`. Simple uppercase mapping is
|
|
380
|
+
taken from the Unicode Character Database, is a plain code-point→code-point table, and is
|
|
381
|
+
**not** a locale-sensitive operation — this satisfies ARCHITECTURE.md §4.4 (`i` → `I`
|
|
382
|
+
always, never `İ`). No other case variation matches: `IN` does not match `in`.
|
|
383
|
+
2. **Left boundary.** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
|
|
384
|
+
**`SENTENCE-DASH`**. Otherwise skip.
|
|
385
|
+
|
|
386
|
+
**A hyphen — U+002D or U+2011 — fails this test, and that is the point.** A hyphen marks an
|
|
387
|
+
_intra-word_ position by construction, so the token after one is not a free-standing word.
|
|
388
|
+
Without this, `из-за дождя` binds twice over: `hyphen` (order 35) produces `из‑за`, and then
|
|
389
|
+
this sub-rule matches the listed short word `за` at the position after the U+2011 and emits
|
|
390
|
+
`из‑за` + U+00A0 + `дождя`. But `за` there is not a preposition — it is the tail of a single
|
|
391
|
+
compound preposition — so the binding is typographically wrong on ordinary Russian prose.
|
|
392
|
+
The same reasoning covers every form `hyphen` produces: compound tails (`под` in `из-под`),
|
|
393
|
+
prefix bodies (`что` in `кое-что`) and suffix bodies (`то`, `таки`) are all word-internal and
|
|
394
|
+
none of them should ever be treated as a word start here.
|
|
395
|
+
`SENTENCE-DASH` keeps the case the boundary was written for: an em or en dash genuinely does
|
|
396
|
+
introduce a new phrase, so `— в Москве` still binds `в`.
|
|
397
|
+
|
|
398
|
+
3. **Right boundary.** `cp[a+k]` must exist and be `SP` or `NBSP`. If it is `NBSP`, emit
|
|
399
|
+
nothing (already correct). If it is anything else — a letter, a punctuation mark, a line
|
|
400
|
+
break, `NNBSP` — skip. (A `NNBSP` there was put by another sub-rule with better
|
|
401
|
+
information.)
|
|
402
|
+
4. **Following-token guard.** `cp[a+k+1]` must exist and be in `ALNUM` or `OPENISH`.
|
|
403
|
+
A short word must not be bound to a punctuation mark or to the end of a line.
|
|
404
|
+
5. Emit an edit replacing `cp[a+k]` with U+00A0.
|
|
405
|
+
|
|
406
|
+
**Longest match wins, and there is no backtracking.** When two entries match at the same index
|
|
407
|
+
`a`, only the longest is considered; entries are unique (`uniqueItems` in the schema), so this
|
|
408
|
+
is a total order. **If the longest matching entry fails a guard, no shorter entry is retried at
|
|
409
|
+
that index** — the sub-rule emits nothing there and the scan moves on. This is normative for
|
|
410
|
+
every list-driven sub-rule in this document (N3, N4, N5, N6, N9, N10) and for `hyphen`
|
|
411
|
+
(`hyphen.md` §3.4). Two reasons: a shorter entry succeeding where a longer one failed is almost
|
|
412
|
+
always wrong (if `mm` was rejected because the following code point is a letter, `m` is wrong
|
|
413
|
+
for the same reason), and backtracking makes the result depend on the interaction between list
|
|
414
|
+
contents and guard order in a way that is tedious to reproduce identically in five runtimes.
|
|
415
|
+
|
|
416
|
+
### 3.6 N4 — `abbreviations` (U+00A0 for every internal space)
|
|
417
|
+
|
|
418
|
+
For each entry `s` in `nbsp.abbreviations`, of `k` code points:
|
|
419
|
+
|
|
420
|
+
1. Find each index `a` where `cp[a … a+k-1]` matches `s` under **space-lenient exact**
|
|
421
|
+
comparison: for every position `j`, either `cp[a+j] = s[j]`, or `s[j]` is `SP` and
|
|
422
|
+
`cp[a+j]` is in `NOBREAK`. Every non-space code point must match exactly — no case
|
|
423
|
+
leniency at all here, because `z. B.` and `Z. B.` are different strings and the locale
|
|
424
|
+
file is expected to list what it means.
|
|
425
|
+
2. **Boundaries.** `cp[a-1]` must not be in `ALNUM`; `cp[a+k]` must not be in `ALNUM`.
|
|
426
|
+
(`z. B.` inside `Xz. B.y` is not a match.)
|
|
427
|
+
3. For every position `j` with `s[j]` = `SP`: if `cp[a+j]` is already `NBSP`, emit nothing for
|
|
428
|
+
that position; otherwise emit an edit replacing `cp[a+j]` with U+00A0.
|
|
429
|
+
|
|
430
|
+
Longest match wins at a given `a`, with no backtracking (§3.5).
|
|
431
|
+
|
|
432
|
+
### 3.7 N5 — `beforeUnits` (U+00A0, conversion only)
|
|
433
|
+
|
|
434
|
+
For each entry `u` in `nbsp.beforeUnits`, of `k` code points:
|
|
435
|
+
|
|
436
|
+
1. Find each index `a` where `cp[a … a+k-1]` matches `u` exactly (code-point equality; no
|
|
437
|
+
case leniency — `Kg` is not `kg`).
|
|
438
|
+
2. **Right boundary.** `cp[a+k]` must not be in `ALNUM`. This is what stops the unit `kg`
|
|
439
|
+
from matching inside `kgf`, and `m` from matching inside `metres`.
|
|
440
|
+
3. **Left side.** `cp[a-1]` must be `SP` or `NBSP`; nothing else qualifies. If it is `NBSP`,
|
|
441
|
+
emit nothing. **If there is no space at all, this sub-rule does nothing** — see §7.2, it
|
|
442
|
+
never inserts.
|
|
443
|
+
4. `cp[a-2]` must be in `DIGIT`. Let `Lrun` be the maximal `DIGIT` run ending at `a-2`,
|
|
444
|
+
starting at index `b`. `cp[b-1]` must not be in `LETTER` (so `H2 O` does not bind).
|
|
445
|
+
5. Emit an edit replacing `cp[a-1]` with U+00A0.
|
|
446
|
+
|
|
447
|
+
Longest match wins at a given `a`, with no backtracking (§3.5), so a locale listing both `m`
|
|
448
|
+
and `mm` binds `5 mm` correctly and does not fall back to `m` when the `mm` match is rejected.
|
|
449
|
+
|
|
450
|
+
### 3.8 N6 — `afterSymbols` (U+00A0, conversion only)
|
|
451
|
+
|
|
452
|
+
For each entry `y` in `nbsp.afterSymbols`, of `k` code points:
|
|
453
|
+
|
|
454
|
+
1. Find each index `a` where `cp[a … a+k-1]` matches `y` exactly.
|
|
455
|
+
2. **Left boundary.** `cp[a-1]` must not be in `ALNUM`.
|
|
456
|
+
3. **Right side.** `cp[a+k]` must be `SP` or `NBSP`. If `NBSP`, emit nothing. If there is no
|
|
457
|
+
space, do nothing (never insert).
|
|
458
|
+
4. `cp[a+k+1]` must be in `DIGIT`.
|
|
459
|
+
5. Emit an edit replacing `cp[a+k]` with U+00A0.
|
|
460
|
+
|
|
461
|
+
### 3.9 N7 — `initialBinding` (U+00A0), only when `nbsp.initialBinding` is not `"none"` (spec 0.6.0)
|
|
462
|
+
|
|
463
|
+
Define an **initial** at index `p` as: `cp[p]` in `UPPER`, `cp[p+1]` = U+002E (.), and
|
|
464
|
+
`cp[p-1]` either `NONE`, in `SPACELIKE`, or in `OPENISH`. (One letter, one full stop, at a
|
|
465
|
+
token start.)
|
|
466
|
+
|
|
467
|
+
Two clauses, both operating on a single space at index `q`:
|
|
468
|
+
|
|
469
|
+
- **C1 — initial on the left.** If `cp[q]` is `SP` or `NBSP`, and there is an initial ending
|
|
470
|
+
at `q-1` (i.e. `cp[q-1]` = U+002E and `cp[q-2]` is an initial's letter satisfying the
|
|
471
|
+
definition above), and `cp[q+1]` is in `UPPER`, **and the mode condition below holds**, then
|
|
472
|
+
bind: if `cp[q]` is `NBSP` emit nothing, else emit an edit replacing `cp[q]` with U+00A0.
|
|
473
|
+
|
|
474
|
+
**Mode condition (spec 0.6.0).** Split `cp[q+1]`'s role into two shapes:
|
|
475
|
+
- **Between-initials** — `cp[q+1]` is *itself* an initial (i.e. `q+1` also satisfies the
|
|
476
|
+
`initial` definition above: `cp[q+1]` in `UPPER`, `cp[q+2]` = U+002E). This shape always
|
|
477
|
+
binds, in every mode except `"none"` — it is unambiguous by construction: two adjacent
|
|
478
|
+
initial-shaped tokens are Chicago's own literal case ("between two or more initials").
|
|
479
|
+
- **Initial-to-word** — `cp[q+1]` is `UPPER` but not itself an initial (no U+002E follows at
|
|
480
|
+
`q+2`), i.e. the space leads into a candidate surname or ordinary word. This shape's
|
|
481
|
+
eligibility depends on `nbsp.initialBinding`:
|
|
482
|
+
- `"single"` — binds unconditionally, the historical shape (this is what `fr`/`fr-CA` need
|
|
483
|
+
for `N. Bourbaki`/`M. Dupont` — André §5.1.3, cited below).
|
|
484
|
+
- `"chain"` — binds **only if** the initial ending at `q-1` (letter at `p = q-2`) is itself
|
|
485
|
+
immediately preceded by another initial: `cp[p-1]` is `SP` or `NBSP`, `cp[p-2]` = U+002E,
|
|
486
|
+
and `p-3` satisfies the `initial` definition. In other words: bind the space leading out
|
|
487
|
+
of a chain into a following word only once the chain already has two or more members —
|
|
488
|
+
Chicago's own "two or more initials" wording, applied to this shape specifically rather
|
|
489
|
+
than to N7 as a whole. `E. B. White`: the space before `White` binds, because `B.` is
|
|
490
|
+
itself preceded by `E.`. A lone `A. Smith`, or `...take the top N. It runs...`, does not:
|
|
491
|
+
neither `A.` nor `N.` is preceded by another initial, so this shape declines. The declined
|
|
492
|
+
case is a **deliberate false negative**, not a bug: this document has no way to tell a
|
|
493
|
+
genuine single-initial name from an ordinary sentence boundary using only local structure
|
|
494
|
+
(§7.9a), and `"chain"` resolves that ambiguity by declining rather than guessing, on the
|
|
495
|
+
same footing P4 in `dashes.md` §3.4 declines rather than guessing at a range.
|
|
496
|
+
- `"none"` — N7 does not run at all (checked once, before either clause).
|
|
497
|
+
|
|
498
|
+
This covers both spaces of `А. С. Пушкин` under `"chain"` (`ru`): the first is
|
|
499
|
+
between-initials (`cp[q+1]` is `С`, itself an initial), the second is initial-to-word but
|
|
500
|
+
eligible because `С.` is preceded by `А.`.
|
|
501
|
+
|
|
502
|
+
**Guard C1-a — lower-case abbreviation tail.** C1 does **not** fire (in any mode, on either
|
|
503
|
+
shape) if, letting `p = q-2` be the initial's letter, `cp[p-1]` is in `SPACELIKE`, `cp[p-2]`
|
|
504
|
+
is U+002E, and `cp[p-3]` is in `LETTER` but **not** in `UPPER`. In other words: an uppercase
|
|
505
|
+
letter plus a dot that is itself preceded by a _lower-case_ letter plus a dot is the second
|
|
506
|
+
token of an abbreviation, not an initial.
|
|
507
|
+
|
|
508
|
+
This is not a refinement, it is a false-positive repair on real German data. With the
|
|
509
|
+
shipped `de-DE` file, `z. B. Berlin` produced `z.⍽B.⍽Berlin`: N4 correctly bound the space
|
|
510
|
+
inside `z. B.`, and then C1 read `B.` as an initial and `Berlin` as a surname and bound the
|
|
511
|
+
space after it too. A false positive on ordinary prose is the ship-blocking class for this
|
|
512
|
+
project (PLAN.md §4, the M4 gate), so this needed a guard rather than a note. §6 example 14
|
|
513
|
+
was right and the rule was wrong.
|
|
514
|
+
`А. С. Пушкин` is unaffected: `cp[p-3]` there is `А`, which is in `UPPER`.
|
|
515
|
+
|
|
516
|
+
- **C2 — two initials on the right.** If `cp[q]` is `SP` or `NBSP`, `cp[q-1]` is in `LETTER`,
|
|
517
|
+
and the code points starting at `q+1` form **two consecutive initials** (`X. Y.`, i.e. an
|
|
518
|
+
initial at `q+1`, a `SP` or `NBSP` at `q+3`, and an initial at `q+4`), then bind `cp[q]` the
|
|
519
|
+
same way, **in every mode except `"none"`** — C2 already requires two initials by
|
|
520
|
+
construction (Chicago's own "two or more"), so it needs no mode split.
|
|
521
|
+
This covers `Пушкин А. С.` — the surname-first order — without binding every
|
|
522
|
+
`word` + `Capital letter + period` pair, which would fire on ordinary sentences.
|
|
523
|
+
|
|
524
|
+
C2 binds only the space before the initials group; the space _inside_ the group is bound by
|
|
525
|
+
C1 (its `cp[q+1]` is the second initial's letter, which is `UPPER`).
|
|
526
|
+
|
|
527
|
+
### 3.10 N8 — `quotes.innerSpace`
|
|
528
|
+
|
|
529
|
+
For each of the two quote pairs `P` ∈ {`quotes.primary`, `quotes.secondary`} with
|
|
530
|
+
`P.innerSpace ≠ "none"`, let `target` be U+00A0 if `P.innerSpace = "nbsp"` and `NARROW-TARGET`
|
|
531
|
+
(§3.1a — U+202F unless the caller substituted it) if `P.innerSpace = "narrow-nbsp"`.
|
|
532
|
+
|
|
533
|
+
**Sidedness precondition.** If `P.open` and `P.close` are the same code point, this sub-rule
|
|
534
|
+
does nothing for that pair, and an implementation must not guess. See §7.5.
|
|
535
|
+
|
|
536
|
+
Otherwise:
|
|
537
|
+
|
|
538
|
+
- **After an opening glyph.** For each index `o` with `cp[o]` = `P.open`:
|
|
539
|
+
- if `cp[o+1]` is `target` → emit nothing;
|
|
540
|
+
- else if `cp[o+1]` is `SP` or in `NOBREAK` → emit an edit replacing `cp[o+1]` with
|
|
541
|
+
`target`;
|
|
542
|
+
- else if `cp[o+1]` is `NONE` or in `BREAK` → skip;
|
|
543
|
+
- else → emit an edit inserting `target` at index `o+1`.
|
|
544
|
+
- **Before a closing glyph.** For each index `c` with `cp[c]` = `P.close`, symmetrically on
|
|
545
|
+
`cp[c-1]`, skipping when `cp[c-1]` is `NONE` or in `BREAK`.
|
|
546
|
+
|
|
547
|
+
**When `innerSpace` is `"none"` this sub-rule does nothing at all, and in particular it does
|
|
548
|
+
not _remove_ a space the author put inside the quotation marks.** This is normative, not an
|
|
549
|
+
oversight: `de-CH` declares `innerSpace: "none"` and `« hallo »` therefore survives as written.
|
|
550
|
+
|
|
551
|
+
The reasoning is that `innerSpace` describes what the pipeline **inserts**, not what it
|
|
552
|
+
enforces. Removing a space is a deletion of author content, and this spec deletes a U+0020 in
|
|
553
|
+
exactly three narrowly-argued places, all of them in `spaces` (§3.2 step 5) and none of them
|
|
554
|
+
next to a quotation mark — `STRIP-BEFORE` excludes quotation glyphs precisely so that their
|
|
555
|
+
inner spacing stays a single rule's business. A tool that silently closed up `« hallo »` in a
|
|
556
|
+
Swiss text would be making an editorial judgement about the author's copy rather than a
|
|
557
|
+
typographic correction to it, and the M4 criterion is zero false positives, not maximum
|
|
558
|
+
conformity. The counter-argument — that `«hallo»` is the only Swiss-conforming form, so leaving
|
|
559
|
+
the spaces leaves the document wrong — is real, and §7.4 records it as the operator's to settle.
|
|
560
|
+
|
|
561
|
+
#### 3.10.1 The N1/N2 oscillation this rule was part of
|
|
562
|
+
|
|
563
|
+
`fr` declares `quotes.primary.innerSpace = "nbsp"` (target U+00A0) and lists `?` in
|
|
564
|
+
`nbsp.narrowBeforePunctuation` (target U+202F). On the input `«?` both N2 and N8 want the
|
|
565
|
+
single index between the two characters, and they want different code points:
|
|
566
|
+
|
|
567
|
+
```
|
|
568
|
+
«? → «⍽? (N8 inserts U+00A0) → «⍹? (N2 converts to U+202F)
|
|
569
|
+
→ «⍽? (N8 converts back) → …
|
|
570
|
+
```
|
|
571
|
+
|
|
572
|
+
First-claim-wins did not resolve it, for the reason set out in §3.2: on the second run N2's
|
|
573
|
+
"already correct" branch fires, emits nothing, and therefore claims nothing — leaving N8 free
|
|
574
|
+
to act on an index N2 still has an opinion about.
|
|
575
|
+
|
|
576
|
+
**Repaired in N1/N2**, not here: the space beside an opening quotation glyph is `innerSpace`,
|
|
577
|
+
full stop, and N1/N2 now decline it outright (§3.3 step 4). N8 is unchanged. The alternative —
|
|
578
|
+
having N8 accept any `NOBREAK` beside a quote glyph as already correct — was rejected because
|
|
579
|
+
it would also make N8 unable to correct a genuinely wrong inner space anywhere else, trading a
|
|
580
|
+
narrow conflict for a broad loss of function.
|
|
581
|
+
|
|
582
|
+
Note the general shape, because it will recur: **the sub-rule that yields must be the one that
|
|
583
|
+
can state its exception locally.** N1/N2 can say "not next to an opening quote glyph" using
|
|
584
|
+
`OPENISH`, which they already have. N8 cannot say "not where a punctuation rule wants a
|
|
585
|
+
different width" without importing both punctuation lists.
|
|
586
|
+
|
|
587
|
+
### 3.11 N9 — `beforeNumber` (U+00A0, conversion only)
|
|
588
|
+
|
|
589
|
+
An abbreviation in `nbsp.beforeNumber` binds forward to a number: `Nr. 5`, `S. 12`,
|
|
590
|
+
`art. 237`, `§§ 12`. Schema minimum length is 2 code points, so a bare initial cannot be
|
|
591
|
+
listed.
|
|
592
|
+
|
|
593
|
+
For each entry `s` in `nbsp.beforeNumber`, of `k` code points:
|
|
594
|
+
|
|
595
|
+
1. Find each index `a` where `cp[a … a+k-1]` matches `s` **exactly**, code point for code
|
|
596
|
+
point. No case leniency at all: `Nr.` is not `nr.`, and a locale that wants both lists
|
|
597
|
+
both. (This differs from N3, where the schema explicitly asks for first-character
|
|
598
|
+
leniency; here it does not, and inventing it would make `S.` match `s.` and bind sentence
|
|
599
|
+
fragments.)
|
|
600
|
+
2. **Left boundary.** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
|
|
601
|
+
`SENTENCE-DASH`. Otherwise skip. This is stronger than "not `ALNUM`" and it is what stops
|
|
602
|
+
`S.` matching inside `Fig.S. 3`. A hyphen fails it, for the reason given in §3.5 step 2 —
|
|
603
|
+
an abbreviation cannot begin immediately after an intra-word hyphen.
|
|
604
|
+
3. **The separator.** `cp[a+k]` must be `SP` or `NBSP`; anything else — a letter, a digit, a
|
|
605
|
+
line terminator, `NNBSP`, a tab — and the sub-rule does nothing. If it is `NBSP`, emit
|
|
606
|
+
nothing: already correct.
|
|
607
|
+
**Exactly one separator.** If `cp[a+k+1]` is also in `SPACELIKE`, skip.
|
|
608
|
+
4. **"A following number" means:** `cp[a+k+1]` is in `DIGIT`. That is the whole definition —
|
|
609
|
+
one code point, tested for membership in U+0030–U+0039. It deliberately does **not**
|
|
610
|
+
require the digit run to be of any particular length, to be followed by anything in
|
|
611
|
+
particular, or to be a "plausible" number: a digit immediately after the separator is
|
|
612
|
+
sufficient evidence, and no false positive of consequence exists for it (`Nr. 5x` is still
|
|
613
|
+
a number after `Nr.`).
|
|
614
|
+
5. Emit an edit replacing `cp[a+k]` with U+00A0.
|
|
615
|
+
|
|
616
|
+
**Never inserts.** `Nr.5` stays `Nr.5`, for the same reason N5 never inserts: adding a space
|
|
617
|
+
is a content change, not a typographic one.
|
|
618
|
+
|
|
619
|
+
**N9 has no equivalent of N10's guard G-D**, so a one-letter abbreviation such as `S.` binds
|
|
620
|
+
forward here even though it is structurally indistinguishable from an initial. That is
|
|
621
|
+
deliberate and safe: N9 fires only when a `DIGIT` follows, and N7 clause C1 fires only when an
|
|
622
|
+
`UPPER` code point follows, so the two are disjoint by their own right-hand tests and cannot
|
|
623
|
+
both claim a space. `S. 12` is a page reference and only N9 can see it; `S. Petrov` is an
|
|
624
|
+
initial and only N7 can. No guard is needed to keep them apart.
|
|
625
|
+
|
|
626
|
+
Longest match wins at a given `a`, with no backtracking (§3.5).
|
|
627
|
+
|
|
628
|
+
### 3.12 N10 — `beforeWord` (U+00A0, conversion only)
|
|
629
|
+
|
|
630
|
+
An abbreviation in `nbsp.beforeWord` binds forward to a word: `г. Москва`, `ул. Ленина`,
|
|
631
|
+
`M. Dupont`, `Mme Hugo`, `Dr. Schmidt`. **This is the highest-false-positive sub-rule in the
|
|
632
|
+
spec** and its guards are correspondingly heavier — an abbreviation that also occurs at the
|
|
633
|
+
end of a sentence will bind across the sentence boundary, and no purely syntactic test can
|
|
634
|
+
separate the two readings. Locale files must list only forms whose binding is normative
|
|
635
|
+
(the schema description says so); this document adds the structural guards.
|
|
636
|
+
|
|
637
|
+
For each entry `s` in `nbsp.beforeWord`, of `k` code points:
|
|
638
|
+
|
|
639
|
+
1. Match `cp[a … a+k-1]` against `s` **exactly**, code point for code point. No case leniency
|
|
640
|
+
in either direction. `г.` matches only lowercase `г.`; `M.` matches only uppercase `M.`;
|
|
641
|
+
`Mme` matches only that capitalisation. A locale wanting both cases lists both, which is
|
|
642
|
+
cheap and explicit and keeps the spec free of any case-folding table on this path.
|
|
643
|
+
2. **Left boundary (G-L).** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
|
|
644
|
+
`SENTENCE-DASH`. Otherwise skip — a hyphen fails it, per §3.5 step 2.
|
|
645
|
+
3. **The separator (G-S).** `cp[a+k]` must be `SP` or `NBSP`. `NBSP` → emit nothing, already
|
|
646
|
+
correct. Anything else → skip. If `cp[a+k+1]` is also in `SPACELIKE`, skip.
|
|
647
|
+
4. **"A following word" means (G-W):** `cp[a+k+1]` is in `LETTER`. Not `ALNUM` — a digit there
|
|
648
|
+
is N9's business, and N9 has already run, so a `beforeWord` entry that is also a
|
|
649
|
+
`beforeNumber` entry never double-claims. Not `OPENISH`, not a quotation glyph, not a
|
|
650
|
+
dash: `г. «Москва»` is left alone, because the binding target is then a quotation and the
|
|
651
|
+
line-break risk this sub-rule exists to remove is not present in the same way.
|
|
652
|
+
5. **Initial-collision guard (G-D).** If `nbsp.initialBinding` is not `"none"`, **and** `s` is
|
|
653
|
+
exactly two code points, **and** the first is in `UPPER`, **and** the second is U+002E, skip.
|
|
654
|
+
An _uppercase_ letter plus a dot is structurally indistinguishable from an initial, and in a
|
|
655
|
+
locale where N7 is active, N7 owns that shape with better evidence (it inspects what follows
|
|
656
|
+
for a second initial or a surname). G-D's own condition does not distinguish `"chain"` from
|
|
657
|
+
`"single"` — it only asks whether N7 is active at all — because G-D's job is routing (which
|
|
658
|
+
sub-rule owns this shape), not deciding whether the space actually ends up bound; that
|
|
659
|
+
decision is C1's alone (§3.9), and under `"chain"` a routed-to-N7 space can still end up
|
|
660
|
+
declined if no chain is confirmed.
|
|
661
|
+
|
|
662
|
+
All three conditions matter:
|
|
663
|
+
- **`UPPER`** — without it, a lower-case two-code-point entry such as `ул.` would be inert
|
|
664
|
+
in an `initialBinding`-active locale, and `ул. Ленина` would not bind. A lower-case letter
|
|
665
|
+
plus a dot is not a plausible initial in any orthography this spec covers, so the guard
|
|
666
|
+
must not capture one. (An earlier revision justified this clause with `г. Москва`, which is
|
|
667
|
+
**not** an example of it: `ru.json` does not list `г.` in `beforeWord` at all — see §6 row
|
|
668
|
+
13a.)
|
|
669
|
+
- **`initialBinding !== "none"`** — in a locale where N7 is switched off, nothing else claims
|
|
670
|
+
the shape and there is no collision to avoid.
|
|
671
|
+
- **exactly two code points** — `Mme`, `ул.`, `art.` are longer and are never initials.
|
|
672
|
+
|
|
673
|
+
Consequence for French, stated plainly because an earlier revision of this document got the
|
|
674
|
+
underlying fact wrong: **`fr` has `initialBinding: "single"`** (spec 0.6.0; cited to Jacques
|
|
675
|
+
André §5.1.3 for `N. Bourbaki`). So a listed `M.` _is_ inert in N10, and `M. Dupont` binds
|
|
676
|
+
through N7 clause C1 instead — which reaches the same U+00A0 by a different route, eligible
|
|
677
|
+
under `"single"`'s unconditional initial-to-word shape (§3.9). The residual gap is that C1
|
|
678
|
+
requires an `UPPER` code point after the space, so `M. dupont` with a lower-case surname does
|
|
679
|
+
not bind at all. That is an acceptable miss; French surnames are capitalised.
|
|
680
|
+
|
|
681
|
+
6. **Line-boundary guard (G-B).** If `cp[a+k]` is in `BREAK`, or if `cp[a+k+1]` is `NONE`,
|
|
682
|
+
skip. Covered by G-S/G-W but stated separately because it is the guard that keeps a
|
|
683
|
+
no-break space off the end of a text unit.
|
|
684
|
+
7. Emit an edit replacing `cp[a+k]` with U+00A0.
|
|
685
|
+
|
|
686
|
+
**Never inserts.** Longest match wins at a given `a`, with no backtracking (§3.5).
|
|
687
|
+
|
|
688
|
+
**What is deliberately _not_ guarded.** There is no sentence-boundary test. In
|
|
689
|
+
`Это было в 1990 г. Москва тогда была другой`, the form `г.` ends the sentence and `Москва`
|
|
690
|
+
begins the next, yet every guard above passes and the two are bound. Distinguishing that from
|
|
691
|
+
`г. Москва` ("the city of Moscow") requires knowing whether `г.` means _год_ or _город_,
|
|
692
|
+
which is a lexical question the spec cannot answer from code points. The mitigation is
|
|
693
|
+
entirely in the locale data — list a form in `beforeWord` only when the wrong reading is rare
|
|
694
|
+
or harmless — and the failure mode is mild: a no-break space where a break was permitted, not
|
|
695
|
+
a changed character. This is recorded as §7.9 and must be reviewed at the M4 gate.
|
|
696
|
+
|
|
697
|
+
### 3.13 Insertions and indices
|
|
698
|
+
|
|
699
|
+
Sub-rules N1, N2 and N8 can _insert_ a code point. Because rules produce edits and the
|
|
700
|
+
pipeline applies them (ARCHITECTURE.md §7.1), every index above refers to the **input**
|
|
701
|
+
array. An insertion is an edit with an empty replaced span at a given index. Two insertions
|
|
702
|
+
at the same index cannot occur: the first-claim-wins rule of §3.2 forbids it.
|
|
703
|
+
|
|
704
|
+
---
|
|
705
|
+
|
|
706
|
+
## 4. Must not touch
|
|
707
|
+
|
|
708
|
+
**Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
|
|
709
|
+
a guarantee of `transform` as a whole — or **[R]** — true of this rule in isolation but capable
|
|
710
|
+
of being falsified by another rule, which is then named.
|
|
711
|
+
|
|
712
|
+
- **[P] A line terminator.** No sub-rule inserts, deletes or crosses one; every guard that could
|
|
713
|
+
reach a `BREAK` skips instead.
|
|
714
|
+
- **[P] The start or end of a text unit.** A no-break space is never inserted at index 0, never
|
|
715
|
+
after the last code point, and never adjacent to a `BREAK`. In `html` mode a text node
|
|
716
|
+
frequently begins or ends at a tag boundary, and a leading U+00A0 there is a visible
|
|
717
|
+
rendering change.
|
|
718
|
+
- **[P] A tab.** U+0009 is `SPACELIKE` for boundary purposes but is never converted.
|
|
719
|
+
- **[R] A space that another sub-rule already claimed** (§3.2).
|
|
720
|
+
- **[P] `http://`, `12:30`, `1:2`** in a locale that lists U+003A in `beforePunctuation`. N1
|
|
721
|
+
guard 2.
|
|
722
|
+
- **[P] The second and later marks of `?!`, `!!!`, `?..`** — N1/N2 guard 1.
|
|
723
|
+
- **[P] A unit with no space before it.** `5km` and `5%` stay exactly as typed; N5 converts, never
|
|
724
|
+
inserts (§7.2).
|
|
725
|
+
- **[P] A number written as `H2O`, `A4`, `MP3`.** N5 step 4's letter guard.
|
|
726
|
+
- **[P] An existing U+00A0 or U+202F that is already the right character.** Every sub-rule's
|
|
727
|
+
"already correct" branch emits nothing, so a correctly-typeset document round-trips
|
|
728
|
+
byte-identically. This is the property PLAN.md §3.4 exists for.
|
|
729
|
+
- **[P] An existing U+2007, U+2009, U+200A, U+2060 or U+FEFF.** None is in `NOBREAK`; none is
|
|
730
|
+
converted. If an author placed a thin space, it stays.
|
|
731
|
+
- **[P] The quotation marks themselves.** N8 only touches the space beside them.
|
|
732
|
+
- **[P] Anything inside a skipped region.** Handled by the mode adapter.
|
|
733
|
+
|
|
734
|
+
---
|
|
735
|
+
|
|
736
|
+
## 5. Idempotency argument
|
|
737
|
+
|
|
738
|
+
Let `T` be the rule and consider `y = T(x)`.
|
|
739
|
+
|
|
740
|
+
**Every sub-rule has an explicit "already correct" branch that emits nothing**, and the
|
|
741
|
+
target of each sub-rule is exactly the state that branch recognises:
|
|
742
|
+
|
|
743
|
+
| Sub-rule | Post-state at the claimed index | Recognised as already correct by |
|
|
744
|
+
| -------- | ----------------------------------------------------- | ------------------------------------------ |
|
|
745
|
+
| N1 | `NBSP` immediately left of the mark | §3.3 step 4, first bullet |
|
|
746
|
+
| N2 | `NARROW-TARGET` immediately left of the mark | §3.4 |
|
|
747
|
+
| N3 | `NBSP` after the short word | §3.5 step 3 |
|
|
748
|
+
| N4 | `NBSP` at every internal position of the abbreviation | §3.6 step 1 (space-lenient match) + step 3 |
|
|
749
|
+
| N5 | `NBSP` before the unit | §3.7 step 3 |
|
|
750
|
+
| N6 | `NBSP` after the symbol | §3.8 step 3 |
|
|
751
|
+
| N7 | `NBSP` at the bound space | §3.9 |
|
|
752
|
+
| N8 | `target` beside the quote glyph | §3.10 |
|
|
753
|
+
| N9 | `NBSP` between the abbreviation and the number | §3.11 step 3 |
|
|
754
|
+
| N10 | `NBSP` between the abbreviation and the word | §3.12 step 3 |
|
|
755
|
+
|
|
756
|
+
So it suffices to show that on the second run **the same sub-rule claims the same index** —
|
|
757
|
+
i.e. that no guard's verdict flips because of an edit made on the first run.
|
|
758
|
+
|
|
759
|
+
Guards test three kinds of thing:
|
|
760
|
+
|
|
761
|
+
0. **The `narrowNbsp` substitution (§3.1a) needs no separate argument.** Every row of the table
|
|
762
|
+
above names a sub-rule's **target**, not a literal code point, and N2's and N8's targets are
|
|
763
|
+
the only ones the option moves. Substituting `NARROW-TARGET` therefore carries the whole
|
|
764
|
+
argument over unchanged: the post-state each branch recognises moves with the character each
|
|
765
|
+
branch writes, which is precisely what post-processing the output could not do (§3.1a). The
|
|
766
|
+
substitution also cannot create a **new** conflict between two sub-rules, because it only ever
|
|
767
|
+
makes N2's and N8's targets **equal to** N1's, never different from it, and no sub-rule's claim
|
|
768
|
+
depends on what another sub-rule's target is — the quote-glyph guard (§3.3 step 3) decides
|
|
769
|
+
ownership of the index beside a quotation glyph by position, whatever character either
|
|
770
|
+
sub-rule would have written there.
|
|
771
|
+
1. **Membership in `SPACELIKE`.** Every conversion `SP → NBSP`, `SP → NNBSP`,
|
|
772
|
+
`NNBSP → NBSP`, `NBSP → NNBSP` stays inside `SPACELIKE` (§3.1). Every insertion adds a
|
|
773
|
+
`NOBREAK`, which is in `SPACELIKE`. So every boundary test that passed on run 1 passes on
|
|
774
|
+
run 2, and every one that failed still fails — **provided** boundary tests never use `SP`
|
|
775
|
+
specifically. They do not: §3.1 states this and every guard above is written against
|
|
776
|
+
`SPACELIKE`, `ALNUM`, `UPPER`, `DIGIT`, `OPENISH`, `CLOSEISH`, `BREAK` or `NONE`, none of
|
|
777
|
+
which gains or loses a member under a `SP`/`NOBREAK` conversion.
|
|
778
|
+
2. **Membership in `ALNUM`, `DIGIT`, `UPPER`, `LETTER`, `OPENISH`, `CLOSEISH`.** No sub-rule
|
|
779
|
+
ever edits a code point in any of these classes; every edit replaces or inserts a space
|
|
780
|
+
character. Unchanged.
|
|
781
|
+
3. **Literal matching (N3, N4, N5, N6, N9, N10).** N4 is space-lenient by construction, so a matched
|
|
782
|
+
abbreviation still matches after its internal spaces become `NBSP`. N3, N5, N6, N9 and N10 match
|
|
783
|
+
only the word/unit/symbol/abbreviation itself, never the adjacent space, so their matches
|
|
784
|
+
are untouched. N9 and N10 additionally require the separator to be `SP` or `NBSP` and
|
|
785
|
+
treat `NBSP` as "already correct", so their second-run verdict is "emit nothing" at the
|
|
786
|
+
very index they claimed on the first run. The _side conditions_ that read the adjacent space accept both `SP` and `NBSP`
|
|
787
|
+
and distinguish them only to decide "convert" versus "emit nothing".
|
|
788
|
+
|
|
789
|
+
Two specific insertion cases need checking because they change lengths:
|
|
790
|
+
|
|
791
|
+
- **N1/N2 insertion.** Run 1 turns `mot!` into `mot` `NNBSP` `!`. On run 2, `left` of the `!`
|
|
792
|
+
is `NNBSP` → "already correct" → no edit. Note the insertion did **not** create a new
|
|
793
|
+
candidate: the inserted character is a space, and no sub-rule's _match_ is a space (N4
|
|
794
|
+
matches a literal containing spaces, but the literal must also match its non-space code
|
|
795
|
+
points, and inserting one space next to a punctuation mark cannot complete an abbreviation
|
|
796
|
+
match that failed before, because the abbreviation's non-space code points are unchanged).
|
|
797
|
+
- **N8 insertion.** Run 1 turns `«mot»` into `«` `NNBSP` `mot` `NNBSP` `»`. On run 2 both
|
|
798
|
+
positions hit the "already `target`" branch. Additionally the inserted `NNBSP` becomes the
|
|
799
|
+
`left` neighbour of nothing punctuation-like, and the `right` neighbour of `«`, which is not
|
|
800
|
+
a candidate for any other sub-rule.
|
|
801
|
+
|
|
802
|
+
Finally, the **first-claim-wins** ordering of §3.2 is deterministic and depends only on the
|
|
803
|
+
sub-rule index, not on the array contents, so the same sub-rule wins on both runs.
|
|
804
|
+
|
|
805
|
+
Hence `T(T(x)) = T(x)`.
|
|
806
|
+
|
|
807
|
+
**What had to be fixed.** Four things, all of them the difference between a rule that works
|
|
808
|
+
and a rule that oscillates:
|
|
809
|
+
|
|
810
|
+
1. **`NOBREAK` had to be inside `SPACELIKE`.** The naive version writes boundary tests
|
|
811
|
+
against U+0020, so after run 1 converts the space, run 2 no longer sees a word boundary,
|
|
812
|
+
and a _different_ sub-rule (or none) claims the position. Symptom: N3 and N5 fight over
|
|
813
|
+
`5 km` in a locale that lists `km` as a unit and has a short word ending in a way that
|
|
814
|
+
overlaps.
|
|
815
|
+
2. **N4 had to match space-leniently.** Matching `z. B.` literally means the converted
|
|
816
|
+
`z.` `NBSP` `B.` no longer matches, which is harmless on its own — but it means the
|
|
817
|
+
abbreviation is invisible to sub-rule ordering on run 2 and a lower-priority sub-rule
|
|
818
|
+
(N3, say, on the word `z`) can claim the same index with a different verdict. Making the
|
|
819
|
+
match space-lenient keeps N4 in control of its own indices forever.
|
|
820
|
+
3. **N5 and N6 had to be conversion-only.** The naive version inserts a space before a unit,
|
|
821
|
+
turning `5km` into `5 km` — a content change, not a typographic one — and, worse, in a
|
|
822
|
+
locale listing `%` it turns `50%` into `50 %`, which is wrong in English and right in
|
|
823
|
+
French, i.e. it is a locale decision that `beforeUnits` alone cannot express. Restricting
|
|
824
|
+
to conversion makes the rule safe and shifts the question to §7.2 where it belongs.
|
|
825
|
+
4. **The conflict policy had to be total and index-based.** "Whichever sub-rule runs first in
|
|
826
|
+
the loop" is not a specification; it is a Go map iteration bug waiting to happen (§3.2).
|
|
827
|
+
|
|
828
|
+
---
|
|
829
|
+
|
|
830
|
+
### Composition obligation
|
|
831
|
+
|
|
832
|
+
Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₈** and runs last,
|
|
833
|
+
so nothing has to preserve _its_ invariant — but it must preserve all seven others, and it is
|
|
834
|
+
the rule with the largest emission surface in the pipeline. It is also the rule that committed
|
|
835
|
+
defect family 2.
|
|
836
|
+
|
|
837
|
+
**What this rule emits.** U+00A0 and U+202F only, either replacing a single space-like code
|
|
838
|
+
point in place or **inserted** at one of three kinds of site: before a listed punctuation
|
|
839
|
+
character (N1, N2), after an opening quote glyph, and before a closing quote glyph (N8). It
|
|
840
|
+
never emits U+0020, never emits a letter, digit, dot, dash or quotation mark, and never deletes
|
|
841
|
+
anything.
|
|
842
|
+
|
|
843
|
+
**Against `I₁` (`spaces`).** Discharged, and this is why the rule never emits U+0020: U+00A0
|
|
844
|
+
and U+202F are `CONTENT` to `spaces`, which touches only U+0020. Converting a U+0020 to a
|
|
845
|
+
no-break space can only _remove_ a potential `spaces` edit, never create one. An insertion adds
|
|
846
|
+
a `CONTENT` code point, which cannot lengthen a U+0020 run or place a U+0020 in a stripping
|
|
847
|
+
position.
|
|
848
|
+
|
|
849
|
+
**Against `I₂` (`ellipsis`).** Discharged: no `DOTLIKE` code point is emitted, and insertions
|
|
850
|
+
only push code points apart, never together.
|
|
851
|
+
|
|
852
|
+
**Against `I₃` (`dashes`).** **This is the obligation the rule failed, twice**, and it is now
|
|
853
|
+
discharged structurally rather than by argument.
|
|
854
|
+
|
|
855
|
+
N8 inserting a `quotes.innerSpace` beside a quoted hyphen turned `«-»` into `«⍽-⍽»`, which
|
|
856
|
+
`dashes` read as a spaced parenthetical dash on the following pass — defect family 2. The
|
|
857
|
+
repair made a right-hand no-break space not count as dash spacing, which left the _asymmetric_
|
|
858
|
+
shape reachable: `«–␣"` gains an inner U+00A0 from N8, and a left-hand no-break space still
|
|
859
|
+
counted, so the token became symmetric and was promoted from en to em — defect (d),
|
|
860
|
+
`dashes.md` §5.2. Both were caused by this sub-rule's insertions, and both repairs were
|
|
861
|
+
attempts to specify which no-break space counts as spacing on which side.
|
|
862
|
+
|
|
863
|
+
**The discharge no longer depends on that.** `E(nbsp) = { U+00A0, U+202F }` — those are the
|
|
864
|
+
only code points this rule can emit or insert, in any sub-rule, under any locale data. `dashes`
|
|
865
|
+
now treats **both** as making an adjacent token inert (`dashes.md` §3.2 step 3), so no emission
|
|
866
|
+
of this rule can create a dash token, change one's spacing verdict, or revive one `dashes`
|
|
867
|
+
declined. That is condition **CO-S** of `pipeline-idempotency.md` §5.1a, and it holds for every
|
|
868
|
+
input rather than for the witnesses anyone happened to test.
|
|
869
|
+
|
|
870
|
+
The Russian direction — N1/N2 promoting the U+0020 before an em dash to U+00A0 — is covered by
|
|
871
|
+
the same statement: `dashes` declines the promoted token at its symmetry guard and emits
|
|
872
|
+
nothing, so the output `Москва⍽— столица` is stable.
|
|
873
|
+
|
|
874
|
+
**The repair stayed in `dashes` rather than moving here**, for the layering reason in
|
|
875
|
+
`pipeline-idempotency.md` §4: `order.json` gives `dashes` `"localeData": ["dash"]` and no access
|
|
876
|
+
to `quotes`, so it cannot recognise a guillemet; and requiring this rule to suppress an
|
|
877
|
+
insertion that might create a dash token would mean encoding `dashes`' admissibility rules in a
|
|
878
|
+
second place, which is how the two rules drifted apart in the first place.
|
|
879
|
+
|
|
880
|
+
**Against `I₄` (`hyphen`).** Discharged, but only just, and the reason is worth stating. A
|
|
881
|
+
listed hyphen form's boundary guards ask whether the neighbouring code point is in `WORDISH`.
|
|
882
|
+
An insertion beside such a form would change that neighbour from whatever it was to a no-break
|
|
883
|
+
space, which is not `WORDISH` — so the guard could go from _reject_ to _accept_ and a form
|
|
884
|
+
could start matching that did not before. That would be an `I₄` violation. It is unreachable
|
|
885
|
+
because every insertion site puts the new space next to a listed punctuation character or a
|
|
886
|
+
quote glyph, never between two letters, and a form whose neighbour is punctuation or a quote
|
|
887
|
+
glyph was already accepted. If a future sub-rule inserts between two letters, this obligation
|
|
888
|
+
must be re-derived.
|
|
889
|
+
|
|
890
|
+
**Against `I₅`, `I₆`, `I₇` (`quotes`, `apostrophe`, `symbols`).** Discharged by case analysis
|
|
891
|
+
over the neighbour tests; the `quotes` half is the one that matters and is written out in
|
|
892
|
+
`quotes.md` §5.6, which shows that neither insertion site can _add_ a capability to a surviving
|
|
893
|
+
straight mark. `apostrophe`'s case ladder reads `ALNUM`, `LETTER`, `DIGIT`, `SPACELIKE`,
|
|
894
|
+
`OPENISH`, `OPENQUOTE` and `CLOSEISH`; an inserted no-break space is `SPACELIKE`, and the only cases it
|
|
895
|
+
could newly satisfy require a `SPACELIKE` **left** neighbour, which is case 4 (leading elision)
|
|
896
|
+
— and case 4 also requires an `ALNUM` right neighbour, which an insertion cannot create.
|
|
897
|
+
`symbols` reads `ALNUM`, `DIGIT` and symmetric spacing; a no-break space is already accepted as
|
|
898
|
+
spacing there (`symbols.md` §3.3 step 1), and converting U+0020 to U+00A0 leaves `lsp`/`rsp` unchanged.
|
|
899
|
+
|
|
900
|
+
---
|
|
901
|
+
|
|
902
|
+
## 6. Worked examples
|
|
903
|
+
|
|
904
|
+
`␣` = U+0020, `⍽` = U+00A0, `⍹` = U+202F, `⟶` = no change. Each block states the fields it uses,
|
|
905
|
+
and **those fields are quoted from the shipped locale file** rather than assumed. That
|
|
906
|
+
distinction is not pedantry: while this preamble said the locale files did not exist yet, rows
|
|
907
|
+
row 13a drifted into asserting behaviour the shipped `ru.json` does not produce — and a
|
|
908
|
+
neighbouring row asserted an `ru` `beforeNumber` binding that has neither data nor a source
|
|
909
|
+
behind it, and has since been deleted rather than softened. Both survived several reviews
|
|
910
|
+
because nothing obliged anyone to check. A row here is a claim
|
|
911
|
+
about the engine **and** about a file in `spec/locales/`, and both halves have to hold.
|
|
912
|
+
|
|
913
|
+
### `fr` — `narrowBeforePunctuation: ["?","!",";"]`, `beforePunctuation: [":"]`, primary `« »` with `innerSpace: "nbsp"`
|
|
914
|
+
|
|
915
|
+
| # | Input | Output | Why |
|
|
916
|
+
| --- | ----------------------------------- | ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
917
|
+
| 1 | `Bonjour!` | `Bonjour⍹!` | N2 insertion; the ordinary space, if any, was already removed by `spaces` |
|
|
918
|
+
| 2 | `Bonjour⍹!` | ⟶ | N2 "already correct" — this is the round-trip case |
|
|
919
|
+
| 3 | `Il a dit «mot».` | `Il a dit «⍽mot⍽».` | N8 inserts on both inner edges. **U+00A0, not U+202F**: `fr.json` sets the primary pair's `innerSpace` to `"nbsp"` — corrected in spec 1.3.0, having claimed the narrow space since this table was written |
|
|
920
|
+
| 4 | `Il a dit «⍹mot⍹».` | `Il a dit «⍽mot⍽».` | and therefore **not** already correct: N8 converts an authored narrow space to its own target. The row claimed a fixed point on the same mistake |
|
|
921
|
+
| 5 | `Voir http://example.org: la suite` | `Voir http://example.org⍽: la suite` | N1 guard 2 rejects the colon in `http://` (next code point is `/`) and accepts the sentence colon (next is a space) |
|
|
922
|
+
| 6 | `Vraiment?!` | `Vraiment⍹?!` | N1/N2 guard 1: only the first mark takes a space |
|
|
923
|
+
| 6a | `Vraiment?…` | `Vraiment⍹?…` | U+2026 is accepted right context (§3.3 step 2). Previously the narrow space was omitted here but not in `Vraiment␣?`, purely because `ellipsis` had already run |
|
|
924
|
+
| 6b | `Voir␣../docs` | ⟶ | the French instance of the `spaces` lone-dot defect; see `spaces.md` §3.4 |
|
|
925
|
+
| 6c | `<strong>gel␣:</strong>␣il` (`html`) | `<strong>gel⍽:</strong>␣il` | **spec 1.2.0.** `spaces` deletes the U+0020; the mark's right neighbour is the span boundary marker, which is in `CLOSEISH`, so N1 inserts U+00A0 at the mark's index, inside the span. Previously produced `gel:</strong>` |
|
|
926
|
+
| 6d | `réglementation␣?<br>Oui` (`html`) | `réglementation⍹?<br>Oui` | same, with N2. `<br>` leaves no line terminator in the gap, so the marker is −1, not −2 |
|
|
927
|
+
| 6e | `<em>non</em>␣!␣Oui` (`html`) | `<em>non</em>⍹!␣Oui` | `spaces` leaves the run alone (its `left` is a span edge), and N2 converts it. The marker two places left of the mark is not `OPENISH`, so step 3 does not decline |
|
|
928
|
+
| 6f | `12:<b>30</b>` (`html`) | `12⍽:<b>30</b>` | **the accepted cost of 6c** — step 2 cannot see past the marker |
|
|
929
|
+
|
|
930
|
+
### `ru` — `afterShortWords: ["в","и","на",…]`, `abbreviations: ["т. д.","и т. п."]`, `initialBinding: "chain"`, `afterSymbols: ["№","§"]`
|
|
931
|
+
|
|
932
|
+
| # | Input | Output | Why |
|
|
933
|
+
| --- | ----------------------- | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
934
|
+
| 7 | `Он живёт в␣Москве` | `Он живёт в⍽Москве` | N3: left boundary is a space, right is `SP`, following token is a letter |
|
|
935
|
+
| 8 | `Он живёт в⍽Москве` | ⟶ | N3 already correct |
|
|
936
|
+
| 9 | `и␣т.␣д.` | `и⍽т.⍽д.` | the leading `и` by N3, the internal space by N4 (`т. д.`) |
|
|
937
|
+
| 10 | `А.␣С.␣Пушкин` | `А.⍽С.⍽Пушкин` | N7 clause C1 twice |
|
|
938
|
+
| 11 | `Пушкин␣А.␣С.` | `Пушкин⍽А.⍽С.` | first space by C2 (two initials follow), second by C1 |
|
|
939
|
+
| 12 | `см. №␣5` | `см. №⍽5` | N6 |
|
|
940
|
+
| 13 | `Иван␣пошёл␣домой` | ⟶ | no short word, no abbreviation, no unit, no initial |
|
|
941
|
+
| 13a | `г.␣Москва, ул.␣Ленина` | `г.␣Москва, ул.⍽Ленина` | N10 binds **only `ул.`**. `ru.json`'s `beforeWord` is `["ул.","пл."]`; `г.` is not in it — it is in `beforeUnits`, where it binds **backwards** to a preceding year (`1990⍽г.`), which is the actual Russian convention. Row checked against the shipped file, not against a hypothetical one |
|
|
942
|
+
| 13c | `г.⍽Москва` | ⟶ | N10 "already correct" |
|
|
943
|
+
| 13d | `г.Москва` | ⟶ | N9/N10 never insert |
|
|
944
|
+
| 13e | `г. «Москва»` | ⟶ | N10 guard G-W: the following code point is a quotation glyph, not a letter |
|
|
945
|
+
| 13f | `из-за␣дождя` | `из‑за␣дождя` | `hyphen` binds the compound; N3 does **not** then bind `за`, because its left neighbour is U+2011 and a hyphen fails the left boundary (§3.5 step 2). Previously produced `из‑за⍽дождя`, a false positive on ordinary prose |
|
|
946
|
+
| 13g | `из-под␣стола` | `из‑под␣стола` | same shape with the other listed compound |
|
|
947
|
+
| 13h | `—␣в␣Москве` | `—␣в⍽Москве` | `SENTENCE-DASH` still opens a phrase, so a genuine preposition after an em dash binds normally |
|
|
948
|
+
| 13i | `в␣<em>Москве</em>` (`html`) | ⟶ | **spec 1.2.0.** N3's following-token guard asks for `ALNUM` or `OPENISH`, and for this rule the span boundary marker is in neither (§3.1). The space stays U+0020. Whether it should bind is §7 item 12 |
|
|
949
|
+
|
|
950
|
+
### `de-DE` — `abbreviations: ["z. B.","d. h."]`, `beforeUnits: ["%","km","°C"]`
|
|
951
|
+
|
|
952
|
+
| # | Input | Output | Why |
|
|
953
|
+
| --- | ----------------------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
954
|
+
| 14 | `z.␣B.␣Berlin` | `z.⍽B.␣Berlin` | N4 binds the space inside the abbreviation. The space after `B.` is **not** bound: guard C1-a sees that the `B.` is preceded by the lower-case `z.` and declines. Previously produced `z.⍽B.⍽Berlin`, a false positive on ordinary prose |
|
|
955
|
+
| 15 | `z.⍽B. Berlin` | ⟶ | N4 space-lenient match, already correct |
|
|
956
|
+
| 16 | `Es sind 20␣km bis 5␣%` | `Es sind 20⍽km bis 5⍽%` | N5 twice |
|
|
957
|
+
| 17 | `Es sind 20km` | ⟶ | N5 never inserts (§7.2) |
|
|
958
|
+
| 18 | `H2␣O ist kein Wert` | ⟶ | N5 step 4: `2` is preceded by the letter `H` |
|
|
959
|
+
| 19 | `siehe Nr.␣5 und S.␣12` | `siehe Nr.⍽5 und S.⍽12` | N9 twice |
|
|
960
|
+
| 20 | `siehe nr.␣5` | ⟶ | N9 matches exactly; `nr.` is not `Nr.` |
|
|
961
|
+
| 21 | `Bonjour :␣oui` (`fr`) | ⟶ | **spec 1.3.0.** N2's character-reference guard: the `;` closes ` `, so no U+202F is inserted and the reference survives. Before 1.3.0 this produced `Bonjour ·:␣oui` |
|
|
962
|
+
| 22 | `Tom␣&␣Jerry` (`fr`) | ⟶ | **spec 1.3.0.** Same guard on a named reference |
|
|
963
|
+
| 23 | `Oui␣;␣non` (`fr`) | `Oui·;␣non` | The guard is not a blanket refusal of `;` — an ordinary semicolon still takes U+202F |
|
|
964
|
+
|
|
965
|
+
#### `narrowNbsp: "nbsp"` (§3.1a, spec 1.3.0)
|
|
966
|
+
|
|
967
|
+
Every row `fr`, and every row measured. `⍽` = U+00A0, `⍹` = U+202F.
|
|
968
|
+
|
|
969
|
+
| # | Input | Default (`narrow`) | `narrowNbsp: "nbsp"` | Why |
|
|
970
|
+
| --- | ---------------------------- | ------------------------------------- | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
|
|
971
|
+
| 31 | `Un délai ?` | `Un délai⍹?` | `Un délai⍽?` | N2's target moves; the index it claims, and the guard that let it claim it, do not |
|
|
972
|
+
| 32 | `Oui⍹?` | ⟶ | `Oui⍽?` | an authored narrow space at a claimed index is normalised to the target, as an authored U+00A0 is in the default configuration |
|
|
973
|
+
| 33 | `Il a dit : « oui » ; puis ?` | `Il a dit⍽: «⍽oui⍽»⍹; puis⍹?` | `Il a dit⍽: «⍽oui⍽»⍽; puis⍽?` | N1 (colon) and N8 (`fr`'s pair, `innerSpace: "nbsp"`) already wrote U+00A0 and do not move; only N2 does |
|
|
974
|
+
| 34 | `12:30 et http://x ; oui` | `12:30 et http://x⍹; oui` | `12:30 et http://x⍽; oui` | the option changes what is written, never what is read: N1's right-context guard still protects the time and the URL |
|
|
975
|
+
|
|
976
|
+
### `fr` — `beforeWord: ["M.","Mme","Mlle"]`
|
|
977
|
+
|
|
978
|
+
| # | Input | Output | Why |
|
|
979
|
+
| --- | ----------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
980
|
+
| 21 | `M.␣Dupont et Mme␣Hugo` | `M.⍽Dupont et Mme⍽Hugo` | `Mme` binds via N10 (three code points, G-D does not apply). `M.` is inert in N10 because `fr` has `initialBinding: "single"` and G-D fires — it binds via **N7 C1** instead, eligible under `"single"`'s unconditional initial-to-word shape, reaching the same U+00A0 |
|
|
981
|
+
| 22 | `M.␣dupont` | ⟶ | lower-case surname: N7 C1 needs `UPPER` after the space, and N10 is inert for `M.`. A documented gap, §7.10 |
|
|
982
|
+
| 23 | `«?` | `«⍽?` | N1/N2 decline the index (quote-glyph guard, §3.3 step 3) and N8 owns it, inserting the `innerSpace` target. Previously this oscillated between U+00A0 and U+202F for ever — §3.10.1 |
|
|
983
|
+
| 24 | `«⍽?` | ⟶ | the fixed point of case 23 |
|
|
984
|
+
| 25 | `mot␣?` | `mot⍹?` | the ordinary case: the code point before the space is a letter, so N2 applies normally |
|
|
985
|
+
|
|
986
|
+
### `el` — every list empty, `initialBinding: "none"`, `quotes.innerSpace: "none"`
|
|
987
|
+
|
|
988
|
+
The first locale for which this rule is a **total no-op**, in the same provable sense `hyphen` is
|
|
989
|
+
a no-op for a locale with empty lists. It is worth a block of its own because the claim needs
|
|
990
|
+
_both_ halves of `order.json`'s `"localeData": ["nbsp", "quotes"]` to hold, and only one of them
|
|
991
|
+
is visible in the `nbsp` object: the eight lists being empty disables N1–N7 and N9–N10, and
|
|
992
|
+
**`quotes.innerSpace: "none"` is what additionally disables N8**. A locale with empty lists but a
|
|
993
|
+
non-`none` `innerSpace` would still edit — `fr` case 23 is exactly that shape — so "all lists
|
|
994
|
+
empty" alone does not license the claim.
|
|
995
|
+
|
|
996
|
+
Greek is also the case that shows why N1/N2 being empty is a positive finding rather than an
|
|
997
|
+
unfilled field. The Greek source denies the French space-before-punctuation pattern **by name**
|
|
998
|
+
(«πράγμα που συμβαίνει, π.χ., στα γαλλικά»), so `spaces` strips the space at order 10 and nothing
|
|
999
|
+
here puts one back. That is a cited decision, not a default.
|
|
1000
|
+
|
|
1001
|
+
| # | Input | Output | Why |
|
|
1002
|
+
| --- | -------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
1003
|
+
| 26 | `Τι κάνεις;` | ⟶ | `beforePunctuation` and `narrowBeforePunctuation` are empty, so N1/N2 have no member to match. Contrast `fr` case 25, where the same shape gains U+202F |
|
|
1004
|
+
| 27 | `10,5␣%␣φέτος` | ⟶ | `beforeUnits` is empty, so N5 is inert and the ordinary space survives as U+0020 |
|
|
1005
|
+
| 28 | `«καλημέρα»` | ⟶ | `innerSpace: "none"`, so N8 inserts nothing. This is the half of the no-op claim that lives in the `quotes` object rather than the `nbsp` one |
|
|
1006
|
+
|
|
1007
|
+
Cases 2, 6b, 8, 13, 13c, 13d, 13e, 13i, 15, 17, 18, 20, 22, 24, 26, 27 and 28 are "no change"
|
|
1008
|
+
cases. (Case 4 left that list in spec 1.3.0, when the row was corrected from a fixed point to a
|
|
1009
|
+
conversion; the §3.1a rows are numbered 31-34 and are outside it.)
|
|
1010
|
+
|
|
1011
|
+
---
|
|
1012
|
+
|
|
1013
|
+
## 7. Open questions
|
|
1014
|
+
|
|
1015
|
+
1. **The schema does not enforce that `beforePunctuation` and `narrowBeforePunctuation` are
|
|
1016
|
+
disjoint**, nor that a character appears in only one of `beforeUnits` / `afterSymbols`.
|
|
1017
|
+
§2 requires an implementation to fail loudly; a `$defs`-level constraint or a CI check
|
|
1018
|
+
would be better. Reported.
|
|
1019
|
+
2. **N5/N6 never insert a space.** `5km` stays `5km` and `50%` stays `50%`. For German,
|
|
1020
|
+
Duden requires a space before `%` and before a unit, so `50%` is arguably _wrong_ input
|
|
1021
|
+
that polytypo declines to fix. Inserting is a content change and the schema has no flag to
|
|
1022
|
+
authorise it per locale (something like `"insertBeforeUnits": true`). Needs an operator
|
|
1023
|
+
decision; I chose the conservative branch because of the zero-false-positive ship
|
|
1024
|
+
criterion.
|
|
1025
|
+
3. **`afterShortWords` matching is "case-insensitive only for the first character"** per the
|
|
1026
|
+
schema description, which I have implemented as the Unicode **simple** uppercase mapping.
|
|
1027
|
+
That is a data table, not a locale operation, so it satisfies §4.4. **The spec pins a Unicode version —
|
|
1028
|
+
`spec/UNICODE` contains `17.0` — and the pin is now normative for the derived tables**
|
|
1029
|
+
(`pipeline-idempotency.md` §6a). §3.1 cites it, as do the other four rules that read general
|
|
1030
|
+
categories. An earlier revision of this item said no pin existed and then that it had no
|
|
1031
|
+
reader; both were true when written and neither is now.
|
|
1032
|
+
4. **`innerSpace: "none"` does not remove an existing space inside quotation marks** — now
|
|
1033
|
+
stated normatively in §3.10 rather than left open, with the deletion-is-not-correction
|
|
1034
|
+
argument. `de-CH` `« hallo »` and `en-US` `“ hello ”` both survive as typed. The live
|
|
1035
|
+
question is only whether the operator wants the opposite: a locale declaring `"none"` could
|
|
1036
|
+
be read as asserting that no inner space is permissible, in which case removal would be a
|
|
1037
|
+
correction and not a mutilation. I did not take that reading, because it is the only place
|
|
1038
|
+
in the spec where a locale field would license a deletion. **Operator decision**; the
|
|
1039
|
+
fixtures currently pin non-removal.
|
|
1040
|
+
5. **A quote pair whose `open` equals `close` cannot have an `innerSpace`.** Finnish and
|
|
1041
|
+
Swedish use U+201D on both sides, and their `innerSpace` is expected to be `"none"`, so the
|
|
1042
|
+
case is currently vacuous — but the schema permits `{"open":"”","close":"”","innerSpace":
|
|
1043
|
+
"nbsp"}`, which this rule cannot implement (it cannot tell an opener from a closer without
|
|
1044
|
+
the pairing information that only `quotes` has, and `quotes` does not pass state to `nbsp`).
|
|
1045
|
+
§3.10 makes it a documented no-op. Either the schema should forbid it or the pipeline
|
|
1046
|
+
should let `quotes` hand its pair positions to `nbsp` — the latter breaks the "rules are
|
|
1047
|
+
independent" property and I did not propose it unilaterally. Reported.
|
|
1048
|
+
6. **`initialBinding` clause C2 requires two consecutive initials, in every mode.** `Пушкин А.`
|
|
1049
|
+
(one initial) is not bound. Restricting it this way avoids binding every `word` + `Capital.`
|
|
1050
|
+
pair, but it is a guess about Russian practice and needs checking against Мильчин.
|
|
1051
|
+
7. _(Settled.)_ `nbsp.beforeNumber` and `nbsp.beforeWord` were added to the schema and are
|
|
1052
|
+
specified as N9 (§3.11) and N10 (§3.12). `Nr. 5`, `S. 12`, `art. 237`, `г. Москва`,
|
|
1053
|
+
`ул. Ленина`, `M. Dupont` are all now expressible.
|
|
1054
|
+
8. _(Settled.)_ U+2011 is produced by the `hyphen` rule at order 35 — see
|
|
1055
|
+
[hyphen.md](hyphen.md). `nbsp` still never produces it, which is correct: a hyphen is not a
|
|
1056
|
+
space and binding it is a different rule with different locale data. U+2060 (word joiner)
|
|
1057
|
+
is produced by `dashes` (order 30) around a tight range dash — see
|
|
1058
|
+
[dashes.md](dashes.md) §3.3.1. `nbsp` still never produces it, and still never converts an
|
|
1059
|
+
author's own U+2060.
|
|
1060
|
+
9. **N10 has no sentence-boundary guard — but with the shipped locale data the risk is not
|
|
1061
|
+
reachable, and an earlier revision of this item was wrong to call it "the single most likely
|
|
1062
|
+
`nbsp` false positive at the M4 gate".** That misdirected the review, which is worse than
|
|
1063
|
+
saying nothing: it pointed the gate at an input the engine cannot produce.
|
|
1064
|
+
|
|
1065
|
+
The mechanism is real in the abstract. If a locale listed `г.` in `beforeWord`, then
|
|
1066
|
+
`в 1990 г. Москва…` would bind across a sentence boundary, because nothing distinguishes
|
|
1067
|
+
*год* from *город* syntactically. **`ru.json` does not list it.** `beforeWord` is
|
|
1068
|
+
`["ул.","пл."]` — two forms that are unambiguously nouns and cannot end a sentence in the
|
|
1069
|
+
relevant sense — and `г.` is instead in `beforeUnits`, binding **backwards** to a preceding
|
|
1070
|
+
year. That is a better answer than any guard this rule could have carried: it encodes the
|
|
1071
|
+
reading (*год*) that actually collides, and it binds in the direction that reading requires,
|
|
1072
|
+
so the forward-binding ambiguity never arises.
|
|
1073
|
+
|
|
1074
|
+
What remains open is the general shape, not this instance: **a locale author may still put an
|
|
1075
|
+
ambiguous form in `beforeWord`**, and this rule will bind it across a sentence boundary
|
|
1076
|
+
without complaint. The mitigation is locale-data curation and the schema description already
|
|
1077
|
+
says so ("Riskier than beforeNumber — list only forms whose binding is normative"). The
|
|
1078
|
+
lesson worth keeping is the one this entry got wrong: **a prose claim about what a rule does
|
|
1079
|
+
to a language must be checked against the shipped locale file**, because the rule alone does
|
|
1080
|
+
not determine it.
|
|
1081
|
+
9a. **N7's own sentence-boundary miss (spec 0.6.0) — resolved for `"chain"`, deliberately left
|
|
1082
|
+
open for `"single"`.** A fresh M4 dogfooding pass surfaced the C1 sibling of item 9's N10
|
|
1083
|
+
miss: `"...take the top N. It runs..."` bound `N.` to `It` across a sentence boundary,
|
|
1084
|
+
because the old boolean `bindInitials` let C1's initial-to-word shape fire on any lone
|
|
1085
|
+
initial next to any uppercase-starting word, with no check that a genuine name — as opposed
|
|
1086
|
+
to an ordinary sentence ending in a single capital letter — was actually present. `"chain"`
|
|
1087
|
+
closes this for en-US/de-DE/de-CH/ru by requiring a confirmed sequence of two or more
|
|
1088
|
+
initials before that shape binds (Chicago's own "two or more initials" wording, applied
|
|
1089
|
+
literally rather than only for turning N7 on at all) — see §3.9's mode condition.
|
|
1090
|
+
|
|
1091
|
+
**This does not close the miss for `"single"` (`fr`/`fr-CA`).** `N. Bourbaki` and
|
|
1092
|
+
`M. Dupont` — both cited, both canonical fixtures — are structurally the identical shape to
|
|
1093
|
+
the English witness: one lone initial, one following capitalized word, no preceding initial
|
|
1094
|
+
and no further initial on the right. No local, non-lexical structural signal distinguishes
|
|
1095
|
+
them (both my English witness and both French citations happen to sit at the very start of
|
|
1096
|
+
their sentence, which is suggestive but not load-bearing evidence, and is not implemented as
|
|
1097
|
+
a guard here — it is not reliable enough to specify without lexical or genuine
|
|
1098
|
+
sentence-segmentation data, which this rule does not have and is not mine to add
|
|
1099
|
+
unilaterally). A French sentence ending in a lone initial immediately followed by a new
|
|
1100
|
+
sentence starting with a capitalized word would misfire under `"single"` exactly as the
|
|
1101
|
+
English witness did before `"chain"` existed. This is a **known, accepted limitation of
|
|
1102
|
+
`"single"` mode**, not a regression: the alternative (applying `"chain"`'s restriction to
|
|
1103
|
+
French too) would break `N. Bourbaki`/`M. Dupont`, which André's own citation supports and
|
|
1104
|
+
which have been canonical fixtures since before this item was written. Revisiting it needs
|
|
1105
|
+
either a genuinely new structural signal (none is known) or a locale-specific operator
|
|
1106
|
+
decision to accept `"single"`'s narrower false-positive-vs-false-negative trade for French,
|
|
1107
|
+
which is the decision already made, recorded here rather than left implicit.
|
|
1108
|
+
10. **G-D is conditional on `initialBinding !== "none"`**, so a `beforeWord` entry's behaviour depends on
|
|
1109
|
+
an unrelated field. That is the only coupling of its kind in the rule and it is mildly
|
|
1110
|
+
unpleasant, but every alternative is worse: an absolute guard makes `г.` inert, and no
|
|
1111
|
+
guard at all lets N7 and N10 both claim `M. Dupont`. The residual cost is §6 case 22 —
|
|
1112
|
+
`M. dupont` binds through neither sub-rule. A cleaner long-term design is for N7 to
|
|
1113
|
+
publish the spans it recognises and for N10 to defer to them by index rather than by
|
|
1114
|
+
shape; that is the same mechanism C1-a needed and would let both guards collapse into one.
|
|
1115
|
+
Worth doing once there are `ru` and `fr` initials fixtures to check it against.
|
|
1116
|
+
11. **Decided refusals, verified against Lebedev's live service.** Recorded so they are not
|
|
1117
|
+
re-litigated:
|
|
1118
|
+
- **Digit-group binding** — `100 000` → `100`+U+00A0+`000`. Refused for v1. It is
|
|
1119
|
+
expressible here in principle (it needs no locale list, only a digit-run test), but it
|
|
1120
|
+
fires on every four-plus-digit number in a document and would need its own
|
|
1121
|
+
false-positive analysis against dates, identifiers and code. Not refused on principle;
|
|
1122
|
+
refused as unscoped.
|
|
1123
|
+
- **U+00A0 before the Russian particles `ли`, `же`, `бы`.** Refused as a rule; if a
|
|
1124
|
+
normative source (Мильчин) supports it, the mechanism already exists —
|
|
1125
|
+
`nbsp.afterShortWords` binds the space _after_ a short word, and a particle needs the
|
|
1126
|
+
space _before_ it, so it would need a new locale field and a citation. Not data we have.
|
|
1127
|
+
- **Degree insertion** (`5 C` → `5 °C`) and **currency substitution** (`1 руб.` → `1 ₽`).
|
|
1128
|
+
Refused: both rewrite content rather than normalising typography.
|
|
1129
|
+
12. **The span boundary marker's class membership was split in spec 1.2.0.** Before 1.2.0,
|
|
1130
|
+
[modes.md](modes.md) §3.3's table put the −1 marker in this rule's `OPENISH` and `CLOSEISH`,
|
|
1131
|
+
while all five runtimes put it in neither. `docs/ROADMAP.md` recorded the discrepancy during the
|
|
1132
|
+
Python port and left it for a `spec-guardian` call. Production French content forced the call
|
|
1133
|
+
(§3.3 step 2, §6 rows 6c–6f): taken literally, the table fixes the lost space but loses the
|
|
1134
|
+
narrow space in `<em>non</em> !` and in `[ceci](url) : oui` through step 3. The runtimes'
|
|
1135
|
+
reading fixes neither. The split, `CLOSEISH` yes and `OPENISH` no, is the one reading that
|
|
1136
|
+
keeps both. Whether `OPENISH` should also include the marker is a separate question. It would
|
|
1137
|
+
widen the left-boundary tests of N3, N7, N9 and N10 and N3's following-token guard, and nobody
|
|
1138
|
+
has measured the false-positive profile of that. The answer in 1.2.0 is no, pinned by the
|
|
1139
|
+
fixture `ru-nbsp-span-boundary-not-openish-short-word` (`в <em>Москве</em>` stays unbound);
|
|
1140
|
+
changing it is a spec change.
|
|
1141
|
+
13. **U+2060 has the same font problem and no option (spec 1.3.0).** `ranges` binds a tight range
|
|
1142
|
+
as `JOINER dash JOINER` (ranges.md §3.3.1), and the same production report that asked for
|
|
1143
|
+
`narrowNbsp` also raised the word joiner — not for a missing glyph, since it is zero-width,
|
|
1144
|
+
but for copy/paste and search indexing: a reader who copies `3–5` out of a page gets two
|
|
1145
|
+
invisible characters with it. Nothing was decided here, deliberately. `ranges` is **off by
|
|
1146
|
+
default**, so a caller who has not asked for range conversion never sees a joiner, and one
|
|
1147
|
+
who has can stop asking; that is a narrower situation than U+202F, which a French locale
|
|
1148
|
+
emits with default options. If a `wordJoiner` option is ever added it belongs in
|
|
1149
|
+
`ranges.md`, not here, and it needs its own answer to the question §3.1a answers for this
|
|
1150
|
+
rule: the joiner is what makes a converted range survive a second pass without regrowing a
|
|
1151
|
+
second joiner (ranges.md §3.1's re-entry condition), so removing it removes that anchor and
|
|
1152
|
+
the re-entry argument has to be rebuilt on the bare dash.
|
|
1153
|
+
14. **U+2011 needs no option, and that is worth stating once.** `hyphen` (order 35) exists only to
|
|
1154
|
+
replace a hyphen with U+2011 in the forms a locale lists, so `rules: { hyphen: false }`
|
|
1155
|
+
already expresses "do not emit U+2011" exactly. The production report asked for
|
|
1156
|
+
`nonBreakingHyphen: false`, which suggests the rule table does not make that obvious — a
|
|
1157
|
+
documentation gap, not a missing option.
|