polytypo 1.2.0 → 1.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +648 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,1157 @@
1
+ # Rule: `nbsp`
2
+
3
+ **Order:** 70 (last). **Default:** on. **Modes:** text, html, markdown, yaml.
4
+ **Spec version:** 1.3.0 (0.1.0 for everything except §3.9's `initialBinding` change (0.6.0),
5
+ §3.3's span-boundary paragraphs (1.2.0) and §3.3's character-reference guard (1.3.0), noted
6
+ inline).
7
+
8
+ ---
9
+
10
+ ## 1. Purpose
11
+
12
+ `nbsp` is the rule that makes a line break fall where the language allows it. It converts
13
+ ordinary spaces to U+00A0 (no-break space) or U+202F (narrow no-break space), and inserts one
14
+ of those where a locale requires a space that the author did not type — before French
15
+ `? ! ; :`, inside French guillemets, between a number and its unit, between a Russian
16
+ preposition and the word it governs, between initials and a surname. It runs **last**, after
17
+ every other rule has settled the characters around each candidate space, which is what makes
18
+ its decisions simple: by the time it runs, `spaces` has normalised every ordinary space run
19
+ to length one, `dashes` has fixed the dash forms, and `quotes` has produced the final quote
20
+ glyphs whose inner spacing this rule then owns. It is also the rule with the largest surface
21
+ for oscillation, so almost all of its design effort goes into being a strict no-op on text
22
+ that already carries the right no-break spaces.
23
+
24
+ ---
25
+
26
+ ## 2. Locale data consumed
27
+
28
+ - `nbsp.beforePunctuation` — array of single characters that take U+00A0 before them
29
+ - `nbsp.narrowBeforePunctuation` — array of single characters that take U+202F before them
30
+ - `nbsp.afterShortWords` — array of lowercase words
31
+ - `nbsp.abbreviations` — array of literal multi-token abbreviations written with U+0020
32
+ - `nbsp.beforeUnits` — array of units that bind to a preceding number
33
+ - `nbsp.beforeNumber` — array of abbreviations that bind to a following **number**
34
+ - `nbsp.beforeWord` — array of abbreviations that bind to a following **word**
35
+ - `nbsp.afterSymbols` — array of symbols that bind to a following number
36
+ - `nbsp.initialBinding` — enum: `"none" | "chain" | "single"` (spec 0.6.0; replaces the boolean
37
+ `bindInitials`). See §3.9.
38
+ - `quotes.primary.open`, `quotes.primary.close`, `quotes.primary.innerSpace`
39
+ - `quotes.secondary.open`, `quotes.secondary.close`, `quotes.secondary.innerSpace`
40
+
41
+ **Constraint Q-P (`quotes.md` §2.1), normative on this rule's own data.** Every code point in
42
+ `nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must be a member of `quotes`'
43
+ `CLOSEISH` (`quotes.md` §3.1). This is what makes `quotes`' Lemma B cover N1/N2 as well as N8:
44
+ Lemma B's case 3 needs `m ∈ CLOSEISH` to know that a mark whose `closeRight` moves from `m` to a
45
+ `SPACELIKE` insertion is accepted either way. The shipped lists (`:`, `;`, `!`, `?` in
46
+ `fr`/`fr-CA`, empty elsewhere) satisfy it; `pipeline-idempotency.md` §7 item 4's schema gap is
47
+ closed by `locale.schema.json`'s `nbspClosePunctuation` check. A locale adding a new
48
+ `beforePunctuation`/`narrowBeforePunctuation` entry outside `CLOSEISH` reopens Lemma B and needs
49
+ a `quotes.md` §5 re-derivation, not just a data change.
50
+
51
+ ### 2.1 A locale attests membership; the rule owns the mechanism
52
+
53
+ **Normative, and general — it governs every list-valued field in `locale.schema.json`, not only
54
+ this rule's.**
55
+
56
+ > A `sources` citation must support exactly what the locale file **can express**: which tokens
57
+ > are in which list, which enum value is chosen, which boolean is set. It is **not** required to
58
+ > support the behaviour the rule attaches to that membership.
59
+
60
+ So a locale populating `beforeUnits` needs a source establishing that **these tokens are units
61
+ and are written with a space after the number**. It does **not** need a source saying that space
62
+ is non-breaking, or that it is U+00A0 rather than U+202F, or that the rule converts an existing
63
+ space rather than inserting a missing one. Those are decisions this document has already made,
64
+ once, for every locale.
65
+
66
+ **The decisive argument is that evidentiary burden must track expressive power.** There is no
67
+ field in which a locale can say "bind these units" or "do not bind them" — the schema offers a
68
+ list and nothing else. A locale therefore cannot be *wrong* about the mechanism, cannot vary it,
69
+ and cannot be asked to cite it. Requiring a citation for a claim the file has no way to make is
70
+ not rigour; it is a category error, and its effect is to empty fields that are correctly
71
+ populated.
72
+
73
+ Four supporting points:
74
+
75
+ - **The schema already says so.** `beforeUnits` is defined as "Units and signs that bind to a
76
+ preceding number with U+00A0". The binding is in the *field's* definition, which is where a
77
+ uniform mechanism belongs.
78
+ - **Every other field works this way, and must.** `afterShortWords` enumerates words and N3
79
+ decides what happens to the following space; nobody asks Мильчин to name U+00A0 by code point.
80
+ `abbreviations` lists strings and N4 converts their internal spaces. `initialBinding` (spec
81
+ 0.6.0; the enum that replaced the boolean `bindInitials`) is still mechanism, not a code-point
82
+ claim — no typographic authority names U+00A0 by code point for it, and this document still
83
+ owns exactly what C1/C2 do — but the enum is more citable than the boolean it replaced: Chicago
84
+ states "two or more initials" (`"chain"`, en-US/de-DE/de-CH/ru) and André states a single
85
+ abbreviated first name (`"single"`, fr/fr-CA) as two *distinguishable* claims a locale file can
86
+ now actually choose between, where the boolean could only say on or off. `"none"` remains the
87
+ uncitable negative it always was.
88
+ - **PLAN.md §6 and ARCHITECTURE.md §5.1 already draw this line.** Declarative facts that vary by
89
+ locale go in the file; algorithms that do not vary go in the rule. *Which tokens are units*
90
+ varies by language. *A measurement does not break between its number and its unit* does not —
91
+ it is the same universal principle that justifies the range binding in `dashes.md` §3.3.1, and
92
+ that one is argued from UAX #14 and Chicago/Мильчин centrally rather than per locale.
93
+ - **Authorities describe how text should look and behave when set, not how it is encoded.** This
94
+ spec has drawn that distinction twice already and in the same direction: for Greek
95
+ (`spaces.md` §3.5 — the guide instructs a typist, it does not ask a tool to delete) and for
96
+ vulgar fractions (`symbols.md` §7.6 — the sources describe how a fraction should *look*). A
97
+ source stating "a space separates the number from the unit", plus a rule stating "that space
98
+ must not break", is not the source being over-read.
99
+
100
+ **What this does not license.** The membership claim still needs a source, and the principle
101
+ makes that burden *sharper*, not looser:
102
+
103
+ - Listing a token the source does not attest is unsupported however the mechanism is owned.
104
+ - `fi`'s `abbreviations: ["fil. maist."]` is right on exactly this test: Kotus attests the
105
+ internal shape of **one** abbreviation, which is a membership claim about a single token that
106
+ happens to contain a space. Joining two tokens a source merely shows side by side would be a
107
+ membership claim the source does not make, and would stay unsupported under either reading.
108
+ - A field whose *shape* the source contradicts is still wrong. This principle relocates the
109
+ burden; it does not lower it.
110
+
111
+ **Consequence, recorded so it is not re-litigated per locale.** A `beforeUnits` list resting on a
112
+ source that states the space but not its non-breaking character is **correctly populated**.
113
+ `fi.beforeUnits` is restored on that footing, and `en-GB`'s and `sv`'s measurement units stand on
114
+ the same one. The reverse reading — requiring each locale to attest the mechanism — would leave
115
+ N5 effectively dead outside `en-US` and `ru`, which is a large behavioural consequence to arrive
116
+ at one file at a time; if it is ever taken, it must be taken deliberately and written here.
117
+
118
+ ---
119
+
120
+ **Precondition.** `nbsp.beforePunctuation` and `nbsp.narrowBeforePunctuation` must be
121
+ disjoint. `locale.schema.json` does not enforce this (reported); an implementation that finds
122
+ a character in both must raise `POLYTYPO_MALFORMED_LOCALE_DATA` rather than pick one.
123
+
124
+ ---
125
+
126
+ ## 3. Algorithm
127
+
128
+ Input is a code-point array `cp[0 … n-1]`.
129
+
130
+ ### 3.1 Character classes
131
+
132
+ | Class | Members |
133
+ | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
134
+ | `SP` | U+0020 (space) |
135
+ | `NBSP` | U+00A0 |
136
+ | `NNBSP` | U+202F |
137
+ | `NOBREAK` | `NBSP` ∪ `NNBSP` |
138
+ | `SPACELIKE` | `SP` ∪ `NOBREAK` ∪ U+0009 ∪ U+2007 ∪ U+2008 ∪ U+2009 ∪ U+200A ∪ U+2000–U+2006 ∪ U+205F ∪ U+3000 ∪ every member of `BREAK` |
139
+ | `OTHER-SPACE` | the fixed-width spaces above: U+2000–U+200A, U+205F, U+3000. In `SPACELIKE`, but **never** converted and never a valid "already correct" state |
140
+ | `BREAK` | U+000A, U+000D, U+000B, U+000C, U+0085, U+2028, U+2029 |
141
+ | `DIGIT` | U+0030–U+0039 |
142
+ | `LETTER` | general category `Lu`, `Ll`, `Lt`, `Lm`, `Lo`, `Mn`, `Mc`, `Me` |
143
+ | `UPPER` | general category `Lu` or `Lt` |
144
+ | `ALNUM` | `LETTER` ∪ `DIGIT` |
145
+ | `OPENISH` | U+0028 `(` U+005B `[` U+007B `{` and every locale `open` glyph |
146
+ | `SENTENCE-DASH` | U+2013, U+2014 only. **Not** U+002D and **not** U+2011 — see §3.5 |
147
+ | `DASHISH` | U+002D, U+2011, U+2013, U+2014. U+2011 is included because `hyphen` (order 35) produces it; a class holding only U+002D would make a word-boundary verdict depend on whether `hyphen` had run (`hyphen.md` §3.2) |
148
+ | `CLOSEISH` | U+0029 `)` U+005D `]` U+007D `}` and every locale `close` glyph |
149
+ | `NONE` | index out of range |
150
+
151
+ **Unicode version.** The general categories and case mappings this rule reads are those of the UCD version pinned in `spec/UNICODE` (`17.0`). The pin is normative for the **derived tables**, not for the host runtime — see [pipeline-idempotency.md](pipeline-idempotency.md) §6a, which also specifies the canary fixtures that make the pin detectable.
152
+
153
+ **The inline span boundary marker is in `CLOSEISH` and not in `OPENISH` (spec 1.2.0).** In `html`
154
+ and `markdown` mode the pipeline runs over spans joined by a −1 marker ([modes.md](modes.md) §3.2).
155
+ For this rule the marker is a member of `CLOSEISH` only; [modes.md](modes.md) §3.3 records the
156
+ split and §3.3 steps 2 and 3 below give the reason for each half. The −2 line marker is a member of
157
+ `BREAK`, as it is for every rule.
158
+
159
+ **`NOBREAK` is a member of `SPACELIKE`.** Every word/number/token boundary test in this rule
160
+ uses `SPACELIKE`, never `SP`. This single decision is what makes the rule idempotent: after a
161
+ conversion the boundary that justified it still reads as a boundary.
162
+
163
+ ### 3.1a The narrow target, and the `narrowNbsp` option (spec 1.3.0)
164
+
165
+ **This rule is the only thing in the spec that emits U+202F**, through exactly two sub-rules: N2
166
+ (§3.4), always, and N8 (§3.10), when the pair's `innerSpace` is `"narrow-nbsp"`. **No shipped
167
+ locale sets `"narrow-nbsp"`** — `fr`, the only locale with a non-empty `narrowBeforePunctuation`,
168
+ sets its primary pair to `"nbsp"` — so today the option's whole visible effect is N2's. The N8
169
+ clause is written anyway, because the two must move together the day a locale does set it, and a
170
+ substitution that covered one emitter and not the other would be a defect nobody could see until
171
+ that locale landed. `quotes` places
172
+ the glyphs and leaves the inner space to N8 (quotes.md §5); no other rule writes a no-break space
173
+ of any width.
174
+
175
+ That makes one substitution expressible without touching anything else:
176
+
177
+ > **`NARROW-TARGET`** is U+202F, **unless** the caller passed `narrowNbsp: "nbsp"`, in which case
178
+ > it is U+00A0.
179
+
180
+ **N2 and N8 read `NARROW-TARGET` everywhere they name U+202F** — the state their "already
181
+ correct" branch recognises, the code point they convert a space to, and the code point they
182
+ insert. Nothing else in this document changes.
183
+
184
+ **Why a caller would ask.** U+202F is missing from many common text faces (Manrope and EB
185
+ Garamond among them), so a browser falls back to another face for that one character and the
186
+ advance width changes mid-line; a PDF renderer with no fallback drops the glyph entirely. That
187
+ is a **rendering** fact about the reader's fonts, not a typographic one about the language, which
188
+ is why it is a caller option and not locale data: `fr` still sets a narrow space before `?`, and
189
+ the locale file still says so.
190
+
191
+ **Why it rewrites the target rather than post-processing the output.** A caller can already write
192
+ `transform(...).replaceAll("\u202f", "\u00a0")`, and that is stable **only as long as it always
193
+ runs**: feed its output back through `transform` and N2 sees U+00A0 where its target is U+202F,
194
+ converts it back, and the two steps disagree for ever. Rewriting the target makes the result a
195
+ fixed point by construction — §5's argument is stated per sub-rule **target**, so substituting the
196
+ target carries it over verbatim, with no new case to check.
197
+
198
+ **Validation, and where it happens.** `narrowNbsp` is `"narrow"` or `"nbsp"`; absent means
199
+ `"narrow"`. Any other value raises `POLYTYPO_INVALID_OPTION` (`ARCHITECTURE.md` §4.6, new in spec
200
+ 1.3.0 and general to every option added from 1.3.0 onward). It is checked **immediately after
201
+ `mode` and before `rules`**, so the full order is `mode` → `narrowNbsp` → `rules` → `locale` →
202
+ `dialect`: the two checks that read nothing but the call itself come first, then rule ids, then
203
+ locale data, then the dialect and the parse. The existing pairs are untouched — an unknown rule
204
+ still wins over an unknown locale — and this order is public, tested behaviour like the rest.
205
+
206
+ The check belongs to the **call**, not to this rule: it runs whether or not `nbsp` is enabled, so
207
+ `rules: { nbsp: false }` with a misspelled `narrowNbsp` still raises rather than silently
208
+ accepting a value that would have mattered.
209
+
210
+ **Conformance.** `spec/fixtures/*.json` carries the option as a case-level `narrowNbsp` field,
211
+ passed through to `transform` exactly as `rules` is; omitting it means `"narrow"`, so every case
212
+ written before spec 1.3.0 keeps its meaning. Note what a fixture **cannot** express: the schema's
213
+ own enum admits only the two valid values, so `POLYTYPO_INVALID_OPTION` has no fixture, the same
214
+ way a missing or misspelled `dialect` has none — a fixture with `mode: "markdown"` is required to
215
+ name a valid dialect. Option-validation errors are each runtime's own test, and the order in the
216
+ paragraph above is what those tests assert.
217
+
218
+ **What it does not do.**
219
+
220
+ - It does not change **which** positions take a no-break space. That is locale data (§2) and is
221
+ untouched: the same indices are claimed, by the same sub-rules, under the same guards.
222
+ - It does not change what the rule **reads**. U+202F stays in `NOBREAK` and `SPACELIKE` (§3.1),
223
+ so every guard still treats an authored narrow space as the no-break space it is. The one
224
+ visible consequence is that an authored U+202F at an index N2 or N8 claims is now converted to
225
+ U+00A0, because the rule normalises a claimed index to its target and the target has moved —
226
+ the same normalisation N2 already performs on an authored U+00A0 in the default configuration.
227
+ An authored U+202F anywhere else is left alone (§4).
228
+ - It does not touch **U+2011**, which `hyphen` (order 35) emits. Binding a hyphen is that rule's
229
+ entire job, so `rules: { hyphen: false }` already asks for exactly this and no new option is
230
+ warranted.
231
+ - It does not touch **U+2060**, which `ranges` (order 25) emits when it binds a tight range
232
+ (ranges.md §3.3.1). `ranges` is off by default, so a caller who has not enabled it never sees
233
+ one; a caller who has can turn it off again. See §7.8.
234
+
235
+ ### 3.2 Structure: ten sub-rules, first claim wins
236
+
237
+ The rule is one left-to-right scan that evaluates **ten** independent sub-rules, N1 through N10. (It was eight before `beforeNumber` and `beforeWord` were added as N9 and N10; a reader who stops at N8 loses both, and with them the abbreviation binding for four locales.) Two sub-rules
238
+ may target the same index with different replacements (French `«` followed by a word that
239
+ starts with `«`… or, more realistically, a unit that is also a listed short word). The
240
+ resolution is fixed and total:
241
+
242
+ > Sub-rules are evaluated in the order **N1, N2, N3, N4, N5, N6, N7, N8, N9, N10** as written
243
+ > below.
244
+ > Each produces candidate edits keyed by the index of the space (or insertion point) it
245
+ > claims. **The first sub-rule to claim an index wins; every later edit at that index is
246
+ > discarded.** No index is ever edited twice.
247
+
248
+ This is stated because a map- or set-based implementation that resolves conflicts by
249
+ iteration order will behave differently in Go (ARCHITECTURE.md §4.5).
250
+
251
+ > **First-claim-wins is not sufficient on its own, and relying on it as though it were is a
252
+ > bug.** It resolves a conflict only when both sub-rules actually _emit an edit_ at the index.
253
+ > A sub-rule whose "already correct" branch fires emits nothing and therefore **claims
254
+ > nothing**, silently yielding the index to a lower-priority sub-rule — which then edits it,
255
+ > after which the higher-priority sub-rule is no longer satisfied, and the two alternate on
256
+ > successive runs. That is not a hypothetical: it is exactly how N2 and N8 oscillated on the
257
+ > French input `«?` (§3.10.1). **Sub-rules that target different code points must therefore be
258
+ > made disjoint by construction, not merely ordered.** Where two sub-rules could want
259
+ > different characters at one index, one of them must be given a guard that makes it decline
260
+ > the index outright, so that its verdict does not depend on what the other did.
261
+
262
+ **Why N9 and N10 are last, and in that order.** `nbsp.abbreviations` (N4) holds multi-token
263
+ forms such as `т. д.`; if a locale also lists `т.` in `beforeWord`, the space inside the
264
+ longer, more specific form must be claimed by N4, so N4 must precede N10. `beforeNumber`
265
+ precedes `beforeWord` so that an abbreviation listed in both — `S.` before `12` and before
266
+ `Petersburg` — resolves to the number reading first; the two are then disjoint by their own
267
+ guards (N9 requires a following digit, N10 a following letter) and the ordering is belt and
268
+ braces rather than a real tie-break. `initialBinding` (N7, spec 0.6.0) precedes both because an initial is a
269
+ narrower, structurally-recognised form than a listed abbreviation, and the two agree on the
270
+ replacement (U+00A0) wherever they overlap.
271
+
272
+ ### 3.3 N1 — `beforePunctuation` (U+00A0)
273
+
274
+ For each index `i` such that `cp[i]` is a member of `nbsp.beforePunctuation`:
275
+
276
+ 1. **Run guard.** If `cp[i-1]` is a member of `beforePunctuation` ∪ `narrowBeforePunctuation`,
277
+ skip. Only the first mark of `?!` or `!!!` takes the space.
278
+ 2. **Right-context guard.** Let `after = cp[i+1]` or `NONE`. If `after` is **not** `NONE`,
279
+ not in `SPACELIKE`, not in `CLOSEISH`, **not U+2026**, and not a member of
280
+ `beforePunctuation` ∪ `narrowBeforePunctuation`, skip.
281
+
282
+ U+2026 is in the accepted set because the guard exists to detect a punctuation mark that is
283
+ _inside a token_ — `http://`, `12:30` — and an ellipsis after a question mark is nothing of
284
+ the kind. Without it, French `Vraiment?…` got no narrow space while `Vraiment ?` did, purely
285
+ because `ellipsis` (order 20) had converted the dots before `nbsp` looked. The space belongs
286
+ before the `?` whatever follows it. This is the guard that protects
287
+ `http://example.org` and `12:30` in a locale that lists U+003A — after the colon comes
288
+ `/` or a digit, so nothing happens.
289
+
290
+ **A span boundary marker after the mark is accepted (spec 1.2.0)**, because the marker is in
291
+ `CLOSEISH` (§3.1). This is what makes the `spaces` round trip (`spaces.md` §1) work when an
292
+ inline element closes or a `<br>` follows the mark. `spaces` deletes the U+0020 in
293
+ `<strong>Résistant au gel :</strong> il` and `Et la réglementation ?<br>Oui`, because in both
294
+ the run has content on each side and `right` is in `STRIP-BEFORE`. This sub-rule must then put
295
+ the no-break form back. Before 1.2.0 every runtime left the marker out of `CLOSEISH`, this
296
+ guard refused, and the author's space was lost for good: `gel:</strong>`, `réglementation?<br>`.
297
+ The insertion is at the mark's own index, inside the span, so the edge-growth rule
298
+ ([modes.md](modes.md) §3.4) does not discard it. When the mark is alone in its span
299
+ (`mot<em>!</em>`) the insertion point is the span edge and is still discarded, as before.
300
+
301
+ The cost is stated here so it is not rediscovered as a bug. The guard cannot see past the
302
+ marker, so a colon that ends one span while a digit or a `/` starts the next is accepted as
303
+ sentence punctuation: `fr` `12:<b>30</b>` becomes `12⍽:<b>30</b>`. That needs a time, a URL
304
+ or a ratio split by markup exactly at the colon. The shape this repairs, `**Label :**` and
305
+ `<strong>Label :</strong>`, is the standard French definition-list and FAQ pattern, and the
306
+ failure it repairs deletes a character.
307
+
308
+ 3. **Quote-glyph guard.** If `cp[i-1]` is in `SPACELIKE` and `cp[i-2]` is in `OPENISH`, or if
309
+ `cp[i-1]` is in `OPENISH`, **skip**. The space immediately after an opening quotation glyph
310
+ belongs to `quotes.innerSpace` and is owned by N8; two sub-rules must not both have an
311
+ opinion about it. Without this guard N2 and N8 alternate for ever on the French input `«?`
312
+ — see §3.10.1, which is the defect this guard repairs.
313
+
314
+ **A span boundary marker never triggers this guard**, because it is not in this rule's
315
+ `OPENISH` (§3.1). N8 matches the literal `P.open` glyph, so it never owns the space beside a
316
+ marker, and nothing is left for the guard to protect. Counting the marker as `OPENISH` would
317
+ make this guard decline ordinary French prose that follows an inline element or a link:
318
+ `Il dit <em>non</em> ! Oui` would keep a plain U+0020 before `!`, and the Markdown
319
+ `Voir [ceci](http://x.org) : oui` a plain U+0020 before `:`. The same membership would also
320
+ widen the left-boundary tests of N3, N7, N9 and N10, and N3's following-token guard, at span
321
+ edges. Spec 1.2.0 does not make that change (§7 item 12).
322
+ 4. **Character-reference guard (spec 1.3.0).** If `cp[i]` is U+003B, and the code points to its
323
+ left have the shape of a character reference, **skip**. Concretely: let `j = i-1` and walk
324
+ left while `cp[j]` is an ASCII letter or an ASCII digit, giving a run of `len` code points;
325
+ if `len` is at least 1 and at most 32, and `cp[j]` is U+0023, decrement `j` once more; then
326
+ if `cp[j]` is U+0026, this `;` terminates a reference and the sub-rule emits nothing.
327
+
328
+ The reason is that in `text` mode there is no markup concept at all (`modes.md` §3.1), so a
329
+ French locale, whose `narrowBeforePunctuation` contains `;`, inserted U+202F before the `;`
330
+ that **ends** the reference: `Bonjour&#160;: oui` became `Bonjour&#160·: oui` and
331
+ `Tom &amp; Jerry` became `Tom &amp· Jerry`. The input stopped being what it was — this is the
332
+ one case found where a rule corrupts the input's own syntax rather than merely typesetting
333
+ something it should not have. In `html` mode the same strings were already safe, because
334
+ §3.6 makes a well-formed reference opaque; the guard makes `text` mode stop destroying them
335
+ too, which is the mode people reach for when they have "just a string".
336
+
337
+ The test is the **shape** of a reference, not membership of the HTML named-reference table.
338
+ That table is thousands of entries and would have to be identical in five runtimes, for no
339
+ gain: the guard's job is to decline, and declining on `&notaname;` costs nothing. The bound
340
+ of 32 code points is the length of the longest named reference (31) plus one, and it keeps
341
+ the left walk bounded rather than open-ended.
342
+
343
+ Accepted cost, pinned: a French semicolon that directly follows a token containing `&` with
344
+ no space, `R&D;`, keeps a plain U+0020 or no space at all rather than gaining U+202F. Such a
345
+ token is not French prose, and the alternative is breaking every character reference in the
346
+ language's own documents.
347
+ 5. Let `left = cp[i-1]` or `NONE`.
348
+ - If `left` is `NBSP` → **already correct**, emit nothing.
349
+ - If `left` is `SP` or `NNBSP` → emit an edit replacing `cp[i-1]` with U+00A0.
350
+ - If `left` is in `OTHER-SPACE` → **skip**. A thin or figure space was placed deliberately;
351
+ §4 promises it is left alone, and treating it as content would insert a _second_ space
352
+ beside it.
353
+ - If `left` is `NONE`, in `BREAK`, or U+0009 → skip. A line must not begin with a no-break
354
+ space.
355
+ - Otherwise (`left` is a content character) → emit an edit **inserting** U+00A0 at
356
+ index `i`.
357
+
358
+ ### 3.4 N2 — `narrowBeforePunctuation` (`NARROW-TARGET`)
359
+
360
+ Identical to N1 with `NARROW-TARGET` (§3.1a — U+202F unless the caller substituted it) in place
361
+ of U+00A0: already-correct means `left` **is** `NARROW-TARGET`; a `SP`, or a `NOBREAK` member
362
+ that is not `NARROW-TARGET`, is converted to it; a content character causes an insertion of it.
363
+ In the default configuration that reads exactly as it always has — already-correct is `NNBSP`,
364
+ and `SP` or `NBSP` converts to U+202F. The
365
+ quote-glyph guard, the `OTHER-SPACE` guard and the character-reference guard apply unchanged —
366
+ the last one matters most here, since `;` is in `narrowBeforePunctuation` for `fr` and nowhere in
367
+ `beforePunctuation` for any shipped locale, so N2 is the sub-rule that was destroying references.
368
+
369
+ Because N1 runs first, a character listed in both arrays would be handled by N1 — which is
370
+ why §2 requires the arrays to be disjoint and requires an implementation to fail loudly
371
+ rather than rely on that precedence.
372
+
373
+ ### 3.5 N3 — `afterShortWords` (U+00A0)
374
+
375
+ For each entry `w` in `nbsp.afterShortWords` (a lowercase word of `k` code points):
376
+
377
+ 1. Find each index `a` where `cp[a … a+k-1]` matches `w` under the **first-character-lenient**
378
+ comparison: `cp[a+j] = w[j]` exactly for `j ≥ 1`, and for `j = 0` either `cp[a] = w[0]` or
379
+ `cp[a]` is the Unicode **simple uppercase mapping** of `w[0]`. Simple uppercase mapping is
380
+ taken from the Unicode Character Database, is a plain code-point→code-point table, and is
381
+ **not** a locale-sensitive operation — this satisfies ARCHITECTURE.md §4.4 (`i` → `I`
382
+ always, never `İ`). No other case variation matches: `IN` does not match `in`.
383
+ 2. **Left boundary.** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
384
+ **`SENTENCE-DASH`**. Otherwise skip.
385
+
386
+ **A hyphen — U+002D or U+2011 — fails this test, and that is the point.** A hyphen marks an
387
+ _intra-word_ position by construction, so the token after one is not a free-standing word.
388
+ Without this, `из-за дождя` binds twice over: `hyphen` (order 35) produces `из‑за`, and then
389
+ this sub-rule matches the listed short word `за` at the position after the U+2011 and emits
390
+ `из‑за` + U+00A0 + `дождя`. But `за` there is not a preposition — it is the tail of a single
391
+ compound preposition — so the binding is typographically wrong on ordinary Russian prose.
392
+ The same reasoning covers every form `hyphen` produces: compound tails (`под` in `из-под`),
393
+ prefix bodies (`что` in `кое-что`) and suffix bodies (`то`, `таки`) are all word-internal and
394
+ none of them should ever be treated as a word start here.
395
+ `SENTENCE-DASH` keeps the case the boundary was written for: an em or en dash genuinely does
396
+ introduce a new phrase, so `— в Москве` still binds `в`.
397
+
398
+ 3. **Right boundary.** `cp[a+k]` must exist and be `SP` or `NBSP`. If it is `NBSP`, emit
399
+ nothing (already correct). If it is anything else — a letter, a punctuation mark, a line
400
+ break, `NNBSP` — skip. (A `NNBSP` there was put by another sub-rule with better
401
+ information.)
402
+ 4. **Following-token guard.** `cp[a+k+1]` must exist and be in `ALNUM` or `OPENISH`.
403
+ A short word must not be bound to a punctuation mark or to the end of a line.
404
+ 5. Emit an edit replacing `cp[a+k]` with U+00A0.
405
+
406
+ **Longest match wins, and there is no backtracking.** When two entries match at the same index
407
+ `a`, only the longest is considered; entries are unique (`uniqueItems` in the schema), so this
408
+ is a total order. **If the longest matching entry fails a guard, no shorter entry is retried at
409
+ that index** — the sub-rule emits nothing there and the scan moves on. This is normative for
410
+ every list-driven sub-rule in this document (N3, N4, N5, N6, N9, N10) and for `hyphen`
411
+ (`hyphen.md` §3.4). Two reasons: a shorter entry succeeding where a longer one failed is almost
412
+ always wrong (if `mm` was rejected because the following code point is a letter, `m` is wrong
413
+ for the same reason), and backtracking makes the result depend on the interaction between list
414
+ contents and guard order in a way that is tedious to reproduce identically in five runtimes.
415
+
416
+ ### 3.6 N4 — `abbreviations` (U+00A0 for every internal space)
417
+
418
+ For each entry `s` in `nbsp.abbreviations`, of `k` code points:
419
+
420
+ 1. Find each index `a` where `cp[a … a+k-1]` matches `s` under **space-lenient exact**
421
+ comparison: for every position `j`, either `cp[a+j] = s[j]`, or `s[j]` is `SP` and
422
+ `cp[a+j]` is in `NOBREAK`. Every non-space code point must match exactly — no case
423
+ leniency at all here, because `z. B.` and `Z. B.` are different strings and the locale
424
+ file is expected to list what it means.
425
+ 2. **Boundaries.** `cp[a-1]` must not be in `ALNUM`; `cp[a+k]` must not be in `ALNUM`.
426
+ (`z. B.` inside `Xz. B.y` is not a match.)
427
+ 3. For every position `j` with `s[j]` = `SP`: if `cp[a+j]` is already `NBSP`, emit nothing for
428
+ that position; otherwise emit an edit replacing `cp[a+j]` with U+00A0.
429
+
430
+ Longest match wins at a given `a`, with no backtracking (§3.5).
431
+
432
+ ### 3.7 N5 — `beforeUnits` (U+00A0, conversion only)
433
+
434
+ For each entry `u` in `nbsp.beforeUnits`, of `k` code points:
435
+
436
+ 1. Find each index `a` where `cp[a … a+k-1]` matches `u` exactly (code-point equality; no
437
+ case leniency — `Kg` is not `kg`).
438
+ 2. **Right boundary.** `cp[a+k]` must not be in `ALNUM`. This is what stops the unit `kg`
439
+ from matching inside `kgf`, and `m` from matching inside `metres`.
440
+ 3. **Left side.** `cp[a-1]` must be `SP` or `NBSP`; nothing else qualifies. If it is `NBSP`,
441
+ emit nothing. **If there is no space at all, this sub-rule does nothing** — see §7.2, it
442
+ never inserts.
443
+ 4. `cp[a-2]` must be in `DIGIT`. Let `Lrun` be the maximal `DIGIT` run ending at `a-2`,
444
+ starting at index `b`. `cp[b-1]` must not be in `LETTER` (so `H2 O` does not bind).
445
+ 5. Emit an edit replacing `cp[a-1]` with U+00A0.
446
+
447
+ Longest match wins at a given `a`, with no backtracking (§3.5), so a locale listing both `m`
448
+ and `mm` binds `5 mm` correctly and does not fall back to `m` when the `mm` match is rejected.
449
+
450
+ ### 3.8 N6 — `afterSymbols` (U+00A0, conversion only)
451
+
452
+ For each entry `y` in `nbsp.afterSymbols`, of `k` code points:
453
+
454
+ 1. Find each index `a` where `cp[a … a+k-1]` matches `y` exactly.
455
+ 2. **Left boundary.** `cp[a-1]` must not be in `ALNUM`.
456
+ 3. **Right side.** `cp[a+k]` must be `SP` or `NBSP`. If `NBSP`, emit nothing. If there is no
457
+ space, do nothing (never insert).
458
+ 4. `cp[a+k+1]` must be in `DIGIT`.
459
+ 5. Emit an edit replacing `cp[a+k]` with U+00A0.
460
+
461
+ ### 3.9 N7 — `initialBinding` (U+00A0), only when `nbsp.initialBinding` is not `"none"` (spec 0.6.0)
462
+
463
+ Define an **initial** at index `p` as: `cp[p]` in `UPPER`, `cp[p+1]` = U+002E (.), and
464
+ `cp[p-1]` either `NONE`, in `SPACELIKE`, or in `OPENISH`. (One letter, one full stop, at a
465
+ token start.)
466
+
467
+ Two clauses, both operating on a single space at index `q`:
468
+
469
+ - **C1 — initial on the left.** If `cp[q]` is `SP` or `NBSP`, and there is an initial ending
470
+ at `q-1` (i.e. `cp[q-1]` = U+002E and `cp[q-2]` is an initial's letter satisfying the
471
+ definition above), and `cp[q+1]` is in `UPPER`, **and the mode condition below holds**, then
472
+ bind: if `cp[q]` is `NBSP` emit nothing, else emit an edit replacing `cp[q]` with U+00A0.
473
+
474
+ **Mode condition (spec 0.6.0).** Split `cp[q+1]`'s role into two shapes:
475
+ - **Between-initials** — `cp[q+1]` is *itself* an initial (i.e. `q+1` also satisfies the
476
+ `initial` definition above: `cp[q+1]` in `UPPER`, `cp[q+2]` = U+002E). This shape always
477
+ binds, in every mode except `"none"` — it is unambiguous by construction: two adjacent
478
+ initial-shaped tokens are Chicago's own literal case ("between two or more initials").
479
+ - **Initial-to-word** — `cp[q+1]` is `UPPER` but not itself an initial (no U+002E follows at
480
+ `q+2`), i.e. the space leads into a candidate surname or ordinary word. This shape's
481
+ eligibility depends on `nbsp.initialBinding`:
482
+ - `"single"` — binds unconditionally, the historical shape (this is what `fr`/`fr-CA` need
483
+ for `N. Bourbaki`/`M. Dupont` — André §5.1.3, cited below).
484
+ - `"chain"` — binds **only if** the initial ending at `q-1` (letter at `p = q-2`) is itself
485
+ immediately preceded by another initial: `cp[p-1]` is `SP` or `NBSP`, `cp[p-2]` = U+002E,
486
+ and `p-3` satisfies the `initial` definition. In other words: bind the space leading out
487
+ of a chain into a following word only once the chain already has two or more members —
488
+ Chicago's own "two or more initials" wording, applied to this shape specifically rather
489
+ than to N7 as a whole. `E. B. White`: the space before `White` binds, because `B.` is
490
+ itself preceded by `E.`. A lone `A. Smith`, or `...take the top N. It runs...`, does not:
491
+ neither `A.` nor `N.` is preceded by another initial, so this shape declines. The declined
492
+ case is a **deliberate false negative**, not a bug: this document has no way to tell a
493
+ genuine single-initial name from an ordinary sentence boundary using only local structure
494
+ (§7.9a), and `"chain"` resolves that ambiguity by declining rather than guessing, on the
495
+ same footing P4 in `dashes.md` §3.4 declines rather than guessing at a range.
496
+ - `"none"` — N7 does not run at all (checked once, before either clause).
497
+
498
+ This covers both spaces of `А. С. Пушкин` under `"chain"` (`ru`): the first is
499
+ between-initials (`cp[q+1]` is `С`, itself an initial), the second is initial-to-word but
500
+ eligible because `С.` is preceded by `А.`.
501
+
502
+ **Guard C1-a — lower-case abbreviation tail.** C1 does **not** fire (in any mode, on either
503
+ shape) if, letting `p = q-2` be the initial's letter, `cp[p-1]` is in `SPACELIKE`, `cp[p-2]`
504
+ is U+002E, and `cp[p-3]` is in `LETTER` but **not** in `UPPER`. In other words: an uppercase
505
+ letter plus a dot that is itself preceded by a _lower-case_ letter plus a dot is the second
506
+ token of an abbreviation, not an initial.
507
+
508
+ This is not a refinement, it is a false-positive repair on real German data. With the
509
+ shipped `de-DE` file, `z. B. Berlin` produced `z.⍽B.⍽Berlin`: N4 correctly bound the space
510
+ inside `z. B.`, and then C1 read `B.` as an initial and `Berlin` as a surname and bound the
511
+ space after it too. A false positive on ordinary prose is the ship-blocking class for this
512
+ project (PLAN.md §4, the M4 gate), so this needed a guard rather than a note. §6 example 14
513
+ was right and the rule was wrong.
514
+ `А. С. Пушкин` is unaffected: `cp[p-3]` there is `А`, which is in `UPPER`.
515
+
516
+ - **C2 — two initials on the right.** If `cp[q]` is `SP` or `NBSP`, `cp[q-1]` is in `LETTER`,
517
+ and the code points starting at `q+1` form **two consecutive initials** (`X. Y.`, i.e. an
518
+ initial at `q+1`, a `SP` or `NBSP` at `q+3`, and an initial at `q+4`), then bind `cp[q]` the
519
+ same way, **in every mode except `"none"`** — C2 already requires two initials by
520
+ construction (Chicago's own "two or more"), so it needs no mode split.
521
+ This covers `Пушкин А. С.` — the surname-first order — without binding every
522
+ `word` + `Capital letter + period` pair, which would fire on ordinary sentences.
523
+
524
+ C2 binds only the space before the initials group; the space _inside_ the group is bound by
525
+ C1 (its `cp[q+1]` is the second initial's letter, which is `UPPER`).
526
+
527
+ ### 3.10 N8 — `quotes.innerSpace`
528
+
529
+ For each of the two quote pairs `P` ∈ {`quotes.primary`, `quotes.secondary`} with
530
+ `P.innerSpace ≠ "none"`, let `target` be U+00A0 if `P.innerSpace = "nbsp"` and `NARROW-TARGET`
531
+ (§3.1a — U+202F unless the caller substituted it) if `P.innerSpace = "narrow-nbsp"`.
532
+
533
+ **Sidedness precondition.** If `P.open` and `P.close` are the same code point, this sub-rule
534
+ does nothing for that pair, and an implementation must not guess. See §7.5.
535
+
536
+ Otherwise:
537
+
538
+ - **After an opening glyph.** For each index `o` with `cp[o]` = `P.open`:
539
+ - if `cp[o+1]` is `target` → emit nothing;
540
+ - else if `cp[o+1]` is `SP` or in `NOBREAK` → emit an edit replacing `cp[o+1]` with
541
+ `target`;
542
+ - else if `cp[o+1]` is `NONE` or in `BREAK` → skip;
543
+ - else → emit an edit inserting `target` at index `o+1`.
544
+ - **Before a closing glyph.** For each index `c` with `cp[c]` = `P.close`, symmetrically on
545
+ `cp[c-1]`, skipping when `cp[c-1]` is `NONE` or in `BREAK`.
546
+
547
+ **When `innerSpace` is `"none"` this sub-rule does nothing at all, and in particular it does
548
+ not _remove_ a space the author put inside the quotation marks.** This is normative, not an
549
+ oversight: `de-CH` declares `innerSpace: "none"` and `« hallo »` therefore survives as written.
550
+
551
+ The reasoning is that `innerSpace` describes what the pipeline **inserts**, not what it
552
+ enforces. Removing a space is a deletion of author content, and this spec deletes a U+0020 in
553
+ exactly three narrowly-argued places, all of them in `spaces` (§3.2 step 5) and none of them
554
+ next to a quotation mark — `STRIP-BEFORE` excludes quotation glyphs precisely so that their
555
+ inner spacing stays a single rule's business. A tool that silently closed up `« hallo »` in a
556
+ Swiss text would be making an editorial judgement about the author's copy rather than a
557
+ typographic correction to it, and the M4 criterion is zero false positives, not maximum
558
+ conformity. The counter-argument — that `«hallo»` is the only Swiss-conforming form, so leaving
559
+ the spaces leaves the document wrong — is real, and §7.4 records it as the operator's to settle.
560
+
561
+ #### 3.10.1 The N1/N2 oscillation this rule was part of
562
+
563
+ `fr` declares `quotes.primary.innerSpace = "nbsp"` (target U+00A0) and lists `?` in
564
+ `nbsp.narrowBeforePunctuation` (target U+202F). On the input `«?` both N2 and N8 want the
565
+ single index between the two characters, and they want different code points:
566
+
567
+ ```
568
+ «? → «⍽? (N8 inserts U+00A0) → «⍹? (N2 converts to U+202F)
569
+ → «⍽? (N8 converts back) → …
570
+ ```
571
+
572
+ First-claim-wins did not resolve it, for the reason set out in §3.2: on the second run N2's
573
+ "already correct" branch fires, emits nothing, and therefore claims nothing — leaving N8 free
574
+ to act on an index N2 still has an opinion about.
575
+
576
+ **Repaired in N1/N2**, not here: the space beside an opening quotation glyph is `innerSpace`,
577
+ full stop, and N1/N2 now decline it outright (§3.3 step 4). N8 is unchanged. The alternative —
578
+ having N8 accept any `NOBREAK` beside a quote glyph as already correct — was rejected because
579
+ it would also make N8 unable to correct a genuinely wrong inner space anywhere else, trading a
580
+ narrow conflict for a broad loss of function.
581
+
582
+ Note the general shape, because it will recur: **the sub-rule that yields must be the one that
583
+ can state its exception locally.** N1/N2 can say "not next to an opening quote glyph" using
584
+ `OPENISH`, which they already have. N8 cannot say "not where a punctuation rule wants a
585
+ different width" without importing both punctuation lists.
586
+
587
+ ### 3.11 N9 — `beforeNumber` (U+00A0, conversion only)
588
+
589
+ An abbreviation in `nbsp.beforeNumber` binds forward to a number: `Nr. 5`, `S. 12`,
590
+ `art. 237`, `§§ 12`. Schema minimum length is 2 code points, so a bare initial cannot be
591
+ listed.
592
+
593
+ For each entry `s` in `nbsp.beforeNumber`, of `k` code points:
594
+
595
+ 1. Find each index `a` where `cp[a … a+k-1]` matches `s` **exactly**, code point for code
596
+ point. No case leniency at all: `Nr.` is not `nr.`, and a locale that wants both lists
597
+ both. (This differs from N3, where the schema explicitly asks for first-character
598
+ leniency; here it does not, and inventing it would make `S.` match `s.` and bind sentence
599
+ fragments.)
600
+ 2. **Left boundary.** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
601
+ `SENTENCE-DASH`. Otherwise skip. This is stronger than "not `ALNUM`" and it is what stops
602
+ `S.` matching inside `Fig.S. 3`. A hyphen fails it, for the reason given in §3.5 step 2 —
603
+ an abbreviation cannot begin immediately after an intra-word hyphen.
604
+ 3. **The separator.** `cp[a+k]` must be `SP` or `NBSP`; anything else — a letter, a digit, a
605
+ line terminator, `NNBSP`, a tab — and the sub-rule does nothing. If it is `NBSP`, emit
606
+ nothing: already correct.
607
+ **Exactly one separator.** If `cp[a+k+1]` is also in `SPACELIKE`, skip.
608
+ 4. **"A following number" means:** `cp[a+k+1]` is in `DIGIT`. That is the whole definition —
609
+ one code point, tested for membership in U+0030–U+0039. It deliberately does **not**
610
+ require the digit run to be of any particular length, to be followed by anything in
611
+ particular, or to be a "plausible" number: a digit immediately after the separator is
612
+ sufficient evidence, and no false positive of consequence exists for it (`Nr. 5x` is still
613
+ a number after `Nr.`).
614
+ 5. Emit an edit replacing `cp[a+k]` with U+00A0.
615
+
616
+ **Never inserts.** `Nr.5` stays `Nr.5`, for the same reason N5 never inserts: adding a space
617
+ is a content change, not a typographic one.
618
+
619
+ **N9 has no equivalent of N10's guard G-D**, so a one-letter abbreviation such as `S.` binds
620
+ forward here even though it is structurally indistinguishable from an initial. That is
621
+ deliberate and safe: N9 fires only when a `DIGIT` follows, and N7 clause C1 fires only when an
622
+ `UPPER` code point follows, so the two are disjoint by their own right-hand tests and cannot
623
+ both claim a space. `S. 12` is a page reference and only N9 can see it; `S. Petrov` is an
624
+ initial and only N7 can. No guard is needed to keep them apart.
625
+
626
+ Longest match wins at a given `a`, with no backtracking (§3.5).
627
+
628
+ ### 3.12 N10 — `beforeWord` (U+00A0, conversion only)
629
+
630
+ An abbreviation in `nbsp.beforeWord` binds forward to a word: `г. Москва`, `ул. Ленина`,
631
+ `M. Dupont`, `Mme Hugo`, `Dr. Schmidt`. **This is the highest-false-positive sub-rule in the
632
+ spec** and its guards are correspondingly heavier — an abbreviation that also occurs at the
633
+ end of a sentence will bind across the sentence boundary, and no purely syntactic test can
634
+ separate the two readings. Locale files must list only forms whose binding is normative
635
+ (the schema description says so); this document adds the structural guards.
636
+
637
+ For each entry `s` in `nbsp.beforeWord`, of `k` code points:
638
+
639
+ 1. Match `cp[a … a+k-1]` against `s` **exactly**, code point for code point. No case leniency
640
+ in either direction. `г.` matches only lowercase `г.`; `M.` matches only uppercase `M.`;
641
+ `Mme` matches only that capitalisation. A locale wanting both cases lists both, which is
642
+ cheap and explicit and keeps the spec free of any case-folding table on this path.
643
+ 2. **Left boundary (G-L).** `cp[a-1]` must be `NONE`, in `SPACELIKE`, in `OPENISH`, or in
644
+ `SENTENCE-DASH`. Otherwise skip — a hyphen fails it, per §3.5 step 2.
645
+ 3. **The separator (G-S).** `cp[a+k]` must be `SP` or `NBSP`. `NBSP` → emit nothing, already
646
+ correct. Anything else → skip. If `cp[a+k+1]` is also in `SPACELIKE`, skip.
647
+ 4. **"A following word" means (G-W):** `cp[a+k+1]` is in `LETTER`. Not `ALNUM` — a digit there
648
+ is N9's business, and N9 has already run, so a `beforeWord` entry that is also a
649
+ `beforeNumber` entry never double-claims. Not `OPENISH`, not a quotation glyph, not a
650
+ dash: `г. «Москва»` is left alone, because the binding target is then a quotation and the
651
+ line-break risk this sub-rule exists to remove is not present in the same way.
652
+ 5. **Initial-collision guard (G-D).** If `nbsp.initialBinding` is not `"none"`, **and** `s` is
653
+ exactly two code points, **and** the first is in `UPPER`, **and** the second is U+002E, skip.
654
+ An _uppercase_ letter plus a dot is structurally indistinguishable from an initial, and in a
655
+ locale where N7 is active, N7 owns that shape with better evidence (it inspects what follows
656
+ for a second initial or a surname). G-D's own condition does not distinguish `"chain"` from
657
+ `"single"` — it only asks whether N7 is active at all — because G-D's job is routing (which
658
+ sub-rule owns this shape), not deciding whether the space actually ends up bound; that
659
+ decision is C1's alone (§3.9), and under `"chain"` a routed-to-N7 space can still end up
660
+ declined if no chain is confirmed.
661
+
662
+ All three conditions matter:
663
+ - **`UPPER`** — without it, a lower-case two-code-point entry such as `ул.` would be inert
664
+ in an `initialBinding`-active locale, and `ул. Ленина` would not bind. A lower-case letter
665
+ plus a dot is not a plausible initial in any orthography this spec covers, so the guard
666
+ must not capture one. (An earlier revision justified this clause with `г. Москва`, which is
667
+ **not** an example of it: `ru.json` does not list `г.` in `beforeWord` at all — see §6 row
668
+ 13a.)
669
+ - **`initialBinding !== "none"`** — in a locale where N7 is switched off, nothing else claims
670
+ the shape and there is no collision to avoid.
671
+ - **exactly two code points** — `Mme`, `ул.`, `art.` are longer and are never initials.
672
+
673
+ Consequence for French, stated plainly because an earlier revision of this document got the
674
+ underlying fact wrong: **`fr` has `initialBinding: "single"`** (spec 0.6.0; cited to Jacques
675
+ André §5.1.3 for `N. Bourbaki`). So a listed `M.` _is_ inert in N10, and `M. Dupont` binds
676
+ through N7 clause C1 instead — which reaches the same U+00A0 by a different route, eligible
677
+ under `"single"`'s unconditional initial-to-word shape (§3.9). The residual gap is that C1
678
+ requires an `UPPER` code point after the space, so `M. dupont` with a lower-case surname does
679
+ not bind at all. That is an acceptable miss; French surnames are capitalised.
680
+
681
+ 6. **Line-boundary guard (G-B).** If `cp[a+k]` is in `BREAK`, or if `cp[a+k+1]` is `NONE`,
682
+ skip. Covered by G-S/G-W but stated separately because it is the guard that keeps a
683
+ no-break space off the end of a text unit.
684
+ 7. Emit an edit replacing `cp[a+k]` with U+00A0.
685
+
686
+ **Never inserts.** Longest match wins at a given `a`, with no backtracking (§3.5).
687
+
688
+ **What is deliberately _not_ guarded.** There is no sentence-boundary test. In
689
+ `Это было в 1990 г. Москва тогда была другой`, the form `г.` ends the sentence and `Москва`
690
+ begins the next, yet every guard above passes and the two are bound. Distinguishing that from
691
+ `г. Москва` ("the city of Moscow") requires knowing whether `г.` means _год_ or _город_,
692
+ which is a lexical question the spec cannot answer from code points. The mitigation is
693
+ entirely in the locale data — list a form in `beforeWord` only when the wrong reading is rare
694
+ or harmless — and the failure mode is mild: a no-break space where a break was permitted, not
695
+ a changed character. This is recorded as §7.9 and must be reviewed at the M4 gate.
696
+
697
+ ### 3.13 Insertions and indices
698
+
699
+ Sub-rules N1, N2 and N8 can _insert_ a code point. Because rules produce edits and the
700
+ pipeline applies them (ARCHITECTURE.md §7.1), every index above refers to the **input**
701
+ array. An insertion is an edit with an empty replaced span at a given index. Two insertions
702
+ at the same index cannot occur: the first-claim-wins rule of §3.2 forbids it.
703
+
704
+ ---
705
+
706
+ ## 4. Must not touch
707
+
708
+ **Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
709
+ a guarantee of `transform` as a whole — or **[R]** — true of this rule in isolation but capable
710
+ of being falsified by another rule, which is then named.
711
+
712
+ - **[P] A line terminator.** No sub-rule inserts, deletes or crosses one; every guard that could
713
+ reach a `BREAK` skips instead.
714
+ - **[P] The start or end of a text unit.** A no-break space is never inserted at index 0, never
715
+ after the last code point, and never adjacent to a `BREAK`. In `html` mode a text node
716
+ frequently begins or ends at a tag boundary, and a leading U+00A0 there is a visible
717
+ rendering change.
718
+ - **[P] A tab.** U+0009 is `SPACELIKE` for boundary purposes but is never converted.
719
+ - **[R] A space that another sub-rule already claimed** (§3.2).
720
+ - **[P] `http://`, `12:30`, `1:2`** in a locale that lists U+003A in `beforePunctuation`. N1
721
+ guard 2.
722
+ - **[P] The second and later marks of `?!`, `!!!`, `?..`** — N1/N2 guard 1.
723
+ - **[P] A unit with no space before it.** `5km` and `5%` stay exactly as typed; N5 converts, never
724
+ inserts (§7.2).
725
+ - **[P] A number written as `H2O`, `A4`, `MP3`.** N5 step 4's letter guard.
726
+ - **[P] An existing U+00A0 or U+202F that is already the right character.** Every sub-rule's
727
+ "already correct" branch emits nothing, so a correctly-typeset document round-trips
728
+ byte-identically. This is the property PLAN.md §3.4 exists for.
729
+ - **[P] An existing U+2007, U+2009, U+200A, U+2060 or U+FEFF.** None is in `NOBREAK`; none is
730
+ converted. If an author placed a thin space, it stays.
731
+ - **[P] The quotation marks themselves.** N8 only touches the space beside them.
732
+ - **[P] Anything inside a skipped region.** Handled by the mode adapter.
733
+
734
+ ---
735
+
736
+ ## 5. Idempotency argument
737
+
738
+ Let `T` be the rule and consider `y = T(x)`.
739
+
740
+ **Every sub-rule has an explicit "already correct" branch that emits nothing**, and the
741
+ target of each sub-rule is exactly the state that branch recognises:
742
+
743
+ | Sub-rule | Post-state at the claimed index | Recognised as already correct by |
744
+ | -------- | ----------------------------------------------------- | ------------------------------------------ |
745
+ | N1 | `NBSP` immediately left of the mark | §3.3 step 4, first bullet |
746
+ | N2 | `NARROW-TARGET` immediately left of the mark | §3.4 |
747
+ | N3 | `NBSP` after the short word | §3.5 step 3 |
748
+ | N4 | `NBSP` at every internal position of the abbreviation | §3.6 step 1 (space-lenient match) + step 3 |
749
+ | N5 | `NBSP` before the unit | §3.7 step 3 |
750
+ | N6 | `NBSP` after the symbol | §3.8 step 3 |
751
+ | N7 | `NBSP` at the bound space | §3.9 |
752
+ | N8 | `target` beside the quote glyph | §3.10 |
753
+ | N9 | `NBSP` between the abbreviation and the number | §3.11 step 3 |
754
+ | N10 | `NBSP` between the abbreviation and the word | §3.12 step 3 |
755
+
756
+ So it suffices to show that on the second run **the same sub-rule claims the same index** —
757
+ i.e. that no guard's verdict flips because of an edit made on the first run.
758
+
759
+ Guards test three kinds of thing:
760
+
761
+ 0. **The `narrowNbsp` substitution (§3.1a) needs no separate argument.** Every row of the table
762
+ above names a sub-rule's **target**, not a literal code point, and N2's and N8's targets are
763
+ the only ones the option moves. Substituting `NARROW-TARGET` therefore carries the whole
764
+ argument over unchanged: the post-state each branch recognises moves with the character each
765
+ branch writes, which is precisely what post-processing the output could not do (§3.1a). The
766
+ substitution also cannot create a **new** conflict between two sub-rules, because it only ever
767
+ makes N2's and N8's targets **equal to** N1's, never different from it, and no sub-rule's claim
768
+ depends on what another sub-rule's target is — the quote-glyph guard (§3.3 step 3) decides
769
+ ownership of the index beside a quotation glyph by position, whatever character either
770
+ sub-rule would have written there.
771
+ 1. **Membership in `SPACELIKE`.** Every conversion `SP → NBSP`, `SP → NNBSP`,
772
+ `NNBSP → NBSP`, `NBSP → NNBSP` stays inside `SPACELIKE` (§3.1). Every insertion adds a
773
+ `NOBREAK`, which is in `SPACELIKE`. So every boundary test that passed on run 1 passes on
774
+ run 2, and every one that failed still fails — **provided** boundary tests never use `SP`
775
+ specifically. They do not: §3.1 states this and every guard above is written against
776
+ `SPACELIKE`, `ALNUM`, `UPPER`, `DIGIT`, `OPENISH`, `CLOSEISH`, `BREAK` or `NONE`, none of
777
+ which gains or loses a member under a `SP`/`NOBREAK` conversion.
778
+ 2. **Membership in `ALNUM`, `DIGIT`, `UPPER`, `LETTER`, `OPENISH`, `CLOSEISH`.** No sub-rule
779
+ ever edits a code point in any of these classes; every edit replaces or inserts a space
780
+ character. Unchanged.
781
+ 3. **Literal matching (N3, N4, N5, N6, N9, N10).** N4 is space-lenient by construction, so a matched
782
+ abbreviation still matches after its internal spaces become `NBSP`. N3, N5, N6, N9 and N10 match
783
+ only the word/unit/symbol/abbreviation itself, never the adjacent space, so their matches
784
+ are untouched. N9 and N10 additionally require the separator to be `SP` or `NBSP` and
785
+ treat `NBSP` as "already correct", so their second-run verdict is "emit nothing" at the
786
+ very index they claimed on the first run. The _side conditions_ that read the adjacent space accept both `SP` and `NBSP`
787
+ and distinguish them only to decide "convert" versus "emit nothing".
788
+
789
+ Two specific insertion cases need checking because they change lengths:
790
+
791
+ - **N1/N2 insertion.** Run 1 turns `mot!` into `mot` `NNBSP` `!`. On run 2, `left` of the `!`
792
+ is `NNBSP` → "already correct" → no edit. Note the insertion did **not** create a new
793
+ candidate: the inserted character is a space, and no sub-rule's _match_ is a space (N4
794
+ matches a literal containing spaces, but the literal must also match its non-space code
795
+ points, and inserting one space next to a punctuation mark cannot complete an abbreviation
796
+ match that failed before, because the abbreviation's non-space code points are unchanged).
797
+ - **N8 insertion.** Run 1 turns `«mot»` into `«` `NNBSP` `mot` `NNBSP` `»`. On run 2 both
798
+ positions hit the "already `target`" branch. Additionally the inserted `NNBSP` becomes the
799
+ `left` neighbour of nothing punctuation-like, and the `right` neighbour of `«`, which is not
800
+ a candidate for any other sub-rule.
801
+
802
+ Finally, the **first-claim-wins** ordering of §3.2 is deterministic and depends only on the
803
+ sub-rule index, not on the array contents, so the same sub-rule wins on both runs.
804
+
805
+ Hence `T(T(x)) = T(x)`.
806
+
807
+ **What had to be fixed.** Four things, all of them the difference between a rule that works
808
+ and a rule that oscillates:
809
+
810
+ 1. **`NOBREAK` had to be inside `SPACELIKE`.** The naive version writes boundary tests
811
+ against U+0020, so after run 1 converts the space, run 2 no longer sees a word boundary,
812
+ and a _different_ sub-rule (or none) claims the position. Symptom: N3 and N5 fight over
813
+ `5 km` in a locale that lists `km` as a unit and has a short word ending in a way that
814
+ overlaps.
815
+ 2. **N4 had to match space-leniently.** Matching `z. B.` literally means the converted
816
+ `z.` `NBSP` `B.` no longer matches, which is harmless on its own — but it means the
817
+ abbreviation is invisible to sub-rule ordering on run 2 and a lower-priority sub-rule
818
+ (N3, say, on the word `z`) can claim the same index with a different verdict. Making the
819
+ match space-lenient keeps N4 in control of its own indices forever.
820
+ 3. **N5 and N6 had to be conversion-only.** The naive version inserts a space before a unit,
821
+ turning `5km` into `5 km` — a content change, not a typographic one — and, worse, in a
822
+ locale listing `%` it turns `50%` into `50 %`, which is wrong in English and right in
823
+ French, i.e. it is a locale decision that `beforeUnits` alone cannot express. Restricting
824
+ to conversion makes the rule safe and shifts the question to §7.2 where it belongs.
825
+ 4. **The conflict policy had to be total and index-based.** "Whichever sub-rule runs first in
826
+ the loop" is not a specification; it is a Go map iteration bug waiting to happen (§3.2).
827
+
828
+ ---
829
+
830
+ ### Composition obligation
831
+
832
+ Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₈** and runs last,
833
+ so nothing has to preserve _its_ invariant — but it must preserve all seven others, and it is
834
+ the rule with the largest emission surface in the pipeline. It is also the rule that committed
835
+ defect family 2.
836
+
837
+ **What this rule emits.** U+00A0 and U+202F only, either replacing a single space-like code
838
+ point in place or **inserted** at one of three kinds of site: before a listed punctuation
839
+ character (N1, N2), after an opening quote glyph, and before a closing quote glyph (N8). It
840
+ never emits U+0020, never emits a letter, digit, dot, dash or quotation mark, and never deletes
841
+ anything.
842
+
843
+ **Against `I₁` (`spaces`).** Discharged, and this is why the rule never emits U+0020: U+00A0
844
+ and U+202F are `CONTENT` to `spaces`, which touches only U+0020. Converting a U+0020 to a
845
+ no-break space can only _remove_ a potential `spaces` edit, never create one. An insertion adds
846
+ a `CONTENT` code point, which cannot lengthen a U+0020 run or place a U+0020 in a stripping
847
+ position.
848
+
849
+ **Against `I₂` (`ellipsis`).** Discharged: no `DOTLIKE` code point is emitted, and insertions
850
+ only push code points apart, never together.
851
+
852
+ **Against `I₃` (`dashes`).** **This is the obligation the rule failed, twice**, and it is now
853
+ discharged structurally rather than by argument.
854
+
855
+ N8 inserting a `quotes.innerSpace` beside a quoted hyphen turned `«-»` into `«⍽-⍽»`, which
856
+ `dashes` read as a spaced parenthetical dash on the following pass — defect family 2. The
857
+ repair made a right-hand no-break space not count as dash spacing, which left the _asymmetric_
858
+ shape reachable: `«–␣"` gains an inner U+00A0 from N8, and a left-hand no-break space still
859
+ counted, so the token became symmetric and was promoted from en to em — defect (d),
860
+ `dashes.md` §5.2. Both were caused by this sub-rule's insertions, and both repairs were
861
+ attempts to specify which no-break space counts as spacing on which side.
862
+
863
+ **The discharge no longer depends on that.** `E(nbsp) = { U+00A0, U+202F }` — those are the
864
+ only code points this rule can emit or insert, in any sub-rule, under any locale data. `dashes`
865
+ now treats **both** as making an adjacent token inert (`dashes.md` §3.2 step 3), so no emission
866
+ of this rule can create a dash token, change one's spacing verdict, or revive one `dashes`
867
+ declined. That is condition **CO-S** of `pipeline-idempotency.md` §5.1a, and it holds for every
868
+ input rather than for the witnesses anyone happened to test.
869
+
870
+ The Russian direction — N1/N2 promoting the U+0020 before an em dash to U+00A0 — is covered by
871
+ the same statement: `dashes` declines the promoted token at its symmetry guard and emits
872
+ nothing, so the output `Москва⍽— столица` is stable.
873
+
874
+ **The repair stayed in `dashes` rather than moving here**, for the layering reason in
875
+ `pipeline-idempotency.md` §4: `order.json` gives `dashes` `"localeData": ["dash"]` and no access
876
+ to `quotes`, so it cannot recognise a guillemet; and requiring this rule to suppress an
877
+ insertion that might create a dash token would mean encoding `dashes`' admissibility rules in a
878
+ second place, which is how the two rules drifted apart in the first place.
879
+
880
+ **Against `I₄` (`hyphen`).** Discharged, but only just, and the reason is worth stating. A
881
+ listed hyphen form's boundary guards ask whether the neighbouring code point is in `WORDISH`.
882
+ An insertion beside such a form would change that neighbour from whatever it was to a no-break
883
+ space, which is not `WORDISH` — so the guard could go from _reject_ to _accept_ and a form
884
+ could start matching that did not before. That would be an `I₄` violation. It is unreachable
885
+ because every insertion site puts the new space next to a listed punctuation character or a
886
+ quote glyph, never between two letters, and a form whose neighbour is punctuation or a quote
887
+ glyph was already accepted. If a future sub-rule inserts between two letters, this obligation
888
+ must be re-derived.
889
+
890
+ **Against `I₅`, `I₆`, `I₇` (`quotes`, `apostrophe`, `symbols`).** Discharged by case analysis
891
+ over the neighbour tests; the `quotes` half is the one that matters and is written out in
892
+ `quotes.md` §5.6, which shows that neither insertion site can _add_ a capability to a surviving
893
+ straight mark. `apostrophe`'s case ladder reads `ALNUM`, `LETTER`, `DIGIT`, `SPACELIKE`,
894
+ `OPENISH`, `OPENQUOTE` and `CLOSEISH`; an inserted no-break space is `SPACELIKE`, and the only cases it
895
+ could newly satisfy require a `SPACELIKE` **left** neighbour, which is case 4 (leading elision)
896
+ — and case 4 also requires an `ALNUM` right neighbour, which an insertion cannot create.
897
+ `symbols` reads `ALNUM`, `DIGIT` and symmetric spacing; a no-break space is already accepted as
898
+ spacing there (`symbols.md` §3.3 step 1), and converting U+0020 to U+00A0 leaves `lsp`/`rsp` unchanged.
899
+
900
+ ---
901
+
902
+ ## 6. Worked examples
903
+
904
+ `␣` = U+0020, `⍽` = U+00A0, `⍹` = U+202F, `⟶` = no change. Each block states the fields it uses,
905
+ and **those fields are quoted from the shipped locale file** rather than assumed. That
906
+ distinction is not pedantry: while this preamble said the locale files did not exist yet, rows
907
+ row 13a drifted into asserting behaviour the shipped `ru.json` does not produce — and a
908
+ neighbouring row asserted an `ru` `beforeNumber` binding that has neither data nor a source
909
+ behind it, and has since been deleted rather than softened. Both survived several reviews
910
+ because nothing obliged anyone to check. A row here is a claim
911
+ about the engine **and** about a file in `spec/locales/`, and both halves have to hold.
912
+
913
+ ### `fr` — `narrowBeforePunctuation: ["?","!",";"]`, `beforePunctuation: [":"]`, primary `« »` with `innerSpace: "nbsp"`
914
+
915
+ | # | Input | Output | Why |
916
+ | --- | ----------------------------------- | ------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
917
+ | 1 | `Bonjour!` | `Bonjour⍹!` | N2 insertion; the ordinary space, if any, was already removed by `spaces` |
918
+ | 2 | `Bonjour⍹!` | ⟶ | N2 "already correct" — this is the round-trip case |
919
+ | 3 | `Il a dit «mot».` | `Il a dit «⍽mot⍽».` | N8 inserts on both inner edges. **U+00A0, not U+202F**: `fr.json` sets the primary pair's `innerSpace` to `"nbsp"` — corrected in spec 1.3.0, having claimed the narrow space since this table was written |
920
+ | 4 | `Il a dit «⍹mot⍹».` | `Il a dit «⍽mot⍽».` | and therefore **not** already correct: N8 converts an authored narrow space to its own target. The row claimed a fixed point on the same mistake |
921
+ | 5 | `Voir http://example.org: la suite` | `Voir http://example.org⍽: la suite` | N1 guard 2 rejects the colon in `http://` (next code point is `/`) and accepts the sentence colon (next is a space) |
922
+ | 6 | `Vraiment?!` | `Vraiment⍹?!` | N1/N2 guard 1: only the first mark takes a space |
923
+ | 6a | `Vraiment?…` | `Vraiment⍹?…` | U+2026 is accepted right context (§3.3 step 2). Previously the narrow space was omitted here but not in `Vraiment␣?`, purely because `ellipsis` had already run |
924
+ | 6b | `Voir␣../docs` | ⟶ | the French instance of the `spaces` lone-dot defect; see `spaces.md` §3.4 |
925
+ | 6c | `<strong>gel␣:</strong>␣il` (`html`) | `<strong>gel⍽:</strong>␣il` | **spec 1.2.0.** `spaces` deletes the U+0020; the mark's right neighbour is the span boundary marker, which is in `CLOSEISH`, so N1 inserts U+00A0 at the mark's index, inside the span. Previously produced `gel:</strong>` |
926
+ | 6d | `réglementation␣?<br>Oui` (`html`) | `réglementation⍹?<br>Oui` | same, with N2. `<br>` leaves no line terminator in the gap, so the marker is −1, not −2 |
927
+ | 6e | `<em>non</em>␣!␣Oui` (`html`) | `<em>non</em>⍹!␣Oui` | `spaces` leaves the run alone (its `left` is a span edge), and N2 converts it. The marker two places left of the mark is not `OPENISH`, so step 3 does not decline |
928
+ | 6f | `12:<b>30</b>` (`html`) | `12⍽:<b>30</b>` | **the accepted cost of 6c** — step 2 cannot see past the marker |
929
+
930
+ ### `ru` — `afterShortWords: ["в","и","на",…]`, `abbreviations: ["т. д.","и т. п."]`, `initialBinding: "chain"`, `afterSymbols: ["№","§"]`
931
+
932
+ | # | Input | Output | Why |
933
+ | --- | ----------------------- | ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
934
+ | 7 | `Он живёт в␣Москве` | `Он живёт в⍽Москве` | N3: left boundary is a space, right is `SP`, following token is a letter |
935
+ | 8 | `Он живёт в⍽Москве` | ⟶ | N3 already correct |
936
+ | 9 | `и␣т.␣д.` | `и⍽т.⍽д.` | the leading `и` by N3, the internal space by N4 (`т. д.`) |
937
+ | 10 | `А.␣С.␣Пушкин` | `А.⍽С.⍽Пушкин` | N7 clause C1 twice |
938
+ | 11 | `Пушкин␣А.␣С.` | `Пушкин⍽А.⍽С.` | first space by C2 (two initials follow), second by C1 |
939
+ | 12 | `см. №␣5` | `см. №⍽5` | N6 |
940
+ | 13 | `Иван␣пошёл␣домой` | ⟶ | no short word, no abbreviation, no unit, no initial |
941
+ | 13a | `г.␣Москва, ул.␣Ленина` | `г.␣Москва, ул.⍽Ленина` | N10 binds **only `ул.`**. `ru.json`'s `beforeWord` is `["ул.","пл."]`; `г.` is not in it — it is in `beforeUnits`, where it binds **backwards** to a preceding year (`1990⍽г.`), which is the actual Russian convention. Row checked against the shipped file, not against a hypothetical one |
942
+ | 13c | `г.⍽Москва` | ⟶ | N10 "already correct" |
943
+ | 13d | `г.Москва` | ⟶ | N9/N10 never insert |
944
+ | 13e | `г. «Москва»` | ⟶ | N10 guard G-W: the following code point is a quotation glyph, not a letter |
945
+ | 13f | `из-за␣дождя` | `из‑за␣дождя` | `hyphen` binds the compound; N3 does **not** then bind `за`, because its left neighbour is U+2011 and a hyphen fails the left boundary (§3.5 step 2). Previously produced `из‑за⍽дождя`, a false positive on ordinary prose |
946
+ | 13g | `из-под␣стола` | `из‑под␣стола` | same shape with the other listed compound |
947
+ | 13h | `—␣в␣Москве` | `—␣в⍽Москве` | `SENTENCE-DASH` still opens a phrase, so a genuine preposition after an em dash binds normally |
948
+ | 13i | `в␣<em>Москве</em>` (`html`) | ⟶ | **spec 1.2.0.** N3's following-token guard asks for `ALNUM` or `OPENISH`, and for this rule the span boundary marker is in neither (§3.1). The space stays U+0020. Whether it should bind is §7 item 12 |
949
+
950
+ ### `de-DE` — `abbreviations: ["z. B.","d. h."]`, `beforeUnits: ["%","km","°C"]`
951
+
952
+ | # | Input | Output | Why |
953
+ | --- | ----------------------- | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
954
+ | 14 | `z.␣B.␣Berlin` | `z.⍽B.␣Berlin` | N4 binds the space inside the abbreviation. The space after `B.` is **not** bound: guard C1-a sees that the `B.` is preceded by the lower-case `z.` and declines. Previously produced `z.⍽B.⍽Berlin`, a false positive on ordinary prose |
955
+ | 15 | `z.⍽B. Berlin` | ⟶ | N4 space-lenient match, already correct |
956
+ | 16 | `Es sind 20␣km bis 5␣%` | `Es sind 20⍽km bis 5⍽%` | N5 twice |
957
+ | 17 | `Es sind 20km` | ⟶ | N5 never inserts (§7.2) |
958
+ | 18 | `H2␣O ist kein Wert` | ⟶ | N5 step 4: `2` is preceded by the letter `H` |
959
+ | 19 | `siehe Nr.␣5 und S.␣12` | `siehe Nr.⍽5 und S.⍽12` | N9 twice |
960
+ | 20 | `siehe nr.␣5` | ⟶ | N9 matches exactly; `nr.` is not `Nr.` |
961
+ | 21 | `Bonjour&#160;:␣oui` (`fr`) | ⟶ | **spec 1.3.0.** N2's character-reference guard: the `;` closes `&#160;`, so no U+202F is inserted and the reference survives. Before 1.3.0 this produced `Bonjour&#160·:␣oui` |
962
+ | 22 | `Tom␣&amp;␣Jerry` (`fr`) | ⟶ | **spec 1.3.0.** Same guard on a named reference |
963
+ | 23 | `Oui␣;␣non` (`fr`) | `Oui·;␣non` | The guard is not a blanket refusal of `;` — an ordinary semicolon still takes U+202F |
964
+
965
+ #### `narrowNbsp: "nbsp"` (§3.1a, spec 1.3.0)
966
+
967
+ Every row `fr`, and every row measured. `⍽` = U+00A0, `⍹` = U+202F.
968
+
969
+ | # | Input | Default (`narrow`) | `narrowNbsp: "nbsp"` | Why |
970
+ | --- | ---------------------------- | ------------------------------------- | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
971
+ | 31 | `Un délai ?` | `Un délai⍹?` | `Un délai⍽?` | N2's target moves; the index it claims, and the guard that let it claim it, do not |
972
+ | 32 | `Oui⍹?` | ⟶ | `Oui⍽?` | an authored narrow space at a claimed index is normalised to the target, as an authored U+00A0 is in the default configuration |
973
+ | 33 | `Il a dit : « oui » ; puis ?` | `Il a dit⍽: «⍽oui⍽»⍹; puis⍹?` | `Il a dit⍽: «⍽oui⍽»⍽; puis⍽?` | N1 (colon) and N8 (`fr`'s pair, `innerSpace: "nbsp"`) already wrote U+00A0 and do not move; only N2 does |
974
+ | 34 | `12:30 et http://x ; oui` | `12:30 et http://x⍹; oui` | `12:30 et http://x⍽; oui` | the option changes what is written, never what is read: N1's right-context guard still protects the time and the URL |
975
+
976
+ ### `fr` — `beforeWord: ["M.","Mme","Mlle"]`
977
+
978
+ | # | Input | Output | Why |
979
+ | --- | ----------------------- | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
980
+ | 21 | `M.␣Dupont et Mme␣Hugo` | `M.⍽Dupont et Mme⍽Hugo` | `Mme` binds via N10 (three code points, G-D does not apply). `M.` is inert in N10 because `fr` has `initialBinding: "single"` and G-D fires — it binds via **N7 C1** instead, eligible under `"single"`'s unconditional initial-to-word shape, reaching the same U+00A0 |
981
+ | 22 | `M.␣dupont` | ⟶ | lower-case surname: N7 C1 needs `UPPER` after the space, and N10 is inert for `M.`. A documented gap, §7.10 |
982
+ | 23 | `«?` | `«⍽?` | N1/N2 decline the index (quote-glyph guard, §3.3 step 3) and N8 owns it, inserting the `innerSpace` target. Previously this oscillated between U+00A0 and U+202F for ever — §3.10.1 |
983
+ | 24 | `«⍽?` | ⟶ | the fixed point of case 23 |
984
+ | 25 | `mot␣?` | `mot⍹?` | the ordinary case: the code point before the space is a letter, so N2 applies normally |
985
+
986
+ ### `el` — every list empty, `initialBinding: "none"`, `quotes.innerSpace: "none"`
987
+
988
+ The first locale for which this rule is a **total no-op**, in the same provable sense `hyphen` is
989
+ a no-op for a locale with empty lists. It is worth a block of its own because the claim needs
990
+ _both_ halves of `order.json`'s `"localeData": ["nbsp", "quotes"]` to hold, and only one of them
991
+ is visible in the `nbsp` object: the eight lists being empty disables N1–N7 and N9–N10, and
992
+ **`quotes.innerSpace: "none"` is what additionally disables N8**. A locale with empty lists but a
993
+ non-`none` `innerSpace` would still edit — `fr` case 23 is exactly that shape — so "all lists
994
+ empty" alone does not license the claim.
995
+
996
+ Greek is also the case that shows why N1/N2 being empty is a positive finding rather than an
997
+ unfilled field. The Greek source denies the French space-before-punctuation pattern **by name**
998
+ («πράγμα που συμβαίνει, π.χ., στα γαλλικά»), so `spaces` strips the space at order 10 and nothing
999
+ here puts one back. That is a cited decision, not a default.
1000
+
1001
+ | # | Input | Output | Why |
1002
+ | --- | -------------- | ------ | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
1003
+ | 26 | `Τι κάνεις;` | ⟶ | `beforePunctuation` and `narrowBeforePunctuation` are empty, so N1/N2 have no member to match. Contrast `fr` case 25, where the same shape gains U+202F |
1004
+ | 27 | `10,5␣%␣φέτος` | ⟶ | `beforeUnits` is empty, so N5 is inert and the ordinary space survives as U+0020 |
1005
+ | 28 | `«καλημέρα»` | ⟶ | `innerSpace: "none"`, so N8 inserts nothing. This is the half of the no-op claim that lives in the `quotes` object rather than the `nbsp` one |
1006
+
1007
+ Cases 2, 6b, 8, 13, 13c, 13d, 13e, 13i, 15, 17, 18, 20, 22, 24, 26, 27 and 28 are "no change"
1008
+ cases. (Case 4 left that list in spec 1.3.0, when the row was corrected from a fixed point to a
1009
+ conversion; the §3.1a rows are numbered 31-34 and are outside it.)
1010
+
1011
+ ---
1012
+
1013
+ ## 7. Open questions
1014
+
1015
+ 1. **The schema does not enforce that `beforePunctuation` and `narrowBeforePunctuation` are
1016
+ disjoint**, nor that a character appears in only one of `beforeUnits` / `afterSymbols`.
1017
+ §2 requires an implementation to fail loudly; a `$defs`-level constraint or a CI check
1018
+ would be better. Reported.
1019
+ 2. **N5/N6 never insert a space.** `5km` stays `5km` and `50%` stays `50%`. For German,
1020
+ Duden requires a space before `%` and before a unit, so `50%` is arguably _wrong_ input
1021
+ that polytypo declines to fix. Inserting is a content change and the schema has no flag to
1022
+ authorise it per locale (something like `"insertBeforeUnits": true`). Needs an operator
1023
+ decision; I chose the conservative branch because of the zero-false-positive ship
1024
+ criterion.
1025
+ 3. **`afterShortWords` matching is "case-insensitive only for the first character"** per the
1026
+ schema description, which I have implemented as the Unicode **simple** uppercase mapping.
1027
+ That is a data table, not a locale operation, so it satisfies §4.4. **The spec pins a Unicode version —
1028
+ `spec/UNICODE` contains `17.0` — and the pin is now normative for the derived tables**
1029
+ (`pipeline-idempotency.md` §6a). §3.1 cites it, as do the other four rules that read general
1030
+ categories. An earlier revision of this item said no pin existed and then that it had no
1031
+ reader; both were true when written and neither is now.
1032
+ 4. **`innerSpace: "none"` does not remove an existing space inside quotation marks** — now
1033
+ stated normatively in §3.10 rather than left open, with the deletion-is-not-correction
1034
+ argument. `de-CH` `« hallo »` and `en-US` `“ hello ”` both survive as typed. The live
1035
+ question is only whether the operator wants the opposite: a locale declaring `"none"` could
1036
+ be read as asserting that no inner space is permissible, in which case removal would be a
1037
+ correction and not a mutilation. I did not take that reading, because it is the only place
1038
+ in the spec where a locale field would license a deletion. **Operator decision**; the
1039
+ fixtures currently pin non-removal.
1040
+ 5. **A quote pair whose `open` equals `close` cannot have an `innerSpace`.** Finnish and
1041
+ Swedish use U+201D on both sides, and their `innerSpace` is expected to be `"none"`, so the
1042
+ case is currently vacuous — but the schema permits `{"open":"”","close":"”","innerSpace":
1043
+ "nbsp"}`, which this rule cannot implement (it cannot tell an opener from a closer without
1044
+ the pairing information that only `quotes` has, and `quotes` does not pass state to `nbsp`).
1045
+ §3.10 makes it a documented no-op. Either the schema should forbid it or the pipeline
1046
+ should let `quotes` hand its pair positions to `nbsp` — the latter breaks the "rules are
1047
+ independent" property and I did not propose it unilaterally. Reported.
1048
+ 6. **`initialBinding` clause C2 requires two consecutive initials, in every mode.** `Пушкин А.`
1049
+ (one initial) is not bound. Restricting it this way avoids binding every `word` + `Capital.`
1050
+ pair, but it is a guess about Russian practice and needs checking against Мильчин.
1051
+ 7. _(Settled.)_ `nbsp.beforeNumber` and `nbsp.beforeWord` were added to the schema and are
1052
+ specified as N9 (§3.11) and N10 (§3.12). `Nr. 5`, `S. 12`, `art. 237`, `г. Москва`,
1053
+ `ул. Ленина`, `M. Dupont` are all now expressible.
1054
+ 8. _(Settled.)_ U+2011 is produced by the `hyphen` rule at order 35 — see
1055
+ [hyphen.md](hyphen.md). `nbsp` still never produces it, which is correct: a hyphen is not a
1056
+ space and binding it is a different rule with different locale data. U+2060 (word joiner)
1057
+ is produced by `dashes` (order 30) around a tight range dash — see
1058
+ [dashes.md](dashes.md) §3.3.1. `nbsp` still never produces it, and still never converts an
1059
+ author's own U+2060.
1060
+ 9. **N10 has no sentence-boundary guard — but with the shipped locale data the risk is not
1061
+ reachable, and an earlier revision of this item was wrong to call it "the single most likely
1062
+ `nbsp` false positive at the M4 gate".** That misdirected the review, which is worse than
1063
+ saying nothing: it pointed the gate at an input the engine cannot produce.
1064
+
1065
+ The mechanism is real in the abstract. If a locale listed `г.` in `beforeWord`, then
1066
+ `в 1990 г. Москва…` would bind across a sentence boundary, because nothing distinguishes
1067
+ *год* from *город* syntactically. **`ru.json` does not list it.** `beforeWord` is
1068
+ `["ул.","пл."]` — two forms that are unambiguously nouns and cannot end a sentence in the
1069
+ relevant sense — and `г.` is instead in `beforeUnits`, binding **backwards** to a preceding
1070
+ year. That is a better answer than any guard this rule could have carried: it encodes the
1071
+ reading (*год*) that actually collides, and it binds in the direction that reading requires,
1072
+ so the forward-binding ambiguity never arises.
1073
+
1074
+ What remains open is the general shape, not this instance: **a locale author may still put an
1075
+ ambiguous form in `beforeWord`**, and this rule will bind it across a sentence boundary
1076
+ without complaint. The mitigation is locale-data curation and the schema description already
1077
+ says so ("Riskier than beforeNumber — list only forms whose binding is normative"). The
1078
+ lesson worth keeping is the one this entry got wrong: **a prose claim about what a rule does
1079
+ to a language must be checked against the shipped locale file**, because the rule alone does
1080
+ not determine it.
1081
+ 9a. **N7's own sentence-boundary miss (spec 0.6.0) — resolved for `"chain"`, deliberately left
1082
+ open for `"single"`.** A fresh M4 dogfooding pass surfaced the C1 sibling of item 9's N10
1083
+ miss: `"...take the top N. It runs..."` bound `N.` to `It` across a sentence boundary,
1084
+ because the old boolean `bindInitials` let C1's initial-to-word shape fire on any lone
1085
+ initial next to any uppercase-starting word, with no check that a genuine name — as opposed
1086
+ to an ordinary sentence ending in a single capital letter — was actually present. `"chain"`
1087
+ closes this for en-US/de-DE/de-CH/ru by requiring a confirmed sequence of two or more
1088
+ initials before that shape binds (Chicago's own "two or more initials" wording, applied
1089
+ literally rather than only for turning N7 on at all) — see §3.9's mode condition.
1090
+
1091
+ **This does not close the miss for `"single"` (`fr`/`fr-CA`).** `N. Bourbaki` and
1092
+ `M. Dupont` — both cited, both canonical fixtures — are structurally the identical shape to
1093
+ the English witness: one lone initial, one following capitalized word, no preceding initial
1094
+ and no further initial on the right. No local, non-lexical structural signal distinguishes
1095
+ them (both my English witness and both French citations happen to sit at the very start of
1096
+ their sentence, which is suggestive but not load-bearing evidence, and is not implemented as
1097
+ a guard here — it is not reliable enough to specify without lexical or genuine
1098
+ sentence-segmentation data, which this rule does not have and is not mine to add
1099
+ unilaterally). A French sentence ending in a lone initial immediately followed by a new
1100
+ sentence starting with a capitalized word would misfire under `"single"` exactly as the
1101
+ English witness did before `"chain"` existed. This is a **known, accepted limitation of
1102
+ `"single"` mode**, not a regression: the alternative (applying `"chain"`'s restriction to
1103
+ French too) would break `N. Bourbaki`/`M. Dupont`, which André's own citation supports and
1104
+ which have been canonical fixtures since before this item was written. Revisiting it needs
1105
+ either a genuinely new structural signal (none is known) or a locale-specific operator
1106
+ decision to accept `"single"`'s narrower false-positive-vs-false-negative trade for French,
1107
+ which is the decision already made, recorded here rather than left implicit.
1108
+ 10. **G-D is conditional on `initialBinding !== "none"`**, so a `beforeWord` entry's behaviour depends on
1109
+ an unrelated field. That is the only coupling of its kind in the rule and it is mildly
1110
+ unpleasant, but every alternative is worse: an absolute guard makes `г.` inert, and no
1111
+ guard at all lets N7 and N10 both claim `M. Dupont`. The residual cost is §6 case 22 —
1112
+ `M. dupont` binds through neither sub-rule. A cleaner long-term design is for N7 to
1113
+ publish the spans it recognises and for N10 to defer to them by index rather than by
1114
+ shape; that is the same mechanism C1-a needed and would let both guards collapse into one.
1115
+ Worth doing once there are `ru` and `fr` initials fixtures to check it against.
1116
+ 11. **Decided refusals, verified against Lebedev's live service.** Recorded so they are not
1117
+ re-litigated:
1118
+ - **Digit-group binding** — `100 000` → `100`+U+00A0+`000`. Refused for v1. It is
1119
+ expressible here in principle (it needs no locale list, only a digit-run test), but it
1120
+ fires on every four-plus-digit number in a document and would need its own
1121
+ false-positive analysis against dates, identifiers and code. Not refused on principle;
1122
+ refused as unscoped.
1123
+ - **U+00A0 before the Russian particles `ли`, `же`, `бы`.** Refused as a rule; if a
1124
+ normative source (Мильчин) supports it, the mechanism already exists —
1125
+ `nbsp.afterShortWords` binds the space _after_ a short word, and a particle needs the
1126
+ space _before_ it, so it would need a new locale field and a citation. Not data we have.
1127
+ - **Degree insertion** (`5 C` → `5 °C`) and **currency substitution** (`1 руб.` → `1 ₽`).
1128
+ Refused: both rewrite content rather than normalising typography.
1129
+ 12. **The span boundary marker's class membership was split in spec 1.2.0.** Before 1.2.0,
1130
+ [modes.md](modes.md) §3.3's table put the −1 marker in this rule's `OPENISH` and `CLOSEISH`,
1131
+ while all five runtimes put it in neither. `docs/ROADMAP.md` recorded the discrepancy during the
1132
+ Python port and left it for a `spec-guardian` call. Production French content forced the call
1133
+ (§3.3 step 2, §6 rows 6c–6f): taken literally, the table fixes the lost space but loses the
1134
+ narrow space in `<em>non</em> !` and in `[ceci](url) : oui` through step 3. The runtimes'
1135
+ reading fixes neither. The split, `CLOSEISH` yes and `OPENISH` no, is the one reading that
1136
+ keeps both. Whether `OPENISH` should also include the marker is a separate question. It would
1137
+ widen the left-boundary tests of N3, N7, N9 and N10 and N3's following-token guard, and nobody
1138
+ has measured the false-positive profile of that. The answer in 1.2.0 is no, pinned by the
1139
+ fixture `ru-nbsp-span-boundary-not-openish-short-word` (`в <em>Москве</em>` stays unbound);
1140
+ changing it is a spec change.
1141
+ 13. **U+2060 has the same font problem and no option (spec 1.3.0).** `ranges` binds a tight range
1142
+ as `JOINER dash JOINER` (ranges.md §3.3.1), and the same production report that asked for
1143
+ `narrowNbsp` also raised the word joiner — not for a missing glyph, since it is zero-width,
1144
+ but for copy/paste and search indexing: a reader who copies `3⁠–⁠5` out of a page gets two
1145
+ invisible characters with it. Nothing was decided here, deliberately. `ranges` is **off by
1146
+ default**, so a caller who has not asked for range conversion never sees a joiner, and one
1147
+ who has can stop asking; that is a narrower situation than U+202F, which a French locale
1148
+ emits with default options. If a `wordJoiner` option is ever added it belongs in
1149
+ `ranges.md`, not here, and it needs its own answer to the question §3.1a answers for this
1150
+ rule: the joiner is what makes a converted range survive a second pass without regrowing a
1151
+ second joiner (ranges.md §3.1's re-entry condition), so removing it removes that anchor and
1152
+ the re-entry argument has to be rebuilt on the bare dash.
1153
+ 14. **U+2011 needs no option, and that is worth stating once.** `hyphen` (order 35) exists only to
1154
+ replace a hyphen with U+2011 in the forms a locale lists, so `rules: { hyphen: false }`
1155
+ already expresses "do not emit U+2011" exactly. The production report asked for
1156
+ `nonBreakingHyphen: false`, which suggests the rule table does not make that obvious — a
1157
+ documentation gap, not a missing option.