polytypo 1.2.0 → 1.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +648 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,649 @@
1
+ # Rule: `spaces`
2
+
3
+ **Order:** 10 (first). **Default:** on. **Modes:** text, html, markdown, yaml.
4
+ **Spec version:** 1.2.0 (0.2.0 for everything except §3.6's mouth side and the clause it adds to
5
+ §3.2 step 5, noted inline, and §3.4's word-start clause, added in 1.2.0).
6
+
7
+ ---
8
+
9
+ ## 1. Purpose
10
+
11
+ `spaces` performs the space hygiene that every later rule depends on: it collapses a run of
12
+ two or more ordinary spaces down to one, and it deletes an ordinary space that sits between
13
+ a word and a following punctuation mark or on the inner edge of a bracket pair. It runs
14
+ first so that the rules after it see one canonical spacing form and never have to consider
15
+ `"a , b"` alongside `"a, b"`. It is deliberately the most conservative rule in the
16
+ pipeline: it only ever _removes_ U+0020 (space), it never inserts anything, it never touches
17
+ any other whitespace character, and it never touches whitespace that carries structural
18
+ meaning (indentation, Markdown hard line breaks, line terminators). Where a locale genuinely
19
+ wants a space before punctuation — French `?` `!` `;` `:` — this rule still removes the
20
+ ordinary space, and the `nbsp` rule (order 70) re-inserts the correct no-break form; that
21
+ round trip is intentional and is what makes the French output deterministic regardless of
22
+ how the author typed it.
23
+
24
+ ---
25
+
26
+ ## 2. Locale data consumed
27
+
28
+ **None.** `order.json` declares `"localeData": []` for this rule. The behaviour of `spaces`
29
+ is identical in every locale. Locale-dependent spacing is entirely the responsibility of
30
+ `nbsp`.
31
+
32
+ ---
33
+
34
+ ## 3. Algorithm
35
+
36
+ The input is a code-point array `cp[0 … n-1]`. Indices below are code-point indices
37
+ (ARCHITECTURE.md §4.2). The rule emits edits; the pipeline applies them.
38
+
39
+ ### 3.1 Character classes
40
+
41
+ | Class | Members |
42
+ | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
43
+ | `SPACE` | U+0020 (space) — **the only** character this rule ever removes |
44
+ | `BREAK` | U+000A (LF), U+000D (CR), U+000B (VT), U+000C (FF), U+0085 (NEL), U+2028 (LS), U+2029 (PS) |
45
+ | `PROTECTED-SPACE` | U+0009 (tab), U+00A0 (nbsp), U+202F (narrow nbsp), U+2007 (figure space), U+2008 (punctuation space), U+2009 (thin space), U+200A (hair space), U+2000–U+2006, U+205F (medium math space), U+3000 (ideographic space), U+200B (zero-width space), U+FEFF |
46
+ | `CONTENT` | any code point that is **not** in `SPACE` and **not** in `BREAK`. Note that every member of `PROTECTED-SPACE` is `CONTENT` for the purposes of this rule — it bounds a space run and is never itself modified. |
47
+ | `STRIP-BEFORE` | exactly six code points: U+002C (,) U+002E (.) U+003B (;) U+003A (:) U+0021 (!) U+003F (?). **U+2026 is deliberately not a member** — see §3.4 |
48
+ | `DOTLIKE` | U+002E (.) and U+2026 (…) |
49
+ | `OPEN-BRACKET` | U+0028 `(` U+005B `[` U+007B `{` |
50
+ | `CLOSE-BRACKET` | U+0029 `)` U+005D `]` U+007D `}` |
51
+ | `EMOTICON-EYE` | U+003A (:), U+003B (;) — a subset of `STRIP-BEFORE`, not a new code point this rule reads outside it. See §3.6 |
52
+ | `EMOTICON-NOSE` | U+002D (-), U+005E (^). Optional. See §3.6 |
53
+ | `EMOTICON-MOUTH` | U+0028 `(`, U+0029 `)`, U+005B `[`, U+005D `]`, U+0044 `D`, U+0064 `d`, U+0050 `P`, U+0070 `p`, U+004F `O`, U+006F `o`, U+002F `/`, U+005C `\`, U+007C `|`, U+002A `*`. U+0028 and U+005B are `OPEN-BRACKET` members too, and U+0029 and U+005D are `CLOSE-BRACKET` members — the overlap is what §3.6's mouth side exists for. See §3.6 |
54
+
55
+ `STRIP-BEFORE` contains only these **six** code points. It deliberately excludes closing
56
+ quotation marks and guillemets: their inner spacing is `nbsp`'s business (`quotes.innerSpace`),
57
+ and stripping there would fight with it.
58
+
59
+ ### 3.2 Scan
60
+
61
+ 1. Set `i = 0`.
62
+ 2. If `cp[i]` is not `SPACE`, emit nothing, set `i = i + 1`, repeat from 2. Terminate when
63
+ `i = n`.
64
+ 3. `cp[i]` is `SPACE`. Find the maximal run: let `s = i`; let `e` be the smallest index
65
+ `> s` such that `cp[e]` is not `SPACE` (or `e = n` if the run reaches the end of the
66
+ array). The run is `cp[s … e-1]`, length `k = e - s`.
67
+ 4. **Boundary guard.** Look one code point left of the run and one right of it:
68
+ - `left = cp[s-1]` if `s > 0`, otherwise `NONE`;
69
+ - `right = cp[e]` if `e < n`, otherwise `NONE`.
70
+ If `left` is `NONE` or is in `BREAK`, **skip the run entirely** (emit nothing, set
71
+ `i = e`, go to 2). This protects leading indentation, including Markdown list and code
72
+ indentation.
73
+ If `right` is `NONE` or is in `BREAK`, **skip the run entirely**. This protects the
74
+ two-space Markdown hard line break (`"foo \n"`) and trailing spaces at the end of a
75
+ text unit, which in `html` mode are frequently the only separator between two inline
76
+ elements.
77
+ Only a run with `CONTENT` on **both** sides is a candidate.
78
+
79
+ **In `html` and `markdown` mode, a span boundary marker counts as `NONE` here.** This is
80
+ the one place in the whole spec where a marker is not opaque content, it is normative, and
81
+ it is specified in [modes.md](modes.md) §3.3 under _Edge tests_ — read the justification
82
+ there before changing either document. In short: this guard exists because deleting
83
+ whitespace at an edge is irreversible and this rule cannot see past the edge, which is
84
+ exactly the situation at a span boundary; and without it `<em>mot</em> !` in `fr` loses its
85
+ space to the `STRIP-BEFORE` branch and gains nothing back, because `nbsp`'s replacement
86
+ insertion is then refused at the span edge.
87
+ 5. **Decide the replacement length in one step** (never two passes — see §5):
88
+ - If the _empty-bracket guard_ (§3.3) fires → replacement length **1**.
89
+ - Else if `left` is in `OPEN-BRACKET` **and the emoticon guard's mouth side (§3.6) does not
90
+ fire** → replacement length 0.
91
+ - Else if `right` is in `CLOSE-BRACKET` → replacement length 0.
92
+ - Else if `right` is in `STRIP-BEFORE` **and the lone-dot condition (§3.4) holds and the
93
+ emoticon guard's eye side (§3.6) does not fire** → replacement length 0.
94
+ - Else → replacement length 1.
95
+ The guard is a clause of this decision, not a separate "skip the run" branch. **This
96
+ reading is normative**; see §3.3.
97
+ 6. If the replacement length equals `k`, emit nothing (the run is already canonical).
98
+ Otherwise emit one edit replacing `cp[s … e-1]` with either the empty sequence or a
99
+ single U+0020.
100
+ 7. Set `i = e` and go to 2.
101
+
102
+ ### 3.3 The empty-bracket guard
103
+
104
+ A space run whose **removal** would produce an empty bracket pair is not removed. Concretely,
105
+ for a run at `[s, e)`:
106
+
107
+ - if `left` is in `OPEN-BRACKET` and `right` is the matching `CLOSE-BRACKET`
108
+ (`(`↔`)`, `[`↔`]`, `{`↔`}`), the guard fires.
109
+
110
+ **Normative reading: the guard forces the replacement length to 1, it does not skip the run.**
111
+ Two readings were possible and they differ on `"( )"` (two spaces between an empty pair):
112
+ "skip the run" leaves `"( )"`, "replacement length 1" collapses it to `"( )"`. The second is
113
+ normative. Reasons: the guard exists to prevent _deletion_, and collapsing a double space is
114
+ the rule's ordinary business everywhere else; a run of two spaces inside an empty bracket pair
115
+ carries no structural meaning in any of the four modes (the GFM task-list marker is
116
+ `"[ ]"` with exactly one space — `"[ ]"` is not a checkbox in any implementation); and the
117
+ one-clause form keeps §3.2 step 5 a single total function of the run's context — `left`, `right`
118
+ and the bounded lookaround of §3.4 and §3.6 — rather than a decision plus a separate skip branch,
119
+ which is what the idempotency argument in §5 relies on. The two-space case is a fixture.
120
+
121
+ The guard exists for one specific ship-blocking reason: the GitHub-Flavoured Markdown
122
+ task-list marker `"- [ ] item"`. Collapsing that space is a no-op; _deleting_ it produces
123
+ `"- [] item"` and silently destroys the checkbox. The guard also protects `"( )"` used as a
124
+ placeholder.
125
+
126
+ ### 3.4 The lone-dot condition
127
+
128
+ `STRIP-BEFORE` exists to remove a space that a typist left before **one terminal punctuation
129
+ mark**. It must not remove a space before a _run_ of dots, because a run of dots is a different
130
+ kind of token — a relative path, a truncation, a typed ellipsis — and deleting the space either
131
+ merges it with a preceding abbreviation dot or silently destroys word spacing.
132
+
133
+ > **Lone-dot condition.** If `right` is U+002E, the replacement length may be 0 **only if** the
134
+ > maximal run of `DOTLIKE` code points beginning at that index has length exactly 1 **and** the
135
+ > code point after that dot, `cp[e+1]`, is neither a `LETTER` (ARCHITECTURE.md §4.1's Unicode
136
+ > category test, as in §3.6) nor an ASCII digit. Otherwise the replacement length is 1.
137
+ >
138
+ > **U+2026 is not in `STRIP-BEFORE` at all**, so a space before an existing ellipsis is never
139
+ > deleted.
140
+
141
+ Both halves are needed and the second is not obvious. Suppose only the first were adopted.
142
+ Then `Wait ...` keeps its space, `ellipsis` converts the run, and the result is `Wait …` — at
143
+ which point a _second_ pipeline pass finds a U+0020 before a U+2026, strips it, and yields
144
+ `Wait…`. That is a two-pass divergence of exactly the kind
145
+ [pipeline-idempotency.md](pipeline-idempotency.md) exists to prevent, introduced by the fix for
146
+ another one. Removing U+2026 from `STRIP-BEFORE` closes it: the author's spacing around an
147
+ ellipsis, however they wrote it, is preserved and is stable.
148
+
149
+ **The word-start clause (spec 1.2.0).** A single dot followed directly by a letter or a digit is
150
+ not terminal punctuation either: it starts a token. `.NET`, `.DWG`, `.gitignore`, `.env` and the
151
+ decimal `.5` are words, and before 1.2.0 the space in front of each was deleted —
152
+ `Use .NET, .NET Core` became `Use.NET,.NET Core`, `CAD files (.DWG, .STEP)` became
153
+ `CAD files (.DWG,.STEP)`. The first half of the condition could not see this, because it measures
154
+ only the dot run. The second half reads one code point further, the same distance and the same
155
+ `LETTER`/ASCII-digit test the emoticon guard's eye side already uses (§3.6 step 3).
156
+
157
+ The cost is stated here so it is not rediscovered as a bug: `end .Next sentence` — a misplaced
158
+ full stop with the following space also missing — keeps its stray space instead of becoming
159
+ `end.Next sentence`. Both readings of that input lose something, and the one given up is the
160
+ destructive one: deleting the space glues two words together, and a later pass cannot tell the
161
+ glued form from a genuine `end.Next`. Keeping it leaves the input as the author typed it. That is
162
+ the same principle as §3.6's trailing check and §7.11 — where deletion and preservation disagree,
163
+ the reading that deletes less wins. A dot followed by anything else — a space, a closing bracket,
164
+ a quotation mark, punctuation, a line terminator, the end of the text, or a span boundary marker
165
+ in `html`/`markdown` mode (which is in neither `LETTER` nor `DIGIT`, [modes.md](modes.md) §3.3) —
166
+ is still a terminal full stop, and the space before it is still deleted.
167
+
168
+ **What this deliberately does not break.** The runs in step 3 are maximal runs **in the input
169
+ array**, and the condition is evaluated against the input. In the Chicago-style spaced ellipsis
170
+ `Hello . . .` every dot is a _lone_ dot at the moment the decision is made — each is followed by
171
+ a space, not by another dot — so all three spaces are still stripped, the dots merge into
172
+ `Hello...`, and `ellipsis` converts them to `Hello…`. The condition costs that case nothing.
173
+
174
+ **What it does change.** `Wait ...` now becomes `Wait …` rather than `Wait…`; the space the
175
+ author typed survives. That is consistent with how this rule treats every other authorial
176
+ spacing choice (§7.5), and it is the price of not deleting the space in `See ../docs`.
177
+
178
+ ### 3.5 Greek compatibility punctuation, and what "match on a code point" means
179
+
180
+ Greek raises the one case where the character a reader sees and the code point a rule matches
181
+ on come apart. Part 1 and part 3 below are properties of the whole rule set and are settled
182
+ here rather than in the locale file; part 2 is about this rule only, and says so, because the
183
+ generalisation is exactly what does not hold (see §7.8).
184
+
185
+ | Reader sees | Recommended code point | Compatibility code point | Canonical relation |
186
+ | ----------------- | ---------------------- | ------------------------ | ------------------ |
187
+ | ερωτηματικό (`;`) | U+003B `;` | U+037E | U+037E ≡ U+003B |
188
+ | άνω τελεία (`·`) | U+00B7 `·` | U+0387 | U+0387 ≡ U+00B7 |
189
+
190
+ Both compatibility characters have a **canonical** decomposition, so NFC and NFD both map them
191
+ away; the Unicode Standard records that **most** vendor code pages never had them, and states
192
+ that their use "is not generally encouraged for representation of Greek punctuation"
193
+ (ch. 7 §7.2.1). The consequence for this spec is that a Greek question mark and a Latin
194
+ semicolon are, at the level a rule operates on, **the same code point** — U+003B — and no
195
+ amount of context lets a single scan distinguish them.
196
+
197
+ **The normative position, in three parts:**
198
+
199
+ 1. **Every rule matches on the literal code points enumerated in its own class tables, and on
200
+ nothing else.** No rule consults canonical equivalence, no rule normalises its input, and no
201
+ rule has a notion of "the same character spelled differently" (ARCHITECTURE.md §4.3). A class
202
+ table is a list of code points, not a list of characters.
203
+
204
+ 2. **For _this_ rule, the U+003B ambiguity is harmless.** U+003B is in `STRIP-BEFORE`. Read as a
205
+ Latin semicolon it takes no preceding space in any locale polytypo supports; read as a Greek
206
+ ερωτηματικό nothing in the Greek sources asks for one either — the nearest statements are
207
+ about the neighbouring marks and deny the French pattern by name for both (`Στα ελληνικά,
208
+ πριν από τη διπλή τελεία δεν πρέπει να υπάρχει διάστημα (πράγμα που συμβαίνει, π.χ., στα
209
+ γαλλικά)`, and the same sentence again for the άνω τελεία). Both readings therefore prescribe
210
+ the same edit and this rule never has to choose: `Τι κάνεις ;` becomes `Τι κάνεις;` whichever
211
+ character the author meant. Note the standard of evidence, because it is weaker than it looks
212
+ — no source examined addresses the ερωτηματικό _specifically_, so the claim is "no source
213
+ contradicts it", not "a source requires it". Nothing rests on the difference here, since
214
+ U+003B would be in `STRIP-BEFORE` for the Latin reading alone.
215
+
216
+ **This does not generalise to the other rules, and §7.8 records where it fails.** The
217
+ ambiguity is harmless in `spaces` because the two readings happen to agree; that is a fact
218
+ about the two conventions, not a property the rule set enjoys everywhere. `ellipsis` is the
219
+ counter-example: its `TERMINAL` set is `{U+0021, U+003F}`, so the Latin reading of U+003B
220
+ wins silently and `Πράγματι;..` — the genuinely Greek spelling — does not get the treatment
221
+ `Πράγματι?..` gets. **Any rule adding U+003B, U+0021 or U+003F to a class table must decide
222
+ the Greek reading explicitly rather than inheriting this paragraph.**
223
+
224
+ 3. **U+037E, U+0387 and U+00B7 are in no class of any rule, and no rule ever emits them.** Text
225
+ already written with a compatibility code point round-trips **untouched** — it is neither
226
+ converted to the recommended spelling nor treated as the character it decomposes to. Three
227
+ separate reasons, each sufficient:
228
+ - Rewriting U+037E → U+003B or U+0387 → U+00B7 **is** normalisation, whatever it is called,
229
+ and §4.3 forbids it. It is also redundant: any downstream consumer applying NFC does it.
230
+ - **U+00B7 must not join `STRIP-BEFORE`.** It is a deliberate spaced separator in real
231
+ content — `Home · About · Contact`, and the Catalan and French interpunct — and stripping
232
+ there would be a false positive in locales that have nothing to do with Greek. This rule
233
+ reads no locale data (§2), so a Greek-only membership is not available to it.
234
+ - **Adding only U+0387 would be the worst of the three options.** Two canonically equivalent
235
+ characters would behave differently, and the one that got the correct treatment would be
236
+ the spelling Unicode discourages, while the recommended U+00B7 kept its stray space. A
237
+ split that punishes the correct input is not a partial fix.
238
+
239
+ The residual cost is exact and small: a space before an άνω τελεία survives. It is a rare
240
+ typo, and the alternative costs other locales real false positives.
241
+
242
+ An earlier revision added, in this sentence, that "the recommended Greek keyboard layouts
243
+ produce U+00B7 and U+003B rather than the compatibility pair". **That clause has been
244
+ deleted.** It is an empirical claim about keyboard layouts, it was stated as fact in
245
+ normative prose, and no source was offered for it — it was the only sentence in this section
246
+ a reviewer could not back with something readable. It may well be true; CLDR keyboard data or
247
+ a vendor layout specification would settle it. Until one is cited the argument does not need
248
+ it, which is why deleting was cheaper than sourcing.
249
+
250
+ ### 3.6 The emoticon guard
251
+
252
+ `STRIP-BEFORE`'s two punctuation-adjacent members U+003A (:) and U+003B (;) are read two ways
253
+ in ordinary text: as sentence punctuation (`Note: read this`, `Wait; think`) and as the eye of a
254
+ Western text emoticon (`:-)`, `:)`, `;-)`). The two readings take opposite spacing: sentence
255
+ punctuation never wants a preceding space (hence `STRIP-BEFORE`), but an emoticon is a token in
256
+ its own right and the space before it is ordinary word spacing that must survive — `Привет :-)`
257
+ must not become `Привет:-)`.
258
+
259
+ Two of the mouth glyphs, U+0028 `(` and U+005B `[`, are also `OPEN-BRACKET` members, so the same
260
+ token has to be protected from the other end as well: the space **after** the mouth is ordinary
261
+ word spacing too, and the opening-bracket clause of §3.2 step 5 must not eat it. The guard
262
+ therefore has two sides. They recognise the same shape — `EMOTICON-EYE`, optional
263
+ `EMOTICON-NOSE`, `EMOTICON-MOUTH` — and differ only in which end of the space run it sits at, and
264
+ in which clause of step 5 they suppress.
265
+
266
+ > **Emoticon guard, eye side.** For a run whose `right` (at index `e`) is in `EMOTICON-EYE`, walk
267
+ > forward:
268
+ >
269
+ > 1. Let `i = e + 1`. If `cp[i]` is in `EMOTICON-NOSE`, set `i = i + 1`.
270
+ > 2. If `cp[i]` is not in `EMOTICON-MOUTH`, the guard does not fire.
271
+ > 3. Otherwise let `after = cp[i + 1]` (or `NONE`). If `after` is a `LETTER` (ARCHITECTURE.md
272
+ > §4.1's Unicode category test, as used throughout this spec) or an ASCII digit, the guard
273
+ > does not fire. Otherwise **the guard fires**, and the `STRIP-BEFORE` clause of §3.2 step 5
274
+ > does not strip the space.
275
+
276
+ > **Emoticon guard, mouth side.** For a run whose `left` (at index `s-1`) is in `EMOTICON-MOUTH`,
277
+ > walk backward:
278
+ >
279
+ > 1. Let `j = s - 2`. If `cp[j]` is in `EMOTICON-NOSE`, set `j = j - 1`.
280
+ > 2. If `cp[j]` is not in `EMOTICON-EYE`, the guard does not fire.
281
+ > 3. Otherwise **the guard fires**, and the `OPEN-BRACKET` clause of §3.2 step 5 does not delete
282
+ > the run.
283
+ >
284
+ > There is no trailing check on this side and none is needed: the code point after the mouth is
285
+ > the space run itself, which is neither a `LETTER` nor an ASCII digit, so the eye side's step 3
286
+ > is satisfied here by construction.
287
+
288
+ **What the mouth side fixes.** Without it the two clauses of step 5 contradicted each other about
289
+ the same character: the eye side recognised `(` as a mouth while the opening-bracket clause went on
290
+ reading it as a bracket. `a :( b` became `a :(b`, gluing the next word onto the emoticon — a false
291
+ positive of exactly the kind the M4 gate (`docs/ROADMAP.md`) forbids — and the damage propagated,
292
+ because with no space after the mouth `after` is a `LETTER`, the eye side stops firing, and a second
293
+ pass strips the space in front of the eye as well: `a:(b`. That is a two-pass divergence, a release
294
+ blocker under [pipeline-idempotency.md](pipeline-idempotency.md), and the eye side's own idempotency
295
+ claim was what had been wrong.
296
+
297
+ **Only the `OPEN-BRACKET` clause is suppressed.** The wider reading — "a run abutting a recognised
298
+ emoticon is always length 1" — was considered and rejected as broader than the defect. It would also
299
+ silence the `CLOSE-BRACKET` clause (`a :) )` would stay as typed instead of becoming `a :))`) and the
300
+ `STRIP-BEFORE` clause (`a :) , b`), neither of which damages the emoticon or the words around it,
301
+ and each of which would need its own composition argument. The mouth side is the smallest change
302
+ that closes the defect.
303
+
304
+ **Why the trailing check.** Without it, `:D` inside an ordinary word — `:Deal with it`,
305
+ `;Design review` — would read as an emoticon and keep a space that sentence punctuation never
306
+ wants. The mouth must be the end of a token, not the start of a capitalised word: `after` is
307
+ checked against `LETTER` and `DIGIT`, not against `SPACE` or `NONE` specifically, so `:-)!`
308
+ (mouth followed by punctuation) and `:-):-)` (two emoticons back to back) both still fire, and
309
+ `10:30` (colon before a digit, no mouth at all — the mouth check in step 2 already declines it
310
+ before this check is reached) and `:Deal` do not.
311
+
312
+ **Scope, deliberately narrow.** This guard recognises the eye-nose-mouth shape of a Western
313
+ text emoticon and nothing else: no East Asian kaomoji (`(^_^)`, whose parenthesis is the frame,
314
+ not the eye), no `=)` (U+003D is not in `STRIP-BEFORE` and has no bug to fix), no emoji. It
315
+ exists because `STRIP-BEFORE`'s membership of U+003A and U+003B was already normative and
316
+ locale-independent (§2), and an emoticon eye is the one shape that class was silently getting
317
+ wrong; it is not a general-purpose emoticon detector and does not try to be one.
318
+
319
+ **Both sides require an eye.** The backward walk is unambiguous because `EMOTICON-NOSE` and
320
+ `EMOTICON-EYE` are disjoint, and a nose on its own is not a face: `a -( b` still becomes `a -(b`
321
+ under the ordinary opening-bracket clause.
322
+
323
+ **Idempotency.** An earlier revision claimed here that the guard "reads only `cp[e]`, `cp[e+1]` and
324
+ `cp[e+2]`, none of which this rule … ever modifies". Both halves were false: with a nose the eye
325
+ side reads through `cp[e+3]`, and step 3's `after` is, in precisely the shape that broke, the U+0020
326
+ run step 5 was about to delete — so the verdict was not a pure function of code points this rule
327
+ does not write. The correct argument has one part per side.
328
+
329
+ - **Mouth side.** It reads `cp[s-1]`, `cp[s-2]` and `cp[s-3]`, and in no position it accepts is that
330
+ a U+0020: an eye, a nose and a mouth are all `CONTENT` and must be adjacent. This rule never
331
+ inserts a code point and never removes a non-space one, so a shape it recognised in the input is
332
+ still contiguous and unchanged in the output. The verdict can therefore only move from "does not
333
+ fire" to "fires" — never the reverse — and firing only ever preserves a run, so neither direction
334
+ can produce a second-pass edit.
335
+ - **Eye side.** Whenever the eye side fires, the mouth side recognises the same three code points
336
+ from the other end: the two walks read `EMOTICON-NOSE` and `EMOTICON-EYE`, which are disjoint, so
337
+ the backward walk's optional step over a nose cannot swallow the eye and the two sides cannot
338
+ disagree about the shape. A run whose `left` is that mouth is therefore protected from the
339
+ `OPEN-BRACKET` clause, and the only clauses of step 5 that can still delete it are the
340
+ `CLOSE-BRACKET` clause and the `STRIP-BEFORE` clause. The code point that comes to sit after the
341
+ mouth is then one of `)` `]` `}` or one of `STRIP-BEFORE`'s six; none of those nine is a `LETTER`
342
+ or an ASCII digit, so step 3 still passes on the re-run. Every other outcome — the empty-bracket
343
+ guard, a collapse to length 1, a run skipped by the boundary guard of §3.2 step 4 — leaves the
344
+ U+0020 in place, so `after` is unchanged. Either way the eye side's verdict survives the re-run.
345
+
346
+ ### 3.7 Worked trace
347
+
348
+ `"a ( b , c ) d"`
349
+
350
+ | run | left | right | decision |
351
+ | ------ | ---- | ----- | -------------------------------------------------------- |
352
+ | `a␣␣(` | `a` | `(` | not bracket-inner, right not in `STRIP-BEFORE` → 1 space |
353
+ | `(␣␣b` | `(` | `b` | left is `OPEN-BRACKET`, guard does not fire → 0 |
354
+ | `b␣␣,` | `b` | `,` | right in `STRIP-BEFORE` → 0 |
355
+ | `,␣␣c` | `,` | `c` | → 1 space |
356
+ | `c␣␣)` | `c` | `)` | right is `CLOSE-BRACKET` → 0 |
357
+ | `)␣␣d` | `)` | `d` | → 1 space |
358
+
359
+ Result: `"a (b, c) d"`.
360
+
361
+ ---
362
+
363
+ ## 4. Must not touch
364
+
365
+ **Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
366
+ a guarantee of `transform` as a whole — or **[R]** — true of this rule alone and capable of
367
+ being falsified by another rule. This rule is R₁, so no _earlier_ rule can invalidate anything
368
+ here; but later rules touch some of the same characters, and those bullets are marked [R].
369
+
370
+ - **[R] Any character other than U+0020.** Tabs (U+0009) are never collapsed, never converted,
371
+ never removed — in Markdown a tab is indentation and the engine has no way to know whether
372
+ it is inside a code block. U+00A0 and U+202F are never collapsed, never removed, and never
373
+ converted to U+0020; a run such as `"a  b"` is left exactly as written.
374
+ U+2000–U+200A, U+205F, U+3000, U+200B and U+FEFF are likewise untouched.
375
+ _[R]: `nbsp` (R₈) may convert an existing U+00A0 to U+202F or the reverse. `transform` does
376
+ not promise an existing no-break space keeps its exact width — only that this rule never
377
+ collapses or deletes one._
378
+ - **[P] Line terminators.** No character in `BREAK` is inserted, removed, or reordered. Space
379
+ runs never merge across a line terminator, because a `BREAK` on either side aborts the
380
+ candidate.
381
+ - **[R] Leading whitespace on a line** (a run at index 0 or directly after a `BREAK`). Markdown
382
+ indented code blocks, nested list indentation and YAML front matter all depend on it.
383
+ _[R]: `nbsp` (R₈) can convert the **last** space of an indentation run when the character
384
+ after it is listed in `beforePunctuation`/`narrowBeforePunctuation` — a French line beginning
385
+ `␣␣?`. The run is never shortened, so indentation width survives; one code point changes
386
+ class._
387
+ - **[P] Trailing whitespace before a line terminator or at the end of the text unit.** This is
388
+ the Markdown hard-break idiom and, in `html` mode, the inter-element separator.
389
+ - **[R] Single spaces in ordinary positions.** `"a b"` is not a candidate for anything.
390
+ - **[P] `"[ ]"`, `"( )"`, `"{ }"`** — see §3.3.
391
+ - **[R] Spaces before a closing quotation mark or guillemet.** Owned by `nbsp` via
392
+ `quotes.innerSpace`.
393
+ - **[R] A space before an emoticon's eye, and a space after an emoticon's mouth.** §3.6.
394
+ `Привет :-)` keeps its space in every locale, and so does `Привет :-( снова`, whose second space
395
+ would otherwise go to the opening-bracket clause. This rule reads no locale data (§2) and neither
396
+ side of the guard is locale-dependent either.
397
+ _[R]: `nbsp` (R₈) may convert the space **before** the eye to a no-break form in a locale whose
398
+ `nbsp.beforePunctuation` lists U+003A or U+003B — French `a :) b` yields `a⍽:) b`, because
399
+ [nbsp.md](nbsp.md) §3.3 step 2's right-context guard admits `CLOSEISH` after the mark. The run is never
400
+ deleted, so word spacing survives; one code point changes class. The space **after** the mouth
401
+ is touched by no later rule._
402
+ - **[P] U+037E, U+0387, U+00B7.** In no class of this rule and of no other. A space before an
403
+ άνω τελεία survives, and a Greek compatibility code point is never rewritten to the character
404
+ it canonically decomposes to — §3.5. Note that the Greek ερωτηματικό is _not_ an exception
405
+ here: it is written U+003B, which is in `STRIP-BEFORE` on its own merits.
406
+ - **[P] Anything inside a skipped region.** Code spans, fenced code, `<pre>`, attributes and
407
+ URLs never reach this rule; the mode adapter (L2) removes them before the pipeline runs.
408
+ This rule contains no code-awareness of its own and must not grow any.
409
+
410
+ ---
411
+
412
+ ## 5. Idempotency argument
413
+
414
+ Let `T` be the transformation described in §3.2.
415
+
416
+ Every edit replaces a maximal `SPACE` run bounded by `CONTENT` on both sides with either
417
+ zero or one U+0020. Consider the output `T(x)` and re-run the scan.
418
+
419
+ - The bounding characters of every run are `CONTENT` and are never modified by this rule, so
420
+ the `left`/`right` classification of any surviving run is unchanged between runs.
421
+ - A run replaced by zero spaces no longer exists; its former neighbours `left` and `right`
422
+ are now adjacent. `right` is a member of `STRIP-BEFORE` or a bracket, and `left` is
423
+ `CONTENT`. No new `SPACE` run has been created — deletion cannot create a space — so
424
+ there is nothing to re-examine at that position.
425
+ - A run replaced by one space is now a run of length `k' = 1` with the same `left` and
426
+ `right`. Re-running step 5 on it yields the same decision — it is a function of `left`, `right`
427
+ and the bounded lookaround of §3.4 and §3.6, each of which argues its own window's stability
428
+ (§3.4's word-start clause: see the paragraph after this list) —
429
+ namely replacement length 1, and step 6 then emits nothing because `k' = 1` already equals the
430
+ replacement length.
431
+ - A skipped run is skipped again for the same reason (its guard condition depends only on
432
+ `left`, `right` and bracket matching, all unchanged).
433
+
434
+ **§3.4's window.** The lone-dot condition reads `cp[e]`, the dot run starting there, and — since
435
+ 1.2.0 — `cp[e+1]`. The dot run is made of non-space code points this rule never writes. `cp[e+1]`
436
+ can change between passes only if it was a U+0020 whose run this pass deleted, bringing its right
437
+ neighbour against the dot. A run is deleted only by the `STRIP-BEFORE`, `CLOSE-BRACKET` and
438
+ `OPEN-BRACKET` clauses of step 5; the run after a dot has the dot as its `left`, so the
439
+ `OPEN-BRACKET` clause cannot apply, and the other two leave behind a `STRIP-BEFORE` member or a
440
+ closing bracket — none of which is a `LETTER` or an ASCII digit. So the word-start clause's verdict
441
+ ("is `cp[e+1]` a letter or a digit?") is the same on both passes, whatever this pass did after the
442
+ dot.
443
+
444
+ Therefore `T(T(x)) = T(x)`.
445
+
446
+ **What had to be fixed to get here.** The naive formulation — "first collapse doubles, then
447
+ strip spaces before punctuation" — is two passes and is _not_ obviously idempotent, and worse,
448
+ it is not obviously order-independent: `"a , b"` collapses to `"a , b"` and then strips to
449
+ `"a, b"`, which requires the second pass to run over the output of the first, i.e. a
450
+ fixed-point loop. Fixed-point loops are exactly what a spec must not require, because two
451
+ implementations will disagree about how many iterations they run. The formulation above
452
+ computes the replacement length for each run **once**, from `left` and `right` alone, so a
453
+ single pass reaches the fixed point directly.
454
+
455
+ The second thing that had to be fixed: the naive rule "collapse every run of ≥2 spaces" is
456
+ not merely non-idempotent-adjacent, it is _destructive_ on Markdown hard breaks and on
457
+ indentation. The `CONTENT`-on-both-sides boundary guard (step 4) is what makes the rule safe,
458
+ and it is a precondition of the argument above, not an optimisation.
459
+
460
+ ---
461
+
462
+ ### Composition obligation
463
+
464
+ Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₁**: nothing runs
465
+ before it, so **the obligation is empty**. There is no earlier invariant it could break.
466
+
467
+ The obligation pointing the other way is not empty, and it is the one that bit. `I₁` — the
468
+ statement that this rule is a no-op, spelled out as S-a … S-d in that document §3 — must be
469
+ preserved by all seven later rules. Only one of them emits U+0020 at all (`dashes`), and it
470
+ violated S-b, S-c and S-d until guard T2 was added. The exact positions from which this rule
471
+ deletes a space are therefore load-bearing for the whole pipeline, and §3.2 step 5 should be
472
+ treated as a published interface rather than an implementation detail.
473
+
474
+ **Spec 1.2.0's word-start clause adds one thing a later rule must not do.** S-b now permits a
475
+ U+0020 before a lone dot whose next code point is a `LETTER` or an ASCII digit. A later rule that
476
+ replaced that letter or digit with something else would turn a permitted space into a forbidden one,
477
+ without emitting any U+0020. None does: `ellipsis` writes only `DOTLIKE` code points; `ranges`,
478
+ `dashes` and `hyphen` replace dashes, hyphens and spaces; `quotes` and `apostrophe` replace
479
+ quotation marks; `symbols` replaces a trademark literal, which starts with `(`, and a `MUL-LETTER`,
480
+ which is always preceded by a digit or a space and so never follows a dot directly; `nbsp` replaces
481
+ or inserts spaces, and inserts only beside a listed punctuation mark or a quote glyph, never between
482
+ a dot and a letter. A new rule that rewrites letters or digits must re-check this.
483
+
484
+ ---
485
+
486
+ ## 6. Worked examples
487
+
488
+ `␣` = U+0020, `⟶` = no change expected, `↵` = U+000A, `⍽` = U+00A0. Every row is
489
+ locale-independent **except 9b**, which is marked with its locale: the rule reads no locale data
490
+ (§2), but that row's point is that a Greek reading and a Latin reading of U+003B reach the same
491
+ verdict here, so naming the locale is what makes the claim checkable.
492
+
493
+ | # | Input | Output | Why |
494
+ | --- | ---------------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
495
+ | 1 | `Hello␣␣␣world.` | `Hello␣world.` | run of 3 with `CONTENT` both sides → 1 |
496
+ | 2 | `Hello␣,␣world␣!` | `Hello,␣world!` | `,` and `!` are in `STRIP-BEFORE`; the run after `,` keeps one space |
497
+ | 3 | `(␣ok␣)␣and␣[␣x␣]` | `(ok)␣and␣[x]` | bracket-inner runs deleted; the guard does not fire because the inner side is not the matching closer |
498
+ | 4 | `-␣[␣]␣buy␣milk` | ⟶ | empty-bracket guard (§3.3): replacement length 1, and the run is already length 1 → no edit. The GFM checkbox survives |
499
+ | 4b | `(␣␣)` | `(␣)` | empty-bracket guard forces length 1, it does not skip — §3.3, normative reading |
500
+ | 5 | `line␣one␣␣↵line␣two` | ⟶ | the two-space run is directly followed by `BREAK` → skipped; the Markdown hard break survives |
501
+ | 6 | `␣␣␣␣indented␣code` | ⟶ | run starts at index 0 → skipped |
502
+ | 7 | `5⍽␣␣km` | `5⍽␣km` | the U+00A0 is `CONTENT`; only the U+0020 run beside it collapses, and the U+00A0 itself is untouched |
503
+ | 8 | `a⍽⍽b` | ⟶ | no U+0020 anywhere; two no-break spaces are never collapsed |
504
+ | 9 | `Bonjour␣!␣Ça␣va␣?` | `Bonjour!␣Ça␣va?` | French input; the ordinary spaces go, and `nbsp` (order 70) later restores `Bonjour⁠<U+202F>!` |
505
+ | 9b | `Τι␣κάνεις␣;` | `Τι␣κάνεις;` | **`el`.** The Greek question mark is written U+003B, which is in `STRIP-BEFORE` on the Latin semicolon's own merits, so the space is stripped with no Greek-specific decision — §3.5 part 2. `nbsp` puts nothing back, because `el`'s punctuation lists are empty |
506
+ | 10 | `See␣p.␣12␣.` | `See␣p.␣12.` | run before the final `.` deleted — a lone dot, so the condition holds |
507
+ | 10a | `See␣../docs` | ⟶ | **lone-dot condition (§3.4).** The dot run has length 2, so the space survives. Previously produced `See../docs`, silently destroying word spacing in a relative path |
508
+ | 10b | `e.g.␣..` | ⟶ | same. Previously produced `e.g...`, which `ellipsis` then legitimately read as a three-dot run and converted to `e.g…` |
509
+ | 10c | `Wait␣...` | ⟶ | the space survives (run length 3). `ellipsis` then yields `Wait␣…`, and because U+2026 is not in `STRIP-BEFORE` that is a fixed point |
510
+ | 10d | `Hello␣.␣.␣.` | `Hello...` | every dot is a **lone** dot in the input, so all three spaces strip and the Chicago-style spaced ellipsis still merges — `ellipsis` converts it to `Hello…` |
511
+ | 10e | `Wait␣…` | ⟶ | U+2026 is not in `STRIP-BEFORE` |
512
+ | 10f | `Use␣.NET,␣.NET␣Core` | ⟶ | **word-start clause (§3.4, spec 1.2.0).** Each dot is followed by a letter, so it starts a word and neither space is deleted. Previously produced `Use.NET,.NET␣Core` |
513
+ | 10g | `(.DWG,␣.STEP)` | ⟶ | same: the space after the comma survives because `.STEP` is a token, not a full stop. The comma run is untouched regardless — its `right` is the dot, not the comma |
514
+ | 10h | `from␣.5␣to␣.9` | ⟶ | same, for an ASCII digit after the dot |
515
+ | 10i | `end␣.Next` | ⟶ | **the accepted cost.** A misplaced full stop followed by a letter keeps its stray space; the alternative glues two words — §3.4 |
516
+ | 10j | `See␣p.␣12␣.␣Next` | `See␣p.␣12.␣Next` | a dot followed by a space is still terminal punctuation, so row 10 is unchanged by the word-start clause |
517
+ | 11 | `foo␣␣␣␣↵␣␣␣␣bar` | ⟶ | first run touches a `BREAK` on the right, second on the left |
518
+ | 12 | `Q:␣␣why␣?␣␣Because␣.` | `Q:␣why?␣Because.` | mixed |
519
+ | 13 | `Привет␣:-)` | ⟶ | **emoticon guard (§3.6).** `:` is `EMOTICON-EYE`, `-` is `EMOTICON-NOSE`, `)` is `EMOTICON-MOUTH`, and there is nothing after it — the guard fires and the space survives |
520
+ | 13a | `Привет␣:)` | ⟶ | same, no nose |
521
+ | 13b | `Hello␣:Deal␣with␣it` | `Hello:Deal␣with␣it` | `D` is `EMOTICON-MOUTH`, but `e` immediately after it is a `LETTER` — the guard does not fire, and ordinary `STRIP-BEFORE` behaviour applies |
522
+ | 13c | `10:30` | ⟶ | no space run at all — outside this rule's scope regardless of the guard |
523
+ | 13d | `See␣you␣at␣10␣:␣30` | `See␣you␣at␣10:␣30` | `:` is `EMOTICON-EYE`, but the very next code point is a space, not `EMOTICON-MOUTH` — the guard does not fire; ordinary `STRIP-BEFORE` strips the leading space, and the trailing space (right neighbour `3`, not `STRIP-BEFORE`) is untouched |
524
+ | 13e | `Sorry␣:(␣it␣happens` | ⟶ | **mouth side (§3.6).** `(` is the mouth of a recognised emoticon, so the opening-bracket clause does not delete the space after it. Previously produced `Sorry␣:(it␣happens`, and then `Sorry:(it␣happens` on a second pass |
525
+ | 13f | `Hmm␣:[␣well` | ⟶ | same, for the other mouth that is also an `OPEN-BRACKET` member — `[` |
526
+ | 13g | `Well␣:-(␣then` | ⟶ | same, with a nose: the backward walk steps over `-` and finds the eye at `cp[s-3]` |
527
+ | 13h | `a␣-(␣b` | `a␣-(b` | a nose with no eye behind it is not a face — the mouth side does not fire and the opening-bracket clause applies as usual |
528
+ | 13i | `word␣(␣note␣)` | `word␣(note)` | the ordinary bracket-inner case, unchanged by the mouth side: `(` here is preceded by a space, not by an eye |
529
+ | 13j | `Note:(␣x␣)` | `Note:(␣x)` | the mouth side puts **no** condition on what precedes the eye, so it fires on a word-attached eye too: the space after the mouth survives while the one before the closer still goes — §7 item 11 |
530
+
531
+ Cases 4, 5, 6, 8, 10a, 10b, 10c, 10e, 10f, 10g, 10h, 10i, 11, 13, 13a, 13e, 13f and 13g are "no change" cases.
532
+
533
+ ---
534
+
535
+ ## 7. Open questions
536
+
537
+ 1. **U+0009 (tab) is completely untouched.** In `text` mode a run of tabs between two words
538
+ is arguably the same typing accident as a run of spaces. I chose safety, because the rule
539
+ cannot distinguish prose from Markdown indentation. If the dogfooding gate (PLAN.md M4)
540
+ shows tab noise in real content, the fix is a separate opt-in rule, not a change here.
541
+ 2. **U+2009 (thin space) and friends are untouched.** Some authors paste text from InDesign
542
+ or LaTeX carrying real thin spaces. Normalising them to U+202F would be locale-dependent
543
+ and is arguably a `nbsp` concern. Unresolved; currently they are simply preserved.
544
+ 3. **`STRIP-BEFORE` includes U+003A (colon).** `"10 : 30"` becomes `"10:30"`. I believe this
545
+ is right for prose but it is a real behaviour change on tabular text. Needs a fixture
546
+ decision from the operator.
547
+ 4. **`STRIP-BEFORE` excludes U+2014/U+2013.** A spaced em dash is a legitimate parenthetical
548
+ form in several locales, so stripping there would be wrong; the `dashes` rule owns dash
549
+ spacing. Confirmed by construction, but worth a fixture.
550
+ 5. **Should a space before an opening bracket be normalised?** `"word(note)"` versus
551
+ `"word (note)"` is an authorial choice, not a typographic error, so this rule does
552
+ nothing. Recorded so nobody adds it later "for symmetry".
553
+ 6. **`html` mode boundary semantics.** A text node ending in a space followed by a sibling
554
+ text node beginning with a space is, at the DOM level, two runs of length 1 that render
555
+ as one collapsed space. This rule sees them separately and leaves both. Whether the mode
556
+ adapter should present adjacent text nodes as one logical unit is an L2 question that
557
+ this document cannot settle; it is flagged here because the answer changes `spaces`
558
+ fixtures for `html` mode.
559
+ 7. **The lone-dot condition (§3.4) changes `Wait␣...` to `Wait␣…` rather than `Wait…`.** The
560
+ author's space survives. I judge that correct — it is the same principle as §7.5, and it is
561
+ what makes `See␣../docs` safe — but it is a visible change from the previous behaviour and
562
+ it deserves a fixture and an operator glance. If tight is preferred, the fix is **not** to
563
+ restore stripping before a dot run (that reopens `See␣../docs`) but to have `ellipsis`
564
+ absorb a preceding space, which is a different rule's business and a different decision.
565
+ 8. **The Greek reading of U+003B is honoured here and ignored by `ellipsis`.** §3.5 part 2 is
566
+ true of this rule and does not generalise; this entry is where the generalisation fails, and
567
+ it was found by review rather than by construction, so assume there are others.
568
+ `ellipsis.md` §3.1 defines `TERMINAL` as `{U+0021, U+003F}`. A Greek author writes a question
569
+ with U+003B, so `Πράγματι;..` keeps its two-dot run while `Πράγματι?..` — which uses a
570
+ character Greek does not use for questions — becomes `Πράγματι?…`. Greek gets the treatment
571
+ only when it is written wrongly. Verified against the engine, and pinned by the fixture
572
+ `el-ellipsis-greek-question-mark-not-terminal`.
573
+
574
+ It is a **miss, not damage**: the input round-trips unchanged, and the two-dot run is left
575
+ exactly as written. That is why it is recorded rather than fixed here.
576
+
577
+ **The obvious fix — adding U+003B to `TERMINAL` — was researched and is REFUSED.** This is a
578
+ settled decision, not an open question, and it is recorded here so it is not reopened:
579
+
580
+ - **It would be a regression in Russian.** The abbreviated two-dot form is tied to two named
581
+ marks, not to a class of "terminal punctuation". Лопатин, «Правила русской орфографии и
582
+ пунктуации. Полный академический справочник» §154: «При сочетании вопросительного или
583
+ восклицательного знака с многоточием знаки эти ставятся на месте первой точки» — and
584
+ §§155–158, which cover every other combination, never mention the semicolon. Розенталь
585
+ §68.1 says the same. So `текст;...` must stay `текст;…` in Russian, and putting U+003B in
586
+ `TERMINAL` would invent a rule that no Russian authority states.
587
+ - **No Greek source asks for it either.** The Ministry of Education grammar and the Κέντρο
588
+ Ελληνικής Γλώσσας materials describe αποσιωπητικά without giving any rule for their
589
+ interaction with the ερωτηματικό, and Greek has no counterpart to the Russian
590
+ dot-absorption convention at all. «πάντοτε τρεις» is stated against runs of four and five
591
+ dots, not against a neighbouring mark.
592
+ - So the change is **unsupported on both sides**: it would break the one locale with a cited
593
+ rule in order to serve a locale whose sources are silent.
594
+
595
+ What remains is the two-dot input `Πράγματι;..`, which no authority addresses in either
596
+ language. It round-trips unchanged, which is the correct behaviour for an unspecified case.
597
+ The correctly-written three-dot form `Πράγματι;...` already converts to U+2026, because a
598
+ three-dot run needs no help from `TERMINAL`.
599
+ 9. **A verified Greek source contradicts §3.4, and the divergence is deliberate.** The EU
600
+ Interinstitutional Style Guide (Greek edition) §10.1.9 ii) states «Μεταξύ των αποσιωπητικών
601
+ και της λέξης που προηγείται δεν αφήνουμε διάστημα» — *no space is left between the ellipsis
602
+ and the word before it*. §3.4 removes U+2026 from `STRIP-BEFORE` and preserves a space before
603
+ a dot run **in every locale**, so `Πράγματι …` is returned as typed. The rule and the source
604
+ disagree, the source was read and verified, and this entry exists so that the disagreement is
605
+ recorded rather than unremarked.
606
+
607
+ **The divergence stands, and the space is preserved.** Three reasons, in the order that
608
+ decided it:
609
+
610
+ - **Honouring it requires a rule that deletes on locale data, and this rule has none.**
611
+ `order.json` gives `spaces` `"localeData": []`; §2 states that its behaviour is identical in
612
+ every locale, which is not decoration but the reason it can be reasoned about at all. The
613
+ alternatives are to give `spaces` locale data — an `order.json` change, and the end of a
614
+ property the whole document relies on — or to make `ellipsis` delete the preceding space,
615
+ which turns a rule that only ever replaces into a rule that deletes, and thereby drags in
616
+ `modes.md` §3.3's edge-test clause, `I₂`, and a fresh CO discharge. Neither is a small change,
617
+ and neither is justified by one locale.
618
+ - **The instruction is about setting text, not about repairing it.** The guide tells a Greek
619
+ typist not to leave a space there. It does not ask a tool to remove one the writer left. That
620
+ distinction is the same one §3.5 part 2 draws for the ερωτηματικό, and it is the reason this
621
+ document is comfortable saying "no source contradicts it" there and must not say "a source
622
+ requires it" here.
623
+ - **The divergence is conservative in the direction the M4 gate cares about.** Preserving
624
+ costs a missed correction on Greek text containing a typo; honouring it would mean deleting
625
+ a character the author typed, in one locale only, through machinery no other locale exercises.
626
+ `dashes` §3.2 step 2a was just narrowed on exactly that principle after 1063 lines of the
627
+ author's corpus were rewritten against his intent.
628
+
629
+ **What would change the answer:** a second locale wanting the same behaviour. One locale does
630
+ not pay for a schema field, a deleting `ellipsis`, and a new composition argument; two might,
631
+ and at that point the right shape is `ellipsis.noSpaceBefore` consumed by `ellipsis` — which
632
+ already reads locale data — rather than anything in this rule. Recorded in `ellipsis.md` §7 and
633
+ in `spec/locales/el.json`'s `ellipsis` note, so that all three say the same thing.
634
+ 10. **An emoticon directly inside a bracket pair still loses the space before its eye.**
635
+ `(␣:(␣b␣)` yields `(:(␣b)`: the opening-bracket clause acts on the outer `(`, and §3.6's eye
636
+ side never sees that run because it only suppresses the `STRIP-BEFORE` clause. Pre-existing
637
+ behaviour, identical for `(␣:)␣b␣)`, and it is a bracket-inner deletion of exactly the kind
638
+ case 3 asks for — so it was left alone rather than folded into the mouth-side fix. Recorded so
639
+ that anyone tempted to widen the guard "for symmetry" starts from the fact that it is the outer
640
+ bracket, not the emoticon, that owns that space.
641
+ 11. **The mouth side asks nothing about what precedes the eye, and that is deliberate.** `Note:( x )`
642
+ yields `Note:( x)` — asymmetric, because the mouth side keeps the inner space on the left while
643
+ the `CLOSE-BRACKET` clause still takes the one on the right. The alternative, requiring the eye
644
+ to begin a token (`SPACE`, `BREAK` or `NONE` before it), would restore the symmetry for that
645
+ input and lose `Hi!:( yes`, where the eye is attached to the preceding word and the shape is
646
+ still a face. Deleting is the irreversible direction, so the reading that deletes less wins —
647
+ the same principle as §3.4 and §7.9. Pinned by the fixture `en-us-spaces-emoticon-mouth-eye-attached-to-word`,
648
+ which is the case that discriminates the two readings; without it a port could take the narrower
649
+ one and still pass the whole suite.