polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,649 @@
|
|
|
1
|
+
# Rule: `spaces`
|
|
2
|
+
|
|
3
|
+
**Order:** 10 (first). **Default:** on. **Modes:** text, html, markdown, yaml.
|
|
4
|
+
**Spec version:** 1.2.0 (0.2.0 for everything except §3.6's mouth side and the clause it adds to
|
|
5
|
+
§3.2 step 5, noted inline, and §3.4's word-start clause, added in 1.2.0).
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Purpose
|
|
10
|
+
|
|
11
|
+
`spaces` performs the space hygiene that every later rule depends on: it collapses a run of
|
|
12
|
+
two or more ordinary spaces down to one, and it deletes an ordinary space that sits between
|
|
13
|
+
a word and a following punctuation mark or on the inner edge of a bracket pair. It runs
|
|
14
|
+
first so that the rules after it see one canonical spacing form and never have to consider
|
|
15
|
+
`"a , b"` alongside `"a, b"`. It is deliberately the most conservative rule in the
|
|
16
|
+
pipeline: it only ever _removes_ U+0020 (space), it never inserts anything, it never touches
|
|
17
|
+
any other whitespace character, and it never touches whitespace that carries structural
|
|
18
|
+
meaning (indentation, Markdown hard line breaks, line terminators). Where a locale genuinely
|
|
19
|
+
wants a space before punctuation — French `?` `!` `;` `:` — this rule still removes the
|
|
20
|
+
ordinary space, and the `nbsp` rule (order 70) re-inserts the correct no-break form; that
|
|
21
|
+
round trip is intentional and is what makes the French output deterministic regardless of
|
|
22
|
+
how the author typed it.
|
|
23
|
+
|
|
24
|
+
---
|
|
25
|
+
|
|
26
|
+
## 2. Locale data consumed
|
|
27
|
+
|
|
28
|
+
**None.** `order.json` declares `"localeData": []` for this rule. The behaviour of `spaces`
|
|
29
|
+
is identical in every locale. Locale-dependent spacing is entirely the responsibility of
|
|
30
|
+
`nbsp`.
|
|
31
|
+
|
|
32
|
+
---
|
|
33
|
+
|
|
34
|
+
## 3. Algorithm
|
|
35
|
+
|
|
36
|
+
The input is a code-point array `cp[0 … n-1]`. Indices below are code-point indices
|
|
37
|
+
(ARCHITECTURE.md §4.2). The rule emits edits; the pipeline applies them.
|
|
38
|
+
|
|
39
|
+
### 3.1 Character classes
|
|
40
|
+
|
|
41
|
+
| Class | Members |
|
|
42
|
+
| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
43
|
+
| `SPACE` | U+0020 (space) — **the only** character this rule ever removes |
|
|
44
|
+
| `BREAK` | U+000A (LF), U+000D (CR), U+000B (VT), U+000C (FF), U+0085 (NEL), U+2028 (LS), U+2029 (PS) |
|
|
45
|
+
| `PROTECTED-SPACE` | U+0009 (tab), U+00A0 (nbsp), U+202F (narrow nbsp), U+2007 (figure space), U+2008 (punctuation space), U+2009 (thin space), U+200A (hair space), U+2000–U+2006, U+205F (medium math space), U+3000 (ideographic space), U+200B (zero-width space), U+FEFF |
|
|
46
|
+
| `CONTENT` | any code point that is **not** in `SPACE` and **not** in `BREAK`. Note that every member of `PROTECTED-SPACE` is `CONTENT` for the purposes of this rule — it bounds a space run and is never itself modified. |
|
|
47
|
+
| `STRIP-BEFORE` | exactly six code points: U+002C (,) U+002E (.) U+003B (;) U+003A (:) U+0021 (!) U+003F (?). **U+2026 is deliberately not a member** — see §3.4 |
|
|
48
|
+
| `DOTLIKE` | U+002E (.) and U+2026 (…) |
|
|
49
|
+
| `OPEN-BRACKET` | U+0028 `(` U+005B `[` U+007B `{` |
|
|
50
|
+
| `CLOSE-BRACKET` | U+0029 `)` U+005D `]` U+007D `}` |
|
|
51
|
+
| `EMOTICON-EYE` | U+003A (:), U+003B (;) — a subset of `STRIP-BEFORE`, not a new code point this rule reads outside it. See §3.6 |
|
|
52
|
+
| `EMOTICON-NOSE` | U+002D (-), U+005E (^). Optional. See §3.6 |
|
|
53
|
+
| `EMOTICON-MOUTH` | U+0028 `(`, U+0029 `)`, U+005B `[`, U+005D `]`, U+0044 `D`, U+0064 `d`, U+0050 `P`, U+0070 `p`, U+004F `O`, U+006F `o`, U+002F `/`, U+005C `\`, U+007C `|`, U+002A `*`. U+0028 and U+005B are `OPEN-BRACKET` members too, and U+0029 and U+005D are `CLOSE-BRACKET` members — the overlap is what §3.6's mouth side exists for. See §3.6 |
|
|
54
|
+
|
|
55
|
+
`STRIP-BEFORE` contains only these **six** code points. It deliberately excludes closing
|
|
56
|
+
quotation marks and guillemets: their inner spacing is `nbsp`'s business (`quotes.innerSpace`),
|
|
57
|
+
and stripping there would fight with it.
|
|
58
|
+
|
|
59
|
+
### 3.2 Scan
|
|
60
|
+
|
|
61
|
+
1. Set `i = 0`.
|
|
62
|
+
2. If `cp[i]` is not `SPACE`, emit nothing, set `i = i + 1`, repeat from 2. Terminate when
|
|
63
|
+
`i = n`.
|
|
64
|
+
3. `cp[i]` is `SPACE`. Find the maximal run: let `s = i`; let `e` be the smallest index
|
|
65
|
+
`> s` such that `cp[e]` is not `SPACE` (or `e = n` if the run reaches the end of the
|
|
66
|
+
array). The run is `cp[s … e-1]`, length `k = e - s`.
|
|
67
|
+
4. **Boundary guard.** Look one code point left of the run and one right of it:
|
|
68
|
+
- `left = cp[s-1]` if `s > 0`, otherwise `NONE`;
|
|
69
|
+
- `right = cp[e]` if `e < n`, otherwise `NONE`.
|
|
70
|
+
If `left` is `NONE` or is in `BREAK`, **skip the run entirely** (emit nothing, set
|
|
71
|
+
`i = e`, go to 2). This protects leading indentation, including Markdown list and code
|
|
72
|
+
indentation.
|
|
73
|
+
If `right` is `NONE` or is in `BREAK`, **skip the run entirely**. This protects the
|
|
74
|
+
two-space Markdown hard line break (`"foo \n"`) and trailing spaces at the end of a
|
|
75
|
+
text unit, which in `html` mode are frequently the only separator between two inline
|
|
76
|
+
elements.
|
|
77
|
+
Only a run with `CONTENT` on **both** sides is a candidate.
|
|
78
|
+
|
|
79
|
+
**In `html` and `markdown` mode, a span boundary marker counts as `NONE` here.** This is
|
|
80
|
+
the one place in the whole spec where a marker is not opaque content, it is normative, and
|
|
81
|
+
it is specified in [modes.md](modes.md) §3.3 under _Edge tests_ — read the justification
|
|
82
|
+
there before changing either document. In short: this guard exists because deleting
|
|
83
|
+
whitespace at an edge is irreversible and this rule cannot see past the edge, which is
|
|
84
|
+
exactly the situation at a span boundary; and without it `<em>mot</em> !` in `fr` loses its
|
|
85
|
+
space to the `STRIP-BEFORE` branch and gains nothing back, because `nbsp`'s replacement
|
|
86
|
+
insertion is then refused at the span edge.
|
|
87
|
+
5. **Decide the replacement length in one step** (never two passes — see §5):
|
|
88
|
+
- If the _empty-bracket guard_ (§3.3) fires → replacement length **1**.
|
|
89
|
+
- Else if `left` is in `OPEN-BRACKET` **and the emoticon guard's mouth side (§3.6) does not
|
|
90
|
+
fire** → replacement length 0.
|
|
91
|
+
- Else if `right` is in `CLOSE-BRACKET` → replacement length 0.
|
|
92
|
+
- Else if `right` is in `STRIP-BEFORE` **and the lone-dot condition (§3.4) holds and the
|
|
93
|
+
emoticon guard's eye side (§3.6) does not fire** → replacement length 0.
|
|
94
|
+
- Else → replacement length 1.
|
|
95
|
+
The guard is a clause of this decision, not a separate "skip the run" branch. **This
|
|
96
|
+
reading is normative**; see §3.3.
|
|
97
|
+
6. If the replacement length equals `k`, emit nothing (the run is already canonical).
|
|
98
|
+
Otherwise emit one edit replacing `cp[s … e-1]` with either the empty sequence or a
|
|
99
|
+
single U+0020.
|
|
100
|
+
7. Set `i = e` and go to 2.
|
|
101
|
+
|
|
102
|
+
### 3.3 The empty-bracket guard
|
|
103
|
+
|
|
104
|
+
A space run whose **removal** would produce an empty bracket pair is not removed. Concretely,
|
|
105
|
+
for a run at `[s, e)`:
|
|
106
|
+
|
|
107
|
+
- if `left` is in `OPEN-BRACKET` and `right` is the matching `CLOSE-BRACKET`
|
|
108
|
+
(`(`↔`)`, `[`↔`]`, `{`↔`}`), the guard fires.
|
|
109
|
+
|
|
110
|
+
**Normative reading: the guard forces the replacement length to 1, it does not skip the run.**
|
|
111
|
+
Two readings were possible and they differ on `"( )"` (two spaces between an empty pair):
|
|
112
|
+
"skip the run" leaves `"( )"`, "replacement length 1" collapses it to `"( )"`. The second is
|
|
113
|
+
normative. Reasons: the guard exists to prevent _deletion_, and collapsing a double space is
|
|
114
|
+
the rule's ordinary business everywhere else; a run of two spaces inside an empty bracket pair
|
|
115
|
+
carries no structural meaning in any of the four modes (the GFM task-list marker is
|
|
116
|
+
`"[ ]"` with exactly one space — `"[ ]"` is not a checkbox in any implementation); and the
|
|
117
|
+
one-clause form keeps §3.2 step 5 a single total function of the run's context — `left`, `right`
|
|
118
|
+
and the bounded lookaround of §3.4 and §3.6 — rather than a decision plus a separate skip branch,
|
|
119
|
+
which is what the idempotency argument in §5 relies on. The two-space case is a fixture.
|
|
120
|
+
|
|
121
|
+
The guard exists for one specific ship-blocking reason: the GitHub-Flavoured Markdown
|
|
122
|
+
task-list marker `"- [ ] item"`. Collapsing that space is a no-op; _deleting_ it produces
|
|
123
|
+
`"- [] item"` and silently destroys the checkbox. The guard also protects `"( )"` used as a
|
|
124
|
+
placeholder.
|
|
125
|
+
|
|
126
|
+
### 3.4 The lone-dot condition
|
|
127
|
+
|
|
128
|
+
`STRIP-BEFORE` exists to remove a space that a typist left before **one terminal punctuation
|
|
129
|
+
mark**. It must not remove a space before a _run_ of dots, because a run of dots is a different
|
|
130
|
+
kind of token — a relative path, a truncation, a typed ellipsis — and deleting the space either
|
|
131
|
+
merges it with a preceding abbreviation dot or silently destroys word spacing.
|
|
132
|
+
|
|
133
|
+
> **Lone-dot condition.** If `right` is U+002E, the replacement length may be 0 **only if** the
|
|
134
|
+
> maximal run of `DOTLIKE` code points beginning at that index has length exactly 1 **and** the
|
|
135
|
+
> code point after that dot, `cp[e+1]`, is neither a `LETTER` (ARCHITECTURE.md §4.1's Unicode
|
|
136
|
+
> category test, as in §3.6) nor an ASCII digit. Otherwise the replacement length is 1.
|
|
137
|
+
>
|
|
138
|
+
> **U+2026 is not in `STRIP-BEFORE` at all**, so a space before an existing ellipsis is never
|
|
139
|
+
> deleted.
|
|
140
|
+
|
|
141
|
+
Both halves are needed and the second is not obvious. Suppose only the first were adopted.
|
|
142
|
+
Then `Wait ...` keeps its space, `ellipsis` converts the run, and the result is `Wait …` — at
|
|
143
|
+
which point a _second_ pipeline pass finds a U+0020 before a U+2026, strips it, and yields
|
|
144
|
+
`Wait…`. That is a two-pass divergence of exactly the kind
|
|
145
|
+
[pipeline-idempotency.md](pipeline-idempotency.md) exists to prevent, introduced by the fix for
|
|
146
|
+
another one. Removing U+2026 from `STRIP-BEFORE` closes it: the author's spacing around an
|
|
147
|
+
ellipsis, however they wrote it, is preserved and is stable.
|
|
148
|
+
|
|
149
|
+
**The word-start clause (spec 1.2.0).** A single dot followed directly by a letter or a digit is
|
|
150
|
+
not terminal punctuation either: it starts a token. `.NET`, `.DWG`, `.gitignore`, `.env` and the
|
|
151
|
+
decimal `.5` are words, and before 1.2.0 the space in front of each was deleted —
|
|
152
|
+
`Use .NET, .NET Core` became `Use.NET,.NET Core`, `CAD files (.DWG, .STEP)` became
|
|
153
|
+
`CAD files (.DWG,.STEP)`. The first half of the condition could not see this, because it measures
|
|
154
|
+
only the dot run. The second half reads one code point further, the same distance and the same
|
|
155
|
+
`LETTER`/ASCII-digit test the emoticon guard's eye side already uses (§3.6 step 3).
|
|
156
|
+
|
|
157
|
+
The cost is stated here so it is not rediscovered as a bug: `end .Next sentence` — a misplaced
|
|
158
|
+
full stop with the following space also missing — keeps its stray space instead of becoming
|
|
159
|
+
`end.Next sentence`. Both readings of that input lose something, and the one given up is the
|
|
160
|
+
destructive one: deleting the space glues two words together, and a later pass cannot tell the
|
|
161
|
+
glued form from a genuine `end.Next`. Keeping it leaves the input as the author typed it. That is
|
|
162
|
+
the same principle as §3.6's trailing check and §7.11 — where deletion and preservation disagree,
|
|
163
|
+
the reading that deletes less wins. A dot followed by anything else — a space, a closing bracket,
|
|
164
|
+
a quotation mark, punctuation, a line terminator, the end of the text, or a span boundary marker
|
|
165
|
+
in `html`/`markdown` mode (which is in neither `LETTER` nor `DIGIT`, [modes.md](modes.md) §3.3) —
|
|
166
|
+
is still a terminal full stop, and the space before it is still deleted.
|
|
167
|
+
|
|
168
|
+
**What this deliberately does not break.** The runs in step 3 are maximal runs **in the input
|
|
169
|
+
array**, and the condition is evaluated against the input. In the Chicago-style spaced ellipsis
|
|
170
|
+
`Hello . . .` every dot is a _lone_ dot at the moment the decision is made — each is followed by
|
|
171
|
+
a space, not by another dot — so all three spaces are still stripped, the dots merge into
|
|
172
|
+
`Hello...`, and `ellipsis` converts them to `Hello…`. The condition costs that case nothing.
|
|
173
|
+
|
|
174
|
+
**What it does change.** `Wait ...` now becomes `Wait …` rather than `Wait…`; the space the
|
|
175
|
+
author typed survives. That is consistent with how this rule treats every other authorial
|
|
176
|
+
spacing choice (§7.5), and it is the price of not deleting the space in `See ../docs`.
|
|
177
|
+
|
|
178
|
+
### 3.5 Greek compatibility punctuation, and what "match on a code point" means
|
|
179
|
+
|
|
180
|
+
Greek raises the one case where the character a reader sees and the code point a rule matches
|
|
181
|
+
on come apart. Part 1 and part 3 below are properties of the whole rule set and are settled
|
|
182
|
+
here rather than in the locale file; part 2 is about this rule only, and says so, because the
|
|
183
|
+
generalisation is exactly what does not hold (see §7.8).
|
|
184
|
+
|
|
185
|
+
| Reader sees | Recommended code point | Compatibility code point | Canonical relation |
|
|
186
|
+
| ----------------- | ---------------------- | ------------------------ | ------------------ |
|
|
187
|
+
| ερωτηματικό (`;`) | U+003B `;` | U+037E | U+037E ≡ U+003B |
|
|
188
|
+
| άνω τελεία (`·`) | U+00B7 `·` | U+0387 | U+0387 ≡ U+00B7 |
|
|
189
|
+
|
|
190
|
+
Both compatibility characters have a **canonical** decomposition, so NFC and NFD both map them
|
|
191
|
+
away; the Unicode Standard records that **most** vendor code pages never had them, and states
|
|
192
|
+
that their use "is not generally encouraged for representation of Greek punctuation"
|
|
193
|
+
(ch. 7 §7.2.1). The consequence for this spec is that a Greek question mark and a Latin
|
|
194
|
+
semicolon are, at the level a rule operates on, **the same code point** — U+003B — and no
|
|
195
|
+
amount of context lets a single scan distinguish them.
|
|
196
|
+
|
|
197
|
+
**The normative position, in three parts:**
|
|
198
|
+
|
|
199
|
+
1. **Every rule matches on the literal code points enumerated in its own class tables, and on
|
|
200
|
+
nothing else.** No rule consults canonical equivalence, no rule normalises its input, and no
|
|
201
|
+
rule has a notion of "the same character spelled differently" (ARCHITECTURE.md §4.3). A class
|
|
202
|
+
table is a list of code points, not a list of characters.
|
|
203
|
+
|
|
204
|
+
2. **For _this_ rule, the U+003B ambiguity is harmless.** U+003B is in `STRIP-BEFORE`. Read as a
|
|
205
|
+
Latin semicolon it takes no preceding space in any locale polytypo supports; read as a Greek
|
|
206
|
+
ερωτηματικό nothing in the Greek sources asks for one either — the nearest statements are
|
|
207
|
+
about the neighbouring marks and deny the French pattern by name for both (`Στα ελληνικά,
|
|
208
|
+
πριν από τη διπλή τελεία δεν πρέπει να υπάρχει διάστημα (πράγμα που συμβαίνει, π.χ., στα
|
|
209
|
+
γαλλικά)`, and the same sentence again for the άνω τελεία). Both readings therefore prescribe
|
|
210
|
+
the same edit and this rule never has to choose: `Τι κάνεις ;` becomes `Τι κάνεις;` whichever
|
|
211
|
+
character the author meant. Note the standard of evidence, because it is weaker than it looks
|
|
212
|
+
— no source examined addresses the ερωτηματικό _specifically_, so the claim is "no source
|
|
213
|
+
contradicts it", not "a source requires it". Nothing rests on the difference here, since
|
|
214
|
+
U+003B would be in `STRIP-BEFORE` for the Latin reading alone.
|
|
215
|
+
|
|
216
|
+
**This does not generalise to the other rules, and §7.8 records where it fails.** The
|
|
217
|
+
ambiguity is harmless in `spaces` because the two readings happen to agree; that is a fact
|
|
218
|
+
about the two conventions, not a property the rule set enjoys everywhere. `ellipsis` is the
|
|
219
|
+
counter-example: its `TERMINAL` set is `{U+0021, U+003F}`, so the Latin reading of U+003B
|
|
220
|
+
wins silently and `Πράγματι;..` — the genuinely Greek spelling — does not get the treatment
|
|
221
|
+
`Πράγματι?..` gets. **Any rule adding U+003B, U+0021 or U+003F to a class table must decide
|
|
222
|
+
the Greek reading explicitly rather than inheriting this paragraph.**
|
|
223
|
+
|
|
224
|
+
3. **U+037E, U+0387 and U+00B7 are in no class of any rule, and no rule ever emits them.** Text
|
|
225
|
+
already written with a compatibility code point round-trips **untouched** — it is neither
|
|
226
|
+
converted to the recommended spelling nor treated as the character it decomposes to. Three
|
|
227
|
+
separate reasons, each sufficient:
|
|
228
|
+
- Rewriting U+037E → U+003B or U+0387 → U+00B7 **is** normalisation, whatever it is called,
|
|
229
|
+
and §4.3 forbids it. It is also redundant: any downstream consumer applying NFC does it.
|
|
230
|
+
- **U+00B7 must not join `STRIP-BEFORE`.** It is a deliberate spaced separator in real
|
|
231
|
+
content — `Home · About · Contact`, and the Catalan and French interpunct — and stripping
|
|
232
|
+
there would be a false positive in locales that have nothing to do with Greek. This rule
|
|
233
|
+
reads no locale data (§2), so a Greek-only membership is not available to it.
|
|
234
|
+
- **Adding only U+0387 would be the worst of the three options.** Two canonically equivalent
|
|
235
|
+
characters would behave differently, and the one that got the correct treatment would be
|
|
236
|
+
the spelling Unicode discourages, while the recommended U+00B7 kept its stray space. A
|
|
237
|
+
split that punishes the correct input is not a partial fix.
|
|
238
|
+
|
|
239
|
+
The residual cost is exact and small: a space before an άνω τελεία survives. It is a rare
|
|
240
|
+
typo, and the alternative costs other locales real false positives.
|
|
241
|
+
|
|
242
|
+
An earlier revision added, in this sentence, that "the recommended Greek keyboard layouts
|
|
243
|
+
produce U+00B7 and U+003B rather than the compatibility pair". **That clause has been
|
|
244
|
+
deleted.** It is an empirical claim about keyboard layouts, it was stated as fact in
|
|
245
|
+
normative prose, and no source was offered for it — it was the only sentence in this section
|
|
246
|
+
a reviewer could not back with something readable. It may well be true; CLDR keyboard data or
|
|
247
|
+
a vendor layout specification would settle it. Until one is cited the argument does not need
|
|
248
|
+
it, which is why deleting was cheaper than sourcing.
|
|
249
|
+
|
|
250
|
+
### 3.6 The emoticon guard
|
|
251
|
+
|
|
252
|
+
`STRIP-BEFORE`'s two punctuation-adjacent members U+003A (:) and U+003B (;) are read two ways
|
|
253
|
+
in ordinary text: as sentence punctuation (`Note: read this`, `Wait; think`) and as the eye of a
|
|
254
|
+
Western text emoticon (`:-)`, `:)`, `;-)`). The two readings take opposite spacing: sentence
|
|
255
|
+
punctuation never wants a preceding space (hence `STRIP-BEFORE`), but an emoticon is a token in
|
|
256
|
+
its own right and the space before it is ordinary word spacing that must survive — `Привет :-)`
|
|
257
|
+
must not become `Привет:-)`.
|
|
258
|
+
|
|
259
|
+
Two of the mouth glyphs, U+0028 `(` and U+005B `[`, are also `OPEN-BRACKET` members, so the same
|
|
260
|
+
token has to be protected from the other end as well: the space **after** the mouth is ordinary
|
|
261
|
+
word spacing too, and the opening-bracket clause of §3.2 step 5 must not eat it. The guard
|
|
262
|
+
therefore has two sides. They recognise the same shape — `EMOTICON-EYE`, optional
|
|
263
|
+
`EMOTICON-NOSE`, `EMOTICON-MOUTH` — and differ only in which end of the space run it sits at, and
|
|
264
|
+
in which clause of step 5 they suppress.
|
|
265
|
+
|
|
266
|
+
> **Emoticon guard, eye side.** For a run whose `right` (at index `e`) is in `EMOTICON-EYE`, walk
|
|
267
|
+
> forward:
|
|
268
|
+
>
|
|
269
|
+
> 1. Let `i = e + 1`. If `cp[i]` is in `EMOTICON-NOSE`, set `i = i + 1`.
|
|
270
|
+
> 2. If `cp[i]` is not in `EMOTICON-MOUTH`, the guard does not fire.
|
|
271
|
+
> 3. Otherwise let `after = cp[i + 1]` (or `NONE`). If `after` is a `LETTER` (ARCHITECTURE.md
|
|
272
|
+
> §4.1's Unicode category test, as used throughout this spec) or an ASCII digit, the guard
|
|
273
|
+
> does not fire. Otherwise **the guard fires**, and the `STRIP-BEFORE` clause of §3.2 step 5
|
|
274
|
+
> does not strip the space.
|
|
275
|
+
|
|
276
|
+
> **Emoticon guard, mouth side.** For a run whose `left` (at index `s-1`) is in `EMOTICON-MOUTH`,
|
|
277
|
+
> walk backward:
|
|
278
|
+
>
|
|
279
|
+
> 1. Let `j = s - 2`. If `cp[j]` is in `EMOTICON-NOSE`, set `j = j - 1`.
|
|
280
|
+
> 2. If `cp[j]` is not in `EMOTICON-EYE`, the guard does not fire.
|
|
281
|
+
> 3. Otherwise **the guard fires**, and the `OPEN-BRACKET` clause of §3.2 step 5 does not delete
|
|
282
|
+
> the run.
|
|
283
|
+
>
|
|
284
|
+
> There is no trailing check on this side and none is needed: the code point after the mouth is
|
|
285
|
+
> the space run itself, which is neither a `LETTER` nor an ASCII digit, so the eye side's step 3
|
|
286
|
+
> is satisfied here by construction.
|
|
287
|
+
|
|
288
|
+
**What the mouth side fixes.** Without it the two clauses of step 5 contradicted each other about
|
|
289
|
+
the same character: the eye side recognised `(` as a mouth while the opening-bracket clause went on
|
|
290
|
+
reading it as a bracket. `a :( b` became `a :(b`, gluing the next word onto the emoticon — a false
|
|
291
|
+
positive of exactly the kind the M4 gate (`docs/ROADMAP.md`) forbids — and the damage propagated,
|
|
292
|
+
because with no space after the mouth `after` is a `LETTER`, the eye side stops firing, and a second
|
|
293
|
+
pass strips the space in front of the eye as well: `a:(b`. That is a two-pass divergence, a release
|
|
294
|
+
blocker under [pipeline-idempotency.md](pipeline-idempotency.md), and the eye side's own idempotency
|
|
295
|
+
claim was what had been wrong.
|
|
296
|
+
|
|
297
|
+
**Only the `OPEN-BRACKET` clause is suppressed.** The wider reading — "a run abutting a recognised
|
|
298
|
+
emoticon is always length 1" — was considered and rejected as broader than the defect. It would also
|
|
299
|
+
silence the `CLOSE-BRACKET` clause (`a :) )` would stay as typed instead of becoming `a :))`) and the
|
|
300
|
+
`STRIP-BEFORE` clause (`a :) , b`), neither of which damages the emoticon or the words around it,
|
|
301
|
+
and each of which would need its own composition argument. The mouth side is the smallest change
|
|
302
|
+
that closes the defect.
|
|
303
|
+
|
|
304
|
+
**Why the trailing check.** Without it, `:D` inside an ordinary word — `:Deal with it`,
|
|
305
|
+
`;Design review` — would read as an emoticon and keep a space that sentence punctuation never
|
|
306
|
+
wants. The mouth must be the end of a token, not the start of a capitalised word: `after` is
|
|
307
|
+
checked against `LETTER` and `DIGIT`, not against `SPACE` or `NONE` specifically, so `:-)!`
|
|
308
|
+
(mouth followed by punctuation) and `:-):-)` (two emoticons back to back) both still fire, and
|
|
309
|
+
`10:30` (colon before a digit, no mouth at all — the mouth check in step 2 already declines it
|
|
310
|
+
before this check is reached) and `:Deal` do not.
|
|
311
|
+
|
|
312
|
+
**Scope, deliberately narrow.** This guard recognises the eye-nose-mouth shape of a Western
|
|
313
|
+
text emoticon and nothing else: no East Asian kaomoji (`(^_^)`, whose parenthesis is the frame,
|
|
314
|
+
not the eye), no `=)` (U+003D is not in `STRIP-BEFORE` and has no bug to fix), no emoji. It
|
|
315
|
+
exists because `STRIP-BEFORE`'s membership of U+003A and U+003B was already normative and
|
|
316
|
+
locale-independent (§2), and an emoticon eye is the one shape that class was silently getting
|
|
317
|
+
wrong; it is not a general-purpose emoticon detector and does not try to be one.
|
|
318
|
+
|
|
319
|
+
**Both sides require an eye.** The backward walk is unambiguous because `EMOTICON-NOSE` and
|
|
320
|
+
`EMOTICON-EYE` are disjoint, and a nose on its own is not a face: `a -( b` still becomes `a -(b`
|
|
321
|
+
under the ordinary opening-bracket clause.
|
|
322
|
+
|
|
323
|
+
**Idempotency.** An earlier revision claimed here that the guard "reads only `cp[e]`, `cp[e+1]` and
|
|
324
|
+
`cp[e+2]`, none of which this rule … ever modifies". Both halves were false: with a nose the eye
|
|
325
|
+
side reads through `cp[e+3]`, and step 3's `after` is, in precisely the shape that broke, the U+0020
|
|
326
|
+
run step 5 was about to delete — so the verdict was not a pure function of code points this rule
|
|
327
|
+
does not write. The correct argument has one part per side.
|
|
328
|
+
|
|
329
|
+
- **Mouth side.** It reads `cp[s-1]`, `cp[s-2]` and `cp[s-3]`, and in no position it accepts is that
|
|
330
|
+
a U+0020: an eye, a nose and a mouth are all `CONTENT` and must be adjacent. This rule never
|
|
331
|
+
inserts a code point and never removes a non-space one, so a shape it recognised in the input is
|
|
332
|
+
still contiguous and unchanged in the output. The verdict can therefore only move from "does not
|
|
333
|
+
fire" to "fires" — never the reverse — and firing only ever preserves a run, so neither direction
|
|
334
|
+
can produce a second-pass edit.
|
|
335
|
+
- **Eye side.** Whenever the eye side fires, the mouth side recognises the same three code points
|
|
336
|
+
from the other end: the two walks read `EMOTICON-NOSE` and `EMOTICON-EYE`, which are disjoint, so
|
|
337
|
+
the backward walk's optional step over a nose cannot swallow the eye and the two sides cannot
|
|
338
|
+
disagree about the shape. A run whose `left` is that mouth is therefore protected from the
|
|
339
|
+
`OPEN-BRACKET` clause, and the only clauses of step 5 that can still delete it are the
|
|
340
|
+
`CLOSE-BRACKET` clause and the `STRIP-BEFORE` clause. The code point that comes to sit after the
|
|
341
|
+
mouth is then one of `)` `]` `}` or one of `STRIP-BEFORE`'s six; none of those nine is a `LETTER`
|
|
342
|
+
or an ASCII digit, so step 3 still passes on the re-run. Every other outcome — the empty-bracket
|
|
343
|
+
guard, a collapse to length 1, a run skipped by the boundary guard of §3.2 step 4 — leaves the
|
|
344
|
+
U+0020 in place, so `after` is unchanged. Either way the eye side's verdict survives the re-run.
|
|
345
|
+
|
|
346
|
+
### 3.7 Worked trace
|
|
347
|
+
|
|
348
|
+
`"a ( b , c ) d"`
|
|
349
|
+
|
|
350
|
+
| run | left | right | decision |
|
|
351
|
+
| ------ | ---- | ----- | -------------------------------------------------------- |
|
|
352
|
+
| `a␣␣(` | `a` | `(` | not bracket-inner, right not in `STRIP-BEFORE` → 1 space |
|
|
353
|
+
| `(␣␣b` | `(` | `b` | left is `OPEN-BRACKET`, guard does not fire → 0 |
|
|
354
|
+
| `b␣␣,` | `b` | `,` | right in `STRIP-BEFORE` → 0 |
|
|
355
|
+
| `,␣␣c` | `,` | `c` | → 1 space |
|
|
356
|
+
| `c␣␣)` | `c` | `)` | right is `CLOSE-BRACKET` → 0 |
|
|
357
|
+
| `)␣␣d` | `)` | `d` | → 1 space |
|
|
358
|
+
|
|
359
|
+
Result: `"a (b, c) d"`.
|
|
360
|
+
|
|
361
|
+
---
|
|
362
|
+
|
|
363
|
+
## 4. Must not touch
|
|
364
|
+
|
|
365
|
+
**Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
|
|
366
|
+
a guarantee of `transform` as a whole — or **[R]** — true of this rule alone and capable of
|
|
367
|
+
being falsified by another rule. This rule is R₁, so no _earlier_ rule can invalidate anything
|
|
368
|
+
here; but later rules touch some of the same characters, and those bullets are marked [R].
|
|
369
|
+
|
|
370
|
+
- **[R] Any character other than U+0020.** Tabs (U+0009) are never collapsed, never converted,
|
|
371
|
+
never removed — in Markdown a tab is indentation and the engine has no way to know whether
|
|
372
|
+
it is inside a code block. U+00A0 and U+202F are never collapsed, never removed, and never
|
|
373
|
+
converted to U+0020; a run such as `"a b"` is left exactly as written.
|
|
374
|
+
U+2000–U+200A, U+205F, U+3000, U+200B and U+FEFF are likewise untouched.
|
|
375
|
+
_[R]: `nbsp` (R₈) may convert an existing U+00A0 to U+202F or the reverse. `transform` does
|
|
376
|
+
not promise an existing no-break space keeps its exact width — only that this rule never
|
|
377
|
+
collapses or deletes one._
|
|
378
|
+
- **[P] Line terminators.** No character in `BREAK` is inserted, removed, or reordered. Space
|
|
379
|
+
runs never merge across a line terminator, because a `BREAK` on either side aborts the
|
|
380
|
+
candidate.
|
|
381
|
+
- **[R] Leading whitespace on a line** (a run at index 0 or directly after a `BREAK`). Markdown
|
|
382
|
+
indented code blocks, nested list indentation and YAML front matter all depend on it.
|
|
383
|
+
_[R]: `nbsp` (R₈) can convert the **last** space of an indentation run when the character
|
|
384
|
+
after it is listed in `beforePunctuation`/`narrowBeforePunctuation` — a French line beginning
|
|
385
|
+
`␣␣?`. The run is never shortened, so indentation width survives; one code point changes
|
|
386
|
+
class._
|
|
387
|
+
- **[P] Trailing whitespace before a line terminator or at the end of the text unit.** This is
|
|
388
|
+
the Markdown hard-break idiom and, in `html` mode, the inter-element separator.
|
|
389
|
+
- **[R] Single spaces in ordinary positions.** `"a b"` is not a candidate for anything.
|
|
390
|
+
- **[P] `"[ ]"`, `"( )"`, `"{ }"`** — see §3.3.
|
|
391
|
+
- **[R] Spaces before a closing quotation mark or guillemet.** Owned by `nbsp` via
|
|
392
|
+
`quotes.innerSpace`.
|
|
393
|
+
- **[R] A space before an emoticon's eye, and a space after an emoticon's mouth.** §3.6.
|
|
394
|
+
`Привет :-)` keeps its space in every locale, and so does `Привет :-( снова`, whose second space
|
|
395
|
+
would otherwise go to the opening-bracket clause. This rule reads no locale data (§2) and neither
|
|
396
|
+
side of the guard is locale-dependent either.
|
|
397
|
+
_[R]: `nbsp` (R₈) may convert the space **before** the eye to a no-break form in a locale whose
|
|
398
|
+
`nbsp.beforePunctuation` lists U+003A or U+003B — French `a :) b` yields `a⍽:) b`, because
|
|
399
|
+
[nbsp.md](nbsp.md) §3.3 step 2's right-context guard admits `CLOSEISH` after the mark. The run is never
|
|
400
|
+
deleted, so word spacing survives; one code point changes class. The space **after** the mouth
|
|
401
|
+
is touched by no later rule._
|
|
402
|
+
- **[P] U+037E, U+0387, U+00B7.** In no class of this rule and of no other. A space before an
|
|
403
|
+
άνω τελεία survives, and a Greek compatibility code point is never rewritten to the character
|
|
404
|
+
it canonically decomposes to — §3.5. Note that the Greek ερωτηματικό is _not_ an exception
|
|
405
|
+
here: it is written U+003B, which is in `STRIP-BEFORE` on its own merits.
|
|
406
|
+
- **[P] Anything inside a skipped region.** Code spans, fenced code, `<pre>`, attributes and
|
|
407
|
+
URLs never reach this rule; the mode adapter (L2) removes them before the pipeline runs.
|
|
408
|
+
This rule contains no code-awareness of its own and must not grow any.
|
|
409
|
+
|
|
410
|
+
---
|
|
411
|
+
|
|
412
|
+
## 5. Idempotency argument
|
|
413
|
+
|
|
414
|
+
Let `T` be the transformation described in §3.2.
|
|
415
|
+
|
|
416
|
+
Every edit replaces a maximal `SPACE` run bounded by `CONTENT` on both sides with either
|
|
417
|
+
zero or one U+0020. Consider the output `T(x)` and re-run the scan.
|
|
418
|
+
|
|
419
|
+
- The bounding characters of every run are `CONTENT` and are never modified by this rule, so
|
|
420
|
+
the `left`/`right` classification of any surviving run is unchanged between runs.
|
|
421
|
+
- A run replaced by zero spaces no longer exists; its former neighbours `left` and `right`
|
|
422
|
+
are now adjacent. `right` is a member of `STRIP-BEFORE` or a bracket, and `left` is
|
|
423
|
+
`CONTENT`. No new `SPACE` run has been created — deletion cannot create a space — so
|
|
424
|
+
there is nothing to re-examine at that position.
|
|
425
|
+
- A run replaced by one space is now a run of length `k' = 1` with the same `left` and
|
|
426
|
+
`right`. Re-running step 5 on it yields the same decision — it is a function of `left`, `right`
|
|
427
|
+
and the bounded lookaround of §3.4 and §3.6, each of which argues its own window's stability
|
|
428
|
+
(§3.4's word-start clause: see the paragraph after this list) —
|
|
429
|
+
namely replacement length 1, and step 6 then emits nothing because `k' = 1` already equals the
|
|
430
|
+
replacement length.
|
|
431
|
+
- A skipped run is skipped again for the same reason (its guard condition depends only on
|
|
432
|
+
`left`, `right` and bracket matching, all unchanged).
|
|
433
|
+
|
|
434
|
+
**§3.4's window.** The lone-dot condition reads `cp[e]`, the dot run starting there, and — since
|
|
435
|
+
1.2.0 — `cp[e+1]`. The dot run is made of non-space code points this rule never writes. `cp[e+1]`
|
|
436
|
+
can change between passes only if it was a U+0020 whose run this pass deleted, bringing its right
|
|
437
|
+
neighbour against the dot. A run is deleted only by the `STRIP-BEFORE`, `CLOSE-BRACKET` and
|
|
438
|
+
`OPEN-BRACKET` clauses of step 5; the run after a dot has the dot as its `left`, so the
|
|
439
|
+
`OPEN-BRACKET` clause cannot apply, and the other two leave behind a `STRIP-BEFORE` member or a
|
|
440
|
+
closing bracket — none of which is a `LETTER` or an ASCII digit. So the word-start clause's verdict
|
|
441
|
+
("is `cp[e+1]` a letter or a digit?") is the same on both passes, whatever this pass did after the
|
|
442
|
+
dot.
|
|
443
|
+
|
|
444
|
+
Therefore `T(T(x)) = T(x)`.
|
|
445
|
+
|
|
446
|
+
**What had to be fixed to get here.** The naive formulation — "first collapse doubles, then
|
|
447
|
+
strip spaces before punctuation" — is two passes and is _not_ obviously idempotent, and worse,
|
|
448
|
+
it is not obviously order-independent: `"a , b"` collapses to `"a , b"` and then strips to
|
|
449
|
+
`"a, b"`, which requires the second pass to run over the output of the first, i.e. a
|
|
450
|
+
fixed-point loop. Fixed-point loops are exactly what a spec must not require, because two
|
|
451
|
+
implementations will disagree about how many iterations they run. The formulation above
|
|
452
|
+
computes the replacement length for each run **once**, from `left` and `right` alone, so a
|
|
453
|
+
single pass reaches the fixed point directly.
|
|
454
|
+
|
|
455
|
+
The second thing that had to be fixed: the naive rule "collapse every run of ≥2 spaces" is
|
|
456
|
+
not merely non-idempotent-adjacent, it is _destructive_ on Markdown hard breaks and on
|
|
457
|
+
indentation. The `CONTENT`-on-both-sides boundary guard (step 4) is what makes the rule safe,
|
|
458
|
+
and it is a precondition of the argument above, not an optimisation.
|
|
459
|
+
|
|
460
|
+
---
|
|
461
|
+
|
|
462
|
+
### Composition obligation
|
|
463
|
+
|
|
464
|
+
Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₁**: nothing runs
|
|
465
|
+
before it, so **the obligation is empty**. There is no earlier invariant it could break.
|
|
466
|
+
|
|
467
|
+
The obligation pointing the other way is not empty, and it is the one that bit. `I₁` — the
|
|
468
|
+
statement that this rule is a no-op, spelled out as S-a … S-d in that document §3 — must be
|
|
469
|
+
preserved by all seven later rules. Only one of them emits U+0020 at all (`dashes`), and it
|
|
470
|
+
violated S-b, S-c and S-d until guard T2 was added. The exact positions from which this rule
|
|
471
|
+
deletes a space are therefore load-bearing for the whole pipeline, and §3.2 step 5 should be
|
|
472
|
+
treated as a published interface rather than an implementation detail.
|
|
473
|
+
|
|
474
|
+
**Spec 1.2.0's word-start clause adds one thing a later rule must not do.** S-b now permits a
|
|
475
|
+
U+0020 before a lone dot whose next code point is a `LETTER` or an ASCII digit. A later rule that
|
|
476
|
+
replaced that letter or digit with something else would turn a permitted space into a forbidden one,
|
|
477
|
+
without emitting any U+0020. None does: `ellipsis` writes only `DOTLIKE` code points; `ranges`,
|
|
478
|
+
`dashes` and `hyphen` replace dashes, hyphens and spaces; `quotes` and `apostrophe` replace
|
|
479
|
+
quotation marks; `symbols` replaces a trademark literal, which starts with `(`, and a `MUL-LETTER`,
|
|
480
|
+
which is always preceded by a digit or a space and so never follows a dot directly; `nbsp` replaces
|
|
481
|
+
or inserts spaces, and inserts only beside a listed punctuation mark or a quote glyph, never between
|
|
482
|
+
a dot and a letter. A new rule that rewrites letters or digits must re-check this.
|
|
483
|
+
|
|
484
|
+
---
|
|
485
|
+
|
|
486
|
+
## 6. Worked examples
|
|
487
|
+
|
|
488
|
+
`␣` = U+0020, `⟶` = no change expected, `↵` = U+000A, `⍽` = U+00A0. Every row is
|
|
489
|
+
locale-independent **except 9b**, which is marked with its locale: the rule reads no locale data
|
|
490
|
+
(§2), but that row's point is that a Greek reading and a Latin reading of U+003B reach the same
|
|
491
|
+
verdict here, so naming the locale is what makes the claim checkable.
|
|
492
|
+
|
|
493
|
+
| # | Input | Output | Why |
|
|
494
|
+
| --- | ---------------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
495
|
+
| 1 | `Hello␣␣␣world.` | `Hello␣world.` | run of 3 with `CONTENT` both sides → 1 |
|
|
496
|
+
| 2 | `Hello␣,␣world␣!` | `Hello,␣world!` | `,` and `!` are in `STRIP-BEFORE`; the run after `,` keeps one space |
|
|
497
|
+
| 3 | `(␣ok␣)␣and␣[␣x␣]` | `(ok)␣and␣[x]` | bracket-inner runs deleted; the guard does not fire because the inner side is not the matching closer |
|
|
498
|
+
| 4 | `-␣[␣]␣buy␣milk` | ⟶ | empty-bracket guard (§3.3): replacement length 1, and the run is already length 1 → no edit. The GFM checkbox survives |
|
|
499
|
+
| 4b | `(␣␣)` | `(␣)` | empty-bracket guard forces length 1, it does not skip — §3.3, normative reading |
|
|
500
|
+
| 5 | `line␣one␣␣↵line␣two` | ⟶ | the two-space run is directly followed by `BREAK` → skipped; the Markdown hard break survives |
|
|
501
|
+
| 6 | `␣␣␣␣indented␣code` | ⟶ | run starts at index 0 → skipped |
|
|
502
|
+
| 7 | `5⍽␣␣km` | `5⍽␣km` | the U+00A0 is `CONTENT`; only the U+0020 run beside it collapses, and the U+00A0 itself is untouched |
|
|
503
|
+
| 8 | `a⍽⍽b` | ⟶ | no U+0020 anywhere; two no-break spaces are never collapsed |
|
|
504
|
+
| 9 | `Bonjour␣!␣Ça␣va␣?` | `Bonjour!␣Ça␣va?` | French input; the ordinary spaces go, and `nbsp` (order 70) later restores `Bonjour<U+202F>!` |
|
|
505
|
+
| 9b | `Τι␣κάνεις␣;` | `Τι␣κάνεις;` | **`el`.** The Greek question mark is written U+003B, which is in `STRIP-BEFORE` on the Latin semicolon's own merits, so the space is stripped with no Greek-specific decision — §3.5 part 2. `nbsp` puts nothing back, because `el`'s punctuation lists are empty |
|
|
506
|
+
| 10 | `See␣p.␣12␣.` | `See␣p.␣12.` | run before the final `.` deleted — a lone dot, so the condition holds |
|
|
507
|
+
| 10a | `See␣../docs` | ⟶ | **lone-dot condition (§3.4).** The dot run has length 2, so the space survives. Previously produced `See../docs`, silently destroying word spacing in a relative path |
|
|
508
|
+
| 10b | `e.g.␣..` | ⟶ | same. Previously produced `e.g...`, which `ellipsis` then legitimately read as a three-dot run and converted to `e.g…` |
|
|
509
|
+
| 10c | `Wait␣...` | ⟶ | the space survives (run length 3). `ellipsis` then yields `Wait␣…`, and because U+2026 is not in `STRIP-BEFORE` that is a fixed point |
|
|
510
|
+
| 10d | `Hello␣.␣.␣.` | `Hello...` | every dot is a **lone** dot in the input, so all three spaces strip and the Chicago-style spaced ellipsis still merges — `ellipsis` converts it to `Hello…` |
|
|
511
|
+
| 10e | `Wait␣…` | ⟶ | U+2026 is not in `STRIP-BEFORE` |
|
|
512
|
+
| 10f | `Use␣.NET,␣.NET␣Core` | ⟶ | **word-start clause (§3.4, spec 1.2.0).** Each dot is followed by a letter, so it starts a word and neither space is deleted. Previously produced `Use.NET,.NET␣Core` |
|
|
513
|
+
| 10g | `(.DWG,␣.STEP)` | ⟶ | same: the space after the comma survives because `.STEP` is a token, not a full stop. The comma run is untouched regardless — its `right` is the dot, not the comma |
|
|
514
|
+
| 10h | `from␣.5␣to␣.9` | ⟶ | same, for an ASCII digit after the dot |
|
|
515
|
+
| 10i | `end␣.Next` | ⟶ | **the accepted cost.** A misplaced full stop followed by a letter keeps its stray space; the alternative glues two words — §3.4 |
|
|
516
|
+
| 10j | `See␣p.␣12␣.␣Next` | `See␣p.␣12.␣Next` | a dot followed by a space is still terminal punctuation, so row 10 is unchanged by the word-start clause |
|
|
517
|
+
| 11 | `foo␣␣␣␣↵␣␣␣␣bar` | ⟶ | first run touches a `BREAK` on the right, second on the left |
|
|
518
|
+
| 12 | `Q:␣␣why␣?␣␣Because␣.` | `Q:␣why?␣Because.` | mixed |
|
|
519
|
+
| 13 | `Привет␣:-)` | ⟶ | **emoticon guard (§3.6).** `:` is `EMOTICON-EYE`, `-` is `EMOTICON-NOSE`, `)` is `EMOTICON-MOUTH`, and there is nothing after it — the guard fires and the space survives |
|
|
520
|
+
| 13a | `Привет␣:)` | ⟶ | same, no nose |
|
|
521
|
+
| 13b | `Hello␣:Deal␣with␣it` | `Hello:Deal␣with␣it` | `D` is `EMOTICON-MOUTH`, but `e` immediately after it is a `LETTER` — the guard does not fire, and ordinary `STRIP-BEFORE` behaviour applies |
|
|
522
|
+
| 13c | `10:30` | ⟶ | no space run at all — outside this rule's scope regardless of the guard |
|
|
523
|
+
| 13d | `See␣you␣at␣10␣:␣30` | `See␣you␣at␣10:␣30` | `:` is `EMOTICON-EYE`, but the very next code point is a space, not `EMOTICON-MOUTH` — the guard does not fire; ordinary `STRIP-BEFORE` strips the leading space, and the trailing space (right neighbour `3`, not `STRIP-BEFORE`) is untouched |
|
|
524
|
+
| 13e | `Sorry␣:(␣it␣happens` | ⟶ | **mouth side (§3.6).** `(` is the mouth of a recognised emoticon, so the opening-bracket clause does not delete the space after it. Previously produced `Sorry␣:(it␣happens`, and then `Sorry:(it␣happens` on a second pass |
|
|
525
|
+
| 13f | `Hmm␣:[␣well` | ⟶ | same, for the other mouth that is also an `OPEN-BRACKET` member — `[` |
|
|
526
|
+
| 13g | `Well␣:-(␣then` | ⟶ | same, with a nose: the backward walk steps over `-` and finds the eye at `cp[s-3]` |
|
|
527
|
+
| 13h | `a␣-(␣b` | `a␣-(b` | a nose with no eye behind it is not a face — the mouth side does not fire and the opening-bracket clause applies as usual |
|
|
528
|
+
| 13i | `word␣(␣note␣)` | `word␣(note)` | the ordinary bracket-inner case, unchanged by the mouth side: `(` here is preceded by a space, not by an eye |
|
|
529
|
+
| 13j | `Note:(␣x␣)` | `Note:(␣x)` | the mouth side puts **no** condition on what precedes the eye, so it fires on a word-attached eye too: the space after the mouth survives while the one before the closer still goes — §7 item 11 |
|
|
530
|
+
|
|
531
|
+
Cases 4, 5, 6, 8, 10a, 10b, 10c, 10e, 10f, 10g, 10h, 10i, 11, 13, 13a, 13e, 13f and 13g are "no change" cases.
|
|
532
|
+
|
|
533
|
+
---
|
|
534
|
+
|
|
535
|
+
## 7. Open questions
|
|
536
|
+
|
|
537
|
+
1. **U+0009 (tab) is completely untouched.** In `text` mode a run of tabs between two words
|
|
538
|
+
is arguably the same typing accident as a run of spaces. I chose safety, because the rule
|
|
539
|
+
cannot distinguish prose from Markdown indentation. If the dogfooding gate (PLAN.md M4)
|
|
540
|
+
shows tab noise in real content, the fix is a separate opt-in rule, not a change here.
|
|
541
|
+
2. **U+2009 (thin space) and friends are untouched.** Some authors paste text from InDesign
|
|
542
|
+
or LaTeX carrying real thin spaces. Normalising them to U+202F would be locale-dependent
|
|
543
|
+
and is arguably a `nbsp` concern. Unresolved; currently they are simply preserved.
|
|
544
|
+
3. **`STRIP-BEFORE` includes U+003A (colon).** `"10 : 30"` becomes `"10:30"`. I believe this
|
|
545
|
+
is right for prose but it is a real behaviour change on tabular text. Needs a fixture
|
|
546
|
+
decision from the operator.
|
|
547
|
+
4. **`STRIP-BEFORE` excludes U+2014/U+2013.** A spaced em dash is a legitimate parenthetical
|
|
548
|
+
form in several locales, so stripping there would be wrong; the `dashes` rule owns dash
|
|
549
|
+
spacing. Confirmed by construction, but worth a fixture.
|
|
550
|
+
5. **Should a space before an opening bracket be normalised?** `"word(note)"` versus
|
|
551
|
+
`"word (note)"` is an authorial choice, not a typographic error, so this rule does
|
|
552
|
+
nothing. Recorded so nobody adds it later "for symmetry".
|
|
553
|
+
6. **`html` mode boundary semantics.** A text node ending in a space followed by a sibling
|
|
554
|
+
text node beginning with a space is, at the DOM level, two runs of length 1 that render
|
|
555
|
+
as one collapsed space. This rule sees them separately and leaves both. Whether the mode
|
|
556
|
+
adapter should present adjacent text nodes as one logical unit is an L2 question that
|
|
557
|
+
this document cannot settle; it is flagged here because the answer changes `spaces`
|
|
558
|
+
fixtures for `html` mode.
|
|
559
|
+
7. **The lone-dot condition (§3.4) changes `Wait␣...` to `Wait␣…` rather than `Wait…`.** The
|
|
560
|
+
author's space survives. I judge that correct — it is the same principle as §7.5, and it is
|
|
561
|
+
what makes `See␣../docs` safe — but it is a visible change from the previous behaviour and
|
|
562
|
+
it deserves a fixture and an operator glance. If tight is preferred, the fix is **not** to
|
|
563
|
+
restore stripping before a dot run (that reopens `See␣../docs`) but to have `ellipsis`
|
|
564
|
+
absorb a preceding space, which is a different rule's business and a different decision.
|
|
565
|
+
8. **The Greek reading of U+003B is honoured here and ignored by `ellipsis`.** §3.5 part 2 is
|
|
566
|
+
true of this rule and does not generalise; this entry is where the generalisation fails, and
|
|
567
|
+
it was found by review rather than by construction, so assume there are others.
|
|
568
|
+
`ellipsis.md` §3.1 defines `TERMINAL` as `{U+0021, U+003F}`. A Greek author writes a question
|
|
569
|
+
with U+003B, so `Πράγματι;..` keeps its two-dot run while `Πράγματι?..` — which uses a
|
|
570
|
+
character Greek does not use for questions — becomes `Πράγματι?…`. Greek gets the treatment
|
|
571
|
+
only when it is written wrongly. Verified against the engine, and pinned by the fixture
|
|
572
|
+
`el-ellipsis-greek-question-mark-not-terminal`.
|
|
573
|
+
|
|
574
|
+
It is a **miss, not damage**: the input round-trips unchanged, and the two-dot run is left
|
|
575
|
+
exactly as written. That is why it is recorded rather than fixed here.
|
|
576
|
+
|
|
577
|
+
**The obvious fix — adding U+003B to `TERMINAL` — was researched and is REFUSED.** This is a
|
|
578
|
+
settled decision, not an open question, and it is recorded here so it is not reopened:
|
|
579
|
+
|
|
580
|
+
- **It would be a regression in Russian.** The abbreviated two-dot form is tied to two named
|
|
581
|
+
marks, not to a class of "terminal punctuation". Лопатин, «Правила русской орфографии и
|
|
582
|
+
пунктуации. Полный академический справочник» §154: «При сочетании вопросительного или
|
|
583
|
+
восклицательного знака с многоточием знаки эти ставятся на месте первой точки» — and
|
|
584
|
+
§§155–158, which cover every other combination, never mention the semicolon. Розенталь
|
|
585
|
+
§68.1 says the same. So `текст;...` must stay `текст;…` in Russian, and putting U+003B in
|
|
586
|
+
`TERMINAL` would invent a rule that no Russian authority states.
|
|
587
|
+
- **No Greek source asks for it either.** The Ministry of Education grammar and the Κέντρο
|
|
588
|
+
Ελληνικής Γλώσσας materials describe αποσιωπητικά without giving any rule for their
|
|
589
|
+
interaction with the ερωτηματικό, and Greek has no counterpart to the Russian
|
|
590
|
+
dot-absorption convention at all. «πάντοτε τρεις» is stated against runs of four and five
|
|
591
|
+
dots, not against a neighbouring mark.
|
|
592
|
+
- So the change is **unsupported on both sides**: it would break the one locale with a cited
|
|
593
|
+
rule in order to serve a locale whose sources are silent.
|
|
594
|
+
|
|
595
|
+
What remains is the two-dot input `Πράγματι;..`, which no authority addresses in either
|
|
596
|
+
language. It round-trips unchanged, which is the correct behaviour for an unspecified case.
|
|
597
|
+
The correctly-written three-dot form `Πράγματι;...` already converts to U+2026, because a
|
|
598
|
+
three-dot run needs no help from `TERMINAL`.
|
|
599
|
+
9. **A verified Greek source contradicts §3.4, and the divergence is deliberate.** The EU
|
|
600
|
+
Interinstitutional Style Guide (Greek edition) §10.1.9 ii) states «Μεταξύ των αποσιωπητικών
|
|
601
|
+
και της λέξης που προηγείται δεν αφήνουμε διάστημα» — *no space is left between the ellipsis
|
|
602
|
+
and the word before it*. §3.4 removes U+2026 from `STRIP-BEFORE` and preserves a space before
|
|
603
|
+
a dot run **in every locale**, so `Πράγματι …` is returned as typed. The rule and the source
|
|
604
|
+
disagree, the source was read and verified, and this entry exists so that the disagreement is
|
|
605
|
+
recorded rather than unremarked.
|
|
606
|
+
|
|
607
|
+
**The divergence stands, and the space is preserved.** Three reasons, in the order that
|
|
608
|
+
decided it:
|
|
609
|
+
|
|
610
|
+
- **Honouring it requires a rule that deletes on locale data, and this rule has none.**
|
|
611
|
+
`order.json` gives `spaces` `"localeData": []`; §2 states that its behaviour is identical in
|
|
612
|
+
every locale, which is not decoration but the reason it can be reasoned about at all. The
|
|
613
|
+
alternatives are to give `spaces` locale data — an `order.json` change, and the end of a
|
|
614
|
+
property the whole document relies on — or to make `ellipsis` delete the preceding space,
|
|
615
|
+
which turns a rule that only ever replaces into a rule that deletes, and thereby drags in
|
|
616
|
+
`modes.md` §3.3's edge-test clause, `I₂`, and a fresh CO discharge. Neither is a small change,
|
|
617
|
+
and neither is justified by one locale.
|
|
618
|
+
- **The instruction is about setting text, not about repairing it.** The guide tells a Greek
|
|
619
|
+
typist not to leave a space there. It does not ask a tool to remove one the writer left. That
|
|
620
|
+
distinction is the same one §3.5 part 2 draws for the ερωτηματικό, and it is the reason this
|
|
621
|
+
document is comfortable saying "no source contradicts it" there and must not say "a source
|
|
622
|
+
requires it" here.
|
|
623
|
+
- **The divergence is conservative in the direction the M4 gate cares about.** Preserving
|
|
624
|
+
costs a missed correction on Greek text containing a typo; honouring it would mean deleting
|
|
625
|
+
a character the author typed, in one locale only, through machinery no other locale exercises.
|
|
626
|
+
`dashes` §3.2 step 2a was just narrowed on exactly that principle after 1063 lines of the
|
|
627
|
+
author's corpus were rewritten against his intent.
|
|
628
|
+
|
|
629
|
+
**What would change the answer:** a second locale wanting the same behaviour. One locale does
|
|
630
|
+
not pay for a schema field, a deleting `ellipsis`, and a new composition argument; two might,
|
|
631
|
+
and at that point the right shape is `ellipsis.noSpaceBefore` consumed by `ellipsis` — which
|
|
632
|
+
already reads locale data — rather than anything in this rule. Recorded in `ellipsis.md` §7 and
|
|
633
|
+
in `spec/locales/el.json`'s `ellipsis` note, so that all three say the same thing.
|
|
634
|
+
10. **An emoticon directly inside a bracket pair still loses the space before its eye.**
|
|
635
|
+
`(␣:(␣b␣)` yields `(:(␣b)`: the opening-bracket clause acts on the outer `(`, and §3.6's eye
|
|
636
|
+
side never sees that run because it only suppresses the `STRIP-BEFORE` clause. Pre-existing
|
|
637
|
+
behaviour, identical for `(␣:)␣b␣)`, and it is a bracket-inner deletion of exactly the kind
|
|
638
|
+
case 3 asks for — so it was left alone rather than folded into the mouth-side fix. Recorded so
|
|
639
|
+
that anyone tempted to widen the guard "for symmetry" starts from the fact that it is the outer
|
|
640
|
+
bracket, not the emoticon, that owns that space.
|
|
641
|
+
11. **The mouth side asks nothing about what precedes the eye, and that is deliberate.** `Note:( x )`
|
|
642
|
+
yields `Note:( x)` — asymmetric, because the mouth side keeps the inner space on the left while
|
|
643
|
+
the `CLOSE-BRACKET` clause still takes the one on the right. The alternative, requiring the eye
|
|
644
|
+
to begin a token (`SPACE`, `BREAK` or `NONE` before it), would restore the symmetry for that
|
|
645
|
+
input and lose `Hi!:( yes`, where the eye is attached to the preceding word and the shape is
|
|
646
|
+
still a face. Deleting is the irreversible direction, so the reading that deletes less wins —
|
|
647
|
+
the same principle as §3.4 and §7.9. Pinned by the fixture `en-us-spaces-emoticon-mouth-eye-attached-to-word`,
|
|
648
|
+
which is the case that discriminates the two readings; without it a port could take the narrower
|
|
649
|
+
one and still pass the whole suite.
|