polytypo 1.1.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +9 -1
- data/lib/polytypo/data/fixtures/de-DE.json +227 -6
- data/lib/polytypo/data/fixtures/el.json +9 -1
- data/lib/polytypo/data/fixtures/en-GB.json +28 -1
- data/lib/polytypo/data/fixtures/en-US.json +690 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +9 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +50 -1
- data/lib/polytypo/data/fixtures/fr.json +282 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +47 -1
- data/lib/polytypo/data/fixtures/sv.json +9 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/apostrophe.rb +10 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +53 -20
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/engine/rules/spaces.rb +8 -1
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
# Rule: `dashes`
|
|
2
2
|
|
|
3
|
-
**Order:** 30. **Default:** on. **Modes:** text, html, markdown.
|
|
3
|
+
**Order:** 30. **Default:** on. **Modes:** text, html, markdown, yaml.
|
|
4
4
|
**Spec version:** 0.6.0 (0.2.0 for everything except the 0.5.0/0.6.0 changes noted inline and in
|
|
5
|
-
§8 History).
|
|
5
|
+
§8 History), amended in **1.3.0**: §1 and §3.4's statement of what belongs to `ranges`, §3.2a's
|
|
6
|
+
re-entry condition, and §3.2 steps 7 and 8, all for [ranges.md](ranges.md) §3.2a's closed-up
|
|
7
|
+
symbols.
|
|
6
8
|
|
|
7
9
|
---
|
|
8
10
|
|
|
@@ -25,9 +27,13 @@ a retired one.
|
|
|
25
27
|
|
|
26
28
|
`dashes` does not touch the ordinary hyphen inside a compound word — which, in the few
|
|
27
29
|
morphological forms where the hyphen must additionally be protected from a line break, belongs
|
|
28
|
-
to `hyphen` at order 35 — and does not touch a
|
|
30
|
+
to `hyphen` at order 35 — and does not touch a **range candidate** at all: that shape belongs to
|
|
29
31
|
`ranges` exclusively, and `dashes` declines it **unconditionally**, whether or not `ranges` is
|
|
30
32
|
enabled (operator decision, spec 0.5.0; see §3.2's note after step 7, and ranges.md §3.2's G1-G5).
|
|
33
|
+
"Range candidate" is [ranges.md](ranges.md) §3.2's term, and as of spec 1.3.0 it is wider than
|
|
34
|
+
"digit-flanked": a closed-up symbol on a flank is walked over first (ranges.md §3.2a), so
|
|
35
|
+
`$15-$20` and `35%-50%` are `ranges`' tokens too, and a token `dashes` used to convert by default
|
|
36
|
+
is now one nothing touches unless the caller enables `ranges`.
|
|
31
37
|
`dashes` never reinterprets a digit-flanked hyphen as a parenthetical dash — that was already true
|
|
32
38
|
in every prior spec version, since the two branches were always mutually exclusive per token; the
|
|
33
39
|
0.5.0 split makes that exclusivity a boundary between two rules instead of two branches of one.
|
|
@@ -80,6 +86,7 @@ Input is a code-point array `cp[0 … n-1]`.
|
|
|
80
86
|
| `ROMAN` | the seven uppercase Roman-numeral letters only: U+0049 `I`, U+0056 `V`, U+0058 `X`, U+004C `L`, U+0043 `C`, U+0044 `D`, U+004D `M`. Lower-case forms are **not** members — see §3.4 P4 |
|
|
81
87
|
| `INERT-DASH` | U+00AD (soft hyphen), U+2011 (non-breaking hyphen), U+2012 (figure dash), U+2015 (horizontal bar), U+FE58, U+FE63, U+FF0D |
|
|
82
88
|
| `JOINER` | U+2060 (word joiner) only. As of spec 0.5.0, produced exclusively by `ranges` (ranges.md §3.3.1) — `dashes` itself never emits one, but still reads through an adjacent run of them (§3.2a, §3.2b), since a `ranges`-produced joiner can sit next to a `dashes` candidate token |
|
|
89
|
+
| `CLOSED-SYMBOL` | **Defined normatively in [ranges.md](ranges.md) §3.2a** (spec 1.3.0), not here: the symbols conventionally written closed up to a number. `dashes` never emits one and never reads one as a candidate, but §3.2 steps 7-8 and §3.4 depend on the class, so it is listed here for a reader transcribing this table. There is exactly one enumeration of it, in ranges.md — do not restate it |
|
|
83
90
|
|
|
84
91
|
`INERT-DASH` members are **never** candidates and are **never** produced. Two of the seven are
|
|
85
92
|
protective markers owned by someone else — U+00AD is invisible formatting, and U+2011 is
|
|
@@ -183,8 +190,8 @@ U+2011 they meant it.
|
|
|
183
190
|
something this rule should be rewriting.
|
|
184
191
|
7. **Cluster guard.** Define a **dash cluster** as a maximal span of code points every one of
|
|
185
192
|
which is in `DASH` ∪ `INERT-DASH` ∪ `DIGIT` ∪ `JOINER` (§3.2b — a joiner `ranges` emitted on
|
|
186
|
-
an earlier pass must not split a cluster it sits inside). (Spaces, letters and punctuation all
|
|
187
|
-
cluster.) Let `C` be the cluster containing this token's run. **If `C` contains two or more
|
|
193
|
+
an earlier pass must not split a cluster it sits inside). (Spaces, letters and punctuation all
|
|
194
|
+
end a cluster.) Let `C` be the cluster containing this token's run. **If `C` contains two or more
|
|
188
195
|
maximal runs of `DASH` ∪ `INERT-DASH`, emit nothing for every token in `C`** — the whole
|
|
189
196
|
cluster is inert.
|
|
190
197
|
This covers idempotency defect (a), §8.2: the classification of a range token reads
|
|
@@ -193,6 +200,13 @@ U+2011 they meant it.
|
|
|
193
200
|
verdict on the next run.
|
|
194
201
|
`2026-08-15`, `978-3-16-148410-0`, `212-555-1234`, `a—0–0` and `1914-1918—annexation` are
|
|
195
202
|
all single clusters with more than one dash run, and are all inert in their entirety.
|
|
203
|
+
**`CLOSED-SYMBOL` is deliberately not in this alphabet (spec 1.3.0).** Widening a range
|
|
204
|
+
member (ranges.md §3.2a) does reach this guard — `a—$15-$20` is two clusters where `a—15-20`
|
|
205
|
+
is one — but the resulting defect is closed by step 8 instead, and at a strictly smaller cost:
|
|
206
|
+
the cluster guard is unconditional, so adding `CLOSED-SYMBOL` here would also make
|
|
207
|
+
`price--$50--drop` inert in an `em-tight` locale that has no such defect to fix. Step 8 fires
|
|
208
|
+
only on the tight-to-spaced transition that can actually disturb a neighbour. See §3.2 step 8
|
|
209
|
+
and [ranges.md](ranges.md) §3.2a.
|
|
196
210
|
**This guard is not sufficient on its own, and it does not subsume G2.** A cluster ends at
|
|
197
211
|
the first space-like code point, so a _spaced_ token's cluster contains only its own run:
|
|
198
212
|
its `cp[L]` is a digit in a neighbouring cluster and its `before` is a dash in a third one,
|
|
@@ -206,14 +220,65 @@ U+2011 they meant it.
|
|
|
206
220
|
- the chosen form is `em-spaced` or `en-spaced` — the replacement will insert a U+0020 on
|
|
207
221
|
each side.
|
|
208
222
|
|
|
209
|
-
In that case, for each side independently
|
|
223
|
+
In that case, for each side independently — reading `L`/`R` and the outward walk as
|
|
224
|
+
**`CLOSED-SYMBOL`-transparent**, per the amendment immediately below:
|
|
210
225
|
- **left:** if `cp[L]` is in `DIGIT`, let `D` be the maximal `DIGIT` run ending at `L` and
|
|
211
226
|
starting at index `d`. Reading **effective neighbours** (§3.2b) outward from `d`: if the
|
|
212
227
|
first is in `DASH` ∪ `INERT-DASH`, **or** the first is in `SPACE` ∪ `NOBREAK-SPACE` and the
|
|
213
228
|
second is in `DASH` ∪ `INERT-DASH`, emit nothing for this token.
|
|
214
229
|
- **right:** if `cp[R]` is in `DIGIT`, let `D` be the maximal `DIGIT` run starting at `R`
|
|
215
|
-
and ending at index `d`.
|
|
216
|
-
|
|
230
|
+
and ending at index `d`. Reading **effective neighbours** (§3.2b) outward from `d` — the
|
|
231
|
+
same reading the left branch uses, and the one §3.2b already requires of "the two-code-point
|
|
232
|
+
reach of T1"; through spec 1.2.0 this branch was written with raw `cp[d+1]`/`cp[d+2]`, which
|
|
233
|
+
disagreed with §3.2b and with every shipped implementation (see below) — if the first is in
|
|
234
|
+
`DASH` ∪ `INERT-DASH`,
|
|
235
|
+
**or** the first is in `SPACE` ∪ `NOBREAK-SPACE` and the second is in `DASH` ∪ `INERT-DASH`,
|
|
236
|
+
emit nothing.
|
|
237
|
+
|
|
238
|
+
**`CLOSED-SYMBOL` transparency (spec 1.3.0).** [ranges.md](ranges.md) §3.2a made a symbol
|
|
239
|
+
written closed up to a digit run part of a **range member**, so the digit run whose verdict
|
|
240
|
+
this guard protects can now sit one code point further out than it used to — on either end.
|
|
241
|
+
At **two** positions per side, one `CLOSED-SYMBOL` is therefore stepped over rather than read:
|
|
242
|
+
|
|
243
|
+
- **p1, between the token and the run.** If `cp[L]` (resp. `cp[R]`) is in `CLOSED-SYMBOL` and
|
|
244
|
+
the code point beyond it is in `DIGIT`, the branch proceeds as though `L` (resp. `R`) were
|
|
245
|
+
that digit. Without p1 the branch is skipped outright, because its gate is `cp[L]`/`cp[R]`
|
|
246
|
+
∈ `DIGIT`: `a--$1 - $1` is the witness.
|
|
247
|
+
- **p2, at the far end of the run.** The outward walk from `d` steps over one `CLOSED-SYMBOL`
|
|
248
|
+
before reading its first and second neighbours. Witness: `a--15% - 20%`, where the walk
|
|
249
|
+
from `d` finds `%` and the dash it exists to find is one code point further out.
|
|
250
|
+
|
|
251
|
+
The two compose on a single side (`a--$15% - $20%` exercises p1 and p2 at once) and are
|
|
252
|
+
independent across sides.
|
|
253
|
+
|
|
254
|
+
**The right branch's raw reading, and what §3.2a changed about it.** The raw text disagreed
|
|
255
|
+
with §3.2b from the start, and the disagreement was always live on the **second** sub-branch:
|
|
256
|
+
a U+0020 ends a cluster, so step 7 never declined `a--15 - 20`, and a port reading raw
|
|
257
|
+
`cp[d+1]` sees the joiner, does not fire, and drifts. (It was unreachable on the first
|
|
258
|
+
sub-branch only — with no space, the token, the digit run, the joiner and the far dash all sat
|
|
259
|
+
inside one cluster, since step 7's alphabet contains `JOINER` and `DIGIT`.) What §3.2a added
|
|
260
|
+
is a second way in: p1 admits `cp[R]` ∈ `CLOSED-SYMBOL`, which is deliberately **not** in the
|
|
261
|
+
cluster alphabet, so `a--$15-$20` splits into two clusters and step 8 stands alone there too.
|
|
262
|
+
Both shapes are now pinned by fixtures, for the same reason §3.2a's re-entry amendment needed
|
|
263
|
+
one: nothing else in the suite separates the two readings. Every shipped implementation
|
|
264
|
+
already read effective neighbours — this is the text catching up, not a behaviour change.
|
|
265
|
+
|
|
266
|
+
**Why two positions are enough, and there is no third.** §3.2a consumes **at most one**
|
|
267
|
+
`CLOSED-SYMBOL` per side — a two-code-point prefix such as `US$` defeats candidacy rather
|
|
268
|
+
than the guard — so each position can hold at most one symbol. A `-spaced` replacement inserts
|
|
269
|
+
exactly one U+0020, so the "first or second neighbour" reach is unchanged in length; the
|
|
270
|
+
symbol shifts *where* that reach starts, never how far it goes. p1 and p2 are the only two
|
|
271
|
+
places a symbol can sit between this token and the far dash, so the amended reach is closed.
|
|
272
|
+
|
|
273
|
+
**The cost, and why it is this guard rather than step 7.** In a locale whose
|
|
274
|
+
`dash.parenthetical` is **spaced**, a tight token next to a closed-up range member no longer
|
|
275
|
+
converts: `Anstieg--50%--war` is left alone in `de-DE`, where spec 1.2.0 produced
|
|
276
|
+
`Anstieg – 50% – war`. That is precisely what the all-digit `Anstieg--50--war` has always
|
|
277
|
+
done, so the two shapes agree, and the cost stops there — a locale with a **tight**
|
|
278
|
+
parenthetical form never reaches this guard at all, so `price--$50--drop` still becomes
|
|
279
|
+
`price—$50—drop` in `en-US`. Putting `CLOSED-SYMBOL` in step 7's cluster alphabet would have
|
|
280
|
+
closed the same defect and taken the `en-US` case with it, because that guard is
|
|
281
|
+
unconditional; it was tried and rejected for exactly that reason.
|
|
217
282
|
|
|
218
283
|
Read plainly: **a tight token must not become spaced when doing so would insert a space
|
|
219
284
|
between itself and a digit run that has another dash on its far side.** That inserted space
|
|
@@ -279,12 +344,18 @@ A `JOINER` adjacent to the token is examined before the branch is chosen:
|
|
|
279
344
|
across a maximal run of `JOINER`. If either walk runs off the array, emit nothing.
|
|
280
345
|
- If no joiner was crossed, `L* = L` and `R* = R` and nothing about the rest of the algorithm
|
|
281
346
|
changes.
|
|
282
|
-
- If a joiner **was** crossed and `cp[L*]` and `cp[R*]` are both in `DIGIT
|
|
283
|
-
|
|
284
|
-
`
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
this
|
|
347
|
+
- If a joiner **was** crossed and `cp[L*]` and `cp[R*]` are both in `DIGIT` — **or, as of spec
|
|
348
|
+
1.3.0, are `DIGIT` after `ranges`' closed-up-symbol walk** ([ranges.md](ranges.md) §3.2a), so
|
|
349
|
+
that `$15–$20` and `35%–50%` reach this clause the same way `1914–1918` does — the token is
|
|
350
|
+
a bound range `ranges` produced on an earlier pass ([ranges.md](ranges.md) §3.3.1): continue
|
|
351
|
+
with `L*`/`R*` in place of `L`/`R`, and extend the token's span to cover the crossed joiners.
|
|
352
|
+
`dashes` itself never reaches this shape as a candidate — extending the span here only ever
|
|
353
|
+
feeds the flank check that routes the token to `ranges`' exclusive territory (§1, §3.3), never
|
|
354
|
+
to this rule's own parenthetical branch. **Without the 1.3.0 clause a bound range carrying
|
|
355
|
+
symbols would fall to the fourth bullet below and be declined by both rules** — stable, but
|
|
356
|
+
wrong in one observable way: a hyphen an author typed between an existing joiner pair,
|
|
357
|
+
`$15-$20`, would never convert. The amendment is what makes that input reach `ranges`, and a
|
|
358
|
+
conformance case pins it, because nothing else in the suite separates the two readings.
|
|
288
359
|
- If a joiner was crossed in any other configuration, **emit nothing**. An author who typed
|
|
289
360
|
U+2060 next to a dash meant it, exactly as with `INERT-DASH` (§3.1).
|
|
290
361
|
|
|
@@ -340,9 +411,10 @@ is [ranges.md](ranges.md) §3.2-§3.3.1 in full** — the admissibility test (`c
|
|
|
340
411
|
verbatim, still reading the same `dash.range` locale field under the same key.
|
|
341
412
|
|
|
342
413
|
**What stays true here, restated because it is `dashes`' own contract now rather than a
|
|
343
|
-
consequence of one rule's two branches:** a
|
|
344
|
-
|
|
345
|
-
|
|
414
|
+
consequence of one rule's two branches:** a range candidate (`cp[L']` and `cp[R']` both `DIGIT`,
|
|
415
|
+
after ranges.md §3.2a's closed-up-symbol walk — through spec 1.2.0 this read `cp[L]`/`cp[R]` and
|
|
416
|
+
meant digit-flanked) is never processed by `dashes` — not converted, not
|
|
417
|
+
declined-and-then-reconsidered, simply never reached. This holds **unconditionally**, whether `ranges` is enabled or not (§1).
|
|
346
418
|
`ranges` disabled does not mean the token falls back to parenthetical treatment; it means nothing
|
|
347
419
|
in the pipeline touches it at all, and `5-10`, `Figure 5-10`, `9-11` and `7-11` are all
|
|
348
420
|
byte-identical no-ops with default options.
|
|
@@ -356,10 +428,10 @@ still occupies its number costs less than a renumbering would.
|
|
|
356
428
|
|
|
357
429
|
### 3.4 Parenthetical branch
|
|
358
430
|
|
|
359
|
-
Reached for every token this rule sees — a
|
|
360
|
-
|
|
361
|
-
|
|
362
|
-
every token that reaches here by construction. Additional guards:
|
|
431
|
+
Reached for every token this rule sees — a range candidate (ranges.md §3.2, §3.2a) is never
|
|
432
|
+
handed to `dashes` at all (§3.3, §1), not merely excluded from this branch, so "is not a range
|
|
433
|
+
candidate" (i.e. at least one flank is not a `DIGIT`, and is not a closed-up symbol matched to
|
|
434
|
+
its opposite member) is true of every token that reaches here by construction. Additional guards:
|
|
363
435
|
|
|
364
436
|
- **P5 (spec 0.6.0) — authored en-dash mark-identity veto.** If `k = 1` and `cp[s]` is U+2013,
|
|
365
437
|
emit nothing — **unconditionally**: every locale, tight or spaced, regardless of
|
|
@@ -525,10 +597,14 @@ of being falsified by another rule, which is then named.
|
|
|
525
597
|
- **[P] A negative number.** `-5` has no space to the left of the digit and a letter/space to
|
|
526
598
|
the left of the hyphen → asymmetric → rejected. This holds for U+002D, U+2010 and U+2212
|
|
527
599
|
alike — the same asymmetry guard, not a glyph-specific one.
|
|
528
|
-
- **[P] Every
|
|
529
|
-
|
|
530
|
-
|
|
531
|
-
|
|
600
|
+
- **[P] Every range candidate, unconditionally** — `5-10`, `Figure 5-10`, `9-11`, and since
|
|
601
|
+
spec 1.3.0 also `$15-$20`, `$15 - $20` and `35%-50%`, where a matched `CLOSED-SYMBOL` sits
|
|
602
|
+
between the stroke and a digit run (ranges.md §3.2a): never reached by `dashes` at all, let
|
|
603
|
+
alone reinterpreted as parenthetical (§1, §3.3). This holds regardless of whether `ranges`'
|
|
604
|
+
own guards would have accepted or declined the token — `dashes` does not evaluate them and
|
|
605
|
+
does not need to. **`$15 - $20` is the one shape this cost anything**: through spec 1.2.0 it
|
|
606
|
+
was `dashes`' token and converted with default options; it is `ranges`' now, and `ranges` is
|
|
607
|
+
off by default.
|
|
532
608
|
- **[P] URLs, code spans, fenced code, HTML attributes.** Removed by the mode adapter before this
|
|
533
609
|
rule sees them. This rule has no notion of a URL and must not grow one.
|
|
534
610
|
- **[P] Line terminators**, which are never inserted, deleted or crossed.
|
|
@@ -537,8 +613,9 @@ of being falsified by another rule, which is then named.
|
|
|
537
613
|
|
|
538
614
|
## 5. Idempotency argument
|
|
539
615
|
|
|
540
|
-
Write `T` for `dashes`. `T` edits only **parenthetical dash tokens**:
|
|
541
|
-
|
|
616
|
+
Write `T` for `dashes`. `T` edits only **parenthetical dash tokens**: tokens that are not range
|
|
617
|
+
candidates (§1, §3.3) — at least one flank is neither a `DIGIT` nor a `CLOSED-SYMBOL` matched on
|
|
618
|
+
the opposite member (ranges.md §3.2a). A range candidate is never reached by `T` on any pass, so
|
|
542
619
|
nothing below needs to reason about one. Each edit `T` makes replaces a span consisting of one
|
|
543
620
|
maximal `DASH` run plus at most one space-like code point on each side, with a span of the same
|
|
544
621
|
shape (`space? dash space?`). So every edit is one of:
|
|
@@ -566,9 +643,10 @@ U+2013/U+2014 no token-level special case, §3.2 step 2a, so this holds regardle
|
|
|
566
643
|
glyph the token holds). A spaced form is re-admitted the same way, with `lsp = rsp = 1` read back
|
|
567
644
|
by §3.2 step 3. A token `nbsp` has since promoted (`ru`: U+00A0 before an em dash) is not
|
|
568
645
|
recomputed at all — §3.2 step 3 makes a no-break-space neighbour space-like, so the isolation
|
|
569
|
-
guard (step 6) declines it and nothing is emitted (§3.6). A
|
|
646
|
+
guard (step 6) declines it and nothing is emitted (§3.6). A range candidate stays outside `T`'s
|
|
570
647
|
domain on every pass, by construction, so it is trivially a fixed point of `T` regardless of what
|
|
571
|
-
`ranges` does to it
|
|
648
|
+
`ranges` does to it — including after `ranges` has bound it, since the binding leaves the flanks
|
|
649
|
+
(and any `CLOSED-SYMBOL` on them) exactly where they were.
|
|
572
650
|
|
|
573
651
|
### 5.2 A declined token stays declined
|
|
574
652
|
|
|
@@ -619,17 +697,25 @@ them, and all three are current:
|
|
|
619
697
|
Range binding ([ranges.md](ranges.md) §3.3.1) is exclusively `ranges`' emission; an interrupting
|
|
620
698
|
parenthetical dash is exactly where a line *may* break, so there is nothing for `dashes` to bind
|
|
621
699
|
(§3.6).
|
|
622
|
-
- **A
|
|
623
|
-
|
|
624
|
-
joiner-crossing walk (§3.2a), so a token `ranges` bound on an
|
|
625
|
-
re-enters across the `JOINER` pair and finds
|
|
626
|
-
|
|
627
|
-
declines it on the same unconditional test.
|
|
628
|
-
|
|
700
|
+
- **A range candidate is never reinterpreted as parenthetical, on any pass.** The candidacy test
|
|
701
|
+
that routes a token to `ranges` instead of `dashes` (§1, §3.3, [ranges.md](ranges.md) §3.2,
|
|
702
|
+
§3.2a) is evaluated after the joiner-crossing walk (§3.2a), so a token `ranges` bound on an
|
|
703
|
+
earlier pass — where the walk re-enters across the `JOINER` pair and finds a candidate on both
|
|
704
|
+
effective neighbours — presents to `dashes` exactly as an unbound one would, and `dashes`
|
|
705
|
+
declines it on the same unconditional test. Since spec 1.3.0 that test reads `cp[L']`/`cp[R']`,
|
|
706
|
+
so a bound range carrying closed-up symbols re-enters on the same footing as an all-digit one.
|
|
707
|
+
`dashes` cannot strip a binding it never reconsiders, and cannot create one, since it never
|
|
708
|
+
emits `JOINER`.
|
|
629
709
|
- **`dashes`' own edits cannot flip a *neighbouring* range token's guard verdict on the next
|
|
630
710
|
pass, even though `dashes` itself never reads that token's guards.** The one edit shape that
|
|
631
711
|
could — a tight token becoming spaced, inserting a U+0020 next to a digit run that has another
|
|
632
712
|
dash on its far side — is exactly what the spacing-transition guard (T1, §3.2 step 8) declines.
|
|
713
|
+
**This clause is true of spec 1.3.0 only because step 8 was amended with it**: ranges.md §3.2a
|
|
714
|
+
widened "next to a digit run" to "next to a range member", which may carry one `CLOSED-SYMBOL`
|
|
715
|
+
at either end, and an unamended T1 reads that symbol as the neighbour and stops one code point
|
|
716
|
+
short of the dash. Four witnesses made the gap concrete — `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
|
|
717
|
+
`$1 - $1--a`, one per position and side — each of which drifted on the second pass before the
|
|
718
|
+
amendment and each of which now behaves exactly as its all-digit analogue always has.
|
|
633
719
|
This is `dashes`' composition obligation *toward* `ranges`, symmetric to the rule-order argument
|
|
634
720
|
in [ranges.md](ranges.md) §4: `ranges` must not see an adjacency `dashes` disturbed, and T1 is
|
|
635
721
|
how `dashes` upholds that on every subsequent pass. The full historical derivation of why this
|
|
@@ -1097,7 +1183,12 @@ form to be `-spaced`. In other words: `A` is tight, `A` is about to become space
|
|
|
1097
1183
|
only, the configuration T1 (§3.2 step 8) declares inert. The mirrored argument gives the `after`
|
|
1098
1184
|
side. T1's two-code-point reach is exactly what the case requires and no more: `T`'s dash sits at
|
|
1099
1185
|
distance 1 from `Lrun` if `T` is tight and distance 2 if `T` is spaced, and a `-spaced` replacement
|
|
1100
|
-
inserts exactly one U+0020, so no other distance is reachable.
|
|
1186
|
+
inserts exactly one U+0020, so no other distance is reachable. **Spec 1.3.0 leaves that reach at
|
|
1187
|
+
two code points and makes it transparent to one `CLOSED-SYMBOL` at each of two positions**
|
|
1188
|
+
(§3.2 step 8): the symbol moves where the reach starts, never how far it goes, because
|
|
1189
|
+
ranges.md §3.2a consumes at most one symbol per side. `JOINER` does not add a position either:
|
|
1190
|
+
it is transparent to this reach by §3.2b, on both branches, so a joiner run between the symbol
|
|
1191
|
+
and the dash collapses rather than counting. The reverse flip — `before` going
|
|
1101
1192
|
from a space to a dash — is benign and needs no guard: it turns a G2 admission into a G2
|
|
1102
1193
|
rejection, and a rejection emits nothing, so a token converted on an earlier pass simply keeps the
|
|
1103
1194
|
form it was given.
|
|
@@ -0,0 +1,271 @@
|
|
|
1
|
+
# Rule: `ellipsis`
|
|
2
|
+
|
|
3
|
+
**Order:** 20. **Default:** on. **Modes:** text, html, markdown, yaml.
|
|
4
|
+
**Spec version:** 0.1.0.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## 1. Purpose
|
|
9
|
+
|
|
10
|
+
`ellipsis` replaces a typed run of full stops with the single character U+2026 (…), and
|
|
11
|
+
implements the abbreviated form that some locales use when an ellipsis follows terminal
|
|
12
|
+
punctuation — Russian writes `?..` and `!..` with two dots rather than three, because the
|
|
13
|
+
question mark or exclamation mark already occupies the first position of the three-dot
|
|
14
|
+
group. The rule is intentionally narrow: a run of exactly two full stops is _never_ touched,
|
|
15
|
+
because `..` occurs in relative paths, in version ranges and in numeric ranges written by
|
|
16
|
+
programmers, and converting it would be a false positive of exactly the class that blocks
|
|
17
|
+
the ship gate (PLAN.md M4). Everything this rule does is a pure function of a maximal run of
|
|
18
|
+
dot-like characters plus, at most, the one code point to its left.
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## 2. Locale data consumed
|
|
23
|
+
|
|
24
|
+
- `ellipsis.abbreviatedAfterTerminal` — boolean, required by `locale.schema.json`.
|
|
25
|
+
|
|
26
|
+
**Evidence for locale data.** A locale attests **membership** — which tokens, which enum value, which boolean — and the rule owns the **mechanism** applied to them. A `sources` citation is not required to name a code point or a behaviour the locale file has no way to vary. Stated once, normatively, in [nbsp.md](nbsp.md) §2.1; it governs every locale field.
|
|
27
|
+
|
|
28
|
+
Nothing else. In particular this rule does not read `nbsp` or `quotes`.
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## 3. Algorithm
|
|
33
|
+
|
|
34
|
+
Input is a code-point array `cp[0 … n-1]`.
|
|
35
|
+
|
|
36
|
+
### 3.1 Character classes
|
|
37
|
+
|
|
38
|
+
| Class | Members |
|
|
39
|
+
| ---------- | ---------------------------- |
|
|
40
|
+
| `DOT` | U+002E (full stop) |
|
|
41
|
+
| `ELL` | U+2026 (horizontal ellipsis) |
|
|
42
|
+
| `DOTLIKE` | `DOT` ∪ `ELL` |
|
|
43
|
+
| `TERMINAL` | U+0021 (!) and U+003F (?) |
|
|
44
|
+
|
|
45
|
+
No other character is examined. U+2025 (‥, two-dot leader), U+22EF, U+FE19 and the CJK
|
|
46
|
+
leaders are **not** members of any class and are never produced or consumed.
|
|
47
|
+
|
|
48
|
+
### 3.2 Ordering dependency
|
|
49
|
+
|
|
50
|
+
`spaces` (order 10) has already run and has deleted every U+0020 that stood immediately before
|
|
51
|
+
a **lone** U+002E. Consequently the spaced form `". . ."` reaches this rule as `"..."` and is
|
|
52
|
+
handled by the ordinary run logic. This rule therefore never needs to look across spaces, and
|
|
53
|
+
must not try to.
|
|
54
|
+
|
|
55
|
+
The word _lone_ is load-bearing and was added after a defect. `spaces` does **not** strip a
|
|
56
|
+
space before a run of two or more dots (`spaces.md` §3.4), so `"e.g. .."` and `"See ../docs"`
|
|
57
|
+
arrive here with their spaces intact. Before that condition existed the space vanished, the
|
|
58
|
+
two-dot run merged with the abbreviation's full stop into a three-dot run, and this rule then
|
|
59
|
+
correctly converted it — producing `e.g…` from input that §4 below promised to leave alone.
|
|
60
|
+
Neither rule was misbehaving; the protection simply was not true of the pipeline.
|
|
61
|
+
|
|
62
|
+
### 3.3 Scan
|
|
63
|
+
|
|
64
|
+
1. Set `i = 0`.
|
|
65
|
+
2. If `cp[i]` is not in `DOTLIKE`, set `i = i + 1` and repeat. Terminate at `i = n`.
|
|
66
|
+
3. Find the maximal run: `s = i`; `e` = the smallest index `> s` with `cp[e]` not in
|
|
67
|
+
`DOTLIKE` (or `e = n`). Let `k = e - s`, and let `d` = the number of `DOT` members and
|
|
68
|
+
`q` = the number of `ELL` members in the run (`d + q = k`).
|
|
69
|
+
4. **Classify the run.**
|
|
70
|
+
- `k = 1` and `d = 1` → an ordinary full stop. Emit nothing. Go to 7.
|
|
71
|
+
- `k = 1` and `q = 1` → an existing ellipsis. Go to 6 (it may still need the abbreviated
|
|
72
|
+
form).
|
|
73
|
+
- `k = 2` and `q = 0` → **two full stops.** Go to 5.
|
|
74
|
+
- `k ≥ 2` and `q ≥ 1` → a mixed or repeated run (`"…."`, `"……"`, `"..…"`). Normalise:
|
|
75
|
+
emit an edit replacing `cp[s … e-1]` with one U+2026. Set the run to that single
|
|
76
|
+
U+2026 and go to 6.
|
|
77
|
+
- `k ≥ 3` and `q = 0` → three or more full stops. Emit an edit replacing
|
|
78
|
+
`cp[s … e-1]` with one U+2026. Set the run to that single U+2026 and go to 6.
|
|
79
|
+
5. **The two-dot case.** Let `left = cp[s-1]` if `s > 0`, else `NONE`.
|
|
80
|
+
- If `ellipsis.abbreviatedAfterTerminal` is `true` → emit nothing, unconditionally. In
|
|
81
|
+
such a locale `"?.."` is the _correct_ output form and must survive re-processing.
|
|
82
|
+
- If `ellipsis.abbreviatedAfterTerminal` is `false` **and** `left` is in `TERMINAL` →
|
|
83
|
+
emit an edit replacing `cp[s … e-1]` with one U+2026. (`"?.."` in an English text is a
|
|
84
|
+
typing slip for `"?…"`.) Go to 7; step 6 cannot fire because the locale flag is false.
|
|
85
|
+
- Otherwise → emit nothing. This is the branch that protects `"../"`, `"1..5"` and
|
|
86
|
+
`"a..b"`.
|
|
87
|
+
Go to 7.
|
|
88
|
+
6. **Abbreviated-after-terminal form.** The run is now a single U+2026 at position `s`
|
|
89
|
+
(either pre-existing or just produced by step 4).
|
|
90
|
+
- If `ellipsis.abbreviatedAfterTerminal` is `false` → emit nothing. Go to 7.
|
|
91
|
+
- Let `left = cp[s-1]` if `s > 0`, else `NONE`. If `left` is **not** in `TERMINAL` →
|
|
92
|
+
emit nothing. Go to 7.
|
|
93
|
+
- Otherwise emit an edit replacing the single U+2026 at `s` with the two code points
|
|
94
|
+
U+002E U+002E. Steps 4 and 6 are stages of one decision about one span: the rule emits
|
|
95
|
+
**exactly one edit per run**, whose replacement is the final form. `"?..."` is therefore
|
|
96
|
+
a single edit replacing three U+002E with two U+002E, not two chained edits.
|
|
97
|
+
7. Set `i = e` and go to 2.
|
|
98
|
+
|
|
99
|
+
### 3.4 Interaction with `!?` and `?!`
|
|
100
|
+
|
|
101
|
+
`left` in step 6 is a single code point. For `"?!..."` the character left of the ellipsis is
|
|
102
|
+
U+0021, which is in `TERMINAL`, so the abbreviated form fires: `"?!.."`. This matches the
|
|
103
|
+
Russian convention (`«Что?!..»`). No lookbehind beyond one code point is required or
|
|
104
|
+
permitted.
|
|
105
|
+
|
|
106
|
+
---
|
|
107
|
+
|
|
108
|
+
## 4. Must not touch
|
|
109
|
+
|
|
110
|
+
**Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
|
|
111
|
+
a guarantee of `transform` as a whole — or **[R]** — true of this rule alone. Every bullet here
|
|
112
|
+
is [P], and two of them only became true when `spaces` gained the lone-dot condition.
|
|
113
|
+
|
|
114
|
+
- **[P] A single U+002E.** Sentence full stops, decimal points, abbreviation dots (`p.`, `z.`),
|
|
115
|
+
and the dot in `"1.2.3"` are all runs of length 1.
|
|
116
|
+
- **[P] A run of exactly two U+002E**, except in the one narrow case in step 5 where the locale
|
|
117
|
+
does _not_ use the abbreviated form and the run directly follows `!` or `?`. This protects
|
|
118
|
+
`"../"`, `"./.."`, `"1..5"`, `"e.g. .."` and every other two-dot idiom.
|
|
119
|
+
_[P] only since `spaces` gained the lone-dot condition. `"e.g. .."` previously became `e.g…`
|
|
120
|
+
and `"See ../docs"` lost its space — the run itself was untouched throughout, which is exactly
|
|
121
|
+
what made the claim look true while the pipeline falsified it._
|
|
122
|
+
- **[P] U+2025 (‥) and the CJK dot leaders.** Not in any class; never produced, never consumed.
|
|
123
|
+
- **[P] Two `DOT` runs separated by anything at all.** Runs are maximal and are never joined
|
|
124
|
+
across an intervening character. After `spaces` has run, a surviving separator between
|
|
125
|
+
dots is a no-break space or a tab, i.e. something the author put there deliberately.
|
|
126
|
+
- **[P] Anything in a skipped region.** `"..."` inside a code span or a URL never reaches this
|
|
127
|
+
rule; the mode adapter removes it. This rule has no code-awareness and must not acquire
|
|
128
|
+
any.
|
|
129
|
+
- **[P] The spacing around an ellipsis.** Whether `"word …"` should carry a no-break space is
|
|
130
|
+
`nbsp`'s decision, driven by `nbsp.beforePunctuation` / `nbsp.narrowBeforePunctuation`,
|
|
131
|
+
which may list U+2026.
|
|
132
|
+
|
|
133
|
+
---
|
|
134
|
+
|
|
135
|
+
## 5. Idempotency argument
|
|
136
|
+
|
|
137
|
+
Write `T` for the rule. The output of `T` contains, at every position where `T` acted, one
|
|
138
|
+
of exactly two forms:
|
|
139
|
+
|
|
140
|
+
- **Form A: a lone U+2026** whose left neighbour is either absent or not in `TERMINAL`, or
|
|
141
|
+
whose locale has `abbreviatedAfterTerminal = false`.
|
|
142
|
+
- **Form B: exactly two U+002E** whose left neighbour is in `TERMINAL`, in a locale with
|
|
143
|
+
`abbreviatedAfterTerminal = true`.
|
|
144
|
+
|
|
145
|
+
Re-running `T`:
|
|
146
|
+
|
|
147
|
+
- Form A is a run with `k = 1`, `q = 1`. Step 4 sends it to step 6. Step 6 emits nothing
|
|
148
|
+
because either the flag is false or `left` is not in `TERMINAL` — both of which are exactly
|
|
149
|
+
the conditions under which Form A was produced. No edit.
|
|
150
|
+
- Form B is a run with `k = 2`, `q = 0`. Step 5 is reached, and the first branch applies
|
|
151
|
+
(flag is `true`) → emit nothing, unconditionally. No edit.
|
|
152
|
+
- Runs the rule declined to touch (single dot, two dots outside the special case) are
|
|
153
|
+
classified identically on the second run, because classification depends only on the run
|
|
154
|
+
itself and on one left neighbour, neither of which this rule altered — the rule never
|
|
155
|
+
edits a character outside a `DOTLIKE` run, and a `DOTLIKE` run is never adjacent to
|
|
156
|
+
another `DOTLIKE` run after `T` (runs are maximal, so merging cannot occur).
|
|
157
|
+
|
|
158
|
+
Hence `T(T(x)) = T(x)`.
|
|
159
|
+
|
|
160
|
+
**What had to be fixed.** The naive formulation is _"`...` → `…`, and in Russian `?…` →
|
|
161
|
+
`?..`"_. That is not idempotent, because the output `?..` is a two-dot run which the same
|
|
162
|
+
naive rule's companion clause _"collapse repeated dots"_ would then re-expand — or, in the
|
|
163
|
+
formulation used by several existing libraries, `?..` → `?…` → `?..` → … oscillating.
|
|
164
|
+
Two changes fix it:
|
|
165
|
+
|
|
166
|
+
1. The two-dot run is **unconditionally inert** when `abbreviatedAfterTerminal` is true. The
|
|
167
|
+
output form is a fixed point by construction rather than by luck.
|
|
168
|
+
2. The two-dot → ellipsis conversion exists **only** in the complementary locale setting and
|
|
169
|
+
**only** after terminal punctuation, so the two branches can never both apply to the same
|
|
170
|
+
text under the same locale. The flag partitions the behaviour; it does not layer it.
|
|
171
|
+
|
|
172
|
+
---
|
|
173
|
+
|
|
174
|
+
### Composition obligation
|
|
175
|
+
|
|
176
|
+
Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₂**, so the
|
|
177
|
+
obligation runs against `spaces` (R₁) only.
|
|
178
|
+
|
|
179
|
+
**What this rule emits.** One U+2026 replacing a run of U+002E/U+2026, or two U+002E replacing
|
|
180
|
+
one U+2026. Nothing else: no spaces, no letters, no dashes, and never a code point outside
|
|
181
|
+
`DOTLIKE`.
|
|
182
|
+
|
|
183
|
+
**Against `I₁` (`spaces`).** Discharged. This rule never emits U+0020, so violations S-a, S-c
|
|
184
|
+
and S-d are unreachable. S-b — a U+0020 whose right neighbour is in `STRIP-BEFORE` — deserves a
|
|
185
|
+
second look, because U+002E and U+2026 are both in `STRIP-BEFORE` and this rule moves them
|
|
186
|
+
about. But every replacement starts at the first code point of the run and the code point to
|
|
187
|
+
its left is unchanged; if that neighbour were a U+0020, `spaces` would already have deleted it
|
|
188
|
+
in the same pass, since the run began with a `DOTLIKE` character then too. Shortening a run
|
|
189
|
+
cannot bring a space into contact with a dot that was not already in contact with one.
|
|
190
|
+
|
|
191
|
+
**`I₂` in the other direction** is preserved by every later rule trivially: none of them emits
|
|
192
|
+
U+002E or U+2026, and none deletes a code point standing between two dot runs.
|
|
193
|
+
|
|
194
|
+
---
|
|
195
|
+
|
|
196
|
+
## 6. Worked examples
|
|
197
|
+
|
|
198
|
+
`⟶` = no change. Both locale settings are shown because the rule is one of only two whose
|
|
199
|
+
output differs on a boolean.
|
|
200
|
+
|
|
201
|
+
### `ellipsis.abbreviatedAfterTerminal = false` (en, fi, sv, de, fr, el)
|
|
202
|
+
|
|
203
|
+
| # | Input | Output | Why |
|
|
204
|
+
| --- | --------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
205
|
+
| 1 | `Wait... what?` | `Wait… what?` | `k = 3`, `q = 0` → single U+2026 |
|
|
206
|
+
| 2 | `Wait…… what?` | `Wait… what?` | mixed run normalised |
|
|
207
|
+
| 3 | `Really?.. ` | `Really?… ` | two dots after `?` in a non-abbreviating locale |
|
|
208
|
+
| 4 | `See ../docs` | ⟶ | two-dot run, left neighbour U+0020 not in `TERMINAL`. **The space also survives**, which is `spaces.md` §3.4's doing, not this rule's — before that condition the pipeline returned `See../docs` |
|
|
209
|
+
| 4a | `e.g. ..` | ⟶ | likewise: the space survives, so no three-dot run is ever formed. Previously `e.g…` |
|
|
210
|
+
| 5 | `Version 1..5` | ⟶ | two-dot run, left neighbour is a digit |
|
|
211
|
+
| 6 | `He left…` | ⟶ | already a lone U+2026 |
|
|
212
|
+
| 7 | `Hmm.....` | `Hmm…` | `k = 5` → one U+2026 |
|
|
213
|
+
| 8 | `Yes. No.` | ⟶ | two independent runs of length 1 |
|
|
214
|
+
|
|
215
|
+
### `ellipsis.abbreviatedAfterTerminal = true` (ru)
|
|
216
|
+
|
|
217
|
+
| # | Input | Output | Why |
|
|
218
|
+
| --- | ------------ | ---------- | ------------------------------------------------- |
|
|
219
|
+
| 9 | `Что?...` | `Что?..` | run → U+2026 → abbreviated after `?` |
|
|
220
|
+
| 10 | `Что?…` | `Что?..` | existing U+2026 after `?` rewritten |
|
|
221
|
+
| 11 | `Что?..` | ⟶ | already the target form; step 5 first branch |
|
|
222
|
+
| 12 | `Что?!...` | `Что?!..` | left neighbour of the run is U+0021 |
|
|
223
|
+
| 13 | `Он ушёл...` | `Он ушёл…` | left neighbour is a letter → ordinary ellipsis |
|
|
224
|
+
| 14 | `см. ../` | ⟶ | two-dot run, unconditionally inert in this locale |
|
|
225
|
+
|
|
226
|
+
Cases 4, 5, 6, 8, 11 and 14 are "no change" cases.
|
|
227
|
+
|
|
228
|
+
---
|
|
229
|
+
|
|
230
|
+
## 7. Open questions
|
|
231
|
+
|
|
232
|
+
1. **Four dots.** Chicago (and several book styles) write `"….",` an ellipsis followed by a
|
|
233
|
+
sentence-ending period, for an omission at the end of a sentence. This rule collapses
|
|
234
|
+
`"...."` to a single `"…"`, losing that distinction. Distinguishing the two requires
|
|
235
|
+
knowing whether the omission ends the sentence, which is not decidable from the code
|
|
236
|
+
points. I chose collapse; an alternative is to leave runs of exactly four alone. Needs an
|
|
237
|
+
operator decision and a fixture either way.
|
|
238
|
+
2. **`abbreviatedAfterTerminal` is a single boolean**, so a locale cannot say "abbreviate
|
|
239
|
+
after `?` but not after `!`". No known locale needs that, but the schema forecloses it.
|
|
240
|
+
Recorded, not proposed.
|
|
241
|
+
3. **The `!..` / `?..` order.** Russian also uses `"..?"` and `"..!"` in some sources for an
|
|
242
|
+
ellipsis _preceding_ terminal punctuation. This rule does not implement that direction,
|
|
243
|
+
and `"..?"` is left alone (two-dot run). Whether Мильчин requires the leading form is a
|
|
244
|
+
locale-research question, and if the answer is yes, the schema needs a second flag.
|
|
245
|
+
4. _(Partly settled.)_ The `nbsp` asymmetry recorded here is fixed on the `nbsp` side: U+2026
|
|
246
|
+
is now accepted right context for N1/N2 (`nbsp.md` §3.3 step 2), so French `Vraiment?…`
|
|
247
|
+
takes its narrow space just as `Vraiment ?` does. What remains open is the original
|
|
248
|
+
observation — **The interaction with `nbsp` is unspecified from this side.** If a locale lists U+2026 in
|
|
249
|
+
`nbsp.beforePunctuation`, then `"word …"` gains a no-break space, but the abbreviated
|
|
250
|
+
Russian form `"?.."` ends in U+002E and would not. That asymmetry may or may not be
|
|
251
|
+
correct; it is a question for the `ru` locale file, not for this rule.
|
|
252
|
+
5. **Two dots in a non-abbreviating locale after `?`/`!`** (case 3) is the only place this
|
|
253
|
+
rule converts a two-dot run. It is conservative and safe, but it is also the only asymmetry
|
|
254
|
+
in the rule. If it produces a false positive in real content, deleting the branch costs
|
|
255
|
+
nothing.
|
|
256
|
+
6. **Greek forbids a space before the ellipsis; polytypo preserves it anyway.** The EU
|
|
257
|
+
Interinstitutional Style Guide (Greek edition) §10.1.9 ii) — «Μεταξύ των αποσιωπητικών και της
|
|
258
|
+
λέξης που προηγείται δεν αφήνουμε διάστημα» — is verified against the source and is not
|
|
259
|
+
honoured. `spaces.md` §3.4 preserves a space before a dot run in every locale, so `Πράγματι …`
|
|
260
|
+
comes back as typed.
|
|
261
|
+
|
|
262
|
+
The full argument is in `spaces.md` §7.9 and is not repeated; the part that belongs **here** is
|
|
263
|
+
why this rule does not fix it instead. This rule reads locale data and could carry an
|
|
264
|
+
`ellipsis.noSpaceBefore` flag, which the schema can express and which fits the declarative-data
|
|
265
|
+
line of PLAN.md §6. What that would cost is a change of kind, not of degree: **this rule
|
|
266
|
+
currently only ever replaces one span of dot-like code points with another, and it would become
|
|
267
|
+
a rule that deletes a character outside its own run.** Deleting brings in `modes.md` §3.3's
|
|
268
|
+
edge-test clause (a deleting rule must treat a span edge as the end of the text), a new
|
|
269
|
+
discharge against `I₁`, and a re-derivation of §5's idempotency argument, which currently
|
|
270
|
+
depends on the run being the only thing touched. That is not worth it for one locale, and it is
|
|
271
|
+
worth it for two. Recorded so the option is costed rather than forgotten.
|