polytypo 1.1.0 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (66) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +9 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +227 -6
  7. data/lib/polytypo/data/fixtures/el.json +9 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +28 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +690 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +9 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +50 -1
  13. data/lib/polytypo/data/fixtures/fr.json +282 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +47 -1
  21. data/lib/polytypo/data/fixtures/sv.json +9 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/apostrophe.rb +10 -1
  55. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  56. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  57. data/lib/polytypo/engine/rules/nbsp.rb +53 -20
  58. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  59. data/lib/polytypo/engine/rules/spaces.rb +8 -1
  60. data/lib/polytypo/errors.rb +3 -0
  61. data/lib/polytypo/modes/runner.rb +17 -0
  62. data/lib/polytypo/modes/spans.rb +30 -2
  63. data/lib/polytypo/modes/yaml.rb +312 -0
  64. data/lib/polytypo/version.rb +1 -1
  65. data/lib/polytypo.rb +126 -15
  66. metadata +31 -1
@@ -1,8 +1,10 @@
1
1
  # Rule: `dashes`
2
2
 
3
- **Order:** 30. **Default:** on. **Modes:** text, html, markdown.
3
+ **Order:** 30. **Default:** on. **Modes:** text, html, markdown, yaml.
4
4
  **Spec version:** 0.6.0 (0.2.0 for everything except the 0.5.0/0.6.0 changes noted inline and in
5
- §8 History).
5
+ §8 History), amended in **1.3.0**: §1 and §3.4's statement of what belongs to `ranges`, §3.2a's
6
+ re-entry condition, and §3.2 steps 7 and 8, all for [ranges.md](ranges.md) §3.2a's closed-up
7
+ symbols.
6
8
 
7
9
  ---
8
10
 
@@ -25,9 +27,13 @@ a retired one.
25
27
 
26
28
  `dashes` does not touch the ordinary hyphen inside a compound word — which, in the few
27
29
  morphological forms where the hyphen must additionally be protected from a line break, belongs
28
- to `hyphen` at order 35 — and does not touch a digit-flanked stroke at all: that shape belongs to
30
+ to `hyphen` at order 35 — and does not touch a **range candidate** at all: that shape belongs to
29
31
  `ranges` exclusively, and `dashes` declines it **unconditionally**, whether or not `ranges` is
30
32
  enabled (operator decision, spec 0.5.0; see §3.2's note after step 7, and ranges.md §3.2's G1-G5).
33
+ "Range candidate" is [ranges.md](ranges.md) §3.2's term, and as of spec 1.3.0 it is wider than
34
+ "digit-flanked": a closed-up symbol on a flank is walked over first (ranges.md §3.2a), so
35
+ `$15-$20` and `35%-50%` are `ranges`' tokens too, and a token `dashes` used to convert by default
36
+ is now one nothing touches unless the caller enables `ranges`.
31
37
  `dashes` never reinterprets a digit-flanked hyphen as a parenthetical dash — that was already true
32
38
  in every prior spec version, since the two branches were always mutually exclusive per token; the
33
39
  0.5.0 split makes that exclusivity a boundary between two rules instead of two branches of one.
@@ -80,6 +86,7 @@ Input is a code-point array `cp[0 … n-1]`.
80
86
  | `ROMAN` | the seven uppercase Roman-numeral letters only: U+0049 `I`, U+0056 `V`, U+0058 `X`, U+004C `L`, U+0043 `C`, U+0044 `D`, U+004D `M`. Lower-case forms are **not** members — see §3.4 P4 |
81
87
  | `INERT-DASH` | U+00AD (soft hyphen), U+2011 (non-breaking hyphen), U+2012 (figure dash), U+2015 (horizontal bar), U+FE58, U+FE63, U+FF0D |
82
88
  | `JOINER` | U+2060 (word joiner) only. As of spec 0.5.0, produced exclusively by `ranges` (ranges.md §3.3.1) — `dashes` itself never emits one, but still reads through an adjacent run of them (§3.2a, §3.2b), since a `ranges`-produced joiner can sit next to a `dashes` candidate token |
89
+ | `CLOSED-SYMBOL` | **Defined normatively in [ranges.md](ranges.md) §3.2a** (spec 1.3.0), not here: the symbols conventionally written closed up to a number. `dashes` never emits one and never reads one as a candidate, but §3.2 steps 7-8 and §3.4 depend on the class, so it is listed here for a reader transcribing this table. There is exactly one enumeration of it, in ranges.md — do not restate it |
83
90
 
84
91
  `INERT-DASH` members are **never** candidates and are **never** produced. Two of the seven are
85
92
  protective markers owned by someone else — U+00AD is invisible formatting, and U+2011 is
@@ -183,8 +190,8 @@ U+2011 they meant it.
183
190
  something this rule should be rewriting.
184
191
  7. **Cluster guard.** Define a **dash cluster** as a maximal span of code points every one of
185
192
  which is in `DASH` ∪ `INERT-DASH` ∪ `DIGIT` ∪ `JOINER` (§3.2b — a joiner `ranges` emitted on
186
- an earlier pass must not split a cluster it sits inside). (Spaces, letters and punctuation all end a
187
- cluster.) Let `C` be the cluster containing this token's run. **If `C` contains two or more
193
+ an earlier pass must not split a cluster it sits inside). (Spaces, letters and punctuation all
194
+ end a cluster.) Let `C` be the cluster containing this token's run. **If `C` contains two or more
188
195
  maximal runs of `DASH` ∪ `INERT-DASH`, emit nothing for every token in `C`** — the whole
189
196
  cluster is inert.
190
197
  This covers idempotency defect (a), §8.2: the classification of a range token reads
@@ -193,6 +200,13 @@ U+2011 they meant it.
193
200
  verdict on the next run.
194
201
  `2026-08-15`, `978-3-16-148410-0`, `212-555-1234`, `a—0–0` and `1914-1918—annexation` are
195
202
  all single clusters with more than one dash run, and are all inert in their entirety.
203
+ **`CLOSED-SYMBOL` is deliberately not in this alphabet (spec 1.3.0).** Widening a range
204
+ member (ranges.md §3.2a) does reach this guard — `a—$15-$20` is two clusters where `a—15-20`
205
+ is one — but the resulting defect is closed by step 8 instead, and at a strictly smaller cost:
206
+ the cluster guard is unconditional, so adding `CLOSED-SYMBOL` here would also make
207
+ `price--$50--drop` inert in an `em-tight` locale that has no such defect to fix. Step 8 fires
208
+ only on the tight-to-spaced transition that can actually disturb a neighbour. See §3.2 step 8
209
+ and [ranges.md](ranges.md) §3.2a.
196
210
  **This guard is not sufficient on its own, and it does not subsume G2.** A cluster ends at
197
211
  the first space-like code point, so a _spaced_ token's cluster contains only its own run:
198
212
  its `cp[L]` is a digit in a neighbouring cluster and its `before` is a dash in a third one,
@@ -206,14 +220,65 @@ U+2011 they meant it.
206
220
  - the chosen form is `em-spaced` or `en-spaced` — the replacement will insert a U+0020 on
207
221
  each side.
208
222
 
209
- In that case, for each side independently:
223
+ In that case, for each side independently — reading `L`/`R` and the outward walk as
224
+ **`CLOSED-SYMBOL`-transparent**, per the amendment immediately below:
210
225
  - **left:** if `cp[L]` is in `DIGIT`, let `D` be the maximal `DIGIT` run ending at `L` and
211
226
  starting at index `d`. Reading **effective neighbours** (§3.2b) outward from `d`: if the
212
227
  first is in `DASH` ∪ `INERT-DASH`, **or** the first is in `SPACE` ∪ `NOBREAK-SPACE` and the
213
228
  second is in `DASH` ∪ `INERT-DASH`, emit nothing for this token.
214
229
  - **right:** if `cp[R]` is in `DIGIT`, let `D` be the maximal `DIGIT` run starting at `R`
215
- and ending at index `d`. If `cp[d+1]` is in `DASH` ∪ `INERT-DASH`, **or** `cp[d+1]` is in
216
- `SPACE` ∪ `NOBREAK-SPACE` and `cp[d+2]` is in `DASH` ∪ `INERT-DASH`, emit nothing.
230
+ and ending at index `d`. Reading **effective neighbours** (§3.2b) outward from `d` — the
231
+ same reading the left branch uses, and the one §3.2b already requires of "the two-code-point
232
+ reach of T1"; through spec 1.2.0 this branch was written with raw `cp[d+1]`/`cp[d+2]`, which
233
+ disagreed with §3.2b and with every shipped implementation (see below) — if the first is in
234
+ `DASH` ∪ `INERT-DASH`,
235
+ **or** the first is in `SPACE` ∪ `NOBREAK-SPACE` and the second is in `DASH` ∪ `INERT-DASH`,
236
+ emit nothing.
237
+
238
+ **`CLOSED-SYMBOL` transparency (spec 1.3.0).** [ranges.md](ranges.md) §3.2a made a symbol
239
+ written closed up to a digit run part of a **range member**, so the digit run whose verdict
240
+ this guard protects can now sit one code point further out than it used to — on either end.
241
+ At **two** positions per side, one `CLOSED-SYMBOL` is therefore stepped over rather than read:
242
+
243
+ - **p1, between the token and the run.** If `cp[L]` (resp. `cp[R]`) is in `CLOSED-SYMBOL` and
244
+ the code point beyond it is in `DIGIT`, the branch proceeds as though `L` (resp. `R`) were
245
+ that digit. Without p1 the branch is skipped outright, because its gate is `cp[L]`/`cp[R]`
246
+ ∈ `DIGIT`: `a--$1 - $1` is the witness.
247
+ - **p2, at the far end of the run.** The outward walk from `d` steps over one `CLOSED-SYMBOL`
248
+ before reading its first and second neighbours. Witness: `a--15% - 20%`, where the walk
249
+ from `d` finds `%` and the dash it exists to find is one code point further out.
250
+
251
+ The two compose on a single side (`a--$15% - $20%` exercises p1 and p2 at once) and are
252
+ independent across sides.
253
+
254
+ **The right branch's raw reading, and what §3.2a changed about it.** The raw text disagreed
255
+ with §3.2b from the start, and the disagreement was always live on the **second** sub-branch:
256
+ a U+0020 ends a cluster, so step 7 never declined `a--15⁠ - 20`, and a port reading raw
257
+ `cp[d+1]` sees the joiner, does not fire, and drifts. (It was unreachable on the first
258
+ sub-branch only — with no space, the token, the digit run, the joiner and the far dash all sat
259
+ inside one cluster, since step 7's alphabet contains `JOINER` and `DIGIT`.) What §3.2a added
260
+ is a second way in: p1 admits `cp[R]` ∈ `CLOSED-SYMBOL`, which is deliberately **not** in the
261
+ cluster alphabet, so `a--$15⁠-⁠$20` splits into two clusters and step 8 stands alone there too.
262
+ Both shapes are now pinned by fixtures, for the same reason §3.2a's re-entry amendment needed
263
+ one: nothing else in the suite separates the two readings. Every shipped implementation
264
+ already read effective neighbours — this is the text catching up, not a behaviour change.
265
+
266
+ **Why two positions are enough, and there is no third.** §3.2a consumes **at most one**
267
+ `CLOSED-SYMBOL` per side — a two-code-point prefix such as `US$` defeats candidacy rather
268
+ than the guard — so each position can hold at most one symbol. A `-spaced` replacement inserts
269
+ exactly one U+0020, so the "first or second neighbour" reach is unchanged in length; the
270
+ symbol shifts *where* that reach starts, never how far it goes. p1 and p2 are the only two
271
+ places a symbol can sit between this token and the far dash, so the amended reach is closed.
272
+
273
+ **The cost, and why it is this guard rather than step 7.** In a locale whose
274
+ `dash.parenthetical` is **spaced**, a tight token next to a closed-up range member no longer
275
+ converts: `Anstieg--50%--war` is left alone in `de-DE`, where spec 1.2.0 produced
276
+ `Anstieg – 50% – war`. That is precisely what the all-digit `Anstieg--50--war` has always
277
+ done, so the two shapes agree, and the cost stops there — a locale with a **tight**
278
+ parenthetical form never reaches this guard at all, so `price--$50--drop` still becomes
279
+ `price—$50—drop` in `en-US`. Putting `CLOSED-SYMBOL` in step 7's cluster alphabet would have
280
+ closed the same defect and taken the `en-US` case with it, because that guard is
281
+ unconditional; it was tried and rejected for exactly that reason.
217
282
 
218
283
  Read plainly: **a tight token must not become spaced when doing so would insert a space
219
284
  between itself and a digit run that has another dash on its far side.** That inserted space
@@ -279,12 +344,18 @@ A `JOINER` adjacent to the token is examined before the branch is chosen:
279
344
  across a maximal run of `JOINER`. If either walk runs off the array, emit nothing.
280
345
  - If no joiner was crossed, `L* = L` and `R* = R` and nothing about the rest of the algorithm
281
346
  changes.
282
- - If a joiner **was** crossed and `cp[L*]` and `cp[R*]` are both in `DIGIT`, the token is a
283
- bound range `ranges` produced on an earlier pass ([ranges.md](ranges.md) §3.3.1): continue with
284
- `L*`/`R*` in place of `L`/`R`, and extend the token's span to cover the crossed joiners. `dashes`
285
- itself never reaches this shape as a candidate — extending the span here only ever feeds the
286
- digit-flanked check that routes the token to `ranges`' exclusive territory (§1, §3.3), never to
287
- this rule's own parenthetical branch.
347
+ - If a joiner **was** crossed and `cp[L*]` and `cp[R*]` are both in `DIGIT` — **or, as of spec
348
+ 1.3.0, are `DIGIT` after `ranges`' closed-up-symbol walk** ([ranges.md](ranges.md) §3.2a), so
349
+ that `$15⁠–⁠$20` and `35%⁠–⁠50%` reach this clause the same way `1914⁠–⁠1918` does — the token is
350
+ a bound range `ranges` produced on an earlier pass ([ranges.md](ranges.md) §3.3.1): continue
351
+ with `L*`/`R*` in place of `L`/`R`, and extend the token's span to cover the crossed joiners.
352
+ `dashes` itself never reaches this shape as a candidate — extending the span here only ever
353
+ feeds the flank check that routes the token to `ranges`' exclusive territory (§1, §3.3), never
354
+ to this rule's own parenthetical branch. **Without the 1.3.0 clause a bound range carrying
355
+ symbols would fall to the fourth bullet below and be declined by both rules** — stable, but
356
+ wrong in one observable way: a hyphen an author typed between an existing joiner pair,
357
+ `$15⁠-⁠$20`, would never convert. The amendment is what makes that input reach `ranges`, and a
358
+ conformance case pins it, because nothing else in the suite separates the two readings.
288
359
  - If a joiner was crossed in any other configuration, **emit nothing**. An author who typed
289
360
  U+2060 next to a dash meant it, exactly as with `INERT-DASH` (§3.1).
290
361
 
@@ -340,9 +411,10 @@ is [ranges.md](ranges.md) §3.2-§3.3.1 in full** — the admissibility test (`c
340
411
  verbatim, still reading the same `dash.range` locale field under the same key.
341
412
 
342
413
  **What stays true here, restated because it is `dashes`' own contract now rather than a
343
- consequence of one rule's two branches:** a digit-flanked dash token (`cp[L]` and `cp[R]` both
344
- `DIGIT`) is never processed by `dashes` — not converted, not declined-and-then-reconsidered,
345
- simply never reached. This holds **unconditionally**, whether `ranges` is enabled or not (§1).
414
+ consequence of one rule's two branches:** a range candidate (`cp[L']` and `cp[R']` both `DIGIT`,
415
+ after ranges.md §3.2a's closed-up-symbol walk — through spec 1.2.0 this read `cp[L]`/`cp[R]` and
416
+ meant digit-flanked) is never processed by `dashes` — not converted, not
417
+ declined-and-then-reconsidered, simply never reached. This holds **unconditionally**, whether `ranges` is enabled or not (§1).
346
418
  `ranges` disabled does not mean the token falls back to parenthetical treatment; it means nothing
347
419
  in the pipeline touches it at all, and `5-10`, `Figure 5-10`, `9-11` and `7-11` are all
348
420
  byte-identical no-ops with default options.
@@ -356,10 +428,10 @@ still occupies its number costs less than a renumbering would.
356
428
 
357
429
  ### 3.4 Parenthetical branch
358
430
 
359
- Reached for every token this rule sees — a digit-flanked token (both `cp[L]` and `cp[R]` in
360
- `DIGIT`) is never handed to `dashes` at all (§3.3, §1), not merely excluded from this branch, so
361
- "is not a range candidate" (i.e. at least one of `cp[L]`, `cp[R]` is not a `DIGIT`) is true of
362
- every token that reaches here by construction. Additional guards:
431
+ Reached for every token this rule sees — a range candidate (ranges.md §3.2, §3.2a) is never
432
+ handed to `dashes` at all (§3.3, §1), not merely excluded from this branch, so "is not a range
433
+ candidate" (i.e. at least one flank is not a `DIGIT`, and is not a closed-up symbol matched to
434
+ its opposite member) is true of every token that reaches here by construction. Additional guards:
363
435
 
364
436
  - **P5 (spec 0.6.0) — authored en-dash mark-identity veto.** If `k = 1` and `cp[s]` is U+2013,
365
437
  emit nothing — **unconditionally**: every locale, tight or spaced, regardless of
@@ -525,10 +597,14 @@ of being falsified by another rule, which is then named.
525
597
  - **[P] A negative number.** `-5` has no space to the left of the digit and a letter/space to
526
598
  the left of the hyphen → asymmetric → rejected. This holds for U+002D, U+2010 and U+2212
527
599
  alike — the same asymmetry guard, not a glyph-specific one.
528
- - **[P] Every digit-flanked stroke, unconditionally** — `5-10`, `Figure 5-10`, `9-11`: never
529
- reached by `dashes` at all, let alone reinterpreted as parenthetical (§1, §3.3). This holds
530
- regardless of whether `ranges`' own guards would have accepted or declined the token —
531
- `dashes` does not evaluate them and does not need to.
600
+ - **[P] Every range candidate, unconditionally** — `5-10`, `Figure 5-10`, `9-11`, and since
601
+ spec 1.3.0 also `$15-$20`, `$15 - $20` and `35%-50%`, where a matched `CLOSED-SYMBOL` sits
602
+ between the stroke and a digit run (ranges.md §3.2a): never reached by `dashes` at all, let
603
+ alone reinterpreted as parenthetical (§1, §3.3). This holds regardless of whether `ranges`'
604
+ own guards would have accepted or declined the token — `dashes` does not evaluate them and
605
+ does not need to. **`$15 - $20` is the one shape this cost anything**: through spec 1.2.0 it
606
+ was `dashes`' token and converted with default options; it is `ranges`' now, and `ranges` is
607
+ off by default.
532
608
  - **[P] URLs, code spans, fenced code, HTML attributes.** Removed by the mode adapter before this
533
609
  rule sees them. This rule has no notion of a URL and must not grow one.
534
610
  - **[P] Line terminators**, which are never inserted, deleted or crossed.
@@ -537,8 +613,9 @@ of being falsified by another rule, which is then named.
537
613
 
538
614
  ## 5. Idempotency argument
539
615
 
540
- Write `T` for `dashes`. `T` edits only **parenthetical dash tokens**: at least one of `cp[L]`,
541
- `cp[R]` is not `DIGIT` (§1, §3.3) — a digit-flanked token is never reached by `T` on any pass, so
616
+ Write `T` for `dashes`. `T` edits only **parenthetical dash tokens**: tokens that are not range
617
+ candidates (§1, §3.3) — at least one flank is neither a `DIGIT` nor a `CLOSED-SYMBOL` matched on
618
+ the opposite member (ranges.md §3.2a). A range candidate is never reached by `T` on any pass, so
542
619
  nothing below needs to reason about one. Each edit `T` makes replaces a span consisting of one
543
620
  maximal `DASH` run plus at most one space-like code point on each side, with a span of the same
544
621
  shape (`space? dash space?`). So every edit is one of:
@@ -566,9 +643,10 @@ U+2013/U+2014 no token-level special case, §3.2 step 2a, so this holds regardle
566
643
  glyph the token holds). A spaced form is re-admitted the same way, with `lsp = rsp = 1` read back
567
644
  by §3.2 step 3. A token `nbsp` has since promoted (`ru`: U+00A0 before an em dash) is not
568
645
  recomputed at all — §3.2 step 3 makes a no-break-space neighbour space-like, so the isolation
569
- guard (step 6) declines it and nothing is emitted (§3.6). A digit-flanked token stays outside `T`'s
646
+ guard (step 6) declines it and nothing is emitted (§3.6). A range candidate stays outside `T`'s
570
647
  domain on every pass, by construction, so it is trivially a fixed point of `T` regardless of what
571
- `ranges` does to it.
648
+ `ranges` does to it — including after `ranges` has bound it, since the binding leaves the flanks
649
+ (and any `CLOSED-SYMBOL` on them) exactly where they were.
572
650
 
573
651
  ### 5.2 A declined token stays declined
574
652
 
@@ -619,17 +697,25 @@ them, and all three are current:
619
697
  Range binding ([ranges.md](ranges.md) §3.3.1) is exclusively `ranges`' emission; an interrupting
620
698
  parenthetical dash is exactly where a line *may* break, so there is nothing for `dashes` to bind
621
699
  (§3.6).
622
- - **A digit-flanked token is never reinterpreted as parenthetical, on any pass.** The `isDigit`
623
- test that routes a token to `ranges` instead of `dashes` (§1, §3.3) is evaluated after the
624
- joiner-crossing walk (§3.2a), so a token `ranges` bound on an earlier pass — where the walk
625
- re-enters across the `JOINER` pair and finds `DIGIT` on both effective neighbours — presents to
626
- `dashes` with digit-flanked `leftCp`/`rightCp` exactly as an unbound one would, and `dashes`
627
- declines it on the same unconditional test. `dashes` cannot strip a binding it never
628
- reconsiders, and cannot create one, since it never emits `JOINER`.
700
+ - **A range candidate is never reinterpreted as parenthetical, on any pass.** The candidacy test
701
+ that routes a token to `ranges` instead of `dashes` (§1, §3.3, [ranges.md](ranges.md) §3.2,
702
+ §3.2a) is evaluated after the joiner-crossing walk (§3.2a), so a token `ranges` bound on an
703
+ earlier pass — where the walk re-enters across the `JOINER` pair and finds a candidate on both
704
+ effective neighbours — presents to `dashes` exactly as an unbound one would, and `dashes`
705
+ declines it on the same unconditional test. Since spec 1.3.0 that test reads `cp[L']`/`cp[R']`,
706
+ so a bound range carrying closed-up symbols re-enters on the same footing as an all-digit one.
707
+ `dashes` cannot strip a binding it never reconsiders, and cannot create one, since it never
708
+ emits `JOINER`.
629
709
  - **`dashes`' own edits cannot flip a *neighbouring* range token's guard verdict on the next
630
710
  pass, even though `dashes` itself never reads that token's guards.** The one edit shape that
631
711
  could — a tight token becoming spaced, inserting a U+0020 next to a digit run that has another
632
712
  dash on its far side — is exactly what the spacing-transition guard (T1, §3.2 step 8) declines.
713
+ **This clause is true of spec 1.3.0 only because step 8 was amended with it**: ranges.md §3.2a
714
+ widened "next to a digit run" to "next to a range member", which may carry one `CLOSED-SYMBOL`
715
+ at either end, and an unamended T1 reads that symbol as the neighbour and stops one code point
716
+ short of the dash. Four witnesses made the gap concrete — `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
717
+ `$1 - $1--a`, one per position and side — each of which drifted on the second pass before the
718
+ amendment and each of which now behaves exactly as its all-digit analogue always has.
633
719
  This is `dashes`' composition obligation *toward* `ranges`, symmetric to the rule-order argument
634
720
  in [ranges.md](ranges.md) §4: `ranges` must not see an adjacency `dashes` disturbed, and T1 is
635
721
  how `dashes` upholds that on every subsequent pass. The full historical derivation of why this
@@ -1097,7 +1183,12 @@ form to be `-spaced`. In other words: `A` is tight, `A` is about to become space
1097
1183
  only, the configuration T1 (§3.2 step 8) declares inert. The mirrored argument gives the `after`
1098
1184
  side. T1's two-code-point reach is exactly what the case requires and no more: `T`'s dash sits at
1099
1185
  distance 1 from `Lrun` if `T` is tight and distance 2 if `T` is spaced, and a `-spaced` replacement
1100
- inserts exactly one U+0020, so no other distance is reachable. The reverse flip — `before` going
1186
+ inserts exactly one U+0020, so no other distance is reachable. **Spec 1.3.0 leaves that reach at
1187
+ two code points and makes it transparent to one `CLOSED-SYMBOL` at each of two positions**
1188
+ (§3.2 step 8): the symbol moves where the reach starts, never how far it goes, because
1189
+ ranges.md §3.2a consumes at most one symbol per side. `JOINER` does not add a position either:
1190
+ it is transparent to this reach by §3.2b, on both branches, so a joiner run between the symbol
1191
+ and the dash collapses rather than counting. The reverse flip — `before` going
1101
1192
  from a space to a dash — is benign and needs no guard: it turns a G2 admission into a G2
1102
1193
  rejection, and a rejection emits nothing, so a token converted on an earlier pass simply keeps the
1103
1194
  form it was given.
@@ -0,0 +1,271 @@
1
+ # Rule: `ellipsis`
2
+
3
+ **Order:** 20. **Default:** on. **Modes:** text, html, markdown, yaml.
4
+ **Spec version:** 0.1.0.
5
+
6
+ ---
7
+
8
+ ## 1. Purpose
9
+
10
+ `ellipsis` replaces a typed run of full stops with the single character U+2026 (…), and
11
+ implements the abbreviated form that some locales use when an ellipsis follows terminal
12
+ punctuation — Russian writes `?..` and `!..` with two dots rather than three, because the
13
+ question mark or exclamation mark already occupies the first position of the three-dot
14
+ group. The rule is intentionally narrow: a run of exactly two full stops is _never_ touched,
15
+ because `..` occurs in relative paths, in version ranges and in numeric ranges written by
16
+ programmers, and converting it would be a false positive of exactly the class that blocks
17
+ the ship gate (PLAN.md M4). Everything this rule does is a pure function of a maximal run of
18
+ dot-like characters plus, at most, the one code point to its left.
19
+
20
+ ---
21
+
22
+ ## 2. Locale data consumed
23
+
24
+ - `ellipsis.abbreviatedAfterTerminal` — boolean, required by `locale.schema.json`.
25
+
26
+ **Evidence for locale data.** A locale attests **membership** — which tokens, which enum value, which boolean — and the rule owns the **mechanism** applied to them. A `sources` citation is not required to name a code point or a behaviour the locale file has no way to vary. Stated once, normatively, in [nbsp.md](nbsp.md) §2.1; it governs every locale field.
27
+
28
+ Nothing else. In particular this rule does not read `nbsp` or `quotes`.
29
+
30
+ ---
31
+
32
+ ## 3. Algorithm
33
+
34
+ Input is a code-point array `cp[0 … n-1]`.
35
+
36
+ ### 3.1 Character classes
37
+
38
+ | Class | Members |
39
+ | ---------- | ---------------------------- |
40
+ | `DOT` | U+002E (full stop) |
41
+ | `ELL` | U+2026 (horizontal ellipsis) |
42
+ | `DOTLIKE` | `DOT` ∪ `ELL` |
43
+ | `TERMINAL` | U+0021 (!) and U+003F (?) |
44
+
45
+ No other character is examined. U+2025 (‥, two-dot leader), U+22EF, U+FE19 and the CJK
46
+ leaders are **not** members of any class and are never produced or consumed.
47
+
48
+ ### 3.2 Ordering dependency
49
+
50
+ `spaces` (order 10) has already run and has deleted every U+0020 that stood immediately before
51
+ a **lone** U+002E. Consequently the spaced form `". . ."` reaches this rule as `"..."` and is
52
+ handled by the ordinary run logic. This rule therefore never needs to look across spaces, and
53
+ must not try to.
54
+
55
+ The word _lone_ is load-bearing and was added after a defect. `spaces` does **not** strip a
56
+ space before a run of two or more dots (`spaces.md` §3.4), so `"e.g. .."` and `"See ../docs"`
57
+ arrive here with their spaces intact. Before that condition existed the space vanished, the
58
+ two-dot run merged with the abbreviation's full stop into a three-dot run, and this rule then
59
+ correctly converted it — producing `e.g…` from input that §4 below promised to leave alone.
60
+ Neither rule was misbehaving; the protection simply was not true of the pipeline.
61
+
62
+ ### 3.3 Scan
63
+
64
+ 1. Set `i = 0`.
65
+ 2. If `cp[i]` is not in `DOTLIKE`, set `i = i + 1` and repeat. Terminate at `i = n`.
66
+ 3. Find the maximal run: `s = i`; `e` = the smallest index `> s` with `cp[e]` not in
67
+ `DOTLIKE` (or `e = n`). Let `k = e - s`, and let `d` = the number of `DOT` members and
68
+ `q` = the number of `ELL` members in the run (`d + q = k`).
69
+ 4. **Classify the run.**
70
+ - `k = 1` and `d = 1` → an ordinary full stop. Emit nothing. Go to 7.
71
+ - `k = 1` and `q = 1` → an existing ellipsis. Go to 6 (it may still need the abbreviated
72
+ form).
73
+ - `k = 2` and `q = 0` → **two full stops.** Go to 5.
74
+ - `k ≥ 2` and `q ≥ 1` → a mixed or repeated run (`"…."`, `"……"`, `"..…"`). Normalise:
75
+ emit an edit replacing `cp[s … e-1]` with one U+2026. Set the run to that single
76
+ U+2026 and go to 6.
77
+ - `k ≥ 3` and `q = 0` → three or more full stops. Emit an edit replacing
78
+ `cp[s … e-1]` with one U+2026. Set the run to that single U+2026 and go to 6.
79
+ 5. **The two-dot case.** Let `left = cp[s-1]` if `s > 0`, else `NONE`.
80
+ - If `ellipsis.abbreviatedAfterTerminal` is `true` → emit nothing, unconditionally. In
81
+ such a locale `"?.."` is the _correct_ output form and must survive re-processing.
82
+ - If `ellipsis.abbreviatedAfterTerminal` is `false` **and** `left` is in `TERMINAL` →
83
+ emit an edit replacing `cp[s … e-1]` with one U+2026. (`"?.."` in an English text is a
84
+ typing slip for `"?…"`.) Go to 7; step 6 cannot fire because the locale flag is false.
85
+ - Otherwise → emit nothing. This is the branch that protects `"../"`, `"1..5"` and
86
+ `"a..b"`.
87
+ Go to 7.
88
+ 6. **Abbreviated-after-terminal form.** The run is now a single U+2026 at position `s`
89
+ (either pre-existing or just produced by step 4).
90
+ - If `ellipsis.abbreviatedAfterTerminal` is `false` → emit nothing. Go to 7.
91
+ - Let `left = cp[s-1]` if `s > 0`, else `NONE`. If `left` is **not** in `TERMINAL` →
92
+ emit nothing. Go to 7.
93
+ - Otherwise emit an edit replacing the single U+2026 at `s` with the two code points
94
+ U+002E U+002E. Steps 4 and 6 are stages of one decision about one span: the rule emits
95
+ **exactly one edit per run**, whose replacement is the final form. `"?..."` is therefore
96
+ a single edit replacing three U+002E with two U+002E, not two chained edits.
97
+ 7. Set `i = e` and go to 2.
98
+
99
+ ### 3.4 Interaction with `!?` and `?!`
100
+
101
+ `left` in step 6 is a single code point. For `"?!..."` the character left of the ellipsis is
102
+ U+0021, which is in `TERMINAL`, so the abbreviated form fires: `"?!.."`. This matches the
103
+ Russian convention (`«Что?!..»`). No lookbehind beyond one code point is required or
104
+ permitted.
105
+
106
+ ---
107
+
108
+ ## 4. Must not touch
109
+
110
+ **Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
111
+ a guarantee of `transform` as a whole — or **[R]** — true of this rule alone. Every bullet here
112
+ is [P], and two of them only became true when `spaces` gained the lone-dot condition.
113
+
114
+ - **[P] A single U+002E.** Sentence full stops, decimal points, abbreviation dots (`p.`, `z.`),
115
+ and the dot in `"1.2.3"` are all runs of length 1.
116
+ - **[P] A run of exactly two U+002E**, except in the one narrow case in step 5 where the locale
117
+ does _not_ use the abbreviated form and the run directly follows `!` or `?`. This protects
118
+ `"../"`, `"./.."`, `"1..5"`, `"e.g. .."` and every other two-dot idiom.
119
+ _[P] only since `spaces` gained the lone-dot condition. `"e.g. .."` previously became `e.g…`
120
+ and `"See ../docs"` lost its space — the run itself was untouched throughout, which is exactly
121
+ what made the claim look true while the pipeline falsified it._
122
+ - **[P] U+2025 (‥) and the CJK dot leaders.** Not in any class; never produced, never consumed.
123
+ - **[P] Two `DOT` runs separated by anything at all.** Runs are maximal and are never joined
124
+ across an intervening character. After `spaces` has run, a surviving separator between
125
+ dots is a no-break space or a tab, i.e. something the author put there deliberately.
126
+ - **[P] Anything in a skipped region.** `"..."` inside a code span or a URL never reaches this
127
+ rule; the mode adapter removes it. This rule has no code-awareness and must not acquire
128
+ any.
129
+ - **[P] The spacing around an ellipsis.** Whether `"word …"` should carry a no-break space is
130
+ `nbsp`'s decision, driven by `nbsp.beforePunctuation` / `nbsp.narrowBeforePunctuation`,
131
+ which may list U+2026.
132
+
133
+ ---
134
+
135
+ ## 5. Idempotency argument
136
+
137
+ Write `T` for the rule. The output of `T` contains, at every position where `T` acted, one
138
+ of exactly two forms:
139
+
140
+ - **Form A: a lone U+2026** whose left neighbour is either absent or not in `TERMINAL`, or
141
+ whose locale has `abbreviatedAfterTerminal = false`.
142
+ - **Form B: exactly two U+002E** whose left neighbour is in `TERMINAL`, in a locale with
143
+ `abbreviatedAfterTerminal = true`.
144
+
145
+ Re-running `T`:
146
+
147
+ - Form A is a run with `k = 1`, `q = 1`. Step 4 sends it to step 6. Step 6 emits nothing
148
+ because either the flag is false or `left` is not in `TERMINAL` — both of which are exactly
149
+ the conditions under which Form A was produced. No edit.
150
+ - Form B is a run with `k = 2`, `q = 0`. Step 5 is reached, and the first branch applies
151
+ (flag is `true`) → emit nothing, unconditionally. No edit.
152
+ - Runs the rule declined to touch (single dot, two dots outside the special case) are
153
+ classified identically on the second run, because classification depends only on the run
154
+ itself and on one left neighbour, neither of which this rule altered — the rule never
155
+ edits a character outside a `DOTLIKE` run, and a `DOTLIKE` run is never adjacent to
156
+ another `DOTLIKE` run after `T` (runs are maximal, so merging cannot occur).
157
+
158
+ Hence `T(T(x)) = T(x)`.
159
+
160
+ **What had to be fixed.** The naive formulation is _"`...` → `…`, and in Russian `?…` →
161
+ `?..`"_. That is not idempotent, because the output `?..` is a two-dot run which the same
162
+ naive rule's companion clause _"collapse repeated dots"_ would then re-expand — or, in the
163
+ formulation used by several existing libraries, `?..` → `?…` → `?..` → … oscillating.
164
+ Two changes fix it:
165
+
166
+ 1. The two-dot run is **unconditionally inert** when `abbreviatedAfterTerminal` is true. The
167
+ output form is a fixed point by construction rather than by luck.
168
+ 2. The two-dot → ellipsis conversion exists **only** in the complementary locale setting and
169
+ **only** after terminal punctuation, so the two branches can never both apply to the same
170
+ text under the same locale. The flag partitions the behaviour; it does not layer it.
171
+
172
+ ---
173
+
174
+ ### Composition obligation
175
+
176
+ Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₂**, so the
177
+ obligation runs against `spaces` (R₁) only.
178
+
179
+ **What this rule emits.** One U+2026 replacing a run of U+002E/U+2026, or two U+002E replacing
180
+ one U+2026. Nothing else: no spaces, no letters, no dashes, and never a code point outside
181
+ `DOTLIKE`.
182
+
183
+ **Against `I₁` (`spaces`).** Discharged. This rule never emits U+0020, so violations S-a, S-c
184
+ and S-d are unreachable. S-b — a U+0020 whose right neighbour is in `STRIP-BEFORE` — deserves a
185
+ second look, because U+002E and U+2026 are both in `STRIP-BEFORE` and this rule moves them
186
+ about. But every replacement starts at the first code point of the run and the code point to
187
+ its left is unchanged; if that neighbour were a U+0020, `spaces` would already have deleted it
188
+ in the same pass, since the run began with a `DOTLIKE` character then too. Shortening a run
189
+ cannot bring a space into contact with a dot that was not already in contact with one.
190
+
191
+ **`I₂` in the other direction** is preserved by every later rule trivially: none of them emits
192
+ U+002E or U+2026, and none deletes a code point standing between two dot runs.
193
+
194
+ ---
195
+
196
+ ## 6. Worked examples
197
+
198
+ `⟶` = no change. Both locale settings are shown because the rule is one of only two whose
199
+ output differs on a boolean.
200
+
201
+ ### `ellipsis.abbreviatedAfterTerminal = false` (en, fi, sv, de, fr, el)
202
+
203
+ | # | Input | Output | Why |
204
+ | --- | --------------- | ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
205
+ | 1 | `Wait... what?` | `Wait… what?` | `k = 3`, `q = 0` → single U+2026 |
206
+ | 2 | `Wait…… what?` | `Wait… what?` | mixed run normalised |
207
+ | 3 | `Really?.. ` | `Really?… ` | two dots after `?` in a non-abbreviating locale |
208
+ | 4 | `See ../docs` | ⟶ | two-dot run, left neighbour U+0020 not in `TERMINAL`. **The space also survives**, which is `spaces.md` §3.4's doing, not this rule's — before that condition the pipeline returned `See../docs` |
209
+ | 4a | `e.g. ..` | ⟶ | likewise: the space survives, so no three-dot run is ever formed. Previously `e.g…` |
210
+ | 5 | `Version 1..5` | ⟶ | two-dot run, left neighbour is a digit |
211
+ | 6 | `He left…` | ⟶ | already a lone U+2026 |
212
+ | 7 | `Hmm.....` | `Hmm…` | `k = 5` → one U+2026 |
213
+ | 8 | `Yes. No.` | ⟶ | two independent runs of length 1 |
214
+
215
+ ### `ellipsis.abbreviatedAfterTerminal = true` (ru)
216
+
217
+ | # | Input | Output | Why |
218
+ | --- | ------------ | ---------- | ------------------------------------------------- |
219
+ | 9 | `Что?...` | `Что?..` | run → U+2026 → abbreviated after `?` |
220
+ | 10 | `Что?…` | `Что?..` | existing U+2026 after `?` rewritten |
221
+ | 11 | `Что?..` | ⟶ | already the target form; step 5 first branch |
222
+ | 12 | `Что?!...` | `Что?!..` | left neighbour of the run is U+0021 |
223
+ | 13 | `Он ушёл...` | `Он ушёл…` | left neighbour is a letter → ordinary ellipsis |
224
+ | 14 | `см. ../` | ⟶ | two-dot run, unconditionally inert in this locale |
225
+
226
+ Cases 4, 5, 6, 8, 11 and 14 are "no change" cases.
227
+
228
+ ---
229
+
230
+ ## 7. Open questions
231
+
232
+ 1. **Four dots.** Chicago (and several book styles) write `"….",` an ellipsis followed by a
233
+ sentence-ending period, for an omission at the end of a sentence. This rule collapses
234
+ `"...."` to a single `"…"`, losing that distinction. Distinguishing the two requires
235
+ knowing whether the omission ends the sentence, which is not decidable from the code
236
+ points. I chose collapse; an alternative is to leave runs of exactly four alone. Needs an
237
+ operator decision and a fixture either way.
238
+ 2. **`abbreviatedAfterTerminal` is a single boolean**, so a locale cannot say "abbreviate
239
+ after `?` but not after `!`". No known locale needs that, but the schema forecloses it.
240
+ Recorded, not proposed.
241
+ 3. **The `!..` / `?..` order.** Russian also uses `"..?"` and `"..!"` in some sources for an
242
+ ellipsis _preceding_ terminal punctuation. This rule does not implement that direction,
243
+ and `"..?"` is left alone (two-dot run). Whether Мильчин requires the leading form is a
244
+ locale-research question, and if the answer is yes, the schema needs a second flag.
245
+ 4. _(Partly settled.)_ The `nbsp` asymmetry recorded here is fixed on the `nbsp` side: U+2026
246
+ is now accepted right context for N1/N2 (`nbsp.md` §3.3 step 2), so French `Vraiment?…`
247
+ takes its narrow space just as `Vraiment ?` does. What remains open is the original
248
+ observation — **The interaction with `nbsp` is unspecified from this side.** If a locale lists U+2026 in
249
+ `nbsp.beforePunctuation`, then `"word …"` gains a no-break space, but the abbreviated
250
+ Russian form `"?.."` ends in U+002E and would not. That asymmetry may or may not be
251
+ correct; it is a question for the `ru` locale file, not for this rule.
252
+ 5. **Two dots in a non-abbreviating locale after `?`/`!`** (case 3) is the only place this
253
+ rule converts a two-dot run. It is conservative and safe, but it is also the only asymmetry
254
+ in the rule. If it produces a false positive in real content, deleting the branch costs
255
+ nothing.
256
+ 6. **Greek forbids a space before the ellipsis; polytypo preserves it anyway.** The EU
257
+ Interinstitutional Style Guide (Greek edition) §10.1.9 ii) — «Μεταξύ των αποσιωπητικών και της
258
+ λέξης που προηγείται δεν αφήνουμε διάστημα» — is verified against the source and is not
259
+ honoured. `spaces.md` §3.4 preserves a space before a dot run in every locale, so `Πράγματι …`
260
+ comes back as typed.
261
+
262
+ The full argument is in `spaces.md` §7.9 and is not repeated; the part that belongs **here** is
263
+ why this rule does not fix it instead. This rule reads locale data and could carry an
264
+ `ellipsis.noSpaceBefore` flag, which the schema can express and which fits the declarative-data
265
+ line of PLAN.md §6. What that would cost is a change of kind, not of degree: **this rule
266
+ currently only ever replaces one span of dot-like code points with another, and it would become
267
+ a rule that deletes a character outside its own run.** Deleting brings in `modes.md` §3.3's
268
+ edge-test clause (a deleting rule must treat a span edge as the end of the text), a new
269
+ discharge against `I₁`, and a re-derivation of §5's idempotency argument, which currently
270
+ depends on the run being the only thing touched. That is not worth it for one locale, and it is
271
+ worth it for two. Recorded so the option is costed rather than forgotten.