polytypo 1.2.0 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +626 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,489 @@
1
+ # Rule: `ranges`
2
+
3
+ **Order:** 25. **Default:** off. **Modes:** text, html, markdown, yaml.
4
+ **Spec version:** 0.5.0 (new rule; split out of `dashes`); §3.2a (closed-up symbols) new in 1.3.0.
5
+
6
+ ---
7
+
8
+ ## 1. Purpose
9
+
10
+ `ranges` recognises the numeric/date **range dash** — `1914–1918` — and renders it in the form
11
+ the locale prescribes. It is the second of the two typographic uses `dashes` (order 30) used to
12
+ own; the two rules together still cover exactly the same input shape `dashes` alone used to
13
+ (spec ≤ 0.4.1), split apart because they now carry **different default-on status**.
14
+
15
+ **Why this is a separate opt-in rule rather than a bounded fix inside `dashes`.** A digit-flanked
16
+ hyphen shaped `5-10` is genuinely ambiguous without knowing the word that precedes it:
17
+
18
+ - `takes 5-10 days`, `aged 9-10 years`, `pages 5-10`, `0-60`, `€30-80` — genuine ranges. Converting
19
+ the hyphen to the locale's dash is correct.
20
+ - `Figure 5-10`, `Table 3-12`, `Section 2-14` — **compound labels**, not ranges (they mean "figure
21
+ 10 in chapter 5" — chapter-and-item, not a span). Converting the hyphen here is a false positive
22
+ that reads oddly to a careful proofreader, even though it is rarely noticed casually.
23
+ - `9-11`, `7-11` — bare proper nouns (a date, a convenience-store chain) that happen to have the
24
+ identical digit-hyphen-digit shape and are neither a label nor a range.
25
+
26
+ Separating the first case from the second and third needs the word immediately before the
27
+ hyphen's left digit run — `Figure`, `Table`, `Section`, `Fig.`, `Abb.`, `рис.` — which is an
28
+ open-ended, per-locale, per-house-style word list with no natural closure
29
+ ([dashes.md](dashes.md) §7.11, carried forward unchanged from when this algorithm lived there).
30
+ `locale.schema.json` could hold such a list as literal strings, but populating it would mean
31
+ either (a) inventing entries with no normative citation — forbidden by this project's evidence
32
+ discipline (PLAN.md §6.1) — or (b) leaving it empty, which is not a fix, only a schema field that
33
+ looks like one. Neither the bare digit-hyphen-digit shape nor a label word list gives a
34
+ **structurally closed, portable** rule that never produces a false positive: this project's rules
35
+ scan code points, not words with real-world referents, and are built never to consult that kind
36
+ of context (ARCHITECTURE.md §4.1).
37
+
38
+ **The resolution is: make range conversion something the caller explicitly asks for, and leave it
39
+ off by default.** With `ranges` off (the default), `5-10`, `Figure 5-10`, `9-11` and `7-11` are
40
+ all byte-identical no-ops — the compound-label false positive cannot occur, because nothing in
41
+ this class converts at all. A caller who explicitly opts in (`{ rules: { ranges: true } }`) is
42
+ choosing to accept the residual, structurally irreducible ambiguity between a genuine range and a
43
+ same-shaped label or proper noun, in exchange for genuine ranges being typeset correctly. This is
44
+ a documented tradeoff, not a claim of reliable label detection — see §5.
45
+
46
+ ---
47
+
48
+ ## 2. Locale data consumed
49
+
50
+ - `dash.range` — one of `"em-tight"`, `"em-spaced"`, `"en-tight"`, `"en-spaced"`, `"none"`. The
51
+ **same field**, under the **same locale JSON key**, `dashes` has always read for range styling
52
+ — this rule's split from `dashes` moves which rule reads it, not what the field means or which
53
+ citations support it. No new locale claim is made by this document (operator decision, spec
54
+ 0.5.0): every `dash.range` value in `spec/locales/*.json` is unchanged from spec 0.4.1.
55
+
56
+ `"none"` means the locale has no verified range convention, and this rule must emit nothing at
57
+ all for a range token in that locale — not a fallback to `dash.parenthetical`, nothing.
58
+
59
+ ---
60
+
61
+ ## 3. Algorithm
62
+
63
+ Input is a code-point array `cp[0 … n-1]`. Character classes (`DASH`, `DIGIT`, `LETTER`, `SPACE`,
64
+ `BREAK`, `NOBREAK-SPACE`, `INERT-DASH`, `JOINER`) are exactly [dashes.md](dashes.md) §3.1's —
65
+ this rule and `dashes` read the same input alphabet, because both scan the same DASH-token shape
66
+ before diverging on what each does with it.
67
+
68
+ ### 3.1 The dash token (shared with `dashes`)
69
+
70
+ Steps 1-7 of [dashes.md](dashes.md) §3.2 — run detection, the length-3 decoration cutoff, outer
71
+ spacing, the symmetry guard, the content-on-both-sides guard, the joiner-crossing walk (§3.2a),
72
+ the isolation guard and the cluster guard (§3.2 step 7) — apply here **identically, unchanged**.
73
+ A token that `dashes`' §3.2 steps would decline is declined here too, for the same reasons; the
74
+ two rules diverge only starting at §3.3 below. The JS reference implementation shares one
75
+ internal module (`src/rules/dash-shared.ts`) between `src/rules/dashes.ts` and
76
+ `src/rules/ranges.ts` for exactly this reason — a port is free to structure its own module
77
+ boundary differently, but must reproduce the same shared guard behaviour in both rules, not two
78
+ independent approximations of it.
79
+
80
+ `JOINER`-crossing re-entry (dashes.md §3.2a's third bullet: "if a joiner was crossed and
81
+ `cp[L*]`/`cp[R*]` are both `DIGIT`, the token is a bound range this rule produced on an earlier
82
+ pass") is **this rule's own** re-entry condition — it is how a tight range this rule already
83
+ converted survives a second pipeline pass without regrowing a second joiner. `dashes` shares the
84
+ same joiner-crossing walk (it must, to correctly decline a token immediately touching a range
85
+ this rule already produced) but never itself owns that re-entry: any joiner adjacent to a token
86
+ that is **not a range candidate** (§3.2, §3.2a — wider than "digit-flanked" as of spec 1.3.0) is
87
+ declined outright (dashes.md §3.2a's fourth bullet).
88
+
89
+ ### 3.2 Range guards (G1-G5)
90
+
91
+ The token is a **range candidate** iff `cp[L']` is in `DIGIT` **and** `cp[R']` is in `DIGIT` —
92
+ `L` and `R` after the joiner-crossing walk of §3.1, and `L'`/`R'` after the closed-up-symbol walk
93
+ of §3.2a, which moves neither index unless a symbol is there to consume. If it is a range
94
+ candidate, this rule owns it exclusively: `dashes` never processes it, whether or not `ranges` is
95
+ enabled (operator decision, spec 0.5.0 — see [dashes.md](dashes.md) §1). If it is **not** a range
96
+ candidate (either flank is not `DIGIT` after that walk), this rule emits nothing for it at all;
97
+ it is `dashes`' concern or no rule's.
98
+
99
+ Compute `Lrun`, `Rrun`, `before`, `after` exactly as [dashes.md](dashes.md) §3.3 (pre-0.5.0)
100
+ specified — reproduced here verbatim since this is now its own rule document, with `L'`/`R'`
101
+ where it said `L`/`R`:
102
+
103
+ - `Lrun` = the maximal run of `DIGIT` ending at `L'` (indices `a … L'`);
104
+ - `Rrun` = the maximal run of `DIGIT` starting at `R'` (indices `R' … b`);
105
+ - `before` = the **effective neighbour** (dashes.md §3.2b) to the left of index `a`, except as
106
+ §3.2a amends it;
107
+ - `after` = the **effective neighbour** to the right of index `b`, except as §3.2a amends it.
108
+
109
+ All five guards must pass:
110
+
111
+ - **G1 — no letter adjacency.** `before` is not in `LETTER`. (`MP3-4`, `H2-2` rejected. `COVID-19`
112
+ is already rejected earlier, because `cp[L]` is `D`, a letter, not a digit.)
113
+ - **G2 — no chain.** `before` is not in `DASH` and not in `INERT-DASH`; `after` is not in `DASH`
114
+ and not in `INERT-DASH`. This protects `2026-08-15`, `978-3-16-148410-0` and `212-555-1234`.
115
+ **This guard reads the input as it stood before this rule (or `dashes`) made any edit in this
116
+ pipeline pass** — see §4's note on why `ranges` runs before `dashes`: evaluating G2 against an
117
+ adjacency `dashes` just created (by removing a space next to this token) rather than against
118
+ what the author actually typed is exactly the defect §4 documents and the reason for the chosen
119
+ order.
120
+ - **G3 — not part of a decimal or path.** `before` is not U+002E, U+002C or U+002F; `after` is
121
+ not U+002F. `1.5-2.5` and `01/02-03/04` are left alone.
122
+ - **G4 — run lengths.** Either `length(Lrun) = length(Rrun)`, **or** `length(Lrun) = 1` and
123
+ `length(Rrun) = 2` and `Rrun` does not begin with U+0030. `1914-1918` (4,4) passes; `5-10` (1,2)
124
+ passes; `0-60` (1,2) passes; `ISO 8859-1` (4,1) fails; `555-1234` (3,4) fails; `1-800` (1,3)
125
+ fails; `2020-24` (4,2) fails; `10-7` (2,1) fails; `9-05` fails on the leading zero. See
126
+ [dashes.md](dashes.md) §3.3's historical discussion of why this exact shape and not a wider one
127
+ — unchanged, reproduced there rather than duplicated here since it is a proof about G5's
128
+ correctness, not new content this rule's split needs to restate.
129
+ - **G5 — non-decreasing.** The integer value of `Lrun` ≤ the integer value of `Rrun`, compared
130
+ digit-by-digit left to right (sound only because G4 guarantees equal length in the only branch
131
+ where G5 does work). `1914-1918` passes; `1234-5678` passes; `20-10` fails.
132
+
133
+ If any guard fails, emit nothing. **The token is not reconsidered by `dashes`.** A stroke between
134
+ two digits is a range or it is nothing; `dashes` never sees it (§3.1 above).
135
+
136
+ ### 3.2a Closed-up symbols on both members (spec 1.3.0)
137
+
138
+ Before spec 1.3.0 a range whose members each carried a symbol — `$15-$20`, `35%-50%` — was not a
139
+ range candidate, because the flank next to the dash was the symbol and not a `DIGIT`. `$15-20`
140
+ converted and `$15-$20` did not; `15-20%` converted and `15%-20%` did not. The asymmetry was a
141
+ consequence of where the symbol sits relative to the digit run, never a decision anyone made.
142
+
143
+ **The distinction the source draws is closed-up versus spaced, not currency versus unit.**
144
+
145
+ > "the abbreviation or symbol is repeated if it is closed up to the number but not if it is
146
+ > separated: 35%–50%" — *The Chicago Manual of Style*, 18th ed., 9.19, quoted in the freely
147
+ > readable [CMOS Online Q&A, "Numbers"](https://www.chicagomanualofstyle.org/qanda/data/faq/topics/Numbers/faq0024.html);
148
+ > the same paragraph is applied to money on
149
+ > [page 5 of that topic](https://www.chicagomanualofstyle.org/qanda/data/faq/topics/Numbers.html?page=5):
150
+ > "in Chicago style an abbreviation or symbol is repeated if it is closed up to a number but not
151
+ > if it is separated by a space: $3–$5 million".
152
+
153
+ So `$15–$20` and `35%–50%` are one rule, and `15 kg–20 kg` is not that rule — `kg` is separated
154
+ by a space, and the source says a separated abbreviation is **not** repeated. The elided forms
155
+ `$3–5 million` and `15–20%` are permitted variants the same answer calls acceptable; they already
156
+ convert and keep converting.
157
+
158
+ **`CLOSED-SYMBOL`** is a literal code-point set, fixed here and not locale data:
159
+
160
+ > U+0024, U+00A2, U+00A3, U+00A4, U+00A5, every code point from U+20A0 through U+20CF inclusive
161
+ > (the whole Currency Symbols block, assigned or not), U+0025, U+2030, U+2031, and U+00B0.
162
+
163
+ `$` and `%` are the source's own examples; the rest are the same class by the source's own
164
+ wording — symbols conventionally written closed up to a number: the other currency signs,
165
+ per-mille and per-ten-thousand, and the degree sign. It is written as literal code points rather
166
+ than a Unicode category test for the same reason `DIGIT` is ASCII-only (§7.1): a category test
167
+ makes the rule's verdict depend on which Unicode version a runtime was built against, and the
168
+ five runtimes must agree. The currency range is the **block's own bounds**, U+20A0–U+20CF, not
169
+ the subset assigned in some Unicode version — block bounds never move, while the assigned subset
170
+ does, and a rule that admitted only today's assignments would drift between runtimes built
171
+ against different UCD releases. Unassigned code points inside the block are members of the set
172
+ and unreachable in valid text; a code point assigned there later is a currency sign by the
173
+ block's own definition, and membership is then already correct without a spec change. `+`, `-`, `#`, `(`, `)` and the quotation marks are **not** members and
174
+ must not be added on the reasoning that they are also written closed up: none is a symbol this
175
+ source's rule is about, and each would admit a shape (`+15-+20`, `#15-#20`, `"15"-"20"`) nobody
176
+ has asked for and no citation supports.
177
+
178
+ **The walk.** Both sides are decided **from the original `L` and `R`, simultaneously**, before
179
+ either index moves — never left-then-right or right-then-left. Otherwise a token whose flanks are
180
+ both in `CLOSED-SYMBOL` (`%15%-%20%`) would have a verdict that depends on evaluation order, and
181
+ two runtimes could disagree while both following this document. After the joiner-crossing walk of
182
+ §3.1 has produced `L` and `R`:
183
+
184
+ - **right:** if `cp[R]` is in `CLOSED-SYMBOL` **and** `cp[R + 1]` exists and is in `DIGIT`, then
185
+ `R' = R + 1` and the **inner right symbol** is `cp[R]`. Otherwise `R' = R` and there is none.
186
+ - **left:** if `cp[L]` is in `CLOSED-SYMBOL` **and** `cp[L - 1]` exists and is in `DIGIT`, then
187
+ `L' = L - 1` and the **inner left symbol** is `cp[L]`. Otherwise `L' = L` and there is none.
188
+
189
+ Each side consumes **at most one** code point, so `US$15-US$20` and `15°C-20°C` are not
190
+ admitted — and the mechanism is candidacy, not a guard: `cp[R]` is `U` and `cp[L]` is `C`,
191
+ neither is in `CLOSED-SYMBOL`, so no walk is taken and the flank is simply not a `DIGIT`. The
192
+ half-written `US$15-$20` **is** a candidate (the right walk matches the outer `$` in front of
193
+ `15`) and is declined by G1 instead, because `before` reads past that `$` and finds `S`. Both are
194
+ recorded in §7.7.
195
+
196
+ **The outer symbols** are the effective neighbour (dashes.md §3.2b) immediately left of `Lrun`'s
197
+ first index `a`, and immediately right of `Rrun`'s last index `b`, when that code point is in
198
+ `CLOSED-SYMBOL`.
199
+
200
+ **Matching, and it is exact.** A side's walk is taken **only if** the symbol it would consume is
201
+ matched on the opposite member: an inner right symbol must be the **same code point** as the
202
+ outer left symbol, and an inner left symbol the same code point as the outer right symbol. Not a
203
+ currency-equivalence table, not a per-locale list: the same code point.
204
+
205
+ **An unmatched symbol means the walk is not taken at all**, so the flank stays a non-`DIGIT`, the
206
+ token is **not** a range candidate, and it remains `dashes`' concern exactly as it was before
207
+ spec 1.3.0. `$15-€20` is a currency conversion, not a range; `15-$20` and `15%-20` are
208
+ half-written. None of the three changes hands, and none of them changes behaviour in this spec
209
+ version.
210
+
211
+ If a side has **no inner symbol**, the corresponding outer symbol is not examined at all, which
212
+ is exactly what keeps `$15-20` and `15-20%` converting as they did before this section existed.
213
+
214
+ **What this does not change.** `Rrun` is the maximal `DIGIT` run starting at `R'` and `Lrun` the
215
+ one ending at `L'`; wherever the shared guards of §3.1 and the replacement of §3.3 say `cp[L]` or
216
+ `cp[R]`, this rule reads `cp[L']` or `cp[R']`. **G4 and G5 are untouched** — they compare digit
217
+ runs and never see a symbol. The edit span is untouched too: the replacement covers the dash run
218
+ (and, when binding, the joiners adjacent to it), so a symbol is never inserted, removed or
219
+ rewritten. This rule changes one dash and nothing else.
220
+
221
+ **G1, G2 and G3 judge the same position they always did.** When an inner right symbol was
222
+ matched, `before` is the effective neighbour left of the **outer left symbol** rather than left of
223
+ `a`; when an inner left symbol was matched, `after` is the effective neighbour right of the
224
+ **outer right symbol** rather than right of `b`. In every other case `before` and `after` are
225
+ unchanged. So `US$15-$20` has `before` = `S`, a `LETTER`, and G1 declines it.
226
+
227
+ **Idempotency, and the two shared guards this section forced open.** A bound `$15⁠–⁠$20`
228
+ re-enters through §3.1's joiner-crossing walk, whose re-entry condition
229
+ ([dashes.md](dashes.md) §3.2a) is amended in this same spec version to read `cp[L']`/`cp[R']`;
230
+ having re-entered, §3.3's identical-replacement test makes the second pass a no-op.
231
+
232
+ That is the easy half. The hard half is that widening what counts as a range member widens what
233
+ a **`dashes`** edit elsewhere in the text can disturb, and `dashes`' spacing-transition guard was
234
+ keyed to `DIGIT`.
235
+
236
+ **T1** ([dashes.md](dashes.md) §3.2 step 8) both gated its branches on a `DIGIT` flank and walked
237
+ outward from the digit run without stepping over a symbol, so a tight `dashes` token next to a
238
+ closed-up range member became spaced, that U+0020 replaced the range's `before` or `after`, and
239
+ the range converted on the **next** pass. Four witnesses, one per position and side, each drifting
240
+ where its all-digit analogue was already inert: `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
241
+ `$1 - $1--a`. T1's reach is `CLOSED-SYMBOL`-transparent at two positions per side as of spec
242
+ 1.3.0, and each witness is pinned by a fixture asserting that the input is now a **no-op** —
243
+ which is what a conformance runner can see, since it cannot run a second pass and compare.
244
+
245
+ **This is deliberately T1 and not the cluster guard** ([dashes.md](dashes.md) §3.2 step 7), which
246
+ would have closed the same defect by putting `CLOSED-SYMBOL` in its alphabet. That guard is
247
+ unconditional, so it would also have made `price--$50--drop` inert in an `em-tight` locale that
248
+ has no defect to fix; T1 fires only on the tight-to-spaced transition that can actually disturb a
249
+ neighbour. The alternative was implemented, measured and rejected on that comparison.
250
+
251
+ **The cost this leaves.** In a locale whose `dash.parenthetical` is **spaced**, a tight token
252
+ next to a closed-up range member no longer converts: `Anstieg--50%--war` is left alone in
253
+ `de-DE`, where spec 1.2.0 gave `Anstieg – 50% – war` (measured). The all-digit
254
+ `Anstieg--50--war` has always been left alone, so the two agree; and a locale with a tight
255
+ parenthetical form never reaches the guard, so `price--$50--drop` still becomes `price—$50—drop`
256
+ in `en-US`. **The remedy is unconditional on `dash.range`, while the defect was not**: the
257
+ conversion loss therefore also lands in `fr`, `fr-CA`, `it`, `nl`, `pt-BR` and `pt-PT`, whose
258
+ `dash.range` is `"none"` and which never had a range verdict to disturb. Making T1 consult
259
+ `dash.range` would make a `dashes` verdict depend on a field `dashes` does not read, which is a
260
+ worse trade than a conversion nobody has asked for in six locales.
261
+
262
+ **The accepted cost, stated plainly.** Widening candidacy moves tokens **out of** `dashes`, and
263
+ `ranges` is off by default, so text that `dashes` used to change now changes only when a caller
264
+ turns `ranges` on — and in the eight locales whose `dash.range` is `"none"`, not even then:
265
+ `$15 - $20` gave `$15 — $20` in `fr` through spec 1.2.0 and is a permanent no-op from 1.3.0. The affected shapes are the **spaced and multi-hyphen** ones, because a lone
266
+ tight hyphen between two non-spaces was never converted by `dashes` either: `$15 - $20` gave
267
+ `$15—$20` in `en-US` before this section and is a no-op with default options after it, while
268
+ `$15-$20` was already a no-op and now converts when `ranges` is on. That is the same behaviour `$15 - 20` has had since spec 0.5.0 — the
269
+ change makes the two consistent rather than introducing an inconsistency — and the alternative is
270
+ worse than the cost: declining `$15-$20` does not leave `$15–20`, it leaves `$15-$20`, a hyphen
271
+ where every reading of every source above wants a dash.
272
+
273
+ ### 3.3 Replacement
274
+
275
+ If all guards pass, emit one edit replacing `cp[s - lsp … e - 1 + rsp]` with:
276
+
277
+ | `dash.range` | replacement |
278
+ | ------------- | -------------------------------------------------------- |
279
+ | `"em-tight"` | U+2014 |
280
+ | `"em-spaced"` | U+0020 U+2014 U+0020 |
281
+ | `"en-tight"` | U+2013 |
282
+ | `"en-spaced"` | U+0020 U+2013 U+0020 |
283
+ | `"none"` | _(emit nothing — the token is left exactly as written)_ |
284
+
285
+ Subject to guards T1 and T2 ([dashes.md](dashes.md) §3.2 steps 8-9, `isSpacingTransitionBlocked`
286
+ and the strip-before/open-bracket checks — shared, unchanged, evaluated here against
287
+ `dash.range`'s chosen form exactly as they were against either branch's form pre-0.5.0). If the
288
+ replacement is identical, code point for code point, to the span it would replace, emit nothing
289
+ instead.
290
+
291
+ #### 3.3.1 Binding a tight range
292
+
293
+ Unchanged from [dashes.md](dashes.md) §3.3.1 (pre-0.5.0), reproduced here since binding is now
294
+ exclusively this rule's behaviour: when the chosen `dash.range` form is `-tight` and all guards
295
+ have passed, the emitted replacement is `JOINER dash JOINER` (U+2060, the dash, U+2060), and the
296
+ edit span covers any `JOINER` already adjacent to the token — the range is a single lexical unit
297
+ (UAX #14 gives U+2013/U+2014 the line-break class `BA`, and Мильчин and Chicago both forbid
298
+ breaking a range in running text). A range already in its exact target form is never bound (the
299
+ invisible-edit test: compute the unbound replacement first; if it is already a no-op, stay
300
+ unbound). `dashes` never emits a joiner — an interrupting parenthetical dash is exactly where a
301
+ line **may** break.
302
+
303
+ ### 3.4 Continue, non-overlap, no-break spaces
304
+
305
+ [dashes.md](dashes.md) §3.5 (continue/non-overlap) and §3.6 (no-break spaces are never touched)
306
+ apply to this rule exactly as written there — both are properties of the shared token shape
307
+ (§3.1), not of which branch/rule a token ends up in.
308
+
309
+ ---
310
+
311
+ ## 4. Rule order: why `ranges` runs before `dashes`
312
+
313
+ `ranges` is order 25, `dashes` is order 30 — `ranges` runs first. This was chosen from an
314
+ observed behavioural difference, not for convenience.
315
+
316
+ The pre-0.5.0 unified `dashes` rule found every dash token's edits in **one scan over the
317
+ unedited input**, then applied all of them together — so every token's guards, including G2's
318
+ cross-token "no chain" check, always read the original, unedited adjacency structure, regardless
319
+ of which other tokens in the same document were also converting. Splitting `ranges` and `dashes`
320
+ into two sequential rule passes means one of them necessarily sees the other's *output*, not the
321
+ original input, for whichever positions the first rule touched.
322
+
323
+ **Concrete case: `a - 5-10`** (a parenthetical dash directly followed, by one space, by a
324
+ genuine range). `en-US`: `dash.parenthetical = "em-tight"`, `dash.range = "en-tight"`.
325
+
326
+ - **`ranges` first:** `ranges` scans the original text; the range token's `before` (G2) is the
327
+ space at index 3, not a dash, so G2 passes and `5-10` becomes `5⁠–⁠10`. `dashes` then scans
328
+ `a - 5⁠–⁠10`; its own token (`a - 5`) is unaffected by the far-away range edit and converts to
329
+ `a—5⁠–⁠10` normally. **Final: `a—5⁠–⁠10`.**
330
+ - **`dashes` first:** `dashes` converts `a - 5` to `a—5` (its `em-tight` form has no spaces, so
331
+ the edit span swallows the space that used to separate the parenthetical dash from the digit
332
+ run). `ranges` then scans `a—5-10`; the range token's `before` (G2) is now the em-dash
333
+ `dashes` just emitted, immediately adjacent (no space) — G2 reads this as a **chain** (the same
334
+ shape as `978-3-16-148410-0`) and declines. **Final: `a—5-10`** — the range is silently never
335
+ converted, purely because of which rule happened to run first, not because of anything either
336
+ rule's own guards were designed to protect.
337
+
338
+ `ranges`-before-`dashes` reproduces the pre-0.5.0 unified rule's output (`a—5⁠–⁠10`) exactly;
339
+ `dashes`-before-`ranges` introduces a new, order-induced decline that never happened before the
340
+ split. `tests/rules/ranges.test.ts` and `tests/rules/dashes.test.ts` pin this case as a fixed
341
+ regression witness. This is a single reproducible counter-example, not an exhaustive proof that
342
+ no input ever favours the opposite order — but it is evidence in one concrete direction, and
343
+ `ranges`-before-`dashes` is kept on that evidence, not on a preference between two otherwise
344
+ indistinguishable choices.
345
+
346
+ ---
347
+
348
+ ## 5. What this rule does not, and cannot, solve
349
+
350
+ **`ranges` converts `Figure 5-10` to `Figure 5⁠–⁠10` when explicitly enabled**, exactly as it
351
+ converts a genuine range. This is not an oversight this document is unaware of — it is the
352
+ residual ambiguity §1 names as the reason the rule is opt-in. Enabling `ranges` is an explicit
353
+ choice to accept that a compound label sharing the digit-hyphen-digit shape will be converted
354
+ alongside genuine ranges; the caller has no bounded, cited mechanism this project can offer to
355
+ tell the two apart (§1). A future spec version could add a cited, closed `dash.labelWords` list
356
+ per locale (dashes.md §7.11's own suggestion) if and when locale-authority research produces
357
+ citable evidence for specific label words in specific locales — no such citation exists as of
358
+ spec 0.5.0, and an empty or invented list is not that fix (§1).
359
+
360
+ ---
361
+
362
+ ## 6. Worked examples
363
+
364
+ `⟨J⟩` = U+2060 (word joiner), `⟶` = no change. **Every row in this table requires
365
+ `{ rules: { ranges: true } }` explicitly** — `ranges` is off by default (§1), so every one of
366
+ these inputs is a byte-identical no-op with default options; that default-options behaviour is
367
+ what canonical fixtures (`spec/fixtures/*.json`) assert. These rows describe this rule's own
368
+ conversion behaviour once explicitly enabled — they were relocated here, verbatim, from
369
+ `dashes.md` §6 (rows 3, 3a-3j, 13, 17), which owned range detection through spec 0.4.1.
370
+
371
+ **A row containing an invisible code point must be checked in the escaped mirror, never read** —
372
+ U+2060 is zero-width, so a row that omits it looks correct on screen and is wrong at the byte
373
+ level (dashes.md §6's own note about this applies identically here).
374
+
375
+ ### `en-US` — `range: "en-tight"`
376
+
377
+ | # | Input | Output | Why |
378
+ | --- | --------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
379
+ | 3 | `1914-1918 and pp. 34-36` | `1914⟨J⟩–⟨J⟩1918 and pp. 34⟨J⟩–⟨J⟩36` | range, G4 equal-length branch (4,4) and (2,2), G5 ✓ |
380
+ | 3a | `Takes 5-10 days` | `Takes 5⟨J⟩–⟨J⟩10 days` | G4's `(1,2)` branch; `Rrun` is `10`, no leading zero. G5 is vacuous here |
381
+ | 3b | `aged 9-10 years` | `aged 9⟨J⟩–⟨J⟩10 years` | `(1,2)`; the largest one-digit run against the smallest two-digit run |
382
+ | 3c | `chapters 1-12` | `chapters 1⟨J⟩–⟨J⟩12` | `(1,2)` |
383
+ | 3d | `0-60 in six seconds` | `0⟨J⟩–⟨J⟩60 in six seconds` | `Lrun` may be `0` — the leading-zero clause constrains `Rrun` only. `before` is `NONE`, so G1 and G3 pass |
384
+ | 3e | `won 10-7` | ⟶ | `(2,1)`: G4's directional branch does not admit it. G5 would also reject it; G4 gets there first |
385
+ | 3f | `code 9-05` | ⟶ | `(1,2)` but `Rrun` begins with U+0030 — the pair G5 could not have caught, since G5 is specified for equal-length runs |
386
+ | 3g | `Call 555-1234` | ⟶ | `(3,4)` — neither branch. The phone-number case G4 exists for |
387
+ | 3h | `Call 1-800 now` | ⟶ | `(1,3)` — the `(1,2)` branch requires `length(Rrun) = 2` exactly |
388
+ | 3i | `the 2020-24 season` | ⟶ | `(4,2)` — the abbreviated year range remains a recorded miss (dashes.md §7.3) |
389
+ | 3j | `Figure 5-10` | `Figure 5⟨J⟩–⟨J⟩10` | **DOCUMENTED LIMITATION, not portable conformance evidence — §5.** A compound label, not a range; this rule has no way to tell the two apart. Never asserted by a canonical fixture; see `tests/rules/ranges.test.ts` |
390
+ | 43 | `1914–1918` | ⟶ | the token is already exactly correct for `en-tight` — glyph, length and spacing — so it is left alone. §3.3's invisible-edit test (compute the unbound form; bind only if that would itself be a visible change) is what decides this, not any property of the original glyph |
391
+ | 44 | `1914-1918` | `1914⟨J⟩–⟨J⟩1918` | the hyphen-typed range still converts and still binds. 43 and 44 are the asymmetry dashes.md §7.14 records |
392
+
393
+ #### Closed-up symbols (§3.2a, spec 1.3.0), `en-US`
394
+
395
+ Every row measured against the reference implementation. `⟨J⟩` = U+2060, as above.
396
+
397
+ | # | Input | Output | Why |
398
+ | --- | ------------------- | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
399
+ | 53 | `$15-$20` | `$15⟨J⟩–⟨J⟩$20` | both members repeat U+0024 closed up; the walk consumes the right flank's symbol and G1-G5 then read the digit runs `15` and `20` |
400
+ | 54 | `35%-50%` | `35%⟨J⟩–⟨J⟩50%` | the suffix mirror — the walk moves the **left** flank, and the match is against the outer symbol after `Rrun`. The literal example in the sentence §3.2a quotes |
401
+ | 55 | `15°-20°` | `15°⟨J⟩–⟨J⟩20°` | U+00B0 is in `CLOSED-SYMBOL` on the same ground as U+0024 and U+0025 |
402
+ | 56 | `$15-€20` | ⟶ | the symbols are different code points: a currency conversion, not a range. The walk is not taken, so the token is not a candidate and stays `dashes`' concern |
403
+ | 57 | `15-$20` | ⟶ | an inner symbol with no matching outer one. The match is required in both directions |
404
+ | 58 | `US$15-US$20` | ⟶ | the walk consumes at most one code point (§7.7); G1 declines it independently, since `before` reads past the matched U+0024 and finds `S` |
405
+ | 59 | `$15-20 and 15-20%` | `$15⟨J⟩–⟨J⟩20 and 15⟨J⟩–⟨J⟩20%` | the elided forms, unchanged by §3.2a — the outer symbol is examined only when there is an inner one to match |
406
+ | 60 | `$15 - $20` | `$15⟨J⟩–⟨J⟩$20` | the spaced form of row 53. **Through spec 1.2.0 this was `dashes`' token and gave `$15—$20` with default options**; it is now a range candidate, so with default options it is a no-op — §3.2a's accepted cost |
407
+
408
+ ### `de-DE` — `range: "en-tight"`
409
+
410
+ | # | Input | Output | Why |
411
+ | --- | ---------------- | ---------------------- | ----- |
412
+ | 13 | `Seiten 34-36` | `Seiten 34⟨J⟩–⟨J⟩36` | range |
413
+
414
+ #### Joiner transparency (dashes.md §3.2b, spec cases 45-47)
415
+
416
+ | # | Input | Output | Why |
417
+ | --- | ----------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
418
+ | 45 | `1-1␣-␣1` | `1⟨J⟩–⟨J⟩1␣-␣1` | **defect (e), dashes.md §8.2 (historical).** The second token's `before` is an effective neighbour that skips the joiner and finds the `–`, so G2 rejects it on every pass. Previously (pre-0.4.1 unified rule) pass 2 produced `1⟨J⟩–⟨J⟩1⟨J⟩–⟨J⟩1` |
419
+ | 46 | `1–1␣-␣1` | ⟶ | the **control**: the first token is already exactly correct for `en-tight` (glyph, length, spacing), so §3.3's invisible-edit test leaves it unbound and unchanged; the second token's G2 still rejects it, because `before` is a real `DASH` regardless of which `DASH` glyph it is. `1—1 - 1` (an em dash where the locale wants en) is **not** a fixed point here — it converts to `1⟨J⟩–⟨J⟩1 - 1`, the same as row 45 |
420
+
421
+ ### `ru` — `range: "em-tight"`
422
+
423
+ | # | Input | Output | Why |
424
+ | --- | ------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
425
+ | 17 | `Годы 1941-1945 были тяжёлыми` | `Годы 1941⟨J⟩—⟨J⟩1945 были тяжёлыми` | Russian sets an **em** dash in ranges; `range: "em-tight"`. Aligned to the shipped fixture — which pair of years the row uses is arbitrary |
426
+ | 47 | `1-1␣-␣1` | `1⟨J⟩—⟨J⟩1␣-␣1` | the defect of row 45 reproduced in every locale whose `range` is not `none`; `ru` sets `em-tight` |
427
+
428
+ ---
429
+
430
+ ## 7. Open questions
431
+
432
+ Every item below is current and open against `ranges` specifically. Items that are shared
433
+ machinery rather than range-specific stay in [dashes.md](dashes.md) §7 (its canonical home,
434
+ §3.1 above); items that are purely historical are [dashes.md](dashes.md) §8.
435
+
436
+ 1. **`DIGIT` is ASCII-only** (the class itself is defined in [dashes.md](dashes.md) §3.1, shared
437
+ with `dashes`). Arabic-Indic, Devanagari and fullwidth digits are not recognised, so
438
+ `١٩١٤-١٩١٨` is left alone. The limitation lands here rather than in `dashes.md` §7 because only
439
+ `ranges`' G4/G5 actually read a digit's *value* — `dashes` never does. Extending to Unicode
440
+ `Nd` would require every runtime to agree on a Unicode version _and_ a digit-value table;
441
+ ASCII-only is the portable choice for the v1 locale set. Recorded so the limitation is
442
+ deliberate rather than accidental.
443
+ 2. **Decimal ranges.** `1.5-2.5` and `1,5-2,5` are legitimate ranges and are currently rejected
444
+ by G3. Supporting them means deciding which of U+002E / U+002C is the decimal separator in a
445
+ given locale, which is data the schema does not carry. Left unsupported.
446
+ 3. **Abbreviated year ranges.** `2020-24` is a real convention (Chicago) and is rejected by G4.
447
+ The false-positive risk of relaxing G4 (it would also accept `8859-1`) was judged worse than
448
+ the miss. Needs an operator decision.
449
+ 4. **Scores and votes.** `5-0`, `2-1` pass all guards and become en dashes. That is arguably
450
+ correct typographically (an en dash is standard for scores), but G5 rejects `5-0` while
451
+ accepting `0-5`, which is inconsistent for this use. Either scores are out of scope or G5
452
+ needs an exception; unresolved.
453
+ 5. **The U+2060 range binding is invisible in both the rendered text and a plain diff, and that
454
+ is a reviewing hazard, not just a curiosity.** Every other change this rule makes is catchable
455
+ by eye — a hyphen becomes a dash, a space appears or disappears. §3.3.1's joiner is zero-width:
456
+ a bound form and its unbound equivalent are pixel-identical in every font, and a unified diff
457
+ shows a changed line with no visible difference on it. **A worked-example or fixture row that
458
+ omits the joiner looks correct on screen and is wrong at the byte level** — §6's own standing
459
+ note says so, and the escaped mirror in `spec/fixtures/.escaped/` is the only rendering in
460
+ which the joiner is visible; check it, never read the row on screen. This is not an argument
461
+ against the binding, which is well-sourced and correct, only a property this rule (and any
462
+ future rule emitting a zero-width or invisible code point — U+00AD, U+200B, a variation
463
+ selector) has to manage explicitly rather than assume a reviewer will catch by eye. The
464
+ historical cost of this exact hazard — 32 fixture cases silently broken, and one idempotency
465
+ defect needing a seven-character witness — is recorded at [dashes.md](dashes.md) §8.6, from
466
+ when this rule's behaviour still lived there.
467
+ 6. **The compound-label ambiguity (`Figure 5-10`, G4's `(1,2)` branch) is already the full current
468
+ statement of this rule's central tradeoff — see §5 above, not duplicated here.**
469
+ 7. **A multi-code-point closed-up symbol is not admitted** (§3.2a). `US$15-US$20` and `R$15-R$20`
470
+ carry a two- or three-code-point prefix, and `15°C-20°C` a two-code-point suffix; §3.2a
471
+ consumes at most one code point per side, so all three are declined. Admitting them means
472
+ deciding how far to walk and what stops the walk, and a walk that crosses `LETTER` would
473
+ collide with G1, which exists to keep `MP3-4` and `H2-2` out. The narrow rule is the one the
474
+ citation supports; the wider one needs its own evidence.
475
+ 8. **A spaced unit is deliberately out of scope, and so is the mirror case.** `15 kg-20 kg` and
476
+ `225 nm-2400 nm` are not admitted, because the source §3.2a quotes says a symbol *separated by
477
+ a space* is **not** repeated — the spaced form is a different construction, and NIST SP 811
478
+ §7.7 goes further and recommends the word "to" rather than a range dash for quantities, on the
479
+ ground that a dash can be read as a minus sign. Separately: in `de-DE`, `fr` and `ru` the
480
+ currency sign follows the amount (`30 EUR`, `800 руб.`), so the money case in those locales is
481
+ shaped `15 €-20 €` — a spaced symbol, not a closed-up one, and therefore out of scope by the
482
+ same clause rather than by oversight. Whether a spaced, repeated unit should be admitted at
483
+ all is an open question with evidence pointing away from it.
484
+ **This is also why a `CLOSED-SYMBOL` member appearing in a locale's `nbsp.beforeUnits` does
485
+ not contradict the locale data.** Several locale files list `°C`, and some list `%`, as units
486
+ the locale sets **after a space** — `20 °C`, `20 %`. That is the spaced construction, out of
487
+ scope here; `15°-20°` and `35%-50%`, with the symbol closed up, are the other one. The two
488
+ never meet in the algorithm either: §3.2a's walk requires a `DIGIT` immediately beside the
489
+ symbol, so `15 %-20 %` is not a candidate at all.