polytypo 1.2.0 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +626 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,489 @@
|
|
|
1
|
+
# Rule: `ranges`
|
|
2
|
+
|
|
3
|
+
**Order:** 25. **Default:** off. **Modes:** text, html, markdown, yaml.
|
|
4
|
+
**Spec version:** 0.5.0 (new rule; split out of `dashes`); §3.2a (closed-up symbols) new in 1.3.0.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## 1. Purpose
|
|
9
|
+
|
|
10
|
+
`ranges` recognises the numeric/date **range dash** — `1914–1918` — and renders it in the form
|
|
11
|
+
the locale prescribes. It is the second of the two typographic uses `dashes` (order 30) used to
|
|
12
|
+
own; the two rules together still cover exactly the same input shape `dashes` alone used to
|
|
13
|
+
(spec ≤ 0.4.1), split apart because they now carry **different default-on status**.
|
|
14
|
+
|
|
15
|
+
**Why this is a separate opt-in rule rather than a bounded fix inside `dashes`.** A digit-flanked
|
|
16
|
+
hyphen shaped `5-10` is genuinely ambiguous without knowing the word that precedes it:
|
|
17
|
+
|
|
18
|
+
- `takes 5-10 days`, `aged 9-10 years`, `pages 5-10`, `0-60`, `€30-80` — genuine ranges. Converting
|
|
19
|
+
the hyphen to the locale's dash is correct.
|
|
20
|
+
- `Figure 5-10`, `Table 3-12`, `Section 2-14` — **compound labels**, not ranges (they mean "figure
|
|
21
|
+
10 in chapter 5" — chapter-and-item, not a span). Converting the hyphen here is a false positive
|
|
22
|
+
that reads oddly to a careful proofreader, even though it is rarely noticed casually.
|
|
23
|
+
- `9-11`, `7-11` — bare proper nouns (a date, a convenience-store chain) that happen to have the
|
|
24
|
+
identical digit-hyphen-digit shape and are neither a label nor a range.
|
|
25
|
+
|
|
26
|
+
Separating the first case from the second and third needs the word immediately before the
|
|
27
|
+
hyphen's left digit run — `Figure`, `Table`, `Section`, `Fig.`, `Abb.`, `рис.` — which is an
|
|
28
|
+
open-ended, per-locale, per-house-style word list with no natural closure
|
|
29
|
+
([dashes.md](dashes.md) §7.11, carried forward unchanged from when this algorithm lived there).
|
|
30
|
+
`locale.schema.json` could hold such a list as literal strings, but populating it would mean
|
|
31
|
+
either (a) inventing entries with no normative citation — forbidden by this project's evidence
|
|
32
|
+
discipline (PLAN.md §6.1) — or (b) leaving it empty, which is not a fix, only a schema field that
|
|
33
|
+
looks like one. Neither the bare digit-hyphen-digit shape nor a label word list gives a
|
|
34
|
+
**structurally closed, portable** rule that never produces a false positive: this project's rules
|
|
35
|
+
scan code points, not words with real-world referents, and are built never to consult that kind
|
|
36
|
+
of context (ARCHITECTURE.md §4.1).
|
|
37
|
+
|
|
38
|
+
**The resolution is: make range conversion something the caller explicitly asks for, and leave it
|
|
39
|
+
off by default.** With `ranges` off (the default), `5-10`, `Figure 5-10`, `9-11` and `7-11` are
|
|
40
|
+
all byte-identical no-ops — the compound-label false positive cannot occur, because nothing in
|
|
41
|
+
this class converts at all. A caller who explicitly opts in (`{ rules: { ranges: true } }`) is
|
|
42
|
+
choosing to accept the residual, structurally irreducible ambiguity between a genuine range and a
|
|
43
|
+
same-shaped label or proper noun, in exchange for genuine ranges being typeset correctly. This is
|
|
44
|
+
a documented tradeoff, not a claim of reliable label detection — see §5.
|
|
45
|
+
|
|
46
|
+
---
|
|
47
|
+
|
|
48
|
+
## 2. Locale data consumed
|
|
49
|
+
|
|
50
|
+
- `dash.range` — one of `"em-tight"`, `"em-spaced"`, `"en-tight"`, `"en-spaced"`, `"none"`. The
|
|
51
|
+
**same field**, under the **same locale JSON key**, `dashes` has always read for range styling
|
|
52
|
+
— this rule's split from `dashes` moves which rule reads it, not what the field means or which
|
|
53
|
+
citations support it. No new locale claim is made by this document (operator decision, spec
|
|
54
|
+
0.5.0): every `dash.range` value in `spec/locales/*.json` is unchanged from spec 0.4.1.
|
|
55
|
+
|
|
56
|
+
`"none"` means the locale has no verified range convention, and this rule must emit nothing at
|
|
57
|
+
all for a range token in that locale — not a fallback to `dash.parenthetical`, nothing.
|
|
58
|
+
|
|
59
|
+
---
|
|
60
|
+
|
|
61
|
+
## 3. Algorithm
|
|
62
|
+
|
|
63
|
+
Input is a code-point array `cp[0 … n-1]`. Character classes (`DASH`, `DIGIT`, `LETTER`, `SPACE`,
|
|
64
|
+
`BREAK`, `NOBREAK-SPACE`, `INERT-DASH`, `JOINER`) are exactly [dashes.md](dashes.md) §3.1's —
|
|
65
|
+
this rule and `dashes` read the same input alphabet, because both scan the same DASH-token shape
|
|
66
|
+
before diverging on what each does with it.
|
|
67
|
+
|
|
68
|
+
### 3.1 The dash token (shared with `dashes`)
|
|
69
|
+
|
|
70
|
+
Steps 1-7 of [dashes.md](dashes.md) §3.2 — run detection, the length-3 decoration cutoff, outer
|
|
71
|
+
spacing, the symmetry guard, the content-on-both-sides guard, the joiner-crossing walk (§3.2a),
|
|
72
|
+
the isolation guard and the cluster guard (§3.2 step 7) — apply here **identically, unchanged**.
|
|
73
|
+
A token that `dashes`' §3.2 steps would decline is declined here too, for the same reasons; the
|
|
74
|
+
two rules diverge only starting at §3.3 below. The JS reference implementation shares one
|
|
75
|
+
internal module (`src/rules/dash-shared.ts`) between `src/rules/dashes.ts` and
|
|
76
|
+
`src/rules/ranges.ts` for exactly this reason — a port is free to structure its own module
|
|
77
|
+
boundary differently, but must reproduce the same shared guard behaviour in both rules, not two
|
|
78
|
+
independent approximations of it.
|
|
79
|
+
|
|
80
|
+
`JOINER`-crossing re-entry (dashes.md §3.2a's third bullet: "if a joiner was crossed and
|
|
81
|
+
`cp[L*]`/`cp[R*]` are both `DIGIT`, the token is a bound range this rule produced on an earlier
|
|
82
|
+
pass") is **this rule's own** re-entry condition — it is how a tight range this rule already
|
|
83
|
+
converted survives a second pipeline pass without regrowing a second joiner. `dashes` shares the
|
|
84
|
+
same joiner-crossing walk (it must, to correctly decline a token immediately touching a range
|
|
85
|
+
this rule already produced) but never itself owns that re-entry: any joiner adjacent to a token
|
|
86
|
+
that is **not a range candidate** (§3.2, §3.2a — wider than "digit-flanked" as of spec 1.3.0) is
|
|
87
|
+
declined outright (dashes.md §3.2a's fourth bullet).
|
|
88
|
+
|
|
89
|
+
### 3.2 Range guards (G1-G5)
|
|
90
|
+
|
|
91
|
+
The token is a **range candidate** iff `cp[L']` is in `DIGIT` **and** `cp[R']` is in `DIGIT` —
|
|
92
|
+
`L` and `R` after the joiner-crossing walk of §3.1, and `L'`/`R'` after the closed-up-symbol walk
|
|
93
|
+
of §3.2a, which moves neither index unless a symbol is there to consume. If it is a range
|
|
94
|
+
candidate, this rule owns it exclusively: `dashes` never processes it, whether or not `ranges` is
|
|
95
|
+
enabled (operator decision, spec 0.5.0 — see [dashes.md](dashes.md) §1). If it is **not** a range
|
|
96
|
+
candidate (either flank is not `DIGIT` after that walk), this rule emits nothing for it at all;
|
|
97
|
+
it is `dashes`' concern or no rule's.
|
|
98
|
+
|
|
99
|
+
Compute `Lrun`, `Rrun`, `before`, `after` exactly as [dashes.md](dashes.md) §3.3 (pre-0.5.0)
|
|
100
|
+
specified — reproduced here verbatim since this is now its own rule document, with `L'`/`R'`
|
|
101
|
+
where it said `L`/`R`:
|
|
102
|
+
|
|
103
|
+
- `Lrun` = the maximal run of `DIGIT` ending at `L'` (indices `a … L'`);
|
|
104
|
+
- `Rrun` = the maximal run of `DIGIT` starting at `R'` (indices `R' … b`);
|
|
105
|
+
- `before` = the **effective neighbour** (dashes.md §3.2b) to the left of index `a`, except as
|
|
106
|
+
§3.2a amends it;
|
|
107
|
+
- `after` = the **effective neighbour** to the right of index `b`, except as §3.2a amends it.
|
|
108
|
+
|
|
109
|
+
All five guards must pass:
|
|
110
|
+
|
|
111
|
+
- **G1 — no letter adjacency.** `before` is not in `LETTER`. (`MP3-4`, `H2-2` rejected. `COVID-19`
|
|
112
|
+
is already rejected earlier, because `cp[L]` is `D`, a letter, not a digit.)
|
|
113
|
+
- **G2 — no chain.** `before` is not in `DASH` and not in `INERT-DASH`; `after` is not in `DASH`
|
|
114
|
+
and not in `INERT-DASH`. This protects `2026-08-15`, `978-3-16-148410-0` and `212-555-1234`.
|
|
115
|
+
**This guard reads the input as it stood before this rule (or `dashes`) made any edit in this
|
|
116
|
+
pipeline pass** — see §4's note on why `ranges` runs before `dashes`: evaluating G2 against an
|
|
117
|
+
adjacency `dashes` just created (by removing a space next to this token) rather than against
|
|
118
|
+
what the author actually typed is exactly the defect §4 documents and the reason for the chosen
|
|
119
|
+
order.
|
|
120
|
+
- **G3 — not part of a decimal or path.** `before` is not U+002E, U+002C or U+002F; `after` is
|
|
121
|
+
not U+002F. `1.5-2.5` and `01/02-03/04` are left alone.
|
|
122
|
+
- **G4 — run lengths.** Either `length(Lrun) = length(Rrun)`, **or** `length(Lrun) = 1` and
|
|
123
|
+
`length(Rrun) = 2` and `Rrun` does not begin with U+0030. `1914-1918` (4,4) passes; `5-10` (1,2)
|
|
124
|
+
passes; `0-60` (1,2) passes; `ISO 8859-1` (4,1) fails; `555-1234` (3,4) fails; `1-800` (1,3)
|
|
125
|
+
fails; `2020-24` (4,2) fails; `10-7` (2,1) fails; `9-05` fails on the leading zero. See
|
|
126
|
+
[dashes.md](dashes.md) §3.3's historical discussion of why this exact shape and not a wider one
|
|
127
|
+
— unchanged, reproduced there rather than duplicated here since it is a proof about G5's
|
|
128
|
+
correctness, not new content this rule's split needs to restate.
|
|
129
|
+
- **G5 — non-decreasing.** The integer value of `Lrun` ≤ the integer value of `Rrun`, compared
|
|
130
|
+
digit-by-digit left to right (sound only because G4 guarantees equal length in the only branch
|
|
131
|
+
where G5 does work). `1914-1918` passes; `1234-5678` passes; `20-10` fails.
|
|
132
|
+
|
|
133
|
+
If any guard fails, emit nothing. **The token is not reconsidered by `dashes`.** A stroke between
|
|
134
|
+
two digits is a range or it is nothing; `dashes` never sees it (§3.1 above).
|
|
135
|
+
|
|
136
|
+
### 3.2a Closed-up symbols on both members (spec 1.3.0)
|
|
137
|
+
|
|
138
|
+
Before spec 1.3.0 a range whose members each carried a symbol — `$15-$20`, `35%-50%` — was not a
|
|
139
|
+
range candidate, because the flank next to the dash was the symbol and not a `DIGIT`. `$15-20`
|
|
140
|
+
converted and `$15-$20` did not; `15-20%` converted and `15%-20%` did not. The asymmetry was a
|
|
141
|
+
consequence of where the symbol sits relative to the digit run, never a decision anyone made.
|
|
142
|
+
|
|
143
|
+
**The distinction the source draws is closed-up versus spaced, not currency versus unit.**
|
|
144
|
+
|
|
145
|
+
> "the abbreviation or symbol is repeated if it is closed up to the number but not if it is
|
|
146
|
+
> separated: 35%–50%" — *The Chicago Manual of Style*, 18th ed., 9.19, quoted in the freely
|
|
147
|
+
> readable [CMOS Online Q&A, "Numbers"](https://www.chicagomanualofstyle.org/qanda/data/faq/topics/Numbers/faq0024.html);
|
|
148
|
+
> the same paragraph is applied to money on
|
|
149
|
+
> [page 5 of that topic](https://www.chicagomanualofstyle.org/qanda/data/faq/topics/Numbers.html?page=5):
|
|
150
|
+
> "in Chicago style an abbreviation or symbol is repeated if it is closed up to a number but not
|
|
151
|
+
> if it is separated by a space: $3–$5 million".
|
|
152
|
+
|
|
153
|
+
So `$15–$20` and `35%–50%` are one rule, and `15 kg–20 kg` is not that rule — `kg` is separated
|
|
154
|
+
by a space, and the source says a separated abbreviation is **not** repeated. The elided forms
|
|
155
|
+
`$3–5 million` and `15–20%` are permitted variants the same answer calls acceptable; they already
|
|
156
|
+
convert and keep converting.
|
|
157
|
+
|
|
158
|
+
**`CLOSED-SYMBOL`** is a literal code-point set, fixed here and not locale data:
|
|
159
|
+
|
|
160
|
+
> U+0024, U+00A2, U+00A3, U+00A4, U+00A5, every code point from U+20A0 through U+20CF inclusive
|
|
161
|
+
> (the whole Currency Symbols block, assigned or not), U+0025, U+2030, U+2031, and U+00B0.
|
|
162
|
+
|
|
163
|
+
`$` and `%` are the source's own examples; the rest are the same class by the source's own
|
|
164
|
+
wording — symbols conventionally written closed up to a number: the other currency signs,
|
|
165
|
+
per-mille and per-ten-thousand, and the degree sign. It is written as literal code points rather
|
|
166
|
+
than a Unicode category test for the same reason `DIGIT` is ASCII-only (§7.1): a category test
|
|
167
|
+
makes the rule's verdict depend on which Unicode version a runtime was built against, and the
|
|
168
|
+
five runtimes must agree. The currency range is the **block's own bounds**, U+20A0–U+20CF, not
|
|
169
|
+
the subset assigned in some Unicode version — block bounds never move, while the assigned subset
|
|
170
|
+
does, and a rule that admitted only today's assignments would drift between runtimes built
|
|
171
|
+
against different UCD releases. Unassigned code points inside the block are members of the set
|
|
172
|
+
and unreachable in valid text; a code point assigned there later is a currency sign by the
|
|
173
|
+
block's own definition, and membership is then already correct without a spec change. `+`, `-`, `#`, `(`, `)` and the quotation marks are **not** members and
|
|
174
|
+
must not be added on the reasoning that they are also written closed up: none is a symbol this
|
|
175
|
+
source's rule is about, and each would admit a shape (`+15-+20`, `#15-#20`, `"15"-"20"`) nobody
|
|
176
|
+
has asked for and no citation supports.
|
|
177
|
+
|
|
178
|
+
**The walk.** Both sides are decided **from the original `L` and `R`, simultaneously**, before
|
|
179
|
+
either index moves — never left-then-right or right-then-left. Otherwise a token whose flanks are
|
|
180
|
+
both in `CLOSED-SYMBOL` (`%15%-%20%`) would have a verdict that depends on evaluation order, and
|
|
181
|
+
two runtimes could disagree while both following this document. After the joiner-crossing walk of
|
|
182
|
+
§3.1 has produced `L` and `R`:
|
|
183
|
+
|
|
184
|
+
- **right:** if `cp[R]` is in `CLOSED-SYMBOL` **and** `cp[R + 1]` exists and is in `DIGIT`, then
|
|
185
|
+
`R' = R + 1` and the **inner right symbol** is `cp[R]`. Otherwise `R' = R` and there is none.
|
|
186
|
+
- **left:** if `cp[L]` is in `CLOSED-SYMBOL` **and** `cp[L - 1]` exists and is in `DIGIT`, then
|
|
187
|
+
`L' = L - 1` and the **inner left symbol** is `cp[L]`. Otherwise `L' = L` and there is none.
|
|
188
|
+
|
|
189
|
+
Each side consumes **at most one** code point, so `US$15-US$20` and `15°C-20°C` are not
|
|
190
|
+
admitted — and the mechanism is candidacy, not a guard: `cp[R]` is `U` and `cp[L]` is `C`,
|
|
191
|
+
neither is in `CLOSED-SYMBOL`, so no walk is taken and the flank is simply not a `DIGIT`. The
|
|
192
|
+
half-written `US$15-$20` **is** a candidate (the right walk matches the outer `$` in front of
|
|
193
|
+
`15`) and is declined by G1 instead, because `before` reads past that `$` and finds `S`. Both are
|
|
194
|
+
recorded in §7.7.
|
|
195
|
+
|
|
196
|
+
**The outer symbols** are the effective neighbour (dashes.md §3.2b) immediately left of `Lrun`'s
|
|
197
|
+
first index `a`, and immediately right of `Rrun`'s last index `b`, when that code point is in
|
|
198
|
+
`CLOSED-SYMBOL`.
|
|
199
|
+
|
|
200
|
+
**Matching, and it is exact.** A side's walk is taken **only if** the symbol it would consume is
|
|
201
|
+
matched on the opposite member: an inner right symbol must be the **same code point** as the
|
|
202
|
+
outer left symbol, and an inner left symbol the same code point as the outer right symbol. Not a
|
|
203
|
+
currency-equivalence table, not a per-locale list: the same code point.
|
|
204
|
+
|
|
205
|
+
**An unmatched symbol means the walk is not taken at all**, so the flank stays a non-`DIGIT`, the
|
|
206
|
+
token is **not** a range candidate, and it remains `dashes`' concern exactly as it was before
|
|
207
|
+
spec 1.3.0. `$15-€20` is a currency conversion, not a range; `15-$20` and `15%-20` are
|
|
208
|
+
half-written. None of the three changes hands, and none of them changes behaviour in this spec
|
|
209
|
+
version.
|
|
210
|
+
|
|
211
|
+
If a side has **no inner symbol**, the corresponding outer symbol is not examined at all, which
|
|
212
|
+
is exactly what keeps `$15-20` and `15-20%` converting as they did before this section existed.
|
|
213
|
+
|
|
214
|
+
**What this does not change.** `Rrun` is the maximal `DIGIT` run starting at `R'` and `Lrun` the
|
|
215
|
+
one ending at `L'`; wherever the shared guards of §3.1 and the replacement of §3.3 say `cp[L]` or
|
|
216
|
+
`cp[R]`, this rule reads `cp[L']` or `cp[R']`. **G4 and G5 are untouched** — they compare digit
|
|
217
|
+
runs and never see a symbol. The edit span is untouched too: the replacement covers the dash run
|
|
218
|
+
(and, when binding, the joiners adjacent to it), so a symbol is never inserted, removed or
|
|
219
|
+
rewritten. This rule changes one dash and nothing else.
|
|
220
|
+
|
|
221
|
+
**G1, G2 and G3 judge the same position they always did.** When an inner right symbol was
|
|
222
|
+
matched, `before` is the effective neighbour left of the **outer left symbol** rather than left of
|
|
223
|
+
`a`; when an inner left symbol was matched, `after` is the effective neighbour right of the
|
|
224
|
+
**outer right symbol** rather than right of `b`. In every other case `before` and `after` are
|
|
225
|
+
unchanged. So `US$15-$20` has `before` = `S`, a `LETTER`, and G1 declines it.
|
|
226
|
+
|
|
227
|
+
**Idempotency, and the two shared guards this section forced open.** A bound `$15–$20`
|
|
228
|
+
re-enters through §3.1's joiner-crossing walk, whose re-entry condition
|
|
229
|
+
([dashes.md](dashes.md) §3.2a) is amended in this same spec version to read `cp[L']`/`cp[R']`;
|
|
230
|
+
having re-entered, §3.3's identical-replacement test makes the second pass a no-op.
|
|
231
|
+
|
|
232
|
+
That is the easy half. The hard half is that widening what counts as a range member widens what
|
|
233
|
+
a **`dashes`** edit elsewhere in the text can disturb, and `dashes`' spacing-transition guard was
|
|
234
|
+
keyed to `DIGIT`.
|
|
235
|
+
|
|
236
|
+
**T1** ([dashes.md](dashes.md) §3.2 step 8) both gated its branches on a `DIGIT` flank and walked
|
|
237
|
+
outward from the digit run without stepping over a symbol, so a tight `dashes` token next to a
|
|
238
|
+
closed-up range member became spaced, that U+0020 replaced the range's `before` or `after`, and
|
|
239
|
+
the range converted on the **next** pass. Four witnesses, one per position and side, each drifting
|
|
240
|
+
where its all-digit analogue was already inert: `a—$15-$20`, `35%-50%—b`, `a--15% - 20%` and
|
|
241
|
+
`$1 - $1--a`. T1's reach is `CLOSED-SYMBOL`-transparent at two positions per side as of spec
|
|
242
|
+
1.3.0, and each witness is pinned by a fixture asserting that the input is now a **no-op** —
|
|
243
|
+
which is what a conformance runner can see, since it cannot run a second pass and compare.
|
|
244
|
+
|
|
245
|
+
**This is deliberately T1 and not the cluster guard** ([dashes.md](dashes.md) §3.2 step 7), which
|
|
246
|
+
would have closed the same defect by putting `CLOSED-SYMBOL` in its alphabet. That guard is
|
|
247
|
+
unconditional, so it would also have made `price--$50--drop` inert in an `em-tight` locale that
|
|
248
|
+
has no defect to fix; T1 fires only on the tight-to-spaced transition that can actually disturb a
|
|
249
|
+
neighbour. The alternative was implemented, measured and rejected on that comparison.
|
|
250
|
+
|
|
251
|
+
**The cost this leaves.** In a locale whose `dash.parenthetical` is **spaced**, a tight token
|
|
252
|
+
next to a closed-up range member no longer converts: `Anstieg--50%--war` is left alone in
|
|
253
|
+
`de-DE`, where spec 1.2.0 gave `Anstieg – 50% – war` (measured). The all-digit
|
|
254
|
+
`Anstieg--50--war` has always been left alone, so the two agree; and a locale with a tight
|
|
255
|
+
parenthetical form never reaches the guard, so `price--$50--drop` still becomes `price—$50—drop`
|
|
256
|
+
in `en-US`. **The remedy is unconditional on `dash.range`, while the defect was not**: the
|
|
257
|
+
conversion loss therefore also lands in `fr`, `fr-CA`, `it`, `nl`, `pt-BR` and `pt-PT`, whose
|
|
258
|
+
`dash.range` is `"none"` and which never had a range verdict to disturb. Making T1 consult
|
|
259
|
+
`dash.range` would make a `dashes` verdict depend on a field `dashes` does not read, which is a
|
|
260
|
+
worse trade than a conversion nobody has asked for in six locales.
|
|
261
|
+
|
|
262
|
+
**The accepted cost, stated plainly.** Widening candidacy moves tokens **out of** `dashes`, and
|
|
263
|
+
`ranges` is off by default, so text that `dashes` used to change now changes only when a caller
|
|
264
|
+
turns `ranges` on — and in the eight locales whose `dash.range` is `"none"`, not even then:
|
|
265
|
+
`$15 - $20` gave `$15 — $20` in `fr` through spec 1.2.0 and is a permanent no-op from 1.3.0. The affected shapes are the **spaced and multi-hyphen** ones, because a lone
|
|
266
|
+
tight hyphen between two non-spaces was never converted by `dashes` either: `$15 - $20` gave
|
|
267
|
+
`$15—$20` in `en-US` before this section and is a no-op with default options after it, while
|
|
268
|
+
`$15-$20` was already a no-op and now converts when `ranges` is on. That is the same behaviour `$15 - 20` has had since spec 0.5.0 — the
|
|
269
|
+
change makes the two consistent rather than introducing an inconsistency — and the alternative is
|
|
270
|
+
worse than the cost: declining `$15-$20` does not leave `$15–20`, it leaves `$15-$20`, a hyphen
|
|
271
|
+
where every reading of every source above wants a dash.
|
|
272
|
+
|
|
273
|
+
### 3.3 Replacement
|
|
274
|
+
|
|
275
|
+
If all guards pass, emit one edit replacing `cp[s - lsp … e - 1 + rsp]` with:
|
|
276
|
+
|
|
277
|
+
| `dash.range` | replacement |
|
|
278
|
+
| ------------- | -------------------------------------------------------- |
|
|
279
|
+
| `"em-tight"` | U+2014 |
|
|
280
|
+
| `"em-spaced"` | U+0020 U+2014 U+0020 |
|
|
281
|
+
| `"en-tight"` | U+2013 |
|
|
282
|
+
| `"en-spaced"` | U+0020 U+2013 U+0020 |
|
|
283
|
+
| `"none"` | _(emit nothing — the token is left exactly as written)_ |
|
|
284
|
+
|
|
285
|
+
Subject to guards T1 and T2 ([dashes.md](dashes.md) §3.2 steps 8-9, `isSpacingTransitionBlocked`
|
|
286
|
+
and the strip-before/open-bracket checks — shared, unchanged, evaluated here against
|
|
287
|
+
`dash.range`'s chosen form exactly as they were against either branch's form pre-0.5.0). If the
|
|
288
|
+
replacement is identical, code point for code point, to the span it would replace, emit nothing
|
|
289
|
+
instead.
|
|
290
|
+
|
|
291
|
+
#### 3.3.1 Binding a tight range
|
|
292
|
+
|
|
293
|
+
Unchanged from [dashes.md](dashes.md) §3.3.1 (pre-0.5.0), reproduced here since binding is now
|
|
294
|
+
exclusively this rule's behaviour: when the chosen `dash.range` form is `-tight` and all guards
|
|
295
|
+
have passed, the emitted replacement is `JOINER dash JOINER` (U+2060, the dash, U+2060), and the
|
|
296
|
+
edit span covers any `JOINER` already adjacent to the token — the range is a single lexical unit
|
|
297
|
+
(UAX #14 gives U+2013/U+2014 the line-break class `BA`, and Мильчин and Chicago both forbid
|
|
298
|
+
breaking a range in running text). A range already in its exact target form is never bound (the
|
|
299
|
+
invisible-edit test: compute the unbound replacement first; if it is already a no-op, stay
|
|
300
|
+
unbound). `dashes` never emits a joiner — an interrupting parenthetical dash is exactly where a
|
|
301
|
+
line **may** break.
|
|
302
|
+
|
|
303
|
+
### 3.4 Continue, non-overlap, no-break spaces
|
|
304
|
+
|
|
305
|
+
[dashes.md](dashes.md) §3.5 (continue/non-overlap) and §3.6 (no-break spaces are never touched)
|
|
306
|
+
apply to this rule exactly as written there — both are properties of the shared token shape
|
|
307
|
+
(§3.1), not of which branch/rule a token ends up in.
|
|
308
|
+
|
|
309
|
+
---
|
|
310
|
+
|
|
311
|
+
## 4. Rule order: why `ranges` runs before `dashes`
|
|
312
|
+
|
|
313
|
+
`ranges` is order 25, `dashes` is order 30 — `ranges` runs first. This was chosen from an
|
|
314
|
+
observed behavioural difference, not for convenience.
|
|
315
|
+
|
|
316
|
+
The pre-0.5.0 unified `dashes` rule found every dash token's edits in **one scan over the
|
|
317
|
+
unedited input**, then applied all of them together — so every token's guards, including G2's
|
|
318
|
+
cross-token "no chain" check, always read the original, unedited adjacency structure, regardless
|
|
319
|
+
of which other tokens in the same document were also converting. Splitting `ranges` and `dashes`
|
|
320
|
+
into two sequential rule passes means one of them necessarily sees the other's *output*, not the
|
|
321
|
+
original input, for whichever positions the first rule touched.
|
|
322
|
+
|
|
323
|
+
**Concrete case: `a - 5-10`** (a parenthetical dash directly followed, by one space, by a
|
|
324
|
+
genuine range). `en-US`: `dash.parenthetical = "em-tight"`, `dash.range = "en-tight"`.
|
|
325
|
+
|
|
326
|
+
- **`ranges` first:** `ranges` scans the original text; the range token's `before` (G2) is the
|
|
327
|
+
space at index 3, not a dash, so G2 passes and `5-10` becomes `5–10`. `dashes` then scans
|
|
328
|
+
`a - 5–10`; its own token (`a - 5`) is unaffected by the far-away range edit and converts to
|
|
329
|
+
`a—5–10` normally. **Final: `a—5–10`.**
|
|
330
|
+
- **`dashes` first:** `dashes` converts `a - 5` to `a—5` (its `em-tight` form has no spaces, so
|
|
331
|
+
the edit span swallows the space that used to separate the parenthetical dash from the digit
|
|
332
|
+
run). `ranges` then scans `a—5-10`; the range token's `before` (G2) is now the em-dash
|
|
333
|
+
`dashes` just emitted, immediately adjacent (no space) — G2 reads this as a **chain** (the same
|
|
334
|
+
shape as `978-3-16-148410-0`) and declines. **Final: `a—5-10`** — the range is silently never
|
|
335
|
+
converted, purely because of which rule happened to run first, not because of anything either
|
|
336
|
+
rule's own guards were designed to protect.
|
|
337
|
+
|
|
338
|
+
`ranges`-before-`dashes` reproduces the pre-0.5.0 unified rule's output (`a—5–10`) exactly;
|
|
339
|
+
`dashes`-before-`ranges` introduces a new, order-induced decline that never happened before the
|
|
340
|
+
split. `tests/rules/ranges.test.ts` and `tests/rules/dashes.test.ts` pin this case as a fixed
|
|
341
|
+
regression witness. This is a single reproducible counter-example, not an exhaustive proof that
|
|
342
|
+
no input ever favours the opposite order — but it is evidence in one concrete direction, and
|
|
343
|
+
`ranges`-before-`dashes` is kept on that evidence, not on a preference between two otherwise
|
|
344
|
+
indistinguishable choices.
|
|
345
|
+
|
|
346
|
+
---
|
|
347
|
+
|
|
348
|
+
## 5. What this rule does not, and cannot, solve
|
|
349
|
+
|
|
350
|
+
**`ranges` converts `Figure 5-10` to `Figure 5–10` when explicitly enabled**, exactly as it
|
|
351
|
+
converts a genuine range. This is not an oversight this document is unaware of — it is the
|
|
352
|
+
residual ambiguity §1 names as the reason the rule is opt-in. Enabling `ranges` is an explicit
|
|
353
|
+
choice to accept that a compound label sharing the digit-hyphen-digit shape will be converted
|
|
354
|
+
alongside genuine ranges; the caller has no bounded, cited mechanism this project can offer to
|
|
355
|
+
tell the two apart (§1). A future spec version could add a cited, closed `dash.labelWords` list
|
|
356
|
+
per locale (dashes.md §7.11's own suggestion) if and when locale-authority research produces
|
|
357
|
+
citable evidence for specific label words in specific locales — no such citation exists as of
|
|
358
|
+
spec 0.5.0, and an empty or invented list is not that fix (§1).
|
|
359
|
+
|
|
360
|
+
---
|
|
361
|
+
|
|
362
|
+
## 6. Worked examples
|
|
363
|
+
|
|
364
|
+
`⟨J⟩` = U+2060 (word joiner), `⟶` = no change. **Every row in this table requires
|
|
365
|
+
`{ rules: { ranges: true } }` explicitly** — `ranges` is off by default (§1), so every one of
|
|
366
|
+
these inputs is a byte-identical no-op with default options; that default-options behaviour is
|
|
367
|
+
what canonical fixtures (`spec/fixtures/*.json`) assert. These rows describe this rule's own
|
|
368
|
+
conversion behaviour once explicitly enabled — they were relocated here, verbatim, from
|
|
369
|
+
`dashes.md` §6 (rows 3, 3a-3j, 13, 17), which owned range detection through spec 0.4.1.
|
|
370
|
+
|
|
371
|
+
**A row containing an invisible code point must be checked in the escaped mirror, never read** —
|
|
372
|
+
U+2060 is zero-width, so a row that omits it looks correct on screen and is wrong at the byte
|
|
373
|
+
level (dashes.md §6's own note about this applies identically here).
|
|
374
|
+
|
|
375
|
+
### `en-US` — `range: "en-tight"`
|
|
376
|
+
|
|
377
|
+
| # | Input | Output | Why |
|
|
378
|
+
| --- | --------------------------- | -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
379
|
+
| 3 | `1914-1918 and pp. 34-36` | `1914⟨J⟩–⟨J⟩1918 and pp. 34⟨J⟩–⟨J⟩36` | range, G4 equal-length branch (4,4) and (2,2), G5 ✓ |
|
|
380
|
+
| 3a | `Takes 5-10 days` | `Takes 5⟨J⟩–⟨J⟩10 days` | G4's `(1,2)` branch; `Rrun` is `10`, no leading zero. G5 is vacuous here |
|
|
381
|
+
| 3b | `aged 9-10 years` | `aged 9⟨J⟩–⟨J⟩10 years` | `(1,2)`; the largest one-digit run against the smallest two-digit run |
|
|
382
|
+
| 3c | `chapters 1-12` | `chapters 1⟨J⟩–⟨J⟩12` | `(1,2)` |
|
|
383
|
+
| 3d | `0-60 in six seconds` | `0⟨J⟩–⟨J⟩60 in six seconds` | `Lrun` may be `0` — the leading-zero clause constrains `Rrun` only. `before` is `NONE`, so G1 and G3 pass |
|
|
384
|
+
| 3e | `won 10-7` | ⟶ | `(2,1)`: G4's directional branch does not admit it. G5 would also reject it; G4 gets there first |
|
|
385
|
+
| 3f | `code 9-05` | ⟶ | `(1,2)` but `Rrun` begins with U+0030 — the pair G5 could not have caught, since G5 is specified for equal-length runs |
|
|
386
|
+
| 3g | `Call 555-1234` | ⟶ | `(3,4)` — neither branch. The phone-number case G4 exists for |
|
|
387
|
+
| 3h | `Call 1-800 now` | ⟶ | `(1,3)` — the `(1,2)` branch requires `length(Rrun) = 2` exactly |
|
|
388
|
+
| 3i | `the 2020-24 season` | ⟶ | `(4,2)` — the abbreviated year range remains a recorded miss (dashes.md §7.3) |
|
|
389
|
+
| 3j | `Figure 5-10` | `Figure 5⟨J⟩–⟨J⟩10` | **DOCUMENTED LIMITATION, not portable conformance evidence — §5.** A compound label, not a range; this rule has no way to tell the two apart. Never asserted by a canonical fixture; see `tests/rules/ranges.test.ts` |
|
|
390
|
+
| 43 | `1914–1918` | ⟶ | the token is already exactly correct for `en-tight` — glyph, length and spacing — so it is left alone. §3.3's invisible-edit test (compute the unbound form; bind only if that would itself be a visible change) is what decides this, not any property of the original glyph |
|
|
391
|
+
| 44 | `1914-1918` | `1914⟨J⟩–⟨J⟩1918` | the hyphen-typed range still converts and still binds. 43 and 44 are the asymmetry dashes.md §7.14 records |
|
|
392
|
+
|
|
393
|
+
#### Closed-up symbols (§3.2a, spec 1.3.0), `en-US`
|
|
394
|
+
|
|
395
|
+
Every row measured against the reference implementation. `⟨J⟩` = U+2060, as above.
|
|
396
|
+
|
|
397
|
+
| # | Input | Output | Why |
|
|
398
|
+
| --- | ------------------- | -------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
399
|
+
| 53 | `$15-$20` | `$15⟨J⟩–⟨J⟩$20` | both members repeat U+0024 closed up; the walk consumes the right flank's symbol and G1-G5 then read the digit runs `15` and `20` |
|
|
400
|
+
| 54 | `35%-50%` | `35%⟨J⟩–⟨J⟩50%` | the suffix mirror — the walk moves the **left** flank, and the match is against the outer symbol after `Rrun`. The literal example in the sentence §3.2a quotes |
|
|
401
|
+
| 55 | `15°-20°` | `15°⟨J⟩–⟨J⟩20°` | U+00B0 is in `CLOSED-SYMBOL` on the same ground as U+0024 and U+0025 |
|
|
402
|
+
| 56 | `$15-€20` | ⟶ | the symbols are different code points: a currency conversion, not a range. The walk is not taken, so the token is not a candidate and stays `dashes`' concern |
|
|
403
|
+
| 57 | `15-$20` | ⟶ | an inner symbol with no matching outer one. The match is required in both directions |
|
|
404
|
+
| 58 | `US$15-US$20` | ⟶ | the walk consumes at most one code point (§7.7); G1 declines it independently, since `before` reads past the matched U+0024 and finds `S` |
|
|
405
|
+
| 59 | `$15-20 and 15-20%` | `$15⟨J⟩–⟨J⟩20 and 15⟨J⟩–⟨J⟩20%` | the elided forms, unchanged by §3.2a — the outer symbol is examined only when there is an inner one to match |
|
|
406
|
+
| 60 | `$15 - $20` | `$15⟨J⟩–⟨J⟩$20` | the spaced form of row 53. **Through spec 1.2.0 this was `dashes`' token and gave `$15—$20` with default options**; it is now a range candidate, so with default options it is a no-op — §3.2a's accepted cost |
|
|
407
|
+
|
|
408
|
+
### `de-DE` — `range: "en-tight"`
|
|
409
|
+
|
|
410
|
+
| # | Input | Output | Why |
|
|
411
|
+
| --- | ---------------- | ---------------------- | ----- |
|
|
412
|
+
| 13 | `Seiten 34-36` | `Seiten 34⟨J⟩–⟨J⟩36` | range |
|
|
413
|
+
|
|
414
|
+
#### Joiner transparency (dashes.md §3.2b, spec cases 45-47)
|
|
415
|
+
|
|
416
|
+
| # | Input | Output | Why |
|
|
417
|
+
| --- | ----------- | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
|
418
|
+
| 45 | `1-1␣-␣1` | `1⟨J⟩–⟨J⟩1␣-␣1` | **defect (e), dashes.md §8.2 (historical).** The second token's `before` is an effective neighbour that skips the joiner and finds the `–`, so G2 rejects it on every pass. Previously (pre-0.4.1 unified rule) pass 2 produced `1⟨J⟩–⟨J⟩1⟨J⟩–⟨J⟩1` |
|
|
419
|
+
| 46 | `1–1␣-␣1` | ⟶ | the **control**: the first token is already exactly correct for `en-tight` (glyph, length, spacing), so §3.3's invisible-edit test leaves it unbound and unchanged; the second token's G2 still rejects it, because `before` is a real `DASH` regardless of which `DASH` glyph it is. `1—1 - 1` (an em dash where the locale wants en) is **not** a fixed point here — it converts to `1⟨J⟩–⟨J⟩1 - 1`, the same as row 45 |
|
|
420
|
+
|
|
421
|
+
### `ru` — `range: "em-tight"`
|
|
422
|
+
|
|
423
|
+
| # | Input | Output | Why |
|
|
424
|
+
| --- | ------------------------------- | --------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
|
|
425
|
+
| 17 | `Годы 1941-1945 были тяжёлыми` | `Годы 1941⟨J⟩—⟨J⟩1945 были тяжёлыми` | Russian sets an **em** dash in ranges; `range: "em-tight"`. Aligned to the shipped fixture — which pair of years the row uses is arbitrary |
|
|
426
|
+
| 47 | `1-1␣-␣1` | `1⟨J⟩—⟨J⟩1␣-␣1` | the defect of row 45 reproduced in every locale whose `range` is not `none`; `ru` sets `em-tight` |
|
|
427
|
+
|
|
428
|
+
---
|
|
429
|
+
|
|
430
|
+
## 7. Open questions
|
|
431
|
+
|
|
432
|
+
Every item below is current and open against `ranges` specifically. Items that are shared
|
|
433
|
+
machinery rather than range-specific stay in [dashes.md](dashes.md) §7 (its canonical home,
|
|
434
|
+
§3.1 above); items that are purely historical are [dashes.md](dashes.md) §8.
|
|
435
|
+
|
|
436
|
+
1. **`DIGIT` is ASCII-only** (the class itself is defined in [dashes.md](dashes.md) §3.1, shared
|
|
437
|
+
with `dashes`). Arabic-Indic, Devanagari and fullwidth digits are not recognised, so
|
|
438
|
+
`١٩١٤-١٩١٨` is left alone. The limitation lands here rather than in `dashes.md` §7 because only
|
|
439
|
+
`ranges`' G4/G5 actually read a digit's *value* — `dashes` never does. Extending to Unicode
|
|
440
|
+
`Nd` would require every runtime to agree on a Unicode version _and_ a digit-value table;
|
|
441
|
+
ASCII-only is the portable choice for the v1 locale set. Recorded so the limitation is
|
|
442
|
+
deliberate rather than accidental.
|
|
443
|
+
2. **Decimal ranges.** `1.5-2.5` and `1,5-2,5` are legitimate ranges and are currently rejected
|
|
444
|
+
by G3. Supporting them means deciding which of U+002E / U+002C is the decimal separator in a
|
|
445
|
+
given locale, which is data the schema does not carry. Left unsupported.
|
|
446
|
+
3. **Abbreviated year ranges.** `2020-24` is a real convention (Chicago) and is rejected by G4.
|
|
447
|
+
The false-positive risk of relaxing G4 (it would also accept `8859-1`) was judged worse than
|
|
448
|
+
the miss. Needs an operator decision.
|
|
449
|
+
4. **Scores and votes.** `5-0`, `2-1` pass all guards and become en dashes. That is arguably
|
|
450
|
+
correct typographically (an en dash is standard for scores), but G5 rejects `5-0` while
|
|
451
|
+
accepting `0-5`, which is inconsistent for this use. Either scores are out of scope or G5
|
|
452
|
+
needs an exception; unresolved.
|
|
453
|
+
5. **The U+2060 range binding is invisible in both the rendered text and a plain diff, and that
|
|
454
|
+
is a reviewing hazard, not just a curiosity.** Every other change this rule makes is catchable
|
|
455
|
+
by eye — a hyphen becomes a dash, a space appears or disappears. §3.3.1's joiner is zero-width:
|
|
456
|
+
a bound form and its unbound equivalent are pixel-identical in every font, and a unified diff
|
|
457
|
+
shows a changed line with no visible difference on it. **A worked-example or fixture row that
|
|
458
|
+
omits the joiner looks correct on screen and is wrong at the byte level** — §6's own standing
|
|
459
|
+
note says so, and the escaped mirror in `spec/fixtures/.escaped/` is the only rendering in
|
|
460
|
+
which the joiner is visible; check it, never read the row on screen. This is not an argument
|
|
461
|
+
against the binding, which is well-sourced and correct, only a property this rule (and any
|
|
462
|
+
future rule emitting a zero-width or invisible code point — U+00AD, U+200B, a variation
|
|
463
|
+
selector) has to manage explicitly rather than assume a reviewer will catch by eye. The
|
|
464
|
+
historical cost of this exact hazard — 32 fixture cases silently broken, and one idempotency
|
|
465
|
+
defect needing a seven-character witness — is recorded at [dashes.md](dashes.md) §8.6, from
|
|
466
|
+
when this rule's behaviour still lived there.
|
|
467
|
+
6. **The compound-label ambiguity (`Figure 5-10`, G4's `(1,2)` branch) is already the full current
|
|
468
|
+
statement of this rule's central tradeoff — see §5 above, not duplicated here.**
|
|
469
|
+
7. **A multi-code-point closed-up symbol is not admitted** (§3.2a). `US$15-US$20` and `R$15-R$20`
|
|
470
|
+
carry a two- or three-code-point prefix, and `15°C-20°C` a two-code-point suffix; §3.2a
|
|
471
|
+
consumes at most one code point per side, so all three are declined. Admitting them means
|
|
472
|
+
deciding how far to walk and what stops the walk, and a walk that crosses `LETTER` would
|
|
473
|
+
collide with G1, which exists to keep `MP3-4` and `H2-2` out. The narrow rule is the one the
|
|
474
|
+
citation supports; the wider one needs its own evidence.
|
|
475
|
+
8. **A spaced unit is deliberately out of scope, and so is the mirror case.** `15 kg-20 kg` and
|
|
476
|
+
`225 nm-2400 nm` are not admitted, because the source §3.2a quotes says a symbol *separated by
|
|
477
|
+
a space* is **not** repeated — the spaced form is a different construction, and NIST SP 811
|
|
478
|
+
§7.7 goes further and recommends the word "to" rather than a range dash for quantities, on the
|
|
479
|
+
ground that a dash can be read as a minus sign. Separately: in `de-DE`, `fr` and `ru` the
|
|
480
|
+
currency sign follows the amount (`30 EUR`, `800 руб.`), so the money case in those locales is
|
|
481
|
+
shaped `15 €-20 €` — a spaced symbol, not a closed-up one, and therefore out of scope by the
|
|
482
|
+
same clause rather than by oversight. Whether a spaced, repeated unit should be admitted at
|
|
483
|
+
all is an open question with evidence pointing away from it.
|
|
484
|
+
**This is also why a `CLOSED-SYMBOL` member appearing in a locale's `nbsp.beforeUnits` does
|
|
485
|
+
not contradict the locale data.** Several locale files list `°C`, and some list `%`, as units
|
|
486
|
+
the locale sets **after a space** — `20 °C`, `20 %`. That is the spaced construction, out of
|
|
487
|
+
scope here; `15°-20°` and `35%-50%`, with the symbol closed up, are the other one. The two
|
|
488
|
+
never meet in the algorithm either: §3.2a's walk requires a `DIGIT` immediately beside the
|
|
489
|
+
symbol, so `15 %-20 %` is not a candidate at all.
|