polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
|
@@ -0,0 +1,353 @@
|
|
|
1
|
+
# Rule: `hyphen`
|
|
2
|
+
|
|
3
|
+
**Order:** 35 (between `dashes` and `quotes`). **Default:** on.
|
|
4
|
+
**Modes:** text, html, markdown, yaml.
|
|
5
|
+
**Spec version:** 0.1.0.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Purpose
|
|
10
|
+
|
|
11
|
+
`hyphen` replaces U+002D (hyphen-minus) with U+2011 (non-breaking hyphen) inside a closed list
|
|
12
|
+
of morphological forms whose hyphen must not be broken across a line. It exists for Russian,
|
|
13
|
+
where `из-под`, `кое-что` and `сделал-таки` are single words that a line-breaking algorithm
|
|
14
|
+
will happily split at the hyphen, producing a line ending in `из-` — which every Russian style
|
|
15
|
+
guide forbids. It is deliberately the most data-bound rule in the pipeline: it does exactly
|
|
16
|
+
nothing except where a locale has listed a literal form, so for `en`, `de`, `fr`, `fi`, `sv`
|
|
17
|
+
and `el` — whose lists are empty — it is a provable total no-op, not merely an unlikely one. It
|
|
18
|
+
changes one code point to another code point of the same width and semantics; it never
|
|
19
|
+
inserts, never deletes, never touches spacing, and never converts anything that is not
|
|
20
|
+
already a hyphen inside a word.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## 2. Locale data consumed
|
|
25
|
+
|
|
26
|
+
- `hyphen.prefixes` — word-initial forms written **with a trailing hyphen**, e.g. `"кое-"`
|
|
27
|
+
- `hyphen.suffixes` — word-final forms written **with a leading hyphen**, e.g. `"-таки"`
|
|
28
|
+
- `hyphen.compounds` — whole hyphenated forms written with U+002D, e.g. `"из-под"`
|
|
29
|
+
|
|
30
|
+
All three are required by `locale.schema.json` and all three may be empty.
|
|
31
|
+
|
|
32
|
+
**Evidence for these lists.** A locale attests **membership** — that these tokens belong in this list — and this rule owns the **mechanism** it applies to them. A `sources` citation is not required to name a code point or a binding the locale file has no way to vary. The principle is stated once, normatively, in [nbsp.md](nbsp.md) §2.1 and governs every list-valued field in `locale.schema.json`. Entries are
|
|
33
|
+
literal lowercase forms of at least two code points. An entry containing no U+002D is
|
|
34
|
+
meaningless and must raise `POLYTYPO_MALFORMED_LOCALE_DATA` — the rule has nothing to convert
|
|
35
|
+
in it.
|
|
36
|
+
|
|
37
|
+
**If all three arrays are empty, the rule emits nothing for any input.** An implementation may
|
|
38
|
+
short-circuit on that condition; the observable behaviour is identical either way.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## 3. Algorithm
|
|
43
|
+
|
|
44
|
+
Input is a code-point array `cp[0 … n-1]`.
|
|
45
|
+
|
|
46
|
+
### 3.1 Character classes
|
|
47
|
+
|
|
48
|
+
| Class | Members |
|
|
49
|
+
| ----------- | -------------------------------------------------------------------------------------------------------------- |
|
|
50
|
+
| `HY` | U+002D (hyphen-minus) |
|
|
51
|
+
| `NBHY` | U+2011 (non-breaking hyphen) |
|
|
52
|
+
| `HYPHENISH` | `HY` ∪ `NBHY` |
|
|
53
|
+
| `LETTER` | general category `Lu`, `Ll`, `Lt`, `Lm`, `Lo`, `Mn`, `Mc`, `Me` (combining marks count as letter-continuation) |
|
|
54
|
+
| `UPPER` | general category `Lu` or `Lt`. Used only by §3.3's first-character leniency, via the simple uppercase mapping — no guard tests it directly |
|
|
55
|
+
| `DIGIT` | U+0030–U+0039 |
|
|
56
|
+
| `ALNUM` | `LETTER` ∪ `DIGIT` |
|
|
57
|
+
| `WORDISH` | `ALNUM` ∪ `HYPHENISH` — the characters that continue a hyphenated word |
|
|
58
|
+
| `NONE` | index out of range |
|
|
59
|
+
|
|
60
|
+
No other class is examined.
|
|
61
|
+
|
|
62
|
+
**Unicode version.** The general categories and case mappings this rule reads are those of the UCD version pinned in `spec/UNICODE` (`17.0`). The pin is normative for the **derived tables**, not for the host runtime — see [pipeline-idempotency.md](pipeline-idempotency.md) §6a, which also specifies the canary fixtures that make the pin detectable. In particular this rule never looks at a space, and never at
|
|
63
|
+
U+2010, U+00AD, U+2013 or U+2014.
|
|
64
|
+
|
|
65
|
+
### 3.2 Why order 35, and what it may assume
|
|
66
|
+
|
|
67
|
+
`hyphen` runs **after `dashes` and before `quotes`**. Both halves of that placement are
|
|
68
|
+
load-bearing.
|
|
69
|
+
|
|
70
|
+
**After `dashes`.** `dashes` guard P1 declines every bare U+002D that is not spaced on both
|
|
71
|
+
sides, so every intra-word hyphen — which is exactly this rule's entire subject matter —
|
|
72
|
+
survives `dashes` untouched. Running `hyphen` first would invert the dependency: `dashes`
|
|
73
|
+
would then encounter U+2011 in positions where it expects U+002D, and while its `INERT-DASH`
|
|
74
|
+
class already covers U+2011, its verdicts would be reached for a different reason than the
|
|
75
|
+
one documented, and the two rules' guards would have to be kept in sync by hand. Running
|
|
76
|
+
second means `hyphen` may assume, and an implementation may assert, that **every U+002D it
|
|
77
|
+
converts has a letter or digit on both sides** — the complement of what `dashes` claims. That
|
|
78
|
+
assumption is what makes §5 short.
|
|
79
|
+
|
|
80
|
+
**Before `quotes`.** This ordering is materially free — the two rules share no characters —
|
|
81
|
+
but it is not entirely free, and the reason is worth recording. `quotes.md` §3.1 lists
|
|
82
|
+
U+002D, U+2013 and U+2014 as acceptable left context for an opening quotation mark
|
|
83
|
+
(`canOpen`). U+2011 was not in that list. If `hyphen` converted a U+002D that sat immediately
|
|
84
|
+
left of a quote candidate, that candidate's `canOpen` would flip from true to false between
|
|
85
|
+
one pipeline run and the next — the exact class of cross-rule drift that the idempotency
|
|
86
|
+
invariant exists to catch. The shape is unreachable in practice (a `prefixes` entry requires a
|
|
87
|
+
`LETTER` after its hyphen, so a quotation mark cannot follow), but relying on unreachability
|
|
88
|
+
across two documents is how ports diverge. **`quotes.md`, `apostrophe.md` and `nbsp.md` have
|
|
89
|
+
therefore been amended to include U+2011 wherever they already list U+002D/U+2013/U+2014 as
|
|
90
|
+
context.** With that amendment, order 35 versus order 45 is genuinely immaterial, and 35 is
|
|
91
|
+
kept because it groups the two hyphen-and-dash rules together.
|
|
92
|
+
|
|
93
|
+
### 3.3 Matching
|
|
94
|
+
|
|
95
|
+
All three lists use the same comparison, called **hyphen-lenient, first-character-lenient
|
|
96
|
+
literal matching**. For a pattern `w` of `k` code points at candidate index `a`:
|
|
97
|
+
|
|
98
|
+
- **hyphen leniency, at every position `j` in `0 … k-1`:** if `w[j]` is in `HY`, then
|
|
99
|
+
`cp[a+j]` may be in `HY` **or** in `NBHY`; otherwise `cp[a+j] = w[j]` exactly.
|
|
100
|
+
The leniency must cover `j = 0` and not merely `j ≥ 1`, because a **suffix** entry _begins_
|
|
101
|
+
with its hyphen: `-таки` converted once becomes `‑таки`, and a pattern that demanded U+002D
|
|
102
|
+
at position 0 would stop matching its own output. §5's round-trip argument depends on this.
|
|
103
|
+
- **first-character case leniency, at `j = 0` only:** either `cp[a] = w[0]`, or `w[0]` is in
|
|
104
|
+
`LETTER` and `cp[a]` is the Unicode **simple uppercase mapping** of `w[0]`. The two
|
|
105
|
+
leniencies never overlap — `w[0]` is a hyphen or a letter, never both.
|
|
106
|
+
|
|
107
|
+
Nothing else matches. `ИЗ-ПОД` does not match `из-под`; `Из-под` does.
|
|
108
|
+
|
|
109
|
+
**Capitalisation without host-locale case folding.** The first-character leniency uses the
|
|
110
|
+
Unicode simple uppercase mapping — a plain code-point→code-point table from the UCD, applied
|
|
111
|
+
to the _pattern_ (which the schema fixes as lowercase), not to the input. It is not
|
|
112
|
+
`toUpperCase()`, not `toLocaleUpperCase()`, not ICU, and it does not depend on the host
|
|
113
|
+
process's locale (ARCHITECTURE.md §4.4). Mapping the pattern rather than the input is
|
|
114
|
+
deliberate: it means at most one extra code point is computed per list entry, it can be
|
|
115
|
+
precomputed once when the locale file is loaded, and the notorious Turkish pair never arises
|
|
116
|
+
because the simple mapping of `i` is `I` unconditionally and no v1 `hyphen` entry begins with
|
|
117
|
+
`i` in any case. This is the same convention `nbsp.afterShortWords` uses (`nbsp.md` §3.5), and
|
|
118
|
+
having one convention across the spec is worth more than the marginal extra coverage a second
|
|
119
|
+
one would buy.
|
|
120
|
+
|
|
121
|
+
All-capitals forms (`ИЗ-ПОД` in a heading) are deliberately **not** matched — see §7.1.
|
|
122
|
+
|
|
123
|
+
### 3.4 Scan
|
|
124
|
+
|
|
125
|
+
Walk `a` from `0` to `n-1`. At each index, select **one** candidate entry: the longest entry
|
|
126
|
+
that matches at `a` across all three lists, with ties broken in the order **compounds,
|
|
127
|
+
prefixes, suffixes**. Entries are unique (`uniqueItems`), so this is a total order and the
|
|
128
|
+
selection is deterministic.
|
|
129
|
+
|
|
130
|
+
**Matching and guarding are separate steps, and there is no backtracking.** The selected entry
|
|
131
|
+
is then checked against its guards. If it passes, it claims the position and the scan continues
|
|
132
|
+
from `a + k`. **If it fails a guard, the sub-rule emits nothing at `a` and no other entry is
|
|
133
|
+
tried there**; the scan advances `a` by one. An earlier revision said "the first entry that
|
|
134
|
+
matches _and passes its guards_", which is backtracking, and it disagreed with `nbsp`'s
|
|
135
|
+
list-driven sub-rules. `nbsp.md` §3.5 now states the same policy, and the two must stay
|
|
136
|
+
aligned: longest match wins, guards are applied to the winner only.
|
|
137
|
+
|
|
138
|
+
Compounds are tried first because a compound is the most specific form: if a locale lists both
|
|
139
|
+
the compound `из-под` and the prefix `из-`, the compound's verdict (bind the hyphen inside the
|
|
140
|
+
whole word) must win, and it happens to produce the same edit — but the scan position it
|
|
141
|
+
consumes differs, and that must be deterministic.
|
|
142
|
+
|
|
143
|
+
**C — compounds.** For an entry `w` of `k` code points matching at `a`:
|
|
144
|
+
|
|
145
|
+
1. **Left boundary.** `cp[a-1]` must be `NONE` or **not** in `WORDISH`.
|
|
146
|
+
2. **Right boundary.** `cp[a+k]` must be `NONE` or **not** in `WORDISH`.
|
|
147
|
+
Together these mean the compound is a whole word: `из-под` matches in `из-под стола` but
|
|
148
|
+
not inside `квазииз-подный` or `из-под-` .
|
|
149
|
+
3. For every position `j` with `w[j]` in `HY`: if `cp[a+j]` is already in `NBHY`, emit nothing
|
|
150
|
+
for that position; otherwise emit an edit replacing `cp[a+j]` with U+2011.
|
|
151
|
+
|
|
152
|
+
**P — prefixes.** An entry `w` ends with U+002D; let `k` be its length.
|
|
153
|
+
|
|
154
|
+
1. **Left boundary.** `cp[a-1]` must be `NONE` or not in `WORDISH` — the prefix starts a word.
|
|
155
|
+
2. **Right boundary.** `cp[a+k]` must be in `LETTER`. A prefix must actually prefix something:
|
|
156
|
+
`кое-что` matches, `кое-` at the end of a line or before a space does not, and neither does
|
|
157
|
+
`кое-2`.
|
|
158
|
+
3. Emit an edit replacing the final code point of the match (`cp[a+k-1]`, the hyphen) with
|
|
159
|
+
U+2011, unless it is already in `NBHY`.
|
|
160
|
+
|
|
161
|
+
**S — suffixes.** An entry `w` begins with U+002D; let `k` be its length. Because a suffix is
|
|
162
|
+
matched at its own start index, the scan finds it at the hyphen, not at the start of the word.
|
|
163
|
+
|
|
164
|
+
1. **Left boundary.** `cp[a-1]` must be in `LETTER`. A suffix must actually suffix something:
|
|
165
|
+
`сделал-таки` matches, a line beginning `-таки` does not.
|
|
166
|
+
2. **Right boundary.** `cp[a+k]` must be `NONE` or not in `WORDISH`.
|
|
167
|
+
3. Emit an edit replacing the first code point of the match (`cp[a]`, the hyphen) with U+2011,
|
|
168
|
+
unless it is already in `NBHY`.
|
|
169
|
+
|
|
170
|
+
Note that the first-character leniency of §3.3 is inert for suffixes (their first code point
|
|
171
|
+
is a hyphen, not a letter) and for compounds it applies to the word's initial letter, which is
|
|
172
|
+
the sentence-start case (`Из-под стола…`).
|
|
173
|
+
|
|
174
|
+
---
|
|
175
|
+
|
|
176
|
+
## 4. Must not touch
|
|
177
|
+
|
|
178
|
+
**Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
|
|
179
|
+
a guarantee of `transform` as a whole — or **[R]** — true of this rule in isolation but capable
|
|
180
|
+
of being falsified by another rule, which is then named.
|
|
181
|
+
|
|
182
|
+
- **[P] Any hyphen not covered by a listed form.** `научно-технический`, `Jean-Luc`, `e-mail`,
|
|
183
|
+
`well-known`, `COVID-19` are all left as U+002D. This rule has no morphology of its own; it
|
|
184
|
+
has a list.
|
|
185
|
+
- **[P] Every locale with empty lists.** `en`, `de`, `fr`, `fi`, `sv`, `el` — the rule is a total no-op
|
|
186
|
+
and must produce byte-identical output (§2).
|
|
187
|
+
- **[R] A hyphen that `dashes` owns.** Anything spaced on both sides was already handled at order
|
|
188
|
+
30; by §3.2 this rule only ever sees intra-word hyphens, and its boundary guards enforce
|
|
189
|
+
that independently rather than trusting it.
|
|
190
|
+
- **[P] U+2010 (hyphen), U+00AD (soft hyphen), U+2012, U+2013, U+2014, U+2212.** Not in
|
|
191
|
+
`HYPHENISH`; never read, never written. A soft hyphen inside `из-под` is the author's
|
|
192
|
+
deliberate break hint and is not this rule's business.
|
|
193
|
+
- **[R] Spacing of any kind.** No space is inserted, removed or converted. Every edit is one code
|
|
194
|
+
point replacing one code point at the same index.
|
|
195
|
+
- **[P] All-capitals forms** — `ИЗ-ПОД` (§7.1).
|
|
196
|
+
- **[P] A listed form appearing inside a longer word.** Boundary guards C1/C2, P1, S2.
|
|
197
|
+
- **[P] Anything inside a skipped region.** A hyphen in a URL, a code span, a fenced block or an
|
|
198
|
+
HTML attribute is removed by the mode adapter (L2) before this rule runs. `из-под` in a
|
|
199
|
+
slug is therefore safe; this rule has no URL-awareness and must not acquire any.
|
|
200
|
+
|
|
201
|
+
---
|
|
202
|
+
|
|
203
|
+
## 5. Idempotency argument
|
|
204
|
+
|
|
205
|
+
Let `T` be the rule.
|
|
206
|
+
|
|
207
|
+
**The output form is recognised as already-correct.** Every edit replaces a U+002D with a
|
|
208
|
+
U+2011 at the same index. §3.3's comparison is **hyphen-lenient**: a pattern's U+002D matches
|
|
209
|
+
either U+002D or U+2011 in the text. So on the second run the same list entry matches at the
|
|
210
|
+
same index `a` with the same length `k` — the match is not lost by the conversion, which is
|
|
211
|
+
the property the leniency exists for. The emission step of each branch then finds `cp[a+j]`
|
|
212
|
+
(or `cp[a+k-1]`, or `cp[a]`) already in `NBHY` and emits nothing.
|
|
213
|
+
|
|
214
|
+
**No guard's verdict changes.** The guards test:
|
|
215
|
+
|
|
216
|
+
- membership in `LETTER`, `UPPER`, `DIGIT` — no edit ever writes a character in any of these
|
|
217
|
+
classes, and no edit ever removes one, since every edit is a 1:1 replacement of a hyphen;
|
|
218
|
+
- membership in `WORDISH` — which contains **both** U+002D and U+2011 by construction
|
|
219
|
+
(`HYPHENISH ⊂ WORDISH`). This is the second place the design has to be deliberate: if
|
|
220
|
+
`WORDISH` contained only U+002D, then converting a hyphen would change a neighbouring
|
|
221
|
+
form's boundary test from "is in `WORDISH`, reject" to "is not in `WORDISH`, accept", and a
|
|
222
|
+
form adjacent to a converted one could start matching on the second run. With `NBHY` in
|
|
223
|
+
`WORDISH`, the boundary verdicts are invariant.
|
|
224
|
+
|
|
225
|
+
**Scan positions are stable.** The scan advances by `a + k` after a claim, and `k` is the
|
|
226
|
+
pattern length, which does not depend on the text. Since every edit is 1:1, no index shifts.
|
|
227
|
+
So the same entries claim the same positions in the same order on every run.
|
|
228
|
+
|
|
229
|
+
Hence `T(T(x)) = T(x)`. ∎
|
|
230
|
+
|
|
231
|
+
**Input that already contains U+2011.** This is the same case as "the rule's own output", and
|
|
232
|
+
it is handled identically: a form written by the author as `из‑под` with a real U+2011 matches
|
|
233
|
+
(hyphen-lenient), passes its guards, and produces **no edit**. A U+2011 outside any listed
|
|
234
|
+
form is never examined at all. So text that is already correctly bound round-trips
|
|
235
|
+
byte-identically, which is the round-trip requirement of ARCHITECTURE.md §6.3 item 4 and the
|
|
236
|
+
reason the leniency is specified rather than left to an implementation.
|
|
237
|
+
|
|
238
|
+
**Stability under the rest of the pipeline.** `dashes` runs before this rule within a single
|
|
239
|
+
pipeline pass, but on the _next_ pass it sees the U+2011 this rule produced. Its verdicts do
|
|
240
|
+
not change: `dashes` §3.1 puts U+2011 in `INERT-DASH`, its §3.2 step 6 rejects a token whose
|
|
241
|
+
neighbour is in `INERT-DASH` **or** in `DASH` alike, and its cluster alphabet contains both —
|
|
242
|
+
so a position that was declined when it held U+002D is declined when it holds U+2011, for the
|
|
243
|
+
same reason. `quotes`, `apostrophe` and `nbsp` list U+2011 alongside U+002D/U+2013/U+2014 in
|
|
244
|
+
their context classes (§3.2), so their classifications are likewise unaffected.
|
|
245
|
+
|
|
246
|
+
**What had to be fixed.** The naive formulation — _"replace the hyphen in each listed form
|
|
247
|
+
with U+2011"_ — is not idempotent, because on the second run the form no longer matches: the
|
|
248
|
+
pattern holds U+002D and the text now holds U+2011, so the rule silently stops recognising its
|
|
249
|
+
own output. That is harmless for the output itself but it is a latent trap, because it means
|
|
250
|
+
the rule's behaviour depends on whether it has run before, and any later guard that consults a
|
|
251
|
+
list membership would then disagree between runs. Making the comparison hyphen-lenient, and
|
|
252
|
+
putting U+2011 into `WORDISH`, are the two changes that make the rule's view of the text
|
|
253
|
+
identical before and after its own edits.
|
|
254
|
+
|
|
255
|
+
---
|
|
256
|
+
|
|
257
|
+
### Composition obligation
|
|
258
|
+
|
|
259
|
+
Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₄**, so the
|
|
260
|
+
obligation runs against `spaces` (R₁), `ellipsis` (R₂) and `dashes` (R₃).
|
|
261
|
+
|
|
262
|
+
**What this rule emits.** U+2011, one code point replacing one U+002D at the same index, inside a
|
|
263
|
+
matched form. It inserts nothing, deletes nothing, and changes no length.
|
|
264
|
+
|
|
265
|
+
**Against `I₁` (`spaces`).** Discharged: no U+0020 is emitted or deleted, and no code point is
|
|
266
|
+
inserted or removed, so no space run can change length or neighbourhood.
|
|
267
|
+
|
|
268
|
+
**Against `I₂` (`ellipsis`).** Discharged: U+2011 is not in `DOTLIKE` and no dot run is
|
|
269
|
+
touched.
|
|
270
|
+
|
|
271
|
+
**Against `I₃` (`dashes`).** This is the only one with content. Converting U+002D to U+2011
|
|
272
|
+
changes a code point that `dashes` classifies — but it moves it from `DASH` to `INERT-DASH`,
|
|
273
|
+
and `dashes` treats the two identically in every place a _neighbouring_ token could read them:
|
|
274
|
+
its isolation guard (§3.2 step 6) rejects both, guard G2 rejects both, and its cluster alphabet
|
|
275
|
+
contains both. The converted hyphen itself was never a `dashes` token: `dashes` guard P1
|
|
276
|
+
declines a bare U+002D that is not spaced on both sides, and this rule only ever claims a
|
|
277
|
+
hyphen with a letter or a listed form's interior on each side. So no verdict changes, and the
|
|
278
|
+
positions this rule edits are exactly the ones `dashes` had already declined.
|
|
279
|
+
|
|
280
|
+
---
|
|
281
|
+
|
|
282
|
+
## 6. Worked examples
|
|
283
|
+
|
|
284
|
+
`⟶` = no change, `‑` = U+2011 (non-breaking hyphen), `-` = U+002D.
|
|
285
|
+
|
|
286
|
+
### `ru` — `compounds: ["из-под","из-за","по-моему"]`, `prefixes: ["кое-"]`, `suffixes: ["-таки","-то","-либо","-нибудь"]`
|
|
287
|
+
|
|
288
|
+
| # | Input | Output | Why |
|
|
289
|
+
| --- | ----------------------------- | --------------------------- | -------------------------------------------------------------------------------- |
|
|
290
|
+
| 1 | `Достал из-под стола` | `Достал из‑под стола` | compound, both boundaries non-`WORDISH` |
|
|
291
|
+
| 2 | `Из-под стола донёсся звук` | `Из‑под стола донёсся звук` | first-character leniency: simple uppercase mapping of `и` |
|
|
292
|
+
| 3 | `ИЗ-ПОД СТОЛА` | ⟶ | all-capitals is not matched (§7.1) |
|
|
293
|
+
| 4 | `кое-что и кое-как` | `кое‑что и кое‑как` | prefix; `cp[a+k]` is a letter in both |
|
|
294
|
+
| 5 | `Он сделал-таки это` | `Он сделал‑таки это` | suffix; `cp[a-1]` is a letter |
|
|
295
|
+
| 6 | `что-нибудь или что-либо` | `что‑нибудь или что‑либо` | two suffixes |
|
|
296
|
+
| 7 | `Достал из‑под стола` | ⟶ | already U+2011 — hyphen-lenient match, no edit. **The round-trip case** |
|
|
297
|
+
| 8 | `научно-технический прогресс` | ⟶ | not in any list |
|
|
298
|
+
| 9 | `кое- и кое-что` | `кое- и кое‑что` | the first `кое-` fails P2 (`cp[a+k]` is a space, not a letter); the second binds |
|
|
299
|
+
| 10 | `квазииз-подный` | ⟶ | compound guard C1: `cp[a-1]` is `и`, in `WORDISH` |
|
|
300
|
+
| 11 | `Москва — из-за дождя` | `Москва — из‑за дождя` | the em dash is `dashes`' business and is untouched here; the compound binds |
|
|
301
|
+
| 12 | `-таки в начале строки` | ⟶ | suffix guard S1: `cp[a-1]` is `NONE` |
|
|
302
|
+
|
|
303
|
+
### `en`, `de`, `fr`, `fi`, `sv`, `el` — all three lists empty
|
|
304
|
+
|
|
305
|
+
| # | Input | Output | Why |
|
|
306
|
+
| --- | ------------------------------------------------- | ------ | ---------------------------- |
|
|
307
|
+
| 13 | `A well-known e-mail address, COVID-19, Jean-Luc` | ⟶ | no entries; total no-op (§2) |
|
|
308
|
+
| 14 | `Any text at all` | ⟶ | same |
|
|
309
|
+
|
|
310
|
+
Cases 3, 7, 8, 10, 12, 13 and 14 are "no change" cases.
|
|
311
|
+
|
|
312
|
+
---
|
|
313
|
+
|
|
314
|
+
## 7. Open questions
|
|
315
|
+
|
|
316
|
+
1. **All-capitals forms are not matched.** `ИЗ-ПОД СТОЛА` in a Russian heading keeps its
|
|
317
|
+
breakable hyphen. Matching it needs either a full uppercase comparison of every code point
|
|
318
|
+
(which requires an explicit case-mapping table in the spec, not just the single first
|
|
319
|
+
character) or a second set of list entries. Russian headings are frequently set in caps, so
|
|
320
|
+
this is a real miss, not a theoretical one. I did not add a full case table because it is a
|
|
321
|
+
cross-cutting decision — `nbsp.afterShortWords` has exactly the same limitation — and
|
|
322
|
+
because ARCHITECTURE.md §4.4 wants any such table to be explicit spec data. Recommend
|
|
323
|
+
deciding it once, for both rules, together with the Unicode-version pinning question.
|
|
324
|
+
2. **Title Case within a form is not matched.** `Из-Под` (a stylised heading) does not match.
|
|
325
|
+
Same root cause as §7.1.
|
|
326
|
+
3. **`hyphen.prefixes` and `hyphen.suffixes` cannot express a constraint on what follows or
|
|
327
|
+
precedes** beyond "a letter". Russian `-то` is a particle in `кто-то` but a conjunction in
|
|
328
|
+
other positions, and `по-` is a prefix in `по-моему` but a preposition elsewhere. The lists
|
|
329
|
+
are blunt instruments; the mitigation is to list only forms whose binding is
|
|
330
|
+
unconditional, and `по-` is deliberately shown in §6 as a **compound** (`по-моему`) rather
|
|
331
|
+
than a prefix for that reason. Whether the Russian locale file should list `по-` as a
|
|
332
|
+
prefix at all is a locale-research question with a normative source (Мильчин), not a
|
|
333
|
+
question about this rule.
|
|
334
|
+
4. **U+2011 has real-world rendering gaps.** A few older fonts lack the glyph and fall back
|
|
335
|
+
visibly. The alternatives are U+2010 plus U+2060 (word joiner), or a zero-width no-break
|
|
336
|
+
space after the hyphen — both worse in other ways, and both invisible in a diff. U+2011 is
|
|
337
|
+
the operator's stated choice and is recorded here as such, but if the M4 dogfooding gate
|
|
338
|
+
shows a rendering problem in the author's own font stack, this is the decision to revisit.
|
|
339
|
+
5. **Nothing prevents a locale from listing an entry with no hyphen in it.** §2 requires
|
|
340
|
+
`POLYTYPO_MALFORMED_LOCALE_DATA`, but `locale.schema.json`'s `wordList` only constrains
|
|
341
|
+
length and uniqueness, so the check is a runtime one rather than a CI one. A
|
|
342
|
+
`"pattern"`-free way to express it in JSON Schema does not exist, so this probably belongs
|
|
343
|
+
in `scripts/validate-spec.mjs` alongside the one-code-point check. **Reported.**
|
|
344
|
+
6. **The interaction with hyphenation is untested.** A CSS `hyphens: auto` renderer may still
|
|
345
|
+
break inside `из‑под` at a syllable boundary; U+2011 only forbids the break _at the hyphen_.
|
|
346
|
+
That is out of scope (PLAN.md §4 excludes hyphenation) but worth knowing before claiming
|
|
347
|
+
the rule "prevents the word from breaking".
|
|
348
|
+
7. **Ordinal repair is refused** — `10-ый` → `10-й`, and the whole family of Russian
|
|
349
|
+
ordinal-ending corrections Lebedev's service performs. This rule binds a hyphen that is
|
|
350
|
+
already correct; it does not rewrite the morphology around one. Changing `ый` to `й` is a
|
|
351
|
+
spelling correction, and PLAN.md §4 refuses spellcheck and typo correction without an
|
|
352
|
+
explicit operator decision. Recorded so the presence of `hyphen.suffixes` is not read as an
|
|
353
|
+
invitation.
|
|
@@ -0,0 +1,239 @@
|
|
|
1
|
+
# Locale resolution
|
|
2
|
+
|
|
3
|
+
**Not a pipeline rule.** This document has no entry in `spec/rules/order.json` and produces no
|
|
4
|
+
edits. It specifies the function that turns the caller's `locale` option into the identifier
|
|
5
|
+
of exactly one locale file, and it lives in `spec/rules/` because it is normative,
|
|
6
|
+
runtime-independent, and fixture-covered like everything else here.
|
|
7
|
+
**Spec version:** 0.1.0.
|
|
8
|
+
|
|
9
|
+
---
|
|
10
|
+
|
|
11
|
+
## 1. Purpose
|
|
12
|
+
|
|
13
|
+
`transform(input, { locale })` takes a locale tag from the caller and must select exactly one
|
|
14
|
+
`spec/locales/<id>.json` file. That selection has to be identical in five runtimes, which is
|
|
15
|
+
precisely why it cannot be delegated to a platform locale-negotiation library: ICU's
|
|
16
|
+
`ResourceBundle`, Go's `golang.org/x/text/language` matcher, PHP's `Locale::lookup`, Python's
|
|
17
|
+
`babel.negotiate_locale` and the JS `Intl.Locale` machinery all implement _different_
|
|
18
|
+
fallback policies, several of them probabilistic in the sense that they will happily return a
|
|
19
|
+
"best effort" match for a language the library has never heard of. polytypo does the opposite:
|
|
20
|
+
it resolves exactly, then by declared alias, then by dropping the region subtag once — and if
|
|
21
|
+
that fails it **throws**. It never falls back to English, never to the first entry in the
|
|
22
|
+
registry, and never to a "closest" language. Silently applying Finnish quotation rules to a
|
|
23
|
+
Portuguese catalogue is the failure mode that destroys trust in a tool like this
|
|
24
|
+
(PLAN.md §5.1), and it is cheaper to make the caller fix a typo than to make them discover
|
|
25
|
+
it in production.
|
|
26
|
+
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
## 2. Data consumed
|
|
30
|
+
|
|
31
|
+
`spec/locales/registry.json`, in full:
|
|
32
|
+
|
|
33
|
+
- `locales` — array of concrete locale identifiers, each of which has a file
|
|
34
|
+
`spec/locales/<id>.json`;
|
|
35
|
+
- `aliases` — object mapping a tag to a member of `locales`.
|
|
36
|
+
|
|
37
|
+
Nothing else. In particular the individual locale files are not read during resolution, and
|
|
38
|
+
the algorithm never inspects the filesystem (ARCHITECTURE.md §3.1: locale data is embedded in
|
|
39
|
+
the artifact).
|
|
40
|
+
|
|
41
|
+
**Registry invariants.** These are preconditions on the data, checked in spec CI, not at
|
|
42
|
+
runtime:
|
|
43
|
+
|
|
44
|
+
- every value of `aliases` is a member of `locales`;
|
|
45
|
+
- no key of `aliases` is a member of `locales` (an alias can never shadow a real locale, and
|
|
46
|
+
a registry that tries to is a mistake, not a precedence question);
|
|
47
|
+
- `aliases` is not transitive — a value is a final answer, never another alias key. The
|
|
48
|
+
invariant above guarantees this.
|
|
49
|
+
|
|
50
|
+
If any invariant is violated at runtime, raise `POLYTYPO_MALFORMED_LOCALE_DATA`.
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## 3. Algorithm
|
|
55
|
+
|
|
56
|
+
The input is the caller's `locale` string. Convert it once to a code-point array
|
|
57
|
+
`t[0 … m-1]` (ARCHITECTURE.md §4.2).
|
|
58
|
+
|
|
59
|
+
### 3.1 Step 1 — presence
|
|
60
|
+
|
|
61
|
+
If `locale` is absent, null, or `m = 0`, throw `POLYTYPO_UNKNOWN_LOCALE`. There is no default
|
|
62
|
+
locale (PLAN.md §5.1: `locale` is required, no default).
|
|
63
|
+
|
|
64
|
+
### 3.2 Step 2 — canonicalisation (ASCII only, host-independent)
|
|
65
|
+
|
|
66
|
+
Produce a canonical array `c` of the same length:
|
|
67
|
+
|
|
68
|
+
1. For each index `j`: if `t[j]` is U+005F (`_`), set `c[j]` = U+002D (`-`); otherwise
|
|
69
|
+
`c[j] = t[j]`.
|
|
70
|
+
2. If `m ≥ 2`: for `j` in `{0, 1}`, if `c[j]` is in U+0041–U+005A (`A`–`Z`), add 32 to get
|
|
71
|
+
the ASCII lowercase form.
|
|
72
|
+
3. If `m = 5` and `c[2]` is U+002D: for `j` in `{3, 4}`, if `c[j]` is in U+0061–U+007A
|
|
73
|
+
(`a`–`z`), subtract 32 to get the ASCII uppercase form.
|
|
74
|
+
|
|
75
|
+
This is an explicit arithmetic table over the ASCII range and involves **no** `toLowerCase`,
|
|
76
|
+
no `toUpperCase`, no `strtolower`, no ICU, and no host locale (ARCHITECTURE.md §4.4). A
|
|
77
|
+
Turkish process resolves `EN-us` exactly as a Finnish one does.
|
|
78
|
+
|
|
79
|
+
Nothing outside U+0041–U+005A, U+0061–U+007A, U+002D and U+005F is altered; a tag containing
|
|
80
|
+
anything else will fail step 3.
|
|
81
|
+
|
|
82
|
+
### 3.3 Step 3 — structural validation
|
|
83
|
+
|
|
84
|
+
`c` must have one of exactly two shapes. Test by index, not by pattern:
|
|
85
|
+
|
|
86
|
+
- **Language only**, `m = 2`: `c[0]` and `c[1]` are both in U+0061–U+007A.
|
|
87
|
+
- **Language + region**, `m = 5`: `c[0]`, `c[1]` in U+0061–U+007A; `c[2]` = U+002D;
|
|
88
|
+
`c[3]`, `c[4]` in U+0041–U+005A.
|
|
89
|
+
|
|
90
|
+
Anything else — `m` of 0, 1, 3, 4, 6 or more; a three-letter code (`eng`); a script subtag
|
|
91
|
+
(`sr-Latn`); a numeric region (`es-419`); a private-use tag (`x-pig`); the wildcard `und` —
|
|
92
|
+
throws `POLYTYPO_UNKNOWN_LOCALE`.
|
|
93
|
+
|
|
94
|
+
`locale.schema.json` constrains the `locale` field of a locale file to
|
|
95
|
+
`^[a-z]{2}(-[A-Z]{2})?$`, which is exactly the same two shapes; the index-based test above is
|
|
96
|
+
the regex-free statement of it (ARCHITECTURE.md §4.1).
|
|
97
|
+
|
|
98
|
+
> A malformed tag and an unrecognised tag both raise `POLYTYPO_UNKNOWN_LOCALE`. There is no
|
|
99
|
+
> separate `POLYTYPO_INVALID_LOCALE` in the taxonomy — see §6.1.
|
|
100
|
+
|
|
101
|
+
### 3.4 Step 4 — resolve
|
|
102
|
+
|
|
103
|
+
Let `tag` be the canonical string built from `c`.
|
|
104
|
+
|
|
105
|
+
1. **Exact match.** If `tag` is a member of `registry.locales`, return `tag`.
|
|
106
|
+
2. **Alias.** If `tag` is a key of `registry.aliases`, let `v = aliases[tag]`. If `v` is not a
|
|
107
|
+
member of `registry.locales`, raise `POLYTYPO_MALFORMED_LOCALE_DATA`. Otherwise return `v`.
|
|
108
|
+
3. **Strip the region and retry, once.** If `m = 5`, let `base` be the two-code-point tag
|
|
109
|
+
`c[0] c[1]`.
|
|
110
|
+
a. If `base` is a member of `registry.locales`, return `base`.
|
|
111
|
+
b. If `base` is a key of `registry.aliases`, let `v = aliases[base]`, validate as in
|
|
112
|
+
step 2, and return `v`.
|
|
113
|
+
4. **Throw** `POLYTYPO_UNKNOWN_LOCALE`.
|
|
114
|
+
|
|
115
|
+
The retry happens **once**. There is no third attempt, no loop, and no chain: a two-letter tag
|
|
116
|
+
has no region to strip, and an alias value is a concrete locale by invariant (§2).
|
|
117
|
+
|
|
118
|
+
Exact match is attempted before the alias table so that a registry which (incorrectly) lists
|
|
119
|
+
a tag in both places can never change behaviour based on lookup order — combined with the §2
|
|
120
|
+
invariant forbidding that overlap, the result is that lookup order is unobservable.
|
|
121
|
+
|
|
122
|
+
### 3.5 Purity and lifetime
|
|
123
|
+
|
|
124
|
+
Resolution is a pure function of `(locale, registry.json)`. It is performed **once per
|
|
125
|
+
`transform` call**, before any rule runs, and the resulting locale object is passed down. It
|
|
126
|
+
must not be cached in mutable module state (ARCHITECTURE.md §7: no global configuration, no
|
|
127
|
+
module-level mutable state, reentrant).
|
|
128
|
+
|
|
129
|
+
---
|
|
130
|
+
|
|
131
|
+
## 4. Must not do
|
|
132
|
+
|
|
133
|
+
- **Never fall back to English**, or to any other default, under any circumstance. Not on an
|
|
134
|
+
unknown language, not on a malformed tag, not on an empty string.
|
|
135
|
+
- **Never fall back to a "nearest" language.** `nb` does not resolve to `sv`, `nl` does not
|
|
136
|
+
resolve to `de`, `pt` does not resolve to `fr`.
|
|
137
|
+
- **Never fall back to the first entry of `registry.locales`.**
|
|
138
|
+
- **Never consult the host environment** — `LANG`, `LC_ALL`, `navigator.language`,
|
|
139
|
+
`Intl.DateTimeFormat().resolvedOptions().locale`, `setlocale`, or a browser's
|
|
140
|
+
`Accept-Language`. `transform` is pure and reads no environment (ARCHITECTURE.md §7).
|
|
141
|
+
- **Never use a platform locale-negotiation library**, even when it appears to agree
|
|
142
|
+
(ARCHITECTURE.md §4.7).
|
|
143
|
+
- **Never strip the region more than once**, and never strip a language subtag.
|
|
144
|
+
- **Never follow an alias to another alias.**
|
|
145
|
+
- **Never lowercase or uppercase with a host-locale-sensitive function** (§3.2).
|
|
146
|
+
- **Never return a locale id that has no file.** Step 2/3b validate the alias target.
|
|
147
|
+
|
|
148
|
+
---
|
|
149
|
+
|
|
150
|
+
## 5. Worked table
|
|
151
|
+
|
|
152
|
+
Against `spec/locales/registry.json` as of spec 0.2.0:
|
|
153
|
+
`locales = ["en-US","en-GB","de-DE","de-CH","fr","fr-CA","ru","fi","sv","el"]`,
|
|
154
|
+
`aliases = {"en":"en-US","de":"de-DE"}`.
|
|
155
|
+
|
|
156
|
+
| Input | Canonical | Path | Result |
|
|
157
|
+
| ------------- | --------- | ---------------------------------------------------- | ----------------------------------- |
|
|
158
|
+
| `en-US` | `en-US` | exact | `en-US` |
|
|
159
|
+
| `en-GB` | `en-GB` | exact | `en-GB` |
|
|
160
|
+
| `en` | `en` | alias | `en-US` |
|
|
161
|
+
| `en-AU` | `en-AU` | no exact, no alias, strip → `en`, no exact, alias | `en-US` |
|
|
162
|
+
| `de` | `de` | alias | `de-DE` |
|
|
163
|
+
| `de-DE` | `de-DE` | exact | `de-DE` |
|
|
164
|
+
| `de-CH` | `de-CH` | exact | `de-CH` |
|
|
165
|
+
| `de-AT` | `de-AT` | no exact, no alias, strip → `de`, no exact, alias | `de-DE` |
|
|
166
|
+
| `fr` | `fr` | exact | `fr` |
|
|
167
|
+
| `fr-CA` | `fr-CA` | exact | `fr-CA` |
|
|
168
|
+
| `fr-BE` | `fr-BE` | no exact, no alias, strip → `fr`, exact | `fr` |
|
|
169
|
+
| `fi` | `fi` | exact | `fi` |
|
|
170
|
+
| `fi-FI` | `fi-FI` | strip → `fi`, exact | `fi` |
|
|
171
|
+
| `sv-FI` | `sv-FI` | strip → `sv`, exact | `sv` (see §6.3) |
|
|
172
|
+
| `ru-BY` | `ru-BY` | strip → `ru`, exact | `ru` |
|
|
173
|
+
| `el` | `el` | exact | `el` |
|
|
174
|
+
| `el-GR` | `el-GR` | strip → `el`, exact | `el` |
|
|
175
|
+
| `EN-us` | `en-US` | canonicalise, exact | `en-US` |
|
|
176
|
+
| `en_US` | `en-US` | canonicalise, exact | `en-US` |
|
|
177
|
+
| `DE` | `de` | canonicalise, alias | `de-DE` |
|
|
178
|
+
| `xx-YY` | `xx-YY` | no exact, no alias, strip → `xx`, no exact, no alias | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
179
|
+
| `xx` | `xx` | no exact, no alias, nothing to strip | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
180
|
+
| `nb-NO` | `nb-NO` | strip → `nb`, miss | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
181
|
+
| `eng` | — | fails §3.3 (length 3) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
182
|
+
| `sr-Latn` | — | fails §3.3 (length 7) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
183
|
+
| `es-419` | — | fails §3.3 (region not `A`–`Z`) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
184
|
+
| `und` | — | fails §3.3 (length 3) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
185
|
+
| `` (empty) | — | fails §3.1 | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
186
|
+
| absent / null | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
187
|
+
| options object absent | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
188
|
+
| non-string tag (number, object, boolean) | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
|
|
189
|
+
|
|
190
|
+
The last three rows carry `tagAbsent: true`, a closed flag in `resolution.schema.json` added so
|
|
191
|
+
that "no tag was supplied at all" is expressible as a fixture rather than only as a JS test.
|
|
192
|
+
`transform("a", {})`, `{locale: null}`, a missing options object and every non-string locale now
|
|
193
|
+
raise the coded error rather than a native `TypeError` — which matters because a `TypeError` is
|
|
194
|
+
not in the taxonomy (ARCHITECTURE.md §4.6) and would differ in each of the five runtimes.
|
|
195
|
+
|
|
196
|
+
Every row is a fixture, not a candidate: **35 resolution cases run today**, covering these rows
|
|
197
|
+
and the malformed-tag rejections. `fixtures.schema.json` was extended with a case shape carrying
|
|
198
|
+
no `mode` and no `out`, whose expected result is a locale id or a thrown code. An earlier
|
|
199
|
+
revision of this paragraph said such rows "cannot be expressed in the existing fixture format";
|
|
200
|
+
that was true when written, and is why the schema was changed.
|
|
201
|
+
|
|
202
|
+
---
|
|
203
|
+
|
|
204
|
+
## 6. Open questions
|
|
205
|
+
|
|
206
|
+
1. **The error taxonomy has no code for a malformed tag.** ARCHITECTURE.md §4.6 lists
|
|
207
|
+
`POLYTYPO_UNKNOWN_LOCALE`, `POLYTYPO_INVALID_MODE` and `POLYTYPO_MALFORMED_LOCALE_DATA`.
|
|
208
|
+
`eng`, `es-419` and `""` are structurally invalid rather than merely unknown, and a caller
|
|
209
|
+
would benefit from telling the two apart (one is a typo in their code, the other is a
|
|
210
|
+
language polytypo does not support yet). I have mapped both to
|
|
211
|
+
`POLYTYPO_UNKNOWN_LOCALE` because inventing a code would change a documented contract.
|
|
212
|
+
Recommend adding `POLYTYPO_INVALID_LOCALE`; operator decision.
|
|
213
|
+
2. *(Closed.)* Resolution **is** fixture-covered: 35 resolution cases run today. The gap this
|
|
214
|
+
item reported — that `fixtures.schema.json` could not express a case with no `mode`, no `out`
|
|
215
|
+
and an expected locale id or thrown code — was closed by extending the schema, and the §5
|
|
216
|
+
table's rows are those cases. (The item also miscounted the pipeline as seven rules; it is
|
|
217
|
+
eight, since `hyphen` was added at order 35.)
|
|
218
|
+
|
|
219
|
+
3. **`sv-FI` (Finland Swedish) silently resolves to `sv`.** The two genuinely differ in some
|
|
220
|
+
conventions, and PLAN.md §7 already flags Swedish quote practice as uncertain. The
|
|
221
|
+
algorithm is correct as specified; the question is whether `sv-FI` should be a locale of
|
|
222
|
+
its own. Not resolvable here — it needs a normative citation (PLAN.md §6.1).
|
|
223
|
+
4. **Canonicalisation accepts `_` as a separator and is case-insensitive.** That is leniency
|
|
224
|
+
in a public API, and PLAN.md §5.1's stance is fail-fast. The arguments for it: Java, Ruby
|
|
225
|
+
and PHP ecosystems routinely produce `en_US`; the mapping is unambiguous; and rejecting it
|
|
226
|
+
converts a trivially-correctable input into a production exception. The argument against:
|
|
227
|
+
every accepted spelling is a spelling someone will depend on. I specified leniency —
|
|
228
|
+
**this is the one behavioural choice in this document that I would want the operator to
|
|
229
|
+
confirm before it becomes contract.**
|
|
230
|
+
5. *(Closed.)* `registry.json` **has** a JSON Schema, compiled and enforced, so the §2
|
|
231
|
+
invariants are machine-checked rather than deferred to a runtime
|
|
232
|
+
`POLYTYPO_MALFORMED_LOCALE_DATA`.
|
|
233
|
+
6. *(Closed.)* The registry and the files in `spec/locales/` **are** cross-checked in both
|
|
234
|
+
directions by `scripts/validate-spec.mjs`: an entry with no file, and a file with no entry,
|
|
235
|
+
both fail the build. This item asked for "a two-line CI check"; it exists.
|
|
236
|
+
|
|
237
|
+
7. **Region-only variation is the only variation supported.** There is no way to express
|
|
238
|
+
`de-CH-1901` or a house-style variant. Deliberate for v1 and consistent with the schema
|
|
239
|
+
pattern, recorded so it is a decision.
|