polytypo 1.2.0 → 1.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (64) hide show
  1. checksums.yaml +4 -4
  2. data/README.md +33 -1
  3. data/lib/polytypo/data/VERSION +1 -1
  4. data/lib/polytypo/data/fixtures/cs.json +161 -0
  5. data/lib/polytypo/data/fixtures/de-CH.json +1 -1
  6. data/lib/polytypo/data/fixtures/de-DE.json +195 -6
  7. data/lib/polytypo/data/fixtures/el.json +1 -1
  8. data/lib/polytypo/data/fixtures/en-GB.json +12 -1
  9. data/lib/polytypo/data/fixtures/en-US.json +626 -1
  10. data/lib/polytypo/data/fixtures/es.json +193 -0
  11. data/lib/polytypo/data/fixtures/fi.json +1 -1
  12. data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
  13. data/lib/polytypo/data/fixtures/fr.json +176 -1
  14. data/lib/polytypo/data/fixtures/it.json +161 -0
  15. data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
  16. data/lib/polytypo/data/fixtures/nl.json +121 -0
  17. data/lib/polytypo/data/fixtures/pl.json +137 -0
  18. data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
  19. data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
  20. data/lib/polytypo/data/fixtures/ru.json +23 -1
  21. data/lib/polytypo/data/fixtures/sv.json +1 -1
  22. data/lib/polytypo/data/fixtures/uk.json +153 -0
  23. data/lib/polytypo/data/locales/cs.json +90 -0
  24. data/lib/polytypo/data/locales/de-DE.json +7 -2
  25. data/lib/polytypo/data/locales/en-US.json +3 -3
  26. data/lib/polytypo/data/locales/es.json +111 -0
  27. data/lib/polytypo/data/locales/fr-CA.json +7 -1
  28. data/lib/polytypo/data/locales/fr.json +7 -1
  29. data/lib/polytypo/data/locales/it.json +95 -0
  30. data/lib/polytypo/data/locales/nl.json +84 -0
  31. data/lib/polytypo/data/locales/pl.json +96 -0
  32. data/lib/polytypo/data/locales/pt-BR.json +82 -0
  33. data/lib/polytypo/data/locales/pt-PT.json +84 -0
  34. data/lib/polytypo/data/locales/registry.json +23 -3
  35. data/lib/polytypo/data/locales/ru.json +2 -2
  36. data/lib/polytypo/data/locales/uk.json +130 -0
  37. data/lib/polytypo/data/rules/analyze.md +157 -0
  38. data/lib/polytypo/data/rules/apostrophe.md +432 -0
  39. data/lib/polytypo/data/rules/dashes.md +128 -37
  40. data/lib/polytypo/data/rules/ellipsis.md +271 -0
  41. data/lib/polytypo/data/rules/hyphen.md +353 -0
  42. data/lib/polytypo/data/rules/locale-resolution.md +239 -0
  43. data/lib/polytypo/data/rules/modes.md +1281 -0
  44. data/lib/polytypo/data/rules/nbsp.md +1157 -0
  45. data/lib/polytypo/data/rules/order.json +11 -11
  46. data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
  47. data/lib/polytypo/data/rules/quotes.md +1324 -0
  48. data/lib/polytypo/data/rules/ranges.md +489 -0
  49. data/lib/polytypo/data/rules/spaces.md +649 -0
  50. data/lib/polytypo/data/rules/symbols.md +540 -0
  51. data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
  52. data/lib/polytypo/engine/origin.rb +75 -0
  53. data/lib/polytypo/engine/pipeline.rb +72 -1
  54. data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
  55. data/lib/polytypo/engine/rules/dashes.rb +4 -1
  56. data/lib/polytypo/engine/rules/nbsp.rb +43 -7
  57. data/lib/polytypo/engine/rules/ranges.rb +24 -20
  58. data/lib/polytypo/errors.rb +3 -0
  59. data/lib/polytypo/modes/runner.rb +17 -0
  60. data/lib/polytypo/modes/spans.rb +30 -2
  61. data/lib/polytypo/modes/yaml.rb +312 -0
  62. data/lib/polytypo/version.rb +1 -1
  63. data/lib/polytypo.rb +126 -15
  64. metadata +31 -1
@@ -0,0 +1,353 @@
1
+ # Rule: `hyphen`
2
+
3
+ **Order:** 35 (between `dashes` and `quotes`). **Default:** on.
4
+ **Modes:** text, html, markdown, yaml.
5
+ **Spec version:** 0.1.0.
6
+
7
+ ---
8
+
9
+ ## 1. Purpose
10
+
11
+ `hyphen` replaces U+002D (hyphen-minus) with U+2011 (non-breaking hyphen) inside a closed list
12
+ of morphological forms whose hyphen must not be broken across a line. It exists for Russian,
13
+ where `из-под`, `кое-что` and `сделал-таки` are single words that a line-breaking algorithm
14
+ will happily split at the hyphen, producing a line ending in `из-` — which every Russian style
15
+ guide forbids. It is deliberately the most data-bound rule in the pipeline: it does exactly
16
+ nothing except where a locale has listed a literal form, so for `en`, `de`, `fr`, `fi`, `sv`
17
+ and `el` — whose lists are empty — it is a provable total no-op, not merely an unlikely one. It
18
+ changes one code point to another code point of the same width and semantics; it never
19
+ inserts, never deletes, never touches spacing, and never converts anything that is not
20
+ already a hyphen inside a word.
21
+
22
+ ---
23
+
24
+ ## 2. Locale data consumed
25
+
26
+ - `hyphen.prefixes` — word-initial forms written **with a trailing hyphen**, e.g. `"кое-"`
27
+ - `hyphen.suffixes` — word-final forms written **with a leading hyphen**, e.g. `"-таки"`
28
+ - `hyphen.compounds` — whole hyphenated forms written with U+002D, e.g. `"из-под"`
29
+
30
+ All three are required by `locale.schema.json` and all three may be empty.
31
+
32
+ **Evidence for these lists.** A locale attests **membership** — that these tokens belong in this list — and this rule owns the **mechanism** it applies to them. A `sources` citation is not required to name a code point or a binding the locale file has no way to vary. The principle is stated once, normatively, in [nbsp.md](nbsp.md) §2.1 and governs every list-valued field in `locale.schema.json`. Entries are
33
+ literal lowercase forms of at least two code points. An entry containing no U+002D is
34
+ meaningless and must raise `POLYTYPO_MALFORMED_LOCALE_DATA` — the rule has nothing to convert
35
+ in it.
36
+
37
+ **If all three arrays are empty, the rule emits nothing for any input.** An implementation may
38
+ short-circuit on that condition; the observable behaviour is identical either way.
39
+
40
+ ---
41
+
42
+ ## 3. Algorithm
43
+
44
+ Input is a code-point array `cp[0 … n-1]`.
45
+
46
+ ### 3.1 Character classes
47
+
48
+ | Class | Members |
49
+ | ----------- | -------------------------------------------------------------------------------------------------------------- |
50
+ | `HY` | U+002D (hyphen-minus) |
51
+ | `NBHY` | U+2011 (non-breaking hyphen) |
52
+ | `HYPHENISH` | `HY` ∪ `NBHY` |
53
+ | `LETTER` | general category `Lu`, `Ll`, `Lt`, `Lm`, `Lo`, `Mn`, `Mc`, `Me` (combining marks count as letter-continuation) |
54
+ | `UPPER` | general category `Lu` or `Lt`. Used only by §3.3's first-character leniency, via the simple uppercase mapping — no guard tests it directly |
55
+ | `DIGIT` | U+0030–U+0039 |
56
+ | `ALNUM` | `LETTER` ∪ `DIGIT` |
57
+ | `WORDISH` | `ALNUM` ∪ `HYPHENISH` — the characters that continue a hyphenated word |
58
+ | `NONE` | index out of range |
59
+
60
+ No other class is examined.
61
+
62
+ **Unicode version.** The general categories and case mappings this rule reads are those of the UCD version pinned in `spec/UNICODE` (`17.0`). The pin is normative for the **derived tables**, not for the host runtime — see [pipeline-idempotency.md](pipeline-idempotency.md) §6a, which also specifies the canary fixtures that make the pin detectable. In particular this rule never looks at a space, and never at
63
+ U+2010, U+00AD, U+2013 or U+2014.
64
+
65
+ ### 3.2 Why order 35, and what it may assume
66
+
67
+ `hyphen` runs **after `dashes` and before `quotes`**. Both halves of that placement are
68
+ load-bearing.
69
+
70
+ **After `dashes`.** `dashes` guard P1 declines every bare U+002D that is not spaced on both
71
+ sides, so every intra-word hyphen — which is exactly this rule's entire subject matter —
72
+ survives `dashes` untouched. Running `hyphen` first would invert the dependency: `dashes`
73
+ would then encounter U+2011 in positions where it expects U+002D, and while its `INERT-DASH`
74
+ class already covers U+2011, its verdicts would be reached for a different reason than the
75
+ one documented, and the two rules' guards would have to be kept in sync by hand. Running
76
+ second means `hyphen` may assume, and an implementation may assert, that **every U+002D it
77
+ converts has a letter or digit on both sides** — the complement of what `dashes` claims. That
78
+ assumption is what makes §5 short.
79
+
80
+ **Before `quotes`.** This ordering is materially free — the two rules share no characters —
81
+ but it is not entirely free, and the reason is worth recording. `quotes.md` §3.1 lists
82
+ U+002D, U+2013 and U+2014 as acceptable left context for an opening quotation mark
83
+ (`canOpen`). U+2011 was not in that list. If `hyphen` converted a U+002D that sat immediately
84
+ left of a quote candidate, that candidate's `canOpen` would flip from true to false between
85
+ one pipeline run and the next — the exact class of cross-rule drift that the idempotency
86
+ invariant exists to catch. The shape is unreachable in practice (a `prefixes` entry requires a
87
+ `LETTER` after its hyphen, so a quotation mark cannot follow), but relying on unreachability
88
+ across two documents is how ports diverge. **`quotes.md`, `apostrophe.md` and `nbsp.md` have
89
+ therefore been amended to include U+2011 wherever they already list U+002D/U+2013/U+2014 as
90
+ context.** With that amendment, order 35 versus order 45 is genuinely immaterial, and 35 is
91
+ kept because it groups the two hyphen-and-dash rules together.
92
+
93
+ ### 3.3 Matching
94
+
95
+ All three lists use the same comparison, called **hyphen-lenient, first-character-lenient
96
+ literal matching**. For a pattern `w` of `k` code points at candidate index `a`:
97
+
98
+ - **hyphen leniency, at every position `j` in `0 … k-1`:** if `w[j]` is in `HY`, then
99
+ `cp[a+j]` may be in `HY` **or** in `NBHY`; otherwise `cp[a+j] = w[j]` exactly.
100
+ The leniency must cover `j = 0` and not merely `j ≥ 1`, because a **suffix** entry _begins_
101
+ with its hyphen: `-таки` converted once becomes `‑таки`, and a pattern that demanded U+002D
102
+ at position 0 would stop matching its own output. §5's round-trip argument depends on this.
103
+ - **first-character case leniency, at `j = 0` only:** either `cp[a] = w[0]`, or `w[0]` is in
104
+ `LETTER` and `cp[a]` is the Unicode **simple uppercase mapping** of `w[0]`. The two
105
+ leniencies never overlap — `w[0]` is a hyphen or a letter, never both.
106
+
107
+ Nothing else matches. `ИЗ-ПОД` does not match `из-под`; `Из-под` does.
108
+
109
+ **Capitalisation without host-locale case folding.** The first-character leniency uses the
110
+ Unicode simple uppercase mapping — a plain code-point→code-point table from the UCD, applied
111
+ to the _pattern_ (which the schema fixes as lowercase), not to the input. It is not
112
+ `toUpperCase()`, not `toLocaleUpperCase()`, not ICU, and it does not depend on the host
113
+ process's locale (ARCHITECTURE.md §4.4). Mapping the pattern rather than the input is
114
+ deliberate: it means at most one extra code point is computed per list entry, it can be
115
+ precomputed once when the locale file is loaded, and the notorious Turkish pair never arises
116
+ because the simple mapping of `i` is `I` unconditionally and no v1 `hyphen` entry begins with
117
+ `i` in any case. This is the same convention `nbsp.afterShortWords` uses (`nbsp.md` §3.5), and
118
+ having one convention across the spec is worth more than the marginal extra coverage a second
119
+ one would buy.
120
+
121
+ All-capitals forms (`ИЗ-ПОД` in a heading) are deliberately **not** matched — see §7.1.
122
+
123
+ ### 3.4 Scan
124
+
125
+ Walk `a` from `0` to `n-1`. At each index, select **one** candidate entry: the longest entry
126
+ that matches at `a` across all three lists, with ties broken in the order **compounds,
127
+ prefixes, suffixes**. Entries are unique (`uniqueItems`), so this is a total order and the
128
+ selection is deterministic.
129
+
130
+ **Matching and guarding are separate steps, and there is no backtracking.** The selected entry
131
+ is then checked against its guards. If it passes, it claims the position and the scan continues
132
+ from `a + k`. **If it fails a guard, the sub-rule emits nothing at `a` and no other entry is
133
+ tried there**; the scan advances `a` by one. An earlier revision said "the first entry that
134
+ matches _and passes its guards_", which is backtracking, and it disagreed with `nbsp`'s
135
+ list-driven sub-rules. `nbsp.md` §3.5 now states the same policy, and the two must stay
136
+ aligned: longest match wins, guards are applied to the winner only.
137
+
138
+ Compounds are tried first because a compound is the most specific form: if a locale lists both
139
+ the compound `из-под` and the prefix `из-`, the compound's verdict (bind the hyphen inside the
140
+ whole word) must win, and it happens to produce the same edit — but the scan position it
141
+ consumes differs, and that must be deterministic.
142
+
143
+ **C — compounds.** For an entry `w` of `k` code points matching at `a`:
144
+
145
+ 1. **Left boundary.** `cp[a-1]` must be `NONE` or **not** in `WORDISH`.
146
+ 2. **Right boundary.** `cp[a+k]` must be `NONE` or **not** in `WORDISH`.
147
+ Together these mean the compound is a whole word: `из-под` matches in `из-под стола` but
148
+ not inside `квазииз-подный` or `из-под-` .
149
+ 3. For every position `j` with `w[j]` in `HY`: if `cp[a+j]` is already in `NBHY`, emit nothing
150
+ for that position; otherwise emit an edit replacing `cp[a+j]` with U+2011.
151
+
152
+ **P — prefixes.** An entry `w` ends with U+002D; let `k` be its length.
153
+
154
+ 1. **Left boundary.** `cp[a-1]` must be `NONE` or not in `WORDISH` — the prefix starts a word.
155
+ 2. **Right boundary.** `cp[a+k]` must be in `LETTER`. A prefix must actually prefix something:
156
+ `кое-что` matches, `кое-` at the end of a line or before a space does not, and neither does
157
+ `кое-2`.
158
+ 3. Emit an edit replacing the final code point of the match (`cp[a+k-1]`, the hyphen) with
159
+ U+2011, unless it is already in `NBHY`.
160
+
161
+ **S — suffixes.** An entry `w` begins with U+002D; let `k` be its length. Because a suffix is
162
+ matched at its own start index, the scan finds it at the hyphen, not at the start of the word.
163
+
164
+ 1. **Left boundary.** `cp[a-1]` must be in `LETTER`. A suffix must actually suffix something:
165
+ `сделал-таки` matches, a line beginning `-таки` does not.
166
+ 2. **Right boundary.** `cp[a+k]` must be `NONE` or not in `WORDISH`.
167
+ 3. Emit an edit replacing the first code point of the match (`cp[a]`, the hyphen) with U+2011,
168
+ unless it is already in `NBHY`.
169
+
170
+ Note that the first-character leniency of §3.3 is inert for suffixes (their first code point
171
+ is a hyphen, not a letter) and for compounds it applies to the word's initial letter, which is
172
+ the sentence-start case (`Из-под стола…`).
173
+
174
+ ---
175
+
176
+ ## 4. Must not touch
177
+
178
+ **Scope.** Per [pipeline-idempotency.md](pipeline-idempotency.md) §5.2 each bullet is **[P]** —
179
+ a guarantee of `transform` as a whole — or **[R]** — true of this rule in isolation but capable
180
+ of being falsified by another rule, which is then named.
181
+
182
+ - **[P] Any hyphen not covered by a listed form.** `научно-технический`, `Jean-Luc`, `e-mail`,
183
+ `well-known`, `COVID-19` are all left as U+002D. This rule has no morphology of its own; it
184
+ has a list.
185
+ - **[P] Every locale with empty lists.** `en`, `de`, `fr`, `fi`, `sv`, `el` — the rule is a total no-op
186
+ and must produce byte-identical output (§2).
187
+ - **[R] A hyphen that `dashes` owns.** Anything spaced on both sides was already handled at order
188
+ 30; by §3.2 this rule only ever sees intra-word hyphens, and its boundary guards enforce
189
+ that independently rather than trusting it.
190
+ - **[P] U+2010 (hyphen), U+00AD (soft hyphen), U+2012, U+2013, U+2014, U+2212.** Not in
191
+ `HYPHENISH`; never read, never written. A soft hyphen inside `из-под` is the author's
192
+ deliberate break hint and is not this rule's business.
193
+ - **[R] Spacing of any kind.** No space is inserted, removed or converted. Every edit is one code
194
+ point replacing one code point at the same index.
195
+ - **[P] All-capitals forms** — `ИЗ-ПОД` (§7.1).
196
+ - **[P] A listed form appearing inside a longer word.** Boundary guards C1/C2, P1, S2.
197
+ - **[P] Anything inside a skipped region.** A hyphen in a URL, a code span, a fenced block or an
198
+ HTML attribute is removed by the mode adapter (L2) before this rule runs. `из-под` in a
199
+ slug is therefore safe; this rule has no URL-awareness and must not acquire any.
200
+
201
+ ---
202
+
203
+ ## 5. Idempotency argument
204
+
205
+ Let `T` be the rule.
206
+
207
+ **The output form is recognised as already-correct.** Every edit replaces a U+002D with a
208
+ U+2011 at the same index. §3.3's comparison is **hyphen-lenient**: a pattern's U+002D matches
209
+ either U+002D or U+2011 in the text. So on the second run the same list entry matches at the
210
+ same index `a` with the same length `k` — the match is not lost by the conversion, which is
211
+ the property the leniency exists for. The emission step of each branch then finds `cp[a+j]`
212
+ (or `cp[a+k-1]`, or `cp[a]`) already in `NBHY` and emits nothing.
213
+
214
+ **No guard's verdict changes.** The guards test:
215
+
216
+ - membership in `LETTER`, `UPPER`, `DIGIT` — no edit ever writes a character in any of these
217
+ classes, and no edit ever removes one, since every edit is a 1:1 replacement of a hyphen;
218
+ - membership in `WORDISH` — which contains **both** U+002D and U+2011 by construction
219
+ (`HYPHENISH ⊂ WORDISH`). This is the second place the design has to be deliberate: if
220
+ `WORDISH` contained only U+002D, then converting a hyphen would change a neighbouring
221
+ form's boundary test from "is in `WORDISH`, reject" to "is not in `WORDISH`, accept", and a
222
+ form adjacent to a converted one could start matching on the second run. With `NBHY` in
223
+ `WORDISH`, the boundary verdicts are invariant.
224
+
225
+ **Scan positions are stable.** The scan advances by `a + k` after a claim, and `k` is the
226
+ pattern length, which does not depend on the text. Since every edit is 1:1, no index shifts.
227
+ So the same entries claim the same positions in the same order on every run.
228
+
229
+ Hence `T(T(x)) = T(x)`. ∎
230
+
231
+ **Input that already contains U+2011.** This is the same case as "the rule's own output", and
232
+ it is handled identically: a form written by the author as `из‑под` with a real U+2011 matches
233
+ (hyphen-lenient), passes its guards, and produces **no edit**. A U+2011 outside any listed
234
+ form is never examined at all. So text that is already correctly bound round-trips
235
+ byte-identically, which is the round-trip requirement of ARCHITECTURE.md §6.3 item 4 and the
236
+ reason the leniency is specified rather than left to an implementation.
237
+
238
+ **Stability under the rest of the pipeline.** `dashes` runs before this rule within a single
239
+ pipeline pass, but on the _next_ pass it sees the U+2011 this rule produced. Its verdicts do
240
+ not change: `dashes` §3.1 puts U+2011 in `INERT-DASH`, its §3.2 step 6 rejects a token whose
241
+ neighbour is in `INERT-DASH` **or** in `DASH` alike, and its cluster alphabet contains both —
242
+ so a position that was declined when it held U+002D is declined when it holds U+2011, for the
243
+ same reason. `quotes`, `apostrophe` and `nbsp` list U+2011 alongside U+002D/U+2013/U+2014 in
244
+ their context classes (§3.2), so their classifications are likewise unaffected.
245
+
246
+ **What had to be fixed.** The naive formulation — _"replace the hyphen in each listed form
247
+ with U+2011"_ — is not idempotent, because on the second run the form no longer matches: the
248
+ pattern holds U+002D and the text now holds U+2011, so the rule silently stops recognising its
249
+ own output. That is harmless for the output itself but it is a latent trap, because it means
250
+ the rule's behaviour depends on whether it has run before, and any later guard that consults a
251
+ list membership would then disagree between runs. Making the comparison hyphen-lenient, and
252
+ putting U+2011 into `WORDISH`, are the two changes that make the rule's view of the text
253
+ identical before and after its own edits.
254
+
255
+ ---
256
+
257
+ ### Composition obligation
258
+
259
+ Per [pipeline-idempotency.md](pipeline-idempotency.md) §5. This rule is **R₄**, so the
260
+ obligation runs against `spaces` (R₁), `ellipsis` (R₂) and `dashes` (R₃).
261
+
262
+ **What this rule emits.** U+2011, one code point replacing one U+002D at the same index, inside a
263
+ matched form. It inserts nothing, deletes nothing, and changes no length.
264
+
265
+ **Against `I₁` (`spaces`).** Discharged: no U+0020 is emitted or deleted, and no code point is
266
+ inserted or removed, so no space run can change length or neighbourhood.
267
+
268
+ **Against `I₂` (`ellipsis`).** Discharged: U+2011 is not in `DOTLIKE` and no dot run is
269
+ touched.
270
+
271
+ **Against `I₃` (`dashes`).** This is the only one with content. Converting U+002D to U+2011
272
+ changes a code point that `dashes` classifies — but it moves it from `DASH` to `INERT-DASH`,
273
+ and `dashes` treats the two identically in every place a _neighbouring_ token could read them:
274
+ its isolation guard (§3.2 step 6) rejects both, guard G2 rejects both, and its cluster alphabet
275
+ contains both. The converted hyphen itself was never a `dashes` token: `dashes` guard P1
276
+ declines a bare U+002D that is not spaced on both sides, and this rule only ever claims a
277
+ hyphen with a letter or a listed form's interior on each side. So no verdict changes, and the
278
+ positions this rule edits are exactly the ones `dashes` had already declined.
279
+
280
+ ---
281
+
282
+ ## 6. Worked examples
283
+
284
+ `⟶` = no change, `‑` = U+2011 (non-breaking hyphen), `-` = U+002D.
285
+
286
+ ### `ru` — `compounds: ["из-под","из-за","по-моему"]`, `prefixes: ["кое-"]`, `suffixes: ["-таки","-то","-либо","-нибудь"]`
287
+
288
+ | # | Input | Output | Why |
289
+ | --- | ----------------------------- | --------------------------- | -------------------------------------------------------------------------------- |
290
+ | 1 | `Достал из-под стола` | `Достал из‑под стола` | compound, both boundaries non-`WORDISH` |
291
+ | 2 | `Из-под стола донёсся звук` | `Из‑под стола донёсся звук` | first-character leniency: simple uppercase mapping of `и` |
292
+ | 3 | `ИЗ-ПОД СТОЛА` | ⟶ | all-capitals is not matched (§7.1) |
293
+ | 4 | `кое-что и кое-как` | `кое‑что и кое‑как` | prefix; `cp[a+k]` is a letter in both |
294
+ | 5 | `Он сделал-таки это` | `Он сделал‑таки это` | suffix; `cp[a-1]` is a letter |
295
+ | 6 | `что-нибудь или что-либо` | `что‑нибудь или что‑либо` | two suffixes |
296
+ | 7 | `Достал из‑под стола` | ⟶ | already U+2011 — hyphen-lenient match, no edit. **The round-trip case** |
297
+ | 8 | `научно-технический прогресс` | ⟶ | not in any list |
298
+ | 9 | `кое- и кое-что` | `кое- и кое‑что` | the first `кое-` fails P2 (`cp[a+k]` is a space, not a letter); the second binds |
299
+ | 10 | `квазииз-подный` | ⟶ | compound guard C1: `cp[a-1]` is `и`, in `WORDISH` |
300
+ | 11 | `Москва — из-за дождя` | `Москва — из‑за дождя` | the em dash is `dashes`' business and is untouched here; the compound binds |
301
+ | 12 | `-таки в начале строки` | ⟶ | suffix guard S1: `cp[a-1]` is `NONE` |
302
+
303
+ ### `en`, `de`, `fr`, `fi`, `sv`, `el` — all three lists empty
304
+
305
+ | # | Input | Output | Why |
306
+ | --- | ------------------------------------------------- | ------ | ---------------------------- |
307
+ | 13 | `A well-known e-mail address, COVID-19, Jean-Luc` | ⟶ | no entries; total no-op (§2) |
308
+ | 14 | `Any text at all` | ⟶ | same |
309
+
310
+ Cases 3, 7, 8, 10, 12, 13 and 14 are "no change" cases.
311
+
312
+ ---
313
+
314
+ ## 7. Open questions
315
+
316
+ 1. **All-capitals forms are not matched.** `ИЗ-ПОД СТОЛА` in a Russian heading keeps its
317
+ breakable hyphen. Matching it needs either a full uppercase comparison of every code point
318
+ (which requires an explicit case-mapping table in the spec, not just the single first
319
+ character) or a second set of list entries. Russian headings are frequently set in caps, so
320
+ this is a real miss, not a theoretical one. I did not add a full case table because it is a
321
+ cross-cutting decision — `nbsp.afterShortWords` has exactly the same limitation — and
322
+ because ARCHITECTURE.md §4.4 wants any such table to be explicit spec data. Recommend
323
+ deciding it once, for both rules, together with the Unicode-version pinning question.
324
+ 2. **Title Case within a form is not matched.** `Из-Под` (a stylised heading) does not match.
325
+ Same root cause as §7.1.
326
+ 3. **`hyphen.prefixes` and `hyphen.suffixes` cannot express a constraint on what follows or
327
+ precedes** beyond "a letter". Russian `-то` is a particle in `кто-то` but a conjunction in
328
+ other positions, and `по-` is a prefix in `по-моему` but a preposition elsewhere. The lists
329
+ are blunt instruments; the mitigation is to list only forms whose binding is
330
+ unconditional, and `по-` is deliberately shown in §6 as a **compound** (`по-моему`) rather
331
+ than a prefix for that reason. Whether the Russian locale file should list `по-` as a
332
+ prefix at all is a locale-research question with a normative source (Мильчин), not a
333
+ question about this rule.
334
+ 4. **U+2011 has real-world rendering gaps.** A few older fonts lack the glyph and fall back
335
+ visibly. The alternatives are U+2010 plus U+2060 (word joiner), or a zero-width no-break
336
+ space after the hyphen — both worse in other ways, and both invisible in a diff. U+2011 is
337
+ the operator's stated choice and is recorded here as such, but if the M4 dogfooding gate
338
+ shows a rendering problem in the author's own font stack, this is the decision to revisit.
339
+ 5. **Nothing prevents a locale from listing an entry with no hyphen in it.** §2 requires
340
+ `POLYTYPO_MALFORMED_LOCALE_DATA`, but `locale.schema.json`'s `wordList` only constrains
341
+ length and uniqueness, so the check is a runtime one rather than a CI one. A
342
+ `"pattern"`-free way to express it in JSON Schema does not exist, so this probably belongs
343
+ in `scripts/validate-spec.mjs` alongside the one-code-point check. **Reported.**
344
+ 6. **The interaction with hyphenation is untested.** A CSS `hyphens: auto` renderer may still
345
+ break inside `из‑под` at a syllable boundary; U+2011 only forbids the break _at the hyphen_.
346
+ That is out of scope (PLAN.md §4 excludes hyphenation) but worth knowing before claiming
347
+ the rule "prevents the word from breaking".
348
+ 7. **Ordinal repair is refused** — `10-ый` → `10-й`, and the whole family of Russian
349
+ ordinal-ending corrections Lebedev's service performs. This rule binds a hyphen that is
350
+ already correct; it does not rewrite the morphology around one. Changing `ый` to `й` is a
351
+ spelling correction, and PLAN.md §4 refuses spellcheck and typo correction without an
352
+ explicit operator decision. Recorded so the presence of `hyphen.suffixes` is not read as an
353
+ invitation.
@@ -0,0 +1,239 @@
1
+ # Locale resolution
2
+
3
+ **Not a pipeline rule.** This document has no entry in `spec/rules/order.json` and produces no
4
+ edits. It specifies the function that turns the caller's `locale` option into the identifier
5
+ of exactly one locale file, and it lives in `spec/rules/` because it is normative,
6
+ runtime-independent, and fixture-covered like everything else here.
7
+ **Spec version:** 0.1.0.
8
+
9
+ ---
10
+
11
+ ## 1. Purpose
12
+
13
+ `transform(input, { locale })` takes a locale tag from the caller and must select exactly one
14
+ `spec/locales/<id>.json` file. That selection has to be identical in five runtimes, which is
15
+ precisely why it cannot be delegated to a platform locale-negotiation library: ICU's
16
+ `ResourceBundle`, Go's `golang.org/x/text/language` matcher, PHP's `Locale::lookup`, Python's
17
+ `babel.negotiate_locale` and the JS `Intl.Locale` machinery all implement _different_
18
+ fallback policies, several of them probabilistic in the sense that they will happily return a
19
+ "best effort" match for a language the library has never heard of. polytypo does the opposite:
20
+ it resolves exactly, then by declared alias, then by dropping the region subtag once — and if
21
+ that fails it **throws**. It never falls back to English, never to the first entry in the
22
+ registry, and never to a "closest" language. Silently applying Finnish quotation rules to a
23
+ Portuguese catalogue is the failure mode that destroys trust in a tool like this
24
+ (PLAN.md §5.1), and it is cheaper to make the caller fix a typo than to make them discover
25
+ it in production.
26
+
27
+ ---
28
+
29
+ ## 2. Data consumed
30
+
31
+ `spec/locales/registry.json`, in full:
32
+
33
+ - `locales` — array of concrete locale identifiers, each of which has a file
34
+ `spec/locales/<id>.json`;
35
+ - `aliases` — object mapping a tag to a member of `locales`.
36
+
37
+ Nothing else. In particular the individual locale files are not read during resolution, and
38
+ the algorithm never inspects the filesystem (ARCHITECTURE.md §3.1: locale data is embedded in
39
+ the artifact).
40
+
41
+ **Registry invariants.** These are preconditions on the data, checked in spec CI, not at
42
+ runtime:
43
+
44
+ - every value of `aliases` is a member of `locales`;
45
+ - no key of `aliases` is a member of `locales` (an alias can never shadow a real locale, and
46
+ a registry that tries to is a mistake, not a precedence question);
47
+ - `aliases` is not transitive — a value is a final answer, never another alias key. The
48
+ invariant above guarantees this.
49
+
50
+ If any invariant is violated at runtime, raise `POLYTYPO_MALFORMED_LOCALE_DATA`.
51
+
52
+ ---
53
+
54
+ ## 3. Algorithm
55
+
56
+ The input is the caller's `locale` string. Convert it once to a code-point array
57
+ `t[0 … m-1]` (ARCHITECTURE.md §4.2).
58
+
59
+ ### 3.1 Step 1 — presence
60
+
61
+ If `locale` is absent, null, or `m = 0`, throw `POLYTYPO_UNKNOWN_LOCALE`. There is no default
62
+ locale (PLAN.md §5.1: `locale` is required, no default).
63
+
64
+ ### 3.2 Step 2 — canonicalisation (ASCII only, host-independent)
65
+
66
+ Produce a canonical array `c` of the same length:
67
+
68
+ 1. For each index `j`: if `t[j]` is U+005F (`_`), set `c[j]` = U+002D (`-`); otherwise
69
+ `c[j] = t[j]`.
70
+ 2. If `m ≥ 2`: for `j` in `{0, 1}`, if `c[j]` is in U+0041–U+005A (`A`–`Z`), add 32 to get
71
+ the ASCII lowercase form.
72
+ 3. If `m = 5` and `c[2]` is U+002D: for `j` in `{3, 4}`, if `c[j]` is in U+0061–U+007A
73
+ (`a`–`z`), subtract 32 to get the ASCII uppercase form.
74
+
75
+ This is an explicit arithmetic table over the ASCII range and involves **no** `toLowerCase`,
76
+ no `toUpperCase`, no `strtolower`, no ICU, and no host locale (ARCHITECTURE.md §4.4). A
77
+ Turkish process resolves `EN-us` exactly as a Finnish one does.
78
+
79
+ Nothing outside U+0041–U+005A, U+0061–U+007A, U+002D and U+005F is altered; a tag containing
80
+ anything else will fail step 3.
81
+
82
+ ### 3.3 Step 3 — structural validation
83
+
84
+ `c` must have one of exactly two shapes. Test by index, not by pattern:
85
+
86
+ - **Language only**, `m = 2`: `c[0]` and `c[1]` are both in U+0061–U+007A.
87
+ - **Language + region**, `m = 5`: `c[0]`, `c[1]` in U+0061–U+007A; `c[2]` = U+002D;
88
+ `c[3]`, `c[4]` in U+0041–U+005A.
89
+
90
+ Anything else — `m` of 0, 1, 3, 4, 6 or more; a three-letter code (`eng`); a script subtag
91
+ (`sr-Latn`); a numeric region (`es-419`); a private-use tag (`x-pig`); the wildcard `und` —
92
+ throws `POLYTYPO_UNKNOWN_LOCALE`.
93
+
94
+ `locale.schema.json` constrains the `locale` field of a locale file to
95
+ `^[a-z]{2}(-[A-Z]{2})?$`, which is exactly the same two shapes; the index-based test above is
96
+ the regex-free statement of it (ARCHITECTURE.md §4.1).
97
+
98
+ > A malformed tag and an unrecognised tag both raise `POLYTYPO_UNKNOWN_LOCALE`. There is no
99
+ > separate `POLYTYPO_INVALID_LOCALE` in the taxonomy — see §6.1.
100
+
101
+ ### 3.4 Step 4 — resolve
102
+
103
+ Let `tag` be the canonical string built from `c`.
104
+
105
+ 1. **Exact match.** If `tag` is a member of `registry.locales`, return `tag`.
106
+ 2. **Alias.** If `tag` is a key of `registry.aliases`, let `v = aliases[tag]`. If `v` is not a
107
+ member of `registry.locales`, raise `POLYTYPO_MALFORMED_LOCALE_DATA`. Otherwise return `v`.
108
+ 3. **Strip the region and retry, once.** If `m = 5`, let `base` be the two-code-point tag
109
+ `c[0] c[1]`.
110
+ a. If `base` is a member of `registry.locales`, return `base`.
111
+ b. If `base` is a key of `registry.aliases`, let `v = aliases[base]`, validate as in
112
+ step 2, and return `v`.
113
+ 4. **Throw** `POLYTYPO_UNKNOWN_LOCALE`.
114
+
115
+ The retry happens **once**. There is no third attempt, no loop, and no chain: a two-letter tag
116
+ has no region to strip, and an alias value is a concrete locale by invariant (§2).
117
+
118
+ Exact match is attempted before the alias table so that a registry which (incorrectly) lists
119
+ a tag in both places can never change behaviour based on lookup order — combined with the §2
120
+ invariant forbidding that overlap, the result is that lookup order is unobservable.
121
+
122
+ ### 3.5 Purity and lifetime
123
+
124
+ Resolution is a pure function of `(locale, registry.json)`. It is performed **once per
125
+ `transform` call**, before any rule runs, and the resulting locale object is passed down. It
126
+ must not be cached in mutable module state (ARCHITECTURE.md §7: no global configuration, no
127
+ module-level mutable state, reentrant).
128
+
129
+ ---
130
+
131
+ ## 4. Must not do
132
+
133
+ - **Never fall back to English**, or to any other default, under any circumstance. Not on an
134
+ unknown language, not on a malformed tag, not on an empty string.
135
+ - **Never fall back to a "nearest" language.** `nb` does not resolve to `sv`, `nl` does not
136
+ resolve to `de`, `pt` does not resolve to `fr`.
137
+ - **Never fall back to the first entry of `registry.locales`.**
138
+ - **Never consult the host environment** — `LANG`, `LC_ALL`, `navigator.language`,
139
+ `Intl.DateTimeFormat().resolvedOptions().locale`, `setlocale`, or a browser's
140
+ `Accept-Language`. `transform` is pure and reads no environment (ARCHITECTURE.md §7).
141
+ - **Never use a platform locale-negotiation library**, even when it appears to agree
142
+ (ARCHITECTURE.md §4.7).
143
+ - **Never strip the region more than once**, and never strip a language subtag.
144
+ - **Never follow an alias to another alias.**
145
+ - **Never lowercase or uppercase with a host-locale-sensitive function** (§3.2).
146
+ - **Never return a locale id that has no file.** Step 2/3b validate the alias target.
147
+
148
+ ---
149
+
150
+ ## 5. Worked table
151
+
152
+ Against `spec/locales/registry.json` as of spec 0.2.0:
153
+ `locales = ["en-US","en-GB","de-DE","de-CH","fr","fr-CA","ru","fi","sv","el"]`,
154
+ `aliases = {"en":"en-US","de":"de-DE"}`.
155
+
156
+ | Input | Canonical | Path | Result |
157
+ | ------------- | --------- | ---------------------------------------------------- | ----------------------------------- |
158
+ | `en-US` | `en-US` | exact | `en-US` |
159
+ | `en-GB` | `en-GB` | exact | `en-GB` |
160
+ | `en` | `en` | alias | `en-US` |
161
+ | `en-AU` | `en-AU` | no exact, no alias, strip → `en`, no exact, alias | `en-US` |
162
+ | `de` | `de` | alias | `de-DE` |
163
+ | `de-DE` | `de-DE` | exact | `de-DE` |
164
+ | `de-CH` | `de-CH` | exact | `de-CH` |
165
+ | `de-AT` | `de-AT` | no exact, no alias, strip → `de`, no exact, alias | `de-DE` |
166
+ | `fr` | `fr` | exact | `fr` |
167
+ | `fr-CA` | `fr-CA` | exact | `fr-CA` |
168
+ | `fr-BE` | `fr-BE` | no exact, no alias, strip → `fr`, exact | `fr` |
169
+ | `fi` | `fi` | exact | `fi` |
170
+ | `fi-FI` | `fi-FI` | strip → `fi`, exact | `fi` |
171
+ | `sv-FI` | `sv-FI` | strip → `sv`, exact | `sv` (see §6.3) |
172
+ | `ru-BY` | `ru-BY` | strip → `ru`, exact | `ru` |
173
+ | `el` | `el` | exact | `el` |
174
+ | `el-GR` | `el-GR` | strip → `el`, exact | `el` |
175
+ | `EN-us` | `en-US` | canonicalise, exact | `en-US` |
176
+ | `en_US` | `en-US` | canonicalise, exact | `en-US` |
177
+ | `DE` | `de` | canonicalise, alias | `de-DE` |
178
+ | `xx-YY` | `xx-YY` | no exact, no alias, strip → `xx`, no exact, no alias | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
179
+ | `xx` | `xx` | no exact, no alias, nothing to strip | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
180
+ | `nb-NO` | `nb-NO` | strip → `nb`, miss | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
181
+ | `eng` | — | fails §3.3 (length 3) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
182
+ | `sr-Latn` | — | fails §3.3 (length 7) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
183
+ | `es-419` | — | fails §3.3 (region not `A`–`Z`) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
184
+ | `und` | — | fails §3.3 (length 3) | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
185
+ | `` (empty) | — | fails §3.1 | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
186
+ | absent / null | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
187
+ | options object absent | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
188
+ | non-string tag (number, object, boolean) | — | fails §3.1, `tagAbsent` | **throw** `POLYTYPO_UNKNOWN_LOCALE` |
189
+
190
+ The last three rows carry `tagAbsent: true`, a closed flag in `resolution.schema.json` added so
191
+ that "no tag was supplied at all" is expressible as a fixture rather than only as a JS test.
192
+ `transform("a", {})`, `{locale: null}`, a missing options object and every non-string locale now
193
+ raise the coded error rather than a native `TypeError` — which matters because a `TypeError` is
194
+ not in the taxonomy (ARCHITECTURE.md §4.6) and would differ in each of the five runtimes.
195
+
196
+ Every row is a fixture, not a candidate: **35 resolution cases run today**, covering these rows
197
+ and the malformed-tag rejections. `fixtures.schema.json` was extended with a case shape carrying
198
+ no `mode` and no `out`, whose expected result is a locale id or a thrown code. An earlier
199
+ revision of this paragraph said such rows "cannot be expressed in the existing fixture format";
200
+ that was true when written, and is why the schema was changed.
201
+
202
+ ---
203
+
204
+ ## 6. Open questions
205
+
206
+ 1. **The error taxonomy has no code for a malformed tag.** ARCHITECTURE.md §4.6 lists
207
+ `POLYTYPO_UNKNOWN_LOCALE`, `POLYTYPO_INVALID_MODE` and `POLYTYPO_MALFORMED_LOCALE_DATA`.
208
+ `eng`, `es-419` and `""` are structurally invalid rather than merely unknown, and a caller
209
+ would benefit from telling the two apart (one is a typo in their code, the other is a
210
+ language polytypo does not support yet). I have mapped both to
211
+ `POLYTYPO_UNKNOWN_LOCALE` because inventing a code would change a documented contract.
212
+ Recommend adding `POLYTYPO_INVALID_LOCALE`; operator decision.
213
+ 2. *(Closed.)* Resolution **is** fixture-covered: 35 resolution cases run today. The gap this
214
+ item reported — that `fixtures.schema.json` could not express a case with no `mode`, no `out`
215
+ and an expected locale id or thrown code — was closed by extending the schema, and the §5
216
+ table's rows are those cases. (The item also miscounted the pipeline as seven rules; it is
217
+ eight, since `hyphen` was added at order 35.)
218
+
219
+ 3. **`sv-FI` (Finland Swedish) silently resolves to `sv`.** The two genuinely differ in some
220
+ conventions, and PLAN.md §7 already flags Swedish quote practice as uncertain. The
221
+ algorithm is correct as specified; the question is whether `sv-FI` should be a locale of
222
+ its own. Not resolvable here — it needs a normative citation (PLAN.md §6.1).
223
+ 4. **Canonicalisation accepts `_` as a separator and is case-insensitive.** That is leniency
224
+ in a public API, and PLAN.md §5.1's stance is fail-fast. The arguments for it: Java, Ruby
225
+ and PHP ecosystems routinely produce `en_US`; the mapping is unambiguous; and rejecting it
226
+ converts a trivially-correctable input into a production exception. The argument against:
227
+ every accepted spelling is a spelling someone will depend on. I specified leniency —
228
+ **this is the one behavioural choice in this document that I would want the operator to
229
+ confirm before it becomes contract.**
230
+ 5. *(Closed.)* `registry.json` **has** a JSON Schema, compiled and enforced, so the §2
231
+ invariants are machine-checked rather than deferred to a runtime
232
+ `POLYTYPO_MALFORMED_LOCALE_DATA`.
233
+ 6. *(Closed.)* The registry and the files in `spec/locales/` **are** cross-checked in both
234
+ directions by `scripts/validate-spec.mjs`: an entry with no file, and a file with no entry,
235
+ both fail the build. This item asked for "a two-line CI check"; it exists.
236
+
237
+ 7. **Region-only variation is the only variation supported.** There is no way to express
238
+ `de-CH-1901` or a house-style variant. Deliberate for v1 and consistent with the schema
239
+ pattern, recorded so it is a decision.