polytypo 1.2.0 → 1.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +33 -1
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +161 -0
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +195 -6
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +12 -1
- data/lib/polytypo/data/fixtures/en-US.json +648 -1
- data/lib/polytypo/data/fixtures/es.json +193 -0
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +25 -1
- data/lib/polytypo/data/fixtures/fr.json +176 -1
- data/lib/polytypo/data/fixtures/it.json +161 -0
- data/lib/polytypo/data/fixtures/locale-resolution.json +76 -4
- data/lib/polytypo/data/fixtures/nl.json +121 -0
- data/lib/polytypo/data/fixtures/pl.json +137 -0
- data/lib/polytypo/data/fixtures/pt-BR.json +156 -0
- data/lib/polytypo/data/fixtures/pt-PT.json +156 -0
- data/lib/polytypo/data/fixtures/ru.json +23 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +153 -0
- data/lib/polytypo/data/locales/cs.json +90 -0
- data/lib/polytypo/data/locales/de-DE.json +7 -2
- data/lib/polytypo/data/locales/en-US.json +3 -3
- data/lib/polytypo/data/locales/es.json +111 -0
- data/lib/polytypo/data/locales/fr-CA.json +7 -1
- data/lib/polytypo/data/locales/fr.json +7 -1
- data/lib/polytypo/data/locales/it.json +95 -0
- data/lib/polytypo/data/locales/nl.json +84 -0
- data/lib/polytypo/data/locales/pl.json +96 -0
- data/lib/polytypo/data/locales/pt-BR.json +82 -0
- data/lib/polytypo/data/locales/pt-PT.json +84 -0
- data/lib/polytypo/data/locales/registry.json +23 -3
- data/lib/polytypo/data/locales/ru.json +2 -2
- data/lib/polytypo/data/locales/uk.json +130 -0
- data/lib/polytypo/data/rules/analyze.md +157 -0
- data/lib/polytypo/data/rules/apostrophe.md +432 -0
- data/lib/polytypo/data/rules/dashes.md +128 -37
- data/lib/polytypo/data/rules/ellipsis.md +271 -0
- data/lib/polytypo/data/rules/hyphen.md +353 -0
- data/lib/polytypo/data/rules/locale-resolution.md +239 -0
- data/lib/polytypo/data/rules/modes.md +1281 -0
- data/lib/polytypo/data/rules/nbsp.md +1157 -0
- data/lib/polytypo/data/rules/order.json +11 -11
- data/lib/polytypo/data/rules/pipeline-idempotency.md +605 -0
- data/lib/polytypo/data/rules/quotes.md +1324 -0
- data/lib/polytypo/data/rules/ranges.md +489 -0
- data/lib/polytypo/data/rules/spaces.md +649 -0
- data/lib/polytypo/data/rules/symbols.md +540 -0
- data/lib/polytypo/data/schema/fixtures.schema.json +18 -3
- data/lib/polytypo/engine/origin.rb +75 -0
- data/lib/polytypo/engine/pipeline.rb +72 -1
- data/lib/polytypo/engine/rules/dash_shared.rb +85 -3
- data/lib/polytypo/engine/rules/dashes.rb +4 -1
- data/lib/polytypo/engine/rules/nbsp.rb +43 -7
- data/lib/polytypo/engine/rules/ranges.rb +24 -20
- data/lib/polytypo/errors.rb +3 -0
- data/lib/polytypo/modes/runner.rb +17 -0
- data/lib/polytypo/modes/spans.rb +30 -2
- data/lib/polytypo/modes/yaml.rb +312 -0
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +126 -15
- metadata +31 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: f23f9a5d1868ea452ed02c3becffa73160c432d050f2af45ce54199a25b3064c
|
|
4
|
+
data.tar.gz: a9195694fa18f75875bb6dd20d5dccffc90c8f3272dd880ad0b11f3164d64b2d
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 3dc6343031405d5195562a8b561395fa111bb664c05649d6859f850f8df5a7d6f3e3945f52793d0adfbe910845cb34032376dd7edb281954273415ea3d0e2487
|
|
7
|
+
data.tar.gz: 05f095ad58cbc58c2247053345675c19e914aba0bc36163ebe108bc0b7751e7905774aa4a7fd3b18b052a1865025bedb30572be0b2b04506cd1e3b919290c0a9
|
data/README.md
CHANGED
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
|
|
22
22
|
This is the Ruby implementation. The full spec — all locales, all rules, worked examples in
|
|
23
23
|
each — lives in [polytypo/polytypo](https://github.com/polytypo/polytypo). This runtime supports
|
|
24
|
-
the `text` and `
|
|
24
|
+
the `text`, `html` and `yaml` modes fully, and `markdown` for the `commonmark` dialect only — `mdx`
|
|
25
25
|
returns `POLYTYPO_INVALID_DIALECT` (no MDX/JSX parser is available for Ruby; see
|
|
26
26
|
[Supported dialects](#supported-dialects)).
|
|
27
27
|
|
|
@@ -69,6 +69,22 @@ fallback to English. `mode:` defaults to `"text"` if omitted.
|
|
|
69
69
|
Polytypo.transform(input, locale: "fr", mode: "markdown", dialect: "commonmark")
|
|
70
70
|
```
|
|
71
71
|
|
|
72
|
+
`yaml` mode is the one that asks something of you, and it asks for a reason. YAML is a data
|
|
73
|
+
format with prose in some of it, so you name the keys whose values are prose; there is no default
|
|
74
|
+
and no guess:
|
|
75
|
+
|
|
76
|
+
```ruby
|
|
77
|
+
Polytypo.transform("summary: Rates -- all of them...\nrun: git diff -- a--b\n",
|
|
78
|
+
locale: "en-US", mode: "yaml", keys: ["summary"])
|
|
79
|
+
# => "summary: Rates—all of them…\nrun: git diff -- a--b\n"
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
Nothing in YAML's syntax separates a sentence from a shell script: `description` holds one and
|
|
83
|
+
`run` holds the other, spelled identically. Quoting, indentation, anchors and a block scalar's
|
|
84
|
+
chomping indicator are never decoded and rewritten — the file is located, not re-emitted — so the
|
|
85
|
+
trailing newlines of a `|+` block come back exactly as you wrote them. An empty `keys:` array is
|
|
86
|
+
legal and processes nothing, and `yaml` mode needs no parser at all.
|
|
87
|
+
|
|
72
88
|
`require "polytypo"` never loads the native `commonmarker` extension unless you actually call
|
|
73
89
|
`Polytypo.transform` with `mode: "markdown"` — the require happens lazily, inside the markdown
|
|
74
90
|
pipeline only. A single `Polytypo.transform` module method (not separate
|
|
@@ -76,6 +92,22 @@ pipeline only. A single `Polytypo.transform` module method (not separate
|
|
|
76
92
|
`require` already gets you the same dependency isolation JS/Python's subpath split exists for,
|
|
77
93
|
without a separate namespace per mode.
|
|
78
94
|
|
|
95
|
+
`Polytypo.analyze` runs the same pipeline and reports what it would do instead of doing it — one
|
|
96
|
+
record per edit, each with the rule that made it and code-point offsets into the input you passed
|
|
97
|
+
(into the **document**, in `html`, `markdown` and `yaml` mode, not into a span):
|
|
98
|
+
|
|
99
|
+
```ruby
|
|
100
|
+
Polytypo.analyze(%q{Wait... "really"?}, locale: "en-US")
|
|
101
|
+
# => [#<data Polytypo::Change rule_id="ellipsis", start=4, end=7, before="...", after="…">,
|
|
102
|
+
# #<data Polytypo::Change rule_id="quotes", start=8, end=9, before="\"", after="“">, ...]
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
It is a report, not a patch. The list is empty exactly when `Polytypo.transform` would return the
|
|
106
|
+
input unchanged, and every `rule_id` is a rule that was enabled for that call — but two rules may
|
|
107
|
+
touch the same original range (French `spaces` deletes the space before `:` and `nbsp` puts a
|
|
108
|
+
no-break one back), so replaying the list is not guaranteed to reproduce the output. Call
|
|
109
|
+
`Polytypo.transform` for the text. Full contract: `spec/rules/analyze.md`.
|
|
110
|
+
|
|
79
111
|
### Errors
|
|
80
112
|
|
|
81
113
|
Every error `Polytypo.transform` raises is a `Polytypo::Error` carrying one of seven stable
|
data/lib/polytypo/data/VERSION
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
1.
|
|
1
|
+
1.3.1
|
|
@@ -0,0 +1,161 @@
|
|
|
1
|
+
{
|
|
2
|
+
"spec": "1.3.1",
|
|
3
|
+
"locale": "cs",
|
|
4
|
+
"cases": [
|
|
5
|
+
{
|
|
6
|
+
"id": "cs-quotes-primary",
|
|
7
|
+
"rule": "quotes",
|
|
8
|
+
"mode": "text",
|
|
9
|
+
"in": "Řekl \"ahoj\".",
|
|
10
|
+
"out": "Řekl „ahoj“.",
|
|
11
|
+
"note": "quotes.primary = U+201E/U+201C, tedy «uvozovky typu 99 66» podle IJP."
|
|
12
|
+
},
|
|
13
|
+
{
|
|
14
|
+
"id": "cs-quotes-not-polish",
|
|
15
|
+
"rule": "quotes",
|
|
16
|
+
"mode": "text",
|
|
17
|
+
"in": "„ahoj”",
|
|
18
|
+
"out": "„ahoj“",
|
|
19
|
+
"note": "ROZLIŠUJÍCÍ PŘÍPAD: zavírací znaménko je U+201C, ne polské U+201D. Vstup je polský pár; výstup ukazuje, že se přesází na českou podobu. V sazbě je ten rozdíl skoro neviditelný, proto má vlastní fixture."
|
|
20
|
+
},
|
|
21
|
+
{
|
|
22
|
+
"id": "cs-quotes-nested",
|
|
23
|
+
"rule": "quotes",
|
|
24
|
+
"mode": "text",
|
|
25
|
+
"in": "\"Toto je 'citát' v citátu\"",
|
|
26
|
+
"out": "„Toto je ‚citát‘ v citátu“",
|
|
27
|
+
"note": "Druhý stupeň jsou jednoduché uvozovky, protože IJP je řadí «v první řadě» pro uvozovky v uvozovkách. Kódy U+201A/U+2018 jsou rozhodnutí operátora (18. 9. 2026) — označení «99 66» zdroj uvádí jen pro dvojité, a tabulka H.1 normy ČSN je placená. V témže případě je vidět N3: jednopísmenné «v» se váže k následujícímu slovu."
|
|
28
|
+
},
|
|
29
|
+
{
|
|
30
|
+
"id": "cs-ellipsis-three-dots",
|
|
31
|
+
"rule": "ellipsis",
|
|
32
|
+
"mode": "text",
|
|
33
|
+
"in": "Počkej... co?",
|
|
34
|
+
"out": "Počkej… co?",
|
|
35
|
+
"note": "IJP «Tři tečky»: výpustka je jeden znak U+2026, nikoli tři samostatné tečky."
|
|
36
|
+
},
|
|
37
|
+
{
|
|
38
|
+
"id": "cs-ellipsis-not-abbreviated",
|
|
39
|
+
"rule": "ellipsis",
|
|
40
|
+
"mode": "text",
|
|
41
|
+
"in": "Opravdu?.. ano",
|
|
42
|
+
"out": "Opravdu?… ano",
|
|
43
|
+
"note": "ZRCADLOVÝ PŘÍPAD k ukrajinskému uk-ellipsis-abbreviated-after-question, zařazený proto, že analogie mezi slovanskými jazyky je tu svůdná a mylná: čeština má abbreviatedAfterTerminal = false, takže dvě tečky po «?» se doplní na plnou výpustku, zatímco ukrajinština je zkracuje."
|
|
44
|
+
},
|
|
45
|
+
{
|
|
46
|
+
"id": "cs-dashes-parenthetical",
|
|
47
|
+
"rule": "dashes",
|
|
48
|
+
"mode": "text",
|
|
49
|
+
"in": "Tato kniha - vydaná ještě před válkou - je opravdu úžasná.",
|
|
50
|
+
"out": "Tato kniha – vydaná ještě před válkou – je opravdu úžasná.",
|
|
51
|
+
"note": "Vlastní příklad IJP. dash.parenthetical = en-spaced: krátká pomlčka U+2013 s mezerami. Délku tu rozhoduje sám zdroj (definiční glyf stránky), na rozdíl od polštiny, kde ji musel rozhodnout operátor."
|
|
52
|
+
},
|
|
53
|
+
{
|
|
54
|
+
"id": "cs-ranges-en-tight",
|
|
55
|
+
"rule": "ranges",
|
|
56
|
+
"mode": "text",
|
|
57
|
+
"rules": {
|
|
58
|
+
"ranges": true
|
|
59
|
+
},
|
|
60
|
+
"in": "strana 23-26",
|
|
61
|
+
"out": "strana 23–26",
|
|
62
|
+
"note": "ranges.md §3: dash.range = en-tight, příklad z IJP «strana 23–26». Potvrzeno i dokumentem ÚJČ k ČSN 01 6910, který mezery v rozsahu popisuje jako výjimku pro víceslovné výrazy."
|
|
63
|
+
},
|
|
64
|
+
{
|
|
65
|
+
"id": "cs-hyphen-noop",
|
|
66
|
+
"rule": "hyphen",
|
|
67
|
+
"mode": "text",
|
|
68
|
+
"in": "česko-slovenský",
|
|
69
|
+
"out": "česko-slovenský",
|
|
70
|
+
"note": "hyphen.md §2: tři prázdné seznamy, tedy dokazatelný no-op. Uzavřený výčet IJP o zalomení řádků spojovník neuvádí — doložená nepřítomnost, ne nenalezená citace. Rozdíl proti ukrajinštině, kde § 64 п. 4 takový seznam má."
|
|
71
|
+
},
|
|
72
|
+
{
|
|
73
|
+
"id": "cs-nbsp-short-words",
|
|
74
|
+
"rule": "nbsp",
|
|
75
|
+
"mode": "text",
|
|
76
|
+
"in": "v Plzni a u babičky",
|
|
77
|
+
"out": "v Plzni a u babičky",
|
|
78
|
+
"note": "N3 afterShortWords: neslabičné předložky k, s, v, z a slabičné o, u se spojkami a, i. Příklady jsou z IJP. Na rozdíl od polštiny je zákaz v prameni bezpodmínečný."
|
|
79
|
+
},
|
|
80
|
+
{
|
|
81
|
+
"id": "cs-nbsp-units",
|
|
82
|
+
"rule": "nbsp",
|
|
83
|
+
"mode": "text",
|
|
84
|
+
"in": "10 ha, 14 % a 19 °C",
|
|
85
|
+
"out": "10 ha, 14 % a 19 °C",
|
|
86
|
+
"note": "N5 na třech značkách z vlastních příkladů IJP. Zároveň je vidět N3 na spojce «a»."
|
|
87
|
+
},
|
|
88
|
+
{
|
|
89
|
+
"id": "cs-nbsp-percent-adjective-inert",
|
|
90
|
+
"rule": "nbsp",
|
|
91
|
+
"mode": "text",
|
|
92
|
+
"in": "20% roztok",
|
|
93
|
+
"out": "20% roztok",
|
|
94
|
+
"note": "Povolená varianta IJP: «20% = 20procentní» se píše bez mezery. N5 existující mezeru pouze nahrazuje a nikdy ji nevkládá, takže jeden seznam obslouží obě doložené podoby."
|
|
95
|
+
},
|
|
96
|
+
{
|
|
97
|
+
"id": "cs-nbsp-paragraph-sign",
|
|
98
|
+
"rule": "nbsp",
|
|
99
|
+
"mode": "text",
|
|
100
|
+
"in": "§ 23",
|
|
101
|
+
"out": "§ 23",
|
|
102
|
+
"note": "N6 afterSymbols. Ze stejného citovaného bodu byly vynechány «*», «†» a «#»: «* 1921» je bajt po bajtu odrážka seznamu následovaná číslicí, což by v režimu markdown byl falešný zásah."
|
|
103
|
+
},
|
|
104
|
+
{
|
|
105
|
+
"id": "cs-nbsp-abbreviations",
|
|
106
|
+
"rule": "nbsp",
|
|
107
|
+
"mode": "text",
|
|
108
|
+
"in": "a. s. a s. r. o.",
|
|
109
|
+
"out": "a. s. a s. r. o.",
|
|
110
|
+
"note": "N4 na obou zkratkách. Zároveň fixuje interakci s N3, která je opačná než u polského «i in.»: po jednopísmenném «a» následuje tečka, ne mezera, takže N3 index nenárokuje a obě mezery dostane N4 — kdežto samostatné «a» uprostřed věty N3 váže."
|
|
111
|
+
},
|
|
112
|
+
{
|
|
113
|
+
"id": "cs-nbsp-before-word",
|
|
114
|
+
"rule": "nbsp",
|
|
115
|
+
"mode": "text",
|
|
116
|
+
"in": "tzv. odborník",
|
|
117
|
+
"out": "tzv. odborník",
|
|
118
|
+
"note": "N10 beforeWord: tj., tzv., tzn. jsou uzavřený a doslovně citovaný seznam."
|
|
119
|
+
},
|
|
120
|
+
{
|
|
121
|
+
"id": "cs-nbsp-initial-single",
|
|
122
|
+
"rule": "nbsp",
|
|
123
|
+
"mode": "text",
|
|
124
|
+
"in": "M. Těšitelová",
|
|
125
|
+
"out": "M. Těšitelová",
|
|
126
|
+
"note": "initialBinding = single. Vlastní příklady zdroje (Fr. Daneš, M. Těšitelová) mají JEDNO zkrácené křestní jméno, takže «chain» by citované pravidlo nechal nenaplněné."
|
|
127
|
+
},
|
|
128
|
+
{
|
|
129
|
+
"id": "cs-nbsp-no-space-before-question",
|
|
130
|
+
"rule": "nbsp",
|
|
131
|
+
"mode": "text",
|
|
132
|
+
"in": "Opravdu?",
|
|
133
|
+
"out": "Opravdu?",
|
|
134
|
+
"note": "Čeština není francouzština: beforePunctuation i narrowBeforePunctuation jsou prázdné a zdroj odstup před «?» výslovně upírá."
|
|
135
|
+
},
|
|
136
|
+
{
|
|
137
|
+
"id": "cs-spaces-brackets-and-final-stop",
|
|
138
|
+
"rule": "spaces",
|
|
139
|
+
"mode": "text",
|
|
140
|
+
"in": "Text ( běžný ) se nemění .",
|
|
141
|
+
"out": "Text (běžný) se nemění.",
|
|
142
|
+
"note": "spaces.md §3: pravidlo nezávislé na locale, případ je tu proto, aby cs mělo po jednom pro každé kanonické pravidlo."
|
|
143
|
+
},
|
|
144
|
+
{
|
|
145
|
+
"id": "cs-symbols-copyright-and-dimensions",
|
|
146
|
+
"rule": "symbols",
|
|
147
|
+
"mode": "text",
|
|
148
|
+
"in": "Copyright (c) 2026, formát 40x60 cm",
|
|
149
|
+
"out": "Copyright © 2026, formát 40×60 cm",
|
|
150
|
+
"note": "symbols.md: «(c)» se stane U+00A9 a «x» mezi číslicemi U+00D7. U+00A0 před «cm» vkládá nbsp N5."
|
|
151
|
+
},
|
|
152
|
+
{
|
|
153
|
+
"id": "cs-apostrophe-elision",
|
|
154
|
+
"rule": "apostrophe",
|
|
155
|
+
"mode": "text",
|
|
156
|
+
"in": "Šel do McDonald's a zpět",
|
|
157
|
+
"out": "Šel do McDonald’s a zpět",
|
|
158
|
+
"note": "Pravidlo apostrophe nečte data locale: U+0027 mezi písmeny se stane U+2019. V téže větě je vidět N3, které váže spojku «a» k následujícímu slovu — dvě nezávislá pravidla na jednom řádku."
|
|
159
|
+
}
|
|
160
|
+
]
|
|
161
|
+
}
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.3.1",
|
|
3
3
|
"locale": "de-DE",
|
|
4
4
|
"cases": [
|
|
5
5
|
{
|
|
@@ -452,16 +452,32 @@
|
|
|
452
452
|
"rule": "nbsp",
|
|
453
453
|
"mode": "text",
|
|
454
454
|
"in": "siehe § 12 und Nr. 5",
|
|
455
|
-
"out": "siehe § 12 und Nr.
|
|
456
|
-
"note": "Two halves with different evidence. `§` is in afterSymbols and binds through N6. `Nr.` binds through nothing: it was removed from the de-DE locale file entirely, because no consulted source states a spacing rule for it — the Schreibweisungen never mention it (it is absent from Rz. 439's closed list of Begriffszeichen and from the Sachregister), Duden classifies it as an abbreviation, which refutes afterSymbols but prescribes no spacing, and DIN 5008 is paywalled and unread. Under PLAN.md §6.1 an entry whose citation does not cover it does not ship. The gap is evidentiary, not typographic."
|
|
455
|
+
"out": "siehe § 12 und Nr. 5",
|
|
456
|
+
"note": "Two halves with different evidence. `§` is in afterSymbols and binds through N6. `Nr.` binds through nothing: it was removed from the de-DE locale file entirely, because no consulted source states a spacing rule for it — the Schreibweisungen never mention it (it is absent from Rz. 439's closed list of Begriffszeichen and from the Sachregister), Duden classifies it as an abbreviation, which refutes afterSymbols but prescribes no spacing, and DIN 5008 is paywalled and unread. Under PLAN.md §6.1 an entry whose citation does not cover it does not ship. The gap is evidentiary, not typographic. Seit Spec 1.3.0 bindet auch `Nr.` (nbsp.md §3.11 N9, nbsp.beforeNumber), sodass beide Bindungen in einem Satz sichtbar sind."
|
|
457
457
|
},
|
|
458
458
|
{
|
|
459
|
-
"id": "de-de-nbsp-nr-
|
|
459
|
+
"id": "de-de-nbsp-nr-binds-number",
|
|
460
460
|
"rule": "nbsp",
|
|
461
461
|
"mode": "text",
|
|
462
462
|
"in": "siehe Nr. 5",
|
|
463
|
-
"out": "siehe Nr.
|
|
464
|
-
"note": "
|
|
463
|
+
"out": "siehe Nr. 5",
|
|
464
|
+
"note": "nbsp.md §3.11 N9: `Nr.` steht seit Spec 1.3.0 in de-DEs nbsp.beforeNumber und bindet an die folgende Zahl. Die Zeile ist eine vereinheitlichte Schreibung, kein Zitat — DIN 5008 ist kostenpflichtig und ungelesen, Duden und Amtliches Regelwerk sprechen keine Abstandsregel aus; die Begründung steht im sources-Eintrag von spec/locales/de-DE.json."
|
|
465
|
+
},
|
|
466
|
+
{
|
|
467
|
+
"id": "de-de-nbsp-nr-needs-a-digit",
|
|
468
|
+
"rule": "nbsp",
|
|
469
|
+
"mode": "text",
|
|
470
|
+
"in": "siehe Nr. fünf",
|
|
471
|
+
"out": "siehe Nr. fünf",
|
|
472
|
+
"note": "nbsp.md §3.11 N9 Schritt 4: „eine folgende Zahl“ heißt, dass cp[a+k+1] in DIGIT liegt. Ein ausgeschriebenes Zahlwort ist keine Ziffer, also bleibt der Abstand trennbar."
|
|
473
|
+
},
|
|
474
|
+
{
|
|
475
|
+
"id": "de-de-nbsp-nr-after-hyphen-unbound",
|
|
476
|
+
"rule": "nbsp",
|
|
477
|
+
"mode": "text",
|
|
478
|
+
"in": "Tarif-Nr. 7",
|
|
479
|
+
"out": "Tarif-Nr. 7",
|
|
480
|
+
"note": "nbsp.md §3.11 N9 Schritt 2: die linke Grenze muss NONE, SPACELIKE, OPENISH oder SENTENCE-DASH sein; ein Bindestrich fällt durch, eine Abkürzung kann nicht unmittelbar nach einem wortinternen Bindestrich beginnen."
|
|
465
481
|
},
|
|
466
482
|
{
|
|
467
483
|
"id": "de-de-nbsp-before-word-sankt",
|
|
@@ -574,6 +590,179 @@
|
|
|
574
590
|
"in": "„Er sagte ‚Hans'“ und ging.",
|
|
575
591
|
"out": "„Er sagte ‚Hans‘“ und ging.",
|
|
576
592
|
"note": "apostrophe.md §3.3 case 3a (spec 1.2.0): the same shape with the outer pair and the opening single quote already curly. `quotes` still pairs the straight closing single quote with ‚, so case 3a never sees it."
|
|
593
|
+
},
|
|
594
|
+
{
|
|
595
|
+
"id": "de-de-ranges-closed-up-currency",
|
|
596
|
+
"rule": "ranges",
|
|
597
|
+
"mode": "text",
|
|
598
|
+
"in": "€15-€20",
|
|
599
|
+
"out": "€15–€20",
|
|
600
|
+
"rules": {
|
|
601
|
+
"ranges": true
|
|
602
|
+
},
|
|
603
|
+
"note": "ranges.md §3.2a (spec 1.3.0): the closed-up walk is locale-independent shape logic — it reads code points, not a currency list, so `de-DE` gets it on the same terms as `en-US`. Note that German normally sets the sign AFTER the amount (`30 EUR`), which is a spaced symbol and therefore out of scope by §7.8; this row asserts the mechanism, not that the shape is idiomatic German."
|
|
604
|
+
},
|
|
605
|
+
{
|
|
606
|
+
"id": "de-de-dashes-closed-up-range-declined",
|
|
607
|
+
"rule": "dashes",
|
|
608
|
+
"mode": "text",
|
|
609
|
+
"in": "$15 - $20",
|
|
610
|
+
"out": "$15 - $20",
|
|
611
|
+
"note": "ranges.md §3.2a and dashes.md §1 (spec 1.3.0): the accepted cost, pinned in a locale whose `dash.parenthetical` is SPACED rather than tight — through spec 1.2.0 this input gave `$15 – $20` with default options, and the loss is the same shape in both kinds of locale."
|
|
612
|
+
},
|
|
613
|
+
{
|
|
614
|
+
"id": "de-de-dashes-t1-closed-up-symbol-adjacent",
|
|
615
|
+
"rule": "dashes",
|
|
616
|
+
"mode": "text",
|
|
617
|
+
"in": "a—$15-$20",
|
|
618
|
+
"out": "a—$15-$20",
|
|
619
|
+
"rules": {
|
|
620
|
+
"ranges": true
|
|
621
|
+
},
|
|
622
|
+
"note": "ranges.md §3.2a, via dashes.md §3.2 step 8 (spec 1.3.0): T1's reach is transparent to one CLOSED-SYMBOL between the token and the digit run. THE IDEMPOTENCY WITNESS — before the amendment T1 read the U+0024 as the neighbour, did not fire, and `dashes` converted the em dash to a spaced en dash; that U+0020 replaced the range's `before`, and the range converted on the SECOND pass. `a—15-20` was inert all along and this input now matches it. de-DE because T1 only applies where `dash.parenthetical` is spaced, which is exactly why the cluster guard's alphabet was NOT widened instead — that guard is unconditional and would have taken em-tight locales with it. To be precise about \"before the amendment\": the drift is a property of the INTERMEDIATE state — ranges.md §3.2a's widened candidacy without T1's transparency — which never shipped. Under spec 1.2.0 this input was stable, because the token was not a range candidate at all."
|
|
623
|
+
},
|
|
624
|
+
{
|
|
625
|
+
"id": "de-de-dashes-t1-closed-up-symbol-right",
|
|
626
|
+
"rule": "dashes",
|
|
627
|
+
"mode": "text",
|
|
628
|
+
"in": "a--15% - 20%",
|
|
629
|
+
"out": "a--15% - 20%",
|
|
630
|
+
"rules": {
|
|
631
|
+
"ranges": true
|
|
632
|
+
},
|
|
633
|
+
"note": "dashes.md §3.2 step 8 (spec 1.3.0): T1's outward walk from the digit run steps over one CLOSED-SYMBOL. The cluster guard cannot catch this one — a cluster ends at the U+0020 — so before the amendment the `--` became a spaced en dash and the range converted on the next pass. To be precise about \"before the amendment\": the drift is a property of the INTERMEDIATE state — ranges.md §3.2a's widened candidacy without T1's transparency — which never shipped. Under spec 1.2.0 this input was stable, because the token was not a range candidate at all."
|
|
634
|
+
},
|
|
635
|
+
{
|
|
636
|
+
"id": "de-de-dashes-t1-closed-up-symbol-left",
|
|
637
|
+
"rule": "dashes",
|
|
638
|
+
"mode": "text",
|
|
639
|
+
"in": "$1 - $1--a",
|
|
640
|
+
"out": "$1 - $1--a",
|
|
641
|
+
"rules": {
|
|
642
|
+
"ranges": true
|
|
643
|
+
},
|
|
644
|
+
"note": "dashes.md §3.2 step 8 (spec 1.3.0): the mirror of the T1 case, with the symbol between the token and the digit run rather than beyond it — `cp[L]` is `1` here but the walk outward from the run meets `$` before the space and the dash. Both steps of the transparency are needed; each of these two cases fails without its own. To be precise about \"before the amendment\": the drift is a property of the INTERMEDIATE state — ranges.md §3.2a's widened candidacy without T1's transparency — which never shipped. Under spec 1.2.0 this input was stable, because the token was not a range candidate at all."
|
|
645
|
+
},
|
|
646
|
+
{
|
|
647
|
+
"id": "de-de-dashes-t1-closed-up-symbol-cost",
|
|
648
|
+
"rule": "dashes",
|
|
649
|
+
"mode": "text",
|
|
650
|
+
"in": "Anstieg--50%--war",
|
|
651
|
+
"out": "Anstieg--50%--war",
|
|
652
|
+
"note": "dashes.md §3.2 step 8 (spec 1.3.0): THE ACCEPTED COST of T1's CLOSED-SYMBOL transparency, with DEFAULT options — no `rules` key here, because the cost lands on `dashes`, which is on by default. Through spec 1.2.0 this gave `Anstieg – 50% – war`; T1 now fires, because the U+0025 no longer hides the far dash from its outward walk. `Anstieg--50--war` has always been left alone, so the two shapes agree. The cost is confined to locales whose `dash.parenthetical` is spaced — see en-us-dashes-t1-closed-up-symbol-unaffected for the other side of that line."
|
|
653
|
+
},
|
|
654
|
+
{
|
|
655
|
+
"id": "de-de-dashes-t1-joiner-and-closed-up-symbol",
|
|
656
|
+
"rule": "dashes",
|
|
657
|
+
"mode": "text",
|
|
658
|
+
"in": "a--$15-$20",
|
|
659
|
+
"out": "a--$15-$20",
|
|
660
|
+
"rules": {
|
|
661
|
+
"ranges": true
|
|
662
|
+
},
|
|
663
|
+
"note": "dashes.md §3.2 step 8 and §3.2b (spec 1.3.0): T1's RIGHT branch reads effective neighbours, not raw indices. Through spec 1.2.0 the branch required a DIGIT flank, which put the token, the digit run, the U+2060 pair and the far dash in one cluster, so step 7 declined before step 8 could disagree with §3.2b; §3.2a's p1 admits a CLOSED-SYMBOL flank, CLOSED-SYMBOL is deliberately not in the cluster alphabet, and step 8 now stands alone. A port transcribing the pre-1.3.0 raw reading does not fire T1 here, converts `--` to a spaced en dash, and the range converts on the SECOND pass. Nothing else in the suite separates the two readings."
|
|
664
|
+
},
|
|
665
|
+
{
|
|
666
|
+
"id": "de-de-dashes-t1-joiner-only",
|
|
667
|
+
"rule": "dashes",
|
|
668
|
+
"mode": "text",
|
|
669
|
+
"in": "a--15 - 20",
|
|
670
|
+
"out": "a--15 - 20",
|
|
671
|
+
"rules": {
|
|
672
|
+
"ranges": true
|
|
673
|
+
},
|
|
674
|
+
"note": "dashes.md §3.2 step 8 and §3.2b: T1's right branch reads effective neighbours, pinned WITHOUT any closed-up symbol involved. The U+0020 ends the cluster before the far dash, so step 7 never declined this shape and the raw-index reading of the right branch was always wrong here — it is spec 1.3.0 that says so out loud (the text carried the raw reading through 1.2.0 while the reference implementation read effective neighbours). Companion to de-de-dashes-t1-joiner-and-closed-up-symbol, which covers the same divergence on the symbol path."
|
|
675
|
+
},
|
|
676
|
+
{
|
|
677
|
+
"id": "de-de-yaml-colon-split-declines-growth",
|
|
678
|
+
"rule": "dashes",
|
|
679
|
+
"mode": "yaml",
|
|
680
|
+
"keys": [
|
|
681
|
+
"description"
|
|
682
|
+
],
|
|
683
|
+
"in": "description: a:--b\n",
|
|
684
|
+
"out": "description: a:--b\n",
|
|
685
|
+
"note": "modes.md §3.8.6: `:` is an opaque unit inside a plain scalar, so the dash token sits at a span extremity and the `-spaced` replacement (2 → 3) is discarded by the edge-growth rule of §3.4. Without the split this becomes `a: – b`, which no longer parses as one scalar."
|
|
686
|
+
},
|
|
687
|
+
{
|
|
688
|
+
"id": "de-de-yaml-hash-split-declines-growth",
|
|
689
|
+
"rule": "dashes",
|
|
690
|
+
"mode": "yaml",
|
|
691
|
+
"keys": [
|
|
692
|
+
"description"
|
|
693
|
+
],
|
|
694
|
+
"in": "description: a--#b\n",
|
|
695
|
+
"out": "description: a--#b\n",
|
|
696
|
+
"note": "modes.md §3.8.6: `#` is an opaque unit inside a plain scalar for the same reason. Without the split this becomes `a – #b`, and everything from the hash becomes a comment."
|
|
697
|
+
},
|
|
698
|
+
{
|
|
699
|
+
"id": "de-de-yaml-block-prose",
|
|
700
|
+
"rule": "quotes",
|
|
701
|
+
"mode": "yaml",
|
|
702
|
+
"keys": [
|
|
703
|
+
"description"
|
|
704
|
+
],
|
|
705
|
+
"in": "description: |\n Er sagte \"hallo\" und ging\n",
|
|
706
|
+
"out": "description: |\n Er sagte „hallo“ und ging\n",
|
|
707
|
+
"note": "modes.md §8 item 4: the same construct in a locale whose quote glyphs are structurally different from en-US, so a fixture is not the same English output with a different locale field."
|
|
708
|
+
},
|
|
709
|
+
{
|
|
710
|
+
"id": "de-de-yaml-unlisted-key-untouched",
|
|
711
|
+
"rule": "dashes",
|
|
712
|
+
"mode": "yaml",
|
|
713
|
+
"keys": [
|
|
714
|
+
"description"
|
|
715
|
+
],
|
|
716
|
+
"in": "run: git diff -- a--b\n",
|
|
717
|
+
"out": "run: git diff -- a--b\n",
|
|
718
|
+
"note": "modes.md §3.8.2: the measured case. A keyless draft rewrote shell inside `run:`; `keys` is what makes it unreachable."
|
|
719
|
+
},
|
|
720
|
+
{
|
|
721
|
+
"id": "de-de-modes-three-dash-span-edge-html",
|
|
722
|
+
"rule": "dashes",
|
|
723
|
+
"mode": "html",
|
|
724
|
+
"in": "a<em>---</em>b",
|
|
725
|
+
"out": "a<em>---</em>b",
|
|
726
|
+
"note": "modes.md §3.4, character clause (spec 1.3.0): `---` → `␣–␣` is 3 → 3, so the length clause misses it while U+0020 lands on both span extremities. Before the character clause this produced `a<em> – </em>b` — an element beginning and ending with a space it never held, which is the harm §7.3 exists to prevent. Two dashes were already declined here; three now are too."
|
|
727
|
+
},
|
|
728
|
+
{
|
|
729
|
+
"id": "de-de-modes-three-dash-span-edge-markdown",
|
|
730
|
+
"rule": "dashes",
|
|
731
|
+
"mode": "markdown",
|
|
732
|
+
"dialect": "commonmark",
|
|
733
|
+
"in": "x *---* y",
|
|
734
|
+
"out": "x *---* y",
|
|
735
|
+
"note": "modes.md §3.4, character clause, and §5 item 2: before it this produced `x * – * y`, which de-flanks the asterisks so that on the next run they are literal content rather than emphasis delimiters — the span partition was not stable."
|
|
736
|
+
},
|
|
737
|
+
{
|
|
738
|
+
"id": "de-de-modes-three-dash-interior-html",
|
|
739
|
+
"rule": "dashes",
|
|
740
|
+
"mode": "html",
|
|
741
|
+
"in": "<p>a---b</p>",
|
|
742
|
+
"out": "<p>a – b</p>",
|
|
743
|
+
"note": "modes.md §3.4: the companion to de-de-modes-three-dash-span-edge-html. Interior to a span the same edit applies — the clause restricts extremities only, and a fixture set without this pair cannot tell the two apart."
|
|
744
|
+
},
|
|
745
|
+
{
|
|
746
|
+
"id": "de-de-yaml-three-dash-against-colon",
|
|
747
|
+
"rule": "dashes",
|
|
748
|
+
"mode": "yaml",
|
|
749
|
+
"keys": [
|
|
750
|
+
"description"
|
|
751
|
+
],
|
|
752
|
+
"in": "description: a:---b\n",
|
|
753
|
+
"out": "description: a:---b\n",
|
|
754
|
+
"note": "modes.md §3.8.6 with §3.4's character clause: without it this became `description: a: – b`, which no longer parses — `: ` is a mapping indicator. The two-dash companion de-de-yaml-colon-split-declines-growth is caught by the length clause; only this one needs the character clause."
|
|
755
|
+
},
|
|
756
|
+
{
|
|
757
|
+
"id": "de-de-yaml-three-dash-against-hash",
|
|
758
|
+
"rule": "dashes",
|
|
759
|
+
"mode": "yaml",
|
|
760
|
+
"keys": [
|
|
761
|
+
"description"
|
|
762
|
+
],
|
|
763
|
+
"in": "description: a---#b\n",
|
|
764
|
+
"out": "description: a---#b\n",
|
|
765
|
+
"note": "modes.md §3.8.6 with §3.4's character clause: without it this became `description: a – #b`, and everything from the hash became a comment — the value silently truncated to `a –`."
|
|
577
766
|
}
|
|
578
767
|
]
|
|
579
768
|
}
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.3.1",
|
|
3
3
|
"locale": "en-GB",
|
|
4
4
|
"cases": [
|
|
5
5
|
{
|
|
@@ -1285,6 +1285,17 @@
|
|
|
1285
1285
|
"in": "dell'“arte” moderna",
|
|
1286
1286
|
"out": "dell’‘arte’ moderna",
|
|
1287
1287
|
"note": "apostrophe.md §3.3 case 3a (spec 1.2.0), §6 row 16: `quotes` re-typesets the curly pair as en-GB's primary single quotes, and the elision before the opening glyph converts to U+2019."
|
|
1288
|
+
},
|
|
1289
|
+
{
|
|
1290
|
+
"id": "en-gb-ranges-closed-up-pound",
|
|
1291
|
+
"rule": "ranges",
|
|
1292
|
+
"mode": "text",
|
|
1293
|
+
"in": "£15-£20",
|
|
1294
|
+
"out": "£15–£20",
|
|
1295
|
+
"rules": {
|
|
1296
|
+
"ranges": true
|
|
1297
|
+
},
|
|
1298
|
+
"note": "ranges.md §3.2a (spec 1.3.0): U+00A3, pinned in the locale that actually writes it. The closed-up walk is locale-independent shape logic; `en-GB` shares `en-US`'s `range: \"en-tight\"`."
|
|
1288
1299
|
}
|
|
1289
1300
|
]
|
|
1290
1301
|
}
|