polytypo 1.6.3 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +25 -0
- data/lib/polytypo/data/.not-vendored +11 -0
- data/lib/polytypo/data/README.md +16 -11
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +1 -1
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +14 -1
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +1 -1
- data/lib/polytypo/data/fixtures/en-US.json +521 -1
- data/lib/polytypo/data/fixtures/es.json +1 -1
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +1 -1
- data/lib/polytypo/data/fixtures/fr.json +68 -1
- data/lib/polytypo/data/fixtures/it.json +1 -1
- data/lib/polytypo/data/fixtures/locale-resolution.json +1 -1
- data/lib/polytypo/data/fixtures/nl.json +1 -1
- data/lib/polytypo/data/fixtures/pl.json +1 -1
- data/lib/polytypo/data/fixtures/pt-BR.json +1 -1
- data/lib/polytypo/data/fixtures/pt-PT.json +1 -1
- data/lib/polytypo/data/fixtures/ru.json +1 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/tr.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +1 -1
- data/lib/polytypo/data/locales/cs.json +56 -37
- data/lib/polytypo/data/locales/de-CH.json +43 -34
- data/lib/polytypo/data/locales/de-DE.json +47 -32
- data/lib/polytypo/data/locales/el.json +51 -1
- data/lib/polytypo/data/locales/en-GB.json +56 -4
- data/lib/polytypo/data/locales/en-US.json +68 -4
- data/lib/polytypo/data/locales/es.json +66 -29
- data/lib/polytypo/data/locales/fi.json +77 -7
- data/lib/polytypo/data/locales/fr-CA.json +63 -36
- data/lib/polytypo/data/locales/fr.json +70 -41
- data/lib/polytypo/data/locales/it.json +65 -22
- data/lib/polytypo/data/locales/nl.json +61 -29
- data/lib/polytypo/data/locales/pl.json +61 -28
- data/lib/polytypo/data/locales/pt-BR.json +53 -28
- data/lib/polytypo/data/locales/pt-PT.json +55 -28
- data/lib/polytypo/data/locales/registry.json +1 -1
- data/lib/polytypo/data/locales/ru.json +45 -30
- data/lib/polytypo/data/locales/sv.json +63 -4
- data/lib/polytypo/data/locales/tr.json +72 -32
- data/lib/polytypo/data/locales/uk.json +55 -27
- data/lib/polytypo/data/rules/modes.md +430 -11
- data/lib/polytypo/data/rules/order.json +1 -1
- data/lib/polytypo/data/schema/fixtures.schema.json +12 -1
- data/lib/polytypo/engine/pipeline.rb +29 -0
- data/lib/polytypo/modes/markdown.rb +256 -31
- data/lib/polytypo/modes/runner.rb +33 -3
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +27 -12
- metadata +2 -1
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.8.0",
|
|
3
3
|
"locale": "fr",
|
|
4
4
|
"cases": [
|
|
5
5
|
{
|
|
@@ -747,6 +747,18 @@
|
|
|
747
747
|
"out": "---\ntitle: Une note ; suite\n---\n\n<Callout>Est-ce vrai ?</Callout>\n",
|
|
748
748
|
"note": "modes.md §3.7.3: the same guarantee under \"mdx\", where frontmatter plus JSX is the ordinary document shape. This is the case the polytypo-js source comment above its frontmatter skip entry described as a spec gap."
|
|
749
749
|
},
|
|
750
|
+
{
|
|
751
|
+
"id": "fr-markdown-commonmark-frontmatter-keys-colon",
|
|
752
|
+
"rule": "nbsp",
|
|
753
|
+
"mode": "markdown",
|
|
754
|
+
"dialect": "commonmark",
|
|
755
|
+
"frontmatterKeys": [
|
|
756
|
+
"title"
|
|
757
|
+
],
|
|
758
|
+
"in": "---\ntitle: \"Chapitre 1: le début\"\ndate: \"2026-09-25\"\n---\n\nVoici: le corps.\n",
|
|
759
|
+
"out": "---\ntitle: \"Chapitre 1 : le début\"\ndate: \"2026-09-25\"\n---\n\nVoici : le corps.\n",
|
|
760
|
+
"note": "modes.md §3.7.4 with the locale that made §3.7.3 skip the block in the first place: inside a listed quoted scalar a colon is content, so fr inserts its narrow no-break space, while the unlisted date field and the key itself keep every byte. That contrast is the whole argument for a caller-named list."
|
|
761
|
+
},
|
|
750
762
|
{
|
|
751
763
|
"id": "fr-nbsp-character-reference-numeric",
|
|
752
764
|
"rule": "nbsp",
|
|
@@ -936,6 +948,61 @@
|
|
|
936
948
|
"in": "Le «mot»'s résumé.",
|
|
937
949
|
"out": "Le « mot »’s résumé.",
|
|
938
950
|
"note": "The one configuration where a rule running AFTER apostrophe touches a CLOSEDELIM neighbour, so nbsp.md §5's I₆ discharge has a witness rather than only prose. N1/N2 put fr's inner no-break space on the glyph's INNER side, so the U+00BB that case 2a reads stays immediately left of the mark and the verdict is identical before and after the insertion. Synthetic rather than natural French — a possessive 's is not French — and deliberately so, in the standing quotes.md §6 gives its own adversarial de-CH witness: what is being pinned is the rule interaction, not an idiom."
|
|
951
|
+
},
|
|
952
|
+
{
|
|
953
|
+
"id": "fr-yaml-plain-elision",
|
|
954
|
+
"rule": "apostrophe",
|
|
955
|
+
"mode": "yaml",
|
|
956
|
+
"keys": [
|
|
957
|
+
"description"
|
|
958
|
+
],
|
|
959
|
+
"in": "description: l'idee et l'autre\n",
|
|
960
|
+
"out": "description: l’idee et l’autre\n",
|
|
961
|
+
"note": "modes.md §3.8.6 in a locale where the apostrophe is elision rather than possession: the same cell as en-us-yaml-plain-apostrophe, and §8 item 4's requirement that a construct be pinned in more than one locale rather than the same English output with the locale field changed."
|
|
962
|
+
},
|
|
963
|
+
{
|
|
964
|
+
"id": "fr-yaml-single-quoted-escape-untouched",
|
|
965
|
+
"rule": "apostrophe",
|
|
966
|
+
"mode": "yaml",
|
|
967
|
+
"keys": [
|
|
968
|
+
"description"
|
|
969
|
+
],
|
|
970
|
+
"in": "description: 'l''idee et l''autre'\n",
|
|
971
|
+
"out": "description: 'l''idee et l''autre'\n",
|
|
972
|
+
"note": "modes.md §3.8.6: two consecutive U+0027 in a single-quoted scalar yield no spans, in every locale — the source spells the content with more characters than it has, and that is a property of YAML, not of the language. This is the cell §7.13 measures at 1040 of 1858 values on a real corpus."
|
|
973
|
+
},
|
|
974
|
+
{
|
|
975
|
+
"id": "fr-yaml-single-quoted-no-escape",
|
|
976
|
+
"rule": "dashes",
|
|
977
|
+
"mode": "yaml",
|
|
978
|
+
"keys": [
|
|
979
|
+
"description"
|
|
980
|
+
],
|
|
981
|
+
"in": "description: 'deux -- trois'\n",
|
|
982
|
+
"out": "description: 'deux — trois'\n",
|
|
983
|
+
"note": "modes.md §3.8.6: a single-quoted scalar with nothing doubled inside it is claimed like any other. The style itself never disables the mode — only the doubling does, which is what the neighbouring case pins."
|
|
984
|
+
},
|
|
985
|
+
{
|
|
986
|
+
"id": "fr-markdown-commonmark-frontmatter-opener-trailing-space",
|
|
987
|
+
"rule": "nbsp",
|
|
988
|
+
"mode": "markdown",
|
|
989
|
+
"dialect": "commonmark",
|
|
990
|
+
"in": "--- \ntitle: Une note : suite\n---\n\nEst-ce vrai ?\n",
|
|
991
|
+
"out": "--- \ntitle: Une note : suite\n---\n\nEst-ce vrai ?\n",
|
|
992
|
+
"note": "modes.md §3.7.3a in the locale that makes the cost visible: without the trailing-space clause this block is prose, and `fr` puts a narrow no-break space in front of the colon of a machine-read field — the exact damage §3.7.3 exists to prevent, reached by one invisible character at the end of a fence."
|
|
993
|
+
},
|
|
994
|
+
{
|
|
995
|
+
"id": "fr-markdown-commonmark-frontmatter-keys-lone-cr-costs-one-line",
|
|
996
|
+
"rule": "quotes",
|
|
997
|
+
"mode": "markdown",
|
|
998
|
+
"dialect": "commonmark",
|
|
999
|
+
"frontmatterKeys": [
|
|
1000
|
+
"title",
|
|
1001
|
+
"slug"
|
|
1002
|
+
],
|
|
1003
|
+
"in": "---\ntitle: \"a\rb\"\nslug: c \"d\"\n---\n\nBody \"q\" here.\n",
|
|
1004
|
+
"out": "---\ntitle: \"a\rb\"\nslug: c « d »\n---\n\nBody « q » here.\n",
|
|
1005
|
+
"note": "modes.md §3.7.4 per-line U+000D bail in a locale whose quotation marks carry their own spacing, so the declined line and the converted one are visibly different constructions rather than the same English output relabelled (§8 item 4)."
|
|
939
1006
|
}
|
|
940
1007
|
]
|
|
941
1008
|
}
|
|
@@ -33,43 +33,62 @@
|
|
|
33
33
|
"nbsp": {
|
|
34
34
|
"beforePunctuation": [],
|
|
35
35
|
"narrowBeforePunctuation": [],
|
|
36
|
-
"afterShortWords": [
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
"k",
|
|
40
|
-
"o",
|
|
41
|
-
"s",
|
|
42
|
-
"u",
|
|
43
|
-
"v",
|
|
44
|
-
"z"
|
|
45
|
-
],
|
|
46
|
-
"abbreviations": [
|
|
47
|
-
"a. s.",
|
|
48
|
-
"s. r. o."
|
|
49
|
-
],
|
|
50
|
-
"beforeUnits": [
|
|
51
|
-
"%",
|
|
52
|
-
"°C",
|
|
53
|
-
"ha",
|
|
54
|
-
"kg",
|
|
55
|
-
"km",
|
|
56
|
-
"cm",
|
|
57
|
-
"mm",
|
|
58
|
-
"m²",
|
|
59
|
-
"kWh",
|
|
60
|
-
"km/h",
|
|
61
|
-
"str.",
|
|
62
|
-
"hod."
|
|
63
|
-
],
|
|
36
|
+
"afterShortWords": ["a", "i", "k", "o", "s", "u", "v", "z"],
|
|
37
|
+
"abbreviations": ["a. s.", "s. r. o."],
|
|
38
|
+
"beforeUnits": ["%", "°C", "ha", "kg", "km", "cm", "mm", "m²", "kWh", "km/h", "str.", "hod."],
|
|
64
39
|
"beforeNumber": [],
|
|
65
|
-
"beforeWord": [
|
|
66
|
-
|
|
67
|
-
"tzv.",
|
|
68
|
-
"tzn."
|
|
69
|
-
],
|
|
70
|
-
"afterSymbols": [
|
|
71
|
-
"§"
|
|
72
|
-
],
|
|
40
|
+
"beforeWord": ["tj.", "tzv.", "tzn."],
|
|
41
|
+
"afterSymbols": ["§"],
|
|
73
42
|
"initialBinding": "single"
|
|
74
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Uvozovky»: «V češtině užíváme různé varianty uvozovek: dvojité „ “, jednoduché ‚ ‘ a boční » «»; «Jako základní se doporučují uvozovky typu 99 66, tj. dvojité „ “»; «Ostatní typy uvozovek (v první řadě jednoduché, až v druhé řadě boční) se užívají především tehdy, pokud potřebujeme do uvozeného textu vložit ještě další uvozovky»; «Uvozovky přiléhají vždy těsně k výrazům, které ohraničují (tedy „takto“)»",
|
|
48
|
+
"url": "https://prirucka.ujc.cas.cz/?id=162",
|
|
49
|
+
"note": "První stupeň: U+201E otevírací, U+201C zavírací — NIKOLI U+201D jako v polštině, a právě to je rozdíl, který se v sazbě přehlédne. Kódy nebyly odečteny z markdownové konverze stránky, která se u uvozovek ukázala jako nespolehlivá; rozhodující je vlastní označení zdroje «uvozovky typu 99 66»: první znaménko má tvar devítky u paty řádku (U+201E), druhé tvar šestky vyzdvižený nahoru (U+201C). Polský pár PWN naproti tomu označuje jako 99 99. Druhý stupeň: jednoduché uvozovky, protože zdroj je výslovně řadí «v první řadě» pro uvozovky v uvozovkách. OMEZENÁ JISTOTA, řečená naplno: označení «99 66» je ve zdroji uvedeno pouze pro dvojité uvozovky, takže kódy jednoduchých (U+201A … U+2018) vycházejí ze stejného vzorce a z vysázených glyfů, ne z doslovného tvrzení. Rozhodla by tabulka H.1 normy ČSN 01 6910:2014, která je však placená (ČSN online od 1 000 Kč/rok). ROZHODNUTÍ OPERÁTORA, 18. 9. 2026: leží-li pravidlo za placenou normou, projekt je formuluje podle převládajícího úzu a označí to jako rozhodnutí, nikoli jako citaci — stejně jako u německého «Nr.». Padne-li placená stěna, nahradí znění normy tento řádek. Q-W platí v obou případech: U+201A, U+2018 i U+2019 patří do NARROW. Ověřeno 18. 9. 2026."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Pomlčka»: «Pomlčku oddělujeme z obou stran mezerami, komplikovanější situace je pouze tehdy, pokud toto znaménko vystupuje ve funkci výrazů a, až, od … do … nebo proti (versus)»; «pomlčka místo čárky: Tato kniha – vydaná ještě před válkou – je opravdu úžasná»; «Pomlčka (–) je dlouhá vodorovná čárka»; «Vedle krátké pomlčky (–) se v textu někdy používá také pomlčka dlouhá (—)»",
|
|
54
|
+
"url": "https://prirucka.ujc.cas.cz/?id=165",
|
|
55
|
+
"note": "Odůvodňuje dash.parenthetical = «en-spaced». Vsuvka není žádná ze čtyř vyjmenovaných výjimek, takže platí obecné pravidlo o mezerách z obou stran. DÉLKA je na rozdíl od polštiny rozhodnuta samotným zdrojem: definiční glyf stránky je krátká pomlčka U+2013 a dlouhá U+2014 je popsána jako to, co se používá «někdy» navíc. U polštiny musel o délce rozhodnout operátor, protože PWN si v korpusu protiřečil; tady stačí číst. PŘÍPUSTNÁ VARIANTA, zaznamenána: U+2014. Ověřeno 18. 9. 2026."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "ranges",
|
|
59
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Pomlčka», vyjádření rozsahu (s významem ‚až‘ či ‚od … do‘): «strana 23–26, v letech 1945–1948, 9–16 h»; ÚJČ (P. Lozan, M. Pravdová), «Otázky a odpovědi k ČSN 01 6910 (2014)», otázka 2: «norma umožňuje i psaní mezery kolem pomlčky v rozsazích a ve významu „a“ a „versus“, jestliže je alespoň jeden z výrazů kolem pomlčky víceslovný»",
|
|
60
|
+
"url": "https://prirucka.ujc.cas.cz/?id=165",
|
|
61
|
+
"note": "Odůvodňuje dash.range = «en-tight»: krátká pomlčka U+2013 bez mezer. Potvrzeno ze dvou stran: příručka sází rozsahy těsně, a dokument ÚJČ k normě popisuje mezery v rozsahu jako povolenou VÝJIMKU pro víceslovné výrazy, což těsnou sazbu předpokládá jako pravidlo. Text samotné ČSN 01 6910 čten nebyl (je placený); citován je výklad jejích vlastních zpracovatelů, což je nejsilnější veřejně dostupná náhrada. Stránka rovněž uvádí, že pomlčka vyjadřující rozsah nemá stát na konci ani na začátku řádku — požadavek, pro který locale.schema.json nemá pole. Pravidlo ranges je ve výchozím stavu vypnuté (spec 0.5.0). Ověřeno 18. 9. 2026."
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"rule": "ellipsis",
|
|
65
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Tři tečky»: «jedná se totiž o jeden znak, nikoli tři samostatné tečky»; «Za třemi tečkami často následují další interpunkční znaménka, která se za tři tečky připojují bez mezery, ale nikdy se však nepřipojuje další tečka (nikoli tedy čtyři tečky vedle sebe)»",
|
|
66
|
+
"url": "https://prirucka.ujc.cas.cz/?id=166",
|
|
67
|
+
"note": "Odůvodňuje abbreviatedAfterTerminal = false. Výpustka je jeden znak U+2026 a otazník či vykřičník stojí ZA ní — ruská dvoutečková forma «?..» není v žádném konzultovaném českém zdroji předvídána. Bylo to ověřeno záměrně: ukrajinština, zkoumaná v témže průchodu, tu formu MÁ, takže analogie mezi slovanskými jazyky by tu byla svůdná a mylná. Ověřeno 18. 9. 2026."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"rule": "hyphen",
|
|
71
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Spojovník»: «píše se bez mezer mezi výrazy, které spojuje»; «Spojovník na začátku dalšího řádku neopakujeme, naznačuje-li rozdělení slova (např. žong- | -lér)»; táž příručka, «Zalomení řádků a nevhodné výrazy na jejich konci», jejíž uzavřený výčet míst, kde k zalomení dojít nemá, spojovník neuvádí",
|
|
72
|
+
"url": "https://prirucka.ujc.cas.cz/?id=164",
|
|
73
|
+
"note": "Odůvodňuje tři prázdné seznamy. Čeština spojovník na konci řádku ani nezakazuje, ani — na rozdíl od polštiny — dělení v jeho místě nepředepisuje; neexistuje uzavřený normativní seznam tvarů, jejichž spojovník se nesmí lámat. To, že jej třináctibodový výčet stránky o zalomení řádků neobsahuje, je doložená nepřítomnost, ne nenalezená citace — a je to zároveň rozdíl proti ukrajinštině, kde § 64 п. 4 takový seznam má. Ověřeno 18. 9. 2026."
|
|
74
|
+
},
|
|
75
|
+
{
|
|
76
|
+
"rule": "nbsp",
|
|
77
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Zalomení řádků a nevhodné výrazy na jejich konci»: k zalomení nemá dojít «ve spojení neslabičných předložek k, s, v, z s následujícím slovem, např. k mostu, s bratrem, v Plzni, z nádraží»; «ve spojení slabičných předložek o, u a spojek a, i s výrazem, který po nich následuje, např. u babičky, o páté»; «ve složených zkratkách a kódech, např. a. s., s. r. o.»; «mezi zkratkami tj., tzv., tzn. a následujícím výrazem»; «mezi zkratkou jména a příjmením»; «mezi číslem a značkou, např. 50 %, § 23»; «V textovém procesoru vkládáme tam, kde nemá dojít k zalomení řádku, místo běžné mezislovní mezery mezeru pevnou»",
|
|
78
|
+
"url": "https://prirucka.ujc.cas.cz/?id=880",
|
|
79
|
+
"note": "Pět polí z jednoho uzavřeného výčtu, který se odvolává na ČSN 01 6910. (1) afterShortWords: osm jednopísmenných výrazů, citovaných doslova. Na rozdíl od polštiny je zákaz v prameni BEZPODMÍNEČNÝ — není vázán na šířku sloupce ani na to, zda jde o titulek — takže zde nebylo co rozhodovat. (2) abbreviations: «a. s.» a «s. r. o.». Interakce s N3 ověřena a je opačná než u polského «i in.»: N3 vyžaduje, aby po jednopísmenném výrazu následovala mezera, a po «a» následuje tečka, takže N3 index nenárokuje a obě mezery dostane N4. «ČSN 01 6910» z výčtu vypuštěno: číslo normy není fakt o jazyce. (3) beforeWord: «tj.», «tzv.», «tzn.». Bod o «zkratce titulu» populován není, protože zdroj u něj žádný seznam neuvádí. (4) afterSymbols = [«§»]. Ze stejného citovaného bodu byly ZÁMĚRNĚ vynechány «*», «†» a «#»: řetězec «* 1921» je bajt po bajtu totožný s odrážkou seznamu následovanou číslicí, takže v režimu markdown by šlo o falešný zásah bez vyvažujícího přínosu. (5) initialBinding = «single»: pramen váže «zkratku jména» k příjmení a jeho vlastní příklady jsou Fr. Daneš a M. Těšitelová, tedy JEDNA zkrácená křestní jméno; «chain» by na ani jednom z nich nezabral a citované pravidlo by zůstalo nenaplněné. Expozici, kterou schéma u «single» samo pojmenovává, přijímáme na základě citace, stejně jako u fr. beforeNumber zůstává prázdné: jediným kandidátem je «strana 2», ale «strana» je víceznačné plnovýznamové slovo, nikoli zkratka — týž argument, jakým ru.json vylučuje «г.». Ověřeno 18. 9. 2026."
|
|
80
|
+
},
|
|
81
|
+
{
|
|
82
|
+
"rule": "nbsp",
|
|
83
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Značky, čísla a číslice»: «Značky se od číselné hodnoty oddělují mezerou, číslo a značka se umísťují na stejný řádek», «např. 10 ha = 10 hektarů, 3 kg = 3 kilogramy, 14 % = 14 procent», «100 kWh», «rychlost 50 km/h», «teplota 12–15 °C»; BIPM, The International System of Units (SI), 9th ed., concise summary: «A single space is always left between the number and the unit»",
|
|
84
|
+
"url": "https://prirucka.ujc.cas.cz/?id=785",
|
|
85
|
+
"note": "Rozdělení rolí jako v en-US, fi, sv, es, pt a pl: BIPM říká, které řetězce jsou značkami jednotek SI, ÚJČ říká, jaká je česká konvence. České znění je přitom silnější, než nbsp.md §2.1 od locale vyžaduje: uvádí nejen mezeru, ale i to, že «číslo a značka se umísťují na stejný řádek». PŘÍPUSTNÁ VARIANTA, zaznamenána: «V případě, že pomocí číslice a značky vyjadřujeme přídavné jméno, mezeru nevkládáme: 8km = 8kilometrový, 20% = 20procentní». Rozpor není třeba řešit: N5 existující mezeru pouze nahrazuje a nikdy ji nevkládá (nbsp.md §3.7), takže «14 %» dostane U+00A0 a «20%» zůstane nedotčeno. Jednoznakové značky vypuštěny, a pro češtinu z konkrétního důvodu, ne ze stylové opatrnosti: «s» je zároveň vyjmenovaná jednopísmenná předložka, takže ve větě «Přišel v 5 s bratrem» by se «5» svázala se «s». Ověřeno 18. 9. 2026."
|
|
86
|
+
},
|
|
87
|
+
{
|
|
88
|
+
"rule": "nbsp",
|
|
89
|
+
"cite": "Ústav pro jazyk český AV ČR, Internetová jazyková příručka, «Otazník»: «Stejně jako naprostá většina interpunkčních znamének se připojuje k předcházejícímu slovu (zkratce, značce) bez mezery, za ním následuje mezera»; táž příručka, «Tři tečky»: «tři tečky následující za slovem se připojují bez mezery»",
|
|
90
|
+
"url": "https://prirucka.ujc.cas.cz/?id=171",
|
|
91
|
+
"note": "Odůvodňuje prázdné beforePunctuation i narrowBeforePunctuation. Čeština před «:», «;», «!», «?» nestaví ani U+00A0, ani U+202F, a zdroj jim odstup výslovně upírá — je to zápor, ne mlčení. Q-P je tím splněno triviálně. Ověřeno 18. 9. 2026."
|
|
92
|
+
}
|
|
93
|
+
]
|
|
75
94
|
}
|
|
@@ -34,39 +34,48 @@
|
|
|
34
34
|
"beforePunctuation": [],
|
|
35
35
|
"narrowBeforePunctuation": [],
|
|
36
36
|
"afterShortWords": [],
|
|
37
|
-
"abbreviations": [
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
"m. a. W.",
|
|
43
|
-
"m. w. H."
|
|
44
|
-
],
|
|
45
|
-
"beforeUnits": [
|
|
46
|
-
"%",
|
|
47
|
-
"‰",
|
|
48
|
-
"°C",
|
|
49
|
-
"km",
|
|
50
|
-
"cm",
|
|
51
|
-
"mm",
|
|
52
|
-
"kg",
|
|
53
|
-
"km/h",
|
|
54
|
-
"kWh"
|
|
55
|
-
],
|
|
56
|
-
"beforeNumber": [
|
|
57
|
-
"Art.",
|
|
58
|
-
"Abs.",
|
|
59
|
-
"Ziff.",
|
|
60
|
-
"Kap.",
|
|
61
|
-
"S."
|
|
62
|
-
],
|
|
63
|
-
"beforeWord": [
|
|
64
|
-
"St."
|
|
65
|
-
],
|
|
66
|
-
"afterSymbols": [
|
|
67
|
-
"§",
|
|
68
|
-
"§§"
|
|
69
|
-
],
|
|
37
|
+
"abbreviations": ["z. B.", "d. h.", "u. a.", "u. U.", "m. a. W.", "m. w. H."],
|
|
38
|
+
"beforeUnits": ["%", "‰", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
39
|
+
"beforeNumber": ["Art.", "Abs.", "Ziff.", "Kap.", "S."],
|
|
40
|
+
"beforeWord": ["St."],
|
|
41
|
+
"afterSymbols": ["§", "§§"],
|
|
70
42
|
"initialBinding": "chain"
|
|
71
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 201–203: zu schreiben sind die Guillemets « » und, für eine Anführung innerhalb einer Anführung, die halben Anführungszeichen ‹ ›; die in Deutschland übliche Schreibung » « ist für amtliche Texte nicht zulässig",
|
|
48
|
+
"url": "https://www.bk.admin.ch/dam/bk/de/dokumente/sprachdienste/sprachdienst_de/schreibweisungen.pdf.download.pdf/schreibweisungen.pdf",
|
|
49
|
+
"note": "Rz. 202 regelt ausdrücklich nur das Leerzeichen VOR dem Anführungszeichen und NACH dem Schlusszeichen; ein Zwischenraum innerhalb der Guillemets wird weder gefordert noch in einem der Beispiele gesetzt («Zukunft für Schweizer Fahrende»). Daher innerSpace = none."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 231, 237, 238: der Gedankenstrich ist der Halbgeviertstrich –; vor und nach dem Gedankenstrich bei Nachträgen und paarigen Einschüben steht ein Leerzeichen",
|
|
54
|
+
"url": "https://www.bk.admin.ch/dam/bk/de/dokumente/sprachdienste/sprachdienst_de/schreibweisungen.pdf.download.pdf/schreibweisungen.pdf",
|
|
55
|
+
"note": "Rz. 236 hält ausdrücklich fest, dass der Geviertstrich (—) nicht zulässig ist."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "dashes",
|
|
59
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 234: der Gedankenstrich als Begriffszeichen für „bis“ steht ohne Leerzeichen — „16–17 Uhr“, „die Artikel 10–12“, „die Jahre 1939–1945“",
|
|
60
|
+
"url": "https://www.bk.admin.ch/dam/bk/de/dokumente/sprachdienste/sprachdienst_de/schreibweisungen.pdf.download.pdf/schreibweisungen.pdf"
|
|
61
|
+
},
|
|
62
|
+
{
|
|
63
|
+
"rule": "nbsp",
|
|
64
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 251 (Festabstand hält Zusammengehöriges zusammen: Ziffer und Masseinheit, Abkürzung und Ortsname, Teile mehrgliedriger Abkürzungen, Gliederungseinheit und Ziffer — „20 km“, „St. Gallen“, „Artikel 35“), Rz. 438 (mehrgliedrige Abkürzungen: d. h. / z. B. / u. a. / u. U. / m. a. W. / m. w. H.), Rz. 549–550 (Festabstand zwischen Zahl und Einheit), Rz. 554 (Festabstand vor % und ‰)",
|
|
65
|
+
"url": "https://www.bk.admin.ch/dam/bk/de/dokumente/sprachdienste/sprachdienst_de/schreibweisungen.pdf.download.pdf/schreibweisungen.pdf",
|
|
66
|
+
"note": "initialBinding: \"chain\" (spec 0.6.0: das Feld hieß zuvor das boolesche bindInitials) ist eine Ableitung, keine wörtliche Weisung: Rz. 251 nennt „eine Abkürzung und ein Ortsname“ (St. Gallen) und Rz. 438 die Teile mehrgliedriger Abkürzungen; die Bindung Initiale–Nachname ist dieselbe Konstruktion, wird aber in den Schreibweisungen nicht ausdrücklich erwähnt. Keine Quelle belegt eine einzelne Initiale (\"single\"); \"chain\" folgt derselben Zwei-oder-mehr-Logik wie de-DE/ru. Die Einheitenliste ist bewusst kurz und enthält keine einbuchstabigen Einheitenzeichen. Der Schweizer Verzicht auf ß (ss statt ß, Amtliches Regelwerk § 25 E2) lässt sich im Schema nicht ausdrücken und ist deshalb hier nicht abgebildet."
|
|
67
|
+
},
|
|
68
|
+
{
|
|
69
|
+
"rule": "nbsp",
|
|
70
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 251 (Festabstand zwischen einer Gliederungseinheit eines Erlasses und der dazugehörigen Ziffer — Beispiele „Artikel 35“, „S. 17 f.“, „St. Gallen“), Rz. 438 (Entsprechendes gilt für die eingliedrige Abkürzung „St.“ für „Sankt“ in Ortsnamen: „St. Gallen“, „St. Moritz“), Rz. 439 (abschliessende Aufzählung der Begriffszeichen: Ziffern, %, ‰, Gedankenstrich, Schrägstrich, §, Währungszeichen, mathematische Zeichen, Einheitenzeichen), Rz. 440 (zwischen einem Begriffszeichen und der dazugehörigen Zahl steht ein Festabstand), Rz. 727 (in verknapptem Text werden die Gliederungseinheiten abgekürzt: „Kap., Art., Abs., Bst. (falsch: Buchst., lit.), Ziff.“), Rz. 730 (Nummerierung der Gliederungseinheiten)",
|
|
71
|
+
"url": "https://www.bk.admin.ch/dam/de/sd-web/YVHazXZRkqKn/schreibweisungen.pdf",
|
|
72
|
+
"note": "Wortlaut im PDF selbst geprüft. Rz. 439 und Rz. 727 entscheiden die Zuordnung von „Art.“: die Aufzählung der Begriffszeichen in Rz. 439 ist abschliessend und enthält „Art.“ nicht, Rz. 727 führt „Art.“ ausdrücklich als Abkürzung einer Gliederungseinheit. „Art.“ gehört damit zu beforeNumber und nicht zu afterSymbols, wo es zuvor stand; die Bindung an die folgende Zahl folgt aus Rz. 251. Aus derselben Liste fehlt nur „Bst.“, weil ihm ein Buchstabe folgt, keine Zahl („Bst. a“). „S.“ ist mit „S. 17 f.“ in Rz. 251 wörtlich belegt. „Nr.“ wurde ersatzlos entfernt: es steht weder in Rz. 439 noch überhaupt im Sachregister (unter N nur „NGO“, „Normen“, „Null“); die einzige Fundstelle ist Rz. 431, wo „Nr. (Nummer), Tarif-Nrn.“ als Beispiel für Deklinationsendungen dient und keine Abstandsregel ausspricht. afterSymbols war damit nachweislich falsch, und für beforeNumber fehlt der Beleg — nach docs/PLAN.md §6.1 zieht das die Streichung nach sich. Zurückkommen kann „Nr.“ mit einer Lesung der DIN 5008. beforeWord enthält nur „St.“: Rz. 438 nennt genau diesen Fall wörtlich, weitere Abkürzung-plus-Wort-Bindungen sind in den Schreibweisungen nicht als geschlossene Liste geregelt („Küssnacht a. R.“ ist eine mehrgliedrige Abkürzung, kein Präfix)."
|
|
73
|
+
},
|
|
74
|
+
{
|
|
75
|
+
"rule": "hyphen",
|
|
76
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 223 und 224: der Bindestrich wird als Ergänzungsstrich für einen eingesparten Wortteil verwendet („Papierproduktion und -handel“); mit dem geschützten Ergänzungsstrich wird verhindert, dass der Ergänzungsstrich am Zeilenende auf der oberen Zeile bleibt, während der zugehörige zweite Wortbestandteil auf die nächste Zeile rutscht",
|
|
77
|
+
"url": "https://www.bk.admin.ch/dam/de/sd-web/YVHazXZRkqKn/schreibweisungen.pdf",
|
|
78
|
+
"note": "Beleg dafür, dass die drei Listen leer bleiben. Das Deutsche kennt keine geschlossene Liste von Morphemen mit unteilbarem Bindestrich; die Trennung am Bindestrich ist zulässig (Rz. 269–271 behandeln nur sinnentstellende Trennungen). Der einzige belegte Fall eines geschützten Bindestrichs ist der Ergänzungsstrich (Rz. 224), und der ist offen: er betrifft jedes beliebige Zweitglied nach „und -“ und lässt sich als literale Wortliste nicht abbilden."
|
|
79
|
+
}
|
|
80
|
+
]
|
|
72
81
|
}
|
|
@@ -34,37 +34,52 @@
|
|
|
34
34
|
"beforePunctuation": [],
|
|
35
35
|
"narrowBeforePunctuation": [],
|
|
36
36
|
"afterShortWords": [],
|
|
37
|
-
"abbreviations": [
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
"z. T.",
|
|
43
|
-
"i. d. R."
|
|
44
|
-
],
|
|
45
|
-
"beforeUnits": [
|
|
46
|
-
"%",
|
|
47
|
-
"‰",
|
|
48
|
-
"€",
|
|
49
|
-
"°C",
|
|
50
|
-
"km",
|
|
51
|
-
"cm",
|
|
52
|
-
"mm",
|
|
53
|
-
"kg",
|
|
54
|
-
"km/h",
|
|
55
|
-
"kWh"
|
|
56
|
-
],
|
|
57
|
-
"beforeNumber": [
|
|
58
|
-
"Nr.",
|
|
59
|
-
"S."
|
|
60
|
-
],
|
|
61
|
-
"beforeWord": [
|
|
62
|
-
"St."
|
|
63
|
-
],
|
|
64
|
-
"afterSymbols": [
|
|
65
|
-
"§",
|
|
66
|
-
"§§"
|
|
67
|
-
],
|
|
37
|
+
"abbreviations": ["z. B.", "d. h.", "u. a.", "u. Ä.", "z. T.", "i. d. R."],
|
|
38
|
+
"beforeUnits": ["%", "‰", "€", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
39
|
+
"beforeNumber": ["Nr.", "S."],
|
|
40
|
+
"beforeWord": ["St."],
|
|
41
|
+
"afterSymbols": ["§", "§§"],
|
|
68
42
|
"initialBinding": "chain"
|
|
69
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Duden, Rechtschreibregeln, „Anführungszeichen“, Regeln D 5 und D 12: Gänsefüßchen „…“ als Anführungszeichen, halbe Anführungszeichen ‚…‘ für eine Anführung innerhalb einer Anführung",
|
|
48
|
+
"url": "https://www.duden.de/sprachwissen/rechtschreibregeln/anfuehrungszeichen",
|
|
49
|
+
"note": "Deckungsgleich mit dem Amtlichen Regelwerk der deutschen Rechtschreibung, § 79 E2 (Anführung innerhalb einer Anführung durch halbe Anführungszeichen)."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Duden, Rechtschreibregeln, „Gedankenstrich“, Regel D 45: der Gedankenstrich (Halbgeviertstrich) steht mit Leerzeichen auf beiden Seiten beim Einschieben eines Zusatzes",
|
|
54
|
+
"url": "https://www.duden.de/sprachwissen/rechtschreibregeln/gedankenstrich"
|
|
55
|
+
},
|
|
56
|
+
{
|
|
57
|
+
"rule": "dashes",
|
|
58
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen (Ausgabe 2013, aktualisiert), Rz. 234: der Gedankenstrich als Begriffszeichen für „bis“ steht ohne Leerzeichen — „die Jahre 1939–1945“",
|
|
59
|
+
"url": "https://www.bk.admin.ch/dam/bk/de/dokumente/sprachdienste/sprachdienst_de/schreibweisungen.pdf.download.pdf/schreibweisungen.pdf",
|
|
60
|
+
"note": "Für den Bis-Strich konnte keine gleichwertig präzise Duden-Onlineregel gefunden werden; die schweizerische Weisung formuliert dieselbe im gesamten deutschen Sprachraum übliche Regel und wird hier als überprüfbare Quelle angegeben."
|
|
61
|
+
},
|
|
62
|
+
{
|
|
63
|
+
"rule": "nbsp",
|
|
64
|
+
"cite": "DIN 5008 (Schreib- und Gestaltungsregeln für die Text- und Informationsverarbeitung): geschütztes Leerzeichen zwischen den Teilen mehrgliedriger Abkürzungen („z. B.“, „d. h.“), bei Initialen („J. K. Rowling“) sowie zwischen Zahl und Einheit, Prozentzeichen und Paragrafenzeichen",
|
|
65
|
+
"note": "Die Norm selbst ist kostenpflichtig; die Regelinhalte wurden über die Duden-Regeln zu Abkürzungen (D 1) und über die von Duden selbst verwendete Schreibung „z. B.“ mit geschütztem Leerzeichen gegengeprüft. Der Einheitenliste liegt keine Normliste zugrunde: sie ist bewusst kurz gehalten und enthält keine einbuchstabigen Einheitenzeichen (m, g, l, s), weil diese ohne Kontextprüfung zu Falschtreffern führen."
|
|
66
|
+
},
|
|
67
|
+
{
|
|
68
|
+
"rule": "nbsp",
|
|
69
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 251: der Festabstand verhindert, dass Zusammengehöriges beim Zeilensprung auseinandergerissen wird — Beispiele „20 km“, „St. Gallen“, „Artikel 35“, „S. 17 f.“, „20. November“; Rz. 438: Entsprechendes gilt für die eingliedrige Abkürzung „St.“ für „Sankt“ in Ortsnamen („St. Gallen“, „St. Moritz“); Rz. 439: abschliessende Aufzählung der Begriffszeichen (Ziffern, %, ‰, Gedankenstrich, Schrägstrich, §, Währungszeichen, mathematische Zeichen, Einheitenzeichen)",
|
|
70
|
+
"url": "https://www.bk.admin.ch/dam/de/sd-web/YVHazXZRkqKn/schreibweisungen.pdf",
|
|
71
|
+
"note": "Wortlaut im PDF selbst geprüft. Für Deutschland regelt dieselbe Konstruktion die DIN 5008 (Schreib- und Gestaltungsregeln), die kostenpflichtig ist und deren Wortlaut hier nicht geprüft werden konnte; deshalb steht die schweizerische Weisung als überprüfbarer Beleg — dasselbe Vorgehen wie bei der Quelle zu Rz. 234 in diesem File. Die Übertragung auf de-DE ist insofern eine Ableitung, keine deutschlandspezifische Weisung. „S.“ ist mit „S. 17 f.“ wörtlich belegt und deckt die von docs/PLAN.md §7 geforderte Schreibung „S.“ + Zahl. „Nr.“ gehört nicht hierher, sondern zu den Abkürzungen: die Aufzählung der Begriffszeichen in Rz. 439 ist abschliessend und enthält „Nr.“ nicht (im Sachregister unter N stehen nur „NGO“, „Normen“, „Null“). Aus afterSymbols ist es deshalb ersatzlos gestrichen. In beforeNumber steht es seit Spec 1.3.0 — nicht als Zitat, sondern als vereinheitlichte Schreibung; die Begründung steht im nächsten Eintrag."
|
|
72
|
+
},
|
|
73
|
+
{
|
|
74
|
+
"rule": "hyphen",
|
|
75
|
+
"cite": "Schweizerische Bundeskanzlei, Schreibweisungen, Rz. 223 und 224: der Bindestrich wird als Ergänzungsstrich für einen eingesparten Wortteil verwendet („Papierproduktion und -handel“); mit dem geschützten Ergänzungsstrich wird verhindert, dass der Ergänzungsstrich am Zeilenende auf der oberen Zeile bleibt, während der zugehörige zweite Wortbestandteil auf die nächste Zeile rutscht",
|
|
76
|
+
"url": "https://www.bk.admin.ch/dam/de/sd-web/YVHazXZRkqKn/schreibweisungen.pdf",
|
|
77
|
+
"note": "Beleg dafür, dass die drei Listen leer bleiben. Das Deutsche kennt keine geschlossene Liste von Morphemen mit unteilbarem Bindestrich; die Trennung am Bindestrich ist zulässig. Der einzige belegte Fall eines geschützten Bindestrichs ist der Ergänzungsstrich (Rz. 224), und der ist offen: er betrifft jedes beliebige Zweitglied nach „und -“ und lässt sich als literale Wortliste nicht abbilden."
|
|
78
|
+
},
|
|
79
|
+
{
|
|
80
|
+
"rule": "nbsp",
|
|
81
|
+
"cite": "Keine frei prüfbare Normquelle für „Nr.“ + Zahl: DIN 5008:2020-03 ist kostenpflichtig und wurde nicht gelesen; Duden und das Amtliche Regelwerk 2024 sprechen keine Abstandsregel aus. Vereinheitlicht nach vorherrschendem Gebrauch (Operatorentscheid, 18.09.2026)",
|
|
82
|
+
"note": "Diese Zeile ist ausdrücklich kein Beleg, sondern eine Festlegung, und wird als solche geführt. Geprüft und ergebnislos: Duden, Rechtschreibregeln „Abkürzungen“ (D 1–D 4, nur der Abkürzungspunkt); Duden, Sprachratgeber „Worttrennung am Zeilenende“; Amtliches Regelwerk 2024, Kapitel F (§§ 84–90, nur Worttrennung) und der Kasten „Sonderzeichen“ auf S. 153, der Typografie ausdrücklich den Konventionen bzw. den DIN-, ÖNORM- und SNV-Normen zuweist, also ausserhalb des Regelwerks. Entscheiden könnte nur DIN 5008:2020-03. Der Projektentscheid dazu: liegt eine Regel hinter einer Bezahlschranke, formuliert polytypo sie nach dem verbreitetsten Gebrauch und hält sie in allen Runtimes gleich — Einheitlichkeit geht hier vor Kanontreue, und die Festlegung wird offen als solche gekennzeichnet statt als Zitat ausgegeben. Der verbreitetste Gebrauch bindet „Nr.“ an die folgende Zahl; die Sekundärquellen, die DIN 5008 referieren, sagen dasselbe. Fällt die Schranke, tritt der Wortlaut der Norm an die Stelle dieser Zeile — auch wenn er ihr widerspricht."
|
|
83
|
+
}
|
|
84
|
+
]
|
|
70
85
|
}
|
|
@@ -40,5 +40,55 @@
|
|
|
40
40
|
"beforeWord": [],
|
|
41
41
|
"afterSymbols": [],
|
|
42
42
|
"initialBinding": "none"
|
|
43
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Υπηρεσία Εκδόσεων της Ευρωπαϊκής Ένωσης, Διοργανικό εγχειρίδιο σύνταξης κειμένων (ελληνική έκδοση 2011, τελευταία ενημέρωση 30.4.2012), Μέρος Τέταρτο «Συμβατικοί κανόνες για την ελληνική γλώσσα», §10.1.7 «Εισαγωγικά»: «Σε εισαγωγικά (στο ελληνικό κείμενο προτιμώνται τα διπλά γωνιώδη εισαγωγικά: « ») κλείνονται κυρίως λόγια ή παραθέματα που αναφέρονται αυτολεξεί. […] Στην περίπτωση που χρειάζονται εισαγωγικά μέσα σε κείμενο που είναι ήδη σε εισαγωγικά, τότε για τα εισαγωγικά αυτά χρησιμοποιούνται τα διπλά ανωφερή εισαγωγικά (“ ”), ενώ σε τρίτο επίπεδο εσωτερικά χρησιμοποιούνται τα μονά ανωφερή εισαγωγικά (‘ ’)»",
|
|
48
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
49
|
+
"note": "Primary U+00AB/U+00BB, secondary U+201C/U+201D. Part Four of this guide is a Greek-specific punctuation chapter, not the multilingual part — the distinction matters and is the reason this source is cited for Greek at all. innerSpace is \"none\": the §6.4 spacing table sets «xx» with an ordinary space only outside the guillemets, and the Greek chapter never asks for one inside. The third nesting level the source describes (U+2018/U+2019) is not expressible in locale.schema.json, which has primary and secondary only; that is a schema limit, not a gap in the source. Corroborated by ΥΠΕΠΘ/ΙΤΥΕ, Γραμματική Νέας Ελληνικής Γλώσσας Α΄–Γ΄ Γυμνασίου, §3.3 («Τα εισαγωγικά ( « » ) σημειώνονται…»), whose examples likewise carry no inner space. Retrieved 2026-08-15."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.8 «Ενωτικό — Παύλα μεσαίου μεγέθους — Μεγάλη παύλα»: για τα αριθμητικά διαστήματα, «Παύλα μεσαίου μεγέθους ή μείον (–) […] γ) για να δηλώσει το διάστημα μεταξύ δύο ορίων (π.χ. άτομα ηλικίας 25–45 ετών) […] δ) στην αναφορά σε περιόδους τουλάχιστον δύο πλήρων ετών (π.χ. 1989–1991) […] Ωστόσο, και στην περίπτωση αυτή μπορεί να χρησιμοποιηθεί και ενωτικό». Ο ίδιος κανόνας για τα αριθμητικά διαστήματα εμφανίζεται ΔΥΟ ΦΟΡΕΣ στην §10.1.8, μία στην υποενότητα «Ενωτικό (-)» και μία στην υποενότητα «Παύλα μεσαίου μεγέθους (–)», και κάθε φορά η υποενότητα παραχωρεί ρητά το άλλο σημείο (paraphrase of the «Ενωτικό» occurrence — its verbatim text was not captured at review time; the «Παύλα» occurrence is quoted above verbatim)· για την παρενθετική χρήση, «Όπως η παρένθεση, η διπλή παύλα δεν χωρίζεται με κενά διαστήματα από τη λέξη, φράση ή πρόταση που περικλείει· αντίθετα, μπαίνουν διαστήματα πριν από την πρώτη και μετά τη δεύτερη παύλα»",
|
|
54
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
55
|
+
"note": "Justifies \"none\" on BOTH fields, which no other locale has — but for TWO DIFFERENT REASONS, and conflating them would misrepresent the source. Range: genuinely ambivalent. The rule appears twice in §10.1.8, once under «Ενωτικό (-)» and once under «Παύλα μεσαίου μεγέθους (–)», and each occurrence explicitly concedes the other mark. That mutual concession is the proof — a source that names both forms in both places is not being vague, it is declining to rank them, and substituting either would express a preference the citation does not carry. Parenthetical: NOT ambivalent. The guide prescribes a U+2014 pair (its own header gives «Alt 0151») with ordinary spaces outside the pair and none on the inner edges. \"none\" here is forced by a SCHEMA LIMIT, not by the source: locale.schema.json's dash enum has no value for asymmetric spacing — \"em-spaced\" puts a space on each side of each dash, which is precisely what this sentence forbids. See spec/rules/dashes.md §6 «el», which states the two justifications separately. SOURCE CONFLICT, recorded and deliberately not settled: this guide prescribes U+2014 for the parenthetical dash, while ΥΠΕΠΘ/ΙΤΥΕ Γραμματική Α΄–Γ΄ Γυμνασίου §3.3 appears to set U+2013 in the same role. The sources array has no way to represent a conflict — it models agreement, not disagreement — so it is recorded here in prose. Resolving it is with the operator and needs a source that ranks the two, not a third that adds a form. Retrieved 2026-08-15."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "ellipsis",
|
|
59
|
+
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.9 «Αποσιωπητικά»: «Τα αποσιωπητικά, που είναι πάντοτε τρεις (και όχι περισσότερες) τελείες, χρησιμοποιούνται κυρίως: […]»· παρατήρηση ii): «Μεταξύ των αποσιωπητικών και της λέξης που προηγείται δεν αφήνουμε διάστημα»· παρατήρηση iii): «Όταν τα αποσιωπητικά βρίσκονται στο τέλος της περιόδου δεν προσθέτουμε τελεία»",
|
|
60
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
61
|
+
"note": "Justifies abbreviatedAfterTerminal = false. The source requires exactly three dots always and nowhere provides for the two-dot form after «?» or «!» that Russian uses, so the Russian branch of spec/rules/ellipsis.md must stay off for Greek. Corroborated by Γραμματική Α΄–Γ΄ Γυμνασίου §3.3: «Οι σημειούμενες τελείες είναι πάντα τρεις». Neither source gives any rule for αποσιωπητικά meeting the ερωτηματικό or the θαυμαστικό — checked specifically, in the Ministry of Education grammar and in the Κέντρο Ελληνικής Γλώσσας materials — so the engine's TERMINAL class is deliberately left as {U+0021, U+003F} and the Greek U+003B is not added to it (spaces.md §7.8; adding it would also regress Russian, where Лопатин §154 ties the two-dot form to «?» and «!» by name). Neither source addresses U+2026 versus three U+002E; that choice is the engine's, made identically in every locale, and is not claimed here. KNOWN DIVERGENCE, recorded because a verified source and the engine disagree and neither may stand unremarked: the same §10.1.9, παρατήρηση ii), states «Μεταξύ των αποσιωπητικών και της λέξης που προηγείται δεν αφήνουμε διάστημα» - no space between the ellipsis and the preceding word - and polytypo does NOT honour it. spec/rules/spaces.md §3.4 preserves a space before a dot run in EVERY locale, so «Πράγματι …» is returned as typed. The divergence is deliberate and argued in spaces.md §7.9 and ellipsis.md §6: honouring it needs either locale data in a rule that has none by design, or an ellipsis rule that deletes rather than replaces, which changes that rule's kind and forces a new composition argument. It would be revisited if a second locale wanted the same behaviour, at which point the shape is an ellipsis.noSpaceBefore flag consumed by the ellipsis rule. The source is right about Greek; the engine is declining to act on it, not disputing it. Retrieved 2026-08-15."
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"rule": "spaces",
|
|
65
|
+
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.4 «Διπλή τελεία», παρατήρηση iv): «Στα ελληνικά, πριν από τη διπλή τελεία δεν πρέπει να υπάρχει διάστημα (πράγμα που συμβαίνει, π.χ., στα γαλλικά)»· §10.1.3 «Άνω τελεία», παρατηρήσεις ii) και iii): «Στα ελληνικά, πριν από την άνω τελεία δεν πρέπει να υπάρχει διάστημα»· «Μετά την άνω τελεία αρχίζουμε με μικρό γράμμα»",
|
|
66
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
67
|
+
"note": "Note the exact standing of this citation, because it is weaker than the claim it is often read as supporting. Neither passage is about the ερωτηματικό: they are about the colon and the άνω τελεία. What they establish is that Greek denies the French space-before-punctuation pattern, by name, for the marks they do cover. No source examined addresses spacing before the Greek question mark specifically, so the position is \"no source contradicts it\", NOT \"a source requires it\". Nothing rests on the difference: the Greek question mark is written U+003B, which spec/rules/spaces.md lists in STRIP-BEFORE for the Latin semicolon on its own merits, so the space is stripped under the Latin reading alone and no Greek-specific decision is being made (spaces.md §3.5 part 2). The §6.4 spacing table is deliberately NOT offered as corroboration here — the nbsp source in this same file rejects §6.4 as publisher house style rather than evidence about Greek, and it cannot be house style there and evidence here. Retrieved 2026-08-15."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"rule": "nbsp",
|
|
71
|
+
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.4 παρατήρηση iv) και §10.1.3 παρατήρηση ii) (ό.π., ρητή αντιπαραβολή με τα γαλλικά)· §10.6 «Συντομογραφίες», γενικός κανόνας γ): «Στις ελληνικές συντομογραφίες με τις οποίες συντέμνονται φράσεις που αποτελούνται από περισσότερες της μίας λέξεις μπαίνει κατά κανόνα τελεία έπειτα από κάθε συντεμνόμενη λέξη» (π.χ. κ.λπ., π.χ., πρβλ., κ.ο.κ.)",
|
|
72
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
73
|
+
"note": "Two empty lists closed by citation rather than by absence of one. beforePunctuation and narrowBeforePunctuation are empty because Greek takes neither U+00A0 nor U+202F before «;» «:» «!» «·» — the source denies the French practice by name. abbreviations is empty because Greek multi-word abbreviations appear to be written with no internal space at all (κ.λπ., not «κ. λπ.»), so there is no U+0020 for the rule to promote. INFERENCE, NOT CITATION: §10.6 γ) states only that a full stop follows each abbreviated word; it says nothing about spacing between them. The no-space form is visible in the guide's own printed examples and in no normative sentence examined. The list would be empty on the absence of a citation alone (docs/PLAN.md §6.1), so nothing rests on the inference — it is labelled because an unlabelled inference in a cite field is the failure mode this field exists to prevent. The remaining lists are empty because no Greek normative source examined states a binding: the one candidate, the §6.4 table binding a number to «%» and «°C», sits in Part Three — «Συμβατικοί κανόνες κοινοί για ΟΛΕΣ τις γλώσσες» — whose own preamble says it replaced diverging national rules for uniform presentation, which makes it publisher house style rather than evidence about Greek. A missing citation is never a licence to guess (docs/PLAN.md §6.1). Observed usage does not contradict that reading, which is worth recording because it was checked rather than assumed: a byte-level scan on 2026-08-15 of the Greek-language home pages of kathimerini.gr, kathimerini.gr/economy, tovima.gr, tanea.gr, efsyn.gr, in.gr, naftemporiki.gr, protagon.gr, lifo.gr, meteo.gr, kedros.gr, onassis.org/el and emst.gr found 183 number-plus-unit occurrences — 140 «%», 41 «°C», 2 «€» — and NOT ONE of them carried a space of any kind, no U+0020, no U+00A0. Greek web practice sets «50%» and «20°C» tight, so beforeUnits would have nothing to promote even if it were populated. Method, so the figure can be re-derived rather than taken on trust: fetch each page, strip script, style and noscript elements and then all tags, decode HTML entities (so that becomes the U+00A0 it denotes rather than disappearing), and count matches of a digit followed by an optional single space character from {U+0020, U+00A0, U+202F, U+2009} followed by the unit. Pages carrying under 200 Greek letters were excluded as not being Greek-language content. That scan is an observation of usage, is not offered as a normative claim, and nothing in this file depends on it: the lists would be empty on the absence of a citation alone. Retrieved 2026-08-15."
|
|
74
|
+
},
|
|
75
|
+
{
|
|
76
|
+
"rule": "hyphen",
|
|
77
|
+
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.8, υποενότητα «Ενωτικό (-)»: «Η μικρή οριζόντια παύλα, το βραχύτερο σε μήκος από τα τρία σημεία […] σημειώνεται χωρίς κενά σε σχέση με ό,τι προηγείται ή έπεται»",
|
|
78
|
+
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
79
|
+
"note": "Justifies all three lists being empty. The source describes what the hyphen does — syllable division, numeric bounds, appositional compounds such as «απόφαση-πλαίσιο» — without defining any closed list of morphological forms whose hyphen must resist a line break, which is the only thing this rule consumes. No Greek source examined defines one. Empty lists make the rule a provable total no-op for Greek (spec/rules/hyphen.md §2), which is the normal case, not a deficiency. Retrieved 2026-08-15."
|
|
80
|
+
},
|
|
81
|
+
{
|
|
82
|
+
"rule": "apostrophe",
|
|
83
|
+
"cite": "ΥΠΕΠΘ / Ινστιτούτο Τεχνολογίας Υπολογιστών και Εκδόσεων, Γραμματική Νέας Ελληνικής Γλώσσας Α΄, Β΄, Γ΄ Γυμνασίου, §3.3 «Η στίξη — Τα σημεία στίξης»: «Η απόστροφος ( ’ ) χρησιμοποιείται για να δηλώσει ότι ένα φωνήεν έχει παραλειφθεί στη γραφή λόγω της προφοράς»· και, για την ταυτότητα του κωδικού σημείου, The Unicode Standard 17.0, Core Specification, §6.2.7 «Apostrophes»: «When text is set, U+2019 RIGHT SINGLE QUOTATION MARK is preferred as apostrophe, but only U+0027 is present on most keyboards»",
|
|
84
|
+
"url": "https://ebooks.edu.gr/ebooks/v/html/8547/2334/Grammatiki-Neas-Ellinikis-Glossas_A-B-G-Gymnasiou_html-apli/index_B_03.html",
|
|
85
|
+
"note": "Two sources for two separate facts. The Greek grammar establishes that elision and apocope (γι’ αυτό, απ’ την, σ’ αυτό) are marked with an apostrophe; the code point comes only from Unicode, because no Greek normative source examined names one. U+02BC is refused: Unicode §6.2.7 reserves it for use as a modifier letter, e.g. a glottal stop in transliteration. U+0384 GREEK TONOS is refused: no source proposes a diacritic in this role. The rule reads no locale data, so this citation certifies rather than configures — but Greek elision is the shape the rule most often meets in Greek text, and it earns a fixture. Retrieved 2026-08-15."
|
|
86
|
+
},
|
|
87
|
+
{
|
|
88
|
+
"rule": "symbols",
|
|
89
|
+
"cite": "Unicode Consortium, The Unicode Standard, Version 17.0, Core Specification, ch. 7 «Europe-I», §7.2.1 Greek, «Compatibility Punctuation», verbatim: «Therefore, use of U+037E and U+0387 is not necessary for interoperating with legacy Greek data, and their use is not generally encouraged for representation of Greek punctuation.» The preceding sentence of that subsection — to the effect that the two characters have canonical equivalences to U+003B and U+00B7, so that normalised Greek text loses the distinction — is given here as PARAPHRASE, not as quotation: it could not be confirmed word for word at review time. Its substance does not rest on the prose in any case; it is entailed directly by UnicodeData.txt, where U+037E carries the canonical decomposition mapping 003B and U+0387 carries 00B7, each as a singleton",
|
|
90
|
+
"url": "https://www.unicode.org/versions/Unicode17.0.0/core-spec/chapter-7/",
|
|
91
|
+
"note": "Code-point identity only, which is the one thing this source is authoritative for. Note the split standing of the citation: one sentence is verbatim, one is paraphrase backed by UnicodeData.txt rather than by the Core Specification's wording, and the file says which is which. Everything the rules do rests on the machine-readable field, not on the prose. The ερωτηματικό is written U+003B, not U+037E; the άνω τελεία is written U+00B7, not U+0387. Consequence for the rules: no rule may emit U+037E or U+0387, and none may rewrite one to the character it decomposes to — that is normalisation, forbidden by docs/ARCHITECTURE.md §4.3 and performed anyway by any downstream NFC pass. The full argument, including why U+00B7 must not join the spaces rule's STRIP-BEFORE set, is in spec/rules/spaces.md §3.5. Retrieved 2026-08-15."
|
|
92
|
+
}
|
|
93
|
+
]
|
|
44
94
|
}
|