polytypo 1.6.3 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/README.md +25 -0
- data/lib/polytypo/data/.not-vendored +11 -0
- data/lib/polytypo/data/README.md +16 -11
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +1 -1
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +14 -1
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +1 -1
- data/lib/polytypo/data/fixtures/en-US.json +521 -1
- data/lib/polytypo/data/fixtures/es.json +1 -1
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +1 -1
- data/lib/polytypo/data/fixtures/fr.json +68 -1
- data/lib/polytypo/data/fixtures/it.json +1 -1
- data/lib/polytypo/data/fixtures/locale-resolution.json +1 -1
- data/lib/polytypo/data/fixtures/nl.json +1 -1
- data/lib/polytypo/data/fixtures/pl.json +1 -1
- data/lib/polytypo/data/fixtures/pt-BR.json +1 -1
- data/lib/polytypo/data/fixtures/pt-PT.json +1 -1
- data/lib/polytypo/data/fixtures/ru.json +1 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/tr.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +1 -1
- data/lib/polytypo/data/locales/cs.json +56 -37
- data/lib/polytypo/data/locales/de-CH.json +43 -34
- data/lib/polytypo/data/locales/de-DE.json +47 -32
- data/lib/polytypo/data/locales/el.json +51 -1
- data/lib/polytypo/data/locales/en-GB.json +56 -4
- data/lib/polytypo/data/locales/en-US.json +68 -4
- data/lib/polytypo/data/locales/es.json +66 -29
- data/lib/polytypo/data/locales/fi.json +77 -7
- data/lib/polytypo/data/locales/fr-CA.json +63 -36
- data/lib/polytypo/data/locales/fr.json +70 -41
- data/lib/polytypo/data/locales/it.json +65 -22
- data/lib/polytypo/data/locales/nl.json +61 -29
- data/lib/polytypo/data/locales/pl.json +61 -28
- data/lib/polytypo/data/locales/pt-BR.json +53 -28
- data/lib/polytypo/data/locales/pt-PT.json +55 -28
- data/lib/polytypo/data/locales/registry.json +1 -1
- data/lib/polytypo/data/locales/ru.json +45 -30
- data/lib/polytypo/data/locales/sv.json +63 -4
- data/lib/polytypo/data/locales/tr.json +72 -32
- data/lib/polytypo/data/locales/uk.json +55 -27
- data/lib/polytypo/data/rules/modes.md +430 -11
- data/lib/polytypo/data/rules/order.json +1 -1
- data/lib/polytypo/data/schema/fixtures.schema.json +12 -1
- data/lib/polytypo/engine/pipeline.rb +29 -0
- data/lib/polytypo/modes/markdown.rb +256 -31
- data/lib/polytypo/modes/runner.rb +33 -3
- data/lib/polytypo/version.rb +1 -1
- data/lib/polytypo.rb +27 -12
- metadata +2 -1
|
@@ -45,47 +45,76 @@
|
|
|
45
45
|
"compounds": []
|
|
46
46
|
},
|
|
47
47
|
"nbsp": {
|
|
48
|
-
"beforePunctuation": [
|
|
49
|
-
|
|
50
|
-
],
|
|
51
|
-
"narrowBeforePunctuation": [
|
|
52
|
-
";",
|
|
53
|
-
"!",
|
|
54
|
-
"?"
|
|
55
|
-
],
|
|
48
|
+
"beforePunctuation": [":"],
|
|
49
|
+
"narrowBeforePunctuation": [";", "!", "?"],
|
|
56
50
|
"afterShortWords": [],
|
|
57
|
-
"abbreviations": [
|
|
58
|
-
|
|
59
|
-
],
|
|
60
|
-
"
|
|
61
|
-
|
|
62
|
-
"‰",
|
|
63
|
-
"€",
|
|
64
|
-
"°C",
|
|
65
|
-
"km",
|
|
66
|
-
"cm",
|
|
67
|
-
"mm",
|
|
68
|
-
"kg",
|
|
69
|
-
"km/h",
|
|
70
|
-
"kWh"
|
|
71
|
-
],
|
|
72
|
-
"beforeNumber": [
|
|
73
|
-
"art.",
|
|
74
|
-
"fig.",
|
|
75
|
-
"n°",
|
|
76
|
-
"N°"
|
|
77
|
-
],
|
|
78
|
-
"beforeWord": [
|
|
79
|
-
"M.",
|
|
80
|
-
"MM.",
|
|
81
|
-
"Mme",
|
|
82
|
-
"Mmes",
|
|
83
|
-
"Mlle",
|
|
84
|
-
"Mlles"
|
|
85
|
-
],
|
|
86
|
-
"afterSymbols": [
|
|
87
|
-
"§"
|
|
88
|
-
],
|
|
51
|
+
"abbreviations": ["p. ex."],
|
|
52
|
+
"beforeUnits": ["%", "‰", "€", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
53
|
+
"beforeNumber": ["art.", "fig.", "n°", "N°"],
|
|
54
|
+
"beforeWord": ["M.", "MM.", "Mme", "Mmes", "Mlle", "Mlles"],
|
|
55
|
+
"afterSymbols": ["§"],
|
|
89
56
|
"initialBinding": "single"
|
|
90
|
-
}
|
|
57
|
+
},
|
|
58
|
+
"sources": [
|
|
59
|
+
{
|
|
60
|
+
"rule": "nbsp",
|
|
61
|
+
"cite": "Jacques André, Petites leçons de typographie, éd. du jobet, révision du 1er juin 2025, § 5.1.1 note 21 et § 5.2.3 : l'espace insécable employée devant les ponctuations doubles « c'est la fine des typographes, en première approximation (contre-exemple : l'espace avant le deux-points est en fait une espace normale insécable) » ; « il faut une espace fine insécable avant le point-virgule »",
|
|
62
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
63
|
+
"note": "C'est le point qui tranche la question posée par docs/PLAN.md §7 : « ; », « ! » et « ? » prennent U+202F ; « : » prend U+00A0. La distinction est donc confirmée, et non simplement répétée. Le tableau 1 (p. 32) donne la saisie complète des signes. Recoupé avec l'article « Espace fine insécable » de la Wikipédia francophone, qui attribue la même règle au Lexique des règles typographiques en usage à l'Imprimerie nationale, 3e éd., 2002 (fine devant ; ? !, « sauf devant deux-points, en France »)."
|
|
64
|
+
},
|
|
65
|
+
{
|
|
66
|
+
"rule": "nbsp",
|
|
67
|
+
"cite": "Jacques André, Petites leçons de typographie, § 2.5 et § 5.1.3 « Autres emplois de l'espace insécable » : espace insécable entre un prénom abrégé et le nom (« N. Bourbaki »), entre une abréviation et le mot qui la suit (« Mme Hugo », « le R.P. Durand »), entre un nombre et ce qu'il quantifie (« 14 francs », « 2 € », « 98 % », « art. 237 »)",
|
|
68
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
69
|
+
"note": "La liste beforeUnits est volontairement courte et exclut les symboles d'une seule lettre (m, g, l, s, A), qui ne peuvent pas être désambiguïsés sans contexte. La liaison « abréviation + mot suivant » (M. Dupont, Mme Hugo) n'est pas exprimable dans ce schéma : afterSymbols lie un symbole à un NOMBRE suivant, pas à un mot ; elle passe par nbsp.initialBinding = \"single\" (spec 0.6.0 : le champ s'appelait auparavant le booléen bindInitials), qui conserve la liaison sur une seule initiale citée ici (« N. Bourbaki »). Contrairement à \"chain\" (en-US, de-DE, de-CH, ru — Chicago exige « two or more initials »), \"single\" ne peut pas distinguer structurellement un prénom abrégé authentique d'une collision de fin de phrase (une initiale isolée suivie d'un mot capitalisé qui commence en fait une nouvelle phrase) ; voir spec/rules/nbsp.md §7."
|
|
70
|
+
},
|
|
71
|
+
{
|
|
72
|
+
"rule": "quotes",
|
|
73
|
+
"cite": "Imprimerie nationale, Lexique des règles typographiques en usage à l'Imprimerie nationale, 3e éd., 2002, entrée « Guillemets » : les guillemets français « … » sont séparés du texte qu'ils encadrent par une espace insécable",
|
|
74
|
+
"url": "https://fr.wikipedia.org/wiki/Espace_fine_ins%C3%A9cable",
|
|
75
|
+
"note": "Le Lexique n'est pas consultable en ligne ; la règle a été vérifiée par recoupement de deux articles de la Wikipédia francophone qui la citent (« Espace fine insécable » : « à l'intérieur des guillemets français […] il est recommandé d'insérer des espaces insécables », attribué au Lexique 2002 ; « Ponctuation », tableau des espacements) et du tableau 1 de Jacques André, qui note une insécable après « « » et avant « » ». Désaccord assumé avec docs/PLAN.md §7, qui annonçait U+202F à l'intérieur des guillemets : la source normative dit espace insécable (U+00A0). L'usage web contemporain, et JoliTypo, emploient souvent U+202F ; le choix retenu ici est celui de la citation, pas celui de l'usage."
|
|
76
|
+
},
|
|
77
|
+
{
|
|
78
|
+
"rule": "quotes",
|
|
79
|
+
"cite": "Jacques André, Petites leçons de typographie, § 2.2 : « en français, les guillemets sont les doubles chevrons « … » et non les (double-)quotes anglaises “…” ni ‘…’ » ; les guillemets anglais “ ” servent de guillemets de second niveau",
|
|
80
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
81
|
+
"note": "Le second niveau est le point le moins solidement établi de ce fichier : l'usage français admet aussi la répétition des chevrons ou les chevrons simples ‹ ›. Les guillemets anglais sont retenus parce qu'ils sont la solution la plus couramment attribuée au Lexique, mais cette ligne mérite d'être revue si la 3e édition papier dit autre chose."
|
|
82
|
+
},
|
|
83
|
+
{
|
|
84
|
+
"rule": "dashes",
|
|
85
|
+
"cite": "Jacques André, Petites leçons de typographie, § 2.2 et tableau 1 (p. 32) : le tiret marquant les incises est le tiret cadratin « — », précédé et suivi d'une espace (« mmm —_mmm … mmm_— mmm »)",
|
|
86
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
87
|
+
"note": "Le même paragraphe signale que « la tendance est d'utiliser à sa place le tiret moyen « – » ». La valeur em-spaced suit la règle citée, pas la tendance ; docs/PLAN.md §7 supposait en-spaced. Aucune valeur de plage n'est vérifiée : les exemples de l'Imprimerie nationale et de Jacques André impriment un trait d'union dans les intervalles de pages (« p. 123-125 »), et l'énumération se fait plutôt « de … à … ». Le schéma a donc reçu la valeur « none » (ne rien substituer), retenue ici tant que le point n'est pas tranché sur le Lexique papier — une citation manquante n'autorise jamais à deviner."
|
|
88
|
+
},
|
|
89
|
+
{
|
|
90
|
+
"rule": "nbsp",
|
|
91
|
+
"cite": "Jacques André, Petites leçons de typographie, § 5.1.3 « Autres emplois de l'espace insécable », liste « Coupure entre les mots » : « entre une abréviation et le mot qui la suit, exemples : “Mme_Hugo, D._Knuth, le R.P._Durand” » ; « entre un nombre et ce qu'il quantifie, par exemple : “14_francs, 2_€, 1_A, t._vii, art._237, fig._3, pages_23 à_25, 98_%” » — le caractère « _ » note l'espace insécable",
|
|
92
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
93
|
+
"note": "Texte relu directement dans le PDF. Cette entrée lève la limitation signalée plus haut dans ce fichier (« la liaison abréviation + mot suivant n'est pas exprimable dans ce schéma ») : le schéma a depuis reçu beforeNumber et beforeWord. beforeNumber ne retient que « art. » et « fig. », cités littéralement devant un nombre ; « t. vii » est suivi d'un chiffre romain et n'a pas été retenu. beforeWord est volontairement réduit à la classe fermée des titres de civilité : la règle citée vaut pour toute abréviation, mais l'énumérer intégralement produirait des faux positifs (« etc. », « p. ex. », « cf. »). « Mme » est cité mot pour mot ; « M., MM., Mmes, Mlle, Mlles » sont les autres membres de la même classe et relèvent donc de la même règle, ce qui reste une inférence de classe et non une citation littérale. Un « M. » qui serait en réalité une initiale de prénom reçoit de toute façon l'espace insécable prescrite au § 5.1.3 pour « N. Bourbaki » : la liaison est correcte dans les deux lectures."
|
|
94
|
+
},
|
|
95
|
+
{
|
|
96
|
+
"rule": "hyphen",
|
|
97
|
+
"cite": "Jacques André, Petites leçons de typographie, tableau 1 (p. 32), ligne « trait d'union » : la saisie est « mmm-mmm », sans espace ni marque d'insécabilité, alors que le même tableau note explicitement l'insécable pour les ponctuations doubles, les guillemets et les tirets d'incise",
|
|
98
|
+
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
99
|
+
"note": "Justification des trois listes vides. Aucune source normative consultée n'impose un trait d'union insécable pour une classe fermée de formes morphologiques françaises ; le § 5.1.3, qui énumère les cas d'insécabilité, ne mentionne aucun trait d'union. Le champ hyphen est donc sans objet en français."
|
|
100
|
+
},
|
|
101
|
+
{
|
|
102
|
+
"rule": "nbsp",
|
|
103
|
+
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Abréviation de numéro » : « L'abréviation courante du nom numéro est no ou No, avec la lettre o minuscule en exposant » ; « On n'abrège le mot numéro que s'il suit immédiatement le nom qu'il détermine » ; exemple : « C'est le billet no 852410 qui a valu le gros lot à sa détentrice »",
|
|
104
|
+
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/25450/les-abreviations-et-les-symboles/les-abreviations/cas-particuliers-dabreviations/abreviation-de-numero",
|
|
105
|
+
"note": "Deux choses distinctes, et il faut les tenir séparées. Ce que la source établit : « n° » est une abréviation, pas un symbole, d'où beforeNumber et non afterSymbols — l'argument est aussi mécanique, un « ° » seul placé dans afterSymbols serait inerte puisque N6 exige que le code point précédent ne soit pas ALNUM, or c'est la lettre « n » (nbsp.md §3.8). Ce que la source n'établit pas : la séquence exacte de code points. L'OQLF prescrit un o minuscule en exposant, un tracé et non un code point, et ne nomme nulle part le signe de degré ni la ligature « numéro » (U+2116). Pour la France, rien : le § 5.1.3 de Jacques André, relu intégralement, ne mentionne ni « numéro » ni « no », et le Lexique de l'Imprimerie nationale est un ouvrage imprimé qui n'a pas pu être consulté. Décision de l'opérateur (18.09.2026) : quand la règle est derrière une source payante, polytypo retient l'usage le plus répandu et l'applique de la même façon partout, l'uniformité primant la conformité au canon, et la décision est signalée comme telle au lieu d'être présentée comme une citation. L'usage le plus répandu est « n » + U+00B0, c'est ce que l'on tape et ce que l'on trouve dans les textes en ligne ; N9 n'ayant aucune tolérance de casse, « n° » et « N° » sont deux entrées. La phrase « on ne sépare pas les numéros par des espacements » de la même page vise les tranches de chiffres à l'intérieur du nombre (« billet no 852410 »), pas l'espace entre l'abréviation et le nombre."
|
|
106
|
+
},
|
|
107
|
+
{
|
|
108
|
+
"rule": "quotes",
|
|
109
|
+
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Cas où l'élision est obligatoire » (dernière mise à jour : 2015) : « L'élision ne touche que des mots grammaticaux, habituellement courts, et que les voyelles a, e et i. » L'élision est obligatoire pour le déterminant et le pronom la et le, pour les pronoms je, me, te, se et ce placés devant le verbe, pour si devant il(s), pour que et les conjonctions qui le contiennent, pour jusque, pour la préposition de et pour l'adverbe de négation ne",
|
|
110
|
+
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/21737/lorthographe/elision-et-apostrophe/elision-obligatoire",
|
|
111
|
+
"note": "Justifie dix des treize entrées de quotes.elisionClitics.before : l (la, le), j (je), m (me), t (te), s (se et si), c (ce), qu (que), jusqu (jusque), d (de), n (ne). Chaque entrée est la plage de lettres maximale à gauche de la marque, d'où jusqu et non qu pour « jusqu'à » : sous une comparaison par plage maximale, une entrée qu ne peut pas apparier « jusqu' ». presqu (presqu'île) et quelqu (quelqu'un) sont volontairement exclus, faute d'attestation dans les articles relus, et l'exclusion est sans coût : dans les deux mots la marque a une lettre de chaque côté. La liste after est vide — la BDL définit l'élision comme l'effacement de la voyelle finale d'un mot, donc la marque suit toujours le mot élidé, et l'enclise française prend le trait d'union (donne-moi, y a-t-il), jamais l'apostrophe. Le cas porté par cette liste est l'élision devant un élément en ligne, l'<em>idée</em> : sans elle, quotes appariait la marque comme une citation et invertissait la paire englobante (issue #53). L'OQLF est l'autorité du Québec : pour fr-CA la citation est directe, pour fr elle tient lieu de substitut, le Lexique de l'Imprimerie nationale et Le Bon Usage n'étant pas consultables en ligne — même lacune que celle déjà signalée ailleurs dans ce fichier. quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
112
|
+
},
|
|
113
|
+
{
|
|
114
|
+
"rule": "quotes",
|
|
115
|
+
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Élision de lorsque, puisque et quoique » (2020) : « les conjonctions lorsque, puisque et quoique s'élident obligatoirement devant il(s), elle(s), on, un et une, et aussi devant la préposition en » ; l'élision généralisée devant tout mot commençant par une voyelle ou un h muet « demeure facultative »",
|
|
116
|
+
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/23623/lorthographe/elision-et-apostrophe/elision-de-lorsque-puisque-et-quoique",
|
|
117
|
+
"note": "Justifie les trois entrées lorsqu, puisqu, quoiqu. L'élision obligatoire devant il(s), elle(s), on, un, une et en suffit à l'appartenance : la liste étant uniquement restrictive, le caractère facultatif de l'élision généralisée ne change rien. Là encore la plage maximale impose des entrées distinctes : qu n'apparie pas « lorsqu' ». quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
118
|
+
}
|
|
119
|
+
]
|
|
91
120
|
}
|
|
@@ -14,11 +14,7 @@
|
|
|
14
14
|
},
|
|
15
15
|
"elisionIdioms": [],
|
|
16
16
|
"elisionClitics": {
|
|
17
|
-
"before": [
|
|
18
|
-
"l",
|
|
19
|
-
"un",
|
|
20
|
-
"d"
|
|
21
|
-
],
|
|
17
|
+
"before": ["l", "un", "d"],
|
|
22
18
|
"after": []
|
|
23
19
|
}
|
|
24
20
|
},
|
|
@@ -39,24 +35,71 @@
|
|
|
39
35
|
"narrowBeforePunctuation": [],
|
|
40
36
|
"afterShortWords": [],
|
|
41
37
|
"abbreviations": [],
|
|
42
|
-
"beforeUnits": [
|
|
43
|
-
|
|
44
|
-
"‰",
|
|
45
|
-
"€",
|
|
46
|
-
"°C",
|
|
47
|
-
"km",
|
|
48
|
-
"cm",
|
|
49
|
-
"mm",
|
|
50
|
-
"kg",
|
|
51
|
-
"km/h",
|
|
52
|
-
"kWh"
|
|
53
|
-
],
|
|
54
|
-
"beforeNumber": [
|
|
55
|
-
"n.",
|
|
56
|
-
"pag."
|
|
57
|
-
],
|
|
38
|
+
"beforeUnits": ["%", "‰", "€", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
39
|
+
"beforeNumber": ["n.", "pag."],
|
|
58
40
|
"beforeWord": [],
|
|
59
41
|
"afterSymbols": [],
|
|
60
42
|
"initialBinding": "none"
|
|
61
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Ufficio delle pubblicazioni dell'Unione europea, Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), punto 4.2.3, riquadro «Virgolette»: «Nella lingua italiana esistono tre livelli di virgolette […]: livello 1 (citazione principale) «…» (Alt 0171/Alt 0187); livello 2 (citazione nella citazione) “…” (Alt 0147/Alt 0148); livello 3 (citazione nella citazione nella citazione) ‘…’ (Alt 0145/Alt 0146)»",
|
|
48
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
49
|
+
"note": "L'attribuzione è per LIVELLO DI ANNIDAMENTO, cioè esattamente la forma che lo schema sa esprimere. Il terzo livello non è rappresentabile: lo schema prevede primary e secondary, non tre — limite dello schema, non lacuna della fonte, come già in el. INFERENZA DICHIARATA sull'identità dei code point: la guida fornisce un «codice alfanumerico» per Windows, non un code point Unicode; Alt 0171/0187/0147/0148 sono CP1252 0xAB/0xBB/0x93/0x94, da cui U+00AB/U+00BB/U+201C/U+201D. Riscontro indipendente: Treccani, Enciclopedia dell'Italiano (2011), voce «Virgolette» (Luca Cignetti): «è consuetudine fare ricorso a tipi diversi di virgoletta, solitamente bassa per la prima citazione e alta per gli altri usi». POSIZIONE CONTRARIA, registrata: l'Accademia della Crusca («La punteggiatura», 16 luglio 2004) rifiuta di ordinare i due segni al primo livello — «Alte e basse si usano indifferentemente» — ma non tratta l'annidamento, quindi non c'è conflitto: la scelta del segno di primo livello è una variante permessa che questo file compie, l'ordine dei livelli è ciò che le altre due fonti stabiliscono. Consultati il 18.09.2026."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "quotes",
|
|
53
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), Parte quarta, punto 10.1.7 «Virgolette»: «Le virgolette non richiedono spazio né in apertura né in chiusura»",
|
|
54
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
55
|
+
"note": "Giustifica innerSpace = \"none\" su entrambe le coppie, ed è il punto che separa l'italiano dal francese. LETTURA DICHIARATA COME TALE: «né in apertura né in chiusura» va inteso come «né dal lato di apertura né dal lato di chiusura del testo racchiuso», cioè nessuno spazio interno; la lettura alternativa è impossibile, perché lo spazio fra la parola precedente e il segno di apertura non è facoltativo. La tabella del punto 6.4 disambiguerebbe la frase ma NON è invocata: il suo preambolo la qualifica come accordo interistituzionale adottato «a vantaggio di una convenzione comune» al posto dei codici tipografici nazionali divergenti, cioè prassi editoriale e non norma italiana — la stessa lettura già applicata in el.json. Conseguenza pratica: l'elisione italiana è al sicuro, perché dell'«amico» non richiede spazio interno. elisionIdioms resta vuoto: l'elisione italiana è un fenomeno clitico, non un idioma della forma {left, elided, right}. Consultati il 18.09.2026."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "dashes",
|
|
59
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), punto 4.2.3, riquadro «Trattini»: «In italiano utilizzare il trattino lungo come eventuale alternativa alle parentesi»; Parte quarta, punto 10.1.10 «Lineetta», con l'esempio stampato «La casa — se casa si poteva definire quel rudere — sorgeva ai piedi del colle»",
|
|
60
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
61
|
+
"note": "Giustifica dash.parenthetical = \"em-spaced\": U+2014 con uno spazio ordinario su CIASCUN lato di CIASCUNA lineetta. La spaziatura è simmetrica, quindi — a differenza dello spagnolo e del greco — l'enum la esprime senza forzature. INFERENZA DICHIARATA sul code point: Alt 0151 è CP1252 0x97, da cui U+2014. Riscontro sulla SPAZIATURA, non sul code point: Treccani, «Trattino [prontuario]» (Stefano Telve, 2011), che del trattino breve scrive «A differenza della lineetta, non è preceduto né seguito da spazi». COSTO ACCETTATO, verificato sul motore e registrato qui perché non venga riscoperto come bug: il punto 10.1.11 prescrive il trattino spaziato per collegare due date di mesi diversi («25 maggio - 10 giugno 1985»), e la guardia P1 di dashes.md §3.2 converte qualunque trattino isolato e spaziato fra parole, quindi il motore produce una lineetta lì. Non è una particolarità italiana: fr, ru, de-DE e en-GB, già pubblicate, si comportano allo stesso modo sulla stessa stringa. Consultati il 18.09.2026."
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"rule": "ranges",
|
|
65
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), Parte quarta, punto 10.3.1 «Numeri arabi»: «I numeri arabi non richiedono la lineetta (lineato) bensì il trattino (divisione)», con l'esempio stampato «la sessione del 15-19 giugno 1985»; punto 10.1.11: «Il trattino è anche impiegato, senza alcuno spazio, per collegare due anni civili o due date di uno stesso mese: 1971-1972, 10-15 settembre 1985»",
|
|
66
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
67
|
+
"note": "Giustifica dash.range = \"none\", e lo fa in POSITIVO, non per assenza di citazione: la fonte prescrive il trattino U+002D — il carattere che l'autore ha già digitato — ed esclude la lineetta per nome. L'enum non ha un valore «trattino»: \"none\" è la codifica esatta di ciò che la fonte richiede. Riscontro: Treccani, «Trattino [prontuario]» (Telve 2011), esempi «il fascicolo del 2-3 luglio», «anno accademico 2010-2011». Voce separata da quella di dashes benché entrambe le regole leggano il campo dash: la citazione è distinta e più forte. Consultati il 18.09.2026."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"rule": "ellipsis",
|
|
71
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), Parte quarta, punto 10.1.8 «Punti di sospensione»: «Segno d'interpunzione costituito da tre punti»; Accademia della Crusca, Consulenza linguistica, «La punteggiatura» (Mara Marzullo, 16 luglio 2004): «I puntini di sospensione si usano sempre nel numero di tre»",
|
|
72
|
+
"url": "https://accademiadellacrusca.it/it/consulenza/la-punteggiatura/143",
|
|
73
|
+
"note": "Giustifica abbreviatedAfterTerminal = false. Due fonti indipendenti affermano che i puntini sono sempre tre, e nessuna prevede la forma a due punti dopo U+003F o U+0021 che usa il russo: il ramo russo di ellipsis.md resta spento per l'italiano. Nessuna delle due si pronuncia su U+2026 contro tre U+002E; quella scelta è del motore ed è identica in ogni locale. DIVERGENZA NOTA, registrata: gli esempi stampati del Manuale non lasciano spazio prima dei puntini, mentre spaces.md §3.4 conserva in ogni locale lo spazio che precede una sequenza di punti; è deliberata e argomentata in spaces.md §7.9, identica a quella già registrata per el. Il Manuale è consultabile all'indirizzo op.europa.eu (cfr. le altre voci di questo file); l'url di questa voce punta alla fonte Crusca, poiché lo schema ne ammette uno solo. Consultati il 18.09.2026."
|
|
74
|
+
},
|
|
75
|
+
{
|
|
76
|
+
"rule": "hyphen",
|
|
77
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), Parte quarta, punto 10.1.11 «Trattino»: «Il trattino («divisione» in linguaggio tipografico) […] si impiega principalmente per dividere la parola in fin di riga»; «Si può anche usare per unire due parole contraendo la prima: trattato italo-francese»",
|
|
78
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
79
|
+
"note": "Giustifica le tre liste vuote. La fonte enumera gli impieghi del trattino — e il primo che nomina è l'OPPOSTO dell'insecabilità — senza definire alcuna classe chiusa di forme morfologiche il cui trattino debba resistere all'andata a capo, che è l'unica cosa che questa regola consuma. Riscontro: Treccani, «Trattino [prontuario]» (Telve 2011), che elenca le funzioni di unione (afro-cubano, porta-finestra, maxi-schermo) senza definire neppure essa un elenco chiuso. La famiglia di prefissi (ex-, anti-, maxi-) è stata verificata espressamente perché è l'unica che lo schema saprebbe contenere: le fonti la trattano come schema produttivo e aperto, e un insieme aperto non è esprimibile in wordList. Consultati il 18.09.2026."
|
|
80
|
+
},
|
|
81
|
+
{
|
|
82
|
+
"rule": "nbsp",
|
|
83
|
+
"cite": "BIPM, The International System of Units (SI), 9th ed. (2019), concise summary: «A single space is always left between the number and the unit»",
|
|
84
|
+
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
85
|
+
"note": "Autorità per l'ELENCO delle unità, non per il legame: che lo spazio sia insecabile e che sia U+00A0 è deciso da nbsp.md §2.1, che una locale non può variare. Stessa impostazione di en-US, en-GB, fi e sv. L'elenco segue la forma prudente di fr e de-DE ed esclude deliberatamente i simboli di una sola lettera (m, g, l, s, A), che non si possono disambiguare senza contesto. La tabella del punto 6.4 del Manuale darebbe lo stesso risultato per «%» e «°C» ma non è invocata, per la ragione detta nella voce quotes. Consultati il 18.09.2026."
|
|
86
|
+
},
|
|
87
|
+
{
|
|
88
|
+
"rule": "nbsp",
|
|
89
|
+
"cite": "Manuale interistituzionale di convenzioni redazionali, ed. 2011 (IT), punto 4.2.3, riquadro «Spazi vuoti protetti»: «Consentono di evitare la troncatura a fine riga di elementi che devono rimanere uniti»; «Da usare solo nei casi seguenti […]: n.• JO L• 10•000 / pag.• JO C• C.•M. Dupont», dove «•» indica lo spazio fisso",
|
|
90
|
+
"url": "https://op.europa.eu/it/publication-detail/-/publication/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b",
|
|
91
|
+
"note": "INFERENZA DICHIARATA, accettata per decisione dell'operatore il 18.09.2026, non citazione piena — ed è esattamente la distinzione che questo campo esiste per non far sparire. I riquadri «Virgolette» e «Trattini» dello stesso punto 4.2.3 dicono «Nella lingua italiana» e «In italiano»: questo riquadro NON lo dice, ed è in parte non adattato, poiché «JO L» e «JO C» sono il Journal officiel francese e «Dupont» è un cognome francese, in un'edizione italiana il cui testo originale era in francese. Ciò che l'italiano c'è davvero: «pag.» è italiano (il francese scrive «p.») e «n.» è uso italiano corrente, quindi il riquadro è stato localizzato almeno in parte — su questo poggia beforeNumber. La sequenza «C.•M. Dupont» non contiene invece alcun materiale italiano, e initialBinding resta perciò \"none\", il valore pulito da citazione. Nessuna tolleranza di caso: N9 non ne ha, e sono stampate solo le forme minuscole «n.» e «pag.». «10•000» è raggruppamento delle migliaia e «JO L•» una serie della Gazzetta ufficiale: nessuno dei due è esprimibile in questo schema. Consultati il 18.09.2026."
|
|
92
|
+
},
|
|
93
|
+
{
|
|
94
|
+
"rule": "nbsp",
|
|
95
|
+
"cite": "Nessuna fonte normativa italiana esaminata prescrive uno spazio prima di un segno d'interpunzione, né un legame per abbreviazioni con spazio interno, simboli o parole seguenti; liste deliberatamente vuote",
|
|
96
|
+
"note": "beforePunctuation e narrowBeforePunctuation sono vuote, e va detto con esattezza su che cosa poggiano: sull'ASSENZA di una fonte italiana che richieda quello spazio, NON su una fonte che lo neghi. A differenza del capitolo greco, che respinge la prassi francese nominandola, la Parte quarta italiana non la nomina mai: i punti 10.1.3 e 10.1.4 descrivono la funzione dei segni e tacciono sulla spaziatura. Il vincolo Q-P rende comunque la lista vuota la scelta strutturalmente sicura. afterShortWords, abbreviations, beforeWord e afterSymbols sono vuote per la stessa ragione: nessuna fonte esaminata le attesta per l'italiano. SEGNALAZIONE ALL'OPERATORE, non risolta qui: la Parte quarta si apre, al punto 10.1, con «Per la spaziatura dei segni d'interpunzione, cfr. punto 6.4», cioè il capitolo squalificato come prassi editoriale in el.json; se quel rinvio si leggesse come un'adozione, la tabella diventerebbe evidenza italiana e rafforzerebbe queste liste vuote. Consultati il 18.09.2026."
|
|
97
|
+
},
|
|
98
|
+
{
|
|
99
|
+
"rule": "quotes",
|
|
100
|
+
"cite": "Accademia della Crusca, Consulenza linguistica : Vera Gheno, « L'articolo indeterminativo » (30 settembre 2002), sulla forma femminile « un' (nei casi in cui si userebbe l'articolo determinativo l', ma l'elisione non è obbligatoria) [...] va usata davanti a parole che iniziano per vocale » ; Raffaella Setti, « Elisione e troncamento nell'italiano contemporaneo » (14 marzo 2008) : con di l'elisione è spesso applicata e in alcuni casi obbligatoria (d'accordo, d'oro), mentre con mi, ti, la, vi, si « l'apostrofo è del tutto facoltativo » e « non c'è nessun obbligo di elidere, ma è opportuno valutare caso per caso »",
|
|
101
|
+
"url": "https://accademiadellacrusca.it/it/consulenza/elisione-e-troncamento-nellitaliano-contemporaneo/174",
|
|
102
|
+
"note": "Italian elision is not a closed set — the Accademia states in terms that it is optional and to be judged case by case — so this entry evidences membership per fragment, exactly as quotes.md §3.2 requires of each elisionIdioms triple, and not the completeness of the list. Three fragments are named literally and are all that is listed: l (definite article l'), un (feminine indefinite un' — Crusca is categorical that masculine un takes no apostrophe, and the mechanism cannot tell the two apart, which is harmless because it only ever declines a pairing), and d (di, with d'accordo and d'oro given as obligatory). dell, nell, all, sull and dall are deliberately absent, being unattested by the consulenze read here; under maximal-run matching an l entry does not cover dell'arte either way, since the run there is dell. after is empty: Italian has no enclitic written with an apostrophe. quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
103
|
+
}
|
|
104
|
+
]
|
|
62
105
|
}
|
|
@@ -15,16 +15,7 @@
|
|
|
15
15
|
"elisionIdioms": [],
|
|
16
16
|
"elisionClitics": {
|
|
17
17
|
"before": [],
|
|
18
|
-
"after": [
|
|
19
|
-
"s",
|
|
20
|
-
"t",
|
|
21
|
-
"ns",
|
|
22
|
-
"k",
|
|
23
|
-
"m",
|
|
24
|
-
"em",
|
|
25
|
-
"r",
|
|
26
|
-
"et"
|
|
27
|
-
]
|
|
18
|
+
"after": ["s", "t", "ns", "k", "m", "em", "r", "et"]
|
|
28
19
|
}
|
|
29
20
|
},
|
|
30
21
|
"dash": {
|
|
@@ -43,26 +34,67 @@
|
|
|
43
34
|
"beforePunctuation": [],
|
|
44
35
|
"narrowBeforePunctuation": [],
|
|
45
36
|
"afterShortWords": [],
|
|
46
|
-
"abbreviations": [
|
|
47
|
-
|
|
48
|
-
],
|
|
49
|
-
"beforeUnits": [
|
|
50
|
-
"%",
|
|
51
|
-
"‰",
|
|
52
|
-
"°C",
|
|
53
|
-
"km",
|
|
54
|
-
"cm",
|
|
55
|
-
"mm",
|
|
56
|
-
"kg",
|
|
57
|
-
"km/h",
|
|
58
|
-
"kWh"
|
|
59
|
-
],
|
|
37
|
+
"abbreviations": ["prof. dr."],
|
|
38
|
+
"beforeUnits": ["%", "‰", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
60
39
|
"beforeNumber": [],
|
|
61
40
|
"beforeWord": [],
|
|
62
|
-
"afterSymbols": [
|
|
63
|
-
"€",
|
|
64
|
-
"$"
|
|
65
|
-
],
|
|
41
|
+
"afterSymbols": ["€", "$"],
|
|
66
42
|
"initialBinding": "none"
|
|
67
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Dubbele of enkele aanhalingstekens bij een citaat»: «Er zijn geen vaste regels voor het gebruik van enkele of dubbele aanhalingstekens. Traditioneel werd aangeraden om bij letterlijk citeren dubbele aanhalingstekens te gebruiken […] We raden aan om consequent voor één systeem te kiezen», met voorbeeld (4a): «De koning zei: [U+201C]Ik herinner me nog dat iemand [U+2018]Vive la république[U+2019] riep[U+201D]»",
|
|
48
|
+
"url": "https://taaladvies.net/dubbele-of-enkele-aanhalingstekens-bij-een-citaat/",
|
|
49
|
+
"note": "BESLISSING VAN DE OPERATOR, 18.09.2026, en uitdrukkelijk geen citaat: de Taalunie WEIGERT te rangschikken — «Er zijn geen vaste regels» — en noemt drie mogelijke nestingen zonder er één te verkiezen. Het schema kent maar één paar per niveau, dus er moet gekozen worden. Gekozen is variant (4a) van de Taalunie zelf: U+201C/U+201D als eerste niveau, U+2018/U+2019 als tweede. Twee gronden, allebei genoemd omdat geen van beide een voorschrift is: het is de vorm die de Taalunie «traditioneel» aan het letterlijke citaat verbindt, en het is de vorm die elk ander locale in dit project al heeft. TEGENGESTELDE AANWIJZING, niet verzwegen: «Aanhalingstekens (algemeen)» merkt op dat «Enkele aanhalingstekens […] het meest gebruikt» worden — een waarneming over frequentie, geen voorschrift, maar hij wijst de andere kant op. GECONTROLEERD, want het was het enige technische argument dat de keuze had kunnen beslissen: de vrees dat U+2019 als sluitteken zou botsen met de Nederlandse apostrof (auto's, 's-Hertogenbosch, A4'tje) blijkt ONGEGROND — beide varianten geven op die zinnen hetzelfde, correcte resultaat. De keuze berust dus op de bron en op eenvormigheid, niet op een motoreigenschap. innerSpace = «none»: geen geraadpleegde bron vraagt om een spatie binnen de tekens. Geraadpleegd op 18.09.2026."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Wel of geen spaties voor en na leestekens en symbolen (algemeen)» (12 mei 2021): «Voor een gedachtestreepje komt altijd een spatie, en ook erna»; en «Gedachtestreepje»: «Als gedachtestreepje wordt doorgaans het zogeheten halve kastlijntje (–) gebruikt. Het halve kastlijntje is wat langer dan het koppelteken of het afbreekteken (-)»",
|
|
54
|
+
"url": "https://taaladvies.net/wel-of-geen-spaties-voor-en-na-leestekens-en-symbolen-algemeen/",
|
|
55
|
+
"note": "Onderbouwt dash.parenthetical = «en-spaced»: een spatie aan weerszijden en een streepje ter lengte van een halve kastlijn, dus U+2013 en niet U+2014. AFLEIDING, als zodanig gemarkeerd: de Taalunie noemt de lengte, niet het codepunt. De toewijzing aan U+2013 steunt op Genootschap Onze Taal, Taalloket, «kort streepje (-) of lang streepje (–)», dat de twee tekens expliciet tegenover elkaar zet en het «hele kastlijntje» (—) als vooral Engels en «in Nederland […] minder gebruikelijk» beschrijft. Rangorde, want de bronnen zijn niet gelijkwaardig: de Taalunie is het officiële taalorgaan, Onze Taal een taalvereniging, en Onze Taal wordt hier alleen gebruikt om vast te stellen welk van twee met name genoemde tekens de term van de Taalunie aanduidt. Geraadpleegd op 18.09.2026."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "ranges",
|
|
59
|
+
"cite": "Genootschap Onze Taal, Taalloket, «kort streepje (-) of lang streepje (–)»: het korte streepje geeft tussen getallen «tot en met» aan — «1940-1945, de categorie 30-45 jaar, pagina 10-12»; Nederlandse Taalunie, «20 tot 30-jarigen of 20-30-jarigen?» (12 mei 2021), dat overal U+002D toont zonder over het streepjestype iets voor te schrijven",
|
|
60
|
+
"url": "https://onzetaal.nl/taalloket/streepje-kort-of-lang",
|
|
61
|
+
"note": "Onderbouwt dash.range = «none» als een keuze en niet als een leemte: het Nederlandse bereikteken is het koppelteken U+002D, dat de auteur al heeft getypt, en het schema kent daarvoor geen enum-waarde. «none» laat de invoer ongemoeid, wat precies is wat de bron beschrijft; «en-tight» zou er een vorm van maken die geen enkele Nederlandse bron voorschrijft. Aparte ingang van die van dashes, omdat de twee velden op verschillende bronnen rusten. De regel «ranges» staat standaard uit (spec 0.5.0). Geraadpleegd op 18.09.2026."
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"rule": "ellipsis",
|
|
65
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Punt na enz. of beletselteken aan het einde van een zin» (12 mei 2021): «na een afkorting met een punt of na een beletselteken aan het eind van een zin, komt geen extra punt», met het voorbeeld «Hou je van jazz, blues, soul …?»",
|
|
66
|
+
"url": "https://taaladvies.net/punt-na-enz-of-beletselteken-aan-het-einde-van-een-zin/",
|
|
67
|
+
"note": "Onderbouwt abbreviatedAfterTerminal = false. Het beletselteken is «een vaste eenheid van drie punten» en het vraagteken volgt erop; voor de Russische tweepuntsvorm «?..» geeft geen geraadpleegde Nederlandse bron een grondslag. Anders dan bij het Grieks valt hier niets af te wijken: het Nederlands schrijft juist wél een spatie vóór het beletselteken, en spaces.md §3.4 laat die spatie in elk locale staan — bron en motor zijn het hier eens. Geraadpleegd op 18.09.2026."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"rule": "hyphen",
|
|
71
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Auto-ongeluk (afbreking)» (12 mei 2021): «Bij woordafbreking valt een koppelteken weg als wordt afgebroken op de plaats waar dat koppelteken in de grondvorm staat»",
|
|
72
|
+
"url": "https://taaladvies.net/auto-ongeluk-afbreking/",
|
|
73
|
+
"note": "Onderbouwt drie lege lijsten. Het Nederlands staat afbreking op het koppelteken niet alleen toe, het laat het koppelteken daarbij wégvallen — het omgekeerde van een claim op U+2011. Geen geraadpleegde bron omschrijft een gesloten klasse vormen waarvan het koppelteken een regelovergang moet weerstaan, en dat is het enige wat deze regel leest. Geraadpleegd op 18.09.2026."
|
|
74
|
+
},
|
|
75
|
+
{
|
|
76
|
+
"rule": "nbsp",
|
|
77
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Wel of geen spaties voor en na leestekens en symbolen (algemeen)» (12 mei 2021): «Vóór een punt, een komma, een dubbele punt, een puntkomma, een uitroepteken en een vraagteken komt geen spatie»; «Bij toevallige, niet-conventionele combinaties van een getal en een afgekorte eenheid […] komt er een spatie tussen het getal en de eenheid»; «Tussen een valutateken en het bijbehorende getal staat een spatie, tenzij het geheel deel uitmaakt van een samenstelling» — «De vaas kostte € 179,–»; BIPM, The International System of Units (SI), 9th ed., concise summary — uitsluitend voor de samenstelling van de symbolenlijst",
|
|
78
|
+
"url": "https://taaladvies.net/wel-of-geen-spaties-voor-en-na-leestekens-en-symbolen-algemeen/",
|
|
79
|
+
"note": "Drie velden uit één pagina. beforePunctuation en narrowBeforePunctuation leeg: alle zes leestekens worden bij name genoemd en krijgen geen spatie — een ontkenning, geen stilzwijgen. beforeUnits: BIPM levert welke reeksen SI-symbolen zijn, de Taalunie de Nederlandse conventie; eenletterige symbolen blijven weg zoals in fr.json. afterSymbols = [«€», «$»]: LET OP BIJ HET OVERNEMEN VAN fr.json — in het Nederlands gaat het valutateken aan het getal vooraf («€ 179»), in het Frans volgt het erop («2 €»), dus het hoort hier in afterSymbols en niet in beforeUnits. De twee leden zijn die welke het kopje «Valutatekens (€, $)» zelf noemt: een klasse-afleiding over een MET NAME GENOEMDE tweeledige klasse, gemarkeerd naar het voorbeeld van fr.json bij M., MM., Mmes. TOEGESTANE VARIANTEN, vastgelegd: voor het procentteken laat de Taalunie zowel «25 %» als «25%» toe, en bij °C schrijft zij dat er «meestal geen spatie» staat terwijl ISO er één voorschrijft. Dat hoeft niet beslecht te worden: N5 zet alleen een bestaande spatie om en voegt er nooit een in (nbsp.md §3.7). afterShortWords, beforeNumber en beforeWord zijn leeg: deze pagina loopt de leestekens en symbolen één voor één langs en zegt daarover niets, en het Nederlands kent geen tegenhanger van de Poolse sierotka-regel. Geraadpleegd op 18.09.2026."
|
|
80
|
+
},
|
|
81
|
+
{
|
|
82
|
+
"rule": "nbsp",
|
|
83
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, «Prof.( )dr.» (12 mei 2021): «Ja, tussen bijvoorbeeld prof. dr. komt een spatie»; «Alleen als de afgekorte woorden tezamen een afkorting vormen, vervalt de spatie: a.s., a.u.b., m.b.v., n.a.v., z.o.z.»; «In verzorgd zetwerk wordt tussen voorletters en in afkortingen als v.d. bij voorkeur een halve spatie (of dunspatie) geplaatst»",
|
|
84
|
+
"url": "http://taaladvies.net/legacylink.php?id=vraag/678",
|
|
85
|
+
"note": "Onderbouwt abbreviations = [«prof. dr.»] — de enige Nederlandse afkorting met een interne U+0020 die letterlijk is aangetroffen. De gesloten vormen (a.s., a.u.b.) hebben geen interne spatie, zodat N4 daar niets te bevorderen heeft; dáárom telt de lijst één ingang en is hij niet leeg. initialBinding = «none»: de enige Nederlandse uitspraak hierover vraagt om een DUNNE spatie, een waarde die het schema niet kent terwijl N7 U+00A0 invoegt — en, belangrijker, geen geraadpleegde bron zegt dat een voorletter niet door een regelovergang van de achternaam gescheiden mag worden, en dát is wat N7 werkelijk doet. Geraadpleegd op 18.09.2026."
|
|
86
|
+
},
|
|
87
|
+
{
|
|
88
|
+
"rule": "quotes",
|
|
89
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, « Wat-ie wil / wat ie wil » (12 mei 2021) : « Gereduceerde vormen van persoonlijke voornaamwoorden komen voor in de gesproken taal. Bij de weergave daarvan in schrift gebruiken we bij vormen als 'k, '(e)m, 'r, d'r en '(e)t een weglatingsteken of apostrof. »",
|
|
90
|
+
"url": "https://taaladvies.net/wat-ie-wil/",
|
|
91
|
+
"note": "Attests five entries of quotes.elisionClitics.after: k ('k), m and em ('(e)m — the Taalunie's own notation covers both 'm and 'em), r ('r), and t with et ('(e)t). These are word-initial omissions: the mark stands at the start of the reduced form, so the letter run is on the mark's right, which is why they belong in after and not before even though the fragment attaches forward rather than back — both lists are positional, per quotes.md §2. before is empty: m'n, z'n, d'r, zo'n and foto's all have a letter on each side of the mark, so the medial-elision veto already declines them and any before entry would be evidenced-inert; keeping it empty also avoids having to adjudicate m, which occurs on both sides of the mark across the language ('m = hem, m'n = mijn). 'n (een) is not listed — it was not attested by a qualifying Taalunie page in this pass. quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
92
|
+
},
|
|
93
|
+
{
|
|
94
|
+
"rule": "quotes",
|
|
95
|
+
"cite": "Nederlandse Taalunie, Taaladvies.net, « 'S avonds / 's Avonds (hoofdletter?) » (12 mei 2021, op grond van de Woordenlijst 2015) : « Als de zin met een apostrof begint, krijgt het eerstvolgende volledige woord in de zin de hoofdletter », met de voorbeelden « 's Avonds werkt John in een café », « 't Was een mooie dag » en « 'ns Kijken hoe we dat doen »",
|
|
96
|
+
"url": "https://taaladvies.net/s-avonds-of-s-avonds-hoofdletter/",
|
|
97
|
+
"note": "Attests s ('s avonds, from the older genitive des avonds), t ('t was) and ns ('ns kijken). The same entries cover 's-Gravenhage and 's-Hertogenbosch, where a hyphen rather than a space follows the letter run."
|
|
98
|
+
}
|
|
99
|
+
]
|
|
68
100
|
}
|
|
@@ -33,35 +33,68 @@
|
|
|
33
33
|
"nbsp": {
|
|
34
34
|
"beforePunctuation": [],
|
|
35
35
|
"narrowBeforePunctuation": [],
|
|
36
|
-
"afterShortWords": [
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
"o",
|
|
40
|
-
"u",
|
|
41
|
-
"w",
|
|
42
|
-
"z"
|
|
43
|
-
],
|
|
44
|
-
"abbreviations": [
|
|
45
|
-
"lek. med.",
|
|
46
|
-
"inż. górn.",
|
|
47
|
-
"bp sufr."
|
|
48
|
-
],
|
|
49
|
-
"beforeUnits": [
|
|
50
|
-
"%",
|
|
51
|
-
"‰",
|
|
52
|
-
"°C",
|
|
53
|
-
"km",
|
|
54
|
-
"cm",
|
|
55
|
-
"mm",
|
|
56
|
-
"kg",
|
|
57
|
-
"km/h",
|
|
58
|
-
"kWh"
|
|
59
|
-
],
|
|
36
|
+
"afterShortWords": ["a", "i", "o", "u", "w", "z"],
|
|
37
|
+
"abbreviations": ["lek. med.", "inż. górn.", "bp sufr."],
|
|
38
|
+
"beforeUnits": ["%", "‰", "°C", "km", "cm", "mm", "kg", "km/h", "kWh"],
|
|
60
39
|
"beforeNumber": [],
|
|
61
40
|
"beforeWord": [],
|
|
62
|
-
"afterSymbols": [
|
|
63
|
-
"§"
|
|
64
|
-
],
|
|
41
|
+
"afterSymbols": ["§"],
|
|
65
42
|
"initialBinding": "none"
|
|
66
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», §98 «Cudzysłów», wstęp: «W druku i w rękopisach używa się najczęściej cudzysłowu o postaci: [U+201E] [U+201D] (tzw. cudzysłów apostrofowy)»; «W specyficznych zastosowaniach pojawia się też cudzysłów ostrokątny: [U+00BB] [U+00AB] lub [U+00AB] [U+00BB]. Cudzysłów o ostrzach skierowanych do środka jest używany […] w przypadku, gdy występuje cudzysłów w cudzysłowie»",
|
|
48
|
+
"url": "https://sjp.pwn.pl/zasady/Cudzyslow;629866.html",
|
|
49
|
+
"note": "Pierwszy stopień: U+201E otwierający, U+201D zamykający — nie U+201C, jak w niemieckim. Drugi stopień: para «o ostrzach skierowanych do środka», czyli U+00BB otwierający i U+00AB zamykający; para U+00AB…U+00BB jest w tym samym paragrafie przypisana do wyodrębniania znaczeń i do dialogów, a nie do cudzysłowu w cudzysłowie — kierunek jest więc odwrotny do tego, czego spodziewa się czytelnik znający francuskie «…». Kody znaków ustalono, pytając o U+XXXX dla każdego glifu z osobna, a nie na oko. innerSpace = «none»: żadne źródło nie żąda odstępu wewnątrz cudzysłowu. WARIANT DOPUSZCZALNY, odnotowany: Wolański, Poradnia Językowa PWN, «Cudzysłowy drugiego stopnia» (26.05.2020) dopuszcza na drugim stopniu zarówno »…«, jak i «…»; Poradnia jest źródłem doradczym niższej rangi niż «Zasady», więc wybrano formę, którą §98 wiąże wprost z cudzysłowem w cudzysłowie. Trzeci stopień zagnieżdżenia nie jest wyrażalny w locale.schema.json. Konsultowano 18.09.2026."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», [403] 93.7: «Wtrącenia ujmujemy w dwa myślniki»; [408] 93.12, UWAGA: «Półpauza używana w funkcji myślnika ma po bokach spacje»; 21.4, z przykładem «Wyrazy: czarny – biały to antonimy»",
|
|
54
|
+
"url": "https://sjp.pwn.pl/zasady/408-93-12-Myslnik-wyznaczajacy-relacje-miedzy-dwoma-wyrazami-lub-wartosciami;629831.html",
|
|
55
|
+
"note": "Uzasadnia dash.parenthetical = «en-spaced». Odstępy są w źródłach jednomyślne — myślnik ma spacje po obu stronach — natomiast DŁUGOŚĆ kreski nie jest przez PWN rozstrzygnięta: korpus [408] składa myślnik jako U+2014, przykład w 21.4 jako U+2013, a Bańko («przedziały liczbowe», 28.02.2012) pisze wprost, że myślnik «określany jest przez obecność pauz po bokach, a nie przez długość kreski». DECYZJA OPERATORA, 18.09.2026, a nie cytat: wybrano półpauzę U+2013, bo to ona jest w UWADZE [408] nazwana po imieniu w funkcji myślnika i bo taką samą wartość mają już de-DE, en-GB, fi i sv — przy źródle, które nie rozstrzyga, jednolitość w obrębie projektu jest lepszym kryterium niż preferencja. Skutek uboczny, odnotowany: dashes.md §3.4 P5 odrzuca ciąg dokładnie jednej U+2013 w każdym locale, więc istniejące polskie wtrącenia z półpauzą pozostają nietknięte, a zamieniane są U+002D i U+2014. Konsultowano 18.09.2026."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "ranges",
|
|
59
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», [408] 93.12, UWAGA: «W druku liczby arabskie oznaczające przedział od — do rozdziela się półpauzą, czyli kreseczką średniej wielkości, niemającą po bokach odstępów, np. 1914–1918»",
|
|
60
|
+
"url": "https://sjp.pwn.pl/zasady/408-93-12-Myslnik-wyznaczajacy-relacje-miedzy-dwoma-wyrazami-lub-wartosciami;629831.html",
|
|
61
|
+
"note": "Uzasadnia dash.range = «en-tight»: półpauza U+2013 bez spacji. Potwierdzone dwukrotnie: Wolański, «Zakres liczbowy» (1.01.2024) — «Spacje wokół znaku przedziału […] to po prostu błąd ortotypograficzny»; oraz Bańko (28.02.2012), który przywołuje «Edycję tekstów» Wolańskiego w tym samym duchu — przywołanie jest cudze, samej książki nie konsultowano i nie jest tu cytowana. Reguła «ranges» jest domyślnie wyłączona (spec 0.5.0). Konsultowano 18.09.2026."
|
|
62
|
+
},
|
|
63
|
+
{
|
|
64
|
+
"rule": "ellipsis",
|
|
65
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», [396] 92.4: gdy wielokropek zbiega się z pytajnikiem lub wykrzyknikiem, «znaki te w tekście umieszczamy», np. «Czy jest ona tu z wami…?», «Jak strasznie gorąco…!»",
|
|
66
|
+
"url": "https://sjp.pwn.pl/zasady/Wielokropek-obok-innych-znakow-interpunkcyjnych;629818.html",
|
|
67
|
+
"note": "Uzasadnia abbreviatedAfterTerminal = false. Wielokropek to zawsze trzy kropki, a pytajnik i wykrzyknik stoją PO nim — rosyjska forma dwukropkowa «?..» nie jest w polszczyźnie przewidziana przez żadne konsultowane źródło. Sprawdzono to celowo: gdyby polski tę formę miał, byłby drugim po rosyjskim locale z wartością true. Konsultowano 18.09.2026."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"rule": "hyphen",
|
|
71
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», [196]: łącznik przenosi się do następnego wiersza, a na końcu poprzedniego stawia się znak podziału, np. «czarno- / -białe fotografie», «województwo warmińsko- / -mazurskie»",
|
|
72
|
+
"url": "https://sjp.pwn.pl/zasady/Podzial-wyrazu-w-miejscu-lacznika;629553.html",
|
|
73
|
+
"note": "Uzasadnia trzy puste listy. Polszczyzna nie tylko dopuszcza podział w miejscu łącznika — ona go normuje, co jest przeciwieństwem żądania U+2011. Puste listy czynią regułę dowodliwym no-opem dla polskiego (hyphen.md §2). Konsultowano 18.09.2026."
|
|
74
|
+
},
|
|
75
|
+
{
|
|
76
|
+
"rule": "nbsp",
|
|
77
|
+
"cite": "Słownik języka polskiego PWN, «Zasady pisowni i interpunkcji», 21.4: «Między literą kończącą wyraz a znakiem interpunkcyjnym […] nie stosujemy spacji»; 21.1.3: «w skrótach od nazw osób […] np. lek. med., inż. górn., bp sufr.»; 21.2: «Spację pomijamy […] między inicjałami, np. J.K. Bielecki»",
|
|
78
|
+
"url": "https://sjp.pwn.pl/zasady/Stosowanie-spacji-w-skrotach-od-dwoch-i-wiecej-wyrazow;4836486.html",
|
|
79
|
+
"note": "Trzy rzeczy z trzech sąsiadujących paragrafów. (1) beforePunctuation i narrowBeforePunctuation puste: polszczyzna nie stawia przed «:», «;», «!», «?» ani U+00A0, ani U+202F, a źródło odmawia im odstępu wprost — to zaprzeczenie, nie milczenie. (2) abbreviations: trzy skróty zacytowane dosłownie wraz z wewnętrzną spacją, czyli twierdzenie o PRZYNALEŻNOŚCI w rozumieniu nbsp.md §2.1. Świadomie pominięto «i in.» i «z d.»: N3 zajmuje te spacje pierwszy (nbsp.md §3.2), więc wpisy byłyby martwe, nie błędne. (3) initialBinding = «none»: polskie inicjały pisze się bez odstępu («J.K. Bielecki»), więc N7 nie ma czego zamieniać; to rzadki przypadek, w którym «none» opiera się na twierdzeniu pozytywnym, a nie na niecytowalnym zaprzeczeniu. beforeNumber i beforeWord pozostają puste: 21.1, 21.2 i 21.4 przeczytano w całości i żaden nie wiąże skrótu z następującą liczbą ani z wyrazem. Konsultowano 18.09.2026."
|
|
80
|
+
},
|
|
81
|
+
{
|
|
82
|
+
"rule": "nbsp",
|
|
83
|
+
"cite": "dr hab. Adam Wolański, Poradnia Językowa PWN, «Sieroty na końcu wiersza» (8.06.2018): «W wierszach tekstu, które liczą powyżej 40 znaków, nie należy zasadniczo pozostawiać na końcu wiersza jednoliterowych spójników i przyimków (a, i, o, u, w, z, A, I, O, U, W, Z)»",
|
|
84
|
+
"url": "https://sjp.pwn.pl/poradnia/haslo/Sieroty-na-koncu-wiersza;18684.html",
|
|
85
|
+
"note": "Uzasadnia afterShortWords. Lista jest zamknięta i zacytowana dosłownie; wersaliki pominięto, bo N3 dopasowuje pierwszą literę bez względu na wielkość (nbsp.md §3.5 krok 1). RANGA ŹRÓDŁA, powiedziana wprost, bo jest tu słabsza niż gdzie indziej w tym pliku: bezwarunkowy zakaz pochodzi z Poradni, która jest źródłem doradczym, natomiast «Zasady» [204] 54.8.1 są ŁAGODNIEJSZE — «W wąskich łamach jednoliterowe spójniki i przyimki mogą pozostawać na końcu wiersza w tekście ciągłym, natomiast w tytułach […] zawsze powinny być przenoszone» — a Bańko («słowa jednoliterowe na końcu wiersza», 16.02.2003) dodaje, że pozostawienie ich «nie stanowi błędu ortograficznego». Obie kwalifikacje — szerokość łamu i to, czy tekst jest tytułem — są dla silnika niewidoczne. DECYZJA OPERATORA, 18.09.2026: reguła jedzie bezwarunkowo, jak rosyjska afterShortWords, ponieważ wiązanie jednoliterowego przyimka jest w polskim składzie praktyką powszechną, a koszt błędnego zadziałania to jeden U+00A0 tam, gdzie i tak nikt nie chce złamania wiersza. Konsultowano 18.09.2026."
|
|
86
|
+
},
|
|
87
|
+
{
|
|
88
|
+
"rule": "nbsp",
|
|
89
|
+
"cite": "Główny Urząd Miar, «Zasady stosowania oznaczeń i nazw jednostek miar», akapit «Spacja»: «należy zostawiać spację między wartością liczbową a oznaczeniem jednostki» (100 g, 2 L; ale nie: 100g); pkt 12: «pomiędzy [symbolem %] a wartością liczbową wielkości pozostawia się spację»; BIPM, The International System of Units (SI), 9th ed., concise summary — dla składu samej listy symboli",
|
|
90
|
+
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
91
|
+
"note": "Podział ról taki sam jak w en-US, fi, sv, es i pt: BIPM mówi, które łańcuchy są symbolami jednostek SI, a GUM i rozporządzenie — jaka jest polska konwencja odstępu. ZASTRZEŻENIE CO DO ŹRÓDŁA: dobór reguł w dokumencie GUM jest zawężony do towarów paczkowanych, natomiast akapit «Spacja» i pkt 12 powołują się na § 15 i § 8 ust. 4 rozporządzenia, które są prawem powszechnym. WARIANT DOPUSZCZALNY, odnotowany, bo dotyczy najczęstszego wpisu: polska tradycja ortotypograficzna pisze procent bez spacji, co Wolański («Jeszcze raz o spacji lub jej braku po znaku %», 26.11.2021) uznaje za równie poprawne. Rozbieżność nie wymaga rozstrzygnięcia: N5 wyłącznie ZAMIENIA istniejącą spację i nigdy jej nie wstawia (nbsp.md §3.7), więc «5 %» dostaje U+00A0, a «5%» zostaje nietknięte — jedna lista obsługuje obie poświadczone formy. Symbole jednoliterowe (m, g, l, s, A) pominięto zgodnie z praktyką fr.json. Konsultowano 18.09.2026."
|
|
92
|
+
},
|
|
93
|
+
{
|
|
94
|
+
"rule": "nbsp",
|
|
95
|
+
"cite": "dr hab. Adam Wolański, Poradnia Językowa PWN, «Symbol paragrafu i wyróżnienia tytułów aktów prawnych» (20.02.2019): «Symbol §, zastępujący wyraz paragraf, może być użyty tylko przed liczbą zapisaną cyframi. Należy go oddzielać od liczby spacją»",
|
|
96
|
+
"url": "https://sjp.pwn.pl/poradnia/haslo/Symbol-paragrafu-i-wyroznienia-tytulow-aktow-prawnych;19239.html",
|
|
97
|
+
"note": "Uzasadnia afterSymbols = [«§»]. Warunek «tylko przed liczbą zapisaną cyframi» pokrywa się co do joty z własnym strażnikiem N6, który wymaga, by cp[a+k+1] należało do klasy DIGIT (nbsp.md §3.8). Poradnia jest źródłem doradczym; «Zasady» kwestii nie poruszają, więc nie ma sprzeczności. Konsultowano 18.09.2026."
|
|
98
|
+
}
|
|
99
|
+
]
|
|
67
100
|
}
|
|
@@ -14,15 +14,7 @@
|
|
|
14
14
|
},
|
|
15
15
|
"elisionIdioms": [],
|
|
16
16
|
"elisionClitics": {
|
|
17
|
-
"before": [
|
|
18
|
-
"d",
|
|
19
|
-
"n",
|
|
20
|
-
"pel",
|
|
21
|
-
"m",
|
|
22
|
-
"t",
|
|
23
|
-
"lh",
|
|
24
|
-
"sant"
|
|
25
|
-
],
|
|
17
|
+
"before": ["d", "n", "pel", "m", "t", "lh", "sant"],
|
|
26
18
|
"after": []
|
|
27
19
|
}
|
|
28
20
|
},
|
|
@@ -42,26 +34,59 @@
|
|
|
42
34
|
"beforePunctuation": [],
|
|
43
35
|
"narrowBeforePunctuation": [],
|
|
44
36
|
"afterShortWords": [],
|
|
45
|
-
"abbreviations": [
|
|
46
|
-
|
|
47
|
-
],
|
|
48
|
-
"beforeUnits": [
|
|
49
|
-
"%",
|
|
50
|
-
"‰",
|
|
51
|
-
"°C",
|
|
52
|
-
"km",
|
|
53
|
-
"cm",
|
|
54
|
-
"mm",
|
|
55
|
-
"kg",
|
|
56
|
-
"kW",
|
|
57
|
-
"kWh",
|
|
58
|
-
"Hz"
|
|
59
|
-
],
|
|
60
|
-
"beforeNumber": [
|
|
61
|
-
"p."
|
|
62
|
-
],
|
|
37
|
+
"abbreviations": ["p. ex."],
|
|
38
|
+
"beforeUnits": ["%", "‰", "°C", "km", "cm", "mm", "kg", "kW", "kWh", "Hz"],
|
|
39
|
+
"beforeNumber": ["p."],
|
|
63
40
|
"beforeWord": [],
|
|
64
41
|
"afterSymbols": [],
|
|
65
42
|
"initialBinding": "none"
|
|
66
|
-
}
|
|
43
|
+
},
|
|
44
|
+
"sources": [
|
|
45
|
+
{
|
|
46
|
+
"rule": "quotes",
|
|
47
|
+
"cite": "Brasil, Presidência da República, Manual de Redação da Presidência da República, 2.ª ed. rev. e atual., Brasília 2002, §9.1.3.2 «Aspas»: «As aspas têm os seguintes empregos: a) usam-se antes e depois de uma citação textual», com os exemplos impressos «A Constituição […] afirma: “Todo o poder emana do povo, que o exerce por meio de representantes eleitos ou diretamente”», «publicado no “Jornal do Brasil”», «a alínea “a” do artigo 146 da Constituição»",
|
|
48
|
+
"url": "https://legis.sigepe.gov.br/sigepe-bgp-ws-legis/legis-service/download/?id=0000356386-ALPDF%2F2018",
|
|
49
|
+
"note": "Primário U+201C/U+201D, secundário U+2018/U+2019, ambos colados ao texto. FONTE DE NÍVEL INFERIOR, dito sem rodeios: é o manual de redação de um editor estatal, não uma academia — a Academia Brasileira de Letras publica um vocabulário ortográfico, que é lexical e não trata de pontuação, e não foi encontrada fonte brasileira de nível académico que enuncie o par primário. O manual nunca usa aspas angulares em nenhuma das páginas lidas e o seu capítulo III declara-se derivado das gramáticas brasileiras (Bechara, Cegalla, Luft, Kury). DECISÃO DO OPERADOR, 18.9.2026: perante um uso regional atestado mas sem fonte de nível académico, ficheiro próprio com a fonte que existe, claramente identificada, em vez de deixar o Brasil cair no fallback de pt-PT e receber aspas angulares que não usa — a mesma decisão registada na issue #29, «corrige-se acrescentando locales». O nível secundário assenta no exemplo de citação dentro de citação, não numa frase de regra, e é por isso a afirmação mais fraca deste ficheiro. Consultado em 18.9.2026."
|
|
50
|
+
},
|
|
51
|
+
{
|
|
52
|
+
"rule": "dashes",
|
|
53
|
+
"cite": "Manual de Redação da Presidência da República (2.ª ed., 2002), §9.1.3.4 «Travessão»: «O travessão, que é um hífen prolongado (–), é empregado nos seguintes casos: a) substitui parênteses, vírgulas, dois-pontos: O controle inflacionário – meta prioritária do Governo – será ainda mais rigoroso»",
|
|
54
|
+
"url": "https://legis.sigepe.gov.br/sigepe-bgp-ws-legis/legis-service/download/?id=0000356386-ALPDF%2F2018",
|
|
55
|
+
"note": "Justifica dash.parenthetical = \"en-spaced\", e é aqui que pt-BR diverge de pt-PT de forma mensurável. A LARGURA não foi julgada a olho: extraiu-se a camada de texto do PDF e o carácter do exemplo é U+2013, ao passo que o mesmo exemplo estrutural no guia europeu dá U+2014. O papel e o espaçamento coincidem nas duas regiões — o travessão substitui parênteses e leva um espaço de cada lado —, diverge a largura. Consultado em 18.9.2026."
|
|
56
|
+
},
|
|
57
|
+
{
|
|
58
|
+
"rule": "ranges",
|
|
59
|
+
"cite": "Nenhuma fonte brasileira consultada prescreve um sinal de intervalo diferente do hífen já escrito; o Manual de Redação da Presidência da República nada diz sobre intervalos",
|
|
60
|
+
"note": "dash.range = \"none\" por ausência de prescrição, e não por prescrição positiva como em pt-PT — a diferença de fundamento fica registada em vez de ser dissimulada. «none» preserva o que o autor escreveu, que é o comportamento seguro quando nenhuma fonte fala. Consultado em 18.9.2026."
|
|
61
|
+
},
|
|
62
|
+
{
|
|
63
|
+
"rule": "ellipsis",
|
|
64
|
+
"cite": "Acordo Ortográfico da Língua Portuguesa (1990) e Código de Redação Interinstitucional (PT, 2022), ponto 10.4.7 «Reticências»: as reticências completas seguem-se imediatamente a «!» e a «?» («É o dianho!…», «porque não vais ter com o padre?…»); o Manual de Redação da Presidência da República não tem secção sobre reticências — o seu capítulo de pontuação (§9.2.4) trata apenas da vírgula, do ponto-e-vírgula, dos dois-pontos e dos pontos de interrogação e exclamação",
|
|
65
|
+
"url": "https://op.europa.eu/pt/publication-detail/-/publication/01ed788a-d266-11ec-a95f-01aa75ed71a1",
|
|
66
|
+
"note": "Justifica abbreviatedAfterTerminal = false. O passo citado escreve as reticências completas imediatamente a seguir a «!» e a «?» — «!…» e «?…» —, e nenhuma fonte consultada prevê a forma abreviada de dois pontos que o russo usa, pelo que esse ramo de ellipsis.md fica desligado para o português. Nenhuma fonte se pronuncia sobre U+2026 contra três U+002E: essa escolha é do motor e é igual em todas as locales. Consultado em 18.9.2026. A fonte brasileira cala-se sobre as reticências, pelo que o valor vem do lado europeu; não há divergência a registar, há silêncio."
|
|
67
|
+
},
|
|
68
|
+
{
|
|
69
|
+
"rule": "hyphen",
|
|
70
|
+
"cite": "Acordo Ortográfico da Língua Portuguesa (1990), Base XX: «deve, por clareza gráfica, repetir-se o hífen no início da linha imediata»; Manual de Redação da Presidência da República (2.ª ed., 2002), §9.1.3.1, Observação: «Hífen de composição vocabular ou de ênclise e mesóclise é repetido quando coincide com translineação: decreto-/-lei, exigem-/-lhe, far-/-se-á»",
|
|
71
|
+
"url": "http://www.portaldalinguaportuguesa.org/?action=acordo&version=1990",
|
|
72
|
+
"note": "Justifica as três listas vazias, e justifica-as pela positiva, não por ausência de fonte. O português não só permite como PRESCREVE a quebra de linha sobre o hífen de um composto, repetindo-o na linha seguinte: uma forma que é normativamente quebrada nesse ponto não pode ser uma forma cujo hífen tenha de resistir à quebra, que é a única coisa que esta regra consome. A regra é, portanto, um no-op total demonstrável para o português (hyphen.md §2), e qualquer runtime que ligue compostos portugueses com U+2011 está errado. Corroborado pelo Código de Redação Interinstitucional (PT, 2022), ponto 10.2 e ponto 10.4.12 («segunda-/-feira», «salmão-do-/-atlântico»), e pelo Manual de Redação da Presidência da República (Brasil, 2.ª ed., 2002), §9.1.3.1 («decreto-/-lei», «far-/-se-á»). É o campo em que Portugal, o Brasil e o próprio Acordo dizem explicitamente a mesma coisa. Consultado em 18.9.2026."
|
|
73
|
+
},
|
|
74
|
+
{
|
|
75
|
+
"rule": "nbsp",
|
|
76
|
+
"cite": "BIPM, The International System of Units (SI), 9.ª ed., resumo conciso: «A single space is always left between the number and the unit»; Código de Redação Interinstitucional (PT, 2022), anexo A3, pontos 2 e 3, para a lista de símbolos",
|
|
77
|
+
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
78
|
+
"note": "Justifica nbsp.beforeUnits. A locale atesta a PERTENÇA — que estes símbolos são unidades e que se escrevem depois do número com um espaço; o mecanismo (converter esse espaço em U+00A0, nunca o inserir) é da regra e está fixado em nbsp.md §2.1. REPARO CITADO, mais importante do que qualquer inclusão: o ponto 10.9.1 exclui «h» por nome — «horas (o símbolo «h» escreve-se sempre sem ponto, sem espaços): eram as 18h30» —, pelo que «h» NÃO entra na lista e há um caso de conformidade a fixá-lo. Os símbolos de uma só letra (m, g, l, t, s, A, W, V, K) ficam igualmente de fora, embora atestados, pela mesma razão que em fr: não são desambiguáveis sem contexto. Note-se que o português escreve «7 %» com espaço, ao contrário do grego. Consultado em 18.9.2026. Para pt-BR a lista é a mesma: nenhuma fonte brasileira consultada a contradiz, e o BIPM não é uma autoridade sobre o português mas sobre que cadeias são símbolos de unidade do SI."
|
|
79
|
+
},
|
|
80
|
+
{
|
|
81
|
+
"rule": "nbsp",
|
|
82
|
+
"cite": "Nenhuma fonte brasileira consultada exige espaço antes de pontuação dupla, nem ligação de abreviaturas a palavras ou iniciais; listas deliberadamente vazias",
|
|
83
|
+
"note": "Justifica as listas vazias e initialBinding = \"none\", e diz com exatidão em que se apoiam: ao contrário do guia grego, a parte portuguesa NÃO contém nenhuma frase que negue pelo nome o uso francês do espaço antes da pontuação dupla. A posição é «nenhuma fonte o exige», e não «uma fonte o proíbe»; o que se observa é que nenhum exemplo do capítulo o escreve, e isso é uma observação, assinalada como tal. beforeNumber contém apenas «p.»: o ponto 10.9.1 atesta a ligação («p. 24», «Lei n.º 123»), mas «n.º» NÃO pode ser listado, porque a nota (3) do ponto 6.4 do mesmo guia prescreve um «o» sobrescrito e RECUSA expressamente tanto U+00BA como U+00B0 — listar «n.º» seria citar uma fonte que proíbe esse mesmo carácter. A decisão do operador de 18.9.2026 sobre o «n°» francês cobre uma fonte SILENCIOSA e não se estende a uma que fala e recusa. beforeWord fica vazio: o anexo A3 enumera «Sr.», «Dr.», «Prof.», mas limita-se a expandi-los e nenhuma fonte os mostra ligados ao nome seguinte. GAP REGISTADO, não resolvido: o ponto 10.9.1 exige espaço protegido DENTRO dos grupos de algarismos («um total de 12 345 euros»; «este espaço é protegido»), e locale.schema.json não tem campo para isso. Consultado em 18.9.2026."
|
|
84
|
+
},
|
|
85
|
+
{
|
|
86
|
+
"rule": "quotes",
|
|
87
|
+
"cite": "Acordo Ortográfico da Língua Portuguesa (1990), Base XVIII « Do apóstrofo », 1.º : a) « Faz-se uso do apóstrofo para cindir graficamente uma contração ou aglutinação vocabular, quando um elemento ou fração respetiva pertence propriamente a um conjunto vocabular distinto: d'Os Lusíadas, d'Os Sertões; n'Os Lusíadas, n'Os Sertões; pel'Os Lusíadas, pel'Os Sertões » ; b) « d'Ele, n'Ele, d'Aquele, n'Aquele, d'O, n'O, pel'O, m'O, t'O, lh'O » e a série feminina correspondente ; c) « Sant'Ana, Sant'Iago » ; d) « borda-d'água, copo-d'água, estrela-d'alva, mãe-d'água, pau-d'alho, pau-d'arco »",
|
|
88
|
+
"url": "https://www.priberam.pt/docs/AcOrtog90.pdf",
|
|
89
|
+
"note": "Base XVIII is a primary normative text common to both orthographies, which is why pt-PT and pt-BR carry the identical list: d, n, pel (1.º a and b), m, t, lh (1.º b) and sant (1.º c). Base XVIII 2.º bounds the set from the other side — the apóstrofo is not admissible in the combinations of de and em with the definite-article forms (do, da, dele, no, na, nele, num and the rest) — so this is an enumerated set rather than an argument from silence. The archaic anthroponymic forms of 1.º c) (Nun'Álvares, Pedr'Eanes) are cited by the same clause but deliberately not listed: both have a letter on each side of the mark, so quotes.md §3.2's medial-elision veto already declines them and an entry would be evidenced-inert. after is empty: Portuguese enclisis takes a hyphen (Base XVII: amá-lo, dá-se), never an apostrophe. The clause that makes this list load-bearing is 1.º a), which is about book titles: a title set in emphasis rather than italics puts a span boundary immediately after the mark, d'<em>Os Lusíadas</em>, which is precisely the shape the veto reads. Retrieved from a faithful reproduction hosted by Priberam/FLiP, including the official text's own errata footnotes, rather than from the official gazette; the Academia Brasileira de Letras PDF has no extractable text layer. quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
90
|
+
}
|
|
91
|
+
]
|
|
67
92
|
}
|