polytypo 1.4.0 → 1.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/lib/polytypo/data/README.md +15 -0
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +1 -1
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +1 -1
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +50 -1
- data/lib/polytypo/data/fixtures/en-US.json +41 -1
- data/lib/polytypo/data/fixtures/es.json +1 -1
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +1 -1
- data/lib/polytypo/data/fixtures/fr.json +9 -1
- data/lib/polytypo/data/fixtures/it.json +1 -1
- data/lib/polytypo/data/fixtures/locale-resolution.json +1 -1
- data/lib/polytypo/data/fixtures/nl.json +1 -1
- data/lib/polytypo/data/fixtures/pl.json +1 -1
- data/lib/polytypo/data/fixtures/pt-BR.json +1 -1
- data/lib/polytypo/data/fixtures/pt-PT.json +1 -1
- data/lib/polytypo/data/fixtures/ru.json +1 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/uk.json +1 -1
- data/lib/polytypo/data/locales/cs.json +37 -56
- data/lib/polytypo/data/locales/de-CH.json +34 -43
- data/lib/polytypo/data/locales/de-DE.json +32 -47
- data/lib/polytypo/data/locales/el.json +1 -51
- data/lib/polytypo/data/locales/en-GB.json +4 -56
- data/lib/polytypo/data/locales/en-US.json +4 -68
- data/lib/polytypo/data/locales/es.json +29 -66
- data/lib/polytypo/data/locales/fi.json +7 -77
- data/lib/polytypo/data/locales/fr-CA.json +36 -63
- data/lib/polytypo/data/locales/fr.json +41 -70
- data/lib/polytypo/data/locales/it.json +22 -65
- data/lib/polytypo/data/locales/nl.json +29 -61
- data/lib/polytypo/data/locales/pl.json +28 -61
- data/lib/polytypo/data/locales/pt-BR.json +28 -53
- data/lib/polytypo/data/locales/pt-PT.json +28 -55
- data/lib/polytypo/data/locales/registry.json +1 -1
- data/lib/polytypo/data/locales/ru.json +30 -45
- data/lib/polytypo/data/locales/sv.json +4 -63
- data/lib/polytypo/data/locales/uk.json +27 -55
- data/lib/polytypo/data/rules/apostrophe.md +183 -20
- data/lib/polytypo/data/rules/modes.md +29 -4
- data/lib/polytypo/data/rules/nbsp.md +20 -6
- data/lib/polytypo/data/rules/order.json +2 -2
- data/lib/polytypo/data/rules/pipeline-idempotency.md +3 -2
- data/lib/polytypo/data/rules/quotes.md +234 -9
- data/lib/polytypo/data/rules/symbols.md +25 -5
- data/lib/polytypo/engine/rules/apostrophe.rb +14 -1
- data/lib/polytypo/version.rb +1 -1
- metadata +1 -1
|
@@ -40,55 +40,5 @@
|
|
|
40
40
|
"beforeWord": [],
|
|
41
41
|
"afterSymbols": [],
|
|
42
42
|
"initialBinding": "none"
|
|
43
|
-
}
|
|
44
|
-
"sources": [
|
|
45
|
-
{
|
|
46
|
-
"rule": "quotes",
|
|
47
|
-
"cite": "Υπηρεσία Εκδόσεων της Ευρωπαϊκής Ένωσης, Διοργανικό εγχειρίδιο σύνταξης κειμένων (ελληνική έκδοση 2011, τελευταία ενημέρωση 30.4.2012), Μέρος Τέταρτο «Συμβατικοί κανόνες για την ελληνική γλώσσα», §10.1.7 «Εισαγωγικά»: «Σε εισαγωγικά (στο ελληνικό κείμενο προτιμώνται τα διπλά γωνιώδη εισαγωγικά: « ») κλείνονται κυρίως λόγια ή παραθέματα που αναφέρονται αυτολεξεί. […] Στην περίπτωση που χρειάζονται εισαγωγικά μέσα σε κείμενο που είναι ήδη σε εισαγωγικά, τότε για τα εισαγωγικά αυτά χρησιμοποιούνται τα διπλά ανωφερή εισαγωγικά (“ ”), ενώ σε τρίτο επίπεδο εσωτερικά χρησιμοποιούνται τα μονά ανωφερή εισαγωγικά (‘ ’)»",
|
|
48
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
49
|
-
"note": "Primary U+00AB/U+00BB, secondary U+201C/U+201D. Part Four of this guide is a Greek-specific punctuation chapter, not the multilingual part — the distinction matters and is the reason this source is cited for Greek at all. innerSpace is \"none\": the §6.4 spacing table sets «xx» with an ordinary space only outside the guillemets, and the Greek chapter never asks for one inside. The third nesting level the source describes (U+2018/U+2019) is not expressible in locale.schema.json, which has primary and secondary only; that is a schema limit, not a gap in the source. Corroborated by ΥΠΕΠΘ/ΙΤΥΕ, Γραμματική Νέας Ελληνικής Γλώσσας Α΄–Γ΄ Γυμνασίου, §3.3 («Τα εισαγωγικά ( « » ) σημειώνονται…»), whose examples likewise carry no inner space. Retrieved 2026-08-15."
|
|
50
|
-
},
|
|
51
|
-
{
|
|
52
|
-
"rule": "dashes",
|
|
53
|
-
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.8 «Ενωτικό — Παύλα μεσαίου μεγέθους — Μεγάλη παύλα»: για τα αριθμητικά διαστήματα, «Παύλα μεσαίου μεγέθους ή μείον (–) […] γ) για να δηλώσει το διάστημα μεταξύ δύο ορίων (π.χ. άτομα ηλικίας 25–45 ετών) […] δ) στην αναφορά σε περιόδους τουλάχιστον δύο πλήρων ετών (π.χ. 1989–1991) […] Ωστόσο, και στην περίπτωση αυτή μπορεί να χρησιμοποιηθεί και ενωτικό». Ο ίδιος κανόνας για τα αριθμητικά διαστήματα εμφανίζεται ΔΥΟ ΦΟΡΕΣ στην §10.1.8, μία στην υποενότητα «Ενωτικό (-)» και μία στην υποενότητα «Παύλα μεσαίου μεγέθους (–)», και κάθε φορά η υποενότητα παραχωρεί ρητά το άλλο σημείο (paraphrase of the «Ενωτικό» occurrence — its verbatim text was not captured at review time; the «Παύλα» occurrence is quoted above verbatim)· για την παρενθετική χρήση, «Όπως η παρένθεση, η διπλή παύλα δεν χωρίζεται με κενά διαστήματα από τη λέξη, φράση ή πρόταση που περικλείει· αντίθετα, μπαίνουν διαστήματα πριν από την πρώτη και μετά τη δεύτερη παύλα»",
|
|
54
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
55
|
-
"note": "Justifies \"none\" on BOTH fields, which no other locale has — but for TWO DIFFERENT REASONS, and conflating them would misrepresent the source. Range: genuinely ambivalent. The rule appears twice in §10.1.8, once under «Ενωτικό (-)» and once under «Παύλα μεσαίου μεγέθους (–)», and each occurrence explicitly concedes the other mark. That mutual concession is the proof — a source that names both forms in both places is not being vague, it is declining to rank them, and substituting either would express a preference the citation does not carry. Parenthetical: NOT ambivalent. The guide prescribes a U+2014 pair (its own header gives «Alt 0151») with ordinary spaces outside the pair and none on the inner edges. \"none\" here is forced by a SCHEMA LIMIT, not by the source: locale.schema.json's dash enum has no value for asymmetric spacing — \"em-spaced\" puts a space on each side of each dash, which is precisely what this sentence forbids. See spec/rules/dashes.md §6 «el», which states the two justifications separately. SOURCE CONFLICT, recorded and deliberately not settled: this guide prescribes U+2014 for the parenthetical dash, while ΥΠΕΠΘ/ΙΤΥΕ Γραμματική Α΄–Γ΄ Γυμνασίου §3.3 appears to set U+2013 in the same role. The sources array has no way to represent a conflict — it models agreement, not disagreement — so it is recorded here in prose. Resolving it is with the operator and needs a source that ranks the two, not a third that adds a form. Retrieved 2026-08-15."
|
|
56
|
-
},
|
|
57
|
-
{
|
|
58
|
-
"rule": "ellipsis",
|
|
59
|
-
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.9 «Αποσιωπητικά»: «Τα αποσιωπητικά, που είναι πάντοτε τρεις (και όχι περισσότερες) τελείες, χρησιμοποιούνται κυρίως: […]»· παρατήρηση ii): «Μεταξύ των αποσιωπητικών και της λέξης που προηγείται δεν αφήνουμε διάστημα»· παρατήρηση iii): «Όταν τα αποσιωπητικά βρίσκονται στο τέλος της περιόδου δεν προσθέτουμε τελεία»",
|
|
60
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
61
|
-
"note": "Justifies abbreviatedAfterTerminal = false. The source requires exactly three dots always and nowhere provides for the two-dot form after «?» or «!» that Russian uses, so the Russian branch of spec/rules/ellipsis.md must stay off for Greek. Corroborated by Γραμματική Α΄–Γ΄ Γυμνασίου §3.3: «Οι σημειούμενες τελείες είναι πάντα τρεις». Neither source gives any rule for αποσιωπητικά meeting the ερωτηματικό or the θαυμαστικό — checked specifically, in the Ministry of Education grammar and in the Κέντρο Ελληνικής Γλώσσας materials — so the engine's TERMINAL class is deliberately left as {U+0021, U+003F} and the Greek U+003B is not added to it (spaces.md §7.8; adding it would also regress Russian, where Лопатин §154 ties the two-dot form to «?» and «!» by name). Neither source addresses U+2026 versus three U+002E; that choice is the engine's, made identically in every locale, and is not claimed here. KNOWN DIVERGENCE, recorded because a verified source and the engine disagree and neither may stand unremarked: the same §10.1.9, παρατήρηση ii), states «Μεταξύ των αποσιωπητικών και της λέξης που προηγείται δεν αφήνουμε διάστημα» - no space between the ellipsis and the preceding word - and polytypo does NOT honour it. spec/rules/spaces.md §3.4 preserves a space before a dot run in EVERY locale, so «Πράγματι …» is returned as typed. The divergence is deliberate and argued in spaces.md §7.9 and ellipsis.md §6: honouring it needs either locale data in a rule that has none by design, or an ellipsis rule that deletes rather than replaces, which changes that rule's kind and forces a new composition argument. It would be revisited if a second locale wanted the same behaviour, at which point the shape is an ellipsis.noSpaceBefore flag consumed by the ellipsis rule. The source is right about Greek; the engine is declining to act on it, not disputing it. Retrieved 2026-08-15."
|
|
62
|
-
},
|
|
63
|
-
{
|
|
64
|
-
"rule": "spaces",
|
|
65
|
-
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.4 «Διπλή τελεία», παρατήρηση iv): «Στα ελληνικά, πριν από τη διπλή τελεία δεν πρέπει να υπάρχει διάστημα (πράγμα που συμβαίνει, π.χ., στα γαλλικά)»· §10.1.3 «Άνω τελεία», παρατηρήσεις ii) και iii): «Στα ελληνικά, πριν από την άνω τελεία δεν πρέπει να υπάρχει διάστημα»· «Μετά την άνω τελεία αρχίζουμε με μικρό γράμμα»",
|
|
66
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
67
|
-
"note": "Note the exact standing of this citation, because it is weaker than the claim it is often read as supporting. Neither passage is about the ερωτηματικό: they are about the colon and the άνω τελεία. What they establish is that Greek denies the French space-before-punctuation pattern, by name, for the marks they do cover. No source examined addresses spacing before the Greek question mark specifically, so the position is \"no source contradicts it\", NOT \"a source requires it\". Nothing rests on the difference: the Greek question mark is written U+003B, which spec/rules/spaces.md lists in STRIP-BEFORE for the Latin semicolon on its own merits, so the space is stripped under the Latin reading alone and no Greek-specific decision is being made (spaces.md §3.5 part 2). The §6.4 spacing table is deliberately NOT offered as corroboration here — the nbsp source in this same file rejects §6.4 as publisher house style rather than evidence about Greek, and it cannot be house style there and evidence here. Retrieved 2026-08-15."
|
|
68
|
-
},
|
|
69
|
-
{
|
|
70
|
-
"rule": "nbsp",
|
|
71
|
-
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.4 παρατήρηση iv) και §10.1.3 παρατήρηση ii) (ό.π., ρητή αντιπαραβολή με τα γαλλικά)· §10.6 «Συντομογραφίες», γενικός κανόνας γ): «Στις ελληνικές συντομογραφίες με τις οποίες συντέμνονται φράσεις που αποτελούνται από περισσότερες της μίας λέξεις μπαίνει κατά κανόνα τελεία έπειτα από κάθε συντεμνόμενη λέξη» (π.χ. κ.λπ., π.χ., πρβλ., κ.ο.κ.)",
|
|
72
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
73
|
-
"note": "Two empty lists closed by citation rather than by absence of one. beforePunctuation and narrowBeforePunctuation are empty because Greek takes neither U+00A0 nor U+202F before «;» «:» «!» «·» — the source denies the French practice by name. abbreviations is empty because Greek multi-word abbreviations appear to be written with no internal space at all (κ.λπ., not «κ. λπ.»), so there is no U+0020 for the rule to promote. INFERENCE, NOT CITATION: §10.6 γ) states only that a full stop follows each abbreviated word; it says nothing about spacing between them. The no-space form is visible in the guide's own printed examples and in no normative sentence examined. The list would be empty on the absence of a citation alone (docs/PLAN.md §6.1), so nothing rests on the inference — it is labelled because an unlabelled inference in a cite field is the failure mode this field exists to prevent. The remaining lists are empty because no Greek normative source examined states a binding: the one candidate, the §6.4 table binding a number to «%» and «°C», sits in Part Three — «Συμβατικοί κανόνες κοινοί για ΟΛΕΣ τις γλώσσες» — whose own preamble says it replaced diverging national rules for uniform presentation, which makes it publisher house style rather than evidence about Greek. A missing citation is never a licence to guess (docs/PLAN.md §6.1). Observed usage does not contradict that reading, which is worth recording because it was checked rather than assumed: a byte-level scan on 2026-08-15 of the Greek-language home pages of kathimerini.gr, kathimerini.gr/economy, tovima.gr, tanea.gr, efsyn.gr, in.gr, naftemporiki.gr, protagon.gr, lifo.gr, meteo.gr, kedros.gr, onassis.org/el and emst.gr found 183 number-plus-unit occurrences — 140 «%», 41 «°C», 2 «€» — and NOT ONE of them carried a space of any kind, no U+0020, no U+00A0. Greek web practice sets «50%» and «20°C» tight, so beforeUnits would have nothing to promote even if it were populated. Method, so the figure can be re-derived rather than taken on trust: fetch each page, strip script, style and noscript elements and then all tags, decode HTML entities (so that becomes the U+00A0 it denotes rather than disappearing), and count matches of a digit followed by an optional single space character from {U+0020, U+00A0, U+202F, U+2009} followed by the unit. Pages carrying under 200 Greek letters were excluded as not being Greek-language content. That scan is an observation of usage, is not offered as a normative claim, and nothing in this file depends on it: the lists would be empty on the absence of a citation alone. Retrieved 2026-08-15."
|
|
74
|
-
},
|
|
75
|
-
{
|
|
76
|
-
"rule": "hyphen",
|
|
77
|
-
"cite": "Διοργανικό εγχειρίδιο σύνταξης κειμένων (EL, 2011), §10.1.8, υποενότητα «Ενωτικό (-)»: «Η μικρή οριζόντια παύλα, το βραχύτερο σε μήκος από τα τρία σημεία […] σημειώνεται χωρίς κενά σε σχέση με ό,τι προηγείται ή έπεται»",
|
|
78
|
-
"url": "http://publications.europa.eu/resource/cellar/e774ea2a-ef84-4bf6-be92-c9ebebf91c1b.0015.03/DOC_2",
|
|
79
|
-
"note": "Justifies all three lists being empty. The source describes what the hyphen does — syllable division, numeric bounds, appositional compounds such as «απόφαση-πλαίσιο» — without defining any closed list of morphological forms whose hyphen must resist a line break, which is the only thing this rule consumes. No Greek source examined defines one. Empty lists make the rule a provable total no-op for Greek (spec/rules/hyphen.md §2), which is the normal case, not a deficiency. Retrieved 2026-08-15."
|
|
80
|
-
},
|
|
81
|
-
{
|
|
82
|
-
"rule": "apostrophe",
|
|
83
|
-
"cite": "ΥΠΕΠΘ / Ινστιτούτο Τεχνολογίας Υπολογιστών και Εκδόσεων, Γραμματική Νέας Ελληνικής Γλώσσας Α΄, Β΄, Γ΄ Γυμνασίου, §3.3 «Η στίξη — Τα σημεία στίξης»: «Η απόστροφος ( ’ ) χρησιμοποιείται για να δηλώσει ότι ένα φωνήεν έχει παραλειφθεί στη γραφή λόγω της προφοράς»· και, για την ταυτότητα του κωδικού σημείου, The Unicode Standard 17.0, Core Specification, §6.2.7 «Apostrophes»: «When text is set, U+2019 RIGHT SINGLE QUOTATION MARK is preferred as apostrophe, but only U+0027 is present on most keyboards»",
|
|
84
|
-
"url": "https://ebooks.edu.gr/ebooks/v/html/8547/2334/Grammatiki-Neas-Ellinikis-Glossas_A-B-G-Gymnasiou_html-apli/index_B_03.html",
|
|
85
|
-
"note": "Two sources for two separate facts. The Greek grammar establishes that elision and apocope (γι’ αυτό, απ’ την, σ’ αυτό) are marked with an apostrophe; the code point comes only from Unicode, because no Greek normative source examined names one. U+02BC is refused: Unicode §6.2.7 reserves it for use as a modifier letter, e.g. a glottal stop in transliteration. U+0384 GREEK TONOS is refused: no source proposes a diacritic in this role. The rule reads no locale data, so this citation certifies rather than configures — but Greek elision is the shape the rule most often meets in Greek text, and it earns a fixture. Retrieved 2026-08-15."
|
|
86
|
-
},
|
|
87
|
-
{
|
|
88
|
-
"rule": "symbols",
|
|
89
|
-
"cite": "Unicode Consortium, The Unicode Standard, Version 17.0, Core Specification, ch. 7 «Europe-I», §7.2.1 Greek, «Compatibility Punctuation», verbatim: «Therefore, use of U+037E and U+0387 is not necessary for interoperating with legacy Greek data, and their use is not generally encouraged for representation of Greek punctuation.» The preceding sentence of that subsection — to the effect that the two characters have canonical equivalences to U+003B and U+00B7, so that normalised Greek text loses the distinction — is given here as PARAPHRASE, not as quotation: it could not be confirmed word for word at review time. Its substance does not rest on the prose in any case; it is entailed directly by UnicodeData.txt, where U+037E carries the canonical decomposition mapping 003B and U+0387 carries 00B7, each as a singleton",
|
|
90
|
-
"url": "https://www.unicode.org/versions/Unicode17.0.0/core-spec/chapter-7/",
|
|
91
|
-
"note": "Code-point identity only, which is the one thing this source is authoritative for. Note the split standing of the citation: one sentence is verbatim, one is paraphrase backed by UnicodeData.txt rather than by the Core Specification's wording, and the file says which is which. Everything the rules do rests on the machine-readable field, not on the prose. The ερωτηματικό is written U+003B, not U+037E; the άνω τελεία is written U+00B7, not U+0387. Consequence for the rules: no rule may emit U+037E or U+0387, and none may rewrite one to the character it decomposes to — that is normalisation, forbidden by docs/ARCHITECTURE.md §4.3 and performed anyway by any downstream NFC pass. The full argument, including why U+00B7 must not join the spaces rule's STRIP-BEFORE set, is in spec/rules/spaces.md §3.5. Retrieved 2026-08-15."
|
|
92
|
-
}
|
|
93
|
-
]
|
|
43
|
+
}
|
|
94
44
|
}
|
|
@@ -15,7 +15,9 @@
|
|
|
15
15
|
"elisionIdioms": [],
|
|
16
16
|
"elisionClitics": {
|
|
17
17
|
"before": [],
|
|
18
|
-
"after": [
|
|
18
|
+
"after": [
|
|
19
|
+
"s"
|
|
20
|
+
]
|
|
19
21
|
}
|
|
20
22
|
},
|
|
21
23
|
"dash": {
|
|
@@ -67,59 +69,5 @@
|
|
|
67
69
|
"beforeWord": [],
|
|
68
70
|
"afterSymbols": [],
|
|
69
71
|
"initialBinding": "none"
|
|
70
|
-
}
|
|
71
|
-
"sources": [
|
|
72
|
-
{
|
|
73
|
-
"rule": "quotes",
|
|
74
|
-
"cite": "University of Oxford Style Guide, section “Quotation marks”",
|
|
75
|
-
"url": "https://www.ox.ac.uk/about/the-university/brand/style-guide/punctuation",
|
|
76
|
-
"note": "Verbatim: “Use single quotation marks for direct speech or a quote, and double quotation marks for direct speech or a quote within that.” Example given: ‘I have never been to Norway,’ he said, ‘but I have heard it described as “the Wales of the North”.’ This settles the ❓ in PLAN.md §7 in favour of single-first. The convention is genuinely divided in British practice — several national newspapers and trade publishers set double-first — but the cited Oxford authority (and New Hart’s Rules behind it) is unambiguous, and citation, not head-count, is the tie-breaker per ARCHITECTURE.md §8."
|
|
77
|
-
},
|
|
78
|
-
{
|
|
79
|
-
"rule": "dashes",
|
|
80
|
-
"cite": "University of Oxford Style Guide, section “Dashes and hyphens”",
|
|
81
|
-
"url": "https://www.ox.ac.uk/about/the-university/brand/style-guide/punctuation",
|
|
82
|
-
"note": "Verbatim, m-dash (—): “Do not use; use an n-dash instead.” n-dash (–): “Use in a pair in place of round brackets or commas, surrounded by spaces” (example: “It was – as far as I could tell – the only example of its kind.”), and “Use to link concepts or ranges of numbers, with no spaces either side” (example: “The salary for the post is £25,000–£30,000.”). Corroborated by CMOS Shop Talk, 23 January 2024: “In British style, spaced en dashes – like this – are more common.”"
|
|
83
|
-
},
|
|
84
|
-
{
|
|
85
|
-
"rule": "ellipsis",
|
|
86
|
-
"cite": "University of Oxford Style Guide, section “Ellipsis”",
|
|
87
|
-
"url": "https://www.ox.ac.uk/about/the-university/brand/style-guide/punctuation",
|
|
88
|
-
"note": "`abbreviatedAfterTerminal` is false. The Oxford guide describes the ellipsis as a single mark used for omitted text, a pause or a trailing off, and notes “an exclamation mark or a question mark can and should follow the ellipsis if required” — i.e. no merged two-dot form after terminal punctuation. That form is Russian-only."
|
|
89
|
-
},
|
|
90
|
-
{
|
|
91
|
-
"rule": "nbsp",
|
|
92
|
-
"cite": "BIPM, The International System of Units (SI), 9th ed. (2019), concise summary, “The language of science: using the SI to express the values of quantities”",
|
|
93
|
-
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
94
|
-
"note": "Verbatim: “A single space is always left between the number and the unit.” Used here because the University of Oxford Style Guide sets measurements closed up (“The average height of a woman in the UK is 1.61m.”, “worth 10% of the available marks”) and is therefore silent on binding a space it does not itself write. The nbsp rule never inserts a space — it only makes an author-typed space non-breaking — so the two positions do not conflict."
|
|
95
|
-
},
|
|
96
|
-
{
|
|
97
|
-
"rule": "nbsp",
|
|
98
|
-
"cite": "University of Oxford Style Guide, “Specific abbreviations — people’s initials”",
|
|
99
|
-
"url": "https://www.ox.ac.uk/about/the-university/brand/style-guide/punctuation",
|
|
100
|
-
"note": "`initialBinding` is \"none\" (spec 0.6.0: the field used to be the boolean bindInitials, false). Oxford says only “Use a space to separate each initial” (examples: J R R Tolkien, C S Lewis — note, without full stops). No British authority consulted prescribes a non-breaking space there, so the true value is set rather than inferred from the American rule. Deliberate divergence from en-US, where Chicago states the nonbreaking space explicitly."
|
|
101
|
-
},
|
|
102
|
-
{
|
|
103
|
-
"rule": "nbsp",
|
|
104
|
-
"cite": "No normative source found for English short-word binding; lists deliberately empty",
|
|
105
|
-
"note": "`afterShortWords`, `beforePunctuation`, `narrowBeforePunctuation`, `abbreviations` and `afterSymbols` are empty by design. Oxford instructs “close up spaces and don’t use full stops in abbreviations (eg 6pm)”, so there are no multi-token abbreviations with internal spaces to bind."
|
|
106
|
-
},
|
|
107
|
-
{
|
|
108
|
-
"rule": "nbsp",
|
|
109
|
-
"cite": "University of Oxford Style Guide, “Abbreviations, contractions and acronyms”",
|
|
110
|
-
"url": "https://www.ox.ac.uk/about/the-university/brand/style-guide/punctuation",
|
|
111
|
-
"note": "`beforeNumber` and `beforeWord` are empty. Oxford says “Don’t use full stops after any abbreviations, contractions or acronyms”, so the British forms this field would hold (p., no., fig., ch., Mr., Dr.) are not even written with a full stop in this style, and nothing in the guide requires a non-breaking space after them. Per PLAN.md §7 the Mr./Dr. + name binding is an optional rule defaulting off; with no per-entry toggle in the schema, and no normative source requiring it, the honest value is an empty list."
|
|
112
|
-
},
|
|
113
|
-
{
|
|
114
|
-
"rule": "hyphen",
|
|
115
|
-
"cite": "No English authority prescribes non-breaking hyphens for particular word forms; lists deliberately empty",
|
|
116
|
-
"note": "`prefixes`, `suffixes` and `compounds` are empty. The field exists for languages with a closed, normative set of hyphenated morphological forms (Russian кое-, -таки, из-под). The Oxford guide’s hyphen section describes when to hyphenate (adjectival phrases before a noun, verb participles) and never restricts breaking at the resulting hyphen. An English list here would be an invented sample."
|
|
117
|
-
},
|
|
118
|
-
{
|
|
119
|
-
"rule": "quotes",
|
|
120
|
-
"cite": "The Chicago Manual of Style, 7.16 “Possessive form of most nouns” (section number and title confirmed against the publisher's own chapter-7 table of contents; cited as 7.16–19 and 7.22 by the CMOS Q&A on the 18th-edition site). Verbatim, as reproduced by CMOS Shop Talk, “Section 7.16 in the Spotlight”, 12 May 2015: “The possessive of most singular nouns is formed by adding an apostrophe and an s.” Reconfirmed for the current edition by CMOS Shop Talk, 31 August 2021: “Chicago adds an apostrophe and an s to form the possessive of a singular noun, including singular nouns ending in s” (see CMOS 7.16)",
|
|
121
|
-
"url": "https://cmosshoptalk.com/2021/08/31/introducing-the-chicago-manual-of-style-for-perfectit/",
|
|
122
|
-
"note": "Supports quotes.elisionClitics.after = [\"s\"], identical to en-US. LABELLED GAP: en-GB rests on an American authority because no British one was reachable. New Hart's Rules: The Oxford Style Guide (ed. Anne Waddingham, OUP) is print and paywalled; ox.ac.uk's style guide returned HTTP 403, OUP's own free companion site for New Hart's Rules returned an empty body, and Oxford Reference is paywalled. Every British statement of the rule found was a third-party paraphrase and therefore not citable. The gap is narrow rather than papered over: the single point on which US and UK possessive practice actually diverges is whether a second s follows a sibilant (Dickens's, Chicago's preference per 7.16–19, versus Dickens', the alternative practice per 7.22), and that divergence cannot reach this entry, because the veto matches the maximal LETTER run to the RIGHT of the mark — Burns's fires on s, Burns' has no letter there and fires nothing. So the two locales take the identical entry not by assumption but because the two practices coincide wherever this field is consulted. Replace with New Hart's Rules if the book is obtained. Note also that en-GB's primary close is U+2019, so this is the locale where a wrong entry could collide with genuine quotation — a one-element list is the minimum that closes issue #53."
|
|
123
|
-
}
|
|
124
|
-
]
|
|
72
|
+
}
|
|
125
73
|
}
|
|
@@ -21,7 +21,9 @@
|
|
|
21
21
|
],
|
|
22
22
|
"elisionClitics": {
|
|
23
23
|
"before": [],
|
|
24
|
-
"after": [
|
|
24
|
+
"after": [
|
|
25
|
+
"s"
|
|
26
|
+
]
|
|
25
27
|
}
|
|
26
28
|
},
|
|
27
29
|
"dash": {
|
|
@@ -73,71 +75,5 @@
|
|
|
73
75
|
"beforeWord": [],
|
|
74
76
|
"afterSymbols": [],
|
|
75
77
|
"initialBinding": "chain"
|
|
76
|
-
}
|
|
77
|
-
"sources": [
|
|
78
|
-
{
|
|
79
|
-
"rule": "quotes",
|
|
80
|
-
"cite": "The Chicago Manual of Style, 18th ed., 6.11 (quotations within quotations); CMOS Q&A, topic “Quotations”",
|
|
81
|
-
"url": "https://www.chicagomanualofstyle.org/qanda/data/faq/topics/Quotations/faq0025.html",
|
|
82
|
-
"note": "US practice: double quotation marks at the first level, single at the second. CMOS 6.11: “When single quotation marks are nested within double quotation marks, and two of the marks appear next to each other, a space between the two marks, though not strictly required, aids legibility.” That optional legibility space is a typesetting refinement and is not expressed in this file."
|
|
83
|
-
},
|
|
84
|
-
{
|
|
85
|
-
"rule": "quotes",
|
|
86
|
-
"cite": "Chicago Manual of Style Online Q&A, topic “Special Characters”: “If you want, for example, an apostrophe rather than an opening single quotation mark at the beginning of a string (e.g., the apostrophe before the n in “rock ’n’ roll”), you must tell the computer that that's what you want.”",
|
|
87
|
-
"url": "https://www.chicagomanualofstyle.org/qanda/data/faq/topics/SpecialCharacters/faq0002.html",
|
|
88
|
-
"note": "Supports quotes.elisionIdioms = [{ left: \"rock\", elided: \"n\", right: \"roll\" }]: the leading mark of the elision must not be resolved as an opening quotation mark. CMOS states the mark's function, not a code point; U+2019 follows from spec/rules/apostrophe.md, which is unchanged. Reproduced verbatim, including CMOS's own U+2019 in “rock ’n’ roll” — not straightened to U+0027 — because it is a direct quotation of the source's printed text. Edition/section not fully cited: the page reports 17th ed. in-body while the site titles itself 18th ed., and the numbered CMOS section is paywalled."
|
|
89
|
-
},
|
|
90
|
-
{
|
|
91
|
-
"rule": "quotes",
|
|
92
|
-
"cite": "American Heritage Dictionary of the English Language, entry “rock and roll”: headword line “rock-and-roll or rock 'n' roll”, the elided variant spaced as rock 'n' roll with a mark on each side of the bare word n.",
|
|
93
|
-
"url": "https://www.ahdictionary.com/word/search.html?q=rock+and+roll",
|
|
94
|
-
"note": "Attests the spaced two-mark orthography of the \"rock\"/\"n\"/\"roll\" idiom only — this citation supports no other elisionIdioms entry. In particular it does not attest \"fish 'n' chips\", which is not listed here and would need its own citation before being added; the field is a closed set of individually evidenced idioms, not a general elision-word list."
|
|
95
|
-
},
|
|
96
|
-
{
|
|
97
|
-
"rule": "dashes",
|
|
98
|
-
"cite": "The Chicago Manual of Style — CMOS Shop Talk, “Hyphens and Dashes: A Refresher”, 23 January 2024",
|
|
99
|
-
"url": "https://cmosshoptalk.com/2024/01/23/hyphens-and-dashes-a-refresher/",
|
|
100
|
-
"note": "Verbatim: “In Chicago style, such dashes consist of em dashes—like this—with no space before or after.” And: “When consecutive digits express a range, however, they are separated in Chicago style not by hyphens but by en dashes, as in the range 3–5.”"
|
|
101
|
-
},
|
|
102
|
-
{
|
|
103
|
-
"rule": "ellipsis",
|
|
104
|
-
"cite": "The Chicago Manual of Style, 18th ed., ellipses (three dots)",
|
|
105
|
-
"url": "https://cmosshoptalk.com/2021/06/15/navigating-spaces-in-manuscripts-and-beyond/",
|
|
106
|
-
"note": "`abbreviatedAfterTerminal` is false: the two-dot form after `!`/`?` is a Russian convention with no counterpart in English usage. Chicago’s spaced ellipsis (. . .) is a print-typesetting variant, not a separate character sequence this spec models."
|
|
107
|
-
},
|
|
108
|
-
{
|
|
109
|
-
"rule": "nbsp",
|
|
110
|
-
"cite": "The Chicago Manual of Style — CMOS Shop Talk, “Navigating Spaces in Manuscripts and Beyond”, 15 June 2021",
|
|
111
|
-
"url": "https://cmosshoptalk.com/2021/06/15/navigating-spaces-in-manuscripts-and-beyond/",
|
|
112
|
-
"note": "Chicago recommends a nonbreaking space “between two or more initials in a name like ‘E. B. White’” (hence initialBinding: \"chain\" — spec 0.6.0: the field used to be the boolean bindInitials; Chicago's own “two or more” wording is the citation for the chain requirement, not merely for turning N7 on) and “between a numeral and an abbreviated unit of measure, as in ‘10 kg’” (hence beforeUnits)."
|
|
113
|
-
},
|
|
114
|
-
{
|
|
115
|
-
"rule": "nbsp",
|
|
116
|
-
"cite": "BIPM, The International System of Units (SI), 9th ed. (2019), concise summary, “The language of science: using the SI to express the values of quantities”",
|
|
117
|
-
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
118
|
-
"note": "Verbatim: “A single space is always left between the number and the unit.” This is the authority for the unit list itself; Chicago is the authority for making that space non-breaking. `%` is included even though Chicago sets percentages closed up (45%): the rule only converts a space the author already typed, and a spaced `45 %` in English is a deliberate scientific-style choice worth binding."
|
|
119
|
-
},
|
|
120
|
-
{
|
|
121
|
-
"rule": "nbsp",
|
|
122
|
-
"cite": "No normative source found for English short-word binding; list deliberately empty",
|
|
123
|
-
"note": "`afterShortWords`, `beforePunctuation` and `narrowBeforePunctuation` are empty by design. English typography has no equivalent of the Russian rule against leaving short prepositions at line end, and no English authority prescribes a space before terminal punctuation. `afterSymbols` is empty because no consulted source states that `§`/`№` bind to a following number in English."
|
|
124
|
-
},
|
|
125
|
-
{
|
|
126
|
-
"rule": "hyphen",
|
|
127
|
-
"cite": "No English authority prescribes non-breaking hyphens for particular word forms; lists deliberately empty",
|
|
128
|
-
"note": "`prefixes`, `suffixes` and `compounds` are empty. This field exists for languages with a closed, normative set of hyphenated morphological forms (Russian кое-, -таки, из-под). English hyphenation is productive and open-ended — self-, ex-, -like, well-known and thousands more — so any list would be an invented sample, not a specification. Chicago’s hyphenation guidance (its hyphenation table) governs whether to hyphenate, not whether the resulting hyphen may break, and Chicago explicitly permits breaking at a hyphen."
|
|
129
|
-
},
|
|
130
|
-
{
|
|
131
|
-
"rule": "nbsp",
|
|
132
|
-
"cite": "The Chicago Manual of Style Online, Q&A, topic “Word Division”: “we don’t include any recommendation that says you must carry an abbreviation that ends in a period over to the next line when the sentence continues beyond the abbreviation. So a line is allowed to end with an ‘a.m.’ or ‘Jr.’ or the like that occurs in the middle of a sentence.”",
|
|
133
|
-
"url": "https://www.chicagomanualofstyle.org/qanda/data/faq/topics/WordDivision/faq0007.html",
|
|
134
|
-
"note": "`beforeNumber` and `beforeWord` are empty, and Chicago does not merely omit reference abbreviations from its nonbreaking-space guidance — it states positively that no such rule exists, so `No. 5`, `p. 12` and `pp. 12–14` are refused on the source’s own words rather than on an argument from silence (issue #30). The same answer attributes the practice to typesetters rather than to Chicago: “some typesetters will use a nonbreaking space to prevent a compound like ‘St. Louis’ or ‘Dr. Smith’ from breaking at the end of a line” — which is why `beforeWord` stays empty too; the named exception is initials in a name, carried by `initialBinding`. CMOS Shop Talk, “Navigating Spaces in Manuscripts and Beyond” (15 June 2021), cited above for `initialBinding` and `beforeUnits`, introduces its four contexts with “can be helpful in the following contexts, among others”, so it must not be described as an exhaustive list."
|
|
135
|
-
},
|
|
136
|
-
{
|
|
137
|
-
"rule": "quotes",
|
|
138
|
-
"cite": "The Chicago Manual of Style, 7.16 “Possessive form of most nouns” (section number and title confirmed against the publisher's own chapter-7 table of contents; cited as 7.16–19 and 7.22 by the CMOS Q&A on the 18th-edition site). Verbatim, as reproduced by CMOS Shop Talk, “Section 7.16 in the Spotlight”, 12 May 2015: “The possessive of most singular nouns is formed by adding an apostrophe and an s.” Reconfirmed for the current edition by CMOS Shop Talk, 31 August 2021: “Chicago adds an apostrophe and an s to form the possessive of a singular noun, including singular nouns ending in s” (see CMOS 7.16)",
|
|
139
|
-
"url": "https://cmosshoptalk.com/2021/08/31/introducing-the-chicago-manual-of-style-for-perfectit/",
|
|
140
|
-
"note": "Supports the single entry quotes.elisionClitics.after = [\"s\"] — the possessive s, not the contraction enclitics. t, ll, re, ve, d and m are deliberately absent: nobody writes `won`'t, so they have no attested occurrence in the shape this veto is scoped to, whereas possessivising a set-off identifier (`x`'s printer) is everyday technical prose and is the shape reported in canonical issue #53. Chicago states the rule, not a code point; U+2019 follows from apostrophe.md §3.3, which is unchanged. Edition and section provenance is mixed deliberately and should not be tidied into a single number: the numbered paragraph itself is paywalled, the verbatim wording above is the 16th edition's (the 2015 post carries an editor's note saying so), and the 2021 post plus the 18th-edition Q&A are what establish that 7.16 is still the general rule. CMOS 7.22's alternative practice (Burns' rather than Burns's) does not affect this entry: the veto matches the maximal LETTER run to the RIGHT of the mark, so Burns's fires on s while Burns' has no letter there and fires nothing. CMOS 7.29 “Possessive with italicized or quoted terms” is a closer fit to the reported shape than 7.16 and would strengthen this citation, but its text is paywalled and was not read — only its title, from the table of contents."
|
|
141
|
-
}
|
|
142
|
-
]
|
|
78
|
+
}
|
|
143
79
|
}
|
|
@@ -44,72 +44,35 @@
|
|
|
44
44
|
"q. e. p. d.",
|
|
45
45
|
"et al."
|
|
46
46
|
],
|
|
47
|
-
"beforeUnits": [
|
|
48
|
-
|
|
49
|
-
|
|
47
|
+
"beforeUnits": [
|
|
48
|
+
"%",
|
|
49
|
+
"‰",
|
|
50
|
+
"°C",
|
|
51
|
+
"km",
|
|
52
|
+
"cm",
|
|
53
|
+
"mm",
|
|
54
|
+
"kg",
|
|
55
|
+
"mg",
|
|
56
|
+
"mL",
|
|
57
|
+
"min",
|
|
58
|
+
"dB",
|
|
59
|
+
"kW",
|
|
60
|
+
"kWh"
|
|
61
|
+
],
|
|
62
|
+
"beforeNumber": [
|
|
63
|
+
"art.",
|
|
64
|
+
"cap.",
|
|
65
|
+
"núm.",
|
|
66
|
+
"nro.",
|
|
67
|
+
"pág.",
|
|
68
|
+
"págs."
|
|
69
|
+
],
|
|
70
|
+
"beforeWord": [
|
|
71
|
+
"Sr.",
|
|
72
|
+
"Sra.",
|
|
73
|
+
"Srta."
|
|
74
|
+
],
|
|
50
75
|
"afterSymbols": [],
|
|
51
76
|
"initialBinding": "chain"
|
|
52
|
-
}
|
|
53
|
-
"sources": [
|
|
54
|
-
{
|
|
55
|
-
"rule": "quotes",
|
|
56
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «comillas», apartado 1: «En los textos impresos, se recomienda utilizar en primera instancia las comillas angulares, reservando los otros tipos para cuando deban entrecomillarse partes de un texto ya entrecomillado. En este caso, las comillas simples se emplearán en último lugar»; «Las comillas inglesas y las simples se escriben en la parte alta del renglón, mientras que las angulares se escriben centradas»",
|
|
57
|
-
"url": "https://www.rae.es/dpd/comillas",
|
|
58
|
-
"note": "El DPD establece QUÉ CLASE de signo corresponde a cada nivel y el orden de anidamiento; no nombra ningún punto de código. La identidad U+201C/U+201D procede de Unicode, autoridad para eso y para nada más. Contra las comillas rectas de máquina de escribir el DPD aporta una descripción de trazado — «en la parte alta del renglón» — y locale.schema.json las prohíbe además por su cuenta. El tercer nivel que el DPD describe (‘ ’) no es expresable: el esquema tiene primary y secondary, no tres; límite del esquema, no laguna de la fuente, igual que en el. ADVERTENCIA DE OBTENCIÓN: rae.es devuelve HTTP 403 a un cliente automático — obstáculo técnico, no obra de pago — y el texto se leyó por un proxy de extracción sobre esta URL canónica; el mismo proxy devolvió la abreviatura de «número» de dos formas distintas en dos pasadas, por lo que ningún campo de este archivo afirma un punto de código no ASCII que no se haya podido fijar por otra vía. Consultado el 18.09.2026."
|
|
59
|
-
},
|
|
60
|
-
{
|
|
61
|
-
"rule": "quotes",
|
|
62
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «comillas», apartado 1: «Las comillas se escriben pegadas a la primera y la última palabra del periodo que enmarcan, y separadas por un espacio de las palabras o signos que las preceden o las siguen»",
|
|
63
|
-
"url": "https://www.rae.es/dpd/comillas",
|
|
64
|
-
"note": "Justifica innerSpace = \"none\" en los dos pares: «pegadas» excluye todo espacio interior, ni U+00A0 ni U+202F. El español no sigue aquí el uso francés. elisionIdioms queda vacío porque el español no tiene ningún modismo de la forma {left, elided, right}: la entrada «apóstrofo» del mismo DPD afirma que el signo «apenas se usa en el español actual» y lo limita a textos antiguos, al habla reproducida y a nombres de otras lenguas. Consultado el 18.09.2026."
|
|
65
|
-
},
|
|
66
|
-
{
|
|
67
|
-
"rule": "dashes",
|
|
68
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «raya», apartado 2: «Como el resto de los signos dobles, se escriben pegadas a la primera y a la última palabra del periodo que enmarcan, y separadas por un espacio de la palabra o signo que las precede o las sigue»",
|
|
69
|
-
"url": "https://www.rae.es/dpd/raya",
|
|
70
|
-
"note": "parenthetical = \"none\" POR UN LÍMITE DEL ESQUEMA, no por falta de fuente: la fuente es explícita y prescribe espacios ORDINARIOS FUERA del par y ninguno en los bordes interiores («Llegó —por fin— a casa»). El enum no expresa esa asimetría: \"em-spaced\" pone un espacio a cada lado de CADA raya, que es lo que la frase prohíbe, y \"em-tight\" suprime los exteriores que la frase exige. Mismo caso que el griego (dashes.md §6) y que el italiano de este mismo lote; al repetirse en tres idiomas independientes deja de ser una rareza y pasa a ser una carencia del esquema, registrada aparte para resolverse en su propia versión. Mientras tanto \"none\" es lo único que el archivo puede decir con verdad, y no autoriza a aproximar. Consultado el 18.09.2026."
|
|
71
|
-
},
|
|
72
|
-
{
|
|
73
|
-
"rule": "ranges",
|
|
74
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «guion», apartado 3 b): «Indica intervalos expresados tanto en números arábigos como en romanos: las páginas 23-45; durante los siglos x-xii»; apartado 1: «Signo ortográfico en forma de pequeña línea horizontal (-)»",
|
|
75
|
-
"url": "https://www.rae.es/dpd/guion",
|
|
76
|
-
"note": "range = \"none\" por una razón DISTINTA de la de parenthetical, y confundirlas tergiversaría las fuentes. El signo español del intervalo es el guion, U+002D — el carácter que el autor ya escribió —, y el enum solo ofrece em, en o none: «no sustituir nada» es aquí a la vez lo honesto y lo correcto, porque «23-45» ya está bien escrito. El enum sí podría expresar un valor equivocado: \"en-tight\" convertiría esa forma en una que el DPD no prescribe. Consultado el 18.09.2026."
|
|
77
|
-
},
|
|
78
|
-
{
|
|
79
|
-
"rule": "ellipsis",
|
|
80
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «puntos suspensivos», apartado 1: «Signo de puntuación formado por tres puntos consecutivos (…) ―y solo tres―»",
|
|
81
|
-
"url": "https://www.rae.es/dpd/puntos%20suspensivos",
|
|
82
|
-
"note": "Justifica abbreviatedAfterTerminal = false: «y solo tres» es una exigencia expresa de número, y el español no dispone en ningún apartado de la forma de dos puntos tras «?» o «!» que sí usa el ruso. El apartado 3.4 de la misma entrada regula el ORDEN de los signos, no su número. La fuente no se pronuncia sobre U+2026 frente a tres U+002E: esa elección es del motor y es igual en todas las locales. Consultado el 18.09.2026."
|
|
83
|
-
},
|
|
84
|
-
{
|
|
85
|
-
"rule": "hyphen",
|
|
86
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «guion», apartados 1, 2 y 3: el guion une palabras y segmentos, divide la palabra al final de línea, indica intervalos y separa grupos de cifras",
|
|
87
|
-
"url": "https://www.rae.es/dpd/guion",
|
|
88
|
-
"note": "Justifica las tres listas vacías. La entrada enumera los usos del guion y en ninguno define una clase cerrada de formas morfológicas cuyo guion no deba partirse al final de línea, que es lo único que consume esta regla; al contrario, en español el guion es el signo del propio corte de línea. Con las listas vacías la regla es un no-op total demostrable (hyphen.md §2), que es el caso normal y no una carencia. Consultado el 18.09.2026."
|
|
89
|
-
},
|
|
90
|
-
{
|
|
91
|
-
"rule": "nbsp",
|
|
92
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «signos de interrogación y exclamación», apartado 1: «Se escriben pegados a la primera y la última palabra del periodo que enmarcan, y separados por un espacio de las palabras que los preceden o los siguen»",
|
|
93
|
-
"url": "https://www.rae.es/dpd/signos%20de%20interrogaci%C3%B3n%20y%20exclamaci%C3%B3n",
|
|
94
|
-
"note": "Justifica beforePunctuation y narrowBeforePunctuation vacíos por dos motivos independientes. PRIMERO, la fuente: no hay espacio antes de «?» ni de «!», ni después de «¿» ni de «¡»; el uso francés queda descartado por la norma, no por silencio. SEGUNDO, y con independencia de lo anterior, el esquema no podría expresar lo contrario: la restricción Q-P exige que todo punto de código de estas listas pertenezca a CLOSEISH, y U+00BF y U+00A1 no lo son — son puntuación de APERTURA, y el único mecanismo de «espacio tras un signo de apertura» del esquema es quotes.innerSpace. NO ES EXPRESABLE, que es distinto de no estar atestiguado. afterShortWords queda vacío por una razón más débil y se dice cuál: ninguna fuente consultada exige que una palabra breve no quede al final de línea; lo resolvería el capítulo «División de palabras al final de línea» de la Ortografía, no consultado. afterSymbols vacío: ninguna fuente española consultada liga «§» a un número. Consultado el 18.09.2026."
|
|
95
|
-
},
|
|
96
|
-
{
|
|
97
|
-
"rule": "nbsp",
|
|
98
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «símbolo», apartado 5.4: «Los símbolos deben escribirse pospuestos a la cifra a la que acompañan y dejando un blanco de separación: 33 dB, 125 m², 4 H»; «se separa de ella y se pega al símbolo de la escala si esta se especifica: 27° (por veintisiete grados), pero 27 °C»; apartado 5.6: «No deben escribirse en líneas diferentes la cifra y el símbolo que la acompaña: ⊗3 / $, ⊗7 / km»",
|
|
99
|
-
"url": "https://www.rae.es/dpd/s%C3%ADmbolo",
|
|
100
|
-
"note": "Justifica beforeUnits, y el apartado 5.6 es además la base normativa de la INDIVISIBILIDAD, dicha por la RAE y en español. El apartado 5.4 impone una exclusión que el archivo respeta: «°» a secas NO figura, porque «27°» va pegado; solo «°C». Dos advertencias de alcance. (1) La lista es MÁS ESTRECHA QUE LA FUENTE: se omiten los símbolos de una sola letra (m, g, s, h, L, A, H), que no se desambiguan sin contexto — criterio de fr —, pese a que el ejemplo del propio apartado es «4 H». (2) No se incluye ningún símbolo monetario, y no por prudencia sino por cita: el apartado 5.5 distingue expresamente el uso de España («3 £», pospuesto y con blanco) del de América («£3», antepuesto y sin blanco), y esta locale no lleva subetiqueta de región. Ese apartado se cita también en POSITIVO: el DPD sabe marcar la divergencia regional donde existe, luego los campos sobre los que no la marca son seguros para un «es» regionalmente neutro. Consultado el 18.09.2026."
|
|
101
|
-
},
|
|
102
|
-
{
|
|
103
|
-
"rule": "nbsp",
|
|
104
|
-
"cite": "BIPM, The International System of Units (SI), 9th ed. (2019), concise summary: «A single space is always left between the number and the unit»",
|
|
105
|
-
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
106
|
-
"note": "Autoridad sobre la COMPOSICIÓN de la lista de unidades, igual que en en-US, en-GB, fi y sv; el DPD «símbolo» 5.4 y 5.6 es la autoridad sobre la convención española y sobre la indivisibilidad. El BIPM no se invoca como autoridad sobre el español — no lo es —, sino sobre qué cadenas son símbolos de unidad del SI. Si ambas fuentes discreparan sobre el español mandaría la RAE; no discrepan. Consultado el 18.09.2026."
|
|
107
|
-
},
|
|
108
|
-
{
|
|
109
|
-
"rule": "nbsp",
|
|
110
|
-
"cite": "RAE y ASALE, Diccionario panhispánico de dudas, entrada «abreviatura»: «Tampoco deben aparecer en renglones diferentes la abreviatura y el término del que esta depende: ⊗15 / págs., ⊗Sr. / Pérez»; «Cuando la abreviatura corresponde a una expresión compleja, se separan mediante un espacio las letras que representan cada una de las palabras que la integran»; «Las iniciales de los componentes del nombre propio de una persona son también abreviaturas formadas por truncamiento»",
|
|
111
|
-
"url": "https://www.rae.es/dpd/abreviatura",
|
|
112
|
-
"note": "Una sola cita para cuatro campos, porque una sola regla los gobierna. abbreviations: la fuente exige el espacio interior y prohíbe partirlo; las formas listadas están atestiguadas en esta entrada o en la Lista de abreviaturas de El buen uso del español. beforeNumber y beforeWord: «⊗15 / págs.» y «⊗Sr. / Pérez» son los dos ejemplos de la propia fuente, uno en cada orden. beforeWord se limita a la clase cerrada de los tratamientos de cortesía: la regla vale para CUALQUIER abreviatura, pero enumerarlas todas dispararía falsos positivos en la subrregla de mayor riesgo del esquema (nbsp.md §3.12). NO SE INCLUYEN «n.º», «D.ª» ni «Sr.ª»: la RAE prescribe una «o volada», es decir un TRAZADO y no un punto de código, y el texto obtenido devolvió dos formas distintas en dos pasadas; adivinar sería peor que omitir. initialBinding = \"chain\" ES UNA DECISIÓN DEL OPERADOR (18.09.2026), NO UNA CITA: el DPD no contiene ninguna calificación de «dos o más» — enuncia la regla general y establece que las iniciales SON abreviaturas, lo que apunta más bien a \"single\" —, pero su único ejemplo con iniciales es «J. A. Pérez», de dos, y ante una fuente que no decide se eligió el valor mayoritario en este proyecto (en-US, de-DE, de-CH, ru) y el de menor superficie de falsos positivos, porque \"chain\" no liga una mayúscula con punto aislada. En «J. A. Pérez» ambos valores dan el mismo resultado; solo difieren ante un inicial suelto. Consultado el 18.09.2026."
|
|
113
|
-
}
|
|
114
|
-
]
|
|
77
|
+
}
|
|
115
78
|
}
|
|
@@ -34,7 +34,9 @@
|
|
|
34
34
|
"beforePunctuation": [],
|
|
35
35
|
"narrowBeforePunctuation": [],
|
|
36
36
|
"afterShortWords": [],
|
|
37
|
-
"abbreviations": [
|
|
37
|
+
"abbreviations": [
|
|
38
|
+
"fil. maist."
|
|
39
|
+
],
|
|
38
40
|
"beforeUnits": [
|
|
39
41
|
"%",
|
|
40
42
|
"‰",
|
|
@@ -60,81 +62,9 @@
|
|
|
60
62
|
],
|
|
61
63
|
"beforeNumber": [],
|
|
62
64
|
"beforeWord": [],
|
|
63
|
-
"afterSymbols": [
|
|
65
|
+
"afterSymbols": [
|
|
66
|
+
"§"
|
|
67
|
+
],
|
|
64
68
|
"initialBinding": "none"
|
|
65
|
-
}
|
|
66
|
-
"sources": [
|
|
67
|
-
{
|
|
68
|
-
"rule": "quotes",
|
|
69
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Lainausmerkit”",
|
|
70
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/lainausmerkit/",
|
|
71
|
-
"note": "Verbatim: “Suomenkielisessä tekstissä käytettävät kokolainausmerkit ovat kaarevat ”, ja ne ovat samanmuotoiset lainatun jakson alussa ja lopussa.” Both the opening and the closing mark are U+201D. The page notes that books and newspapers sometimes use angle marks (»…») as a design choice; the Kotus recommendation for Finnish text is ”…”."
|
|
72
|
-
},
|
|
73
|
-
{
|
|
74
|
-
"rule": "quotes",
|
|
75
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Puolilainausmerkki”",
|
|
76
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/puolilainausmerkki/",
|
|
77
|
-
"note": "Verbatim: “Puolilainausmerkkiä käytetään myös lainausmerkkinä kokolainausmerkkien sisällä”, and “Suomenkielisissä teksteissä käytettävä puolilainausmerkki on ’, ei ‛ eikä '.” The page gives the Windows input sequence Alt+0146, which is U+2019, confirming the code point for both the opening and the closing secondary mark."
|
|
78
|
-
},
|
|
79
|
-
{
|
|
80
|
-
"rule": "dashes",
|
|
81
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Ajatusviiva virkkeen välimerkkinä”",
|
|
82
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/ajatusviiva-virkkeen-valimerkkina/",
|
|
83
|
-
"note": "Verbatim: “Tällaisen virkkeen välimerkkinä käytetyn ajatusviivan molemmin puolin tulee välilyönti.” Hence `parenthetical: en-spaced` — Finnish uses the en dash (ajatusviiva, n-viiva), not an em dash."
|
|
84
|
-
},
|
|
85
|
-
{
|
|
86
|
-
"rule": "dashes",
|
|
87
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Ajanilmaukset: aikavälit (1.1.–31.1.)”, citing standard SFS 4175",
|
|
88
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/ajanilmaukset-aikavalit-1-1-31-1/",
|
|
89
|
-
"note": "Verbatim: “Numero- ja merkkistandardi SFS 4175 suosittaa rajakohdan merkkinä lyhempää ajatusviivaa eli ns. n-viivaa”, set tight against the endpoints (ma–pe). Known exception not expressible in this schema: when the endpoints are themselves multi-word, Kotus spaces the dash (“ma 10.4. – pe 21.4.”). `range: en-tight` therefore describes the numeric case only."
|
|
90
|
-
},
|
|
91
|
-
{
|
|
92
|
-
"rule": "ellipsis",
|
|
93
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Kolme pistettä”",
|
|
94
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/kolme-pistetta/",
|
|
95
|
-
"note": "`abbreviatedAfterTerminal` is false. Kotus describes the mark as three dots marking an unfinished sentence, a missing element or a continuing list; the two-dot form after `!`/`?` is a Russian convention and appears in no Kotus guidance."
|
|
96
|
-
},
|
|
97
|
-
{
|
|
98
|
-
"rule": "nbsp",
|
|
99
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Prosenttimerkki (%)”",
|
|
100
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/prosenttimerkki/",
|
|
101
|
-
"note": "Finnish writes a space between the numeral and the percent sign (10,5 %), unlike English, which sets it closed up. This attests membership: `%` is a Finnish unit-like sign that follows a numeral after a space, which is what `beforeUnits` expresses. `‰` is included on the same reading."
|
|
102
|
-
},
|
|
103
|
-
{
|
|
104
|
-
"rule": "nbsp",
|
|
105
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Pykälät ja pykälämerkki (§)”",
|
|
106
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/pykalat-ja-pykalamerkki/",
|
|
107
|
-
"note": "Kotus gives **three** notations as alternatives, not two: `6. §` (general language, with the ordinal full stop), `§ 6`, and `6 §` (legal language, where the ordinal stop is conventionally dropped). An earlier revision of this note named only the last two and read as though Finnish never sets the sign before the number; that was wrong. A space separates the number and the sign in all three. Because Kotus attests both orders, `§` is a member of **both** lists: `beforeUnits` covers `6 §` (N5, binds to a preceding number) and `afterSymbols` covers `§ 6` (N6, binds to a following number). The two sub-rules are disjoint by their own left/right tests, so the double membership is safe rather than contradictory. The third form, `6. §`, binds under neither — N5 requires a digit immediately before the space and there is an ordinal full stop there — which is an algorithmic gap held with the spec author, not something this file can express. Contrast Finnish with Swedish, where Språkrådet’s example is `§ 7` and only the sign-first order is attested."
|
|
108
|
-
},
|
|
109
|
-
{
|
|
110
|
-
"rule": "nbsp",
|
|
111
|
-
"cite": "Oikeusministeriö, Lainkirjoittajan opas §24.4 “Merkeistä ja taivutusmuodoista lakikielessä”",
|
|
112
|
-
"url": "https://lainkirjoittaja.finlex.fi/24-lakikieli/24-4/",
|
|
113
|
-
"note": "Second source for the legal form: statutory drafting sets `2 §` and inflects it with a colon (`3–5 §:ssä`). Corroborates the number-first notation Kotus gives, and with it `§`’s membership of `beforeUnits`."
|
|
114
|
-
},
|
|
115
|
-
{
|
|
116
|
-
"rule": "nbsp",
|
|
117
|
-
"cite": "BIPM, The International System of Units (SI), 9th ed. (2019), concise summary, “The language of science: using the SI to express the values of quantities”; Kotus, Kielitoimiston ohjepankki: “Prosenttimerkki (%)” and “Pykälät ja pykälämerkki (§)”",
|
|
118
|
-
"url": "https://www.bipm.org/documents/20126/41483022/SI-Brochure-9-concise-EN.pdf",
|
|
119
|
-
"note": "Membership claim, per nbsp.md §2.1: these tokens are units of measurement in Finnish and are written after a numeral with a space. BIPM verbatim: “A single space is always left between the number and the unit.” Kotus supplies the Finnish-specific members that SI does not cover — the percent sign takes a space (10,5 %), unlike English, and the section sign follows the number (6 §). That is the whole of what this list can express: which tokens are units. Whether the space is then made non-breaking is N5’s mechanism, fixed for every locale by the rule and not variable per locale, so no citation is owed for it. **Correction kept on the record, because it was real:** an earlier revision of this file justified the list with the sentence “Numeron ja lyhenteen väliin tulee välilyönti”, attributed to the Kotus page on numeroiden ryhmittely. That sentence is not on that page — it came from a search-engine summary and was never checked against the source. What that page’s “Sitova välilyönti” section actually says is “Tekstinkäsittelyohjelmissa pitkät luvut saa pysymään samalla rivillä käyttämällä sitovaa eli yhdistävää välilyöntiä”, which is about holding the digit groups of one long number together (100 000) — the class nbsp.md §7.11 declines to implement — and not about a number and a following unit. The page is no longer cited here. Kielikello and the rest of the ohjepankki were searched for a Finnish number+unit line-breaking statement and none was found; SFS 4175 may contain one, but it is a paid standard whose text was not read, and it is not cited on the strength of a secondary summary."
|
|
120
|
-
},
|
|
121
|
-
{
|
|
122
|
-
"rule": "nbsp",
|
|
123
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Lyhenteet: pisteelliset”",
|
|
124
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/lyhenteet-pisteelliset/",
|
|
125
|
-
"note": "Verbatim: “Yhdyssanan lyhennettyjen osien väliin ei tule välilyöntiä, toisin kuin sanaliiton osien väliin” — a compound word abbreviates closed up (sos.dem., dipl.ins.) while a word-group abbreviation keeps the space (fil. maist.). The `abbreviations` list is therefore deliberately minimal: only the one form quoted verbatim on that page is included. Kotus’s full Lyhenneluettelo contains further space-bearing academic-title abbreviations (fil. tri, valt. maist. and similar); they are omitted rather than guessed at, and should be added by whoever can transcribe that list directly. The membership claim here is precise: `fil. maist.` is **one abbreviation that contains a space**, and Kotus states its internal shape directly. A form that a source merely shows adjacent to another token would not qualify."
|
|
126
|
-
},
|
|
127
|
-
{
|
|
128
|
-
"rule": "nbsp",
|
|
129
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: välimerkit and lyhenteet, read for membership in each remaining list",
|
|
130
|
-
"url": "https://kielitoimistonohjepankki.fi/asiasana/valimerkit/",
|
|
131
|
-
"note": "`beforePunctuation` and `narrowBeforePunctuation` are empty because Finnish sets no space before a punctuation mark — there are no members to list. `afterShortWords` is empty because Finnish has no closed class of short words that may not end a line; its function words are inflectional suffixes rather than separate particles, so the list has no candidates. `beforeWord` is empty for want of an attested membership claim: no Kotus page shows titles such as hra or prof. as forms that bind to a following name. `beforeNumber` is a genuine open question rather than a settled empty: Kotus does write s. 12 and kuva 3, which under nbsp.md §2.1 is a membership claim of the right shape, and the earlier justification for leaving it empty (that Kotus does not require the space to be non-breaking) is no longer a valid reason. It is left empty here only because populating it is a live behaviour change that belongs in its own round with N9’s false-positive surface reviewed. `initialBinding` is \"none\" (spec 0.6.0: the field used to be the boolean bindInitials, false), and §2.1 makes plain that this field is unciteable in principle — it expresses nothing but mechanism, and no typographic authority states it in terms of a code point. Kotus attests the space in J. K. Paasikivi and says nothing further. The value is therefore a judgement, held at \"none\" to match the same call made for en-GB, and it should be settled by the operator across all locales at once rather than per file."
|
|
132
|
-
},
|
|
133
|
-
{
|
|
134
|
-
"rule": "hyphen",
|
|
135
|
-
"cite": "Kotimaisten kielten keskus (Kotus), Kielitoimiston ohjepankki: “Yhdysmerkki eli yhdysviiva”",
|
|
136
|
-
"url": "https://kielitoimistonohjepankki.fi/ohje/yhdysmerkki-eli-yhdysviiva/",
|
|
137
|
-
"note": "`prefixes`, `suffixes` and `compounds` are empty. Kotus’s hyphen guidance is entirely about when to write a hyphen (compounds with numerals or abbreviations, identical adjoining vowels, word-group compounds such as “avaimet käteen -sopimus”), never about forbidding a line break at one. Finnish compounding is productive, so there is no closed normative list to encode — unlike Russian, where кое-, -таки and из-под are a fixed inventory. Nothing here, and no Kotus page found, prescribes U+2011."
|
|
138
|
-
}
|
|
139
|
-
]
|
|
69
|
+
}
|
|
140
70
|
}
|
|
@@ -45,70 +45,43 @@
|
|
|
45
45
|
"compounds": []
|
|
46
46
|
},
|
|
47
47
|
"nbsp": {
|
|
48
|
-
"beforePunctuation": [
|
|
48
|
+
"beforePunctuation": [
|
|
49
|
+
":"
|
|
50
|
+
],
|
|
49
51
|
"narrowBeforePunctuation": [],
|
|
50
52
|
"afterShortWords": [],
|
|
51
|
-
"abbreviations": [
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
"
|
|
55
|
-
|
|
53
|
+
"abbreviations": [
|
|
54
|
+
"p. ex."
|
|
55
|
+
],
|
|
56
|
+
"beforeUnits": [
|
|
57
|
+
"%",
|
|
58
|
+
"‰",
|
|
59
|
+
"€",
|
|
60
|
+
"°C",
|
|
61
|
+
"km",
|
|
62
|
+
"cm",
|
|
63
|
+
"mm",
|
|
64
|
+
"kg",
|
|
65
|
+
"km/h",
|
|
66
|
+
"kWh"
|
|
67
|
+
],
|
|
68
|
+
"beforeNumber": [
|
|
69
|
+
"art.",
|
|
70
|
+
"fig.",
|
|
71
|
+
"n°",
|
|
72
|
+
"N°"
|
|
73
|
+
],
|
|
74
|
+
"beforeWord": [
|
|
75
|
+
"M.",
|
|
76
|
+
"MM.",
|
|
77
|
+
"Mme",
|
|
78
|
+
"Mmes",
|
|
79
|
+
"Mlle",
|
|
80
|
+
"Mlles"
|
|
81
|
+
],
|
|
82
|
+
"afterSymbols": [
|
|
83
|
+
"§"
|
|
84
|
+
],
|
|
56
85
|
"initialBinding": "single"
|
|
57
|
-
}
|
|
58
|
-
"sources": [
|
|
59
|
-
{
|
|
60
|
-
"rule": "nbsp",
|
|
61
|
-
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Espacement avant et après les signes de ponctuation et les symboles » : « L'Office québécois de la langue française opte pour l'absence d'espace devant le point-virgule, le point d'exclamation et le point d'interrogation », l'espace fine restant admise « lorsqu'elle est disponible » ; le deux-points est précédé d'une espace insécable",
|
|
62
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/22039/la-typographie/espacement/espacement-avant-et-apres-les-signes-de-ponctuation-et-les-symboles",
|
|
63
|
-
"note": "C'est le point qui distingue fr-CA de fr (France) : l'usage de France, documenté dans spec/locales/fr.json à partir de Jacques André, insère une espace fine insécable (U+202F) avant « ; », « ! » et « ? ». L'OQLF choisit explicitement l'absence d'espace pour l'usage québécois ; l'espace fine n'est mentionnée que comme variante disponible, pas comme la règle par défaut. Le champ narrowBeforePunctuation est donc vide : la règle « spaces » retire toute espace ordinaire tapée devant ces signes (comportement indépendant de la locale, spec/rules/spaces.md), et aucune espace n'est réinsérée ici — le résultat net est « Bonjour! », sans espace, ce qui correspond au choix explicite de l'OQLF plutôt qu'à un oubli. Le deux-points reste précédé d'une espace insécable ordinaire (U+00A0), comme en France : les deux sources OQLF consultées s'accordent sur ce point avec Jacques André."
|
|
64
|
-
},
|
|
65
|
-
{
|
|
66
|
-
"rule": "nbsp",
|
|
67
|
-
"cite": "Jacques André, Petites leçons de typographie, éd. du jobet, § 5.1.3 « Autres emplois de l'espace insécable » — voir spec/locales/fr.json pour la citation complète",
|
|
68
|
-
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
69
|
-
"note": "beforeUnits, beforeNumber, beforeWord, abbreviations et initialBinding (spec 0.6.0 : le champ s'appelait auparavant le booléen bindInitials ; fr conserve \"single\", cité au § 5.1.3 d'André pour « N. Bourbaki ») sont repris tels quels de spec/locales/fr.json : aucune source spécifiquement québécoise n'a été consultée pour ces cas (seuls le point-virgule/point d'exclamation/point d'interrogation et le deux-points ont été vérifiés séparément auprès de l'OQLF, voir l'entrée précédente). L'OQLF confirme par ailleurs, sur la même page que l'entrée précédente, une espace insécable devant les symboles d'unités SI et le pourcentage, ce qui corrobore — sans le prouver intégralement — le maintien de ces listes pour fr-CA."
|
|
70
|
-
},
|
|
71
|
-
{
|
|
72
|
-
"rule": "quotes",
|
|
73
|
-
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Généralités sur les guillemets » : « Les guillemets français (« »), appelés chevrons à cause de leur forme, sont ceux que l'on utilise normalement dans un texte français » ; « Une espace insécable sépare les guillemets ouvrants et fermants du texte » ; pour une citation à l'intérieur d'une citation, « on utilise successivement les guillemets anglais doubles, puis anglais simples »",
|
|
74
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/23363/la-ponctuation/guillemets/generalites-sur-les-guillemets",
|
|
75
|
-
"note": "Chevrons, second niveau et espace intérieure identiques à fr (France) ; aucun écart québécois trouvé. La largeur exacte de l'espace intérieure (U+00A0 ordinaire vs U+202F fine) n'est pas non plus tranchée ici — l'OQLF emploie le terme générique « espace insécable » ; la valeur « nbsp » est reprise de fr.json en attendant que ce point soit résolu pour les deux locales."
|
|
76
|
-
},
|
|
77
|
-
{
|
|
78
|
-
"rule": "dashes",
|
|
79
|
-
"cite": "Jacques André, Petites leçons de typographie, § 2.2 et tableau 1 — voir spec/locales/fr.json pour la citation complète ; corroboré par l'OQLF, « Tiret : mise en valeur » : « Les tirets sont précédés et suivis d'un espacement. »",
|
|
80
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/index.php?id=23378",
|
|
81
|
-
"note": "Repris de fr : tiret cadratin espacé pour l'incise, aucune valeur de plage vérifiée (range: none). L'OQLF confirme l'espacement mais ne distingue pas le cadratin du demi-cadratin ; aucun écart québécois trouvé sur ce point."
|
|
82
|
-
},
|
|
83
|
-
{
|
|
84
|
-
"rule": "ellipsis",
|
|
85
|
-
"cite": "Jacques André, Petites leçons de typographie, § 5.1.2 et tableau 2 — voir spec/locales/fr.json pour la citation complète",
|
|
86
|
-
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
87
|
-
"note": "Repris tel quel de fr ; aucune source québécoise consultée séparément pour ce point, aucun écart identifié dans les sources OQLF lues."
|
|
88
|
-
},
|
|
89
|
-
{
|
|
90
|
-
"rule": "hyphen",
|
|
91
|
-
"cite": "Jacques André, Petites leçons de typographie, tableau 1 — voir spec/locales/fr.json pour la citation complète",
|
|
92
|
-
"url": "http://jacques-andre.fr/faqtypo/lessons.pdf",
|
|
93
|
-
"note": "Repris tel quel de fr : aucune source consultée n'impose un trait d'union insécable pour une classe fermée de formes morphologiques françaises, en France comme au Québec."
|
|
94
|
-
},
|
|
95
|
-
{
|
|
96
|
-
"rule": "nbsp",
|
|
97
|
-
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Abréviation de numéro » : « L'abréviation courante du nom numéro est no ou No, avec la lettre o minuscule en exposant » ; « On n'abrège le mot numéro que s'il suit immédiatement le nom qu'il détermine » ; exemple : « C'est le billet no 852410 qui a valu le gros lot à sa détentrice »",
|
|
98
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/25450/les-abreviations-et-les-symboles/les-abreviations/cas-particuliers-dabreviations/abreviation-de-numero",
|
|
99
|
-
"note": "Deux choses distinctes, et il faut les tenir séparées. Ce que la source établit : « n° » est une abréviation, pas un symbole, d'où beforeNumber et non afterSymbols — l'argument est aussi mécanique, un « ° » seul placé dans afterSymbols serait inerte puisque N6 exige que le code point précédent ne soit pas ALNUM, or c'est la lettre « n » (nbsp.md §3.8). Ce que la source n'établit pas : la séquence exacte de code points. L'OQLF prescrit un o minuscule en exposant, un tracé et non un code point, et ne nomme nulle part le signe de degré ni la ligature « numéro » (U+2116). Pour la France, rien : le § 5.1.3 de Jacques André, relu intégralement, ne mentionne ni « numéro » ni « no », et le Lexique de l'Imprimerie nationale est un ouvrage imprimé qui n'a pas pu être consulté. Décision de l'opérateur (18.09.2026) : quand la règle est derrière une source payante, polytypo retient l'usage le plus répandu et l'applique de la même façon partout, l'uniformité primant la conformité au canon, et la décision est signalée comme telle au lieu d'être présentée comme une citation. L'usage le plus répandu est « n » + U+00B0, c'est ce que l'on tape et ce que l'on trouve dans les textes en ligne ; N9 n'ayant aucune tolérance de casse, « n° » et « N° » sont deux entrées. La phrase « on ne sépare pas les numéros par des espacements » de la même page vise les tranches de chiffres à l'intérieur du nombre (« billet no 852410 »), pas l'espace entre l'abréviation et le nombre."
|
|
100
|
-
},
|
|
101
|
-
{
|
|
102
|
-
"rule": "quotes",
|
|
103
|
-
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Cas où l'élision est obligatoire » (dernière mise à jour : 2015) : « L'élision ne touche que des mots grammaticaux, habituellement courts, et que les voyelles a, e et i. » L'élision est obligatoire pour le déterminant et le pronom la et le, pour les pronoms je, me, te, se et ce placés devant le verbe, pour si devant il(s), pour que et les conjonctions qui le contiennent, pour jusque, pour la préposition de et pour l'adverbe de négation ne",
|
|
104
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/21737/lorthographe/elision-et-apostrophe/elision-obligatoire",
|
|
105
|
-
"note": "Justifie dix des treize entrées de quotes.elisionClitics.before : l (la, le), j (je), m (me), t (te), s (se et si), c (ce), qu (que), jusqu (jusque), d (de), n (ne). Chaque entrée est la plage de lettres maximale à gauche de la marque, d'où jusqu et non qu pour « jusqu'à » : sous une comparaison par plage maximale, une entrée qu ne peut pas apparier « jusqu' ». presqu (presqu'île) et quelqu (quelqu'un) sont volontairement exclus, faute d'attestation dans les articles relus, et l'exclusion est sans coût : dans les deux mots la marque a une lettre de chaque côté. La liste after est vide — la BDL définit l'élision comme l'effacement de la voyelle finale d'un mot, donc la marque suit toujours le mot élidé, et l'enclise française prend le trait d'union (donne-moi, y a-t-il), jamais l'apostrophe. Le cas porté par cette liste est l'élision devant un élément en ligne, l'<em>idée</em> : sans elle, quotes appariait la marque comme une citation et invertissait la paire englobante (issue #53). L'OQLF est l'autorité du Québec : pour fr-CA la citation est directe, pour fr elle tient lieu de substitut, le Lexique de l'Imprimerie nationale et Le Bon Usage n'étant pas consultables en ligne — même lacune que celle déjà signalée ailleurs dans ce fichier. quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
106
|
-
},
|
|
107
|
-
{
|
|
108
|
-
"rule": "quotes",
|
|
109
|
-
"cite": "Office québécois de la langue française, Banque de dépannage linguistique, « Élision de lorsque, puisque et quoique » (2020) : « les conjonctions lorsque, puisque et quoique s'élident obligatoirement devant il(s), elle(s), on, un et une, et aussi devant la préposition en » ; l'élision généralisée devant tout mot commençant par une voyelle ou un h muet « demeure facultative »",
|
|
110
|
-
"url": "https://vitrinelinguistique.oqlf.gouv.qc.ca/23623/lorthographe/elision-et-apostrophe/elision-de-lorsque-puisque-et-quoique",
|
|
111
|
-
"note": "Justifie les trois entrées lorsqu, puisqu, quoiqu. L'élision obligatoire devant il(s), elle(s), on, un, une et en suffit à l'appartenance : la liste étant uniquement restrictive, le caractère facultatif de l'élision généralisée ne change rien. Là encore la plage maximale impose des entrées distinctes : qu n'apparie pas « lorsqu' ». quotes.md §3.2's span-boundary elision veto (spec 1.4.0) reads this list only when the mark's other literal neighbour is modes.md §3.2's inline span boundary, so every entry is matched against the maximal LETTER run on one side of the mark and nothing else. The list is decline-only: the worst it can do is refuse a pairing quotes would otherwise have formed."
|
|
112
|
-
}
|
|
113
|
-
]
|
|
86
|
+
}
|
|
114
87
|
}
|