polytypo 1.5.0 → 1.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/lib/polytypo/data/.not-canonical +4 -0
- data/lib/polytypo/data/README.md +21 -6
- data/lib/polytypo/data/VERSION +1 -1
- data/lib/polytypo/data/fixtures/cs.json +1 -1
- data/lib/polytypo/data/fixtures/de-CH.json +1 -1
- data/lib/polytypo/data/fixtures/de-DE.json +1 -1
- data/lib/polytypo/data/fixtures/el.json +1 -1
- data/lib/polytypo/data/fixtures/en-GB.json +1 -1
- data/lib/polytypo/data/fixtures/en-US.json +1 -1
- data/lib/polytypo/data/fixtures/es.json +1 -1
- data/lib/polytypo/data/fixtures/fi.json +1 -1
- data/lib/polytypo/data/fixtures/fr-CA.json +1 -1
- data/lib/polytypo/data/fixtures/fr.json +1 -1
- data/lib/polytypo/data/fixtures/it.json +1 -1
- data/lib/polytypo/data/fixtures/locale-resolution.json +13 -1
- data/lib/polytypo/data/fixtures/nl.json +1 -1
- data/lib/polytypo/data/fixtures/pl.json +1 -1
- data/lib/polytypo/data/fixtures/pt-BR.json +1 -1
- data/lib/polytypo/data/fixtures/pt-PT.json +1 -1
- data/lib/polytypo/data/fixtures/ru.json +1 -1
- data/lib/polytypo/data/fixtures/sv.json +1 -1
- data/lib/polytypo/data/fixtures/tr.json +250 -0
- data/lib/polytypo/data/fixtures/uk.json +1 -1
- data/lib/polytypo/data/locales/registry.json +3 -2
- data/lib/polytypo/data/locales/tr.json +72 -0
- data/lib/polytypo/data/rules/apostrophe.md +56 -14
- data/lib/polytypo/data/rules/dashes.md +4 -1
- data/lib/polytypo/data/rules/locale-resolution.md +3 -3
- data/lib/polytypo/data/rules/order.json +1 -1
- data/lib/polytypo/data/rules/quotes.md +32 -5
- data/lib/polytypo/version.rb +1 -1
- metadata +4 -1
checksums.yaml
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
SHA256:
|
|
3
|
-
metadata.gz:
|
|
4
|
-
data.tar.gz:
|
|
3
|
+
metadata.gz: 0bb01ab8478817898277fc5b6d1f1654cc847a3f82ab1975e043e1a58409d2db
|
|
4
|
+
data.tar.gz: 559fdc9ca8756129e28d9762f2f3cc5666384f4c470960d84f181448dba56a0b
|
|
5
5
|
SHA512:
|
|
6
|
-
metadata.gz:
|
|
7
|
-
data.tar.gz:
|
|
6
|
+
metadata.gz: 69173cb92815bc03d120ea753d9cf1304c31786bdc576f7ac7f358679cbcac558b789d1d742622728d9cc23486de58b7185832e5e768c939032a24a7b8a94cca
|
|
7
|
+
data.tar.gz: ed92be4e7fe8d8d026fee4aaff03936553c4d086fe44a1eccf6bd8f890ba17c58710856e0b08a486722022d034ac432c623ccac28610c7c76f2994240e1e83c7
|
data/lib/polytypo/data/README.md
CHANGED
|
@@ -1,10 +1,12 @@
|
|
|
1
1
|
# Vendored spec subset
|
|
2
2
|
|
|
3
|
-
This directory is a manually-synced copy of
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
`
|
|
3
|
+
This directory is a manually-synced copy of a subset of `polytypo/polytypo`'s canonical `spec/`:
|
|
4
|
+
`locales/`, `fixtures/`, the whole of `rules/`, `schema/`, `VERSION` and `UNICODE`. The gem and its
|
|
5
|
+
specs read only part of that — the locale and fixture data, `rules/order.json` and
|
|
6
|
+
`rules/dashes.md`; the other twelve rule documents are carried as the normative prose for the
|
|
7
|
+
behaviour the data drives, next to the data. It is **not** the canonical spec: `validate-spec.mjs`
|
|
8
|
+
lives only in `polytypo/polytypo`, and so does the authority — a change starts there and arrives
|
|
9
|
+
here by re-copying, never the other way round.
|
|
8
10
|
|
|
9
11
|
Named `lib/polytypo/data/`, not `vendor/polytypo-spec/` like the JS/Python ports: a Ruby
|
|
10
12
|
gemspec's `files` list can include any path directly, so there is no need for a separate
|
|
@@ -14,7 +16,20 @@ one copy, shipped as part of the gem's own `lib/` payload and loaded via `File.r
|
|
|
14
16
|
`__dir__` at runtime — never a network or external-filesystem read.
|
|
15
17
|
|
|
16
18
|
Editing a file here does not change the spec; it only drifts this copy from canonical. When
|
|
17
|
-
canonical's `spec/` changes, re-copy the affected files here.
|
|
19
|
+
canonical's `spec/` changes, re-copy the affected files here.
|
|
20
|
+
|
|
21
|
+
CI checks that it was done. `script/check-vendored-spec.sh` compares every file in this
|
|
22
|
+
directory against canonical `polytypo/polytypo` at tag `spec-v` + this directory's own
|
|
23
|
+
`VERSION`, and fails on any difference. Three details: the files this repository authors
|
|
24
|
+
itself are listed in `.not-canonical` and skipped; `locales/*.json` are compared with the
|
|
25
|
+
`sources` array dropped from both sides, which is the one field a vendored copy may
|
|
26
|
+
legitimately differ in; and a file here with no canonical counterpart is a failure, so a
|
|
27
|
+
canonical rename cannot pass unnoticed. Completeness is deliberately not checked — each
|
|
28
|
+
runtime vendors its own subset, and the subsets differ.
|
|
29
|
+
|
|
30
|
+
Before that check existed, this half of the tree went stale unnoticed in four of the five
|
|
31
|
+
ports at once: the data half is proved by the test suite, and nothing at all read the prose.
|
|
32
|
+
See `polytypo/polytypo` issue #56. How this vendoring will work
|
|
18
33
|
long-term (submodule, per-ecosystem spec package, or something else) is an open decision tracked
|
|
19
34
|
in `polytypo/polytypo`'s roadmap; this is the interim, manually-synced form — the same status
|
|
20
35
|
every other port's vendored copy has.
|
data/lib/polytypo/data/VERSION
CHANGED
|
@@ -1 +1 @@
|
|
|
1
|
-
1.
|
|
1
|
+
1.6.1
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.6.1",
|
|
3
3
|
"cases": [
|
|
4
4
|
{
|
|
5
5
|
"id": "exact-en-us",
|
|
@@ -276,6 +276,18 @@
|
|
|
276
276
|
"tag": "pt-AO",
|
|
277
277
|
"resolves": "pt-PT",
|
|
278
278
|
"note": "section 3.4 step 3b: the strip lands on an alias key, so the alias is followed — the one path through step 3 that steps 1 and 2 cannot reach."
|
|
279
|
+
},
|
|
280
|
+
{
|
|
281
|
+
"id": "exact-tr",
|
|
282
|
+
"tag": "tr",
|
|
283
|
+
"resolves": "tr",
|
|
284
|
+
"note": "section 3.4 step 1. Added with the locale in spec 1.6.0, for the reason exact-es records: a registry member with no resolution case is a member nothing proves resolves rather than throws."
|
|
285
|
+
},
|
|
286
|
+
{
|
|
287
|
+
"id": "strip-tr-tr",
|
|
288
|
+
"tag": "tr-TR",
|
|
289
|
+
"resolves": "tr",
|
|
290
|
+
"note": "section 3.4 step 3a, and the tag a Turkish caller actually sends. tr needs no registry alias precisely because the region strip reaches it — this case is what makes that claim testable rather than asserted."
|
|
279
291
|
}
|
|
280
292
|
]
|
|
281
293
|
}
|
|
@@ -0,0 +1,250 @@
|
|
|
1
|
+
{
|
|
2
|
+
"spec": "1.6.1",
|
|
3
|
+
"locale": "tr",
|
|
4
|
+
"cases": [
|
|
5
|
+
{
|
|
6
|
+
"id": "tr-quotes-primary-and-nested",
|
|
7
|
+
"rule": "quotes",
|
|
8
|
+
"mode": "text",
|
|
9
|
+
"in": "Sordu: \"Bu, 'köşedeki' dedikleri dükkân mı?\"",
|
|
10
|
+
"out": "Sordu: “Bu, ‘köşedeki’ dedikleri dükkân mı?”",
|
|
11
|
+
"note": "Birinci derece U+201C/U+201D, ikinci derece U+2018/U+2019 — TDK «Tek Tırnak İşareti» kuralının iki derecesi tek bir cümlede. Kod noktaları Türk Dili dergisinin metin katmanından okundu, kural ise TDK Yazım Kuralları sayfasından; ikisi de tr.json içindeki sources alanında."
|
|
12
|
+
},
|
|
13
|
+
{
|
|
14
|
+
"id": "tr-quotes-tdk-nested-example",
|
|
15
|
+
"rule": "quotes",
|
|
16
|
+
"mode": "text",
|
|
17
|
+
"in": "\"Atatürk henüz 'Gazi Mustafa Kemal Paşa' idi.\"",
|
|
18
|
+
"out": "“Atatürk henüz ‘Gazi Mustafa Kemal Paşa’ idi.”",
|
|
19
|
+
"note": "TDK Yazım Kılavuzu'nun tek tırnak kuralı için verdiği örneğin kendisi (Falih Rıfkı Atay), düz işaretlerle yazılıp motora verildi. Çıktı, kaynağın bastığı iç içe biçimle aynıdır."
|
|
20
|
+
},
|
|
21
|
+
{
|
|
22
|
+
"id": "tr-quotes-not-guillemets",
|
|
23
|
+
"rule": "quotes",
|
|
24
|
+
"mode": "text",
|
|
25
|
+
"in": "«Tanrı, Musa'ya söz söylemiştir.»",
|
|
26
|
+
"out": "“Tanrı, Musa’ya söz söylemiştir.”",
|
|
27
|
+
"note": "Issue #51'de bildirilen «…» biçimi Türkçenin normu DEĞİLDİR: Türk Dili, «Tırnak İşareti Üzerine», s. 32 onu «bazı yazarlar»ın tercihi olarak tanıtır. quotes.md §0 mandate 1 uyarınca var olan her tırnak yeniden dizilebilir, bu yüzden yan tırnak TDK'nin üstten çift tırnağına çevrilir. Bu satır, hangi biçimin kazandığını sabitler."
|
|
28
|
+
},
|
|
29
|
+
{
|
|
30
|
+
"id": "tr-quotes-already-correct",
|
|
31
|
+
"rule": "quotes",
|
|
32
|
+
"mode": "text",
|
|
33
|
+
"in": "“Kendi Gök Kubbemiz” adı altında çıktı.",
|
|
34
|
+
"out": "“Kendi Gök Kubbemiz” adı altında çıktı.",
|
|
35
|
+
"note": "Doğru dizilmiş metin sabit noktadır: locale'in kendi çifti yeniden yazılmaz."
|
|
36
|
+
},
|
|
37
|
+
{
|
|
38
|
+
"id": "tr-apostrophe-suffix-after-abbreviation-and-numeral",
|
|
39
|
+
"rule": "apostrophe",
|
|
40
|
+
"mode": "text",
|
|
41
|
+
"in": "TBMM'nin kararı 1985'te açıklandı.",
|
|
42
|
+
"out": "TBMM’nin kararı 1985’te açıklandı.",
|
|
43
|
+
"note": "TDK «Kesme İşareti» 3 ve 4: kısaltmalara ve sayılara getirilen ekler kesme işaretiyle ayrılır. Her iki işaretin de iki yanı ALNUM olduğundan apostrophe.md §3.3 case 2 çalışır — bu satır locale verisi okumaz, yapısal olarak doğrudur, ve Türkçenin en sık biçimini sabitler."
|
|
44
|
+
},
|
|
45
|
+
{
|
|
46
|
+
"id": "tr-apostrophe-suffix-after-proper-noun",
|
|
47
|
+
"rule": "apostrophe",
|
|
48
|
+
"mode": "text",
|
|
49
|
+
"in": "Türkiye'nin ve Atatürk'ün",
|
|
50
|
+
"out": "Türkiye’nin ve Atatürk’ün",
|
|
51
|
+
"note": "Aynı kural, özel adlarda. «Atatürk'ün» parçası «ün» ile başlar; apostrophe kuralı hiçbir locale verisi okumadığı için bu, harf katlamasından bağımsız olarak çalışır."
|
|
52
|
+
},
|
|
53
|
+
{
|
|
54
|
+
"id": "tr-apostrophe-suffix-across-span-boundary",
|
|
55
|
+
"rule": "apostrophe",
|
|
56
|
+
"mode": "html",
|
|
57
|
+
"in": "<b>TBMM</b>'nin kararı",
|
|
58
|
+
"out": "<b>TBMM</b>’nin kararı",
|
|
59
|
+
"note": "Ek, satır içi bir span sınırına dayandığında: modes.md §3.2'nin işaretçisi apostrophe.md §3.1'in OPENISH kümesindedir, dolayısıyla işaret case 4'e düşer ve aynı U+2019'u verir. Burada eşleşecek ikinci bir tırnak olmadığı için quotes zaten bir şey yapamaz; çevreleyen bir alıntının olduğu biçim ayrı bir satırda sabitlenmiştir."
|
|
60
|
+
},
|
|
61
|
+
{
|
|
62
|
+
"id": "tr-ellipsis-abbreviated-after-question",
|
|
63
|
+
"rule": "ellipsis",
|
|
64
|
+
"mode": "text",
|
|
65
|
+
"in": "Nasıl da akşam oldu?...",
|
|
66
|
+
"out": "Nasıl da akşam oldu?..",
|
|
67
|
+
"note": "ellipsis.abbreviatedAfterTerminal = true. TDK «Üç Nokta» UYARI'sı soru ve ünlem işaretinden sonra iki noktayı yeterli sayar ve bütün örnekleri iki noktalıdır; §3 adım 4→6 üç noktayı o biçime normalleştirir. Değerin neden true olduğu ölçümle gerekçelendirilmiştir: false olsaydı «?..» girdisi «?…» olurdu, yani Kılavuz'un bastığı biçim bozulurdu."
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"id": "tr-ellipsis-abbreviated-is-a-fixed-point",
|
|
71
|
+
"rule": "ellipsis",
|
|
72
|
+
"mode": "text",
|
|
73
|
+
"in": "Gök ekini biçer gibi!.. Başaklar daha dolmadan.",
|
|
74
|
+
"out": "Gök ekini biçer gibi!.. Başaklar daha dolmadan.",
|
|
75
|
+
"note": "TDK'nin kendi örneği (Tarık Buğra), olduğu gibi. Doğru biçim yeniden işlemeye dayanır — abbreviatedAfterTerminal = true olan bir locale'de iki nokta çıktının kendisidir, girdi hatası değil."
|
|
76
|
+
},
|
|
77
|
+
{
|
|
78
|
+
"id": "tr-ellipsis-plain-run",
|
|
79
|
+
"rule": "ellipsis",
|
|
80
|
+
"mode": "text",
|
|
81
|
+
"in": "Bekledik ... ve gitti.",
|
|
82
|
+
"out": "Bekledik … ve gitti.",
|
|
83
|
+
"note": "Terminal işaret olmadan üç nokta tek U+2026 olur; abbreviatedAfterTerminal yalnızca soru ve ünlemden SONRA devreye girer (§3 adım 6'nın TERMINAL testi)."
|
|
84
|
+
},
|
|
85
|
+
{
|
|
86
|
+
"id": "tr-dashes-parenthetical-none-is-a-total-noop",
|
|
87
|
+
"rule": "dashes",
|
|
88
|
+
"mode": "text",
|
|
89
|
+
"in": "Küçük bir sürü -dört inekle birkaç koyun- köye girdi.",
|
|
90
|
+
"out": "Küçük bir sürü -dört inekle birkaç koyun- köye girdi.",
|
|
91
|
+
"note": "dash.parenthetical = «none»: TDK ara sözü KISA ÇİZGİ ile ve bitişik yazar, locale.schema.json'ın enum'u ise yalnızca em/en uzunluklarını tanır, dolayısıyla bu gelenek ifade edilemez ve kural sessiz bırakılmıştır. Örnek TDK'nin kendisinindir ve olduğu gibi kalır. tr, dashes kuralının kanıtlanabilir biçimde total no-op olduğu üçüncü locale'dir, el ve es'ten sonra (dashes.md §6)."
|
|
92
|
+
},
|
|
93
|
+
{
|
|
94
|
+
"id": "tr-dashes-existing-em-dash-untouched",
|
|
95
|
+
"rule": "dashes",
|
|
96
|
+
"mode": "text",
|
|
97
|
+
"in": "Plan — eğer varsa — başarısız.",
|
|
98
|
+
"out": "Plan — eğer varsa — başarısız.",
|
|
99
|
+
"note": "Aynı «none» değeri, var olan bir tam kare çizgiye de dokunmaz: kuralın hiçbir emisyonu yoktur, bu yüzden yeniden dizme de yapmaz. Bu satır, «none»'ın «kısa çizgiye çevir» anlamına GELMEDİĞİNİ sabitler."
|
|
100
|
+
},
|
|
101
|
+
{
|
|
102
|
+
"id": "tr-ranges-none-even-when-the-rule-is-enabled",
|
|
103
|
+
"rule": "ranges",
|
|
104
|
+
"mode": "text",
|
|
105
|
+
"rules": {
|
|
106
|
+
"ranges": true
|
|
107
|
+
},
|
|
108
|
+
"in": "1914-1918 Birinci Dünya Savaşı",
|
|
109
|
+
"out": "1914-1918 Birinci Dünya Savaşı",
|
|
110
|
+
"note": "ranges.md §2: dash.range = «none» olduğunda kural hiçbir şey yayamaz. Girdi TDK «Kısa Çizgi» 7'nin kendi örneğidir — aralık normludur ama işaret kısa çizgidir, yani doğrulanmış bir en/em geleneği yoktur. Kural varsayılan olarak kapalı olduğundan bu satır onu açıkça açar; açıkken de çıktı değişmez, ki asıl sabitlenen budur."
|
|
111
|
+
},
|
|
112
|
+
{
|
|
113
|
+
"id": "tr-nbsp-before-units",
|
|
114
|
+
"rule": "nbsp",
|
|
115
|
+
"mode": "text",
|
|
116
|
+
"in": "15 °C, 20 kg ve 5 cm² ölçüldü.",
|
|
117
|
+
"out": "15 °C, 20 kg ve 5 cm² ölçüldü.",
|
|
118
|
+
"note": "N5, TDK SSS'nin «15 °C» ve «20 kg» örnekleriyle. «cm²» girişinin «cm» ile birlikte listelenmesi güvenlidir: aynı konumda en uzun eşleşme kazanır (nbsp.md §3.7)."
|
|
119
|
+
},
|
|
120
|
+
{
|
|
121
|
+
"id": "tr-nbsp-ton-and-mm",
|
|
122
|
+
"rule": "nbsp",
|
|
123
|
+
"mode": "text",
|
|
124
|
+
"in": "350 ton yük ve 12 mm kalınlık.",
|
|
125
|
+
"out": "350 ton yük ve 12 mm kalınlık.",
|
|
126
|
+
"note": "«ton» uluslararası bir simge değil bir kelimedir ve yalnızca TDK SSS'nin düzyazı örneğinden gelir; listede tutulmasının nedeni sağ sınır testinin onu «tonluk» içinde eşleştirmemesidir. Kaynak rütbesi tr.json'da açıkça yazılıdır."
|
|
127
|
+
},
|
|
128
|
+
{
|
|
129
|
+
"id": "tr-nbsp-percent-is-not-bound",
|
|
130
|
+
"rule": "nbsp",
|
|
131
|
+
"mode": "text",
|
|
132
|
+
"in": "%25 ve ‰50 oranları.",
|
|
133
|
+
"out": "%25 ve ‰50 oranları.",
|
|
134
|
+
"note": "TDK «Sayıların Yazılışı»: yüzde ve binde işaretleri sayıdan ÖNCE ve bitişik yazılır. beforeUnits ve afterSymbols'in ikisi de bu işaret için boştur ve N5/N6 yalnızca var olan bir boşluğu dönüştürür, asla eklemez — bu yüzden biçim olduğu gibi kalır. pl.json'un «%» girişinin tersi, bilerek."
|
|
135
|
+
},
|
|
136
|
+
{
|
|
137
|
+
"id": "tr-nbsp-abbreviation-internal-space",
|
|
138
|
+
"rule": "nbsp",
|
|
139
|
+
"mode": "text",
|
|
140
|
+
"in": "Kur. Bşk. ve Nö. Sb. geldi.",
|
|
141
|
+
"out": "Kur. Bşk. ve Nö. Sb. geldi.",
|
|
142
|
+
"note": "N4, TDK Kısaltmalar Dizini'nden harfi harfine alınan iki giriş. Ölçüt: her iki parçası da noktalı kısaltma olan girişler; akronim içeren biçimler dışarıda bırakıldı."
|
|
143
|
+
},
|
|
144
|
+
{
|
|
145
|
+
"id": "tr-nbsp-short-word-inert",
|
|
146
|
+
"rule": "nbsp",
|
|
147
|
+
"mode": "text",
|
|
148
|
+
"in": "ve bir de o geldi",
|
|
149
|
+
"out": "ve bir de o geldi",
|
|
150
|
+
"note": "afterShortWords boş, ve bu satır onu kanıtlanabilir yapar: Türkçede tek harfli bağlaç veya edat yoktur ve TDK asılı kalan kısa sözcükler için bir kural vermez. ru ve pl'nin listeleri buraya taşınmamıştır."
|
|
151
|
+
},
|
|
152
|
+
{
|
|
153
|
+
"id": "tr-hyphen-noop",
|
|
154
|
+
"rule": "hyphen",
|
|
155
|
+
"mode": "text",
|
|
156
|
+
"in": "Ural-Altay dil grubu ve Türk-Alman ilişkileri",
|
|
157
|
+
"out": "Ural-Altay dil grubu ve Türk-Alman ilişkileri",
|
|
158
|
+
"note": "Üç liste de boş: TDK satır sonunda bölmeyi normlar, yasaklamaz, dolayısıyla hiçbir biçim U+2011 talep etmez. Girdi TDK «Kısa Çizgi» 7'nin kendi örneğidir ve kural onun için kanıtlanabilir bir no-op'tur (hyphen.md §2)."
|
|
159
|
+
},
|
|
160
|
+
{
|
|
161
|
+
"id": "tr-symbols-and-spaces",
|
|
162
|
+
"rule": "symbols",
|
|
163
|
+
"mode": "text",
|
|
164
|
+
"in": "Copyright (c) 2026, baskı 40x60 cm.",
|
|
165
|
+
"out": "Copyright © 2026, baskı 40×60 cm.",
|
|
166
|
+
"note": "Locale'den bağımsız kurallar Türkçe metinde de çalışır: (c) → U+00A9, x → U+00D7, ve «60 cm» beforeUnits üzerinden bağlanır."
|
|
167
|
+
},
|
|
168
|
+
{
|
|
169
|
+
"id": "tr-spaces-punctuation-is-closed-up",
|
|
170
|
+
"rule": "spaces",
|
|
171
|
+
"mode": "text",
|
|
172
|
+
"in": "Nasıl gidiyorsun ? İyiyim , sağ ol .",
|
|
173
|
+
"out": "Nasıl gidiyorsun? İyiyim, sağ ol.",
|
|
174
|
+
"note": "TDK «Noktalama İşaretleri (Açıklamalar)» giriş hükmü: işaretler ait oldukları kelimelere bitişik yazılır ve boşluk işaretten SONRA gelir. spaces kuralı bunu tam olarak uygular — işaretten önceki boşluk silinir, ikili boşluk teke iner — ve nbsp.beforePunctuation'ın bu locale'de boş olmasının nedeni de aynı hükümdür: Fransızca «mot !» geleneği Türkçede yoktur, dolayısıyla silinen boşluğun geri konması istenmez."
|
|
175
|
+
},
|
|
176
|
+
{
|
|
177
|
+
"id": "tr-html-span-boundary-suffix-declined",
|
|
178
|
+
"rule": "quotes",
|
|
179
|
+
"mode": "html",
|
|
180
|
+
"in": "Dedi: '<b>TBMM</b>'nin kararı doğru.' Bitti.",
|
|
181
|
+
"out": "Dedi: “<b>TBMM</b>’nin kararı doğru.” Bitti.",
|
|
182
|
+
"note": "quotes.md §3.2'nin span-boundary elizyon vetosu (spec 1.4.0), Türkçe için canlı: ek kesme işaretiyle bir satır içi span sınırına dayanır, modes.md §3.2'nin işaretçisi «M» harfinin yerini tutar ve medial-elizyon vetosu çalışamaz. quotes.elisionClitics.after'daki «nin» girişi işareti reddeder, apostrophe (order 50) onu case 4 ile U+2019 yapar, ve yazarın kendi çifti birincil glifleri korur. Liste boş olsaydı — spec 1.6.0'dan önce her locale için olduğu gibi — ek alıntıyı açar, açılış işareti düz U+0027 kalır ve kapanış işareti kapanış glifi olurdu; ölçülmüştür (canonical issue #53). Türkçede bu biçim nadir değil, dilbilgisel olarak zorunludur."
|
|
183
|
+
},
|
|
184
|
+
{
|
|
185
|
+
"id": "tr-markdown-commonmark-span-boundary-suffix-declined",
|
|
186
|
+
"rule": "quotes",
|
|
187
|
+
"mode": "markdown",
|
|
188
|
+
"dialect": "commonmark",
|
|
189
|
+
"in": "Dedi: '**Ankara**'da olacak.' Bitti.\n",
|
|
190
|
+
"out": "Dedi: “**Ankara**’da olacak.” Bitti.\n",
|
|
191
|
+
"note": "Aynı mekanizma markdown'da ve «da» girişiyle: veto işaretçiyi okur, onu üreten span türünü değil. Yer adlarına gelen bulunma eki, Türkçe metinde en sık rastlanan biçimlerden biridir."
|
|
192
|
+
},
|
|
193
|
+
{
|
|
194
|
+
"id": "tr-html-span-boundary-listed-fragment-quoted",
|
|
195
|
+
"rule": "quotes",
|
|
196
|
+
"mode": "html",
|
|
197
|
+
"in": "Ek <em>'de'</em> biçiminde yazılır.",
|
|
198
|
+
"out": "Ek <em>’de’</em> biçiminde yazılır.",
|
|
199
|
+
"note": "Kabul edilen yanlış pozitif, birinin metninde bırakılmak yerine suite'e yazılmıştır: içeriği TAMAMEN listelenmiş bir parça olan ve satır içi bir span sınırına dayanan bir alıntı, ek sanılır. Bu, spec 1.4.0'ın İngilizce «s» için kabul ettiği bedelin aynısıdır ve Türkçede riski daha yüksektir, çünkü «de» ayrıca ayrı yazılan bir bağlaçtır. İki hafifletici tr.json'da kayıtlıdır: TDK'nin kendi anma biçimi başta kısa çizgi taşır ve kısa çizgi LETTER olmadığı için veto çalışamaz, ayrıca TDK bağlacı tırnaksız yazar."
|
|
200
|
+
},
|
|
201
|
+
{
|
|
202
|
+
"id": "tr-html-span-boundary-uppercase-suffix-not-matched",
|
|
203
|
+
"rule": "quotes",
|
|
204
|
+
"mode": "html",
|
|
205
|
+
"in": "Dedi: '<b>TBMM</b>'NİN kararı.' Bitti.",
|
|
206
|
+
"out": "Dedi: '<b>TBMM</b>“NİN kararı.” Bitti.",
|
|
207
|
+
"note": "Kapsam sınırı, adıyla sabitlenmiştir: büyük harfli ek eşleşmez, çünkü katlama yalnızca ilk kod noktasını ve yalnızca ASCII A-Z'yi kapsar — «N»→«n» katlanır ama kuyrukta U+0130 «İ» ile U+0069 «i» ayrı kod noktalarıdır. Katlamayı genişletmek ARCHITECTURE.md §4.4'ün noktasız ı (U+0131) yüzünden açıkça yasakladığı yerdir, dolayısıyla bu kayıp giderilmez, kabul edilir ve görünür tutulur."
|
|
208
|
+
},
|
|
209
|
+
{
|
|
210
|
+
"id": "tr-html-span-boundary-unlisted-suffix-not-matched",
|
|
211
|
+
"rule": "quotes",
|
|
212
|
+
"mode": "html",
|
|
213
|
+
"in": "Dedi: '<b>Irak</b>'ta olacak.' Bitti.",
|
|
214
|
+
"out": "Dedi: '<b>Irak</b>“ta olacak.” Bitti.",
|
|
215
|
+
"note": "İkinci kapsam sınırı: «ta» parçası, TDK'nin okunan sayfalarında kesme işaretinden sonraki dizinin TAMAMI olarak basılmadığı için listeye girmemiştir — ölçüt paradigma değil, basılı örnektir (tr.json). Liste yalnızca reddeder, dolayısıyla eksik bir giriş yanlış bir giriş değildir; bu satır hangi biçimlerin henüz onarılmadığını görünür kılar."
|
|
216
|
+
},
|
|
217
|
+
{
|
|
218
|
+
"id": "tr-html-span-boundary-quotation-outside-the-span",
|
|
219
|
+
"rule": "quotes",
|
|
220
|
+
"mode": "html",
|
|
221
|
+
"in": "Bağlaç olan '<em>de</em>' ayrı yazılır.",
|
|
222
|
+
"out": "Bağlaç olan “<em>de</em>” ayrı yazılır.",
|
|
223
|
+
"note": "Yanlış pozitifin negatif kontrolü, ve maliyetin NE KADAR DAR olduğunu gösterir. Alıntı işaretleri span'in DIŞINDA olduğunda hiçbir işaret span sınırına dayanmaz, veto hiç çalışmaz ve cümle doğru dizilir. Terimi vurgulayıp alıntılamanın olağan yazılışı budur; veto ancak yazar işaretleri span'in İÇİNE koyduğunda yanılır."
|
|
224
|
+
},
|
|
225
|
+
{
|
|
226
|
+
"id": "tr-html-span-boundary-tdk-hyphen-citation-form",
|
|
227
|
+
"rule": "quotes",
|
|
228
|
+
"mode": "html",
|
|
229
|
+
"in": "<em>'-de'</em> eki bitişik yazılır.",
|
|
230
|
+
"out": "<em>“-de”</em> eki bitişik yazılır.",
|
|
231
|
+
"note": "TDK'nin ekleri anma biçiminin kendisi, baştaki kısa çizgiyle («-da / -de / -ta / -te»). Kısa çizgi LETTER olmadığı için işaretin sağındaki dizi boş kalır, hiçbir giriş eşleşemez ve alıntı bozulmadan dizilir. Yanlış pozitifin iki ölçülen hafifleticisinden biri, prozada iddia olarak değil suite'te satır olarak."
|
|
232
|
+
},
|
|
233
|
+
{
|
|
234
|
+
"id": "tr-html-span-boundary-non-ascii-initial-fragment",
|
|
235
|
+
"rule": "quotes",
|
|
236
|
+
"mode": "html",
|
|
237
|
+
"in": "Dedi: '<b>Atatürk</b>'üm dedi.' Bitti.",
|
|
238
|
+
"out": "Dedi: “<b>Atatürk</b>’üm dedi.” Bitti.",
|
|
239
|
+
"note": "Baş harfi ASCII olmayan bir parça — «üm», U+00FC ile başlar. Eşleşir, çünkü kesme işaretinden sonraki dizi zaten küçük harflidir ve TAM karşılaştırılır; katlama yalnızca büyük harfli bir dizi için gerekir ve yalnızca ASCII A-Z'yi kapsar. Bu satır, ARCHITECTURE.md §4.4'ün yasağının bu alanı KISITLAMADIĞINI sabitler: küçük harfli Türkçe parçalar hiçbir katlamaya ihtiyaç duymaz, kayıp yalnızca büyük harfte gerçekleşir (ayrı satır)."
|
|
240
|
+
},
|
|
241
|
+
{
|
|
242
|
+
"id": "tr-html-span-boundary-ordinal-fragment",
|
|
243
|
+
"rule": "quotes",
|
|
244
|
+
"mode": "html",
|
|
245
|
+
"in": "Dedi: '<b>2</b>'nci kat.' Bitti.",
|
|
246
|
+
"out": "Dedi: “<b>2</b>’nci kat.” Bitti.",
|
|
247
|
+
"note": "Sayılara gelen sıra eki, TDK «Kesme İşareti» 4'ün kendi örneğinden («2'nci kat»). Aynı maddeden gelen «7,65'lik» ve «Atatürk'üm» de listededir; üçü de basılı örnekte kesme işaretinden sonraki dizinin TAMAMI olduğu ve çıplak bir sözcükle çakışmadığı için alınmıştır."
|
|
248
|
+
}
|
|
249
|
+
]
|
|
250
|
+
}
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.6.1",
|
|
3
3
|
"$comment": "Locale resolution input. Algorithm is specified in spec/rules/locale-resolution.md and is identical in every runtime — never delegate it to a platform locale-negotiation library. \"spec\" here must track spec/VERSION exactly — it is not itself the global version source; scripts/validate-spec.mjs enforces the match.",
|
|
4
4
|
"locales": [
|
|
5
5
|
"en-US",
|
|
@@ -19,7 +19,8 @@
|
|
|
19
19
|
"nl",
|
|
20
20
|
"pl",
|
|
21
21
|
"uk",
|
|
22
|
-
"cs"
|
|
22
|
+
"cs",
|
|
23
|
+
"tr"
|
|
23
24
|
],
|
|
24
25
|
"aliases": {
|
|
25
26
|
"en": "en-US",
|
|
@@ -0,0 +1,72 @@
|
|
|
1
|
+
{
|
|
2
|
+
"locale": "tr",
|
|
3
|
+
"name": "Turkish",
|
|
4
|
+
"quotes": {
|
|
5
|
+
"primary": {
|
|
6
|
+
"open": "“",
|
|
7
|
+
"close": "”",
|
|
8
|
+
"innerSpace": "none"
|
|
9
|
+
},
|
|
10
|
+
"secondary": {
|
|
11
|
+
"open": "‘",
|
|
12
|
+
"close": "’",
|
|
13
|
+
"innerSpace": "none"
|
|
14
|
+
},
|
|
15
|
+
"elisionIdioms": [],
|
|
16
|
+
"elisionClitics": {
|
|
17
|
+
"before": [],
|
|
18
|
+
"after": [
|
|
19
|
+
"nin",
|
|
20
|
+
"nın",
|
|
21
|
+
"de",
|
|
22
|
+
"da",
|
|
23
|
+
"te",
|
|
24
|
+
"ye",
|
|
25
|
+
"yle",
|
|
26
|
+
"nı",
|
|
27
|
+
"dan",
|
|
28
|
+
"den",
|
|
29
|
+
"lik",
|
|
30
|
+
"nci",
|
|
31
|
+
"üm"
|
|
32
|
+
]
|
|
33
|
+
}
|
|
34
|
+
},
|
|
35
|
+
"dash": {
|
|
36
|
+
"parenthetical": "none",
|
|
37
|
+
"range": "none"
|
|
38
|
+
},
|
|
39
|
+
"ellipsis": {
|
|
40
|
+
"abbreviatedAfterTerminal": true
|
|
41
|
+
},
|
|
42
|
+
"hyphen": {
|
|
43
|
+
"prefixes": [],
|
|
44
|
+
"suffixes": [],
|
|
45
|
+
"compounds": []
|
|
46
|
+
},
|
|
47
|
+
"nbsp": {
|
|
48
|
+
"beforePunctuation": [],
|
|
49
|
+
"narrowBeforePunctuation": [],
|
|
50
|
+
"afterShortWords": [],
|
|
51
|
+
"abbreviations": [
|
|
52
|
+
"Kur. Bşk.",
|
|
53
|
+
"Nö. Sb."
|
|
54
|
+
],
|
|
55
|
+
"beforeUnits": [
|
|
56
|
+
"mm",
|
|
57
|
+
"cm",
|
|
58
|
+
"km",
|
|
59
|
+
"kg",
|
|
60
|
+
"mg",
|
|
61
|
+
"hl",
|
|
62
|
+
"m²",
|
|
63
|
+
"cm²",
|
|
64
|
+
"°C",
|
|
65
|
+
"ton"
|
|
66
|
+
],
|
|
67
|
+
"beforeNumber": [],
|
|
68
|
+
"beforeWord": [],
|
|
69
|
+
"afterSymbols": [],
|
|
70
|
+
"initialBinding": "none"
|
|
71
|
+
}
|
|
72
|
+
}
|
|
@@ -571,12 +571,37 @@ layouts and in text pasted from older systems. Converting them is not authorised
|
|
|
571
571
|
gives `A ‘quoted’’s meaning`: the closing quotation mark and the possessive end up flush, and
|
|
572
572
|
at text size the pair reads as one double quote.
|
|
573
573
|
|
|
574
|
-
|
|
575
|
-
|
|
576
|
-
|
|
577
|
-
|
|
578
|
-
|
|
579
|
-
|
|
574
|
+
The mark case 2a adds is always U+2019; the mark it abuts is whichever `CLOSEDELIM` member
|
|
575
|
+
the locale closed with, and three of that class's four quotation members are reachable as a
|
|
576
|
+
locale's primary close, so this is **not an `en-GB` phenomenon**. Grouped by `quotes.primary.close`, seventeen of the
|
|
577
|
+
nineteen locales reach it through case 2a: `’’s` in `en-GB`; `”’s` in `en-US`, `fi`, `sv`,
|
|
578
|
+
`nl`, `pl`, `pt-BR` and `tr`; `»’s` in `de-CH`, `fr`, `fr-CA`, `ru`, `el`, `es`, `it`,
|
|
579
|
+
`pt-PT` and `uk`. Only the first two are plausibly confusable, and **which pairs those are is
|
|
580
|
+
a property of the glyph pair, not of the language** — so no existing locale field expresses
|
|
581
|
+
the distinction a separator rule would have to make.
|
|
582
|
+
|
|
583
|
+
**`de-DE` and `cs` look like this and are not it.** They close with U+201C, which §3.1 puts
|
|
584
|
+
in `OPENISH` and deliberately *not* in `CLOSEDELIM`, so `A „quoted“'s meaning` reaches its
|
|
585
|
+
U+2019 through **case 4** (leading elision: `OPENISH` left, `ALNUM` right), which is 0.4.1
|
|
586
|
+
vintage — the same U+201C-is-not-closeish asymmetry case 3a records for its own reason.
|
|
587
|
+
Measured on the published packages: 1.4.0 and 1.6.0 both emit `A „quoted“’s meaning`, while
|
|
588
|
+
`en-US` emits `A “quoted”'s meaning` on 1.4.0 and `A “quoted”’s meaning` on 1.6.0. The
|
|
589
|
+
`“’s` shape predates case 2a and is unchanged by it, so it is not this item's accepted cost.
|
|
590
|
+
|
|
591
|
+
**What the authority says, and what it does not.** The remedy and the character are both
|
|
592
|
+
named, but by a *CMOS Shop Talk* post — "When Quotation Marks and Apostrophes Collide",
|
|
593
|
+
published 2020-01-14, updated 2025-12-16 — and what that post describes
|
|
594
|
+
is the typesetting *CMOS* Online applies **to its own website**, not a rule stated for English
|
|
595
|
+
text. It enumerates the separator by code point (U+00A0, thin space U+2009, hair space U+200A
|
|
596
|
+
in print, U+202F online) and cross-references 18th ed. §6.11. **That cross-reference does not
|
|
597
|
+
survive a check against §6.11's own text, which this repository already holds.**
|
|
598
|
+
`spec/locales/en-US.json` quotes it verbatim: “When single quotation marks are nested within
|
|
599
|
+
double quotation marks, and two of the marks appear next to each other, a space between the
|
|
600
|
+
two marks, though not strictly required, aids legibility.” That is about *nested quotation
|
|
601
|
+
marks*, not an apostrophe, and it is permissive (“not strictly required”) rather than
|
|
602
|
+
prescriptive. The 18th edition's index files apostrophe adjacency somewhere else entirely,
|
|
603
|
+
at `apostrophes: other punctuation with, 6.126`; §6.126 and §6.128 are behind the
|
|
604
|
+
subscription and unread.
|
|
580
605
|
|
|
581
606
|
**polytypo emits the right character and inserts nothing.** This rule *cannot* insert: §1 and
|
|
582
607
|
§4 make every edit one code point replacing one code point at the same index, and that is
|
|
@@ -585,11 +610,28 @@ layouts and in text pasted from older systems. Converting them is not authorised
|
|
|
585
610
|
between them is unaddressed by every rule in `order.json` — recorded here as a decision
|
|
586
611
|
rather than left as an omission, in the standing of items 2 and 3.
|
|
587
612
|
|
|
588
|
-
**
|
|
589
|
-
|
|
590
|
-
apostrophe
|
|
591
|
-
|
|
592
|
-
|
|
593
|
-
|
|
594
|
-
|
|
595
|
-
|
|
613
|
+
**No separator is inserted, and that is now a decision rather than an open question
|
|
614
|
+
(2026-09-24).** No normative source addresses the ordering case 2a produces — closing mark,
|
|
615
|
+
then apostrophe, then `s`. What exists is one mark-order away from it and is not normative
|
|
616
|
+
anyway: the attested passage separates a title's *own* trailing apostrophe from a *following*
|
|
617
|
+
closing quotation mark (the song title *Ain't Misbehavin'* in single quotation marks —
|
|
618
|
+
apostrophe first), and it describes a publisher's rendering practice for its own website. The
|
|
619
|
+
paragraph that would govern the possessive ordering, `CMOS` §7.29 ("Possessive with
|
|
620
|
+
italicized or quoted terms"), is behind a subscription. `MHRA` was read in full and is silent;
|
|
621
|
+
Kotus names no code points at all; the Swedish genitive takes no apostrophe, so the shape is
|
|
622
|
+
not Swedish. **Operator decision: where no exact answer exists, choose one behaviour and pin
|
|
623
|
+
it rather than leave the rule undefined — consistency across the five runtimes is the
|
|
624
|
+
product, not conformance to a style guide nobody can read.** The pinned choice is *no
|
|
625
|
+
separator*: the marks sit flush.
|
|
626
|
+
|
|
627
|
+
**What makes that the cheap side of the choice.** The pair a reader could actually misread as
|
|
628
|
+
one mark is `’’`, and it is **unreachable from ordinary typing** — exhaustive sweep over the
|
|
629
|
+
alphabet `'` `"` `a` `s` space, 370,975 inputs to length 6 across all nineteen locales and
|
|
630
|
+
1,464,825 to length 8 across `en-GB`, `en-US` and `fi`: zero produce it. Two straight marks
|
|
631
|
+
are vetoed by `quotes` V1, and a straight `"…"` pair in a `NARROW`-primary locale is declined
|
|
632
|
+
by the certification gate for the same adjacency, so `’’` needs a closing mark that was
|
|
633
|
+
*already* U+2019 in the source. The pairs reachable from real input — `”’s`, `“’s`, `»’s`
|
|
634
|
+
— are legible as two marks. Inserting a space would mean a new `nbsp` insertion *position*
|
|
635
|
+
class, which `nbsp.md` §5's `I₆` and `quotes.md` §5's Lemma B would both have to be
|
|
636
|
+
re-derived for, in exchange for a cosmetic gain on text someone has already typeset. Recorded
|
|
637
|
+
as the standing answer; do not reopen it on a citation hunt.
|
|
@@ -900,7 +900,10 @@ never reaches a digit-flanked token at all.
|
|
|
900
900
|
### `el` — `parenthetical: "none"` (`range` is now [ranges.md](ranges.md)'s field, also `"none"`, not this rule's)
|
|
901
901
|
|
|
902
902
|
The first locale with `"none"` on **both** fields, so the rule is a **total no-op** for it in the
|
|
903
|
-
same provable sense `hyphen` is a no-op for a locale with empty lists.
|
|
903
|
+
same provable sense `hyphen` is a no-op for a locale with empty lists. It is no longer the only
|
|
904
|
+
one: `es` (spec 1.3.0) and `tr` (spec 1.6.0) are both-`none` too, each for its own reason and each
|
|
905
|
+
recorded in its own locale file. This section stays written about `el` because its two fields are
|
|
906
|
+
`"none"` for two *different* reasons, which is what makes it worth reading. Tokens are still
|
|
904
907
|
classified — §3.3's "a range token is never reconsidered as a parenthetical" still holds — but no
|
|
905
908
|
classification has an emission to make. The rows below are therefore all "no change" rows by
|
|
906
909
|
construction, and they are worth pinning precisely because nothing else in the suite exercises
|
|
@@ -193,7 +193,7 @@ that "no tag was supplied at all" is expressible as a fixture rather than only a
|
|
|
193
193
|
raise the coded error rather than a native `TypeError` — which matters because a `TypeError` is
|
|
194
194
|
not in the taxonomy (ARCHITECTURE.md §4.6) and would differ in each of the five runtimes.
|
|
195
195
|
|
|
196
|
-
Every row is a fixture, not a candidate: **
|
|
196
|
+
Every row is a fixture, not a candidate: **50 resolution cases run today**, covering these rows
|
|
197
197
|
and the malformed-tag rejections. `fixtures.schema.json` was extended with a case shape carrying
|
|
198
198
|
no `mode` and no `out`, whose expected result is a locale id or a thrown code. An earlier
|
|
199
199
|
revision of this paragraph said such rows "cannot be expressed in the existing fixture format";
|
|
@@ -210,11 +210,11 @@ that was true when written, and is why the schema was changed.
|
|
|
210
210
|
language polytypo does not support yet). I have mapped both to
|
|
211
211
|
`POLYTYPO_UNKNOWN_LOCALE` because inventing a code would change a documented contract.
|
|
212
212
|
Recommend adding `POLYTYPO_INVALID_LOCALE`; operator decision.
|
|
213
|
-
2. *(Closed.)* Resolution **is** fixture-covered:
|
|
213
|
+
2. *(Closed.)* Resolution **is** fixture-covered: 50 resolution cases run today. The gap this
|
|
214
214
|
item reported — that `fixtures.schema.json` could not express a case with no `mode`, no `out`
|
|
215
215
|
and an expected locale id or thrown code — was closed by extending the schema, and the §5
|
|
216
216
|
table's rows are those cases. (The item also miscounted the pipeline as seven rules; it is
|
|
217
|
-
|
|
217
|
+
nine — `hyphen` was added at order 35 and `ranges` at 37.)
|
|
218
218
|
|
|
219
219
|
3. **`sv-FI` (Finland Swedish) silently resolves to `sv`.** The two genuinely differ in some
|
|
220
220
|
conventions, and PLAN.md §7 already flags Swedish quote practice as uncertain. The
|
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"spec": "1.
|
|
2
|
+
"spec": "1.6.1",
|
|
3
3
|
"$comment": "Single source of truth for pipeline order. Rules run in ascending `order`. Disabling a rule removes it from the sequence and never reorders the rest. Rule ids are public API (see docs/ARCHITECTURE.md section 5). \"spec\" here must track spec/VERSION exactly — it is not itself the global version source; scripts/validate-spec.mjs enforces the match.",
|
|
4
4
|
"rules": [
|
|
5
5
|
{
|
|
@@ -595,14 +595,39 @@ typesets as `A ’stray mark here.` / `Then ‘inner’ text.`, and appending `S
|
|
|
595
595
|
re-pairs the first mark with the new one, making the middle pair nested and changing a line the
|
|
596
596
|
edit did not touch. Measured on 1.3.1 and unchanged by this veto.
|
|
597
597
|
|
|
598
|
-
**This is issue #53 §3,
|
|
598
|
+
**This is issue #53 §3, and as of 2026-09-24 it is decided rather than open: the behaviour
|
|
599
|
+
stands.** The report's own trigger lines are
|
|
599
600
|
`` `quotes`' locale data `` and `` `dashes`' own admissibility `` — plural possessives, exactly
|
|
600
601
|
this shape — so the cross-paragraph effect it documents in a real article is *not* repaired by
|
|
601
602
|
spec 1.4.0. Measured on this change, `markdown`/`commonmark`, `en-GB`:
|
|
602
603
|
`It is 'the `` `quotes`' `` locale data' here.` gives
|
|
603
604
|
`It is ‘the `` `quotes`’ `` locale data’ here.` — the plural possessive closes the quotation two
|
|
604
605
|
words early and the author's own closing mark is left to `apostrophe` as a stray U+2019. §1 of
|
|
605
|
-
that issue is closed
|
|
606
|
+
that issue is closed by the veto above; §3 is not, and is not going to be.
|
|
607
|
+
|
|
608
|
+
**Why §3 is decided and not merely unfixed.** Two measurements, both on the published 1.6.0
|
|
609
|
+
package. First, **the glyph at the possessive's own position is correct either way** — U+2019,
|
|
610
|
+
whether the mark is taken as a closing quotation mark or left unmatched for `apostrophe`. What
|
|
611
|
+
the defect costs is never the character a reader sees there; it is the *pairing it consumes*,
|
|
612
|
+
which is why the visible symptom always appears somewhere else (a stray U+0027 at the real
|
|
613
|
+
closing mark, or a middle paragraph re-nested from `‘ ’` to `“ ”`). Second, **separating the two
|
|
614
|
+
readings was implemented and measured, and it is worse.** The narrowest local veto that can do
|
|
615
|
+
it — decline a `NARROW` mark whose literal neighbour is `MARKER` and whose attaching `LETTER`
|
|
616
|
+
run is empty — breaks **7 of 2801 conformance cases**, five of them `tr` cases that are correct
|
|
617
|
+
today: `tr-html-span-boundary-suffix-declined`,
|
|
618
|
+
`tr-markdown-commonmark-span-boundary-suffix-declined`,
|
|
619
|
+
`tr-html-span-boundary-quotation-outside-the-span`,
|
|
620
|
+
`tr-html-span-boundary-non-ascii-initial-fragment` and `tr-html-span-boundary-ordinal-fragment`.
|
|
621
|
+
That is this section's own claim — no test over these neighbours can separate a plural
|
|
622
|
+
possessive from a closing mark after a span — turned from an argument into a number.
|
|
623
|
+
|
|
624
|
+
**Operator decision (2026-09-24): where no exact answer exists, choose one behaviour and pin it
|
|
625
|
+
rather than leave the rule undefined.** Five runtimes agreeing byte-for-byte is the product; an
|
|
626
|
+
undecided rule is worse than an arbitrary decided one. The pinned behaviour is the one specified
|
|
627
|
+
here and fixtured at §6 row S5. A future change would have to be a **pairing-preference**
|
|
628
|
+
mechanism, not a classification test, and §3.3's three-outcome partition is the premise §5's
|
|
629
|
+
certification theorem rests on — so it replaces that premise rather than extending it, and needs
|
|
630
|
+
its own idempotency argument. Nobody should spend that on this without a new reason.
|
|
606
631
|
|
|
607
632
|
**V1 — same-V1-identity adjacency veto** (both widths):
|
|
608
633
|
|
|
@@ -1330,7 +1355,7 @@ what a port should be able to re-derive from §3.2 alone.
|
|
|
1330
1355
|
| S2 | `fr` | `Il dit 'l'<em>idée</em> est bonne.' Fin.` | `Il dit «⍽l’<em>idée</em> est bonne.⍽» Fin.` | the mirror direction: `Rlit` is the `MARKER`, the run left of the mark is `l`, a cited `before` entry. Through 1.3.1 this gave `Il dit «⍽l⍽»<em>idée</em> est bonne.' Fin.` — guillemets around one letter, the real closing mark abandoned |
|
|
1331
1356
|
| S3 | `fr` | `Il dit 'jusqu'<em>ici</em> tout va bien.' Fin.` | `Il dit «⍽jusqu’<em>ici</em> tout va bien.⍽» Fin.` | the run is compared **whole**: it is `jusqu`, so an entry `qu` could not match it. This is why `jusqu`, `lorsqu`, `puisqu` and `quoiqu` are listed in their own right |
|
|
1332
1357
|
| S4 | `fr` | `Il dit <em>'oui'</em> ici.` | `Il dit <em>«⍽oui⍽»</em> ici.` | the negative control the mechanism exists to preserve. `Llit` is the `MARKER` here too — what separates this from S2 is only that `oui` is not a listed fragment |
|
|
1333
|
-
| S5 | `en-GB` | `` He says 'avoid `xs`' printer.' Done. `` | `` He says ‘avoid `xs`’ printer.' Done. `` | **not closed.** The plural possessive's run is empty, so neither test applies, and these code points are also a closing mark after a span. The second mark takes the pairing, the third is abandoned as U+0027. §4, §7 item 10, issue #54 |
|
|
1358
|
+
| S5 | `en-GB` | `` He says 'avoid `xs`' printer.' Done. `` | `` He says ‘avoid `xs`’ printer.' Done. `` | **not closed.** The plural possessive's run is empty, so neither test applies, and these code points are also a closing mark after a span. The second mark takes the pairing, the third is abandoned as U+0027. **Decided behaviour, not a pending fix** — §3.2. §4, §7 item 10, issue #54 |
|
|
1334
1359
|
| S6 | `fr` | `Il dit 'QU'<em>il</em> vienne.' Fin.` | `Il dit «⍽QU⍽»<em>il</em> vienne.' Fin.` | **not closed.** The fold reaches only the run's first code point, and only ASCII `A`–`Z`, so `QU` does not match `qu` and the inversion survives. Same limit as the listed veto's context words |
|
|
1335
1360
|
| S7 | `en-US` | `He said <em>'s'</em> loudly.` | `He said <em>’s’</em> loudly.` | the accepted false positive: a quotation inside a span whose whole content is a listed fragment. Narrower than the medial-`n` veto's own accepted `The letter 'n' is common.`, since it needs the boundary as well |
|
|
1336
1361
|
| S8 | `nl` | `Hij zegt 'ik kom <em>vroeg </em>'s avonds terug.' Klaar.` | `Hij zegt “ik kom <em>vroeg </em>’s avonds terug.” Klaar.` | `'s` is a word-*initial* omission, so its run lies to the mark's right and the entry is an `after` one — the clearest case for reading both lists positionally (§2) |
|
|
@@ -1480,7 +1505,8 @@ Ordered by how much this matters.
|
|
|
1480
1505
|
rule. Two shapes stay open in **every** locale, both for the same reason: the attaching
|
|
1481
1506
|
fragment is not there to read. A plural possessive after a span (`` `xs`' ``) has an empty
|
|
1482
1507
|
run and is byte-identical to a closing mark after a span (§4) — **this is issue #53 §3, the
|
|
1483
|
-
cross-paragraph witness that motivated the report
|
|
1508
|
+
cross-paragraph witness that motivated the report, and it is decided rather than open — §3.2
|
|
1509
|
+
gives the two measurements and issue #54 records the close**; and a fragment written in
|
|
1484
1510
|
capitals (`QU'<em>il</em>`) fails the first-code-point folding this rule shares with item 8's
|
|
1485
1511
|
listed veto. Neither is a candidate for a wider mechanism: widening the folding is
|
|
1486
1512
|
`ARCHITECTURE.md` §4.4's banned territory, and the plural possessive has no local evidence at
|
|
@@ -1539,7 +1565,8 @@ consulted only when the mark is flush against an inline span boundary. It closes
|
|
|
1539
1565
|
§1** — a possessive or elision written against a span (`` `x`'s ``, `l'<em>idée</em>`) was
|
|
1540
1566
|
classified as a quotation candidate, took the pairing from the author's own mark inside a
|
|
1541
1567
|
quotation, and inverted the pair in every locale. **It does not close issue #53 §3**, the
|
|
1542
|
-
cross-paragraph damage that motivated the report, which
|
|
1568
|
+
cross-paragraph damage that motivated the report, which was tracked separately as issue #54 and
|
|
1569
|
+
is decided rather than fixed (2026-09-24, §3.2): that witness is a *plural* possessive after a
|
|
1543
1570
|
span, and §3.2 records why no veto over these neighbours can reach it. The simpler design — letting the boundary marker satisfy the
|
|
1544
1571
|
medial-elision veto's `ALNUM` test — was implemented, measured, and rejected before any release:
|
|
1545
1572
|
it breaks a quotation that legitimately begins or ends at a span boundary, including the
|
data/lib/polytypo/version.rb
CHANGED
metadata
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
--- !ruby/object:Gem::Specification
|
|
2
2
|
name: polytypo
|
|
3
3
|
version: !ruby/object:Gem::Version
|
|
4
|
-
version: 1.
|
|
4
|
+
version: 1.6.1
|
|
5
5
|
platform: ruby
|
|
6
6
|
authors:
|
|
7
7
|
- Iurii Rogulia
|
|
@@ -35,6 +35,7 @@ files:
|
|
|
35
35
|
- LICENSE
|
|
36
36
|
- README.md
|
|
37
37
|
- lib/polytypo.rb
|
|
38
|
+
- lib/polytypo/data/.not-canonical
|
|
38
39
|
- lib/polytypo/data/README.md
|
|
39
40
|
- lib/polytypo/data/UNICODE
|
|
40
41
|
- lib/polytypo/data/VERSION
|
|
@@ -56,6 +57,7 @@ files:
|
|
|
56
57
|
- lib/polytypo/data/fixtures/pt-PT.json
|
|
57
58
|
- lib/polytypo/data/fixtures/ru.json
|
|
58
59
|
- lib/polytypo/data/fixtures/sv.json
|
|
60
|
+
- lib/polytypo/data/fixtures/tr.json
|
|
59
61
|
- lib/polytypo/data/fixtures/uk.json
|
|
60
62
|
- lib/polytypo/data/locales/cs.json
|
|
61
63
|
- lib/polytypo/data/locales/de-CH.json
|
|
@@ -75,6 +77,7 @@ files:
|
|
|
75
77
|
- lib/polytypo/data/locales/registry.json
|
|
76
78
|
- lib/polytypo/data/locales/ru.json
|
|
77
79
|
- lib/polytypo/data/locales/sv.json
|
|
80
|
+
- lib/polytypo/data/locales/tr.json
|
|
78
81
|
- lib/polytypo/data/locales/uk.json
|
|
79
82
|
- lib/polytypo/data/rules/analyze.md
|
|
80
83
|
- lib/polytypo/data/rules/apostrophe.md
|