tdk-api-wrapper 1.2.1 → 1.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +14 -3
- package/dist/{chunk-6BTOGV2M.mjs → chunk-SLNXKZKR.mjs} +284 -21
- package/dist/cli.js +329 -23
- package/dist/cli.mjs +46 -3
- package/dist/index.d.mts +97 -10
- package/dist/index.d.ts +97 -10
- package/dist/index.js +284 -21
- package/dist/index.mjs +1 -1
- package/package.json +1 -1
- package/src/cli.ts +41 -1
- package/src/tdk.ts +298 -20
- package/src/types.ts +10 -0
package/README.md
CHANGED
|
@@ -35,6 +35,9 @@ tdk kural kısaltmalar
|
|
|
35
35
|
tdk karsilastir kalem kağıt
|
|
36
36
|
tdk analiz "Bu güzel kalem masanın üstünde duruyor"
|
|
37
37
|
tdk oneri kale
|
|
38
|
+
tdk kubbealti merhaba
|
|
39
|
+
tdk nisanyan merhaba
|
|
40
|
+
tdk viki merhaba
|
|
38
41
|
```
|
|
39
42
|
|
|
40
43
|
Herhangi bir komuta `--json` bayrağı eklendiğinde çıktı, insan-okunur metin yerine tek satırlık JSON olarak basılır (script/otomasyon kullanımı için):
|
|
@@ -68,7 +71,7 @@ Aşağıdaki metotlar `TDK` sınıfı üzerinden statik olarak erişilebilir dur
|
|
|
68
71
|
- **`TDK.syllabicate(word)`**: Kelimeyi Türkçe heceleme kurallarına göre doğru hecelerine ayırır (Örn: `['mu', 'vaf', 'fa', 'ki', 'yet']`). API isteği atmaz, çok hızlıdır.
|
|
69
72
|
- **`TDK.checkVowelHarmony(word)`**: Kelimenin büyük ünlü uyumuna uyup uymadığını (boolean) kontrol eder.
|
|
70
73
|
- **`TDK.getPartOfSpeech(word)`**: Kelimenin sözcük türünü (isim, sıfat, zarf vb.) döndürür.
|
|
71
|
-
- **`TDK.checkSpelling(word)`**: Kelimenin doğru yazılıp yazılmadığını kontrol eder. Önce TDK'nin "sık yapılan yanlışlar" listesinde tam eşleşme arar; bulamazsa TDK'nin ~81 bin kelimelik tam madde listesi üzerinde edit-distance
|
|
74
|
+
- **`TDK.checkSpelling(word)`**: Kelimenin doğru yazılıp yazılmadığını kontrol eder. Önce TDK'nin "sık yapılan yanlışlar" listesinde tam eşleşme arar; bulamazsa TDK'nin ~81 bin kelimelik tam madde listesi üzerinde Damerau-Levenshtein edit-distance ile en yakın kelimeyi önerir (bitişik harf yer değiştirmelerini de tek düzeltme sayar; örn. `herkez` → `herkes`, `mektub` → `mektup`, `yanlız` → `yalnız`). Kelime sıklığı verisi olmadığı için nadiren aynı mesafedeki iki aday arasında beklenenden farklı biri seçilebilir.
|
|
72
75
|
- **`TDK.getCompoundWords(word)`**: Aranan kelime ile oluşturulmuş birleşik kelimeleri (Örn: dolma kalem) listeler.
|
|
73
76
|
|
|
74
77
|
### 3. Edebi ve Kültürel Analiz
|
|
@@ -79,7 +82,7 @@ Aşağıdaki metotlar `TDK` sınıfı üzerinden statik olarak erişilebilir dur
|
|
|
79
82
|
- **`TDK.groupByOrigin(words)`**: Bir kelime listesini etimolojik kökenlerine göre gruplar (bulunamayanlar `"Bilinmiyor"` altında toplanır).
|
|
80
83
|
- **`TDK.getSynonyms(word)`** / **`TDK.getAntonyms(word)`**: Kelimenin eş/zıt anlamlılarını döner (undocumented `gts-yeni` endpoint'i üzerinden; sonuç bulunamazsa `[]`).
|
|
81
84
|
- **`TDK.compareWords(a, b)`**: İki kelimeyi anlam sayısı, köken, hece bölünüşü ve büyük ünlü uyumu açısından yan yana karşılaştırır.
|
|
82
|
-
- **`TDK.analyzeText(text)`**: Bir metindeki (Türkçe bağlaçlar/edatlar hariç) her benzersiz kelimeyi tek tek arayıp ilk anlamını ve kökenini döner.
|
|
85
|
+
- **`TDK.analyzeText(text)`**: Bir metindeki (Türkçe bağlaçlar/edatlar hariç) her benzersiz kelimeyi tek tek arayıp ilk anlamını ve kökenini döner. Not: TDK yalnızca yalın (sözlük) biçimleri indeksliyor, morfolojik analiz yapmıyor — bu yüzden "evde", "dildir" gibi ek almış kelimeler kökleri (`ev`, `dil`) sözlükte olsa bile `found: false` döner; bu veri kaynağının doğal bir sınırlılığıdır.
|
|
83
86
|
|
|
84
87
|
### 4. Yardımcı Metotlar
|
|
85
88
|
- **`TDK.getSuggestions(prefix)`**: TDK'nin ~81 bin kelimelik tam madde listesi üzerinden önek bazlı otomatik tamamlama önerileri döner (ilk çağrıda listeyi indirip önbelleğe alır, sonraki çağrılar anlıktır).
|
|
@@ -88,7 +91,15 @@ Aşağıdaki metotlar `TDK` sınıfı üzerinden statik olarak erişilebilir dur
|
|
|
88
91
|
- **`TDK.getWordOfTheDay()`**: `getDailyContent()`'in üzerine ince bir katman; günün kelimesini ve tüm anlamlarını `{ word, meanings }` şeklinde döner.
|
|
89
92
|
- **`TDK.getRandomWord()`**: Günün içeriğindeki kelime ve atasözü havuzundan rastgele bir tanesini `{ type: "kelime" | "atasoz", madde, anlam }` şeklinde seçer (not: tüm sözlük değil, sadece o günkü içerik havuzundan seçim yapar).
|
|
90
93
|
- **`TDK.getKurallar()`**: TDK'nin `/icerik` akışının o an döndürdüğü yazım kuralı sayfa(lar)ını `{ adi, url }` şeklinde listeler. Not: bu sabit bir katalog değildir — `/icerik` her istekte, yaklaşık yirmi kurallık bir havuzdan rastgele tek bir kural döndürür.
|
|
91
|
-
- **`TDK.getRule(name)`**: Adı verilen (küçük/büyük harf duyarsız, alt dize eşleşmesi) yazım kuralının tam metnini `tdk.gov.tr`'den çekip düz metne çevirir. `getKurallar()`'ın rastgeleliği yüzünden istenen kuralı bulana kadar
|
|
94
|
+
- **`TDK.getRule(name)`**: Adı verilen (küçük/büyük harf duyarsız, alt dize eşleşmesi) yazım kuralının tam metnini `tdk.gov.tr`'den çekip düz metne çevirir. `getKurallar()`'ın rastgeleliği yüzünden istenen kuralı bulana kadar eşzamanlı gruplar hâlinde (toplam en fazla 25 deneme, ~5 round-trip'e sığdırılmış) yeniden dener; bulamazsa veya sayfa ayrıştırılamazsa `null` döner.
|
|
95
|
+
|
|
96
|
+
### 5. Diğer Sözlük Kaynakları
|
|
97
|
+
|
|
98
|
+
TDK dışındaki bu üç kaynak da her zaman kullanılabilir/dokümante edilmiş resmî API'ler değildir; her biri **fragile scraping** (kırılgan, dokümante edilmemiş entegrasyon) — kaynak taraflarında bir değişiklik olursa `null`/`[]` dönerler, hataya düşmezler. Verinin telif/kullanım koşulları kaynağa göre farklıdır: Wiktionary içeriği CC BY-SA lisanslıdır (açık); Nişanyan Sözlük ücretsiz, açık bir kişisel/akademik kaynaktır; **Kubbealtı Lugatı ise ticari bir sözlük ürünüdür** — bu kütüphane onu da dokümante edilmemiş bir uç noktadan çekebiliyor olsa da, kullanımınızı Kubbealtı'nın kendi kullanım şartlarına göre değerlendirmeniz önerilir.
|
|
99
|
+
|
|
100
|
+
- **`TDK.getKubbealti(word)`**: Kubbealtı Lugatı'nın ("Misalli Büyük Türkçe Sözlük") verilerini `{ kelime, anlam }` dizisi olarak döner (`anlam` zengin tipografi içeren ham HTML'dir). `getKubbealtiMeanings(word)` aynı veriyi düz metne çevirir. `getKubbealtiSuggestions(prefix)` Kubbealtı'nın kendi otomatik tamamlama uç noktasını kullanır (TDK'nin `getSuggestions()`'ından bağımsız, ayrı bir veri kaynağı). Not: Kubbealtı'nın veri sunucusu (`eski.lugatim.com`) sertifika zincirini eksik gönderiyor; bu kütüphane eksik ara sertifikaları ekleyerek zinciri düzgün doğruluyor (doğrulamayı kapatmıyor) — Let's Encrypt bu ara sertifikayı döndürürse bu entegrasyon `null` dönmeye başlar.
|
|
101
|
+
- **`TDK.getNisanyan(word)`**: Nişanyan Sözlük'ten kelimenin etimoloji paragrafını düz metin olarak döner; kelime bulunamazsa `null`.
|
|
102
|
+
- **`TDK.getWiktionary(word)`**: Türkçe Vikisözlük'ten (`tr.wiktionary.org`) resmî MediaWiki API'si (`action=query&prop=extracts`) üzerinden veri çeker — bu üçü arasında scraping olmayan, resmî ve en kararlı olanı. `{ raw, sections }` döner; `sections` metni `== Köken ==`, `=== Söyleniş ===` gibi başlıklara göre bir sözlüğe ayırır. `getWiktionarySection(word, sectionName)` tek bir bölümü (örn. `"Köken"`) büyük/küçük harf duyarsız süzer.
|
|
92
103
|
|
|
93
104
|
## Hata Yönetimi
|
|
94
105
|
|
|
@@ -30,9 +30,87 @@ import * as fs from "fs";
|
|
|
30
30
|
import * as path from "path";
|
|
31
31
|
import * as os from "os";
|
|
32
32
|
import * as https from "https";
|
|
33
|
+
import * as tls from "tls";
|
|
33
34
|
var TDK = class {
|
|
34
35
|
static BASE_URL = "https://sozluk.gov.tr";
|
|
35
36
|
static AUDIO_API_HOST = "api.sozluk.gov.tr";
|
|
37
|
+
static KUBBEALTI_HOST = "eski.lugatim.com";
|
|
38
|
+
/**
|
|
39
|
+
* `eski.lugatim.com` (Kubbealtı Lugatı's data API) sends only its leaf
|
|
40
|
+
* certificate during the TLS handshake, omitting the intermediates a
|
|
41
|
+
* correctly configured server would include — a server-side misconfiguration,
|
|
42
|
+
* not something we should paper over by disabling verification. These are
|
|
43
|
+
* the two certificates the server *should* be sending (fetched from the
|
|
44
|
+
* leaf's own Authority Information Access URLs), supplied here so Node can
|
|
45
|
+
* still build a full, properly verified chain up to a root it already
|
|
46
|
+
* trusts (ISRG Root X1). If Let's Encrypt rotates this intermediate, this
|
|
47
|
+
* stops working and every Kubbealtı call fails closed to `null` — same
|
|
48
|
+
* fail-closed contract as the rest of this file's fragile integrations.
|
|
49
|
+
*/
|
|
50
|
+
static KUBBEALTI_EXTRA_CA = [
|
|
51
|
+
`-----BEGIN CERTIFICATE-----
|
|
52
|
+
MIIE2jCCAsKgAwIBAgIQTr0klH4k05SALYSlL9WzGTANBgkqhkiG9w0BAQsFADAu
|
|
53
|
+
MQswCQYDVQQGEwJVUzENMAsGA1UEChMESVNSRzEQMA4GA1UEAxMHUm9vdCBZUjAe
|
|
54
|
+
Fw0yNTA5MDMwMDAwMDBaFw0yODA5MDIyMzU5NTlaMDMxCzAJBgNVBAYTAlVTMRYw
|
|
55
|
+
FAYDVQQKEw1MZXQncyBFbmNyeXB0MQwwCgYDVQQDEwNZUjIwggEiMA0GCSqGSIb3
|
|
56
|
+
DQEBAQUAA4IBDwAwggEKAoIBAQDZ0LxwBppqh84luqMerV/eeL/fXQ7mLQQv1Lnp
|
|
57
|
+
WKZbyvGpx6wh6AfnslAnF6ewTkcHA+gSOoBvm3Dfm06AuGiF+KRut4fAcowqnAQQ
|
|
58
|
+
CW98+QPP/eOv/wug7Iyk4NkOxf2I6g2f55T6nJoOTLFcukeRq80JGQEYan+dPFr9
|
|
59
|
+
OGUgQK2hGKgNkW87pappsOAuUJcroYhRt5uUis4qaZireiseu32gzDJNBAiKtsvd
|
|
60
|
+
6HX4v25bpkRNcS/B/Gtc9kVbUpD+2PLPxdei3Tim55k4tfAEXwD2qyiPTxrTNq6l
|
|
61
|
+
N+AMr5g2c1dNqkOTwjxeV6L5lpP1rGiYvLnRaPlOqyZRPW+5AgMBAAGjge4wgesw
|
|
62
|
+
DgYDVR0PAQH/BAQDAgGGMBMGA1UdJQQMMAoGCCsGAQUFBwMBMBIGA1UdEwEB/wQI
|
|
63
|
+
MAYBAf8CAQAwHQYDVR0OBBYEFEAVLSZ57TIgnt+ach3WMh+BDIEMMB8GA1UdIwQY
|
|
64
|
+
MBaAFN7nW2DQIm1AKH0/DQH+pLVStFGUMDIGCCsGAQUFBwEBBCYwJDAiBggrBgEF
|
|
65
|
+
BQcwAoYWaHR0cDovL3lyLmkubGVuY3Iub3JnLzATBgNVHSAEDDAKMAgGBmeBDAEC
|
|
66
|
+
ATAnBgNVHR8EIDAeMBygGqAYhhZodHRwOi8veXIuYy5sZW5jci5vcmcvMA0GCSqG
|
|
67
|
+
SIb3DQEBCwUAA4ICAQB0ZUQWZ9/Yn9COEpo+JfecMnB0h0vwDm/M66IqXqw3LoaL
|
|
68
|
+
mx9lZvRTeDIS67PUeI3yCA2W6PKRD0/FE/G57lOmS+Xy5AaaL00ICGOqjNcCaMWW
|
|
69
|
+
8o8nevHOd4i4lqgtznE/28QwlcdJyF8yBiWHpnyjhEpmNWJURgOCOg2xpwRMBCsj
|
|
70
|
+
MScqYPtOhBeuYQvSwAEeTML2Ukh6uGuX4E14q65Ja8cdjF5bAldnP1eE4FBaAwsZ
|
|
71
|
+
G2fOqqrKV03Y85Nw2btedP1AtliQuJZs/Jo/gXxXdc7LrH3McgnpnbTiAncX7yES
|
|
72
|
+
hP6kzQejllqMCIt52HOjxDGWafS7Xw+DKwqmH+Eqy8dcbOuag/1AYlQoKNVK3F5q
|
|
73
|
+
Hh6tEDiMqQcLIibGKteE6iHo4A/bIScbzrhXUYuism42ZYzmc48FMVIH3qy4L84E
|
|
74
|
+
TdAH2gtxw0PAhvRVXp8HP7wfngpzsN/8xOTpeRSbM4+Qbc56G6+Bifmv6sk1ieQb
|
|
75
|
+
NA3wJdl4DDUuQSV8hBgx6zoI1ZSGORprDFux7c6rhc77QZMSRrEgomBeklervEve
|
|
76
|
+
86ylWmZ3WWHV6RLMi8xNvjd71r4EPIGgY7BZU/VPBkq+uA7Gb6mbJnFgV43uh3xy
|
|
77
|
+
LRFgxIAphIukwTGSMZZR+AI+Qnp0BYTWovHXozOf3H8r6hozEoT02JHn0AeTfA==
|
|
78
|
+
-----END CERTIFICATE-----`,
|
|
79
|
+
`-----BEGIN CERTIFICATE-----
|
|
80
|
+
MIIF9DCCA9ygAwIBAgIRAPJLbRf52a18scn+p4eCaZ8wDQYJKoZIhvcNAQELBQAw
|
|
81
|
+
TzELMAkGA1UEBhMCVVMxKTAnBgNVBAoTIEludGVybmV0IFNlY3VyaXR5IFJlc2Vh
|
|
82
|
+
cmNoIEdyb3VwMRUwEwYDVQQDEwxJU1JHIFJvb3QgWDEwHhcNMjYwNTEzMDAwMDAw
|
|
83
|
+
WhcNMzIwOTAyMjM1OTU5WjAuMQswCQYDVQQGEwJVUzENMAsGA1UEChMESVNSRzEQ
|
|
84
|
+
MA4GA1UEAxMHUm9vdCBZUjCCAiIwDQYJKoZIhvcNAQEBBQADggIPADCCAgoCggIB
|
|
85
|
+
ANvGJnN78CTJdWL3+eGfsLN5TrNBJs+VH9hRXqRbwxu9sGNiB0BD1fcOxbSUQCJI
|
|
86
|
+
M1xE13Db+5Cw1w0s0EBYsvuIP/6joF0w8cuImbgR1OGgYbSQ4OpzI+DG8SGuTlcE
|
|
87
|
+
873OCS+kh3srlo6vl43M5OJg4Aeo1sfHp6kTJDoIiFBNJAY+OKfX/FUvYKuhjT+n
|
|
88
|
+
o49lmqmupSBI5PkBQiqrEGtWU5uxU/cQWHGu8jSjFBznZqvbNPLMXMLFxCb3WTfr
|
|
89
|
+
JBXXjqvWG+v4bjzxjjeAtOlU7qarRDvNOyAuQYLln904M+faKx8hnLCpJ15ZqaEg
|
|
90
|
+
cNlY+9MMWcC5yvL2A2j3l9+2buggZX+dOE91zYmIdawTvSZuVvlbRrAlLxIB6pwM
|
|
91
|
+
BjneXCjYQ8+3BCCjssbSNpZU3hTcBDdhfAlEDlYr6pEatnMdmDT5BqnKC92bd0Eh
|
|
92
|
+
M1fbLHioLccLCuievT8ZkPhZrq7Mii7gNXAcUEAR8+lzYal+9zTg7C5DALyVOeG/
|
|
93
|
+
CqfRAMn1KSHCR0NSA6P8tn/mGRlnCct5rtVCLnVySVpU6H1qGg3DgTOuskf8eahT
|
|
94
|
+
MiYbI5ezPJmO5ertalskQ1utp74+eDy92PI4ftHKTbq9IWhH4YZKh3WnJEIt+oQv
|
|
95
|
+
lYZbY8tpEroKrFB6PFGzrJIDRyts4HqvuH52RFj2zv/BAgMBAAGjgeswgegwDgYD
|
|
96
|
+
VR0PAQH/BAQDAgEGMBMGA1UdJQQMMAoGCCsGAQUFBwMBMA8GA1UdEwEB/wQFMAMB
|
|
97
|
+
Af8wHQYDVR0OBBYEFN7nW2DQIm1AKH0/DQH+pLVStFGUMB8GA1UdIwQYMBaAFHm0
|
|
98
|
+
WeZ7tuXkAXOACIjIGlj26ZtuMDIGCCsGAQUFBwEBBCYwJDAiBggrBgEFBQcwAoYW
|
|
99
|
+
aHR0cDovL3gxLmkubGVuY3Iub3JnLzATBgNVHSAEDDAKMAgGBmeBDAECATAnBgNV
|
|
100
|
+
HR8EIDAeMBygGqAYhhZodHRwOi8veDEuYy5sZW5jci5vcmcvMA0GCSqGSIb3DQEB
|
|
101
|
+
CwUAA4ICAQA8spSI95KKfn2W6GMmDpHBJSPaLbsS3W93cijJCRCYAc1fsJgL1FIL
|
|
102
|
+
7C0C9ecPOdcwB2fi0Dk2p94j9iTJCxmt5CFSKLRWwnXT2MMSXexVxqoVB79BdWPx
|
|
103
|
+
VXETkVme/qYSAuKVHh5Ps+5BixgmwS1JkjSAc+MfrUbNssVEEnH0aEiAh+rotXAV
|
|
104
|
+
JSP/Ye7LJPEwD9DWG72vVWbhAcuOf5OLjz57Ctk7MgQHynZ7+PlHJtajroCaIbtC
|
|
105
|
+
r6tcZZaAwUQm+jQyeWdV+2hv9deOYFmKeQyjjcSrN5Nadrw+L9DZJLbA1HqeNvLh
|
|
106
|
+
BgqpP0fvJq2N6EtD574N6eMI7uMsJTnji2UDz9el5XLSv9fqJMuDQtYVb2oTNoKp
|
|
107
|
+
oUqhxPVC0aq4eG5MESaIdn8b5ZGSSeAJLMHXljEdlNza+ncfkviXk1POLnnFdvx8
|
|
108
|
+
/gk6M374WbLWFXw8N141B/Rl/tINGfl1TxOIiqtiMYkL02RSGb1kq34BL9NPP27z
|
|
109
|
+
RGMuHGnzS3hFIrRTfKxrzUZ9RzQWzEG3K6fJ3r2nqSltkeytis9DIBoFY9VmVyjL
|
|
110
|
+
M71DMi+y1+TRSJVClEMwvA4yL++7q9XZx5r5wBRWB4kQTKH5qyoZnDw7iiuh1lID
|
|
111
|
+
yDFx8r7i9vIJU5HS3moZLkYWAOilMaV9N56A9Bgb6dNcHkvg3NoaYA==
|
|
112
|
+
-----END CERTIFICATE-----`
|
|
113
|
+
];
|
|
36
114
|
// Cache Mechanism
|
|
37
115
|
static isCacheEnabled = false;
|
|
38
116
|
static wordCache = /* @__PURE__ */ new Map();
|
|
@@ -399,11 +477,16 @@ var TDK = class {
|
|
|
399
477
|
for (const candidate of this.autocompleteCache) {
|
|
400
478
|
if (candidate.includes(" ") || candidate !== candidate.toLocaleLowerCase("tr-TR"))
|
|
401
479
|
continue;
|
|
402
|
-
|
|
403
|
-
|
|
404
|
-
|
|
405
|
-
|
|
406
|
-
|
|
480
|
+
if (Math.abs(candidate.length - cleanWord.length) > 2)
|
|
481
|
+
continue;
|
|
482
|
+
const distance = this.damerauLevenshtein(cleanWord, candidate);
|
|
483
|
+
if (distance === 0)
|
|
484
|
+
continue;
|
|
485
|
+
const firstMismatch = candidate[0] === cleanWord[0] ? 0 : 1;
|
|
486
|
+
const lengthMismatch = candidate.length === cleanWord.length ? 0 : 1;
|
|
487
|
+
const better = !best || distance < best.distance || distance === best.distance && firstMismatch < best.firstMismatch || distance === best.distance && firstMismatch === best.firstMismatch && lengthMismatch < best.lengthMismatch;
|
|
488
|
+
if (better) {
|
|
489
|
+
best = { candidate, distance, firstMismatch, lengthMismatch };
|
|
407
490
|
}
|
|
408
491
|
}
|
|
409
492
|
if (best && best.distance <= 2) {
|
|
@@ -482,24 +565,33 @@ var TDK = class {
|
|
|
482
565
|
* case-insensitively, substring match) from `tdk.gov.tr`. Since `/icerik`
|
|
483
566
|
* hands back a single randomly-rotated rule per request (out of a pool of
|
|
484
567
|
* roughly twenty) rather than a fixed catalog, a single `getKurallar()`
|
|
485
|
-
* draw would rarely match a given name — this re-draws
|
|
486
|
-
*
|
|
487
|
-
*
|
|
488
|
-
*
|
|
489
|
-
*
|
|
490
|
-
*
|
|
491
|
-
*
|
|
568
|
+
* draw would rarely match a given name — this re-draws until it finds a
|
|
569
|
+
* match or gives up. Draws happen in concurrent batches (each `/icerik`
|
|
570
|
+
* request is independent and stateless) rather than one-at-a-time with a
|
|
571
|
+
* delay: same total sample size (25) and hit probability as a sequential
|
|
572
|
+
* loop, but bounded to a handful of round-trips instead of 25 of them, so
|
|
573
|
+
* a miss resolves in roughly one round-trip time instead of several
|
|
574
|
+
* seconds. Every draw bypasses `dailyContentCache` — without that, once
|
|
575
|
+
* `enableCache(true)` is on, every attempt would just re-read the same
|
|
576
|
+
* cached `/icerik` response and could never find a rule outside whatever
|
|
577
|
+
* the first draw happened to be. Returns `null` if no match turns up
|
|
578
|
+
* within the attempt budget or the matched page can't be parsed.
|
|
492
579
|
*/
|
|
493
580
|
static async getRule(name) {
|
|
494
581
|
if (!name || name.trim() === "")
|
|
495
582
|
return null;
|
|
496
583
|
const target = name.trim().toLocaleLowerCase("tr-TR");
|
|
497
|
-
|
|
498
|
-
|
|
499
|
-
|
|
500
|
-
|
|
501
|
-
|
|
502
|
-
|
|
584
|
+
const BATCH_SIZE = 5;
|
|
585
|
+
const ROUNDS = 5;
|
|
586
|
+
for (let round = 0; round < ROUNDS; round++) {
|
|
587
|
+
const batches = await Promise.all(
|
|
588
|
+
Array.from({ length: BATCH_SIZE }, () => this.getKurallar(true))
|
|
589
|
+
);
|
|
590
|
+
for (const rules of batches) {
|
|
591
|
+
const match = rules.find((r) => r.adi.toLocaleLowerCase("tr-TR").includes(target));
|
|
592
|
+
if (match)
|
|
593
|
+
return this.fetchRuleText(match.url);
|
|
594
|
+
}
|
|
503
595
|
}
|
|
504
596
|
return null;
|
|
505
597
|
}
|
|
@@ -528,7 +620,166 @@ var TDK = class {
|
|
|
528
620
|
}
|
|
529
621
|
}
|
|
530
622
|
static htmlToPlainText(html) {
|
|
531
|
-
return html.replace(/<br\s*\/?>/gi, "\n").replace(/<\/(p|div)>/gi, "\n\n").replace(/<[^>]+>/g, "").replace(/ /gi, " ").replace(/&
|
|
623
|
+
return html.replace(/<br\s*\/?>/gi, "\n").replace(/<\/(p|div)>/gi, "\n\n").replace(/<[^>]+>/g, "").replace(/ /gi, " ").replace(/</gi, "<").replace(/>/gi, ">").replace(/"/gi, '"').replace(/'|’/gi, "'").replace(/&/gi, "&").replace(/[ \t]+/g, " ").replace(/[ \t]*\n[ \t]*/g, "\n").replace(/\n{3,}/g, "\n\n").trim();
|
|
624
|
+
}
|
|
625
|
+
/**
|
|
626
|
+
* GETs a JSON path from Kubbealtı Lugatı's data API (`eski.lugatim.com`),
|
|
627
|
+
* supplying `KUBBEALTI_EXTRA_CA` to work around that host's incomplete
|
|
628
|
+
* certificate chain (see the constant's doc comment). Fails closed to
|
|
629
|
+
* `null` on any error — network, TLS, HTTP, or JSON parse.
|
|
630
|
+
*/
|
|
631
|
+
static fetchKubbealtiJson(path2) {
|
|
632
|
+
return new Promise((resolve) => {
|
|
633
|
+
const req = https.request(
|
|
634
|
+
{
|
|
635
|
+
hostname: this.KUBBEALTI_HOST,
|
|
636
|
+
path: path2,
|
|
637
|
+
method: "GET",
|
|
638
|
+
ca: [...tls.rootCertificates, ...this.KUBBEALTI_EXTRA_CA],
|
|
639
|
+
headers: {
|
|
640
|
+
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/128.0.0.0 Safari/537.36"
|
|
641
|
+
}
|
|
642
|
+
},
|
|
643
|
+
(res) => {
|
|
644
|
+
if (res.statusCode !== 200) {
|
|
645
|
+
res.resume();
|
|
646
|
+
resolve(null);
|
|
647
|
+
return;
|
|
648
|
+
}
|
|
649
|
+
let body = "";
|
|
650
|
+
res.on("data", (chunk) => body += chunk);
|
|
651
|
+
res.on("end", () => {
|
|
652
|
+
try {
|
|
653
|
+
resolve(JSON.parse(body));
|
|
654
|
+
} catch {
|
|
655
|
+
resolve(null);
|
|
656
|
+
}
|
|
657
|
+
});
|
|
658
|
+
}
|
|
659
|
+
);
|
|
660
|
+
req.on("error", () => resolve(null));
|
|
661
|
+
req.end();
|
|
662
|
+
});
|
|
663
|
+
}
|
|
664
|
+
/**
|
|
665
|
+
* Returns Kubbealtı Lugatı ("Misalli Büyük Türkçe Sözlük") entries for a
|
|
666
|
+
* word, scraped from the site's own data API — undocumented, and Kubbealtı
|
|
667
|
+
* Lugatı is a commercial dictionary product, unlike TDK's or Wiktionary's
|
|
668
|
+
* openly-published data, so use this in line with their terms. `anlam` is
|
|
669
|
+
* raw HTML (rich typography markup); use `getKubbealtiMeanings()` for
|
|
670
|
+
* plain text. Returns `null` on any fetch/parse failure, `[]` if the word
|
|
671
|
+
* isn't found.
|
|
672
|
+
*/
|
|
673
|
+
static async getKubbealti(word) {
|
|
674
|
+
if (!word || word.trim() === "")
|
|
675
|
+
return null;
|
|
676
|
+
const data = await this.fetchKubbealtiJson(`/rest/s/${encodeURIComponent(word.trim())}/`);
|
|
677
|
+
if (!data || !Array.isArray(data.content))
|
|
678
|
+
return null;
|
|
679
|
+
return data.content.map((entry) => ({ kelime: entry.kelime, anlam: entry.anlam }));
|
|
680
|
+
}
|
|
681
|
+
/**
|
|
682
|
+
* Same as `getKubbealti()` but with each entry's `anlam` HTML stripped to
|
|
683
|
+
* plain text via `htmlToPlainText()`.
|
|
684
|
+
*/
|
|
685
|
+
static async getKubbealtiMeanings(word) {
|
|
686
|
+
const entries = await this.getKubbealti(word);
|
|
687
|
+
if (!entries)
|
|
688
|
+
return null;
|
|
689
|
+
return entries.map((e) => this.htmlToPlainText(e.anlam));
|
|
690
|
+
}
|
|
691
|
+
/**
|
|
692
|
+
* Autocomplete suggestions from Kubbealtı Lugatı's own typeahead endpoint
|
|
693
|
+
* (separate from `getSuggestions()`, which uses TDK's data).
|
|
694
|
+
*/
|
|
695
|
+
static async getKubbealtiSuggestions(prefix) {
|
|
696
|
+
if (!prefix || prefix.trim() === "")
|
|
697
|
+
return [];
|
|
698
|
+
const data = await this.fetchKubbealtiJson(`/rest/word-search/${encodeURIComponent(prefix.trim())}`);
|
|
699
|
+
if (!Array.isArray(data))
|
|
700
|
+
return [];
|
|
701
|
+
return data.map((item) => item.display).filter(Boolean);
|
|
702
|
+
}
|
|
703
|
+
/**
|
|
704
|
+
* Returns the etymology paragraph for a word from Nişanyan Sözlük, scraped
|
|
705
|
+
* from that page's server-rendered `<meta name="description">` tag (the
|
|
706
|
+
* page already puts the full etymology text there for SEO, so no need to
|
|
707
|
+
* parse the site's internal SvelteKit data format). Returns `null` if the
|
|
708
|
+
* word isn't found (the page falls back to a generic site tagline in that
|
|
709
|
+
* case) or the request fails.
|
|
710
|
+
*/
|
|
711
|
+
static async getNisanyan(word) {
|
|
712
|
+
if (!word || word.trim() === "")
|
|
713
|
+
return null;
|
|
714
|
+
try {
|
|
715
|
+
const response = await fetch(
|
|
716
|
+
`https://www.nisanyansozluk.com/kelime/${encodeURIComponent(word.trim().toLocaleLowerCase("tr-TR"))}`,
|
|
717
|
+
{ headers: { "User-Agent": "TDK-API-Nodejs-Wrapper/1.0" } }
|
|
718
|
+
);
|
|
719
|
+
if (!response.ok)
|
|
720
|
+
return null;
|
|
721
|
+
const html = await response.text();
|
|
722
|
+
const match = html.match(/<meta name="description" content="([^"]*)"/);
|
|
723
|
+
if (!match)
|
|
724
|
+
return null;
|
|
725
|
+
const description = this.htmlToPlainText(match[1]);
|
|
726
|
+
if (description === "\xC7a\u011Fda\u015F T\xFCrk\xE7enin Etimolojisi")
|
|
727
|
+
return null;
|
|
728
|
+
return description;
|
|
729
|
+
} catch {
|
|
730
|
+
return null;
|
|
731
|
+
}
|
|
732
|
+
}
|
|
733
|
+
/**
|
|
734
|
+
* Returns the Turkish Wiktionary (`tr.wiktionary.org`) entry for a word,
|
|
735
|
+
* via MediaWiki's official Action API (`action=query&prop=extracts`) — no
|
|
736
|
+
* scraping involved, this is a stable, documented public API. `sections`
|
|
737
|
+
* splits the plain-text extract on its `== Heading ==`/`=== Heading ===`
|
|
738
|
+
* markers (e.g. "Köken", "Söyleniş", "Ad") for convenience; `raw` has the
|
|
739
|
+
* unsplit text. Returns `null` if the page doesn't exist or the request
|
|
740
|
+
* fails.
|
|
741
|
+
*/
|
|
742
|
+
static async getWiktionary(word) {
|
|
743
|
+
if (!word || word.trim() === "")
|
|
744
|
+
return null;
|
|
745
|
+
try {
|
|
746
|
+
const url = `https://tr.wiktionary.org/w/api.php?action=query&prop=extracts&titles=${encodeURIComponent(
|
|
747
|
+
word.trim()
|
|
748
|
+
)}&format=json&explaintext=1&formatversion=2`;
|
|
749
|
+
const response = await fetch(url, { headers: { "User-Agent": "TDK-API-Nodejs-Wrapper/1.0" } });
|
|
750
|
+
if (!response.ok)
|
|
751
|
+
return null;
|
|
752
|
+
const data = await response.json();
|
|
753
|
+
const page = data?.query?.pages?.[0];
|
|
754
|
+
if (!page || page.missing || !page.extract)
|
|
755
|
+
return null;
|
|
756
|
+
const raw = page.extract;
|
|
757
|
+
const sections = {};
|
|
758
|
+
const parts = raw.split(/\n(={2,4})\s*(.+?)\s*\1\n/);
|
|
759
|
+
for (let i = 1; i < parts.length; i += 3) {
|
|
760
|
+
const title = parts[i + 1]?.trim();
|
|
761
|
+
const content = parts[i + 2]?.trim();
|
|
762
|
+
if (title)
|
|
763
|
+
sections[title] = content ?? "";
|
|
764
|
+
}
|
|
765
|
+
return { raw, sections };
|
|
766
|
+
} catch {
|
|
767
|
+
return null;
|
|
768
|
+
}
|
|
769
|
+
}
|
|
770
|
+
/**
|
|
771
|
+
* Convenience filter over `getWiktionary()`: returns just one section's
|
|
772
|
+
* text (e.g. `getWiktionarySection(word, "Köken")` for etymology), matched
|
|
773
|
+
* case-insensitively. Returns `null` if the word or the section isn't found.
|
|
774
|
+
*/
|
|
775
|
+
static async getWiktionarySection(word, sectionName) {
|
|
776
|
+
const entry = await this.getWiktionary(word);
|
|
777
|
+
if (!entry)
|
|
778
|
+
return null;
|
|
779
|
+
const key = Object.keys(entry.sections).find(
|
|
780
|
+
(k) => k.toLocaleLowerCase("tr-TR") === sectionName.trim().toLocaleLowerCase("tr-TR")
|
|
781
|
+
);
|
|
782
|
+
return key ? entry.sections[key] : null;
|
|
532
783
|
}
|
|
533
784
|
/**
|
|
534
785
|
* Returns compound words that contain this word.
|
|
@@ -647,6 +898,11 @@ var TDK = class {
|
|
|
647
898
|
* Analyzes every distinct word in a text (Turkish stopwords filtered out),
|
|
648
899
|
* returning each word's first meaning and etymological origin if found.
|
|
649
900
|
* Looks each word up individually (throttled), so scales with text length.
|
|
901
|
+
* TDK only indexes dictionary (dictionary/root) forms, not inflected ones —
|
|
902
|
+
* it does no morphological analysis, and neither does this method: a
|
|
903
|
+
* suffixed word like "evde" or "dildir" (root "ev"/"dil" plus a case/verb
|
|
904
|
+
* suffix) will come back `found: false` even though the root is a real
|
|
905
|
+
* headword. This is an inherent limitation of the data source, not a bug.
|
|
650
906
|
*/
|
|
651
907
|
static async analyzeText(text) {
|
|
652
908
|
const words = text.toLocaleLowerCase("tr-TR").replace(/[^\p{L}\s]/gu, " ").split(/\s+/).filter((w) => w.length > 1 && !this.STOPWORDS.has(w));
|
|
@@ -666,9 +922,13 @@ var TDK = class {
|
|
|
666
922
|
return analyses;
|
|
667
923
|
}
|
|
668
924
|
/**
|
|
669
|
-
*
|
|
925
|
+
* Damerau-Levenshtein edit-distance (optimal string alignment variant):
|
|
926
|
+
* like classic Levenshtein but also counts an adjacent-character
|
|
927
|
+
* transposition (e.g. "yanlız" -> "yalnız") as a single edit instead of
|
|
928
|
+
* two substitutions — a very common class of typo that plain Levenshtein
|
|
929
|
+
* otherwise misses.
|
|
670
930
|
*/
|
|
671
|
-
static
|
|
931
|
+
static damerauLevenshtein(a, b) {
|
|
672
932
|
const dp = Array.from({ length: a.length + 1 }, () => new Array(b.length + 1).fill(0));
|
|
673
933
|
for (let i = 0; i <= a.length; i++)
|
|
674
934
|
dp[i][0] = i;
|
|
@@ -678,6 +938,9 @@ var TDK = class {
|
|
|
678
938
|
for (let j = 1; j <= b.length; j++) {
|
|
679
939
|
const cost = a[i - 1] === b[j - 1] ? 0 : 1;
|
|
680
940
|
dp[i][j] = Math.min(dp[i - 1][j] + 1, dp[i][j - 1] + 1, dp[i - 1][j - 1] + cost);
|
|
941
|
+
if (i > 1 && j > 1 && a[i - 1] === b[j - 2] && a[i - 2] === b[j - 1]) {
|
|
942
|
+
dp[i][j] = Math.min(dp[i][j], dp[i - 2][j - 2] + cost);
|
|
943
|
+
}
|
|
681
944
|
}
|
|
682
945
|
}
|
|
683
946
|
return dp[a.length][b.length];
|