thai-rtgs 0.0.0-stage → 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md ADDED
@@ -0,0 +1,13 @@
1
+ # Changelog
2
+
3
+ ## 0.1.0 — 2026-10-06
4
+
5
+ First release.
6
+
7
+ - `romanize`, `romanizeName`, `analyze`, `syllabify` and `createRomanizer` for Thai → RTGS (Royal Institute, 1999)
8
+ - Rule-based syllable parser: clusters, leading ห/อ, อักษรนำ, ร หัน, karan, silent ร and final vowels, inferred a/o vowels, ฤ/ฦ
9
+ - Pali/Sanskrit combining forms for names and compounds (ธนกฤต thanakrit, ภัทรกมล phattharakamon)
10
+ - Built-in dictionary: 77 provinces, 50 Bangkok districts, irregular words, common given names and name parts; user and per-call dictionaries
11
+ - Word segmentation with `Intl.Segmenter`, with dictionary-only and custom segmenter options
12
+ - Options for letter case, word and syllable separators, ambiguity hyphens, Thai digits, ๆ and ฯ
13
+ - ESM + CommonJS builds with TypeScript types; no runtime dependencies
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 PhaiKub
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md CHANGED
@@ -1,3 +1,127 @@
1
- # Temporary Holding Version
1
+ # thai-rtgs — ถอดอักษรไทยเป็นอักษรโรมันแบบ RTGS
2
2
 
3
- This version is a temporary placeholder for this package. An operational version to replace this has been submitted for review and is awaiting a staged release.
3
+ ระบบถอดอักษรไทยเป็นอักษรโรมันตาม **Royal Thai General System of Transcription (RTGS)** หรือหลักเกณฑ์การถอดอักษรไทยเป็นอักษรโรมันแบบถ่ายเสียงของราชบัณฑิตยสถาน (พ.ศ. 2542) เป็น library ภาษา TypeScript ไม่มี dependency ใช้ได้ทั้ง Node และเบราว์เซอร์ (ESM + CJS + type definitions)
4
+
5
+ ```ts
6
+ import { romanize } from 'thai-rtgs';
7
+
8
+ romanize('กรุงเทพมหานคร', { case: 'title' }); // 'Krung Thep Maha Nakhon'
9
+ romanize('สวัสดีครับ ยินดีต้อนรับสู่ประเทศไทย'); // 'sawatdi khrap yindi tonrap su prathet thai'
10
+ romanize('ถนน', { syllableSeparator: '-' }); // 'tha-non'
11
+ ```
12
+
13
+ ## ทำงานอย่างไร
14
+
15
+ 1. **Normalize** — แก้รูปแบบการพิมพ์ที่ต่างกัน (`ํา`→`ำ`, `เเ`→`แ`, วรรณยุกต์ที่พิมพ์ก่อนสระ, อักขระที่มองไม่เห็น)
16
+ 2. **Tokenize** — แยกข้อความไทย อังกฤษ ตัวเลข เครื่องหมาย `ๆ` `ฯ` และ `ฯลฯ`
17
+ 3. **ตัดคำ** — ใช้ `Intl.Segmenter('th')` แล้วรวม segment ที่ต่อกันเป็นคำในพจนานุกรม (เช่น กรุงเทพมหานคร)
18
+ 4. **พจนานุกรม** — มีชื่อ 77 จังหวัด, 50 เขตของกรุงเทพฯ, คำที่อ่านไม่ตรงกฎ (ราชการ, ผลไม้, ประวัติศาสตร์, ก็, …) และชื่อคน/ส่วนประกอบของนามสกุล (สกุล, ตระกูล, สวัสดิ์, …) คำในพจนานุกรมในตัวยังถูกจับคู่ภายในคำที่ยาวกว่าได้ด้วย (มี|สวัสดิ์ → misawat) ผู้ใช้เพิ่มหรือแก้คำเองได้
19
+ 5. **แยกพยางค์ตามกฎ** — สำหรับคำที่ไม่มีในพจนานุกรม หาวิธีแบ่งพยางค์ที่ "ต้นทุนต่ำสุด" (dynamic programming) จาก template ของสระ พยัญชนะต้น ควบกล้ำ ห นำ ตัวสะกด การันต์ และ ร หัน สระที่ไม่ได้เขียนรูป (ถนน → tha-non) มีต้นทุน จึงถูกเลือกเฉพาะเมื่อตัวสะกดบังคับให้ต้องมี นอกจากนี้ยังรู้จัก **combining forms** ของคำบาลี-สันสกฤตที่อ่านมีพยางค์เชื่อมเมื่อมีส่วนอื่นตามมา (ธน|กฤต → thanakrit, ภัทร|กมล → phattharakamon) และอักษรนำที่เขียนสระหน้าไว้ก่อน (เจริญ → charoen, เสนอ → sanoe)
20
+ 6. **จัดรูปแบบ** — ต่อพยางค์และคำ ใส่ยัติภังค์ (`-`) เมื่ออ่านกำกวม (สะอาด → sa-at) และจัดตัวพิมพ์เล็ก/ใหญ่
21
+
22
+ รายละเอียดกฎทั้งหมดอยู่ใน [docs/rtgs-rules.md](docs/rtgs-rules.md)
23
+
24
+ ## ความแม่นยำ
25
+
26
+ `npm run eval` วัดผลกับชุดทดสอบใน [test/fixtures/golden.ts](test/fixtures/golden.ts):
27
+
28
+ | ชุดทดสอบ | ผล |
29
+ |---|---|
30
+ | คำทั่วไป อ่านด้วยกฎล้วน (ปิดพจนานุกรม) | 225/225 |
31
+ | คำศัพท์ประจำวันที่ไม่ได้ใช้ระหว่างปรับกฎ (เปิดพจนานุกรม) | 190/190 |
32
+ | ชื่อจริงคนไทย (ชื่อจริงอย่างเดียว ไม่รวมนามสกุล) | 124/124 |
33
+ | คำที่ต้องใช้พจนานุกรม | 25/25 |
34
+ | จังหวัด (Title case) | 28/28 |
35
+ | ประโยค | 8/8 |
36
+
37
+ ตัวเลขนี้วัดกับคำที่ใช้ปรับกฎเอง จึงสูงกว่าความแม่นยำจริง ตัวเลขที่บอกความแม่นยำจริงได้ดีกว่าคือผลก่อนปรับ: ชุดคำศัพท์ประจำวันได้ 184/190 (96.8%) และชุดชื่อจริงได้ 106/124 (85%) คำที่ผิดเกือบทั้งหมดเป็นคำบาลี-สันสกฤตที่มีพยางค์เชื่อม (เอกสาร ek-ka-san, ธนกฤต tha-na-krit) ซึ่งกฎอย่างเดียวคาดไม่ได้
38
+
39
+ ทดลองเพิ่มเติมกับรายชื่อจริงของนักเรียนราว 1,900 ชื่อและ 2,400 นามสกุล (ข้อมูลไม่ได้เก็บไว้ใน repo) สุ่มตรวจ 120 รายการที่ไม่ได้ใช้ปรับกฎ ได้ชื่อจริงถูกประมาณ 80% และนามสกุลถูกประมาณ 93% จากนั้นแก้กฎและเพิ่มส่วนประกอบชื่อที่พบบ่อยเข้าไปแล้ว ชื่อและนามสกุลที่ไม่มีในชุดทดสอบอาจยังผิดได้ ให้ตรวจผลที่มี `confidence: "medium"` และเพิ่มคำผ่าน `dictionary`
40
+
41
+ ## Library
42
+
43
+ ```bash
44
+ npm install thai-rtgs
45
+ ```
46
+
47
+ ```ts
48
+ import { romanize, romanizeName, analyze, syllabify, createRomanizer } from 'thai-rtgs';
49
+
50
+ // ตัวเลือก
51
+ romanize('สะอาด'); // 'sa-at'
52
+ romanize('สะอาด', { hyphenateAmbiguous: false }); // 'saat'
53
+ romanize('เด็กๆ เล่นกัน', { case: 'sentence' }); // 'Dek dek len kan'
54
+ romanize('วันที่ ๑๒ มกราคม ๒๕๖๗'); // 'wan thi 12 mokkarakhom 2567'
55
+ romanize('ฉันชอบ iPhone 15', { keepNonThai: false }); // 'chan chop'
56
+
57
+ // ชื่อคน: อ่านแต่ละส่วนที่คั่นด้วยเว้นวรรคเป็นหนึ่งคำ แยกคำนำหน้าออก และขึ้นต้นตัวใหญ่
58
+ romanizeName('นางสาวภัทรกมล สุขสวัสดิ์'); // 'Nangsao Phattharakamon Suksawat'
59
+
60
+ // คำเฉพาะสำหรับการเรียกครั้งเดียว (- คั่นพยางค์ เว้นวรรคคั่นคำ)
61
+ romanize('สมชาย ใจดี', { case: 'title', dictionary: { สมชาย: 'som-chai', ใจดี: 'chai-di' } }); // 'Somchai Chaidi'
62
+
63
+ // instance ที่มีพจนานุกรมและค่าตั้งต้นของตัวเอง
64
+ const rtgs = createRomanizer({ case: 'title' });
65
+ rtgs.addWord('ภัทรา', 'phat-tha-ra');
66
+ rtgs.romanize('ภัทรา'); // 'Phatthara'
67
+ rtgs.removeWord('ราชบุรี'); // ซ่อนคำในพจนานุกรมในตัว (เฉพาะ instance นี้)
68
+
69
+ // รายละเอียดทีละคำ/พยางค์
70
+ analyze('ถนนราชบุรี').tokens[0].words;
71
+ // [{ thai: 'ถนน', roman: 'thanon', source: 'rules', confidence: 'medium', syllables: [...] },
72
+ // { thai: 'ราชบุรี', roman: 'ratchaburi', source: 'dictionary', confidence: 'dictionary', ... }]
73
+
74
+ syllabify('เปลี่ยน');
75
+ // [{ thai: 'เปลี่ยน', roman: 'plian', initial: { thai: 'ปล', roman: 'pl' },
76
+ // vowel: { thai: 'เ-ีย', roman: 'ia', implicit: false }, final: { thai: 'น', roman: 'n' },
77
+ // silent: '', rules: ['true-cluster'] }]
78
+ ```
79
+
80
+ ### ตัวเลือก (`RomanizeOptions`)
81
+
82
+ | ตัวเลือก | ค่า | ค่าเริ่มต้น | ความหมาย |
83
+ |---|---|---|---|
84
+ | `case` | `lower` `upper` `title` `sentence` | `lower` | ตัวพิมพ์ `title` ขึ้นต้นตัวใหญ่ทุกคำ เหมาะกับชื่อเฉพาะ |
85
+ | `wordSeparator` | string | `' '` | ตัวคั่นระหว่างคำไทย |
86
+ | `syllableSeparator` | string | `''` | ตัวคั่นพยางค์ เช่น `'-'` |
87
+ | `hyphenateAmbiguous` | boolean | `true` | ใส่ `-` เมื่อพยางค์ต่อกันแล้วอ่านกำกวม (sa-at, sa-nga) |
88
+ | `thaiDigits` | `arabic` `keep` | `arabic` | เลขไทย ๐-๙ |
89
+ | `maiYamok` | `repeat` `keep` `drop` | `repeat` | ไม้ยมก `ๆ` |
90
+ | `paiyannoi` | `drop` `keep` | `drop` | ไปยาลน้อย `ฯ` (`ฯลฯ` → `etc.`) |
91
+ | `keepNonThai` | boolean | `true` | เก็บข้อความภาษาอื่น ตัวเลข และเครื่องหมาย |
92
+ | `useBuiltinDictionary` | boolean | `true` | ใช้พจนานุกรมในตัว |
93
+ | `dictionary` | `Record<string, string>` | — | คำเพิ่มเติมสำหรับการเรียกครั้งนั้น |
94
+ | `segmenter` | `intl` `dictionary` `none` หรือ function | `intl` | วิธีตัดคำ (ถ้าไม่มี `Intl.Segmenter` จะใช้ `dictionary`) |
95
+
96
+ ค่า `confidence` ของแต่ละคำ: `dictionary` (จากพจนานุกรม), `high` (สระเขียนรูปครบ), `medium` (มีสระที่ไม่เขียนรูป ควรตรวจถ้าเป็นชื่อเฉพาะ), `low` (มีอักษรที่แยกพยางค์ไม่ได้ และมี warning `unparsed`)
97
+
98
+ ## พัฒนา
99
+
100
+ ```bash
101
+ npm install
102
+ npm test # vitest
103
+ npm run typecheck
104
+ npm run eval # รายงานความแม่นยำ
105
+ npm run build # dist/ (ESM + CJS + .d.ts)
106
+ ```
107
+
108
+ รองรับ Node.js 18 ขึ้นไป (การพัฒนาใช้ Node 20 ขึ้นไป)
109
+
110
+ ## ข้อจำกัด
111
+
112
+ - RTGS ไม่มีวรรณยุกต์และความยาวสระ จึงแปลงกลับเป็นอักษรไทยไม่ได้
113
+ - คำบาลี-สันสกฤตที่มีพยางค์เชื่อม (ราชบุรี rat-cha-bu-ri, ผลไม้ phon-la-mai) ชื่อคน และคำทับศัพท์ ต้องพึ่งพจนานุกรม คำที่กฎอ่านผิดให้เพิ่มผ่าน `dictionary` หรือ `addWord`
114
+ - ชื่อคนที่อยู่ในประโยคอาจถูกตัดเป็นหลายคำ (`Intl.Segmenter` ไม่รู้จักชื่อ) ให้ใช้ `romanizeName()` กับชื่อ ซึ่งอ่านแต่ละส่วนเป็นหนึ่งคำ
115
+ - ผลการตัดคำของ `Intl.Segmenter` ขึ้นกับเวอร์ชัน ICU ของ runtime อาจต่างกันเล็กน้อย คำประสมที่ ICU ไม่แยก (เช่น รถไฟฟ้า) จะออกมาติดกัน (`rotfaifa`)
116
+ - ชื่อจังหวัดและเขตเขียนตามแบบราชบัณฑิตยสถาน (เช่น Chon Buri, Buri Ram) ซึ่งอาจต่างจากตัวสะกดที่ใช้กันทั่วไป (Chonburi, Buriram) แก้ได้ผ่าน `dictionary`
117
+
118
+ ---
119
+
120
+ ## English quickstart
121
+
122
+ `thai-rtgs` converts Thai script to Latin letters with the Royal Thai General System of Transcription. Words not in its dictionary go through a rule-based syllable parser, which finds the lowest-cost segmentation of Thai orthographic syllables. A dictionary covers the provinces, the Bangkok districts and words with irregular or Pali/Sanskrit readings. It has no dependencies and runs in Node and browsers.
123
+
124
+ ```ts
125
+ import { romanize } from 'thai-rtgs';
126
+ romanize('เชียงใหม่', { case: 'title' }); // 'Chiang Mai'
127
+ ```