@laisuk/opencc-fmmseg-wasm 0.3.8 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,731 +1,809 @@
1
- # opencc-fmmseg-wasm
2
-
3
- [![npm version](https://img.shields.io/npm/v/@laisuk/opencc-fmmseg-wasm)](https://www.npmjs.com/package/@laisuk/opencc-fmmseg-wasm)
4
- [![npm downloads](https://img.shields.io/npm/dm/@laisuk/opencc-fmmseg-wasm)](https://www.npmjs.com/package/@laisuk/opencc-fmmseg-wasm)
5
- [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
6
- [![WebAssembly](https://img.shields.io/badge/WebAssembly-enabled-blue)](https://webassembly.org/)
7
-
8
- OpenCC FMM segmentation WebAssembly bindings for browsers and JavaScript runtimes.
9
-
10
- This package provides high-quality Simplified Chinese ↔ Traditional Chinese conversion powered by the Rust [
11
- `opencc-fmmseg`](https://github.com/laisuk/opencc-fmmseg) engine.
12
-
13
- Features:
14
-
15
- * OpenCC-compatible conversion configs
16
- * Pure WebAssembly (no native binaries)
17
- * Browser-friendly
18
- * TypeScript-friendly APIs
19
- * Fast Rust backend
20
- * FMM-based phrase segmentation
21
- * Traditional Chinese regional variants
22
- * Japanese Shinjitai conversion support
23
- * Chinese script detection (`zho_check`)
24
- * Optional CJK Compatibility Ideograph normalization
25
- * In-memory Office / EPUB document conversion
26
- * Zero-dependency Node.js CLI
27
-
28
- Package profile:
29
-
30
- * 0 runtime dependencies
31
- * 1 WASM file
32
- * 18 conversion configs
33
- * 100% offline
34
-
35
- ---
36
-
37
- ## Installation
38
-
39
- ```bash
40
- npm install @laisuk/opencc-fmmseg-wasm
41
- ```
42
-
43
- ---
44
-
45
- ## Quick Start
46
-
47
- ```javascript
48
- import init, {
49
- OpenccWasm,
50
- DetofuLevelWasm
51
- } from "@laisuk/opencc-fmmseg-wasm";
52
-
53
- await init();
54
-
55
- const cc = new OpenccWasm("s2t");
56
-
57
- console.log(cc.convert("汉字", false));
58
- // 漢字
59
-
60
- console.log(cc.convertDetofu("儼驂騑於上路", false, DetofuLevelWasm.ExtB));
61
- // 俨骖騑于上路
62
- ```
63
-
64
- ---
65
-
66
- ## Using Config Enums
67
-
68
- ```javascript
69
- import init, {
70
- OpenccWasm,
71
- OpenccConfigWasm
72
- } from "@laisuk/opencc-fmmseg-wasm";
73
-
74
- await init();
75
-
76
- const cc = OpenccWasm.newWithEnum(
77
- OpenccConfigWasm.S2hkp
78
- );
79
-
80
- console.log(cc.convert("别随便录影侵犯个人隐私权", false));
81
- // 別隨便錄影侵犯個人私隱權
82
- ```
83
-
84
- ---
85
-
86
- ## Supported Configs
87
-
88
- | Config | Enum | Description |
89
- |---------|--------------------------|------------------------------------------------------|
90
- | `s2t` | `OpenccConfigWasm.S2t` | Simplified Chinese → Traditional Chinese |
91
- | `s2tw` | `OpenccConfigWasm.S2tw` | Simplified Chinese → Taiwan Traditional |
92
- | `s2twp` | `OpenccConfigWasm.S2twp` | Simplified Chinese → Taiwan Traditional (phrases) |
93
- | `s2hk` | `OpenccConfigWasm.S2hk` | Simplified Chinese → Hong Kong Traditional |
94
- | `s2hkp` | `OpenccConfigWasm.S2hkp` | Simplified Chinese → Hong Kong Traditional (phrases) |
95
- | `t2s` | `OpenccConfigWasm.T2s` | Traditional Chinese → Simplified Chinese |
96
- | `t2tw` | `OpenccConfigWasm.T2tw` | Traditional Chinese → Taiwan Traditional |
97
- | `t2twp` | `OpenccConfigWasm.T2twp` | Traditional Chinese → Taiwan Traditional (phrases) |
98
- | `t2hk` | `OpenccConfigWasm.T2hk` | Traditional Chinese → Hong Kong Traditional |
99
- | `t2hkp` | `OpenccConfigWasm.T2hkp` | Traditional Chinese → Hong Kong Traditional (phrases)|
100
- | `tw2s` | `OpenccConfigWasm.Tw2s` | Taiwan Traditional → Simplified Chinese |
101
- | `tw2sp` | `OpenccConfigWasm.Tw2sp` | Taiwan Traditional → Simplified Chinese (phrases) |
102
- | `tw2t` | `OpenccConfigWasm.Tw2t` | Taiwan Traditional → Traditional Chinese |
103
- | `tw2tp` | `OpenccConfigWasm.Tw2tp` | Taiwan Traditional → Traditional Chinese (phrases) |
104
- | `hk2s` | `OpenccConfigWasm.Hk2s` | Hong Kong Traditional → Simplified Chinese |
105
- | `hk2sp` | `OpenccConfigWasm.Hk2sp` | Hong Kong Traditional → Simplified Chinese (phrases) |
106
- | `hk2t` | `OpenccConfigWasm.Hk2t` | Hong Kong Traditional → Traditional Chinese |
107
- | `hk2tp` | `OpenccConfigWasm.Hk2tp` | Hong Kong Traditional → Traditional Chinese (phrases)|
108
- | `jp2t` | `OpenccConfigWasm.Jp2t` | Japanese Shinjitai → Traditional Chinese |
109
- | `t2jp` | `OpenccConfigWasm.T2jp` | Traditional Chinese → Japanese Shinjitai |
110
-
111
- The numeric enum values match the vendored Rust backend. Existing values are unchanged; `S2hkp = 17`, `Hk2sp = 18`, `T2hkp = 19`, and `Hk2tp = 20`.
112
-
113
- ---
114
-
115
- ## API
116
-
117
- ### Constructor
118
-
119
- ```javascript
120
- const cc = new OpenccWasm("s2t");
121
- ```
122
-
123
- Parameters:
124
-
125
- * `config` (optional): OpenCC config string
126
- * default: `"s2t"`
127
-
128
- Example:
129
-
130
- ```javascript
131
- const cc = new OpenccWasm("t2s");
132
- ```
133
-
134
- Hong Kong phrase config example:
135
-
136
- ```javascript
137
- const cc = new OpenccWasm("s2hkp");
138
-
139
- cc.convert("别随便录影侵犯个人隐私权", false);
140
- // 別隨便錄影侵犯個人私隱權
141
- ```
142
-
143
- ---
144
-
145
- ### convert
146
-
147
- ```javascript
148
- cc.convert(text, punctuation)
149
- ```
150
-
151
- Parameters:
152
-
153
- * `text`: input string
154
- * `punctuation`: whether to convert punctuation variants
155
-
156
- Returns:
157
-
158
- * converted string
159
-
160
- Example:
161
-
162
- ```javascript
163
- cc.convert("汉字", false);
164
- ```
165
-
166
- ---
167
-
168
- ### normalizeCompat
169
-
170
- Normalize Unicode CJK Compatibility Ideographs before conversion.
171
-
172
- ```javascript
173
- cc.normalizeCompat(text)
174
- ```
175
-
176
- Parameters:
177
-
178
- * `text`: input string
179
-
180
- Returns:
181
-
182
- * normalized string
183
-
184
- Example:
185
-
186
- ```javascript
187
- const cc = new OpenccWasm("t2s");
188
-
189
- const input = "天龍八部書裡的喬峰是契丹人";
190
- const normalized = cc.normalizeCompat(input);
191
-
192
- console.log(normalized);
193
- // 天龍八部書裡的喬峰是契丹人
194
-
195
- console.log(cc.convert(normalized, false));
196
- // 天龙八部书里的乔峰是契丹人
197
- ```
198
-
199
- This is an optional pre-conversion pass for text that contains compatibility ideographs from Unicode compatibility ranges. Unmapped characters are preserved unchanged. Normal OpenCC conversion does not automatically run this pass, so call it explicitly when compatibility normalization is desired.
200
-
201
- ---
202
-
203
- ### detofu
204
-
205
- Replace tofu-risk rare CJK extension characters with display-compatible fallbacks.
206
-
207
- ```javascript
208
- cc.detofu(text, level)
209
- ```
210
-
211
- Parameters:
212
-
213
- * `text`: input string
214
- * `level`: `DetofuLevelWasm` threshold for the CJK extension ranges to replace
215
-
216
- Returns:
217
-
218
- * detofu-safe string
219
-
220
- Supported levels:
221
-
222
- | Enum | CLI value |
223
- |------------------------|-----------|
224
- | `DetofuLevelWasm.ExtB` | `ext-b` |
225
- | `DetofuLevelWasm.ExtC` | `ext-c` |
226
- | `DetofuLevelWasm.ExtD` | `ext-d` |
227
- | `DetofuLevelWasm.ExtE` | `ext-e` |
228
- | `DetofuLevelWasm.ExtF` | `ext-f` |
229
- | `DetofuLevelWasm.ExtG` | `ext-g` |
230
- | `DetofuLevelWasm.ExtH` | `ext-h` |
231
- | `DetofuLevelWasm.ExtI` | `ext-i` |
232
-
233
- Example:
234
-
235
- ```javascript
236
- import init, {
237
- OpenccWasm,
238
- DetofuLevelWasm
239
- } from "@laisuk/opencc-fmmseg-wasm";
240
-
241
- await init();
242
-
243
- const cc = new OpenccWasm("t2s");
244
- const converted = cc.convert("儼驂騑於上路", false);
245
-
246
- console.log(converted);
247
- // 俨骖𬴂于上路
248
-
249
- console.log(cc.detofu(converted, DetofuLevelWasm.ExtB));
250
- // 俨骖騑于上路
251
- ```
252
-
253
- ---
254
-
255
- ### convertDetofu
256
-
257
- Convert text and apply detofu in one call.
258
-
259
- ```javascript
260
- cc.convertDetofu(text, punctuation, level)
261
- ```
262
-
263
- Parameters:
264
-
265
- * `text`: input string
266
- * `punctuation`: whether to convert punctuation variants
267
- * `level`: `DetofuLevelWasm` threshold for the CJK extension ranges to replace
268
-
269
- Returns:
270
-
271
- * converted detofu-safe string
272
-
273
- Example:
274
-
275
- ```javascript
276
- cc.convertDetofu("儼驂騑於上路", false, DetofuLevelWasm.ExtB);
277
- // 俨骖騑于上路
278
- ```
279
-
280
- ---
281
-
282
- ### setConfig
283
-
284
- ```javascript
285
- cc.setConfig("t2s");
286
- ```
287
-
288
- Returns:
289
-
290
- * `true` if valid
291
- * `false` if invalid
292
-
293
- ---
294
-
295
- ### getConfig
296
-
297
- ```javascript
298
- cc.getConfig();
299
- ```
300
-
301
- Returns current config string.
302
-
303
- ---
304
-
305
- ### isValidConfig
306
-
307
- ```javascript
308
- OpenccWasm.isValidConfig("s2t");
309
- ```
310
-
311
- ---
312
-
313
- ### getSupportedConfigs
314
-
315
- ```javascript
316
- OpenccWasm.getSupportedConfigs();
317
- ```
318
-
319
- Returns all supported config strings.
320
-
321
- Includes `s2hkp`, `hk2sp`, `t2hkp`, and `hk2tp`.
322
-
323
- ---
324
-
325
- ### zhoCheck
326
-
327
- Detect Chinese script type.
328
-
329
- ```javascript
330
- cc.zhoCheck(text);
331
- ```
332
-
333
- Returns:
334
-
335
- | Value | Meaning |
336
- |-------|---------------------|
337
- | `0` | Unknown / mixed |
338
- | `1` | Traditional Chinese |
339
- | `2` | Simplified Chinese |
340
-
341
- ---
342
-
343
- ### newWithCustomDicts
344
-
345
- Construct a converter with in-memory custom dictionary pairs.
346
-
347
- ```javascript
348
- const cc = OpenccWasm.newWithCustomDicts(config, specs);
349
- ```
350
-
351
- Parameters:
352
-
353
- * `config`: OpenCC config string, such as `"s2t"`
354
- * `specs`: array of custom dictionary specs
355
-
356
- TypeScript-style spec shape:
357
-
358
- ```typescript
359
- type WasmCustomDictSpec = {
360
- slot: string;
361
- mode?: "Append" | "Override";
362
- pairs: Array<[string, string]>;
363
- };
364
- ```
365
-
366
- `mode` defaults to `"Append"` when omitted.
367
-
368
- Each `pairs` entry is a `[source, target]` string tuple for the selected slot.
369
-
370
- TypeScript example:
371
-
372
- ```typescript
373
- import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
374
-
375
- await init();
376
-
377
- const specs: WasmCustomDictSpec[] = [
378
- {
379
- slot: "STPhrases",
380
- pairs: [
381
- ["云端", "雲端"]
382
- ]
383
- }
384
- ];
385
-
386
- const cc = OpenccWasm.newWithCustomDicts("s2t", specs);
387
- ```
388
-
389
- Practical example:
390
-
391
- ```javascript
392
- import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
393
-
394
- await init();
395
-
396
- const cc = OpenccWasm.newWithCustomDicts("s2t", [
397
- {
398
- slot: "STPhrases",
399
- mode: "Append",
400
- pairs: [
401
- ["帕兰蒂尔", "柏蘭蒂爾"],
402
- ["软件", "軟體"]
403
- ]
404
- }
405
- ]);
406
-
407
- console.log(cc.convert("帕兰蒂尔软件", false));
408
- // 柏蘭蒂爾軟體
409
- ```
410
-
411
- Override example:
412
-
413
- ```javascript
414
- const cc = OpenccWasm.newWithCustomDicts("s2t", [
415
- {
416
- slot: "STPhrases",
417
- mode: "Override",
418
- pairs: [
419
- ["软件", "軟體"]
420
- ]
421
- }
422
- ]);
423
- ```
424
-
425
- `Override` replaces the selected slot before inserting the provided pairs. It is powerful and should be used only when
426
- the caller intentionally wants to discard built-in entries for that slot.
427
-
428
- Custom dictionary specs identify the target dictionary slot by `DictSlot` name. Slot names are trimmed and normalized
429
- case-insensitively for the known slots, so `"stphrases"`, `" STPhrases "`, and `"STPhrases"` all select
430
- `STPhrases`. Canonical names are recommended in TypeScript code and docs:
431
-
432
- ```text
433
- STPhrases
434
- TSPhrases
435
- STCharacters
436
- TSCharacters
437
- TWPhrases
438
- TWPhrasesRev
439
- HKPhrases
440
- HKPhrasesRev
441
- TWVariants
442
- TWVariantsPhrases
443
- TWVariantsRev
444
- TWVariantsRevPhrases
445
- HKVariants
446
- HKVariantsPhrases
447
- HKVariantsRev
448
- HKVariantsRevPhrases
449
- JPSCharacters
450
- JPSCharactersRev
451
- JPSPhrases
452
- STPunctuations
453
- TSPunctuations
454
- ```
455
-
456
- Suffixes such as `.txt` are not accepted, even though case and surrounding whitespace are normalized. Use
457
- `"STPhrases"` or `"stphrases"`, not `"STPhrases.txt"`.
458
-
459
- Merge contract:
460
-
461
- * Custom dictionaries are loaded from in-memory pairs only; no file I/O is involved.
462
- * The embedded compressed CBOR dictionary is loaded first.
463
- * Custom specs are applied to `DictionaryMaxlength` before `OpenCC::from_dictionary(...)`.
464
- * Conversion hot paths remain immutable after construction.
465
- * `Append` mode merges into the selected slot.
466
- * Duplicate or conflicting keys use last-wins semantics.
467
- * `Override` mode clears the selected slot first, then inserts the provided custom pairs.
468
- * Multiple specs are applied in array order.
469
-
470
- This API is useful for browser apps, user-defined terminology, database-loaded terms, generated dictionaries,
471
- `localStorage` or `IndexedDB` terms, testing, and embedded WASM environments. Customization happens at construction
472
- time, not during conversion.
473
-
474
- ---
475
-
476
- ## Office / EPUB Conversion
477
-
478
- Office and EPUB conversion runs fully locally in the browser or Node.js. Files are passed in and returned as bytes;
479
- nothing is uploaded to a backend server.
480
-
481
- This is useful for converting text inside:
482
-
483
- ```text
484
- docx, xlsx, pptx, odt, ods, odp, epub
485
- ```
486
-
487
- File size is limited by available browser or Node.js memory, but there is no upload or server-side limit. Font
488
- preservation is supported with the `keepFont` option.
489
-
490
- Use the instance method when possible. It reuses the converter configuration and any custom dictionaries already held by
491
- the `OpenccWasm` instance.
492
-
493
- ```javascript
494
- cc.convertOfficeBytes(inputBytes, format, punctuation, keepFont)
495
- ```
496
-
497
- Parameters:
498
-
499
- * `inputBytes`: `Uint8Array` document bytes
500
- * `format`: `docx`, `xlsx`, `pptx`, `odt`, `ods`, `odp`, or `epub`
501
- * `punctuation`: whether to convert punctuation variants
502
- * `keepFont`: whether to preserve font declarations where supported
503
-
504
- Returns:
505
-
506
- * converted output bytes
507
-
508
- The older free function remains available for compatibility:
509
-
510
- ```javascript
511
- convert_office_bytes(inputBytes, format, config, punctuation, keepFont)
512
- ```
513
-
514
- ### Browser Office Example
515
-
516
- ```javascript
517
- import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
518
-
519
- await init();
520
-
521
- const cc = new OpenccWasm("s2t");
522
- const file = document.querySelector("input[type=file]").files[0];
523
- const inputBytes = new Uint8Array(await file.arrayBuffer());
524
-
525
- const outputBytes = cc.convertOfficeBytes(
526
- inputBytes,
527
- "docx",
528
- true,
529
- true
530
- );
531
-
532
- const blob = new Blob([outputBytes], {
533
- type: "application/vnd.openxmlformats-officedocument.wordprocessingml.document"
534
- });
535
-
536
- const a = document.createElement("a");
537
- a.href = URL.createObjectURL(blob);
538
- a.download = "converted.docx";
539
- a.click();
540
- URL.revokeObjectURL(a.href);
541
- ```
542
-
543
- ### Node.js Office Example
544
-
545
- ```javascript
546
- import fs from "fs";
547
- import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
548
-
549
- await init();
550
-
551
- const cc = new OpenccWasm("s2t");
552
- const inputBytes = fs.readFileSync("input.docx");
553
-
554
- const outputBytes = cc.convertOfficeBytes(
555
- inputBytes,
556
- "docx",
557
- true,
558
- true
559
- );
560
-
561
- fs.writeFileSync("output.docx", outputBytes);
562
- ```
563
-
564
- ---
565
-
566
- ## Browser Example
567
-
568
- ```html
569
- <!DOCTYPE html>
570
- <html lang="en">
571
- <head>
572
- <meta charset="UTF-8">
573
- <title>OpenCC WASM Demo</title>
574
- </head>
575
- <body>
576
-
577
- <script type="module">
578
- import init, {
579
- OpenccWasm
580
- } from "./pkg/opencc_fmmseg_wasm.js";
581
-
582
- await init();
583
-
584
- const cc = new OpenccWasm("s2t");
585
-
586
- console.log(
587
- cc.convert("汉字", false)
588
- );
589
- </script>
590
-
591
- </body>
592
- </html>
593
- ```
594
-
595
- > **Note**
596
- >
597
- > Normally, `await init();` is sufficient when using the published npm package.
598
- >
599
- > When running directly from a local repository checkout (for example in tests
600
- > or development scripts), initialize using explicit WASM bytes:
601
- >
602
- > ```javascript
603
- > import fs from "fs";
604
- > import init from "../pkg/opencc_fmmseg_wasm.js";
605
- >
606
- > const wasmBytes = fs.readFileSync(
607
- > "../pkg/opencc_fmmseg_wasm_bg.wasm"
608
- > );
609
- >
610
- > await init({
611
- > module_or_path: wasmBytes
612
- > });
613
- > ```
614
-
615
- ---
616
-
617
- ## Node.js CLI
618
-
619
- The package includes a zero-dependency Node.js CLI:
620
-
621
- ```bash
622
- opencc-fmmseg convert -i input.txt -o output.txt -c s2t -p
623
- opencc-fmmseg convert -i input.txt -o output.txt -c t2s -p --detofu all
624
- echo "别随便录影侵犯个人隐私权" | opencc-fmmseg convert -c s2hkp
625
- echo "天龍八部書裡的喬峰是契丹人" | opencc-fmmseg convert -c t2s --norm-compat
626
- // 天龙八部书里的乔峰是契丹人
627
- echo "這個細路哥很靈活" | opencc-fmmseg convert -c hk2sp --custom-dict hkphrasesrev:append:my_hk_dict.txt
628
- // 这个小男孩很灵活
629
- ```
630
-
631
- my_hk_dict.txt:
632
-
633
- ```
634
- # Custom Dictionary
635
-
636
- 細路哥 小男孩
637
- ```
638
-
639
- ```bash
640
- opencc-fmmseg office -i input.docx -o output.docx -c s2t -p --keep-font
641
- ```
642
-
643
- ### Text Conversion Options
644
-
645
- ```text
646
- -i, --input <file> Input text file; stdin if omitted
647
- -o, --output <file> Output text file; stdout if omitted
648
- -c, --config <conversion> Conversion config (default: s2t)
649
- -p, --punct Enable punctuation conversion
650
- --detofu [level] Replace tofu-risk rare CJK extension chars after conversion
651
- level: all | ext-b | ext-c | ext-d | ext-e | ext-f | ext-g | ext-h | ext-i
652
- default when omitted value: all
653
- --keep-ids Preserve complete IDS expressions during conversion (default: false)
654
- -n, --norm-compat Normalize CJK Compatibility Ideographs before conversion (default: false)
655
- -D, --custom-dict <slot:mode:file>
656
- Load a custom dictionary.
657
- May be specified multiple times.
658
- Examples:
659
- --custom-dict hkphrasesrev:append:my_hk_dict.txt
660
- --custom-dict stphrases:override:terms.txt
661
- --in-enc <encoding> Input encoding (default: utf8)
662
- --out-enc <encoding> Output encoding (default: utf8)
663
- ```
664
-
665
- Supported conversion configs:
666
-
667
- ```text
668
- s2t, s2tw, s2twp, s2hk, s2hkp, t2s, t2tw, t2twp, t2hk, t2hkp,
669
- tw2s, tw2sp, tw2t, tw2tp, hk2s, hk2sp, hk2t, hk2tp, jp2t, t2jp
670
- ```
671
-
672
- ### Office / EPUB Options
673
-
674
- ```text
675
- -i, --input <file> Input Office / EPUB file
676
- -o, --output <file> Output file
677
- -c, --config <conversion> Conversion config (default: s2t)
678
- -p, --punct Enable punctuation conversion
679
- -f, --format <format> docx | xlsx | pptx | odt | ods | odp | epub
680
- -F, --convert-filename Convert generated output filename stem (default: false)
681
- --keep-font Preserve font-family information (default)
682
- --no-keep-font Do not preserve font-family information
683
- --custom-dict <slot:mode:file>
684
- Load a custom dictionary.
685
- May be specified multiple times.
686
- Examples:
687
- --custom-dict hkphrasesrev:append:my_hk_dict.txt
688
- --custom-dict stphrases:override:terms.txt
689
- ```
690
-
691
- For `office`, the format is inferred from the input file extension when `--format` is omitted.
692
-
693
- If `-o, --output` is omitted, `office` writes:
694
-
695
- ```text
696
- <input-name>_converted.<ext>
697
- ```
698
-
699
- ---
700
-
701
- ## TypeScript Support
702
-
703
- The package includes generated TypeScript definitions from `wasm-bindgen`.
704
-
705
- The WASM-facing enum is exported as `OpenccConfigWasm`, alongside `OpenccWasm`.
706
-
707
- `OpenccConfigWasm.S2hkp`, `OpenccConfigWasm.Hk2sp`, `OpenccConfigWasm.T2hkp`, and
708
- `OpenccConfigWasm.Hk2tp` are available for Hong Kong phrase conversions and map to backend config IDs `17` through `20`.
709
-
710
- ---
711
-
712
- ## Performance Notes
713
-
714
- * WebAssembly build disables Rayon parallelism by default.
715
- * Dictionaries are embedded into the WASM binary.
716
- * Browser caching significantly improves subsequent loads.
717
-
718
- ---
719
-
720
- ## Related Projects
721
-
722
- * Rust backend: https://github.com/laisuk/opencc-fmmseg
723
- * C API: https://github.com/laisuk/opencc-fmmseg/tree/master/capi/opencc-fmmseg-capi
724
- * .NET: https://github.com/laisuk/OpenccNet
725
- * Python: https://github.com/laisuk/opencc_purepy
726
-
727
- ---
728
-
729
- ## License
730
-
731
- MIT
1
+ # opencc-fmmseg-wasm
2
+
3
+ [![npm version](https://img.shields.io/npm/v/@laisuk/opencc-fmmseg-wasm)](https://www.npmjs.com/package/@laisuk/opencc-fmmseg-wasm)
4
+ [![npm downloads](https://img.shields.io/npm/dm/@laisuk/opencc-fmmseg-wasm)](https://www.npmjs.com/package/@laisuk/opencc-fmmseg-wasm)
5
+ [![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
6
+ [![WebAssembly](https://img.shields.io/badge/WebAssembly-enabled-blue)](https://webassembly.org/)
7
+
8
+ OpenCC FMM segmentation WebAssembly bindings for browsers and JavaScript runtimes.
9
+
10
+ This package provides high-quality Simplified Chinese ↔ Traditional Chinese conversion powered by the Rust [
11
+ `opencc-fmmseg`](https://github.com/laisuk/opencc-fmmseg) engine.
12
+
13
+ Features:
14
+
15
+ - OpenCC-compatible conversion configs
16
+ - Pure WebAssembly (no native binaries)
17
+ - Browser-friendly
18
+ - TypeScript-friendly APIs
19
+ - Fast Rust backend
20
+ - FMM-based phrase segmentation
21
+ - Traditional Chinese regional variants
22
+ - Japanese Shinjitai conversion support
23
+ - Chinese script detection (`zho_check`)
24
+ - Optional CJK Compatibility Ideograph normalization
25
+ - In-memory Office / EPUB document conversion
26
+ - Zero-dependency Node.js CLI
27
+
28
+ Package profile:
29
+
30
+ - 0 runtime dependencies
31
+ - 1 WASM file
32
+ - 20 conversion configs
33
+ - 100% offline
34
+
35
+ ---
36
+
37
+ ## Installation
38
+
39
+ ```bash
40
+ npm install @laisuk/opencc-fmmseg-wasm
41
+ ```
42
+
43
+ ---
44
+
45
+ ## Quick Start
46
+
47
+ ```javascript
48
+ import init, {
49
+ OpenccWasm,
50
+ DetofuLevelWasm
51
+ } from "@laisuk/opencc-fmmseg-wasm";
52
+
53
+ await init();
54
+
55
+ const cc = new OpenccWasm("t2s");
56
+
57
+ console.log(cc.convert("漢字", false));
58
+ // 汉字
59
+
60
+ console.log(cc.convertDetofu("儼驂騑於上路", false, DetofuLevelWasm.ExtB));
61
+ // 俨骖騑于上路
62
+ ```
63
+
64
+ ---
65
+
66
+ ## API
67
+
68
+ ### Constructor
69
+
70
+ ```javascript
71
+ const cc = new OpenccWasm("s2t");
72
+ ```
73
+
74
+ Parameters:
75
+
76
+ - `config` (optional): OpenCC config string
77
+ - default: `"s2t"`
78
+
79
+ Example:
80
+
81
+ ```javascript
82
+ const cc = new OpenccWasm("t2s");
83
+ ```
84
+
85
+ Hong Kong phrase config example:
86
+
87
+ ```javascript
88
+ const cc = new OpenccWasm("s2hkp");
89
+
90
+ cc.convert("别随便录影侵犯个人隐私权", false);
91
+ // 別隨便錄影侵犯個人私隱權
92
+ ```
93
+
94
+ ---
95
+
96
+ ### convert
97
+
98
+ ```javascript
99
+ cc.convert(text, punctuation)
100
+ ```
101
+
102
+ Parameters:
103
+
104
+ - `text`: input string
105
+ - `punctuation`: whether to convert punctuation variants
106
+
107
+ Returns:
108
+
109
+ - converted string
110
+
111
+ Example:
112
+
113
+ ```javascript
114
+ cc.convert("汉字", false);
115
+ ```
116
+
117
+ ---
118
+
119
+ ### setConfig
120
+
121
+ ```javascript
122
+ cc.setConfig("t2s");
123
+ ```
124
+
125
+ Returns:
126
+
127
+ - `true` if valid
128
+ - `false` if invalid
129
+
130
+ ---
131
+
132
+ ### getConfig
133
+
134
+ ```javascript
135
+ cc.getConfig();
136
+ ```
137
+
138
+ Returns current config string.
139
+
140
+ ---
141
+
142
+ ### isValidConfig
143
+
144
+ ```javascript
145
+ OpenccWasm.isValidConfig("s2t");
146
+ ```
147
+
148
+ ---
149
+
150
+ ### getSupportedConfigs
151
+
152
+ ```javascript
153
+ OpenccWasm.getSupportedConfigs();
154
+ ```
155
+
156
+ Returns all supported config strings.
157
+
158
+ Includes `s2hkp`, `hk2sp`, `t2hkp`, and `hk2tp`.
159
+
160
+ ---
161
+
162
+ ### getAvailableSlots
163
+
164
+ ```javascript
165
+ const slots = OpenccWasm.getAvailableSlots();
166
+ ```
167
+
168
+ Returns all canonical dictionary slot names accepted by `newWithCustomDicts` as a string array. The list is sourced from
169
+ the core `DictSlot` definitions, so callers can use it to populate selectors or validate custom dictionary input without
170
+ maintaining their own slot list.
171
+
172
+ ```javascript
173
+ if (!OpenccWasm.getAvailableSlots().includes(slot)) {
174
+ throw new Error(`Unsupported dictionary slot: ${slot}`);
175
+ }
176
+ ```
177
+
178
+ ---
179
+
180
+ ### zhoCheck
181
+
182
+ Detect Chinese script type.
183
+
184
+ ```javascript
185
+ cc.zhoCheck(text);
186
+ ```
187
+
188
+ Returns:
189
+
190
+ | Value | Meaning |
191
+ |-------|---------------------|
192
+ | `0` | Unknown / mixed |
193
+ | `1` | Traditional Chinese |
194
+ | `2` | Simplified Chinese |
195
+
196
+ ---
197
+
198
+ ### normalizeCompat
199
+
200
+ Normalize Unicode CJK Compatibility Ideographs before conversion.
201
+
202
+ ```javascript
203
+ cc.normalizeCompat(text)
204
+ ```
205
+
206
+ Parameters:
207
+
208
+ - `text`: input string
209
+
210
+ Returns:
211
+
212
+ - normalized string
213
+
214
+ Example:
215
+
216
+ ```javascript
217
+ const cc = new OpenccWasm("t2s");
218
+
219
+ const input = "天龍八部書裡的喬峰是契丹人";
220
+ const normalized = cc.normalizeCompat(input);
221
+
222
+ console.log(normalized);
223
+ // 天龍八部書裡的喬峰是契丹人
224
+
225
+ console.log(cc.convert(normalized, false));
226
+ // 天龙八部书里的乔峰是契丹人
227
+ ```
228
+
229
+ This is an optional pre-conversion pass for text that contains CJK Compatibility Ideographs. Unmapped characters are
230
+ preserved unchanged. Normal OpenCC conversion does not automatically run this pass, so call it explicitly when
231
+ compatibility normalization is desired.
232
+
233
+ ### normalizeUnicodeCompat
234
+
235
+ Normalize additional Unicode compatibility forms, CJK radicals, allographs, legacy glyphs, and selected
236
+ compatibility-like punctuation before conversion.
237
+
238
+ ```javascript
239
+ cc.normalizeUnicodeCompat(text)
240
+ ```
241
+
242
+ Parameters:
243
+
244
+ - `text`: input string
245
+
246
+ Returns:
247
+
248
+ - normalized string
249
+
250
+ Example:
251
+
252
+ ```javascript
253
+ const cc = new OpenccWasm("t2s");
254
+
255
+ const input = "聼聼竒羙⽟䂖甁噐⾳";
256
+ const normalized = cc.normalizeUnicodeCompat(input);
257
+
258
+ console.log(normalized);
259
+ // 聽聽奇美玉石瓶器音
260
+
261
+ console.log(cc.convert(normalized, false));
262
+ // 听听奇美玉石瓶器音
263
+ ```
264
+
265
+ This pass uses the extended Unicode compatibility table and is separate from `normalizeCompat()`. It is useful for text
266
+ containing radical forms, historical or allographic Han forms, and other compatibility-like characters that are not
267
+ covered by the CJK Compatibility Ideograph ranges.
268
+
269
+ Unmapped characters are preserved unchanged.
270
+
271
+ ### normalizeCompatExtended
272
+
273
+ Apply complete compatibility normalization before conversion.
274
+
275
+ ```javascript
276
+ cc.normalizeCompatExtended(text)
277
+ ```
278
+
279
+ Parameters:
280
+
281
+ - `text`: input string
282
+
283
+ Returns:
284
+
285
+ - normalized string
286
+
287
+ Example:
288
+
289
+ ```javascript
290
+ const cc = new OpenccWasm("t2s");
291
+
292
+ const input = "天龍八部書裡的聼眾";
293
+ const normalized = cc.normalizeCompatExtended(input);
294
+
295
+ console.log(normalized);
296
+ // 天龍八部書裡的聽眾
297
+
298
+ console.log(cc.convert(normalized, false));
299
+ // 天龙八部书里的听众
300
+ ```
301
+
302
+ `normalizeCompatExtended()` combines the extended Unicode compatibility table with CJK Compatibility Ideograph
303
+ normalization. Use this when input may contain characters handled by either normalization set.
304
+
305
+ The normalization order is:
306
+
307
+ 1. extended Unicode compatibility normalization;
308
+ 2. CJK Compatibility Ideograph normalization.
309
+
310
+ Normal OpenCC conversion does not automatically perform compatibility normalization. Call this method explicitly before
311
+ `convert()` when complete compatibility normalization is desired.
312
+
313
+ ---
314
+
315
+ ### detofu
316
+
317
+ Replace tofu-risk rare CJK extension characters with display-compatible fallbacks.
318
+
319
+ ```javascript
320
+ cc.detofu(text, level)
321
+ ```
322
+
323
+ Parameters:
324
+
325
+ - `text`: input string
326
+ - `level`: `DetofuLevelWasm` threshold for the CJK extension ranges to replace
327
+
328
+ Returns:
329
+
330
+ - detofu-safe string
331
+
332
+ Supported levels:
333
+
334
+ | Enum | CLI value |
335
+ |------------------------|-----------|
336
+ | `DetofuLevelWasm.ExtB` | `ext-b` |
337
+ | `DetofuLevelWasm.ExtC` | `ext-c` |
338
+ | `DetofuLevelWasm.ExtD` | `ext-d` |
339
+ | `DetofuLevelWasm.ExtE` | `ext-e` |
340
+ | `DetofuLevelWasm.ExtF` | `ext-f` |
341
+ | `DetofuLevelWasm.ExtG` | `ext-g` |
342
+ | `DetofuLevelWasm.ExtH` | `ext-h` |
343
+ | `DetofuLevelWasm.ExtI` | `ext-i` |
344
+
345
+ Example:
346
+
347
+ ```javascript
348
+ import init, {
349
+ OpenccWasm,
350
+ DetofuLevelWasm
351
+ } from "@laisuk/opencc-fmmseg-wasm";
352
+
353
+ await init();
354
+
355
+ const cc = new OpenccWasm("t2s");
356
+ const converted = cc.convert("儼驂騑於上路", false);
357
+
358
+ console.log(converted);
359
+ // 俨骖𬴂于上路
360
+
361
+ console.log(cc.detofu(converted, DetofuLevelWasm.ExtB));
362
+ // 俨骖騑于上路
363
+ ```
364
+
365
+ ---
366
+
367
+ ### convertDetofu
368
+
369
+ Convert text and apply detofu in one call.
370
+
371
+ ```javascript
372
+ cc.convertDetofu(text, punctuation, level)
373
+ ```
374
+
375
+ Parameters:
376
+
377
+ - `text`: input string
378
+ - `punctuation`: whether to convert punctuation variants
379
+ - `level`: `DetofuLevelWasm` threshold for the CJK extension ranges to replace
380
+
381
+ Returns:
382
+
383
+ - converted detofu-safe string
384
+
385
+ Example:
386
+
387
+ ```javascript
388
+ cc.convertDetofu("儼驂騑於上路", false, DetofuLevelWasm.ExtB);
389
+ // 俨骖騑于上路
390
+ ```
391
+
392
+ ---
393
+
394
+ ### newWithCustomDicts
395
+
396
+ Construct a converter with in-memory custom dictionary pairs.
397
+
398
+ ```javascript
399
+ const cc = OpenccWasm.newWithCustomDicts(config, specs);
400
+ ```
401
+
402
+ Parameters:
403
+
404
+ - `config`: OpenCC config string, such as `"s2t"`
405
+ - `specs`: array of custom dictionary specs
406
+
407
+ TypeScript-style spec shape:
408
+
409
+ ```typescript
410
+ type WasmCustomDictSpec = {
411
+ slot: string;
412
+ mode?: "Append" | "Override";
413
+ pairs: Array<[string, string]>;
414
+ };
415
+ ```
416
+
417
+ `mode` defaults to `"Append"` when omitted.
418
+
419
+ Each `pairs` entry is a `[source, target]` string tuple for the selected slot.
420
+
421
+ TypeScript example:
422
+
423
+ ```typescript
424
+ import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
425
+
426
+ await init();
427
+
428
+ const specs: WasmCustomDictSpec[] = [
429
+ {
430
+ slot: "STPhrases",
431
+ pairs: [
432
+ ["云端", "雲端"]
433
+ ]
434
+ }
435
+ ];
436
+
437
+ const cc = OpenccWasm.newWithCustomDicts("s2t", specs);
438
+ ```
439
+
440
+ Practical example:
441
+
442
+ ```javascript
443
+ import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
444
+
445
+ await init();
446
+
447
+ const cc = OpenccWasm.newWithCustomDicts("s2t", [
448
+ {
449
+ slot: "STPhrases",
450
+ mode: "Append",
451
+ pairs: [
452
+ ["帕兰蒂尔", "柏蘭蒂爾"],
453
+ ["软件", "軟體"]
454
+ ]
455
+ }
456
+ ]);
457
+
458
+ console.log(cc.convert("帕兰蒂尔软件", false));
459
+ // 柏蘭蒂爾軟體
460
+ ```
461
+
462
+ Override example:
463
+
464
+ ```javascript
465
+ const cc = OpenccWasm.newWithCustomDicts("s2t", [
466
+ {
467
+ slot: "STPhrases",
468
+ mode: "Override",
469
+ pairs: [
470
+ ["软件", "軟體"]
471
+ ]
472
+ }
473
+ ]);
474
+ ```
475
+
476
+ `Override` replaces the selected slot before inserting the provided pairs. It is powerful and should be used only when
477
+ the caller intentionally wants to discard built-in entries for that slot.
478
+
479
+ Custom dictionary specs identify the target dictionary slot by `DictSlot` name. Slot names are trimmed and matched
480
+ case-insensitively, so `"stphrases"`, `" STPhrases "`, and `"STPhrases"` all select `STPhrases`. Canonical names are
481
+ recommended in TypeScript code and docs; use `OpenccWasm.getAvailableSlots()` to retrieve the current list.
482
+
483
+ Suffixes such as `.txt` are not accepted, even though case and surrounding whitespace are normalized. Use
484
+ `"STPhrases"` or `"stphrases"`, not `"STPhrases.txt"`.
485
+
486
+ Merge contract:
487
+
488
+ - Custom dictionaries are loaded from in-memory pairs only; no file I/O is involved.
489
+ - The embedded compressed CBOR dictionary is loaded first.
490
+ - Custom specs are applied to `DictionaryMaxlength` before `OpenCC::from_dictionary(...)`.
491
+ - Conversion hot paths remain immutable after construction.
492
+ - `Append` mode merges into the selected slot.
493
+ - Duplicate or conflicting keys use last-wins semantics.
494
+ - `Override` mode clears the selected slot first, then inserts the provided custom pairs.
495
+ - Multiple specs are applied in array order.
496
+
497
+ This API is useful for browser apps, user-defined terminology, database-loaded terms, generated dictionaries,
498
+ `localStorage` or `IndexedDB` terms, testing, and embedded WASM environments. Customization happens at construction
499
+ time, not during conversion.
500
+
501
+ ---
502
+
503
+ ## Supported Configs
504
+
505
+ | Config | Enum | Description |
506
+ |---------|--------------------------|-------------------------------------------------------|
507
+ | `s2t` | `OpenccConfigWasm.S2t` | Simplified Chinese → Traditional Chinese |
508
+ | `s2tw` | `OpenccConfigWasm.S2tw` | Simplified Chinese → Taiwan Traditional |
509
+ | `s2twp` | `OpenccConfigWasm.S2twp` | Simplified Chinese → Taiwan Traditional (phrases) |
510
+ | `s2hk` | `OpenccConfigWasm.S2hk` | Simplified Chinese → Hong Kong Traditional |
511
+ | `s2hkp` | `OpenccConfigWasm.S2hkp` | Simplified Chinese → Hong Kong Traditional (phrases) |
512
+ | `t2s` | `OpenccConfigWasm.T2s` | Traditional Chinese → Simplified Chinese |
513
+ | `t2tw` | `OpenccConfigWasm.T2tw` | Traditional Chinese → Taiwan Traditional |
514
+ | `t2twp` | `OpenccConfigWasm.T2twp` | Traditional Chinese → Taiwan Traditional (phrases) |
515
+ | `t2hk` | `OpenccConfigWasm.T2hk` | Traditional Chinese → Hong Kong Traditional |
516
+ | `t2hkp` | `OpenccConfigWasm.T2hkp` | Traditional Chinese → Hong Kong Traditional (phrases) |
517
+ | `tw2s` | `OpenccConfigWasm.Tw2s` | Taiwan Traditional → Simplified Chinese |
518
+ | `tw2sp` | `OpenccConfigWasm.Tw2sp` | Taiwan Traditional → Simplified Chinese (phrases) |
519
+ | `tw2t` | `OpenccConfigWasm.Tw2t` | Taiwan Traditional → Traditional Chinese |
520
+ | `tw2tp` | `OpenccConfigWasm.Tw2tp` | Taiwan Traditional → Traditional Chinese (phrases) |
521
+ | `hk2s` | `OpenccConfigWasm.Hk2s` | Hong Kong Traditional → Simplified Chinese |
522
+ | `hk2sp` | `OpenccConfigWasm.Hk2sp` | Hong Kong Traditional → Simplified Chinese (phrases) |
523
+ | `hk2t` | `OpenccConfigWasm.Hk2t` | Hong Kong Traditional → Traditional Chinese |
524
+ | `hk2tp` | `OpenccConfigWasm.Hk2tp` | Hong Kong Traditional → Traditional Chinese (phrases) |
525
+ | `jp2t` | `OpenccConfigWasm.Jp2t` | Japanese Shinjitai → Traditional Chinese |
526
+ | `t2jp` | `OpenccConfigWasm.T2jp` | Traditional Chinese → Japanese Shinjitai |
527
+
528
+ The numeric enum values match the vendored Rust backend. Existing values are unchanged; `S2hkp = 17`, `Hk2sp = 18`,
529
+ `T2hkp = 19`, and `Hk2tp = 20`.
530
+
531
+ ---
532
+
533
+ ## Using Config Enums
534
+
535
+ ```javascript
536
+ import init, {
537
+ OpenccWasm,
538
+ OpenccConfigWasm
539
+ } from "@laisuk/opencc-fmmseg-wasm";
540
+
541
+ await init();
542
+
543
+ const cc = OpenccWasm.newWithEnum(
544
+ OpenccConfigWasm.S2hkp
545
+ );
546
+
547
+ console.log(cc.convert("别随便录影侵犯个人隐私权", false));
548
+ // 別隨便錄影侵犯個人私隱權
549
+ ```
550
+
551
+ ---
552
+
553
+ ## Office / EPUB Conversion
554
+
555
+ Office and EPUB conversion runs fully locally in the browser or Node.js. Files are passed in and returned as bytes;
556
+ nothing is uploaded to a backend server.
557
+
558
+ This is useful for converting text inside:
559
+
560
+ ```text
561
+ docx, xlsx, pptx, odt, ods, odp, epub
562
+ ```
563
+
564
+ File size is limited by available browser or Node.js memory, but there is no upload or server-side limit. Font
565
+ preservation is supported with the `keepFont` option.
566
+
567
+ Use the instance method when possible. It reuses the converter configuration and any custom dictionaries already held by
568
+ the `OpenccWasm` instance.
569
+
570
+ ```javascript
571
+ cc.convertOfficeBytes(inputBytes, format, punctuation, keepFont)
572
+ ```
573
+
574
+ Parameters:
575
+
576
+ - `inputBytes`: `Uint8Array` document bytes
577
+ - `format`: `docx`, `xlsx`, `pptx`, `odt`, `ods`, `odp`, or `epub`
578
+ - `punctuation`: whether to convert punctuation variants
579
+ - `keepFont`: whether to preserve font declarations where supported
580
+
581
+ Returns:
582
+
583
+ - converted output bytes
584
+
585
+ The older free function remains available for compatibility:
586
+
587
+ ```javascript
588
+ convert_office_bytes(inputBytes, format, config, punctuation, keepFont)
589
+ ```
590
+
591
+ ### Browser Office Example
592
+
593
+ ```javascript
594
+ import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
595
+
596
+ await init();
597
+
598
+ const cc = new OpenccWasm("s2t");
599
+ const file = document.querySelector("input[type=file]").files[0];
600
+ const inputBytes = new Uint8Array(await file.arrayBuffer());
601
+
602
+ const outputBytes = cc.convertOfficeBytes(
603
+ inputBytes,
604
+ "docx",
605
+ true,
606
+ true
607
+ );
608
+
609
+ const blob = new Blob([outputBytes], {
610
+ type: "application/vnd.openxmlformats-officedocument.wordprocessingml.document"
611
+ });
612
+
613
+ const a = document.createElement("a");
614
+ a.href = URL.createObjectURL(blob);
615
+ a.download = "converted.docx";
616
+ a.click();
617
+ URL.revokeObjectURL(a.href);
618
+ ```
619
+
620
+ ### Node.js Office Example
621
+
622
+ ```javascript
623
+ import fs from "fs";
624
+ import init, {OpenccWasm} from "@laisuk/opencc-fmmseg-wasm";
625
+
626
+ await init();
627
+
628
+ const cc = new OpenccWasm("s2t");
629
+ const inputBytes = fs.readFileSync("input.docx");
630
+
631
+ const outputBytes = cc.convertOfficeBytes(
632
+ inputBytes,
633
+ "docx",
634
+ true,
635
+ true
636
+ );
637
+
638
+ fs.writeFileSync("output.docx", outputBytes);
639
+ ```
640
+
641
+ ---
642
+
643
+ ## Browser Example
644
+
645
+ ```html
646
+ <!DOCTYPE html>
647
+ <html lang="en">
648
+ <head>
649
+ <meta charset="UTF-8">
650
+ <title>OpenCC WASM Demo</title>
651
+ </head>
652
+ <body>
653
+
654
+ <script type="module">
655
+ import init, {
656
+ OpenccWasm
657
+ } from "./pkg/opencc_fmmseg_wasm.js";
658
+
659
+ await init();
660
+
661
+ const cc = new OpenccWasm("s2t");
662
+
663
+ console.log(
664
+ cc.convert("汉字", false)
665
+ );
666
+ </script>
667
+
668
+ </body>
669
+ </html>
670
+ ```
671
+
672
+ > **Note**
673
+ >
674
+ > Normally, `await init();` is sufficient when using the published npm package.
675
+ >
676
+ > When running directly from a local repository checkout (for example in tests
677
+ > or development scripts), initialize using explicit WASM bytes:
678
+ >
679
+ > ```javascript
680
+ > import fs from "fs";
681
+ > import init from "../pkg/opencc_fmmseg_wasm.js";
682
+ >
683
+ > const wasmBytes = fs.readFileSync(
684
+ > "../pkg/opencc_fmmseg_wasm_bg.wasm"
685
+ > );
686
+ >
687
+ > await init({
688
+ > module_or_path: wasmBytes
689
+ > });
690
+ > ```
691
+
692
+ ---
693
+
694
+ ## Node.js CLI
695
+
696
+ The package includes a zero-dependency Node.js CLI:
697
+
698
+ ```bash
699
+ opencc-fmmseg convert -i input.txt -o output.txt -c s2t -p
700
+ opencc-fmmseg convert -i input.txt -o output.txt -c t2s -p --detofu all
701
+ echo "别随便录影侵犯个人隐私权" | opencc-fmmseg convert -c s2hkp
702
+ echo "天龍八部書裡的喬峰是契丹人" | opencc-fmmseg convert -c t2s --norm-compat
703
+ // 天龙八部书里的乔峰是契丹人
704
+ echo "這個細路哥很靈活" | opencc-fmmseg convert -c hk2sp --custom-dict hkphrasesrev:append:my_hk_dict.txt
705
+ // 这个小男孩很灵活
706
+ ```
707
+
708
+ my_hk_dict.txt:
709
+
710
+ ```
711
+ # Custom Dictionary
712
+
713
+ 細路哥 小男孩
714
+ ```
715
+
716
+ ```bash
717
+ opencc-fmmseg office -i input.docx -o output.docx -c s2t -p --keep-font
718
+ ```
719
+
720
+ ### Text Conversion Options
721
+
722
+ ```text
723
+ -i, --input <file> Input text file; stdin if omitted
724
+ -o, --output <file> Output text file; stdout if omitted
725
+ -c, --config <conversion> Conversion config (default: s2t)
726
+ -p, --punct Enable punctuation conversion
727
+ --detofu [level] Replace tofu-risk rare CJK extension chars after conversion
728
+ level: all | ext-b | ext-c | ext-d | ext-e | ext-f | ext-g | ext-h | ext-i
729
+ default when omitted value: all
730
+ --keep-ids Preserve complete IDS expressions during conversion (default: false)
731
+ -n, --norm-compat Normalize CJK Compatibility Ideographs before conversion (default: false)
732
+ -E, --norm-compat-extended Normalize extended Unicode compatibility forms before conversion (default: false)
733
+ -D, --custom-dict <slot:mode:file>
734
+ Load a custom dictionary.
735
+ May be specified multiple times.
736
+ Examples:
737
+ --custom-dict hkphrasesrev:append:my_hk_dict.txt
738
+ --custom-dict stphrases:override:terms.txt
739
+ --in-enc <encoding> Input encoding (default: utf8)
740
+ --out-enc <encoding> Output encoding (default: utf8)
741
+ ```
742
+
743
+ Supported conversion configs:
744
+
745
+ ```text
746
+ s2t, s2tw, s2twp, s2hk, s2hkp, t2s, t2tw, t2twp, t2hk, t2hkp,
747
+ tw2s, tw2sp, tw2t, tw2tp, hk2s, hk2sp, hk2t, hk2tp, jp2t, t2jp
748
+ ```
749
+
750
+ ### Office / EPUB Options
751
+
752
+ ```text
753
+ -i, --input <file> Input Office / EPUB file
754
+ -o, --output <file> Output file
755
+ -c, --config <conversion> Conversion config (default: s2t)
756
+ -p, --punct Enable punctuation conversion
757
+ -f, --format <format> docx | xlsx | pptx | odt | ods | odp | epub
758
+ -F, --convert-filename Convert generated output filename stem (default: false)
759
+ --keep-font Preserve font-family information (default)
760
+ --no-keep-font Do not preserve font-family information
761
+ --custom-dict <slot:mode:file>
762
+ Load a custom dictionary.
763
+ May be specified multiple times.
764
+ Examples:
765
+ --custom-dict hkphrasesrev:append:my_hk_dict.txt
766
+ --custom-dict stphrases:override:terms.txt
767
+ ```
768
+
769
+ For `office`, the format is inferred from the input file extension when `--format` is omitted.
770
+
771
+ If `-o, --output` is omitted, `office` writes:
772
+
773
+ ```text
774
+ <input-name>_converted.<ext>
775
+ ```
776
+
777
+ ---
778
+
779
+ ## TypeScript Support
780
+
781
+ The package includes generated TypeScript definitions from `wasm-bindgen`.
782
+
783
+ The WASM-facing enum is exported as `OpenccConfigWasm`, alongside `OpenccWasm`.
784
+
785
+ `OpenccConfigWasm.S2hkp`, `OpenccConfigWasm.Hk2sp`, `OpenccConfigWasm.T2hkp`, and
786
+ `OpenccConfigWasm.Hk2tp` are available for Hong Kong phrase conversions and map to backend config IDs `17` through `20`.
787
+
788
+ ---
789
+
790
+ ## Performance Notes
791
+
792
+ - WebAssembly build disables Rayon parallelism by default.
793
+ - Dictionaries are embedded into the WASM binary.
794
+ - Browser caching significantly improves subsequent loads.
795
+
796
+ ---
797
+
798
+ ## Related Projects
799
+
800
+ - Rust backend: https://github.com/laisuk/opencc-fmmseg
801
+ - C API: https://github.com/laisuk/opencc-fmmseg/tree/master/capi/opencc-fmmseg-capi
802
+ - .NET: https://github.com/laisuk/OpenccNet
803
+ - Python: https://github.com/laisuk/opencc_purepy
804
+
805
+ ---
806
+
807
+ ## License
808
+
809
+ MIT