@singapore-editor/spellcheck 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +106 -0
- package/THIRD_PARTY_NOTICES +611 -0
- package/dist/assets/spellcheck.worker-Dx-kt277.js +8971 -0
- package/dist/assets/spellcheck.worker-Dx-kt277.js.map +1 -0
- package/dist/controller.d.ts +51 -0
- package/dist/controller.d.ts.map +1 -0
- package/dist/controller.js +307 -0
- package/dist/controller.js.map +1 -0
- package/dist/dictionaryAssets.d.ts +4 -0
- package/dist/dictionaryAssets.d.ts.map +1 -0
- package/dist/dictionaryData.d.ts +13 -0
- package/dist/dictionaryData.d.ts.map +1 -0
- package/dist/engine.d.ts +18 -0
- package/dist/engine.d.ts.map +1 -0
- package/dist/feature.d.ts +19 -0
- package/dist/feature.d.ts.map +1 -0
- package/dist/feature.js +8 -0
- package/dist/feature.js.map +1 -0
- package/dist/index.d.ts +7 -0
- package/dist/index.d.ts.map +1 -0
- package/dist/index.js +5 -0
- package/dist/plugin.d.ts +12 -0
- package/dist/plugin.d.ts.map +1 -0
- package/dist/plugin.js +27 -0
- package/dist/plugin.js.map +1 -0
- package/dist/proseRanges.d.ts +25 -0
- package/dist/proseRanges.d.ts.map +1 -0
- package/dist/proseRanges.js +75 -0
- package/dist/proseRanges.js.map +1 -0
- package/dist/protocol.d.ts +27 -0
- package/dist/protocol.d.ts.map +1 -0
- package/dist/service.d.ts +38 -0
- package/dist/service.d.ts.map +1 -0
- package/dist/service.js +128 -0
- package/dist/service.js.map +1 -0
- package/dist/spellcheck.worker.d.ts +2 -0
- package/dist/spellcheck.worker.d.ts.map +1 -0
- package/dist/styles.d.ts +7 -0
- package/dist/styles.d.ts.map +1 -0
- package/dist/styles.js +13 -0
- package/dist/styles.js.map +1 -0
- package/dist/tokenizer.d.ts +23 -0
- package/dist/tokenizer.d.ts.map +1 -0
- package/dist/tokenizer.js +140 -0
- package/dist/tokenizer.js.map +1 -0
- package/package.json +52 -0
package/README.md
ADDED
|
@@ -0,0 +1,106 @@
|
|
|
1
|
+
# @singapore-editor/spellcheck
|
|
2
|
+
|
|
3
|
+
English spellcheck engine for text the editor paints itself. A worker holds the dictionary; the
|
|
4
|
+
page creates one `SpellcheckService` and shares it between editors.
|
|
5
|
+
|
|
6
|
+
US and British spellings are both accepted, plus software vocabulary. The dictionary is built from
|
|
7
|
+
SCOWL and cspell's word lists by `bun run build:dictionaries`; see `THIRD_PARTY_NOTICES`.
|
|
8
|
+
|
|
9
|
+
## Usage
|
|
10
|
+
|
|
11
|
+
In an editor: one service per page, one plugin per editor.
|
|
12
|
+
|
|
13
|
+
```ts
|
|
14
|
+
import { Editor } from '@singapore-editor/core/editor'
|
|
15
|
+
import {
|
|
16
|
+
createSpellcheckPlugin,
|
|
17
|
+
EDITOR_SPELLCHECK_FEATURE,
|
|
18
|
+
SpellcheckService,
|
|
19
|
+
} from '@singapore-editor/spellcheck'
|
|
20
|
+
|
|
21
|
+
const service = new SpellcheckService()
|
|
22
|
+
const editor = new Editor(element, { plugins: [createSpellcheckPlugin({ service })] })
|
|
23
|
+
const spelling = editor.getFeature(EDITOR_SPELLCHECK_FEATURE)
|
|
24
|
+
const issue = spelling?.issueAt(offset) // { start, end, word } or null
|
|
25
|
+
const suggestions = await spelling?.suggestions(offset)
|
|
26
|
+
spelling?.replace(offset, suggestions[0]) // one undoable edit
|
|
27
|
+
```
|
|
28
|
+
|
|
29
|
+
Plain text and Markdown prose are checked; `scope: 'proseAndCode'` adds comments and strings in
|
|
30
|
+
code. Markdown code, link targets and labels, and anything an inline replacement stands in for are
|
|
31
|
+
skipped. The word being typed is not marked until the caret leaves it.
|
|
32
|
+
|
|
33
|
+
On its own:
|
|
34
|
+
|
|
35
|
+
```ts
|
|
36
|
+
import { SpellcheckService, tokenizeSpellWords } from '@singapore-editor/spellcheck'
|
|
37
|
+
|
|
38
|
+
const spellcheck = new SpellcheckService()
|
|
39
|
+
const text = 'the list settles befor the cursor'
|
|
40
|
+
const words = tokenizeSpellWords(text)
|
|
41
|
+
const misspelled = new Set(await spellcheck.check(words.map((word) => word.word)))
|
|
42
|
+
const marks = words.filter((word) => misspelled.has(word.word))
|
|
43
|
+
const suggestions = await spellcheck.suggest('befor') // ['before', …]
|
|
44
|
+
spellcheck.setAcceptedWords(['fregat'])
|
|
45
|
+
spellcheck.dispose()
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## Exports
|
|
49
|
+
|
|
50
|
+
- `createSpellcheckPlugin({ service, scope })` and `EDITOR_SPELLCHECK_FEATURE`: `issueAt(offset)`,
|
|
51
|
+
`suggestions(offset, limit)`, `replace(offset, word)` and `setAcceptedWords(words)`.
|
|
52
|
+
- `SpellcheckService`: `check(words)`, `suggest(word, limit)`, `setAcceptedWords(words)` and
|
|
53
|
+
`dispose()`. The worker starts on the first request.
|
|
54
|
+
- `tokenizeSpellWords(text, { mode, excluded })`: the words to check, with offsets. It skips
|
|
55
|
+
acronyms, camelCase, words with digits or `_`, runs containing letters outside ASCII, URLs, email addresses, paths and
|
|
56
|
+
the `excluded` ranges. `mode: 'code'` splits camelCase and snake_case instead.
|
|
57
|
+
|
|
58
|
+
## Language and work limits
|
|
59
|
+
|
|
60
|
+
The bundled dictionary checks English. There is no natural-language detection. Prose candidates
|
|
61
|
+
contain ASCII letters and internal straight or typographic apostrophes. `hola mundo` is submitted
|
|
62
|
+
to the English checker; `שלום`, `привет` and `café` are skipped. This also skips accented words
|
|
63
|
+
that occur in English. Programming-language selection controls which editor regions are prose.
|
|
64
|
+
|
|
65
|
+
Whitespace-delimited chunks longer than 256 UTF-16 code units are skipped before structured-text
|
|
66
|
+
classification. The editor skips prose lines and code regions longer than 16,384 code units before
|
|
67
|
+
reading them. This bounds synchronous work on generated text and large pastes. Short surrounding
|
|
68
|
+
chunks remain checkable through `tokenizeSpellWords`.
|
|
69
|
+
|
|
70
|
+
Worker setup and posting failures reject `check` and `suggest`, settle every outstanding request,
|
|
71
|
+
and terminate that worker. A later service request creates a fresh worker and resends accepted
|
|
72
|
+
words. An editor that sees a check failure stops requesting checks for its lifetime so typing
|
|
73
|
+
cannot start a retry loop. Failed accepted-word synchronization retains the local list for the
|
|
74
|
+
next worker and notifies listeners. Disposal rejects outstanding requests and prevents restart.
|
|
75
|
+
|
|
76
|
+
## Language-support follow-up
|
|
77
|
+
|
|
78
|
+
Add an explicit `languages` selection separate from syntax language and default it to English.
|
|
79
|
+
Start with manually selected dictionaries; mixed-language documents accept a word found in any
|
|
80
|
+
selected dictionary, and merge/deduplicate suggestions in configured language order. Automatic
|
|
81
|
+
language detection is outside this scope.
|
|
82
|
+
|
|
83
|
+
Load dictionaries lazily in the shared service from a language manifest. Each entry must record
|
|
84
|
+
its source, version, attribution and verified permissive data license. Review these for each
|
|
85
|
+
selected dictionary; the engine license does not cover dictionary data. Bundle English only by
|
|
86
|
+
default and load other selected assets on demand.
|
|
87
|
+
|
|
88
|
+
Change tokenization with dictionary selection: Unicode letters/marks, language-specific word
|
|
89
|
+
segmentation, apostrophe policy and NFC normalization must preserve original UTF-16 offsets.
|
|
90
|
+
Keep script selection explicit so mixed scripts are checked only when a selected dictionary
|
|
91
|
+
supports them. Define case folding per dictionary rather than globally stripping accents.
|
|
92
|
+
|
|
93
|
+
Treat a language change as a new dictionary generation. Clear controller verdicts, pending-word
|
|
94
|
+
sets, cached line tokenization and painted issues; ignore replies from the previous generation.
|
|
95
|
+
Only publish the new generation once all selected dictionaries load. A load failure leaves the
|
|
96
|
+
previous complete selection active and reports the failed selection to the host. Accepted words
|
|
97
|
+
remain user-owned and are synchronized into the new generation.
|
|
98
|
+
|
|
99
|
+
## Benchmark interpretation
|
|
100
|
+
|
|
101
|
+
`bun run bench:engine` evaluates the common Norvig spell-testset1/2 pairs. It reports total pairs,
|
|
102
|
+
target-word coverage, and misspellings accepted despite a covered target separately. Suggestion
|
|
103
|
+
ranking runs only on pairs where the dictionary accepts the target and rejects the typo.
|
|
104
|
+
`eligibleRankingPairs` is the denominator of `eligibleTop1Percent` and `eligibleTop5Percent`.
|
|
105
|
+
Compare engines using the same pair set and report coverage and missed typos alongside ranking.
|
|
106
|
+
These conditional ranking percentages do not measure overall typo-detection accuracy.
|