@singapore-editor/spellcheck 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (46) hide show
  1. package/README.md +106 -0
  2. package/THIRD_PARTY_NOTICES +611 -0
  3. package/dist/assets/spellcheck.worker-Dx-kt277.js +8971 -0
  4. package/dist/assets/spellcheck.worker-Dx-kt277.js.map +1 -0
  5. package/dist/controller.d.ts +51 -0
  6. package/dist/controller.d.ts.map +1 -0
  7. package/dist/controller.js +307 -0
  8. package/dist/controller.js.map +1 -0
  9. package/dist/dictionaryAssets.d.ts +4 -0
  10. package/dist/dictionaryAssets.d.ts.map +1 -0
  11. package/dist/dictionaryData.d.ts +13 -0
  12. package/dist/dictionaryData.d.ts.map +1 -0
  13. package/dist/engine.d.ts +18 -0
  14. package/dist/engine.d.ts.map +1 -0
  15. package/dist/feature.d.ts +19 -0
  16. package/dist/feature.d.ts.map +1 -0
  17. package/dist/feature.js +8 -0
  18. package/dist/feature.js.map +1 -0
  19. package/dist/index.d.ts +7 -0
  20. package/dist/index.d.ts.map +1 -0
  21. package/dist/index.js +5 -0
  22. package/dist/plugin.d.ts +12 -0
  23. package/dist/plugin.d.ts.map +1 -0
  24. package/dist/plugin.js +27 -0
  25. package/dist/plugin.js.map +1 -0
  26. package/dist/proseRanges.d.ts +25 -0
  27. package/dist/proseRanges.d.ts.map +1 -0
  28. package/dist/proseRanges.js +75 -0
  29. package/dist/proseRanges.js.map +1 -0
  30. package/dist/protocol.d.ts +27 -0
  31. package/dist/protocol.d.ts.map +1 -0
  32. package/dist/service.d.ts +38 -0
  33. package/dist/service.d.ts.map +1 -0
  34. package/dist/service.js +128 -0
  35. package/dist/service.js.map +1 -0
  36. package/dist/spellcheck.worker.d.ts +2 -0
  37. package/dist/spellcheck.worker.d.ts.map +1 -0
  38. package/dist/styles.d.ts +7 -0
  39. package/dist/styles.d.ts.map +1 -0
  40. package/dist/styles.js +13 -0
  41. package/dist/styles.js.map +1 -0
  42. package/dist/tokenizer.d.ts +23 -0
  43. package/dist/tokenizer.d.ts.map +1 -0
  44. package/dist/tokenizer.js +140 -0
  45. package/dist/tokenizer.js.map +1 -0
  46. package/package.json +52 -0
package/README.md ADDED
@@ -0,0 +1,106 @@
1
+ # @singapore-editor/spellcheck
2
+
3
+ English spellcheck engine for text the editor paints itself. A worker holds the dictionary; the
4
+ page creates one `SpellcheckService` and shares it between editors.
5
+
6
+ US and British spellings are both accepted, plus software vocabulary. The dictionary is built from
7
+ SCOWL and cspell's word lists by `bun run build:dictionaries`; see `THIRD_PARTY_NOTICES`.
8
+
9
+ ## Usage
10
+
11
+ In an editor: one service per page, one plugin per editor.
12
+
13
+ ```ts
14
+ import { Editor } from '@singapore-editor/core/editor'
15
+ import {
16
+ createSpellcheckPlugin,
17
+ EDITOR_SPELLCHECK_FEATURE,
18
+ SpellcheckService,
19
+ } from '@singapore-editor/spellcheck'
20
+
21
+ const service = new SpellcheckService()
22
+ const editor = new Editor(element, { plugins: [createSpellcheckPlugin({ service })] })
23
+ const spelling = editor.getFeature(EDITOR_SPELLCHECK_FEATURE)
24
+ const issue = spelling?.issueAt(offset) // { start, end, word } or null
25
+ const suggestions = await spelling?.suggestions(offset)
26
+ spelling?.replace(offset, suggestions[0]) // one undoable edit
27
+ ```
28
+
29
+ Plain text and Markdown prose are checked; `scope: 'proseAndCode'` adds comments and strings in
30
+ code. Markdown code, link targets and labels, and anything an inline replacement stands in for are
31
+ skipped. The word being typed is not marked until the caret leaves it.
32
+
33
+ On its own:
34
+
35
+ ```ts
36
+ import { SpellcheckService, tokenizeSpellWords } from '@singapore-editor/spellcheck'
37
+
38
+ const spellcheck = new SpellcheckService()
39
+ const text = 'the list settles befor the cursor'
40
+ const words = tokenizeSpellWords(text)
41
+ const misspelled = new Set(await spellcheck.check(words.map((word) => word.word)))
42
+ const marks = words.filter((word) => misspelled.has(word.word))
43
+ const suggestions = await spellcheck.suggest('befor') // ['before', …]
44
+ spellcheck.setAcceptedWords(['fregat'])
45
+ spellcheck.dispose()
46
+ ```
47
+
48
+ ## Exports
49
+
50
+ - `createSpellcheckPlugin({ service, scope })` and `EDITOR_SPELLCHECK_FEATURE`: `issueAt(offset)`,
51
+ `suggestions(offset, limit)`, `replace(offset, word)` and `setAcceptedWords(words)`.
52
+ - `SpellcheckService`: `check(words)`, `suggest(word, limit)`, `setAcceptedWords(words)` and
53
+ `dispose()`. The worker starts on the first request.
54
+ - `tokenizeSpellWords(text, { mode, excluded })`: the words to check, with offsets. It skips
55
+ acronyms, camelCase, words with digits or `_`, runs containing letters outside ASCII, URLs, email addresses, paths and
56
+ the `excluded` ranges. `mode: 'code'` splits camelCase and snake_case instead.
57
+
58
+ ## Language and work limits
59
+
60
+ The bundled dictionary checks English. There is no natural-language detection. Prose candidates
61
+ contain ASCII letters and internal straight or typographic apostrophes. `hola mundo` is submitted
62
+ to the English checker; `שלום`, `привет` and `café` are skipped. This also skips accented words
63
+ that occur in English. Programming-language selection controls which editor regions are prose.
64
+
65
+ Whitespace-delimited chunks longer than 256 UTF-16 code units are skipped before structured-text
66
+ classification. The editor skips prose lines and code regions longer than 16,384 code units before
67
+ reading them. This bounds synchronous work on generated text and large pastes. Short surrounding
68
+ chunks remain checkable through `tokenizeSpellWords`.
69
+
70
+ Worker setup and posting failures reject `check` and `suggest`, settle every outstanding request,
71
+ and terminate that worker. A later service request creates a fresh worker and resends accepted
72
+ words. An editor that sees a check failure stops requesting checks for its lifetime so typing
73
+ cannot start a retry loop. Failed accepted-word synchronization retains the local list for the
74
+ next worker and notifies listeners. Disposal rejects outstanding requests and prevents restart.
75
+
76
+ ## Language-support follow-up
77
+
78
+ Add an explicit `languages` selection separate from syntax language and default it to English.
79
+ Start with manually selected dictionaries; mixed-language documents accept a word found in any
80
+ selected dictionary, and merge/deduplicate suggestions in configured language order. Automatic
81
+ language detection is outside this scope.
82
+
83
+ Load dictionaries lazily in the shared service from a language manifest. Each entry must record
84
+ its source, version, attribution and verified permissive data license. Review these for each
85
+ selected dictionary; the engine license does not cover dictionary data. Bundle English only by
86
+ default and load other selected assets on demand.
87
+
88
+ Change tokenization with dictionary selection: Unicode letters/marks, language-specific word
89
+ segmentation, apostrophe policy and NFC normalization must preserve original UTF-16 offsets.
90
+ Keep script selection explicit so mixed scripts are checked only when a selected dictionary
91
+ supports them. Define case folding per dictionary rather than globally stripping accents.
92
+
93
+ Treat a language change as a new dictionary generation. Clear controller verdicts, pending-word
94
+ sets, cached line tokenization and painted issues; ignore replies from the previous generation.
95
+ Only publish the new generation once all selected dictionaries load. A load failure leaves the
96
+ previous complete selection active and reports the failed selection to the host. Accepted words
97
+ remain user-owned and are synchronized into the new generation.
98
+
99
+ ## Benchmark interpretation
100
+
101
+ `bun run bench:engine` evaluates the common Norvig spell-testset1/2 pairs. It reports total pairs,
102
+ target-word coverage, and misspellings accepted despite a covered target separately. Suggestion
103
+ ranking runs only on pairs where the dictionary accepts the target and rejects the typo.
104
+ `eligibleRankingPairs` is the denominator of `eligibleTop1Percent` and `eligibleTop5Percent`.
105
+ Compare engines using the same pair set and report coverage and missed typos alongside ranking.
106
+ These conditional ranking percentages do not measure overall typo-detection accuracy.