kotoshu-server 0.1.0 → 0.1.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
checksums.yaml CHANGED
@@ -1,7 +1,7 @@
1
1
  ---
2
2
  SHA256:
3
- metadata.gz: 80541f5f06f9c4f0c41cbdecfbf195fec21d630f9e89711c061f4a3796753cfc
4
- data.tar.gz: dae5536d6035f67546136f12ffe1fc52d602bdffc8dd5eeba120fec2313c4cd7
3
+ metadata.gz: 4a1120bdda9dcd7ecc468c95d5386127efefa6509e7aa1a03f2fa12a5e28c33f
4
+ data.tar.gz: b8cbe5d0db18715009576fd104be80bb06e35ba1b730cbd2564178755a05e6a1
5
5
  SHA512:
6
- metadata.gz: b3bcab28c17474fbc07c721bb7686ae6d7505580a70e03042438aef75d032b846efc9576520351f1967a4e5fc04d8e9b21590a4961b7f9bbc420bb3ff10505b6
7
- data.tar.gz: 8cf8d948c812075b068479e25737a0d43999831f2dea87a018d6d62fe8c1049d96bb27076f88173ed91d4bd4ae1939669d768402ee9e8ce8e5c215c14014d909
6
+ metadata.gz: 6c001de9ec24f78e48eac1c400b0ecf9038b47fdeed7115ad3ed0635e100adc7b3ae9e241a7474c8956636f1690f3dea39b425c125d122577aa67dfd097af2a4
7
+ data.tar.gz: 8696ca46e3728fbd2b2f32b3e9468be4e544290d83badb6837b2133bb76700fd7241c687446b7926e9d0a485eb95fbae22d44183503da64cb84cada4227080d7
data/LICENSE ADDED
@@ -0,0 +1,26 @@
1
+ BSD 2-Clause License
2
+
3
+ Copyright (c) 2025-2026, Kotoshu contributors
4
+ All rights reserved.
5
+
6
+ Redistribution and use in source and binary forms, with or without
7
+ modification, are permitted provided that the following conditions are met:
8
+
9
+ 1. Redistributions of source code must retain the above copyright notice, this
10
+ list of conditions and the following disclaimer.
11
+
12
+ 2. Redistributions in binary form must reproduce the above copyright notice,
13
+ this list of conditions and the following disclaimer in the documentation
14
+ and/or other materials provided with the distribution.
15
+
16
+ THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
17
+ AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
18
+ IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
19
+ DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
20
+ FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
21
+ DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
22
+ SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
23
+ CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
24
+ OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
25
+ OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
26
+
data/README.md ADDED
@@ -0,0 +1,163 @@
1
+ # kotoshu-server
2
+
3
+ Self-hostable HTTP API wrapping the [Kotoshu](https://github.com/kotoshu/kotoshu) spell checker.
4
+
5
+ ## Status
6
+
7
+ MVP. Six endpoints over JSON. Rack/Sinatra + Puma. Pre-warms
8
+ languages on boot. Designed as the deployment surface for non-Ruby
9
+ SDKs (`kotoshu-python`, `kotoshu-js`, `kotoshu-go`).
10
+
11
+ See `TODO.impl/64-http-api-and-sdks.md` for the full plan.
12
+
13
+ ## Install
14
+
15
+ ```bash
16
+ gem install kotoshu-server # >= 0.1.1 (the 0.1.0 gem was empty — a gemspec file-list bug)
17
+ kotoshu-server
18
+ ```
19
+
20
+ Or run from source:
21
+
22
+ ```bash
23
+ cd kotoshu-server && bundle install && bundle exec exe/kotoshu-server
24
+ ```
25
+
26
+ ## Run
27
+
28
+ ```bash
29
+ # Default: localhost:9292, English pre-warmed
30
+ kotoshu-server
31
+
32
+ # Custom port + multi-language pre-warm
33
+ KOTOSHU_SERVER_PORT=8080 \
34
+ KOTOSHU_SERVER_LANGUAGES="en de fr" \
35
+ kotoshu-server
36
+ ```
37
+
38
+ ## Semantic models (opt-in)
39
+
40
+ By default the server is dictionary-only — exactly the 0.1.0
41
+ behavior. Semantic reranking is a boot-time opt-in; the server never
42
+ downloads a model implicitly:
43
+
44
+ ```bash
45
+ KOTOSHU_SERVER_LANGUAGES="en de" \
46
+ KOTOSHU_SERVER_MODEL_LANGS="en" \
47
+ kotoshu-server
48
+ ```
49
+
50
+ - `KOTOSHU_SERVER_MODEL_LANGS` — languages the pre-warm thread sets
51
+ up with spelling + semantic model (space separated). Unset means no
52
+ models, no downloads, no behavior change.
53
+ - `KOTOSHU_SERVER_MODEL_TIER` — model tier to set up and resolve:
54
+ `fluency` (default), `full`, or `mini`.
55
+
56
+ **Requires kotoshu >= 0.7.0.** Model tiers, the resource registry,
57
+ and the confidence cascade ship in the 0.7.0 cut. The gemspec keeps
58
+ its `kotoshu ~> 0.6` constraint (dependency floors are the owner's
59
+ decision), so the server checks the *installed* gem instead: setting
60
+ `KOTOSHU_SERVER_MODEL_LANGS` with an older kotoshu fails fast at
61
+ boot with a clear error, and explicit `"model": true` requests
62
+ return 503.
63
+
64
+ Once a language has a model set up, `POST /v1/check` accepts an
65
+ optional `"model": true|false` flag (default: whether the language
66
+ has a model set up server-side). With it on, each error's
67
+ suggestions are reranked by the semantic analyzer; the gem's
68
+ confidence cascade (`KOTOSHU_SEMANTIC_CASCADE_THRESHOLD`) decides per
69
+ word whether the ONNX rerank actually runs. `GET /v1/languages`
70
+ reports `"model": {"en": true}` per cached language.
71
+
72
+ Memory and latency: budget roughly 15 MB resident per language at
73
+ the default fluency tier (the full tier is far larger), plus
74
+ one-time model load on the first model-enabled request. Listing the
75
+ language in `KOTOSHU_SERVER_MODEL_LANGS` warms it at boot instead.
76
+
77
+ ## Language detection
78
+
79
+ `POST /v1/detect` reports which engine answered in the `engine`
80
+ field:
81
+
82
+ - `"lid-176"` — the 176-language lid.176 model served through the
83
+ kotoshu native extension (kotoshu >= 0.10.0). The model is set up
84
+ lazily on the first detect — one download, never at boot, and
85
+ never at all under `KOTOSHU_OFFLINE=1` without a cache (Docker
86
+ defaults to offline).
87
+ - `"heuristic"` — the 7-language character-set heuristic (en, de,
88
+ es, fr, pt, ru, ja), the fallback whenever lid-176 cannot serve:
89
+ kotoshu < 0.10.0, a pure-Ruby install or `KOTOSHU_BACKEND=ruby`,
90
+ or the model missing after a failed setup.
91
+
92
+ Set `KOTOSHU_DETECT=heuristic` to pin the heuristic regardless of
93
+ the installed gem — the 0.1.1 behavior, byte for byte.
94
+
95
+ ## Endpoints
96
+
97
+ | Method | Path | Body | Returns |
98
+ |---|---|---|---|
99
+ | `GET` | `/` | — | service metadata |
100
+ | `GET` | `/v1/health` | — | `{ status, ready, timestamp }` |
101
+ | `GET` | `/v1/version` | — | `{ server, kotoshu, ruby }` |
102
+ | `GET` | `/v1/languages` | — | `{ cached, supported, model }` |
103
+ | `POST` | `/v1/check` | `{ text, language?, format?, model? }` | `{ file, word_count, errors: [...] }` |
104
+ | `POST` | `/v1/suggest` | `{ word, language?, max? }` | `{ word, suggestions: [...] }` |
105
+ | `POST` | `/v1/detect` | `{ text }` | `{ language, confidence, engine }` |
106
+
107
+ Each `error` and `suggestion` mirrors the lutaml-model serialization
108
+ (`word`, `distance`, `confidence`, `source`).
109
+
110
+ ## curl examples
111
+
112
+ ```bash
113
+ # Check a document
114
+ curl -X POST http://localhost:9292/v1/check \
115
+ -H "Content-Type: application/json" \
116
+ -d '{"text":"helo wrold","language":"en"}'
117
+
118
+ # Suggestions for one word
119
+ curl -X POST http://localhost:9292/v1/suggest \
120
+ -H "Content-Type: application/json" \
121
+ -d '{"word":"helo","language":"en","max":3}'
122
+
123
+ # Detect language
124
+ curl -X POST http://localhost:9292/v1/detect \
125
+ -H "Content-Type: application/json" \
126
+ -d '{"text":"bonjour le monde"}'
127
+ ```
128
+
129
+ ## Docker
130
+
131
+ ```bash
132
+ docker build -t kotoshu-server .
133
+
134
+ docker run --rm -p 9292:9292 \
135
+ -e KOTOSHU_SERVER_LANGUAGES="en de" \
136
+ kotoshu-server
137
+ ```
138
+
139
+ Healthcheck probes `/v1/health` every 30s.
140
+
141
+ ## Configuration
142
+
143
+ | Env var | Default | Purpose |
144
+ |---|---|---|
145
+ | `KOTOSHU_SERVER_PORT` | `9292` | Listen port |
146
+ | `KOTOSHU_SERVER_BIND` | `0.0.0.0` | Bind address |
147
+ | `KOTOSHU_SERVER_LANGUAGES` | `en` | Languages to pre-warm on boot |
148
+ | `KOTOSHU_SERVER_MODEL_LANGS` | unset | Languages to set up with a semantic model (requires kotoshu >= 0.7) |
149
+ | `KOTOSHU_SERVER_MODEL_TIER` | `fluency` | Model tier for `KOTOSHU_SERVER_MODEL_LANGS` |
150
+ | `KOTOSHU_SERVER_LAZY` | `0` | Skip pre-warm; load on first request |
151
+ | `KOTOSHU_SERVER_DEFAULT_LANG` | `en` | When client omits `language` |
152
+ | `KOTOSHU_SERVER_LOG_LEVEL` | `info` | `debug`/`info`/`warn`/`error` |
153
+ | `KOTOSHU_DETECT` | `auto` | `heuristic` pins /v1/detect to the heuristic engine |
154
+ | `KOTOSHU_OFFLINE` | `1` (in Docker) | Never trigger downloads |
155
+
156
+ ## OpenAPI
157
+
158
+ The full OpenAPI 3.1 spec is at [`openapi.yaml`](./openapi.yaml). Use
159
+ `openapi-generator` to spin up SDKs in Rust / .NET / Java.
160
+
161
+ ## License
162
+
163
+ BSD-2-Clause, same as Kotoshu.
@@ -0,0 +1,7 @@
1
+ #!/usr/bin/env ruby
2
+ # frozen_string_literal: true
3
+
4
+ $LOAD_PATH.unshift(File.expand_path("../lib", __dir__))
5
+ require "kotoshu/server"
6
+
7
+ Kotoshu::Server.run!
@@ -0,0 +1,613 @@
1
+ # frozen_string_literal: true
2
+
3
+ require "json"
4
+ require "logger"
5
+ require "sinatra/base"
6
+ require "kotoshu"
7
+
8
+ module Kotoshu
9
+ module Server
10
+ # Raised at boot when KOTOSHU_SERVER_MODEL_LANGS / _TIER are set
11
+ # but cannot be honored (kotoshu gem too old, unknown tier).
12
+ class ModelConfigError < StandardError; end
13
+
14
+ class App < Sinatra::Base
15
+ # The release version, from lib/kotoshu/server/version.rb —
16
+ # the only source of truth (the release workflow bumps it).
17
+ VERSION = Kotoshu::Server::VERSION
18
+
19
+ # Semantic models (tiers, registry, confidence cascade) ship in
20
+ # kotoshu 0.7.0. The gemspec keeps its `kotoshu ~> 0.6`
21
+ # constraint (dependency floors are the owner's call), so the
22
+ # server validates the *installed* gem at boot instead.
23
+ MODEL_MIN_KOTOSHU = Gem::Version.new("0.7.0")
24
+
25
+ # Native language identification (plan 106) ships in kotoshu
26
+ # 0.10.0: Kotoshu.detect_language returning a
27
+ # Language::Detection backed by the lid-176 model through the
28
+ # native extension. Same policy as MODEL_MIN_KOTOSHU — the
29
+ # gemspec floor stays `kotoshu ~> 0.6` (older gems keep the
30
+ # heuristic), so the installed gem is checked per request.
31
+ DETECT_MIN_KOTOSHU = Gem::Version.new("0.10.0")
32
+
33
+ # Languages to set up at boot, from KOTOSHU_SERVER_LANGUAGES
34
+ # (space separated). Single source of truth for the env var: the
35
+ # boot pre-warm and /v1/health both read it here.
36
+ #
37
+ # @return [Array<String>] Language codes to pre-warm
38
+ def self.configured_languages
39
+ ENV.fetch("KOTOSHU_SERVER_LANGUAGES", "en").split
40
+ end
41
+
42
+ # Languages to set up with spelling + semantic model at boot,
43
+ # from KOTOSHU_SERVER_MODEL_LANGS (space separated, e.g. "en de").
44
+ # Empty when unset — boot-time opt-in only, never an implicit
45
+ # download (the gem's two-stage promise).
46
+ #
47
+ # @return [Array<String>] Language codes to set up with a model
48
+ def self.model_languages
49
+ ENV.fetch("KOTOSHU_SERVER_MODEL_LANGS", "").split
50
+ end
51
+
52
+ # Model tier used for setup and resolution, from
53
+ # KOTOSHU_SERVER_MODEL_TIER. Defaults to "fluency", the
54
+ # ecosystem default (owner decision 2026-09-04).
55
+ #
56
+ # @return [String] "full", "fluency", or "mini"
57
+ def self.model_tier
58
+ ENV.fetch("KOTOSHU_SERVER_MODEL_TIER", "fluency")
59
+ end
60
+
61
+ # Whether the installed kotoshu gem carries the semantic-model
62
+ # surface the server wires against (0.7.0+).
63
+ #
64
+ # @return [Boolean]
65
+ def self.semantic_models_supported?
66
+ Gem::Version.new(Kotoshu::VERSION) >= MODEL_MIN_KOTOSHU
67
+ end
68
+
69
+ # Whether a semantic model is actually set up server-side for a
70
+ # language (cache-only lookup; false on gems without model
71
+ # support, so the /v1/check default flag never turns models on
72
+ # underneath an old install).
73
+ #
74
+ # @param language [String] Language code
75
+ # @return [Boolean]
76
+ def self.model_available?(language)
77
+ return false unless semantic_models_supported?
78
+
79
+ Kotoshu.setup?(language.to_sym, :model)
80
+ rescue StandardError
81
+ false
82
+ end
83
+
84
+ # Validate the semantic-model environment before anything is
85
+ # set up. Raises {ModelConfigError} with an actionable message
86
+ # when KOTOSHU_SERVER_MODEL_LANGS is set but the installed
87
+ # kotoshu gem predates model support, or the configured tier is
88
+ # unknown. A no-op when the vars are unset — boot behavior is
89
+ # then exactly today's.
90
+ #
91
+ # @raise [ModelConfigError] on an unsatisfiable configuration
92
+ # @return [void]
93
+ def self.validate_model_config!
94
+ return if model_languages.empty?
95
+
96
+ unless semantic_models_supported?
97
+ raise ModelConfigError,
98
+ "KOTOSHU_SERVER_MODEL_LANGS requires kotoshu >= 0.7.0 " \
99
+ "(installed: #{Kotoshu::VERSION}). Semantic models ship in " \
100
+ "kotoshu 0.7.0; unset KOTOSHU_SERVER_MODEL_LANGS or upgrade " \
101
+ "the kotoshu gem."
102
+ end
103
+
104
+ begin
105
+ Kotoshu::Cache::ModelCache.normalize_tier(model_tier)
106
+ rescue ArgumentError => e
107
+ raise ModelConfigError, "KOTOSHU_SERVER_MODEL_TIER: #{e.message}"
108
+ end
109
+ end
110
+
111
+ # Synchronously set up the given languages (downloads on a cold
112
+ # or expired cache). Runs on the pre-warm thread; also usable
113
+ # directly by embedders that want a blocking warm-up.
114
+ #
115
+ # Languages listed in KOTOSHU_SERVER_MODEL_LANGS are set up with
116
+ # spelling + model (at KOTOSHU_SERVER_MODEL_TIER) in one setup
117
+ # call; the rest get spelling only, as before. Both lists are
118
+ # unioned, and setup is idempotent, so a language in both is
119
+ # set up once with the model.
120
+ #
121
+ # @param languages [Array<String>] Language codes
122
+ # @return [void]
123
+ def self.prewarm!(languages)
124
+ logger = Logger.new($stderr)
125
+ model_langs = model_languages
126
+ tier = model_tier
127
+ (languages | model_langs).each do |lang|
128
+ with_model = model_langs.include?(lang)
129
+ logger.info("pre-warming #{lang}#{with_model ? " (spelling + model, #{tier} tier)" : ""}")
130
+ begin
131
+ if with_model
132
+ Kotoshu.setup(lang.to_sym, want: %i[spelling model], tier: tier)
133
+ else
134
+ Kotoshu.setup(lang.to_sym)
135
+ end
136
+ logger.info("pre-warm #{lang} complete")
137
+ rescue StandardError => e
138
+ logger.warn("pre-warm #{lang} failed: #{e.message}")
139
+ end
140
+ end
141
+ end
142
+
143
+ # Start pre-warming in a detached background thread and return
144
+ # immediately. The server must bind and serve within seconds of
145
+ # boot, while resource setup can block on the network
146
+ # (download retries, and DNS resolution that no Net::HTTP
147
+ # timeout bounds) — so setup never runs on the boot path.
148
+ # Progress and completion are logged; /v1/health reports
149
+ # per-language readiness while it runs.
150
+ #
151
+ # @param languages [Array<String>] Language codes
152
+ # @return [Thread] the detached pre-warm thread
153
+ def self.prewarm_async!(languages)
154
+ validate_model_config! # fail fast, on the caller's thread
155
+ Thread.new do
156
+ Thread.current.name = "kotoshu-server-prewarm"
157
+ prewarm!(languages)
158
+ end
159
+ end
160
+
161
+ # ---- Language detection (/v1/detect) ----
162
+
163
+ @lid_setup_mutex = Mutex.new
164
+ @lid_setup_attempted = false
165
+
166
+ # Whether the installed kotoshu gem carries the lid-176
167
+ # detection surface (0.10.0+): Kotoshu.detect_language returning
168
+ # a Language::Detection, plus setup_lid / LidDetector.available?.
169
+ #
170
+ # @return [Boolean]
171
+ def self.lid_supported?
172
+ Gem::Version.new(Kotoshu::VERSION) >= DETECT_MIN_KOTOSHU
173
+ end
174
+
175
+ # Whether KOTOSHU_DETECT=heuristic pins /v1/detect to the
176
+ # 7-language heuristic, bypassing the gem engine selection (the
177
+ # 0.1.1 behavior, e.g. to reproduce earlier results).
178
+ #
179
+ # @return [Boolean]
180
+ def self.heuristic_detect?
181
+ ENV.fetch("KOTOSHU_DETECT", nil) == "heuristic"
182
+ end
183
+
184
+ # Detect for /v1/detect: the language code, its score, and the
185
+ # engine that served — "lid-176" or "heuristic".
186
+ #
187
+ # Engine selection, in order:
188
+ # 1. KOTOSHU_DETECT=heuristic — the heuristic, pinned.
189
+ # 2. kotoshu < 0.10.0 — the heuristic, as in 0.1.1; the server
190
+ # keeps working on old gems.
191
+ # 3. Otherwise the gem's own selection: lid-176 through the
192
+ # native extension when it is built, the backend is not
193
+ # explicitly ruby, and the artifact pair is cached; the
194
+ # heuristic otherwise (KOTOSHU_BACKEND=ruby, pure-Ruby
195
+ # install, missing model).
196
+ #
197
+ # The lid model is set up lazily on the first detect reaching
198
+ # case 3 — one download, off the boot path, honoring
199
+ # KOTOSHU_OFFLINE through the gem; a failed setup (offline with
200
+ # no cache, no registry entry, checksum mismatch) logs once and
201
+ # detection answers from whatever the gem can serve.
202
+ #
203
+ # @param text [String] the text to analyze
204
+ # @return [Array(String, Float, String)] code (nil when the
205
+ # heuristic is uncertain), score in [0, 1], engine name
206
+ def self.detect_language_with_engine(text)
207
+ return heuristic_detection(text) if heuristic_detect? || !lid_supported?
208
+
209
+ ensure_lid_setup!
210
+ detection = Kotoshu.detect_language(text)
211
+ [detection.code, detection.score, lid_engine]
212
+ end
213
+
214
+ # The heuristic detection (Language::Detector, the same engine
215
+ # /v1/detect served in 0.1.1, unchanged across gem versions).
216
+ #
217
+ # @param text [String]
218
+ # @return [Array(String, Float, String)]
219
+ def self.heuristic_detection(text)
220
+ code, confidence = Kotoshu::Language::Detector.detect_with_confidence(text)
221
+ [code, confidence, "heuristic"]
222
+ end
223
+
224
+ # Which engine the gem serves detect from: lid-176 when the
225
+ # native model is loadable, the heuristic otherwise. Asked
226
+ # after a detect, so it always names the engine behind the
227
+ # returned code.
228
+ #
229
+ # @return [String]
230
+ def self.lid_engine
231
+ Kotoshu::Language::LidDetector.available? ? "lid-176" : "heuristic"
232
+ end
233
+
234
+ # Run Kotoshu.setup_lid once per process, on the first detect.
235
+ # Idempotent in the gem; the once-guard keeps concurrent
236
+ # requests from racing the download. Any failure is logged and
237
+ # swallowed — detect then falls back inside the gem to whatever
238
+ # is cached (typically the heuristic).
239
+ #
240
+ # @return [void]
241
+ def self.ensure_lid_setup!
242
+ @lid_setup_mutex.synchronize do
243
+ return if @lid_setup_attempted
244
+
245
+ @lid_setup_attempted = true
246
+ Kotoshu.setup_lid
247
+ end
248
+ rescue StandardError => e
249
+ Logger.new($stderr).warn(
250
+ "lid model setup failed, falling back to the heuristic: #{e.class}: #{e.message}"
251
+ )
252
+ end
253
+
254
+ # ---- Semantic analyzers (memoized per language + model file) ----
255
+
256
+ @semantic_analyzers = {}
257
+ @semantic_analyzers_mutex = Mutex.new
258
+
259
+ # A memoized {Kotoshu::Analyzers::SemanticAnalyzer} for the model
260
+ # file the ResourceManager resolved for `language`. Loading a
261
+ # model (vocab + ONNX session) is expensive; one analyzer per
262
+ # language is kept for the process lifetime. A changed model
263
+ # file (different path) loads a fresh analyzer.
264
+ #
265
+ # @param language [String] Language code
266
+ # @param model_info [Hash] bundle.model from ResourceManager
267
+ # (needs :model_path)
268
+ # @return [Kotoshu::Analyzers::SemanticAnalyzer]
269
+ def self.semantic_analyzer_for(language, model_info)
270
+ path = model_info[:model_path]
271
+ @semantic_analyzers_mutex.synchronize do
272
+ @semantic_analyzers[[language.to_s, path]] ||= begin
273
+ model = Kotoshu::Models::OnnxModel.from_file(path, language_code: language.to_s)
274
+ Kotoshu::Analyzers::SemanticAnalyzer.new(model)
275
+ end
276
+ end
277
+ end
278
+
279
+ configure do
280
+ set :logging, false
281
+ set :show_exceptions, false
282
+
283
+ app_logger = Logger.new($stderr)
284
+ app_logger.level = ENV.fetch("KOTOSHU_SERVER_LOG_LEVEL", Logger::INFO)
285
+ set :app_logger, app_logger
286
+ end
287
+
288
+ # ---- Endpoints ----
289
+
290
+ get "/" do
291
+ content_type :json
292
+ {
293
+ name: "kotoshu-server",
294
+ version: VERSION,
295
+ kotoshu_version: Kotoshu::VERSION,
296
+ docs: "/v1/health, /v1/version, /v1/languages, /v1/check, /v1/suggest, /v1/detect"
297
+ }.to_json
298
+ end
299
+
300
+ get "/v1/health" do
301
+ content_type :json
302
+ ready = self.class.configured_languages.map { |l| [l, Kotoshu.setup?(l.to_sym, :spelling)] }.to_h
303
+ { status: "ok", ready: ready, timestamp: Time.now.utc.iso8601 }.to_json
304
+ end
305
+
306
+ get "/v1/version" do
307
+ content_type :json
308
+ { server: VERSION, kotoshu: Kotoshu::VERSION, ruby: RUBY_VERSION }.to_json
309
+ end
310
+
311
+ get "/v1/languages" do
312
+ content_type :json
313
+ cached = Kotoshu.languages_setup
314
+ model = cached.to_h { |lang| [lang, self.class.model_available?(lang)] }
315
+ { cached: cached, supported: cached, model: model }.to_json
316
+ end
317
+
318
+ post "/v1/check" do
319
+ body = parse_json_body(request.body.read)
320
+ text = body["text"]
321
+ language = body["language"] || ENV.fetch("KOTOSHU_SERVER_DEFAULT_LANG", "en")
322
+ format_hint = body["format"] || "full"
323
+ want_model = model_request_flag(body["model"], language)
324
+
325
+ halt_with_error(400, "missing 'text'") unless text.is_a?(String)
326
+
327
+ want = want_model ? %i[spelling model] : %i[spelling]
328
+ result = with_resource(language, want: want) do |checker, bundle|
329
+ if want_model && bundle.model.nil?
330
+ # The gem resolves a nil model for languages that cannot
331
+ # have one; an explicit request must not silently degrade.
332
+ halt 422, { "Content-Type" => "application/json" },
333
+ {
334
+ error: "resource_not_setup",
335
+ message: "no semantic model set up for language '#{language}'",
336
+ hint: "list it in KOTOSHU_SERVER_MODEL_LANGS and restart"
337
+ }.to_json
338
+ end
339
+ check_result = checker.check(text)
340
+ want_model ? rerank_semantic(language, check_result, text, bundle.model) : check_result
341
+ end
342
+
343
+ content_type :json
344
+ case format_hint
345
+ when "errors" then serialize_errors(result).to_json
346
+ else serialize_full(result).to_json
347
+ end
348
+ end
349
+
350
+ post "/v1/suggest" do
351
+ body = parse_json_body(request.body.read)
352
+ word = body["word"]
353
+ language = body["language"] || ENV.fetch("KOTOSHU_SERVER_DEFAULT_LANG", "en")
354
+ max = body["max"]&.to_i
355
+
356
+ halt_with_error(400, "missing 'word'") unless word.is_a?(String)
357
+
358
+ result = with_resource(language) do |checker|
359
+ checker.suggest(word, max_suggestions: max)
360
+ end
361
+
362
+ content_type :json
363
+ { word: word, suggestions: serialize_suggestions(result) }.to_json
364
+ end
365
+
366
+ post "/v1/detect" do
367
+ body = parse_json_body(request.body.read)
368
+ text = body["text"]
369
+ halt_with_error(400, "missing 'text'") unless text.is_a?(String)
370
+
371
+ language, confidence, engine = self.class.detect_language_with_engine(text)
372
+ content_type :json
373
+ { language: language, confidence: confidence, engine: engine }.to_json
374
+ end
375
+
376
+ # ---- Error handling ----
377
+
378
+ error Kotoshu::Server::ModelConfigError do
379
+ status 500
380
+ content_type :json
381
+ { error: "model_config", message: env["sinatra.error"].message }.to_json
382
+ end
383
+
384
+ error Kotoshu::Models::OnnxModel::OnnxUnavailable do
385
+ status 503
386
+ content_type :json
387
+ { error: "onnx_unavailable", message: env["sinatra.error"].message,
388
+ hint: "install the onnxruntime gem, or stop setting model langs" }.to_json
389
+ end
390
+
391
+ error Kotoshu::ResourceNotSetupError do
392
+ e = env["sinatra.error"]
393
+ status 422
394
+ content_type :json
395
+ { error: "resource_not_setup", message: e.message,
396
+ hint: "POST /v1/admin/setup with {language: \"...\"} (admin only)" }.to_json
397
+ end
398
+
399
+ error JSON::ParserError do
400
+ status 400
401
+ content_type :json
402
+ { error: "invalid_json", message: env["sinatra.error"].message }.to_json
403
+ end
404
+
405
+ error StandardError do
406
+ e = env["sinatra.error"]
407
+ settings.app_logger&.error("unhandled: #{e.class}: #{e.message}")
408
+ status 500
409
+ content_type :json
410
+ { error: "internal", message: e.message }.to_json
411
+ end
412
+
413
+ # ---- Helpers ----
414
+
415
+ helpers do
416
+ def parse_json_body(body)
417
+ return {} if body.nil? || body.empty?
418
+
419
+ JSON.parse(body)
420
+ end
421
+
422
+ def halt_with_error(code, message)
423
+ halt code, { "Content-Type" => "application/json" },
424
+ { error: "invalid_request", message: message }.to_json
425
+ end
426
+
427
+ def with_resource(language, want: %i[spelling])
428
+ Kotoshu.reset_spellchecker if Kotoshu.instance_variable_get(:@spellcheckers).nil?
429
+ bundle = Kotoshu::ResourceManager.resolve(language: language, want: want)
430
+ checker = Kotoshu::Spellchecker.new(resource_bundle: bundle)
431
+ yield checker, bundle
432
+ end
433
+
434
+ # Resolve the effective "model" flag for /v1/check: an explicit
435
+ # request value wins; otherwise the default is whether the
436
+ # language has a model set up server-side (cache-only). An
437
+ # explicit true on a kotoshu install without model support is
438
+ # a 503 naming the requirement, not a silent dictionary-only
439
+ # answer.
440
+ #
441
+ # @param value [Boolean, nil] the request's "model" field
442
+ # @param language [String] resolved request language
443
+ # @return [Boolean]
444
+ def model_request_flag(value, language)
445
+ case value
446
+ when nil then self.class.model_available?(language)
447
+ when true
448
+ unless self.class.semantic_models_supported?
449
+ halt 503, { "Content-Type" => "application/json" },
450
+ {
451
+ error: "model_unsupported",
452
+ message: "semantic models require kotoshu >= 0.7.0 " \
453
+ "(installed: #{Kotoshu::VERSION})"
454
+ }.to_json
455
+ end
456
+ true
457
+ when false then false
458
+ else halt_with_error(400, "'model' must be true or false")
459
+ end
460
+ end
461
+
462
+ # Rerank a traditional check result with the gem's semantic
463
+ # analyzer (dictionary verdict first, neural rerank for
464
+ # uncertain candidates — the gem's confidence cascade decides
465
+ # per word whether the ONNX rerank runs at all). Semantic
466
+ # candidates lead the merged suggestion list; traditional
467
+ # candidates follow, deduplicated by word. Any per-word
468
+ # analyzer failure keeps the traditional suggestions for that
469
+ # word — one bad word never fails the request.
470
+ #
471
+ # @param language [String] Language code
472
+ # @param result [Kotoshu::Models::Result::DocumentResult]
473
+ # @param text [String] the checked text (context source)
474
+ # @param model_info [Hash, nil] bundle.model from the resolve
475
+ # @return [Kotoshu::Models::Result::DocumentResult]
476
+ def rerank_semantic(language, result, text, model_info)
477
+ return result if model_info.nil? || model_info[:model_path].nil?
478
+ errors = Array(result.errors)
479
+ return result if errors.empty?
480
+
481
+ analyzer = self.class.semantic_analyzer_for(language, model_info)
482
+ cascade = Kotoshu::Suggestions::SemanticCascade.from_configuration(Kotoshu.configuration)
483
+
484
+ reranked = errors.map { |error| rerank_error(analyzer, cascade, text, error) }
485
+
486
+ Kotoshu::Models::Result::DocumentResult.new(
487
+ file: result.file,
488
+ errors: reranked,
489
+ suppressed_errors: result.suppressed_errors,
490
+ word_count: result.word_count
491
+ )
492
+ end
493
+
494
+ # Rerank one error's suggestions, or return the error unchanged
495
+ # when the cascade skips it or the analyzer has nothing better
496
+ # (including an analyzer failure on this word — traditional
497
+ # suggestions survive).
498
+ #
499
+ # @param analyzer [Kotoshu::Analyzers::SemanticAnalyzer]
500
+ # @param cascade [Kotoshu::Suggestions::SemanticCascade]
501
+ # @param text [String] the checked text
502
+ # @param error [Kotoshu::Models::Result::WordResult]
503
+ # @return [Kotoshu::Models::Result::WordResult]
504
+ def rerank_error(analyzer, cascade, text, error)
505
+ traditional = error.suggestions.to_a
506
+ return error if traditional.empty? || cascade.skip?(traditional)
507
+
508
+ semantic = begin
509
+ analyzer.suggest_corrections(error.word, context: semantic_context(text, error))
510
+ rescue StandardError
511
+ []
512
+ end
513
+ return error if semantic.empty?
514
+
515
+ Kotoshu::Models::Result::WordResult.new(
516
+ word: error.word,
517
+ correct: false,
518
+ position: error.position,
519
+ suggestions: merge_suggestions(semantic, traditional),
520
+ suppressed: error.suppressed,
521
+ suppressed_by: error.suppressed_by
522
+ )
523
+ end
524
+
525
+ # Merge analyzer suggestions (semantic, context-ranked) ahead
526
+ # of the traditional ones, deduplicated by word, mapped to the
527
+ # wire Suggestion shape /v1/check already serializes.
528
+ #
529
+ # @param semantic [Array<Kotoshu::Models::Suggestion>]
530
+ # @param traditional [Array<Kotoshu::Suggestions::Suggestion>]
531
+ # @return [Array<Kotoshu::Suggestions::Suggestion>]
532
+ def merge_suggestions(semantic, traditional)
533
+ merged = []
534
+ seen = {}
535
+ semantic.each do |s|
536
+ next if seen[s.word]
537
+
538
+ seen[s.word] = true
539
+ merged << Kotoshu::Suggestions::Suggestion.new(
540
+ word: s.word,
541
+ distance: s.metadata[:distance].to_i,
542
+ confidence: s.confidence,
543
+ source: :semantic
544
+ )
545
+ end
546
+ traditional.each do |s|
547
+ next if seen[s.word]
548
+
549
+ seen[s.word] = true
550
+ merged << s
551
+ end
552
+ merged
553
+ end
554
+
555
+ # The text window around an error, in the shape the analyzer's
556
+ # own context ranking expects ({Models::Context} before /
557
+ # current / after slices; nil when the error has no position).
558
+ #
559
+ # @param text [String] the checked text
560
+ # @param error [Kotoshu::Models::Result::WordResult]
561
+ # @return [Kotoshu::Models::Context, nil]
562
+ def semantic_context(text, error)
563
+ pos = error.position
564
+ return nil unless pos.is_a?(Integer)
565
+
566
+ window = 32
567
+ word_end = pos + error.word.length
568
+ Kotoshu::Models::Context.new(
569
+ before: text[[pos - window, 0].max...pos] || "",
570
+ current: text[pos...word_end] || "",
571
+ after: text[word_end...(word_end + window)] || "",
572
+ location: nil
573
+ )
574
+ end
575
+
576
+ def serialize_full(result)
577
+ {
578
+ file: result.file,
579
+ word_count: result.word_count,
580
+ errors: serialize_errors(result)
581
+ }
582
+ end
583
+
584
+ def serialize_errors(result)
585
+ (result.errors || []).map do |err|
586
+ {
587
+ word: err.word,
588
+ position: err.position,
589
+ suggestions: err.suggestions.map { |s| serialize_suggestion(s) }
590
+ }
591
+ end
592
+ end
593
+
594
+ def serialize_suggestions(set)
595
+ return [] unless set.respond_to?(:suggestions)
596
+
597
+ set.suggestions.map { |s| serialize_suggestion(s) }
598
+ end
599
+
600
+ def serialize_suggestion(s)
601
+ {
602
+ word: s.word,
603
+ distance: s.distance,
604
+ confidence: s.confidence,
605
+ source: s.source
606
+ }
607
+ end
608
+ end
609
+ end
610
+ end
611
+ end
612
+
613
+ require "time"
@@ -0,0 +1,7 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Kotoshu
4
+ module Server
5
+ VERSION = "0.1.2"
6
+ end
7
+ end
@@ -0,0 +1,27 @@
1
+ # frozen_string_literal: true
2
+
3
+ require_relative "server/version"
4
+ require_relative "server/app"
5
+
6
+ module Kotoshu
7
+ module Server
8
+ # Boot the server.
9
+ #
10
+ # Pre-warm starts in a detached thread and never blocks: Puma binds
11
+ # and serves within seconds regardless of how long resource setup
12
+ # (cold-cache downloads) takes. KOTOSHU_SERVER_LAZY=1 skips the
13
+ # pre-warm entirely; /v1/health reports readiness while it runs.
14
+ #
15
+ # Semantic-model env vars (KOTOSHU_SERVER_MODEL_LANGS / _TIER) are
16
+ # validated before the bind, so an unsatisfiable model config
17
+ # fails fast with a clear error instead of surprising per-request
18
+ # failures. Unset, boot is exactly as before.
19
+ def self.run!(port: ENV.fetch("KOTOSHU_SERVER_PORT", 9292), bind: ENV.fetch("KOTOSHU_SERVER_BIND", "0.0.0.0"))
20
+ App.validate_model_config!
21
+ unless ENV.fetch("KOTOSHU_SERVER_LAZY", "0") == "1"
22
+ App.prewarm_async!(App.configured_languages)
23
+ end
24
+ App.run!(port: port.to_i, bind: bind)
25
+ end
26
+ end
27
+ end
data/openapi.yaml ADDED
@@ -0,0 +1,251 @@
1
+ openapi: 3.1.0
2
+ info:
3
+ title: Kotoshu HTTP API
4
+ version: 0.1.0
5
+ description: |
6
+ Self-hostable HTTP API wrapping the Kotoshu spell checker.
7
+ Used by the Python, JavaScript, and Go SDKs.
8
+ license:
9
+ name: BSD-2-Clause
10
+ url: https://github.com/kotoshu/kotoshu/blob/main/LICENSE
11
+ servers:
12
+ - url: http://localhost:9292
13
+ description: Local server
14
+ paths:
15
+ /:
16
+ get:
17
+ summary: Service metadata
18
+ responses:
19
+ "200":
20
+ description: OK
21
+ content:
22
+ application/json:
23
+ schema:
24
+ $ref: "#/components/schemas/Metadata"
25
+ /v1/health:
26
+ get:
27
+ summary: Liveness probe + cache state
28
+ responses:
29
+ "200":
30
+ description: OK
31
+ content:
32
+ application/json:
33
+ schema:
34
+ $ref: "#/components/schemas/Health"
35
+ /v1/version:
36
+ get:
37
+ summary: Server + kotoshu versions
38
+ responses:
39
+ "200":
40
+ description: OK
41
+ content:
42
+ application/json:
43
+ schema:
44
+ $ref: "#/components/schemas/Version"
45
+ /v1/languages:
46
+ get:
47
+ summary: Languages with cached dictionaries and model availability
48
+ responses:
49
+ "200":
50
+ description: OK
51
+ content:
52
+ application/json:
53
+ schema:
54
+ $ref: "#/components/schemas/Languages"
55
+ /v1/check:
56
+ post:
57
+ summary: Check text for spelling errors
58
+ description: |
59
+ Optional `"model": true` reranks each error's suggestions with
60
+ the semantic model set up server-side for the language
61
+ (requires KOTOSHU_SERVER_MODEL_LANGS at boot, kotoshu >= 0.7).
62
+ The default is whether a model is set up for the language.
63
+ requestBody:
64
+ required: true
65
+ content:
66
+ application/json:
67
+ schema:
68
+ $ref: "#/components/schemas/CheckRequest"
69
+ responses:
70
+ "200":
71
+ description: Result
72
+ content:
73
+ application/json:
74
+ schema:
75
+ $ref: "#/components/schemas/DocumentResult"
76
+ "400":
77
+ $ref: "#/components/responses/BadRequest"
78
+ "422":
79
+ $ref: "#/components/responses/ResourceNotSetup"
80
+ "503":
81
+ $ref: "#/components/responses/ModelUnsupported"
82
+ /v1/suggest:
83
+ post:
84
+ summary: Get suggestions for a single word
85
+ requestBody:
86
+ required: true
87
+ content:
88
+ application/json:
89
+ schema:
90
+ $ref: "#/components/schemas/SuggestRequest"
91
+ responses:
92
+ "200":
93
+ description: Suggestions
94
+ content:
95
+ application/json:
96
+ schema:
97
+ $ref: "#/components/schemas/Suggestions"
98
+ "400":
99
+ $ref: "#/components/responses/BadRequest"
100
+ /v1/detect:
101
+ post:
102
+ summary: Detect the language of text
103
+ requestBody:
104
+ required: true
105
+ content:
106
+ application/json:
107
+ schema:
108
+ $ref: "#/components/schemas/DetectRequest"
109
+ responses:
110
+ "200":
111
+ description: Detected language
112
+ content:
113
+ application/json:
114
+ schema:
115
+ $ref: "#/components/schemas/Detection"
116
+ "400":
117
+ $ref: "#/components/responses/BadRequest"
118
+ components:
119
+ responses:
120
+ BadRequest:
121
+ description: Malformed request
122
+ content:
123
+ application/json:
124
+ schema:
125
+ $ref: "#/components/schemas/Error"
126
+ ResourceNotSetup:
127
+ description: Language not set up on the server
128
+ content:
129
+ application/json:
130
+ schema:
131
+ $ref: "#/components/schemas/Error"
132
+ ModelUnsupported:
133
+ description: Model requested but the installed kotoshu gem is older than 0.7.0
134
+ content:
135
+ application/json:
136
+ schema:
137
+ $ref: "#/components/schemas/Error"
138
+ schemas:
139
+ Metadata:
140
+ type: object
141
+ required: [name, version, kotoshu_version]
142
+ properties:
143
+ name: { type: string, example: kotoshu-server }
144
+ version: { type: string, example: 0.1.0 }
145
+ kotoshu_version: { type: string, example: 0.6.0 }
146
+ docs: { type: string }
147
+ Health:
148
+ type: object
149
+ required: [status, ready]
150
+ properties:
151
+ status: { type: string, example: ok }
152
+ ready:
153
+ type: object
154
+ additionalProperties: { type: boolean }
155
+ example: { en: true }
156
+ timestamp: { type: string, format: date-time }
157
+ Version:
158
+ type: object
159
+ required: [server, kotoshu]
160
+ properties:
161
+ server: { type: string }
162
+ kotoshu: { type: string }
163
+ ruby: { type: string }
164
+ Languages:
165
+ type: object
166
+ required: [cached, model]
167
+ properties:
168
+ cached:
169
+ type: array
170
+ items: { type: string }
171
+ supported:
172
+ type: array
173
+ items: { type: string }
174
+ model:
175
+ type: object
176
+ description: Whether a semantic model is set up per cached language
177
+ additionalProperties: { type: boolean }
178
+ example: { en: true, de: false }
179
+ CheckRequest:
180
+ type: object
181
+ required: [text]
182
+ properties:
183
+ text: { type: string }
184
+ language: { type: string, default: en }
185
+ format: { type: string, enum: [full, errors], default: full }
186
+ model:
187
+ type: boolean
188
+ description: |
189
+ Rerank suggestions with the language's semantic model.
190
+ Default is whether a model is set up server-side. true on a
191
+ language without a cached model returns 422.
192
+ SuggestRequest:
193
+ type: object
194
+ required: [word]
195
+ properties:
196
+ word: { type: string }
197
+ language: { type: string, default: en }
198
+ max: { type: integer, minimum: 1, maximum: 50 }
199
+ DetectRequest:
200
+ type: object
201
+ required: [text]
202
+ properties:
203
+ text: { type: string }
204
+ DocumentResult:
205
+ type: object
206
+ properties:
207
+ file: { type: string, nullable: true }
208
+ word_count: { type: integer }
209
+ errors:
210
+ type: array
211
+ items: { $ref: "#/components/schemas/WordError" }
212
+ WordError:
213
+ type: object
214
+ properties:
215
+ word: { type: string }
216
+ position: { type: integer, nullable: true }
217
+ suggestions:
218
+ type: array
219
+ items: { $ref: "#/components/schemas/Suggestion" }
220
+ Suggestions:
221
+ type: object
222
+ properties:
223
+ word: { type: string }
224
+ suggestions:
225
+ type: array
226
+ items: { $ref: "#/components/schemas/Suggestion" }
227
+ Suggestion:
228
+ type: object
229
+ properties:
230
+ word: { type: string }
231
+ distance: { type: integer }
232
+ confidence: { type: number, format: float }
233
+ source: { type: string }
234
+ Detection:
235
+ type: object
236
+ properties:
237
+ language: { type: string, nullable: true }
238
+ confidence: { type: number, format: float }
239
+ engine:
240
+ type: string
241
+ enum: [lid-176, heuristic]
242
+ description: >
243
+ Which engine served the detection: the 176-language lid-176
244
+ model (kotoshu >= 0.10.0 with the native extension and the
245
+ model set up) or the 7-language heuristic.
246
+ Error:
247
+ type: object
248
+ properties:
249
+ error: { type: string }
250
+ message: { type: string }
251
+ hint: { type: string }
metadata CHANGED
@@ -1,13 +1,14 @@
1
1
  --- !ruby/object:Gem::Specification
2
2
  name: kotoshu-server
3
3
  version: !ruby/object:Gem::Version
4
- version: 0.1.0
4
+ version: 0.1.2
5
5
  platform: ruby
6
6
  authors:
7
7
  - Ribose Inc.
8
+ autorequire:
8
9
  bindir: exe
9
10
  cert_chain: []
10
- date: 1980-01-02 00:00:00.000000000 Z
11
+ date: 2026-09-07 00:00:00.000000000 Z
11
12
  dependencies:
12
13
  - !ruby/object:Gem::Dependency
13
14
  name: kotoshu
@@ -69,10 +70,18 @@ description: Self-hostable HTTP API that exposes Kotoshu's check / suggest / det
69
70
  over JSON. Designed as the deployment surface for non-Ruby SDKs (Python, JS, Go).
70
71
  email:
71
72
  - open.source@ribose.com
72
- executables: []
73
+ executables:
74
+ - kotoshu-server
73
75
  extensions: []
74
76
  extra_rdoc_files: []
75
- files: []
77
+ files:
78
+ - LICENSE
79
+ - README.md
80
+ - exe/kotoshu-server
81
+ - lib/kotoshu/server.rb
82
+ - lib/kotoshu/server/app.rb
83
+ - lib/kotoshu/server/version.rb
84
+ - openapi.yaml
76
85
  homepage: https://github.com/kotoshu/kotoshu-server
77
86
  licenses:
78
87
  - BSD-2-Clause
@@ -80,6 +89,7 @@ metadata:
80
89
  homepage_uri: https://github.com/kotoshu/kotoshu-server
81
90
  source_code_uri: https://github.com/kotoshu/kotoshu-server/tree/main
82
91
  rubygems_mfa_required: 'true'
92
+ post_install_message:
83
93
  rdoc_options: []
84
94
  require_paths:
85
95
  - lib
@@ -94,7 +104,8 @@ required_rubygems_version: !ruby/object:Gem::Requirement
94
104
  - !ruby/object:Gem::Version
95
105
  version: '0'
96
106
  requirements: []
97
- rubygems_version: 4.0.16
107
+ rubygems_version: 3.5.22
108
+ signing_key:
98
109
  specification_version: 4
99
110
  summary: HTTP API server wrapping the Kotoshu spell checker
100
111
  test_files: []