arcane-os 0.5.10 → 0.5.11

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/CHANGELOG.md +19 -0
  2. package/README.md +117 -26
  3. package/browser-runtime/ai/browser-speech-providers.mjs +1 -1
  4. package/docs/architecture.md +303 -0
  5. package/docs/compatibility.md +38 -0
  6. package/docs/event-manager.md +263 -0
  7. package/docs/platform-targets.md +104 -0
  8. package/docs/publishing.md +126 -0
  9. package/docs/reference/README.md +206 -0
  10. package/docs/reference/ai/browser-speech.md +813 -0
  11. package/docs/reference/ai/browser-wasm.md +637 -0
  12. package/docs/reference/ai/twin-cloud.md +156 -0
  13. package/docs/reference/arcane-ollama.md +288 -0
  14. package/docs/reference/availability-and-normalization.md +224 -0
  15. package/docs/reference/behavioral-testing.md +129 -0
  16. package/docs/reference/cli.md +820 -0
  17. package/docs/reference/core/README.md +61 -0
  18. package/docs/reference/core/arcane-ai-contracts.md +907 -0
  19. package/docs/reference/core/arcane-api.md +601 -0
  20. package/docs/reference/core/arcane-entities.md +59 -0
  21. package/docs/reference/core/arcane-events.md +134 -0
  22. package/docs/reference/core/ollama-module.md +181 -0
  23. package/docs/reference/core/reference/arcane-api/ai-and-ollama.md +1909 -0
  24. package/docs/reference/core/reference/arcane-api/applications-terminal-capabilities.md +1057 -0
  25. package/docs/reference/core/reference/arcane-api/core-and-events.md +320 -0
  26. package/docs/reference/core/reference/arcane-api/filesystem-storage-preferences-appearance.md +610 -0
  27. package/docs/reference/core/reference/arcane-api/namespaces.md +1157 -0
  28. package/docs/reference/core/reference/arcane-api/platform-installation-users-system.md +1423 -0
  29. package/docs/reference/core/reference/arcane-api/session-provisioning-diagnostics-development.md +315 -0
  30. package/docs/reference/event-manager.md +1409 -0
  31. package/docs/reference/inventory/package-api.json +3194 -0
  32. package/docs/reference/inventory/runtime-components.json +1015 -0
  33. package/docs/reference/inventory/runtime-entities.json +25 -0
  34. package/docs/reference/inventory/runtime-modules.json +1367 -0
  35. package/docs/reference/mail.md +309 -0
  36. package/docs/reference/protocols.md +749 -0
  37. package/docs/reference/runtime-components.md +1529 -0
  38. package/docs/reference/runtime-entities.md +305 -0
  39. package/docs/reference/runtime-modules.md +3275 -0
  40. package/docs/reference/sdk-api.md +6733 -0
  41. package/docs/roadmap.md +79 -0
  42. package/docs/work-amplification.md +66 -0
  43. package/examples/wasm-ai-demo/README.md +80 -0
  44. package/examples/wasm-ai-demo/app.js +787 -0
  45. package/examples/wasm-ai-demo/index.html +343 -0
  46. package/examples/wasm-ai-demo/profile-tools.js +217 -0
  47. package/examples/wasm-ai-demo/profiles/BOSS.Modelfile +106 -0
  48. package/examples/wasm-ai-demo/profiles/PreCrisis.Modelfile +693 -0
  49. package/examples/wasm-ai-demo/rag/boss-library.json +3006 -0
  50. package/examples/wasm-ai-demo/rag.js +295 -0
  51. package/examples/wasm-ai-demo/server.mjs +71 -0
  52. package/package.json +10 -1
  53. package/runtime/arcane/modules/AI.js +1 -1
  54. package/runtime/arcane/modules/AIProviderRuntime.js +26 -5
@@ -0,0 +1,813 @@
1
+ # Browser speech providers
2
+
3
+ `arcane-os/ai/browser-speech` is the browser-only SDK boundary for
4
+ caller-selected Whisper speech-to-text and Kokoro text-to-speech runtimes. It
5
+ provides artifact storage, live module routing, role Workers, provider/2
6
+ adapters, bounded parallel TTS synthesis, audio normalization, cancellation,
7
+ and cleanup.
8
+
9
+ ## Quick start: say one sentence
10
+
11
+ Use this in a browser application served by `arcane dev`, where the generated
12
+ import map resolves `arcane/AI` and `arcane/DBOPFS`. These browser modules are
13
+ not Node inference APIs. To create an application:
14
+
15
+ ```bash
16
+ npx arcane-os@0.5.11 new hello-speech --path ./hello-speech --target browser
17
+ cd hello-speech
18
+ npm install
19
+ npm run dev
20
+ ```
21
+
22
+ Keep the generated page's Arcane theme and import map. The examples below go in
23
+ `apps/hello-speech/modules/App.js` and the adjacent `speech-selection.js`.
24
+ Follow the development server's printed URL. Installing the SDK supplies its
25
+ provider, storage, and Worker code; it does not install a speech model or choose
26
+ an upstream speech runtime for your application.
27
+
28
+ First create **`speech-selection.js`**, the one application-owned configuration
29
+ file used throughout this guide. This concrete selection is also used by the
30
+ [maintained WASM voice-chat example](https://github.com/TheWizardNexus/arcane-os-sdk/tree/main/examples/wasm-ai-demo).
31
+ Your application owns these runtime/model versions, URLs, dtype, and voice.
32
+ Loading this selection uses those upstream publishers' downloads and caches.
33
+
34
+ ```javascript
35
+ export const speechSelection = {
36
+ model: {
37
+ id: 'onnx-community/Kokoro-82M-v1.0-ONNX',
38
+ repository: 'onnx-community/Kokoro-82M-v1.0-ONNX',
39
+ revision: '1939ad2a8e416c0acfeecc08a694d14ef25f2231',
40
+ dtype: 'q8',
41
+ defaultVoice: 'af_heart'
42
+ },
43
+ runtime: {
44
+ adapter: 'kokoro-js',
45
+ version: '1.2.1',
46
+ revision: '664c76a704021239ba59c84dcbaa4d3dece01fe9',
47
+ entry: 'kokoro.web.js',
48
+ wasmPaths: 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@3.5.1/dist/',
49
+ files: [{
50
+ path: 'kokoro.web.js',
51
+ url: 'https://cdn.jsdelivr.net/npm/kokoro-js@1.2.1/dist/kokoro.web.js',
52
+ mediaType: 'text/javascript'
53
+ }]
54
+ }
55
+ };
56
+ ```
57
+
58
+ Then use this **`App.js`**. The application creates and owns the DBOPFS
59
+ instance. Configuration selects the provider without loading it; the button
60
+ explicitly loads and unmutes TTS before requesting speech.
61
+
62
+ ```javascript
63
+ import arcaneThemeReady from 'arcane/ThemeBootstrap';
64
+ import AI, { AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL } from 'arcane/AI';
65
+ import DBOPFS from 'arcane/DBOPFS';
66
+ import { speechSelection } from './speech-selection.js';
67
+
68
+ await arcaneThemeReady;
69
+ const dbopfs = new DBOPFS();
70
+ await dbopfs.readyPromise;
71
+ const ai = new AI();
72
+
73
+ await ai.configureBrowserSpeech({
74
+ protocol: AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL,
75
+ id: 'hello-speech',
76
+ dbopfs,
77
+ tts: {
78
+ providerId: 'hello-kokoro',
79
+ model: speechSelection.model,
80
+ runtime: speechSelection.runtime,
81
+ offline: false
82
+ }
83
+ });
84
+
85
+ const speakButton = document.createElement('button');
86
+ speakButton.textContent = 'Load voice and say hello';
87
+ document.body.append(speakButton);
88
+ speakButton.addEventListener('click', async function sayHello() {
89
+ speakButton.disabled = true;
90
+ try {
91
+ await ai.setSpeechMuted(false); // Loads the selected TTS provider.
92
+ const complete = await ai.streamTTS('Hello from Arcane. ', true);
93
+ console.log('Speech preparation completed:', complete);
94
+ } catch (error) {
95
+ console.error(error.code, error.message);
96
+ } finally {
97
+ speakButton.disabled = false;
98
+ }
99
+ });
100
+ ```
101
+
102
+ The first user action may download the selected runtime, model, and voice.
103
+ The browser may require another audio-unlock gesture after a long load; the SDK
104
+ retains prepared audio for that gesture. `streamTTS()` prepares and schedules
105
+ playback; its promise is not proof that a listener heard the sound. It returns
106
+ `false` for muted, stopped, or failed work, and the SDK reports full synthesis
107
+ or playback failures in its console diagnostics and `ai-tts-failure` event.
108
+
109
+ To display complete high-level playback errors in this page, observe its
110
+ existing event. The listener belongs to this example's one `ai` instance:
111
+
112
+ ```javascript
113
+ const speechEvents = new AbortController();
114
+ window.addEventListener('ai-tts-failure', function reportSpeechFailure(event) {
115
+ if (event.detail.ai !== ai) return;
116
+ console.error(event.detail.error.code, event.detail.error.message);
117
+ }, { signal: speechEvents.signal });
118
+ ```
119
+
120
+ Call `speechEvents.abort()` when disposing that interface to remove the listener.
121
+
122
+ ## Four synthesis slots and exact-order playback
123
+
124
+ Capacity 4 means up to four segments synthesize at once. Segment 5 and later
125
+ wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out of
126
+ order, but playback waits for earlier segments and plays exact input order.
127
+ Each slot owns a Worker/model session, so raising capacity trades memory for
128
+ latency.
129
+
130
+ The high-level `AI` route owns the FIFO queue. Calling a low-level Kokoro
131
+ provider directly beyond its capacity returns `ARCANE_AI_PROVIDER_BUSY`.
132
+ Ready adjacent audio buffers use contiguous AudioContext scheduling. Browser
133
+ audio scheduling and selected WebGPU status do not prove physical GPU kernel
134
+ overlap or audio quality. LLM and Whisper/STT capacity remains one.
135
+
136
+ ## Stream chunks as they arrive
137
+
138
+ Use the configured `ai` created above. Run this snippet from an owned user
139
+ action, such as your Speak button, and catch errors with the earlier
140
+ `error.code` / `error.message` pattern. First load/unmute, then accept chunks.
141
+ Call `streamTTS(chunk)` inside the producer's chunk callback immediately;
142
+ waiting for each speech promise there would serialize synthesis. This tiny
143
+ example uses three arriving chunks. Replace the three `onTextChunk(...)` calls
144
+ with your actual text stream callback, and flush after that producer ends.
145
+
146
+ ```javascript
147
+ await ai.setSpeechMuted(false);
148
+ const pendingSpeech = [];
149
+
150
+ function onTextChunk(chunk) {
151
+ const pending = ai.streamTTS(chunk);
152
+ // Attach both handlers immediately, so later failure cannot be unhandled.
153
+ pendingSpeech.push(pending.then(
154
+ function speechPrepared(value) { return { status: 'fulfilled', value }; },
155
+ function speechRejected(reason) { return { status: 'rejected', reason }; }
156
+ ));
157
+ }
158
+
159
+ onTextChunk('First sentence. ');
160
+ onTextChunk('Second sentence. ');
161
+ onTextChunk('Third sentence.');
162
+
163
+ // After the text producer ends, flush trailing text and settle every call.
164
+ const finalPrepared = await ai.finishTTS();
165
+ const outcomes = await Promise.all(pendingSpeech);
166
+ for (const outcome of outcomes) {
167
+ if (outcome.status === 'rejected') {
168
+ console.error(outcome.reason.code, outcome.reason.message);
169
+ } else if (outcome.value === false) {
170
+ console.log('Speech was muted, stopped, or failed; inspect SDK diagnostics.');
171
+ }
172
+ }
173
+ if (finalPrepared === false) {
174
+ console.log('Final speech was muted, stopped, or failed; inspect SDK diagnostics.');
175
+ }
176
+ ```
177
+
178
+ The segmentation default uses sentence punctuation. To submit smaller complete
179
+ segments, call `ai.configureTTSSegmentation({punctuation:'any',wordCadence:4})`
180
+ before feeding the stream. Chunk boundaries themselves do not force a sentence
181
+ boundary; `finishTTS()` flushes any remaining text. It is not a playback-ended
182
+ notification. Do not mute or dispose immediately after it if playback should
183
+ continue.
184
+
185
+ ## Choose a device or reduce memory use
186
+
187
+ Omitting `tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`.
188
+ These are three alternative configurations, not a sequence of required loads:
189
+
190
+ ```javascript
191
+ async function selectSpeechExecution(execution) {
192
+ await ai.configureBrowserSpeech({
193
+ protocol: AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL,
194
+ id: 'hello-speech',
195
+ dbopfs,
196
+ tts: {
197
+ providerId: 'hello-kokoro',
198
+ model: speechSelection.model,
199
+ runtime: speechSelection.runtime,
200
+ offline: false,
201
+ execution
202
+ }
203
+ });
204
+ }
205
+
206
+ // Choose and call one from your application settings action:
207
+ // await selectSpeechExecution({ device: 'auto' });
208
+ // await selectSpeechExecution({ device: 'webgpu' });
209
+ // await selectSpeechExecution({ device: 'wasm', maxConcurrentRequests: 1 });
210
+ ```
211
+
212
+ The override accepts integers 1, 2, 3, or 4. `auto` tries a complete WebGPU pool
213
+ when available, then recreates a complete WASM pool if that load cannot finish.
214
+ Explicit `webgpu` reports a load error when unavailable; explicit `wasm` never
215
+ attempts WebGPU. The same application-selected model and dtype apply on both
216
+ devices. A configuration change leaves TTS muted; explicitly load/unmute again.
217
+
218
+ ## Inspect the requested and selected device
219
+
220
+ Request the execution projection on the existing public runtime status after
221
+ loading. This explicitly reads the selected provider's current report; ordinary
222
+ `status()` retains its existing sticky snapshot and identity. There is no
223
+ separate execution-state event subscription.
224
+
225
+ ```javascript
226
+ function printSpeechStatus() {
227
+ const status = ai.providerRuntime.status('tts', { execution: true });
228
+ const execution = status.execution;
229
+ console.log('TTS state:', status.state);
230
+ if (execution) {
231
+ console.log('Requested device:', execution.requestedDevice);
232
+ console.log('Selected device:', execution.selectedDevice);
233
+ console.log('Capacity:', execution.maxConcurrentRequests);
234
+ console.log('Active synthesis requests:', execution.activeRequestCount);
235
+ console.log('Automatic WASM fallback:',
236
+ execution.requestedDevice === 'auto' && execution.selectedDevice === 'wasm');
237
+ }
238
+ }
239
+ ```
240
+
241
+ Call `printSpeechStatus()` after the load in `sayHello()` or from your status
242
+ button. The same projection is at
243
+ `ai.providerRuntime.status(null, {execution:true}).roles.tts.execution`.
244
+ `selectedDevice` is `null`
245
+ until a pool is selected and returns to `null` on unload. Providers without an
246
+ execution report omit `execution`; do not infer a device from `navigator.gpu`
247
+ or a configured preference alone. An explicit inspection can throw a provider
248
+ status error; handle it with the same `error.code` / `error.message` pattern.
249
+
250
+ ## Stop, mute, cancel, and release
251
+
252
+ These are actions for your own controls, using the same `ai` instance:
253
+
254
+ ```javascript
255
+ function stopSpeech() {
256
+ ai.stopAudio(); // Cancels queued speech and stops scheduled/playing audio.
257
+ }
258
+
259
+ async function muteSpeech() {
260
+ await ai.setSpeechMuted(true); // Stops audio and unloads TTS.
261
+ }
262
+
263
+ async function unmuteSpeech() {
264
+ await ai.setSpeechMuted(false); // Loads TTS and permits playback.
265
+ }
266
+
267
+ async function unloadSpeech() {
268
+ ai.stopAudio();
269
+ await ai.providerRuntime.unload('tts'); // Keeps the provider configured.
270
+ }
271
+
272
+ async function disposeSpeech() {
273
+ await ai.disposeBrowserSpeech(); // Releases SDK-owned speech providers.
274
+ }
275
+ ```
276
+
277
+ STT has its own `ai.providerRuntime.load('stt')` and `unload('stt')` lifecycle
278
+ when configured; muting TTS does not unload STT or the LLM. A directly created
279
+ provider similarly exposes `load()`, `unload()`, and final `dispose()`.
280
+ `disposeBrowserSpeech()` leaves application-owned DBOPFS and stored artifacts
281
+ in place; it does not erase the application's data.
282
+
283
+ For a cancellable individual synthesis, unmute first. A fresh browser speech
284
+ configuration is muted, so calling `providerRuntime.load('tts')` directly at
285
+ that point rejects with `ARCANE_AI_TTS_MUTED`. `fetchTTS()` accepts an
286
+ `AbortSignal` as its second argument and returns a WAV `Blob` without playing it:
287
+
288
+ ```javascript
289
+ const synthesisController = new AbortController();
290
+
291
+ async function synthesizeOneSentence() {
292
+ try {
293
+ await ai.setSpeechMuted(false);
294
+ const result = await ai.fetchTTS({
295
+ input: 'This request can be cancelled.',
296
+ responseFormat: 'wav'
297
+ }, synthesisController.signal);
298
+ console.log(result); // The complete WAV Blob, with type 'audio/wav'.
299
+ return result;
300
+ } catch (error) {
301
+ console.error(error.code, error.message); // Keep the complete message.
302
+ throw error;
303
+ }
304
+ }
305
+
306
+ function cancelSynthesis() {
307
+ synthesisController.abort();
308
+ }
309
+ ```
310
+
311
+ Call `synthesizeOneSentence()` from an owned UI action and catch its rejection;
312
+ wire `cancelSynthesis()` to its Cancel control. The controller cancels the
313
+ individual `fetchTTS()` request; it does not control the preceding
314
+ `setSpeechMuted(false)` load/unmute lifecycle. Use a fresh controller for each
315
+ new operation. Configuration accepts `configureBrowserSpeech(configuration,
316
+ {signal})`; disposal accepts `disposeBrowserSpeech({signal})`. The streaming
317
+ playback methods do not accept a caller signal; wire
318
+ your stream's abort action to `ai.stopAudio()` as well as aborting its producer.
319
+ Cancellation suppresses late results, but upstream Kokoro may finish active
320
+ engine work before the affected slot is reusable.
321
+
322
+ ```javascript
323
+ const textStreamController = new AbortController();
324
+ textStreamController.signal.addEventListener('abort', function stopStreamAudio() {
325
+ ai.stopAudio();
326
+ }, { once: true });
327
+
328
+ function cancelTextAndSpeech() {
329
+ textStreamController.abort();
330
+ }
331
+ ```
332
+
333
+ Pass that same `textStreamController.signal` to your text producer's supported
334
+ signal option. Connect `cancelTextAndSpeech()` to Cancel; use a new controller
335
+ for the next stream.
336
+
337
+ ## Advanced provider and artifact reference
338
+
339
+ The SDK does not choose a runtime, model, voice, catalog, prompt, or product
340
+ policy. Applications keep those choices. Nothing is downloaded or activated
341
+ until the application explicitly calls `load()`.
342
+
343
+ Ordinary speech operation is the complete functional path. It uses the selected
344
+ upstream Transformers or Kokoro package and the browser's normal networking,
345
+ Worker, and Cache APIs. The records returned by this entrypoint are ordinary
346
+ JavaScript objects and arrays. Callers may copy, extend, and present complete
347
+ records; this contract does not freeze them or shorten their content.
348
+
349
+ ## Availability
350
+
351
+ | Host | Availability | Notes |
352
+ | --- | --- | --- |
353
+ | Browser | Shipped | Requires Workers, Fetch, Blob/File, object URLs, DBOPFS/OPFS, and Web Locks. Blob/File STT requests also require the browser audio decoder. |
354
+ | Native WebView | Conditional | Available when the WebView exposes the browser APIs above. It does not invoke Core speech. |
355
+ | Node | Importable, execution unavailable | The ESM subpath imports, but the SDK supplies no Node speech storage, Worker, or audio-decoder host. |
356
+ | Cloud | Not offered | The SDK's built-in speech profile is device-only: Whisper owns STT and Kokoro owns TTS. |
357
+
358
+ STT and TTS own independent provider lifecycles. A failure or cancellation in
359
+ one role does not disable the other role or authorize a fallback provider.
360
+
361
+ ## Public exports
362
+
363
+ ```javascript
364
+ import {
365
+ BROWSER_SPEECH_ARTIFACT_GRAPH_PROTOCOL,
366
+ BROWSER_SPEECH_ARTIFACT_PROTOCOL,
367
+ createBrowserKokoroProvider,
368
+ createBrowserSpeechArtifactGraph,
369
+ createBrowserSpeechAuthority,
370
+ createBrowserWhisperProvider,
371
+ createDbopfsSpeechArtifactStore
372
+ } from 'arcane-os/ai/browser-speech';
373
+ ```
374
+
375
+ Importing this entrypoint downloads nothing, opens no cache, creates no Worker,
376
+ and publishes no event.
377
+
378
+ ## Protocol identifiers
379
+
380
+ These exact strings identify the current public contracts:
381
+
382
+ | Subject | Exact value |
383
+ | --- | --- |
384
+ | Artifact-store protocol | `arcane-ai-browser-speech-artifacts/1` |
385
+ | Artifact-graph protocol | `arcane-ai-browser-speech-artifact-graph/1` |
386
+ | Graph `kind` and prepared `runtime.moduleGraph` | `browser-speech-authenticated-artifact-graph` |
387
+ | Single-module `runtime.moduleGraph` | `self-contained` |
388
+ | Model authority | `arcane-ai-model-authority/1` |
389
+ | Provider | `arcane-ai-provider/2` |
390
+ | Worker | `arcane-ai-speech-worker/1` |
391
+ | Worker error envelope | `arcane-ai-speech-worker-error/1` |
392
+ | Nested module Worker | `arcane-ai-browser-speech-artifact-module-worker/1` |
393
+
394
+ The word `authenticated` in the graph discriminator does not activate an
395
+ authentication, admission, or isolation stage. It is the current protocol
396
+ value.
397
+
398
+ ## `createBrowserSpeechArtifactGraph()`
399
+
400
+ An artifact graph describes the caller-selected runtime, model, voice, and
401
+ supporting files that the SDK stores and materializes. It is a routing and
402
+ selection record, not an execution permission list.
403
+
404
+ ```javascript
405
+ const graph = createBrowserSpeechArtifactGraph({
406
+ providerId: 'my-whisper',
407
+ role: 'stt',
408
+ model: {
409
+ id: 'whisper-small',
410
+ repository: 'publisher/whisper-small',
411
+ revision: 'selected-model-revision',
412
+ dtype: 'q8',
413
+ inputSampleRate: 16000
414
+ },
415
+ runtime: {
416
+ adapter: 'transformers-whisper',
417
+ version: 'selected-runtime-version',
418
+ revision: 'selected-runtime-revision',
419
+ entrypoint: 'runtime/transformers.js',
420
+ onnxWasm: {
421
+ namespace: 'transformers-env-backends-onnx-wasm',
422
+ mjsPath: 'runtime/ort-wasm.mjs',
423
+ wasmPath: 'runtime/ort-wasm.wasm'
424
+ }
425
+ },
426
+ files: [
427
+ {
428
+ kind: 'runtime-entrypoint-javascript',
429
+ path: 'runtime/transformers.js',
430
+ sourceUrl: 'https://publisher.example/transformers.js',
431
+ revision: 'selected-runtime-revision',
432
+ mediaType: 'text/javascript'
433
+ },
434
+ {
435
+ kind: 'runtime-auxiliary-javascript',
436
+ path: 'runtime/ort-wasm.mjs',
437
+ sourceUrl: 'https://publisher.example/ort-wasm.mjs',
438
+ revision: 'selected-runtime-revision',
439
+ mediaType: 'text/javascript'
440
+ },
441
+ {
442
+ kind: 'runtime-wasm-binary',
443
+ path: 'runtime/ort-wasm.wasm',
444
+ sourceUrl: 'https://publisher.example/ort-wasm.wasm',
445
+ revision: 'selected-runtime-revision',
446
+ mediaType: 'application/wasm'
447
+ },
448
+ {
449
+ kind: 'model-onnx-binary',
450
+ path: 'model/encoder.onnx',
451
+ sourceUrl: 'https://publisher.example/encoder.onnx',
452
+ revision: 'selected-model-revision',
453
+ mediaType: 'application/octet-stream',
454
+ runtimeRequestUrls: [
455
+ 'https://publisher.example/model/encoder.onnx'
456
+ ]
457
+ }
458
+ ]
459
+ });
460
+ ```
461
+
462
+ ### Model and runtime selection
463
+
464
+ `role` is `stt` or `tts`. The runtime adapter is
465
+ `transformers-whisper` for STT and `kokoro-js` for TTS. The caller supplies the
466
+ model id, repository, revision, dtype, sample rate, and runtime version and
467
+ revision.
468
+
469
+ STT requires `model.inputSampleRate`. TTS requires
470
+ `model.outputSampleRate`, `model.defaultVoice`, and a nonempty
471
+ `model.voices` array of `{id,path}` records. Each voice path names a declared
472
+ `voice-style-binary` file. `runtime.onnxWasm` names the selected ONNX module and
473
+ WASM files. `numThreads` is optional for Transformers and is not inferred from
474
+ hardware. Kokoro does not expose that field.
475
+
476
+ ### File records
477
+
478
+ Each file record uses:
479
+
480
+ ```text
481
+ {
482
+ kind,
483
+ path,
484
+ sourceUrl,
485
+ revision,
486
+ license?,
487
+ mediaType,
488
+ sourceMediaType?,
489
+ runtimeRequestUrls?
490
+ }
491
+ ```
492
+
493
+ `path` is a normalized relative path. `sourceUrl` is the caller-selected source
494
+ used for installation. `runtimeRequestUrls` lists aliases used by upstream
495
+ runtime code for the same stored file. `mediaType` becomes the materialized
496
+ Blob type; `sourceMediaType` may describe a different upstream response type.
497
+ The SDK does not require or interpret legal metadata at runtime. If the caller
498
+ includes `license`, the graph preserves that complete value as inert metadata;
499
+ runtime materialization never treats it as capability or admission data.
500
+
501
+ Runtime file kinds are `runtime-entrypoint-javascript`,
502
+ `runtime-auxiliary-javascript`, `runtime-wasm-binary`, and
503
+ `runtime-opaque-data`. Model/data kinds are `model-configuration-json`,
504
+ `model-generation-configuration-json`, `model-onnx-binary`,
505
+ `model-onnx-external-data`, `model-preprocessor-json`,
506
+ `model-tokenizer-json`, `model-opaque-data`, and `voice-style-binary`.
507
+
508
+ Graph construction requires paths and route aliases to be unambiguous so one
509
+ known URL maps to at most one stored file. The runtime router is independently
510
+ permissive: if ambiguous routing metadata nevertheless reaches it, that
511
+ URL is left unmapped and uses the native browser operation.
512
+
513
+ ## `createDbopfsSpeechArtifactStore()`
514
+
515
+ ```javascript
516
+ const store = createDbopfsSpeechArtifactStore({
517
+ dbopfs,
518
+ tableName: 'arcane_ai_browser_speech'
519
+ });
520
+ ```
521
+
522
+ The store exposes `{protocol,tableName,prepare,remove}`. It serializes updates
523
+ to one selected authority with Web Locks, downloads each caller-selected file
524
+ after explicit provider activation, stores it in DBOPFS, and reopens the stored
525
+ file before materialization. Missing storage, an unreadable response, a failed
526
+ HTTP request, a missing stored file, or cancellation rejects honestly.
527
+
528
+ The store writes ordinary mutable selection metadata before the selected files.
529
+ On a later load, a changed file inventory or source mapping is a cache miss and
530
+ is downloaded again. The selection metadata is not a completion, integrity, or
531
+ publication receipt and never blocks ordinary loading.
532
+
533
+ `prepare(authority,{signal,onProgress,offline=false,security})` accepts an
534
+ SDK-created artifact graph or upstream-package authority. `offline:true` uses
535
+ only existing DBOPFS state. Preparation returns the selected runtime/model
536
+ configuration, `cache` as `installed` or `cached`, object URLs, and a `release()`
537
+ function that revokes the materialized URLs.
538
+
539
+ `security` records caller intent only. Ordinary preparation does not
540
+ forward a security payload to the Worker and performs no security work. Passing
541
+ `{secure:true}` does not activate hardening. Any future hardening stage requires
542
+ a separate user review and an explicit implementation change before it may execute.
543
+
544
+ ## Ordinary module routing
545
+
546
+ Artifact-graph preparation reads each stored JavaScript module, discovers
547
+ ordinary module operations, and materializes the complete stored files as
548
+ object URLs. The prepared runtime retains the existing
549
+ `browser-speech-authenticated-artifact-graph` discriminator.
550
+
551
+ The Worker installs one private module router before importing the entrypoint:
552
+
553
+ - a static import whose target is a known stored file uses that file's
554
+ materialized URL;
555
+ - a dynamic import whose target is known imports the materialized URL;
556
+ - a fetch whose target is known reads the materialized URL;
557
+ - a Worker whose target is known starts the SDK role Worker and imports the
558
+ materialized target there; and
559
+ - a Cache match whose target is known returns the materialized file.
560
+
561
+ Every unmapped operation keeps ordinary browser behavior:
562
+
563
+ - an unmapped relative or URL-like import resolves against the calling module's
564
+ original source URL, while a bare specifier remains unchanged for native
565
+ import-map resolution;
566
+ - an unmapped fetch calls native `fetch` and preserves the caller's options;
567
+ - an unmapped Worker calls the native `Worker` constructor and preserves the
568
+ caller's options; and
569
+ - an unmapped Cache operation delegates to native Cache Storage.
570
+
571
+ Cache `put`, `add`, `addAll`, `delete`, and `keys` are normal mutable browser
572
+ operations. The SDK does not replace them with a read-only cache. Relative
573
+ requests are resolved from the calling module's original source URL before the
574
+ native Cache operation.
575
+
576
+ Routing discovery is best effort and is not an admission gate. If the scanner
577
+ cannot interpret a module, that module is left unchanged and follows its native
578
+ URLs. A static-import cycle may likewise retain an original source URL where a
579
+ target has not yet been materialized. These fallbacks preserve functionality;
580
+ they do not silently convert into a rejection policy.
581
+
582
+ ## ONNX runtime configuration
583
+
584
+ The Worker applies only the selected runtime settings needed to run the chosen
585
+ provider:
586
+
587
+ - Kokoro forwards the pool's selected `webgpu` or `wasm` device to
588
+ `KokoroTTS.from_pretrained()`. Its configured dtype remains exactly the
589
+ caller-selected dtype on both paths. The WASM path also uses
590
+ `namespace.env.wasmPaths = {mjs,wasm}`.
591
+ - Transformers uses
592
+ `namespace.env.backends.onnx.wasm.wasmPaths = {mjs,wasm}`, keeps remote model
593
+ loading enabled, and applies caller-selected `numThreads` when present.
594
+
595
+ Transformers and Kokoro keep their normal provider downloads and Cache behavior
596
+ for routes not materialized by the SDK. Kokoro voice aliases and Transformers
597
+ model aliases may be listed in `runtimeRequestUrls` so an upstream request for a
598
+ known file resolves to the already materialized local file. Model, voice, and
599
+ runtime selection remains with the application and upstream publisher.
600
+
601
+ ## Providers
602
+
603
+ ```javascript
604
+ const whisper = createBrowserWhisperProvider({
605
+ id: 'my-whisper',
606
+ graph,
607
+ store
608
+ });
609
+
610
+ const kokoro = createBrowserKokoroProvider({
611
+ id: 'my-kokoro',
612
+ graph: kokoroGraph,
613
+ store,
614
+ execution: {
615
+ device: 'auto',
616
+ maxConcurrentRequests: 4
617
+ }
618
+ });
619
+ ```
620
+
621
+ The constructors also accept ordinary `model` and `runtime` descriptors instead
622
+ of `graph`; the two forms are mutually exclusive. Both forms require an
623
+ SDK-created DBOPFS speech artifact store. `localOnly` remains `true`.
624
+
625
+ Kokoro additionally accepts the exact `execution` record
626
+ `{device,maxConcurrentRequests}`. `device` is `auto`, `webgpu`, or `wasm`;
627
+ `maxConcurrentRequests` is an integer from 1 through 4. Omission defaults to
628
+ `{device:'auto',maxConcurrentRequests:4}`. `auto` attempts a complete WebGPU
629
+ Worker pool only when the browser exposes WebGPU. If that pool cannot load, the
630
+ SDK tears it down and creates a complete WASM pool with the same caller-selected
631
+ model and dtype. Explicit `webgpu` rejects when WebGPU cannot load; explicit
632
+ `wasm` never attempts GPU. Whisper remains one WASM Worker and does not accept
633
+ this option.
634
+
635
+ Each constructor returns an `arcane-ai-provider/2` object with:
636
+
637
+ ```text
638
+ {
639
+ protocol,
640
+ role,
641
+ id,
642
+ localOnly,
643
+ maxConcurrentRequests,
644
+ catalog,
645
+ inspect,
646
+ status,
647
+ load,
648
+ request,
649
+ unload,
650
+ dispose
651
+ }
652
+ ```
653
+
654
+ `catalog()`, `inspect()`, and `status()` do not activate a provider. `load()` is
655
+ the explicit activation boundary. The caller supplies model/profile policy and
656
+ may display the provider's lifecycle status. `unload()` releases the Worker and
657
+ materialized URLs. `dispose()` performs final teardown and prevents later use.
658
+
659
+ ### Upstream-package authority
660
+
661
+ `createBrowserSpeechAuthority({providerId,role,model,runtime,security})` creates
662
+ the ordinary single-entrypoint authority. The model descriptor is
663
+ `{id,repository,revision,dtype?,defaultVoice?,files?}`. The runtime descriptor is
664
+ `{adapter,version,revision,entry,wasmPaths?,files}`. A file is
665
+ `{path,url,mediaType?}`. Model files may be omitted so the selected upstream
666
+ provider performs its normal model and voice downloads after explicit use.
667
+
668
+ The authority record is a mutable complete record. An omitted ordinary security
669
+ option produces no security field. A present `{secure:true}` value records only
670
+ the caller's future intent and does not change loading or routing behavior.
671
+
672
+ ### Whisper STT
673
+
674
+ ```javascript
675
+ const result = await whisper.request({
676
+ role: 'stt',
677
+ operation: 'transcribe',
678
+ signal,
679
+ payload: {
680
+ audio: pcmFloat32,
681
+ sampleRate: 16000
682
+ }
683
+ });
684
+ ```
685
+
686
+ The provider-native payload is mono `Float32Array` PCM at the selected input
687
+ sample rate, and the result is `{text}`. The shared AI form accepts
688
+ `{audio:Blob|File,mimeType,model}` and uses the browser decoder to produce the
689
+ same mono input. The SDK preserves the complete returned transcript.
690
+
691
+ ### Kokoro TTS
692
+
693
+ ```javascript
694
+ const result = await kokoro.request({
695
+ role: 'tts',
696
+ operation: 'synthesize',
697
+ signal,
698
+ payload: {
699
+ text: 'Hello from Arcane.',
700
+ voice: 'caller-voice-id',
701
+ speed: 1
702
+ }
703
+ });
704
+ ```
705
+
706
+ The voice belongs to the caller-selected inventory; omission uses that model's
707
+ default voice. The provider-native result is
708
+ `{audio:Float32Array,sampleRate,voice}`. The provider/2 shared request form accepts
709
+ `{model,input,responseFormat:'wav',voice?,speed?}` and returns
710
+ `{audio:Uint8Array,contentType:'audio/wav'}`. High-level `AI.fetchTTS()` wraps
711
+ that provider result in a WAV `Blob`. Returned provider records remain ordinary
712
+ mutable values.
713
+
714
+ ## Lifecycle and cancellation
715
+
716
+ Provider states are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
717
+ `disposed`. `status()` includes role, provider/model ids, state, lifecycle
718
+ status and reason, active operation, loaded/busy flags, generation, error code,
719
+ cache state, and warnings. Kokoro status also includes an `execution` record
720
+ with requested and selected device, request limit, and active request count. A
721
+ successful `selectedDevice:'webgpu'` reports the execution provider selected by
722
+ the upstream model load; it does not claim that browser, driver, or GPU kernels
723
+ overlap physically. A security field is absent in ordinary mode.
724
+
725
+ The provider/2 load context accepts an optional progress callback for interface
726
+ compatibility, but the current browser-speech artifact and Worker transport
727
+ publishes no progress records. Consumers present the explicit lifecycle states
728
+ instead of inventing a numeric total.
729
+
730
+ Compatible concurrent loads share one underlying preparation and pool load
731
+ while each caller retains its own cancellation signal. One observer cannot
732
+ cancel another still-active observer; cancellation of the final observer stops
733
+ the shared load.
734
+
735
+ Whisper retains one active role request. Kokoro admits synthesis requests up to
736
+ its declared capacity and rejects a direct over-capacity call with
737
+ `ARCANE_AI_PROVIDER_BUSY`; the provider-neutral runtime keeps overflow in FIFO
738
+ order. Each Kokoro slot owns a distinct Worker and loaded model session because
739
+ the selected browser adapter serializes inference inside one JavaScript
740
+ isolate. The SDK prepares the artifact URLs once and shares that same prepared
741
+ selection across the bounded pool. Pool activation completes the first model
742
+ session before it starts the remaining Workers, then loads those remaining
743
+ sessions concurrently. This avoids multiplying simultaneous cold artifact
744
+ acquisition while still making the configured synthesis capacity ready in
745
+ parallel.
746
+
747
+ Cancellation of one active Kokoro synthesis suppresses only that request and
748
+ sends the Worker's targeted cancel control. The selected upstream Kokoro
749
+ version may finish already-running engine work before that slot can run its
750
+ next request; the SDK does not claim stronger per-call preemption. Whisper
751
+ cancellation and role unload/dispose retain destructive Worker teardown.
752
+ Kokoro unload/dispose abort every active request, terminates every pool Worker,
753
+ and releases materialized URLs once. Late results cannot settle a cancelled or
754
+ superseded operation.
755
+
756
+ The provider owns no event bus. Applications may project promises and status
757
+ into the SDK's shared event/state owner. Mute and unmute are likewise
758
+ application state: mute may await `unload()`, and unmute may explicitly call
759
+ `load()`.
760
+
761
+ ## Errors
762
+
763
+ Errors retain the normal `code`, `message`, `reason`, and `cause` fields used by
764
+ the browser AI runtime.
765
+
766
+ The Worker error envelope carries `cause` as an optional mutable diagnostic
767
+ record. It preserves complete nested messages, stacks, codes, reasons, details,
768
+ own properties, and cycles without a depth or content cap. Worker and client
769
+ sources must come from the same SDK revision so their `/1` envelope shape is
770
+ updated atomically; a current client still accepts the cause-free four-field form.
771
+ If the platform cannot clone an exotic diagnostic value, the Worker keeps the
772
+ complete raw failure in its console diagnostics and retries the response with
773
+ that cause-free four-field envelope so the caller still receives an error.
774
+
775
+ Representative stable codes include:
776
+
777
+ - `ARCANE_AI_INVALID_REQUEST`
778
+ - `ARCANE_AI_MODEL_AUTHORITY_REQUIRED`
779
+ - `ARCANE_AI_PROVIDER_UNAVAILABLE`
780
+ - `ARCANE_AI_PROVIDER_BUSY`
781
+ - `ARCANE_AI_PROVIDER_DISPOSED`
782
+ - `ARCANE_AI_REQUEST_ABORTED`
783
+ - `ARCANE_AI_OPERATION_SUPERSEDED`
784
+ - `ARCANE_AI_ARTIFACT_DOWNLOAD_FAILED`
785
+ - `ARCANE_AI_ARTIFACT_OFFLINE_MISS`
786
+ - `ARCANE_AI_STORAGE_BUSY`
787
+ - `ARCANE_AI_STORAGE_UNAVAILABLE`
788
+ - `ARCANE_AI_WORKER_MESSAGE_ERROR`
789
+
790
+ Malformed selected descriptors, missing required files, unreadable responses,
791
+ unsupported provider namespace shapes, and unavailable browser APIs reject at
792
+ their functional owner. An unmapped runtime route retains the native browser
793
+ operation.
794
+
795
+ ## Ownership
796
+
797
+ - Applications own model, runtime, dtype, sample-rate, voice, profile, prompt,
798
+ catalog, activation, optional TTS execution override, and presentation
799
+ policy.
800
+ - Upstream publishers own their runtime, model, voice, and license delivery.
801
+ - The SDK owns storage, materialization, routing, Worker lifecycle, normalized
802
+ provider contracts, cancellation, and cleanup.
803
+ - The SDK redistributes no third-party runtime, model, or voice package through
804
+ this entrypoint.
805
+
806
+ ## Related
807
+
808
+ - [Normalized AI](../README.md#normalized-ai)
809
+ - [Browser-WASM LLM](browser-wasm.md)
810
+ - [AIProviderRuntime.js](../runtime-modules.md#aiproviderruntimejs)
811
+ - [AIRuntimeState.js](../runtime-modules.md#airuntimestatejs)
812
+ - [Availability and normalization](../availability-and-normalization.md)
813
+ - [Protocol architecture](../protocols.md#portable-ai-provider-runtime)