arcane-os 0.5.10 → 0.5.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (54) hide show
  1. package/CHANGELOG.md +41 -0
  2. package/README.md +117 -26
  3. package/browser-runtime/ai/browser-speech-providers.mjs +1 -1
  4. package/docs/architecture.md +303 -0
  5. package/docs/compatibility.md +38 -0
  6. package/docs/event-manager.md +263 -0
  7. package/docs/platform-targets.md +104 -0
  8. package/docs/publishing.md +126 -0
  9. package/docs/reference/README.md +206 -0
  10. package/docs/reference/ai/browser-speech.md +879 -0
  11. package/docs/reference/ai/browser-wasm.md +637 -0
  12. package/docs/reference/ai/twin-cloud.md +156 -0
  13. package/docs/reference/arcane-ollama.md +288 -0
  14. package/docs/reference/availability-and-normalization.md +224 -0
  15. package/docs/reference/behavioral-testing.md +129 -0
  16. package/docs/reference/cli.md +820 -0
  17. package/docs/reference/core/README.md +61 -0
  18. package/docs/reference/core/arcane-ai-contracts.md +906 -0
  19. package/docs/reference/core/arcane-api.md +601 -0
  20. package/docs/reference/core/arcane-entities.md +59 -0
  21. package/docs/reference/core/arcane-events.md +134 -0
  22. package/docs/reference/core/ollama-module.md +181 -0
  23. package/docs/reference/core/reference/arcane-api/ai-and-ollama.md +1909 -0
  24. package/docs/reference/core/reference/arcane-api/applications-terminal-capabilities.md +1057 -0
  25. package/docs/reference/core/reference/arcane-api/core-and-events.md +320 -0
  26. package/docs/reference/core/reference/arcane-api/filesystem-storage-preferences-appearance.md +610 -0
  27. package/docs/reference/core/reference/arcane-api/namespaces.md +1157 -0
  28. package/docs/reference/core/reference/arcane-api/platform-installation-users-system.md +1423 -0
  29. package/docs/reference/core/reference/arcane-api/session-provisioning-diagnostics-development.md +315 -0
  30. package/docs/reference/event-manager.md +1409 -0
  31. package/docs/reference/inventory/package-api.json +3194 -0
  32. package/docs/reference/inventory/runtime-components.json +1015 -0
  33. package/docs/reference/inventory/runtime-entities.json +25 -0
  34. package/docs/reference/inventory/runtime-modules.json +1367 -0
  35. package/docs/reference/mail.md +309 -0
  36. package/docs/reference/protocols.md +749 -0
  37. package/docs/reference/runtime-components.md +1532 -0
  38. package/docs/reference/runtime-entities.md +305 -0
  39. package/docs/reference/runtime-modules.md +3310 -0
  40. package/docs/reference/sdk-api.md +6733 -0
  41. package/docs/roadmap.md +79 -0
  42. package/docs/work-amplification.md +66 -0
  43. package/examples/wasm-ai-demo/README.md +80 -0
  44. package/examples/wasm-ai-demo/app.js +787 -0
  45. package/examples/wasm-ai-demo/index.html +343 -0
  46. package/examples/wasm-ai-demo/profile-tools.js +217 -0
  47. package/examples/wasm-ai-demo/profiles/BOSS.Modelfile +106 -0
  48. package/examples/wasm-ai-demo/profiles/PreCrisis.Modelfile +693 -0
  49. package/examples/wasm-ai-demo/rag/boss-library.json +3006 -0
  50. package/examples/wasm-ai-demo/rag.js +295 -0
  51. package/examples/wasm-ai-demo/server.mjs +71 -0
  52. package/package.json +10 -1
  53. package/runtime/arcane/modules/AI.js +60 -11
  54. package/runtime/arcane/modules/AIProviderRuntime.js +26 -5
@@ -0,0 +1,879 @@
1
+ # Browser speech providers
2
+
3
+ `arcane-os/ai/browser-speech` is the browser-only SDK boundary for
4
+ caller-selected Whisper speech-to-text and Kokoro text-to-speech runtimes. It
5
+ provides artifact storage, live module routing, role Workers, provider/2
6
+ adapters, bounded parallel TTS synthesis, audio normalization, cancellation,
7
+ and cleanup.
8
+
9
+ ## Quick start: say one sentence
10
+
11
+ Use this in a browser application served by `arcane dev`, where the generated
12
+ import map resolves `arcane/AI` and `arcane/DBOPFS`. These browser modules are
13
+ not Node inference APIs. To create an application:
14
+
15
+ ```bash
16
+ npx arcane-os@0.5.12 new hello-speech --path ./hello-speech --target browser
17
+ cd hello-speech
18
+ npm install
19
+ npm run dev
20
+ ```
21
+
22
+ Keep the generated page's Arcane theme and import map. The examples below go in
23
+ `apps/hello-speech/modules/App.js` and the adjacent `speech-selection.js`.
24
+ Follow the development server's printed URL. Installing the SDK supplies its
25
+ provider, storage, and Worker code; it does not install a speech model or choose
26
+ an upstream speech runtime for your application.
27
+
28
+ First create **`speech-selection.js`**, the one application-owned configuration
29
+ file used throughout this guide. This concrete selection is also used by the
30
+ [maintained WASM voice-chat example](https://github.com/TheWizardNexus/arcane-os-sdk/tree/main/examples/wasm-ai-demo).
31
+ Your application owns these runtime/model versions, URLs, dtype, and voice.
32
+ Loading this selection uses those upstream publishers' downloads and caches.
33
+
34
+ ```javascript
35
+ export const speechSelection = {
36
+ model: {
37
+ id: 'onnx-community/Kokoro-82M-v1.0-ONNX',
38
+ repository: 'onnx-community/Kokoro-82M-v1.0-ONNX',
39
+ revision: '1939ad2a8e416c0acfeecc08a694d14ef25f2231',
40
+ dtype: 'q8',
41
+ defaultVoice: 'af_heart'
42
+ },
43
+ runtime: {
44
+ adapter: 'kokoro-js',
45
+ version: '1.2.1',
46
+ revision: '664c76a704021239ba59c84dcbaa4d3dece01fe9',
47
+ entry: 'kokoro.web.js',
48
+ wasmPaths: 'https://cdn.jsdelivr.net/npm/@huggingface/transformers@3.5.1/dist/',
49
+ files: [{
50
+ path: 'kokoro.web.js',
51
+ url: 'https://cdn.jsdelivr.net/npm/kokoro-js@1.2.1/dist/kokoro.web.js',
52
+ mediaType: 'text/javascript'
53
+ }]
54
+ }
55
+ };
56
+ ```
57
+
58
+ Then use this **`App.js`**. The application creates and owns the DBOPFS
59
+ instance. Configuration selects the provider without loading it; the button
60
+ explicitly loads and unmutes TTS before requesting speech.
61
+
62
+ ```javascript
63
+ import arcaneThemeReady from 'arcane/ThemeBootstrap';
64
+ import AI, { AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL } from 'arcane/AI';
65
+ import DBOPFS from 'arcane/DBOPFS';
66
+ import { speechSelection } from './speech-selection.js';
67
+
68
+ await arcaneThemeReady;
69
+ const dbopfs = new DBOPFS();
70
+ await dbopfs.readyPromise;
71
+ const ai = new AI();
72
+
73
+ await ai.configureBrowserSpeech({
74
+ protocol: AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL,
75
+ id: 'hello-speech',
76
+ dbopfs,
77
+ tts: {
78
+ providerId: 'hello-kokoro',
79
+ model: speechSelection.model,
80
+ runtime: speechSelection.runtime,
81
+ offline: false
82
+ }
83
+ });
84
+
85
+ const speakButton = document.createElement('button');
86
+ speakButton.textContent = 'Load voice and say hello';
87
+ document.body.append(speakButton);
88
+ speakButton.addEventListener('click', async function sayHello() {
89
+ speakButton.disabled = true;
90
+ try {
91
+ await ai.setSpeechMuted(false); // Loads the selected TTS provider.
92
+ const prepared = await ai.streamTTS('Hello from Arcane. ', true);
93
+ console.log('Speech preparation completed:', prepared);
94
+ } catch (error) {
95
+ console.error(error.code, error.message);
96
+ } finally {
97
+ speakButton.disabled = false;
98
+ }
99
+ });
100
+ ```
101
+
102
+ The first user action may download the selected runtime, model, and voice.
103
+ The browser may require another audio-unlock gesture after a long load; the SDK
104
+ retains prepared audio for that gesture. The two-argument `streamTTS()` call
105
+ resolves after preparing audio for scheduling; it does not wait for playback
106
+ to end. It returns `false` when that preparation is muted, stopped, or fails.
107
+ Later playback failures still reach the SDK's complete console diagnostics
108
+ and `ai-tts-failure` event. Use the optional playback mode below when you need
109
+ to wait for the submitted audio to end. Neither mode proves a listener heard it.
110
+
111
+ To display complete high-level playback errors in this page, observe its
112
+ existing event. The listener belongs to this example's one `ai` instance:
113
+
114
+ ```javascript
115
+ const speechEvents = new AbortController();
116
+ window.addEventListener('ai-tts-failure', function reportSpeechFailure(event) {
117
+ if (event.detail.ai !== ai) return;
118
+ console.error(event.detail.error.code, event.detail.error.message);
119
+ }, { signal: speechEvents.signal });
120
+ ```
121
+
122
+ Call `speechEvents.abort()` when disposing that interface to remove the listener.
123
+
124
+ ## Four synthesis slots and exact-order playback
125
+
126
+ Capacity 4 means up to four segments synthesize at once. Segment 5 and later
127
+ wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out of
128
+ order, but playback waits for earlier segments and plays exact input order.
129
+ Each slot owns a Worker/model session, so raising capacity trades memory for
130
+ latency.
131
+
132
+ The high-level `AI` route owns the FIFO queue. Calling a low-level Kokoro
133
+ provider directly beyond its capacity returns `ARCANE_AI_PROVIDER_BUSY`.
134
+ Ready adjacent audio buffers use contiguous AudioContext scheduling. Browser
135
+ audio scheduling and selected WebGPU status do not prove physical GPU kernel
136
+ overlap or audio quality. LLM and Whisper/STT capacity remains one.
137
+
138
+ ## Queue complete passages and wait for playback
139
+
140
+ These options are available in SDK `0.5.12`.
141
+
142
+ For a complete page or passage, call
143
+ `ai.streamTTS(text, true, {voice, speed, pauseAfterMs, waitForPlayback:true})`.
144
+ The existing AI queue owns segmentation, concurrent synthesis, ordered
145
+ playback, and cancellation. Supply the exact text; there is no need for an
146
+ application sentence queue, audio cache, or playback scheduler.
147
+
148
+ This function uses the configured `ai` above. Its application-supplied
149
+ `passages` argument is an ordered array of `{text, voice?, speed?, pauseAfterMs?}`
150
+ records. An omitted voice uses the selected model's default voice, an omitted
151
+ speed uses `ai.voiceSpeed`, and an omitted pause is zero. Each supplied voice
152
+ must be supported by the selected model; speed must be positive. A pause is
153
+ finite, nonnegative milliseconds and applies only after that passage's final
154
+ extracted segment. These options do not change the instance defaults.
155
+
156
+ ```javascript
157
+ async function speakPassages(passages) {
158
+ await ai.setSpeechMuted(false);
159
+ const pending = passages.map(
160
+ function queuePassage(passage) {
161
+ return ai.streamTTS(
162
+ passage.text,
163
+ true,
164
+ {
165
+ voice: passage.voice,
166
+ speed: passage.speed,
167
+ pauseAfterMs: passage.pauseAfterMs,
168
+ waitForPlayback: true
169
+ }
170
+ );
171
+ }
172
+ );
173
+ return Promise.all(pending);
174
+ }
175
+ ```
176
+
177
+ Call `speakPassages(...)` from your owned user action and handle errors with
178
+ the earlier `error.code` / `error.message` pattern. The `map` submits every
179
+ passage synchronously before `Promise.all` waits, so synthesis can use the
180
+ provider's available capacity. The returned array has one boolean per passage
181
+ in input order: `true` after all its extracted audio buffers naturally end,
182
+ or `false` after terminal cancellation or failure. `ai.stopAudio()` cancels
183
+ all speech owned by that AI instance and settles pending playback results
184
+ `false`.
185
+
186
+ The selected voice and speed are captured for segments extracted by that call.
187
+ That includes any text left in the same AI instance's partial-stream buffer;
188
+ finish the previous producer before starting a separate complete passage.
189
+ Options are not retained with an unfinished `end:false` remainder. A later
190
+ call supplies its own options, and `finishTTS()` uses defaults.
191
+ A call extracting no segments returns `true` without waiting for earlier jobs.
192
+ An already muted call returns `false`.
193
+
194
+ Autoplay permission waiting and recoverable audio-resume attempts leave the
195
+ playback promise pending until playback completes or is stopped. A failed
196
+ resume of a closed `AudioContext` terminates the affected jobs and settles
197
+ their playback results `false`. A trailing
198
+ pause delays the next queued audio on the existing `AudioContext` clock; it
199
+ does not delay the preceding promise after that passage's last buffer ends.
200
+ The promise is a playback result, not a listener acknowledgement.
201
+
202
+ ## Stream chunks as they arrive
203
+
204
+ Use the configured `ai` created above. Run this snippet from an owned user
205
+ action, such as your Speak button, and catch errors with the earlier
206
+ `error.code` / `error.message` pattern. First load/unmute, then accept chunks.
207
+ Call `streamTTS(chunk)` inside the producer's chunk callback immediately;
208
+ waiting for each speech promise there would serialize synthesis. This tiny
209
+ example uses three arriving chunks. Replace the three `onTextChunk(...)` calls
210
+ with your actual text stream callback, and flush after that producer ends.
211
+
212
+ ```javascript
213
+ await ai.setSpeechMuted(false);
214
+ const pendingSpeech = [];
215
+
216
+ function onTextChunk(chunk) {
217
+ const pending = ai.streamTTS(chunk);
218
+ // Attach both handlers immediately, so later failure cannot be unhandled.
219
+ pendingSpeech.push(pending.then(
220
+ function speechPrepared(value) { return { status: 'fulfilled', value }; },
221
+ function speechRejected(reason) { return { status: 'rejected', reason }; }
222
+ ));
223
+ }
224
+
225
+ onTextChunk('First sentence. ');
226
+ onTextChunk('Second sentence. ');
227
+ onTextChunk('Third sentence.');
228
+
229
+ // After the text producer ends, flush trailing text and settle every call.
230
+ const finalPrepared = await ai.finishTTS();
231
+ const outcomes = await Promise.all(pendingSpeech);
232
+ for (const outcome of outcomes) {
233
+ if (outcome.status === 'rejected') {
234
+ console.error(outcome.reason.code, outcome.reason.message);
235
+ } else if (outcome.value === false) {
236
+ console.log('Speech was muted, stopped, or failed; inspect SDK diagnostics.');
237
+ }
238
+ }
239
+ if (finalPrepared === false) {
240
+ console.log('Final speech was muted, stopped, or failed; inspect SDK diagnostics.');
241
+ }
242
+ ```
243
+
244
+ The segmentation default uses sentence punctuation. To submit smaller complete
245
+ segments, call `ai.configureTTSSegmentation({punctuation:'any',wordCadence:4})`
246
+ before feeding the stream. Chunk boundaries themselves do not force a sentence
247
+ boundary; `finishTTS()` flushes any remaining text. It is not a playback-ended
248
+ notification. Do not mute or dispose immediately after it if playback should
249
+ continue.
250
+
251
+ ## Choose a device or reduce memory use
252
+
253
+ Omitting `tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`.
254
+ These are three alternative configurations, not a sequence of required loads:
255
+
256
+ ```javascript
257
+ async function selectSpeechExecution(execution) {
258
+ await ai.configureBrowserSpeech({
259
+ protocol: AI_BROWSER_SPEECH_CONFIGURATION_PROTOCOL,
260
+ id: 'hello-speech',
261
+ dbopfs,
262
+ tts: {
263
+ providerId: 'hello-kokoro',
264
+ model: speechSelection.model,
265
+ runtime: speechSelection.runtime,
266
+ offline: false,
267
+ execution
268
+ }
269
+ });
270
+ }
271
+
272
+ // Choose and call one from your application settings action:
273
+ // await selectSpeechExecution({ device: 'auto' });
274
+ // await selectSpeechExecution({ device: 'webgpu' });
275
+ // await selectSpeechExecution({ device: 'wasm', maxConcurrentRequests: 1 });
276
+ ```
277
+
278
+ The override accepts integers 1, 2, 3, or 4. `auto` tries a complete WebGPU pool
279
+ when available, then recreates a complete WASM pool if that load cannot finish.
280
+ Explicit `webgpu` reports a load error when unavailable; explicit `wasm` never
281
+ attempts WebGPU. The same application-selected model and dtype apply on both
282
+ devices. A configuration change leaves TTS muted; explicitly load/unmute again.
283
+
284
+ ## Inspect the requested and selected device
285
+
286
+ Request the execution projection on the existing public runtime status after
287
+ loading. This explicitly reads the selected provider's current report; ordinary
288
+ `status()` retains its existing sticky snapshot and identity. There is no
289
+ separate execution-state event subscription.
290
+
291
+ ```javascript
292
+ function printSpeechStatus() {
293
+ const status = ai.providerRuntime.status('tts', { execution: true });
294
+ const execution = status.execution;
295
+ console.log('TTS state:', status.state);
296
+ if (execution) {
297
+ console.log('Requested device:', execution.requestedDevice);
298
+ console.log('Selected device:', execution.selectedDevice);
299
+ console.log('Capacity:', execution.maxConcurrentRequests);
300
+ console.log('Active synthesis requests:', execution.activeRequestCount);
301
+ console.log('Automatic WASM fallback:',
302
+ execution.requestedDevice === 'auto' && execution.selectedDevice === 'wasm');
303
+ }
304
+ }
305
+ ```
306
+
307
+ Call `printSpeechStatus()` after the load in `sayHello()` or from your status
308
+ button. The same projection is at
309
+ `ai.providerRuntime.status(null, {execution:true}).roles.tts.execution`.
310
+ `selectedDevice` is `null`
311
+ until a pool is selected and returns to `null` on unload. Providers without an
312
+ execution report omit `execution`; do not infer a device from `navigator.gpu`
313
+ or a configured preference alone. An explicit inspection can throw a provider
314
+ status error; handle it with the same `error.code` / `error.message` pattern.
315
+
316
+ ## Stop, mute, cancel, and release
317
+
318
+ These are actions for your own controls, using the same `ai` instance:
319
+
320
+ ```javascript
321
+ function stopSpeech() {
322
+ ai.stopAudio(); // Cancels queued speech and stops scheduled/playing audio.
323
+ }
324
+
325
+ async function muteSpeech() {
326
+ await ai.setSpeechMuted(true); // Stops audio and unloads TTS.
327
+ }
328
+
329
+ async function unmuteSpeech() {
330
+ await ai.setSpeechMuted(false); // Loads TTS and permits playback.
331
+ }
332
+
333
+ async function unloadSpeech() {
334
+ ai.stopAudio();
335
+ await ai.providerRuntime.unload('tts'); // Keeps the provider configured.
336
+ }
337
+
338
+ async function disposeSpeech() {
339
+ await ai.disposeBrowserSpeech(); // Releases SDK-owned speech providers.
340
+ }
341
+ ```
342
+
343
+ STT has its own `ai.providerRuntime.load('stt')` and `unload('stt')` lifecycle
344
+ when configured; muting TTS does not unload STT or the LLM. A directly created
345
+ provider similarly exposes `load()`, `unload()`, and final `dispose()`.
346
+ `disposeBrowserSpeech()` leaves application-owned DBOPFS and stored artifacts
347
+ in place; it does not erase the application's data.
348
+
349
+ For a cancellable individual synthesis, unmute first. A fresh browser speech
350
+ configuration is muted, so calling `providerRuntime.load('tts')` directly at
351
+ that point rejects with `ARCANE_AI_TTS_MUTED`. `fetchTTS()` accepts an
352
+ `AbortSignal` as its second argument and returns a WAV `Blob` without playing it:
353
+
354
+ ```javascript
355
+ const synthesisController = new AbortController();
356
+
357
+ async function synthesizeOneSentence() {
358
+ try {
359
+ await ai.setSpeechMuted(false);
360
+ const result = await ai.fetchTTS({
361
+ input: 'This request can be cancelled.',
362
+ responseFormat: 'wav'
363
+ }, synthesisController.signal);
364
+ console.log(result); // The complete WAV Blob, with type 'audio/wav'.
365
+ return result;
366
+ } catch (error) {
367
+ console.error(error.code, error.message); // Keep the complete message.
368
+ throw error;
369
+ }
370
+ }
371
+
372
+ function cancelSynthesis() {
373
+ synthesisController.abort();
374
+ }
375
+ ```
376
+
377
+ Call `synthesizeOneSentence()` from an owned UI action and catch its rejection;
378
+ wire `cancelSynthesis()` to its Cancel control. The controller cancels the
379
+ individual `fetchTTS()` request; it does not control the preceding
380
+ `setSpeechMuted(false)` load/unmute lifecycle. Use a fresh controller for each
381
+ new operation. Configuration accepts `configureBrowserSpeech(configuration,
382
+ {signal})`; disposal accepts `disposeBrowserSpeech({signal})`. The streaming
383
+ playback methods do not accept a caller signal; wire
384
+ your stream's abort action to `ai.stopAudio()` as well as aborting its producer.
385
+ Cancellation suppresses late results, but upstream Kokoro may finish active
386
+ engine work before the affected slot is reusable.
387
+
388
+ ```javascript
389
+ const textStreamController = new AbortController();
390
+ textStreamController.signal.addEventListener('abort', function stopStreamAudio() {
391
+ ai.stopAudio();
392
+ }, { once: true });
393
+
394
+ function cancelTextAndSpeech() {
395
+ textStreamController.abort();
396
+ }
397
+ ```
398
+
399
+ Pass that same `textStreamController.signal` to your text producer's supported
400
+ signal option. Connect `cancelTextAndSpeech()` to Cancel; use a new controller
401
+ for the next stream.
402
+
403
+ ## Advanced provider and artifact reference
404
+
405
+ The SDK does not choose a runtime, model, voice, catalog, prompt, or product
406
+ policy. Applications keep those choices. Nothing is downloaded or activated
407
+ until the application explicitly calls `load()`.
408
+
409
+ Ordinary speech operation is the complete functional path. It uses the selected
410
+ upstream Transformers or Kokoro package and the browser's normal networking,
411
+ Worker, and Cache APIs. The records returned by this entrypoint are ordinary
412
+ JavaScript objects and arrays. Callers may copy, extend, and present complete
413
+ records; this contract does not freeze them or shorten their content.
414
+
415
+ ## Availability
416
+
417
+ | Host | Availability | Notes |
418
+ | --- | --- | --- |
419
+ | Browser | Shipped | Requires Workers, Fetch, Blob/File, object URLs, DBOPFS/OPFS, and Web Locks. Blob/File STT requests also require the browser audio decoder. |
420
+ | Native WebView | Conditional | Available when the WebView exposes the browser APIs above. It does not invoke Core speech. |
421
+ | Node | Importable, execution unavailable | The ESM subpath imports, but the SDK supplies no Node speech storage, Worker, or audio-decoder host. |
422
+ | Cloud | Not offered | The SDK's built-in speech profile is device-only: Whisper owns STT and Kokoro owns TTS. |
423
+
424
+ STT and TTS own independent provider lifecycles. A failure or cancellation in
425
+ one role does not disable the other role or authorize a fallback provider.
426
+
427
+ ## Public exports
428
+
429
+ ```javascript
430
+ import {
431
+ BROWSER_SPEECH_ARTIFACT_GRAPH_PROTOCOL,
432
+ BROWSER_SPEECH_ARTIFACT_PROTOCOL,
433
+ createBrowserKokoroProvider,
434
+ createBrowserSpeechArtifactGraph,
435
+ createBrowserSpeechAuthority,
436
+ createBrowserWhisperProvider,
437
+ createDbopfsSpeechArtifactStore
438
+ } from 'arcane-os/ai/browser-speech';
439
+ ```
440
+
441
+ Importing this entrypoint downloads nothing, opens no cache, creates no Worker,
442
+ and publishes no event.
443
+
444
+ ## Protocol identifiers
445
+
446
+ These exact strings identify the current public contracts:
447
+
448
+ | Subject | Exact value |
449
+ | --- | --- |
450
+ | Artifact-store protocol | `arcane-ai-browser-speech-artifacts/1` |
451
+ | Artifact-graph protocol | `arcane-ai-browser-speech-artifact-graph/1` |
452
+ | Graph `kind` and prepared `runtime.moduleGraph` | `browser-speech-authenticated-artifact-graph` |
453
+ | Single-module `runtime.moduleGraph` | `self-contained` |
454
+ | Model authority | `arcane-ai-model-authority/1` |
455
+ | Provider | `arcane-ai-provider/2` |
456
+ | Worker | `arcane-ai-speech-worker/1` |
457
+ | Worker error envelope | `arcane-ai-speech-worker-error/1` |
458
+ | Nested module Worker | `arcane-ai-browser-speech-artifact-module-worker/1` |
459
+
460
+ The word `authenticated` in the graph discriminator does not activate an
461
+ authentication, admission, or isolation stage. It is the current protocol
462
+ value.
463
+
464
+ ## `createBrowserSpeechArtifactGraph()`
465
+
466
+ An artifact graph describes the caller-selected runtime, model, voice, and
467
+ supporting files that the SDK stores and materializes. It is a routing and
468
+ selection record, not an execution permission list.
469
+
470
+ ```javascript
471
+ const graph = createBrowserSpeechArtifactGraph({
472
+ providerId: 'my-whisper',
473
+ role: 'stt',
474
+ model: {
475
+ id: 'whisper-small',
476
+ repository: 'publisher/whisper-small',
477
+ revision: 'selected-model-revision',
478
+ dtype: 'q8',
479
+ inputSampleRate: 16000
480
+ },
481
+ runtime: {
482
+ adapter: 'transformers-whisper',
483
+ version: 'selected-runtime-version',
484
+ revision: 'selected-runtime-revision',
485
+ entrypoint: 'runtime/transformers.js',
486
+ onnxWasm: {
487
+ namespace: 'transformers-env-backends-onnx-wasm',
488
+ mjsPath: 'runtime/ort-wasm.mjs',
489
+ wasmPath: 'runtime/ort-wasm.wasm'
490
+ }
491
+ },
492
+ files: [
493
+ {
494
+ kind: 'runtime-entrypoint-javascript',
495
+ path: 'runtime/transformers.js',
496
+ sourceUrl: 'https://publisher.example/transformers.js',
497
+ revision: 'selected-runtime-revision',
498
+ mediaType: 'text/javascript'
499
+ },
500
+ {
501
+ kind: 'runtime-auxiliary-javascript',
502
+ path: 'runtime/ort-wasm.mjs',
503
+ sourceUrl: 'https://publisher.example/ort-wasm.mjs',
504
+ revision: 'selected-runtime-revision',
505
+ mediaType: 'text/javascript'
506
+ },
507
+ {
508
+ kind: 'runtime-wasm-binary',
509
+ path: 'runtime/ort-wasm.wasm',
510
+ sourceUrl: 'https://publisher.example/ort-wasm.wasm',
511
+ revision: 'selected-runtime-revision',
512
+ mediaType: 'application/wasm'
513
+ },
514
+ {
515
+ kind: 'model-onnx-binary',
516
+ path: 'model/encoder.onnx',
517
+ sourceUrl: 'https://publisher.example/encoder.onnx',
518
+ revision: 'selected-model-revision',
519
+ mediaType: 'application/octet-stream',
520
+ runtimeRequestUrls: [
521
+ 'https://publisher.example/model/encoder.onnx'
522
+ ]
523
+ }
524
+ ]
525
+ });
526
+ ```
527
+
528
+ ### Model and runtime selection
529
+
530
+ `role` is `stt` or `tts`. The runtime adapter is
531
+ `transformers-whisper` for STT and `kokoro-js` for TTS. The caller supplies the
532
+ model id, repository, revision, dtype, sample rate, and runtime version and
533
+ revision.
534
+
535
+ STT requires `model.inputSampleRate`. TTS requires
536
+ `model.outputSampleRate`, `model.defaultVoice`, and a nonempty
537
+ `model.voices` array of `{id,path}` records. Each voice path names a declared
538
+ `voice-style-binary` file. `runtime.onnxWasm` names the selected ONNX module and
539
+ WASM files. `numThreads` is optional for Transformers and is not inferred from
540
+ hardware. Kokoro does not expose that field.
541
+
542
+ ### File records
543
+
544
+ Each file record uses:
545
+
546
+ ```text
547
+ {
548
+ kind,
549
+ path,
550
+ sourceUrl,
551
+ revision,
552
+ license?,
553
+ mediaType,
554
+ sourceMediaType?,
555
+ runtimeRequestUrls?
556
+ }
557
+ ```
558
+
559
+ `path` is a normalized relative path. `sourceUrl` is the caller-selected source
560
+ used for installation. `runtimeRequestUrls` lists aliases used by upstream
561
+ runtime code for the same stored file. `mediaType` becomes the materialized
562
+ Blob type; `sourceMediaType` may describe a different upstream response type.
563
+ The SDK does not require or interpret legal metadata at runtime. If the caller
564
+ includes `license`, the graph preserves that complete value as inert metadata;
565
+ runtime materialization never treats it as capability or admission data.
566
+
567
+ Runtime file kinds are `runtime-entrypoint-javascript`,
568
+ `runtime-auxiliary-javascript`, `runtime-wasm-binary`, and
569
+ `runtime-opaque-data`. Model/data kinds are `model-configuration-json`,
570
+ `model-generation-configuration-json`, `model-onnx-binary`,
571
+ `model-onnx-external-data`, `model-preprocessor-json`,
572
+ `model-tokenizer-json`, `model-opaque-data`, and `voice-style-binary`.
573
+
574
+ Graph construction requires paths and route aliases to be unambiguous so one
575
+ known URL maps to at most one stored file. The runtime router is independently
576
+ permissive: if ambiguous routing metadata nevertheless reaches it, that
577
+ URL is left unmapped and uses the native browser operation.
578
+
579
+ ## `createDbopfsSpeechArtifactStore()`
580
+
581
+ ```javascript
582
+ const store = createDbopfsSpeechArtifactStore({
583
+ dbopfs,
584
+ tableName: 'arcane_ai_browser_speech'
585
+ });
586
+ ```
587
+
588
+ The store exposes `{protocol,tableName,prepare,remove}`. It serializes updates
589
+ to one selected authority with Web Locks, downloads each caller-selected file
590
+ after explicit provider activation, stores it in DBOPFS, and reopens the stored
591
+ file before materialization. Missing storage, an unreadable response, a failed
592
+ HTTP request, a missing stored file, or cancellation rejects honestly.
593
+
594
+ The store writes ordinary mutable selection metadata before the selected files.
595
+ On a later load, a changed file inventory or source mapping is a cache miss and
596
+ is downloaded again. The selection metadata is not a completion, integrity, or
597
+ publication receipt and never blocks ordinary loading.
598
+
599
+ `prepare(authority,{signal,onProgress,offline=false,security})` accepts an
600
+ SDK-created artifact graph or upstream-package authority. `offline:true` uses
601
+ only existing DBOPFS state. Preparation returns the selected runtime/model
602
+ configuration, `cache` as `installed` or `cached`, object URLs, and a `release()`
603
+ function that revokes the materialized URLs.
604
+
605
+ `security` records caller intent only. Ordinary preparation does not
606
+ forward a security payload to the Worker and performs no security work. Passing
607
+ `{secure:true}` does not activate hardening. Any future hardening stage requires
608
+ a separate user review and an explicit implementation change before it may execute.
609
+
610
+ ## Ordinary module routing
611
+
612
+ Artifact-graph preparation reads each stored JavaScript module, discovers
613
+ ordinary module operations, and materializes the complete stored files as
614
+ object URLs. The prepared runtime retains the existing
615
+ `browser-speech-authenticated-artifact-graph` discriminator.
616
+
617
+ The Worker installs one private module router before importing the entrypoint:
618
+
619
+ - a static import whose target is a known stored file uses that file's
620
+ materialized URL;
621
+ - a dynamic import whose target is known imports the materialized URL;
622
+ - a fetch whose target is known reads the materialized URL;
623
+ - a Worker whose target is known starts the SDK role Worker and imports the
624
+ materialized target there; and
625
+ - a Cache match whose target is known returns the materialized file.
626
+
627
+ Every unmapped operation keeps ordinary browser behavior:
628
+
629
+ - an unmapped relative or URL-like import resolves against the calling module's
630
+ original source URL, while a bare specifier remains unchanged for native
631
+ import-map resolution;
632
+ - an unmapped fetch calls native `fetch` and preserves the caller's options;
633
+ - an unmapped Worker calls the native `Worker` constructor and preserves the
634
+ caller's options; and
635
+ - an unmapped Cache operation delegates to native Cache Storage.
636
+
637
+ Cache `put`, `add`, `addAll`, `delete`, and `keys` are normal mutable browser
638
+ operations. The SDK does not replace them with a read-only cache. Relative
639
+ requests are resolved from the calling module's original source URL before the
640
+ native Cache operation.
641
+
642
+ Routing discovery is best effort and is not an admission gate. If the scanner
643
+ cannot interpret a module, that module is left unchanged and follows its native
644
+ URLs. A static-import cycle may likewise retain an original source URL where a
645
+ target has not yet been materialized. These fallbacks preserve functionality;
646
+ they do not silently convert into a rejection policy.
647
+
648
+ ## ONNX runtime configuration
649
+
650
+ The Worker applies only the selected runtime settings needed to run the chosen
651
+ provider:
652
+
653
+ - Kokoro forwards the pool's selected `webgpu` or `wasm` device to
654
+ `KokoroTTS.from_pretrained()`. Its configured dtype remains exactly the
655
+ caller-selected dtype on both paths. The WASM path also uses
656
+ `namespace.env.wasmPaths = {mjs,wasm}`.
657
+ - Transformers uses
658
+ `namespace.env.backends.onnx.wasm.wasmPaths = {mjs,wasm}`, keeps remote model
659
+ loading enabled, and applies caller-selected `numThreads` when present.
660
+
661
+ Transformers and Kokoro keep their normal provider downloads and Cache behavior
662
+ for routes not materialized by the SDK. Kokoro voice aliases and Transformers
663
+ model aliases may be listed in `runtimeRequestUrls` so an upstream request for a
664
+ known file resolves to the already materialized local file. Model, voice, and
665
+ runtime selection remains with the application and upstream publisher.
666
+
667
+ ## Providers
668
+
669
+ ```javascript
670
+ const whisper = createBrowserWhisperProvider({
671
+ id: 'my-whisper',
672
+ graph,
673
+ store
674
+ });
675
+
676
+ const kokoro = createBrowserKokoroProvider({
677
+ id: 'my-kokoro',
678
+ graph: kokoroGraph,
679
+ store,
680
+ execution: {
681
+ device: 'auto',
682
+ maxConcurrentRequests: 4
683
+ }
684
+ });
685
+ ```
686
+
687
+ The constructors also accept ordinary `model` and `runtime` descriptors instead
688
+ of `graph`; the two forms are mutually exclusive. Both forms require an
689
+ SDK-created DBOPFS speech artifact store. `localOnly` remains `true`.
690
+
691
+ Kokoro additionally accepts the exact `execution` record
692
+ `{device,maxConcurrentRequests}`. `device` is `auto`, `webgpu`, or `wasm`;
693
+ `maxConcurrentRequests` is an integer from 1 through 4. Omission defaults to
694
+ `{device:'auto',maxConcurrentRequests:4}`. `auto` attempts a complete WebGPU
695
+ Worker pool only when the browser exposes WebGPU. If that pool cannot load, the
696
+ SDK tears it down and creates a complete WASM pool with the same caller-selected
697
+ model and dtype. Explicit `webgpu` rejects when WebGPU cannot load; explicit
698
+ `wasm` never attempts GPU. Whisper remains one WASM Worker and does not accept
699
+ this option.
700
+
701
+ Each constructor returns an `arcane-ai-provider/2` object with:
702
+
703
+ ```text
704
+ {
705
+ protocol,
706
+ role,
707
+ id,
708
+ localOnly,
709
+ maxConcurrentRequests,
710
+ catalog,
711
+ inspect,
712
+ status,
713
+ load,
714
+ request,
715
+ unload,
716
+ dispose
717
+ }
718
+ ```
719
+
720
+ `catalog()`, `inspect()`, and `status()` do not activate a provider. `load()` is
721
+ the explicit activation boundary. The caller supplies model/profile policy and
722
+ may display the provider's lifecycle status. `unload()` releases the Worker and
723
+ materialized URLs. `dispose()` performs final teardown and prevents later use.
724
+
725
+ ### Upstream-package authority
726
+
727
+ `createBrowserSpeechAuthority({providerId,role,model,runtime,security})` creates
728
+ the ordinary single-entrypoint authority. The model descriptor is
729
+ `{id,repository,revision,dtype?,defaultVoice?,files?}`. The runtime descriptor is
730
+ `{adapter,version,revision,entry,wasmPaths?,files}`. A file is
731
+ `{path,url,mediaType?}`. Model files may be omitted so the selected upstream
732
+ provider performs its normal model and voice downloads after explicit use.
733
+
734
+ The authority record is a mutable complete record. An omitted ordinary security
735
+ option produces no security field. A present `{secure:true}` value records only
736
+ the caller's future intent and does not change loading or routing behavior.
737
+
738
+ ### Whisper STT
739
+
740
+ ```javascript
741
+ const result = await whisper.request({
742
+ role: 'stt',
743
+ operation: 'transcribe',
744
+ signal,
745
+ payload: {
746
+ audio: pcmFloat32,
747
+ sampleRate: 16000
748
+ }
749
+ });
750
+ ```
751
+
752
+ The provider-native payload is mono `Float32Array` PCM at the selected input
753
+ sample rate, and the result is `{text}`. The shared AI form accepts
754
+ `{audio:Blob|File,mimeType,model}` and uses the browser decoder to produce the
755
+ same mono input. The SDK preserves the complete returned transcript.
756
+
757
+ ### Kokoro TTS
758
+
759
+ ```javascript
760
+ const result = await kokoro.request({
761
+ role: 'tts',
762
+ operation: 'synthesize',
763
+ signal,
764
+ payload: {
765
+ text: 'Hello from Arcane.',
766
+ voice: 'caller-voice-id',
767
+ speed: 1
768
+ }
769
+ });
770
+ ```
771
+
772
+ The voice belongs to the caller-selected inventory; omission uses that model's
773
+ default voice. The provider-native result is
774
+ `{audio:Float32Array,sampleRate,voice}`. The provider/2 shared request form accepts
775
+ `{model,input,responseFormat:'wav',voice?,speed?}` and returns
776
+ `{audio:Uint8Array,contentType:'audio/wav'}`. High-level `AI.fetchTTS()` wraps
777
+ that provider result in a WAV `Blob`. Returned provider records remain ordinary
778
+ mutable values.
779
+
780
+ ## Lifecycle and cancellation
781
+
782
+ Provider states are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
783
+ `disposed`. `status()` includes role, provider/model ids, state, lifecycle
784
+ status and reason, active operation, loaded/busy flags, generation, error code,
785
+ cache state, and warnings. Kokoro status also includes an `execution` record
786
+ with requested and selected device, request limit, and active request count. A
787
+ successful `selectedDevice:'webgpu'` reports the execution provider selected by
788
+ the upstream model load; it does not claim that browser, driver, or GPU kernels
789
+ overlap physically. A security field is absent in ordinary mode.
790
+
791
+ The provider/2 load context accepts an optional progress callback for interface
792
+ compatibility, but the current browser-speech artifact and Worker transport
793
+ publishes no progress records. Consumers present the explicit lifecycle states
794
+ instead of inventing a numeric total.
795
+
796
+ Compatible concurrent loads share one underlying preparation and pool load
797
+ while each caller retains its own cancellation signal. One observer cannot
798
+ cancel another still-active observer; cancellation of the final observer stops
799
+ the shared load.
800
+
801
+ Whisper retains one active role request. Kokoro admits synthesis requests up to
802
+ its declared capacity and rejects a direct over-capacity call with
803
+ `ARCANE_AI_PROVIDER_BUSY`; the provider-neutral runtime keeps overflow in FIFO
804
+ order. Each Kokoro slot owns a distinct Worker and loaded model session because
805
+ the selected browser adapter serializes inference inside one JavaScript
806
+ isolate. The SDK prepares the artifact URLs once and shares that same prepared
807
+ selection across the bounded pool. Pool activation completes the first model
808
+ session before it starts the remaining Workers, then loads those remaining
809
+ sessions concurrently. This avoids multiplying simultaneous cold artifact
810
+ acquisition while still making the configured synthesis capacity ready in
811
+ parallel.
812
+
813
+ Cancellation of one active Kokoro synthesis suppresses only that request and
814
+ sends the Worker's targeted cancel control. The selected upstream Kokoro
815
+ version may finish already-running engine work before that slot can run its
816
+ next request; the SDK does not claim stronger per-call preemption. Whisper
817
+ cancellation and role unload/dispose retain destructive Worker teardown.
818
+ Kokoro unload/dispose abort every active request, terminates every pool Worker,
819
+ and releases materialized URLs once. Late results cannot settle a cancelled or
820
+ superseded operation.
821
+
822
+ The provider owns no event bus. Applications may project promises and status
823
+ into the SDK's shared event/state owner. Mute and unmute are likewise
824
+ application state: mute may await `unload()`, and unmute may explicitly call
825
+ `load()`.
826
+
827
+ ## Errors
828
+
829
+ Errors retain the normal `code`, `message`, `reason`, and `cause` fields used by
830
+ the browser AI runtime.
831
+
832
+ The Worker error envelope carries `cause` as an optional mutable diagnostic
833
+ record. It preserves complete nested messages, stacks, codes, reasons, details,
834
+ own properties, and cycles without a depth or content cap. Worker and client
835
+ sources must come from the same SDK revision so their `/1` envelope shape is
836
+ updated atomically; a current client still accepts the cause-free four-field form.
837
+ If the platform cannot clone an exotic diagnostic value, the Worker keeps the
838
+ complete raw failure in its console diagnostics and retries the response with
839
+ that cause-free four-field envelope so the caller still receives an error.
840
+
841
+ Representative stable codes include:
842
+
843
+ - `ARCANE_AI_INVALID_REQUEST`
844
+ - `ARCANE_AI_MODEL_AUTHORITY_REQUIRED`
845
+ - `ARCANE_AI_PROVIDER_UNAVAILABLE`
846
+ - `ARCANE_AI_PROVIDER_BUSY`
847
+ - `ARCANE_AI_PROVIDER_DISPOSED`
848
+ - `ARCANE_AI_REQUEST_ABORTED`
849
+ - `ARCANE_AI_OPERATION_SUPERSEDED`
850
+ - `ARCANE_AI_ARTIFACT_DOWNLOAD_FAILED`
851
+ - `ARCANE_AI_ARTIFACT_OFFLINE_MISS`
852
+ - `ARCANE_AI_STORAGE_BUSY`
853
+ - `ARCANE_AI_STORAGE_UNAVAILABLE`
854
+ - `ARCANE_AI_WORKER_MESSAGE_ERROR`
855
+
856
+ Malformed selected descriptors, missing required files, unreadable responses,
857
+ unsupported provider namespace shapes, and unavailable browser APIs reject at
858
+ their functional owner. An unmapped runtime route retains the native browser
859
+ operation.
860
+
861
+ ## Ownership
862
+
863
+ - Applications own model, runtime, dtype, sample-rate, voice, profile, prompt,
864
+ catalog, activation, optional TTS execution override, and presentation
865
+ policy.
866
+ - Upstream publishers own their runtime, model, voice, and license delivery.
867
+ - The SDK owns storage, materialization, routing, Worker lifecycle, normalized
868
+ provider contracts, cancellation, and cleanup.
869
+ - The SDK redistributes no third-party runtime, model, or voice package through
870
+ this entrypoint.
871
+
872
+ ## Related
873
+
874
+ - [Normalized AI](../README.md#normalized-ai)
875
+ - [Browser-WASM LLM](browser-wasm.md)
876
+ - [AIProviderRuntime.js](../runtime-modules.md#aiproviderruntimejs)
877
+ - [AIRuntimeState.js](../runtime-modules.md#airuntimestatejs)
878
+ - [Availability and normalization](../availability-and-normalization.md)
879
+ - [Protocol architecture](../protocols.md#portable-ai-provider-runtime)