arcane-os 0.5.19 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,20 @@
1
1
  # Changelog
2
2
 
3
+ ## 0.6.0
4
+
5
+ - Prefer WebNN NPU, then WebGPU, then CPU through WASM for browser Whisper
6
+ transcription and Kokoro synthesis. Skip absent accelerator APIs and replace
7
+ failed model-load Workers before trying the next backend with the same
8
+ application-selected model and precision.
9
+ - Add Whisper execution configuration and NPU selection for both speech roles.
10
+ Explicit `webnn-npu`, `webgpu`, and `wasm` choices report failure without
11
+ falling back. Whisper retains one slot; Kokoro retains four by default and
12
+ accepts capacities from one through four.
13
+ - Report requested and successfully selected backends for both roles through
14
+ provider execution status. Selection reports upstream session loading;
15
+ actual accelerator use and speech quality depend on the selected runtime,
16
+ model, browser, drivers, and hardware. Wllama remains on its WebGPU backend.
17
+
3
18
  ## 0.5.19
4
19
 
5
20
  - Report observed Wllama initialization stages and runtime activity through
package/README.md CHANGED
@@ -19,7 +19,7 @@ version-locked SDK runtime, while an integrated Arcane checkout uses its live
19
19
  `arcane/` runtime. Both profiles preserve the same app URLs, theme, packaging,
20
20
  event, cancellation, and browser run contracts.
21
21
 
22
- This checkout defines the `0.5.19` SDK contract. Applications pin one exact npm
22
+ This checkout defines the `0.6.0` SDK contract. Applications pin one exact npm
23
23
  version and lockfile; registry state is deliberately not baked into application
24
24
  artifacts.
25
25
 
@@ -35,7 +35,7 @@ Create one browser application, install its pinned SDK, and start its source
35
35
  server:
36
36
 
37
37
  ```bash
38
- npx arcane-os@0.5.19 new hello-speech --path ./hello-speech --target browser
38
+ npx arcane-os@0.6.0 new hello-speech --path ./hello-speech --target browser
39
39
  cd hello-speech
40
40
  npm install
41
41
  npm run dev
@@ -299,10 +299,16 @@ provider factories. The package contains the plain-JavaScript provider and
299
299
  Worker machinery, not speech runtimes, models, voices, or a CDN default. An app
300
300
  must supply each runtime/model selection explicitly. Speech roles
301
301
  load, cancel, unload, fail, and recover independently, so speech failure never
302
- silently falls back or prevents text chat. Kokoro defaults to a four-slot Worker
303
- and model-session pool: it selects WebGPU when the browser can load the complete
304
- pool and otherwise recreates that pool on WASM. Apps may select `webgpu` or
305
- `wasm` explicitly and may set the bounded TTS capacity from one through four.
302
+ switches to a cloud provider or prevents text chat. Both roles default to
303
+ NPU, GPU, then CPU automatic selection: WebNN NPU when its browser API is
304
+ exposed, then WebGPU, then WASM. Failed model loads release their Workers before
305
+ the next backend uses the same model and precision. Apps may select
306
+ `webnn-npu`, `webgpu`, or `wasm` explicitly without automatic fallback.
307
+ Whisper keeps one transcription slot; Kokoro defaults to four Worker/model
308
+ sessions and accepts capacities from one through four. Both roles expose
309
+ requested and selected devices through execution status. A selected backend
310
+ reports successful upstream loading, not physical accelerator use for every
311
+ operation; model and hardware compatibility remain upstream.
306
312
 
307
313
  Applications that need faster spoken-response onset can configure the shared
308
314
  TTS stream without taking over synthesis or playback:
@@ -357,7 +363,7 @@ uses the same controller for automatic memory extraction.
357
363
  Create a new repository-shaped Arcane application with the exact stable SDK:
358
364
 
359
365
  ```bash
360
- npx arcane-os@0.5.19 new my-app --path ./my-app --target portable --git
366
+ npx arcane-os@0.6.0 new my-app --path ./my-app --target portable --git
361
367
  cd my-app
362
368
  npm install
363
369
  npm run dev
@@ -367,7 +373,7 @@ To enroll an existing repository, install the exact SDK and initialize only
367
373
  missing Arcane files:
368
374
 
369
375
  ```bash
370
- npm install --save-dev --save-exact arcane-os@0.5.19
376
+ npm install --save-dev --save-exact arcane-os@0.6.0
371
377
  npm exec -- arcane init my-app --target portable
372
378
  ```
373
379
 
@@ -383,7 +389,7 @@ npm exec -- arcane-os targets
383
389
  No global SDK install or standalone Arcane CLI is required. The application
384
390
  repository's exact npm dependency and lockfile own the CLI and toolchain version.
385
391
 
386
- Use `npx arcane-os@0.5.19` for the initial bootstrap because it names this npm
392
+ Use `npx arcane-os@0.6.0` for the initial bootstrap because it names this npm
387
393
  package explicitly; bare `npx arcane` outside an installed project could resolve
388
394
  a different package. Both installed commands invoke the same headless toolchain.
389
395
  Project-local npm scripts use the SDK pinned by that app's `package-lock.json`,
@@ -403,7 +409,7 @@ node ./bin/arcane.mjs new local-app --path ../local-app --target portable --git
403
409
 
404
410
  # From the generated app repository
405
411
  cd ../local-app
406
- npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.5.19.tgz
412
+ npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.6.0.tgz
407
413
  npm ci
408
414
  ```
409
415
 
@@ -412,7 +418,7 @@ same location. The lockfile retains the selected package dependency while
412
418
  Arcane uses the installed package name and version. Local directory `file:` dependencies are not
413
419
  accepted because npm may install them as links; use a packed `.tgz`. A GitHub
414
420
  runner also needs that tarball at the locked path. After publication, replace
415
- the local declaration with the exact `arcane-os@0.5.19` registry package and
421
+ the local declaration with the exact `arcane-os@0.6.0` registry package and
416
422
  commit the regenerated lock.
417
423
 
418
424
  Generated repositories use `npm ci --ignore-scripts` in CI. Run dependency
@@ -551,7 +557,7 @@ package installation, or assertions.
551
557
 
552
558
  ## Current target support
553
559
 
554
- Version `0.5.19` exposes one browser target and five explicitly paired
560
+ Version `0.6.0` exposes one browser target and five explicitly paired
555
561
  native development targets: a non-runnable portable directory, a
556
562
  Windows x64 unsigned-local-test EXE bundle, Linux x64 and Linux ARM64
557
563
  unsigned-local-test DEBs, and an Android development-signed APK. The
@@ -23,8 +23,8 @@ const ROLE_OPERATION = completeValue({ stt: "transcribe", tts: "synthesize" });
23
23
  const STT_SAMPLE_RATE = 16_000;
24
24
  const TTS_SAMPLE_RATE = 24_000;
25
25
  const TTS_RESPONSE_FORMAT = "wav";
26
- const TTS_EXECUTION_DEVICES = new Set(["auto", "webgpu", "wasm"]);
27
- const DEFAULT_TTS_EXECUTION_DEVICE = "auto";
26
+ const SPEECH_EXECUTION_DEVICES = new Set(["auto", "webnn-npu", "webgpu", "wasm"]);
27
+ const DEFAULT_SPEECH_EXECUTION_DEVICE = "auto";
28
28
  const DEFAULT_TTS_MAX_CONCURRENT_REQUESTS = 4;
29
29
  const MAX_TTS_CONCURRENT_REQUESTS = 4;
30
30
  const ROLE_REQUEST_REASON = completeValue({
@@ -150,24 +150,21 @@ function requiredIdentifier(value, label) {
150
150
  }
151
151
 
152
152
  function normalizeSpeechExecution(role, execution) {
153
- if (role === "stt") {
154
- if (execution !== undefined) {
155
- throw new TypeError("Browser Whisper does not accept an execution option.");
156
- }
157
- return completeValue({ device: "wasm", maxConcurrentRequests: 1 });
158
- }
153
+ const label = role === "stt" ? "Browser Whisper" : "Browser Kokoro";
154
+ const defaultConcurrency = role === "stt" ? 1 : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
155
+ const maximumConcurrency = role === "stt" ? 1 : MAX_TTS_CONCURRENT_REQUESTS;
159
156
  if (execution === undefined) {
160
157
  return completeValue({
161
- device: DEFAULT_TTS_EXECUTION_DEVICE,
162
- maxConcurrentRequests: DEFAULT_TTS_MAX_CONCURRENT_REQUESTS,
158
+ device: DEFAULT_SPEECH_EXECUTION_DEVICE,
159
+ maxConcurrentRequests: defaultConcurrency,
163
160
  });
164
161
  }
165
162
  if (!execution || !is.object(execution) || is.array(execution)) {
166
- throw new TypeError("Browser Kokoro execution must be a plain data record.");
163
+ throw new TypeError(`${label} execution must be a plain data record.`);
167
164
  }
168
165
  const prototype = Object.getPrototypeOf(execution);
169
166
  if (prototype !== Object.prototype && prototype !== null) {
170
- throw new TypeError("Browser Kokoro execution must be a plain data record.");
167
+ throw new TypeError(`${label} execution must be a plain data record.`);
171
168
  }
172
169
  const descriptors = Object.getOwnPropertyDescriptors(execution);
173
170
  for (const key of Reflect.ownKeys(descriptors)) {
@@ -175,34 +172,35 @@ function normalizeSpeechExecution(role, execution) {
175
172
  (key !== "device" && key !== "maxConcurrentRequests")
176
173
  || !Object.hasOwn(descriptors[key], "value")
177
174
  ) {
178
- throw new TypeError("Browser Kokoro execution contains an unsupported or accessor field.");
175
+ throw new TypeError(`${label} execution contains an unsupported or accessor field.`);
179
176
  }
180
177
  }
181
178
  const device = Object.hasOwn(descriptors, "device")
182
179
  ? descriptors.device.value
183
- : DEFAULT_TTS_EXECUTION_DEVICE;
180
+ : DEFAULT_SPEECH_EXECUTION_DEVICE;
184
181
  const maxConcurrentRequests = Object.hasOwn(descriptors, "maxConcurrentRequests")
185
182
  ? descriptors.maxConcurrentRequests.value
186
- : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
187
- if (!TTS_EXECUTION_DEVICES.has(device)) {
188
- throw new TypeError('Browser Kokoro execution.device must be "auto", "webgpu", or "wasm".');
183
+ : defaultConcurrency;
184
+ if (!SPEECH_EXECUTION_DEVICES.has(device)) {
185
+ throw new TypeError(`${label} execution.device must be "auto", "webnn-npu", "webgpu", or "wasm".`);
189
186
  }
190
187
  if (
191
188
  !is.safeInteger(maxConcurrentRequests)
192
189
  || maxConcurrentRequests < 1
193
- || maxConcurrentRequests > MAX_TTS_CONCURRENT_REQUESTS
190
+ || maxConcurrentRequests > maximumConcurrency
194
191
  ) {
195
- throw new TypeError("Browser Kokoro execution.maxConcurrentRequests must be a safe integer from 1 through 4.");
192
+ throw new TypeError(`${label} execution.maxConcurrentRequests must be a safe integer from 1 through ${maximumConcurrency}.`);
196
193
  }
197
194
  return completeValue({ device, maxConcurrentRequests });
198
195
  }
199
196
 
200
- function navigatorHasWebGpu() {
201
- try {
202
- return Boolean(globalThis.navigator?.gpu);
203
- } catch {
204
- return false;
205
- }
197
+ function speechExecutionDevices(requestedDevice) {
198
+ if (requestedDevice !== "auto") return [requestedDevice];
199
+ const devices = [];
200
+ if (is.function(globalThis.navigator?.ml?.createContext)) devices.push("webnn-npu");
201
+ if (globalThis.navigator?.gpu) devices.push("webgpu");
202
+ devices.push("wasm");
203
+ return devices;
206
204
  }
207
205
 
208
206
  function createProviderAuthority({
@@ -1187,14 +1185,12 @@ function createBrowserSpeechProvider({
1187
1185
  cache,
1188
1186
  lifecycleReason,
1189
1187
  activeOperation,
1190
- execution: role === "tts"
1191
- ? completeValue({
1192
- requestedDevice: speechExecution.device,
1193
- selectedDevice,
1194
- maxConcurrentRequests: speechExecution.maxConcurrentRequests,
1195
- activeRequestCount: requestOperations.size,
1196
- })
1197
- : null,
1188
+ execution: completeValue({
1189
+ requestedDevice: speechExecution.device,
1190
+ selectedDevice,
1191
+ maxConcurrentRequests: speechExecution.maxConcurrentRequests,
1192
+ activeRequestCount: requestOperations.size,
1193
+ }),
1198
1194
  secureIntent,
1199
1195
  warnings: providerWarnings(runtimeWarnings),
1200
1196
  });
@@ -1535,34 +1531,42 @@ function createBrowserSpeechProvider({
1535
1531
  `${role}-load-superseded-by-unload`,
1536
1532
  );
1537
1533
  }
1538
- const primaryDevice = role === "stt"
1539
- ? "wasm"
1540
- : speechExecution.device === "auto"
1541
- ? navigatorHasWebGpu() ? "webgpu" : "wasm"
1542
- : speechExecution.device;
1543
- try {
1544
- pool = await loadWorkerPool(
1545
- preparation,
1546
- primaryDevice,
1547
- record.warnings,
1548
- linked.controller.signal,
1549
- );
1550
- } catch (error) {
1551
- if (
1552
- role !== "tts"
1553
- || speechExecution.device !== "auto"
1554
- || primaryDevice !== "webgpu"
1555
- || linked.controller.signal.aborted
1556
- || operationGeneration !== generation
1557
- ) {
1558
- throw error;
1534
+ const devices = speechExecutionDevices(speechExecution.device);
1535
+ for (const [index, device] of devices.entries()) {
1536
+ throwIfAborted(linked.controller.signal, `${role}-load-cancelled`);
1537
+ if (operationGeneration !== generation) {
1538
+ throw providerError(
1539
+ "ARCANE_AI_OPERATION_SUPERSEDED",
1540
+ "Browser speech loading was superseded.",
1541
+ undefined,
1542
+ `${role}-load-superseded-by-unload`,
1543
+ );
1544
+ }
1545
+ try {
1546
+ pool = await loadWorkerPool(
1547
+ preparation,
1548
+ device,
1549
+ record.warnings,
1550
+ linked.controller.signal,
1551
+ );
1552
+ break;
1553
+ } catch (error) {
1554
+ if (
1555
+ index === devices.length - 1
1556
+ || linked.controller.signal.aborted
1557
+ || operationGeneration !== generation
1558
+ || error?.code === "ARCANE_AI_REQUEST_ABORTED"
1559
+ || error?.code === "ARCANE_AI_OPERATION_SUPERSEDED"
1560
+ ) {
1561
+ throw error;
1562
+ }
1563
+ // A fresh Worker releases a failed upstream initialization before
1564
+ // the next backend reuses the same prepared model and precision.
1565
+ const warning = `Browser ${role} ${device} loading failed; trying ${devices[index + 1]}.`;
1566
+ record.warnings = [...record.warnings, warning];
1567
+ lastWarnings = record.warnings;
1568
+ console.warn(warning, error);
1559
1569
  }
1560
- pool = await loadWorkerPool(
1561
- preparation,
1562
- "wasm",
1563
- record.warnings,
1564
- linked.controller.signal,
1565
- );
1566
1570
  }
1567
1571
  throwIfAborted(
1568
1572
  linked.controller.signal,
@@ -738,10 +738,9 @@ function validateConfiguration(configuration, role) {
738
738
  if (
739
739
  descriptorMismatch
740
740
  || !Object.hasOwn(descriptors, "device")
741
- || (role === "stt" && descriptors.device.value !== "wasm")
742
- || (role === "tts"
743
- && descriptors.device.value !== "wasm"
744
- && descriptors.device.value !== "webgpu")
741
+ || (descriptors.device.value !== "wasm"
742
+ && descriptors.device.value !== "webgpu"
743
+ && descriptors.device.value !== "webnn-npu")
745
744
  ) {
746
745
  throw workerError(
747
746
  "ARCANE_AI_INVALID_REQUEST",
@@ -1367,7 +1366,7 @@ async function createWhisperEngine(namespace, configuration, signal, report) {
1367
1366
  "automatic-speech-recognition",
1368
1367
  configuration.model.repository,
1369
1368
  {
1370
- device: "wasm",
1369
+ device: configuration.execution?.device ?? "wasm",
1371
1370
  dtype: configuration.model.dtype ?? "fp32",
1372
1371
  revision: configuration.model.revision,
1373
1372
  progress_callback: report,
@@ -279,7 +279,7 @@ paths are withheld from the native provider. The provider copies the complete
279
279
  selected release rather than accepting an unrelated source path. Verification
280
280
  is a separate explicit operation for a selected release artifact.
281
281
 
282
- The SDK `0.5.19` runtime requires Arcane `0.8.12` or newer. Compatibility
282
+ The SDK `0.6.0` runtime requires Arcane `0.8.12` or newer. Compatibility
283
283
  is contractual rather than exact-version pinning: the prepared Core must meet
284
284
  the highest minimum declared by the runtime, selected app, and bundled app
285
285
  dependencies; keep each app's Arcane protocol generation; and provide every
@@ -55,11 +55,11 @@ export const speechSelection = {
55
55
  };
56
56
  ```
57
57
 
58
- The omitted execution record below uses the SDK's WebGPU-first automatic
59
- selection, so this basic configuration uses `fp32` because
58
+ The omitted execution record below uses the SDK's NPU, GPU, then CPU automatic
59
+ selection. This basic configuration uses `fp32` because
60
60
  [Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage).
61
- Automatic fallback carries this same selected model and dtype to WASM; the SDK
62
- does not rewrite the application selection. If you intentionally choose
61
+ Automatic fallback carries this same selected model and dtype between devices;
62
+ the SDK does not rewrite the application selection. If you intentionally choose
63
63
  another dtype, evaluate that exact model, browser, and execution route.
64
64
  `selectedDevice` reports routing after load, not pronunciation, text fidelity,
65
65
  or audio quality.
@@ -229,9 +229,10 @@ playback. `replay()` keeps completed and pending provider segments and retries
229
229
  only failed missing segments.
230
230
 
231
231
  The default `{device:'auto',maxConcurrentRequests:4}` attempts the full ONNX
232
- Worker/session pool on WebGPU, then recreates that pool on WASM only if WebGPU
233
- loading fails. The basic configuration above keeps `fp32` for this WebGPU-first
234
- path and any automatic WASM fallback. Use the status example below to read
232
+ Worker/session pool on WebNN NPU, then WebGPU, then CPU through WASM. It skips
233
+ an accelerator when its browser API is absent and replaces a failed candidate
234
+ with a fresh pool before trying the next device. The basic configuration above
235
+ keeps `fp32` throughout that sequence. Use the status example below to read
235
236
  `selectedDevice`; console node-assignment warnings alone do not identify the
236
237
  selected execution device or assess the generated audio.
237
238
 
@@ -588,8 +589,20 @@ console.log(speechText.append('ing', true)); // ing
588
589
 
589
590
  ## Choose a device or reduce memory use
590
591
 
591
- Omitting `tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`.
592
- These are three alternative configurations, not a sequence of required loads:
592
+ Both speech roles accept `execution:{device,maxConcurrentRequests}`. Omitting
593
+ `stt.execution` selects `{device:'auto',maxConcurrentRequests:1}`; omitting
594
+ `tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`. Whisper keeps
595
+ one transcription slot. Kokoro accepts capacities 1 through 4.
596
+
597
+ Automatic selection tries `webnn-npu` when `navigator.ml.createContext` is
598
+ exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
599
+ API exposure only determines which upstream backend to attempt; it
600
+ does not establish compatible hardware, operators, or model shapes. A failed
601
+ candidate is fully cleaned up before the next candidate uses fresh Workers
602
+ with the same prepared model and dtype. An explicit `webnn-npu`, `webgpu`, or
603
+ `wasm` selection attempts only that backend and reports its failure.
604
+
605
+ These are four alternative TTS configurations:
593
606
 
594
607
  ```javascript
595
608
  async function selectSpeechExecution(execution) {
@@ -609,15 +622,22 @@ async function selectSpeechExecution(execution) {
609
622
 
610
623
  // Choose and call one from your application settings action:
611
624
  // await selectSpeechExecution({ device: 'auto' });
625
+ // await selectSpeechExecution({ device: 'webnn-npu' });
612
626
  // await selectSpeechExecution({ device: 'webgpu' });
613
627
  // await selectSpeechExecution({ device: 'wasm', maxConcurrentRequests: 1 });
614
628
  ```
615
629
 
616
- The override accepts integers 1, 2, 3, or 4. `auto` tries a complete WebGPU pool
617
- when available, then recreates a complete WASM pool if that load cannot finish.
618
- Explicit `webgpu` reports a load error when unavailable; explicit `wasm` never
619
- attempts WebGPU. The same application-selected model and dtype apply on both
620
- devices. A configuration change leaves TTS muted; explicitly load/unmute again.
630
+ The TTS capacity override accepts integers 1, 2, 3, or 4; an STT capacity
631
+ override accepts only 1. Apply the same device choices to the configured
632
+ `stt.execution` record. A configuration change leaves TTS muted; explicitly
633
+ load/unmute again.
634
+
635
+ The selected upstream versions expose WebNN NPU through
636
+ [Transformers.js device selection](https://github.com/huggingface/transformers.js/blob/4.2.0/packages/transformers/src/backends/onnx.js)
637
+ and [Kokoro.js device forwarding](https://github.com/hexgrad/kokoro/blob/664c76a704021239ba59c84dcbaa4d3dece01fe9/kokoro.js/src/kokoro.js).
638
+ WebNN compatibility depends on the model's shapes and operations, browser,
639
+ drivers, and hardware. Unsupported operations may run through WASM even after
640
+ an NPU session loads; see the [ONNX Runtime WebNN contract](https://onnxruntime.ai/docs/tutorials/web/ep-webnn.html).
621
641
 
622
642
  ## Inspect the requested and selected device
623
643
 
@@ -627,15 +647,15 @@ loading. This explicitly reads the selected provider's current report; ordinary
627
647
  separate execution-state event subscription.
628
648
 
629
649
  ```javascript
630
- function printSpeechStatus() {
631
- const status = ai.providerRuntime.status('tts', { execution: true });
650
+ function printSpeechStatus(role = 'tts') {
651
+ const status = ai.providerRuntime.status(role, { execution: true });
632
652
  const execution = status.execution;
633
- console.log('TTS state:', status.state);
653
+ console.log('Speech role and state:', role, status.state);
634
654
  if (execution) {
635
655
  console.log('Requested device:', execution.requestedDevice);
636
656
  console.log('Selected device:', execution.selectedDevice);
637
657
  console.log('Capacity:', execution.maxConcurrentRequests);
638
- console.log('Active synthesis requests:', execution.activeRequestCount);
658
+ console.log('Active requests:', execution.activeRequestCount);
639
659
  console.log('Automatic WASM fallback:',
640
660
  execution.requestedDevice === 'auto' && execution.selectedDevice === 'wasm');
641
661
  }
@@ -645,13 +665,18 @@ function printSpeechStatus() {
645
665
  Call `printSpeechStatus()` after the load in `sayHello()` or from your status
646
666
  button. The same projection is at
647
667
  `ai.providerRuntime.status(null, {execution:true}).roles.tts.execution`.
668
+ For configured Whisper, call `printSpeechStatus('stt')` or read the corresponding
669
+ `roles.stt.execution` projection.
648
670
  `selectedDevice` is `null`
649
671
  until a pool is selected and returns to `null` on unload. Providers without an
650
672
  execution report omit `execution`; do not infer a device from `navigator.gpu`
651
673
  or a configured preference alone. An explicit inspection can throw a provider
652
674
  status error; handle it with the same `error.code` / `error.message` pattern.
653
- Treat `selectedDevice` as route status only; evaluate actual speech output for
654
- the model, dtype, browser, and device combinations your application supports.
675
+ `selectedDevice` names the backend requested by the successful upstream session
676
+ load. It does not prove that every operation ran on a physical NPU or GPU, or
677
+ establish transcription correctness, pronunciation, or audio quality. Evaluate
678
+ actual speech output for the model, dtype, browser, and device combinations your
679
+ application supports.
655
680
 
656
681
  ## Stop, mute, cancel, and release
657
682
 
@@ -991,11 +1016,12 @@ they do not silently convert into a rejection policy.
991
1016
  The Worker applies only the selected runtime settings needed to run the chosen
992
1017
  provider:
993
1018
 
994
- - Kokoro forwards the pool's selected `webgpu` or `wasm` device to
1019
+ - Kokoro forwards the pool's selected `webnn-npu`, `webgpu`, or `wasm` device to
995
1020
  `KokoroTTS.from_pretrained()`. Its configured dtype remains exactly the
996
- caller-selected dtype on both paths. The WASM path also uses
1021
+ caller-selected dtype on every path. The WASM path also uses
997
1022
  `namespace.env.wasmPaths = {mjs,wasm}`.
998
- - Transformers uses
1023
+ - Transformers forwards the selected device to its speech-recognition pipeline,
1024
+ preserves the caller-selected dtype, and uses
999
1025
  `namespace.env.backends.onnx.wasm.wasmPaths = {mjs,wasm}`, keeps remote model
1000
1026
  loading enabled, and applies caller-selected `numThreads` when present.
1001
1027
 
@@ -1029,15 +1055,15 @@ The constructors also accept ordinary `model` and `runtime` descriptors instead
1029
1055
  of `graph`; the two forms are mutually exclusive. Both forms require an
1030
1056
  SDK-created DBOPFS speech artifact store. `localOnly` remains `true`.
1031
1057
 
1032
- Kokoro additionally accepts the exact `execution` record
1033
- `{device,maxConcurrentRequests}`. `device` is `auto`, `webgpu`, or `wasm`;
1034
- `maxConcurrentRequests` is an integer from 1 through 4. Omission defaults to
1035
- `{device:'auto',maxConcurrentRequests:4}`. `auto` attempts a complete WebGPU
1036
- Worker pool only when the browser exposes WebGPU. If that pool cannot load, the
1037
- SDK tears it down and creates a complete WASM pool with the same caller-selected
1038
- model and dtype. Explicit `webgpu` rejects when WebGPU cannot load; explicit
1039
- `wasm` never attempts GPU. Whisper remains one WASM Worker and does not accept
1040
- this option.
1058
+ Both constructors accept the exact `execution` record
1059
+ `{device,maxConcurrentRequests}`. `device` is `auto`, `webnn-npu`, `webgpu`, or
1060
+ `wasm`. Whisper permits capacity 1 and defaults to
1061
+ `{device:'auto',maxConcurrentRequests:1}`. Kokoro permits integer capacities 1
1062
+ through 4 and defaults to `{device:'auto',maxConcurrentRequests:4}`. `auto`
1063
+ tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when
1064
+ `navigator.gpu` is exposed, then WASM. Before advancing after a failed load,
1065
+ the SDK tears down the candidate and creates fresh Workers using the same
1066
+ prepared model and dtype. Explicit device selections never fall back.
1041
1067
 
1042
1068
  Each constructor returns an `arcane-ai-provider/2` object with:
1043
1069
 
@@ -1123,12 +1149,13 @@ mutable values.
1123
1149
  Provider states are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
1124
1150
  `disposed`. `status()` includes role, provider/model ids, state, lifecycle
1125
1151
  status and reason, active operation, loaded/busy flags, generation, error code,
1126
- cache state, and warnings. Kokoro status also includes an `execution` record
1152
+ cache state, and warnings. Both speech roles include an `execution` record
1127
1153
  with requested and selected device, request limit, and active request count. A
1128
- successful `selectedDevice:'webgpu'` reports the execution provider selected by
1129
- the upstream model load; it does not claim that browser, driver, or GPU kernels
1130
- overlap physically or that generated audio has been quality-validated. A
1131
- security field is absent in ordinary mode.
1154
+ successful `selectedDevice` reports the backend requested by the upstream model
1155
+ load. It does not prove that every operation ran on a physical NPU or GPU,
1156
+ that accelerator kernels overlap, or that generated speech is correct. WebNN
1157
+ may execute unsupported operations through WASM. A security field is absent
1158
+ in ordinary mode.
1132
1159
 
1133
1160
  The provider/2 load context accepts an optional progress callback for interface
1134
1161
  compatibility, but the current browser-speech artifact and Worker transport
@@ -1203,7 +1230,7 @@ operation.
1203
1230
  ## Ownership
1204
1231
 
1205
1232
  - Applications own model, runtime, dtype, sample-rate, voice, profile, prompt,
1206
- catalog, activation, optional TTS execution override, and presentation
1233
+ catalog, activation, optional STT/TTS execution override, and presentation
1207
1234
  policy.
1208
1235
  - Upstream publishers own their runtime, model, voice, and license delivery.
1209
1236
  - The SDK owns storage, materialization, routing, Worker lifecycle, normalized
@@ -41,7 +41,7 @@ development evidence; it is not native artifact or release acceptance.
41
41
  | Browser runtime modules | Every shipped ESM module parses and its export inventory matches the catalog; pure helpers run focused success/error cases. | DOM, OPFS, media, and Web Component journeys use a browser harness. |
42
42
  | Provider-neutral AI runtime and chat/speech activation | Provider/2 registration, three-role configuration, TWiN Cloud LLM readiness, on-device Whisper/Kokoro selection, opt-in STT startup, Core speech readiness, independent LLM/STT/TTS load/unload/status, capacity-1 FIFO settlement for LLM/STT, bounded parallel synthesis with FIFO admission for an explicitly capable TTS provider, owned STT signals, TTS mute lifecycle, route-owned voice defaults, immediate chunk synthesis admission, original-order audio-clock scheduling, sticky-state-only readiness for both speech components, selected-unloaded activation request/cancellation/error behavior, programmatic voice recording, transcript-replacement supersession of late settlement, direct `AI.fetchSTT` result delivery, rejection of non-local speech configuration, and absence of silent provider fallback are represented against complete providers and host callbacks. | Real model/runtime availability remains the selected provider's evidence boundary; provider-promise settlement, state, an abort signal, or an activation event does not by itself prove underlying provider work stopped or native, cloud, or browser-model availability. |
43
43
  | Browser-WASM local AI | The exported namespace, canonical ordered `{id, files:[{name?,url},...]}` descriptor, nonempty provider `sources` catalog, default `secure:false`, dormant `secure:true` intent, public AI API module lifecycle, lazy/manual policy, successful Wllama-load requirement, abort normalization, complete output and reasoning, all-choice validation, required structural-call `arguments.message`, exact call identity, and matching tool-result sequencing are represented in deterministic fixture sources. | A real Chrome exercise may load the selected Wllama runtime and model only after explicit user action. It is not an ordinary publication gate or an implicit model download. |
44
- | Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, GPU-first device selection with explicit WASM fallback, bounded Kokoro Worker/session pools, out-of-order synthesis settlement with FIFO admission, request-targeted TTS cancellation, destructive lifecycle cleanup, Blob/File STT conversion, WAV TTS conversion, complete text/audio, and no cloud fallback are represented with synthetic artifacts and adapters. | A real runtime/model/voice download, WebGPU model load, physical accelerator behavior, and actual transcription or synthesis use the application's selected upstream packages/providers, browser media support, and explicit user action. |
44
+ | Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, NPU-to-GPU-to-CPU automatic selection for both roles, fresh Workers after a failed device load, explicit-device failure without fallback, bounded Kokoro Worker/session pools, out-of-order synthesis settlement with FIFO admission, request-targeted TTS cancellation, destructive lifecycle cleanup, Blob/File STT conversion, WAV TTS conversion, complete text/audio, and no cloud fallback are represented with synthetic artifacts and adapters. | A real runtime/model/voice download, WebNN or WebGPU model load, physical accelerator behavior, and actual transcription or synthesis use the application's selected upstream packages/providers, browser media support, and explicit user action. A selected backend is not proof that every operation ran on that physical accelerator; WebNN may execute unsupported operations through WASM. |
45
45
  | Persistent chat and document context | Atomic in-memory/history commit, explicit per-turn persistence, streamed/non-stream fallback, session-owned callback fields, complete data callbacks, per-turn request options, exact per-choice streamed/terminal call correlation before publication, terminal-only call acceptance, ordered parallel-call sequencing, atomic all-ID nonblank executed/declined/cancelled/not-executed result batches, readable unmodified malformed pre-existing rows, complete UI transcript metadata, generic visible failure outcomes with complete console diagnostics, BFCache-preserving component lifecycle, complete bootstrap/search/context, caller-source evaluation, cancellation, and partial-read handling are represented with app-scoped adapters. | Live Core/provider inference and durable browser storage remain separate authorities; tests never treat a fake chat function or in-memory adapter as host/storage proof. |
46
46
  | Core bridge docs | Canonical namespace/method/event/entity inventories match their one-per-member guides and required sections. | Live Core conformance belongs in Arcane OS because Core implementation is not shipped as SDK source. |
47
47
  | Arcane Ollama wrapper | Missing-host error, method forwarding, text/readiness normalization, unload request, and stream-option forwarding run against a deterministic fake `Arcane.ollama`. | Real managed-service, model download/create, GPU/resource admission, and service restart require an admitted Arcane host. |
@@ -5,7 +5,7 @@
5
5
  "repository": "https://github.com/TheWizardNexus/arcane-os-sdk.git",
6
6
  "branch": "main",
7
7
  "path": "runtime/arcane",
8
- "sdkVersion": "0.5.18",
8
+ "sdkVersion": "0.6.0",
9
9
  "protocol": "arcane/1"
10
10
  },
11
11
  "artifactCount": 85,
@@ -29,7 +29,7 @@
29
29
  "summary": "Owns provider-selectable chat and the one-time caller-authority browser STT/TTS configuration, lifecycle, synthesis, transcription, and playback boundary.",
30
30
  "availability": "Browser + native bridge + TWiN Cloud",
31
31
  "protocol": "arcane-ai-browser-speech-configuration/1, AIProviderRuntime arcane-ai-provider/2 routes, globalThis.arcaneEvents, TWiN Cloud HTTPS, Arcane.ollama, Arcane.speech, Android WebView bridge",
32
- "normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Kokoro TTS execution with default capacity 4, explicit provider-neutral execution snapshots through status(role,{execution:true}), sticky readiness, canonical window.user readiness without compatibility-event object-identity admission or polling, exact ordered structural calls with required arguments.message, atomic nonblank all-ID tool-result sequencing, complete all-choice streaming/data validation, native Ollama structural adaptation, complete-response then all validated tool callbacks then completion ordering, observed async callbacks, Blob/File transcription, automatic repeated-formatting-mark removal from cloned outbound TTS input while caller and visible content stays exact, immediate exact-segment synthesis, per-call voice/speed capture, optional final-segment pause on the existing audio clock, optional per-call terminal playback completion, detached preparation with complete semantic DBOPFS audio reuse and same-owner pending request sharing, independent prepared playback without a model load for cached audio, ordered audio-clock scheduling, playable audio, active-generation TTS operation-failure routing, mute, and cancellation are normalized; provider/model/runtime/voice policy remains caller-owned.",
32
+ "normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Whisper STT and Kokoro TTS execution with default capacities 1 and 4 respectively, explicit provider-neutral execution snapshots through status(role,{execution:true}), sticky readiness, canonical window.user readiness without compatibility-event object-identity admission or polling, exact ordered structural calls with required arguments.message, atomic nonblank all-ID tool-result sequencing, complete all-choice streaming/data validation, native Ollama structural adaptation, complete-response then all validated tool callbacks then completion ordering, observed async callbacks, Blob/File transcription, automatic repeated-formatting-mark removal from cloned outbound TTS input while caller and visible content stays exact, immediate exact-segment synthesis, per-call voice/speed capture, optional final-segment pause on the existing audio clock, optional per-call terminal playback completion, detached preparation with complete semantic DBOPFS audio reuse and same-owner pending request sharing, independent prepared playback without a model load for cached audio, ordered audio-clock scheduling, playable audio, active-generation TTS operation-failure routing, mute, and cancellation are normalized; provider/model/runtime/voice policy remains caller-owned. Both speech roles default execution.device to auto: webnn-npu when navigator.ml.createContext is exposed, then webgpu when navigator.gpu is exposed, then wasm for CPU execution. Failed backend loads clean up the pool before trying fresh Workers with the same prepared model and dtype; explicit devices never fall back. STT capacity only accepts 1; TTS accepts 1 through 4. Execution snapshots expose requestedDevice, selectedDevice, maxConcurrentRequests, and activeRequestCount for both roles; selectedDevice is null while unloaded and names the successful upstream session backend request when loaded. It does not prove every operation used the physical NPU: WebNN unsupported operations may use WASM, and compatibility depends on the browser, driver, hardware, and graph.",
33
33
  "surface": "Browser-speech protocol/event/error/reason constants; default `AI`; read-only `providerRuntime`, `browserSpeechConfiguration`, and `browserSpeechDescriptor`; explicit `providerRuntime.status(role,{execution:true})` snapshots; `streamTTS(text,end,options={})` with optional voice, speed, pauseAfterMs, and waitForPlayback; automatic speech-input cleanup with SDK-internal preparation metadata; `prepareTTS({parts,storage,identity,signal,onState})` returning ordered segments, state, ready, getAudio(index), and cancel(); `playPreparedTTS(prepared,{signal,onState})` returning state, error, finished, pause(), resume(), and stop() with one playback lane per AI and observational state/error callbacks; configure/dispose speech, route lifecycle, declaration-validated chat/stream requests with exact ordered calls, synthesis/transcription, and playback controls; initializes from current canonical user readiness, installs `window.ai`, and projects `ai-ready` plus active-generation `ai-tts-failure`."
34
34
  },
35
35
  {
@@ -253,10 +253,15 @@ user activation intent exposed by the shared speech component.
253
253
  `setSpeechMuted(false)` records the public unmuted state only after the selected
254
254
  TTS route reaches ready; a failed load leaves the public state muted. In contrast,
255
255
  `setSpeechMuted(true)` cancels active TTS work and unloads that role.
256
- The optional browser-speech `tts.execution` record selects
257
- `device:'auto'|'webgpu'|'wasm'` and a `maxConcurrentRequests` integer from 1
258
- through 4. Omission uses GPU-first automatic selection with four bounded Kokoro
259
- Worker/session slots; STT remains one WASM Worker.
256
+ The optional browser-speech `stt.execution` and `tts.execution` records select
257
+ `device:'auto'|'webnn-npu'|'webgpu'|'wasm'` and `maxConcurrentRequests`.
258
+ Omission uses NPU, GPU, then CPU automatic selection with one Whisper slot and
259
+ four bounded Kokoro Worker/session slots. Whisper accepts only capacity 1;
260
+ Kokoro accepts integers 1 through 4. Automatic loading attempts WebNN NPU when
261
+ `navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
262
+ exposed, then WASM. A failed candidate is cleaned up before fresh Workers try
263
+ the next device with the same prepared model and dtype. Explicit device
264
+ selections report failure without falling back.
260
265
  Capacity 4 means up to four segments synthesize at once. Segment 5 and later
261
266
  wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out
262
267
  of order, but playback waits for earlier segments and plays exact input order.
@@ -265,11 +270,15 @@ latency. This capacity does not establish physical GPU kernel overlap.
265
270
 
266
271
  After configuration, explicitly inspect execution through
267
272
  `ai.providerRuntime.status('tts', {execution:true}).execution`. When supplied
268
- by the selected provider, this read returns its execution snapshot. Kokoro
269
- reports `requestedDevice`, `selectedDevice`, `maxConcurrentRequests`, and
273
+ by the selected provider, this read returns its execution snapshot. Whisper and
274
+ Kokoro report `requestedDevice`, `selectedDevice`, `maxConcurrentRequests`, and
270
275
  `activeRequestCount`. `selectedDevice` is `null` before load and after unload.
271
276
  `requestedDevice === 'auto' && selectedDevice === 'wasm'` identifies automatic
272
- WASM fallback after a successful load. Calling `status()` without options keeps
277
+ WASM fallback after a successful load. Read the same fields for Whisper with
278
+ `status('stt', {execution:true})`. The selected device is the backend requested
279
+ by a successful upstream session load, not proof that every operation ran on a
280
+ physical accelerator; WebNN may execute unsupported operations through WASM.
281
+ Calling `status()` without options keeps
273
282
  the existing sticky lifecycle snapshot and does not inspect provider execution.
274
283
  Provider inspection failures are surfaced to the caller.
275
284
  `fetchTTS({model,voice,input,responseFormat,speed},signal,preparation={})` accepts the public
@@ -501,9 +510,11 @@ await ai.disposeBrowserSpeech({signal});
501
510
 
502
511
  The record is a mutable plain data record with exactly
503
512
  `{protocol,id,dbopfs,tableName?,stt?,tts?}` and at least one role. Each supplied
504
- mutable STT role is exactly `{providerId,graph,security?,offline}` or
505
- `{providerId,model,runtime,security?,offline}`. TTS accepts the corresponding
506
- shape plus optional `execution:{device,maxConcurrentRequests}`. The graph and
513
+ mutable STT or TTS role is exactly
514
+ `{providerId,graph,security?,offline,execution?}` or
515
+ `{providerId,model,runtime,security?,offline,execution?}`. The optional
516
+ `execution` record contains `{device,maxConcurrentRequests}` with the
517
+ role-specific defaults and capacities described above. The graph and
507
518
  direct authority forms are mutually exclusive; `providerId` and `id` are nonblank exact strings,
508
519
  `graph` is the role-matching graph returned by the SDK browser
509
520
  speech artifact API, and `offline` is boolean. The direct form forwards its
@@ -526,8 +537,9 @@ or reproduce DBOPFS cache logic.
526
537
 
527
538
  The returned descriptor is exactly `{protocol,configurationId,stt,tts}`; an
528
539
  external, unmanaged role is `null`. A managed STT descriptor is
529
- `{role:'stt',providerId,modelId,artifactGraphId?,offline}`; TTS adds
530
- `defaultVoice` and the normalized `execution` record. `artifactGraphId` is present only for the graph form.
540
+ `{role:'stt',providerId,modelId,artifactGraphId?,offline,execution}`; TTS adds
541
+ `defaultVoice`. Both include the normalized `execution` record.
542
+ `artifactGraphId` is present only for the graph form.
531
543
  `browserSpeechConfiguration` returns the exact caller-owned record when no
532
544
  managed role is carried. After a partial replacement that carries another
533
545
  managed role, it returns a mutable merged record with the replacement call's
@@ -764,16 +776,19 @@ never loads or downloads a model. `load()` forwards provider progress into the
764
776
  sticky role record; `unload()` and `dispose()` abort owned work, await exposed
765
777
  settlement, and verify provider status before publishing terminal state.
766
778
 
767
- `status('tts', {execution:true})` explicitly reads the selected provider and
768
- adds its optional `execution` snapshot to a copy of the role record.
779
+ `status('stt', {execution:true})` or `status('tts', {execution:true})` explicitly
780
+ reads the selected provider and adds its optional `execution` snapshot to a
781
+ copy of the role record.
769
782
  `status(null, {execution:true})` provides the equivalent projection under
770
783
  `roles.llm`, `roles.stt`, and `roles.tts`. Providers that do not supply execution
771
784
  omit that field. No provider load or sticky-state event is triggered; default
772
785
  `status()` keeps its existing identity and behavior. A provider inspection
773
- error propagates. Kokoro's execution contains `requestedDevice`,
786
+ error propagates. Each Whisper or Kokoro execution report contains `requestedDevice`,
774
787
  `selectedDevice` (`null` while unloaded), `maxConcurrentRequests`, and
775
- `activeRequestCount`; these describe provider execution, not physical GPU
776
- kernel overlap.
788
+ `activeRequestCount`. The selected device names the backend requested by a
789
+ successful upstream session load. It does not prove that every operation ran
790
+ on a physical NPU or GPU, or that accelerator kernels overlap; WebNN may use
791
+ WASM for unsupported operations.
777
792
 
778
793
  `validateSpeechConfiguration(value)` returns one mutable two-role selection
779
794
  record without committing it, where `value` is the closed `{stt,tts}` record.
@@ -6437,8 +6437,10 @@ createBrowserWhisperProvider(options={})
6437
6437
 
6438
6438
  The recognized options are
6439
6439
  `{id='arcane-browser-whisper',localOnly=true,graph,model,runtime,appSecurity,
6440
- security,store,offline=false}`. `graph` is mutually exclusive with `model` and
6441
- `runtime`. The mutable result is
6440
+ security,store,offline=false,execution={device:'auto',maxConcurrentRequests:1}}`.
6441
+ `execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
6442
+ exactly 1. `graph` is mutually exclusive with `model` and `runtime`. The mutable
6443
+ result is
6442
6444
  `{protocol:'arcane-ai-provider/2',role:'stt',id,localOnly:true,
6443
6445
  maxConcurrentRequests:1,catalog,inspect,status,load,request,unload,dispose}`.
6444
6446
  The only request operation is
@@ -6448,8 +6450,22 @@ re-freeze the cloned record.
6448
6450
 
6449
6451
  `status()` returns
6450
6452
  `{role,providerId,modelId,state,lifecycleStatus,lifecycleReason,activeOperation,
6451
- loaded,busy,generation,errorCode,cache,warnings}` and includes `security` only
6452
- for an explicit secure intent.
6453
+ loaded,busy,generation,errorCode,cache,warnings,execution}` and includes
6454
+ `security` only for an explicit secure intent.
6455
+ `execution` reports `requestedDevice`, `selectedDevice`,
6456
+ `maxConcurrentRequests`, and `activeRequestCount`; `selectedDevice` is `null`
6457
+ before load and after unload. The high-level projection is
6458
+ `ai.providerRuntime.status('stt', {execution:true}).execution`.
6459
+
6460
+ Automatic loading tries `webnn-npu` when `navigator.ml.createContext` is
6461
+ exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
6462
+ A failed candidate is cleaned up before a fresh Worker tries the next backend
6463
+ with the same prepared model and dtype. Explicit selections do not fall back.
6464
+ `selectedDevice` names the backend requested by the successful upstream session
6465
+ load; it does not prove that every operation ran on a physical accelerator.
6466
+ WebNN may execute unsupported operations through WASM, and exact model,
6467
+ browser, driver, and hardware compatibility remains upstream.
6468
+
6453
6469
  States are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
6454
6470
  `disposed`. Compatible concurrent loads coalesce; concurrent requests fail as
6455
6471
  `ARCANE_AI_PROVIDER_BUSY`. Cancellation after the Worker request begins
@@ -6523,8 +6539,8 @@ createBrowserKokoroProvider(options={})
6523
6539
  The recognized options are
6524
6540
  `{id='arcane-browser-kokoro',localOnly=true,graph,model,runtime,appSecurity,
6525
6541
  security,store,offline=false,execution={device:'auto',maxConcurrentRequests:4}}`.
6526
- `execution.device` is `auto`, `webgpu`, or `wasm`; its capacity is an integer
6527
- from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
6542
+ `execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
6543
+ an integer from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
6528
6544
  mutable result is
6529
6545
  `{protocol:'arcane-ai-provider/2',role:'tts',id,localOnly:true,
6530
6546
  maxConcurrentRequests,catalog,inspect,status,load,request,unload,dispose}`. The
@@ -6543,10 +6559,11 @@ mono 16-bit PCM. Unsupported formats fail
6543
6559
  `ARCANE_AI_UNSUPPORTED_RESPONSE_FORMAT`; malformed adapter audio fails
6544
6560
  `ARCANE_AI_INVALID_PROVIDER_RESULT`. Unknown fields and accessors reject as malformed.
6545
6561
 
6546
- Automatic execution attempts a complete WebGPU pool when the browser exposes
6547
- WebGPU and falls back by replacing the complete candidate pool with WASM when
6548
- WebGPU model loading rejects. Explicit `webgpu` does not fall back. Every slot
6549
- loads the same caller-selected model and dtype in a distinct Worker so the
6562
+ Automatic execution tries a complete WebNN NPU pool when
6563
+ `navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
6564
+ exposed, then CPU through WASM. A failed candidate is cleaned up before a fresh
6565
+ pool tries the next backend. Explicit device selections do not fall back. Every
6566
+ slot loads the same caller-selected model and dtype in a distinct Worker so the
6550
6567
  selected adapter's per-isolate inference serialization does not serialize the
6551
6568
  pool. Direct provider `status().execution` reports `requestedDevice`,
6552
6569
  `selectedDevice`, `maxConcurrentRequests`, and `activeRequestCount`.
@@ -6555,7 +6572,11 @@ by `AI.configureBrowserSpeech()`, use
6555
6572
  `ai.providerRuntime.status('tts', {execution:true}).execution`. Requested
6556
6573
  `auto` with selected `wasm` means automatic fallback occurred. Inspection
6557
6574
  errors propagate; the default runtime `status()` remains a sticky lifecycle
6558
- read. These fields do not prove physical GPU kernel overlap.
6575
+ read. `selectedDevice` names the backend requested by the successful upstream
6576
+ session load. These fields do not prove that every operation ran on a physical
6577
+ NPU or GPU, or that accelerator kernels overlap. WebNN may execute unsupported
6578
+ operations through WASM; exact model, browser, driver, and hardware
6579
+ compatibility remains upstream.
6559
6580
 
6560
6581
  Capacity 4 means up to four segments synthesize at once. Segment 5 and later
6561
6582
  wait in the SDK's provider-neutral FIFO queue; they are not dropped. Synthesis
@@ -6,13 +6,15 @@ For the smallest first request, start with the [browser speech quick start](../.
6
6
 
7
7
  ## Speech defaults and inspection
8
8
 
9
- The demo deliberately omits `tts.execution` in `speechConfiguration(dbopfs)` so it consumes the SDK default `{device:'auto',maxConcurrentRequests:4}`. Change that application's TTS record to `execution:{device:'wasm',maxConcurrentRequests:1}` to use less memory, or choose `webgpu` explicitly. Allowed capacities are 1 through 4; Whisper and the LLM retain capacity one.
9
+ The demo omits both speech execution records in `speechConfiguration(dbopfs)` so it consumes the SDK defaults: `{device:'auto',maxConcurrentRequests:1}` for Whisper STT and `{device:'auto',maxConcurrentRequests:4}` for Kokoro TTS. Auto tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is exposed, then CPU through WASM. Failed loads are cleaned up before fresh Workers try the next device with the same prepared model and dtype. Explicit `webnn-npu`, `webgpu`, and `wasm` selections do not fall back.
10
10
 
11
- The maintained Kokoro selection uses `fp32` because [Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage). Automatic fallback keeps the same dtype on WASM. `selectedDevice` identifies the loaded route; it does not validate speech correctness or audio quality, so evaluate actual output for each browser and device combination your application supports.
11
+ Change the application's TTS record to `execution:{device:'wasm',maxConcurrentRequests:1}` to use fewer synthesis sessions. TTS accepts capacities 1 through 4; Whisper accepts only 1 and the LLM retains capacity one.
12
+
13
+ The maintained Kokoro selection uses `fp32` because [Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage). Automatic fallback keeps that dtype on every device. `selectedDevice` identifies the backend requested by the successful upstream session load. It does not prove every operation ran on a physical NPU or GPU; WebNN may use WASM for unsupported operations. Exact model, browser, driver, and hardware compatibility and actual speech quality require evaluation on that combination.
12
14
 
13
15
  Capacity 4 means up to four segments synthesize at once. Segment 5 and later wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out of order, but playback waits for earlier segments and plays exact input order. Each slot owns a Worker/model session, so raising capacity trades memory for latency.
14
16
 
15
- Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected; `webgpu` reports the upstream execution-provider selection and does not prove physical GPU kernel overlap or audio quality. Auto may fall back to WASM with the same selected model/dtype; explicit WebGPU reports an error when it cannot load.
17
+ Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected. The corresponding Whisper report is available through `ai.providerRuntime.status('stt',{execution:true}).execution`; reading either report does not load a model.
16
18
 
17
19
  Shared Speech still owns mute, stop, and voice controls. The example's status button inspects the public report without loading a model or reaching into private providers.
18
20
 
@@ -434,7 +434,7 @@ function speechConfiguration(dbopfs) {
434
434
  offline: false,
435
435
  },
436
436
  tts: {
437
- // Omit execution to use the SDK's auto device and four synthesis slots.
437
+ // Omit execution for NPU, GPU, then CPU selection and four synthesis slots.
438
438
  // Kokoro.js recommends fp32 for the WebGPU route attempted by auto.
439
439
  providerId: "wasm-ai-demo-browser-kokoro",
440
440
  model: {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "arcane-os",
3
- "version": "0.5.19",
3
+ "version": "0.6.0",
4
4
  "description": "Arcane OS JavaScript SDK, project-local CLI, browser runtime, and repository-portable application packager.",
5
5
  "type": "module",
6
6
  "main": "./src/index.mjs",
@@ -27,11 +27,8 @@ const DEFAULT_TTS_SEGMENTATION={
27
27
  punctuation:'sentence',
28
28
  wordCadence:null
29
29
  };
30
- const DEFAULT_BROWSER_TTS_EXECUTION={
31
- device:'auto',
32
- maxConcurrentRequests:4
33
- };
34
- const BROWSER_TTS_EXECUTION_DEVICES=new Set(['auto','webgpu','wasm']);
30
+ const DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE='auto';
31
+ const BROWSER_SPEECH_EXECUTION_DEVICES=new Set(['auto','webnn-npu','webgpu','wasm']);
35
32
  const MAX_BROWSER_TTS_CONCURRENT_REQUESTS=4;
36
33
  const TTS_PUNCTUATION_MODES=new Set(['sentence','any','none']);
37
34
  // A complete punctuation run is a boundary unless the whole run consists of
@@ -817,36 +814,45 @@ function browserSpeechIdentifier(value,label){
817
814
  return value;
818
815
  }
819
816
 
820
- function normalizeBrowserTTSExecution(value){
817
+ function normalizeBrowserSpeechExecution(role,value){
818
+ const label=`AI browser speech ${role}.execution`;
819
+ const maximumConcurrentRequests=role==='stt'
820
+ ?1
821
+ :MAX_BROWSER_TTS_CONCURRENT_REQUESTS;
821
822
  if(value===undefined){
822
- return completeValue({...DEFAULT_BROWSER_TTS_EXECUTION});
823
+ return completeValue({
824
+ device:DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE,
825
+ maxConcurrentRequests:maximumConcurrentRequests
826
+ });
823
827
  }
824
828
  const descriptors=closedRecord(
825
829
  value,
826
830
  ['device','maxConcurrentRequests'],
827
831
  [],
828
- 'AI browser speech tts.execution'
832
+ label
829
833
  );
830
834
  const device=descriptors.device
831
835
  ?descriptors.device.value
832
- :DEFAULT_BROWSER_TTS_EXECUTION.device;
836
+ :DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE;
833
837
  const maxConcurrentRequests=descriptors.maxConcurrentRequests
834
838
  ?descriptors.maxConcurrentRequests.value
835
- :DEFAULT_BROWSER_TTS_EXECUTION.maxConcurrentRequests;
836
- if(!is.string(device)||!BROWSER_TTS_EXECUTION_DEVICES.has(device)){
839
+ :maximumConcurrentRequests;
840
+ if(!is.string(device)||!BROWSER_SPEECH_EXECUTION_DEVICES.has(device)){
837
841
  throw aiBrowserSpeechError(
838
842
  AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
839
843
  AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
840
- 'AI browser speech tts.execution.device must be auto, webgpu, or wasm.'
844
+ `${label}.device must be auto, webnn-npu, webgpu, or wasm.`
841
845
  );
842
846
  }
843
847
  if(!is.safeInteger(maxConcurrentRequests)
844
848
  ||maxConcurrentRequests<1
845
- ||maxConcurrentRequests>MAX_BROWSER_TTS_CONCURRENT_REQUESTS){
849
+ ||maxConcurrentRequests>maximumConcurrentRequests){
846
850
  throw aiBrowserSpeechError(
847
851
  AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
848
852
  AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
849
- `AI browser speech tts.execution.maxConcurrentRequests must be an integer from 1 through ${MAX_BROWSER_TTS_CONCURRENT_REQUESTS}.`
853
+ role==='stt'
854
+ ?`${label}.maxConcurrentRequests must be 1.`
855
+ :`${label}.maxConcurrentRequests must be an integer from 1 through ${maximumConcurrentRequests}.`
850
856
  );
851
857
  }
852
858
  return completeValue({device,maxConcurrentRequests});
@@ -917,13 +923,6 @@ function normalizeBrowserSpeechRole(value,role){
917
923
  `${label}.offline must be a boolean.`
918
924
  );
919
925
  }
920
- if(role==='stt'&&descriptors.execution){
921
- throw aiBrowserSpeechError(
922
- AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
923
- AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
924
- 'AI browser speech execution policy is available only for TTS.'
925
- );
926
- }
927
926
  return completeValue({
928
927
  providerId,
929
928
  ...(hasGraph
@@ -937,11 +936,10 @@ function normalizeBrowserSpeechRole(value,role){
937
936
  ...(secure?{security:{secure:true}}:{})
938
937
  }),
939
938
  offline:descriptors.offline.value,
940
- ...(role==='tts'
941
- ?{execution:normalizeBrowserTTSExecution(
942
- descriptors.execution?.value
943
- )}
944
- :{})
939
+ execution:normalizeBrowserSpeechExecution(
940
+ role,
941
+ descriptors.execution?.value
942
+ )
945
943
  });
946
944
  }
947
945
 
@@ -2770,10 +2768,10 @@ class AI {
2770
2768
  ?{artifactGraphId:catalog.artifactGraphId}
2771
2769
  :{}),
2772
2770
  offline:configured.offline,
2771
+ execution:configured.execution,
2773
2772
  ...(role==='tts'
2774
2773
  ?{
2775
- defaultVoice:catalog.defaultVoice,
2776
- execution:configured.execution
2774
+ defaultVoice:catalog.defaultVoice
2777
2775
  }
2778
2776
  :{})
2779
2777
  });
@@ -3145,7 +3143,7 @@ class AI {
3145
3143
  }),
3146
3144
  store,
3147
3145
  offline:configured.offline,
3148
- ...(role==='tts'?{execution:configured.execution}:{})
3146
+ execution:configured.execution
3149
3147
  });
3150
3148
  }
3151
3149
  providers=completeValue({...candidateProviders});