arcane-os 0.5.18 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,6 +1,32 @@
1
1
  # Changelog
2
2
 
3
- ## Unreleased
3
+ ## 0.6.0
4
+
5
+ - Prefer WebNN NPU, then WebGPU, then CPU through WASM for browser Whisper
6
+ transcription and Kokoro synthesis. Skip absent accelerator APIs and replace
7
+ failed model-load Workers before trying the next backend with the same
8
+ application-selected model and precision.
9
+ - Add Whisper execution configuration and NPU selection for both speech roles.
10
+ Explicit `webnn-npu`, `webgpu`, and `wasm` choices report failure without
11
+ falling back. Whisper retains one slot; Kokoro retains four by default and
12
+ accepts capacities from one through four.
13
+ - Report requested and successfully selected backends for both roles through
14
+ provider execution status. Selection reports upstream session loading;
15
+ actual accelerator use and speech quality depend on the selected runtime,
16
+ model, browser, drivers, and hardware. Wllama remains on its WebGPU backend.
17
+
18
+ ## 0.5.19
19
+
20
+ - Report observed Wllama initialization stages and runtime activity through
21
+ direct provider loading and the shared AI lifecycle. Keep initialization
22
+ indeterminate when the runtime supplies no meaningful completion total.
23
+ - Show the current initialization stage, stage duration, and time since runtime
24
+ activity in shared chat instead of a completed file count during activation.
25
+ Preserve download progress, cancellation, and complete runtime logging.
26
+ - Remove an adjacent duplicate cancellation check while retaining cancellation
27
+ handling before initialization and after loading.
28
+
29
+ ## 0.5.18
4
30
 
5
31
  - Use `strong-type` predicates throughout SDK-owned toolchain, runtime,
6
32
  browser-provider, and component code while retaining existing defaults,
@@ -12,8 +38,6 @@
12
38
  - Remove `AIResponseLength` exports; applications own response verbosity.
13
39
  Parse URL-audit HTML with the native HTML parser.
14
40
 
15
- ## 0.5.18
16
-
17
41
  - Add `AI.prepareTTS()` for detached punctuation-segmented synthesis, optional
18
42
  DBOPFS audio storage, complete semantic-input reuse, shared pending work and
19
43
  independent preparation cancellation. Persist MIME metadata with each audio
package/README.md CHANGED
@@ -19,7 +19,7 @@ version-locked SDK runtime, while an integrated Arcane checkout uses its live
19
19
  `arcane/` runtime. Both profiles preserve the same app URLs, theme, packaging,
20
20
  event, cancellation, and browser run contracts.
21
21
 
22
- This checkout defines the `0.5.18` SDK contract. Applications pin one exact npm
22
+ This checkout defines the `0.6.0` SDK contract. Applications pin one exact npm
23
23
  version and lockfile; registry state is deliberately not baked into application
24
24
  artifacts.
25
25
 
@@ -35,7 +35,7 @@ Create one browser application, install its pinned SDK, and start its source
35
35
  server:
36
36
 
37
37
  ```bash
38
- npx arcane-os@0.5.18 new hello-speech --path ./hello-speech --target browser
38
+ npx arcane-os@0.6.0 new hello-speech --path ./hello-speech --target browser
39
39
  cd hello-speech
40
40
  npm install
41
41
  npm run dev
@@ -299,10 +299,16 @@ provider factories. The package contains the plain-JavaScript provider and
299
299
  Worker machinery, not speech runtimes, models, voices, or a CDN default. An app
300
300
  must supply each runtime/model selection explicitly. Speech roles
301
301
  load, cancel, unload, fail, and recover independently, so speech failure never
302
- silently falls back or prevents text chat. Kokoro defaults to a four-slot Worker
303
- and model-session pool: it selects WebGPU when the browser can load the complete
304
- pool and otherwise recreates that pool on WASM. Apps may select `webgpu` or
305
- `wasm` explicitly and may set the bounded TTS capacity from one through four.
302
+ switches to a cloud provider or prevents text chat. Both roles default to
303
+ NPU, GPU, then CPU automatic selection: WebNN NPU when its browser API is
304
+ exposed, then WebGPU, then WASM. Failed model loads release their Workers before
305
+ the next backend uses the same model and precision. Apps may select
306
+ `webnn-npu`, `webgpu`, or `wasm` explicitly without automatic fallback.
307
+ Whisper keeps one transcription slot; Kokoro defaults to four Worker/model
308
+ sessions and accepts capacities from one through four. Both roles expose
309
+ requested and selected devices through execution status. A selected backend
310
+ reports successful upstream loading, not physical accelerator use for every
311
+ operation; model and hardware compatibility remain upstream.
306
312
 
307
313
  Applications that need faster spoken-response onset can configure the shared
308
314
  TTS stream without taking over synthesis or playback:
@@ -357,7 +363,7 @@ uses the same controller for automatic memory extraction.
357
363
  Create a new repository-shaped Arcane application with the exact stable SDK:
358
364
 
359
365
  ```bash
360
- npx arcane-os@0.5.18 new my-app --path ./my-app --target portable --git
366
+ npx arcane-os@0.6.0 new my-app --path ./my-app --target portable --git
361
367
  cd my-app
362
368
  npm install
363
369
  npm run dev
@@ -367,7 +373,7 @@ To enroll an existing repository, install the exact SDK and initialize only
367
373
  missing Arcane files:
368
374
 
369
375
  ```bash
370
- npm install --save-dev --save-exact arcane-os@0.5.18
376
+ npm install --save-dev --save-exact arcane-os@0.6.0
371
377
  npm exec -- arcane init my-app --target portable
372
378
  ```
373
379
 
@@ -383,7 +389,7 @@ npm exec -- arcane-os targets
383
389
  No global SDK install or standalone Arcane CLI is required. The application
384
390
  repository's exact npm dependency and lockfile own the CLI and toolchain version.
385
391
 
386
- Use `npx arcane-os@0.5.18` for the initial bootstrap because it names this npm
392
+ Use `npx arcane-os@0.6.0` for the initial bootstrap because it names this npm
387
393
  package explicitly; bare `npx arcane` outside an installed project could resolve
388
394
  a different package. Both installed commands invoke the same headless toolchain.
389
395
  Project-local npm scripts use the SDK pinned by that app's `package-lock.json`,
@@ -403,7 +409,7 @@ node ./bin/arcane.mjs new local-app --path ../local-app --target portable --git
403
409
 
404
410
  # From the generated app repository
405
411
  cd ../local-app
406
- npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.5.18.tgz
412
+ npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.6.0.tgz
407
413
  npm ci
408
414
  ```
409
415
 
@@ -412,7 +418,7 @@ same location. The lockfile retains the selected package dependency while
412
418
  Arcane uses the installed package name and version. Local directory `file:` dependencies are not
413
419
  accepted because npm may install them as links; use a packed `.tgz`. A GitHub
414
420
  runner also needs that tarball at the locked path. After publication, replace
415
- the local declaration with the exact `arcane-os@0.5.18` registry package and
421
+ the local declaration with the exact `arcane-os@0.6.0` registry package and
416
422
  commit the regenerated lock.
417
423
 
418
424
  Generated repositories use `npm ci --ignore-scripts` in CI. Run dependency
@@ -551,7 +557,7 @@ package installation, or assertions.
551
557
 
552
558
  ## Current target support
553
559
 
554
- Version `0.5.18` exposes one browser target and five explicitly paired
560
+ Version `0.6.0` exposes one browser target and five explicitly paired
555
561
  native development targets: a non-runnable portable directory, a
556
562
  Windows x64 unsigned-local-test EXE bundle, Linux x64 and Linux ARM64
557
563
  unsigned-local-test DEBs, and an Android development-signed APK. The
@@ -23,8 +23,8 @@ const ROLE_OPERATION = completeValue({ stt: "transcribe", tts: "synthesize" });
23
23
  const STT_SAMPLE_RATE = 16_000;
24
24
  const TTS_SAMPLE_RATE = 24_000;
25
25
  const TTS_RESPONSE_FORMAT = "wav";
26
- const TTS_EXECUTION_DEVICES = new Set(["auto", "webgpu", "wasm"]);
27
- const DEFAULT_TTS_EXECUTION_DEVICE = "auto";
26
+ const SPEECH_EXECUTION_DEVICES = new Set(["auto", "webnn-npu", "webgpu", "wasm"]);
27
+ const DEFAULT_SPEECH_EXECUTION_DEVICE = "auto";
28
28
  const DEFAULT_TTS_MAX_CONCURRENT_REQUESTS = 4;
29
29
  const MAX_TTS_CONCURRENT_REQUESTS = 4;
30
30
  const ROLE_REQUEST_REASON = completeValue({
@@ -150,24 +150,21 @@ function requiredIdentifier(value, label) {
150
150
  }
151
151
 
152
152
  function normalizeSpeechExecution(role, execution) {
153
- if (role === "stt") {
154
- if (execution !== undefined) {
155
- throw new TypeError("Browser Whisper does not accept an execution option.");
156
- }
157
- return completeValue({ device: "wasm", maxConcurrentRequests: 1 });
158
- }
153
+ const label = role === "stt" ? "Browser Whisper" : "Browser Kokoro";
154
+ const defaultConcurrency = role === "stt" ? 1 : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
155
+ const maximumConcurrency = role === "stt" ? 1 : MAX_TTS_CONCURRENT_REQUESTS;
159
156
  if (execution === undefined) {
160
157
  return completeValue({
161
- device: DEFAULT_TTS_EXECUTION_DEVICE,
162
- maxConcurrentRequests: DEFAULT_TTS_MAX_CONCURRENT_REQUESTS,
158
+ device: DEFAULT_SPEECH_EXECUTION_DEVICE,
159
+ maxConcurrentRequests: defaultConcurrency,
163
160
  });
164
161
  }
165
162
  if (!execution || !is.object(execution) || is.array(execution)) {
166
- throw new TypeError("Browser Kokoro execution must be a plain data record.");
163
+ throw new TypeError(`${label} execution must be a plain data record.`);
167
164
  }
168
165
  const prototype = Object.getPrototypeOf(execution);
169
166
  if (prototype !== Object.prototype && prototype !== null) {
170
- throw new TypeError("Browser Kokoro execution must be a plain data record.");
167
+ throw new TypeError(`${label} execution must be a plain data record.`);
171
168
  }
172
169
  const descriptors = Object.getOwnPropertyDescriptors(execution);
173
170
  for (const key of Reflect.ownKeys(descriptors)) {
@@ -175,34 +172,35 @@ function normalizeSpeechExecution(role, execution) {
175
172
  (key !== "device" && key !== "maxConcurrentRequests")
176
173
  || !Object.hasOwn(descriptors[key], "value")
177
174
  ) {
178
- throw new TypeError("Browser Kokoro execution contains an unsupported or accessor field.");
175
+ throw new TypeError(`${label} execution contains an unsupported or accessor field.`);
179
176
  }
180
177
  }
181
178
  const device = Object.hasOwn(descriptors, "device")
182
179
  ? descriptors.device.value
183
- : DEFAULT_TTS_EXECUTION_DEVICE;
180
+ : DEFAULT_SPEECH_EXECUTION_DEVICE;
184
181
  const maxConcurrentRequests = Object.hasOwn(descriptors, "maxConcurrentRequests")
185
182
  ? descriptors.maxConcurrentRequests.value
186
- : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
187
- if (!TTS_EXECUTION_DEVICES.has(device)) {
188
- throw new TypeError('Browser Kokoro execution.device must be "auto", "webgpu", or "wasm".');
183
+ : defaultConcurrency;
184
+ if (!SPEECH_EXECUTION_DEVICES.has(device)) {
185
+ throw new TypeError(`${label} execution.device must be "auto", "webnn-npu", "webgpu", or "wasm".`);
189
186
  }
190
187
  if (
191
188
  !is.safeInteger(maxConcurrentRequests)
192
189
  || maxConcurrentRequests < 1
193
- || maxConcurrentRequests > MAX_TTS_CONCURRENT_REQUESTS
190
+ || maxConcurrentRequests > maximumConcurrency
194
191
  ) {
195
- throw new TypeError("Browser Kokoro execution.maxConcurrentRequests must be a safe integer from 1 through 4.");
192
+ throw new TypeError(`${label} execution.maxConcurrentRequests must be a safe integer from 1 through ${maximumConcurrency}.`);
196
193
  }
197
194
  return completeValue({ device, maxConcurrentRequests });
198
195
  }
199
196
 
200
- function navigatorHasWebGpu() {
201
- try {
202
- return Boolean(globalThis.navigator?.gpu);
203
- } catch {
204
- return false;
205
- }
197
+ function speechExecutionDevices(requestedDevice) {
198
+ if (requestedDevice !== "auto") return [requestedDevice];
199
+ const devices = [];
200
+ if (is.function(globalThis.navigator?.ml?.createContext)) devices.push("webnn-npu");
201
+ if (globalThis.navigator?.gpu) devices.push("webgpu");
202
+ devices.push("wasm");
203
+ return devices;
206
204
  }
207
205
 
208
206
  function createProviderAuthority({
@@ -1187,14 +1185,12 @@ function createBrowserSpeechProvider({
1187
1185
  cache,
1188
1186
  lifecycleReason,
1189
1187
  activeOperation,
1190
- execution: role === "tts"
1191
- ? completeValue({
1192
- requestedDevice: speechExecution.device,
1193
- selectedDevice,
1194
- maxConcurrentRequests: speechExecution.maxConcurrentRequests,
1195
- activeRequestCount: requestOperations.size,
1196
- })
1197
- : null,
1188
+ execution: completeValue({
1189
+ requestedDevice: speechExecution.device,
1190
+ selectedDevice,
1191
+ maxConcurrentRequests: speechExecution.maxConcurrentRequests,
1192
+ activeRequestCount: requestOperations.size,
1193
+ }),
1198
1194
  secureIntent,
1199
1195
  warnings: providerWarnings(runtimeWarnings),
1200
1196
  });
@@ -1535,34 +1531,42 @@ function createBrowserSpeechProvider({
1535
1531
  `${role}-load-superseded-by-unload`,
1536
1532
  );
1537
1533
  }
1538
- const primaryDevice = role === "stt"
1539
- ? "wasm"
1540
- : speechExecution.device === "auto"
1541
- ? navigatorHasWebGpu() ? "webgpu" : "wasm"
1542
- : speechExecution.device;
1543
- try {
1544
- pool = await loadWorkerPool(
1545
- preparation,
1546
- primaryDevice,
1547
- record.warnings,
1548
- linked.controller.signal,
1549
- );
1550
- } catch (error) {
1551
- if (
1552
- role !== "tts"
1553
- || speechExecution.device !== "auto"
1554
- || primaryDevice !== "webgpu"
1555
- || linked.controller.signal.aborted
1556
- || operationGeneration !== generation
1557
- ) {
1558
- throw error;
1534
+ const devices = speechExecutionDevices(speechExecution.device);
1535
+ for (const [index, device] of devices.entries()) {
1536
+ throwIfAborted(linked.controller.signal, `${role}-load-cancelled`);
1537
+ if (operationGeneration !== generation) {
1538
+ throw providerError(
1539
+ "ARCANE_AI_OPERATION_SUPERSEDED",
1540
+ "Browser speech loading was superseded.",
1541
+ undefined,
1542
+ `${role}-load-superseded-by-unload`,
1543
+ );
1544
+ }
1545
+ try {
1546
+ pool = await loadWorkerPool(
1547
+ preparation,
1548
+ device,
1549
+ record.warnings,
1550
+ linked.controller.signal,
1551
+ );
1552
+ break;
1553
+ } catch (error) {
1554
+ if (
1555
+ index === devices.length - 1
1556
+ || linked.controller.signal.aborted
1557
+ || operationGeneration !== generation
1558
+ || error?.code === "ARCANE_AI_REQUEST_ABORTED"
1559
+ || error?.code === "ARCANE_AI_OPERATION_SUPERSEDED"
1560
+ ) {
1561
+ throw error;
1562
+ }
1563
+ // A fresh Worker releases a failed upstream initialization before
1564
+ // the next backend reuses the same prepared model and precision.
1565
+ const warning = `Browser ${role} ${device} loading failed; trying ${devices[index + 1]}.`;
1566
+ record.warnings = [...record.warnings, warning];
1567
+ lastWarnings = record.warnings;
1568
+ console.warn(warning, error);
1559
1569
  }
1560
- pool = await loadWorkerPool(
1561
- preparation,
1562
- "wasm",
1563
- record.warnings,
1564
- linked.controller.signal,
1565
- );
1566
1570
  }
1567
1571
  throwIfAborted(
1568
1572
  linked.controller.signal,
@@ -3000,24 +3000,47 @@ export function createBrowserWasmLlmProvider({
3000
3000
  ? context.reportProgress
3001
3001
  : options.onProgress ?? null;
3002
3002
  const progressStartedAt = Date.now();
3003
+ let progressStageStartedAt = progressStartedAt;
3004
+ let progressUpdatedAt = progressStartedAt;
3003
3005
  let currentProgress = null;
3004
3006
  let progressHeartbeat = null;
3005
3007
 
3006
3008
  function publishModelLoadProgress(progress) {
3007
- if (!reportProgress) return;
3009
+ if (!reportProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
3010
+ const now = Date.now();
3011
+ if (currentProgress?.phase !== progress.phase || currentProgress?.stage !== progress.stage) {
3012
+ progressStageStartedAt = now;
3013
+ }
3014
+ progressUpdatedAt = now;
3008
3015
  currentProgress = { ...progress, heartbeat: false };
3009
3016
  reportProgress(completeValue({
3010
3017
  ...currentProgress,
3011
- elapsedMs: Math.max(0, Date.now() - progressStartedAt),
3018
+ elapsedMs: Math.max(0, now - progressStartedAt),
3019
+ ...(progress.phase === "initialize" ? {
3020
+ phaseElapsedMs: Math.max(0, now - progressStageStartedAt),
3021
+ activityElapsedMs: 0,
3022
+ } : {}),
3012
3023
  }));
3013
3024
  }
3014
3025
 
3026
+ function publishModelInitializationProgress(progress) {
3027
+ if (!reportProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
3028
+ progressUpdatedAt = Date.now();
3029
+ if (!progress || (currentProgress?.stage === progress.stage && currentProgress?.message === progress.message)) return;
3030
+ publishModelLoadProgress(progress);
3031
+ }
3032
+
3015
3033
  function publishModelLoadHeartbeat() {
3016
- if (!reportProgress || !currentProgress) return;
3034
+ if (!reportProgress || !currentProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
3035
+ const now = Date.now();
3017
3036
  reportProgress(completeValue({
3018
3037
  ...currentProgress,
3019
3038
  heartbeat: true,
3020
- elapsedMs: Math.max(0, Date.now() - progressStartedAt),
3039
+ elapsedMs: Math.max(0, now - progressStartedAt),
3040
+ ...(currentProgress.phase === "initialize" ? {
3041
+ phaseElapsedMs: Math.max(0, now - progressStageStartedAt),
3042
+ activityElapsedMs: Math.max(0, now - progressUpdatedAt),
3043
+ } : {}),
3021
3044
  }));
3022
3045
  }
3023
3046
 
@@ -3047,16 +3070,12 @@ export function createBrowserWasmLlmProvider({
3047
3070
  if (generation !== lifecycleGeneration || state !== "loading") {
3048
3071
  throw fail("ARCANE_AI_OPERATION_SUPERSEDED", "The model load was superseded by unload.");
3049
3072
  }
3050
- throwIfAborted(signal, "load");
3051
- if (generation !== lifecycleGeneration || state !== "loading") {
3052
- throw fail("ARCANE_AI_OPERATION_SUPERSEDED", "The model load was superseded by unload.");
3053
- }
3054
3073
  const members = sourceMetadata(activeSource).files;
3055
3074
  publishModelLoadProgress({
3056
3075
  phase: "initialize",
3057
- completed: members.length,
3058
- total: members.length,
3059
- unit: "files",
3076
+ stage: "runtime",
3077
+ message: "Starting the WebAssembly runtime and opening model files",
3078
+ total: null,
3060
3079
  heartbeat: false,
3061
3080
  });
3062
3081
  const modelFiles = admitted.files.map((file, index) => (
@@ -3074,6 +3093,7 @@ export function createBrowserWasmLlmProvider({
3074
3093
  ...runtimeOptions,
3075
3094
  ...activeLoadPlan,
3076
3095
  signal,
3096
+ onProgress: publishModelInitializationProgress,
3077
3097
  });
3078
3098
  if (!runtime.isLoaded()) {
3079
3099
  throw fail(
@@ -94,6 +94,59 @@ function createEvidenceLogger(logger) {
94
94
  let offload = null;
95
95
  let invalid = false;
96
96
  let completionCapture = null;
97
+ let loadProgress = null;
98
+ let tensorCount = null;
99
+
100
+ function observeLoadLine(line) {
101
+ if (!loadProgress) return;
102
+ let stage;
103
+ let message;
104
+ const metadata = line.match(/^llama_model_loader: loaded meta data with \d+ key-value pairs and (\d+) tensors from /u);
105
+ const layers = line.match(GPU_OFFLOAD_PATTERN);
106
+ if (line.startsWith('Loading "wllama.wasm" from ')) {
107
+ stage = "runtime";
108
+ message = "Loading the WebAssembly runtime";
109
+ } else if (line === "Calling wllamaStart...") {
110
+ stage = "backend";
111
+ message = "Starting the inference engine";
112
+ } else if (line === "Loading model...") {
113
+ stage = "metadata";
114
+ message = "Reading model metadata";
115
+ } else if (WEBGPU_ADAPTER_PATTERN.test(line)) {
116
+ stage = "gpu";
117
+ message = "Graphics device initialized; preparing the model";
118
+ } else if (metadata) {
119
+ tensorCount = Number(metadata[1]);
120
+ stage = "metadata";
121
+ message = `Model metadata read: ${tensorCount} tensors`;
122
+ } else if (line.startsWith("load_tensors: loading model tensors,")) {
123
+ stage = "weights";
124
+ message = tensorCount === null
125
+ ? "Loading model weights"
126
+ : `Loading model weights for ${tensorCount} tensors`;
127
+ } else if (layers) {
128
+ // Upstream reports layer assignment before the weight reads finish.
129
+ stage = "weights";
130
+ message = `Loading model weights; ${layers[1]} of ${layers[2]} layers assigned to the GPU`;
131
+ } else if (line === "llama_context: constructing llama_context") {
132
+ stage = "context";
133
+ message = "Preparing the inference context";
134
+ } else if (/^[^:]+:\s+graph (?:nodes|splits)\s+=\s+\d+/u.test(line)) {
135
+ stage = "graph";
136
+ message = "Preparing the inference graph";
137
+ } else if (/^cmn\s+common_init_:\s+warming up the model with an empty run\b/u.test(line)) {
138
+ stage = "warmup";
139
+ message = "Warming up the model with an empty run";
140
+ }
141
+ if (stage) {
142
+ loadProgress({
143
+ phase: "initialize",
144
+ stage,
145
+ message,
146
+ total: null,
147
+ });
148
+ }
149
+ }
97
150
 
98
151
  function same(left, right) {
99
152
  return JSON.stringify(left) === JSON.stringify(right);
@@ -119,6 +172,7 @@ function createEvidenceLogger(logger) {
119
172
  observeCompletionLine(level, value);
120
173
  const line = String(value).trim();
121
174
  if (!line) return;
175
+ observeLoadLine(line);
122
176
  const adapterMatch = line.match(WEBGPU_ADAPTER_PATTERN);
123
177
  if (adapterMatch) {
124
178
  const next = completeValue({
@@ -144,6 +198,8 @@ function createEvidenceLogger(logger) {
144
198
  }
145
199
 
146
200
  function observe(level, args) {
201
+ // Activity without a new stage updates its age without repainting the UI.
202
+ loadProgress?.(null);
147
203
  for (const value of args) {
148
204
  if (!is.string(value)) continue;
149
205
  for (const line of value.split(/\r?\n/u)) observeLine(level, line);
@@ -160,6 +216,13 @@ function createEvidenceLogger(logger) {
160
216
 
161
217
  return completeValue({
162
218
  logger: completeValue(wrapped),
219
+ beginLoadProgress(report) {
220
+ loadProgress = report;
221
+ tensorCount = null;
222
+ return function releaseLoadProgress() {
223
+ loadProgress = null;
224
+ };
225
+ },
163
226
  beginCompletionCapture() {
164
227
  if (completionCapture) {
165
228
  throw runtimeFailure(
@@ -469,6 +532,18 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
469
532
  cleanup: null,
470
533
  });
471
534
  const loadController = new AbortController();
535
+ let progressFailure = null;
536
+ const releaseLoadProgress = sessionObservers.get(next).beginLoadProgress(
537
+ function reportRuntimeLoadProgress(progress) {
538
+ if (loadController.signal.aborted || !is.function(options.onProgress)) return;
539
+ try {
540
+ options.onProgress(progress);
541
+ } catch (error) {
542
+ progressFailure = error;
543
+ loadController.abort(error);
544
+ }
545
+ },
546
+ );
472
547
  const loadOperation = Promise.resolve().then(() => (
473
548
  next.arcaneLoadModel(files, loadOptions, loadController.signal)
474
549
  ));
@@ -484,6 +559,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
484
559
  else signal?.addEventListener?.("abort", onAbort, { once: true });
485
560
  try {
486
561
  await loadOperation;
562
+ if (progressFailure) throw progressFailure;
487
563
  if (pending?.engine !== next) throw new Error("Wllama load was cancelled.");
488
564
  if (!is.function(next.isModelLoaded) || next.isModelLoaded() !== true) {
489
565
  throw runtimeFailure(
@@ -496,6 +572,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
496
572
  engine = next;
497
573
  publishEvidence({ state: "ready", webgpu, cancellation: null, cleanup: null });
498
574
  } catch (error) {
575
+ releaseLoadProgress();
499
576
  let cleanupFailure = null;
500
577
  try {
501
578
  await exitSession(next);
@@ -513,6 +590,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
513
590
  if (cleanupFailure) throw cleanupFailure;
514
591
  throw error;
515
592
  } finally {
593
+ releaseLoadProgress();
516
594
  signal?.removeEventListener?.("abort", onAbort);
517
595
  }
518
596
 
@@ -738,10 +738,9 @@ function validateConfiguration(configuration, role) {
738
738
  if (
739
739
  descriptorMismatch
740
740
  || !Object.hasOwn(descriptors, "device")
741
- || (role === "stt" && descriptors.device.value !== "wasm")
742
- || (role === "tts"
743
- && descriptors.device.value !== "wasm"
744
- && descriptors.device.value !== "webgpu")
741
+ || (descriptors.device.value !== "wasm"
742
+ && descriptors.device.value !== "webgpu"
743
+ && descriptors.device.value !== "webnn-npu")
745
744
  ) {
746
745
  throw workerError(
747
746
  "ARCANE_AI_INVALID_REQUEST",
@@ -1367,7 +1366,7 @@ async function createWhisperEngine(namespace, configuration, signal, report) {
1367
1366
  "automatic-speech-recognition",
1368
1367
  configuration.model.repository,
1369
1368
  {
1370
- device: "wasm",
1369
+ device: configuration.execution?.device ?? "wasm",
1371
1370
  dtype: configuration.model.dtype ?? "fp32",
1372
1371
  revision: configuration.model.revision,
1373
1372
  progress_callback: report,
@@ -279,7 +279,7 @@ paths are withheld from the native provider. The provider copies the complete
279
279
  selected release rather than accepting an unrelated source path. Verification
280
280
  is a separate explicit operation for a selected release artifact.
281
281
 
282
- The SDK `0.5.17` runtime requires Arcane `0.8.12` or newer. Compatibility
282
+ The SDK `0.6.0` runtime requires Arcane `0.8.12` or newer. Compatibility
283
283
  is contractual rather than exact-version pinning: the prepared Core must meet
284
284
  the highest minimum declared by the runtime, selected app, and bundled app
285
285
  dependencies; keep each app's Arcane protocol generation; and provide every
@@ -16,7 +16,13 @@ dependency and invoke its local CLI with `npm exec -- arcane`. A separate
16
16
  global installer, standalone SDK executable, NuGet package, Homebrew formula,
17
17
  or OS package is not part of this release surface.
18
18
 
19
- Publication checks run only after the user explicitly selects an npm release.
19
+ The user's standing instruction selects publication when a coherent SDK change
20
+ is complete. Publish every ready change that preserves required functionality
21
+ and remains relevant, while excluding unfinished concurrent or explicitly
22
+ deferred work. Default to a patch release and assess whether a new capability
23
+ warrants a minor revision. Honor an explicit no-publish instruction.
24
+
25
+ Publication checks run only for that selected npm release output.
20
26
  That selected-release workflow validates package metadata, the executable and
21
27
  `.gitattributes` boundary, the complete package inventory, version/channel
22
28
  agreement, and required license notices. One unprivileged producer packs one
@@ -94,9 +100,9 @@ locked installation, runs one selected `arcane package`, creates one selected
94
100
  identities, receipts, provenance records, or attestation sidecars and does not
95
101
  run a second admission job.
96
102
 
97
- Stable versioning, the npm `latest` tag, and an official GitHub release remain a
98
- separate explicit release decision. Current `main` development does not
99
- silently convert a `-dev` package into an official release. A stable release
103
+ Stable versioning, the npm `latest` tag, and an official GitHub release follow
104
+ the selected release decision, including the standing completed-work authority
105
+ above. Unfinished `main` development does not select a release. A stable release
100
106
  must publish the exact selected Check artifact under `latest`; GitHub may then
101
107
  attach that same package. Its Git
102
108
  tag and GitHub release title must both be the same bare numeric