arcane-os 0.5.19 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +15 -0
- package/README.md +18 -12
- package/browser-runtime/ai/browser-speech-providers.mjs +64 -60
- package/browser-runtime/ai/speech-worker-runtime.mjs +4 -5
- package/docs/architecture.md +1 -1
- package/docs/reference/ai/browser-speech.md +65 -38
- package/docs/reference/behavioral-testing.md +1 -1
- package/docs/reference/inventory/runtime-modules.json +2 -2
- package/docs/reference/runtime-modules.md +32 -17
- package/docs/reference/sdk-api.md +32 -11
- package/examples/wasm-ai-demo/README.md +5 -3
- package/examples/wasm-ai-demo/app.js +1 -1
- package/package.json +1 -1
- package/runtime/arcane/modules/AI.js +27 -29
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,20 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 0.6.0
|
|
4
|
+
|
|
5
|
+
- Prefer WebNN NPU, then WebGPU, then CPU through WASM for browser Whisper
|
|
6
|
+
transcription and Kokoro synthesis. Skip absent accelerator APIs and replace
|
|
7
|
+
failed model-load Workers before trying the next backend with the same
|
|
8
|
+
application-selected model and precision.
|
|
9
|
+
- Add Whisper execution configuration and NPU selection for both speech roles.
|
|
10
|
+
Explicit `webnn-npu`, `webgpu`, and `wasm` choices report failure without
|
|
11
|
+
falling back. Whisper retains one slot; Kokoro retains four by default and
|
|
12
|
+
accepts capacities from one through four.
|
|
13
|
+
- Report requested and successfully selected backends for both roles through
|
|
14
|
+
provider execution status. Selection reports upstream session loading;
|
|
15
|
+
actual accelerator use and speech quality depend on the selected runtime,
|
|
16
|
+
model, browser, drivers, and hardware. Wllama remains on its WebGPU backend.
|
|
17
|
+
|
|
3
18
|
## 0.5.19
|
|
4
19
|
|
|
5
20
|
- Report observed Wllama initialization stages and runtime activity through
|
package/README.md
CHANGED
|
@@ -19,7 +19,7 @@ version-locked SDK runtime, while an integrated Arcane checkout uses its live
|
|
|
19
19
|
`arcane/` runtime. Both profiles preserve the same app URLs, theme, packaging,
|
|
20
20
|
event, cancellation, and browser run contracts.
|
|
21
21
|
|
|
22
|
-
This checkout defines the `0.
|
|
22
|
+
This checkout defines the `0.6.0` SDK contract. Applications pin one exact npm
|
|
23
23
|
version and lockfile; registry state is deliberately not baked into application
|
|
24
24
|
artifacts.
|
|
25
25
|
|
|
@@ -35,7 +35,7 @@ Create one browser application, install its pinned SDK, and start its source
|
|
|
35
35
|
server:
|
|
36
36
|
|
|
37
37
|
```bash
|
|
38
|
-
npx arcane-os@0.
|
|
38
|
+
npx arcane-os@0.6.0 new hello-speech --path ./hello-speech --target browser
|
|
39
39
|
cd hello-speech
|
|
40
40
|
npm install
|
|
41
41
|
npm run dev
|
|
@@ -299,10 +299,16 @@ provider factories. The package contains the plain-JavaScript provider and
|
|
|
299
299
|
Worker machinery, not speech runtimes, models, voices, or a CDN default. An app
|
|
300
300
|
must supply each runtime/model selection explicitly. Speech roles
|
|
301
301
|
load, cancel, unload, fail, and recover independently, so speech failure never
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
302
|
+
switches to a cloud provider or prevents text chat. Both roles default to
|
|
303
|
+
NPU, GPU, then CPU automatic selection: WebNN NPU when its browser API is
|
|
304
|
+
exposed, then WebGPU, then WASM. Failed model loads release their Workers before
|
|
305
|
+
the next backend uses the same model and precision. Apps may select
|
|
306
|
+
`webnn-npu`, `webgpu`, or `wasm` explicitly without automatic fallback.
|
|
307
|
+
Whisper keeps one transcription slot; Kokoro defaults to four Worker/model
|
|
308
|
+
sessions and accepts capacities from one through four. Both roles expose
|
|
309
|
+
requested and selected devices through execution status. A selected backend
|
|
310
|
+
reports successful upstream loading, not physical accelerator use for every
|
|
311
|
+
operation; model and hardware compatibility remain upstream.
|
|
306
312
|
|
|
307
313
|
Applications that need faster spoken-response onset can configure the shared
|
|
308
314
|
TTS stream without taking over synthesis or playback:
|
|
@@ -357,7 +363,7 @@ uses the same controller for automatic memory extraction.
|
|
|
357
363
|
Create a new repository-shaped Arcane application with the exact stable SDK:
|
|
358
364
|
|
|
359
365
|
```bash
|
|
360
|
-
npx arcane-os@0.
|
|
366
|
+
npx arcane-os@0.6.0 new my-app --path ./my-app --target portable --git
|
|
361
367
|
cd my-app
|
|
362
368
|
npm install
|
|
363
369
|
npm run dev
|
|
@@ -367,7 +373,7 @@ To enroll an existing repository, install the exact SDK and initialize only
|
|
|
367
373
|
missing Arcane files:
|
|
368
374
|
|
|
369
375
|
```bash
|
|
370
|
-
npm install --save-dev --save-exact arcane-os@0.
|
|
376
|
+
npm install --save-dev --save-exact arcane-os@0.6.0
|
|
371
377
|
npm exec -- arcane init my-app --target portable
|
|
372
378
|
```
|
|
373
379
|
|
|
@@ -383,7 +389,7 @@ npm exec -- arcane-os targets
|
|
|
383
389
|
No global SDK install or standalone Arcane CLI is required. The application
|
|
384
390
|
repository's exact npm dependency and lockfile own the CLI and toolchain version.
|
|
385
391
|
|
|
386
|
-
Use `npx arcane-os@0.
|
|
392
|
+
Use `npx arcane-os@0.6.0` for the initial bootstrap because it names this npm
|
|
387
393
|
package explicitly; bare `npx arcane` outside an installed project could resolve
|
|
388
394
|
a different package. Both installed commands invoke the same headless toolchain.
|
|
389
395
|
Project-local npm scripts use the SDK pinned by that app's `package-lock.json`,
|
|
@@ -403,7 +409,7 @@ node ./bin/arcane.mjs new local-app --path ../local-app --target portable --git
|
|
|
403
409
|
|
|
404
410
|
# From the generated app repository
|
|
405
411
|
cd ../local-app
|
|
406
|
-
npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.
|
|
412
|
+
npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.6.0.tgz
|
|
407
413
|
npm ci
|
|
408
414
|
```
|
|
409
415
|
|
|
@@ -412,7 +418,7 @@ same location. The lockfile retains the selected package dependency while
|
|
|
412
418
|
Arcane uses the installed package name and version. Local directory `file:` dependencies are not
|
|
413
419
|
accepted because npm may install them as links; use a packed `.tgz`. A GitHub
|
|
414
420
|
runner also needs that tarball at the locked path. After publication, replace
|
|
415
|
-
the local declaration with the exact `arcane-os@0.
|
|
421
|
+
the local declaration with the exact `arcane-os@0.6.0` registry package and
|
|
416
422
|
commit the regenerated lock.
|
|
417
423
|
|
|
418
424
|
Generated repositories use `npm ci --ignore-scripts` in CI. Run dependency
|
|
@@ -551,7 +557,7 @@ package installation, or assertions.
|
|
|
551
557
|
|
|
552
558
|
## Current target support
|
|
553
559
|
|
|
554
|
-
Version `0.
|
|
560
|
+
Version `0.6.0` exposes one browser target and five explicitly paired
|
|
555
561
|
native development targets: a non-runnable portable directory, a
|
|
556
562
|
Windows x64 unsigned-local-test EXE bundle, Linux x64 and Linux ARM64
|
|
557
563
|
unsigned-local-test DEBs, and an Android development-signed APK. The
|
|
@@ -23,8 +23,8 @@ const ROLE_OPERATION = completeValue({ stt: "transcribe", tts: "synthesize" });
|
|
|
23
23
|
const STT_SAMPLE_RATE = 16_000;
|
|
24
24
|
const TTS_SAMPLE_RATE = 24_000;
|
|
25
25
|
const TTS_RESPONSE_FORMAT = "wav";
|
|
26
|
-
const
|
|
27
|
-
const
|
|
26
|
+
const SPEECH_EXECUTION_DEVICES = new Set(["auto", "webnn-npu", "webgpu", "wasm"]);
|
|
27
|
+
const DEFAULT_SPEECH_EXECUTION_DEVICE = "auto";
|
|
28
28
|
const DEFAULT_TTS_MAX_CONCURRENT_REQUESTS = 4;
|
|
29
29
|
const MAX_TTS_CONCURRENT_REQUESTS = 4;
|
|
30
30
|
const ROLE_REQUEST_REASON = completeValue({
|
|
@@ -150,24 +150,21 @@ function requiredIdentifier(value, label) {
|
|
|
150
150
|
}
|
|
151
151
|
|
|
152
152
|
function normalizeSpeechExecution(role, execution) {
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
}
|
|
157
|
-
return completeValue({ device: "wasm", maxConcurrentRequests: 1 });
|
|
158
|
-
}
|
|
153
|
+
const label = role === "stt" ? "Browser Whisper" : "Browser Kokoro";
|
|
154
|
+
const defaultConcurrency = role === "stt" ? 1 : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
|
|
155
|
+
const maximumConcurrency = role === "stt" ? 1 : MAX_TTS_CONCURRENT_REQUESTS;
|
|
159
156
|
if (execution === undefined) {
|
|
160
157
|
return completeValue({
|
|
161
|
-
device:
|
|
162
|
-
maxConcurrentRequests:
|
|
158
|
+
device: DEFAULT_SPEECH_EXECUTION_DEVICE,
|
|
159
|
+
maxConcurrentRequests: defaultConcurrency,
|
|
163
160
|
});
|
|
164
161
|
}
|
|
165
162
|
if (!execution || !is.object(execution) || is.array(execution)) {
|
|
166
|
-
throw new TypeError(
|
|
163
|
+
throw new TypeError(`${label} execution must be a plain data record.`);
|
|
167
164
|
}
|
|
168
165
|
const prototype = Object.getPrototypeOf(execution);
|
|
169
166
|
if (prototype !== Object.prototype && prototype !== null) {
|
|
170
|
-
throw new TypeError(
|
|
167
|
+
throw new TypeError(`${label} execution must be a plain data record.`);
|
|
171
168
|
}
|
|
172
169
|
const descriptors = Object.getOwnPropertyDescriptors(execution);
|
|
173
170
|
for (const key of Reflect.ownKeys(descriptors)) {
|
|
@@ -175,34 +172,35 @@ function normalizeSpeechExecution(role, execution) {
|
|
|
175
172
|
(key !== "device" && key !== "maxConcurrentRequests")
|
|
176
173
|
|| !Object.hasOwn(descriptors[key], "value")
|
|
177
174
|
) {
|
|
178
|
-
throw new TypeError(
|
|
175
|
+
throw new TypeError(`${label} execution contains an unsupported or accessor field.`);
|
|
179
176
|
}
|
|
180
177
|
}
|
|
181
178
|
const device = Object.hasOwn(descriptors, "device")
|
|
182
179
|
? descriptors.device.value
|
|
183
|
-
:
|
|
180
|
+
: DEFAULT_SPEECH_EXECUTION_DEVICE;
|
|
184
181
|
const maxConcurrentRequests = Object.hasOwn(descriptors, "maxConcurrentRequests")
|
|
185
182
|
? descriptors.maxConcurrentRequests.value
|
|
186
|
-
:
|
|
187
|
-
if (!
|
|
188
|
-
throw new TypeError(
|
|
183
|
+
: defaultConcurrency;
|
|
184
|
+
if (!SPEECH_EXECUTION_DEVICES.has(device)) {
|
|
185
|
+
throw new TypeError(`${label} execution.device must be "auto", "webnn-npu", "webgpu", or "wasm".`);
|
|
189
186
|
}
|
|
190
187
|
if (
|
|
191
188
|
!is.safeInteger(maxConcurrentRequests)
|
|
192
189
|
|| maxConcurrentRequests < 1
|
|
193
|
-
|| maxConcurrentRequests >
|
|
190
|
+
|| maxConcurrentRequests > maximumConcurrency
|
|
194
191
|
) {
|
|
195
|
-
throw new TypeError(
|
|
192
|
+
throw new TypeError(`${label} execution.maxConcurrentRequests must be a safe integer from 1 through ${maximumConcurrency}.`);
|
|
196
193
|
}
|
|
197
194
|
return completeValue({ device, maxConcurrentRequests });
|
|
198
195
|
}
|
|
199
196
|
|
|
200
|
-
function
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
197
|
+
function speechExecutionDevices(requestedDevice) {
|
|
198
|
+
if (requestedDevice !== "auto") return [requestedDevice];
|
|
199
|
+
const devices = [];
|
|
200
|
+
if (is.function(globalThis.navigator?.ml?.createContext)) devices.push("webnn-npu");
|
|
201
|
+
if (globalThis.navigator?.gpu) devices.push("webgpu");
|
|
202
|
+
devices.push("wasm");
|
|
203
|
+
return devices;
|
|
206
204
|
}
|
|
207
205
|
|
|
208
206
|
function createProviderAuthority({
|
|
@@ -1187,14 +1185,12 @@ function createBrowserSpeechProvider({
|
|
|
1187
1185
|
cache,
|
|
1188
1186
|
lifecycleReason,
|
|
1189
1187
|
activeOperation,
|
|
1190
|
-
execution:
|
|
1191
|
-
|
|
1192
|
-
|
|
1193
|
-
|
|
1194
|
-
|
|
1195
|
-
|
|
1196
|
-
})
|
|
1197
|
-
: null,
|
|
1188
|
+
execution: completeValue({
|
|
1189
|
+
requestedDevice: speechExecution.device,
|
|
1190
|
+
selectedDevice,
|
|
1191
|
+
maxConcurrentRequests: speechExecution.maxConcurrentRequests,
|
|
1192
|
+
activeRequestCount: requestOperations.size,
|
|
1193
|
+
}),
|
|
1198
1194
|
secureIntent,
|
|
1199
1195
|
warnings: providerWarnings(runtimeWarnings),
|
|
1200
1196
|
});
|
|
@@ -1535,34 +1531,42 @@ function createBrowserSpeechProvider({
|
|
|
1535
1531
|
`${role}-load-superseded-by-unload`,
|
|
1536
1532
|
);
|
|
1537
1533
|
}
|
|
1538
|
-
const
|
|
1539
|
-
|
|
1540
|
-
|
|
1541
|
-
|
|
1542
|
-
|
|
1543
|
-
|
|
1544
|
-
|
|
1545
|
-
|
|
1546
|
-
|
|
1547
|
-
|
|
1548
|
-
|
|
1549
|
-
|
|
1550
|
-
|
|
1551
|
-
|
|
1552
|
-
|
|
1553
|
-
|
|
1554
|
-
|
|
1555
|
-
|
|
1556
|
-
|
|
1557
|
-
) {
|
|
1558
|
-
|
|
1534
|
+
const devices = speechExecutionDevices(speechExecution.device);
|
|
1535
|
+
for (const [index, device] of devices.entries()) {
|
|
1536
|
+
throwIfAborted(linked.controller.signal, `${role}-load-cancelled`);
|
|
1537
|
+
if (operationGeneration !== generation) {
|
|
1538
|
+
throw providerError(
|
|
1539
|
+
"ARCANE_AI_OPERATION_SUPERSEDED",
|
|
1540
|
+
"Browser speech loading was superseded.",
|
|
1541
|
+
undefined,
|
|
1542
|
+
`${role}-load-superseded-by-unload`,
|
|
1543
|
+
);
|
|
1544
|
+
}
|
|
1545
|
+
try {
|
|
1546
|
+
pool = await loadWorkerPool(
|
|
1547
|
+
preparation,
|
|
1548
|
+
device,
|
|
1549
|
+
record.warnings,
|
|
1550
|
+
linked.controller.signal,
|
|
1551
|
+
);
|
|
1552
|
+
break;
|
|
1553
|
+
} catch (error) {
|
|
1554
|
+
if (
|
|
1555
|
+
index === devices.length - 1
|
|
1556
|
+
|| linked.controller.signal.aborted
|
|
1557
|
+
|| operationGeneration !== generation
|
|
1558
|
+
|| error?.code === "ARCANE_AI_REQUEST_ABORTED"
|
|
1559
|
+
|| error?.code === "ARCANE_AI_OPERATION_SUPERSEDED"
|
|
1560
|
+
) {
|
|
1561
|
+
throw error;
|
|
1562
|
+
}
|
|
1563
|
+
// A fresh Worker releases a failed upstream initialization before
|
|
1564
|
+
// the next backend reuses the same prepared model and precision.
|
|
1565
|
+
const warning = `Browser ${role} ${device} loading failed; trying ${devices[index + 1]}.`;
|
|
1566
|
+
record.warnings = [...record.warnings, warning];
|
|
1567
|
+
lastWarnings = record.warnings;
|
|
1568
|
+
console.warn(warning, error);
|
|
1559
1569
|
}
|
|
1560
|
-
pool = await loadWorkerPool(
|
|
1561
|
-
preparation,
|
|
1562
|
-
"wasm",
|
|
1563
|
-
record.warnings,
|
|
1564
|
-
linked.controller.signal,
|
|
1565
|
-
);
|
|
1566
1570
|
}
|
|
1567
1571
|
throwIfAborted(
|
|
1568
1572
|
linked.controller.signal,
|
|
@@ -738,10 +738,9 @@ function validateConfiguration(configuration, role) {
|
|
|
738
738
|
if (
|
|
739
739
|
descriptorMismatch
|
|
740
740
|
|| !Object.hasOwn(descriptors, "device")
|
|
741
|
-
|| (
|
|
742
|
-
|
|
743
|
-
&& descriptors.device.value !== "
|
|
744
|
-
&& descriptors.device.value !== "webgpu")
|
|
741
|
+
|| (descriptors.device.value !== "wasm"
|
|
742
|
+
&& descriptors.device.value !== "webgpu"
|
|
743
|
+
&& descriptors.device.value !== "webnn-npu")
|
|
745
744
|
) {
|
|
746
745
|
throw workerError(
|
|
747
746
|
"ARCANE_AI_INVALID_REQUEST",
|
|
@@ -1367,7 +1366,7 @@ async function createWhisperEngine(namespace, configuration, signal, report) {
|
|
|
1367
1366
|
"automatic-speech-recognition",
|
|
1368
1367
|
configuration.model.repository,
|
|
1369
1368
|
{
|
|
1370
|
-
device: "wasm",
|
|
1369
|
+
device: configuration.execution?.device ?? "wasm",
|
|
1371
1370
|
dtype: configuration.model.dtype ?? "fp32",
|
|
1372
1371
|
revision: configuration.model.revision,
|
|
1373
1372
|
progress_callback: report,
|
package/docs/architecture.md
CHANGED
|
@@ -279,7 +279,7 @@ paths are withheld from the native provider. The provider copies the complete
|
|
|
279
279
|
selected release rather than accepting an unrelated source path. Verification
|
|
280
280
|
is a separate explicit operation for a selected release artifact.
|
|
281
281
|
|
|
282
|
-
The SDK `0.
|
|
282
|
+
The SDK `0.6.0` runtime requires Arcane `0.8.12` or newer. Compatibility
|
|
283
283
|
is contractual rather than exact-version pinning: the prepared Core must meet
|
|
284
284
|
the highest minimum declared by the runtime, selected app, and bundled app
|
|
285
285
|
dependencies; keep each app's Arcane protocol generation; and provide every
|
|
@@ -55,11 +55,11 @@ export const speechSelection = {
|
|
|
55
55
|
};
|
|
56
56
|
```
|
|
57
57
|
|
|
58
|
-
The omitted execution record below uses the SDK's
|
|
59
|
-
selection
|
|
58
|
+
The omitted execution record below uses the SDK's NPU, GPU, then CPU automatic
|
|
59
|
+
selection. This basic configuration uses `fp32` because
|
|
60
60
|
[Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage).
|
|
61
|
-
Automatic fallback carries this same selected model and dtype
|
|
62
|
-
does not rewrite the application selection. If you intentionally choose
|
|
61
|
+
Automatic fallback carries this same selected model and dtype between devices;
|
|
62
|
+
the SDK does not rewrite the application selection. If you intentionally choose
|
|
63
63
|
another dtype, evaluate that exact model, browser, and execution route.
|
|
64
64
|
`selectedDevice` reports routing after load, not pronunciation, text fidelity,
|
|
65
65
|
or audio quality.
|
|
@@ -229,9 +229,10 @@ playback. `replay()` keeps completed and pending provider segments and retries
|
|
|
229
229
|
only failed missing segments.
|
|
230
230
|
|
|
231
231
|
The default `{device:'auto',maxConcurrentRequests:4}` attempts the full ONNX
|
|
232
|
-
Worker/session pool on
|
|
233
|
-
|
|
234
|
-
|
|
232
|
+
Worker/session pool on WebNN NPU, then WebGPU, then CPU through WASM. It skips
|
|
233
|
+
an accelerator when its browser API is absent and replaces a failed candidate
|
|
234
|
+
with a fresh pool before trying the next device. The basic configuration above
|
|
235
|
+
keeps `fp32` throughout that sequence. Use the status example below to read
|
|
235
236
|
`selectedDevice`; console node-assignment warnings alone do not identify the
|
|
236
237
|
selected execution device or assess the generated audio.
|
|
237
238
|
|
|
@@ -588,8 +589,20 @@ console.log(speechText.append('ing', true)); // ing
|
|
|
588
589
|
|
|
589
590
|
## Choose a device or reduce memory use
|
|
590
591
|
|
|
591
|
-
|
|
592
|
-
|
|
592
|
+
Both speech roles accept `execution:{device,maxConcurrentRequests}`. Omitting
|
|
593
|
+
`stt.execution` selects `{device:'auto',maxConcurrentRequests:1}`; omitting
|
|
594
|
+
`tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`. Whisper keeps
|
|
595
|
+
one transcription slot. Kokoro accepts capacities 1 through 4.
|
|
596
|
+
|
|
597
|
+
Automatic selection tries `webnn-npu` when `navigator.ml.createContext` is
|
|
598
|
+
exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
|
|
599
|
+
API exposure only determines which upstream backend to attempt; it
|
|
600
|
+
does not establish compatible hardware, operators, or model shapes. A failed
|
|
601
|
+
candidate is fully cleaned up before the next candidate uses fresh Workers
|
|
602
|
+
with the same prepared model and dtype. An explicit `webnn-npu`, `webgpu`, or
|
|
603
|
+
`wasm` selection attempts only that backend and reports its failure.
|
|
604
|
+
|
|
605
|
+
These are four alternative TTS configurations:
|
|
593
606
|
|
|
594
607
|
```javascript
|
|
595
608
|
async function selectSpeechExecution(execution) {
|
|
@@ -609,15 +622,22 @@ async function selectSpeechExecution(execution) {
|
|
|
609
622
|
|
|
610
623
|
// Choose and call one from your application settings action:
|
|
611
624
|
// await selectSpeechExecution({ device: 'auto' });
|
|
625
|
+
// await selectSpeechExecution({ device: 'webnn-npu' });
|
|
612
626
|
// await selectSpeechExecution({ device: 'webgpu' });
|
|
613
627
|
// await selectSpeechExecution({ device: 'wasm', maxConcurrentRequests: 1 });
|
|
614
628
|
```
|
|
615
629
|
|
|
616
|
-
The override accepts integers 1, 2, 3, or 4
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
620
|
-
|
|
630
|
+
The TTS capacity override accepts integers 1, 2, 3, or 4; an STT capacity
|
|
631
|
+
override accepts only 1. Apply the same device choices to the configured
|
|
632
|
+
`stt.execution` record. A configuration change leaves TTS muted; explicitly
|
|
633
|
+
load/unmute again.
|
|
634
|
+
|
|
635
|
+
The selected upstream versions expose WebNN NPU through
|
|
636
|
+
[Transformers.js device selection](https://github.com/huggingface/transformers.js/blob/4.2.0/packages/transformers/src/backends/onnx.js)
|
|
637
|
+
and [Kokoro.js device forwarding](https://github.com/hexgrad/kokoro/blob/664c76a704021239ba59c84dcbaa4d3dece01fe9/kokoro.js/src/kokoro.js).
|
|
638
|
+
WebNN compatibility depends on the model's shapes and operations, browser,
|
|
639
|
+
drivers, and hardware. Unsupported operations may run through WASM even after
|
|
640
|
+
an NPU session loads; see the [ONNX Runtime WebNN contract](https://onnxruntime.ai/docs/tutorials/web/ep-webnn.html).
|
|
621
641
|
|
|
622
642
|
## Inspect the requested and selected device
|
|
623
643
|
|
|
@@ -627,15 +647,15 @@ loading. This explicitly reads the selected provider's current report; ordinary
|
|
|
627
647
|
separate execution-state event subscription.
|
|
628
648
|
|
|
629
649
|
```javascript
|
|
630
|
-
function printSpeechStatus() {
|
|
631
|
-
const status = ai.providerRuntime.status(
|
|
650
|
+
function printSpeechStatus(role = 'tts') {
|
|
651
|
+
const status = ai.providerRuntime.status(role, { execution: true });
|
|
632
652
|
const execution = status.execution;
|
|
633
|
-
console.log('
|
|
653
|
+
console.log('Speech role and state:', role, status.state);
|
|
634
654
|
if (execution) {
|
|
635
655
|
console.log('Requested device:', execution.requestedDevice);
|
|
636
656
|
console.log('Selected device:', execution.selectedDevice);
|
|
637
657
|
console.log('Capacity:', execution.maxConcurrentRequests);
|
|
638
|
-
console.log('Active
|
|
658
|
+
console.log('Active requests:', execution.activeRequestCount);
|
|
639
659
|
console.log('Automatic WASM fallback:',
|
|
640
660
|
execution.requestedDevice === 'auto' && execution.selectedDevice === 'wasm');
|
|
641
661
|
}
|
|
@@ -645,13 +665,18 @@ function printSpeechStatus() {
|
|
|
645
665
|
Call `printSpeechStatus()` after the load in `sayHello()` or from your status
|
|
646
666
|
button. The same projection is at
|
|
647
667
|
`ai.providerRuntime.status(null, {execution:true}).roles.tts.execution`.
|
|
668
|
+
For configured Whisper, call `printSpeechStatus('stt')` or read the corresponding
|
|
669
|
+
`roles.stt.execution` projection.
|
|
648
670
|
`selectedDevice` is `null`
|
|
649
671
|
until a pool is selected and returns to `null` on unload. Providers without an
|
|
650
672
|
execution report omit `execution`; do not infer a device from `navigator.gpu`
|
|
651
673
|
or a configured preference alone. An explicit inspection can throw a provider
|
|
652
674
|
status error; handle it with the same `error.code` / `error.message` pattern.
|
|
653
|
-
|
|
654
|
-
|
|
675
|
+
`selectedDevice` names the backend requested by the successful upstream session
|
|
676
|
+
load. It does not prove that every operation ran on a physical NPU or GPU, or
|
|
677
|
+
establish transcription correctness, pronunciation, or audio quality. Evaluate
|
|
678
|
+
actual speech output for the model, dtype, browser, and device combinations your
|
|
679
|
+
application supports.
|
|
655
680
|
|
|
656
681
|
## Stop, mute, cancel, and release
|
|
657
682
|
|
|
@@ -991,11 +1016,12 @@ they do not silently convert into a rejection policy.
|
|
|
991
1016
|
The Worker applies only the selected runtime settings needed to run the chosen
|
|
992
1017
|
provider:
|
|
993
1018
|
|
|
994
|
-
- Kokoro forwards the pool's selected `webgpu
|
|
1019
|
+
- Kokoro forwards the pool's selected `webnn-npu`, `webgpu`, or `wasm` device to
|
|
995
1020
|
`KokoroTTS.from_pretrained()`. Its configured dtype remains exactly the
|
|
996
|
-
caller-selected dtype on
|
|
1021
|
+
caller-selected dtype on every path. The WASM path also uses
|
|
997
1022
|
`namespace.env.wasmPaths = {mjs,wasm}`.
|
|
998
|
-
- Transformers
|
|
1023
|
+
- Transformers forwards the selected device to its speech-recognition pipeline,
|
|
1024
|
+
preserves the caller-selected dtype, and uses
|
|
999
1025
|
`namespace.env.backends.onnx.wasm.wasmPaths = {mjs,wasm}`, keeps remote model
|
|
1000
1026
|
loading enabled, and applies caller-selected `numThreads` when present.
|
|
1001
1027
|
|
|
@@ -1029,15 +1055,15 @@ The constructors also accept ordinary `model` and `runtime` descriptors instead
|
|
|
1029
1055
|
of `graph`; the two forms are mutually exclusive. Both forms require an
|
|
1030
1056
|
SDK-created DBOPFS speech artifact store. `localOnly` remains `true`.
|
|
1031
1057
|
|
|
1032
|
-
|
|
1033
|
-
`{device,maxConcurrentRequests}`. `device` is `auto`, `webgpu`, or
|
|
1034
|
-
`
|
|
1035
|
-
`{device:'auto',maxConcurrentRequests:
|
|
1036
|
-
|
|
1037
|
-
|
|
1038
|
-
|
|
1039
|
-
|
|
1040
|
-
|
|
1058
|
+
Both constructors accept the exact `execution` record
|
|
1059
|
+
`{device,maxConcurrentRequests}`. `device` is `auto`, `webnn-npu`, `webgpu`, or
|
|
1060
|
+
`wasm`. Whisper permits capacity 1 and defaults to
|
|
1061
|
+
`{device:'auto',maxConcurrentRequests:1}`. Kokoro permits integer capacities 1
|
|
1062
|
+
through 4 and defaults to `{device:'auto',maxConcurrentRequests:4}`. `auto`
|
|
1063
|
+
tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when
|
|
1064
|
+
`navigator.gpu` is exposed, then WASM. Before advancing after a failed load,
|
|
1065
|
+
the SDK tears down the candidate and creates fresh Workers using the same
|
|
1066
|
+
prepared model and dtype. Explicit device selections never fall back.
|
|
1041
1067
|
|
|
1042
1068
|
Each constructor returns an `arcane-ai-provider/2` object with:
|
|
1043
1069
|
|
|
@@ -1123,12 +1149,13 @@ mutable values.
|
|
|
1123
1149
|
Provider states are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
|
|
1124
1150
|
`disposed`. `status()` includes role, provider/model ids, state, lifecycle
|
|
1125
1151
|
status and reason, active operation, loaded/busy flags, generation, error code,
|
|
1126
|
-
cache state, and warnings.
|
|
1152
|
+
cache state, and warnings. Both speech roles include an `execution` record
|
|
1127
1153
|
with requested and selected device, request limit, and active request count. A
|
|
1128
|
-
successful `selectedDevice
|
|
1129
|
-
|
|
1130
|
-
overlap
|
|
1131
|
-
security field is absent
|
|
1154
|
+
successful `selectedDevice` reports the backend requested by the upstream model
|
|
1155
|
+
load. It does not prove that every operation ran on a physical NPU or GPU,
|
|
1156
|
+
that accelerator kernels overlap, or that generated speech is correct. WebNN
|
|
1157
|
+
may execute unsupported operations through WASM. A security field is absent
|
|
1158
|
+
in ordinary mode.
|
|
1132
1159
|
|
|
1133
1160
|
The provider/2 load context accepts an optional progress callback for interface
|
|
1134
1161
|
compatibility, but the current browser-speech artifact and Worker transport
|
|
@@ -1203,7 +1230,7 @@ operation.
|
|
|
1203
1230
|
## Ownership
|
|
1204
1231
|
|
|
1205
1232
|
- Applications own model, runtime, dtype, sample-rate, voice, profile, prompt,
|
|
1206
|
-
catalog, activation, optional TTS execution override, and presentation
|
|
1233
|
+
catalog, activation, optional STT/TTS execution override, and presentation
|
|
1207
1234
|
policy.
|
|
1208
1235
|
- Upstream publishers own their runtime, model, voice, and license delivery.
|
|
1209
1236
|
- The SDK owns storage, materialization, routing, Worker lifecycle, normalized
|
|
@@ -41,7 +41,7 @@ development evidence; it is not native artifact or release acceptance.
|
|
|
41
41
|
| Browser runtime modules | Every shipped ESM module parses and its export inventory matches the catalog; pure helpers run focused success/error cases. | DOM, OPFS, media, and Web Component journeys use a browser harness. |
|
|
42
42
|
| Provider-neutral AI runtime and chat/speech activation | Provider/2 registration, three-role configuration, TWiN Cloud LLM readiness, on-device Whisper/Kokoro selection, opt-in STT startup, Core speech readiness, independent LLM/STT/TTS load/unload/status, capacity-1 FIFO settlement for LLM/STT, bounded parallel synthesis with FIFO admission for an explicitly capable TTS provider, owned STT signals, TTS mute lifecycle, route-owned voice defaults, immediate chunk synthesis admission, original-order audio-clock scheduling, sticky-state-only readiness for both speech components, selected-unloaded activation request/cancellation/error behavior, programmatic voice recording, transcript-replacement supersession of late settlement, direct `AI.fetchSTT` result delivery, rejection of non-local speech configuration, and absence of silent provider fallback are represented against complete providers and host callbacks. | Real model/runtime availability remains the selected provider's evidence boundary; provider-promise settlement, state, an abort signal, or an activation event does not by itself prove underlying provider work stopped or native, cloud, or browser-model availability. |
|
|
43
43
|
| Browser-WASM local AI | The exported namespace, canonical ordered `{id, files:[{name?,url},...]}` descriptor, nonempty provider `sources` catalog, default `secure:false`, dormant `secure:true` intent, public AI API module lifecycle, lazy/manual policy, successful Wllama-load requirement, abort normalization, complete output and reasoning, all-choice validation, required structural-call `arguments.message`, exact call identity, and matching tool-result sequencing are represented in deterministic fixture sources. | A real Chrome exercise may load the selected Wllama runtime and model only after explicit user action. It is not an ordinary publication gate or an implicit model download. |
|
|
44
|
-
| Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, GPU-
|
|
44
|
+
| Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, NPU-to-GPU-to-CPU automatic selection for both roles, fresh Workers after a failed device load, explicit-device failure without fallback, bounded Kokoro Worker/session pools, out-of-order synthesis settlement with FIFO admission, request-targeted TTS cancellation, destructive lifecycle cleanup, Blob/File STT conversion, WAV TTS conversion, complete text/audio, and no cloud fallback are represented with synthetic artifacts and adapters. | A real runtime/model/voice download, WebNN or WebGPU model load, physical accelerator behavior, and actual transcription or synthesis use the application's selected upstream packages/providers, browser media support, and explicit user action. A selected backend is not proof that every operation ran on that physical accelerator; WebNN may execute unsupported operations through WASM. |
|
|
45
45
|
| Persistent chat and document context | Atomic in-memory/history commit, explicit per-turn persistence, streamed/non-stream fallback, session-owned callback fields, complete data callbacks, per-turn request options, exact per-choice streamed/terminal call correlation before publication, terminal-only call acceptance, ordered parallel-call sequencing, atomic all-ID nonblank executed/declined/cancelled/not-executed result batches, readable unmodified malformed pre-existing rows, complete UI transcript metadata, generic visible failure outcomes with complete console diagnostics, BFCache-preserving component lifecycle, complete bootstrap/search/context, caller-source evaluation, cancellation, and partial-read handling are represented with app-scoped adapters. | Live Core/provider inference and durable browser storage remain separate authorities; tests never treat a fake chat function or in-memory adapter as host/storage proof. |
|
|
46
46
|
| Core bridge docs | Canonical namespace/method/event/entity inventories match their one-per-member guides and required sections. | Live Core conformance belongs in Arcane OS because Core implementation is not shipped as SDK source. |
|
|
47
47
|
| Arcane Ollama wrapper | Missing-host error, method forwarding, text/readiness normalization, unload request, and stream-option forwarding run against a deterministic fake `Arcane.ollama`. | Real managed-service, model download/create, GPU/resource admission, and service restart require an admitted Arcane host. |
|
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
"repository": "https://github.com/TheWizardNexus/arcane-os-sdk.git",
|
|
6
6
|
"branch": "main",
|
|
7
7
|
"path": "runtime/arcane",
|
|
8
|
-
"sdkVersion": "0.
|
|
8
|
+
"sdkVersion": "0.6.0",
|
|
9
9
|
"protocol": "arcane/1"
|
|
10
10
|
},
|
|
11
11
|
"artifactCount": 85,
|
|
@@ -29,7 +29,7 @@
|
|
|
29
29
|
"summary": "Owns provider-selectable chat and the one-time caller-authority browser STT/TTS configuration, lifecycle, synthesis, transcription, and playback boundary.",
|
|
30
30
|
"availability": "Browser + native bridge + TWiN Cloud",
|
|
31
31
|
"protocol": "arcane-ai-browser-speech-configuration/1, AIProviderRuntime arcane-ai-provider/2 routes, globalThis.arcaneEvents, TWiN Cloud HTTPS, Arcane.ollama, Arcane.speech, Android WebView bridge",
|
|
32
|
-
"normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Kokoro TTS execution with default
|
|
32
|
+
"normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Whisper STT and Kokoro TTS execution with default capacities 1 and 4 respectively, explicit provider-neutral execution snapshots through status(role,{execution:true}), sticky readiness, canonical window.user readiness without compatibility-event object-identity admission or polling, exact ordered structural calls with required arguments.message, atomic nonblank all-ID tool-result sequencing, complete all-choice streaming/data validation, native Ollama structural adaptation, complete-response then all validated tool callbacks then completion ordering, observed async callbacks, Blob/File transcription, automatic repeated-formatting-mark removal from cloned outbound TTS input while caller and visible content stays exact, immediate exact-segment synthesis, per-call voice/speed capture, optional final-segment pause on the existing audio clock, optional per-call terminal playback completion, detached preparation with complete semantic DBOPFS audio reuse and same-owner pending request sharing, independent prepared playback without a model load for cached audio, ordered audio-clock scheduling, playable audio, active-generation TTS operation-failure routing, mute, and cancellation are normalized; provider/model/runtime/voice policy remains caller-owned. Both speech roles default execution.device to auto: webnn-npu when navigator.ml.createContext is exposed, then webgpu when navigator.gpu is exposed, then wasm for CPU execution. Failed backend loads clean up the pool before trying fresh Workers with the same prepared model and dtype; explicit devices never fall back. STT capacity only accepts 1; TTS accepts 1 through 4. Execution snapshots expose requestedDevice, selectedDevice, maxConcurrentRequests, and activeRequestCount for both roles; selectedDevice is null while unloaded and names the successful upstream session backend request when loaded. It does not prove every operation used the physical NPU: WebNN unsupported operations may use WASM, and compatibility depends on the browser, driver, hardware, and graph.",
|
|
33
33
|
"surface": "Browser-speech protocol/event/error/reason constants; default `AI`; read-only `providerRuntime`, `browserSpeechConfiguration`, and `browserSpeechDescriptor`; explicit `providerRuntime.status(role,{execution:true})` snapshots; `streamTTS(text,end,options={})` with optional voice, speed, pauseAfterMs, and waitForPlayback; automatic speech-input cleanup with SDK-internal preparation metadata; `prepareTTS({parts,storage,identity,signal,onState})` returning ordered segments, state, ready, getAudio(index), and cancel(); `playPreparedTTS(prepared,{signal,onState})` returning state, error, finished, pause(), resume(), and stop() with one playback lane per AI and observational state/error callbacks; configure/dispose speech, route lifecycle, declaration-validated chat/stream requests with exact ordered calls, synthesis/transcription, and playback controls; initializes from current canonical user readiness, installs `window.ai`, and projects `ai-ready` plus active-generation `ai-tts-failure`."
|
|
34
34
|
},
|
|
35
35
|
{
|
|
@@ -253,10 +253,15 @@ user activation intent exposed by the shared speech component.
|
|
|
253
253
|
`setSpeechMuted(false)` records the public unmuted state only after the selected
|
|
254
254
|
TTS route reaches ready; a failed load leaves the public state muted. In contrast,
|
|
255
255
|
`setSpeechMuted(true)` cancels active TTS work and unloads that role.
|
|
256
|
-
The optional browser-speech `tts.execution`
|
|
257
|
-
`device:'auto'|'webgpu'|'wasm'` and
|
|
258
|
-
|
|
259
|
-
Worker/session slots
|
|
256
|
+
The optional browser-speech `stt.execution` and `tts.execution` records select
|
|
257
|
+
`device:'auto'|'webnn-npu'|'webgpu'|'wasm'` and `maxConcurrentRequests`.
|
|
258
|
+
Omission uses NPU, GPU, then CPU automatic selection with one Whisper slot and
|
|
259
|
+
four bounded Kokoro Worker/session slots. Whisper accepts only capacity 1;
|
|
260
|
+
Kokoro accepts integers 1 through 4. Automatic loading attempts WebNN NPU when
|
|
261
|
+
`navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
|
|
262
|
+
exposed, then WASM. A failed candidate is cleaned up before fresh Workers try
|
|
263
|
+
the next device with the same prepared model and dtype. Explicit device
|
|
264
|
+
selections report failure without falling back.
|
|
260
265
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later
|
|
261
266
|
wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out
|
|
262
267
|
of order, but playback waits for earlier segments and plays exact input order.
|
|
@@ -265,11 +270,15 @@ latency. This capacity does not establish physical GPU kernel overlap.
|
|
|
265
270
|
|
|
266
271
|
After configuration, explicitly inspect execution through
|
|
267
272
|
`ai.providerRuntime.status('tts', {execution:true}).execution`. When supplied
|
|
268
|
-
by the selected provider, this read returns its execution snapshot.
|
|
269
|
-
|
|
273
|
+
by the selected provider, this read returns its execution snapshot. Whisper and
|
|
274
|
+
Kokoro report `requestedDevice`, `selectedDevice`, `maxConcurrentRequests`, and
|
|
270
275
|
`activeRequestCount`. `selectedDevice` is `null` before load and after unload.
|
|
271
276
|
`requestedDevice === 'auto' && selectedDevice === 'wasm'` identifies automatic
|
|
272
|
-
WASM fallback after a successful load.
|
|
277
|
+
WASM fallback after a successful load. Read the same fields for Whisper with
|
|
278
|
+
`status('stt', {execution:true})`. The selected device is the backend requested
|
|
279
|
+
by a successful upstream session load, not proof that every operation ran on a
|
|
280
|
+
physical accelerator; WebNN may execute unsupported operations through WASM.
|
|
281
|
+
Calling `status()` without options keeps
|
|
273
282
|
the existing sticky lifecycle snapshot and does not inspect provider execution.
|
|
274
283
|
Provider inspection failures are surfaced to the caller.
|
|
275
284
|
`fetchTTS({model,voice,input,responseFormat,speed},signal,preparation={})` accepts the public
|
|
@@ -501,9 +510,11 @@ await ai.disposeBrowserSpeech({signal});
|
|
|
501
510
|
|
|
502
511
|
The record is a mutable plain data record with exactly
|
|
503
512
|
`{protocol,id,dbopfs,tableName?,stt?,tts?}` and at least one role. Each supplied
|
|
504
|
-
mutable STT role is exactly
|
|
505
|
-
`{providerId,
|
|
506
|
-
|
|
513
|
+
mutable STT or TTS role is exactly
|
|
514
|
+
`{providerId,graph,security?,offline,execution?}` or
|
|
515
|
+
`{providerId,model,runtime,security?,offline,execution?}`. The optional
|
|
516
|
+
`execution` record contains `{device,maxConcurrentRequests}` with the
|
|
517
|
+
role-specific defaults and capacities described above. The graph and
|
|
507
518
|
direct authority forms are mutually exclusive; `providerId` and `id` are nonblank exact strings,
|
|
508
519
|
`graph` is the role-matching graph returned by the SDK browser
|
|
509
520
|
speech artifact API, and `offline` is boolean. The direct form forwards its
|
|
@@ -526,8 +537,9 @@ or reproduce DBOPFS cache logic.
|
|
|
526
537
|
|
|
527
538
|
The returned descriptor is exactly `{protocol,configurationId,stt,tts}`; an
|
|
528
539
|
external, unmanaged role is `null`. A managed STT descriptor is
|
|
529
|
-
`{role:'stt',providerId,modelId,artifactGraphId?,offline}`; TTS adds
|
|
530
|
-
`defaultVoice
|
|
540
|
+
`{role:'stt',providerId,modelId,artifactGraphId?,offline,execution}`; TTS adds
|
|
541
|
+
`defaultVoice`. Both include the normalized `execution` record.
|
|
542
|
+
`artifactGraphId` is present only for the graph form.
|
|
531
543
|
`browserSpeechConfiguration` returns the exact caller-owned record when no
|
|
532
544
|
managed role is carried. After a partial replacement that carries another
|
|
533
545
|
managed role, it returns a mutable merged record with the replacement call's
|
|
@@ -764,16 +776,19 @@ never loads or downloads a model. `load()` forwards provider progress into the
|
|
|
764
776
|
sticky role record; `unload()` and `dispose()` abort owned work, await exposed
|
|
765
777
|
settlement, and verify provider status before publishing terminal state.
|
|
766
778
|
|
|
767
|
-
`status('
|
|
768
|
-
adds its optional `execution` snapshot to a
|
|
779
|
+
`status('stt', {execution:true})` or `status('tts', {execution:true})` explicitly
|
|
780
|
+
reads the selected provider and adds its optional `execution` snapshot to a
|
|
781
|
+
copy of the role record.
|
|
769
782
|
`status(null, {execution:true})` provides the equivalent projection under
|
|
770
783
|
`roles.llm`, `roles.stt`, and `roles.tts`. Providers that do not supply execution
|
|
771
784
|
omit that field. No provider load or sticky-state event is triggered; default
|
|
772
785
|
`status()` keeps its existing identity and behavior. A provider inspection
|
|
773
|
-
error propagates. Kokoro
|
|
786
|
+
error propagates. Each Whisper or Kokoro execution report contains `requestedDevice`,
|
|
774
787
|
`selectedDevice` (`null` while unloaded), `maxConcurrentRequests`, and
|
|
775
|
-
`activeRequestCount
|
|
776
|
-
|
|
788
|
+
`activeRequestCount`. The selected device names the backend requested by a
|
|
789
|
+
successful upstream session load. It does not prove that every operation ran
|
|
790
|
+
on a physical NPU or GPU, or that accelerator kernels overlap; WebNN may use
|
|
791
|
+
WASM for unsupported operations.
|
|
777
792
|
|
|
778
793
|
`validateSpeechConfiguration(value)` returns one mutable two-role selection
|
|
779
794
|
record without committing it, where `value` is the closed `{stt,tts}` record.
|
|
@@ -6437,8 +6437,10 @@ createBrowserWhisperProvider(options={})
|
|
|
6437
6437
|
|
|
6438
6438
|
The recognized options are
|
|
6439
6439
|
`{id='arcane-browser-whisper',localOnly=true,graph,model,runtime,appSecurity,
|
|
6440
|
-
security,store,offline=false}`.
|
|
6441
|
-
`
|
|
6440
|
+
security,store,offline=false,execution={device:'auto',maxConcurrentRequests:1}}`.
|
|
6441
|
+
`execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
|
|
6442
|
+
exactly 1. `graph` is mutually exclusive with `model` and `runtime`. The mutable
|
|
6443
|
+
result is
|
|
6442
6444
|
`{protocol:'arcane-ai-provider/2',role:'stt',id,localOnly:true,
|
|
6443
6445
|
maxConcurrentRequests:1,catalog,inspect,status,load,request,unload,dispose}`.
|
|
6444
6446
|
The only request operation is
|
|
@@ -6448,8 +6450,22 @@ re-freeze the cloned record.
|
|
|
6448
6450
|
|
|
6449
6451
|
`status()` returns
|
|
6450
6452
|
`{role,providerId,modelId,state,lifecycleStatus,lifecycleReason,activeOperation,
|
|
6451
|
-
loaded,busy,generation,errorCode,cache,warnings}` and includes
|
|
6452
|
-
for an explicit secure intent.
|
|
6453
|
+
loaded,busy,generation,errorCode,cache,warnings,execution}` and includes
|
|
6454
|
+
`security` only for an explicit secure intent.
|
|
6455
|
+
`execution` reports `requestedDevice`, `selectedDevice`,
|
|
6456
|
+
`maxConcurrentRequests`, and `activeRequestCount`; `selectedDevice` is `null`
|
|
6457
|
+
before load and after unload. The high-level projection is
|
|
6458
|
+
`ai.providerRuntime.status('stt', {execution:true}).execution`.
|
|
6459
|
+
|
|
6460
|
+
Automatic loading tries `webnn-npu` when `navigator.ml.createContext` is
|
|
6461
|
+
exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
|
|
6462
|
+
A failed candidate is cleaned up before a fresh Worker tries the next backend
|
|
6463
|
+
with the same prepared model and dtype. Explicit selections do not fall back.
|
|
6464
|
+
`selectedDevice` names the backend requested by the successful upstream session
|
|
6465
|
+
load; it does not prove that every operation ran on a physical accelerator.
|
|
6466
|
+
WebNN may execute unsupported operations through WASM, and exact model,
|
|
6467
|
+
browser, driver, and hardware compatibility remains upstream.
|
|
6468
|
+
|
|
6453
6469
|
States are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
|
|
6454
6470
|
`disposed`. Compatible concurrent loads coalesce; concurrent requests fail as
|
|
6455
6471
|
`ARCANE_AI_PROVIDER_BUSY`. Cancellation after the Worker request begins
|
|
@@ -6523,8 +6539,8 @@ createBrowserKokoroProvider(options={})
|
|
|
6523
6539
|
The recognized options are
|
|
6524
6540
|
`{id='arcane-browser-kokoro',localOnly=true,graph,model,runtime,appSecurity,
|
|
6525
6541
|
security,store,offline=false,execution={device:'auto',maxConcurrentRequests:4}}`.
|
|
6526
|
-
`execution.device` is `auto`, `webgpu`, or `wasm`; its capacity is
|
|
6527
|
-
from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
|
|
6542
|
+
`execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
|
|
6543
|
+
an integer from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
|
|
6528
6544
|
mutable result is
|
|
6529
6545
|
`{protocol:'arcane-ai-provider/2',role:'tts',id,localOnly:true,
|
|
6530
6546
|
maxConcurrentRequests,catalog,inspect,status,load,request,unload,dispose}`. The
|
|
@@ -6543,10 +6559,11 @@ mono 16-bit PCM. Unsupported formats fail
|
|
|
6543
6559
|
`ARCANE_AI_UNSUPPORTED_RESPONSE_FORMAT`; malformed adapter audio fails
|
|
6544
6560
|
`ARCANE_AI_INVALID_PROVIDER_RESULT`. Unknown fields and accessors reject as malformed.
|
|
6545
6561
|
|
|
6546
|
-
Automatic execution
|
|
6547
|
-
|
|
6548
|
-
|
|
6549
|
-
|
|
6562
|
+
Automatic execution tries a complete WebNN NPU pool when
|
|
6563
|
+
`navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
|
|
6564
|
+
exposed, then CPU through WASM. A failed candidate is cleaned up before a fresh
|
|
6565
|
+
pool tries the next backend. Explicit device selections do not fall back. Every
|
|
6566
|
+
slot loads the same caller-selected model and dtype in a distinct Worker so the
|
|
6550
6567
|
selected adapter's per-isolate inference serialization does not serialize the
|
|
6551
6568
|
pool. Direct provider `status().execution` reports `requestedDevice`,
|
|
6552
6569
|
`selectedDevice`, `maxConcurrentRequests`, and `activeRequestCount`.
|
|
@@ -6555,7 +6572,11 @@ by `AI.configureBrowserSpeech()`, use
|
|
|
6555
6572
|
`ai.providerRuntime.status('tts', {execution:true}).execution`. Requested
|
|
6556
6573
|
`auto` with selected `wasm` means automatic fallback occurred. Inspection
|
|
6557
6574
|
errors propagate; the default runtime `status()` remains a sticky lifecycle
|
|
6558
|
-
read.
|
|
6575
|
+
read. `selectedDevice` names the backend requested by the successful upstream
|
|
6576
|
+
session load. These fields do not prove that every operation ran on a physical
|
|
6577
|
+
NPU or GPU, or that accelerator kernels overlap. WebNN may execute unsupported
|
|
6578
|
+
operations through WASM; exact model, browser, driver, and hardware
|
|
6579
|
+
compatibility remains upstream.
|
|
6559
6580
|
|
|
6560
6581
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later
|
|
6561
6582
|
wait in the SDK's provider-neutral FIFO queue; they are not dropped. Synthesis
|
|
@@ -6,13 +6,15 @@ For the smallest first request, start with the [browser speech quick start](../.
|
|
|
6
6
|
|
|
7
7
|
## Speech defaults and inspection
|
|
8
8
|
|
|
9
|
-
The demo
|
|
9
|
+
The demo omits both speech execution records in `speechConfiguration(dbopfs)` so it consumes the SDK defaults: `{device:'auto',maxConcurrentRequests:1}` for Whisper STT and `{device:'auto',maxConcurrentRequests:4}` for Kokoro TTS. Auto tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is exposed, then CPU through WASM. Failed loads are cleaned up before fresh Workers try the next device with the same prepared model and dtype. Explicit `webnn-npu`, `webgpu`, and `wasm` selections do not fall back.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
Change the application's TTS record to `execution:{device:'wasm',maxConcurrentRequests:1}` to use fewer synthesis sessions. TTS accepts capacities 1 through 4; Whisper accepts only 1 and the LLM retains capacity one.
|
|
12
|
+
|
|
13
|
+
The maintained Kokoro selection uses `fp32` because [Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage). Automatic fallback keeps that dtype on every device. `selectedDevice` identifies the backend requested by the successful upstream session load. It does not prove every operation ran on a physical NPU or GPU; WebNN may use WASM for unsupported operations. Exact model, browser, driver, and hardware compatibility and actual speech quality require evaluation on that combination.
|
|
12
14
|
|
|
13
15
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out of order, but playback waits for earlier segments and plays exact input order. Each slot owns a Worker/model session, so raising capacity trades memory for latency.
|
|
14
16
|
|
|
15
|
-
Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected
|
|
17
|
+
Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected. The corresponding Whisper report is available through `ai.providerRuntime.status('stt',{execution:true}).execution`; reading either report does not load a model.
|
|
16
18
|
|
|
17
19
|
Shared Speech still owns mute, stop, and voice controls. The example's status button inspects the public report without loading a model or reaching into private providers.
|
|
18
20
|
|
|
@@ -434,7 +434,7 @@ function speechConfiguration(dbopfs) {
|
|
|
434
434
|
offline: false,
|
|
435
435
|
},
|
|
436
436
|
tts: {
|
|
437
|
-
// Omit execution
|
|
437
|
+
// Omit execution for NPU, GPU, then CPU selection and four synthesis slots.
|
|
438
438
|
// Kokoro.js recommends fp32 for the WebGPU route attempted by auto.
|
|
439
439
|
providerId: "wasm-ai-demo-browser-kokoro",
|
|
440
440
|
model: {
|
package/package.json
CHANGED
|
@@ -27,11 +27,8 @@ const DEFAULT_TTS_SEGMENTATION={
|
|
|
27
27
|
punctuation:'sentence',
|
|
28
28
|
wordCadence:null
|
|
29
29
|
};
|
|
30
|
-
const
|
|
31
|
-
|
|
32
|
-
maxConcurrentRequests:4
|
|
33
|
-
};
|
|
34
|
-
const BROWSER_TTS_EXECUTION_DEVICES=new Set(['auto','webgpu','wasm']);
|
|
30
|
+
const DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE='auto';
|
|
31
|
+
const BROWSER_SPEECH_EXECUTION_DEVICES=new Set(['auto','webnn-npu','webgpu','wasm']);
|
|
35
32
|
const MAX_BROWSER_TTS_CONCURRENT_REQUESTS=4;
|
|
36
33
|
const TTS_PUNCTUATION_MODES=new Set(['sentence','any','none']);
|
|
37
34
|
// A complete punctuation run is a boundary unless the whole run consists of
|
|
@@ -817,36 +814,45 @@ function browserSpeechIdentifier(value,label){
|
|
|
817
814
|
return value;
|
|
818
815
|
}
|
|
819
816
|
|
|
820
|
-
function
|
|
817
|
+
function normalizeBrowserSpeechExecution(role,value){
|
|
818
|
+
const label=`AI browser speech ${role}.execution`;
|
|
819
|
+
const maximumConcurrentRequests=role==='stt'
|
|
820
|
+
?1
|
|
821
|
+
:MAX_BROWSER_TTS_CONCURRENT_REQUESTS;
|
|
821
822
|
if(value===undefined){
|
|
822
|
-
return completeValue({
|
|
823
|
+
return completeValue({
|
|
824
|
+
device:DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE,
|
|
825
|
+
maxConcurrentRequests:maximumConcurrentRequests
|
|
826
|
+
});
|
|
823
827
|
}
|
|
824
828
|
const descriptors=closedRecord(
|
|
825
829
|
value,
|
|
826
830
|
['device','maxConcurrentRequests'],
|
|
827
831
|
[],
|
|
828
|
-
|
|
832
|
+
label
|
|
829
833
|
);
|
|
830
834
|
const device=descriptors.device
|
|
831
835
|
?descriptors.device.value
|
|
832
|
-
:
|
|
836
|
+
:DEFAULT_BROWSER_SPEECH_EXECUTION_DEVICE;
|
|
833
837
|
const maxConcurrentRequests=descriptors.maxConcurrentRequests
|
|
834
838
|
?descriptors.maxConcurrentRequests.value
|
|
835
|
-
:
|
|
836
|
-
if(!is.string(device)||!
|
|
839
|
+
:maximumConcurrentRequests;
|
|
840
|
+
if(!is.string(device)||!BROWSER_SPEECH_EXECUTION_DEVICES.has(device)){
|
|
837
841
|
throw aiBrowserSpeechError(
|
|
838
842
|
AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
|
|
839
843
|
AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
|
|
840
|
-
|
|
844
|
+
`${label}.device must be auto, webnn-npu, webgpu, or wasm.`
|
|
841
845
|
);
|
|
842
846
|
}
|
|
843
847
|
if(!is.safeInteger(maxConcurrentRequests)
|
|
844
848
|
||maxConcurrentRequests<1
|
|
845
|
-
||maxConcurrentRequests>
|
|
849
|
+
||maxConcurrentRequests>maximumConcurrentRequests){
|
|
846
850
|
throw aiBrowserSpeechError(
|
|
847
851
|
AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
|
|
848
852
|
AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
|
|
849
|
-
|
|
853
|
+
role==='stt'
|
|
854
|
+
?`${label}.maxConcurrentRequests must be 1.`
|
|
855
|
+
:`${label}.maxConcurrentRequests must be an integer from 1 through ${maximumConcurrentRequests}.`
|
|
850
856
|
);
|
|
851
857
|
}
|
|
852
858
|
return completeValue({device,maxConcurrentRequests});
|
|
@@ -917,13 +923,6 @@ function normalizeBrowserSpeechRole(value,role){
|
|
|
917
923
|
`${label}.offline must be a boolean.`
|
|
918
924
|
);
|
|
919
925
|
}
|
|
920
|
-
if(role==='stt'&&descriptors.execution){
|
|
921
|
-
throw aiBrowserSpeechError(
|
|
922
|
-
AI_BROWSER_SPEECH_ERROR_CODES.configurationContractMismatch,
|
|
923
|
-
AI_BROWSER_SPEECH_REASONS.configurationContractMismatch,
|
|
924
|
-
'AI browser speech execution policy is available only for TTS.'
|
|
925
|
-
);
|
|
926
|
-
}
|
|
927
926
|
return completeValue({
|
|
928
927
|
providerId,
|
|
929
928
|
...(hasGraph
|
|
@@ -937,11 +936,10 @@ function normalizeBrowserSpeechRole(value,role){
|
|
|
937
936
|
...(secure?{security:{secure:true}}:{})
|
|
938
937
|
}),
|
|
939
938
|
offline:descriptors.offline.value,
|
|
940
|
-
|
|
941
|
-
|
|
942
|
-
|
|
943
|
-
|
|
944
|
-
:{})
|
|
939
|
+
execution:normalizeBrowserSpeechExecution(
|
|
940
|
+
role,
|
|
941
|
+
descriptors.execution?.value
|
|
942
|
+
)
|
|
945
943
|
});
|
|
946
944
|
}
|
|
947
945
|
|
|
@@ -2770,10 +2768,10 @@ class AI {
|
|
|
2770
2768
|
?{artifactGraphId:catalog.artifactGraphId}
|
|
2771
2769
|
:{}),
|
|
2772
2770
|
offline:configured.offline,
|
|
2771
|
+
execution:configured.execution,
|
|
2773
2772
|
...(role==='tts'
|
|
2774
2773
|
?{
|
|
2775
|
-
defaultVoice:catalog.defaultVoice
|
|
2776
|
-
execution:configured.execution
|
|
2774
|
+
defaultVoice:catalog.defaultVoice
|
|
2777
2775
|
}
|
|
2778
2776
|
:{})
|
|
2779
2777
|
});
|
|
@@ -3145,7 +3143,7 @@ class AI {
|
|
|
3145
3143
|
}),
|
|
3146
3144
|
store,
|
|
3147
3145
|
offline:configured.offline,
|
|
3148
|
-
|
|
3146
|
+
execution:configured.execution
|
|
3149
3147
|
});
|
|
3150
3148
|
}
|
|
3151
3149
|
providers=completeValue({...candidateProviders});
|