arcane-os 0.5.19 → 0.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +26 -0
- package/README.md +18 -12
- package/browser-runtime/ai/browser-device-settings.mjs +83 -0
- package/browser-runtime/ai/browser-speech-providers.mjs +64 -60
- package/browser-runtime/ai/browser-wasm-llm-provider.mjs +4 -48
- package/browser-runtime/ai/speech-worker-runtime.mjs +4 -5
- package/docs/architecture.md +1 -1
- package/docs/reference/ai/browser-speech.md +95 -38
- package/docs/reference/behavioral-testing.md +1 -1
- package/docs/reference/inventory/runtime-modules.json +2 -2
- package/docs/reference/runtime-components.md +53 -0
- package/docs/reference/runtime-modules.md +32 -17
- package/docs/reference/sdk-api.md +32 -11
- package/examples/wasm-ai-demo/README.md +5 -3
- package/examples/wasm-ai-demo/app.js +1 -1
- package/package.json +1 -1
- package/runtime/arcane/components/browser-ai-setup.html +212 -0
- package/runtime/arcane/modules/AI.js +27 -29
|
@@ -55,11 +55,11 @@ export const speechSelection = {
|
|
|
55
55
|
};
|
|
56
56
|
```
|
|
57
57
|
|
|
58
|
-
The omitted execution record below uses the SDK's
|
|
59
|
-
selection
|
|
58
|
+
The omitted execution record below uses the SDK's NPU, GPU, then CPU automatic
|
|
59
|
+
selection. This basic configuration uses `fp32` because
|
|
60
60
|
[Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage).
|
|
61
|
-
Automatic fallback carries this same selected model and dtype
|
|
62
|
-
does not rewrite the application selection. If you intentionally choose
|
|
61
|
+
Automatic fallback carries this same selected model and dtype between devices;
|
|
62
|
+
the SDK does not rewrite the application selection. If you intentionally choose
|
|
63
63
|
another dtype, evaluate that exact model, browser, and execution route.
|
|
64
64
|
`selectedDevice` reports routing after load, not pronunciation, text fidelity,
|
|
65
65
|
or audio quality.
|
|
@@ -130,6 +130,36 @@ window.addEventListener('ai-tts-failure', function reportSpeechFailure(event) {
|
|
|
130
130
|
|
|
131
131
|
Call `speechEvents.abort()` when disposing that interface to remove the listener.
|
|
132
132
|
|
|
133
|
+
## Browser NPU setup
|
|
134
|
+
|
|
135
|
+
Applications can place the shared `browser-ai-setup.html` component in their
|
|
136
|
+
profile or settings page. It reports whether this page exposes WebNN and
|
|
137
|
+
WebGPU, without loading a model or creating an accelerator context:
|
|
138
|
+
|
|
139
|
+
```html
|
|
140
|
+
<html-import
|
|
141
|
+
id="browserAISetup"
|
|
142
|
+
href="/arcane/components/browser-ai-setup.html">
|
|
143
|
+
</html-import>
|
|
144
|
+
```
|
|
145
|
+
|
|
146
|
+
**Set up NPU** uses the same browser-settings approach as the existing
|
|
147
|
+
high-performance GPU notice: attempt to open the browser's flags page, then
|
|
148
|
+
show instructions including the full address to paste if navigation was
|
|
149
|
+
blocked. Chrome uses `chrome://flags/#web-machine-learning-neural-network`;
|
|
150
|
+
Edge uses `edge://flags/#web-machine-learning-neural-network`. Unrecognized
|
|
151
|
+
browsers receive explicit Chrome and Edge choices instead of an assumed target.
|
|
152
|
+
The [ONNX Runtime WebNN guide](https://onnxruntime.ai/docs/tutorials/web/ep-webnn.html)
|
|
153
|
+
documents the **Enables WebNN API** flag and model/operator requirements.
|
|
154
|
+
|
|
155
|
+
This control does not save an execution preference, change browser settings,
|
|
156
|
+
restart the browser, or report that the NPU is active. A WebNN API presence
|
|
157
|
+
result is not proof of NPU hardware, a compatible model, or physical execution.
|
|
158
|
+
The profile also cannot report a different chat page's loaded runtime. The
|
|
159
|
+
existing automatic speech route remains NPU, then GPU, then CPU; upstream
|
|
160
|
+
sessions can still place unsupported operators on CPU. Use the owning model
|
|
161
|
+
runtime's evidence to determine actual accelerator execution.
|
|
162
|
+
|
|
133
163
|
## Developer diagnostics
|
|
134
164
|
|
|
135
165
|
The shared logging API and speech traces are available in SDK `0.5.14`.
|
|
@@ -229,9 +259,10 @@ playback. `replay()` keeps completed and pending provider segments and retries
|
|
|
229
259
|
only failed missing segments.
|
|
230
260
|
|
|
231
261
|
The default `{device:'auto',maxConcurrentRequests:4}` attempts the full ONNX
|
|
232
|
-
Worker/session pool on
|
|
233
|
-
|
|
234
|
-
|
|
262
|
+
Worker/session pool on WebNN NPU, then WebGPU, then CPU through WASM. It skips
|
|
263
|
+
an accelerator when its browser API is absent and replaces a failed candidate
|
|
264
|
+
with a fresh pool before trying the next device. The basic configuration above
|
|
265
|
+
keeps `fp32` throughout that sequence. Use the status example below to read
|
|
235
266
|
`selectedDevice`; console node-assignment warnings alone do not identify the
|
|
236
267
|
selected execution device or assess the generated audio.
|
|
237
268
|
|
|
@@ -588,8 +619,20 @@ console.log(speechText.append('ing', true)); // ing
|
|
|
588
619
|
|
|
589
620
|
## Choose a device or reduce memory use
|
|
590
621
|
|
|
591
|
-
|
|
592
|
-
|
|
622
|
+
Both speech roles accept `execution:{device,maxConcurrentRequests}`. Omitting
|
|
623
|
+
`stt.execution` selects `{device:'auto',maxConcurrentRequests:1}`; omitting
|
|
624
|
+
`tts.execution` selects `{device:'auto',maxConcurrentRequests:4}`. Whisper keeps
|
|
625
|
+
one transcription slot. Kokoro accepts capacities 1 through 4.
|
|
626
|
+
|
|
627
|
+
Automatic selection tries `webnn-npu` when `navigator.ml.createContext` is
|
|
628
|
+
exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
|
|
629
|
+
API exposure only determines which upstream backend to attempt; it
|
|
630
|
+
does not establish compatible hardware, operators, or model shapes. A failed
|
|
631
|
+
candidate is fully cleaned up before the next candidate uses fresh Workers
|
|
632
|
+
with the same prepared model and dtype. An explicit `webnn-npu`, `webgpu`, or
|
|
633
|
+
`wasm` selection attempts only that backend and reports its failure.
|
|
634
|
+
|
|
635
|
+
These are four alternative TTS configurations:
|
|
593
636
|
|
|
594
637
|
```javascript
|
|
595
638
|
async function selectSpeechExecution(execution) {
|
|
@@ -609,15 +652,22 @@ async function selectSpeechExecution(execution) {
|
|
|
609
652
|
|
|
610
653
|
// Choose and call one from your application settings action:
|
|
611
654
|
// await selectSpeechExecution({ device: 'auto' });
|
|
655
|
+
// await selectSpeechExecution({ device: 'webnn-npu' });
|
|
612
656
|
// await selectSpeechExecution({ device: 'webgpu' });
|
|
613
657
|
// await selectSpeechExecution({ device: 'wasm', maxConcurrentRequests: 1 });
|
|
614
658
|
```
|
|
615
659
|
|
|
616
|
-
The override accepts integers 1, 2, 3, or 4
|
|
617
|
-
|
|
618
|
-
|
|
619
|
-
|
|
620
|
-
|
|
660
|
+
The TTS capacity override accepts integers 1, 2, 3, or 4; an STT capacity
|
|
661
|
+
override accepts only 1. Apply the same device choices to the configured
|
|
662
|
+
`stt.execution` record. A configuration change leaves TTS muted; explicitly
|
|
663
|
+
load/unmute again.
|
|
664
|
+
|
|
665
|
+
The selected upstream versions expose WebNN NPU through
|
|
666
|
+
[Transformers.js device selection](https://github.com/huggingface/transformers.js/blob/4.2.0/packages/transformers/src/backends/onnx.js)
|
|
667
|
+
and [Kokoro.js device forwarding](https://github.com/hexgrad/kokoro/blob/664c76a704021239ba59c84dcbaa4d3dece01fe9/kokoro.js/src/kokoro.js).
|
|
668
|
+
WebNN compatibility depends on the model's shapes and operations, browser,
|
|
669
|
+
drivers, and hardware. Unsupported operations may run through WASM even after
|
|
670
|
+
an NPU session loads; see the [ONNX Runtime WebNN contract](https://onnxruntime.ai/docs/tutorials/web/ep-webnn.html).
|
|
621
671
|
|
|
622
672
|
## Inspect the requested and selected device
|
|
623
673
|
|
|
@@ -627,15 +677,15 @@ loading. This explicitly reads the selected provider's current report; ordinary
|
|
|
627
677
|
separate execution-state event subscription.
|
|
628
678
|
|
|
629
679
|
```javascript
|
|
630
|
-
function printSpeechStatus() {
|
|
631
|
-
const status = ai.providerRuntime.status(
|
|
680
|
+
function printSpeechStatus(role = 'tts') {
|
|
681
|
+
const status = ai.providerRuntime.status(role, { execution: true });
|
|
632
682
|
const execution = status.execution;
|
|
633
|
-
console.log('
|
|
683
|
+
console.log('Speech role and state:', role, status.state);
|
|
634
684
|
if (execution) {
|
|
635
685
|
console.log('Requested device:', execution.requestedDevice);
|
|
636
686
|
console.log('Selected device:', execution.selectedDevice);
|
|
637
687
|
console.log('Capacity:', execution.maxConcurrentRequests);
|
|
638
|
-
console.log('Active
|
|
688
|
+
console.log('Active requests:', execution.activeRequestCount);
|
|
639
689
|
console.log('Automatic WASM fallback:',
|
|
640
690
|
execution.requestedDevice === 'auto' && execution.selectedDevice === 'wasm');
|
|
641
691
|
}
|
|
@@ -645,13 +695,18 @@ function printSpeechStatus() {
|
|
|
645
695
|
Call `printSpeechStatus()` after the load in `sayHello()` or from your status
|
|
646
696
|
button. The same projection is at
|
|
647
697
|
`ai.providerRuntime.status(null, {execution:true}).roles.tts.execution`.
|
|
698
|
+
For configured Whisper, call `printSpeechStatus('stt')` or read the corresponding
|
|
699
|
+
`roles.stt.execution` projection.
|
|
648
700
|
`selectedDevice` is `null`
|
|
649
701
|
until a pool is selected and returns to `null` on unload. Providers without an
|
|
650
702
|
execution report omit `execution`; do not infer a device from `navigator.gpu`
|
|
651
703
|
or a configured preference alone. An explicit inspection can throw a provider
|
|
652
704
|
status error; handle it with the same `error.code` / `error.message` pattern.
|
|
653
|
-
|
|
654
|
-
|
|
705
|
+
`selectedDevice` names the backend requested by the successful upstream session
|
|
706
|
+
load. It does not prove that every operation ran on a physical NPU or GPU, or
|
|
707
|
+
establish transcription correctness, pronunciation, or audio quality. Evaluate
|
|
708
|
+
actual speech output for the model, dtype, browser, and device combinations your
|
|
709
|
+
application supports.
|
|
655
710
|
|
|
656
711
|
## Stop, mute, cancel, and release
|
|
657
712
|
|
|
@@ -991,11 +1046,12 @@ they do not silently convert into a rejection policy.
|
|
|
991
1046
|
The Worker applies only the selected runtime settings needed to run the chosen
|
|
992
1047
|
provider:
|
|
993
1048
|
|
|
994
|
-
- Kokoro forwards the pool's selected `webgpu
|
|
1049
|
+
- Kokoro forwards the pool's selected `webnn-npu`, `webgpu`, or `wasm` device to
|
|
995
1050
|
`KokoroTTS.from_pretrained()`. Its configured dtype remains exactly the
|
|
996
|
-
caller-selected dtype on
|
|
1051
|
+
caller-selected dtype on every path. The WASM path also uses
|
|
997
1052
|
`namespace.env.wasmPaths = {mjs,wasm}`.
|
|
998
|
-
- Transformers
|
|
1053
|
+
- Transformers forwards the selected device to its speech-recognition pipeline,
|
|
1054
|
+
preserves the caller-selected dtype, and uses
|
|
999
1055
|
`namespace.env.backends.onnx.wasm.wasmPaths = {mjs,wasm}`, keeps remote model
|
|
1000
1056
|
loading enabled, and applies caller-selected `numThreads` when present.
|
|
1001
1057
|
|
|
@@ -1029,15 +1085,15 @@ The constructors also accept ordinary `model` and `runtime` descriptors instead
|
|
|
1029
1085
|
of `graph`; the two forms are mutually exclusive. Both forms require an
|
|
1030
1086
|
SDK-created DBOPFS speech artifact store. `localOnly` remains `true`.
|
|
1031
1087
|
|
|
1032
|
-
|
|
1033
|
-
`{device,maxConcurrentRequests}`. `device` is `auto`, `webgpu`, or
|
|
1034
|
-
`
|
|
1035
|
-
`{device:'auto',maxConcurrentRequests:
|
|
1036
|
-
|
|
1037
|
-
|
|
1038
|
-
|
|
1039
|
-
|
|
1040
|
-
|
|
1088
|
+
Both constructors accept the exact `execution` record
|
|
1089
|
+
`{device,maxConcurrentRequests}`. `device` is `auto`, `webnn-npu`, `webgpu`, or
|
|
1090
|
+
`wasm`. Whisper permits capacity 1 and defaults to
|
|
1091
|
+
`{device:'auto',maxConcurrentRequests:1}`. Kokoro permits integer capacities 1
|
|
1092
|
+
through 4 and defaults to `{device:'auto',maxConcurrentRequests:4}`. `auto`
|
|
1093
|
+
tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when
|
|
1094
|
+
`navigator.gpu` is exposed, then WASM. Before advancing after a failed load,
|
|
1095
|
+
the SDK tears down the candidate and creates fresh Workers using the same
|
|
1096
|
+
prepared model and dtype. Explicit device selections never fall back.
|
|
1041
1097
|
|
|
1042
1098
|
Each constructor returns an `arcane-ai-provider/2` object with:
|
|
1043
1099
|
|
|
@@ -1123,12 +1179,13 @@ mutable values.
|
|
|
1123
1179
|
Provider states are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
|
|
1124
1180
|
`disposed`. `status()` includes role, provider/model ids, state, lifecycle
|
|
1125
1181
|
status and reason, active operation, loaded/busy flags, generation, error code,
|
|
1126
|
-
cache state, and warnings.
|
|
1182
|
+
cache state, and warnings. Both speech roles include an `execution` record
|
|
1127
1183
|
with requested and selected device, request limit, and active request count. A
|
|
1128
|
-
successful `selectedDevice
|
|
1129
|
-
|
|
1130
|
-
overlap
|
|
1131
|
-
security field is absent
|
|
1184
|
+
successful `selectedDevice` reports the backend requested by the upstream model
|
|
1185
|
+
load. It does not prove that every operation ran on a physical NPU or GPU,
|
|
1186
|
+
that accelerator kernels overlap, or that generated speech is correct. WebNN
|
|
1187
|
+
may execute unsupported operations through WASM. A security field is absent
|
|
1188
|
+
in ordinary mode.
|
|
1132
1189
|
|
|
1133
1190
|
The provider/2 load context accepts an optional progress callback for interface
|
|
1134
1191
|
compatibility, but the current browser-speech artifact and Worker transport
|
|
@@ -1203,7 +1260,7 @@ operation.
|
|
|
1203
1260
|
## Ownership
|
|
1204
1261
|
|
|
1205
1262
|
- Applications own model, runtime, dtype, sample-rate, voice, profile, prompt,
|
|
1206
|
-
catalog, activation, optional TTS execution override, and presentation
|
|
1263
|
+
catalog, activation, optional STT/TTS execution override, and presentation
|
|
1207
1264
|
policy.
|
|
1208
1265
|
- Upstream publishers own their runtime, model, voice, and license delivery.
|
|
1209
1266
|
- The SDK owns storage, materialization, routing, Worker lifecycle, normalized
|
|
@@ -41,7 +41,7 @@ development evidence; it is not native artifact or release acceptance.
|
|
|
41
41
|
| Browser runtime modules | Every shipped ESM module parses and its export inventory matches the catalog; pure helpers run focused success/error cases. | DOM, OPFS, media, and Web Component journeys use a browser harness. |
|
|
42
42
|
| Provider-neutral AI runtime and chat/speech activation | Provider/2 registration, three-role configuration, TWiN Cloud LLM readiness, on-device Whisper/Kokoro selection, opt-in STT startup, Core speech readiness, independent LLM/STT/TTS load/unload/status, capacity-1 FIFO settlement for LLM/STT, bounded parallel synthesis with FIFO admission for an explicitly capable TTS provider, owned STT signals, TTS mute lifecycle, route-owned voice defaults, immediate chunk synthesis admission, original-order audio-clock scheduling, sticky-state-only readiness for both speech components, selected-unloaded activation request/cancellation/error behavior, programmatic voice recording, transcript-replacement supersession of late settlement, direct `AI.fetchSTT` result delivery, rejection of non-local speech configuration, and absence of silent provider fallback are represented against complete providers and host callbacks. | Real model/runtime availability remains the selected provider's evidence boundary; provider-promise settlement, state, an abort signal, or an activation event does not by itself prove underlying provider work stopped or native, cloud, or browser-model availability. |
|
|
43
43
|
| Browser-WASM local AI | The exported namespace, canonical ordered `{id, files:[{name?,url},...]}` descriptor, nonempty provider `sources` catalog, default `secure:false`, dormant `secure:true` intent, public AI API module lifecycle, lazy/manual policy, successful Wllama-load requirement, abort normalization, complete output and reasoning, all-choice validation, required structural-call `arguments.message`, exact call identity, and matching tool-result sequencing are represented in deterministic fixture sources. | A real Chrome exercise may load the selected Wllama runtime and model only after explicit user action. It is not an ordinary publication gate or an implicit model download. |
|
|
44
|
-
| Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, GPU-
|
|
44
|
+
| Browser speech | Caller-owned Whisper/Kokoro selection, independent STT/TTS routes, ordinary direct upstream runtime/model use, dormant `secure:true` intent, DBOPFS cache, materialized known-route/native-fallback behavior, NPU-to-GPU-to-CPU automatic selection for both roles, fresh Workers after a failed device load, explicit-device failure without fallback, bounded Kokoro Worker/session pools, out-of-order synthesis settlement with FIFO admission, request-targeted TTS cancellation, destructive lifecycle cleanup, Blob/File STT conversion, WAV TTS conversion, complete text/audio, and no cloud fallback are represented with synthetic artifacts and adapters. | A real runtime/model/voice download, WebNN or WebGPU model load, physical accelerator behavior, and actual transcription or synthesis use the application's selected upstream packages/providers, browser media support, and explicit user action. A selected backend is not proof that every operation ran on that physical accelerator; WebNN may execute unsupported operations through WASM. |
|
|
45
45
|
| Persistent chat and document context | Atomic in-memory/history commit, explicit per-turn persistence, streamed/non-stream fallback, session-owned callback fields, complete data callbacks, per-turn request options, exact per-choice streamed/terminal call correlation before publication, terminal-only call acceptance, ordered parallel-call sequencing, atomic all-ID nonblank executed/declined/cancelled/not-executed result batches, readable unmodified malformed pre-existing rows, complete UI transcript metadata, generic visible failure outcomes with complete console diagnostics, BFCache-preserving component lifecycle, complete bootstrap/search/context, caller-source evaluation, cancellation, and partial-read handling are represented with app-scoped adapters. | Live Core/provider inference and durable browser storage remain separate authorities; tests never treat a fake chat function or in-memory adapter as host/storage proof. |
|
|
46
46
|
| Core bridge docs | Canonical namespace/method/event/entity inventories match their one-per-member guides and required sections. | Live Core conformance belongs in Arcane OS because Core implementation is not shipped as SDK source. |
|
|
47
47
|
| Arcane Ollama wrapper | Missing-host error, method forwarding, text/readiness normalization, unload request, and stream-option forwarding run against a deterministic fake `Arcane.ollama`. | Real managed-service, model download/create, GPU/resource admission, and service restart require an admitted Arcane host. |
|
|
@@ -5,7 +5,7 @@
|
|
|
5
5
|
"repository": "https://github.com/TheWizardNexus/arcane-os-sdk.git",
|
|
6
6
|
"branch": "main",
|
|
7
7
|
"path": "runtime/arcane",
|
|
8
|
-
"sdkVersion": "0.
|
|
8
|
+
"sdkVersion": "0.6.0",
|
|
9
9
|
"protocol": "arcane/1"
|
|
10
10
|
},
|
|
11
11
|
"artifactCount": 85,
|
|
@@ -29,7 +29,7 @@
|
|
|
29
29
|
"summary": "Owns provider-selectable chat and the one-time caller-authority browser STT/TTS configuration, lifecycle, synthesis, transcription, and playback boundary.",
|
|
30
30
|
"availability": "Browser + native bridge + TWiN Cloud",
|
|
31
31
|
"protocol": "arcane-ai-browser-speech-configuration/1, AIProviderRuntime arcane-ai-provider/2 routes, globalThis.arcaneEvents, TWiN Cloud HTTPS, Arcane.ollama, Arcane.speech, Android WebView bridge",
|
|
32
|
-
"normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Kokoro TTS execution with default
|
|
32
|
+
"normalization": "Complete mutable caller-owned browser speech authority, SDK-owned provider registration/replacement/disposal, explicit STT/TTS activation, normalized Whisper STT and Kokoro TTS execution with default capacities 1 and 4 respectively, explicit provider-neutral execution snapshots through status(role,{execution:true}), sticky readiness, canonical window.user readiness without compatibility-event object-identity admission or polling, exact ordered structural calls with required arguments.message, atomic nonblank all-ID tool-result sequencing, complete all-choice streaming/data validation, native Ollama structural adaptation, complete-response then all validated tool callbacks then completion ordering, observed async callbacks, Blob/File transcription, automatic repeated-formatting-mark removal from cloned outbound TTS input while caller and visible content stays exact, immediate exact-segment synthesis, per-call voice/speed capture, optional final-segment pause on the existing audio clock, optional per-call terminal playback completion, detached preparation with complete semantic DBOPFS audio reuse and same-owner pending request sharing, independent prepared playback without a model load for cached audio, ordered audio-clock scheduling, playable audio, active-generation TTS operation-failure routing, mute, and cancellation are normalized; provider/model/runtime/voice policy remains caller-owned. Both speech roles default execution.device to auto: webnn-npu when navigator.ml.createContext is exposed, then webgpu when navigator.gpu is exposed, then wasm for CPU execution. Failed backend loads clean up the pool before trying fresh Workers with the same prepared model and dtype; explicit devices never fall back. STT capacity only accepts 1; TTS accepts 1 through 4. Execution snapshots expose requestedDevice, selectedDevice, maxConcurrentRequests, and activeRequestCount for both roles; selectedDevice is null while unloaded and names the successful upstream session backend request when loaded. It does not prove every operation used the physical NPU: WebNN unsupported operations may use WASM, and compatibility depends on the browser, driver, hardware, and graph.",
|
|
33
33
|
"surface": "Browser-speech protocol/event/error/reason constants; default `AI`; read-only `providerRuntime`, `browserSpeechConfiguration`, and `browserSpeechDescriptor`; explicit `providerRuntime.status(role,{execution:true})` snapshots; `streamTTS(text,end,options={})` with optional voice, speed, pauseAfterMs, and waitForPlayback; automatic speech-input cleanup with SDK-internal preparation metadata; `prepareTTS({parts,storage,identity,signal,onState})` returning ordered segments, state, ready, getAudio(index), and cancel(); `playPreparedTTS(prepared,{signal,onState})` returning state, error, finished, pause(), resume(), and stop() with one playback lane per AI and observational state/error callbacks; configure/dispose speech, route lifecycle, declaration-validated chat/stream requests with exact ordered calls, synthesis/transcription, and playback controls; initializes from current canonical user readiness, installs `window.ai`, and projects `ai-ready` plus active-generation `ai-tts-failure`."
|
|
34
34
|
},
|
|
35
35
|
{
|
|
@@ -60,6 +60,7 @@ appropriate.
|
|
|
60
60
|
| --- | --- | --- | --- | --- |
|
|
61
61
|
| [`app-bar.html`](#app-barhtml) | Responsive application navigation, route state, status, and trailing actions. | `setNavigation()`<br>`setActiveRoute()`<br>`setStatus()`<br>`refresh()`<br>`destroy()` | `app-bar-ready` | DOM-normalized |
|
|
62
62
|
| [`assistant-panel.html`](#assistant-panelhtml) | Reusable assistant drawer, message area, composer, pending/streaming/empty/error state, and actions. | `open()`<br>`close()`<br>`toggle()`<br>`send()`<br>`clear()`<br>`setState()`<br>`focusComposer()`<br>`scrollToEnd()`<br>`destroy()` | `assistant-ready`<br>`assistant-opened`<br>`assistant-closed`<br>`assistant-send`<br>`assistant-clear` | DOM-normalized; caller/provider results remain external |
|
|
63
|
+
| [`browser-ai-setup.html`](#browser-ai-setuphtml) | Browser API availability and NPU setup instructions with browser-specific flags actions. | `refresh()`<br>`open()`<br>`destroy()`<br>`ready` | `browser-ai-setup-ready` | API presence only; settings navigation and hardware execution remain browser-owned |
|
|
63
64
|
| [`calculator.html`](#calculatorhtml) | Calculator keypad and result/error event surface backed by CalculatorEngine. | `calculate()`<br>`destroy()` | `calculator-ready`<br>`calculation-complete`<br>`calculation-error` | Normalized Calculation/error events |
|
|
64
65
|
| [`chart.html`](#charthtml) | Accessible uPlot line, area, or point chart with normalized options and rows. | `configure()`<br>`populate()`<br>`setData()`<br>`addData()`<br>`update()`<br>`destroy()` | `chart-ready`<br>`chart-remove` | Options/rows normalized; uPlot rendering is vendor-native |
|
|
65
66
|
| [`chat.html`](#chathtml) | Shared chat, visible selected-model activation request, file upload, streaming, structural tool settlement, speech, language, availability, and conversation-timebox surface. | `streamMessage()`<br>`setMessageProgress()`<br>`setAIAvailability()`<br>`setInitialSpeechMuted()`<br>`setConversationComplete()`<br>`bindConversationTimebox()`<br>`bindSession()`<br>`submitMessage()`<br>`submitToolResult()`<br>`submitToolResults()`<br>`sendMessage()`<br>`languageChanged()`<br>`requestAIActivation()`<br>`destroy()` | `chat-ready`<br>`chat-session-bound`<br>`chat-session-message`<br>`chat-session-error`<br>`chat-send-message`<br>`chat-send-error`<br>`chat-file-uploaded`<br>`chat-file-upload-error`<br>`chat-language-changed`<br>`chat-language-change-error`<br>`chat-ai-activation-request`<br>`chat-ai-activation-error`<br>`chat-speech-synthesis-error`<br>`conversation-timebox-error` | UI/runtime state, explicit user activation intent, and honest structural-call settlement normalized; AI/storage/media behavior mixed |
|
|
@@ -153,6 +154,58 @@ Slots: `title`, `subtitle`, `identity`, `messages/message`, `composer`, `actions
|
|
|
153
154
|
</html-import>
|
|
154
155
|
```
|
|
155
156
|
|
|
157
|
+
## browser-ai-setup.html
|
|
158
|
+
|
|
159
|
+
### Overview
|
|
160
|
+
|
|
161
|
+
Displays WebNN and WebGPU API availability for the current page and explains how
|
|
162
|
+
to enable WebNN for NPU use. Its setup action uses the same browser-flags opening
|
|
163
|
+
attempt and alert fallback as the existing high-performance GPU notice. Chrome
|
|
164
|
+
and Edge receive their own flags addresses; other or unidentified browsers show
|
|
165
|
+
both explicit choices.
|
|
166
|
+
|
|
167
|
+
### Public surface
|
|
168
|
+
|
|
169
|
+
Methods/properties: `refresh()`, `open(browserId?)`, `destroy()`, `ready`.
|
|
170
|
+
|
|
171
|
+
Events: `browser-ai-setup-ready`.
|
|
172
|
+
|
|
173
|
+
`refresh()` synchronously updates API presence and browser guidance and returns
|
|
174
|
+
the current settings record: `browserId`, `name`, `webnnFlagsURL`,
|
|
175
|
+
`highPerformanceGpu`, `webnnAvailable`, and `webgpuAvailable`. It returns `false`
|
|
176
|
+
after destruction. Availability means that the browser exposes the API to this
|
|
177
|
+
page; it does not confirm a working adapter, loaded model, or hardware execution.
|
|
178
|
+
|
|
179
|
+
`open()` refreshes the display, attempts to open the detected Chrome or Edge flags
|
|
180
|
+
page, and presents an alert containing the full address and enable/relaunch
|
|
181
|
+
instructions. Pass `"chrome"` or `"edge"` to choose explicitly. Invoke it from a
|
|
182
|
+
user action. It returns `false` when destroyed or when no supported target is
|
|
183
|
+
selected, and otherwise returns `undefined`; it never reports that navigation
|
|
184
|
+
succeeded. Browsers may block internal-page navigation, so the full addresses
|
|
185
|
+
also remain visible and copyable in the component.
|
|
186
|
+
|
|
187
|
+
`destroy()` aborts owned listeners, disposes the event source, marks `ready`
|
|
188
|
+
false, and suppresses UI updates from a pending clipboard operation. It returns
|
|
189
|
+
`true` the first time and `false` thereafter. A BFCache-persisted `pagehide`
|
|
190
|
+
preserves the component; nonpersisted `pagehide` destroys it.
|
|
191
|
+
|
|
192
|
+
### Availability and normalization
|
|
193
|
+
|
|
194
|
+
**Browser and supported native WebViews.** API presence and setup instructions
|
|
195
|
+
are normalized; browser flags, clipboard support, and device execution remain
|
|
196
|
+
platform-owned. Mounting or refreshing creates no model, GPU adapter, or WebNN
|
|
197
|
+
context. The component saves no preferences, changes no browser settings, and
|
|
198
|
+
does not restart the browser. It cannot report model state owned by another page.
|
|
199
|
+
|
|
200
|
+
### Example
|
|
201
|
+
|
|
202
|
+
```html
|
|
203
|
+
<html-import
|
|
204
|
+
id="browser-ai-setup"
|
|
205
|
+
href="/arcane/components/browser-ai-setup.html">
|
|
206
|
+
</html-import>
|
|
207
|
+
```
|
|
208
|
+
|
|
156
209
|
## calculator.html
|
|
157
210
|
|
|
158
211
|
### Overview
|
|
@@ -253,10 +253,15 @@ user activation intent exposed by the shared speech component.
|
|
|
253
253
|
`setSpeechMuted(false)` records the public unmuted state only after the selected
|
|
254
254
|
TTS route reaches ready; a failed load leaves the public state muted. In contrast,
|
|
255
255
|
`setSpeechMuted(true)` cancels active TTS work and unloads that role.
|
|
256
|
-
The optional browser-speech `tts.execution`
|
|
257
|
-
`device:'auto'|'webgpu'|'wasm'` and
|
|
258
|
-
|
|
259
|
-
Worker/session slots
|
|
256
|
+
The optional browser-speech `stt.execution` and `tts.execution` records select
|
|
257
|
+
`device:'auto'|'webnn-npu'|'webgpu'|'wasm'` and `maxConcurrentRequests`.
|
|
258
|
+
Omission uses NPU, GPU, then CPU automatic selection with one Whisper slot and
|
|
259
|
+
four bounded Kokoro Worker/session slots. Whisper accepts only capacity 1;
|
|
260
|
+
Kokoro accepts integers 1 through 4. Automatic loading attempts WebNN NPU when
|
|
261
|
+
`navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
|
|
262
|
+
exposed, then WASM. A failed candidate is cleaned up before fresh Workers try
|
|
263
|
+
the next device with the same prepared model and dtype. Explicit device
|
|
264
|
+
selections report failure without falling back.
|
|
260
265
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later
|
|
261
266
|
wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out
|
|
262
267
|
of order, but playback waits for earlier segments and plays exact input order.
|
|
@@ -265,11 +270,15 @@ latency. This capacity does not establish physical GPU kernel overlap.
|
|
|
265
270
|
|
|
266
271
|
After configuration, explicitly inspect execution through
|
|
267
272
|
`ai.providerRuntime.status('tts', {execution:true}).execution`. When supplied
|
|
268
|
-
by the selected provider, this read returns its execution snapshot.
|
|
269
|
-
|
|
273
|
+
by the selected provider, this read returns its execution snapshot. Whisper and
|
|
274
|
+
Kokoro report `requestedDevice`, `selectedDevice`, `maxConcurrentRequests`, and
|
|
270
275
|
`activeRequestCount`. `selectedDevice` is `null` before load and after unload.
|
|
271
276
|
`requestedDevice === 'auto' && selectedDevice === 'wasm'` identifies automatic
|
|
272
|
-
WASM fallback after a successful load.
|
|
277
|
+
WASM fallback after a successful load. Read the same fields for Whisper with
|
|
278
|
+
`status('stt', {execution:true})`. The selected device is the backend requested
|
|
279
|
+
by a successful upstream session load, not proof that every operation ran on a
|
|
280
|
+
physical accelerator; WebNN may execute unsupported operations through WASM.
|
|
281
|
+
Calling `status()` without options keeps
|
|
273
282
|
the existing sticky lifecycle snapshot and does not inspect provider execution.
|
|
274
283
|
Provider inspection failures are surfaced to the caller.
|
|
275
284
|
`fetchTTS({model,voice,input,responseFormat,speed},signal,preparation={})` accepts the public
|
|
@@ -501,9 +510,11 @@ await ai.disposeBrowserSpeech({signal});
|
|
|
501
510
|
|
|
502
511
|
The record is a mutable plain data record with exactly
|
|
503
512
|
`{protocol,id,dbopfs,tableName?,stt?,tts?}` and at least one role. Each supplied
|
|
504
|
-
mutable STT role is exactly
|
|
505
|
-
`{providerId,
|
|
506
|
-
|
|
513
|
+
mutable STT or TTS role is exactly
|
|
514
|
+
`{providerId,graph,security?,offline,execution?}` or
|
|
515
|
+
`{providerId,model,runtime,security?,offline,execution?}`. The optional
|
|
516
|
+
`execution` record contains `{device,maxConcurrentRequests}` with the
|
|
517
|
+
role-specific defaults and capacities described above. The graph and
|
|
507
518
|
direct authority forms are mutually exclusive; `providerId` and `id` are nonblank exact strings,
|
|
508
519
|
`graph` is the role-matching graph returned by the SDK browser
|
|
509
520
|
speech artifact API, and `offline` is boolean. The direct form forwards its
|
|
@@ -526,8 +537,9 @@ or reproduce DBOPFS cache logic.
|
|
|
526
537
|
|
|
527
538
|
The returned descriptor is exactly `{protocol,configurationId,stt,tts}`; an
|
|
528
539
|
external, unmanaged role is `null`. A managed STT descriptor is
|
|
529
|
-
`{role:'stt',providerId,modelId,artifactGraphId?,offline}`; TTS adds
|
|
530
|
-
`defaultVoice
|
|
540
|
+
`{role:'stt',providerId,modelId,artifactGraphId?,offline,execution}`; TTS adds
|
|
541
|
+
`defaultVoice`. Both include the normalized `execution` record.
|
|
542
|
+
`artifactGraphId` is present only for the graph form.
|
|
531
543
|
`browserSpeechConfiguration` returns the exact caller-owned record when no
|
|
532
544
|
managed role is carried. After a partial replacement that carries another
|
|
533
545
|
managed role, it returns a mutable merged record with the replacement call's
|
|
@@ -764,16 +776,19 @@ never loads or downloads a model. `load()` forwards provider progress into the
|
|
|
764
776
|
sticky role record; `unload()` and `dispose()` abort owned work, await exposed
|
|
765
777
|
settlement, and verify provider status before publishing terminal state.
|
|
766
778
|
|
|
767
|
-
`status('
|
|
768
|
-
adds its optional `execution` snapshot to a
|
|
779
|
+
`status('stt', {execution:true})` or `status('tts', {execution:true})` explicitly
|
|
780
|
+
reads the selected provider and adds its optional `execution` snapshot to a
|
|
781
|
+
copy of the role record.
|
|
769
782
|
`status(null, {execution:true})` provides the equivalent projection under
|
|
770
783
|
`roles.llm`, `roles.stt`, and `roles.tts`. Providers that do not supply execution
|
|
771
784
|
omit that field. No provider load or sticky-state event is triggered; default
|
|
772
785
|
`status()` keeps its existing identity and behavior. A provider inspection
|
|
773
|
-
error propagates. Kokoro
|
|
786
|
+
error propagates. Each Whisper or Kokoro execution report contains `requestedDevice`,
|
|
774
787
|
`selectedDevice` (`null` while unloaded), `maxConcurrentRequests`, and
|
|
775
|
-
`activeRequestCount
|
|
776
|
-
|
|
788
|
+
`activeRequestCount`. The selected device names the backend requested by a
|
|
789
|
+
successful upstream session load. It does not prove that every operation ran
|
|
790
|
+
on a physical NPU or GPU, or that accelerator kernels overlap; WebNN may use
|
|
791
|
+
WASM for unsupported operations.
|
|
777
792
|
|
|
778
793
|
`validateSpeechConfiguration(value)` returns one mutable two-role selection
|
|
779
794
|
record without committing it, where `value` is the closed `{stt,tts}` record.
|
|
@@ -6437,8 +6437,10 @@ createBrowserWhisperProvider(options={})
|
|
|
6437
6437
|
|
|
6438
6438
|
The recognized options are
|
|
6439
6439
|
`{id='arcane-browser-whisper',localOnly=true,graph,model,runtime,appSecurity,
|
|
6440
|
-
security,store,offline=false}`.
|
|
6441
|
-
`
|
|
6440
|
+
security,store,offline=false,execution={device:'auto',maxConcurrentRequests:1}}`.
|
|
6441
|
+
`execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
|
|
6442
|
+
exactly 1. `graph` is mutually exclusive with `model` and `runtime`. The mutable
|
|
6443
|
+
result is
|
|
6442
6444
|
`{protocol:'arcane-ai-provider/2',role:'stt',id,localOnly:true,
|
|
6443
6445
|
maxConcurrentRequests:1,catalog,inspect,status,load,request,unload,dispose}`.
|
|
6444
6446
|
The only request operation is
|
|
@@ -6448,8 +6450,22 @@ re-freeze the cloned record.
|
|
|
6448
6450
|
|
|
6449
6451
|
`status()` returns
|
|
6450
6452
|
`{role,providerId,modelId,state,lifecycleStatus,lifecycleReason,activeOperation,
|
|
6451
|
-
loaded,busy,generation,errorCode,cache,warnings}` and includes
|
|
6452
|
-
for an explicit secure intent.
|
|
6453
|
+
loaded,busy,generation,errorCode,cache,warnings,execution}` and includes
|
|
6454
|
+
`security` only for an explicit secure intent.
|
|
6455
|
+
`execution` reports `requestedDevice`, `selectedDevice`,
|
|
6456
|
+
`maxConcurrentRequests`, and `activeRequestCount`; `selectedDevice` is `null`
|
|
6457
|
+
before load and after unload. The high-level projection is
|
|
6458
|
+
`ai.providerRuntime.status('stt', {execution:true}).execution`.
|
|
6459
|
+
|
|
6460
|
+
Automatic loading tries `webnn-npu` when `navigator.ml.createContext` is
|
|
6461
|
+
exposed, then `webgpu` when `navigator.gpu` is exposed, then CPU through `wasm`.
|
|
6462
|
+
A failed candidate is cleaned up before a fresh Worker tries the next backend
|
|
6463
|
+
with the same prepared model and dtype. Explicit selections do not fall back.
|
|
6464
|
+
`selectedDevice` names the backend requested by the successful upstream session
|
|
6465
|
+
load; it does not prove that every operation ran on a physical accelerator.
|
|
6466
|
+
WebNN may execute unsupported operations through WASM, and exact model,
|
|
6467
|
+
browser, driver, and hardware compatibility remains upstream.
|
|
6468
|
+
|
|
6453
6469
|
States are `unloaded`, `loading`, `ready`, `unloading`, `error`, and
|
|
6454
6470
|
`disposed`. Compatible concurrent loads coalesce; concurrent requests fail as
|
|
6455
6471
|
`ARCANE_AI_PROVIDER_BUSY`. Cancellation after the Worker request begins
|
|
@@ -6523,8 +6539,8 @@ createBrowserKokoroProvider(options={})
|
|
|
6523
6539
|
The recognized options are
|
|
6524
6540
|
`{id='arcane-browser-kokoro',localOnly=true,graph,model,runtime,appSecurity,
|
|
6525
6541
|
security,store,offline=false,execution={device:'auto',maxConcurrentRequests:4}}`.
|
|
6526
|
-
`execution.device` is `auto`, `webgpu`, or `wasm`; its capacity is
|
|
6527
|
-
from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
|
|
6542
|
+
`execution.device` is `auto`, `webnn-npu`, `webgpu`, or `wasm`; its capacity is
|
|
6543
|
+
an integer from 1 through 4. `graph` is mutually exclusive with `model` and `runtime`. The
|
|
6528
6544
|
mutable result is
|
|
6529
6545
|
`{protocol:'arcane-ai-provider/2',role:'tts',id,localOnly:true,
|
|
6530
6546
|
maxConcurrentRequests,catalog,inspect,status,load,request,unload,dispose}`. The
|
|
@@ -6543,10 +6559,11 @@ mono 16-bit PCM. Unsupported formats fail
|
|
|
6543
6559
|
`ARCANE_AI_UNSUPPORTED_RESPONSE_FORMAT`; malformed adapter audio fails
|
|
6544
6560
|
`ARCANE_AI_INVALID_PROVIDER_RESULT`. Unknown fields and accessors reject as malformed.
|
|
6545
6561
|
|
|
6546
|
-
Automatic execution
|
|
6547
|
-
|
|
6548
|
-
|
|
6549
|
-
|
|
6562
|
+
Automatic execution tries a complete WebNN NPU pool when
|
|
6563
|
+
`navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is
|
|
6564
|
+
exposed, then CPU through WASM. A failed candidate is cleaned up before a fresh
|
|
6565
|
+
pool tries the next backend. Explicit device selections do not fall back. Every
|
|
6566
|
+
slot loads the same caller-selected model and dtype in a distinct Worker so the
|
|
6550
6567
|
selected adapter's per-isolate inference serialization does not serialize the
|
|
6551
6568
|
pool. Direct provider `status().execution` reports `requestedDevice`,
|
|
6552
6569
|
`selectedDevice`, `maxConcurrentRequests`, and `activeRequestCount`.
|
|
@@ -6555,7 +6572,11 @@ by `AI.configureBrowserSpeech()`, use
|
|
|
6555
6572
|
`ai.providerRuntime.status('tts', {execution:true}).execution`. Requested
|
|
6556
6573
|
`auto` with selected `wasm` means automatic fallback occurred. Inspection
|
|
6557
6574
|
errors propagate; the default runtime `status()` remains a sticky lifecycle
|
|
6558
|
-
read.
|
|
6575
|
+
read. `selectedDevice` names the backend requested by the successful upstream
|
|
6576
|
+
session load. These fields do not prove that every operation ran on a physical
|
|
6577
|
+
NPU or GPU, or that accelerator kernels overlap. WebNN may execute unsupported
|
|
6578
|
+
operations through WASM; exact model, browser, driver, and hardware
|
|
6579
|
+
compatibility remains upstream.
|
|
6559
6580
|
|
|
6560
6581
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later
|
|
6561
6582
|
wait in the SDK's provider-neutral FIFO queue; they are not dropped. Synthesis
|
|
@@ -6,13 +6,15 @@ For the smallest first request, start with the [browser speech quick start](../.
|
|
|
6
6
|
|
|
7
7
|
## Speech defaults and inspection
|
|
8
8
|
|
|
9
|
-
The demo
|
|
9
|
+
The demo omits both speech execution records in `speechConfiguration(dbopfs)` so it consumes the SDK defaults: `{device:'auto',maxConcurrentRequests:1}` for Whisper STT and `{device:'auto',maxConcurrentRequests:4}` for Kokoro TTS. Auto tries WebNN NPU when `navigator.ml.createContext` is exposed, then WebGPU when `navigator.gpu` is exposed, then CPU through WASM. Failed loads are cleaned up before fresh Workers try the next device with the same prepared model and dtype. Explicit `webnn-npu`, `webgpu`, and `wasm` selections do not fall back.
|
|
10
10
|
|
|
11
|
-
|
|
11
|
+
Change the application's TTS record to `execution:{device:'wasm',maxConcurrentRequests:1}` to use fewer synthesis sessions. TTS accepts capacities 1 through 4; Whisper accepts only 1 and the LLM retains capacity one.
|
|
12
|
+
|
|
13
|
+
The maintained Kokoro selection uses `fp32` because [Kokoro.js recommends `fp32` when using WebGPU](https://github.com/hexgrad/kokoro/tree/main/kokoro.js#usage). Automatic fallback keeps that dtype on every device. `selectedDevice` identifies the backend requested by the successful upstream session load. It does not prove every operation ran on a physical NPU or GPU; WebNN may use WASM for unsupported operations. Exact model, browser, driver, and hardware compatibility and actual speech quality require evaluation on that combination.
|
|
12
14
|
|
|
13
15
|
Capacity 4 means up to four segments synthesize at once. Segment 5 and later wait in the SDK's FIFO queue; they are not dropped. Synthesis may finish out of order, but playback waits for earlier segments and plays exact input order. Each slot owns a Worker/model session, so raising capacity trades memory for latency.
|
|
14
16
|
|
|
15
|
-
Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected
|
|
17
|
+
Use the shared Speech component to load/unmute speech, then select **Inspect speech execution** below Chat. The button reads `ai.providerRuntime.status('tts',{execution:true}).execution` and displays requested device, selected device, capacity, active requests, and automatic WASM fallback. It is a snapshot at the moment clicked. `selectedDevice:null` means no pool is selected. The corresponding Whisper report is available through `ai.providerRuntime.status('stt',{execution:true}).execution`; reading either report does not load a model.
|
|
16
18
|
|
|
17
19
|
Shared Speech still owns mute, stop, and voice controls. The example's status button inspects the public report without loading a model or reaching into private providers.
|
|
18
20
|
|
|
@@ -434,7 +434,7 @@ function speechConfiguration(dbopfs) {
|
|
|
434
434
|
offline: false,
|
|
435
435
|
},
|
|
436
436
|
tts: {
|
|
437
|
-
// Omit execution
|
|
437
|
+
// Omit execution for NPU, GPU, then CPU selection and four synthesis slots.
|
|
438
438
|
// Kokoro.js recommends fp32 for the WebGPU route attempted by auto.
|
|
439
439
|
providerId: "wasm-ai-demo-browser-kokoro",
|
|
440
440
|
model: {
|
package/package.json
CHANGED