arcane-os 0.5.18 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +27 -3
- package/README.md +18 -12
- package/browser-runtime/ai/browser-speech-providers.mjs +64 -60
- package/browser-runtime/ai/browser-wasm-llm-provider.mjs +31 -11
- package/browser-runtime/ai/browser-wllama-runtime.mjs +78 -0
- package/browser-runtime/ai/speech-worker-runtime.mjs +4 -5
- package/docs/architecture.md +1 -1
- package/docs/publishing.md +10 -4
- package/docs/reference/ai/browser-speech.md +65 -38
- package/docs/reference/ai/browser-wasm.md +29 -6
- package/docs/reference/behavioral-testing.md +1 -1
- package/docs/reference/inventory/runtime-modules.json +2 -2
- package/docs/reference/runtime-modules.md +32 -17
- package/docs/reference/sdk-api.md +32 -11
- package/examples/wasm-ai-demo/README.md +5 -3
- package/examples/wasm-ai-demo/app.js +1 -1
- package/package.json +1 -1
- package/runtime/arcane/components/chat.html +53 -10
- package/runtime/arcane/modules/AI.js +27 -29
package/CHANGELOG.md
CHANGED
|
@@ -1,6 +1,32 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
-
##
|
|
3
|
+
## 0.6.0
|
|
4
|
+
|
|
5
|
+
- Prefer WebNN NPU, then WebGPU, then CPU through WASM for browser Whisper
|
|
6
|
+
transcription and Kokoro synthesis. Skip absent accelerator APIs and replace
|
|
7
|
+
failed model-load Workers before trying the next backend with the same
|
|
8
|
+
application-selected model and precision.
|
|
9
|
+
- Add Whisper execution configuration and NPU selection for both speech roles.
|
|
10
|
+
Explicit `webnn-npu`, `webgpu`, and `wasm` choices report failure without
|
|
11
|
+
falling back. Whisper retains one slot; Kokoro retains four by default and
|
|
12
|
+
accepts capacities from one through four.
|
|
13
|
+
- Report requested and successfully selected backends for both roles through
|
|
14
|
+
provider execution status. Selection reports upstream session loading;
|
|
15
|
+
actual accelerator use and speech quality depend on the selected runtime,
|
|
16
|
+
model, browser, drivers, and hardware. Wllama remains on its WebGPU backend.
|
|
17
|
+
|
|
18
|
+
## 0.5.19
|
|
19
|
+
|
|
20
|
+
- Report observed Wllama initialization stages and runtime activity through
|
|
21
|
+
direct provider loading and the shared AI lifecycle. Keep initialization
|
|
22
|
+
indeterminate when the runtime supplies no meaningful completion total.
|
|
23
|
+
- Show the current initialization stage, stage duration, and time since runtime
|
|
24
|
+
activity in shared chat instead of a completed file count during activation.
|
|
25
|
+
Preserve download progress, cancellation, and complete runtime logging.
|
|
26
|
+
- Remove an adjacent duplicate cancellation check while retaining cancellation
|
|
27
|
+
handling before initialization and after loading.
|
|
28
|
+
|
|
29
|
+
## 0.5.18
|
|
4
30
|
|
|
5
31
|
- Use `strong-type` predicates throughout SDK-owned toolchain, runtime,
|
|
6
32
|
browser-provider, and component code while retaining existing defaults,
|
|
@@ -12,8 +38,6 @@
|
|
|
12
38
|
- Remove `AIResponseLength` exports; applications own response verbosity.
|
|
13
39
|
Parse URL-audit HTML with the native HTML parser.
|
|
14
40
|
|
|
15
|
-
## 0.5.18
|
|
16
|
-
|
|
17
41
|
- Add `AI.prepareTTS()` for detached punctuation-segmented synthesis, optional
|
|
18
42
|
DBOPFS audio storage, complete semantic-input reuse, shared pending work and
|
|
19
43
|
independent preparation cancellation. Persist MIME metadata with each audio
|
package/README.md
CHANGED
|
@@ -19,7 +19,7 @@ version-locked SDK runtime, while an integrated Arcane checkout uses its live
|
|
|
19
19
|
`arcane/` runtime. Both profiles preserve the same app URLs, theme, packaging,
|
|
20
20
|
event, cancellation, and browser run contracts.
|
|
21
21
|
|
|
22
|
-
This checkout defines the `0.
|
|
22
|
+
This checkout defines the `0.6.0` SDK contract. Applications pin one exact npm
|
|
23
23
|
version and lockfile; registry state is deliberately not baked into application
|
|
24
24
|
artifacts.
|
|
25
25
|
|
|
@@ -35,7 +35,7 @@ Create one browser application, install its pinned SDK, and start its source
|
|
|
35
35
|
server:
|
|
36
36
|
|
|
37
37
|
```bash
|
|
38
|
-
npx arcane-os@0.
|
|
38
|
+
npx arcane-os@0.6.0 new hello-speech --path ./hello-speech --target browser
|
|
39
39
|
cd hello-speech
|
|
40
40
|
npm install
|
|
41
41
|
npm run dev
|
|
@@ -299,10 +299,16 @@ provider factories. The package contains the plain-JavaScript provider and
|
|
|
299
299
|
Worker machinery, not speech runtimes, models, voices, or a CDN default. An app
|
|
300
300
|
must supply each runtime/model selection explicitly. Speech roles
|
|
301
301
|
load, cancel, unload, fail, and recover independently, so speech failure never
|
|
302
|
-
|
|
303
|
-
|
|
304
|
-
|
|
305
|
-
|
|
302
|
+
switches to a cloud provider or prevents text chat. Both roles default to
|
|
303
|
+
NPU, GPU, then CPU automatic selection: WebNN NPU when its browser API is
|
|
304
|
+
exposed, then WebGPU, then WASM. Failed model loads release their Workers before
|
|
305
|
+
the next backend uses the same model and precision. Apps may select
|
|
306
|
+
`webnn-npu`, `webgpu`, or `wasm` explicitly without automatic fallback.
|
|
307
|
+
Whisper keeps one transcription slot; Kokoro defaults to four Worker/model
|
|
308
|
+
sessions and accepts capacities from one through four. Both roles expose
|
|
309
|
+
requested and selected devices through execution status. A selected backend
|
|
310
|
+
reports successful upstream loading, not physical accelerator use for every
|
|
311
|
+
operation; model and hardware compatibility remain upstream.
|
|
306
312
|
|
|
307
313
|
Applications that need faster spoken-response onset can configure the shared
|
|
308
314
|
TTS stream without taking over synthesis or playback:
|
|
@@ -357,7 +363,7 @@ uses the same controller for automatic memory extraction.
|
|
|
357
363
|
Create a new repository-shaped Arcane application with the exact stable SDK:
|
|
358
364
|
|
|
359
365
|
```bash
|
|
360
|
-
npx arcane-os@0.
|
|
366
|
+
npx arcane-os@0.6.0 new my-app --path ./my-app --target portable --git
|
|
361
367
|
cd my-app
|
|
362
368
|
npm install
|
|
363
369
|
npm run dev
|
|
@@ -367,7 +373,7 @@ To enroll an existing repository, install the exact SDK and initialize only
|
|
|
367
373
|
missing Arcane files:
|
|
368
374
|
|
|
369
375
|
```bash
|
|
370
|
-
npm install --save-dev --save-exact arcane-os@0.
|
|
376
|
+
npm install --save-dev --save-exact arcane-os@0.6.0
|
|
371
377
|
npm exec -- arcane init my-app --target portable
|
|
372
378
|
```
|
|
373
379
|
|
|
@@ -383,7 +389,7 @@ npm exec -- arcane-os targets
|
|
|
383
389
|
No global SDK install or standalone Arcane CLI is required. The application
|
|
384
390
|
repository's exact npm dependency and lockfile own the CLI and toolchain version.
|
|
385
391
|
|
|
386
|
-
Use `npx arcane-os@0.
|
|
392
|
+
Use `npx arcane-os@0.6.0` for the initial bootstrap because it names this npm
|
|
387
393
|
package explicitly; bare `npx arcane` outside an installed project could resolve
|
|
388
394
|
a different package. Both installed commands invoke the same headless toolchain.
|
|
389
395
|
Project-local npm scripts use the SDK pinned by that app's `package-lock.json`,
|
|
@@ -403,7 +409,7 @@ node ./bin/arcane.mjs new local-app --path ../local-app --target portable --git
|
|
|
403
409
|
|
|
404
410
|
# From the generated app repository
|
|
405
411
|
cd ../local-app
|
|
406
|
-
npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.
|
|
412
|
+
npm install --save-dev --save-exact ../arcane-os-sdk/arcane-os-0.6.0.tgz
|
|
407
413
|
npm ci
|
|
408
414
|
```
|
|
409
415
|
|
|
@@ -412,7 +418,7 @@ same location. The lockfile retains the selected package dependency while
|
|
|
412
418
|
Arcane uses the installed package name and version. Local directory `file:` dependencies are not
|
|
413
419
|
accepted because npm may install them as links; use a packed `.tgz`. A GitHub
|
|
414
420
|
runner also needs that tarball at the locked path. After publication, replace
|
|
415
|
-
the local declaration with the exact `arcane-os@0.
|
|
421
|
+
the local declaration with the exact `arcane-os@0.6.0` registry package and
|
|
416
422
|
commit the regenerated lock.
|
|
417
423
|
|
|
418
424
|
Generated repositories use `npm ci --ignore-scripts` in CI. Run dependency
|
|
@@ -551,7 +557,7 @@ package installation, or assertions.
|
|
|
551
557
|
|
|
552
558
|
## Current target support
|
|
553
559
|
|
|
554
|
-
Version `0.
|
|
560
|
+
Version `0.6.0` exposes one browser target and five explicitly paired
|
|
555
561
|
native development targets: a non-runnable portable directory, a
|
|
556
562
|
Windows x64 unsigned-local-test EXE bundle, Linux x64 and Linux ARM64
|
|
557
563
|
unsigned-local-test DEBs, and an Android development-signed APK. The
|
|
@@ -23,8 +23,8 @@ const ROLE_OPERATION = completeValue({ stt: "transcribe", tts: "synthesize" });
|
|
|
23
23
|
const STT_SAMPLE_RATE = 16_000;
|
|
24
24
|
const TTS_SAMPLE_RATE = 24_000;
|
|
25
25
|
const TTS_RESPONSE_FORMAT = "wav";
|
|
26
|
-
const
|
|
27
|
-
const
|
|
26
|
+
const SPEECH_EXECUTION_DEVICES = new Set(["auto", "webnn-npu", "webgpu", "wasm"]);
|
|
27
|
+
const DEFAULT_SPEECH_EXECUTION_DEVICE = "auto";
|
|
28
28
|
const DEFAULT_TTS_MAX_CONCURRENT_REQUESTS = 4;
|
|
29
29
|
const MAX_TTS_CONCURRENT_REQUESTS = 4;
|
|
30
30
|
const ROLE_REQUEST_REASON = completeValue({
|
|
@@ -150,24 +150,21 @@ function requiredIdentifier(value, label) {
|
|
|
150
150
|
}
|
|
151
151
|
|
|
152
152
|
function normalizeSpeechExecution(role, execution) {
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
156
|
-
}
|
|
157
|
-
return completeValue({ device: "wasm", maxConcurrentRequests: 1 });
|
|
158
|
-
}
|
|
153
|
+
const label = role === "stt" ? "Browser Whisper" : "Browser Kokoro";
|
|
154
|
+
const defaultConcurrency = role === "stt" ? 1 : DEFAULT_TTS_MAX_CONCURRENT_REQUESTS;
|
|
155
|
+
const maximumConcurrency = role === "stt" ? 1 : MAX_TTS_CONCURRENT_REQUESTS;
|
|
159
156
|
if (execution === undefined) {
|
|
160
157
|
return completeValue({
|
|
161
|
-
device:
|
|
162
|
-
maxConcurrentRequests:
|
|
158
|
+
device: DEFAULT_SPEECH_EXECUTION_DEVICE,
|
|
159
|
+
maxConcurrentRequests: defaultConcurrency,
|
|
163
160
|
});
|
|
164
161
|
}
|
|
165
162
|
if (!execution || !is.object(execution) || is.array(execution)) {
|
|
166
|
-
throw new TypeError(
|
|
163
|
+
throw new TypeError(`${label} execution must be a plain data record.`);
|
|
167
164
|
}
|
|
168
165
|
const prototype = Object.getPrototypeOf(execution);
|
|
169
166
|
if (prototype !== Object.prototype && prototype !== null) {
|
|
170
|
-
throw new TypeError(
|
|
167
|
+
throw new TypeError(`${label} execution must be a plain data record.`);
|
|
171
168
|
}
|
|
172
169
|
const descriptors = Object.getOwnPropertyDescriptors(execution);
|
|
173
170
|
for (const key of Reflect.ownKeys(descriptors)) {
|
|
@@ -175,34 +172,35 @@ function normalizeSpeechExecution(role, execution) {
|
|
|
175
172
|
(key !== "device" && key !== "maxConcurrentRequests")
|
|
176
173
|
|| !Object.hasOwn(descriptors[key], "value")
|
|
177
174
|
) {
|
|
178
|
-
throw new TypeError(
|
|
175
|
+
throw new TypeError(`${label} execution contains an unsupported or accessor field.`);
|
|
179
176
|
}
|
|
180
177
|
}
|
|
181
178
|
const device = Object.hasOwn(descriptors, "device")
|
|
182
179
|
? descriptors.device.value
|
|
183
|
-
:
|
|
180
|
+
: DEFAULT_SPEECH_EXECUTION_DEVICE;
|
|
184
181
|
const maxConcurrentRequests = Object.hasOwn(descriptors, "maxConcurrentRequests")
|
|
185
182
|
? descriptors.maxConcurrentRequests.value
|
|
186
|
-
:
|
|
187
|
-
if (!
|
|
188
|
-
throw new TypeError(
|
|
183
|
+
: defaultConcurrency;
|
|
184
|
+
if (!SPEECH_EXECUTION_DEVICES.has(device)) {
|
|
185
|
+
throw new TypeError(`${label} execution.device must be "auto", "webnn-npu", "webgpu", or "wasm".`);
|
|
189
186
|
}
|
|
190
187
|
if (
|
|
191
188
|
!is.safeInteger(maxConcurrentRequests)
|
|
192
189
|
|| maxConcurrentRequests < 1
|
|
193
|
-
|| maxConcurrentRequests >
|
|
190
|
+
|| maxConcurrentRequests > maximumConcurrency
|
|
194
191
|
) {
|
|
195
|
-
throw new TypeError(
|
|
192
|
+
throw new TypeError(`${label} execution.maxConcurrentRequests must be a safe integer from 1 through ${maximumConcurrency}.`);
|
|
196
193
|
}
|
|
197
194
|
return completeValue({ device, maxConcurrentRequests });
|
|
198
195
|
}
|
|
199
196
|
|
|
200
|
-
function
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
204
|
-
|
|
205
|
-
|
|
197
|
+
function speechExecutionDevices(requestedDevice) {
|
|
198
|
+
if (requestedDevice !== "auto") return [requestedDevice];
|
|
199
|
+
const devices = [];
|
|
200
|
+
if (is.function(globalThis.navigator?.ml?.createContext)) devices.push("webnn-npu");
|
|
201
|
+
if (globalThis.navigator?.gpu) devices.push("webgpu");
|
|
202
|
+
devices.push("wasm");
|
|
203
|
+
return devices;
|
|
206
204
|
}
|
|
207
205
|
|
|
208
206
|
function createProviderAuthority({
|
|
@@ -1187,14 +1185,12 @@ function createBrowserSpeechProvider({
|
|
|
1187
1185
|
cache,
|
|
1188
1186
|
lifecycleReason,
|
|
1189
1187
|
activeOperation,
|
|
1190
|
-
execution:
|
|
1191
|
-
|
|
1192
|
-
|
|
1193
|
-
|
|
1194
|
-
|
|
1195
|
-
|
|
1196
|
-
})
|
|
1197
|
-
: null,
|
|
1188
|
+
execution: completeValue({
|
|
1189
|
+
requestedDevice: speechExecution.device,
|
|
1190
|
+
selectedDevice,
|
|
1191
|
+
maxConcurrentRequests: speechExecution.maxConcurrentRequests,
|
|
1192
|
+
activeRequestCount: requestOperations.size,
|
|
1193
|
+
}),
|
|
1198
1194
|
secureIntent,
|
|
1199
1195
|
warnings: providerWarnings(runtimeWarnings),
|
|
1200
1196
|
});
|
|
@@ -1535,34 +1531,42 @@ function createBrowserSpeechProvider({
|
|
|
1535
1531
|
`${role}-load-superseded-by-unload`,
|
|
1536
1532
|
);
|
|
1537
1533
|
}
|
|
1538
|
-
const
|
|
1539
|
-
|
|
1540
|
-
|
|
1541
|
-
|
|
1542
|
-
|
|
1543
|
-
|
|
1544
|
-
|
|
1545
|
-
|
|
1546
|
-
|
|
1547
|
-
|
|
1548
|
-
|
|
1549
|
-
|
|
1550
|
-
|
|
1551
|
-
|
|
1552
|
-
|
|
1553
|
-
|
|
1554
|
-
|
|
1555
|
-
|
|
1556
|
-
|
|
1557
|
-
) {
|
|
1558
|
-
|
|
1534
|
+
const devices = speechExecutionDevices(speechExecution.device);
|
|
1535
|
+
for (const [index, device] of devices.entries()) {
|
|
1536
|
+
throwIfAborted(linked.controller.signal, `${role}-load-cancelled`);
|
|
1537
|
+
if (operationGeneration !== generation) {
|
|
1538
|
+
throw providerError(
|
|
1539
|
+
"ARCANE_AI_OPERATION_SUPERSEDED",
|
|
1540
|
+
"Browser speech loading was superseded.",
|
|
1541
|
+
undefined,
|
|
1542
|
+
`${role}-load-superseded-by-unload`,
|
|
1543
|
+
);
|
|
1544
|
+
}
|
|
1545
|
+
try {
|
|
1546
|
+
pool = await loadWorkerPool(
|
|
1547
|
+
preparation,
|
|
1548
|
+
device,
|
|
1549
|
+
record.warnings,
|
|
1550
|
+
linked.controller.signal,
|
|
1551
|
+
);
|
|
1552
|
+
break;
|
|
1553
|
+
} catch (error) {
|
|
1554
|
+
if (
|
|
1555
|
+
index === devices.length - 1
|
|
1556
|
+
|| linked.controller.signal.aborted
|
|
1557
|
+
|| operationGeneration !== generation
|
|
1558
|
+
|| error?.code === "ARCANE_AI_REQUEST_ABORTED"
|
|
1559
|
+
|| error?.code === "ARCANE_AI_OPERATION_SUPERSEDED"
|
|
1560
|
+
) {
|
|
1561
|
+
throw error;
|
|
1562
|
+
}
|
|
1563
|
+
// A fresh Worker releases a failed upstream initialization before
|
|
1564
|
+
// the next backend reuses the same prepared model and precision.
|
|
1565
|
+
const warning = `Browser ${role} ${device} loading failed; trying ${devices[index + 1]}.`;
|
|
1566
|
+
record.warnings = [...record.warnings, warning];
|
|
1567
|
+
lastWarnings = record.warnings;
|
|
1568
|
+
console.warn(warning, error);
|
|
1559
1569
|
}
|
|
1560
|
-
pool = await loadWorkerPool(
|
|
1561
|
-
preparation,
|
|
1562
|
-
"wasm",
|
|
1563
|
-
record.warnings,
|
|
1564
|
-
linked.controller.signal,
|
|
1565
|
-
);
|
|
1566
1570
|
}
|
|
1567
1571
|
throwIfAborted(
|
|
1568
1572
|
linked.controller.signal,
|
|
@@ -3000,24 +3000,47 @@ export function createBrowserWasmLlmProvider({
|
|
|
3000
3000
|
? context.reportProgress
|
|
3001
3001
|
: options.onProgress ?? null;
|
|
3002
3002
|
const progressStartedAt = Date.now();
|
|
3003
|
+
let progressStageStartedAt = progressStartedAt;
|
|
3004
|
+
let progressUpdatedAt = progressStartedAt;
|
|
3003
3005
|
let currentProgress = null;
|
|
3004
3006
|
let progressHeartbeat = null;
|
|
3005
3007
|
|
|
3006
3008
|
function publishModelLoadProgress(progress) {
|
|
3007
|
-
if (!reportProgress) return;
|
|
3009
|
+
if (!reportProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
|
|
3010
|
+
const now = Date.now();
|
|
3011
|
+
if (currentProgress?.phase !== progress.phase || currentProgress?.stage !== progress.stage) {
|
|
3012
|
+
progressStageStartedAt = now;
|
|
3013
|
+
}
|
|
3014
|
+
progressUpdatedAt = now;
|
|
3008
3015
|
currentProgress = { ...progress, heartbeat: false };
|
|
3009
3016
|
reportProgress(completeValue({
|
|
3010
3017
|
...currentProgress,
|
|
3011
|
-
elapsedMs: Math.max(0,
|
|
3018
|
+
elapsedMs: Math.max(0, now - progressStartedAt),
|
|
3019
|
+
...(progress.phase === "initialize" ? {
|
|
3020
|
+
phaseElapsedMs: Math.max(0, now - progressStageStartedAt),
|
|
3021
|
+
activityElapsedMs: 0,
|
|
3022
|
+
} : {}),
|
|
3012
3023
|
}));
|
|
3013
3024
|
}
|
|
3014
3025
|
|
|
3026
|
+
function publishModelInitializationProgress(progress) {
|
|
3027
|
+
if (!reportProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
|
|
3028
|
+
progressUpdatedAt = Date.now();
|
|
3029
|
+
if (!progress || (currentProgress?.stage === progress.stage && currentProgress?.message === progress.message)) return;
|
|
3030
|
+
publishModelLoadProgress(progress);
|
|
3031
|
+
}
|
|
3032
|
+
|
|
3015
3033
|
function publishModelLoadHeartbeat() {
|
|
3016
|
-
if (!reportProgress || !currentProgress) return;
|
|
3034
|
+
if (!reportProgress || !currentProgress || signal.aborted || generation !== lifecycleGeneration || state !== "loading") return;
|
|
3035
|
+
const now = Date.now();
|
|
3017
3036
|
reportProgress(completeValue({
|
|
3018
3037
|
...currentProgress,
|
|
3019
3038
|
heartbeat: true,
|
|
3020
|
-
elapsedMs: Math.max(0,
|
|
3039
|
+
elapsedMs: Math.max(0, now - progressStartedAt),
|
|
3040
|
+
...(currentProgress.phase === "initialize" ? {
|
|
3041
|
+
phaseElapsedMs: Math.max(0, now - progressStageStartedAt),
|
|
3042
|
+
activityElapsedMs: Math.max(0, now - progressUpdatedAt),
|
|
3043
|
+
} : {}),
|
|
3021
3044
|
}));
|
|
3022
3045
|
}
|
|
3023
3046
|
|
|
@@ -3047,16 +3070,12 @@ export function createBrowserWasmLlmProvider({
|
|
|
3047
3070
|
if (generation !== lifecycleGeneration || state !== "loading") {
|
|
3048
3071
|
throw fail("ARCANE_AI_OPERATION_SUPERSEDED", "The model load was superseded by unload.");
|
|
3049
3072
|
}
|
|
3050
|
-
throwIfAborted(signal, "load");
|
|
3051
|
-
if (generation !== lifecycleGeneration || state !== "loading") {
|
|
3052
|
-
throw fail("ARCANE_AI_OPERATION_SUPERSEDED", "The model load was superseded by unload.");
|
|
3053
|
-
}
|
|
3054
3073
|
const members = sourceMetadata(activeSource).files;
|
|
3055
3074
|
publishModelLoadProgress({
|
|
3056
3075
|
phase: "initialize",
|
|
3057
|
-
|
|
3058
|
-
|
|
3059
|
-
|
|
3076
|
+
stage: "runtime",
|
|
3077
|
+
message: "Starting the WebAssembly runtime and opening model files",
|
|
3078
|
+
total: null,
|
|
3060
3079
|
heartbeat: false,
|
|
3061
3080
|
});
|
|
3062
3081
|
const modelFiles = admitted.files.map((file, index) => (
|
|
@@ -3074,6 +3093,7 @@ export function createBrowserWasmLlmProvider({
|
|
|
3074
3093
|
...runtimeOptions,
|
|
3075
3094
|
...activeLoadPlan,
|
|
3076
3095
|
signal,
|
|
3096
|
+
onProgress: publishModelInitializationProgress,
|
|
3077
3097
|
});
|
|
3078
3098
|
if (!runtime.isLoaded()) {
|
|
3079
3099
|
throw fail(
|
|
@@ -94,6 +94,59 @@ function createEvidenceLogger(logger) {
|
|
|
94
94
|
let offload = null;
|
|
95
95
|
let invalid = false;
|
|
96
96
|
let completionCapture = null;
|
|
97
|
+
let loadProgress = null;
|
|
98
|
+
let tensorCount = null;
|
|
99
|
+
|
|
100
|
+
function observeLoadLine(line) {
|
|
101
|
+
if (!loadProgress) return;
|
|
102
|
+
let stage;
|
|
103
|
+
let message;
|
|
104
|
+
const metadata = line.match(/^llama_model_loader: loaded meta data with \d+ key-value pairs and (\d+) tensors from /u);
|
|
105
|
+
const layers = line.match(GPU_OFFLOAD_PATTERN);
|
|
106
|
+
if (line.startsWith('Loading "wllama.wasm" from ')) {
|
|
107
|
+
stage = "runtime";
|
|
108
|
+
message = "Loading the WebAssembly runtime";
|
|
109
|
+
} else if (line === "Calling wllamaStart...") {
|
|
110
|
+
stage = "backend";
|
|
111
|
+
message = "Starting the inference engine";
|
|
112
|
+
} else if (line === "Loading model...") {
|
|
113
|
+
stage = "metadata";
|
|
114
|
+
message = "Reading model metadata";
|
|
115
|
+
} else if (WEBGPU_ADAPTER_PATTERN.test(line)) {
|
|
116
|
+
stage = "gpu";
|
|
117
|
+
message = "Graphics device initialized; preparing the model";
|
|
118
|
+
} else if (metadata) {
|
|
119
|
+
tensorCount = Number(metadata[1]);
|
|
120
|
+
stage = "metadata";
|
|
121
|
+
message = `Model metadata read: ${tensorCount} tensors`;
|
|
122
|
+
} else if (line.startsWith("load_tensors: loading model tensors,")) {
|
|
123
|
+
stage = "weights";
|
|
124
|
+
message = tensorCount === null
|
|
125
|
+
? "Loading model weights"
|
|
126
|
+
: `Loading model weights for ${tensorCount} tensors`;
|
|
127
|
+
} else if (layers) {
|
|
128
|
+
// Upstream reports layer assignment before the weight reads finish.
|
|
129
|
+
stage = "weights";
|
|
130
|
+
message = `Loading model weights; ${layers[1]} of ${layers[2]} layers assigned to the GPU`;
|
|
131
|
+
} else if (line === "llama_context: constructing llama_context") {
|
|
132
|
+
stage = "context";
|
|
133
|
+
message = "Preparing the inference context";
|
|
134
|
+
} else if (/^[^:]+:\s+graph (?:nodes|splits)\s+=\s+\d+/u.test(line)) {
|
|
135
|
+
stage = "graph";
|
|
136
|
+
message = "Preparing the inference graph";
|
|
137
|
+
} else if (/^cmn\s+common_init_:\s+warming up the model with an empty run\b/u.test(line)) {
|
|
138
|
+
stage = "warmup";
|
|
139
|
+
message = "Warming up the model with an empty run";
|
|
140
|
+
}
|
|
141
|
+
if (stage) {
|
|
142
|
+
loadProgress({
|
|
143
|
+
phase: "initialize",
|
|
144
|
+
stage,
|
|
145
|
+
message,
|
|
146
|
+
total: null,
|
|
147
|
+
});
|
|
148
|
+
}
|
|
149
|
+
}
|
|
97
150
|
|
|
98
151
|
function same(left, right) {
|
|
99
152
|
return JSON.stringify(left) === JSON.stringify(right);
|
|
@@ -119,6 +172,7 @@ function createEvidenceLogger(logger) {
|
|
|
119
172
|
observeCompletionLine(level, value);
|
|
120
173
|
const line = String(value).trim();
|
|
121
174
|
if (!line) return;
|
|
175
|
+
observeLoadLine(line);
|
|
122
176
|
const adapterMatch = line.match(WEBGPU_ADAPTER_PATTERN);
|
|
123
177
|
if (adapterMatch) {
|
|
124
178
|
const next = completeValue({
|
|
@@ -144,6 +198,8 @@ function createEvidenceLogger(logger) {
|
|
|
144
198
|
}
|
|
145
199
|
|
|
146
200
|
function observe(level, args) {
|
|
201
|
+
// Activity without a new stage updates its age without repainting the UI.
|
|
202
|
+
loadProgress?.(null);
|
|
147
203
|
for (const value of args) {
|
|
148
204
|
if (!is.string(value)) continue;
|
|
149
205
|
for (const line of value.split(/\r?\n/u)) observeLine(level, line);
|
|
@@ -160,6 +216,13 @@ function createEvidenceLogger(logger) {
|
|
|
160
216
|
|
|
161
217
|
return completeValue({
|
|
162
218
|
logger: completeValue(wrapped),
|
|
219
|
+
beginLoadProgress(report) {
|
|
220
|
+
loadProgress = report;
|
|
221
|
+
tensorCount = null;
|
|
222
|
+
return function releaseLoadProgress() {
|
|
223
|
+
loadProgress = null;
|
|
224
|
+
};
|
|
225
|
+
},
|
|
163
226
|
beginCompletionCapture() {
|
|
164
227
|
if (completionCapture) {
|
|
165
228
|
throw runtimeFailure(
|
|
@@ -469,6 +532,18 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
|
|
|
469
532
|
cleanup: null,
|
|
470
533
|
});
|
|
471
534
|
const loadController = new AbortController();
|
|
535
|
+
let progressFailure = null;
|
|
536
|
+
const releaseLoadProgress = sessionObservers.get(next).beginLoadProgress(
|
|
537
|
+
function reportRuntimeLoadProgress(progress) {
|
|
538
|
+
if (loadController.signal.aborted || !is.function(options.onProgress)) return;
|
|
539
|
+
try {
|
|
540
|
+
options.onProgress(progress);
|
|
541
|
+
} catch (error) {
|
|
542
|
+
progressFailure = error;
|
|
543
|
+
loadController.abort(error);
|
|
544
|
+
}
|
|
545
|
+
},
|
|
546
|
+
);
|
|
472
547
|
const loadOperation = Promise.resolve().then(() => (
|
|
473
548
|
next.arcaneLoadModel(files, loadOptions, loadController.signal)
|
|
474
549
|
));
|
|
@@ -484,6 +559,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
|
|
|
484
559
|
else signal?.addEventListener?.("abort", onAbort, { once: true });
|
|
485
560
|
try {
|
|
486
561
|
await loadOperation;
|
|
562
|
+
if (progressFailure) throw progressFailure;
|
|
487
563
|
if (pending?.engine !== next) throw new Error("Wllama load was cancelled.");
|
|
488
564
|
if (!is.function(next.isModelLoaded) || next.isModelLoaded() !== true) {
|
|
489
565
|
throw runtimeFailure(
|
|
@@ -496,6 +572,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
|
|
|
496
572
|
engine = next;
|
|
497
573
|
publishEvidence({ state: "ready", webgpu, cancellation: null, cleanup: null });
|
|
498
574
|
} catch (error) {
|
|
575
|
+
releaseLoadProgress();
|
|
499
576
|
let cleanupFailure = null;
|
|
500
577
|
try {
|
|
501
578
|
await exitSession(next);
|
|
@@ -513,6 +590,7 @@ export function createPackagedWllamaRuntime({ logger = arcaneLogging } = {}) {
|
|
|
513
590
|
if (cleanupFailure) throw cleanupFailure;
|
|
514
591
|
throw error;
|
|
515
592
|
} finally {
|
|
593
|
+
releaseLoadProgress();
|
|
516
594
|
signal?.removeEventListener?.("abort", onAbort);
|
|
517
595
|
}
|
|
518
596
|
|
|
@@ -738,10 +738,9 @@ function validateConfiguration(configuration, role) {
|
|
|
738
738
|
if (
|
|
739
739
|
descriptorMismatch
|
|
740
740
|
|| !Object.hasOwn(descriptors, "device")
|
|
741
|
-
|| (
|
|
742
|
-
|
|
743
|
-
&& descriptors.device.value !== "
|
|
744
|
-
&& descriptors.device.value !== "webgpu")
|
|
741
|
+
|| (descriptors.device.value !== "wasm"
|
|
742
|
+
&& descriptors.device.value !== "webgpu"
|
|
743
|
+
&& descriptors.device.value !== "webnn-npu")
|
|
745
744
|
) {
|
|
746
745
|
throw workerError(
|
|
747
746
|
"ARCANE_AI_INVALID_REQUEST",
|
|
@@ -1367,7 +1366,7 @@ async function createWhisperEngine(namespace, configuration, signal, report) {
|
|
|
1367
1366
|
"automatic-speech-recognition",
|
|
1368
1367
|
configuration.model.repository,
|
|
1369
1368
|
{
|
|
1370
|
-
device: "wasm",
|
|
1369
|
+
device: configuration.execution?.device ?? "wasm",
|
|
1371
1370
|
dtype: configuration.model.dtype ?? "fp32",
|
|
1372
1371
|
revision: configuration.model.revision,
|
|
1373
1372
|
progress_callback: report,
|
package/docs/architecture.md
CHANGED
|
@@ -279,7 +279,7 @@ paths are withheld from the native provider. The provider copies the complete
|
|
|
279
279
|
selected release rather than accepting an unrelated source path. Verification
|
|
280
280
|
is a separate explicit operation for a selected release artifact.
|
|
281
281
|
|
|
282
|
-
The SDK `0.
|
|
282
|
+
The SDK `0.6.0` runtime requires Arcane `0.8.12` or newer. Compatibility
|
|
283
283
|
is contractual rather than exact-version pinning: the prepared Core must meet
|
|
284
284
|
the highest minimum declared by the runtime, selected app, and bundled app
|
|
285
285
|
dependencies; keep each app's Arcane protocol generation; and provide every
|
package/docs/publishing.md
CHANGED
|
@@ -16,7 +16,13 @@ dependency and invoke its local CLI with `npm exec -- arcane`. A separate
|
|
|
16
16
|
global installer, standalone SDK executable, NuGet package, Homebrew formula,
|
|
17
17
|
or OS package is not part of this release surface.
|
|
18
18
|
|
|
19
|
-
|
|
19
|
+
The user's standing instruction selects publication when a coherent SDK change
|
|
20
|
+
is complete. Publish every ready change that preserves required functionality
|
|
21
|
+
and remains relevant, while excluding unfinished concurrent or explicitly
|
|
22
|
+
deferred work. Default to a patch release and assess whether a new capability
|
|
23
|
+
warrants a minor revision. Honor an explicit no-publish instruction.
|
|
24
|
+
|
|
25
|
+
Publication checks run only for that selected npm release output.
|
|
20
26
|
That selected-release workflow validates package metadata, the executable and
|
|
21
27
|
`.gitattributes` boundary, the complete package inventory, version/channel
|
|
22
28
|
agreement, and required license notices. One unprivileged producer packs one
|
|
@@ -94,9 +100,9 @@ locked installation, runs one selected `arcane package`, creates one selected
|
|
|
94
100
|
identities, receipts, provenance records, or attestation sidecars and does not
|
|
95
101
|
run a second admission job.
|
|
96
102
|
|
|
97
|
-
Stable versioning, the npm `latest` tag, and an official GitHub release
|
|
98
|
-
|
|
99
|
-
|
|
103
|
+
Stable versioning, the npm `latest` tag, and an official GitHub release follow
|
|
104
|
+
the selected release decision, including the standing completed-work authority
|
|
105
|
+
above. Unfinished `main` development does not select a release. A stable release
|
|
100
106
|
must publish the exact selected Check artifact under `latest`; GitHub may then
|
|
101
107
|
attach that same package. Its Git
|
|
102
108
|
tag and GitHub release title must both be the same bare numeric
|