omnindicator 0.1.0 → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +24 -24
- package/npm/omnindicator/lib/installer.js +1 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -215,8 +215,9 @@ before TTS, stops the comprehension service to make room, streams decoder PCM
|
|
|
215
215
|
with a small startup lead, and starts restoring comprehension while audio is
|
|
216
216
|
still playing. If the user interrupts, playback ducks and then pauses/fades;
|
|
217
217
|
the microphone remains active throughout. This is a safe memory arrangement,
|
|
218
|
-
not the theoretical minimum-latency arrangement.
|
|
219
|
-
|
|
218
|
+
not the theoretical minimum-latency arrangement. The managed runtime now uses
|
|
219
|
+
the same safety invariant on every host: the TTS graph is loaded for one
|
|
220
|
+
utterance and shed when that response completes.
|
|
220
221
|
|
|
221
222
|
The essential constrained-host overrides are:
|
|
222
223
|
|
|
@@ -359,7 +360,7 @@ details hidden behind the adapter.
|
|
|
359
360
|
| Media comprehension | Qwen3-Omni `llama-server` | Prompt caching is disabled so stale multimodal embeddings cannot cross turns |
|
|
360
361
|
| Semantic bridge | Adapter-generated tagged observation | Speech transcript, non-speech acoustics, and visual evidence stay separate and remain untrusted data |
|
|
361
362
|
| Language, reasoning, tool choice | Selected Ollama base or configured OpenAI-compatible worker | Reasoning remains in `message.thinking`; unresolved tool calls cannot enter TTS |
|
|
362
|
-
| Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, and
|
|
363
|
+
| Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, blocks exact duplicate execution, and lets recovery continue until final/timeout/disconnect |
|
|
363
364
|
| Computer use | `portal/browser.py` + `portal/gui.py` | Opens visible Chromium on the desktop, observes rendered screenshots, clicks/types through native DevTools input, and can see/control the wider workspace through `xdotool` plus fresh desktop screenshots |
|
|
364
365
|
| Speech | Patched Qwen3-TTS worker | Emits ordered decoder PCM; generation state is reset between prompts and never leaks one utterance into the next |
|
|
365
366
|
| Local conversation | `harness/` | VAD, interruption, ReSpeaker state/direction, camera capture, history, deferred memory writes, and foreground scheduling |
|
|
@@ -588,10 +589,11 @@ dynamic inside the allocated KV window; new work is admission-gated, and the
|
|
|
588
589
|
next supervised start reselects its tier from live memory. This avoids process
|
|
589
590
|
churn without reverting to a board-specific context limit.
|
|
590
591
|
|
|
591
|
-
The trained-bridge runtime keeps TTS and
|
|
592
|
-
|
|
592
|
+
The trained-bridge runtime keeps comprehension resident while TTS and pointing
|
|
593
|
+
weights are request-scoped. Guided deployment does not block the desktop on
|
|
594
|
+
generation probes. It marks
|
|
593
595
|
the core ready after local component health and starts the indicator service as
|
|
594
|
-
soon as the core unit starts. The full ASR/cloned-TTS/
|
|
596
|
+
soon as the core unit starts. The full ASR/cloned-TTS/on-demand-residency smoke remains
|
|
595
597
|
available as an explicit diagnostic. Legacy/manual constrained profiles may
|
|
596
598
|
still opt into explicit eviction callbacks and non-persistent TTS. Service
|
|
597
599
|
managers restart failed workers; the harness waits
|
|
@@ -766,9 +768,9 @@ The core daemon's optional smoke uses a tracked speech fixture and requires a
|
|
|
766
768
|
tagged transcript, a direct ASR-to-cloned-TTS route using the shipped default
|
|
767
769
|
speaker reference, valid 24 kHz mono PCM16 output, the normal streamed TTS gate,
|
|
768
770
|
and another tagged-ASR pass after speech. It then proves that the original
|
|
769
|
-
comprehension PID and
|
|
770
|
-
|
|
771
|
-
gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
|
|
771
|
+
comprehension PID remains GPU-resident and the clone-profile TTS child has
|
|
772
|
+
exited after synthesis. A generic sound observation or unconditioned WAV cannot
|
|
773
|
+
satisfy the gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
|
|
772
774
|
Guided deployment sets it to `0`, starts the indicator in parallel with core
|
|
773
775
|
readiness, and does not wait for inference. On Jetson diagnostic requests run
|
|
774
776
|
against the installed arm64/CUDA workers;
|
|
@@ -809,12 +811,11 @@ These environment variables are worth knowing:
|
|
|
809
811
|
| `OMNI_CALL_LOG_CONTENT` | Opt in to exact structured heard/generated/TTS traces; disabled by default |
|
|
810
812
|
| `OMNI_UPDATE_INTERVAL_SECONDS` | Indicator Git update polling interval; minimum 60 seconds, default 900 |
|
|
811
813
|
|
|
812
|
-
|
|
813
|
-
|
|
814
|
-
|
|
815
|
-
|
|
816
|
-
|
|
817
|
-
prevents the kernel from overcommitting the machine.
|
|
814
|
+
`OMNI_TTS_PERSISTENT=1` selects the framed worker protocol within one synthesis
|
|
815
|
+
request; it no longer means idle GPU residency. The worker exits after the WAV
|
|
816
|
+
or PCM stream completes. `OMNI_TTS_PERSISTENT=0` keeps the isolated single-shot
|
|
817
|
+
fallback. Legacy deployments that cannot fit comprehension and active TTS may
|
|
818
|
+
still set `OMNI_CALL_SPEECH_EVICT_UNIT` to swap comprehension around synthesis.
|
|
818
819
|
|
|
819
820
|
Tool resource admission is declared beside each tool in `context.json`.
|
|
820
821
|
Standard work must clear the soft floor, bounded continuations may run within
|
|
@@ -867,9 +868,9 @@ separate so environmental sounds are never misrouted as the user's words.
|
|
|
867
868
|
- A wrench toggle, off by default, exposes server-pinned structured tools
|
|
868
869
|
for local-browser public-web discovery/fetch, attached-document retrieval,
|
|
869
870
|
current time/capabilities, on-demand host snapshots, and temporary session
|
|
870
|
-
web/memory recall and isolated text-only sub-agent delegation.
|
|
871
|
-
|
|
872
|
-
|
|
871
|
+
web/memory recall and isolated text-only sub-agent delegation. Tool chains
|
|
872
|
+
continue until a final answer, request timeout, or client disconnect. Exact
|
|
873
|
+
duplicate side effects are blocked without terminating recovery; live collapsible execution
|
|
873
874
|
evidence appears in the response and phone UI. No hosted search API is used.
|
|
874
875
|
- Same-origin IndexedDB restores messages, drafts, pending attachments, reply
|
|
875
876
|
audio, and bounded image/video previews after reload. It is keyed by a
|
|
@@ -879,12 +880,11 @@ separate so environmental sounds are never misrouted as the user's words.
|
|
|
879
880
|
The document index follows the same session partition and expiry policy.
|
|
880
881
|
- Long speech is split before the per-generation codec-frame ceiling, streamed
|
|
881
882
|
with continuous sequence numbers, and assembled into one complete final WAV.
|
|
882
|
-
- Trained-bridge runtime
|
|
883
|
-
|
|
884
|
-
|
|
885
|
-
|
|
886
|
-
|
|
887
|
-
escape hatches. Guided startup does not run a blocking generation gate.
|
|
883
|
+
- Trained-bridge runtime loads the matching shipped Qwen3-TTS voice profile for
|
|
884
|
+
one utterance, emits two codec frames (about 160 ms) per stream window, and
|
|
885
|
+
sheds the TTS child when the response completes. Moondream follows the same
|
|
886
|
+
request-scoped rule for explicit point/observe calls. Comprehension remains
|
|
887
|
+
resident. Guided startup does not run a blocking generation gate.
|
|
888
888
|
- Ordinary turns receive only a compact stable behavioral system policy. With
|
|
889
889
|
tools enabled, `get_system_snapshot` can explicitly sample current date/time,
|
|
890
890
|
OS/architecture, CPU/load, RAM, interface counters, and NVIDIA utilization.
|
|
@@ -7,7 +7,7 @@ const { spawnSync } = require("node:child_process");
|
|
|
7
7
|
const { getModel } = require("./models");
|
|
8
8
|
|
|
9
9
|
const REPOSITORY_URL = "https://github.com/robit-man/qwen-omni-adapters.git";
|
|
10
|
-
const RELEASE_REF = "npm-v0.1.
|
|
10
|
+
const RELEASE_REF = "npm-v0.1.1";
|
|
11
11
|
|
|
12
12
|
function dataRoot(platform = process.platform, env = process.env) {
|
|
13
13
|
if (env.OMNINDICATOR_HOME) return path.resolve(env.OMNINDICATOR_HOME);
|