omnindicator 0.1.0 → 0.1.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -215,8 +215,9 @@ before TTS, stops the comprehension service to make room, streams decoder PCM
215
215
  with a small startup lead, and starts restoring comprehension while audio is
216
216
  still playing. If the user interrupts, playback ducks and then pauses/fades;
217
217
  the microphone remains active throughout. This is a safe memory arrangement,
218
- not the theoretical minimum-latency arrangement. Hosts with enough memory keep
219
- matching TTS and language workers resident.
218
+ not the theoretical minimum-latency arrangement. The managed runtime now uses
219
+ the same safety invariant on every host: the TTS graph is loaded for one
220
+ utterance and shed when that response completes.
220
221
 
221
222
  The essential constrained-host overrides are:
222
223
 
@@ -359,7 +360,7 @@ details hidden behind the adapter.
359
360
  | Media comprehension | Qwen3-Omni `llama-server` | Prompt caching is disabled so stale multimodal embeddings cannot cross turns |
360
361
  | Semantic bridge | Adapter-generated tagged observation | Speech transcript, non-speech acoustics, and visual evidence stay separate and remain untrusted data |
361
362
  | Language, reasoning, tool choice | Selected Ollama base or configured OpenAI-compatible worker | Reasoning remains in `message.thinking`; unresolved tool calls cannot enter TTS |
362
- | Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, and rejects repeated nonproductive calls |
363
+ | Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, blocks exact duplicate execution, and lets recovery continue until final/timeout/disconnect |
363
364
  | Computer use | `portal/browser.py` + `portal/gui.py` | Opens visible Chromium on the desktop, observes rendered screenshots, clicks/types through native DevTools input, and can see/control the wider workspace through `xdotool` plus fresh desktop screenshots |
364
365
  | Speech | Patched Qwen3-TTS worker | Emits ordered decoder PCM; generation state is reset between prompts and never leaks one utterance into the next |
365
366
  | Local conversation | `harness/` | VAD, interruption, ReSpeaker state/direction, camera capture, history, deferred memory writes, and foreground scheduling |
@@ -588,10 +589,11 @@ dynamic inside the allocated KV window; new work is admission-gated, and the
588
589
  next supervised start reselects its tier from live memory. This avoids process
589
590
  churn without reverting to a board-specific context limit.
590
591
 
591
- The trained-bridge runtime keeps TTS and comprehension simultaneously resident,
592
- but guided deployment does not block the desktop on generation probes. It marks
592
+ The trained-bridge runtime keeps comprehension resident while TTS and pointing
593
+ weights are request-scoped. Guided deployment does not block the desktop on
594
+ generation probes. It marks
593
595
  the core ready after local component health and starts the indicator service as
594
- soon as the core unit starts. The full ASR/cloned-TTS/co-residency smoke remains
596
+ soon as the core unit starts. The full ASR/cloned-TTS/on-demand-residency smoke remains
595
597
  available as an explicit diagnostic. Legacy/manual constrained profiles may
596
598
  still opt into explicit eviction callbacks and non-persistent TTS. Service
597
599
  managers restart failed workers; the harness waits
@@ -766,9 +768,9 @@ The core daemon's optional smoke uses a tracked speech fixture and requires a
766
768
  tagged transcript, a direct ASR-to-cloned-TTS route using the shipped default
767
769
  speaker reference, valid 24 kHz mono PCM16 output, the normal streamed TTS gate,
768
770
  and another tagged-ASR pass after speech. It then proves that the original
769
- comprehension PID and persistent clone-profile TTS PID remain GPU-resident
770
- together. A generic sound observation or unconditioned WAV cannot satisfy the
771
- gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
771
+ comprehension PID remains GPU-resident and the clone-profile TTS child has
772
+ exited after synthesis. A generic sound observation or unconditioned WAV cannot
773
+ satisfy the gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
772
774
  Guided deployment sets it to `0`, starts the indicator in parallel with core
773
775
  readiness, and does not wait for inference. On Jetson diagnostic requests run
774
776
  against the installed arm64/CUDA workers;
@@ -809,12 +811,11 @@ These environment variables are worth knowing:
809
811
  | `OMNI_CALL_LOG_CONTENT` | Opt in to exact structured heard/generated/TTS traces; disabled by default |
810
812
  | `OMNI_UPDATE_INTERVAL_SECONDS` | Indicator Git update polling interval; minimum 60 seconds, default 900 |
811
813
 
812
- Keep `OMNI_TTS_PERSISTENT=1` only when speech and comprehension genuinely fit
813
- together. On constrained unified-memory hosts, use `OMNI_TTS_PERSISTENT=0` and
814
- set `OMNI_CALL_SPEECH_EVICT_UNIT`: the harness completes hearing, reasoning and
815
- tools as text, stops comprehension, synthesizes once, lets TTS exit, and restores
816
- comprehension before listening again. This is slower than resident TTS, but it
817
- prevents the kernel from overcommitting the machine.
814
+ `OMNI_TTS_PERSISTENT=1` selects the framed worker protocol within one synthesis
815
+ request; it no longer means idle GPU residency. The worker exits after the WAV
816
+ or PCM stream completes. `OMNI_TTS_PERSISTENT=0` keeps the isolated single-shot
817
+ fallback. Legacy deployments that cannot fit comprehension and active TTS may
818
+ still set `OMNI_CALL_SPEECH_EVICT_UNIT` to swap comprehension around synthesis.
818
819
 
819
820
  Tool resource admission is declared beside each tool in `context.json`.
820
821
  Standard work must clear the soft floor, bounded continuations may run within
@@ -867,9 +868,9 @@ separate so environmental sounds are never misrouted as the user's words.
867
868
  - A wrench toggle, off by default, exposes server-pinned structured tools
868
869
  for local-browser public-web discovery/fetch, attached-document retrieval,
869
870
  current time/capabilities, on-demand host snapshots, and temporary session
870
- web/memory recall and isolated text-only sub-agent delegation. Productive
871
- tool chains continue until a final answer, while exact duplicates and
872
- repeated nonproductive rounds stop safely; live collapsible execution
871
+ web/memory recall and isolated text-only sub-agent delegation. Tool chains
872
+ continue until a final answer, request timeout, or client disconnect. Exact
873
+ duplicate side effects are blocked without terminating recovery; live collapsible execution
873
874
  evidence appears in the response and phone UI. No hosted search API is used.
874
875
  - Same-origin IndexedDB restores messages, drafts, pending attachments, reply
875
876
  audio, and bounded image/video previews after reload. It is keyed by a
@@ -879,12 +880,11 @@ separate so environmental sounds are never misrouted as the user's words.
879
880
  The document index follows the same session partition and expiry policy.
880
881
  - Long speech is split before the per-generation codec-frame ceiling, streamed
881
882
  with continuous sequence numbers, and assembled into one complete final WAV.
882
- - Trained-bridge runtime keeps the matching shipped Qwen3-TTS voice profile
883
- resident alongside comprehension on its
884
- assigned GPU and emits two codec frames (about 160 ms) per stream window by
885
- default. A voice-profile change intentionally replaces the resident worker.
886
- Non-persistent workers and explicit residency handoff remain legacy/manual
887
- escape hatches. Guided startup does not run a blocking generation gate.
883
+ - Trained-bridge runtime loads the matching shipped Qwen3-TTS voice profile for
884
+ one utterance, emits two codec frames (about 160 ms) per stream window, and
885
+ sheds the TTS child when the response completes. Moondream follows the same
886
+ request-scoped rule for explicit point/observe calls. Comprehension remains
887
+ resident. Guided startup does not run a blocking generation gate.
888
888
  - Ordinary turns receive only a compact stable behavioral system policy. With
889
889
  tools enabled, `get_system_snapshot` can explicitly sample current date/time,
890
890
  OS/architecture, CPU/load, RAM, interface counters, and NVIDIA utilization.
@@ -7,7 +7,7 @@ const { spawnSync } = require("node:child_process");
7
7
  const { getModel } = require("./models");
8
8
 
9
9
  const REPOSITORY_URL = "https://github.com/robit-man/qwen-omni-adapters.git";
10
- const RELEASE_REF = "npm-v0.1.0";
10
+ const RELEASE_REF = "npm-v0.1.1";
11
11
 
12
12
  function dataRoot(platform = process.platform, env = process.env) {
13
13
  if (env.OMNINDICATOR_HOME) return path.resolve(env.OMNINDICATOR_HOME);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "omnindicator",
3
- "version": "0.1.0",
3
+ "version": "0.1.1",
4
4
  "description": "Guarded host doctor and installer for the Qwen Omni desktop indicator runtime",
5
5
  "license": "Apache-2.0",
6
6
  "author": "Robit",