omnindicator 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md ADDED
@@ -0,0 +1,959 @@
1
+ # Qwen Omni Adapters
2
+
3
+ Standalone runtime, protocol, and deployment tooling for logical Ollama Omni
4
+ models. The guided Jetson deployer offers these verified reduced profiles:
5
+
6
+ ```text
7
+ robit/qwen3.8-27b-e03-obliterated-omni-audio-bridge:q4km
8
+ robit/ornith-1.5-omni-audio-bridge:q4km
9
+ ```
10
+
11
+ The repository turns that one Ollama tag into one authenticated, Ollama-shaped
12
+ API for text, tools, optional thinking, images, audio/ASR, environmental sound
13
+ analysis, video understanding, and Qwen3-TTS speech. It also includes the
14
+ phone-first validation portal used to exercise microphone, camera, allowlisted
15
+ Female/Male voice presets, request-local voice clone,
16
+ streamed playback, call mode, and concurrent isolated sessions.
17
+
18
+ For a host that should simply listen, `harness/` runs that same call mode
19
+ locally: microphone in, speakers out, state in the desktop's top bar, no
20
+ browser involved. See [Always-listening call harness](#always-listening-call-harness).
21
+
22
+ This runtime is also the speech, perception, and local-agent integration used
23
+ by [EGG — Experimental Generalized Gateway](https://github.com/robit-man/EGG),
24
+ Robit's open-source edge-AI hardware and software platform. EGG is the larger
25
+ robot/peripheral system; this repository is the independently installable Omni
26
+ model runtime.
27
+
28
+ ## npm guided installer
29
+
30
+ The published `omnindicator` package provides a guarded one-command entry point
31
+ for a new host:
32
+
33
+ ```bash
34
+ npx omnindicator@latest
35
+ ```
36
+
37
+ It detects the platform, architecture, Tegra unified-memory versus discrete
38
+ NVIDIA topology, broker availability, host/GPU memory class, free disk space,
39
+ desktop-session signal, and required tools **before** cloning this repository
40
+ or pulling model weights. It then selects one of the two trained bridge
41
+ profiles and hands off to the native deployment script. On Linux the default
42
+ is the core service plus the visible always-listening AppIndicator harness;
43
+ only `--core-only` opts out.
44
+
45
+ The package does not pretend that platform support is identical. macOS and
46
+ Windows can install the accelerated core runtime, managed service, and portal,
47
+ but this repository does not yet ship a native tray indicator for either, so
48
+ the installer discloses that boundary and requires a core-only acknowledgement.
49
+ It also does not bundle weights or bypass the deployer's exact live-memory,
50
+ GPU-residency, readiness, and rollback checks. See the complete
51
+ [npm installer scope and capacity policy](docs/npm-installer.md).
52
+
53
+ For a persistent command instead of `npx`:
54
+
55
+ ```bash
56
+ npm install --global omnindicator
57
+ omnindicator
58
+ ```
59
+
60
+ ## Agent quick start
61
+
62
+ An automation agent starting from a clean checkout should read `AGENTS.md`,
63
+ select a model profile, bootstrap, run the deployment doctor, validate, and
64
+ only then start services. Do not copy a CUDA configuration from another host:
65
+ the launcher distinguishes broker-managed discrete GPUs from unified-memory
66
+ NVIDIA Tegra systems.
67
+
68
+ Required host tools are Python 3.10+, Node.js, Git, CMake, a working NVIDIA
69
+ CUDA or Apple Metal toolchain, FFmpeg, and a running Ollama installation.
70
+ Cloudflared is optional; without it the portal remains available only on
71
+ loopback.
72
+
73
+ ### Guided clone and install
74
+
75
+ Clone the repository and run the guided installer. On a Jetson it detects the
76
+ Tegra SoC, unified-memory size, existing runtime and managed-service state,
77
+ then presents arrow-key menus for install/upgrade, model, and local voice
78
+ harness. The Enter-through/default path always enables its desktop indicator;
79
+ core-only deployment requires the explicit `--no-harness` opt-out:
80
+
81
+ ```bash
82
+ git clone https://github.com/robit-man/qwen-omni-adapters.git
83
+ cd qwen-omni-adapters
84
+ ./deploy.sh
85
+ ```
86
+
87
+ The installer pulls the selected logical Ollama tag, validates its trained
88
+ audio-bridge sidecar, builds or upgrades the runtime, runs doctor and regression
89
+ gates, persists the exact profile, and installs/restarts the systemd service.
90
+ The default desktop harness deployment also installs the Ubuntu
91
+ GTK/AppIndicator and PulseAudio client bindings, waits for a real top-bar
92
+ indicator, and does not declare the harness ready until it reads a microphone
93
+ frame. It does not require a separate adjacent language-model download.
94
+
95
+ ### Validate and start
96
+
97
+ ```bash
98
+ export OMNI_MODEL=robit/ornith-1.5-omni-audio-bridge:q4km
99
+ export OMNI_LANGUAGE_MODEL=$OMNI_MODEL
100
+ cat AGENTS.md
101
+ .venv/bin/qwen-omni doctor --deployment
102
+ ./scripts/validate.sh
103
+ ./portal/start.sh --daemon
104
+ ./portal/start.sh --status
105
+ ```
106
+
107
+ `doctor` must report the intended accelerator and no missing required
108
+ components. Validation must finish green. `start.sh --status` prints component
109
+ health and the portal URL; never publish the URL's `#access=...` fragment,
110
+ because the fragment is the portal credential.
111
+
112
+ To install managed Linux services, including the always-listening desktop
113
+ harness:
114
+
115
+ ```bash
116
+ ./services/linux/install.sh --with-harness
117
+ ```
118
+
119
+ Use `./portal/start.sh --stop` for a foreground/staged deployment, or the
120
+ platform service manager after a service install. Do not delete
121
+ `runtime-data/components` while any worker is running.
122
+
123
+ ### One-line install and launch
124
+
125
+ For non-interactive automation, the short profile names select the same two
126
+ bridge releases and deploy the managed service:
127
+
128
+
129
+ ```bash
130
+ git clone https://github.com/robit-man/qwen-omni-adapters.git && cd qwen-omni-adapters && ./deploy.sh ornith15
131
+ ```
132
+
133
+ Available profiles are:
134
+
135
+ ```bash
136
+ ./deploy.sh ornith15 # standard Ornith 1.5 9B bridge; about 8.15 GiB
137
+ ./deploy.sh qwen38 # Qwen3.8 27B E03 bridge; about 18.33 GiB
138
+ ```
139
+
140
+ On a broker-managed GPU host the installed service uses the scoped broker. On
141
+ Tegra it uses the direct supervisor and proves residency from `nvgpu`/`nvmap`
142
+ handles. The Ornith bridge is the recommended 32 GB Jetson choice. Qwen's
143
+ 18.33 GiB artifact set fits nominally, but an eviction-free production peak is
144
+ not claimed until measured on the target board with its real context and TTS
145
+ policy.
146
+
147
+ Upgrades perform a live-runtime handoff before replacing the unit: the
148
+ deployer identifies and stops recognized old Omni listeners, unloads their
149
+ relevant Ollama runners, verifies the ports are free, and samples Jetson GPU
150
+ load and unified-memory headroom. The old configuration, unit, and managed
151
+ services are restored if the replacement does not reach ready state.
152
+
153
+ The first run creates `.venv`, installs the Python package, clones a pinned
154
+ llama.cpp revision, applies the Qwen3-TTS PCM streaming and resident-worker
155
+ patches, builds the two
156
+ CUDA binaries, pulls the selected Ollama tag, validates the attached sidecar,
157
+ materializes its disposable TTS views, installs the managed service, runs local
158
+ smoke gates, and records the authenticated portal URL in the protected daemon
159
+ status.
160
+
161
+ For a staged installation:
162
+
163
+ ```bash
164
+ ./scripts/bootstrap.sh
165
+ .venv/bin/qwen-omni doctor --deployment
166
+ ./portal/start.sh --daemon
167
+ ./portal/start.sh --status
168
+ ./portal/start.sh --stop
169
+ ```
170
+
171
+ On an arm64 NVIDIA Jetson, the launcher selects the direct managed service
172
+ because a Tegra module has an integrated GPU and no GPU broker. Run without a
173
+ profile to choose interactively:
174
+
175
+ ```bash
176
+ ./deploy.sh
177
+ ```
178
+
179
+ See [arm64 and NVIDIA Jetson](docs/arm-jetson.md) for build architecture
180
+ pinning, residency evidence, and unified-memory guidance.
181
+
182
+ Platform service installs are also one command after cloning:
183
+
184
+ ```bash
185
+ ./deploy-macos.sh # macOS Metal + launchd
186
+ ```
187
+
188
+ ```powershell
189
+ .\deploy.ps1 # Windows CUDA + managed user task
190
+ .\deploy.ps1 -Mode Service # true pywin32 Windows Service
191
+ ```
192
+
193
+ Do not expose the URL including its `#access=...` fragment publicly. The
194
+ fragment is the portal credential.
195
+
196
+ ## Current implementation state
197
+
198
+ The guided deployment uses a trained audio bridge: the selected Qwen3.8 or
199
+ standard Ornith trunk handles audio/ASR, native vision, language, and tools in
200
+ one llama.cpp server, while Qwen3-TTS provides speech. Legacy full-Omni bundles
201
+ remain supported and use Qwen3-Omni for media comprehension plus a selected
202
+ language base. The previously validated legacy 32 GB AGX Orin deployment uses
203
+ a tighter residency profile because all three graphs cannot safely coexist in
204
+ its 29.98 GiB unified-memory pool:
205
+
206
+ | Component | Current constrained-host role | Residency |
207
+ |---|---|---|
208
+ | Qwen3-Omni + projector | Speech/audio/image/video comprehension **and** language/tool reasoning through its OpenAI-compatible endpoint | Resident; context chosen from live memory and published to the adapter; post-speech recovery live-validated at 4K and 8K under desktop load |
209
+ | Ornith 1.5 base | Logical release/base option, but not loaded by the constrained profile | Not resident |
210
+ | Qwen3-TTS + codec projector | Final 24 kHz PCM16 speech | Loaded only after text/tools finish; exits after the utterance |
211
+ | Nomic text embedder | Passive semantic conversation memory | Admitted only while foreground, background-agent, TTS, and restoration work are idle and live headroom permits |
212
+
213
+ On that profile the harness finishes comprehension, reasoning, and tool calls
214
+ before TTS, stops the comprehension service to make room, streams decoder PCM
215
+ with a small startup lead, and starts restoring comprehension while audio is
216
+ still playing. If the user interrupts, playback ducks and then pauses/fades;
217
+ the microphone remains active throughout. This is a safe memory arrangement,
218
+ not the theoretical minimum-latency arrangement. Hosts with enough memory keep
219
+ matching TTS and language workers resident.
220
+
221
+ The essential constrained-host overrides are:
222
+
223
+ ```bash
224
+ # Run Qwen3-Omni as an independently supervised worker on port 8901.
225
+ export OMNI_ENABLE_COMPREHENSION=0
226
+ export OMNI_COMPREHENSION_URL=http://127.0.0.1:8901/v1/chat/completions
227
+
228
+ # Reuse that resident worker for language instead of loading Ornith beside it.
229
+ export OMNI_LANGUAGE_API=openai
230
+ export OMNI_LANGUAGE_URL=http://127.0.0.1:8901/v1/chat/completions
231
+
232
+ # Do not retain TTS beside comprehension on a 29.98 GiB pool.
233
+ export OMNI_TTS_PERSISTENT=0
234
+ export OMNI_CALL_SPEECH_EVICT_UNIT=egg-omni-comprehension.service
235
+ export OMNI_CALL_COMPREHENSION_HEALTH=http://127.0.0.1:8901/health
236
+ ```
237
+
238
+ The comprehension service uses `runtime/comprehension_launcher.py` rather
239
+ than a fixed `-c` value. Set a service `MemoryMax=` as an independent final
240
+ guard; live admission and automatic downshift remain the primary mechanism.
241
+
242
+ Current local-call behavior also includes:
243
+
244
+ - model-requested camera capture rather than transcript keyword heuristics;
245
+ - a compact discovery tool plus on-demand schemas, with raw shell available to
246
+ the trusted local harness;
247
+ - screenshot-grounded control of a real, visible Chromium window and the full
248
+ Ubuntu desktop, with every action followed by fresh visual evidence;
249
+ - crash-safe persistent background tasks that checkpoint after each
250
+ bounded inference/tool slice, finalize only through referenced tool evidence,
251
+ yield cancellable inference to foreground speech, accept later spoken
252
+ guidance, keep knowledge/environment/controller state outside renewable model
253
+ transcripts, and use native thinking without speaking or storing it;
254
+ - bounded conversational history with time-based relevance reduction and
255
+ explicit current-query memory tools rather than next-turn prefetch;
256
+ - live memory-derived comprehension context selection and automatic
257
+ downshifting rather than a board-specific fixed context size;
258
+ - continuous PCM playback, timing/starvation diagnostics, ReSpeaker echo
259
+ handling, and automatic service/microphone-loop recovery.
260
+ - a resident, typed Laya System-1 decision plane with batched routing,
261
+ pre-action, post-action, and context-relevance waves; it is shadow-only until
262
+ real calibration proves a per-family fast path, and failures always preserve
263
+ the deliberative route. See [`docs/decision-plane.md`](docs/decision-plane.md).
264
+
265
+ ## What the model tag contains
266
+
267
+ The release is one logical Ollama model, not one graph that stock Ollama can
268
+ execute end to end:
269
+
270
+ ```text
271
+ logical Ollama tag
272
+ ├── standard model/projector/template layers
273
+ │ └── selected Qwen-family base: text, image vision, tools, optional thinking
274
+ └── application/vnd.robit.ollama.omni.bundle.v1+gguf
275
+ ├── Qwen3-Omni comprehension model + projector
276
+ └── Qwen3-TTS model + codec/projector
277
+ ```
278
+
279
+ Stock Ollama handles the standard layers. This adapter resolves the custom
280
+ sidecar layer, reconstructs byte-preserving executable component views, and
281
+ runs audio/video comprehension and TTS with the pinned llama.cpp build. The
282
+ public request remains Ollama-shaped and names the one logical tag.
283
+
284
+ The runtime also accepts the reduced `robit.ollama-audio-bridge.v1` profile.
285
+ There the standard projector combines target-native vision with the frozen
286
+ Omni audio encoder and its trained final projection. The sidecar carries only
287
+ TTS, and one local llama.cpp server is both the multimodal comprehension path
288
+ and the sole language/tool trunk. This removes the full secondary Omni Thinker
289
+ and avoids loading an adjacent Ollama language copy.
290
+
291
+ Legacy full-Omni tags remain semantic routers because Qwen3.8, Qwen3-Omni, and
292
+ Qwen3-TTS do not share compatible hidden-state interfaces. The trained bridge
293
+ profiles are different: they contain the frozen Omni audio encoder plus a
294
+ trained final projection into the selected trunk, while keeping evidence tags
295
+ and TTS as explicit runtime boundaries. Standard Ornith and Qwen3.8 E03 tensors
296
+ and Ollama tags are never interchangeable.
297
+
298
+ ## Architecture breakdown
299
+
300
+ ### End-to-end data plane
301
+
302
+ ```text
303
+ browser / API client local always-listening harness
304
+ │ │
305
+ └──────────────────┬───────────────────────────────┘
306
+ ▼
307
+ authenticated portal (optional)
308
+ session isolation · admission queue
309
+ tool execution · diagnostics · NDJSON relay
310
+ │
311
+ ▼
312
+ unified adapter API
313
+ │
314
+ ┌────────────────┴─────────────────┐
315
+ │ │
316
+ text-only request current media request
317
+ │ audio · image · video
318
+ │ │
319
+ │ validation / normalization
320
+ │ │
321
+ │ ▼
322
+ │ Qwen3-Omni comprehension
323
+ │ │
324
+ │ tagged, untrusted evidence
325
+ │ ┌────────────┼────────────┐
326
+ │ │ │ │
327
+ │ transcript sound context visual evidence
328
+ │ └────────────┼────────────┘
329
+ └──────────────────────────────────┘
330
+ ▼
331
+ selected language backend
332
+ Qwen3.8 · Ornith · Qwen3-Omni
333
+ │
334
+ ┌──────────────┴──────────────┐
335
+ │ unresolved structured calls?│
336
+ ▼ │
337
+ portal tool/discovery loop ─────────────┘
338
+ │
339
+ ▼
340
+ final answer text
341
+ │
342
+ speech requested and safe?
343
+ ▼
344
+ Qwen3-TTS
345
+ │
346
+ ordered 24 kHz mono PCM16
347
+ ```
348
+
349
+ The public model name remains the logical Ollama tag throughout. Component
350
+ URLs, process placement, model eviction, and derived GGUF views are deployment
351
+ details hidden behind the adapter.
352
+
353
+ ### Stage ownership
354
+
355
+ | Stage | Owner | Important invariant |
356
+ |---|---|---|
357
+ | Request parsing | `qwen_omni_adapters.contract` | Bounds and validates media before inference; preserves native `think`, tools, options, and response modalities |
358
+ | Audio/image/video preparation | `runtime/adapter_server.py` + FFmpeg where needed | Only the newest attachment is current evidence; video duration/frame count and decoded audio are bounded |
359
+ | Media comprehension | Qwen3-Omni `llama-server` | Prompt caching is disabled so stale multimodal embeddings cannot cross turns |
360
+ | Semantic bridge | Adapter-generated tagged observation | Speech transcript, non-speech acoustics, and visual evidence stay separate and remain untrusted data |
361
+ | Language, reasoning, tool choice | Selected Ollama base or configured OpenAI-compatible worker | Reasoning remains in `message.thinking`; unresolved tool calls cannot enter TTS |
362
+ | Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, and rejects repeated nonproductive calls |
363
+ | Computer use | `portal/browser.py` + `portal/gui.py` | Opens visible Chromium on the desktop, observes rendered screenshots, clicks/types through native DevTools input, and can see/control the wider workspace through `xdotool` plus fresh desktop screenshots |
364
+ | Speech | Patched Qwen3-TTS worker | Emits ordered decoder PCM; generation state is reset between prompts and never leaks one utterance into the next |
365
+ | Local conversation | `harness/` | VAD, interruption, ReSpeaker state/direction, camera capture, history, deferred memory writes, and foreground scheduling |
366
+ | Persistent work | `harness/background_agent.py` + `portal/background_tasks.py` | Durable checkpoints and leases survive process restarts; long jobs may yield sparse spoken milestones, terminal speech is durable, and foreground speech wins every scheduling boundary |
367
+
368
+ ### A normal spoken turn
369
+
370
+ 1. The microphone loop continuously captures audio. Adaptive VAD accepts real
371
+ near-end speech, joins brief continuation segments, and keeps the newest
372
+ bounded turn rather than filling an inference queue.
373
+ 2. The harness submits one audio-bearing request. Qwen3-Omni produces a tagged
374
+ transcript and any non-speech observation; silence or a cough stops before
375
+ language and TTS when no speech was found.
376
+ 3. The configured language backend receives the recognized speech as the latest
377
+ user message, with bounded text history and non-speech observations as
378
+ secondary tagged evidence. Empty negative acoustic boilerplate such as “no
379
+ non-speech sounds” remains available in adapter diagnostics but is omitted
380
+ from the language prompt so it cannot contradict a valid transcript. Visual evidence is added only after a structured
381
+ camera request, or when the client explicitly attached media. The
382
+ local voice harness enables the portal's real tool allowlist on every turn
383
+ without a separate classification pass: the model either answers
384
+ conversationally or calls the smallest tool that accomplishes the request,
385
+ the portal auto-executes it, and the spoken answer is the grounded reply that
386
+ follows the completed work. On the constrained 32 GB profile this is the same
387
+ resident Qwen3-Omni worker; a normal profile uses the selected Ollama base.
388
+ 4. A durable-task request goes through the same pass via the `background_task`
389
+ tool, which writes the objective to the durable store before any
390
+ acknowledgment. Camera-enabled embodied turns expose a capture bridge but do
391
+ not attach an ambient still unless the request needs current physical-scene
392
+ evidence; web and other current-information requests complete a real tool
393
+ call in the pass. No answer is eligible for speech before the work it claims
394
+ has actually completed. Once transcript-aware routing finds concrete matching
395
+ schemas, the language trunk receives their names in a compact required-action
396
+ context and still chooses the tool and arguments itself.
397
+ 5. Only final answer text is sent to TTS. PCM is played as decoder windows
398
+ arrive, with one small initial lead to absorb packet jitter rather than
399
+ waiting for the complete WAV.
400
+ 6. Conversation persistence and embeddings are queued after answer, tools, and
401
+ speech. They cannot delay the turn.
402
+
403
+ The local harness keeps foreground reasoning disabled by default for natural
404
+ response latency. Persistent background tasks enable native thinking because
405
+ deliberation is more valuable than sub-second response there; the thinking
406
+ channel is never synthesized or written into the task transcript.
407
+
408
+ ### Camera and video flow
409
+
410
+ The browser may attach a current image or bounded clip directly. The local
411
+ harness deliberately does not attach a room image to every spoken question:
412
+ doing that biases the model into describing the scene even when the user asked
413
+ about something else. Instead, the model calls `request_camera_view`; the
414
+ harness then captures every configured V4L2 camera at the same moment, stitches
415
+ and downscales them, and performs a grounded multimodal follow-up. A motion
416
+ request captures a bounded clip. Internet uses of words such as “look up” are
417
+ therefore routed by the model to web tools rather than intercepted by a local
418
+ keyword list. The grounded follow-up carries the fresh visual evidence but no
419
+ second camera bridge, preventing a recapture loop. The pre-capture placeholder
420
+ is never spoken, logged as generated dialogue, or kept in conversation history;
421
+ the grounded pass is the sole answer. Broad casual questions get a short casual
422
+ overview, while questions about a particular item or feature stay focused on
423
+ that target instead of inventorying the rest of the scene.
424
+
425
+ ### Tools and long-horizon work
426
+
427
+ The browser portal keeps powerful tools opt-in. The trusted local voice harness
428
+ exposes the persistent `background_task` bridge beside tool discovery; its
429
+ worker, rather than the latency-critical spoken pass, owns unrestricted shell.
430
+ This structural split prevents a small model from entering a synchronous shell
431
+ retry loop and stranding the conversation. A request that inspects or mutates
432
+ files or the system, needs verification or retry, or spans multiple commands is
433
+ accepted into a crash-safe JSON store and acknowledged immediately, after which
434
+ the background worker:
435
+
436
+ ```text
437
+ claim lease → reason once → execute/discover one or more tools
438
+ → inspect results → checkpoint → yield to speech → continue
439
+ ```
440
+
441
+ Shell commands return stdout, stderr, exit status, timeouts, and truncation
442
+ markers. Unsupported “done” claims are rejected until at least one concrete
443
+ action produced evidence. The worker has no fixed step horizon. On a genuinely
444
+ long task it may pause after a meaningful verified milestone, speak a short
445
+ progress explanation, then resume with the same task context. A later spoken
446
+ update is appended as authoritative task guidance before the next step,
447
+ including a final race check so stale completion cannot beat a new instruction.
448
+ Typed focus records advertise no paging operation while their detailed receipts
449
+ are still resident. After compaction they gain structured pointers to the
450
+ separate `task_expand` control function; that control name is never presented as
451
+ an action or argument of `workspace_file` or another external tool. The focus
452
+ ledger is a live working set, not another copy of history: its size follows the
453
+ currently resident comprehension window, repeated inspections and failures are
454
+ coalesced, a changed artifact supersedes its older resident version, and omitted
455
+ record counts remain visible. Exact receipts and every superseded version remain
456
+ append-only in task evidence storage. A compacted request carries this ledger
457
+ once in the pinned system contract rather than duplicating it in a recurrent
458
+ checkpoint. If an omitted record's evidence ID is no longer resident,
459
+ `task_expand` also accepts a distinctive exact path, URL, symbol, error, or
460
+ other query and returns the highest-scoring immutable receipts.
461
+ It is a one-shot page-in, never capability discovery: after either a successful
462
+ expansion or `evidence_not_found`, it leaves the next action surface until a new
463
+ concrete result makes older evidence relevant again. Missing tool schemas route
464
+ through `tool_search`, preventing a failed lexical page query from becoming a
465
+ maintenance loop.
466
+ The ledger applies the same evidence authority as checkpoints: search results
467
+ marked discovery-only never become acquired-source successes, and an empty
468
+ `about:blank` browser snapshot remains an inspection. A blank visible browser
469
+ explicitly requires `navigate` to a grounded URL and cannot divert the next
470
+ round into pixel clicking or unrelated capability discovery.
471
+ At the 4K/8K tiers the worker also projects the already-selected tool schema to
472
+ its executable JSON constraints: names, types, enums, required fields, bounds,
473
+ and `additionalProperties` remain exact while repeated prose annotations are
474
+ removed. Only one selected concrete capability is exposed per action round;
475
+ discovery is suppressed until that leaf receives one concrete attempt, then
476
+ returns as the route to a different capability. At constrained tiers discovery
477
+ and evidence expansion alternate instead of crowding the same envelope. A
478
+ background request that still cannot fit the live pack fails explicitly and retries from
479
+ its checkpoint—it never silently falls back to native FIFO truncation.
480
+ Discovery queries describe only the missing interaction mechanism (for example,
481
+ public-web search, file editing, shell execution, or visible-browser control),
482
+ not the task topic. The deterministic router also recognizes product/SaaS,
483
+ comparison, interface, and design research as web discovery, preventing a word
484
+ such as “service” in the subject from accidentally selecting system shell.
485
+ Background web fetches are provenance-bound as well: a target must appear in
486
+ the accepted task/user input or prior search, browser, crawl, or fetch evidence.
487
+ An invented address is rejected before network access and returned beside the
488
+ exact admissible URLs so the next model step can choose a grounded source.
489
+ The pinned task envelope also carries an explicit execution frontier: work
490
+ advances the earliest unmet prerequisite in the user's stated order, and a
491
+ failed downstream probe returns the worker to that prerequisite instead of
492
+ encouraging variants of the same premature verification.
493
+ Completed or blocked work ends with a brief natural spoken status only when the
494
+ live conversation is idle. Detailed reports, evidence IDs, paths, and worker
495
+ self-assessment stay in the indicator and task archive and are never passed
496
+ verbatim to TTS. The pending-delivery flag survives a harness restart, and the
497
+ speech can be interrupted like any other reply. A direct `synthesize` request
498
+ is a literal text-to-speech transport pass: it retains the selected voice-clone
499
+ profile but bypasses conversation policy, tools, documents, and virtual-memory
500
+ indexing/replay so recalled text cannot be appended to the utterance.
501
+
502
+ Rendered computer work is not reduced to a fetched-text corpus. The
503
+ `browser_interact` tool launches a real Chromium window on the active desktop,
504
+ returns both a screenshot and a bounded accessibility/element map, performs
505
+ real pointer/keyboard input, then observes the changed page. `gui_interact`
506
+ extends the same screenshot → action → screenshot loop to Ubuntu. Its
507
+ default screenshot is the active window and its coordinates are relative to
508
+ that image. The runtime captures root-window pixels and crops them to the exact
509
+ X11 geometry used for pointer translation, so window-manager decorations cannot
510
+ offset clicks. Full-screen mode is explicit for panels and workspace navigation;
511
+ when a later pointer call omits its coordinate space, it remains bound to the
512
+ newest returned frame. Active-window clicks fail safely if focus changed after
513
+ observation. Every result also reports whether the coarse visual state materially
514
+ changed, which lets the worker reject a missed click as non-progress.
515
+ Chromium's rendered network-error documents remain visible evidence, but are
516
+ tagged as failed navigation and release the sticky browser action space; a
517
+ connection-refused page can never count as successful GUI verification.
518
+ Actionable DOM controls are always re-resolved through live CDP geometry. For
519
+ canvas, challenge, and image targets on Jetson, the conversational vision pass
520
+ supplies a concise referring expression and an isolated resident Moondream 2
521
+ point head resolves it against the exact CDP viewport. The browser re-captures
522
+ and compares that viewport after point inference before sending input, so a
523
+ coordinate is never carried onto changed pixels. Multiple returned matches are
524
+ disambiguated by the conversational model's coarse current-frame point; the
525
+ point head's structured coordinate remains the executed value. If the point
526
+ head is not configured, the existing bounded crop-refinement loop remains the
527
+ fail-safe fallback rather than widening hit tolerances.
528
+ With the point head active, every visual click requires a concise target phrase.
529
+ If the planner's coarse crop misses, the executor searches the same horizontal
530
+ band in bounded overlapping tiles before widening to a bounded tile grid; the
531
+ point head still selects every executable coordinate.
532
+ An explicit verification snapshot or completed visual click also asks that same
533
+ resident visual worker to transcribe visible status text and exact completion
534
+ identifiers from the exact returned CDP frame. That bounded reading is stored as
535
+ current-frame evidence, while a compact action receipt preserves the exact
536
+ grounded target, so terminal state and action reporting survive image compaction
537
+ without relying on an invented marker or renamed target.
538
+ Screenshot bytes are shown to the multimodal model for one reasoning
539
+ pass and then removed from the durable transcript so long tasks retain visual
540
+ grounding without filling their context with base64. Local HTTP pages and all
541
+ desktop/file tools remain usable offline; public sites naturally require a
542
+ working network.
543
+
544
+ ### Conversation state and durable memory
545
+
546
+ The adapter itself is stateless between requests. The browser owns its
547
+ cookie-isolated conversation record; the local harness owns a bounded recent
548
+ text history with time-based falloff. Current media is never replayed from that
549
+ history.
550
+
551
+ Optional durable voice memory stores completed text exchanges in SQLite and
552
+ uses a small semantic encoder rather than forcing chat-model hidden states into
553
+ an embedding role. Storage runs on a daemon worker; retrieval is an explicit
554
+ tool call against the current request, never a result automatically carried
555
+ from one turn into the next. Admission checks
556
+ foreground activity, background-agent work, comprehension readiness, the
557
+ encoder's installed payload, current `MemAvailable`, and the active model's
558
+ measured KV slope. If any check fails, deferred storage waits; conversation
559
+ never waits for it.
560
+
561
+ ### Context, residency, and recovery
562
+
563
+ On unified-memory hosts, `runtime/comprehension_launcher.py` reads the installed
564
+ GGUF, derives KV bytes per token, samples live available memory, and chooses the
565
+ largest context tier that fits with the runtime memory reserve. Compact models
566
+ may have a 262,144-token native positional range, but the managed compact
567
+ profiles intentionally expose a 16,384-token resident working-set ceiling.
568
+ The guided Ornith-on-Tegra profile uses its live-qualified q8 KV cache; other
569
+ model/platform pairs retain f16 until they pass the same answer, multimodal,
570
+ voice, tool, and memory-pressure gates. The configured 16K value is a ceiling:
571
+ a longer live action soak selected 8K and then 4K as the rest of the resident
572
+ stack consumed unified memory. The live tier is authoritative.
573
+ Longer history is paged through the lossless virtual-context layer instead of
574
+ preallocating a nominal 256K KV cache. First load uses the conservative complete
575
+ component-byte footprint; later loads also use measured residency. It records
576
+ before/after residency and automatically caps the next load below a tier that
577
+ exits or leaves too little memory. The adapter reads the chosen window per
578
+ request and sheds old history/tool evidence before llama.cpp can reject an
579
+ oversized prompt.
580
+
581
+ On discrete-memory hosts, sampling continues after readiness: a continuous
582
+ low-headroom interval downshifts one tier, and sustained surplus performs the
583
+ inverse only while every inference slot is idle and the cooldown has elapsed.
584
+ On Tegra, the startup-selected llama.cpp process stays pinned for the service
585
+ session because affected JetPack kernels can panic while closing GPU character
586
+ devices during an otherwise controlled worker restart. Prompt budgets remain
587
+ dynamic inside the allocated KV window; new work is admission-gated, and the
588
+ next supervised start reselects its tier from live memory. This avoids process
589
+ churn without reverting to a board-specific context limit.
590
+
591
+ The trained-bridge runtime keeps TTS and comprehension simultaneously resident,
592
+ but guided deployment does not block the desktop on generation probes. It marks
593
+ the core ready after local component health and starts the indicator service as
594
+ soon as the core unit starts. The full ASR/cloned-TTS/co-residency smoke remains
595
+ available as an explicit diagnostic. Legacy/manual constrained profiles may
596
+ still opt into explicit eviction callbacks and non-persistent TTS. Service
597
+ managers restart failed workers; the harness waits
598
+ through token rotation and adapter/comprehension restoration, reopens a failed
599
+ microphone capture loop, and resumes expired background-task leases. On Tegra,
600
+ GPU residency is proven from each process's `nvgpu`/`nvmap` handles rather than
601
+ the unsupported `nvidia-smi` process table.
602
+
603
+ ### Trust and isolation boundaries
604
+
605
+ - The portal binds locally, requires a generated bearer capability, and may
606
+ publish an authenticated Cloudflare tunnel. The URL fragment is secret.
607
+ - Browser documents, web cache, memory, location, diagnostics, and UI state are
608
+ cookie-scoped and expire after disconnect or are removed immediately by
609
+ Trash.
610
+ - Media text, OCR, transcripts, web pages, tool results, and model observations
611
+ are evidence, never system instructions or authorization.
612
+ - Host snapshots exclude hostnames, addresses, routes, sockets, processes,
613
+ credentials, and session content.
614
+ - CUDA execution has no silent CPU fallback; startup verifies the platform's
615
+ actual GPU-residency evidence.
616
+
617
+ ## Capability map
618
+
619
+ | Capability | Runtime owner | Available through this adapter |
620
+ |---|---|---|
621
+ | Text and Markdown, including responsive GFM tables | Selected Qwen3.8/Ornith Ollama base or configured Qwen3-Omni language worker | Yes |
622
+ | Structured tools | Selected language backend + portal executor | Yes |
623
+ | Portal web/document/session-memory tools | Explicit opt-in allowlisted portal loop with no-key DuckDuckGo HTML discovery | Yes |
624
+ | Visible Chromium interaction | Persistent rendered browser, screenshot + element evidence, click/type/scroll/back | Yes |
625
+ | Full desktop computer use | Fresh active-window screenshots with image-relative input; explicit full-screen workspace mode | Yes |
626
+ | Trusted local shell | Unrestricted execution in the checkpointed voice-task worker; bounded output and timeout | Yes |
627
+ | Persistent background tasks | Checkpointed long-horizon worker with status/update/cancel, sparse spoken milestones, durable terminal speech, and restart recovery | Yes |
628
+ | Passive semantic voice memory | Idle-only encoder worker + SQLite; never gates a foreground turn | Yes |
629
+ | Thinking | Native backend `think` control, separate from answer text and never spoken | Yes |
630
+ | Image understanding | Qwen3.8 or Omni comprehension path | Yes |
631
+ | Speech transcription | Qwen3-Omni comprehension | Yes |
632
+ | Environmental audio interpretation | Qwen3-Omni tagged observation | Yes |
633
+ | Video understanding | Qwen3-Omni bounded `input_video` | Yes |
634
+ | Silent video and animated GIF | FFmpeg probe/normalization → Qwen3-Omni | Yes |
635
+ | PDF/DOCX/text retrieval | Session-isolated portal extraction/index | Yes |
636
+ | Spoken response | Qwen3-TTS, 24 kHz mono PCM16 | Yes |
637
+ | Voice reference cloning | Qwen3-TTS Base speaker embedding path | Yes |
638
+ | Live-call turns | Adaptive VAD + bounded single-flight speech consolidation + streamed text/PCM | Yes |
639
+ | Always-listening local call mode | `harness/`: local mic/speakers, mandatory GNOME top-bar state for its managed desktop unit | Yes |
640
+ | Every camera at once, on request | `harness/camera.py` snaps all V4L2 devices together, stitches and downscales to one image or clip only after a visual-evidence request | Yes |
641
+ | ReSpeaker ring and direction | Used when the array is attached, ignored when it is not | Yes |
642
+ | Tool execution by an external loop | `GET /api/tools`, `POST /api/tools/<name>/call` | Yes |
643
+ | Video generation | No component is shipped | No |
644
+
645
+ ## Always-listening call harness
646
+
647
+ `harness/` turns a host that has the model into a host you can talk to. It
648
+ drives the same endpoint and the same request shape as the portal's browser
649
+ call mode, through the machine's own microphone and speakers.
650
+
651
+ ```bash
652
+ # The adapter must be running; the harness waits for it either way.
653
+ PYTHONPATH=. .venv/bin/python -m harness
654
+ ```
655
+
656
+ It listens continuously, answers out loud, and shows what it is doing in the
657
+ GNOME top bar (`Omni ●` listening, `◉` hearing, `◍` thinking, `▶` speaking).
658
+ The indicator's menu mutes the microphone, toggles tools, reasoning and
659
+ cameras, copies the public link when the portal is published through a tunnel,
660
+ opens a loopback-only live view of every camera stitched left-to-right, and
661
+ cleanly reloads the voice service. It polls `origin/main` in the background and
662
+ shows an **Update available — install and restart** action when a verified
663
+ fast-forward exists. Clicking it preserves untracked local files, refuses
664
+ tracked edits or diverged history, refreshes the runtime environment without
665
+ implicitly downloading or changing model weights, and asks the managed runtime
666
+ and indicator to restart on the new revision. Its
667
+ **Models** submenu lists both compact
668
+ audio bridges and both legacy full bundles. Missing tags expose **Download**
669
+ with live percentage/status text; downloaded tags expose **Activate**, **Load
670
+ into Ollama**, **Unload from Ollama**, and confirmed **Delete local copy**
671
+ actions as applicable. Activation atomically selects the one logical language/
672
+ Omni tag, requests a controlled core restart, and lets the indicator reconnect
673
+ with the new service environment. An active direct-daemon model is labelled
674
+ separately from an optional Ollama runner so the UI never hides a duplicate
675
+ allocation on a 32 GB Jetson. The adjacent **Voice** submenu selects **Default
676
+ female**, **Default male**, or any previously imported custom clone reference.
677
+ **Add custom voice…** opens the native audio file picker, normalizes the owned
678
+ clip to a bounded 16 kHz mono WAV, stores it under untracked `runtime-data`,
679
+ adds it to the persistent list, and selects it for the next reply. Voice-profile
680
+ changes hot-reload in the portal; only the TTS worker changes speaker profile,
681
+ so the language and comprehension weights stay resident. It also appends
682
+ live/recent durable tasks as native submenus, so inspecting a task does not
683
+ close the whole menu. Each
684
+ submenu shows the current-stage spinner, exact bounded tool-call arguments and
685
+ outcomes, retained checkpoints, and terminal result. Long action rows wrap and
686
+ ellipsize within the menu while preserving the complete text in their tooltip,
687
+ so tool arguments cannot widen the indicator beyond the screen. Live tasks expose
688
+ **Cancel task** and every record exposes **Clear task record**. **Clear finished
689
+ tasks** moves all terminal records into the human-readable archive, and **Open
690
+ task archive** opens that log in the desktop editor. A real themed state icon
691
+ sits to the left of `Omni`. A manual invocation may run headless and log instead.
692
+ The managed desktop service fails closed if GTK, the StatusNotifier host, the
693
+ desktop audio server, or a usable microphone/sink is unavailable; it never
694
+ silently leaves an always-listening service running without its indicator.
695
+
696
+ Defaults are chosen for a spoken conversation:
697
+
698
+ - **Reasoning off.** A hidden chain of thought is silence the other person has
699
+ to sit through.
700
+ - **Tools on, in the answering pass.** This is the same server-side chain used
701
+ by the cloudflared portal: one tiny discovery schema rides with the turn,
702
+ only the matching concrete contracts appear on the next tool round, and
703
+ requested calls execute until the model has a grounded final answer. The
704
+ trusted local harness also carries compact shell and persistent-task bridges.
705
+ The full catalog no longer displaces conversation or memory context.
706
+ - **Every camera together, only when relevant.** The audio-only pass may call
707
+ `request_camera_view` when the answer depends on the current physical scene.
708
+ Only then are V4L2 devices opened and probed, all working cameras are snapped
709
+ at the same moment, and their frames are stitched into one left-to-right row.
710
+ The indicator's explicit live-view action uses the same horizontal composition. Clips
711
+ work the same way for temporal questions. Merely starting the harness never
712
+ launches FFmpeg or activates a camera privacy indicator.
713
+ - **Observation does not force speech.** Sound-only events are retained as
714
+ bounded context without language or TTS. The language model may also leave a
715
+ transcribed room utterance unanswered when current evidence shows it was
716
+ addressed elsewhere; if gaze would resolve genuine ambiguity, it can request
717
+ a fresh still before deciding. Empty intentional responses do not invoke TTS.
718
+ - **ReSpeaker when present.** Its ring follows the conversation and the
719
+ direction a voice came from is attached to the turn as evidence. With no
720
+ array attached the default microphone is used and nothing else changes.
721
+ - **Memory storage is passive.** Completed exchanges are embedded on a daemon
722
+ worker only after the answer, tools and speech finish. Recall uses an explicit
723
+ current-query portal tool, so a result selected for one utterance cannot
724
+ contaminate the next one. Storage never gates hearing, answering, reasoning,
725
+ tools, or speech. The
726
+ Ornith/Omni chat weights have no embedding head and measured poorly when
727
+ forced into that role, so the small dedicated encoder remains the deliberate
728
+ exception and unloads after each job.
729
+ - **Long work is externally audited, checkpointed, and self-checked.** The foreground turn can hand a sustained job
730
+ to the persistent worker, acknowledge immediately, and keep listening. The
731
+ worker reasons and uses tools between speech turns, while a deterministic
732
+ state ledger separately records learned evidence, verified environment
733
+ mutations, controller transitions, and repeated-state stagnation. Accepted
734
+ phase boundaries renew the executor context from that ledger. Spoken
735
+ refinements retain exact user provenance, and completion after a mutation
736
+ requires fresh verification against the current environment version.
737
+ - **Live host state is explicit.** Ordinary turns carry no eager clock,
738
+ location, network, battery, or process blob. Current time, client-browser
739
+ approximate location, and bounded hardware/load facts come from their
740
+ dedicated tools only when the request needs them. The local voice client
741
+ warms the same privacy-filtered browser lookup while the model loads.
742
+
743
+ The speech detector is a port of `portal/static/call_vad.js` with its constants
744
+ intact, so the same room behaves the same way in the browser and here.
745
+
746
+ ### Running it as a service
747
+
748
+ ```bash
749
+ services/linux/install.sh --with-harness
750
+ ```
751
+
752
+ That installs the adapter as a system service and the harness as a **user**
753
+ service. The distinction matters: the harness needs the desktop session it
754
+ speaks into -- its audio devices, its top bar, and the USB permissions the
755
+ logged-in user already has -- so installing it system-wide would leave it
756
+ listening on behalf of nobody.
757
+
758
+ The unit template is `services/linux/omni-call-harness.service.in`. It is a
759
+ desktop-session user unit, while the adapter is a system unit, so it waits for
760
+ the loopback portal rather than declaring an invalid cross-manager dependency.
761
+ Its preflight runs in the exact user-service environment and requires the real
762
+ indicator and audio stack. The unit caps its own memory: the adapter holds the
763
+ weights, this process only moves audio.
764
+
765
+ The core daemon's optional smoke uses a tracked speech fixture and requires a
766
+ tagged transcript, a direct ASR-to-cloned-TTS route using the shipped default
767
+ speaker reference, valid 24 kHz mono PCM16 output, the normal streamed TTS gate,
768
+ and another tagged-ASR pass after speech. It then proves that the original
769
+ comprehension PID and persistent clone-profile TTS PID remain GPU-resident
770
+ together. A generic sound observation or unconditioned WAV cannot satisfy the
771
+ gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
772
+ Guided deployment sets it to `0`, starts the indicator in parallel with core
773
+ readiness, and does not wait for inference. On Jetson diagnostic requests run
774
+ against the installed arm64/CUDA workers;
775
+ desktop-host unit tests do not substitute for that device gate.
776
+
777
+ After deployment, run the non-blocking post-training tool-routing suite
778
+ explicitly without delaying normal boot:
779
+
780
+ ```bash
781
+ .venv/bin/python portal/smoke.py \
782
+ --endpoint http://127.0.0.1:8920 \
783
+ --token-file runtime-data/state/access-token.txt \
784
+ --model "$(awk -F= '$1 == "OMNI_MODEL" {print $2}' .env)" \
785
+ --tool-suite
786
+ ```
787
+
788
+ It requires real structured calls for portal capabilities, arithmetic, current
789
+ runtime state, and time; a prose capability disclaimer does not pass.
790
+
791
+ These environment variables are worth knowing:
792
+
793
+ | Variable | Effect |
794
+ |---|---|
795
+ | `OMNI_PORTAL_URL` | Where the portal is (default `http://127.0.0.1:8920`) |
796
+ | `OMNI_CALL_CAMERA` | A single camera to use instead of every one found |
797
+ | `OMNI_CALL_MEMORY` | Persistent passive-memory SQLite path |
798
+ | `OMNI_CALL_SPEECH_EVICT_UNIT` | User service to stop before TTS and restore afterward |
799
+ | `OMNI_CALL_COMPREHENSION_HEALTH` | Readiness URL used after restoring that service |
800
+ | `OMNI_MEMORY_GOVERNOR` | Enable (`1`) or disable (`0`) generic runtime memory admission; enabled automatically on Tegra |
801
+ | `OMNI_MEMORY_SOFT_FLOOR_GIB` | Free-memory floor retained before work establishes new model/tool residency (default `3`) |
802
+ | `OMNI_MEMORY_HARD_FLOOR_GIB` | Emergency floor that cancels cancellable work before kernel OOM (default `2`) |
803
+ | `OMNI_MEMORY_OPERATION_RESERVE_GIB` | Additional per-operation reserve above the soft floor (default `1`) |
804
+ | `OMNI_COMPREHENSION_PRESSURE_GRACE_SECONDS` | Continuous low-headroom interval before a context downshift (default `8`) |
805
+ | `OMNI_COMPREHENSION_RUNTIME_RESIZE` | Restart llama.cpp to change its allocated KV tier at runtime; defaults off on Tegra to avoid unsafe JetPack GPU-device teardown and on elsewhere |
806
+ | `OMNI_COMPREHENSION_EXPANSION_GRACE_SECONDS` | Continuous idle-surplus interval before a one-tier expansion (default `60`) |
807
+ | `OMNI_COMPREHENSION_EXPANSION_COOLDOWN_SECONDS` | Minimum delay after a failed/pressured tier before retrying it (default `900`) |
808
+ | `OMNI_CONTEXT_FILE` | Optional complete `robit.omni.context.v1` catalog override; defaults to the packaged context catalog |
809
+ | `OMNI_CALL_LOG_CONTENT` | Opt in to exact structured heard/generated/TTS traces; disabled by default |
810
+ | `OMNI_UPDATE_INTERVAL_SECONDS` | Indicator Git update polling interval; minimum 60 seconds, default 900 |
811
+
812
+ Keep `OMNI_TTS_PERSISTENT=1` only when speech and comprehension genuinely fit
813
+ together. On constrained unified-memory hosts, use `OMNI_TTS_PERSISTENT=0` and
814
+ set `OMNI_CALL_SPEECH_EVICT_UNIT`: the harness completes hearing, reasoning and
815
+ tools as text, stops comprehension, synthesizes once, lets TTS exit, and restores
816
+ comprehension before listening again. This is slower than resident TTS, but it
817
+ prevents the kernel from overcommitting the machine.
818
+
819
+ Tool resource admission is declared beside each tool in `context.json`.
820
+ Standard work must clear the soft floor, bounded continuations may run within
821
+ the soft-to-hard safety band, executors dynamically distinguish new residency
822
+ from reuse, and control-plane work remains available so stalled work can be
823
+ inspected or cancelled. Every cancellable operation still stops at the hard
824
+ floor. An executor may declare a measured `memory_reserve_gib`; otherwise its
825
+ new residency uses the conservative standard admission boundary.
826
+
827
+ ## Request example
828
+
829
+ Adapter v1 is `robit.ollama.omni-adapter.v1` and its portable route requires
830
+ `stream:false`:
831
+
832
+ ```json
833
+ {
834
+ "model": "robit/qwen3.8-27b-e03-obliterated-omni-audio-bridge:q4km",
835
+ "messages": [{
836
+ "role": "user",
837
+ "content": "What happened, and answer aloud.",
838
+ "audios": [{
839
+ "mime_type": "audio/wav",
840
+ "encoding": "base64",
841
+ "data": "<16 kHz mono PCM16 RIFF/WAVE>"
842
+ }]
843
+ }],
844
+ "omni": {
845
+ "schema": "robit.ollama.omni-adapter.v1",
846
+ "task": "chat"
847
+ },
848
+ "response_modalities": ["text", "audio"],
849
+ "speech_mode": "always",
850
+ "think": false,
851
+ "stream": false
852
+ }
853
+ ```
854
+
855
+ Speech returns under `message.audio` as a tagged base64 RIFF/WAVE envelope.
856
+ Transcripts, non-speech acoustic observations, and visual observations remain
857
+ separate so environmental sounds are never misrouted as the user's words.
858
+
859
+ ## Runtime guarantees
860
+
861
+ - Every comprehension request sets `cache_prompt:false`; a prior audio/video
862
+ embedding cannot be reused for a new clip.
863
+ - Media turns send only the current attachment as present-tense perceptual
864
+ evidence while retaining bounded prior text dialogue for natural continuity.
865
+ - The portal defaults to one active GPU lane and four admitted active/queued
866
+ requests, with request-local media, tools, voice settings, and streams.
867
+ - A wrench toggle, off by default, exposes server-pinned structured tools
868
+ for local-browser public-web discovery/fetch, attached-document retrieval,
869
+ current time/capabilities, on-demand host snapshots, and temporary session
870
+ web/memory recall and isolated text-only sub-agent delegation. Productive
871
+ tool chains continue until a final answer, while exact duplicates and
872
+ repeated nonproductive rounds stop safely; live collapsible execution
873
+ evidence appears in the response and phone UI. No hosted search API is used.
874
+ - Same-origin IndexedDB restores messages, drafts, pending attachments, reply
875
+ audio, and bounded image/video previews after reload. It is keyed by a
876
+ one-way cookie-derived scope, begins a five-minute expiry on page leave, and
877
+ is deleted immediately by trash. Restored media is display-only and is never
878
+ submitted automatically. The server has no shared model conversation state.
879
+ The document index follows the same session partition and expiry policy.
880
+ - Long speech is split before the per-generation codec-frame ceiling, streamed
881
+ with continuous sequence numbers, and assembled into one complete final WAV.
882
+ - Trained-bridge runtime keeps the matching shipped Qwen3-TTS voice profile
883
+ resident alongside comprehension on its
884
+ assigned GPU and emits two codec frames (about 160 ms) per stream window by
885
+ default. A voice-profile change intentionally replaces the resident worker.
886
+ Non-persistent workers and explicit residency handoff remain legacy/manual
887
+ escape hatches. Guided startup does not run a blocking generation gate.
888
+ - Ordinary turns receive only a compact stable behavioral system policy. With
889
+ tools enabled, `get_system_snapshot` can explicitly sample current date/time,
890
+ OS/architecture, CPU/load, RAM, interface counters, and NVIDIA utilization.
891
+ It excludes hostnames, addresses, routes, sockets, processes, credentials,
892
+ and session content, and describes the portal host—not the user's device.
893
+ - Call cognition is single-flight per browser call. Rapid confirmed segments
894
+ merge into one bounded latest-turn buffer; continuing speech cancels stale
895
+ unanswered inference and preserves its input instead of filling the server
896
+ queue.
897
+ - Sound-only call captures stop after comprehension, render as dim **Audio
898
+ context**, and retain at most six bounded environmental observations for the
899
+ next actual spoken turn. They never invoke language or TTS.
900
+ - `get_user_location` uses a browser-side HTTPS lookup and exposes only
901
+ sanitized, approximate session geography. Raw IP and network metadata never
902
+ reach the portal or model.
903
+ - Adapter responses state current media modalities and whether current visual
904
+ input exists. Location/search/fetch results carry tool/source authority and
905
+ claim limits, so tool evidence cannot legitimately be presented as vision.
906
+ - Reasoning is off until the client sends native `think:true`. Thinking is
907
+ returned separately and is never synthesized.
908
+ - CUDA media inference has no CPU fallback. Broker allocation and exact UUID
909
+ residency are deployment gates on the managed host; direct NVIDIA mode also
910
+ verifies comprehension and every TTS PID with `nvidia-smi`. On an NVIDIA
911
+ Tegra module, whose driver publishes no compute-app accounting at all, the
912
+ same gate is proven from each worker's own integrated-GPU device handles.
913
+ - Session diagnostics are content-redacted, partitioned by an opaque cookie,
914
+ deleted by the trash control, and expire five minutes after a client leaves.
915
+
916
+ ## Repository map
917
+
918
+ | Path | Purpose |
919
+ |---|---|
920
+ | `src/qwen_omni_adapters/` | Wire contract, audio validation, GGUF views, Ollama sidecar resolver, accelerator/residency probes, CLI |
921
+ | `runtime/adapter_server.py` | Unified comprehension → language → optional TTS router |
922
+ | `runtime/tts_server.py` | CUDA-only Qwen3-TTS wrapper and PCM stream endpoint |
923
+ | `portal/` | Authenticated phone UI, proxy, supervisor, safe tools, persistent-task store, smoke tests, VAD harness |
924
+ | `harness/` | Always-listening local call mode: VAD, audio I/O, cameras, ReSpeaker, deferred memory storage, background agent, top-bar indicator |
925
+ | `clients/` | Minimal Python and JavaScript request examples |
926
+ | `docs/` | Protocol, architecture, runtime, deployment, ABI, testing, release evidence |
927
+ | `patches/` | Pinned llama.cpp Qwen3-TTS streaming and persistent-worker patches |
928
+ | `scripts/` | Bootstrap, build, validation, and scoped cleanup |
929
+ | `tests/` | Contract, routing, isolation, diagnostics, GGUF, and portal regression tests |
930
+
931
+ ## Documentation
932
+
933
+ - [Virtual context](docs/virtual-context.md) — lossless long-history storage,
934
+ recursive retrieval, exact replay, and the bounded 16K working-set contract.
935
+
936
+ - [Architecture and ownership](docs/architecture.md)
937
+ - [Agent runbook](docs/agent-runbook.md)
938
+ - [Runtime guide](docs/runtime.md)
939
+ - [Context engineering and tool/phase map](docs/context-engineering.md)
940
+ - [Verified model profiles](docs/model-profiles.md)
941
+ - [Phone deployment](docs/phone-portal.md)
942
+ - [Linux, macOS, and Windows services](docs/services.md)
943
+ - [arm64 and NVIDIA Jetson](docs/arm-jetson.md)
944
+ - [Wire protocol](docs/protocol.md)
945
+ - [Portal tools and tool chaining](docs/tools.md)
946
+ - [GGUF/Ollama sidecar ABI](docs/gguf-abi.md)
947
+ - [Testing](docs/testing.md)
948
+ - [Deferred component candidates](docs/candidate-components.md)
949
+ - [Cleanup and storage safety](docs/cleanup.md)
950
+ - [Security model](SECURITY.md)
951
+
952
+ The original implementation remains in
953
+ [`robit-man/fine_tuning_suite`](https://github.com/robit-man/fine_tuning_suite)
954
+ for build and model-development workflows. This repository is the smaller,
955
+ stable runtime and integration baseline.
956
+
957
+ For the complete edge device, enclosure, peripherals, telepresence, and
958
+ orchestration project, see
959
+ [`robit-man/EGG`](https://github.com/robit-man/EGG).