omnindicator 0.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +201 -0
- package/README.md +959 -0
- package/docs/npm-installer.md +91 -0
- package/npm/omnindicator/bin/omnindicator.js +11 -0
- package/npm/omnindicator/lib/cli.js +267 -0
- package/npm/omnindicator/lib/host.js +207 -0
- package/npm/omnindicator/lib/installer.js +178 -0
- package/npm/omnindicator/lib/models.js +114 -0
- package/package.json +40 -0
package/README.md
ADDED
|
@@ -0,0 +1,959 @@
|
|
|
1
|
+
# Qwen Omni Adapters
|
|
2
|
+
|
|
3
|
+
Standalone runtime, protocol, and deployment tooling for logical Ollama Omni
|
|
4
|
+
models. The guided Jetson deployer offers these verified reduced profiles:
|
|
5
|
+
|
|
6
|
+
```text
|
|
7
|
+
robit/qwen3.8-27b-e03-obliterated-omni-audio-bridge:q4km
|
|
8
|
+
robit/ornith-1.5-omni-audio-bridge:q4km
|
|
9
|
+
```
|
|
10
|
+
|
|
11
|
+
The repository turns that one Ollama tag into one authenticated, Ollama-shaped
|
|
12
|
+
API for text, tools, optional thinking, images, audio/ASR, environmental sound
|
|
13
|
+
analysis, video understanding, and Qwen3-TTS speech. It also includes the
|
|
14
|
+
phone-first validation portal used to exercise microphone, camera, allowlisted
|
|
15
|
+
Female/Male voice presets, request-local voice clone,
|
|
16
|
+
streamed playback, call mode, and concurrent isolated sessions.
|
|
17
|
+
|
|
18
|
+
For a host that should simply listen, `harness/` runs that same call mode
|
|
19
|
+
locally: microphone in, speakers out, state in the desktop's top bar, no
|
|
20
|
+
browser involved. See [Always-listening call harness](#always-listening-call-harness).
|
|
21
|
+
|
|
22
|
+
This runtime is also the speech, perception, and local-agent integration used
|
|
23
|
+
by [EGG — Experimental Generalized Gateway](https://github.com/robit-man/EGG),
|
|
24
|
+
Robit's open-source edge-AI hardware and software platform. EGG is the larger
|
|
25
|
+
robot/peripheral system; this repository is the independently installable Omni
|
|
26
|
+
model runtime.
|
|
27
|
+
|
|
28
|
+
## npm guided installer
|
|
29
|
+
|
|
30
|
+
The published `omnindicator` package provides a guarded one-command entry point
|
|
31
|
+
for a new host:
|
|
32
|
+
|
|
33
|
+
```bash
|
|
34
|
+
npx omnindicator@latest
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
It detects the platform, architecture, Tegra unified-memory versus discrete
|
|
38
|
+
NVIDIA topology, broker availability, host/GPU memory class, free disk space,
|
|
39
|
+
desktop-session signal, and required tools **before** cloning this repository
|
|
40
|
+
or pulling model weights. It then selects one of the two trained bridge
|
|
41
|
+
profiles and hands off to the native deployment script. On Linux the default
|
|
42
|
+
is the core service plus the visible always-listening AppIndicator harness;
|
|
43
|
+
only `--core-only` opts out.
|
|
44
|
+
|
|
45
|
+
The package does not pretend that platform support is identical. macOS and
|
|
46
|
+
Windows can install the accelerated core runtime, managed service, and portal,
|
|
47
|
+
but this repository does not yet ship a native tray indicator for either, so
|
|
48
|
+
the installer discloses that boundary and requires a core-only acknowledgement.
|
|
49
|
+
It also does not bundle weights or bypass the deployer's exact live-memory,
|
|
50
|
+
GPU-residency, readiness, and rollback checks. See the complete
|
|
51
|
+
[npm installer scope and capacity policy](docs/npm-installer.md).
|
|
52
|
+
|
|
53
|
+
For a persistent command instead of `npx`:
|
|
54
|
+
|
|
55
|
+
```bash
|
|
56
|
+
npm install --global omnindicator
|
|
57
|
+
omnindicator
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
## Agent quick start
|
|
61
|
+
|
|
62
|
+
An automation agent starting from a clean checkout should read `AGENTS.md`,
|
|
63
|
+
select a model profile, bootstrap, run the deployment doctor, validate, and
|
|
64
|
+
only then start services. Do not copy a CUDA configuration from another host:
|
|
65
|
+
the launcher distinguishes broker-managed discrete GPUs from unified-memory
|
|
66
|
+
NVIDIA Tegra systems.
|
|
67
|
+
|
|
68
|
+
Required host tools are Python 3.10+, Node.js, Git, CMake, a working NVIDIA
|
|
69
|
+
CUDA or Apple Metal toolchain, FFmpeg, and a running Ollama installation.
|
|
70
|
+
Cloudflared is optional; without it the portal remains available only on
|
|
71
|
+
loopback.
|
|
72
|
+
|
|
73
|
+
### Guided clone and install
|
|
74
|
+
|
|
75
|
+
Clone the repository and run the guided installer. On a Jetson it detects the
|
|
76
|
+
Tegra SoC, unified-memory size, existing runtime and managed-service state,
|
|
77
|
+
then presents arrow-key menus for install/upgrade, model, and local voice
|
|
78
|
+
harness. The Enter-through/default path always enables its desktop indicator;
|
|
79
|
+
core-only deployment requires the explicit `--no-harness` opt-out:
|
|
80
|
+
|
|
81
|
+
```bash
|
|
82
|
+
git clone https://github.com/robit-man/qwen-omni-adapters.git
|
|
83
|
+
cd qwen-omni-adapters
|
|
84
|
+
./deploy.sh
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
The installer pulls the selected logical Ollama tag, validates its trained
|
|
88
|
+
audio-bridge sidecar, builds or upgrades the runtime, runs doctor and regression
|
|
89
|
+
gates, persists the exact profile, and installs/restarts the systemd service.
|
|
90
|
+
The default desktop harness deployment also installs the Ubuntu
|
|
91
|
+
GTK/AppIndicator and PulseAudio client bindings, waits for a real top-bar
|
|
92
|
+
indicator, and does not declare the harness ready until it reads a microphone
|
|
93
|
+
frame. It does not require a separate adjacent language-model download.
|
|
94
|
+
|
|
95
|
+
### Validate and start
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
export OMNI_MODEL=robit/ornith-1.5-omni-audio-bridge:q4km
|
|
99
|
+
export OMNI_LANGUAGE_MODEL=$OMNI_MODEL
|
|
100
|
+
cat AGENTS.md
|
|
101
|
+
.venv/bin/qwen-omni doctor --deployment
|
|
102
|
+
./scripts/validate.sh
|
|
103
|
+
./portal/start.sh --daemon
|
|
104
|
+
./portal/start.sh --status
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
`doctor` must report the intended accelerator and no missing required
|
|
108
|
+
components. Validation must finish green. `start.sh --status` prints component
|
|
109
|
+
health and the portal URL; never publish the URL's `#access=...` fragment,
|
|
110
|
+
because the fragment is the portal credential.
|
|
111
|
+
|
|
112
|
+
To install managed Linux services, including the always-listening desktop
|
|
113
|
+
harness:
|
|
114
|
+
|
|
115
|
+
```bash
|
|
116
|
+
./services/linux/install.sh --with-harness
|
|
117
|
+
```
|
|
118
|
+
|
|
119
|
+
Use `./portal/start.sh --stop` for a foreground/staged deployment, or the
|
|
120
|
+
platform service manager after a service install. Do not delete
|
|
121
|
+
`runtime-data/components` while any worker is running.
|
|
122
|
+
|
|
123
|
+
### One-line install and launch
|
|
124
|
+
|
|
125
|
+
For non-interactive automation, the short profile names select the same two
|
|
126
|
+
bridge releases and deploy the managed service:
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
git clone https://github.com/robit-man/qwen-omni-adapters.git && cd qwen-omni-adapters && ./deploy.sh ornith15
|
|
131
|
+
```
|
|
132
|
+
|
|
133
|
+
Available profiles are:
|
|
134
|
+
|
|
135
|
+
```bash
|
|
136
|
+
./deploy.sh ornith15 # standard Ornith 1.5 9B bridge; about 8.15 GiB
|
|
137
|
+
./deploy.sh qwen38 # Qwen3.8 27B E03 bridge; about 18.33 GiB
|
|
138
|
+
```
|
|
139
|
+
|
|
140
|
+
On a broker-managed GPU host the installed service uses the scoped broker. On
|
|
141
|
+
Tegra it uses the direct supervisor and proves residency from `nvgpu`/`nvmap`
|
|
142
|
+
handles. The Ornith bridge is the recommended 32 GB Jetson choice. Qwen's
|
|
143
|
+
18.33 GiB artifact set fits nominally, but an eviction-free production peak is
|
|
144
|
+
not claimed until measured on the target board with its real context and TTS
|
|
145
|
+
policy.
|
|
146
|
+
|
|
147
|
+
Upgrades perform a live-runtime handoff before replacing the unit: the
|
|
148
|
+
deployer identifies and stops recognized old Omni listeners, unloads their
|
|
149
|
+
relevant Ollama runners, verifies the ports are free, and samples Jetson GPU
|
|
150
|
+
load and unified-memory headroom. The old configuration, unit, and managed
|
|
151
|
+
services are restored if the replacement does not reach ready state.
|
|
152
|
+
|
|
153
|
+
The first run creates `.venv`, installs the Python package, clones a pinned
|
|
154
|
+
llama.cpp revision, applies the Qwen3-TTS PCM streaming and resident-worker
|
|
155
|
+
patches, builds the two
|
|
156
|
+
CUDA binaries, pulls the selected Ollama tag, validates the attached sidecar,
|
|
157
|
+
materializes its disposable TTS views, installs the managed service, runs local
|
|
158
|
+
smoke gates, and records the authenticated portal URL in the protected daemon
|
|
159
|
+
status.
|
|
160
|
+
|
|
161
|
+
For a staged installation:
|
|
162
|
+
|
|
163
|
+
```bash
|
|
164
|
+
./scripts/bootstrap.sh
|
|
165
|
+
.venv/bin/qwen-omni doctor --deployment
|
|
166
|
+
./portal/start.sh --daemon
|
|
167
|
+
./portal/start.sh --status
|
|
168
|
+
./portal/start.sh --stop
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
On an arm64 NVIDIA Jetson, the launcher selects the direct managed service
|
|
172
|
+
because a Tegra module has an integrated GPU and no GPU broker. Run without a
|
|
173
|
+
profile to choose interactively:
|
|
174
|
+
|
|
175
|
+
```bash
|
|
176
|
+
./deploy.sh
|
|
177
|
+
```
|
|
178
|
+
|
|
179
|
+
See [arm64 and NVIDIA Jetson](docs/arm-jetson.md) for build architecture
|
|
180
|
+
pinning, residency evidence, and unified-memory guidance.
|
|
181
|
+
|
|
182
|
+
Platform service installs are also one command after cloning:
|
|
183
|
+
|
|
184
|
+
```bash
|
|
185
|
+
./deploy-macos.sh # macOS Metal + launchd
|
|
186
|
+
```
|
|
187
|
+
|
|
188
|
+
```powershell
|
|
189
|
+
.\deploy.ps1 # Windows CUDA + managed user task
|
|
190
|
+
.\deploy.ps1 -Mode Service # true pywin32 Windows Service
|
|
191
|
+
```
|
|
192
|
+
|
|
193
|
+
Do not expose the URL including its `#access=...` fragment publicly. The
|
|
194
|
+
fragment is the portal credential.
|
|
195
|
+
|
|
196
|
+
## Current implementation state
|
|
197
|
+
|
|
198
|
+
The guided deployment uses a trained audio bridge: the selected Qwen3.8 or
|
|
199
|
+
standard Ornith trunk handles audio/ASR, native vision, language, and tools in
|
|
200
|
+
one llama.cpp server, while Qwen3-TTS provides speech. Legacy full-Omni bundles
|
|
201
|
+
remain supported and use Qwen3-Omni for media comprehension plus a selected
|
|
202
|
+
language base. The previously validated legacy 32 GB AGX Orin deployment uses
|
|
203
|
+
a tighter residency profile because all three graphs cannot safely coexist in
|
|
204
|
+
its 29.98 GiB unified-memory pool:
|
|
205
|
+
|
|
206
|
+
| Component | Current constrained-host role | Residency |
|
|
207
|
+
|---|---|---|
|
|
208
|
+
| Qwen3-Omni + projector | Speech/audio/image/video comprehension **and** language/tool reasoning through its OpenAI-compatible endpoint | Resident; context chosen from live memory and published to the adapter; post-speech recovery live-validated at 4K and 8K under desktop load |
|
|
209
|
+
| Ornith 1.5 base | Logical release/base option, but not loaded by the constrained profile | Not resident |
|
|
210
|
+
| Qwen3-TTS + codec projector | Final 24 kHz PCM16 speech | Loaded only after text/tools finish; exits after the utterance |
|
|
211
|
+
| Nomic text embedder | Passive semantic conversation memory | Admitted only while foreground, background-agent, TTS, and restoration work are idle and live headroom permits |
|
|
212
|
+
|
|
213
|
+
On that profile the harness finishes comprehension, reasoning, and tool calls
|
|
214
|
+
before TTS, stops the comprehension service to make room, streams decoder PCM
|
|
215
|
+
with a small startup lead, and starts restoring comprehension while audio is
|
|
216
|
+
still playing. If the user interrupts, playback ducks and then pauses/fades;
|
|
217
|
+
the microphone remains active throughout. This is a safe memory arrangement,
|
|
218
|
+
not the theoretical minimum-latency arrangement. Hosts with enough memory keep
|
|
219
|
+
matching TTS and language workers resident.
|
|
220
|
+
|
|
221
|
+
The essential constrained-host overrides are:
|
|
222
|
+
|
|
223
|
+
```bash
|
|
224
|
+
# Run Qwen3-Omni as an independently supervised worker on port 8901.
|
|
225
|
+
export OMNI_ENABLE_COMPREHENSION=0
|
|
226
|
+
export OMNI_COMPREHENSION_URL=http://127.0.0.1:8901/v1/chat/completions
|
|
227
|
+
|
|
228
|
+
# Reuse that resident worker for language instead of loading Ornith beside it.
|
|
229
|
+
export OMNI_LANGUAGE_API=openai
|
|
230
|
+
export OMNI_LANGUAGE_URL=http://127.0.0.1:8901/v1/chat/completions
|
|
231
|
+
|
|
232
|
+
# Do not retain TTS beside comprehension on a 29.98 GiB pool.
|
|
233
|
+
export OMNI_TTS_PERSISTENT=0
|
|
234
|
+
export OMNI_CALL_SPEECH_EVICT_UNIT=egg-omni-comprehension.service
|
|
235
|
+
export OMNI_CALL_COMPREHENSION_HEALTH=http://127.0.0.1:8901/health
|
|
236
|
+
```
|
|
237
|
+
|
|
238
|
+
The comprehension service uses `runtime/comprehension_launcher.py` rather
|
|
239
|
+
than a fixed `-c` value. Set a service `MemoryMax=` as an independent final
|
|
240
|
+
guard; live admission and automatic downshift remain the primary mechanism.
|
|
241
|
+
|
|
242
|
+
Current local-call behavior also includes:
|
|
243
|
+
|
|
244
|
+
- model-requested camera capture rather than transcript keyword heuristics;
|
|
245
|
+
- a compact discovery tool plus on-demand schemas, with raw shell available to
|
|
246
|
+
the trusted local harness;
|
|
247
|
+
- screenshot-grounded control of a real, visible Chromium window and the full
|
|
248
|
+
Ubuntu desktop, with every action followed by fresh visual evidence;
|
|
249
|
+
- crash-safe persistent background tasks that checkpoint after each
|
|
250
|
+
bounded inference/tool slice, finalize only through referenced tool evidence,
|
|
251
|
+
yield cancellable inference to foreground speech, accept later spoken
|
|
252
|
+
guidance, keep knowledge/environment/controller state outside renewable model
|
|
253
|
+
transcripts, and use native thinking without speaking or storing it;
|
|
254
|
+
- bounded conversational history with time-based relevance reduction and
|
|
255
|
+
explicit current-query memory tools rather than next-turn prefetch;
|
|
256
|
+
- live memory-derived comprehension context selection and automatic
|
|
257
|
+
downshifting rather than a board-specific fixed context size;
|
|
258
|
+
- continuous PCM playback, timing/starvation diagnostics, ReSpeaker echo
|
|
259
|
+
handling, and automatic service/microphone-loop recovery.
|
|
260
|
+
- a resident, typed Laya System-1 decision plane with batched routing,
|
|
261
|
+
pre-action, post-action, and context-relevance waves; it is shadow-only until
|
|
262
|
+
real calibration proves a per-family fast path, and failures always preserve
|
|
263
|
+
the deliberative route. See [`docs/decision-plane.md`](docs/decision-plane.md).
|
|
264
|
+
|
|
265
|
+
## What the model tag contains
|
|
266
|
+
|
|
267
|
+
The release is one logical Ollama model, not one graph that stock Ollama can
|
|
268
|
+
execute end to end:
|
|
269
|
+
|
|
270
|
+
```text
|
|
271
|
+
logical Ollama tag
|
|
272
|
+
├── standard model/projector/template layers
|
|
273
|
+
│ └── selected Qwen-family base: text, image vision, tools, optional thinking
|
|
274
|
+
└── application/vnd.robit.ollama.omni.bundle.v1+gguf
|
|
275
|
+
├── Qwen3-Omni comprehension model + projector
|
|
276
|
+
└── Qwen3-TTS model + codec/projector
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
Stock Ollama handles the standard layers. This adapter resolves the custom
|
|
280
|
+
sidecar layer, reconstructs byte-preserving executable component views, and
|
|
281
|
+
runs audio/video comprehension and TTS with the pinned llama.cpp build. The
|
|
282
|
+
public request remains Ollama-shaped and names the one logical tag.
|
|
283
|
+
|
|
284
|
+
The runtime also accepts the reduced `robit.ollama-audio-bridge.v1` profile.
|
|
285
|
+
There the standard projector combines target-native vision with the frozen
|
|
286
|
+
Omni audio encoder and its trained final projection. The sidecar carries only
|
|
287
|
+
TTS, and one local llama.cpp server is both the multimodal comprehension path
|
|
288
|
+
and the sole language/tool trunk. This removes the full secondary Omni Thinker
|
|
289
|
+
and avoids loading an adjacent Ollama language copy.
|
|
290
|
+
|
|
291
|
+
Legacy full-Omni tags remain semantic routers because Qwen3.8, Qwen3-Omni, and
|
|
292
|
+
Qwen3-TTS do not share compatible hidden-state interfaces. The trained bridge
|
|
293
|
+
profiles are different: they contain the frozen Omni audio encoder plus a
|
|
294
|
+
trained final projection into the selected trunk, while keeping evidence tags
|
|
295
|
+
and TTS as explicit runtime boundaries. Standard Ornith and Qwen3.8 E03 tensors
|
|
296
|
+
and Ollama tags are never interchangeable.
|
|
297
|
+
|
|
298
|
+
## Architecture breakdown
|
|
299
|
+
|
|
300
|
+
### End-to-end data plane
|
|
301
|
+
|
|
302
|
+
```text
|
|
303
|
+
browser / API client local always-listening harness
|
|
304
|
+
│ │
|
|
305
|
+
└──────────────────┬───────────────────────────────┘
|
|
306
|
+
▼
|
|
307
|
+
authenticated portal (optional)
|
|
308
|
+
session isolation · admission queue
|
|
309
|
+
tool execution · diagnostics · NDJSON relay
|
|
310
|
+
│
|
|
311
|
+
▼
|
|
312
|
+
unified adapter API
|
|
313
|
+
│
|
|
314
|
+
┌────────────────┴─────────────────┐
|
|
315
|
+
│ │
|
|
316
|
+
text-only request current media request
|
|
317
|
+
│ audio · image · video
|
|
318
|
+
│ │
|
|
319
|
+
│ validation / normalization
|
|
320
|
+
│ │
|
|
321
|
+
│ ▼
|
|
322
|
+
│ Qwen3-Omni comprehension
|
|
323
|
+
│ │
|
|
324
|
+
│ tagged, untrusted evidence
|
|
325
|
+
│ ┌────────────┼────────────┐
|
|
326
|
+
│ │ │ │
|
|
327
|
+
│ transcript sound context visual evidence
|
|
328
|
+
│ └────────────┼────────────┘
|
|
329
|
+
└──────────────────────────────────┘
|
|
330
|
+
▼
|
|
331
|
+
selected language backend
|
|
332
|
+
Qwen3.8 · Ornith · Qwen3-Omni
|
|
333
|
+
│
|
|
334
|
+
┌──────────────┴──────────────┐
|
|
335
|
+
│ unresolved structured calls?│
|
|
336
|
+
▼ │
|
|
337
|
+
portal tool/discovery loop ─────────────┘
|
|
338
|
+
│
|
|
339
|
+
▼
|
|
340
|
+
final answer text
|
|
341
|
+
│
|
|
342
|
+
speech requested and safe?
|
|
343
|
+
▼
|
|
344
|
+
Qwen3-TTS
|
|
345
|
+
│
|
|
346
|
+
ordered 24 kHz mono PCM16
|
|
347
|
+
```
|
|
348
|
+
|
|
349
|
+
The public model name remains the logical Ollama tag throughout. Component
|
|
350
|
+
URLs, process placement, model eviction, and derived GGUF views are deployment
|
|
351
|
+
details hidden behind the adapter.
|
|
352
|
+
|
|
353
|
+
### Stage ownership
|
|
354
|
+
|
|
355
|
+
| Stage | Owner | Important invariant |
|
|
356
|
+
|---|---|---|
|
|
357
|
+
| Request parsing | `qwen_omni_adapters.contract` | Bounds and validates media before inference; preserves native `think`, tools, options, and response modalities |
|
|
358
|
+
| Audio/image/video preparation | `runtime/adapter_server.py` + FFmpeg where needed | Only the newest attachment is current evidence; video duration/frame count and decoded audio are bounded |
|
|
359
|
+
| Media comprehension | Qwen3-Omni `llama-server` | Prompt caching is disabled so stale multimodal embeddings cannot cross turns |
|
|
360
|
+
| Semantic bridge | Adapter-generated tagged observation | Speech transcript, non-speech acoustics, and visual evidence stay separate and remain untrusted data |
|
|
361
|
+
| Language, reasoning, tool choice | Selected Ollama base or configured OpenAI-compatible worker | Reasoning remains in `message.thinking`; unresolved tool calls cannot enter TTS |
|
|
362
|
+
| Tool execution | Authenticated portal | Starts with compact discovery, exposes only relevant concrete schemas, records bounded evidence, and rejects repeated nonproductive calls |
|
|
363
|
+
| Computer use | `portal/browser.py` + `portal/gui.py` | Opens visible Chromium on the desktop, observes rendered screenshots, clicks/types through native DevTools input, and can see/control the wider workspace through `xdotool` plus fresh desktop screenshots |
|
|
364
|
+
| Speech | Patched Qwen3-TTS worker | Emits ordered decoder PCM; generation state is reset between prompts and never leaks one utterance into the next |
|
|
365
|
+
| Local conversation | `harness/` | VAD, interruption, ReSpeaker state/direction, camera capture, history, deferred memory writes, and foreground scheduling |
|
|
366
|
+
| Persistent work | `harness/background_agent.py` + `portal/background_tasks.py` | Durable checkpoints and leases survive process restarts; long jobs may yield sparse spoken milestones, terminal speech is durable, and foreground speech wins every scheduling boundary |
|
|
367
|
+
|
|
368
|
+
### A normal spoken turn
|
|
369
|
+
|
|
370
|
+
1. The microphone loop continuously captures audio. Adaptive VAD accepts real
|
|
371
|
+
near-end speech, joins brief continuation segments, and keeps the newest
|
|
372
|
+
bounded turn rather than filling an inference queue.
|
|
373
|
+
2. The harness submits one audio-bearing request. Qwen3-Omni produces a tagged
|
|
374
|
+
transcript and any non-speech observation; silence or a cough stops before
|
|
375
|
+
language and TTS when no speech was found.
|
|
376
|
+
3. The configured language backend receives the recognized speech as the latest
|
|
377
|
+
user message, with bounded text history and non-speech observations as
|
|
378
|
+
secondary tagged evidence. Empty negative acoustic boilerplate such as “no
|
|
379
|
+
non-speech sounds” remains available in adapter diagnostics but is omitted
|
|
380
|
+
from the language prompt so it cannot contradict a valid transcript. Visual evidence is added only after a structured
|
|
381
|
+
camera request, or when the client explicitly attached media. The
|
|
382
|
+
local voice harness enables the portal's real tool allowlist on every turn
|
|
383
|
+
without a separate classification pass: the model either answers
|
|
384
|
+
conversationally or calls the smallest tool that accomplishes the request,
|
|
385
|
+
the portal auto-executes it, and the spoken answer is the grounded reply that
|
|
386
|
+
follows the completed work. On the constrained 32 GB profile this is the same
|
|
387
|
+
resident Qwen3-Omni worker; a normal profile uses the selected Ollama base.
|
|
388
|
+
4. A durable-task request goes through the same pass via the `background_task`
|
|
389
|
+
tool, which writes the objective to the durable store before any
|
|
390
|
+
acknowledgment. Camera-enabled embodied turns expose a capture bridge but do
|
|
391
|
+
not attach an ambient still unless the request needs current physical-scene
|
|
392
|
+
evidence; web and other current-information requests complete a real tool
|
|
393
|
+
call in the pass. No answer is eligible for speech before the work it claims
|
|
394
|
+
has actually completed. Once transcript-aware routing finds concrete matching
|
|
395
|
+
schemas, the language trunk receives their names in a compact required-action
|
|
396
|
+
context and still chooses the tool and arguments itself.
|
|
397
|
+
5. Only final answer text is sent to TTS. PCM is played as decoder windows
|
|
398
|
+
arrive, with one small initial lead to absorb packet jitter rather than
|
|
399
|
+
waiting for the complete WAV.
|
|
400
|
+
6. Conversation persistence and embeddings are queued after answer, tools, and
|
|
401
|
+
speech. They cannot delay the turn.
|
|
402
|
+
|
|
403
|
+
The local harness keeps foreground reasoning disabled by default for natural
|
|
404
|
+
response latency. Persistent background tasks enable native thinking because
|
|
405
|
+
deliberation is more valuable than sub-second response there; the thinking
|
|
406
|
+
channel is never synthesized or written into the task transcript.
|
|
407
|
+
|
|
408
|
+
### Camera and video flow
|
|
409
|
+
|
|
410
|
+
The browser may attach a current image or bounded clip directly. The local
|
|
411
|
+
harness deliberately does not attach a room image to every spoken question:
|
|
412
|
+
doing that biases the model into describing the scene even when the user asked
|
|
413
|
+
about something else. Instead, the model calls `request_camera_view`; the
|
|
414
|
+
harness then captures every configured V4L2 camera at the same moment, stitches
|
|
415
|
+
and downscales them, and performs a grounded multimodal follow-up. A motion
|
|
416
|
+
request captures a bounded clip. Internet uses of words such as “look up” are
|
|
417
|
+
therefore routed by the model to web tools rather than intercepted by a local
|
|
418
|
+
keyword list. The grounded follow-up carries the fresh visual evidence but no
|
|
419
|
+
second camera bridge, preventing a recapture loop. The pre-capture placeholder
|
|
420
|
+
is never spoken, logged as generated dialogue, or kept in conversation history;
|
|
421
|
+
the grounded pass is the sole answer. Broad casual questions get a short casual
|
|
422
|
+
overview, while questions about a particular item or feature stay focused on
|
|
423
|
+
that target instead of inventorying the rest of the scene.
|
|
424
|
+
|
|
425
|
+
### Tools and long-horizon work
|
|
426
|
+
|
|
427
|
+
The browser portal keeps powerful tools opt-in. The trusted local voice harness
|
|
428
|
+
exposes the persistent `background_task` bridge beside tool discovery; its
|
|
429
|
+
worker, rather than the latency-critical spoken pass, owns unrestricted shell.
|
|
430
|
+
This structural split prevents a small model from entering a synchronous shell
|
|
431
|
+
retry loop and stranding the conversation. A request that inspects or mutates
|
|
432
|
+
files or the system, needs verification or retry, or spans multiple commands is
|
|
433
|
+
accepted into a crash-safe JSON store and acknowledged immediately, after which
|
|
434
|
+
the background worker:
|
|
435
|
+
|
|
436
|
+
```text
|
|
437
|
+
claim lease → reason once → execute/discover one or more tools
|
|
438
|
+
→ inspect results → checkpoint → yield to speech → continue
|
|
439
|
+
```
|
|
440
|
+
|
|
441
|
+
Shell commands return stdout, stderr, exit status, timeouts, and truncation
|
|
442
|
+
markers. Unsupported “done” claims are rejected until at least one concrete
|
|
443
|
+
action produced evidence. The worker has no fixed step horizon. On a genuinely
|
|
444
|
+
long task it may pause after a meaningful verified milestone, speak a short
|
|
445
|
+
progress explanation, then resume with the same task context. A later spoken
|
|
446
|
+
update is appended as authoritative task guidance before the next step,
|
|
447
|
+
including a final race check so stale completion cannot beat a new instruction.
|
|
448
|
+
Typed focus records advertise no paging operation while their detailed receipts
|
|
449
|
+
are still resident. After compaction they gain structured pointers to the
|
|
450
|
+
separate `task_expand` control function; that control name is never presented as
|
|
451
|
+
an action or argument of `workspace_file` or another external tool. The focus
|
|
452
|
+
ledger is a live working set, not another copy of history: its size follows the
|
|
453
|
+
currently resident comprehension window, repeated inspections and failures are
|
|
454
|
+
coalesced, a changed artifact supersedes its older resident version, and omitted
|
|
455
|
+
record counts remain visible. Exact receipts and every superseded version remain
|
|
456
|
+
append-only in task evidence storage. A compacted request carries this ledger
|
|
457
|
+
once in the pinned system contract rather than duplicating it in a recurrent
|
|
458
|
+
checkpoint. If an omitted record's evidence ID is no longer resident,
|
|
459
|
+
`task_expand` also accepts a distinctive exact path, URL, symbol, error, or
|
|
460
|
+
other query and returns the highest-scoring immutable receipts.
|
|
461
|
+
It is a one-shot page-in, never capability discovery: after either a successful
|
|
462
|
+
expansion or `evidence_not_found`, it leaves the next action surface until a new
|
|
463
|
+
concrete result makes older evidence relevant again. Missing tool schemas route
|
|
464
|
+
through `tool_search`, preventing a failed lexical page query from becoming a
|
|
465
|
+
maintenance loop.
|
|
466
|
+
The ledger applies the same evidence authority as checkpoints: search results
|
|
467
|
+
marked discovery-only never become acquired-source successes, and an empty
|
|
468
|
+
`about:blank` browser snapshot remains an inspection. A blank visible browser
|
|
469
|
+
explicitly requires `navigate` to a grounded URL and cannot divert the next
|
|
470
|
+
round into pixel clicking or unrelated capability discovery.
|
|
471
|
+
At the 4K/8K tiers the worker also projects the already-selected tool schema to
|
|
472
|
+
its executable JSON constraints: names, types, enums, required fields, bounds,
|
|
473
|
+
and `additionalProperties` remain exact while repeated prose annotations are
|
|
474
|
+
removed. Only one selected concrete capability is exposed per action round;
|
|
475
|
+
discovery is suppressed until that leaf receives one concrete attempt, then
|
|
476
|
+
returns as the route to a different capability. At constrained tiers discovery
|
|
477
|
+
and evidence expansion alternate instead of crowding the same envelope. A
|
|
478
|
+
background request that still cannot fit the live pack fails explicitly and retries from
|
|
479
|
+
its checkpoint—it never silently falls back to native FIFO truncation.
|
|
480
|
+
Discovery queries describe only the missing interaction mechanism (for example,
|
|
481
|
+
public-web search, file editing, shell execution, or visible-browser control),
|
|
482
|
+
not the task topic. The deterministic router also recognizes product/SaaS,
|
|
483
|
+
comparison, interface, and design research as web discovery, preventing a word
|
|
484
|
+
such as “service” in the subject from accidentally selecting system shell.
|
|
485
|
+
Background web fetches are provenance-bound as well: a target must appear in
|
|
486
|
+
the accepted task/user input or prior search, browser, crawl, or fetch evidence.
|
|
487
|
+
An invented address is rejected before network access and returned beside the
|
|
488
|
+
exact admissible URLs so the next model step can choose a grounded source.
|
|
489
|
+
The pinned task envelope also carries an explicit execution frontier: work
|
|
490
|
+
advances the earliest unmet prerequisite in the user's stated order, and a
|
|
491
|
+
failed downstream probe returns the worker to that prerequisite instead of
|
|
492
|
+
encouraging variants of the same premature verification.
|
|
493
|
+
Completed or blocked work ends with a brief natural spoken status only when the
|
|
494
|
+
live conversation is idle. Detailed reports, evidence IDs, paths, and worker
|
|
495
|
+
self-assessment stay in the indicator and task archive and are never passed
|
|
496
|
+
verbatim to TTS. The pending-delivery flag survives a harness restart, and the
|
|
497
|
+
speech can be interrupted like any other reply. A direct `synthesize` request
|
|
498
|
+
is a literal text-to-speech transport pass: it retains the selected voice-clone
|
|
499
|
+
profile but bypasses conversation policy, tools, documents, and virtual-memory
|
|
500
|
+
indexing/replay so recalled text cannot be appended to the utterance.
|
|
501
|
+
|
|
502
|
+
Rendered computer work is not reduced to a fetched-text corpus. The
|
|
503
|
+
`browser_interact` tool launches a real Chromium window on the active desktop,
|
|
504
|
+
returns both a screenshot and a bounded accessibility/element map, performs
|
|
505
|
+
real pointer/keyboard input, then observes the changed page. `gui_interact`
|
|
506
|
+
extends the same screenshot → action → screenshot loop to Ubuntu. Its
|
|
507
|
+
default screenshot is the active window and its coordinates are relative to
|
|
508
|
+
that image. The runtime captures root-window pixels and crops them to the exact
|
|
509
|
+
X11 geometry used for pointer translation, so window-manager decorations cannot
|
|
510
|
+
offset clicks. Full-screen mode is explicit for panels and workspace navigation;
|
|
511
|
+
when a later pointer call omits its coordinate space, it remains bound to the
|
|
512
|
+
newest returned frame. Active-window clicks fail safely if focus changed after
|
|
513
|
+
observation. Every result also reports whether the coarse visual state materially
|
|
514
|
+
changed, which lets the worker reject a missed click as non-progress.
|
|
515
|
+
Chromium's rendered network-error documents remain visible evidence, but are
|
|
516
|
+
tagged as failed navigation and release the sticky browser action space; a
|
|
517
|
+
connection-refused page can never count as successful GUI verification.
|
|
518
|
+
Actionable DOM controls are always re-resolved through live CDP geometry. For
|
|
519
|
+
canvas, challenge, and image targets on Jetson, the conversational vision pass
|
|
520
|
+
supplies a concise referring expression and an isolated resident Moondream 2
|
|
521
|
+
point head resolves it against the exact CDP viewport. The browser re-captures
|
|
522
|
+
and compares that viewport after point inference before sending input, so a
|
|
523
|
+
coordinate is never carried onto changed pixels. Multiple returned matches are
|
|
524
|
+
disambiguated by the conversational model's coarse current-frame point; the
|
|
525
|
+
point head's structured coordinate remains the executed value. If the point
|
|
526
|
+
head is not configured, the existing bounded crop-refinement loop remains the
|
|
527
|
+
fail-safe fallback rather than widening hit tolerances.
|
|
528
|
+
With the point head active, every visual click requires a concise target phrase.
|
|
529
|
+
If the planner's coarse crop misses, the executor searches the same horizontal
|
|
530
|
+
band in bounded overlapping tiles before widening to a bounded tile grid; the
|
|
531
|
+
point head still selects every executable coordinate.
|
|
532
|
+
An explicit verification snapshot or completed visual click also asks that same
|
|
533
|
+
resident visual worker to transcribe visible status text and exact completion
|
|
534
|
+
identifiers from the exact returned CDP frame. That bounded reading is stored as
|
|
535
|
+
current-frame evidence, while a compact action receipt preserves the exact
|
|
536
|
+
grounded target, so terminal state and action reporting survive image compaction
|
|
537
|
+
without relying on an invented marker or renamed target.
|
|
538
|
+
Screenshot bytes are shown to the multimodal model for one reasoning
|
|
539
|
+
pass and then removed from the durable transcript so long tasks retain visual
|
|
540
|
+
grounding without filling their context with base64. Local HTTP pages and all
|
|
541
|
+
desktop/file tools remain usable offline; public sites naturally require a
|
|
542
|
+
working network.
|
|
543
|
+
|
|
544
|
+
### Conversation state and durable memory
|
|
545
|
+
|
|
546
|
+
The adapter itself is stateless between requests. The browser owns its
|
|
547
|
+
cookie-isolated conversation record; the local harness owns a bounded recent
|
|
548
|
+
text history with time-based falloff. Current media is never replayed from that
|
|
549
|
+
history.
|
|
550
|
+
|
|
551
|
+
Optional durable voice memory stores completed text exchanges in SQLite and
|
|
552
|
+
uses a small semantic encoder rather than forcing chat-model hidden states into
|
|
553
|
+
an embedding role. Storage runs on a daemon worker; retrieval is an explicit
|
|
554
|
+
tool call against the current request, never a result automatically carried
|
|
555
|
+
from one turn into the next. Admission checks
|
|
556
|
+
foreground activity, background-agent work, comprehension readiness, the
|
|
557
|
+
encoder's installed payload, current `MemAvailable`, and the active model's
|
|
558
|
+
measured KV slope. If any check fails, deferred storage waits; conversation
|
|
559
|
+
never waits for it.
|
|
560
|
+
|
|
561
|
+
### Context, residency, and recovery
|
|
562
|
+
|
|
563
|
+
On unified-memory hosts, `runtime/comprehension_launcher.py` reads the installed
|
|
564
|
+
GGUF, derives KV bytes per token, samples live available memory, and chooses the
|
|
565
|
+
largest context tier that fits with the runtime memory reserve. Compact models
|
|
566
|
+
may have a 262,144-token native positional range, but the managed compact
|
|
567
|
+
profiles intentionally expose a 16,384-token resident working-set ceiling.
|
|
568
|
+
The guided Ornith-on-Tegra profile uses its live-qualified q8 KV cache; other
|
|
569
|
+
model/platform pairs retain f16 until they pass the same answer, multimodal,
|
|
570
|
+
voice, tool, and memory-pressure gates. The configured 16K value is a ceiling:
|
|
571
|
+
a longer live action soak selected 8K and then 4K as the rest of the resident
|
|
572
|
+
stack consumed unified memory. The live tier is authoritative.
|
|
573
|
+
Longer history is paged through the lossless virtual-context layer instead of
|
|
574
|
+
preallocating a nominal 256K KV cache. First load uses the conservative complete
|
|
575
|
+
component-byte footprint; later loads also use measured residency. It records
|
|
576
|
+
before/after residency and automatically caps the next load below a tier that
|
|
577
|
+
exits or leaves too little memory. The adapter reads the chosen window per
|
|
578
|
+
request and sheds old history/tool evidence before llama.cpp can reject an
|
|
579
|
+
oversized prompt.
|
|
580
|
+
|
|
581
|
+
On discrete-memory hosts, sampling continues after readiness: a continuous
|
|
582
|
+
low-headroom interval downshifts one tier, and sustained surplus performs the
|
|
583
|
+
inverse only while every inference slot is idle and the cooldown has elapsed.
|
|
584
|
+
On Tegra, the startup-selected llama.cpp process stays pinned for the service
|
|
585
|
+
session because affected JetPack kernels can panic while closing GPU character
|
|
586
|
+
devices during an otherwise controlled worker restart. Prompt budgets remain
|
|
587
|
+
dynamic inside the allocated KV window; new work is admission-gated, and the
|
|
588
|
+
next supervised start reselects its tier from live memory. This avoids process
|
|
589
|
+
churn without reverting to a board-specific context limit.
|
|
590
|
+
|
|
591
|
+
The trained-bridge runtime keeps TTS and comprehension simultaneously resident,
|
|
592
|
+
but guided deployment does not block the desktop on generation probes. It marks
|
|
593
|
+
the core ready after local component health and starts the indicator service as
|
|
594
|
+
soon as the core unit starts. The full ASR/cloned-TTS/co-residency smoke remains
|
|
595
|
+
available as an explicit diagnostic. Legacy/manual constrained profiles may
|
|
596
|
+
still opt into explicit eviction callbacks and non-persistent TTS. Service
|
|
597
|
+
managers restart failed workers; the harness waits
|
|
598
|
+
through token rotation and adapter/comprehension restoration, reopens a failed
|
|
599
|
+
microphone capture loop, and resumes expired background-task leases. On Tegra,
|
|
600
|
+
GPU residency is proven from each process's `nvgpu`/`nvmap` handles rather than
|
|
601
|
+
the unsupported `nvidia-smi` process table.
|
|
602
|
+
|
|
603
|
+
### Trust and isolation boundaries
|
|
604
|
+
|
|
605
|
+
- The portal binds locally, requires a generated bearer capability, and may
|
|
606
|
+
publish an authenticated Cloudflare tunnel. The URL fragment is secret.
|
|
607
|
+
- Browser documents, web cache, memory, location, diagnostics, and UI state are
|
|
608
|
+
cookie-scoped and expire after disconnect or are removed immediately by
|
|
609
|
+
Trash.
|
|
610
|
+
- Media text, OCR, transcripts, web pages, tool results, and model observations
|
|
611
|
+
are evidence, never system instructions or authorization.
|
|
612
|
+
- Host snapshots exclude hostnames, addresses, routes, sockets, processes,
|
|
613
|
+
credentials, and session content.
|
|
614
|
+
- CUDA execution has no silent CPU fallback; startup verifies the platform's
|
|
615
|
+
actual GPU-residency evidence.
|
|
616
|
+
|
|
617
|
+
## Capability map
|
|
618
|
+
|
|
619
|
+
| Capability | Runtime owner | Available through this adapter |
|
|
620
|
+
|---|---|---|
|
|
621
|
+
| Text and Markdown, including responsive GFM tables | Selected Qwen3.8/Ornith Ollama base or configured Qwen3-Omni language worker | Yes |
|
|
622
|
+
| Structured tools | Selected language backend + portal executor | Yes |
|
|
623
|
+
| Portal web/document/session-memory tools | Explicit opt-in allowlisted portal loop with no-key DuckDuckGo HTML discovery | Yes |
|
|
624
|
+
| Visible Chromium interaction | Persistent rendered browser, screenshot + element evidence, click/type/scroll/back | Yes |
|
|
625
|
+
| Full desktop computer use | Fresh active-window screenshots with image-relative input; explicit full-screen workspace mode | Yes |
|
|
626
|
+
| Trusted local shell | Unrestricted execution in the checkpointed voice-task worker; bounded output and timeout | Yes |
|
|
627
|
+
| Persistent background tasks | Checkpointed long-horizon worker with status/update/cancel, sparse spoken milestones, durable terminal speech, and restart recovery | Yes |
|
|
628
|
+
| Passive semantic voice memory | Idle-only encoder worker + SQLite; never gates a foreground turn | Yes |
|
|
629
|
+
| Thinking | Native backend `think` control, separate from answer text and never spoken | Yes |
|
|
630
|
+
| Image understanding | Qwen3.8 or Omni comprehension path | Yes |
|
|
631
|
+
| Speech transcription | Qwen3-Omni comprehension | Yes |
|
|
632
|
+
| Environmental audio interpretation | Qwen3-Omni tagged observation | Yes |
|
|
633
|
+
| Video understanding | Qwen3-Omni bounded `input_video` | Yes |
|
|
634
|
+
| Silent video and animated GIF | FFmpeg probe/normalization → Qwen3-Omni | Yes |
|
|
635
|
+
| PDF/DOCX/text retrieval | Session-isolated portal extraction/index | Yes |
|
|
636
|
+
| Spoken response | Qwen3-TTS, 24 kHz mono PCM16 | Yes |
|
|
637
|
+
| Voice reference cloning | Qwen3-TTS Base speaker embedding path | Yes |
|
|
638
|
+
| Live-call turns | Adaptive VAD + bounded single-flight speech consolidation + streamed text/PCM | Yes |
|
|
639
|
+
| Always-listening local call mode | `harness/`: local mic/speakers, mandatory GNOME top-bar state for its managed desktop unit | Yes |
|
|
640
|
+
| Every camera at once, on request | `harness/camera.py` snaps all V4L2 devices together, stitches and downscales to one image or clip only after a visual-evidence request | Yes |
|
|
641
|
+
| ReSpeaker ring and direction | Used when the array is attached, ignored when it is not | Yes |
|
|
642
|
+
| Tool execution by an external loop | `GET /api/tools`, `POST /api/tools/<name>/call` | Yes |
|
|
643
|
+
| Video generation | No component is shipped | No |
|
|
644
|
+
|
|
645
|
+
## Always-listening call harness
|
|
646
|
+
|
|
647
|
+
`harness/` turns a host that has the model into a host you can talk to. It
|
|
648
|
+
drives the same endpoint and the same request shape as the portal's browser
|
|
649
|
+
call mode, through the machine's own microphone and speakers.
|
|
650
|
+
|
|
651
|
+
```bash
|
|
652
|
+
# The adapter must be running; the harness waits for it either way.
|
|
653
|
+
PYTHONPATH=. .venv/bin/python -m harness
|
|
654
|
+
```
|
|
655
|
+
|
|
656
|
+
It listens continuously, answers out loud, and shows what it is doing in the
|
|
657
|
+
GNOME top bar (`Omni ●` listening, `◉` hearing, `◍` thinking, `▶` speaking).
|
|
658
|
+
The indicator's menu mutes the microphone, toggles tools, reasoning and
|
|
659
|
+
cameras, copies the public link when the portal is published through a tunnel,
|
|
660
|
+
opens a loopback-only live view of every camera stitched left-to-right, and
|
|
661
|
+
cleanly reloads the voice service. It polls `origin/main` in the background and
|
|
662
|
+
shows an **Update available — install and restart** action when a verified
|
|
663
|
+
fast-forward exists. Clicking it preserves untracked local files, refuses
|
|
664
|
+
tracked edits or diverged history, refreshes the runtime environment without
|
|
665
|
+
implicitly downloading or changing model weights, and asks the managed runtime
|
|
666
|
+
and indicator to restart on the new revision. Its
|
|
667
|
+
**Models** submenu lists both compact
|
|
668
|
+
audio bridges and both legacy full bundles. Missing tags expose **Download**
|
|
669
|
+
with live percentage/status text; downloaded tags expose **Activate**, **Load
|
|
670
|
+
into Ollama**, **Unload from Ollama**, and confirmed **Delete local copy**
|
|
671
|
+
actions as applicable. Activation atomically selects the one logical language/
|
|
672
|
+
Omni tag, requests a controlled core restart, and lets the indicator reconnect
|
|
673
|
+
with the new service environment. An active direct-daemon model is labelled
|
|
674
|
+
separately from an optional Ollama runner so the UI never hides a duplicate
|
|
675
|
+
allocation on a 32 GB Jetson. The adjacent **Voice** submenu selects **Default
|
|
676
|
+
female**, **Default male**, or any previously imported custom clone reference.
|
|
677
|
+
**Add custom voice…** opens the native audio file picker, normalizes the owned
|
|
678
|
+
clip to a bounded 16 kHz mono WAV, stores it under untracked `runtime-data`,
|
|
679
|
+
adds it to the persistent list, and selects it for the next reply. Voice-profile
|
|
680
|
+
changes hot-reload in the portal; only the TTS worker changes speaker profile,
|
|
681
|
+
so the language and comprehension weights stay resident. It also appends
|
|
682
|
+
live/recent durable tasks as native submenus, so inspecting a task does not
|
|
683
|
+
close the whole menu. Each
|
|
684
|
+
submenu shows the current-stage spinner, exact bounded tool-call arguments and
|
|
685
|
+
outcomes, retained checkpoints, and terminal result. Long action rows wrap and
|
|
686
|
+
ellipsize within the menu while preserving the complete text in their tooltip,
|
|
687
|
+
so tool arguments cannot widen the indicator beyond the screen. Live tasks expose
|
|
688
|
+
**Cancel task** and every record exposes **Clear task record**. **Clear finished
|
|
689
|
+
tasks** moves all terminal records into the human-readable archive, and **Open
|
|
690
|
+
task archive** opens that log in the desktop editor. A real themed state icon
|
|
691
|
+
sits to the left of `Omni`. A manual invocation may run headless and log instead.
|
|
692
|
+
The managed desktop service fails closed if GTK, the StatusNotifier host, the
|
|
693
|
+
desktop audio server, or a usable microphone/sink is unavailable; it never
|
|
694
|
+
silently leaves an always-listening service running without its indicator.
|
|
695
|
+
|
|
696
|
+
Defaults are chosen for a spoken conversation:
|
|
697
|
+
|
|
698
|
+
- **Reasoning off.** A hidden chain of thought is silence the other person has
|
|
699
|
+
to sit through.
|
|
700
|
+
- **Tools on, in the answering pass.** This is the same server-side chain used
|
|
701
|
+
by the cloudflared portal: one tiny discovery schema rides with the turn,
|
|
702
|
+
only the matching concrete contracts appear on the next tool round, and
|
|
703
|
+
requested calls execute until the model has a grounded final answer. The
|
|
704
|
+
trusted local harness also carries compact shell and persistent-task bridges.
|
|
705
|
+
The full catalog no longer displaces conversation or memory context.
|
|
706
|
+
- **Every camera together, only when relevant.** The audio-only pass may call
|
|
707
|
+
`request_camera_view` when the answer depends on the current physical scene.
|
|
708
|
+
Only then are V4L2 devices opened and probed, all working cameras are snapped
|
|
709
|
+
at the same moment, and their frames are stitched into one left-to-right row.
|
|
710
|
+
The indicator's explicit live-view action uses the same horizontal composition. Clips
|
|
711
|
+
work the same way for temporal questions. Merely starting the harness never
|
|
712
|
+
launches FFmpeg or activates a camera privacy indicator.
|
|
713
|
+
- **Observation does not force speech.** Sound-only events are retained as
|
|
714
|
+
bounded context without language or TTS. The language model may also leave a
|
|
715
|
+
transcribed room utterance unanswered when current evidence shows it was
|
|
716
|
+
addressed elsewhere; if gaze would resolve genuine ambiguity, it can request
|
|
717
|
+
a fresh still before deciding. Empty intentional responses do not invoke TTS.
|
|
718
|
+
- **ReSpeaker when present.** Its ring follows the conversation and the
|
|
719
|
+
direction a voice came from is attached to the turn as evidence. With no
|
|
720
|
+
array attached the default microphone is used and nothing else changes.
|
|
721
|
+
- **Memory storage is passive.** Completed exchanges are embedded on a daemon
|
|
722
|
+
worker only after the answer, tools and speech finish. Recall uses an explicit
|
|
723
|
+
current-query portal tool, so a result selected for one utterance cannot
|
|
724
|
+
contaminate the next one. Storage never gates hearing, answering, reasoning,
|
|
725
|
+
tools, or speech. The
|
|
726
|
+
Ornith/Omni chat weights have no embedding head and measured poorly when
|
|
727
|
+
forced into that role, so the small dedicated encoder remains the deliberate
|
|
728
|
+
exception and unloads after each job.
|
|
729
|
+
- **Long work is externally audited, checkpointed, and self-checked.** The foreground turn can hand a sustained job
|
|
730
|
+
to the persistent worker, acknowledge immediately, and keep listening. The
|
|
731
|
+
worker reasons and uses tools between speech turns, while a deterministic
|
|
732
|
+
state ledger separately records learned evidence, verified environment
|
|
733
|
+
mutations, controller transitions, and repeated-state stagnation. Accepted
|
|
734
|
+
phase boundaries renew the executor context from that ledger. Spoken
|
|
735
|
+
refinements retain exact user provenance, and completion after a mutation
|
|
736
|
+
requires fresh verification against the current environment version.
|
|
737
|
+
- **Live host state is explicit.** Ordinary turns carry no eager clock,
|
|
738
|
+
location, network, battery, or process blob. Current time, client-browser
|
|
739
|
+
approximate location, and bounded hardware/load facts come from their
|
|
740
|
+
dedicated tools only when the request needs them. The local voice client
|
|
741
|
+
warms the same privacy-filtered browser lookup while the model loads.
|
|
742
|
+
|
|
743
|
+
The speech detector is a port of `portal/static/call_vad.js` with its constants
|
|
744
|
+
intact, so the same room behaves the same way in the browser and here.
|
|
745
|
+
|
|
746
|
+
### Running it as a service
|
|
747
|
+
|
|
748
|
+
```bash
|
|
749
|
+
services/linux/install.sh --with-harness
|
|
750
|
+
```
|
|
751
|
+
|
|
752
|
+
That installs the adapter as a system service and the harness as a **user**
|
|
753
|
+
service. The distinction matters: the harness needs the desktop session it
|
|
754
|
+
speaks into -- its audio devices, its top bar, and the USB permissions the
|
|
755
|
+
logged-in user already has -- so installing it system-wide would leave it
|
|
756
|
+
listening on behalf of nobody.
|
|
757
|
+
|
|
758
|
+
The unit template is `services/linux/omni-call-harness.service.in`. It is a
|
|
759
|
+
desktop-session user unit, while the adapter is a system unit, so it waits for
|
|
760
|
+
the loopback portal rather than declaring an invalid cross-manager dependency.
|
|
761
|
+
Its preflight runs in the exact user-service environment and requires the real
|
|
762
|
+
indicator and audio stack. The unit caps its own memory: the adapter holds the
|
|
763
|
+
weights, this process only moves audio.
|
|
764
|
+
|
|
765
|
+
The core daemon's optional smoke uses a tracked speech fixture and requires a
|
|
766
|
+
tagged transcript, a direct ASR-to-cloned-TTS route using the shipped default
|
|
767
|
+
speaker reference, valid 24 kHz mono PCM16 output, the normal streamed TTS gate,
|
|
768
|
+
and another tagged-ASR pass after speech. It then proves that the original
|
|
769
|
+
comprehension PID and persistent clone-profile TTS PID remain GPU-resident
|
|
770
|
+
together. A generic sound observation or unconditioned WAV cannot satisfy the
|
|
771
|
+
gate. Set `OMNI_STARTUP_SMOKE=1` only when this blocking diagnostic is wanted.
|
|
772
|
+
Guided deployment sets it to `0`, starts the indicator in parallel with core
|
|
773
|
+
readiness, and does not wait for inference. On Jetson diagnostic requests run
|
|
774
|
+
against the installed arm64/CUDA workers;
|
|
775
|
+
desktop-host unit tests do not substitute for that device gate.
|
|
776
|
+
|
|
777
|
+
After deployment, run the non-blocking post-training tool-routing suite
|
|
778
|
+
explicitly without delaying normal boot:
|
|
779
|
+
|
|
780
|
+
```bash
|
|
781
|
+
.venv/bin/python portal/smoke.py \
|
|
782
|
+
--endpoint http://127.0.0.1:8920 \
|
|
783
|
+
--token-file runtime-data/state/access-token.txt \
|
|
784
|
+
--model "$(awk -F= '$1 == "OMNI_MODEL" {print $2}' .env)" \
|
|
785
|
+
--tool-suite
|
|
786
|
+
```
|
|
787
|
+
|
|
788
|
+
It requires real structured calls for portal capabilities, arithmetic, current
|
|
789
|
+
runtime state, and time; a prose capability disclaimer does not pass.
|
|
790
|
+
|
|
791
|
+
These environment variables are worth knowing:
|
|
792
|
+
|
|
793
|
+
| Variable | Effect |
|
|
794
|
+
|---|---|
|
|
795
|
+
| `OMNI_PORTAL_URL` | Where the portal is (default `http://127.0.0.1:8920`) |
|
|
796
|
+
| `OMNI_CALL_CAMERA` | A single camera to use instead of every one found |
|
|
797
|
+
| `OMNI_CALL_MEMORY` | Persistent passive-memory SQLite path |
|
|
798
|
+
| `OMNI_CALL_SPEECH_EVICT_UNIT` | User service to stop before TTS and restore afterward |
|
|
799
|
+
| `OMNI_CALL_COMPREHENSION_HEALTH` | Readiness URL used after restoring that service |
|
|
800
|
+
| `OMNI_MEMORY_GOVERNOR` | Enable (`1`) or disable (`0`) generic runtime memory admission; enabled automatically on Tegra |
|
|
801
|
+
| `OMNI_MEMORY_SOFT_FLOOR_GIB` | Free-memory floor retained before work establishes new model/tool residency (default `3`) |
|
|
802
|
+
| `OMNI_MEMORY_HARD_FLOOR_GIB` | Emergency floor that cancels cancellable work before kernel OOM (default `2`) |
|
|
803
|
+
| `OMNI_MEMORY_OPERATION_RESERVE_GIB` | Additional per-operation reserve above the soft floor (default `1`) |
|
|
804
|
+
| `OMNI_COMPREHENSION_PRESSURE_GRACE_SECONDS` | Continuous low-headroom interval before a context downshift (default `8`) |
|
|
805
|
+
| `OMNI_COMPREHENSION_RUNTIME_RESIZE` | Restart llama.cpp to change its allocated KV tier at runtime; defaults off on Tegra to avoid unsafe JetPack GPU-device teardown and on elsewhere |
|
|
806
|
+
| `OMNI_COMPREHENSION_EXPANSION_GRACE_SECONDS` | Continuous idle-surplus interval before a one-tier expansion (default `60`) |
|
|
807
|
+
| `OMNI_COMPREHENSION_EXPANSION_COOLDOWN_SECONDS` | Minimum delay after a failed/pressured tier before retrying it (default `900`) |
|
|
808
|
+
| `OMNI_CONTEXT_FILE` | Optional complete `robit.omni.context.v1` catalog override; defaults to the packaged context catalog |
|
|
809
|
+
| `OMNI_CALL_LOG_CONTENT` | Opt in to exact structured heard/generated/TTS traces; disabled by default |
|
|
810
|
+
| `OMNI_UPDATE_INTERVAL_SECONDS` | Indicator Git update polling interval; minimum 60 seconds, default 900 |
|
|
811
|
+
|
|
812
|
+
Keep `OMNI_TTS_PERSISTENT=1` only when speech and comprehension genuinely fit
|
|
813
|
+
together. On constrained unified-memory hosts, use `OMNI_TTS_PERSISTENT=0` and
|
|
814
|
+
set `OMNI_CALL_SPEECH_EVICT_UNIT`: the harness completes hearing, reasoning and
|
|
815
|
+
tools as text, stops comprehension, synthesizes once, lets TTS exit, and restores
|
|
816
|
+
comprehension before listening again. This is slower than resident TTS, but it
|
|
817
|
+
prevents the kernel from overcommitting the machine.
|
|
818
|
+
|
|
819
|
+
Tool resource admission is declared beside each tool in `context.json`.
|
|
820
|
+
Standard work must clear the soft floor, bounded continuations may run within
|
|
821
|
+
the soft-to-hard safety band, executors dynamically distinguish new residency
|
|
822
|
+
from reuse, and control-plane work remains available so stalled work can be
|
|
823
|
+
inspected or cancelled. Every cancellable operation still stops at the hard
|
|
824
|
+
floor. An executor may declare a measured `memory_reserve_gib`; otherwise its
|
|
825
|
+
new residency uses the conservative standard admission boundary.
|
|
826
|
+
|
|
827
|
+
## Request example
|
|
828
|
+
|
|
829
|
+
Adapter v1 is `robit.ollama.omni-adapter.v1` and its portable route requires
|
|
830
|
+
`stream:false`:
|
|
831
|
+
|
|
832
|
+
```json
|
|
833
|
+
{
|
|
834
|
+
"model": "robit/qwen3.8-27b-e03-obliterated-omni-audio-bridge:q4km",
|
|
835
|
+
"messages": [{
|
|
836
|
+
"role": "user",
|
|
837
|
+
"content": "What happened, and answer aloud.",
|
|
838
|
+
"audios": [{
|
|
839
|
+
"mime_type": "audio/wav",
|
|
840
|
+
"encoding": "base64",
|
|
841
|
+
"data": "<16 kHz mono PCM16 RIFF/WAVE>"
|
|
842
|
+
}]
|
|
843
|
+
}],
|
|
844
|
+
"omni": {
|
|
845
|
+
"schema": "robit.ollama.omni-adapter.v1",
|
|
846
|
+
"task": "chat"
|
|
847
|
+
},
|
|
848
|
+
"response_modalities": ["text", "audio"],
|
|
849
|
+
"speech_mode": "always",
|
|
850
|
+
"think": false,
|
|
851
|
+
"stream": false
|
|
852
|
+
}
|
|
853
|
+
```
|
|
854
|
+
|
|
855
|
+
Speech returns under `message.audio` as a tagged base64 RIFF/WAVE envelope.
|
|
856
|
+
Transcripts, non-speech acoustic observations, and visual observations remain
|
|
857
|
+
separate so environmental sounds are never misrouted as the user's words.
|
|
858
|
+
|
|
859
|
+
## Runtime guarantees
|
|
860
|
+
|
|
861
|
+
- Every comprehension request sets `cache_prompt:false`; a prior audio/video
|
|
862
|
+
embedding cannot be reused for a new clip.
|
|
863
|
+
- Media turns send only the current attachment as present-tense perceptual
|
|
864
|
+
evidence while retaining bounded prior text dialogue for natural continuity.
|
|
865
|
+
- The portal defaults to one active GPU lane and four admitted active/queued
|
|
866
|
+
requests, with request-local media, tools, voice settings, and streams.
|
|
867
|
+
- A wrench toggle, off by default, exposes server-pinned structured tools
|
|
868
|
+
for local-browser public-web discovery/fetch, attached-document retrieval,
|
|
869
|
+
current time/capabilities, on-demand host snapshots, and temporary session
|
|
870
|
+
web/memory recall and isolated text-only sub-agent delegation. Productive
|
|
871
|
+
tool chains continue until a final answer, while exact duplicates and
|
|
872
|
+
repeated nonproductive rounds stop safely; live collapsible execution
|
|
873
|
+
evidence appears in the response and phone UI. No hosted search API is used.
|
|
874
|
+
- Same-origin IndexedDB restores messages, drafts, pending attachments, reply
|
|
875
|
+
audio, and bounded image/video previews after reload. It is keyed by a
|
|
876
|
+
one-way cookie-derived scope, begins a five-minute expiry on page leave, and
|
|
877
|
+
is deleted immediately by trash. Restored media is display-only and is never
|
|
878
|
+
submitted automatically. The server has no shared model conversation state.
|
|
879
|
+
The document index follows the same session partition and expiry policy.
|
|
880
|
+
- Long speech is split before the per-generation codec-frame ceiling, streamed
|
|
881
|
+
with continuous sequence numbers, and assembled into one complete final WAV.
|
|
882
|
+
- Trained-bridge runtime keeps the matching shipped Qwen3-TTS voice profile
|
|
883
|
+
resident alongside comprehension on its
|
|
884
|
+
assigned GPU and emits two codec frames (about 160 ms) per stream window by
|
|
885
|
+
default. A voice-profile change intentionally replaces the resident worker.
|
|
886
|
+
Non-persistent workers and explicit residency handoff remain legacy/manual
|
|
887
|
+
escape hatches. Guided startup does not run a blocking generation gate.
|
|
888
|
+
- Ordinary turns receive only a compact stable behavioral system policy. With
|
|
889
|
+
tools enabled, `get_system_snapshot` can explicitly sample current date/time,
|
|
890
|
+
OS/architecture, CPU/load, RAM, interface counters, and NVIDIA utilization.
|
|
891
|
+
It excludes hostnames, addresses, routes, sockets, processes, credentials,
|
|
892
|
+
and session content, and describes the portal host—not the user's device.
|
|
893
|
+
- Call cognition is single-flight per browser call. Rapid confirmed segments
|
|
894
|
+
merge into one bounded latest-turn buffer; continuing speech cancels stale
|
|
895
|
+
unanswered inference and preserves its input instead of filling the server
|
|
896
|
+
queue.
|
|
897
|
+
- Sound-only call captures stop after comprehension, render as dim **Audio
|
|
898
|
+
context**, and retain at most six bounded environmental observations for the
|
|
899
|
+
next actual spoken turn. They never invoke language or TTS.
|
|
900
|
+
- `get_user_location` uses a browser-side HTTPS lookup and exposes only
|
|
901
|
+
sanitized, approximate session geography. Raw IP and network metadata never
|
|
902
|
+
reach the portal or model.
|
|
903
|
+
- Adapter responses state current media modalities and whether current visual
|
|
904
|
+
input exists. Location/search/fetch results carry tool/source authority and
|
|
905
|
+
claim limits, so tool evidence cannot legitimately be presented as vision.
|
|
906
|
+
- Reasoning is off until the client sends native `think:true`. Thinking is
|
|
907
|
+
returned separately and is never synthesized.
|
|
908
|
+
- CUDA media inference has no CPU fallback. Broker allocation and exact UUID
|
|
909
|
+
residency are deployment gates on the managed host; direct NVIDIA mode also
|
|
910
|
+
verifies comprehension and every TTS PID with `nvidia-smi`. On an NVIDIA
|
|
911
|
+
Tegra module, whose driver publishes no compute-app accounting at all, the
|
|
912
|
+
same gate is proven from each worker's own integrated-GPU device handles.
|
|
913
|
+
- Session diagnostics are content-redacted, partitioned by an opaque cookie,
|
|
914
|
+
deleted by the trash control, and expire five minutes after a client leaves.
|
|
915
|
+
|
|
916
|
+
## Repository map
|
|
917
|
+
|
|
918
|
+
| Path | Purpose |
|
|
919
|
+
|---|---|
|
|
920
|
+
| `src/qwen_omni_adapters/` | Wire contract, audio validation, GGUF views, Ollama sidecar resolver, accelerator/residency probes, CLI |
|
|
921
|
+
| `runtime/adapter_server.py` | Unified comprehension → language → optional TTS router |
|
|
922
|
+
| `runtime/tts_server.py` | CUDA-only Qwen3-TTS wrapper and PCM stream endpoint |
|
|
923
|
+
| `portal/` | Authenticated phone UI, proxy, supervisor, safe tools, persistent-task store, smoke tests, VAD harness |
|
|
924
|
+
| `harness/` | Always-listening local call mode: VAD, audio I/O, cameras, ReSpeaker, deferred memory storage, background agent, top-bar indicator |
|
|
925
|
+
| `clients/` | Minimal Python and JavaScript request examples |
|
|
926
|
+
| `docs/` | Protocol, architecture, runtime, deployment, ABI, testing, release evidence |
|
|
927
|
+
| `patches/` | Pinned llama.cpp Qwen3-TTS streaming and persistent-worker patches |
|
|
928
|
+
| `scripts/` | Bootstrap, build, validation, and scoped cleanup |
|
|
929
|
+
| `tests/` | Contract, routing, isolation, diagnostics, GGUF, and portal regression tests |
|
|
930
|
+
|
|
931
|
+
## Documentation
|
|
932
|
+
|
|
933
|
+
- [Virtual context](docs/virtual-context.md) — lossless long-history storage,
|
|
934
|
+
recursive retrieval, exact replay, and the bounded 16K working-set contract.
|
|
935
|
+
|
|
936
|
+
- [Architecture and ownership](docs/architecture.md)
|
|
937
|
+
- [Agent runbook](docs/agent-runbook.md)
|
|
938
|
+
- [Runtime guide](docs/runtime.md)
|
|
939
|
+
- [Context engineering and tool/phase map](docs/context-engineering.md)
|
|
940
|
+
- [Verified model profiles](docs/model-profiles.md)
|
|
941
|
+
- [Phone deployment](docs/phone-portal.md)
|
|
942
|
+
- [Linux, macOS, and Windows services](docs/services.md)
|
|
943
|
+
- [arm64 and NVIDIA Jetson](docs/arm-jetson.md)
|
|
944
|
+
- [Wire protocol](docs/protocol.md)
|
|
945
|
+
- [Portal tools and tool chaining](docs/tools.md)
|
|
946
|
+
- [GGUF/Ollama sidecar ABI](docs/gguf-abi.md)
|
|
947
|
+
- [Testing](docs/testing.md)
|
|
948
|
+
- [Deferred component candidates](docs/candidate-components.md)
|
|
949
|
+
- [Cleanup and storage safety](docs/cleanup.md)
|
|
950
|
+
- [Security model](SECURITY.md)
|
|
951
|
+
|
|
952
|
+
The original implementation remains in
|
|
953
|
+
[`robit-man/fine_tuning_suite`](https://github.com/robit-man/fine_tuning_suite)
|
|
954
|
+
for build and model-development workflows. This repository is the smaller,
|
|
955
|
+
stable runtime and integration baseline.
|
|
956
|
+
|
|
957
|
+
For the complete edge device, enclosure, peripherals, telepresence, and
|
|
958
|
+
orchestration project, see
|
|
959
|
+
[`robit-man/EGG`](https://github.com/robit-man/EGG).
|