omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,160 @@
|
|
|
1
|
+
# Trajectory Grounding
|
|
2
|
+
|
|
3
|
+
## Purpose
|
|
4
|
+
|
|
5
|
+
Trajectory Grounding gives a running agent a compact, evidence-bound account of
|
|
6
|
+
where a project stands, what changed, what remains uncertain, and the safest
|
|
7
|
+
next action. Before the first main-agent request, Omnius asks the selected model
|
|
8
|
+
to form a specific, structured orientation from that evidence. This is a
|
|
9
|
+
control-plane checkpoint, not free-form self-talk and not a request to reveal
|
|
10
|
+
private reasoning.
|
|
11
|
+
|
|
12
|
+
The feature exists to prevent loops where changing a payload appears to be
|
|
13
|
+
progress even though the underlying prerequisite has not changed. The
|
|
14
|
+
myactuator repeated full-file write incident is the reference failure mode:
|
|
15
|
+
an existing target rejected unguarded `file_write` calls while generic recovery
|
|
16
|
+
guidance kept steering toward another full write.
|
|
17
|
+
|
|
18
|
+
## Invariants
|
|
19
|
+
|
|
20
|
+
- Every completed, failed, file-state, or verifier claim must point to a
|
|
21
|
+
sanitized evidence reference: a tool result, file hash, verifier outcome, or
|
|
22
|
+
child-agent result.
|
|
23
|
+
- Unknown or stale state is rendered as an open question, never as completed
|
|
24
|
+
work.
|
|
25
|
+
- One current checkpoint is regenerated into the active context frame. It must
|
|
26
|
+
not accumulate as independent system messages.
|
|
27
|
+
- The first active frame always attempts a tool-less model grounding pass. Its
|
|
28
|
+
concise situation assessment must cite supplied evidence labels and be visible
|
|
29
|
+
in the live trajectory box; if inference is unavailable, the UI labels the
|
|
30
|
+
deterministic safety orientation rather than pretending it was generative.
|
|
31
|
+
- Focus-supervisor and explicit safety contracts override generic trajectory
|
|
32
|
+
suggestions.
|
|
33
|
+
- The model must use the checkpoint to choose a next tool action, but must not
|
|
34
|
+
echo it or generate a reasoning transcript for the user.
|
|
35
|
+
- Child agents receive only a scoped slice: parent goal, delegated scope,
|
|
36
|
+
relevant constraint, and exit evidence requirement.
|
|
37
|
+
|
|
38
|
+
## Data contract
|
|
39
|
+
|
|
40
|
+
```ts
|
|
41
|
+
type TrajectoryAssessment =
|
|
42
|
+
| "on_trajectory"
|
|
43
|
+
| "recovery_required"
|
|
44
|
+
| "verification_due"
|
|
45
|
+
| "blocked";
|
|
46
|
+
|
|
47
|
+
interface TrajectoryCheckpoint {
|
|
48
|
+
schemaVersion: 1;
|
|
49
|
+
id: string;
|
|
50
|
+
revision: number;
|
|
51
|
+
turn: number;
|
|
52
|
+
trigger: string;
|
|
53
|
+
assessment: TrajectoryAssessment;
|
|
54
|
+
goal: string;
|
|
55
|
+
currentStep?: string;
|
|
56
|
+
situationAssessment?: string;
|
|
57
|
+
groundingSource?: "model" | "deterministic";
|
|
58
|
+
groundingEvidenceRefs?: string[];
|
|
59
|
+
completedWork: string[];
|
|
60
|
+
groundedFacts: Array<{
|
|
61
|
+
statement: string;
|
|
62
|
+
evidence: string;
|
|
63
|
+
freshness: "fresh" | "stale" | "unknown";
|
|
64
|
+
}>;
|
|
65
|
+
openQuestions: string[];
|
|
66
|
+
nextAction: string;
|
|
67
|
+
successEvidence: string;
|
|
68
|
+
doNotRepeat: string[];
|
|
69
|
+
}
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
The model-facing rendering is capped and follows this form:
|
|
73
|
+
|
|
74
|
+
```text
|
|
75
|
+
[TRAJECTORY CHECKPOINT]
|
|
76
|
+
Goal: ...
|
|
77
|
+
Assessment: recovery_required
|
|
78
|
+
Reasoned situation: the full-file write was rejected and there is no fresh read
|
|
79
|
+
of the target, so a changed payload would not resolve the actual prerequisite.
|
|
80
|
+
Grounded facts:
|
|
81
|
+
- [turn 18/tool_result] file_write on pal.cpp was blocked: existing target
|
|
82
|
+
lacks fresh overwrite/hash evidence.
|
|
83
|
+
Open question: current pal.cpp bytes have not been freshly read.
|
|
84
|
+
Required next action: file_read pal.cpp, then use file_edit/file_patch.
|
|
85
|
+
Success evidence: a targeted mutation or explicit blocker evidence.
|
|
86
|
+
Do not repeat: file_write pal.cpp with a changed full-file payload.
|
|
87
|
+
```
|
|
88
|
+
|
|
89
|
+
## Integration path
|
|
90
|
+
|
|
91
|
+
1. `packages/orchestrator/src/trajectory-checkpoint.ts` owns pure types,
|
|
92
|
+
deterministic safety assessment, model-output validation/merging, rendering,
|
|
93
|
+
fingerprints, and scoped child slices.
|
|
94
|
+
2. `AgenticRunner` records bounded sanitized observations for both executed
|
|
95
|
+
tools and synthetic/preflight blocks. It runs the bounded tool-less
|
|
96
|
+
grounding request before the first main frame, then rebuilds checkpoints in
|
|
97
|
+
`_buildTurnContextFrame()`, which serves normal and brute-force loops.
|
|
98
|
+
3. `ContextFabric` receives a first-class `trajectory_checkpoint` signal with
|
|
99
|
+
high priority, one-turn TTL, and a single semantic conflict group.
|
|
100
|
+
4. `context-compiler` recognizes the block for last-surface deduplication.
|
|
101
|
+
5. Tier prompts tell the model to reconcile its next tool call against the
|
|
102
|
+
checkpoint without narrating it.
|
|
103
|
+
6. The runner emits a typed trajectory event. TUI, API events, debug artifacts,
|
|
104
|
+
session handoffs, and child prompts consume bounded projections of it.
|
|
105
|
+
7. Large whole-file reads hand their isolated extraction branch an agentic
|
|
106
|
+
request built from the current trajectory, active read action, and trigger
|
|
107
|
+
evidence. The branch receives the specific unresolved question and return
|
|
108
|
+
contract, never a raw copy of the original user prompt as its query.
|
|
109
|
+
|
|
110
|
+
## Trigger policy
|
|
111
|
+
|
|
112
|
+
The checkpoint is recomputed before each model request, but only emits a new
|
|
113
|
+
revision when content or direction changes. The model grounding pass runs at
|
|
114
|
+
task start and, in adaptive mode, after material failures, safety-direction
|
|
115
|
+
changes, compaction, or user steering; ordinary successful reads do not create
|
|
116
|
+
an extra planning call. Important triggers are task start, tool failure,
|
|
117
|
+
synthetic safety block, mutation, verifier result, child result, compaction,
|
|
118
|
+
user steering, focus-directive change, and adaptive cadence.
|
|
119
|
+
|
|
120
|
+
Use total tool-call count for cadence rather than a loop-local turn number so
|
|
121
|
+
brute-force re-engagement cannot reset the schedule.
|
|
122
|
+
|
|
123
|
+
## Exposure
|
|
124
|
+
|
|
125
|
+
- **Model context:** a first-class active-frame signal placed after goal/user
|
|
126
|
+
steering and before generic guidance.
|
|
127
|
+
- **TUI:** a compact, collapsible main-view dynamic block and a direction detail
|
|
128
|
+
in the live stage/footer; show updates only when the revision changes.
|
|
129
|
+
- **Sub-agents:** a scoped, non-authoritative parent slice in their prompt.
|
|
130
|
+
- **API/debug:** typed event plus debug artifact on direction change.
|
|
131
|
+
- **Handoff/session diary:** persist the latest meaningful checkpoint only at
|
|
132
|
+
compaction, direction changes, handoff, or completion. Restored checkpoints
|
|
133
|
+
are orientation, not fresh file evidence.
|
|
134
|
+
|
|
135
|
+
## Safety and rollout
|
|
136
|
+
|
|
137
|
+
`OMNIUS_TRAJECTORY_CHECKPOINT=off|shadow|adaptive|always` controls checkpoint
|
|
138
|
+
visibility and cadence. `OMNIUS_TRAJECTORY_GROUNDING=off|initial|adaptive`
|
|
139
|
+
controls model orientation: `initial` always grounds the first main-agent
|
|
140
|
+
request; `adaptive` also regrounds material direction changes. The main TUI
|
|
141
|
+
uses `adaptive`; child agents receive the parent slice rather than recursively
|
|
142
|
+
running a second grounding pass.
|
|
143
|
+
`shadow` computes and records checkpoints without model injection. `adaptive`
|
|
144
|
+
is the default target: inject on meaningful state changes and bounded cadence.
|
|
145
|
+
|
|
146
|
+
No mandatory `trajectory_update` tool is introduced. A tool would create
|
|
147
|
+
ceremony and failure loops for smaller models. Existing todo/workboard state is
|
|
148
|
+
updated only when the trajectory materially changes.
|
|
149
|
+
|
|
150
|
+
## Acceptance criteria
|
|
151
|
+
|
|
152
|
+
- Changed-payload full-write retries produce `recovery_required`, a fresh read
|
|
153
|
+
prerequisite, and a no-repeat constraint.
|
|
154
|
+
- Cached/replayed reads do not satisfy a fresh-read prerequisite.
|
|
155
|
+
- Child completion yields verification work, not an unsupported completion
|
|
156
|
+
claim.
|
|
157
|
+
- Context contains at most one current checkpoint and compaction retains its
|
|
158
|
+
current evidence references.
|
|
159
|
+
- TUI/API/debug views all agree on checkpoint revision and assessment.
|
|
160
|
+
- The feature is bounded, sanitized, and disabled cleanly by configuration.
|
|
@@ -0,0 +1,489 @@
|
|
|
1
|
+
# Voice TTS Flow — Architecture & Implementation Guide
|
|
2
|
+
|
|
3
|
+
This document describes the voice synthesis system in omnius: how it's built,
|
|
4
|
+
how all the pieces connect, and how to add new voice features following the same patterns.
|
|
5
|
+
|
|
6
|
+
---
|
|
7
|
+
|
|
8
|
+
## System Overview
|
|
9
|
+
|
|
10
|
+
```
|
|
11
|
+
┌─────────────┐ ┌──────────────────┐ ┌──────────────┐
|
|
12
|
+
│ interactive │────▶│ VoiceEngine │────▶│ System Audio │
|
|
13
|
+
│ .ts │ │ voice.ts │ │ (afplay/ │
|
|
14
|
+
│ │ │ │ │ paplay) │
|
|
15
|
+
│ ┌────────┐ │ │ ┌────────────┐ │ └──────────────┘
|
|
16
|
+
│ │ Agent │──┤ │ │ ONNX Model │ │
|
|
17
|
+
│ │ Events │ │ │ │ Session │ │ ┌──────────────┐
|
|
18
|
+
│ └────────┘ │ │ └────────────┘ │────▶│ WebSocket │
|
|
19
|
+
│ │ │ ┌────────────┐ │ │ Clients │
|
|
20
|
+
│ ┌────────┐ │ │ │ MLX Audio │ │ │ (voice- │
|
|
21
|
+
│ │Emotion │──┤ │ │ Backend │ │ │ session.ts) │
|
|
22
|
+
│ │Context │ │ │ └────────────┘ │ └──────────────┘
|
|
23
|
+
│ └────────┘ │ └──────────────────┘
|
|
24
|
+
│ │ ▲
|
|
25
|
+
│ ┌────────┐ │ │
|
|
26
|
+
│ │Narrate │──┘ ┌────────┴────────┐
|
|
27
|
+
│ │ Engine │ │ Model Registry │
|
|
28
|
+
│ └────────┘ │ glados/over- │
|
|
29
|
+
└─────────────┘ │ watch/kokoro │
|
|
30
|
+
└─────────────────┘
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
## File Map
|
|
34
|
+
|
|
35
|
+
| File | Purpose | Lines |
|
|
36
|
+
|------|---------|-------|
|
|
37
|
+
| `packages/cli/src/tui/voice.ts` | VoiceEngine class + narration engine | ~2260 |
|
|
38
|
+
| `packages/cli/src/tui/voice-session.ts` | WebSocket streaming + cloudflared tunnel | ~885 |
|
|
39
|
+
| `packages/cli/src/tui/interactive.ts` | Wiring: instantiation, events, narration calls | ~2300 |
|
|
40
|
+
| `packages/cli/src/tui/commands.ts` | `/voice` command handler | ~1850 |
|
|
41
|
+
| `packages/cli/src/tui/render.ts` | `renderInfo`/`renderWarning` + content-write hook | ~670 |
|
|
42
|
+
| `packages/cli/tests/voice-narration.test.ts` | Narration unit tests | varies |
|
|
43
|
+
| `packages/cli/tests/voice-session.test.ts` | Session unit tests | varies |
|
|
44
|
+
|
|
45
|
+
---
|
|
46
|
+
|
|
47
|
+
## 1. Model Registry
|
|
48
|
+
|
|
49
|
+
All voice models are defined in `VOICE_MODELS` (voice.ts, top of file):
|
|
50
|
+
|
|
51
|
+
```typescript
|
|
52
|
+
const VOICE_MODELS: Record<string, VoiceModel> = {
|
|
53
|
+
glados: { id: "glados", backend: "onnx", onnxUrl: "...", configUrl: "..." },
|
|
54
|
+
overwatch:{ id: "overwatch", backend: "onnx", onnxUrl: "...", configUrl: "..." },
|
|
55
|
+
kokoro: { id: "kokoro", backend: "mlx", mlxModelId: "mlx-community/Kokoro-82M-bf16", mlxVoice: "af_heart" },
|
|
56
|
+
// ... more kokoro voices
|
|
57
|
+
};
|
|
58
|
+
```
|
|
59
|
+
|
|
60
|
+
**To add a new model:**
|
|
61
|
+
|
|
62
|
+
1. Add an entry to `VOICE_MODELS` with the appropriate `backend` field
|
|
63
|
+
2. For ONNX: provide `onnxUrl` + `configUrl` (Piper ONNX format)
|
|
64
|
+
3. For MLX: provide `mlxModelId` (Hugging Face repo) + `mlxVoice` + `mlxLangCode`
|
|
65
|
+
4. The model is immediately available via `/voice <id>`
|
|
66
|
+
|
|
67
|
+
---
|
|
68
|
+
|
|
69
|
+
## 2. VoiceEngine Lifecycle
|
|
70
|
+
|
|
71
|
+
### Instantiation
|
|
72
|
+
|
|
73
|
+
A single `VoiceEngine` instance is created in `interactive.ts:1359`:
|
|
74
|
+
|
|
75
|
+
```typescript
|
|
76
|
+
const voiceEngine = new VoiceEngine();
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
### Startup
|
|
80
|
+
|
|
81
|
+
If the user previously enabled voice (persisted in settings):
|
|
82
|
+
|
|
83
|
+
```typescript
|
|
84
|
+
if (savedSettings.voice) {
|
|
85
|
+
voiceEngine.toggle(); // Enable TTS
|
|
86
|
+
if (savedSettings.voiceModel)
|
|
87
|
+
voiceEngine.setModel(savedSettings.voiceModel); // Load specific model
|
|
88
|
+
}
|
|
89
|
+
```
|
|
90
|
+
|
|
91
|
+
### toggle() Flow
|
|
92
|
+
|
|
93
|
+
```
|
|
94
|
+
toggle()
|
|
95
|
+
├─ If disabling: set enabled=false, killPlayback(), done
|
|
96
|
+
└─ If enabling:
|
|
97
|
+
├─ MLX model?
|
|
98
|
+
│ └─ ensureMlxAudio() → pip install mlx-audio
|
|
99
|
+
└─ ONNX model?
|
|
100
|
+
├─ ensureRuntime() → npm install onnxruntime-node + phonemizer
|
|
101
|
+
├─ ensureModel(id) → download .onnx + .json from GitHub
|
|
102
|
+
└─ loadSession() → create InferenceSession
|
|
103
|
+
set enabled=true, ready=true
|
|
104
|
+
```
|
|
105
|
+
|
|
106
|
+
### setModel(id) Flow
|
|
107
|
+
|
|
108
|
+
```
|
|
109
|
+
setModel(id)
|
|
110
|
+
├─ Validate id exists in VOICE_MODELS
|
|
111
|
+
├─ Reset: session=null, config=null, ready=false
|
|
112
|
+
└─ If currently enabled:
|
|
113
|
+
└─ Load the new model (same as toggle enable path)
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
### Shutdown
|
|
117
|
+
|
|
118
|
+
```typescript
|
|
119
|
+
voiceEngine.dispose(); // in interactive.ts cleanup
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
---
|
|
123
|
+
|
|
124
|
+
## 3. Synthesis Pipeline
|
|
125
|
+
|
|
126
|
+
### ONNX Path (glados, overwatch)
|
|
127
|
+
|
|
128
|
+
```
|
|
129
|
+
text
|
|
130
|
+
│
|
|
131
|
+
├── chunkText() Split >200 char text on newlines + sentence boundaries
|
|
132
|
+
│
|
|
133
|
+
├── For each chunk:
|
|
134
|
+
│ ├── textToPhonemes() espeak-ng WASM phonemization
|
|
135
|
+
│ ├── phonemesToIds() Map phonemes → integer IDs via config.phoneme_id_map
|
|
136
|
+
│ │ Format: BOS → PAD → (phoneme + PAD)* → EOS
|
|
137
|
+
│ ├── Build tensors: input (int64), input_lengths (int64), scales (float32)
|
|
138
|
+
│ ├── session.run() ONNX inference → Float32 audio samples
|
|
139
|
+
│ └── Concatenate with 180ms silence gaps between sentences
|
|
140
|
+
│
|
|
141
|
+
├── Apply volume scaling (0.0–1.0 multiplier per sample)
|
|
142
|
+
├── Apply pitch shift (linear-interpolation resampling)
|
|
143
|
+
│
|
|
144
|
+
├── Stream PCM to WebSocket clients (if onPCMOutput wired)
|
|
145
|
+
├── Write WAV to temp file
|
|
146
|
+
├── Play via system command (afplay / paplay / pw-play / aplay)
|
|
147
|
+
└── Delete temp file
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
### MLX Path (kokoro, kokoro:af_heart, etc.)
|
|
151
|
+
|
|
152
|
+
```
|
|
153
|
+
text
|
|
154
|
+
│
|
|
155
|
+
├── Clean markdown (* removal)
|
|
156
|
+
├── Build Python command:
|
|
157
|
+
│ python3 -c "from mlx_audio.tts import generate; generate.main([...])"
|
|
158
|
+
│ --model mlx-community/Kokoro-82M-bf16
|
|
159
|
+
│ --text "..."
|
|
160
|
+
│ --voice af_heart
|
|
161
|
+
│ --lang_code a
|
|
162
|
+
│ --audio_path /tmp/omnius-mlx-{ts}.wav
|
|
163
|
+
│
|
|
164
|
+
├── execSync() with 60s timeout
|
|
165
|
+
│ (fallback: python3 -m mlx_audio.tts.generate CLI)
|
|
166
|
+
│
|
|
167
|
+
├── Apply volume scaling (rewrite WAV PCM samples)
|
|
168
|
+
├── Stream PCM to WebSocket clients (parse WAV header for sample rate)
|
|
169
|
+
├── Play via system command
|
|
170
|
+
└── Delete temp file
|
|
171
|
+
```
|
|
172
|
+
|
|
173
|
+
---
|
|
174
|
+
|
|
175
|
+
## 4. Narration Engine
|
|
176
|
+
|
|
177
|
+
The narration system generates context-aware spoken descriptions of agent activity.
|
|
178
|
+
It lives in voice.ts below the VoiceEngine class (~line 1268+).
|
|
179
|
+
|
|
180
|
+
### Personality Levels
|
|
181
|
+
|
|
182
|
+
```
|
|
183
|
+
1 = minimal "Reading file.ts"
|
|
184
|
+
2 = brief "Reading file.ts"
|
|
185
|
+
3 = conv "Let me take a look at file.ts"
|
|
186
|
+
4 = chatty "Alright, let's crack open file.ts"
|
|
187
|
+
5 = theatrical "Alright, let's crack open file.ts and see what we're working with"
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
Mapped from the agent's personality preset:
|
|
191
|
+
```typescript
|
|
192
|
+
{ concise: 1, balanced: 3, verbose: 4, pedagogical: 5 }
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
### describeToolCall(toolName, args, level, emotion)
|
|
196
|
+
|
|
197
|
+
Called from `interactive.ts` on every `tool_call` event:
|
|
198
|
+
|
|
199
|
+
```typescript
|
|
200
|
+
if (voice?.enabled) {
|
|
201
|
+
const desc = describeToolCall(event.toolName, event.toolArgs, vLevel, emoCtx);
|
|
202
|
+
voice.speakSubordinate(desc, emoCtx); // 55% volume, 0.92x pitch
|
|
203
|
+
}
|
|
204
|
+
```
|
|
205
|
+
|
|
206
|
+
**Variant pools** — each tool has 3 tiers of phrasings:
|
|
207
|
+
- `FILE_READ_VARIANTS.terse` / `.conv` / `.chatty`
|
|
208
|
+
- `FILE_WRITE_VARIANTS`, `FILE_EDIT_VARIANTS`, `GREP_VARIANTS`, etc.
|
|
209
|
+
- `SHELL_VARIANTS` — categorized by shell command type (git, npm, test, etc.)
|
|
210
|
+
|
|
211
|
+
**Context modifiers** applied based on narration state:
|
|
212
|
+
- After errors: prefix with "Okay, " or "Right, "
|
|
213
|
+
- Same file again: "Back to " or "Still working on "
|
|
214
|
+
- Revisiting a file: "Coming back to " or "Revisiting "
|
|
215
|
+
- Progress beats every 8 tools: "Making good progress. "
|
|
216
|
+
|
|
217
|
+
### describeToolResult(toolName, success, level, content, emotion)
|
|
218
|
+
|
|
219
|
+
Called on every `tool_result` event. Generates success/failure descriptions:
|
|
220
|
+
- Success: "Got it", "Done", "Found it"
|
|
221
|
+
- Failure: "That didn't work", "Hit a snag"
|
|
222
|
+
|
|
223
|
+
**Content-aware extraction** via `extractResultDigest()`:
|
|
224
|
+
- ETH balances, test results, error messages, wallet addresses, file paths
|
|
225
|
+
|
|
226
|
+
### describeTaskComplete(summary, complete, level)
|
|
227
|
+
|
|
228
|
+
Announces when the agent finishes a task:
|
|
229
|
+
- "Task complete", "Got it done", "All set"
|
|
230
|
+
|
|
231
|
+
### Narration State
|
|
232
|
+
|
|
233
|
+
```typescript
|
|
234
|
+
interface NarrationState {
|
|
235
|
+
toolCount: number; // Total tools this session
|
|
236
|
+
toolCounts: Record<string, number>; // Per-tool counts
|
|
237
|
+
consecutiveErrors: number; // Reset on success
|
|
238
|
+
totalErrors: number;
|
|
239
|
+
lastTool: string;
|
|
240
|
+
lastFile: string;
|
|
241
|
+
filesSeen: Set<string>; // All files visited
|
|
242
|
+
lastVariantIdx: Record<string, number>; // Avoid repeats
|
|
243
|
+
lastResultDigest: string; // Last tool result summary
|
|
244
|
+
}
|
|
245
|
+
```
|
|
246
|
+
|
|
247
|
+
`resetNarrationContext()` is called at the start of each task.
|
|
248
|
+
|
|
249
|
+
### pick(key, variants)
|
|
250
|
+
|
|
251
|
+
Selects a random variant from a pool, avoiding the last-used index for that key.
|
|
252
|
+
Ensures you never hear the same phrasing twice in a row.
|
|
253
|
+
|
|
254
|
+
---
|
|
255
|
+
|
|
256
|
+
## 5. Emotion Modulation
|
|
257
|
+
|
|
258
|
+
The emotion engine provides valence-arousal context to the voice:
|
|
259
|
+
|
|
260
|
+
```typescript
|
|
261
|
+
interface VoiceEmotionContext {
|
|
262
|
+
valence: number; // -1 (sad) to +1 (happy)
|
|
263
|
+
arousal: number; // 0 (calm) to 1 (activated)
|
|
264
|
+
label: string; // "excited", "focused", etc.
|
|
265
|
+
emoji: string;
|
|
266
|
+
}
|
|
267
|
+
```
|
|
268
|
+
|
|
269
|
+
### Pitch Bias
|
|
270
|
+
|
|
271
|
+
`emotionToPitchBias(emotion)` converts emotion to a pitch adjustment:
|
|
272
|
+
|
|
273
|
+
```
|
|
274
|
+
pitch_bias = valence × 0.6 + (arousal - 0.5) × 0.4
|
|
275
|
+
clamped to [-0.10, +0.10]
|
|
276
|
+
```
|
|
277
|
+
|
|
278
|
+
- **Excited** (high valence + high arousal) → voice pitch rises
|
|
279
|
+
- **Dejected** (low valence + low arousal) → voice pitch drops
|
|
280
|
+
|
|
281
|
+
Applied in `speak()` and `speakSubordinate()`:
|
|
282
|
+
|
|
283
|
+
```typescript
|
|
284
|
+
speak(text, emotion): pitchFactor = 1.0 + pitchBias
|
|
285
|
+
speakSubordinate(text): pitchFactor = 0.92 + pitchBias (lower base pitch)
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
### Emotion Coloring
|
|
289
|
+
|
|
290
|
+
At personality >= 3, ~30% of narrations get emotion-colored prefixes:
|
|
291
|
+
- Excited: "Feeling good about this"
|
|
292
|
+
- Stressed: "Pushing through"
|
|
293
|
+
- Calm: "Nice and steady"
|
|
294
|
+
- Subdued: "Being careful here"
|
|
295
|
+
|
|
296
|
+
---
|
|
297
|
+
|
|
298
|
+
## 6. Queue & Playback
|
|
299
|
+
|
|
300
|
+
### Speech Queue
|
|
301
|
+
|
|
302
|
+
```typescript
|
|
303
|
+
private speakQueue: SpeakItem[] = [];
|
|
304
|
+
```
|
|
305
|
+
|
|
306
|
+
Items are queued FIFO. `drainQueue()` processes them sequentially:
|
|
307
|
+
1. Pop item from front
|
|
308
|
+
2. Synthesize + play to completion
|
|
309
|
+
3. 250ms silence gap
|
|
310
|
+
4. Next item
|
|
311
|
+
|
|
312
|
+
Queue overflow protection: if > 30 items backed up, clear the queue.
|
|
313
|
+
|
|
314
|
+
### Volume Levels
|
|
315
|
+
|
|
316
|
+
```
|
|
317
|
+
speak() → volume 1.0 (full)
|
|
318
|
+
speakSubordinate() → volume 0.55 (reduced for tool narration)
|
|
319
|
+
```
|
|
320
|
+
|
|
321
|
+
### Playback
|
|
322
|
+
|
|
323
|
+
WAV temp file → system audio command:
|
|
324
|
+
- macOS: `afplay`
|
|
325
|
+
- Linux: `paplay` → `pw-play` → `aplay` (tries in order)
|
|
326
|
+
- Windows: PowerShell `Media.SoundPlayer`
|
|
327
|
+
|
|
328
|
+
15-second safety timeout per playback. `killPlayback()` sends SIGTERM.
|
|
329
|
+
|
|
330
|
+
---
|
|
331
|
+
|
|
332
|
+
## 7. WebSocket Voice Session
|
|
333
|
+
|
|
334
|
+
`VoiceSession` (voice-session.ts) enables real-time voice interaction via browser:
|
|
335
|
+
|
|
336
|
+
```
|
|
337
|
+
Browser Client ←──WebSocket──→ VoiceSession ←──PCM──→ VoiceEngine
|
|
338
|
+
(mic) binary PCM (server) onPCMOutput (TTS)
|
|
339
|
+
(speaker) binary PCM
|
|
340
|
+
```
|
|
341
|
+
|
|
342
|
+
### Setup
|
|
343
|
+
|
|
344
|
+
```
|
|
345
|
+
start()
|
|
346
|
+
├── Start HTTP server (serves HTML single-page app)
|
|
347
|
+
├── Start WebSocket server (ws)
|
|
348
|
+
├── Launch cloudflared tunnel for public URL
|
|
349
|
+
└── Wire VoiceEngine.onPCMOutput → broadcast to all clients
|
|
350
|
+
```
|
|
351
|
+
|
|
352
|
+
### Audio Format
|
|
353
|
+
|
|
354
|
+
- 16kHz, 16-bit, mono PCM (Int16Array)
|
|
355
|
+
- Binary WebSocket frames
|
|
356
|
+
- Echo cancellation: suppress mic input while TTS is playing
|
|
357
|
+
|
|
358
|
+
### Frontend
|
|
359
|
+
|
|
360
|
+
Embedded HTML with:
|
|
361
|
+
- Braille waveform animator (color-coded: idle/listening/speaking)
|
|
362
|
+
- WebAudio API for mic capture + speaker playback
|
|
363
|
+
- Transcript view (user + agent messages)
|
|
364
|
+
- Start/stop mic button
|
|
365
|
+
|
|
366
|
+
---
|
|
367
|
+
|
|
368
|
+
## 8. Settings Persistence
|
|
369
|
+
|
|
370
|
+
Voice settings are saved per-project or globally via `resolveSettings()`:
|
|
371
|
+
|
|
372
|
+
```typescript
|
|
373
|
+
{
|
|
374
|
+
voice: boolean, // Enabled/disabled
|
|
375
|
+
voiceModel: string // Model ID ("glados", "kokoro:af_heart", etc.)
|
|
376
|
+
}
|
|
377
|
+
```
|
|
378
|
+
|
|
379
|
+
The `/voice` command handler in commands.ts:
|
|
380
|
+
```typescript
|
|
381
|
+
case "voice":
|
|
382
|
+
if (arg) {
|
|
383
|
+
ctx.voiceSetModel(arg); // /voice kokoro
|
|
384
|
+
save({ voice: true, voiceModel: arg });
|
|
385
|
+
} else {
|
|
386
|
+
ctx.voiceToggle(); // /voice (toggle on/off)
|
|
387
|
+
save({ voice: isOn });
|
|
388
|
+
}
|
|
389
|
+
```
|
|
390
|
+
|
|
391
|
+
---
|
|
392
|
+
|
|
393
|
+
## 9. Content-Write Hook (TUI Safety)
|
|
394
|
+
|
|
395
|
+
Voice operations can trigger `renderInfo()` / `renderWarning()` messages asynchronously.
|
|
396
|
+
To prevent these from overwriting the TUI input area, a global content-write hook
|
|
397
|
+
brackets all render calls with scroll-region management:
|
|
398
|
+
|
|
399
|
+
```typescript
|
|
400
|
+
// render.ts
|
|
401
|
+
export function renderInfo(message: string): void {
|
|
402
|
+
_contentWriteHook?.begin(); // → statusBar.beginContentWrite()
|
|
403
|
+
process.stdout.write(`ℹ ${message}\n`);
|
|
404
|
+
_contentWriteHook?.end(); // → statusBar.endContentWrite()
|
|
405
|
+
}
|
|
406
|
+
```
|
|
407
|
+
|
|
408
|
+
Registered once in interactive.ts after StatusBar activation:
|
|
409
|
+
```typescript
|
|
410
|
+
setContentWriteHook({
|
|
411
|
+
begin: () => statusBar.beginContentWrite(),
|
|
412
|
+
end: () => statusBar.endContentWrite(),
|
|
413
|
+
});
|
|
414
|
+
```
|
|
415
|
+
|
|
416
|
+
**Rule:** Never use raw `process.stdout.write()` in voice.ts for user-facing messages.
|
|
417
|
+
Always use `renderInfo()` / `renderWarning()` / `renderError()` so the content hook
|
|
418
|
+
routes output to the scroll region above the status bar and input area.
|
|
419
|
+
|
|
420
|
+
---
|
|
421
|
+
|
|
422
|
+
## 10. How to Add a New Voice Feature
|
|
423
|
+
|
|
424
|
+
### Adding a New ONNX Voice Model
|
|
425
|
+
|
|
426
|
+
1. Host the `.onnx` + `.onnx.json` files (Piper format) on a public URL
|
|
427
|
+
2. Add entry to `VOICE_MODELS`:
|
|
428
|
+
```typescript
|
|
429
|
+
myvoice: {
|
|
430
|
+
id: "myvoice", label: "My Voice", backend: "onnx",
|
|
431
|
+
onnxUrl: "https://...", configUrl: "https://...",
|
|
432
|
+
},
|
|
433
|
+
```
|
|
434
|
+
3. Done. `/voice myvoice` works immediately.
|
|
435
|
+
|
|
436
|
+
### Adding a New MLX Voice
|
|
437
|
+
|
|
438
|
+
1. Find the Hugging Face model ID (must be MLX-compatible)
|
|
439
|
+
2. Add entry to `VOICE_MODELS`:
|
|
440
|
+
```typescript
|
|
441
|
+
"newmodel:voice_name": {
|
|
442
|
+
id: "newmodel:voice_name", label: "New Model (MLX)", backend: "mlx",
|
|
443
|
+
mlxModelId: "mlx-community/NewModel", mlxVoice: "voice_name", mlxLangCode: "a",
|
|
444
|
+
onnxUrl: "", configUrl: "",
|
|
445
|
+
},
|
|
446
|
+
```
|
|
447
|
+
3. Done. `/voice newmodel:voice_name` works on macOS Apple Silicon.
|
|
448
|
+
|
|
449
|
+
### Adding a New TTS Backend
|
|
450
|
+
|
|
451
|
+
1. Add a new `backend` value to the `VoiceModel` interface
|
|
452
|
+
2. Add a method like `ensureNewBackend()` for installation
|
|
453
|
+
3. Add a method like `synthesizeWithNewBackend()` for synthesis
|
|
454
|
+
4. Gate in `toggle()`, `setModel()`, `synthesizeAndPlay()`, `synthesizeToBuffer()`, `synthesizeToPCM()`
|
|
455
|
+
5. Follow the pattern: the backend outputs a WAV file, then existing `playWav()` handles playback
|
|
456
|
+
|
|
457
|
+
### Adding New Narration Variants
|
|
458
|
+
|
|
459
|
+
1. Find the relevant variant pool (e.g., `FILE_READ_VARIANTS`)
|
|
460
|
+
2. Add new strings to the `terse`, `conv`, and `chatty` tiers
|
|
461
|
+
3. The `pick()` function automatically rotates through variants
|
|
462
|
+
|
|
463
|
+
### Adding Emotion-Aware Features
|
|
464
|
+
|
|
465
|
+
1. Receive `VoiceEmotionContext` from the emotion engine
|
|
466
|
+
2. Use `emotionToPitchBias()` for pitch modulation
|
|
467
|
+
3. Use `emotionColor()` for spoken prefixes
|
|
468
|
+
4. Valence drives tone (happy/sad), arousal drives energy (calm/activated)
|
|
469
|
+
|
|
470
|
+
---
|
|
471
|
+
|
|
472
|
+
## Quick Reference: Key Functions
|
|
473
|
+
|
|
474
|
+
| Function | Location | Purpose |
|
|
475
|
+
|----------|----------|---------|
|
|
476
|
+
| `VoiceEngine.toggle()` | voice.ts | Enable/disable TTS |
|
|
477
|
+
| `VoiceEngine.speak()` | voice.ts | Queue speech (full volume) |
|
|
478
|
+
| `VoiceEngine.speakSubordinate()` | voice.ts | Queue speech (55% volume) |
|
|
479
|
+
| `VoiceEngine.synthesizeAndPlay()` | voice.ts | Core synthesis + playback |
|
|
480
|
+
| `VoiceEngine.synthesizeWithMlx()` | voice.ts | MLX backend synthesis |
|
|
481
|
+
| `VoiceEngine.ensureRuntime()` | voice.ts | Install ONNX runtime |
|
|
482
|
+
| `VoiceEngine.ensureMlxAudio()` | voice.ts | Install mlx-audio pip package |
|
|
483
|
+
| `describeToolCall()` | voice.ts | Generate tool narration text |
|
|
484
|
+
| `describeToolResult()` | voice.ts | Generate result narration text |
|
|
485
|
+
| `describeTaskComplete()` | voice.ts | Generate completion narration |
|
|
486
|
+
| `emotionToPitchBias()` | voice.ts | Emotion → pitch modulation |
|
|
487
|
+
| `pick()` | voice.ts | Random variant selection (no repeats) |
|
|
488
|
+
| `setContentWriteHook()` | render.ts | Register TUI scroll-region safety |
|
|
489
|
+
| `VoiceSession.start()` | voice-session.ts | Start WebSocket voice server |
|