omnius 1.0.591 → 1.0.592

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/.aiwg/addons/omnius-docs/README.md +15 -1
  2. package/.aiwg/addons/omnius-docs/manifest.json +28 -68
  3. package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
  4. package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
  5. package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
  6. package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
  7. package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
  8. package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
  9. package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
  10. package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
  11. package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
  12. package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
  13. package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
  14. package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
  15. package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
  16. package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
  17. package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
  18. package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
  19. package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
  20. package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
  21. package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
  22. package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
  23. package/README.md +36 -0
  24. package/dist/discovery.d.ts +50 -0
  25. package/dist/index.js +5975 -4021
  26. package/dist/library.d.ts +7 -0
  27. package/dist/library.js +950 -0
  28. package/dist/postinstall-daemon.cjs +18 -0
  29. package/dist/providerRegistry.d.ts +80 -0
  30. package/dist/service-version.d.ts +35 -0
  31. package/docs/.vitepress/config.mts +8 -0
  32. package/docs/DISCOVERY.json +20224 -0
  33. package/docs/DISCOVERY.md +648 -0
  34. package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
  35. package/docs/agent-memory/INDEX.md +9 -4
  36. package/docs/agent-memory/index.md +7 -0
  37. package/docs/concept-relational-language.md +869 -0
  38. package/docs/context-management-medium-models-proposal.md +449 -0
  39. package/docs/dedup-false-positive-meta-analysis.md +96 -0
  40. package/docs/discovery/catalog-overrides.json +724 -0
  41. package/docs/duplicate-calls-root-cause-analysis.md +91 -0
  42. package/docs/duplicate-calls-root-cause-deep.md +155 -0
  43. package/docs/ephemeral-skill-pack-small-context.md +57 -0
  44. package/docs/explorations/context-window-todo-association.md +156 -0
  45. package/docs/explorations/todo-association-verify.json +30 -0
  46. package/docs/explorations/verification-ledger.json +45 -0
  47. package/docs/explorations/verify-todo-association.sh +30 -0
  48. package/docs/flowstate.md +806 -0
  49. package/docs/getting-started/install.md +24 -0
  50. package/docs/getting-started/model-providers.md +13 -0
  51. package/docs/guides/agent-integration.md +87 -0
  52. package/docs/guides/bring-your-own-inference.md +126 -0
  53. package/docs/guides/tools-and-web-search.md +95 -0
  54. package/docs/index.md +14 -0
  55. package/docs/longhaul-35b-workorders.md +496 -0
  56. package/docs/memory-integration-analysis.md +303 -0
  57. package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
  58. package/docs/multimodal-identity-memory-implementation.md +76 -0
  59. package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
  60. package/docs/opencode-agentic-loop-comparison.md +290 -0
  61. package/docs/operations/security-and-remote-access.md +2 -2
  62. package/docs/operations/version-compatibility.md +63 -0
  63. package/docs/proposals/git-progress-tracking-strategy.md +289 -0
  64. package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
  65. package/docs/proposals/opencode-modules/childSession.ts +288 -0
  66. package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
  67. package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
  68. package/docs/proposals/opencode-modules/runner.ts +258 -0
  69. package/docs/reference/auth-map.md +87 -196
  70. package/docs/reference/configuration.md +27 -0
  71. package/docs/reference/rest-api.md +7 -0
  72. package/docs/reference/slash-commands.md +125 -2
  73. package/docs/research/_archived/README.md +18 -0
  74. package/docs/research/_archived/context_window_attention_model.py +418 -0
  75. package/docs/research/_archived/context_window_attention_spec.md +55 -0
  76. package/docs/research/_archived/context_window_attention_weights.json +68 -0
  77. package/docs/research/k-splanifolds.pdf +0 -0
  78. package/docs/research/personality-verbosity-control.md +293 -0
  79. package/docs/rest/INDEX.md +7 -0
  80. package/docs/rest/QUICKREF.md +18 -0
  81. package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
  82. package/docs/rest/auth-and-scopes.md +7 -1
  83. package/docs/rest/endpoints/discovery.md +44 -0
  84. package/docs/rest/endpoints/events.md +5 -0
  85. package/docs/rest/endpoints/tools.md +9 -0
  86. package/docs/reviews/adversary-system-review.md +42 -0
  87. package/docs/sana-and-video-generation-integration-plan.md +712 -0
  88. package/docs/session-diary-llm-training-analysis.md +218 -0
  89. package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
  90. package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
  91. package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
  92. package/docs/telegram-unified-tooling-architecture.md +332 -0
  93. package/docs/threat-model.md +868 -0
  94. package/docs/trajectory-grounding.md +160 -0
  95. package/docs/voice-flow-architecture.md +489 -0
  96. package/docs/work-orders/WO-AM-GAPS.md +638 -0
  97. package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
  98. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
  99. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
  100. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
  101. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
  102. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
  103. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
  104. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
  105. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
  106. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
  107. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
  108. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
  109. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
  110. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
  111. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
  112. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
  113. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
  114. package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
  115. package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
  116. package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
  117. package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
  118. package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
  119. package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
  120. package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
  121. package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
  122. package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
  123. package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
  124. package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
  125. package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
  126. package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
  127. package/docs/x402-remote-inference-plan.md +323 -0
  128. package/npm-shrinkwrap.json +108 -117
  129. package/package.json +7 -6
  130. package/templates/AGENTS.md +6 -0
  131. package/templates/OMNIUS.md +20 -0
@@ -0,0 +1,160 @@
1
+ # Trajectory Grounding
2
+
3
+ ## Purpose
4
+
5
+ Trajectory Grounding gives a running agent a compact, evidence-bound account of
6
+ where a project stands, what changed, what remains uncertain, and the safest
7
+ next action. Before the first main-agent request, Omnius asks the selected model
8
+ to form a specific, structured orientation from that evidence. This is a
9
+ control-plane checkpoint, not free-form self-talk and not a request to reveal
10
+ private reasoning.
11
+
12
+ The feature exists to prevent loops where changing a payload appears to be
13
+ progress even though the underlying prerequisite has not changed. The
14
+ myactuator repeated full-file write incident is the reference failure mode:
15
+ an existing target rejected unguarded `file_write` calls while generic recovery
16
+ guidance kept steering toward another full write.
17
+
18
+ ## Invariants
19
+
20
+ - Every completed, failed, file-state, or verifier claim must point to a
21
+ sanitized evidence reference: a tool result, file hash, verifier outcome, or
22
+ child-agent result.
23
+ - Unknown or stale state is rendered as an open question, never as completed
24
+ work.
25
+ - One current checkpoint is regenerated into the active context frame. It must
26
+ not accumulate as independent system messages.
27
+ - The first active frame always attempts a tool-less model grounding pass. Its
28
+ concise situation assessment must cite supplied evidence labels and be visible
29
+ in the live trajectory box; if inference is unavailable, the UI labels the
30
+ deterministic safety orientation rather than pretending it was generative.
31
+ - Focus-supervisor and explicit safety contracts override generic trajectory
32
+ suggestions.
33
+ - The model must use the checkpoint to choose a next tool action, but must not
34
+ echo it or generate a reasoning transcript for the user.
35
+ - Child agents receive only a scoped slice: parent goal, delegated scope,
36
+ relevant constraint, and exit evidence requirement.
37
+
38
+ ## Data contract
39
+
40
+ ```ts
41
+ type TrajectoryAssessment =
42
+ | "on_trajectory"
43
+ | "recovery_required"
44
+ | "verification_due"
45
+ | "blocked";
46
+
47
+ interface TrajectoryCheckpoint {
48
+ schemaVersion: 1;
49
+ id: string;
50
+ revision: number;
51
+ turn: number;
52
+ trigger: string;
53
+ assessment: TrajectoryAssessment;
54
+ goal: string;
55
+ currentStep?: string;
56
+ situationAssessment?: string;
57
+ groundingSource?: "model" | "deterministic";
58
+ groundingEvidenceRefs?: string[];
59
+ completedWork: string[];
60
+ groundedFacts: Array<{
61
+ statement: string;
62
+ evidence: string;
63
+ freshness: "fresh" | "stale" | "unknown";
64
+ }>;
65
+ openQuestions: string[];
66
+ nextAction: string;
67
+ successEvidence: string;
68
+ doNotRepeat: string[];
69
+ }
70
+ ```
71
+
72
+ The model-facing rendering is capped and follows this form:
73
+
74
+ ```text
75
+ [TRAJECTORY CHECKPOINT]
76
+ Goal: ...
77
+ Assessment: recovery_required
78
+ Reasoned situation: the full-file write was rejected and there is no fresh read
79
+ of the target, so a changed payload would not resolve the actual prerequisite.
80
+ Grounded facts:
81
+ - [turn 18/tool_result] file_write on pal.cpp was blocked: existing target
82
+ lacks fresh overwrite/hash evidence.
83
+ Open question: current pal.cpp bytes have not been freshly read.
84
+ Required next action: file_read pal.cpp, then use file_edit/file_patch.
85
+ Success evidence: a targeted mutation or explicit blocker evidence.
86
+ Do not repeat: file_write pal.cpp with a changed full-file payload.
87
+ ```
88
+
89
+ ## Integration path
90
+
91
+ 1. `packages/orchestrator/src/trajectory-checkpoint.ts` owns pure types,
92
+ deterministic safety assessment, model-output validation/merging, rendering,
93
+ fingerprints, and scoped child slices.
94
+ 2. `AgenticRunner` records bounded sanitized observations for both executed
95
+ tools and synthetic/preflight blocks. It runs the bounded tool-less
96
+ grounding request before the first main frame, then rebuilds checkpoints in
97
+ `_buildTurnContextFrame()`, which serves normal and brute-force loops.
98
+ 3. `ContextFabric` receives a first-class `trajectory_checkpoint` signal with
99
+ high priority, one-turn TTL, and a single semantic conflict group.
100
+ 4. `context-compiler` recognizes the block for last-surface deduplication.
101
+ 5. Tier prompts tell the model to reconcile its next tool call against the
102
+ checkpoint without narrating it.
103
+ 6. The runner emits a typed trajectory event. TUI, API events, debug artifacts,
104
+ session handoffs, and child prompts consume bounded projections of it.
105
+ 7. Large whole-file reads hand their isolated extraction branch an agentic
106
+ request built from the current trajectory, active read action, and trigger
107
+ evidence. The branch receives the specific unresolved question and return
108
+ contract, never a raw copy of the original user prompt as its query.
109
+
110
+ ## Trigger policy
111
+
112
+ The checkpoint is recomputed before each model request, but only emits a new
113
+ revision when content or direction changes. The model grounding pass runs at
114
+ task start and, in adaptive mode, after material failures, safety-direction
115
+ changes, compaction, or user steering; ordinary successful reads do not create
116
+ an extra planning call. Important triggers are task start, tool failure,
117
+ synthetic safety block, mutation, verifier result, child result, compaction,
118
+ user steering, focus-directive change, and adaptive cadence.
119
+
120
+ Use total tool-call count for cadence rather than a loop-local turn number so
121
+ brute-force re-engagement cannot reset the schedule.
122
+
123
+ ## Exposure
124
+
125
+ - **Model context:** a first-class active-frame signal placed after goal/user
126
+ steering and before generic guidance.
127
+ - **TUI:** a compact, collapsible main-view dynamic block and a direction detail
128
+ in the live stage/footer; show updates only when the revision changes.
129
+ - **Sub-agents:** a scoped, non-authoritative parent slice in their prompt.
130
+ - **API/debug:** typed event plus debug artifact on direction change.
131
+ - **Handoff/session diary:** persist the latest meaningful checkpoint only at
132
+ compaction, direction changes, handoff, or completion. Restored checkpoints
133
+ are orientation, not fresh file evidence.
134
+
135
+ ## Safety and rollout
136
+
137
+ `OMNIUS_TRAJECTORY_CHECKPOINT=off|shadow|adaptive|always` controls checkpoint
138
+ visibility and cadence. `OMNIUS_TRAJECTORY_GROUNDING=off|initial|adaptive`
139
+ controls model orientation: `initial` always grounds the first main-agent
140
+ request; `adaptive` also regrounds material direction changes. The main TUI
141
+ uses `adaptive`; child agents receive the parent slice rather than recursively
142
+ running a second grounding pass.
143
+ `shadow` computes and records checkpoints without model injection. `adaptive`
144
+ is the default target: inject on meaningful state changes and bounded cadence.
145
+
146
+ No mandatory `trajectory_update` tool is introduced. A tool would create
147
+ ceremony and failure loops for smaller models. Existing todo/workboard state is
148
+ updated only when the trajectory materially changes.
149
+
150
+ ## Acceptance criteria
151
+
152
+ - Changed-payload full-write retries produce `recovery_required`, a fresh read
153
+ prerequisite, and a no-repeat constraint.
154
+ - Cached/replayed reads do not satisfy a fresh-read prerequisite.
155
+ - Child completion yields verification work, not an unsupported completion
156
+ claim.
157
+ - Context contains at most one current checkpoint and compaction retains its
158
+ current evidence references.
159
+ - TUI/API/debug views all agree on checkpoint revision and assessment.
160
+ - The feature is bounded, sanitized, and disabled cleanly by configuration.
@@ -0,0 +1,489 @@
1
+ # Voice TTS Flow — Architecture & Implementation Guide
2
+
3
+ This document describes the voice synthesis system in omnius: how it's built,
4
+ how all the pieces connect, and how to add new voice features following the same patterns.
5
+
6
+ ---
7
+
8
+ ## System Overview
9
+
10
+ ```
11
+ ┌─────────────┐ ┌──────────────────┐ ┌──────────────┐
12
+ │ interactive │────▶│ VoiceEngine │────▶│ System Audio │
13
+ │ .ts │ │ voice.ts │ │ (afplay/ │
14
+ │ │ │ │ │ paplay) │
15
+ │ ┌────────┐ │ │ ┌────────────┐ │ └──────────────┘
16
+ │ │ Agent │──┤ │ │ ONNX Model │ │
17
+ │ │ Events │ │ │ │ Session │ │ ┌──────────────┐
18
+ │ └────────┘ │ │ └────────────┘ │────▶│ WebSocket │
19
+ │ │ │ ┌────────────┐ │ │ Clients │
20
+ │ ┌────────┐ │ │ │ MLX Audio │ │ │ (voice- │
21
+ │ │Emotion │──┤ │ │ Backend │ │ │ session.ts) │
22
+ │ │Context │ │ │ └────────────┘ │ └──────────────┘
23
+ │ └────────┘ │ └──────────────────┘
24
+ │ │ ▲
25
+ │ ┌────────┐ │ │
26
+ │ │Narrate │──┘ ┌────────┴────────┐
27
+ │ │ Engine │ │ Model Registry │
28
+ │ └────────┘ │ glados/over- │
29
+ └─────────────┘ │ watch/kokoro │
30
+ └─────────────────┘
31
+ ```
32
+
33
+ ## File Map
34
+
35
+ | File | Purpose | Lines |
36
+ |------|---------|-------|
37
+ | `packages/cli/src/tui/voice.ts` | VoiceEngine class + narration engine | ~2260 |
38
+ | `packages/cli/src/tui/voice-session.ts` | WebSocket streaming + cloudflared tunnel | ~885 |
39
+ | `packages/cli/src/tui/interactive.ts` | Wiring: instantiation, events, narration calls | ~2300 |
40
+ | `packages/cli/src/tui/commands.ts` | `/voice` command handler | ~1850 |
41
+ | `packages/cli/src/tui/render.ts` | `renderInfo`/`renderWarning` + content-write hook | ~670 |
42
+ | `packages/cli/tests/voice-narration.test.ts` | Narration unit tests | varies |
43
+ | `packages/cli/tests/voice-session.test.ts` | Session unit tests | varies |
44
+
45
+ ---
46
+
47
+ ## 1. Model Registry
48
+
49
+ All voice models are defined in `VOICE_MODELS` (voice.ts, top of file):
50
+
51
+ ```typescript
52
+ const VOICE_MODELS: Record<string, VoiceModel> = {
53
+ glados: { id: "glados", backend: "onnx", onnxUrl: "...", configUrl: "..." },
54
+ overwatch:{ id: "overwatch", backend: "onnx", onnxUrl: "...", configUrl: "..." },
55
+ kokoro: { id: "kokoro", backend: "mlx", mlxModelId: "mlx-community/Kokoro-82M-bf16", mlxVoice: "af_heart" },
56
+ // ... more kokoro voices
57
+ };
58
+ ```
59
+
60
+ **To add a new model:**
61
+
62
+ 1. Add an entry to `VOICE_MODELS` with the appropriate `backend` field
63
+ 2. For ONNX: provide `onnxUrl` + `configUrl` (Piper ONNX format)
64
+ 3. For MLX: provide `mlxModelId` (Hugging Face repo) + `mlxVoice` + `mlxLangCode`
65
+ 4. The model is immediately available via `/voice <id>`
66
+
67
+ ---
68
+
69
+ ## 2. VoiceEngine Lifecycle
70
+
71
+ ### Instantiation
72
+
73
+ A single `VoiceEngine` instance is created in `interactive.ts:1359`:
74
+
75
+ ```typescript
76
+ const voiceEngine = new VoiceEngine();
77
+ ```
78
+
79
+ ### Startup
80
+
81
+ If the user previously enabled voice (persisted in settings):
82
+
83
+ ```typescript
84
+ if (savedSettings.voice) {
85
+ voiceEngine.toggle(); // Enable TTS
86
+ if (savedSettings.voiceModel)
87
+ voiceEngine.setModel(savedSettings.voiceModel); // Load specific model
88
+ }
89
+ ```
90
+
91
+ ### toggle() Flow
92
+
93
+ ```
94
+ toggle()
95
+ ├─ If disabling: set enabled=false, killPlayback(), done
96
+ └─ If enabling:
97
+ ├─ MLX model?
98
+ │ └─ ensureMlxAudio() → pip install mlx-audio
99
+ └─ ONNX model?
100
+ ├─ ensureRuntime() → npm install onnxruntime-node + phonemizer
101
+ ├─ ensureModel(id) → download .onnx + .json from GitHub
102
+ └─ loadSession() → create InferenceSession
103
+ set enabled=true, ready=true
104
+ ```
105
+
106
+ ### setModel(id) Flow
107
+
108
+ ```
109
+ setModel(id)
110
+ ├─ Validate id exists in VOICE_MODELS
111
+ ├─ Reset: session=null, config=null, ready=false
112
+ └─ If currently enabled:
113
+ └─ Load the new model (same as toggle enable path)
114
+ ```
115
+
116
+ ### Shutdown
117
+
118
+ ```typescript
119
+ voiceEngine.dispose(); // in interactive.ts cleanup
120
+ ```
121
+
122
+ ---
123
+
124
+ ## 3. Synthesis Pipeline
125
+
126
+ ### ONNX Path (glados, overwatch)
127
+
128
+ ```
129
+ text
130
+ │
131
+ ├── chunkText() Split >200 char text on newlines + sentence boundaries
132
+ │
133
+ ├── For each chunk:
134
+ │ ├── textToPhonemes() espeak-ng WASM phonemization
135
+ │ ├── phonemesToIds() Map phonemes → integer IDs via config.phoneme_id_map
136
+ │ │ Format: BOS → PAD → (phoneme + PAD)* → EOS
137
+ │ ├── Build tensors: input (int64), input_lengths (int64), scales (float32)
138
+ │ ├── session.run() ONNX inference → Float32 audio samples
139
+ │ └── Concatenate with 180ms silence gaps between sentences
140
+ │
141
+ ├── Apply volume scaling (0.0–1.0 multiplier per sample)
142
+ ├── Apply pitch shift (linear-interpolation resampling)
143
+ │
144
+ ├── Stream PCM to WebSocket clients (if onPCMOutput wired)
145
+ ├── Write WAV to temp file
146
+ ├── Play via system command (afplay / paplay / pw-play / aplay)
147
+ └── Delete temp file
148
+ ```
149
+
150
+ ### MLX Path (kokoro, kokoro:af_heart, etc.)
151
+
152
+ ```
153
+ text
154
+ │
155
+ ├── Clean markdown (* removal)
156
+ ├── Build Python command:
157
+ │ python3 -c "from mlx_audio.tts import generate; generate.main([...])"
158
+ │ --model mlx-community/Kokoro-82M-bf16
159
+ │ --text "..."
160
+ │ --voice af_heart
161
+ │ --lang_code a
162
+ │ --audio_path /tmp/omnius-mlx-{ts}.wav
163
+ │
164
+ ├── execSync() with 60s timeout
165
+ │ (fallback: python3 -m mlx_audio.tts.generate CLI)
166
+ │
167
+ ├── Apply volume scaling (rewrite WAV PCM samples)
168
+ ├── Stream PCM to WebSocket clients (parse WAV header for sample rate)
169
+ ├── Play via system command
170
+ └── Delete temp file
171
+ ```
172
+
173
+ ---
174
+
175
+ ## 4. Narration Engine
176
+
177
+ The narration system generates context-aware spoken descriptions of agent activity.
178
+ It lives in voice.ts below the VoiceEngine class (~line 1268+).
179
+
180
+ ### Personality Levels
181
+
182
+ ```
183
+ 1 = minimal "Reading file.ts"
184
+ 2 = brief "Reading file.ts"
185
+ 3 = conv "Let me take a look at file.ts"
186
+ 4 = chatty "Alright, let's crack open file.ts"
187
+ 5 = theatrical "Alright, let's crack open file.ts and see what we're working with"
188
+ ```
189
+
190
+ Mapped from the agent's personality preset:
191
+ ```typescript
192
+ { concise: 1, balanced: 3, verbose: 4, pedagogical: 5 }
193
+ ```
194
+
195
+ ### describeToolCall(toolName, args, level, emotion)
196
+
197
+ Called from `interactive.ts` on every `tool_call` event:
198
+
199
+ ```typescript
200
+ if (voice?.enabled) {
201
+ const desc = describeToolCall(event.toolName, event.toolArgs, vLevel, emoCtx);
202
+ voice.speakSubordinate(desc, emoCtx); // 55% volume, 0.92x pitch
203
+ }
204
+ ```
205
+
206
+ **Variant pools** — each tool has 3 tiers of phrasings:
207
+ - `FILE_READ_VARIANTS.terse` / `.conv` / `.chatty`
208
+ - `FILE_WRITE_VARIANTS`, `FILE_EDIT_VARIANTS`, `GREP_VARIANTS`, etc.
209
+ - `SHELL_VARIANTS` — categorized by shell command type (git, npm, test, etc.)
210
+
211
+ **Context modifiers** applied based on narration state:
212
+ - After errors: prefix with "Okay, " or "Right, "
213
+ - Same file again: "Back to " or "Still working on "
214
+ - Revisiting a file: "Coming back to " or "Revisiting "
215
+ - Progress beats every 8 tools: "Making good progress. "
216
+
217
+ ### describeToolResult(toolName, success, level, content, emotion)
218
+
219
+ Called on every `tool_result` event. Generates success/failure descriptions:
220
+ - Success: "Got it", "Done", "Found it"
221
+ - Failure: "That didn't work", "Hit a snag"
222
+
223
+ **Content-aware extraction** via `extractResultDigest()`:
224
+ - ETH balances, test results, error messages, wallet addresses, file paths
225
+
226
+ ### describeTaskComplete(summary, complete, level)
227
+
228
+ Announces when the agent finishes a task:
229
+ - "Task complete", "Got it done", "All set"
230
+
231
+ ### Narration State
232
+
233
+ ```typescript
234
+ interface NarrationState {
235
+ toolCount: number; // Total tools this session
236
+ toolCounts: Record<string, number>; // Per-tool counts
237
+ consecutiveErrors: number; // Reset on success
238
+ totalErrors: number;
239
+ lastTool: string;
240
+ lastFile: string;
241
+ filesSeen: Set<string>; // All files visited
242
+ lastVariantIdx: Record<string, number>; // Avoid repeats
243
+ lastResultDigest: string; // Last tool result summary
244
+ }
245
+ ```
246
+
247
+ `resetNarrationContext()` is called at the start of each task.
248
+
249
+ ### pick(key, variants)
250
+
251
+ Selects a random variant from a pool, avoiding the last-used index for that key.
252
+ Ensures you never hear the same phrasing twice in a row.
253
+
254
+ ---
255
+
256
+ ## 5. Emotion Modulation
257
+
258
+ The emotion engine provides valence-arousal context to the voice:
259
+
260
+ ```typescript
261
+ interface VoiceEmotionContext {
262
+ valence: number; // -1 (sad) to +1 (happy)
263
+ arousal: number; // 0 (calm) to 1 (activated)
264
+ label: string; // "excited", "focused", etc.
265
+ emoji: string;
266
+ }
267
+ ```
268
+
269
+ ### Pitch Bias
270
+
271
+ `emotionToPitchBias(emotion)` converts emotion to a pitch adjustment:
272
+
273
+ ```
274
+ pitch_bias = valence × 0.6 + (arousal - 0.5) × 0.4
275
+ clamped to [-0.10, +0.10]
276
+ ```
277
+
278
+ - **Excited** (high valence + high arousal) → voice pitch rises
279
+ - **Dejected** (low valence + low arousal) → voice pitch drops
280
+
281
+ Applied in `speak()` and `speakSubordinate()`:
282
+
283
+ ```typescript
284
+ speak(text, emotion): pitchFactor = 1.0 + pitchBias
285
+ speakSubordinate(text): pitchFactor = 0.92 + pitchBias (lower base pitch)
286
+ ```
287
+
288
+ ### Emotion Coloring
289
+
290
+ At personality >= 3, ~30% of narrations get emotion-colored prefixes:
291
+ - Excited: "Feeling good about this"
292
+ - Stressed: "Pushing through"
293
+ - Calm: "Nice and steady"
294
+ - Subdued: "Being careful here"
295
+
296
+ ---
297
+
298
+ ## 6. Queue & Playback
299
+
300
+ ### Speech Queue
301
+
302
+ ```typescript
303
+ private speakQueue: SpeakItem[] = [];
304
+ ```
305
+
306
+ Items are queued FIFO. `drainQueue()` processes them sequentially:
307
+ 1. Pop item from front
308
+ 2. Synthesize + play to completion
309
+ 3. 250ms silence gap
310
+ 4. Next item
311
+
312
+ Queue overflow protection: if > 30 items backed up, clear the queue.
313
+
314
+ ### Volume Levels
315
+
316
+ ```
317
+ speak() → volume 1.0 (full)
318
+ speakSubordinate() → volume 0.55 (reduced for tool narration)
319
+ ```
320
+
321
+ ### Playback
322
+
323
+ WAV temp file → system audio command:
324
+ - macOS: `afplay`
325
+ - Linux: `paplay` → `pw-play` → `aplay` (tries in order)
326
+ - Windows: PowerShell `Media.SoundPlayer`
327
+
328
+ 15-second safety timeout per playback. `killPlayback()` sends SIGTERM.
329
+
330
+ ---
331
+
332
+ ## 7. WebSocket Voice Session
333
+
334
+ `VoiceSession` (voice-session.ts) enables real-time voice interaction via browser:
335
+
336
+ ```
337
+ Browser Client ←──WebSocket──→ VoiceSession ←──PCM──→ VoiceEngine
338
+ (mic) binary PCM (server) onPCMOutput (TTS)
339
+ (speaker) binary PCM
340
+ ```
341
+
342
+ ### Setup
343
+
344
+ ```
345
+ start()
346
+ ├── Start HTTP server (serves HTML single-page app)
347
+ ├── Start WebSocket server (ws)
348
+ ├── Launch cloudflared tunnel for public URL
349
+ └── Wire VoiceEngine.onPCMOutput → broadcast to all clients
350
+ ```
351
+
352
+ ### Audio Format
353
+
354
+ - 16kHz, 16-bit, mono PCM (Int16Array)
355
+ - Binary WebSocket frames
356
+ - Echo cancellation: suppress mic input while TTS is playing
357
+
358
+ ### Frontend
359
+
360
+ Embedded HTML with:
361
+ - Braille waveform animator (color-coded: idle/listening/speaking)
362
+ - WebAudio API for mic capture + speaker playback
363
+ - Transcript view (user + agent messages)
364
+ - Start/stop mic button
365
+
366
+ ---
367
+
368
+ ## 8. Settings Persistence
369
+
370
+ Voice settings are saved per-project or globally via `resolveSettings()`:
371
+
372
+ ```typescript
373
+ {
374
+ voice: boolean, // Enabled/disabled
375
+ voiceModel: string // Model ID ("glados", "kokoro:af_heart", etc.)
376
+ }
377
+ ```
378
+
379
+ The `/voice` command handler in commands.ts:
380
+ ```typescript
381
+ case "voice":
382
+ if (arg) {
383
+ ctx.voiceSetModel(arg); // /voice kokoro
384
+ save({ voice: true, voiceModel: arg });
385
+ } else {
386
+ ctx.voiceToggle(); // /voice (toggle on/off)
387
+ save({ voice: isOn });
388
+ }
389
+ ```
390
+
391
+ ---
392
+
393
+ ## 9. Content-Write Hook (TUI Safety)
394
+
395
+ Voice operations can trigger `renderInfo()` / `renderWarning()` messages asynchronously.
396
+ To prevent these from overwriting the TUI input area, a global content-write hook
397
+ brackets all render calls with scroll-region management:
398
+
399
+ ```typescript
400
+ // render.ts
401
+ export function renderInfo(message: string): void {
402
+ _contentWriteHook?.begin(); // → statusBar.beginContentWrite()
403
+ process.stdout.write(`ℹ ${message}\n`);
404
+ _contentWriteHook?.end(); // → statusBar.endContentWrite()
405
+ }
406
+ ```
407
+
408
+ Registered once in interactive.ts after StatusBar activation:
409
+ ```typescript
410
+ setContentWriteHook({
411
+ begin: () => statusBar.beginContentWrite(),
412
+ end: () => statusBar.endContentWrite(),
413
+ });
414
+ ```
415
+
416
+ **Rule:** Never use raw `process.stdout.write()` in voice.ts for user-facing messages.
417
+ Always use `renderInfo()` / `renderWarning()` / `renderError()` so the content hook
418
+ routes output to the scroll region above the status bar and input area.
419
+
420
+ ---
421
+
422
+ ## 10. How to Add a New Voice Feature
423
+
424
+ ### Adding a New ONNX Voice Model
425
+
426
+ 1. Host the `.onnx` + `.onnx.json` files (Piper format) on a public URL
427
+ 2. Add entry to `VOICE_MODELS`:
428
+ ```typescript
429
+ myvoice: {
430
+ id: "myvoice", label: "My Voice", backend: "onnx",
431
+ onnxUrl: "https://...", configUrl: "https://...",
432
+ },
433
+ ```
434
+ 3. Done. `/voice myvoice` works immediately.
435
+
436
+ ### Adding a New MLX Voice
437
+
438
+ 1. Find the Hugging Face model ID (must be MLX-compatible)
439
+ 2. Add entry to `VOICE_MODELS`:
440
+ ```typescript
441
+ "newmodel:voice_name": {
442
+ id: "newmodel:voice_name", label: "New Model (MLX)", backend: "mlx",
443
+ mlxModelId: "mlx-community/NewModel", mlxVoice: "voice_name", mlxLangCode: "a",
444
+ onnxUrl: "", configUrl: "",
445
+ },
446
+ ```
447
+ 3. Done. `/voice newmodel:voice_name` works on macOS Apple Silicon.
448
+
449
+ ### Adding a New TTS Backend
450
+
451
+ 1. Add a new `backend` value to the `VoiceModel` interface
452
+ 2. Add a method like `ensureNewBackend()` for installation
453
+ 3. Add a method like `synthesizeWithNewBackend()` for synthesis
454
+ 4. Gate in `toggle()`, `setModel()`, `synthesizeAndPlay()`, `synthesizeToBuffer()`, `synthesizeToPCM()`
455
+ 5. Follow the pattern: the backend outputs a WAV file, then existing `playWav()` handles playback
456
+
457
+ ### Adding New Narration Variants
458
+
459
+ 1. Find the relevant variant pool (e.g., `FILE_READ_VARIANTS`)
460
+ 2. Add new strings to the `terse`, `conv`, and `chatty` tiers
461
+ 3. The `pick()` function automatically rotates through variants
462
+
463
+ ### Adding Emotion-Aware Features
464
+
465
+ 1. Receive `VoiceEmotionContext` from the emotion engine
466
+ 2. Use `emotionToPitchBias()` for pitch modulation
467
+ 3. Use `emotionColor()` for spoken prefixes
468
+ 4. Valence drives tone (happy/sad), arousal drives energy (calm/activated)
469
+
470
+ ---
471
+
472
+ ## Quick Reference: Key Functions
473
+
474
+ | Function | Location | Purpose |
475
+ |----------|----------|---------|
476
+ | `VoiceEngine.toggle()` | voice.ts | Enable/disable TTS |
477
+ | `VoiceEngine.speak()` | voice.ts | Queue speech (full volume) |
478
+ | `VoiceEngine.speakSubordinate()` | voice.ts | Queue speech (55% volume) |
479
+ | `VoiceEngine.synthesizeAndPlay()` | voice.ts | Core synthesis + playback |
480
+ | `VoiceEngine.synthesizeWithMlx()` | voice.ts | MLX backend synthesis |
481
+ | `VoiceEngine.ensureRuntime()` | voice.ts | Install ONNX runtime |
482
+ | `VoiceEngine.ensureMlxAudio()` | voice.ts | Install mlx-audio pip package |
483
+ | `describeToolCall()` | voice.ts | Generate tool narration text |
484
+ | `describeToolResult()` | voice.ts | Generate result narration text |
485
+ | `describeTaskComplete()` | voice.ts | Generate completion narration |
486
+ | `emotionToPitchBias()` | voice.ts | Emotion → pitch modulation |
487
+ | `pick()` | voice.ts | Random variant selection (no repeats) |
488
+ | `setContentWriteHook()` | render.ts | Register TUI scroll-region safety |
489
+ | `VoiceSession.start()` | voice-session.ts | Start WebSocket voice server |