omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,799 @@
|
|
|
1
|
+
# Model Capability Awareness And Multimodal Identity Memory Root Fix
|
|
2
|
+
|
|
3
|
+
Date: 2026-05-15
|
|
4
|
+
|
|
5
|
+
Status: planning and integration tracker
|
|
6
|
+
|
|
7
|
+
Latest integration pass: 2026-05-15
|
|
8
|
+
|
|
9
|
+
- Implemented the Telegram/TUI visual identity memory root path first.
|
|
10
|
+
- Added `identity_memory` in `packages/cli/src/tui/identity-memory-tool.ts`.
|
|
11
|
+
- Wired it into TUI sub-agent tools and Telegram admin/public action tools.
|
|
12
|
+
- Repaired CLIP-only zettelkasten neighbor linking in
|
|
13
|
+
`packages/memory/src/zettelkasten.ts`.
|
|
14
|
+
- Added regression tests:
|
|
15
|
+
`packages/memory/tests/zettelkasten.test.ts` and
|
|
16
|
+
`packages/cli/tests/identity-memory-tool.test.ts`.
|
|
17
|
+
|
|
18
|
+
## Purpose
|
|
19
|
+
|
|
20
|
+
Omnius already has substantial capability metadata for creative models, voice,
|
|
21
|
+
image/audio generation, Telegram scoped media, CLIP embeddings, face memory,
|
|
22
|
+
and graph-backed multimodal identity. The current weakness is that this data is
|
|
23
|
+
not consistently exposed to the model as runtime-aware, inspectable context.
|
|
24
|
+
|
|
25
|
+
The root fix is to make the agent loop aware of:
|
|
26
|
+
|
|
27
|
+
- which voice, image, music, and sound models are available;
|
|
28
|
+
- which of those models are currently selected for this invocation;
|
|
29
|
+
- model details such as backend, size class, quality notes, hardware fit,
|
|
30
|
+
cached/downloaded status, setup/prewarm commands, and runtime limitations;
|
|
31
|
+
- which tools are actually available in the current invocation after policy
|
|
32
|
+
filtering;
|
|
33
|
+
- how to inspect deeper structures behind those tools without relying on
|
|
34
|
+
brittle keyword heuristics;
|
|
35
|
+
- how to write natural multimodal identity memories through model-decided tool
|
|
36
|
+
calls, not regex parsing.
|
|
37
|
+
|
|
38
|
+
This document is the tracking artifact for that integration.
|
|
39
|
+
|
|
40
|
+
## Hard Constraints
|
|
41
|
+
|
|
42
|
+
- Do not solve natural memory requests with regex or keyword aliases.
|
|
43
|
+
- The model must decide to call identity/capability tools from full context.
|
|
44
|
+
- Tool implementations may validate inputs and evidence, but must not silently
|
|
45
|
+
infer identity without explicit user-provided evidence.
|
|
46
|
+
- Public Telegram scope must remain per-chat and policy filtered.
|
|
47
|
+
- Private/admin/TUI/global context must not leak into public group context.
|
|
48
|
+
- Capability awareness must describe the current invocation, not a stale static
|
|
49
|
+
list.
|
|
50
|
+
- Shared services should feed TUI, Telegram, GUI/API, and future voice sessions.
|
|
51
|
+
|
|
52
|
+
## Existing Anchors
|
|
53
|
+
|
|
54
|
+
### Creative Model Registries
|
|
55
|
+
|
|
56
|
+
Image model metadata already exists in:
|
|
57
|
+
|
|
58
|
+
- `packages/execution/src/tools/image-generate.ts:24`
|
|
59
|
+
- `ImageGenerationPreset`
|
|
60
|
+
- fields: `id`, `label`, `backend`, `install`, `category`, `sizeClass`,
|
|
61
|
+
`quality`, `minVramGB`, `recommendedVramGB`, `deployment`, `steps`,
|
|
62
|
+
`guidance`, `width`, `height`, `fallbackFor`, `note`
|
|
63
|
+
- `packages/execution/src/tools/image-generate.ts:59`
|
|
64
|
+
- `DEFAULT_DIFFUSERS_IMAGE_MODEL`
|
|
65
|
+
- `packages/execution/src/tools/image-generate.ts:120`
|
|
66
|
+
- `IMAGE_GENERATION_MODEL_PRESETS`
|
|
67
|
+
- `packages/execution/src/tools/image-generate.ts:727`
|
|
68
|
+
- `getImageGenerationPreset`
|
|
69
|
+
- `packages/execution/src/tools/image-generate.ts:733`
|
|
70
|
+
- `imageGenerationQualityLadder`
|
|
71
|
+
- `packages/execution/src/tools/image-generate.ts:739`
|
|
72
|
+
- `imageGenerationFallbackAlternates`
|
|
73
|
+
- `packages/execution/src/tools/image-generate.ts:1145`
|
|
74
|
+
- `ImageGenerateTool`
|
|
75
|
+
- `packages/execution/src/tools/image-generate.ts:1248`
|
|
76
|
+
- `ImageGenerateTool` supports `action=list_models`, but the schema still
|
|
77
|
+
marks `prompt` as required and output is shallow.
|
|
78
|
+
|
|
79
|
+
Sound/music model metadata already exists in:
|
|
80
|
+
|
|
81
|
+
- `packages/execution/src/tools/audio-generate.ts:27`
|
|
82
|
+
- `AudioGenerationPreset`
|
|
83
|
+
- fields: `id`, `label`, `kind`, `backend`, `install`, `category`,
|
|
84
|
+
`sizeClass`, `quality`, `output`, `bestUse`, `minVramGB`,
|
|
85
|
+
`recommendedVramGB`, `deployment`, `defaultDurationSec`, `defaultSteps`,
|
|
86
|
+
`note`
|
|
87
|
+
- `packages/execution/src/tools/audio-generate.ts:79`
|
|
88
|
+
- `DEFAULT_SOUND_MODEL`
|
|
89
|
+
- `packages/execution/src/tools/audio-generate.ts:80`
|
|
90
|
+
- `DEFAULT_MUSIC_MODEL`
|
|
91
|
+
- `packages/execution/src/tools/audio-generate.ts:135`
|
|
92
|
+
- `AUDIO_GENERATION_MODEL_PRESETS`
|
|
93
|
+
- `packages/execution/src/tools/audio-generate.ts:1420`
|
|
94
|
+
- `getAudioGenerationPreset`
|
|
95
|
+
- `packages/execution/src/tools/audio-generate.ts:1426`
|
|
96
|
+
- `audioGenerationQualityLadder`
|
|
97
|
+
- `packages/execution/src/tools/audio-generate.ts:1648`
|
|
98
|
+
- `AudioGenerateTool`
|
|
99
|
+
- `packages/execution/src/tools/audio-generate.ts:1790`
|
|
100
|
+
- `AudioGenerateTool` supports `action=list_models`, but the schema does not
|
|
101
|
+
expose `action`.
|
|
102
|
+
|
|
103
|
+
Voice model metadata exists in:
|
|
104
|
+
|
|
105
|
+
- `packages/cli/src/tui/voice.ts:171`
|
|
106
|
+
- `VOICE_MODELS`
|
|
107
|
+
- `packages/cli/src/tui/voice.ts:247`
|
|
108
|
+
- `listVoiceModels`
|
|
109
|
+
- `packages/cli/src/tui/voice.ts:271`
|
|
110
|
+
- `getSupertonicVoiceOptions`
|
|
111
|
+
- `packages/cli/src/api/serve.ts:5355`
|
|
112
|
+
- `GET /v1/voice/models`
|
|
113
|
+
- `packages/cli/src/api/serve.ts:5359`
|
|
114
|
+
- `GET /v1/voice/supertonic-settings`
|
|
115
|
+
- `packages/cli/src/api/serve.ts:5390`
|
|
116
|
+
- `POST /v1/voice/models/switch`
|
|
117
|
+
- `packages/execution/src/tools/audio-playback.ts:610`
|
|
118
|
+
- `AudioPlaybackTool`
|
|
119
|
+
- `packages/execution/src/tools/audio-playback.ts:797`
|
|
120
|
+
- `audio_playback action=list_voices`
|
|
121
|
+
- `packages/execution/src/tools/audio-playback.ts:1096`
|
|
122
|
+
- `TtsGenerateTool`
|
|
123
|
+
|
|
124
|
+
### Current Model Selections
|
|
125
|
+
|
|
126
|
+
Settings schema:
|
|
127
|
+
|
|
128
|
+
- `packages/cli/src/tui/omnius-directory.ts:171`
|
|
129
|
+
- `OmniusSettings`
|
|
130
|
+
- `packages/cli/src/tui/omnius-directory.ts:182`
|
|
131
|
+
- `voiceModel`
|
|
132
|
+
- `packages/cli/src/tui/omnius-directory.ts:216`
|
|
133
|
+
- `imageModel`
|
|
134
|
+
- `packages/cli/src/tui/omnius-directory.ts:218`
|
|
135
|
+
- `imageBackend`
|
|
136
|
+
- `packages/cli/src/tui/omnius-directory.ts:220`
|
|
137
|
+
- `soundModel`
|
|
138
|
+
- `packages/cli/src/tui/omnius-directory.ts:222`
|
|
139
|
+
- `soundBackend`
|
|
140
|
+
- `packages/cli/src/tui/omnius-directory.ts:224`
|
|
141
|
+
- `musicModel`
|
|
142
|
+
- `packages/cli/src/tui/omnius-directory.ts:226`
|
|
143
|
+
- `musicBackend`
|
|
144
|
+
- `packages/cli/src/tui/omnius-directory.ts:282`
|
|
145
|
+
- `resolveSettings`
|
|
146
|
+
|
|
147
|
+
TUI configured defaults:
|
|
148
|
+
|
|
149
|
+
- `packages/cli/src/tui/interactive.ts:982`
|
|
150
|
+
- `imageGenerationDefaultsForRepo`
|
|
151
|
+
- `packages/cli/src/tui/interactive.ts:992`
|
|
152
|
+
- `createConfiguredImageGenerateTool`
|
|
153
|
+
- `packages/cli/src/tui/interactive.ts:996`
|
|
154
|
+
- `audioGenerationDefaultsForRepo`
|
|
155
|
+
- `packages/cli/src/tui/interactive.ts:1010`
|
|
156
|
+
- `createConfiguredAudioGenerateTool`
|
|
157
|
+
|
|
158
|
+
Telegram configured defaults:
|
|
159
|
+
|
|
160
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5106`
|
|
161
|
+
- `imageGenerationDefaultsForRepo`
|
|
162
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5116`
|
|
163
|
+
- `audioGenerationDefaultsForRepo`
|
|
164
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4939`
|
|
165
|
+
- admin Telegram tools instantiate `ImageGenerateTool`
|
|
166
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4940`
|
|
167
|
+
- admin Telegram tools instantiate `AudioGenerateTool`
|
|
168
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4975`
|
|
169
|
+
- public Telegram appends scoped creative tools
|
|
170
|
+
|
|
171
|
+
### Hardware Rating And Disk Status
|
|
172
|
+
|
|
173
|
+
The richer model scoring is trapped inside TUI command rendering:
|
|
174
|
+
|
|
175
|
+
- `packages/cli/src/tui/commands.ts:8689`
|
|
176
|
+
- `rateImagePresetForHardware`
|
|
177
|
+
- `packages/cli/src/tui/commands.ts:8772`
|
|
178
|
+
- `renderImageModelList`
|
|
179
|
+
- `packages/cli/src/tui/commands.ts:8876`
|
|
180
|
+
- `fetchOllamaModelSizes`
|
|
181
|
+
- `packages/cli/src/tui/commands.ts:8904`
|
|
182
|
+
- `imageModelDiskStats`
|
|
183
|
+
- `packages/cli/src/tui/commands.ts:8914`
|
|
184
|
+
- `audioModelDiskStats`
|
|
185
|
+
- `packages/cli/src/tui/commands.ts:9193`
|
|
186
|
+
- `rateAudioPresetForHardware`
|
|
187
|
+
- `packages/cli/src/tui/commands.ts:9240`
|
|
188
|
+
- `renderAudioModelList`
|
|
189
|
+
|
|
190
|
+
This logic needs to move into a shared service that can be used by:
|
|
191
|
+
|
|
192
|
+
- TUI slash commands;
|
|
193
|
+
- Telegram context injection;
|
|
194
|
+
- GUI/API endpoints;
|
|
195
|
+
- agent-facing capability inspection tool.
|
|
196
|
+
|
|
197
|
+
### Agent Tool Awareness
|
|
198
|
+
|
|
199
|
+
Static discovery:
|
|
200
|
+
|
|
201
|
+
- `packages/execution/src/tools/explore-tools.ts:20`
|
|
202
|
+
- static `TOOL_CATALOG`
|
|
203
|
+
- `packages/execution/src/tools/explore-tools.ts:82`
|
|
204
|
+
- `ExploreToolsTool`
|
|
205
|
+
|
|
206
|
+
Runtime-aware discovery:
|
|
207
|
+
|
|
208
|
+
- `packages/orchestrator/src/agenticRunner.ts:1372`
|
|
209
|
+
- promoted/deferred tool tracking
|
|
210
|
+
- `packages/orchestrator/src/agenticRunner.ts:17376`
|
|
211
|
+
- `buildToolDefinitions`
|
|
212
|
+
- `packages/orchestrator/src/agenticRunner.ts:17630`
|
|
213
|
+
- `tool_search` pseudo-tool
|
|
214
|
+
- `packages/orchestrator/src/agenticRunner.ts:17659`
|
|
215
|
+
- runtime handler for `tool_search`
|
|
216
|
+
|
|
217
|
+
Post-compaction awareness:
|
|
218
|
+
|
|
219
|
+
- `packages/orchestrator/src/agenticRunner.ts:15506`
|
|
220
|
+
- post-compaction skill and MCP tool re-injection
|
|
221
|
+
|
|
222
|
+
Current gap:
|
|
223
|
+
|
|
224
|
+
- `tool_search` reflects runtime registered tools, but only as deferred tool
|
|
225
|
+
one-liners.
|
|
226
|
+
- `ExploreToolsTool` is static and can drift from the actual invocation.
|
|
227
|
+
- Neither path exposes selected models, model detail, embedding status, scoped
|
|
228
|
+
media affordances, or deep inspection hooks.
|
|
229
|
+
|
|
230
|
+
### Invocation Context Injection
|
|
231
|
+
|
|
232
|
+
TUI dynamic context:
|
|
233
|
+
|
|
234
|
+
- `packages/cli/src/tui/interactive.ts:2163`
|
|
235
|
+
- project context loaded
|
|
236
|
+
- `packages/cli/src/tui/interactive.ts:2165`
|
|
237
|
+
- `dynamicContext`
|
|
238
|
+
- `packages/cli/src/tui/interactive.ts:2188`
|
|
239
|
+
- `environmentProvider`
|
|
240
|
+
- `packages/cli/src/tui/interactive.ts:2277`
|
|
241
|
+
- self-modify settings loaded
|
|
242
|
+
- `packages/cli/src/tui/interactive.ts:2312`
|
|
243
|
+
- vision capability guidance injection
|
|
244
|
+
- `packages/cli/src/tui/interactive.ts:2585`
|
|
245
|
+
- passed into `AgenticRunner`
|
|
246
|
+
|
|
247
|
+
Telegram dynamic context:
|
|
248
|
+
|
|
249
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4123`
|
|
250
|
+
- quick-chat message construction
|
|
251
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4145`
|
|
252
|
+
- runtime context in quick-chat system prompt
|
|
253
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4248`
|
|
254
|
+
- action/chat sub-agent session context
|
|
255
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4266`
|
|
256
|
+
- `AgenticRunner` creation
|
|
257
|
+
- `packages/cli/src/tui/telegram-bridge.ts:4276`
|
|
258
|
+
- Telegram `dynamicContext`
|
|
259
|
+
- `packages/cli/src/tui/telegram-bridge.ts:3278`
|
|
260
|
+
- `buildTelegramSessionContext`
|
|
261
|
+
- `packages/cli/src/tui/telegram-bridge.ts:3297`
|
|
262
|
+
- Telegram runtime context section
|
|
263
|
+
- `packages/cli/src/tui/telegram-bridge.ts:3315`
|
|
264
|
+
- admin group safety/context section
|
|
265
|
+
- `packages/cli/src/tui/telegram-bridge.ts:3317`
|
|
266
|
+
- public group safety/context section
|
|
267
|
+
|
|
268
|
+
Telegram reply/media context anchors:
|
|
269
|
+
|
|
270
|
+
- `packages/cli/src/tui/telegram-bridge.ts:204`
|
|
271
|
+
- `TelegramReplyContext`
|
|
272
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2110`
|
|
273
|
+
- `resolveTelegramReplyContext`
|
|
274
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2160`
|
|
275
|
+
- `buildTelegramCurrentReplyContext`
|
|
276
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2203`
|
|
277
|
+
- `formatTelegramCurrentMessageForPrompt`
|
|
278
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2793`
|
|
279
|
+
- `buildTelegramConversationContextStream`
|
|
280
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2853`
|
|
281
|
+
- recent media context section
|
|
282
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5048`
|
|
283
|
+
- `buildTelegramMediaRecentTool`
|
|
284
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5527`
|
|
285
|
+
- `processMediaContextForMessage`
|
|
286
|
+
|
|
287
|
+
### Multimodal Memory And Identity
|
|
288
|
+
|
|
289
|
+
Graph-backed multimodal identity:
|
|
290
|
+
|
|
291
|
+
- `packages/memory/src/multimodalIdentity.ts:6`
|
|
292
|
+
- source/scope types
|
|
293
|
+
- `packages/memory/src/multimodalIdentity.ts:35`
|
|
294
|
+
- media refs
|
|
295
|
+
- `packages/memory/src/multimodalIdentity.ts:45`
|
|
296
|
+
- `MultimodalIdentityAssertion`
|
|
297
|
+
- `packages/memory/src/multimodalIdentity.ts:53`
|
|
298
|
+
- `MultimodalEmbeddingBundle`
|
|
299
|
+
- `packages/memory/src/multimodalIdentity.ts:167`
|
|
300
|
+
- `MultimodalIdentityService`
|
|
301
|
+
- `packages/memory/src/multimodalIdentity.ts:178`
|
|
302
|
+
- `ingest`
|
|
303
|
+
- `packages/memory/src/multimodalIdentity.ts:269`
|
|
304
|
+
- identity assertion application
|
|
305
|
+
- `packages/memory/src/multimodalIdentity.ts:298`
|
|
306
|
+
- `applyIdentityAssertions`
|
|
307
|
+
- `packages/memory/src/multimodalIdentity.ts:330`
|
|
308
|
+
- embedding persistence
|
|
309
|
+
|
|
310
|
+
Knowledge graph relations:
|
|
311
|
+
|
|
312
|
+
- `packages/memory/src/temporalGraph.ts:24`
|
|
313
|
+
- `KGRelation`
|
|
314
|
+
- `packages/memory/src/temporalGraph.ts:27`
|
|
315
|
+
- reply/media identity relations: `authored_by`, `uploaded_by`,
|
|
316
|
+
`replied_to`, `depicts`, `named_as`
|
|
317
|
+
- `packages/memory/src/temporalGraph.ts:28`
|
|
318
|
+
- `face_match`, `voice_sample_of`, `speaker_candidate`,
|
|
319
|
+
`same_person_candidate`
|
|
320
|
+
|
|
321
|
+
Episode store:
|
|
322
|
+
|
|
323
|
+
- `packages/memory/src/episodeStore.ts:69`
|
|
324
|
+
- `Episode`
|
|
325
|
+
- `packages/memory/src/episodeStore.ts:80`
|
|
326
|
+
- native `embedding`
|
|
327
|
+
- `packages/memory/src/episodeStore.ts:83`
|
|
328
|
+
- `clipEmbedding`
|
|
329
|
+
- `packages/memory/src/episodeStore.ts:337`
|
|
330
|
+
- `clip_embedding` column
|
|
331
|
+
- `packages/memory/src/episodeStore.ts:397`
|
|
332
|
+
- automatic zettelkasten linking on insert
|
|
333
|
+
- `packages/memory/src/episodeStore.ts:625`
|
|
334
|
+
- `setEmbedding`
|
|
335
|
+
- `packages/memory/src/episodeStore.ts:631`
|
|
336
|
+
- `setClipEmbedding`
|
|
337
|
+
|
|
338
|
+
Zettelkasten:
|
|
339
|
+
|
|
340
|
+
- `packages/memory/src/zettelkasten.ts:72`
|
|
341
|
+
- `findNeighbors`
|
|
342
|
+
- `packages/memory/src/zettelkasten.ts:78`
|
|
343
|
+
- currently returns empty if the source episode has no native embedding
|
|
344
|
+
- `packages/memory/src/zettelkasten.ts:90`
|
|
345
|
+
- CLIP fallback only runs when native embedding lengths mismatch
|
|
346
|
+
- `packages/memory/src/zettelkasten.ts:114`
|
|
347
|
+
- `linkEpisode`
|
|
348
|
+
- `packages/memory/src/zettelkasten.ts:177`
|
|
349
|
+
- `batchLink`
|
|
350
|
+
|
|
351
|
+
Embedding workers:
|
|
352
|
+
|
|
353
|
+
- `packages/cli/src/api/embedding-workers.ts:41`
|
|
354
|
+
- `startEmbeddingWorkers`
|
|
355
|
+
- `packages/cli/src/api/embedding-workers.ts:64`
|
|
356
|
+
- scans visual/audio episodes missing native embeddings
|
|
357
|
+
- `packages/cli/src/api/embedding-workers.ts:107`
|
|
358
|
+
- stores native embedding
|
|
359
|
+
- `packages/cli/src/api/embedding-workers.ts:121`
|
|
360
|
+
- creates `EmbeddingAligner`
|
|
361
|
+
- `packages/cli/src/api/embedding-workers.ts:143`
|
|
362
|
+
- stores aligned embedding as `clipEmbedding`
|
|
363
|
+
- `packages/cli/src/api/embedding-workers.ts:148`
|
|
364
|
+
- coarse KG node insertion currently labels visual episodes as `person`
|
|
365
|
+
- `packages/cli/src/api/embedding-workers.ts:165`
|
|
366
|
+
- stale `person_node_id` co-occurrence metadata path
|
|
367
|
+
|
|
368
|
+
Embedding scripts:
|
|
369
|
+
|
|
370
|
+
- `scripts/embed-image.py:23`
|
|
371
|
+
- OpenCLIP ViT-B-32 image embedding
|
|
372
|
+
- `scripts/embed-text.py:14`
|
|
373
|
+
- OpenCLIP ViT-B-32 text embedding
|
|
374
|
+
- `scripts/embed-audio.py:45`
|
|
375
|
+
- SpeechBrain ECAPA speaker embedding
|
|
376
|
+
- `packages/cli/src/api/py-embed.ts:16`
|
|
377
|
+
- `ensureEmbedDeps`
|
|
378
|
+
- `packages/cli/src/api/py-embed.ts:41`
|
|
379
|
+
- full embedding deps require `OMNIUS_INSTALL_FULL_EMBED_DEPS=1`
|
|
380
|
+
- `packages/cli/src/api/py-embed.ts:55`
|
|
381
|
+
- `runEmbedImage`
|
|
382
|
+
- `packages/cli/src/api/py-embed.ts:66`
|
|
383
|
+
- `runEmbedAudio`
|
|
384
|
+
- `packages/cli/src/api/py-embed.ts:77`
|
|
385
|
+
- `runEmbedText`
|
|
386
|
+
|
|
387
|
+
Older visual/multimodal memory tools:
|
|
388
|
+
|
|
389
|
+
- `packages/execution/src/tools/visual-memory.ts:44`
|
|
390
|
+
- `VisualMemoryTool`
|
|
391
|
+
- `packages/execution/src/tools/visual-memory.ts:58`
|
|
392
|
+
- actions: `detect`, `enroll`, `identify`, `teach`, `recognize`, `list`,
|
|
393
|
+
`forget`
|
|
394
|
+
- `packages/execution/src/tools/visual-memory.ts:161`
|
|
395
|
+
- face enrollment
|
|
396
|
+
- `packages/execution/src/tools/multimodal-memory.ts:74`
|
|
397
|
+
- `MultimodalMemoryTool`
|
|
398
|
+
- `packages/execution/src/tools/multimodal-memory.ts:89`
|
|
399
|
+
- actions: `capture`, `meet`, `recall`, `timeline`
|
|
400
|
+
- `packages/execution/src/tools/multimodal-memory.ts:639`
|
|
401
|
+
- stores visual CLIP into `EpisodeStore.setClipEmbedding`
|
|
402
|
+
|
|
403
|
+
Current gap:
|
|
404
|
+
|
|
405
|
+
- The new graph-backed identity path is central, but older face and
|
|
406
|
+
multimodal-memory stores are still separate silos.
|
|
407
|
+
- Telegram media ingest writes scoped episodes, but there is no model-facing
|
|
408
|
+
identity tool for explicit user assertions such as "this person is Amy".
|
|
409
|
+
- Zettelkasten linking does not fully exploit CLIP-only visual episodes.
|
|
410
|
+
|
|
411
|
+
### Telegram Memory Ingest
|
|
412
|
+
|
|
413
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2227`
|
|
414
|
+
- sender normalization
|
|
415
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2247`
|
|
416
|
+
- memory scope
|
|
417
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2255`
|
|
418
|
+
- reply ref
|
|
419
|
+
- `packages/cli/src/tui/telegram-bridge.ts:2272`
|
|
420
|
+
- `telegramMemoryIngestPayload`
|
|
421
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5477`
|
|
422
|
+
- image ingest POST to `/v1/memory/ingest`
|
|
423
|
+
- `packages/cli/src/tui/telegram-bridge.ts:5506`
|
|
424
|
+
- audio/voice ingest POST to `/v1/memory/ingest`
|
|
425
|
+
- `packages/cli/src/api/serve.ts:8324`
|
|
426
|
+
- `/v1/memory/ingest`
|
|
427
|
+
- `packages/cli/src/api/serve.ts:8907`
|
|
428
|
+
- `handleMemoryIngest`
|
|
429
|
+
- `packages/cli/src/api/serve.ts:8947`
|
|
430
|
+
- identity assertion extraction from API payload
|
|
431
|
+
- `packages/cli/src/api/serve.ts:8997`
|
|
432
|
+
- `MultimodalIdentityService`
|
|
433
|
+
|
|
434
|
+
## Root Fix Architecture
|
|
435
|
+
|
|
436
|
+
### 1. Shared Runtime Capability Snapshot
|
|
437
|
+
|
|
438
|
+
Create a shared capability module, likely under one of:
|
|
439
|
+
|
|
440
|
+
- `packages/execution/src/model-capabilities.ts`
|
|
441
|
+
- `packages/cli/src/tui/model-capabilities.ts`
|
|
442
|
+
|
|
443
|
+
Preferred direction:
|
|
444
|
+
|
|
445
|
+
- put pure registry/scoring logic in `@omnius/execution` if it does not need
|
|
446
|
+
TUI renderer state;
|
|
447
|
+
- keep CLI-only disk probes and settings loaders in CLI helpers if necessary;
|
|
448
|
+
- expose plain JSON structures, not formatted TUI lines.
|
|
449
|
+
|
|
450
|
+
Proposed shape:
|
|
451
|
+
|
|
452
|
+
```ts
|
|
453
|
+
export type CapabilityModelKind = "voice" | "image" | "sound" | "music";
|
|
454
|
+
|
|
455
|
+
export interface RuntimeModelSelection {
|
|
456
|
+
kind: CapabilityModelKind;
|
|
457
|
+
selectedModel?: string;
|
|
458
|
+
selectedBackend?: string;
|
|
459
|
+
source: "project" | "global" | "default" | "runtime";
|
|
460
|
+
}
|
|
461
|
+
|
|
462
|
+
export interface RuntimeModelFit {
|
|
463
|
+
score?: number;
|
|
464
|
+
label?: string;
|
|
465
|
+
note?: string;
|
|
466
|
+
minVramGB?: number;
|
|
467
|
+
recommendedVramGB?: number;
|
|
468
|
+
}
|
|
469
|
+
|
|
470
|
+
export interface RuntimeModelDiskState {
|
|
471
|
+
downloaded: boolean;
|
|
472
|
+
bytes: number;
|
|
473
|
+
paths: string[];
|
|
474
|
+
}
|
|
475
|
+
|
|
476
|
+
export interface RuntimeModelCapability {
|
|
477
|
+
kind: CapabilityModelKind;
|
|
478
|
+
id: string;
|
|
479
|
+
label: string;
|
|
480
|
+
backend: string;
|
|
481
|
+
category?: string;
|
|
482
|
+
sizeClass?: string;
|
|
483
|
+
quality?: string;
|
|
484
|
+
output?: string;
|
|
485
|
+
bestUse?: string;
|
|
486
|
+
deployment?: string;
|
|
487
|
+
install?: string;
|
|
488
|
+
defaults?: Record<string, unknown>;
|
|
489
|
+
fit?: RuntimeModelFit;
|
|
490
|
+
disk?: RuntimeModelDiskState;
|
|
491
|
+
selected?: boolean;
|
|
492
|
+
notes: string[];
|
|
493
|
+
}
|
|
494
|
+
|
|
495
|
+
export interface RuntimeCapabilitySnapshot {
|
|
496
|
+
generatedAt: string;
|
|
497
|
+
repoRoot: string;
|
|
498
|
+
hardware: {
|
|
499
|
+
totalRamGB: number;
|
|
500
|
+
availableRamGB: number;
|
|
501
|
+
gpuVramGB: number;
|
|
502
|
+
gpuName?: string;
|
|
503
|
+
};
|
|
504
|
+
selections: RuntimeModelSelection[];
|
|
505
|
+
models: RuntimeModelCapability[];
|
|
506
|
+
embedding: RuntimeEmbeddingStatus;
|
|
507
|
+
tools?: RuntimeToolAwareness;
|
|
508
|
+
}
|
|
509
|
+
```
|
|
510
|
+
|
|
511
|
+
### 2. Agent-Facing Capability Tool
|
|
512
|
+
|
|
513
|
+
Add a tool such as `model_capabilities` or `capability_inspect`.
|
|
514
|
+
|
|
515
|
+
Actions:
|
|
516
|
+
|
|
517
|
+
- `summary`
|
|
518
|
+
- current selected voice/image/sound/music plus compact fit.
|
|
519
|
+
- `list`
|
|
520
|
+
- filter by `kind=voice|image|sound|music|embedding|tools`.
|
|
521
|
+
- `inspect`
|
|
522
|
+
- inspect one model/tool/embedding provider in depth.
|
|
523
|
+
- `current`
|
|
524
|
+
- exact current selections for this invocation.
|
|
525
|
+
- `embedding_status`
|
|
526
|
+
- OpenCLIP, text CLIP, speaker embedding, transcript embedding, installed
|
|
527
|
+
status, vector spaces, counts.
|
|
528
|
+
- `tool_context`
|
|
529
|
+
- actual registered tools in this invocation after policy filtering.
|
|
530
|
+
|
|
531
|
+
This tool must be registered in:
|
|
532
|
+
|
|
533
|
+
- TUI `buildTools` / sub-agent tool set.
|
|
534
|
+
- Telegram admin and public contexts, with public output scoped and redacted.
|
|
535
|
+
- GUI/API agent runs.
|
|
536
|
+
|
|
537
|
+
### 3. Compact Capability Context Injection
|
|
538
|
+
|
|
539
|
+
Every invocation should receive a compact, non-overwhelming block:
|
|
540
|
+
|
|
541
|
+
```text
|
|
542
|
+
<runtime-capabilities>
|
|
543
|
+
Selected models:
|
|
544
|
+
- voice: supertonic (backend supertonic)
|
|
545
|
+
- image: stabilityai/sdxl-turbo (diffusers, fit 88/100 excellent)
|
|
546
|
+
- sound: cvssp/audioldm-s-full-v2 (diffusers, fit 74/100 comfortable)
|
|
547
|
+
- music: facebook/musicgen-small (transformers, fit 70/100 comfortable)
|
|
548
|
+
|
|
549
|
+
Use model_capabilities(action="inspect", kind="image", id="...") for details.
|
|
550
|
+
Use model_capabilities(action="list", kind="music") before switching models.
|
|
551
|
+
</runtime-capabilities>
|
|
552
|
+
```
|
|
553
|
+
|
|
554
|
+
Injection anchors:
|
|
555
|
+
|
|
556
|
+
- TUI: `packages/cli/src/tui/interactive.ts:2165`
|
|
557
|
+
- TUI environment refresh: `packages/cli/src/tui/interactive.ts:2188`
|
|
558
|
+
- Telegram quick chat: `packages/cli/src/tui/telegram-bridge.ts:4149`
|
|
559
|
+
- Telegram action context: `packages/cli/src/tui/telegram-bridge.ts:3297`
|
|
560
|
+
- API/GUI agent subprocess: agent launch path under `packages/cli/src/api/serve.ts`
|
|
561
|
+
|
|
562
|
+
The context should remain compact. Deep detail belongs in the tool.
|
|
563
|
+
|
|
564
|
+
### 4. Runtime Tool Awareness
|
|
565
|
+
|
|
566
|
+
Improve agent self-awareness with the actual invocation tool set.
|
|
567
|
+
|
|
568
|
+
Targets:
|
|
569
|
+
|
|
570
|
+
- keep `tool_search` as the schema promotion mechanism;
|
|
571
|
+
- replace or augment static `ExploreToolsTool` with a runtime-provided tool
|
|
572
|
+
index where possible;
|
|
573
|
+
- expose policy-filtered tool names/categories in the capability snapshot;
|
|
574
|
+
- after compaction, re-inject the compact tool/category summary alongside MCP
|
|
575
|
+
names.
|
|
576
|
+
|
|
577
|
+
Important anchors:
|
|
578
|
+
|
|
579
|
+
- `packages/execution/src/tools/explore-tools.ts:20`
|
|
580
|
+
- `packages/orchestrator/src/agenticRunner.ts:17376`
|
|
581
|
+
- `packages/orchestrator/src/agenticRunner.ts:17630`
|
|
582
|
+
- `packages/orchestrator/src/agenticRunner.ts:15506`
|
|
583
|
+
|
|
584
|
+
### 5. Natural Identity Memory Tool
|
|
585
|
+
|
|
586
|
+
Add an agent-facing tool, likely `identity_memory`.
|
|
587
|
+
|
|
588
|
+
Actions:
|
|
589
|
+
|
|
590
|
+
- `assert_identity`
|
|
591
|
+
- user has explicitly provided identity information.
|
|
592
|
+
- `enroll_face`
|
|
593
|
+
- enroll a face embedding using image evidence.
|
|
594
|
+
- `identify`
|
|
595
|
+
- identify known faces/voices in supplied media.
|
|
596
|
+
- `recall`
|
|
597
|
+
- retrieve identity evidence by name/person/media.
|
|
598
|
+
- `inspect`
|
|
599
|
+
- show identity node, evidence, modalities, confidence, scope.
|
|
600
|
+
|
|
601
|
+
Inputs:
|
|
602
|
+
|
|
603
|
+
```ts
|
|
604
|
+
interface IdentityMemoryArgs {
|
|
605
|
+
action: "assert_identity" | "enroll_face" | "identify" | "recall" | "inspect";
|
|
606
|
+
name?: string;
|
|
607
|
+
relation?: "depicts" | "speaker" | "named_as" | "same_person_candidate";
|
|
608
|
+
media?: "reply" | "latest" | string;
|
|
609
|
+
message_id?: string | number;
|
|
610
|
+
scope?: "current" | "private" | "group" | "global";
|
|
611
|
+
confidence?: number;
|
|
612
|
+
evidence_note?: string;
|
|
613
|
+
}
|
|
614
|
+
```
|
|
615
|
+
|
|
616
|
+
Validation rules:
|
|
617
|
+
|
|
618
|
+
- The model may call this naturally when the user asks to remember a face,
|
|
619
|
+
states a person name, identifies a speaker, or asks who/what is remembered.
|
|
620
|
+
- The tool must require explicit `name` for identity assertion writes.
|
|
621
|
+
- The tool must resolve media aliases through scoped media providers, not local
|
|
622
|
+
filesystem guessing.
|
|
623
|
+
- If there is no usable image/voice evidence, store the textual identity
|
|
624
|
+
assertion but mark it as lacking biometric embedding evidence.
|
|
625
|
+
- If a face is present, write through `MultimodalIdentityService` and mirror
|
|
626
|
+
enrollment into `VisualMemoryTool` or a shared face embedding helper.
|
|
627
|
+
- In public Telegram, writes are scoped to the current chat.
|
|
628
|
+
|
|
629
|
+
### 6. CLIP And Associative Memory Repair
|
|
630
|
+
|
|
631
|
+
Required fixes:
|
|
632
|
+
|
|
633
|
+
- Allow zettelkasten linking when an episode has `clipEmbedding` but no native
|
|
634
|
+
`embedding`.
|
|
635
|
+
- Make `findNeighbors` consider candidates with compatible CLIP embeddings even
|
|
636
|
+
when native embeddings are absent.
|
|
637
|
+
- Avoid labeling every visual episode node as `person` in embedding workers.
|
|
638
|
+
- Replace stale `person_node_id` co-occurrence assumptions with graph relations
|
|
639
|
+
created by `MultimodalIdentityService`.
|
|
640
|
+
- Add status/introspection for embedding providers and vector counts.
|
|
641
|
+
|
|
642
|
+
Important anchors:
|
|
643
|
+
|
|
644
|
+
- `packages/memory/src/zettelkasten.ts:78`
|
|
645
|
+
- `packages/memory/src/zettelkasten.ts:90`
|
|
646
|
+
- `packages/cli/src/api/embedding-workers.ts:148`
|
|
647
|
+
- `packages/cli/src/api/embedding-workers.ts:165`
|
|
648
|
+
- `packages/memory/src/multimodalIdentity.ts:330`
|
|
649
|
+
|
|
650
|
+
## Implementation Checklist
|
|
651
|
+
|
|
652
|
+
### Phase 0 - Tracking Document
|
|
653
|
+
|
|
654
|
+
- [x] Create root-fix document with anchors and artifacts.
|
|
655
|
+
- [x] Update this checklist as each implementation step lands.
|
|
656
|
+
- [ ] Add links from related docs if the implementation spans multiple passes.
|
|
657
|
+
|
|
658
|
+
### Phase 1 - Shared Model Capability Snapshot
|
|
659
|
+
|
|
660
|
+
- [ ] Extract image hardware fit logic from TUI command rendering into a shared
|
|
661
|
+
pure function.
|
|
662
|
+
- [ ] Extract sound/music hardware fit logic from TUI command rendering into a
|
|
663
|
+
shared pure function.
|
|
664
|
+
- [ ] Add shared disk/cache probe helpers or a CLI wrapper around shared model
|
|
665
|
+
metadata.
|
|
666
|
+
- [ ] Add `buildRuntimeCapabilitySnapshot(repoRoot, options)`.
|
|
667
|
+
- [ ] Include current selected voice/image/sound/music settings.
|
|
668
|
+
- [ ] Include defaults when no explicit setting is present.
|
|
669
|
+
- [ ] Include voice model metadata from `listVoiceModels`.
|
|
670
|
+
- [ ] Include image/music/sound model metadata and fit scores.
|
|
671
|
+
- [ ] Include downloaded/cache size where available.
|
|
672
|
+
- [ ] Add unit tests for snapshot shape and selected/default resolution.
|
|
673
|
+
|
|
674
|
+
### Phase 2 - Agent-Facing Capability Tool
|
|
675
|
+
|
|
676
|
+
- [ ] Add `ModelCapabilitiesTool` or `CapabilityInspectTool`.
|
|
677
|
+
- [ ] Support `summary`, `current`, `list`, `inspect`, `embedding_status`,
|
|
678
|
+
and `tool_context`.
|
|
679
|
+
- [ ] Register the tool in TUI main tool set.
|
|
680
|
+
- [ ] Register the tool in TUI sub-agent tool set.
|
|
681
|
+
- [ ] Register the tool in Telegram admin DM tools.
|
|
682
|
+
- [ ] Register a scoped/redacted version in Telegram public tools.
|
|
683
|
+
- [ ] Register the tool for GUI/API agent runs.
|
|
684
|
+
- [ ] Add tests for public redaction and full admin output.
|
|
685
|
+
|
|
686
|
+
### Phase 3 - Capability Context Injection
|
|
687
|
+
|
|
688
|
+
- [ ] Add compact `<runtime-capabilities>` renderer.
|
|
689
|
+
- [ ] Inject into TUI `dynamicContext`.
|
|
690
|
+
- [ ] Inject into TUI `environmentProvider` if selections can change mid-run.
|
|
691
|
+
- [ ] Inject into Telegram quick-chat system context.
|
|
692
|
+
- [ ] Inject into Telegram action sub-agent dynamic context.
|
|
693
|
+
- [ ] Inject into GUI/API agent subprocess context.
|
|
694
|
+
- [ ] Add post-compaction reinjection for compact capability awareness.
|
|
695
|
+
- [ ] Verify context remains compact for small/medium models.
|
|
696
|
+
|
|
697
|
+
### Phase 4 - Existing Tool Schema Cleanup
|
|
698
|
+
|
|
699
|
+
- [ ] Fix `ImageGenerateTool` schema so `action=list_models` does not require
|
|
700
|
+
`prompt`.
|
|
701
|
+
- [ ] Expand `ImageGenerateTool` `list_models` output or delegate to the new
|
|
702
|
+
capability snapshot.
|
|
703
|
+
- [ ] Add `action` to `AudioGenerateTool` schema.
|
|
704
|
+
- [ ] Ensure `AudioGenerateTool action=list_models kind=music|sound` works
|
|
705
|
+
without `prompt`.
|
|
706
|
+
- [ ] Ensure `generate_tts`/`audio_playback list_voices` points to the same
|
|
707
|
+
voice capability metadata.
|
|
708
|
+
- [ ] Update `/help` and README model/tool descriptions after implementation.
|
|
709
|
+
|
|
710
|
+
### Phase 5 - Natural Identity Memory Tool
|
|
711
|
+
|
|
712
|
+
- [x] Add `IdentityMemoryTool`.
|
|
713
|
+
- [x] Integrate with `MultimodalIdentityService`.
|
|
714
|
+
- [x] Resolve scoped media aliases: `reply`, `latest`, and explicit
|
|
715
|
+
Telegram/media paths.
|
|
716
|
+
- [x] Validate explicit name/evidence on identity assertion writes.
|
|
717
|
+
- [x] Store scoped graph evidence for `depicts`, `named_as`,
|
|
718
|
+
`speaker_candidate`, and `same_person_candidate`.
|
|
719
|
+
- [ ] Add voice embedding extraction so explicit speaker assertions can also
|
|
720
|
+
create `voice_sample_of` evidence.
|
|
721
|
+
- [x] Mirror face enrollment into visual memory or a shared face embedding
|
|
722
|
+
backend.
|
|
723
|
+
- [ ] Support recall/inspect by sender, media, voice, and scope. Person-name
|
|
724
|
+
recall/inspect is implemented.
|
|
725
|
+
- [x] Add Telegram public scoped version.
|
|
726
|
+
- [x] Add tests for no-regex behavior: tool only writes when called with
|
|
727
|
+
explicit structured args.
|
|
728
|
+
|
|
729
|
+
Implemented anchors:
|
|
730
|
+
|
|
731
|
+
- `packages/cli/src/tui/identity-memory-tool.ts`
|
|
732
|
+
- `IdentityMemoryTool`
|
|
733
|
+
- `formatIdentityMemoryContext`
|
|
734
|
+
- media resolver interface for Telegram, TUI, GUI, and future voice/session
|
|
735
|
+
surfaces.
|
|
736
|
+
- `packages/cli/src/tui/telegram-bridge.ts`
|
|
737
|
+
- `buildTelegramIdentityMemoryTool`
|
|
738
|
+
- `resolveTelegramIdentityMedia`
|
|
739
|
+
- `formatIdentityMemoryContext` injection into Telegram action context.
|
|
740
|
+
- `packages/cli/src/tui/interactive.ts`
|
|
741
|
+
- TUI main-agent and sub-agent registration of `IdentityMemoryTool`
|
|
742
|
+
- compact identity memory contract injection.
|
|
743
|
+
|
|
744
|
+
### Phase 6 - CLIP/Zettelkasten Repair
|
|
745
|
+
|
|
746
|
+
- [x] Update `findNeighbors` to handle CLIP-only source episodes.
|
|
747
|
+
- [x] Update candidate filtering to include CLIP-compatible candidates.
|
|
748
|
+
- [x] Ensure CLIP image/text embeddings use clearly named vector spaces.
|
|
749
|
+
- [x] Add regression tests for visual episode with only `clipEmbedding`.
|
|
750
|
+
- [ ] Remove or replace stale `person_node_id` coupling in embedding workers.
|
|
751
|
+
- [ ] Stop labeling all visual embedding worker nodes as `person`.
|
|
752
|
+
- [x] Link identity assertions to media/message/person nodes through the graph.
|
|
753
|
+
|
|
754
|
+
### Phase 7 - GUI/API And Documentation
|
|
755
|
+
|
|
756
|
+
- [ ] Add `/v1/capabilities` or `/v1/model-capabilities`.
|
|
757
|
+
- [ ] Add `/v1/capabilities/current`.
|
|
758
|
+
- [ ] Add `/v1/capabilities/models?kind=image|sound|music|voice`.
|
|
759
|
+
- [ ] Add GUI display for selected/current models using the same endpoint.
|
|
760
|
+
- [ ] Add README section for model capability awareness.
|
|
761
|
+
- [ ] Add `/help` entries for the new capability and identity memory tools.
|
|
762
|
+
- [ ] Update existing docs that mention embedding auto-installs to reflect
|
|
763
|
+
`OMNIUS_INSTALL_FULL_EMBED_DEPS=1`.
|
|
764
|
+
|
|
765
|
+
### Phase 8 - Verification
|
|
766
|
+
|
|
767
|
+
- [ ] `pnpm -r build`
|
|
768
|
+
- [x] targeted memory package tests
|
|
769
|
+
- [x] targeted CLI identity tool tests
|
|
770
|
+
- [x] targeted CLI and memory builds
|
|
771
|
+
- [ ] Telegram public scoped smoke test
|
|
772
|
+
- [ ] Telegram admin DM smoke test
|
|
773
|
+
- [ ] TUI smoke test for capability summary
|
|
774
|
+
- [ ] GUI/API smoke test for capability endpoint
|
|
775
|
+
- [ ] Manual test: ask "what image model are you using?"
|
|
776
|
+
- [ ] Manual test: ask "list available music models and pick the best fit"
|
|
777
|
+
- [ ] Manual test: reply to an image with "this is <name>" and confirm graph,
|
|
778
|
+
visual memory, and recall evidence
|
|
779
|
+
|
|
780
|
+
## Expected End State
|
|
781
|
+
|
|
782
|
+
The agent should be able to answer, without guessing:
|
|
783
|
+
|
|
784
|
+
- "What voice model are you using right now?"
|
|
785
|
+
- "Which image models can you use, and which fits this GPU best?"
|
|
786
|
+
- "List music models by quality and size."
|
|
787
|
+
- "Can you inspect your sound model setup?"
|
|
788
|
+
- "Do you remember this face?"
|
|
789
|
+
- "This person is Amy. Remember that."
|
|
790
|
+
- "Who is in this image based on prior memory?"
|
|
791
|
+
|
|
792
|
+
For those questions, the model should either:
|
|
793
|
+
|
|
794
|
+
- answer from compact runtime capability context;
|
|
795
|
+
- call the capability inspection tool for details;
|
|
796
|
+
- call the identity memory tool with structured evidence;
|
|
797
|
+
- ask a clarifying question if the identity evidence is ambiguous.
|
|
798
|
+
|
|
799
|
+
No part of this should depend on regex phrase triggers.
|