omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,38 @@
|
|
|
1
|
+
# Hermes Architecture Deltas Workorder Index
|
|
2
|
+
|
|
3
|
+
Status: ready for implementation planning. These documents are workorders only; they do not claim the features are implemented.
|
|
4
|
+
|
|
5
|
+
Source baseline:
|
|
6
|
+
- Omnius repository: `/home/robit/Documents/repositories/open-agents-1`
|
|
7
|
+
- Hermes comparison checkout: `/home/robit/Documents/repositories/hermes-agent`
|
|
8
|
+
- Hermes commit inspected: `6f6eb871d83415fe2980f3483cc41a435ba22196`
|
|
9
|
+
- Hermes version inspected: `0.15.1`
|
|
10
|
+
- AIWG framing used: `sdlc/architecture-evolution`, `sdlc/flow-handoff-checklist`, `sdlc/flow-requirements-evolution`
|
|
11
|
+
|
|
12
|
+
## Feature Workorders
|
|
13
|
+
|
|
14
|
+
| Folder | Workorder | Purpose | Current State |
|
|
15
|
+
| --- | --- | --- | --- |
|
|
16
|
+
| `01-public-scrutiny-provenance-control/` | [WORKORDER.md](01-public-scrutiny-provenance-control/WORKORDER.md) | Runtime public Telegram scrutiny modes with human and agent-controlled provenance escalation. | Planned |
|
|
17
|
+
| `02-context-engine-plugin-boundary/` | [WORKORDER.md](02-context-engine-plugin-boundary/WORKORDER.md) | Extract context assembly, compaction, and evidence framing behind a pluggable context engine boundary. | Planned |
|
|
18
|
+
| `03-typed-gateway-event-stream/` | [WORKORDER.md](03-typed-gateway-event-stream/WORKORDER.md) | Separate runner facts from Telegram/TUI rendering through typed events and adapter dispatch. | Planned |
|
|
19
|
+
| `04-task-local-gateway-context/` | [WORKORDER.md](04-task-local-gateway-context/WORKORDER.md) | Add task-local gateway/session context using `AsyncLocalStorage` to prevent cross-chat and cross-run bleed. | Planned |
|
|
20
|
+
| `05-process-lifecycle-monitoring-notifications/` | [WORKORDER.md](05-process-lifecycle-monitoring-notifications/WORKORDER.md) | Keep Omnius's lease registry as authority and add watcher, adoption, notification, and legacy kill-path cleanup. | Planned |
|
|
21
|
+
| `06-vision-evidence-routing-ladder/` | [WORKORDER.md](06-vision-evidence-routing-ladder/WORKORDER.md) | Route visual evidence from cheap observation to OCR to auxiliary vision based on request-derived information need. | Planned |
|
|
22
|
+
| `07-durable-multi-agent-kanban/` | [WORKORDER.md](07-durable-multi-agent-kanban/WORKORDER.md) | Add durable, inspectable work boards for long admin tasks and disputed public factual work. | Planned |
|
|
23
|
+
| `08-completion-critic-reconciliation-ledger/` | [WORKORDER.md](08-completion-critic-reconciliation-ledger/WORKORDER.md) | Turn completion review into a durable claim/evidence ledger with critic reconciliation packets. | Planned |
|
|
24
|
+
|
|
25
|
+
## Carryover Rules
|
|
26
|
+
|
|
27
|
+
- Do not encode scenario-specific phrase filters as completion or scrutiny logic.
|
|
28
|
+
- Any escalation or completion gate must be driven by the current request, proposed claims, observed tool results, and explicit human or agent-declared state.
|
|
29
|
+
- Keep public Telegram bounded by default; strict review modes must be intentional, reversible, and time-limited.
|
|
30
|
+
- Never let a renderer label backend `incomplete_verification` as completed.
|
|
31
|
+
- Never authorize process termination from command-name matching alone; lease ownership and identity verification remain mandatory.
|
|
32
|
+
|
|
33
|
+
## Global Completion Metrics
|
|
34
|
+
|
|
35
|
+
- Every feature folder has an `INDEX.md` and `WORKORDER.md`.
|
|
36
|
+
- Every workorder names code anchors in Omnius and Hermes.
|
|
37
|
+
- Every workorder includes phases, file-by-file implementation notes, tests, rollout controls, and measurable completion criteria.
|
|
38
|
+
- A future agent can update progress checkboxes without rereading this chat.
|
|
@@ -0,0 +1,281 @@
|
|
|
1
|
+
# Omnius Context Engineering Behavior Fixes
|
|
2
|
+
|
|
3
|
+
Date: 2026-06-29
|
|
4
|
+
|
|
5
|
+
Source observation: current Omnius behavior while supervising a noclip/earth sidebar cleanup task. The observed run mixed internal cognitive-agent artifacts with user-task artifacts, repeated stale edit attempts after the file had changed, and overclaimed completion despite stale or missing verification.
|
|
6
|
+
|
|
7
|
+
## Work Order
|
|
8
|
+
|
|
9
|
+
### 1. Isolate Internal Runners From User-Task Artifacts
|
|
10
|
+
|
|
11
|
+
Status: implemented in this work session.
|
|
12
|
+
|
|
13
|
+
Problem: DMN, emotion, and SNR evaluator runners instantiate `AgenticRunner` as normal task runners. Their internal prompts can write handoffs, reflections, completion ledgers, provenance, and task memory that later pollute user-task context.
|
|
14
|
+
|
|
15
|
+
Implementation:
|
|
16
|
+
- Add `artifactMode?: "user-task" | "internal"` to `AgenticRunnerOptions`.
|
|
17
|
+
- Treat `subAgent: true` as isolated for backward compatibility, but use `artifactMode: "internal"` for cognitive helper runners.
|
|
18
|
+
- Disable automatic cross-task handoff read/write in internal mode.
|
|
19
|
+
- Disable task reflection read/write in internal mode.
|
|
20
|
+
- Avoid creating completion contracts/ledgers for internal mode.
|
|
21
|
+
- Avoid writing consolidation/provenance/task-summary artifacts for internal mode.
|
|
22
|
+
- Pass `artifactMode: "internal"` from DMN, emotion, and SNR internal runners.
|
|
23
|
+
|
|
24
|
+
Acceptance:
|
|
25
|
+
- Internal runner prompts cannot become `.omnius/handoffs/latest.json`.
|
|
26
|
+
- Internal zero-turn failures cannot become task reflections.
|
|
27
|
+
- User-task runners continue writing normal artifacts.
|
|
28
|
+
|
|
29
|
+
### 2. Replace Loop Escape Equals `task_complete`
|
|
30
|
+
|
|
31
|
+
Status: implemented in this work session.
|
|
32
|
+
|
|
33
|
+
Problem: the loop intervention still injects a mandatory `task_complete` directive after detecting repetition. This encourages partial or false completion.
|
|
34
|
+
|
|
35
|
+
Implementation:
|
|
36
|
+
- Remove the mandatory `task_complete` wording from loop intervention.
|
|
37
|
+
- Instruct the model to stop repeating the implicated tool family.
|
|
38
|
+
- Prefer verification when files changed, reread the exact target once for stale edits, ask the user if genuinely blocked, or report incomplete/blocking status.
|
|
39
|
+
|
|
40
|
+
Acceptance:
|
|
41
|
+
- Loop intervention text never says to call `task_complete` with whatever is available.
|
|
42
|
+
- Existing `ask_user` circuit-breaker remains the high-tier escape path.
|
|
43
|
+
|
|
44
|
+
### 3. Hard-Block Stale Edit Loops
|
|
45
|
+
|
|
46
|
+
Status: implemented in this work session.
|
|
47
|
+
|
|
48
|
+
Problem: exact duplicate detection does not catch stale edit families. The observed run repeatedly retried an old `file_edit` even after the target text was already absent.
|
|
49
|
+
|
|
50
|
+
Implementation:
|
|
51
|
+
- Add a stale edit fingerprint for `file_edit`, `file_patch`, and `batch_edit`.
|
|
52
|
+
- Track repeated failed edit families by path, normalized target text, and error class.
|
|
53
|
+
- After the threshold, block the edit before dispatch and provide a concrete recovery instruction.
|
|
54
|
+
- When the tool output indicates the old text is absent or no replacement occurred, record the failure family.
|
|
55
|
+
|
|
56
|
+
Acceptance:
|
|
57
|
+
- A repeated `old_string not found` edit is blocked after repeated failures.
|
|
58
|
+
- The block tells the model to read the current target once, verify if already satisfied, or choose a different patch.
|
|
59
|
+
|
|
60
|
+
### 4. Make Completion State Truth-Based, Not Summary-Based
|
|
61
|
+
|
|
62
|
+
Status: implemented in this work session.
|
|
63
|
+
|
|
64
|
+
Problem: the completion ledger records tool evidence but only creates claims from model-proposed completion text. Timeout/open runs can end with evidence but no truth-state synthesis.
|
|
65
|
+
|
|
66
|
+
Implementation:
|
|
67
|
+
- Track mutation and verification evidence from tool results.
|
|
68
|
+
- Add finalization helpers that derive unresolved items when mutations occur after the last successful verification.
|
|
69
|
+
- Mark repeated stale edit blocks as unresolved evidence.
|
|
70
|
+
- Finalize open ledgers as `incomplete_verification` when unresolved evidence exists.
|
|
71
|
+
|
|
72
|
+
Acceptance:
|
|
73
|
+
- A run with file mutations after its last verification cannot look cleanly complete.
|
|
74
|
+
- A timeout/open run can still explain unresolved evidence without a `task_complete` summary.
|
|
75
|
+
|
|
76
|
+
### 5. Stop Todo Chunker Overclaiming Completion
|
|
77
|
+
|
|
78
|
+
Status: implemented in this work session.
|
|
79
|
+
|
|
80
|
+
Problem: todo context chunks are marked `completed` even when summaries contain unresolved work, failed edit loops, or stale verification.
|
|
81
|
+
|
|
82
|
+
Implementation:
|
|
83
|
+
- Add `partial` and `unverified` chunk statuses.
|
|
84
|
+
- Compute chunk status from unresolved and verification text instead of always setting `completed`.
|
|
85
|
+
- Make heuristic fallback conservative: never invent completion when evidence is ambiguous.
|
|
86
|
+
|
|
87
|
+
Acceptance:
|
|
88
|
+
- A chunk containing unresolved failures is not marked `completed`.
|
|
89
|
+
- A chunk lacking verification is `unverified` unless explicit verification evidence exists.
|
|
90
|
+
|
|
91
|
+
### 6. Add Defensive Handoff and Reflection Quality Gates
|
|
92
|
+
|
|
93
|
+
Status: implemented in this work session.
|
|
94
|
+
|
|
95
|
+
Problem: even with internal isolation, stale/low-quality handoffs and reflections should be rejected at the storage boundary.
|
|
96
|
+
|
|
97
|
+
Implementation:
|
|
98
|
+
- Add handoff metadata for artifact mode and quality.
|
|
99
|
+
- Reject handoffs for internal prompt families, zero-turn/no-tool runs, empty summaries with no files/tools, and known SNR/DMN/emotion evaluator prompts.
|
|
100
|
+
- Reject task reflections for the same internal prompt families and zero-progress helper failures.
|
|
101
|
+
- Fix duplicate `failedPaths` assignment in reflection serialization.
|
|
102
|
+
|
|
103
|
+
Acceptance:
|
|
104
|
+
- Internal prompt families cannot be persisted as user-task handoff/reflection artifacts.
|
|
105
|
+
- Empty zero-tool helper failures are quarantined instead of injected into later user tasks.
|
|
106
|
+
|
|
107
|
+
## Verification Plan
|
|
108
|
+
|
|
109
|
+
- Add focused unit coverage for artifact isolation, stale edit blocking, loop intervention wording, handoff/reflection quality gates, completion ledger finalization, and todo chunk status.
|
|
110
|
+
- Run the affected test files first:
|
|
111
|
+
- `packages/orchestrator/tests/critic.test.ts`
|
|
112
|
+
- `packages/orchestrator/tests/completionLedger.test.ts`
|
|
113
|
+
- `packages/orchestrator/tests/reflection.test.ts`
|
|
114
|
+
- `packages/cli/tests/task-handoff.test.ts`
|
|
115
|
+
- `packages/cli/tests/dmn-engine.test.ts`
|
|
116
|
+
- `packages/cli/tests/emotion-engine.test.ts`
|
|
117
|
+
- Run the relevant package test command after focused tests pass.
|
|
118
|
+
|
|
119
|
+
## Live Multimodal Regression Work Order
|
|
120
|
+
|
|
121
|
+
Date: 2026-06-30
|
|
122
|
+
|
|
123
|
+
Source observation: aggressive live testing surfaced visual-memory CLIP wrapper failures, active todo overclaiming after failed tool calls, JSON/headless lifecycle drift, and noisy global context recall. Follow-up correction: context selection must be semantic/vector/inference based, not hard-coded English keyword filters.
|
|
124
|
+
|
|
125
|
+
### 7. Fix CLIP/SigLIP Feature Output Unwrapping
|
|
126
|
+
|
|
127
|
+
Status: implemented.
|
|
128
|
+
|
|
129
|
+
Problem: `visual_memory teach` and related multimodal CLIP paths failed when transformers returned `BaseModelOutputWithPooling` instead of a tensor, causing `.norm(...)` to crash.
|
|
130
|
+
|
|
131
|
+
Implementation:
|
|
132
|
+
- Add a shared Python helper that unwraps `image_embeds`, `text_embeds`, `pooler_output`, `last_hidden_state[:, 0]`, tuple/list outputs, or plain tensors before normalization.
|
|
133
|
+
- Inject the helper into `visual_memory` object teach/recognize paths.
|
|
134
|
+
- Inject the helper into multimodal episode image/text CLIP embedding paths.
|
|
135
|
+
- Add source-level unit coverage for wrapper support.
|
|
136
|
+
|
|
137
|
+
### 8. Make Vision Failures Actionable First
|
|
138
|
+
|
|
139
|
+
Status: implemented.
|
|
140
|
+
|
|
141
|
+
Problem: warnings and dependency noise could hide the actual traceback root cause from the model.
|
|
142
|
+
|
|
143
|
+
Implementation:
|
|
144
|
+
- Export and harden `summarizeProcessFailure`.
|
|
145
|
+
- Put the final exception line first as `Root cause: ...`.
|
|
146
|
+
- Preserve traceback context and tail output below it.
|
|
147
|
+
|
|
148
|
+
### 9. Truth-Reconcile Active Todos From Tool Evidence
|
|
149
|
+
|
|
150
|
+
Status: implemented.
|
|
151
|
+
|
|
152
|
+
Problem: the live agent marked “Teach each object to visual_memory” completed even though every matching `visual_memory teach` call failed.
|
|
153
|
+
|
|
154
|
+
Implementation:
|
|
155
|
+
- Add `todoTruth` reconciliation.
|
|
156
|
+
- Downgrade completed todos to `blocked` when matching tool-family evidence failed and no later same-family success exists.
|
|
157
|
+
- Run reconciliation immediately after successful `todo_write`, before verification nudges and todo chunking.
|
|
158
|
+
- Add a system message instructing the model to reconfigure the affected subtask instead of overclaiming.
|
|
159
|
+
|
|
160
|
+
### 10. First-Class Nested Todo Decomposition
|
|
161
|
+
|
|
162
|
+
Status: implemented.
|
|
163
|
+
|
|
164
|
+
Problem: `parentId` existed in storage but was not an explicit task decomposition contract and parents could overclaim child completion.
|
|
165
|
+
|
|
166
|
+
Implementation:
|
|
167
|
+
- Expand `todo_write` instructions for stable ids, `parentId`, leaf-subtask work, and evidence-driven tree rewrites.
|
|
168
|
+
- Add parent truth reconciliation: a parent cannot be completed while any child is pending, in progress, or blocked.
|
|
169
|
+
- Render nested todos in the TUI with indentation.
|
|
170
|
+
- Add tests for parent/child persistence and parent downgrade behavior.
|
|
171
|
+
|
|
172
|
+
### 11. Semantic Context Selection, Not English Keyword Filters
|
|
173
|
+
|
|
174
|
+
Status: implemented.
|
|
175
|
+
|
|
176
|
+
Problem: persisted failure modes and preflight memories must not be admitted by hard-coded English regex/stopword heuristics.
|
|
177
|
+
|
|
178
|
+
Implementation:
|
|
179
|
+
- Remove semantic keyword filtering from `failureHandoff`.
|
|
180
|
+
- Add runner-side task-to-failure-pattern selection using embedding batch cosine similarity.
|
|
181
|
+
- Add inference fallback for failure-pattern relevance when embeddings are unavailable.
|
|
182
|
+
- If neither vector nor inference scoring is available, inject no persisted failure patterns rather than guessing.
|
|
183
|
+
- Convert preflight memory recall to async embedding retrieval and cosine scoring over stored episode embeddings.
|
|
184
|
+
- Add source guards to prevent reintroducing the keyword matcher/internal-recall filter.
|
|
185
|
+
|
|
186
|
+
### 12. Explicit Degraded Completion Contract
|
|
187
|
+
|
|
188
|
+
Status: implemented.
|
|
189
|
+
|
|
190
|
+
Problem: the resolution gate blocked completion even when the original user task explicitly allowed documenting a tool failure as the desired fallback outcome.
|
|
191
|
+
|
|
192
|
+
Implementation:
|
|
193
|
+
- Add an auxiliary inference verdict for degraded-completion allowance.
|
|
194
|
+
- Accept degraded completion only when the inference verdict says the original request permits it, the summary discloses the failure, and fallback evidence exists.
|
|
195
|
+
- Update the verifier prompt to understand explicitly permitted degraded fallback.
|
|
196
|
+
|
|
197
|
+
### 13. Headless JSON Lifecycle
|
|
198
|
+
|
|
199
|
+
Status: implemented.
|
|
200
|
+
|
|
201
|
+
Problem: `omnius --json` could emit a final JSON object but keep running post-task trajectory/memory side effects and could report completed without the actual runner result.
|
|
202
|
+
|
|
203
|
+
Implementation:
|
|
204
|
+
- Add `onRunResult` callback to `runWithTUI`.
|
|
205
|
+
- JSON mode uses runner `completed/status/summary/turns/toolCalls/filesEdited/testsRun`.
|
|
206
|
+
- Headless mode skips non-critical post-run identity/archive/trajectory/cohere side effects.
|
|
207
|
+
- Headless errors throw back to JSON mode instead of directly exiting inside `runWithTUI`.
|
|
208
|
+
- Add controlled force-exit for JSON CLI with test opt-out.
|
|
209
|
+
|
|
210
|
+
### 14. Model Resolution Telemetry
|
|
211
|
+
|
|
212
|
+
Status: implemented.
|
|
213
|
+
|
|
214
|
+
Problem: live testing showed ambiguity between requested visible model and models used by auxiliary/background calls.
|
|
215
|
+
|
|
216
|
+
Implementation:
|
|
217
|
+
- Emit status telemetry for main runner model resolution.
|
|
218
|
+
- Emit status telemetry for completion-resolution and failure-pattern relevance inference calls.
|
|
219
|
+
|
|
220
|
+
### 15. Crossmodal Self-Test Isolation
|
|
221
|
+
|
|
222
|
+
Status: implemented.
|
|
223
|
+
|
|
224
|
+
Problem: crossmodal self-test used the live repo `.omnius/memory.db`, printed false retrieval results, and still exited successfully.
|
|
225
|
+
|
|
226
|
+
Implementation:
|
|
227
|
+
- Use temporary isolated memory DBs.
|
|
228
|
+
- Fail when image/text embeddings are unavailable or retrieval top result is false.
|
|
229
|
+
- Close stores and remove temp directories.
|
|
230
|
+
|
|
231
|
+
### 16. Truth-Based Terminal Completion After Verified Todo Closure
|
|
232
|
+
|
|
233
|
+
Status: implemented in source; pending publish/live reinstall verification.
|
|
234
|
+
|
|
235
|
+
Problem: live headless testing with `robit/ornith:35b` completed the nested todo tree, edited the requested files, ran the declared verification command successfully, and then stalled after the REG-31 "call task_complete" prompt. The runner had enough structured truth evidence but still depended on another model turn for the final terminal action.
|
|
236
|
+
|
|
237
|
+
Implementation:
|
|
238
|
+
- Add a pure truth-based completion decision module.
|
|
239
|
+
- Require all todos completed, zero unresolved verification failures, and a fresh successful declared `verifyCommand`.
|
|
240
|
+
- Prefer structural matching against `todo.verifyCommand` for validation success.
|
|
241
|
+
- Move generic command-name completion detection behind `OMNIUS_ENABLE_GENERIC_COMPLETION_COMMAND_HEURISTIC=1`.
|
|
242
|
+
- Synthesize terminal completion before another model request, while still running existing completion/provenance/backward-pass gates.
|
|
243
|
+
|
|
244
|
+
Acceptance:
|
|
245
|
+
- A fully verified todo tree can finish headless JSON without waiting on a final model-generated `task_complete`.
|
|
246
|
+
- Partial work, open children, stale validation, or unresolved verification failures remain incomplete.
|
|
247
|
+
- Completion automation is anchored to declared verification evidence, not English command-name filters.
|
|
248
|
+
|
|
249
|
+
### 17. Keep Verify-Command Evidence Separate From Artifact Evidence
|
|
250
|
+
|
|
251
|
+
Status: implemented in source; pending publish/live reinstall verification.
|
|
252
|
+
|
|
253
|
+
Problem: a live local run marked a final todo completed with a combined `verifyCommand` it had not run exactly. REG-37 recorded the missing verifier, but REG-38 artifact inspection then cleared the same aggregate failure because the files existed. The model then called `task_complete` and overclaimed the combined verifier as proven.
|
|
254
|
+
|
|
255
|
+
Implementation:
|
|
256
|
+
- Add structural verification-command matching.
|
|
257
|
+
- Accept exact verifier commands and reliable `verify && echo ...` suffixes.
|
|
258
|
+
- Reject `verify && echo ok || echo failed` wrappers because they can mask failure with exit 0.
|
|
259
|
+
- Allow a conjunctive verifier to be satisfied by separately successful reliable component commands.
|
|
260
|
+
- Track verify-command failures and artifact-inspection failures separately, then combine them into the completion gate.
|
|
261
|
+
- Block `task_complete` while any todo verification failure remains unresolved.
|
|
262
|
+
- Replace todo-write verification nudge content regex with structured `verifyCommand` / `declaredArtifacts` evidence checks.
|
|
263
|
+
|
|
264
|
+
Acceptance:
|
|
265
|
+
- Artifact existence cannot clear a missing `verifyCommand`.
|
|
266
|
+
- A `task_complete` call is held when the todo tree still has unresolved verification evidence.
|
|
267
|
+
- Verification nudges are structural, not English content heuristics.
|
|
268
|
+
|
|
269
|
+
### 18. Preserve Phase Archives in ESM Builds
|
|
270
|
+
|
|
271
|
+
Status: implemented in source; pending publish/live reinstall verification.
|
|
272
|
+
|
|
273
|
+
Problem: live local JSON runs emitted `Phase archive failed (non-fatal): require is not defined`. Phase archives are context-continuity artifacts for long tasks; losing them weakens later phase recall and forensic review.
|
|
274
|
+
|
|
275
|
+
Implementation:
|
|
276
|
+
- Replace runtime `require("node:fs")` and `require("node:path")` uses in `agenticRunner.ts` with existing ESM-safe top-level imports.
|
|
277
|
+
- Cover phase archives, KG summary writes/pruning, and checkpoint persistence in the same pass.
|
|
278
|
+
|
|
279
|
+
Acceptance:
|
|
280
|
+
- ESM-built CLI runs can write `.omnius/phases/` archives without `require` failures.
|
|
281
|
+
- Long-task context contraction keeps disk-backed phase recovery available.
|
|
@@ -0,0 +1,202 @@
|
|
|
1
|
+
# Telegram Dropbear Context Engineering RCA Work Order
|
|
2
|
+
|
|
3
|
+
Date: 2026-07-01
|
|
4
|
+
|
|
5
|
+
Observed run: `/home/roko/Documents/Projects/Adjacent/telegram_test/.omnius`, run id `1782873796963-i5r7mv`.
|
|
6
|
+
|
|
7
|
+
Task under observation: get a meaningful gear-sonic training run with observed progress on the Dropbear MJCF/URDF after addressing closed-loop knee components.
|
|
8
|
+
|
|
9
|
+
## Evidence Snapshot
|
|
10
|
+
|
|
11
|
+
- The run made one successful mutation: `dropbear_mjcf/dropbear_mjcf.xml` changed `RL_Revolute28` to `<joint name="RL_Revolute28" type="weld" />`.
|
|
12
|
+
- Immediately afterward, the model retried the pre-mutation `old_string` and received repeated `old_string not found` failures.
|
|
13
|
+
- The focus supervisor then forced `required_next_action=update_todos`, which blocked useful recovery reads, patches, and verification attempts.
|
|
14
|
+
- The run kept emitting blocked edits, sed/python rewrites, a partial `task_complete`, and more blocked reads through at least turn 28.
|
|
15
|
+
- Main context grew from about 73k estimated tokens after the successful edit to about 92.7k estimated tokens by turn 28.
|
|
16
|
+
- System prompt content grew from about 142k chars to about 230k chars.
|
|
17
|
+
- The final main dump had dozens of `[RUN EVIDENCE]` and focus-supervisor system messages, including repeated synthetic failures.
|
|
18
|
+
- The workboard existed but had no cards, decisions, or diagnostics.
|
|
19
|
+
- The completion ledger remained `open`; no meaningful gear-sonic training verification was run.
|
|
20
|
+
|
|
21
|
+
## Work Orders
|
|
22
|
+
|
|
23
|
+
### WO-TD-01: Stale Edit Recovery Must Read Current Truth, Not Replan Forever
|
|
24
|
+
|
|
25
|
+
Status: implemented
|
|
26
|
+
|
|
27
|
+
Root cause: repeated stale edit failures currently force `update_todos` for non-shell tools. For an `old_string not found` family, this is the wrong recovery state.
|
|
28
|
+
|
|
29
|
+
Anchors:
|
|
30
|
+
- `packages/orchestrator/src/focusSupervisor.ts`
|
|
31
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
32
|
+
- `packages/orchestrator/tests/focusSupervisor.test.ts`
|
|
33
|
+
|
|
34
|
+
Implementation:
|
|
35
|
+
- Classify stale edit failure samples in the focus supervisor.
|
|
36
|
+
- Repeated stale `file_edit`, `file_patch`, or `batch_edit` failures must set `requiredNextAction=read_authoritative_target`.
|
|
37
|
+
- After the authoritative read succeeds, clear that directive and permit a fresh target-specific edit or verification.
|
|
38
|
+
- Do not use `update_todos` as the default escape hatch for stale edit families.
|
|
39
|
+
|
|
40
|
+
Acceptance:
|
|
41
|
+
- Two `old_string not found` failures produce a `read_authoritative_target` directive.
|
|
42
|
+
- A `file_read` of the target satisfies the directive.
|
|
43
|
+
- A different edit target in the same file is not blocked just because an old edit target failed.
|
|
44
|
+
|
|
45
|
+
Verification:
|
|
46
|
+
- Covered by `tests/focusSupervisor.test.ts`.
|
|
47
|
+
|
|
48
|
+
### WO-TD-02: Directive Escalation Must Terminate Cleanly After Repeated Ignoring
|
|
49
|
+
|
|
50
|
+
Status: implemented
|
|
51
|
+
|
|
52
|
+
Root cause: the observed run ignored 27 focus directives and kept accumulating blocked tool calls. The supervisor continued asking for the same action without forcing a terminal incomplete/blocker path.
|
|
53
|
+
|
|
54
|
+
Anchors:
|
|
55
|
+
- `packages/orchestrator/src/focusSupervisor.ts`
|
|
56
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
57
|
+
- `packages/orchestrator/tests/focusSupervisor.test.ts`
|
|
58
|
+
|
|
59
|
+
Implementation:
|
|
60
|
+
- Add an ignored-directive escalation threshold.
|
|
61
|
+
- Once the threshold is crossed with no satisfying mutation/read/verification, force `requiredNextAction=report_incomplete`.
|
|
62
|
+
- In terminal incomplete state, permit `task_complete` only as an incomplete/blocker report.
|
|
63
|
+
- Stop appending broad blocked-family lists indefinitely.
|
|
64
|
+
|
|
65
|
+
Acceptance:
|
|
66
|
+
- Repeatedly ignoring a recovery directive transitions to `terminal_incomplete`.
|
|
67
|
+
- After terminal incomplete, edit/read/shell variants are blocked.
|
|
68
|
+
- A reporting tool is allowed so the run can end truthfully.
|
|
69
|
+
|
|
70
|
+
Verification:
|
|
71
|
+
- Covered by `tests/focusSupervisor.test.ts`.
|
|
72
|
+
|
|
73
|
+
### WO-TD-03: Focus Action Families Must Be Target-Specific For Edits
|
|
74
|
+
|
|
75
|
+
Status: implemented
|
|
76
|
+
|
|
77
|
+
Root cause: focus supervisor action families collapse all edits to `tool:path`. In the observed run, a stale `RL_Revolute28` edit contaminated later distinct joint edits and line patches in the same XML.
|
|
78
|
+
|
|
79
|
+
Anchors:
|
|
80
|
+
- `packages/orchestrator/src/focusSupervisor.ts`
|
|
81
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
82
|
+
- `packages/orchestrator/tests/focusSupervisor.test.ts`
|
|
83
|
+
|
|
84
|
+
Implementation:
|
|
85
|
+
- Build edit action families from tool name, path, normalized edit target, patch mode/offset/limit, and error class where available.
|
|
86
|
+
- Keep shell action families command-based, but trim and hash long commands rather than injecting huge command prefixes.
|
|
87
|
+
- Preserve a separate coarse path lock only for full-file overwrite repair blocks.
|
|
88
|
+
|
|
89
|
+
Acceptance:
|
|
90
|
+
- A repeated stale edit for `RL_Revolute28` blocks only that edit family.
|
|
91
|
+
- A fresh `file_patch` for a different offset/target in the same XML can pass after the required authoritative read.
|
|
92
|
+
- The directive frame no longer grows with long sed/python command bodies.
|
|
93
|
+
|
|
94
|
+
Verification:
|
|
95
|
+
- Covered by `tests/focusSupervisor.test.ts`.
|
|
96
|
+
|
|
97
|
+
### WO-TD-04: Context Evidence Must Fold Repeats And Demote Synthetic Failures
|
|
98
|
+
|
|
99
|
+
Status: implemented
|
|
100
|
+
|
|
101
|
+
Root cause: the context engine injects every recorded tool event as high-authority `[RUN EVIDENCE]`, including repeated focus blocks and stale discovery reads.
|
|
102
|
+
|
|
103
|
+
Anchors:
|
|
104
|
+
- `packages/orchestrator/src/contextEngine.ts`
|
|
105
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
106
|
+
- `packages/orchestrator/tests/contextEngine.test.ts`
|
|
107
|
+
|
|
108
|
+
Implementation:
|
|
109
|
+
- Fold tool evidence by family before rendering system evidence.
|
|
110
|
+
- Keep successes, latest mutations, latest verification results, and first/last unresolved failures.
|
|
111
|
+
- Collapse repeated focus-supervisor synthetic failures into one counted summary.
|
|
112
|
+
- Cap evidence injection separately from raw conversation/tool results.
|
|
113
|
+
|
|
114
|
+
Acceptance:
|
|
115
|
+
- Repeated identical focus blocks render as one summary with a repeat count.
|
|
116
|
+
- Evidence diagnostics report retained and dropped/folded counts.
|
|
117
|
+
- Large early discovery listings do not stay as repeated high-authority system evidence once more specific evidence exists.
|
|
118
|
+
|
|
119
|
+
Verification:
|
|
120
|
+
- Covered by `tests/contextEngine.test.ts`.
|
|
121
|
+
|
|
122
|
+
### WO-TD-05: Context Engine Compaction Must Affect Live Requests
|
|
123
|
+
|
|
124
|
+
Status: implemented
|
|
125
|
+
|
|
126
|
+
Root cause: `AgenticRunner` calls `contextEngine.compact()` but discards its output, while `contextEngine.build()` still prepends extra evidence to the live request.
|
|
127
|
+
|
|
128
|
+
Anchors:
|
|
129
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
130
|
+
- `packages/orchestrator/src/contextEngine.ts`
|
|
131
|
+
- `packages/orchestrator/tests/agenticRunner-context-behavior.test.ts`
|
|
132
|
+
- `packages/orchestrator/tests/contextEngine.test.ts`
|
|
133
|
+
|
|
134
|
+
Implementation:
|
|
135
|
+
- Apply context-engine compaction output to the live message list when it actually compacts.
|
|
136
|
+
- Avoid duplicating engine-added evidence already present in the active context.
|
|
137
|
+
- Carry latest SNR/evidence metrics into focus-supervisor proposed-call snapshots.
|
|
138
|
+
|
|
139
|
+
Acceptance:
|
|
140
|
+
- A compacted context-engine output reduces the live message list before backend calls.
|
|
141
|
+
- `_buildFocusContextSnapshot()` includes raw discovery chars, active evidence chars, active frame chars, and signal-to-noise ratio.
|
|
142
|
+
- Context dumps and focus decisions see the same pressure/noise picture.
|
|
143
|
+
|
|
144
|
+
Verification:
|
|
145
|
+
- Covered by `tests/contextEngine.test.ts`, `tests/agenticRunner-context-behavior.test.ts`, and package build.
|
|
146
|
+
|
|
147
|
+
### WO-TD-06: Completion Ledger Must Finalize Open Failure Runs Truthfully
|
|
148
|
+
|
|
149
|
+
Status: implemented
|
|
150
|
+
|
|
151
|
+
Root cause: the ledger contained evidence of mutation, stale edit failures, and no final verification, but stayed `open`.
|
|
152
|
+
|
|
153
|
+
Anchors:
|
|
154
|
+
- `packages/orchestrator/src/completionLedger.ts`
|
|
155
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
156
|
+
- `packages/orchestrator/tests/completionLedger.test.ts`
|
|
157
|
+
|
|
158
|
+
Implementation:
|
|
159
|
+
- Finalize open ledgers using evidence when a run exits or hits a terminal incomplete state.
|
|
160
|
+
- Treat unresolved stale edit families, mutations after last verification, and missing verification after code mutations as incomplete verification.
|
|
161
|
+
- Record focus terminal-incomplete events as unresolved ledger evidence.
|
|
162
|
+
|
|
163
|
+
Acceptance:
|
|
164
|
+
- A ledger with an unverified code mutation finalizes as `incomplete_verification`.
|
|
165
|
+
- A ledger with unresolved stale edit blocks finalizes as `incomplete_verification`.
|
|
166
|
+
- A run that reaches terminal incomplete writes the ledger status before exit.
|
|
167
|
+
|
|
168
|
+
Verification:
|
|
169
|
+
- Covered by `tests/completionLedger.test.ts`; runner bridge typechecked by package build.
|
|
170
|
+
|
|
171
|
+
### WO-TD-07: Complex Tasks Must Seed A Durable Workboard Or Todo Tree
|
|
172
|
+
|
|
173
|
+
Status: implemented
|
|
174
|
+
|
|
175
|
+
Root cause: this robotics/training task had an empty workboard despite requiring discovery, model repair, integration, training, and verification.
|
|
176
|
+
|
|
177
|
+
Anchors:
|
|
178
|
+
- `packages/orchestrator/src/agenticRunner.ts`
|
|
179
|
+
- `packages/execution/src/tools/workboard.ts`
|
|
180
|
+
- `packages/execution/tests/workboard.test.ts`
|
|
181
|
+
- `packages/orchestrator/tests/agenticRunner-context-behavior.test.ts`
|
|
182
|
+
|
|
183
|
+
Implementation:
|
|
184
|
+
- Detect complex implementation/training tasks and seed a minimal board or nested todo skeleton before freeform tool use.
|
|
185
|
+
- Require cards or nested todos for repair, integration, training run, and observed verification.
|
|
186
|
+
- Attach tool evidence to active cards so completion review has structured state.
|
|
187
|
+
|
|
188
|
+
Acceptance:
|
|
189
|
+
- A multi-step training/robotics task starts with non-empty durable decomposition.
|
|
190
|
+
- Workboard cards or nested todos update as evidence is collected.
|
|
191
|
+
- Completion review sees unresolved cards/blockers when verification has not happened.
|
|
192
|
+
|
|
193
|
+
Verification:
|
|
194
|
+
- Runner seed-card path guarded by `tests/agenticRunner-context-behavior.test.ts`.
|
|
195
|
+
|
|
196
|
+
## Verification Checklist
|
|
197
|
+
|
|
198
|
+
- [x] `pnpm --filter @omnius/orchestrator test -- tests/focusSupervisor.test.ts`
|
|
199
|
+
- [x] `pnpm --filter @omnius/orchestrator test -- tests/contextEngine.test.ts`
|
|
200
|
+
- [x] `pnpm --filter @omnius/orchestrator test -- tests/completionLedger.test.ts`
|
|
201
|
+
- [x] `pnpm --filter @omnius/orchestrator test -- tests/agenticRunner-context-behavior.test.ts`
|
|
202
|
+
- [x] `pnpm --filter @omnius/orchestrator build`
|