omnius 1.0.591 → 1.0.592

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. package/.aiwg/addons/omnius-docs/README.md +15 -1
  2. package/.aiwg/addons/omnius-docs/manifest.json +28 -68
  3. package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
  4. package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
  5. package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
  6. package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
  7. package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
  8. package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
  9. package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
  10. package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
  11. package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
  12. package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
  13. package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
  14. package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
  15. package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
  16. package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
  17. package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
  18. package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
  19. package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
  20. package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
  21. package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
  22. package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
  23. package/README.md +36 -0
  24. package/dist/discovery.d.ts +50 -0
  25. package/dist/index.js +5975 -4021
  26. package/dist/library.d.ts +7 -0
  27. package/dist/library.js +950 -0
  28. package/dist/postinstall-daemon.cjs +18 -0
  29. package/dist/providerRegistry.d.ts +80 -0
  30. package/dist/service-version.d.ts +35 -0
  31. package/docs/.vitepress/config.mts +8 -0
  32. package/docs/DISCOVERY.json +20224 -0
  33. package/docs/DISCOVERY.md +648 -0
  34. package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
  35. package/docs/agent-memory/INDEX.md +9 -4
  36. package/docs/agent-memory/index.md +7 -0
  37. package/docs/concept-relational-language.md +869 -0
  38. package/docs/context-management-medium-models-proposal.md +449 -0
  39. package/docs/dedup-false-positive-meta-analysis.md +96 -0
  40. package/docs/discovery/catalog-overrides.json +724 -0
  41. package/docs/duplicate-calls-root-cause-analysis.md +91 -0
  42. package/docs/duplicate-calls-root-cause-deep.md +155 -0
  43. package/docs/ephemeral-skill-pack-small-context.md +57 -0
  44. package/docs/explorations/context-window-todo-association.md +156 -0
  45. package/docs/explorations/todo-association-verify.json +30 -0
  46. package/docs/explorations/verification-ledger.json +45 -0
  47. package/docs/explorations/verify-todo-association.sh +30 -0
  48. package/docs/flowstate.md +806 -0
  49. package/docs/getting-started/install.md +24 -0
  50. package/docs/getting-started/model-providers.md +13 -0
  51. package/docs/guides/agent-integration.md +87 -0
  52. package/docs/guides/bring-your-own-inference.md +126 -0
  53. package/docs/guides/tools-and-web-search.md +95 -0
  54. package/docs/index.md +14 -0
  55. package/docs/longhaul-35b-workorders.md +496 -0
  56. package/docs/memory-integration-analysis.md +303 -0
  57. package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
  58. package/docs/multimodal-identity-memory-implementation.md +76 -0
  59. package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
  60. package/docs/opencode-agentic-loop-comparison.md +290 -0
  61. package/docs/operations/security-and-remote-access.md +2 -2
  62. package/docs/operations/version-compatibility.md +63 -0
  63. package/docs/proposals/git-progress-tracking-strategy.md +289 -0
  64. package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
  65. package/docs/proposals/opencode-modules/childSession.ts +288 -0
  66. package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
  67. package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
  68. package/docs/proposals/opencode-modules/runner.ts +258 -0
  69. package/docs/reference/auth-map.md +87 -196
  70. package/docs/reference/configuration.md +27 -0
  71. package/docs/reference/rest-api.md +7 -0
  72. package/docs/reference/slash-commands.md +125 -2
  73. package/docs/research/_archived/README.md +18 -0
  74. package/docs/research/_archived/context_window_attention_model.py +418 -0
  75. package/docs/research/_archived/context_window_attention_spec.md +55 -0
  76. package/docs/research/_archived/context_window_attention_weights.json +68 -0
  77. package/docs/research/k-splanifolds.pdf +0 -0
  78. package/docs/research/personality-verbosity-control.md +293 -0
  79. package/docs/rest/INDEX.md +7 -0
  80. package/docs/rest/QUICKREF.md +18 -0
  81. package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
  82. package/docs/rest/auth-and-scopes.md +7 -1
  83. package/docs/rest/endpoints/discovery.md +44 -0
  84. package/docs/rest/endpoints/events.md +5 -0
  85. package/docs/rest/endpoints/tools.md +9 -0
  86. package/docs/reviews/adversary-system-review.md +42 -0
  87. package/docs/sana-and-video-generation-integration-plan.md +712 -0
  88. package/docs/session-diary-llm-training-analysis.md +218 -0
  89. package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
  90. package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
  91. package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
  92. package/docs/telegram-unified-tooling-architecture.md +332 -0
  93. package/docs/threat-model.md +868 -0
  94. package/docs/trajectory-grounding.md +160 -0
  95. package/docs/voice-flow-architecture.md +489 -0
  96. package/docs/work-orders/WO-AM-GAPS.md +638 -0
  97. package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
  98. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
  99. package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
  100. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
  101. package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
  102. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
  103. package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
  104. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
  105. package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
  106. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
  107. package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
  108. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
  109. package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
  110. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
  111. package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
  112. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
  113. package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
  114. package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
  115. package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
  116. package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
  117. package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
  118. package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
  119. package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
  120. package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
  121. package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
  122. package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
  123. package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
  124. package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
  125. package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
  126. package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
  127. package/docs/x402-remote-inference-plan.md +323 -0
  128. package/npm-shrinkwrap.json +108 -117
  129. package/package.json +7 -6
  130. package/templates/AGENTS.md +6 -0
  131. package/templates/OMNIUS.md +20 -0
@@ -0,0 +1,199 @@
1
+ # Workorder: Vision Evidence Routing Ladder
2
+
3
+ ## Objective
4
+
5
+ Build a generic visual evidence router that resolves scoped media reliably and chooses an analysis ladder based on the current request's information need:
6
+
7
+ 1. low-cost visual observation or native model image path when sufficient,
8
+ 2. OCR when text extraction is material,
9
+ 3. auxiliary vision such as Moondream/full vision when richer object, scene, UI, identity, or ambiguous evidence is needed.
10
+
11
+ This directly addresses the Telegram failure mode where the vision tool received missing `image` arguments and the agent fell back to unsupported assumptions.
12
+
13
+ ## Problem Statement
14
+
15
+ Telegram exposes media through scoped chat storage and helper tools, but the model can still call `vision` without a valid `image` parameter. Public prompt text describes a media stack, yet the executable path is distributed across Telegram media resolution, image analysis tool wrappers, OCR, and vision tools.
16
+
17
+ The desired behavior is not scenario-specific. The router should derive the needed evidence from the user's current request and the available media, then run the lowest sufficient stage and escalate only when the evidence need is not met.
18
+
19
+ ## Existing Code Anchors
20
+
21
+ Omnius:
22
+ - `packages/cli/src/tui/telegram-bridge.ts:1275` describes the public Telegram vision/media stack.
23
+ - `packages/cli/src/tui/telegram-bridge.ts:16807` builds `telegram_image_analyze`.
24
+ - `packages/cli/src/tui/telegram-bridge.ts:16840` resolves scoped Telegram media.
25
+ - `packages/cli/src/tui/telegram-bridge.ts:16848` calls staged image analysis.
26
+ - `packages/cli/src/tui/telegram-bridge.ts:16949` reports failed image evidence extraction.
27
+ - `packages/cli/src/tui/telegram-bridge.ts:17393` processes media from messages.
28
+ - `packages/execution/src/tools/vision.ts` owns the core vision tool.
29
+ - `packages/execution/src/tools/ocr-image-advanced.ts` owns OCR.
30
+ - `packages/cli/tests/telegram-media-evidence.test.ts` covers media evidence behavior.
31
+ - `packages/execution/tests/vision-action-loop.test.ts` and `packages/execution/tests/ocr-image-advanced.test.ts` are likely execution-side regression homes.
32
+
33
+ Hermes:
34
+ - `/home/robit/Documents/repositories/hermes-agent/tools/computer_use/vision_routing.py:118` decides whether capture should route to auxiliary vision.
35
+ - `/home/robit/Documents/repositories/hermes-agent/tools/vision_tools.py:527` checks whether providers accept media in tool results.
36
+ - `/home/robit/Documents/repositories/hermes-agent/tools/vision_tools.py:599` checks native vision fast path.
37
+ - `/home/robit/Documents/repositories/hermes-agent/tools/vision_tools.py:797` implements auxiliary vision analysis.
38
+
39
+ ## Target Architecture
40
+
41
+ Add `packages/execution/src/vision-evidence-router.ts` or `packages/cli/src/tui/telegram-vision-router.ts` if Telegram-specific media scope must remain in CLI.
42
+
43
+ Suggested contracts:
44
+
45
+ ```ts
46
+ export interface VisualEvidenceRequest {
47
+ prompt: string;
48
+ mediaRefs: VisualMediaRef[];
49
+ surface: "telegram" | "tui" | "api";
50
+ currentModel?: ModelCapabilitySummary;
51
+ requestedDetail?: "auto" | "low" | "text" | "full";
52
+ }
53
+
54
+ export interface VisualEvidencePlan {
55
+ stages: VisualEvidenceStage[];
56
+ reason: string;
57
+ }
58
+
59
+ export type VisualEvidenceStage =
60
+ | { kind: "resolve_media"; required: true }
61
+ | { kind: "low_fidelity_observation"; required: boolean }
62
+ | { kind: "ocr"; required: boolean }
63
+ | { kind: "auxiliary_vision"; required: boolean; modelPreference?: string };
64
+ ```
65
+
66
+ Evidence output:
67
+ - resolved media path or safe alias,
68
+ - stage outputs,
69
+ - confidence/coverage notes,
70
+ - unresolved fields,
71
+ - final evidence summary for the model.
72
+
73
+ ## Routing Principles
74
+
75
+ - Media resolution always happens before vision/OCR execution.
76
+ - Missing `image` should be repaired from current scoped media when possible.
77
+ - If multiple recent images exist, ask the model/tool caller to identify which one or use reply-to media first.
78
+ - OCR is selected when text, numbers, labels, UI copy, documents, screenshots with text, or exact wording matters.
79
+ - Auxiliary vision is selected when low-cost observation/OCR cannot answer the claim.
80
+ - The decision can be model-guided, but the router must record why each stage was selected.
81
+ - No hard-coded news/event/person scenario filters.
82
+
83
+ ## Phased Implementation
84
+
85
+ ### Phase 0: Failure Reproduction
86
+
87
+ Tasks:
88
+ - Add a regression test for Telegram image analysis where the model calls `vision` or `telegram_image_analyze` with no `image` while a replied-to image exists.
89
+ - Add a test for no available image: tool returns a clear actionable error, not a generic missing parameter error.
90
+
91
+ Completion metrics:
92
+ - Failure is reproduced before fix.
93
+ - Error message names available media aliases when they exist.
94
+
95
+ ### Phase 1: Media Resolution API
96
+
97
+ Tasks:
98
+ - Extract media resolution into a reusable function:
99
+ - reply-to image first,
100
+ - explicit alias/path second,
101
+ - latest image in current chat third,
102
+ - otherwise no-media error.
103
+ - Return structured result with provenance:
104
+ - source message id,
105
+ - chat/session key,
106
+ - media kind,
107
+ - safe alias,
108
+ - local path hidden from public output.
109
+
110
+ File-by-file notes:
111
+ - `packages/cli/src/tui/telegram-bridge.ts`: extract from `resolveTelegramScopedMediaPath`.
112
+ - `packages/cli/tests/telegram-media-evidence.test.ts`: add scoped resolution tests.
113
+
114
+ Completion metrics:
115
+ - Missing `image` argument is auto-filled only from scoped Telegram media.
116
+ - Public responses never expose local paths.
117
+
118
+ ### Phase 2: Evidence Need Decomposition
119
+
120
+ Tasks:
121
+ - Implement a small generic function that derives a `VisualEvidencePlan` from request text and available media metadata.
122
+ - It can use simple semantic categories, but not scenario-specific topics.
123
+ - If confidence is low, include multiple stages rather than making unsupported assumptions.
124
+
125
+ Completion metrics:
126
+ - "What text is in this image?" plans OCR.
127
+ - "What is happening in this photo?" plans low observation then auxiliary vision if needed.
128
+ - "Who/what is shown?" plans auxiliary vision when low observation is insufficient.
129
+ - Tests inspect plan stages and reasons.
130
+
131
+ ### Phase 3: Stage Executor
132
+
133
+ Tasks:
134
+ - Execute planned stages in order.
135
+ - Stop early only if output satisfies the requested evidence need.
136
+ - Normalize outputs into a single evidence packet.
137
+ - Add stage timings and model/tool identifiers.
138
+
139
+ Completion metrics:
140
+ - OCR output is available to the agent before final answer when text was needed.
141
+ - Auxiliary vision output is available when OCR/low observation is insufficient.
142
+ - Tool result includes evidence provenance and unresolved gaps.
143
+
144
+ ### Phase 4: Tool Integration
145
+
146
+ Tasks:
147
+ - Update `telegram_image_analyze` to call the router.
148
+ - Update core `vision` tool to require explicit `image`, but offer a Telegram wrapper that can resolve scoped media.
149
+ - Add aliases in `telegram_media_recent` output that teach the model safe media refs.
150
+ - Ensure public prompt describes the router as an evidence ladder, not as local filesystem access.
151
+
152
+ Completion metrics:
153
+ - Telegram bots no longer see `Missing required parameter(s): image` when a scoped image exists.
154
+ - If no scoped image exists, the error tells the user to send/reply to an image.
155
+ - Core non-Telegram vision remains strict about explicit paths.
156
+
157
+ ### Phase 5: Model Capability Routing
158
+
159
+ Tasks:
160
+ - If current model can accept image input natively, use native fast path for low-cost observation when supported.
161
+ - Otherwise use auxiliary model configured for vision.
162
+ - Prefer configured local Moondream/full vision when available and suitable.
163
+ - Record the selected route in the evidence packet.
164
+
165
+ Completion metrics:
166
+ - Tests mock capability true/false and assert stage selection.
167
+ - Fallback to OCR/auxiliary vision is deterministic in tests.
168
+
169
+ ## Required Tests
170
+
171
+ - Missing image argument repaired from reply-to media.
172
+ - Missing image argument repaired from latest chat image.
173
+ - No image available returns actionable scoped error.
174
+ - OCR selected for text extraction.
175
+ - Auxiliary vision selected for scene/object ambiguity.
176
+ - Native vision fast path selected when model supports media tool results.
177
+ - Public output hides local path.
178
+ - Evidence packet includes stage provenance.
179
+
180
+ ## Rollout Plan
181
+
182
+ 1. Add resolution and tests.
183
+ 2. Route Telegram image analysis through resolver.
184
+ 3. Add evidence planner and staged executor.
185
+ 4. Add model capability routing.
186
+ 5. Tighten prompt/tool descriptions after behavior is proven.
187
+
188
+ ## Risks
189
+
190
+ - Overusing expensive/full vision. Mitigation: low-cost and OCR stages first when sufficient.
191
+ - Exposing local paths in public. Mitigation: structured evidence packet with public-safe aliases.
192
+ - Router becoming heuristic-heavy. Mitigation: restrict rules to evidence modality needs, not scenario content.
193
+
194
+ ## Definition of Done
195
+
196
+ - Telegram media calls no longer fail with missing `image` when scoped media exists.
197
+ - Visual analysis escalates through documented stages based on evidence need.
198
+ - The agent receives structured evidence before answering.
199
+ - Tests cover no-media, OCR, native vision, and auxiliary vision paths.
@@ -0,0 +1,20 @@
1
+ # Durable Multi-Agent Kanban Index
2
+
3
+ Progress: not started.
4
+
5
+ Files in this folder:
6
+ - [WORKORDER.md](WORKORDER.md) - Full phased implementation workorder.
7
+
8
+ Primary Omnius anchors:
9
+ - `packages/orchestrator/src/agenticRunner.ts:1414` - completion hold loop regression tests.
10
+ - `packages/orchestrator/src/agenticRunner.ts:10011` - todo planning protocol.
11
+ - `packages/cli/src/api/task-manager-singleton.ts` - API task management surface.
12
+ - `packages/cli/src/tui/telegram-bridge.ts:13368` - Telegram sub-agent lifecycle.
13
+
14
+ Primary Hermes anchors:
15
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_swarm.py:151` - swarm spec.
16
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_diagnostics.py:326` - hallucinated completion guard.
17
+ - `/home/robit/Documents/repositories/hermes-agent/docs/kanban/multi-gateway.md:9` - single dispatcher ownership.
18
+
19
+ Completion target:
20
+ - Long or disputed work can be decomposed into durable, inspectable cards with worker/verifier/synthesizer roles and explicit completion evidence.
@@ -0,0 +1,174 @@
1
+ # Workorder: Durable Multi-Agent Kanban
2
+
3
+ ## Objective
4
+
5
+ Add a durable multi-agent work board for long-running admin tasks and public disputed factual work. The board should support decomposition, worker assignment, verification, synthesis, diagnostics, and completion evidence without relying on one runner's transient todo list.
6
+
7
+ ## Problem Statement
8
+
9
+ Omnius has todos and sub-agents, but large tasks still depend on a single runner maintaining context and self-verifying. Hermes added a Kanban-style durable board, swarm worker specs, dispatcher ownership, and diagnostics for hallucinated card completions.
10
+
11
+ Omnius should adopt the architectural shape where it improves carryover:
12
+ - durable work decomposition,
13
+ - worker/verifier/synthesizer lanes,
14
+ - explicit evidence attached to card completion,
15
+ - diagnostics when workers claim cards that do not exist or lack evidence.
16
+
17
+ ## Existing Code Anchors
18
+
19
+ Omnius:
20
+ - `packages/orchestrator/src/agenticRunner.ts:10011` builds todo guidance.
21
+ - `packages/orchestrator/src/agenticRunner.ts:1414` covers persistent completion-provenance hold regression tests.
22
+ - `packages/orchestrator/src/completionContract.ts:233` builds completion scenario decomposition.
23
+ - `packages/cli/src/api/task-manager-singleton.ts` is a task management anchor.
24
+ - `packages/cli/src/tui/telegram-bridge.ts:13368` creates Telegram sub-agent state.
25
+ - `packages/execution/src/tools/full-sub-agent.ts` is a likely worker spawning reference.
26
+
27
+ Hermes:
28
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_swarm.py:151` builds swarm metadata.
29
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_swarm.py:164` creates worker-assigned tasks.
30
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_diagnostics.py:326` handles hallucinated completion.
31
+ - `/home/robit/Documents/repositories/hermes-agent/hermes_cli/kanban_diagnostics.py:520` diagnoses repeated failure loops.
32
+ - `/home/robit/Documents/repositories/hermes-agent/docs/kanban/multi-gateway.md:9` documents single dispatcher ownership.
33
+
34
+ ## Target Architecture
35
+
36
+ Add an Omnius board under `.omnius/workboards/` with JSONL event log and compact active snapshot.
37
+
38
+ Entities:
39
+ - `Board`: run/project scope, owner, status, created/updated.
40
+ - `Card`: title, description, lane, assignee, status, dependencies, evidence requirements.
41
+ - `Evidence`: tool result references, files, URLs, messages, verification notes.
42
+ - `Decision`: why card was split, blocked, reassigned, or completed.
43
+
44
+ Roles:
45
+ - Decomposer: turns a goal into cards and evidence expectations.
46
+ - Worker: executes a card.
47
+ - Verifier: checks evidence and marks card verified or needs changes.
48
+ - Synthesizer: builds final answer from verified cards and unresolved blockers.
49
+
50
+ Important constraint:
51
+ - The board must not become a hard-coded scenario matrix. Card schemas are generic; card content is generated from the actual request and observations.
52
+
53
+ ## Phased Implementation
54
+
55
+ ### Phase 0: Scope and Data Model
56
+
57
+ Tasks:
58
+ - Define board/card/evidence JSON schemas.
59
+ - Decide storage path:
60
+ - `.omnius/workboards/<run-id>/events.jsonl`
61
+ - `.omnius/workboards/<run-id>/active.json`
62
+ - Add pure model tests for event replay and status transitions.
63
+
64
+ Completion metrics:
65
+ - Event replay reconstructs active board.
66
+ - Invalid transition is rejected.
67
+ - Evidence attachment is required before verified completion.
68
+
69
+ ### Phase 1: Board Tools
70
+
71
+ Tasks:
72
+ - Add execution tools:
73
+ - `workboard_create`
74
+ - `workboard_add_card`
75
+ - `workboard_update_card`
76
+ - `workboard_attach_evidence`
77
+ - `workboard_complete_card`
78
+ - `workboard_show`
79
+ - Keep tools generic and run-scoped by default.
80
+
81
+ File-by-file notes:
82
+ - `packages/execution/src/tools/workboard.ts`: new tool module.
83
+ - `packages/execution/src/tools/index.ts` or manifest registry: register tools.
84
+ - `packages/execution/tests/workboard.test.ts`: model/tool tests.
85
+
86
+ Completion metrics:
87
+ - Tools mutate only the scoped board.
88
+ - Completion without evidence returns structured failure.
89
+ - Board show is compact enough for model context.
90
+
91
+ ### Phase 2: Runner Integration
92
+
93
+ Tasks:
94
+ - For long admin runs, expose board tools and inject a compact board status.
95
+ - Keep normal simple tasks on existing todo path unless the agent chooses board decomposition.
96
+ - Add an option to create a board at run start for admin tasks above a complexity threshold selected by the model or user, not by scenario keywords.
97
+
98
+ Completion metrics:
99
+ - Simple tasks are not forced into board flow.
100
+ - Long tasks can create board and survive compaction.
101
+ - Board state appears in completion critic context.
102
+
103
+ ### Phase 3: Worker and Verifier Roles
104
+
105
+ Tasks:
106
+ - Add worker prompt templates for a single card.
107
+ - Add verifier prompt template that reviews card evidence.
108
+ - Ensure worker completion cannot mark card verified directly unless policy allows.
109
+ - Add synthesizer prompt that composes final answer only from verified cards and named blockers.
110
+
111
+ Completion metrics:
112
+ - Worker can complete card with evidence.
113
+ - Verifier can request changes with specific missing evidence.
114
+ - Synthesizer refuses to claim unverified cards as done.
115
+
116
+ ### Phase 4: Diagnostics
117
+
118
+ Tasks:
119
+ - Add diagnostics similar to Hermes:
120
+ - card id not found,
121
+ - repeated worker failure,
122
+ - assignee unavailable,
123
+ - completion claim without evidence,
124
+ - dependency blocked.
125
+ - Add `/workboard diagnostics` or API endpoint.
126
+
127
+ Completion metrics:
128
+ - A hallucinated card id is caught and reported.
129
+ - Repeated failures produce suggested next actions.
130
+ - Diagnostics do not auto-complete blocked work.
131
+
132
+ ### Phase 5: Telegram/Public Use
133
+
134
+ Tasks:
135
+ - Public group disputed facts can optionally create a small board when scrutiny is strict and the agent declares the need.
136
+ - Board details remain internal unless the final answer intentionally summarizes sources/evidence.
137
+ - Admin can inspect board via Telegram DM or TUI.
138
+
139
+ Completion metrics:
140
+ - Public answers remain concise.
141
+ - Admin can inspect board state.
142
+ - Public strict scrutiny can use board evidence without leaking internal paths.
143
+
144
+ ## Required Tests
145
+
146
+ - Board event replay and compaction.
147
+ - Card cannot verify without evidence.
148
+ - Worker completion attaches evidence.
149
+ - Verifier request_changes keeps card open.
150
+ - Synthesizer excludes unverified cards.
151
+ - Hallucinated card id diagnostic.
152
+ - Repeated worker failure diagnostic.
153
+ - Runner includes board status in critic context.
154
+
155
+ ## Rollout Plan
156
+
157
+ 1. Add board storage and tools.
158
+ 2. Add optional admin-run integration.
159
+ 3. Add verifier/synthesizer roles.
160
+ 4. Add diagnostics.
161
+ 5. Add public strict-scrutiny integration.
162
+
163
+ ## Risks
164
+
165
+ - Adds overhead to simple tasks. Mitigation: optional and model/user-selected.
166
+ - Creates another planning surface beside todos. Mitigation: board is for durable multi-agent work; todos remain lightweight.
167
+ - Worker self-certification. Mitigation: verifier lane and evidence-required transitions.
168
+
169
+ ## Definition of Done
170
+
171
+ - Durable board state survives runner compaction and process restarts.
172
+ - Cards require evidence for verified completion.
173
+ - Diagnostics catch hallucinated or unsupported completion.
174
+ - Runner and critic can consume board state generically.
@@ -0,0 +1,22 @@
1
+ # Completion Critic Reconciliation Ledger Index
2
+
3
+ Progress: not started.
4
+
5
+ Files in this folder:
6
+ - [WORKORDER.md](WORKORDER.md) - Full phased implementation workorder.
7
+
8
+ Primary Omnius anchors:
9
+ - `packages/orchestrator/src/completionContract.ts:128` - generic completion contract inference.
10
+ - `packages/orchestrator/src/completionContract.ts:233` - completion meta-decomposition.
11
+ - `packages/orchestrator/src/backward-pass-critic.ts:188` - reconciliation packet rendering.
12
+ - `packages/orchestrator/src/agenticRunner.ts:3684` - critic evidence collection.
13
+ - `packages/orchestrator/src/agenticRunner.ts:3729` - review reconciliation packet.
14
+ - `packages/orchestrator/src/agenticRunner.ts:15748` - final run status selection.
15
+
16
+ Primary Hermes anchors:
17
+ - `/home/robit/Documents/repositories/hermes-agent/agent/tool_result_classification.py` - tool-result classification reference.
18
+ - `/home/robit/Documents/repositories/hermes-agent/tools/tool_result_storage.py:122` - persisted tool-result storage.
19
+ - `/home/robit/Documents/repositories/hermes-agent/agent/background_review.py` - background review reference.
20
+
21
+ Completion target:
22
+ - Completion review is driven by a durable claim/evidence ledger and critic reconciliation packets, not by optional model wording or hard-coded scenario profiles.
@@ -0,0 +1,226 @@
1
+ # Workorder: Completion Critic Reconciliation Ledger
2
+
3
+ ## Objective
4
+
5
+ Convert completion verification into a durable claim/evidence ledger. The critic must receive:
6
+ - proposed final claims,
7
+ - evidence observed before the critique,
8
+ - prior critique feedback,
9
+ - evidence observed after the critique,
10
+ - unresolved loops and failed observations,
11
+ - final status constraints.
12
+
13
+ The system must block unsupported success claims but avoid unbounded loops by ending as `incomplete_verification` when repeated attempts cannot satisfy the ledger.
14
+
15
+ ## Problem Statement
16
+
17
+ Recent work improved completion contracts and critic reconciliation, but the architecture still needs a durable source of truth. Evidence currently flows through in-memory tool call logs and prompt packets. The next step is to make claim/evidence state explicit so the critic and renderer cannot drift apart.
18
+
19
+ The design must not use shortcut wording or scenario-specific keyword lists. It should derive completion criteria from the user request, proposed summary, and observed data.
20
+
21
+ ## Existing Code Anchors
22
+
23
+ Omnius:
24
+ - `packages/orchestrator/src/completionContract.ts:128` creates generic completion contracts.
25
+ - `packages/orchestrator/src/completionContract.ts:233` builds completion scenario decomposition.
26
+ - `packages/orchestrator/src/completionContract.ts:280` instructs reviewer to invent a situation-specific scenario from request and observations.
27
+ - `packages/orchestrator/src/backward-pass-critic.ts:188` renders reconciliation packets.
28
+ - `packages/orchestrator/src/backward-pass-critic.ts:224` tells critic how to compare prior critique and new evidence.
29
+ - `packages/orchestrator/src/agenticRunner.ts:3684` collects critic tool evidence.
30
+ - `packages/orchestrator/src/agenticRunner.ts:3729` builds `reviewReconciliation`.
31
+ - `packages/orchestrator/src/agenticRunner.ts:3816` handles request_changes or blocked.
32
+ - `packages/orchestrator/src/agenticRunner.ts:15748` selects final run status.
33
+ - `packages/orchestrator/tests/agenticRunner.test.ts:1418` tests persistent provenance hold ending as incomplete verification.
34
+ - `packages/orchestrator/tests/completionContract.test.ts:117` tests incomplete verification contract text.
35
+
36
+ Hermes:
37
+ - `/home/robit/Documents/repositories/hermes-agent/agent/tool_result_classification.py` classifies tool result significance.
38
+ - `/home/robit/Documents/repositories/hermes-agent/tools/tool_result_storage.py:122` persists tool results.
39
+ - `/home/robit/Documents/repositories/hermes-agent/agent/background_review.py` provides background review reference patterns.
40
+
41
+ ## Target Architecture
42
+
43
+ Add a run-local ledger:
44
+
45
+ ```ts
46
+ interface CompletionLedger {
47
+ runId: string;
48
+ goal: string;
49
+ createdAtIso: string;
50
+ updatedAtIso: string;
51
+ proposedClaims: CompletionClaim[];
52
+ evidence: CompletionEvidence[];
53
+ critiques: CompletionCritiqueRecord[];
54
+ unresolved: CompletionUnresolvedItem[];
55
+ status: "open" | "approved" | "request_changes" | "blocked" | "incomplete_verification";
56
+ }
57
+
58
+ interface CompletionClaim {
59
+ id: string;
60
+ text: string;
61
+ source: "task_complete_summary" | "assistant_visible_text" | "derived";
62
+ materiality: "low" | "medium" | "high";
63
+ evidenceRequirement: string;
64
+ evidenceIds: string[];
65
+ status: "supported" | "unsupported" | "contradicted" | "unverified" | "blocked";
66
+ }
67
+
68
+ interface CompletionEvidence {
69
+ id: string;
70
+ kind: "tool_result" | "file_change" | "delivery_result" | "runtime_observation" | "human_input" | "critic_feedback";
71
+ observedAtIso: string;
72
+ toolName?: string;
73
+ success?: boolean;
74
+ summary: string;
75
+ rawRef?: string;
76
+ }
77
+ ```
78
+
79
+ Storage:
80
+ - `.omnius/completion-ledgers/<run-id>.json`
81
+ - append-only JSONL optional for recovery.
82
+
83
+ ## Phased Implementation
84
+
85
+ ### Phase 0: Current Behavior Tests
86
+
87
+ Tasks:
88
+ - Add tests for:
89
+ - request_changes blocks completion and loops,
90
+ - evidence after request_changes is present in next critic prompt,
91
+ - blocked critic ends as incomplete verification,
92
+ - repeated holds end as incomplete verification, not completed,
93
+ - simple public Telegram task still completes when guard disabled.
94
+
95
+ Completion metrics:
96
+ - Tests lock the current repaired behavior before ledger extraction.
97
+
98
+ ### Phase 1: Ledger Data Model
99
+
100
+ Tasks:
101
+ - Add `completionLedger.ts` with pure functions:
102
+ - `createCompletionLedger`
103
+ - `recordCompletionEvidence`
104
+ - `deriveClaimsFromProposedText`
105
+ - `reconcileClaimsWithEvidence`
106
+ - `recordCritique`
107
+ - `buildCriticPacketFromLedger`
108
+ - Use generic claim extraction prompts or deterministic scaffolding. Do not encode scenario-specific evidence profiles.
109
+
110
+ File-by-file notes:
111
+ - `packages/orchestrator/src/completionLedger.ts`: pure model and serializer.
112
+ - `packages/orchestrator/tests/completionLedger.test.ts`: model tests.
113
+
114
+ Completion metrics:
115
+ - Claims and evidence can be recorded and serialized.
116
+ - Reconciliation packet can be produced without an `AgenticRunner` instance.
117
+
118
+ ### Phase 2: Evidence Ingestion
119
+
120
+ Tasks:
121
+ - Convert tool call log entries into ledger evidence as tools finish.
122
+ - Record mutated files and delivery results.
123
+ - Record failed tool results, not only successful evidence.
124
+ - Preserve evidence ids in typed run events if the typed event stream workorder has landed.
125
+
126
+ Completion metrics:
127
+ - Ledger contains shell/API/browser/delivery evidence with success flags.
128
+ - Failed observations are visible to critic.
129
+ - Evidence after prior critique is distinguishable by timestamp or sequence.
130
+
131
+ ### Phase 3: Claim Extraction and Scenario Decomposition
132
+
133
+ Tasks:
134
+ - When `task_complete` is attempted, derive claims from:
135
+ - task_complete summary,
136
+ - visible assistant answer,
137
+ - original user request,
138
+ - open todos/board cards.
139
+ - Reuse `buildCompletionScenarioDecomposition`, but make it consume ledger evidence.
140
+ - Store generated decomposition in the ledger for audit.
141
+
142
+ Completion metrics:
143
+ - Completion scenario is generated from actual claims and observations.
144
+ - Unsupported claims are named as unverified/blocked instead of silently ignored.
145
+ - Unit tests use different task domains without hard-coded scenario branches.
146
+
147
+ ### Phase 4: Critic Packet and Re-review
148
+
149
+ Tasks:
150
+ - Build critic prompt from ledger:
151
+ - all material claims,
152
+ - evidence mapped to each claim,
153
+ - prior critique,
154
+ - evidence recorded after prior critique,
155
+ - unresolved loops.
156
+ - On `request_changes`, store critique and feedback.
157
+ - On next completion attempt, include reconciliation packet.
158
+ - If new evidence resolves the prior critique, critic prompt must explicitly tell critic not to repeat resolved objections.
159
+
160
+ Completion metrics:
161
+ - Test critic prompt includes prior critique and post-critique evidence.
162
+ - Test resolved evidence appears under the right prior critique.
163
+ - Test no evidence after critique is clearly represented.
164
+
165
+ ### Phase 5: Terminal State Discipline
166
+
167
+ Tasks:
168
+ - Only `approve` can set `completed=true`.
169
+ - `request_changes` continues unless cycle budget is exhausted.
170
+ - `blocked` sets `incomplete_verification`.
171
+ - Cycle exhaustion sets `incomplete_verification`, never `completed`.
172
+ - Renderer receives status separately from human-facing summary.
173
+
174
+ Completion metrics:
175
+ - `AgenticResult.status` is accurate.
176
+ - Telegram admin wrapper refuses "completed" for `incomplete_verification`.
177
+ - Public simple tasks still complete when configured fast.
178
+
179
+ ### Phase 6: User-Facing Provenance Summary
180
+
181
+ Tasks:
182
+ - Admin final response sections:
183
+ - Done,
184
+ - Verified,
185
+ - Not verified / blocked,
186
+ - Evidence.
187
+ - Public strict mode can cite evidence compactly without exposing internals.
188
+ - TUI can show ledger path for inspection.
189
+
190
+ Completion metrics:
191
+ - Admin final text includes verified and unverified sections.
192
+ - Ledger file path appears in TUI/admin diagnostics, not normal public chat.
193
+
194
+ ## Required Tests
195
+
196
+ - Ledger serializes and replays.
197
+ - Claim extraction handles unrelated task domains without scenario keyword logic.
198
+ - Evidence map attaches tool results to proposed claims.
199
+ - Critic prompt includes prior critique and evidence since prior critique.
200
+ - Request_changes blocks completion.
201
+ - Blocked critic yields `incomplete_verification`.
202
+ - Cycle exhaustion yields `incomplete_verification`.
203
+ - Telegram admin does not render incomplete verification as completed.
204
+ - Public fast simple task still reaches task_complete.
205
+
206
+ ## Rollout Plan
207
+
208
+ 1. Add ledger pure model and tests.
209
+ 2. Record evidence into ledger while preserving current critic behavior.
210
+ 3. Build critic packets from ledger.
211
+ 4. Use ledger for final status and admin provenance rendering.
212
+ 5. Remove duplicate in-memory-only evidence paths.
213
+
214
+ ## Risks
215
+
216
+ - Ledger bloat. Mitigation: store raw refs and compact summaries, not full large outputs.
217
+ - Critic over-strictness. Mitigation: materiality, cycle budget, and incomplete verification terminal status.
218
+ - Reward-hacking through claim wording. Mitigation: derive claims from visible answer and task summary, not only self-declared fields.
219
+
220
+ ## Definition of Done
221
+
222
+ - Completion verification has a durable ledger per run.
223
+ - Critic receives a complete reconciliation packet on re-review.
224
+ - Unsupported success claims cannot complete.
225
+ - Repeated unresolved holds terminate as incomplete verification.
226
+ - Tests prove simple Telegram tasks still complete under fast public mode.