omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,76 @@
|
|
|
1
|
+
# Multimodal Identity Memory Implementation Tracker
|
|
2
|
+
|
|
3
|
+
Goal: make voice, text, image, Telegram reply context, CLIP-like embeddings, zettelkasten links, and social profiles converge on one evidence-based identity substrate.
|
|
4
|
+
|
|
5
|
+
## Scope
|
|
6
|
+
|
|
7
|
+
- [ ] Route Telegram, TUI, GUI, voice, and API media ingestion through one identity/evidence layer.
|
|
8
|
+
- [ ] Keep private/public/terminal/gui scopes explicit on every assertion and retrieval path.
|
|
9
|
+
- [ ] Represent uploader/sender, depicted person, speaker, reply target, media asset, and message as separate graph atoms.
|
|
10
|
+
- [ ] Preserve CLIP/image, CLIP/text, face, speaker, transcript, and aligned embeddings with explicit vector-space metadata.
|
|
11
|
+
- [ ] Use agentic structured identity extraction later for "this is X" style assertions; avoid hard-coded naming heuristics.
|
|
12
|
+
- [ ] Feed durable interaction/profile facts into social memory and zettelkasten instead of isolated JSON-only stores.
|
|
13
|
+
|
|
14
|
+
## Code Anchors
|
|
15
|
+
|
|
16
|
+
- `MMID-001` - central multimodal identity/evidence service.
|
|
17
|
+
- `MMID-002` - API memory ingest uses shared orchestrator DB names.
|
|
18
|
+
- `MMID-003` - Telegram media ingest sends structured source/scope/sender/message metadata.
|
|
19
|
+
- `MMID-004` - embedding-space handling stores OpenCLIP-compatible vectors as `clipEmbedding`.
|
|
20
|
+
- `MMID-005` - scoped graph atoms for sender, message, media asset, reply target, and identity assertion.
|
|
21
|
+
- `MMID-006` - dependency/worker path for image/audio/text embedding providers.
|
|
22
|
+
- `MMID-007` - Telegram/TUI tool exposure and context injection.
|
|
23
|
+
- `MMID-008` - backward-pass validation checklist and tests.
|
|
24
|
+
- `MMID-009` - scoped automatic prior-enrolled visual identity association and recall injection.
|
|
25
|
+
- `MMID-010` - unified `.omnius/episodes.db` / `.omnius/knowledge.db` paths for identity, reflection, and API media ingest.
|
|
26
|
+
- `MMID-011` - natural chronology support for explicit identity-before-image assertions and unknown-face steering.
|
|
27
|
+
|
|
28
|
+
## Implementation Checklist
|
|
29
|
+
|
|
30
|
+
- [x] Create this tracking document before code changes.
|
|
31
|
+
- [x] Add multimodal identity service in `@omnius/memory`.
|
|
32
|
+
- [x] Extend graph relation vocabulary for identity evidence.
|
|
33
|
+
- [x] Make `/v1/memory/ingest` use `episodes.db` and `knowledge.db`.
|
|
34
|
+
- [x] Replace legacy `labels`-as-person behavior with structured sender/media/assertion handling.
|
|
35
|
+
- [x] Store supplied image/text CLIP vectors in `clipEmbedding`.
|
|
36
|
+
- [x] Store supplied speaker/audio vectors in native `embedding` with vector-space metadata.
|
|
37
|
+
- [x] Pass Telegram source scope, sender, message id, media id, caption, and reply context into ingest.
|
|
38
|
+
- [x] Keep legacy ingest clients compatible.
|
|
39
|
+
- [x] Add tests for visual media not misbinding uploader as depicted person.
|
|
40
|
+
- [x] Add tests for explicit identity assertion creating scoped `named_as` / `depicts` evidence.
|
|
41
|
+
- [x] Add tests for audio voice sample linking to sender without global leakage.
|
|
42
|
+
- [x] Add structured JSON output mode for `visual_memory identify` so association code does not parse display text.
|
|
43
|
+
- [x] Add scoped visual identity association service that records prior enrolled face matches as graph evidence.
|
|
44
|
+
- [x] Inject scoped visual identity recall into Telegram and TUI image ingress.
|
|
45
|
+
- [x] Return visual identity recall metadata from `/v1/memory/ingest` for GUI/API callers.
|
|
46
|
+
- [x] Wire GUI attachment upload into scoped media ingest and prepend returned identity context to the next chat turn.
|
|
47
|
+
- [x] Move identity/reflection stores onto the same `.omnius/episodes.db` and `.omnius/knowledge.db` files used by the orchestrator and API.
|
|
48
|
+
- [x] Add `identity_memory action='stage_identity'` for explicit "next image/media is <name>" chronology without regex parsing.
|
|
49
|
+
- [x] Apply pending same-scope visual identity assertions when a later image arrives and face enrollment succeeds.
|
|
50
|
+
- [x] Add unknown-face context when structured detection sees a face but no enrolled identity matches.
|
|
51
|
+
- [x] Run focused builds/tests.
|
|
52
|
+
- [x] Backward pass: verify each item above against code anchors and test coverage.
|
|
53
|
+
|
|
54
|
+
## Backward Pass Verification
|
|
55
|
+
|
|
56
|
+
- `MMID-001`: implemented `MultimodalIdentityService` and exported it from `@omnius/memory`.
|
|
57
|
+
- `MMID-002`: API ingest, API memory search/entities, route-v1 memory stores, embedding workers, and `multimodal_memory` bridge now use `episodes.db` / `knowledge.db`.
|
|
58
|
+
- `MMID-003`: Telegram media ingest now sends structured source/scope/sender/message/reply/media payloads.
|
|
59
|
+
- `MMID-004`: OpenCLIP-compatible vectors are stored via `setClipEmbedding`; speaker/audio vectors are stored as native embeddings.
|
|
60
|
+
- `MMID-005`: tests cover distinct sender, media, message, reply/identity evidence graph atoms.
|
|
61
|
+
- `MMID-006`: provider hooks are used when configured; full CLIP/ECAPA dependency install is opt-in via `OMNIUS_INSTALL_FULL_EMBED_DEPS=1`.
|
|
62
|
+
- `MMID-007`: identity memory remains an agent-facing tool; explicit name/enrollment requests are tool calls, not regex captures.
|
|
63
|
+
- `MMID-008`: verified with `pnpm -r build`, `pnpm --filter omnius test -- tests/identity-memory-tool.test.ts tests/visual-identity-association.test.ts`, `pnpm --filter omnius test -- tests/telegram-reflection-corpus.test.ts tests/telegram-reflection-extraction.test.ts`, `pnpm --filter omnius test -- tests/telegram-bot-api-10.test.ts`, and `git diff --check`. The visual association tests now cover prior-match recall, cross-scope isolation, pending next-image application, and unknown-face steering.
|
|
64
|
+
- `MMID-009`: `associateVisualIdentityFromImage` calls `visual_memory identify` with `format=json`, commits `same_person_candidate` / `depicts` evidence for prior enrolled matches, and formats same-scope recall context.
|
|
65
|
+
- `MMID-010`: `IdentityMemoryTool`, Telegram reflection corpus/extraction, API ingest, memory menus, and orchestrator memory now converge on `.omnius/episodes.db` / `.omnius/knowledge.db`.
|
|
66
|
+
- `MMID-011`: `stageVisualIdentityAssertion` records explicit pending next-image identity evidence; later image ingress consumes it only in the same scope/session and only after face enrollment succeeds. Unknown face detection produces a model-facing prompt to ask the user instead of guessing.
|
|
67
|
+
|
|
68
|
+
## Follow-Up Work After First Slice
|
|
69
|
+
|
|
70
|
+
- [ ] Add agentic structured identity assertion extraction from reply/message/media context.
|
|
71
|
+
- [x] Wire `visual_memory` face enrollment and prior-enrolled recognition into the identity graph.
|
|
72
|
+
- [ ] Extend the same structured association path to taught object CLIP recognition.
|
|
73
|
+
- [ ] Wire ECAPA speaker embeddings into Telegram voice/audio ingest.
|
|
74
|
+
- [x] Add retrieval context injection for scoped visual identity graph neighborhoods.
|
|
75
|
+
- [ ] Mirror scoped behavior/personality facts into `SocialMemoryStore`.
|
|
76
|
+
- [ ] Expose scoped Telegram tools for identity recall/enrollment with privacy guards.
|
|
@@ -0,0 +1,169 @@
|
|
|
1
|
+
# Omnius Self-Edit Evaluation — Uncommitted Changes & Duplicate-Tool-Call Failure Mode
|
|
2
|
+
|
|
3
|
+
> Eval of the working-tree changes against `docs/opencode-agentic-loop-comparison.md`.
|
|
4
|
+
> Subject: an Omnius instance editing its own orchestrator to implement the P0–P10
|
|
5
|
+
> learnings. Generated 2026-06-10.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## TL;DR
|
|
10
|
+
|
|
11
|
+
1. **The self-edit failed.** The agent attempted the full P0–P10 refactor, produced
|
|
12
|
+
6 new modules plus in-place edits, but **the orchestrator package does not compile**
|
|
13
|
+
(`tsc --noEmit` → exit 2, ~40+ errors). The 6 new modules are **dead code** — not
|
|
14
|
+
imported anywhere except one broken import in `agent-tool.ts`.
|
|
15
|
+
2. **The live instance is unaffected — by luck.** `dist/agenticRunner.js` was built at
|
|
16
|
+
02:35; the broken source edits landed at 05:06 and never recompiled. The running
|
|
17
|
+
self-editing instance is still executing the **old, working dist**. The build silently
|
|
18
|
+
stayed green-at-runtime because the broken code was never built.
|
|
19
|
+
3. **The reported duplicate-tool-call / duplicate-read failure mode is NOT caused by
|
|
20
|
+
these edits.** It lives in the existing (working) dedup architecture. Root cause:
|
|
21
|
+
dedup is **reactive, cross-turn-blind, and result-agnostic** — duplicates are executed
|
|
22
|
+
then cleaned up, and repeated *failures* have their evidence pruned away.
|
|
23
|
+
4. **The diff is itself a symptom of the same pathology.** The agent pasted **three
|
|
24
|
+
near-identical compaction blocks** into one loop and hallucinated an entire `index.ts`
|
|
25
|
+
export surface. "Duplicate, low-grounding output" shows up at both the tool-call layer
|
|
26
|
+
(runtime) and the code-generation layer (this diff).
|
|
27
|
+
|
|
28
|
+
---
|
|
29
|
+
|
|
30
|
+
## Part 1 — How well did the self-editing agent do?
|
|
31
|
+
|
|
32
|
+
### Verdict: **Poor / non-shippable.** Ambitious scope, no verification, broken build.
|
|
33
|
+
|
|
34
|
+
The agent violated the project's own `executable-feedback` rule (code-generating agents
|
|
35
|
+
must compile/test before returning). None of this compiles.
|
|
36
|
+
|
|
37
|
+
### 1.1 What was produced
|
|
38
|
+
|
|
39
|
+
| Module | P-item | Lines | Wired in? | Quality |
|
|
40
|
+
| --- | --- | --- | --- | --- |
|
|
41
|
+
| `runner.ts` | P3/P7 | 258 | ❌ no | Toy. Regex `<tool_call>` parsing (Omnius uses native OpenAI tool calls). Name **collides** with the real `AgenticRunner`. Re-calls the LLM after tools and discards the result. |
|
|
42
|
+
| `orchestrator.ts` | P4/P7 | 387 | ❌ no | Toy. `checkCompletion` = substring match on last message. `prune` strategy keeps "every other message" — would split `tool_call`/`tool_result` pairs. |
|
|
43
|
+
| `childSession.ts` | P0/P10 | 288 | ❌ (import is broken) | Best of the six. Plausible permission-derivation + `task_id` resume shape. Glob matcher is naive. |
|
|
44
|
+
| `permissionRuleset.ts` | P1/P6 | 166 | ❌ no | Reasonable. `evaluatePermission` cascade + `checkDoomLoop`. Never called. |
|
|
45
|
+
| `compactionAgent.ts` | P9 | 101 | ⚠️ imported, wrong shape | **Placeholder** — `compress()` never calls an LLM; returns "last 10 messages" stub. Imports `ChatMessage` which `agent-types.ts` does not export. |
|
|
46
|
+
| `backendAdapter.ts` | P8 | 443 | ❌ no | Three backend adapters. **Single-tool-call only** (`currentToolName`/`accumulatedArgs` are scalars) — contradicts the parallel-batching premise. Ollama branch double-accumulates `toolCallBuffer`. |
|
|
47
|
+
|
|
48
|
+
### 1.2 What broke (compile errors, by file)
|
|
49
|
+
|
|
50
|
+
- **`agenticRunner.ts`** — the loop got **three stacked compaction blocks** (lines ~9474,
|
|
51
|
+
~9489, ~9495), each calling a different non-existent API:
|
|
52
|
+
- `compactionResult.success` / `.compressedContext` — real shape is `{ compacted, messages, summary }`.
|
|
53
|
+
- `this.messages` — the field is not named `messages` on the class (`TS2339` ×7).
|
|
54
|
+
- `new CompactionAgent({ model })` — `CompactionAgent` is **not imported** here and is a
|
|
55
|
+
const object, not a class (`TS2304`); its method is `compress`, not `compact`;
|
|
56
|
+
`result.tokenCount` doesn't exist; `this.options.model` doesn't exist.
|
|
57
|
+
- `estimateMessagesTokens` imported from `textSanitize.js` **and** redeclared → `TS2440`.
|
|
58
|
+
- **`index.ts`** — rewritten into a "clean" barrel that re-exports a **hallucinated API
|
|
59
|
+
surface**: `SessionManager`, `SessionConfig` (module exports `SessionOrchestrator`),
|
|
60
|
+
`AgentType`, `AgentMode`, `AgentConfig`, `AgentState`, `AgentStatus`, `CompactionConfig`,
|
|
61
|
+
`AgentTask` — none exist. Plus duplicate `LLMBackend` and `export type` violations under
|
|
62
|
+
`verbatimModuleSyntax`. This **also deleted real exports** the rest of the repo depends on
|
|
63
|
+
(`OllamaPool`, `discoverSystemOllamaModelStore`, `modelProfile` helpers, …) — `cli` imports
|
|
64
|
+
these from `@omnius/orchestrator` in ~30 files.
|
|
65
|
+
- **`compactionAgent.ts`** — `import { ChatMessage }` fails (`TS2305`).
|
|
66
|
+
- **`agent-tool.ts`** — import line injected **into the file's top doc-comment** (line 2),
|
|
67
|
+
then duplicated correctly at line 28. Cross-package deep import
|
|
68
|
+
`../../orchestrator/src/childSession.js` from `packages/execution` breaks package boundaries.
|
|
69
|
+
Calls `this.callbacks.getParentSessionId?.()` / `resolveParentPermissions?.()` callbacks
|
|
70
|
+
that don't exist on the interface.
|
|
71
|
+
|
|
72
|
+
### 1.3 What was actually fine
|
|
73
|
+
|
|
74
|
+
A few in-place edits are individually sound and worth salvaging:
|
|
75
|
+
- `agent-types.ts` — adding `mode: "primary" | "subagent" | "all"` + `canSpawnAsSubagent` /
|
|
76
|
+
`canUserSelect` (P1) is clean. **But** it renamed `plan` → `architect` and dropped
|
|
77
|
+
`getAgentTypeRegistry` / `resolveAgentTools` / `registerAgentType` / `buildAgentTypeSummary`,
|
|
78
|
+
which are referenced elsewhere — breaking callers.
|
|
79
|
+
- `tool-batching.ts` — Zod schemas + `validateToolArgs` (P5) are self-contained and harmless.
|
|
80
|
+
- `agenticRunner.ts` — the `finish_reason` plumbing on the Ollama backend (P3) is correct and
|
|
81
|
+
is the one genuinely useful, low-risk change in the whole diff.
|
|
82
|
+
- `textSanitize.ts` — `estimateMessagesTokens` helper is fine (just double-imported elsewhere).
|
|
83
|
+
|
|
84
|
+
### 1.4 Behavioral read
|
|
85
|
+
|
|
86
|
+
Classic over-eager refactor by a capable model with **no feedback loop**:
|
|
87
|
+
- Wrote the code it *wished* existed (idealized `index.ts`) rather than the code that matches
|
|
88
|
+
the modules it actually wrote.
|
|
89
|
+
- Built parallel "clean-room" modules instead of integrating into the 25k-line monolith — the
|
|
90
|
+
hard part (wiring) was skipped, so 100% of the architectural value is unrealized.
|
|
91
|
+
- Never ran `tsc`. The `executable-feedback` / `anti-laziness` rules were not honored.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Part 2 — Duplicate failed tool calls & duplicate reads
|
|
96
|
+
|
|
97
|
+
### This is a pre-existing runtime issue in the working code, independent of Part 1.
|
|
98
|
+
|
|
99
|
+
The dedup architecture in the live `agenticRunner.ts` is large and thoughtful, but its design
|
|
100
|
+
has three structural gaps that together produce the observed loop.
|
|
101
|
+
|
|
102
|
+
### 2.1 What exists today
|
|
103
|
+
|
|
104
|
+
1. **`_dedupeToolCallsForResponse` (line 7397)** — drops exact-duplicate tool calls
|
|
105
|
+
**within a single response batch**, before execution. Good, but **intra-turn only**.
|
|
106
|
+
2. **`proactivePrune` (line 4937)** — post-hoc context hygiene. On each turn it walks history
|
|
107
|
+
and, for repeated fingerprints, **replaces the *earlier* result** with
|
|
108
|
+
`[deduped — same call as turn N]`. Also ages out `file_read` results older than 20 turns
|
|
109
|
+
(`AGED_FILE_READ_TURNS`) and successful shells older than 12.
|
|
110
|
+
3. **Failure-learning injectors** — `_recentFailures`, `_argCohorts`, `_errorPatterns`
|
|
111
|
+
(WO-NC-07 pre-action guidance), `_failureReflections` (REG-26 Reflexion), REG-32 opaque-error
|
|
112
|
+
hints. These *can* inject "you tried this and it failed" before re-dispatch.
|
|
113
|
+
|
|
114
|
+
### 2.2 Root causes (ranked)
|
|
115
|
+
|
|
116
|
+
**RC-1 — No cross-turn pre-execution dedup (primary).**
|
|
117
|
+
`_dedupeToolCallsForResponse` only looks within one response. A `file_read(X)` on turn 5 and an
|
|
118
|
+
identical `file_read(X)` on turn 25 **both execute**. The fingerprint
|
|
119
|
+
(`_buildToolFingerprint` = name + exact args, line 7389) is computed but never consulted as an
|
|
120
|
+
emission-time gate across turns. Duplicates are *cleaned up*, never *prevented*.
|
|
121
|
+
|
|
122
|
+
**RC-2 — Aging causes re-reads (acknowledged in-code).**
|
|
123
|
+
`proactivePrune` ages out `file_read` results after 20 turns. The REG-64 comment (line 4940)
|
|
124
|
+
admits: *"too-aggressive aging causes the model to re-read the same file because it forgot the
|
|
125
|
+
content, creating duplicate calls."* Bumping 10→20 mitigated but did not remove this — long runs
|
|
126
|
+
still strip content the model then re-fetches.
|
|
127
|
+
|
|
128
|
+
**RC-3 — Repeated *failures* have their evidence pruned (why failed calls loop).**
|
|
129
|
+
The fingerprint is **result-agnostic**: a failed call and its retry share a fingerprint, so
|
|
130
|
+
proactivePrune replaces the earlier *failure* with `[deduped — same call as turn N — duplicate
|
|
131
|
+
file_read() call]`. That stub **does not record that the prior call failed**. The model's inline
|
|
132
|
+
history therefore looks *cleaner* after each retry, removing the strongest natural signal
|
|
133
|
+
("I've failed this exact call twice") and making another retry *more* likely. The system leans
|
|
134
|
+
entirely on the separate `_recentFailures`/`_errorPatterns` injectors to re-surface the pattern —
|
|
135
|
+
and those are gated by per-turn one-shot flags (`_errorGuidanceInjected`,
|
|
136
|
+
`_reflectionsInjectedThisTurn`, `_opaqueErrorHintInjected`) and **stem matching**, which misses
|
|
137
|
+
when the model varies args slightly or pivots stems (the exact case REG-32 was added for).
|
|
138
|
+
|
|
139
|
+
### 2.3 Where to patch
|
|
140
|
+
|
|
141
|
+
| # | Patch | Location | Effort |
|
|
142
|
+
| --- | --- | --- | --- |
|
|
143
|
+
| **P-A** | **Cross-turn pre-execution dedup gate.** Maintain a run-level `Map<fingerprint, { turn, ok, resultRef }>`. Before dispatch, if a *successful* identical fingerprint exists, **short-circuit**: skip execution and inject `"already executed at turn N — see prior result"` instead of re-running. Reuse `_buildToolFingerprint`. This is the single highest-leverage fix. | new gate in the dispatch path alongside `_dedupeToolCallsForResponse` (`agenticRunner.ts:7397`) | M |
|
|
144
|
+
| **P-B** | **Make dedup result-aware.** When pruning a duplicate in `proactivePrune`, if the pruned call **failed**, preserve that in the stub: `[deduped — same FAILED call as turn N: <error 1-liner>]`. Keep the failure signal visible instead of erasing it. | `proactivePrune` dedupe branch (`agenticRunner.ts:5020-5031`) | S |
|
|
145
|
+
| **P-C** | **Escalate on Nth identical failure → hard stop, not soft nudge.** Wire `permissionRuleset.checkDoomLoop` (already written, P6) — after 3 identical fingerprints, switch from "inject guidance" to "refuse the call and force a different action or `task_complete`." This is OpenCode's `doom_loop: ask` pattern. | consume `_recentFailures`/`_argCohorts` at dispatch; the helper already exists in `permissionRuleset.ts:113` | M |
|
|
146
|
+
| **P-D** | **Don't age out a `file_read` unless it can be cheaply re-summarized in place.** Instead of clearing aged reads to a stub that triggers re-reads, replace with a **content digest** (path + line range + a few key symbols) so the model rarely needs the raw bytes again. | `proactivePrune` aged-file branch (`agenticRunner.ts:5096+`) | M |
|
|
147
|
+
| **P-E** | **Loosen one-shot injection gating for repeated failures.** The per-turn dedup flags suppress re-injection across turns when the model pivots stems; allow re-injection when `_argCohorts[fp].failure >= 2` regardless of stem. | injectors around `_errorPatterns` / REG-32 | S |
|
|
148
|
+
|
|
149
|
+
### 2.4 Note on the new modules
|
|
150
|
+
|
|
151
|
+
`permissionRuleset.checkDoomLoop` and `childSession`'s permission model are exactly the
|
|
152
|
+
primitives P-C wants — but they are **dead code**. The agent built the right tool for this fix and
|
|
153
|
+
then never connected it. Wiring `checkDoomLoop` into the live dispatch path (P-C) is the one piece
|
|
154
|
+
of the self-edit worth rescuing immediately.
|
|
155
|
+
|
|
156
|
+
---
|
|
157
|
+
|
|
158
|
+
## Recommended next actions
|
|
159
|
+
|
|
160
|
+
1. **Do not commit the working tree as-is** — it does not build and deletes exports the CLI needs.
|
|
161
|
+
2. **Salvage list** (cherry-pick into compiling changes): `agent-types.ts` `mode` field (restore
|
|
162
|
+
the removed registry functions + keep `plan`), `tool-batching.ts` Zod schemas, Ollama
|
|
163
|
+
`finish_reason` plumbing.
|
|
164
|
+
3. **Quarantine** the 6 new modules into a `proposals/` or feature branch — they're design sketches,
|
|
165
|
+
not integrations. Fix `compactionAgent` to actually call an LLM before anyone imports it.
|
|
166
|
+
4. **Ship the duplicate-call fix independently** of the P0–P10 refactor: P-A + P-B + P-C give the
|
|
167
|
+
biggest behavior win and don't depend on the broken modules.
|
|
168
|
+
5. **Add a build gate to the self-edit loop** — the agent must run `tsc -p packages/orchestrator`
|
|
169
|
+
and only mark `task_complete` on exit 0. This run would have caught everything in Part 1.
|
|
@@ -0,0 +1,290 @@
|
|
|
1
|
+
# OpenCode → Omnius: Agentic Loop & Sub-Agent Delegation Comparison
|
|
2
|
+
|
|
3
|
+
> Exhaustive comparison between `https://github.com/anomalyco/opencode/tree/dev` and
|
|
4
|
+
> `packages/orchestrator/src/agenticRunner.ts` (omnius).
|
|
5
|
+
> Generated 2026-06-10.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Main Agent Loop Structure
|
|
10
|
+
|
|
11
|
+
### OpenCode — `src/session/prompt.ts:1190-1390`
|
|
12
|
+
|
|
13
|
+
- **`while(true)`** loop with **natural exit** conditions — breaks when `lastAssistant.finish` is not `"tool-calls"` AND no pending tool calls
|
|
14
|
+
- Each iteration: `filterCompacted` → `latest` (find last user/assistant/finished/tasks) → resolve tools → call LLM → decide next step
|
|
15
|
+
- No fixed turn cap — bounded by agent `steps` config (default varies by agent)
|
|
16
|
+
- Result triage: `"stop"`, `"compact"`, or `"continue"` — compaction is an **inline loop control** decision
|
|
17
|
+
|
|
18
|
+
### Omnius — `orchestrator/src/agenticRunner.ts:9311`
|
|
19
|
+
|
|
20
|
+
- **`for (let turn = 0; turn < turnCap; turn++)`** with explicit turn cap (default 60)
|
|
21
|
+
- Much more complex per-turn preamble: REG-35 DoVer checkpoints, REG-46 world-state, REG-58/60/61 stagnation, REG-44 stuck, REG-50 write-thrash, REG-53 edit-fail-thrash, REG-18 stagnation window, etc.
|
|
22
|
+
- Exit only via `task_complete` tool with 3 identical handler implementations (streaming/unhandled/batch paths)
|
|
23
|
+
- No natural "model ran out of things to do" exit — model must explicitly call `task_complete`
|
|
24
|
+
|
|
25
|
+
### Learnings
|
|
26
|
+
|
|
27
|
+
- **Replace fixed `for` loop with natural-exit `while(true)`** — let the model exhaust its intent naturally. Keep turn cap as a safety fuse only (e.g., `maxTurns`). The 3× duplicated `task_complete` handler is a code-smell; unify into one path.
|
|
28
|
+
- **Adopt OpenCode's `finish` reason pattern** — use the LLM's `stop_reason` / `finish_reason` to detect natural completion instead of requiring an explicit `task_complete` tool call for every exit.
|
|
29
|
+
|
|
30
|
+
---
|
|
31
|
+
|
|
32
|
+
## 2. Sub-Agent Delegation
|
|
33
|
+
|
|
34
|
+
### OpenCode — `src/tool/task.ts`
|
|
35
|
+
|
|
36
|
+
- **`sessions.create({ parentID, title, agent, permission })`** — creates a **persisted child session** in the DB with full lifecycle
|
|
37
|
+
- **Permission derivation** (`src/agent/subagent-permissions.ts:1-46`): forwards parent agent `edit` denies, parent session denies, defaults-deny `todowrite`/`task` unless subagent explicitly permits
|
|
38
|
+
- **Background mode** (`task.ts:background.start()`): returns immediately, result injected as synthetic message
|
|
39
|
+
- **`task_id` resumption**: passing a prior `task_id` continues the same subagent session with accumulated context
|
|
40
|
+
- **Agent types are declarative** (`src/agent/agent.ts`): `mode: "subagent" | "primary" | "all"` — subagents can only be spawned via `task` tool, never selected by user
|
|
41
|
+
- **Concurrency hint in prompt** (`src/tool/task.txt`): "Launch multiple agents concurrently whenever possible"
|
|
42
|
+
|
|
43
|
+
### Omnius — `execution/src/tools/agent-tool.ts`
|
|
44
|
+
|
|
45
|
+
- **Three delegation mechanisms**:
|
|
46
|
+
1. `agent-tool.ts:329` — **worktree isolation**: spawns subprocess via `spawnSubprocess`
|
|
47
|
+
2. `agent-tool.ts:351` — **background in-process**: returns task ID, spawns in-process agent
|
|
48
|
+
3. `agent-tool.ts:390` — **foreground in-process**: blocks until complete
|
|
49
|
+
- **No persisted child session** — sub-agents share the parent session's message array
|
|
50
|
+
- **No permission derivation** — sub-agents get same permissions as parent (except optional `toolNames` filter)
|
|
51
|
+
- **No `task_id` resumption** — each spawn is fresh context
|
|
52
|
+
- **Also has legacy**: `full-sub-agent.ts` (separate OS process), `sub_agent` (routed through agent-tool.ts now)
|
|
53
|
+
- **Coordinator pattern** (`orchestrator/src/coordinator.ts:106`): `CoordinatorManager` enforces `maxConcurrentWorkers` (5) and `maxTotalWorkers` (20), but limited to coordinator mode
|
|
54
|
+
|
|
55
|
+
### Learnings
|
|
56
|
+
|
|
57
|
+
- **Implement child sessions with permission derivation** — this is the single biggest gap. Sub-agents currently unbounded — can do anything parent can. No state isolation.
|
|
58
|
+
- **Add `task_id` resumption** — deduplicate sub-agent invocations by allowing the model to continue a prior sub-agent's context. This avoids redundant exploration and enables iterative refinement.
|
|
59
|
+
- **Adopt declarative agent types with `mode`** — replace the programmatic `agent-types.ts` registry with a schema-driven system where `mode: "subagent"` agents can only be called via the `agent` tool (not activated directly by user), matching OpenCode's separation of concerns.
|
|
60
|
+
- **Permission forwarding**: parent agent's `edit`-deny rules (e.g., plan mode) must cascade to sub-agents automatically.
|
|
61
|
+
|
|
62
|
+
---
|
|
63
|
+
|
|
64
|
+
## 3. Tool Batching / Parallel Execution
|
|
65
|
+
|
|
66
|
+
### OpenCode — `src/session/processor.ts`
|
|
67
|
+
|
|
68
|
+
- **No explicit batching** — relies on LLM provider's native ability to emit multiple `tool-call` events in one stream
|
|
69
|
+
- **AI SDK `streamText()`** (`src/session/llm/ai-sdk.ts`) executes tools concurrently internally
|
|
70
|
+
- **Per-tool `Deferred<void>`** (`processor.ts:ensureToolCall`) — each tool call gets a deferred that resolves when result arrives
|
|
71
|
+
- **Cleanup**: `Effect.forEach(Object.values(ctx.toolcalls), ..., { concurrency: "unbounded" })` with 250ms timeout per call
|
|
72
|
+
- **Ordering**: events arrive in temporal order from the LLM stream; results are interleaved arbitrarily
|
|
73
|
+
|
|
74
|
+
### Omnius — `orchestrator/src/tool-batching.ts`
|
|
75
|
+
|
|
76
|
+
- **Explicit batching**: `partitionToolCalls()` groups concurrent-safe tools into parallel batches, serial tools each get single-item batches
|
|
77
|
+
- **`executeBatch()` with worker pool** (`tool-batching.ts:212-233`): `withConcurrencyLimit(fns, limit=8)` — N concurrent workers
|
|
78
|
+
- **Dual dispatch paths** (`agenticRunner.ts:15281-15705`):
|
|
79
|
+
- Streaming: `StreamingToolExecutor` with `queue → finalize → waitAll → drainCompleted` + ordering guarantees
|
|
80
|
+
- Non-streaming: `partitionToolCalls → executeBatch` with REG-24 fingerprint dedup
|
|
81
|
+
- **Streaming executor state machine** (`streaming-executor.ts:292`): `canExecute()` checks — concurrent-safe tools run in parallel, exclusive tools run alone (ordering: stops at first exclusive)
|
|
82
|
+
- **Duplicate detection** (`streaming-executor.ts:327-393`): `entryFingerprint()` / `findPriorEquivalent()` / `mirrorPriorEquivalent()` — identical tool calls within same stream share results
|
|
83
|
+
|
|
84
|
+
### Learnings
|
|
85
|
+
|
|
86
|
+
- **Simplify to single dispatch path** — the streaming/non-streaming bifurcation leads to code duplication (3 `task_complete` handlers, parallel dispatch logic). OpenCode's unified event-stream model avoids this entirely.
|
|
87
|
+
- **Adopt event-based tool tracking** — replace the `rawToolCalls` array + result-mapping with a per-call `Deferred`/`Promise` map. This eliminates the need for `drainCompleted()` ordering logic and the dual-path complexity.
|
|
88
|
+
- **OpenCode's simplicity is instructive** — it doesn't need `partitionToolCalls` or concurrency-safe classifications because the AI SDK handles it. Consider delegating concurrency to the backend layer.
|
|
89
|
+
|
|
90
|
+
---
|
|
91
|
+
|
|
92
|
+
## 4. Session / Runner State Management
|
|
93
|
+
|
|
94
|
+
### OpenCode — Three-Layer Architecture
|
|
95
|
+
|
|
96
|
+
1. **`Runner`** (`src/effect/runner.ts:1-220`): Per-session state machine with 4 states — `Idle`, `Running`, `Shell`, `ShellThenRun`. Run and shell are **mutually exclusive**. Shell has priority; runs queue as `ShellThenRun` and auto-start when shell finishes.
|
|
97
|
+
2. **`RunCoordinator`** (`packages/core/src/session/run-coordinator.ts:1-200`): Per-key drain coordination with coalescing. At most **1 active + 1 pending** per session. `run` (explicit) dominates `wake` (advisory). Interrupts suppress stale wakes.
|
|
98
|
+
3. **`SessionPrompt.runLoop`** (`src/session/prompt.ts`): High-level loop orchestration. `ensureRunning()` guarantees at most one loop iteration per session.
|
|
99
|
+
|
|
100
|
+
### Omnius — `orchestrator/src/agenticRunner.ts:1583-1600`
|
|
101
|
+
|
|
102
|
+
- **Single-class monolith**: `AgenticRunner` class with 200+ private properties, all in one file (~25,159 lines)
|
|
103
|
+
- **No state machine** for run/shell concurrency — the `run()` method is called directly
|
|
104
|
+
- **No drain coordination** — multiple calls to `run()` would stack, not coalesce
|
|
105
|
+
- **No explicit interruption model** — abort via `options.abortSignal` at `agenticRunner.ts:9330` (passthrough, not session-aware)
|
|
106
|
+
|
|
107
|
+
### Learnings
|
|
108
|
+
|
|
109
|
+
- **Decompose `AgenticRunner`** — split into layers: Runner (state machine), Orchestrator (loop control), Coordinator (drain + interruption). The monolith approach makes reasoning about concurrency impossible.
|
|
110
|
+
- **Add Run/Shell mutual exclusion** — when a shell command is running, agent loop should defer; when loop is running, shell should either fail or queue. OpenCode's `ShellThenRun` pattern is correct.
|
|
111
|
+
- **Add drain coalescing** — if the agent completes a turn and the session has new work, coalesce into one chain instead of stacking multiple `run()` calls.
|
|
112
|
+
|
|
113
|
+
---
|
|
114
|
+
|
|
115
|
+
## 5. Tool Definition & Registration
|
|
116
|
+
|
|
117
|
+
### OpenCode — `packages/core/src/tool/tool.ts`
|
|
118
|
+
|
|
119
|
+
- **`Tool.make(config)`**: Zod-schema-validated input/output, `toModelOutput` mapping, opaque branded `Definition<Input, Output>` type
|
|
120
|
+
- **`definition(name)` + `settle(call, context)`** stored in `WeakMap<AnyTool, Runtime>` — decoupled from AI SDK format
|
|
121
|
+
- **`withPermission(tool, name)`** — decorator pattern, produces a new tool object with an attached permission string
|
|
122
|
+
- **`ToolRegistry.materialize()`** (`packages/core/src/tool/registry.ts`): resolves full tool definitions for a given permission ruleset
|
|
123
|
+
- **`SessionTools.resolve()`** (`packages/opencode/src/session/tools.ts`): wraps tools as AI SDK `tool()` objects with permission checking, plugin hooks, truncation
|
|
124
|
+
|
|
125
|
+
### Omnius — `execution/src/index.ts`
|
|
126
|
+
|
|
127
|
+
- **`AgenticTool` interface**: `{ name, description, execute(args): Promise<string>, parameters?, isConcurrencySafe?, isReadOnly?, executeStream? }`
|
|
128
|
+
- **Flat catalog registration** — all tools exported from a single massive `index.ts` (~100+ tools)
|
|
129
|
+
- **No schema validation** — `execute` receives `Record<string, unknown>`, manual parsing inside each tool
|
|
130
|
+
- **No permission system** — tools are either available or not, no granular `allow`/`deny`/`ask` rules
|
|
131
|
+
- **No AI SDK compatibility layer** — tool definitions are constructed manually for each backend format
|
|
132
|
+
|
|
133
|
+
### Learnings
|
|
134
|
+
|
|
135
|
+
- **Add schema-validated tool definitions** — Zod schemas for input/output give free validation, documentation generation, and type safety. The current `Record<string, unknown>` pattern is error-prone.
|
|
136
|
+
- **Add a permission decorator system** — `withPermission` pattern allows reusable, composable permission rules on any tool without modifying the tool itself. This cascades naturally to sub-agent permissions.
|
|
137
|
+
- **Decouple tool definition from backend format** — like OpenCode's `WeakMap<AnyTool, Runtime>`, have a canonical internal format and adapters per backend (OpenAI, vLLM, Ollama).
|
|
138
|
+
|
|
139
|
+
---
|
|
140
|
+
|
|
141
|
+
## 6. Context Compaction / Overflow
|
|
142
|
+
|
|
143
|
+
### OpenCode — `src/session/prompt.ts` + `compaction` agent
|
|
144
|
+
|
|
145
|
+
- **Inline overflow detection**: `runLoop` checks `lastFinished` context size against budget; if exceeded, `compaction.create()` and `continue`
|
|
146
|
+
- **Dedicated compaction agent** (`src/agent/agent.ts`): `compaction` — hidden primary agent with all tools denied, runs `PROMPT_COMPACTION` via LLM
|
|
147
|
+
- **Message filtering**: `MessageV2.filterCompactedEffect(sessionID)` — compacted messages are filtered out before each loop iteration
|
|
148
|
+
- **Compact or break**: compaction returns `"stop"` or `"continue"` — if context is too tight even after compaction, the loop terminates naturally
|
|
149
|
+
|
|
150
|
+
### Omnius — `orchestrator/src/context-compressor.ts`
|
|
151
|
+
|
|
152
|
+
- **Two-phase compression**:
|
|
153
|
+
1. Phase 1 (cheap): prune old tool results (no LLM call)
|
|
154
|
+
2. Phase 2 (expensive): LLM-generated structured summaries with sections (Goal, Constraints, Progress, Key Decisions, Relevant Files, Next Steps, Critical Context)
|
|
155
|
+
- **`DefaultContextEngine`** (`orchestrator/src/contextEngine.ts`): simple budget-based pruning — drops oldest non-evidence messages
|
|
156
|
+
- **Context assembly** (`agenticRunner.ts:assembleContext()`): sections `c_instr, c_state, c_identity, c_know, c_retrieval, c_graph, c_plan, c_todos, c_lessons, c_workboard`
|
|
157
|
+
- **No compaction agent** — compression is performed in-process, not delegated to a dedicated agent call
|
|
158
|
+
|
|
159
|
+
### Learnings
|
|
160
|
+
|
|
161
|
+
- **Consider a dedicated compaction agent** — OpenCode's approach of delegating context compression to a separate LLM call with zero-tool agent avoids polluting the main agent's context with compression overhead and allows the compression to be a real summarization pass rather than a pruning pass.
|
|
162
|
+
- **Add post-compaction exit** — when compression has been applied but token budget is still exceeded, the loop should terminate gracefully (signal "context too tightly coupled to summarize") instead of entering a compaction loop.
|
|
163
|
+
- **Pre-compute budget before LLM call** — detect overflow before calling the LLM, not after. OpenCode does this in `runLoop` via the `overflow` task check.
|
|
164
|
+
|
|
165
|
+
---
|
|
166
|
+
|
|
167
|
+
## 7. Model Calling / Streaming
|
|
168
|
+
|
|
169
|
+
### OpenCode — `src/session/llm.ts`
|
|
170
|
+
|
|
171
|
+
- **Dual runtime architecture**: Native (`@opencode-ai/llm`) or AI SDK (`streamText`) fallback
|
|
172
|
+
- **Unified `LLMEvent` stream**: normalized event types — `text-start/delta/end`, `reasoning-start/delta/end`, `tool-input-start/delta/end`, `tool-call`, `tool-result`, `tool-error`, `step-start/finish`, `finish`
|
|
173
|
+
- **Tool call streaming**: `tool-input-start/delta/end` events let the processor track partial tool arguments before the call is finalized
|
|
174
|
+
- **`createStructuredOutputTool()`** (`prompt.ts`): injects a json_schema tool via `tool_choice: "required"` for structured outputs
|
|
175
|
+
|
|
176
|
+
### Omnius — `backend-vllm/src/VllmBackend.ts`, `OllamaBackend.ts`
|
|
177
|
+
|
|
178
|
+
- **Simple `chatCompletion({ messages, tools })`** — standard OpenAI-compatible API
|
|
179
|
+
- **Streaming path** (`agenticRunner.ts:23035-23392`): custom SSE accumulation via `streamingRequest()` — manually accumulates `tool_call_delta` chunks, repairs JSON, manages multi-tool ordering
|
|
180
|
+
- **No normalized event stream** — tool parsing is done after the entire response is received (non-streaming) or via ad-hoc accumulator logic (streaming)
|
|
181
|
+
- **No tool-call streaming** — partial tool arguments are not exposed to the executor (cannot start executing before args are fully received, unlike OpenCode's `tool-input-start`)
|
|
182
|
+
|
|
183
|
+
### Learnings
|
|
184
|
+
|
|
185
|
+
- **Normalize the event stream** — build an `LLMEvent` union type that all backends emit. This eliminates the streaming/non-streaming bifurcation and the two separate dispatch paths. Omnius already has `runEvents.ts` types but doesn't use them as a unified stream.
|
|
186
|
+
- **Adopt tool-input streaming** — expose partial tool argument events upstream so the streaming executor can start executing tools before all arguments are fully received (especially valuable for large writes/reads with many arguments).
|
|
187
|
+
- **Unify backend adapter interface** — currently each backend returns a different format. A normalized `Stream<LLMEvent>` return type would allow the entire dispatch pipeline to be backend-agnostic.
|
|
188
|
+
|
|
189
|
+
---
|
|
190
|
+
|
|
191
|
+
## 8. Adversary / Critic / Post-Turn Analysis
|
|
192
|
+
|
|
193
|
+
### OpenCode
|
|
194
|
+
|
|
195
|
+
- **No adversary system** — OpenCode does not have a dedicated post-turn meta-analysis layer. It relies on:
|
|
196
|
+
- **Permission `ask` prompts** for user-in-the-loop decisions (`doom_loop` detection at `processor.ts` — checks last 3 identical tool calls)
|
|
197
|
+
- **The model's own reasoning** to self-correct
|
|
198
|
+
- **Compaction agent** for context-level correction
|
|
199
|
+
- **Error handling**: `tool-error` → `failToolCall()` → marks part as `"error"` + sets `ctx.blocked` for permission rejections
|
|
200
|
+
|
|
201
|
+
### Omnius — `orchestrator/src/agenticRunner.ts:20395-20758`
|
|
202
|
+
|
|
203
|
+
- **Extensive adversary system** with 4 detections:
|
|
204
|
+
1. `adversaryObserve` false failure claim (line 20541)
|
|
205
|
+
2. False success claim (line 20595)
|
|
206
|
+
3. **Redundant action** — repeated same tool+args that already succeeded (line 20659)
|
|
207
|
+
4. **Idle think** — runaway output without input (line 20743)
|
|
208
|
+
- **Redundant action signal** (`_adversaryRedundantSignals`): forwarded to `Critic.evaluate()` in `executeSingle` (line 12735-12738)
|
|
209
|
+
- **Escalation**: dedup `_adversaryRecentFlags` with system-role injection after 3 repeats
|
|
210
|
+
- **Adversary mode**: `"backseat"` (events only), `"skillcoach"` (injects critiques), `"both"` (default)
|
|
211
|
+
- **Also**: REG-24 fingerprint dedup, REG-47 backward-pass critic, REG-11 semantic shell failure detection
|
|
212
|
+
|
|
213
|
+
### Learnings
|
|
214
|
+
|
|
215
|
+
- **Omnius's adversary system is genuinely innovative** — OpenCode has no equivalent. The redundant action detection, false-failure-correction, and depth-gauge integration are unique differentiators.
|
|
216
|
+
- **But complexity is high** — the adversary adds 500+ lines of stateful logic across multiple methods. Consider whether some adversarial patterns could be simplified to OpenCode's permission `ask` pattern (user-in-the-loop) instead of autonomous critique generation, at least for lower-confidence detections.
|
|
217
|
+
- **Unify with Critic** — the adversary `_adversaryRedundantSignals` is consumed by `Critic.evaluate()` in `executeSingle`. This cross-cutting coupling makes it hard to understand the full redundant-action flow. Consider making the adversary emit typed events that the critic subscribes to.
|
|
218
|
+
|
|
219
|
+
---
|
|
220
|
+
|
|
221
|
+
## 9. Agent Type System
|
|
222
|
+
|
|
223
|
+
### OpenCode — `src/agent/agent.ts`
|
|
224
|
+
|
|
225
|
+
- **Declarative schema**: `Info = { name, description, mode: "subagent"|"primary"|"all", permission: Ruleset[], model, prompt, steps, color, ... }`
|
|
226
|
+
- **Built-in registration inside `layer`**: build (primary), plan (primary), general (subagent), explore (subagent), compaction/title/summary (hidden primaries)
|
|
227
|
+
- **User-defined**: read from user config (`cfg.agent`), merged into `agents` map
|
|
228
|
+
- **`generate()`**: creates new agents via LLM `generateObject` using `PROMPT_GENERATE`
|
|
229
|
+
- **Permission rulesets per agent**: `explore` denies everything except grep/glob/list/bash/webfetch/websearch/read; `plan` denies all edit tools
|
|
230
|
+
- **Hidden agents**: `compaction`, `title`, `summary` — `native: true, hidden: true`, not shown in agent picker, used internally
|
|
231
|
+
|
|
232
|
+
### Omnius — `orchestrator/src/agent-types.ts`
|
|
233
|
+
|
|
234
|
+
- **Programmatic registry**: `AGENT_TYPES = Record<string, AgentType>` with `registerAgentType(name, config)`
|
|
235
|
+
- **Built-in**: `general` (full access), `explore` (read-only), `plan` (read-only + planning), `coordinator` (orchestration + spawning)
|
|
236
|
+
- **Agent shape**: `{ allowedTools, disallowedTools, maxTurns, model, canSpawnAgents, description, systemPromptAddition }`
|
|
237
|
+
- **No `mode` field** — agents aren't classified as primary/subagent. Any agent can be activated by the user.
|
|
238
|
+
- **No user-defined agents** — no config-based agent creation
|
|
239
|
+
- **No agent generation** — no LLM-based agent creation
|
|
240
|
+
|
|
241
|
+
### Learnings
|
|
242
|
+
|
|
243
|
+
- **Add `mode` to agent types** — explicitly declare which agents are `"primary"` (user-selectable) vs `"subagent"` (task-tool-only) vs `"all"`. This enforces the architectural constraint that subagents can only be spawned via the `agent` tool.
|
|
244
|
+
- **Make agent definitions declarative/configurable** — allow users to define custom agents in config files. The `generate()` pattern (LLM creates an agent definition) is an interesting power feature but secondary.
|
|
245
|
+
- **Use permission rulesets per agent type** — instead of `allowedTools`/`disallowedTools` arrays, adopt OpenCode's `PermissionV1.Ruleset` pattern. This enables granular `allow`/`deny`/`ask` per tool/permission and natural subagent inheritance.
|
|
246
|
+
|
|
247
|
+
---
|
|
248
|
+
|
|
249
|
+
## 10. Permissions & Safety
|
|
250
|
+
|
|
251
|
+
### OpenCode — `packages/core/src/session/permission.ts`
|
|
252
|
+
|
|
253
|
+
- **Granular rulesets**: `PermissionV1.Ruleset = Array<{ permission: string, pattern: string, action: "allow" | "deny" | "ask" }>`
|
|
254
|
+
- **`ask` prompts**: user-in-the-loop permission checks (e.g., `doom_loop: ask` — after 3 identical tool calls, ask user if they want to allow/deny)
|
|
255
|
+
- **Parent→child forwarding** (`src/agent/subagent-permissions.ts`): parent agent denies + parent session denies + `external_directory` rules are forwarded to subagent sessions
|
|
256
|
+
- **Default deny for subagents**: `todowrite` and `task` are denied in subagents unless explicitly permitted
|
|
257
|
+
|
|
258
|
+
### Omnius — `orchestrator/src/agenticRunner.ts` + tool implementations
|
|
259
|
+
|
|
260
|
+
- **Constraint system** (`execution/src/workingNotes.ts`, `constraints`):
|
|
261
|
+
- `formatSecurityNotice()`, `checkConstraints()`, `formatViolationWarning()`
|
|
262
|
+
- Constraints checked in `executeSingle` before tool dispatch
|
|
263
|
+
- **Tool allow/disallow lists**: per agent type, simple string arrays
|
|
264
|
+
- **No granular permission rules** — no `ask` prompts, no `deny`/`allow`/`ask` ternary
|
|
265
|
+
- **No subagent permission derivation** — sub-agents (via `agent` tool) inherit parent's full permission set
|
|
266
|
+
- **No doom-loop detection** — the adversary has `idle_think` (consecutive short outputs) but no permission-based intervention
|
|
267
|
+
|
|
268
|
+
### Learnings
|
|
269
|
+
|
|
270
|
+
- **Implement granular permission rulesets** — replace flat allow/disallow arrays with structured `{ permission, pattern, action }` rules. This enables: (a) `ask` for sensitive operations, (b) pattern-based denials (e.g., `edit:*.env` → deny), (c) forwarding to subagents, (d) user-configurable permission overrides.
|
|
271
|
+
- **Add doom-loop detection with `ask`** — when the model repeats the same tool with same args 3+ times, prompt the user instead of silently injecting an adversary critique. The user can decide allow/deny/interrupt.
|
|
272
|
+
- **Constraint system is good, but integrate with permissions** — the `checkConstraints()` system is powerful but operates independently of the agent type system. Constraints should be part of the permission ruleset, not a separate pre-hook.
|
|
273
|
+
|
|
274
|
+
---
|
|
275
|
+
|
|
276
|
+
## Summary: Highest-Impact Improvements (Stack Ranked)
|
|
277
|
+
|
|
278
|
+
| Priority | Improvement | Why |
|
|
279
|
+
| -------- | ----------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
|
280
|
+
| **P0** | Child sessions with permission derivation | Sub-agents currently unbounded — can do anything parent can. No state isolation. |
|
|
281
|
+
| **P1** | Declarative agent types with `mode` | Enables architectural enforcement (primary vs subagent), user-configurable agents, LLM-generated agents |
|
|
282
|
+
| **P2** | Unified event-stream dispatch | Eliminates streaming/non-streaming bifurcation + 3× duplicated `task_complete` handlers |
|
|
283
|
+
| **P3** | Natural exit condition | Replace `task_complete`-required exit with LLM `stop_reason` detection; keep `task_complete` as optional |
|
|
284
|
+
| **P4** | Runner state machine | Run/Shell mutual exclusion, drain coalescing, interruption model — prevents subtle concurrency bugs |
|
|
285
|
+
| **P5** | Tool definition schema validation | Replace `Record<string, unknown>` with Zod schemas for free validation + documented schemas |
|
|
286
|
+
| **P6** | Granular permission rulesets | `allow`/`deny`/`ask` ternary — cascade to subagents, user-in-the-loop for sensitive ops |
|
|
287
|
+
| **P7** | Decompose AgenticRunner | Split into Runner + Orchestrator + Coordinator layers (currently 25K-line monolith) |
|
|
288
|
+
| **P8** | Unify backend adapter interface | Normalized `LLMEvent` stream type eliminates per-backend dispatch differences |
|
|
289
|
+
| **P9** | Dedicated compaction agent | Offload context compression to a zero-tool agent to avoid pollution in main context |
|
|
290
|
+
| **P10** | `task_id` resumption for sub-agents | Allow continuing prior sub-agent sessions (dedup exploration, iterative refinement) |
|
|
@@ -14,7 +14,7 @@ Bind to a network interface only when you also configure auth:
|
|
|
14
14
|
|
|
15
15
|
```bash
|
|
16
16
|
OMNIUS_HOST=0.0.0.0:11435 \
|
|
17
|
-
|
|
17
|
+
OMNIUS_REST_API_KEYS="read-key:read:grafana,run-key:run:ci:60:100000:3,admin-key:admin:ops" \
|
|
18
18
|
omnius serve
|
|
19
19
|
```
|
|
20
20
|
|
|
@@ -59,7 +59,7 @@ For a LAN daemon, prefer a scoped key set:
|
|
|
59
59
|
|
|
60
60
|
```bash
|
|
61
61
|
OMNIUS_HOST=0.0.0.0:11435 \
|
|
62
|
-
|
|
62
|
+
OMNIUS_REST_API_KEYS="dash:read:grafana:600::,ci:run:ci:60:100000:3,ops:admin:ops:120:500000:10" \
|
|
63
63
|
omnius serve
|
|
64
64
|
```
|
|
65
65
|
|
|
@@ -0,0 +1,63 @@
|
|
|
1
|
+
# Service Version Compatibility
|
|
2
|
+
|
|
3
|
+
A client may be newer than the long-running Omnius daemon it reaches. The
|
|
4
|
+
runtime version gate prevents a stale service from quietly accepting work
|
|
5
|
+
whose provider, tool, or request contract it does not understand.
|
|
6
|
+
|
|
7
|
+
## Inspect The Service
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
curl -s http://127.0.0.1:11435/version
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
The response contains `package_version`, `api_version`, and
|
|
14
|
+
`discovery_schema_version`. Compatibility fields `version`, `boot_version`,
|
|
15
|
+
`boot_package_hash`, `node`, and `platform` remain for existing clients.
|
|
16
|
+
|
|
17
|
+
Inspection endpoints remain available even when the service is too old to run
|
|
18
|
+
a new client request: health, version, help, OpenAPI, and discovery.
|
|
19
|
+
|
|
20
|
+
## Require A Minimum Version
|
|
21
|
+
|
|
22
|
+
Send a SemVer minimum on execution requests:
|
|
23
|
+
|
|
24
|
+
```text
|
|
25
|
+
X-Omnius-Min-Version: <minimum-compatible-package-version>
|
|
26
|
+
```
|
|
27
|
+
|
|
28
|
+
The gate applies to:
|
|
29
|
+
|
|
30
|
+
- `POST /v1/run`;
|
|
31
|
+
- `POST /v1/chat`, `/api/chat`, and `/v1/chat/completions`;
|
|
32
|
+
- `POST /v1/generate` and `/api/generate`;
|
|
33
|
+
- `POST /v1/tools/{name}/call`;
|
|
34
|
+
- `POST /v1/commands/{cmd}`.
|
|
35
|
+
|
|
36
|
+
The service checks the precondition before the request handler creates state,
|
|
37
|
+
loads a model, calls a provider, or invokes a command/tool.
|
|
38
|
+
|
|
39
|
+
Responses:
|
|
40
|
+
|
|
41
|
+
| Condition | Result |
|
|
42
|
+
| --- | --- |
|
|
43
|
+
| header omitted | normal backward-compatible behavior |
|
|
44
|
+
| valid minimum, service satisfies it | request proceeds |
|
|
45
|
+
| invalid SemVer | RFC 7807 `400 Bad Request` |
|
|
46
|
+
| service older than minimum | RFC 7807 `412 Precondition Failed` |
|
|
47
|
+
|
|
48
|
+
A `412` is not retryable against the same daemon. Update or reconnect to a
|
|
49
|
+
verified compatible Omnius service, inspect `/version` again, and resubmit
|
|
50
|
+
only after the precondition can succeed.
|
|
51
|
+
|
|
52
|
+
## Why Both Check And Header
|
|
53
|
+
|
|
54
|
+
`GET /version` gives a useful compatibility diagnostic. The header is the
|
|
55
|
+
execution-time guard and closes the race between checking a service and
|
|
56
|
+
submitting work after that service has been replaced or routed elsewhere.
|
|
57
|
+
|
|
58
|
+
## Client Rule
|
|
59
|
+
|
|
60
|
+
Keep the required version in the adapter that constructs Omnius execution
|
|
61
|
+
requests. Do not rely on a human-readable npm version, cached install
|
|
62
|
+
metadata, or the client's own package version. Compare against the daemon's
|
|
63
|
+
reported package version and preserve the execution precondition.
|