omnius 1.0.591 → 1.0.592
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.aiwg/addons/omnius-docs/README.md +15 -1
- package/.aiwg/addons/omnius-docs/manifest.json +28 -68
- package/.aiwg/addons/omnius-docs/skills/agent-failure-recovery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/browser-interaction-validation/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/evidence-directed-delivery/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/hardware-evidence-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/omnius-docs/SKILL.md +17 -7
- package/.aiwg/addons/omnius-docs/skills/omnius-inference-docs/SKILL.md +27 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-integration-docs/SKILL.md +21 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-ops-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-realtime-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-sponsor-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-telegram-docs/SKILL.md +2 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-tools-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/omnius-version-compatibility-docs/SKILL.md +23 -0
- package/.aiwg/addons/omnius-docs/skills/runtime-provenance-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/secrets-and-config-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/test-surface-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-docs/skills/workspace-reality-audit/SKILL.md +2 -1
- package/.aiwg/addons/omnius-rest-docs/README.md +3 -0
- package/.aiwg/addons/omnius-rest-docs/manifest.json +27 -20
- package/.aiwg/addons/omnius-rest-docs/skills/omnius-rest-docs/SKILL.md +9 -5
- package/README.md +36 -0
- package/dist/discovery.d.ts +50 -0
- package/dist/index.js +5975 -4021
- package/dist/library.d.ts +7 -0
- package/dist/library.js +950 -0
- package/dist/postinstall-daemon.cjs +18 -0
- package/dist/providerRegistry.d.ts +80 -0
- package/dist/service-version.d.ts +35 -0
- package/docs/.vitepress/config.mts +8 -0
- package/docs/DISCOVERY.json +20224 -0
- package/docs/DISCOVERY.md +648 -0
- package/docs/HANDOFF-crl-encoder-decoder-fix.md +129 -0
- package/docs/agent-memory/INDEX.md +9 -4
- package/docs/agent-memory/index.md +7 -0
- package/docs/concept-relational-language.md +869 -0
- package/docs/context-management-medium-models-proposal.md +449 -0
- package/docs/dedup-false-positive-meta-analysis.md +96 -0
- package/docs/discovery/catalog-overrides.json +724 -0
- package/docs/duplicate-calls-root-cause-analysis.md +91 -0
- package/docs/duplicate-calls-root-cause-deep.md +155 -0
- package/docs/ephemeral-skill-pack-small-context.md +57 -0
- package/docs/explorations/context-window-todo-association.md +156 -0
- package/docs/explorations/todo-association-verify.json +30 -0
- package/docs/explorations/verification-ledger.json +45 -0
- package/docs/explorations/verify-todo-association.sh +30 -0
- package/docs/flowstate.md +806 -0
- package/docs/getting-started/install.md +24 -0
- package/docs/getting-started/model-providers.md +13 -0
- package/docs/guides/agent-integration.md +87 -0
- package/docs/guides/bring-your-own-inference.md +126 -0
- package/docs/guides/tools-and-web-search.md +95 -0
- package/docs/index.md +14 -0
- package/docs/longhaul-35b-workorders.md +496 -0
- package/docs/memory-integration-analysis.md +303 -0
- package/docs/model-capability-awareness-and-multimodal-memory-root-fix.md +799 -0
- package/docs/multimodal-identity-memory-implementation.md +76 -0
- package/docs/omnius-self-edit-eval-2026-06-10.md +169 -0
- package/docs/opencode-agentic-loop-comparison.md +290 -0
- package/docs/operations/security-and-remote-access.md +2 -2
- package/docs/operations/version-compatibility.md +63 -0
- package/docs/proposals/git-progress-tracking-strategy.md +289 -0
- package/docs/proposals/opencode-modules/backendAdapter.ts +443 -0
- package/docs/proposals/opencode-modules/childSession.ts +288 -0
- package/docs/proposals/opencode-modules/compactionAgent.ts +101 -0
- package/docs/proposals/opencode-modules/orchestrator.ts +387 -0
- package/docs/proposals/opencode-modules/runner.ts +258 -0
- package/docs/reference/auth-map.md +87 -196
- package/docs/reference/configuration.md +27 -0
- package/docs/reference/rest-api.md +7 -0
- package/docs/reference/slash-commands.md +125 -2
- package/docs/research/_archived/README.md +18 -0
- package/docs/research/_archived/context_window_attention_model.py +418 -0
- package/docs/research/_archived/context_window_attention_spec.md +55 -0
- package/docs/research/_archived/context_window_attention_weights.json +68 -0
- package/docs/research/k-splanifolds.pdf +0 -0
- package/docs/research/personality-verbosity-control.md +293 -0
- package/docs/rest/INDEX.md +7 -0
- package/docs/rest/QUICKREF.md +18 -0
- package/docs/rest/REST-DOCS-MANIFEST.json +1 -0
- package/docs/rest/auth-and-scopes.md +7 -1
- package/docs/rest/endpoints/discovery.md +44 -0
- package/docs/rest/endpoints/events.md +5 -0
- package/docs/rest/endpoints/tools.md +9 -0
- package/docs/reviews/adversary-system-review.md +42 -0
- package/docs/sana-and-video-generation-integration-plan.md +712 -0
- package/docs/session-diary-llm-training-analysis.md +218 -0
- package/docs/telegram-dmn-curiosity-outreach-scaffold.md +91 -0
- package/docs/telegram-mid-horizon-download-loop-handoff.md +468 -0
- package/docs/telegram-reflection-corpus-integration-plan.md +306 -0
- package/docs/telegram-unified-tooling-architecture.md +332 -0
- package/docs/threat-model.md +868 -0
- package/docs/trajectory-grounding.md +160 -0
- package/docs/voice-flow-architecture.md +489 -0
- package/docs/work-orders/WO-AM-GAPS.md +638 -0
- package/docs/work-orders/daemon-hud-ui-overhaul.md +82 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/INDEX.md +21 -0
- package/docs/work-orders/hermes-architecture-deltas/01-public-scrutiny-provenance-control/WORKORDER.md +225 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/02-context-engine-plugin-boundary/WORKORDER.md +198 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/03-typed-gateway-event-stream/WORKORDER.md +172 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/INDEX.md +19 -0
- package/docs/work-orders/hermes-architecture-deltas/04-task-local-gateway-context/WORKORDER.md +169 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/05-process-lifecycle-monitoring-notifications/WORKORDER.md +189 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/06-vision-evidence-routing-ladder/WORKORDER.md +199 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/INDEX.md +20 -0
- package/docs/work-orders/hermes-architecture-deltas/07-durable-multi-agent-kanban/WORKORDER.md +174 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/INDEX.md +22 -0
- package/docs/work-orders/hermes-architecture-deltas/08-completion-critic-reconciliation-ledger/WORKORDER.md +226 -0
- package/docs/work-orders/hermes-architecture-deltas/INDEX.md +38 -0
- package/docs/work-orders/omnius-context-engineering-behavior-fixes.md +281 -0
- package/docs/work-orders/telegram-dropbear-context-rca-workorder.md +202 -0
- package/docs/work-orders/world-class-memory-compiler/README.md +162 -0
- package/docs/work-orders/world-class-memory-compiler/TRACKER.md +179 -0
- package/docs/work-orders/world-class-memory-compiler/WO-01-exact-request-budget.md +79 -0
- package/docs/work-orders/world-class-memory-compiler/WO-02-typed-memory-fabric.md +65 -0
- package/docs/work-orders/world-class-memory-compiler/WO-03-dependency-working-set.md +55 -0
- package/docs/work-orders/world-class-memory-compiler/WO-04-inference-memory-compiler.md +67 -0
- package/docs/work-orders/world-class-memory-compiler/WO-05-artifact-fidelity-materialization.md +72 -0
- package/docs/work-orders/world-class-memory-compiler/WO-06-temporal-hybrid-retrieval.md +49 -0
- package/docs/work-orders/world-class-memory-compiler/WO-07-evaluation-harness.md +45 -0
- package/docs/work-orders/world-class-memory-compiler/WO-08-rollout-legacy-removal.md +45 -0
- package/docs/x402-remote-inference-plan.md +323 -0
- package/npm-shrinkwrap.json +108 -117
- package/package.json +7 -6
- package/templates/AGENTS.md +6 -0
- package/templates/OMNIUS.md +20 -0
|
@@ -0,0 +1,449 @@
|
|
|
1
|
+
# Context Management for Medium Models (30-40B) — Implementation Proposal
|
|
2
|
+
|
|
3
|
+
**Date:** 2026-04-25
|
|
4
|
+
**Problem:** Medium-tier models (~35B params) exhibit excessive repetition at ~35% context fill despite 256K context windows
|
|
5
|
+
**Root Cause:** Monolithic context structure + attention degradation + tool schema bloat
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Literature Summary
|
|
10
|
+
|
|
11
|
+
### Key Findings
|
|
12
|
+
|
|
13
|
+
| Paper | Key Insight | Application |
|
|
14
|
+
|-------|-------------|-------------|
|
|
15
|
+
| **Lost in the Middle** (Liu et al., 2024) | U-shaped attention: strong at start/end, 40% accuracy drop in middle 50% | Position critical info at edges |
|
|
16
|
+
| **RECOMP** (ICLR 2024) | Context compressed to 6% with minimal quality loss via observation masking | Aggressive tool output masking |
|
|
17
|
+
| **AgentFold** (arXiv:2510.24699) | Multi-scale folding prevents exponential fact decay (0.99^100 = 36.6%) | Progressive summarization with locked blocks |
|
|
18
|
+
| **ARC** (arXiv:2601.12030) | Active revision + reflection = 11% accuracy gain | Structural preservation through compaction |
|
|
19
|
+
| **NATURAL PLAN** (arXiv:2406.04520) | GPT-4 only 31% on planning even with full context | Efficient use > naive expansion |
|
|
20
|
+
| **ToolLLM DFSDT** (arXiv:2307.16789) | Error preservation + backtracking = +35pp success | Error-preserving compaction strategy |
|
|
21
|
+
| **SPRINT** (arXiv:2506.05745) | Parallel sub-calls distribute reasoning | Sub-agent delegation for independent tasks |
|
|
22
|
+
| **Recursive LMs** (arXiv:2512.24601) | Externalize to REPL, recursive chunk analysis | RLM context OS layer |
|
|
23
|
+
|
|
24
|
+
### Medium Model Specific Issues
|
|
25
|
+
|
|
26
|
+
1. **Earlier attention degradation** - Starts at 35% vs 50% for large models
|
|
27
|
+
2. **Tool schema bloat** - 64+ tools = ~15K tokens (6% of 256K)
|
|
28
|
+
3. **Repetition loops** - Model re-reads files it already has in context
|
|
29
|
+
4. **Completed task persistence** - Finished todos still occupy context
|
|
30
|
+
|
|
31
|
+
---
|
|
32
|
+
|
|
33
|
+
## 2. Current Omnius Implementation Analysis
|
|
34
|
+
|
|
35
|
+
### What We Have (Strengths)
|
|
36
|
+
|
|
37
|
+
```
|
|
38
|
+
┌─────────────────────────────────────────────────────────────────┐
|
|
39
|
+
│ CURRENT CONTEXT ARCHITECTURE │
|
|
40
|
+
├─────────────────────────────────────────────────────────────────┤
|
|
41
|
+
│ │
|
|
42
|
+
│ [System Prompt] ─┬─> [Head: preserved] │
|
|
43
|
+
│ │ │
|
|
44
|
+
│ [User Task] ─────┘ │
|
|
45
|
+
│ │
|
|
46
|
+
│ [Middle Messages] ──> Compaction ──> [Summary Block] │
|
|
47
|
+
│ │ │ │
|
|
48
|
+
│ │ ┌─────┴─────┐ │
|
|
49
|
+
│ │ │ Strategies│ │
|
|
50
|
+
│ │ │ - default │ │
|
|
51
|
+
│ │ │ - aggressive │
|
|
52
|
+
│ │ │ - decisions │
|
|
53
|
+
│ │ │ - errors │ │
|
|
54
|
+
│ │ │ - summary │ │
|
|
55
|
+
│ │ │ - structured │
|
|
56
|
+
│ │ └───────────┘ │
|
|
57
|
+
│ │ │ │
|
|
58
|
+
│ └────────────────────┴──> [Memex Archive] │
|
|
59
|
+
│ │
|
|
60
|
+
│ [Recent Messages] ──> Preserved verbatim (4-12 msgs) │
|
|
61
|
+
│ │
|
|
62
|
+
└─────────────────────────────────────────────────────────────────┘
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
**Compaction Thresholds (tier-aware):**
|
|
66
|
+
- Small (≤7B): 65% of context window
|
|
67
|
+
- Medium (8-29B): 70% of context window
|
|
68
|
+
- Large (≥30B): 75% of context window
|
|
69
|
+
- Deep mode: 85% of context window
|
|
70
|
+
|
|
71
|
+
**Features:**
|
|
72
|
+
- ✅ Progressive summarization (AgentFold-inspired)
|
|
73
|
+
- ✅ Observation masking (RECOMP-inspired)
|
|
74
|
+
- ✅ Goal re-injection after compaction
|
|
75
|
+
- ✅ Anti-repetition reminders for small/medium
|
|
76
|
+
- ✅ Tool calling reminders for small models
|
|
77
|
+
- ✅ Memex archive for large tool outputs
|
|
78
|
+
- ✅ SNR (Signal-to-Noise Ratio) tracking
|
|
79
|
+
- ✅ Task state preservation through compaction
|
|
80
|
+
- ✅ File registry with access counts
|
|
81
|
+
- ✅ Sub-agent types with tool restrictions
|
|
82
|
+
|
|
83
|
+
### What We Lack (Gaps)
|
|
84
|
+
|
|
85
|
+
```
|
|
86
|
+
┌─────────────────────────────────────────────────────────────────┐
|
|
87
|
+
│ IDENTIFIED GAPS │
|
|
88
|
+
├─────────────────────────────────────────────────────────────────┤
|
|
89
|
+
│ │
|
|
90
|
+
│ 1. MONOLITHIC CONTEXT │
|
|
91
|
+
│ └─> Flat message array, no hierarchical structure │
|
|
92
|
+
│ └─> All context equally weighted │
|
|
93
|
+
│ │
|
|
94
|
+
│ 2. NO TASK-PHASE AWARENESS │
|
|
95
|
+
│ └─> Completed todos still in context │
|
|
96
|
+
│ └─> No dynamic expansion/contraction based on phase │
|
|
97
|
+
│ │
|
|
98
|
+
│ 3. TOOL SCHEMA BLOAT │
|
|
99
|
+
│ └─> All 64+ tools loaded even when not needed │
|
|
100
|
+
│ └─> ~15K tokens for tool schemas │
|
|
101
|
+
│ │
|
|
102
|
+
│ 4. SUB-AGENT CONTEXT LEAKAGE │
|
|
103
|
+
│ └─> Sub-agents inherit parent context unnecessarily │
|
|
104
|
+
│ └─> No clean context boundary for delegation │
|
|
105
|
+
│ │
|
|
106
|
+
│ 5. REACTIVE ONLY │
|
|
107
|
+
│ └─> Compaction only at threshold │
|
|
108
|
+
│ └─> No proactive pruning of completed work │
|
|
109
|
+
│ │
|
|
110
|
+
│ 6. NO ANCHOR SURFACING │
|
|
111
|
+
│ └─> Previous task context always present │
|
|
112
|
+
│ └─> Should surface only when needed │
|
|
113
|
+
│ │
|
|
114
|
+
└─────────────────────────────────────────────────────────────────┘
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
---
|
|
118
|
+
|
|
119
|
+
## 3. Proposed Implementation Roadmap
|
|
120
|
+
|
|
121
|
+
### Phase 1: Task-Phase-Aware Context Tree (PRIORITY: HIGH)
|
|
122
|
+
|
|
123
|
+
**Problem:** Context is monolithic; completed work persists unnecessarily
|
|
124
|
+
|
|
125
|
+
**Solution:** Hierarchical context tree with phase-based expansion/contraction
|
|
126
|
+
|
|
127
|
+
```typescript
|
|
128
|
+
interface ContextTree {
|
|
129
|
+
// Root: always present
|
|
130
|
+
root: {
|
|
131
|
+
systemPrompt: string; // Condensed for medium models
|
|
132
|
+
activeGoal: string; // Current task
|
|
133
|
+
toolSubset: string[]; // Only relevant tools
|
|
134
|
+
};
|
|
135
|
+
|
|
136
|
+
// Phase nodes: expand when active, contract when complete
|
|
137
|
+
phases: {
|
|
138
|
+
explore: ContextNode; // Active during exploration
|
|
139
|
+
plan: ContextNode; // Active during planning
|
|
140
|
+
implement: ContextNode; // Active during implementation
|
|
141
|
+
verify: ContextNode; // Active during verification
|
|
142
|
+
};
|
|
143
|
+
|
|
144
|
+
// Archive: accessible via memex_retrieve
|
|
145
|
+
archive: {
|
|
146
|
+
completedPhases: string[]; // Hash IDs
|
|
147
|
+
priorTasks: string[]; // Hash IDs
|
|
148
|
+
};
|
|
149
|
+
}
|
|
150
|
+
|
|
151
|
+
interface ContextNode {
|
|
152
|
+
status: "active" | "contracted" | "archived";
|
|
153
|
+
messages: ChatMessage[];
|
|
154
|
+
summary?: string; // Generated when contracted
|
|
155
|
+
anchors: string[]; // Key facts to preserve
|
|
156
|
+
}
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
**Implementation:**
|
|
160
|
+
|
|
161
|
+
1. **Phase detection** - Analyze recent messages to determine current phase
|
|
162
|
+
2. **Automatic contraction** - When phase completes, summarize + archive
|
|
163
|
+
3. **Anchor extraction** - Preserve only critical facts from completed phases
|
|
164
|
+
4. **Dynamic expansion** - When returning to a phase, restore from archive
|
|
165
|
+
|
|
166
|
+
**Code locations:**
|
|
167
|
+
- `packages/orchestrator/src/agenticRunner.ts:5357` (compactMessages)
|
|
168
|
+
- `packages/orchestrator/src/agenticRunner.ts:1143` (contextLimits)
|
|
169
|
+
|
|
170
|
+
### Phase 2: Tool Schema Lazy Loading (PRIORITY: HIGH)
|
|
171
|
+
|
|
172
|
+
**Problem:** 64+ tools = ~15K tokens loaded at start
|
|
173
|
+
|
|
174
|
+
**Solution:** Tier-aware tool subsets with on-demand expansion
|
|
175
|
+
|
|
176
|
+
```typescript
|
|
177
|
+
// Medium model tool subsets
|
|
178
|
+
const MEDIUM_CORE_TOOLS = [
|
|
179
|
+
"file_read", "file_write", "file_edit",
|
|
180
|
+
"shell", "grep_search", "find_files",
|
|
181
|
+
"task_complete", "memory_read", "memory_write"
|
|
182
|
+
];
|
|
183
|
+
|
|
184
|
+
const MEDIUM_EXPANDABLE_TOOLS = {
|
|
185
|
+
web: ["web_search", "web_fetch", "web_crawl"],
|
|
186
|
+
code: ["file_patch", "file_explore", "batch_edit"],
|
|
187
|
+
agent: ["agent", "sub_agent", "background_run"],
|
|
188
|
+
memory: ["memory_search", "working_notes"],
|
|
189
|
+
// ... more subsets
|
|
190
|
+
};
|
|
191
|
+
|
|
192
|
+
// On-demand loading via explore_tools
|
|
193
|
+
function expandToolSubset(subset: string): void {
|
|
194
|
+
// Add tools to active set
|
|
195
|
+
// Update tool schema in context
|
|
196
|
+
}
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
**Implementation:**
|
|
200
|
+
|
|
201
|
+
1. **Core tool set** - Only 9-12 tools loaded initially for medium models
|
|
202
|
+
2. **Subset expansion** - `explore_tools("web")` loads web tools
|
|
203
|
+
3. **Schema caching** - Tool schemas cached, not re-sent
|
|
204
|
+
4. **Usage tracking** - Auto-unload unused tools after N turns
|
|
205
|
+
|
|
206
|
+
**Code locations:**
|
|
207
|
+
- `packages/execution/src/index.ts` (tool exports)
|
|
208
|
+
- `packages/orchestrator/src/agenticRunner.ts:870` (tool schema injection)
|
|
209
|
+
|
|
210
|
+
### Phase 3: Completed Todo Pruning (PRIORITY: MEDIUM)
|
|
211
|
+
|
|
212
|
+
**Problem:** Completed todos persist in context, wasting tokens
|
|
213
|
+
|
|
214
|
+
**Solution:** Automatic pruning with anchor extraction
|
|
215
|
+
|
|
216
|
+
```typescript
|
|
217
|
+
interface TodoState {
|
|
218
|
+
active: Todo[]; // In context
|
|
219
|
+
completed: TodoAnchor[]; // Summarized
|
|
220
|
+
archived: string[]; // Hash IDs in Memex
|
|
221
|
+
}
|
|
222
|
+
|
|
223
|
+
interface TodoAnchor {
|
|
224
|
+
id: string;
|
|
225
|
+
summary: string; // One-line summary
|
|
226
|
+
keyFiles: string[]; // Files touched
|
|
227
|
+
outcome: "success" | "blocked" | "delegated";
|
|
228
|
+
}
|
|
229
|
+
|
|
230
|
+
// Pruning logic
|
|
231
|
+
function pruneCompletedTodos(): void {
|
|
232
|
+
const completed = todos.filter(t => t.status === "completed");
|
|
233
|
+
const anchors = completed.map(t => extractAnchor(t));
|
|
234
|
+
|
|
235
|
+
// Add anchors to context (compact)
|
|
236
|
+
// Archive full todos to Memex
|
|
237
|
+
// Update todo list
|
|
238
|
+
}
|
|
239
|
+
```
|
|
240
|
+
|
|
241
|
+
**Implementation:**
|
|
242
|
+
|
|
243
|
+
1. **Completion detection** - When todo marked completed
|
|
244
|
+
2. **Anchor extraction** - Generate one-line summary + key files
|
|
245
|
+
3. **Context update** - Replace full todo with anchor
|
|
246
|
+
4. **Memex archive** - Store full todo for retrieval
|
|
247
|
+
|
|
248
|
+
**Code locations:**
|
|
249
|
+
- `packages/execution/src/tools/todo.ts`
|
|
250
|
+
- `packages/orchestrator/src/agenticRunner.ts:5554` (formatTaskState)
|
|
251
|
+
|
|
252
|
+
### Phase 4: Sub-Agent Context Isolation (PRIORITY: MEDIUM)
|
|
253
|
+
|
|
254
|
+
**Problem:** Sub-agents inherit parent context, causing bloat
|
|
255
|
+
|
|
256
|
+
**Solution:** Clean context boundary with explicit handoff
|
|
257
|
+
|
|
258
|
+
```typescript
|
|
259
|
+
interface SubAgentContext {
|
|
260
|
+
// Minimal context for sub-agent
|
|
261
|
+
task: string; // Delegated task only
|
|
262
|
+
relevantFiles: string[]; // Only files needed
|
|
263
|
+
toolSubset: string[]; // Only tools needed
|
|
264
|
+
|
|
265
|
+
// NOT included
|
|
266
|
+
// - Parent conversation history
|
|
267
|
+
// - Completed todos
|
|
268
|
+
// - Other sub-agent results
|
|
269
|
+
}
|
|
270
|
+
|
|
271
|
+
interface HandoffProtocol {
|
|
272
|
+
// Parent → Sub-agent
|
|
273
|
+
handoff: {
|
|
274
|
+
task: string;
|
|
275
|
+
files: FileContent[]; // Pre-loaded files
|
|
276
|
+
constraints: string[]; // Rules to follow
|
|
277
|
+
};
|
|
278
|
+
|
|
279
|
+
// Sub-agent → Parent
|
|
280
|
+
return: {
|
|
281
|
+
summary: string; // What was done
|
|
282
|
+
filesModified: string[]; // Files changed
|
|
283
|
+
artifacts: string[]; // Memex IDs for results
|
|
284
|
+
};
|
|
285
|
+
}
|
|
286
|
+
```
|
|
287
|
+
|
|
288
|
+
**Implementation:**
|
|
289
|
+
|
|
290
|
+
1. **Context stripping** - Remove parent context before spawn
|
|
291
|
+
2. **File pre-loading** - Only relevant files passed
|
|
292
|
+
3. **Result summarization** - Sub-agent returns summary, not full context
|
|
293
|
+
4. **Parent integration** - Summary added to parent context
|
|
294
|
+
|
|
295
|
+
**Code locations:**
|
|
296
|
+
- `packages/execution/src/tools/agent-tool.ts`
|
|
297
|
+
- `packages/orchestrator/src/agent-types.ts`
|
|
298
|
+
|
|
299
|
+
### Phase 5: Proactive Context Pruning (PRIORITY: LOW)
|
|
300
|
+
|
|
301
|
+
**Problem:** Compaction only at threshold, no proactive cleanup
|
|
302
|
+
|
|
303
|
+
**Solution:** Background pruning of low-value context
|
|
304
|
+
|
|
305
|
+
```typescript
|
|
306
|
+
interface PruningRules {
|
|
307
|
+
// Auto-prune after N turns
|
|
308
|
+
duplicateToolCalls: { maxOccurrences: 2 };
|
|
309
|
+
oldFileReads: { maxAge: 10 }; // Turns
|
|
310
|
+
successfulTests: { keepSummary: true };
|
|
311
|
+
|
|
312
|
+
// Never prune
|
|
313
|
+
errors: { preserve: true };
|
|
314
|
+
decisions: { preserve: true };
|
|
315
|
+
activeFiles: { preserve: true };
|
|
316
|
+
}
|
|
317
|
+
|
|
318
|
+
// Background pruning (runs every 5 turns)
|
|
319
|
+
function proactivePrune(): void {
|
|
320
|
+
// Remove duplicate tool calls
|
|
321
|
+
// Summarize old file reads
|
|
322
|
+
// Archive successful test runs
|
|
323
|
+
// Update SNR
|
|
324
|
+
}
|
|
325
|
+
```
|
|
326
|
+
|
|
327
|
+
**Implementation:**
|
|
328
|
+
|
|
329
|
+
1. **Turn counter** - Track message age
|
|
330
|
+
2. **Value scoring** - Score each message for relevance
|
|
331
|
+
3. **Background pruning** - Run every N turns
|
|
332
|
+
4. **SNR update** - Recalculate after pruning
|
|
333
|
+
|
|
334
|
+
**Code locations:**
|
|
335
|
+
- `packages/orchestrator/src/agenticRunner.ts:5357` (compactMessages)
|
|
336
|
+
|
|
337
|
+
### Phase 6: Anchor Surfacing (PRIORITY: LOW)
|
|
338
|
+
|
|
339
|
+
**Problem:** Previous task context always present
|
|
340
|
+
|
|
341
|
+
**Solution:** Demand-driven anchor retrieval
|
|
342
|
+
|
|
343
|
+
```typescript
|
|
344
|
+
interface AnchorStore {
|
|
345
|
+
// Lightweight anchors always present
|
|
346
|
+
anchors: Map<string, Anchor>;
|
|
347
|
+
|
|
348
|
+
// Full context retrieved on demand
|
|
349
|
+
archive: Map<string, string>; // Memex IDs
|
|
350
|
+
}
|
|
351
|
+
|
|
352
|
+
interface Anchor {
|
|
353
|
+
id: string;
|
|
354
|
+
type: "file" | "decision" | "error" | "task";
|
|
355
|
+
summary: string; // One line
|
|
356
|
+
keywords: string[]; // For retrieval
|
|
357
|
+
memexId?: string; // Full context
|
|
358
|
+
}
|
|
359
|
+
|
|
360
|
+
// Retrieval triggered by:
|
|
361
|
+
// - Keyword match in current task
|
|
362
|
+
// - File path reference
|
|
363
|
+
// - Explicit memex_retrieve call
|
|
364
|
+
```
|
|
365
|
+
|
|
366
|
+
**Implementation:**
|
|
367
|
+
|
|
368
|
+
1. **Anchor extraction** - During compaction, extract anchors
|
|
369
|
+
2. **Keyword indexing** - Index anchors by keywords
|
|
370
|
+
3. **Demand retrieval** - Surface when keywords match
|
|
371
|
+
4. **Context injection** - Add retrieved context to recent messages
|
|
372
|
+
|
|
373
|
+
**Code locations:**
|
|
374
|
+
- `packages/orchestrator/src/agenticRunner.ts:5561` (formatFileRegistry)
|
|
375
|
+
- `packages/memory/src/episodeStore.ts` (semantic search)
|
|
376
|
+
|
|
377
|
+
---
|
|
378
|
+
|
|
379
|
+
## 4. Implementation Priority Matrix
|
|
380
|
+
|
|
381
|
+
| Phase | Impact | Effort | Priority | Dependencies |
|
|
382
|
+
|-------|--------|--------|----------|--------------|
|
|
383
|
+
| Phase 1: Context Tree | HIGH | HIGH | P0 | None |
|
|
384
|
+
| Phase 2: Tool Lazy Loading | HIGH | MEDIUM | P0 | None |
|
|
385
|
+
| Phase 3: Todo Pruning | MEDIUM | LOW | P1 | Phase 1 |
|
|
386
|
+
| Phase 4: Sub-Agent Isolation | MEDIUM | MEDIUM | P1 | Phase 1 |
|
|
387
|
+
| Phase 5: Proactive Pruning | LOW | MEDIUM | P2 | Phase 1 |
|
|
388
|
+
| Phase 6: Anchor Surfacing | LOW | HIGH | P2 | Phase 1, Phase 3 |
|
|
389
|
+
|
|
390
|
+
**Recommended Order:**
|
|
391
|
+
1. **Phase 2** (Tool Lazy Loading) - Quick win, immediate token savings
|
|
392
|
+
2. **Phase 1** (Context Tree) - Foundation for all other phases
|
|
393
|
+
3. **Phase 3** (Todo Pruning) - Builds on Phase 1
|
|
394
|
+
4. **Phase 4** (Sub-Agent Isolation) - Independent, medium effort
|
|
395
|
+
5. **Phase 5** (Proactive Pruning) - Enhancement
|
|
396
|
+
6. **Phase 6** (Anchor Surfacing) - Enhancement
|
|
397
|
+
|
|
398
|
+
---
|
|
399
|
+
|
|
400
|
+
## 5. Expected Outcomes
|
|
401
|
+
|
|
402
|
+
### Token Savings (Estimated)
|
|
403
|
+
|
|
404
|
+
| Component | Before | After | Savings |
|
|
405
|
+
|-----------|--------|-------|---------|
|
|
406
|
+
| Tool schemas | ~15K | ~3K | 80% |
|
|
407
|
+
| Completed todos | ~5K | ~500 | 90% |
|
|
408
|
+
| Old file reads | ~10K | ~2K | 80% |
|
|
409
|
+
| Sub-agent context | ~8K | ~2K | 75% |
|
|
410
|
+
| **Total** | ~38K | ~7.5K | **80%** |
|
|
411
|
+
|
|
412
|
+
### Repetition Loop Reduction
|
|
413
|
+
|
|
414
|
+
- **Before:** 35% repetition rate at 35% context fill
|
|
415
|
+
- **After:** <10% repetition rate at 70% context fill
|
|
416
|
+
|
|
417
|
+
### Context Utilization
|
|
418
|
+
|
|
419
|
+
- **Before:** Monolithic, degrades at 35%
|
|
420
|
+
- **After:** Hierarchical, maintains quality to 70%
|
|
421
|
+
|
|
422
|
+
---
|
|
423
|
+
|
|
424
|
+
## 6. Research References
|
|
425
|
+
|
|
426
|
+
1. Lost in the Middle: https://arxiv.org/abs/2307.03172
|
|
427
|
+
2. RECOMP: https://arxiv.org/abs/2310.04408
|
|
428
|
+
3. AgentFold: https://arxiv.org/abs/2510.24699
|
|
429
|
+
4. ARC: https://arxiv.org/abs/2601.12030
|
|
430
|
+
5. NATURAL PLAN: https://arxiv.org/abs/2406.04520
|
|
431
|
+
6. ToolLLM DFSDT: https://arxiv.org/abs/2307.16789
|
|
432
|
+
7. SPRINT: https://arxiv.org/abs/2506.05745
|
|
433
|
+
8. Recursive LMs: https://arxiv.org/abs/2512.24601
|
|
434
|
+
9. MASS (multi-agent topology): https://arxiv.org/abs/2502.11578
|
|
435
|
+
10. ExpeL (experience learning): https://arxiv.org/abs/2308.10144
|
|
436
|
+
|
|
437
|
+
---
|
|
438
|
+
|
|
439
|
+
## 7. Next Steps
|
|
440
|
+
|
|
441
|
+
1. **Review this proposal** with team
|
|
442
|
+
2. **Prototype Phase 2** (Tool Lazy Loading) - 2-3 days
|
|
443
|
+
3. **Design Phase 1** (Context Tree) - 1 week
|
|
444
|
+
4. **Implement Phase 1** - 2-3 weeks
|
|
445
|
+
5. **Iterate on remaining phases** based on learnings
|
|
446
|
+
|
|
447
|
+
---
|
|
448
|
+
|
|
449
|
+
*Generated by Omnius — Context Management Deep Dive (2026-04-25)*
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
# Meta-Analysis: False Positive Duplicate Tool Call Detection
|
|
2
|
+
|
|
3
|
+
**Date:** 2026-05-14
|
|
4
|
+
**Scope:** `packages/orchestrator/src/agenticRunner.ts` — `proactivePrune()` + `_buildResourceKey()`
|
|
5
|
+
|
|
6
|
+
## Executive Summary
|
|
7
|
+
|
|
8
|
+
Two interrelated bugs were discovered in the proactive context pruning system:
|
|
9
|
+
|
|
10
|
+
1. **REG-62 semantic dedup was designed but never wired** — the `seenResource` Map and `_buildResourceKey()` method existed but were dead code
|
|
11
|
+
2. **`_buildResourceKey` for `file_read` ignored offset/limit** — all reads of the same file mapped to the same resource key regardless of which lines were being read
|
|
12
|
+
|
|
13
|
+
If REG-62 had been wired WITHOUT fixing #2, it would have caused **catastrophic false positives**: reading different sections of a large file would be treated as the same resource, causing the older section to be pruned, then re-requested, then flagged as a duplicate — an infinite loop of false positive detection.
|
|
14
|
+
|
|
15
|
+
## Root Cause Analysis
|
|
16
|
+
|
|
17
|
+
### Bug 1: Dead Code — REG-62 Never Wired
|
|
18
|
+
|
|
19
|
+
```
|
|
20
|
+
Line 2993: const seenResource = new Map<string, { turn: number; idx: number }>();
|
|
21
|
+
```
|
|
22
|
+
|
|
23
|
+
This Map was declared but never populated or checked. The `_buildResourceKey()` method (line ~4611) was defined but never called from `proactivePrune()`. The entire semantic dedup layer — designed to catch cross-tool reads of the same resource (e.g., `file_read("foo.ts")` + `shell("cat foo.ts")`) — was inert.
|
|
24
|
+
|
|
25
|
+
**Impact:** Without semantic dedup, the system relied solely on exact fingerprint matching (`_buildToolFingerprint`), which includes ALL arguments. This meant `file_read("foo.ts", offset=10)` and `file_read("foo.ts", offset=50)` were correctly treated as different calls by exact dedup. However, the semantic layer that should catch cross-tool duplicates was missing entirely.
|
|
26
|
+
|
|
27
|
+
### Bug 2: Resource Key Ignored Line Range
|
|
28
|
+
|
|
29
|
+
```typescript
|
|
30
|
+
// BEFORE (broken):
|
|
31
|
+
if (name === "file_read") {
|
|
32
|
+
const p = String(a.path ?? a.file ?? "");
|
|
33
|
+
return p ? `resource:file:${p}` : "";
|
|
34
|
+
}
|
|
35
|
+
```
|
|
36
|
+
|
|
37
|
+
All `file_read` calls to the same path produced `resource:file:foo.ts` — regardless of offset/limit. This is the exact false positive the user reported: reads of different lines in larger files would be seen as the same resource.
|
|
38
|
+
|
|
39
|
+
**Why this matters:** When semantic dedup IS active, two calls with the same resource key but different fingerprints trigger pruning of the older call. Without offset/limit in the key, reading line 10-50 and then reading line 200-250 of the same file would prune the first read — even though both sections contain unique, needed information.
|
|
40
|
+
|
|
41
|
+
## Fixes Applied
|
|
42
|
+
|
|
43
|
+
### REG-67: Include offset/limit in resource key
|
|
44
|
+
|
|
45
|
+
```typescript
|
|
46
|
+
// AFTER (fixed):
|
|
47
|
+
if (name === "file_read") {
|
|
48
|
+
const p = String(a.path ?? a.file ?? "");
|
|
49
|
+
if (!p) return "";
|
|
50
|
+
const offset = a.offset != null ? Number(a.offset) : undefined;
|
|
51
|
+
const limit = a.limit != null ? Number(a.limit) : undefined;
|
|
52
|
+
if (offset !== undefined || limit !== undefined) {
|
|
53
|
+
return `resource:file:${p}@${offset ?? 0}:${limit ?? "end"}`;
|
|
54
|
+
}
|
|
55
|
+
return `resource:file:${p}`;
|
|
56
|
+
}
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
Now `file_read("big.ts", offset=10, limit=50)` → `resource:file:big.ts@10:50`
|
|
60
|
+
And `file_read("big.ts", offset=200, limit=50)` → `resource:file:big.ts@200:50`
|
|
61
|
+
Different keys → no false positive dedup.
|
|
62
|
+
|
|
63
|
+
### REG-62 Wired: Semantic dedup now active
|
|
64
|
+
|
|
65
|
+
The `seenResource` Map is now populated and checked in `proactivePrune()`. When two different exact calls target the same resource (same `rkey`), the older one is pruned. This correctly catches:
|
|
66
|
+
|
|
67
|
+
- `file_read("foo.ts")` followed by `file_read("foo.ts")` — same resource, prune older
|
|
68
|
+
- `shell("cat foo.ts")` followed by `file_read("foo.ts")` — same resource, prune older
|
|
69
|
+
- `file_read("foo.ts", offset=10)` followed by `file_read("foo.ts", offset=10)` — same resource+range, prune older
|
|
70
|
+
|
|
71
|
+
And correctly does NOT prune:
|
|
72
|
+
|
|
73
|
+
- `file_read("foo.ts", offset=10)` followed by `file_read("foo.ts", offset=200)` — different ranges, keep both
|
|
74
|
+
|
|
75
|
+
## Memory System Self-Evaluation
|
|
76
|
+
|
|
77
|
+
### What the memory systems got right
|
|
78
|
+
- Previous sessions correctly identified the duplicate tool call problem and stored findings in `project/bug_fixes`
|
|
79
|
+
- The `inner_critic_findings` topic flagged the agenticRunner.ts as a high-risk file for regressions
|
|
80
|
+
- Task lessons from prior sessions guided toward reading before editing
|
|
81
|
+
|
|
82
|
+
### What the memory systems missed
|
|
83
|
+
- No memory entry flagged that REG-62 was dead code — the `seenResource` Map was declared but never used
|
|
84
|
+
- The `project/architecture` topic lists 88 tools but doesn't document the proactivePrune dedup layers
|
|
85
|
+
- No cross-reference between the "duplicate calls" bug pattern and the specific dedup mechanism that causes it
|
|
86
|
+
|
|
87
|
+
### Recommendations for memory improvement
|
|
88
|
+
1. **Dead code detection**: When storing architecture knowledge, flag declared-but-unused variables as potential bugs
|
|
89
|
+
2. **Layer documentation**: The dedup system has 3 layers (exact fingerprint, semantic resource, aged file) — this should be documented in `project/architecture`
|
|
90
|
+
3. **False positive tracking**: Add a `project/false_positive_patterns` topic for known scenarios where dedup incorrectly flags unique calls
|
|
91
|
+
|
|
92
|
+
## Verification
|
|
93
|
+
|
|
94
|
+
- TypeScript compilation: ✅ `tsc --noEmit` passes
|
|
95
|
+
- Grep verification: ✅ dedup/duplicate references found in 11 orchestrator source files
|
|
96
|
+
- Edit verification: ✅ REG-67 and REG-62 comments present in agenticRunner.ts
|