dsh-omni-router 2.0.0 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -71,7 +71,8 @@
71
71
  - **Evidence Engine / Write Locks**:`src/evidence-engine.mjs` 统一证据协议与存储;DAG scheduler 通过 `selectReadyBatch` 避免 write overlap 并行冲突。
72
72
  - **Stage 3/5**:Recursive Project Index、contextBudget 接入、Per-task Context、Missing capability handling、OmniBench v2 prompt generator。
73
73
  - **Stage 4(部分)**:`compileDagToWorkflow` 将 Mission DAG 编译为 DSH-native workflow outline。
74
- - **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 scaffold(manifest schema + runner stub + prepare script + DAG-to-plan compile)。
74
+ - **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 runner(manifest schema + prepare + prompt generation + `--exec` 本地执行/结果采集 + DAG-to-plan compile)。
75
+ - **2.1 收敛增强**:Mission Resume(`omni_mission_resume` 跨 session 继续 DAG,`runDagLoop` 支持 `onProgress` 周期保存)、Harness evidence extraction(`extractHarnessEvidence` 从 commands/tests/toolCalls/EVIDENCE_JSON 提取机器可验证证据)、context token budget(`maxContextTokens` 真正限制检索输出)、semantic failure retrieval + evidence-backed skill distillation、OmniBench v2 real runner。
75
76
  - **TaskDecision**:`createTaskDecision` 生成唯一决策对象,Policy/Runtime 统一消费,避免重复 classify 分叉。
76
77
  - **Evidence Protocol**:`src/evidence.mjs` 用结构化 command/file/test/finding 证据判定 PASS/FAIL;`isQaPass` 已支持从 QA 输出解析 JSON evidence,不再只信文本。
77
78
  - **Task Compiler**:`compileTask` 生成 objective / constraints / non-goals / acceptance / hidden assumptions / ambiguities / invariants / risk / artifacts;`compileTaskWithLLM` 可用 LLM 增强。
@@ -153,7 +154,7 @@ Then restart DSH and select **Omni Router** in a new session.
153
154
  - `/omni direct` — enter direct mode.
154
155
  - `/omni mode spec|react|balanced` — set thinking mode.
155
156
  - `/omni reroute plan|direct` — adaptively reroute current task.
156
- - Model tools: `omni_status`, `omni_plan`, `omni_direct`, `omni_mode`, `omni_reroute`, `omni_delegate`, `omni_memory`, `omni_benchmark`, `omni_mission_run`, `omni_visual_check`.
157
+ - Model tools: `omni_status`, `omni_plan`, `omni_direct`, `omni_mode`, `omni_reroute`, `omni_delegate`, `omni_memory`, `omni_benchmark`, `omni_mission_run`, `omni_mission_resume`, `omni_visual_check`.
157
158
 
158
159
  ## Configuration
159
160
 
package/README.zh-CN.md CHANGED
@@ -71,7 +71,8 @@
71
71
  - **Evidence Engine / Write Locks**:`src/evidence-engine.mjs` 统一证据协议与存储;DAG scheduler 通过 `selectReadyBatch` 避免 write overlap 并行冲突。
72
72
  - **Stage 3/5**:Recursive Project Index、contextBudget 接入、Per-task Context、Missing capability handling、OmniBench v2 prompt generator。
73
73
  - **Stage 4(部分)**:`compileDagToWorkflow` 将 Mission DAG 编译为 DSH-native workflow outline。
74
- - **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 scaffold(manifest schema + runner stub + prepare script + DAG-to-plan compile)。
74
+ - **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 runner(manifest schema + prepare + prompt generation + `--exec` 本地执行/结果采集 + DAG-to-plan compile)。
75
+ - **2.1 收敛增强**:Mission Resume(`omni_mission_resume` 跨 session 继续 DAG,`runDagLoop` 支持 `onProgress` 周期保存)、Harness evidence extraction(`extractHarnessEvidence` 从 commands/tests/toolCalls/EVIDENCE_JSON 提取机器可验证证据)、context token budget(`maxContextTokens` 真正限制检索输出)、semantic failure retrieval + evidence-backed skill distillation、OmniBench v2 real runner。
75
76
  - **TaskDecision**:`createTaskDecision` 生成唯一决策对象,Policy/Runtime 统一消费,避免重复 classify 分叉。
76
77
  - **Evidence Protocol**:`src/evidence.mjs` 用结构化 command/file/test/finding 证据判定 PASS/FAIL;`isQaPass` 已支持从 QA 输出解析 JSON evidence,不再只信文本。
77
78
  - **Task Compiler**:`compileTask` 生成 objective / constraints / non-goals / acceptance / hidden assumptions / ambiguities / invariants / risk / artifacts;`compileTaskWithLLM` 可用 LLM 增强。
@@ -153,7 +154,7 @@ node scripts/install-preset.mjs
153
154
  - `/omni direct` — 进入直接执行模式。
154
155
  - `/omni mode spec|react|balanced` — 设置思维模式。
155
156
  - `/omni reroute plan|direct` — 动态切换当前任务路由。
156
- - 模型工具:`omni_status` / `omni_plan` / `omni_direct` / `omni_mode` / `omni_reroute` / `omni_delegate` / `omni_memory` / `omni_benchmark` / `omni_mission_run` / `omni_visual_check`。
157
+ - 模型工具:`omni_status` / `omni_plan` / `omni_direct` / `omni_mode` / `omni_reroute` / `omni_delegate` / `omni_memory` / `omni_benchmark` / `omni_mission_run` / `omni_mission_resume` / `omni_visual_check`。
157
158
 
158
159
  ## 配置
159
160
 
@@ -1,8 +1,8 @@
1
- # OmniBench v2 (scaffold)
1
+ # OmniBench v2
2
2
 
3
3
  Goal: move from 10 small fixture tasks to a reproducible multi-repo benchmark.
4
4
 
5
- ## Requirements (from 优化5.md)
5
+ ## Requirements (from 优化5.md / 优化6.md)
6
6
 
7
7
  - 30~100 real GitHub repos
8
8
  - fixed starting commit
@@ -25,18 +25,18 @@ Goal: move from 10 small fixture tasks to a reproducible multi-repo benchmark.
25
25
  "task": "Fix intermittent login timeout under concurrent session refresh.",
26
26
  "acceptance": ["regression test added", "all tests pass"],
27
27
  "hiddenTests": ["tests/hidden/session-concurrent.test.js"],
28
+ "setupCommand": "npm ci",
29
+ "baselineCommand": "npm test",
30
+ "agentCommand": "node path/to/dsh-agent-runner.mjs",
31
+ "verifyCommand": "npm run test:hidden",
28
32
  "timeoutMs": 7200000,
29
33
  "runs": 5
30
34
  }
31
35
  ```
32
36
 
33
- See `manifest.example.json`.
34
-
35
- ## Runner stub
36
-
37
- `node benchmark/omnibench-v2/runner-stub.mjs <manifest.json>` validates the
38
- manifest and prints a run plan. Actual execution will be driven by DSH
39
- subagents in a future integration.
37
+ `setupCommand`, `baselineCommand`, `agentCommand`, and `verifyCommand` are
38
+ optional. `agentCommand` receives the generated prompt as a single quoted
39
+ argument; when present, `--exec` executes it inside the checked-out repo.
40
40
 
41
41
  ## Generate DSH prompts
42
42
 
@@ -46,7 +46,29 @@ npm run omnibench:v2:generate -- benchmark/omnibench-v2/manifest.local.example.j
46
46
 
47
47
  This writes one prompt per repo/arm/run under `benchmark/omnibench-v2/prompts/`.
48
48
  Open each prompt in a DSH Desktop session with the Omni Router preset to
49
- execute, then collect the results.
49
+ execute manually, or wire `agentCommand` to your DSH CLI/runner for automation.
50
+
51
+ ## Run locally (automated)
52
+
53
+ ```bash
54
+ # Plan + prompt generation only (default)
55
+ node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json
56
+
57
+ # Execute each run with the manifest's agentCommand
58
+ node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json --exec
59
+
60
+ # Override the agent command for every run
61
+ node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json --exec --agent-command "node agent.mjs"
62
+ ```
63
+
64
+ The runner clones/checks out each repo, optionally runs setup/baseline/verify
65
+ commands, captures exit codes and output, and writes a timestamped JSON result
66
+ file to `benchmark/omnibench-v2/results/`.
67
+
68
+ ## Runner stub
69
+
70
+ `node benchmark/omnibench-v2/runner-stub.mjs <manifest.json>` still validates
71
+ the manifest and prints the run plan for quick checks.
50
72
 
51
73
  ## Local / offline usage
52
74
 
@@ -57,6 +79,7 @@ existing `Omni group` fixture:
57
79
  ```bash
58
80
  npm run omnibench:v2:prepare -- benchmark/omnibench-v2/manifest.local.example.json
59
81
  npm run omnibench:v2:plan -- benchmark/omnibench-v2/manifest.local.example.json
82
+ npm run omnibench:v2:generate -- benchmark/omnibench-v2/manifest.local.example.json
60
83
  ```
61
84
 
62
85
  `git clone` works with local paths, so no network is required. Replace `repo`
@@ -2,20 +2,24 @@
2
2
  /**
3
3
  * OmniBench v2 runner / prompt generator.
4
4
  *
5
- * Generates per-repo per-arm per-run DSH prompts. Actual execution must happen
6
- * inside DSH Desktop; this script makes that reproducible.
5
+ * Generates per-repo per-arm per-run DSH prompts. With `--exec` and a
6
+ * configured `agentCommand` (or `--agent-command`), it also executes the run
7
+ * locally, runs baseline/verify commands, and writes machine-readable results.
7
8
  *
8
9
  * Usage:
9
- * node benchmark/omnibench-v2/run.mjs <manifest.json>
10
+ * node benchmark/omnibench-v2/run.mjs <manifest.json> # plan + prompts
11
+ * node benchmark/omnibench-v2/run.mjs <manifest.json> --exec # execute agentCommand runs
12
+ * node benchmark/omnibench-v2/run.mjs <manifest.json> --exec --agent-command "node agent.mjs"
10
13
  */
11
14
 
12
15
  import fs from 'node:fs'
13
16
  import path from 'node:path'
14
- import { fileURLToPath } from 'node:url'
17
+ import { execSync } from 'node:child_process'
18
+ import { fileURLToPath, pathToFileURL } from 'node:url'
15
19
 
16
20
  const here = path.dirname(fileURLToPath(import.meta.url))
17
21
 
18
- function buildPrompt(manifest, arm) {
22
+ export function buildPrompt(manifest, arm) {
19
23
  const criteria = (manifest.acceptance || []).map((c) => `- ${c}`).join('\n')
20
24
  const task = `Task:\n${manifest.task}\n\nAcceptance criteria:\n${criteria}`
21
25
  if (arm === 'raw') {
@@ -24,10 +28,54 @@ function buildPrompt(manifest, arm) {
24
28
  return `${task}\n\nYou are using the Omni control plane. Follow engineering methodology, verify with real evidence, and do not claim completion without checks. When done, reply exactly "BENCHMARK: PASS" if you verified all criteria, otherwise "BENCHMARK: FAIL".`
25
29
  }
26
30
 
27
- function main() {
28
- const manifestPath = process.argv[2] || path.join(here, 'manifest.local.example.json')
31
+ export function readManifests(manifestPath) {
29
32
  const manifests = JSON.parse(fs.readFileSync(manifestPath, 'utf8'))
30
- const list = Array.isArray(manifests) ? manifests : [manifests]
33
+ return Array.isArray(manifests) ? manifests : [manifests]
34
+ }
35
+
36
+ export function validateManifests(list) {
37
+ const errors = []
38
+ for (const m of list) {
39
+ if (!m.id) errors.push('missing id')
40
+ if (!m.repo) errors.push(`${m.id || '?'}: missing repo`)
41
+ if (!m.commit) errors.push(`${m.id || '?'}: missing commit`)
42
+ if (!m.task) errors.push(`${m.id || '?'}: missing task`)
43
+ if (!Array.isArray(m.acceptance) || !m.acceptance.length) errors.push(`${m.id || '?'}: missing acceptance`)
44
+ if (!m.runs || m.runs < 3) errors.push(`${m.id || '?'}: runs should be >= 3`)
45
+ }
46
+ return errors
47
+ }
48
+
49
+ export function ensureRepo(m) {
50
+ const workDir = path.resolve(process.env.OMNIBENCH_REPOS || path.join(here, 'repos'))
51
+ const dir = path.join(workDir, m.id)
52
+ fs.mkdirSync(workDir, { recursive: true })
53
+ if (!fs.existsSync(path.join(dir, '.git'))) {
54
+ console.log(`cloning ${m.repo} -> ${dir}`)
55
+ execSync(`git clone "${m.repo}" "${dir}"`, { stdio: 'inherit' })
56
+ }
57
+ console.log(`checkout ${m.id} @ ${m.commit}`)
58
+ execSync(`git -C "${dir}" checkout ${m.commit}`, { stdio: 'inherit' })
59
+ return dir
60
+ }
61
+
62
+ export function runCommand(command, cwd, label) {
63
+ if (!command) return { skipped: true }
64
+ const start = Date.now()
65
+ try {
66
+ const output = execSync(command, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] })
67
+ return { skipped: false, exitCode: 0, output, durationMs: Date.now() - start }
68
+ } catch (error) {
69
+ return {
70
+ skipped: false,
71
+ exitCode: error.status ?? 1,
72
+ output: String(error.stdout || '') + String(error.stderr || ''),
73
+ durationMs: Date.now() - start,
74
+ }
75
+ }
76
+ }
77
+
78
+ export function generatePrompts(list) {
31
79
  const outDir = path.join(here, 'prompts')
32
80
  fs.mkdirSync(outDir, { recursive: true })
33
81
  let count = 0
@@ -40,8 +88,115 @@ function main() {
40
88
  }
41
89
  }
42
90
  }
43
- console.log(`Generated ${count} DSH prompts in ${outDir}`)
44
- console.log('Run each prompt inside a DSH Desktop session with the Omni Router preset, then collect results.')
91
+ return count
92
+ }
93
+
94
+ export function writeResults(results, resultsDir) {
95
+ fs.mkdirSync(resultsDir, { recursive: true })
96
+ const file = path.join(resultsDir, `omnibench-v2-${Date.now()}.json`)
97
+ fs.writeFileSync(file, JSON.stringify(results, null, 2), 'utf8')
98
+ return file
99
+ }
100
+
101
+ function main() {
102
+ const manifestPath = process.argv[2] || path.join(here, 'manifest.local.example.json')
103
+ const execMode = process.argv.includes('--exec')
104
+ const agentCommandOverride = (() => {
105
+ const idx = process.argv.indexOf('--agent-command')
106
+ return idx !== -1 ? process.argv[idx + 1] : null
107
+ })()
108
+ const resultsDir = path.resolve(process.env.OMNIBENCH_RESULTS || path.join(here, 'results'))
109
+
110
+ const list = readManifests(manifestPath)
111
+ const errors = validateManifests(list)
112
+ if (errors.length) {
113
+ console.error('Invalid manifest:')
114
+ for (const e of errors) console.error(`- ${e}`)
115
+ process.exit(1)
116
+ }
117
+ const promptCount = generatePrompts(list)
118
+
119
+ let totalRuns = 0
120
+ for (const m of list) totalRuns += (m.runs || 1) * 2
121
+ console.log(`Generated ${promptCount} DSH prompts in ${path.join(here, 'prompts')}`)
122
+ console.log(`OmniBench v2 run plan (${totalRuns} agent runs)`)
123
+
124
+ if (!execMode) {
125
+ for (const m of list) {
126
+ console.log(`- ${m.id}: ${m.repo} @ ${m.commit} (${m.language}/${m.framework}) runs=${m.runs} -> raw ${m.runs} + omni ${m.runs}`)
127
+ }
128
+ console.log('Pass --exec and configure agentCommand to execute runs locally.')
129
+ return
130
+ }
131
+
132
+ const results = []
133
+ for (const m of list) {
134
+ const repoDir = ensureRepo(m)
135
+ for (const arm of ['raw', 'omni']) {
136
+ for (let i = 1; i <= (m.runs || 3); i++) {
137
+ const agentCommand = agentCommandOverride || m.agentCommand
138
+ const promptFile = path.join(here, 'prompts', `${m.id}-${arm}-${i}.txt`)
139
+ const prompt = fs.readFileSync(promptFile, 'utf8')
140
+ console.log(`\n=== ${m.id} ${arm} run ${i} ===`)
141
+ if (m.setupCommand) {
142
+ const setup = runCommand(m.setupCommand, repoDir, 'setup')
143
+ console.log(`setup exit=${setup.exitCode ?? 'skip'}`)
144
+ }
145
+ const baseline = m.baselineCommand ? runCommand(m.baselineCommand, repoDir, 'baseline') : null
146
+ if (baseline) console.log(`baseline exit=${baseline.exitCode ?? 'skip'}`)
147
+
148
+ let agent = null
149
+ if (agentCommand) {
150
+ const command = `${agentCommand} ${JSON.stringify(prompt)}`
151
+ agent = runCommand(command, repoDir, 'agent')
152
+ console.log(`agent exit=${agent.exitCode ?? 'skip'} duration=${agent.durationMs ?? 0}ms`)
153
+ } else {
154
+ console.log(`agentCommand not configured; run prompt manually in DSH Desktop:\n ${promptFile}`)
155
+ }
156
+
157
+ const verify = m.verifyCommand ? runCommand(m.verifyCommand, repoDir, 'verify') : null
158
+ if (verify) console.log(`verify exit=${verify.exitCode ?? 'skip'}`)
159
+
160
+ let success = null
161
+ if (verify && verify.skipped !== true) {
162
+ success = verify.exitCode === 0
163
+ } else if (agent && agent.skipped !== true) {
164
+ success = agent.exitCode === 0 && /BENCHMARK:\s*PASS/i.test(agent.output || '')
165
+ }
166
+
167
+ results.push({
168
+ id: m.id,
169
+ arm,
170
+ run: i,
171
+ repo: m.repo,
172
+ commit: m.commit,
173
+ task: m.task,
174
+ success,
175
+ agentExitCode: agent?.exitCode ?? null,
176
+ verifyExitCode: verify?.exitCode ?? null,
177
+ durationMs: (agent?.durationMs || 0) + (verify?.durationMs || 0),
178
+ baselineExitCode: baseline?.exitCode ?? null,
179
+ agentOutput: agent?.output?.slice(0, 4000) || '',
180
+ verifyOutput: verify?.output?.slice(0, 4000) || '',
181
+ metrics: {
182
+ tokens: 0,
183
+ cost: 0,
184
+ toolCalls: 0,
185
+ },
186
+ promptFile,
187
+ ranAt: new Date().toISOString(),
188
+ })
189
+ }
190
+ }
191
+ }
192
+
193
+ const file = writeResults(results, resultsDir)
194
+ const executed = results.filter((r) => r.success !== null).length
195
+ const passed = results.filter((r) => r.success === true).length
196
+ console.log(`\nWrote ${results.length} result record(s) to ${file}`)
197
+ console.log(`Executed ${executed}/${results.length} runs; passed ${passed}.`)
45
198
  }
46
199
 
47
- main()
200
+ if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
201
+ main()
202
+ }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-omni-router",
3
- "version": "2.0.0",
3
+ "version": "2.1.0",
4
4
  "description": "Omni - a reliability and orchestration control plane for DeepSeek Harness, designed to turn fast models into dependable long-horizon software-engineering agents.",
5
5
  "type": "module",
6
6
  "main": "./lib/index.js",
@@ -41,7 +41,7 @@
41
41
  }
42
42
  },
43
43
  "scripts": {
44
- "test": "node test/omni-router.test.mjs && node test/agent-chain.test.mjs && node test/skill-suggest.test.mjs && node test/methodology.test.mjs && node test/compat.test.mjs && node test/project-brain.test.mjs && node test/mission-planner.test.mjs && node test/agent-runtime.test.mjs && node test/judge.test.mjs && node test/memory.test.mjs && node test/memory-v3.test.mjs && node test/memory-engine.test.mjs && node test/mission-resume.test.mjs && node test/engineering-benchmark.test.mjs && node test/benchmark-runner.test.mjs && node test/benchmark-results.test.mjs && node test/continuous-eval.test.mjs && node --no-warnings test/project-brain-v2.test.mjs && node --no-warnings test/project-brain-v3.test.mjs && node test/visual-qa.test.mjs && node test/evidence.test.mjs && node test/evidence-store.test.mjs && node test/evidence-engine.test.mjs && node test/task-decision.test.mjs && node test/task-compiler.test.mjs && node test/capability-brain.test.mjs && node test/capability-manifest.test.mjs && node test/capability-sandbox.test.mjs && node test/planner-dag.test.mjs && node test/dynamic-context.test.mjs && node test/mission-dag.test.mjs && node test/failure-taxonomy.test.mjs && node test/context-expansion.test.mjs && node test/hybrid-retrieval.test.mjs && node test/dependency-graph.test.mjs && node test/ast-provider.test.mjs && node test/experience-router.test.mjs && node test/benchmark-tasks.test.mjs && node test/bundle-installer.test.mjs",
44
+ "test": "node test/omni-router.test.mjs && node test/agent-chain.test.mjs && node test/skill-suggest.test.mjs && node test/methodology.test.mjs && node test/compat.test.mjs && node test/project-brain.test.mjs && node test/mission-planner.test.mjs && node test/agent-runtime.test.mjs && node test/judge.test.mjs && node test/memory.test.mjs && node test/memory-v3.test.mjs && node test/memory-engine.test.mjs && node test/mission-resume.test.mjs && node test/engineering-benchmark.test.mjs && node test/benchmark-runner.test.mjs && node test/benchmark-results.test.mjs && node test/continuous-eval.test.mjs && node --no-warnings test/project-brain-v2.test.mjs && node --no-warnings test/project-brain-v3.test.mjs && node test/visual-qa.test.mjs && node test/evidence.test.mjs && node test/evidence-store.test.mjs && node test/evidence-engine.test.mjs && node test/task-decision.test.mjs && node test/task-compiler.test.mjs && node test/capability-brain.test.mjs && node test/capability-manifest.test.mjs && node test/capability-sandbox.test.mjs && node test/planner-dag.test.mjs && node test/dynamic-context.test.mjs && node test/mission-dag.test.mjs && node test/failure-taxonomy.test.mjs && node test/context-expansion.test.mjs && node test/hybrid-retrieval.test.mjs && node test/dependency-graph.test.mjs && node test/ast-provider.test.mjs && node test/experience-router.test.mjs && node test/benchmark-tasks.test.mjs && node test/bundle-installer.test.mjs && node test/omnibench-runner.test.mjs",
45
45
  "benchmark": "node benchmark/run.mjs",
46
46
  "benchmark:engineering": "node benchmark/engineering-run.mjs",
47
47
  "benchmark:compare": "node benchmark/compare.mjs",
@@ -50,6 +50,7 @@
50
50
  "omnibench:v2:plan": "node benchmark/omnibench-v2/runner-stub.mjs",
51
51
  "omnibench:v2:prepare": "node benchmark/omnibench-v2/prepare.mjs",
52
52
  "omnibench:v2:generate": "node benchmark/omnibench-v2/run.mjs",
53
+ "omnibench:v2:run": "node benchmark/omnibench-v2/run.mjs --exec",
53
54
  "analyze": "node benchmark/analyze.mjs",
54
55
  "llm-eval": "node benchmark/llm-eval.mjs",
55
56
  "install:preset": "node scripts/install-preset.mjs"
@@ -165,6 +165,7 @@ export async function runDagLoop(dag, {
165
165
  maxRepairs = 0,
166
166
  maxSameActionRetries = 0,
167
167
  budget = {},
168
+ onProgress = null,
168
169
  } = {}) {
169
170
  const b = {
170
171
  maxSteps: maxSteps ?? budget.steps ?? 50,
@@ -231,6 +232,18 @@ export async function runDagLoop(dag, {
231
232
  actions.push({ taskId: task.id, observation })
232
233
  }
233
234
  step += 1
235
+ if (typeof onProgress === 'function') {
236
+ try {
237
+ await onProgress({
238
+ dag: current,
239
+ step,
240
+ actions,
241
+ metrics: { step, replanCount, repairCount, sameActionCount, tokenUsage, cost, toolCalls },
242
+ })
243
+ } catch {
244
+ // Progress callbacks are best-effort; they must never abort the mission.
245
+ }
246
+ }
234
247
  }
235
248
 
236
249
  const done = isMissionDagComplete(current)
@@ -54,7 +54,15 @@ export function buildProgressiveContext(taskText, entries, files = {}, options =
54
54
  const tests = [...names].filter((n) => /\.(test|spec)\./i.test(n))
55
55
  if (tests.length) parts.push('', 'Tests:', ...tests.map((t) => `- ${t}`))
56
56
  }
57
- return parts.join('\n')
57
+ return applyContextBudget(parts.join('\n'), options)
58
+ }
59
+
60
+ function applyContextBudget(text, options = {}) {
61
+ const maxTokens = Number(options.maxContextTokens) || 0
62
+ if (maxTokens <= 0) return text
63
+ const maxChars = Math.max(32, maxTokens * 4)
64
+ if (text.length <= maxChars) return text
65
+ return `${text.slice(0, maxChars)}\n… (truncated by context budget ${maxTokens} tokens)`
58
66
  }
59
67
 
60
68
  export function shouldExpand(uncertainty = 0, threshold = 0.5) {
@@ -33,6 +33,7 @@ export function buildDynamicContext(taskText, entries, files = {}, options = {})
33
33
  level,
34
34
  maxFiles: options.maxFiles || 8,
35
35
  maxFileChars: options.maxFileChars || 2000,
36
+ maxContextTokens: options.maxContextTokens || 0,
36
37
  graph: options.graph,
37
38
  }),
38
39
  }
package/src/evidence.mjs CHANGED
@@ -70,6 +70,92 @@ export function evidencePass(evidence = {}) {
70
70
  return true
71
71
  }
72
72
 
73
+ /**
74
+ * Extract machine-checkable evidence from a subagent/tool result.
75
+ *
76
+ * Accepts either a structured `result.evidence` object, top-level arrays on the
77
+ * result (`commands`, `tests`, `files`, `findings`), tool-call records, or an
78
+ * embedded `EVIDENCE_JSON` block in the result text. This lets the runtime
79
+ * prefer Harness-captured facts over agent prose whenever the Harness provides
80
+ * them.
81
+ */
82
+ export function extractHarnessEvidence(result = {}) {
83
+ let evidence = createEvidence()
84
+ const source = (result && typeof result === 'object' && result.evidence) || result || {}
85
+
86
+ const apply = (key, fn) => {
87
+ for (const item of Array.isArray(source[key]) ? source[key] : []) {
88
+ try {
89
+ evidence = fn(evidence, item)
90
+ } catch {
91
+ // Skip malformed evidence entries; never fail the mission.
92
+ }
93
+ }
94
+ }
95
+
96
+ apply('commands', addCommandEvidence)
97
+ apply('tests', addTestEvidence)
98
+ apply('files', addFileEvidence)
99
+ apply('findings', addFindingEvidence)
100
+
101
+ if (Array.isArray(result?.toolCalls)) {
102
+ for (const call of result.toolCalls) {
103
+ try {
104
+ if (call?.command || call?.name) {
105
+ evidence = addCommandEvidence(evidence, {
106
+ command: call.command || call.name,
107
+ exitCode: call.exitCode ?? (call.ok === false ? 1 : 0),
108
+ output: call.output || call.stdout || '',
109
+ durationMs: call.durationMs || 0,
110
+ })
111
+ }
112
+ } catch {
113
+ // Ignore malformed tool-call evidence.
114
+ }
115
+ }
116
+ }
117
+
118
+ const embedded = parseEmbeddedEvidence(result?.output || result?.text || '')
119
+ if (embedded) {
120
+ for (const item of embedded.commands || []) {
121
+ try { evidence = addCommandEvidence(evidence, item) } catch { /* ignore */ }
122
+ }
123
+ for (const item of embedded.tests || []) {
124
+ try { evidence = addTestEvidence(evidence, item) } catch { /* ignore */ }
125
+ }
126
+ for (const item of embedded.files || []) {
127
+ try { evidence = addFileEvidence(evidence, item) } catch { /* ignore */ }
128
+ }
129
+ for (const item of embedded.findings || []) {
130
+ try { evidence = addFindingEvidence(evidence, item) } catch { /* ignore */ }
131
+ }
132
+ }
133
+
134
+ return evidence
135
+ }
136
+
137
+ function parseEmbeddedEvidence(text) {
138
+ const source = String(text || '')
139
+ const markers = ['EVIDENCE_JSON', 'EVIDENCE:']
140
+ for (const marker of markers) {
141
+ const start = source.indexOf(marker)
142
+ if (start === -1) continue
143
+ const jsonStart = source.indexOf('{', start)
144
+ if (jsonStart === -1) continue
145
+ const jsonEnd = source.lastIndexOf('}')
146
+ if (jsonEnd === -1 || jsonEnd <= jsonStart) continue
147
+ try {
148
+ const parsed = JSON.parse(source.slice(jsonStart, jsonEnd + 1))
149
+ if (parsed && (Array.isArray(parsed.commands) || Array.isArray(parsed.tests) || Array.isArray(parsed.files) || Array.isArray(parsed.findings))) {
150
+ return parsed
151
+ }
152
+ } catch {
153
+ // Not valid JSON; fall through to prose-based heuristics.
154
+ }
155
+ }
156
+ return null
157
+ }
158
+
73
159
  export function summarizeEvidence(evidence = {}) {
74
160
  const lines = []
75
161
  if (evidence.commands?.length) {
package/src/memory-v3.mjs CHANGED
@@ -17,10 +17,11 @@ export function createMemoryV3(base = {}) {
17
17
  }
18
18
  }
19
19
 
20
- function countOutcomes(trajectories = []) {
20
+ function countOutcomes(trajectories = [], options = {}) {
21
21
  let successCount = 0
22
22
  let failureCount = 0
23
23
  for (const t of trajectories || []) {
24
+ if (options.requireEvidence && !hasEvidenceBacking(t)) continue
24
25
  const text = String(t?.text || t?.outcome || '').toLowerCase()
25
26
  if (/success|done|pass|完成|ok/.test(text)) successCount += 1
26
27
  else if (/fail|error|blocked|失败|错误/.test(text)) failureCount += 1
@@ -28,13 +29,28 @@ function countOutcomes(trajectories = []) {
28
29
  return { successCount, failureCount }
29
30
  }
30
31
 
32
+ function hasEvidenceBacking(t = {}) {
33
+ const evidence = t.evidence || t.verification || t.verified
34
+ if (evidence === true) return true
35
+ if (evidence && typeof evidence === 'object') {
36
+ const e = evidence
37
+ return Boolean(
38
+ (Array.isArray(e.commands) && e.commands.length) ||
39
+ (Array.isArray(e.tests) && e.tests.length) ||
40
+ (Array.isArray(e.files) && e.files.length) ||
41
+ (Array.isArray(e.findings) && e.findings.length),
42
+ )
43
+ }
44
+ return false
45
+ }
46
+
31
47
  export function distillSkill(memory, trajectories = [], options = {}) {
32
48
  const minSuccesses = Number(options.minSuccesses) || 3
33
- const { successCount, failureCount } = countOutcomes(trajectories)
49
+ const { successCount, failureCount } = countOutcomes(trajectories, options)
34
50
  if (successCount < minSuccesses) return { memory, promoted: false, reason: `need ${minSuccesses} successes, got ${successCount}` }
35
51
 
36
52
  const successfulSteps = (trajectories || [])
37
- .filter((t) => /success|done|pass|完成|ok/i.test(String(t?.text || t?.outcome || '')))
53
+ .filter((t) => (!options.requireEvidence || hasEvidenceBacking(t)) && /success|done|pass|完成|ok/i.test(String(t?.text || t?.outcome || '')))
38
54
  .map((t) => t?.text || t?.step || '')
39
55
  .filter(Boolean)
40
56
 
@@ -78,11 +94,27 @@ export function recordSkillOutcome(memory, skillName, success) {
78
94
  return { ...memory, learnedSkills: skills }
79
95
  }
80
96
 
81
- export function retrieveHistoricalFailures(memory, taskText = '') {
97
+ export function retrieveHistoricalFailures(memory, taskText = '', options = {}) {
82
98
  const text = String(taskText || '').toLowerCase()
83
99
  const failures = memory.failures || []
84
100
  if (!text) return failures
85
- return failures.filter((f) => String(f.text || f.reason || '').toLowerCase().includes(text))
101
+ const minScore = options.minScore === undefined ? 0.1 : Number(options.minScore) || 0
102
+ const tokens = tokenize(text)
103
+ return failures
104
+ .map((f) => {
105
+ const hay = `${f.text || ''} ${f.reason || ''} ${f.category || ''} ${f.type || ''}`.toLowerCase()
106
+ const hayTokens = tokenize(hay)
107
+ const overlap = hayTokens.filter((t) => tokens.includes(t)).length
108
+ const score = overlap / Math.max(1, Math.min(tokens.length, hayTokens.length))
109
+ return { ...f, _score: Math.round(score * 1000) / 1000 }
110
+ })
111
+ .filter((f) => f._score >= minScore || String(f.text || f.reason || '').toLowerCase().includes(text))
112
+ .sort((a, b) => b._score - a._score)
113
+ .map(({ _score, ...f }) => f)
114
+ }
115
+
116
+ function tokenize(text) {
117
+ return String(text || '').toLowerCase().match(/[a-z0-9_]+/g) || []
86
118
  }
87
119
 
88
120
  export function recordExecutionPolicy(memory, policy = {}) {
@@ -30,9 +30,9 @@ import { runDagLoop } from './agent-runtime.mjs'
30
30
  import { buildVisualQaPrompt, buildVisualQaStepRequirement, callVisionApi, isFrontendTask, parseVisualQaResponse } from './visual-qa.mjs'
31
31
  import { createTaskDecision, buildPolicyFromTaskDecision } from './task-decision.mjs'
32
32
  import { compileTaskWithLLM } from './task-compiler.mjs'
33
- import { bindCapabilitiesToDag, createMissionDag, formatMissionDag } from './mission-dag.mjs'
33
+ import { bindCapabilitiesToDag, createMissionDag, formatMissionDag, isMissionDagComplete } from './mission-dag.mjs'
34
34
  import { generateMissionDag, roleForTask } from './planner-dag.mjs'
35
- import { autoPopulateCapabilityBrain, createCapabilityBrain } from './capability-brain.mjs'
35
+ import { autoPopulateCapabilityBrain, createCapabilityBrain, recordCapabilityOutcome } from './capability-brain.mjs'
36
36
  import { loadCapabilityManifests } from './capability-manifest.mjs'
37
37
  import { capabilityToolFilter } from './capability-sandbox.mjs'
38
38
  import { buildProgressiveContext } from './context-expansion.mjs'
@@ -42,7 +42,8 @@ import { retrieveContext } from './hybrid-retrieval.mjs'
42
42
  import { formatMemory, recordDecision, recordFailure, recordProject, recordTrajectory, summarizeMemory } from './memory.mjs'
43
43
  import { createMemoryEngine, loadMemoryEngine, saveMemoryEngine } from './memory-engine.mjs'
44
44
  import { captureEvidence, createEvidenceStore, evidenceSummary } from './evidence-store.mjs'
45
- import { saveMissionState } from './mission-resume.mjs'
45
+ import { evidencePass, extractHarnessEvidence } from './evidence.mjs'
46
+ import { listMissionStates, loadMissionState, saveMissionState } from './mission-resume.mjs'
46
47
  import { collectResults, formatResultSummary, importBenchmarkRecord, missingTaskIds, summarizeResults } from './benchmark-results.mjs'
47
48
  import { buildAstGraph, collectSourceFiles } from './ast-provider.mjs'
48
49
 
@@ -695,6 +696,7 @@ export function apply(ctx, config = {}) {
695
696
  uncertainty: Number(config.contextUncertainty) || 0.6,
696
697
  maxFiles,
697
698
  maxFileChars: 2000,
699
+ maxContextTokens: ctxBudget.retrievalBudget,
698
700
  graph: graphAdj,
699
701
  })
700
702
  return dynamic.context
@@ -1359,6 +1361,113 @@ export function apply(ctx, config = {}) {
1359
1361
  },
1360
1362
  })
1361
1363
 
1364
+ async function createMissionExecutor({ session, agent, subagents, task, taskType, frontend, brief, capabilityBrain, evidenceRecords, resumeKey }) {
1365
+ const cwd = session?.meta?.cwd || session?.header?.cwd
1366
+ let brain = capabilityBrain
1367
+
1368
+ const saveProgress = async (snapshot = {}) => {
1369
+ if (!cwd || !resumeKey) return
1370
+ const dag = snapshot.dag
1371
+ saveMissionState(cwd, resumeKey, {
1372
+ status: dag ? (isMissionDagComplete(dag) ? 'completed' : 'active') : 'active',
1373
+ dag,
1374
+ evidence: evidenceRecords,
1375
+ metrics: snapshot.metrics || {},
1376
+ capabilityBrain: brain,
1377
+ brief,
1378
+ task,
1379
+ taskType,
1380
+ savedAt: new Date().toISOString(),
1381
+ })
1382
+ }
1383
+
1384
+ const act = async (action) => {
1385
+ const goal = action.task?.goal || action.taskId || ''
1386
+ const visualQa = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
1387
+ ? `\n\n${buildVisualQaStepRequirement()}`
1388
+ : ''
1389
+ const briefText = `\n\nTask brief:\nObjective: ${brief.objective}\nAcceptance: ${brief.acceptanceCriteria.join('; ')}`
1390
+ const taskContext = await getProjectContext(session, taskType, goal).catch(() => '')
1391
+ const prompt = `Mission: ${task}\nTask: ${action.taskId} — ${goal}\n\nExecute this step. Reply with a short result and evidence.${briefText}${taskContext ? `\n\nContext:\n${taskContext}` : ''}${visualQa}`
1392
+ const role = roleForTask(action.task || {})
1393
+ const sandbox = capabilityToolFilter(brain, action.task?.requiredCapabilities || [], role)
1394
+ const run = await subagents.start('spawn', {
1395
+ label: `mission-${action.taskId}`,
1396
+ prompt: [{ type: 'text', text: prompt }],
1397
+ parent: agent,
1398
+ maxDepth: 1,
1399
+ ...(sandbox.allow.length || sandbox.deny.length ? { toolFilter: sandbox } : {}),
1400
+ })
1401
+ const result = await run.result
1402
+ const output = (Array.isArray(result.output) ? result.output : [])
1403
+ .filter((block) => block?.type === 'text' && typeof block.text === 'string')
1404
+ .map((block) => block.text)
1405
+ .join('')
1406
+ try { await run.dispose() } catch { /* best-effort */ }
1407
+
1408
+ const harness = extractHarnessEvidence({ ...result, output })
1409
+ const hasStructured = harness.commands.length > 0 || harness.tests.length > 0 || harness.files.length > 0 || harness.findings.length > 0
1410
+ let record
1411
+ if (hasStructured) {
1412
+ record = {
1413
+ id: `E-${Date.now().toString(36)}-${Math.random().toString(36).slice(2, 6)}`,
1414
+ type: 'harness',
1415
+ source: action.taskId,
1416
+ value: output.slice(0, 2000),
1417
+ ok: evidencePass(harness),
1418
+ evidence: harness,
1419
+ at: new Date().toISOString(),
1420
+ }
1421
+ } else {
1422
+ const captured = captureEvidence({ entries: evidenceRecords }, {
1423
+ type: 'agent_output',
1424
+ source: action.taskId,
1425
+ value: output.slice(0, 2000),
1426
+ ok: !/FAIL|error|失败|not ok/i.test(output),
1427
+ })
1428
+ record = captured.record
1429
+ }
1430
+ evidenceRecords.push(record)
1431
+ return {
1432
+ output,
1433
+ evidenceId: record.id,
1434
+ evidence: harness,
1435
+ hasStructuredEvidence: hasStructured,
1436
+ tokenUsage: result?.tokenUsage ?? result?.usage?.totalTokens ?? 0,
1437
+ cost: result?.cost ?? 0,
1438
+ toolCalls: result?.toolCalls ?? result?.usage?.toolCalls ?? 0,
1439
+ }
1440
+ }
1441
+
1442
+ const observe = async (result, _current, action) => {
1443
+ const goal = action?.task?.goal || ''
1444
+ const needsVisual = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
1445
+ if (needsVisual && !/VISUAL_QA:\s*PASS/i.test(result.output || '')) return { type: 'test_failure', reason: 'visual QA failed' }
1446
+
1447
+ const harness = result.evidence || {}
1448
+ const hasStructured = result.hasStructuredEvidence === true || harness.commands?.length || harness.tests?.length || harness.files?.length || harness.findings?.length
1449
+ if (hasStructured) {
1450
+ const passed = evidencePass(harness)
1451
+ if (!passed) {
1452
+ brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], false)
1453
+ return { type: 'test_failure', reason: 'structured harness evidence failed', evidence: harness }
1454
+ }
1455
+ brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], true)
1456
+ return { type: 'step_done', evidence: harness }
1457
+ }
1458
+
1459
+ if (/FAIL|error|失败|not ok/i.test(result.output || '')) {
1460
+ const failure = classifyFailure({ type: 'unknown', reason: result.output || '' })
1461
+ brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], false)
1462
+ return { type: failure.category, reason: failure.recovery, detail: String(result.output || '').slice(0, 500) }
1463
+ }
1464
+ brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], true)
1465
+ return { type: 'step_done' }
1466
+ }
1467
+
1468
+ return { act, observe, saveProgress, getCapabilityBrain: () => brain }
1469
+ }
1470
+
1362
1471
  registerTool({
1363
1472
  name: 'omni_mission_run',
1364
1473
  description: 'Run a Mission Planner loop with real subagents: Observe → Think → Act → Replan until completed or maxSteps.',
@@ -1397,56 +1506,25 @@ export function apply(ctx, config = {}) {
1397
1506
  capabilityBrain = loadCapabilityManifests(capabilityBrain, config.capabilityManifests || [])
1398
1507
  const dag = bindCapabilitiesToDag(generateMissionDag(mission, brief, { taskType }), capabilityBrain)
1399
1508
  const evidenceRecords = []
1509
+ const cwd = session?.meta?.cwd || session?.header?.cwd
1510
+ const resumeKey = cwd ? `mission-${Date.now().toString(36)}` : null
1511
+ const executor = await createMissionExecutor({
1512
+ session,
1513
+ agent,
1514
+ subagents,
1515
+ task,
1516
+ taskType,
1517
+ frontend,
1518
+ brief,
1519
+ capabilityBrain,
1520
+ evidenceRecords,
1521
+ resumeKey,
1522
+ })
1400
1523
 
1401
1524
  const finalDag = await runDagLoop(dag, {
1402
- act: async (action) => {
1403
- const goal = action.task?.goal || action.taskId || ''
1404
- const visualQa = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
1405
- ? `\n\n${buildVisualQaStepRequirement()}`
1406
- : ''
1407
- const briefText = `\n\nTask brief:\nObjective: ${brief.objective}\nAcceptance: ${brief.acceptanceCriteria.join('; ')}`
1408
- const taskContext = await getProjectContext(session, taskType, goal).catch(() => '')
1409
- const prompt = `Mission: ${task}\nTask: ${action.taskId} — ${goal}\n\nExecute this step. Reply with a short result and evidence.${briefText}${taskContext ? `\n\nContext:\n${taskContext}` : ''}${visualQa}`
1410
- const role = roleForTask(action.task || {})
1411
- const sandbox = capabilityToolFilter(capabilityBrain, action.task?.requiredCapabilities || [], role)
1412
- const run = await subagents.start('spawn', {
1413
- label: `mission-${action.taskId}`,
1414
- prompt: [{ type: 'text', text: prompt }],
1415
- parent: agent,
1416
- maxDepth: 1,
1417
- ...(sandbox.allow.length || sandbox.deny.length ? { toolFilter: sandbox } : {}),
1418
- })
1419
- const result = await run.result
1420
- const output = (Array.isArray(result.output) ? result.output : [])
1421
- .filter((block) => block?.type === 'text' && typeof block.text === 'string')
1422
- .map((block) => block.text)
1423
- .join('')
1424
- try { await run.dispose() } catch { /* best-effort */ }
1425
- const captured = captureEvidence({ entries: evidenceRecords }, {
1426
- type: 'agent_output',
1427
- source: action.taskId,
1428
- value: output.slice(0, 2000),
1429
- ok: !/FAIL|error|失败|not ok/i.test(output),
1430
- })
1431
- evidenceRecords.push(captured.record)
1432
- return {
1433
- output,
1434
- evidenceId: captured.record.id,
1435
- tokenUsage: result?.tokenUsage ?? result?.usage?.totalTokens ?? 0,
1436
- cost: result?.cost ?? 0,
1437
- toolCalls: result?.toolCalls ?? result?.usage?.toolCalls ?? 0,
1438
- }
1439
- },
1440
- observe: async (result, _current, action) => {
1441
- const goal = action?.task?.goal || ''
1442
- const needsVisual = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
1443
- if (needsVisual && !/VISUAL_QA:\s*PASS/i.test(result.output || '')) return { type: 'test_failure', reason: 'visual QA failed' }
1444
- if (/FAIL|error|失败|not ok/i.test(result.output || '')) {
1445
- const failure = classifyFailure({ type: 'unknown', reason: result.output || '' })
1446
- return { type: failure.category, reason: failure.recovery, detail: String(result.output || '').slice(0, 500) }
1447
- }
1448
- return { type: 'step_done' }
1449
- },
1525
+ act: executor.act,
1526
+ observe: executor.observe,
1527
+ onProgress: executor.saveProgress,
1450
1528
  maxSteps,
1451
1529
  maxParallel: Number(args?.maxParallel) || 1,
1452
1530
  maxTokens: Number(args?.maxTokens) || 0,
@@ -1456,19 +1534,21 @@ export function apply(ctx, config = {}) {
1456
1534
  maxReplans: Number(args?.maxReplans) || 0,
1457
1535
  maxWallClockMs: Number(args?.maxWallClockMs) || 0,
1458
1536
  })
1537
+ capabilityBrain = executor.getCapabilityBrain()
1459
1538
 
1460
1539
  const metrics = finalDag.metrics || {}
1461
1540
  const evSummary = evidenceSummary({ entries: evidenceRecords })
1462
- const cwd = session?.meta?.cwd || session?.header?.cwd
1463
1541
  let resumeInfo = ''
1464
- if (cwd) {
1465
- const resumeKey = `mission-${Date.now().toString(36)}`
1542
+ if (cwd && resumeKey) {
1466
1543
  const saved = saveMissionState(cwd, resumeKey, {
1467
1544
  status: finalDag.status,
1468
1545
  dag: finalDag.dag,
1469
1546
  evidence: evidenceRecords,
1470
1547
  metrics,
1548
+ capabilityBrain,
1549
+ brief,
1471
1550
  task,
1551
+ taskType,
1472
1552
  savedAt: new Date().toISOString(),
1473
1553
  })
1474
1554
  resumeInfo = `\nSaved mission state: ${saved}`
@@ -1487,6 +1567,107 @@ export function apply(ctx, config = {}) {
1487
1567
  },
1488
1568
  })
1489
1569
 
1570
+ registerTool({
1571
+ name: 'omni_mission_resume',
1572
+ description: 'Resume a saved Mission DAG from .omni/missions/<key>.json. Without a key, lists saved missions.',
1573
+ parameters: {
1574
+ type: 'object',
1575
+ properties: {
1576
+ key: { type: 'string', description: 'Saved mission state key (from omni_mission_run output)' },
1577
+ maxSteps: { type: 'number', description: 'Max additional loop steps (default 20)' },
1578
+ maxParallel: { type: 'number', description: 'Max tasks to run concurrently (default 1)' },
1579
+ maxTokens: { type: 'number', description: 'Max token budget for the resumed portion (0 = unlimited)' },
1580
+ maxCost: { type: 'number', description: 'Max cost budget for the resumed portion (0 = unlimited)' },
1581
+ maxToolCalls: { type: 'number', description: 'Max tool calls for the resumed portion (0 = unlimited)' },
1582
+ maxRepairs: { type: 'number', description: 'Max repairs before blocking (0 = unlimited)' },
1583
+ maxReplans: { type: 'number', description: 'Max replans before blocking (0 = unlimited)' },
1584
+ maxWallClockMs: { type: 'number', description: 'Max wall-clock time in ms (0 = unlimited)' },
1585
+ },
1586
+ required: [],
1587
+ },
1588
+ async execute(args) {
1589
+ const session = currentSession()
1590
+ const agent = session && agentFor(session)
1591
+ const subagents = ctx.get('subagents') || ctx.subagents
1592
+ const cwd = session?.meta?.cwd || session?.header?.cwd
1593
+ if (!cwd) return 'No workspace cwd found.'
1594
+ const key = String(args?.key || '').trim()
1595
+ if (!key) {
1596
+ const keys = listMissionStates(cwd)
1597
+ return keys.length
1598
+ ? `Saved missions:\n${keys.map((k) => `- ${k}`).join('\n')}\n\nPass "key" to resume one.`
1599
+ : 'No saved missions found in .omni/missions.'
1600
+ }
1601
+ const saved = loadMissionState(cwd, key)
1602
+ if (!saved) {
1603
+ const keys = listMissionStates(cwd)
1604
+ return `No saved mission found for key "${key}". Available: ${keys.join(', ') || '(none)'}`
1605
+ }
1606
+ if (saved.status === 'completed') return `Mission "${key}" is already completed.`
1607
+ if (!session || !agent || !subagents?.start) return 'Mission resume requires an active session with subagents.'
1608
+
1609
+ const task = saved.task || saved.dag?.mission?.task || ''
1610
+ const taskType = saved.taskType || classifyTaskType(task)
1611
+ const frontend = isFrontendTask(task)
1612
+ const brief = saved.brief || { objective: task, acceptanceCriteria: [] }
1613
+ const evidenceRecords = Array.isArray(saved.evidence) ? saved.evidence : []
1614
+ const toolsService = ctx.get('tools') || ctx.tools
1615
+ let capabilityBrain = saved.capabilityBrain || autoPopulateCapabilityBrain(createCapabilityBrain(), await collectToolNames(toolsService))
1616
+ capabilityBrain = loadCapabilityManifests(capabilityBrain, config.capabilityManifests || [])
1617
+
1618
+ const executor = await createMissionExecutor({
1619
+ session,
1620
+ agent,
1621
+ subagents,
1622
+ task,
1623
+ taskType,
1624
+ frontend,
1625
+ brief,
1626
+ capabilityBrain,
1627
+ evidenceRecords,
1628
+ resumeKey: key,
1629
+ })
1630
+
1631
+ const finalDag = await runDagLoop(saved.dag, {
1632
+ act: executor.act,
1633
+ observe: executor.observe,
1634
+ onProgress: executor.saveProgress,
1635
+ maxSteps: Number(args?.maxSteps) || 20,
1636
+ maxParallel: Number(args?.maxParallel) || 1,
1637
+ maxTokens: Number(args?.maxTokens) || 0,
1638
+ maxCost: Number(args?.maxCost) || 0,
1639
+ maxToolCalls: Number(args?.maxToolCalls) || 0,
1640
+ maxRepairs: Number(args?.maxRepairs) || 0,
1641
+ maxReplans: Number(args?.maxReplans) || 0,
1642
+ maxWallClockMs: Number(args?.maxWallClockMs) || 0,
1643
+ })
1644
+
1645
+ const metrics = finalDag.metrics || {}
1646
+ saveMissionState(cwd, key, {
1647
+ status: finalDag.status,
1648
+ dag: finalDag.dag,
1649
+ evidence: evidenceRecords,
1650
+ metrics,
1651
+ capabilityBrain: executor.getCapabilityBrain(),
1652
+ brief,
1653
+ task,
1654
+ taskType,
1655
+ savedAt: new Date().toISOString(),
1656
+ })
1657
+ const evSummary = evidenceSummary({ entries: evidenceRecords })
1658
+ return [
1659
+ `Mission resume: ${finalDag.status}`,
1660
+ `Tasks done: ${finalDag.dag.tasks.filter((t) => t.status === 'done').length}/${finalDag.dag.tasks.length}`,
1661
+ `Additional steps: ${finalDag.actions.length}`,
1662
+ `Replans: ${metrics.replanCount || 0}, Repairs: ${metrics.repairCount || 0}, ToolCalls: ${metrics.toolCalls || 0}, Tokens: ${metrics.tokenUsage || 0}, Cost: ${metrics.cost || 0}`,
1663
+ `Evidence captured: ${evSummary.total} (failed=${evSummary.failed})`,
1664
+ `Mission: ${task}`,
1665
+ '',
1666
+ formatMissionDag(finalDag.dag),
1667
+ ].join('\n')
1668
+ },
1669
+ })
1670
+
1490
1671
  registerTool({
1491
1672
  name: 'omni_plan',
1492
1673
  description: 'Force the current session into plan-first mode (enter built-in plan mode).',
@@ -146,3 +146,20 @@ test('runDagLoop respects maxReplans', async () => {
146
146
  })
147
147
  assert.equal(result.status, 'max_replans')
148
148
  })
149
+
150
+ test('runDagLoop calls onProgress after each batch', async () => {
151
+ const dag = createMissionDag(buildMission('实现退款', { taskType: 'feature' }))
152
+ const snapshots = []
153
+ const result = await runDagLoop(dag, {
154
+ act: async () => ({ ok: true }),
155
+ observe: async () => ({ type: 'step_done' }),
156
+ maxSteps: 100,
157
+ maxParallel: 1,
158
+ onProgress: async (snapshot) => {
159
+ snapshots.push(snapshot)
160
+ },
161
+ })
162
+ assert.equal(result.status, 'completed')
163
+ assert.ok(snapshots.length >= dag.tasks.length)
164
+ assert.ok(snapshots.every((s) => s.dag && Array.isArray(s.actions)))
165
+ })
@@ -36,3 +36,14 @@ test('buildDynamicContext expands when uncertainty is high', () => {
36
36
  assert.equal(result.level, 2)
37
37
  assert.match(result.context, /Implementations:/)
38
38
  })
39
+
40
+ test('buildDynamicContext enforces maxContextTokens budget', () => {
41
+ const result = buildDynamicContext('fix login', entries, files, {
42
+ level: 4,
43
+ uncertainty: 0.8,
44
+ threshold: 0.5,
45
+ maxContextTokens: 10,
46
+ })
47
+ assert.ok(result.context.length <= 10 * 4 + 64)
48
+ assert.match(result.context, /truncated by context budget/)
49
+ })
@@ -8,6 +8,7 @@ import {
8
8
  addTestEvidence,
9
9
  createEvidence,
10
10
  evidencePass,
11
+ extractHarnessEvidence,
11
12
  summarizeEvidence,
12
13
  } from '../src/evidence.mjs'
13
14
 
@@ -44,3 +45,37 @@ test('summarizeEvidence renders sections', () => {
44
45
  assert.match(text, /Commands/)
45
46
  assert.match(text, /Tests/)
46
47
  })
48
+
49
+ test('extractHarnessEvidence normalizes structured result evidence', () => {
50
+ const evidence = extractHarnessEvidence({
51
+ commands: [{ command: 'npm test', exitCode: 0, output: 'ok' }],
52
+ tests: [{ command: 'npm test', exitCode: 0, total: 3, passed: 3, failed: 0 }],
53
+ files: [{ file: 'src/a.ts', lines: [1], beforeHash: 'a', afterHash: 'b', diffHash: 'c' }],
54
+ findings: [{ finding: 'nit', severity: 'low', file: 'src/a.ts', line: 2 }],
55
+ })
56
+ assert.equal(evidence.commands.length, 1)
57
+ assert.equal(evidence.tests.length, 1)
58
+ assert.equal(evidence.files.length, 1)
59
+ assert.equal(evidence.findings.length, 1)
60
+ assert.equal(evidencePass(evidence), true)
61
+ })
62
+
63
+ test('extractHarnessEvidence reads toolCalls as command evidence', () => {
64
+ const evidence = extractHarnessEvidence({
65
+ toolCalls: [
66
+ { name: 'npm test', exitCode: 0, output: 'all pass' },
67
+ { name: 'node run.js', exitCode: 1, output: 'boom' },
68
+ ],
69
+ })
70
+ assert.equal(evidence.commands.length, 2)
71
+ assert.equal(evidence.commands[0].exitCode, 0)
72
+ assert.equal(evidence.commands[1].exitCode, 1)
73
+ assert.equal(evidencePass(evidence), false)
74
+ })
75
+
76
+ test('extractHarnessEvidence parses embedded EVIDENCE_JSON block', () => {
77
+ const output = 'I ran the checks.\nEVIDENCE_JSON\n{"tests":[{"command":"npm test","exitCode":0,"total":2,"passed":2,"failed":0}]}'
78
+ const evidence = extractHarnessEvidence({ output })
79
+ assert.equal(evidence.tests.length, 1)
80
+ assert.equal(evidencePass(evidence), true)
81
+ })
@@ -37,6 +37,38 @@ test('retrieveHistoricalFailures filters by task text', () => {
37
37
  assert.equal(hits.length, 1)
38
38
  })
39
39
 
40
+ test('retrieveHistoricalFailures ranks by token overlap instead of exact substring only', () => {
41
+ const memory = createMemoryV3({
42
+ failures: [
43
+ { text: 'login timeout under concurrency', reason: 'session refresh race' },
44
+ { text: 'payment rounding', reason: 'float math' },
45
+ ],
46
+ })
47
+ const hits = retrieveHistoricalFailures(memory, 'session refresh race timeout')
48
+ assert.ok(hits.length >= 1)
49
+ assert.match(hits[0].text, /login timeout/)
50
+ })
51
+
52
+ test('distillSkill can require evidence-backed successes before promoting', () => {
53
+ const memory = createMemoryV3()
54
+ const noEvidence = distillSkill(
55
+ memory,
56
+ [{ text: 'step ok' }, { text: 'step ok' }, { text: 'step ok' }],
57
+ { name: 'no-evidence', minSuccesses: 3, requireEvidence: true },
58
+ )
59
+ assert.equal(noEvidence.promoted, false)
60
+ const withEvidence = distillSkill(
61
+ memory,
62
+ [
63
+ { text: 'step ok', evidence: { tests: [{ command: 'npm test', exitCode: 0, total: 1, passed: 1, failed: 0 }] } },
64
+ { text: 'step ok', evidence: { commands: [{ command: 'npm test', exitCode: 0 }] } },
65
+ { text: 'step ok', evidence: { tests: [{ command: 'npm test', exitCode: 0, total: 1, passed: 1, failed: 0 }] } },
66
+ ],
67
+ { name: 'evidence-skill', minSuccesses: 3, requireEvidence: true },
68
+ )
69
+ assert.equal(withEvidence.promoted, true)
70
+ })
71
+
40
72
  test('execution policies and cross-session strategies persist in memory', () => {
41
73
  let memory = createMemoryV3()
42
74
  memory = recordExecutionPolicy(memory, { name: 'schema-change', trigger: 'migration' })
@@ -0,0 +1,68 @@
1
+ import test from 'node:test'
2
+ import assert from 'node:assert/strict'
3
+ import fs from 'node:fs'
4
+ import os from 'node:os'
5
+ import path from 'node:path'
6
+
7
+ import {
8
+ buildPrompt,
9
+ readManifests,
10
+ runCommand,
11
+ validateManifests,
12
+ writeResults,
13
+ } from '../benchmark/omnibench-v2/run.mjs'
14
+
15
+ test('buildPrompt includes task and BENCHMARK marker', () => {
16
+ const manifest = { task: 'Fix pagination', acceptance: ['all tests pass'] }
17
+ const raw = buildPrompt(manifest, 'raw')
18
+ assert.match(raw, /Fix pagination/)
19
+ assert.match(raw, /BENCHMARK: PASS/)
20
+ const omni = buildPrompt(manifest, 'omni')
21
+ assert.match(omni, /Omni control plane/)
22
+ })
23
+
24
+ test('validateManifests returns errors for invalid manifest', () => {
25
+ const errors = validateManifests([{ id: 'x' }])
26
+ assert.ok(errors.length >= 3)
27
+ assert.equal(validateManifests([{
28
+ id: 'x',
29
+ repo: 'r',
30
+ commit: 'c',
31
+ task: 't',
32
+ acceptance: ['a'],
33
+ runs: 3,
34
+ }]).length, 0)
35
+ })
36
+
37
+ test('readManifests supports single object and array', () => {
38
+ const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'omnibench-read-'))
39
+ try {
40
+ const file = path.join(dir, 'manifest.json')
41
+ fs.writeFileSync(file, JSON.stringify({ id: 'one', repo: 'r', commit: 'c', task: 't', acceptance: ['a'], runs: 3 }))
42
+ assert.equal(readManifests(file).length, 1)
43
+ fs.writeFileSync(file, JSON.stringify([{ id: 'one', repo: 'r', commit: 'c', task: 't', acceptance: ['a'], runs: 3 }]))
44
+ assert.equal(readManifests(file).length, 1)
45
+ } finally {
46
+ fs.rmSync(dir, { recursive: true, force: true })
47
+ }
48
+ })
49
+
50
+ test('runCommand captures exit code and output', () => {
51
+ const ok = runCommand('node -e "console.log(1)"', process.cwd())
52
+ assert.equal(ok.exitCode, 0)
53
+ assert.match(ok.output, /1/)
54
+ const bad = runCommand('node -e "process.exit(3)"', process.cwd())
55
+ assert.equal(bad.exitCode, 3)
56
+ })
57
+
58
+ test('writeResults writes JSON result file', () => {
59
+ const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'omnibench-write-'))
60
+ try {
61
+ const file = writeResults([{ id: 'x', success: true }], dir)
62
+ assert.ok(fs.existsSync(file))
63
+ const parsed = JSON.parse(fs.readFileSync(file, 'utf8'))
64
+ assert.equal(parsed[0].success, true)
65
+ } finally {
66
+ fs.rmSync(dir, { recursive: true, force: true })
67
+ }
68
+ })