dsh-omni-router 2.0.0 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +3 -2
- package/README.zh-CN.md +3 -2
- package/benchmark/omnibench-v2/README.md +33 -10
- package/benchmark/omnibench-v2/run.mjs +166 -11
- package/package.json +3 -2
- package/src/agent-runtime.mjs +13 -0
- package/src/context-expansion.mjs +9 -1
- package/src/dynamic-context.mjs +1 -0
- package/src/evidence.mjs +86 -0
- package/src/memory-v3.mjs +37 -5
- package/src/omni-router.mjs +235 -54
- package/test/agent-runtime.test.mjs +17 -0
- package/test/dynamic-context.test.mjs +11 -0
- package/test/evidence.test.mjs +35 -0
- package/test/memory-v3.test.mjs +32 -0
- package/test/omnibench-runner.test.mjs +68 -0
package/README.md
CHANGED
|
@@ -71,7 +71,8 @@
|
|
|
71
71
|
- **Evidence Engine / Write Locks**:`src/evidence-engine.mjs` 统一证据协议与存储;DAG scheduler 通过 `selectReadyBatch` 避免 write overlap 并行冲突。
|
|
72
72
|
- **Stage 3/5**:Recursive Project Index、contextBudget 接入、Per-task Context、Missing capability handling、OmniBench v2 prompt generator。
|
|
73
73
|
- **Stage 4(部分)**:`compileDagToWorkflow` 将 Mission DAG 编译为 DSH-native workflow outline。
|
|
74
|
-
- **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2
|
|
74
|
+
- **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 runner(manifest schema + prepare + prompt generation + `--exec` 本地执行/结果采集 + DAG-to-plan compile)。
|
|
75
|
+
- **2.1 收敛增强**:Mission Resume(`omni_mission_resume` 跨 session 继续 DAG,`runDagLoop` 支持 `onProgress` 周期保存)、Harness evidence extraction(`extractHarnessEvidence` 从 commands/tests/toolCalls/EVIDENCE_JSON 提取机器可验证证据)、context token budget(`maxContextTokens` 真正限制检索输出)、semantic failure retrieval + evidence-backed skill distillation、OmniBench v2 real runner。
|
|
75
76
|
- **TaskDecision**:`createTaskDecision` 生成唯一决策对象,Policy/Runtime 统一消费,避免重复 classify 分叉。
|
|
76
77
|
- **Evidence Protocol**:`src/evidence.mjs` 用结构化 command/file/test/finding 证据判定 PASS/FAIL;`isQaPass` 已支持从 QA 输出解析 JSON evidence,不再只信文本。
|
|
77
78
|
- **Task Compiler**:`compileTask` 生成 objective / constraints / non-goals / acceptance / hidden assumptions / ambiguities / invariants / risk / artifacts;`compileTaskWithLLM` 可用 LLM 增强。
|
|
@@ -153,7 +154,7 @@ Then restart DSH and select **Omni Router** in a new session.
|
|
|
153
154
|
- `/omni direct` — enter direct mode.
|
|
154
155
|
- `/omni mode spec|react|balanced` — set thinking mode.
|
|
155
156
|
- `/omni reroute plan|direct` — adaptively reroute current task.
|
|
156
|
-
- Model tools: `omni_status`, `omni_plan`, `omni_direct`, `omni_mode`, `omni_reroute`, `omni_delegate`, `omni_memory`, `omni_benchmark`, `omni_mission_run`, `omni_visual_check`.
|
|
157
|
+
- Model tools: `omni_status`, `omni_plan`, `omni_direct`, `omni_mode`, `omni_reroute`, `omni_delegate`, `omni_memory`, `omni_benchmark`, `omni_mission_run`, `omni_mission_resume`, `omni_visual_check`.
|
|
157
158
|
|
|
158
159
|
## Configuration
|
|
159
160
|
|
package/README.zh-CN.md
CHANGED
|
@@ -71,7 +71,8 @@
|
|
|
71
71
|
- **Evidence Engine / Write Locks**:`src/evidence-engine.mjs` 统一证据协议与存储;DAG scheduler 通过 `selectReadyBatch` 避免 write overlap 并行冲突。
|
|
72
72
|
- **Stage 3/5**:Recursive Project Index、contextBudget 接入、Per-task Context、Missing capability handling、OmniBench v2 prompt generator。
|
|
73
73
|
- **Stage 4(部分)**:`compileDagToWorkflow` 将 Mission DAG 编译为 DSH-native workflow outline。
|
|
74
|
-
- **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2
|
|
74
|
+
- **P1/P3 能力**:Real capability sandbox(已接入 `omni_mission_run`,按 role + capability 生成 toolFilter)、Capability Resolver v2(risk/reliability 综合打分)、OmniBench v2 runner(manifest schema + prepare + prompt generation + `--exec` 本地执行/结果采集 + DAG-to-plan compile)。
|
|
75
|
+
- **2.1 收敛增强**:Mission Resume(`omni_mission_resume` 跨 session 继续 DAG,`runDagLoop` 支持 `onProgress` 周期保存)、Harness evidence extraction(`extractHarnessEvidence` 从 commands/tests/toolCalls/EVIDENCE_JSON 提取机器可验证证据)、context token budget(`maxContextTokens` 真正限制检索输出)、semantic failure retrieval + evidence-backed skill distillation、OmniBench v2 real runner。
|
|
75
76
|
- **TaskDecision**:`createTaskDecision` 生成唯一决策对象,Policy/Runtime 统一消费,避免重复 classify 分叉。
|
|
76
77
|
- **Evidence Protocol**:`src/evidence.mjs` 用结构化 command/file/test/finding 证据判定 PASS/FAIL;`isQaPass` 已支持从 QA 输出解析 JSON evidence,不再只信文本。
|
|
77
78
|
- **Task Compiler**:`compileTask` 生成 objective / constraints / non-goals / acceptance / hidden assumptions / ambiguities / invariants / risk / artifacts;`compileTaskWithLLM` 可用 LLM 增强。
|
|
@@ -153,7 +154,7 @@ node scripts/install-preset.mjs
|
|
|
153
154
|
- `/omni direct` — 进入直接执行模式。
|
|
154
155
|
- `/omni mode spec|react|balanced` — 设置思维模式。
|
|
155
156
|
- `/omni reroute plan|direct` — 动态切换当前任务路由。
|
|
156
|
-
- 模型工具:`omni_status` / `omni_plan` / `omni_direct` / `omni_mode` / `omni_reroute` / `omni_delegate` / `omni_memory` / `omni_benchmark` / `omni_mission_run` / `omni_visual_check`。
|
|
157
|
+
- 模型工具:`omni_status` / `omni_plan` / `omni_direct` / `omni_mode` / `omni_reroute` / `omni_delegate` / `omni_memory` / `omni_benchmark` / `omni_mission_run` / `omni_mission_resume` / `omni_visual_check`。
|
|
157
158
|
|
|
158
159
|
## 配置
|
|
159
160
|
|
|
@@ -1,8 +1,8 @@
|
|
|
1
|
-
# OmniBench v2
|
|
1
|
+
# OmniBench v2
|
|
2
2
|
|
|
3
3
|
Goal: move from 10 small fixture tasks to a reproducible multi-repo benchmark.
|
|
4
4
|
|
|
5
|
-
## Requirements (from 优化5.md)
|
|
5
|
+
## Requirements (from 优化5.md / 优化6.md)
|
|
6
6
|
|
|
7
7
|
- 30~100 real GitHub repos
|
|
8
8
|
- fixed starting commit
|
|
@@ -25,18 +25,18 @@ Goal: move from 10 small fixture tasks to a reproducible multi-repo benchmark.
|
|
|
25
25
|
"task": "Fix intermittent login timeout under concurrent session refresh.",
|
|
26
26
|
"acceptance": ["regression test added", "all tests pass"],
|
|
27
27
|
"hiddenTests": ["tests/hidden/session-concurrent.test.js"],
|
|
28
|
+
"setupCommand": "npm ci",
|
|
29
|
+
"baselineCommand": "npm test",
|
|
30
|
+
"agentCommand": "node path/to/dsh-agent-runner.mjs",
|
|
31
|
+
"verifyCommand": "npm run test:hidden",
|
|
28
32
|
"timeoutMs": 7200000,
|
|
29
33
|
"runs": 5
|
|
30
34
|
}
|
|
31
35
|
```
|
|
32
36
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
`node benchmark/omnibench-v2/runner-stub.mjs <manifest.json>` validates the
|
|
38
|
-
manifest and prints a run plan. Actual execution will be driven by DSH
|
|
39
|
-
subagents in a future integration.
|
|
37
|
+
`setupCommand`, `baselineCommand`, `agentCommand`, and `verifyCommand` are
|
|
38
|
+
optional. `agentCommand` receives the generated prompt as a single quoted
|
|
39
|
+
argument; when present, `--exec` executes it inside the checked-out repo.
|
|
40
40
|
|
|
41
41
|
## Generate DSH prompts
|
|
42
42
|
|
|
@@ -46,7 +46,29 @@ npm run omnibench:v2:generate -- benchmark/omnibench-v2/manifest.local.example.j
|
|
|
46
46
|
|
|
47
47
|
This writes one prompt per repo/arm/run under `benchmark/omnibench-v2/prompts/`.
|
|
48
48
|
Open each prompt in a DSH Desktop session with the Omni Router preset to
|
|
49
|
-
execute,
|
|
49
|
+
execute manually, or wire `agentCommand` to your DSH CLI/runner for automation.
|
|
50
|
+
|
|
51
|
+
## Run locally (automated)
|
|
52
|
+
|
|
53
|
+
```bash
|
|
54
|
+
# Plan + prompt generation only (default)
|
|
55
|
+
node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json
|
|
56
|
+
|
|
57
|
+
# Execute each run with the manifest's agentCommand
|
|
58
|
+
node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json --exec
|
|
59
|
+
|
|
60
|
+
# Override the agent command for every run
|
|
61
|
+
node benchmark/omnibench-v2/run.mjs benchmark/omnibench-v2/manifest.local.example.json --exec --agent-command "node agent.mjs"
|
|
62
|
+
```
|
|
63
|
+
|
|
64
|
+
The runner clones/checks out each repo, optionally runs setup/baseline/verify
|
|
65
|
+
commands, captures exit codes and output, and writes a timestamped JSON result
|
|
66
|
+
file to `benchmark/omnibench-v2/results/`.
|
|
67
|
+
|
|
68
|
+
## Runner stub
|
|
69
|
+
|
|
70
|
+
`node benchmark/omnibench-v2/runner-stub.mjs <manifest.json>` still validates
|
|
71
|
+
the manifest and prints the run plan for quick checks.
|
|
50
72
|
|
|
51
73
|
## Local / offline usage
|
|
52
74
|
|
|
@@ -57,6 +79,7 @@ existing `Omni group` fixture:
|
|
|
57
79
|
```bash
|
|
58
80
|
npm run omnibench:v2:prepare -- benchmark/omnibench-v2/manifest.local.example.json
|
|
59
81
|
npm run omnibench:v2:plan -- benchmark/omnibench-v2/manifest.local.example.json
|
|
82
|
+
npm run omnibench:v2:generate -- benchmark/omnibench-v2/manifest.local.example.json
|
|
60
83
|
```
|
|
61
84
|
|
|
62
85
|
`git clone` works with local paths, so no network is required. Replace `repo`
|
|
@@ -2,20 +2,24 @@
|
|
|
2
2
|
/**
|
|
3
3
|
* OmniBench v2 runner / prompt generator.
|
|
4
4
|
*
|
|
5
|
-
* Generates per-repo per-arm per-run DSH prompts.
|
|
6
|
-
*
|
|
5
|
+
* Generates per-repo per-arm per-run DSH prompts. With `--exec` and a
|
|
6
|
+
* configured `agentCommand` (or `--agent-command`), it also executes the run
|
|
7
|
+
* locally, runs baseline/verify commands, and writes machine-readable results.
|
|
7
8
|
*
|
|
8
9
|
* Usage:
|
|
9
|
-
* node benchmark/omnibench-v2/run.mjs <manifest.json>
|
|
10
|
+
* node benchmark/omnibench-v2/run.mjs <manifest.json> # plan + prompts
|
|
11
|
+
* node benchmark/omnibench-v2/run.mjs <manifest.json> --exec # execute agentCommand runs
|
|
12
|
+
* node benchmark/omnibench-v2/run.mjs <manifest.json> --exec --agent-command "node agent.mjs"
|
|
10
13
|
*/
|
|
11
14
|
|
|
12
15
|
import fs from 'node:fs'
|
|
13
16
|
import path from 'node:path'
|
|
14
|
-
import {
|
|
17
|
+
import { execSync } from 'node:child_process'
|
|
18
|
+
import { fileURLToPath, pathToFileURL } from 'node:url'
|
|
15
19
|
|
|
16
20
|
const here = path.dirname(fileURLToPath(import.meta.url))
|
|
17
21
|
|
|
18
|
-
function buildPrompt(manifest, arm) {
|
|
22
|
+
export function buildPrompt(manifest, arm) {
|
|
19
23
|
const criteria = (manifest.acceptance || []).map((c) => `- ${c}`).join('\n')
|
|
20
24
|
const task = `Task:\n${manifest.task}\n\nAcceptance criteria:\n${criteria}`
|
|
21
25
|
if (arm === 'raw') {
|
|
@@ -24,10 +28,54 @@ function buildPrompt(manifest, arm) {
|
|
|
24
28
|
return `${task}\n\nYou are using the Omni control plane. Follow engineering methodology, verify with real evidence, and do not claim completion without checks. When done, reply exactly "BENCHMARK: PASS" if you verified all criteria, otherwise "BENCHMARK: FAIL".`
|
|
25
29
|
}
|
|
26
30
|
|
|
27
|
-
function
|
|
28
|
-
const manifestPath = process.argv[2] || path.join(here, 'manifest.local.example.json')
|
|
31
|
+
export function readManifests(manifestPath) {
|
|
29
32
|
const manifests = JSON.parse(fs.readFileSync(manifestPath, 'utf8'))
|
|
30
|
-
|
|
33
|
+
return Array.isArray(manifests) ? manifests : [manifests]
|
|
34
|
+
}
|
|
35
|
+
|
|
36
|
+
export function validateManifests(list) {
|
|
37
|
+
const errors = []
|
|
38
|
+
for (const m of list) {
|
|
39
|
+
if (!m.id) errors.push('missing id')
|
|
40
|
+
if (!m.repo) errors.push(`${m.id || '?'}: missing repo`)
|
|
41
|
+
if (!m.commit) errors.push(`${m.id || '?'}: missing commit`)
|
|
42
|
+
if (!m.task) errors.push(`${m.id || '?'}: missing task`)
|
|
43
|
+
if (!Array.isArray(m.acceptance) || !m.acceptance.length) errors.push(`${m.id || '?'}: missing acceptance`)
|
|
44
|
+
if (!m.runs || m.runs < 3) errors.push(`${m.id || '?'}: runs should be >= 3`)
|
|
45
|
+
}
|
|
46
|
+
return errors
|
|
47
|
+
}
|
|
48
|
+
|
|
49
|
+
export function ensureRepo(m) {
|
|
50
|
+
const workDir = path.resolve(process.env.OMNIBENCH_REPOS || path.join(here, 'repos'))
|
|
51
|
+
const dir = path.join(workDir, m.id)
|
|
52
|
+
fs.mkdirSync(workDir, { recursive: true })
|
|
53
|
+
if (!fs.existsSync(path.join(dir, '.git'))) {
|
|
54
|
+
console.log(`cloning ${m.repo} -> ${dir}`)
|
|
55
|
+
execSync(`git clone "${m.repo}" "${dir}"`, { stdio: 'inherit' })
|
|
56
|
+
}
|
|
57
|
+
console.log(`checkout ${m.id} @ ${m.commit}`)
|
|
58
|
+
execSync(`git -C "${dir}" checkout ${m.commit}`, { stdio: 'inherit' })
|
|
59
|
+
return dir
|
|
60
|
+
}
|
|
61
|
+
|
|
62
|
+
export function runCommand(command, cwd, label) {
|
|
63
|
+
if (!command) return { skipped: true }
|
|
64
|
+
const start = Date.now()
|
|
65
|
+
try {
|
|
66
|
+
const output = execSync(command, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] })
|
|
67
|
+
return { skipped: false, exitCode: 0, output, durationMs: Date.now() - start }
|
|
68
|
+
} catch (error) {
|
|
69
|
+
return {
|
|
70
|
+
skipped: false,
|
|
71
|
+
exitCode: error.status ?? 1,
|
|
72
|
+
output: String(error.stdout || '') + String(error.stderr || ''),
|
|
73
|
+
durationMs: Date.now() - start,
|
|
74
|
+
}
|
|
75
|
+
}
|
|
76
|
+
}
|
|
77
|
+
|
|
78
|
+
export function generatePrompts(list) {
|
|
31
79
|
const outDir = path.join(here, 'prompts')
|
|
32
80
|
fs.mkdirSync(outDir, { recursive: true })
|
|
33
81
|
let count = 0
|
|
@@ -40,8 +88,115 @@ function main() {
|
|
|
40
88
|
}
|
|
41
89
|
}
|
|
42
90
|
}
|
|
43
|
-
|
|
44
|
-
|
|
91
|
+
return count
|
|
92
|
+
}
|
|
93
|
+
|
|
94
|
+
export function writeResults(results, resultsDir) {
|
|
95
|
+
fs.mkdirSync(resultsDir, { recursive: true })
|
|
96
|
+
const file = path.join(resultsDir, `omnibench-v2-${Date.now()}.json`)
|
|
97
|
+
fs.writeFileSync(file, JSON.stringify(results, null, 2), 'utf8')
|
|
98
|
+
return file
|
|
99
|
+
}
|
|
100
|
+
|
|
101
|
+
function main() {
|
|
102
|
+
const manifestPath = process.argv[2] || path.join(here, 'manifest.local.example.json')
|
|
103
|
+
const execMode = process.argv.includes('--exec')
|
|
104
|
+
const agentCommandOverride = (() => {
|
|
105
|
+
const idx = process.argv.indexOf('--agent-command')
|
|
106
|
+
return idx !== -1 ? process.argv[idx + 1] : null
|
|
107
|
+
})()
|
|
108
|
+
const resultsDir = path.resolve(process.env.OMNIBENCH_RESULTS || path.join(here, 'results'))
|
|
109
|
+
|
|
110
|
+
const list = readManifests(manifestPath)
|
|
111
|
+
const errors = validateManifests(list)
|
|
112
|
+
if (errors.length) {
|
|
113
|
+
console.error('Invalid manifest:')
|
|
114
|
+
for (const e of errors) console.error(`- ${e}`)
|
|
115
|
+
process.exit(1)
|
|
116
|
+
}
|
|
117
|
+
const promptCount = generatePrompts(list)
|
|
118
|
+
|
|
119
|
+
let totalRuns = 0
|
|
120
|
+
for (const m of list) totalRuns += (m.runs || 1) * 2
|
|
121
|
+
console.log(`Generated ${promptCount} DSH prompts in ${path.join(here, 'prompts')}`)
|
|
122
|
+
console.log(`OmniBench v2 run plan (${totalRuns} agent runs)`)
|
|
123
|
+
|
|
124
|
+
if (!execMode) {
|
|
125
|
+
for (const m of list) {
|
|
126
|
+
console.log(`- ${m.id}: ${m.repo} @ ${m.commit} (${m.language}/${m.framework}) runs=${m.runs} -> raw ${m.runs} + omni ${m.runs}`)
|
|
127
|
+
}
|
|
128
|
+
console.log('Pass --exec and configure agentCommand to execute runs locally.')
|
|
129
|
+
return
|
|
130
|
+
}
|
|
131
|
+
|
|
132
|
+
const results = []
|
|
133
|
+
for (const m of list) {
|
|
134
|
+
const repoDir = ensureRepo(m)
|
|
135
|
+
for (const arm of ['raw', 'omni']) {
|
|
136
|
+
for (let i = 1; i <= (m.runs || 3); i++) {
|
|
137
|
+
const agentCommand = agentCommandOverride || m.agentCommand
|
|
138
|
+
const promptFile = path.join(here, 'prompts', `${m.id}-${arm}-${i}.txt`)
|
|
139
|
+
const prompt = fs.readFileSync(promptFile, 'utf8')
|
|
140
|
+
console.log(`\n=== ${m.id} ${arm} run ${i} ===`)
|
|
141
|
+
if (m.setupCommand) {
|
|
142
|
+
const setup = runCommand(m.setupCommand, repoDir, 'setup')
|
|
143
|
+
console.log(`setup exit=${setup.exitCode ?? 'skip'}`)
|
|
144
|
+
}
|
|
145
|
+
const baseline = m.baselineCommand ? runCommand(m.baselineCommand, repoDir, 'baseline') : null
|
|
146
|
+
if (baseline) console.log(`baseline exit=${baseline.exitCode ?? 'skip'}`)
|
|
147
|
+
|
|
148
|
+
let agent = null
|
|
149
|
+
if (agentCommand) {
|
|
150
|
+
const command = `${agentCommand} ${JSON.stringify(prompt)}`
|
|
151
|
+
agent = runCommand(command, repoDir, 'agent')
|
|
152
|
+
console.log(`agent exit=${agent.exitCode ?? 'skip'} duration=${agent.durationMs ?? 0}ms`)
|
|
153
|
+
} else {
|
|
154
|
+
console.log(`agentCommand not configured; run prompt manually in DSH Desktop:\n ${promptFile}`)
|
|
155
|
+
}
|
|
156
|
+
|
|
157
|
+
const verify = m.verifyCommand ? runCommand(m.verifyCommand, repoDir, 'verify') : null
|
|
158
|
+
if (verify) console.log(`verify exit=${verify.exitCode ?? 'skip'}`)
|
|
159
|
+
|
|
160
|
+
let success = null
|
|
161
|
+
if (verify && verify.skipped !== true) {
|
|
162
|
+
success = verify.exitCode === 0
|
|
163
|
+
} else if (agent && agent.skipped !== true) {
|
|
164
|
+
success = agent.exitCode === 0 && /BENCHMARK:\s*PASS/i.test(agent.output || '')
|
|
165
|
+
}
|
|
166
|
+
|
|
167
|
+
results.push({
|
|
168
|
+
id: m.id,
|
|
169
|
+
arm,
|
|
170
|
+
run: i,
|
|
171
|
+
repo: m.repo,
|
|
172
|
+
commit: m.commit,
|
|
173
|
+
task: m.task,
|
|
174
|
+
success,
|
|
175
|
+
agentExitCode: agent?.exitCode ?? null,
|
|
176
|
+
verifyExitCode: verify?.exitCode ?? null,
|
|
177
|
+
durationMs: (agent?.durationMs || 0) + (verify?.durationMs || 0),
|
|
178
|
+
baselineExitCode: baseline?.exitCode ?? null,
|
|
179
|
+
agentOutput: agent?.output?.slice(0, 4000) || '',
|
|
180
|
+
verifyOutput: verify?.output?.slice(0, 4000) || '',
|
|
181
|
+
metrics: {
|
|
182
|
+
tokens: 0,
|
|
183
|
+
cost: 0,
|
|
184
|
+
toolCalls: 0,
|
|
185
|
+
},
|
|
186
|
+
promptFile,
|
|
187
|
+
ranAt: new Date().toISOString(),
|
|
188
|
+
})
|
|
189
|
+
}
|
|
190
|
+
}
|
|
191
|
+
}
|
|
192
|
+
|
|
193
|
+
const file = writeResults(results, resultsDir)
|
|
194
|
+
const executed = results.filter((r) => r.success !== null).length
|
|
195
|
+
const passed = results.filter((r) => r.success === true).length
|
|
196
|
+
console.log(`\nWrote ${results.length} result record(s) to ${file}`)
|
|
197
|
+
console.log(`Executed ${executed}/${results.length} runs; passed ${passed}.`)
|
|
45
198
|
}
|
|
46
199
|
|
|
47
|
-
|
|
200
|
+
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) {
|
|
201
|
+
main()
|
|
202
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "dsh-omni-router",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.1.0",
|
|
4
4
|
"description": "Omni - a reliability and orchestration control plane for DeepSeek Harness, designed to turn fast models into dependable long-horizon software-engineering agents.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"main": "./lib/index.js",
|
|
@@ -41,7 +41,7 @@
|
|
|
41
41
|
}
|
|
42
42
|
},
|
|
43
43
|
"scripts": {
|
|
44
|
-
"test": "node test/omni-router.test.mjs && node test/agent-chain.test.mjs && node test/skill-suggest.test.mjs && node test/methodology.test.mjs && node test/compat.test.mjs && node test/project-brain.test.mjs && node test/mission-planner.test.mjs && node test/agent-runtime.test.mjs && node test/judge.test.mjs && node test/memory.test.mjs && node test/memory-v3.test.mjs && node test/memory-engine.test.mjs && node test/mission-resume.test.mjs && node test/engineering-benchmark.test.mjs && node test/benchmark-runner.test.mjs && node test/benchmark-results.test.mjs && node test/continuous-eval.test.mjs && node --no-warnings test/project-brain-v2.test.mjs && node --no-warnings test/project-brain-v3.test.mjs && node test/visual-qa.test.mjs && node test/evidence.test.mjs && node test/evidence-store.test.mjs && node test/evidence-engine.test.mjs && node test/task-decision.test.mjs && node test/task-compiler.test.mjs && node test/capability-brain.test.mjs && node test/capability-manifest.test.mjs && node test/capability-sandbox.test.mjs && node test/planner-dag.test.mjs && node test/dynamic-context.test.mjs && node test/mission-dag.test.mjs && node test/failure-taxonomy.test.mjs && node test/context-expansion.test.mjs && node test/hybrid-retrieval.test.mjs && node test/dependency-graph.test.mjs && node test/ast-provider.test.mjs && node test/experience-router.test.mjs && node test/benchmark-tasks.test.mjs && node test/bundle-installer.test.mjs",
|
|
44
|
+
"test": "node test/omni-router.test.mjs && node test/agent-chain.test.mjs && node test/skill-suggest.test.mjs && node test/methodology.test.mjs && node test/compat.test.mjs && node test/project-brain.test.mjs && node test/mission-planner.test.mjs && node test/agent-runtime.test.mjs && node test/judge.test.mjs && node test/memory.test.mjs && node test/memory-v3.test.mjs && node test/memory-engine.test.mjs && node test/mission-resume.test.mjs && node test/engineering-benchmark.test.mjs && node test/benchmark-runner.test.mjs && node test/benchmark-results.test.mjs && node test/continuous-eval.test.mjs && node --no-warnings test/project-brain-v2.test.mjs && node --no-warnings test/project-brain-v3.test.mjs && node test/visual-qa.test.mjs && node test/evidence.test.mjs && node test/evidence-store.test.mjs && node test/evidence-engine.test.mjs && node test/task-decision.test.mjs && node test/task-compiler.test.mjs && node test/capability-brain.test.mjs && node test/capability-manifest.test.mjs && node test/capability-sandbox.test.mjs && node test/planner-dag.test.mjs && node test/dynamic-context.test.mjs && node test/mission-dag.test.mjs && node test/failure-taxonomy.test.mjs && node test/context-expansion.test.mjs && node test/hybrid-retrieval.test.mjs && node test/dependency-graph.test.mjs && node test/ast-provider.test.mjs && node test/experience-router.test.mjs && node test/benchmark-tasks.test.mjs && node test/bundle-installer.test.mjs && node test/omnibench-runner.test.mjs",
|
|
45
45
|
"benchmark": "node benchmark/run.mjs",
|
|
46
46
|
"benchmark:engineering": "node benchmark/engineering-run.mjs",
|
|
47
47
|
"benchmark:compare": "node benchmark/compare.mjs",
|
|
@@ -50,6 +50,7 @@
|
|
|
50
50
|
"omnibench:v2:plan": "node benchmark/omnibench-v2/runner-stub.mjs",
|
|
51
51
|
"omnibench:v2:prepare": "node benchmark/omnibench-v2/prepare.mjs",
|
|
52
52
|
"omnibench:v2:generate": "node benchmark/omnibench-v2/run.mjs",
|
|
53
|
+
"omnibench:v2:run": "node benchmark/omnibench-v2/run.mjs --exec",
|
|
53
54
|
"analyze": "node benchmark/analyze.mjs",
|
|
54
55
|
"llm-eval": "node benchmark/llm-eval.mjs",
|
|
55
56
|
"install:preset": "node scripts/install-preset.mjs"
|
package/src/agent-runtime.mjs
CHANGED
|
@@ -165,6 +165,7 @@ export async function runDagLoop(dag, {
|
|
|
165
165
|
maxRepairs = 0,
|
|
166
166
|
maxSameActionRetries = 0,
|
|
167
167
|
budget = {},
|
|
168
|
+
onProgress = null,
|
|
168
169
|
} = {}) {
|
|
169
170
|
const b = {
|
|
170
171
|
maxSteps: maxSteps ?? budget.steps ?? 50,
|
|
@@ -231,6 +232,18 @@ export async function runDagLoop(dag, {
|
|
|
231
232
|
actions.push({ taskId: task.id, observation })
|
|
232
233
|
}
|
|
233
234
|
step += 1
|
|
235
|
+
if (typeof onProgress === 'function') {
|
|
236
|
+
try {
|
|
237
|
+
await onProgress({
|
|
238
|
+
dag: current,
|
|
239
|
+
step,
|
|
240
|
+
actions,
|
|
241
|
+
metrics: { step, replanCount, repairCount, sameActionCount, tokenUsage, cost, toolCalls },
|
|
242
|
+
})
|
|
243
|
+
} catch {
|
|
244
|
+
// Progress callbacks are best-effort; they must never abort the mission.
|
|
245
|
+
}
|
|
246
|
+
}
|
|
234
247
|
}
|
|
235
248
|
|
|
236
249
|
const done = isMissionDagComplete(current)
|
|
@@ -54,7 +54,15 @@ export function buildProgressiveContext(taskText, entries, files = {}, options =
|
|
|
54
54
|
const tests = [...names].filter((n) => /\.(test|spec)\./i.test(n))
|
|
55
55
|
if (tests.length) parts.push('', 'Tests:', ...tests.map((t) => `- ${t}`))
|
|
56
56
|
}
|
|
57
|
-
return parts.join('\n')
|
|
57
|
+
return applyContextBudget(parts.join('\n'), options)
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
function applyContextBudget(text, options = {}) {
|
|
61
|
+
const maxTokens = Number(options.maxContextTokens) || 0
|
|
62
|
+
if (maxTokens <= 0) return text
|
|
63
|
+
const maxChars = Math.max(32, maxTokens * 4)
|
|
64
|
+
if (text.length <= maxChars) return text
|
|
65
|
+
return `${text.slice(0, maxChars)}\n… (truncated by context budget ${maxTokens} tokens)`
|
|
58
66
|
}
|
|
59
67
|
|
|
60
68
|
export function shouldExpand(uncertainty = 0, threshold = 0.5) {
|
package/src/dynamic-context.mjs
CHANGED
|
@@ -33,6 +33,7 @@ export function buildDynamicContext(taskText, entries, files = {}, options = {})
|
|
|
33
33
|
level,
|
|
34
34
|
maxFiles: options.maxFiles || 8,
|
|
35
35
|
maxFileChars: options.maxFileChars || 2000,
|
|
36
|
+
maxContextTokens: options.maxContextTokens || 0,
|
|
36
37
|
graph: options.graph,
|
|
37
38
|
}),
|
|
38
39
|
}
|
package/src/evidence.mjs
CHANGED
|
@@ -70,6 +70,92 @@ export function evidencePass(evidence = {}) {
|
|
|
70
70
|
return true
|
|
71
71
|
}
|
|
72
72
|
|
|
73
|
+
/**
|
|
74
|
+
* Extract machine-checkable evidence from a subagent/tool result.
|
|
75
|
+
*
|
|
76
|
+
* Accepts either a structured `result.evidence` object, top-level arrays on the
|
|
77
|
+
* result (`commands`, `tests`, `files`, `findings`), tool-call records, or an
|
|
78
|
+
* embedded `EVIDENCE_JSON` block in the result text. This lets the runtime
|
|
79
|
+
* prefer Harness-captured facts over agent prose whenever the Harness provides
|
|
80
|
+
* them.
|
|
81
|
+
*/
|
|
82
|
+
export function extractHarnessEvidence(result = {}) {
|
|
83
|
+
let evidence = createEvidence()
|
|
84
|
+
const source = (result && typeof result === 'object' && result.evidence) || result || {}
|
|
85
|
+
|
|
86
|
+
const apply = (key, fn) => {
|
|
87
|
+
for (const item of Array.isArray(source[key]) ? source[key] : []) {
|
|
88
|
+
try {
|
|
89
|
+
evidence = fn(evidence, item)
|
|
90
|
+
} catch {
|
|
91
|
+
// Skip malformed evidence entries; never fail the mission.
|
|
92
|
+
}
|
|
93
|
+
}
|
|
94
|
+
}
|
|
95
|
+
|
|
96
|
+
apply('commands', addCommandEvidence)
|
|
97
|
+
apply('tests', addTestEvidence)
|
|
98
|
+
apply('files', addFileEvidence)
|
|
99
|
+
apply('findings', addFindingEvidence)
|
|
100
|
+
|
|
101
|
+
if (Array.isArray(result?.toolCalls)) {
|
|
102
|
+
for (const call of result.toolCalls) {
|
|
103
|
+
try {
|
|
104
|
+
if (call?.command || call?.name) {
|
|
105
|
+
evidence = addCommandEvidence(evidence, {
|
|
106
|
+
command: call.command || call.name,
|
|
107
|
+
exitCode: call.exitCode ?? (call.ok === false ? 1 : 0),
|
|
108
|
+
output: call.output || call.stdout || '',
|
|
109
|
+
durationMs: call.durationMs || 0,
|
|
110
|
+
})
|
|
111
|
+
}
|
|
112
|
+
} catch {
|
|
113
|
+
// Ignore malformed tool-call evidence.
|
|
114
|
+
}
|
|
115
|
+
}
|
|
116
|
+
}
|
|
117
|
+
|
|
118
|
+
const embedded = parseEmbeddedEvidence(result?.output || result?.text || '')
|
|
119
|
+
if (embedded) {
|
|
120
|
+
for (const item of embedded.commands || []) {
|
|
121
|
+
try { evidence = addCommandEvidence(evidence, item) } catch { /* ignore */ }
|
|
122
|
+
}
|
|
123
|
+
for (const item of embedded.tests || []) {
|
|
124
|
+
try { evidence = addTestEvidence(evidence, item) } catch { /* ignore */ }
|
|
125
|
+
}
|
|
126
|
+
for (const item of embedded.files || []) {
|
|
127
|
+
try { evidence = addFileEvidence(evidence, item) } catch { /* ignore */ }
|
|
128
|
+
}
|
|
129
|
+
for (const item of embedded.findings || []) {
|
|
130
|
+
try { evidence = addFindingEvidence(evidence, item) } catch { /* ignore */ }
|
|
131
|
+
}
|
|
132
|
+
}
|
|
133
|
+
|
|
134
|
+
return evidence
|
|
135
|
+
}
|
|
136
|
+
|
|
137
|
+
function parseEmbeddedEvidence(text) {
|
|
138
|
+
const source = String(text || '')
|
|
139
|
+
const markers = ['EVIDENCE_JSON', 'EVIDENCE:']
|
|
140
|
+
for (const marker of markers) {
|
|
141
|
+
const start = source.indexOf(marker)
|
|
142
|
+
if (start === -1) continue
|
|
143
|
+
const jsonStart = source.indexOf('{', start)
|
|
144
|
+
if (jsonStart === -1) continue
|
|
145
|
+
const jsonEnd = source.lastIndexOf('}')
|
|
146
|
+
if (jsonEnd === -1 || jsonEnd <= jsonStart) continue
|
|
147
|
+
try {
|
|
148
|
+
const parsed = JSON.parse(source.slice(jsonStart, jsonEnd + 1))
|
|
149
|
+
if (parsed && (Array.isArray(parsed.commands) || Array.isArray(parsed.tests) || Array.isArray(parsed.files) || Array.isArray(parsed.findings))) {
|
|
150
|
+
return parsed
|
|
151
|
+
}
|
|
152
|
+
} catch {
|
|
153
|
+
// Not valid JSON; fall through to prose-based heuristics.
|
|
154
|
+
}
|
|
155
|
+
}
|
|
156
|
+
return null
|
|
157
|
+
}
|
|
158
|
+
|
|
73
159
|
export function summarizeEvidence(evidence = {}) {
|
|
74
160
|
const lines = []
|
|
75
161
|
if (evidence.commands?.length) {
|
package/src/memory-v3.mjs
CHANGED
|
@@ -17,10 +17,11 @@ export function createMemoryV3(base = {}) {
|
|
|
17
17
|
}
|
|
18
18
|
}
|
|
19
19
|
|
|
20
|
-
function countOutcomes(trajectories = []) {
|
|
20
|
+
function countOutcomes(trajectories = [], options = {}) {
|
|
21
21
|
let successCount = 0
|
|
22
22
|
let failureCount = 0
|
|
23
23
|
for (const t of trajectories || []) {
|
|
24
|
+
if (options.requireEvidence && !hasEvidenceBacking(t)) continue
|
|
24
25
|
const text = String(t?.text || t?.outcome || '').toLowerCase()
|
|
25
26
|
if (/success|done|pass|完成|ok/.test(text)) successCount += 1
|
|
26
27
|
else if (/fail|error|blocked|失败|错误/.test(text)) failureCount += 1
|
|
@@ -28,13 +29,28 @@ function countOutcomes(trajectories = []) {
|
|
|
28
29
|
return { successCount, failureCount }
|
|
29
30
|
}
|
|
30
31
|
|
|
32
|
+
function hasEvidenceBacking(t = {}) {
|
|
33
|
+
const evidence = t.evidence || t.verification || t.verified
|
|
34
|
+
if (evidence === true) return true
|
|
35
|
+
if (evidence && typeof evidence === 'object') {
|
|
36
|
+
const e = evidence
|
|
37
|
+
return Boolean(
|
|
38
|
+
(Array.isArray(e.commands) && e.commands.length) ||
|
|
39
|
+
(Array.isArray(e.tests) && e.tests.length) ||
|
|
40
|
+
(Array.isArray(e.files) && e.files.length) ||
|
|
41
|
+
(Array.isArray(e.findings) && e.findings.length),
|
|
42
|
+
)
|
|
43
|
+
}
|
|
44
|
+
return false
|
|
45
|
+
}
|
|
46
|
+
|
|
31
47
|
export function distillSkill(memory, trajectories = [], options = {}) {
|
|
32
48
|
const minSuccesses = Number(options.minSuccesses) || 3
|
|
33
|
-
const { successCount, failureCount } = countOutcomes(trajectories)
|
|
49
|
+
const { successCount, failureCount } = countOutcomes(trajectories, options)
|
|
34
50
|
if (successCount < minSuccesses) return { memory, promoted: false, reason: `need ${minSuccesses} successes, got ${successCount}` }
|
|
35
51
|
|
|
36
52
|
const successfulSteps = (trajectories || [])
|
|
37
|
-
.filter((t) => /success|done|pass|完成|ok/i.test(String(t?.text || t?.outcome || '')))
|
|
53
|
+
.filter((t) => (!options.requireEvidence || hasEvidenceBacking(t)) && /success|done|pass|完成|ok/i.test(String(t?.text || t?.outcome || '')))
|
|
38
54
|
.map((t) => t?.text || t?.step || '')
|
|
39
55
|
.filter(Boolean)
|
|
40
56
|
|
|
@@ -78,11 +94,27 @@ export function recordSkillOutcome(memory, skillName, success) {
|
|
|
78
94
|
return { ...memory, learnedSkills: skills }
|
|
79
95
|
}
|
|
80
96
|
|
|
81
|
-
export function retrieveHistoricalFailures(memory, taskText = '') {
|
|
97
|
+
export function retrieveHistoricalFailures(memory, taskText = '', options = {}) {
|
|
82
98
|
const text = String(taskText || '').toLowerCase()
|
|
83
99
|
const failures = memory.failures || []
|
|
84
100
|
if (!text) return failures
|
|
85
|
-
|
|
101
|
+
const minScore = options.minScore === undefined ? 0.1 : Number(options.minScore) || 0
|
|
102
|
+
const tokens = tokenize(text)
|
|
103
|
+
return failures
|
|
104
|
+
.map((f) => {
|
|
105
|
+
const hay = `${f.text || ''} ${f.reason || ''} ${f.category || ''} ${f.type || ''}`.toLowerCase()
|
|
106
|
+
const hayTokens = tokenize(hay)
|
|
107
|
+
const overlap = hayTokens.filter((t) => tokens.includes(t)).length
|
|
108
|
+
const score = overlap / Math.max(1, Math.min(tokens.length, hayTokens.length))
|
|
109
|
+
return { ...f, _score: Math.round(score * 1000) / 1000 }
|
|
110
|
+
})
|
|
111
|
+
.filter((f) => f._score >= minScore || String(f.text || f.reason || '').toLowerCase().includes(text))
|
|
112
|
+
.sort((a, b) => b._score - a._score)
|
|
113
|
+
.map(({ _score, ...f }) => f)
|
|
114
|
+
}
|
|
115
|
+
|
|
116
|
+
function tokenize(text) {
|
|
117
|
+
return String(text || '').toLowerCase().match(/[a-z0-9_]+/g) || []
|
|
86
118
|
}
|
|
87
119
|
|
|
88
120
|
export function recordExecutionPolicy(memory, policy = {}) {
|
package/src/omni-router.mjs
CHANGED
|
@@ -30,9 +30,9 @@ import { runDagLoop } from './agent-runtime.mjs'
|
|
|
30
30
|
import { buildVisualQaPrompt, buildVisualQaStepRequirement, callVisionApi, isFrontendTask, parseVisualQaResponse } from './visual-qa.mjs'
|
|
31
31
|
import { createTaskDecision, buildPolicyFromTaskDecision } from './task-decision.mjs'
|
|
32
32
|
import { compileTaskWithLLM } from './task-compiler.mjs'
|
|
33
|
-
import { bindCapabilitiesToDag, createMissionDag, formatMissionDag } from './mission-dag.mjs'
|
|
33
|
+
import { bindCapabilitiesToDag, createMissionDag, formatMissionDag, isMissionDagComplete } from './mission-dag.mjs'
|
|
34
34
|
import { generateMissionDag, roleForTask } from './planner-dag.mjs'
|
|
35
|
-
import { autoPopulateCapabilityBrain, createCapabilityBrain } from './capability-brain.mjs'
|
|
35
|
+
import { autoPopulateCapabilityBrain, createCapabilityBrain, recordCapabilityOutcome } from './capability-brain.mjs'
|
|
36
36
|
import { loadCapabilityManifests } from './capability-manifest.mjs'
|
|
37
37
|
import { capabilityToolFilter } from './capability-sandbox.mjs'
|
|
38
38
|
import { buildProgressiveContext } from './context-expansion.mjs'
|
|
@@ -42,7 +42,8 @@ import { retrieveContext } from './hybrid-retrieval.mjs'
|
|
|
42
42
|
import { formatMemory, recordDecision, recordFailure, recordProject, recordTrajectory, summarizeMemory } from './memory.mjs'
|
|
43
43
|
import { createMemoryEngine, loadMemoryEngine, saveMemoryEngine } from './memory-engine.mjs'
|
|
44
44
|
import { captureEvidence, createEvidenceStore, evidenceSummary } from './evidence-store.mjs'
|
|
45
|
-
import {
|
|
45
|
+
import { evidencePass, extractHarnessEvidence } from './evidence.mjs'
|
|
46
|
+
import { listMissionStates, loadMissionState, saveMissionState } from './mission-resume.mjs'
|
|
46
47
|
import { collectResults, formatResultSummary, importBenchmarkRecord, missingTaskIds, summarizeResults } from './benchmark-results.mjs'
|
|
47
48
|
import { buildAstGraph, collectSourceFiles } from './ast-provider.mjs'
|
|
48
49
|
|
|
@@ -695,6 +696,7 @@ export function apply(ctx, config = {}) {
|
|
|
695
696
|
uncertainty: Number(config.contextUncertainty) || 0.6,
|
|
696
697
|
maxFiles,
|
|
697
698
|
maxFileChars: 2000,
|
|
699
|
+
maxContextTokens: ctxBudget.retrievalBudget,
|
|
698
700
|
graph: graphAdj,
|
|
699
701
|
})
|
|
700
702
|
return dynamic.context
|
|
@@ -1359,6 +1361,113 @@ export function apply(ctx, config = {}) {
|
|
|
1359
1361
|
},
|
|
1360
1362
|
})
|
|
1361
1363
|
|
|
1364
|
+
async function createMissionExecutor({ session, agent, subagents, task, taskType, frontend, brief, capabilityBrain, evidenceRecords, resumeKey }) {
|
|
1365
|
+
const cwd = session?.meta?.cwd || session?.header?.cwd
|
|
1366
|
+
let brain = capabilityBrain
|
|
1367
|
+
|
|
1368
|
+
const saveProgress = async (snapshot = {}) => {
|
|
1369
|
+
if (!cwd || !resumeKey) return
|
|
1370
|
+
const dag = snapshot.dag
|
|
1371
|
+
saveMissionState(cwd, resumeKey, {
|
|
1372
|
+
status: dag ? (isMissionDagComplete(dag) ? 'completed' : 'active') : 'active',
|
|
1373
|
+
dag,
|
|
1374
|
+
evidence: evidenceRecords,
|
|
1375
|
+
metrics: snapshot.metrics || {},
|
|
1376
|
+
capabilityBrain: brain,
|
|
1377
|
+
brief,
|
|
1378
|
+
task,
|
|
1379
|
+
taskType,
|
|
1380
|
+
savedAt: new Date().toISOString(),
|
|
1381
|
+
})
|
|
1382
|
+
}
|
|
1383
|
+
|
|
1384
|
+
const act = async (action) => {
|
|
1385
|
+
const goal = action.task?.goal || action.taskId || ''
|
|
1386
|
+
const visualQa = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
|
|
1387
|
+
? `\n\n${buildVisualQaStepRequirement()}`
|
|
1388
|
+
: ''
|
|
1389
|
+
const briefText = `\n\nTask brief:\nObjective: ${brief.objective}\nAcceptance: ${brief.acceptanceCriteria.join('; ')}`
|
|
1390
|
+
const taskContext = await getProjectContext(session, taskType, goal).catch(() => '')
|
|
1391
|
+
const prompt = `Mission: ${task}\nTask: ${action.taskId} — ${goal}\n\nExecute this step. Reply with a short result and evidence.${briefText}${taskContext ? `\n\nContext:\n${taskContext}` : ''}${visualQa}`
|
|
1392
|
+
const role = roleForTask(action.task || {})
|
|
1393
|
+
const sandbox = capabilityToolFilter(brain, action.task?.requiredCapabilities || [], role)
|
|
1394
|
+
const run = await subagents.start('spawn', {
|
|
1395
|
+
label: `mission-${action.taskId}`,
|
|
1396
|
+
prompt: [{ type: 'text', text: prompt }],
|
|
1397
|
+
parent: agent,
|
|
1398
|
+
maxDepth: 1,
|
|
1399
|
+
...(sandbox.allow.length || sandbox.deny.length ? { toolFilter: sandbox } : {}),
|
|
1400
|
+
})
|
|
1401
|
+
const result = await run.result
|
|
1402
|
+
const output = (Array.isArray(result.output) ? result.output : [])
|
|
1403
|
+
.filter((block) => block?.type === 'text' && typeof block.text === 'string')
|
|
1404
|
+
.map((block) => block.text)
|
|
1405
|
+
.join('')
|
|
1406
|
+
try { await run.dispose() } catch { /* best-effort */ }
|
|
1407
|
+
|
|
1408
|
+
const harness = extractHarnessEvidence({ ...result, output })
|
|
1409
|
+
const hasStructured = harness.commands.length > 0 || harness.tests.length > 0 || harness.files.length > 0 || harness.findings.length > 0
|
|
1410
|
+
let record
|
|
1411
|
+
if (hasStructured) {
|
|
1412
|
+
record = {
|
|
1413
|
+
id: `E-${Date.now().toString(36)}-${Math.random().toString(36).slice(2, 6)}`,
|
|
1414
|
+
type: 'harness',
|
|
1415
|
+
source: action.taskId,
|
|
1416
|
+
value: output.slice(0, 2000),
|
|
1417
|
+
ok: evidencePass(harness),
|
|
1418
|
+
evidence: harness,
|
|
1419
|
+
at: new Date().toISOString(),
|
|
1420
|
+
}
|
|
1421
|
+
} else {
|
|
1422
|
+
const captured = captureEvidence({ entries: evidenceRecords }, {
|
|
1423
|
+
type: 'agent_output',
|
|
1424
|
+
source: action.taskId,
|
|
1425
|
+
value: output.slice(0, 2000),
|
|
1426
|
+
ok: !/FAIL|error|失败|not ok/i.test(output),
|
|
1427
|
+
})
|
|
1428
|
+
record = captured.record
|
|
1429
|
+
}
|
|
1430
|
+
evidenceRecords.push(record)
|
|
1431
|
+
return {
|
|
1432
|
+
output,
|
|
1433
|
+
evidenceId: record.id,
|
|
1434
|
+
evidence: harness,
|
|
1435
|
+
hasStructuredEvidence: hasStructured,
|
|
1436
|
+
tokenUsage: result?.tokenUsage ?? result?.usage?.totalTokens ?? 0,
|
|
1437
|
+
cost: result?.cost ?? 0,
|
|
1438
|
+
toolCalls: result?.toolCalls ?? result?.usage?.toolCalls ?? 0,
|
|
1439
|
+
}
|
|
1440
|
+
}
|
|
1441
|
+
|
|
1442
|
+
const observe = async (result, _current, action) => {
|
|
1443
|
+
const goal = action?.task?.goal || ''
|
|
1444
|
+
const needsVisual = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
|
|
1445
|
+
if (needsVisual && !/VISUAL_QA:\s*PASS/i.test(result.output || '')) return { type: 'test_failure', reason: 'visual QA failed' }
|
|
1446
|
+
|
|
1447
|
+
const harness = result.evidence || {}
|
|
1448
|
+
const hasStructured = result.hasStructuredEvidence === true || harness.commands?.length || harness.tests?.length || harness.files?.length || harness.findings?.length
|
|
1449
|
+
if (hasStructured) {
|
|
1450
|
+
const passed = evidencePass(harness)
|
|
1451
|
+
if (!passed) {
|
|
1452
|
+
brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], false)
|
|
1453
|
+
return { type: 'test_failure', reason: 'structured harness evidence failed', evidence: harness }
|
|
1454
|
+
}
|
|
1455
|
+
brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], true)
|
|
1456
|
+
return { type: 'step_done', evidence: harness }
|
|
1457
|
+
}
|
|
1458
|
+
|
|
1459
|
+
if (/FAIL|error|失败|not ok/i.test(result.output || '')) {
|
|
1460
|
+
const failure = classifyFailure({ type: 'unknown', reason: result.output || '' })
|
|
1461
|
+
brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], false)
|
|
1462
|
+
return { type: failure.category, reason: failure.recovery, detail: String(result.output || '').slice(0, 500) }
|
|
1463
|
+
}
|
|
1464
|
+
brain = recordCapabilityOutcome(brain, action?.task?.allowedTools?.[0], true)
|
|
1465
|
+
return { type: 'step_done' }
|
|
1466
|
+
}
|
|
1467
|
+
|
|
1468
|
+
return { act, observe, saveProgress, getCapabilityBrain: () => brain }
|
|
1469
|
+
}
|
|
1470
|
+
|
|
1362
1471
|
registerTool({
|
|
1363
1472
|
name: 'omni_mission_run',
|
|
1364
1473
|
description: 'Run a Mission Planner loop with real subagents: Observe → Think → Act → Replan until completed or maxSteps.',
|
|
@@ -1397,56 +1506,25 @@ export function apply(ctx, config = {}) {
|
|
|
1397
1506
|
capabilityBrain = loadCapabilityManifests(capabilityBrain, config.capabilityManifests || [])
|
|
1398
1507
|
const dag = bindCapabilitiesToDag(generateMissionDag(mission, brief, { taskType }), capabilityBrain)
|
|
1399
1508
|
const evidenceRecords = []
|
|
1509
|
+
const cwd = session?.meta?.cwd || session?.header?.cwd
|
|
1510
|
+
const resumeKey = cwd ? `mission-${Date.now().toString(36)}` : null
|
|
1511
|
+
const executor = await createMissionExecutor({
|
|
1512
|
+
session,
|
|
1513
|
+
agent,
|
|
1514
|
+
subagents,
|
|
1515
|
+
task,
|
|
1516
|
+
taskType,
|
|
1517
|
+
frontend,
|
|
1518
|
+
brief,
|
|
1519
|
+
capabilityBrain,
|
|
1520
|
+
evidenceRecords,
|
|
1521
|
+
resumeKey,
|
|
1522
|
+
})
|
|
1400
1523
|
|
|
1401
1524
|
const finalDag = await runDagLoop(dag, {
|
|
1402
|
-
act:
|
|
1403
|
-
|
|
1404
|
-
|
|
1405
|
-
? `\n\n${buildVisualQaStepRequirement()}`
|
|
1406
|
-
: ''
|
|
1407
|
-
const briefText = `\n\nTask brief:\nObjective: ${brief.objective}\nAcceptance: ${brief.acceptanceCriteria.join('; ')}`
|
|
1408
|
-
const taskContext = await getProjectContext(session, taskType, goal).catch(() => '')
|
|
1409
|
-
const prompt = `Mission: ${task}\nTask: ${action.taskId} — ${goal}\n\nExecute this step. Reply with a short result and evidence.${briefText}${taskContext ? `\n\nContext:\n${taskContext}` : ''}${visualQa}`
|
|
1410
|
-
const role = roleForTask(action.task || {})
|
|
1411
|
-
const sandbox = capabilityToolFilter(capabilityBrain, action.task?.requiredCapabilities || [], role)
|
|
1412
|
-
const run = await subagents.start('spawn', {
|
|
1413
|
-
label: `mission-${action.taskId}`,
|
|
1414
|
-
prompt: [{ type: 'text', text: prompt }],
|
|
1415
|
-
parent: agent,
|
|
1416
|
-
maxDepth: 1,
|
|
1417
|
-
...(sandbox.allow.length || sandbox.deny.length ? { toolFilter: sandbox } : {}),
|
|
1418
|
-
})
|
|
1419
|
-
const result = await run.result
|
|
1420
|
-
const output = (Array.isArray(result.output) ? result.output : [])
|
|
1421
|
-
.filter((block) => block?.type === 'text' && typeof block.text === 'string')
|
|
1422
|
-
.map((block) => block.text)
|
|
1423
|
-
.join('')
|
|
1424
|
-
try { await run.dispose() } catch { /* best-effort */ }
|
|
1425
|
-
const captured = captureEvidence({ entries: evidenceRecords }, {
|
|
1426
|
-
type: 'agent_output',
|
|
1427
|
-
source: action.taskId,
|
|
1428
|
-
value: output.slice(0, 2000),
|
|
1429
|
-
ok: !/FAIL|error|失败|not ok/i.test(output),
|
|
1430
|
-
})
|
|
1431
|
-
evidenceRecords.push(captured.record)
|
|
1432
|
-
return {
|
|
1433
|
-
output,
|
|
1434
|
-
evidenceId: captured.record.id,
|
|
1435
|
-
tokenUsage: result?.tokenUsage ?? result?.usage?.totalTokens ?? 0,
|
|
1436
|
-
cost: result?.cost ?? 0,
|
|
1437
|
-
toolCalls: result?.toolCalls ?? result?.usage?.toolCalls ?? 0,
|
|
1438
|
-
}
|
|
1439
|
-
},
|
|
1440
|
-
observe: async (result, _current, action) => {
|
|
1441
|
-
const goal = action?.task?.goal || ''
|
|
1442
|
-
const needsVisual = frontend && /validate|verify|visual|ui/i.test(goal) && config.autoVisualQA !== false
|
|
1443
|
-
if (needsVisual && !/VISUAL_QA:\s*PASS/i.test(result.output || '')) return { type: 'test_failure', reason: 'visual QA failed' }
|
|
1444
|
-
if (/FAIL|error|失败|not ok/i.test(result.output || '')) {
|
|
1445
|
-
const failure = classifyFailure({ type: 'unknown', reason: result.output || '' })
|
|
1446
|
-
return { type: failure.category, reason: failure.recovery, detail: String(result.output || '').slice(0, 500) }
|
|
1447
|
-
}
|
|
1448
|
-
return { type: 'step_done' }
|
|
1449
|
-
},
|
|
1525
|
+
act: executor.act,
|
|
1526
|
+
observe: executor.observe,
|
|
1527
|
+
onProgress: executor.saveProgress,
|
|
1450
1528
|
maxSteps,
|
|
1451
1529
|
maxParallel: Number(args?.maxParallel) || 1,
|
|
1452
1530
|
maxTokens: Number(args?.maxTokens) || 0,
|
|
@@ -1456,19 +1534,21 @@ export function apply(ctx, config = {}) {
|
|
|
1456
1534
|
maxReplans: Number(args?.maxReplans) || 0,
|
|
1457
1535
|
maxWallClockMs: Number(args?.maxWallClockMs) || 0,
|
|
1458
1536
|
})
|
|
1537
|
+
capabilityBrain = executor.getCapabilityBrain()
|
|
1459
1538
|
|
|
1460
1539
|
const metrics = finalDag.metrics || {}
|
|
1461
1540
|
const evSummary = evidenceSummary({ entries: evidenceRecords })
|
|
1462
|
-
const cwd = session?.meta?.cwd || session?.header?.cwd
|
|
1463
1541
|
let resumeInfo = ''
|
|
1464
|
-
if (cwd) {
|
|
1465
|
-
const resumeKey = `mission-${Date.now().toString(36)}`
|
|
1542
|
+
if (cwd && resumeKey) {
|
|
1466
1543
|
const saved = saveMissionState(cwd, resumeKey, {
|
|
1467
1544
|
status: finalDag.status,
|
|
1468
1545
|
dag: finalDag.dag,
|
|
1469
1546
|
evidence: evidenceRecords,
|
|
1470
1547
|
metrics,
|
|
1548
|
+
capabilityBrain,
|
|
1549
|
+
brief,
|
|
1471
1550
|
task,
|
|
1551
|
+
taskType,
|
|
1472
1552
|
savedAt: new Date().toISOString(),
|
|
1473
1553
|
})
|
|
1474
1554
|
resumeInfo = `\nSaved mission state: ${saved}`
|
|
@@ -1487,6 +1567,107 @@ export function apply(ctx, config = {}) {
|
|
|
1487
1567
|
},
|
|
1488
1568
|
})
|
|
1489
1569
|
|
|
1570
|
+
registerTool({
|
|
1571
|
+
name: 'omni_mission_resume',
|
|
1572
|
+
description: 'Resume a saved Mission DAG from .omni/missions/<key>.json. Without a key, lists saved missions.',
|
|
1573
|
+
parameters: {
|
|
1574
|
+
type: 'object',
|
|
1575
|
+
properties: {
|
|
1576
|
+
key: { type: 'string', description: 'Saved mission state key (from omni_mission_run output)' },
|
|
1577
|
+
maxSteps: { type: 'number', description: 'Max additional loop steps (default 20)' },
|
|
1578
|
+
maxParallel: { type: 'number', description: 'Max tasks to run concurrently (default 1)' },
|
|
1579
|
+
maxTokens: { type: 'number', description: 'Max token budget for the resumed portion (0 = unlimited)' },
|
|
1580
|
+
maxCost: { type: 'number', description: 'Max cost budget for the resumed portion (0 = unlimited)' },
|
|
1581
|
+
maxToolCalls: { type: 'number', description: 'Max tool calls for the resumed portion (0 = unlimited)' },
|
|
1582
|
+
maxRepairs: { type: 'number', description: 'Max repairs before blocking (0 = unlimited)' },
|
|
1583
|
+
maxReplans: { type: 'number', description: 'Max replans before blocking (0 = unlimited)' },
|
|
1584
|
+
maxWallClockMs: { type: 'number', description: 'Max wall-clock time in ms (0 = unlimited)' },
|
|
1585
|
+
},
|
|
1586
|
+
required: [],
|
|
1587
|
+
},
|
|
1588
|
+
async execute(args) {
|
|
1589
|
+
const session = currentSession()
|
|
1590
|
+
const agent = session && agentFor(session)
|
|
1591
|
+
const subagents = ctx.get('subagents') || ctx.subagents
|
|
1592
|
+
const cwd = session?.meta?.cwd || session?.header?.cwd
|
|
1593
|
+
if (!cwd) return 'No workspace cwd found.'
|
|
1594
|
+
const key = String(args?.key || '').trim()
|
|
1595
|
+
if (!key) {
|
|
1596
|
+
const keys = listMissionStates(cwd)
|
|
1597
|
+
return keys.length
|
|
1598
|
+
? `Saved missions:\n${keys.map((k) => `- ${k}`).join('\n')}\n\nPass "key" to resume one.`
|
|
1599
|
+
: 'No saved missions found in .omni/missions.'
|
|
1600
|
+
}
|
|
1601
|
+
const saved = loadMissionState(cwd, key)
|
|
1602
|
+
if (!saved) {
|
|
1603
|
+
const keys = listMissionStates(cwd)
|
|
1604
|
+
return `No saved mission found for key "${key}". Available: ${keys.join(', ') || '(none)'}`
|
|
1605
|
+
}
|
|
1606
|
+
if (saved.status === 'completed') return `Mission "${key}" is already completed.`
|
|
1607
|
+
if (!session || !agent || !subagents?.start) return 'Mission resume requires an active session with subagents.'
|
|
1608
|
+
|
|
1609
|
+
const task = saved.task || saved.dag?.mission?.task || ''
|
|
1610
|
+
const taskType = saved.taskType || classifyTaskType(task)
|
|
1611
|
+
const frontend = isFrontendTask(task)
|
|
1612
|
+
const brief = saved.brief || { objective: task, acceptanceCriteria: [] }
|
|
1613
|
+
const evidenceRecords = Array.isArray(saved.evidence) ? saved.evidence : []
|
|
1614
|
+
const toolsService = ctx.get('tools') || ctx.tools
|
|
1615
|
+
let capabilityBrain = saved.capabilityBrain || autoPopulateCapabilityBrain(createCapabilityBrain(), await collectToolNames(toolsService))
|
|
1616
|
+
capabilityBrain = loadCapabilityManifests(capabilityBrain, config.capabilityManifests || [])
|
|
1617
|
+
|
|
1618
|
+
const executor = await createMissionExecutor({
|
|
1619
|
+
session,
|
|
1620
|
+
agent,
|
|
1621
|
+
subagents,
|
|
1622
|
+
task,
|
|
1623
|
+
taskType,
|
|
1624
|
+
frontend,
|
|
1625
|
+
brief,
|
|
1626
|
+
capabilityBrain,
|
|
1627
|
+
evidenceRecords,
|
|
1628
|
+
resumeKey: key,
|
|
1629
|
+
})
|
|
1630
|
+
|
|
1631
|
+
const finalDag = await runDagLoop(saved.dag, {
|
|
1632
|
+
act: executor.act,
|
|
1633
|
+
observe: executor.observe,
|
|
1634
|
+
onProgress: executor.saveProgress,
|
|
1635
|
+
maxSteps: Number(args?.maxSteps) || 20,
|
|
1636
|
+
maxParallel: Number(args?.maxParallel) || 1,
|
|
1637
|
+
maxTokens: Number(args?.maxTokens) || 0,
|
|
1638
|
+
maxCost: Number(args?.maxCost) || 0,
|
|
1639
|
+
maxToolCalls: Number(args?.maxToolCalls) || 0,
|
|
1640
|
+
maxRepairs: Number(args?.maxRepairs) || 0,
|
|
1641
|
+
maxReplans: Number(args?.maxReplans) || 0,
|
|
1642
|
+
maxWallClockMs: Number(args?.maxWallClockMs) || 0,
|
|
1643
|
+
})
|
|
1644
|
+
|
|
1645
|
+
const metrics = finalDag.metrics || {}
|
|
1646
|
+
saveMissionState(cwd, key, {
|
|
1647
|
+
status: finalDag.status,
|
|
1648
|
+
dag: finalDag.dag,
|
|
1649
|
+
evidence: evidenceRecords,
|
|
1650
|
+
metrics,
|
|
1651
|
+
capabilityBrain: executor.getCapabilityBrain(),
|
|
1652
|
+
brief,
|
|
1653
|
+
task,
|
|
1654
|
+
taskType,
|
|
1655
|
+
savedAt: new Date().toISOString(),
|
|
1656
|
+
})
|
|
1657
|
+
const evSummary = evidenceSummary({ entries: evidenceRecords })
|
|
1658
|
+
return [
|
|
1659
|
+
`Mission resume: ${finalDag.status}`,
|
|
1660
|
+
`Tasks done: ${finalDag.dag.tasks.filter((t) => t.status === 'done').length}/${finalDag.dag.tasks.length}`,
|
|
1661
|
+
`Additional steps: ${finalDag.actions.length}`,
|
|
1662
|
+
`Replans: ${metrics.replanCount || 0}, Repairs: ${metrics.repairCount || 0}, ToolCalls: ${metrics.toolCalls || 0}, Tokens: ${metrics.tokenUsage || 0}, Cost: ${metrics.cost || 0}`,
|
|
1663
|
+
`Evidence captured: ${evSummary.total} (failed=${evSummary.failed})`,
|
|
1664
|
+
`Mission: ${task}`,
|
|
1665
|
+
'',
|
|
1666
|
+
formatMissionDag(finalDag.dag),
|
|
1667
|
+
].join('\n')
|
|
1668
|
+
},
|
|
1669
|
+
})
|
|
1670
|
+
|
|
1490
1671
|
registerTool({
|
|
1491
1672
|
name: 'omni_plan',
|
|
1492
1673
|
description: 'Force the current session into plan-first mode (enter built-in plan mode).',
|
|
@@ -146,3 +146,20 @@ test('runDagLoop respects maxReplans', async () => {
|
|
|
146
146
|
})
|
|
147
147
|
assert.equal(result.status, 'max_replans')
|
|
148
148
|
})
|
|
149
|
+
|
|
150
|
+
test('runDagLoop calls onProgress after each batch', async () => {
|
|
151
|
+
const dag = createMissionDag(buildMission('实现退款', { taskType: 'feature' }))
|
|
152
|
+
const snapshots = []
|
|
153
|
+
const result = await runDagLoop(dag, {
|
|
154
|
+
act: async () => ({ ok: true }),
|
|
155
|
+
observe: async () => ({ type: 'step_done' }),
|
|
156
|
+
maxSteps: 100,
|
|
157
|
+
maxParallel: 1,
|
|
158
|
+
onProgress: async (snapshot) => {
|
|
159
|
+
snapshots.push(snapshot)
|
|
160
|
+
},
|
|
161
|
+
})
|
|
162
|
+
assert.equal(result.status, 'completed')
|
|
163
|
+
assert.ok(snapshots.length >= dag.tasks.length)
|
|
164
|
+
assert.ok(snapshots.every((s) => s.dag && Array.isArray(s.actions)))
|
|
165
|
+
})
|
|
@@ -36,3 +36,14 @@ test('buildDynamicContext expands when uncertainty is high', () => {
|
|
|
36
36
|
assert.equal(result.level, 2)
|
|
37
37
|
assert.match(result.context, /Implementations:/)
|
|
38
38
|
})
|
|
39
|
+
|
|
40
|
+
test('buildDynamicContext enforces maxContextTokens budget', () => {
|
|
41
|
+
const result = buildDynamicContext('fix login', entries, files, {
|
|
42
|
+
level: 4,
|
|
43
|
+
uncertainty: 0.8,
|
|
44
|
+
threshold: 0.5,
|
|
45
|
+
maxContextTokens: 10,
|
|
46
|
+
})
|
|
47
|
+
assert.ok(result.context.length <= 10 * 4 + 64)
|
|
48
|
+
assert.match(result.context, /truncated by context budget/)
|
|
49
|
+
})
|
package/test/evidence.test.mjs
CHANGED
|
@@ -8,6 +8,7 @@ import {
|
|
|
8
8
|
addTestEvidence,
|
|
9
9
|
createEvidence,
|
|
10
10
|
evidencePass,
|
|
11
|
+
extractHarnessEvidence,
|
|
11
12
|
summarizeEvidence,
|
|
12
13
|
} from '../src/evidence.mjs'
|
|
13
14
|
|
|
@@ -44,3 +45,37 @@ test('summarizeEvidence renders sections', () => {
|
|
|
44
45
|
assert.match(text, /Commands/)
|
|
45
46
|
assert.match(text, /Tests/)
|
|
46
47
|
})
|
|
48
|
+
|
|
49
|
+
test('extractHarnessEvidence normalizes structured result evidence', () => {
|
|
50
|
+
const evidence = extractHarnessEvidence({
|
|
51
|
+
commands: [{ command: 'npm test', exitCode: 0, output: 'ok' }],
|
|
52
|
+
tests: [{ command: 'npm test', exitCode: 0, total: 3, passed: 3, failed: 0 }],
|
|
53
|
+
files: [{ file: 'src/a.ts', lines: [1], beforeHash: 'a', afterHash: 'b', diffHash: 'c' }],
|
|
54
|
+
findings: [{ finding: 'nit', severity: 'low', file: 'src/a.ts', line: 2 }],
|
|
55
|
+
})
|
|
56
|
+
assert.equal(evidence.commands.length, 1)
|
|
57
|
+
assert.equal(evidence.tests.length, 1)
|
|
58
|
+
assert.equal(evidence.files.length, 1)
|
|
59
|
+
assert.equal(evidence.findings.length, 1)
|
|
60
|
+
assert.equal(evidencePass(evidence), true)
|
|
61
|
+
})
|
|
62
|
+
|
|
63
|
+
test('extractHarnessEvidence reads toolCalls as command evidence', () => {
|
|
64
|
+
const evidence = extractHarnessEvidence({
|
|
65
|
+
toolCalls: [
|
|
66
|
+
{ name: 'npm test', exitCode: 0, output: 'all pass' },
|
|
67
|
+
{ name: 'node run.js', exitCode: 1, output: 'boom' },
|
|
68
|
+
],
|
|
69
|
+
})
|
|
70
|
+
assert.equal(evidence.commands.length, 2)
|
|
71
|
+
assert.equal(evidence.commands[0].exitCode, 0)
|
|
72
|
+
assert.equal(evidence.commands[1].exitCode, 1)
|
|
73
|
+
assert.equal(evidencePass(evidence), false)
|
|
74
|
+
})
|
|
75
|
+
|
|
76
|
+
test('extractHarnessEvidence parses embedded EVIDENCE_JSON block', () => {
|
|
77
|
+
const output = 'I ran the checks.\nEVIDENCE_JSON\n{"tests":[{"command":"npm test","exitCode":0,"total":2,"passed":2,"failed":0}]}'
|
|
78
|
+
const evidence = extractHarnessEvidence({ output })
|
|
79
|
+
assert.equal(evidence.tests.length, 1)
|
|
80
|
+
assert.equal(evidencePass(evidence), true)
|
|
81
|
+
})
|
package/test/memory-v3.test.mjs
CHANGED
|
@@ -37,6 +37,38 @@ test('retrieveHistoricalFailures filters by task text', () => {
|
|
|
37
37
|
assert.equal(hits.length, 1)
|
|
38
38
|
})
|
|
39
39
|
|
|
40
|
+
test('retrieveHistoricalFailures ranks by token overlap instead of exact substring only', () => {
|
|
41
|
+
const memory = createMemoryV3({
|
|
42
|
+
failures: [
|
|
43
|
+
{ text: 'login timeout under concurrency', reason: 'session refresh race' },
|
|
44
|
+
{ text: 'payment rounding', reason: 'float math' },
|
|
45
|
+
],
|
|
46
|
+
})
|
|
47
|
+
const hits = retrieveHistoricalFailures(memory, 'session refresh race timeout')
|
|
48
|
+
assert.ok(hits.length >= 1)
|
|
49
|
+
assert.match(hits[0].text, /login timeout/)
|
|
50
|
+
})
|
|
51
|
+
|
|
52
|
+
test('distillSkill can require evidence-backed successes before promoting', () => {
|
|
53
|
+
const memory = createMemoryV3()
|
|
54
|
+
const noEvidence = distillSkill(
|
|
55
|
+
memory,
|
|
56
|
+
[{ text: 'step ok' }, { text: 'step ok' }, { text: 'step ok' }],
|
|
57
|
+
{ name: 'no-evidence', minSuccesses: 3, requireEvidence: true },
|
|
58
|
+
)
|
|
59
|
+
assert.equal(noEvidence.promoted, false)
|
|
60
|
+
const withEvidence = distillSkill(
|
|
61
|
+
memory,
|
|
62
|
+
[
|
|
63
|
+
{ text: 'step ok', evidence: { tests: [{ command: 'npm test', exitCode: 0, total: 1, passed: 1, failed: 0 }] } },
|
|
64
|
+
{ text: 'step ok', evidence: { commands: [{ command: 'npm test', exitCode: 0 }] } },
|
|
65
|
+
{ text: 'step ok', evidence: { tests: [{ command: 'npm test', exitCode: 0, total: 1, passed: 1, failed: 0 }] } },
|
|
66
|
+
],
|
|
67
|
+
{ name: 'evidence-skill', minSuccesses: 3, requireEvidence: true },
|
|
68
|
+
)
|
|
69
|
+
assert.equal(withEvidence.promoted, true)
|
|
70
|
+
})
|
|
71
|
+
|
|
40
72
|
test('execution policies and cross-session strategies persist in memory', () => {
|
|
41
73
|
let memory = createMemoryV3()
|
|
42
74
|
memory = recordExecutionPolicy(memory, { name: 'schema-change', trigger: 'migration' })
|
|
@@ -0,0 +1,68 @@
|
|
|
1
|
+
import test from 'node:test'
|
|
2
|
+
import assert from 'node:assert/strict'
|
|
3
|
+
import fs from 'node:fs'
|
|
4
|
+
import os from 'node:os'
|
|
5
|
+
import path from 'node:path'
|
|
6
|
+
|
|
7
|
+
import {
|
|
8
|
+
buildPrompt,
|
|
9
|
+
readManifests,
|
|
10
|
+
runCommand,
|
|
11
|
+
validateManifests,
|
|
12
|
+
writeResults,
|
|
13
|
+
} from '../benchmark/omnibench-v2/run.mjs'
|
|
14
|
+
|
|
15
|
+
test('buildPrompt includes task and BENCHMARK marker', () => {
|
|
16
|
+
const manifest = { task: 'Fix pagination', acceptance: ['all tests pass'] }
|
|
17
|
+
const raw = buildPrompt(manifest, 'raw')
|
|
18
|
+
assert.match(raw, /Fix pagination/)
|
|
19
|
+
assert.match(raw, /BENCHMARK: PASS/)
|
|
20
|
+
const omni = buildPrompt(manifest, 'omni')
|
|
21
|
+
assert.match(omni, /Omni control plane/)
|
|
22
|
+
})
|
|
23
|
+
|
|
24
|
+
test('validateManifests returns errors for invalid manifest', () => {
|
|
25
|
+
const errors = validateManifests([{ id: 'x' }])
|
|
26
|
+
assert.ok(errors.length >= 3)
|
|
27
|
+
assert.equal(validateManifests([{
|
|
28
|
+
id: 'x',
|
|
29
|
+
repo: 'r',
|
|
30
|
+
commit: 'c',
|
|
31
|
+
task: 't',
|
|
32
|
+
acceptance: ['a'],
|
|
33
|
+
runs: 3,
|
|
34
|
+
}]).length, 0)
|
|
35
|
+
})
|
|
36
|
+
|
|
37
|
+
test('readManifests supports single object and array', () => {
|
|
38
|
+
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'omnibench-read-'))
|
|
39
|
+
try {
|
|
40
|
+
const file = path.join(dir, 'manifest.json')
|
|
41
|
+
fs.writeFileSync(file, JSON.stringify({ id: 'one', repo: 'r', commit: 'c', task: 't', acceptance: ['a'], runs: 3 }))
|
|
42
|
+
assert.equal(readManifests(file).length, 1)
|
|
43
|
+
fs.writeFileSync(file, JSON.stringify([{ id: 'one', repo: 'r', commit: 'c', task: 't', acceptance: ['a'], runs: 3 }]))
|
|
44
|
+
assert.equal(readManifests(file).length, 1)
|
|
45
|
+
} finally {
|
|
46
|
+
fs.rmSync(dir, { recursive: true, force: true })
|
|
47
|
+
}
|
|
48
|
+
})
|
|
49
|
+
|
|
50
|
+
test('runCommand captures exit code and output', () => {
|
|
51
|
+
const ok = runCommand('node -e "console.log(1)"', process.cwd())
|
|
52
|
+
assert.equal(ok.exitCode, 0)
|
|
53
|
+
assert.match(ok.output, /1/)
|
|
54
|
+
const bad = runCommand('node -e "process.exit(3)"', process.cwd())
|
|
55
|
+
assert.equal(bad.exitCode, 3)
|
|
56
|
+
})
|
|
57
|
+
|
|
58
|
+
test('writeResults writes JSON result file', () => {
|
|
59
|
+
const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'omnibench-write-'))
|
|
60
|
+
try {
|
|
61
|
+
const file = writeResults([{ id: 'x', success: true }], dir)
|
|
62
|
+
assert.ok(fs.existsSync(file))
|
|
63
|
+
const parsed = JSON.parse(fs.readFileSync(file, 'utf8'))
|
|
64
|
+
assert.equal(parsed[0].success, true)
|
|
65
|
+
} finally {
|
|
66
|
+
fs.rmSync(dir, { recursive: true, force: true })
|
|
67
|
+
}
|
|
68
|
+
})
|