@hecer/yoke 1.10.0 → 1.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +398 -379
- package/README.md +915 -915
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +41 -41
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +56 -56
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/gemini-rtk-hook.mjs +25 -25
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/contracts.js +1 -1
- package/dist/agents/host.js +4 -0
- package/dist/agents/providers.js +13 -0
- package/dist/agents/telemetry.js +33 -0
- package/dist/canon/manifest.js +1 -1
- package/dist/cli.js +9 -9
- package/dist/dashboard/discovery.js +73 -0
- package/dist/dashboard/page.js +122 -122
- package/dist/dashboard/panels.js +91 -91
- package/dist/goals/command.js +2 -2
- package/dist/loop/claims.js +1 -1
- package/dist/loop/decision.js +2 -2
- package/dist/loop/prd.js +1 -1
- package/dist/prd/command.js +17 -17
- package/dist/quality/types.js +1 -1
- package/dist/retrofit/plan.js +2 -0
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/planners/qwen.js +73 -0
- package/dist/retrofit/preserve.js +2 -2
- package/dist/retrofit/skill-actions.js +1 -0
- package/dist/review/command.js +1 -1
- package/dist/routing/capability.js +1 -1
- package/dist/routing/router.js +1 -1
- package/dist/setup/command.js +9 -3
- package/docs/CAPABILITY-ROUTING.md +51 -51
- package/docs/DASHBOARD-EVOLUTION.md +33 -33
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PRODUCT-DIRECTION-2026-09-05.md +210 -210
- package/docs/PUBLISHING.md +114 -114
- package/docs/VERIFIED-PROJECTS-VALIDATION.md +29 -29
- package/docs/VERIFIED-PROJECTS.md +167 -167
- package/docs/community-outreach-2026-08-20.md +85 -0
- package/docs/launch-copy-2026-08-21.md +193 -0
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/plans/2026-09-05-verified-projects.md +83 -83
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +87 -87
|
@@ -1,46 +1,46 @@
|
|
|
1
|
-
{
|
|
2
|
-
"schemaVersion": 1,
|
|
3
|
-
"fixtureVersion": "string-kit@1",
|
|
4
|
-
"runner": "gemini",
|
|
5
|
-
"sampleLabel": "matrix-2026-07-27T18:03:26.202Z",
|
|
6
|
-
"permissionProfile": "safe",
|
|
7
|
-
"yokeVersion": "1.0.0",
|
|
8
|
-
"fixture": "string-kit",
|
|
9
|
-
"startedAt": "2026-07-27T18:03:45.223Z",
|
|
10
|
-
"wallClockMs": 13844,
|
|
11
|
-
"exitCode": 1,
|
|
12
|
-
"finalState": "blocked",
|
|
13
|
-
"verdict": "blocked",
|
|
14
|
-
"blocker": "Gemini runner exited before implementation; fixture verification remained red.",
|
|
15
|
-
"conflicts": 0,
|
|
16
|
-
"iterations": 1,
|
|
17
|
-
"finalTestsPass": false,
|
|
18
|
-
"progress": {
|
|
19
|
-
"passed": 0,
|
|
20
|
-
"total": 3
|
|
21
|
-
},
|
|
22
|
-
"usageAvailable": false,
|
|
23
|
-
"modelAvailable": false,
|
|
24
|
-
"tokens": null,
|
|
25
|
-
"stories": [
|
|
26
|
-
{
|
|
27
|
-
"id": "STORY-1",
|
|
28
|
-
"durationMs": 7040,
|
|
29
|
-
"iterations": 1,
|
|
30
|
-
"finalTestsPass": false
|
|
31
|
-
},
|
|
32
|
-
{
|
|
33
|
-
"id": "STORY-2",
|
|
34
|
-
"durationMs": null,
|
|
35
|
-
"iterations": 0,
|
|
36
|
-
"finalTestsPass": false
|
|
37
|
-
},
|
|
38
|
-
{
|
|
39
|
-
"id": "STORY-3",
|
|
40
|
-
"durationMs": null,
|
|
41
|
-
"iterations": 0,
|
|
42
|
-
"finalTestsPass": false
|
|
43
|
-
}
|
|
44
|
-
],
|
|
45
|
-
"srcLoc": 3
|
|
46
|
-
}
|
|
1
|
+
{
|
|
2
|
+
"schemaVersion": 1,
|
|
3
|
+
"fixtureVersion": "string-kit@1",
|
|
4
|
+
"runner": "gemini",
|
|
5
|
+
"sampleLabel": "matrix-2026-07-27T18:03:26.202Z",
|
|
6
|
+
"permissionProfile": "safe",
|
|
7
|
+
"yokeVersion": "1.0.0",
|
|
8
|
+
"fixture": "string-kit",
|
|
9
|
+
"startedAt": "2026-07-27T18:03:45.223Z",
|
|
10
|
+
"wallClockMs": 13844,
|
|
11
|
+
"exitCode": 1,
|
|
12
|
+
"finalState": "blocked",
|
|
13
|
+
"verdict": "blocked",
|
|
14
|
+
"blocker": "Gemini runner exited before implementation; fixture verification remained red.",
|
|
15
|
+
"conflicts": 0,
|
|
16
|
+
"iterations": 1,
|
|
17
|
+
"finalTestsPass": false,
|
|
18
|
+
"progress": {
|
|
19
|
+
"passed": 0,
|
|
20
|
+
"total": 3
|
|
21
|
+
},
|
|
22
|
+
"usageAvailable": false,
|
|
23
|
+
"modelAvailable": false,
|
|
24
|
+
"tokens": null,
|
|
25
|
+
"stories": [
|
|
26
|
+
{
|
|
27
|
+
"id": "STORY-1",
|
|
28
|
+
"durationMs": 7040,
|
|
29
|
+
"iterations": 1,
|
|
30
|
+
"finalTestsPass": false
|
|
31
|
+
},
|
|
32
|
+
{
|
|
33
|
+
"id": "STORY-2",
|
|
34
|
+
"durationMs": null,
|
|
35
|
+
"iterations": 0,
|
|
36
|
+
"finalTestsPass": false
|
|
37
|
+
},
|
|
38
|
+
{
|
|
39
|
+
"id": "STORY-3",
|
|
40
|
+
"durationMs": null,
|
|
41
|
+
"iterations": 0,
|
|
42
|
+
"finalTestsPass": false
|
|
43
|
+
}
|
|
44
|
+
],
|
|
45
|
+
"srcLoc": 3
|
|
46
|
+
}
|
package/bench/run-matrix.mjs
CHANGED
|
@@ -1,26 +1,26 @@
|
|
|
1
|
-
#!/usr/bin/env node
|
|
2
|
-
import { spawnSync } from 'node:child_process'
|
|
3
|
-
import { mkdirSync, writeFileSync } from 'node:fs'
|
|
4
|
-
import { dirname, join } from 'node:path'
|
|
5
|
-
import { fileURLToPath } from 'node:url'
|
|
6
|
-
import { validateResult } from './result-schema.mjs'
|
|
7
|
-
|
|
8
|
-
const benchDir = dirname(fileURLToPath(import.meta.url))
|
|
9
|
-
if (process.argv.includes('--help')) {
|
|
10
|
-
console.log('usage: node bench/run-matrix.mjs [--label=sample]')
|
|
11
|
-
process.exit(0)
|
|
12
|
-
}
|
|
13
|
-
const label = process.argv.find(arg => arg.startsWith('--label='))?.slice(8) ?? `matrix-${new Date().toISOString()}`
|
|
14
|
-
mkdirSync(join(benchDir, 'results'), { recursive: true })
|
|
15
|
-
for (const runner of ['claude', 'codex', 'gemini']) {
|
|
16
|
-
const probe = spawnSync(runner, ['--version'], { encoding: 'utf8', shell: process.platform === 'win32', timeout: 20_000 })
|
|
17
|
-
if (probe.status !== 0) {
|
|
18
|
-
const row = validateResult({ schemaVersion: 1, fixtureVersion: 'string-kit@1', runner, sampleLabel: label, permissionProfile: 'safe', usageAvailable: false, modelAvailable: false, verdict: 'unavailable', blocker: (probe.stderr || probe.error?.message || 'CLI unavailable').trim(), conflicts: 0, wallClockMs: null, iterations: 0, finalTestsPass: false })
|
|
19
|
-
const out = join(benchDir, 'results', `${runner}-unavailable-${Date.now()}.json`)
|
|
20
|
-
writeFileSync(out, JSON.stringify(row, null, 2) + '\n')
|
|
21
|
-
console.log(JSON.stringify(row))
|
|
22
|
-
continue
|
|
23
|
-
}
|
|
24
|
-
const run = spawnSync(process.execPath, [join(benchDir, 'run.mjs'), `--runner=${runner}`, `--label=${label}`], { encoding: 'utf8', stdio: ['ignore', 'pipe', 'inherit'] })
|
|
25
|
-
if (run.stdout) process.stdout.write(run.stdout)
|
|
26
|
-
}
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
import { spawnSync } from 'node:child_process'
|
|
3
|
+
import { mkdirSync, writeFileSync } from 'node:fs'
|
|
4
|
+
import { dirname, join } from 'node:path'
|
|
5
|
+
import { fileURLToPath } from 'node:url'
|
|
6
|
+
import { validateResult } from './result-schema.mjs'
|
|
7
|
+
|
|
8
|
+
const benchDir = dirname(fileURLToPath(import.meta.url))
|
|
9
|
+
if (process.argv.includes('--help')) {
|
|
10
|
+
console.log('usage: node bench/run-matrix.mjs [--label=sample]')
|
|
11
|
+
process.exit(0)
|
|
12
|
+
}
|
|
13
|
+
const label = process.argv.find(arg => arg.startsWith('--label='))?.slice(8) ?? `matrix-${new Date().toISOString()}`
|
|
14
|
+
mkdirSync(join(benchDir, 'results'), { recursive: true })
|
|
15
|
+
for (const runner of ['claude', 'codex', 'gemini']) {
|
|
16
|
+
const probe = spawnSync(runner, ['--version'], { encoding: 'utf8', shell: process.platform === 'win32', timeout: 20_000 })
|
|
17
|
+
if (probe.status !== 0) {
|
|
18
|
+
const row = validateResult({ schemaVersion: 1, fixtureVersion: 'string-kit@1', runner, sampleLabel: label, permissionProfile: 'safe', usageAvailable: false, modelAvailable: false, verdict: 'unavailable', blocker: (probe.stderr || probe.error?.message || 'CLI unavailable').trim(), conflicts: 0, wallClockMs: null, iterations: 0, finalTestsPass: false })
|
|
19
|
+
const out = join(benchDir, 'results', `${runner}-unavailable-${Date.now()}.json`)
|
|
20
|
+
writeFileSync(out, JSON.stringify(row, null, 2) + '\n')
|
|
21
|
+
console.log(JSON.stringify(row))
|
|
22
|
+
continue
|
|
23
|
+
}
|
|
24
|
+
const run = spawnSync(process.execPath, [join(benchDir, 'run.mjs'), `--runner=${runner}`, `--label=${label}`], { encoding: 'utf8', stdio: ['ignore', 'pipe', 'inherit'] })
|
|
25
|
+
if (run.stdout) process.stdout.write(run.stdout)
|
|
26
|
+
}
|
package/bench/run.mjs
CHANGED
|
@@ -1,35 +1,35 @@
|
|
|
1
|
-
#!/usr/bin/env node
|
|
2
|
-
// Yoke benchmark harness.
|
|
3
|
-
//
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Yoke benchmark harness.
|
|
3
|
+
//
|
|
4
4
|
// node bench/run.mjs --runner=claude [--fixture=string-kit] [--routing=on|off|config]
|
|
5
|
-
//
|
|
6
|
-
// Copies the fixture into bench/.runs/<runner>-<stamp>, git-inits it, then drives
|
|
7
|
-
// `yoke loop run --json` and measures from the OUTSIDE (the loop itself records no
|
|
5
|
+
//
|
|
6
|
+
// Copies the fixture into bench/.runs/<runner>-<stamp>, git-inits it, then drives
|
|
7
|
+
// `yoke loop run --json` and measures from the OUTSIDE (the loop itself records no
|
|
8
8
|
// durations): per-story wall-clock from NDJSON event timestamps, tokens/model from
|
|
9
9
|
// provider telemetry when available, and quality as the fixture's own
|
|
10
|
-
// pre-written tests — run per story AFTER the loop finishes, on the final tree.
|
|
11
|
-
import { spawn, spawnSync } from 'node:child_process'
|
|
10
|
+
// pre-written tests — run per story AFTER the loop finishes, on the final tree.
|
|
11
|
+
import { spawn, spawnSync } from 'node:child_process'
|
|
12
12
|
import { cpSync, existsSync, mkdirSync, readFileSync, writeFileSync, readdirSync, statSync } from 'node:fs'
|
|
13
|
-
import { join, dirname } from 'node:path'
|
|
14
|
-
import { fileURLToPath } from 'node:url'
|
|
13
|
+
import { join, dirname } from 'node:path'
|
|
14
|
+
import { fileURLToPath } from 'node:url'
|
|
15
15
|
import { validateResult } from './result-schema.mjs'
|
|
16
16
|
import { parse } from 'yaml'
|
|
17
|
-
|
|
18
|
-
const benchDir = dirname(fileURLToPath(import.meta.url))
|
|
19
|
-
const repoRoot = dirname(benchDir)
|
|
20
|
-
const cli = join(repoRoot, 'dist', 'cli.js')
|
|
21
|
-
|
|
22
|
-
const args = Object.fromEntries(
|
|
23
|
-
process.argv.slice(2).filter(a => a.startsWith('--')).map(a => {
|
|
24
|
-
const [k, v] = a.slice(2).split('=')
|
|
25
|
-
return [k, v ?? true]
|
|
26
|
-
}),
|
|
27
|
-
)
|
|
28
|
-
const runner = args.runner
|
|
29
|
-
if (!['claude', 'codex', 'gemini'].includes(runner)) {
|
|
17
|
+
|
|
18
|
+
const benchDir = dirname(fileURLToPath(import.meta.url))
|
|
19
|
+
const repoRoot = dirname(benchDir)
|
|
20
|
+
const cli = join(repoRoot, 'dist', 'cli.js')
|
|
21
|
+
|
|
22
|
+
const args = Object.fromEntries(
|
|
23
|
+
process.argv.slice(2).filter(a => a.startsWith('--')).map(a => {
|
|
24
|
+
const [k, v] = a.slice(2).split('=')
|
|
25
|
+
return [k, v ?? true]
|
|
26
|
+
}),
|
|
27
|
+
)
|
|
28
|
+
const runner = args.runner
|
|
29
|
+
if (!['claude', 'codex', 'gemini'].includes(runner)) {
|
|
30
30
|
console.error('usage: node bench/run.mjs --runner=<claude|codex|gemini> [--fixture=string-kit] [--routing=on|off|config] [--run-root=path] [--unsafe] [--max=6] [--timeout=10] [--label=note]')
|
|
31
|
-
process.exit(2)
|
|
32
|
-
}
|
|
31
|
+
process.exit(2)
|
|
32
|
+
}
|
|
33
33
|
const max = Number(args.max ?? 6)
|
|
34
34
|
const timeout = Number(args.timeout ?? 10)
|
|
35
35
|
const fixture = String(args.fixture ?? 'string-kit')
|
|
@@ -49,100 +49,100 @@ const runRoot = args['run-root'] ? String(args['run-root']) : join(benchDir, '.r
|
|
|
49
49
|
const runDir = join(runRoot, `${fixture}-${runner}-routing-${routing}-${stamp}`)
|
|
50
50
|
mkdirSync(runDir, { recursive: true })
|
|
51
51
|
cpSync(fixtureDir, runDir, { recursive: true })
|
|
52
|
-
|
|
53
|
-
const git = (...a) => {
|
|
54
|
-
const r = spawnSync('git', ['-C', runDir, ...a], { encoding: 'utf8' })
|
|
55
|
-
if (r.status !== 0) throw new Error(`git ${a.join(' ')} failed: ${r.stderr}`)
|
|
56
|
-
}
|
|
57
|
-
git('init', '-q')
|
|
58
|
-
git('-c', 'user.name=bench', '-c', 'user.email=bench@yoke', 'add', '-A')
|
|
59
|
-
git('-c', 'user.name=bench', '-c', 'user.email=bench@yoke', 'commit', '-q', '-m', 'bench: fixture baseline')
|
|
60
|
-
|
|
61
|
-
// A nested Claude Code session refuses some operations; scrub session markers.
|
|
62
|
-
const env = { ...process.env }
|
|
63
|
-
for (const k of Object.keys(env)) if (k.startsWith('CLAUDE_CODE') || k === 'CLAUDECODE') delete env[k]
|
|
64
|
-
|
|
52
|
+
|
|
53
|
+
const git = (...a) => {
|
|
54
|
+
const r = spawnSync('git', ['-C', runDir, ...a], { encoding: 'utf8' })
|
|
55
|
+
if (r.status !== 0) throw new Error(`git ${a.join(' ')} failed: ${r.stderr}`)
|
|
56
|
+
}
|
|
57
|
+
git('init', '-q')
|
|
58
|
+
git('-c', 'user.name=bench', '-c', 'user.email=bench@yoke', 'add', '-A')
|
|
59
|
+
git('-c', 'user.name=bench', '-c', 'user.email=bench@yoke', 'commit', '-q', '-m', 'bench: fixture baseline')
|
|
60
|
+
|
|
61
|
+
// A nested Claude Code session refuses some operations; scrub session markers.
|
|
62
|
+
const env = { ...process.env }
|
|
63
|
+
for (const k of Object.keys(env)) if (k.startsWith('CLAUDE_CODE') || k === 'CLAUDECODE') delete env[k]
|
|
64
|
+
|
|
65
65
|
console.error(`[bench] ${runner} · fixture=${fixture} · routing=${routing} → ${runDir}`)
|
|
66
|
-
const t0 = Date.now()
|
|
67
|
-
const events = []
|
|
66
|
+
const t0 = Date.now()
|
|
67
|
+
const events = []
|
|
68
68
|
const routingFlag = routing === 'on' ? ['--routing'] : routing === 'off' ? ['--no-routing'] : []
|
|
69
69
|
const permissionFlag = args.unsafe ? ['--unsafe'] : []
|
|
70
70
|
const child = spawn(process.execPath, [cli, 'loop', 'run', runDir, '--json', `--runner=${runner}`, `--max=${max}`, `--timeout=${timeout}`, ...routingFlag, ...permissionFlag], {
|
|
71
|
-
env, stdio: ['ignore', 'pipe', 'inherit'],
|
|
72
|
-
})
|
|
73
|
-
let buf = ''
|
|
74
|
-
child.stdout.on('data', d => {
|
|
75
|
-
buf += d
|
|
76
|
-
let i
|
|
77
|
-
while ((i = buf.indexOf('\n')) >= 0) {
|
|
78
|
-
const line = buf.slice(0, i).trim()
|
|
79
|
-
buf = buf.slice(i + 1)
|
|
80
|
-
if (!line) continue
|
|
81
|
-
try { events.push({ at: Date.now(), ...JSON.parse(line) }) } catch { /* non-JSON noise */ }
|
|
82
|
-
}
|
|
83
|
-
})
|
|
84
|
-
const exitCode = await new Promise(res => child.on('close', res))
|
|
85
|
-
const wallClockMs = Date.now() - t0
|
|
86
|
-
|
|
87
|
-
// Per-story duration: first event mentioning the story -> first event mentioning the next story (or end).
|
|
71
|
+
env, stdio: ['ignore', 'pipe', 'inherit'],
|
|
72
|
+
})
|
|
73
|
+
let buf = ''
|
|
74
|
+
child.stdout.on('data', d => {
|
|
75
|
+
buf += d
|
|
76
|
+
let i
|
|
77
|
+
while ((i = buf.indexOf('\n')) >= 0) {
|
|
78
|
+
const line = buf.slice(0, i).trim()
|
|
79
|
+
buf = buf.slice(i + 1)
|
|
80
|
+
if (!line) continue
|
|
81
|
+
try { events.push({ at: Date.now(), ...JSON.parse(line) }) } catch { /* non-JSON noise */ }
|
|
82
|
+
}
|
|
83
|
+
})
|
|
84
|
+
const exitCode = await new Promise(res => child.on('close', res))
|
|
85
|
+
const wallClockMs = Date.now() - t0
|
|
86
|
+
|
|
87
|
+
// Per-story duration: first event mentioning the story -> first event mentioning the next story (or end).
|
|
88
88
|
const prd = parse(readFileSync(join(runDir, '.yoke', 'prd.yaml'), 'utf8'))
|
|
89
89
|
const storyIds = Array.isArray(prd) ? prd.map(story => String(story.id)) : []
|
|
90
90
|
const benchConfig = parse(readFileSync(join(runDir, '.yoke', 'config.yaml'), 'utf8'))
|
|
91
91
|
const requestedParentModel = benchConfig?.runner?.model ?? null
|
|
92
|
-
const firstSeen = {}
|
|
93
|
-
for (const e of events) if (e.story && !(e.story in firstSeen)) firstSeen[e.story] = e.at
|
|
94
|
-
const stories = storyIds.map((id, idx) => {
|
|
95
|
-
const start = firstSeen[id]
|
|
96
|
-
const next = storyIds.slice(idx + 1).map(n => firstSeen[n]).find(v => v !== undefined)
|
|
97
|
-
const durationMs = start === undefined ? null : (next ?? t0 + wallClockMs) - start
|
|
98
|
-
const iterations = new Set(events.filter(e => e.story === id).map(e => e.iteration)).size
|
|
99
|
-
// Quality: the fixture's own tests for this story, on the final tree.
|
|
100
|
-
const q = spawnSync(process.execPath, ['--test', `tests/${id}.test.mjs`], { cwd: runDir, encoding: 'utf8' })
|
|
101
|
-
return { id, durationMs, iterations, finalTestsPass: q.status === 0 }
|
|
102
|
-
})
|
|
103
|
-
|
|
104
|
-
const last = events[events.length - 1] ?? {}
|
|
105
|
-
let status = {}
|
|
106
|
-
try { status = JSON.parse(readFileSync(join(runDir, '.yoke', 'loop-status.json'), 'utf8')) } catch { /* loop may have refused before writing status */ }
|
|
107
|
-
|
|
108
|
-
// Source size (LOC in src/) as a code-economy proxy.
|
|
109
|
-
const loc = (dir) => readdirSync(dir).reduce((n, f) => {
|
|
110
|
-
const p = join(dir, f)
|
|
111
|
-
if (statSync(p).isDirectory()) return n + loc(p)
|
|
112
|
-
return n + readFileSync(p, 'utf8').split('\n').filter(l => l.trim() !== '').length
|
|
113
|
-
}, 0)
|
|
114
|
-
|
|
115
|
-
const result = {
|
|
116
|
-
schemaVersion: 1,
|
|
92
|
+
const firstSeen = {}
|
|
93
|
+
for (const e of events) if (e.story && !(e.story in firstSeen)) firstSeen[e.story] = e.at
|
|
94
|
+
const stories = storyIds.map((id, idx) => {
|
|
95
|
+
const start = firstSeen[id]
|
|
96
|
+
const next = storyIds.slice(idx + 1).map(n => firstSeen[n]).find(v => v !== undefined)
|
|
97
|
+
const durationMs = start === undefined ? null : (next ?? t0 + wallClockMs) - start
|
|
98
|
+
const iterations = new Set(events.filter(e => e.story === id).map(e => e.iteration)).size
|
|
99
|
+
// Quality: the fixture's own tests for this story, on the final tree.
|
|
100
|
+
const q = spawnSync(process.execPath, ['--test', `tests/${id}.test.mjs`], { cwd: runDir, encoding: 'utf8' })
|
|
101
|
+
return { id, durationMs, iterations, finalTestsPass: q.status === 0 }
|
|
102
|
+
})
|
|
103
|
+
|
|
104
|
+
const last = events[events.length - 1] ?? {}
|
|
105
|
+
let status = {}
|
|
106
|
+
try { status = JSON.parse(readFileSync(join(runDir, '.yoke', 'loop-status.json'), 'utf8')) } catch { /* loop may have refused before writing status */ }
|
|
107
|
+
|
|
108
|
+
// Source size (LOC in src/) as a code-economy proxy.
|
|
109
|
+
const loc = (dir) => readdirSync(dir).reduce((n, f) => {
|
|
110
|
+
const p = join(dir, f)
|
|
111
|
+
if (statSync(p).isDirectory()) return n + loc(p)
|
|
112
|
+
return n + readFileSync(p, 'utf8').split('\n').filter(l => l.trim() !== '').length
|
|
113
|
+
}, 0)
|
|
114
|
+
|
|
115
|
+
const result = {
|
|
116
|
+
schemaVersion: 1,
|
|
117
117
|
fixtureVersion: `${fixture}@1`,
|
|
118
|
-
runner,
|
|
119
|
-
sampleLabel: String(args.label ?? `${runner}-${stamp}`),
|
|
120
|
-
permissionProfile: args.unsafe ? 'unsafe' : 'safe',
|
|
121
|
-
yokeVersion: JSON.parse(readFileSync(join(repoRoot, 'package.json'), 'utf8')).version,
|
|
118
|
+
runner,
|
|
119
|
+
sampleLabel: String(args.label ?? `${runner}-${stamp}`),
|
|
120
|
+
permissionProfile: args.unsafe ? 'unsafe' : 'safe',
|
|
121
|
+
yokeVersion: JSON.parse(readFileSync(join(repoRoot, 'package.json'), 'utf8')).version,
|
|
122
122
|
fixture,
|
|
123
123
|
routing,
|
|
124
|
-
startedAt: new Date(t0).toISOString(),
|
|
125
|
-
wallClockMs,
|
|
126
|
-
exitCode,
|
|
127
|
-
finalState: last.state ?? null,
|
|
128
|
-
verdict: exitCode === 0 ? 'completed' : (/api key|login|auth/i.test(String(status.reason ?? '')) ? 'auth-failed' : 'blocked'),
|
|
129
|
-
blocker: exitCode === 0 ? null : (status.reason ?? 'runner exited without a diagnostic'),
|
|
130
|
-
conflicts: events.filter(e => /conflict/i.test(String(e.reason ?? e.summary ?? ''))).length,
|
|
131
|
-
iterations: stories.reduce((sum, story) => sum + story.iterations, 0),
|
|
132
|
-
finalTestsPass: stories.every(story => story.finalTestsPass),
|
|
133
|
-
progress: last.progress ?? null,
|
|
134
|
-
usageAvailable: Number(status.tokens?.inputTokens ?? 0) + Number(status.tokens?.outputTokens ?? 0) > 0,
|
|
124
|
+
startedAt: new Date(t0).toISOString(),
|
|
125
|
+
wallClockMs,
|
|
126
|
+
exitCode,
|
|
127
|
+
finalState: last.state ?? null,
|
|
128
|
+
verdict: exitCode === 0 ? 'completed' : (/api key|login|auth/i.test(String(status.reason ?? '')) ? 'auth-failed' : 'blocked'),
|
|
129
|
+
blocker: exitCode === 0 ? null : (status.reason ?? 'runner exited without a diagnostic'),
|
|
130
|
+
conflicts: events.filter(e => /conflict/i.test(String(e.reason ?? e.summary ?? ''))).length,
|
|
131
|
+
iterations: stories.reduce((sum, story) => sum + story.iterations, 0),
|
|
132
|
+
finalTestsPass: stories.every(story => story.finalTestsPass),
|
|
133
|
+
progress: last.progress ?? null,
|
|
134
|
+
usageAvailable: Number(status.tokens?.inputTokens ?? 0) + Number(status.tokens?.outputTokens ?? 0) > 0,
|
|
135
135
|
modelAvailable: (typeof status.tokens?.model === 'string' && status.tokens.model !== '<synthetic>') || requestedParentModel !== null,
|
|
136
136
|
requestedParentModel,
|
|
137
137
|
tokens: status.tokens ?? null,
|
|
138
138
|
modelCalls: status.tokens?.calls ?? [],
|
|
139
|
-
stories,
|
|
140
|
-
srcLoc: loc(join(runDir, 'src')),
|
|
141
|
-
}
|
|
142
|
-
validateResult(result)
|
|
143
|
-
|
|
144
|
-
mkdirSync(join(benchDir, 'results'), { recursive: true })
|
|
139
|
+
stories,
|
|
140
|
+
srcLoc: loc(join(runDir, 'src')),
|
|
141
|
+
}
|
|
142
|
+
validateResult(result)
|
|
143
|
+
|
|
144
|
+
mkdirSync(join(benchDir, 'results'), { recursive: true })
|
|
145
145
|
const out = join(benchDir, 'results', `${fixture}-${runner}-routing-${routing}-${stamp}.json`)
|
|
146
|
-
writeFileSync(out, JSON.stringify(result, null, 2) + '\n')
|
|
147
|
-
console.error(`[bench] done: ${out}`)
|
|
148
|
-
console.log(JSON.stringify(result, null, 2))
|
|
146
|
+
writeFileSync(out, JSON.stringify(result, null, 2) + '\n')
|
|
147
|
+
console.error(`[bench] done: ${out}`)
|
|
148
|
+
console.log(JSON.stringify(result, null, 2))
|
package/canon/AGENTS.md
CHANGED
|
@@ -1,30 +1,30 @@
|
|
|
1
|
-
# Yoke Harness — Agent Baseline
|
|
2
|
-
|
|
3
|
-
You are operating in a project retrofitted by Yoke. Follow these always:
|
|
4
|
-
|
|
5
|
-
- **Quality first:** Test-driven development is the default. No production code without a failing test first. See skill `tdd`.
|
|
6
|
-
- **Stop-the-Line:** Do not start implementation until Definition of Done / Acceptance Criteria are written. See `policy/gates.md`.
|
|
7
|
-
- **Role separation:** The agent that implements does not self-review, self-merge, or self-audit security. See `policy/roles.md`.
|
|
8
|
-
- **Context efficiency:** Prefer the wired tools (rtk for command output, the code-graph for symbol lookup) over reading whole files. See `tools/`.
|
|
9
|
-
|
|
10
|
-
This file is the portable baseline. Agent-specific instructions are generated alongside it (CLAUDE.md, GEMINI.md).
|
|
11
|
-
|
|
12
|
-
## Skill routing & precedence
|
|
13
|
-
|
|
14
|
-
When several skills could match the same task, resolve deterministically:
|
|
15
|
-
|
|
16
|
-
1. **Methodology before role.** Skills that decide *how* to work (`brainstorming`, `writing-plans`,
|
|
17
|
-
`tdd`, `subagent-driven-development`, `systematic-debugging`, …) take precedence and set the
|
|
18
|
-
process. Role skills (`review`, `ship`, `health`, `retro`, …) add a perspective on top.
|
|
19
|
-
2. **One canonical entrypoint per concern** — pick the most specific:
|
|
20
|
-
- Set up or update Yoke → `yoke-retrofit`
|
|
21
|
-
- Yoke-owned planning + autonomous story execution → `yoke-workflow`
|
|
22
|
-
- Plan-time architecture review → `plan-eng-review`
|
|
23
|
-
- Plan-time product / scope review → `plan-ceo-review`
|
|
24
|
-
- **Pre-merge code review → `review`** (the single canonical one)
|
|
25
|
-
- Requesting a review (dispatch a reviewer) → `requesting-code-review`
|
|
26
|
-
- Handling review feedback → `receiving-code-review`
|
|
27
|
-
- Overall order of operations (idea → deploy) → `workflow`
|
|
28
|
-
3. **Don't double-run.** These skills declare their own triggers aggressively; when more than one
|
|
29
|
-
matches, the precedence above and the most-specific entrypoint decide. Do not run two skills
|
|
30
|
-
that serve the same concern on the same task.
|
|
1
|
+
# Yoke Harness — Agent Baseline
|
|
2
|
+
|
|
3
|
+
You are operating in a project retrofitted by Yoke. Follow these always:
|
|
4
|
+
|
|
5
|
+
- **Quality first:** Test-driven development is the default. No production code without a failing test first. See skill `tdd`.
|
|
6
|
+
- **Stop-the-Line:** Do not start implementation until Definition of Done / Acceptance Criteria are written. See `policy/gates.md`.
|
|
7
|
+
- **Role separation:** The agent that implements does not self-review, self-merge, or self-audit security. See `policy/roles.md`.
|
|
8
|
+
- **Context efficiency:** Prefer the wired tools (rtk for command output, the code-graph for symbol lookup) over reading whole files. See `tools/`.
|
|
9
|
+
|
|
10
|
+
This file is the portable baseline. Agent-specific instructions are generated alongside it (CLAUDE.md, GEMINI.md).
|
|
11
|
+
|
|
12
|
+
## Skill routing & precedence
|
|
13
|
+
|
|
14
|
+
When several skills could match the same task, resolve deterministically:
|
|
15
|
+
|
|
16
|
+
1. **Methodology before role.** Skills that decide *how* to work (`brainstorming`, `writing-plans`,
|
|
17
|
+
`tdd`, `subagent-driven-development`, `systematic-debugging`, …) take precedence and set the
|
|
18
|
+
process. Role skills (`review`, `ship`, `health`, `retro`, …) add a perspective on top.
|
|
19
|
+
2. **One canonical entrypoint per concern** — pick the most specific:
|
|
20
|
+
- Set up or update Yoke → `yoke-retrofit`
|
|
21
|
+
- Yoke-owned planning + autonomous story execution → `yoke-workflow`
|
|
22
|
+
- Plan-time architecture review → `plan-eng-review`
|
|
23
|
+
- Plan-time product / scope review → `plan-ceo-review`
|
|
24
|
+
- **Pre-merge code review → `review`** (the single canonical one)
|
|
25
|
+
- Requesting a review (dispatch a reviewer) → `requesting-code-review`
|
|
26
|
+
- Handling review feedback → `receiving-code-review`
|
|
27
|
+
- Overall order of operations (idea → deploy) → `workflow`
|
|
28
|
+
3. **Don't double-run.** These skills declare their own triggers aggressively; when more than one
|
|
29
|
+
matches, the precedence above and the most-specific entrypoint decide. Do not run two skills
|
|
30
|
+
that serve the same concern on the same task.
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Decisions
|
|
2
|
-
|
|
3
|
-
> Append-only ledger. The Yoke loop adds an entry per completed story; agents add
|
|
4
|
-
> entries for non-obvious calls made during interactive work. Newest at the bottom.
|
|
1
|
+
# Decisions
|
|
2
|
+
|
|
3
|
+
> Append-only ledger. The Yoke loop adds an entry per completed story; agents add
|
|
4
|
+
> entries for non-obvious calls made during interactive work. Newest at the bottom.
|
|
@@ -1,11 +1,11 @@
|
|
|
1
|
-
# Glossary
|
|
2
|
-
|
|
3
|
-
Record canonical project-specific domain terms here when they are settled.
|
|
4
|
-
|
|
5
|
-
## Language
|
|
6
|
-
|
|
7
|
-
<!--
|
|
8
|
-
**Canonical term**
|
|
9
|
-
: One or two sentences defining what the term is in this project.
|
|
10
|
-
_Avoid_: ambiguous alias, rejected synonym
|
|
11
|
-
-->
|
|
1
|
+
# Glossary
|
|
2
|
+
|
|
3
|
+
Record canonical project-specific domain terms here when they are settled.
|
|
4
|
+
|
|
5
|
+
## Language
|
|
6
|
+
|
|
7
|
+
<!--
|
|
8
|
+
**Canonical term**
|
|
9
|
+
: One or two sentences defining what the term is in this project.
|
|
10
|
+
_Avoid_: ambiguous alias, rejected synonym
|
|
11
|
+
-->
|
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Knowledge
|
|
2
|
-
|
|
3
|
-
> Reusable gotchas, conventions, and learnings. Add a bullet whenever you discover
|
|
4
|
-
> something a future agent would waste time rediscovering.
|
|
1
|
+
# Knowledge
|
|
2
|
+
|
|
3
|
+
> Reusable gotchas, conventions, and learnings. Add a bullet whenever you discover
|
|
4
|
+
> something a future agent would waste time rediscovering.
|
package/canon/context/PROJECT.md
CHANGED
|
@@ -1,15 +1,15 @@
|
|
|
1
|
-
# Project — North Star
|
|
2
|
-
|
|
3
|
-
> Edit this file. It is the durable goal every Yoke agent reads before implementing.
|
|
4
|
-
|
|
5
|
-
## Goal
|
|
6
|
-
<!-- One paragraph: what this project is and the outcome it must produce. -->
|
|
7
|
-
|
|
8
|
-
## Constraints
|
|
9
|
-
<!-- Hard limits: stack, platforms, performance, compliance. -->
|
|
10
|
-
|
|
11
|
-
## Non-goals
|
|
12
|
-
<!-- Explicitly out of scope. The most valuable section for preventing drift. -->
|
|
13
|
-
|
|
14
|
-
## Success criteria
|
|
15
|
-
<!-- How we know it works. -->
|
|
1
|
+
# Project — North Star
|
|
2
|
+
|
|
3
|
+
> Edit this file. It is the durable goal every Yoke agent reads before implementing.
|
|
4
|
+
|
|
5
|
+
## Goal
|
|
6
|
+
<!-- One paragraph: what this project is and the outcome it must produce. -->
|
|
7
|
+
|
|
8
|
+
## Constraints
|
|
9
|
+
<!-- Hard limits: stack, platforms, performance, compliance. -->
|
|
10
|
+
|
|
11
|
+
## Non-goals
|
|
12
|
+
<!-- Explicitly out of scope. The most valuable section for preventing drift. -->
|
|
13
|
+
|
|
14
|
+
## Success criteria
|
|
15
|
+
<!-- How we know it works. -->
|