@hecer/yoke 1.18.0 → 1.20.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +46 -0
- package/README.md +86 -910
- package/bench/RESULTS.md +30 -0
- package/bench/analyze-codex-comparison.mjs +22 -0
- package/bench/compare-codex.mjs +71 -0
- package/bench/fixtures/independent-utils/.yoke/prd.yaml +30 -0
- package/bench/fixtures/independent-utils/bench-verify.mjs +4 -0
- package/bench/fixtures/independent-utils/package.json +1 -0
- package/bench/fixtures/independent-utils/src/chunk.mjs +1 -0
- package/bench/fixtures/independent-utils/src/range.mjs +1 -0
- package/bench/fixtures/independent-utils/src/unique.mjs +1 -0
- package/bench/fixtures/independent-utils/tests/STORY-1.test.mjs +8 -0
- package/bench/fixtures/independent-utils/tests/STORY-2.test.mjs +8 -0
- package/bench/fixtures/independent-utils/tests/STORY-3.test.mjs +8 -0
- package/bench/probe-goal-integration.mjs +93 -0
- package/bench/probe-parallel-gate-context.mjs +29 -0
- package/bench/results/codex-comparison-2026-09-29.json +305 -0
- package/bench/results/goal-integration-probes-2026-09-29.json +107 -0
- package/bench/results/native-goal-validation-2026-09-30.json +122 -0
- package/bench/results/parallel-validation-2026-09-30.json +86 -0
- package/canon/manifest.yaml +1 -1
- package/dist/agents/process-record-identity.js +33 -0
- package/dist/agents/process.js +42 -11
- package/dist/agents/providers.js +11 -2
- package/dist/agents/sol-pi-runtime.js +76 -0
- package/dist/check/command.js +218 -0
- package/dist/cli.js +11 -7
- package/dist/code-intelligence/adapters/graphify.js +1 -1
- package/dist/code-intelligence/adapters/mcp.js +25 -4
- package/dist/code-intelligence/adapters/serena.js +1 -1
- package/dist/code-intelligence/coordinator.js +2 -1
- package/dist/code-intelligence/evidence.js +12 -1
- package/dist/code-intelligence/mcp-client.js +2 -2
- package/dist/dashboard/page.js +1 -1
- package/dist/dashboard/panels.js +25 -3
- package/dist/dashboard/server.js +88 -7
- package/dist/goals/codex-native.js +340 -0
- package/dist/goals/command.js +230 -29
- package/dist/loop/cleanup.js +10 -1
- package/dist/loop/dispatcher.js +82 -10
- package/dist/loop/git.js +1 -1
- package/dist/loop/parallel-adapters.js +53 -7
- package/dist/loop/parallel-command.js +8 -1
- package/dist/loop/recovery.js +81 -2
- package/dist/loop/reporter.js +3 -0
- package/dist/loop/resource-pool.js +21 -15
- package/dist/loop/run-command.js +91 -25
- package/dist/loop/run-state.js +116 -0
- package/dist/prd/explore.js +12 -4
- package/dist/retrofit/config.js +29 -0
- package/dist/retrofit/gitignore.js +3 -0
- package/docs/CODEX-COMPARISON-2026-09-29.md +27 -0
- package/docs/CONTINUOUS-EXPLORATION.md +6 -2
- package/docs/GOALS-RESOURCE-AUDIT-2026-09-29.md +92 -0
- package/docs/GOALS.md +109 -0
- package/docs/HARNESSES.md +1 -1
- package/docs/RELEASE-VALIDATION-1.20.0.md +100 -0
- package/docs/SOL-PI.md +41 -0
- package/docs/VERIFIED-PROJECTS.md +4 -4
- package/docs/code-intelligence-stdio-fix-2026-09-27.md +8 -0
- package/docs/parallel-execution.md +15 -0
- package/docs/superpowers/plans/2026-09-30-goals-resources-release.md +78 -0
- package/gemini-extension.json +1 -1
- package/package.json +5 -1
package/docs/SOL-PI.md
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
1
|
+
# SoL-Pi with Yoke
|
|
2
|
+
|
|
3
|
+
Yoke can add NVIDIA's SoL-Pi extension to opted-in Pi invocations. The extension is off by default. This guide records the published evidence and the limits of Yoke's integration so you can decide whether to enable it.
|
|
4
|
+
|
|
5
|
+
## What the paper measured
|
|
6
|
+
|
|
7
|
+
The [SoL-Pi paper](https://arxiv.org/abs/2609.20519) reports benchmark runs on fixed tasks, models, harness configurations, and API prices. “Token traffic” below is recorded model-token traffic; scores are benchmark scores. The authors tested 51 publicly released tasks from EdgeBench, a subset of its 134 tasks.
|
|
8
|
+
|
|
9
|
+
On EdgeBench with GPT-5.6 Sol, the paper reports:
|
|
10
|
+
|
|
11
|
+
| Configuration | Recorded token traffic | API cost | Average score |
|
|
12
|
+
| --- | ---: | ---: | ---: |
|
|
13
|
+
| Pi baseline | 2.1538 billion | $1,339 | 44.833 |
|
|
14
|
+
| SoL-Pi Efficiency, all four mechanisms | 1.0990 billion (49.0% lower) | $894 (33.2% lower) | 42.003 (93.7% of Pi) |
|
|
15
|
+
| SoL-Pi Performance, ObservationPack alone | 2.0224 billion (6.1% lower) | $1,271 | 47.208 (5.3% above Pi) |
|
|
16
|
+
|
|
17
|
+
The complete efficiency configuration used fewer tokens and cost less in those runs, but scored lower than Pi. The separate performance configuration was the best-scoring single mechanism for GPT-5.6 Sol; it was not the full four-mechanism stack.
|
|
18
|
+
|
|
19
|
+
The paper also ran the complete stack with Opus 5, a backend not used to search for the mechanisms. It reports 44.7% less token traffic and 33.5% lower API cost than Pi ($1,158 versus $1,741), with an average score of 42.224 versus Pi's 44.756 (94.3% of Pi's score).
|
|
20
|
+
|
|
21
|
+
Results vary by benchmark. On 63 CPU-only Terminal-Bench 4 tasks, SoL-Pi solved 15 of 63 while Pi solved 18 of 63; total API cost was $211.12 versus $286.45 (26.3% lower), and cost per solved task was $14.07 versus $15.91 (11.6% lower). On six IMO 2026 problems, both systems passed three; SoL-Pi's reported cost per passed problem was $20.90 versus Pi's $25.32.
|
|
22
|
+
|
|
23
|
+
These are the paper's measurements, not measurements of Yoke's pinned extension build. They are not a project-level prediction or guarantee of savings. Your tasks, models, provider prices, cache billing, and enabled mechanisms can produce different cost and capability results. Yoke does not measure or promise a savings percentage for an individual project.
|
|
24
|
+
|
|
25
|
+
## Defaults and runtime
|
|
26
|
+
|
|
27
|
+
The `solpi.enabled` project setting defaults to `false`. Yoke adds the pinned SoL-Pi extension only to a Pi invocation when that setting is enabled. Each mechanism—`actionFusion`, `observationPack`, `evidencePreservingReducer`, and `onlineContextCompact`—also defaults to `false`; enabling the extension does not turn those mechanisms on. `cacheWriteReadRatio` defaults to `12.5` and is an input to the compaction decision, not a live price lookup or bill estimate.
|
|
28
|
+
|
|
29
|
+
Yoke pins SoL-Pi to [`d7ecfc089944f0d04b80122a0a9a6ca0d786f3d0`](https://github.com/NVlabs/SoL-Pi/tree/d7ecfc089944f0d04b80122a0a9a6ca0d786f3d0). The pinned upstream lists Node.js 22.19 or newer and `@earendil-works/pi-coding-agent@0.84.2` as its tested runtime baseline. Yoke checks the Node.js version but does not enforce the Pi version; other Pi releases are unverified. Yoke does not install Pi or set up provider authentication. Review the [pinned extension source](https://github.com/NVlabs/SoL-Pi/tree/d7ecfc089944f0d04b80122a0a9a6ca0d786f3d0) before enabling it.
|
|
30
|
+
|
|
31
|
+
## Trust, permissions, and data
|
|
32
|
+
|
|
33
|
+
Pi project trust and extension approval remain Pi's responsibility. Yoke does not grant trust or add an automatic approval flag. If Pi requires the project to be trusted before loading project-local settings or an extension, trust it through Pi yourself.
|
|
34
|
+
|
|
35
|
+
SoL-Pi runs inside the Pi process with that process's filesystem, process, network, and credential permissions. It is not a separate sandbox or permission boundary. Pi's selected tool permissions still apply: Yoke's safe Pi profile allows `read,bash,edit,write`, while its read-only profile allows `read,grep,find,ls`; the unsafe profile does not add an explicit tool allowlist. Action Fusion can run a model-requested validation command through Pi's shell tool. Review the extension at its pinned revision and choose Pi permissions appropriate for the project before enabling it.
|
|
36
|
+
|
|
37
|
+
When enabled, Evidence-Preserving Reducer may send eligible diagnostic-log content to its configured reducer model using Pi-managed authentication. Logs can contain project data. Its evidence checks validate the reduced receipt against the archived source; they do not make the source private or act as a complete redaction step. Leave the reducer off for logs that must stay local, and follow the data rules for the provider that receives the request. See the upstream [SoL-Pi security policy](https://github.com/NVlabs/SoL-Pi/blob/d7ecfc089944f0d04b80122a0a9a6ca0d786f3d0/SECURITY.md).
|
|
38
|
+
|
|
39
|
+
## Configuration ownership
|
|
40
|
+
|
|
41
|
+
The project's `.yoke/config.yaml` is the durable source for Yoke's SoL-Pi settings; the dashboard edits the same project settings. For an enabled Pi invocation, Yoke writes the corresponding temporary extension configuration to `.pi/sol-pi.json` in that invocation's workspace, the project-local path SoL-Pi reads. When the invocation ends, Yoke restores the file's previous contents or removes the temporary file and the `.pi` directory if Yoke created it. Pi only loads project-local extension configuration for a trusted project. Keep persistent SoL-Pi choices in Yoke's project configuration. Pi installation, project trust, and provider credentials remain owned by Pi and the user. See the pinned upstream [configuration contract](https://github.com/NVlabs/SoL-Pi/blob/d7ecfc089944f0d04b80122a0a9a6ca0d786f3d0/docs/configuration.md).
|
|
@@ -33,7 +33,7 @@ After intentionally editing protected tests, use `yoke check . --protect --refre
|
|
|
33
33
|
## Durable goals across providers
|
|
34
34
|
|
|
35
35
|
```sh
|
|
36
|
-
yoke goal set . --objective="Finish guest checkout" --attempts=3 --minutes=30
|
|
36
|
+
yoke goal set . --objective="Finish guest checkout" --criteria=guest-checkout --attempts=3 --minutes=30
|
|
37
37
|
yoke goal run . --runner=codex
|
|
38
38
|
yoke goal resume . --runner=claude --model=<installed-model-id>
|
|
39
39
|
yoke goal resume . --runner=gemini --model=<installed-model-id>
|
|
@@ -42,9 +42,9 @@ yoke goal handoff .
|
|
|
42
42
|
yoke goal pause .
|
|
43
43
|
```
|
|
44
44
|
|
|
45
|
-
Goals keep objective, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock
|
|
45
|
+
Goals keep objective, explicit acceptance binding, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock and worker capacity. A completed goal is checked again on a later run. Failed work stays in the project; goal execution itself does not commit or publish it. Native Codex goals are optional through `--native-goal`; see the [goal guide](GOALS.md) for capability negotiation and migration of unbound goals.
|
|
46
46
|
|
|
47
|
-
`--minutes` limits cumulative
|
|
47
|
+
`--minutes` limits cumulative admitted provider execution time. Optional `--wall-minutes` also includes capacity waits and independent checks. `--tokens=N` blocks further execution when measured usage is exhausted or unknown, and reports post-call overruns even when checks pass; providers with end-of-call telemetry cannot promise a hard mid-call cap. Interrupted attempts are charged conservatively and recorded process ownership is reconciled before continuation. A pause requests cancellation of Yoke-owned asynchronous execution and retains unfinished work and evidence.
|
|
48
48
|
|
|
49
49
|
Explicitly extend total budgets without deleting history:
|
|
50
50
|
|
|
@@ -65,7 +65,7 @@ Failed or paused isolated story worktrees are retained. Resume the same story wi
|
|
|
65
65
|
yoke loop run . --isolate --resume-worktree --parallel=1 --candidates=1
|
|
66
66
|
```
|
|
67
67
|
|
|
68
|
-
Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery
|
|
68
|
+
Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery. Parallel retained candidates resume through the dispatcher without reimplementation when their recovery binding remains valid. `yoke loop cleanup` is an explicit discard operation. Existing projects should rerun retrofit to add runtime ignore entries for checks, events, goals, saved runs and recovery evidence.
|
|
69
69
|
|
|
70
70
|
## Spend fewer model calls
|
|
71
71
|
|
|
@@ -0,0 +1,8 @@
|
|
|
1
|
+
# Code Intelligence stdio correction
|
|
2
|
+
|
|
3
|
+
Acceptance criteria:
|
|
4
|
+
- Serena and Graphify adapters initialize and call tools against newline-delimited JSON-RPC stdio servers.
|
|
5
|
+
- Existing Graft and client tests remain green; TypeScript builds.
|
|
6
|
+
- ReferDock uses explicit installed executable paths where Windows shims cannot be spawned.
|
|
7
|
+
- Real read-only probes report backend failures honestly; no cloud graph processing or package upgrades are required.
|
|
8
|
+
- Installed runtime changes retain backups and receive independent review before release.
|
|
@@ -80,6 +80,21 @@ one serial integration lane when integration measurements exist. Older runs with
|
|
|
80
80
|
remain visible as missing integration history; forecasts are empirical ranges, not deadlines, and do
|
|
81
81
|
not predict future contention from other projects.
|
|
82
82
|
|
|
83
|
+
### Rejected integration recovery
|
|
84
|
+
|
|
85
|
+
Worker and integrated-tree gates receive the same `YOKE_STORY` context. A rejected
|
|
86
|
+
candidate is retained with its reason and an `integration-recovery.json` proof record;
|
|
87
|
+
it is not silently discarded and regenerated. The next parallel run can reuse it
|
|
88
|
+
without a new implementation model call when canonical project/worktree ownership,
|
|
89
|
+
Git registration, target base and PRD digest still match. Integration gates run again.
|
|
90
|
+
Changed target or PRD state blocks recovery and requires explicit reconciliation.
|
|
91
|
+
Generated worktree names are shorter and Windows path limits are checked before setup.
|
|
92
|
+
|
|
93
|
+
Goals also use the worker pool. Their asynchronous checks use a separate default-one
|
|
94
|
+
check pool; see [goal resources](GOALS.md). These are concurrency permits, not hard CPU
|
|
95
|
+
or RAM quotas. Current controlled comparisons do not establish general efficiency gains
|
|
96
|
+
over direct Codex; see the [measured comparison](CODEX-COMPARISON-2026-09-29.md).
|
|
97
|
+
|
|
83
98
|
## Safe task decomposition
|
|
84
99
|
|
|
85
100
|
Make a preview with:
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# Goals, verification and recovery implementation plan
|
|
2
|
+
|
|
3
|
+
> **For agentic workers:** Implement independent parallel and resume components with dispatching-parallel-agents; integrate goal policy centrally. Write each regression before changing production behavior.
|
|
4
|
+
|
|
5
|
+
**Goal:** Correct the reproduced completion/resource/integration failures, provide optional persistent native Codex Goals, and prepare a validated 1.20.0 release.
|
|
6
|
+
|
|
7
|
+
**Architecture:** Yoke owns objective-specific executable acceptance and project integration. A negotiated optional native adapter owns a Codex thread, paused at verification boundaries. All actual provider calls use the same admission pool; local goal checks have abortable execution and a separate shared check budget. Durable run identity restores the correct mode and limits.
|
|
8
|
+
|
|
9
|
+
**Tech stack:** TypeScript, Node 20+, Zod, Vitest, native CLI/app-server stdio; no new dependency.
|
|
10
|
+
|
|
11
|
+
**Spec:** docs/GOALS-RESOURCE-AUDIT-2026-09-29.md and docs/CODEX-COMPARISON-2026-09-29.md. Recommendations in those dated reports become implementation scope here; historic results remain historic.
|
|
12
|
+
|
|
13
|
+
## Constraints and decisions
|
|
14
|
+
|
|
15
|
+
- Preserve the original dirty checkout; use the existing codex/benchmark-goals-2026-09-29 worktree.
|
|
16
|
+
- Native goals remain optional, capability negotiated, and never required for other providers. No global CLI configuration changes.
|
|
17
|
+
- Bind each goal to explicitly selected criterion IDs and a stable objective/contract digest. Old unbound goals require explicit binding before accepting completion; provide CLI migration command.
|
|
18
|
+
- Preserve --minutes as cumulative agent-time budget; add explicit cumulative wall-time budget. Report token overshoot and unknown consumption, never silently reset them.
|
|
19
|
+
- Disable unmanaged native multi-agent delegation for managed provider calls; reserve one pool unit per actual call, not an outer unit plus nested units.
|
|
20
|
+
- Keep protected acceptance, scope controls, commits and safe-boundary exploration semantics. No weaker test gates or fabricated CPU/USD savings.
|
|
21
|
+
- Prepare the release locally; tagging, npm publication and GitHub release are not performed by a preparation request.
|
|
22
|
+
|
|
23
|
+
## Review focus
|
|
24
|
+
|
|
25
|
+
Goal contract unrelated to new objective; stale/edited contracts; abort while admission is queued; planner/worker token accounting; dashboard resuming a stale goal instead of active explore run; expired exploration deadline; unavailable native protocol versus genuine authentication errors; native complete while Yoke check fails; retained worktree ownership; Windows full generated path length.
|
|
26
|
+
|
|
27
|
+
## Task 1 — Parallel integration and recoverable candidates
|
|
28
|
+
|
|
29
|
+
Files: src/loop/dispatcher.ts, parallel-command.ts, parallel-adapters.ts, merge-queue.ts; tests/loop parallel/dispatcher suites; reporter integration centrally.
|
|
30
|
+
|
|
31
|
+
- [x] Add failing story-environment regression using independent scopes and an environment-dependent verifier.
|
|
32
|
+
- [x] Preserve verification environment across worker and integration phases and restore previous environment on errors.
|
|
33
|
+
- [x] Persist concrete rejection reasons and recoverable candidate ownership; retain implemented work safely and enable existing recovery paths.
|
|
34
|
+
- [x] Shorten generated worktree names without weakening ownership validation; test Windows path failure handling.
|
|
35
|
+
- [x] Run covering tests and deterministic control probe.
|
|
36
|
+
|
|
37
|
+
## Task 2 — Durable run identity and resume
|
|
38
|
+
|
|
39
|
+
Files: src/loop/run-state.ts, run-command.ts, src/dashboard/server.ts, run-state/dashboard tests.
|
|
40
|
+
|
|
41
|
+
- [x] Add failing mode/provider/selection/deadline and unrelated-goal resume tests.
|
|
42
|
+
- [x] Persist validated safe run options under owned project lock; restore the interrupted run's mode.
|
|
43
|
+
- [x] Preserve absolute exploration deadline and user pauses; explicitly fresh CLI run gets a new deadline.
|
|
44
|
+
- [x] Verify expired deadline starts no provider and unsafe permissions do not silently carry over.
|
|
45
|
+
|
|
46
|
+
## Task 3 — Objective acceptance, policy and budgets
|
|
47
|
+
|
|
48
|
+
Files: src/goals/command.ts, contracts.ts/admission.ts as needed, src/check/command.ts async counterpart, src/cli.ts, src/retrofit/config.ts, goals/check tests.
|
|
49
|
+
|
|
50
|
+
- [x] Add failing unrelated-green-contract test, explicit binding/digest mutation tests, saturated pool and configured-selection tests.
|
|
51
|
+
- [x] Bind goal objective to selected criteria and contract revision; add goal bind CLI and documented legacy migration.
|
|
52
|
+
- [x] Resolve explicit provider/model/effort/bare overrides then project runner configuration; centrally disable unmanaged delegation.
|
|
53
|
+
- [x] Admit actual implementation and planner calls through the pool and release in finally.
|
|
54
|
+
- [x] Add abortable checks with shared local check admission; preserve fingerprints and protected contract evidence.
|
|
55
|
+
- [x] Add cumulative wall-time option and post-attempt overshoot/unknown-usage accounting, streamed cancellation where supported; preserve safe pause and interrupted work.
|
|
56
|
+
|
|
57
|
+
## Task 4 — Optional native Codex adapter
|
|
58
|
+
|
|
59
|
+
Files: src/goals/codex-native.ts, tests/goals/codex-native.test.ts; central goal integration/config.
|
|
60
|
+
|
|
61
|
+
- [x] Test initialized RPC transport, unsupported capability, malformed replies, thread start/resume, abort and bounded cleanup.
|
|
62
|
+
- [x] Persist objective-bound thread identity; execute bounded turns with native goals paused between turns so independent checks arbitrate continuation.
|
|
63
|
+
- [x] Synchronize completion/pause/budget/blocker states after Yoke checks; unavailable method can fall back, real execution/auth errors cannot silently downgrade.
|
|
64
|
+
- [x] Test provider switches and existing-thread/model mismatch; prevent competing auto-continuation.
|
|
65
|
+
|
|
66
|
+
## Task 5 — Review, measurements and release preparation
|
|
67
|
+
|
|
68
|
+
- [x] Review merged components, fix covering regressions and rerun deterministic audit probes with updated expected outcomes.
|
|
69
|
+
- [x] Run authenticated corrected parallel benchmark and report accepted quality/time/tokens with original sample limits.
|
|
70
|
+
- [x] Run TypeScript/build, full tests, metadata checks and package dry run in isolated state.
|
|
71
|
+
- [x] Add dated 1.20.0 CHANGELOG entry with implemented behavior, binding migration, default/native limits and actual validation limits.
|
|
72
|
+
- [x] Synchronize package/lock/provider manifest/README versions; link compact guides for goals/resources/recovery.
|
|
73
|
+
- [x] Read-only provenance audit for changed reports; prepare matching release notes and build archive.
|
|
74
|
+
|
|
75
|
+
## Progress ledger
|
|
76
|
+
|
|
77
|
+
- 2026-09-30: Plan created; three independent implementers assigned parallel correctness, durable resume, native adapter. Controller owns goal policy, check execution, integration, release preparation. Historic benchmark failures remain in reports.
|
|
78
|
+
- 2026-09-30: Implementation and review complete. Final prepublish pipeline passed (1,418 tests, 2 skipped, 155 files), lint/build/Canon/metadata/package checks and zero-vulnerability audit passed. Authenticated parallel and native accounting fixtures succeeded with their stated limits. Local 1.20.0 archive and changelog-derived notes prepared; publication remains out of scope.
|
package/gemini-extension.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.20.0",
|
|
4
4
|
"description": "Cross-agent coding harness for eight supported CLIs: curated skill canon, mechanical safety gates, autonomous loop with proof artifacts. CLI: npm i -g @hecer/yoke",
|
|
5
5
|
"contextFileName": "GEMINI-EXTENSION.md"
|
|
6
6
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@hecer/yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.20.0",
|
|
4
4
|
"description": "One harness, eight agents, zero trust in \"done\" — cross-agent coding harness for Claude Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi and Hermes: one skill canon, mechanical safety gates, an autonomous loop with screenshot/video proofs.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -23,6 +23,10 @@
|
|
|
23
23
|
"bench/run-matrix.mjs",
|
|
24
24
|
"bench/run-parallel-matrix.mjs",
|
|
25
25
|
"bench/analyze-routing-study.mjs",
|
|
26
|
+
"bench/compare-codex.mjs",
|
|
27
|
+
"bench/analyze-codex-comparison.mjs",
|
|
28
|
+
"bench/probe-goal-integration.mjs",
|
|
29
|
+
"bench/probe-parallel-gate-context.mjs",
|
|
26
30
|
"bench/fixtures",
|
|
27
31
|
"bench/results",
|
|
28
32
|
"docs",
|