@hecer/yoke 1.19.0 → 1.21.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +46 -0
- package/README.md +4 -1
- package/bench/RESULTS.md +30 -0
- package/bench/analyze-codex-comparison.mjs +22 -0
- package/bench/compare-codex.mjs +71 -0
- package/bench/fixtures/independent-utils/.yoke/prd.yaml +30 -0
- package/bench/fixtures/independent-utils/bench-verify.mjs +4 -0
- package/bench/fixtures/independent-utils/package.json +1 -0
- package/bench/fixtures/independent-utils/src/chunk.mjs +1 -0
- package/bench/fixtures/independent-utils/src/range.mjs +1 -0
- package/bench/fixtures/independent-utils/src/unique.mjs +1 -0
- package/bench/fixtures/independent-utils/tests/STORY-1.test.mjs +8 -0
- package/bench/fixtures/independent-utils/tests/STORY-2.test.mjs +8 -0
- package/bench/fixtures/independent-utils/tests/STORY-3.test.mjs +8 -0
- package/bench/probe-goal-integration.mjs +93 -0
- package/bench/probe-parallel-gate-context.mjs +29 -0
- package/bench/results/codex-comparison-2026-09-29.json +305 -0
- package/bench/results/goal-integration-probes-2026-09-29.json +107 -0
- package/bench/results/native-goal-validation-2026-09-30.json +122 -0
- package/bench/results/parallel-validation-2026-09-30.json +86 -0
- package/dist/agents/process-record-identity.js +33 -0
- package/dist/check/command.js +218 -0
- package/dist/cli.js +85 -12
- package/dist/dashboard/server.js +26 -6
- package/dist/goals/codex-native.js +340 -0
- package/dist/goals/command.js +230 -29
- package/dist/loop/cleanup.js +110 -1
- package/dist/loop/dispatcher.js +82 -10
- package/dist/loop/git.js +1 -1
- package/dist/loop/parallel-adapters.js +53 -7
- package/dist/loop/parallel-command.js +8 -1
- package/dist/loop/recovery.js +81 -2
- package/dist/loop/reporter.js +3 -0
- package/dist/loop/resource-pool.js +21 -15
- package/dist/loop/run-command.js +98 -27
- package/dist/loop/run-state.js +116 -0
- package/dist/prd/explore.js +12 -4
- package/dist/retrofit/command.js +22 -0
- package/dist/retrofit/config.js +1 -0
- package/dist/retrofit/gitignore.js +3 -0
- package/dist/retrofit/planners/claude.js +4 -4
- package/dist/retrofit/wsl.js +23 -3
- package/dist/setup/command.js +47 -5
- package/docs/CODEX-COMPARISON-2026-09-29.md +27 -0
- package/docs/CONTINUOUS-EXPLORATION.md +6 -2
- package/docs/GOALS-RESOURCE-AUDIT-2026-09-29.md +92 -0
- package/docs/GOALS.md +109 -0
- package/docs/RELEASE-VALIDATION-1.20.0.md +100 -0
- package/docs/VERIFIED-PROJECTS.md +4 -4
- package/docs/parallel-execution.md +15 -0
- package/docs/superpowers/plans/2026-09-30-goals-resources-release.md +78 -0
- package/gemini-extension.json +1 -1
- package/package.json +5 -1
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
# Direct Codex comparison — 2026-09-29
|
|
2
|
+
|
|
3
|
+
AI-assisted authenticated evaluation of Yoke **1.19.0**, commit `05fac3db1a51412563005e07f80e6e5a9aab8e06`. A fresh origin fetch confirmed that revision remains the remote main head.
|
|
4
|
+
|
|
5
|
+
All arms requested **gpt-6.1-sol / low**, disabled routing and native multi-agent delegation, and used disposable Git fixtures with the same source and acceptance tests within each comparison. User config was disabled; removal of every host skill/plugin was not certified. The direct Codex arm received all story requirements in one call. Yoke used its normal per-story implementation, verification and commit workflow. This compares complete user workflows, not just model-call latency. Permissions allowed changes in the disposable fixture projects. The protected seed tests were replayed after execution; they were visible to agents, not hidden tests.
|
|
6
|
+
|
|
7
|
+
| Fixture | Arm | Accepted runs | Median elapsed seconds | Fresh input tokens | Output tokens |
|
|
8
|
+
| --- | --- | --- | --- | --- | --- |
|
|
9
|
+
| routing-queue | codex | 2/2 | 118.3 | 25,292 | 2,162 |
|
|
10
|
+
| routing-queue | yoke-serial | 2/2 | 274.9 | 48,524 | 4,822 |
|
|
11
|
+
| independent-utils | codex | 1/1 | 71.3 | 10,263 | 797 |
|
|
12
|
+
| independent-utils | yoke-serial | 1/1 | 442.2 | 92,330 | 3,905 |
|
|
13
|
+
| independent-utils | yoke-parallel | 0/1 | 117.1 | 81,497 | 3,487 |
|
|
14
|
+
|
|
15
|
+
Queue: **two alternating pairs**, each evaluated against ten original tests. Both arms passed both times. Yoke's serial workflow took **2.32 times** the elapsed time, with **91.9% more fresh input tokens**. Input totals also include cached tokens; fresh input is aggregate input minus cached input. This is not a USD cost calculation.
|
|
16
|
+
|
|
17
|
+
Independent utilities: **one run per arm**, with three disjoint source scopes and fifteen original tests. Direct Codex and Yoke serial passed. Parallel Yoke started three real workers but accepted no stories after its integration checks lost the `YOKE_STORY` context. It stopped at the explicit three-attempt cap. **Its shorter failure time is not a performance improvement.** The deterministic control probe reproduced the context loss independently of model quality.
|
|
18
|
+
|
|
19
|
+
The initial independent fixture lacked runtime ignore rules; its serial run stopped after one story. Its parallel attempt also hit Git's Windows worktree path limit before model execution. The fixture was corrected and all three arms rerun under shorter paths. Those initial records are retained separately as setup diagnostics, not silently discarded or mixed into accepted performance statistics.
|
|
20
|
+
|
|
21
|
+
These are small fixtures on one Windows machine with uncontrolled background load. CPU, memory and USD cost were not measured. Routing, cheap-worker selection, native Codex subagents, large repositories and prolonged exploration were not compared. There is no evidence here that Yoke generally outperforms direct Codex. The observed serial overhead and parallel failure warrant fixing workflow/context correctness before promoting universal efficiency claims.
|
|
22
|
+
|
|
23
|
+
See the [Goals/resource audit](GOALS-RESOURCE-AUDIT-2026-09-29.md) for implementation priorities. Machine-readable records: [comparison results](../bench/results/codex-comparison-2026-09-29.json), [integration probes](../bench/results/goal-integration-probes-2026-09-29.json). Full raw logs remain in the local testground; they are not included in these summaries.
|
|
24
|
+
|
|
25
|
+
Reproduce with `bench/compare-codex.mjs`, then `bench/analyze-codex-comparison.mjs`. Use fresh, short output roots on Windows. The new summary tests verify median arithmetic, cache subtraction, unknown usage, incompatible policies and invalid measurements.
|
|
26
|
+
|
|
27
|
+
Provenance: this report is disclosed as AI-assisted. Read-only text scans cannot establish human authorship or verify proprietary keyed watermarks; cryptographic verification and signer trust remain unknown without the corresponding verifier and trust policy.
|
|
@@ -53,8 +53,12 @@ when quoted, for example `--explore-limit="2 days"`. When it expires, Yoke pause
|
|
|
53
53
|
code `3`. It stops launching new workers, lets already active workers finish their acceptance,
|
|
54
54
|
verification and integration gates, and preserves resumable work. Finite runs process bounded task
|
|
55
55
|
batches so an unfinished large backlog does not run past the deadline by draining the whole PRD.
|
|
56
|
-
Omitting `--explore-limit`
|
|
57
|
-
|
|
56
|
+
Omitting `--explore-limit` on a fresh run keeps the supervisor unbounded. Dashboard
|
|
57
|
+
resume restores the saved absolute deadline and consumed iteration budget, so pausing
|
|
58
|
+
does not grant another duration. An already expired saved run remains paused. Starting
|
|
59
|
+
an explicit new CLI invocation with a new duration creates a new run and deadline.
|
|
60
|
+
The supervisor retains the project lock during idle waits, preventing a second runner
|
|
61
|
+
from changing the same project between exploration batches.
|
|
58
62
|
|
|
59
63
|
The supervisor can retry failures while its process remains alive. Run it in a background process
|
|
60
64
|
for long sessions. An operating-system shutdown, forced process termination, exhausted credentials
|
|
@@ -0,0 +1,92 @@
|
|
|
1
|
+
# Goals and resource audit — 2026-09-29
|
|
2
|
+
|
|
3
|
+
AI-assisted inspection and authenticated tests of Yoke **1.19.0**, commit `05fac3db1a51412563005e07f80e6e5a9aab8e06`. Test tooling and this report were added in a separate worktree and branch. The original checkout and its local changes were preserved. Recommendations below are proposals, not shipped changes.
|
|
4
|
+
|
|
5
|
+
## What is already reliable
|
|
6
|
+
|
|
7
|
+
This report preserves the **1.19.0** findings. The implementation following this
|
|
8
|
+
audit is documented in [verified goals](GOALS.md) and the **1.20.0** changelog;
|
|
9
|
+
its release validation is recorded separately. Historical benchmark timings are
|
|
10
|
+
not measurements of the corrected version.
|
|
11
|
+
|
|
12
|
+
Yoke persists a provider-neutral objective, attempt history, independent check IDs, and interrupted-attempt evidence in `.yoke/goal.json`. Goals and story loops share an exclusive project lock. Goals require executable acceptance and protected test infrastructure. Each attempt is independently checked, and an earlier completion is checked again when resumed. Failed work is retained; a goal run does not automatically commit or publish it. These protections are useful even when a provider confidently claims completion.
|
|
13
|
+
|
|
14
|
+
The latest release's complete prepublication validation previously passed: TypeScript lint/build, **148 test files; 1,337 passed and 2 skipped**, docs checks and package dry run. In this audit, the final focused goal, shared worker-pool, dashboard and comparison-summary suites passed: **40 tests**, including 11 goal tests and five new summary tests. A passing suite does not establish coverage for the gaps below.
|
|
15
|
+
|
|
16
|
+
## Reproduced gaps
|
|
17
|
+
|
|
18
|
+
| Gap | Evidence | Consequence | Proposed fix |
|
|
19
|
+
| --- | --- | --- | --- |
|
|
20
|
+
| Objective is not bound to relevant acceptance | A disposable project had passing protected checks for an old feature. A new objective requested `new-feature.txt`, but its criteria still only checked the old feature. Yoke marked the new goal `complete` with **zero attempts**, without calling the fake executor or creating the requested artifact. | Checks prove the supplied contract, which can be unrelated to the newly requested work. Existing green tests alone do not prove a new objective. | Persist an explicit goal-to-criterion mapping and contract revision. Require criteria that cover the new objective, generated or selected before autonomous execution and then protected. Keep regression-suite success separate from objective-specific acceptance; do not infer relevance from a free-text objective. |
|
|
21
|
+
| Goals bypass the shared resource pool | With an isolated pool limited to one worker and its permit occupied, a goal's fake executor still started and completed. Default execution and capability-planner calls also directly start providers without requesting a permit. | Concurrent projects can exceed the intended Yoke worker limit; goal routing can introduce additional unaccounted calls. | Route all goal implementation and planner calls through the shared admission service. Disable unmanaged native delegation or account for its children explicitly. Release permits in `finally`, including abort, provider failure and routing escalation. |
|
|
22
|
+
| No native Codex goal binding | The persisted goal schema contains no native thread reference. `goalHandoff` supplies text, and execution launches a fresh CLI invocation. There are no native goal protocol calls. | Native and Yoke goal statuses, pause controls, budgets and continuation can disagree. Context is restarted between attempts. | Add a capability-negotiated Codex adapter with persisted thread binding, native goal set/get/update notifications, pause synchronization and session recovery. Keep provider-neutral fallback for CLIs without this capability. |
|
|
23
|
+
| Token budget is checked between attempts | A disposable goal with a budget of **1 token**, a fake executor reporting **100 tokens**, and passing acceptance finished as `complete`. This is a mechanical control-flow test, not an actual 100-token model call. | A single attempt can exceed the budget. The option is an admission guard for subsequent attempts, not a strict live token cap. | Pass supported live limits into the native runtime, consume streamed usage, reserve a bounded allowance for controller and verification calls, and report overshoot explicitly. Do not equate native budget accounting with provider aggregate tokens until their units are verified. |
|
|
24
|
+
| Time budget excludes verification | With a **60 ms agent-time budget**, intentionally delayed protected checks and a fake fast executor, the final recorded probe completed after **652 ms** while recording **3 ms** agent work. | Users cannot use `--minutes` as a total elapsed-time or CPU-resource limit. The current handoff explicitly says total agent work; this is a limitation rather than proof the documented budget is incorrect. | Preserve the existing meaning for compatibility and add an explicit elapsed-time deadline. Record admission wait, model work, verification and integration separately. Allow safe-boundary overrun only when documented. |
|
|
25
|
+
| Goal CLI does not inherit the configured runner selection | `src/cli.ts` defaults the goal provider to Codex and passes only `--model`; `runProjectGoal` loads config for routing but does not merge `runner.reasoningEffort`, `bare`, or native delegation policy into the direct selection. | Goal execution can use different startup context, effort and resource policy than the same project's story loop. | Resolve one common execution policy: explicit CLI overrides, then project runner config, then provider defaults. Test parity across goal, loop and dashboard resume. |
|
|
26
|
+
|
|
27
|
+
Source references: `src/goals/command.ts:21,69,83,115,128,134,136,147,152`; `src/cli.ts:202`; `src/loop/resource-pool.ts`; `src/agents/providers.ts`.
|
|
28
|
+
|
|
29
|
+
## Native capability verified on this machine
|
|
30
|
+
|
|
31
|
+
Installed Codex: **0.161.0-alpha.2**, authenticated with ChatGPT; its `goals` feature is enabled. The local CLI emitted experimental app-server schemas for `thread/goal/set`, `thread/goal/get`, and `thread/goal/clear`. In a disposable thread the actual app-server successfully performed:
|
|
32
|
+
|
|
33
|
+
1. Set an active objective with a 1,000-token budget.
|
|
34
|
+
2. Read the same objective and budget.
|
|
35
|
+
3. Set status `paused` and read it back.
|
|
36
|
+
4. Clear the goal and verify `goal: null`.
|
|
37
|
+
|
|
38
|
+
The thread was archived. This protocol probe did **not** request a model development turn; it verifies the local lifecycle API, not autonomous completion quality. Protocol availability must be negotiated for other installed versions. Do not enable global features or silently create a native goal in the user's current conversation.
|
|
39
|
+
|
|
40
|
+
OpenAI describes native goals as persistent, thread-scoped objectives with continuation subject to completion, interruption, budgets and blockers. This is not a guarantee of endless productive development. See the [official Goals cookbook](https://developers.openai.com/cookbook/examples/codex/using_goals_in_codex).
|
|
41
|
+
|
|
42
|
+
## Unify continuation without competing controllers
|
|
43
|
+
|
|
44
|
+
Keep Yoke authoritative for executable project acceptance, evidence and safe integration. Let the native goal manage work inside a provider thread. When the provider reports completion, run Yoke's independent check against the current tree. Only synchronize successful completion when that check passes. When it fails, supply bounded findings and continue the same objective within the remaining budget. Persist the native completion claim and the failed evidence separately rather than discarding either.
|
|
45
|
+
|
|
46
|
+
Store a run identity, native thread reference, objective revision, acceptance digest, provider/model selection, budget units and stop reason. Recovery must reconcile native and Yoke state before resuming. An explicit user pause must remain paused across crashes and restarts. Temporary rate limits, exhausted budgets, unresolved decisions and infrastructure errors need distinguishable reasons and retry rules; they should not all become an indistinguishable `blocked` status.
|
|
47
|
+
|
|
48
|
+
Checks currently use synchronous command execution with a default timeout of up to ten minutes **per command** (`src/loop/verify.ts:42`). The goal's agent-time timer has already been cleared before those checks. A total elapsed-time limit therefore needs an asynchronous, abortable verification path or an external supervisor; adding only another JavaScript timer around synchronous checks would not enforce it. Synchronous checks also block the dashboard process's event loop when a goal is resumed inside that process.
|
|
49
|
+
|
|
50
|
+
Continuous exploration is already an optional story-loop supervisor; it does not consume or synchronize `.yoke/goal.json`. Keep discovery separate from a finite goal contract. Each accepted exploration batch should have a bounded objective and verified acceptance, then return to discovery if the user still wants exploration. A waiting or blocked provider must never cause two independent supervisors to issue overlapping work.
|
|
51
|
+
|
|
52
|
+
Dashboard resume currently chooses a goal when an unfinished `.yoke/goal.json` exists; otherwise it starts `runLoopCommand(root, {})`. It does not persist and restore the original exploration flag, exploration deadline or full run selection. This is an observed source-level integration gap, not a crash-recovery test. Persist the actual execution mode and options, select the interrupted run by identity, and explicitly choose whether a resumed exploration limit preserves the original deadline. Test pause/resume both before and during model calls and integration.
|
|
53
|
+
|
|
54
|
+
## Resource efficiency priorities
|
|
55
|
+
|
|
56
|
+
**An actual parallel integration failure was reproduced.** The corrected independent-task fixture ran three authenticated Codex workers, but integrated **0/3** stories and ended at the explicit three-attempt cap. The worker gate sets `YOKE_STORY`; the dispatcher's integration gate directly invokes the verifier without setting it (`src/loop/worker.ts:60`, `src/loop/parallel-command.ts:113`, `src/loop/dispatcher.ts:133`). The fixture's scoped verifier therefore ran all three tests against each partial candidate at integration, rejecting it because the other independent modules were still unimplemented.
|
|
57
|
+
|
|
58
|
+
A deterministic diagnostic exercised the real loop/dispatcher with fake agents and Git adapters: worker observations were `A,B,C`, integration observations were `null,null,null`. The always-green control accepted **3/3**; the story-context-dependent verifier accepted **0/3**. Preserve identical verification context across worker and integration phases, preferably with explicit subprocess environment rather than ambient global mutation. Record rejection reasons and retain recoverable candidate evidence: this run's final status only reported `cap-reached`, while worker directories were removed. Lost context and repeated implementation after rejected integration consume resources without accepted work. The failed authenticated parallel run remains in the result set and is not reported as a speedup.
|
|
59
|
+
|
|
60
|
+
The first independent-task comparison exposed a Windows path limit before any parallel provider call: `git worktree add` failed with `fatal: '$GIT_DIR' too big`. Yoke's generated candidate path included a 64-character story digest and a UUID below the project path. The fixture was rerun under a shorter root. A production improvement is to preflight the complete Git worktree path, use shorter collision-safe candidate names or an explicit short worktree root, and report recoverable setup failures without a raw stack trace. Separately, the new fixture initially lacked runtime-file ignore rules and its serial run stopped on a dirty worktree after one story; this was corrected in the fixture. Those initial failed attempts are retained as setup diagnostics, not counted as successful performance samples.
|
|
61
|
+
|
|
62
|
+
The existing pool limits worker units, not measured CPU, memory or test subprocesses. Model request concurrency and local compilation/testing are different constraints. A useful next step is separate admission for model calls and expensive local checks, with per-project fairness and a shared cross-project ceiling. Include goal planners, repair attempts, exploration and native children in the accounting. Do not hold an implementation permit while merely waiting for integration if the integration lane can acquire its own permit safely.
|
|
63
|
+
|
|
64
|
+
For small coherent tasks, repeated context loading, separate sessions, worktree creation and repeated gates can cost more than they save. Introduce an explicit lightweight path or batching of tightly related criteria where the final independent acceptance remains protected. Parallelism is useful when write scopes are independent and model work dominates; it does not remove serial integration and can increase total tokens. Preserve mandatory final acceptance, but reuse test results only when the relevant code, toolchain, environment and test contract digests match. Missing telemetry is unknown, not zero cost.
|
|
65
|
+
|
|
66
|
+
## Concrete implementation order and acceptance
|
|
67
|
+
|
|
68
|
+
1. **Correct verification context, objective acceptance and shared execution policy.** Assert worker and integration gates receive the same story context and that rejection reasons survive in status/events. Test a new objective against an old green contract; it must not be accepted without an explicit relevant criterion binding. Preserve legitimate already-satisfied objective handling. Saturate a one-worker pool; assert goals and their planner wait. Exercise abort while queued, failure after acquisition and escalation. Assert no leaked permit and consistent selection in CLI/dashboard paths.
|
|
69
|
+
2. **Optional Codex native-goal adapter.** Test supported and unsupported schemas, persistent thread binding, independent-completion reconciliation, provider switches, objective revision mismatch and stale native state. Never require native Codex goals for Hermes or other providers.
|
|
70
|
+
3. **Budgets, pauses and recovery.** Test live token accounting, unknown usage, single-attempt overshoot reporting, elapsed-time limits during checks, user pause while active, restart after interruption and resumption without resetting measured consumption.
|
|
71
|
+
4. **Exploration and dashboard identity.** Persist mode/options and connect each bounded batch to its own verified goal. Test that a requested pause survives restart and dashboard resume preserves the chosen mode and documented deadline policy.
|
|
72
|
+
5. **Measured efficiency.** Run repeated real comparisons at equal accepted quality, model and effort. Include dependent tasks, independent tasks and a larger build/test workload. Report elapsed time, fresh/cached tokens, acceptance, retries and actual process-tree CPU/memory when measured. Decide lightweight/default policies from that evidence.
|
|
73
|
+
|
|
74
|
+
## Reproduce
|
|
75
|
+
|
|
76
|
+
Run from the validation branch after building Yoke:
|
|
77
|
+
|
|
78
|
+
```powershell
|
|
79
|
+
rtk npx vitest run tests/goals/command.test.ts tests/loop/resource-pool.test.ts
|
|
80
|
+
rtk proxy node bench/probe-goal-integration.mjs G:/NN-Developed/Yoke-Testground/goal-probe-NEW
|
|
81
|
+
rtk proxy node bench/probe-parallel-gate-context.mjs G:/NN-Developed/Yoke-Testground/parallel-probe-NEW
|
|
82
|
+
rtk proxy node bench/compare-codex.mjs --fixture=routing-queue --repeats=2 --root=G:/NN-Developed/Yoke-Testground/comparison-NEW
|
|
83
|
+
rtk proxy node bench/compare-codex.mjs --fixture=independent-utils --repeats=1 --root=G:/NN-Developed/Yoke-Testground/independent-NEW
|
|
84
|
+
```
|
|
85
|
+
|
|
86
|
+
Use a fresh output directory. The comparisons make authenticated model calls and permit changes inside disposable fixture projects. Routing is disabled and native multi-agent delegation is disabled in both arms. The Yoke parallel arm uses three Yoke workers. The original fixture tests are replayed after execution, so agent-modified tests do not define success. They remain visible to the agents; this is not hidden-test evaluation. User config is disabled, but the harness does not certify removal of every installed plugin or host-discovered skill. CPU and memory are not measured. Full logs remain in the local testground; summarized results accompany this report.
|
|
87
|
+
|
|
88
|
+
These runs share one Windows machine with other activity; background resource load was not controlled or measured. Small-sample elapsed times should not be generalized to all projects or treated as isolated CPU benchmarks.
|
|
89
|
+
|
|
90
|
+
The [authenticated comparison report](CODEX-COMPARISON-2026-09-29.md) contains timings, tokens, the failed parallel outcome and explicit sample limits. [Machine-readable probes](../bench/results/goal-integration-probes-2026-09-29.json) retain the mechanical findings and protocol results without private thread identifiers.
|
|
91
|
+
|
|
92
|
+
Provenance: AI assistance is disclosed above. Read-only scans found no supported provenance carrier or observable text-watermark signal and inspected the full text. Cryptographic verification, signer trust, and metadata privacy remain unknown for this text without supported verifier/trust inputs. Proprietary keyed watermark detectors were unavailable; absence of a signal is not evidence of human authorship. No authorship classification is made.
|
package/docs/GOALS.md
ADDED
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
# Verified project goals
|
|
2
|
+
|
|
3
|
+
A goal is a durable objective with explicit executable acceptance. Yoke owns the
|
|
4
|
+
completion decision; a model's claim or an unrelated green suite cannot complete it.
|
|
5
|
+
|
|
6
|
+
## Define and bind acceptance
|
|
7
|
+
|
|
8
|
+
Create `.yoke/acceptance.yaml` and protect the tests it runs:
|
|
9
|
+
|
|
10
|
+
```yaml
|
|
11
|
+
version: 1
|
|
12
|
+
protected: [tests/checkout.test.mjs]
|
|
13
|
+
criteria:
|
|
14
|
+
- id: checkout
|
|
15
|
+
text: Checkout saves an order and rejects invalid input
|
|
16
|
+
commands: [node --test tests/checkout.test.mjs]
|
|
17
|
+
```
|
|
18
|
+
|
|
19
|
+
```sh
|
|
20
|
+
yoke goal set . --objective="Implement checkout" --criteria=checkout
|
|
21
|
+
yoke goal run . --runner=codex --model=gpt-6.1-sol --effort=low
|
|
22
|
+
yoke goal status .
|
|
23
|
+
yoke goal pause .
|
|
24
|
+
yoke goal resume .
|
|
25
|
+
```
|
|
26
|
+
|
|
27
|
+
`--criteria` asserts that the named checks actually prove this objective. Yoke validates
|
|
28
|
+
IDs and executable commands and binds the objective to a digest of the acceptance
|
|
29
|
+
manifest. It cannot determine whether your tests fully express a free-text request.
|
|
30
|
+
All project acceptance and configured regression checks must still pass.
|
|
31
|
+
|
|
32
|
+
Existing goals without a binding now stop before any model call. Review their tests,
|
|
33
|
+
then explicitly bind them with `yoke goal bind . --criteria=checkout`. Attempts and
|
|
34
|
+
measured consumption remain intact. Changed acceptance infrastructure still requires
|
|
35
|
+
an explicit protection refresh after review; binding never refreshes protected tests.
|
|
36
|
+
|
|
37
|
+
## Resource and budget controls
|
|
38
|
+
|
|
39
|
+
Goals and story loops share the project lock and the user-level worker pool. Routing
|
|
40
|
+
planner and implementation calls each acquire a permit; native model delegation is
|
|
41
|
+
disabled. Explicit runner/model/effort/bare options override project runner defaults.
|
|
42
|
+
Execution uses Yoke's safe permission profile.
|
|
43
|
+
|
|
44
|
+
| Setting | Meaning |
|
|
45
|
+
| --- | --- |
|
|
46
|
+
| `--attempts=N` on `set` or `budget` | Maximum cumulative attempts, default 3 |
|
|
47
|
+
| `--minutes=N` | Cumulative admitted provider work, default 30 minutes; excludes capacity waits and independent checks |
|
|
48
|
+
| `--wall-minutes=N` | Optional cumulative run time including admission, checks and state synchronization |
|
|
49
|
+
| `--tokens=N` | Cumulative measured input + output tokens; unknown consumption stops further budgeted work |
|
|
50
|
+
| `YOKE_MAX_PARALLEL_WORKERS=1..8` | Shared model/integration permit ceiling, default 3 |
|
|
51
|
+
| `YOKE_MAX_PARALLEL_CHECKS=1..8` | Separate shared asynchronous goal verification ceiling, default 1 |
|
|
52
|
+
|
|
53
|
+
After a measured token overrun, Yoke retains evidence and work and reports `blocked`
|
|
54
|
+
even if acceptance passed. Update the budget explicitly with `yoke goal budget .
|
|
55
|
+
--tokens=50000`; `--clear-token-budget` explicitly removes that ceiling. Providers
|
|
56
|
+
reporting usage only after a call cannot provide a hard mid-call token cap. Native
|
|
57
|
+
Codex usage units and Yoke's token accounting are recorded separately. Native
|
|
58
|
+
stream updates use cumulative turn deltas and request cancellation when measured
|
|
59
|
+
consumption exceeds the goal ceiling. Notifications arrive after model requests,
|
|
60
|
+
so this can still overshoot within a request; progress updates are never added
|
|
61
|
+
twice to persisted final usage.
|
|
62
|
+
|
|
63
|
+
The asynchronous goal checker can interrupt its own command tree at a deadline or
|
|
64
|
+
pause. Synchronous `yoke check` and story gates retain their existing execution
|
|
65
|
+
contracts; this check pool is not an operating-system CPU, RAM, or per-test scheduler.
|
|
66
|
+
After an interrupted process, unknown consumption is charged conservatively rather
|
|
67
|
+
than silently granting a fresh budget.
|
|
68
|
+
|
|
69
|
+
## Optional native Codex goals
|
|
70
|
+
|
|
71
|
+
```sh
|
|
72
|
+
yoke goal run . --native-goal
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
Or set `goals.nativeCodex: true` in `.yoke/config.yaml`. The default is off;
|
|
76
|
+
`--no-native-goal` overrides the project setting. Other providers remain supported.
|
|
77
|
+
|
|
78
|
+
Yoke negotiates the local Codex app-server goal capability, stores the thread binding,
|
|
79
|
+
pauses native automatic continuation, and explicitly starts one bounded development
|
|
80
|
+
turn. The same thread resumes on later attempts. Only Yoke's independent acceptance
|
|
81
|
+
can synchronize it to `complete`. Provider/model changes create a new binding.
|
|
82
|
+
An unsupported native goal method falls back to ordinary Codex execution; authentication,
|
|
83
|
+
transport and execution failures remain visible. No global Codex configuration changes
|
|
84
|
+
are needed. Support depends on the installed CLI, not on the model name alone.
|
|
85
|
+
|
|
86
|
+
Codex app-server does not currently support exec's bare startup option. Combining
|
|
87
|
+
native goals with `runner.bare: true` or `--bare` stops before dispatch rather than
|
|
88
|
+
silently loading user configuration. Use `--no-native-goal` to retain bare execution,
|
|
89
|
+
or explicitly disable bare startup when opting into native goals.
|
|
90
|
+
|
|
91
|
+
## Dashboard resume
|
|
92
|
+
|
|
93
|
+
The dashboard restores the recorded mode, goal identity and safe execution settings
|
|
94
|
+
from `.yoke/run-state.json`. It rejects a mismatched goal identity. A saved story or
|
|
95
|
+
exploration run takes precedence over an unrelated legacy unfinished goal. See
|
|
96
|
+
[continuous exploration](CONTINUOUS-EXPLORATION.md) for absolute deadline behavior.
|
|
97
|
+
|
|
98
|
+
Resume restores validated safe options, not every original authorization override.
|
|
99
|
+
It uses safe permissions and defaults for self-review overrides and unbounded quality
|
|
100
|
+
repair. Reviewer, quality and routing policy still follow the current project config;
|
|
101
|
+
implementation and exploration-planner provider/model selections are saved separately.
|
|
102
|
+
|
|
103
|
+
If process-tree termination cannot be confirmed, Yoke blocks subsequent execution and
|
|
104
|
+
retains its process record and concurrency permit. A late confirmed termination can
|
|
105
|
+
release the retained permit; otherwise inspect the recorded ownership before restarting
|
|
106
|
+
the holder. Uncertainty is not treated as successful cleanup.
|
|
107
|
+
|
|
108
|
+
Documentation and regression tooling were prepared with AI assistance; test evidence
|
|
109
|
+
and release validation are recorded separately from product guarantees.
|
|
@@ -0,0 +1,100 @@
|
|
|
1
|
+
# Yoke 1.20.0 release preparation — 2026-09-30
|
|
2
|
+
|
|
3
|
+
AI-assisted implementation and validation in an isolated worktree based on
|
|
4
|
+
`05fac3db1a51412563005e07f80e6e5a9aab8e06`. The original checkout and its local
|
|
5
|
+
changes were preserved. This report records local release preparation, not a
|
|
6
|
+
published GitHub or npm release.
|
|
7
|
+
|
|
8
|
+
## Implemented and checked
|
|
9
|
+
|
|
10
|
+
- Goal-to-criterion binding, contract digests and migration for unbound goals.
|
|
11
|
+
- Shared model admission, separate asynchronous check admission, token overrun
|
|
12
|
+
reporting, cumulative provider/wall budgets, cancellation and retained ownership.
|
|
13
|
+
- Optional Codex native goal threads with paused automatic continuation, independent
|
|
14
|
+
completion, durable usage baselines and cumulative turn accounting.
|
|
15
|
+
- Story context through integration gates, retained parallel candidates, validated
|
|
16
|
+
recovery, bounded metadata and shorter Windows worktree paths.
|
|
17
|
+
- Saved run mode, identity, implementation/planner selections, absolute exploration
|
|
18
|
+
deadline and durable attempt charging before scheduling work.
|
|
19
|
+
- Runtime claims and nested worker checkouts excluded from cleanliness and staging.
|
|
20
|
+
|
|
21
|
+
Regression tests reproduced the original failures before correction. Review also
|
|
22
|
+
covered telemetry failures during retention, reporter failures after admission,
|
|
23
|
+
interrupted counters, unsafe resume defaults and process-tree cleanup uncertainty.
|
|
24
|
+
|
|
25
|
+
## Actual execution evidence
|
|
26
|
+
|
|
27
|
+
The corrected authenticated parallel smoke used `gpt-6.1-sol`, low effort, three
|
|
28
|
+
workers and routing disabled. It accepted **3/3 stories**, exited `0`, and passed
|
|
29
|
+
**15/15 original tests** after replaying immutable verification files. All four
|
|
30
|
+
protected verification files matched the seed. Duration: **112,949 ms**. Recorded
|
|
31
|
+
input: **456,659**, including **372,096 cached**; output: **3,217**. The initial
|
|
32
|
+
failed sample is preserved rather than excluded from the report.
|
|
33
|
+
|
|
34
|
+
See [parallel validation data](../bench/results/parallel-validation-2026-09-30.json).
|
|
35
|
+
This is a single corrected arm, not a fresh direct-Codex comparison. CPU, RAM and
|
|
36
|
+
USD costs were not measured, and host load was not controlled. It establishes this
|
|
37
|
+
fixture's acceptance, not general efficiency superiority.
|
|
38
|
+
|
|
39
|
+
The integrated native goal fixture completed its protected acceptance in one model
|
|
40
|
+
attempt and resumed the same goal/thread without another implementation attempt.
|
|
41
|
+
The final accounting trial measured **153,736 input** and **606 output** tokens,
|
|
42
|
+
with **134,528 cached input**. Its persisted baseline matched the cumulative
|
|
43
|
+
provider total; native paused-goal accounting separately reported zero. Model and
|
|
44
|
+
check permits and process records were empty after completion.
|
|
45
|
+
|
|
46
|
+
See [native goal validation data](../bench/results/native-goal-validation-2026-09-30.json).
|
|
47
|
+
|
|
48
|
+
Native accounting uses cumulative thread snapshots and a durable baseline. Missing,
|
|
49
|
+
malformed or decreasing counters remain unknown; a decreasing counter marks the
|
|
50
|
+
binding untrusted for subsequent accounting. No token usage is fabricated from the
|
|
51
|
+
absence of telemetry. Paused native goals do not provide a hard token cap for a
|
|
52
|
+
manual turn; Yoke's independent accounting guards further work and reports overruns.
|
|
53
|
+
Native cumulative progress notifications also request cancellation when the ceiling
|
|
54
|
+
is exceeded. Focused parent/adapter regressions cover streaming, buffered replies
|
|
55
|
+
and foreign-turn rejection; the authenticated sample above used a generous budget
|
|
56
|
+
and establishes completion/accounting, not a live small-budget cancellation trial.
|
|
57
|
+
|
|
58
|
+
## Release checks
|
|
59
|
+
|
|
60
|
+
Environment: Windows, Node.js **24.13.0**. Native protocol and model execution used
|
|
61
|
+
installed Codex **0.161.0-alpha.2**. TypeScript lint and build passed. Dependency
|
|
62
|
+
audit passed with **0 vulnerabilities**. The complete test run passed **1,418 tests**
|
|
63
|
+
across **155 files**, with **2 skipped** (1,420 defined), in **342.26 seconds**.
|
|
64
|
+
Canon, README metadata and package dry-run checks passed. The complete
|
|
65
|
+
`prepublishOnly` pipeline exited **0**. The package contains **362 files**;
|
|
66
|
+
temporary runtime state and raw model transcripts are excluded.
|
|
67
|
+
|
|
68
|
+
The local build archive `hecer-yoke-1.20.0.tgz` was created and inspected directly.
|
|
69
|
+
All four packaged version manifests match; the native adapter and durable run-state
|
|
70
|
+
modules are included. Matching changelog-derived `RELEASE-NOTES.md`, `SHA256SUMS`
|
|
71
|
+
and `PACKAGE-VALIDATION.json` accompany the archive in the local release output.
|
|
72
|
+
|
|
73
|
+
Versions are synchronized to **1.20.0** in package/lockfile, Claude/Codex manifests,
|
|
74
|
+
Gemini extension and README. The dated changelog entry is the source of release
|
|
75
|
+
notes. No remote tag, GitHub release, CI run or npm publication is claimed here;
|
|
76
|
+
the repository's GitHub matrix and npm verification remain publication gates.
|
|
77
|
+
|
|
78
|
+
## Documentation provenance
|
|
79
|
+
|
|
80
|
+
Read-only provenance inspection fully scanned README, CHANGELOG, this report, the
|
|
81
|
+
goal guide and sanitized validation reports. No supported provenance carrier or
|
|
82
|
+
observable text marker was found.
|
|
83
|
+
Cryptographic verification, signer trust and text metadata privacy remain unknown:
|
|
84
|
+
no conforming verifier/trust policy was supplied, and keyed provider watermarks
|
|
85
|
+
are unavailable. Absence of a marker is not evidence of human authorship. Raw model
|
|
86
|
+
transcripts and native thread identifiers are kept outside the published reports.
|
|
87
|
+
|
|
88
|
+
## Operational limits
|
|
89
|
+
|
|
90
|
+
- Native goals are opt-in. App-server lacks exec's bare startup control, so native
|
|
91
|
+
plus bare stops before dispatch. Ordinary Codex execution remains available.
|
|
92
|
+
- Worker/check permits limit concurrency, not CPU/RAM consumption inside a process.
|
|
93
|
+
- Synchronous check/story execution retains its existing contracts.
|
|
94
|
+
- Process termination uncertainty blocks follow-up and retains ownership; a late
|
|
95
|
+
confirmed tree stop can release the retained permit.
|
|
96
|
+
- Dashboard resume restores validated safe options. Reviewer/quality/routing policy
|
|
97
|
+
still follows current config; self-review/unbounded-quality overrides use defaults.
|
|
98
|
+
- Retained parallel candidates require validated repository, contract and ownership
|
|
99
|
+
state. Stale target history requires reconciliation; serial recovery does not
|
|
100
|
+
automatically adopt a parallel checkout.
|
|
@@ -33,7 +33,7 @@ After intentionally editing protected tests, use `yoke check . --protect --refre
|
|
|
33
33
|
## Durable goals across providers
|
|
34
34
|
|
|
35
35
|
```sh
|
|
36
|
-
yoke goal set . --objective="Finish guest checkout" --attempts=3 --minutes=30
|
|
36
|
+
yoke goal set . --objective="Finish guest checkout" --criteria=guest-checkout --attempts=3 --minutes=30
|
|
37
37
|
yoke goal run . --runner=codex
|
|
38
38
|
yoke goal resume . --runner=claude --model=<installed-model-id>
|
|
39
39
|
yoke goal resume . --runner=gemini --model=<installed-model-id>
|
|
@@ -42,9 +42,9 @@ yoke goal handoff .
|
|
|
42
42
|
yoke goal pause .
|
|
43
43
|
```
|
|
44
44
|
|
|
45
|
-
Goals keep objective, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock
|
|
45
|
+
Goals keep objective, explicit acceptance binding, attempts, failure context and check IDs in `.yoke/goal.json`. They share the story-loop lock and worker capacity. A completed goal is checked again on a later run. Failed work stays in the project; goal execution itself does not commit or publish it. Native Codex goals are optional through `--native-goal`; see the [goal guide](GOALS.md) for capability negotiation and migration of unbound goals.
|
|
46
46
|
|
|
47
|
-
`--minutes` limits cumulative
|
|
47
|
+
`--minutes` limits cumulative admitted provider execution time. Optional `--wall-minutes` also includes capacity waits and independent checks. `--tokens=N` blocks further execution when measured usage is exhausted or unknown, and reports post-call overruns even when checks pass; providers with end-of-call telemetry cannot promise a hard mid-call cap. Interrupted attempts are charged conservatively and recorded process ownership is reconciled before continuation. A pause requests cancellation of Yoke-owned asynchronous execution and retains unfinished work and evidence.
|
|
48
48
|
|
|
49
49
|
Explicitly extend total budgets without deleting history:
|
|
50
50
|
|
|
@@ -65,7 +65,7 @@ Failed or paused isolated story worktrees are retained. Resume the same story wi
|
|
|
65
65
|
yoke loop run . --isolate --resume-worktree --parallel=1 --candidates=1
|
|
66
66
|
```
|
|
67
67
|
|
|
68
|
-
Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery
|
|
68
|
+
Yoke validates the registered worktree, repository, original target commit and PRD digest. Changed target/PRD state is refused; inspect retained edits and reconcile deliberately. This flag is for serial isolated recovery. Parallel retained candidates resume through the dispatcher without reimplementation when their recovery binding remains valid. `yoke loop cleanup` is an explicit discard operation. Existing projects should rerun retrofit to add runtime ignore entries for checks, events, goals, saved runs and recovery evidence.
|
|
69
69
|
|
|
70
70
|
## Spend fewer model calls
|
|
71
71
|
|
|
@@ -80,6 +80,21 @@ one serial integration lane when integration measurements exist. Older runs with
|
|
|
80
80
|
remain visible as missing integration history; forecasts are empirical ranges, not deadlines, and do
|
|
81
81
|
not predict future contention from other projects.
|
|
82
82
|
|
|
83
|
+
### Rejected integration recovery
|
|
84
|
+
|
|
85
|
+
Worker and integrated-tree gates receive the same `YOKE_STORY` context. A rejected
|
|
86
|
+
candidate is retained with its reason and an `integration-recovery.json` proof record;
|
|
87
|
+
it is not silently discarded and regenerated. The next parallel run can reuse it
|
|
88
|
+
without a new implementation model call when canonical project/worktree ownership,
|
|
89
|
+
Git registration, target base and PRD digest still match. Integration gates run again.
|
|
90
|
+
Changed target or PRD state blocks recovery and requires explicit reconciliation.
|
|
91
|
+
Generated worktree names are shorter and Windows path limits are checked before setup.
|
|
92
|
+
|
|
93
|
+
Goals also use the worker pool. Their asynchronous checks use a separate default-one
|
|
94
|
+
check pool; see [goal resources](GOALS.md). These are concurrency permits, not hard CPU
|
|
95
|
+
or RAM quotas. Current controlled comparisons do not establish general efficiency gains
|
|
96
|
+
over direct Codex; see the [measured comparison](CODEX-COMPARISON-2026-09-29.md).
|
|
97
|
+
|
|
83
98
|
## Safe task decomposition
|
|
84
99
|
|
|
85
100
|
Make a preview with:
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# Goals, verification and recovery implementation plan
|
|
2
|
+
|
|
3
|
+
> **For agentic workers:** Implement independent parallel and resume components with dispatching-parallel-agents; integrate goal policy centrally. Write each regression before changing production behavior.
|
|
4
|
+
|
|
5
|
+
**Goal:** Correct the reproduced completion/resource/integration failures, provide optional persistent native Codex Goals, and prepare a validated 1.20.0 release.
|
|
6
|
+
|
|
7
|
+
**Architecture:** Yoke owns objective-specific executable acceptance and project integration. A negotiated optional native adapter owns a Codex thread, paused at verification boundaries. All actual provider calls use the same admission pool; local goal checks have abortable execution and a separate shared check budget. Durable run identity restores the correct mode and limits.
|
|
8
|
+
|
|
9
|
+
**Tech stack:** TypeScript, Node 20+, Zod, Vitest, native CLI/app-server stdio; no new dependency.
|
|
10
|
+
|
|
11
|
+
**Spec:** docs/GOALS-RESOURCE-AUDIT-2026-09-29.md and docs/CODEX-COMPARISON-2026-09-29.md. Recommendations in those dated reports become implementation scope here; historic results remain historic.
|
|
12
|
+
|
|
13
|
+
## Constraints and decisions
|
|
14
|
+
|
|
15
|
+
- Preserve the original dirty checkout; use the existing codex/benchmark-goals-2026-09-29 worktree.
|
|
16
|
+
- Native goals remain optional, capability negotiated, and never required for other providers. No global CLI configuration changes.
|
|
17
|
+
- Bind each goal to explicitly selected criterion IDs and a stable objective/contract digest. Old unbound goals require explicit binding before accepting completion; provide CLI migration command.
|
|
18
|
+
- Preserve --minutes as cumulative agent-time budget; add explicit cumulative wall-time budget. Report token overshoot and unknown consumption, never silently reset them.
|
|
19
|
+
- Disable unmanaged native multi-agent delegation for managed provider calls; reserve one pool unit per actual call, not an outer unit plus nested units.
|
|
20
|
+
- Keep protected acceptance, scope controls, commits and safe-boundary exploration semantics. No weaker test gates or fabricated CPU/USD savings.
|
|
21
|
+
- Prepare the release locally; tagging, npm publication and GitHub release are not performed by a preparation request.
|
|
22
|
+
|
|
23
|
+
## Review focus
|
|
24
|
+
|
|
25
|
+
Goal contract unrelated to new objective; stale/edited contracts; abort while admission is queued; planner/worker token accounting; dashboard resuming a stale goal instead of active explore run; expired exploration deadline; unavailable native protocol versus genuine authentication errors; native complete while Yoke check fails; retained worktree ownership; Windows full generated path length.
|
|
26
|
+
|
|
27
|
+
## Task 1 — Parallel integration and recoverable candidates
|
|
28
|
+
|
|
29
|
+
Files: src/loop/dispatcher.ts, parallel-command.ts, parallel-adapters.ts, merge-queue.ts; tests/loop parallel/dispatcher suites; reporter integration centrally.
|
|
30
|
+
|
|
31
|
+
- [x] Add failing story-environment regression using independent scopes and an environment-dependent verifier.
|
|
32
|
+
- [x] Preserve verification environment across worker and integration phases and restore previous environment on errors.
|
|
33
|
+
- [x] Persist concrete rejection reasons and recoverable candidate ownership; retain implemented work safely and enable existing recovery paths.
|
|
34
|
+
- [x] Shorten generated worktree names without weakening ownership validation; test Windows path failure handling.
|
|
35
|
+
- [x] Run covering tests and deterministic control probe.
|
|
36
|
+
|
|
37
|
+
## Task 2 — Durable run identity and resume
|
|
38
|
+
|
|
39
|
+
Files: src/loop/run-state.ts, run-command.ts, src/dashboard/server.ts, run-state/dashboard tests.
|
|
40
|
+
|
|
41
|
+
- [x] Add failing mode/provider/selection/deadline and unrelated-goal resume tests.
|
|
42
|
+
- [x] Persist validated safe run options under owned project lock; restore the interrupted run's mode.
|
|
43
|
+
- [x] Preserve absolute exploration deadline and user pauses; explicitly fresh CLI run gets a new deadline.
|
|
44
|
+
- [x] Verify expired deadline starts no provider and unsafe permissions do not silently carry over.
|
|
45
|
+
|
|
46
|
+
## Task 3 — Objective acceptance, policy and budgets
|
|
47
|
+
|
|
48
|
+
Files: src/goals/command.ts, contracts.ts/admission.ts as needed, src/check/command.ts async counterpart, src/cli.ts, src/retrofit/config.ts, goals/check tests.
|
|
49
|
+
|
|
50
|
+
- [x] Add failing unrelated-green-contract test, explicit binding/digest mutation tests, saturated pool and configured-selection tests.
|
|
51
|
+
- [x] Bind goal objective to selected criteria and contract revision; add goal bind CLI and documented legacy migration.
|
|
52
|
+
- [x] Resolve explicit provider/model/effort/bare overrides then project runner configuration; centrally disable unmanaged delegation.
|
|
53
|
+
- [x] Admit actual implementation and planner calls through the pool and release in finally.
|
|
54
|
+
- [x] Add abortable checks with shared local check admission; preserve fingerprints and protected contract evidence.
|
|
55
|
+
- [x] Add cumulative wall-time option and post-attempt overshoot/unknown-usage accounting, streamed cancellation where supported; preserve safe pause and interrupted work.
|
|
56
|
+
|
|
57
|
+
## Task 4 — Optional native Codex adapter
|
|
58
|
+
|
|
59
|
+
Files: src/goals/codex-native.ts, tests/goals/codex-native.test.ts; central goal integration/config.
|
|
60
|
+
|
|
61
|
+
- [x] Test initialized RPC transport, unsupported capability, malformed replies, thread start/resume, abort and bounded cleanup.
|
|
62
|
+
- [x] Persist objective-bound thread identity; execute bounded turns with native goals paused between turns so independent checks arbitrate continuation.
|
|
63
|
+
- [x] Synchronize completion/pause/budget/blocker states after Yoke checks; unavailable method can fall back, real execution/auth errors cannot silently downgrade.
|
|
64
|
+
- [x] Test provider switches and existing-thread/model mismatch; prevent competing auto-continuation.
|
|
65
|
+
|
|
66
|
+
## Task 5 — Review, measurements and release preparation
|
|
67
|
+
|
|
68
|
+
- [x] Review merged components, fix covering regressions and rerun deterministic audit probes with updated expected outcomes.
|
|
69
|
+
- [x] Run authenticated corrected parallel benchmark and report accepted quality/time/tokens with original sample limits.
|
|
70
|
+
- [x] Run TypeScript/build, full tests, metadata checks and package dry run in isolated state.
|
|
71
|
+
- [x] Add dated 1.20.0 CHANGELOG entry with implemented behavior, binding migration, default/native limits and actual validation limits.
|
|
72
|
+
- [x] Synchronize package/lock/provider manifest/README versions; link compact guides for goals/resources/recovery.
|
|
73
|
+
- [x] Read-only provenance audit for changed reports; prepare matching release notes and build archive.
|
|
74
|
+
|
|
75
|
+
## Progress ledger
|
|
76
|
+
|
|
77
|
+
- 2026-09-30: Plan created; three independent implementers assigned parallel correctness, durable resume, native adapter. Controller owns goal policy, check execution, integration, release preparation. Historic benchmark failures remain in reports.
|
|
78
|
+
- 2026-09-30: Implementation and review complete. Final prepublish pipeline passed (1,418 tests, 2 skipped, 155 files), lint/build/Canon/metadata/package checks and zero-vulnerability audit passed. Authenticated parallel and native accounting fixtures succeeded with their stated limits. Local 1.20.0 archive and changelog-derived notes prepared; publication remains out of scope.
|
package/gemini-extension.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.21.1",
|
|
4
4
|
"description": "Cross-agent coding harness for eight supported CLIs: curated skill canon, mechanical safety gates, autonomous loop with proof artifacts. CLI: npm i -g @hecer/yoke",
|
|
5
5
|
"contextFileName": "GEMINI-EXTENSION.md"
|
|
6
6
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@hecer/yoke",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.21.1",
|
|
4
4
|
"description": "One harness, eight agents, zero trust in \"done\" — cross-agent coding harness for Claude Code, Codex CLI, Gemini CLI, Qwen Code, OpenCode, Kilo, Pi and Hermes: one skill canon, mechanical safety gates, an autonomous loop with screenshot/video proofs.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -23,6 +23,10 @@
|
|
|
23
23
|
"bench/run-matrix.mjs",
|
|
24
24
|
"bench/run-parallel-matrix.mjs",
|
|
25
25
|
"bench/analyze-routing-study.mjs",
|
|
26
|
+
"bench/compare-codex.mjs",
|
|
27
|
+
"bench/analyze-codex-comparison.mjs",
|
|
28
|
+
"bench/probe-goal-integration.mjs",
|
|
29
|
+
"bench/probe-parallel-gate-context.mjs",
|
|
26
30
|
"bench/fixtures",
|
|
27
31
|
"bench/results",
|
|
28
32
|
"docs",
|