@hecer/yoke 1.20.0 → 1.22.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (67) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/.codex-plugin/plugin.json +1 -1
  3. package/CHANGELOG.md +46 -0
  4. package/README.md +6 -1
  5. package/bench/analyze-codex-comparison.mjs +90 -17
  6. package/bench/compare-codex.mjs +159 -36
  7. package/bench/result-schema.mjs +132 -0
  8. package/canon/manifest.yaml +1 -1
  9. package/dist/agents/pi-telemetry.js +2 -1
  10. package/dist/agents/process-streams.js +12 -64
  11. package/dist/agents/provider-selection.js +12 -0
  12. package/dist/agents/telemetry.js +52 -52
  13. package/dist/change/inbox.js +8 -3
  14. package/dist/check/command.js +69 -17
  15. package/dist/check/delivery.js +121 -0
  16. package/dist/cli.js +86 -8
  17. package/dist/code-intelligence/budgets.js +138 -0
  18. package/dist/code-intelligence/contracts.js +2 -0
  19. package/dist/code-intelligence/coordinator.js +156 -84
  20. package/dist/code-intelligence/evidence.js +87 -34
  21. package/dist/code-intelligence/mcp-client.js +10 -2
  22. package/dist/code-intelligence/mcp-server.js +10 -10
  23. package/dist/dashboard/analytics.js +5 -3
  24. package/dist/goals/command.js +183 -53
  25. package/dist/goals/usage.js +87 -0
  26. package/dist/loop/candidate-cleanup.js +47 -17
  27. package/dist/loop/candidates.js +17 -11
  28. package/dist/loop/cleanup.js +100 -0
  29. package/dist/loop/dispatcher.js +89 -26
  30. package/dist/loop/failure.js +104 -0
  31. package/dist/loop/gate-snapshot.js +19 -0
  32. package/dist/loop/git.js +1 -1
  33. package/dist/loop/loop.js +100 -66
  34. package/dist/loop/parallel-adapters.js +22 -4
  35. package/dist/loop/parallel-command.js +32 -5
  36. package/dist/loop/recovery.js +23 -5
  37. package/dist/loop/reporter.js +21 -4
  38. package/dist/loop/run-command.js +102 -48
  39. package/dist/loop/runner.js +5 -4
  40. package/dist/loop/worker.js +145 -91
  41. package/dist/observability/history.js +1 -1
  42. package/dist/observability/invocation.js +42 -0
  43. package/dist/observability/usage.js +17 -0
  44. package/dist/prd/command.js +20 -7
  45. package/dist/prd/decompose.js +5 -2
  46. package/dist/retrofit/command.js +22 -0
  47. package/dist/retrofit/config.js +26 -1
  48. package/dist/retrofit/gitignore.js +2 -0
  49. package/dist/retrofit/planners/claude.js +4 -4
  50. package/dist/retrofit/wsl.js +23 -3
  51. package/dist/routing/attempts.js +241 -0
  52. package/dist/routing/capability.js +13 -9
  53. package/dist/routing/optimization.js +73 -0
  54. package/dist/routing/registry.js +7 -1
  55. package/dist/routing/router.js +280 -127
  56. package/dist/setup/command.js +54 -6
  57. package/dist/smoke/command.js +302 -74
  58. package/docs/BENCHMARK-MANIFEST.md +131 -0
  59. package/docs/CODE-INTELLIGENCE.md +43 -1
  60. package/docs/CODEX-COMPARISON-2026-09-29.md +15 -0
  61. package/docs/DELIVERY-JOURNEYS.md +199 -0
  62. package/docs/ECONOMIC-ROUTING.md +180 -0
  63. package/docs/GOALS.md +61 -4
  64. package/docs/RELEASE-VALIDATION-1.22.0.md +115 -0
  65. package/docs/parallel-execution.md +37 -9
  66. package/gemini-extension.json +1 -1
  67. package/package.json +1 -1
@@ -0,0 +1,199 @@
1
+ # Delivery artifacts and executable user journeys
2
+
3
+ Use executable acceptance criteria to decide whether a release is ready. A build command shows that an artifact can be produced; a journey checks a specific user interaction. Yoke can now record the relationship between those criteria, named journeys and the exact artifact files present during a check.
4
+
5
+ The delivery declaration belongs in `.yoke/acceptance.yaml`. Browser smoke steps belong in `.yoke/config.yaml`. Existing smoke flows containing only `name`, `path` and an optional `landmark` continue to work.
6
+
7
+ ## Optional browser steps
8
+
9
+ `yoke flow-smoke` uses the target project's Playwright installation and its Chromium browser. Start the application or preview server before running the command. Yoke does not start a development server or install a browser automatically.
10
+
11
+ For example, add this section to the project's existing `.yoke/config.yaml`:
12
+
13
+ ```yaml
14
+ smoke:
15
+ baseUrl: http://localhost:3000
16
+ flows:
17
+ - name: profile-survives-reload
18
+ path: /login
19
+ landmark: main
20
+ timeoutMs: 60000
21
+ steps:
22
+ - action: fill
23
+ selector: '[data-testid="email"]'
24
+ value: smoke-user@example.invalid
25
+ - action: fill
26
+ selector: '[data-testid="password"]'
27
+ valueEnv: SMOKE_TEST_PASSWORD
28
+ - action: click
29
+ selector: '[data-testid="sign-in"]'
30
+ - action: expect-url
31
+ url: /profile
32
+ - action: fill
33
+ selector: '[data-testid="display-name"]'
34
+ value: Smoke User
35
+ - action: press
36
+ selector: '[data-testid="display-name"]'
37
+ key: Tab
38
+ - action: click
39
+ selector: '[data-testid="save-profile"]'
40
+ - action: expect-visible
41
+ selector: '[data-testid="save-confirmation"]'
42
+ - action: expect-text
43
+ selector: '[data-testid="save-confirmation"]'
44
+ text: Profile saved
45
+ exact: true
46
+ timeoutMs: 10000
47
+ - action: reload
48
+ - action: expect-text
49
+ selector: '[data-testid="profile-name"]'
50
+ text: Smoke User
51
+ exact: true
52
+ ```
53
+
54
+ Supply `SMOKE_TEST_PASSWORD` through the environment using credentials for a dedicated test account, then run:
55
+
56
+ ```sh
57
+ yoke flow-smoke . --label=profile-release
58
+ ```
59
+
60
+ `--url=http://localhost:4173` overrides `smoke.baseUrl`. Flow navigation keeps the existing `baseUrl + path` behavior, so configure their slashes consistently. Each flow gets a fresh browser context with a 1280 × 720 viewport. Steps within one flow share that context; different flows do not share their browser session.
61
+
62
+ ### Supported steps
63
+
64
+ | Action | Fields | Behavior |
65
+ | --- | --- | --- |
66
+ | `click` | `selector` | Click the matching element using Playwright's action waiting. |
67
+ | `fill` | `selector`, exactly one of `value` or `valueEnv` | Fill a form control with a literal value or a captured environment value. Missing environment values fail the step. |
68
+ | `press` | `selector`, `key` | Press a key, such as `Enter` or `Tab`, on the matching element. |
69
+ | `expect-visible` | `selector` | Wait for the matching element to be visible. |
70
+ | `expect-text` | `selector`, `text`, optional `exact` | Wait for visible text. Whitespace is normalized. Matching is case sensitive and uses a literal substring by default; `exact: true` compares the complete normalized text. |
71
+ | `expect-url` | `url` | Wait for the exact resolved URL. A relative URL resolves against the effective base URL. This is not a glob or regular expression. |
72
+ | `reload` | No additional action fields | Reload the page, wait for its load event and reject a non-successful HTTP response. |
73
+
74
+ Each step may specify `timeoutMs`. Unknown actions and extra action fields are rejected before browser launch. There is no arbitrary JavaScript action, general scripting language, branch or loop. Use a project E2E test command for more complex behavior, multiple pages, downloads, native applications or specialized fixtures.
75
+
76
+ ### Limits and failure behavior
77
+
78
+ A flow may contain **1–50 steps**. Its default total timeout is **60,000 ms**, with a configurable maximum of **120,000 ms**. This budget includes browser-context preparation, navigation, the optional landmark and all steps. Navigation has its own maximum of 30,000 ms. A landmark has a maximum of 10,000 ms.
79
+
80
+ A step defaults to **10,000 ms** and may request at most **30,000 ms**. Its actual allowance is capped by the remaining flow budget. Text waiting has both an attempt limit and a deadline. Screenshot capture, context closing and video finalization each have a separate 5,000 ms allowance after interaction ends; browser startup and shutdown also have bounded waits. The flow timeout is therefore not a promise that the whole multi-flow command, including evidence capture and cleanup, ends within that duration.
81
+
82
+ The first failed step stops the remaining steps of that flow. The report marks them `skipped`. Later flows still run. Non-successful navigation, page errors and error-level console events also fail a flow. A screenshot is required for a passing result. Failure video remains best-effort evidence and does not replace a missing screenshot.
83
+
84
+ Fill values are limited to 8,192 characters, including values read from the environment. Environment-backed fill values are captured once when the run begins.
85
+
86
+ ## Evidence from a smoke run
87
+
88
+ Evidence is written to `.yoke/proof/<label>/`:
89
+
90
+ - A screenshot for each flow where screenshot capture succeeds.
91
+ - A `.webm` video for a failed flow when Playwright can finalize it.
92
+ - `report.json`, a versioned machine-readable report.
93
+
94
+ The default label is `YOKE_STORY` when present, otherwise `latest`; `--label` takes priority. Labels and filenames are sanitized, and colliding sanitized flow names receive distinct filenames. Running the same label replaces its previous evidence. A run that cannot start because its configuration, browser, source identity or proof location is unavailable does not erase the prior evidence during that preflight.
95
+
96
+ The report contains:
97
+
98
+ - Per-flow status, navigation status, optional landmark status and duration.
99
+ - Each configured step's index, action, status, duration and a bounded failure category.
100
+ - Relative screenshot and failure-video filenames.
101
+ - Source fingerprints before and after the run, and whether they match.
102
+ - A digest of the effective smoke configuration, URL override and captured fill inputs.
103
+ - Node version, operating-system platform, architecture, configured viewport, browser version when available and the effective URL's origin.
104
+
105
+ Raw configuration, selectors, expected text, fill values, environment-variable names and browser exception call logs are not serialized into the report. URL credentials, query strings and fragments are not included in its recorded origin. Driver errors are represented by categories such as `timeout`, `step-failed` or `navigation-failed`, because their original messages can contain form values. Screenshots and video still show what the application visibly displays; use test data appropriate for those visual artifacts.
106
+
107
+ Exit code `0` means all flows passed, their screenshots were captured, the source fingerprints before and after the run matched and no browser-cleanup failure was reported. Exit code `1` represents a failed run or invalidated evidence. Exit code `2` covers unavailable prerequisites or evidence storage. Invalid step configuration is rejected before browser launch.
108
+
109
+ The local source fingerprint does **not** prove that a remote server is running that source revision. Arrange the project's preview or deployment checks so the browser exercises the intended build. Record a separate deployment verification criterion when that relationship matters.
110
+
111
+ ## Bind release artifacts and journeys to acceptance criteria
112
+
113
+ Add an optional `delivery` section to `.yoke/acceptance.yaml`:
114
+
115
+ ```yaml
116
+ version: 1
117
+ protected:
118
+ - .yoke/config.yaml
119
+ - playwright.config.ts
120
+ - tests/e2e/profile.spec.ts
121
+ criteria:
122
+ - id: profile-e2e
123
+ text: A signed-in user can save a profile and recover it after reloading.
124
+ commands:
125
+ - npm run test:e2e -- tests/e2e/profile.spec.ts
126
+ - id: profile-smoke
127
+ text: The release preview completes the configured profile journey.
128
+ commands:
129
+ - yoke flow-smoke . --label=profile-release
130
+ delivery:
131
+ version: 1
132
+ artifacts:
133
+ - path: dist/release.zip
134
+ criteria: [profile-e2e, profile-smoke]
135
+ journeys:
136
+ - id: profile-update
137
+ text: Sign in, change a profile field, save it, reload and observe the saved value.
138
+ criteria: [profile-e2e, profile-smoke]
139
+ environment:
140
+ name: Local release preview with Chromium
141
+ url: http://localhost:3000
142
+ ```
143
+
144
+ The script names, test files and artifact path in this example must exist in the application project. `delivery.journeys` names the intended user experience and references executable criteria; it does not execute browser steps by itself. The example executes smoke steps through the `profile-smoke` criterion's command.
145
+
146
+ **Build declared artifacts before starting `yoke check`:**
147
+
148
+ ```sh
149
+ npm run build:release
150
+ # Start the project's preview server for that release in another terminal.
151
+ yoke check . --json
152
+ ```
153
+
154
+ `build:release` here denotes the project's build script that creates `dist/release.zip`. A criterion that creates the declared artifact during `yoke check` is too late: the artifact must already exist for its initial snapshot.
155
+
156
+ The check report records artifact SHA-256 hashes and sizes before and after the commands. An artifact missing at either snapshot, differing between snapshots or using an unsafe path fails artifact binding. Declared artifacts must be regular project-relative files without symbolic links or path escapes. Explicit artifact hashing also covers ordinary ignored build outputs.
157
+
158
+ Artifact hashing accepts at most **512 MiB per file**. Each snapshot, before or after verification, has a combined **1 GiB byte budget** and a **10-second deadline**; an earlier enclosing check deadline also applies. Exceeding a limit or cancelling the check prevents a passing binding. Data is streamed in bounded chunks, with deadline and cancellation checks between synchronous reads. An individual blocked filesystem read is not preempted by that deadline. The declaration accepts at most 32 artifacts and 100 named journeys.
159
+
160
+ Artifact and journey status comes from their referenced criteria. All referenced criteria must pass for a passing status; a failed criterion fails the binding, while an unexecuted criterion remains unverified. Source-integrity, cancellation or artifact-binding problems also prevent a passing delivery result. Unknown criterion IDs and duplicate artifact paths or journey IDs are rejected.
161
+
162
+ The report labels this relationship **`binding: declared-criteria`**. Matching artifact hashes plus green criteria establish the declared relationship and the same artifact content at both snapshots. This is not continuous monitoring of the file between snapshots. The project command must actually exercise the intended artifact; Yoke cannot infer that a test which ignores its APK, ZIP or binary argument has verified that file.
163
+
164
+ Protect the acceptance manifest and the test infrastructure with the existing `yoke check . --protect` workflow after reviewing the contract. See [Verified projects](VERIFIED-PROJECTS.md) for protection and intentional-baseline-refresh behavior.
165
+
166
+ ## Android example: verify the built APK on a specified device
167
+
168
+ A browser journey does not install an APK or validate Android permission handling. Use an application-owned executable test for that task and reference it from acceptance:
169
+
170
+ ```yaml
171
+ version: 1
172
+ protected:
173
+ - tools/check-apk-install.mjs
174
+ - tests/android/tracking-fixture.json
175
+ criteria:
176
+ - id: tracking-installed-apk
177
+ text: The declared APK installs and completes the tracking fixture on the selected emulator.
178
+ commands:
179
+ - node tools/check-apk-install.mjs --apk=android/app/build/outputs/apk/debug/app-debug.apk --serial=emulator-5554
180
+ delivery:
181
+ version: 1
182
+ artifacts:
183
+ - path: android/app/build/outputs/apk/debug/app-debug.apk
184
+ criteria: [tracking-installed-apk]
185
+ journeys:
186
+ - id: record-and-reopen
187
+ text: Grant the required permission, start tracking, receive fixture GPS points, stop, save and reopen the route after restarting the app.
188
+ criteria: [tracking-installed-apk]
189
+ environment:
190
+ name: Android emulator emulator-5554 with the project's configured system image
191
+ ```
192
+
193
+ `check-apk-install.mjs` and the fixture are **project-supplied tests**, not bundled Yoke commands. The script must install that exact APK, select and verify its device, perform the required permission and tracking actions, inspect the resulting route and process logs, and exit nonzero when an assertion fails. Build the APK first, prepare the emulator, then run `yoke check`.
194
+
195
+ The declared environment name is metadata. It does not detect the actual emulator image, phone model or OS version. Have the project test verify those properties when they are acceptance requirements. Passing on one emulator does not establish correct background tracking on a particular Vivo phone, and passing against one deployment does not certify every deployment.
196
+
197
+ ## Validation scope
198
+
199
+ The smoke engine's repository tests use filesystem-backed browser doubles to verify step order, failure handling, timeout budgets, redaction, source/configuration binding and evidence persistence. They do not install browsers or certify an application. Application-level evidence comes from successfully running the configured commands with the project's actual browser, build, server or device.
@@ -0,0 +1,180 @@
1
+ # Measured capability routing
2
+
3
+ Yoke 1.22 adds optional economic selection to capability routing. It compares
4
+ observed, independently verified execution sequences. It does not assign an
5
+ estimated success probability to a model or assume that a profile labelled
6
+ `low` has the lowest cost of completing a task.
7
+
8
+ ## Configuration and compatibility
9
+
10
+ ```yaml
11
+ routing:
12
+ enabled: true
13
+ strategy: capability
14
+ optimization:
15
+ version: 1
16
+ objective: balanced
17
+ minSamples: 20
18
+ ```
19
+
20
+ This fragment supplements the project's existing routing profiles and limits.
21
+ `minSamples` defaults to 20 and has a minimum of 10. Existing configurations
22
+ without `routing.optimization` retain their previous capability ordering. New
23
+ setup-generated capability configurations enable the conservative `balanced`
24
+ policy; this does not change selection until sufficient comparable measurements
25
+ exist.
26
+
27
+ The planner's task/risk assessment still determines the minimum capability tier.
28
+ Configured provider constraints, allowed roles, `maxTier`, parent-fallback policy,
29
+ and the existing gate-failure reliability filter remain in force. Economic
30
+ selection only compares profiles that survive those checks. Explicit project
31
+ routing rules retain their precedence.
32
+
33
+ ## What a sample measures
34
+
35
+ The comparison unit is a bounded **execution sequence for one task contract**,
36
+ attributed to the profile that started it. If a small model needs a targeted
37
+ repair and then escalation to a stronger model, all those measured attempts
38
+ belong to the initial profile's sequence. Charging only the final successful
39
+ model would hide the cost of the failed attempts.
40
+
41
+ Within a fully accounted execution attempt, the durable call ledger includes:
42
+
43
+ - The routing/planning call made inside that execution, when one was needed.
44
+ - The implementation call, including reported usage from interrupted calls.
45
+ - Additional reviewer, critic, repair and model-driven quality calls joined by
46
+ the loop reporter before its final outcome.
47
+
48
+ Stable call IDs prevent a worker's usage from being counted again when its
49
+ aggregate reaches the reporter. An explicit attempt ID can join a call directly.
50
+ Without one, a role call is joined only if that story has exactly one open
51
+ attempt. Competing candidates make the attribution ambiguous; all affected
52
+ samples become incomplete instead of assigning the bill by guesswork.
53
+
54
+ The measured scope is `execution-attempt`. It is not the total cost of producing
55
+ or shipping a product. Separately prepared PRD drafts, batch assessments, change
56
+ planning, human work and later operational costs are not amortized into these
57
+ samples. Their separately emitted telemetry remains useful for project-level
58
+ analysis. The duration covers each admitted attempt through its recorded gate
59
+ outcome, summed across the sequence; it does not include idle time between
60
+ separate attempts or a developer's review time.
61
+
62
+ Only native loop paths with the complete role-accounting contract enable this
63
+ scope. Injected runners/gates and ambiguous candidate races remain `worker`
64
+ scope. Their measurements are still available diagnostically, but cannot supply
65
+ the complete evidence required for economic selection.
66
+
67
+ ## Evidence and missing data
68
+
69
+ For each eligible starting profile, the router examines its most recent
70
+ `minSamples` comparable sequences. They must share the project, implementation
71
+ role, task-assessment dimensions, effective execution-policy key, requested
72
+ provider/model/effort/variant, and the current reported concrete starting model.
73
+ The policy key binds the effective gate and review configuration, including
74
+ commands and retries. A change to the concrete model starts a new evidence
75
+ window even when its requested alias remains the same.
76
+
77
+ A sequence is complete only when every reserved attempt has an outcome and the
78
+ sequence has either passed independent gates or exhausted its bounded attempt
79
+ allowance. Every attempt must have known token coverage, complete reported cost,
80
+ the full execution scope, and the same policy key. Infrastructure failures are
81
+ not treated as evidence of model reasoning quality and do not qualify for the
82
+ economic comparison.
83
+
84
+ One recent unknown or partial sequence keeps its window ineligible. The router
85
+ does not discard that row and search for an older successful subset. Open
86
+ sequence records are replaced by their later completion record, and retries do
87
+ not turn into additional independent samples. The optional observation registry
88
+ uses a 30-day evidence window and reads at most its latest 1,000 events; losing
89
+ economic history leads back to conservative selection.
90
+
91
+ Unknown usage is distinct from measured zero. Known partial token or dollar
92
+ amounts are retained with incomplete-coverage flags. For example, a fully
93
+ measured worker remains visible when its planner supplied no usage. OpenCode
94
+ and Kilo step summaries require coverage of every completed step before they
95
+ claim complete token totals; absent optional cache or cost fields are not
96
+ converted into measured zeros.
97
+
98
+ No reported dollars means no cost ranking. Yoke does not invent a bill from
99
+ declared cost tiers, token counts or a presumed provider price table. This can
100
+ leave economic routing at its conservative baseline for providers that do not
101
+ report sufficient telemetry.
102
+
103
+ ## The three objectives
104
+
105
+ The baseline is the first eligible profile under the existing tier, declared
106
+ cost-tier and profile-ID ordering. Both the baseline and an alternative require
107
+ a complete evidence window. An alternative must have at least as many accepted
108
+ sequences in that equally sized window as the baseline.
109
+
110
+ The router computes two observed ratios:
111
+
112
+ ```text
113
+ cost per accepted sequence = total measured sequence dollars / accepted sequences
114
+ time per accepted sequence = total measured sequence duration / accepted sequences
115
+ ```
116
+
117
+ The numerators include completed failed sequences and every measured repair or
118
+ escalation in the window. A profile with zero accepted sequences is ineligible.
119
+
120
+ | Objective | Condition for choosing an alternative |
121
+ | --- | --- |
122
+ | `cost` | Lower observed dollars per accepted sequence, with no lower observed acceptance count. |
123
+ | `speed` | Lower observed duration per accepted sequence, with no lower observed acceptance count. |
124
+ | `balanced` | No higher cost or duration, and at least one strictly lower, with no lower observed acceptance count. |
125
+
126
+ `balanced` uses this conservative comparison instead of inventing a dollar
127
+ value for a second of runtime. Cost and speed can disagree; use the corresponding
128
+ explicit objective if that tradeoff is intentional. Equal scores preserve the
129
+ existing ordering.
130
+
131
+ Decision reasons identify the objective, sample count, observed acceptance
132
+ count, and measured ratios. When the baseline lacks enough data, they explain
133
+ the missing evidence and retain its existing ordering. These are observational
134
+ measurements, not calibrated forecasts or statistical confidence guarantees.
135
+ Task mix, repeated correlated failures and changing model behavior can still
136
+ affect the comparison. The minimum window is a conservative product policy,
137
+ not a claim that twenty examples prove superiority.
138
+
139
+ Reviewer and critic approval frequency is not independent evidence that their
140
+ reviews were good. Economic reordering of those roles therefore remains disabled
141
+ until an independent role-quality measure is available; their risk-based floors
142
+ and explicit model selections continue to apply.
143
+
144
+ ## Persistent attempt admission
145
+
146
+ Capability routing reserves an attempt before invoking its implementation
147
+ provider. Configured bounded legacy routes use the same account. The account
148
+ is stored under the stable project root, keyed by story and bound task/plan
149
+ contract, and is independent of the optional global observation registry.
150
+ Exclusive immutable reservation files prevent two processes from replacing the
151
+ same slot. An unreadable or unwritable authoritative account blocks admission.
152
+
153
+ The capability allowance is the smaller of `maxAttempts` and the existing
154
+ tier-dependent bound: five attempts from `light`, four from `standard`, three
155
+ from `strong`, and two from `frontier`. Independent outer loop, goal and local
156
+ retry limits can stop execution sooner. Infrastructure or interrupted attempts
157
+ still occupy their reserved slots, but do not become verified reasoning failures
158
+ that automatically demand a stronger model.
159
+
160
+ Deleting or filling the optional registry, restarting a runner, or executing in
161
+ another disposable worktree does not restore those slots. Repeated preparation
162
+ or budget refusals before an implementation is admitted do not consume an
163
+ implementation attempt. A planner may itself incur cost before a later worker
164
+ is refused; that paid call is still reported. Goal callers recheck their durable
165
+ budget after planning and before implementation reservation.
166
+
167
+ A changed task/plan contract has a new account. Resume an unchanged blocked task
168
+ only after inspecting its reason and performing the required replan or explicit
169
+ configuration change. Do not delete authoritative attempt state to disguise
170
+ retries as a new run.
171
+
172
+ ## Validation limits
173
+
174
+ The release tests use deterministic fake calls and synthetic observations. They
175
+ cover registry write failure and history eviction, restart-safe reservations,
176
+ known partial usage, synchronous/asynchronous provider telemetry, paid early
177
+ blocks, callback/write failures, complete escalation accounting, cold starts,
178
+ model/policy changes and objective selection. They are regression tests for the
179
+ accounting and decision rules, not provider performance benchmarks. No claimed
180
+ cost or speed improvement follows from those test fixtures.
package/docs/GOALS.md CHANGED
@@ -29,11 +29,44 @@ IDs and executable commands and binds the objective to a digest of the acceptanc
29
29
  manifest. It cannot determine whether your tests fully express a free-text request.
30
30
  All project acceptance and configured regression checks must still pass.
31
31
 
32
+ Each independent check writes its full report, including declared artifact and
33
+ journey evidence. Goal state keeps the check ID in `lastCheck` and its report path
34
+ in `lastCheckEvidencePath`; `goal handoff` includes both. The report binds its
35
+ findings to the checked workspace. Declaring an artifact or journey is not itself
36
+ proof that a criterion passed.
37
+
32
38
  Existing goals without a binding now stop before any model call. Review their tests,
33
39
  then explicitly bind them with `yoke goal bind . --criteria=checkout`. Attempts and
34
40
  measured consumption remain intact. Changed acceptance infrastructure still requires
35
41
  an explicit protection refresh after review; binding never refreshes protected tests.
36
42
 
43
+ ## Prepare routing explicitly
44
+
45
+ Capability routing uses the same goal ID, objective and structured executable
46
+ criteria as goal acceptance. It follows the project's `routing.assessmentPolicy`,
47
+ explicit `routing.rules`, profile limits and optional optimization settings.
48
+
49
+ With `routing.assessmentPolicy: prepared`, first prepare the bound goal:
50
+
51
+ ```sh
52
+ yoke goal assess . --runner=codex
53
+ yoke goal run . --runner=codex
54
+ ```
55
+
56
+ `goal assess` makes at most one read-only planning call and saves its assessment.
57
+ It reuses an existing assessment for the current contract. It does not implement
58
+ the objective or create an implementation attempt. Preparation consumes the same
59
+ durable token and provider-time budgets as execution. The planning provider and
60
+ model follow the project's planning configuration.
61
+
62
+ The following run reuses that assessment. A missing or stale prepared assessment
63
+ blocks execution before any model call, including when an explicit routing rule
64
+ exists. Changes to the objective, executable criteria, selected runner or approved
65
+ `.yoke/plan.md` invalidate the contract key. Run `goal assess` again after reviewing
66
+ those changes; acceptance changes still require explicit binding and protection
67
+ review. The `on-demand` policy may assess during execution; explicit rules avoid
68
+ the selection-planner call under that policy.
69
+
37
70
  ## Resource and budget controls
38
71
 
39
72
  Goals and story loops share the project lock and the user-level worker pool. Routing
@@ -41,24 +74,45 @@ planner and implementation calls each acquire a permit; native model delegation
41
74
  disabled. Explicit runner/model/effort/bare options override project runner defaults.
42
75
  Execution uses Yoke's safe permission profile.
43
76
 
77
+ Runner defaults apply only to the provider that owns them. For example,
78
+ `--runner=claude` does not inherit a Codex model, effort, provider, variant or bare
79
+ setting from `runner`. Explicit command options still apply; when keeping the same
80
+ runner, omitted options retain its configured defaults.
81
+
44
82
  | Setting | Meaning |
45
83
  | --- | --- |
46
84
  | `--attempts=N` on `set` or `budget` | Maximum cumulative attempts, default 3 |
47
85
  | `--minutes=N` | Cumulative admitted provider work, default 30 minutes; excludes capacity waits and independent checks |
48
86
  | `--wall-minutes=N` | Optional cumulative run time including admission, checks and state synchronization |
49
- | `--tokens=N` | Cumulative measured input + output tokens; unknown consumption stops further budgeted work |
87
+ | `--tokens=N` | Cumulative measured input + output tokens, checked before every planning and implementation call; unknown consumption stops further budgeted work |
50
88
  | `YOKE_MAX_PARALLEL_WORKERS=1..8` | Shared model/integration permit ceiling, default 3 |
51
89
  | `YOKE_MAX_PARALLEL_CHECKS=1..8` | Separate shared asynchronous goal verification ceiling, default 1 |
52
90
 
91
+ Planner consumption is persisted before admitting its implementation worker. If
92
+ planning reaches the ceiling, overruns it, or leaves usage unknown, no next model
93
+ call is admitted. A completed objective exactly at its measured ceiling may finish;
94
+ the same balance cannot authorize another model call. Native Codex receives only
95
+ the remaining token allowance after earlier planning and implementation calls.
96
+
53
97
  After a measured token overrun, Yoke retains evidence and work and reports `blocked`
54
98
  even if acceptance passed. Update the budget explicitly with `yoke goal budget .
55
99
  --tokens=50000`; `--clear-token-budget` explicitly removes that ceiling. Providers
56
- reporting usage only after a call cannot provide a hard mid-call token cap. Native
100
+ reporting usage only after a call cannot provide a hard mid-call token cap. This is
101
+ a call-admission limit plus cancellation where telemetry permits it. Native
57
102
  Codex usage units and Yoke's token accounting are recorded separately. Native
58
103
  stream updates use cumulative turn deltas and request cancellation when measured
59
104
  consumption exceeds the goal ceiling. Notifications arrive after model requests,
60
105
  so this can still overshoot within a request; progress updates are never added
61
- twice to persisted final usage.
106
+ twice to persisted final usage. Partial counts remain a known lower bound, with
107
+ incomplete measurement explicitly recorded. Raising a token ceiling does not make
108
+ unknown consumption known; explicitly removing the ceiling is a separate decision.
109
+
110
+ The goal's `usageLedger` records each call before dispatch, then records its final
111
+ usage and duration. Recovery retains finished planner calls and marks unfinished
112
+ calls as interrupted with unknown final usage. Each call has a stable ID and its
113
+ own usage event; attempt summaries are not emitted again as additional token usage.
114
+ Earlier goal files retain their existing attempt totals once, and new calls are
115
+ charged separately. Preparation and resume never reset the previous consumption.
62
116
 
63
117
  The asynchronous goal checker can interrupt its own command tree at a deadline or
64
118
  pause. Synchronous `yoke check` and story gates retain their existing execution
@@ -78,7 +132,10 @@ Or set `goals.nativeCodex: true` in `.yoke/config.yaml`. The default is off;
78
132
  Yoke negotiates the local Codex app-server goal capability, stores the thread binding,
79
133
  pauses native automatic continuation, and explicitly starts one bounded development
80
134
  turn. The same thread resumes on later attempts. Only Yoke's independent acceptance
81
- can synchronize it to `complete`. Provider/model changes create a new binding.
135
+ can synchronize it to `complete`. Provider/model changes detach the previous binding
136
+ and retain it in `detachedNativeBindings`. A later Codex execution creates a binding
137
+ for the selected model; a goal finished by another provider does not need the old
138
+ Codex runtime to synchronize a historical thread.
82
139
  An unsupported native goal method falls back to ordinary Codex execution; authentication,
83
140
  transport and execution failures remain visible. No global Codex configuration changes
84
141
  are needed. Support depends on the installed CLI, not on the model name alone.
@@ -0,0 +1,115 @@
1
+ # Yoke 1.22.0 validation — 2026-10-03
2
+
3
+ Implementation was based on `f2d9180f4e72afe79c9930c62625874d734db3c6`
4
+ (1.21.1). Parallel implementation and independent reviews covered routing and
5
+ telemetry, goals, worker recovery, Code Intelligence, benchmark provenance and
6
+ delivery evidence. This record separates deterministic regression evidence from
7
+ provider performance or publication claims.
8
+
9
+ ## Local release gates
10
+
11
+ Environment: Linux, Node.js **24.19.0**, npm **11.9.0**.
12
+
13
+ | Gate | Result |
14
+ | --- | --- |
15
+ | Baseline before changes | 1,426 tests in 156 files passed. |
16
+ | `npm run prepublishOnly` | Passed: lint, build, complete test suite, README metadata and package dry-run. |
17
+ | Complete corrected suite | **1,610 tests in 177 files passed**, no skipped tests; 63.75 seconds in this run. |
18
+ | Canon validation | Passed with `node --import tsx src/cli.ts validate canon`. |
19
+ | Dependency audit | `npm audit --audit-level=high`: **0 vulnerabilities**. |
20
+ | Whitespace validation | `git diff --check` passed. |
21
+ | Tarball installation | Passed in a fresh prefix with lifecycle scripts disabled: installed binary starts and validates its packaged Canon. |
22
+ | Package contents | 376 files; all five packaged version manifests report 1.22.0; new runtime modules are included, with no root test suite, dependency tree or `.yoke` runtime state. |
23
+
24
+ The usual `npm run yoke -- validate canon` wrapper was blocked by this host's
25
+ restriction on the tsx CLI's local IPC socket (`EPERM`). Loading the same source
26
+ CLI through Node's tsx import hook passed. No project permission or verification
27
+ policy was weakened to accommodate that environment restriction.
28
+
29
+ The first integrated runs exposed issues added during concurrent implementation;
30
+ the corrected final suite above passed. A registry-eviction test additionally
31
+ needed an explicit clock advance: equal-millisecond event filenames have UUID
32
+ tie ordering, so wall-clock timing could leave an older observation in its
33
+ fixture. The test now establishes eviction deterministically and still verifies
34
+ that the independent durable attempt account remains exhausted.
35
+
36
+ The first GitHub matrix passed both Linux jobs and exposed a Windows-only fixture
37
+ issue in the new recovery tests. Windows TEMP can use a DOS short-name alias while
38
+ recovery intentionally persists canonical paths. The JavaScript `realpathSync`
39
+ implementation can retain that alias, so an initial fixture correction was
40
+ insufficient. Both fixtures now use `realpathSync.native`, matching recovery's
41
+ handle-based resolution, before asserting exact ownership paths. Production
42
+ recovery validation remains unchanged. The full matrix is rerun on that correction.
43
+
44
+ ## Regressions covered
45
+
46
+ - Durable implementation admission across restarts, registry write failures and
47
+ eviction of more than 1,000 optional observations; concurrent reservations and
48
+ interrupted attempts retain their spent slots.
49
+ - Per-call goal budget checks before dispatch, planning consumption before worker
50
+ admission, native/ordinary provider selection and complete or partial usage
51
+ through failure paths.
52
+ - Stable usage IDs, reviewer/critic/repair joins, partial provider telemetry and
53
+ dashboard aggregates that preserve known child calls.
54
+ - Economic selection of comparable complete execution sequences, including
55
+ failed attempts and escalation, model/policy changes, cold starts, missing
56
+ costs and all three objectives. Synthetic observations test the decision
57
+ rules; they do not measure model superiority.
58
+ - Real temporary Git repositories for incomplete-worker recovery and candidate
59
+ retention. Pause, cancellation, failed gates and selection errors preserve
60
+ work; ownership, contract, base and repository validation guard reuse.
61
+ - Unchanged gate reuse, invalidation after relevant changes, typed completion
62
+ repair, persistent no-progress blocking and separate candidate scopes.
63
+ - Code Intelligence reference direction, unknown freshness, incomplete traversal,
64
+ total response budgets, shared deadlines and pre-mutation receipt admission.
65
+ - Delivery artifact/source/acceptance binding, bounded hashing, unsafe paths,
66
+ command outcomes and secret-safe multi-step browser proof reports.
67
+ - Benchmark manifests binding actual source/build identity, fixture, acceptance,
68
+ model and startup policy; incompatible and legacy rows stay out of verified
69
+ comparison groups.
70
+
71
+ ## Limits and operational behavior
72
+
73
+ No paid LLM calls, authenticated provider benchmarks, production deployments or
74
+ real-browser end-to-end application runs were performed for this release. Browser
75
+ journey tests use deterministic browser seams. Model cost/latency observations
76
+ are fixtures, and there is no claimed percentage saving or general speedup.
77
+
78
+ Old routing configurations retain their ranking unless optimization is enabled.
79
+ New setup configurations choose `balanced`, requiring 20 comparable complete
80
+ samples for baseline and alternative; the supported minimum is 10. Cost tiers
81
+ and versioned setup profiles are configuration priors. Missing actual costs,
82
+ partial usage or ambiguous role attribution prevent economic promotion.
83
+
84
+ Provider counters reported only at call completion cannot impose an in-flight
85
+ hard token cap. Known overruns are retained and block further dispatch; unknown
86
+ interrupted usage needs an explicit budget decision. Separately prepared project
87
+ planning and human time are not amortized into execution-sequence comparisons.
88
+
89
+ Recovery preserves code and evidence, not a previous process's callback closures.
90
+ A recovered integration cannot manufacture a missing economic success outcome.
91
+ Candidate races resume the first retained alternative through ordinary recovery;
92
+ additional alternatives remain available for inspection and explicit cleanup.
93
+ Custom lifecycles without `retain` preserve their prior cleanup contract.
94
+
95
+ Artifact hashes compare pre-check and post-check contents. They do not detect a
96
+ change reverted between snapshots, establish reproducible builds or prove that a
97
+ particular deployed binary or device was exercised. Hashing is bounded to 512 MiB
98
+ per artifact, 1 GiB and 10 seconds per snapshot; checks run between synchronous
99
+ reads and cannot preempt an individual blocked filesystem read. Delivery commands
100
+ and declared environment labels remain project-authored evidence.
101
+
102
+ Code Intelligence uses a conservative UTF-8 byte-based token estimate, not
103
+ provider token accounting. Tiny budgets may only return the documented bounded
104
+ error envelope. Backend freshness remains unknown without independent evidence.
105
+
106
+ ## Remote verification and publication
107
+
108
+ The repository CI matrix runs Linux and Windows with Node 20 and 24. Its status
109
+ is attached to the pushed release commit and pull request in GitHub Actions;
110
+ local success alone does not establish cross-platform success.
111
+
112
+ The dated [changelog](../CHANGELOG.md#1220--2026-10-03) is the source of release
113
+ notes. A Git commit, push or tag does not imply npm publication. Publishing a
114
+ GitHub Release triggers the separate trusted npm workflow; publication must be
115
+ verified independently before reporting this version as available from npm.
@@ -80,15 +80,43 @@ one serial integration lane when integration measurements exist. Older runs with
80
80
  remain visible as missing integration history; forecasts are empirical ranges, not deadlines, and do
81
81
  not predict future contention from other projects.
82
82
 
83
- ### Rejected integration recovery
84
-
85
- Worker and integrated-tree gates receive the same `YOKE_STORY` context. A rejected
86
- candidate is retained with its reason and an `integration-recovery.json` proof record;
87
- it is not silently discarded and regenerated. The next parallel run can reuse it
88
- without a new implementation model call when canonical project/worktree ownership,
89
- Git registration, target base and PRD digest still match. Integration gates run again.
90
- Changed target or PRD state blocks recovery and requires explicit reconciliation.
91
- Generated worktree names are shorter and Windows path limits are checked before setup.
83
+ ### Interrupted work and rejected integration recovery
84
+
85
+ Worker and integrated-tree gates receive the same `YOKE_STORY` context. Production
86
+ parallel execution retains useful work after failed gates, pause, cancellation and
87
+ decision boundaries. Versioned records under `.yoke/integration-recovery/` bind the
88
+ worktree, original owner, Git base, PRD contract and failure reason to a recovery phase:
89
+
90
+ - `implementation` resumes the retained implementation with the previous failure as
91
+ feedback. An unchanged or descendant target is allowed after ownership and history
92
+ validation; integration still rebases and runs its gates.
93
+ - `integration` resumes an already prepared candidate without another implementation
94
+ call. It requires the unchanged target base and reruns integration gates.
95
+
96
+ Both phases require a valid registered worktree in the same repository and the same
97
+ PRD contract. Untrusted or stale state requires explicit reconciliation. Generated
98
+ worktree names are shorter and Windows path limits are checked before setup.
99
+
100
+ Optional candidate races defer destructive loser cleanup until successful selection.
101
+ If every candidate fails, the race pauses or is cancelled, or selection cannot finish,
102
+ materialized candidates are retained after process cleanup. The first retained
103
+ candidate occupies the ordinary story recovery record and resumes on the next run.
104
+ Additional alternatives have separate records and remain available for inspection
105
+ and explicit worktree cleanup; Yoke does not automatically schedule another race to
106
+ compare those retained alternatives. Successful selection still removes unselected
107
+ alternatives. A failed retention write reports recovery information and leaves the
108
+ files in place. Custom or injected lifecycles without the optional `retain` hook keep
109
+ their own existing cleanup contract.
110
+
111
+ Retention preserves implementation and evidence, but consumes disk until integration
112
+ or intentional cleanup. It does not reconstruct a previous process's routing callback:
113
+ a recovered integration cannot manufacture a missing economic success observation.
114
+
115
+ Repeated failures are tracked persistently against source, failure and approved
116
+ contract identity. The first unchanged failure permits retry, the second requests a
117
+ focused diagnosis, and the third blocks as `no-progress`. A real source, assertion,
118
+ failure or approved-plan change resets the sequence. Volatile test durations and
119
+ reporter timestamps do not reset it; concurrent candidates have separate scopes.
92
120
 
93
121
  Goals also use the worker pool. Their asynchronous checks use a separate default-one
94
122
  check pool; see [goal resources](GOALS.md). These are concurrency permits, not hard CPU