@hecer/yoke 1.21.1 → 1.23.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (103) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/.codex-plugin/plugin.json +1 -1
  3. package/CHANGELOG.md +48 -0
  4. package/README.md +8 -1
  5. package/TODOS.md +6 -0
  6. package/bench/analyze-codex-comparison.mjs +90 -17
  7. package/bench/compare-codex.mjs +159 -36
  8. package/bench/result-schema.mjs +132 -0
  9. package/canon/manifest.yaml +1 -1
  10. package/canon/skills/visual-verification/SKILL.md +25 -2
  11. package/canon/tools/codex-rtk-hook.mjs +6 -16
  12. package/dist/agents/pi-telemetry.js +2 -1
  13. package/dist/agents/process-streams.js +12 -64
  14. package/dist/agents/provider-selection.js +12 -0
  15. package/dist/agents/telemetry.js +52 -52
  16. package/dist/change/inbox.js +8 -3
  17. package/dist/check/command.js +69 -17
  18. package/dist/check/delivery.js +121 -0
  19. package/dist/cli.js +91 -3
  20. package/dist/code-intelligence/adapters/mcp.js +1 -0
  21. package/dist/code-intelligence/budgets.js +138 -0
  22. package/dist/code-intelligence/contracts.js +2 -0
  23. package/dist/code-intelligence/coordinator.js +159 -85
  24. package/dist/code-intelligence/evidence.js +87 -34
  25. package/dist/code-intelligence/index.js +1 -0
  26. package/dist/code-intelligence/mcp-client.js +25 -6
  27. package/dist/code-intelligence/mcp-server.js +14 -11
  28. package/dist/code-intelligence/preflight.js +71 -0
  29. package/dist/dashboard/analytics.js +5 -3
  30. package/dist/goals/command.js +183 -53
  31. package/dist/goals/usage.js +87 -0
  32. package/dist/loop/cache-isolation.js +36 -0
  33. package/dist/loop/candidate-cleanup.js +47 -17
  34. package/dist/loop/candidates.js +17 -11
  35. package/dist/loop/dispatcher.js +89 -26
  36. package/dist/loop/failure.js +104 -0
  37. package/dist/loop/gate-snapshot.js +19 -0
  38. package/dist/loop/git.js +1 -1
  39. package/dist/loop/loop.js +124 -70
  40. package/dist/loop/parallel-adapters.js +57 -6
  41. package/dist/loop/parallel-command.js +49 -7
  42. package/dist/loop/proof-retention.js +70 -0
  43. package/dist/loop/recovery.js +23 -5
  44. package/dist/loop/reporter.js +22 -5
  45. package/dist/loop/run-command.js +101 -47
  46. package/dist/loop/runner.js +6 -5
  47. package/dist/loop/worker.js +152 -91
  48. package/dist/observability/history.js +2 -1
  49. package/dist/observability/invocation.js +42 -0
  50. package/dist/observability/local-report.js +120 -0
  51. package/dist/observability/usage.js +19 -0
  52. package/dist/prd/command.js +20 -7
  53. package/dist/prd/decompose.js +5 -2
  54. package/dist/retrofit/config.js +29 -2
  55. package/dist/retrofit/gitignore.js +12 -0
  56. package/dist/retrofit/planners/codex.js +20 -20
  57. package/dist/routing/attempts.js +241 -0
  58. package/dist/routing/capability.js +13 -9
  59. package/dist/routing/optimization.js +73 -0
  60. package/dist/routing/registry.js +7 -1
  61. package/dist/routing/router.js +282 -127
  62. package/dist/setup/command.js +8 -2
  63. package/dist/smoke/command.js +387 -85
  64. package/dist/update/check.js +1 -1
  65. package/docs/BENCHMARK-MANIFEST.md +131 -0
  66. package/docs/CODE-INTELLIGENCE.md +43 -1
  67. package/docs/CODEX-COMPARISON-2026-09-29.md +15 -0
  68. package/docs/DELIVERY-JOURNEYS.md +206 -0
  69. package/docs/ECONOMIC-ROUTING.md +180 -0
  70. package/docs/GOALS.md +61 -4
  71. package/docs/RELEASE-VALIDATION-1.22.0.md +115 -0
  72. package/docs/RELEASE-VALIDATION-1.23.0.md +39 -0
  73. package/docs/benchmarks/2026-10-04-efficiency/ANALYSE.md +182 -0
  74. package/docs/benchmarks/2026-10-04-efficiency/compare-help.py +55 -0
  75. package/docs/benchmarks/2026-10-04-efficiency/manifest.json +125 -0
  76. package/docs/benchmarks/2026-10-04-efficiency/provenance-analysis.json +90 -0
  77. package/docs/benchmarks/2026-10-04-efficiency/provenance-design.json +90 -0
  78. package/docs/benchmarks/2026-10-04-efficiency/provenance-original-report.json +90 -0
  79. package/docs/benchmarks/2026-10-04-efficiency/raw/DEVELOPMENT_ANALYSIS.md +142 -0
  80. package/docs/benchmarks/2026-10-04-efficiency/raw/RESULT.md +21 -0
  81. package/docs/benchmarks/2026-10-04-efficiency/raw/commands.jsonl +26 -0
  82. package/docs/benchmarks/2026-10-04-efficiency/raw/environment.json +31 -0
  83. package/docs/benchmarks/2026-10-04-efficiency/raw/final-yoke-smoke.json +40 -0
  84. package/docs/benchmarks/2026-10-04-efficiency/raw/model-purpose-hints.csv +19 -0
  85. package/docs/benchmarks/2026-10-04-efficiency/raw/observations.jsonl +21 -0
  86. package/docs/benchmarks/2026-10-04-efficiency/raw/observer-command-phases.csv +12 -0
  87. package/docs/benchmarks/2026-10-04-efficiency/raw/roles.csv +5 -0
  88. package/docs/benchmarks/2026-10-04-efficiency/raw/shell-categories.csv +8 -0
  89. package/docs/benchmarks/2026-10-04-efficiency/raw/stories.csv +8 -0
  90. package/docs/benchmarks/2026-10-04-efficiency/raw/summary.json +469 -0
  91. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-history.jsonl +104 -0
  92. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-1.log +58 -0
  93. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-2.log +29 -0
  94. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-3.log +5 -0
  95. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-4.log +12 -0
  96. package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-phases.csv +10 -0
  97. package/docs/benchmarks/2026-10-04-efficiency/regression-comparison.json +104 -0
  98. package/docs/parallel-execution.md +37 -9
  99. package/docs/superpowers/plans/2026-10-04-yoke-1.23-efficiency-prd.json +11 -0
  100. package/docs/superpowers/plans/2026-10-04-yoke-1.23-efficiency.md +83 -0
  101. package/docs/superpowers/specs/2026-10-04-yoke-1.23-efficiency-design.md +120 -0
  102. package/gemini-extension.json +1 -1
  103. package/package.json +1 -1
@@ -28,9 +28,51 @@ Install and pin these tools in the project or developer environment before enabl
28
28
 
29
29
  - Every request is tied to a workspace and a content-addressed snapshot of committed and uncommitted, non-sensitive files.
30
30
  - Paths are workspace-relative; traversal, symlinks, `.env`/key material and policy-excluded paths are rejected.
31
- - Read calls are bounded by timeout and output budgets. Results carry backend, version, source path, content hash, resolution and freshness.
31
+ - Read calls share a request deadline and a complete-response output budget. Results distinguish backend identity, source location, semantic resolution and unverified index freshness. See the limits and evidence contracts below.
32
32
  - Edits run in a disposable Yoke sandbox. Preview writes a durable plan and diff but requires an approval record created by the Yoke control layer.
33
33
  - Apply rechecks the snapshot, plan expiry, plan hash, an exclusive transaction lock and idempotency. It produces a managed worktree and does not commit or merge; the existing Yoke loop and gates remain authoritative for integration.
34
34
  - Network access is not needed by the facade. Backend installation and any backend-specific network behavior remain an explicit environment concern.
35
35
 
36
36
  For a host that has no native MCP support (currently Pi), use the normal Yoke loop and bounded CLI/skill adapter. Yoke does not invent a second MCP transport for that host.
37
+
38
+ ## Evidence and trace contracts
39
+
40
+ `code_trace` treats Serena's incoming reference list as references to the requested symbol. Each validated referencing symbol has an edge **to the requested target**. Two adjacent results never imply an edge between those results. Incoming implementation lists similarly use the `implements` relation. A bare symbol name is resolved to a unique path-bound Serena identity before requesting these relations; ambiguous targets remain unresolved.
41
+
42
+ The pinned Graft trace interface returns text. Until an explicit edge format is validated, its output is retained as structural candidates, without invented call edges. Incoming references cannot establish outgoing calls, imports or inheritance: unsupported direction/relation combinations return partial results and explain the missing contract. `require_resolved` excludes unresolved structural candidates; it does not turn an unsupported query into a resolved graph. Semantic trace results currently cover one hop. Requests for greater depth are marked unverified instead of pretending traversal was completed.
43
+
44
+ `frontier_remaining` counts known entries omitted by result limits, and `traversal_complete` is false when exhaustive traversal is unverified. Zero known omissions does not prove that no other references exist. Node limits preserve graph endpoints: returned edges never point at removed nodes.
45
+
46
+ The workspace snapshot is content-addressed and an explicitly supplied snapshot is checked against the current workspace before processing. This check does **not** prove that Graft's index, Graphify's graph or the language server has processed those exact file contents. The supported wire contracts do not provide an independently validated index-to-snapshot binding. Consequently backend evidence uses `freshness: "unknown"` and `content_hash: null`; an untrusted backend field claiming `current` is not accepted as proof. The facade no longer stamps a current workspace file hash onto potentially older evidence.
47
+
48
+ Recognized, path-bound Serena symbol records can have `resolution: "resolved"` while freshness remains unknown. Unrecognized records remain unresolved and cannot create symbols or edges. Read responses use `status: "partial"` and partial coverage for available backends because index freshness and exhaustive coverage are not established. A successful process call alone does not imply complete coverage. Missing backends remain separately identified in `backends_missing`.
49
+
50
+ Each retained evidence link has a matching `provenance.evidence_id`. This additive field allows callers to resolve evidence IDs and lets output reduction remove unused provenance without breaking retained links. No response cache is used; no stale evidence is reused under a new snapshot or query.
51
+
52
+ ## Output and time limits
53
+
54
+ The facade forwards all four configured limits:
55
+
56
+ ```yaml
57
+ codeIntelligence:
58
+ mode: shadow
59
+ limits:
60
+ tokenBudget: 2400
61
+ timeoutMs: 10000
62
+ maxBytes: 2000000
63
+ maxBackends: 3
64
+ ```
65
+
66
+ The effective output cap is the smaller of `maxBytes` and the effective token budget. A request's `token_budget`, where supported, can lower the configured cap, but cannot raise it. The default configured token budget is 2400. The budget covers the entire serialized JSON response in MCP `structuredContent`: result data, provenance, warnings, identifiers, coverage, errors and the metrics themselves. The outer JSON-RPC framing and the short MCP text summary are outside that response.
67
+
68
+ **Token accounting is deliberately conservative:** one UTF-8 byte consumes one budget unit. `metrics.result_tokens` is therefore an estimate with `token_count_kind: "estimated"` and `token_count_method: "utf8_bytes_conservative"`. It is not provider-measured usage, a tokenizer-specific exact count or a billing estimate. `metrics.returned_bytes` is the actual UTF-8 size of the serialized structured response, including the metrics. This replaces the earlier four-bytes-per-token approximation and makes the same numeric budget stricter; increase the configured budget explicitly if a workflow requires larger context, rather than treating the old number as an exact token allowance.
69
+
70
+ When an answer does not fit, text is shortened at Unicode-safe boundaries and lower-ranked entries are removed. Arrays and operation receipt fields keep their schema, retained graph edges have valid endpoints, and unused provenance is removed. The answer reports `partial`, adds a truncation warning and never claims complete coverage. Narrow the query or raise the configured budget to obtain omitted content; `next_cursor: null` does not imply a complete search.
71
+
72
+ A budget can be smaller than the mandatory response envelope. For example, 128 budget units cannot encode all required identifiers and error fields. Such requests fail with `BUDGET_EXCEEDED` **before a backend is invoked**. The small, fixed-shape error envelope is the only output-cap exception; its size is reported accurately, and it contains no backend payload or unbounded error text. Apply additionally reserves space for its entire concrete receipt, including all changed paths and a possible deadline warning, before starting a transaction. An insufficient receipt budget rejects the apply without mutating. A completed apply keeps its transaction receipt even if synchronous work crossed the deadline.
73
+
74
+ The effective deadline is the smaller of configured `timeoutMs` (10000 ms by default) and the request's `timeout_ms`/tool default. All backend calls, semantic target resolution and preview operations consume the same remaining time. MCP initialization and its subsequent tool call also share that allowance. A backend which ignores its timeout is closed, late results are discarded, and no subsequent backend starts after expiry. Backend shutdown has a separate bounded cleanup grace; synchronous filesystem snapshot/transaction operations cannot be preempted mid-operation, so the deadline is not an exact end-to-end wall-clock promise. A final deadline check after normalization and output reduction marks delayed results partial and keeps any completed operation's receipt. A preview needing the former 120000 ms tool default must also have a configured timeout of at least that size.
75
+
76
+ ## Validation scope
77
+
78
+ The regression suite uses model-free adapters and local synthetic MCP processes. It verifies reference direction and endpoints, unknown payload handling, honest freshness and coverage, aggregate output limits, Unicode boundaries, evidence link preservation, tiny-budget failures, limit propagation, and shared deadlines including initialization. It does not assert a successful run against installed authenticated backends or an exact provider-token saving.
@@ -25,3 +25,18 @@ See the [Goals/resource audit](GOALS-RESOURCE-AUDIT-2026-09-29.md) for implement
25
25
  Reproduce with `bench/compare-codex.mjs`, then `bench/analyze-codex-comparison.mjs`. Use fresh, short output roots on Windows. The new summary tests verify median arithmetic, cache subtraction, unknown usage, incompatible policies and invalid measurements.
26
26
 
27
27
  Provenance: this report is disclosed as AI-assisted. Read-only text scans cannot establish human authorship or verify proprietary keyed watermarks; cryptographic verification and signer trust remain unknown without the corresponding verifier and trust policy.
28
+
29
+ ## Later tooling update — 1.22.0
30
+
31
+ The measurements above remain the historical 1.19.0 results. The current analyzer
32
+ reads their aggregate JSON as legacy/unverified because those runs did not record
33
+ the new complete comparison manifest. It preserves the reported values and does
34
+ not manufacture missing provenance or reinterpret them as 1.22.0 measurements.
35
+
36
+ New runs record source/build, fixture, acceptance and submitted-input digests,
37
+ requested and reported model identities, and explicit startup conditions.
38
+ The historical direct-arm ignore-rules difference is now an expressly declared
39
+ workflow variable. Undeclared cross-arm differences, missing provenance and
40
+ failed acceptance do not produce performance comparison groups.
41
+ See [benchmark manifests](BENCHMARK-MANIFEST.md) for the contract and limitations.
42
+ This is a tooling change; no new authenticated comparison was performed for it.
@@ -0,0 +1,206 @@
1
+ # Delivery artifacts and executable user journeys
2
+
3
+ Use executable acceptance criteria to decide whether a release is ready. A build command shows that an artifact can be produced; a journey checks a specific user interaction. Yoke can now record the relationship between those criteria, named journeys and the exact artifact files present during a check.
4
+
5
+ The delivery declaration belongs in `.yoke/acceptance.yaml`. Browser smoke steps belong in `.yoke/config.yaml`. Existing flow definitions containing only `name`, `path` and an optional `landmark` remain valid, but production browser runs now also require `smoke.sourceIdentity`.
6
+
7
+ ## Optional browser steps
8
+
9
+ `yoke flow-smoke` uses the target project's Playwright installation and its Chromium browser. Start the application or preview server before running the command. Yoke does not start a development server or install a browser automatically.
10
+
11
+ For example, add this section to the project's existing `.yoke/config.yaml`:
12
+
13
+ ```yaml
14
+ smoke:
15
+ baseUrl: http://localhost:3000
16
+ sourceIdentity:
17
+ path: /assets/app-specific-static-source.js
18
+ sha256: "<replace with SHA-256 of the served bytes>"
19
+ flows:
20
+ - name: profile-survives-reload
21
+ path: /login
22
+ landmark: main
23
+ timeoutMs: 60000
24
+ steps:
25
+ - action: fill
26
+ selector: '[data-testid="email"]'
27
+ value: smoke-user@example.invalid
28
+ - action: fill
29
+ selector: '[data-testid="password"]'
30
+ valueEnv: SMOKE_TEST_PASSWORD
31
+ - action: click
32
+ selector: '[data-testid="sign-in"]'
33
+ - action: expect-url
34
+ url: /profile
35
+ - action: fill
36
+ selector: '[data-testid="display-name"]'
37
+ value: Smoke User
38
+ - action: press
39
+ selector: '[data-testid="display-name"]'
40
+ key: Tab
41
+ - action: click
42
+ selector: '[data-testid="save-profile"]'
43
+ - action: expect-visible
44
+ selector: '[data-testid="save-confirmation"]'
45
+ - action: expect-text
46
+ selector: '[data-testid="save-confirmation"]'
47
+ text: Profile saved
48
+ exact: true
49
+ timeoutMs: 10000
50
+ - action: reload
51
+ - action: expect-text
52
+ selector: '[data-testid="profile-name"]'
53
+ text: Smoke User
54
+ exact: true
55
+ ```
56
+
57
+ Replace the source resource path and hash placeholder before running this example. Confirm the intended server's checkout/build and port, then fetch a stable, app-specific source resource from the effective `baseUrl` origin and calculate the SHA-256 of the exact served body bytes. Dev servers may transform local source, so hashing a local file alone is insufficient. Set `sha256` to the resulting 64-character lowercase hexadecimal digest. The resource must return a successful HTTP response without redirects, finish within five seconds and contain at most 1 MiB. It must remain stable through the run. Production smoke checks the pin before browser launch and again after the flows; a missing or mismatching pin prevents valid smoke evidence.
58
+
59
+ After an intentional source/build change, verify the server serves that change and explicitly refresh the pin. Do not update it automatically on a mismatch or accept a foreign process occupying the expected port. Choose a resource that changes with the relevant app source, rather than a shared health response or an unchanged marker. The static pin checks only that resource's bytes before and after the run; it does not independently prove the identity of every module, backend or dynamic response. Source fingerprints and other acceptance checks remain necessary.
60
+
61
+ Supply `SMOKE_TEST_PASSWORD` through the environment using credentials for a dedicated test account, then run:
62
+
63
+ ```sh
64
+ yoke flow-smoke . --label=profile-release
65
+ ```
66
+
67
+ `--url=http://localhost:4173` overrides `smoke.baseUrl`. Flow navigation keeps the existing `baseUrl + path` behavior, so configure their slashes consistently. Each flow gets a fresh browser context with a 1280 × 720 viewport. Steps within one flow share that context; different flows do not share their browser session.
68
+
69
+ ### Supported steps
70
+
71
+ | Action | Fields | Behavior |
72
+ | --- | --- | --- |
73
+ | `click` | `selector` | Click the matching element using Playwright's action waiting. |
74
+ | `fill` | `selector`, exactly one of `value` or `valueEnv` | Fill a form control with a literal value or a captured environment value. Missing environment values fail the step. |
75
+ | `press` | `selector`, `key` | Press a key, such as `Enter` or `Tab`, on the matching element. |
76
+ | `expect-visible` | `selector` | Wait for the matching element to be visible. |
77
+ | `expect-text` | `selector`, `text`, optional `exact` | Wait for visible text. Whitespace is normalized. Matching is case sensitive and uses a literal substring by default; `exact: true` compares the complete normalized text. |
78
+ | `expect-url` | `url` | Wait for the exact resolved URL. A relative URL resolves against the effective base URL. This is not a glob or regular expression. |
79
+ | `reload` | No additional action fields | Reload the page, wait for its load event and reject a non-successful HTTP response. |
80
+
81
+ Each step may specify `timeoutMs`. Unknown actions and extra action fields are rejected before browser launch. There is no arbitrary JavaScript action, general scripting language, branch or loop. Use a project E2E test command for more complex behavior, multiple pages, downloads, native applications or specialized fixtures.
82
+
83
+ ### Limits and failure behavior
84
+
85
+ A flow may contain **1–50 steps**. Its default total timeout is **60,000 ms**, with a configurable maximum of **120,000 ms**. This budget includes browser-context preparation, navigation, the optional landmark and all steps. Navigation has its own maximum of 30,000 ms. A landmark has a maximum of 10,000 ms.
86
+
87
+ A step defaults to **10,000 ms** and may request at most **30,000 ms**. Its actual allowance is capped by the remaining flow budget. Text waiting has both an attempt limit and a deadline. Screenshot capture, context closing and video finalization each have a separate 5,000 ms allowance after interaction ends; browser startup and shutdown also have bounded waits. The flow timeout is therefore not a promise that the whole multi-flow command, including evidence capture and cleanup, ends within that duration.
88
+
89
+ The first failed step stops the remaining steps of that flow. The report marks them `skipped`. Later flows still run. Non-successful navigation, page errors and error-level console events also fail a flow. A screenshot is required for a passing result. Failure video remains best-effort evidence and does not replace a missing screenshot.
90
+
91
+ Fill values are limited to 8,192 characters, including values read from the environment. Environment-backed fill values are captured once when the run begins.
92
+
93
+ ## Evidence from a smoke run
94
+
95
+ Evidence is written to `.yoke/proof/<label>/`:
96
+
97
+ - A screenshot for each flow where screenshot capture succeeds.
98
+ - A `.webm` video for a failed flow when Playwright can finalize it.
99
+ - `report.json`, a versioned machine-readable report.
100
+
101
+ The default label is `YOKE_STORY` when present, otherwise `latest`; `--label` takes priority. Labels and filenames are sanitized, and colliding sanitized flow names receive distinct filenames. Running the same label replaces its previous evidence. A run that cannot start because its configuration, browser, source identity or proof location is unavailable does not erase the prior evidence during that preflight.
102
+
103
+ The report contains:
104
+
105
+ - Per-flow status, navigation status, optional landmark status and duration.
106
+ - Each configured step's index, action, status, duration and a bounded failure category.
107
+ - Relative screenshot and failure-video filenames.
108
+ - Source fingerprints before and after the run, and whether they match.
109
+ - A digest of the effective smoke configuration, URL override and captured fill inputs.
110
+ - Node version, operating-system platform, architecture, configured viewport, browser version when available and the effective URL's origin.
111
+
112
+ Raw configuration, selectors, expected text, fill values, environment-variable names and browser exception call logs are not serialized into the report. URL credentials, query strings and fragments are not included in its recorded origin. Driver errors are represented by categories such as `timeout`, `step-failed` or `navigation-failed`, because their original messages can contain form values. Screenshots and video still show what the application visibly displays; use test data appropriate for those visual artifacts.
113
+
114
+ Exit code `0` means all flows passed, their screenshots were captured, the source fingerprints before and after the run matched and no browser-cleanup failure was reported. Exit code `1` represents a failed run or invalidated evidence. Exit code `2` covers unavailable prerequisites or evidence storage. Invalid step configuration is rejected before browser launch.
115
+
116
+ The local source fingerprint does **not** prove that a remote server is running that source revision. Arrange the project's preview or deployment checks so the browser exercises the intended build. Record a separate deployment verification criterion when that relationship matters.
117
+
118
+ ## Bind release artifacts and journeys to acceptance criteria
119
+
120
+ Add an optional `delivery` section to `.yoke/acceptance.yaml`:
121
+
122
+ ```yaml
123
+ version: 1
124
+ protected:
125
+ - .yoke/config.yaml
126
+ - playwright.config.ts
127
+ - tests/e2e/profile.spec.ts
128
+ criteria:
129
+ - id: profile-e2e
130
+ text: A signed-in user can save a profile and recover it after reloading.
131
+ commands:
132
+ - npm run test:e2e -- tests/e2e/profile.spec.ts
133
+ - id: profile-smoke
134
+ text: The release preview completes the configured profile journey.
135
+ commands:
136
+ - yoke flow-smoke . --label=profile-release
137
+ delivery:
138
+ version: 1
139
+ artifacts:
140
+ - path: dist/release.zip
141
+ criteria: [profile-e2e, profile-smoke]
142
+ journeys:
143
+ - id: profile-update
144
+ text: Sign in, change a profile field, save it, reload and observe the saved value.
145
+ criteria: [profile-e2e, profile-smoke]
146
+ environment:
147
+ name: Local release preview with Chromium
148
+ url: http://localhost:3000
149
+ ```
150
+
151
+ The script names, test files and artifact path in this example must exist in the application project. `delivery.journeys` names the intended user experience and references executable criteria; it does not execute browser steps by itself. The example executes smoke steps through the `profile-smoke` criterion's command.
152
+
153
+ **Build declared artifacts before starting `yoke check`:**
154
+
155
+ ```sh
156
+ npm run build:release
157
+ # Start the project's preview server for that release in another terminal.
158
+ yoke check . --json
159
+ ```
160
+
161
+ `build:release` here denotes the project's build script that creates `dist/release.zip`. A criterion that creates the declared artifact during `yoke check` is too late: the artifact must already exist for its initial snapshot.
162
+
163
+ The check report records artifact SHA-256 hashes and sizes before and after the commands. An artifact missing at either snapshot, differing between snapshots or using an unsafe path fails artifact binding. Declared artifacts must be regular project-relative files without symbolic links or path escapes. Explicit artifact hashing also covers ordinary ignored build outputs.
164
+
165
+ Artifact hashing accepts at most **512 MiB per file**. Each snapshot, before or after verification, has a combined **1 GiB byte budget** and a **10-second deadline**; an earlier enclosing check deadline also applies. Exceeding a limit or cancelling the check prevents a passing binding. Data is streamed in bounded chunks, with deadline and cancellation checks between synchronous reads. An individual blocked filesystem read is not preempted by that deadline. The declaration accepts at most 32 artifacts and 100 named journeys.
166
+
167
+ Artifact and journey status comes from their referenced criteria. All referenced criteria must pass for a passing status; a failed criterion fails the binding, while an unexecuted criterion remains unverified. Source-integrity, cancellation or artifact-binding problems also prevent a passing delivery result. Unknown criterion IDs and duplicate artifact paths or journey IDs are rejected.
168
+
169
+ The report labels this relationship **`binding: declared-criteria`**. Matching artifact hashes plus green criteria establish the declared relationship and the same artifact content at both snapshots. This is not continuous monitoring of the file between snapshots. The project command must actually exercise the intended artifact; Yoke cannot infer that a test which ignores its APK, ZIP or binary argument has verified that file.
170
+
171
+ Protect the acceptance manifest and the test infrastructure with the existing `yoke check . --protect` workflow after reviewing the contract. See [Verified projects](VERIFIED-PROJECTS.md) for protection and intentional-baseline-refresh behavior.
172
+
173
+ ## Android example: verify the built APK on a specified device
174
+
175
+ A browser journey does not install an APK or validate Android permission handling. Use an application-owned executable test for that task and reference it from acceptance:
176
+
177
+ ```yaml
178
+ version: 1
179
+ protected:
180
+ - tools/check-apk-install.mjs
181
+ - tests/android/tracking-fixture.json
182
+ criteria:
183
+ - id: tracking-installed-apk
184
+ text: The declared APK installs and completes the tracking fixture on the selected emulator.
185
+ commands:
186
+ - node tools/check-apk-install.mjs --apk=android/app/build/outputs/apk/debug/app-debug.apk --serial=emulator-5554
187
+ delivery:
188
+ version: 1
189
+ artifacts:
190
+ - path: android/app/build/outputs/apk/debug/app-debug.apk
191
+ criteria: [tracking-installed-apk]
192
+ journeys:
193
+ - id: record-and-reopen
194
+ text: Grant the required permission, start tracking, receive fixture GPS points, stop, save and reopen the route after restarting the app.
195
+ criteria: [tracking-installed-apk]
196
+ environment:
197
+ name: Android emulator emulator-5554 with the project's configured system image
198
+ ```
199
+
200
+ `check-apk-install.mjs` and the fixture are **project-supplied tests**, not bundled Yoke commands. The script must install that exact APK, select and verify its device, perform the required permission and tracking actions, inspect the resulting route and process logs, and exit nonzero when an assertion fails. Build the APK first, prepare the emulator, then run `yoke check`.
201
+
202
+ The declared environment name is metadata. It does not detect the actual emulator image, phone model or OS version. Have the project test verify those properties when they are acceptance requirements. Passing on one emulator does not establish correct background tracking on a particular Vivo phone, and passing against one deployment does not certify every deployment.
203
+
204
+ ## Validation scope
205
+
206
+ The smoke engine's repository tests use filesystem-backed browser doubles to verify step order, failure handling, timeout budgets, redaction, source/configuration binding and evidence persistence. They do not install browsers or certify an application. Application-level evidence comes from successfully running the configured commands with the project's actual browser, build, server or device.
@@ -0,0 +1,180 @@
1
+ # Measured capability routing
2
+
3
+ Yoke 1.22 adds optional economic selection to capability routing. It compares
4
+ observed, independently verified execution sequences. It does not assign an
5
+ estimated success probability to a model or assume that a profile labelled
6
+ `low` has the lowest cost of completing a task.
7
+
8
+ ## Configuration and compatibility
9
+
10
+ ```yaml
11
+ routing:
12
+ enabled: true
13
+ strategy: capability
14
+ optimization:
15
+ version: 1
16
+ objective: balanced
17
+ minSamples: 20
18
+ ```
19
+
20
+ This fragment supplements the project's existing routing profiles and limits.
21
+ `minSamples` defaults to 20 and has a minimum of 10. Existing configurations
22
+ without `routing.optimization` retain their previous capability ordering. New
23
+ setup-generated capability configurations enable the conservative `balanced`
24
+ policy; this does not change selection until sufficient comparable measurements
25
+ exist.
26
+
27
+ The planner's task/risk assessment still determines the minimum capability tier.
28
+ Configured provider constraints, allowed roles, `maxTier`, parent-fallback policy,
29
+ and the existing gate-failure reliability filter remain in force. Economic
30
+ selection only compares profiles that survive those checks. Explicit project
31
+ routing rules retain their precedence.
32
+
33
+ ## What a sample measures
34
+
35
+ The comparison unit is a bounded **execution sequence for one task contract**,
36
+ attributed to the profile that started it. If a small model needs a targeted
37
+ repair and then escalation to a stronger model, all those measured attempts
38
+ belong to the initial profile's sequence. Charging only the final successful
39
+ model would hide the cost of the failed attempts.
40
+
41
+ Within a fully accounted execution attempt, the durable call ledger includes:
42
+
43
+ - The routing/planning call made inside that execution, when one was needed.
44
+ - The implementation call, including reported usage from interrupted calls.
45
+ - Additional reviewer, critic, repair and model-driven quality calls joined by
46
+ the loop reporter before its final outcome.
47
+
48
+ Stable call IDs prevent a worker's usage from being counted again when its
49
+ aggregate reaches the reporter. An explicit attempt ID can join a call directly.
50
+ Without one, a role call is joined only if that story has exactly one open
51
+ attempt. Competing candidates make the attribution ambiguous; all affected
52
+ samples become incomplete instead of assigning the bill by guesswork.
53
+
54
+ The measured scope is `execution-attempt`. It is not the total cost of producing
55
+ or shipping a product. Separately prepared PRD drafts, batch assessments, change
56
+ planning, human work and later operational costs are not amortized into these
57
+ samples. Their separately emitted telemetry remains useful for project-level
58
+ analysis. The duration covers each admitted attempt through its recorded gate
59
+ outcome, summed across the sequence; it does not include idle time between
60
+ separate attempts or a developer's review time.
61
+
62
+ Only native loop paths with the complete role-accounting contract enable this
63
+ scope. Injected runners/gates and ambiguous candidate races remain `worker`
64
+ scope. Their measurements are still available diagnostically, but cannot supply
65
+ the complete evidence required for economic selection.
66
+
67
+ ## Evidence and missing data
68
+
69
+ For each eligible starting profile, the router examines its most recent
70
+ `minSamples` comparable sequences. They must share the project, implementation
71
+ role, task-assessment dimensions, effective execution-policy key, requested
72
+ provider/model/effort/variant, and the current reported concrete starting model.
73
+ The policy key binds the effective gate and review configuration, including
74
+ commands and retries. A change to the concrete model starts a new evidence
75
+ window even when its requested alias remains the same.
76
+
77
+ A sequence is complete only when every reserved attempt has an outcome and the
78
+ sequence has either passed independent gates or exhausted its bounded attempt
79
+ allowance. Every attempt must have known token coverage, complete reported cost,
80
+ the full execution scope, and the same policy key. Infrastructure failures are
81
+ not treated as evidence of model reasoning quality and do not qualify for the
82
+ economic comparison.
83
+
84
+ One recent unknown or partial sequence keeps its window ineligible. The router
85
+ does not discard that row and search for an older successful subset. Open
86
+ sequence records are replaced by their later completion record, and retries do
87
+ not turn into additional independent samples. The optional observation registry
88
+ uses a 30-day evidence window and reads at most its latest 1,000 events; losing
89
+ economic history leads back to conservative selection.
90
+
91
+ Unknown usage is distinct from measured zero. Known partial token or dollar
92
+ amounts are retained with incomplete-coverage flags. For example, a fully
93
+ measured worker remains visible when its planner supplied no usage. OpenCode
94
+ and Kilo step summaries require coverage of every completed step before they
95
+ claim complete token totals; absent optional cache or cost fields are not
96
+ converted into measured zeros.
97
+
98
+ No reported dollars means no cost ranking. Yoke does not invent a bill from
99
+ declared cost tiers, token counts or a presumed provider price table. This can
100
+ leave economic routing at its conservative baseline for providers that do not
101
+ report sufficient telemetry.
102
+
103
+ ## The three objectives
104
+
105
+ The baseline is the first eligible profile under the existing tier, declared
106
+ cost-tier and profile-ID ordering. Both the baseline and an alternative require
107
+ a complete evidence window. An alternative must have at least as many accepted
108
+ sequences in that equally sized window as the baseline.
109
+
110
+ The router computes two observed ratios:
111
+
112
+ ```text
113
+ cost per accepted sequence = total measured sequence dollars / accepted sequences
114
+ time per accepted sequence = total measured sequence duration / accepted sequences
115
+ ```
116
+
117
+ The numerators include completed failed sequences and every measured repair or
118
+ escalation in the window. A profile with zero accepted sequences is ineligible.
119
+
120
+ | Objective | Condition for choosing an alternative |
121
+ | --- | --- |
122
+ | `cost` | Lower observed dollars per accepted sequence, with no lower observed acceptance count. |
123
+ | `speed` | Lower observed duration per accepted sequence, with no lower observed acceptance count. |
124
+ | `balanced` | No higher cost or duration, and at least one strictly lower, with no lower observed acceptance count. |
125
+
126
+ `balanced` uses this conservative comparison instead of inventing a dollar
127
+ value for a second of runtime. Cost and speed can disagree; use the corresponding
128
+ explicit objective if that tradeoff is intentional. Equal scores preserve the
129
+ existing ordering.
130
+
131
+ Decision reasons identify the objective, sample count, observed acceptance
132
+ count, and measured ratios. When the baseline lacks enough data, they explain
133
+ the missing evidence and retain its existing ordering. These are observational
134
+ measurements, not calibrated forecasts or statistical confidence guarantees.
135
+ Task mix, repeated correlated failures and changing model behavior can still
136
+ affect the comparison. The minimum window is a conservative product policy,
137
+ not a claim that twenty examples prove superiority.
138
+
139
+ Reviewer and critic approval frequency is not independent evidence that their
140
+ reviews were good. Economic reordering of those roles therefore remains disabled
141
+ until an independent role-quality measure is available; their risk-based floors
142
+ and explicit model selections continue to apply.
143
+
144
+ ## Persistent attempt admission
145
+
146
+ Capability routing reserves an attempt before invoking its implementation
147
+ provider. Configured bounded legacy routes use the same account. The account
148
+ is stored under the stable project root, keyed by story and bound task/plan
149
+ contract, and is independent of the optional global observation registry.
150
+ Exclusive immutable reservation files prevent two processes from replacing the
151
+ same slot. An unreadable or unwritable authoritative account blocks admission.
152
+
153
+ The capability allowance is the smaller of `maxAttempts` and the existing
154
+ tier-dependent bound: five attempts from `light`, four from `standard`, three
155
+ from `strong`, and two from `frontier`. Independent outer loop, goal and local
156
+ retry limits can stop execution sooner. Infrastructure or interrupted attempts
157
+ still occupy their reserved slots, but do not become verified reasoning failures
158
+ that automatically demand a stronger model.
159
+
160
+ Deleting or filling the optional registry, restarting a runner, or executing in
161
+ another disposable worktree does not restore those slots. Repeated preparation
162
+ or budget refusals before an implementation is admitted do not consume an
163
+ implementation attempt. A planner may itself incur cost before a later worker
164
+ is refused; that paid call is still reported. Goal callers recheck their durable
165
+ budget after planning and before implementation reservation.
166
+
167
+ A changed task/plan contract has a new account. Resume an unchanged blocked task
168
+ only after inspecting its reason and performing the required replan or explicit
169
+ configuration change. Do not delete authoritative attempt state to disguise
170
+ retries as a new run.
171
+
172
+ ## Validation limits
173
+
174
+ The release tests use deterministic fake calls and synthetic observations. They
175
+ cover registry write failure and history eviction, restart-safe reservations,
176
+ known partial usage, synchronous/asynchronous provider telemetry, paid early
177
+ blocks, callback/write failures, complete escalation accounting, cold starts,
178
+ model/policy changes and objective selection. They are regression tests for the
179
+ accounting and decision rules, not provider performance benchmarks. No claimed
180
+ cost or speed improvement follows from those test fixtures.
package/docs/GOALS.md CHANGED
@@ -29,11 +29,44 @@ IDs and executable commands and binds the objective to a digest of the acceptanc
29
29
  manifest. It cannot determine whether your tests fully express a free-text request.
30
30
  All project acceptance and configured regression checks must still pass.
31
31
 
32
+ Each independent check writes its full report, including declared artifact and
33
+ journey evidence. Goal state keeps the check ID in `lastCheck` and its report path
34
+ in `lastCheckEvidencePath`; `goal handoff` includes both. The report binds its
35
+ findings to the checked workspace. Declaring an artifact or journey is not itself
36
+ proof that a criterion passed.
37
+
32
38
  Existing goals without a binding now stop before any model call. Review their tests,
33
39
  then explicitly bind them with `yoke goal bind . --criteria=checkout`. Attempts and
34
40
  measured consumption remain intact. Changed acceptance infrastructure still requires
35
41
  an explicit protection refresh after review; binding never refreshes protected tests.
36
42
 
43
+ ## Prepare routing explicitly
44
+
45
+ Capability routing uses the same goal ID, objective and structured executable
46
+ criteria as goal acceptance. It follows the project's `routing.assessmentPolicy`,
47
+ explicit `routing.rules`, profile limits and optional optimization settings.
48
+
49
+ With `routing.assessmentPolicy: prepared`, first prepare the bound goal:
50
+
51
+ ```sh
52
+ yoke goal assess . --runner=codex
53
+ yoke goal run . --runner=codex
54
+ ```
55
+
56
+ `goal assess` makes at most one read-only planning call and saves its assessment.
57
+ It reuses an existing assessment for the current contract. It does not implement
58
+ the objective or create an implementation attempt. Preparation consumes the same
59
+ durable token and provider-time budgets as execution. The planning provider and
60
+ model follow the project's planning configuration.
61
+
62
+ The following run reuses that assessment. A missing or stale prepared assessment
63
+ blocks execution before any model call, including when an explicit routing rule
64
+ exists. Changes to the objective, executable criteria, selected runner or approved
65
+ `.yoke/plan.md` invalidate the contract key. Run `goal assess` again after reviewing
66
+ those changes; acceptance changes still require explicit binding and protection
67
+ review. The `on-demand` policy may assess during execution; explicit rules avoid
68
+ the selection-planner call under that policy.
69
+
37
70
  ## Resource and budget controls
38
71
 
39
72
  Goals and story loops share the project lock and the user-level worker pool. Routing
@@ -41,24 +74,45 @@ planner and implementation calls each acquire a permit; native model delegation
41
74
  disabled. Explicit runner/model/effort/bare options override project runner defaults.
42
75
  Execution uses Yoke's safe permission profile.
43
76
 
77
+ Runner defaults apply only to the provider that owns them. For example,
78
+ `--runner=claude` does not inherit a Codex model, effort, provider, variant or bare
79
+ setting from `runner`. Explicit command options still apply; when keeping the same
80
+ runner, omitted options retain its configured defaults.
81
+
44
82
  | Setting | Meaning |
45
83
  | --- | --- |
46
84
  | `--attempts=N` on `set` or `budget` | Maximum cumulative attempts, default 3 |
47
85
  | `--minutes=N` | Cumulative admitted provider work, default 30 minutes; excludes capacity waits and independent checks |
48
86
  | `--wall-minutes=N` | Optional cumulative run time including admission, checks and state synchronization |
49
- | `--tokens=N` | Cumulative measured input + output tokens; unknown consumption stops further budgeted work |
87
+ | `--tokens=N` | Cumulative measured input + output tokens, checked before every planning and implementation call; unknown consumption stops further budgeted work |
50
88
  | `YOKE_MAX_PARALLEL_WORKERS=1..8` | Shared model/integration permit ceiling, default 3 |
51
89
  | `YOKE_MAX_PARALLEL_CHECKS=1..8` | Separate shared asynchronous goal verification ceiling, default 1 |
52
90
 
91
+ Planner consumption is persisted before admitting its implementation worker. If
92
+ planning reaches the ceiling, overruns it, or leaves usage unknown, no next model
93
+ call is admitted. A completed objective exactly at its measured ceiling may finish;
94
+ the same balance cannot authorize another model call. Native Codex receives only
95
+ the remaining token allowance after earlier planning and implementation calls.
96
+
53
97
  After a measured token overrun, Yoke retains evidence and work and reports `blocked`
54
98
  even if acceptance passed. Update the budget explicitly with `yoke goal budget .
55
99
  --tokens=50000`; `--clear-token-budget` explicitly removes that ceiling. Providers
56
- reporting usage only after a call cannot provide a hard mid-call token cap. Native
100
+ reporting usage only after a call cannot provide a hard mid-call token cap. This is
101
+ a call-admission limit plus cancellation where telemetry permits it. Native
57
102
  Codex usage units and Yoke's token accounting are recorded separately. Native
58
103
  stream updates use cumulative turn deltas and request cancellation when measured
59
104
  consumption exceeds the goal ceiling. Notifications arrive after model requests,
60
105
  so this can still overshoot within a request; progress updates are never added
61
- twice to persisted final usage.
106
+ twice to persisted final usage. Partial counts remain a known lower bound, with
107
+ incomplete measurement explicitly recorded. Raising a token ceiling does not make
108
+ unknown consumption known; explicitly removing the ceiling is a separate decision.
109
+
110
+ The goal's `usageLedger` records each call before dispatch, then records its final
111
+ usage and duration. Recovery retains finished planner calls and marks unfinished
112
+ calls as interrupted with unknown final usage. Each call has a stable ID and its
113
+ own usage event; attempt summaries are not emitted again as additional token usage.
114
+ Earlier goal files retain their existing attempt totals once, and new calls are
115
+ charged separately. Preparation and resume never reset the previous consumption.
62
116
 
63
117
  The asynchronous goal checker can interrupt its own command tree at a deadline or
64
118
  pause. Synchronous `yoke check` and story gates retain their existing execution
@@ -78,7 +132,10 @@ Or set `goals.nativeCodex: true` in `.yoke/config.yaml`. The default is off;
78
132
  Yoke negotiates the local Codex app-server goal capability, stores the thread binding,
79
133
  pauses native automatic continuation, and explicitly starts one bounded development
80
134
  turn. The same thread resumes on later attempts. Only Yoke's independent acceptance
81
- can synchronize it to `complete`. Provider/model changes create a new binding.
135
+ can synchronize it to `complete`. Provider/model changes detach the previous binding
136
+ and retain it in `detachedNativeBindings`. A later Codex execution creates a binding
137
+ for the selected model; a goal finished by another provider does not need the old
138
+ Codex runtime to synchronize a historical thread.
82
139
  An unsupported native goal method falls back to ordinary Codex execution; authentication,
83
140
  transport and execution failures remain visible. No global Codex configuration changes
84
141
  are needed. Support depends on the installed CLI, not on the model name alone.