@hecer/yoke 1.21.1 → 1.23.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/.codex-plugin/plugin.json +1 -1
- package/CHANGELOG.md +48 -0
- package/README.md +8 -1
- package/TODOS.md +6 -0
- package/bench/analyze-codex-comparison.mjs +90 -17
- package/bench/compare-codex.mjs +159 -36
- package/bench/result-schema.mjs +132 -0
- package/canon/manifest.yaml +1 -1
- package/canon/skills/visual-verification/SKILL.md +25 -2
- package/canon/tools/codex-rtk-hook.mjs +6 -16
- package/dist/agents/pi-telemetry.js +2 -1
- package/dist/agents/process-streams.js +12 -64
- package/dist/agents/provider-selection.js +12 -0
- package/dist/agents/telemetry.js +52 -52
- package/dist/change/inbox.js +8 -3
- package/dist/check/command.js +69 -17
- package/dist/check/delivery.js +121 -0
- package/dist/cli.js +91 -3
- package/dist/code-intelligence/adapters/mcp.js +1 -0
- package/dist/code-intelligence/budgets.js +138 -0
- package/dist/code-intelligence/contracts.js +2 -0
- package/dist/code-intelligence/coordinator.js +159 -85
- package/dist/code-intelligence/evidence.js +87 -34
- package/dist/code-intelligence/index.js +1 -0
- package/dist/code-intelligence/mcp-client.js +25 -6
- package/dist/code-intelligence/mcp-server.js +14 -11
- package/dist/code-intelligence/preflight.js +71 -0
- package/dist/dashboard/analytics.js +5 -3
- package/dist/goals/command.js +183 -53
- package/dist/goals/usage.js +87 -0
- package/dist/loop/cache-isolation.js +36 -0
- package/dist/loop/candidate-cleanup.js +47 -17
- package/dist/loop/candidates.js +17 -11
- package/dist/loop/dispatcher.js +89 -26
- package/dist/loop/failure.js +104 -0
- package/dist/loop/gate-snapshot.js +19 -0
- package/dist/loop/git.js +1 -1
- package/dist/loop/loop.js +124 -70
- package/dist/loop/parallel-adapters.js +57 -6
- package/dist/loop/parallel-command.js +49 -7
- package/dist/loop/proof-retention.js +70 -0
- package/dist/loop/recovery.js +23 -5
- package/dist/loop/reporter.js +22 -5
- package/dist/loop/run-command.js +101 -47
- package/dist/loop/runner.js +6 -5
- package/dist/loop/worker.js +152 -91
- package/dist/observability/history.js +2 -1
- package/dist/observability/invocation.js +42 -0
- package/dist/observability/local-report.js +120 -0
- package/dist/observability/usage.js +19 -0
- package/dist/prd/command.js +20 -7
- package/dist/prd/decompose.js +5 -2
- package/dist/retrofit/config.js +29 -2
- package/dist/retrofit/gitignore.js +12 -0
- package/dist/retrofit/planners/codex.js +20 -20
- package/dist/routing/attempts.js +241 -0
- package/dist/routing/capability.js +13 -9
- package/dist/routing/optimization.js +73 -0
- package/dist/routing/registry.js +7 -1
- package/dist/routing/router.js +282 -127
- package/dist/setup/command.js +8 -2
- package/dist/smoke/command.js +387 -85
- package/dist/update/check.js +1 -1
- package/docs/BENCHMARK-MANIFEST.md +131 -0
- package/docs/CODE-INTELLIGENCE.md +43 -1
- package/docs/CODEX-COMPARISON-2026-09-29.md +15 -0
- package/docs/DELIVERY-JOURNEYS.md +206 -0
- package/docs/ECONOMIC-ROUTING.md +180 -0
- package/docs/GOALS.md +61 -4
- package/docs/RELEASE-VALIDATION-1.22.0.md +115 -0
- package/docs/RELEASE-VALIDATION-1.23.0.md +39 -0
- package/docs/benchmarks/2026-10-04-efficiency/ANALYSE.md +182 -0
- package/docs/benchmarks/2026-10-04-efficiency/compare-help.py +55 -0
- package/docs/benchmarks/2026-10-04-efficiency/manifest.json +125 -0
- package/docs/benchmarks/2026-10-04-efficiency/provenance-analysis.json +90 -0
- package/docs/benchmarks/2026-10-04-efficiency/provenance-design.json +90 -0
- package/docs/benchmarks/2026-10-04-efficiency/provenance-original-report.json +90 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/DEVELOPMENT_ANALYSIS.md +142 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/RESULT.md +21 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/commands.jsonl +26 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/environment.json +31 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/final-yoke-smoke.json +40 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/model-purpose-hints.csv +19 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/observations.jsonl +21 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/observer-command-phases.csv +12 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/roles.csv +5 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/shell-categories.csv +8 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/stories.csv +8 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/summary.json +469 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-history.jsonl +104 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-1.log +58 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-2.log +29 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-3.log +5 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-loop-4.log +12 -0
- package/docs/benchmarks/2026-10-04-efficiency/raw/yoke-phases.csv +10 -0
- package/docs/benchmarks/2026-10-04-efficiency/regression-comparison.json +104 -0
- package/docs/parallel-execution.md +37 -9
- package/docs/superpowers/plans/2026-10-04-yoke-1.23-efficiency-prd.json +11 -0
- package/docs/superpowers/plans/2026-10-04-yoke-1.23-efficiency.md +83 -0
- package/docs/superpowers/specs/2026-10-04-yoke-1.23-efficiency-design.md +120 -0
- package/gemini-extension.json +1 -1
- package/package.json +1 -1
|
@@ -28,9 +28,51 @@ Install and pin these tools in the project or developer environment before enabl
|
|
|
28
28
|
|
|
29
29
|
- Every request is tied to a workspace and a content-addressed snapshot of committed and uncommitted, non-sensitive files.
|
|
30
30
|
- Paths are workspace-relative; traversal, symlinks, `.env`/key material and policy-excluded paths are rejected.
|
|
31
|
-
- Read calls
|
|
31
|
+
- Read calls share a request deadline and a complete-response output budget. Results distinguish backend identity, source location, semantic resolution and unverified index freshness. See the limits and evidence contracts below.
|
|
32
32
|
- Edits run in a disposable Yoke sandbox. Preview writes a durable plan and diff but requires an approval record created by the Yoke control layer.
|
|
33
33
|
- Apply rechecks the snapshot, plan expiry, plan hash, an exclusive transaction lock and idempotency. It produces a managed worktree and does not commit or merge; the existing Yoke loop and gates remain authoritative for integration.
|
|
34
34
|
- Network access is not needed by the facade. Backend installation and any backend-specific network behavior remain an explicit environment concern.
|
|
35
35
|
|
|
36
36
|
For a host that has no native MCP support (currently Pi), use the normal Yoke loop and bounded CLI/skill adapter. Yoke does not invent a second MCP transport for that host.
|
|
37
|
+
|
|
38
|
+
## Evidence and trace contracts
|
|
39
|
+
|
|
40
|
+
`code_trace` treats Serena's incoming reference list as references to the requested symbol. Each validated referencing symbol has an edge **to the requested target**. Two adjacent results never imply an edge between those results. Incoming implementation lists similarly use the `implements` relation. A bare symbol name is resolved to a unique path-bound Serena identity before requesting these relations; ambiguous targets remain unresolved.
|
|
41
|
+
|
|
42
|
+
The pinned Graft trace interface returns text. Until an explicit edge format is validated, its output is retained as structural candidates, without invented call edges. Incoming references cannot establish outgoing calls, imports or inheritance: unsupported direction/relation combinations return partial results and explain the missing contract. `require_resolved` excludes unresolved structural candidates; it does not turn an unsupported query into a resolved graph. Semantic trace results currently cover one hop. Requests for greater depth are marked unverified instead of pretending traversal was completed.
|
|
43
|
+
|
|
44
|
+
`frontier_remaining` counts known entries omitted by result limits, and `traversal_complete` is false when exhaustive traversal is unverified. Zero known omissions does not prove that no other references exist. Node limits preserve graph endpoints: returned edges never point at removed nodes.
|
|
45
|
+
|
|
46
|
+
The workspace snapshot is content-addressed and an explicitly supplied snapshot is checked against the current workspace before processing. This check does **not** prove that Graft's index, Graphify's graph or the language server has processed those exact file contents. The supported wire contracts do not provide an independently validated index-to-snapshot binding. Consequently backend evidence uses `freshness: "unknown"` and `content_hash: null`; an untrusted backend field claiming `current` is not accepted as proof. The facade no longer stamps a current workspace file hash onto potentially older evidence.
|
|
47
|
+
|
|
48
|
+
Recognized, path-bound Serena symbol records can have `resolution: "resolved"` while freshness remains unknown. Unrecognized records remain unresolved and cannot create symbols or edges. Read responses use `status: "partial"` and partial coverage for available backends because index freshness and exhaustive coverage are not established. A successful process call alone does not imply complete coverage. Missing backends remain separately identified in `backends_missing`.
|
|
49
|
+
|
|
50
|
+
Each retained evidence link has a matching `provenance.evidence_id`. This additive field allows callers to resolve evidence IDs and lets output reduction remove unused provenance without breaking retained links. No response cache is used; no stale evidence is reused under a new snapshot or query.
|
|
51
|
+
|
|
52
|
+
## Output and time limits
|
|
53
|
+
|
|
54
|
+
The facade forwards all four configured limits:
|
|
55
|
+
|
|
56
|
+
```yaml
|
|
57
|
+
codeIntelligence:
|
|
58
|
+
mode: shadow
|
|
59
|
+
limits:
|
|
60
|
+
tokenBudget: 2400
|
|
61
|
+
timeoutMs: 10000
|
|
62
|
+
maxBytes: 2000000
|
|
63
|
+
maxBackends: 3
|
|
64
|
+
```
|
|
65
|
+
|
|
66
|
+
The effective output cap is the smaller of `maxBytes` and the effective token budget. A request's `token_budget`, where supported, can lower the configured cap, but cannot raise it. The default configured token budget is 2400. The budget covers the entire serialized JSON response in MCP `structuredContent`: result data, provenance, warnings, identifiers, coverage, errors and the metrics themselves. The outer JSON-RPC framing and the short MCP text summary are outside that response.
|
|
67
|
+
|
|
68
|
+
**Token accounting is deliberately conservative:** one UTF-8 byte consumes one budget unit. `metrics.result_tokens` is therefore an estimate with `token_count_kind: "estimated"` and `token_count_method: "utf8_bytes_conservative"`. It is not provider-measured usage, a tokenizer-specific exact count or a billing estimate. `metrics.returned_bytes` is the actual UTF-8 size of the serialized structured response, including the metrics. This replaces the earlier four-bytes-per-token approximation and makes the same numeric budget stricter; increase the configured budget explicitly if a workflow requires larger context, rather than treating the old number as an exact token allowance.
|
|
69
|
+
|
|
70
|
+
When an answer does not fit, text is shortened at Unicode-safe boundaries and lower-ranked entries are removed. Arrays and operation receipt fields keep their schema, retained graph edges have valid endpoints, and unused provenance is removed. The answer reports `partial`, adds a truncation warning and never claims complete coverage. Narrow the query or raise the configured budget to obtain omitted content; `next_cursor: null` does not imply a complete search.
|
|
71
|
+
|
|
72
|
+
A budget can be smaller than the mandatory response envelope. For example, 128 budget units cannot encode all required identifiers and error fields. Such requests fail with `BUDGET_EXCEEDED` **before a backend is invoked**. The small, fixed-shape error envelope is the only output-cap exception; its size is reported accurately, and it contains no backend payload or unbounded error text. Apply additionally reserves space for its entire concrete receipt, including all changed paths and a possible deadline warning, before starting a transaction. An insufficient receipt budget rejects the apply without mutating. A completed apply keeps its transaction receipt even if synchronous work crossed the deadline.
|
|
73
|
+
|
|
74
|
+
The effective deadline is the smaller of configured `timeoutMs` (10000 ms by default) and the request's `timeout_ms`/tool default. All backend calls, semantic target resolution and preview operations consume the same remaining time. MCP initialization and its subsequent tool call also share that allowance. A backend which ignores its timeout is closed, late results are discarded, and no subsequent backend starts after expiry. Backend shutdown has a separate bounded cleanup grace; synchronous filesystem snapshot/transaction operations cannot be preempted mid-operation, so the deadline is not an exact end-to-end wall-clock promise. A final deadline check after normalization and output reduction marks delayed results partial and keeps any completed operation's receipt. A preview needing the former 120000 ms tool default must also have a configured timeout of at least that size.
|
|
75
|
+
|
|
76
|
+
## Validation scope
|
|
77
|
+
|
|
78
|
+
The regression suite uses model-free adapters and local synthetic MCP processes. It verifies reference direction and endpoints, unknown payload handling, honest freshness and coverage, aggregate output limits, Unicode boundaries, evidence link preservation, tiny-budget failures, limit propagation, and shared deadlines including initialization. It does not assert a successful run against installed authenticated backends or an exact provider-token saving.
|
|
@@ -25,3 +25,18 @@ See the [Goals/resource audit](GOALS-RESOURCE-AUDIT-2026-09-29.md) for implement
|
|
|
25
25
|
Reproduce with `bench/compare-codex.mjs`, then `bench/analyze-codex-comparison.mjs`. Use fresh, short output roots on Windows. The new summary tests verify median arithmetic, cache subtraction, unknown usage, incompatible policies and invalid measurements.
|
|
26
26
|
|
|
27
27
|
Provenance: this report is disclosed as AI-assisted. Read-only text scans cannot establish human authorship or verify proprietary keyed watermarks; cryptographic verification and signer trust remain unknown without the corresponding verifier and trust policy.
|
|
28
|
+
|
|
29
|
+
## Later tooling update — 1.22.0
|
|
30
|
+
|
|
31
|
+
The measurements above remain the historical 1.19.0 results. The current analyzer
|
|
32
|
+
reads their aggregate JSON as legacy/unverified because those runs did not record
|
|
33
|
+
the new complete comparison manifest. It preserves the reported values and does
|
|
34
|
+
not manufacture missing provenance or reinterpret them as 1.22.0 measurements.
|
|
35
|
+
|
|
36
|
+
New runs record source/build, fixture, acceptance and submitted-input digests,
|
|
37
|
+
requested and reported model identities, and explicit startup conditions.
|
|
38
|
+
The historical direct-arm ignore-rules difference is now an expressly declared
|
|
39
|
+
workflow variable. Undeclared cross-arm differences, missing provenance and
|
|
40
|
+
failed acceptance do not produce performance comparison groups.
|
|
41
|
+
See [benchmark manifests](BENCHMARK-MANIFEST.md) for the contract and limitations.
|
|
42
|
+
This is a tooling change; no new authenticated comparison was performed for it.
|
|
@@ -0,0 +1,206 @@
|
|
|
1
|
+
# Delivery artifacts and executable user journeys
|
|
2
|
+
|
|
3
|
+
Use executable acceptance criteria to decide whether a release is ready. A build command shows that an artifact can be produced; a journey checks a specific user interaction. Yoke can now record the relationship between those criteria, named journeys and the exact artifact files present during a check.
|
|
4
|
+
|
|
5
|
+
The delivery declaration belongs in `.yoke/acceptance.yaml`. Browser smoke steps belong in `.yoke/config.yaml`. Existing flow definitions containing only `name`, `path` and an optional `landmark` remain valid, but production browser runs now also require `smoke.sourceIdentity`.
|
|
6
|
+
|
|
7
|
+
## Optional browser steps
|
|
8
|
+
|
|
9
|
+
`yoke flow-smoke` uses the target project's Playwright installation and its Chromium browser. Start the application or preview server before running the command. Yoke does not start a development server or install a browser automatically.
|
|
10
|
+
|
|
11
|
+
For example, add this section to the project's existing `.yoke/config.yaml`:
|
|
12
|
+
|
|
13
|
+
```yaml
|
|
14
|
+
smoke:
|
|
15
|
+
baseUrl: http://localhost:3000
|
|
16
|
+
sourceIdentity:
|
|
17
|
+
path: /assets/app-specific-static-source.js
|
|
18
|
+
sha256: "<replace with SHA-256 of the served bytes>"
|
|
19
|
+
flows:
|
|
20
|
+
- name: profile-survives-reload
|
|
21
|
+
path: /login
|
|
22
|
+
landmark: main
|
|
23
|
+
timeoutMs: 60000
|
|
24
|
+
steps:
|
|
25
|
+
- action: fill
|
|
26
|
+
selector: '[data-testid="email"]'
|
|
27
|
+
value: smoke-user@example.invalid
|
|
28
|
+
- action: fill
|
|
29
|
+
selector: '[data-testid="password"]'
|
|
30
|
+
valueEnv: SMOKE_TEST_PASSWORD
|
|
31
|
+
- action: click
|
|
32
|
+
selector: '[data-testid="sign-in"]'
|
|
33
|
+
- action: expect-url
|
|
34
|
+
url: /profile
|
|
35
|
+
- action: fill
|
|
36
|
+
selector: '[data-testid="display-name"]'
|
|
37
|
+
value: Smoke User
|
|
38
|
+
- action: press
|
|
39
|
+
selector: '[data-testid="display-name"]'
|
|
40
|
+
key: Tab
|
|
41
|
+
- action: click
|
|
42
|
+
selector: '[data-testid="save-profile"]'
|
|
43
|
+
- action: expect-visible
|
|
44
|
+
selector: '[data-testid="save-confirmation"]'
|
|
45
|
+
- action: expect-text
|
|
46
|
+
selector: '[data-testid="save-confirmation"]'
|
|
47
|
+
text: Profile saved
|
|
48
|
+
exact: true
|
|
49
|
+
timeoutMs: 10000
|
|
50
|
+
- action: reload
|
|
51
|
+
- action: expect-text
|
|
52
|
+
selector: '[data-testid="profile-name"]'
|
|
53
|
+
text: Smoke User
|
|
54
|
+
exact: true
|
|
55
|
+
```
|
|
56
|
+
|
|
57
|
+
Replace the source resource path and hash placeholder before running this example. Confirm the intended server's checkout/build and port, then fetch a stable, app-specific source resource from the effective `baseUrl` origin and calculate the SHA-256 of the exact served body bytes. Dev servers may transform local source, so hashing a local file alone is insufficient. Set `sha256` to the resulting 64-character lowercase hexadecimal digest. The resource must return a successful HTTP response without redirects, finish within five seconds and contain at most 1 MiB. It must remain stable through the run. Production smoke checks the pin before browser launch and again after the flows; a missing or mismatching pin prevents valid smoke evidence.
|
|
58
|
+
|
|
59
|
+
After an intentional source/build change, verify the server serves that change and explicitly refresh the pin. Do not update it automatically on a mismatch or accept a foreign process occupying the expected port. Choose a resource that changes with the relevant app source, rather than a shared health response or an unchanged marker. The static pin checks only that resource's bytes before and after the run; it does not independently prove the identity of every module, backend or dynamic response. Source fingerprints and other acceptance checks remain necessary.
|
|
60
|
+
|
|
61
|
+
Supply `SMOKE_TEST_PASSWORD` through the environment using credentials for a dedicated test account, then run:
|
|
62
|
+
|
|
63
|
+
```sh
|
|
64
|
+
yoke flow-smoke . --label=profile-release
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
`--url=http://localhost:4173` overrides `smoke.baseUrl`. Flow navigation keeps the existing `baseUrl + path` behavior, so configure their slashes consistently. Each flow gets a fresh browser context with a 1280 × 720 viewport. Steps within one flow share that context; different flows do not share their browser session.
|
|
68
|
+
|
|
69
|
+
### Supported steps
|
|
70
|
+
|
|
71
|
+
| Action | Fields | Behavior |
|
|
72
|
+
| --- | --- | --- |
|
|
73
|
+
| `click` | `selector` | Click the matching element using Playwright's action waiting. |
|
|
74
|
+
| `fill` | `selector`, exactly one of `value` or `valueEnv` | Fill a form control with a literal value or a captured environment value. Missing environment values fail the step. |
|
|
75
|
+
| `press` | `selector`, `key` | Press a key, such as `Enter` or `Tab`, on the matching element. |
|
|
76
|
+
| `expect-visible` | `selector` | Wait for the matching element to be visible. |
|
|
77
|
+
| `expect-text` | `selector`, `text`, optional `exact` | Wait for visible text. Whitespace is normalized. Matching is case sensitive and uses a literal substring by default; `exact: true` compares the complete normalized text. |
|
|
78
|
+
| `expect-url` | `url` | Wait for the exact resolved URL. A relative URL resolves against the effective base URL. This is not a glob or regular expression. |
|
|
79
|
+
| `reload` | No additional action fields | Reload the page, wait for its load event and reject a non-successful HTTP response. |
|
|
80
|
+
|
|
81
|
+
Each step may specify `timeoutMs`. Unknown actions and extra action fields are rejected before browser launch. There is no arbitrary JavaScript action, general scripting language, branch or loop. Use a project E2E test command for more complex behavior, multiple pages, downloads, native applications or specialized fixtures.
|
|
82
|
+
|
|
83
|
+
### Limits and failure behavior
|
|
84
|
+
|
|
85
|
+
A flow may contain **1–50 steps**. Its default total timeout is **60,000 ms**, with a configurable maximum of **120,000 ms**. This budget includes browser-context preparation, navigation, the optional landmark and all steps. Navigation has its own maximum of 30,000 ms. A landmark has a maximum of 10,000 ms.
|
|
86
|
+
|
|
87
|
+
A step defaults to **10,000 ms** and may request at most **30,000 ms**. Its actual allowance is capped by the remaining flow budget. Text waiting has both an attempt limit and a deadline. Screenshot capture, context closing and video finalization each have a separate 5,000 ms allowance after interaction ends; browser startup and shutdown also have bounded waits. The flow timeout is therefore not a promise that the whole multi-flow command, including evidence capture and cleanup, ends within that duration.
|
|
88
|
+
|
|
89
|
+
The first failed step stops the remaining steps of that flow. The report marks them `skipped`. Later flows still run. Non-successful navigation, page errors and error-level console events also fail a flow. A screenshot is required for a passing result. Failure video remains best-effort evidence and does not replace a missing screenshot.
|
|
90
|
+
|
|
91
|
+
Fill values are limited to 8,192 characters, including values read from the environment. Environment-backed fill values are captured once when the run begins.
|
|
92
|
+
|
|
93
|
+
## Evidence from a smoke run
|
|
94
|
+
|
|
95
|
+
Evidence is written to `.yoke/proof/<label>/`:
|
|
96
|
+
|
|
97
|
+
- A screenshot for each flow where screenshot capture succeeds.
|
|
98
|
+
- A `.webm` video for a failed flow when Playwright can finalize it.
|
|
99
|
+
- `report.json`, a versioned machine-readable report.
|
|
100
|
+
|
|
101
|
+
The default label is `YOKE_STORY` when present, otherwise `latest`; `--label` takes priority. Labels and filenames are sanitized, and colliding sanitized flow names receive distinct filenames. Running the same label replaces its previous evidence. A run that cannot start because its configuration, browser, source identity or proof location is unavailable does not erase the prior evidence during that preflight.
|
|
102
|
+
|
|
103
|
+
The report contains:
|
|
104
|
+
|
|
105
|
+
- Per-flow status, navigation status, optional landmark status and duration.
|
|
106
|
+
- Each configured step's index, action, status, duration and a bounded failure category.
|
|
107
|
+
- Relative screenshot and failure-video filenames.
|
|
108
|
+
- Source fingerprints before and after the run, and whether they match.
|
|
109
|
+
- A digest of the effective smoke configuration, URL override and captured fill inputs.
|
|
110
|
+
- Node version, operating-system platform, architecture, configured viewport, browser version when available and the effective URL's origin.
|
|
111
|
+
|
|
112
|
+
Raw configuration, selectors, expected text, fill values, environment-variable names and browser exception call logs are not serialized into the report. URL credentials, query strings and fragments are not included in its recorded origin. Driver errors are represented by categories such as `timeout`, `step-failed` or `navigation-failed`, because their original messages can contain form values. Screenshots and video still show what the application visibly displays; use test data appropriate for those visual artifacts.
|
|
113
|
+
|
|
114
|
+
Exit code `0` means all flows passed, their screenshots were captured, the source fingerprints before and after the run matched and no browser-cleanup failure was reported. Exit code `1` represents a failed run or invalidated evidence. Exit code `2` covers unavailable prerequisites or evidence storage. Invalid step configuration is rejected before browser launch.
|
|
115
|
+
|
|
116
|
+
The local source fingerprint does **not** prove that a remote server is running that source revision. Arrange the project's preview or deployment checks so the browser exercises the intended build. Record a separate deployment verification criterion when that relationship matters.
|
|
117
|
+
|
|
118
|
+
## Bind release artifacts and journeys to acceptance criteria
|
|
119
|
+
|
|
120
|
+
Add an optional `delivery` section to `.yoke/acceptance.yaml`:
|
|
121
|
+
|
|
122
|
+
```yaml
|
|
123
|
+
version: 1
|
|
124
|
+
protected:
|
|
125
|
+
- .yoke/config.yaml
|
|
126
|
+
- playwright.config.ts
|
|
127
|
+
- tests/e2e/profile.spec.ts
|
|
128
|
+
criteria:
|
|
129
|
+
- id: profile-e2e
|
|
130
|
+
text: A signed-in user can save a profile and recover it after reloading.
|
|
131
|
+
commands:
|
|
132
|
+
- npm run test:e2e -- tests/e2e/profile.spec.ts
|
|
133
|
+
- id: profile-smoke
|
|
134
|
+
text: The release preview completes the configured profile journey.
|
|
135
|
+
commands:
|
|
136
|
+
- yoke flow-smoke . --label=profile-release
|
|
137
|
+
delivery:
|
|
138
|
+
version: 1
|
|
139
|
+
artifacts:
|
|
140
|
+
- path: dist/release.zip
|
|
141
|
+
criteria: [profile-e2e, profile-smoke]
|
|
142
|
+
journeys:
|
|
143
|
+
- id: profile-update
|
|
144
|
+
text: Sign in, change a profile field, save it, reload and observe the saved value.
|
|
145
|
+
criteria: [profile-e2e, profile-smoke]
|
|
146
|
+
environment:
|
|
147
|
+
name: Local release preview with Chromium
|
|
148
|
+
url: http://localhost:3000
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
The script names, test files and artifact path in this example must exist in the application project. `delivery.journeys` names the intended user experience and references executable criteria; it does not execute browser steps by itself. The example executes smoke steps through the `profile-smoke` criterion's command.
|
|
152
|
+
|
|
153
|
+
**Build declared artifacts before starting `yoke check`:**
|
|
154
|
+
|
|
155
|
+
```sh
|
|
156
|
+
npm run build:release
|
|
157
|
+
# Start the project's preview server for that release in another terminal.
|
|
158
|
+
yoke check . --json
|
|
159
|
+
```
|
|
160
|
+
|
|
161
|
+
`build:release` here denotes the project's build script that creates `dist/release.zip`. A criterion that creates the declared artifact during `yoke check` is too late: the artifact must already exist for its initial snapshot.
|
|
162
|
+
|
|
163
|
+
The check report records artifact SHA-256 hashes and sizes before and after the commands. An artifact missing at either snapshot, differing between snapshots or using an unsafe path fails artifact binding. Declared artifacts must be regular project-relative files without symbolic links or path escapes. Explicit artifact hashing also covers ordinary ignored build outputs.
|
|
164
|
+
|
|
165
|
+
Artifact hashing accepts at most **512 MiB per file**. Each snapshot, before or after verification, has a combined **1 GiB byte budget** and a **10-second deadline**; an earlier enclosing check deadline also applies. Exceeding a limit or cancelling the check prevents a passing binding. Data is streamed in bounded chunks, with deadline and cancellation checks between synchronous reads. An individual blocked filesystem read is not preempted by that deadline. The declaration accepts at most 32 artifacts and 100 named journeys.
|
|
166
|
+
|
|
167
|
+
Artifact and journey status comes from their referenced criteria. All referenced criteria must pass for a passing status; a failed criterion fails the binding, while an unexecuted criterion remains unverified. Source-integrity, cancellation or artifact-binding problems also prevent a passing delivery result. Unknown criterion IDs and duplicate artifact paths or journey IDs are rejected.
|
|
168
|
+
|
|
169
|
+
The report labels this relationship **`binding: declared-criteria`**. Matching artifact hashes plus green criteria establish the declared relationship and the same artifact content at both snapshots. This is not continuous monitoring of the file between snapshots. The project command must actually exercise the intended artifact; Yoke cannot infer that a test which ignores its APK, ZIP or binary argument has verified that file.
|
|
170
|
+
|
|
171
|
+
Protect the acceptance manifest and the test infrastructure with the existing `yoke check . --protect` workflow after reviewing the contract. See [Verified projects](VERIFIED-PROJECTS.md) for protection and intentional-baseline-refresh behavior.
|
|
172
|
+
|
|
173
|
+
## Android example: verify the built APK on a specified device
|
|
174
|
+
|
|
175
|
+
A browser journey does not install an APK or validate Android permission handling. Use an application-owned executable test for that task and reference it from acceptance:
|
|
176
|
+
|
|
177
|
+
```yaml
|
|
178
|
+
version: 1
|
|
179
|
+
protected:
|
|
180
|
+
- tools/check-apk-install.mjs
|
|
181
|
+
- tests/android/tracking-fixture.json
|
|
182
|
+
criteria:
|
|
183
|
+
- id: tracking-installed-apk
|
|
184
|
+
text: The declared APK installs and completes the tracking fixture on the selected emulator.
|
|
185
|
+
commands:
|
|
186
|
+
- node tools/check-apk-install.mjs --apk=android/app/build/outputs/apk/debug/app-debug.apk --serial=emulator-5554
|
|
187
|
+
delivery:
|
|
188
|
+
version: 1
|
|
189
|
+
artifacts:
|
|
190
|
+
- path: android/app/build/outputs/apk/debug/app-debug.apk
|
|
191
|
+
criteria: [tracking-installed-apk]
|
|
192
|
+
journeys:
|
|
193
|
+
- id: record-and-reopen
|
|
194
|
+
text: Grant the required permission, start tracking, receive fixture GPS points, stop, save and reopen the route after restarting the app.
|
|
195
|
+
criteria: [tracking-installed-apk]
|
|
196
|
+
environment:
|
|
197
|
+
name: Android emulator emulator-5554 with the project's configured system image
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
`check-apk-install.mjs` and the fixture are **project-supplied tests**, not bundled Yoke commands. The script must install that exact APK, select and verify its device, perform the required permission and tracking actions, inspect the resulting route and process logs, and exit nonzero when an assertion fails. Build the APK first, prepare the emulator, then run `yoke check`.
|
|
201
|
+
|
|
202
|
+
The declared environment name is metadata. It does not detect the actual emulator image, phone model or OS version. Have the project test verify those properties when they are acceptance requirements. Passing on one emulator does not establish correct background tracking on a particular Vivo phone, and passing against one deployment does not certify every deployment.
|
|
203
|
+
|
|
204
|
+
## Validation scope
|
|
205
|
+
|
|
206
|
+
The smoke engine's repository tests use filesystem-backed browser doubles to verify step order, failure handling, timeout budgets, redaction, source/configuration binding and evidence persistence. They do not install browsers or certify an application. Application-level evidence comes from successfully running the configured commands with the project's actual browser, build, server or device.
|
|
@@ -0,0 +1,180 @@
|
|
|
1
|
+
# Measured capability routing
|
|
2
|
+
|
|
3
|
+
Yoke 1.22 adds optional economic selection to capability routing. It compares
|
|
4
|
+
observed, independently verified execution sequences. It does not assign an
|
|
5
|
+
estimated success probability to a model or assume that a profile labelled
|
|
6
|
+
`low` has the lowest cost of completing a task.
|
|
7
|
+
|
|
8
|
+
## Configuration and compatibility
|
|
9
|
+
|
|
10
|
+
```yaml
|
|
11
|
+
routing:
|
|
12
|
+
enabled: true
|
|
13
|
+
strategy: capability
|
|
14
|
+
optimization:
|
|
15
|
+
version: 1
|
|
16
|
+
objective: balanced
|
|
17
|
+
minSamples: 20
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
This fragment supplements the project's existing routing profiles and limits.
|
|
21
|
+
`minSamples` defaults to 20 and has a minimum of 10. Existing configurations
|
|
22
|
+
without `routing.optimization` retain their previous capability ordering. New
|
|
23
|
+
setup-generated capability configurations enable the conservative `balanced`
|
|
24
|
+
policy; this does not change selection until sufficient comparable measurements
|
|
25
|
+
exist.
|
|
26
|
+
|
|
27
|
+
The planner's task/risk assessment still determines the minimum capability tier.
|
|
28
|
+
Configured provider constraints, allowed roles, `maxTier`, parent-fallback policy,
|
|
29
|
+
and the existing gate-failure reliability filter remain in force. Economic
|
|
30
|
+
selection only compares profiles that survive those checks. Explicit project
|
|
31
|
+
routing rules retain their precedence.
|
|
32
|
+
|
|
33
|
+
## What a sample measures
|
|
34
|
+
|
|
35
|
+
The comparison unit is a bounded **execution sequence for one task contract**,
|
|
36
|
+
attributed to the profile that started it. If a small model needs a targeted
|
|
37
|
+
repair and then escalation to a stronger model, all those measured attempts
|
|
38
|
+
belong to the initial profile's sequence. Charging only the final successful
|
|
39
|
+
model would hide the cost of the failed attempts.
|
|
40
|
+
|
|
41
|
+
Within a fully accounted execution attempt, the durable call ledger includes:
|
|
42
|
+
|
|
43
|
+
- The routing/planning call made inside that execution, when one was needed.
|
|
44
|
+
- The implementation call, including reported usage from interrupted calls.
|
|
45
|
+
- Additional reviewer, critic, repair and model-driven quality calls joined by
|
|
46
|
+
the loop reporter before its final outcome.
|
|
47
|
+
|
|
48
|
+
Stable call IDs prevent a worker's usage from being counted again when its
|
|
49
|
+
aggregate reaches the reporter. An explicit attempt ID can join a call directly.
|
|
50
|
+
Without one, a role call is joined only if that story has exactly one open
|
|
51
|
+
attempt. Competing candidates make the attribution ambiguous; all affected
|
|
52
|
+
samples become incomplete instead of assigning the bill by guesswork.
|
|
53
|
+
|
|
54
|
+
The measured scope is `execution-attempt`. It is not the total cost of producing
|
|
55
|
+
or shipping a product. Separately prepared PRD drafts, batch assessments, change
|
|
56
|
+
planning, human work and later operational costs are not amortized into these
|
|
57
|
+
samples. Their separately emitted telemetry remains useful for project-level
|
|
58
|
+
analysis. The duration covers each admitted attempt through its recorded gate
|
|
59
|
+
outcome, summed across the sequence; it does not include idle time between
|
|
60
|
+
separate attempts or a developer's review time.
|
|
61
|
+
|
|
62
|
+
Only native loop paths with the complete role-accounting contract enable this
|
|
63
|
+
scope. Injected runners/gates and ambiguous candidate races remain `worker`
|
|
64
|
+
scope. Their measurements are still available diagnostically, but cannot supply
|
|
65
|
+
the complete evidence required for economic selection.
|
|
66
|
+
|
|
67
|
+
## Evidence and missing data
|
|
68
|
+
|
|
69
|
+
For each eligible starting profile, the router examines its most recent
|
|
70
|
+
`minSamples` comparable sequences. They must share the project, implementation
|
|
71
|
+
role, task-assessment dimensions, effective execution-policy key, requested
|
|
72
|
+
provider/model/effort/variant, and the current reported concrete starting model.
|
|
73
|
+
The policy key binds the effective gate and review configuration, including
|
|
74
|
+
commands and retries. A change to the concrete model starts a new evidence
|
|
75
|
+
window even when its requested alias remains the same.
|
|
76
|
+
|
|
77
|
+
A sequence is complete only when every reserved attempt has an outcome and the
|
|
78
|
+
sequence has either passed independent gates or exhausted its bounded attempt
|
|
79
|
+
allowance. Every attempt must have known token coverage, complete reported cost,
|
|
80
|
+
the full execution scope, and the same policy key. Infrastructure failures are
|
|
81
|
+
not treated as evidence of model reasoning quality and do not qualify for the
|
|
82
|
+
economic comparison.
|
|
83
|
+
|
|
84
|
+
One recent unknown or partial sequence keeps its window ineligible. The router
|
|
85
|
+
does not discard that row and search for an older successful subset. Open
|
|
86
|
+
sequence records are replaced by their later completion record, and retries do
|
|
87
|
+
not turn into additional independent samples. The optional observation registry
|
|
88
|
+
uses a 30-day evidence window and reads at most its latest 1,000 events; losing
|
|
89
|
+
economic history leads back to conservative selection.
|
|
90
|
+
|
|
91
|
+
Unknown usage is distinct from measured zero. Known partial token or dollar
|
|
92
|
+
amounts are retained with incomplete-coverage flags. For example, a fully
|
|
93
|
+
measured worker remains visible when its planner supplied no usage. OpenCode
|
|
94
|
+
and Kilo step summaries require coverage of every completed step before they
|
|
95
|
+
claim complete token totals; absent optional cache or cost fields are not
|
|
96
|
+
converted into measured zeros.
|
|
97
|
+
|
|
98
|
+
No reported dollars means no cost ranking. Yoke does not invent a bill from
|
|
99
|
+
declared cost tiers, token counts or a presumed provider price table. This can
|
|
100
|
+
leave economic routing at its conservative baseline for providers that do not
|
|
101
|
+
report sufficient telemetry.
|
|
102
|
+
|
|
103
|
+
## The three objectives
|
|
104
|
+
|
|
105
|
+
The baseline is the first eligible profile under the existing tier, declared
|
|
106
|
+
cost-tier and profile-ID ordering. Both the baseline and an alternative require
|
|
107
|
+
a complete evidence window. An alternative must have at least as many accepted
|
|
108
|
+
sequences in that equally sized window as the baseline.
|
|
109
|
+
|
|
110
|
+
The router computes two observed ratios:
|
|
111
|
+
|
|
112
|
+
```text
|
|
113
|
+
cost per accepted sequence = total measured sequence dollars / accepted sequences
|
|
114
|
+
time per accepted sequence = total measured sequence duration / accepted sequences
|
|
115
|
+
```
|
|
116
|
+
|
|
117
|
+
The numerators include completed failed sequences and every measured repair or
|
|
118
|
+
escalation in the window. A profile with zero accepted sequences is ineligible.
|
|
119
|
+
|
|
120
|
+
| Objective | Condition for choosing an alternative |
|
|
121
|
+
| --- | --- |
|
|
122
|
+
| `cost` | Lower observed dollars per accepted sequence, with no lower observed acceptance count. |
|
|
123
|
+
| `speed` | Lower observed duration per accepted sequence, with no lower observed acceptance count. |
|
|
124
|
+
| `balanced` | No higher cost or duration, and at least one strictly lower, with no lower observed acceptance count. |
|
|
125
|
+
|
|
126
|
+
`balanced` uses this conservative comparison instead of inventing a dollar
|
|
127
|
+
value for a second of runtime. Cost and speed can disagree; use the corresponding
|
|
128
|
+
explicit objective if that tradeoff is intentional. Equal scores preserve the
|
|
129
|
+
existing ordering.
|
|
130
|
+
|
|
131
|
+
Decision reasons identify the objective, sample count, observed acceptance
|
|
132
|
+
count, and measured ratios. When the baseline lacks enough data, they explain
|
|
133
|
+
the missing evidence and retain its existing ordering. These are observational
|
|
134
|
+
measurements, not calibrated forecasts or statistical confidence guarantees.
|
|
135
|
+
Task mix, repeated correlated failures and changing model behavior can still
|
|
136
|
+
affect the comparison. The minimum window is a conservative product policy,
|
|
137
|
+
not a claim that twenty examples prove superiority.
|
|
138
|
+
|
|
139
|
+
Reviewer and critic approval frequency is not independent evidence that their
|
|
140
|
+
reviews were good. Economic reordering of those roles therefore remains disabled
|
|
141
|
+
until an independent role-quality measure is available; their risk-based floors
|
|
142
|
+
and explicit model selections continue to apply.
|
|
143
|
+
|
|
144
|
+
## Persistent attempt admission
|
|
145
|
+
|
|
146
|
+
Capability routing reserves an attempt before invoking its implementation
|
|
147
|
+
provider. Configured bounded legacy routes use the same account. The account
|
|
148
|
+
is stored under the stable project root, keyed by story and bound task/plan
|
|
149
|
+
contract, and is independent of the optional global observation registry.
|
|
150
|
+
Exclusive immutable reservation files prevent two processes from replacing the
|
|
151
|
+
same slot. An unreadable or unwritable authoritative account blocks admission.
|
|
152
|
+
|
|
153
|
+
The capability allowance is the smaller of `maxAttempts` and the existing
|
|
154
|
+
tier-dependent bound: five attempts from `light`, four from `standard`, three
|
|
155
|
+
from `strong`, and two from `frontier`. Independent outer loop, goal and local
|
|
156
|
+
retry limits can stop execution sooner. Infrastructure or interrupted attempts
|
|
157
|
+
still occupy their reserved slots, but do not become verified reasoning failures
|
|
158
|
+
that automatically demand a stronger model.
|
|
159
|
+
|
|
160
|
+
Deleting or filling the optional registry, restarting a runner, or executing in
|
|
161
|
+
another disposable worktree does not restore those slots. Repeated preparation
|
|
162
|
+
or budget refusals before an implementation is admitted do not consume an
|
|
163
|
+
implementation attempt. A planner may itself incur cost before a later worker
|
|
164
|
+
is refused; that paid call is still reported. Goal callers recheck their durable
|
|
165
|
+
budget after planning and before implementation reservation.
|
|
166
|
+
|
|
167
|
+
A changed task/plan contract has a new account. Resume an unchanged blocked task
|
|
168
|
+
only after inspecting its reason and performing the required replan or explicit
|
|
169
|
+
configuration change. Do not delete authoritative attempt state to disguise
|
|
170
|
+
retries as a new run.
|
|
171
|
+
|
|
172
|
+
## Validation limits
|
|
173
|
+
|
|
174
|
+
The release tests use deterministic fake calls and synthetic observations. They
|
|
175
|
+
cover registry write failure and history eviction, restart-safe reservations,
|
|
176
|
+
known partial usage, synchronous/asynchronous provider telemetry, paid early
|
|
177
|
+
blocks, callback/write failures, complete escalation accounting, cold starts,
|
|
178
|
+
model/policy changes and objective selection. They are regression tests for the
|
|
179
|
+
accounting and decision rules, not provider performance benchmarks. No claimed
|
|
180
|
+
cost or speed improvement follows from those test fixtures.
|
package/docs/GOALS.md
CHANGED
|
@@ -29,11 +29,44 @@ IDs and executable commands and binds the objective to a digest of the acceptanc
|
|
|
29
29
|
manifest. It cannot determine whether your tests fully express a free-text request.
|
|
30
30
|
All project acceptance and configured regression checks must still pass.
|
|
31
31
|
|
|
32
|
+
Each independent check writes its full report, including declared artifact and
|
|
33
|
+
journey evidence. Goal state keeps the check ID in `lastCheck` and its report path
|
|
34
|
+
in `lastCheckEvidencePath`; `goal handoff` includes both. The report binds its
|
|
35
|
+
findings to the checked workspace. Declaring an artifact or journey is not itself
|
|
36
|
+
proof that a criterion passed.
|
|
37
|
+
|
|
32
38
|
Existing goals without a binding now stop before any model call. Review their tests,
|
|
33
39
|
then explicitly bind them with `yoke goal bind . --criteria=checkout`. Attempts and
|
|
34
40
|
measured consumption remain intact. Changed acceptance infrastructure still requires
|
|
35
41
|
an explicit protection refresh after review; binding never refreshes protected tests.
|
|
36
42
|
|
|
43
|
+
## Prepare routing explicitly
|
|
44
|
+
|
|
45
|
+
Capability routing uses the same goal ID, objective and structured executable
|
|
46
|
+
criteria as goal acceptance. It follows the project's `routing.assessmentPolicy`,
|
|
47
|
+
explicit `routing.rules`, profile limits and optional optimization settings.
|
|
48
|
+
|
|
49
|
+
With `routing.assessmentPolicy: prepared`, first prepare the bound goal:
|
|
50
|
+
|
|
51
|
+
```sh
|
|
52
|
+
yoke goal assess . --runner=codex
|
|
53
|
+
yoke goal run . --runner=codex
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
`goal assess` makes at most one read-only planning call and saves its assessment.
|
|
57
|
+
It reuses an existing assessment for the current contract. It does not implement
|
|
58
|
+
the objective or create an implementation attempt. Preparation consumes the same
|
|
59
|
+
durable token and provider-time budgets as execution. The planning provider and
|
|
60
|
+
model follow the project's planning configuration.
|
|
61
|
+
|
|
62
|
+
The following run reuses that assessment. A missing or stale prepared assessment
|
|
63
|
+
blocks execution before any model call, including when an explicit routing rule
|
|
64
|
+
exists. Changes to the objective, executable criteria, selected runner or approved
|
|
65
|
+
`.yoke/plan.md` invalidate the contract key. Run `goal assess` again after reviewing
|
|
66
|
+
those changes; acceptance changes still require explicit binding and protection
|
|
67
|
+
review. The `on-demand` policy may assess during execution; explicit rules avoid
|
|
68
|
+
the selection-planner call under that policy.
|
|
69
|
+
|
|
37
70
|
## Resource and budget controls
|
|
38
71
|
|
|
39
72
|
Goals and story loops share the project lock and the user-level worker pool. Routing
|
|
@@ -41,24 +74,45 @@ planner and implementation calls each acquire a permit; native model delegation
|
|
|
41
74
|
disabled. Explicit runner/model/effort/bare options override project runner defaults.
|
|
42
75
|
Execution uses Yoke's safe permission profile.
|
|
43
76
|
|
|
77
|
+
Runner defaults apply only to the provider that owns them. For example,
|
|
78
|
+
`--runner=claude` does not inherit a Codex model, effort, provider, variant or bare
|
|
79
|
+
setting from `runner`. Explicit command options still apply; when keeping the same
|
|
80
|
+
runner, omitted options retain its configured defaults.
|
|
81
|
+
|
|
44
82
|
| Setting | Meaning |
|
|
45
83
|
| --- | --- |
|
|
46
84
|
| `--attempts=N` on `set` or `budget` | Maximum cumulative attempts, default 3 |
|
|
47
85
|
| `--minutes=N` | Cumulative admitted provider work, default 30 minutes; excludes capacity waits and independent checks |
|
|
48
86
|
| `--wall-minutes=N` | Optional cumulative run time including admission, checks and state synchronization |
|
|
49
|
-
| `--tokens=N` | Cumulative measured input + output tokens; unknown consumption stops further budgeted work |
|
|
87
|
+
| `--tokens=N` | Cumulative measured input + output tokens, checked before every planning and implementation call; unknown consumption stops further budgeted work |
|
|
50
88
|
| `YOKE_MAX_PARALLEL_WORKERS=1..8` | Shared model/integration permit ceiling, default 3 |
|
|
51
89
|
| `YOKE_MAX_PARALLEL_CHECKS=1..8` | Separate shared asynchronous goal verification ceiling, default 1 |
|
|
52
90
|
|
|
91
|
+
Planner consumption is persisted before admitting its implementation worker. If
|
|
92
|
+
planning reaches the ceiling, overruns it, or leaves usage unknown, no next model
|
|
93
|
+
call is admitted. A completed objective exactly at its measured ceiling may finish;
|
|
94
|
+
the same balance cannot authorize another model call. Native Codex receives only
|
|
95
|
+
the remaining token allowance after earlier planning and implementation calls.
|
|
96
|
+
|
|
53
97
|
After a measured token overrun, Yoke retains evidence and work and reports `blocked`
|
|
54
98
|
even if acceptance passed. Update the budget explicitly with `yoke goal budget .
|
|
55
99
|
--tokens=50000`; `--clear-token-budget` explicitly removes that ceiling. Providers
|
|
56
|
-
reporting usage only after a call cannot provide a hard mid-call token cap.
|
|
100
|
+
reporting usage only after a call cannot provide a hard mid-call token cap. This is
|
|
101
|
+
a call-admission limit plus cancellation where telemetry permits it. Native
|
|
57
102
|
Codex usage units and Yoke's token accounting are recorded separately. Native
|
|
58
103
|
stream updates use cumulative turn deltas and request cancellation when measured
|
|
59
104
|
consumption exceeds the goal ceiling. Notifications arrive after model requests,
|
|
60
105
|
so this can still overshoot within a request; progress updates are never added
|
|
61
|
-
twice to persisted final usage.
|
|
106
|
+
twice to persisted final usage. Partial counts remain a known lower bound, with
|
|
107
|
+
incomplete measurement explicitly recorded. Raising a token ceiling does not make
|
|
108
|
+
unknown consumption known; explicitly removing the ceiling is a separate decision.
|
|
109
|
+
|
|
110
|
+
The goal's `usageLedger` records each call before dispatch, then records its final
|
|
111
|
+
usage and duration. Recovery retains finished planner calls and marks unfinished
|
|
112
|
+
calls as interrupted with unknown final usage. Each call has a stable ID and its
|
|
113
|
+
own usage event; attempt summaries are not emitted again as additional token usage.
|
|
114
|
+
Earlier goal files retain their existing attempt totals once, and new calls are
|
|
115
|
+
charged separately. Preparation and resume never reset the previous consumption.
|
|
62
116
|
|
|
63
117
|
The asynchronous goal checker can interrupt its own command tree at a deadline or
|
|
64
118
|
pause. Synchronous `yoke check` and story gates retain their existing execution
|
|
@@ -78,7 +132,10 @@ Or set `goals.nativeCodex: true` in `.yoke/config.yaml`. The default is off;
|
|
|
78
132
|
Yoke negotiates the local Codex app-server goal capability, stores the thread binding,
|
|
79
133
|
pauses native automatic continuation, and explicitly starts one bounded development
|
|
80
134
|
turn. The same thread resumes on later attempts. Only Yoke's independent acceptance
|
|
81
|
-
can synchronize it to `complete`. Provider/model changes
|
|
135
|
+
can synchronize it to `complete`. Provider/model changes detach the previous binding
|
|
136
|
+
and retain it in `detachedNativeBindings`. A later Codex execution creates a binding
|
|
137
|
+
for the selected model; a goal finished by another provider does not need the old
|
|
138
|
+
Codex runtime to synchronize a historical thread.
|
|
82
139
|
An unsupported native goal method falls back to ordinary Codex execution; authentication,
|
|
83
140
|
transport and execution failures remain visible. No global Codex configuration changes
|
|
84
141
|
are needed. Support depends on the installed CLI, not on the model name alone.
|