driftproof 0.6.0 β†’ 0.7.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,7 +1,7 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
2
  # Driftproof
3
3
 
4
- **Continuous, model-version-bound verification of agent skills.**
4
+ **A dated proof that this skill, this hash, this model, still helps.**
5
5
 
6
6
  [![driftproof](https://img.shields.io/endpoint?url=https://driftproofhq.com/badges/commit-message-conventions.json)](https://driftproofhq.com)
7
7
  &nbsp;β€” live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
@@ -15,11 +15,11 @@ across model releases and you get a **drift report**.
15
15
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
16
16
  format; it does not invent its own.
17
17
 
18
- πŸ“Š **Six published reports** (each re-derived from committed receipts, nothing
19
- hand-entered), spanning five published report types, the fifth being revision
20
- drift. All six share one band-based, floor-gated verdict rule and differ in what
21
- moves underneath the skill β€” or, in the value report, in which axes are
22
- measured:
18
+ πŸ“Š **Seven published reports** (each re-derived from committed receipts, nothing
19
+ hand-entered), spanning six published report types, the newest being instrument
20
+ re-measurement. All seven share one band-based, floor-gated verdict rule and
21
+ differ in what moves underneath the skill β€” or, in the value report, in which
22
+ axes are measured; or, in the newest, in the instrument itself:
23
23
 
24
24
  - **[Report #001](https://driftproofhq.com/reports/001/)** β€” *release drift*: ten
25
25
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
@@ -51,11 +51,23 @@ measured:
51
51
  established β€” and the receipt spec gains generation sampling as a result. **No
52
52
  cause is asserted**; the control proves non-reproduction and cannot say why.
53
53
  *The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
54
+ - **[Report #007](https://driftproofhq.com/reports/007/)** β€” *instrument
55
+ re-measurement*: the three cells Report #005 published for these skills, run
56
+ again with the generation sampled adaptively instead of once and with the call
57
+ timeout the surface policy declares. **No cell separates**: every lift is
58
+ smaller than its own band, and the 21 cases read 3 improved, 0 regressed,
59
+ 18 no effect, 0 not measured. Five of six comparisons against the archive were
60
+ **refused** on their baseline control. The instrument defect it reports is its
61
+ own: a declared 300 s timeout had been shadowed by a `120000` literal since
62
+ 2026-07-27, and the truncated run it caused measured *less* variance than the
63
+ clean re-run, which is the direction that flatters an instrument. Both runs are
64
+ published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
54
65
 
55
66
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
56
- skills](https://driftproofhq.com/writing/three-releases/)**, reads all six reports
57
- together: what moves underneath a skill, and what the skill costs to run. Revised
58
- 2026-08-29; every figure in it is gate-checked against the report page it cites.
67
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all seven reports
68
+ together: what moves underneath a skill, what the skill costs to run, and what a
69
+ corrected instrument did to three published results. Revised 2026-09-01; every
70
+ figure in it is gate-checked against the report page it cites.
59
71
 
60
72
  ## Why
61
73
 
@@ -67,7 +79,7 @@ re-checks.
67
79
 
68
80
  **Verdicts age because the substrate moves.** Driftproof exists to keep the
69
81
  verdict current: cheap, repeatable, hash-stamped measurements bound to a specific
70
- model version, so "does this skill still help?" has a dated, reproducible answer
82
+ model version, so "does this skill still help?" has a dated, verifiable answer
71
83
  instead of a stale one.
72
84
 
73
85
  The hard part isn't running an eval once β€” it's making the number **credible
@@ -100,8 +112,10 @@ npx driftproof init my-skill
100
112
  export CLAUDE_PROVIDER=api
101
113
  read -rsp "Anthropic API key: " ANTHROPIC_API_KEY && export ANTHROPIC_API_KEY
102
114
 
103
- # 4. Run the suite with a hard $2 budget (the run is refused up front if it would exceed it)
104
- npx driftproof run my-skill --models claude-haiku-4-5 --max-usd 2
115
+ # 4. Run the suite. The shipped defaults (DEV_MAX_USD / DEV_MAX_CALLS in config.js)
116
+ # refuse the run up front if the projection exceeds either; --max-usd and
117
+ # --max-calls override them.
118
+ npx driftproof run my-skill --models claude-haiku-4-5
105
119
 
106
120
  # 5. Read the receipt + human summary written to ./receipts/
107
121
  cat receipts/*.summary.md
@@ -187,11 +201,16 @@ The surface and judge settings used are recorded in every receipt.
187
201
 
188
202
  ### Cost guard
189
203
 
190
- Sampling multiplies calls. Each case costs `2 + 2 Γ— samples` model calls
191
- (two generations, each judged `samples` times) β€” 12 calls per case at the default
192
- `--samples 5`. Driftproof **projects the whole run up front, prints the count, and
193
- refuses before spending anything** if it would exceed `--max-calls` (default 200).
194
- The default model list is `haiku` only.
204
+ Sampling multiplies calls on two axes, and the second one is easy to miss. Each
205
+ case costs `SAMPLING.max Γ— (2 + 2 Γ— samples)` model calls: two arms, each **drawn
206
+ up to `SAMPLING.max` times** (generation sampling, receipt spec v0.5), and every
207
+ draw judged `samples` times. Read at a single draw, that formula understates a run
208
+ by the whole draw factor β€” which is how the shipped caps came to sit a release
209
+ behind the estimator. Driftproof **projects the whole run up front, prints the
210
+ count, and refuses before spending anything** if the projection exceeds the
211
+ per-model call cap (`DEV_MAX_CALLS`) or the dollar budget (`DEV_MAX_USD`) β€” both
212
+ declared with their derivation in [`config.js`](config.js), both overridable with
213
+ `--max-calls` / `--max-usd`. The default model list is `haiku` only.
195
214
 
196
215
  ## Receipt anatomy
197
216
 
@@ -201,7 +220,7 @@ A receipt is the unit of evidence β€” one JSON document conforming to
201
220
 
202
221
  ```jsonc
203
222
  {
204
- "schema_version": "0.4",
223
+ "schema_version": "0.5",
205
224
  "skill": { "name": "commit-message-conventions", "version": "0.2.0",
206
225
  "content_hash": "…sha256 over SKILL.md + bundled files…" },
207
226
  "suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
@@ -210,7 +229,7 @@ A receipt is the unit of evidence β€” one JSON document conforming to
210
229
  "model_release_date": "2025-10-01",
211
230
  "provider": "anthropic",
212
231
  "surface": "claude-cli",
213
- "runner_version": "0.6.0",
232
+ "runner_version": "0.7.2",
214
233
  "date_utc": "2026-07-27T…Z",
215
234
  "registry": "registered",
216
235
  "transcripts": "hashes-only",
@@ -270,7 +289,7 @@ receipts. The rule that keeps it honest: a **regression** (or improvement) is
270
289
  claimed **only when the two bands do not overlap**. Overlapping bands are reported
271
290
  as **within noise** and never counted as a regression.
272
291
 
273
- ## Continuous verification in CI (GitHub Action + badge)
292
+ ## Verification in CI (GitHub Action + badge)
274
293
 
275
294
  Wire drift detection into a repo so the skill is re-checked on every push and when
276
295
  the model underneath it changes.
@@ -284,12 +303,13 @@ jobs:
284
303
  runs-on: ubuntu-latest
285
304
  steps:
286
305
  - uses: actions/checkout@v4
287
- - uses: driftproofhq/driftproof@v0.6.0
306
+ - uses: driftproofhq/driftproof@v0.7.2
288
307
  with:
289
308
  skill-dir: skills/my-skill
290
309
  models: claude-haiku-4-5
291
- max-usd: '2'
292
310
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
311
+ # max-usd: <n> # override the dollar budget (default: DEV_MAX_USD in config.js)
312
+ # max-calls: <n> # override the per-model call cap (default: DEV_MAX_CALLS)
293
313
  # fail-on-regression: 'true' # (default) fail the job if the skill REGRESSED
294
314
  ```
295
315
 
@@ -307,7 +327,7 @@ runs).
307
327
  JSON object. Commit it somewhere public and point a shields endpoint URL at it:
308
328
 
309
329
  ```bash
310
- npx driftproof run skills/my-skill --models claude-haiku-4-5 --max-usd 2 --out receipts
330
+ npx driftproof run skills/my-skill --models claude-haiku-4-5 --out receipts
311
331
  npx driftproof badge receipts/*.json --out badges/my-skill.json
312
332
  git add badges/my-skill.json && git commit -m "chore: driftproof badge"
313
333
  ```
@@ -323,7 +343,7 @@ site, so it reflects a real dated run, not a hand-set color.
323
343
 
324
344
  ## Reports
325
345
 
326
- Six reports are published, spanning five report types. A report page lives at a
346
+ Seven reports are published, spanning six report types. A report page lives at a
327
347
  draft path β€” `docs/reports/NNN-draft/` β€” until the publish sequence renames it, and
328
348
  `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
329
349
  (see the roll at the top of this README, and
package/bin/driftproof CHANGED
@@ -4,9 +4,10 @@
4
4
 
5
5
  const fs = require('fs');
6
6
  const path = require('path');
7
- const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = require('../config');
7
+ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD, DEV_MAX_CALLS } = require('../config');
8
8
  const { loadSkill } = require('../lib/skill');
9
9
  const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
10
+ const { SAMPLING } = require('../lib/sampling');
10
11
  const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
11
12
  const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
12
13
  const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
@@ -91,8 +92,12 @@ ENV
91
92
  NOTES
92
93
  - Default models list is 'haiku' only (cheap dev default).
93
94
  - Sampled judging: each case does 2 generations + 2Γ—samples judge calls.
94
- Default --samples 5 β†’ 12 calls/case; --samples 1 disables bands.
95
- - --max-calls is a hard per-run call cap (default 200). The whole run is
95
+ Default --samples 5 β†’ 12 calls/case; --samples 1 disables bands AND the
96
+ variance ratio: with one judge sample the per-draw judge spread is zero by
97
+ construction, so the receipt records variance_ratio: null with
98
+ variance_ratio_unavailable: "single_judge_sample". Legal, but a run whose
99
+ ratio you intend to publish needs --samples 2 or more.
100
+ - --max-calls is a hard per-run call cap (default ${DEV_MAX_CALLS}). The whole run is
96
101
  projected up front and REFUSED before any call if it would exceed the cap.
97
102
  - --max-usd is a hard dollar budget (default ${DEV_MAX_USD} for dev runs). The projected
98
103
  cost is printed up front and the run is REFUSED on BOTH surfaces if it would
@@ -126,7 +131,7 @@ async function cmdRun(positional, flags) {
126
131
  const rc = loadRc(skillDir);
127
132
  const models = String(flags.models || rc.models || 'haiku').split(',').map((s) => s.trim()).filter(Boolean);
128
133
  const maxCases = flags['max-cases'] ? parseInt(flags['max-cases'], 10) : (rc.max_cases != null ? parseInt(rc.max_cases, 10) : null);
129
- const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : 200);
134
+ const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : DEV_MAX_CALLS);
130
135
  const samples = flags.samples ? parseInt(flags.samples, 10) : (rc.samples != null ? parseInt(rc.samples, 10) : DEFAULT_JUDGE_SAMPLES);
131
136
  const judgeModel = flags['judge-model'] || rc.judge_model || null;
132
137
  const concurrency = flags.concurrency ? parseInt(flags.concurrency, 10) : (rc.concurrency != null ? parseInt(rc.concurrency, 10) : 1);
@@ -137,14 +142,20 @@ async function cmdRun(positional, flags) {
137
142
 
138
143
  const skill = loadSkill(skillDir);
139
144
  const nCases = maxCases ? Math.min(maxCases, skill.suite.caseCount) : skill.suite.caseCount;
140
- const perModelCalls = projectCalls(nCases, samples);
145
+ // v0.5 draws the generation up to SAMPLING.max times per arm, so BOTH the
146
+ // printed projection and the dollar guard below must be scaled by it. Left at
147
+ // draws=1 the guard would admit a run costing up to ten times its own
148
+ // projection β€” the display would say one number and the runner would refuse
149
+ // quoting another. Conservative on purpose: refusing a run that would have fit
150
+ // is recoverable, overspending is not.
151
+ const perModelCalls = projectCalls(nCases, samples, SAMPLING.max);
141
152
  const totalProjected = perModelCalls * models.length;
142
153
 
143
154
  // Dollar cost guard (Week 3): project the METERED USD cost up front. On the
144
155
  // `api` surface this is real spend, so we refuse if it would exceed --max-usd.
145
156
  // On `claude-cli` the metered spend is $0 (subscription); the figure is the
146
157
  // hypothetical "if run on the metered API" cost β€” printed, never blocks.
147
- const cost = estimateRunCostUSD({ caseCount: nCases, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
158
+ const cost = estimateRunCostUSD({ caseCount: nCases, draws: SAMPLING.max, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
148
159
  // Surface is per-model now (a run may mix an Anthropic and an OpenAI target).
149
160
  const surfaces = [...new Set(models.map((m) => surfaceForModel(m)))];
150
161
  const surface = surfaces.join(', ');
@@ -315,7 +326,7 @@ function cmdInit(positional) {
315
326
  console.log(`\nNext:
316
327
  1. Edit ${rel(path.join(dir, 'SKILL.md'))} with your skill's instructions.
317
328
  2. Edit ${rel(path.join(dir, 'evals', 'evals.json'))} β€” replace the 3 example cases (rubrics anchored at 0.80).
318
- 3. driftproof run ${rel(dir)} --models claude-haiku-4-5 --max-usd 2
329
+ 3. driftproof run ${rel(dir)} --models claude-haiku-4-5
319
330
  See AUTHORING.md for how to write a fair suite.`);
320
331
  }
321
332
 
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.6.0';
12
+ const RUNNER_VERSION = '0.7.2';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
@@ -29,7 +29,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
29
29
  // at run time so derived dollars stay reproducible), and the derived
30
30
  // `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
31
31
  // v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
32
- const RECEIPT_SCHEMA_VERSION = '0.4';
32
+ const RECEIPT_SCHEMA_VERSION = '0.5';
33
33
 
34
34
  // Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
35
35
  // these. The projection is refused before any call if it exceeds the cap, on
@@ -37,8 +37,33 @@ const RECEIPT_SCHEMA_VERSION = '0.4';
37
37
  // dev β€” an interactive `driftproof run`
38
38
  // report β€” a full Report-#001-style suite run (scripts/run-report-001.js)
39
39
  // trigger β€” a release-trigger-initiated prepare-report run
40
- const DEV_MAX_USD = 2;
41
- const REPORT_MAX_USD = 40;
40
+ // Recalibrated 2026-09-01 (spec 018). BOTH literals below were set against a
41
+ // draws=1 projection and were never rescaled when spec 014 (5e08ba2) made the
42
+ // runner draw the generation up to SAMPLING.max times per arm and spec 016 made
43
+ // estimateRunCostUSD require the draw factor. The projection grew 10x; these did
44
+ // not, so `npx driftproof run` refused every suite of two or more cases and the
45
+ // v0.7.0 Action self-test aborted with exit 3. REPORT_MAX_USD was raised 40->300
46
+ // on 2026-08-31 for exactly this reason; that pass missed the dev path, which is
47
+ // the one the Action and every npx user take.
48
+ //
49
+ // DERIVATION β€” the headroom the pre-sampling defaults carried is preserved, not
50
+ // widened. Bundled 10-case example, 5 judge samples, per-case factor 2+2*5 = 12:
51
+ // draws=1 -> 120 calls, $0.3525 headroom 200/120 = 1.67x, $2/$0.3525 = 5.67x
52
+ // draws=10 -> 1200 calls, $3.5250 1200 * 1.67 = 2000, $3.5250 * 5.67 = $20.00
53
+ // A 2000-call cap under draws=10 is exactly as tight as 200 was under draws=1.
54
+ // LITERALS ON PURPOSE (spec 018 AC-2): a default derived from the suite in hand
55
+ // can never fire, which retires the guard instead of recalibrating it.
56
+ const DEV_MAX_USD = 20;
57
+ const DEV_MAX_CALLS = 2000;
58
+ // Raised from $40 on 2026-08-31 by the repository owner, in daylight, recorded in
59
+ // DECISIONS. The run did not get more expensive β€” the projection got honest:
60
+ // spec 016 made estimateRunCostUSD require a draw factor, and v0.5 has drawn the
61
+ // generation up to SAMPLING.max times per arm since spec 014, so the old cap had
62
+ // been passing on a figure up to ten times too small. Set above the ceiling this
63
+ // instrument can currently reach (#007's measured basis at the ceiling is
64
+ // $268.83) rather than just above today's staging figure, because a cap tripped
65
+ // by the next honest run teaches everyone to nudge it.
66
+ const REPORT_MAX_USD = 300;
42
67
  const TRIGGER_MAX_USD = 25;
43
68
 
44
69
  // Default number of judge samples per case in sampled mode. Sampling is what
@@ -56,6 +81,24 @@ const DEFAULT_JUDGE_SAMPLES = 5;
56
81
  // spec/RECEIPT.md Β§ "Drift verdict rule" and the report methodology.
57
82
  const EFFECT_FLOOR = 0.05;
58
83
 
84
+ // ── generation sampling (receipt spec v0.5) ─────────────────────────────────
85
+ //
86
+ // Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
87
+ // instrument sampled only the judge. These are the policy that samples the
88
+ // other axis. Constants, not literals in the runner: lib/sampling.js reads
89
+ // them, so moving one here moves the policy, and the gate asserts that by
90
+ // value rather than by grepping for a name.
91
+ //
92
+ // MIN is 3 because two draws give an sd that is barely a measurement and one
93
+ // gives none at all. MAX is 10: #006's probe used 20 by hand and found the
94
+ // shape at well under half of that, and a per-case ceiling bounds the spend.
95
+ // The SD threshold is the effect floor β€” a spread wider than the smallest move
96
+ // the verdict rule will call real is exactly when more draws are owed.
97
+ const GENERATION_SAMPLES_MIN = 3;
98
+ const GENERATION_SAMPLES_MAX = 10;
99
+ const GENERATION_SD_THRESHOLD = EFFECT_FLOOR;
100
+ const GENERATION_STABILITY_EPS = 0.01;
101
+
59
102
  // Report #002 is the first CROSS-PROVIDER report: the same suites, the same fixed
60
103
  // Haiku judge, run on two substrates β€” a Claude flagship and a GPT flagship. The
61
104
  // GPT flagship is a config constant (not hard-coded across scripts) so a future
@@ -90,7 +133,8 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
90
133
 
91
134
  module.exports = {
92
135
  PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
93
- EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
136
+ EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
137
+ GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
94
138
  REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
95
139
  REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
96
140
  REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
package/lib/canary.js ADDED
@@ -0,0 +1,27 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ 'use strict';
3
+
4
+ // Suite canary β€” receipt spec v0.5.
5
+ //
6
+ // Terminal-Bench 4.0 records a canary string so that a benchmark appearing in
7
+ // training data is detectable. Adopted, translated: our unit of publication is
8
+ // the SUITE, so the canary is per suite, derived from the suite's identity and
9
+ // its case ids rather than randomly assigned, so the same suite yields the same
10
+ // canary on every machine and no registry has to be kept.
11
+ //
12
+ // It is a detection aid, not a control: a canary tells you a suite leaked, it
13
+ // does not stop the leak, and it cannot prove the absence of one.
14
+
15
+ const crypto = require('crypto');
16
+
17
+ const NAMESPACE = 'driftproof/suite-canary/v1';
18
+
19
+ function suiteCanary(suite) {
20
+ const id = String((suite && (suite.id || suite.name)) || '');
21
+ const caseIds = ((suite && suite.cases) || []).map((c) => String((c && c.id) || '')).sort();
22
+ const h = crypto.createHash('sha256').update(`${NAMESPACE}\n${id}\n${caseIds.join('\n')}`).digest('hex');
23
+ // Formatted as a GUID so it is recognisable as one in a corpus scan.
24
+ return [h.slice(0, 8), h.slice(8, 12), h.slice(12, 16), h.slice(16, 20), h.slice(20, 32)].join('-');
25
+ }
26
+
27
+ module.exports = { suiteCanary, NAMESPACE };
package/lib/cost.js CHANGED
@@ -58,7 +58,25 @@ function perCallCostUSD(modelId, kind) {
58
58
  // generated on each target model, judged `samples` times each on `judgeModel`.
59
59
  //
60
60
  // returns { totalUSD, perModel: [{ model, usd }], judgeUSD, genUSD, assumptions }
61
- function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
61
+ // THE DRAW FACTOR IS REQUIRED AND HAS NO DEFAULT (spec 016 AC-7).
62
+ //
63
+ // v0.5 draws the generation up to `SAMPLING.max` times per arm, so a run costs
64
+ // its one-draw estimate times the number of draws. `projectCalls` learned this in
65
+ // spec 014 with a `draws` argument DEFAULTING TO 1 "so every existing caller
66
+ // projects exactly what it projected before" β€” and every existing caller then
67
+ // went on projecting a tenth of the run, silently, for two more loops. Spec 015
68
+ // corrected the CALL projection in three report scripts and left the DOLLAR
69
+ // projection at one draw everywhere, which is the figure a human actually reads
70
+ // before authorising a paid run.
71
+ //
72
+ // A default is what made that invisible, so there is none. Omitting the factor
73
+ // throws, which turns a silent understatement into a loud stop β€” and every call
74
+ // site has to say what it means, including the ones that legitimately mean 1.
75
+ function estimateRunCostUSD({ caseCount, samples, models, judgeModel, draws }) {
76
+ if (!Number.isFinite(draws) || draws < 1) {
77
+ throw new Error('estimateRunCostUSD: `draws` is required and must be >= 1 β€” pass SAMPLING.max to project a v0.5 run, or 1 to price a single draw deliberately. It is not defaulted, because a default is how the dollar projection stayed at one draw through two loops.');
78
+ }
79
+ caseCount = caseCount * draws;
62
80
  const judgePrice = priceFor(judgeModel);
63
81
  const perModel = [];
64
82
  let judgeUSD = 0;
@@ -80,7 +98,7 @@ function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
80
98
  perModel,
81
99
  judgeUSD: round4(judgeUSD),
82
100
  genUSD: round4(genUSD),
83
- assumptions: { tokens: TOKENS, judgeModel, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
101
+ assumptions: { tokens: TOKENS, judgeModel, draws, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
84
102
  };
85
103
  }
86
104
 
package/lib/diff.js CHANGED
@@ -4,6 +4,7 @@
4
4
  const { bandVerdict, round } = require('./stats');
5
5
  const { EFFECT_FLOOR } = require('../config');
6
6
  const { revisionHeadline } = require('./revision');
7
+ const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
7
8
 
8
9
  // Practical-significance gate applied ON TOP of band separation. bandVerdict()
9
10
  // stays a pure geometry test (kept that way so its unit checks are unambiguous);
@@ -32,11 +33,23 @@ function isWithinNoise(v) { return v === 'within noise' || v === WITHIN_NOISE_FL
32
33
 
33
34
  // Map case-id β†’ { mean, stddev } for a receipt's with_skill cases. Falls back to
34
35
  // score/0 for v0.1 receipts that have no per-case band.
36
+ // NO RENDERING SWITCHES. An earlier revision carried two const-true flags so each
37
+ // half of F-015-B's fix could be removed in a probe β€” and the dead branch of one
38
+ // of them contained, verbatim, the false headline AC-4 exists to forbid. Nothing
39
+ // could fire it, and it was still a defect-restoring branch living in the tree
40
+ // that gets published. The mutation probes patch this source in a disposable copy.
41
+
35
42
  function withSkillBands(receipt) {
36
43
  const out = {};
37
44
  for (const c of receipt.results.cases) {
38
45
  if (c.mode !== 'with_skill') continue;
39
- out[c.id] = { mean: c.mean != null ? c.mean : c.score, stddev: c.stddev || 0 };
46
+ // ONE BAND DEFINITION (spec 016 AC-1). This built its own from the v0.4-shaped
47
+ // fields, which worked β€” and that is the point: three copies existed, two
48
+ // reading the legacy shape and one reading only v0.5, and the one that read
49
+ // only v0.5 was the one a cross-version control depended on (F-015-C). Routing
50
+ // every comparison path through `bandOf` means a future shape is added once.
51
+ const b = bandOf(c);
52
+ if (b) out[c.id] = { mean: b.mean, stddev: b.sd, source: b.source };
40
53
  }
41
54
  return out;
42
55
  }
@@ -51,7 +64,26 @@ function aggWithBand(receipt) {
51
64
 
52
65
  function fmt(n) { return n == null ? 'n/a' : (n >= 0 ? '+' : '') + n.toFixed(3); }
53
66
  function pct(n) { return n == null ? 'n/a' : n.toFixed(3); }
54
- function bandStr(x) { return `${x.mean.toFixed(3)} Β± ${x.stddev.toFixed(3)}`; }
67
+ // THE BAND SAYS WHICH BAND IT IS (spec 017 AC-7).
68
+ //
69
+ // `bandOf` computes `source: 'legacy' | 'generation'` and carries it onto every
70
+ // band; nothing rendered it, so a reader comparing an archived receipt with a
71
+ // v0.5 one was comparing a JUDGE-SAMPLE spread against an ACROSS-DRAW spread
72
+ // with nothing on the page saying so. They are different statistics over
73
+ // different things, and the comparison is still the only one v0.4 admits β€” which
74
+ // is exactly why the page has to name them rather than leave them to look alike.
75
+ //
76
+ // THE MARKER IS THE RECEIPT'S OWN WORD β€” `legacy` or `generation`, exactly as
77
+ // `bandOf` records it β€” rather than a prettier synonym. A reader who greps the
78
+ // page for what a receipt says should find the same token; a rendering that
79
+ // renames the thing it is disclosing has disclosed a different thing. Omitted
80
+ // when a band carries no source (an aggregate band is computed from case means,
81
+ // not from one case's draws).
82
+
83
+ function bandStr(x) {
84
+ const label = x && x.source;
85
+ return `${x.mean.toFixed(3)} Β± ${x.stddev.toFixed(3)}${label ? ` (${label})` : ''}`;
86
+ }
55
87
  function short(h) { return h ? String(h).slice(0, 12) : 'n/a'; }
56
88
 
57
89
  // The headline is a SUMMARY of the per-case band-overlap verdicts β€” the credibility
@@ -85,6 +117,25 @@ function revisionPairProblem(a, b) {
85
117
 
86
118
  function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
87
119
  const revision = mode === 'revision';
120
+
121
+ // AC-6 (spec 014) β€” THE PRECONDITION, IN THE PATH THAT EMITS THE VERDICT.
122
+ //
123
+ // A comparison across schema versions assumes the two runs measured the same
124
+ // thing. The assumption is checkable: the no-skill arm contains no skill text,
125
+ // so nothing about a skill or a spec revision can move it. If it fails to
126
+ // reproduce, the comparison has no ground, and Report #006 is what that looks
127
+ // like when it is checked β€” 3 cells, 0 measured, 3 refused.
128
+ //
129
+ // SCOPED TO CROSS-VERSION PAIRS, which is what AC-6 names. Revision-mode pairs
130
+ // already carry their own register through revisionPairProblem(); widening
131
+ // this to every same-version comparison is a live question, recorded in
132
+ // tasks.md as OPEN-QUESTION-3 rather than decided here.
133
+ const crossVersion = String(a.schema_version || '') !== String(b.schema_version || '');
134
+ let refusal = null;
135
+ if (crossVersion) {
136
+ const pre = baselineReproduces(a, b);
137
+ if (!pre.ok) refusal = { key: pre.key, reason: REFUSAL_REASONS[pre.key](pre) };
138
+ }
88
139
  const aB = withSkillBands(a);
89
140
  const bB = withSkillBands(b);
90
141
  const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
@@ -95,18 +146,21 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
95
146
  // improvement verdicts are never computed from evidence we did not run.
96
147
  const levelOf = (r) => r.verification_level || 'TESTED';
97
148
  const belowTested = [[labelA, a], [labelB, b]].filter(([, r]) => levelOf(r) !== 'TESTED');
98
- const measured = belowTested.length === 0;
149
+ // A REFUSAL IS A RESULT, and it suppresses the verdict exactly as an
150
+ // untested input does: no delta is asserted, and the reason travels with it.
151
+ const measured = belowTested.length === 0 && !refusal;
99
152
 
100
153
  const perCase = ids.map((id) => {
101
154
  const before = aB[id] || null;
102
155
  const after = bB[id] || null;
103
156
  const delta = (before && after) ? round(after.mean - before.mean) : null;
104
- const verdict = !measured ? 'not measured'
157
+ const verdict = refusal ? 'refused'
158
+ : !measured ? 'not measured'
105
159
  : (before && after) ? verdictWithFloor(before, after, delta) : 'n/a';
106
160
  return { id, before, after, delta, verdict };
107
161
  });
108
162
  // Sort worst-first: regressions, then by delta.
109
- const order = { regression: 0, 'within noise': 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3 };
163
+ const order = { regression: 0, 'within noise': 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3, refused: 3 };
110
164
  perCase.sort((x, y) => (order[x.verdict] - order[y.verdict]) || ((x.delta || 0) - (y.delta || 0)));
111
165
 
112
166
  const aAgg = aggWithBand(a);
@@ -172,13 +226,47 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
172
226
  if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β€” the eval suite changed; per-case comparison may be misleading.');
173
227
  if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β€” comparison is not meaningful.`);
174
228
  if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β€” non-overlap can only be trusted when both sides are sampled.');
175
- if (!measured) warnings.push(`verdicts NOT computed β€” ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
229
+ // WHAT SUPPRESSED THE VERDICT IS SAID, AND IT IS SAID CORRECTLY (spec 016 AC-3
230
+ // and AC-4, closing F-015-B).
231
+ //
232
+ // TWO DIFFERENT THINGS can suppress a verdict, and this line used to describe
233
+ // only one of them. A REFUSAL β€” the baseline-reproduction precondition β€” set
234
+ // `measured` false and then the caveat rendered `belowTested`, which on a
235
+ // refused pair is EMPTY: the page read `verdicts NOT computed β€” .` under a
236
+ // headline claiming `0 receipt(s) below TESTED`, on a pair where both receipts
237
+ // were TESTED. An empty reason and a false statement, in the artifact a report's
238
+ // verdicts come from. The refusal's own cause-honest reason was computed into
239
+ // `refusal.reason` and thrown away by the renderer, so nothing a reader could
240
+ // see said why the comparison had stopped.
241
+ //
242
+ // Each half is proved load-bearing by a mutation that patches THIS source in a
243
+ // disposable copy. An earlier revision guarded them with const-true flags and
244
+ // this sentence described those; the flags were removed because a
245
+ // defect-restoring branch resident in the published tree is a hazard, and the
246
+ // sentence outlived them by one commit.
247
+ if (refusal) {
248
+ warnings.push(`verdicts NOT computed β€” the comparison was REFUSED before any verdict was formed: ${refusal.reason}`);
249
+ }
250
+ if (belowTested.length) {
251
+ warnings.push(`verdicts NOT computed β€” ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
252
+ }
176
253
  // Cross-provider / cross-surface disclosure (Phase 6). A comparison across
177
254
  // providers is a skill-DURABILITY comparison across substrates, not model drift
178
255
  // over time; across surfaces, sampling control differs. Both are flagged so a
179
256
  // reader never mistakes one for the other (see docs/neutrality.html).
180
257
  if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β€” this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
181
258
  if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β€” sampling control differs between surfaces; compare with care.`);
259
+ // The legend for the band markers, printed whenever either side carries one.
260
+ // A two-letter marker a reader cannot decode is worse than no marker.
261
+ // GATED ON WHETHER A LABEL IS ACTUALLY RENDERED, not on whether the data
262
+ // carries a source. A legend explains markers on the page; if `bandStr` emits
263
+ // none, the legend is describing something the reader cannot see. Asked of
264
+ // `bandStr` itself rather than recomputed, so the two cannot disagree.
265
+ const anyLabelRendered = [...Object.values(aB), ...Object.values(bB)]
266
+ .some((x) => x && /\([a-z]+\)\s*$/.test(bandStr(x)));
267
+ if (anyLabelRendered) {
268
+ warnings.push('band provenance: `(generation)` is an ACROSS-DRAW spread β€” receipt spec v0.5, n generation draws per arm. `(legacy)` is a JUDGE-SAMPLE spread over a single generation, which is what v0.4 and earlier recorded. They are different statistics. The comparison is the only one the older receipt admits, and it is not like for like.');
269
+ }
182
270
  if (warnings.length) {
183
271
  L.push('> **⚠ Caveats**');
184
272
  for (const w of warnings) L.push(`> - ${w}`);
@@ -190,7 +278,20 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
190
278
  L.push(`## Headline`);
191
279
  L.push('');
192
280
  if (!measured) {
193
- L.push(`**NOT MEASURED β€” ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
281
+ // THE HEADLINE NAMES THE CAUSE THAT ACTUALLY APPLIES. `belowTested.length`
282
+ // was printed unconditionally, so a refused pair got "0 receipt(s) below
283
+ // TESTED" β€” a statement measurably false of the receipts it was given.
284
+ if (refusal) {
285
+ // The reason is a complete sentence and already ends by saying no verdict
286
+ // is asserted; prefixing that again produced "REFUSED β€” no verdict is
287
+ // asserted. the baseline arm did not reproduce…" β€” a duplicated clause and
288
+ // a lower-case sentence start, in the artifact a published report quotes.
289
+ L.push(`**REFUSED β€” ${refusal.reason}**`);
290
+ } else if (belowTested.length) {
291
+ L.push(`**NOT MEASURED β€” ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
292
+ } else {
293
+ L.push('**NOT MEASURED β€” no verdict is asserted.**');
294
+ }
194
295
  L.push('');
195
296
  } else {
196
297
  L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
@@ -218,7 +319,12 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
218
319
  L.push('');
219
320
  }
220
321
 
221
- return { markdown: L.join('\n'), perCase, headlineDelta, regressions };
322
+ return {
323
+ markdown: L.join('\n'), perCase, headlineDelta, regressions,
324
+ refused: !!refusal,
325
+ refusal_reason: refusal ? refusal.reason : null,
326
+ refusal_key: refusal ? refusal.key : null,
327
+ };
222
328
  }
223
329
 
224
330
  module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
package/lib/judge.js CHANGED
@@ -99,7 +99,12 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
99
99
  // { samples:[scores], mean, stddev, reason, judge_settings, model_id, rubric_hash }
100
100
  // `mean` Β± `stddev` is the per-case confidence band (raw spread of the N scores)
101
101
  // used by the borderline-outcome rule and per-case drift band-overlap logic.
102
- async function gradeSamples({ task, response, rubric, model, samples = 5, timeoutMs = 120000 }) {
102
+ // NO DEFAULT TIMEOUT HERE (spec 017 AC-2). This defaulted to 120000, which
103
+ // outranked the per-surface policy exactly as lib/run.js's literal did β€” so the
104
+ // JUDGE calls timed out on the api policy while running on a CLI surface, which
105
+ // is the second shadowing site and the one #007's prep session had not found.
106
+ // Passing `undefined` through lets lib/provider.js resolve the declared policy.
107
+ async function gradeSamples({ task, response, rubric, model, samples = 5, timeoutMs }) {
103
108
  const settings = judgeSettings(samples, model);
104
109
  const scores = [];
105
110
  const reasons = [];