driftproof 0.5.0 → 0.7.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,7 +1,7 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
2
  # Driftproof
3
3
 
4
- **Continuous, model-version-bound verification of agent skills.**
4
+ **A dated proof that this skill, this hash, this model, still helps.**
5
5
 
6
6
  [![driftproof](https://img.shields.io/endpoint?url=https://driftproofhq.com/badges/commit-message-conventions.json)](https://driftproofhq.com)
7
7
  &nbsp;— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
@@ -9,15 +9,17 @@
9
9
  Driftproof is an open **receipt spec** plus a **runner** that measures whether an
10
10
  agent skill actually helps — by running the skill's eval suite **with** and
11
11
  **without** the skill on a named model version, judging each case several times to
12
- get a confidence band, and emitting a signed, dated **receipt**. Diff two receipts
12
+ get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
13
13
  across model releases and you get a **drift report**.
14
14
 
15
15
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
16
16
  format; it does not invent its own.
17
17
 
18
- 📊 **Four published reports** (each re-derived from committed receipts, nothing
19
- hand-entered), spanning three report types that share one band-based, floor-gated
20
- verdict rule and differ only in what moves underneath the skill:
18
+ 📊 **Seven published reports** (each re-derived from committed receipts, nothing
19
+ hand-entered), spanning six published report types, the newest being instrument
20
+ re-measurement. All seven share one band-based, floor-gated verdict rule and
21
+ differ in what moves underneath the skill — or, in the value report, in which
22
+ axes are measured; or, in the newest, in the instrument itself:
21
23
 
22
24
  - **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
23
25
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
@@ -31,10 +33,41 @@ verdict rule and differ only in what moves underneath the skill:
31
33
  `claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
32
34
  3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
33
35
  survives the frontier tier.
36
+ - **[Report #005](https://driftproofhq.com/reports/005/)** — *value*: what a skill
37
+ *costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
38
+ three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) —
39
+ 14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
40
+ saving instead, having improved quality while reducing cost.
41
+ *Three of those cells carry an amendment (v1.1, applied when Report #006
42
+ published): their lifts rest on single-draw baselines since shown
43
+ unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
44
+ substrate-disagreement findings below are unaffected.*
45
+ - **[Report #006](https://driftproofhq.com/reports/006/)** — *revision drift*: the pinned skill revision against the one
46
+ upstream ships today, on a held substrate. **The reuse premise was tested and
47
+ refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
48
+ control. A 120-call probe found generation-level sampling noise 3.2× and 7.5×
49
+ larger than the judge-level noise this instrument actually samples — enough to
50
+ account for every gap the controls saw without any other cause being
51
+ established — and the receipt spec gains generation sampling as a result. **No
52
+ cause is asserted**; the control proves non-reproduction and cannot say why.
53
+ *The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
54
+ - **[Report #007](https://driftproofhq.com/reports/007/)** — *instrument
55
+ re-measurement*: the three cells Report #005 published for these skills, run
56
+ again with the generation sampled adaptively instead of once and with the call
57
+ timeout the surface policy declares. **No cell separates**: every lift is
58
+ smaller than its own band, and the 21 cases read 3 improved, 0 regressed,
59
+ 18 no effect, 0 not measured. Five of six comparisons against the archive were
60
+ **refused** on their baseline control. The instrument defect it reports is its
61
+ own: a declared 300 s timeout had been shadowed by a `120000` literal since
62
+ 2026-07-27, and the truncated run it caused measured *less* variance than the
63
+ clean re-run, which is the direction that flatters an instrument. Both runs are
64
+ published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
34
65
 
35
66
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
36
- skills](https://driftproofhq.com/writing/three-releases/)**, reads the first three
37
- reports together.
67
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all seven reports
68
+ together: what moves underneath a skill, what the skill costs to run, and what a
69
+ corrected instrument did to three published results. Revised 2026-09-01; every
70
+ figure in it is gate-checked against the report page it cites.
38
71
 
39
72
  ## Why
40
73
 
@@ -46,7 +79,7 @@ re-checks.
46
79
 
47
80
  **Verdicts age because the substrate moves.** Driftproof exists to keep the
48
81
  verdict current: cheap, repeatable, hash-stamped measurements bound to a specific
49
- model version, so "does this skill still help?" has a dated, reproducible answer
82
+ model version, so "does this skill still help?" has a dated, verifiable answer
50
83
  instead of a stale one.
51
84
 
52
85
  The hard part isn't running an eval once — it's making the number **credible
@@ -55,6 +88,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
55
88
  the judge and report confidence bands**, and to **only claim a regression when the
56
89
  bands don't overlap**. A tool that cries wolf is worse than no tool.
57
90
 
91
+ **A verdict without a price is half an answer.** The same receipts price the
92
+ marginal cost of a skill firing, and Report #005 found the dominant cost driver is
93
+ not the skill's own text but the input it causes the model to pull in: across those
94
+ 30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
95
+ tracks it at only `r = +0.33`, and one 738-token skill drew 34× its own size in
96
+ extra input. Identical token deltas also price very differently across substrates —
97
+ the same skill at near-identical deltas costs 3.3× more on `claude-fable-5` than on
98
+ `claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
99
+
58
100
  ## Quickstart — receipt for your own skill in ~10 minutes
59
101
 
60
102
  You need Node ≥ 22 and an `ANTHROPIC_API_KEY`.
@@ -70,8 +112,10 @@ npx driftproof init my-skill
70
112
  export CLAUDE_PROVIDER=api
71
113
  read -rsp "Anthropic API key: " ANTHROPIC_API_KEY && export ANTHROPIC_API_KEY
72
114
 
73
- # 4. Run the suite with a hard $2 budget (the run is refused up front if it would exceed it)
74
- npx driftproof run my-skill --models claude-haiku-4-5 --max-usd 2
115
+ # 4. Run the suite. The shipped defaults (DEV_MAX_USD / DEV_MAX_CALLS in config.js)
116
+ # refuse the run up front if the projection exceeds either; --max-usd and
117
+ # --max-calls override them.
118
+ npx driftproof run my-skill --models claude-haiku-4-5
75
119
 
76
120
  # 5. Read the receipt + human summary written to ./receipts/
77
121
  cat receipts/*.summary.md
@@ -86,6 +130,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
86
130
  minimum real move required before a change counts as more than noise — band
87
131
  separation *plus* a floor-sized delta, never either alone.
88
132
 
133
+ **What the band does not cover.** A verdict rests on **one generation draw per
134
+ arm**: the band is the spread of the *judge* re-scoring that single response, not
135
+ the spread of the model writing a different one. Report #006 measured the second
136
+ directly and found it larger — draw-to-draw spread up to **sd 0.186** on the 0–1
137
+ scale, against judge-level noise several times smaller. So treat a surprising
138
+ single-run verdict as **provisional and worth re-running** before you act on it.
139
+ Generation sampling lands in the next receipt spec; until it does, this is a
140
+ limit of the instrument, stated rather than implied.
141
+
89
142
  ### Install
90
143
 
91
144
  ```bash
@@ -148,11 +201,16 @@ The surface and judge settings used are recorded in every receipt.
148
201
 
149
202
  ### Cost guard
150
203
 
151
- Sampling multiplies calls. Each case costs `2 + 2 × samples` model calls
152
- (two generations, each judged `samples` times) — 12 calls per case at the default
153
- `--samples 5`. Driftproof **projects the whole run up front, prints the count, and
154
- refuses before spending anything** if it would exceed `--max-calls` (default 200).
155
- The default model list is `haiku` only.
204
+ Sampling multiplies calls on two axes, and the second one is easy to miss. Each
205
+ case costs `SAMPLING.max × (2 + 2 × samples)` model calls: two arms, each **drawn
206
+ up to `SAMPLING.max` times** (generation sampling, receipt spec v0.5), and every
207
+ draw judged `samples` times. Read at a single draw, that formula understates a run
208
+ by the whole draw factor — which is how the shipped caps came to sit a release
209
+ behind the estimator. Driftproof **projects the whole run up front, prints the
210
+ count, and refuses before spending anything** if the projection exceeds the
211
+ per-model call cap (`DEV_MAX_CALLS`) or the dollar budget (`DEV_MAX_USD`) — both
212
+ declared with their derivation in [`config.js`](config.js), both overridable with
213
+ `--max-calls` / `--max-usd`. The default model list is `haiku` only.
156
214
 
157
215
  ## Receipt anatomy
158
216
 
@@ -162,7 +220,7 @@ A receipt is the unit of evidence — one JSON document conforming to
162
220
 
163
221
  ```jsonc
164
222
  {
165
- "schema_version": "0.4",
223
+ "schema_version": "0.5",
166
224
  "skill": { "name": "commit-message-conventions", "version": "0.2.0",
167
225
  "content_hash": "…sha256 over SKILL.md + bundled files…" },
168
226
  "suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
@@ -171,7 +229,7 @@ A receipt is the unit of evidence — one JSON document conforming to
171
229
  "model_release_date": "2025-10-01",
172
230
  "provider": "anthropic",
173
231
  "surface": "claude-cli",
174
- "runner_version": "0.5.0",
232
+ "runner_version": "0.7.1",
175
233
  "date_utc": "2026-07-27T…Z",
176
234
  "registry": "registered",
177
235
  "transcripts": "hashes-only",
@@ -194,6 +252,12 @@ A receipt is the unit of evidence — one JSON document conforming to
194
252
  },
195
253
  "comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
196
254
  "delta": 0.39, "delta_uncertainty": 0.036 },
255
+ // v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
256
+ // freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
257
+ // holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
258
+ // mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
259
+ // skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
260
+ // output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
197
261
  "verification_level": "TESTED",
198
262
  "receipt_hash": "…sha256 of the canonical receipt with this field removed…"
199
263
  }
@@ -209,6 +273,12 @@ Key ideas:
209
273
  - **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
210
274
  - **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
211
275
  `TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
276
+ - **`economics`** is *derived*, never a second measurement: the token delta is the
277
+ durable fact, and the dollars are exactly those tokens at the rates frozen into
278
+ `run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
279
+ Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
280
+ figure (`judge_excluded` is `const true`) — measuring the skill is our cost, not
281
+ the skill's.
212
282
  - **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
213
283
  signature — see the spec's open questions).
214
284
 
@@ -219,7 +289,7 @@ receipts. The rule that keeps it honest: a **regression** (or improvement) is
219
289
  claimed **only when the two bands do not overlap**. Overlapping bands are reported
220
290
  as **within noise** and never counted as a regression.
221
291
 
222
- ## Continuous verification in CI (GitHub Action + badge)
292
+ ## Verification in CI (GitHub Action + badge)
223
293
 
224
294
  Wire drift detection into a repo so the skill is re-checked on every push and when
225
295
  the model underneath it changes.
@@ -233,12 +303,13 @@ jobs:
233
303
  runs-on: ubuntu-latest
234
304
  steps:
235
305
  - uses: actions/checkout@v4
236
- - uses: driftproofhq/driftproof@v0.5.0
306
+ - uses: driftproofhq/driftproof@v0.7.1
237
307
  with:
238
308
  skill-dir: skills/my-skill
239
309
  models: claude-haiku-4-5
240
- max-usd: '2'
241
310
  api-key: ${{ secrets.ANTHROPIC_API_KEY }}
311
+ # max-usd: <n> # override the dollar budget (default: DEV_MAX_USD in config.js)
312
+ # max-calls: <n> # override the per-model call cap (default: DEV_MAX_CALLS)
242
313
  # fail-on-regression: 'true' # (default) fail the job if the skill REGRESSED
243
314
  ```
244
315
 
@@ -256,7 +327,7 @@ runs).
256
327
  JSON object. Commit it somewhere public and point a shields endpoint URL at it:
257
328
 
258
329
  ```bash
259
- npx driftproof run skills/my-skill --models claude-haiku-4-5 --max-usd 2 --out receipts
330
+ npx driftproof run skills/my-skill --models claude-haiku-4-5 --out receipts
260
331
  npx driftproof badge receipts/*.json --out badges/my-skill.json
261
332
  git add badges/my-skill.json && git commit -m "chore: driftproof badge"
262
333
  ```
@@ -272,11 +343,14 @@ site, so it reflects a real dated run, not a hand-set color.
272
343
 
273
344
  ## Reports
274
345
 
275
- Four reports are published, spanning three report types (see the roll at the top
276
- of this README, and [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome
277
- every report inherits). Each report and every verdict in it are **re-derived from
278
- the receipts** committed under [`receipts/`](receipts/)
279
- (`receipts/report-001/` … `receipts/report-004/`) — nothing is hand-entered.
346
+ Seven reports are published, spanning six report types. A report page lives at a
347
+ draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
348
+ `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
349
+ (see the roll at the top of this README, and
350
+ [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
351
+ Each report and every verdict in it are **re-derived from the receipts** committed
352
+ under [`receipts/`](receipts/)
353
+ (`receipts/report-001/` … `receipts/report-006/`) — nothing is hand-entered.
280
354
 
281
355
  Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
282
356
  fetched at run time from a pinned commit and verified by sha256 against
@@ -290,9 +364,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models × with/basel
290
364
  node scripts/build-report-001.js # re-derive the report from the receipts
291
365
  ```
292
366
 
293
- Reports #002–#004 have their own runners
294
- (`scripts/prepare-report-00N.js`) following the same
295
- fetch → run → re-derive shape.
367
+ Reports #002–#005 have their own runners
368
+ (`scripts/prepare-report-00N.js` — Report #005's is
369
+ [`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
370
+ same fetch → run → re-derive shape.
296
371
 
297
372
  **Model-release triggers are live**: `scripts/release-watch.js` (keyless — it
298
373
  reads the public models registry) notices a new model release, re-runs the
package/bin/driftproof CHANGED
@@ -4,11 +4,12 @@
4
4
 
5
5
  const fs = require('fs');
6
6
  const path = require('path');
7
- const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = require('../config');
7
+ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD, DEV_MAX_CALLS } = require('../config');
8
8
  const { loadSkill } = require('../lib/skill');
9
9
  const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
10
+ const { SAMPLING } = require('../lib/sampling');
10
11
  const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
11
- const { buildDriftReport } = require('../lib/diff');
12
+ const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
12
13
  const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
13
14
  const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
14
15
  const { registryStatus } = require('../lib/models');
@@ -76,7 +77,7 @@ USAGE
76
77
  ${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
77
78
  [--judge-model M] [--concurrency N] [--max-usd N]
78
79
  [--keep-transcripts] [--out DIR]
79
- ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
80
+ ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
80
81
  ${PROJECT_NAME} validate <receipt.json>
81
82
  ${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
82
83
  ${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
@@ -91,8 +92,12 @@ ENV
91
92
  NOTES
92
93
  - Default models list is 'haiku' only (cheap dev default).
93
94
  - Sampled judging: each case does 2 generations + 2×samples judge calls.
94
- Default --samples 5 → 12 calls/case; --samples 1 disables bands.
95
- - --max-calls is a hard per-run call cap (default 200). The whole run is
95
+ Default --samples 5 → 12 calls/case; --samples 1 disables bands AND the
96
+ variance ratio: with one judge sample the per-draw judge spread is zero by
97
+ construction, so the receipt records variance_ratio: null with
98
+ variance_ratio_unavailable: "single_judge_sample". Legal, but a run whose
99
+ ratio you intend to publish needs --samples 2 or more.
100
+ - --max-calls is a hard per-run call cap (default ${DEV_MAX_CALLS}). The whole run is
96
101
  projected up front and REFUSED before any call if it would exceed the cap.
97
102
  - --max-usd is a hard dollar budget (default ${DEV_MAX_USD} for dev runs). The projected
98
103
  cost is printed up front and the run is REFUSED on BOTH surfaces if it would
@@ -126,7 +131,7 @@ async function cmdRun(positional, flags) {
126
131
  const rc = loadRc(skillDir);
127
132
  const models = String(flags.models || rc.models || 'haiku').split(',').map((s) => s.trim()).filter(Boolean);
128
133
  const maxCases = flags['max-cases'] ? parseInt(flags['max-cases'], 10) : (rc.max_cases != null ? parseInt(rc.max_cases, 10) : null);
129
- const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : 200);
134
+ const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : DEV_MAX_CALLS);
130
135
  const samples = flags.samples ? parseInt(flags.samples, 10) : (rc.samples != null ? parseInt(rc.samples, 10) : DEFAULT_JUDGE_SAMPLES);
131
136
  const judgeModel = flags['judge-model'] || rc.judge_model || null;
132
137
  const concurrency = flags.concurrency ? parseInt(flags.concurrency, 10) : (rc.concurrency != null ? parseInt(rc.concurrency, 10) : 1);
@@ -137,14 +142,20 @@ async function cmdRun(positional, flags) {
137
142
 
138
143
  const skill = loadSkill(skillDir);
139
144
  const nCases = maxCases ? Math.min(maxCases, skill.suite.caseCount) : skill.suite.caseCount;
140
- const perModelCalls = projectCalls(nCases, samples);
145
+ // v0.5 draws the generation up to SAMPLING.max times per arm, so BOTH the
146
+ // printed projection and the dollar guard below must be scaled by it. Left at
147
+ // draws=1 the guard would admit a run costing up to ten times its own
148
+ // projection — the display would say one number and the runner would refuse
149
+ // quoting another. Conservative on purpose: refusing a run that would have fit
150
+ // is recoverable, overspending is not.
151
+ const perModelCalls = projectCalls(nCases, samples, SAMPLING.max);
141
152
  const totalProjected = perModelCalls * models.length;
142
153
 
143
154
  // Dollar cost guard (Week 3): project the METERED USD cost up front. On the
144
155
  // `api` surface this is real spend, so we refuse if it would exceed --max-usd.
145
156
  // On `claude-cli` the metered spend is $0 (subscription); the figure is the
146
157
  // hypothetical "if run on the metered API" cost — printed, never blocks.
147
- const cost = estimateRunCostUSD({ caseCount: nCases, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
158
+ const cost = estimateRunCostUSD({ caseCount: nCases, draws: SAMPLING.max, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
148
159
  // Surface is per-model now (a run may mix an Anthropic and an OpenAI target).
149
160
  const surfaces = [...new Set(models.map((m) => surfaceForModel(m)))];
150
161
  const surface = surfaces.join(', ');
@@ -250,9 +261,32 @@ function cmdDiff(positional, flags) {
250
261
  if (!verifyReceiptHash(r)) console.error(` ⚠ ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
251
262
  }
252
263
 
253
- const labelA = dateStamp(a.run.date_utc);
254
- const labelB = dateStamp(b.run.date_utc);
255
- const { markdown } = buildDriftReport(a, b, { labelA, labelB });
264
+ // --mode revision inverts the axis: the skill text is the variable under test
265
+ // and the substrate is the control. The fields release drift merely warns about
266
+ // are preconditions here, so a pair that is not a revision pair is REFUSED
267
+ // (exit 6) rather than rendered with a caveat nobody reads. A differing model
268
+ // would be release drift wearing a revision label — the one confound this mode
269
+ // exists to exclude — and an EQUAL content_hash has no revision to measure.
270
+ const mode = flags.mode || 'release';
271
+ if (mode !== 'release' && mode !== 'revision') {
272
+ console.error(`unknown --mode "${mode}" — supported: release, revision`);
273
+ process.exit(2);
274
+ }
275
+ if (mode === 'revision') {
276
+ const problem = revisionPairProblem(a, b);
277
+ if (problem) {
278
+ const why = problem === 'skill.content_hash'
279
+ ? 'the two receipts carry the SAME skill.content_hash — there is no revision between them to measure'
280
+ : `${problem} differs between the two receipts — revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
281
+ console.error(` ✗ REFUSED (--mode revision): ${why}.`);
282
+ console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
283
+ process.exit(6);
284
+ }
285
+ }
286
+
287
+ const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
288
+ const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
289
+ const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
256
290
 
257
291
  if (flags.out) {
258
292
  fs.writeFileSync(path.resolve(flags.out), markdown);
@@ -292,7 +326,7 @@ function cmdInit(positional) {
292
326
  console.log(`\nNext:
293
327
  1. Edit ${rel(path.join(dir, 'SKILL.md'))} with your skill's instructions.
294
328
  2. Edit ${rel(path.join(dir, 'evals', 'evals.json'))} — replace the 3 example cases (rubrics anchored at 0.80).
295
- 3. driftproof run ${rel(dir)} --models claude-haiku-4-5 --max-usd 2
329
+ 3. driftproof run ${rel(dir)} --models claude-haiku-4-5
296
330
  See AUTHORING.md for how to write a fair suite.`);
297
331
  }
298
332
 
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.5.0';
12
+ const RUNNER_VERSION = '0.7.1';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
@@ -29,7 +29,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
29
29
  // at run time so derived dollars stay reproducible), and the derived
30
30
  // `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
31
31
  // v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
32
- const RECEIPT_SCHEMA_VERSION = '0.4';
32
+ const RECEIPT_SCHEMA_VERSION = '0.5';
33
33
 
34
34
  // Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
35
35
  // these. The projection is refused before any call if it exceeds the cap, on
@@ -37,8 +37,33 @@ const RECEIPT_SCHEMA_VERSION = '0.4';
37
37
  // dev — an interactive `driftproof run`
38
38
  // report — a full Report-#001-style suite run (scripts/run-report-001.js)
39
39
  // trigger — a release-trigger-initiated prepare-report run
40
- const DEV_MAX_USD = 2;
41
- const REPORT_MAX_USD = 40;
40
+ // Recalibrated 2026-09-01 (spec 018). BOTH literals below were set against a
41
+ // draws=1 projection and were never rescaled when spec 014 (5e08ba2) made the
42
+ // runner draw the generation up to SAMPLING.max times per arm and spec 016 made
43
+ // estimateRunCostUSD require the draw factor. The projection grew 10x; these did
44
+ // not, so `npx driftproof run` refused every suite of two or more cases and the
45
+ // v0.7.0 Action self-test aborted with exit 3. REPORT_MAX_USD was raised 40->300
46
+ // on 2026-08-31 for exactly this reason; that pass missed the dev path, which is
47
+ // the one the Action and every npx user take.
48
+ //
49
+ // DERIVATION — the headroom the pre-sampling defaults carried is preserved, not
50
+ // widened. Bundled 10-case example, 5 judge samples, per-case factor 2+2*5 = 12:
51
+ // draws=1 -> 120 calls, $0.3525 headroom 200/120 = 1.67x, $2/$0.3525 = 5.67x
52
+ // draws=10 -> 1200 calls, $3.5250 1200 * 1.67 = 2000, $3.5250 * 5.67 = $20.00
53
+ // A 2000-call cap under draws=10 is exactly as tight as 200 was under draws=1.
54
+ // LITERALS ON PURPOSE (spec 018 AC-2): a default derived from the suite in hand
55
+ // can never fire, which retires the guard instead of recalibrating it.
56
+ const DEV_MAX_USD = 20;
57
+ const DEV_MAX_CALLS = 2000;
58
+ // Raised from $40 on 2026-08-31 by the repository owner, in daylight, recorded in
59
+ // DECISIONS. The run did not get more expensive — the projection got honest:
60
+ // spec 016 made estimateRunCostUSD require a draw factor, and v0.5 has drawn the
61
+ // generation up to SAMPLING.max times per arm since spec 014, so the old cap had
62
+ // been passing on a figure up to ten times too small. Set above the ceiling this
63
+ // instrument can currently reach (#007's measured basis at the ceiling is
64
+ // $268.83) rather than just above today's staging figure, because a cap tripped
65
+ // by the next honest run teaches everyone to nudge it.
66
+ const REPORT_MAX_USD = 300;
42
67
  const TRIGGER_MAX_USD = 25;
43
68
 
44
69
  // Default number of judge samples per case in sampled mode. Sampling is what
@@ -56,6 +81,24 @@ const DEFAULT_JUDGE_SAMPLES = 5;
56
81
  // spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
57
82
  const EFFECT_FLOOR = 0.05;
58
83
 
84
+ // ── generation sampling (receipt spec v0.5) ─────────────────────────────────
85
+ //
86
+ // Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
87
+ // instrument sampled only the judge. These are the policy that samples the
88
+ // other axis. Constants, not literals in the runner: lib/sampling.js reads
89
+ // them, so moving one here moves the policy, and the gate asserts that by
90
+ // value rather than by grepping for a name.
91
+ //
92
+ // MIN is 3 because two draws give an sd that is barely a measurement and one
93
+ // gives none at all. MAX is 10: #006's probe used 20 by hand and found the
94
+ // shape at well under half of that, and a per-case ceiling bounds the spend.
95
+ // The SD threshold is the effect floor — a spread wider than the smallest move
96
+ // the verdict rule will call real is exactly when more draws are owed.
97
+ const GENERATION_SAMPLES_MIN = 3;
98
+ const GENERATION_SAMPLES_MAX = 10;
99
+ const GENERATION_SD_THRESHOLD = EFFECT_FLOOR;
100
+ const GENERATION_STABILITY_EPS = 0.01;
101
+
59
102
  // Report #002 is the first CROSS-PROVIDER report: the same suites, the same fixed
60
103
  // Haiku judge, run on two substrates — a Claude flagship and a GPT flagship. The
61
104
  // GPT flagship is a config constant (not hard-coded across scripts) so a future
@@ -90,7 +133,8 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
90
133
 
91
134
  module.exports = {
92
135
  PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
93
- EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
136
+ EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
137
+ GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
94
138
  REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
95
139
  REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
96
140
  REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
package/lib/canary.js ADDED
@@ -0,0 +1,27 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ 'use strict';
3
+
4
+ // Suite canary — receipt spec v0.5.
5
+ //
6
+ // Terminal-Bench 4.0 records a canary string so that a benchmark appearing in
7
+ // training data is detectable. Adopted, translated: our unit of publication is
8
+ // the SUITE, so the canary is per suite, derived from the suite's identity and
9
+ // its case ids rather than randomly assigned, so the same suite yields the same
10
+ // canary on every machine and no registry has to be kept.
11
+ //
12
+ // It is a detection aid, not a control: a canary tells you a suite leaked, it
13
+ // does not stop the leak, and it cannot prove the absence of one.
14
+
15
+ const crypto = require('crypto');
16
+
17
+ const NAMESPACE = 'driftproof/suite-canary/v1';
18
+
19
+ function suiteCanary(suite) {
20
+ const id = String((suite && (suite.id || suite.name)) || '');
21
+ const caseIds = ((suite && suite.cases) || []).map((c) => String((c && c.id) || '')).sort();
22
+ const h = crypto.createHash('sha256').update(`${NAMESPACE}\n${id}\n${caseIds.join('\n')}`).digest('hex');
23
+ // Formatted as a GUID so it is recognisable as one in a corpus scan.
24
+ return [h.slice(0, 8), h.slice(8, 12), h.slice(12, 16), h.slice(16, 20), h.slice(20, 32)].join('-');
25
+ }
26
+
27
+ module.exports = { suiteCanary, NAMESPACE };
package/lib/cost.js CHANGED
@@ -58,7 +58,25 @@ function perCallCostUSD(modelId, kind) {
58
58
  // generated on each target model, judged `samples` times each on `judgeModel`.
59
59
  //
60
60
  // returns { totalUSD, perModel: [{ model, usd }], judgeUSD, genUSD, assumptions }
61
- function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
61
+ // THE DRAW FACTOR IS REQUIRED AND HAS NO DEFAULT (spec 016 AC-7).
62
+ //
63
+ // v0.5 draws the generation up to `SAMPLING.max` times per arm, so a run costs
64
+ // its one-draw estimate times the number of draws. `projectCalls` learned this in
65
+ // spec 014 with a `draws` argument DEFAULTING TO 1 "so every existing caller
66
+ // projects exactly what it projected before" — and every existing caller then
67
+ // went on projecting a tenth of the run, silently, for two more loops. Spec 015
68
+ // corrected the CALL projection in three report scripts and left the DOLLAR
69
+ // projection at one draw everywhere, which is the figure a human actually reads
70
+ // before authorising a paid run.
71
+ //
72
+ // A default is what made that invisible, so there is none. Omitting the factor
73
+ // throws, which turns a silent understatement into a loud stop — and every call
74
+ // site has to say what it means, including the ones that legitimately mean 1.
75
+ function estimateRunCostUSD({ caseCount, samples, models, judgeModel, draws }) {
76
+ if (!Number.isFinite(draws) || draws < 1) {
77
+ throw new Error('estimateRunCostUSD: `draws` is required and must be >= 1 — pass SAMPLING.max to project a v0.5 run, or 1 to price a single draw deliberately. It is not defaulted, because a default is how the dollar projection stayed at one draw through two loops.');
78
+ }
79
+ caseCount = caseCount * draws;
62
80
  const judgePrice = priceFor(judgeModel);
63
81
  const perModel = [];
64
82
  let judgeUSD = 0;
@@ -80,7 +98,7 @@ function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
80
98
  perModel,
81
99
  judgeUSD: round4(judgeUSD),
82
100
  genUSD: round4(genUSD),
83
- assumptions: { tokens: TOKENS, judgeModel, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
101
+ assumptions: { tokens: TOKENS, judgeModel, draws, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
84
102
  };
85
103
  }
86
104