driftproof 0.5.0 → 0.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +104 -29
- package/bin/driftproof +46 -12
- package/config.js +49 -5
- package/lib/canary.js +27 -0
- package/lib/cost.js +20 -2
- package/lib/diff.js +172 -15
- package/lib/hygiene.js +113 -0
- package/lib/judge.js +6 -1
- package/lib/receipt.js +76 -4
- package/lib/reuse.js +134 -0
- package/lib/revision.js +169 -0
- package/lib/run.js +165 -45
- package/lib/sampling.js +110 -0
- package/lib/stats.js +16 -1
- package/lib/value.js +27 -2
- package/package.json +2 -2
- package/spec/RECEIPT.md +102 -3
- package/spec/receipt.schema.json +360 -3
- package/spec/receipt.v0.4.schema.json +961 -0
package/README.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
2
|
# Driftproof
|
|
3
3
|
|
|
4
|
-
**
|
|
4
|
+
**A dated proof that this skill, this hash, this model, still helps.**
|
|
5
5
|
|
|
6
6
|
[](https://driftproofhq.com)
|
|
7
7
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
@@ -9,15 +9,17 @@
|
|
|
9
9
|
Driftproof is an open **receipt spec** plus a **runner** that measures whether an
|
|
10
10
|
agent skill actually helps — by running the skill's eval suite **with** and
|
|
11
11
|
**without** the skill on a named model version, judging each case several times to
|
|
12
|
-
get a confidence band, and emitting a
|
|
12
|
+
get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
|
|
13
13
|
across model releases and you get a **drift report**.
|
|
14
14
|
|
|
15
15
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
16
16
|
format; it does not invent its own.
|
|
17
17
|
|
|
18
|
-
📊 **
|
|
19
|
-
hand-entered), spanning
|
|
20
|
-
|
|
18
|
+
📊 **Seven published reports** (each re-derived from committed receipts, nothing
|
|
19
|
+
hand-entered), spanning six published report types, the newest being instrument
|
|
20
|
+
re-measurement. All seven share one band-based, floor-gated verdict rule and
|
|
21
|
+
differ in what moves underneath the skill — or, in the value report, in which
|
|
22
|
+
axes are measured; or, in the newest, in the instrument itself:
|
|
21
23
|
|
|
22
24
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
23
25
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
|
|
@@ -31,10 +33,41 @@ verdict rule and differ only in what moves underneath the skill:
|
|
|
31
33
|
`claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
|
|
32
34
|
3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
|
|
33
35
|
survives the frontier tier.
|
|
36
|
+
- **[Report #005](https://driftproofhq.com/reports/005/)** — *value*: what a skill
|
|
37
|
+
*costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
|
|
38
|
+
three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) —
|
|
39
|
+
14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
|
|
40
|
+
saving instead, having improved quality while reducing cost.
|
|
41
|
+
*Three of those cells carry an amendment (v1.1, applied when Report #006
|
|
42
|
+
published): their lifts rest on single-draw baselines since shown
|
|
43
|
+
unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
|
|
44
|
+
substrate-disagreement findings below are unaffected.*
|
|
45
|
+
- **[Report #006](https://driftproofhq.com/reports/006/)** — *revision drift*: the pinned skill revision against the one
|
|
46
|
+
upstream ships today, on a held substrate. **The reuse premise was tested and
|
|
47
|
+
refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
|
|
48
|
+
control. A 120-call probe found generation-level sampling noise 3.2× and 7.5×
|
|
49
|
+
larger than the judge-level noise this instrument actually samples — enough to
|
|
50
|
+
account for every gap the controls saw without any other cause being
|
|
51
|
+
established — and the receipt spec gains generation sampling as a result. **No
|
|
52
|
+
cause is asserted**; the control proves non-reproduction and cannot say why.
|
|
53
|
+
*The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
|
|
54
|
+
- **[Report #007](https://driftproofhq.com/reports/007/)** — *instrument
|
|
55
|
+
re-measurement*: the three cells Report #005 published for these skills, run
|
|
56
|
+
again with the generation sampled adaptively instead of once and with the call
|
|
57
|
+
timeout the surface policy declares. **No cell separates**: every lift is
|
|
58
|
+
smaller than its own band, and the 21 cases read 3 improved, 0 regressed,
|
|
59
|
+
18 no effect, 0 not measured. Five of six comparisons against the archive were
|
|
60
|
+
**refused** on their baseline control. The instrument defect it reports is its
|
|
61
|
+
own: a declared 300 s timeout had been shadowed by a `120000` literal since
|
|
62
|
+
2026-07-27, and the truncated run it caused measured *less* variance than the
|
|
63
|
+
clean re-run, which is the direction that flatters an instrument. Both runs are
|
|
64
|
+
published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
|
|
34
65
|
|
|
35
66
|
✍️ The launch essay, **[Three model releases later: what actually happens to agent
|
|
36
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads
|
|
37
|
-
|
|
67
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all seven reports
|
|
68
|
+
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
69
|
+
corrected instrument did to three published results. Revised 2026-09-01; every
|
|
70
|
+
figure in it is gate-checked against the report page it cites.
|
|
38
71
|
|
|
39
72
|
## Why
|
|
40
73
|
|
|
@@ -46,7 +79,7 @@ re-checks.
|
|
|
46
79
|
|
|
47
80
|
**Verdicts age because the substrate moves.** Driftproof exists to keep the
|
|
48
81
|
verdict current: cheap, repeatable, hash-stamped measurements bound to a specific
|
|
49
|
-
model version, so "does this skill still help?" has a dated,
|
|
82
|
+
model version, so "does this skill still help?" has a dated, verifiable answer
|
|
50
83
|
instead of a stale one.
|
|
51
84
|
|
|
52
85
|
The hard part isn't running an eval once — it's making the number **credible
|
|
@@ -55,6 +88,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
|
|
|
55
88
|
the judge and report confidence bands**, and to **only claim a regression when the
|
|
56
89
|
bands don't overlap**. A tool that cries wolf is worse than no tool.
|
|
57
90
|
|
|
91
|
+
**A verdict without a price is half an answer.** The same receipts price the
|
|
92
|
+
marginal cost of a skill firing, and Report #005 found the dominant cost driver is
|
|
93
|
+
not the skill's own text but the input it causes the model to pull in: across those
|
|
94
|
+
30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
|
|
95
|
+
tracks it at only `r = +0.33`, and one 738-token skill drew 34× its own size in
|
|
96
|
+
extra input. Identical token deltas also price very differently across substrates —
|
|
97
|
+
the same skill at near-identical deltas costs 3.3× more on `claude-fable-5` than on
|
|
98
|
+
`claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
|
|
99
|
+
|
|
58
100
|
## Quickstart — receipt for your own skill in ~10 minutes
|
|
59
101
|
|
|
60
102
|
You need Node ≥ 22 and an `ANTHROPIC_API_KEY`.
|
|
@@ -70,8 +112,10 @@ npx driftproof init my-skill
|
|
|
70
112
|
export CLAUDE_PROVIDER=api
|
|
71
113
|
read -rsp "Anthropic API key: " ANTHROPIC_API_KEY && export ANTHROPIC_API_KEY
|
|
72
114
|
|
|
73
|
-
# 4. Run the suite
|
|
74
|
-
|
|
115
|
+
# 4. Run the suite. The shipped defaults (DEV_MAX_USD / DEV_MAX_CALLS in config.js)
|
|
116
|
+
# refuse the run up front if the projection exceeds either; --max-usd and
|
|
117
|
+
# --max-calls override them.
|
|
118
|
+
npx driftproof run my-skill --models claude-haiku-4-5
|
|
75
119
|
|
|
76
120
|
# 5. Read the receipt + human summary written to ./receipts/
|
|
77
121
|
cat receipts/*.summary.md
|
|
@@ -86,6 +130,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
|
|
|
86
130
|
minimum real move required before a change counts as more than noise — band
|
|
87
131
|
separation *plus* a floor-sized delta, never either alone.
|
|
88
132
|
|
|
133
|
+
**What the band does not cover.** A verdict rests on **one generation draw per
|
|
134
|
+
arm**: the band is the spread of the *judge* re-scoring that single response, not
|
|
135
|
+
the spread of the model writing a different one. Report #006 measured the second
|
|
136
|
+
directly and found it larger — draw-to-draw spread up to **sd 0.186** on the 0–1
|
|
137
|
+
scale, against judge-level noise several times smaller. So treat a surprising
|
|
138
|
+
single-run verdict as **provisional and worth re-running** before you act on it.
|
|
139
|
+
Generation sampling lands in the next receipt spec; until it does, this is a
|
|
140
|
+
limit of the instrument, stated rather than implied.
|
|
141
|
+
|
|
89
142
|
### Install
|
|
90
143
|
|
|
91
144
|
```bash
|
|
@@ -148,11 +201,16 @@ The surface and judge settings used are recorded in every receipt.
|
|
|
148
201
|
|
|
149
202
|
### Cost guard
|
|
150
203
|
|
|
151
|
-
Sampling multiplies calls
|
|
152
|
-
(
|
|
153
|
-
|
|
154
|
-
|
|
155
|
-
|
|
204
|
+
Sampling multiplies calls on two axes, and the second one is easy to miss. Each
|
|
205
|
+
case costs `SAMPLING.max × (2 + 2 × samples)` model calls: two arms, each **drawn
|
|
206
|
+
up to `SAMPLING.max` times** (generation sampling, receipt spec v0.5), and every
|
|
207
|
+
draw judged `samples` times. Read at a single draw, that formula understates a run
|
|
208
|
+
by the whole draw factor — which is how the shipped caps came to sit a release
|
|
209
|
+
behind the estimator. Driftproof **projects the whole run up front, prints the
|
|
210
|
+
count, and refuses before spending anything** if the projection exceeds the
|
|
211
|
+
per-model call cap (`DEV_MAX_CALLS`) or the dollar budget (`DEV_MAX_USD`) — both
|
|
212
|
+
declared with their derivation in [`config.js`](config.js), both overridable with
|
|
213
|
+
`--max-calls` / `--max-usd`. The default model list is `haiku` only.
|
|
156
214
|
|
|
157
215
|
## Receipt anatomy
|
|
158
216
|
|
|
@@ -162,7 +220,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
162
220
|
|
|
163
221
|
```jsonc
|
|
164
222
|
{
|
|
165
|
-
"schema_version": "0.
|
|
223
|
+
"schema_version": "0.5",
|
|
166
224
|
"skill": { "name": "commit-message-conventions", "version": "0.2.0",
|
|
167
225
|
"content_hash": "…sha256 over SKILL.md + bundled files…" },
|
|
168
226
|
"suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
|
|
@@ -171,7 +229,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
171
229
|
"model_release_date": "2025-10-01",
|
|
172
230
|
"provider": "anthropic",
|
|
173
231
|
"surface": "claude-cli",
|
|
174
|
-
"runner_version": "0.
|
|
232
|
+
"runner_version": "0.7.1",
|
|
175
233
|
"date_utc": "2026-07-27T…Z",
|
|
176
234
|
"registry": "registered",
|
|
177
235
|
"transcripts": "hashes-only",
|
|
@@ -194,6 +252,12 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
194
252
|
},
|
|
195
253
|
"comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
|
|
196
254
|
"delta": 0.39, "delta_uncertainty": 0.036 },
|
|
255
|
+
// v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
|
|
256
|
+
// freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
|
|
257
|
+
// holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
|
|
258
|
+
// mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
|
|
259
|
+
// skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
|
|
260
|
+
// output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
|
|
197
261
|
"verification_level": "TESTED",
|
|
198
262
|
"receipt_hash": "…sha256 of the canonical receipt with this field removed…"
|
|
199
263
|
}
|
|
@@ -209,6 +273,12 @@ Key ideas:
|
|
|
209
273
|
- **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
|
|
210
274
|
- **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
|
|
211
275
|
`TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
|
|
276
|
+
- **`economics`** is *derived*, never a second measurement: the token delta is the
|
|
277
|
+
durable fact, and the dollars are exactly those tokens at the rates frozen into
|
|
278
|
+
`run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
|
|
279
|
+
Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
|
|
280
|
+
figure (`judge_excluded` is `const true`) — measuring the skill is our cost, not
|
|
281
|
+
the skill's.
|
|
212
282
|
- **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
|
|
213
283
|
signature — see the spec's open questions).
|
|
214
284
|
|
|
@@ -219,7 +289,7 @@ receipts. The rule that keeps it honest: a **regression** (or improvement) is
|
|
|
219
289
|
claimed **only when the two bands do not overlap**. Overlapping bands are reported
|
|
220
290
|
as **within noise** and never counted as a regression.
|
|
221
291
|
|
|
222
|
-
##
|
|
292
|
+
## Verification in CI (GitHub Action + badge)
|
|
223
293
|
|
|
224
294
|
Wire drift detection into a repo so the skill is re-checked on every push and when
|
|
225
295
|
the model underneath it changes.
|
|
@@ -233,12 +303,13 @@ jobs:
|
|
|
233
303
|
runs-on: ubuntu-latest
|
|
234
304
|
steps:
|
|
235
305
|
- uses: actions/checkout@v4
|
|
236
|
-
- uses: driftproofhq/driftproof@v0.
|
|
306
|
+
- uses: driftproofhq/driftproof@v0.7.1
|
|
237
307
|
with:
|
|
238
308
|
skill-dir: skills/my-skill
|
|
239
309
|
models: claude-haiku-4-5
|
|
240
|
-
max-usd: '2'
|
|
241
310
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
311
|
+
# max-usd: <n> # override the dollar budget (default: DEV_MAX_USD in config.js)
|
|
312
|
+
# max-calls: <n> # override the per-model call cap (default: DEV_MAX_CALLS)
|
|
242
313
|
# fail-on-regression: 'true' # (default) fail the job if the skill REGRESSED
|
|
243
314
|
```
|
|
244
315
|
|
|
@@ -256,7 +327,7 @@ runs).
|
|
|
256
327
|
JSON object. Commit it somewhere public and point a shields endpoint URL at it:
|
|
257
328
|
|
|
258
329
|
```bash
|
|
259
|
-
npx driftproof run skills/my-skill --models claude-haiku-4-5 --
|
|
330
|
+
npx driftproof run skills/my-skill --models claude-haiku-4-5 --out receipts
|
|
260
331
|
npx driftproof badge receipts/*.json --out badges/my-skill.json
|
|
261
332
|
git add badges/my-skill.json && git commit -m "chore: driftproof badge"
|
|
262
333
|
```
|
|
@@ -272,11 +343,14 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
272
343
|
|
|
273
344
|
## Reports
|
|
274
345
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
the
|
|
279
|
-
(
|
|
346
|
+
Seven reports are published, spanning six report types. A report page lives at a
|
|
347
|
+
draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
|
|
348
|
+
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
349
|
+
(see the roll at the top of this README, and
|
|
350
|
+
[REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
|
|
351
|
+
Each report and every verdict in it are **re-derived from the receipts** committed
|
|
352
|
+
under [`receipts/`](receipts/)
|
|
353
|
+
(`receipts/report-001/` … `receipts/report-006/`) — nothing is hand-entered.
|
|
280
354
|
|
|
281
355
|
Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
|
|
282
356
|
fetched at run time from a pinned commit and verified by sha256 against
|
|
@@ -290,9 +364,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models × with/basel
|
|
|
290
364
|
node scripts/build-report-001.js # re-derive the report from the receipts
|
|
291
365
|
```
|
|
292
366
|
|
|
293
|
-
Reports #002–#
|
|
294
|
-
(`scripts/prepare-report-00N.js`
|
|
295
|
-
|
|
367
|
+
Reports #002–#005 have their own runners
|
|
368
|
+
(`scripts/prepare-report-00N.js` — Report #005's is
|
|
369
|
+
[`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
|
|
370
|
+
same fetch → run → re-derive shape.
|
|
296
371
|
|
|
297
372
|
**Model-release triggers are live**: `scripts/release-watch.js` (keyless — it
|
|
298
373
|
reads the public models registry) notices a new model release, re-runs the
|
package/bin/driftproof
CHANGED
|
@@ -4,11 +4,12 @@
|
|
|
4
4
|
|
|
5
5
|
const fs = require('fs');
|
|
6
6
|
const path = require('path');
|
|
7
|
-
const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = require('../config');
|
|
7
|
+
const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD, DEV_MAX_CALLS } = require('../config');
|
|
8
8
|
const { loadSkill } = require('../lib/skill');
|
|
9
9
|
const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
|
|
10
|
+
const { SAMPLING } = require('../lib/sampling');
|
|
10
11
|
const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
|
|
11
|
-
const { buildDriftReport } = require('../lib/diff');
|
|
12
|
+
const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
|
|
12
13
|
const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
|
|
13
14
|
const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
|
|
14
15
|
const { registryStatus } = require('../lib/models');
|
|
@@ -76,7 +77,7 @@ USAGE
|
|
|
76
77
|
${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
|
|
77
78
|
[--judge-model M] [--concurrency N] [--max-usd N]
|
|
78
79
|
[--keep-transcripts] [--out DIR]
|
|
79
|
-
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
|
|
80
|
+
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
|
|
80
81
|
${PROJECT_NAME} validate <receipt.json>
|
|
81
82
|
${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
|
|
82
83
|
${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
|
|
@@ -91,8 +92,12 @@ ENV
|
|
|
91
92
|
NOTES
|
|
92
93
|
- Default models list is 'haiku' only (cheap dev default).
|
|
93
94
|
- Sampled judging: each case does 2 generations + 2×samples judge calls.
|
|
94
|
-
Default --samples 5 → 12 calls/case; --samples 1 disables bands
|
|
95
|
-
|
|
95
|
+
Default --samples 5 → 12 calls/case; --samples 1 disables bands AND the
|
|
96
|
+
variance ratio: with one judge sample the per-draw judge spread is zero by
|
|
97
|
+
construction, so the receipt records variance_ratio: null with
|
|
98
|
+
variance_ratio_unavailable: "single_judge_sample". Legal, but a run whose
|
|
99
|
+
ratio you intend to publish needs --samples 2 or more.
|
|
100
|
+
- --max-calls is a hard per-run call cap (default ${DEV_MAX_CALLS}). The whole run is
|
|
96
101
|
projected up front and REFUSED before any call if it would exceed the cap.
|
|
97
102
|
- --max-usd is a hard dollar budget (default ${DEV_MAX_USD} for dev runs). The projected
|
|
98
103
|
cost is printed up front and the run is REFUSED on BOTH surfaces if it would
|
|
@@ -126,7 +131,7 @@ async function cmdRun(positional, flags) {
|
|
|
126
131
|
const rc = loadRc(skillDir);
|
|
127
132
|
const models = String(flags.models || rc.models || 'haiku').split(',').map((s) => s.trim()).filter(Boolean);
|
|
128
133
|
const maxCases = flags['max-cases'] ? parseInt(flags['max-cases'], 10) : (rc.max_cases != null ? parseInt(rc.max_cases, 10) : null);
|
|
129
|
-
const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) :
|
|
134
|
+
const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : DEV_MAX_CALLS);
|
|
130
135
|
const samples = flags.samples ? parseInt(flags.samples, 10) : (rc.samples != null ? parseInt(rc.samples, 10) : DEFAULT_JUDGE_SAMPLES);
|
|
131
136
|
const judgeModel = flags['judge-model'] || rc.judge_model || null;
|
|
132
137
|
const concurrency = flags.concurrency ? parseInt(flags.concurrency, 10) : (rc.concurrency != null ? parseInt(rc.concurrency, 10) : 1);
|
|
@@ -137,14 +142,20 @@ async function cmdRun(positional, flags) {
|
|
|
137
142
|
|
|
138
143
|
const skill = loadSkill(skillDir);
|
|
139
144
|
const nCases = maxCases ? Math.min(maxCases, skill.suite.caseCount) : skill.suite.caseCount;
|
|
140
|
-
|
|
145
|
+
// v0.5 draws the generation up to SAMPLING.max times per arm, so BOTH the
|
|
146
|
+
// printed projection and the dollar guard below must be scaled by it. Left at
|
|
147
|
+
// draws=1 the guard would admit a run costing up to ten times its own
|
|
148
|
+
// projection — the display would say one number and the runner would refuse
|
|
149
|
+
// quoting another. Conservative on purpose: refusing a run that would have fit
|
|
150
|
+
// is recoverable, overspending is not.
|
|
151
|
+
const perModelCalls = projectCalls(nCases, samples, SAMPLING.max);
|
|
141
152
|
const totalProjected = perModelCalls * models.length;
|
|
142
153
|
|
|
143
154
|
// Dollar cost guard (Week 3): project the METERED USD cost up front. On the
|
|
144
155
|
// `api` surface this is real spend, so we refuse if it would exceed --max-usd.
|
|
145
156
|
// On `claude-cli` the metered spend is $0 (subscription); the figure is the
|
|
146
157
|
// hypothetical "if run on the metered API" cost — printed, never blocks.
|
|
147
|
-
const cost = estimateRunCostUSD({ caseCount: nCases, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
|
|
158
|
+
const cost = estimateRunCostUSD({ caseCount: nCases, draws: SAMPLING.max, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
|
|
148
159
|
// Surface is per-model now (a run may mix an Anthropic and an OpenAI target).
|
|
149
160
|
const surfaces = [...new Set(models.map((m) => surfaceForModel(m)))];
|
|
150
161
|
const surface = surfaces.join(', ');
|
|
@@ -250,9 +261,32 @@ function cmdDiff(positional, flags) {
|
|
|
250
261
|
if (!verifyReceiptHash(r)) console.error(` ⚠ ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
|
|
251
262
|
}
|
|
252
263
|
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
264
|
+
// --mode revision inverts the axis: the skill text is the variable under test
|
|
265
|
+
// and the substrate is the control. The fields release drift merely warns about
|
|
266
|
+
// are preconditions here, so a pair that is not a revision pair is REFUSED
|
|
267
|
+
// (exit 6) rather than rendered with a caveat nobody reads. A differing model
|
|
268
|
+
// would be release drift wearing a revision label — the one confound this mode
|
|
269
|
+
// exists to exclude — and an EQUAL content_hash has no revision to measure.
|
|
270
|
+
const mode = flags.mode || 'release';
|
|
271
|
+
if (mode !== 'release' && mode !== 'revision') {
|
|
272
|
+
console.error(`unknown --mode "${mode}" — supported: release, revision`);
|
|
273
|
+
process.exit(2);
|
|
274
|
+
}
|
|
275
|
+
if (mode === 'revision') {
|
|
276
|
+
const problem = revisionPairProblem(a, b);
|
|
277
|
+
if (problem) {
|
|
278
|
+
const why = problem === 'skill.content_hash'
|
|
279
|
+
? 'the two receipts carry the SAME skill.content_hash — there is no revision between them to measure'
|
|
280
|
+
: `${problem} differs between the two receipts — revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
|
|
281
|
+
console.error(` ✗ REFUSED (--mode revision): ${why}.`);
|
|
282
|
+
console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
|
|
283
|
+
process.exit(6);
|
|
284
|
+
}
|
|
285
|
+
}
|
|
286
|
+
|
|
287
|
+
const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
|
|
288
|
+
const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
|
|
289
|
+
const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
|
|
256
290
|
|
|
257
291
|
if (flags.out) {
|
|
258
292
|
fs.writeFileSync(path.resolve(flags.out), markdown);
|
|
@@ -292,7 +326,7 @@ function cmdInit(positional) {
|
|
|
292
326
|
console.log(`\nNext:
|
|
293
327
|
1. Edit ${rel(path.join(dir, 'SKILL.md'))} with your skill's instructions.
|
|
294
328
|
2. Edit ${rel(path.join(dir, 'evals', 'evals.json'))} — replace the 3 example cases (rubrics anchored at 0.80).
|
|
295
|
-
3. driftproof run ${rel(dir)} --models claude-haiku-4-5
|
|
329
|
+
3. driftproof run ${rel(dir)} --models claude-haiku-4-5
|
|
296
330
|
See AUTHORING.md for how to write a fair suite.`);
|
|
297
331
|
}
|
|
298
332
|
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.7.1';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
|
@@ -29,7 +29,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
|
|
|
29
29
|
// at run time so derived dollars stay reproducible), and the derived
|
|
30
30
|
// `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
|
|
31
31
|
// v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
|
|
32
|
-
const RECEIPT_SCHEMA_VERSION = '0.
|
|
32
|
+
const RECEIPT_SCHEMA_VERSION = '0.5';
|
|
33
33
|
|
|
34
34
|
// Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
|
|
35
35
|
// these. The projection is refused before any call if it exceeds the cap, on
|
|
@@ -37,8 +37,33 @@ const RECEIPT_SCHEMA_VERSION = '0.4';
|
|
|
37
37
|
// dev — an interactive `driftproof run`
|
|
38
38
|
// report — a full Report-#001-style suite run (scripts/run-report-001.js)
|
|
39
39
|
// trigger — a release-trigger-initiated prepare-report run
|
|
40
|
-
|
|
41
|
-
|
|
40
|
+
// Recalibrated 2026-09-01 (spec 018). BOTH literals below were set against a
|
|
41
|
+
// draws=1 projection and were never rescaled when spec 014 (5e08ba2) made the
|
|
42
|
+
// runner draw the generation up to SAMPLING.max times per arm and spec 016 made
|
|
43
|
+
// estimateRunCostUSD require the draw factor. The projection grew 10x; these did
|
|
44
|
+
// not, so `npx driftproof run` refused every suite of two or more cases and the
|
|
45
|
+
// v0.7.0 Action self-test aborted with exit 3. REPORT_MAX_USD was raised 40->300
|
|
46
|
+
// on 2026-08-31 for exactly this reason; that pass missed the dev path, which is
|
|
47
|
+
// the one the Action and every npx user take.
|
|
48
|
+
//
|
|
49
|
+
// DERIVATION — the headroom the pre-sampling defaults carried is preserved, not
|
|
50
|
+
// widened. Bundled 10-case example, 5 judge samples, per-case factor 2+2*5 = 12:
|
|
51
|
+
// draws=1 -> 120 calls, $0.3525 headroom 200/120 = 1.67x, $2/$0.3525 = 5.67x
|
|
52
|
+
// draws=10 -> 1200 calls, $3.5250 1200 * 1.67 = 2000, $3.5250 * 5.67 = $20.00
|
|
53
|
+
// A 2000-call cap under draws=10 is exactly as tight as 200 was under draws=1.
|
|
54
|
+
// LITERALS ON PURPOSE (spec 018 AC-2): a default derived from the suite in hand
|
|
55
|
+
// can never fire, which retires the guard instead of recalibrating it.
|
|
56
|
+
const DEV_MAX_USD = 20;
|
|
57
|
+
const DEV_MAX_CALLS = 2000;
|
|
58
|
+
// Raised from $40 on 2026-08-31 by the repository owner, in daylight, recorded in
|
|
59
|
+
// DECISIONS. The run did not get more expensive — the projection got honest:
|
|
60
|
+
// spec 016 made estimateRunCostUSD require a draw factor, and v0.5 has drawn the
|
|
61
|
+
// generation up to SAMPLING.max times per arm since spec 014, so the old cap had
|
|
62
|
+
// been passing on a figure up to ten times too small. Set above the ceiling this
|
|
63
|
+
// instrument can currently reach (#007's measured basis at the ceiling is
|
|
64
|
+
// $268.83) rather than just above today's staging figure, because a cap tripped
|
|
65
|
+
// by the next honest run teaches everyone to nudge it.
|
|
66
|
+
const REPORT_MAX_USD = 300;
|
|
42
67
|
const TRIGGER_MAX_USD = 25;
|
|
43
68
|
|
|
44
69
|
// Default number of judge samples per case in sampled mode. Sampling is what
|
|
@@ -56,6 +81,24 @@ const DEFAULT_JUDGE_SAMPLES = 5;
|
|
|
56
81
|
// spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
|
|
57
82
|
const EFFECT_FLOOR = 0.05;
|
|
58
83
|
|
|
84
|
+
// ── generation sampling (receipt spec v0.5) ─────────────────────────────────
|
|
85
|
+
//
|
|
86
|
+
// Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
|
|
87
|
+
// instrument sampled only the judge. These are the policy that samples the
|
|
88
|
+
// other axis. Constants, not literals in the runner: lib/sampling.js reads
|
|
89
|
+
// them, so moving one here moves the policy, and the gate asserts that by
|
|
90
|
+
// value rather than by grepping for a name.
|
|
91
|
+
//
|
|
92
|
+
// MIN is 3 because two draws give an sd that is barely a measurement and one
|
|
93
|
+
// gives none at all. MAX is 10: #006's probe used 20 by hand and found the
|
|
94
|
+
// shape at well under half of that, and a per-case ceiling bounds the spend.
|
|
95
|
+
// The SD threshold is the effect floor — a spread wider than the smallest move
|
|
96
|
+
// the verdict rule will call real is exactly when more draws are owed.
|
|
97
|
+
const GENERATION_SAMPLES_MIN = 3;
|
|
98
|
+
const GENERATION_SAMPLES_MAX = 10;
|
|
99
|
+
const GENERATION_SD_THRESHOLD = EFFECT_FLOOR;
|
|
100
|
+
const GENERATION_STABILITY_EPS = 0.01;
|
|
101
|
+
|
|
59
102
|
// Report #002 is the first CROSS-PROVIDER report: the same suites, the same fixed
|
|
60
103
|
// Haiku judge, run on two substrates — a Claude flagship and a GPT flagship. The
|
|
61
104
|
// GPT flagship is a config constant (not hard-coded across scripts) so a future
|
|
@@ -90,7 +133,8 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
|
90
133
|
|
|
91
134
|
module.exports = {
|
|
92
135
|
PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
|
|
93
|
-
EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
136
|
+
EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
137
|
+
GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
|
|
94
138
|
REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
|
|
95
139
|
REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
|
|
96
140
|
REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
|
package/lib/canary.js
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
'use strict';
|
|
3
|
+
|
|
4
|
+
// Suite canary — receipt spec v0.5.
|
|
5
|
+
//
|
|
6
|
+
// Terminal-Bench 4.0 records a canary string so that a benchmark appearing in
|
|
7
|
+
// training data is detectable. Adopted, translated: our unit of publication is
|
|
8
|
+
// the SUITE, so the canary is per suite, derived from the suite's identity and
|
|
9
|
+
// its case ids rather than randomly assigned, so the same suite yields the same
|
|
10
|
+
// canary on every machine and no registry has to be kept.
|
|
11
|
+
//
|
|
12
|
+
// It is a detection aid, not a control: a canary tells you a suite leaked, it
|
|
13
|
+
// does not stop the leak, and it cannot prove the absence of one.
|
|
14
|
+
|
|
15
|
+
const crypto = require('crypto');
|
|
16
|
+
|
|
17
|
+
const NAMESPACE = 'driftproof/suite-canary/v1';
|
|
18
|
+
|
|
19
|
+
function suiteCanary(suite) {
|
|
20
|
+
const id = String((suite && (suite.id || suite.name)) || '');
|
|
21
|
+
const caseIds = ((suite && suite.cases) || []).map((c) => String((c && c.id) || '')).sort();
|
|
22
|
+
const h = crypto.createHash('sha256').update(`${NAMESPACE}\n${id}\n${caseIds.join('\n')}`).digest('hex');
|
|
23
|
+
// Formatted as a GUID so it is recognisable as one in a corpus scan.
|
|
24
|
+
return [h.slice(0, 8), h.slice(8, 12), h.slice(12, 16), h.slice(16, 20), h.slice(20, 32)].join('-');
|
|
25
|
+
}
|
|
26
|
+
|
|
27
|
+
module.exports = { suiteCanary, NAMESPACE };
|
package/lib/cost.js
CHANGED
|
@@ -58,7 +58,25 @@ function perCallCostUSD(modelId, kind) {
|
|
|
58
58
|
// generated on each target model, judged `samples` times each on `judgeModel`.
|
|
59
59
|
//
|
|
60
60
|
// returns { totalUSD, perModel: [{ model, usd }], judgeUSD, genUSD, assumptions }
|
|
61
|
-
|
|
61
|
+
// THE DRAW FACTOR IS REQUIRED AND HAS NO DEFAULT (spec 016 AC-7).
|
|
62
|
+
//
|
|
63
|
+
// v0.5 draws the generation up to `SAMPLING.max` times per arm, so a run costs
|
|
64
|
+
// its one-draw estimate times the number of draws. `projectCalls` learned this in
|
|
65
|
+
// spec 014 with a `draws` argument DEFAULTING TO 1 "so every existing caller
|
|
66
|
+
// projects exactly what it projected before" — and every existing caller then
|
|
67
|
+
// went on projecting a tenth of the run, silently, for two more loops. Spec 015
|
|
68
|
+
// corrected the CALL projection in three report scripts and left the DOLLAR
|
|
69
|
+
// projection at one draw everywhere, which is the figure a human actually reads
|
|
70
|
+
// before authorising a paid run.
|
|
71
|
+
//
|
|
72
|
+
// A default is what made that invisible, so there is none. Omitting the factor
|
|
73
|
+
// throws, which turns a silent understatement into a loud stop — and every call
|
|
74
|
+
// site has to say what it means, including the ones that legitimately mean 1.
|
|
75
|
+
function estimateRunCostUSD({ caseCount, samples, models, judgeModel, draws }) {
|
|
76
|
+
if (!Number.isFinite(draws) || draws < 1) {
|
|
77
|
+
throw new Error('estimateRunCostUSD: `draws` is required and must be >= 1 — pass SAMPLING.max to project a v0.5 run, or 1 to price a single draw deliberately. It is not defaulted, because a default is how the dollar projection stayed at one draw through two loops.');
|
|
78
|
+
}
|
|
79
|
+
caseCount = caseCount * draws;
|
|
62
80
|
const judgePrice = priceFor(judgeModel);
|
|
63
81
|
const perModel = [];
|
|
64
82
|
let judgeUSD = 0;
|
|
@@ -80,7 +98,7 @@ function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
|
|
|
80
98
|
perModel,
|
|
81
99
|
judgeUSD: round4(judgeUSD),
|
|
82
100
|
genUSD: round4(genUSD),
|
|
83
|
-
assumptions: { tokens: TOKENS, judgeModel, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
|
|
101
|
+
assumptions: { tokens: TOKENS, judgeModel, draws, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
|
|
84
102
|
};
|
|
85
103
|
}
|
|
86
104
|
|