driftproof 0.6.0 β 0.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +44 -24
- package/bin/driftproof +18 -7
- package/config.js +49 -5
- package/lib/canary.js +27 -0
- package/lib/cost.js +20 -2
- package/lib/diff.js +114 -8
- package/lib/judge.js +6 -1
- package/lib/receipt.js +76 -4
- package/lib/reuse.js +134 -0
- package/lib/revision.js +8 -1
- package/lib/run.js +165 -45
- package/lib/sampling.js +110 -0
- package/lib/stats.js +16 -1
- package/lib/value.js +27 -2
- package/package.json +2 -2
- package/spec/RECEIPT.md +100 -1
- package/spec/receipt.schema.json +360 -3
- package/spec/receipt.v0.4.schema.json +961 -0
package/README.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
2
|
# Driftproof
|
|
3
3
|
|
|
4
|
-
**
|
|
4
|
+
**A dated proof that this skill, this hash, this model, still helps.**
|
|
5
5
|
|
|
6
6
|
[](https://driftproofhq.com)
|
|
7
7
|
β live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
@@ -15,11 +15,11 @@ across model releases and you get a **drift report**.
|
|
|
15
15
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
16
16
|
format; it does not invent its own.
|
|
17
17
|
|
|
18
|
-
π **
|
|
19
|
-
hand-entered), spanning
|
|
20
|
-
|
|
21
|
-
moves underneath the skill β or, in the value report, in which
|
|
22
|
-
measured:
|
|
18
|
+
π **Seven published reports** (each re-derived from committed receipts, nothing
|
|
19
|
+
hand-entered), spanning six published report types, the newest being instrument
|
|
20
|
+
re-measurement. All seven share one band-based, floor-gated verdict rule and
|
|
21
|
+
differ in what moves underneath the skill β or, in the value report, in which
|
|
22
|
+
axes are measured; or, in the newest, in the instrument itself:
|
|
23
23
|
|
|
24
24
|
- **[Report #001](https://driftproofhq.com/reports/001/)** β *release drift*: ten
|
|
25
25
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
|
|
@@ -51,11 +51,23 @@ measured:
|
|
|
51
51
|
established β and the receipt spec gains generation sampling as a result. **No
|
|
52
52
|
cause is asserted**; the control proves non-reproduction and cannot say why.
|
|
53
53
|
*The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
|
|
54
|
+
- **[Report #007](https://driftproofhq.com/reports/007/)** β *instrument
|
|
55
|
+
re-measurement*: the three cells Report #005 published for these skills, run
|
|
56
|
+
again with the generation sampled adaptively instead of once and with the call
|
|
57
|
+
timeout the surface policy declares. **No cell separates**: every lift is
|
|
58
|
+
smaller than its own band, and the 21 cases read 3 improved, 0 regressed,
|
|
59
|
+
18 no effect, 0 not measured. Five of six comparisons against the archive were
|
|
60
|
+
**refused** on their baseline control. The instrument defect it reports is its
|
|
61
|
+
own: a declared 300 s timeout had been shadowed by a `120000` literal since
|
|
62
|
+
2026-07-27, and the truncated run it caused measured *less* variance than the
|
|
63
|
+
clean re-run, which is the direction that flatters an instrument. Both runs are
|
|
64
|
+
published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
|
|
54
65
|
|
|
55
66
|
βοΈ The launch essay, **[Three model releases later: what actually happens to agent
|
|
56
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads all
|
|
57
|
-
together: what moves underneath a skill,
|
|
58
|
-
|
|
67
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all seven reports
|
|
68
|
+
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
69
|
+
corrected instrument did to three published results. Revised 2026-09-01; every
|
|
70
|
+
figure in it is gate-checked against the report page it cites.
|
|
59
71
|
|
|
60
72
|
## Why
|
|
61
73
|
|
|
@@ -67,7 +79,7 @@ re-checks.
|
|
|
67
79
|
|
|
68
80
|
**Verdicts age because the substrate moves.** Driftproof exists to keep the
|
|
69
81
|
verdict current: cheap, repeatable, hash-stamped measurements bound to a specific
|
|
70
|
-
model version, so "does this skill still help?" has a dated,
|
|
82
|
+
model version, so "does this skill still help?" has a dated, verifiable answer
|
|
71
83
|
instead of a stale one.
|
|
72
84
|
|
|
73
85
|
The hard part isn't running an eval once β it's making the number **credible
|
|
@@ -100,8 +112,10 @@ npx driftproof init my-skill
|
|
|
100
112
|
export CLAUDE_PROVIDER=api
|
|
101
113
|
read -rsp "Anthropic API key: " ANTHROPIC_API_KEY && export ANTHROPIC_API_KEY
|
|
102
114
|
|
|
103
|
-
# 4. Run the suite
|
|
104
|
-
|
|
115
|
+
# 4. Run the suite. The shipped defaults (DEV_MAX_USD / DEV_MAX_CALLS in config.js)
|
|
116
|
+
# refuse the run up front if the projection exceeds either; --max-usd and
|
|
117
|
+
# --max-calls override them.
|
|
118
|
+
npx driftproof run my-skill --models claude-haiku-4-5
|
|
105
119
|
|
|
106
120
|
# 5. Read the receipt + human summary written to ./receipts/
|
|
107
121
|
cat receipts/*.summary.md
|
|
@@ -187,11 +201,16 @@ The surface and judge settings used are recorded in every receipt.
|
|
|
187
201
|
|
|
188
202
|
### Cost guard
|
|
189
203
|
|
|
190
|
-
Sampling multiplies calls
|
|
191
|
-
(
|
|
192
|
-
|
|
193
|
-
|
|
194
|
-
|
|
204
|
+
Sampling multiplies calls on two axes, and the second one is easy to miss. Each
|
|
205
|
+
case costs `SAMPLING.max Γ (2 + 2 Γ samples)` model calls: two arms, each **drawn
|
|
206
|
+
up to `SAMPLING.max` times** (generation sampling, receipt spec v0.5), and every
|
|
207
|
+
draw judged `samples` times. Read at a single draw, that formula understates a run
|
|
208
|
+
by the whole draw factor β which is how the shipped caps came to sit a release
|
|
209
|
+
behind the estimator. Driftproof **projects the whole run up front, prints the
|
|
210
|
+
count, and refuses before spending anything** if the projection exceeds the
|
|
211
|
+
per-model call cap (`DEV_MAX_CALLS`) or the dollar budget (`DEV_MAX_USD`) β both
|
|
212
|
+
declared with their derivation in [`config.js`](config.js), both overridable with
|
|
213
|
+
`--max-calls` / `--max-usd`. The default model list is `haiku` only.
|
|
195
214
|
|
|
196
215
|
## Receipt anatomy
|
|
197
216
|
|
|
@@ -201,7 +220,7 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
201
220
|
|
|
202
221
|
```jsonc
|
|
203
222
|
{
|
|
204
|
-
"schema_version": "0.
|
|
223
|
+
"schema_version": "0.5",
|
|
205
224
|
"skill": { "name": "commit-message-conventions", "version": "0.2.0",
|
|
206
225
|
"content_hash": "β¦sha256 over SKILL.md + bundled filesβ¦" },
|
|
207
226
|
"suite": { "format": "agentskills.io/evals", "suite_hash": "β¦", "case_count": 10 },
|
|
@@ -210,7 +229,7 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
210
229
|
"model_release_date": "2025-10-01",
|
|
211
230
|
"provider": "anthropic",
|
|
212
231
|
"surface": "claude-cli",
|
|
213
|
-
"runner_version": "0.
|
|
232
|
+
"runner_version": "0.7.1",
|
|
214
233
|
"date_utc": "2026-07-27Tβ¦Z",
|
|
215
234
|
"registry": "registered",
|
|
216
235
|
"transcripts": "hashes-only",
|
|
@@ -270,7 +289,7 @@ receipts. The rule that keeps it honest: a **regression** (or improvement) is
|
|
|
270
289
|
claimed **only when the two bands do not overlap**. Overlapping bands are reported
|
|
271
290
|
as **within noise** and never counted as a regression.
|
|
272
291
|
|
|
273
|
-
##
|
|
292
|
+
## Verification in CI (GitHub Action + badge)
|
|
274
293
|
|
|
275
294
|
Wire drift detection into a repo so the skill is re-checked on every push and when
|
|
276
295
|
the model underneath it changes.
|
|
@@ -284,12 +303,13 @@ jobs:
|
|
|
284
303
|
runs-on: ubuntu-latest
|
|
285
304
|
steps:
|
|
286
305
|
- uses: actions/checkout@v4
|
|
287
|
-
- uses: driftproofhq/driftproof@v0.
|
|
306
|
+
- uses: driftproofhq/driftproof@v0.7.1
|
|
288
307
|
with:
|
|
289
308
|
skill-dir: skills/my-skill
|
|
290
309
|
models: claude-haiku-4-5
|
|
291
|
-
max-usd: '2'
|
|
292
310
|
api-key: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
311
|
+
# max-usd: <n> # override the dollar budget (default: DEV_MAX_USD in config.js)
|
|
312
|
+
# max-calls: <n> # override the per-model call cap (default: DEV_MAX_CALLS)
|
|
293
313
|
# fail-on-regression: 'true' # (default) fail the job if the skill REGRESSED
|
|
294
314
|
```
|
|
295
315
|
|
|
@@ -307,7 +327,7 @@ runs).
|
|
|
307
327
|
JSON object. Commit it somewhere public and point a shields endpoint URL at it:
|
|
308
328
|
|
|
309
329
|
```bash
|
|
310
|
-
npx driftproof run skills/my-skill --models claude-haiku-4-5 --
|
|
330
|
+
npx driftproof run skills/my-skill --models claude-haiku-4-5 --out receipts
|
|
311
331
|
npx driftproof badge receipts/*.json --out badges/my-skill.json
|
|
312
332
|
git add badges/my-skill.json && git commit -m "chore: driftproof badge"
|
|
313
333
|
```
|
|
@@ -323,7 +343,7 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
323
343
|
|
|
324
344
|
## Reports
|
|
325
345
|
|
|
326
|
-
|
|
346
|
+
Seven reports are published, spanning six report types. A report page lives at a
|
|
327
347
|
draft path β `docs/reports/NNN-draft/` β until the publish sequence renames it, and
|
|
328
348
|
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
329
349
|
(see the roll at the top of this README, and
|
package/bin/driftproof
CHANGED
|
@@ -4,9 +4,10 @@
|
|
|
4
4
|
|
|
5
5
|
const fs = require('fs');
|
|
6
6
|
const path = require('path');
|
|
7
|
-
const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = require('../config');
|
|
7
|
+
const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD, DEV_MAX_CALLS } = require('../config');
|
|
8
8
|
const { loadSkill } = require('../lib/skill');
|
|
9
9
|
const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
|
|
10
|
+
const { SAMPLING } = require('../lib/sampling');
|
|
10
11
|
const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
|
|
11
12
|
const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
|
|
12
13
|
const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
|
|
@@ -91,8 +92,12 @@ ENV
|
|
|
91
92
|
NOTES
|
|
92
93
|
- Default models list is 'haiku' only (cheap dev default).
|
|
93
94
|
- Sampled judging: each case does 2 generations + 2Γsamples judge calls.
|
|
94
|
-
Default --samples 5 β 12 calls/case; --samples 1 disables bands
|
|
95
|
-
|
|
95
|
+
Default --samples 5 β 12 calls/case; --samples 1 disables bands AND the
|
|
96
|
+
variance ratio: with one judge sample the per-draw judge spread is zero by
|
|
97
|
+
construction, so the receipt records variance_ratio: null with
|
|
98
|
+
variance_ratio_unavailable: "single_judge_sample". Legal, but a run whose
|
|
99
|
+
ratio you intend to publish needs --samples 2 or more.
|
|
100
|
+
- --max-calls is a hard per-run call cap (default ${DEV_MAX_CALLS}). The whole run is
|
|
96
101
|
projected up front and REFUSED before any call if it would exceed the cap.
|
|
97
102
|
- --max-usd is a hard dollar budget (default ${DEV_MAX_USD} for dev runs). The projected
|
|
98
103
|
cost is printed up front and the run is REFUSED on BOTH surfaces if it would
|
|
@@ -126,7 +131,7 @@ async function cmdRun(positional, flags) {
|
|
|
126
131
|
const rc = loadRc(skillDir);
|
|
127
132
|
const models = String(flags.models || rc.models || 'haiku').split(',').map((s) => s.trim()).filter(Boolean);
|
|
128
133
|
const maxCases = flags['max-cases'] ? parseInt(flags['max-cases'], 10) : (rc.max_cases != null ? parseInt(rc.max_cases, 10) : null);
|
|
129
|
-
const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) :
|
|
134
|
+
const maxCalls = flags['max-calls'] ? parseInt(flags['max-calls'], 10) : (rc.max_calls != null ? parseInt(rc.max_calls, 10) : DEV_MAX_CALLS);
|
|
130
135
|
const samples = flags.samples ? parseInt(flags.samples, 10) : (rc.samples != null ? parseInt(rc.samples, 10) : DEFAULT_JUDGE_SAMPLES);
|
|
131
136
|
const judgeModel = flags['judge-model'] || rc.judge_model || null;
|
|
132
137
|
const concurrency = flags.concurrency ? parseInt(flags.concurrency, 10) : (rc.concurrency != null ? parseInt(rc.concurrency, 10) : 1);
|
|
@@ -137,14 +142,20 @@ async function cmdRun(positional, flags) {
|
|
|
137
142
|
|
|
138
143
|
const skill = loadSkill(skillDir);
|
|
139
144
|
const nCases = maxCases ? Math.min(maxCases, skill.suite.caseCount) : skill.suite.caseCount;
|
|
140
|
-
|
|
145
|
+
// v0.5 draws the generation up to SAMPLING.max times per arm, so BOTH the
|
|
146
|
+
// printed projection and the dollar guard below must be scaled by it. Left at
|
|
147
|
+
// draws=1 the guard would admit a run costing up to ten times its own
|
|
148
|
+
// projection β the display would say one number and the runner would refuse
|
|
149
|
+
// quoting another. Conservative on purpose: refusing a run that would have fit
|
|
150
|
+
// is recoverable, overspending is not.
|
|
151
|
+
const perModelCalls = projectCalls(nCases, samples, SAMPLING.max);
|
|
141
152
|
const totalProjected = perModelCalls * models.length;
|
|
142
153
|
|
|
143
154
|
// Dollar cost guard (Week 3): project the METERED USD cost up front. On the
|
|
144
155
|
// `api` surface this is real spend, so we refuse if it would exceed --max-usd.
|
|
145
156
|
// On `claude-cli` the metered spend is $0 (subscription); the figure is the
|
|
146
157
|
// hypothetical "if run on the metered API" cost β printed, never blocks.
|
|
147
|
-
const cost = estimateRunCostUSD({ caseCount: nCases, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
|
|
158
|
+
const cost = estimateRunCostUSD({ caseCount: nCases, draws: SAMPLING.max, samples, models: models.map((m) => require('../lib/provider').resolveModel(m)), judgeModel: judgeModel || 'haiku' });
|
|
148
159
|
// Surface is per-model now (a run may mix an Anthropic and an OpenAI target).
|
|
149
160
|
const surfaces = [...new Set(models.map((m) => surfaceForModel(m)))];
|
|
150
161
|
const surface = surfaces.join(', ');
|
|
@@ -315,7 +326,7 @@ function cmdInit(positional) {
|
|
|
315
326
|
console.log(`\nNext:
|
|
316
327
|
1. Edit ${rel(path.join(dir, 'SKILL.md'))} with your skill's instructions.
|
|
317
328
|
2. Edit ${rel(path.join(dir, 'evals', 'evals.json'))} β replace the 3 example cases (rubrics anchored at 0.80).
|
|
318
|
-
3. driftproof run ${rel(dir)} --models claude-haiku-4-5
|
|
329
|
+
3. driftproof run ${rel(dir)} --models claude-haiku-4-5
|
|
319
330
|
See AUTHORING.md for how to write a fair suite.`);
|
|
320
331
|
}
|
|
321
332
|
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.7.1';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
|
@@ -29,7 +29,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
|
|
|
29
29
|
// at run time so derived dollars stay reproducible), and the derived
|
|
30
30
|
// `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
|
|
31
31
|
// v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
|
|
32
|
-
const RECEIPT_SCHEMA_VERSION = '0.
|
|
32
|
+
const RECEIPT_SCHEMA_VERSION = '0.5';
|
|
33
33
|
|
|
34
34
|
// Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
|
|
35
35
|
// these. The projection is refused before any call if it exceeds the cap, on
|
|
@@ -37,8 +37,33 @@ const RECEIPT_SCHEMA_VERSION = '0.4';
|
|
|
37
37
|
// dev β an interactive `driftproof run`
|
|
38
38
|
// report β a full Report-#001-style suite run (scripts/run-report-001.js)
|
|
39
39
|
// trigger β a release-trigger-initiated prepare-report run
|
|
40
|
-
|
|
41
|
-
|
|
40
|
+
// Recalibrated 2026-09-01 (spec 018). BOTH literals below were set against a
|
|
41
|
+
// draws=1 projection and were never rescaled when spec 014 (5e08ba2) made the
|
|
42
|
+
// runner draw the generation up to SAMPLING.max times per arm and spec 016 made
|
|
43
|
+
// estimateRunCostUSD require the draw factor. The projection grew 10x; these did
|
|
44
|
+
// not, so `npx driftproof run` refused every suite of two or more cases and the
|
|
45
|
+
// v0.7.0 Action self-test aborted with exit 3. REPORT_MAX_USD was raised 40->300
|
|
46
|
+
// on 2026-08-31 for exactly this reason; that pass missed the dev path, which is
|
|
47
|
+
// the one the Action and every npx user take.
|
|
48
|
+
//
|
|
49
|
+
// DERIVATION β the headroom the pre-sampling defaults carried is preserved, not
|
|
50
|
+
// widened. Bundled 10-case example, 5 judge samples, per-case factor 2+2*5 = 12:
|
|
51
|
+
// draws=1 -> 120 calls, $0.3525 headroom 200/120 = 1.67x, $2/$0.3525 = 5.67x
|
|
52
|
+
// draws=10 -> 1200 calls, $3.5250 1200 * 1.67 = 2000, $3.5250 * 5.67 = $20.00
|
|
53
|
+
// A 2000-call cap under draws=10 is exactly as tight as 200 was under draws=1.
|
|
54
|
+
// LITERALS ON PURPOSE (spec 018 AC-2): a default derived from the suite in hand
|
|
55
|
+
// can never fire, which retires the guard instead of recalibrating it.
|
|
56
|
+
const DEV_MAX_USD = 20;
|
|
57
|
+
const DEV_MAX_CALLS = 2000;
|
|
58
|
+
// Raised from $40 on 2026-08-31 by the repository owner, in daylight, recorded in
|
|
59
|
+
// DECISIONS. The run did not get more expensive β the projection got honest:
|
|
60
|
+
// spec 016 made estimateRunCostUSD require a draw factor, and v0.5 has drawn the
|
|
61
|
+
// generation up to SAMPLING.max times per arm since spec 014, so the old cap had
|
|
62
|
+
// been passing on a figure up to ten times too small. Set above the ceiling this
|
|
63
|
+
// instrument can currently reach (#007's measured basis at the ceiling is
|
|
64
|
+
// $268.83) rather than just above today's staging figure, because a cap tripped
|
|
65
|
+
// by the next honest run teaches everyone to nudge it.
|
|
66
|
+
const REPORT_MAX_USD = 300;
|
|
42
67
|
const TRIGGER_MAX_USD = 25;
|
|
43
68
|
|
|
44
69
|
// Default number of judge samples per case in sampled mode. Sampling is what
|
|
@@ -56,6 +81,24 @@ const DEFAULT_JUDGE_SAMPLES = 5;
|
|
|
56
81
|
// spec/RECEIPT.md Β§ "Drift verdict rule" and the report methodology.
|
|
57
82
|
const EFFECT_FLOOR = 0.05;
|
|
58
83
|
|
|
84
|
+
// ββ generation sampling (receipt spec v0.5) βββββββββββββββββββββββββββββββββ
|
|
85
|
+
//
|
|
86
|
+
// Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
|
|
87
|
+
// instrument sampled only the judge. These are the policy that samples the
|
|
88
|
+
// other axis. Constants, not literals in the runner: lib/sampling.js reads
|
|
89
|
+
// them, so moving one here moves the policy, and the gate asserts that by
|
|
90
|
+
// value rather than by grepping for a name.
|
|
91
|
+
//
|
|
92
|
+
// MIN is 3 because two draws give an sd that is barely a measurement and one
|
|
93
|
+
// gives none at all. MAX is 10: #006's probe used 20 by hand and found the
|
|
94
|
+
// shape at well under half of that, and a per-case ceiling bounds the spend.
|
|
95
|
+
// The SD threshold is the effect floor β a spread wider than the smallest move
|
|
96
|
+
// the verdict rule will call real is exactly when more draws are owed.
|
|
97
|
+
const GENERATION_SAMPLES_MIN = 3;
|
|
98
|
+
const GENERATION_SAMPLES_MAX = 10;
|
|
99
|
+
const GENERATION_SD_THRESHOLD = EFFECT_FLOOR;
|
|
100
|
+
const GENERATION_STABILITY_EPS = 0.01;
|
|
101
|
+
|
|
59
102
|
// Report #002 is the first CROSS-PROVIDER report: the same suites, the same fixed
|
|
60
103
|
// Haiku judge, run on two substrates β a Claude flagship and a GPT flagship. The
|
|
61
104
|
// GPT flagship is a config constant (not hard-coded across scripts) so a future
|
|
@@ -90,7 +133,8 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
|
90
133
|
|
|
91
134
|
module.exports = {
|
|
92
135
|
PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
|
|
93
|
-
EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
136
|
+
EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
137
|
+
GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
|
|
94
138
|
REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
|
|
95
139
|
REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
|
|
96
140
|
REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
|
package/lib/canary.js
ADDED
|
@@ -0,0 +1,27 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
'use strict';
|
|
3
|
+
|
|
4
|
+
// Suite canary β receipt spec v0.5.
|
|
5
|
+
//
|
|
6
|
+
// Terminal-Bench 4.0 records a canary string so that a benchmark appearing in
|
|
7
|
+
// training data is detectable. Adopted, translated: our unit of publication is
|
|
8
|
+
// the SUITE, so the canary is per suite, derived from the suite's identity and
|
|
9
|
+
// its case ids rather than randomly assigned, so the same suite yields the same
|
|
10
|
+
// canary on every machine and no registry has to be kept.
|
|
11
|
+
//
|
|
12
|
+
// It is a detection aid, not a control: a canary tells you a suite leaked, it
|
|
13
|
+
// does not stop the leak, and it cannot prove the absence of one.
|
|
14
|
+
|
|
15
|
+
const crypto = require('crypto');
|
|
16
|
+
|
|
17
|
+
const NAMESPACE = 'driftproof/suite-canary/v1';
|
|
18
|
+
|
|
19
|
+
function suiteCanary(suite) {
|
|
20
|
+
const id = String((suite && (suite.id || suite.name)) || '');
|
|
21
|
+
const caseIds = ((suite && suite.cases) || []).map((c) => String((c && c.id) || '')).sort();
|
|
22
|
+
const h = crypto.createHash('sha256').update(`${NAMESPACE}\n${id}\n${caseIds.join('\n')}`).digest('hex');
|
|
23
|
+
// Formatted as a GUID so it is recognisable as one in a corpus scan.
|
|
24
|
+
return [h.slice(0, 8), h.slice(8, 12), h.slice(12, 16), h.slice(16, 20), h.slice(20, 32)].join('-');
|
|
25
|
+
}
|
|
26
|
+
|
|
27
|
+
module.exports = { suiteCanary, NAMESPACE };
|
package/lib/cost.js
CHANGED
|
@@ -58,7 +58,25 @@ function perCallCostUSD(modelId, kind) {
|
|
|
58
58
|
// generated on each target model, judged `samples` times each on `judgeModel`.
|
|
59
59
|
//
|
|
60
60
|
// returns { totalUSD, perModel: [{ model, usd }], judgeUSD, genUSD, assumptions }
|
|
61
|
-
|
|
61
|
+
// THE DRAW FACTOR IS REQUIRED AND HAS NO DEFAULT (spec 016 AC-7).
|
|
62
|
+
//
|
|
63
|
+
// v0.5 draws the generation up to `SAMPLING.max` times per arm, so a run costs
|
|
64
|
+
// its one-draw estimate times the number of draws. `projectCalls` learned this in
|
|
65
|
+
// spec 014 with a `draws` argument DEFAULTING TO 1 "so every existing caller
|
|
66
|
+
// projects exactly what it projected before" β and every existing caller then
|
|
67
|
+
// went on projecting a tenth of the run, silently, for two more loops. Spec 015
|
|
68
|
+
// corrected the CALL projection in three report scripts and left the DOLLAR
|
|
69
|
+
// projection at one draw everywhere, which is the figure a human actually reads
|
|
70
|
+
// before authorising a paid run.
|
|
71
|
+
//
|
|
72
|
+
// A default is what made that invisible, so there is none. Omitting the factor
|
|
73
|
+
// throws, which turns a silent understatement into a loud stop β and every call
|
|
74
|
+
// site has to say what it means, including the ones that legitimately mean 1.
|
|
75
|
+
function estimateRunCostUSD({ caseCount, samples, models, judgeModel, draws }) {
|
|
76
|
+
if (!Number.isFinite(draws) || draws < 1) {
|
|
77
|
+
throw new Error('estimateRunCostUSD: `draws` is required and must be >= 1 β pass SAMPLING.max to project a v0.5 run, or 1 to price a single draw deliberately. It is not defaulted, because a default is how the dollar projection stayed at one draw through two loops.');
|
|
78
|
+
}
|
|
79
|
+
caseCount = caseCount * draws;
|
|
62
80
|
const judgePrice = priceFor(judgeModel);
|
|
63
81
|
const perModel = [];
|
|
64
82
|
let judgeUSD = 0;
|
|
@@ -80,7 +98,7 @@ function estimateRunCostUSD({ caseCount, samples, models, judgeModel }) {
|
|
|
80
98
|
perModel,
|
|
81
99
|
judgeUSD: round4(judgeUSD),
|
|
82
100
|
genUSD: round4(genUSD),
|
|
83
|
-
assumptions: { tokens: TOKENS, judgeModel, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
|
|
101
|
+
assumptions: { tokens: TOKENS, judgeModel, draws, note: 'rough upper-bound estimate; registry per-MTok pricing; not measured with count_tokens' },
|
|
84
102
|
};
|
|
85
103
|
}
|
|
86
104
|
|
package/lib/diff.js
CHANGED
|
@@ -4,6 +4,7 @@
|
|
|
4
4
|
const { bandVerdict, round } = require('./stats');
|
|
5
5
|
const { EFFECT_FLOOR } = require('../config');
|
|
6
6
|
const { revisionHeadline } = require('./revision');
|
|
7
|
+
const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
|
|
7
8
|
|
|
8
9
|
// Practical-significance gate applied ON TOP of band separation. bandVerdict()
|
|
9
10
|
// stays a pure geometry test (kept that way so its unit checks are unambiguous);
|
|
@@ -32,11 +33,23 @@ function isWithinNoise(v) { return v === 'within noise' || v === WITHIN_NOISE_FL
|
|
|
32
33
|
|
|
33
34
|
// Map case-id β { mean, stddev } for a receipt's with_skill cases. Falls back to
|
|
34
35
|
// score/0 for v0.1 receipts that have no per-case band.
|
|
36
|
+
// NO RENDERING SWITCHES. An earlier revision carried two const-true flags so each
|
|
37
|
+
// half of F-015-B's fix could be removed in a probe β and the dead branch of one
|
|
38
|
+
// of them contained, verbatim, the false headline AC-4 exists to forbid. Nothing
|
|
39
|
+
// could fire it, and it was still a defect-restoring branch living in the tree
|
|
40
|
+
// that gets published. The mutation probes patch this source in a disposable copy.
|
|
41
|
+
|
|
35
42
|
function withSkillBands(receipt) {
|
|
36
43
|
const out = {};
|
|
37
44
|
for (const c of receipt.results.cases) {
|
|
38
45
|
if (c.mode !== 'with_skill') continue;
|
|
39
|
-
|
|
46
|
+
// ONE BAND DEFINITION (spec 016 AC-1). This built its own from the v0.4-shaped
|
|
47
|
+
// fields, which worked β and that is the point: three copies existed, two
|
|
48
|
+
// reading the legacy shape and one reading only v0.5, and the one that read
|
|
49
|
+
// only v0.5 was the one a cross-version control depended on (F-015-C). Routing
|
|
50
|
+
// every comparison path through `bandOf` means a future shape is added once.
|
|
51
|
+
const b = bandOf(c);
|
|
52
|
+
if (b) out[c.id] = { mean: b.mean, stddev: b.sd, source: b.source };
|
|
40
53
|
}
|
|
41
54
|
return out;
|
|
42
55
|
}
|
|
@@ -51,7 +64,26 @@ function aggWithBand(receipt) {
|
|
|
51
64
|
|
|
52
65
|
function fmt(n) { return n == null ? 'n/a' : (n >= 0 ? '+' : '') + n.toFixed(3); }
|
|
53
66
|
function pct(n) { return n == null ? 'n/a' : n.toFixed(3); }
|
|
54
|
-
|
|
67
|
+
// THE BAND SAYS WHICH BAND IT IS (spec 017 AC-7).
|
|
68
|
+
//
|
|
69
|
+
// `bandOf` computes `source: 'legacy' | 'generation'` and carries it onto every
|
|
70
|
+
// band; nothing rendered it, so a reader comparing an archived receipt with a
|
|
71
|
+
// v0.5 one was comparing a JUDGE-SAMPLE spread against an ACROSS-DRAW spread
|
|
72
|
+
// with nothing on the page saying so. They are different statistics over
|
|
73
|
+
// different things, and the comparison is still the only one v0.4 admits β which
|
|
74
|
+
// is exactly why the page has to name them rather than leave them to look alike.
|
|
75
|
+
//
|
|
76
|
+
// THE MARKER IS THE RECEIPT'S OWN WORD β `legacy` or `generation`, exactly as
|
|
77
|
+
// `bandOf` records it β rather than a prettier synonym. A reader who greps the
|
|
78
|
+
// page for what a receipt says should find the same token; a rendering that
|
|
79
|
+
// renames the thing it is disclosing has disclosed a different thing. Omitted
|
|
80
|
+
// when a band carries no source (an aggregate band is computed from case means,
|
|
81
|
+
// not from one case's draws).
|
|
82
|
+
|
|
83
|
+
function bandStr(x) {
|
|
84
|
+
const label = x && x.source;
|
|
85
|
+
return `${x.mean.toFixed(3)} Β± ${x.stddev.toFixed(3)}${label ? ` (${label})` : ''}`;
|
|
86
|
+
}
|
|
55
87
|
function short(h) { return h ? String(h).slice(0, 12) : 'n/a'; }
|
|
56
88
|
|
|
57
89
|
// The headline is a SUMMARY of the per-case band-overlap verdicts β the credibility
|
|
@@ -85,6 +117,25 @@ function revisionPairProblem(a, b) {
|
|
|
85
117
|
|
|
86
118
|
function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
|
|
87
119
|
const revision = mode === 'revision';
|
|
120
|
+
|
|
121
|
+
// AC-6 (spec 014) β THE PRECONDITION, IN THE PATH THAT EMITS THE VERDICT.
|
|
122
|
+
//
|
|
123
|
+
// A comparison across schema versions assumes the two runs measured the same
|
|
124
|
+
// thing. The assumption is checkable: the no-skill arm contains no skill text,
|
|
125
|
+
// so nothing about a skill or a spec revision can move it. If it fails to
|
|
126
|
+
// reproduce, the comparison has no ground, and Report #006 is what that looks
|
|
127
|
+
// like when it is checked β 3 cells, 0 measured, 3 refused.
|
|
128
|
+
//
|
|
129
|
+
// SCOPED TO CROSS-VERSION PAIRS, which is what AC-6 names. Revision-mode pairs
|
|
130
|
+
// already carry their own register through revisionPairProblem(); widening
|
|
131
|
+
// this to every same-version comparison is a live question, recorded in
|
|
132
|
+
// tasks.md as OPEN-QUESTION-3 rather than decided here.
|
|
133
|
+
const crossVersion = String(a.schema_version || '') !== String(b.schema_version || '');
|
|
134
|
+
let refusal = null;
|
|
135
|
+
if (crossVersion) {
|
|
136
|
+
const pre = baselineReproduces(a, b);
|
|
137
|
+
if (!pre.ok) refusal = { key: pre.key, reason: REFUSAL_REASONS[pre.key](pre) };
|
|
138
|
+
}
|
|
88
139
|
const aB = withSkillBands(a);
|
|
89
140
|
const bB = withSkillBands(b);
|
|
90
141
|
const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
|
|
@@ -95,18 +146,21 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
95
146
|
// improvement verdicts are never computed from evidence we did not run.
|
|
96
147
|
const levelOf = (r) => r.verification_level || 'TESTED';
|
|
97
148
|
const belowTested = [[labelA, a], [labelB, b]].filter(([, r]) => levelOf(r) !== 'TESTED');
|
|
98
|
-
|
|
149
|
+
// A REFUSAL IS A RESULT, and it suppresses the verdict exactly as an
|
|
150
|
+
// untested input does: no delta is asserted, and the reason travels with it.
|
|
151
|
+
const measured = belowTested.length === 0 && !refusal;
|
|
99
152
|
|
|
100
153
|
const perCase = ids.map((id) => {
|
|
101
154
|
const before = aB[id] || null;
|
|
102
155
|
const after = bB[id] || null;
|
|
103
156
|
const delta = (before && after) ? round(after.mean - before.mean) : null;
|
|
104
|
-
const verdict =
|
|
157
|
+
const verdict = refusal ? 'refused'
|
|
158
|
+
: !measured ? 'not measured'
|
|
105
159
|
: (before && after) ? verdictWithFloor(before, after, delta) : 'n/a';
|
|
106
160
|
return { id, before, after, delta, verdict };
|
|
107
161
|
});
|
|
108
162
|
// Sort worst-first: regressions, then by delta.
|
|
109
|
-
const order = { regression: 0, 'within noise': 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3 };
|
|
163
|
+
const order = { regression: 0, 'within noise': 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3, refused: 3 };
|
|
110
164
|
perCase.sort((x, y) => (order[x.verdict] - order[y.verdict]) || ((x.delta || 0) - (y.delta || 0)));
|
|
111
165
|
|
|
112
166
|
const aAgg = aggWithBand(a);
|
|
@@ -172,13 +226,47 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
172
226
|
if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β the eval suite changed; per-case comparison may be misleading.');
|
|
173
227
|
if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β comparison is not meaningful.`);
|
|
174
228
|
if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β non-overlap can only be trusted when both sides are sampled.');
|
|
175
|
-
|
|
229
|
+
// WHAT SUPPRESSED THE VERDICT IS SAID, AND IT IS SAID CORRECTLY (spec 016 AC-3
|
|
230
|
+
// and AC-4, closing F-015-B).
|
|
231
|
+
//
|
|
232
|
+
// TWO DIFFERENT THINGS can suppress a verdict, and this line used to describe
|
|
233
|
+
// only one of them. A REFUSAL β the baseline-reproduction precondition β set
|
|
234
|
+
// `measured` false and then the caveat rendered `belowTested`, which on a
|
|
235
|
+
// refused pair is EMPTY: the page read `verdicts NOT computed β .` under a
|
|
236
|
+
// headline claiming `0 receipt(s) below TESTED`, on a pair where both receipts
|
|
237
|
+
// were TESTED. An empty reason and a false statement, in the artifact a report's
|
|
238
|
+
// verdicts come from. The refusal's own cause-honest reason was computed into
|
|
239
|
+
// `refusal.reason` and thrown away by the renderer, so nothing a reader could
|
|
240
|
+
// see said why the comparison had stopped.
|
|
241
|
+
//
|
|
242
|
+
// Each half is proved load-bearing by a mutation that patches THIS source in a
|
|
243
|
+
// disposable copy. An earlier revision guarded them with const-true flags and
|
|
244
|
+
// this sentence described those; the flags were removed because a
|
|
245
|
+
// defect-restoring branch resident in the published tree is a hazard, and the
|
|
246
|
+
// sentence outlived them by one commit.
|
|
247
|
+
if (refusal) {
|
|
248
|
+
warnings.push(`verdicts NOT computed β the comparison was REFUSED before any verdict was formed: ${refusal.reason}`);
|
|
249
|
+
}
|
|
250
|
+
if (belowTested.length) {
|
|
251
|
+
warnings.push(`verdicts NOT computed β ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
|
|
252
|
+
}
|
|
176
253
|
// Cross-provider / cross-surface disclosure (Phase 6). A comparison across
|
|
177
254
|
// providers is a skill-DURABILITY comparison across substrates, not model drift
|
|
178
255
|
// over time; across surfaces, sampling control differs. Both are flagged so a
|
|
179
256
|
// reader never mistakes one for the other (see docs/neutrality.html).
|
|
180
257
|
if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
|
|
181
258
|
if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β sampling control differs between surfaces; compare with care.`);
|
|
259
|
+
// The legend for the band markers, printed whenever either side carries one.
|
|
260
|
+
// A two-letter marker a reader cannot decode is worse than no marker.
|
|
261
|
+
// GATED ON WHETHER A LABEL IS ACTUALLY RENDERED, not on whether the data
|
|
262
|
+
// carries a source. A legend explains markers on the page; if `bandStr` emits
|
|
263
|
+
// none, the legend is describing something the reader cannot see. Asked of
|
|
264
|
+
// `bandStr` itself rather than recomputed, so the two cannot disagree.
|
|
265
|
+
const anyLabelRendered = [...Object.values(aB), ...Object.values(bB)]
|
|
266
|
+
.some((x) => x && /\([a-z]+\)\s*$/.test(bandStr(x)));
|
|
267
|
+
if (anyLabelRendered) {
|
|
268
|
+
warnings.push('band provenance: `(generation)` is an ACROSS-DRAW spread β receipt spec v0.5, n generation draws per arm. `(legacy)` is a JUDGE-SAMPLE spread over a single generation, which is what v0.4 and earlier recorded. They are different statistics. The comparison is the only one the older receipt admits, and it is not like for like.');
|
|
269
|
+
}
|
|
182
270
|
if (warnings.length) {
|
|
183
271
|
L.push('> **β Caveats**');
|
|
184
272
|
for (const w of warnings) L.push(`> - ${w}`);
|
|
@@ -190,7 +278,20 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
190
278
|
L.push(`## Headline`);
|
|
191
279
|
L.push('');
|
|
192
280
|
if (!measured) {
|
|
193
|
-
|
|
281
|
+
// THE HEADLINE NAMES THE CAUSE THAT ACTUALLY APPLIES. `belowTested.length`
|
|
282
|
+
// was printed unconditionally, so a refused pair got "0 receipt(s) below
|
|
283
|
+
// TESTED" β a statement measurably false of the receipts it was given.
|
|
284
|
+
if (refusal) {
|
|
285
|
+
// The reason is a complete sentence and already ends by saying no verdict
|
|
286
|
+
// is asserted; prefixing that again produced "REFUSED β no verdict is
|
|
287
|
+
// asserted. the baseline arm did not reproduceβ¦" β a duplicated clause and
|
|
288
|
+
// a lower-case sentence start, in the artifact a published report quotes.
|
|
289
|
+
L.push(`**REFUSED β ${refusal.reason}**`);
|
|
290
|
+
} else if (belowTested.length) {
|
|
291
|
+
L.push(`**NOT MEASURED β ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
|
|
292
|
+
} else {
|
|
293
|
+
L.push('**NOT MEASURED β no verdict is asserted.**');
|
|
294
|
+
}
|
|
194
295
|
L.push('');
|
|
195
296
|
} else {
|
|
196
297
|
L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
|
|
@@ -218,7 +319,12 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
218
319
|
L.push('');
|
|
219
320
|
}
|
|
220
321
|
|
|
221
|
-
return {
|
|
322
|
+
return {
|
|
323
|
+
markdown: L.join('\n'), perCase, headlineDelta, regressions,
|
|
324
|
+
refused: !!refusal,
|
|
325
|
+
refusal_reason: refusal ? refusal.reason : null,
|
|
326
|
+
refusal_key: refusal ? refusal.key : null,
|
|
327
|
+
};
|
|
222
328
|
}
|
|
223
329
|
|
|
224
330
|
module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
|
package/lib/judge.js
CHANGED
|
@@ -99,7 +99,12 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
|
|
|
99
99
|
// { samples:[scores], mean, stddev, reason, judge_settings, model_id, rubric_hash }
|
|
100
100
|
// `mean` Β± `stddev` is the per-case confidence band (raw spread of the N scores)
|
|
101
101
|
// used by the borderline-outcome rule and per-case drift band-overlap logic.
|
|
102
|
-
|
|
102
|
+
// NO DEFAULT TIMEOUT HERE (spec 017 AC-2). This defaulted to 120000, which
|
|
103
|
+
// outranked the per-surface policy exactly as lib/run.js's literal did β so the
|
|
104
|
+
// JUDGE calls timed out on the api policy while running on a CLI surface, which
|
|
105
|
+
// is the second shadowing site and the one #007's prep session had not found.
|
|
106
|
+
// Passing `undefined` through lets lib/provider.js resolve the declared policy.
|
|
107
|
+
async function gradeSamples({ task, response, rubric, model, samples = 5, timeoutMs }) {
|
|
103
108
|
const settings = judgeSettings(samples, model);
|
|
104
109
|
const scores = [];
|
|
105
110
|
const reasons = [];
|