driftproof 0.10.0 → 0.10.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +40 -22
- package/config.js +4 -4
- package/lib/diff.js +61 -22
- package/lib/judge.js +14 -6
- package/lib/receipt.js +1 -1
- package/lib/revision.js +20 -11
- package/lib/stats.js +17 -8
- package/lib/value.js +5 -5
- package/lib/verdict.js +1 -1
- package/package.json +1 -1
- package/spec/RECEIPT.md +11 -5
package/README.md
CHANGED
|
@@ -25,13 +25,13 @@ differ in what moves underneath the skill — or, in the value report, in which
|
|
|
25
25
|
axes are measured; or, in the instrument re-measurement, in the instrument itself:
|
|
26
26
|
|
|
27
27
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
28
|
-
public agent skills across a current-vs-previous Sonnet release; 9 of 10
|
|
29
|
-
|
|
28
|
+
public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
|
|
29
|
+
a separation detected under the rule. ([markdown](reports/report-001.md))
|
|
30
30
|
- **[Report #002](https://driftproofhq.com/reports/002/)** — *substrate
|
|
31
31
|
durability*: the same suites across two vendors' CLIs (Claude vs Codex);
|
|
32
32
|
3 durable, 4 substrate-dependent, 2 regressed, 1 no effect.
|
|
33
33
|
- **[Report #003](https://driftproofhq.com/reports/003/)** — *release drift*:
|
|
34
|
-
`claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4
|
|
34
|
+
`claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 with no separation detected.
|
|
35
35
|
- **[Report #004](https://driftproofhq.com/reports/004/)** — *capability gap*:
|
|
36
36
|
`claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
|
|
37
37
|
3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
|
|
@@ -68,8 +68,9 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
|
|
|
68
68
|
- **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
|
|
69
69
|
two of Report #007's cells re-measured on `claude-fable-5-1` against
|
|
70
70
|
`claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
|
|
71
|
-
identical before the first call. **
|
|
72
|
-
14 cases read 0 improved, 0 regressed, 14
|
|
71
|
+
identical before the first call. **Neither cell showed a separation detected
|
|
72
|
+
under the rule**: the 14 cases read 0 improved, 0 regressed, 14 with no separation
|
|
73
|
+
detected, 0 not measured, which is not evidence that nothing changed. The
|
|
73
74
|
first release pair in this project where both sides are generation-sampled
|
|
74
75
|
receipts, which is what makes the delta attributable to the model rather than
|
|
75
76
|
to the instrument. One case sits inside the verdict on the effect floor alone
|
|
@@ -97,8 +98,10 @@ instead of a stale one.
|
|
|
97
98
|
The hard part isn't running an eval once — it's making the number **credible
|
|
98
99
|
enough to act on**. An LLM judge is noisy, so a naive score can swing run to run
|
|
99
100
|
by more than the drift you're trying to detect. Driftproof's answer is to **sample
|
|
100
|
-
|
|
101
|
-
|
|
101
|
+
and report bands**, each a descriptive spread (the mean plus or minus one sample
|
|
102
|
+
standard deviation), and to **only claim a regression when the bands don't
|
|
103
|
+
overlap**: a separation detected under that rule, not proof. A tool that cries wolf
|
|
104
|
+
is worse than no tool.
|
|
102
105
|
|
|
103
106
|
**A verdict without a price is half an answer.** The same receipts price the
|
|
104
107
|
marginal cost of a skill firing, and Report #005 found the dominant cost driver is
|
|
@@ -139,17 +142,24 @@ effect floor, `NO_EFFECT` when it doesn't, `REGRESSED` when the skill hurts. Eac
|
|
|
139
142
|
case is judged several times, so it carries a **band** (`mean ± stddev`) instead of
|
|
140
143
|
one fragile number, and a case is only ever called regressed/improved when its two
|
|
141
144
|
bands don't overlap. The **effect floor** (0.05, one judge quantization step) is the
|
|
142
|
-
minimum
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
**What the band
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
145
|
+
minimum move required before a separation is called a change: band separation
|
|
146
|
+
*plus* a floor-sized delta, never either alone.
|
|
147
|
+
|
|
148
|
+
**What the band covers.** Since receipt spec v0.5 a run **samples the generation**
|
|
149
|
+
as well as the judge: each arm is drawn at least 3 and at most 10 times
|
|
150
|
+
(`GENERATION_SAMPLES_MIN` and `GENERATION_SAMPLES_MAX` in `config.js`, applied by
|
|
151
|
+
`lib/sampling.js`), and each draw is judged several times. A run stops at 3 draws
|
|
152
|
+
when the across-draw spread is no wider than the 0.05 effect floor; otherwise it
|
|
153
|
+
keeps drawing until two successive estimates of that spread agree to within 0.01, or
|
|
154
|
+
the ceiling is reached, and each case records its `stopping_reason`. So a band on a
|
|
155
|
+
generation-sampled receipt is the spread of the model writing different responses,
|
|
156
|
+
not only of the judge re-scoring one. Report #006 measured that spread directly and
|
|
157
|
+
found it larger than judge-level noise, draw-to-draw **sd 0.186** at the most, which
|
|
158
|
+
is why v0.5 samples it. A receipt from before v0.5 carries a **legacy** band, one
|
|
159
|
+
generation per arm re-scored by the judge, and a comparison report labels each band
|
|
160
|
+
`(legacy)` or `(generation)`. Either way a band describes the draws a run made, at
|
|
161
|
+
the sample size it used, so treat a surprising verdict as **worth re-running**
|
|
162
|
+
before you act on it.
|
|
153
163
|
|
|
154
164
|
### Install
|
|
155
165
|
|
|
@@ -179,6 +189,13 @@ optional and pulled in only for `CLAUDE_PROVIDER=api`. The CLI resolves its spec
|
|
|
179
189
|
schema, and model registry from inside the package, so it runs the same from a
|
|
180
190
|
global/npx install as it does in a checkout.
|
|
181
191
|
|
|
192
|
+
**Platforms.** Driftproof is tested on Linux and macOS. In an outside retest of
|
|
193
|
+
0.10.1 on macOS the shipped gate passed 592 of 611, with 19 failed and 1 not
|
|
194
|
+
applicable, and all nineteen failures are in the publishing helper, which needs
|
|
195
|
+
GNU `realpath -m`. Windows is untested: from Node's source the plugin would not
|
|
196
|
+
find `npx` there, but that failure has never been observed. A CI matrix across
|
|
197
|
+
Linux, macOS and Windows is planned.
|
|
198
|
+
|
|
182
199
|
### Other commands
|
|
183
200
|
|
|
184
201
|
```bash
|
|
@@ -287,7 +304,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
287
304
|
"model_release_date": "2025-10-01",
|
|
288
305
|
"provider": "anthropic",
|
|
289
306
|
"surface": "claude-cli",
|
|
290
|
-
"runner_version": "0.10.
|
|
307
|
+
"runner_version": "0.10.2",
|
|
291
308
|
"date_utc": "2026-07-27T…Z",
|
|
292
309
|
"registry": "registered",
|
|
293
310
|
"transcripts": "hashes-only",
|
|
@@ -329,7 +346,7 @@ Key ideas:
|
|
|
329
346
|
|
|
330
347
|
- **`content_hash` / `suite_hash`** are computed over a canonical JSON form, so the
|
|
331
348
|
same skill and suite hash identically on any machine — receipts are comparable.
|
|
332
|
-
- **Per-case `samples` / `mean` / `stddev`** give each case a
|
|
349
|
+
- **Per-case `samples` / `mean` / `stddev`** give each case a band, a descriptive spread of one sample standard deviation. An
|
|
333
350
|
`outcome` of **`borderline`** means the pass threshold sits *inside* the band —
|
|
334
351
|
the run can't confidently call it pass or fail.
|
|
335
352
|
- **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
|
|
@@ -349,7 +366,8 @@ Key ideas:
|
|
|
349
366
|
`driftproof diff A.json B.json` compares the with-skill bands per case across two
|
|
350
367
|
receipts. The rule that keeps it honest: a **regression** (or improvement) is
|
|
351
368
|
claimed **only when the two bands do not overlap**. Overlapping bands are reported
|
|
352
|
-
as **
|
|
369
|
+
as **no separation detected** at the sample size used: never counted as a
|
|
370
|
+
regression, and never evidence that nothing changed.
|
|
353
371
|
|
|
354
372
|
## Verification in CI (GitHub Action + badge)
|
|
355
373
|
|
|
@@ -365,7 +383,7 @@ jobs:
|
|
|
365
383
|
runs-on: ubuntu-latest
|
|
366
384
|
steps:
|
|
367
385
|
- uses: actions/checkout@v4
|
|
368
|
-
- uses: driftproofhq/driftproof@v0.10.
|
|
386
|
+
- uses: driftproofhq/driftproof@v0.10.2
|
|
369
387
|
with:
|
|
370
388
|
skill-dir: skills/my-skill
|
|
371
389
|
models: claude-haiku-4-5
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.10.
|
|
12
|
+
const RUNNER_VERSION = '0.10.2';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
|
@@ -99,8 +99,8 @@ const DEFAULT_JUDGE_SAMPLES = 5;
|
|
|
99
99
|
// to a coarse ~0.05–0.1 grid, so a confident grade often has stddev 0 (a "point
|
|
100
100
|
// band"). Two point bands that differ by a single quantum (e.g. 0.60 vs 0.64)
|
|
101
101
|
// are technically non-overlapping yet represent no meaningful behaviour change.
|
|
102
|
-
// The floor
|
|
103
|
-
// "
|
|
102
|
+
// The floor keeps such separated but trivial moves from being called a verdict
|
|
103
|
+
// ("below effect floor" in the per-case table). 0.05 = one judge quantum. Documented in
|
|
104
104
|
// spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
|
|
105
105
|
const EFFECT_FLOOR = 0.05;
|
|
106
106
|
|
|
@@ -131,7 +131,7 @@ const REPORT_002_GPT_MODEL = 'gpt-5.6-sol';
|
|
|
131
131
|
const REPORT_002_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
132
132
|
|
|
133
133
|
// Report #003 is a RELEASE DRIFT report (Report #001's type): one provider, two
|
|
134
|
-
// model versions,
|
|
134
|
+
// model versions, a per-skill verdict label (regressed, improved, mixed, or none separated) under
|
|
135
135
|
// the per-case band-separation rule + effect floor. New vs its family predecessor.
|
|
136
136
|
const REPORT_003_NEW_MODEL = 'claude-opus-5';
|
|
137
137
|
const REPORT_003_OLD_MODEL = 'claude-opus-4-8';
|
package/lib/diff.js
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
// SPDX-License-Identifier: Apache-2.0
|
|
2
2
|
'use strict';
|
|
3
3
|
|
|
4
|
-
const { bandVerdict, round } = require('./stats');
|
|
4
|
+
const { bandVerdict, round, WITHIN_NOISE } = require('./stats');
|
|
5
5
|
const { EFFECT_FLOOR } = require('../config');
|
|
6
6
|
const { revisionHeadline } = require('./revision');
|
|
7
7
|
const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
|
|
@@ -9,8 +9,9 @@ const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
|
|
|
9
9
|
// Practical-significance gate applied ON TOP of band separation. bandVerdict()
|
|
10
10
|
// stays a pure geometry test (kept that way so its unit checks are unambiguous);
|
|
11
11
|
// this wrapper additionally requires |delta| >= EFFECT_FLOOR before a verdict is
|
|
12
|
-
// claimed. A separated
|
|
13
|
-
//
|
|
12
|
+
// claimed. A separated but trivial move (below the judge's quantization floor)
|
|
13
|
+
// gets the verdict value WITHIN_NOISE_FLOOR, which the table labels "below effect
|
|
14
|
+
// floor".
|
|
14
15
|
const WITHIN_NOISE_FLOOR = 'within noise (below effect floor)';
|
|
15
16
|
function verdictWithFloor(before, after, delta) {
|
|
16
17
|
const raw = bandVerdict(before.mean, before.stddev, after.mean, after.stddev);
|
|
@@ -19,17 +20,18 @@ function verdictWithFloor(before, after, delta) {
|
|
|
19
20
|
}
|
|
20
21
|
return raw;
|
|
21
22
|
}
|
|
22
|
-
function isWithinNoise(v) { return v ===
|
|
23
|
+
function isWithinNoise(v) { return v === WITHIN_NOISE || v === WITHIN_NOISE_FLOOR; }
|
|
23
24
|
|
|
24
25
|
// Build a drift report (markdown) between two receipts.
|
|
25
26
|
//
|
|
26
|
-
// CREDIBILITY CORE
|
|
27
|
-
// improvement) is claimed ONLY when (1) the two
|
|
28
|
-
// AND (2) the mean moved by at
|
|
29
|
-
// see config.js). When the
|
|
30
|
-
//
|
|
31
|
-
//
|
|
32
|
-
//
|
|
27
|
+
// CREDIBILITY CORE, the anti-false-positive rule: a per-case regression (or
|
|
28
|
+
// improvement) is claimed ONLY when (1) the two bands, each the mean plus or
|
|
29
|
+
// minus one sample standard deviation, do NOT overlap AND (2) the mean moved by at
|
|
30
|
+
// least EFFECT_FLOOR (practical-significance floor, see config.js). When the
|
|
31
|
+
// bands overlap OR the move is below the floor, no separation is detected at the
|
|
32
|
+
// sample size used, and the case is NOT counted as a regression. A tool that cries
|
|
33
|
+
// wolf is worse than useless, so band separation PLUS a real-sized delta, not
|
|
34
|
+
// either alone, is what triggers a verdict.
|
|
33
35
|
|
|
34
36
|
// Map case-id → { mean, stddev } for a receipt's with_skill cases. Falls back to
|
|
35
37
|
// score/0 for v0.1 receipts that have no per-case band.
|
|
@@ -124,14 +126,21 @@ function judgeDiffers(a, b, ids, aB, bB) {
|
|
|
124
126
|
// It deliberately does NOT run a separate band test on the aggregate mean: the
|
|
125
127
|
// aggregate band is suite dispersion, and a separate test there would either cry
|
|
126
128
|
// wolf (if too tight) or mask real per-case drift (if too wide).
|
|
129
|
+
//
|
|
130
|
+
// THE WORDING (spec 031 A-031-20). A band is a descriptive spread, the mean plus or
|
|
131
|
+
// minus one sample standard deviation. A separation is stated as detected under
|
|
132
|
+
// the rule, never as proof that the skill moved; no separation is stated as none
|
|
133
|
+
// detected at this sample size, never as evidence that nothing changed. The
|
|
134
|
+
// leading word is the label; the verdict values it summarises are unchanged.
|
|
127
135
|
function headlineVerdict(perCase) {
|
|
128
136
|
const reg = perCase.filter((r) => r.verdict === 'regression').length;
|
|
129
137
|
const imp = perCase.filter((r) => r.verdict === 'improvement').length;
|
|
130
138
|
const s = (n) => (n === 1 ? '' : 's');
|
|
131
|
-
|
|
132
|
-
if (reg) return `
|
|
133
|
-
if (
|
|
134
|
-
return
|
|
139
|
+
const rule = 'bands do not overlap and the move clears the effect floor';
|
|
140
|
+
if (reg && imp) return `MIXED: separation detected under the rule on ${reg} case${s(reg)} downward and ${imp} case${s(imp)} upward (${rule}).`;
|
|
141
|
+
if (reg) return `DRIFT: separation detected under the rule on ${reg} case${s(reg)}, downward (${rule}); none upward.`;
|
|
142
|
+
if (imp) return `IMPROVED: separation detected under the rule on ${imp} case${s(imp)}, upward (${rule}); none downward.`;
|
|
143
|
+
return 'NO SEPARATION DETECTED: no case separated under the rule at this sample size (band = mean \u00b1 1 sd); this is not evidence that nothing changed.';
|
|
135
144
|
}
|
|
136
145
|
|
|
137
146
|
// A revision pair is a pair in which the SKILL TEXT is the only thing that
|
|
@@ -199,7 +208,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
199
208
|
return { id, before, after, delta, verdict };
|
|
200
209
|
});
|
|
201
210
|
// Sort worst-first: regressions, then by delta.
|
|
202
|
-
const order = { regression: 0,
|
|
211
|
+
const order = { regression: 0, [WITHIN_NOISE]: 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3, refused: 3 };
|
|
203
212
|
perCase.sort((x, y) => (order[x.verdict] - order[y.verdict]) || ((x.delta || 0) - (y.delta || 0)));
|
|
204
213
|
|
|
205
214
|
const aAgg = aggWithBand(a);
|
|
@@ -310,10 +319,33 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
310
319
|
// carries a source. A legend explains markers on the page; if `bandStr` emits
|
|
311
320
|
// none, the legend is describing something the reader cannot see. Asked of
|
|
312
321
|
// `bandStr` itself rather than recomputed, so the two cannot disagree.
|
|
313
|
-
|
|
314
|
-
|
|
315
|
-
|
|
316
|
-
|
|
322
|
+
//
|
|
323
|
+
// THE LEGEND DESCRIBES ONLY THE SOURCES ON THE PAGE, AND THE PAIR IS CALLED
|
|
324
|
+
// NOT LIKE FOR LIKE ONLY WHEN IT IS NOT (spec 031 AC-6). This used to print one
|
|
325
|
+
// sentence whenever any label rendered, naming both sources and saying the
|
|
326
|
+
// comparison was the only one an older receipt admits and was not like for
|
|
327
|
+
// like. Every v0.4 and v0.5 band carries a label, so two v0.5 receipts with
|
|
328
|
+
// `generation` bands on both sides were told they were not like for like, and
|
|
329
|
+
// shown a `(legacy)` source neither carried. The sources are read per receipt,
|
|
330
|
+
// from the markers each side renders; the difference is stated only when the
|
|
331
|
+
// two sides' sets differ, naming which side carries which.
|
|
332
|
+
const renderedSources = (bands) => [...new Set(Object.values(bands)
|
|
333
|
+
.map((x) => (x && (/\(([a-z]+)\)\s*$/.exec(bandStr(x)) || [])[1]) || null).filter(Boolean))].sort();
|
|
334
|
+
const sourcesA = renderedSources(aB);
|
|
335
|
+
const sourcesB = renderedSources(bB);
|
|
336
|
+
const sourcesOnPage = [...new Set([...sourcesA, ...sourcesB])].sort();
|
|
337
|
+
const sourcesDiffer = sourcesA.join() !== sourcesB.join();
|
|
338
|
+
if (sourcesOnPage.length) {
|
|
339
|
+
const LEGEND = {
|
|
340
|
+
generation: '`(generation)` is an ACROSS-DRAW spread: the standard deviation of the case\'s score across n generation draws per arm.',
|
|
341
|
+
legacy: '`(legacy)` is a JUDGE-SAMPLE spread over a single generation: the standard deviation across the judge\'s samples of that one text.',
|
|
342
|
+
};
|
|
343
|
+
const legend = sourcesOnPage.map((s) => LEGEND[s] || `\`(${s})\` is a band source this differ does not describe.`);
|
|
344
|
+
if (sourcesDiffer) {
|
|
345
|
+
const side = (label, s) => `${label} carries ${s.length ? s.map((x) => `\`(${x})\``).join(' and ') : 'no labelled'} bands`;
|
|
346
|
+
legend.push(`The two receipts' bands come from different sources: ${side(labelA, sourcesA)}, ${side(labelB, sourcesB)}. They are different statistics, so a per-case comparison across them is not like for like.`);
|
|
347
|
+
}
|
|
348
|
+
warnings.push(`band provenance: ${legend.join(' ')}`);
|
|
317
349
|
}
|
|
318
350
|
if (warnings.length) {
|
|
319
351
|
L.push('> **⚠ Caveats**');
|
|
@@ -350,7 +382,14 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
350
382
|
} else {
|
|
351
383
|
L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
|
|
352
384
|
L.push('');
|
|
353
|
-
|
|
385
|
+
// ONE LABEL PER CASE, THE TABLE'S (F-5 of
|
|
386
|
+
// specs/031-artefact-claims/evidence/approval-20260915T032341Z.md). A case whose
|
|
387
|
+
// bands do not overlap but whose move is below the floor is labelled below
|
|
388
|
+
// effect floor in the per-case table. This line counted it among the cases with
|
|
389
|
+
// no separation detected and then called the same cases band-separated; it now
|
|
390
|
+
// counts the two apart, as the table labels them. Neither is a separation under
|
|
391
|
+
// the rule, which is what the headline above says.
|
|
392
|
+
L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} → ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin - nFloor} with no separation detected${nFloor ? `, ${nFloor} below the ${EFFECT_FLOOR} effect floor` : ''}.`);
|
|
354
393
|
L.push('');
|
|
355
394
|
}
|
|
356
395
|
|
|
@@ -359,7 +398,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
|
|
|
359
398
|
L.push(`| case | ${labelA} (mean ± sd) | ${labelB} (mean ± sd) | Δ | verdict |`);
|
|
360
399
|
L.push(`|---|---|---|---|---|`);
|
|
361
400
|
for (const r of perCase) {
|
|
362
|
-
const flag = r.verdict === 'regression' ? '🔻 regression' : r.verdict === 'improvement' ? '🔼 improvement' : r.verdict === WITHIN_NOISE_FLOOR ? '
|
|
401
|
+
const flag = r.verdict === 'regression' ? '🔻 regression' : r.verdict === 'improvement' ? '🔼 improvement' : r.verdict === WITHIN_NOISE_FLOOR ? 'below effect floor' : r.verdict === WITHIN_NOISE ? 'no separation detected' : r.verdict === 'not measured' ? 'not measured' : 'n/a';
|
|
363
402
|
L.push(`| \`${r.id}\` | ${r.before ? bandStr(r.before) : 'n/a'} | ${r.after ? bandStr(r.after) : 'n/a'} | ${fmt(r.delta)} | ${flag} |`);
|
|
364
403
|
}
|
|
365
404
|
L.push('');
|
package/lib/judge.js
CHANGED
|
@@ -78,8 +78,15 @@ function promptTemplateHash() {
|
|
|
78
78
|
const TRUNCATION_STOP_REASONS = new Set(['max_tokens', 'length']);
|
|
79
79
|
function isTruncated(stopReason) { return TRUNCATION_STOP_REASONS.has(String(stopReason || '')); }
|
|
80
80
|
|
|
81
|
+
// Only a JSON number is a score (spec 032 AC-1). Number() turned null, false,
|
|
82
|
+
// "" and [] into 0 and true into 1, and each of those became a measured sample:
|
|
83
|
+
// a judge that said nothing scored the worst possible grade, and two receipts
|
|
84
|
+
// over the same generations read as a regression. Every non-number becomes NaN
|
|
85
|
+
// here and is refused by the isFinite line below, which stays the one line every
|
|
86
|
+
// non-number reaches (spec 026's clamp mutation is planted on it).
|
|
81
87
|
function scoreOf(parsed) {
|
|
82
|
-
const
|
|
88
|
+
const raw = parsed ? parsed.score : undefined;
|
|
89
|
+
const x = typeof raw === 'number' ? raw : NaN;
|
|
83
90
|
if (!Number.isFinite(x)) return null;
|
|
84
91
|
if (x < 0 || x > 1) return null;
|
|
85
92
|
return x;
|
|
@@ -128,10 +135,11 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
|
|
|
128
135
|
}
|
|
129
136
|
const score = scoreOf(parsed);
|
|
130
137
|
if (score === null) {
|
|
131
|
-
const x = parsed
|
|
132
|
-
const reason = x === undefined
|
|
133
|
-
:
|
|
134
|
-
: `judge score ${x}
|
|
138
|
+
const x = parsed ? parsed.score : undefined;
|
|
139
|
+
const reason = x === undefined ? 'judge output carries no numeric score'
|
|
140
|
+
: typeof x !== 'number' ? `judge output carries no numeric score (score ${JSON.stringify(x)})`
|
|
141
|
+
: !Number.isFinite(x) ? `judge output carries no numeric score (score ${String(x)})`
|
|
142
|
+
: `judge score ${x} outside [0, 1]`;
|
|
135
143
|
return { ...base, unmeasured: true, reason };
|
|
136
144
|
}
|
|
137
145
|
return { ...base, score, reason: String(parsed.reason || '').slice(0, 300) };
|
|
@@ -143,7 +151,7 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
|
|
|
143
151
|
// samples: a partial sample set must never become a band, and a draw one of
|
|
144
152
|
// whose judge samples carried no score is unmeasured as a whole (spec 026
|
|
145
153
|
// AC-3). The remaining samples are not taken.
|
|
146
|
-
// `mean` ± `stddev` is the per-case
|
|
154
|
+
// `mean` ± `stddev` is the per-case band, a descriptive spread (one sample sd of the N scores)
|
|
147
155
|
// used by the borderline-outcome rule and per-case drift band-overlap logic.
|
|
148
156
|
// NO DEFAULT TIMEOUT HERE (spec 017 AC-2). This defaulted to 120000, which
|
|
149
157
|
// outranked the per-surface policy exactly as lib/run.js's literal did — so the
|
package/lib/receipt.js
CHANGED
|
@@ -114,7 +114,7 @@ function comparisonOf(aggWith, aggBase) {
|
|
|
114
114
|
baseline_score: aggBase.mean_score,
|
|
115
115
|
delta: noCases ? null : round(aggWith.mean_score - aggBase.mean_score),
|
|
116
116
|
// Combined uncertainty of the delta: quadrature sum of the two aggregate
|
|
117
|
-
// bands.
|
|
117
|
+
// bands. A figure beside the delta; the per-case band rule decides the headline.
|
|
118
118
|
delta_uncertainty: noCases ? null : combineUncertainty(aggWith.stddev, aggBase.stddev),
|
|
119
119
|
};
|
|
120
120
|
if (noCases) cmp.delta_uncertainty_unavailable = 'no_cases';
|
package/lib/revision.js
CHANGED
|
@@ -21,29 +21,38 @@ const { EFFECT_FLOOR } = require('../config');
|
|
|
21
21
|
|
|
22
22
|
// ── the cell headline ────────────────────────────────────────────────────────
|
|
23
23
|
// A summary of the per-case band-overlap verdicts, worded about the REVISION.
|
|
24
|
-
// The release-drift headline
|
|
25
|
-
//
|
|
24
|
+
// The release-drift headline is a sentence about a skill under a moving model.
|
|
25
|
+
// Here the model is the control. Worded under spec 031 A-031-20: a separation is
|
|
26
|
+
// detected under the rule, never proof; none detected is not evidence of sameness.
|
|
26
27
|
function revisionHeadline(perCase) {
|
|
27
28
|
const reg = perCase.filter((r) => r.verdict === 'regression').length;
|
|
28
29
|
const imp = perCase.filter((r) => r.verdict === 'improvement').length;
|
|
29
30
|
const s = (n) => (n === 1 ? '' : 's');
|
|
30
31
|
if (reg && imp) {
|
|
31
|
-
return `MIXED
|
|
32
|
+
return `MIXED: separation detected under the rule on ${imp} case${s(imp)} upward and ${reg} downward under the current upstream text (bands do not overlap).`;
|
|
32
33
|
}
|
|
33
34
|
if (reg) {
|
|
34
|
-
return `REVISION REGRESSED
|
|
35
|
+
return `REVISION REGRESSED: separation detected under the rule on ${reg} case${s(reg)}, lower under the current upstream text (bands do not overlap).`;
|
|
35
36
|
}
|
|
36
37
|
if (imp) {
|
|
37
|
-
return `REVISION IMPROVED
|
|
38
|
+
return `REVISION IMPROVED: separation detected under the rule on ${imp} case${s(imp)}, higher under the current upstream text (bands do not overlap); none lower.`;
|
|
38
39
|
}
|
|
39
|
-
return '
|
|
40
|
+
return 'NO SEPARATION DETECTED: no case separated under the rule at this sample size (band = mean \u00b1 1 sd); this is not evidence that the pinned text and the current text behave alike.';
|
|
40
41
|
}
|
|
41
42
|
|
|
42
|
-
// Classification
|
|
43
|
-
//
|
|
43
|
+
// Classification value for a cell. Kept separate so a caller can branch on the
|
|
44
|
+
// class without parsing prose. It used to be the headline's first word, cut at
|
|
45
|
+
// its dash; the headline's wording is now a display matter (spec 031 A-031-20),
|
|
46
|
+
// so the values are stated here and are unchanged: scripts/prepare-report-006.js
|
|
47
|
+
// and spec 009's gate read them. CLASS_NOT_SEPARATED is the value, not a label.
|
|
48
|
+
const CLASS_NOT_SEPARATED = 'WITHIN NOISE';
|
|
44
49
|
function revisionClass(perCase) {
|
|
45
|
-
const
|
|
46
|
-
|
|
50
|
+
const reg = perCase.filter((r) => r.verdict === 'regression').length;
|
|
51
|
+
const imp = perCase.filter((r) => r.verdict === 'improvement').length;
|
|
52
|
+
if (reg && imp) return 'MIXED';
|
|
53
|
+
if (reg) return 'REVISION REGRESSED';
|
|
54
|
+
if (imp) return 'REVISION IMPROVED';
|
|
55
|
+
return CLASS_NOT_SEPARATED;
|
|
47
56
|
}
|
|
48
57
|
|
|
49
58
|
// ── the fairness sentence ────────────────────────────────────────────────────
|
|
@@ -54,7 +63,7 @@ function revisionClass(perCase) {
|
|
|
54
63
|
// one-directional courtesy — the symmetry is a property of the code, and the
|
|
55
64
|
// gate asserts it.
|
|
56
65
|
//
|
|
57
|
-
// A cell
|
|
66
|
+
// A cell with no separation detected gets NO sentence. #005's figure stands unamended, because
|
|
58
67
|
// nothing was measured that would amend it, and manufacturing a hedge for a null
|
|
59
68
|
// result is how a report launders noise into a finding.
|
|
60
69
|
function fairnessSentence({ slug, classification, report005Delta, measuredDelta }) {
|
package/lib/stats.js
CHANGED
|
@@ -1,8 +1,9 @@
|
|
|
1
1
|
// SPDX-License-Identifier: Apache-2.0
|
|
2
2
|
'use strict';
|
|
3
3
|
|
|
4
|
-
// Small statistics helpers for sampled
|
|
5
|
-
//
|
|
4
|
+
// Small statistics helpers for sampled scores and their bands. A band is a
|
|
5
|
+
// descriptive spread, the mean plus or minus one sample standard deviation; it
|
|
6
|
+
// carries no coverage probability. Kept dependency-free and deterministic.
|
|
6
7
|
|
|
7
8
|
function round(n, dp = 6) { const f = Math.pow(10, dp); return Math.round(n * f) / f; }
|
|
8
9
|
|
|
@@ -60,16 +61,24 @@ function aggregateBands(cases) {
|
|
|
60
61
|
return { mean: mean(means), stddev: means.length < 2 ? null : stddev(means) };
|
|
61
62
|
}
|
|
62
63
|
|
|
63
|
-
//
|
|
64
|
-
//
|
|
65
|
-
//
|
|
66
|
-
//
|
|
64
|
+
// THE NOT-SEPARATED VERDICT VALUE, which is a value and not a label (spec 031
|
|
65
|
+
// A-031-20). Callers compare against it, and every display label is mapped from
|
|
66
|
+
// it where it is rendered: lib/diff.js prints "no separation detected". The value
|
|
67
|
+
// itself is unchanged pending the verdict-rule spec.
|
|
68
|
+
const WITHIN_NOISE = 'within noise';
|
|
69
|
+
|
|
70
|
+
// Do two bands (mean ± half-width, the half-width one sample standard deviation)
|
|
71
|
+
// fail to overlap, and in which direction? Returns 'regression' (b below a),
|
|
72
|
+
// 'improvement' (b above a), or WITHIN_NOISE (the bands touch or overlap). This
|
|
73
|
+
// is the anti-false-positive rule: a separation is detected only when the bands
|
|
74
|
+
// are fully apart, and overlap is the absence of a detected separation at the
|
|
75
|
+
// sample size used, not evidence that nothing changed.
|
|
67
76
|
// regression : meanB + hwB < meanA - hwA
|
|
68
77
|
// improvement : meanB - hwB > meanA + hwA
|
|
69
78
|
function bandVerdict(meanA, hwA, meanB, hwB) {
|
|
70
79
|
if (meanB + hwB < meanA - hwA) return 'regression';
|
|
71
80
|
if (meanB - hwB > meanA + hwA) return 'improvement';
|
|
72
|
-
return
|
|
81
|
+
return WITHIN_NOISE;
|
|
73
82
|
}
|
|
74
83
|
|
|
75
84
|
// The variance ratio: how much larger the GENERATION-level spread is than the
|
|
@@ -87,4 +96,4 @@ function varianceRatio(generationSd, judgeSdMean) {
|
|
|
87
96
|
return round(g / j);
|
|
88
97
|
}
|
|
89
98
|
|
|
90
|
-
module.exports = { mean, stddev, stderr, combineUncertainty, aggregateBands, bandVerdict, round, varianceRatio };
|
|
99
|
+
module.exports = { mean, stddev, stderr, combineUncertainty, aggregateBands, bandVerdict, round, varianceRatio, WITHIN_NOISE };
|
package/lib/value.js
CHANGED
|
@@ -25,8 +25,8 @@ const { EFFECT_FLOOR } = require('../config');
|
|
|
25
25
|
//
|
|
26
26
|
// RATIO FRAMINGS ARE FLOOR-GATED. A ratio like "dollars per 0.01 lift" is
|
|
27
27
|
// only meaningful when the benefit it prices is a real move. When the lift did not clear
|
|
28
|
-
// the effect floor (or its bands overlap), the ratio renders "n/a (
|
|
29
|
-
//
|
|
28
|
+
// the effect floor (or its bands overlap), the ratio renders "n/a (no separation
|
|
29
|
+
// detected)", never a number, however tempting the arithmetic. Dividing noise by a cost
|
|
30
30
|
// produces a precise-looking figure with no evidence under it.
|
|
31
31
|
//
|
|
32
32
|
// PRICING IS FROZEN AT RUN TIME. Registry prices change (vendors cut prices; we
|
|
@@ -47,14 +47,14 @@ const COST_BASIS = {
|
|
|
47
47
|
const LATENCY_DISCLOSURE =
|
|
48
48
|
'observed on subscription CLI surface, indicative';
|
|
49
49
|
|
|
50
|
-
const NOISE_CELL = 'n/a (
|
|
50
|
+
const NOISE_CELL = 'n/a (no separation detected)';
|
|
51
51
|
|
|
52
52
|
// A floor-clearing NEGATIVE lift: the skill measurably hurt, so there is no
|
|
53
53
|
// benefit to put a price on.
|
|
54
54
|
const REGRESSED_CELL = 'n/a (skill regressed)';
|
|
55
55
|
|
|
56
56
|
// A cell with at least one separated, floor-clearing DRIVER whose AGGREGATE lift
|
|
57
|
-
// does not clear the floor. It is not
|
|
57
|
+
// does not clear the floor. It is not the no-separation cell: the QA re-derivation found
|
|
58
58
|
// six such cells rendering the noise string while the Verdict basis on the same
|
|
59
59
|
// page named their drivers (QA V-4, 2026-08-19). The aggregate stays the
|
|
60
60
|
// denominator (spec 002 AC-4, DECISIONS #9), so no price is quoted, but the cell
|
|
@@ -461,7 +461,7 @@ function costPerLiftPoint({ lift, separated, incrementalCostPer1kCalls }) {
|
|
|
461
461
|
// Two different absences, two different strings. `separated` means at least
|
|
462
462
|
// one case cleared the floor with non-overlapping bands; when that holds and
|
|
463
463
|
// only the aggregate falls short, the evidence exists and is listed under
|
|
464
|
-
// Verdict basis
|
|
464
|
+
// Verdict basis, and the no-separation cell there would contradict the same page.
|
|
465
465
|
return separated ? DRIVER_ONLY_CELL : NOISE_CELL;
|
|
466
466
|
}
|
|
467
467
|
const cost = Number(incrementalCostPer1kCalls);
|
package/lib/verdict.js
CHANGED
|
@@ -15,7 +15,7 @@ const { EFFECT_FLOOR } = require('../config');
|
|
|
15
15
|
// judge's quantization grid is reported as "no effect", never as "passing".
|
|
16
16
|
//
|
|
17
17
|
// delta >= EFFECT_FLOOR → PASSED (the skill measurably helps)
|
|
18
|
-
// |delta| < EFFECT_FLOOR → NO_EFFECT (
|
|
18
|
+
// |delta| < EFFECT_FLOOR → NO_EFFECT (below the effect floor, one judge quantum)
|
|
19
19
|
// delta <= -EFFECT_FLOOR → REGRESSED (the skill measurably hurts on this model)
|
|
20
20
|
|
|
21
21
|
const VERDICTS = {
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.10.
|
|
3
|
+
"version": "0.10.2",
|
|
4
4
|
"description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
package/spec/RECEIPT.md
CHANGED
|
@@ -3,7 +3,8 @@
|
|
|
3
3
|
|
|
4
4
|
A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
|
|
5
5
|
**with** and **without** the skill on one model version, with the judge **sampled**
|
|
6
|
-
so every score carries a
|
|
6
|
+
so every score carries a band: a descriptive spread, the mean plus or minus one
|
|
7
|
+
sample standard deviation, with no coverage probability. Receipts are the unit of evidence
|
|
7
8
|
Driftproof produces; diffing two receipts across model releases yields a **drift
|
|
8
9
|
report**.
|
|
9
10
|
|
|
@@ -429,10 +430,13 @@ A verdict requires **two** conditions — band separation **and** a minimum effe
|
|
|
429
430
|
1. **Band separation** (geometry):
|
|
430
431
|
- **regression** — `meanB + stddevB < meanA − stddevA` (B's band is entirely below A's)
|
|
431
432
|
- **improvement** — `meanB − stddevB > meanA + stddevA`
|
|
432
|
-
-
|
|
433
|
+
- otherwise no separation is detected (the bands touch or overlap), and a report
|
|
434
|
+
labels the case *no separation detected*; the verdict value the differ records
|
|
435
|
+
is unchanged pending the verdict-rule spec
|
|
433
436
|
2. **Minimum effect floor** — `|meanB − meanA| ≥ EFFECT_FLOOR` (default **0.05**, see
|
|
434
437
|
`config.js`). A regression/improvement from step 1 whose mean moved by *less*
|
|
435
|
-
than the floor
|
|
438
|
+
than the floor keeps a verdict value of its own, labelled *below effect floor*: a
|
|
439
|
+
separation was detected, and it is too small to call.
|
|
436
440
|
|
|
437
441
|
**Why the floor.** The LLM judge quantizes scores to a coarse ~0.05–0.1 grid, so a
|
|
438
442
|
confident grade frequently has `stddev 0` — a zero-width "point band." Two point
|
|
@@ -442,8 +446,10 @@ but-practically-trivial move from being reported as drift. `0.05` ≈ one judge
|
|
|
442
446
|
quantization step; the comparison is inclusive (a move of exactly 0.05 counts).
|
|
443
447
|
|
|
444
448
|
A regression is claimed **only** when both conditions hold. The **headline verdict
|
|
445
|
-
summarizes the per-case verdicts** (e.g. "DRIFT
|
|
446
|
-
|
|
449
|
+
summarizes the per-case verdicts** (e.g. "DRIFT: separation detected under the rule
|
|
450
|
+
on 2 cases, downward", or "NO SEPARATION DETECTED" when no case separated under the
|
|
451
|
+
rule at the sample size used, which is not evidence that nothing changed); it does
|
|
452
|
+
not run a separate, tighter
|
|
447
453
|
test on the aggregate mean. This is the anti-false-positive core: band separation
|
|
448
454
|
**plus a real-sized delta**, not either alone, is what triggers a verdict.
|
|
449
455
|
|