driftproof 0.10.1 → 0.10.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -25,13 +25,13 @@ differ in what moves underneath the skill — or, in the value report, in which
25
25
  axes are measured; or, in the instrument re-measurement, in the instrument itself:
26
26
 
27
27
  - **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
28
- public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
29
- beyond noise. ([markdown](reports/report-001.md))
28
+ public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
29
+ a separation detected under the rule. ([markdown](reports/report-001.md))
30
30
  - **[Report #002](https://driftproofhq.com/reports/002/)** — *substrate
31
31
  durability*: the same suites across two vendors' CLIs (Claude vs Codex);
32
32
  3 durable, 4 substrate-dependent, 2 regressed, 1 no effect.
33
33
  - **[Report #003](https://driftproofhq.com/reports/003/)** — *release drift*:
34
- `claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 within noise.
34
+ `claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 with no separation detected.
35
35
  - **[Report #004](https://driftproofhq.com/reports/004/)** — *capability gap*:
36
36
  `claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
37
37
  3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
@@ -68,8 +68,9 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
68
68
  - **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
69
69
  two of Report #007's cells re-measured on `claude-fable-5-1` against
70
70
  `claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
71
- identical before the first call. **Both cells came back within noise**: the
72
- 14 cases read 0 improved, 0 regressed, 14 within noise, 0 not measured. The
71
+ identical before the first call. **Neither cell showed a separation detected
72
+ under the rule**: the 14 cases read 0 improved, 0 regressed, 14 with no separation
73
+ detected, 0 not measured, which is not evidence that nothing changed. The
73
74
  first release pair in this project where both sides are generation-sampled
74
75
  receipts, which is what makes the delta attributable to the model rather than
75
76
  to the instrument. One case sits inside the verdict on the effect floor alone
@@ -97,8 +98,10 @@ instead of a stale one.
97
98
  The hard part isn't running an eval once — it's making the number **credible
98
99
  enough to act on**. An LLM judge is noisy, so a naive score can swing run to run
99
100
  by more than the drift you're trying to detect. Driftproof's answer is to **sample
100
- the judge and report confidence bands**, and to **only claim a regression when the
101
- bands don't overlap**. A tool that cries wolf is worse than no tool.
101
+ and report bands**, each a descriptive spread (the mean plus or minus one sample
102
+ standard deviation), and to **only claim a regression when the bands don't
103
+ overlap**: a separation detected under that rule, not proof. A tool that cries wolf
104
+ is worse than no tool.
102
105
 
103
106
  **A verdict without a price is half an answer.** The same receipts price the
104
107
  marginal cost of a skill firing, and Report #005 found the dominant cost driver is
@@ -139,17 +142,24 @@ effect floor, `NO_EFFECT` when it doesn't, `REGRESSED` when the skill hurts. Eac
139
142
  case is judged several times, so it carries a **band** (`mean ± stddev`) instead of
140
143
  one fragile number, and a case is only ever called regressed/improved when its two
141
144
  bands don't overlap. The **effect floor** (0.05, one judge quantization step) is the
142
- minimum real move required before a change counts as more than noise — band
143
- separation *plus* a floor-sized delta, never either alone.
144
-
145
- **What the band does not cover.** A verdict rests on **one generation draw per
146
- arm**: the band is the spread of the *judge* re-scoring that single response, not
147
- the spread of the model writing a different one. Report #006 measured the second
148
- directly and found it larger — draw-to-draw spread up to **sd 0.186** on the 0–1
149
- scale, against judge-level noise several times smaller. So treat a surprising
150
- single-run verdict as **provisional and worth re-running** before you act on it.
151
- Generation sampling lands in the next receipt spec; until it does, this is a
152
- limit of the instrument, stated rather than implied.
145
+ minimum move required before a separation is called a change: band separation
146
+ *plus* a floor-sized delta, never either alone.
147
+
148
+ **What the band covers.** Since receipt spec v0.5 a run **samples the generation**
149
+ as well as the judge: each arm is drawn at least 3 and at most 10 times
150
+ (`GENERATION_SAMPLES_MIN` and `GENERATION_SAMPLES_MAX` in `config.js`, applied by
151
+ `lib/sampling.js`), and each draw is judged several times. A run stops at 3 draws
152
+ when the across-draw spread is no wider than the 0.05 effect floor; otherwise it
153
+ keeps drawing until two successive estimates of that spread agree to within 0.01, or
154
+ the ceiling is reached, and each case records its `stopping_reason`. So a band on a
155
+ generation-sampled receipt is the spread of the model writing different responses,
156
+ not only of the judge re-scoring one. Report #006 measured that spread directly and
157
+ found it larger than judge-level noise, draw-to-draw **sd 0.186** at the most, which
158
+ is why v0.5 samples it. A receipt from before v0.5 carries a **legacy** band, one
159
+ generation per arm re-scored by the judge, and a comparison report labels each band
160
+ `(legacy)` or `(generation)`. Either way a band describes the draws a run made, at
161
+ the sample size it used, so treat a surprising verdict as **worth re-running**
162
+ before you act on it.
153
163
 
154
164
  ### Install
155
165
 
@@ -179,6 +189,13 @@ optional and pulled in only for `CLAUDE_PROVIDER=api`. The CLI resolves its spec
179
189
  schema, and model registry from inside the package, so it runs the same from a
180
190
  global/npx install as it does in a checkout.
181
191
 
192
+ **Platforms.** Driftproof is tested on Linux and macOS. In an outside retest of
193
+ 0.10.1 on macOS the shipped gate passed 592 of 611, with 19 failed and 1 not
194
+ applicable, and all nineteen failures are in the publishing helper, which needs
195
+ GNU `realpath -m`. Windows is untested: from Node's source the plugin would not
196
+ find `npx` there, but that failure has never been observed. A CI matrix across
197
+ Linux, macOS and Windows is planned.
198
+
182
199
  ### Other commands
183
200
 
184
201
  ```bash
@@ -287,7 +304,7 @@ A receipt is the unit of evidence — one JSON document conforming to
287
304
  "model_release_date": "2025-10-01",
288
305
  "provider": "anthropic",
289
306
  "surface": "claude-cli",
290
- "runner_version": "0.10.1",
307
+ "runner_version": "0.10.2",
291
308
  "date_utc": "2026-07-27T…Z",
292
309
  "registry": "registered",
293
310
  "transcripts": "hashes-only",
@@ -329,7 +346,7 @@ Key ideas:
329
346
 
330
347
  - **`content_hash` / `suite_hash`** are computed over a canonical JSON form, so the
331
348
  same skill and suite hash identically on any machine — receipts are comparable.
332
- - **Per-case `samples` / `mean` / `stddev`** give each case a confidence band. An
349
+ - **Per-case `samples` / `mean` / `stddev`** give each case a band, a descriptive spread of one sample standard deviation. An
333
350
  `outcome` of **`borderline`** means the pass threshold sits *inside* the band —
334
351
  the run can't confidently call it pass or fail.
335
352
  - **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
@@ -349,7 +366,8 @@ Key ideas:
349
366
  `driftproof diff A.json B.json` compares the with-skill bands per case across two
350
367
  receipts. The rule that keeps it honest: a **regression** (or improvement) is
351
368
  claimed **only when the two bands do not overlap**. Overlapping bands are reported
352
- as **within noise** and never counted as a regression.
369
+ as **no separation detected** at the sample size used: never counted as a
370
+ regression, and never evidence that nothing changed.
353
371
 
354
372
  ## Verification in CI (GitHub Action + badge)
355
373
 
@@ -365,7 +383,7 @@ jobs:
365
383
  runs-on: ubuntu-latest
366
384
  steps:
367
385
  - uses: actions/checkout@v4
368
- - uses: driftproofhq/driftproof@v0.10.1
386
+ - uses: driftproofhq/driftproof@v0.10.2
369
387
  with:
370
388
  skill-dir: skills/my-skill
371
389
  models: claude-haiku-4-5
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.10.1';
12
+ const RUNNER_VERSION = '0.10.2';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
@@ -99,8 +99,8 @@ const DEFAULT_JUDGE_SAMPLES = 5;
99
99
  // to a coarse ~0.05–0.1 grid, so a confident grade often has stddev 0 (a "point
100
100
  // band"). Two point bands that differ by a single quantum (e.g. 0.60 vs 0.64)
101
101
  // are technically non-overlapping yet represent no meaningful behaviour change.
102
- // The floor turns such statistically-separated-but-trivial moves into
103
- // "within noise (below effect floor)". 0.05 = one judge quantum. Documented in
102
+ // The floor keeps such separated but trivial moves from being called a verdict
103
+ // ("below effect floor" in the per-case table). 0.05 = one judge quantum. Documented in
104
104
  // spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
105
105
  const EFFECT_FLOOR = 0.05;
106
106
 
@@ -131,7 +131,7 @@ const REPORT_002_GPT_MODEL = 'gpt-5.6-sol';
131
131
  const REPORT_002_JUDGE_MODEL = 'claude-haiku-4-5';
132
132
 
133
133
  // Report #003 is a RELEASE DRIFT report (Report #001's type): one provider, two
134
- // model versions, verdicts REGRESSED/IMPROVED/MIXED/WITHIN NOISE per skill under
134
+ // model versions, a per-skill verdict label (regressed, improved, mixed, or none separated) under
135
135
  // the per-case band-separation rule + effect floor. New vs its family predecessor.
136
136
  const REPORT_003_NEW_MODEL = 'claude-opus-5';
137
137
  const REPORT_003_OLD_MODEL = 'claude-opus-4-8';
package/lib/diff.js CHANGED
@@ -1,7 +1,7 @@
1
1
  // SPDX-License-Identifier: Apache-2.0
2
2
  'use strict';
3
3
 
4
- const { bandVerdict, round } = require('./stats');
4
+ const { bandVerdict, round, WITHIN_NOISE } = require('./stats');
5
5
  const { EFFECT_FLOOR } = require('../config');
6
6
  const { revisionHeadline } = require('./revision');
7
7
  const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
@@ -9,8 +9,9 @@ const { baselineReproduces, REFUSAL_REASONS, bandOf } = require('./reuse');
9
9
  // Practical-significance gate applied ON TOP of band separation. bandVerdict()
10
10
  // stays a pure geometry test (kept that way so its unit checks are unambiguous);
11
11
  // this wrapper additionally requires |delta| >= EFFECT_FLOOR before a verdict is
12
- // claimed. A separated-but-trivial move (below the judge's quantization floor)
13
- // becomes "within noise (below effect floor)".
12
+ // claimed. A separated but trivial move (below the judge's quantization floor)
13
+ // gets the verdict value WITHIN_NOISE_FLOOR, which the table labels "below effect
14
+ // floor".
14
15
  const WITHIN_NOISE_FLOOR = 'within noise (below effect floor)';
15
16
  function verdictWithFloor(before, after, delta) {
16
17
  const raw = bandVerdict(before.mean, before.stddev, after.mean, after.stddev);
@@ -19,17 +20,18 @@ function verdictWithFloor(before, after, delta) {
19
20
  }
20
21
  return raw;
21
22
  }
22
- function isWithinNoise(v) { return v === 'within noise' || v === WITHIN_NOISE_FLOOR; }
23
+ function isWithinNoise(v) { return v === WITHIN_NOISE || v === WITHIN_NOISE_FLOOR; }
23
24
 
24
25
  // Build a drift report (markdown) between two receipts.
25
26
  //
26
- // CREDIBILITY CORE — the anti-false-positive rule: a per-case regression (or
27
- // improvement) is claimed ONLY when (1) the two confidence bands do NOT overlap
28
- // AND (2) the mean moved by at least EFFECT_FLOOR (practical-significance floor,
29
- // see config.js). When the bands overlap OR the move is below the floor, the
30
- // change is reported as "within noise" and is NOT counted as a regression. A tool
31
- // that cries wolf is worse than useless, so band separation PLUS a real-sized
32
- // delta — not either alone — is what triggers a verdict.
27
+ // CREDIBILITY CORE, the anti-false-positive rule: a per-case regression (or
28
+ // improvement) is claimed ONLY when (1) the two bands, each the mean plus or
29
+ // minus one sample standard deviation, do NOT overlap AND (2) the mean moved by at
30
+ // least EFFECT_FLOOR (practical-significance floor, see config.js). When the
31
+ // bands overlap OR the move is below the floor, no separation is detected at the
32
+ // sample size used, and the case is NOT counted as a regression. A tool that cries
33
+ // wolf is worse than useless, so band separation PLUS a real-sized delta, not
34
+ // either alone, is what triggers a verdict.
33
35
 
34
36
  // Map case-id → { mean, stddev } for a receipt's with_skill cases. Falls back to
35
37
  // score/0 for v0.1 receipts that have no per-case band.
@@ -124,14 +126,21 @@ function judgeDiffers(a, b, ids, aB, bB) {
124
126
  // It deliberately does NOT run a separate band test on the aggregate mean: the
125
127
  // aggregate band is suite dispersion, and a separate test there would either cry
126
128
  // wolf (if too tight) or mask real per-case drift (if too wide).
129
+ //
130
+ // THE WORDING (spec 031 A-031-20). A band is a descriptive spread, the mean plus or
131
+ // minus one sample standard deviation. A separation is stated as detected under
132
+ // the rule, never as proof that the skill moved; no separation is stated as none
133
+ // detected at this sample size, never as evidence that nothing changed. The
134
+ // leading word is the label; the verdict values it summarises are unchanged.
127
135
  function headlineVerdict(perCase) {
128
136
  const reg = perCase.filter((r) => r.verdict === 'regression').length;
129
137
  const imp = perCase.filter((r) => r.verdict === 'improvement').length;
130
138
  const s = (n) => (n === 1 ? '' : 's');
131
- if (reg && imp) return `MIXED — ${reg} case regression${s(reg)} and ${imp} improvement${s(imp)} on non-overlapping bands.`;
132
- if (reg) return `DRIFT — ${reg} case${s(reg)} regressed (bands do not overlap); the skill is measurably weaker on ${reg} case${s(reg)}.`;
133
- if (imp) return `IMPROVED — ${imp} case${s(imp)} improved (bands do not overlap); none regressed.`;
134
- return 'WITHIN NOISE — no case moved beyond its confidence band; the skill holds up.';
139
+ const rule = 'bands do not overlap and the move clears the effect floor';
140
+ if (reg && imp) return `MIXED: separation detected under the rule on ${reg} case${s(reg)} downward and ${imp} case${s(imp)} upward (${rule}).`;
141
+ if (reg) return `DRIFT: separation detected under the rule on ${reg} case${s(reg)}, downward (${rule}); none upward.`;
142
+ if (imp) return `IMPROVED: separation detected under the rule on ${imp} case${s(imp)}, upward (${rule}); none downward.`;
143
+ return 'NO SEPARATION DETECTED: no case separated under the rule at this sample size (band = mean \u00b1 1 sd); this is not evidence that nothing changed.';
135
144
  }
136
145
 
137
146
  // A revision pair is a pair in which the SKILL TEXT is the only thing that
@@ -199,7 +208,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
199
208
  return { id, before, after, delta, verdict };
200
209
  });
201
210
  // Sort worst-first: regressions, then by delta.
202
- const order = { regression: 0, 'within noise': 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3, refused: 3 };
211
+ const order = { regression: 0, [WITHIN_NOISE]: 1, [WITHIN_NOISE_FLOOR]: 1, improvement: 2, 'n/a': 3, 'not measured': 3, refused: 3 };
203
212
  perCase.sort((x, y) => (order[x.verdict] - order[y.verdict]) || ((x.delta || 0) - (y.delta || 0)));
204
213
 
205
214
  const aAgg = aggWithBand(a);
@@ -310,10 +319,33 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
310
319
  // carries a source. A legend explains markers on the page; if `bandStr` emits
311
320
  // none, the legend is describing something the reader cannot see. Asked of
312
321
  // `bandStr` itself rather than recomputed, so the two cannot disagree.
313
- const anyLabelRendered = [...Object.values(aB), ...Object.values(bB)]
314
- .some((x) => x && /\([a-z]+\)\s*$/.test(bandStr(x)));
315
- if (anyLabelRendered) {
316
- warnings.push('band provenance: `(generation)` is an ACROSS-DRAW spread — receipt spec v0.5, n generation draws per arm. `(legacy)` is a JUDGE-SAMPLE spread over a single generation, which is what v0.4 and earlier recorded. They are different statistics. The comparison is the only one the older receipt admits, and it is not like for like.');
322
+ //
323
+ // THE LEGEND DESCRIBES ONLY THE SOURCES ON THE PAGE, AND THE PAIR IS CALLED
324
+ // NOT LIKE FOR LIKE ONLY WHEN IT IS NOT (spec 031 AC-6). This used to print one
325
+ // sentence whenever any label rendered, naming both sources and saying the
326
+ // comparison was the only one an older receipt admits and was not like for
327
+ // like. Every v0.4 and v0.5 band carries a label, so two v0.5 receipts with
328
+ // `generation` bands on both sides were told they were not like for like, and
329
+ // shown a `(legacy)` source neither carried. The sources are read per receipt,
330
+ // from the markers each side renders; the difference is stated only when the
331
+ // two sides' sets differ, naming which side carries which.
332
+ const renderedSources = (bands) => [...new Set(Object.values(bands)
333
+ .map((x) => (x && (/\(([a-z]+)\)\s*$/.exec(bandStr(x)) || [])[1]) || null).filter(Boolean))].sort();
334
+ const sourcesA = renderedSources(aB);
335
+ const sourcesB = renderedSources(bB);
336
+ const sourcesOnPage = [...new Set([...sourcesA, ...sourcesB])].sort();
337
+ const sourcesDiffer = sourcesA.join() !== sourcesB.join();
338
+ if (sourcesOnPage.length) {
339
+ const LEGEND = {
340
+ generation: '`(generation)` is an ACROSS-DRAW spread: the standard deviation of the case\'s score across n generation draws per arm.',
341
+ legacy: '`(legacy)` is a JUDGE-SAMPLE spread over a single generation: the standard deviation across the judge\'s samples of that one text.',
342
+ };
343
+ const legend = sourcesOnPage.map((s) => LEGEND[s] || `\`(${s})\` is a band source this differ does not describe.`);
344
+ if (sourcesDiffer) {
345
+ const side = (label, s) => `${label} carries ${s.length ? s.map((x) => `\`(${x})\``).join(' and ') : 'no labelled'} bands`;
346
+ legend.push(`The two receipts' bands come from different sources: ${side(labelA, sourcesA)}, ${side(labelB, sourcesB)}. They are different statistics, so a per-case comparison across them is not like for like.`);
347
+ }
348
+ warnings.push(`band provenance: ${legend.join(' ')}`);
317
349
  }
318
350
  if (warnings.length) {
319
351
  L.push('> **⚠ Caveats**');
@@ -350,7 +382,14 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
350
382
  } else {
351
383
  L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
352
384
  L.push('');
353
- L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} → ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin} within noise${nFloor ? ` (${nFloor} of them band-separated but below the ${EFFECT_FLOOR} effect floor)` : ''}.`);
385
+ // ONE LABEL PER CASE, THE TABLE'S (F-5 of
386
+ // specs/031-artefact-claims/evidence/approval-20260915T032341Z.md). A case whose
387
+ // bands do not overlap but whose move is below the floor is labelled below
388
+ // effect floor in the per-case table. This line counted it among the cases with
389
+ // no separation detected and then called the same cases band-separated; it now
390
+ // counts the two apart, as the table labels them. Neither is a separation under
391
+ // the rule, which is what the headline above says.
392
+ L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} → ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin - nFloor} with no separation detected${nFloor ? `, ${nFloor} below the ${EFFECT_FLOOR} effect floor` : ''}.`);
354
393
  L.push('');
355
394
  }
356
395
 
@@ -359,7 +398,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' }
359
398
  L.push(`| case | ${labelA} (mean ± sd) | ${labelB} (mean ± sd) | Δ | verdict |`);
360
399
  L.push(`|---|---|---|---|---|`);
361
400
  for (const r of perCase) {
362
- const flag = r.verdict === 'regression' ? '🔻 regression' : r.verdict === 'improvement' ? '🔼 improvement' : r.verdict === WITHIN_NOISE_FLOOR ? 'within noise (below floor)' : r.verdict === 'within noise' ? 'within noise' : r.verdict === 'not measured' ? 'not measured' : 'n/a';
401
+ const flag = r.verdict === 'regression' ? '🔻 regression' : r.verdict === 'improvement' ? '🔼 improvement' : r.verdict === WITHIN_NOISE_FLOOR ? 'below effect floor' : r.verdict === WITHIN_NOISE ? 'no separation detected' : r.verdict === 'not measured' ? 'not measured' : 'n/a';
363
402
  L.push(`| \`${r.id}\` | ${r.before ? bandStr(r.before) : 'n/a'} | ${r.after ? bandStr(r.after) : 'n/a'} | ${fmt(r.delta)} | ${flag} |`);
364
403
  }
365
404
  L.push('');
package/lib/judge.js CHANGED
@@ -78,8 +78,15 @@ function promptTemplateHash() {
78
78
  const TRUNCATION_STOP_REASONS = new Set(['max_tokens', 'length']);
79
79
  function isTruncated(stopReason) { return TRUNCATION_STOP_REASONS.has(String(stopReason || '')); }
80
80
 
81
+ // Only a JSON number is a score (spec 032 AC-1). Number() turned null, false,
82
+ // "" and [] into 0 and true into 1, and each of those became a measured sample:
83
+ // a judge that said nothing scored the worst possible grade, and two receipts
84
+ // over the same generations read as a regression. Every non-number becomes NaN
85
+ // here and is refused by the isFinite line below, which stays the one line every
86
+ // non-number reaches (spec 026's clamp mutation is planted on it).
81
87
  function scoreOf(parsed) {
82
- const x = Number(parsed && parsed.score);
88
+ const raw = parsed ? parsed.score : undefined;
89
+ const x = typeof raw === 'number' ? raw : NaN;
83
90
  if (!Number.isFinite(x)) return null;
84
91
  if (x < 0 || x > 1) return null;
85
92
  return x;
@@ -128,10 +135,11 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
128
135
  }
129
136
  const score = scoreOf(parsed);
130
137
  if (score === null) {
131
- const x = parsed && parsed.score;
132
- const reason = x === undefined || x === null ? 'judge output carries no numeric score'
133
- : !Number.isFinite(Number(x)) ? `judge output carries no numeric score (score ${JSON.stringify(x)})`
134
- : `judge score ${x} outside [0, 1]`;
138
+ const x = parsed ? parsed.score : undefined;
139
+ const reason = x === undefined ? 'judge output carries no numeric score'
140
+ : typeof x !== 'number' ? `judge output carries no numeric score (score ${JSON.stringify(x)})`
141
+ : !Number.isFinite(x) ? `judge output carries no numeric score (score ${String(x)})`
142
+ : `judge score ${x} outside [0, 1]`;
135
143
  return { ...base, unmeasured: true, reason };
136
144
  }
137
145
  return { ...base, score, reason: String(parsed.reason || '').slice(0, 300) };
@@ -143,7 +151,7 @@ async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature
143
151
  // samples: a partial sample set must never become a band, and a draw one of
144
152
  // whose judge samples carried no score is unmeasured as a whole (spec 026
145
153
  // AC-3). The remaining samples are not taken.
146
- // `mean` ± `stddev` is the per-case confidence band (raw spread of the N scores)
154
+ // `mean` ± `stddev` is the per-case band, a descriptive spread (one sample sd of the N scores)
147
155
  // used by the borderline-outcome rule and per-case drift band-overlap logic.
148
156
  // NO DEFAULT TIMEOUT HERE (spec 017 AC-2). This defaulted to 120000, which
149
157
  // outranked the per-surface policy exactly as lib/run.js's literal did — so the
package/lib/receipt.js CHANGED
@@ -114,7 +114,7 @@ function comparisonOf(aggWith, aggBase) {
114
114
  baseline_score: aggBase.mean_score,
115
115
  delta: noCases ? null : round(aggWith.mean_score - aggBase.mean_score),
116
116
  // Combined uncertainty of the delta: quadrature sum of the two aggregate
117
- // bands. Diff uses this for the headline "within noise" vs real-move rule.
117
+ // bands. A figure beside the delta; the per-case band rule decides the headline.
118
118
  delta_uncertainty: noCases ? null : combineUncertainty(aggWith.stddev, aggBase.stddev),
119
119
  };
120
120
  if (noCases) cmp.delta_uncertainty_unavailable = 'no_cases';
package/lib/revision.js CHANGED
@@ -21,29 +21,38 @@ const { EFFECT_FLOOR } = require('../config');
21
21
 
22
22
  // ── the cell headline ────────────────────────────────────────────────────────
23
23
  // A summary of the per-case band-overlap verdicts, worded about the REVISION.
24
- // The release-drift headline says "the skill is measurably weaker", which is a
25
- // sentence about a skill under a moving model. Here the model is the control.
24
+ // The release-drift headline is a sentence about a skill under a moving model.
25
+ // Here the model is the control. Worded under spec 031 A-031-20: a separation is
26
+ // detected under the rule, never proof; none detected is not evidence of sameness.
26
27
  function revisionHeadline(perCase) {
27
28
  const reg = perCase.filter((r) => r.verdict === 'regression').length;
28
29
  const imp = perCase.filter((r) => r.verdict === 'improvement').length;
29
30
  const s = (n) => (n === 1 ? '' : 's');
30
31
  if (reg && imp) {
31
- return `MIXED — the revision improved ${imp} case${s(imp)} and regressed ${reg} on non-overlapping bands.`;
32
+ return `MIXED: separation detected under the rule on ${imp} case${s(imp)} upward and ${reg} downward under the current upstream text (bands do not overlap).`;
32
33
  }
33
34
  if (reg) {
34
- return `REVISION REGRESSED — ${reg} case${s(reg)} scored lower under the current upstream text (bands do not overlap).`;
35
+ return `REVISION REGRESSED: separation detected under the rule on ${reg} case${s(reg)}, lower under the current upstream text (bands do not overlap).`;
35
36
  }
36
37
  if (imp) {
37
- return `REVISION IMPROVED — ${imp} case${s(imp)} scored higher under the current upstream text (bands do not overlap); none regressed.`;
38
+ return `REVISION IMPROVED: separation detected under the rule on ${imp} case${s(imp)}, higher under the current upstream text (bands do not overlap); none lower.`;
38
39
  }
39
- return 'WITHIN NOISE — the revision moved no case beyond its confidence band; the pinned text and the current text measure the same.';
40
+ return 'NO SEPARATION DETECTED: no case separated under the rule at this sample size (band = mean \u00b1 1 sd); this is not evidence that the pinned text and the current text behave alike.';
40
41
  }
41
42
 
42
- // Classification word for a cell, from its headline. Kept separate so a caller
43
- // can branch on the class without parsing prose.
43
+ // Classification value for a cell. Kept separate so a caller can branch on the
44
+ // class without parsing prose. It used to be the headline's first word, cut at
45
+ // its dash; the headline's wording is now a display matter (spec 031 A-031-20),
46
+ // so the values are stated here and are unchanged: scripts/prepare-report-006.js
47
+ // and spec 009's gate read them. CLASS_NOT_SEPARATED is the value, not a label.
48
+ const CLASS_NOT_SEPARATED = 'WITHIN NOISE';
44
49
  function revisionClass(perCase) {
45
- const h = revisionHeadline(perCase);
46
- return h.split(' —')[0];
50
+ const reg = perCase.filter((r) => r.verdict === 'regression').length;
51
+ const imp = perCase.filter((r) => r.verdict === 'improvement').length;
52
+ if (reg && imp) return 'MIXED';
53
+ if (reg) return 'REVISION REGRESSED';
54
+ if (imp) return 'REVISION IMPROVED';
55
+ return CLASS_NOT_SEPARATED;
47
56
  }
48
57
 
49
58
  // ── the fairness sentence ────────────────────────────────────────────────────
@@ -54,7 +63,7 @@ function revisionClass(perCase) {
54
63
  // one-directional courtesy — the symmetry is a property of the code, and the
55
64
  // gate asserts it.
56
65
  //
57
- // A cell within noise gets NO sentence. #005's figure stands unamended, because
66
+ // A cell with no separation detected gets NO sentence. #005's figure stands unamended, because
58
67
  // nothing was measured that would amend it, and manufacturing a hedge for a null
59
68
  // result is how a report launders noise into a finding.
60
69
  function fairnessSentence({ slug, classification, report005Delta, measuredDelta }) {
package/lib/stats.js CHANGED
@@ -1,8 +1,9 @@
1
1
  // SPDX-License-Identifier: Apache-2.0
2
2
  'use strict';
3
3
 
4
- // Small statistics helpers for sampled judging and confidence bands. Kept
5
- // dependency-free and deterministic.
4
+ // Small statistics helpers for sampled scores and their bands. A band is a
5
+ // descriptive spread, the mean plus or minus one sample standard deviation; it
6
+ // carries no coverage probability. Kept dependency-free and deterministic.
6
7
 
7
8
  function round(n, dp = 6) { const f = Math.pow(10, dp); return Math.round(n * f) / f; }
8
9
 
@@ -60,16 +61,24 @@ function aggregateBands(cases) {
60
61
  return { mean: mean(means), stddev: means.length < 2 ? null : stddev(means) };
61
62
  }
62
63
 
63
- // Do two confidence bands (mean ± half-width) fail to overlap, and in which
64
- // direction? Returns 'regression' (b below a), 'improvement' (b above a), or
65
- // 'within noise' (bands touch/overlap). This is the anti-false-positive rule:
66
- // a change is only claimed when the bands are fully separated.
64
+ // THE NOT-SEPARATED VERDICT VALUE, which is a value and not a label (spec 031
65
+ // A-031-20). Callers compare against it, and every display label is mapped from
66
+ // it where it is rendered: lib/diff.js prints "no separation detected". The value
67
+ // itself is unchanged pending the verdict-rule spec.
68
+ const WITHIN_NOISE = 'within noise';
69
+
70
+ // Do two bands (mean ± half-width, the half-width one sample standard deviation)
71
+ // fail to overlap, and in which direction? Returns 'regression' (b below a),
72
+ // 'improvement' (b above a), or WITHIN_NOISE (the bands touch or overlap). This
73
+ // is the anti-false-positive rule: a separation is detected only when the bands
74
+ // are fully apart, and overlap is the absence of a detected separation at the
75
+ // sample size used, not evidence that nothing changed.
67
76
  // regression : meanB + hwB < meanA - hwA
68
77
  // improvement : meanB - hwB > meanA + hwA
69
78
  function bandVerdict(meanA, hwA, meanB, hwB) {
70
79
  if (meanB + hwB < meanA - hwA) return 'regression';
71
80
  if (meanB - hwB > meanA + hwA) return 'improvement';
72
- return 'within noise';
81
+ return WITHIN_NOISE;
73
82
  }
74
83
 
75
84
  // The variance ratio: how much larger the GENERATION-level spread is than the
@@ -87,4 +96,4 @@ function varianceRatio(generationSd, judgeSdMean) {
87
96
  return round(g / j);
88
97
  }
89
98
 
90
- module.exports = { mean, stddev, stderr, combineUncertainty, aggregateBands, bandVerdict, round, varianceRatio };
99
+ module.exports = { mean, stddev, stderr, combineUncertainty, aggregateBands, bandVerdict, round, varianceRatio, WITHIN_NOISE };
package/lib/value.js CHANGED
@@ -25,8 +25,8 @@ const { EFFECT_FLOOR } = require('../config');
25
25
  //
26
26
  // RATIO FRAMINGS ARE FLOOR-GATED. A ratio like "dollars per 0.01 lift" is
27
27
  // only meaningful when the benefit it prices is a real move. When the lift did not clear
28
- // the effect floor (or its bands overlap), the ratio renders "n/a (within noise)"
29
- // — never a number, however tempting the arithmetic. Dividing noise by a cost
28
+ // the effect floor (or its bands overlap), the ratio renders "n/a (no separation
29
+ // detected)", never a number, however tempting the arithmetic. Dividing noise by a cost
30
30
  // produces a precise-looking figure with no evidence under it.
31
31
  //
32
32
  // PRICING IS FROZEN AT RUN TIME. Registry prices change (vendors cut prices; we
@@ -47,14 +47,14 @@ const COST_BASIS = {
47
47
  const LATENCY_DISCLOSURE =
48
48
  'observed on subscription CLI surface, indicative';
49
49
 
50
- const NOISE_CELL = 'n/a (within noise)';
50
+ const NOISE_CELL = 'n/a (no separation detected)';
51
51
 
52
52
  // A floor-clearing NEGATIVE lift: the skill measurably hurt, so there is no
53
53
  // benefit to put a price on.
54
54
  const REGRESSED_CELL = 'n/a (skill regressed)';
55
55
 
56
56
  // A cell with at least one separated, floor-clearing DRIVER whose AGGREGATE lift
57
- // does not clear the floor. It is not "within noise" — the QA re-derivation found
57
+ // does not clear the floor. It is not the no-separation cell: the QA re-derivation found
58
58
  // six such cells rendering the noise string while the Verdict basis on the same
59
59
  // page named their drivers (QA V-4, 2026-08-19). The aggregate stays the
60
60
  // denominator (spec 002 AC-4, DECISIONS #9), so no price is quoted, but the cell
@@ -461,7 +461,7 @@ function costPerLiftPoint({ lift, separated, incrementalCostPer1kCalls }) {
461
461
  // Two different absences, two different strings. `separated` means at least
462
462
  // one case cleared the floor with non-overlapping bands; when that holds and
463
463
  // only the aggregate falls short, the evidence exists and is listed under
464
- // Verdict basis — saying "within noise" there contradicts the same page.
464
+ // Verdict basis, and the no-separation cell there would contradict the same page.
465
465
  return separated ? DRIVER_ONLY_CELL : NOISE_CELL;
466
466
  }
467
467
  const cost = Number(incrementalCostPer1kCalls);
package/lib/verdict.js CHANGED
@@ -15,7 +15,7 @@ const { EFFECT_FLOOR } = require('../config');
15
15
  // judge's quantization grid is reported as "no effect", never as "passing".
16
16
  //
17
17
  // delta >= EFFECT_FLOOR → PASSED (the skill measurably helps)
18
- // |delta| < EFFECT_FLOOR → NO_EFFECT (within the judge's noise floor)
18
+ // |delta| < EFFECT_FLOOR → NO_EFFECT (below the effect floor, one judge quantum)
19
19
  // delta <= -EFFECT_FLOOR → REGRESSED (the skill measurably hurts on this model)
20
20
 
21
21
  const VERDICTS = {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "driftproof",
3
- "version": "0.10.1",
3
+ "version": "0.10.2",
4
4
  "description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
package/spec/RECEIPT.md CHANGED
@@ -3,7 +3,8 @@
3
3
 
4
4
  A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
5
5
  **with** and **without** the skill on one model version, with the judge **sampled**
6
- so every score carries a confidence band. Receipts are the unit of evidence
6
+ so every score carries a band: a descriptive spread, the mean plus or minus one
7
+ sample standard deviation, with no coverage probability. Receipts are the unit of evidence
7
8
  Driftproof produces; diffing two receipts across model releases yields a **drift
8
9
  report**.
9
10
 
@@ -429,10 +430,13 @@ A verdict requires **two** conditions — band separation **and** a minimum effe
429
430
  1. **Band separation** (geometry):
430
431
  - **regression** — `meanB + stddevB < meanA − stddevA` (B's band is entirely below A's)
431
432
  - **improvement** — `meanB − stddevB > meanA + stddevA`
432
- - **within noise** — otherwise (the bands touch or overlap)
433
+ - otherwise no separation is detected (the bands touch or overlap), and a report
434
+ labels the case *no separation detected*; the verdict value the differ records
435
+ is unchanged pending the verdict-rule spec
433
436
  2. **Minimum effect floor** — `|meanB − meanA| ≥ EFFECT_FLOOR` (default **0.05**, see
434
437
  `config.js`). A regression/improvement from step 1 whose mean moved by *less*
435
- than the floor is downgraded to **`within noise (below effect floor)`**.
438
+ than the floor keeps a verdict value of its own, labelled *below effect floor*: a
439
+ separation was detected, and it is too small to call.
436
440
 
437
441
  **Why the floor.** The LLM judge quantizes scores to a coarse ~0.05–0.1 grid, so a
438
442
  confident grade frequently has `stddev 0` — a zero-width "point band." Two point
@@ -442,8 +446,10 @@ but-practically-trivial move from being reported as drift. `0.05` ≈ one judge
442
446
  quantization step; the comparison is inclusive (a move of exactly 0.05 counts).
443
447
 
444
448
  A regression is claimed **only** when both conditions hold. The **headline verdict
445
- summarizes the per-case verdicts** (e.g. "DRIFT — 2 cases regressed", or "WITHIN
446
- NOISE" when none moved beyond its band + floor); it does not run a separate, tighter
449
+ summarizes the per-case verdicts** (e.g. "DRIFT: separation detected under the rule
450
+ on 2 cases, downward", or "NO SEPARATION DETECTED" when no case separated under the
451
+ rule at the sample size used, which is not evidence that nothing changed); it does
452
+ not run a separate, tighter
447
453
  test on the aggregate mean. This is the anti-false-positive core: band separation
448
454
  **plus a real-sized delta**, not either alone, is what triggers a verdict.
449
455