driftproof 0.4.0 β†’ 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -9,15 +9,17 @@
9
9
  Driftproof is an open **receipt spec** plus a **runner** that measures whether an
10
10
  agent skill actually helps β€” by running the skill's eval suite **with** and
11
11
  **without** the skill on a named model version, judging each case several times to
12
- get a confidence band, and emitting a signed, dated **receipt**. Diff two receipts
12
+ get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
13
13
  across model releases and you get a **drift report**.
14
14
 
15
15
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
16
16
  format; it does not invent its own.
17
17
 
18
- πŸ“Š **Four published reports** (each re-derived from committed receipts, nothing
19
- hand-entered), spanning three report types that share one band-based, floor-gated
20
- verdict rule and differ only in what moves underneath the skill:
18
+ πŸ“Š **Six published reports** (each re-derived from committed receipts, nothing
19
+ hand-entered), spanning five published report types, the fifth being revision
20
+ drift. All six share one band-based, floor-gated verdict rule and differ in what
21
+ moves underneath the skill β€” or, in the value report, in which axes are
22
+ measured:
21
23
 
22
24
  - **[Report #001](https://driftproofhq.com/reports/001/)** β€” *release drift*: ten
23
25
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
@@ -31,10 +33,29 @@ verdict rule and differ only in what moves underneath the skill:
31
33
  `claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
32
34
  3 durable, 5 tier-dependent, 0 regressions, 2 no effect β€” encoded expertise
33
35
  survives the frontier tier.
36
+ - **[Report #005](https://driftproofhq.com/reports/005/)** β€” *value*: what a skill
37
+ *costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
38
+ three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) β€”
39
+ 14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
40
+ saving instead, having improved quality while reducing cost.
41
+ *Three of those cells carry an amendment (v1.1, applied when Report #006
42
+ published): their lifts rest on single-draw baselines since shown
43
+ unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
44
+ substrate-disagreement findings below are unaffected.*
45
+ - **[Report #006](https://driftproofhq.com/reports/006/)** β€” *revision drift*: the pinned skill revision against the one
46
+ upstream ships today, on a held substrate. **The reuse premise was tested and
47
+ refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
48
+ control. A 120-call probe found generation-level sampling noise 3.2Γ— and 7.5Γ—
49
+ larger than the judge-level noise this instrument actually samples β€” enough to
50
+ account for every gap the controls saw without any other cause being
51
+ established β€” and the receipt spec gains generation sampling as a result. **No
52
+ cause is asserted**; the control proves non-reproduction and cannot say why.
53
+ *The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
34
54
 
35
55
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
36
- skills](https://driftproofhq.com/writing/three-releases/)**, reads the first three
37
- reports together.
56
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all six reports
57
+ together: what moves underneath a skill, and what the skill costs to run. Revised
58
+ 2026-08-29; every figure in it is gate-checked against the report page it cites.
38
59
 
39
60
  ## Why
40
61
 
@@ -55,6 +76,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
55
76
  the judge and report confidence bands**, and to **only claim a regression when the
56
77
  bands don't overlap**. A tool that cries wolf is worse than no tool.
57
78
 
79
+ **A verdict without a price is half an answer.** The same receipts price the
80
+ marginal cost of a skill firing, and Report #005 found the dominant cost driver is
81
+ not the skill's own text but the input it causes the model to pull in: across those
82
+ 30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
83
+ tracks it at only `r = +0.33`, and one 738-token skill drew 34Γ— its own size in
84
+ extra input. Identical token deltas also price very differently across substrates β€”
85
+ the same skill at near-identical deltas costs 3.3Γ— more on `claude-fable-5` than on
86
+ `claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
87
+
58
88
  ## Quickstart β€” receipt for your own skill in ~10 minutes
59
89
 
60
90
  You need Node β‰₯ 22 and an `ANTHROPIC_API_KEY`.
@@ -86,6 +116,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
86
116
  minimum real move required before a change counts as more than noise β€” band
87
117
  separation *plus* a floor-sized delta, never either alone.
88
118
 
119
+ **What the band does not cover.** A verdict rests on **one generation draw per
120
+ arm**: the band is the spread of the *judge* re-scoring that single response, not
121
+ the spread of the model writing a different one. Report #006 measured the second
122
+ directly and found it larger β€” draw-to-draw spread up to **sd 0.186** on the 0–1
123
+ scale, against judge-level noise several times smaller. So treat a surprising
124
+ single-run verdict as **provisional and worth re-running** before you act on it.
125
+ Generation sampling lands in the next receipt spec; until it does, this is a
126
+ limit of the instrument, stated rather than implied.
127
+
89
128
  ### Install
90
129
 
91
130
  ```bash
@@ -162,7 +201,7 @@ A receipt is the unit of evidence β€” one JSON document conforming to
162
201
 
163
202
  ```jsonc
164
203
  {
165
- "schema_version": "0.3.1",
204
+ "schema_version": "0.4",
166
205
  "skill": { "name": "commit-message-conventions", "version": "0.2.0",
167
206
  "content_hash": "…sha256 over SKILL.md + bundled files…" },
168
207
  "suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
@@ -171,7 +210,7 @@ A receipt is the unit of evidence β€” one JSON document conforming to
171
210
  "model_release_date": "2025-10-01",
172
211
  "provider": "anthropic",
173
212
  "surface": "claude-cli",
174
- "runner_version": "0.4.0",
213
+ "runner_version": "0.6.0",
175
214
  "date_utc": "2026-07-27T…Z",
176
215
  "registry": "registered",
177
216
  "transcripts": "hashes-only",
@@ -194,6 +233,12 @@ A receipt is the unit of evidence β€” one JSON document conforming to
194
233
  },
195
234
  "comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
196
235
  "delta": 0.39, "delta_uncertainty": 0.036 },
236
+ // v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
237
+ // freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
238
+ // holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
239
+ // mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
240
+ // skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
241
+ // output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
197
242
  "verification_level": "TESTED",
198
243
  "receipt_hash": "…sha256 of the canonical receipt with this field removed…"
199
244
  }
@@ -209,6 +254,12 @@ Key ideas:
209
254
  - **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
210
255
  - **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
211
256
  `TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
257
+ - **`economics`** is *derived*, never a second measurement: the token delta is the
258
+ durable fact, and the dollars are exactly those tokens at the rates frozen into
259
+ `run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
260
+ Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
261
+ figure (`judge_excluded` is `const true`) β€” measuring the skill is our cost, not
262
+ the skill's.
212
263
  - **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
213
264
  signature β€” see the spec's open questions).
214
265
 
@@ -233,7 +284,7 @@ jobs:
233
284
  runs-on: ubuntu-latest
234
285
  steps:
235
286
  - uses: actions/checkout@v4
236
- - uses: driftproofhq/driftproof@v0.4.0
287
+ - uses: driftproofhq/driftproof@v0.6.0
237
288
  with:
238
289
  skill-dir: skills/my-skill
239
290
  models: claude-haiku-4-5
@@ -272,11 +323,14 @@ site, so it reflects a real dated run, not a hand-set color.
272
323
 
273
324
  ## Reports
274
325
 
275
- Four reports are published, spanning three report types (see the roll at the top
276
- of this README, and [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome
277
- every report inherits). Each report and every verdict in it are **re-derived from
278
- the receipts** committed under [`receipts/`](receipts/)
279
- (`receipts/report-001/` … `receipts/report-004/`) β€” nothing is hand-entered.
326
+ Six reports are published, spanning five report types. A report page lives at a
327
+ draft path β€” `docs/reports/NNN-draft/` β€” until the publish sequence renames it, and
328
+ `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
329
+ (see the roll at the top of this README, and
330
+ [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
331
+ Each report and every verdict in it are **re-derived from the receipts** committed
332
+ under [`receipts/`](receipts/)
333
+ (`receipts/report-001/` … `receipts/report-006/`) β€” nothing is hand-entered.
280
334
 
281
335
  Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
282
336
  fetched at run time from a pinned commit and verified by sha256 against
@@ -290,9 +344,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models Γ— with/basel
290
344
  node scripts/build-report-001.js # re-derive the report from the receipts
291
345
  ```
292
346
 
293
- Reports #002–#004 have their own runners
294
- (`scripts/prepare-report-00N.js`) following the same
295
- fetch β†’ run β†’ re-derive shape.
347
+ Reports #002–#005 have their own runners
348
+ (`scripts/prepare-report-00N.js` β€” Report #005's is
349
+ [`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
350
+ same fetch β†’ run β†’ re-derive shape.
296
351
 
297
352
  **Model-release triggers are live**: `scripts/release-watch.js` (keyless β€” it
298
353
  reads the public models registry) notices a new model release, re-runs the
package/bin/driftproof CHANGED
@@ -8,7 +8,7 @@ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = req
8
8
  const { loadSkill } = require('../lib/skill');
9
9
  const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
10
10
  const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
11
- const { buildDriftReport } = require('../lib/diff');
11
+ const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
12
12
  const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
13
13
  const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
14
14
  const { registryStatus } = require('../lib/models');
@@ -76,7 +76,7 @@ USAGE
76
76
  ${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
77
77
  [--judge-model M] [--concurrency N] [--max-usd N]
78
78
  [--keep-transcripts] [--out DIR]
79
- ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
79
+ ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
80
80
  ${PROJECT_NAME} validate <receipt.json>
81
81
  ${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
82
82
  ${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
@@ -250,9 +250,32 @@ function cmdDiff(positional, flags) {
250
250
  if (!verifyReceiptHash(r)) console.error(` ⚠ ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
251
251
  }
252
252
 
253
- const labelA = dateStamp(a.run.date_utc);
254
- const labelB = dateStamp(b.run.date_utc);
255
- const { markdown } = buildDriftReport(a, b, { labelA, labelB });
253
+ // --mode revision inverts the axis: the skill text is the variable under test
254
+ // and the substrate is the control. The fields release drift merely warns about
255
+ // are preconditions here, so a pair that is not a revision pair is REFUSED
256
+ // (exit 6) rather than rendered with a caveat nobody reads. A differing model
257
+ // would be release drift wearing a revision label β€” the one confound this mode
258
+ // exists to exclude β€” and an EQUAL content_hash has no revision to measure.
259
+ const mode = flags.mode || 'release';
260
+ if (mode !== 'release' && mode !== 'revision') {
261
+ console.error(`unknown --mode "${mode}" β€” supported: release, revision`);
262
+ process.exit(2);
263
+ }
264
+ if (mode === 'revision') {
265
+ const problem = revisionPairProblem(a, b);
266
+ if (problem) {
267
+ const why = problem === 'skill.content_hash'
268
+ ? 'the two receipts carry the SAME skill.content_hash β€” there is no revision between them to measure'
269
+ : `${problem} differs between the two receipts β€” revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
270
+ console.error(` βœ— REFUSED (--mode revision): ${why}.`);
271
+ console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
272
+ process.exit(6);
273
+ }
274
+ }
275
+
276
+ const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
277
+ const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
278
+ const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
256
279
 
257
280
  if (flags.out) {
258
281
  fs.writeFileSync(path.resolve(flags.out), markdown);
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.4.0';
12
+ const RUNNER_VERSION = '0.6.0';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
@@ -22,7 +22,14 @@ const SUITE_FORMAT = 'agentskills.io/evals';
22
22
  // surface enums (openai-api/openai-cli), optional run.surface_overhead_note,
23
23
  // optional per-case checks[] (deterministic post-checks), and optional
24
24
  // skill.tokens (value-per-token axis). v0.1/v0.2/v0.3 receipts still load.
25
- const RECEIPT_SCHEMA_VERSION = '0.3.1';
25
+ // v0.4 (additive over v0.3.1) adds the ECONOMICS axis: per-case, per-arm
26
+ // generation `usage` (input/output/cached tokens + measured wall_ms), a
27
+ // separate per-case `judge_usage` (measurement overhead, excluded from
28
+ // every skill-value figure), run.pricing_snapshot (registry prices frozen
29
+ // at run time so derived dollars stay reproducible), and the derived
30
+ // `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
31
+ // v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
32
+ const RECEIPT_SCHEMA_VERSION = '0.4';
26
33
 
27
34
  // Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
28
35
  // these. The projection is refused before any call if it exceeds the cap, on
@@ -73,10 +80,19 @@ const REPORT_004_BASE_MODEL = 'claude-opus-5'; // flagship tier
73
80
  const REPORT_004_FRONTIER_MODEL = 'claude-fable-5'; // frontier tier (full id β€” no alias)
74
81
  const REPORT_004_JUDGE_MODEL = 'claude-haiku-4-5';
75
82
 
83
+ // Report #005 is a VALUE report β€” the fourth report type. It asks what a skill
84
+ // COSTS to run alongside whether it helps, over three substrates, and shows the
85
+ // three axes (accuracy lift / Ξ”cost / Ξ”latency) side by side and never combined.
86
+ // Same suites, same fixed judge; the substrate list spans two providers so the
87
+ // economics are read across surfaces, not within one vendor's pricing.
88
+ const REPORT_005_MODELS = ['claude-sonnet-5', 'claude-fable-5', 'gpt-5.6-sol'];
89
+ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
90
+
76
91
  module.exports = {
77
92
  PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
78
93
  EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
79
94
  REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
80
95
  REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
81
96
  REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
97
+ REPORT_005_MODELS, REPORT_005_JUDGE_MODEL,
82
98
  };
package/lib/diff.js CHANGED
@@ -3,6 +3,7 @@
3
3
 
4
4
  const { bandVerdict, round } = require('./stats');
5
5
  const { EFFECT_FLOOR } = require('../config');
6
+ const { revisionHeadline } = require('./revision');
6
7
 
7
8
  // Practical-significance gate applied ON TOP of band separation. bandVerdict()
8
9
  // stays a pure geometry test (kept that way so its unit checks are unambiguous);
@@ -68,7 +69,22 @@ function headlineVerdict(perCase) {
68
69
  return 'WITHIN NOISE β€” no case moved beyond its confidence band; the skill holds up.';
69
70
  }
70
71
 
71
- function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
72
+ // A revision pair is a pair in which the SKILL TEXT is the only thing that
73
+ // moved. `diff` was built for release drift, where the model varies and the text
74
+ // is fixed; this inverts it, so the fields that release drift merely warns about
75
+ // become the preconditions of the comparison. Returns the offending field name,
76
+ // or null when the pair is a valid revision pair.
77
+ function revisionPairProblem(a, b) {
78
+ if ((a.run.model_id || '') !== (b.run.model_id || '')) return 'run.model_id';
79
+ if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) return 'run.provider';
80
+ if ((a.run.surface || '') !== (b.run.surface || '')) return 'run.surface';
81
+ if ((a.suite.suite_hash || '') !== (b.suite.suite_hash || '')) return 'suite.suite_hash';
82
+ if ((a.skill.content_hash || '') === (b.skill.content_hash || '')) return 'skill.content_hash';
83
+ return null;
84
+ }
85
+
86
+ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
87
+ const revision = mode === 'revision';
72
88
  const aB = withSkillBands(a);
73
89
  const bB = withSkillBands(b);
74
90
  const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
@@ -99,6 +115,32 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
99
115
  const regressions = perCase.filter((r) => r.verdict === 'regression');
100
116
 
101
117
  const L = [];
118
+ if (revision) {
119
+ // The axis leads. In a release-drift report the model is what moved and it
120
+ // belongs at the top; here the model is the control and the skill's own text
121
+ // is the finding, so content_hash is the first row a reader meets.
122
+ L.push(`# Revision drift report`);
123
+ L.push('');
124
+ L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
125
+ L.push('');
126
+ L.push(`**Report type:** revision drift β€” the skill's own text is the variable under test; the substrate is held fixed.`);
127
+ L.push('');
128
+ L.push(`Left column is the **pinned** revision; right column is the **current** upstream revision.`);
129
+ L.push('');
130
+ L.push(`| | ${labelA} | ${labelB} |`);
131
+ L.push(`|---|---|---|`);
132
+ L.push(`| skill content_hash | \`${short(a.skill.content_hash)}\` | \`${short(b.skill.content_hash)}\` |`);
133
+ L.push(`| model (held) | \`${a.run.model_id}\` | \`${b.run.model_id}\` |`);
134
+ L.push(`| provider (held) | ${a.run.provider || 'anthropic'} | ${b.run.provider || 'anthropic'} |`);
135
+ L.push(`| surface (held) | ${a.run.surface} | ${b.run.surface} |`);
136
+ L.push(`| suite_hash (held) | \`${short(a.suite.suite_hash)}\` | \`${short(b.suite.suite_hash)}\` |`);
137
+ L.push(`| run date (UTC) | ${a.run.date_utc} | ${b.run.date_utc} |`);
138
+ L.push(`| judge samples/case | ${(a.run.judge || {}).samples || 1} | ${(b.run.judge || {}).samples || 1} |`);
139
+ L.push(`| with_skill (mean Β± band) | ${bandStr(aAgg)} | ${bandStr(bAgg)} |`);
140
+ L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
141
+ L.push(`| skill lift (Ξ”) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
142
+ L.push('');
143
+ } else {
102
144
  L.push(`# Drift report`);
103
145
  L.push('');
104
146
  L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
@@ -115,10 +157,19 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
115
157
  L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
116
158
  L.push(`| skill lift (Ξ”) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
117
159
  L.push('');
160
+ }
118
161
 
119
162
  const warnings = [];
120
- if (a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β€” the skill itself changed between receipts, so drift mixes skill edits with model drift.');
121
- if (a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β€” the eval suite changed; per-case comparison may be misleading.');
163
+ // In revision mode the changed skill text is the INDEPENDENT VARIABLE, not
164
+ // contamination, and the held-constant substrate is what makes the comparison
165
+ // valid. The release-drift caveat below says the opposite of both, so it is
166
+ // replaced rather than suppressed: a reader is told what is held, and why a
167
+ // separated band is attributable to the revision.
168
+ if (revision) {
169
+ warnings.push(`model \`${a.run.model_id}\`, provider ${a.run.provider || 'anthropic'}, surface ${a.run.surface} and suite_hash \`${short(a.suite.suite_hash)}\` are held constant across both receipts β€” the skill text is the only variable under test, so a band-separated move above the ${EFFECT_FLOOR} floor is attributable to the revision.`);
170
+ }
171
+ if (!revision && a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β€” the skill itself changed between receipts, so drift mixes skill edits with model drift.');
172
+ if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β€” the eval suite changed; per-case comparison may be misleading.');
122
173
  if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β€” comparison is not meaningful.`);
123
174
  if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β€” non-overlap can only be trusted when both sides are sampled.');
124
175
  if (!measured) warnings.push(`verdicts NOT computed β€” ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
@@ -126,8 +177,8 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
126
177
  // providers is a skill-DURABILITY comparison across substrates, not model drift
127
178
  // over time; across surfaces, sampling control differs. Both are flagged so a
128
179
  // reader never mistakes one for the other (see docs/neutrality.html).
129
- if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β€” this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
130
- if (a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β€” sampling control differs between surfaces; compare with care.`);
180
+ if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β€” this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
181
+ if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β€” sampling control differs between surfaces; compare with care.`);
131
182
  if (warnings.length) {
132
183
  L.push('> **⚠ Caveats**');
133
184
  for (const w of warnings) L.push(`> - ${w}`);
@@ -142,7 +193,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
142
193
  L.push(`**NOT MEASURED β€” ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
143
194
  L.push('');
144
195
  } else {
145
- L.push(`**${headlineVerdict(perCase)}**`);
196
+ L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
146
197
  L.push('');
147
198
  L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} β†’ ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin} within noise${nFloor ? ` (${nFloor} of them band-separated but below the ${EFFECT_FLOOR} effect floor)` : ''}.`);
148
199
  L.push('');
@@ -170,4 +221,4 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
170
221
  return { markdown: L.join('\n'), perCase, headlineDelta, regressions };
171
222
  }
172
223
 
173
- module.exports = { buildDriftReport, withSkillBands };
224
+ module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
package/lib/hygiene.js ADDED
@@ -0,0 +1,113 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ //
3
+ // The single definition of what counts as a leak.
4
+ //
5
+ // Two callers scan for the same classes of secret: `tests/gate.js`, which walks
6
+ // the working tree at gate time, and `scripts/merge-check.js`, which walks the
7
+ // candidate tree read out of the commit before permitting a merge. They must
8
+ // agree. Two copies of a deny-list drift, and the copy that quietly stops
9
+ // matching still reads as protection β€” the same argument the acknowledgment
10
+ // deny-list bite test already makes, applied to the scanner itself.
11
+ //
12
+ // Backlog B-8 records four approval records that reached `dev` carrying a host
13
+ // path. The recorded cause β€” "the scan walks tracked files" β€” is false, and this
14
+ // module's existence does not depend on it: `listFiles()` in the repo gate has
15
+ // always enumerated untracked files too, and a planted path in an untracked
16
+ // evidence record does take the repo gate red. The real window is VANTAGE. An
17
+ // approval session measures in a disposable copy made before it writes its
18
+ // record into the shared checkout, so the record is simply not in the tree that
19
+ // was scanned. Scanning the working tree harder cannot close that. Scanning the
20
+ // commit that is about to merge does, and that is what merge-check now uses this
21
+ // for.
22
+ //
23
+ // Patterns are written so a regex literal cannot match its own text: a
24
+ // metacharacter or a character class follows each fixed prefix.
25
+
26
+ const dec = (b64) => JSON.parse(Buffer.from(b64, 'base64').toString('utf8'));
27
+
28
+ // Build-box host + private IP, base64 so the plaintext never lands in a file.
29
+ const HOSTIP = dec('WyJpcC0xNzItMzEtNDAtMTU5LmFwLXNvdXRoZWFzdC0xLmNvbXB1dGUuaW50ZXJuYWwiLCIxNzIuMzEuNDAuMTU5Il0=');
30
+
31
+ const EMAIL_ALLOW = new Set(['example.com', 'example.org', 'driftproofhq.com']);
32
+
33
+ const PATTERNS = [
34
+ { name: 'home-path', re: /\/home\/[a-z0-9_-]+\/|\/Users\/[A-Za-z0-9_-]+\// },
35
+ { name: 'ec2-internal-host', re: /ip-\d+-\d+-\d+-\d+\.[a-z0-9.-]*compute\.(internal|amazonaws\.com)/i },
36
+ { name: 'private-ip', re: /\b(10\.\d{1,3}\.\d{1,3}\.\d{1,3}|172\.(1[6-9]|2\d|3[01])\.\d{1,3}\.\d{1,3}|192\.168\.\d{1,3}\.\d{1,3})\b/ },
37
+ { name: 'anthropic-key', re: /sk-ant-[A-Za-z0-9_-]{8,}/ },
38
+ { name: 'github-token', re: /gh[pousr]_[A-Za-z0-9]{20,}/ },
39
+ { name: 'aws-key', re: /\bAKIA[0-9A-Z]{16}\b/ },
40
+ { name: 'private-key-block', re: /-----BEGIN [A-Z ]*PRIVATE KEY-----/ },
41
+ { name: 'env-secret-assignment', re: /\b(ANTHROPIC_API_KEY|AWS_SECRET_ACCESS_KEY|OPENAI_API_KEY)\s*=\s*\S+/ },
42
+ // Codex subscription auth material (~/.codex/auth.json contents) must NEVER
43
+ // land in a committed file: the id_token/access_token are JWTs, and an OpenAI
44
+ // secret key is sk-proj-/sk-svcacct-/sk-admin-. We ban the CONTENTS (tokens),
45
+ // not the documented path string (`~/.codex/auth.json` is referenced in help
46
+ // text and docs by design).
47
+ { name: 'jwt-token', re: /\beyJ[A-Za-z0-9_=-]{10,}\.eyJ[A-Za-z0-9_=-]{10,}\.[A-Za-z0-9_=-]{6,}/ },
48
+ { name: 'openai-secret-key', re: /\bsk-(proj|svcacct|admin)-[A-Za-z0-9_-]{20,}/ },
49
+ ];
50
+
51
+ // A path is a hit on its own name, with no content read: a committed .env file
52
+ // is a leak whatever it happens to contain.
53
+ function scanPath(rel) {
54
+ return /(^|\/)\.env(\.|$)/.test(rel) ? [{ file: rel, kind: 'env-file' }] : [];
55
+ }
56
+
57
+ function scanContent(rel, content) {
58
+ const hits = [];
59
+ for (const p of PATTERNS) {
60
+ const m = content.match(p.re);
61
+ if (m) hits.push({ file: rel, kind: p.name, sample: m[0].slice(0, 24) });
62
+ }
63
+ for (const hv of HOSTIP) if (content.includes(hv)) hits.push({ file: rel, kind: 'build-host-or-ip' });
64
+ const emails = content.match(/\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g) || [];
65
+ for (const e of emails) {
66
+ const dom = e.split('@')[1].toLowerCase();
67
+ if (!EMAIL_ALLOW.has(dom) && !dom.endsWith('.example')) hits.push({ file: rel, kind: 'email', sample: e });
68
+ }
69
+ return hits;
70
+ }
71
+
72
+ // scanFiles(files, read, readLink) β€” `read(rel)` returns the file's text, or
73
+ // throws/returns null for anything unreadable (binary, deleted, a symlink to
74
+ // nowhere), which is skipped exactly as the gate has always skipped it.
75
+ //
76
+ // `readLink(rel)` is optional and returns a symbolic link's TARGET PATH as a
77
+ // string, or null for an entry that is not a link. Where it is supplied, that
78
+ // target is scanned as the entry's content through `scanContent`, so every
79
+ // pattern above applies to it with no second deny-list to drift (spec 011,
80
+ // AC-2). A link's target is a path, and a path is exactly the class of leak the
81
+ // `home-path` pattern exists to catch: P-1 shipped `node_modules ->
82
+ // /home/<user>/... ` into a public tree past a scan that read the link as an
83
+ // unreadable file and skipped it.
84
+ //
85
+ // The argument is optional so a caller supplying nothing behaves exactly as
86
+ // before. That compatibility is deliberate, and it is also the standing risk: a
87
+ // caller added later is blind by default. The repo gate asserts that every call
88
+ // site supplies a reader, with `scripts/merge-check.js` the one exception β€” its
89
+ // reader is `git show <ref>:<path>`, and git stores a link's blob as its target
90
+ // string, so the target already arrives as content there.
91
+ function scanFiles(files, read, readLink) {
92
+ const hits = [];
93
+ for (const rel of files) {
94
+ hits.push(...scanPath(rel));
95
+ // The link branch is a single condition on purpose: it is the mutation seam
96
+ // the spec-011 gate neutralises to prove the scanner goes blind without it.
97
+ if (typeof readLink === 'function') {
98
+ let target;
99
+ try { target = readLink(rel); } catch (_e) { target = null; }
100
+ if (typeof target === 'string') {
101
+ hits.push(...scanContent(rel, target));
102
+ continue;
103
+ }
104
+ }
105
+ let c;
106
+ try { c = read(rel); } catch (_e) { continue; }
107
+ if (typeof c !== 'string') continue;
108
+ hits.push(...scanContent(rel, c));
109
+ }
110
+ return hits;
111
+ }
112
+
113
+ module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, scanPath, scanContent, scanFiles };
package/lib/judge.js CHANGED
@@ -5,6 +5,7 @@ const { complete, surfaceForModel } = require('./provider');
5
5
  const { extractJsonObject } = require('./json');
6
6
  const { sha256 } = require('./canonical');
7
7
  const { mean, stddev } = require('./stats');
8
+ const { sumUsage } = require('./usage');
8
9
 
9
10
  // Rubric-based LLM judge.
10
11
  //
@@ -82,16 +83,16 @@ function judgeSettings(samples, judgeModel) {
82
83
  // transcript auditability, and optionally retained under --keep-transcripts).
83
84
  async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature }) {
84
85
  const prompt = buildJudgePrompt({ task, response, rubric });
85
- const { text, attempts } = await complete({ system: JUDGE_SYSTEM, prompt, model, maxTokens: 400, timeoutMs, temperature });
86
+ const { text, attempts, usage } = await complete({ system: JUDGE_SYSTEM, prompt, model, maxTokens: 400, timeoutMs, temperature });
86
87
  let parsed;
87
88
  try {
88
89
  parsed = extractJsonObject(text);
89
90
  } catch (_e) {
90
91
  // Unsalvageable judge output β†’ conservative 0 (a judge that can't be parsed
91
92
  // must never silently "pass"), tagged so the caller can see it happened.
92
- return { score: 0, reason: 'judge output unparseable', unparsed: true, raw: String(text || ''), attempts: attempts || 1 };
93
+ return { score: 0, reason: 'judge output unparseable', unparsed: true, raw: String(text || ''), attempts: attempts || 1, usage };
93
94
  }
94
- return { score: clamp01(parsed.score), reason: String(parsed.reason || '').slice(0, 300), raw: String(text || ''), attempts: attempts || 1 };
95
+ return { score: clamp01(parsed.score), reason: String(parsed.reason || '').slice(0, 300), raw: String(text || ''), attempts: attempts || 1, usage };
95
96
  }
96
97
 
97
98
  // Grade a response N times and return the sampled distribution:
@@ -103,6 +104,7 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
103
104
  const scores = [];
104
105
  const reasons = [];
105
106
  const rawTexts = [];
107
+ const usages = [];
106
108
  let attemptsTotal = 0;
107
109
  for (let i = 0; i < samples; i++) {
108
110
  let r;
@@ -116,6 +118,7 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
116
118
  throw e;
117
119
  }
118
120
  attemptsTotal += r.attempts || 1;
121
+ usages.push(r.usage || null);
119
122
  scores.push(r.score);
120
123
  reasons.push(r.reason);
121
124
  rawTexts.push(r.raw || '');
@@ -134,6 +137,11 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
134
137
  model_id: model,
135
138
  rubric_hash: rubricHash(rubric),
136
139
  attempts: attemptsTotal,
140
+ // v0.4: the measurement overhead of grading this one case β€” the SUM over all
141
+ // N judge calls. Recorded in the receipt as the case's `judge_usage` and
142
+ // EXCLUDED from every skill-value figure (lib/value.js): it is a cost we
143
+ // impose to measure, not a cost of running the skill.
144
+ usage: sumUsage(usages),
137
145
  };
138
146
  }
139
147