driftproof 0.5.0 β†’ 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -9,15 +9,17 @@
9
9
  Driftproof is an open **receipt spec** plus a **runner** that measures whether an
10
10
  agent skill actually helps β€” by running the skill's eval suite **with** and
11
11
  **without** the skill on a named model version, judging each case several times to
12
- get a confidence band, and emitting a signed, dated **receipt**. Diff two receipts
12
+ get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
13
13
  across model releases and you get a **drift report**.
14
14
 
15
15
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
16
16
  format; it does not invent its own.
17
17
 
18
- πŸ“Š **Four published reports** (each re-derived from committed receipts, nothing
19
- hand-entered), spanning three report types that share one band-based, floor-gated
20
- verdict rule and differ only in what moves underneath the skill:
18
+ πŸ“Š **Six published reports** (each re-derived from committed receipts, nothing
19
+ hand-entered), spanning five published report types, the fifth being revision
20
+ drift. All six share one band-based, floor-gated verdict rule and differ in what
21
+ moves underneath the skill β€” or, in the value report, in which axes are
22
+ measured:
21
23
 
22
24
  - **[Report #001](https://driftproofhq.com/reports/001/)** β€” *release drift*: ten
23
25
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
@@ -31,10 +33,29 @@ verdict rule and differ only in what moves underneath the skill:
31
33
  `claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
32
34
  3 durable, 5 tier-dependent, 0 regressions, 2 no effect β€” encoded expertise
33
35
  survives the frontier tier.
36
+ - **[Report #005](https://driftproofhq.com/reports/005/)** β€” *value*: what a skill
37
+ *costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
38
+ three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) β€”
39
+ 14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
40
+ saving instead, having improved quality while reducing cost.
41
+ *Three of those cells carry an amendment (v1.1, applied when Report #006
42
+ published): their lifts rest on single-draw baselines since shown
43
+ unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
44
+ substrate-disagreement findings below are unaffected.*
45
+ - **[Report #006](https://driftproofhq.com/reports/006/)** β€” *revision drift*: the pinned skill revision against the one
46
+ upstream ships today, on a held substrate. **The reuse premise was tested and
47
+ refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
48
+ control. A 120-call probe found generation-level sampling noise 3.2Γ— and 7.5Γ—
49
+ larger than the judge-level noise this instrument actually samples β€” enough to
50
+ account for every gap the controls saw without any other cause being
51
+ established β€” and the receipt spec gains generation sampling as a result. **No
52
+ cause is asserted**; the control proves non-reproduction and cannot say why.
53
+ *The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
34
54
 
35
55
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
36
- skills](https://driftproofhq.com/writing/three-releases/)**, reads the first three
37
- reports together.
56
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all six reports
57
+ together: what moves underneath a skill, and what the skill costs to run. Revised
58
+ 2026-08-29; every figure in it is gate-checked against the report page it cites.
38
59
 
39
60
  ## Why
40
61
 
@@ -55,6 +76,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
55
76
  the judge and report confidence bands**, and to **only claim a regression when the
56
77
  bands don't overlap**. A tool that cries wolf is worse than no tool.
57
78
 
79
+ **A verdict without a price is half an answer.** The same receipts price the
80
+ marginal cost of a skill firing, and Report #005 found the dominant cost driver is
81
+ not the skill's own text but the input it causes the model to pull in: across those
82
+ 30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
83
+ tracks it at only `r = +0.33`, and one 738-token skill drew 34Γ— its own size in
84
+ extra input. Identical token deltas also price very differently across substrates β€”
85
+ the same skill at near-identical deltas costs 3.3Γ— more on `claude-fable-5` than on
86
+ `claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
87
+
58
88
  ## Quickstart β€” receipt for your own skill in ~10 minutes
59
89
 
60
90
  You need Node β‰₯ 22 and an `ANTHROPIC_API_KEY`.
@@ -86,6 +116,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
86
116
  minimum real move required before a change counts as more than noise β€” band
87
117
  separation *plus* a floor-sized delta, never either alone.
88
118
 
119
+ **What the band does not cover.** A verdict rests on **one generation draw per
120
+ arm**: the band is the spread of the *judge* re-scoring that single response, not
121
+ the spread of the model writing a different one. Report #006 measured the second
122
+ directly and found it larger β€” draw-to-draw spread up to **sd 0.186** on the 0–1
123
+ scale, against judge-level noise several times smaller. So treat a surprising
124
+ single-run verdict as **provisional and worth re-running** before you act on it.
125
+ Generation sampling lands in the next receipt spec; until it does, this is a
126
+ limit of the instrument, stated rather than implied.
127
+
89
128
  ### Install
90
129
 
91
130
  ```bash
@@ -171,7 +210,7 @@ A receipt is the unit of evidence β€” one JSON document conforming to
171
210
  "model_release_date": "2025-10-01",
172
211
  "provider": "anthropic",
173
212
  "surface": "claude-cli",
174
- "runner_version": "0.5.0",
213
+ "runner_version": "0.6.0",
175
214
  "date_utc": "2026-07-27T…Z",
176
215
  "registry": "registered",
177
216
  "transcripts": "hashes-only",
@@ -194,6 +233,12 @@ A receipt is the unit of evidence β€” one JSON document conforming to
194
233
  },
195
234
  "comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
196
235
  "delta": 0.39, "delta_uncertainty": 0.036 },
236
+ // v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
237
+ // freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
238
+ // holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
239
+ // mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
240
+ // skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
241
+ // output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
197
242
  "verification_level": "TESTED",
198
243
  "receipt_hash": "…sha256 of the canonical receipt with this field removed…"
199
244
  }
@@ -209,6 +254,12 @@ Key ideas:
209
254
  - **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
210
255
  - **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
211
256
  `TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
257
+ - **`economics`** is *derived*, never a second measurement: the token delta is the
258
+ durable fact, and the dollars are exactly those tokens at the rates frozen into
259
+ `run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
260
+ Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
261
+ figure (`judge_excluded` is `const true`) β€” measuring the skill is our cost, not
262
+ the skill's.
212
263
  - **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
213
264
  signature β€” see the spec's open questions).
214
265
 
@@ -233,7 +284,7 @@ jobs:
233
284
  runs-on: ubuntu-latest
234
285
  steps:
235
286
  - uses: actions/checkout@v4
236
- - uses: driftproofhq/driftproof@v0.5.0
287
+ - uses: driftproofhq/driftproof@v0.6.0
237
288
  with:
238
289
  skill-dir: skills/my-skill
239
290
  models: claude-haiku-4-5
@@ -272,11 +323,14 @@ site, so it reflects a real dated run, not a hand-set color.
272
323
 
273
324
  ## Reports
274
325
 
275
- Four reports are published, spanning three report types (see the roll at the top
276
- of this README, and [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome
277
- every report inherits). Each report and every verdict in it are **re-derived from
278
- the receipts** committed under [`receipts/`](receipts/)
279
- (`receipts/report-001/` … `receipts/report-004/`) β€” nothing is hand-entered.
326
+ Six reports are published, spanning five report types. A report page lives at a
327
+ draft path β€” `docs/reports/NNN-draft/` β€” until the publish sequence renames it, and
328
+ `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
329
+ (see the roll at the top of this README, and
330
+ [REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
331
+ Each report and every verdict in it are **re-derived from the receipts** committed
332
+ under [`receipts/`](receipts/)
333
+ (`receipts/report-001/` … `receipts/report-006/`) β€” nothing is hand-entered.
280
334
 
281
335
  Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
282
336
  fetched at run time from a pinned commit and verified by sha256 against
@@ -290,9 +344,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models Γ— with/basel
290
344
  node scripts/build-report-001.js # re-derive the report from the receipts
291
345
  ```
292
346
 
293
- Reports #002–#004 have their own runners
294
- (`scripts/prepare-report-00N.js`) following the same
295
- fetch β†’ run β†’ re-derive shape.
347
+ Reports #002–#005 have their own runners
348
+ (`scripts/prepare-report-00N.js` β€” Report #005's is
349
+ [`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
350
+ same fetch β†’ run β†’ re-derive shape.
296
351
 
297
352
  **Model-release triggers are live**: `scripts/release-watch.js` (keyless β€” it
298
353
  reads the public models registry) notices a new model release, re-runs the
package/bin/driftproof CHANGED
@@ -8,7 +8,7 @@ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = req
8
8
  const { loadSkill } = require('../lib/skill');
9
9
  const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
10
10
  const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
11
- const { buildDriftReport } = require('../lib/diff');
11
+ const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
12
12
  const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
13
13
  const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
14
14
  const { registryStatus } = require('../lib/models');
@@ -76,7 +76,7 @@ USAGE
76
76
  ${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
77
77
  [--judge-model M] [--concurrency N] [--max-usd N]
78
78
  [--keep-transcripts] [--out DIR]
79
- ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
79
+ ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
80
80
  ${PROJECT_NAME} validate <receipt.json>
81
81
  ${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
82
82
  ${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
@@ -250,9 +250,32 @@ function cmdDiff(positional, flags) {
250
250
  if (!verifyReceiptHash(r)) console.error(` ⚠ ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
251
251
  }
252
252
 
253
- const labelA = dateStamp(a.run.date_utc);
254
- const labelB = dateStamp(b.run.date_utc);
255
- const { markdown } = buildDriftReport(a, b, { labelA, labelB });
253
+ // --mode revision inverts the axis: the skill text is the variable under test
254
+ // and the substrate is the control. The fields release drift merely warns about
255
+ // are preconditions here, so a pair that is not a revision pair is REFUSED
256
+ // (exit 6) rather than rendered with a caveat nobody reads. A differing model
257
+ // would be release drift wearing a revision label β€” the one confound this mode
258
+ // exists to exclude β€” and an EQUAL content_hash has no revision to measure.
259
+ const mode = flags.mode || 'release';
260
+ if (mode !== 'release' && mode !== 'revision') {
261
+ console.error(`unknown --mode "${mode}" β€” supported: release, revision`);
262
+ process.exit(2);
263
+ }
264
+ if (mode === 'revision') {
265
+ const problem = revisionPairProblem(a, b);
266
+ if (problem) {
267
+ const why = problem === 'skill.content_hash'
268
+ ? 'the two receipts carry the SAME skill.content_hash β€” there is no revision between them to measure'
269
+ : `${problem} differs between the two receipts β€” revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
270
+ console.error(` βœ— REFUSED (--mode revision): ${why}.`);
271
+ console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
272
+ process.exit(6);
273
+ }
274
+ }
275
+
276
+ const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
277
+ const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
278
+ const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
256
279
 
257
280
  if (flags.out) {
258
281
  fs.writeFileSync(path.resolve(flags.out), markdown);
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.5.0';
12
+ const RUNNER_VERSION = '0.6.0';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
package/lib/diff.js CHANGED
@@ -3,6 +3,7 @@
3
3
 
4
4
  const { bandVerdict, round } = require('./stats');
5
5
  const { EFFECT_FLOOR } = require('../config');
6
+ const { revisionHeadline } = require('./revision');
6
7
 
7
8
  // Practical-significance gate applied ON TOP of band separation. bandVerdict()
8
9
  // stays a pure geometry test (kept that way so its unit checks are unambiguous);
@@ -68,7 +69,22 @@ function headlineVerdict(perCase) {
68
69
  return 'WITHIN NOISE β€” no case moved beyond its confidence band; the skill holds up.';
69
70
  }
70
71
 
71
- function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
72
+ // A revision pair is a pair in which the SKILL TEXT is the only thing that
73
+ // moved. `diff` was built for release drift, where the model varies and the text
74
+ // is fixed; this inverts it, so the fields that release drift merely warns about
75
+ // become the preconditions of the comparison. Returns the offending field name,
76
+ // or null when the pair is a valid revision pair.
77
+ function revisionPairProblem(a, b) {
78
+ if ((a.run.model_id || '') !== (b.run.model_id || '')) return 'run.model_id';
79
+ if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) return 'run.provider';
80
+ if ((a.run.surface || '') !== (b.run.surface || '')) return 'run.surface';
81
+ if ((a.suite.suite_hash || '') !== (b.suite.suite_hash || '')) return 'suite.suite_hash';
82
+ if ((a.skill.content_hash || '') === (b.skill.content_hash || '')) return 'skill.content_hash';
83
+ return null;
84
+ }
85
+
86
+ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
87
+ const revision = mode === 'revision';
72
88
  const aB = withSkillBands(a);
73
89
  const bB = withSkillBands(b);
74
90
  const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
@@ -99,6 +115,32 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
99
115
  const regressions = perCase.filter((r) => r.verdict === 'regression');
100
116
 
101
117
  const L = [];
118
+ if (revision) {
119
+ // The axis leads. In a release-drift report the model is what moved and it
120
+ // belongs at the top; here the model is the control and the skill's own text
121
+ // is the finding, so content_hash is the first row a reader meets.
122
+ L.push(`# Revision drift report`);
123
+ L.push('');
124
+ L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
125
+ L.push('');
126
+ L.push(`**Report type:** revision drift β€” the skill's own text is the variable under test; the substrate is held fixed.`);
127
+ L.push('');
128
+ L.push(`Left column is the **pinned** revision; right column is the **current** upstream revision.`);
129
+ L.push('');
130
+ L.push(`| | ${labelA} | ${labelB} |`);
131
+ L.push(`|---|---|---|`);
132
+ L.push(`| skill content_hash | \`${short(a.skill.content_hash)}\` | \`${short(b.skill.content_hash)}\` |`);
133
+ L.push(`| model (held) | \`${a.run.model_id}\` | \`${b.run.model_id}\` |`);
134
+ L.push(`| provider (held) | ${a.run.provider || 'anthropic'} | ${b.run.provider || 'anthropic'} |`);
135
+ L.push(`| surface (held) | ${a.run.surface} | ${b.run.surface} |`);
136
+ L.push(`| suite_hash (held) | \`${short(a.suite.suite_hash)}\` | \`${short(b.suite.suite_hash)}\` |`);
137
+ L.push(`| run date (UTC) | ${a.run.date_utc} | ${b.run.date_utc} |`);
138
+ L.push(`| judge samples/case | ${(a.run.judge || {}).samples || 1} | ${(b.run.judge || {}).samples || 1} |`);
139
+ L.push(`| with_skill (mean Β± band) | ${bandStr(aAgg)} | ${bandStr(bAgg)} |`);
140
+ L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
141
+ L.push(`| skill lift (Ξ”) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
142
+ L.push('');
143
+ } else {
102
144
  L.push(`# Drift report`);
103
145
  L.push('');
104
146
  L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
@@ -115,10 +157,19 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
115
157
  L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
116
158
  L.push(`| skill lift (Ξ”) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
117
159
  L.push('');
160
+ }
118
161
 
119
162
  const warnings = [];
120
- if (a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β€” the skill itself changed between receipts, so drift mixes skill edits with model drift.');
121
- if (a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β€” the eval suite changed; per-case comparison may be misleading.');
163
+ // In revision mode the changed skill text is the INDEPENDENT VARIABLE, not
164
+ // contamination, and the held-constant substrate is what makes the comparison
165
+ // valid. The release-drift caveat below says the opposite of both, so it is
166
+ // replaced rather than suppressed: a reader is told what is held, and why a
167
+ // separated band is attributable to the revision.
168
+ if (revision) {
169
+ warnings.push(`model \`${a.run.model_id}\`, provider ${a.run.provider || 'anthropic'}, surface ${a.run.surface} and suite_hash \`${short(a.suite.suite_hash)}\` are held constant across both receipts β€” the skill text is the only variable under test, so a band-separated move above the ${EFFECT_FLOOR} floor is attributable to the revision.`);
170
+ }
171
+ if (!revision && a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β€” the skill itself changed between receipts, so drift mixes skill edits with model drift.');
172
+ if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β€” the eval suite changed; per-case comparison may be misleading.');
122
173
  if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β€” comparison is not meaningful.`);
123
174
  if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β€” non-overlap can only be trusted when both sides are sampled.');
124
175
  if (!measured) warnings.push(`verdicts NOT computed β€” ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
@@ -126,8 +177,8 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
126
177
  // providers is a skill-DURABILITY comparison across substrates, not model drift
127
178
  // over time; across surfaces, sampling control differs. Both are flagged so a
128
179
  // reader never mistakes one for the other (see docs/neutrality.html).
129
- if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β€” this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
130
- if (a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β€” sampling control differs between surfaces; compare with care.`);
180
+ if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β€” this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
181
+ if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β€” sampling control differs between surfaces; compare with care.`);
131
182
  if (warnings.length) {
132
183
  L.push('> **⚠ Caveats**');
133
184
  for (const w of warnings) L.push(`> - ${w}`);
@@ -142,7 +193,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
142
193
  L.push(`**NOT MEASURED β€” ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
143
194
  L.push('');
144
195
  } else {
145
- L.push(`**${headlineVerdict(perCase)}**`);
196
+ L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
146
197
  L.push('');
147
198
  L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} β†’ ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin} within noise${nFloor ? ` (${nFloor} of them band-separated but below the ${EFFECT_FLOOR} effect floor)` : ''}.`);
148
199
  L.push('');
@@ -170,4 +221,4 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
170
221
  return { markdown: L.join('\n'), perCase, headlineDelta, regressions };
171
222
  }
172
223
 
173
- module.exports = { buildDriftReport, withSkillBands };
224
+ module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
package/lib/hygiene.js ADDED
@@ -0,0 +1,113 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ //
3
+ // The single definition of what counts as a leak.
4
+ //
5
+ // Two callers scan for the same classes of secret: `tests/gate.js`, which walks
6
+ // the working tree at gate time, and `scripts/merge-check.js`, which walks the
7
+ // candidate tree read out of the commit before permitting a merge. They must
8
+ // agree. Two copies of a deny-list drift, and the copy that quietly stops
9
+ // matching still reads as protection β€” the same argument the acknowledgment
10
+ // deny-list bite test already makes, applied to the scanner itself.
11
+ //
12
+ // Backlog B-8 records four approval records that reached `dev` carrying a host
13
+ // path. The recorded cause β€” "the scan walks tracked files" β€” is false, and this
14
+ // module's existence does not depend on it: `listFiles()` in the repo gate has
15
+ // always enumerated untracked files too, and a planted path in an untracked
16
+ // evidence record does take the repo gate red. The real window is VANTAGE. An
17
+ // approval session measures in a disposable copy made before it writes its
18
+ // record into the shared checkout, so the record is simply not in the tree that
19
+ // was scanned. Scanning the working tree harder cannot close that. Scanning the
20
+ // commit that is about to merge does, and that is what merge-check now uses this
21
+ // for.
22
+ //
23
+ // Patterns are written so a regex literal cannot match its own text: a
24
+ // metacharacter or a character class follows each fixed prefix.
25
+
26
+ const dec = (b64) => JSON.parse(Buffer.from(b64, 'base64').toString('utf8'));
27
+
28
+ // Build-box host + private IP, base64 so the plaintext never lands in a file.
29
+ const HOSTIP = dec('WyJpcC0xNzItMzEtNDAtMTU5LmFwLXNvdXRoZWFzdC0xLmNvbXB1dGUuaW50ZXJuYWwiLCIxNzIuMzEuNDAuMTU5Il0=');
30
+
31
+ const EMAIL_ALLOW = new Set(['example.com', 'example.org', 'driftproofhq.com']);
32
+
33
+ const PATTERNS = [
34
+ { name: 'home-path', re: /\/home\/[a-z0-9_-]+\/|\/Users\/[A-Za-z0-9_-]+\// },
35
+ { name: 'ec2-internal-host', re: /ip-\d+-\d+-\d+-\d+\.[a-z0-9.-]*compute\.(internal|amazonaws\.com)/i },
36
+ { name: 'private-ip', re: /\b(10\.\d{1,3}\.\d{1,3}\.\d{1,3}|172\.(1[6-9]|2\d|3[01])\.\d{1,3}\.\d{1,3}|192\.168\.\d{1,3}\.\d{1,3})\b/ },
37
+ { name: 'anthropic-key', re: /sk-ant-[A-Za-z0-9_-]{8,}/ },
38
+ { name: 'github-token', re: /gh[pousr]_[A-Za-z0-9]{20,}/ },
39
+ { name: 'aws-key', re: /\bAKIA[0-9A-Z]{16}\b/ },
40
+ { name: 'private-key-block', re: /-----BEGIN [A-Z ]*PRIVATE KEY-----/ },
41
+ { name: 'env-secret-assignment', re: /\b(ANTHROPIC_API_KEY|AWS_SECRET_ACCESS_KEY|OPENAI_API_KEY)\s*=\s*\S+/ },
42
+ // Codex subscription auth material (~/.codex/auth.json contents) must NEVER
43
+ // land in a committed file: the id_token/access_token are JWTs, and an OpenAI
44
+ // secret key is sk-proj-/sk-svcacct-/sk-admin-. We ban the CONTENTS (tokens),
45
+ // not the documented path string (`~/.codex/auth.json` is referenced in help
46
+ // text and docs by design).
47
+ { name: 'jwt-token', re: /\beyJ[A-Za-z0-9_=-]{10,}\.eyJ[A-Za-z0-9_=-]{10,}\.[A-Za-z0-9_=-]{6,}/ },
48
+ { name: 'openai-secret-key', re: /\bsk-(proj|svcacct|admin)-[A-Za-z0-9_-]{20,}/ },
49
+ ];
50
+
51
+ // A path is a hit on its own name, with no content read: a committed .env file
52
+ // is a leak whatever it happens to contain.
53
+ function scanPath(rel) {
54
+ return /(^|\/)\.env(\.|$)/.test(rel) ? [{ file: rel, kind: 'env-file' }] : [];
55
+ }
56
+
57
+ function scanContent(rel, content) {
58
+ const hits = [];
59
+ for (const p of PATTERNS) {
60
+ const m = content.match(p.re);
61
+ if (m) hits.push({ file: rel, kind: p.name, sample: m[0].slice(0, 24) });
62
+ }
63
+ for (const hv of HOSTIP) if (content.includes(hv)) hits.push({ file: rel, kind: 'build-host-or-ip' });
64
+ const emails = content.match(/\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g) || [];
65
+ for (const e of emails) {
66
+ const dom = e.split('@')[1].toLowerCase();
67
+ if (!EMAIL_ALLOW.has(dom) && !dom.endsWith('.example')) hits.push({ file: rel, kind: 'email', sample: e });
68
+ }
69
+ return hits;
70
+ }
71
+
72
+ // scanFiles(files, read, readLink) β€” `read(rel)` returns the file's text, or
73
+ // throws/returns null for anything unreadable (binary, deleted, a symlink to
74
+ // nowhere), which is skipped exactly as the gate has always skipped it.
75
+ //
76
+ // `readLink(rel)` is optional and returns a symbolic link's TARGET PATH as a
77
+ // string, or null for an entry that is not a link. Where it is supplied, that
78
+ // target is scanned as the entry's content through `scanContent`, so every
79
+ // pattern above applies to it with no second deny-list to drift (spec 011,
80
+ // AC-2). A link's target is a path, and a path is exactly the class of leak the
81
+ // `home-path` pattern exists to catch: P-1 shipped `node_modules ->
82
+ // /home/<user>/... ` into a public tree past a scan that read the link as an
83
+ // unreadable file and skipped it.
84
+ //
85
+ // The argument is optional so a caller supplying nothing behaves exactly as
86
+ // before. That compatibility is deliberate, and it is also the standing risk: a
87
+ // caller added later is blind by default. The repo gate asserts that every call
88
+ // site supplies a reader, with `scripts/merge-check.js` the one exception β€” its
89
+ // reader is `git show <ref>:<path>`, and git stores a link's blob as its target
90
+ // string, so the target already arrives as content there.
91
+ function scanFiles(files, read, readLink) {
92
+ const hits = [];
93
+ for (const rel of files) {
94
+ hits.push(...scanPath(rel));
95
+ // The link branch is a single condition on purpose: it is the mutation seam
96
+ // the spec-011 gate neutralises to prove the scanner goes blind without it.
97
+ if (typeof readLink === 'function') {
98
+ let target;
99
+ try { target = readLink(rel); } catch (_e) { target = null; }
100
+ if (typeof target === 'string') {
101
+ hits.push(...scanContent(rel, target));
102
+ continue;
103
+ }
104
+ }
105
+ let c;
106
+ try { c = read(rel); } catch (_e) { continue; }
107
+ if (typeof c !== 'string') continue;
108
+ hits.push(...scanContent(rel, c));
109
+ }
110
+ return hits;
111
+ }
112
+
113
+ module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, scanPath, scanContent, scanFiles };
@@ -0,0 +1,162 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ 'use strict';
3
+
4
+ const { bandVerdict, round } = require('./stats');
5
+ const { EFFECT_FLOOR } = require('../config');
6
+
7
+ // Revision drift β€” the fifth report type (spec 009, Report #006).
8
+ //
9
+ // Every other report type holds the skill text fixed and moves something
10
+ // underneath it: the model release (#001, #003), the vendor surface (#002), the
11
+ // capability tier (#004), the axes and the price (#005). This one inverts the
12
+ // design. The substrate is held still β€” same model, same provider, same surface,
13
+ // same suite, same fixed judge, same sampling β€” and the SKILL'S OWN TEXT moves,
14
+ // from the revision a report pinned to the revision upstream ships today.
15
+ //
16
+ // This module holds the report type's LANGUAGE as code rather than as
17
+ // hand-written page copy. That is deliberate: the fairness rule below is the one
18
+ // a measurement project is most tempted to apply in one direction only, and a
19
+ // rule that lives in prose cannot be gated before the run that would tempt it.
20
+
21
+ // ── the cell headline ────────────────────────────────────────────────────────
22
+ // A summary of the per-case band-overlap verdicts, worded about the REVISION.
23
+ // The release-drift headline says "the skill is measurably weaker", which is a
24
+ // sentence about a skill under a moving model. Here the model is the control.
25
+ function revisionHeadline(perCase) {
26
+ const reg = perCase.filter((r) => r.verdict === 'regression').length;
27
+ const imp = perCase.filter((r) => r.verdict === 'improvement').length;
28
+ const s = (n) => (n === 1 ? '' : 's');
29
+ if (reg && imp) {
30
+ return `MIXED β€” the revision improved ${imp} case${s(imp)} and regressed ${reg} on non-overlapping bands.`;
31
+ }
32
+ if (reg) {
33
+ return `REVISION REGRESSED β€” ${reg} case${s(reg)} scored lower under the current upstream text (bands do not overlap).`;
34
+ }
35
+ if (imp) {
36
+ return `REVISION IMPROVED β€” ${imp} case${s(imp)} scored higher under the current upstream text (bands do not overlap); none regressed.`;
37
+ }
38
+ return 'WITHIN NOISE β€” the revision moved no case beyond its confidence band; the pinned text and the current text measure the same.';
39
+ }
40
+
41
+ // Classification word for a cell, from its headline. Kept separate so a caller
42
+ // can branch on the class without parsing prose.
43
+ function revisionClass(perCase) {
44
+ const h = revisionHeadline(perCase);
45
+ return h.split(' β€”')[0];
46
+ }
47
+
48
+ // ── the fairness sentence ────────────────────────────────────────────────────
49
+ // Spec 009 Β§ Fairness, clauses 1 and 4. Where a revision IMPROVED a skill, the
50
+ // published #005 figure understates the pack a reader can install today; where it
51
+ // REGRESSED one, #005 overstates it. Both sentences are generated by the same
52
+ // function, from the same template, so the disclosure cannot quietly become a
53
+ // one-directional courtesy β€” the symmetry is a property of the code, and the
54
+ // gate asserts it.
55
+ //
56
+ // A cell within noise gets NO sentence. #005's figure stands unamended, because
57
+ // nothing was measured that would amend it, and manufacturing a hedge for a null
58
+ // result is how a report launders noise into a finding.
59
+ function fairnessSentence({ slug, classification, report005Delta, measuredDelta }) {
60
+ const cls = String(classification || '');
61
+ if (cls !== 'REVISION IMPROVED' && cls !== 'REVISION REGRESSED') return null;
62
+ const improved = cls === 'REVISION IMPROVED';
63
+ const direction = improved ? 'understates' : 'overstates';
64
+ const d = (n) => (n == null ? 'n/a' : (n >= 0 ? '+' : '') + Number(n).toFixed(3));
65
+ return `Report #005 measured ${slug} at ${d(report005Delta)} on the text it had pinned. `
66
+ + `Report #006 measures the current upstream revision at ${d(measuredDelta)} on the same substrate and the same suite. `
67
+ + `#005's published figure therefore ${direction} the pack upstream ships today for this skill, `
68
+ + `and is amended by this report rather than corrected in place.`;
69
+ }
70
+
71
+ // ── per-cell scoping disclosures ─────────────────────────────────────────────
72
+ // Spec 009 AC-10. One cell in Report #006 measures something other than what its
73
+ // upstream author changed it to do, and the reader looking at that row is the
74
+ // reader who needs to be told.
75
+ const SCOPING_NOTES = {
76
+ 'git-workflow-and-versioning':
77
+ 'This revision changes the frontmatter `description:` line. In a skill runtime a description is a '
78
+ + 'routing trigger: it decides whether the skill loads, and never reaches the model as guidance. '
79
+ + 'Driftproof makes no routing decision β€” it always injects the skill, and passes the whole file, '
80
+ + 'frontmatter included, as the system prompt. This cell therefore measures the revision as added '
81
+ + 'context and cannot measure it as a trigger.',
82
+ };
83
+ function scopingNote(slug) {
84
+ return SCOPING_NOTES[slug] || null;
85
+ }
86
+
87
+ // ── the baseline-reproduction control ────────────────────────────────────────
88
+ // Spec 009 AC-6, and the thing that makes the free pinned arm honest.
89
+ //
90
+ // Report #006 reuses #005's receipts as the pinned-text arm. That is valid only
91
+ // if the substrate has not moved, and `run.model_release_date` is null on every
92
+ // #005 receipt, so id equality is the only version evidence a receipt carries. A
93
+ // provider that re-points a concrete id at a new snapshot is invisible to it.
94
+ //
95
+ // It does not have to be. Every fresh run emits a BASELINE arm: the same cases,
96
+ // the same substrate, and no skill text at all. The revision cannot touch it by
97
+ // construction, so comparing the fresh baseline against the reused receipt's
98
+ // baseline re-measures exactly the thing id equality could not prove β€” at no
99
+ // extra cost, because that arm is already paid for.
100
+ //
101
+ // A cell whose baselines do not reproduce is NOT MEASURED. The reuse is a tested
102
+ // prediction, not an assumption the report asks the reader to grant.
103
+ function baselineBands(receipt) {
104
+ const out = {};
105
+ for (const c of receipt.results.cases) {
106
+ if (c.mode !== 'baseline') continue;
107
+ out[c.id] = { mean: c.mean != null ? c.mean : c.score, stddev: c.stddev || 0 };
108
+ }
109
+ return out;
110
+ }
111
+
112
+ function baselineControl(reused, fresh) {
113
+ const A = baselineBands(reused);
114
+ const B = baselineBands(fresh);
115
+ const ids = [...new Set([...Object.keys(A), ...Object.keys(B)])];
116
+
117
+ const perCase = ids.map((id) => {
118
+ const before = A[id] || null;
119
+ const after = B[id] || null;
120
+ if (!before || !after) return { id, before, after, delta: null, moved: false, missing: true };
121
+ const delta = round(after.mean - before.mean);
122
+ // The same rule the study uses everywhere else: band separation AND the
123
+ // effect floor. A baseline that wobbles inside its band has not moved.
124
+ const raw = bandVerdict(before.mean, before.stddev, after.mean, after.stddev);
125
+ const separated = raw === 'regression' || raw === 'improvement';
126
+ return { id, before, after, delta, moved: separated && Math.abs(delta) >= EFFECT_FLOOR, missing: false };
127
+ });
128
+
129
+ const movedCases = perCase.filter((r) => r.moved);
130
+ const missing = perCase.filter((r) => r.missing);
131
+ const reproduced = movedCases.length === 0 && missing.length === 0;
132
+ const aggDelta = round(
133
+ (fresh.comparison && fresh.comparison.baseline_score != null ? fresh.comparison.baseline_score : 0)
134
+ - (reused.comparison && reused.comparison.baseline_score != null ? reused.comparison.baseline_score : 0),
135
+ );
136
+
137
+ return {
138
+ reproduced,
139
+ blocked: !reproduced,
140
+ verdict: reproduced ? 'MEASURED' : 'NOT MEASURED',
141
+ moved_cases: movedCases.map((r) => r.id),
142
+ missing_cases: missing.map((r) => r.id),
143
+ aggregate_baseline_delta: aggDelta,
144
+ floor: EFFECT_FLOOR,
145
+ // THE CONTROL PROVES NON-REPRODUCTION. IT CANNOT SAY WHY. These strings used
146
+ // to read 'the substrate moved' and 'the substrate held still' β€” a cause,
147
+ // asserted by a comparison that measures two baseline arms and nothing else.
148
+ // A 120-call stability probe then found generation-level sampling noise large
149
+ // enough to account for every gap this control saw, with no substrate movement
150
+ // required, and the report page retracted the claim while three committed
151
+ // control records still carried it (approval finding F-009-N). Reason strings
152
+ // only: no score, sample, hash or verdict changed with this edit.
153
+ reason: reproduced
154
+ ? 'the fresh baseline reproduces the reused receipt\'s baseline within the band and the floor, so the pinned-arm reuse stands for this cell'
155
+ : `the fresh baseline does not reproduce the reused receipt's baseline (${movedCases.length} case(s) moved beyond the band and the ${EFFECT_FLOOR} floor${missing.length ? `, ${missing.length} case(s) absent on one side` : ''}) β€” the reused pinned arm is not comparable to the fresh arm, so revision drift cannot be separated from whatever else changed in this cell; the control establishes non-reproduction and does not identify a cause`,
156
+ perCase,
157
+ };
158
+ }
159
+
160
+ module.exports = {
161
+ revisionHeadline, revisionClass, fairnessSentence, scopingNote, baselineControl,
162
+ };
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "driftproof",
3
- "version": "0.5.0",
4
- "description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit signed dated receipts, and diff receipts into drift reports.",
3
+ "version": "0.6.0",
4
+ "description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
7
7
  "agent-skills",
package/spec/RECEIPT.md CHANGED
@@ -1,7 +1,7 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
2
  # Driftproof receipt β€” spec v0.4
3
3
 
4
- A **receipt** is a signed, dated record of running one agent skill's eval suite
4
+ A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
5
5
  **with** and **without** the skill on one model version, with the judge **sampled**
6
6
  so every score carries a confidence band. Receipts are the unit of evidence
7
7
  Driftproof produces; diffing two receipts across model releases yields a **drift
@@ -164,7 +164,7 @@ The v0.3.1 schema gained an **additive interop revision** so receipts can be
164
164
 
165
165
  ## Fields
166
166
 
167
- ### `schema_version` (string, required) β€” `"0.3.1"`.
167
+ ### `schema_version` (string, required) β€” `"0.4"`.
168
168
 
169
169
  ### `skill` (object, required)
170
170
  | field | type | notes |
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://json-schema.org/draft/2020-12/schema",
3
3
  "$id": "https://driftproofhq.com/spec/receipt.schema.json",
4
4
  "title": "driftproof receipt",
5
- "description": "A signed, dated record of running one agent skill's eval suite with and without the skill on one model version, with sampled judge scores and confidence bands. Receipt spec v0.4 β€” additive over v0.3.1: adds per-case, per-arm generation `usage` (input/output/cached tokens + measured wall_ms) captured from the surfaces that report it, a separate per-case `judge_usage` (measurement overhead, EXCLUDED from every skill-value figure by construction β€” `economics.judge_excluded` is const true), a run-level `run.pricing_snapshot` freezing the registry prices the derived dollar figures were computed from (so a receipt keeps its meaning when prices later change), and a derived `economics` block (per-arm mean cost/call, skill incremental cost per call and per 1k calls, output-length delta, median wall_ms with IQR). The three value axes β€” accuracy lift, cost, latency β€” are recorded separately and NEVER combined into a composite score. All v0.3.1 semantics are unchanged and every prior receipt still validates against its own frozen schema (v0.1, v0.2, v0.3, v0.3.1). The TESTED tightening (see allOf) is unchanged: the interop relaxations remain available only below TESTED.",
5
+ "description": "A hash-verified, dated record of running one agent skill's eval suite with and without the skill on one model version, with sampled judge scores and confidence bands. Receipt spec v0.4 β€” additive over v0.3.1: adds per-case, per-arm generation `usage` (input/output/cached tokens + measured wall_ms) captured from the surfaces that report it, a separate per-case `judge_usage` (measurement overhead, EXCLUDED from every skill-value figure by construction β€” `economics.judge_excluded` is const true), a run-level `run.pricing_snapshot` freezing the registry prices the derived dollar figures were computed from (so a receipt keeps its meaning when prices later change), and a derived `economics` block (per-arm mean cost/call, skill incremental cost per call and per 1k calls, output-length delta, median wall_ms with IQR). The three value axes β€” accuracy lift, cost, latency β€” are recorded separately and NEVER combined into a composite score. All v0.3.1 semantics are unchanged and every prior receipt still validates against its own frozen schema (v0.1, v0.2, v0.3, v0.3.1). The TESTED tightening (see allOf) is unchanged: the interop relaxations remain available only below TESTED.",
6
6
  "type": "object",
7
7
  "additionalProperties": false,
8
8
  "required": [