driftproof 0.4.0 β 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +72 -17
- package/bin/driftproof +28 -5
- package/config.js +18 -2
- package/lib/diff.js +58 -7
- package/lib/hygiene.js +113 -0
- package/lib/judge.js +11 -3
- package/lib/provider.js +55 -25
- package/lib/receipt.js +8 -2
- package/lib/revision.js +162 -0
- package/lib/run.js +32 -5
- package/lib/stub.js +24 -3
- package/lib/usage.js +168 -0
- package/lib/value.js +502 -0
- package/package.json +2 -2
- package/spec/RECEIPT.md +52 -9
- package/spec/receipt.schema.json +703 -76
- package/spec/receipt.v0.3.1.schema.json +642 -0
package/README.md
CHANGED
|
@@ -9,15 +9,17 @@
|
|
|
9
9
|
Driftproof is an open **receipt spec** plus a **runner** that measures whether an
|
|
10
10
|
agent skill actually helps β by running the skill's eval suite **with** and
|
|
11
11
|
**without** the skill on a named model version, judging each case several times to
|
|
12
|
-
get a confidence band, and emitting a
|
|
12
|
+
get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
|
|
13
13
|
across model releases and you get a **drift report**.
|
|
14
14
|
|
|
15
15
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
16
16
|
format; it does not invent its own.
|
|
17
17
|
|
|
18
|
-
π **
|
|
19
|
-
hand-entered), spanning
|
|
20
|
-
verdict rule and differ
|
|
18
|
+
π **Six published reports** (each re-derived from committed receipts, nothing
|
|
19
|
+
hand-entered), spanning five published report types, the fifth being revision
|
|
20
|
+
drift. All six share one band-based, floor-gated verdict rule and differ in what
|
|
21
|
+
moves underneath the skill β or, in the value report, in which axes are
|
|
22
|
+
measured:
|
|
21
23
|
|
|
22
24
|
- **[Report #001](https://driftproofhq.com/reports/001/)** β *release drift*: ten
|
|
23
25
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
|
|
@@ -31,10 +33,29 @@ verdict rule and differ only in what moves underneath the skill:
|
|
|
31
33
|
`claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
|
|
32
34
|
3 durable, 5 tier-dependent, 0 regressions, 2 no effect β encoded expertise
|
|
33
35
|
survives the frontier tier.
|
|
36
|
+
- **[Report #005](https://driftproofhq.com/reports/005/)** β *value*: what a skill
|
|
37
|
+
*costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
|
|
38
|
+
three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) β
|
|
39
|
+
14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
|
|
40
|
+
saving instead, having improved quality while reducing cost.
|
|
41
|
+
*Three of those cells carry an amendment (v1.1, applied when Report #006
|
|
42
|
+
published): their lifts rest on single-draw baselines since shown
|
|
43
|
+
unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
|
|
44
|
+
substrate-disagreement findings below are unaffected.*
|
|
45
|
+
- **[Report #006](https://driftproofhq.com/reports/006/)** β *revision drift*: the pinned skill revision against the one
|
|
46
|
+
upstream ships today, on a held substrate. **The reuse premise was tested and
|
|
47
|
+
refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
|
|
48
|
+
control. A 120-call probe found generation-level sampling noise 3.2Γ and 7.5Γ
|
|
49
|
+
larger than the judge-level noise this instrument actually samples β enough to
|
|
50
|
+
account for every gap the controls saw without any other cause being
|
|
51
|
+
established β and the receipt spec gains generation sampling as a result. **No
|
|
52
|
+
cause is asserted**; the control proves non-reproduction and cannot say why.
|
|
53
|
+
*The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
|
|
34
54
|
|
|
35
55
|
βοΈ The launch essay, **[Three model releases later: what actually happens to agent
|
|
36
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads
|
|
37
|
-
|
|
56
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all six reports
|
|
57
|
+
together: what moves underneath a skill, and what the skill costs to run. Revised
|
|
58
|
+
2026-08-29; every figure in it is gate-checked against the report page it cites.
|
|
38
59
|
|
|
39
60
|
## Why
|
|
40
61
|
|
|
@@ -55,6 +76,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
|
|
|
55
76
|
the judge and report confidence bands**, and to **only claim a regression when the
|
|
56
77
|
bands don't overlap**. A tool that cries wolf is worse than no tool.
|
|
57
78
|
|
|
79
|
+
**A verdict without a price is half an answer.** The same receipts price the
|
|
80
|
+
marginal cost of a skill firing, and Report #005 found the dominant cost driver is
|
|
81
|
+
not the skill's own text but the input it causes the model to pull in: across those
|
|
82
|
+
30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
|
|
83
|
+
tracks it at only `r = +0.33`, and one 738-token skill drew 34Γ its own size in
|
|
84
|
+
extra input. Identical token deltas also price very differently across substrates β
|
|
85
|
+
the same skill at near-identical deltas costs 3.3Γ more on `claude-fable-5` than on
|
|
86
|
+
`claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
|
|
87
|
+
|
|
58
88
|
## Quickstart β receipt for your own skill in ~10 minutes
|
|
59
89
|
|
|
60
90
|
You need Node β₯ 22 and an `ANTHROPIC_API_KEY`.
|
|
@@ -86,6 +116,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
|
|
|
86
116
|
minimum real move required before a change counts as more than noise β band
|
|
87
117
|
separation *plus* a floor-sized delta, never either alone.
|
|
88
118
|
|
|
119
|
+
**What the band does not cover.** A verdict rests on **one generation draw per
|
|
120
|
+
arm**: the band is the spread of the *judge* re-scoring that single response, not
|
|
121
|
+
the spread of the model writing a different one. Report #006 measured the second
|
|
122
|
+
directly and found it larger β draw-to-draw spread up to **sd 0.186** on the 0β1
|
|
123
|
+
scale, against judge-level noise several times smaller. So treat a surprising
|
|
124
|
+
single-run verdict as **provisional and worth re-running** before you act on it.
|
|
125
|
+
Generation sampling lands in the next receipt spec; until it does, this is a
|
|
126
|
+
limit of the instrument, stated rather than implied.
|
|
127
|
+
|
|
89
128
|
### Install
|
|
90
129
|
|
|
91
130
|
```bash
|
|
@@ -162,7 +201,7 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
162
201
|
|
|
163
202
|
```jsonc
|
|
164
203
|
{
|
|
165
|
-
"schema_version": "0.
|
|
204
|
+
"schema_version": "0.4",
|
|
166
205
|
"skill": { "name": "commit-message-conventions", "version": "0.2.0",
|
|
167
206
|
"content_hash": "β¦sha256 over SKILL.md + bundled filesβ¦" },
|
|
168
207
|
"suite": { "format": "agentskills.io/evals", "suite_hash": "β¦", "case_count": 10 },
|
|
@@ -171,7 +210,7 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
171
210
|
"model_release_date": "2025-10-01",
|
|
172
211
|
"provider": "anthropic",
|
|
173
212
|
"surface": "claude-cli",
|
|
174
|
-
"runner_version": "0.
|
|
213
|
+
"runner_version": "0.6.0",
|
|
175
214
|
"date_utc": "2026-07-27Tβ¦Z",
|
|
176
215
|
"registry": "registered",
|
|
177
216
|
"transcripts": "hashes-only",
|
|
@@ -194,6 +233,12 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
194
233
|
},
|
|
195
234
|
"comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
|
|
196
235
|
"delta": 0.39, "delta_uncertainty": 0.036 },
|
|
236
|
+
// v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
|
|
237
|
+
// freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
|
|
238
|
+
// holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
|
|
239
|
+
// mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
|
|
240
|
+
// skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
|
|
241
|
+
// output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
|
|
197
242
|
"verification_level": "TESTED",
|
|
198
243
|
"receipt_hash": "β¦sha256 of the canonical receipt with this field removedβ¦"
|
|
199
244
|
}
|
|
@@ -209,6 +254,12 @@ Key ideas:
|
|
|
209
254
|
- **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
|
|
210
255
|
- **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
|
|
211
256
|
`TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
|
|
257
|
+
- **`economics`** is *derived*, never a second measurement: the token delta is the
|
|
258
|
+
durable fact, and the dollars are exactly those tokens at the rates frozen into
|
|
259
|
+
`run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
|
|
260
|
+
Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
|
|
261
|
+
figure (`judge_excluded` is `const true`) β measuring the skill is our cost, not
|
|
262
|
+
the skill's.
|
|
212
263
|
- **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
|
|
213
264
|
signature β see the spec's open questions).
|
|
214
265
|
|
|
@@ -233,7 +284,7 @@ jobs:
|
|
|
233
284
|
runs-on: ubuntu-latest
|
|
234
285
|
steps:
|
|
235
286
|
- uses: actions/checkout@v4
|
|
236
|
-
- uses: driftproofhq/driftproof@v0.
|
|
287
|
+
- uses: driftproofhq/driftproof@v0.6.0
|
|
237
288
|
with:
|
|
238
289
|
skill-dir: skills/my-skill
|
|
239
290
|
models: claude-haiku-4-5
|
|
@@ -272,11 +323,14 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
272
323
|
|
|
273
324
|
## Reports
|
|
274
325
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
the
|
|
279
|
-
(
|
|
326
|
+
Six reports are published, spanning five report types. A report page lives at a
|
|
327
|
+
draft path β `docs/reports/NNN-draft/` β until the publish sequence renames it, and
|
|
328
|
+
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
329
|
+
(see the roll at the top of this README, and
|
|
330
|
+
[REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
|
|
331
|
+
Each report and every verdict in it are **re-derived from the receipts** committed
|
|
332
|
+
under [`receipts/`](receipts/)
|
|
333
|
+
(`receipts/report-001/` β¦ `receipts/report-006/`) β nothing is hand-entered.
|
|
280
334
|
|
|
281
335
|
Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
|
|
282
336
|
fetched at run time from a pinned commit and verified by sha256 against
|
|
@@ -290,9 +344,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models Γ with/basel
|
|
|
290
344
|
node scripts/build-report-001.js # re-derive the report from the receipts
|
|
291
345
|
```
|
|
292
346
|
|
|
293
|
-
Reports #002β#
|
|
294
|
-
(`scripts/prepare-report-00N.js`
|
|
295
|
-
|
|
347
|
+
Reports #002β#005 have their own runners
|
|
348
|
+
(`scripts/prepare-report-00N.js` β Report #005's is
|
|
349
|
+
[`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
|
|
350
|
+
same fetch β run β re-derive shape.
|
|
296
351
|
|
|
297
352
|
**Model-release triggers are live**: `scripts/release-watch.js` (keyless β it
|
|
298
353
|
reads the public models registry) notices a new model release, re-runs the
|
package/bin/driftproof
CHANGED
|
@@ -8,7 +8,7 @@ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = req
|
|
|
8
8
|
const { loadSkill } = require('../lib/skill');
|
|
9
9
|
const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
|
|
10
10
|
const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
|
|
11
|
-
const { buildDriftReport } = require('../lib/diff');
|
|
11
|
+
const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
|
|
12
12
|
const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
|
|
13
13
|
const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
|
|
14
14
|
const { registryStatus } = require('../lib/models');
|
|
@@ -76,7 +76,7 @@ USAGE
|
|
|
76
76
|
${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
|
|
77
77
|
[--judge-model M] [--concurrency N] [--max-usd N]
|
|
78
78
|
[--keep-transcripts] [--out DIR]
|
|
79
|
-
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
|
|
79
|
+
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
|
|
80
80
|
${PROJECT_NAME} validate <receipt.json>
|
|
81
81
|
${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
|
|
82
82
|
${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
|
|
@@ -250,9 +250,32 @@ function cmdDiff(positional, flags) {
|
|
|
250
250
|
if (!verifyReceiptHash(r)) console.error(` β ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
|
|
251
251
|
}
|
|
252
252
|
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
253
|
+
// --mode revision inverts the axis: the skill text is the variable under test
|
|
254
|
+
// and the substrate is the control. The fields release drift merely warns about
|
|
255
|
+
// are preconditions here, so a pair that is not a revision pair is REFUSED
|
|
256
|
+
// (exit 6) rather than rendered with a caveat nobody reads. A differing model
|
|
257
|
+
// would be release drift wearing a revision label β the one confound this mode
|
|
258
|
+
// exists to exclude β and an EQUAL content_hash has no revision to measure.
|
|
259
|
+
const mode = flags.mode || 'release';
|
|
260
|
+
if (mode !== 'release' && mode !== 'revision') {
|
|
261
|
+
console.error(`unknown --mode "${mode}" β supported: release, revision`);
|
|
262
|
+
process.exit(2);
|
|
263
|
+
}
|
|
264
|
+
if (mode === 'revision') {
|
|
265
|
+
const problem = revisionPairProblem(a, b);
|
|
266
|
+
if (problem) {
|
|
267
|
+
const why = problem === 'skill.content_hash'
|
|
268
|
+
? 'the two receipts carry the SAME skill.content_hash β there is no revision between them to measure'
|
|
269
|
+
: `${problem} differs between the two receipts β revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
|
|
270
|
+
console.error(` β REFUSED (--mode revision): ${why}.`);
|
|
271
|
+
console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
|
|
272
|
+
process.exit(6);
|
|
273
|
+
}
|
|
274
|
+
}
|
|
275
|
+
|
|
276
|
+
const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
|
|
277
|
+
const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
|
|
278
|
+
const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
|
|
256
279
|
|
|
257
280
|
if (flags.out) {
|
|
258
281
|
fs.writeFileSync(path.resolve(flags.out), markdown);
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.6.0';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
|
@@ -22,7 +22,14 @@ const SUITE_FORMAT = 'agentskills.io/evals';
|
|
|
22
22
|
// surface enums (openai-api/openai-cli), optional run.surface_overhead_note,
|
|
23
23
|
// optional per-case checks[] (deterministic post-checks), and optional
|
|
24
24
|
// skill.tokens (value-per-token axis). v0.1/v0.2/v0.3 receipts still load.
|
|
25
|
-
|
|
25
|
+
// v0.4 (additive over v0.3.1) adds the ECONOMICS axis: per-case, per-arm
|
|
26
|
+
// generation `usage` (input/output/cached tokens + measured wall_ms), a
|
|
27
|
+
// separate per-case `judge_usage` (measurement overhead, excluded from
|
|
28
|
+
// every skill-value figure), run.pricing_snapshot (registry prices frozen
|
|
29
|
+
// at run time so derived dollars stay reproducible), and the derived
|
|
30
|
+
// `economics` block. v0.3.1 is frozen as receipt.v0.3.1.schema.json;
|
|
31
|
+
// v0.1/v0.2/v0.3/v0.3.1 receipts all still load.
|
|
32
|
+
const RECEIPT_SCHEMA_VERSION = '0.4';
|
|
26
33
|
|
|
27
34
|
// Hard USD budget defaults per entry point (Week 4). --max-usd overrides any of
|
|
28
35
|
// these. The projection is refused before any call if it exceeds the cap, on
|
|
@@ -73,10 +80,19 @@ const REPORT_004_BASE_MODEL = 'claude-opus-5'; // flagship tier
|
|
|
73
80
|
const REPORT_004_FRONTIER_MODEL = 'claude-fable-5'; // frontier tier (full id β no alias)
|
|
74
81
|
const REPORT_004_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
75
82
|
|
|
83
|
+
// Report #005 is a VALUE report β the fourth report type. It asks what a skill
|
|
84
|
+
// COSTS to run alongside whether it helps, over three substrates, and shows the
|
|
85
|
+
// three axes (accuracy lift / Ξcost / Ξlatency) side by side and never combined.
|
|
86
|
+
// Same suites, same fixed judge; the substrate list spans two providers so the
|
|
87
|
+
// economics are read across surfaces, not within one vendor's pricing.
|
|
88
|
+
const REPORT_005_MODELS = ['claude-sonnet-5', 'claude-fable-5', 'gpt-5.6-sol'];
|
|
89
|
+
const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
90
|
+
|
|
76
91
|
module.exports = {
|
|
77
92
|
PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
|
|
78
93
|
EFFECT_FLOOR, DEV_MAX_USD, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
79
94
|
REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
|
|
80
95
|
REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
|
|
81
96
|
REPORT_004_BASE_MODEL, REPORT_004_FRONTIER_MODEL, REPORT_004_JUDGE_MODEL,
|
|
97
|
+
REPORT_005_MODELS, REPORT_005_JUDGE_MODEL,
|
|
82
98
|
};
|
package/lib/diff.js
CHANGED
|
@@ -3,6 +3,7 @@
|
|
|
3
3
|
|
|
4
4
|
const { bandVerdict, round } = require('./stats');
|
|
5
5
|
const { EFFECT_FLOOR } = require('../config');
|
|
6
|
+
const { revisionHeadline } = require('./revision');
|
|
6
7
|
|
|
7
8
|
// Practical-significance gate applied ON TOP of band separation. bandVerdict()
|
|
8
9
|
// stays a pure geometry test (kept that way so its unit checks are unambiguous);
|
|
@@ -68,7 +69,22 @@ function headlineVerdict(perCase) {
|
|
|
68
69
|
return 'WITHIN NOISE β no case moved beyond its confidence band; the skill holds up.';
|
|
69
70
|
}
|
|
70
71
|
|
|
71
|
-
|
|
72
|
+
// A revision pair is a pair in which the SKILL TEXT is the only thing that
|
|
73
|
+
// moved. `diff` was built for release drift, where the model varies and the text
|
|
74
|
+
// is fixed; this inverts it, so the fields that release drift merely warns about
|
|
75
|
+
// become the preconditions of the comparison. Returns the offending field name,
|
|
76
|
+
// or null when the pair is a valid revision pair.
|
|
77
|
+
function revisionPairProblem(a, b) {
|
|
78
|
+
if ((a.run.model_id || '') !== (b.run.model_id || '')) return 'run.model_id';
|
|
79
|
+
if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) return 'run.provider';
|
|
80
|
+
if ((a.run.surface || '') !== (b.run.surface || '')) return 'run.surface';
|
|
81
|
+
if ((a.suite.suite_hash || '') !== (b.suite.suite_hash || '')) return 'suite.suite_hash';
|
|
82
|
+
if ((a.skill.content_hash || '') === (b.skill.content_hash || '')) return 'skill.content_hash';
|
|
83
|
+
return null;
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
|
|
87
|
+
const revision = mode === 'revision';
|
|
72
88
|
const aB = withSkillBands(a);
|
|
73
89
|
const bB = withSkillBands(b);
|
|
74
90
|
const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
|
|
@@ -99,6 +115,32 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
99
115
|
const regressions = perCase.filter((r) => r.verdict === 'regression');
|
|
100
116
|
|
|
101
117
|
const L = [];
|
|
118
|
+
if (revision) {
|
|
119
|
+
// The axis leads. In a release-drift report the model is what moved and it
|
|
120
|
+
// belongs at the top; here the model is the control and the skill's own text
|
|
121
|
+
// is the finding, so content_hash is the first row a reader meets.
|
|
122
|
+
L.push(`# Revision drift report`);
|
|
123
|
+
L.push('');
|
|
124
|
+
L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
|
|
125
|
+
L.push('');
|
|
126
|
+
L.push(`**Report type:** revision drift β the skill's own text is the variable under test; the substrate is held fixed.`);
|
|
127
|
+
L.push('');
|
|
128
|
+
L.push(`Left column is the **pinned** revision; right column is the **current** upstream revision.`);
|
|
129
|
+
L.push('');
|
|
130
|
+
L.push(`| | ${labelA} | ${labelB} |`);
|
|
131
|
+
L.push(`|---|---|---|`);
|
|
132
|
+
L.push(`| skill content_hash | \`${short(a.skill.content_hash)}\` | \`${short(b.skill.content_hash)}\` |`);
|
|
133
|
+
L.push(`| model (held) | \`${a.run.model_id}\` | \`${b.run.model_id}\` |`);
|
|
134
|
+
L.push(`| provider (held) | ${a.run.provider || 'anthropic'} | ${b.run.provider || 'anthropic'} |`);
|
|
135
|
+
L.push(`| surface (held) | ${a.run.surface} | ${b.run.surface} |`);
|
|
136
|
+
L.push(`| suite_hash (held) | \`${short(a.suite.suite_hash)}\` | \`${short(b.suite.suite_hash)}\` |`);
|
|
137
|
+
L.push(`| run date (UTC) | ${a.run.date_utc} | ${b.run.date_utc} |`);
|
|
138
|
+
L.push(`| judge samples/case | ${(a.run.judge || {}).samples || 1} | ${(b.run.judge || {}).samples || 1} |`);
|
|
139
|
+
L.push(`| with_skill (mean Β± band) | ${bandStr(aAgg)} | ${bandStr(bAgg)} |`);
|
|
140
|
+
L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
|
|
141
|
+
L.push(`| skill lift (Ξ) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
|
|
142
|
+
L.push('');
|
|
143
|
+
} else {
|
|
102
144
|
L.push(`# Drift report`);
|
|
103
145
|
L.push('');
|
|
104
146
|
L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
|
|
@@ -115,10 +157,19 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
115
157
|
L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
|
|
116
158
|
L.push(`| skill lift (Ξ) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
|
|
117
159
|
L.push('');
|
|
160
|
+
}
|
|
118
161
|
|
|
119
162
|
const warnings = [];
|
|
120
|
-
|
|
121
|
-
|
|
163
|
+
// In revision mode the changed skill text is the INDEPENDENT VARIABLE, not
|
|
164
|
+
// contamination, and the held-constant substrate is what makes the comparison
|
|
165
|
+
// valid. The release-drift caveat below says the opposite of both, so it is
|
|
166
|
+
// replaced rather than suppressed: a reader is told what is held, and why a
|
|
167
|
+
// separated band is attributable to the revision.
|
|
168
|
+
if (revision) {
|
|
169
|
+
warnings.push(`model \`${a.run.model_id}\`, provider ${a.run.provider || 'anthropic'}, surface ${a.run.surface} and suite_hash \`${short(a.suite.suite_hash)}\` are held constant across both receipts β the skill text is the only variable under test, so a band-separated move above the ${EFFECT_FLOOR} floor is attributable to the revision.`);
|
|
170
|
+
}
|
|
171
|
+
if (!revision && a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β the skill itself changed between receipts, so drift mixes skill edits with model drift.');
|
|
172
|
+
if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β the eval suite changed; per-case comparison may be misleading.');
|
|
122
173
|
if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β comparison is not meaningful.`);
|
|
123
174
|
if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β non-overlap can only be trusted when both sides are sampled.');
|
|
124
175
|
if (!measured) warnings.push(`verdicts NOT computed β ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
|
|
@@ -126,8 +177,8 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
126
177
|
// providers is a skill-DURABILITY comparison across substrates, not model drift
|
|
127
178
|
// over time; across surfaces, sampling control differs. Both are flagged so a
|
|
128
179
|
// reader never mistakes one for the other (see docs/neutrality.html).
|
|
129
|
-
if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
|
|
130
|
-
if (a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β sampling control differs between surfaces; compare with care.`);
|
|
180
|
+
if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
|
|
181
|
+
if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β sampling control differs between surfaces; compare with care.`);
|
|
131
182
|
if (warnings.length) {
|
|
132
183
|
L.push('> **β Caveats**');
|
|
133
184
|
for (const w of warnings) L.push(`> - ${w}`);
|
|
@@ -142,7 +193,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
142
193
|
L.push(`**NOT MEASURED β ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
|
|
143
194
|
L.push('');
|
|
144
195
|
} else {
|
|
145
|
-
L.push(`**${headlineVerdict(perCase)}**`);
|
|
196
|
+
L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
|
|
146
197
|
L.push('');
|
|
147
198
|
L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} β ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin} within noise${nFloor ? ` (${nFloor} of them band-separated but below the ${EFFECT_FLOOR} effect floor)` : ''}.`);
|
|
148
199
|
L.push('');
|
|
@@ -170,4 +221,4 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
170
221
|
return { markdown: L.join('\n'), perCase, headlineDelta, regressions };
|
|
171
222
|
}
|
|
172
223
|
|
|
173
|
-
module.exports = { buildDriftReport, withSkillBands };
|
|
224
|
+
module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
|
package/lib/hygiene.js
ADDED
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
//
|
|
3
|
+
// The single definition of what counts as a leak.
|
|
4
|
+
//
|
|
5
|
+
// Two callers scan for the same classes of secret: `tests/gate.js`, which walks
|
|
6
|
+
// the working tree at gate time, and `scripts/merge-check.js`, which walks the
|
|
7
|
+
// candidate tree read out of the commit before permitting a merge. They must
|
|
8
|
+
// agree. Two copies of a deny-list drift, and the copy that quietly stops
|
|
9
|
+
// matching still reads as protection β the same argument the acknowledgment
|
|
10
|
+
// deny-list bite test already makes, applied to the scanner itself.
|
|
11
|
+
//
|
|
12
|
+
// Backlog B-8 records four approval records that reached `dev` carrying a host
|
|
13
|
+
// path. The recorded cause β "the scan walks tracked files" β is false, and this
|
|
14
|
+
// module's existence does not depend on it: `listFiles()` in the repo gate has
|
|
15
|
+
// always enumerated untracked files too, and a planted path in an untracked
|
|
16
|
+
// evidence record does take the repo gate red. The real window is VANTAGE. An
|
|
17
|
+
// approval session measures in a disposable copy made before it writes its
|
|
18
|
+
// record into the shared checkout, so the record is simply not in the tree that
|
|
19
|
+
// was scanned. Scanning the working tree harder cannot close that. Scanning the
|
|
20
|
+
// commit that is about to merge does, and that is what merge-check now uses this
|
|
21
|
+
// for.
|
|
22
|
+
//
|
|
23
|
+
// Patterns are written so a regex literal cannot match its own text: a
|
|
24
|
+
// metacharacter or a character class follows each fixed prefix.
|
|
25
|
+
|
|
26
|
+
const dec = (b64) => JSON.parse(Buffer.from(b64, 'base64').toString('utf8'));
|
|
27
|
+
|
|
28
|
+
// Build-box host + private IP, base64 so the plaintext never lands in a file.
|
|
29
|
+
const HOSTIP = dec('WyJpcC0xNzItMzEtNDAtMTU5LmFwLXNvdXRoZWFzdC0xLmNvbXB1dGUuaW50ZXJuYWwiLCIxNzIuMzEuNDAuMTU5Il0=');
|
|
30
|
+
|
|
31
|
+
const EMAIL_ALLOW = new Set(['example.com', 'example.org', 'driftproofhq.com']);
|
|
32
|
+
|
|
33
|
+
const PATTERNS = [
|
|
34
|
+
{ name: 'home-path', re: /\/home\/[a-z0-9_-]+\/|\/Users\/[A-Za-z0-9_-]+\// },
|
|
35
|
+
{ name: 'ec2-internal-host', re: /ip-\d+-\d+-\d+-\d+\.[a-z0-9.-]*compute\.(internal|amazonaws\.com)/i },
|
|
36
|
+
{ name: 'private-ip', re: /\b(10\.\d{1,3}\.\d{1,3}\.\d{1,3}|172\.(1[6-9]|2\d|3[01])\.\d{1,3}\.\d{1,3}|192\.168\.\d{1,3}\.\d{1,3})\b/ },
|
|
37
|
+
{ name: 'anthropic-key', re: /sk-ant-[A-Za-z0-9_-]{8,}/ },
|
|
38
|
+
{ name: 'github-token', re: /gh[pousr]_[A-Za-z0-9]{20,}/ },
|
|
39
|
+
{ name: 'aws-key', re: /\bAKIA[0-9A-Z]{16}\b/ },
|
|
40
|
+
{ name: 'private-key-block', re: /-----BEGIN [A-Z ]*PRIVATE KEY-----/ },
|
|
41
|
+
{ name: 'env-secret-assignment', re: /\b(ANTHROPIC_API_KEY|AWS_SECRET_ACCESS_KEY|OPENAI_API_KEY)\s*=\s*\S+/ },
|
|
42
|
+
// Codex subscription auth material (~/.codex/auth.json contents) must NEVER
|
|
43
|
+
// land in a committed file: the id_token/access_token are JWTs, and an OpenAI
|
|
44
|
+
// secret key is sk-proj-/sk-svcacct-/sk-admin-. We ban the CONTENTS (tokens),
|
|
45
|
+
// not the documented path string (`~/.codex/auth.json` is referenced in help
|
|
46
|
+
// text and docs by design).
|
|
47
|
+
{ name: 'jwt-token', re: /\beyJ[A-Za-z0-9_=-]{10,}\.eyJ[A-Za-z0-9_=-]{10,}\.[A-Za-z0-9_=-]{6,}/ },
|
|
48
|
+
{ name: 'openai-secret-key', re: /\bsk-(proj|svcacct|admin)-[A-Za-z0-9_-]{20,}/ },
|
|
49
|
+
];
|
|
50
|
+
|
|
51
|
+
// A path is a hit on its own name, with no content read: a committed .env file
|
|
52
|
+
// is a leak whatever it happens to contain.
|
|
53
|
+
function scanPath(rel) {
|
|
54
|
+
return /(^|\/)\.env(\.|$)/.test(rel) ? [{ file: rel, kind: 'env-file' }] : [];
|
|
55
|
+
}
|
|
56
|
+
|
|
57
|
+
function scanContent(rel, content) {
|
|
58
|
+
const hits = [];
|
|
59
|
+
for (const p of PATTERNS) {
|
|
60
|
+
const m = content.match(p.re);
|
|
61
|
+
if (m) hits.push({ file: rel, kind: p.name, sample: m[0].slice(0, 24) });
|
|
62
|
+
}
|
|
63
|
+
for (const hv of HOSTIP) if (content.includes(hv)) hits.push({ file: rel, kind: 'build-host-or-ip' });
|
|
64
|
+
const emails = content.match(/\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g) || [];
|
|
65
|
+
for (const e of emails) {
|
|
66
|
+
const dom = e.split('@')[1].toLowerCase();
|
|
67
|
+
if (!EMAIL_ALLOW.has(dom) && !dom.endsWith('.example')) hits.push({ file: rel, kind: 'email', sample: e });
|
|
68
|
+
}
|
|
69
|
+
return hits;
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
// scanFiles(files, read, readLink) β `read(rel)` returns the file's text, or
|
|
73
|
+
// throws/returns null for anything unreadable (binary, deleted, a symlink to
|
|
74
|
+
// nowhere), which is skipped exactly as the gate has always skipped it.
|
|
75
|
+
//
|
|
76
|
+
// `readLink(rel)` is optional and returns a symbolic link's TARGET PATH as a
|
|
77
|
+
// string, or null for an entry that is not a link. Where it is supplied, that
|
|
78
|
+
// target is scanned as the entry's content through `scanContent`, so every
|
|
79
|
+
// pattern above applies to it with no second deny-list to drift (spec 011,
|
|
80
|
+
// AC-2). A link's target is a path, and a path is exactly the class of leak the
|
|
81
|
+
// `home-path` pattern exists to catch: P-1 shipped `node_modules ->
|
|
82
|
+
// /home/<user>/... ` into a public tree past a scan that read the link as an
|
|
83
|
+
// unreadable file and skipped it.
|
|
84
|
+
//
|
|
85
|
+
// The argument is optional so a caller supplying nothing behaves exactly as
|
|
86
|
+
// before. That compatibility is deliberate, and it is also the standing risk: a
|
|
87
|
+
// caller added later is blind by default. The repo gate asserts that every call
|
|
88
|
+
// site supplies a reader, with `scripts/merge-check.js` the one exception β its
|
|
89
|
+
// reader is `git show <ref>:<path>`, and git stores a link's blob as its target
|
|
90
|
+
// string, so the target already arrives as content there.
|
|
91
|
+
function scanFiles(files, read, readLink) {
|
|
92
|
+
const hits = [];
|
|
93
|
+
for (const rel of files) {
|
|
94
|
+
hits.push(...scanPath(rel));
|
|
95
|
+
// The link branch is a single condition on purpose: it is the mutation seam
|
|
96
|
+
// the spec-011 gate neutralises to prove the scanner goes blind without it.
|
|
97
|
+
if (typeof readLink === 'function') {
|
|
98
|
+
let target;
|
|
99
|
+
try { target = readLink(rel); } catch (_e) { target = null; }
|
|
100
|
+
if (typeof target === 'string') {
|
|
101
|
+
hits.push(...scanContent(rel, target));
|
|
102
|
+
continue;
|
|
103
|
+
}
|
|
104
|
+
}
|
|
105
|
+
let c;
|
|
106
|
+
try { c = read(rel); } catch (_e) { continue; }
|
|
107
|
+
if (typeof c !== 'string') continue;
|
|
108
|
+
hits.push(...scanContent(rel, c));
|
|
109
|
+
}
|
|
110
|
+
return hits;
|
|
111
|
+
}
|
|
112
|
+
|
|
113
|
+
module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, scanPath, scanContent, scanFiles };
|
package/lib/judge.js
CHANGED
|
@@ -5,6 +5,7 @@ const { complete, surfaceForModel } = require('./provider');
|
|
|
5
5
|
const { extractJsonObject } = require('./json');
|
|
6
6
|
const { sha256 } = require('./canonical');
|
|
7
7
|
const { mean, stddev } = require('./stats');
|
|
8
|
+
const { sumUsage } = require('./usage');
|
|
8
9
|
|
|
9
10
|
// Rubric-based LLM judge.
|
|
10
11
|
//
|
|
@@ -82,16 +83,16 @@ function judgeSettings(samples, judgeModel) {
|
|
|
82
83
|
// transcript auditability, and optionally retained under --keep-transcripts).
|
|
83
84
|
async function gradeOnce({ task, response, rubric, model, timeoutMs, temperature }) {
|
|
84
85
|
const prompt = buildJudgePrompt({ task, response, rubric });
|
|
85
|
-
const { text, attempts } = await complete({ system: JUDGE_SYSTEM, prompt, model, maxTokens: 400, timeoutMs, temperature });
|
|
86
|
+
const { text, attempts, usage } = await complete({ system: JUDGE_SYSTEM, prompt, model, maxTokens: 400, timeoutMs, temperature });
|
|
86
87
|
let parsed;
|
|
87
88
|
try {
|
|
88
89
|
parsed = extractJsonObject(text);
|
|
89
90
|
} catch (_e) {
|
|
90
91
|
// Unsalvageable judge output β conservative 0 (a judge that can't be parsed
|
|
91
92
|
// must never silently "pass"), tagged so the caller can see it happened.
|
|
92
|
-
return { score: 0, reason: 'judge output unparseable', unparsed: true, raw: String(text || ''), attempts: attempts || 1 };
|
|
93
|
+
return { score: 0, reason: 'judge output unparseable', unparsed: true, raw: String(text || ''), attempts: attempts || 1, usage };
|
|
93
94
|
}
|
|
94
|
-
return { score: clamp01(parsed.score), reason: String(parsed.reason || '').slice(0, 300), raw: String(text || ''), attempts: attempts || 1 };
|
|
95
|
+
return { score: clamp01(parsed.score), reason: String(parsed.reason || '').slice(0, 300), raw: String(text || ''), attempts: attempts || 1, usage };
|
|
95
96
|
}
|
|
96
97
|
|
|
97
98
|
// Grade a response N times and return the sampled distribution:
|
|
@@ -103,6 +104,7 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
|
|
|
103
104
|
const scores = [];
|
|
104
105
|
const reasons = [];
|
|
105
106
|
const rawTexts = [];
|
|
107
|
+
const usages = [];
|
|
106
108
|
let attemptsTotal = 0;
|
|
107
109
|
for (let i = 0; i < samples; i++) {
|
|
108
110
|
let r;
|
|
@@ -116,6 +118,7 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
|
|
|
116
118
|
throw e;
|
|
117
119
|
}
|
|
118
120
|
attemptsTotal += r.attempts || 1;
|
|
121
|
+
usages.push(r.usage || null);
|
|
119
122
|
scores.push(r.score);
|
|
120
123
|
reasons.push(r.reason);
|
|
121
124
|
rawTexts.push(r.raw || '');
|
|
@@ -134,6 +137,11 @@ async function gradeSamples({ task, response, rubric, model, samples = 5, timeou
|
|
|
134
137
|
model_id: model,
|
|
135
138
|
rubric_hash: rubricHash(rubric),
|
|
136
139
|
attempts: attemptsTotal,
|
|
140
|
+
// v0.4: the measurement overhead of grading this one case β the SUM over all
|
|
141
|
+
// N judge calls. Recorded in the receipt as the case's `judge_usage` and
|
|
142
|
+
// EXCLUDED from every skill-value figure (lib/value.js): it is a cost we
|
|
143
|
+
// impose to measure, not a cost of running the skill.
|
|
144
|
+
usage: sumUsage(usages),
|
|
137
145
|
};
|
|
138
146
|
}
|
|
139
147
|
|