driftproof 0.5.0 β 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +71 -16
- package/bin/driftproof +28 -5
- package/config.js +1 -1
- package/lib/diff.js +58 -7
- package/lib/hygiene.js +113 -0
- package/lib/revision.js +162 -0
- package/package.json +2 -2
- package/spec/RECEIPT.md +2 -2
- package/spec/receipt.schema.json +1 -1
package/README.md
CHANGED
|
@@ -9,15 +9,17 @@
|
|
|
9
9
|
Driftproof is an open **receipt spec** plus a **runner** that measures whether an
|
|
10
10
|
agent skill actually helps β by running the skill's eval suite **with** and
|
|
11
11
|
**without** the skill on a named model version, judging each case several times to
|
|
12
|
-
get a confidence band, and emitting a
|
|
12
|
+
get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
|
|
13
13
|
across model releases and you get a **drift report**.
|
|
14
14
|
|
|
15
15
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
16
16
|
format; it does not invent its own.
|
|
17
17
|
|
|
18
|
-
π **
|
|
19
|
-
hand-entered), spanning
|
|
20
|
-
verdict rule and differ
|
|
18
|
+
π **Six published reports** (each re-derived from committed receipts, nothing
|
|
19
|
+
hand-entered), spanning five published report types, the fifth being revision
|
|
20
|
+
drift. All six share one band-based, floor-gated verdict rule and differ in what
|
|
21
|
+
moves underneath the skill β or, in the value report, in which axes are
|
|
22
|
+
measured:
|
|
21
23
|
|
|
22
24
|
- **[Report #001](https://driftproofhq.com/reports/001/)** β *release drift*: ten
|
|
23
25
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
|
|
@@ -31,10 +33,29 @@ verdict rule and differ only in what moves underneath the skill:
|
|
|
31
33
|
`claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
|
|
32
34
|
3 durable, 5 tier-dependent, 0 regressions, 2 no effect β encoded expertise
|
|
33
35
|
survives the frontier tier.
|
|
36
|
+
- **[Report #005](https://driftproofhq.com/reports/005/)** β *value*: what a skill
|
|
37
|
+
*costs* to run, on three axes (accuracy, cost, latency); the same ten suites on
|
|
38
|
+
three substrates (`claude-sonnet-5`, `claude-fable-5`, `gpt-5.6-sol`) β
|
|
39
|
+
14 of 30 cells cleared the floor on aggregate: 10 carry a price and 4 report a
|
|
40
|
+
saving instead, having improved quality while reducing cost.
|
|
41
|
+
*Three of those cells carry an amendment (v1.1, applied when Report #006
|
|
42
|
+
published): their lifts rest on single-draw baselines since shown
|
|
43
|
+
unstable. Cause-agnostic, no corrected figures offered, and the cost-driver and
|
|
44
|
+
substrate-disagreement findings below are unaffected.*
|
|
45
|
+
- **[Report #006](https://driftproofhq.com/reports/006/)** β *revision drift*: the pinned skill revision against the one
|
|
46
|
+
upstream ships today, on a held substrate. **The reuse premise was tested and
|
|
47
|
+
refused: 3 of 3 cells returned no verdict**, each blocked by its own baseline
|
|
48
|
+
control. A 120-call probe found generation-level sampling noise 3.2Γ and 7.5Γ
|
|
49
|
+
larger than the judge-level noise this instrument actually samples β enough to
|
|
50
|
+
account for every gap the controls saw without any other cause being
|
|
51
|
+
established β and the receipt spec gains generation sampling as a result. **No
|
|
52
|
+
cause is asserted**; the control proves non-reproduction and cannot say why.
|
|
53
|
+
*The tally a refusal carries: 3 cells, 0 measured, 3 refused.*
|
|
34
54
|
|
|
35
55
|
βοΈ The launch essay, **[Three model releases later: what actually happens to agent
|
|
36
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads
|
|
37
|
-
|
|
56
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all six reports
|
|
57
|
+
together: what moves underneath a skill, and what the skill costs to run. Revised
|
|
58
|
+
2026-08-29; every figure in it is gate-checked against the report page it cites.
|
|
38
59
|
|
|
39
60
|
## Why
|
|
40
61
|
|
|
@@ -55,6 +76,15 @@ by more than the drift you're trying to detect. Driftproof's answer is to **samp
|
|
|
55
76
|
the judge and report confidence bands**, and to **only claim a regression when the
|
|
56
77
|
bands don't overlap**. A tool that cries wolf is worse than no tool.
|
|
57
78
|
|
|
79
|
+
**A verdict without a price is half an answer.** The same receipts price the
|
|
80
|
+
marginal cost of a skill firing, and Report #005 found the dominant cost driver is
|
|
81
|
+
not the skill's own text but the input it causes the model to pull in: across those
|
|
82
|
+
30 cells the input delta tracks cost at `r = +0.92` while the skill's own length
|
|
83
|
+
tracks it at only `r = +0.33`, and one 738-token skill drew 34Γ its own size in
|
|
84
|
+
extra input. Identical token deltas also price very differently across substrates β
|
|
85
|
+
the same skill at near-identical deltas costs 3.3Γ more on `claude-fable-5` than on
|
|
86
|
+
`claude-sonnet-5`, which is exactly their input-rate ratio in the frozen snapshot.
|
|
87
|
+
|
|
58
88
|
## Quickstart β receipt for your own skill in ~10 minutes
|
|
59
89
|
|
|
60
90
|
You need Node β₯ 22 and an `ANTHROPIC_API_KEY`.
|
|
@@ -86,6 +116,15 @@ bands don't overlap. The **effect floor** (0.05, one judge quantization step) is
|
|
|
86
116
|
minimum real move required before a change counts as more than noise β band
|
|
87
117
|
separation *plus* a floor-sized delta, never either alone.
|
|
88
118
|
|
|
119
|
+
**What the band does not cover.** A verdict rests on **one generation draw per
|
|
120
|
+
arm**: the band is the spread of the *judge* re-scoring that single response, not
|
|
121
|
+
the spread of the model writing a different one. Report #006 measured the second
|
|
122
|
+
directly and found it larger β draw-to-draw spread up to **sd 0.186** on the 0β1
|
|
123
|
+
scale, against judge-level noise several times smaller. So treat a surprising
|
|
124
|
+
single-run verdict as **provisional and worth re-running** before you act on it.
|
|
125
|
+
Generation sampling lands in the next receipt spec; until it does, this is a
|
|
126
|
+
limit of the instrument, stated rather than implied.
|
|
127
|
+
|
|
89
128
|
### Install
|
|
90
129
|
|
|
91
130
|
```bash
|
|
@@ -171,7 +210,7 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
171
210
|
"model_release_date": "2025-10-01",
|
|
172
211
|
"provider": "anthropic",
|
|
173
212
|
"surface": "claude-cli",
|
|
174
|
-
"runner_version": "0.
|
|
213
|
+
"runner_version": "0.6.0",
|
|
175
214
|
"date_utc": "2026-07-27Tβ¦Z",
|
|
176
215
|
"registry": "registered",
|
|
177
216
|
"transcripts": "hashes-only",
|
|
@@ -194,6 +233,12 @@ A receipt is the unit of evidence β one JSON document conforming to
|
|
|
194
233
|
},
|
|
195
234
|
"comparison": { "with_skill_score": 0.81, "baseline_score": 0.42,
|
|
196
235
|
"delta": 0.39, "delta_uncertainty": 0.036 },
|
|
236
|
+
// v0.4 economics, all derived and never composited into one score: "run.pricing_snapshot"
|
|
237
|
+
// freezes the rates; each case carries "usage" and a separate "judge_usage"; "economics"
|
|
238
|
+
// holds basis, surface, with_skill/baseline (call_count, mean_input_tokens,
|
|
239
|
+
// mean_output_tokens, mean_cost_usd_per_call, median_wall_ms + p25/p75/IQR),
|
|
240
|
+
// skill_incremental_cost_usd_per_call, skill_incremental_cost_usd_per_1k_calls,
|
|
241
|
+
// output_tokens_delta, median_wall_ms_delta, judge_excluded (const true), judge_overhead.
|
|
197
242
|
"verification_level": "TESTED",
|
|
198
243
|
"receipt_hash": "β¦sha256 of the canonical receipt with this field removedβ¦"
|
|
199
244
|
}
|
|
@@ -209,6 +254,12 @@ Key ideas:
|
|
|
209
254
|
- **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
|
|
210
255
|
- **`verification_level`** uses the community lattice: `UNVERIFIED` / `DECLARED` /
|
|
211
256
|
`TESTED` (Driftproof emits `TESTED`). `FORMAL` is reserved.
|
|
257
|
+
- **`economics`** is *derived*, never a second measurement: the token delta is the
|
|
258
|
+
durable fact, and the dollars are exactly those tokens at the rates frozen into
|
|
259
|
+
`run.pricing_snapshot`, so a receipt keeps its meaning after a vendor reprices.
|
|
260
|
+
Judge cost is recorded apart as `judge_usage` and excluded from every skill-value
|
|
261
|
+
figure (`judge_excluded` is `const true`) β measuring the skill is our cost, not
|
|
262
|
+
the skill's.
|
|
212
263
|
- **`receipt_hash`** is a self-hash for tamper-evidence (integrity, not yet a key
|
|
213
264
|
signature β see the spec's open questions).
|
|
214
265
|
|
|
@@ -233,7 +284,7 @@ jobs:
|
|
|
233
284
|
runs-on: ubuntu-latest
|
|
234
285
|
steps:
|
|
235
286
|
- uses: actions/checkout@v4
|
|
236
|
-
- uses: driftproofhq/driftproof@v0.
|
|
287
|
+
- uses: driftproofhq/driftproof@v0.6.0
|
|
237
288
|
with:
|
|
238
289
|
skill-dir: skills/my-skill
|
|
239
290
|
models: claude-haiku-4-5
|
|
@@ -272,11 +323,14 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
272
323
|
|
|
273
324
|
## Reports
|
|
274
325
|
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
the
|
|
279
|
-
(
|
|
326
|
+
Six reports are published, spanning five report types. A report page lives at a
|
|
327
|
+
draft path β `docs/reports/NNN-draft/` β until the publish sequence renames it, and
|
|
328
|
+
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
329
|
+
(see the roll at the top of this README, and
|
|
330
|
+
[REPORT-STYLE.md](REPORT-STYLE.md) for the shared chrome every report inherits).
|
|
331
|
+
Each report and every verdict in it are **re-derived from the receipts** committed
|
|
332
|
+
under [`receipts/`](receipts/)
|
|
333
|
+
(`receipts/report-001/` β¦ `receipts/report-006/`) β nothing is hand-entered.
|
|
280
334
|
|
|
281
335
|
Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
|
|
282
336
|
fetched at run time from a pinned commit and verified by sha256 against
|
|
@@ -290,9 +344,10 @@ node scripts/run-report-001.js --concurrency 5 # run both models Γ with/basel
|
|
|
290
344
|
node scripts/build-report-001.js # re-derive the report from the receipts
|
|
291
345
|
```
|
|
292
346
|
|
|
293
|
-
Reports #002β#
|
|
294
|
-
(`scripts/prepare-report-00N.js`
|
|
295
|
-
|
|
347
|
+
Reports #002β#005 have their own runners
|
|
348
|
+
(`scripts/prepare-report-00N.js` β Report #005's is
|
|
349
|
+
[`scripts/prepare-report-005.js`](scripts/prepare-report-005.js)) following the
|
|
350
|
+
same fetch β run β re-derive shape.
|
|
296
351
|
|
|
297
352
|
**Model-release triggers are live**: `scripts/release-watch.js` (keyless β it
|
|
298
353
|
reads the public models registry) notices a new model release, re-runs the
|
package/bin/driftproof
CHANGED
|
@@ -8,7 +8,7 @@ const { PROJECT_NAME, RUNNER_VERSION, DEFAULT_JUDGE_SAMPLES, DEV_MAX_USD } = req
|
|
|
8
8
|
const { loadSkill } = require('../lib/skill');
|
|
9
9
|
const { runSkillOnModel, summarizeReceipt, projectCalls } = require('../lib/run');
|
|
10
10
|
const { validateReceipt, verifyReceiptHash } = require('../lib/receipt');
|
|
11
|
-
const { buildDriftReport } = require('../lib/diff');
|
|
11
|
+
const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
|
|
12
12
|
const { surfaceForModel, isSubscriptionSurface, resolveModel } = require('../lib/provider');
|
|
13
13
|
const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
|
|
14
14
|
const { registryStatus } = require('../lib/models');
|
|
@@ -76,7 +76,7 @@ USAGE
|
|
|
76
76
|
${PROJECT_NAME} run <skill-dir> [--models a,b] [--samples N] [--max-cases N] [--max-calls N]
|
|
77
77
|
[--judge-model M] [--concurrency N] [--max-usd N]
|
|
78
78
|
[--keep-transcripts] [--out DIR]
|
|
79
|
-
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE]
|
|
79
|
+
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
|
|
80
80
|
${PROJECT_NAME} validate <receipt.json>
|
|
81
81
|
${PROJECT_NAME} badge <receipt.json> [--out FILE] [--github-output]
|
|
82
82
|
${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
|
|
@@ -250,9 +250,32 @@ function cmdDiff(positional, flags) {
|
|
|
250
250
|
if (!verifyReceiptHash(r)) console.error(` β ${path.basename(p)}: receipt_hash does not verify (tampered or hand-edited)`);
|
|
251
251
|
}
|
|
252
252
|
|
|
253
|
-
|
|
254
|
-
|
|
255
|
-
|
|
253
|
+
// --mode revision inverts the axis: the skill text is the variable under test
|
|
254
|
+
// and the substrate is the control. The fields release drift merely warns about
|
|
255
|
+
// are preconditions here, so a pair that is not a revision pair is REFUSED
|
|
256
|
+
// (exit 6) rather than rendered with a caveat nobody reads. A differing model
|
|
257
|
+
// would be release drift wearing a revision label β the one confound this mode
|
|
258
|
+
// exists to exclude β and an EQUAL content_hash has no revision to measure.
|
|
259
|
+
const mode = flags.mode || 'release';
|
|
260
|
+
if (mode !== 'release' && mode !== 'revision') {
|
|
261
|
+
console.error(`unknown --mode "${mode}" β supported: release, revision`);
|
|
262
|
+
process.exit(2);
|
|
263
|
+
}
|
|
264
|
+
if (mode === 'revision') {
|
|
265
|
+
const problem = revisionPairProblem(a, b);
|
|
266
|
+
if (problem) {
|
|
267
|
+
const why = problem === 'skill.content_hash'
|
|
268
|
+
? 'the two receipts carry the SAME skill.content_hash β there is no revision between them to measure'
|
|
269
|
+
: `${problem} differs between the two receipts β revision drift requires the substrate to be held fixed, and a differing ${problem} would confound the revision with release drift`;
|
|
270
|
+
console.error(` β REFUSED (--mode revision): ${why}.`);
|
|
271
|
+
console.error(` Compare these two with the default release mode, or supply a pair that differs only in skill.content_hash.`);
|
|
272
|
+
process.exit(6);
|
|
273
|
+
}
|
|
274
|
+
}
|
|
275
|
+
|
|
276
|
+
const labelA = mode === 'revision' ? `pinned (${dateStamp(a.run.date_utc)})` : dateStamp(a.run.date_utc);
|
|
277
|
+
const labelB = mode === 'revision' ? `current (${dateStamp(b.run.date_utc)})` : dateStamp(b.run.date_utc);
|
|
278
|
+
const { markdown } = buildDriftReport(a, b, { labelA, labelB, mode });
|
|
256
279
|
|
|
257
280
|
if (flags.out) {
|
|
258
281
|
fs.writeFileSync(path.resolve(flags.out), markdown);
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.6.0';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
package/lib/diff.js
CHANGED
|
@@ -3,6 +3,7 @@
|
|
|
3
3
|
|
|
4
4
|
const { bandVerdict, round } = require('./stats');
|
|
5
5
|
const { EFFECT_FLOOR } = require('../config');
|
|
6
|
+
const { revisionHeadline } = require('./revision');
|
|
6
7
|
|
|
7
8
|
// Practical-significance gate applied ON TOP of band separation. bandVerdict()
|
|
8
9
|
// stays a pure geometry test (kept that way so its unit checks are unambiguous);
|
|
@@ -68,7 +69,22 @@ function headlineVerdict(perCase) {
|
|
|
68
69
|
return 'WITHIN NOISE β no case moved beyond its confidence band; the skill holds up.';
|
|
69
70
|
}
|
|
70
71
|
|
|
71
|
-
|
|
72
|
+
// A revision pair is a pair in which the SKILL TEXT is the only thing that
|
|
73
|
+
// moved. `diff` was built for release drift, where the model varies and the text
|
|
74
|
+
// is fixed; this inverts it, so the fields that release drift merely warns about
|
|
75
|
+
// become the preconditions of the comparison. Returns the offending field name,
|
|
76
|
+
// or null when the pair is a valid revision pair.
|
|
77
|
+
function revisionPairProblem(a, b) {
|
|
78
|
+
if ((a.run.model_id || '') !== (b.run.model_id || '')) return 'run.model_id';
|
|
79
|
+
if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) return 'run.provider';
|
|
80
|
+
if ((a.run.surface || '') !== (b.run.surface || '')) return 'run.surface';
|
|
81
|
+
if ((a.suite.suite_hash || '') !== (b.suite.suite_hash || '')) return 'suite.suite_hash';
|
|
82
|
+
if ((a.skill.content_hash || '') === (b.skill.content_hash || '')) return 'skill.content_hash';
|
|
83
|
+
return null;
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
function buildDriftReport(a, b, { labelA = 'A', labelB = 'B', mode = 'release' } = {}) {
|
|
87
|
+
const revision = mode === 'revision';
|
|
72
88
|
const aB = withSkillBands(a);
|
|
73
89
|
const bB = withSkillBands(b);
|
|
74
90
|
const ids = [...new Set([...Object.keys(aB), ...Object.keys(bB)])];
|
|
@@ -99,6 +115,32 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
99
115
|
const regressions = perCase.filter((r) => r.verdict === 'regression');
|
|
100
116
|
|
|
101
117
|
const L = [];
|
|
118
|
+
if (revision) {
|
|
119
|
+
// The axis leads. In a release-drift report the model is what moved and it
|
|
120
|
+
// belongs at the top; here the model is the control and the skill's own text
|
|
121
|
+
// is the finding, so content_hash is the first row a reader meets.
|
|
122
|
+
L.push(`# Revision drift report`);
|
|
123
|
+
L.push('');
|
|
124
|
+
L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
|
|
125
|
+
L.push('');
|
|
126
|
+
L.push(`**Report type:** revision drift β the skill's own text is the variable under test; the substrate is held fixed.`);
|
|
127
|
+
L.push('');
|
|
128
|
+
L.push(`Left column is the **pinned** revision; right column is the **current** upstream revision.`);
|
|
129
|
+
L.push('');
|
|
130
|
+
L.push(`| | ${labelA} | ${labelB} |`);
|
|
131
|
+
L.push(`|---|---|---|`);
|
|
132
|
+
L.push(`| skill content_hash | \`${short(a.skill.content_hash)}\` | \`${short(b.skill.content_hash)}\` |`);
|
|
133
|
+
L.push(`| model (held) | \`${a.run.model_id}\` | \`${b.run.model_id}\` |`);
|
|
134
|
+
L.push(`| provider (held) | ${a.run.provider || 'anthropic'} | ${b.run.provider || 'anthropic'} |`);
|
|
135
|
+
L.push(`| surface (held) | ${a.run.surface} | ${b.run.surface} |`);
|
|
136
|
+
L.push(`| suite_hash (held) | \`${short(a.suite.suite_hash)}\` | \`${short(b.suite.suite_hash)}\` |`);
|
|
137
|
+
L.push(`| run date (UTC) | ${a.run.date_utc} | ${b.run.date_utc} |`);
|
|
138
|
+
L.push(`| judge samples/case | ${(a.run.judge || {}).samples || 1} | ${(b.run.judge || {}).samples || 1} |`);
|
|
139
|
+
L.push(`| with_skill (mean Β± band) | ${bandStr(aAgg)} | ${bandStr(bAgg)} |`);
|
|
140
|
+
L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
|
|
141
|
+
L.push(`| skill lift (Ξ) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
|
|
142
|
+
L.push('');
|
|
143
|
+
} else {
|
|
102
144
|
L.push(`# Drift report`);
|
|
103
145
|
L.push('');
|
|
104
146
|
L.push(`**Skill:** ${a.skill.name} \`${a.skill.version}\``);
|
|
@@ -115,10 +157,19 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
115
157
|
L.push(`| baseline score | ${pct(a.comparison.baseline_score)} | ${pct(b.comparison.baseline_score)} |`);
|
|
116
158
|
L.push(`| skill lift (Ξ) | ${fmt(a.comparison.delta)} | ${fmt(b.comparison.delta)} |`);
|
|
117
159
|
L.push('');
|
|
160
|
+
}
|
|
118
161
|
|
|
119
162
|
const warnings = [];
|
|
120
|
-
|
|
121
|
-
|
|
163
|
+
// In revision mode the changed skill text is the INDEPENDENT VARIABLE, not
|
|
164
|
+
// contamination, and the held-constant substrate is what makes the comparison
|
|
165
|
+
// valid. The release-drift caveat below says the opposite of both, so it is
|
|
166
|
+
// replaced rather than suppressed: a reader is told what is held, and why a
|
|
167
|
+
// separated band is attributable to the revision.
|
|
168
|
+
if (revision) {
|
|
169
|
+
warnings.push(`model \`${a.run.model_id}\`, provider ${a.run.provider || 'anthropic'}, surface ${a.run.surface} and suite_hash \`${short(a.suite.suite_hash)}\` are held constant across both receipts β the skill text is the only variable under test, so a band-separated move above the ${EFFECT_FLOOR} floor is attributable to the revision.`);
|
|
170
|
+
}
|
|
171
|
+
if (!revision && a.skill.content_hash !== b.skill.content_hash) warnings.push('skill content_hash differs β the skill itself changed between receipts, so drift mixes skill edits with model drift.');
|
|
172
|
+
if (!revision && a.suite.suite_hash !== b.suite.suite_hash) warnings.push('suite_hash differs β the eval suite changed; per-case comparison may be misleading.');
|
|
122
173
|
if (a.skill.name !== b.skill.name) warnings.push(`different skills (${a.skill.name} vs ${b.skill.name}) β comparison is not meaningful.`);
|
|
123
174
|
if ((a.run.judge || {}).samples <= 1 || (b.run.judge || {}).samples <= 1) warnings.push('one or both receipts are single-sample (no bands) β non-overlap can only be trusted when both sides are sampled.');
|
|
124
175
|
if (!measured) warnings.push(`verdicts NOT computed β ${belowTested.map(([l, r]) => `${l} is ${levelOf(r)}${r.run && r.run.source ? ` (${r.run.source})` : ''}`).join('; ')}. Drift verdicts require TESTED receipts on both sides; declared numbers are shown as context only (see /interop.html).`);
|
|
@@ -126,8 +177,8 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
126
177
|
// providers is a skill-DURABILITY comparison across substrates, not model drift
|
|
127
178
|
// over time; across surfaces, sampling control differs. Both are flagged so a
|
|
128
179
|
// reader never mistakes one for the other (see docs/neutrality.html).
|
|
129
|
-
if ((a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
|
|
130
|
-
if (a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β sampling control differs between surfaces; compare with care.`);
|
|
180
|
+
if (!revision && (a.run.provider || 'anthropic') !== (b.run.provider || 'anthropic')) warnings.push(`different providers (${a.run.provider || 'anthropic'} vs ${b.run.provider || 'anthropic'}) β this is a cross-substrate durability comparison, not model drift over time; read the delta, not absolute scores (see the neutrality policy).`);
|
|
181
|
+
if (!revision && a.run.surface !== b.run.surface) warnings.push(`different surfaces (${a.run.surface} vs ${b.run.surface}) β sampling control differs between surfaces; compare with care.`);
|
|
131
182
|
if (warnings.length) {
|
|
132
183
|
L.push('> **β Caveats**');
|
|
133
184
|
for (const w of warnings) L.push(`> - ${w}`);
|
|
@@ -142,7 +193,7 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
142
193
|
L.push(`**NOT MEASURED β ${belowTested.length} receipt(s) below TESTED. Drift verdicts require TESTED receipts on both sides; the declared numbers above are context, not band-verified evidence.**`);
|
|
143
194
|
L.push('');
|
|
144
195
|
} else {
|
|
145
|
-
L.push(`**${headlineVerdict(perCase)}**`);
|
|
196
|
+
L.push(`**${revision ? revisionHeadline(perCase) : headlineVerdict(perCase)}**`);
|
|
146
197
|
L.push('');
|
|
147
198
|
L.push(`with_skill mean moved ${fmt(headlineDelta)} (${bandStr(aAgg)} β ${bandStr(bAgg)}; band = suite dispersion). Per-case band-overlap verdicts: ${regressions.length} regression(s), ${perCase.filter((r) => r.verdict === 'improvement').length} improvement(s), ${nWithin} within noise${nFloor ? ` (${nFloor} of them band-separated but below the ${EFFECT_FLOOR} effect floor)` : ''}.`);
|
|
148
199
|
L.push('');
|
|
@@ -170,4 +221,4 @@ function buildDriftReport(a, b, { labelA = 'A', labelB = 'B' } = {}) {
|
|
|
170
221
|
return { markdown: L.join('\n'), perCase, headlineDelta, regressions };
|
|
171
222
|
}
|
|
172
223
|
|
|
173
|
-
module.exports = { buildDriftReport, withSkillBands };
|
|
224
|
+
module.exports = { buildDriftReport, withSkillBands, revisionPairProblem };
|
package/lib/hygiene.js
ADDED
|
@@ -0,0 +1,113 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
//
|
|
3
|
+
// The single definition of what counts as a leak.
|
|
4
|
+
//
|
|
5
|
+
// Two callers scan for the same classes of secret: `tests/gate.js`, which walks
|
|
6
|
+
// the working tree at gate time, and `scripts/merge-check.js`, which walks the
|
|
7
|
+
// candidate tree read out of the commit before permitting a merge. They must
|
|
8
|
+
// agree. Two copies of a deny-list drift, and the copy that quietly stops
|
|
9
|
+
// matching still reads as protection β the same argument the acknowledgment
|
|
10
|
+
// deny-list bite test already makes, applied to the scanner itself.
|
|
11
|
+
//
|
|
12
|
+
// Backlog B-8 records four approval records that reached `dev` carrying a host
|
|
13
|
+
// path. The recorded cause β "the scan walks tracked files" β is false, and this
|
|
14
|
+
// module's existence does not depend on it: `listFiles()` in the repo gate has
|
|
15
|
+
// always enumerated untracked files too, and a planted path in an untracked
|
|
16
|
+
// evidence record does take the repo gate red. The real window is VANTAGE. An
|
|
17
|
+
// approval session measures in a disposable copy made before it writes its
|
|
18
|
+
// record into the shared checkout, so the record is simply not in the tree that
|
|
19
|
+
// was scanned. Scanning the working tree harder cannot close that. Scanning the
|
|
20
|
+
// commit that is about to merge does, and that is what merge-check now uses this
|
|
21
|
+
// for.
|
|
22
|
+
//
|
|
23
|
+
// Patterns are written so a regex literal cannot match its own text: a
|
|
24
|
+
// metacharacter or a character class follows each fixed prefix.
|
|
25
|
+
|
|
26
|
+
const dec = (b64) => JSON.parse(Buffer.from(b64, 'base64').toString('utf8'));
|
|
27
|
+
|
|
28
|
+
// Build-box host + private IP, base64 so the plaintext never lands in a file.
|
|
29
|
+
const HOSTIP = dec('WyJpcC0xNzItMzEtNDAtMTU5LmFwLXNvdXRoZWFzdC0xLmNvbXB1dGUuaW50ZXJuYWwiLCIxNzIuMzEuNDAuMTU5Il0=');
|
|
30
|
+
|
|
31
|
+
const EMAIL_ALLOW = new Set(['example.com', 'example.org', 'driftproofhq.com']);
|
|
32
|
+
|
|
33
|
+
const PATTERNS = [
|
|
34
|
+
{ name: 'home-path', re: /\/home\/[a-z0-9_-]+\/|\/Users\/[A-Za-z0-9_-]+\// },
|
|
35
|
+
{ name: 'ec2-internal-host', re: /ip-\d+-\d+-\d+-\d+\.[a-z0-9.-]*compute\.(internal|amazonaws\.com)/i },
|
|
36
|
+
{ name: 'private-ip', re: /\b(10\.\d{1,3}\.\d{1,3}\.\d{1,3}|172\.(1[6-9]|2\d|3[01])\.\d{1,3}\.\d{1,3}|192\.168\.\d{1,3}\.\d{1,3})\b/ },
|
|
37
|
+
{ name: 'anthropic-key', re: /sk-ant-[A-Za-z0-9_-]{8,}/ },
|
|
38
|
+
{ name: 'github-token', re: /gh[pousr]_[A-Za-z0-9]{20,}/ },
|
|
39
|
+
{ name: 'aws-key', re: /\bAKIA[0-9A-Z]{16}\b/ },
|
|
40
|
+
{ name: 'private-key-block', re: /-----BEGIN [A-Z ]*PRIVATE KEY-----/ },
|
|
41
|
+
{ name: 'env-secret-assignment', re: /\b(ANTHROPIC_API_KEY|AWS_SECRET_ACCESS_KEY|OPENAI_API_KEY)\s*=\s*\S+/ },
|
|
42
|
+
// Codex subscription auth material (~/.codex/auth.json contents) must NEVER
|
|
43
|
+
// land in a committed file: the id_token/access_token are JWTs, and an OpenAI
|
|
44
|
+
// secret key is sk-proj-/sk-svcacct-/sk-admin-. We ban the CONTENTS (tokens),
|
|
45
|
+
// not the documented path string (`~/.codex/auth.json` is referenced in help
|
|
46
|
+
// text and docs by design).
|
|
47
|
+
{ name: 'jwt-token', re: /\beyJ[A-Za-z0-9_=-]{10,}\.eyJ[A-Za-z0-9_=-]{10,}\.[A-Za-z0-9_=-]{6,}/ },
|
|
48
|
+
{ name: 'openai-secret-key', re: /\bsk-(proj|svcacct|admin)-[A-Za-z0-9_-]{20,}/ },
|
|
49
|
+
];
|
|
50
|
+
|
|
51
|
+
// A path is a hit on its own name, with no content read: a committed .env file
|
|
52
|
+
// is a leak whatever it happens to contain.
|
|
53
|
+
function scanPath(rel) {
|
|
54
|
+
return /(^|\/)\.env(\.|$)/.test(rel) ? [{ file: rel, kind: 'env-file' }] : [];
|
|
55
|
+
}
|
|
56
|
+
|
|
57
|
+
function scanContent(rel, content) {
|
|
58
|
+
const hits = [];
|
|
59
|
+
for (const p of PATTERNS) {
|
|
60
|
+
const m = content.match(p.re);
|
|
61
|
+
if (m) hits.push({ file: rel, kind: p.name, sample: m[0].slice(0, 24) });
|
|
62
|
+
}
|
|
63
|
+
for (const hv of HOSTIP) if (content.includes(hv)) hits.push({ file: rel, kind: 'build-host-or-ip' });
|
|
64
|
+
const emails = content.match(/\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b/g) || [];
|
|
65
|
+
for (const e of emails) {
|
|
66
|
+
const dom = e.split('@')[1].toLowerCase();
|
|
67
|
+
if (!EMAIL_ALLOW.has(dom) && !dom.endsWith('.example')) hits.push({ file: rel, kind: 'email', sample: e });
|
|
68
|
+
}
|
|
69
|
+
return hits;
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
// scanFiles(files, read, readLink) β `read(rel)` returns the file's text, or
|
|
73
|
+
// throws/returns null for anything unreadable (binary, deleted, a symlink to
|
|
74
|
+
// nowhere), which is skipped exactly as the gate has always skipped it.
|
|
75
|
+
//
|
|
76
|
+
// `readLink(rel)` is optional and returns a symbolic link's TARGET PATH as a
|
|
77
|
+
// string, or null for an entry that is not a link. Where it is supplied, that
|
|
78
|
+
// target is scanned as the entry's content through `scanContent`, so every
|
|
79
|
+
// pattern above applies to it with no second deny-list to drift (spec 011,
|
|
80
|
+
// AC-2). A link's target is a path, and a path is exactly the class of leak the
|
|
81
|
+
// `home-path` pattern exists to catch: P-1 shipped `node_modules ->
|
|
82
|
+
// /home/<user>/... ` into a public tree past a scan that read the link as an
|
|
83
|
+
// unreadable file and skipped it.
|
|
84
|
+
//
|
|
85
|
+
// The argument is optional so a caller supplying nothing behaves exactly as
|
|
86
|
+
// before. That compatibility is deliberate, and it is also the standing risk: a
|
|
87
|
+
// caller added later is blind by default. The repo gate asserts that every call
|
|
88
|
+
// site supplies a reader, with `scripts/merge-check.js` the one exception β its
|
|
89
|
+
// reader is `git show <ref>:<path>`, and git stores a link's blob as its target
|
|
90
|
+
// string, so the target already arrives as content there.
|
|
91
|
+
function scanFiles(files, read, readLink) {
|
|
92
|
+
const hits = [];
|
|
93
|
+
for (const rel of files) {
|
|
94
|
+
hits.push(...scanPath(rel));
|
|
95
|
+
// The link branch is a single condition on purpose: it is the mutation seam
|
|
96
|
+
// the spec-011 gate neutralises to prove the scanner goes blind without it.
|
|
97
|
+
if (typeof readLink === 'function') {
|
|
98
|
+
let target;
|
|
99
|
+
try { target = readLink(rel); } catch (_e) { target = null; }
|
|
100
|
+
if (typeof target === 'string') {
|
|
101
|
+
hits.push(...scanContent(rel, target));
|
|
102
|
+
continue;
|
|
103
|
+
}
|
|
104
|
+
}
|
|
105
|
+
let c;
|
|
106
|
+
try { c = read(rel); } catch (_e) { continue; }
|
|
107
|
+
if (typeof c !== 'string') continue;
|
|
108
|
+
hits.push(...scanContent(rel, c));
|
|
109
|
+
}
|
|
110
|
+
return hits;
|
|
111
|
+
}
|
|
112
|
+
|
|
113
|
+
module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, scanPath, scanContent, scanFiles };
|
package/lib/revision.js
ADDED
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
'use strict';
|
|
3
|
+
|
|
4
|
+
const { bandVerdict, round } = require('./stats');
|
|
5
|
+
const { EFFECT_FLOOR } = require('../config');
|
|
6
|
+
|
|
7
|
+
// Revision drift β the fifth report type (spec 009, Report #006).
|
|
8
|
+
//
|
|
9
|
+
// Every other report type holds the skill text fixed and moves something
|
|
10
|
+
// underneath it: the model release (#001, #003), the vendor surface (#002), the
|
|
11
|
+
// capability tier (#004), the axes and the price (#005). This one inverts the
|
|
12
|
+
// design. The substrate is held still β same model, same provider, same surface,
|
|
13
|
+
// same suite, same fixed judge, same sampling β and the SKILL'S OWN TEXT moves,
|
|
14
|
+
// from the revision a report pinned to the revision upstream ships today.
|
|
15
|
+
//
|
|
16
|
+
// This module holds the report type's LANGUAGE as code rather than as
|
|
17
|
+
// hand-written page copy. That is deliberate: the fairness rule below is the one
|
|
18
|
+
// a measurement project is most tempted to apply in one direction only, and a
|
|
19
|
+
// rule that lives in prose cannot be gated before the run that would tempt it.
|
|
20
|
+
|
|
21
|
+
// ββ the cell headline ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
|
22
|
+
// A summary of the per-case band-overlap verdicts, worded about the REVISION.
|
|
23
|
+
// The release-drift headline says "the skill is measurably weaker", which is a
|
|
24
|
+
// sentence about a skill under a moving model. Here the model is the control.
|
|
25
|
+
function revisionHeadline(perCase) {
|
|
26
|
+
const reg = perCase.filter((r) => r.verdict === 'regression').length;
|
|
27
|
+
const imp = perCase.filter((r) => r.verdict === 'improvement').length;
|
|
28
|
+
const s = (n) => (n === 1 ? '' : 's');
|
|
29
|
+
if (reg && imp) {
|
|
30
|
+
return `MIXED β the revision improved ${imp} case${s(imp)} and regressed ${reg} on non-overlapping bands.`;
|
|
31
|
+
}
|
|
32
|
+
if (reg) {
|
|
33
|
+
return `REVISION REGRESSED β ${reg} case${s(reg)} scored lower under the current upstream text (bands do not overlap).`;
|
|
34
|
+
}
|
|
35
|
+
if (imp) {
|
|
36
|
+
return `REVISION IMPROVED β ${imp} case${s(imp)} scored higher under the current upstream text (bands do not overlap); none regressed.`;
|
|
37
|
+
}
|
|
38
|
+
return 'WITHIN NOISE β the revision moved no case beyond its confidence band; the pinned text and the current text measure the same.';
|
|
39
|
+
}
|
|
40
|
+
|
|
41
|
+
// Classification word for a cell, from its headline. Kept separate so a caller
|
|
42
|
+
// can branch on the class without parsing prose.
|
|
43
|
+
function revisionClass(perCase) {
|
|
44
|
+
const h = revisionHeadline(perCase);
|
|
45
|
+
return h.split(' β')[0];
|
|
46
|
+
}
|
|
47
|
+
|
|
48
|
+
// ββ the fairness sentence ββββββββββββββββββββββββββββββββββββββββββββββββββββ
|
|
49
|
+
// Spec 009 Β§ Fairness, clauses 1 and 4. Where a revision IMPROVED a skill, the
|
|
50
|
+
// published #005 figure understates the pack a reader can install today; where it
|
|
51
|
+
// REGRESSED one, #005 overstates it. Both sentences are generated by the same
|
|
52
|
+
// function, from the same template, so the disclosure cannot quietly become a
|
|
53
|
+
// one-directional courtesy β the symmetry is a property of the code, and the
|
|
54
|
+
// gate asserts it.
|
|
55
|
+
//
|
|
56
|
+
// A cell within noise gets NO sentence. #005's figure stands unamended, because
|
|
57
|
+
// nothing was measured that would amend it, and manufacturing a hedge for a null
|
|
58
|
+
// result is how a report launders noise into a finding.
|
|
59
|
+
function fairnessSentence({ slug, classification, report005Delta, measuredDelta }) {
|
|
60
|
+
const cls = String(classification || '');
|
|
61
|
+
if (cls !== 'REVISION IMPROVED' && cls !== 'REVISION REGRESSED') return null;
|
|
62
|
+
const improved = cls === 'REVISION IMPROVED';
|
|
63
|
+
const direction = improved ? 'understates' : 'overstates';
|
|
64
|
+
const d = (n) => (n == null ? 'n/a' : (n >= 0 ? '+' : '') + Number(n).toFixed(3));
|
|
65
|
+
return `Report #005 measured ${slug} at ${d(report005Delta)} on the text it had pinned. `
|
|
66
|
+
+ `Report #006 measures the current upstream revision at ${d(measuredDelta)} on the same substrate and the same suite. `
|
|
67
|
+
+ `#005's published figure therefore ${direction} the pack upstream ships today for this skill, `
|
|
68
|
+
+ `and is amended by this report rather than corrected in place.`;
|
|
69
|
+
}
|
|
70
|
+
|
|
71
|
+
// ββ per-cell scoping disclosures βββββββββββββββββββββββββββββββββββββββββββββ
|
|
72
|
+
// Spec 009 AC-10. One cell in Report #006 measures something other than what its
|
|
73
|
+
// upstream author changed it to do, and the reader looking at that row is the
|
|
74
|
+
// reader who needs to be told.
|
|
75
|
+
const SCOPING_NOTES = {
|
|
76
|
+
'git-workflow-and-versioning':
|
|
77
|
+
'This revision changes the frontmatter `description:` line. In a skill runtime a description is a '
|
|
78
|
+
+ 'routing trigger: it decides whether the skill loads, and never reaches the model as guidance. '
|
|
79
|
+
+ 'Driftproof makes no routing decision β it always injects the skill, and passes the whole file, '
|
|
80
|
+
+ 'frontmatter included, as the system prompt. This cell therefore measures the revision as added '
|
|
81
|
+
+ 'context and cannot measure it as a trigger.',
|
|
82
|
+
};
|
|
83
|
+
function scopingNote(slug) {
|
|
84
|
+
return SCOPING_NOTES[slug] || null;
|
|
85
|
+
}
|
|
86
|
+
|
|
87
|
+
// ββ the baseline-reproduction control ββββββββββββββββββββββββββββββββββββββββ
|
|
88
|
+
// Spec 009 AC-6, and the thing that makes the free pinned arm honest.
|
|
89
|
+
//
|
|
90
|
+
// Report #006 reuses #005's receipts as the pinned-text arm. That is valid only
|
|
91
|
+
// if the substrate has not moved, and `run.model_release_date` is null on every
|
|
92
|
+
// #005 receipt, so id equality is the only version evidence a receipt carries. A
|
|
93
|
+
// provider that re-points a concrete id at a new snapshot is invisible to it.
|
|
94
|
+
//
|
|
95
|
+
// It does not have to be. Every fresh run emits a BASELINE arm: the same cases,
|
|
96
|
+
// the same substrate, and no skill text at all. The revision cannot touch it by
|
|
97
|
+
// construction, so comparing the fresh baseline against the reused receipt's
|
|
98
|
+
// baseline re-measures exactly the thing id equality could not prove β at no
|
|
99
|
+
// extra cost, because that arm is already paid for.
|
|
100
|
+
//
|
|
101
|
+
// A cell whose baselines do not reproduce is NOT MEASURED. The reuse is a tested
|
|
102
|
+
// prediction, not an assumption the report asks the reader to grant.
|
|
103
|
+
function baselineBands(receipt) {
|
|
104
|
+
const out = {};
|
|
105
|
+
for (const c of receipt.results.cases) {
|
|
106
|
+
if (c.mode !== 'baseline') continue;
|
|
107
|
+
out[c.id] = { mean: c.mean != null ? c.mean : c.score, stddev: c.stddev || 0 };
|
|
108
|
+
}
|
|
109
|
+
return out;
|
|
110
|
+
}
|
|
111
|
+
|
|
112
|
+
function baselineControl(reused, fresh) {
|
|
113
|
+
const A = baselineBands(reused);
|
|
114
|
+
const B = baselineBands(fresh);
|
|
115
|
+
const ids = [...new Set([...Object.keys(A), ...Object.keys(B)])];
|
|
116
|
+
|
|
117
|
+
const perCase = ids.map((id) => {
|
|
118
|
+
const before = A[id] || null;
|
|
119
|
+
const after = B[id] || null;
|
|
120
|
+
if (!before || !after) return { id, before, after, delta: null, moved: false, missing: true };
|
|
121
|
+
const delta = round(after.mean - before.mean);
|
|
122
|
+
// The same rule the study uses everywhere else: band separation AND the
|
|
123
|
+
// effect floor. A baseline that wobbles inside its band has not moved.
|
|
124
|
+
const raw = bandVerdict(before.mean, before.stddev, after.mean, after.stddev);
|
|
125
|
+
const separated = raw === 'regression' || raw === 'improvement';
|
|
126
|
+
return { id, before, after, delta, moved: separated && Math.abs(delta) >= EFFECT_FLOOR, missing: false };
|
|
127
|
+
});
|
|
128
|
+
|
|
129
|
+
const movedCases = perCase.filter((r) => r.moved);
|
|
130
|
+
const missing = perCase.filter((r) => r.missing);
|
|
131
|
+
const reproduced = movedCases.length === 0 && missing.length === 0;
|
|
132
|
+
const aggDelta = round(
|
|
133
|
+
(fresh.comparison && fresh.comparison.baseline_score != null ? fresh.comparison.baseline_score : 0)
|
|
134
|
+
- (reused.comparison && reused.comparison.baseline_score != null ? reused.comparison.baseline_score : 0),
|
|
135
|
+
);
|
|
136
|
+
|
|
137
|
+
return {
|
|
138
|
+
reproduced,
|
|
139
|
+
blocked: !reproduced,
|
|
140
|
+
verdict: reproduced ? 'MEASURED' : 'NOT MEASURED',
|
|
141
|
+
moved_cases: movedCases.map((r) => r.id),
|
|
142
|
+
missing_cases: missing.map((r) => r.id),
|
|
143
|
+
aggregate_baseline_delta: aggDelta,
|
|
144
|
+
floor: EFFECT_FLOOR,
|
|
145
|
+
// THE CONTROL PROVES NON-REPRODUCTION. IT CANNOT SAY WHY. These strings used
|
|
146
|
+
// to read 'the substrate moved' and 'the substrate held still' β a cause,
|
|
147
|
+
// asserted by a comparison that measures two baseline arms and nothing else.
|
|
148
|
+
// A 120-call stability probe then found generation-level sampling noise large
|
|
149
|
+
// enough to account for every gap this control saw, with no substrate movement
|
|
150
|
+
// required, and the report page retracted the claim while three committed
|
|
151
|
+
// control records still carried it (approval finding F-009-N). Reason strings
|
|
152
|
+
// only: no score, sample, hash or verdict changed with this edit.
|
|
153
|
+
reason: reproduced
|
|
154
|
+
? 'the fresh baseline reproduces the reused receipt\'s baseline within the band and the floor, so the pinned-arm reuse stands for this cell'
|
|
155
|
+
: `the fresh baseline does not reproduce the reused receipt's baseline (${movedCases.length} case(s) moved beyond the band and the ${EFFECT_FLOOR} floor${missing.length ? `, ${missing.length} case(s) absent on one side` : ''}) β the reused pinned arm is not comparable to the fresh arm, so revision drift cannot be separated from whatever else changed in this cell; the control establishes non-reproduction and does not identify a cause`,
|
|
156
|
+
perCase,
|
|
157
|
+
};
|
|
158
|
+
}
|
|
159
|
+
|
|
160
|
+
module.exports = {
|
|
161
|
+
revisionHeadline, revisionClass, fairnessSentence, scopingNote, baselineControl,
|
|
162
|
+
};
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit
|
|
3
|
+
"version": "0.6.0",
|
|
4
|
+
"description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
|
7
7
|
"agent-skills",
|
package/spec/RECEIPT.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
2
|
# Driftproof receipt β spec v0.4
|
|
3
3
|
|
|
4
|
-
A **receipt** is a
|
|
4
|
+
A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
|
|
5
5
|
**with** and **without** the skill on one model version, with the judge **sampled**
|
|
6
6
|
so every score carries a confidence band. Receipts are the unit of evidence
|
|
7
7
|
Driftproof produces; diffing two receipts across model releases yields a **drift
|
|
@@ -164,7 +164,7 @@ The v0.3.1 schema gained an **additive interop revision** so receipts can be
|
|
|
164
164
|
|
|
165
165
|
## Fields
|
|
166
166
|
|
|
167
|
-
### `schema_version` (string, required) β `"0.
|
|
167
|
+
### `schema_version` (string, required) β `"0.4"`.
|
|
168
168
|
|
|
169
169
|
### `skill` (object, required)
|
|
170
170
|
| field | type | notes |
|
package/spec/receipt.schema.json
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
|
3
3
|
"$id": "https://driftproofhq.com/spec/receipt.schema.json",
|
|
4
4
|
"title": "driftproof receipt",
|
|
5
|
-
"description": "A
|
|
5
|
+
"description": "A hash-verified, dated record of running one agent skill's eval suite with and without the skill on one model version, with sampled judge scores and confidence bands. Receipt spec v0.4 β additive over v0.3.1: adds per-case, per-arm generation `usage` (input/output/cached tokens + measured wall_ms) captured from the surfaces that report it, a separate per-case `judge_usage` (measurement overhead, EXCLUDED from every skill-value figure by construction β `economics.judge_excluded` is const true), a run-level `run.pricing_snapshot` freezing the registry prices the derived dollar figures were computed from (so a receipt keeps its meaning when prices later change), and a derived `economics` block (per-arm mean cost/call, skill incremental cost per call and per 1k calls, output-length delta, median wall_ms with IQR). The three value axes β accuracy lift, cost, latency β are recorded separately and NEVER combined into a composite score. All v0.3.1 semantics are unchanged and every prior receipt still validates against its own frozen schema (v0.1, v0.2, v0.3, v0.3.1). The TESTED tightening (see allOf) is unchanged: the interop relaxations remain available only below TESTED.",
|
|
6
6
|
"type": "object",
|
|
7
7
|
"additionalProperties": false,
|
|
8
8
|
"required": [
|