driftproof 0.11.0 → 0.11.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -11
- package/config.js +1 -1
- package/lib/hygiene.js +67 -2
- package/package.json +2 -2
- package/spec/RECEIPT.md +1 -1
package/README.md
CHANGED
|
@@ -13,17 +13,19 @@ there were too few draws to tell, the receipt says so.
|
|
|
13
13
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
14
14
|
|
|
15
15
|
**[Quickstart](https://driftproofhq.com/#quickstart)** ·
|
|
16
|
-
**[Latest report](https://driftproofhq.com/reports/
|
|
16
|
+
**[Latest report](https://driftproofhq.com/reports/010/)** ·
|
|
17
17
|
[driftproofhq.com](https://driftproofhq.com)
|
|
18
18
|
|
|
19
19
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
20
20
|
format; it does not invent its own.
|
|
21
21
|
|
|
22
|
-
📊 **
|
|
23
|
-
hand-entered
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
22
|
+
📊 **Ten published reports** (each re-derived from committed files, nothing
|
|
23
|
+
hand-entered: Driftproof's receipts, or for Report #010 the upstream harness's own
|
|
24
|
+
output), spanning seven published report types, the newest being instrument
|
|
25
|
+
comparison. The nine that measure with Driftproof read its own arms by one
|
|
26
|
+
band-based, floor-gated verdict rule and differ in what moves underneath the
|
|
27
|
+
skill — or, in the value report, in which axes are measured; or, in the instrument re-measurement, in the
|
|
28
|
+
instrument itself; or, in the instrument comparison, in which instrument measures:
|
|
27
29
|
|
|
28
30
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
29
31
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
|
|
@@ -76,11 +78,25 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
|
|
|
76
78
|
receipts, which is what makes the delta attributable to the model rather than
|
|
77
79
|
to the instrument. One case sits inside the verdict on the effect floor alone
|
|
78
80
|
and the report names it.
|
|
81
|
+
- **[Report #009](https://driftproofhq.com/reports/009/)** — *instrument
|
|
82
|
+
comparison*: three skills from one plugin, one case each, measured by Claude
|
|
83
|
+
Code's native plugin eval and by Driftproof on the same SKILL.md bytes, task
|
|
84
|
+
prompts and rubrics. The two tools apply different treatments and grade
|
|
85
|
+
differently, so the report reads them side by side, ranks neither, and states
|
|
86
|
+
what each can and cannot establish. Both tools judged with `claude-opus-5`,
|
|
87
|
+
which departs from the judge policy, so its figures are not comparable with
|
|
88
|
+
Reports #001 to #008.
|
|
89
|
+
- **[Report #010](https://driftproofhq.com/reports/010/)** — *instrument
|
|
90
|
+
comparison*: one skill's own behavioural eval from its upstream repository, run
|
|
91
|
+
five times under each of two Claude Code configurations at one commit: 8 passed
|
|
92
|
+
and 2 failed, both failures on one expectation the skill's ADR instructions do
|
|
93
|
+
not state. No Driftproof measurement was taken and no Driftproof judge ran, and
|
|
94
|
+
the report does not isolate what caused the variation.
|
|
79
95
|
|
|
80
96
|
✍️ The launch essay, **[Three model releases later: what actually happens to agent
|
|
81
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads all
|
|
97
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all ten reports
|
|
82
98
|
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
83
|
-
corrected instrument did to three published results. Revised 2026-09-
|
|
99
|
+
corrected instrument did to three published results. Revised 2026-09-22; every
|
|
84
100
|
figure in it is gate-checked against the report page it cites.
|
|
85
101
|
|
|
86
102
|
## Why
|
|
@@ -305,7 +321,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
305
321
|
"model_release_date": "2025-10-01",
|
|
306
322
|
"provider": "anthropic",
|
|
307
323
|
"surface": "claude-cli",
|
|
308
|
-
"runner_version": "0.11.
|
|
324
|
+
"runner_version": "0.11.2",
|
|
309
325
|
"date_utc": "2026-07-27T…Z",
|
|
310
326
|
"registry": "registered",
|
|
311
327
|
"transcripts": "hashes-only",
|
|
@@ -384,7 +400,7 @@ jobs:
|
|
|
384
400
|
runs-on: ubuntu-latest
|
|
385
401
|
steps:
|
|
386
402
|
- uses: actions/checkout@v4
|
|
387
|
-
- uses: driftproofhq/driftproof@v0.11.
|
|
403
|
+
- uses: driftproofhq/driftproof@v0.11.2
|
|
388
404
|
with:
|
|
389
405
|
skill-dir: skills/my-skill
|
|
390
406
|
models: claude-haiku-4-5
|
|
@@ -479,7 +495,7 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
479
495
|
|
|
480
496
|
## Reports
|
|
481
497
|
|
|
482
|
-
|
|
498
|
+
Ten reports are published, spanning seven report types. A report page lives at a
|
|
483
499
|
draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
|
|
484
500
|
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
485
501
|
(see the roll at the top of this README, and
|
|
@@ -487,6 +503,8 @@ draft path — `docs/reports/NNN-draft/` — until the publish sequence renames
|
|
|
487
503
|
Each report and every verdict in it are **re-derived from the receipts** committed
|
|
488
504
|
under [`receipts/`](receipts/)
|
|
489
505
|
(`receipts/report-001/` … `receipts/report-006/`) — nothing is hand-entered.
|
|
506
|
+
Report #010 takes no Driftproof measurement and has no receipts: it is re-derived
|
|
507
|
+
from the upstream harness's output files, published beside its page.
|
|
490
508
|
|
|
491
509
|
Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
|
|
492
510
|
fetched at run time from a pinned commit and verified by sha256 against
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.11.
|
|
12
|
+
const RUNNER_VERSION = '0.11.2';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
package/lib/hygiene.js
CHANGED
|
@@ -23,6 +23,8 @@
|
|
|
23
23
|
// Patterns are written so a regex literal cannot match its own text: a
|
|
24
24
|
// metacharacter or a character class follows each fixed prefix.
|
|
25
25
|
|
|
26
|
+
const zlib = require('zlib');
|
|
27
|
+
|
|
26
28
|
const dec = (b64) => JSON.parse(Buffer.from(b64, 'base64').toString('utf8'));
|
|
27
29
|
|
|
28
30
|
// Build-box host + private IP, base64 so the plaintext never lands in a file.
|
|
@@ -54,7 +56,70 @@ function scanPath(rel) {
|
|
|
54
56
|
return /(^|\/)\.env(\.|$)/.test(rel) ? [{ file: rel, kind: 'env-file' }] : [];
|
|
55
57
|
}
|
|
56
58
|
|
|
59
|
+
// A PNG'S TEXT-BEARING REGIONS (spec 038 A-038-5, the classification spec 027 already made for the
|
|
60
|
+
// repository gate's confidentiality scan, applied here).
|
|
61
|
+
//
|
|
62
|
+
// THE THREAT MODEL IS ACCIDENTAL PLAINTEXT DISCLOSURE by a person or an agent writing the tree
|
|
63
|
+
// (CONSTITUTION § Threat model). A PNG carries text in its tEXt, iTXt and zTXt chunks, and carries
|
|
64
|
+
// pixels in IDAT, which is deflate output: nobody types into it, and a run of bytes that happens to
|
|
65
|
+
// decode as an address there is a coincidence of compression, not a disclosure. Scanned as one
|
|
66
|
+
// string, a 2 MB screenshot is two million chances of that coincidence, and it has fired twice on
|
|
67
|
+
// this repository's own committed captures.
|
|
68
|
+
//
|
|
69
|
+
// So: every chunk but IDAT is scanned, chunk type names included, and IDAT is not. A name RENDERED
|
|
70
|
+
// INTO a screenshot is pixels, which no byte scan sees either way - that is the gap this scan has
|
|
71
|
+
// always had, and it is unchanged. Everything that is not a PNG is scanned whole, as before.
|
|
72
|
+
//
|
|
73
|
+
// FAIL CLOSED WHERE THE FILE STOPS BEING A PNG (spec 038 approval F-3). Bytes after IEND are no chunk
|
|
74
|
+
// at all, and a chunk whose length runs past the end of the file is a truncated file: either way the
|
|
75
|
+
// remainder, from where the walk lost the structure, is scanned whole, as every file was before this
|
|
76
|
+
// classification. And a text chunk is read in the encoding it carries text in: a zTXt chunk's text
|
|
77
|
+
// and a compressed iTXt chunk's text are inflated and scanned as well as their raw bytes. A chunk
|
|
78
|
+
// that does not inflate is scanned raw, which it already was.
|
|
79
|
+
const PNG_SIG = Buffer.from([0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a]);
|
|
80
|
+
const INFLATE_CAP = 16 * 1024 * 1024;
|
|
81
|
+
function inflated(data) {
|
|
82
|
+
try { return zlib.inflateSync(data, { maxOutputLength: INFLATE_CAP }).toString('latin1'); } catch (_e) { return null; }
|
|
83
|
+
}
|
|
84
|
+
function pngTextBearing(buf) {
|
|
85
|
+
const parts = [];
|
|
86
|
+
const rest = (from) => { parts.push(buf.slice(from).toString('latin1')); return parts.join('\n'); };
|
|
87
|
+
let off = 8;
|
|
88
|
+
while (off + 8 <= buf.length) {
|
|
89
|
+
const len = buf.readUInt32BE(off);
|
|
90
|
+
const type = buf.slice(off + 4, off + 8).toString('latin1');
|
|
91
|
+
const start = off + 8;
|
|
92
|
+
const end = start + len;
|
|
93
|
+
if (end + 4 > buf.length) return rest(off); // a truncated chunk: it and all after it, whole
|
|
94
|
+
const data = buf.slice(start, end);
|
|
95
|
+
if (type !== 'IDAT') parts.push(type, data.toString('latin1'));
|
|
96
|
+
if (type === 'zTXt') parts.push(zTXtText(data));
|
|
97
|
+
if (type === 'iTXt') parts.push(iTXtText(data));
|
|
98
|
+
off = end + 4;
|
|
99
|
+
if (type === 'IEND') return rest(off); // bytes after IEND, whole
|
|
100
|
+
}
|
|
101
|
+
return rest(off); // fewer than a chunk header's bytes, no IEND
|
|
102
|
+
}
|
|
103
|
+
// keyword \0 method data
|
|
104
|
+
function zTXtText(data) {
|
|
105
|
+
const z = data.indexOf(0);
|
|
106
|
+
return (z >= 0 && inflated(data.slice(z + 2))) || '';
|
|
107
|
+
}
|
|
108
|
+
// keyword \0 flag method language \0 translated keyword \0 text; the text is deflated when flag is 1
|
|
109
|
+
function iTXtText(data) {
|
|
110
|
+
const z = data.indexOf(0);
|
|
111
|
+
if (z < 0 || data[z + 1] !== 1) return '';
|
|
112
|
+
const lang = data.indexOf(0, z + 3);
|
|
113
|
+
const tr = lang >= 0 ? data.indexOf(0, lang + 1) : -1;
|
|
114
|
+
return (tr >= 0 && inflated(data.slice(tr + 1))) || '';
|
|
115
|
+
}
|
|
116
|
+
|
|
57
117
|
function scanContent(rel, content) {
|
|
118
|
+
// A reader may hand this bytes rather than text. A PNG is then read as its text-bearing regions;
|
|
119
|
+
// anything else is decoded and scanned whole.
|
|
120
|
+
if (Buffer.isBuffer(content)) {
|
|
121
|
+
return scanContent(rel, content.slice(0, 8).equals(PNG_SIG) ? pngTextBearing(content) : content.toString('utf8'));
|
|
122
|
+
}
|
|
58
123
|
const hits = [];
|
|
59
124
|
for (const p of PATTERNS) {
|
|
60
125
|
const m = content.match(p.re);
|
|
@@ -104,10 +169,10 @@ function scanFiles(files, read, readLink) {
|
|
|
104
169
|
}
|
|
105
170
|
let c;
|
|
106
171
|
try { c = read(rel); } catch (_e) { continue; }
|
|
107
|
-
if (typeof c !== 'string') continue;
|
|
172
|
+
if (typeof c !== 'string' && !Buffer.isBuffer(c)) continue;
|
|
108
173
|
hits.push(...scanContent(rel, c));
|
|
109
174
|
}
|
|
110
175
|
return hits;
|
|
111
176
|
}
|
|
112
177
|
|
|
113
|
-
module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, scanPath, scanContent, scanFiles };
|
|
178
|
+
module.exports = { PATTERNS, EMAIL_ALLOW, HOSTIP, PNG_SIG, pngTextBearing, scanPath, scanContent, scanFiles };
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.11.
|
|
3
|
+
"version": "0.11.2",
|
|
4
4
|
"description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
|
@@ -36,7 +36,7 @@
|
|
|
36
36
|
"diff": "node bin/driftproof diff"
|
|
37
37
|
},
|
|
38
38
|
"dependencies": {
|
|
39
|
-
"ajv": "^8.
|
|
39
|
+
"ajv": "^8.18.0"
|
|
40
40
|
},
|
|
41
41
|
"optionalDependencies": {
|
|
42
42
|
"@anthropic-ai/sdk": "^0.39.0"
|
package/spec/RECEIPT.md
CHANGED
|
@@ -183,7 +183,7 @@ carry a numeric `mean`.
|
|
|
183
183
|
## Which schema validates which receipt (the coexistence rule)
|
|
184
184
|
|
|
185
185
|
- **A producer emits the current version.** `config.js` `RECEIPT_SCHEMA_VERSION`
|
|
186
|
-
decides it, and at this revision that is **v0.
|
|
186
|
+
decides it, and at this revision that is **v0.7**.
|
|
187
187
|
- **A reader validates against the receipt's own `schema_version`**, never
|
|
188
188
|
against the newest schema it happens to have. `validateReceipt()` selects the
|
|
189
189
|
schema by that field, which is why a v0.1 receipt from the first report still
|