driftproof 0.11.1 → 0.11.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +29 -11
- package/config.js +1 -1
- package/package.json +2 -2
- package/spec/RECEIPT.md +1 -1
package/README.md
CHANGED
|
@@ -13,17 +13,19 @@ there were too few draws to tell, the receipt says so.
|
|
|
13
13
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
14
14
|
|
|
15
15
|
**[Quickstart](https://driftproofhq.com/#quickstart)** ·
|
|
16
|
-
**[Latest report](https://driftproofhq.com/reports/
|
|
16
|
+
**[Latest report](https://driftproofhq.com/reports/010/)** ·
|
|
17
17
|
[driftproofhq.com](https://driftproofhq.com)
|
|
18
18
|
|
|
19
19
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
20
20
|
format; it does not invent its own.
|
|
21
21
|
|
|
22
|
-
📊 **
|
|
23
|
-
hand-entered
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
|
|
22
|
+
📊 **Ten published reports** (each re-derived from committed files, nothing
|
|
23
|
+
hand-entered: Driftproof's receipts, or for Report #010 the upstream harness's own
|
|
24
|
+
output), spanning seven published report types, the newest being instrument
|
|
25
|
+
comparison. The nine that measure with Driftproof read its own arms by one
|
|
26
|
+
band-based, floor-gated verdict rule and differ in what moves underneath the
|
|
27
|
+
skill — or, in the value report, in which axes are measured; or, in the instrument re-measurement, in the
|
|
28
|
+
instrument itself; or, in the instrument comparison, in which instrument measures:
|
|
27
29
|
|
|
28
30
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
29
31
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
|
|
@@ -76,11 +78,25 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
|
|
|
76
78
|
receipts, which is what makes the delta attributable to the model rather than
|
|
77
79
|
to the instrument. One case sits inside the verdict on the effect floor alone
|
|
78
80
|
and the report names it.
|
|
81
|
+
- **[Report #009](https://driftproofhq.com/reports/009/)** — *instrument
|
|
82
|
+
comparison*: three skills from one plugin, one case each, measured by Claude
|
|
83
|
+
Code's native plugin eval and by Driftproof on the same SKILL.md bytes, task
|
|
84
|
+
prompts and rubrics. The two tools apply different treatments and grade
|
|
85
|
+
differently, so the report reads them side by side, ranks neither, and states
|
|
86
|
+
what each can and cannot establish. Both tools judged with `claude-opus-5`,
|
|
87
|
+
which departs from the judge policy, so its figures are not comparable with
|
|
88
|
+
Reports #001 to #008.
|
|
89
|
+
- **[Report #010](https://driftproofhq.com/reports/010/)** — *instrument
|
|
90
|
+
comparison*: one skill's own behavioural eval from its upstream repository, run
|
|
91
|
+
five times under each of two Claude Code configurations at one commit: 8 passed
|
|
92
|
+
and 2 failed, both failures on one expectation the skill's ADR instructions do
|
|
93
|
+
not state. No Driftproof measurement was taken and no Driftproof judge ran, and
|
|
94
|
+
the report does not isolate what caused the variation.
|
|
79
95
|
|
|
80
96
|
✍️ The launch essay, **[Three model releases later: what actually happens to agent
|
|
81
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads all
|
|
97
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all ten reports
|
|
82
98
|
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
83
|
-
corrected instrument did to three published results. Revised 2026-09-
|
|
99
|
+
corrected instrument did to three published results. Revised 2026-09-22; every
|
|
84
100
|
figure in it is gate-checked against the report page it cites.
|
|
85
101
|
|
|
86
102
|
## Why
|
|
@@ -305,7 +321,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
305
321
|
"model_release_date": "2025-10-01",
|
|
306
322
|
"provider": "anthropic",
|
|
307
323
|
"surface": "claude-cli",
|
|
308
|
-
"runner_version": "0.11.
|
|
324
|
+
"runner_version": "0.11.2",
|
|
309
325
|
"date_utc": "2026-07-27T…Z",
|
|
310
326
|
"registry": "registered",
|
|
311
327
|
"transcripts": "hashes-only",
|
|
@@ -384,7 +400,7 @@ jobs:
|
|
|
384
400
|
runs-on: ubuntu-latest
|
|
385
401
|
steps:
|
|
386
402
|
- uses: actions/checkout@v4
|
|
387
|
-
- uses: driftproofhq/driftproof@v0.11.
|
|
403
|
+
- uses: driftproofhq/driftproof@v0.11.2
|
|
388
404
|
with:
|
|
389
405
|
skill-dir: skills/my-skill
|
|
390
406
|
models: claude-haiku-4-5
|
|
@@ -479,7 +495,7 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
479
495
|
|
|
480
496
|
## Reports
|
|
481
497
|
|
|
482
|
-
|
|
498
|
+
Ten reports are published, spanning seven report types. A report page lives at a
|
|
483
499
|
draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
|
|
484
500
|
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
485
501
|
(see the roll at the top of this README, and
|
|
@@ -487,6 +503,8 @@ draft path — `docs/reports/NNN-draft/` — until the publish sequence renames
|
|
|
487
503
|
Each report and every verdict in it are **re-derived from the receipts** committed
|
|
488
504
|
under [`receipts/`](receipts/)
|
|
489
505
|
(`receipts/report-001/` … `receipts/report-006/`) — nothing is hand-entered.
|
|
506
|
+
Report #010 takes no Driftproof measurement and has no receipts: it is re-derived
|
|
507
|
+
from the upstream harness's output files, published beside its page.
|
|
490
508
|
|
|
491
509
|
Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
|
|
492
510
|
fetched at run time from a pinned commit and verified by sha256 against
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.11.
|
|
12
|
+
const RUNNER_VERSION = '0.11.2';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.11.
|
|
3
|
+
"version": "0.11.2",
|
|
4
4
|
"description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
|
@@ -36,7 +36,7 @@
|
|
|
36
36
|
"diff": "node bin/driftproof diff"
|
|
37
37
|
},
|
|
38
38
|
"dependencies": {
|
|
39
|
-
"ajv": "^8.
|
|
39
|
+
"ajv": "^8.18.0"
|
|
40
40
|
},
|
|
41
41
|
"optionalDependencies": {
|
|
42
42
|
"@anthropic-ai/sdk": "^0.39.0"
|
package/spec/RECEIPT.md
CHANGED
|
@@ -183,7 +183,7 @@ carry a numeric `mean`.
|
|
|
183
183
|
## Which schema validates which receipt (the coexistence rule)
|
|
184
184
|
|
|
185
185
|
- **A producer emits the current version.** `config.js` `RECEIPT_SCHEMA_VERSION`
|
|
186
|
-
decides it, and at this revision that is **v0.
|
|
186
|
+
decides it, and at this revision that is **v0.7**.
|
|
187
187
|
- **A reader validates against the receipt's own `schema_version`**, never
|
|
188
188
|
against the newest schema it happens to have. `validateReceipt()` selects the
|
|
189
189
|
schema by that field, which is why a v0.1 receipt from the first report still
|