driftproof 0.11.1 → 0.11.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -13,17 +13,19 @@ there were too few draws to tell, the receipt says so.
13
13
   — live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
14
14
 
15
15
  **[Quickstart](https://driftproofhq.com/#quickstart)** ·
16
- **[Latest report](https://driftproofhq.com/reports/008/)** ·
16
+ **[Latest report](https://driftproofhq.com/reports/010/)** ·
17
17
  [driftproofhq.com](https://driftproofhq.com)
18
18
 
19
19
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
20
20
  format; it does not invent its own.
21
21
 
22
- 📊 **Eight published reports** (each re-derived from committed receipts, nothing
23
- hand-entered), spanning six published report types, the newest being release
24
- drift. All eight share one band-based, floor-gated verdict rule and
25
- differ in what moves underneath the skill — or, in the value report, in which
26
- axes are measured; or, in the instrument re-measurement, in the instrument itself:
22
+ 📊 **Ten published reports** (each re-derived from committed files, nothing
23
+ hand-entered: Driftproof's receipts, or for Report #010 the upstream harness's own
24
+ output), spanning seven published report types, the newest being instrument
25
+ comparison. The nine that measure with Driftproof read its own arms by one
26
+ band-based, floor-gated verdict rule and differ in what moves underneath the
27
+ skill — or, in the value report, in which axes are measured; or, in the instrument re-measurement, in the
28
+ instrument itself; or, in the instrument comparison, in which instrument measures:
27
29
 
28
30
  - **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
29
31
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
@@ -76,11 +78,25 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
76
78
  receipts, which is what makes the delta attributable to the model rather than
77
79
  to the instrument. One case sits inside the verdict on the effect floor alone
78
80
  and the report names it.
81
+ - **[Report #009](https://driftproofhq.com/reports/009/)** — *instrument
82
+ comparison*: three skills from one plugin, one case each, measured by Claude
83
+ Code's native plugin eval and by Driftproof on the same SKILL.md bytes, task
84
+ prompts and rubrics. The two tools apply different treatments and grade
85
+ differently, so the report reads them side by side, ranks neither, and states
86
+ what each can and cannot establish. Both tools judged with `claude-opus-5`,
87
+ which departs from the judge policy, so its figures are not comparable with
88
+ Reports #001 to #008.
89
+ - **[Report #010](https://driftproofhq.com/reports/010/)** — *instrument
90
+ comparison*: one skill's own behavioural eval from its upstream repository, run
91
+ five times under each of two Claude Code configurations at one commit: 8 passed
92
+ and 2 failed, both failures on one expectation the skill's ADR instructions do
93
+ not state. No Driftproof measurement was taken and no Driftproof judge ran, and
94
+ the report does not isolate what caused the variation.
79
95
 
80
96
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
81
- skills](https://driftproofhq.com/writing/three-releases/)**, reads all eight reports
97
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all ten reports
82
98
  together: what moves underneath a skill, what the skill costs to run, and what a
83
- corrected instrument did to three published results. Revised 2026-09-01; every
99
+ corrected instrument did to three published results. Revised 2026-09-22; every
84
100
  figure in it is gate-checked against the report page it cites.
85
101
 
86
102
  ## Why
@@ -305,7 +321,7 @@ A receipt is the unit of evidence — one JSON document conforming to
305
321
  "model_release_date": "2025-10-01",
306
322
  "provider": "anthropic",
307
323
  "surface": "claude-cli",
308
- "runner_version": "0.11.1",
324
+ "runner_version": "0.11.2",
309
325
  "date_utc": "2026-07-27T…Z",
310
326
  "registry": "registered",
311
327
  "transcripts": "hashes-only",
@@ -384,7 +400,7 @@ jobs:
384
400
  runs-on: ubuntu-latest
385
401
  steps:
386
402
  - uses: actions/checkout@v4
387
- - uses: driftproofhq/driftproof@v0.11.1
403
+ - uses: driftproofhq/driftproof@v0.11.2
388
404
  with:
389
405
  skill-dir: skills/my-skill
390
406
  models: claude-haiku-4-5
@@ -479,7 +495,7 @@ site, so it reflects a real dated run, not a hand-set color.
479
495
 
480
496
  ## Reports
481
497
 
482
- Eight reports are published, spanning six report types. A report page lives at a
498
+ Ten reports are published, spanning seven report types. A report page lives at a
483
499
  draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
484
500
  `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
485
501
  (see the roll at the top of this README, and
@@ -487,6 +503,8 @@ draft path — `docs/reports/NNN-draft/` — until the publish sequence renames
487
503
  Each report and every verdict in it are **re-derived from the receipts** committed
488
504
  under [`receipts/`](receipts/)
489
505
  (`receipts/report-001/` … `receipts/report-006/`) — nothing is hand-entered.
506
+ Report #010 takes no Driftproof measurement and has no receipts: it is re-derived
507
+ from the upstream harness's output files, published beside its page.
490
508
 
491
509
  Driftproof does **not** commit third-party skill content. Each `SKILL.md` is
492
510
  fetched at run time from a pinned commit and verified by sha256 against
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.11.1';
12
+ const RUNNER_VERSION = '0.11.2';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "driftproof",
3
- "version": "0.11.1",
3
+ "version": "0.11.2",
4
4
  "description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
@@ -36,7 +36,7 @@
36
36
  "diff": "node bin/driftproof diff"
37
37
  },
38
38
  "dependencies": {
39
- "ajv": "^8.17.1"
39
+ "ajv": "^8.18.0"
40
40
  },
41
41
  "optionalDependencies": {
42
42
  "@anthropic-ai/sdk": "^0.39.0"
package/spec/RECEIPT.md CHANGED
@@ -183,7 +183,7 @@ carry a numeric `mean`.
183
183
  ## Which schema validates which receipt (the coexistence rule)
184
184
 
185
185
  - **A producer emits the current version.** `config.js` `RECEIPT_SCHEMA_VERSION`
186
- decides it, and at this revision that is **v0.6**.
186
+ decides it, and at this revision that is **v0.7**.
187
187
  - **A reader validates against the receipt's own `schema_version`**, never
188
188
  against the newest schema it happens to have. `validateReceipt()` selects the
189
189
  schema by that field, which is why a v0.1 receipt from the first report still