driftproof 0.11.2 → 0.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -13,16 +13,16 @@ there were too few draws to tell, the receipt says so.
13
13
   — live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
14
14
 
15
15
  **[Quickstart](https://driftproofhq.com/#quickstart)** ·
16
- **[Latest report](https://driftproofhq.com/reports/010/)** ·
16
+ **[Latest report](https://driftproofhq.com/reports/011/)** ·
17
17
  [driftproofhq.com](https://driftproofhq.com)
18
18
 
19
19
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
20
20
  format; it does not invent its own.
21
21
 
22
- 📊 **Ten published reports** (each re-derived from committed files, nothing
22
+ 📊 **Eleven published reports** (each re-derived from committed files, nothing
23
23
  hand-entered: Driftproof's receipts, or for Report #010 the upstream harness's own
24
24
  output), spanning seven published report types, the newest being instrument
25
- comparison. The nine that measure with Driftproof read its own arms by one
25
+ comparison. The ten that measure with Driftproof read its own arms by one
26
26
  band-based, floor-gated verdict rule and differ in what moves underneath the
27
27
  skill — or, in the value report, in which axes are measured; or, in the instrument re-measurement, in the
28
28
  instrument itself; or, in the instrument comparison, in which instrument measures:
@@ -92,11 +92,21 @@ instrument itself; or, in the instrument comparison, in which instrument measure
92
92
  and 2 failed, both failures on one expectation the skill's ADR instructions do
93
93
  not state. No Driftproof measurement was taken and no Driftproof judge ran, and
94
94
  the report does not isolate what caused the variation.
95
+ - **[Report #011](https://driftproofhq.com/reports/011/)** — *release drift*:
96
+ Claude Opus 5.5 on its release day, measured on Report #009's three skills and
97
+ cases, one case each, beside a fresh Claude Opus 5 arm on the same Claude Code
98
+ version. Read by the runner's own comparison of the with-skill arms, the three
99
+ skills read 1 with no separation detected and 2 with not enough draws to conclude
100
+ at the effect floor, which is not evidence that nothing changed. A second table
101
+ sets Report #009's own Claude Opus 5 receipts beside this run's; between them the
102
+ Claude Code version, its host build and the date all changed. Every draw was
103
+ judged with `claude-opus-5`, which departs from the judge policy, so its figures
104
+ are not comparable with Reports #001 to #008.
95
105
 
96
106
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
97
- skills](https://driftproofhq.com/writing/three-releases/)**, reads all ten reports
107
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all eleven reports
98
108
  together: what moves underneath a skill, what the skill costs to run, and what a
99
- corrected instrument did to three published results. Revised 2026-09-22; every
109
+ corrected instrument did to three published results. Revised 2026-09-23; every
100
110
  figure in it is gate-checked against the report page it cites.
101
111
 
102
112
  ## Why
@@ -229,6 +239,7 @@ npx driftproof badge receipt.json --out badges/my-skill.json
229
239
  # epistemics — no fabricated hashes, excluded from drift verdicts), and emit
230
240
  # the minimal stable summary other tools can consume. See docs/interop.md.
231
241
  npx driftproof import results.json --from agent-skills-eval # or: skillgrade
242
+ npx driftproof import evals/results/<timestamp>/ --from claude-plugin-eval # or: skill-creator
232
243
  npx driftproof export receipt.json --to summary-json
233
244
  ```
234
245
 
@@ -245,7 +256,9 @@ A skill directory is expected to look like:
245
256
  my-skill/
246
257
  SKILL.md # the skill instructions (required)
247
258
  evals/evals.json # agentskills.io/evals suite (required)
248
- .driftproofrc # optional per-project run defaults (models, samples, max_usd)
259
+ .driftproofrc # optional run defaults (models, max_usd); samples, max_cases
260
+ # and judge_model are read only from the working directory's rc
261
+ # (the GitHub Action reads no working-directory rc at all)
249
262
  ... # any bundled files (contribute to content_hash)
250
263
  ```
251
264
 
@@ -312,7 +325,7 @@ A receipt is the unit of evidence — one JSON document conforming to
312
325
 
313
326
  ```jsonc
314
327
  {
315
- "schema_version": "0.7",
328
+ "schema_version": "0.9",
316
329
  "skill": { "name": "commit-message-conventions", "version": "0.2.0",
317
330
  "content_hash": "…sha256 over SKILL.md + bundled files…" },
318
331
  "suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
@@ -321,7 +334,7 @@ A receipt is the unit of evidence — one JSON document conforming to
321
334
  "model_release_date": "2025-10-01",
322
335
  "provider": "anthropic",
323
336
  "surface": "claude-cli",
324
- "runner_version": "0.11.2",
337
+ "runner_version": "0.12.0",
325
338
  "date_utc": "2026-07-27T…Z",
326
339
  "registry": "registered",
327
340
  "transcripts": "hashes-only",
@@ -400,7 +413,7 @@ jobs:
400
413
  runs-on: ubuntu-latest
401
414
  steps:
402
415
  - uses: actions/checkout@v4
403
- - uses: driftproofhq/driftproof@v0.11.2
416
+ - uses: driftproofhq/driftproof@v0.12.0
404
417
  with:
405
418
  skill-dir: skills/my-skill
406
419
  models: claude-haiku-4-5
@@ -416,6 +429,13 @@ model, uploads the whole receipt directory as a build artifact, and renders the
416
429
  run on three surfaces: the **badge**, a **job summary** carrying one row per
417
430
  requested model, and the **check title** of the enforcement step.
418
431
 
432
+ **The checkout's `.driftproofrc` is not read.** The action starts the run in an
433
+ empty working directory, so no key of a `.driftproofrc` at the repository root
434
+ is read, and a skill directory's `.driftproofrc` may not set `max_cases`,
435
+ `samples` or `judge_model`. Both files are pull-request content, and a pull
436
+ request may not narrow the run that measures it. The run takes its models and
437
+ caps from the inputs above.
438
+
419
439
  **The decision is taken over every receipt the run produced.** Each requested
420
440
  model gets one decision state, and the run's `verdict` is the **worst** of them —
421
441
  worst first, in this order:
@@ -448,9 +468,15 @@ the skill hurt — but they **never render as success** on any of the three
448
468
  surfaces: not in the badge, not in the summary row, and not in the check title,
449
469
  which carries a `::warning` naming the state and the models it came from.
450
470
 
471
+ A receipt set that does not say one thing is `REFUSED` as well. Two receipts for
472
+ one requested model, or a receipt with two rows for one case and arm, fail the job
473
+ naming the files or the case, because a decision that depends on which one was read
474
+ is not a pass. Each invocation of the Action writes to its own directory and uploads
475
+ its own artifact, so a job may run it more than once, one skill per step.
476
+
451
477
  Step outputs: `verdict`, `delta` (of the model the worst decision came from),
452
- `worst_state`, `regressed_models`, `missing_models` and `receipts_dir`. Each is
453
- described in [`action.yml`](action.yml).
478
+ `worst_state`, `regressed_models`, `missing_models`, `receipts_dir` and
479
+ `artifact_name`. Each is described in [`action.yml`](action.yml).
454
480
 
455
481
  For a free CI dry-run with **zero model calls**, set `DRIFTPROOF_STUB=1` in the job
456
482
  env — the runner returns canned receipts so the wiring can be tested without spend
@@ -473,6 +499,47 @@ surface. The CLI has carried its own input contract at its own door since
473
499
  `npx driftproof` refuses a malformed cap or model id before it projects a run.
474
500
  Neither statement covers the other. Said of the Action, of CI, or of the runner, "hostile input is refused" is only ever true of the one surface it was measured on.
475
501
 
502
+ ### Scheduled stale check
503
+
504
+ `driftproof stale` says whether each receipt's conclusion still stands under the model, harness,
505
+ skill, suite and judge that would run today. The staleness check runs it on a schedule in your
506
+ repository. While any receipt needs a rerun or a regrade, one issue labelled `driftproof-stale` lists
507
+ each one, what moved, and the command to run next. Later runs update that issue, and the first run
508
+ that finds everything current closes it. It makes no model call and needs no API key.
509
+
510
+ Copy [`examples/workflows/driftproof-stale.yml`](examples/workflows/driftproof-stale.yml) into
511
+ `.github/workflows/`. It runs weekly and on demand, with these permissions and no others:
512
+
513
+ ```yaml
514
+ permissions:
515
+ contents: read
516
+ issues: write
517
+ # ...
518
+ - uses: driftproofhq/driftproof/stale@v0.12.0
519
+ with:
520
+ receipts: receipts/**/*.json
521
+ skill: skills/my-skill
522
+ ```
523
+
524
+ | Input | Default | What it does |
525
+ |---|---|---|
526
+ | `receipts` | `receipts/**/*.json` | Glob of receipt files, one per line for several. |
527
+ | `skill` | none | The skill directory as it is today. Without it, the skill axis reads unknown. |
528
+ | `suite` | the skill's `evals/evals.json` | The eval suite file. |
529
+ | `model`, `judge` | `.driftproofrc` | The model and the judge that would run today. |
530
+ | `harness-version` | `latest` | A Claude Code version, `latest` (read from npm), or `none` (not checked). |
531
+ | `strict` | `false` | Count an unknown axis or an advisory as stale. |
532
+ | `fail-on-stale` | `false` | Fail the job while anything is stale. |
533
+ | `open-issue` | `true` | Keep the issue. |
534
+ | `issue-label` | `driftproof-stale` | Give each check its own label to keep separate issues. |
535
+ | `github-token` | the workflow's token | Used for the issue only. |
536
+
537
+ The job fails on an error, such as a receipt that does not validate, and on stale only when
538
+ `fail-on-stale` is `'true'`. A patch release of the harness alone (2.1.280 to 2.1.281) is an
539
+ advisory: the job summary reports it and nothing fails. An axis that cannot be known is reported as
540
+ unknown, never as current. Outputs: `result` (`current`, `advisory`, `stale` or `error`),
541
+ `stale-count`, `issue-number` and `report-dir`.
542
+
476
543
  ### Badge
477
544
 
478
545
  `driftproof badge <receipt>` emits a [shields.io endpoint](https://shields.io/badges/endpoint-badge)
@@ -495,7 +562,7 @@ site, so it reflects a real dated run, not a hand-set color.
495
562
 
496
563
  ## Reports
497
564
 
498
- Ten reports are published, spanning seven report types. A report page lives at a
565
+ Eleven reports are published, spanning seven report types. A report page lives at a
499
566
  draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
500
567
  `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
501
568
  (see the roll at the top of this README, and