driftproof 0.11.2 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +79 -12
- package/bin/driftproof +343 -39
- package/config/models.json +3 -2
- package/config.js +2 -2
- package/lib/counts.js +104 -0
- package/lib/decision.js +72 -31
- package/lib/diff.js +12 -3
- package/lib/export.js +6 -2
- package/lib/importers-anthropic.js +451 -0
- package/lib/importers.js +50 -13
- package/lib/init.js +3 -2
- package/lib/receipt.js +182 -10
- package/lib/regrade.js +266 -0
- package/lib/reuse.js +144 -9
- package/lib/run.js +39 -4
- package/lib/skill.js +70 -15
- package/lib/stale.js +194 -0
- package/lib/verdict.js +50 -16
- package/package.json +1 -1
- package/spec/RECEIPT.md +103 -13
- package/spec/receipt.schema.json +330 -15
- package/spec/receipt.v0.7.schema.json +1563 -0
- package/spec/receipt.v0.8.schema.json +1677 -0
- package/spec/stale.v1.schema.json +353 -0
package/README.md
CHANGED
|
@@ -13,16 +13,16 @@ there were too few draws to tell, the receipt says so.
|
|
|
13
13
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
14
14
|
|
|
15
15
|
**[Quickstart](https://driftproofhq.com/#quickstart)** ·
|
|
16
|
-
**[Latest report](https://driftproofhq.com/reports/
|
|
16
|
+
**[Latest report](https://driftproofhq.com/reports/011/)** ·
|
|
17
17
|
[driftproofhq.com](https://driftproofhq.com)
|
|
18
18
|
|
|
19
19
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
20
20
|
format; it does not invent its own.
|
|
21
21
|
|
|
22
|
-
📊 **
|
|
22
|
+
📊 **Eleven published reports** (each re-derived from committed files, nothing
|
|
23
23
|
hand-entered: Driftproof's receipts, or for Report #010 the upstream harness's own
|
|
24
24
|
output), spanning seven published report types, the newest being instrument
|
|
25
|
-
comparison. The
|
|
25
|
+
comparison. The ten that measure with Driftproof read its own arms by one
|
|
26
26
|
band-based, floor-gated verdict rule and differ in what moves underneath the
|
|
27
27
|
skill — or, in the value report, in which axes are measured; or, in the instrument re-measurement, in the
|
|
28
28
|
instrument itself; or, in the instrument comparison, in which instrument measures:
|
|
@@ -92,11 +92,21 @@ instrument itself; or, in the instrument comparison, in which instrument measure
|
|
|
92
92
|
and 2 failed, both failures on one expectation the skill's ADR instructions do
|
|
93
93
|
not state. No Driftproof measurement was taken and no Driftproof judge ran, and
|
|
94
94
|
the report does not isolate what caused the variation.
|
|
95
|
+
- **[Report #011](https://driftproofhq.com/reports/011/)** — *release drift*:
|
|
96
|
+
Claude Opus 5.5 on its release day, measured on Report #009's three skills and
|
|
97
|
+
cases, one case each, beside a fresh Claude Opus 5 arm on the same Claude Code
|
|
98
|
+
version. Read by the runner's own comparison of the with-skill arms, the three
|
|
99
|
+
skills read 1 with no separation detected and 2 with not enough draws to conclude
|
|
100
|
+
at the effect floor, which is not evidence that nothing changed. A second table
|
|
101
|
+
sets Report #009's own Claude Opus 5 receipts beside this run's; between them the
|
|
102
|
+
Claude Code version, its host build and the date all changed. Every draw was
|
|
103
|
+
judged with `claude-opus-5`, which departs from the judge policy, so its figures
|
|
104
|
+
are not comparable with Reports #001 to #008.
|
|
95
105
|
|
|
96
106
|
✍️ The launch essay, **[Three model releases later: what actually happens to agent
|
|
97
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads all
|
|
107
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all eleven reports
|
|
98
108
|
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
99
|
-
corrected instrument did to three published results. Revised 2026-09-
|
|
109
|
+
corrected instrument did to three published results. Revised 2026-09-23; every
|
|
100
110
|
figure in it is gate-checked against the report page it cites.
|
|
101
111
|
|
|
102
112
|
## Why
|
|
@@ -229,6 +239,7 @@ npx driftproof badge receipt.json --out badges/my-skill.json
|
|
|
229
239
|
# epistemics — no fabricated hashes, excluded from drift verdicts), and emit
|
|
230
240
|
# the minimal stable summary other tools can consume. See docs/interop.md.
|
|
231
241
|
npx driftproof import results.json --from agent-skills-eval # or: skillgrade
|
|
242
|
+
npx driftproof import evals/results/<timestamp>/ --from claude-plugin-eval # or: skill-creator
|
|
232
243
|
npx driftproof export receipt.json --to summary-json
|
|
233
244
|
```
|
|
234
245
|
|
|
@@ -245,7 +256,9 @@ A skill directory is expected to look like:
|
|
|
245
256
|
my-skill/
|
|
246
257
|
SKILL.md # the skill instructions (required)
|
|
247
258
|
evals/evals.json # agentskills.io/evals suite (required)
|
|
248
|
-
.driftproofrc # optional
|
|
259
|
+
.driftproofrc # optional run defaults (models, max_usd); samples, max_cases
|
|
260
|
+
# and judge_model are read only from the working directory's rc
|
|
261
|
+
# (the GitHub Action reads no working-directory rc at all)
|
|
249
262
|
... # any bundled files (contribute to content_hash)
|
|
250
263
|
```
|
|
251
264
|
|
|
@@ -312,7 +325,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
312
325
|
|
|
313
326
|
```jsonc
|
|
314
327
|
{
|
|
315
|
-
"schema_version": "0.
|
|
328
|
+
"schema_version": "0.9",
|
|
316
329
|
"skill": { "name": "commit-message-conventions", "version": "0.2.0",
|
|
317
330
|
"content_hash": "…sha256 over SKILL.md + bundled files…" },
|
|
318
331
|
"suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
|
|
@@ -321,7 +334,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
321
334
|
"model_release_date": "2025-10-01",
|
|
322
335
|
"provider": "anthropic",
|
|
323
336
|
"surface": "claude-cli",
|
|
324
|
-
"runner_version": "0.
|
|
337
|
+
"runner_version": "0.12.0",
|
|
325
338
|
"date_utc": "2026-07-27T…Z",
|
|
326
339
|
"registry": "registered",
|
|
327
340
|
"transcripts": "hashes-only",
|
|
@@ -400,7 +413,7 @@ jobs:
|
|
|
400
413
|
runs-on: ubuntu-latest
|
|
401
414
|
steps:
|
|
402
415
|
- uses: actions/checkout@v4
|
|
403
|
-
- uses: driftproofhq/driftproof@v0.
|
|
416
|
+
- uses: driftproofhq/driftproof@v0.12.0
|
|
404
417
|
with:
|
|
405
418
|
skill-dir: skills/my-skill
|
|
406
419
|
models: claude-haiku-4-5
|
|
@@ -416,6 +429,13 @@ model, uploads the whole receipt directory as a build artifact, and renders the
|
|
|
416
429
|
run on three surfaces: the **badge**, a **job summary** carrying one row per
|
|
417
430
|
requested model, and the **check title** of the enforcement step.
|
|
418
431
|
|
|
432
|
+
**The checkout's `.driftproofrc` is not read.** The action starts the run in an
|
|
433
|
+
empty working directory, so no key of a `.driftproofrc` at the repository root
|
|
434
|
+
is read, and a skill directory's `.driftproofrc` may not set `max_cases`,
|
|
435
|
+
`samples` or `judge_model`. Both files are pull-request content, and a pull
|
|
436
|
+
request may not narrow the run that measures it. The run takes its models and
|
|
437
|
+
caps from the inputs above.
|
|
438
|
+
|
|
419
439
|
**The decision is taken over every receipt the run produced.** Each requested
|
|
420
440
|
model gets one decision state, and the run's `verdict` is the **worst** of them —
|
|
421
441
|
worst first, in this order:
|
|
@@ -448,9 +468,15 @@ the skill hurt — but they **never render as success** on any of the three
|
|
|
448
468
|
surfaces: not in the badge, not in the summary row, and not in the check title,
|
|
449
469
|
which carries a `::warning` naming the state and the models it came from.
|
|
450
470
|
|
|
471
|
+
A receipt set that does not say one thing is `REFUSED` as well. Two receipts for
|
|
472
|
+
one requested model, or a receipt with two rows for one case and arm, fail the job
|
|
473
|
+
naming the files or the case, because a decision that depends on which one was read
|
|
474
|
+
is not a pass. Each invocation of the Action writes to its own directory and uploads
|
|
475
|
+
its own artifact, so a job may run it more than once, one skill per step.
|
|
476
|
+
|
|
451
477
|
Step outputs: `verdict`, `delta` (of the model the worst decision came from),
|
|
452
|
-
`worst_state`, `regressed_models`, `missing_models
|
|
453
|
-
described in [`action.yml`](action.yml).
|
|
478
|
+
`worst_state`, `regressed_models`, `missing_models`, `receipts_dir` and
|
|
479
|
+
`artifact_name`. Each is described in [`action.yml`](action.yml).
|
|
454
480
|
|
|
455
481
|
For a free CI dry-run with **zero model calls**, set `DRIFTPROOF_STUB=1` in the job
|
|
456
482
|
env — the runner returns canned receipts so the wiring can be tested without spend
|
|
@@ -473,6 +499,47 @@ surface. The CLI has carried its own input contract at its own door since
|
|
|
473
499
|
`npx driftproof` refuses a malformed cap or model id before it projects a run.
|
|
474
500
|
Neither statement covers the other. Said of the Action, of CI, or of the runner, "hostile input is refused" is only ever true of the one surface it was measured on.
|
|
475
501
|
|
|
502
|
+
### Scheduled stale check
|
|
503
|
+
|
|
504
|
+
`driftproof stale` says whether each receipt's conclusion still stands under the model, harness,
|
|
505
|
+
skill, suite and judge that would run today. The staleness check runs it on a schedule in your
|
|
506
|
+
repository. While any receipt needs a rerun or a regrade, one issue labelled `driftproof-stale` lists
|
|
507
|
+
each one, what moved, and the command to run next. Later runs update that issue, and the first run
|
|
508
|
+
that finds everything current closes it. It makes no model call and needs no API key.
|
|
509
|
+
|
|
510
|
+
Copy [`examples/workflows/driftproof-stale.yml`](examples/workflows/driftproof-stale.yml) into
|
|
511
|
+
`.github/workflows/`. It runs weekly and on demand, with these permissions and no others:
|
|
512
|
+
|
|
513
|
+
```yaml
|
|
514
|
+
permissions:
|
|
515
|
+
contents: read
|
|
516
|
+
issues: write
|
|
517
|
+
# ...
|
|
518
|
+
- uses: driftproofhq/driftproof/stale@v0.12.0
|
|
519
|
+
with:
|
|
520
|
+
receipts: receipts/**/*.json
|
|
521
|
+
skill: skills/my-skill
|
|
522
|
+
```
|
|
523
|
+
|
|
524
|
+
| Input | Default | What it does |
|
|
525
|
+
|---|---|---|
|
|
526
|
+
| `receipts` | `receipts/**/*.json` | Glob of receipt files, one per line for several. |
|
|
527
|
+
| `skill` | none | The skill directory as it is today. Without it, the skill axis reads unknown. |
|
|
528
|
+
| `suite` | the skill's `evals/evals.json` | The eval suite file. |
|
|
529
|
+
| `model`, `judge` | `.driftproofrc` | The model and the judge that would run today. |
|
|
530
|
+
| `harness-version` | `latest` | A Claude Code version, `latest` (read from npm), or `none` (not checked). |
|
|
531
|
+
| `strict` | `false` | Count an unknown axis or an advisory as stale. |
|
|
532
|
+
| `fail-on-stale` | `false` | Fail the job while anything is stale. |
|
|
533
|
+
| `open-issue` | `true` | Keep the issue. |
|
|
534
|
+
| `issue-label` | `driftproof-stale` | Give each check its own label to keep separate issues. |
|
|
535
|
+
| `github-token` | the workflow's token | Used for the issue only. |
|
|
536
|
+
|
|
537
|
+
The job fails on an error, such as a receipt that does not validate, and on stale only when
|
|
538
|
+
`fail-on-stale` is `'true'`. A patch release of the harness alone (2.1.280 to 2.1.281) is an
|
|
539
|
+
advisory: the job summary reports it and nothing fails. An axis that cannot be known is reported as
|
|
540
|
+
unknown, never as current. Outputs: `result` (`current`, `advisory`, `stale` or `error`),
|
|
541
|
+
`stale-count`, `issue-number` and `report-dir`.
|
|
542
|
+
|
|
476
543
|
### Badge
|
|
477
544
|
|
|
478
545
|
`driftproof badge <receipt>` emits a [shields.io endpoint](https://shields.io/badges/endpoint-badge)
|
|
@@ -495,7 +562,7 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
495
562
|
|
|
496
563
|
## Reports
|
|
497
564
|
|
|
498
|
-
|
|
565
|
+
Eleven reports are published, spanning seven report types. A report page lives at a
|
|
499
566
|
draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
|
|
500
567
|
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
501
568
|
(see the roll at the top of this README, and
|