driftproof 0.7.2 → 0.8.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,25 +1,28 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
2
  # Driftproof
3
3
 
4
- **A dated proof that this skill, this hash, this model, still helps.**
4
+ **Your skill passed. On which model? On what date?**
5
+
6
+ A SKILL.md teaches an AI coding agent how you like things done. When a new model
7
+ ships, the same file can stop helping, or start hurting. Driftproof re-runs the
8
+ skill's tests on the new model and hands you a dated, hash-verified receipt
9
+ saying whether it still helps.
5
10
 
6
11
  [![driftproof](https://img.shields.io/endpoint?url=https://driftproofhq.com/badges/commit-message-conventions.json)](https://driftproofhq.com)
7
12
  &nbsp;— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
8
13
 
9
- Driftproof is an open **receipt spec** plus a **runner** that measures whether an
10
- agent skill actually helps — by running the skill's eval suite **with** and
11
- **without** the skill on a named model version, judging each case several times to
12
- get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
13
- across model releases and you get a **drift report**.
14
+ **[Quickstart](https://driftproofhq.com/#quickstart)** ·
15
+ **[Latest report](https://driftproofhq.com/reports/008/)** ·
16
+ [driftproofhq.com](https://driftproofhq.com)
14
17
 
15
18
  Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
16
19
  format; it does not invent its own.
17
20
 
18
- 📊 **Seven published reports** (each re-derived from committed receipts, nothing
19
- hand-entered), spanning six published report types, the newest being instrument
20
- re-measurement. All seven share one band-based, floor-gated verdict rule and
21
+ 📊 **Eight published reports** (each re-derived from committed receipts, nothing
22
+ hand-entered), spanning six published report types, the newest being release
23
+ drift. All eight share one band-based, floor-gated verdict rule and
21
24
  differ in what moves underneath the skill — or, in the value report, in which
22
- axes are measured; or, in the newest, in the instrument itself:
25
+ axes are measured; or, in the instrument re-measurement, in the instrument itself:
23
26
 
24
27
  - **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
25
28
  public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
@@ -62,9 +65,18 @@ axes are measured; or, in the newest, in the instrument itself:
62
65
  2026-07-27, and the truncated run it caused measured *less* variance than the
63
66
  clean re-run, which is the direction that flatters an instrument. Both runs are
64
67
  published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
68
+ - **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
69
+ two of Report #007's cells re-measured on `claude-fable-5-1` against
70
+ `claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
71
+ identical before the first call. **Both cells came back within noise**: the
72
+ 14 cases read 0 improved, 0 regressed, 14 within noise, 0 not measured. The
73
+ first release pair in this project where both sides are generation-sampled
74
+ receipts, which is what makes the delta attributable to the model rather than
75
+ to the instrument. One case sits inside the verdict on the effect floor alone
76
+ and the report names it.
65
77
 
66
78
  ✍️ The launch essay, **[Three model releases later: what actually happens to agent
67
- skills](https://driftproofhq.com/writing/three-releases/)**, reads all seven reports
79
+ skills](https://driftproofhq.com/writing/three-releases/)**, reads all eight reports
68
80
  together: what moves underneath a skill, what the skill costs to run, and what a
69
81
  corrected instrument did to three published results. Revised 2026-09-01; every
70
82
  figure in it is gate-checked against the report page it cites.
@@ -229,7 +241,7 @@ A receipt is the unit of evidence — one JSON document conforming to
229
241
  "model_release_date": "2025-10-01",
230
242
  "provider": "anthropic",
231
243
  "surface": "claude-cli",
232
- "runner_version": "0.7.2",
244
+ "runner_version": "0.8.0",
233
245
  "date_utc": "2026-07-27T…Z",
234
246
  "registry": "registered",
235
247
  "transcripts": "hashes-only",
@@ -303,7 +315,7 @@ jobs:
303
315
  runs-on: ubuntu-latest
304
316
  steps:
305
317
  - uses: actions/checkout@v4
306
- - uses: driftproofhq/driftproof@v0.7.2
318
+ - uses: driftproofhq/driftproof@v0.8.0
307
319
  with:
308
320
  skill-dir: skills/my-skill
309
321
  models: claude-haiku-4-5
@@ -343,7 +355,7 @@ site, so it reflects a real dated run, not a hand-set color.
343
355
 
344
356
  ## Reports
345
357
 
346
- Seven reports are published, spanning six report types. A report page lives at a
358
+ Eight reports are published, spanning six report types. A report page lives at a
347
359
  draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
348
360
  `scripts/build-public.sh` excludes every `*-draft/` path from the published tree
349
361
  (see the roll at the top of this README, and
@@ -1,5 +1,5 @@
1
1
  {
2
- "_comment": "Driftproof model registry. The runner resolves --models ids against this list; an unknown id still runs but its receipt is marked registry:\"unregistered\" and its cost is estimated with the conservative default price in lib/models.js. Prices are STANDARD first-party USD per 1,000,000 tokens (input/output). Anthropic prices are from platform.claude.com (July 2026); Sonnet 5's introductory rate is deliberately NOT used so projections stay an upper bound. OpenAI prices are the public standard per-MTok rates as of July 2026 (see reports/phase-6-providers.md for the fetched source): GPT-5.6 Sol $5/$30, Terra $2.50/$15, Luna $1/$6; GPT-5.5 $5/$30; GPT-5.4 $2.50/$15. gpt-5.4-mini is NOT in the published table (which lists a Nano tier at $0.20/$1.25), so its price here is a deliberately conservative UPPER-bound estimate flagged price_estimate:true. `released` is best-effort (spec RECEIPT.md open question #4). `tier` drives reporting; `judge_eligible` encodes the fixed-judge policy (docs/judge-policy.html) — only the cheap Haiku judge is eligible, and NO OpenAI model is judge-eligible (the judge stays claude-haiku across providers so a cross-substrate report varies only the target model). `provider` is the two-axis provider; per-provider surface/base_url config is under `providers`. The release trigger (scripts/release-watch.js) appends auto-discovered ids here with auto_added:true.",
2
+ "_comment": "Driftproof model registry. The runner resolves --models ids against this list; an unknown id still runs but its receipt is marked registry:\"unregistered\" and its cost is estimated with the conservative default price in lib/models.js. Prices are STANDARD first-party USD per 1,000,000 tokens (input/output). Anthropic prices are from platform.claude.com (July 2026); Sonnet 5's introductory rate is deliberately NOT used so projections stay an upper bound. OpenAI prices are the public standard per-MTok rates as of July 2026 (see reports/phase-6-providers.md for the fetched source): GPT-5.6 Sol $5/$30, Terra $2.50/$15, Luna $1/$6; GPT-5.5 $5/$30; GPT-5.4 $2.50/$15. gpt-5.4-mini is NOT in the published table (which lists a Nano tier at $0.20/$1.25), so its price here is a deliberately conservative UPPER-bound estimate flagged price_estimate:true. `released` is best-effort (spec RECEIPT.md open question #4). `tier` drives reporting; `judge_eligible` encodes the fixed-judge policy (docs/judge-policy.html) — only the cheap Haiku judge is eligible, and NO OpenAI model is judge-eligible (the judge stays claude-haiku across providers so a cross-substrate report varies only the target model). `provider` is the two-axis provider; per-provider surface/base_url config is under `providers`. The release trigger (scripts/release-watch.js) appends auto-discovered ids here with auto_added:true. Three optional per-row annotations were added by spec-019b (2026-09-01) and are read by NO code path in lib/: `cache_read_price` is the documented cache-read rate per MTok, recorded for provenance and NOT consumed -- lib/cost.js prices input and output only, and the 019b gate asserts that it stays that way, so wiring caching into the projection turns that assertion red rather than leaving this note quietly false; `lifecycle: legacy` records that the docs moved an id to the Legacy list, and never means removed (claude-fable-5 is referenced by receipts/report-007 and receipts/report-007-rerun); `surface_verified` records that a live call reached the id on that surface on `surface_verified_at`, which is an OBSERVATION, not a routing setting -- the surface is chosen at call time by lib/provider.js from CLAUDE_PROVIDER. claude-fable-5-1 prices are sourced to specs/019b-fable-5-1-registration/evidence/docs-pricing-snapshot-2026-09-01.json.",
3
3
  "registry_version": "1.1",
4
4
  "provider": "anthropic",
5
5
  "providers": {
@@ -8,7 +8,8 @@
8
8
  },
9
9
  "default_price_note": "Unregistered ids are costed at the most expensive known tier (Fable, $10/$50) — a budget guard must never under-estimate. See lib/models.js DEFAULT_PRICE.",
10
10
  "models": [
11
- { "id": "claude-fable-5", "family": "fable", "provider": "anthropic", "released": "2026-06-09", "input_price": 10.0, "output_price": 50.0, "tier": "frontier", "judge_eligible": false },
11
+ { "id": "claude-fable-5-1", "family": "fable", "provider": "anthropic", "released": "2026-09-01", "input_price": 10.0, "output_price": 50.0, "cache_read_price": 0.25, "tier": "frontier", "judge_eligible": false, "surface_verified": "claude-cli", "surface_verified_at": "2026-09-01" },
12
+ { "id": "claude-fable-5", "family": "fable", "provider": "anthropic", "released": "2026-06-09", "input_price": 10.0, "output_price": 50.0, "tier": "frontier", "judge_eligible": false, "lifecycle": "legacy" },
12
13
  { "id": "claude-opus-5", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
13
14
  { "id": "claude-opus-4-8", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
14
15
  { "id": "claude-opus-4-7", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.7.2';
12
+ const RUNNER_VERSION = '0.8.0';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
package/package.json CHANGED
@@ -1,14 +1,16 @@
1
1
  {
2
2
  "name": "driftproof",
3
- "version": "0.7.2",
3
+ "version": "0.8.0",
4
4
  "description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
7
7
  "agent-skills",
8
- "evals",
9
- "llm",
8
+ "skill-md",
9
+ "skill-eval",
10
+ "llm-evaluation",
11
+ "claude-code",
10
12
  "drift",
11
- "verification"
13
+ "receipts"
12
14
  ],
13
15
  "homepage": "https://driftproofhq.com",
14
16
  "repository": {
@@ -41,5 +43,8 @@
41
43
  },
42
44
  "engines": {
43
45
  "node": ">=22"
46
+ },
47
+ "devDependencies": {
48
+ "html-validate": "^11.11.0"
44
49
  }
45
50
  }