driftproof 0.7.2 → 0.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +26 -14
- package/config/models.json +3 -2
- package/config.js +1 -1
- package/package.json +9 -4
package/README.md
CHANGED
|
@@ -1,25 +1,28 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
2
|
# Driftproof
|
|
3
3
|
|
|
4
|
-
**
|
|
4
|
+
**Your skill passed. On which model? On what date?**
|
|
5
|
+
|
|
6
|
+
A SKILL.md teaches an AI coding agent how you like things done. When a new model
|
|
7
|
+
ships, the same file can stop helping, or start hurting. Driftproof re-runs the
|
|
8
|
+
skill's tests on the new model and hands you a dated, hash-verified receipt
|
|
9
|
+
saying whether it still helps.
|
|
5
10
|
|
|
6
11
|
[](https://driftproofhq.com)
|
|
7
12
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
8
13
|
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
get a confidence band, and emitting a **hash-verified**, dated **receipt**. Diff two receipts
|
|
13
|
-
across model releases and you get a **drift report**.
|
|
14
|
+
**[Quickstart](https://driftproofhq.com/#quickstart)** ·
|
|
15
|
+
**[Latest report](https://driftproofhq.com/reports/008/)** ·
|
|
16
|
+
[driftproofhq.com](https://driftproofhq.com)
|
|
14
17
|
|
|
15
18
|
Driftproof **consumes** the [`agentskills.io/evals`](https://agentskills.io) eval
|
|
16
19
|
format; it does not invent its own.
|
|
17
20
|
|
|
18
|
-
📊 **
|
|
19
|
-
hand-entered), spanning six published report types, the newest being
|
|
20
|
-
|
|
21
|
+
📊 **Eight published reports** (each re-derived from committed receipts, nothing
|
|
22
|
+
hand-entered), spanning six published report types, the newest being release
|
|
23
|
+
drift. All eight share one band-based, floor-gated verdict rule and
|
|
21
24
|
differ in what moves underneath the skill — or, in the value report, in which
|
|
22
|
-
axes are measured; or, in the
|
|
25
|
+
axes are measured; or, in the instrument re-measurement, in the instrument itself:
|
|
23
26
|
|
|
24
27
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
25
28
|
public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
|
|
@@ -62,9 +65,18 @@ axes are measured; or, in the newest, in the instrument itself:
|
|
|
62
65
|
2026-07-27, and the truncated run it caused measured *less* variance than the
|
|
63
66
|
clean re-run, which is the direction that flatters an instrument. Both runs are
|
|
64
67
|
published, the broken one as evidence. Amends #005 to v1.2 and #006 to v1.1.
|
|
68
|
+
- **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
|
|
69
|
+
two of Report #007's cells re-measured on `claude-fable-5-1` against
|
|
70
|
+
`claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
|
|
71
|
+
identical before the first call. **Both cells came back within noise**: the
|
|
72
|
+
14 cases read 0 improved, 0 regressed, 14 within noise, 0 not measured. The
|
|
73
|
+
first release pair in this project where both sides are generation-sampled
|
|
74
|
+
receipts, which is what makes the delta attributable to the model rather than
|
|
75
|
+
to the instrument. One case sits inside the verdict on the effect floor alone
|
|
76
|
+
and the report names it.
|
|
65
77
|
|
|
66
78
|
✍️ The launch essay, **[Three model releases later: what actually happens to agent
|
|
67
|
-
skills](https://driftproofhq.com/writing/three-releases/)**, reads all
|
|
79
|
+
skills](https://driftproofhq.com/writing/three-releases/)**, reads all eight reports
|
|
68
80
|
together: what moves underneath a skill, what the skill costs to run, and what a
|
|
69
81
|
corrected instrument did to three published results. Revised 2026-09-01; every
|
|
70
82
|
figure in it is gate-checked against the report page it cites.
|
|
@@ -229,7 +241,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
229
241
|
"model_release_date": "2025-10-01",
|
|
230
242
|
"provider": "anthropic",
|
|
231
243
|
"surface": "claude-cli",
|
|
232
|
-
"runner_version": "0.
|
|
244
|
+
"runner_version": "0.8.0",
|
|
233
245
|
"date_utc": "2026-07-27T…Z",
|
|
234
246
|
"registry": "registered",
|
|
235
247
|
"transcripts": "hashes-only",
|
|
@@ -303,7 +315,7 @@ jobs:
|
|
|
303
315
|
runs-on: ubuntu-latest
|
|
304
316
|
steps:
|
|
305
317
|
- uses: actions/checkout@v4
|
|
306
|
-
- uses: driftproofhq/driftproof@v0.
|
|
318
|
+
- uses: driftproofhq/driftproof@v0.8.0
|
|
307
319
|
with:
|
|
308
320
|
skill-dir: skills/my-skill
|
|
309
321
|
models: claude-haiku-4-5
|
|
@@ -343,7 +355,7 @@ site, so it reflects a real dated run, not a hand-set color.
|
|
|
343
355
|
|
|
344
356
|
## Reports
|
|
345
357
|
|
|
346
|
-
|
|
358
|
+
Eight reports are published, spanning six report types. A report page lives at a
|
|
347
359
|
draft path — `docs/reports/NNN-draft/` — until the publish sequence renames it, and
|
|
348
360
|
`scripts/build-public.sh` excludes every `*-draft/` path from the published tree
|
|
349
361
|
(see the roll at the top of this README, and
|
package/config/models.json
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
{
|
|
2
|
-
"_comment": "Driftproof model registry. The runner resolves --models ids against this list; an unknown id still runs but its receipt is marked registry:\"unregistered\" and its cost is estimated with the conservative default price in lib/models.js. Prices are STANDARD first-party USD per 1,000,000 tokens (input/output). Anthropic prices are from platform.claude.com (July 2026); Sonnet 5's introductory rate is deliberately NOT used so projections stay an upper bound. OpenAI prices are the public standard per-MTok rates as of July 2026 (see reports/phase-6-providers.md for the fetched source): GPT-5.6 Sol $5/$30, Terra $2.50/$15, Luna $1/$6; GPT-5.5 $5/$30; GPT-5.4 $2.50/$15. gpt-5.4-mini is NOT in the published table (which lists a Nano tier at $0.20/$1.25), so its price here is a deliberately conservative UPPER-bound estimate flagged price_estimate:true. `released` is best-effort (spec RECEIPT.md open question #4). `tier` drives reporting; `judge_eligible` encodes the fixed-judge policy (docs/judge-policy.html) — only the cheap Haiku judge is eligible, and NO OpenAI model is judge-eligible (the judge stays claude-haiku across providers so a cross-substrate report varies only the target model). `provider` is the two-axis provider; per-provider surface/base_url config is under `providers`. The release trigger (scripts/release-watch.js) appends auto-discovered ids here with auto_added:true.",
|
|
2
|
+
"_comment": "Driftproof model registry. The runner resolves --models ids against this list; an unknown id still runs but its receipt is marked registry:\"unregistered\" and its cost is estimated with the conservative default price in lib/models.js. Prices are STANDARD first-party USD per 1,000,000 tokens (input/output). Anthropic prices are from platform.claude.com (July 2026); Sonnet 5's introductory rate is deliberately NOT used so projections stay an upper bound. OpenAI prices are the public standard per-MTok rates as of July 2026 (see reports/phase-6-providers.md for the fetched source): GPT-5.6 Sol $5/$30, Terra $2.50/$15, Luna $1/$6; GPT-5.5 $5/$30; GPT-5.4 $2.50/$15. gpt-5.4-mini is NOT in the published table (which lists a Nano tier at $0.20/$1.25), so its price here is a deliberately conservative UPPER-bound estimate flagged price_estimate:true. `released` is best-effort (spec RECEIPT.md open question #4). `tier` drives reporting; `judge_eligible` encodes the fixed-judge policy (docs/judge-policy.html) — only the cheap Haiku judge is eligible, and NO OpenAI model is judge-eligible (the judge stays claude-haiku across providers so a cross-substrate report varies only the target model). `provider` is the two-axis provider; per-provider surface/base_url config is under `providers`. The release trigger (scripts/release-watch.js) appends auto-discovered ids here with auto_added:true. Three optional per-row annotations were added by spec-019b (2026-09-01) and are read by NO code path in lib/: `cache_read_price` is the documented cache-read rate per MTok, recorded for provenance and NOT consumed -- lib/cost.js prices input and output only, and the 019b gate asserts that it stays that way, so wiring caching into the projection turns that assertion red rather than leaving this note quietly false; `lifecycle: legacy` records that the docs moved an id to the Legacy list, and never means removed (claude-fable-5 is referenced by receipts/report-007 and receipts/report-007-rerun); `surface_verified` records that a live call reached the id on that surface on `surface_verified_at`, which is an OBSERVATION, not a routing setting -- the surface is chosen at call time by lib/provider.js from CLAUDE_PROVIDER. claude-fable-5-1 prices are sourced to specs/019b-fable-5-1-registration/evidence/docs-pricing-snapshot-2026-09-01.json.",
|
|
3
3
|
"registry_version": "1.1",
|
|
4
4
|
"provider": "anthropic",
|
|
5
5
|
"providers": {
|
|
@@ -8,7 +8,8 @@
|
|
|
8
8
|
},
|
|
9
9
|
"default_price_note": "Unregistered ids are costed at the most expensive known tier (Fable, $10/$50) — a budget guard must never under-estimate. See lib/models.js DEFAULT_PRICE.",
|
|
10
10
|
"models": [
|
|
11
|
-
{ "id": "claude-fable-5",
|
|
11
|
+
{ "id": "claude-fable-5-1", "family": "fable", "provider": "anthropic", "released": "2026-09-01", "input_price": 10.0, "output_price": 50.0, "cache_read_price": 0.25, "tier": "frontier", "judge_eligible": false, "surface_verified": "claude-cli", "surface_verified_at": "2026-09-01" },
|
|
12
|
+
{ "id": "claude-fable-5", "family": "fable", "provider": "anthropic", "released": "2026-06-09", "input_price": 10.0, "output_price": 50.0, "tier": "frontier", "judge_eligible": false, "lifecycle": "legacy" },
|
|
12
13
|
{ "id": "claude-opus-5", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
|
|
13
14
|
{ "id": "claude-opus-4-8", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
|
|
14
15
|
{ "id": "claude-opus-4-7", "family": "opus", "provider": "anthropic", "released": null, "input_price": 5.0, "output_price": 25.0, "tier": "frontier", "judge_eligible": false },
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.8.0';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
package/package.json
CHANGED
|
@@ -1,14 +1,16 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.8.0",
|
|
4
4
|
"description": "A dated proof that this skill, this hash, this model, still helps: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
|
7
7
|
"agent-skills",
|
|
8
|
-
"
|
|
9
|
-
"
|
|
8
|
+
"skill-md",
|
|
9
|
+
"skill-eval",
|
|
10
|
+
"llm-evaluation",
|
|
11
|
+
"claude-code",
|
|
10
12
|
"drift",
|
|
11
|
-
"
|
|
13
|
+
"receipts"
|
|
12
14
|
],
|
|
13
15
|
"homepage": "https://driftproofhq.com",
|
|
14
16
|
"repository": {
|
|
@@ -41,5 +43,8 @@
|
|
|
41
43
|
},
|
|
42
44
|
"engines": {
|
|
43
45
|
"node": ">=22"
|
|
46
|
+
},
|
|
47
|
+
"devDependencies": {
|
|
48
|
+
"html-validate": "^11.11.0"
|
|
44
49
|
}
|
|
45
50
|
}
|