driftproof 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +63 -19
- package/bin/driftproof +66 -4
- package/config/models.json +14 -4
- package/config.js +55 -5
- package/lib/checks.js +50 -0
- package/lib/diff.js +28 -7
- package/lib/export.js +52 -0
- package/lib/importers.js +207 -0
- package/lib/judge.js +35 -13
- package/lib/models.js +58 -8
- package/lib/provider.js +318 -45
- package/lib/receipt.js +29 -4
- package/lib/run.js +108 -27
- package/lib/skill.js +12 -2
- package/lib/skillCost.js +31 -0
- package/lib/stub.js +24 -3
- package/lib/usage.js +168 -0
- package/lib/value.js +502 -0
- package/lib/verdict.js +10 -1
- package/package.json +1 -1
- package/spec/RECEIPT.md +116 -8
- package/spec/receipt.schema.json +809 -59
- package/spec/receipt.v0.3.1.schema.json +642 -0
- package/spec/receipt.v0.3.schema.json +211 -0
package/spec/RECEIPT.md
CHANGED
|
@@ -1,5 +1,5 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
|
-
# Driftproof receipt — spec v0.
|
|
2
|
+
# Driftproof receipt — spec v0.4
|
|
3
3
|
|
|
4
4
|
A **receipt** is a signed, dated record of running one agent skill's eval suite
|
|
5
5
|
**with** and **without** the skill on one model version, with the judge **sampled**
|
|
@@ -11,13 +11,117 @@ The machine-readable contract is [`receipt.schema.json`](./receipt.schema.json)
|
|
|
11
11
|
(JSON Schema, draft 2020-12). This document is the human companion. Where they
|
|
12
12
|
disagree, the schema wins.
|
|
13
13
|
|
|
14
|
-
**Versioning.** The current schema is v0.
|
|
15
|
-
([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept as
|
|
14
|
+
**Versioning.** The current schema is v0.4
|
|
15
|
+
([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept frozen as
|
|
16
|
+
[`receipt.v0.3.1.schema.json`](./receipt.v0.3.1.schema.json),
|
|
17
|
+
[`receipt.v0.3.schema.json`](./receipt.v0.3.schema.json),
|
|
16
18
|
[`receipt.v0.2.schema.json`](./receipt.v0.2.schema.json) and
|
|
17
19
|
[`receipt.v0.1.schema.json`](./receipt.v0.1.schema.json); the validator picks the
|
|
18
|
-
schema by the receipt's own `schema_version`, so v0.1
|
|
19
|
-
|
|
20
|
-
|
|
20
|
+
schema by the receipt's own `schema_version`, so v0.1, v0.2, v0.3 and v0.3.1
|
|
21
|
+
receipts still load and validate — including every receipt behind the four
|
|
22
|
+
published reports, which the gate asserts on each run. v0.4 is an **additive**
|
|
23
|
+
bump: everything it adds is optional, and no earlier receipt is invalidated.
|
|
24
|
+
|
|
25
|
+
## What changed in v0.4 (economics — what a skill costs to run)
|
|
26
|
+
|
|
27
|
+
Reports #001–#004 could say whether a skill still helps. They could not say what
|
|
28
|
+
it costs, because the two CLI surfaces we run on were reporting token usage that
|
|
29
|
+
the runner discarded. v0.4 captures it and derives the economics from it.
|
|
30
|
+
|
|
31
|
+
- **`results.cases[].usage`** — the GENERATION call's usage for that (case, mode)
|
|
32
|
+
row: `{ input_tokens, output_tokens, cached_tokens, wall_ms }`. `input_tokens`
|
|
33
|
+
is normalized to the TOTAL input including any cached portion, because the
|
|
34
|
+
surfaces disagree natively (`claude -p` reports it excluding cache; `codex`
|
|
35
|
+
includes it). `cached_tokens` is `null` — never `0` — where a surface does not
|
|
36
|
+
report it. `wall_ms` is measured by the runner around the successful attempt,
|
|
37
|
+
so it means the same thing on every lane.
|
|
38
|
+
- **`results.cases[].judge_usage`** — the summed usage of the N judge calls that
|
|
39
|
+
graded that row. This is **measurement overhead we impose, not a cost of
|
|
40
|
+
running the skill**, and it is excluded from every derived value figure. The
|
|
41
|
+
schema pins `economics.judge_excluded` to `const: true`, so a receipt cannot
|
|
42
|
+
claim otherwise.
|
|
43
|
+
- **`run.pricing_snapshot`** — the registry prices **frozen at run time** for the
|
|
44
|
+
models this run touched. Every derived dollar figure is computed from this
|
|
45
|
+
snapshot and never from the live registry, so a published receipt does not
|
|
46
|
+
silently change meaning when a vendor cuts prices later.
|
|
47
|
+
- **`economics`** — the derived block: per-arm mean cost per call, mean input and
|
|
48
|
+
output tokens, median `wall_ms` with its interquartile range; and the deltas
|
|
49
|
+
that matter — `skill_incremental_cost_usd_per_call`,
|
|
50
|
+
`skill_incremental_cost_usd_per_1k_calls`, `output_tokens_delta`,
|
|
51
|
+
`median_wall_ms_delta`. `basis` states `metered` or `metered-equivalent`
|
|
52
|
+
(subscription surfaces, where actual metered spend is $0).
|
|
53
|
+
|
|
54
|
+
**Three axes, never combined.** Accuracy lift, cost, and latency are recorded and
|
|
55
|
+
reported separately. There is deliberately no composite "value score": the axes
|
|
56
|
+
have different units, different error bars, and different owners, so collapsing
|
|
57
|
+
them would manufacture a number no reader could trace back to evidence. The
|
|
58
|
+
presentation rules that follow from this — ratio framings gated on the effect
|
|
59
|
+
floor, latency always carrying its disclosure — are in
|
|
60
|
+
[`REPORT-STYLE.md`](../REPORT-STYLE.md) § "Value-axis presentation rules".
|
|
61
|
+
|
|
62
|
+
**Read the incremental figures, not the absolute ones.** Both CLI surfaces
|
|
63
|
+
prepend a large fixed harness preamble we do not control (observed ~25k input
|
|
64
|
+
tokens on `claude-cli`, ~11k on `codex`). It inflates absolute per-call cost, but
|
|
65
|
+
it is identical in the with-skill and baseline arms, so it cancels in every Δ.
|
|
66
|
+
|
|
67
|
+
## Interop-additive revision (Phase 7 — receipts as an open format)
|
|
68
|
+
|
|
69
|
+
The v0.3.1 schema gained an **additive interop revision** so receipts can be
|
|
70
|
+
**imported** from neighboring eval tools with honest epistemics (see
|
|
71
|
+
[`docs/interop.md`](../docs/interop.md) and the site's `/interop.html`):
|
|
72
|
+
|
|
73
|
+
- `run.surface` and `run.judge.surface` gain **`"external"`** (the run happened
|
|
74
|
+
on another tool's harness); `run.transcripts` gains **`"none"`** (nothing
|
|
75
|
+
retained, not even hashes — only honest on an import).
|
|
76
|
+
- New optional **`run.source`** — provenance of a converted receipt, e.g.
|
|
77
|
+
`"imported/agent-skills-eval"` or `"imported/skillgrade"`.
|
|
78
|
+
- `skill.content_hash`, `suite.suite_hash`, and per-case `judge.rubric_hash` may
|
|
79
|
+
be **`null`**, per-case `generation_hash`/`judge_sample_hashes` may be
|
|
80
|
+
**omitted**, and `comparison.baseline_score`/`delta`/`delta_uncertainty` may be
|
|
81
|
+
**`null`** (a source tool with no baseline mode) — hashes and baselines are
|
|
82
|
+
**never fabricated**. `suite.format` may name a non-agentskills format (e.g.
|
|
83
|
+
`"skillgrade/eval.yaml"`).
|
|
84
|
+
- **The TESTED tightening** (a top-level schema conditional) makes every one of
|
|
85
|
+
those relaxations available **only below `TESTED`**: a receipt claiming
|
|
86
|
+
`verification_level: "TESTED"` must still carry the full evidence chain —
|
|
87
|
+
content/suite hashes, per-case generation + judge-sample hashes, a
|
|
88
|
+
non-`external` surface, numeric comparison — exactly as before. No previously
|
|
89
|
+
issued receipt is invalidated, and `TESTED` keeps its meaning.
|
|
90
|
+
- Structural consequence, enforced in code: **drift verdicts require `TESTED`
|
|
91
|
+
on both sides** — `diff` reports **NOT MEASURED** against a `DECLARED`
|
|
92
|
+
(imported) receipt, and a `DECLARED` receipt's badge reads *not measured*.
|
|
93
|
+
|
|
94
|
+
## What changed from v0.3 (multi-provider — spec v0.3.1)
|
|
95
|
+
|
|
96
|
+
- **`run.provider`** (required) — the two-axis provider the target model ran on:
|
|
97
|
+
`"anthropic"` or `"openai"` (from the registry `provider`, else inferred from the
|
|
98
|
+
id). Driftproof is now two-provider: the same suites and the same fixed judge,
|
|
99
|
+
run on Claude and GPT substrates.
|
|
100
|
+
- **`run.surface`** gains the OpenAI lanes: the enum is now
|
|
101
|
+
`"api" | "claude-cli" | "openai-api" | "openai-cli"`. `openai-api` = a
|
|
102
|
+
Chat-Completions-compatible API (base_url configurable); `openai-cli` = the Codex
|
|
103
|
+
subscription surface (`codex exec`). `run.judge.surface` accepts the same set.
|
|
104
|
+
- **`run.surface_overhead_note`** (optional) — present on the `openai-cli` surface:
|
|
105
|
+
states the fixed Codex base-instruction preamble (~12–15k input tokens per call)
|
|
106
|
+
that the harness prepends and does not control. The model id is set by us via
|
|
107
|
+
`-m` (it is not echoed in the Codex JSONL stream).
|
|
108
|
+
- **Per-case `checks[]`** (optional) — deterministic post-check results, each
|
|
109
|
+
`{ name, kind, pass }` with `kind ∈ regex | contains | not_contains | min_length`.
|
|
110
|
+
Structural/regex assertions run on the model output **alongside** the judge and
|
|
111
|
+
reported as a **separate column** — **supplementary evidence only, never folded
|
|
112
|
+
into the `outcome`/band verdict.**
|
|
113
|
+
- **`skill.tokens`** (optional) — the estimated token size of the skill's SKILL.md
|
|
114
|
+
(a coarse `chars/4` proxy, not a model tokenizer), used for the **value-per-token**
|
|
115
|
+
axis: `delta` per 1k skill tokens. Method documented on the methodology page.
|
|
116
|
+
- **Per-case `case_status`** (optional; default `"ok"`) — `"failed_timeout"` marks a
|
|
117
|
+
case whose model/judge call persistently timed out after retries. Such a case is
|
|
118
|
+
**recorded WITHOUT fabricated samples/hashes** and is **excluded from the
|
|
119
|
+
aggregates** — a band is never invented from a case that did not complete. When
|
|
120
|
+
any case failed, the run is stamped **`run.status: "incomplete"`** (+
|
|
121
|
+
`run.failed_case_count`), and a drift/durability report **must exclude an
|
|
122
|
+
incomplete receipt from verdicts** (listing it honestly as "not measured").
|
|
123
|
+
- These are additive; a v0.3 or earlier receipt reads unchanged against its own
|
|
124
|
+
frozen schema.
|
|
21
125
|
|
|
22
126
|
## Design goals
|
|
23
127
|
|
|
@@ -60,7 +164,7 @@ fresh run, but older receipts are read unchanged against their own schema.
|
|
|
60
164
|
|
|
61
165
|
## Fields
|
|
62
166
|
|
|
63
|
-
### `schema_version` (string, required) — `"0.3"`.
|
|
167
|
+
### `schema_version` (string, required) — `"0.3.1"`.
|
|
64
168
|
|
|
65
169
|
### `skill` (object, required)
|
|
66
170
|
| field | type | notes |
|
|
@@ -68,6 +172,7 @@ fresh run, but older receipts are read unchanged against their own schema.
|
|
|
68
172
|
| `name` | string | From SKILL.md front-matter or its H1. |
|
|
69
173
|
| `version` | string | Skill's declared version. |
|
|
70
174
|
| `content_hash` | sha256 hex | Over `SKILL.md` + every bundled file (the `evals/` dir is **excluded** — hashed separately as `suite_hash`), path + content, path-sorted. |
|
|
175
|
+
| `tokens` | integer | **v0.3.1, optional.** Estimated SKILL.md token size (coarse `chars/4` proxy) for the value-per-token axis. |
|
|
71
176
|
|
|
72
177
|
### `suite` (object, required)
|
|
73
178
|
| field | type | notes |
|
|
@@ -81,7 +186,9 @@ fresh run, but older receipts are read unchanged against their own schema.
|
|
|
81
186
|
|---|---|---|
|
|
82
187
|
| `model_id` | string | Canonical model id run on. |
|
|
83
188
|
| `model_release_date` | ISO date or `null` | Release date **if known**, else `null`. |
|
|
84
|
-
| `
|
|
189
|
+
| `provider` | `"anthropic"` \| `"openai"` | **v0.3.1.** The two-axis provider (registry `provider`, else inferred from the id). |
|
|
190
|
+
| `surface` | `"api"` \| `"claude-cli"` \| `"openai-api"` \| `"openai-cli"` | `api` = Anthropic Messages API; `claude-cli` = spawned `claude -p` (subscription); `openai-api` = Chat-Completions-compatible API (base_url configurable); `openai-cli` = Codex subscription (`codex exec`). |
|
|
191
|
+
| `surface_overhead_note` | string | **v0.3.1, optional.** On `openai-cli`: the fixed Codex base-instruction preamble (~12–15k input tokens/call) the harness prepends and does not control. |
|
|
85
192
|
| `runner_version` | string | Runner version, for reproducibility. |
|
|
86
193
|
| `date_utc` | ISO 8601 UTC | When the run finished. |
|
|
87
194
|
| `registry` | `"registered"` \| `"unregistered"` | **v0.3.** Whether `model_id` resolved in the model registry (`config/models.json`). |
|
|
@@ -111,6 +218,7 @@ receipt says so.
|
|
|
111
218
|
| `score` | number 0–1 | Alias of `mean` (kept for v0.1 readers). |
|
|
112
219
|
| `threshold` | number or `null` | The case's pass threshold (or `null`). |
|
|
113
220
|
| `reason` | string | One-line judge rationale (optional). |
|
|
221
|
+
| `checks` | array | **v0.3.1, optional.** Deterministic post-check results, each `{ name, kind, pass }` (`kind ∈ regex/contains/not_contains/min_length`). Run alongside the judge; a **separate column**, **not** folded into `outcome`/the band verdict. |
|
|
114
222
|
| `judge` | object | `{ model_id, rubric_hash }` — who graded and a hash binding the grade to the exact rubric + judge system prompt. |
|
|
115
223
|
- **`aggregates`** — `{ with_skill, baseline }`, each
|
|
116
224
|
`{ case_count, pass_count, borderline_count?, mean_score, stddev }`. Here
|