driftproof 0.3.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/spec/RECEIPT.md CHANGED
@@ -1,5 +1,5 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
- # Driftproof receipt — spec v0.3
2
+ # Driftproof receipt — spec v0.4
3
3
 
4
4
  A **receipt** is a signed, dated record of running one agent skill's eval suite
5
5
  **with** and **without** the skill on one model version, with the judge **sampled**
@@ -11,13 +11,117 @@ The machine-readable contract is [`receipt.schema.json`](./receipt.schema.json)
11
11
  (JSON Schema, draft 2020-12). This document is the human companion. Where they
12
12
  disagree, the schema wins.
13
13
 
14
- **Versioning.** The current schema is v0.3
15
- ([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept as
14
+ **Versioning.** The current schema is v0.4
15
+ ([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept frozen as
16
+ [`receipt.v0.3.1.schema.json`](./receipt.v0.3.1.schema.json),
17
+ [`receipt.v0.3.schema.json`](./receipt.v0.3.schema.json),
16
18
  [`receipt.v0.2.schema.json`](./receipt.v0.2.schema.json) and
17
19
  [`receipt.v0.1.schema.json`](./receipt.v0.1.schema.json); the validator picks the
18
- schema by the receipt's own `schema_version`, so v0.1 and v0.2 receipts still load
19
- and validate. v0.3 is an **additive** bump — every field it adds is required in a
20
- fresh run, but older receipts are read unchanged against their own schema.
20
+ schema by the receipt's own `schema_version`, so v0.1, v0.2, v0.3 and v0.3.1
21
+ receipts still load and validate — including every receipt behind the four
22
+ published reports, which the gate asserts on each run. v0.4 is an **additive**
23
+ bump: everything it adds is optional, and no earlier receipt is invalidated.
24
+
25
+ ## What changed in v0.4 (economics — what a skill costs to run)
26
+
27
+ Reports #001–#004 could say whether a skill still helps. They could not say what
28
+ it costs, because the two CLI surfaces we run on were reporting token usage that
29
+ the runner discarded. v0.4 captures it and derives the economics from it.
30
+
31
+ - **`results.cases[].usage`** — the GENERATION call's usage for that (case, mode)
32
+ row: `{ input_tokens, output_tokens, cached_tokens, wall_ms }`. `input_tokens`
33
+ is normalized to the TOTAL input including any cached portion, because the
34
+ surfaces disagree natively (`claude -p` reports it excluding cache; `codex`
35
+ includes it). `cached_tokens` is `null` — never `0` — where a surface does not
36
+ report it. `wall_ms` is measured by the runner around the successful attempt,
37
+ so it means the same thing on every lane.
38
+ - **`results.cases[].judge_usage`** — the summed usage of the N judge calls that
39
+ graded that row. This is **measurement overhead we impose, not a cost of
40
+ running the skill**, and it is excluded from every derived value figure. The
41
+ schema pins `economics.judge_excluded` to `const: true`, so a receipt cannot
42
+ claim otherwise.
43
+ - **`run.pricing_snapshot`** — the registry prices **frozen at run time** for the
44
+ models this run touched. Every derived dollar figure is computed from this
45
+ snapshot and never from the live registry, so a published receipt does not
46
+ silently change meaning when a vendor cuts prices later.
47
+ - **`economics`** — the derived block: per-arm mean cost per call, mean input and
48
+ output tokens, median `wall_ms` with its interquartile range; and the deltas
49
+ that matter — `skill_incremental_cost_usd_per_call`,
50
+ `skill_incremental_cost_usd_per_1k_calls`, `output_tokens_delta`,
51
+ `median_wall_ms_delta`. `basis` states `metered` or `metered-equivalent`
52
+ (subscription surfaces, where actual metered spend is $0).
53
+
54
+ **Three axes, never combined.** Accuracy lift, cost, and latency are recorded and
55
+ reported separately. There is deliberately no composite "value score": the axes
56
+ have different units, different error bars, and different owners, so collapsing
57
+ them would manufacture a number no reader could trace back to evidence. The
58
+ presentation rules that follow from this — ratio framings gated on the effect
59
+ floor, latency always carrying its disclosure — are in
60
+ [`REPORT-STYLE.md`](../REPORT-STYLE.md) § "Value-axis presentation rules".
61
+
62
+ **Read the incremental figures, not the absolute ones.** Both CLI surfaces
63
+ prepend a large fixed harness preamble we do not control (observed ~25k input
64
+ tokens on `claude-cli`, ~11k on `codex`). It inflates absolute per-call cost, but
65
+ it is identical in the with-skill and baseline arms, so it cancels in every Δ.
66
+
67
+ ## Interop-additive revision (Phase 7 — receipts as an open format)
68
+
69
+ The v0.3.1 schema gained an **additive interop revision** so receipts can be
70
+ **imported** from neighboring eval tools with honest epistemics (see
71
+ [`docs/interop.md`](../docs/interop.md) and the site's `/interop.html`):
72
+
73
+ - `run.surface` and `run.judge.surface` gain **`"external"`** (the run happened
74
+ on another tool's harness); `run.transcripts` gains **`"none"`** (nothing
75
+ retained, not even hashes — only honest on an import).
76
+ - New optional **`run.source`** — provenance of a converted receipt, e.g.
77
+ `"imported/agent-skills-eval"` or `"imported/skillgrade"`.
78
+ - `skill.content_hash`, `suite.suite_hash`, and per-case `judge.rubric_hash` may
79
+ be **`null`**, per-case `generation_hash`/`judge_sample_hashes` may be
80
+ **omitted**, and `comparison.baseline_score`/`delta`/`delta_uncertainty` may be
81
+ **`null`** (a source tool with no baseline mode) — hashes and baselines are
82
+ **never fabricated**. `suite.format` may name a non-agentskills format (e.g.
83
+ `"skillgrade/eval.yaml"`).
84
+ - **The TESTED tightening** (a top-level schema conditional) makes every one of
85
+ those relaxations available **only below `TESTED`**: a receipt claiming
86
+ `verification_level: "TESTED"` must still carry the full evidence chain —
87
+ content/suite hashes, per-case generation + judge-sample hashes, a
88
+ non-`external` surface, numeric comparison — exactly as before. No previously
89
+ issued receipt is invalidated, and `TESTED` keeps its meaning.
90
+ - Structural consequence, enforced in code: **drift verdicts require `TESTED`
91
+ on both sides** — `diff` reports **NOT MEASURED** against a `DECLARED`
92
+ (imported) receipt, and a `DECLARED` receipt's badge reads *not measured*.
93
+
94
+ ## What changed from v0.3 (multi-provider — spec v0.3.1)
95
+
96
+ - **`run.provider`** (required) — the two-axis provider the target model ran on:
97
+ `"anthropic"` or `"openai"` (from the registry `provider`, else inferred from the
98
+ id). Driftproof is now two-provider: the same suites and the same fixed judge,
99
+ run on Claude and GPT substrates.
100
+ - **`run.surface`** gains the OpenAI lanes: the enum is now
101
+ `"api" | "claude-cli" | "openai-api" | "openai-cli"`. `openai-api` = a
102
+ Chat-Completions-compatible API (base_url configurable); `openai-cli` = the Codex
103
+ subscription surface (`codex exec`). `run.judge.surface` accepts the same set.
104
+ - **`run.surface_overhead_note`** (optional) — present on the `openai-cli` surface:
105
+ states the fixed Codex base-instruction preamble (~12–15k input tokens per call)
106
+ that the harness prepends and does not control. The model id is set by us via
107
+ `-m` (it is not echoed in the Codex JSONL stream).
108
+ - **Per-case `checks[]`** (optional) — deterministic post-check results, each
109
+ `{ name, kind, pass }` with `kind ∈ regex | contains | not_contains | min_length`.
110
+ Structural/regex assertions run on the model output **alongside** the judge and
111
+ reported as a **separate column** — **supplementary evidence only, never folded
112
+ into the `outcome`/band verdict.**
113
+ - **`skill.tokens`** (optional) — the estimated token size of the skill's SKILL.md
114
+ (a coarse `chars/4` proxy, not a model tokenizer), used for the **value-per-token**
115
+ axis: `delta` per 1k skill tokens. Method documented on the methodology page.
116
+ - **Per-case `case_status`** (optional; default `"ok"`) — `"failed_timeout"` marks a
117
+ case whose model/judge call persistently timed out after retries. Such a case is
118
+ **recorded WITHOUT fabricated samples/hashes** and is **excluded from the
119
+ aggregates** — a band is never invented from a case that did not complete. When
120
+ any case failed, the run is stamped **`run.status: "incomplete"`** (+
121
+ `run.failed_case_count`), and a drift/durability report **must exclude an
122
+ incomplete receipt from verdicts** (listing it honestly as "not measured").
123
+ - These are additive; a v0.3 or earlier receipt reads unchanged against its own
124
+ frozen schema.
21
125
 
22
126
  ## Design goals
23
127
 
@@ -60,7 +164,7 @@ fresh run, but older receipts are read unchanged against their own schema.
60
164
 
61
165
  ## Fields
62
166
 
63
- ### `schema_version` (string, required) — `"0.3"`.
167
+ ### `schema_version` (string, required) — `"0.3.1"`.
64
168
 
65
169
  ### `skill` (object, required)
66
170
  | field | type | notes |
@@ -68,6 +172,7 @@ fresh run, but older receipts are read unchanged against their own schema.
68
172
  | `name` | string | From SKILL.md front-matter or its H1. |
69
173
  | `version` | string | Skill's declared version. |
70
174
  | `content_hash` | sha256 hex | Over `SKILL.md` + every bundled file (the `evals/` dir is **excluded** — hashed separately as `suite_hash`), path + content, path-sorted. |
175
+ | `tokens` | integer | **v0.3.1, optional.** Estimated SKILL.md token size (coarse `chars/4` proxy) for the value-per-token axis. |
71
176
 
72
177
  ### `suite` (object, required)
73
178
  | field | type | notes |
@@ -81,7 +186,9 @@ fresh run, but older receipts are read unchanged against their own schema.
81
186
  |---|---|---|
82
187
  | `model_id` | string | Canonical model id run on. |
83
188
  | `model_release_date` | ISO date or `null` | Release date **if known**, else `null`. |
84
- | `surface` | `"api"` \| `"claude-cli"` | `api` = Messages API; `claude-cli` = spawned `claude -p` (subscription). |
189
+ | `provider` | `"anthropic"` \| `"openai"` | **v0.3.1.** The two-axis provider (registry `provider`, else inferred from the id). |
190
+ | `surface` | `"api"` \| `"claude-cli"` \| `"openai-api"` \| `"openai-cli"` | `api` = Anthropic Messages API; `claude-cli` = spawned `claude -p` (subscription); `openai-api` = Chat-Completions-compatible API (base_url configurable); `openai-cli` = Codex subscription (`codex exec`). |
191
+ | `surface_overhead_note` | string | **v0.3.1, optional.** On `openai-cli`: the fixed Codex base-instruction preamble (~12–15k input tokens/call) the harness prepends and does not control. |
85
192
  | `runner_version` | string | Runner version, for reproducibility. |
86
193
  | `date_utc` | ISO 8601 UTC | When the run finished. |
87
194
  | `registry` | `"registered"` \| `"unregistered"` | **v0.3.** Whether `model_id` resolved in the model registry (`config/models.json`). |
@@ -111,6 +218,7 @@ receipt says so.
111
218
  | `score` | number 0–1 | Alias of `mean` (kept for v0.1 readers). |
112
219
  | `threshold` | number or `null` | The case's pass threshold (or `null`). |
113
220
  | `reason` | string | One-line judge rationale (optional). |
221
+ | `checks` | array | **v0.3.1, optional.** Deterministic post-check results, each `{ name, kind, pass }` (`kind ∈ regex/contains/not_contains/min_length`). Run alongside the judge; a **separate column**, **not** folded into `outcome`/the band verdict. |
114
222
  | `judge` | object | `{ model_id, rubric_hash }` — who graded and a hash binding the grade to the exact rubric + judge system prompt. |
115
223
  - **`aggregates`** — `{ with_skill, baseline }`, each
116
224
  `{ case_count, pass_count, borderline_count?, mean_score, stddev }`. Here