driftproof 0.4.0 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/lib/value.js ADDED
@@ -0,0 +1,502 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ 'use strict';
3
+
4
+ const { EFFECT_FLOOR } = require('../config');
5
+
6
+ // Economics of a run — what a skill COSTS to run, alongside whether it helps.
7
+ //
8
+ // Driftproof has always answered "does this skill still help?". Spec v0.4 adds
9
+ // the other half a reader needs to act: what it costs in money and in latency.
10
+ // The three axes are deliberately kept SEPARATE and always shown TOGETHER:
11
+ //
12
+ // 1. accuracy lift — with-skill vs baseline, with bands (the existing verdict)
13
+ // 2. Δ cost — metered-equivalent dollars the skill adds per call
14
+ // 3. Δ latency — wall-clock the skill adds, observed on the run's surface
15
+ //
16
+ // THERE IS NO COMPOSITE SCORE, ANYWHERE, EVER. Collapsing three axes with
17
+ // different units, different error bars, and different decision-owners into one
18
+ // "value score" would manufacture a number nobody can trace back to evidence —
19
+ // exactly the cry-wolf failure the band rule exists to prevent. A reader weighs
20
+ // the three against their own constraints; the report refuses to weigh them on
21
+ // the reader's behalf. This module therefore exports no such function, and the
22
+ // gate asserts none appears.
23
+ //
24
+ // RATIO FRAMINGS ARE FLOOR-GATED. A ratio like "dollars per 0.01 lift" is
25
+ // only meaningful when the benefit it prices is a real move. When the lift did not clear
26
+ // the effect floor (or its bands overlap), the ratio renders "n/a (within noise)"
27
+ // — never a number, however tempting the arithmetic. Dividing noise by a cost
28
+ // produces a precise-looking figure with no evidence under it.
29
+ //
30
+ // PRICING IS FROZEN AT RUN TIME. Registry prices change (vendors cut prices; we
31
+ // correct estimates). A receipt that recomputed its costs against today's
32
+ // registry would silently change meaning after publication, so every derived
33
+ // dollar figure is computed from the run's own `run.pricing_snapshot` and NEVER
34
+ // from the live registry. Functions here take the snapshot as an argument and
35
+ // have no access to lib/models — that is the enforcement, not a convention.
36
+
37
+ // What each surface's dollar figure MEANS. On a subscription CLI the actual
38
+ // metered spend is $0; the figure is what the same tokens would have cost on the
39
+ // metered API, which is the only cross-substrate-comparable basis available.
40
+ const COST_BASIS = {
41
+ metered: 'metered',
42
+ equivalent: 'metered-equivalent',
43
+ };
44
+
45
+ const LATENCY_DISCLOSURE =
46
+ 'observed on subscription CLI surface, indicative';
47
+
48
+ const NOISE_CELL = 'n/a (within noise)';
49
+
50
+ // A floor-clearing NEGATIVE lift: the skill measurably hurt, so there is no
51
+ // benefit to put a price on.
52
+ const REGRESSED_CELL = 'n/a (skill regressed)';
53
+
54
+ // A cell with at least one separated, floor-clearing DRIVER whose AGGREGATE lift
55
+ // does not clear the floor. It is not "within noise" — the QA re-derivation found
56
+ // six such cells rendering the noise string while the Verdict basis on the same
57
+ // page named their drivers (QA V-4, 2026-08-19). The aggregate stays the
58
+ // denominator (spec 002 AC-4, DECISIONS #9), so no price is quoted, but the cell
59
+ // must not claim an absence of evidence that the page itself contradicts.
60
+ const DRIVER_ONLY_CELL = 'n/a (driver-only)';
61
+
62
+ // The unit of benefit the cost is expressed against. 0.01 is one fifth of the
63
+ // 0.05 effect floor — small enough to read as "a point of lift", large enough
64
+ // that the quoted cost stays in human range.
65
+ const LIFT_POINT = 0.01;
66
+
67
+ // Absolute per-call cost on a CLI surface is inflated by a fixed harness preamble
68
+ // (~25k input tokens on claude-cli, ~11k on codex) that we neither control nor
69
+ // can strip. It is identical in both arms, so it CANCELS in every Δ figure —
70
+ // which is why the incremental columns are the honest ones.
71
+ const ABSOLUTE_COST_CAVEAT =
72
+ 'Absolute per-call cost on a CLI surface includes a fixed harness preamble '
73
+ + '(observed ~25k input tokens on claude-cli, ~11k on codex) that we do not control. '
74
+ + 'It is identical in the with-skill and baseline arms, so it cancels in the Δ '
75
+ + 'figures — read the incremental columns, not the absolute ones.';
76
+
77
+ // Cached input is costed at the full list rate here (an over-estimate on
78
+ // cache-heavy surfaces). Cache-discount multipliers are vendor- and tier-specific
79
+ // and would be a guess; `cached_tokens` is recorded on every case so a reader can
80
+ // recompute under their own assumption.
81
+ const CACHE_PRICING_NOTE =
82
+ 'Cached input tokens are costed at the list input rate (an over-estimate where '
83
+ + 'caching is heavy); cached_tokens is recorded per call so a reader can recompute.';
84
+
85
+ // ── pricing snapshot ──────────────────────────────────────────────────────────
86
+ // Freeze the prices for exactly the models a run touches. `lookup` is injected
87
+ // (lib/models.priceForModel) so this module never reads the registry itself.
88
+ function buildPricingSnapshot({ models, lookup, nowIso }) {
89
+ const entries = {};
90
+ for (const m of models || []) {
91
+ if (!m || entries[m]) continue;
92
+ const p = lookup(m) || {};
93
+ entries[m] = {
94
+ input_per_mtok: Number(p.input),
95
+ output_per_mtok: Number(p.output),
96
+ registered: !!p.registered,
97
+ };
98
+ }
99
+ return {
100
+ frozen_at: nowIso,
101
+ source: 'config/models.json',
102
+ currency: 'USD',
103
+ models: entries,
104
+ note: 'Prices frozen at run time. Every derived cost in this receipt is computed '
105
+ + 'from THIS snapshot, never from the live registry, so the receipt keeps its '
106
+ + 'meaning when registry prices later change.',
107
+ };
108
+ }
109
+
110
+ // USD for one call's usage at snapshot prices. Returns null when either the
111
+ // usage or the price is unknown — a missing number is never imputed as zero.
112
+ function costForUsage(usage, price) {
113
+ if (!usage || !price) return null;
114
+ if (usage.input_tokens == null && usage.output_tokens == null) return null;
115
+ const inRate = Number(price.input_per_mtok);
116
+ const outRate = Number(price.output_per_mtok);
117
+ if (!Number.isFinite(inRate) || !Number.isFinite(outRate)) return null;
118
+ const usd = (Number(usage.input_tokens || 0) / 1e6) * inRate
119
+ + (Number(usage.output_tokens || 0) / 1e6) * outRate;
120
+ return round6(usd);
121
+ }
122
+
123
+ // ── small stats (median / IQR, kept here so latency needs no new dependency) ──
124
+ function median(nums) {
125
+ const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number).sort((a, b) => a - b);
126
+ if (!xs.length) return null;
127
+ const mid = Math.floor(xs.length / 2);
128
+ return xs.length % 2 ? xs[mid] : round2((xs[mid - 1] + xs[mid]) / 2);
129
+ }
130
+
131
+ // Quartiles by the "exclusive halves" convention (Tukey hinges): the median is
132
+ // not included in either half when n is odd.
133
+ function quartiles(nums) {
134
+ const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number).sort((a, b) => a - b);
135
+ if (xs.length < 4) return { p25: null, p75: null, iqr: null };
136
+ const mid = Math.floor(xs.length / 2);
137
+ const lower = xs.slice(0, mid);
138
+ const upper = xs.slice(xs.length % 2 ? mid + 1 : mid);
139
+ const p25 = median(lower);
140
+ const p75 = median(upper);
141
+ return { p25, p75, iqr: p25 == null || p75 == null ? null : round2(p75 - p25) };
142
+ }
143
+
144
+ function meanOf(nums) {
145
+ const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number);
146
+ if (!xs.length) return null;
147
+ return xs.reduce((a, b) => a + b, 0) / xs.length;
148
+ }
149
+
150
+ // ── per-arm economics ─────────────────────────────────────────────────────────
151
+ // One arm = all completed case rows of one mode (with_skill | baseline). Only
152
+ // the GENERATION usage of each row feeds this; judge usage is measurement
153
+ // overhead and is never passed in (see computeEconomics).
154
+ function armEconomics(rows, price) {
155
+ const usages = rows.map((r) => r.usage).filter(Boolean);
156
+ const costs = usages.map((u) => costForUsage(u, price)).filter((c) => c != null);
157
+ const wall = usages.map((u) => (u.wall_ms == null ? null : u.wall_ms)).filter((x) => x != null);
158
+ const q = quartiles(wall);
159
+ const mIn = meanOf(usages.map((u) => u.input_tokens));
160
+ const mOut = meanOf(usages.map((u) => u.output_tokens));
161
+ return {
162
+ call_count: usages.length,
163
+ mean_input_tokens: mIn == null ? null : round2(mIn),
164
+ mean_output_tokens: mOut == null ? null : round2(mOut),
165
+ mean_cost_usd_per_call: costs.length ? round6(meanOf(costs)) : null,
166
+ median_wall_ms: median(wall),
167
+ wall_ms_p25: q.p25,
168
+ wall_ms_p75: q.p75,
169
+ wall_ms_iqr: q.iqr,
170
+ };
171
+ }
172
+
173
+ // Build the receipt's `economics` block from its case rows.
174
+ //
175
+ // cases results.cases (failed_timeout rows are excluded)
176
+ // modelId the run's generation model — priced from the snapshot
177
+ // judgeModelId the judge model — priced ONLY for the excluded-overhead line
178
+ // pricingSnapshot run.pricing_snapshot (the frozen prices; never the registry)
179
+ // surface the run surface, to label the cost basis honestly
180
+ function computeEconomics({ cases, modelId, judgeModelId, pricingSnapshot, surface, meteredSurface }) {
181
+ const price = (pricingSnapshot && pricingSnapshot.models && pricingSnapshot.models[modelId]) || null;
182
+ const ok = (cases || []).filter((c) => c.case_status !== 'failed_timeout');
183
+ const withRows = ok.filter((c) => c.mode === 'with_skill');
184
+ const baseRows = ok.filter((c) => c.mode === 'baseline');
185
+
186
+ const w = armEconomics(withRows, price);
187
+ const b = armEconomics(baseRows, price);
188
+
189
+ const incrementalPerCall = (w.mean_cost_usd_per_call == null || b.mean_cost_usd_per_call == null)
190
+ ? null : round6(w.mean_cost_usd_per_call - b.mean_cost_usd_per_call);
191
+ const outputDelta = (w.mean_output_tokens == null || b.mean_output_tokens == null)
192
+ ? null : round2(w.mean_output_tokens - b.mean_output_tokens);
193
+ const wallDelta = (w.median_wall_ms == null || b.median_wall_ms == null)
194
+ ? null : round2(w.median_wall_ms - b.median_wall_ms);
195
+
196
+ // Judge cost is computed for DISCLOSURE only — it is measurement overhead we
197
+ // impose, not a cost of running the skill, so it never touches the fields above.
198
+ const judgePrice = (pricingSnapshot && pricingSnapshot.models && pricingSnapshot.models[judgeModelId]) || null;
199
+ const judgeUsages = ok.map((c) => c.judge_usage).filter(Boolean);
200
+ const judgeCosts = judgeUsages.map((u) => costForUsage(u, judgePrice)).filter((c) => c != null);
201
+
202
+ return {
203
+ basis: meteredSurface ? COST_BASIS.metered : COST_BASIS.equivalent,
204
+ surface,
205
+ with_skill: w,
206
+ baseline: b,
207
+ skill_incremental_cost_usd_per_call: incrementalPerCall,
208
+ skill_incremental_cost_usd_per_1k_calls: incrementalPerCall == null ? null : round4(incrementalPerCall * 1000),
209
+ output_tokens_delta: outputDelta,
210
+ median_wall_ms_delta: wallDelta,
211
+ judge_excluded: true,
212
+ judge_overhead: {
213
+ note: 'Measurement overhead imposed by Driftproof, NOT a cost of running the skill. '
214
+ + 'Excluded from every field above.',
215
+ total_cost_usd: judgeCosts.length ? round6(judgeCosts.reduce((a, c) => a + c, 0)) : null,
216
+ case_rows_measured: judgeUsages.length,
217
+ },
218
+ notes: {
219
+ absolute_cost: ABSOLUTE_COST_CAVEAT,
220
+ cache_pricing: CACHE_PRICING_NOTE,
221
+ latency: LATENCY_DISCLOSURE,
222
+ },
223
+ };
224
+ }
225
+
226
+ // Every dollar figure must be re-derivable from the receipt's OWN frozen pricing.
227
+ //
228
+ // Tokens are the durable measurement — what the skill actually consumed. Dollars
229
+ // are a dated derived view of those tokens, and provider pricing moves
230
+ // independently of model behaviour. So a dollar figure is only publishable if a
231
+ // reader can recompute it from the tokens and the frozen rates recorded alongside
232
+ // it; anything else is an authoritative-looking number with no derivation behind
233
+ // it. This is the check, used by the gate rather than trusted by convention.
234
+ //
235
+ // arms { with_skill: {mean_input_tokens, mean_output_tokens, mean_cost_usd_per_call}, baseline: {…} }
236
+ // rates { input_per_mtok, output_per_mtok } — from run.pricing_snapshot
237
+ //
238
+ // Cost is linear in tokens, so the mean of per-call costs equals the cost of the
239
+ // mean tokens exactly; the tolerance below absorbs only rounding of the stored
240
+ // figures, not disagreement.
241
+ function dollarsTraceable({ arms, rates, incrementalPer1kCalls, tolerance = 1e-5 }) {
242
+ const mismatches = [];
243
+ if (!arms || !rates || !Number.isFinite(Number(rates.input_per_mtok)) || !Number.isFinite(Number(rates.output_per_mtok))) {
244
+ return { traceable: false, mismatches: [{ reason: 'no frozen rates to derive from' }] };
245
+ }
246
+ const derivedArm = {};
247
+ for (const arm of ['with_skill', 'baseline']) {
248
+ const a = arms[arm];
249
+ if (!a) continue;
250
+ if (a.mean_cost_usd_per_call == null) continue; // nothing claimed, nothing to trace
251
+ const derived = (Number(a.mean_input_tokens || 0) / 1e6) * Number(rates.input_per_mtok)
252
+ + (Number(a.mean_output_tokens || 0) / 1e6) * Number(rates.output_per_mtok);
253
+ derivedArm[arm] = derived;
254
+ const diff = Math.abs(derived - Number(a.mean_cost_usd_per_call));
255
+ if (!(diff <= tolerance)) {
256
+ mismatches.push({ arm, recorded: a.mean_cost_usd_per_call, derived: round6(derived), diff: round6(diff) });
257
+ }
258
+ }
259
+
260
+ // THE PUBLISHED FIGURE, not merely its inputs. Four consecutive approvals
261
+ // flagged the same gap: verifying the two per-arm means leaves the subtraction
262
+ // and the ×1000 scaling that actually produce the rendered
263
+ // `derived: $X/1k calls` outside the traced chain — so a wrong-but-non-zero
264
+ // increment passed everything. The increment is re-derived here from the
265
+ // DERIVED arm costs (never from the recorded ones), so an error anywhere in
266
+ // tokens → rates → arm cost → subtraction → scaling is caught.
267
+ if (incrementalPer1kCalls != null) {
268
+ if (derivedArm.with_skill == null || derivedArm.baseline == null) {
269
+ mismatches.push({ field: 'skill_incremental_cost_usd_per_1k_calls', reason: 'an arm cost could not be derived, so the increment cannot be traced' });
270
+ } else {
271
+ const derivedIncrement = (derivedArm.with_skill - derivedArm.baseline) * 1000;
272
+ const diff = Math.abs(derivedIncrement - Number(incrementalPer1kCalls));
273
+ // Scaled by 1000, so the tolerance scales with it.
274
+ if (!(diff <= tolerance * 1000)) {
275
+ mismatches.push({
276
+ field: 'skill_incremental_cost_usd_per_1k_calls',
277
+ recorded: incrementalPer1kCalls, derived: round6(derivedIncrement), diff: round6(diff),
278
+ });
279
+ }
280
+ }
281
+ }
282
+ return { traceable: mismatches.length === 0, mismatches };
283
+ }
284
+
285
+ // Re-derive a receipt's own generation and judge spend from its CASE ROWS —
286
+ // tokens × the rates frozen into that same receipt — rather than reading the
287
+ // aggregates it recorded. Same adjacency disease as F2: a figure read verbatim is
288
+ // a figure nobody checked. The judge model is taken per row from `case.judge.model_id`,
289
+ // so a run whose judge changed mid-flight still prices each row correctly.
290
+ function receiptCostBreakdown(receipt) {
291
+ const snap = receipt && receipt.run && receipt.run.pricing_snapshot;
292
+ const models = (snap && snap.models) || null;
293
+ if (!models) return { derivable: false, generation_usd: null, judge_usd: null, reason: 'no frozen pricing' };
294
+ const genRate = models[receipt.run.model_id];
295
+ if (!genRate) return { derivable: false, generation_usd: null, judge_usd: null, reason: 'generation model absent from the frozen snapshot' };
296
+ let generation = 0;
297
+ let judge = 0;
298
+ let missingJudgeRate = null;
299
+ for (const c of ((receipt.results || {}).cases || [])) {
300
+ if (c.case_status === 'failed_timeout') continue;
301
+ const g = costForUsage(c.usage, genRate);
302
+ if (g != null) generation += g;
303
+ if (c.judge_usage) {
304
+ const jid = (c.judge || {}).model_id;
305
+ const jRate = jid ? models[jid] : null;
306
+ if (!jRate) { missingJudgeRate = jid || '(unnamed judge)'; continue; }
307
+ const j = costForUsage(c.judge_usage, jRate);
308
+ if (j != null) judge += j;
309
+ }
310
+ }
311
+ if (missingJudgeRate) {
312
+ return { derivable: false, generation_usd: null, judge_usd: null, reason: `judge model ${missingJudgeRate} absent from the frozen snapshot` };
313
+ }
314
+ return { derivable: true, generation_usd: round6(generation), judge_usd: round6(judge) };
315
+ }
316
+
317
+ // The run's total metered-equivalent spend, derived from the receipts themselves.
318
+ //
319
+ // The obvious source for a run-total is the up-front projection — and it is the
320
+ // wrong one. A projection prices ASSUMED per-call token constants at whatever the
321
+ // registry says TODAY; it is explicitly "a rough upper bound, not an invoice". Put
322
+ // on a page next to a disclosure promising every dollar re-derives from frozen
323
+ // rates, it makes that promise false. This derives the total from what was
324
+ // actually measured: each receipt's per-arm mean cost (already computed at that
325
+ // receipt's own frozen rates) times its call count, plus the judge overhead the
326
+ // same receipt recorded.
327
+ //
328
+ // Judge cost is INCLUDED here and excluded from skill-value figures — different
329
+ // questions. "What did this run cost to perform" includes the measuring; "what
330
+ // does this skill cost to run" does not.
331
+ function runTotalFromReceipts(receipts) {
332
+ let generation = 0;
333
+ let judge = 0;
334
+ let counted = 0;
335
+ const untraceable = [];
336
+ for (const r of receipts || []) {
337
+ const e = (r && r.economics) || null;
338
+ const snap = r && r.run && r.run.pricing_snapshot;
339
+ if (!e || !snap) { untraceable.push({ model: r && r.run && r.run.model_id, reason: 'no economics or no frozen pricing' }); continue; }
340
+ // B1: both halves are RE-DERIVED from the receipt's case rows at its frozen
341
+ // rates — not read from the aggregates. The judge half in particular was
342
+ // previously taken verbatim from `judge_overhead.total_cost_usd`, so a
343
+ // write-time mispricing would have flowed into the published headline
344
+ // untested.
345
+ const bd = receiptCostBreakdown(r);
346
+ if (!bd.derivable) { untraceable.push({ model: r.run.model_id, reason: bd.reason }); continue; }
347
+ generation += bd.generation_usd;
348
+ judge += bd.judge_usd;
349
+ counted += 1;
350
+ }
351
+ // CONSTITUTION invariant 1: a published number states its verification level
352
+ // rather than implying one. The total inherits the WEAKEST level among the
353
+ // receipts it was derived from — a sum is only as verified as its least-verified
354
+ // term, so a single DECLARED (e.g. imported) receipt drags the headline down
355
+ // rather than hiding behind the TESTED majority.
356
+ const level = weakestVerificationLevel(receipts);
357
+
358
+ // Display triple, rounded to cents and made ARITHMETICALLY EXACT as displayed:
359
+ // the total is the sum of the two rounded components, so a reader adding the
360
+ // printed figures gets the printed total. Rounding each independently produced
361
+ // "$60.73 + $3.44 = $64.18" — off by a cent on a page whose disclosure promises
362
+ // every dollar re-derives.
363
+ const genDisplay = round2(generation);
364
+ const judgeDisplay = round2(judge);
365
+ return {
366
+ total_usd: round6(generation + judge),
367
+ generation_usd: round6(generation),
368
+ judge_usd: round6(judge),
369
+ display: {
370
+ generation_usd: genDisplay,
371
+ judge_usd: judgeDisplay,
372
+ total_usd: round2(genDisplay + judgeDisplay),
373
+ },
374
+ verification_level: level,
375
+ receipts_counted: counted,
376
+ traceable: untraceable.length === 0 && counted > 0,
377
+ untraceable,
378
+ };
379
+ }
380
+
381
+ // The community lattice, weakest first. A receipt with no stated level is treated
382
+ // as UNVERIFIED — absence is not evidence of verification.
383
+ const VERIFICATION_LATTICE = ['UNVERIFIED', 'DECLARED', 'TESTED', 'FORMAL'];
384
+ function weakestVerificationLevel(receipts) {
385
+ const list = (receipts || []).map((r) => (r && r.verification_level) || 'UNVERIFIED');
386
+ if (!list.length) return 'UNVERIFIED';
387
+ return list.reduce((weakest, lvl) => {
388
+ const a = VERIFICATION_LATTICE.indexOf(weakest);
389
+ const b = VERIFICATION_LATTICE.indexOf(lvl);
390
+ return (b < 0 || b < a) ? (b < 0 ? 'UNVERIFIED' : lvl) : weakest;
391
+ }, 'FORMAL');
392
+ }
393
+
394
+ // ── presentation rules ────────────────────────────────────────────────────────
395
+ // Whether a lift is allowed to appear as the numerator of a ratio: it must have
396
+ // cleared the effect floor AND been a real (band-separated) move. `separated` is
397
+ // supplied by the caller from the same per-case band rule the verdict uses — this
398
+ // function never re-derives a verdict.
399
+ function liftIsReportable({ lift, separated }) {
400
+ if (typeof lift !== 'number' || !Number.isFinite(lift)) return false;
401
+ if (separated === false) return false;
402
+ return Math.abs(lift) >= EFFECT_FLOOR;
403
+ }
404
+
405
+ // A rendered cell "looks zero" when its leading numeric token is zero — the
406
+ // failure this unit was changed to fix. `+0.00 /$/1k` next to a real, floor-
407
+ // clearing lift states the opposite of the measurement. Exported so the spec gate
408
+ // and the repo gate share ONE definition of the defect rather than each carrying
409
+ // its own regex. A cell with no numeric token at all (the noise string) is not
410
+ // zero-looking — it makes no numeric claim.
411
+ function isZeroLooking(cell) {
412
+ const m = String(cell == null ? '' : cell).match(/-?\d+(?:\.\d+)?/);
413
+ if (!m) return false;
414
+ return Number(m[0]) === 0;
415
+ }
416
+
417
+ // Cost per unit of measured benefit: dollars per 0.01 lift.
418
+ //
419
+ // cost_per_0.01_lift = incremental_cost_per_1k_calls × 0.01 ÷ |lift|
420
+ //
421
+ // This replaced "lift per dollar per 1,000 calls", which was arithmetically fine
422
+ // and practically useless: real incremental costs are $24–70 per 1k calls against
423
+ // lifts of 0.03–0.6, so every real cell rendered `+0.00` — reading as "no benefit
424
+ // per dollar" beside a case that had moved +0.636. Inverting the ratio puts the
425
+ // number in human range and asks the question a reader actually has: what does
426
+ // this benefit cost me? It also cannot collapse toward zero as costs rise.
427
+ //
428
+ // The floor gate is unchanged and still decides whether ANY number renders: a
429
+ // lift that did not clear the effect floor with separated bands returns the noise
430
+ // cell, never a figure. Returns a rendered STRING, never a bare number, so a
431
+ // caller cannot format its way around the rule.
432
+ function costPerLiftPoint({ lift, separated, incrementalCostPer1kCalls }) {
433
+ if (!liftIsReportable({ lift, separated })) {
434
+ // Two different absences, two different strings. `separated` means at least
435
+ // one case cleared the floor with non-overlapping bands; when that holds and
436
+ // only the aggregate falls short, the evidence exists and is listed under
437
+ // Verdict basis — saying "within noise" there contradicts the same page.
438
+ return separated ? DRIVER_ONLY_CELL : NOISE_CELL;
439
+ }
440
+ const cost = Number(incrementalCostPer1kCalls);
441
+ if (!Number.isFinite(cost) || cost === 0) return NOISE_CELL;
442
+ // A negative lift can clear the floor — the skill measurably HURT. There is no
443
+ // benefit to price, and a negative price would read as a refund, so the cell
444
+ // states the regression instead of pricing it.
445
+ if (lift < 0) return REGRESSED_CELL;
446
+ // A negative COST is the mirror case: the skill improved measured quality AND
447
+ // made the call cheaper. "What one unit of benefit costs" has no meaning when
448
+ // the benefit is free, and a negative price sorts backwards against every other
449
+ // cell in the column (more negative = better), which is the same reading
450
+ // failure `+0.00 /$/1k` had. Four cells reach this on the #005 data.
451
+ if (cost < 0) return `saves $${Math.abs(cost).toFixed(2)}/1k calls`;
452
+ const perPoint = (cost * LIFT_POINT) / Math.abs(lift);
453
+ if (!Number.isFinite(perPoint)) return NOISE_CELL;
454
+ // Two decimals down to a cent; below that, two significant figures, so a small
455
+ // but real cost never displays as $0.00 — the very defect this replaced.
456
+ const shown = perPoint >= 0.01 ? perPoint.toFixed(2) : Number(perPoint.toPrecision(2)).toString();
457
+ return `$${shown} per ${LIFT_POINT} lift`;
458
+ }
459
+
460
+ // ── T4 (AC-7) — what measurement cost in TIME ────────────────────────────────
461
+ // The judge is excluded from every value figure by construction, which makes it
462
+ // easy to forget how much of a run it is. In dollars it is the smaller half on
463
+ // two of three substrates; in wall-clock it dominates. Receipts carry per-call
464
+ // `wall_ms` on both the generation and the judge side, so this is measured, not
465
+ // modelled. There is no run finish stamp on a receipt — see the run record, which
466
+ // labels any finish as derived.
467
+ function runWallClockFromReceipts(receipts) {
468
+ let genMs = 0, judgeMs = 0, calls = 0, rows = 0;
469
+ for (const r of receipts || []) {
470
+ for (const c of ((r.results || {}).cases || [])) {
471
+ const u = c.usage, j = c.judge_usage;
472
+ if (u && u.wall_ms != null) { genMs += Number(u.wall_ms); calls++; }
473
+ if (j && j.wall_ms != null) { judgeMs += Number(j.wall_ms); rows++; }
474
+ }
475
+ }
476
+ const totalMs = genMs + judgeMs;
477
+ const hours = (ms) => round2(ms / 3.6e6);
478
+ return {
479
+ generation_hours: hours(genMs),
480
+ judge_hours: hours(judgeMs),
481
+ total_hours: hours(totalMs),
482
+ judge_share_pct: totalMs > 0 ? Math.round((100 * judgeMs) / totalMs) : null,
483
+ generation_calls: calls,
484
+ judged_rows: rows,
485
+ measured: totalMs > 0,
486
+ };
487
+ }
488
+
489
+ function round2(n) { return Math.round(Number(n) * 100) / 100; }
490
+ function round4(n) { return Math.round(Number(n) * 1e4) / 1e4; }
491
+ function round6(n) { return Math.round(Number(n) * 1e6) / 1e6; }
492
+
493
+ module.exports = {
494
+ buildPricingSnapshot, costForUsage, computeEconomics, armEconomics,
495
+ liftIsReportable, costPerLiftPoint, isZeroLooking, dollarsTraceable, runTotalFromReceipts,
496
+ runWallClockFromReceipts,
497
+ receiptCostBreakdown,
498
+ weakestVerificationLevel, VERIFICATION_LATTICE,
499
+ median, quartiles,
500
+ COST_BASIS, LATENCY_DISCLOSURE, NOISE_CELL, REGRESSED_CELL, DRIVER_ONLY_CELL, LIFT_POINT,
501
+ ABSOLUTE_COST_CAVEAT, CACHE_PRICING_NOTE,
502
+ };
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "driftproof",
3
- "version": "0.4.0",
4
- "description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit signed dated receipts, and diff receipts into drift reports.",
3
+ "version": "0.6.0",
4
+ "description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
5
5
  "license": "Apache-2.0",
6
6
  "keywords": [
7
7
  "agent-skills",
package/spec/RECEIPT.md CHANGED
@@ -1,7 +1,7 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
- # Driftproof receipt — spec v0.3.1
2
+ # Driftproof receipt — spec v0.4
3
3
 
4
- A **receipt** is a signed, dated record of running one agent skill's eval suite
4
+ A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
5
5
  **with** and **without** the skill on one model version, with the judge **sampled**
6
6
  so every score carries a confidence band. Receipts are the unit of evidence
7
7
  Driftproof produces; diffing two receipts across model releases yields a **drift
@@ -11,15 +11,58 @@ The machine-readable contract is [`receipt.schema.json`](./receipt.schema.json)
11
11
  (JSON Schema, draft 2020-12). This document is the human companion. Where they
12
12
  disagree, the schema wins.
13
13
 
14
- **Versioning.** The current schema is v0.3.1
15
- ([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept as
14
+ **Versioning.** The current schema is v0.4
15
+ ([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept frozen as
16
+ [`receipt.v0.3.1.schema.json`](./receipt.v0.3.1.schema.json),
16
17
  [`receipt.v0.3.schema.json`](./receipt.v0.3.schema.json),
17
18
  [`receipt.v0.2.schema.json`](./receipt.v0.2.schema.json) and
18
19
  [`receipt.v0.1.schema.json`](./receipt.v0.1.schema.json); the validator picks the
19
- schema by the receipt's own `schema_version`, so v0.1, v0.2 and v0.3 receipts still
20
- load and validate. v0.3.1 is an **additive** bump — `run.provider` is required in a
21
- fresh run and the rest is optional, but older receipts are read unchanged against
22
- their own schema.
20
+ schema by the receipt's own `schema_version`, so v0.1, v0.2, v0.3 and v0.3.1
21
+ receipts still load and validate — including every receipt behind the four
22
+ published reports, which the gate asserts on each run. v0.4 is an **additive**
23
+ bump: everything it adds is optional, and no earlier receipt is invalidated.
24
+
25
+ ## What changed in v0.4 (economics — what a skill costs to run)
26
+
27
+ Reports #001–#004 could say whether a skill still helps. They could not say what
28
+ it costs, because the two CLI surfaces we run on were reporting token usage that
29
+ the runner discarded. v0.4 captures it and derives the economics from it.
30
+
31
+ - **`results.cases[].usage`** — the GENERATION call's usage for that (case, mode)
32
+ row: `{ input_tokens, output_tokens, cached_tokens, wall_ms }`. `input_tokens`
33
+ is normalized to the TOTAL input including any cached portion, because the
34
+ surfaces disagree natively (`claude -p` reports it excluding cache; `codex`
35
+ includes it). `cached_tokens` is `null` — never `0` — where a surface does not
36
+ report it. `wall_ms` is measured by the runner around the successful attempt,
37
+ so it means the same thing on every lane.
38
+ - **`results.cases[].judge_usage`** — the summed usage of the N judge calls that
39
+ graded that row. This is **measurement overhead we impose, not a cost of
40
+ running the skill**, and it is excluded from every derived value figure. The
41
+ schema pins `economics.judge_excluded` to `const: true`, so a receipt cannot
42
+ claim otherwise.
43
+ - **`run.pricing_snapshot`** — the registry prices **frozen at run time** for the
44
+ models this run touched. Every derived dollar figure is computed from this
45
+ snapshot and never from the live registry, so a published receipt does not
46
+ silently change meaning when a vendor cuts prices later.
47
+ - **`economics`** — the derived block: per-arm mean cost per call, mean input and
48
+ output tokens, median `wall_ms` with its interquartile range; and the deltas
49
+ that matter — `skill_incremental_cost_usd_per_call`,
50
+ `skill_incremental_cost_usd_per_1k_calls`, `output_tokens_delta`,
51
+ `median_wall_ms_delta`. `basis` states `metered` or `metered-equivalent`
52
+ (subscription surfaces, where actual metered spend is $0).
53
+
54
+ **Three axes, never combined.** Accuracy lift, cost, and latency are recorded and
55
+ reported separately. There is deliberately no composite "value score": the axes
56
+ have different units, different error bars, and different owners, so collapsing
57
+ them would manufacture a number no reader could trace back to evidence. The
58
+ presentation rules that follow from this — ratio framings gated on the effect
59
+ floor, latency always carrying its disclosure — are in
60
+ [`REPORT-STYLE.md`](../REPORT-STYLE.md) § "Value-axis presentation rules".
61
+
62
+ **Read the incremental figures, not the absolute ones.** Both CLI surfaces
63
+ prepend a large fixed harness preamble we do not control (observed ~25k input
64
+ tokens on `claude-cli`, ~11k on `codex`). It inflates absolute per-call cost, but
65
+ it is identical in the with-skill and baseline arms, so it cancels in every Δ.
23
66
 
24
67
  ## Interop-additive revision (Phase 7 — receipts as an open format)
25
68
 
@@ -121,7 +164,7 @@ The v0.3.1 schema gained an **additive interop revision** so receipts can be
121
164
 
122
165
  ## Fields
123
166
 
124
- ### `schema_version` (string, required) — `"0.3.1"`.
167
+ ### `schema_version` (string, required) — `"0.4"`.
125
168
 
126
169
  ### `skill` (object, required)
127
170
  | field | type | notes |