driftproof 0.4.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +72 -17
- package/bin/driftproof +28 -5
- package/config.js +18 -2
- package/lib/diff.js +58 -7
- package/lib/hygiene.js +113 -0
- package/lib/judge.js +11 -3
- package/lib/provider.js +55 -25
- package/lib/receipt.js +8 -2
- package/lib/revision.js +162 -0
- package/lib/run.js +32 -5
- package/lib/stub.js +24 -3
- package/lib/usage.js +168 -0
- package/lib/value.js +502 -0
- package/package.json +2 -2
- package/spec/RECEIPT.md +52 -9
- package/spec/receipt.schema.json +703 -76
- package/spec/receipt.v0.3.1.schema.json +642 -0
package/lib/value.js
ADDED
|
@@ -0,0 +1,502 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
'use strict';
|
|
3
|
+
|
|
4
|
+
const { EFFECT_FLOOR } = require('../config');
|
|
5
|
+
|
|
6
|
+
// Economics of a run — what a skill COSTS to run, alongside whether it helps.
|
|
7
|
+
//
|
|
8
|
+
// Driftproof has always answered "does this skill still help?". Spec v0.4 adds
|
|
9
|
+
// the other half a reader needs to act: what it costs in money and in latency.
|
|
10
|
+
// The three axes are deliberately kept SEPARATE and always shown TOGETHER:
|
|
11
|
+
//
|
|
12
|
+
// 1. accuracy lift — with-skill vs baseline, with bands (the existing verdict)
|
|
13
|
+
// 2. Δ cost — metered-equivalent dollars the skill adds per call
|
|
14
|
+
// 3. Δ latency — wall-clock the skill adds, observed on the run's surface
|
|
15
|
+
//
|
|
16
|
+
// THERE IS NO COMPOSITE SCORE, ANYWHERE, EVER. Collapsing three axes with
|
|
17
|
+
// different units, different error bars, and different decision-owners into one
|
|
18
|
+
// "value score" would manufacture a number nobody can trace back to evidence —
|
|
19
|
+
// exactly the cry-wolf failure the band rule exists to prevent. A reader weighs
|
|
20
|
+
// the three against their own constraints; the report refuses to weigh them on
|
|
21
|
+
// the reader's behalf. This module therefore exports no such function, and the
|
|
22
|
+
// gate asserts none appears.
|
|
23
|
+
//
|
|
24
|
+
// RATIO FRAMINGS ARE FLOOR-GATED. A ratio like "dollars per 0.01 lift" is
|
|
25
|
+
// only meaningful when the benefit it prices is a real move. When the lift did not clear
|
|
26
|
+
// the effect floor (or its bands overlap), the ratio renders "n/a (within noise)"
|
|
27
|
+
// — never a number, however tempting the arithmetic. Dividing noise by a cost
|
|
28
|
+
// produces a precise-looking figure with no evidence under it.
|
|
29
|
+
//
|
|
30
|
+
// PRICING IS FROZEN AT RUN TIME. Registry prices change (vendors cut prices; we
|
|
31
|
+
// correct estimates). A receipt that recomputed its costs against today's
|
|
32
|
+
// registry would silently change meaning after publication, so every derived
|
|
33
|
+
// dollar figure is computed from the run's own `run.pricing_snapshot` and NEVER
|
|
34
|
+
// from the live registry. Functions here take the snapshot as an argument and
|
|
35
|
+
// have no access to lib/models — that is the enforcement, not a convention.
|
|
36
|
+
|
|
37
|
+
// What each surface's dollar figure MEANS. On a subscription CLI the actual
|
|
38
|
+
// metered spend is $0; the figure is what the same tokens would have cost on the
|
|
39
|
+
// metered API, which is the only cross-substrate-comparable basis available.
|
|
40
|
+
const COST_BASIS = {
|
|
41
|
+
metered: 'metered',
|
|
42
|
+
equivalent: 'metered-equivalent',
|
|
43
|
+
};
|
|
44
|
+
|
|
45
|
+
const LATENCY_DISCLOSURE =
|
|
46
|
+
'observed on subscription CLI surface, indicative';
|
|
47
|
+
|
|
48
|
+
const NOISE_CELL = 'n/a (within noise)';
|
|
49
|
+
|
|
50
|
+
// A floor-clearing NEGATIVE lift: the skill measurably hurt, so there is no
|
|
51
|
+
// benefit to put a price on.
|
|
52
|
+
const REGRESSED_CELL = 'n/a (skill regressed)';
|
|
53
|
+
|
|
54
|
+
// A cell with at least one separated, floor-clearing DRIVER whose AGGREGATE lift
|
|
55
|
+
// does not clear the floor. It is not "within noise" — the QA re-derivation found
|
|
56
|
+
// six such cells rendering the noise string while the Verdict basis on the same
|
|
57
|
+
// page named their drivers (QA V-4, 2026-08-19). The aggregate stays the
|
|
58
|
+
// denominator (spec 002 AC-4, DECISIONS #9), so no price is quoted, but the cell
|
|
59
|
+
// must not claim an absence of evidence that the page itself contradicts.
|
|
60
|
+
const DRIVER_ONLY_CELL = 'n/a (driver-only)';
|
|
61
|
+
|
|
62
|
+
// The unit of benefit the cost is expressed against. 0.01 is one fifth of the
|
|
63
|
+
// 0.05 effect floor — small enough to read as "a point of lift", large enough
|
|
64
|
+
// that the quoted cost stays in human range.
|
|
65
|
+
const LIFT_POINT = 0.01;
|
|
66
|
+
|
|
67
|
+
// Absolute per-call cost on a CLI surface is inflated by a fixed harness preamble
|
|
68
|
+
// (~25k input tokens on claude-cli, ~11k on codex) that we neither control nor
|
|
69
|
+
// can strip. It is identical in both arms, so it CANCELS in every Δ figure —
|
|
70
|
+
// which is why the incremental columns are the honest ones.
|
|
71
|
+
const ABSOLUTE_COST_CAVEAT =
|
|
72
|
+
'Absolute per-call cost on a CLI surface includes a fixed harness preamble '
|
|
73
|
+
+ '(observed ~25k input tokens on claude-cli, ~11k on codex) that we do not control. '
|
|
74
|
+
+ 'It is identical in the with-skill and baseline arms, so it cancels in the Δ '
|
|
75
|
+
+ 'figures — read the incremental columns, not the absolute ones.';
|
|
76
|
+
|
|
77
|
+
// Cached input is costed at the full list rate here (an over-estimate on
|
|
78
|
+
// cache-heavy surfaces). Cache-discount multipliers are vendor- and tier-specific
|
|
79
|
+
// and would be a guess; `cached_tokens` is recorded on every case so a reader can
|
|
80
|
+
// recompute under their own assumption.
|
|
81
|
+
const CACHE_PRICING_NOTE =
|
|
82
|
+
'Cached input tokens are costed at the list input rate (an over-estimate where '
|
|
83
|
+
+ 'caching is heavy); cached_tokens is recorded per call so a reader can recompute.';
|
|
84
|
+
|
|
85
|
+
// ── pricing snapshot ──────────────────────────────────────────────────────────
|
|
86
|
+
// Freeze the prices for exactly the models a run touches. `lookup` is injected
|
|
87
|
+
// (lib/models.priceForModel) so this module never reads the registry itself.
|
|
88
|
+
function buildPricingSnapshot({ models, lookup, nowIso }) {
|
|
89
|
+
const entries = {};
|
|
90
|
+
for (const m of models || []) {
|
|
91
|
+
if (!m || entries[m]) continue;
|
|
92
|
+
const p = lookup(m) || {};
|
|
93
|
+
entries[m] = {
|
|
94
|
+
input_per_mtok: Number(p.input),
|
|
95
|
+
output_per_mtok: Number(p.output),
|
|
96
|
+
registered: !!p.registered,
|
|
97
|
+
};
|
|
98
|
+
}
|
|
99
|
+
return {
|
|
100
|
+
frozen_at: nowIso,
|
|
101
|
+
source: 'config/models.json',
|
|
102
|
+
currency: 'USD',
|
|
103
|
+
models: entries,
|
|
104
|
+
note: 'Prices frozen at run time. Every derived cost in this receipt is computed '
|
|
105
|
+
+ 'from THIS snapshot, never from the live registry, so the receipt keeps its '
|
|
106
|
+
+ 'meaning when registry prices later change.',
|
|
107
|
+
};
|
|
108
|
+
}
|
|
109
|
+
|
|
110
|
+
// USD for one call's usage at snapshot prices. Returns null when either the
|
|
111
|
+
// usage or the price is unknown — a missing number is never imputed as zero.
|
|
112
|
+
function costForUsage(usage, price) {
|
|
113
|
+
if (!usage || !price) return null;
|
|
114
|
+
if (usage.input_tokens == null && usage.output_tokens == null) return null;
|
|
115
|
+
const inRate = Number(price.input_per_mtok);
|
|
116
|
+
const outRate = Number(price.output_per_mtok);
|
|
117
|
+
if (!Number.isFinite(inRate) || !Number.isFinite(outRate)) return null;
|
|
118
|
+
const usd = (Number(usage.input_tokens || 0) / 1e6) * inRate
|
|
119
|
+
+ (Number(usage.output_tokens || 0) / 1e6) * outRate;
|
|
120
|
+
return round6(usd);
|
|
121
|
+
}
|
|
122
|
+
|
|
123
|
+
// ── small stats (median / IQR, kept here so latency needs no new dependency) ──
|
|
124
|
+
function median(nums) {
|
|
125
|
+
const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number).sort((a, b) => a - b);
|
|
126
|
+
if (!xs.length) return null;
|
|
127
|
+
const mid = Math.floor(xs.length / 2);
|
|
128
|
+
return xs.length % 2 ? xs[mid] : round2((xs[mid - 1] + xs[mid]) / 2);
|
|
129
|
+
}
|
|
130
|
+
|
|
131
|
+
// Quartiles by the "exclusive halves" convention (Tukey hinges): the median is
|
|
132
|
+
// not included in either half when n is odd.
|
|
133
|
+
function quartiles(nums) {
|
|
134
|
+
const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number).sort((a, b) => a - b);
|
|
135
|
+
if (xs.length < 4) return { p25: null, p75: null, iqr: null };
|
|
136
|
+
const mid = Math.floor(xs.length / 2);
|
|
137
|
+
const lower = xs.slice(0, mid);
|
|
138
|
+
const upper = xs.slice(xs.length % 2 ? mid + 1 : mid);
|
|
139
|
+
const p25 = median(lower);
|
|
140
|
+
const p75 = median(upper);
|
|
141
|
+
return { p25, p75, iqr: p25 == null || p75 == null ? null : round2(p75 - p25) };
|
|
142
|
+
}
|
|
143
|
+
|
|
144
|
+
function meanOf(nums) {
|
|
145
|
+
const xs = (nums || []).filter((x) => Number.isFinite(Number(x))).map(Number);
|
|
146
|
+
if (!xs.length) return null;
|
|
147
|
+
return xs.reduce((a, b) => a + b, 0) / xs.length;
|
|
148
|
+
}
|
|
149
|
+
|
|
150
|
+
// ── per-arm economics ─────────────────────────────────────────────────────────
|
|
151
|
+
// One arm = all completed case rows of one mode (with_skill | baseline). Only
|
|
152
|
+
// the GENERATION usage of each row feeds this; judge usage is measurement
|
|
153
|
+
// overhead and is never passed in (see computeEconomics).
|
|
154
|
+
function armEconomics(rows, price) {
|
|
155
|
+
const usages = rows.map((r) => r.usage).filter(Boolean);
|
|
156
|
+
const costs = usages.map((u) => costForUsage(u, price)).filter((c) => c != null);
|
|
157
|
+
const wall = usages.map((u) => (u.wall_ms == null ? null : u.wall_ms)).filter((x) => x != null);
|
|
158
|
+
const q = quartiles(wall);
|
|
159
|
+
const mIn = meanOf(usages.map((u) => u.input_tokens));
|
|
160
|
+
const mOut = meanOf(usages.map((u) => u.output_tokens));
|
|
161
|
+
return {
|
|
162
|
+
call_count: usages.length,
|
|
163
|
+
mean_input_tokens: mIn == null ? null : round2(mIn),
|
|
164
|
+
mean_output_tokens: mOut == null ? null : round2(mOut),
|
|
165
|
+
mean_cost_usd_per_call: costs.length ? round6(meanOf(costs)) : null,
|
|
166
|
+
median_wall_ms: median(wall),
|
|
167
|
+
wall_ms_p25: q.p25,
|
|
168
|
+
wall_ms_p75: q.p75,
|
|
169
|
+
wall_ms_iqr: q.iqr,
|
|
170
|
+
};
|
|
171
|
+
}
|
|
172
|
+
|
|
173
|
+
// Build the receipt's `economics` block from its case rows.
|
|
174
|
+
//
|
|
175
|
+
// cases results.cases (failed_timeout rows are excluded)
|
|
176
|
+
// modelId the run's generation model — priced from the snapshot
|
|
177
|
+
// judgeModelId the judge model — priced ONLY for the excluded-overhead line
|
|
178
|
+
// pricingSnapshot run.pricing_snapshot (the frozen prices; never the registry)
|
|
179
|
+
// surface the run surface, to label the cost basis honestly
|
|
180
|
+
function computeEconomics({ cases, modelId, judgeModelId, pricingSnapshot, surface, meteredSurface }) {
|
|
181
|
+
const price = (pricingSnapshot && pricingSnapshot.models && pricingSnapshot.models[modelId]) || null;
|
|
182
|
+
const ok = (cases || []).filter((c) => c.case_status !== 'failed_timeout');
|
|
183
|
+
const withRows = ok.filter((c) => c.mode === 'with_skill');
|
|
184
|
+
const baseRows = ok.filter((c) => c.mode === 'baseline');
|
|
185
|
+
|
|
186
|
+
const w = armEconomics(withRows, price);
|
|
187
|
+
const b = armEconomics(baseRows, price);
|
|
188
|
+
|
|
189
|
+
const incrementalPerCall = (w.mean_cost_usd_per_call == null || b.mean_cost_usd_per_call == null)
|
|
190
|
+
? null : round6(w.mean_cost_usd_per_call - b.mean_cost_usd_per_call);
|
|
191
|
+
const outputDelta = (w.mean_output_tokens == null || b.mean_output_tokens == null)
|
|
192
|
+
? null : round2(w.mean_output_tokens - b.mean_output_tokens);
|
|
193
|
+
const wallDelta = (w.median_wall_ms == null || b.median_wall_ms == null)
|
|
194
|
+
? null : round2(w.median_wall_ms - b.median_wall_ms);
|
|
195
|
+
|
|
196
|
+
// Judge cost is computed for DISCLOSURE only — it is measurement overhead we
|
|
197
|
+
// impose, not a cost of running the skill, so it never touches the fields above.
|
|
198
|
+
const judgePrice = (pricingSnapshot && pricingSnapshot.models && pricingSnapshot.models[judgeModelId]) || null;
|
|
199
|
+
const judgeUsages = ok.map((c) => c.judge_usage).filter(Boolean);
|
|
200
|
+
const judgeCosts = judgeUsages.map((u) => costForUsage(u, judgePrice)).filter((c) => c != null);
|
|
201
|
+
|
|
202
|
+
return {
|
|
203
|
+
basis: meteredSurface ? COST_BASIS.metered : COST_BASIS.equivalent,
|
|
204
|
+
surface,
|
|
205
|
+
with_skill: w,
|
|
206
|
+
baseline: b,
|
|
207
|
+
skill_incremental_cost_usd_per_call: incrementalPerCall,
|
|
208
|
+
skill_incremental_cost_usd_per_1k_calls: incrementalPerCall == null ? null : round4(incrementalPerCall * 1000),
|
|
209
|
+
output_tokens_delta: outputDelta,
|
|
210
|
+
median_wall_ms_delta: wallDelta,
|
|
211
|
+
judge_excluded: true,
|
|
212
|
+
judge_overhead: {
|
|
213
|
+
note: 'Measurement overhead imposed by Driftproof, NOT a cost of running the skill. '
|
|
214
|
+
+ 'Excluded from every field above.',
|
|
215
|
+
total_cost_usd: judgeCosts.length ? round6(judgeCosts.reduce((a, c) => a + c, 0)) : null,
|
|
216
|
+
case_rows_measured: judgeUsages.length,
|
|
217
|
+
},
|
|
218
|
+
notes: {
|
|
219
|
+
absolute_cost: ABSOLUTE_COST_CAVEAT,
|
|
220
|
+
cache_pricing: CACHE_PRICING_NOTE,
|
|
221
|
+
latency: LATENCY_DISCLOSURE,
|
|
222
|
+
},
|
|
223
|
+
};
|
|
224
|
+
}
|
|
225
|
+
|
|
226
|
+
// Every dollar figure must be re-derivable from the receipt's OWN frozen pricing.
|
|
227
|
+
//
|
|
228
|
+
// Tokens are the durable measurement — what the skill actually consumed. Dollars
|
|
229
|
+
// are a dated derived view of those tokens, and provider pricing moves
|
|
230
|
+
// independently of model behaviour. So a dollar figure is only publishable if a
|
|
231
|
+
// reader can recompute it from the tokens and the frozen rates recorded alongside
|
|
232
|
+
// it; anything else is an authoritative-looking number with no derivation behind
|
|
233
|
+
// it. This is the check, used by the gate rather than trusted by convention.
|
|
234
|
+
//
|
|
235
|
+
// arms { with_skill: {mean_input_tokens, mean_output_tokens, mean_cost_usd_per_call}, baseline: {…} }
|
|
236
|
+
// rates { input_per_mtok, output_per_mtok } — from run.pricing_snapshot
|
|
237
|
+
//
|
|
238
|
+
// Cost is linear in tokens, so the mean of per-call costs equals the cost of the
|
|
239
|
+
// mean tokens exactly; the tolerance below absorbs only rounding of the stored
|
|
240
|
+
// figures, not disagreement.
|
|
241
|
+
function dollarsTraceable({ arms, rates, incrementalPer1kCalls, tolerance = 1e-5 }) {
|
|
242
|
+
const mismatches = [];
|
|
243
|
+
if (!arms || !rates || !Number.isFinite(Number(rates.input_per_mtok)) || !Number.isFinite(Number(rates.output_per_mtok))) {
|
|
244
|
+
return { traceable: false, mismatches: [{ reason: 'no frozen rates to derive from' }] };
|
|
245
|
+
}
|
|
246
|
+
const derivedArm = {};
|
|
247
|
+
for (const arm of ['with_skill', 'baseline']) {
|
|
248
|
+
const a = arms[arm];
|
|
249
|
+
if (!a) continue;
|
|
250
|
+
if (a.mean_cost_usd_per_call == null) continue; // nothing claimed, nothing to trace
|
|
251
|
+
const derived = (Number(a.mean_input_tokens || 0) / 1e6) * Number(rates.input_per_mtok)
|
|
252
|
+
+ (Number(a.mean_output_tokens || 0) / 1e6) * Number(rates.output_per_mtok);
|
|
253
|
+
derivedArm[arm] = derived;
|
|
254
|
+
const diff = Math.abs(derived - Number(a.mean_cost_usd_per_call));
|
|
255
|
+
if (!(diff <= tolerance)) {
|
|
256
|
+
mismatches.push({ arm, recorded: a.mean_cost_usd_per_call, derived: round6(derived), diff: round6(diff) });
|
|
257
|
+
}
|
|
258
|
+
}
|
|
259
|
+
|
|
260
|
+
// THE PUBLISHED FIGURE, not merely its inputs. Four consecutive approvals
|
|
261
|
+
// flagged the same gap: verifying the two per-arm means leaves the subtraction
|
|
262
|
+
// and the ×1000 scaling that actually produce the rendered
|
|
263
|
+
// `derived: $X/1k calls` outside the traced chain — so a wrong-but-non-zero
|
|
264
|
+
// increment passed everything. The increment is re-derived here from the
|
|
265
|
+
// DERIVED arm costs (never from the recorded ones), so an error anywhere in
|
|
266
|
+
// tokens → rates → arm cost → subtraction → scaling is caught.
|
|
267
|
+
if (incrementalPer1kCalls != null) {
|
|
268
|
+
if (derivedArm.with_skill == null || derivedArm.baseline == null) {
|
|
269
|
+
mismatches.push({ field: 'skill_incremental_cost_usd_per_1k_calls', reason: 'an arm cost could not be derived, so the increment cannot be traced' });
|
|
270
|
+
} else {
|
|
271
|
+
const derivedIncrement = (derivedArm.with_skill - derivedArm.baseline) * 1000;
|
|
272
|
+
const diff = Math.abs(derivedIncrement - Number(incrementalPer1kCalls));
|
|
273
|
+
// Scaled by 1000, so the tolerance scales with it.
|
|
274
|
+
if (!(diff <= tolerance * 1000)) {
|
|
275
|
+
mismatches.push({
|
|
276
|
+
field: 'skill_incremental_cost_usd_per_1k_calls',
|
|
277
|
+
recorded: incrementalPer1kCalls, derived: round6(derivedIncrement), diff: round6(diff),
|
|
278
|
+
});
|
|
279
|
+
}
|
|
280
|
+
}
|
|
281
|
+
}
|
|
282
|
+
return { traceable: mismatches.length === 0, mismatches };
|
|
283
|
+
}
|
|
284
|
+
|
|
285
|
+
// Re-derive a receipt's own generation and judge spend from its CASE ROWS —
|
|
286
|
+
// tokens × the rates frozen into that same receipt — rather than reading the
|
|
287
|
+
// aggregates it recorded. Same adjacency disease as F2: a figure read verbatim is
|
|
288
|
+
// a figure nobody checked. The judge model is taken per row from `case.judge.model_id`,
|
|
289
|
+
// so a run whose judge changed mid-flight still prices each row correctly.
|
|
290
|
+
function receiptCostBreakdown(receipt) {
|
|
291
|
+
const snap = receipt && receipt.run && receipt.run.pricing_snapshot;
|
|
292
|
+
const models = (snap && snap.models) || null;
|
|
293
|
+
if (!models) return { derivable: false, generation_usd: null, judge_usd: null, reason: 'no frozen pricing' };
|
|
294
|
+
const genRate = models[receipt.run.model_id];
|
|
295
|
+
if (!genRate) return { derivable: false, generation_usd: null, judge_usd: null, reason: 'generation model absent from the frozen snapshot' };
|
|
296
|
+
let generation = 0;
|
|
297
|
+
let judge = 0;
|
|
298
|
+
let missingJudgeRate = null;
|
|
299
|
+
for (const c of ((receipt.results || {}).cases || [])) {
|
|
300
|
+
if (c.case_status === 'failed_timeout') continue;
|
|
301
|
+
const g = costForUsage(c.usage, genRate);
|
|
302
|
+
if (g != null) generation += g;
|
|
303
|
+
if (c.judge_usage) {
|
|
304
|
+
const jid = (c.judge || {}).model_id;
|
|
305
|
+
const jRate = jid ? models[jid] : null;
|
|
306
|
+
if (!jRate) { missingJudgeRate = jid || '(unnamed judge)'; continue; }
|
|
307
|
+
const j = costForUsage(c.judge_usage, jRate);
|
|
308
|
+
if (j != null) judge += j;
|
|
309
|
+
}
|
|
310
|
+
}
|
|
311
|
+
if (missingJudgeRate) {
|
|
312
|
+
return { derivable: false, generation_usd: null, judge_usd: null, reason: `judge model ${missingJudgeRate} absent from the frozen snapshot` };
|
|
313
|
+
}
|
|
314
|
+
return { derivable: true, generation_usd: round6(generation), judge_usd: round6(judge) };
|
|
315
|
+
}
|
|
316
|
+
|
|
317
|
+
// The run's total metered-equivalent spend, derived from the receipts themselves.
|
|
318
|
+
//
|
|
319
|
+
// The obvious source for a run-total is the up-front projection — and it is the
|
|
320
|
+
// wrong one. A projection prices ASSUMED per-call token constants at whatever the
|
|
321
|
+
// registry says TODAY; it is explicitly "a rough upper bound, not an invoice". Put
|
|
322
|
+
// on a page next to a disclosure promising every dollar re-derives from frozen
|
|
323
|
+
// rates, it makes that promise false. This derives the total from what was
|
|
324
|
+
// actually measured: each receipt's per-arm mean cost (already computed at that
|
|
325
|
+
// receipt's own frozen rates) times its call count, plus the judge overhead the
|
|
326
|
+
// same receipt recorded.
|
|
327
|
+
//
|
|
328
|
+
// Judge cost is INCLUDED here and excluded from skill-value figures — different
|
|
329
|
+
// questions. "What did this run cost to perform" includes the measuring; "what
|
|
330
|
+
// does this skill cost to run" does not.
|
|
331
|
+
function runTotalFromReceipts(receipts) {
|
|
332
|
+
let generation = 0;
|
|
333
|
+
let judge = 0;
|
|
334
|
+
let counted = 0;
|
|
335
|
+
const untraceable = [];
|
|
336
|
+
for (const r of receipts || []) {
|
|
337
|
+
const e = (r && r.economics) || null;
|
|
338
|
+
const snap = r && r.run && r.run.pricing_snapshot;
|
|
339
|
+
if (!e || !snap) { untraceable.push({ model: r && r.run && r.run.model_id, reason: 'no economics or no frozen pricing' }); continue; }
|
|
340
|
+
// B1: both halves are RE-DERIVED from the receipt's case rows at its frozen
|
|
341
|
+
// rates — not read from the aggregates. The judge half in particular was
|
|
342
|
+
// previously taken verbatim from `judge_overhead.total_cost_usd`, so a
|
|
343
|
+
// write-time mispricing would have flowed into the published headline
|
|
344
|
+
// untested.
|
|
345
|
+
const bd = receiptCostBreakdown(r);
|
|
346
|
+
if (!bd.derivable) { untraceable.push({ model: r.run.model_id, reason: bd.reason }); continue; }
|
|
347
|
+
generation += bd.generation_usd;
|
|
348
|
+
judge += bd.judge_usd;
|
|
349
|
+
counted += 1;
|
|
350
|
+
}
|
|
351
|
+
// CONSTITUTION invariant 1: a published number states its verification level
|
|
352
|
+
// rather than implying one. The total inherits the WEAKEST level among the
|
|
353
|
+
// receipts it was derived from — a sum is only as verified as its least-verified
|
|
354
|
+
// term, so a single DECLARED (e.g. imported) receipt drags the headline down
|
|
355
|
+
// rather than hiding behind the TESTED majority.
|
|
356
|
+
const level = weakestVerificationLevel(receipts);
|
|
357
|
+
|
|
358
|
+
// Display triple, rounded to cents and made ARITHMETICALLY EXACT as displayed:
|
|
359
|
+
// the total is the sum of the two rounded components, so a reader adding the
|
|
360
|
+
// printed figures gets the printed total. Rounding each independently produced
|
|
361
|
+
// "$60.73 + $3.44 = $64.18" — off by a cent on a page whose disclosure promises
|
|
362
|
+
// every dollar re-derives.
|
|
363
|
+
const genDisplay = round2(generation);
|
|
364
|
+
const judgeDisplay = round2(judge);
|
|
365
|
+
return {
|
|
366
|
+
total_usd: round6(generation + judge),
|
|
367
|
+
generation_usd: round6(generation),
|
|
368
|
+
judge_usd: round6(judge),
|
|
369
|
+
display: {
|
|
370
|
+
generation_usd: genDisplay,
|
|
371
|
+
judge_usd: judgeDisplay,
|
|
372
|
+
total_usd: round2(genDisplay + judgeDisplay),
|
|
373
|
+
},
|
|
374
|
+
verification_level: level,
|
|
375
|
+
receipts_counted: counted,
|
|
376
|
+
traceable: untraceable.length === 0 && counted > 0,
|
|
377
|
+
untraceable,
|
|
378
|
+
};
|
|
379
|
+
}
|
|
380
|
+
|
|
381
|
+
// The community lattice, weakest first. A receipt with no stated level is treated
|
|
382
|
+
// as UNVERIFIED — absence is not evidence of verification.
|
|
383
|
+
const VERIFICATION_LATTICE = ['UNVERIFIED', 'DECLARED', 'TESTED', 'FORMAL'];
|
|
384
|
+
function weakestVerificationLevel(receipts) {
|
|
385
|
+
const list = (receipts || []).map((r) => (r && r.verification_level) || 'UNVERIFIED');
|
|
386
|
+
if (!list.length) return 'UNVERIFIED';
|
|
387
|
+
return list.reduce((weakest, lvl) => {
|
|
388
|
+
const a = VERIFICATION_LATTICE.indexOf(weakest);
|
|
389
|
+
const b = VERIFICATION_LATTICE.indexOf(lvl);
|
|
390
|
+
return (b < 0 || b < a) ? (b < 0 ? 'UNVERIFIED' : lvl) : weakest;
|
|
391
|
+
}, 'FORMAL');
|
|
392
|
+
}
|
|
393
|
+
|
|
394
|
+
// ── presentation rules ────────────────────────────────────────────────────────
|
|
395
|
+
// Whether a lift is allowed to appear as the numerator of a ratio: it must have
|
|
396
|
+
// cleared the effect floor AND been a real (band-separated) move. `separated` is
|
|
397
|
+
// supplied by the caller from the same per-case band rule the verdict uses — this
|
|
398
|
+
// function never re-derives a verdict.
|
|
399
|
+
function liftIsReportable({ lift, separated }) {
|
|
400
|
+
if (typeof lift !== 'number' || !Number.isFinite(lift)) return false;
|
|
401
|
+
if (separated === false) return false;
|
|
402
|
+
return Math.abs(lift) >= EFFECT_FLOOR;
|
|
403
|
+
}
|
|
404
|
+
|
|
405
|
+
// A rendered cell "looks zero" when its leading numeric token is zero — the
|
|
406
|
+
// failure this unit was changed to fix. `+0.00 /$/1k` next to a real, floor-
|
|
407
|
+
// clearing lift states the opposite of the measurement. Exported so the spec gate
|
|
408
|
+
// and the repo gate share ONE definition of the defect rather than each carrying
|
|
409
|
+
// its own regex. A cell with no numeric token at all (the noise string) is not
|
|
410
|
+
// zero-looking — it makes no numeric claim.
|
|
411
|
+
function isZeroLooking(cell) {
|
|
412
|
+
const m = String(cell == null ? '' : cell).match(/-?\d+(?:\.\d+)?/);
|
|
413
|
+
if (!m) return false;
|
|
414
|
+
return Number(m[0]) === 0;
|
|
415
|
+
}
|
|
416
|
+
|
|
417
|
+
// Cost per unit of measured benefit: dollars per 0.01 lift.
|
|
418
|
+
//
|
|
419
|
+
// cost_per_0.01_lift = incremental_cost_per_1k_calls × 0.01 ÷ |lift|
|
|
420
|
+
//
|
|
421
|
+
// This replaced "lift per dollar per 1,000 calls", which was arithmetically fine
|
|
422
|
+
// and practically useless: real incremental costs are $24–70 per 1k calls against
|
|
423
|
+
// lifts of 0.03–0.6, so every real cell rendered `+0.00` — reading as "no benefit
|
|
424
|
+
// per dollar" beside a case that had moved +0.636. Inverting the ratio puts the
|
|
425
|
+
// number in human range and asks the question a reader actually has: what does
|
|
426
|
+
// this benefit cost me? It also cannot collapse toward zero as costs rise.
|
|
427
|
+
//
|
|
428
|
+
// The floor gate is unchanged and still decides whether ANY number renders: a
|
|
429
|
+
// lift that did not clear the effect floor with separated bands returns the noise
|
|
430
|
+
// cell, never a figure. Returns a rendered STRING, never a bare number, so a
|
|
431
|
+
// caller cannot format its way around the rule.
|
|
432
|
+
function costPerLiftPoint({ lift, separated, incrementalCostPer1kCalls }) {
|
|
433
|
+
if (!liftIsReportable({ lift, separated })) {
|
|
434
|
+
// Two different absences, two different strings. `separated` means at least
|
|
435
|
+
// one case cleared the floor with non-overlapping bands; when that holds and
|
|
436
|
+
// only the aggregate falls short, the evidence exists and is listed under
|
|
437
|
+
// Verdict basis — saying "within noise" there contradicts the same page.
|
|
438
|
+
return separated ? DRIVER_ONLY_CELL : NOISE_CELL;
|
|
439
|
+
}
|
|
440
|
+
const cost = Number(incrementalCostPer1kCalls);
|
|
441
|
+
if (!Number.isFinite(cost) || cost === 0) return NOISE_CELL;
|
|
442
|
+
// A negative lift can clear the floor — the skill measurably HURT. There is no
|
|
443
|
+
// benefit to price, and a negative price would read as a refund, so the cell
|
|
444
|
+
// states the regression instead of pricing it.
|
|
445
|
+
if (lift < 0) return REGRESSED_CELL;
|
|
446
|
+
// A negative COST is the mirror case: the skill improved measured quality AND
|
|
447
|
+
// made the call cheaper. "What one unit of benefit costs" has no meaning when
|
|
448
|
+
// the benefit is free, and a negative price sorts backwards against every other
|
|
449
|
+
// cell in the column (more negative = better), which is the same reading
|
|
450
|
+
// failure `+0.00 /$/1k` had. Four cells reach this on the #005 data.
|
|
451
|
+
if (cost < 0) return `saves $${Math.abs(cost).toFixed(2)}/1k calls`;
|
|
452
|
+
const perPoint = (cost * LIFT_POINT) / Math.abs(lift);
|
|
453
|
+
if (!Number.isFinite(perPoint)) return NOISE_CELL;
|
|
454
|
+
// Two decimals down to a cent; below that, two significant figures, so a small
|
|
455
|
+
// but real cost never displays as $0.00 — the very defect this replaced.
|
|
456
|
+
const shown = perPoint >= 0.01 ? perPoint.toFixed(2) : Number(perPoint.toPrecision(2)).toString();
|
|
457
|
+
return `$${shown} per ${LIFT_POINT} lift`;
|
|
458
|
+
}
|
|
459
|
+
|
|
460
|
+
// ── T4 (AC-7) — what measurement cost in TIME ────────────────────────────────
|
|
461
|
+
// The judge is excluded from every value figure by construction, which makes it
|
|
462
|
+
// easy to forget how much of a run it is. In dollars it is the smaller half on
|
|
463
|
+
// two of three substrates; in wall-clock it dominates. Receipts carry per-call
|
|
464
|
+
// `wall_ms` on both the generation and the judge side, so this is measured, not
|
|
465
|
+
// modelled. There is no run finish stamp on a receipt — see the run record, which
|
|
466
|
+
// labels any finish as derived.
|
|
467
|
+
function runWallClockFromReceipts(receipts) {
|
|
468
|
+
let genMs = 0, judgeMs = 0, calls = 0, rows = 0;
|
|
469
|
+
for (const r of receipts || []) {
|
|
470
|
+
for (const c of ((r.results || {}).cases || [])) {
|
|
471
|
+
const u = c.usage, j = c.judge_usage;
|
|
472
|
+
if (u && u.wall_ms != null) { genMs += Number(u.wall_ms); calls++; }
|
|
473
|
+
if (j && j.wall_ms != null) { judgeMs += Number(j.wall_ms); rows++; }
|
|
474
|
+
}
|
|
475
|
+
}
|
|
476
|
+
const totalMs = genMs + judgeMs;
|
|
477
|
+
const hours = (ms) => round2(ms / 3.6e6);
|
|
478
|
+
return {
|
|
479
|
+
generation_hours: hours(genMs),
|
|
480
|
+
judge_hours: hours(judgeMs),
|
|
481
|
+
total_hours: hours(totalMs),
|
|
482
|
+
judge_share_pct: totalMs > 0 ? Math.round((100 * judgeMs) / totalMs) : null,
|
|
483
|
+
generation_calls: calls,
|
|
484
|
+
judged_rows: rows,
|
|
485
|
+
measured: totalMs > 0,
|
|
486
|
+
};
|
|
487
|
+
}
|
|
488
|
+
|
|
489
|
+
function round2(n) { return Math.round(Number(n) * 100) / 100; }
|
|
490
|
+
function round4(n) { return Math.round(Number(n) * 1e4) / 1e4; }
|
|
491
|
+
function round6(n) { return Math.round(Number(n) * 1e6) / 1e6; }
|
|
492
|
+
|
|
493
|
+
module.exports = {
|
|
494
|
+
buildPricingSnapshot, costForUsage, computeEconomics, armEconomics,
|
|
495
|
+
liftIsReportable, costPerLiftPoint, isZeroLooking, dollarsTraceable, runTotalFromReceipts,
|
|
496
|
+
runWallClockFromReceipts,
|
|
497
|
+
receiptCostBreakdown,
|
|
498
|
+
weakestVerificationLevel, VERIFICATION_LATTICE,
|
|
499
|
+
median, quartiles,
|
|
500
|
+
COST_BASIS, LATENCY_DISCLOSURE, NOISE_CELL, REGRESSED_CELL, DRIVER_ONLY_CELL, LIFT_POINT,
|
|
501
|
+
ABSOLUTE_COST_CAVEAT, CACHE_PRICING_NOTE,
|
|
502
|
+
};
|
package/package.json
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "driftproof",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit
|
|
3
|
+
"version": "0.6.0",
|
|
4
|
+
"description": "Continuous verification of agent skills: run a skill's eval suite with and without the skill across model versions, emit hash-verified dated receipts, and diff receipts into drift reports.",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
6
|
"keywords": [
|
|
7
7
|
"agent-skills",
|
package/spec/RECEIPT.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
|
-
# Driftproof receipt — spec v0.
|
|
2
|
+
# Driftproof receipt — spec v0.4
|
|
3
3
|
|
|
4
|
-
A **receipt** is a
|
|
4
|
+
A **receipt** is a **hash-verified**, dated record of running one agent skill's eval suite
|
|
5
5
|
**with** and **without** the skill on one model version, with the judge **sampled**
|
|
6
6
|
so every score carries a confidence band. Receipts are the unit of evidence
|
|
7
7
|
Driftproof produces; diffing two receipts across model releases yields a **drift
|
|
@@ -11,15 +11,58 @@ The machine-readable contract is [`receipt.schema.json`](./receipt.schema.json)
|
|
|
11
11
|
(JSON Schema, draft 2020-12). This document is the human companion. Where they
|
|
12
12
|
disagree, the schema wins.
|
|
13
13
|
|
|
14
|
-
**Versioning.** The current schema is v0.
|
|
15
|
-
([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept as
|
|
14
|
+
**Versioning.** The current schema is v0.4
|
|
15
|
+
([`receipt.schema.json`](./receipt.schema.json)). Prior schemas are kept frozen as
|
|
16
|
+
[`receipt.v0.3.1.schema.json`](./receipt.v0.3.1.schema.json),
|
|
16
17
|
[`receipt.v0.3.schema.json`](./receipt.v0.3.schema.json),
|
|
17
18
|
[`receipt.v0.2.schema.json`](./receipt.v0.2.schema.json) and
|
|
18
19
|
[`receipt.v0.1.schema.json`](./receipt.v0.1.schema.json); the validator picks the
|
|
19
|
-
schema by the receipt's own `schema_version`, so v0.1, v0.2 and v0.3
|
|
20
|
-
load and validate
|
|
21
|
-
|
|
22
|
-
|
|
20
|
+
schema by the receipt's own `schema_version`, so v0.1, v0.2, v0.3 and v0.3.1
|
|
21
|
+
receipts still load and validate — including every receipt behind the four
|
|
22
|
+
published reports, which the gate asserts on each run. v0.4 is an **additive**
|
|
23
|
+
bump: everything it adds is optional, and no earlier receipt is invalidated.
|
|
24
|
+
|
|
25
|
+
## What changed in v0.4 (economics — what a skill costs to run)
|
|
26
|
+
|
|
27
|
+
Reports #001–#004 could say whether a skill still helps. They could not say what
|
|
28
|
+
it costs, because the two CLI surfaces we run on were reporting token usage that
|
|
29
|
+
the runner discarded. v0.4 captures it and derives the economics from it.
|
|
30
|
+
|
|
31
|
+
- **`results.cases[].usage`** — the GENERATION call's usage for that (case, mode)
|
|
32
|
+
row: `{ input_tokens, output_tokens, cached_tokens, wall_ms }`. `input_tokens`
|
|
33
|
+
is normalized to the TOTAL input including any cached portion, because the
|
|
34
|
+
surfaces disagree natively (`claude -p` reports it excluding cache; `codex`
|
|
35
|
+
includes it). `cached_tokens` is `null` — never `0` — where a surface does not
|
|
36
|
+
report it. `wall_ms` is measured by the runner around the successful attempt,
|
|
37
|
+
so it means the same thing on every lane.
|
|
38
|
+
- **`results.cases[].judge_usage`** — the summed usage of the N judge calls that
|
|
39
|
+
graded that row. This is **measurement overhead we impose, not a cost of
|
|
40
|
+
running the skill**, and it is excluded from every derived value figure. The
|
|
41
|
+
schema pins `economics.judge_excluded` to `const: true`, so a receipt cannot
|
|
42
|
+
claim otherwise.
|
|
43
|
+
- **`run.pricing_snapshot`** — the registry prices **frozen at run time** for the
|
|
44
|
+
models this run touched. Every derived dollar figure is computed from this
|
|
45
|
+
snapshot and never from the live registry, so a published receipt does not
|
|
46
|
+
silently change meaning when a vendor cuts prices later.
|
|
47
|
+
- **`economics`** — the derived block: per-arm mean cost per call, mean input and
|
|
48
|
+
output tokens, median `wall_ms` with its interquartile range; and the deltas
|
|
49
|
+
that matter — `skill_incremental_cost_usd_per_call`,
|
|
50
|
+
`skill_incremental_cost_usd_per_1k_calls`, `output_tokens_delta`,
|
|
51
|
+
`median_wall_ms_delta`. `basis` states `metered` or `metered-equivalent`
|
|
52
|
+
(subscription surfaces, where actual metered spend is $0).
|
|
53
|
+
|
|
54
|
+
**Three axes, never combined.** Accuracy lift, cost, and latency are recorded and
|
|
55
|
+
reported separately. There is deliberately no composite "value score": the axes
|
|
56
|
+
have different units, different error bars, and different owners, so collapsing
|
|
57
|
+
them would manufacture a number no reader could trace back to evidence. The
|
|
58
|
+
presentation rules that follow from this — ratio framings gated on the effect
|
|
59
|
+
floor, latency always carrying its disclosure — are in
|
|
60
|
+
[`REPORT-STYLE.md`](../REPORT-STYLE.md) § "Value-axis presentation rules".
|
|
61
|
+
|
|
62
|
+
**Read the incremental figures, not the absolute ones.** Both CLI surfaces
|
|
63
|
+
prepend a large fixed harness preamble we do not control (observed ~25k input
|
|
64
|
+
tokens on `claude-cli`, ~11k on `codex`). It inflates absolute per-call cost, but
|
|
65
|
+
it is identical in the with-skill and baseline arms, so it cancels in every Δ.
|
|
23
66
|
|
|
24
67
|
## Interop-additive revision (Phase 7 — receipts as an open format)
|
|
25
68
|
|
|
@@ -121,7 +164,7 @@ The v0.3.1 schema gained an **additive interop revision** so receipts can be
|
|
|
121
164
|
|
|
122
165
|
## Fields
|
|
123
166
|
|
|
124
|
-
### `schema_version` (string, required) — `"0.
|
|
167
|
+
### `schema_version` (string, required) — `"0.4"`.
|
|
125
168
|
|
|
126
169
|
### `skill` (object, required)
|
|
127
170
|
| field | type | notes |
|