driftproof 0.10.1 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,12 +1,13 @@
1
1
  <!-- SPDX-License-Identifier: Apache-2.0 -->
2
2
  # Driftproof
3
3
 
4
- **Your skill passed. On which model? On what date?**
4
+ **Is the gap your skill makes real, or noise? Did it hold on the last model release?**
5
5
 
6
- A SKILL.md teaches an AI coding agent how you like things done. When a new model
7
- ships, the same file can stop helping, or start hurting. Driftproof re-runs the
8
- skill's tests on the new model and hands you a dated, hash-verified receipt
9
- saying whether it still helps.
6
+ A SKILL.md teaches an AI coding agent how you like things done. A skill's own
7
+ tests passing is one answer. Driftproof asks two more: whether the scores with
8
+ the skill and without it separate beyond their spread, and whether that held
9
+ when the model changed. Each answer is a dated, hash-verified receipt, and when
10
+ there were too few draws to tell, the receipt says so.
10
11
 
11
12
  [![driftproof](https://img.shields.io/endpoint?url=https://driftproofhq.com/badges/commit-message-conventions.json)](https://driftproofhq.com)
12
13
  &nbsp;— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
@@ -25,13 +26,13 @@ differ in what moves underneath the skill — or, in the value report, in which
25
26
  axes are measured; or, in the instrument re-measurement, in the instrument itself:
26
27
 
27
28
  - **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
28
- public agent skills across a current-vs-previous Sonnet release; 9 of 10 moved
29
- beyond noise. ([markdown](reports/report-001.md))
29
+ public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
30
+ a separation detected under the rule. ([markdown](reports/report-001.md))
30
31
  - **[Report #002](https://driftproofhq.com/reports/002/)** — *substrate
31
32
  durability*: the same suites across two vendors' CLIs (Claude vs Codex);
32
33
  3 durable, 4 substrate-dependent, 2 regressed, 1 no effect.
33
34
  - **[Report #003](https://driftproofhq.com/reports/003/)** — *release drift*:
34
- `claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 within noise.
35
+ `claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 with no separation detected.
35
36
  - **[Report #004](https://driftproofhq.com/reports/004/)** — *capability gap*:
36
37
  `claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
37
38
  3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
@@ -68,8 +69,9 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
68
69
  - **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
69
70
  two of Report #007's cells re-measured on `claude-fable-5-1` against
70
71
  `claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
71
- identical before the first call. **Both cells came back within noise**: the
72
- 14 cases read 0 improved, 0 regressed, 14 within noise, 0 not measured. The
72
+ identical before the first call. **Neither cell showed a separation detected
73
+ under the rule**: the 14 cases read 0 improved, 0 regressed, 14 with no separation
74
+ detected, 0 not measured, which is not evidence that nothing changed. The
73
75
  first release pair in this project where both sides are generation-sampled
74
76
  receipts, which is what makes the delta attributable to the model rather than
75
77
  to the instrument. One case sits inside the verdict on the effect floor alone
@@ -97,8 +99,10 @@ instead of a stale one.
97
99
  The hard part isn't running an eval once — it's making the number **credible
98
100
  enough to act on**. An LLM judge is noisy, so a naive score can swing run to run
99
101
  by more than the drift you're trying to detect. Driftproof's answer is to **sample
100
- the judge and report confidence bands**, and to **only claim a regression when the
101
- bands don't overlap**. A tool that cries wolf is worse than no tool.
102
+ and report bands**, each a descriptive spread (the mean plus or minus one sample
103
+ standard deviation), and to **only claim a regression when the bands don't
104
+ overlap**: a separation detected under that rule, not proof. A tool that cries wolf
105
+ is worse than no tool.
102
106
 
103
107
  **A verdict without a price is half an answer.** The same receipts price the
104
108
  marginal cost of a skill firing, and Report #005 found the dominant cost driver is
@@ -139,17 +143,24 @@ effect floor, `NO_EFFECT` when it doesn't, `REGRESSED` when the skill hurts. Eac
139
143
  case is judged several times, so it carries a **band** (`mean ± stddev`) instead of
140
144
  one fragile number, and a case is only ever called regressed/improved when its two
141
145
  bands don't overlap. The **effect floor** (0.05, one judge quantization step) is the
142
- minimum real move required before a change counts as more than noise — band
143
- separation *plus* a floor-sized delta, never either alone.
144
-
145
- **What the band does not cover.** A verdict rests on **one generation draw per
146
- arm**: the band is the spread of the *judge* re-scoring that single response, not
147
- the spread of the model writing a different one. Report #006 measured the second
148
- directly and found it larger — draw-to-draw spread up to **sd 0.186** on the 0–1
149
- scale, against judge-level noise several times smaller. So treat a surprising
150
- single-run verdict as **provisional and worth re-running** before you act on it.
151
- Generation sampling lands in the next receipt spec; until it does, this is a
152
- limit of the instrument, stated rather than implied.
146
+ minimum move required before a separation is called a change: band separation
147
+ *plus* a floor-sized delta, never either alone.
148
+
149
+ **What the band covers.** Since receipt spec v0.5 a run **samples the generation**
150
+ as well as the judge: each arm is drawn at least 3 and at most 10 times
151
+ (`GENERATION_SAMPLES_MIN` and `GENERATION_SAMPLES_MAX` in `config.js`, applied by
152
+ `lib/sampling.js`), and each draw is judged several times. A run stops at 3 draws
153
+ when the across-draw spread is no wider than the 0.05 effect floor; otherwise it
154
+ keeps drawing until two successive estimates of that spread agree to within 0.01, or
155
+ the ceiling is reached, and each case records its `stopping_reason`. So a band on a
156
+ generation-sampled receipt is the spread of the model writing different responses,
157
+ not only of the judge re-scoring one. Report #006 measured that spread directly and
158
+ found it larger than judge-level noise, draw-to-draw **sd 0.186** at the most, which
159
+ is why v0.5 samples it. A receipt from before v0.5 carries a **legacy** band, one
160
+ generation per arm re-scored by the judge, and a comparison report labels each band
161
+ `(legacy)` or `(generation)`. Either way a band describes the draws a run made, at
162
+ the sample size it used, so treat a surprising verdict as **worth re-running**
163
+ before you act on it.
153
164
 
154
165
  ### Install
155
166
 
@@ -179,6 +190,13 @@ optional and pulled in only for `CLAUDE_PROVIDER=api`. The CLI resolves its spec
179
190
  schema, and model registry from inside the package, so it runs the same from a
180
191
  global/npx install as it does in a checkout.
181
192
 
193
+ **Platforms.** Driftproof is tested on Linux and macOS. In an outside retest of
194
+ 0.10.1 on macOS the shipped gate passed 592 of 611, with 19 failed and 1 not
195
+ applicable, and all nineteen failures are in the publishing helper, which needs
196
+ GNU `realpath -m`. Windows is untested: from Node's source the plugin would not
197
+ find `npx` there, but that failure has never been observed. A CI matrix across
198
+ Linux, macOS and Windows is planned.
199
+
182
200
  ### Other commands
183
201
 
184
202
  ```bash
@@ -278,7 +296,7 @@ A receipt is the unit of evidence — one JSON document conforming to
278
296
 
279
297
  ```jsonc
280
298
  {
281
- "schema_version": "0.6",
299
+ "schema_version": "0.7",
282
300
  "skill": { "name": "commit-message-conventions", "version": "0.2.0",
283
301
  "content_hash": "…sha256 over SKILL.md + bundled files…" },
284
302
  "suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
@@ -287,7 +305,7 @@ A receipt is the unit of evidence — one JSON document conforming to
287
305
  "model_release_date": "2025-10-01",
288
306
  "provider": "anthropic",
289
307
  "surface": "claude-cli",
290
- "runner_version": "0.10.1",
308
+ "runner_version": "0.11.0",
291
309
  "date_utc": "2026-07-27T…Z",
292
310
  "registry": "registered",
293
311
  "transcripts": "hashes-only",
@@ -329,7 +347,7 @@ Key ideas:
329
347
 
330
348
  - **`content_hash` / `suite_hash`** are computed over a canonical JSON form, so the
331
349
  same skill and suite hash identically on any machine — receipts are comparable.
332
- - **Per-case `samples` / `mean` / `stddev`** give each case a confidence band. An
350
+ - **Per-case `samples` / `mean` / `stddev`** give each case a band, a descriptive spread of one sample standard deviation. An
333
351
  `outcome` of **`borderline`** means the pass threshold sits *inside* the band —
334
352
  the run can't confidently call it pass or fail.
335
353
  - **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
@@ -349,7 +367,8 @@ Key ideas:
349
367
  `driftproof diff A.json B.json` compares the with-skill bands per case across two
350
368
  receipts. The rule that keeps it honest: a **regression** (or improvement) is
351
369
  claimed **only when the two bands do not overlap**. Overlapping bands are reported
352
- as **within noise** and never counted as a regression.
370
+ as **no separation detected** at the sample size used: never counted as a
371
+ regression, and never evidence that nothing changed.
353
372
 
354
373
  ## Verification in CI (GitHub Action + badge)
355
374
 
@@ -365,7 +384,7 @@ jobs:
365
384
  runs-on: ubuntu-latest
366
385
  steps:
367
386
  - uses: actions/checkout@v4
368
- - uses: driftproofhq/driftproof@v0.10.1
387
+ - uses: driftproofhq/driftproof@v0.11.0
369
388
  with:
370
389
  skill-dir: skills/my-skill
371
390
  models: claude-haiku-4-5
package/bin/driftproof CHANGED
@@ -13,7 +13,7 @@ const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
13
13
  const { surfaceForModel, isSubscriptionSurface, resolveModel, evalUser } = require('../lib/provider');
14
14
  const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
15
15
  const { registryStatus, assertRegistered, REGISTRY_PATH } = require('../lib/models');
16
- const { verdictFromReceipt, badgeEndpoint, githubOutputLines } = require('../lib/verdict');
16
+ const { verdictFromReceipt, badgeEndpoint, githubOutputLines, drawsLine, UNDERPOWERED_LINE } = require('../lib/verdict');
17
17
  const decision = require('../lib/decision');
18
18
  const { scaffoldInit } = require('../lib/init');
19
19
 
@@ -75,7 +75,7 @@ function loadRc(skillDir) {
75
75
 
76
76
  // Flags that never take a value, so `--trusted-skill <skill-dir>` keeps the dir
77
77
  // positional instead of swallowing it as the flag's value.
78
- const BOOLEAN_FLAGS = new Set(['trusted-skill']);
78
+ const BOOLEAN_FLAGS = new Set(['trusted-skill', 'svg']);
79
79
 
80
80
  // ── THE INPUT CONTRACT (spec 026 AC-9, F5) ────────────────────────────────────
81
81
  // A COPY of action/lib.sh's RE_* rules, in the shape spec 028's fixture holds
@@ -177,7 +177,7 @@ USAGE
177
177
  [--keep-transcripts] [--out DIR] [--trusted-skill]
178
178
  ${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
179
179
  ${PROJECT_NAME} validate <receipt.json>
180
- ${PROJECT_NAME} badge <receipt.json|receipts-dir> [--models a,b] [--out FILE] [--github-output]
180
+ ${PROJECT_NAME} badge <receipt.json|receipts-dir> [--models a,b] [--out FILE] [--github-output] [--svg [--href URL]]
181
181
  ${PROJECT_NAME} decide <receipts-dir> --models a,b [--github-output] [--badge FILE]
182
182
  [--summary FILE] [--enforce] [--fail-on-regression true|false]
183
183
  ${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
@@ -515,17 +515,35 @@ function cmdBadge(positional, flags) {
515
515
  if (fs.existsSync(p) && fs.statSync(p).isDirectory()) return cmdBadgeSet(p, flags);
516
516
  const receipt = JSON.parse(fs.readFileSync(p, 'utf8'));
517
517
  refuseUnverified(receipt, p, 'badge');
518
+ if (flags.svg) {
519
+ // SPEC 036. The drawn badge: state, model, date, lift and uncertainty, with the
520
+ // machine token in its data attributes. After the same hash check as the JSON.
521
+ const svg = require('../lib/badge-svg').badgeSvg(receipt, { href: typeof flags.href === 'string' ? flags.href : null });
522
+ if (flags.out) {
523
+ const out = path.resolve(flags.out);
524
+ fs.mkdirSync(path.dirname(out), { recursive: true });
525
+ fs.writeFileSync(out, svg);
526
+ console.log(`badge written to ${flags.out} (${verdictFromReceipt(receipt).verdict}, svg)`);
527
+ } else process.stdout.write(svg);
528
+ return;
529
+ }
518
530
  if (flags['github-output']) { console.log(githubOutputLines(receipt)); return; }
519
531
  const badge = badgeEndpoint(receipt);
520
532
  const json = JSON.stringify(badge, null, 2);
533
+ const v = verdictFromReceipt(receipt);
534
+ // SPEC 035 AC-3. An UNDERPOWERED receipt says so in words and says what it would
535
+ // have needed. Beside JSON on stdout the two lines go to stderr, so a caller that
536
+ // redirects the badge still gets a file that parses.
537
+ const power = v.verdict === 'UNDERPOWERED' ? [`${UNDERPOWERED_LINE}.`, drawsLine(v.drawsNeeded)] : [];
521
538
  if (flags.out) {
522
539
  const out = path.resolve(flags.out);
523
540
  fs.mkdirSync(path.dirname(out), { recursive: true });
524
541
  fs.writeFileSync(out, json + '\n');
525
- const v = verdictFromReceipt(receipt);
526
542
  console.log(`badge written to ${flags.out} (${v.verdict}: ${badge.message}, ${badge.color})`);
543
+ for (const line of power) console.log(line);
527
544
  } else {
528
545
  console.log(json);
546
+ for (const line of power) console.error(line);
529
547
  }
530
548
  }
531
549
 
package/config.js CHANGED
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
9
9
  // Bumped whenever the runner's behaviour or receipt-generation semantics change
10
10
  // in a way that could affect results. Recorded into every receipt as
11
11
  // run.runner_version so a receipt is reproducible against a known engine.
12
- const RUNNER_VERSION = '0.10.1';
12
+ const RUNNER_VERSION = '0.11.0';
13
13
 
14
14
  // The eval format we CONSUME (we deliberately do not invent our own).
15
15
  const SUITE_FORMAT = 'agentskills.io/evals';
@@ -38,7 +38,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
38
38
  // results.aggregates.band_rule, and bands that are null where the formula
39
39
  // cannot form. NOT additive for the validator: a v0.5 receipt restamped 0.6
40
40
  // is refused. v0.5 is frozen as receipt.v0.5.schema.json.
41
- const RECEIPT_SCHEMA_VERSION = '0.6';
41
+ const RECEIPT_SCHEMA_VERSION = '0.7';
42
42
 
43
43
  // Input bounds on the skill loader and the post-checks (spec 026 AC-13, audit
44
44
  // A3/A7). lib/skill.js loaded every bundled file with no count, depth, byte,
@@ -99,11 +99,18 @@ const DEFAULT_JUDGE_SAMPLES = 5;
99
99
  // to a coarse ~0.05–0.1 grid, so a confident grade often has stddev 0 (a "point
100
100
  // band"). Two point bands that differ by a single quantum (e.g. 0.60 vs 0.64)
101
101
  // are technically non-overlapping yet represent no meaningful behaviour change.
102
- // The floor turns such statistically-separated-but-trivial moves into
103
- // "within noise (below effect floor)". 0.05 = one judge quantum. Documented in
102
+ // The floor keeps such separated but trivial moves from being called a verdict
103
+ // ("below effect floor" in the per-case table). 0.05 = one judge quantum. Documented in
104
104
  // spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
105
105
  const EFFECT_FLOOR = 0.05;
106
106
 
107
+ // How many standard errors of the difference between two means a case's resolution
108
+ // adds to its two spreads (spec 035 R-2). A case the band rule does not separate is
109
+ // UNDERPOWERED when its spreads plus this many standard errors reach the floor: at
110
+ // those spreads and draws, a true shift of the floor's size could have gone unseparated.
111
+ // Named, so a later decision moves it by one line and one amendment.
112
+ const POWER_Z = 2;
113
+
107
114
  // ── generation sampling (receipt spec v0.5) ─────────────────────────────────
108
115
  //
109
116
  // Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
@@ -131,7 +138,7 @@ const REPORT_002_GPT_MODEL = 'gpt-5.6-sol';
131
138
  const REPORT_002_JUDGE_MODEL = 'claude-haiku-4-5';
132
139
 
133
140
  // Report #003 is a RELEASE DRIFT report (Report #001's type): one provider, two
134
- // model versions, verdicts REGRESSED/IMPROVED/MIXED/WITHIN NOISE per skill under
141
+ // model versions, a per-skill verdict label (regressed, improved, mixed, or none separated) under
135
142
  // the per-case band-separation rule + effect floor. New vs its family predecessor.
136
143
  const REPORT_003_NEW_MODEL = 'claude-opus-5';
137
144
  const REPORT_003_OLD_MODEL = 'claude-opus-4-8';
@@ -157,7 +164,7 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
157
164
  module.exports = {
158
165
  PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
159
166
  SKILL_MAX_FILES, SKILL_MAX_BYTES, SKILL_MAX_DEPTH, SUITE_MAX_CASES, CASE_MAX_CHARS, CHECK_MAX_PATTERN,
160
- EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
167
+ EFFECT_FLOOR, POWER_Z, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
161
168
  GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
162
169
  REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
163
170
  REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
@@ -0,0 +1,130 @@
1
+ // SPDX-License-Identifier: Apache-2.0
2
+ 'use strict';
3
+
4
+ // The badge, drawn (spec 036).
5
+ //
6
+ // The shields endpoint JSON (lib/verdict.js badgeEndpoint) carries a label, a message
7
+ // and a colour, which is all a shields badge can hold. This draws the badge itself, so
8
+ // it can show what a reader needs at a glance: the state, the model, the date, the lift
9
+ // and its uncertainty. Four segments, left to right:
10
+ //
11
+ // driftproof | <state label> | <model> · <date> | lift <+0.000> <± 0.000 | no ±>
12
+ //
13
+ // The fourth segment carries the lift ONLY for a state that measured one. UNDERPOWERED and
14
+ // NOT_MEASURED carry the draws taken against the draws needed instead (A-036-2):
15
+ //
16
+ // driftproof | not enough draws | <model> · <date> | 3 draws of 87 needed
17
+ //
18
+ // THE WORDS ARE HUMAN; THE TOKENS ARE DATA. What a reader sees is a label (passing, not
19
+ // enough draws); the machine token the Action and the differ publish sits on the root
20
+ // element as data-verdict, with the receipt hash beside it, and never in the text.
21
+ //
22
+ // UNDERPOWERED IS DRAWN AS THE HONEST STATE, not as a failure: a calm indigo no other
23
+ // state uses, an open hatch over it, a dashed edge around it, and the words "not enough
24
+ // draws". Every state's text is white on a fill it clears 4.5:1 against, hatch included.
25
+ //
26
+ // Text widths are fixed with textLength, so a text's rendered box is the width this file
27
+ // allots it whatever font the viewer has, and it cannot spill out of its segment.
28
+
29
+ const { receiptVerdict, shortModel, receiptDrawsTaken } = require('./verdict');
30
+
31
+ const INK = '#ffffff';
32
+ const BRAND = '#24292f';
33
+ const DETAIL = '#4b5563';
34
+ const EFFECT = '#374151';
35
+ const STATE = {
36
+ PASSED: { fill: '#1a7f37' },
37
+ REGRESSED: { fill: '#b42318' },
38
+ UNDERPOWERED: { fill: '#3538cd', hatch: '#444ce7', edge: '#a4bcfd' },
39
+ NO_EFFECT: { fill: '#57606a' },
40
+ NOT_MEASURED: { fill: '#656d76' },
41
+ };
42
+ const LABELS = {
43
+ PASSED: 'passing',
44
+ REGRESSED: 'regressed',
45
+ UNDERPOWERED: 'not enough draws',
46
+ NO_EFFECT: 'no separation detected',
47
+ NOT_MEASURED: 'not measured',
48
+ };
49
+
50
+ const H = 20;
51
+ const PAD = 7;
52
+ const CHAR = 6.6;
53
+ const esc = (s) => String(s).replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;').replace(/"/g, '&quot;');
54
+ const textW = (s) => Math.round(String(s).length * CHAR * 10) / 10;
55
+
56
+ // A LIFT FIGURE IS A MEASUREMENT, so only a state that measured one may print it
57
+ // (spec 036 A-036-2). UNDERPOWERED says the receipt's draws cannot resolve a shift of the
58
+ // effect floor, and NOT_MEASURED says the receipt carries no verdict at all. A lift beside
59
+ // either is the badge answering, in its most legible segment, a question the state has just
60
+ // said it cannot answer -- and a reader who takes one figure from a badge takes that one.
61
+ // Those two states print what the reader can act on instead: the draws taken against the
62
+ // draws that would have been needed.
63
+ function effectText(receipt, v) {
64
+ const c = receipt.comparison || {};
65
+ if (v.verdict === 'UNDERPOWERED' || v.verdict === 'NOT_MEASURED') return drawsText(receipt, v);
66
+ const lift = typeof c.delta === 'number' ? `lift ${c.delta >= 0 ? '+' : '-'}${Math.abs(c.delta).toFixed(3)}` : 'no lift';
67
+ const unc = typeof c.delta_uncertainty === 'number' ? `± ${c.delta_uncertainty.toFixed(3)}`
68
+ : c.delta_uncertainty_unavailable === 'single_case' ? 'no ± (1 case)' : 'no ±';
69
+ return `${lift} ${unc}`;
70
+ }
71
+
72
+ // Draws taken against draws needed, in the badge's words. The needed count is
73
+ // receiptDrawsNeeded's, the same figure drawsLine prints in a sentence, so the badge and
74
+ // the page cannot say two different numbers. Where no count exists the badge says which
75
+ // of the two reasons it is, rather than printing a bare figure or nothing at all.
76
+ //
77
+ // BOTH NUMBERS ARE ONE CASE'S (A-036-7). Where a case sets the draws needed, the draws taken
78
+ // are that case's own, which receiptDrawsNeeded carries as `taken`. This read the smallest
79
+ // over every readable case, so a receipt whose binding case drew 4 and whose other case drew
80
+ // 2 badged "2 draws of 5 needed": a 2 from a case the 5 does not belong to. Where no case sets
81
+ // the draws needed (NOT_MEASURED carries no verdict to compute one from) there is no binding
82
+ // case, and the smallest over every readable case is the figure.
83
+ function drawsText(receipt, v) {
84
+ const d = v.drawsNeeded;
85
+ const taken = d ? d.taken : receiptDrawsTaken(receipt);
86
+ if (taken === null || taken === undefined) return 'no draws recorded';
87
+ const draws = `${taken} ${taken === 1 ? 'draw' : 'draws'}`;
88
+ if (!d) return `${draws}, needed not computed`;
89
+ if (typeof d.value === 'number') return `${draws} of ${d.value} needed`;
90
+ return `${draws}, no count at these spreads`;
91
+ }
92
+
93
+ function badgeSvg(receipt, { href = null } = {}) {
94
+ const v = receiptVerdict(receipt);
95
+ const state = STATE[v.verdict];
96
+ const label = LABELS[v.verdict];
97
+ const model = shortModel((receipt.run || {}).model_id);
98
+ const date = String((receipt.run || {}).date_utc || '').slice(0, 10);
99
+ const segs = [
100
+ { seg: 'brand', text: 'driftproof', fill: BRAND },
101
+ { seg: 'state', text: label, fill: state.fill },
102
+ { seg: 'detail', text: `${model} · ${date}`, fill: DETAIL },
103
+ { seg: 'effect', text: effectText(receipt, v), fill: EFFECT },
104
+ ];
105
+ let x = 0;
106
+ for (const s of segs) { s.w = Math.round(textW(s.text) + 2 * PAD); s.x = x; x += s.w; }
107
+ const width = x;
108
+ const title = `driftproof: ${label} on ${model}, ${date}, ${effectText(receipt, v)}`;
109
+ const L = [];
110
+ L.push(`<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${H}" viewBox="0 0 ${width} ${H}" role="img" aria-label="${esc(title)}" data-verdict="${v.verdict}" data-receipt-hash="${esc(receipt.receipt_hash || '')}">`);
111
+ L.push(`<title>${esc(title)}</title>`);
112
+ if (state.hatch) {
113
+ L.push(`<defs><pattern id="hatch" width="6" height="6" patternUnits="userSpaceOnUse" patternTransform="rotate(45)"><line x1="0" y1="0" x2="0" y2="6" stroke="${state.hatch}" stroke-width="3"/></pattern></defs>`);
114
+ }
115
+ if (href) L.push(`<a href="${esc(href)}">`);
116
+ for (const s of segs) {
117
+ const isState = s.seg === 'state';
118
+ const edge = isState && state.edge ? ` stroke="${state.edge}" stroke-width="1" stroke-dasharray="3 2"` : '';
119
+ L.push(`<rect data-seg="${s.seg}" x="${s.x}" y="0" width="${s.w}" height="${H}" fill="${s.fill}"${edge}/>`);
120
+ if (isState && state.hatch) L.push(`<rect data-seg="state-hatch" x="${s.x}" y="0" width="${s.w}" height="${H}" fill="url(#hatch)" opacity="0.35"/>`);
121
+ }
122
+ for (const s of segs) {
123
+ L.push(`<text data-seg="${s.seg}" x="${s.x + PAD}" y="14" fill="${INK}" font-family="Verdana,DejaVu Sans,sans-serif" font-size="11" textLength="${textW(s.text)}" lengthAdjust="spacingAndGlyphs">${esc(s.text)}</text>`);
124
+ }
125
+ if (href) L.push('</a>');
126
+ L.push('</svg>');
127
+ return L.join('\n') + '\n';
128
+ }
129
+
130
+ module.exports = { badgeSvg, LABELS, STATE };
package/lib/decision.js CHANGED
@@ -20,7 +20,7 @@
20
20
  const fs = require('fs');
21
21
  const path = require('path');
22
22
  const { EFFECT_FLOOR } = require('../config');
23
- const { shortModel, githubOutputEntry } = require('./verdict');
23
+ const { shortModel, githubOutputEntry, receiptVerdict, drawsLine, UNDERPOWERED_LINE } = require('./verdict');
24
24
 
25
25
  // The six decision states (spec 030 AC-4).
26
26
  //
@@ -29,7 +29,12 @@ const { shortModel, githubOutputEntry } = require('./verdict');
29
29
  // that says so. `no detected effect` is false about such a model - the effect
30
30
  // was detected and it cleared the floor - and widening it to mean "not a
31
31
  // regression" would put a measured lift and a measured nothing in one cell.
32
- const STATES = ['regression', 'refused', 'inconclusive', 'not measured', 'no detected effect', 'helped'];
32
+ //
33
+ // `underpowered` is the seventh (spec 035): the receipt measured, and its own spreads
34
+ // and draws could not resolve a shift of the effect floor. It sits after `not
35
+ // measured`, which measured nothing, and before `no detected effect`, which measured
36
+ // with a resolution below the floor.
37
+ const STATES = ['regression', 'refused', 'inconclusive', 'not measured', 'underpowered', 'no detected effect', 'helped'];
33
38
 
34
39
  // WORST FIRST. This is the single definition of the ordering; AC-1's
35
40
  // enforcement and AC-3's badge both read it, so they cannot disagree about
@@ -38,7 +43,7 @@ const STATE_ORDER = STATES.slice();
38
43
 
39
44
  // States that must never render as success on any surface the action writes -
40
45
  // the badge, the summary row, or the check title (AC-4).
41
- const NEVER_SUCCESS = ['inconclusive', 'not measured'];
46
+ const NEVER_SUCCESS = ['inconclusive', 'not measured', 'underpowered'];
42
47
 
43
48
  // States that fail the job. `refused` fails CLOSED (AC-2): a model that
44
49
  // produced no receipt is not a model that passed. `regression` fails subject to
@@ -86,11 +91,15 @@ function decisionState(receipt) {
86
91
  if (level !== 'TESTED' || kind !== 'model') return 'not measured';
87
92
  const cmp = receipt.comparison || {};
88
93
  if ((receipt.run && receipt.run.status === 'incomplete') || typeof cmp.delta !== 'number') return 'inconclusive';
89
- if (cmp.delta <= -EFFECT_FLOOR) return 'regression';
90
- if (cmp.delta >= EFFECT_FLOOR) return 'helped';
91
- return 'no detected effect';
94
+ // Spec 035: the measured states are the receipt verdict's, read per case by the
95
+ // band rule (lib/verdict.js receiptVerdict), not the aggregate delta against the floor.
96
+ // receiptVerdict refuses on the same four conditions the two clauses above split
97
+ // between `not measured` and `inconclusive`, so it cannot answer NOT_MEASURED here.
98
+ return MEASURED_STATE[receiptVerdict(receipt).verdict];
92
99
  }
93
100
 
101
+ const MEASURED_STATE = { REGRESSED: 'regression', PASSED: 'helped', UNDERPOWERED: 'underpowered', NO_EFFECT: 'no detected effect' };
102
+
94
103
  // The worst state in a set, by STATE_ORDER. An empty set has no decision, and
95
104
  // says so with null rather than defaulting to something benign.
96
105
  function worstState(states) {
@@ -182,6 +191,8 @@ function decideSet(dir, requestedCsv, { failOnRegression = true } = {}) {
182
191
  delta: receipt && receipt.comparison && typeof receipt.comparison.delta === 'number'
183
192
  ? receipt.comparison.delta : null,
184
193
  file: hit ? path.basename(hit.file) : (unreadable ? path.basename(unreadable.file) : null),
194
+ // Spec 035 AC-3: what an underpowered row would have needed, from the same reading.
195
+ drawsNeeded: state === 'underpowered' ? receiptVerdict(receipt).drawsNeeded : null,
185
196
  // The distinction is carried on the ROW, not left in a sentence, because
186
197
  // the enforcement message has to say which of the two happened (AC-2's
187
198
  // mutation class, absence-vs-unreadable).
@@ -246,6 +257,7 @@ const RENDER = {
246
257
  refused: { word: 'refused', color: 'red', marker: '\u274c' },
247
258
  inconclusive: { word: 'inconclusive', color: 'yellow', marker: '\u26a0\ufe0f' },
248
259
  'not measured': { word: 'not measured', color: 'lightgrey', marker: '\u26a0\ufe0f' },
260
+ underpowered: { word: 'not enough draws', color: 'blue', marker: '\u26a0\ufe0f' },
249
261
  'no detected effect': { word: 'no effect', color: 'lightgrey', marker: '\u2014' },
250
262
  helped: { word: 'passing', color: 'brightgreen', marker: '\u2705' },
251
263
  };
@@ -283,7 +295,7 @@ function badgeEndpointForSet(d, { label = 'driftproof' } = {}) {
283
295
  // `worst`, `regressed` and `missing` are new and carry what it could not say.
284
296
  const VERDICT_WORD = {
285
297
  regression: 'REGRESSED', refused: 'REFUSED', inconclusive: 'INCONCLUSIVE',
286
- 'not measured': 'NOT_MEASURED', 'no detected effect': 'NO_EFFECT', helped: 'PASSED',
298
+ 'not measured': 'NOT_MEASURED', underpowered: 'UNDERPOWERED', 'no detected effect': 'NO_EFFECT', helped: 'PASSED',
287
299
  };
288
300
 
289
301
  function githubOutputLines(d) {
@@ -299,6 +311,10 @@ function githubOutputLines(d) {
299
311
  ['regressed_models', d.regressed.join(',')],
300
312
  ['missing_models', d.missing.join(',')],
301
313
  ['receipt_count', d.receiptCount],
314
+ // Spec 035 AC-3: the draws the worst row would have needed, when that row is
315
+ // underpowered; `none` when no draw count reaches the floor at its spreads, and
316
+ // empty otherwise.
317
+ ['draws_needed', worstRow && worstRow.drawsNeeded ? (worstRow.drawsNeeded.value === null ? 'none' : worstRow.drawsNeeded.value) : ''],
302
318
  ].map(([k, v]) => githubOutputEntry(k, v)).join('\n');
303
319
  }
304
320
 
@@ -324,7 +340,8 @@ function summaryMarkdown(d) {
324
340
  const r = RENDER[row.state];
325
341
  const delta = row.delta === null ? 'n/a' : (row.delta >= 0 ? '+' : '') + row.delta.toFixed(3);
326
342
  const note = row.reason ? ` <br><sub>${cell(row.reason)}</sub>` : '';
327
- L.push(`| \`${cell(row.model)}\` | ${r.marker} ${row.state} | ${delta} | ${cell(row.file) || '\u2014'}${note} |`);
343
+ const power = row.state === 'underpowered' ? ` <br><sub>${cell(UNDERPOWERED_LINE)}. ${cell(drawsLine(row.drawsNeeded))}</sub>` : '';
344
+ L.push(`| \`${cell(row.model)}\` | ${r.marker} ${row.state}${power} | ${delta} | ${cell(row.file) || '\u2014'}${note} |`);
328
345
  }
329
346
  L.push('');
330
347
  L.push(`Decided over ${d.receiptCount} receipt(s) for ${d.requestedCount} requested model(s). `