driftproof 0.10.1 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +47 -28
- package/bin/driftproof +22 -4
- package/config.js +13 -6
- package/lib/badge-svg.js +130 -0
- package/lib/decision.js +25 -8
- package/lib/diff.js +92 -24
- package/lib/judge.js +14 -6
- package/lib/receipt.js +6 -2
- package/lib/revision.js +20 -11
- package/lib/stats.js +17 -8
- package/lib/value.js +5 -5
- package/lib/verdict.js +197 -26
- package/package.json +1 -1
- package/spec/RECEIPT.md +60 -15
- package/spec/receipt.schema.json +2 -2
- package/spec/receipt.v0.6.schema.json +1563 -0
package/README.md
CHANGED
|
@@ -1,12 +1,13 @@
|
|
|
1
1
|
<!-- SPDX-License-Identifier: Apache-2.0 -->
|
|
2
2
|
# Driftproof
|
|
3
3
|
|
|
4
|
-
**
|
|
4
|
+
**Is the gap your skill makes real, or noise? Did it hold on the last model release?**
|
|
5
5
|
|
|
6
|
-
A SKILL.md teaches an AI coding agent how you like things done.
|
|
7
|
-
|
|
8
|
-
skill
|
|
9
|
-
|
|
6
|
+
A SKILL.md teaches an AI coding agent how you like things done. A skill's own
|
|
7
|
+
tests passing is one answer. Driftproof asks two more: whether the scores with
|
|
8
|
+
the skill and without it separate beyond their spread, and whether that held
|
|
9
|
+
when the model changed. Each answer is a dated, hash-verified receipt, and when
|
|
10
|
+
there were too few draws to tell, the receipt says so.
|
|
10
11
|
|
|
11
12
|
[](https://driftproofhq.com)
|
|
12
13
|
— live badge for the bundled `commit-message-conventions` example, generated from its own receipt.
|
|
@@ -25,13 +26,13 @@ differ in what moves underneath the skill — or, in the value report, in which
|
|
|
25
26
|
axes are measured; or, in the instrument re-measurement, in the instrument itself:
|
|
26
27
|
|
|
27
28
|
- **[Report #001](https://driftproofhq.com/reports/001/)** — *release drift*: ten
|
|
28
|
-
public agent skills across a current-vs-previous Sonnet release; 9 of 10
|
|
29
|
-
|
|
29
|
+
public agent skills across a current-vs-previous Sonnet release; 9 of 10 showed
|
|
30
|
+
a separation detected under the rule. ([markdown](reports/report-001.md))
|
|
30
31
|
- **[Report #002](https://driftproofhq.com/reports/002/)** — *substrate
|
|
31
32
|
durability*: the same suites across two vendors' CLIs (Claude vs Codex);
|
|
32
33
|
3 durable, 4 substrate-dependent, 2 regressed, 1 no effect.
|
|
33
34
|
- **[Report #003](https://driftproofhq.com/reports/003/)** — *release drift*:
|
|
34
|
-
`claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4
|
|
35
|
+
`claude-opus-4-8` → `claude-opus-5`; 4 improved, 2 regressed, 4 with no separation detected.
|
|
35
36
|
- **[Report #004](https://driftproofhq.com/reports/004/)** — *capability gap*:
|
|
36
37
|
`claude-opus-5` (flagship) vs `claude-fable-5` (frontier tier);
|
|
37
38
|
3 durable, 5 tier-dependent, 0 regressions, 2 no effect — encoded expertise
|
|
@@ -68,8 +69,9 @@ axes are measured; or, in the instrument re-measurement, in the instrument itsel
|
|
|
68
69
|
- **[Report #008](https://driftproofhq.com/reports/008/)** — *release drift*:
|
|
69
70
|
two of Report #007's cells re-measured on `claude-fable-5-1` against
|
|
70
71
|
`claude-fable-5`, with the skill `content_hash` and `suite_hash` asserted
|
|
71
|
-
identical before the first call. **
|
|
72
|
-
14 cases read 0 improved, 0 regressed, 14
|
|
72
|
+
identical before the first call. **Neither cell showed a separation detected
|
|
73
|
+
under the rule**: the 14 cases read 0 improved, 0 regressed, 14 with no separation
|
|
74
|
+
detected, 0 not measured, which is not evidence that nothing changed. The
|
|
73
75
|
first release pair in this project where both sides are generation-sampled
|
|
74
76
|
receipts, which is what makes the delta attributable to the model rather than
|
|
75
77
|
to the instrument. One case sits inside the verdict on the effect floor alone
|
|
@@ -97,8 +99,10 @@ instead of a stale one.
|
|
|
97
99
|
The hard part isn't running an eval once — it's making the number **credible
|
|
98
100
|
enough to act on**. An LLM judge is noisy, so a naive score can swing run to run
|
|
99
101
|
by more than the drift you're trying to detect. Driftproof's answer is to **sample
|
|
100
|
-
|
|
101
|
-
|
|
102
|
+
and report bands**, each a descriptive spread (the mean plus or minus one sample
|
|
103
|
+
standard deviation), and to **only claim a regression when the bands don't
|
|
104
|
+
overlap**: a separation detected under that rule, not proof. A tool that cries wolf
|
|
105
|
+
is worse than no tool.
|
|
102
106
|
|
|
103
107
|
**A verdict without a price is half an answer.** The same receipts price the
|
|
104
108
|
marginal cost of a skill firing, and Report #005 found the dominant cost driver is
|
|
@@ -139,17 +143,24 @@ effect floor, `NO_EFFECT` when it doesn't, `REGRESSED` when the skill hurts. Eac
|
|
|
139
143
|
case is judged several times, so it carries a **band** (`mean ± stddev`) instead of
|
|
140
144
|
one fragile number, and a case is only ever called regressed/improved when its two
|
|
141
145
|
bands don't overlap. The **effect floor** (0.05, one judge quantization step) is the
|
|
142
|
-
minimum
|
|
143
|
-
|
|
144
|
-
|
|
145
|
-
**What the band
|
|
146
|
-
|
|
147
|
-
|
|
148
|
-
|
|
149
|
-
|
|
150
|
-
|
|
151
|
-
|
|
152
|
-
|
|
146
|
+
minimum move required before a separation is called a change: band separation
|
|
147
|
+
*plus* a floor-sized delta, never either alone.
|
|
148
|
+
|
|
149
|
+
**What the band covers.** Since receipt spec v0.5 a run **samples the generation**
|
|
150
|
+
as well as the judge: each arm is drawn at least 3 and at most 10 times
|
|
151
|
+
(`GENERATION_SAMPLES_MIN` and `GENERATION_SAMPLES_MAX` in `config.js`, applied by
|
|
152
|
+
`lib/sampling.js`), and each draw is judged several times. A run stops at 3 draws
|
|
153
|
+
when the across-draw spread is no wider than the 0.05 effect floor; otherwise it
|
|
154
|
+
keeps drawing until two successive estimates of that spread agree to within 0.01, or
|
|
155
|
+
the ceiling is reached, and each case records its `stopping_reason`. So a band on a
|
|
156
|
+
generation-sampled receipt is the spread of the model writing different responses,
|
|
157
|
+
not only of the judge re-scoring one. Report #006 measured that spread directly and
|
|
158
|
+
found it larger than judge-level noise, draw-to-draw **sd 0.186** at the most, which
|
|
159
|
+
is why v0.5 samples it. A receipt from before v0.5 carries a **legacy** band, one
|
|
160
|
+
generation per arm re-scored by the judge, and a comparison report labels each band
|
|
161
|
+
`(legacy)` or `(generation)`. Either way a band describes the draws a run made, at
|
|
162
|
+
the sample size it used, so treat a surprising verdict as **worth re-running**
|
|
163
|
+
before you act on it.
|
|
153
164
|
|
|
154
165
|
### Install
|
|
155
166
|
|
|
@@ -179,6 +190,13 @@ optional and pulled in only for `CLAUDE_PROVIDER=api`. The CLI resolves its spec
|
|
|
179
190
|
schema, and model registry from inside the package, so it runs the same from a
|
|
180
191
|
global/npx install as it does in a checkout.
|
|
181
192
|
|
|
193
|
+
**Platforms.** Driftproof is tested on Linux and macOS. In an outside retest of
|
|
194
|
+
0.10.1 on macOS the shipped gate passed 592 of 611, with 19 failed and 1 not
|
|
195
|
+
applicable, and all nineteen failures are in the publishing helper, which needs
|
|
196
|
+
GNU `realpath -m`. Windows is untested: from Node's source the plugin would not
|
|
197
|
+
find `npx` there, but that failure has never been observed. A CI matrix across
|
|
198
|
+
Linux, macOS and Windows is planned.
|
|
199
|
+
|
|
182
200
|
### Other commands
|
|
183
201
|
|
|
184
202
|
```bash
|
|
@@ -278,7 +296,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
278
296
|
|
|
279
297
|
```jsonc
|
|
280
298
|
{
|
|
281
|
-
"schema_version": "0.
|
|
299
|
+
"schema_version": "0.7",
|
|
282
300
|
"skill": { "name": "commit-message-conventions", "version": "0.2.0",
|
|
283
301
|
"content_hash": "…sha256 over SKILL.md + bundled files…" },
|
|
284
302
|
"suite": { "format": "agentskills.io/evals", "suite_hash": "…", "case_count": 10 },
|
|
@@ -287,7 +305,7 @@ A receipt is the unit of evidence — one JSON document conforming to
|
|
|
287
305
|
"model_release_date": "2025-10-01",
|
|
288
306
|
"provider": "anthropic",
|
|
289
307
|
"surface": "claude-cli",
|
|
290
|
-
"runner_version": "0.
|
|
308
|
+
"runner_version": "0.11.0",
|
|
291
309
|
"date_utc": "2026-07-27T…Z",
|
|
292
310
|
"registry": "registered",
|
|
293
311
|
"transcripts": "hashes-only",
|
|
@@ -329,7 +347,7 @@ Key ideas:
|
|
|
329
347
|
|
|
330
348
|
- **`content_hash` / `suite_hash`** are computed over a canonical JSON form, so the
|
|
331
349
|
same skill and suite hash identically on any machine — receipts are comparable.
|
|
332
|
-
- **Per-case `samples` / `mean` / `stddev`** give each case a
|
|
350
|
+
- **Per-case `samples` / `mean` / `stddev`** give each case a band, a descriptive spread of one sample standard deviation. An
|
|
333
351
|
`outcome` of **`borderline`** means the pass threshold sits *inside* the band —
|
|
334
352
|
the run can't confidently call it pass or fail.
|
|
335
353
|
- **`delta_uncertainty`** is the combined band on the with-skill-vs-baseline lift.
|
|
@@ -349,7 +367,8 @@ Key ideas:
|
|
|
349
367
|
`driftproof diff A.json B.json` compares the with-skill bands per case across two
|
|
350
368
|
receipts. The rule that keeps it honest: a **regression** (or improvement) is
|
|
351
369
|
claimed **only when the two bands do not overlap**. Overlapping bands are reported
|
|
352
|
-
as **
|
|
370
|
+
as **no separation detected** at the sample size used: never counted as a
|
|
371
|
+
regression, and never evidence that nothing changed.
|
|
353
372
|
|
|
354
373
|
## Verification in CI (GitHub Action + badge)
|
|
355
374
|
|
|
@@ -365,7 +384,7 @@ jobs:
|
|
|
365
384
|
runs-on: ubuntu-latest
|
|
366
385
|
steps:
|
|
367
386
|
- uses: actions/checkout@v4
|
|
368
|
-
- uses: driftproofhq/driftproof@v0.
|
|
387
|
+
- uses: driftproofhq/driftproof@v0.11.0
|
|
369
388
|
with:
|
|
370
389
|
skill-dir: skills/my-skill
|
|
371
390
|
models: claude-haiku-4-5
|
package/bin/driftproof
CHANGED
|
@@ -13,7 +13,7 @@ const { buildDriftReport, revisionPairProblem } = require('../lib/diff');
|
|
|
13
13
|
const { surfaceForModel, isSubscriptionSurface, resolveModel, evalUser } = require('../lib/provider');
|
|
14
14
|
const { estimateRunCostUSD, BudgetTracker } = require('../lib/cost');
|
|
15
15
|
const { registryStatus, assertRegistered, REGISTRY_PATH } = require('../lib/models');
|
|
16
|
-
const { verdictFromReceipt, badgeEndpoint, githubOutputLines } = require('../lib/verdict');
|
|
16
|
+
const { verdictFromReceipt, badgeEndpoint, githubOutputLines, drawsLine, UNDERPOWERED_LINE } = require('../lib/verdict');
|
|
17
17
|
const decision = require('../lib/decision');
|
|
18
18
|
const { scaffoldInit } = require('../lib/init');
|
|
19
19
|
|
|
@@ -75,7 +75,7 @@ function loadRc(skillDir) {
|
|
|
75
75
|
|
|
76
76
|
// Flags that never take a value, so `--trusted-skill <skill-dir>` keeps the dir
|
|
77
77
|
// positional instead of swallowing it as the flag's value.
|
|
78
|
-
const BOOLEAN_FLAGS = new Set(['trusted-skill']);
|
|
78
|
+
const BOOLEAN_FLAGS = new Set(['trusted-skill', 'svg']);
|
|
79
79
|
|
|
80
80
|
// ── THE INPUT CONTRACT (spec 026 AC-9, F5) ────────────────────────────────────
|
|
81
81
|
// A COPY of action/lib.sh's RE_* rules, in the shape spec 028's fixture holds
|
|
@@ -177,7 +177,7 @@ USAGE
|
|
|
177
177
|
[--keep-transcripts] [--out DIR] [--trusted-skill]
|
|
178
178
|
${PROJECT_NAME} diff <receiptA.json> <receiptB.json> [--out FILE] [--mode release|revision]
|
|
179
179
|
${PROJECT_NAME} validate <receipt.json>
|
|
180
|
-
${PROJECT_NAME} badge <receipt.json|receipts-dir> [--models a,b] [--out FILE] [--github-output]
|
|
180
|
+
${PROJECT_NAME} badge <receipt.json|receipts-dir> [--models a,b] [--out FILE] [--github-output] [--svg [--href URL]]
|
|
181
181
|
${PROJECT_NAME} decide <receipts-dir> --models a,b [--github-output] [--badge FILE]
|
|
182
182
|
[--summary FILE] [--enforce] [--fail-on-regression true|false]
|
|
183
183
|
${PROJECT_NAME} import <results.json> --from agent-skills-eval|skillgrade [--out DIR]
|
|
@@ -515,17 +515,35 @@ function cmdBadge(positional, flags) {
|
|
|
515
515
|
if (fs.existsSync(p) && fs.statSync(p).isDirectory()) return cmdBadgeSet(p, flags);
|
|
516
516
|
const receipt = JSON.parse(fs.readFileSync(p, 'utf8'));
|
|
517
517
|
refuseUnverified(receipt, p, 'badge');
|
|
518
|
+
if (flags.svg) {
|
|
519
|
+
// SPEC 036. The drawn badge: state, model, date, lift and uncertainty, with the
|
|
520
|
+
// machine token in its data attributes. After the same hash check as the JSON.
|
|
521
|
+
const svg = require('../lib/badge-svg').badgeSvg(receipt, { href: typeof flags.href === 'string' ? flags.href : null });
|
|
522
|
+
if (flags.out) {
|
|
523
|
+
const out = path.resolve(flags.out);
|
|
524
|
+
fs.mkdirSync(path.dirname(out), { recursive: true });
|
|
525
|
+
fs.writeFileSync(out, svg);
|
|
526
|
+
console.log(`badge written to ${flags.out} (${verdictFromReceipt(receipt).verdict}, svg)`);
|
|
527
|
+
} else process.stdout.write(svg);
|
|
528
|
+
return;
|
|
529
|
+
}
|
|
518
530
|
if (flags['github-output']) { console.log(githubOutputLines(receipt)); return; }
|
|
519
531
|
const badge = badgeEndpoint(receipt);
|
|
520
532
|
const json = JSON.stringify(badge, null, 2);
|
|
533
|
+
const v = verdictFromReceipt(receipt);
|
|
534
|
+
// SPEC 035 AC-3. An UNDERPOWERED receipt says so in words and says what it would
|
|
535
|
+
// have needed. Beside JSON on stdout the two lines go to stderr, so a caller that
|
|
536
|
+
// redirects the badge still gets a file that parses.
|
|
537
|
+
const power = v.verdict === 'UNDERPOWERED' ? [`${UNDERPOWERED_LINE}.`, drawsLine(v.drawsNeeded)] : [];
|
|
521
538
|
if (flags.out) {
|
|
522
539
|
const out = path.resolve(flags.out);
|
|
523
540
|
fs.mkdirSync(path.dirname(out), { recursive: true });
|
|
524
541
|
fs.writeFileSync(out, json + '\n');
|
|
525
|
-
const v = verdictFromReceipt(receipt);
|
|
526
542
|
console.log(`badge written to ${flags.out} (${v.verdict}: ${badge.message}, ${badge.color})`);
|
|
543
|
+
for (const line of power) console.log(line);
|
|
527
544
|
} else {
|
|
528
545
|
console.log(json);
|
|
546
|
+
for (const line of power) console.error(line);
|
|
529
547
|
}
|
|
530
548
|
}
|
|
531
549
|
|
package/config.js
CHANGED
|
@@ -9,7 +9,7 @@ const PROJECT_NAME = 'driftproof';
|
|
|
9
9
|
// Bumped whenever the runner's behaviour or receipt-generation semantics change
|
|
10
10
|
// in a way that could affect results. Recorded into every receipt as
|
|
11
11
|
// run.runner_version so a receipt is reproducible against a known engine.
|
|
12
|
-
const RUNNER_VERSION = '0.
|
|
12
|
+
const RUNNER_VERSION = '0.11.0';
|
|
13
13
|
|
|
14
14
|
// The eval format we CONSUME (we deliberately do not invent our own).
|
|
15
15
|
const SUITE_FORMAT = 'agentskills.io/evals';
|
|
@@ -38,7 +38,7 @@ const SUITE_FORMAT = 'agentskills.io/evals';
|
|
|
38
38
|
// results.aggregates.band_rule, and bands that are null where the formula
|
|
39
39
|
// cannot form. NOT additive for the validator: a v0.5 receipt restamped 0.6
|
|
40
40
|
// is refused. v0.5 is frozen as receipt.v0.5.schema.json.
|
|
41
|
-
const RECEIPT_SCHEMA_VERSION = '0.
|
|
41
|
+
const RECEIPT_SCHEMA_VERSION = '0.7';
|
|
42
42
|
|
|
43
43
|
// Input bounds on the skill loader and the post-checks (spec 026 AC-13, audit
|
|
44
44
|
// A3/A7). lib/skill.js loaded every bundled file with no count, depth, byte,
|
|
@@ -99,11 +99,18 @@ const DEFAULT_JUDGE_SAMPLES = 5;
|
|
|
99
99
|
// to a coarse ~0.05–0.1 grid, so a confident grade often has stddev 0 (a "point
|
|
100
100
|
// band"). Two point bands that differ by a single quantum (e.g. 0.60 vs 0.64)
|
|
101
101
|
// are technically non-overlapping yet represent no meaningful behaviour change.
|
|
102
|
-
// The floor
|
|
103
|
-
// "
|
|
102
|
+
// The floor keeps such separated but trivial moves from being called a verdict
|
|
103
|
+
// ("below effect floor" in the per-case table). 0.05 = one judge quantum. Documented in
|
|
104
104
|
// spec/RECEIPT.md § "Drift verdict rule" and the report methodology.
|
|
105
105
|
const EFFECT_FLOOR = 0.05;
|
|
106
106
|
|
|
107
|
+
// How many standard errors of the difference between two means a case's resolution
|
|
108
|
+
// adds to its two spreads (spec 035 R-2). A case the band rule does not separate is
|
|
109
|
+
// UNDERPOWERED when its spreads plus this many standard errors reach the floor: at
|
|
110
|
+
// those spreads and draws, a true shift of the floor's size could have gone unseparated.
|
|
111
|
+
// Named, so a later decision moves it by one line and one amendment.
|
|
112
|
+
const POWER_Z = 2;
|
|
113
|
+
|
|
107
114
|
// ── generation sampling (receipt spec v0.5) ─────────────────────────────────
|
|
108
115
|
//
|
|
109
116
|
// Report #006 measured across-draw spread at sd 0.186 and 0.183 while the
|
|
@@ -131,7 +138,7 @@ const REPORT_002_GPT_MODEL = 'gpt-5.6-sol';
|
|
|
131
138
|
const REPORT_002_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
132
139
|
|
|
133
140
|
// Report #003 is a RELEASE DRIFT report (Report #001's type): one provider, two
|
|
134
|
-
// model versions,
|
|
141
|
+
// model versions, a per-skill verdict label (regressed, improved, mixed, or none separated) under
|
|
135
142
|
// the per-case band-separation rule + effect floor. New vs its family predecessor.
|
|
136
143
|
const REPORT_003_NEW_MODEL = 'claude-opus-5';
|
|
137
144
|
const REPORT_003_OLD_MODEL = 'claude-opus-4-8';
|
|
@@ -157,7 +164,7 @@ const REPORT_005_JUDGE_MODEL = 'claude-haiku-4-5';
|
|
|
157
164
|
module.exports = {
|
|
158
165
|
PROJECT_NAME, RUNNER_VERSION, SUITE_FORMAT, RECEIPT_SCHEMA_VERSION, DEFAULT_JUDGE_SAMPLES,
|
|
159
166
|
SKILL_MAX_FILES, SKILL_MAX_BYTES, SKILL_MAX_DEPTH, SUITE_MAX_CASES, CASE_MAX_CHARS, CHECK_MAX_PATTERN,
|
|
160
|
-
EFFECT_FLOOR, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
167
|
+
EFFECT_FLOOR, POWER_Z, DEV_MAX_USD, DEV_MAX_CALLS, REPORT_MAX_USD, TRIGGER_MAX_USD,
|
|
161
168
|
GENERATION_SAMPLES_MIN, GENERATION_SAMPLES_MAX, GENERATION_SD_THRESHOLD, GENERATION_STABILITY_EPS,
|
|
162
169
|
REPORT_002_CLAUDE_MODEL, REPORT_002_GPT_MODEL, REPORT_002_JUDGE_MODEL,
|
|
163
170
|
REPORT_003_NEW_MODEL, REPORT_003_OLD_MODEL, REPORT_003_JUDGE_MODEL,
|
package/lib/badge-svg.js
ADDED
|
@@ -0,0 +1,130 @@
|
|
|
1
|
+
// SPDX-License-Identifier: Apache-2.0
|
|
2
|
+
'use strict';
|
|
3
|
+
|
|
4
|
+
// The badge, drawn (spec 036).
|
|
5
|
+
//
|
|
6
|
+
// The shields endpoint JSON (lib/verdict.js badgeEndpoint) carries a label, a message
|
|
7
|
+
// and a colour, which is all a shields badge can hold. This draws the badge itself, so
|
|
8
|
+
// it can show what a reader needs at a glance: the state, the model, the date, the lift
|
|
9
|
+
// and its uncertainty. Four segments, left to right:
|
|
10
|
+
//
|
|
11
|
+
// driftproof | <state label> | <model> · <date> | lift <+0.000> <± 0.000 | no ±>
|
|
12
|
+
//
|
|
13
|
+
// The fourth segment carries the lift ONLY for a state that measured one. UNDERPOWERED and
|
|
14
|
+
// NOT_MEASURED carry the draws taken against the draws needed instead (A-036-2):
|
|
15
|
+
//
|
|
16
|
+
// driftproof | not enough draws | <model> · <date> | 3 draws of 87 needed
|
|
17
|
+
//
|
|
18
|
+
// THE WORDS ARE HUMAN; THE TOKENS ARE DATA. What a reader sees is a label (passing, not
|
|
19
|
+
// enough draws); the machine token the Action and the differ publish sits on the root
|
|
20
|
+
// element as data-verdict, with the receipt hash beside it, and never in the text.
|
|
21
|
+
//
|
|
22
|
+
// UNDERPOWERED IS DRAWN AS THE HONEST STATE, not as a failure: a calm indigo no other
|
|
23
|
+
// state uses, an open hatch over it, a dashed edge around it, and the words "not enough
|
|
24
|
+
// draws". Every state's text is white on a fill it clears 4.5:1 against, hatch included.
|
|
25
|
+
//
|
|
26
|
+
// Text widths are fixed with textLength, so a text's rendered box is the width this file
|
|
27
|
+
// allots it whatever font the viewer has, and it cannot spill out of its segment.
|
|
28
|
+
|
|
29
|
+
const { receiptVerdict, shortModel, receiptDrawsTaken } = require('./verdict');
|
|
30
|
+
|
|
31
|
+
const INK = '#ffffff';
|
|
32
|
+
const BRAND = '#24292f';
|
|
33
|
+
const DETAIL = '#4b5563';
|
|
34
|
+
const EFFECT = '#374151';
|
|
35
|
+
const STATE = {
|
|
36
|
+
PASSED: { fill: '#1a7f37' },
|
|
37
|
+
REGRESSED: { fill: '#b42318' },
|
|
38
|
+
UNDERPOWERED: { fill: '#3538cd', hatch: '#444ce7', edge: '#a4bcfd' },
|
|
39
|
+
NO_EFFECT: { fill: '#57606a' },
|
|
40
|
+
NOT_MEASURED: { fill: '#656d76' },
|
|
41
|
+
};
|
|
42
|
+
const LABELS = {
|
|
43
|
+
PASSED: 'passing',
|
|
44
|
+
REGRESSED: 'regressed',
|
|
45
|
+
UNDERPOWERED: 'not enough draws',
|
|
46
|
+
NO_EFFECT: 'no separation detected',
|
|
47
|
+
NOT_MEASURED: 'not measured',
|
|
48
|
+
};
|
|
49
|
+
|
|
50
|
+
const H = 20;
|
|
51
|
+
const PAD = 7;
|
|
52
|
+
const CHAR = 6.6;
|
|
53
|
+
const esc = (s) => String(s).replace(/&/g, '&').replace(/</g, '<').replace(/>/g, '>').replace(/"/g, '"');
|
|
54
|
+
const textW = (s) => Math.round(String(s).length * CHAR * 10) / 10;
|
|
55
|
+
|
|
56
|
+
// A LIFT FIGURE IS A MEASUREMENT, so only a state that measured one may print it
|
|
57
|
+
// (spec 036 A-036-2). UNDERPOWERED says the receipt's draws cannot resolve a shift of the
|
|
58
|
+
// effect floor, and NOT_MEASURED says the receipt carries no verdict at all. A lift beside
|
|
59
|
+
// either is the badge answering, in its most legible segment, a question the state has just
|
|
60
|
+
// said it cannot answer -- and a reader who takes one figure from a badge takes that one.
|
|
61
|
+
// Those two states print what the reader can act on instead: the draws taken against the
|
|
62
|
+
// draws that would have been needed.
|
|
63
|
+
function effectText(receipt, v) {
|
|
64
|
+
const c = receipt.comparison || {};
|
|
65
|
+
if (v.verdict === 'UNDERPOWERED' || v.verdict === 'NOT_MEASURED') return drawsText(receipt, v);
|
|
66
|
+
const lift = typeof c.delta === 'number' ? `lift ${c.delta >= 0 ? '+' : '-'}${Math.abs(c.delta).toFixed(3)}` : 'no lift';
|
|
67
|
+
const unc = typeof c.delta_uncertainty === 'number' ? `± ${c.delta_uncertainty.toFixed(3)}`
|
|
68
|
+
: c.delta_uncertainty_unavailable === 'single_case' ? 'no ± (1 case)' : 'no ±';
|
|
69
|
+
return `${lift} ${unc}`;
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
// Draws taken against draws needed, in the badge's words. The needed count is
|
|
73
|
+
// receiptDrawsNeeded's, the same figure drawsLine prints in a sentence, so the badge and
|
|
74
|
+
// the page cannot say two different numbers. Where no count exists the badge says which
|
|
75
|
+
// of the two reasons it is, rather than printing a bare figure or nothing at all.
|
|
76
|
+
//
|
|
77
|
+
// BOTH NUMBERS ARE ONE CASE'S (A-036-7). Where a case sets the draws needed, the draws taken
|
|
78
|
+
// are that case's own, which receiptDrawsNeeded carries as `taken`. This read the smallest
|
|
79
|
+
// over every readable case, so a receipt whose binding case drew 4 and whose other case drew
|
|
80
|
+
// 2 badged "2 draws of 5 needed": a 2 from a case the 5 does not belong to. Where no case sets
|
|
81
|
+
// the draws needed (NOT_MEASURED carries no verdict to compute one from) there is no binding
|
|
82
|
+
// case, and the smallest over every readable case is the figure.
|
|
83
|
+
function drawsText(receipt, v) {
|
|
84
|
+
const d = v.drawsNeeded;
|
|
85
|
+
const taken = d ? d.taken : receiptDrawsTaken(receipt);
|
|
86
|
+
if (taken === null || taken === undefined) return 'no draws recorded';
|
|
87
|
+
const draws = `${taken} ${taken === 1 ? 'draw' : 'draws'}`;
|
|
88
|
+
if (!d) return `${draws}, needed not computed`;
|
|
89
|
+
if (typeof d.value === 'number') return `${draws} of ${d.value} needed`;
|
|
90
|
+
return `${draws}, no count at these spreads`;
|
|
91
|
+
}
|
|
92
|
+
|
|
93
|
+
function badgeSvg(receipt, { href = null } = {}) {
|
|
94
|
+
const v = receiptVerdict(receipt);
|
|
95
|
+
const state = STATE[v.verdict];
|
|
96
|
+
const label = LABELS[v.verdict];
|
|
97
|
+
const model = shortModel((receipt.run || {}).model_id);
|
|
98
|
+
const date = String((receipt.run || {}).date_utc || '').slice(0, 10);
|
|
99
|
+
const segs = [
|
|
100
|
+
{ seg: 'brand', text: 'driftproof', fill: BRAND },
|
|
101
|
+
{ seg: 'state', text: label, fill: state.fill },
|
|
102
|
+
{ seg: 'detail', text: `${model} · ${date}`, fill: DETAIL },
|
|
103
|
+
{ seg: 'effect', text: effectText(receipt, v), fill: EFFECT },
|
|
104
|
+
];
|
|
105
|
+
let x = 0;
|
|
106
|
+
for (const s of segs) { s.w = Math.round(textW(s.text) + 2 * PAD); s.x = x; x += s.w; }
|
|
107
|
+
const width = x;
|
|
108
|
+
const title = `driftproof: ${label} on ${model}, ${date}, ${effectText(receipt, v)}`;
|
|
109
|
+
const L = [];
|
|
110
|
+
L.push(`<svg xmlns="http://www.w3.org/2000/svg" width="${width}" height="${H}" viewBox="0 0 ${width} ${H}" role="img" aria-label="${esc(title)}" data-verdict="${v.verdict}" data-receipt-hash="${esc(receipt.receipt_hash || '')}">`);
|
|
111
|
+
L.push(`<title>${esc(title)}</title>`);
|
|
112
|
+
if (state.hatch) {
|
|
113
|
+
L.push(`<defs><pattern id="hatch" width="6" height="6" patternUnits="userSpaceOnUse" patternTransform="rotate(45)"><line x1="0" y1="0" x2="0" y2="6" stroke="${state.hatch}" stroke-width="3"/></pattern></defs>`);
|
|
114
|
+
}
|
|
115
|
+
if (href) L.push(`<a href="${esc(href)}">`);
|
|
116
|
+
for (const s of segs) {
|
|
117
|
+
const isState = s.seg === 'state';
|
|
118
|
+
const edge = isState && state.edge ? ` stroke="${state.edge}" stroke-width="1" stroke-dasharray="3 2"` : '';
|
|
119
|
+
L.push(`<rect data-seg="${s.seg}" x="${s.x}" y="0" width="${s.w}" height="${H}" fill="${s.fill}"${edge}/>`);
|
|
120
|
+
if (isState && state.hatch) L.push(`<rect data-seg="state-hatch" x="${s.x}" y="0" width="${s.w}" height="${H}" fill="url(#hatch)" opacity="0.35"/>`);
|
|
121
|
+
}
|
|
122
|
+
for (const s of segs) {
|
|
123
|
+
L.push(`<text data-seg="${s.seg}" x="${s.x + PAD}" y="14" fill="${INK}" font-family="Verdana,DejaVu Sans,sans-serif" font-size="11" textLength="${textW(s.text)}" lengthAdjust="spacingAndGlyphs">${esc(s.text)}</text>`);
|
|
124
|
+
}
|
|
125
|
+
if (href) L.push('</a>');
|
|
126
|
+
L.push('</svg>');
|
|
127
|
+
return L.join('\n') + '\n';
|
|
128
|
+
}
|
|
129
|
+
|
|
130
|
+
module.exports = { badgeSvg, LABELS, STATE };
|
package/lib/decision.js
CHANGED
|
@@ -20,7 +20,7 @@
|
|
|
20
20
|
const fs = require('fs');
|
|
21
21
|
const path = require('path');
|
|
22
22
|
const { EFFECT_FLOOR } = require('../config');
|
|
23
|
-
const { shortModel, githubOutputEntry } = require('./verdict');
|
|
23
|
+
const { shortModel, githubOutputEntry, receiptVerdict, drawsLine, UNDERPOWERED_LINE } = require('./verdict');
|
|
24
24
|
|
|
25
25
|
// The six decision states (spec 030 AC-4).
|
|
26
26
|
//
|
|
@@ -29,7 +29,12 @@ const { shortModel, githubOutputEntry } = require('./verdict');
|
|
|
29
29
|
// that says so. `no detected effect` is false about such a model - the effect
|
|
30
30
|
// was detected and it cleared the floor - and widening it to mean "not a
|
|
31
31
|
// regression" would put a measured lift and a measured nothing in one cell.
|
|
32
|
-
|
|
32
|
+
//
|
|
33
|
+
// `underpowered` is the seventh (spec 035): the receipt measured, and its own spreads
|
|
34
|
+
// and draws could not resolve a shift of the effect floor. It sits after `not
|
|
35
|
+
// measured`, which measured nothing, and before `no detected effect`, which measured
|
|
36
|
+
// with a resolution below the floor.
|
|
37
|
+
const STATES = ['regression', 'refused', 'inconclusive', 'not measured', 'underpowered', 'no detected effect', 'helped'];
|
|
33
38
|
|
|
34
39
|
// WORST FIRST. This is the single definition of the ordering; AC-1's
|
|
35
40
|
// enforcement and AC-3's badge both read it, so they cannot disagree about
|
|
@@ -38,7 +43,7 @@ const STATE_ORDER = STATES.slice();
|
|
|
38
43
|
|
|
39
44
|
// States that must never render as success on any surface the action writes -
|
|
40
45
|
// the badge, the summary row, or the check title (AC-4).
|
|
41
|
-
const NEVER_SUCCESS = ['inconclusive', 'not measured'];
|
|
46
|
+
const NEVER_SUCCESS = ['inconclusive', 'not measured', 'underpowered'];
|
|
42
47
|
|
|
43
48
|
// States that fail the job. `refused` fails CLOSED (AC-2): a model that
|
|
44
49
|
// produced no receipt is not a model that passed. `regression` fails subject to
|
|
@@ -86,11 +91,15 @@ function decisionState(receipt) {
|
|
|
86
91
|
if (level !== 'TESTED' || kind !== 'model') return 'not measured';
|
|
87
92
|
const cmp = receipt.comparison || {};
|
|
88
93
|
if ((receipt.run && receipt.run.status === 'incomplete') || typeof cmp.delta !== 'number') return 'inconclusive';
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
94
|
+
// Spec 035: the measured states are the receipt verdict's, read per case by the
|
|
95
|
+
// band rule (lib/verdict.js receiptVerdict), not the aggregate delta against the floor.
|
|
96
|
+
// receiptVerdict refuses on the same four conditions the two clauses above split
|
|
97
|
+
// between `not measured` and `inconclusive`, so it cannot answer NOT_MEASURED here.
|
|
98
|
+
return MEASURED_STATE[receiptVerdict(receipt).verdict];
|
|
92
99
|
}
|
|
93
100
|
|
|
101
|
+
const MEASURED_STATE = { REGRESSED: 'regression', PASSED: 'helped', UNDERPOWERED: 'underpowered', NO_EFFECT: 'no detected effect' };
|
|
102
|
+
|
|
94
103
|
// The worst state in a set, by STATE_ORDER. An empty set has no decision, and
|
|
95
104
|
// says so with null rather than defaulting to something benign.
|
|
96
105
|
function worstState(states) {
|
|
@@ -182,6 +191,8 @@ function decideSet(dir, requestedCsv, { failOnRegression = true } = {}) {
|
|
|
182
191
|
delta: receipt && receipt.comparison && typeof receipt.comparison.delta === 'number'
|
|
183
192
|
? receipt.comparison.delta : null,
|
|
184
193
|
file: hit ? path.basename(hit.file) : (unreadable ? path.basename(unreadable.file) : null),
|
|
194
|
+
// Spec 035 AC-3: what an underpowered row would have needed, from the same reading.
|
|
195
|
+
drawsNeeded: state === 'underpowered' ? receiptVerdict(receipt).drawsNeeded : null,
|
|
185
196
|
// The distinction is carried on the ROW, not left in a sentence, because
|
|
186
197
|
// the enforcement message has to say which of the two happened (AC-2's
|
|
187
198
|
// mutation class, absence-vs-unreadable).
|
|
@@ -246,6 +257,7 @@ const RENDER = {
|
|
|
246
257
|
refused: { word: 'refused', color: 'red', marker: '\u274c' },
|
|
247
258
|
inconclusive: { word: 'inconclusive', color: 'yellow', marker: '\u26a0\ufe0f' },
|
|
248
259
|
'not measured': { word: 'not measured', color: 'lightgrey', marker: '\u26a0\ufe0f' },
|
|
260
|
+
underpowered: { word: 'not enough draws', color: 'blue', marker: '\u26a0\ufe0f' },
|
|
249
261
|
'no detected effect': { word: 'no effect', color: 'lightgrey', marker: '\u2014' },
|
|
250
262
|
helped: { word: 'passing', color: 'brightgreen', marker: '\u2705' },
|
|
251
263
|
};
|
|
@@ -283,7 +295,7 @@ function badgeEndpointForSet(d, { label = 'driftproof' } = {}) {
|
|
|
283
295
|
// `worst`, `regressed` and `missing` are new and carry what it could not say.
|
|
284
296
|
const VERDICT_WORD = {
|
|
285
297
|
regression: 'REGRESSED', refused: 'REFUSED', inconclusive: 'INCONCLUSIVE',
|
|
286
|
-
'not measured': 'NOT_MEASURED', 'no detected effect': 'NO_EFFECT', helped: 'PASSED',
|
|
298
|
+
'not measured': 'NOT_MEASURED', underpowered: 'UNDERPOWERED', 'no detected effect': 'NO_EFFECT', helped: 'PASSED',
|
|
287
299
|
};
|
|
288
300
|
|
|
289
301
|
function githubOutputLines(d) {
|
|
@@ -299,6 +311,10 @@ function githubOutputLines(d) {
|
|
|
299
311
|
['regressed_models', d.regressed.join(',')],
|
|
300
312
|
['missing_models', d.missing.join(',')],
|
|
301
313
|
['receipt_count', d.receiptCount],
|
|
314
|
+
// Spec 035 AC-3: the draws the worst row would have needed, when that row is
|
|
315
|
+
// underpowered; `none` when no draw count reaches the floor at its spreads, and
|
|
316
|
+
// empty otherwise.
|
|
317
|
+
['draws_needed', worstRow && worstRow.drawsNeeded ? (worstRow.drawsNeeded.value === null ? 'none' : worstRow.drawsNeeded.value) : ''],
|
|
302
318
|
].map(([k, v]) => githubOutputEntry(k, v)).join('\n');
|
|
303
319
|
}
|
|
304
320
|
|
|
@@ -324,7 +340,8 @@ function summaryMarkdown(d) {
|
|
|
324
340
|
const r = RENDER[row.state];
|
|
325
341
|
const delta = row.delta === null ? 'n/a' : (row.delta >= 0 ? '+' : '') + row.delta.toFixed(3);
|
|
326
342
|
const note = row.reason ? ` <br><sub>${cell(row.reason)}</sub>` : '';
|
|
327
|
-
|
|
343
|
+
const power = row.state === 'underpowered' ? ` <br><sub>${cell(UNDERPOWERED_LINE)}. ${cell(drawsLine(row.drawsNeeded))}</sub>` : '';
|
|
344
|
+
L.push(`| \`${cell(row.model)}\` | ${r.marker} ${row.state}${power} | ${delta} | ${cell(row.file) || '\u2014'}${note} |`);
|
|
328
345
|
}
|
|
329
346
|
L.push('');
|
|
330
347
|
L.push(`Decided over ${d.receiptCount} receipt(s) for ${d.requestedCount} requested model(s). `
|