shapeup-sdlc 3.12.0 → 3.14.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +2 -2
- package/kernel/compile.mjs +96 -2
- package/kernel/gate.mjs +28 -2
- package/kernel/probe/digest.mjs +10 -0
- package/kernel/probe/stats.mjs +4 -1
- package/kernel/reduce/ship.mjs +1 -1
- package/kernel/schemas/domain.schema.json +14 -0
- package/kernel/schemas/work-order.schema.json +44 -13
- package/kernel/verify/t0.mjs +105 -2
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +5 -3
- package/skills/scope-architect/SKILL.md +3 -1
- package/skills/scope-hammer/SKILL.md +7 -5
- package/skills/solution-architect/SKILL.md +3 -1
- package/skills/spec-evaluator/SKILL.md +10 -5
- package/skills/task-executor/SKILL.md +5 -3
- package/skills/tech-lead/workflows/shapeup-run.js +1 -1
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.14.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -37,7 +37,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
|
|
|
37
37
|
| Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
|
|
38
38
|
| Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
|
|
39
39
|
| EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
|
|
40
|
-
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
|
|
40
|
+
| FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
|
|
41
41
|
|
|
42
42
|
✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
|
|
43
43
|
|
|
@@ -69,7 +69,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
69
69
|
## Setup & Execution
|
|
70
70
|
|
|
71
71
|
- Orders/results live in `.shapeup/<slug>/orders|results/`; the envelope schemas ship with the plugin runtime, not with any individual skill, so every worker validates against the same copy.
|
|
72
|
-
- The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf.
|
|
72
|
+
- The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf. Two facts about how grants reach the run's workers, both measured: a worker dispatched inside a run takes its permissions from the project's settings file, not from flags given to the launching session, so a tool the workers need belongs in `permissions.allow`; and every work order names the plugin's kernel by absolute path, because a command that spells it through a variable is refused in a headless session before any permission rule is read.
|
|
73
73
|
- The grant is necessary but sits under two more layers this plugin cannot reach either. A fresh
|
|
74
74
|
checkout is an **untrusted workspace**, and Claude Code discards the whole permission grant — every
|
|
75
75
|
rule in it, not only this one — until the workspace is trusted; the installer detects that state and
|
package/kernel/compile.mjs
CHANGED
|
@@ -37,7 +37,7 @@ import { readRunId, dispatchReceipts, legLedger, readReceipt, receipt } from "./
|
|
|
37
37
|
// --spec-overridden directory, and the import is the convention-derived default.
|
|
38
38
|
import {
|
|
39
39
|
tasksDir, specDir as defaultSpecDir, roundLedger, trials, verdictsDir, ordersDir,
|
|
40
|
-
relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir,
|
|
40
|
+
relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir, localRoot,
|
|
41
41
|
} from "./lib/paths.mjs";
|
|
42
42
|
import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT, reqId } from "./lib/contract.mjs";
|
|
43
43
|
import { writeActiveOrder } from "./probe/resume.mjs";
|
|
@@ -527,6 +527,57 @@ export function verdictBugs(cwd, slug, round) {
|
|
|
527
527
|
return v.bugs.filter((b) => !refuted.has(String(b?.id)) && !refuted.has(String(b?.criterion)));
|
|
528
528
|
}
|
|
529
529
|
|
|
530
|
+
/**
|
|
531
|
+
* The previous round's FAILED CRITERIA that the judge filed no bug for, as bug entries.
|
|
532
|
+
*
|
|
533
|
+
* A verdict carries two lists: the criteria it graded and the bugs it filed. Only the second
|
|
534
|
+
* reached the next round, so a criterion graded FAIL — "no evidence: no check names this row", a
|
|
535
|
+
* device row the judge could not observe — with no matching bug handed the fix round nothing to do.
|
|
536
|
+
* Measured: three fix rounds in a row compiled with no bugs over a FAIL verdict, and the run ended
|
|
537
|
+
* where it started. Every failed criterion is a thing the round must address, so each one without a
|
|
538
|
+
* bug becomes one, addressed to the scope whose contract lists the use case the criterion names.
|
|
539
|
+
*
|
|
540
|
+
* @param {string} cwd - Project root.
|
|
541
|
+
* @param {string} slug - Feature slug.
|
|
542
|
+
* @param {number} [round] - The round being compiled.
|
|
543
|
+
* @returns {Array<object>} `{criterion, severity, expected, actual, source:"criterion", scope_id?}`
|
|
544
|
+
* per uncovered FAIL criterion; [] for round 1, a PASS, or an unreadable result.
|
|
545
|
+
*/
|
|
546
|
+
export function criteriaBugs(cwd, slug, round) {
|
|
547
|
+
if (!round || round < 2) return [];
|
|
548
|
+
const p = join(resultsDir(cwd, slug), `evaluate-r${round - 1}.json`);
|
|
549
|
+
if (!existsSync(p)) return [];
|
|
550
|
+
let v;
|
|
551
|
+
try { v = JSON.parse(readFileSync(p, "utf8"))?.verdict; } catch { return []; }
|
|
552
|
+
if (v?.overall !== "FAIL" || !Array.isArray(v.criteria)) return [];
|
|
553
|
+
const filed = (Array.isArray(v.bugs) ? v.bugs : []).map((b) => String(b?.criterion ?? ""));
|
|
554
|
+
const refuted = new Set((Array.isArray(v.refuted) ? v.refuted : [])
|
|
555
|
+
.flatMap((r) => [r?.id, r?.ac_id, r?.criterion, typeof r === "string" ? r : null]).filter(Boolean).map(String));
|
|
556
|
+
const owners = new Map();
|
|
557
|
+
for (const { contract, id } of readAllContracts(scopesDir(cwd, slug))) {
|
|
558
|
+
const sid = contract?.scope_id || id;
|
|
559
|
+
for (const uc of Array.isArray(contract?.use_cases) ? contract.use_cases : []) if (!owners.has(uc)) owners.set(uc, sid);
|
|
560
|
+
}
|
|
561
|
+
const out = [];
|
|
562
|
+
for (const c of v.criteria) {
|
|
563
|
+
if (c?.verdict !== "FAIL" || typeof c.criterion !== "string") continue;
|
|
564
|
+
const name = c.criterion;
|
|
565
|
+
if (refuted.has(name)) continue;
|
|
566
|
+
if (filed.some((f) => f && (f === name || f.includes(name) || name.includes(f)))) continue;
|
|
567
|
+
const uc = (name.match(/\bUC-[A-Za-z0-9_-]+/) || [])[0];
|
|
568
|
+
const owner = uc ? owners.get(uc) : undefined;
|
|
569
|
+
out.push({
|
|
570
|
+
criterion: name,
|
|
571
|
+
severity: "major",
|
|
572
|
+
expected: "the criterion is met, with evidence the judge can cite",
|
|
573
|
+
actual: String(c.evidence ?? "graded FAIL").slice(0, 400),
|
|
574
|
+
source: "criterion",
|
|
575
|
+
...(owner ? { scope_id: owner } : {}),
|
|
576
|
+
});
|
|
577
|
+
}
|
|
578
|
+
return out;
|
|
579
|
+
}
|
|
580
|
+
|
|
530
581
|
/**
|
|
531
582
|
* The previous round's RED BUILD GATE, as bug entries the fix round can act on.
|
|
532
583
|
*
|
|
@@ -748,6 +799,38 @@ export function t0ArtifactsFor(cwd, slug, round) {
|
|
|
748
799
|
return { artifacts, missing };
|
|
749
800
|
}
|
|
750
801
|
|
|
802
|
+
// --- checks a scope rewrote between a failing trial and a passing one --------------------------
|
|
803
|
+
|
|
804
|
+
/**
|
|
805
|
+
* Every revised check any trial of the round recorded, with the scope it belongs to.
|
|
806
|
+
*
|
|
807
|
+
* Read across all of the round's trials rather than off the cited verdict alone: a check rewritten
|
|
808
|
+
* on attempt 2 and re-run unchanged on attempt 3 leaves no mark on attempt 3's verdict, which is the
|
|
809
|
+
* one the judge cites.
|
|
810
|
+
*
|
|
811
|
+
* @param {string} cwd - Project root.
|
|
812
|
+
* @param {string} slug - Feature slug.
|
|
813
|
+
* @param {number} [round] - The round being evaluated; omitted, every round.
|
|
814
|
+
* @returns {Array<{scope_id:string, id:string, file:string}>} Deduplicated by scope, id and file.
|
|
815
|
+
*/
|
|
816
|
+
export function revisedChecksFor(cwd, slug, round) {
|
|
817
|
+
const seen = new Set();
|
|
818
|
+
const out = [];
|
|
819
|
+
for (const tr of readTrials(trials(cwd, slug))) {
|
|
820
|
+
if (round != null && tr.round !== round) continue;
|
|
821
|
+
if (!tr.artifact) continue;
|
|
822
|
+
let v;
|
|
823
|
+
try { v = JSON.parse(readFileSync(join(localRoot(cwd, slug), tr.artifact), "utf8")); } catch { continue; }
|
|
824
|
+
for (const r of Array.isArray(v?.revised_checks) ? v.revised_checks : []) {
|
|
825
|
+
const key = `${tr.scope_id}\u0000${r.id}\u0000${r.file}`;
|
|
826
|
+
if (seen.has(key)) continue;
|
|
827
|
+
seen.add(key);
|
|
828
|
+
out.push({ scope_id: tr.scope_id, id: r.id, file: r.file });
|
|
829
|
+
}
|
|
830
|
+
}
|
|
831
|
+
return out;
|
|
832
|
+
}
|
|
833
|
+
|
|
751
834
|
// --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
|
|
752
835
|
//
|
|
753
836
|
// WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
|
|
@@ -851,6 +934,13 @@ export function compileOrder({
|
|
|
851
934
|
const order = {
|
|
852
935
|
schema_version: 1,
|
|
853
936
|
order_id: `${slug}/${suffix}`,
|
|
937
|
+
// WHERE THE KERNEL IS, AS A PATH A PERMISSION RULE CAN MATCH. Workers were told to run
|
|
938
|
+
// `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" …`. In a headless session that variable is not
|
|
939
|
+
// expanded for them, and a command carrying an unexpanded variable is refused before any rule is
|
|
940
|
+
// consulted ("Contains expansion") — so every kernel query a worker's contract requires was
|
|
941
|
+
// refused, and workers reported it as "not permitted". The absolute path of this very file's
|
|
942
|
+
// kernel is known here, and quoted it matches the grant `init` writes.
|
|
943
|
+
kernel: join(HERE, "harness.mjs"),
|
|
854
944
|
// THE TWO ANALYTIC FIELDS, and why they are on the ORDER rather than the result.
|
|
855
945
|
//
|
|
856
946
|
// `order_id` identifies a dispatch within a run and repeats across runs of the same slug, so
|
|
@@ -1009,7 +1099,7 @@ export async function cli(rawArgv) {
|
|
|
1009
1099
|
// this line, and none of them can pass a payload to a build order (see the banner above).
|
|
1010
1100
|
// Two sources, one channel: the judge's cited defects and the build gate's failing steps.
|
|
1011
1101
|
const bugs = scope
|
|
1012
|
-
? bugsForScope([...verdictBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
|
|
1102
|
+
? bugsForScope([...verdictBugs(cwd, slug, round), ...criteriaBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
|
|
1013
1103
|
: [];
|
|
1014
1104
|
|
|
1015
1105
|
// A ROUND CARRYING CITED DEFECTS IS A `fix`, AND THE ORDER HAS TO SAY SO.
|
|
@@ -1131,6 +1221,10 @@ export async function cli(rawArgv) {
|
|
|
1131
1221
|
}
|
|
1132
1222
|
// The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
|
|
1133
1223
|
// `--payload` value outranks the derivation.
|
|
1224
|
+
if (operation === "evaluate" && payloadExtra.revised_checks === undefined) {
|
|
1225
|
+
const revised = revisedChecksFor(cwd, slug, round);
|
|
1226
|
+
if (revised.length) payloadExtra.revised_checks = revised;
|
|
1227
|
+
}
|
|
1134
1228
|
if (operation === "evaluate") {
|
|
1135
1229
|
const ev = launchEvidenceFor(cwd, slug, round);
|
|
1136
1230
|
if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
|
package/kernel/gate.mjs
CHANGED
|
@@ -364,7 +364,7 @@ export function discover({ cwd = process.cwd(), file = null, preset = null, slug
|
|
|
364
364
|
export const ARGV_SPEC = {
|
|
365
365
|
usage: "harness.mjs gate (--init | --list | --verify | --resolve <gate-id>) [--preset <name>] " +
|
|
366
366
|
"[--file <path>] [--slug <slug>] [--cwd <dir>] [--out <path>] [--by <who>] " +
|
|
367
|
-
"[--auto-level <level>] [--tiny] [--no-qa] [--round <n>]",
|
|
367
|
+
"[--auto-level <level>] [--tiny] [--no-qa] [--round <n>] [--verdict <pass|fail|…>]",
|
|
368
368
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
369
369
|
cwd: { type: "path" },
|
|
370
370
|
init: { type: "flag" },
|
|
@@ -383,8 +383,32 @@ export const ARGV_SPEC = {
|
|
|
383
383
|
// ledger row needs it to key a `GateDecision` node uniquely per crossing. Round-independent gates
|
|
384
384
|
// (L0, L1a, …) simply omit it and the row carries `round: null`.
|
|
385
385
|
round: { type: "int", min: 1 },
|
|
386
|
+
// The verdict the round being crossed actually reached, when there is one. A preset answers L3
|
|
387
|
+
// before any verdict exists — "loop" on a FAIL — and its note says so; recorded without the
|
|
388
|
+
// verdict, a PASS round's row read "FAIL → fix round", the opposite of what happened.
|
|
389
|
+
verdict: { type: "str" },
|
|
386
390
|
};
|
|
387
391
|
|
|
392
|
+
/**
|
|
393
|
+
* The note a gate row carries — the answer's own, unless the verdict it was crossed over makes that
|
|
394
|
+
* note false.
|
|
395
|
+
*
|
|
396
|
+
* L3's answers are written for a verdict that has not happened yet: `loop` means "on a FAIL, run
|
|
397
|
+
* the next round". Crossed over a PASS, nothing loops, and the preset's "FAIL → fix round" note
|
|
398
|
+
* would put a failed round on the record for a round that passed.
|
|
399
|
+
*
|
|
400
|
+
* @param {{gate:string, decision?:string, note?:string, reason?:string}} r - The resolved answer.
|
|
401
|
+
* @param {(string|null)} verdict - The round's verdict, lower-cased, or null when none was given.
|
|
402
|
+
* @returns {(string|null)} The note to record.
|
|
403
|
+
*/
|
|
404
|
+
export function gateNote(r, verdict) {
|
|
405
|
+
const own = r.note ?? r.reason ?? null;
|
|
406
|
+
if (r.gate === "L3" && verdict === "pass") {
|
|
407
|
+
return `verdict PASS — the "${r.decision}" answer applies only to a failed round; the run goes on to QA and GATE H`;
|
|
408
|
+
}
|
|
409
|
+
return own;
|
|
410
|
+
}
|
|
411
|
+
|
|
388
412
|
function out(obj, code = 0) {
|
|
389
413
|
console.log(JSON.stringify(obj, null, 2));
|
|
390
414
|
process.exit(code);
|
|
@@ -468,11 +492,13 @@ export function cli(rawArgv) {
|
|
|
468
492
|
// THE ROW CARRIES THE RUN KEY. It did not, and the export stamped the current run's key onto
|
|
469
493
|
// every row it found — a prior run's sign-off became this run's in the one table that answers
|
|
470
494
|
// "was this ship signed off". Driven on a two-run fixture before it was fixed.
|
|
495
|
+
const verdict = args.verdict ? String(args.verdict).toLowerCase() : null;
|
|
471
496
|
appendGateLedger(cwd, args.slug, {
|
|
472
497
|
at: new Date().toISOString(), run_id: readRunId(cwd, args.slug),
|
|
473
498
|
gate: r.gate, status: r.status, decision: r.decision ?? null,
|
|
474
|
-
source: r.source ?? found.source, note: r
|
|
499
|
+
source: r.source ?? found.source, note: gateNote(r, verdict),
|
|
475
500
|
round: args.round ?? null,
|
|
501
|
+
...(verdict ? { verdict } : {}),
|
|
476
502
|
});
|
|
477
503
|
}
|
|
478
504
|
if (r.status === "ask") out({ ...r, ok: false }, 4);
|
package/kernel/probe/digest.mjs
CHANGED
|
@@ -33,6 +33,11 @@ const PATTERNS = [
|
|
|
33
33
|
// ("ERROR in the build pipeline", "ERROR in test suite failed to run") is left unmatched
|
|
34
34
|
// instead of handing back a fabricated file.
|
|
35
35
|
{ re: /^(?:ERROR|WARNING)\s+in\s+(\.{1,2}\/[^\s:]*|[^\s:]+\.[A-Za-z0-9]{1,10})\b/i, kind: "compiler-diagnostic" },
|
|
36
|
+
// hvigor / ArkTS compiler: the message and the location arrive on ONE line,
|
|
37
|
+
// "Error Message: Expected 5 arguments, but got 3. At File: /abs/path/Foo.test.ets:75:33"
|
|
38
|
+
// — the commonest failure on that toolchain, and one no other pattern here anchored, so every red
|
|
39
|
+
// compile handed the next attempt an empty error list.
|
|
40
|
+
{ re: /^\s*Error Message:\s*(.+?)\s+At File:\s*(.+?):(\d+):\d+\s*$/, kind: "arkts-compiler" },
|
|
36
41
|
// A test that FAILED BY NAME, with no file:line: "FAIL TS-05-05 step 4: no text 'Bread' on screen"
|
|
37
42
|
// or jest's "FAIL src/cart.test.js". Runners that drive an app from outside it (a device flow, an
|
|
38
43
|
// end-to-end script) report a case this way and nothing else, and the line is the whole signal —
|
|
@@ -77,6 +82,11 @@ export function digest(rawText) {
|
|
|
77
82
|
pendingMessage = coreMessage(m[1]);
|
|
78
83
|
continue; // wait for the stack frame that follows to get a file:line
|
|
79
84
|
}
|
|
85
|
+
if (kind === "arkts-compiler") {
|
|
86
|
+
triples.push({ file: m[2].trim(), line: Number(m[3]), core_message: coreMessage(m[1]), kind });
|
|
87
|
+
pendingMessage = null;
|
|
88
|
+
break;
|
|
89
|
+
}
|
|
80
90
|
const named = kind === "named-test-failure";
|
|
81
91
|
const file = named ? (/[\\/]|\.[A-Za-z0-9]{1,10}$/.test(m[1]) ? m[1] : null) : m[1]?.trim();
|
|
82
92
|
const lineNo = !named && m[2] ? Number(m[2]) : null;
|
package/kernel/probe/stats.mjs
CHANGED
|
@@ -194,7 +194,10 @@ export function ratchetReport(trials) {
|
|
|
194
194
|
}
|
|
195
195
|
if (scopeMonotone) monotone++;
|
|
196
196
|
}
|
|
197
|
-
|
|
197
|
+
// A score with no `regressions` field was not measured for regressions — T0 writes one only when
|
|
198
|
+
// it has a baseline to regress against — which is not the same as having some. Requiring an
|
|
199
|
+
// explicit 0 reported "no scope reached green" over a run whose every scope was green.
|
|
200
|
+
const greenAt = seq.findIndex((t) => t.score && t.score.fixtures_total > 0 && t.score.fixtures_passed === t.score.fixtures_total && (t.score.regressions ?? 0) === 0);
|
|
198
201
|
if (greenAt !== -1) toGreen.push(greenAt + 1);
|
|
199
202
|
per_scope.push({
|
|
200
203
|
scope_id, trials: seq.length,
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -280,7 +280,7 @@ export function buildReport(facts) {
|
|
|
280
280
|
L.push("| scope | fixtures | regressions | trials | last status | delta |", "|---|---|---|---|---|---|");
|
|
281
281
|
for (const s of t0) {
|
|
282
282
|
const f = s.score ? `${s.score.fixtures_passed}/${s.score.fixtures_total}` : "—";
|
|
283
|
-
const r = s.score ? String(s.score.regressions) : "—";
|
|
283
|
+
const r = s.score?.regressions != null ? String(s.score.regressions) : "—";
|
|
284
284
|
L.push(`| ${s.scope_id} | ${f} | ${r} | ${s.trials} | ${s.status} | ${s.delta || "—"} |`);
|
|
285
285
|
}
|
|
286
286
|
L.push("");
|
|
@@ -362,6 +362,7 @@
|
|
|
362
362
|
"run_cmd",
|
|
363
363
|
"launch_cmd",
|
|
364
364
|
"build_gate",
|
|
365
|
+
"revised_checks",
|
|
365
366
|
"t0_artifacts",
|
|
366
367
|
"browser",
|
|
367
368
|
"tasks"
|
|
@@ -2429,6 +2430,19 @@
|
|
|
2429
2430
|
"type": "string",
|
|
2430
2431
|
"description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
|
|
2431
2432
|
},
|
|
2433
|
+
"revised_checks": {
|
|
2434
|
+
"type": "array",
|
|
2435
|
+
"items": {
|
|
2436
|
+
"type": "object",
|
|
2437
|
+
"properties": {
|
|
2438
|
+
"scope_id": { "type": "string" },
|
|
2439
|
+
"id": { "type": "string" },
|
|
2440
|
+
"file": { "type": "string" }
|
|
2441
|
+
},
|
|
2442
|
+
"required": ["id", "file"]
|
|
2443
|
+
},
|
|
2444
|
+
"description": "spec-evaluator: rows that FAILed in one of the round's T0 trials and PASS in a later one whose own check file (named by the row id) changed in between. Derived by harness compile. Each must be read against its row before its PASS counts; absent when none."
|
|
2445
|
+
},
|
|
2432
2446
|
"build_gate": {
|
|
2433
2447
|
"type": "string",
|
|
2434
2448
|
"description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
|
|
@@ -1,30 +1,61 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$id": "work-order.schema.json",
|
|
3
3
|
"title": "WorkOrder",
|
|
4
|
-
"description": "The orchestrator
|
|
4
|
+
"description": "The orchestrator \u2192 worker envelope (pure-skill architecture v1.0). Compiled by kernel/compile.mjs, validated by harness verify envelope before any worker dispatch. A worker depends only on this envelope \u2014 never on filesystem topology, run-state format, board schema, or another worker. Path: .shapeup/<slug>/orders/r<N>-a<M>.json (or <slug>/orders/<operation>.json for non-attempt work). Every record type and payload field is DEFINED CENTRALLY in domain.schema.json ($defs + x-payload-by-worker) \u2014 this file only shapes the envelope; it never re-defines a domain entity.",
|
|
5
5
|
"type": "object",
|
|
6
|
-
"required": [
|
|
6
|
+
"required": [
|
|
7
|
+
"schema_version",
|
|
8
|
+
"order_id",
|
|
9
|
+
"worker",
|
|
10
|
+
"mode",
|
|
11
|
+
"payload"
|
|
12
|
+
],
|
|
7
13
|
"properties": {
|
|
8
|
-
"schema_version": {
|
|
14
|
+
"schema_version": {
|
|
15
|
+
"type": "integer",
|
|
16
|
+
"enum": [
|
|
17
|
+
1
|
|
18
|
+
]
|
|
19
|
+
},
|
|
9
20
|
"order_id": {
|
|
10
21
|
"type": "string",
|
|
11
|
-
"description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run
|
|
22
|
+
"description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run \u2014 it repeats across runs of the same slug, which is why run_id exists.",
|
|
12
23
|
"pattern": "^[a-z0-9][a-z0-9-]*/[a-z0-9][A-Za-z0-9.-]*$"
|
|
13
24
|
},
|
|
25
|
+
"kernel": {
|
|
26
|
+
"type": "string",
|
|
27
|
+
"description": "OPTIONAL \u2014 the absolute path of the kernel entry point (harness.mjs) of the plugin copy that compiled this order. A worker runs every kernel command as node \"<kernel>\" \u2026: a command naming a variable instead (${CLAUDE_PLUGIN_ROOT}) is refused in a headless session before any permission rule is consulted."
|
|
28
|
+
},
|
|
14
29
|
"run_id": {
|
|
15
30
|
"type": "string",
|
|
16
|
-
"description": "OPTIONAL (v1.8)
|
|
31
|
+
"description": "OPTIONAL (v1.8) \u2014 the run this dispatch belongs to, stamped by harness compile from the receipt (kernel/lib/paths.mjs). The join key the analysis plane groups on: order_id alone collides across runs of the same slug, so without this no record the pipeline writes can be attributed to a run. Absent when no readable receipt exists (a standalone dispatch in a workspace with no open run) \u2014 never invented, and never required, because an analytic field must not be able to block a build.",
|
|
17
32
|
"pattern": "^[a-z0-9][a-z0-9-]*-[0-9]{8}T[0-9]{6}Z-[0-9a-f]{8}$"
|
|
18
33
|
},
|
|
19
34
|
"compiled_at": {
|
|
20
35
|
"type": "string",
|
|
21
|
-
"description": "OPTIONAL (v1.8)
|
|
22
|
-
},
|
|
23
|
-
"worker": {
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
"
|
|
27
|
-
|
|
28
|
-
|
|
36
|
+
"description": "OPTIONAL (v1.8) \u2014 ISO timestamp of compilation. The dispatch record's own time dimension: WorkResult carries none, and the journal's timings exist only on the workflow lane, so without this a dispatch on the prose lane is timeless."
|
|
37
|
+
},
|
|
38
|
+
"worker": {
|
|
39
|
+
"$ref": "domain.schema.json#/$defs/WorkerName"
|
|
40
|
+
},
|
|
41
|
+
"mode": {
|
|
42
|
+
"type": "string",
|
|
43
|
+
"enum": [
|
|
44
|
+
"orchestrated",
|
|
45
|
+
"standalone"
|
|
46
|
+
]
|
|
47
|
+
},
|
|
48
|
+
"operation": {
|
|
49
|
+
"$ref": "domain.schema.json#/$defs/Operation"
|
|
50
|
+
},
|
|
51
|
+
"interaction": {
|
|
52
|
+
"$ref": "domain.schema.json#/$defs/Interaction"
|
|
53
|
+
},
|
|
54
|
+
"substrate": {
|
|
55
|
+
"$ref": "domain.schema.json#/$defs/Substrate"
|
|
56
|
+
},
|
|
57
|
+
"payload": {
|
|
58
|
+
"$ref": "domain.schema.json#/$defs/WorkOrderPayload"
|
|
59
|
+
}
|
|
29
60
|
}
|
|
30
61
|
}
|
package/kernel/verify/t0.mjs
CHANGED
|
@@ -34,8 +34,8 @@
|
|
|
34
34
|
//
|
|
35
35
|
// Exit code: 0 = overall green, 1 = overall red (mirrors the oracle convention), 2 = bad argv.
|
|
36
36
|
|
|
37
|
-
import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync } from "node:fs";
|
|
38
|
-
import { join, dirname } from "node:path";
|
|
37
|
+
import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync, statSync } from "node:fs";
|
|
38
|
+
import { join, dirname, basename, extname, resolve as resolvePath, relative } from "node:path";
|
|
39
39
|
import { spawnSync } from "node:child_process";
|
|
40
40
|
import { createHash } from "node:crypto";
|
|
41
41
|
import { digest } from "../probe/digest.mjs";
|
|
@@ -389,6 +389,97 @@ export function nextTrialNo(dir, round, attempt) {
|
|
|
389
389
|
return max + 1;
|
|
390
390
|
}
|
|
391
391
|
|
|
392
|
+
// --- a check rewritten to pass ----------------------------------------------------------------
|
|
393
|
+
//
|
|
394
|
+
// A fixture that names Test Surface rows (`PASS TS-02-04`, `FAIL TS-02-04 step 10: …`) reads its
|
|
395
|
+
// checks from files the scope itself writes. Measured on a live run: a row FAILed on attempt 1 —
|
|
396
|
+
// the check renamed a list and found the old name still on screen — and went green on attempt 2
|
|
397
|
+
// because the check had been rewritten to stop renaming, while the code was unchanged. Nothing
|
|
398
|
+
// recorded that the evidence had moved. So each verdict records the checks it read, by digest, and
|
|
399
|
+
// the named results they printed; a row that failed in this scope's previous trial and passes now,
|
|
400
|
+
// whose own check file changed in between, is a REVISED check — a pass the judge must read against
|
|
401
|
+
// its row before it can count. Recorded, never refused: rewriting a check that overreached is
|
|
402
|
+
// legitimate, and only a person or the judge can tell which it was.
|
|
403
|
+
|
|
404
|
+
/**
|
|
405
|
+
* The files a fixture's arguments name, with their digests — a named directory contributes its
|
|
406
|
+
* direct children. Paths resolve against the project root and are recorded relative to it.
|
|
407
|
+
*
|
|
408
|
+
* @param {string[]} commands - The fixture command lines.
|
|
409
|
+
* @param {string} cwd - Project root.
|
|
410
|
+
* @returns {Object<string,string>} Relative path → sha256 of its bytes. Empty when no argument names
|
|
411
|
+
* an existing path.
|
|
412
|
+
*/
|
|
413
|
+
export function checkFiles(commands, cwd) {
|
|
414
|
+
const out = {};
|
|
415
|
+
/**
|
|
416
|
+
* Record one file's digest under its project-relative path.
|
|
417
|
+
* @param {string} abs - Absolute path of the file.
|
|
418
|
+
* @returns {void} Nothing; an unreadable file is skipped, since it is not evidence.
|
|
419
|
+
*/
|
|
420
|
+
const add = (abs) => {
|
|
421
|
+
try { out[relative(cwd, abs)] = createHash("sha256").update(readFileSync(abs)).digest("hex"); } catch { /* unreadable — not evidence */ }
|
|
422
|
+
};
|
|
423
|
+
for (const cmd of commands || []) {
|
|
424
|
+
const tokens = String(cmd).match(/"[^"]*"|'[^']*'|\S+/g) || [];
|
|
425
|
+
for (const raw of tokens.slice(1)) {
|
|
426
|
+
const tok = raw.replace(/^["']|["']$/g, "");
|
|
427
|
+
if (!tok || tok.startsWith("-")) continue;
|
|
428
|
+
const abs = resolvePath(cwd, tok);
|
|
429
|
+
let st; try { st = statSync(abs); } catch { continue; }
|
|
430
|
+
if (st.isFile()) add(abs);
|
|
431
|
+
else if (st.isDirectory()) {
|
|
432
|
+
for (const f of readdirSync(abs).sort()) {
|
|
433
|
+
const child = join(abs, f);
|
|
434
|
+
try { if (statSync(child).isFile()) add(child); } catch { /* skip */ }
|
|
435
|
+
}
|
|
436
|
+
}
|
|
437
|
+
}
|
|
438
|
+
}
|
|
439
|
+
return out;
|
|
440
|
+
}
|
|
441
|
+
|
|
442
|
+
/**
|
|
443
|
+
* The rows a fixture's output names, and how each came out.
|
|
444
|
+
*
|
|
445
|
+
* @param {Array<{stdout?:string}>} results - Fixture results with their full stdout.
|
|
446
|
+
* @returns {Object<string,("PASS"|"FAIL")>} Row id → its result; a FAIL anywhere wins over a PASS.
|
|
447
|
+
*/
|
|
448
|
+
export function namedResults(results) {
|
|
449
|
+
const out = {};
|
|
450
|
+
for (const r of results || []) {
|
|
451
|
+
for (const line of String(r?.stdout || "").split(/\r?\n/)) {
|
|
452
|
+
const m = line.match(/^(PASS|FAIL)\s+(\S+)/);
|
|
453
|
+
if (!m) continue;
|
|
454
|
+
if (m[1] === "FAIL" || out[m[2]] !== "FAIL") out[m[2]] = m[1];
|
|
455
|
+
}
|
|
456
|
+
}
|
|
457
|
+
return out;
|
|
458
|
+
}
|
|
459
|
+
|
|
460
|
+
/**
|
|
461
|
+
* Rows that failed in the previous trial and pass now while their own check file changed.
|
|
462
|
+
*
|
|
463
|
+
* A row's check is the file whose name, without extension, is the row id — the convention a
|
|
464
|
+
* per-row check follows. A row with no such file is not reported: nothing can be said about it.
|
|
465
|
+
*
|
|
466
|
+
* @param {({check_files?:Object<string,string>, named_results?:Object<string,string>}|null)} prev
|
|
467
|
+
* The scope's previous verdict in this round, or null.
|
|
468
|
+
* @param {{check_files:Object<string,string>, named_results:Object<string,string>}} curr - This one.
|
|
469
|
+
* @returns {Array<{id:string, file:string}>} One entry per revised check.
|
|
470
|
+
*/
|
|
471
|
+
export function revisedChecks(prev, curr) {
|
|
472
|
+
if (!prev?.named_results || !prev?.check_files) return [];
|
|
473
|
+
const out = [];
|
|
474
|
+
for (const [id, now] of Object.entries(curr.named_results || {})) {
|
|
475
|
+
if (now !== "PASS" || prev.named_results[id] !== "FAIL") continue;
|
|
476
|
+
const file = Object.keys(curr.check_files || {}).find((f) => basename(f, extname(f)) === id);
|
|
477
|
+
if (!file) continue;
|
|
478
|
+
if (prev.check_files[file] !== curr.check_files[file]) out.push({ id, file });
|
|
479
|
+
}
|
|
480
|
+
return out;
|
|
481
|
+
}
|
|
482
|
+
|
|
392
483
|
/**
|
|
393
484
|
* Distill every failing command's output into AEGIS {file,line,core_message} triples.
|
|
394
485
|
* @param {{fixtures:{results:Array<{pass:boolean,stdout:string,stderr:string}>},
|
|
@@ -514,6 +605,10 @@ export async function cli(rawArgv) {
|
|
|
514
605
|
const dbProbe = runDbProbe(contract.db_probe, cwd);
|
|
515
606
|
const verdict = computeVerdict({ fixtures, dbProbe });
|
|
516
607
|
const discovered = verdict.overall === "red" ? digestFailures({ fixtures, dbProbe }) : [];
|
|
608
|
+
const checks = {
|
|
609
|
+
check_files: checkFiles(fixtures.results.map((r) => r.cmd), cwd),
|
|
610
|
+
named_results: namedResults(fixtures.results),
|
|
611
|
+
};
|
|
517
612
|
|
|
518
613
|
// ---- the ratchet ---------------------------------------------------------------------
|
|
519
614
|
// `current` is the incumbent: the score of the most recent trial whose TREE is the one on disk
|
|
@@ -526,6 +621,11 @@ export async function cli(rawArgv) {
|
|
|
526
621
|
const verdictBetter = better(s, baseline ? baseline.score : null);
|
|
527
622
|
const crashed = fixtures.results.some((r) => r.error) || !!dbProbe?.error;
|
|
528
623
|
const { status, action } = decideStatus(verdictBetter, crashed);
|
|
624
|
+
// The scope's previous trial in this round, read back for the checks it recorded.
|
|
625
|
+
const prevTrial = [...priorTrials].reverse().find((tr) => tr.round === round);
|
|
626
|
+
let prevVerdict = null;
|
|
627
|
+
if (prevTrial?.artifact) { try { prevVerdict = JSON.parse(readFileSync(join(outDir, prevTrial.artifact), "utf8")); } catch { /* none */ } }
|
|
628
|
+
const revised = revisedChecks(prevVerdict, checks);
|
|
529
629
|
|
|
530
630
|
// The run key, read from the receipt that lives in the run root this script was pointed at.
|
|
531
631
|
// `--out` IS that root, so identity comes from the receipt rather than from parsing a slug back
|
|
@@ -550,6 +650,9 @@ export async function cli(rawArgv) {
|
|
|
550
650
|
...verdict,
|
|
551
651
|
score: s,
|
|
552
652
|
discovered_tasks: discovered,
|
|
653
|
+
...(Object.keys(checks.check_files).length ? { check_files: checks.check_files } : {}),
|
|
654
|
+
...(Object.keys(checks.named_results).length ? { named_results: checks.named_results } : {}),
|
|
655
|
+
...(revised.length ? { revised_checks: revised } : {}),
|
|
553
656
|
});
|
|
554
657
|
|
|
555
658
|
// The tree operation. `--no-ratchet` leaves the working tree exactly as the attempt left it —
|
package/package.json
CHANGED
|
@@ -20,6 +20,8 @@ same artifacts out.
|
|
|
20
20
|
|
|
21
21
|
## Input contract — the WorkOrder
|
|
22
22
|
|
|
23
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
24
|
+
|
|
23
25
|
Invoked as `--order <path>`. Fields you may rely on (absent = unknown; surface it, never guess):
|
|
24
26
|
|
|
25
27
|
| Field | What it is |
|
|
@@ -69,10 +71,10 @@ its phase; templates live in `assets/templates/`.
|
|
|
69
71
|
6 TASKS atomic, ordered, executable → tasks/ (LOCAL root; the one uncommitted branch
|
|
70
72
|
of the tree — regenerable, machine-local) [references/task-generation.md]
|
|
71
73
|
7 DERIVE+LINT mechanical, not yours to grade:
|
|
72
|
-
node "
|
|
74
|
+
node "<kernel>" reduce board --slug <slug> --write
|
|
73
75
|
(unlocks = depends_on inverse; Σ hours; critical path; appetite arithmetic —
|
|
74
76
|
overflow is a fact you REPORT for the caller's HAMMER gate, never resolve)
|
|
75
|
-
node "
|
|
77
|
+
node "<kernel>" verify spec --slug <slug>
|
|
76
78
|
(structure, wikilinks, edge symmetry — fix reds, then re-run; you never
|
|
77
79
|
self-grade with a hand-walked checklist. BREADBOARD-PLACE / BREADBOARD-UI:
|
|
78
80
|
add the screen or defer the Place; never fold it into another screen)
|
|
@@ -187,7 +189,7 @@ status flips for built work (ingest's job), scope contracts (scope-architect's),
|
|
|
187
189
|
/ba-pitch-analyzer --order .shapeup/checkout-vnpay/orders/analyze.json
|
|
188
190
|
|
|
189
191
|
# Standalone — the preamble shim compiles the order (mode: standalone, pause_gates: true):
|
|
190
|
-
# node "
|
|
192
|
+
# node "<kernel>" compile --operation analyze --slug <slug> \
|
|
191
193
|
# --worker ba-pitch-analyzer --payload '{"pitch": "docs/pitch.md", "lens": "standard"}'
|
|
192
194
|
/ba-pitch-analyzer docs/pitch.md # operation: analyze, lens judged
|
|
193
195
|
/ba-pitch-analyzer --lens standard docs/pitch.md # lens pinned
|
|
@@ -17,6 +17,8 @@ the ship report's census table.
|
|
|
17
17
|
|
|
18
18
|
## Input contract — the WorkOrder
|
|
19
19
|
|
|
20
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
21
|
+
|
|
20
22
|
| Field | What it is |
|
|
21
23
|
|---|---|
|
|
22
24
|
| `operation` | `map-scopes` — the only operation this skill has. It covers first slicing after the board exists, folding discovered items in, and re-slicing a stuck scope; the payload says which of those you are doing |
|
|
@@ -95,7 +97,7 @@ the ship report's census table.
|
|
|
95
97
|
hill_phase: "UPHILL_UNKNOWN" — ALWAYS; phase is derived from
|
|
96
98
|
T0/T1 facts later,
|
|
97
99
|
never authored
|
|
98
|
-
4 LINT node "
|
|
100
|
+
4 LINT node "<kernel>" verify spec --slug <slug>
|
|
99
101
|
→ PA1 (directory alignment), PA2 (>~15 files), DISJOINT (undeclared overlap),
|
|
100
102
|
SCOPE-ANCHOR (empty/unresolvable use_cases), TIER-DIRECTION (a task id in a
|
|
101
103
|
committed contract), SCOPE-DEPS (depends_on naming a scope that isn't here).
|
|
@@ -52,13 +52,15 @@ INPUT: run's finished/unfinished scopes + baseline + census sources
|
|
|
52
52
|
|
|
53
53
|
## GATE H0 — Census
|
|
54
54
|
|
|
55
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
56
|
+
|
|
55
57
|
**Purpose:** Gather every open item into one list before judging any of them. Never judge
|
|
56
58
|
piecemeal — a partial view produces a wrong cut.
|
|
57
59
|
|
|
58
60
|
```
|
|
59
61
|
H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns X" or "X is
|
|
60
62
|
scope Y's", run
|
|
61
|
-
node "
|
|
63
|
+
node "<kernel>" probe owner --slug <slug> [--path <p>]...
|
|
62
64
|
and cite its row. With no --path it answers for every engine and entry call site the wiring
|
|
63
65
|
map names plus the profile's entry point; `writers: []` is an unowned seam and `missing`
|
|
64
66
|
lists seams the wiring names that are not on disk — owned but never written. A census that
|
|
@@ -70,7 +72,7 @@ H0.1 Unresolved scopes (breaker cases only):
|
|
|
70
72
|
- scopes with hammer_proposals (attempt_budget exhausted) → CARRY candidates. Exhaustion
|
|
71
73
|
is DERIVED, never read off `t0/verdicts/*.json` directly — a compiled order or a T0
|
|
72
74
|
verdict is writable by the very scope being judged and proves nothing on its own. Run
|
|
73
|
-
node "
|
|
75
|
+
node "<kernel>" probe attempts --slug <slug> \
|
|
74
76
|
--scope <scope-id> --round <n> --attempt-budget <n>
|
|
75
77
|
and cite its `spent`/`tripped` fields (exit 1 = tripped) — an attempt counts only when a
|
|
76
78
|
dispatch receipt AND either a leg-completion row or a WorkResult attest it, so a leg still
|
|
@@ -81,7 +83,7 @@ H0.2 QA findings (qa-edge-hunter's hunt-report.md, when present) — all `~` by
|
|
|
81
83
|
H0.3 Discovered-task ledger entries still open (discovery/ledger.md, `[+]`/`~` unresolved).
|
|
82
84
|
H0.4 Attempt-budget hammer proposals (scopes that exhausted their T0 attempts during BUILD).
|
|
83
85
|
H0.4b Requirements with no PASS evidence — the pitch clauses the run never showed working. Run
|
|
84
|
-
node "
|
|
86
|
+
node "<kernel>" probe requirements --slug <slug> --format table
|
|
85
87
|
and take its `no evidence` rows; cite the row, the same way H0.0 cites ownership. Each is a
|
|
86
88
|
census item carrying its source clause (`REQ-12 ← shaping.md R12`). A `cut` row is an answer
|
|
87
89
|
the PO already gave — not an item. An inconsistency row (a criterion anchored to a
|
|
@@ -200,8 +202,8 @@ with a green census could only be recorded as `ask`. Still a proposal: the file
|
|
|
200
202
|
/scope-hammer --slug checkout-vnpay --unattended
|
|
201
203
|
|
|
202
204
|
# The ownership query every census claim cites (H0.0)
|
|
203
|
-
node "
|
|
204
|
-
node "
|
|
205
|
+
node "<kernel>" probe owner --slug checkout-vnpay --format table
|
|
206
|
+
node "<kernel>" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
|
|
205
207
|
```
|
|
206
208
|
|
|
207
209
|
### Flags
|
|
@@ -38,6 +38,8 @@ return as a WorkResult.
|
|
|
38
38
|
|
|
39
39
|
## Input contract — the WorkOrder
|
|
40
40
|
|
|
41
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
42
|
+
|
|
41
43
|
| Field | What it is |
|
|
42
44
|
|---|---|
|
|
43
45
|
| `operation` | `wire` (author/refresh the wiring map after `analyze`, before `map-scopes`) |
|
|
@@ -152,5 +154,5 @@ seam, or an engine with no attachment path, and why). You never touch spec docs,
|
|
|
152
154
|
# The reachability oracle is the ORCHESTRATOR's, run advisory at L1b — not part of your craft.
|
|
153
155
|
# Standalone, you MAY preview it after writing the map (it self-skips arms whose artifacts are
|
|
154
156
|
# absent, and is near-vacuous pre-build since the engine code does not exist yet):
|
|
155
|
-
# node "
|
|
157
|
+
# node "<kernel>" verify trace --slug checkout-vnpay
|
|
156
158
|
```
|
|
@@ -29,6 +29,8 @@ criterion with no collected evidence is a **FAIL**, never a pass-by-assumption.
|
|
|
29
29
|
|
|
30
30
|
## Input contract — the WorkOrder
|
|
31
31
|
|
|
32
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
33
|
+
|
|
32
34
|
Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inferred):
|
|
33
35
|
|
|
34
36
|
| Field | What it is |
|
|
@@ -39,6 +41,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
|
|
|
39
41
|
| `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
|
|
40
42
|
| `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
|
|
41
43
|
| `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
|
|
44
|
+
| `payload.revised_checks[]` | Rows that FAILed in one trial of this round and PASS in a later one whose own check file changed in between — a pass obtained by editing the check. Read each listed file against its row's Expect before you let its PASS count; a check that no longer asserts what the row asks makes that row a FAIL, and the bug names the check. Absent → no check was rewritten |
|
|
42
45
|
| `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
|
|
43
46
|
| `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
|
|
44
47
|
| `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
|
|
@@ -85,9 +88,11 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
|
|
|
85
88
|
- Contract work: send real requests, compare field-by-field.
|
|
86
89
|
- **A fixture that names a row is evidence for that row.** When a T0 artifact you cite, or the
|
|
87
90
|
build gate, carries output that names a Test Surface row by id — `PASS TS-05-05`, or
|
|
88
|
-
`FAIL TS-05-05 step 4: …` — grade that row on it
|
|
89
|
-
|
|
90
|
-
|
|
91
|
+
`FAIL TS-05-05 step 4: …` — grade that row on it, after reading the check that printed it: a
|
|
92
|
+
named PASS confirms only when that check asserts the row's Expect; a check weaker than its row
|
|
93
|
+
(it opens the dialog, the row says pre-filled) is not evidence for the row, and grading it PASS
|
|
94
|
+
is the generator grading itself. A named FAIL is a FAIL whose bug is that line. This is how a
|
|
95
|
+
device row is evidenced when you cannot drive the app; never a reason not to when you can.
|
|
91
96
|
- No evidence collected = recorded "NO EVIDENCE" → FAILs at verdict.
|
|
92
97
|
|
|
93
98
|
**VERDICT.**
|
|
@@ -222,7 +227,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
|
|
|
222
227
|
/spec-evaluator --order .shapeup/checkout-vnpay/orders/evaluate-r2.json
|
|
223
228
|
|
|
224
229
|
# Standalone — the preamble shim compiles a minimal order, then the single code path runs:
|
|
225
|
-
# node "
|
|
230
|
+
# node "<kernel>" compile --operation evaluate --slug <slug> \
|
|
226
231
|
# --worker spec-evaluator [--payload '{"dimensions": [...], "run_cmd": "..."}']
|
|
227
232
|
/spec-evaluator --spec shapeup/checkout-vnpay/spec/ --task TASK-007
|
|
228
233
|
/spec-evaluator --spec shapeup/checkout-vnpay/spec/ --feature checkout-vnpay --single-pass
|
|
@@ -230,7 +235,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
|
|
|
230
235
|
|
|
231
236
|
Standalone keeps `--task` (per-task check, not round-gated) and `--single-pass` (feature-level)
|
|
232
237
|
— the shim maps them onto the order's payload; missing run command → ask. After writing the
|
|
233
|
-
WorkResult, run `node "
|
|
238
|
+
WorkResult, run `node "<kernel>" reduce ingest <result path>` and show its
|
|
234
239
|
summary — standalone has no orchestrator to ingest for you.
|
|
235
240
|
|
|
236
241
|
---
|
|
@@ -16,6 +16,8 @@ it does not exist for you.
|
|
|
16
16
|
|
|
17
17
|
## Input contract — the WorkOrder
|
|
18
18
|
|
|
19
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
20
|
+
|
|
19
21
|
You are invoked as `--order <path>` pointing at a schema-valid WorkOrder. Fields you may
|
|
20
22
|
rely on (anything absent = **unknown**; never invent it):
|
|
21
23
|
|
|
@@ -194,8 +196,8 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
|
|
|
194
196
|
|
|
195
197
|
# Standalone — the preamble shim compiles a minimal WorkOrder from the flags, then the
|
|
196
198
|
# single code path above runs. Requires the harness scripts (plugin install):
|
|
197
|
-
# node "
|
|
198
|
-
# node "
|
|
199
|
+
# node "<kernel>" compile --task TASK-003 --slug checkout-vnpay
|
|
200
|
+
# node "<kernel>" compile --next --slug checkout-vnpay
|
|
199
201
|
/task-executor --spec shapeup/checkout-vnpay/spec/ --task TASK-003
|
|
200
202
|
/task-executor --spec shapeup/checkout-vnpay/spec/ --next
|
|
201
203
|
```
|
|
@@ -203,6 +205,6 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
|
|
|
203
205
|
Standalone shim: derive `<slug>` from the `--spec` path (`shapeup/<slug>/spec`),
|
|
204
206
|
run `harness compile` with the matching flags (mode becomes `standalone`), then proceed
|
|
205
207
|
against the compiled order exactly as if dispatched. After writing the WorkResult, run
|
|
206
|
-
`node "
|
|
208
|
+
`node "<kernel>" reduce ingest <result path>` yourself and show the user its
|
|
207
209
|
summary — standalone has no orchestrator to ingest for you. One code path inside; two entry
|
|
208
210
|
points outside.
|
|
@@ -898,7 +898,7 @@ async function crossGate(gateId, phaseName, validDecisions, ctx) {
|
|
|
898
898
|
// The gate ledger keys a per-round crossing (L2, L3) on gate id + round, so it needs the round
|
|
899
899
|
// whenever the caller already has one to show in the block — the same value `ctx.round` carries
|
|
900
900
|
// for display, threaded through rather than re-derived.
|
|
901
|
-
const roundFlag = ctx?.round != null ? ` --round ${ctx.round}` : "";
|
|
901
|
+
const roundFlag = (ctx?.round != null ? ` --round ${ctx.round}` : "") + (ctx?.verdict ? ` --verdict ${ctx.verdict}` : "");
|
|
902
902
|
const g = await cmd(`gate --resolve ${gateId} --slug ${slug}${roundFlag} ${answersFlag(args.answers)}`.trim(), phaseName, `gate:${gateId}`);
|
|
903
903
|
if (g.exit_code === 4) return { stop: paused(gateId, validDecisions, ctx) };
|
|
904
904
|
if (g.exit_code === 5) return { stop: aborted(gateId, g.detail || `GATE ${gateId} aborted`) };
|