shapeup-sdlc 3.12.0 → 3.14.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.12.0",
4
+ "version": "3.14.0",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -37,7 +37,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
37
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
38
38
  | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
39
39
  | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade |
40
- | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC |
40
+ | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
41
41
 
42
42
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
43
43
 
@@ -69,7 +69,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
69
69
  ## Setup & Execution
70
70
 
71
71
  - Orders/results live in `.shapeup/<slug>/orders|results/`; the envelope schemas ship with the plugin runtime, not with any individual skill, so every worker validates against the same copy.
72
- - The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf.
72
+ - The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf. Two facts about how grants reach the run's workers, both measured: a worker dispatched inside a run takes its permissions from the project's settings file, not from flags given to the launching session, so a tool the workers need belongs in `permissions.allow`; and every work order names the plugin's kernel by absolute path, because a command that spells it through a variable is refused in a headless session before any permission rule is read.
73
73
  - The grant is necessary but sits under two more layers this plugin cannot reach either. A fresh
74
74
  checkout is an **untrusted workspace**, and Claude Code discards the whole permission grant — every
75
75
  rule in it, not only this one — until the workspace is trusted; the installer detects that state and
@@ -37,7 +37,7 @@ import { readRunId, dispatchReceipts, legLedger, readReceipt, receipt } from "./
37
37
  // --spec-overridden directory, and the import is the convention-derived default.
38
38
  import {
39
39
  tasksDir, specDir as defaultSpecDir, roundLedger, trials, verdictsDir, ordersDir,
40
- relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir,
40
+ relShared, relLocal, globLocal, globShared, relKnowledgeBase, resultsDir, scopesDir, localRoot,
41
41
  } from "./lib/paths.mjs";
42
42
  import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT, reqId } from "./lib/contract.mjs";
43
43
  import { writeActiveOrder } from "./probe/resume.mjs";
@@ -527,6 +527,57 @@ export function verdictBugs(cwd, slug, round) {
527
527
  return v.bugs.filter((b) => !refuted.has(String(b?.id)) && !refuted.has(String(b?.criterion)));
528
528
  }
529
529
 
530
+ /**
531
+ * The previous round's FAILED CRITERIA that the judge filed no bug for, as bug entries.
532
+ *
533
+ * A verdict carries two lists: the criteria it graded and the bugs it filed. Only the second
534
+ * reached the next round, so a criterion graded FAIL — "no evidence: no check names this row", a
535
+ * device row the judge could not observe — with no matching bug handed the fix round nothing to do.
536
+ * Measured: three fix rounds in a row compiled with no bugs over a FAIL verdict, and the run ended
537
+ * where it started. Every failed criterion is a thing the round must address, so each one without a
538
+ * bug becomes one, addressed to the scope whose contract lists the use case the criterion names.
539
+ *
540
+ * @param {string} cwd - Project root.
541
+ * @param {string} slug - Feature slug.
542
+ * @param {number} [round] - The round being compiled.
543
+ * @returns {Array<object>} `{criterion, severity, expected, actual, source:"criterion", scope_id?}`
544
+ * per uncovered FAIL criterion; [] for round 1, a PASS, or an unreadable result.
545
+ */
546
+ export function criteriaBugs(cwd, slug, round) {
547
+ if (!round || round < 2) return [];
548
+ const p = join(resultsDir(cwd, slug), `evaluate-r${round - 1}.json`);
549
+ if (!existsSync(p)) return [];
550
+ let v;
551
+ try { v = JSON.parse(readFileSync(p, "utf8"))?.verdict; } catch { return []; }
552
+ if (v?.overall !== "FAIL" || !Array.isArray(v.criteria)) return [];
553
+ const filed = (Array.isArray(v.bugs) ? v.bugs : []).map((b) => String(b?.criterion ?? ""));
554
+ const refuted = new Set((Array.isArray(v.refuted) ? v.refuted : [])
555
+ .flatMap((r) => [r?.id, r?.ac_id, r?.criterion, typeof r === "string" ? r : null]).filter(Boolean).map(String));
556
+ const owners = new Map();
557
+ for (const { contract, id } of readAllContracts(scopesDir(cwd, slug))) {
558
+ const sid = contract?.scope_id || id;
559
+ for (const uc of Array.isArray(contract?.use_cases) ? contract.use_cases : []) if (!owners.has(uc)) owners.set(uc, sid);
560
+ }
561
+ const out = [];
562
+ for (const c of v.criteria) {
563
+ if (c?.verdict !== "FAIL" || typeof c.criterion !== "string") continue;
564
+ const name = c.criterion;
565
+ if (refuted.has(name)) continue;
566
+ if (filed.some((f) => f && (f === name || f.includes(name) || name.includes(f)))) continue;
567
+ const uc = (name.match(/\bUC-[A-Za-z0-9_-]+/) || [])[0];
568
+ const owner = uc ? owners.get(uc) : undefined;
569
+ out.push({
570
+ criterion: name,
571
+ severity: "major",
572
+ expected: "the criterion is met, with evidence the judge can cite",
573
+ actual: String(c.evidence ?? "graded FAIL").slice(0, 400),
574
+ source: "criterion",
575
+ ...(owner ? { scope_id: owner } : {}),
576
+ });
577
+ }
578
+ return out;
579
+ }
580
+
530
581
  /**
531
582
  * The previous round's RED BUILD GATE, as bug entries the fix round can act on.
532
583
  *
@@ -748,6 +799,38 @@ export function t0ArtifactsFor(cwd, slug, round) {
748
799
  return { artifacts, missing };
749
800
  }
750
801
 
802
+ // --- checks a scope rewrote between a failing trial and a passing one --------------------------
803
+
804
+ /**
805
+ * Every revised check any trial of the round recorded, with the scope it belongs to.
806
+ *
807
+ * Read across all of the round's trials rather than off the cited verdict alone: a check rewritten
808
+ * on attempt 2 and re-run unchanged on attempt 3 leaves no mark on attempt 3's verdict, which is the
809
+ * one the judge cites.
810
+ *
811
+ * @param {string} cwd - Project root.
812
+ * @param {string} slug - Feature slug.
813
+ * @param {number} [round] - The round being evaluated; omitted, every round.
814
+ * @returns {Array<{scope_id:string, id:string, file:string}>} Deduplicated by scope, id and file.
815
+ */
816
+ export function revisedChecksFor(cwd, slug, round) {
817
+ const seen = new Set();
818
+ const out = [];
819
+ for (const tr of readTrials(trials(cwd, slug))) {
820
+ if (round != null && tr.round !== round) continue;
821
+ if (!tr.artifact) continue;
822
+ let v;
823
+ try { v = JSON.parse(readFileSync(join(localRoot(cwd, slug), tr.artifact), "utf8")); } catch { continue; }
824
+ for (const r of Array.isArray(v?.revised_checks) ? v.revised_checks : []) {
825
+ const key = `${tr.scope_id}\u0000${r.id}\u0000${r.file}`;
826
+ if (seen.has(key)) continue;
827
+ seen.add(key);
828
+ out.push({ scope_id: tr.scope_id, id: r.id, file: r.file });
829
+ }
830
+ }
831
+ return out;
832
+ }
833
+
751
834
  // --- the launch evidence the judge grades `[ui]` rows against ---------------------------------
752
835
  //
753
836
  // WHY THE KERNEL DERIVES IT. A `[ui]` criterion is graded on the RUNNING app, and the evaluator's
@@ -851,6 +934,13 @@ export function compileOrder({
851
934
  const order = {
852
935
  schema_version: 1,
853
936
  order_id: `${slug}/${suffix}`,
937
+ // WHERE THE KERNEL IS, AS A PATH A PERMISSION RULE CAN MATCH. Workers were told to run
938
+ // `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" …`. In a headless session that variable is not
939
+ // expanded for them, and a command carrying an unexpanded variable is refused before any rule is
940
+ // consulted ("Contains expansion") — so every kernel query a worker's contract requires was
941
+ // refused, and workers reported it as "not permitted". The absolute path of this very file's
942
+ // kernel is known here, and quoted it matches the grant `init` writes.
943
+ kernel: join(HERE, "harness.mjs"),
854
944
  // THE TWO ANALYTIC FIELDS, and why they are on the ORDER rather than the result.
855
945
  //
856
946
  // `order_id` identifies a dispatch within a run and repeats across runs of the same slug, so
@@ -1009,7 +1099,7 @@ export async function cli(rawArgv) {
1009
1099
  // this line, and none of them can pass a payload to a build order (see the banner above).
1010
1100
  // Two sources, one channel: the judge's cited defects and the build gate's failing steps.
1011
1101
  const bugs = scope
1012
- ? bugsForScope([...verdictBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
1102
+ ? bugsForScope([...verdictBugs(cwd, slug, round), ...criteriaBugs(cwd, slug, round), ...buildBugs(cwd, slug, round)], scope.scope_id, scopeSubstrates(cwd, slug))
1013
1103
  : [];
1014
1104
 
1015
1105
  // A ROUND CARRYING CITED DEFECTS IS A `fix`, AND THE ORDER HAS TO SAY SO.
@@ -1131,6 +1221,10 @@ export async function cli(rawArgv) {
1131
1221
  }
1132
1222
  // The launch evidence, for every lane (see launchEvidenceFor). Same rule as above: an explicit
1133
1223
  // `--payload` value outranks the derivation.
1224
+ if (operation === "evaluate" && payloadExtra.revised_checks === undefined) {
1225
+ const revised = revisedChecksFor(cwd, slug, round);
1226
+ if (revised.length) payloadExtra.revised_checks = revised;
1227
+ }
1134
1228
  if (operation === "evaluate") {
1135
1229
  const ev = launchEvidenceFor(cwd, slug, round);
1136
1230
  if (ev.build_gate !== undefined && payloadExtra.build_gate === undefined) payloadExtra.build_gate = ev.build_gate;
package/kernel/gate.mjs CHANGED
@@ -364,7 +364,7 @@ export function discover({ cwd = process.cwd(), file = null, preset = null, slug
364
364
  export const ARGV_SPEC = {
365
365
  usage: "harness.mjs gate (--init | --list | --verify | --resolve <gate-id>) [--preset <name>] " +
366
366
  "[--file <path>] [--slug <slug>] [--cwd <dir>] [--out <path>] [--by <who>] " +
367
- "[--auto-level <level>] [--tiny] [--no-qa] [--round <n>]",
367
+ "[--auto-level <level>] [--tiny] [--no-qa] [--round <n>] [--verdict <pass|fail|…>]",
368
368
  _: { arity: 0, max: 0, name: "(no positional operands)" },
369
369
  cwd: { type: "path" },
370
370
  init: { type: "flag" },
@@ -383,8 +383,32 @@ export const ARGV_SPEC = {
383
383
  // ledger row needs it to key a `GateDecision` node uniquely per crossing. Round-independent gates
384
384
  // (L0, L1a, …) simply omit it and the row carries `round: null`.
385
385
  round: { type: "int", min: 1 },
386
+ // The verdict the round being crossed actually reached, when there is one. A preset answers L3
387
+ // before any verdict exists — "loop" on a FAIL — and its note says so; recorded without the
388
+ // verdict, a PASS round's row read "FAIL → fix round", the opposite of what happened.
389
+ verdict: { type: "str" },
386
390
  };
387
391
 
392
+ /**
393
+ * The note a gate row carries — the answer's own, unless the verdict it was crossed over makes that
394
+ * note false.
395
+ *
396
+ * L3's answers are written for a verdict that has not happened yet: `loop` means "on a FAIL, run
397
+ * the next round". Crossed over a PASS, nothing loops, and the preset's "FAIL → fix round" note
398
+ * would put a failed round on the record for a round that passed.
399
+ *
400
+ * @param {{gate:string, decision?:string, note?:string, reason?:string}} r - The resolved answer.
401
+ * @param {(string|null)} verdict - The round's verdict, lower-cased, or null when none was given.
402
+ * @returns {(string|null)} The note to record.
403
+ */
404
+ export function gateNote(r, verdict) {
405
+ const own = r.note ?? r.reason ?? null;
406
+ if (r.gate === "L3" && verdict === "pass") {
407
+ return `verdict PASS — the "${r.decision}" answer applies only to a failed round; the run goes on to QA and GATE H`;
408
+ }
409
+ return own;
410
+ }
411
+
388
412
  function out(obj, code = 0) {
389
413
  console.log(JSON.stringify(obj, null, 2));
390
414
  process.exit(code);
@@ -468,11 +492,13 @@ export function cli(rawArgv) {
468
492
  // THE ROW CARRIES THE RUN KEY. It did not, and the export stamped the current run's key onto
469
493
  // every row it found — a prior run's sign-off became this run's in the one table that answers
470
494
  // "was this ship signed off". Driven on a two-run fixture before it was fixed.
495
+ const verdict = args.verdict ? String(args.verdict).toLowerCase() : null;
471
496
  appendGateLedger(cwd, args.slug, {
472
497
  at: new Date().toISOString(), run_id: readRunId(cwd, args.slug),
473
498
  gate: r.gate, status: r.status, decision: r.decision ?? null,
474
- source: r.source ?? found.source, note: r.note ?? r.reason ?? null,
499
+ source: r.source ?? found.source, note: gateNote(r, verdict),
475
500
  round: args.round ?? null,
501
+ ...(verdict ? { verdict } : {}),
476
502
  });
477
503
  }
478
504
  if (r.status === "ask") out({ ...r, ok: false }, 4);
@@ -33,6 +33,11 @@ const PATTERNS = [
33
33
  // ("ERROR in the build pipeline", "ERROR in test suite failed to run") is left unmatched
34
34
  // instead of handing back a fabricated file.
35
35
  { re: /^(?:ERROR|WARNING)\s+in\s+(\.{1,2}\/[^\s:]*|[^\s:]+\.[A-Za-z0-9]{1,10})\b/i, kind: "compiler-diagnostic" },
36
+ // hvigor / ArkTS compiler: the message and the location arrive on ONE line,
37
+ // "Error Message: Expected 5 arguments, but got 3. At File: /abs/path/Foo.test.ets:75:33"
38
+ // — the commonest failure on that toolchain, and one no other pattern here anchored, so every red
39
+ // compile handed the next attempt an empty error list.
40
+ { re: /^\s*Error Message:\s*(.+?)\s+At File:\s*(.+?):(\d+):\d+\s*$/, kind: "arkts-compiler" },
36
41
  // A test that FAILED BY NAME, with no file:line: "FAIL TS-05-05 step 4: no text 'Bread' on screen"
37
42
  // or jest's "FAIL src/cart.test.js". Runners that drive an app from outside it (a device flow, an
38
43
  // end-to-end script) report a case this way and nothing else, and the line is the whole signal —
@@ -77,6 +82,11 @@ export function digest(rawText) {
77
82
  pendingMessage = coreMessage(m[1]);
78
83
  continue; // wait for the stack frame that follows to get a file:line
79
84
  }
85
+ if (kind === "arkts-compiler") {
86
+ triples.push({ file: m[2].trim(), line: Number(m[3]), core_message: coreMessage(m[1]), kind });
87
+ pendingMessage = null;
88
+ break;
89
+ }
80
90
  const named = kind === "named-test-failure";
81
91
  const file = named ? (/[\\/]|\.[A-Za-z0-9]{1,10}$/.test(m[1]) ? m[1] : null) : m[1]?.trim();
82
92
  const lineNo = !named && m[2] ? Number(m[2]) : null;
@@ -194,7 +194,10 @@ export function ratchetReport(trials) {
194
194
  }
195
195
  if (scopeMonotone) monotone++;
196
196
  }
197
- const greenAt = seq.findIndex((t) => t.score && t.score.fixtures_total > 0 && t.score.fixtures_passed === t.score.fixtures_total && t.score.regressions === 0);
197
+ // A score with no `regressions` field was not measured for regressions — T0 writes one only when
198
+ // it has a baseline to regress against — which is not the same as having some. Requiring an
199
+ // explicit 0 reported "no scope reached green" over a run whose every scope was green.
200
+ const greenAt = seq.findIndex((t) => t.score && t.score.fixtures_total > 0 && t.score.fixtures_passed === t.score.fixtures_total && (t.score.regressions ?? 0) === 0);
198
201
  if (greenAt !== -1) toGreen.push(greenAt + 1);
199
202
  per_scope.push({
200
203
  scope_id, trials: seq.length,
@@ -280,7 +280,7 @@ export function buildReport(facts) {
280
280
  L.push("| scope | fixtures | regressions | trials | last status | delta |", "|---|---|---|---|---|---|");
281
281
  for (const s of t0) {
282
282
  const f = s.score ? `${s.score.fixtures_passed}/${s.score.fixtures_total}` : "—";
283
- const r = s.score ? String(s.score.regressions) : "—";
283
+ const r = s.score?.regressions != null ? String(s.score.regressions) : "—";
284
284
  L.push(`| ${s.scope_id} | ${f} | ${r} | ${s.trials} | ${s.status} | ${s.delta || "—"} |`);
285
285
  }
286
286
  L.push("");
@@ -362,6 +362,7 @@
362
362
  "run_cmd",
363
363
  "launch_cmd",
364
364
  "build_gate",
365
+ "revised_checks",
365
366
  "t0_artifacts",
366
367
  "browser",
367
368
  "tasks"
@@ -2429,6 +2430,19 @@
2429
2430
  "type": "string",
2430
2431
  "description": "spec-evaluator: the project profile's launch probe — installs the built artifact, starts it and asserts the first screen. Derived by `harness compile` from project-profile.md, absent when the profile declares none. Where `run_cmd` is only a build, this is how the app is brought up; a non-zero exit is a finding, not a reason to guess another way."
2431
2432
  },
2433
+ "revised_checks": {
2434
+ "type": "array",
2435
+ "items": {
2436
+ "type": "object",
2437
+ "properties": {
2438
+ "scope_id": { "type": "string" },
2439
+ "id": { "type": "string" },
2440
+ "file": { "type": "string" }
2441
+ },
2442
+ "required": ["id", "file"]
2443
+ },
2444
+ "description": "spec-evaluator: rows that FAILed in one of the round's T0 trials and PASS in a later one whose own check file (named by the row id) changed in between. Derived by harness compile. Each must be read against its row before its PASS counts; absent when none."
2445
+ },
2432
2446
  "build_gate": {
2433
2447
  "type": "string",
2434
2448
  "description": "spec-evaluator: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
@@ -1,30 +1,61 @@
1
1
  {
2
2
  "$id": "work-order.schema.json",
3
3
  "title": "WorkOrder",
4
- "description": "The orchestrator → worker envelope (pure-skill architecture v1.0). Compiled by kernel/compile.mjs, validated by harness verify envelope before any worker dispatch. A worker depends only on this envelope — never on filesystem topology, run-state format, board schema, or another worker. Path: .shapeup/<slug>/orders/r<N>-a<M>.json (or <slug>/orders/<operation>.json for non-attempt work). Every record type and payload field is DEFINED CENTRALLY in domain.schema.json ($defs + x-payload-by-worker) — this file only shapes the envelope; it never re-defines a domain entity.",
4
+ "description": "The orchestrator \u2192 worker envelope (pure-skill architecture v1.0). Compiled by kernel/compile.mjs, validated by harness verify envelope before any worker dispatch. A worker depends only on this envelope \u2014 never on filesystem topology, run-state format, board schema, or another worker. Path: .shapeup/<slug>/orders/r<N>-a<M>.json (or <slug>/orders/<operation>.json for non-attempt work). Every record type and payload field is DEFINED CENTRALLY in domain.schema.json ($defs + x-payload-by-worker) \u2014 this file only shapes the envelope; it never re-defines a domain entity.",
5
5
  "type": "object",
6
- "required": ["schema_version", "order_id", "worker", "mode", "payload"],
6
+ "required": [
7
+ "schema_version",
8
+ "order_id",
9
+ "worker",
10
+ "mode",
11
+ "payload"
12
+ ],
7
13
  "properties": {
8
- "schema_version": { "type": "integer", "enum": [1] },
14
+ "schema_version": {
15
+ "type": "integer",
16
+ "enum": [
17
+ 1
18
+ ]
19
+ },
9
20
  "order_id": {
10
21
  "type": "string",
11
- "description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run — it repeats across runs of the same slug, which is why run_id exists.",
22
+ "description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run \u2014 it repeats across runs of the same slug, which is why run_id exists.",
12
23
  "pattern": "^[a-z0-9][a-z0-9-]*/[a-z0-9][A-Za-z0-9.-]*$"
13
24
  },
25
+ "kernel": {
26
+ "type": "string",
27
+ "description": "OPTIONAL \u2014 the absolute path of the kernel entry point (harness.mjs) of the plugin copy that compiled this order. A worker runs every kernel command as node \"<kernel>\" \u2026: a command naming a variable instead (${CLAUDE_PLUGIN_ROOT}) is refused in a headless session before any permission rule is consulted."
28
+ },
14
29
  "run_id": {
15
30
  "type": "string",
16
- "description": "OPTIONAL (v1.8) — the run this dispatch belongs to, stamped by harness compile from the receipt (kernel/lib/paths.mjs). The join key the analysis plane groups on: order_id alone collides across runs of the same slug, so without this no record the pipeline writes can be attributed to a run. Absent when no readable receipt exists (a standalone dispatch in a workspace with no open run) — never invented, and never required, because an analytic field must not be able to block a build.",
31
+ "description": "OPTIONAL (v1.8) \u2014 the run this dispatch belongs to, stamped by harness compile from the receipt (kernel/lib/paths.mjs). The join key the analysis plane groups on: order_id alone collides across runs of the same slug, so without this no record the pipeline writes can be attributed to a run. Absent when no readable receipt exists (a standalone dispatch in a workspace with no open run) \u2014 never invented, and never required, because an analytic field must not be able to block a build.",
17
32
  "pattern": "^[a-z0-9][a-z0-9-]*-[0-9]{8}T[0-9]{6}Z-[0-9a-f]{8}$"
18
33
  },
19
34
  "compiled_at": {
20
35
  "type": "string",
21
- "description": "OPTIONAL (v1.8) — ISO timestamp of compilation. The dispatch record's own time dimension: WorkResult carries none, and the journal's timings exist only on the workflow lane, so without this a dispatch on the prose lane is timeless."
22
- },
23
- "worker": { "$ref": "domain.schema.json#/$defs/WorkerName" },
24
- "mode": { "type": "string", "enum": ["orchestrated", "standalone"] },
25
- "operation": { "$ref": "domain.schema.json#/$defs/Operation" },
26
- "interaction": { "$ref": "domain.schema.json#/$defs/Interaction" },
27
- "substrate": { "$ref": "domain.schema.json#/$defs/Substrate" },
28
- "payload": { "$ref": "domain.schema.json#/$defs/WorkOrderPayload" }
36
+ "description": "OPTIONAL (v1.8) \u2014 ISO timestamp of compilation. The dispatch record's own time dimension: WorkResult carries none, and the journal's timings exist only on the workflow lane, so without this a dispatch on the prose lane is timeless."
37
+ },
38
+ "worker": {
39
+ "$ref": "domain.schema.json#/$defs/WorkerName"
40
+ },
41
+ "mode": {
42
+ "type": "string",
43
+ "enum": [
44
+ "orchestrated",
45
+ "standalone"
46
+ ]
47
+ },
48
+ "operation": {
49
+ "$ref": "domain.schema.json#/$defs/Operation"
50
+ },
51
+ "interaction": {
52
+ "$ref": "domain.schema.json#/$defs/Interaction"
53
+ },
54
+ "substrate": {
55
+ "$ref": "domain.schema.json#/$defs/Substrate"
56
+ },
57
+ "payload": {
58
+ "$ref": "domain.schema.json#/$defs/WorkOrderPayload"
59
+ }
29
60
  }
30
61
  }
@@ -34,8 +34,8 @@
34
34
  //
35
35
  // Exit code: 0 = overall green, 1 = overall red (mirrors the oracle convention), 2 = bad argv.
36
36
 
37
- import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync } from "node:fs";
38
- import { join, dirname } from "node:path";
37
+ import { readFileSync, writeFileSync, appendFileSync, mkdirSync, existsSync, readdirSync, statSync } from "node:fs";
38
+ import { join, dirname, basename, extname, resolve as resolvePath, relative } from "node:path";
39
39
  import { spawnSync } from "node:child_process";
40
40
  import { createHash } from "node:crypto";
41
41
  import { digest } from "../probe/digest.mjs";
@@ -389,6 +389,97 @@ export function nextTrialNo(dir, round, attempt) {
389
389
  return max + 1;
390
390
  }
391
391
 
392
+ // --- a check rewritten to pass ----------------------------------------------------------------
393
+ //
394
+ // A fixture that names Test Surface rows (`PASS TS-02-04`, `FAIL TS-02-04 step 10: …`) reads its
395
+ // checks from files the scope itself writes. Measured on a live run: a row FAILed on attempt 1 —
396
+ // the check renamed a list and found the old name still on screen — and went green on attempt 2
397
+ // because the check had been rewritten to stop renaming, while the code was unchanged. Nothing
398
+ // recorded that the evidence had moved. So each verdict records the checks it read, by digest, and
399
+ // the named results they printed; a row that failed in this scope's previous trial and passes now,
400
+ // whose own check file changed in between, is a REVISED check — a pass the judge must read against
401
+ // its row before it can count. Recorded, never refused: rewriting a check that overreached is
402
+ // legitimate, and only a person or the judge can tell which it was.
403
+
404
+ /**
405
+ * The files a fixture's arguments name, with their digests — a named directory contributes its
406
+ * direct children. Paths resolve against the project root and are recorded relative to it.
407
+ *
408
+ * @param {string[]} commands - The fixture command lines.
409
+ * @param {string} cwd - Project root.
410
+ * @returns {Object<string,string>} Relative path → sha256 of its bytes. Empty when no argument names
411
+ * an existing path.
412
+ */
413
+ export function checkFiles(commands, cwd) {
414
+ const out = {};
415
+ /**
416
+ * Record one file's digest under its project-relative path.
417
+ * @param {string} abs - Absolute path of the file.
418
+ * @returns {void} Nothing; an unreadable file is skipped, since it is not evidence.
419
+ */
420
+ const add = (abs) => {
421
+ try { out[relative(cwd, abs)] = createHash("sha256").update(readFileSync(abs)).digest("hex"); } catch { /* unreadable — not evidence */ }
422
+ };
423
+ for (const cmd of commands || []) {
424
+ const tokens = String(cmd).match(/"[^"]*"|'[^']*'|\S+/g) || [];
425
+ for (const raw of tokens.slice(1)) {
426
+ const tok = raw.replace(/^["']|["']$/g, "");
427
+ if (!tok || tok.startsWith("-")) continue;
428
+ const abs = resolvePath(cwd, tok);
429
+ let st; try { st = statSync(abs); } catch { continue; }
430
+ if (st.isFile()) add(abs);
431
+ else if (st.isDirectory()) {
432
+ for (const f of readdirSync(abs).sort()) {
433
+ const child = join(abs, f);
434
+ try { if (statSync(child).isFile()) add(child); } catch { /* skip */ }
435
+ }
436
+ }
437
+ }
438
+ }
439
+ return out;
440
+ }
441
+
442
+ /**
443
+ * The rows a fixture's output names, and how each came out.
444
+ *
445
+ * @param {Array<{stdout?:string}>} results - Fixture results with their full stdout.
446
+ * @returns {Object<string,("PASS"|"FAIL")>} Row id → its result; a FAIL anywhere wins over a PASS.
447
+ */
448
+ export function namedResults(results) {
449
+ const out = {};
450
+ for (const r of results || []) {
451
+ for (const line of String(r?.stdout || "").split(/\r?\n/)) {
452
+ const m = line.match(/^(PASS|FAIL)\s+(\S+)/);
453
+ if (!m) continue;
454
+ if (m[1] === "FAIL" || out[m[2]] !== "FAIL") out[m[2]] = m[1];
455
+ }
456
+ }
457
+ return out;
458
+ }
459
+
460
+ /**
461
+ * Rows that failed in the previous trial and pass now while their own check file changed.
462
+ *
463
+ * A row's check is the file whose name, without extension, is the row id — the convention a
464
+ * per-row check follows. A row with no such file is not reported: nothing can be said about it.
465
+ *
466
+ * @param {({check_files?:Object<string,string>, named_results?:Object<string,string>}|null)} prev
467
+ * The scope's previous verdict in this round, or null.
468
+ * @param {{check_files:Object<string,string>, named_results:Object<string,string>}} curr - This one.
469
+ * @returns {Array<{id:string, file:string}>} One entry per revised check.
470
+ */
471
+ export function revisedChecks(prev, curr) {
472
+ if (!prev?.named_results || !prev?.check_files) return [];
473
+ const out = [];
474
+ for (const [id, now] of Object.entries(curr.named_results || {})) {
475
+ if (now !== "PASS" || prev.named_results[id] !== "FAIL") continue;
476
+ const file = Object.keys(curr.check_files || {}).find((f) => basename(f, extname(f)) === id);
477
+ if (!file) continue;
478
+ if (prev.check_files[file] !== curr.check_files[file]) out.push({ id, file });
479
+ }
480
+ return out;
481
+ }
482
+
392
483
  /**
393
484
  * Distill every failing command's output into AEGIS {file,line,core_message} triples.
394
485
  * @param {{fixtures:{results:Array<{pass:boolean,stdout:string,stderr:string}>},
@@ -514,6 +605,10 @@ export async function cli(rawArgv) {
514
605
  const dbProbe = runDbProbe(contract.db_probe, cwd);
515
606
  const verdict = computeVerdict({ fixtures, dbProbe });
516
607
  const discovered = verdict.overall === "red" ? digestFailures({ fixtures, dbProbe }) : [];
608
+ const checks = {
609
+ check_files: checkFiles(fixtures.results.map((r) => r.cmd), cwd),
610
+ named_results: namedResults(fixtures.results),
611
+ };
517
612
 
518
613
  // ---- the ratchet ---------------------------------------------------------------------
519
614
  // `current` is the incumbent: the score of the most recent trial whose TREE is the one on disk
@@ -526,6 +621,11 @@ export async function cli(rawArgv) {
526
621
  const verdictBetter = better(s, baseline ? baseline.score : null);
527
622
  const crashed = fixtures.results.some((r) => r.error) || !!dbProbe?.error;
528
623
  const { status, action } = decideStatus(verdictBetter, crashed);
624
+ // The scope's previous trial in this round, read back for the checks it recorded.
625
+ const prevTrial = [...priorTrials].reverse().find((tr) => tr.round === round);
626
+ let prevVerdict = null;
627
+ if (prevTrial?.artifact) { try { prevVerdict = JSON.parse(readFileSync(join(outDir, prevTrial.artifact), "utf8")); } catch { /* none */ } }
628
+ const revised = revisedChecks(prevVerdict, checks);
529
629
 
530
630
  // The run key, read from the receipt that lives in the run root this script was pointed at.
531
631
  // `--out` IS that root, so identity comes from the receipt rather than from parsing a slug back
@@ -550,6 +650,9 @@ export async function cli(rawArgv) {
550
650
  ...verdict,
551
651
  score: s,
552
652
  discovered_tasks: discovered,
653
+ ...(Object.keys(checks.check_files).length ? { check_files: checks.check_files } : {}),
654
+ ...(Object.keys(checks.named_results).length ? { named_results: checks.named_results } : {}),
655
+ ...(revised.length ? { revised_checks: revised } : {}),
553
656
  });
554
657
 
555
658
  // The tree operation. `--no-ratchet` leaves the working tree exactly as the attempt left it —
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.12.0",
3
+ "version": "3.14.0",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -20,6 +20,8 @@ same artifacts out.
20
20
 
21
21
  ## Input contract — the WorkOrder
22
22
 
23
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
24
+
23
25
  Invoked as `--order <path>`. Fields you may rely on (absent = unknown; surface it, never guess):
24
26
 
25
27
  | Field | What it is |
@@ -69,10 +71,10 @@ its phase; templates live in `assets/templates/`.
69
71
  6 TASKS atomic, ordered, executable → tasks/ (LOCAL root; the one uncommitted branch
70
72
  of the tree — regenerable, machine-local) [references/task-generation.md]
71
73
  7 DERIVE+LINT mechanical, not yours to grade:
72
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" reduce board --slug <slug> --write
74
+ node "<kernel>" reduce board --slug <slug> --write
73
75
  (unlocks = depends_on inverse; Σ hours; critical path; appetite arithmetic —
74
76
  overflow is a fact you REPORT for the caller's HAMMER gate, never resolve)
75
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify spec --slug <slug>
77
+ node "<kernel>" verify spec --slug <slug>
76
78
  (structure, wikilinks, edge symmetry — fix reds, then re-run; you never
77
79
  self-grade with a hand-walked checklist. BREADBOARD-PLACE / BREADBOARD-UI:
78
80
  add the screen or defer the Place; never fold it into another screen)
@@ -187,7 +189,7 @@ status flips for built work (ingest's job), scope contracts (scope-architect's),
187
189
  /ba-pitch-analyzer --order .shapeup/checkout-vnpay/orders/analyze.json
188
190
 
189
191
  # Standalone — the preamble shim compiles the order (mode: standalone, pause_gates: true):
190
- # node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" compile --operation analyze --slug <slug> \
192
+ # node "<kernel>" compile --operation analyze --slug <slug> \
191
193
  # --worker ba-pitch-analyzer --payload '{"pitch": "docs/pitch.md", "lens": "standard"}'
192
194
  /ba-pitch-analyzer docs/pitch.md # operation: analyze, lens judged
193
195
  /ba-pitch-analyzer --lens standard docs/pitch.md # lens pinned
@@ -17,6 +17,8 @@ the ship report's census table.
17
17
 
18
18
  ## Input contract — the WorkOrder
19
19
 
20
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
21
+
20
22
  | Field | What it is |
21
23
  |---|---|
22
24
  | `operation` | `map-scopes` — the only operation this skill has. It covers first slicing after the board exists, folding discovered items in, and re-slicing a stuck scope; the payload says which of those you are doing |
@@ -95,7 +97,7 @@ the ship report's census table.
95
97
  hill_phase: "UPHILL_UNKNOWN" — ALWAYS; phase is derived from
96
98
  T0/T1 facts later,
97
99
  never authored
98
- 4 LINT node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify spec --slug <slug>
100
+ 4 LINT node "<kernel>" verify spec --slug <slug>
99
101
  → PA1 (directory alignment), PA2 (>~15 files), DISJOINT (undeclared overlap),
100
102
  SCOPE-ANCHOR (empty/unresolvable use_cases), TIER-DIRECTION (a task id in a
101
103
  committed contract), SCOPE-DEPS (depends_on naming a scope that isn't here).
@@ -52,13 +52,15 @@ INPUT: run's finished/unfinished scopes + baseline + census sources
52
52
 
53
53
  ## GATE H0 — Census
54
54
 
55
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
56
+
55
57
  **Purpose:** Gather every open item into one list before judging any of them. Never judge
56
58
  piecemeal — a partial view produces a wrong cut.
57
59
 
58
60
  ```
59
61
  H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns X" or "X is
60
62
  scope Y's", run
61
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug <slug> [--path <p>]...
63
+ node "<kernel>" probe owner --slug <slug> [--path <p>]...
62
64
  and cite its row. With no --path it answers for every engine and entry call site the wiring
63
65
  map names plus the profile's entry point; `writers: []` is an unowned seam and `missing`
64
66
  lists seams the wiring names that are not on disk — owned but never written. A census that
@@ -70,7 +72,7 @@ H0.1 Unresolved scopes (breaker cases only):
70
72
  - scopes with hammer_proposals (attempt_budget exhausted) → CARRY candidates. Exhaustion
71
73
  is DERIVED, never read off `t0/verdicts/*.json` directly — a compiled order or a T0
72
74
  verdict is writable by the very scope being judged and proves nothing on its own. Run
73
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe attempts --slug <slug> \
75
+ node "<kernel>" probe attempts --slug <slug> \
74
76
  --scope <scope-id> --round <n> --attempt-budget <n>
75
77
  and cite its `spent`/`tripped` fields (exit 1 = tripped) — an attempt counts only when a
76
78
  dispatch receipt AND either a leg-completion row or a WorkResult attest it, so a leg still
@@ -81,7 +83,7 @@ H0.2 QA findings (qa-edge-hunter's hunt-report.md, when present) — all `~` by
81
83
  H0.3 Discovered-task ledger entries still open (discovery/ledger.md, `[+]`/`~` unresolved).
82
84
  H0.4 Attempt-budget hammer proposals (scopes that exhausted their T0 attempts during BUILD).
83
85
  H0.4b Requirements with no PASS evidence — the pitch clauses the run never showed working. Run
84
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe requirements --slug <slug> --format table
86
+ node "<kernel>" probe requirements --slug <slug> --format table
85
87
  and take its `no evidence` rows; cite the row, the same way H0.0 cites ownership. Each is a
86
88
  census item carrying its source clause (`REQ-12 ← shaping.md R12`). A `cut` row is an answer
87
89
  the PO already gave — not an item. An inconsistency row (a criterion anchored to a
@@ -200,8 +202,8 @@ with a green census could only be recorded as `ask`. Still a proposal: the file
200
202
  /scope-hammer --slug checkout-vnpay --unattended
201
203
 
202
204
  # The ownership query every census claim cites (H0.0)
203
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --format table
204
- node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
205
+ node "<kernel>" probe owner --slug checkout-vnpay --format table
206
+ node "<kernel>" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
205
207
  ```
206
208
 
207
209
  ### Flags
@@ -38,6 +38,8 @@ return as a WorkResult.
38
38
 
39
39
  ## Input contract — the WorkOrder
40
40
 
41
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
42
+
41
43
  | Field | What it is |
42
44
  |---|---|
43
45
  | `operation` | `wire` (author/refresh the wiring map after `analyze`, before `map-scopes`) |
@@ -152,5 +154,5 @@ seam, or an engine with no attachment path, and why). You never touch spec docs,
152
154
  # The reachability oracle is the ORCHESTRATOR's, run advisory at L1b — not part of your craft.
153
155
  # Standalone, you MAY preview it after writing the map (it self-skips arms whose artifacts are
154
156
  # absent, and is near-vacuous pre-build since the engine code does not exist yet):
155
- # node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify trace --slug checkout-vnpay
157
+ # node "<kernel>" verify trace --slug checkout-vnpay
156
158
  ```
@@ -29,6 +29,8 @@ criterion with no collected evidence is a **FAIL**, never a pass-by-assumption.
29
29
 
30
30
  ## Input contract — the WorkOrder
31
31
 
32
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
33
+
32
34
  Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inferred):
33
35
 
34
36
  | Field | What it is |
@@ -39,6 +41,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
39
41
  | `payload.run_cmd` | How to start the running app. On a stack where building and launching are different acts it is only the build — prefer `payload.launch_cmd` when the order carries one. Absent standalone → ask; absent orchestrated with no `launch_cmd` either → ESCALATE, do not guess |
40
42
  | `payload.launch_cmd` | The project profile's launch probe: installs the built artifact, starts it and asserts the first screen. This is how the app is brought up for `[ui]` probing. A non-zero exit is a finding to cite, not a reason to try another way. Absent → the profile declares none; fall back to `run_cmd` |
41
43
  | `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
44
+ | `payload.revised_checks[]` | Rows that FAILed in one trial of this round and PASS in a later one whose own check file changed in between — a pass obtained by editing the check. Read each listed file against its row's Expect before you let its PASS count; a check that no longer asserts what the row asks makes that row a FAIL, and the bug names the check. Absent → no check was rewritten |
42
45
  | `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
43
46
  | `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
44
47
  | `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
@@ -85,9 +88,11 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
85
88
  - Contract work: send real requests, compare field-by-field.
86
89
  - **A fixture that names a row is evidence for that row.** When a T0 artifact you cite, or the
87
90
  build gate, carries output that names a Test Surface row by id — `PASS TS-05-05`, or
88
- `FAIL TS-05-05 step 4: …` — grade that row on it: a named PASS confirms, a named FAIL is a
89
- FAIL whose bug is that line. This is how a device row is evidenced when you cannot drive the
90
- app yourself; it is never a reason to skip driving it when you can.
91
+ `FAIL TS-05-05 step 4: …` — grade that row on it, after reading the check that printed it: a
92
+ named PASS confirms only when that check asserts the row's Expect; a check weaker than its row
93
+ (it opens the dialog, the row says pre-filled) is not evidence for the row, and grading it PASS
94
+ is the generator grading itself. A named FAIL is a FAIL whose bug is that line. This is how a
95
+ device row is evidenced when you cannot drive the app; never a reason not to when you can.
91
96
  - No evidence collected = recorded "NO EVIDENCE" → FAILs at verdict.
92
97
 
93
98
  **VERDICT.**
@@ -222,7 +227,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
222
227
  /spec-evaluator --order .shapeup/checkout-vnpay/orders/evaluate-r2.json
223
228
 
224
229
  # Standalone — the preamble shim compiles a minimal order, then the single code path runs:
225
- # node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" compile --operation evaluate --slug <slug> \
230
+ # node "<kernel>" compile --operation evaluate --slug <slug> \
226
231
  # --worker spec-evaluator [--payload '{"dimensions": [...], "run_cmd": "..."}']
227
232
  /spec-evaluator --spec shapeup/checkout-vnpay/spec/ --task TASK-007
228
233
  /spec-evaluator --spec shapeup/checkout-vnpay/spec/ --feature checkout-vnpay --single-pass
@@ -230,7 +235,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
230
235
 
231
236
  Standalone keeps `--task` (per-task check, not round-gated) and `--single-pass` (feature-level)
232
237
  — the shim maps them onto the order's payload; missing run command → ask. After writing the
233
- WorkResult, run `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" reduce ingest <result path>` and show its
238
+ WorkResult, run `node "<kernel>" reduce ingest <result path>` and show its
234
239
  summary — standalone has no orchestrator to ingest for you.
235
240
 
236
241
  ---
@@ -16,6 +16,8 @@ it does not exist for you.
16
16
 
17
17
  ## Input contract — the WorkOrder
18
18
 
19
+ **Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
20
+
19
21
  You are invoked as `--order <path>` pointing at a schema-valid WorkOrder. Fields you may
20
22
  rely on (anything absent = **unknown**; never invent it):
21
23
 
@@ -194,8 +196,8 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
194
196
 
195
197
  # Standalone — the preamble shim compiles a minimal WorkOrder from the flags, then the
196
198
  # single code path above runs. Requires the harness scripts (plugin install):
197
- # node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" compile --task TASK-003 --slug checkout-vnpay
198
- # node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" compile --next --slug checkout-vnpay
199
+ # node "<kernel>" compile --task TASK-003 --slug checkout-vnpay
200
+ # node "<kernel>" compile --next --slug checkout-vnpay
199
201
  /task-executor --spec shapeup/checkout-vnpay/spec/ --task TASK-003
200
202
  /task-executor --spec shapeup/checkout-vnpay/spec/ --next
201
203
  ```
@@ -203,6 +205,6 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
203
205
  Standalone shim: derive `<slug>` from the `--spec` path (`shapeup/<slug>/spec`),
204
206
  run `harness compile` with the matching flags (mode becomes `standalone`), then proceed
205
207
  against the compiled order exactly as if dispatched. After writing the WorkResult, run
206
- `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" reduce ingest <result path>` yourself and show the user its
208
+ `node "<kernel>" reduce ingest <result path>` yourself and show the user its
207
209
  summary — standalone has no orchestrator to ingest for you. One code path inside; two entry
208
210
  points outside.
@@ -898,7 +898,7 @@ async function crossGate(gateId, phaseName, validDecisions, ctx) {
898
898
  // The gate ledger keys a per-round crossing (L2, L3) on gate id + round, so it needs the round
899
899
  // whenever the caller already has one to show in the block — the same value `ctx.round` carries
900
900
  // for display, threaded through rather than re-derived.
901
- const roundFlag = ctx?.round != null ? ` --round ${ctx.round}` : "";
901
+ const roundFlag = (ctx?.round != null ? ` --round ${ctx.round}` : "") + (ctx?.verdict ? ` --verdict ${ctx.verdict}` : "");
902
902
  const g = await cmd(`gate --resolve ${gateId} --slug ${slug}${roundFlag} ${answersFlag(args.answers)}`.trim(), phaseName, `gate:${gateId}`);
903
903
  if (g.exit_code === 4) return { stop: paused(gateId, validDecisions, ctx) };
904
904
  if (g.exit_code === 5) return { stop: aborted(gateId, g.detail || `GATE ${gateId} aborted`) };