shapeup-sdlc 3.17.0 → 3.17.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "shapeup-sdlc-plugin",
3
3
  "displayName": "ShapeUp SDLC Plugin",
4
- "version": "3.17.0",
4
+ "version": "3.17.2",
5
5
  "description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
6
6
  "author": {
7
7
  "name": "Liberty Nguyen",
package/AGENTS.md CHANGED
@@ -36,7 +36,7 @@ Betting Table: PO decides; rejected pitches loop back to raw idea.
36
36
  | Wire | ⏸ **L1a.5** — Wiring Review ✚ | `/solution-architect` (`wire`): sole writer of committed `wiring-map.md` — per-UC engine → seam → entry-point call site → affordance, per `project-profile.md` |
37
37
  | Map Scopes | ⏸ **L1b** — Board Review (+ substrate disjointness lint) | `/scope-architect` (scope contracts ✦ — sole writer); traceability oracle advisory ✚. A registered requirement that no acceptance criterion grades and no scope claims is **red** here, and L1b prints the `REQ → AC` table: the two ways out are an AC carrying `(covers: REQ-…)` or the PO marking the clause `CUT (PO-approved)`. Red only where the plan is still cheap to change — after L1b nobody re-reads the pitch |
38
38
  | Build Vertically | ⏸ **L2** — Board 100% ✅ + T0-green ✦ | per dispatch: compile order → `/task-executor` (--order) → ingest result; T0-verified per attempt (fixtures + DB probe ✦), substrate-sandboxed ✦. Scopes build **concurrently** ✦ — `--parallel-scopes N` caps it (default 4), a scope is released the moment its own dependencies are green, and a scope green in this round is skipped rather than rebuilt. Then the **round build gate** ⚙: the ledger's run command, then the profile's `build_probe` and `launch_probe`, run once per round before EVAL — a red gate ends the round with no verdict and its failing step is compiled into the next round's orders as bugs; a `mobile` profile with no `launch_probe` is warned about every round, so the install/launch risk has an owner |
39
- | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason |
39
+ | EVAL (once per round) | ⏸ **L3** — Verdict | `/spec-evaluator` (--order), only over a round whose build gate ⚙ is not red: spec- + test-surface-conformance ★, T0 citation ✦; refuted boxes/verdict applied by ingest. The order names the round's build gate artifact and the profile's `launch_probe`, so a `[ui]` row is graded on the app the gate launched rather than recorded as no evidence — a project with no `launch_probe` still gets no on-device grade. A PASS names every Test Surface row as its own criterion: one that grades the surface as a group is refused, and the judge is sent back once with the reason. A scope that did not go green in the round has its rows graded FAIL, and the order names such scopes, so one red scope never makes the whole round ungradeable |
40
40
  | FAIL → round r+1 | — | regression rule ★: bugs + full Test Surface of touched UC. Every criterion the verdict graded FAIL reaches the round, whether or not the judge filed a bug for it, addressed to the scope that owns its use case — a round with a FAIL verdict and nothing to fix is how a run stalls. A row whose check the scope rewrote between a failing trial and a passing one is listed for the judge, who reads it against the row before the pass counts |
41
41
 
42
42
  ✦ = requires scope contracts (`shapeup/<slug>/scopes/*.md`); ✚ = requires the spine artifacts (`requirements.md`, `wiring-map.md`, `project-profile.md`). Traceability stays advisory until `covers:` is populated. Absent artifact ⇒ arm skipped (non-regression).
@@ -62,7 +62,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
62
62
  - **Ledger = single source of truth** — every discovery flow writes only its own section.
63
63
  - **QA is a level-up, not a gate** — `--no-qa` skips it; circuit breaker outranks the Hunter.
64
64
  - **Role separation** — Evaluator grades, task-executor fixes, QA discovers.
65
- - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
65
+ - **The requirements matrix is a projection, never a verdict** — `REQ → AC → criterion → verdict` is derived from files for one named run (the registry, the board's `covers:` clauses, the run's verdict rows, the T0 citations), never narrated and never passed in. `covers:` is the authoritative join: a criterion anchored to a requirement no AC covers is printed for reconciliation and counted as nothing. A verdict that anchors no criterion at all, beside a board whose ACs carry `covers:`, is sent back to the judge once. L4 reads one line off it, GATE H's census takes the clauses with no evidence, `REPORT.md` freezes the table — and none of that blocks a ship. A clause with no PASS evidence is a fact the baseline comparison weighs, not a veto.
66
66
  - **Hill phase is mechanical ✦** — derived only from T0/T1 artifacts, never self-reported, and a T0-green from a round whose build gate ⚙ is red moves no dot (a green fixture in a round the feature did not build is evidence about the fixture); the evaluator cites a T0 artifact it re-hashes itself, from the list its order carries. A scoped verdict citing none is refused: its round stays open and is evaluated again, never advanced.
67
67
  - **Envelope port (v1.0)** — every dispatch is WorkOrder in / WorkResult out; shared state has exactly one writer (the ingest step); malformed envelopes are hook-denied. Workers: stateless, craft-only, pipeline-blind.
68
68
 
@@ -1267,6 +1267,11 @@ export async function cli(rawArgv) {
1267
1267
  if (operation === "evaluate" && payloadExtra.t0_artifacts === undefined) {
1268
1268
  const { artifacts, missing } = t0ArtifactsFor(cwd, slug, round);
1269
1269
  if (artifacts.length) payloadExtra.t0_artifacts = artifacts;
1270
+ // A scope that did not go green leaves its rows ungraded, not the round: the judge refused a whole
1271
+ // round over one such scope and the run aborted at L3, before the next round or GATE H's census
1272
+ // could act on it. With at least one citation the round is gradeable, and the order names the
1273
+ // scopes whose rows are FAILs for want of a green T0.
1274
+ if (artifacts.length && missing.length) payloadExtra.scopes_without_t0 = missing;
1270
1275
  // On stderr, never stdout: stdout is the order path the caller consumes.
1271
1276
  if (missing.length) {
1272
1277
  console.error(`compile-order: warning — no green T0 verdict${round ? ` in round ${round}` : ""} for ` +
@@ -36,6 +36,8 @@ import { join, resolve, resolve as resolvePath, sep } from "node:path";
36
36
  import { createHash } from "node:crypto";
37
37
  import { runArgs } from "../lib/argv.mjs";
38
38
  import { resultsDir, scopesDir, readRunId, verdictsDir, specDir } from "../lib/paths.mjs";
39
+ import { readBoard } from "../compile.mjs";
40
+ import { coveringAcs } from "./requirements.mjs";
39
41
 
40
42
  /** Longest `reason` reported. A deviation is prose written by a worker and can run to paragraphs. */
41
43
  const REASON_MAX = 400;
@@ -253,6 +255,33 @@ export function coverageProblem(cwd, slug, verdict) {
253
255
  `each row is its own criterion, its id in the criterion or its traces_to; ungraded: ${shown}`;
254
256
  }
255
257
 
258
+ /**
259
+ * Whether a verdict anchors any criterion to a requirement when the board says which ones it could.
260
+ *
261
+ * `traces_to` is copied from the `(covers: REQ-…)` clauses of the acceptance criteria a criterion
262
+ * grades, and the requirements matrix is joined along it. Two judges on the same spec and the same
263
+ * instruction differed: one anchored every row and the matrix read 5/5, the next anchored none and it
264
+ * read 0/5 over the same passes. Held only to the all-empty case: which requirement a criterion
265
+ * traces to is the judge's reading, but none at all, beside a board that covers some, is a copy
266
+ * step skipped.
267
+ *
268
+ * @param {string} cwd - Project root.
269
+ * @param {string} slug - Feature slug.
270
+ * @param {object} verdict - The WorkResult's `verdict`.
271
+ * @returns {(string|null)} A reason phrased for the judge, or null.
272
+ */
273
+ export function tracesProblem(cwd, slug, verdict) {
274
+ if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
275
+ const criteria = Array.isArray(verdict.criteria) ? verdict.criteria : [];
276
+ if (!criteria.length) return null;
277
+ if (criteria.some((c) => Array.isArray(c?.traces_to) && c.traces_to.length)) return null;
278
+ let covered = 0;
279
+ try { covered = coveringAcs(readBoard(cwd, slug)).size; } catch { return null; }
280
+ if (!covered) return null;
281
+ return `the verdict anchors none of its ${criteria.length} criteria to a requirement while the board's acceptance ` +
282
+ `criteria cover ${covered} — copy each graded AC's (covers: REQ-…) clause into that criterion's traces_to`;
283
+ }
284
+
256
285
  export function citationProblem(cwd, slug, verdict, { round = null } = {}) {
257
286
  if (verdict?.overall !== "PASS" && verdict?.overall !== "FAIL") return null;
258
287
  if (!isScoped(cwd, slug)) return null;
@@ -297,7 +326,7 @@ export function evalVerdict(cwd, slug, round) {
297
326
  ? `the evaluator returned ${status || "no status"}: ${first}`
298
327
  : `status ${status || "unknown"} with no PASS/FAIL verdict`), status);
299
328
  }
300
- const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) || citationProblem(cwd, slug, v, { round });
329
+ const problem = verdictProblem(v) || coverageProblem(cwd, slug, v) || tracesProblem(cwd, slug, v) || citationProblem(cwd, slug, v, { round });
301
330
  if (problem) return unfit(problem, status, overall);
302
331
  return {
303
332
  found: true,
@@ -37,7 +37,7 @@ import { fileURLToPath } from "node:url";
37
37
  import { validate } from "../verify/envelope.mjs";
38
38
  import { runArgs } from "../lib/argv.mjs";
39
39
  import { tasksDir, localRoot, dispatchReceipts, legLedger, readRunId, qaDir } from "../lib/paths.mjs";
40
- import { citationProblem, coverageProblem, verdictProblem } from "../probe/eval.mjs";
40
+ import { citationProblem, coverageProblem, tracesProblem, verdictProblem } from "../probe/eval.mjs";
41
41
  import { huntProblem } from "../probe/hunt.mjs";
42
42
 
43
43
  const HERE = dirname(fileURLToPath(import.meta.url));
@@ -702,7 +702,7 @@ export async function cli(rawArgv) {
702
702
  const evalSlug = String(result.order_id).split("/")[0];
703
703
  const evalRound = Number((String(result.order_id).match(/-r(\d+)$/) || [])[1]) || null;
704
704
  const problem = verdictProblem(result.verdict) || coverageProblem(cwd, evalSlug, result.verdict)
705
- || citationProblem(cwd, evalSlug, result.verdict, { round: evalRound });
705
+ || tracesProblem(cwd, evalSlug, result.verdict) || citationProblem(cwd, evalSlug, result.verdict, { round: evalRound });
706
706
  if (problem) {
707
707
  console.error(`ingest-result: result refused — ${problem}.`);
708
708
  console.error(` The round stays open: re-dispatch the evaluator against its order, which lists`);
@@ -363,6 +363,7 @@
363
363
  "launch_cmd",
364
364
  "build_gate",
365
365
  "revised_checks",
366
+ "scopes_without_t0",
366
367
  "t0_artifacts",
367
368
  "browser",
368
369
  "tasks"
@@ -2445,6 +2446,11 @@
2445
2446
  },
2446
2447
  "description": "spec-evaluator: rows that FAILed in one of the round's T0 trials and PASS in a later one whose own check file (named by the row id) changed in between. Derived by harness compile. Each must be read against its row before its PASS counts; absent when none."
2447
2448
  },
2449
+ "scopes_without_t0": {
2450
+ "type": "array",
2451
+ "items": { "type": "string" },
2452
+ "description": "spec-evaluator: scopes whose contract exists but which hold no green T0 verdict for the round, while at least one other scope does. Derived by harness compile. The round is still gradeable: each row such a scope owns is a FAIL citing the scope contract, never a reason to refuse the round. Absent when every scope is green."
2453
+ },
2448
2454
  "build_gate": {
2449
2455
  "type": "string",
2450
2456
  "description": "spec-evaluator, qa-edge-hunter: this run's newest round build gate artifact (build/r<N>-t<T>.json) — each step's exit code and output tail, the record that the build ran and the app launched. Derived by `harness compile`, absent when the gate never ran. Read it before grading a [ui] row NO EVIDENCE."
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "shapeup-sdlc",
3
- "version": "3.17.0",
3
+ "version": "3.17.2",
4
4
  "description": "Shape Up for coding agents \u2014 with gates the agent can't talk its way past. Harness for Claude Code.",
5
5
  "bin": {
6
6
  "shapeup-sdlc": "bin/init.mjs"
@@ -43,6 +43,7 @@ Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inf
43
43
  | `payload.build_gate` | This run's newest round build gate artifact — each step's exit code and output tail. It records that the build ran and whether the app launched, and the launch step's output names what it captured. Read it before grading any `[ui]` row NO EVIDENCE: a launch that succeeded is evidence the app can be probed, and the row is graded on the app, not on its absence. Absent → the gate never ran |
44
44
  | `payload.revised_checks[]` | Rows that FAILed in one trial of this round and PASS in a later one whose own check file changed in between — a pass obtained by editing the check. Read each listed file against its row's Expect before you let its PASS count; a check that no longer asserts what the row asks makes that row a FAIL, and the bug names the check. Absent → no check was rewritten |
45
45
  | `payload.t0_artifacts[]` | Per-scope T0 verdict paths for this round (scoped specs), compiled from each scope's green verdict. An artifact listed but missing/red on disk, or a scoped spec with none listed → the round is NOT gradeable: return `status: failed` with the reason, naming the scope, as your FIRST deviation — a structural precondition, not a criterion |
46
+ | `payload.scopes_without_t0[]` | Scopes with no green T0 this round while others have one. The round IS gradeable: grade every row such a scope owns as FAIL, evidence "no green T0 for scope <id>" with the scope contract as its locator (`…/scopes/<id>.md:1`). Never refuse the round for them — a refusal ends the run before the next round or GATE H can act. Absent → every scope is green |
46
47
  | `payload.browser` | `cli` (default, ~4x cheaper) \| `mcp` \| `none` |
47
48
  | `payload.tasks[]` | Traceability only (which UCs a task claims): NEVER a grading source — the committed UC text is the criterion, a paraphrase mismatch is a finding |
48
49
  | `substrate.allowed` | Your only write surface: `.shapeup/<slug>/evaluation/**` (the report + evidence) |
@@ -187,7 +188,7 @@ clause yields no anchor — leave the array empty rather than guessing, and neve
187
188
  supply one. This changes nothing you grade: the anchor is a navigation path, never a grading input,
188
189
  and a criterion passes or fails on its evidence exactly as before. It matters downstream because
189
190
  the requirement matrix at GATE L4 and the census at GATE H are projected from these anchors; a
190
- verdict that drops them grades the build and says nothing about what the pitch asked for.
191
+ verdict that drops them grades the build and says nothing about what the pitch asked for. Ingest refuses a verdict that anchors no criterion at all while the board's ACs carry `covers:` clauses, and you are sent back once.
191
192
 
192
193
  **Every FAIL criterion's `evidence` MUST carry a `file:line` locator** — schema-enforced, not
193
194
  advice: the envelope is validated against `work-result.schema.json` at ingest and a locatorless