shapeup-sdlc 3.1.2 → 3.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/.claude-plugin/plugin.json +1 -1
  2. package/AGENTS.md +7 -6
  3. package/README.md +1 -1
  4. package/SECURITY.md +4 -1
  5. package/bin/init.mjs +3 -0
  6. package/commands/retro.md +19 -2
  7. package/commands/ship.md +5 -0
  8. package/hooks/dispatch-receipt.mjs +6 -3
  9. package/hooks/gate-intake.mjs +1 -1
  10. package/hooks/gate-zerowork.mjs +5 -2
  11. package/hooks/lib/decision.mjs +49 -6
  12. package/hooks/safety-spine.mjs +8 -5
  13. package/hooks/sandbox-guard.mjs +13 -6
  14. package/kernel/compile.mjs +112 -6
  15. package/kernel/harness.mjs +11 -5
  16. package/kernel/init/run.mjs +122 -6
  17. package/kernel/lib/breadboard.mjs +165 -0
  18. package/kernel/lib/paths.mjs +15 -3
  19. package/kernel/probe/owner.mjs +139 -0
  20. package/kernel/probe/resume.mjs +7 -1
  21. package/kernel/probe/stats.mjs +49 -2
  22. package/kernel/reduce/hill.mjs +12 -2
  23. package/kernel/verify/build.mjs +319 -0
  24. package/kernel/verify/spec.mjs +190 -2
  25. package/package.json +1 -1
  26. package/skills/ba-pitch-analyzer/SKILL.md +11 -5
  27. package/skills/ba-pitch-analyzer/assets/templates/_index.tmpl.md +2 -1
  28. package/skills/ba-pitch-analyzer/assets/templates/ux-behavior.tmpl.md +12 -2
  29. package/skills/ba-pitch-analyzer/references/doc-schemas.md +3 -0
  30. package/skills/ba-pitch-analyzer/references/ux-behavior-patterns.md +9 -0
  31. package/skills/coach/SKILL.md +232 -43
  32. package/skills/orient/SKILL.md +15 -5
  33. package/skills/qa-edge-hunter/SKILL.md +3 -2
  34. package/skills/scope-architect/SKILL.md +7 -0
  35. package/skills/scope-hammer/SKILL.md +13 -0
  36. package/skills/solution-architect/SKILL.md +7 -2
  37. package/skills/tech-lead/SKILL.md +7 -7
  38. package/skills/tech-lead/references/gates.md +69 -17
  39. package/skills/tech-lead/references/protocol.md +33 -8
  40. package/skills/tech-lead/schemas/domain.schema.json +180 -9
  41. package/skills/tech-lead/workflows/shapeup-run.js +54 -12
@@ -56,6 +56,14 @@ INPUT: run's finished/unfinished scopes + baseline + census sources
56
56
  piecemeal — a partial view produces a wrong cut.
57
57
 
58
58
  ```
59
+ H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns X" or "X is
60
+ scope Y's", run
61
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug <slug> [--path <p>]...
62
+ and cite its row. With no --path it answers for every engine and entry call site the wiring
63
+ map names plus the profile's entry point; `writers: []` is an unowned seam and `missing`
64
+ lists seams the wiring names that are not on disk — owned but never written. A census that
65
+ narrated ownership from memory once told the PO no scope owned a screen directory that a
66
+ committed contract listed in plain sight — and pointed the ship decision at the wrong gap.
59
67
  H0.1 Unresolved scopes (breaker cases only):
60
68
  - uphill/downhill scopes when round_budget hit 0 → CARRY candidates (their own hill
61
69
  phase + open unknowns, from hill/<scope-id>.yml)
@@ -160,6 +168,10 @@ The WorkResult may carry only `files_touched`, `artifacts`, `assumptions`, `devi
160
168
 
161
169
  # Headless — no PO available; still refuses to auto-ship a ship-blocking item
162
170
  /scope-hammer --slug checkout-vnpay --unattended
171
+
172
+ # The ownership query every census claim cites (H0.0)
173
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --format table
174
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
163
175
  ```
164
176
 
165
177
  ### Flags
@@ -181,6 +193,7 @@ The WorkResult may carry only `files_touched`, `artifacts`, `assumptions`, `devi
181
193
  | Default classification is NICE-TO-HAVE unless traced to a pitch boundary or business_goal | A generous must-have list defeats the point of hammering |
182
194
  | A MUST-HAVE that fails H1.2 is never cut silently | The one case where scope-hammer refuses to make the run "look" shippable |
183
195
  | Cuts are proposals; the PO confirms every one | This skill never overrides the human at the ship gate |
196
+ | Every ownership claim in the census cites `harness probe owner` | Ownership is what the contracts' substrates say — the same election `harness compile` uses to address a bug — never what the report remembers |
184
197
  | Cut items are carried to the discovery ledger, never silently dropped | Cool-down must stay debt-free — an idea deferred is still recorded |
185
198
  | Never sets status: done, never deploys, never ships unilaterally | Judge/doer/advisor separation holds even at the very last gate |
186
199
  | An overridden ship-blocking item is logged explicitly in the ship report | "Shipped" must never quietly mean "shipped with a known must-have gap" |
@@ -42,6 +42,8 @@ return as a WorkResult.
42
42
  |---|---|
43
43
  | `operation` | `wire` (author/refresh the wiring map after `analyze`, before `map-scopes`) |
44
44
  | `payload.feature` / `payload.spec_folder` | Slug + committed spec — read `usecases/` for the UCs and the engine each one needs, `domain-model.md`/`synthesis.md` for the module surface |
45
+ | `payload.breadboard` | When present, name each UC's `affordance` by its U# and Place |
46
+ | `payload.kb_rules_path` | Team guidelines (read if the file exists) — the seams this codebase actually wires through, entry points that are not where the template says. Steering, never spec: the profile's `entry_point` still wins, and a guideline that disagrees with it is reported in `deviations`, not applied |
45
47
  | `payload.project_profile` | Path to the SHARED `project-profile.md`. Its `entry_point` is the composition root every engine must attach to — **archetype-specific** (a client-only game's `main.js` is not a web-service's `src/server.ts`). Read it; never guess the entry point |
46
48
  | `substrate.allowed` | `wiring-map.md` — your ONLY write surface (the spec core, scopes, and the profile are frozen) |
47
49
 
@@ -52,7 +54,8 @@ guessed `main.js` would make the later oracle certify nothing.
52
54
  ## Core process
53
55
 
54
56
  ```
55
- 1 READ the project profile → entry_point + archetype. Read every use case in usecases/.
57
+ 1 READ the team guidelines at payload.kb_rules_path if the file exists (absent = none),
58
+ then the project profile → entry_point + archetype. Read every use case in usecases/.
56
59
  For each UC, identify the engine module that carries its core logic (the file that
57
60
  WILL exist, named from the domain model / synthesis surface — not a guess at a folder).
58
61
  2 DESIGN for each UC, design the integration path from the entry_point inward:
@@ -69,7 +72,9 @@ guessed `main.js` would make the later oracle certify nothing.
69
72
  file:line is a build-time fact (the oracle proves reachability by
70
73
  the import graph, it does not parse this field)
71
74
  affordance the player-visible thing this UC exposes once wired (the human
72
- end of the chain — what a user can DO, not an internal call)
75
+ end of the chain — what a user can DO, not an internal call);
76
+ with a breadboard, name it by U# and Place —
77
+ `U1 Pay (P1) → P2 Payment Sheet`
73
78
  3 WRITE shapeup/<slug>/wiring-map.md (WiringMap): frontmatter for schema_version, feature
74
79
  and entry_point (echo of the profile), then entries[] as ONE MARKDOWN TABLE under a
75
80
  `## Wiring` heading — this exact shape, because it is the only one the reader parses:
@@ -23,7 +23,7 @@ turns fighting shell quoting):
23
23
  node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run \
24
24
  --slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> \
25
25
  --auto-level <interactive|auto|unattended> \
26
- [--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3]
26
+ [--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
27
27
  ```
28
28
 
29
29
  **After a compaction, or in a fresh session over an open run, re-derive before you act.** One
@@ -43,13 +43,13 @@ counts.
43
43
  plugin and need a one-time permission grant (`npx shapeup-sdlc init` writes it). Do not route
44
44
  around it, and do not silently hand-build the feature instead.
45
45
 
46
- **Language gate (delegated to `translator`, not this skill):** at GATE L0, before Step 2, dispatch
47
- an Agent (model: exec) that calls `Skill(shapeup-sdlc-plugin:translator) --check <intake>`.
48
- English → proceed as-is. Non-English → dispatch a second Agent (`--auto` under auto/unattended)
49
- and orchestrate against the produced `<name>.en.md`. The tech lead detects and sequences; it never
50
- translates itself.
46
+ **Language gate (delegated to `translator`, not this skill):** before Step 1 opens the run, dispatch
47
+ an Agent (model: exec) calling `Skill(shapeup-sdlc-plugin:translator) --check` on the pitch *and* its
48
+ breadboard. English → proceed as-is. Non-English → a second Agent translates both (`--auto` under
49
+ auto/unattended); Step 1 then names the `.en.md` files (a run already open on the original: re-open
50
+ it with `--force` — nothing is dispatched yet). The tech lead detects and sequences; never translates.
51
51
 
52
- **Step 2 — pin GATE L0, then launch.** Collect the L0.1–L0.9 config (spec folder, lens, stack,
52
+ **Step 2 — pin GATE L0, then launch.** Collect the L0.1–L0.10 config (spec folder, lens, stack,
53
53
  eval dims, max_rounds, the model/budget matrix — see `references/gates.md` GATE L0 for the full
54
54
  collect-list), write the SHARED `project-profile.md` yourself (`{schema_version:1, archetype,
55
55
  entry_point}` — `shapeup-run.js` has no filesystem of its own), emit the `⏸ GATE L0` block, then
@@ -31,13 +31,19 @@ escalates, writes nothing, and every relaunch re-dispatches it.
31
31
 
32
32
  ```
33
33
  Collect (explicit — never inferred):
34
- L0.1 Kicked-off pitch source: path to a shaping.md / pitch.md (already shaped + bet by PO).
34
+ L0.1 Kicked-off pitch source: `shaping.md` — and its `breadboard.md`, which the run finds
35
+ beside it (or in `shaping/`) or takes from `--breadboard`; a `pitch.md` may carry the
36
+ breadboard inline. Already shaped + bet by PO.
35
37
  Not a raw idea — shaping (1-4) / betting (5) / kick-off (6) are PO-personal, upstream.
36
- L0.1a Language gate: Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator) --check <intake>.
37
- English → use intake as-is.
38
- non-English → Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator) <intake>
39
- (--auto under auto/unattended), then use the produced <name>.en.md as
40
- the ORIENT/MAP-SCOPES input. Log in ledger.
38
+ L0.1a Language gate, BEFORE init run: Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator)
39
+ --check over the pitch AND its breadboard.
40
+ English → use both as-is.
41
+ non-English → Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator) <pitch> <breadboard>
42
+ (--auto under auto/unattended), then open the run on the produced
43
+ <name>.en.md — init run prefers a breadboard.en.md beside it, and warns
44
+ when it can find only the untranslated breadboard. The run's intake is
45
+ what every planning worker reads, so a translation made after the run
46
+ opened reaches nobody: re-open with --force naming the .en.md. Log in ledger.
41
47
  L0.1b Appetite: read the `appetite` field from the pitch's YAML frontmatter (set by /shapeup).
42
48
  Surface it in the gate output. Use it to:
43
49
  - Contextualise the scope at L1b (right-size cuts to the budget).
@@ -90,6 +96,21 @@ Collect (explicit — never inferred):
90
96
  same GATE H proposal. attempt_budget counts ATTEMPTS and cannot see that the last two
91
97
  produced nothing; this term can, and on a flailing scope it saves three of five
92
98
  attempts. Set per scope (`no_progress_k` on the contract) or per run in the payload.
99
+ L0.10 knowledge base (read, never obeyed): if `shapeup/knowledge-base/tech-lead.md` exists,
100
+ read it now. Its **Workflow guidance** may add a question, a check or a warning line to
101
+ any gate block below and may name the spike to insist on at L1a; its **Suggested run
102
+ config** lines are PROPOSALS for L0.2–L0.9 and the profile — confirm each with the PO
103
+ before pinning it, and record `(source: knowledge-base)` beside a value taken from there.
104
+ Nothing in that file answers a gate: the answer set resolves exactly as it would without
105
+ it, and a line that could only be honoured by skipping, reordering or relaxing a gate is
106
+ reported as a suspected harness defect, not applied. If `shapeup/knowledge-base/` has no
107
+ file at all, offer the optional seed once — `/retro --scan` (the coach's scan operation)
108
+ drafts guidelines from the project on disk, or `/retro --research <stack>` (its
109
+ research operation) from the platform's official documentation when the project has no
110
+ build file to scan; both put every rule through GATE COACH-1 — and continue whether or
111
+ not the PO takes it. A Suggested run config line with `web-research` provenance has
112
+ never run in this project: confirm it like any other, and expect the first round build
113
+ gate to be its first execution.
93
114
  ```
94
115
 
95
116
  **L0.9b — the launch record.** Every switch the operator typed becomes a `RunArgs` field, or it
@@ -134,9 +155,11 @@ Intake lang : [English | translated via /translator → <name>.en.md]
134
155
  Appetite : [~1 week | ~2 weeks | ~6 weeks | ⚠️ missing — scope uncapped]
135
156
  Spec folder : [path] (lens: [lite|standard])
136
157
  Eval dims : [spec-conformance] max_rounds: [N, appetite-informed] auto: [interactive|auto|unattended]
137
- Run commands : [web: ... | api: ... | mobile: ...]
158
+ Run commands : [web: ... | api: ... | mobile: ...] (run_cmd → the round build gate, every round before EVAL)
159
+ Build gate : build_probe [set | —] launch_probe [set | — ⚠ mobile: the install/launch risk has no owner]
138
160
  Model matrix : orch=[model] exec=[model] eval=[model] qa=[model] digester=[script|sonnet] (source: [flags|settings.local|settings.json|default])
139
161
  Budgets : round_budget=[N] (outer) attempt_budget=[N] (inner, per scope)
162
+ Knowledge : [tech-lead.md — N workflow rules, M suggested values (confirmed above) | none — `/retro --scan` or `/retro --research <stack>` seeds it (optional)]
140
163
  ```
141
164
  Do NOT start ORIENT until confirmed (interactive/auto). Under --unattended, proceed.
142
165
 
@@ -167,8 +190,9 @@ committing to a scope map. This is the first Hill read (area-level — slices do
167
190
  ```
168
191
  Read .shapeup/<slug>/orient/. Render the 🗻 Hill from hill-signal.md (see protocol.md "Hill report"):
169
192
  - each suspected area → uphill (open unknowns) | crest (approach proven by the spike) | downhill
170
- Print: the code-surface headline (where it lands), the spiked area + result, the riskiest
171
- open unknowns going into mapping.
193
+ Print: `Breadboard: <source> | none` (how init run found the pitch's breadboard — flag, sibling,
194
+ shaping-dir, shared-root, embedded — or none), the code-surface headline (where it lands),
195
+ the spiked area + result, the riskiest open unknowns going into mapping.
172
196
  Ask (max 2): is the riskiest area the right one to have spiked? any unknown that must be
173
197
  resolved (another spike) before we map scopes?
174
198
  ```
@@ -182,11 +206,17 @@ Do NOT enter MAP SCOPES until Orient is accepted.
182
206
 
183
207
  ```
184
208
  1. PROFILE (you write it at L0 — compile-order stays pipeline-blind): SHARED project-profile.md
185
- = {schema_version:1, archetype, entry_point}. archetype ∈ {client-only-game|web-service|
186
- mobile|library|data-pipeline}; entry_point is the reachability seam (a game's main.js is NOT a
187
- service's src/server.ts). Validate the enum — a typo must fail, not silently disable the check.
209
+ = {schema_version:1, archetype, entry_point, build_probe?, launch_probe?}. archetype ∈
210
+ {client-only-game|web-service|mobile|library|data-pipeline}; entry_point is the reachability
211
+ seam (a game's main.js is NOT a service's src/server.ts). Validate the enum — a typo must fail,
212
+ not silently disable the check. The two probes feed the round build gate (`harness verify
213
+ build`, every round before EVAL): build_probe asserts the BUILT ARTIFACT covers what the run
214
+ wrote (a green exit code is not proof the feature compiled when the toolchain compiles only what
215
+ an entry point reaches); launch_probe installs, starts and asserts the first screen. A `mobile`
216
+ profile without a launch_probe is warned about every round — nothing else in the loop launches
217
+ the app.
188
218
  2. WIRE — compile-order --operation wire --slug <slug> (worker→solution-architect), payload
189
- {project_profile}. Sole writer of committed wiring-map.md (per-UC engine → seam → entry-point
219
+ {project_profile, breadboard?}. Sole writer of committed wiring-map.md (per-UC engine → seam → entry-point
190
220
  call site → affordance). ⏸ GATE L1a.5: confirm each UC has a declared seam before slicing.
191
221
  ⟐ PRECONDITION: MAP SCOPES step 1 (ANALYZE) has already run and usecases/ is
192
222
  populated. WIRE writes one entry per use case, so dispatching it against an empty spec folder
@@ -212,7 +242,7 @@ against the seams WIRE declared. Sequence: ORIENT → L1a → **ANALYZE** → **
212
242
  ```
213
243
  Two orders, two workers, one step (both model: exec — see references/protocol.md):
214
244
  1. ANALYZE + BOARD — compile-order --operation analyze --slug <slug> --worker ba-pitch-analyzer
215
- --payload '{"pitch": "<path>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
245
+ --payload '{"pitch": "<path>", "breadboard": "<the path init run printed, when it printed one>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
216
246
  dispatch: Skill(shapeup-sdlc-plugin:ba-pitch-analyzer) --order <path>. The order hands it
217
247
  code-surface.md (Phase-1 ingest consumes the map, does not re-scan), discovered-seed.md
218
248
  (task gen starts from reality), spike-<area>.md (feasibility/contracts).
@@ -265,6 +295,8 @@ Scope contracts present:
265
295
  - scope board: scope_id, topology_type, substrate file count (scopes/*.md / scope-board.md)
266
296
  - any SPIKE blockers (scope-summary.md)
267
297
  - scope-summary "Done when" headline statements
298
+ - the Deferred Places from ux-behavior.md (breadboard Places this shape will not build) —
299
+ each one needs the PO's yes; a rejected deferral goes back to the planner as a screen
268
300
  No scope contracts (pre-v0.3.0, unchanged from v0.2.6):
269
301
  Read tasks/_index.md (LOCAL root). Print:
270
302
  - task count by package/variant (.shared / .be / .web / .mobile / .e2e)
@@ -280,7 +312,11 @@ this is the orchestrator's own re-confirmation before committing to a build sequ
280
312
  waiting to happen), PA1 (directory-aligned scope), PA2 (size cap), SCOPE-ANCHOR (a scope
281
313
  naming no committed use case, or one that does not resolve), TIER-DIRECTION (a committed
282
314
  contract naming LOCAL task ids), SCOPE-DEPS (a build-order id naming a scope that is not
283
- in this run). Any red → HARD STOP, past a 🔴 at the architect's own checkpoint.
315
+ in this run), BREADBOARD-PLACE (a breadboard Place with UI affordances has no
316
+ `## Screen: … (P#)` in ux-behavior.md and is not deferred), BREADBOARD-UI (a U# not
317
+ specified on a screen of its own Place). Any red → HARD STOP, past a 🔴 at the
318
+ architect's own checkpoint. The breadboard reds are the planner's to fix — add the screen
319
+ or defer the Place; never fold it into another screen.
284
320
  - Lock the build SEQUENCE riskiest-first: order scopes by open-unknowns count (from
285
321
  hill/<scope-id>.yml if present, else the orient hill signal), not by file count or
286
322
  alphabetical — Shape Up's "solve in the right sequence" (step 10).
@@ -315,8 +351,16 @@ Do NOT enter BUILD until the board is accepted.
315
351
  Feature : [slug]
316
352
  Round : [r]
317
353
  Scopes : [N] green, [M] queued for hammer
354
+ Build gate: [green | red — <failing step> | undeclared]
318
355
  ```
319
356
 
357
+ The build gate is `harness verify build --slug <slug> --round <r>` — the ledger's `run_cmd`, then
358
+ the profile's `build_probe` and `launch_probe`, stopping at the first failure. It runs before this
359
+ gate so the block shows it; `red` means EVAL is NOT dispatched this round (the judge grades a running
360
+ feature) and the failing step is compiled into round r+1's orders as `payload.bugs`. `undeclared`
361
+ means L0 pinned no run command and the profile names no probe — the round proceeds over an
362
+ unproven build, and the block says so.
363
+
320
364
  Under `--interactive` / `--auto`, the hook warns if the board is not truly green (advisory) and requires explicit PO approval to proceed. Under `--unattended`, it automatically aborts on a red board or proceeds on a green one.
321
365
 
322
366
  ---
@@ -344,6 +388,10 @@ PASS:
344
388
  → --no-qa or skill absent: proceed straight to SHIP; ledger records `qa: skipped`.
345
389
 
346
390
  FAIL:
391
+ → a round whose build gate was red never reached the judge: the block carries `build_gate: red`,
392
+ the "bug list" is the gate's failing step (its command, exit and output tail), and round r+1
393
+ is a fix round over exactly that. Nothing the evaluator would have said is missing — it was
394
+ never asked.
347
395
  → print the bug list grouped by task/severity. For each bug: task ID + failed Done-when criterion + repro.
348
396
  DO NOT prescribe fix options or root cause hypotheses — that is the implementer's job.
349
397
  The tech lead names scope; the implementer diagnoses and fixes.
@@ -378,6 +426,9 @@ S.0 GATE H — delegate to scope-hammer (this IS Shape Up's "Decide When to Sto
378
426
  (no --breaker flag) normal stop — all scopes FINISHED, post-QA-hunt
379
427
  Feeds it: qa/hunt-report.md findings (when present), discovery/ledger.md open items,
380
428
  the hammer-proposal queue from BUILD (attempt-budget exhaustions).
429
+ Ownership facts in its census — "no scope owns X", "X belongs to scope Y" — come from
430
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug <slug> [--path <p>]...
431
+ which elects the owner from the committed contracts' substrates, never from prose.
381
432
  Reads back: its GATE H0/H1/H2 output — census, baseline comparison, cut list + verdict.
382
433
  Authority: scope-hammer proposes; the tech lead records the PO's decision in
383
434
  round-ledger.md and performs the actual close (S.1 onward). It never ships on its own.
@@ -414,7 +465,8 @@ S.6 Harvest one signal row → append to `.shapeup/metrics/<machine-id>.jsonl`
414
465
  deliberately without colliding on one filename. The read plane is
415
466
  `harness probe stats`, or `cat .shapeup/metrics/*.jsonl`).
416
467
  Copy fields that ALREADY exist as structured output (run-state, final EVAL report,
417
- discovery ledger, qa/hunt-report, breadboard B5). Two hard rules:
468
+ discovery ledger, qa/hunt-report, and the receipt's `breadboard.ids.V` for `slice_count` —
469
+ the slices init run counted in the staged breadboard, never a hand copy). Two hard rules:
418
470
  1. Harvest only fields that already exist at ship time — never evaluate something new.
419
471
  2. Record facts, never compute a new verdict (no `run_quality_score` — that would be
420
472
  a second judge behind spec-evaluator). The eval suite interprets; harvest records.
@@ -465,7 +517,7 @@ Ledger : harness-run.md
465
517
  ```
466
518
  Question (max 1): "Anything to record before I close the run? (y/n) or provide feedback for the next sprint."
467
519
  On confirm:
468
- - If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable skills: `task-executor`, `ba-pitch-analyzer`, `qa-edge-hunter`; each reads its own file at the top of its next run. The tech lead does not categorize the feedback itself — that is the coach's gate, by design (no assumptions).
520
+ - If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable: `task-executor`, `ba-pitch-analyzer`, `qa-edge-hunter`, `orient`, `scope-architect`, `solution-architect` (each reads its own file at the top of its next run) and `tech-lead` (workflow guidance, read at the next GATE L0). Guidance never decides a gate: a filed rule may add a question or a check to a gate block, never an answer. The tech lead does not categorize the feedback itself — that is the coach's gate, by design (no assumptions).
469
521
  - Then output → `✅ [slug] [shipped & deployed | built & verified, deploy pending] — [r] rounds, verdict PASS.`
470
522
 
471
523
  ---
@@ -68,7 +68,9 @@ approval of the new tasks and estimates before resuming the BUILD loop.
68
68
 
69
69
  ## The EVAL timing rule (the core constraint)
70
70
 
71
- EVAL fires **once** per round and **only** when GATE L2 has confirmed the board is 100% done.
71
+ EVAL fires **once** per round and **only** when GATE L2 has confirmed the board is 100% done
72
+ and the round build gate (§3d) is not red — a feature that does not build or launch has nothing
73
+ for a judge to grade, and the gate's failing step is what round r+1 fixes.
72
74
  It is never:
73
75
  - called per task,
74
76
  - called inside the BUILD loop,
@@ -311,14 +313,15 @@ workers keep only their own product-idempotency key and emit domain artifacts.
311
313
 
312
314
  ## 0. LANGUAGE GATE → translator (GATE L0, only if non-English)
313
315
  ```
314
- Invoke via Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) --check "<intake path>"
316
+ Invoke via Agent (model: exec), BEFORE init run: Skill(shapeup-sdlc-plugin:translator) --check "<intake path>" "<breadboard path>"
315
317
  # detect-only, writes nothing
316
318
  English → skip; ORIENT against the original.
317
- non-English → Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) "<intake path>" [--auto]
318
- # full pass
319
- Writes: <name>.en.md (English copy; original untouched) + glossary.md
319
+ non-English → Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) "<intake path>" "<breadboard path>" [--auto]
320
+ # full pass — the pitch AND its breadboard, before init run
321
+ Writes: <name>.en.md per file (English copy; original untouched) + glossary.md
320
322
  + translation-report.md.
321
- ORIENT against the <name>.en.md copy.
323
+ Open the run on the .en.md pitch: init run stages breadboard.en.md beside it,
324
+ and every planning worker reads what init run pinned — never a later copy.
322
325
  Read back: the detect table (--check) / the .en.md path + residual scan result (full pass).
323
326
  Authority: translator normalizes language only — it does not orient/plan/build/judge. The tech
324
327
  lead never translates itself; it only detects and sequences this step before ORIENT.
@@ -340,7 +343,7 @@ Authority: pure worker — no code, no board, no run-state, no reporting.
340
343
  ```
341
344
  Order A (the spec tree + board):
342
345
  compile-order --operation analyze --slug <slug> --worker ba-pitch-analyzer
343
- --payload '{"pitch": "<path>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
346
+ --payload '{"pitch": "<path>", "breadboard": "<the path init run printed, when it printed one>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
344
347
  Agent (model: exec): Skill(shapeup-sdlc-plugin:ba-pitch-analyzer) --order <path>
345
348
  The order hands it code-surface.md (Phase-1 ingest, no re-scan), discovered-seed.md (task
346
349
  gen from reality), spike-<area>.md (feasibility/contracts).
@@ -462,6 +465,27 @@ Read back: the stdout JSON — {path, sha256, trial, overall, regression, score,
462
465
  (harness verify t0 calls its sibling harness probe digest internally on failure).
463
466
  ```
464
467
 
468
+ ## 3d. Round build gate → `harness verify build` (once per round, before EVAL)
469
+ ```
470
+ Invoke via Bash directly — deterministic tooling, not a worker:
471
+ node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify build --slug <slug> --round <N>
472
+ Effect: runs, in order and stopping at the first failure, the ledger's `run_cmd` (the build), then
473
+ project-profile.md's `build_probe` (the built artifact covers what the run wrote — a green
474
+ exit code is not proof the feature compiled when the toolchain compiles only what an entry
475
+ point reaches) and `launch_probe` (install, start, assert the first screen, fail on fatal
476
+ logs). Writes .shapeup/<slug>/build/r<N>-t<T>.json, immutable per run of the gate.
477
+ Read back: stdout JSON — {overall: green|red, steps[], warnings[], failed_step?, stderr_tail?}.
478
+ Exit 0 green · 1 red · 3 nothing declared (no run_cmd, no probes — logged, never green).
479
+ A `mobile` profile with no launch_probe is warned about on stderr every round, and so is
480
+ every scope none of whose fixtures invoke the tool run_cmd builds with — advisory, because a
481
+ green T0 from such fixtures is not evidence the scope compiles.
482
+ Consequences, both mechanical and both read off the artifact, never off this prose:
483
+ red → EVAL is not dispatched this round; `harness compile` turns each failing step into a
484
+ `payload.bugs` entry for round N+1, addressed to the scope whose substrate holds the
485
+ files the tool's output names (unowned → every scope, marked).
486
+ red → `reduce hill` withholds DOWNHILL_EXECUTION from every T0-green verdict of round N.
487
+ ```
488
+
465
489
  ## 4. EVAL → spec-evaluator (once per round)
466
490
  ```
467
491
  compile-order --operation evaluate --slug <slug> --worker spec-evaluator --round <r>
@@ -526,6 +550,7 @@ Read back: the proposed cut list + verdict (SHIP now | SHIP after fixing ship-bl
526
550
  | `harness-run.md` | **tech lead (sole writer)** | tech lead (round ledger + Hill + run-state), PO (audit) |
527
551
  | `scopes/<scope-id>.md` | `scope-architect` (sole writer) | tech lead (substrate/sequence), sandbox hook (write-whitelist), compile-order (inlined into orders) |
528
552
  | `t0/verdicts/r<N>-a<M>-t<T>.json` | `harness verify t0` (skill-local, mechanical — not a worker) | spec-evaluator (required citation), tech lead (hill derivation), compile-order (digested errors) |
553
+ | `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1) |
529
554
  | `t0/trials.jsonl` (the ratchet ledger, append-only, `baseline_trial` as the parent link) | `harness verify t0` (one row per attempt: score, status, delta, tree_ref) | compile-order (`trial_history` into the next order), ship-report (T0 + Ratchet sections), `harness probe stats --ratchet` |
530
555
  | `round-ledger.md` | **tech lead (sole writer)** | compile-order (decisions into every order), PO (audit) |
531
556
  | `hill/<scope-id>.yml` + `hill-chart.md` | **tech lead (sole writer)** | PO ("status without asking"), scope-hammer (H0 census) |
@@ -826,7 +851,7 @@ fixtures run in isolation and do not consume it.
826
851
  | `spike_unresolved_count` | `SPIKE-UNRESOLVED` markers at bet | shaping quality — open risk into bet |
827
852
  | `scope_cut_count` | `~` items cut at SHIP S.0 | appetite pressure / scope hammer |
828
853
  | `qa_findings` | `.shapeup/<slug>/qa/hunt-report.md` + triage → `{total, promoted, held}` | edge quality |
829
- | `slice_count` | breadboard B5 (≤9) | **normalizer / denominator** |
854
+ | `slice_count` | `receipt.json` → `breadboard.ids.V` (the V# slices init run counted in the staged breadboard; omit the field when `breadboard` is null) | **normalizer / denominator** |
830
855
  | `sources` | path to each **SHARED** source artifact — never a LOCAL `.shapeup/` path (the run-trace is superseded run by run, so a LOCAL path dangles by the time anyone reads the row; SHARED paths resolve on any clone — tier-direction rule) | auditability |
831
856
 
832
857
  - `slice_count` is the **denominator**: `round_count=4` on a 2-slice feature is alarming,