shapeup-sdlc 3.1.2 → 3.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +7 -6
- package/README.md +1 -1
- package/SECURITY.md +4 -1
- package/bin/init.mjs +3 -0
- package/commands/retro.md +19 -2
- package/commands/ship.md +5 -0
- package/hooks/dispatch-receipt.mjs +6 -3
- package/hooks/gate-intake.mjs +1 -1
- package/hooks/gate-zerowork.mjs +5 -2
- package/hooks/lib/decision.mjs +49 -6
- package/hooks/safety-spine.mjs +8 -5
- package/hooks/sandbox-guard.mjs +13 -6
- package/kernel/compile.mjs +112 -6
- package/kernel/harness.mjs +11 -5
- package/kernel/init/run.mjs +122 -6
- package/kernel/lib/breadboard.mjs +165 -0
- package/kernel/lib/paths.mjs +15 -3
- package/kernel/probe/owner.mjs +139 -0
- package/kernel/probe/resume.mjs +7 -1
- package/kernel/probe/stats.mjs +49 -2
- package/kernel/reduce/hill.mjs +12 -2
- package/kernel/verify/build.mjs +319 -0
- package/kernel/verify/spec.mjs +190 -2
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +11 -5
- package/skills/ba-pitch-analyzer/assets/templates/_index.tmpl.md +2 -1
- package/skills/ba-pitch-analyzer/assets/templates/ux-behavior.tmpl.md +12 -2
- package/skills/ba-pitch-analyzer/references/doc-schemas.md +3 -0
- package/skills/ba-pitch-analyzer/references/ux-behavior-patterns.md +9 -0
- package/skills/coach/SKILL.md +232 -43
- package/skills/orient/SKILL.md +15 -5
- package/skills/qa-edge-hunter/SKILL.md +3 -2
- package/skills/scope-architect/SKILL.md +7 -0
- package/skills/scope-hammer/SKILL.md +13 -0
- package/skills/solution-architect/SKILL.md +7 -2
- package/skills/tech-lead/SKILL.md +7 -7
- package/skills/tech-lead/references/gates.md +69 -17
- package/skills/tech-lead/references/protocol.md +33 -8
- package/skills/tech-lead/schemas/domain.schema.json +180 -9
- package/skills/tech-lead/workflows/shapeup-run.js +54 -12
|
@@ -56,6 +56,14 @@ INPUT: run's finished/unfinished scopes + baseline + census sources
|
|
|
56
56
|
piecemeal — a partial view produces a wrong cut.
|
|
57
57
|
|
|
58
58
|
```
|
|
59
|
+
H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns X" or "X is
|
|
60
|
+
scope Y's", run
|
|
61
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug <slug> [--path <p>]...
|
|
62
|
+
and cite its row. With no --path it answers for every engine and entry call site the wiring
|
|
63
|
+
map names plus the profile's entry point; `writers: []` is an unowned seam and `missing`
|
|
64
|
+
lists seams the wiring names that are not on disk — owned but never written. A census that
|
|
65
|
+
narrated ownership from memory once told the PO no scope owned a screen directory that a
|
|
66
|
+
committed contract listed in plain sight — and pointed the ship decision at the wrong gap.
|
|
59
67
|
H0.1 Unresolved scopes (breaker cases only):
|
|
60
68
|
- uphill/downhill scopes when round_budget hit 0 → CARRY candidates (their own hill
|
|
61
69
|
phase + open unknowns, from hill/<scope-id>.yml)
|
|
@@ -160,6 +168,10 @@ The WorkResult may carry only `files_touched`, `artifacts`, `assumptions`, `devi
|
|
|
160
168
|
|
|
161
169
|
# Headless — no PO available; still refuses to auto-ship a ship-blocking item
|
|
162
170
|
/scope-hammer --slug checkout-vnpay --unattended
|
|
171
|
+
|
|
172
|
+
# The ownership query every census claim cites (H0.0)
|
|
173
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --format table
|
|
174
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
|
|
163
175
|
```
|
|
164
176
|
|
|
165
177
|
### Flags
|
|
@@ -181,6 +193,7 @@ The WorkResult may carry only `files_touched`, `artifacts`, `assumptions`, `devi
|
|
|
181
193
|
| Default classification is NICE-TO-HAVE unless traced to a pitch boundary or business_goal | A generous must-have list defeats the point of hammering |
|
|
182
194
|
| A MUST-HAVE that fails H1.2 is never cut silently | The one case where scope-hammer refuses to make the run "look" shippable |
|
|
183
195
|
| Cuts are proposals; the PO confirms every one | This skill never overrides the human at the ship gate |
|
|
196
|
+
| Every ownership claim in the census cites `harness probe owner` | Ownership is what the contracts' substrates say — the same election `harness compile` uses to address a bug — never what the report remembers |
|
|
184
197
|
| Cut items are carried to the discovery ledger, never silently dropped | Cool-down must stay debt-free — an idea deferred is still recorded |
|
|
185
198
|
| Never sets status: done, never deploys, never ships unilaterally | Judge/doer/advisor separation holds even at the very last gate |
|
|
186
199
|
| An overridden ship-blocking item is logged explicitly in the ship report | "Shipped" must never quietly mean "shipped with a known must-have gap" |
|
|
@@ -42,6 +42,8 @@ return as a WorkResult.
|
|
|
42
42
|
|---|---|
|
|
43
43
|
| `operation` | `wire` (author/refresh the wiring map after `analyze`, before `map-scopes`) |
|
|
44
44
|
| `payload.feature` / `payload.spec_folder` | Slug + committed spec — read `usecases/` for the UCs and the engine each one needs, `domain-model.md`/`synthesis.md` for the module surface |
|
|
45
|
+
| `payload.breadboard` | When present, name each UC's `affordance` by its U# and Place |
|
|
46
|
+
| `payload.kb_rules_path` | Team guidelines (read if the file exists) — the seams this codebase actually wires through, entry points that are not where the template says. Steering, never spec: the profile's `entry_point` still wins, and a guideline that disagrees with it is reported in `deviations`, not applied |
|
|
45
47
|
| `payload.project_profile` | Path to the SHARED `project-profile.md`. Its `entry_point` is the composition root every engine must attach to — **archetype-specific** (a client-only game's `main.js` is not a web-service's `src/server.ts`). Read it; never guess the entry point |
|
|
46
48
|
| `substrate.allowed` | `wiring-map.md` — your ONLY write surface (the spec core, scopes, and the profile are frozen) |
|
|
47
49
|
|
|
@@ -52,7 +54,8 @@ guessed `main.js` would make the later oracle certify nothing.
|
|
|
52
54
|
## Core process
|
|
53
55
|
|
|
54
56
|
```
|
|
55
|
-
1 READ the
|
|
57
|
+
1 READ the team guidelines at payload.kb_rules_path if the file exists (absent = none),
|
|
58
|
+
then the project profile → entry_point + archetype. Read every use case in usecases/.
|
|
56
59
|
For each UC, identify the engine module that carries its core logic (the file that
|
|
57
60
|
WILL exist, named from the domain model / synthesis surface — not a guess at a folder).
|
|
58
61
|
2 DESIGN for each UC, design the integration path from the entry_point inward:
|
|
@@ -69,7 +72,9 @@ guessed `main.js` would make the later oracle certify nothing.
|
|
|
69
72
|
file:line is a build-time fact (the oracle proves reachability by
|
|
70
73
|
the import graph, it does not parse this field)
|
|
71
74
|
affordance the player-visible thing this UC exposes once wired (the human
|
|
72
|
-
end of the chain — what a user can DO, not an internal call)
|
|
75
|
+
end of the chain — what a user can DO, not an internal call);
|
|
76
|
+
with a breadboard, name it by U# and Place —
|
|
77
|
+
`U1 Pay (P1) → P2 Payment Sheet`
|
|
73
78
|
3 WRITE shapeup/<slug>/wiring-map.md (WiringMap): frontmatter for schema_version, feature
|
|
74
79
|
and entry_point (echo of the profile), then entries[] as ONE MARKDOWN TABLE under a
|
|
75
80
|
`## Wiring` heading — this exact shape, because it is the only one the reader parses:
|
|
@@ -23,7 +23,7 @@ turns fighting shell quoting):
|
|
|
23
23
|
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" init run \
|
|
24
24
|
--slug <slug-from-the-request> --intake-file <path/to/the/requirement.md> \
|
|
25
25
|
--auto-level <interactive|auto|unattended> \
|
|
26
|
-
[--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3]
|
|
26
|
+
[--dimensions <a,b>] [--gate-answers <ci|guarded|path.json>] [--wall-clock-budget <seconds>] [--max-rounds 3] [--breadboard <path>]
|
|
27
27
|
```
|
|
28
28
|
|
|
29
29
|
**After a compaction, or in a fresh session over an open run, re-derive before you act.** One
|
|
@@ -43,13 +43,13 @@ counts.
|
|
|
43
43
|
plugin and need a one-time permission grant (`npx shapeup-sdlc init` writes it). Do not route
|
|
44
44
|
around it, and do not silently hand-build the feature instead.
|
|
45
45
|
|
|
46
|
-
**Language gate (delegated to `translator`, not this skill):**
|
|
47
|
-
an Agent (model: exec)
|
|
48
|
-
English → proceed as-is. Non-English →
|
|
49
|
-
|
|
50
|
-
|
|
46
|
+
**Language gate (delegated to `translator`, not this skill):** before Step 1 opens the run, dispatch
|
|
47
|
+
an Agent (model: exec) calling `Skill(shapeup-sdlc-plugin:translator) --check` on the pitch *and* its
|
|
48
|
+
breadboard. English → proceed as-is. Non-English → a second Agent translates both (`--auto` under
|
|
49
|
+
auto/unattended); Step 1 then names the `.en.md` files (a run already open on the original: re-open
|
|
50
|
+
it with `--force` — nothing is dispatched yet). The tech lead detects and sequences; never translates.
|
|
51
51
|
|
|
52
|
-
**Step 2 — pin GATE L0, then launch.** Collect the L0.1–L0.
|
|
52
|
+
**Step 2 — pin GATE L0, then launch.** Collect the L0.1–L0.10 config (spec folder, lens, stack,
|
|
53
53
|
eval dims, max_rounds, the model/budget matrix — see `references/gates.md` GATE L0 for the full
|
|
54
54
|
collect-list), write the SHARED `project-profile.md` yourself (`{schema_version:1, archetype,
|
|
55
55
|
entry_point}` — `shapeup-run.js` has no filesystem of its own), emit the `⏸ GATE L0` block, then
|
|
@@ -31,13 +31,19 @@ escalates, writes nothing, and every relaunch re-dispatches it.
|
|
|
31
31
|
|
|
32
32
|
```
|
|
33
33
|
Collect (explicit — never inferred):
|
|
34
|
-
L0.1 Kicked-off pitch source:
|
|
34
|
+
L0.1 Kicked-off pitch source: `shaping.md` — and its `breadboard.md`, which the run finds
|
|
35
|
+
beside it (or in `shaping/`) or takes from `--breadboard`; a `pitch.md` may carry the
|
|
36
|
+
breadboard inline. Already shaped + bet by PO.
|
|
35
37
|
Not a raw idea — shaping (1-4) / betting (5) / kick-off (6) are PO-personal, upstream.
|
|
36
|
-
L0.1a Language gate: Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator)
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
38
|
+
L0.1a Language gate, BEFORE init run: Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator)
|
|
39
|
+
--check over the pitch AND its breadboard.
|
|
40
|
+
English → use both as-is.
|
|
41
|
+
non-English → Agent (model: exec) → Skill(shapeup-sdlc-plugin:translator) <pitch> <breadboard>
|
|
42
|
+
(--auto under auto/unattended), then open the run on the produced
|
|
43
|
+
<name>.en.md — init run prefers a breadboard.en.md beside it, and warns
|
|
44
|
+
when it can find only the untranslated breadboard. The run's intake is
|
|
45
|
+
what every planning worker reads, so a translation made after the run
|
|
46
|
+
opened reaches nobody: re-open with --force naming the .en.md. Log in ledger.
|
|
41
47
|
L0.1b Appetite: read the `appetite` field from the pitch's YAML frontmatter (set by /shapeup).
|
|
42
48
|
Surface it in the gate output. Use it to:
|
|
43
49
|
- Contextualise the scope at L1b (right-size cuts to the budget).
|
|
@@ -90,6 +96,21 @@ Collect (explicit — never inferred):
|
|
|
90
96
|
same GATE H proposal. attempt_budget counts ATTEMPTS and cannot see that the last two
|
|
91
97
|
produced nothing; this term can, and on a flailing scope it saves three of five
|
|
92
98
|
attempts. Set per scope (`no_progress_k` on the contract) or per run in the payload.
|
|
99
|
+
L0.10 knowledge base (read, never obeyed): if `shapeup/knowledge-base/tech-lead.md` exists,
|
|
100
|
+
read it now. Its **Workflow guidance** may add a question, a check or a warning line to
|
|
101
|
+
any gate block below and may name the spike to insist on at L1a; its **Suggested run
|
|
102
|
+
config** lines are PROPOSALS for L0.2–L0.9 and the profile — confirm each with the PO
|
|
103
|
+
before pinning it, and record `(source: knowledge-base)` beside a value taken from there.
|
|
104
|
+
Nothing in that file answers a gate: the answer set resolves exactly as it would without
|
|
105
|
+
it, and a line that could only be honoured by skipping, reordering or relaxing a gate is
|
|
106
|
+
reported as a suspected harness defect, not applied. If `shapeup/knowledge-base/` has no
|
|
107
|
+
file at all, offer the optional seed once — `/retro --scan` (the coach's scan operation)
|
|
108
|
+
drafts guidelines from the project on disk, or `/retro --research <stack>` (its
|
|
109
|
+
research operation) from the platform's official documentation when the project has no
|
|
110
|
+
build file to scan; both put every rule through GATE COACH-1 — and continue whether or
|
|
111
|
+
not the PO takes it. A Suggested run config line with `web-research` provenance has
|
|
112
|
+
never run in this project: confirm it like any other, and expect the first round build
|
|
113
|
+
gate to be its first execution.
|
|
93
114
|
```
|
|
94
115
|
|
|
95
116
|
**L0.9b — the launch record.** Every switch the operator typed becomes a `RunArgs` field, or it
|
|
@@ -134,9 +155,11 @@ Intake lang : [English | translated via /translator → <name>.en.md]
|
|
|
134
155
|
Appetite : [~1 week | ~2 weeks | ~6 weeks | ⚠️ missing — scope uncapped]
|
|
135
156
|
Spec folder : [path] (lens: [lite|standard])
|
|
136
157
|
Eval dims : [spec-conformance] max_rounds: [N, appetite-informed] auto: [interactive|auto|unattended]
|
|
137
|
-
Run commands : [web: ... | api: ... | mobile: ...]
|
|
158
|
+
Run commands : [web: ... | api: ... | mobile: ...] (run_cmd → the round build gate, every round before EVAL)
|
|
159
|
+
Build gate : build_probe [set | —] launch_probe [set | — ⚠ mobile: the install/launch risk has no owner]
|
|
138
160
|
Model matrix : orch=[model] exec=[model] eval=[model] qa=[model] digester=[script|sonnet] (source: [flags|settings.local|settings.json|default])
|
|
139
161
|
Budgets : round_budget=[N] (outer) attempt_budget=[N] (inner, per scope)
|
|
162
|
+
Knowledge : [tech-lead.md — N workflow rules, M suggested values (confirmed above) | none — `/retro --scan` or `/retro --research <stack>` seeds it (optional)]
|
|
140
163
|
```
|
|
141
164
|
Do NOT start ORIENT until confirmed (interactive/auto). Under --unattended, proceed.
|
|
142
165
|
|
|
@@ -167,8 +190,9 @@ committing to a scope map. This is the first Hill read (area-level — slices do
|
|
|
167
190
|
```
|
|
168
191
|
Read .shapeup/<slug>/orient/. Render the 🗻 Hill from hill-signal.md (see protocol.md "Hill report"):
|
|
169
192
|
- each suspected area → uphill (open unknowns) | crest (approach proven by the spike) | downhill
|
|
170
|
-
Print:
|
|
171
|
-
|
|
193
|
+
Print: `Breadboard: <source> | none` (how init run found the pitch's breadboard — flag, sibling,
|
|
194
|
+
shaping-dir, shared-root, embedded — or none), the code-surface headline (where it lands),
|
|
195
|
+
the spiked area + result, the riskiest open unknowns going into mapping.
|
|
172
196
|
Ask (max 2): is the riskiest area the right one to have spiked? any unknown that must be
|
|
173
197
|
resolved (another spike) before we map scopes?
|
|
174
198
|
```
|
|
@@ -182,11 +206,17 @@ Do NOT enter MAP SCOPES until Orient is accepted.
|
|
|
182
206
|
|
|
183
207
|
```
|
|
184
208
|
1. PROFILE (you write it at L0 — compile-order stays pipeline-blind): SHARED project-profile.md
|
|
185
|
-
= {schema_version:1, archetype, entry_point}. archetype ∈
|
|
186
|
-
mobile|library|data-pipeline}; entry_point is the reachability
|
|
187
|
-
service's src/server.ts). Validate the enum — a typo must fail,
|
|
209
|
+
= {schema_version:1, archetype, entry_point, build_probe?, launch_probe?}. archetype ∈
|
|
210
|
+
{client-only-game|web-service|mobile|library|data-pipeline}; entry_point is the reachability
|
|
211
|
+
seam (a game's main.js is NOT a service's src/server.ts). Validate the enum — a typo must fail,
|
|
212
|
+
not silently disable the check. The two probes feed the round build gate (`harness verify
|
|
213
|
+
build`, every round before EVAL): build_probe asserts the BUILT ARTIFACT covers what the run
|
|
214
|
+
wrote (a green exit code is not proof the feature compiled when the toolchain compiles only what
|
|
215
|
+
an entry point reaches); launch_probe installs, starts and asserts the first screen. A `mobile`
|
|
216
|
+
profile without a launch_probe is warned about every round — nothing else in the loop launches
|
|
217
|
+
the app.
|
|
188
218
|
2. WIRE — compile-order --operation wire --slug <slug> (worker→solution-architect), payload
|
|
189
|
-
{project_profile}. Sole writer of committed wiring-map.md (per-UC engine → seam → entry-point
|
|
219
|
+
{project_profile, breadboard?}. Sole writer of committed wiring-map.md (per-UC engine → seam → entry-point
|
|
190
220
|
call site → affordance). ⏸ GATE L1a.5: confirm each UC has a declared seam before slicing.
|
|
191
221
|
⟐ PRECONDITION: MAP SCOPES step 1 (ANALYZE) has already run and usecases/ is
|
|
192
222
|
populated. WIRE writes one entry per use case, so dispatching it against an empty spec folder
|
|
@@ -212,7 +242,7 @@ against the seams WIRE declared. Sequence: ORIENT → L1a → **ANALYZE** → **
|
|
|
212
242
|
```
|
|
213
243
|
Two orders, two workers, one step (both model: exec — see references/protocol.md):
|
|
214
244
|
1. ANALYZE + BOARD — compile-order --operation analyze --slug <slug> --worker ba-pitch-analyzer
|
|
215
|
-
--payload '{"pitch": "<path>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
|
|
245
|
+
--payload '{"pitch": "<path>", "breadboard": "<the path init run printed, when it printed one>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
|
|
216
246
|
dispatch: Skill(shapeup-sdlc-plugin:ba-pitch-analyzer) --order <path>. The order hands it
|
|
217
247
|
code-surface.md (Phase-1 ingest consumes the map, does not re-scan), discovered-seed.md
|
|
218
248
|
(task gen starts from reality), spike-<area>.md (feasibility/contracts).
|
|
@@ -265,6 +295,8 @@ Scope contracts present:
|
|
|
265
295
|
- scope board: scope_id, topology_type, substrate file count (scopes/*.md / scope-board.md)
|
|
266
296
|
- any SPIKE blockers (scope-summary.md)
|
|
267
297
|
- scope-summary "Done when" headline statements
|
|
298
|
+
- the Deferred Places from ux-behavior.md (breadboard Places this shape will not build) —
|
|
299
|
+
each one needs the PO's yes; a rejected deferral goes back to the planner as a screen
|
|
268
300
|
No scope contracts (pre-v0.3.0, unchanged from v0.2.6):
|
|
269
301
|
Read tasks/_index.md (LOCAL root). Print:
|
|
270
302
|
- task count by package/variant (.shared / .be / .web / .mobile / .e2e)
|
|
@@ -280,7 +312,11 @@ this is the orchestrator's own re-confirmation before committing to a build sequ
|
|
|
280
312
|
waiting to happen), PA1 (directory-aligned scope), PA2 (size cap), SCOPE-ANCHOR (a scope
|
|
281
313
|
naming no committed use case, or one that does not resolve), TIER-DIRECTION (a committed
|
|
282
314
|
contract naming LOCAL task ids), SCOPE-DEPS (a build-order id naming a scope that is not
|
|
283
|
-
in this run)
|
|
315
|
+
in this run), BREADBOARD-PLACE (a breadboard Place with UI affordances has no
|
|
316
|
+
`## Screen: … (P#)` in ux-behavior.md and is not deferred), BREADBOARD-UI (a U# not
|
|
317
|
+
specified on a screen of its own Place). Any red → HARD STOP, past a 🔴 at the
|
|
318
|
+
architect's own checkpoint. The breadboard reds are the planner's to fix — add the screen
|
|
319
|
+
or defer the Place; never fold it into another screen.
|
|
284
320
|
- Lock the build SEQUENCE riskiest-first: order scopes by open-unknowns count (from
|
|
285
321
|
hill/<scope-id>.yml if present, else the orient hill signal), not by file count or
|
|
286
322
|
alphabetical — Shape Up's "solve in the right sequence" (step 10).
|
|
@@ -315,8 +351,16 @@ Do NOT enter BUILD until the board is accepted.
|
|
|
315
351
|
Feature : [slug]
|
|
316
352
|
Round : [r]
|
|
317
353
|
Scopes : [N] green, [M] queued for hammer
|
|
354
|
+
Build gate: [green | red — <failing step> | undeclared]
|
|
318
355
|
```
|
|
319
356
|
|
|
357
|
+
The build gate is `harness verify build --slug <slug> --round <r>` — the ledger's `run_cmd`, then
|
|
358
|
+
the profile's `build_probe` and `launch_probe`, stopping at the first failure. It runs before this
|
|
359
|
+
gate so the block shows it; `red` means EVAL is NOT dispatched this round (the judge grades a running
|
|
360
|
+
feature) and the failing step is compiled into round r+1's orders as `payload.bugs`. `undeclared`
|
|
361
|
+
means L0 pinned no run command and the profile names no probe — the round proceeds over an
|
|
362
|
+
unproven build, and the block says so.
|
|
363
|
+
|
|
320
364
|
Under `--interactive` / `--auto`, the hook warns if the board is not truly green (advisory) and requires explicit PO approval to proceed. Under `--unattended`, it automatically aborts on a red board or proceeds on a green one.
|
|
321
365
|
|
|
322
366
|
---
|
|
@@ -344,6 +388,10 @@ PASS:
|
|
|
344
388
|
→ --no-qa or skill absent: proceed straight to SHIP; ledger records `qa: skipped`.
|
|
345
389
|
|
|
346
390
|
FAIL:
|
|
391
|
+
→ a round whose build gate was red never reached the judge: the block carries `build_gate: red`,
|
|
392
|
+
the "bug list" is the gate's failing step (its command, exit and output tail), and round r+1
|
|
393
|
+
is a fix round over exactly that. Nothing the evaluator would have said is missing — it was
|
|
394
|
+
never asked.
|
|
347
395
|
→ print the bug list grouped by task/severity. For each bug: task ID + failed Done-when criterion + repro.
|
|
348
396
|
DO NOT prescribe fix options or root cause hypotheses — that is the implementer's job.
|
|
349
397
|
The tech lead names scope; the implementer diagnoses and fixes.
|
|
@@ -378,6 +426,9 @@ S.0 GATE H — delegate to scope-hammer (this IS Shape Up's "Decide When to Sto
|
|
|
378
426
|
(no --breaker flag) normal stop — all scopes FINISHED, post-QA-hunt
|
|
379
427
|
Feeds it: qa/hunt-report.md findings (when present), discovery/ledger.md open items,
|
|
380
428
|
the hammer-proposal queue from BUILD (attempt-budget exhaustions).
|
|
429
|
+
Ownership facts in its census — "no scope owns X", "X belongs to scope Y" — come from
|
|
430
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" probe owner --slug <slug> [--path <p>]...
|
|
431
|
+
which elects the owner from the committed contracts' substrates, never from prose.
|
|
381
432
|
Reads back: its GATE H0/H1/H2 output — census, baseline comparison, cut list + verdict.
|
|
382
433
|
Authority: scope-hammer proposes; the tech lead records the PO's decision in
|
|
383
434
|
round-ledger.md and performs the actual close (S.1 onward). It never ships on its own.
|
|
@@ -414,7 +465,8 @@ S.6 Harvest one signal row → append to `.shapeup/metrics/<machine-id>.jsonl`
|
|
|
414
465
|
deliberately without colliding on one filename. The read plane is
|
|
415
466
|
`harness probe stats`, or `cat .shapeup/metrics/*.jsonl`).
|
|
416
467
|
Copy fields that ALREADY exist as structured output (run-state, final EVAL report,
|
|
417
|
-
discovery ledger, qa/hunt-report, breadboard
|
|
468
|
+
discovery ledger, qa/hunt-report, and the receipt's `breadboard.ids.V` for `slice_count` —
|
|
469
|
+
the slices init run counted in the staged breadboard, never a hand copy). Two hard rules:
|
|
418
470
|
1. Harvest only fields that already exist at ship time — never evaluate something new.
|
|
419
471
|
2. Record facts, never compute a new verdict (no `run_quality_score` — that would be
|
|
420
472
|
a second judge behind spec-evaluator). The eval suite interprets; harvest records.
|
|
@@ -465,7 +517,7 @@ Ledger : harness-run.md
|
|
|
465
517
|
```
|
|
466
518
|
Question (max 1): "Anything to record before I close the run? (y/n) or provide feedback for the next sprint."
|
|
467
519
|
On confirm:
|
|
468
|
-
- If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable
|
|
520
|
+
- If the PO provides substantive feedback (not just 'y' or empty) → automatically delegate via Agent (model: exec — see references/protocol.md "Invocation mechanism"): Skill(shapeup-sdlc-plugin:coach) with the provided feedback for RLHF. The coach runs its own GATE COACH-1 to have the PO categorize each rule, then files it under the responsible skill in `shapeup/knowledge-base/<skill>.md` (committed → team-shared). Coachable: `task-executor`, `ba-pitch-analyzer`, `qa-edge-hunter`, `orient`, `scope-architect`, `solution-architect` (each reads its own file at the top of its next run) and `tech-lead` (workflow guidance, read at the next GATE L0). Guidance never decides a gate: a filed rule may add a question or a check to a gate block, never an answer. The tech lead does not categorize the feedback itself — that is the coach's gate, by design (no assumptions).
|
|
469
521
|
- Then output → `✅ [slug] [shipped & deployed | built & verified, deploy pending] — [r] rounds, verdict PASS.`
|
|
470
522
|
|
|
471
523
|
---
|
|
@@ -68,7 +68,9 @@ approval of the new tasks and estimates before resuming the BUILD loop.
|
|
|
68
68
|
|
|
69
69
|
## The EVAL timing rule (the core constraint)
|
|
70
70
|
|
|
71
|
-
EVAL fires **once** per round and **only** when GATE L2 has confirmed the board is 100% done
|
|
71
|
+
EVAL fires **once** per round and **only** when GATE L2 has confirmed the board is 100% done
|
|
72
|
+
and the round build gate (§3d) is not red — a feature that does not build or launch has nothing
|
|
73
|
+
for a judge to grade, and the gate's failing step is what round r+1 fixes.
|
|
72
74
|
It is never:
|
|
73
75
|
- called per task,
|
|
74
76
|
- called inside the BUILD loop,
|
|
@@ -311,14 +313,15 @@ workers keep only their own product-idempotency key and emit domain artifacts.
|
|
|
311
313
|
|
|
312
314
|
## 0. LANGUAGE GATE → translator (GATE L0, only if non-English)
|
|
313
315
|
```
|
|
314
|
-
Invoke via Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) --check "<intake path>"
|
|
316
|
+
Invoke via Agent (model: exec), BEFORE init run: Skill(shapeup-sdlc-plugin:translator) --check "<intake path>" "<breadboard path>"
|
|
315
317
|
# detect-only, writes nothing
|
|
316
318
|
English → skip; ORIENT against the original.
|
|
317
|
-
non-English → Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) "<intake path>" [--auto]
|
|
318
|
-
# full pass
|
|
319
|
-
Writes: <name>.en.md (English copy; original untouched) + glossary.md
|
|
319
|
+
non-English → Agent (model: exec): Skill(shapeup-sdlc-plugin:translator) "<intake path>" "<breadboard path>" [--auto]
|
|
320
|
+
# full pass — the pitch AND its breadboard, before init run
|
|
321
|
+
Writes: <name>.en.md per file (English copy; original untouched) + glossary.md
|
|
320
322
|
+ translation-report.md.
|
|
321
|
-
|
|
323
|
+
Open the run on the .en.md pitch: init run stages breadboard.en.md beside it,
|
|
324
|
+
and every planning worker reads what init run pinned — never a later copy.
|
|
322
325
|
Read back: the detect table (--check) / the .en.md path + residual scan result (full pass).
|
|
323
326
|
Authority: translator normalizes language only — it does not orient/plan/build/judge. The tech
|
|
324
327
|
lead never translates itself; it only detects and sequences this step before ORIENT.
|
|
@@ -340,7 +343,7 @@ Authority: pure worker — no code, no board, no run-state, no reporting.
|
|
|
340
343
|
```
|
|
341
344
|
Order A (the spec tree + board):
|
|
342
345
|
compile-order --operation analyze --slug <slug> --worker ba-pitch-analyzer
|
|
343
|
-
--payload '{"pitch": "<path>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
|
|
346
|
+
--payload '{"pitch": "<path>", "breadboard": "<the path init run printed, when it printed one>", "lens": "<lens>", "orient_dir": ".shapeup/<slug>/orient/"}'
|
|
344
347
|
Agent (model: exec): Skill(shapeup-sdlc-plugin:ba-pitch-analyzer) --order <path>
|
|
345
348
|
The order hands it code-surface.md (Phase-1 ingest, no re-scan), discovered-seed.md (task
|
|
346
349
|
gen from reality), spike-<area>.md (feasibility/contracts).
|
|
@@ -462,6 +465,27 @@ Read back: the stdout JSON — {path, sha256, trial, overall, regression, score,
|
|
|
462
465
|
(harness verify t0 calls its sibling harness probe digest internally on failure).
|
|
463
466
|
```
|
|
464
467
|
|
|
468
|
+
## 3d. Round build gate → `harness verify build` (once per round, before EVAL)
|
|
469
|
+
```
|
|
470
|
+
Invoke via Bash directly — deterministic tooling, not a worker:
|
|
471
|
+
node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" verify build --slug <slug> --round <N>
|
|
472
|
+
Effect: runs, in order and stopping at the first failure, the ledger's `run_cmd` (the build), then
|
|
473
|
+
project-profile.md's `build_probe` (the built artifact covers what the run wrote — a green
|
|
474
|
+
exit code is not proof the feature compiled when the toolchain compiles only what an entry
|
|
475
|
+
point reaches) and `launch_probe` (install, start, assert the first screen, fail on fatal
|
|
476
|
+
logs). Writes .shapeup/<slug>/build/r<N>-t<T>.json, immutable per run of the gate.
|
|
477
|
+
Read back: stdout JSON — {overall: green|red, steps[], warnings[], failed_step?, stderr_tail?}.
|
|
478
|
+
Exit 0 green · 1 red · 3 nothing declared (no run_cmd, no probes — logged, never green).
|
|
479
|
+
A `mobile` profile with no launch_probe is warned about on stderr every round, and so is
|
|
480
|
+
every scope none of whose fixtures invoke the tool run_cmd builds with — advisory, because a
|
|
481
|
+
green T0 from such fixtures is not evidence the scope compiles.
|
|
482
|
+
Consequences, both mechanical and both read off the artifact, never off this prose:
|
|
483
|
+
red → EVAL is not dispatched this round; `harness compile` turns each failing step into a
|
|
484
|
+
`payload.bugs` entry for round N+1, addressed to the scope whose substrate holds the
|
|
485
|
+
files the tool's output names (unowned → every scope, marked).
|
|
486
|
+
red → `reduce hill` withholds DOWNHILL_EXECUTION from every T0-green verdict of round N.
|
|
487
|
+
```
|
|
488
|
+
|
|
465
489
|
## 4. EVAL → spec-evaluator (once per round)
|
|
466
490
|
```
|
|
467
491
|
compile-order --operation evaluate --slug <slug> --worker spec-evaluator --round <r>
|
|
@@ -526,6 +550,7 @@ Read back: the proposed cut list + verdict (SHIP now | SHIP after fixing ship-bl
|
|
|
526
550
|
| `harness-run.md` | **tech lead (sole writer)** | tech lead (round ledger + Hill + run-state), PO (audit) |
|
|
527
551
|
| `scopes/<scope-id>.md` | `scope-architect` (sole writer) | tech lead (substrate/sequence), sandbox hook (write-whitelist), compile-order (inlined into orders) |
|
|
528
552
|
| `t0/verdicts/r<N>-a<M>-t<T>.json` | `harness verify t0` (skill-local, mechanical — not a worker) | spec-evaluator (required citation), tech lead (hill derivation), compile-order (digested errors) |
|
|
553
|
+
| `build/r<N>-t<T>.json` (the round build gate) | `harness verify build` (mechanical — run_cmd + build_probe + launch_probe, once per round before EVAL) | tech lead (GATE L2 block), `harness reduce hill` (a red round moves no dot), compile-order (`payload.bugs` for round N+1) |
|
|
529
554
|
| `t0/trials.jsonl` (the ratchet ledger, append-only, `baseline_trial` as the parent link) | `harness verify t0` (one row per attempt: score, status, delta, tree_ref) | compile-order (`trial_history` into the next order), ship-report (T0 + Ratchet sections), `harness probe stats --ratchet` |
|
|
530
555
|
| `round-ledger.md` | **tech lead (sole writer)** | compile-order (decisions into every order), PO (audit) |
|
|
531
556
|
| `hill/<scope-id>.yml` + `hill-chart.md` | **tech lead (sole writer)** | PO ("status without asking"), scope-hammer (H0 census) |
|
|
@@ -826,7 +851,7 @@ fixtures run in isolation and do not consume it.
|
|
|
826
851
|
| `spike_unresolved_count` | `SPIKE-UNRESOLVED` markers at bet | shaping quality — open risk into bet |
|
|
827
852
|
| `scope_cut_count` | `~` items cut at SHIP S.0 | appetite pressure / scope hammer |
|
|
828
853
|
| `qa_findings` | `.shapeup/<slug>/qa/hunt-report.md` + triage → `{total, promoted, held}` | edge quality |
|
|
829
|
-
| `slice_count` | breadboard
|
|
854
|
+
| `slice_count` | `receipt.json` → `breadboard.ids.V` (the V# slices init run counted in the staged breadboard; omit the field when `breadboard` is null) | **normalizer / denominator** |
|
|
830
855
|
| `sources` | path to each **SHARED** source artifact — never a LOCAL `.shapeup/` path (the run-trace is superseded run by run, so a LOCAL path dangles by the time anyone reads the row; SHARED paths resolve on any clone — tier-direction rule) | auditability |
|
|
831
856
|
|
|
832
857
|
- `slice_count` is the **denominator**: `round_count=4` on a 2-slice feature is alarming,
|