shapeup-sdlc 3.11.0 → 3.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +1 -1
- package/SECURITY.md +4 -1
- package/hooks/safety-spine.mjs +29 -2
- package/kernel/compile.mjs +7 -0
- package/kernel/gate.mjs +28 -2
- package/kernel/probe/digest.mjs +9 -2
- package/kernel/probe/stats.mjs +4 -1
- package/kernel/reduce/ship.mjs +1 -1
- package/kernel/schemas/work-order.schema.json +44 -13
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +5 -3
- package/skills/scope-architect/SKILL.md +3 -1
- package/skills/scope-hammer/SKILL.md +7 -5
- package/skills/solution-architect/SKILL.md +3 -1
- package/skills/spec-evaluator/SKILL.md +11 -2
- package/skills/task-executor/SKILL.md +6 -4
- package/skills/tech-lead/workflows/shapeup-run.js +1 -1
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.13.0",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -69,7 +69,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
69
69
|
## Setup & Execution
|
|
70
70
|
|
|
71
71
|
- Orders/results live in `.shapeup/<slug>/orders|results/`; the envelope schemas ship with the plugin runtime, not with any individual skill, so every worker validates against the same copy.
|
|
72
|
-
- The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf.
|
|
72
|
+
- The plugin's run entry points need a one-time permission grant — `npx shapeup-sdlc init` writes it into `.claude/settings.json` (`permissions.allow`); without it a headless run stalls at step one. That grant is necessary, not sufficient: it covers the run's own deterministic entry points, not the generic file edits every worker skill makes constantly, or any command a worker reaches for beyond the grant's own exact shape. A truly unattended run also needs a Claude Code permission mode that covers those (`acceptEdits` at minimum) — the plugin cannot grant that on your behalf. Two facts about how grants reach the run's workers, both measured: a worker dispatched inside a run takes its permissions from the project's settings file, not from flags given to the launching session, so a tool the workers need belongs in `permissions.allow`; and every work order names the plugin's kernel by absolute path, because a command that spells it through a variable is refused in a headless session before any permission rule is read.
|
|
73
73
|
- The grant is necessary but sits under two more layers this plugin cannot reach either. A fresh
|
|
74
74
|
checkout is an **untrusted workspace**, and Claude Code discards the whole permission grant — every
|
|
75
75
|
rule in it, not only this one — until the workspace is trusted; the installer detects that state and
|
package/SECURITY.md
CHANGED
|
@@ -58,7 +58,10 @@ test against machines you don't own.
|
|
|
58
58
|
6. **Every hook decision is recorded.** `hooks/lib/decision.mjs` is the only exit path a hook
|
|
59
59
|
has, so allow, deny, block and error each leave a row in `.shapeup/decisions.jsonl`. An
|
|
60
60
|
inert hook and a permitting hook are therefore distinguishable — which matters, because
|
|
61
|
-
"exit 0, no output" is what both used to look like.
|
|
61
|
+
"exit 0, no output" is what both used to look like. A permitted Bash call's row names the
|
|
62
|
+
programs the command runs, by basename (`hdc`, `curl | jq`) — never the arguments, which is
|
|
63
|
+
where a secret, a token or a private path would be. A denied call's row keeps the first 200
|
|
64
|
+
characters of the command, as it always has, because a denial must be reviewable.
|
|
62
65
|
|
|
63
66
|
If you find any of these to be false, that is a vulnerability — report it as claim #ⁿ.
|
|
64
67
|
|
package/hooks/safety-spine.mjs
CHANGED
|
@@ -96,6 +96,32 @@ function tokens(segment) {
|
|
|
96
96
|
.map((t) => t.replace(/^['"]|['"]$/g, ""));
|
|
97
97
|
}
|
|
98
98
|
|
|
99
|
+
/**
|
|
100
|
+
* The programs a command runs — the executable of each `&&`/`;`/`|` segment, by basename, and
|
|
101
|
+
* nothing else.
|
|
102
|
+
*
|
|
103
|
+
* WHY AN ALLOW ROW NAMES THE PROGRAM. Every permitted Bash call used to be recorded with no subject,
|
|
104
|
+
* so "did this worker ever call the device tool, and was it refused?" had no answer in the ledger:
|
|
105
|
+
* a sub-agent that never tried and one whose call was stopped above this hook left the same rows.
|
|
106
|
+
* The program names answer it. Arguments are never recorded — they are where a secret, a token or a
|
|
107
|
+
* private path would be — and a basename says which tool ran without saying where it lives.
|
|
108
|
+
*
|
|
109
|
+
* @param {string} command - The raw Bash command.
|
|
110
|
+
* @returns {(string|null)} Program basenames joined by " | ", in order, duplicates kept; null when
|
|
111
|
+
* no segment yields one.
|
|
112
|
+
*/
|
|
113
|
+
export function programsOf(command) {
|
|
114
|
+
if (!command || typeof command !== "string") return null;
|
|
115
|
+
const names = [];
|
|
116
|
+
for (const segment of command.split(/\s*(?:\|\||&&|;|\||\n)\s*/).filter(Boolean)) {
|
|
117
|
+
const first = commandTokens(segment)[0];
|
|
118
|
+
if (!first) continue;
|
|
119
|
+
const base = first.replace(/^["']|["']$/g, "").split("/").pop();
|
|
120
|
+
if (base && !/^[({]$/.test(base)) names.push(base.slice(0, 64));
|
|
121
|
+
}
|
|
122
|
+
return names.length ? names.join(" | ") : null;
|
|
123
|
+
}
|
|
124
|
+
|
|
99
125
|
/** Strip leading env assignments and privilege/no-op wrappers to find the real command. */
|
|
100
126
|
function commandTokens(segment) {
|
|
101
127
|
const ts = tokens(segment);
|
|
@@ -217,8 +243,9 @@ async function main() {
|
|
|
217
243
|
const raw = await readStdin();
|
|
218
244
|
let p;
|
|
219
245
|
/** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
|
|
220
|
-
const defer = (reason, rule) => settle({
|
|
246
|
+
const defer = (reason, rule, subject = null) => settle({
|
|
221
247
|
verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
|
|
248
|
+
...(subject ? { subject } : {}),
|
|
222
249
|
});
|
|
223
250
|
try { p = JSON.parse(raw || "{}"); }
|
|
224
251
|
catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
|
|
@@ -258,7 +285,7 @@ async function main() {
|
|
|
258
285
|
if (p.tool_name === "Bash") {
|
|
259
286
|
const command = p.tool_input?.command || "";
|
|
260
287
|
const verdict = classifyCommand(command, overrides);
|
|
261
|
-
if (!verdict.deny) defer("command matched no destructive rule — inspected and permitted", "bash-clean");
|
|
288
|
+
if (!verdict.deny) defer("command matched no destructive rule — inspected and permitted", "bash-clean", programsOf(command));
|
|
262
289
|
if (commandOverridden(command, overrides)) {
|
|
263
290
|
// Exercised override: allowed, but never invisible.
|
|
264
291
|
logPathology(metricsPath, {
|
package/kernel/compile.mjs
CHANGED
|
@@ -851,6 +851,13 @@ export function compileOrder({
|
|
|
851
851
|
const order = {
|
|
852
852
|
schema_version: 1,
|
|
853
853
|
order_id: `${slug}/${suffix}`,
|
|
854
|
+
// WHERE THE KERNEL IS, AS A PATH A PERMISSION RULE CAN MATCH. Workers were told to run
|
|
855
|
+
// `node "${CLAUDE_PLUGIN_ROOT}/kernel/harness.mjs" …`. In a headless session that variable is not
|
|
856
|
+
// expanded for them, and a command carrying an unexpanded variable is refused before any rule is
|
|
857
|
+
// consulted ("Contains expansion") — so every kernel query a worker's contract requires was
|
|
858
|
+
// refused, and workers reported it as "not permitted". The absolute path of this very file's
|
|
859
|
+
// kernel is known here, and quoted it matches the grant `init` writes.
|
|
860
|
+
kernel: join(HERE, "harness.mjs"),
|
|
854
861
|
// THE TWO ANALYTIC FIELDS, and why they are on the ORDER rather than the result.
|
|
855
862
|
//
|
|
856
863
|
// `order_id` identifies a dispatch within a run and repeats across runs of the same slug, so
|
package/kernel/gate.mjs
CHANGED
|
@@ -364,7 +364,7 @@ export function discover({ cwd = process.cwd(), file = null, preset = null, slug
|
|
|
364
364
|
export const ARGV_SPEC = {
|
|
365
365
|
usage: "harness.mjs gate (--init | --list | --verify | --resolve <gate-id>) [--preset <name>] " +
|
|
366
366
|
"[--file <path>] [--slug <slug>] [--cwd <dir>] [--out <path>] [--by <who>] " +
|
|
367
|
-
"[--auto-level <level>] [--tiny] [--no-qa] [--round <n>]",
|
|
367
|
+
"[--auto-level <level>] [--tiny] [--no-qa] [--round <n>] [--verdict <pass|fail|…>]",
|
|
368
368
|
_: { arity: 0, max: 0, name: "(no positional operands)" },
|
|
369
369
|
cwd: { type: "path" },
|
|
370
370
|
init: { type: "flag" },
|
|
@@ -383,8 +383,32 @@ export const ARGV_SPEC = {
|
|
|
383
383
|
// ledger row needs it to key a `GateDecision` node uniquely per crossing. Round-independent gates
|
|
384
384
|
// (L0, L1a, …) simply omit it and the row carries `round: null`.
|
|
385
385
|
round: { type: "int", min: 1 },
|
|
386
|
+
// The verdict the round being crossed actually reached, when there is one. A preset answers L3
|
|
387
|
+
// before any verdict exists — "loop" on a FAIL — and its note says so; recorded without the
|
|
388
|
+
// verdict, a PASS round's row read "FAIL → fix round", the opposite of what happened.
|
|
389
|
+
verdict: { type: "str" },
|
|
386
390
|
};
|
|
387
391
|
|
|
392
|
+
/**
|
|
393
|
+
* The note a gate row carries — the answer's own, unless the verdict it was crossed over makes that
|
|
394
|
+
* note false.
|
|
395
|
+
*
|
|
396
|
+
* L3's answers are written for a verdict that has not happened yet: `loop` means "on a FAIL, run
|
|
397
|
+
* the next round". Crossed over a PASS, nothing loops, and the preset's "FAIL → fix round" note
|
|
398
|
+
* would put a failed round on the record for a round that passed.
|
|
399
|
+
*
|
|
400
|
+
* @param {{gate:string, decision?:string, note?:string, reason?:string}} r - The resolved answer.
|
|
401
|
+
* @param {(string|null)} verdict - The round's verdict, lower-cased, or null when none was given.
|
|
402
|
+
* @returns {(string|null)} The note to record.
|
|
403
|
+
*/
|
|
404
|
+
export function gateNote(r, verdict) {
|
|
405
|
+
const own = r.note ?? r.reason ?? null;
|
|
406
|
+
if (r.gate === "L3" && verdict === "pass") {
|
|
407
|
+
return `verdict PASS — the "${r.decision}" answer applies only to a failed round; the run goes on to QA and GATE H`;
|
|
408
|
+
}
|
|
409
|
+
return own;
|
|
410
|
+
}
|
|
411
|
+
|
|
388
412
|
function out(obj, code = 0) {
|
|
389
413
|
console.log(JSON.stringify(obj, null, 2));
|
|
390
414
|
process.exit(code);
|
|
@@ -468,11 +492,13 @@ export function cli(rawArgv) {
|
|
|
468
492
|
// THE ROW CARRIES THE RUN KEY. It did not, and the export stamped the current run's key onto
|
|
469
493
|
// every row it found — a prior run's sign-off became this run's in the one table that answers
|
|
470
494
|
// "was this ship signed off". Driven on a two-run fixture before it was fixed.
|
|
495
|
+
const verdict = args.verdict ? String(args.verdict).toLowerCase() : null;
|
|
471
496
|
appendGateLedger(cwd, args.slug, {
|
|
472
497
|
at: new Date().toISOString(), run_id: readRunId(cwd, args.slug),
|
|
473
498
|
gate: r.gate, status: r.status, decision: r.decision ?? null,
|
|
474
|
-
source: r.source ?? found.source, note: r
|
|
499
|
+
source: r.source ?? found.source, note: gateNote(r, verdict),
|
|
475
500
|
round: args.round ?? null,
|
|
501
|
+
...(verdict ? { verdict } : {}),
|
|
476
502
|
});
|
|
477
503
|
}
|
|
478
504
|
if (r.status === "ask") out({ ...r, ok: false }, 4);
|
package/kernel/probe/digest.mjs
CHANGED
|
@@ -33,6 +33,12 @@ const PATTERNS = [
|
|
|
33
33
|
// ("ERROR in the build pipeline", "ERROR in test suite failed to run") is left unmatched
|
|
34
34
|
// instead of handing back a fabricated file.
|
|
35
35
|
{ re: /^(?:ERROR|WARNING)\s+in\s+(\.{1,2}\/[^\s:]*|[^\s:]+\.[A-Za-z0-9]{1,10})\b/i, kind: "compiler-diagnostic" },
|
|
36
|
+
// A test that FAILED BY NAME, with no file:line: "FAIL TS-05-05 step 4: no text 'Bread' on screen"
|
|
37
|
+
// or jest's "FAIL src/cart.test.js". Runners that drive an app from outside it (a device flow, an
|
|
38
|
+
// end-to-end script) report a case this way and nothing else, and the line is the whole signal —
|
|
39
|
+
// dropping it handed the next attempt an empty error list over a red fixture. The name is kept as
|
|
40
|
+
// the file only when it looks like a path; an id like TS-05-05 is not one.
|
|
41
|
+
{ re: /^(?:FAIL|FAILED)\s+(\S+)(?:\s+.*)?$/, kind: "named-test-failure" },
|
|
36
42
|
// Generic "Error: message" line followed later by a stack — capture the message alone.
|
|
37
43
|
{ re: /^\s*(?:Error|TypeError|ReferenceError|AssertionError)\s*:\s*(.+)$/, kind: "error-message" },
|
|
38
44
|
];
|
|
@@ -71,8 +77,9 @@ export function digest(rawText) {
|
|
|
71
77
|
pendingMessage = coreMessage(m[1]);
|
|
72
78
|
continue; // wait for the stack frame that follows to get a file:line
|
|
73
79
|
}
|
|
74
|
-
const
|
|
75
|
-
const
|
|
80
|
+
const named = kind === "named-test-failure";
|
|
81
|
+
const file = named ? (/[\\/]|\.[A-Za-z0-9]{1,10}$/.test(m[1]) ? m[1] : null) : m[1]?.trim();
|
|
82
|
+
const lineNo = !named && m[2] ? Number(m[2]) : null;
|
|
76
83
|
triples.push({
|
|
77
84
|
file: file || null,
|
|
78
85
|
line: lineNo,
|
package/kernel/probe/stats.mjs
CHANGED
|
@@ -194,7 +194,10 @@ export function ratchetReport(trials) {
|
|
|
194
194
|
}
|
|
195
195
|
if (scopeMonotone) monotone++;
|
|
196
196
|
}
|
|
197
|
-
|
|
197
|
+
// A score with no `regressions` field was not measured for regressions — T0 writes one only when
|
|
198
|
+
// it has a baseline to regress against — which is not the same as having some. Requiring an
|
|
199
|
+
// explicit 0 reported "no scope reached green" over a run whose every scope was green.
|
|
200
|
+
const greenAt = seq.findIndex((t) => t.score && t.score.fixtures_total > 0 && t.score.fixtures_passed === t.score.fixtures_total && (t.score.regressions ?? 0) === 0);
|
|
198
201
|
if (greenAt !== -1) toGreen.push(greenAt + 1);
|
|
199
202
|
per_scope.push({
|
|
200
203
|
scope_id, trials: seq.length,
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -280,7 +280,7 @@ export function buildReport(facts) {
|
|
|
280
280
|
L.push("| scope | fixtures | regressions | trials | last status | delta |", "|---|---|---|---|---|---|");
|
|
281
281
|
for (const s of t0) {
|
|
282
282
|
const f = s.score ? `${s.score.fixtures_passed}/${s.score.fixtures_total}` : "—";
|
|
283
|
-
const r = s.score ? String(s.score.regressions) : "—";
|
|
283
|
+
const r = s.score?.regressions != null ? String(s.score.regressions) : "—";
|
|
284
284
|
L.push(`| ${s.scope_id} | ${f} | ${r} | ${s.trials} | ${s.status} | ${s.delta || "—"} |`);
|
|
285
285
|
}
|
|
286
286
|
L.push("");
|
|
@@ -1,30 +1,61 @@
|
|
|
1
1
|
{
|
|
2
2
|
"$id": "work-order.schema.json",
|
|
3
3
|
"title": "WorkOrder",
|
|
4
|
-
"description": "The orchestrator
|
|
4
|
+
"description": "The orchestrator \u2192 worker envelope (pure-skill architecture v1.0). Compiled by kernel/compile.mjs, validated by harness verify envelope before any worker dispatch. A worker depends only on this envelope \u2014 never on filesystem topology, run-state format, board schema, or another worker. Path: .shapeup/<slug>/orders/r<N>-a<M>.json (or <slug>/orders/<operation>.json for non-attempt work). Every record type and payload field is DEFINED CENTRALLY in domain.schema.json ($defs + x-payload-by-worker) \u2014 this file only shapes the envelope; it never re-defines a domain entity.",
|
|
5
5
|
"type": "object",
|
|
6
|
-
"required": [
|
|
6
|
+
"required": [
|
|
7
|
+
"schema_version",
|
|
8
|
+
"order_id",
|
|
9
|
+
"worker",
|
|
10
|
+
"mode",
|
|
11
|
+
"payload"
|
|
12
|
+
],
|
|
7
13
|
"properties": {
|
|
8
|
-
"schema_version": {
|
|
14
|
+
"schema_version": {
|
|
15
|
+
"type": "integer",
|
|
16
|
+
"enum": [
|
|
17
|
+
1
|
|
18
|
+
]
|
|
19
|
+
},
|
|
9
20
|
"order_id": {
|
|
10
21
|
"type": "string",
|
|
11
|
-
"description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run
|
|
22
|
+
"description": "\"<slug>/r<N>-a<M>\" for build attempts, \"<slug>/<operation>[-r<N>]\" otherwise. Identifies a dispatch WITHIN a run \u2014 it repeats across runs of the same slug, which is why run_id exists.",
|
|
12
23
|
"pattern": "^[a-z0-9][a-z0-9-]*/[a-z0-9][A-Za-z0-9.-]*$"
|
|
13
24
|
},
|
|
25
|
+
"kernel": {
|
|
26
|
+
"type": "string",
|
|
27
|
+
"description": "OPTIONAL \u2014 the absolute path of the kernel entry point (harness.mjs) of the plugin copy that compiled this order. A worker runs every kernel command as node \"<kernel>\" \u2026: a command naming a variable instead (${CLAUDE_PLUGIN_ROOT}) is refused in a headless session before any permission rule is consulted."
|
|
28
|
+
},
|
|
14
29
|
"run_id": {
|
|
15
30
|
"type": "string",
|
|
16
|
-
"description": "OPTIONAL (v1.8)
|
|
31
|
+
"description": "OPTIONAL (v1.8) \u2014 the run this dispatch belongs to, stamped by harness compile from the receipt (kernel/lib/paths.mjs). The join key the analysis plane groups on: order_id alone collides across runs of the same slug, so without this no record the pipeline writes can be attributed to a run. Absent when no readable receipt exists (a standalone dispatch in a workspace with no open run) \u2014 never invented, and never required, because an analytic field must not be able to block a build.",
|
|
17
32
|
"pattern": "^[a-z0-9][a-z0-9-]*-[0-9]{8}T[0-9]{6}Z-[0-9a-f]{8}$"
|
|
18
33
|
},
|
|
19
34
|
"compiled_at": {
|
|
20
35
|
"type": "string",
|
|
21
|
-
"description": "OPTIONAL (v1.8)
|
|
22
|
-
},
|
|
23
|
-
"worker": {
|
|
24
|
-
|
|
25
|
-
|
|
26
|
-
"
|
|
27
|
-
|
|
28
|
-
|
|
36
|
+
"description": "OPTIONAL (v1.8) \u2014 ISO timestamp of compilation. The dispatch record's own time dimension: WorkResult carries none, and the journal's timings exist only on the workflow lane, so without this a dispatch on the prose lane is timeless."
|
|
37
|
+
},
|
|
38
|
+
"worker": {
|
|
39
|
+
"$ref": "domain.schema.json#/$defs/WorkerName"
|
|
40
|
+
},
|
|
41
|
+
"mode": {
|
|
42
|
+
"type": "string",
|
|
43
|
+
"enum": [
|
|
44
|
+
"orchestrated",
|
|
45
|
+
"standalone"
|
|
46
|
+
]
|
|
47
|
+
},
|
|
48
|
+
"operation": {
|
|
49
|
+
"$ref": "domain.schema.json#/$defs/Operation"
|
|
50
|
+
},
|
|
51
|
+
"interaction": {
|
|
52
|
+
"$ref": "domain.schema.json#/$defs/Interaction"
|
|
53
|
+
},
|
|
54
|
+
"substrate": {
|
|
55
|
+
"$ref": "domain.schema.json#/$defs/Substrate"
|
|
56
|
+
},
|
|
57
|
+
"payload": {
|
|
58
|
+
"$ref": "domain.schema.json#/$defs/WorkOrderPayload"
|
|
59
|
+
}
|
|
29
60
|
}
|
|
30
61
|
}
|
package/package.json
CHANGED
|
@@ -20,6 +20,8 @@ same artifacts out.
|
|
|
20
20
|
|
|
21
21
|
## Input contract — the WorkOrder
|
|
22
22
|
|
|
23
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
24
|
+
|
|
23
25
|
Invoked as `--order <path>`. Fields you may rely on (absent = unknown; surface it, never guess):
|
|
24
26
|
|
|
25
27
|
| Field | What it is |
|
|
@@ -69,10 +71,10 @@ its phase; templates live in `assets/templates/`.
|
|
|
69
71
|
6 TASKS atomic, ordered, executable → tasks/ (LOCAL root; the one uncommitted branch
|
|
70
72
|
of the tree — regenerable, machine-local) [references/task-generation.md]
|
|
71
73
|
7 DERIVE+LINT mechanical, not yours to grade:
|
|
72
|
-
node "
|
|
74
|
+
node "<kernel>" reduce board --slug <slug> --write
|
|
73
75
|
(unlocks = depends_on inverse; Σ hours; critical path; appetite arithmetic —
|
|
74
76
|
overflow is a fact you REPORT for the caller's HAMMER gate, never resolve)
|
|
75
|
-
node "
|
|
77
|
+
node "<kernel>" verify spec --slug <slug>
|
|
76
78
|
(structure, wikilinks, edge symmetry — fix reds, then re-run; you never
|
|
77
79
|
self-grade with a hand-walked checklist. BREADBOARD-PLACE / BREADBOARD-UI:
|
|
78
80
|
add the screen or defer the Place; never fold it into another screen)
|
|
@@ -187,7 +189,7 @@ status flips for built work (ingest's job), scope contracts (scope-architect's),
|
|
|
187
189
|
/ba-pitch-analyzer --order .shapeup/checkout-vnpay/orders/analyze.json
|
|
188
190
|
|
|
189
191
|
# Standalone — the preamble shim compiles the order (mode: standalone, pause_gates: true):
|
|
190
|
-
# node "
|
|
192
|
+
# node "<kernel>" compile --operation analyze --slug <slug> \
|
|
191
193
|
# --worker ba-pitch-analyzer --payload '{"pitch": "docs/pitch.md", "lens": "standard"}'
|
|
192
194
|
/ba-pitch-analyzer docs/pitch.md # operation: analyze, lens judged
|
|
193
195
|
/ba-pitch-analyzer --lens standard docs/pitch.md # lens pinned
|
|
@@ -17,6 +17,8 @@ the ship report's census table.
|
|
|
17
17
|
|
|
18
18
|
## Input contract — the WorkOrder
|
|
19
19
|
|
|
20
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
21
|
+
|
|
20
22
|
| Field | What it is |
|
|
21
23
|
|---|---|
|
|
22
24
|
| `operation` | `map-scopes` — the only operation this skill has. It covers first slicing after the board exists, folding discovered items in, and re-slicing a stuck scope; the payload says which of those you are doing |
|
|
@@ -95,7 +97,7 @@ the ship report's census table.
|
|
|
95
97
|
hill_phase: "UPHILL_UNKNOWN" — ALWAYS; phase is derived from
|
|
96
98
|
T0/T1 facts later,
|
|
97
99
|
never authored
|
|
98
|
-
4 LINT node "
|
|
100
|
+
4 LINT node "<kernel>" verify spec --slug <slug>
|
|
99
101
|
→ PA1 (directory alignment), PA2 (>~15 files), DISJOINT (undeclared overlap),
|
|
100
102
|
SCOPE-ANCHOR (empty/unresolvable use_cases), TIER-DIRECTION (a task id in a
|
|
101
103
|
committed contract), SCOPE-DEPS (depends_on naming a scope that isn't here).
|
|
@@ -52,13 +52,15 @@ INPUT: run's finished/unfinished scopes + baseline + census sources
|
|
|
52
52
|
|
|
53
53
|
## GATE H0 — Census
|
|
54
54
|
|
|
55
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
56
|
+
|
|
55
57
|
**Purpose:** Gather every open item into one list before judging any of them. Never judge
|
|
56
58
|
piecemeal — a partial view produces a wrong cut.
|
|
57
59
|
|
|
58
60
|
```
|
|
59
61
|
H0.0 Ownership is DERIVED, never stated. Before the census says "no scope owns X" or "X is
|
|
60
62
|
scope Y's", run
|
|
61
|
-
node "
|
|
63
|
+
node "<kernel>" probe owner --slug <slug> [--path <p>]...
|
|
62
64
|
and cite its row. With no --path it answers for every engine and entry call site the wiring
|
|
63
65
|
map names plus the profile's entry point; `writers: []` is an unowned seam and `missing`
|
|
64
66
|
lists seams the wiring names that are not on disk — owned but never written. A census that
|
|
@@ -70,7 +72,7 @@ H0.1 Unresolved scopes (breaker cases only):
|
|
|
70
72
|
- scopes with hammer_proposals (attempt_budget exhausted) → CARRY candidates. Exhaustion
|
|
71
73
|
is DERIVED, never read off `t0/verdicts/*.json` directly — a compiled order or a T0
|
|
72
74
|
verdict is writable by the very scope being judged and proves nothing on its own. Run
|
|
73
|
-
node "
|
|
75
|
+
node "<kernel>" probe attempts --slug <slug> \
|
|
74
76
|
--scope <scope-id> --round <n> --attempt-budget <n>
|
|
75
77
|
and cite its `spent`/`tripped` fields (exit 1 = tripped) — an attempt counts only when a
|
|
76
78
|
dispatch receipt AND either a leg-completion row or a WorkResult attest it, so a leg still
|
|
@@ -81,7 +83,7 @@ H0.2 QA findings (qa-edge-hunter's hunt-report.md, when present) — all `~` by
|
|
|
81
83
|
H0.3 Discovered-task ledger entries still open (discovery/ledger.md, `[+]`/`~` unresolved).
|
|
82
84
|
H0.4 Attempt-budget hammer proposals (scopes that exhausted their T0 attempts during BUILD).
|
|
83
85
|
H0.4b Requirements with no PASS evidence — the pitch clauses the run never showed working. Run
|
|
84
|
-
node "
|
|
86
|
+
node "<kernel>" probe requirements --slug <slug> --format table
|
|
85
87
|
and take its `no evidence` rows; cite the row, the same way H0.0 cites ownership. Each is a
|
|
86
88
|
census item carrying its source clause (`REQ-12 ← shaping.md R12`). A `cut` row is an answer
|
|
87
89
|
the PO already gave — not an item. An inconsistency row (a criterion anchored to a
|
|
@@ -200,8 +202,8 @@ with a green census could only be recorded as `ask`. Still a proposal: the file
|
|
|
200
202
|
/scope-hammer --slug checkout-vnpay --unattended
|
|
201
203
|
|
|
202
204
|
# The ownership query every census claim cites (H0.0)
|
|
203
|
-
node "
|
|
204
|
-
node "
|
|
205
|
+
node "<kernel>" probe owner --slug checkout-vnpay --format table
|
|
206
|
+
node "<kernel>" probe owner --slug checkout-vnpay --path src/pages/Cart.ets
|
|
205
207
|
```
|
|
206
208
|
|
|
207
209
|
### Flags
|
|
@@ -38,6 +38,8 @@ return as a WorkResult.
|
|
|
38
38
|
|
|
39
39
|
## Input contract — the WorkOrder
|
|
40
40
|
|
|
41
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
42
|
+
|
|
41
43
|
| Field | What it is |
|
|
42
44
|
|---|---|
|
|
43
45
|
| `operation` | `wire` (author/refresh the wiring map after `analyze`, before `map-scopes`) |
|
|
@@ -152,5 +154,5 @@ seam, or an engine with no attachment path, and why). You never touch spec docs,
|
|
|
152
154
|
# The reachability oracle is the ORCHESTRATOR's, run advisory at L1b — not part of your craft.
|
|
153
155
|
# Standalone, you MAY preview it after writing the map (it self-skips arms whose artifacts are
|
|
154
156
|
# absent, and is near-vacuous pre-build since the engine code does not exist yet):
|
|
155
|
-
# node "
|
|
157
|
+
# node "<kernel>" verify trace --slug checkout-vnpay
|
|
156
158
|
```
|
|
@@ -29,6 +29,8 @@ criterion with no collected evidence is a **FAIL**, never a pass-by-assumption.
|
|
|
29
29
|
|
|
30
30
|
## Input contract — the WorkOrder
|
|
31
31
|
|
|
32
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
33
|
+
|
|
32
34
|
Invoked as `--order <path>`. Fields you may rely on (absent = unknown, never inferred):
|
|
33
35
|
|
|
34
36
|
| Field | What it is |
|
|
@@ -83,6 +85,13 @@ Done-when statements; `_index.md` Non-Go list. Which UCs are in scope comes from
|
|
|
83
85
|
freeze through the judge). Ugly-but-correct PASSes; pretty-but-wrong-`data-state` FAILs.
|
|
84
86
|
- `[data]`: query the DB/storage, capture actual state.
|
|
85
87
|
- Contract work: send real requests, compare field-by-field.
|
|
88
|
+
- **A fixture that names a row is evidence for that row.** When a T0 artifact you cite, or the
|
|
89
|
+
build gate, carries output that names a Test Surface row by id — `PASS TS-05-05`, or
|
|
90
|
+
`FAIL TS-05-05 step 4: …` — grade that row on it, after reading the check that printed it: a
|
|
91
|
+
named PASS confirms only when that check asserts the row's Expect; a check weaker than its row
|
|
92
|
+
(it opens the dialog, the row says pre-filled) is not evidence for the row, and grading it PASS
|
|
93
|
+
is the generator grading itself. A named FAIL is a FAIL whose bug is that line. This is how a
|
|
94
|
+
device row is evidenced when you cannot drive the app; never a reason not to when you can.
|
|
86
95
|
- No evidence collected = recorded "NO EVIDENCE" → FAILs at verdict.
|
|
87
96
|
|
|
88
97
|
**VERDICT.**
|
|
@@ -217,7 +226,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
|
|
|
217
226
|
/spec-evaluator --order .shapeup/checkout-vnpay/orders/evaluate-r2.json
|
|
218
227
|
|
|
219
228
|
# Standalone — the preamble shim compiles a minimal order, then the single code path runs:
|
|
220
|
-
# node "
|
|
229
|
+
# node "<kernel>" compile --operation evaluate --slug <slug> \
|
|
221
230
|
# --worker spec-evaluator [--payload '{"dimensions": [...], "run_cmd": "..."}']
|
|
222
231
|
/spec-evaluator --spec shapeup/checkout-vnpay/spec/ --task TASK-007
|
|
223
232
|
/spec-evaluator --spec shapeup/checkout-vnpay/spec/ --feature checkout-vnpay --single-pass
|
|
@@ -225,7 +234,7 @@ bug_template). Adding one (e.g. security) = write `references/dimensions/securit
|
|
|
225
234
|
|
|
226
235
|
Standalone keeps `--task` (per-task check, not round-gated) and `--single-pass` (feature-level)
|
|
227
236
|
— the shim maps them onto the order's payload; missing run command → ask. After writing the
|
|
228
|
-
WorkResult, run `node "
|
|
237
|
+
WorkResult, run `node "<kernel>" reduce ingest <result path>` and show its
|
|
229
238
|
summary — standalone has no orchestrator to ingest for you.
|
|
230
239
|
|
|
231
240
|
---
|
|
@@ -16,6 +16,8 @@ it does not exist for you.
|
|
|
16
16
|
|
|
17
17
|
## Input contract — the WorkOrder
|
|
18
18
|
|
|
19
|
+
**Kernel commands.** `<kernel>` below is the absolute path in your order's `kernel` field: run every kernel command as `node "<kernel>" …`. Never spell it `${CLAUDE_PLUGIN_ROOT}` — a command carrying a variable is refused in a headless session before any permission rule is read, and the query your contract requires never runs. With no order (invoked by hand), the kernel is `kernel/harness.mjs` two directories above this skill's base directory.
|
|
20
|
+
|
|
19
21
|
You are invoked as `--order <path>` pointing at a schema-valid WorkOrder. Fields you may
|
|
20
22
|
rely on (anything absent = **unknown**; never invent it):
|
|
21
23
|
|
|
@@ -25,7 +27,7 @@ rely on (anything absent = **unknown**; never invent it):
|
|
|
25
27
|
| `payload.scope_contract` | The active scope: `affordance_manifest`, `e2e_verification_fixtures`, topology |
|
|
26
28
|
| `substrate.allowed` / `substrate.shared` | The ONLY globs you may write. A needed file outside them → ESCALATE, never a write (a sandbox hook blocks it anyway) |
|
|
27
29
|
| `payload.decisions[]` | Adjudicated answers from prior escalations — binding precedent, apply them |
|
|
28
|
-
| `payload.digested_errors[]` | `{file, line, core_message}` triples from the previous attempt's failed verification — your starting bug list |
|
|
30
|
+
| `payload.digested_errors[]` | `{file, line, core_message}` triples from the previous attempt's failed verification (a test that failed by name with no file:line — `FAIL TS-05-05 step 4: …` — arrives with `file: null` and the whole line as its message) — your starting bug list |
|
|
29
31
|
| `payload.trial_history[]` | Up to 8 prior attempts on this scope, oldest first, CROSSING the round boundary: `{score, status, delta, digest}`. `status: "reverted"` is a change that was tried and made things WORSE — do not re-propose it. `status: "kept"` with a still-red score is the tree you are building ON, not a failure to undo. Absent on the first attempt |
|
|
30
32
|
| `payload.verify.test_cmd` | The command that verifies your work. No test_cmd → command-verifiable ACs still need *some* observable check; say what you used |
|
|
31
33
|
| `payload.kb_rules_path` | Team guidelines (read if the file exists) — steering, never spec; conflict → the AC wins, note it in `deviations` |
|
|
@@ -194,8 +196,8 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
|
|
|
194
196
|
|
|
195
197
|
# Standalone — the preamble shim compiles a minimal WorkOrder from the flags, then the
|
|
196
198
|
# single code path above runs. Requires the harness scripts (plugin install):
|
|
197
|
-
# node "
|
|
198
|
-
# node "
|
|
199
|
+
# node "<kernel>" compile --task TASK-003 --slug checkout-vnpay
|
|
200
|
+
# node "<kernel>" compile --next --slug checkout-vnpay
|
|
199
201
|
/task-executor --spec shapeup/checkout-vnpay/spec/ --task TASK-003
|
|
200
202
|
/task-executor --spec shapeup/checkout-vnpay/spec/ --next
|
|
201
203
|
```
|
|
@@ -203,6 +205,6 @@ orchestrator's `harness reduce ingest` does all of that from your envelope.
|
|
|
203
205
|
Standalone shim: derive `<slug>` from the `--spec` path (`shapeup/<slug>/spec`),
|
|
204
206
|
run `harness compile` with the matching flags (mode becomes `standalone`), then proceed
|
|
205
207
|
against the compiled order exactly as if dispatched. After writing the WorkResult, run
|
|
206
|
-
`node "
|
|
208
|
+
`node "<kernel>" reduce ingest <result path>` yourself and show the user its
|
|
207
209
|
summary — standalone has no orchestrator to ingest for you. One code path inside; two entry
|
|
208
210
|
points outside.
|
|
@@ -898,7 +898,7 @@ async function crossGate(gateId, phaseName, validDecisions, ctx) {
|
|
|
898
898
|
// The gate ledger keys a per-round crossing (L2, L3) on gate id + round, so it needs the round
|
|
899
899
|
// whenever the caller already has one to show in the block — the same value `ctx.round` carries
|
|
900
900
|
// for display, threaded through rather than re-derived.
|
|
901
|
-
const roundFlag = ctx?.round != null ? ` --round ${ctx.round}` : "";
|
|
901
|
+
const roundFlag = (ctx?.round != null ? ` --round ${ctx.round}` : "") + (ctx?.verdict ? ` --verdict ${ctx.verdict}` : "");
|
|
902
902
|
const g = await cmd(`gate --resolve ${gateId} --slug ${slug}${roundFlag} ${answersFlag(args.answers)}`.trim(), phaseName, `gate:${gateId}`);
|
|
903
903
|
if (g.exit_code === 4) return { stop: paused(gateId, validDecisions, ctx) };
|
|
904
904
|
if (g.exit_code === 5) return { stop: aborted(gateId, g.detail || `GATE ${gateId} aborted`) };
|