shapeup-sdlc 3.7.0 → 3.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +3 -2
- package/README.md +11 -4
- package/SECURITY.md +3 -2
- package/hooks/hooks.json +10 -0
- package/hooks/tier-guard.mjs +162 -0
- package/kernel/compile.mjs +15 -1
- package/kernel/gate.mjs +67 -2
- package/kernel/probe/attempts.mjs +22 -5
- package/kernel/probe/resume.mjs +45 -9
- package/kernel/reduce/ship.mjs +39 -3
- package/kernel/schemas/domain.schema.json +15 -2
- package/kernel/verify/spec.mjs +52 -20
- package/kernel/verify/t0.mjs +67 -2
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/references/doc-schemas.md +4 -0
- package/skills/tech-lead/references/tiny-lane.md +2 -1
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.7.
|
|
4
|
+
"version": "3.7.1",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -7,7 +7,8 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
|
|
|
7
7
|
|
|
8
8
|
Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
|
|
9
9
|
|
|
10
|
-
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, stopping a run with no receipt (the run's first act writes one).
|
|
10
|
+
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
|
|
11
|
+
- **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
|
|
11
12
|
- Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
|
|
12
13
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
13
14
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
@@ -80,7 +81,7 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
80
81
|
through the chained launch, until the run can resume unattended again.
|
|
81
82
|
- Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
|
|
82
83
|
- **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
|
|
83
|
-
- **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright —
|
|
84
|
+
- **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
|
|
84
85
|
- Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
|
|
85
86
|
- Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
|
|
86
87
|
`reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
|
package/README.md
CHANGED
|
@@ -210,7 +210,7 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
210
210
|
| **Runtime** — the kernel and the run script | When the run goes through the harness. A schema rejection or a non-zero exit stops the step. | A lane that never calls the kernel is never checked — which is why the hooks below cover the doors, not the steps. |
|
|
211
211
|
| **Advisory** — a report section | When somebody reads the artifact. | Silently, if nobody does. It is a cleanup list, never a verdict. |
|
|
212
212
|
|
|
213
|
-
**
|
|
213
|
+
**Five walls.** These are hooks because nothing in the runtime can substitute for them:
|
|
214
214
|
|
|
215
215
|
- `PreToolUse` (`Skill`) — **`hooks/gate-intake.mjs` denies a `tech-lead` dispatch that carries no
|
|
216
216
|
pitch, no spec folder, and no requirement text.** Observed, not theorized: when the requirement
|
|
@@ -227,6 +227,13 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
227
227
|
checked first and outranks everything, across every live contract — including the carve-out that
|
|
228
228
|
otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
|
|
229
229
|
frozen wherever it lives.
|
|
230
|
+
- `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
|
|
231
|
+
whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
|
|
232
|
+
GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
|
|
233
|
+
own lint then rejected, and one of them did it twice in one session, an hour apart, after fixing
|
|
234
|
+
the first occurrence itself. A worker carries no lesson across a dispatch, so the remedy is a
|
|
235
|
+
refusal the writer can act on while it still knows what it meant to say. The knowledge base is
|
|
236
|
+
outside it — those files are instructions, not references a reader has to resolve.
|
|
230
237
|
- `PreToolUse` (`Bash|Read|Write|Edit|MultiEdit`) — **`hooks/safety-spine.mjs` denies destructive
|
|
231
238
|
commands** (`rm -rf` on unrecoverable targets, force-push/push-to-main, `git reset --hard`,
|
|
232
239
|
`DROP TABLE`) and secret-file reads. A machine guard, not a pipeline guard; the escape hatch is
|
|
@@ -265,7 +272,7 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
265
272
|
| `anti-rationalization` (claims the facts contradict) | The ship report's census, derived from the board and the T0 artifacts | The facts are in an artifact a teammate finds on `git pull`, not in a transcript nobody re-reads. |
|
|
266
273
|
| `slop-cleaner` (TODO/`console.log` leftovers) | The ship report's **Leftovers** section | Same scan, same added-lines-only rule; it lands somewhere checkable. |
|
|
267
274
|
|
|
268
|
-
**Nothing load-bearing depends on permission mode.** The
|
|
275
|
+
**Nothing load-bearing depends on permission mode.** The five walls plus the zero-work gate run
|
|
269
276
|
under every mode. The kernel needs a grant to be *invoked* — two Bash lines `npx shapeup-sdlc init`
|
|
270
277
|
writes — but a session that never gets that grant is a session that cannot run the pipeline at all,
|
|
271
278
|
not one that runs it unguarded.
|
|
@@ -338,8 +345,8 @@ kernel/{verify,reduce,probe,init,report}/ # its subcommands, plus compile and
|
|
|
338
345
|
kernel/lib/ # argv (the typed CLI boundary), paths (+ the run key), contract (shape)
|
|
339
346
|
kernel/schemas/ # the envelope port: WorkOrder, WorkResult, domain registry
|
|
340
347
|
commands/*.md # slash commands (/ship + the 9 phase commands)
|
|
341
|
-
hooks/ # hooks.json + the
|
|
342
|
-
# (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
|
|
348
|
+
hooks/ # hooks.json + the five walls: safety-spine, gate-intake, sandbox-guard,
|
|
349
|
+
# tier-guard (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
|
|
343
350
|
# + dispatch-receipt (PostToolUse, denies nothing, attests which skill ran)
|
|
344
351
|
# + lib/decision.mjs (every hook records allow / deny / error)
|
|
345
352
|
oracles/ # the evaluation-contract oracle registry (test · snapshot · http · process)
|
package/SECURITY.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Security
|
|
2
2
|
|
|
3
|
-
This plugin installs **
|
|
4
|
-
`PreToolUse` position, **all
|
|
3
|
+
This plugin installs **eight hook entries (seven Node scripts + one `echo`)**: five in a
|
|
4
|
+
`PreToolUse` position, **all five of which can deny a tool call**, one `PostToolUse` hook that has no
|
|
5
5
|
deny path at all, and one `Stop`-position hook that can block a session from ending. That is the
|
|
6
6
|
product — and it is also exactly the kind of surface a careful reviewer should want spelled out
|
|
7
7
|
before installing. This page is that spelling-out.
|
|
@@ -73,6 +73,7 @@ sitting, and reading them is the recommended review.
|
|
|
73
73
|
| [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
|
|
74
74
|
| [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
|
|
75
75
|
| [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
|
|
76
|
+
| [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
|
|
76
77
|
| [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
|
|
77
78
|
| [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
|
|
78
79
|
|
package/hooks/hooks.json
CHANGED
|
@@ -50,6 +50,16 @@
|
|
|
50
50
|
"timeout": 10
|
|
51
51
|
}
|
|
52
52
|
]
|
|
53
|
+
},
|
|
54
|
+
{
|
|
55
|
+
"matcher": "Edit|Write|MultiEdit",
|
|
56
|
+
"hooks": [
|
|
57
|
+
{
|
|
58
|
+
"type": "command",
|
|
59
|
+
"command": "node \"${CLAUDE_PLUGIN_ROOT}/hooks/tier-guard.mjs\"",
|
|
60
|
+
"timeout": 10
|
|
61
|
+
}
|
|
62
|
+
]
|
|
53
63
|
}
|
|
54
64
|
],
|
|
55
65
|
"PostToolUse": [
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Tier guard — PreToolUse hook: the committed tier's write boundary.
|
|
3
|
+
//
|
|
4
|
+
// Refuses an Edit/Write/MultiEdit into `shapeup/<slug>/` whose CONTENT carries a reference the
|
|
5
|
+
// committed tier cannot hold — a path into the gitignored run trace, or a machine-local board id.
|
|
6
|
+
// It is the same rule spec-lint's TIER-DIRECTION already enforces, asked at a different moment.
|
|
7
|
+
//
|
|
8
|
+
// WHY A SECOND ENFORCEMENT POINT FOR A RULE THAT WAS NEVER IN DOUBT. Four different producers wrote
|
|
9
|
+
// committed files this harness's own lint then reds: the requirements registry, the ship report, the
|
|
10
|
+
// project profile, the coverage clauses. The lint was right every time and caught every one of them
|
|
11
|
+
// — at GATE L1b, a whole phase after the sentence was written, addressed to a worker that no longer
|
|
12
|
+
// holds the context needed to rephrase it. The run pays for the round trip, and the fix lands in
|
|
13
|
+
// whatever wording the next reader guesses at.
|
|
14
|
+
//
|
|
15
|
+
// TEACHING DOES NOT CLOSE IT, and that is measured rather than assumed. One of the four recurred
|
|
16
|
+
// INSIDE A SINGLE SESSION, an hour apart, on a different line of the same file, after that same
|
|
17
|
+
// author had fixed the first occurrence and watched its own lint go green. A worker carries no
|
|
18
|
+
// lesson across a dispatch. So the remedy has to be a precondition the model cannot talk past, and
|
|
19
|
+
// it has to speak AT THE WRITE — the one moment the writer still knows what it meant to say.
|
|
20
|
+
//
|
|
21
|
+
// THE PREDICATE IS IMPORTED, NEVER RESTATED. `tierLeaks` is the lint's own scanner
|
|
22
|
+
// (kernel/verify/spec.mjs), so this hook cannot become narrower than the rule it fronts — which
|
|
23
|
+
// would let the defect through to L1b exactly as before — nor wider, which is how a guard earns
|
|
24
|
+
// being switched off. The structural suite executes both over one corpus and requires agreement.
|
|
25
|
+
//
|
|
26
|
+
// Design (deliberately conservative, the shape `sandbox-guard.mjs` argues for):
|
|
27
|
+
// • Fail-OPEN on everything it cannot positively prove: not a write tool, no resolvable path, a
|
|
28
|
+
// path outside `shapeup/<slug>/`, a file form the lint does not scan, or a payload carrying no
|
|
29
|
+
// content to read. A guard that blocks legitimate work gets disabled, and a disabled guard
|
|
30
|
+
// enforces nothing.
|
|
31
|
+
// • `shapeup/knowledge-base/` is OUTSIDE by construction, exactly as it is outside the lint's
|
|
32
|
+
// walk: those files are instructions read by a worker at runtime, not references a reader is
|
|
33
|
+
// expected to resolve, and the defect register they hold has every reason to quote a local path.
|
|
34
|
+
// • Fail-CLOSED only on a leak the lint would red, with the file, the line, the offending token,
|
|
35
|
+
// why it is wrong and what to write instead — all in the denial, because a denial the writer
|
|
36
|
+
// cannot act on immediately just becomes the L1b round trip with extra steps.
|
|
37
|
+
//
|
|
38
|
+
// WHAT IT DOES NOT COVER, stated because the gap is the same one the substrate fence has: this is a
|
|
39
|
+
// PreToolUse hook, so it sees this assistant's own edit path and nothing else. A kernel subcommand
|
|
40
|
+
// writing a committed file, a shell heredoc, or any other editor writes straight through it. Those
|
|
41
|
+
// producers are answered where they are written; this closes the channel the four measured
|
|
42
|
+
// recurrences actually came through.
|
|
43
|
+
//
|
|
44
|
+
// Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path, content | new_string |
|
|
45
|
+
// edits[]}, cwd }. Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny",
|
|
46
|
+
// permissionDecisionReason } }.
|
|
47
|
+
|
|
48
|
+
import { resolve, relative, sep } from "node:path";
|
|
49
|
+
import { isMain } from "../kernel/lib/argv.mjs";
|
|
50
|
+
import { SHARED } from "../kernel/lib/paths.mjs";
|
|
51
|
+
import { SCANNED, tierLeaks } from "../kernel/verify/spec.mjs";
|
|
52
|
+
import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
|
|
53
|
+
|
|
54
|
+
/** Committed-tier trees that hold instructions rather than references — outside the lint's walk. */
|
|
55
|
+
const NOT_A_SLUG = new Set(["knowledge-base"]);
|
|
56
|
+
|
|
57
|
+
/**
|
|
58
|
+
* The (path, content) pairs one write tool call is about to commit to disk.
|
|
59
|
+
*
|
|
60
|
+
* THE THREE WRITE TOOLS SPELL CONTENT THREE WAYS, and a guard that reads only `content` is blind to
|
|
61
|
+
* the Edit that adds the same line — which is the form the measured recurrence took, since the file
|
|
62
|
+
* already existed by then.
|
|
63
|
+
*
|
|
64
|
+
* @param {object} toolInput - The PreToolUse `tool_input` block.
|
|
65
|
+
* @returns {Array<{path:string, content:string}>} Pairs with a usable path and string content.
|
|
66
|
+
*/
|
|
67
|
+
export function extractWrites(toolInput) {
|
|
68
|
+
const out = [];
|
|
69
|
+
const base = toolInput?.file_path;
|
|
70
|
+
if (typeof toolInput?.content === "string" && base) out.push({ path: base, content: toolInput.content });
|
|
71
|
+
if (typeof toolInput?.new_string === "string" && base) out.push({ path: base, content: toolInput.new_string });
|
|
72
|
+
if (Array.isArray(toolInput?.edits)) {
|
|
73
|
+
for (const e of toolInput.edits) {
|
|
74
|
+
const p = e?.file_path || base;
|
|
75
|
+
if (p && typeof e?.new_string === "string") out.push({ path: p, content: e.new_string });
|
|
76
|
+
}
|
|
77
|
+
}
|
|
78
|
+
return out;
|
|
79
|
+
}
|
|
80
|
+
|
|
81
|
+
/**
|
|
82
|
+
* Is this path inside the committed tier the lint walks — `shapeup/<slug>/…`, scanned form?
|
|
83
|
+
*
|
|
84
|
+
* @param {string} rel - Path relative to the project root, as the lint would name it.
|
|
85
|
+
* @returns {boolean} True when a leak here is a leak the lint would red.
|
|
86
|
+
*/
|
|
87
|
+
export function inCommittedTier(rel) {
|
|
88
|
+
if (!rel || rel.startsWith("..") || rel.startsWith(sep)) return false;
|
|
89
|
+
const parts = rel.split(/[\\/]/);
|
|
90
|
+
// `shapeup/<slug>/<file>` — the lint walks a slug's tree, so a file at the tier root belongs to
|
|
91
|
+
// no run and is left alone, and `knowledge-base` is a sibling of the slugs, not one of them.
|
|
92
|
+
if (parts[0] !== SHARED || parts.length < 3 || NOT_A_SLUG.has(parts[1])) return false;
|
|
93
|
+
return SCANNED.test(rel);
|
|
94
|
+
}
|
|
95
|
+
|
|
96
|
+
async function main() {
|
|
97
|
+
await runHook("tier-guard", async () => {
|
|
98
|
+
const raw = await readStdin();
|
|
99
|
+
let p;
|
|
100
|
+
/** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
|
|
101
|
+
const defer = (reason, rule) => settle({
|
|
102
|
+
verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
|
|
103
|
+
});
|
|
104
|
+
try { p = JSON.parse(raw || "{}"); }
|
|
105
|
+
catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
|
|
106
|
+
|
|
107
|
+
if (!["Edit", "Write", "MultiEdit"].includes(p.tool_name)) {
|
|
108
|
+
defer(`${p.tool_name ?? "no tool_name"} is not a write tool — out of scope`, "not-write-tool");
|
|
109
|
+
}
|
|
110
|
+
|
|
111
|
+
// The shell's cwd is where the call fired; the tier lives at the project root. Same split the
|
|
112
|
+
// substrate fence has to make, and for the same reason: a worker that `cd`s into a sub-folder
|
|
113
|
+
// must not thereby leave the tier unguarded. Raw tool paths still resolve against the shell.
|
|
114
|
+
const cwd = p.cwd || process.cwd();
|
|
115
|
+
const root = projectRoot(cwd);
|
|
116
|
+
|
|
117
|
+
const writes = extractWrites(p.tool_input);
|
|
118
|
+
if (writes.length === 0) defer("no readable path+content pair in the tool input", "no-content");
|
|
119
|
+
|
|
120
|
+
const inTier = writes
|
|
121
|
+
.map((w) => ({ ...w, rel: relative(root, resolve(cwd, w.path)) }))
|
|
122
|
+
.filter((w) => inCommittedTier(w.rel));
|
|
123
|
+
if (inTier.length === 0) defer(`${writes.length} write(s), none inside ${SHARED}/<slug>/ in a scanned form`, "not-committed-tier");
|
|
124
|
+
|
|
125
|
+
const blocked = [];
|
|
126
|
+
for (const w of inTier) {
|
|
127
|
+
// THE TOKEN IS QUOTED BACK, which the lint's own message does not do and does not need to —
|
|
128
|
+
// it reports against a file on disk the reader can open at that line. Here the line does not
|
|
129
|
+
// exist yet, so "line 1 points into the local tier" leaves the writer hunting through a
|
|
130
|
+
// fragment it is holding in its head. Naming the exact string is the difference between a
|
|
131
|
+
// denial that is acted on and one that is retried verbatim.
|
|
132
|
+
for (const leak of tierLeaks(w.content)) blocked.push(`${w.rel}:${leak.line} — \`${leak.token}\` ${leak.detail}`);
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
if (blocked.length === 0) {
|
|
136
|
+
defer(`${inTier.length} committed-tier write(s) carry no tier-direction leak — permitted`, "tier-clean");
|
|
137
|
+
}
|
|
138
|
+
|
|
139
|
+
return {
|
|
140
|
+
verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: inTier[0].rel, cwd: root,
|
|
141
|
+
rule: "tier-direction",
|
|
142
|
+
reason: `${blocked.length} tier-direction leak(s) refused at the write boundary: ${blocked.join("; ")}`,
|
|
143
|
+
payload: {
|
|
144
|
+
hookSpecificOutput: {
|
|
145
|
+
hookEventName: "PreToolUse",
|
|
146
|
+
permissionDecision: "deny",
|
|
147
|
+
permissionDecisionReason:
|
|
148
|
+
"Tier guard (TIER-DIRECTION) — this write would put a reference into the committed tier that " +
|
|
149
|
+
"cannot survive the trip to another machine:\n" +
|
|
150
|
+
`${blocked.join("\n")}\n` +
|
|
151
|
+
`Refused here rather than at GATE L1b, where spec-lint reds the same file after the phase is over. ` +
|
|
152
|
+
"Rephrase the line now: cite the committed artifact, the use case or the scope_id, or describe the " +
|
|
153
|
+
"tier without naming a path.",
|
|
154
|
+
},
|
|
155
|
+
},
|
|
156
|
+
};
|
|
157
|
+
});
|
|
158
|
+
}
|
|
159
|
+
|
|
160
|
+
if (isMain(import.meta.url)) {
|
|
161
|
+
main();
|
|
162
|
+
}
|
package/kernel/compile.mjs
CHANGED
|
@@ -1007,7 +1007,21 @@ export async function cli(rawArgv) {
|
|
|
1007
1007
|
// than the flailing it detects. It advises; the orchestrator queues the GATE H proposal.
|
|
1008
1008
|
if (scope?.scope_id) {
|
|
1009
1009
|
const k = Number(scope.no_progress_k ?? payloadExtra.no_progress_k ?? 2);
|
|
1010
|
-
|
|
1010
|
+
// SCOPED TO THIS RUN, for the same reason the attempt census is. `t0/trials.jsonl` is per-slug
|
|
1011
|
+
// and append-only, so a streak survives the run that produced it. Measured on a consumer: two
|
|
1012
|
+
// non-kept trials from earlier runs — one of them graded against an order no worker was ever
|
|
1013
|
+
// dispatched for — read as a stagnation streak that every later run of that scope tripped on,
|
|
1014
|
+
// seven hours after the tree they graded had been fixed. The breaker escalated on evidence that
|
|
1015
|
+
// predated its own fix, and no attempt of the tripping run had been dispatched at all.
|
|
1016
|
+
//
|
|
1017
|
+
// A trial with no run key counts for no run; an unresolvable current run matches nothing. Both
|
|
1018
|
+
// leave the breaker un-tripped, which is the fail-open direction this repo's guards take when
|
|
1019
|
+
// the bad state cannot be positively proven.
|
|
1020
|
+
const myRunId = readRunId(cwd, slug);
|
|
1021
|
+
const st = stagnation(
|
|
1022
|
+
allTrials.filter((t) => t.scope_id === scope.scope_id && myRunId != null && t.run_id === myRunId),
|
|
1023
|
+
k,
|
|
1024
|
+
);
|
|
1011
1025
|
if (st.stagnant) {
|
|
1012
1026
|
console.error(JSON.stringify({
|
|
1013
1027
|
breaker: "stagnation", scope_id: scope.scope_id, streak: st.streak, no_progress_k: st.k,
|
package/kernel/gate.mjs
CHANGED
|
@@ -53,7 +53,7 @@
|
|
|
53
53
|
import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } from "node:fs";
|
|
54
54
|
import { join, dirname } from "node:path";
|
|
55
55
|
import { runArgs } from "./lib/argv.mjs";
|
|
56
|
-
import { gateAnswerCandidates, gates as gatesPath, LOCAL } from "./lib/paths.mjs";
|
|
56
|
+
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir } from "./lib/paths.mjs";
|
|
57
57
|
|
|
58
58
|
export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
|
|
59
59
|
|
|
@@ -168,6 +168,71 @@ export function validate(set) {
|
|
|
168
168
|
* Returns { gate, decision, source, note, status } where status ∈ ok | ask | abort.
|
|
169
169
|
* Pure: the CLI turns `status` into the exit code, nothing here exits.
|
|
170
170
|
*/
|
|
171
|
+
/** The three verdicts `scope-hammer` may return (shapeup-run.js's HAMMER schema). */
|
|
172
|
+
export const HAMMER_VERDICTS = ["ship-now", "ship-after-fixes", "cannot-ship"];
|
|
173
|
+
|
|
174
|
+
/**
|
|
175
|
+
* The census verdict on disk for one run, or null when no census has run.
|
|
176
|
+
*
|
|
177
|
+
* Derived from the hammer's own WorkResult rather than taken from a caller: the verdict is the
|
|
178
|
+
* product of a dispatch, and a gate that accepts it as an argument accepts whatever the caller
|
|
179
|
+
* believes. `null` is a real answer here — "scope-hammer has not run" — and is the state a run that
|
|
180
|
+
* returned `gate_h` from a circuit breaker is in, because the breaker returns long before the
|
|
181
|
+
* hammer is dispatched.
|
|
182
|
+
*
|
|
183
|
+
* @param {string} cwd - Project root.
|
|
184
|
+
* @param {string} slug - Feature slug.
|
|
185
|
+
* @returns {(string|null)} The verdict, or null when there is no readable census.
|
|
186
|
+
*/
|
|
187
|
+
export function censusVerdict(cwd, slug) {
|
|
188
|
+
if (!slug) return null;
|
|
189
|
+
try {
|
|
190
|
+
const p = join(resultsDir(cwd, slug), "hammer.json");
|
|
191
|
+
if (!existsSync(p)) return null;
|
|
192
|
+
const r = JSON.parse(readFileSync(p, "utf8"));
|
|
193
|
+
const v = r?.payload?.verdict ?? r?.verdict ?? null;
|
|
194
|
+
return HAMMER_VERDICTS.includes(v) ? v : null;
|
|
195
|
+
} catch { return null; }
|
|
196
|
+
}
|
|
197
|
+
|
|
198
|
+
/**
|
|
199
|
+
* Narrow a resolved gate answer to what the run's own evidence supports.
|
|
200
|
+
*
|
|
201
|
+
* THE RULE, and it is about answers rather than callers. An answer set chooses among the answers a
|
|
202
|
+
* gate ALLOWS; it can never supply the evidence that makes one allowed. `ship` at L4 asserts the
|
|
203
|
+
* census cleared the run, so `ship` is only on the menu when a census exists and did clear it.
|
|
204
|
+
*
|
|
205
|
+
* Measured on a consumer: a run whose dispatch receipts carried `orient` and `task-executor` and
|
|
206
|
+
* nothing else — `scope-hammer` never dispatched — recorded `H → accept-cut-list` and `L4 → ship`
|
|
207
|
+
* from the `ci` preset, while its own ledger read `status: escalated`, `final_verdict: ~`, every
|
|
208
|
+
* requirement at `no evidence`. `shapeup-run.js` does guard this, but only after the hammer
|
|
209
|
+
* dispatch, so a run that returns `gate_h` from a breaker reaches L4 by a route the guard does not
|
|
210
|
+
* cover. A check on one call site is not a property.
|
|
211
|
+
*
|
|
212
|
+
* Absence of a census disqualifies `ship` on its own: "no one looked" and "someone looked and it
|
|
213
|
+
* was fine" are different facts, and only the second warrants a ship.
|
|
214
|
+
*
|
|
215
|
+
* @param {object} r - A resolved gate result from {@link resolve}.
|
|
216
|
+
* @param {string} cwd - Project root.
|
|
217
|
+
* @param {(string|null)} slug - Feature slug, when the gate is being resolved inside a run.
|
|
218
|
+
* @returns {object} The result, or an `ask` carrying why `ship` was not available.
|
|
219
|
+
*/
|
|
220
|
+
export function narrowToEvidence(r, cwd, slug) {
|
|
221
|
+
if (r?.gate !== "L4" || r?.status !== "ok" || r?.decision !== "ship") return r;
|
|
222
|
+
const verdict = censusVerdict(cwd, slug);
|
|
223
|
+
if (verdict === "ship-now" || verdict === "ship-after-fixes") return r;
|
|
224
|
+
const why = verdict === "cannot-ship"
|
|
225
|
+
? "scope-hammer's census returned CANNOT SHIP"
|
|
226
|
+
: "scope-hammer has not run, so no census exists";
|
|
227
|
+
return {
|
|
228
|
+
gate: r.gate, status: "ask", source: r.source, decision: "ask", note: r.note,
|
|
229
|
+
refused: "ship", census: verdict,
|
|
230
|
+
reason: `GATE L4 cannot be answered "ship": ${why}. An answer set chooses among the answers a ` +
|
|
231
|
+
`gate allows; it cannot supply the evidence that makes one allowed. Put the block to the ` +
|
|
232
|
+
`PO, or run the census first.`,
|
|
233
|
+
};
|
|
234
|
+
}
|
|
235
|
+
|
|
171
236
|
export function resolve(set, gate, source) {
|
|
172
237
|
if (!GATE_IDS.includes(gate)) {
|
|
173
238
|
return { gate, status: "error", reason: `unknown gate "${gate}" — known: ${GATE_IDS.join(", ")}` };
|
|
@@ -362,7 +427,7 @@ export function cli(rawArgv) {
|
|
|
362
427
|
const gate = args.resolve ?? null;
|
|
363
428
|
if (!gate) die("nothing to do — pass --init, --list, --verify, or --resolve <gate-id>");
|
|
364
429
|
|
|
365
|
-
const r = resolve(found.set, gate, found.source);
|
|
430
|
+
const r = narrowToEvidence(resolve(found.set, gate, found.source), cwd, args.slug ?? null);
|
|
366
431
|
if (r.status === "error") die(r.reason);
|
|
367
432
|
// A gate with no `--slug` (e.g. `--file` used ad hoc, outside any run) has nowhere to file a
|
|
368
433
|
// per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
|
|
@@ -30,7 +30,7 @@
|
|
|
30
30
|
import { existsSync, readdirSync, readFileSync } from "node:fs";
|
|
31
31
|
import { join, resolve } from "node:path";
|
|
32
32
|
import { runArgs } from "../lib/argv.mjs";
|
|
33
|
-
import { dispatchReceipts, legLedger, resultsDir } from "../lib/paths.mjs";
|
|
33
|
+
import { dispatchReceipts, legLedger, resultsDir, readRunId } from "../lib/paths.mjs";
|
|
34
34
|
import { readLegs } from "./leg.mjs";
|
|
35
35
|
import { greenVerdict } from "./t0.mjs";
|
|
36
36
|
|
|
@@ -64,10 +64,26 @@ export function readReceipts(path) {
|
|
|
64
64
|
* @param {object[]} legs - Pre-read `legs.jsonl` rows.
|
|
65
65
|
* @returns {{orderId:string, hasReceipt:boolean, hasResult:boolean, hasLeg:boolean, state:("unattested"|"in-flight"|"spent")}}
|
|
66
66
|
*/
|
|
67
|
-
export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs) {
|
|
67
|
+
export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, legs, runId = null) {
|
|
68
68
|
const orderId = `${slug}/${scopeId}-r${round}-a${attempt}`;
|
|
69
|
-
|
|
70
|
-
|
|
69
|
+
// SCOPED TO THIS RUN, and that qualifier is the whole correction. `receipts/dispatch.jsonl` and
|
|
70
|
+
// `legs.jsonl` are per-slug and append-only, so they accumulate across every run of a pitch,
|
|
71
|
+
// while `order_id` repeats — the run key is the only thing that separates two runs of one
|
|
72
|
+
// feature. Matching on `order_id` alone answered "was this attempt spent?" with a PREVIOUS run's
|
|
73
|
+
// receipt: measured on a consumer, a run that dispatched nothing was told its first attempt was
|
|
74
|
+
// spent, complete with leg and result, by rows two launches old.
|
|
75
|
+
//
|
|
76
|
+
// A row carrying no run key belongs to NO run rather than to this one, and an unresolvable
|
|
77
|
+
// current run (no receipt on disk) matches nothing. Both directions under-count rather than
|
|
78
|
+
// over-count, which is the safe way to be wrong here: an under-count leaves a breaker un-tripped
|
|
79
|
+
// and the round continues, where an over-count stops work that was never done.
|
|
80
|
+
const mine = (r) => r?.order_id === orderId && runId != null && r?.run_id === runId;
|
|
81
|
+
const hasReceipt = receipts.some(mine);
|
|
82
|
+
const hasLeg = legs.some(mine);
|
|
83
|
+
// A WorkResult carries no run key and reaches one only through its `order_id`, which repeats —
|
|
84
|
+
// so this stays a file check and is deliberately NOT sufficient on its own. It can only turn an
|
|
85
|
+
// already run-scoped receipt into `spent`; a result left behind by an earlier run cannot attest
|
|
86
|
+
// an attempt this run never dispatched.
|
|
71
87
|
const hasResult = existsSync(join(resultsDir(cwd, slug), `${scopeId}-r${round}-a${attempt}.json`));
|
|
72
88
|
const state = !hasReceipt ? "unattested" : (hasResult || hasLeg) ? "spent" : "in-flight";
|
|
73
89
|
return { orderId, hasReceipt, hasResult, hasLeg, state };
|
|
@@ -91,10 +107,11 @@ export function attemptEvidence(cwd, slug, scopeId, round, attempt, receipts, le
|
|
|
91
107
|
export function scopeAttempts(cwd, slug, scopeId, round, attemptBudget) {
|
|
92
108
|
const receipts = readReceipts(dispatchReceipts(cwd, slug));
|
|
93
109
|
const legs = readLegs(legLedger(cwd, slug));
|
|
110
|
+
const runId = readRunId(cwd, slug);
|
|
94
111
|
const attempts = [];
|
|
95
112
|
let spent = 0, inFlight = 0, unattested = 0;
|
|
96
113
|
for (let a = 1; a <= attemptBudget; a++) {
|
|
97
|
-
const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs);
|
|
114
|
+
const ev = attemptEvidence(cwd, slug, scopeId, round, a, receipts, legs, runId);
|
|
98
115
|
if (ev.state === "spent") spent++;
|
|
99
116
|
else if (ev.state === "in-flight") inFlight++;
|
|
100
117
|
else unattested++;
|
package/kernel/probe/resume.mjs
CHANGED
|
@@ -56,14 +56,14 @@
|
|
|
56
56
|
// Exit: 0 ok · 2 malformed argv (nothing ran) · 3 the target the operation needs is not on disk ·
|
|
57
57
|
// 6 the required phase's artifact is NOT on disk (the phase did not complete).
|
|
58
58
|
|
|
59
|
-
import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync } from "node:fs";
|
|
59
|
+
import { existsSync, readdirSync, readFileSync, writeFileSync, mkdirSync, rmSync } from "node:fs";
|
|
60
60
|
import { dirname, join, resolve } from "node:path";
|
|
61
61
|
import { runArgs } from "../lib/argv.mjs";
|
|
62
62
|
import { splitFrontmatter, uncoerce } from "../lib/contract.mjs";
|
|
63
63
|
import { globToRegExp } from "../verify/spec.mjs";
|
|
64
64
|
import {
|
|
65
65
|
intake, harnessRun, wiringMap, projectProfile, scopesDir, resultsDir, ordersDir,
|
|
66
|
-
orientDir, activeOrder, usecasesDir, breadboard, receipt, readReceipt, requirements,
|
|
66
|
+
orientDir, activeOrder, activeScope, usecasesDir, breadboard, receipt, readReceipt, requirements,
|
|
67
67
|
exportRunDir,
|
|
68
68
|
} from "../lib/paths.mjs";
|
|
69
69
|
import { evalVerdict } from "./eval.mjs";
|
|
@@ -679,10 +679,46 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
679
679
|
// The Ship phase (`shapeup-run.js`) already exports a shipped run itself, before this call ever
|
|
680
680
|
// runs — Stage 2 adds the endings that wrote nothing, and leaves that path untouched.
|
|
681
681
|
const shouldExport = withExport && status !== "shipped";
|
|
682
|
-
|
|
683
|
-
|
|
684
|
-
|
|
685
|
-
|
|
682
|
+
|
|
683
|
+
/**
|
|
684
|
+
* Everything a close owes the checkout once the ledger line is written: export the run's
|
|
685
|
+
* records, then retire its pointers.
|
|
686
|
+
*
|
|
687
|
+
* The export runs FIRST, but not because it has to: `exportOnClose` is handed the slug and keys
|
|
688
|
+
* its output by the receipt's `run_id`, so it never reads the pointers this retires. (The bare
|
|
689
|
+
* `reduce export` CLI does read `active-scope` — only to work out which run the operator meant
|
|
690
|
+
* when they named none.) Reporting before teardown is a defensive default, not a correctness
|
|
691
|
+
* requirement, and it is recorded as such so the next reader does not defend an ordering that
|
|
692
|
+
* carries nothing. Measured: swapping the two leaves every check green.
|
|
693
|
+
*
|
|
694
|
+
* WHY RETIRE AT ALL. The substrate fence is enforced while an order is compiled and unanswered,
|
|
695
|
+
* and a run that ends any way other than shipping leaves exactly that by construction. Until this
|
|
696
|
+
* ran, the fence outlived the run: after a close that exited 0 and recorded everything, an
|
|
697
|
+
* ordinary write anywhere in the project was still denied, with no dispatch in flight. The
|
|
698
|
+
* operator's obvious remedy did not help either — `init run --force` is documented as "abandon
|
|
699
|
+
* the open run and start over", never as "release a stuck fence".
|
|
700
|
+
*
|
|
701
|
+
* AND RETIRING IS NOT ANSWERING. The pointer says "a run is in flight"; the order's missing
|
|
702
|
+
* result says "nobody came back". Only the first is untrue after a close. The abandoned order
|
|
703
|
+
* stays unanswered, the attempt census still sees nothing spent on it, and `init run --force`
|
|
704
|
+
* remains the one thing that writes a synthetic result — because a close that quietly claimed the
|
|
705
|
+
* work was answered would spend an attempt budget on work nobody did.
|
|
706
|
+
*
|
|
707
|
+
* Best-effort, like the export: a pointer that cannot be removed degrades the close and is
|
|
708
|
+
* reported in its return, but never turns a close into a non-close.
|
|
709
|
+
*/
|
|
710
|
+
const finishClose = (result) => {
|
|
711
|
+
const warning = shouldExport ? exportOnClose(cwd, slug) : null;
|
|
712
|
+
const stuck = [];
|
|
713
|
+
for (const pointer of [activeOrder(cwd), activeScope(cwd)]) {
|
|
714
|
+
try { rmSync(pointer, { force: true }); } catch { /* fall through to the check below */ }
|
|
715
|
+
if (existsSync(pointer)) stuck.push(pointer);
|
|
716
|
+
}
|
|
717
|
+
return {
|
|
718
|
+
...result,
|
|
719
|
+
...(warning ? { export_warning: warning } : {}),
|
|
720
|
+
...(stuck.length ? { pointer_warning: `could not retire ${stuck.join(", ")} — the substrate fence may still deny writes until it is removed by hand` } : {}),
|
|
721
|
+
};
|
|
686
722
|
};
|
|
687
723
|
|
|
688
724
|
// Truncated, not elided: a cause this long has already done its job in the run's own log — the
|
|
@@ -699,7 +735,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
699
735
|
if (priorClosedStatus && priorClosedAt) {
|
|
700
736
|
if (priorClosedStatus === status && normCause === priorCause) {
|
|
701
737
|
// The identical fact, restated — a retried or duplicated call costs nothing.
|
|
702
|
-
return
|
|
738
|
+
return finishClose({ ok: true, path: p, status, closed_at: priorClosedAt, cause: priorCause, decision: "idempotent", reason: `already closed as "${status}" at ${priorClosedAt} — idempotent no-op` });
|
|
703
739
|
}
|
|
704
740
|
if (priorClosedStatus !== status) {
|
|
705
741
|
// A DIFFERENT terminal status over an already-closed run — refused outright, the original
|
|
@@ -724,7 +760,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
724
760
|
if (afterSup.status !== status || !afterSup.closed_at || afterSup.closed_at === "~") {
|
|
725
761
|
return { ok: false, path: p, status, reason: `wrote the superseding close but the ledger reads back status="${afterSup.status}" closed_at="${afterSup.closed_at}" — the write did not take` };
|
|
726
762
|
}
|
|
727
|
-
return
|
|
763
|
+
return finishClose({
|
|
728
764
|
ok: true, path: p, status, closed_at: afterSup.closed_at, cause: afterSup.close_cause ?? null,
|
|
729
765
|
superseded: true, decision: "superseded", prior_cause: priorCause, prior_closed_at: priorClosedAt,
|
|
730
766
|
});
|
|
@@ -739,7 +775,7 @@ export function closeRun(cwd, slug, { status, cause = null, withExport = true }
|
|
|
739
775
|
if (after.status !== status || !after.closed_at || after.closed_at === "~") {
|
|
740
776
|
return { ok: false, path: p, status, reason: `wrote the close but the ledger reads back status="${after.status}" closed_at="${after.closed_at}" — the write did not take` };
|
|
741
777
|
}
|
|
742
|
-
return
|
|
778
|
+
return finishClose({ ok: true, path: p, status, closed_at: after.closed_at, cause: after.close_cause ?? null, decision: "closed" });
|
|
743
779
|
}
|
|
744
780
|
|
|
745
781
|
/**
|
package/kernel/reduce/ship.mjs
CHANGED
|
@@ -78,7 +78,7 @@ export function frontmatter(text) {
|
|
|
78
78
|
*/
|
|
79
79
|
export function boardCensus(cwd, slug) {
|
|
80
80
|
const dir = tasksDir(cwd, slug);
|
|
81
|
-
const out = { total: 0, done: 0, unfinished: [] };
|
|
81
|
+
const out = { total: 0, done: 0, unfinished: [], anchors: {} };
|
|
82
82
|
if (!existsSync(dir)) return out;
|
|
83
83
|
for (const f of readdirSync(dir)) {
|
|
84
84
|
if (!/^TASK-[\w.-]+\.md$/i.test(f)) continue;
|
|
@@ -86,6 +86,12 @@ export function boardCensus(cwd, slug) {
|
|
|
86
86
|
const fm = frontmatter(body);
|
|
87
87
|
const id = fm.id || f.replace(/\.md$/, "");
|
|
88
88
|
out.total++;
|
|
89
|
+
// The COMMITTED anchor for this board id. `use_case_refs` is the tier-direction rule's own
|
|
90
|
+
// sanctioned direction (LOCAL names SHARED), and it is what the frozen report cites instead of
|
|
91
|
+
// the id — boards renumber per machine, use cases do not.
|
|
92
|
+
const ucs = String(fm.use_case_refs ?? "").replace(/^\[|\]$/g, "")
|
|
93
|
+
.split(",").map((x) => x.trim()).filter(Boolean);
|
|
94
|
+
out.anchors[id] = ucs;
|
|
89
95
|
if (fm.status === "done") out.done++;
|
|
90
96
|
else out.unfinished.push(id);
|
|
91
97
|
}
|
|
@@ -93,6 +99,31 @@ export function boardCensus(cwd, slug) {
|
|
|
93
99
|
return out;
|
|
94
100
|
}
|
|
95
101
|
|
|
102
|
+
/**
|
|
103
|
+
* Replace every board id in free prose with its committed anchor.
|
|
104
|
+
*
|
|
105
|
+
* THE WRITE BOUNDARY, not the column, and that is the whole point. Board ids reached the frozen
|
|
106
|
+
* report three different ways — the unfinished-task callout, the covering-AC column's own prefix,
|
|
107
|
+
* and INSIDE acceptance-criterion prose a planner wrote ("given the seeded todos (TASK-006)"). The
|
|
108
|
+
* third is upstream free text, so a fix that only changes what the columns interpolate still
|
|
109
|
+
* commits a file the next run's spec-lint reds. Everything written into the committed report passes
|
|
110
|
+
* through here.
|
|
111
|
+
*
|
|
112
|
+
* An id with no resolvable use case becomes a neutral phrase rather than the id: the report loses a
|
|
113
|
+
* pointer that never resolved off this machine anyway, and keeps the sentence around it.
|
|
114
|
+
*
|
|
115
|
+
* @param {*} text - Any value destined for the committed report.
|
|
116
|
+
* @param {Record<string, string[]>} anchors - Board id → its `use_case_refs`.
|
|
117
|
+
* @returns {string} The text with every `TASK-…` replaced by a stable anchor.
|
|
118
|
+
*/
|
|
119
|
+
export function deboard(text, anchors = {}) {
|
|
120
|
+
return String(text ?? "").replace(/\bTASK-[A-Za-z0-9][\w.-]*/g, (id) => {
|
|
121
|
+
const ucs = anchors[id];
|
|
122
|
+
if (ucs && ucs.length) return ucs.join("/");
|
|
123
|
+
return "a board task";
|
|
124
|
+
});
|
|
125
|
+
}
|
|
126
|
+
|
|
96
127
|
/**
|
|
97
128
|
* Per-scope T0 outcome, reduced from the trial ledger.
|
|
98
129
|
*
|
|
@@ -194,7 +225,10 @@ export function buildReport(facts) {
|
|
|
194
225
|
L.push("");
|
|
195
226
|
|
|
196
227
|
if (board.unfinished.length) {
|
|
197
|
-
|
|
228
|
+
// Anchored, never enumerated by board id: the ids renumber per machine, and a committed file
|
|
229
|
+
// carrying one reds the NEXT run of this pitch at L1b.
|
|
230
|
+
const unfinishedAnchors = [...new Set(board.unfinished.flatMap((id) => board.anchors?.[id] ?? []))];
|
|
231
|
+
L.push(`> **${board.unfinished.length} task(s) did not finish**${unfinishedAnchors.length ? ` — use cases: ${unfinishedAnchors.join(", ")}` : ""}.`,
|
|
198
232
|
"> The verdict above grades what was built, not what was planned.", "");
|
|
199
233
|
}
|
|
200
234
|
|
|
@@ -241,7 +275,9 @@ export function buildReport(facts) {
|
|
|
241
275
|
const cell = (s) => String(s).replace(/\|/g, "\\|");
|
|
242
276
|
L.push("| REQ | source | evidence | covering AC | criterion | T0 |", "|---|---|---|---|---|---|");
|
|
243
277
|
for (const r of requirements.rows) {
|
|
244
|
-
const ac = r.covering_acs.length
|
|
278
|
+
const ac = r.covering_acs.length
|
|
279
|
+
? `${deboard(r.covering_acs[0].task_id, facts.board?.anchors)}: ${deboard(r.covering_acs[0].ac, facts.board?.anchors)}${r.covering_acs.length > 1 ? ` (+${r.covering_acs.length - 1})` : ""}`
|
|
280
|
+
: "—";
|
|
245
281
|
const crit = r.criteria.length ? `${r.criteria[0].criterion}${r.criteria.length > 1 ? ` (+${r.criteria.length - 1})` : ""} → ${r.criteria.map((c) => c.verdict).join(",")}` : "—";
|
|
246
282
|
const t0h = r.t0.length ? r.t0.map((h) => String(h).slice(0, 12)).join(", ") : "—";
|
|
247
283
|
L.push(`| ${r.id} | ${cell(r.source || "—")} | ${r.evidence} | ${cell(ac)} | ${cell(crit)} | ${t0h} |`);
|
|
@@ -1215,7 +1215,7 @@
|
|
|
1215
1215
|
}
|
|
1216
1216
|
},
|
|
1217
1217
|
"CommandResult": {
|
|
1218
|
-
"description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it.",
|
|
1218
|
+
"description": "One executed command's outcome inside a T0 artifact — produced by actually running the command; no agent can fabricate it. It carries the EVIDENCE, not only the score: `exit` maps a command that never started onto the same 1 a real failure returns, so without `error` a refused or timed-out command and a broken build are the same record everywhere downstream, and without the output a verdict asserting `exit 1` is an assertion nobody can check.",
|
|
1219
1219
|
"x-tier": "EMBEDDED",
|
|
1220
1220
|
"type": "object",
|
|
1221
1221
|
"properties": {
|
|
@@ -1223,10 +1223,23 @@
|
|
|
1223
1223
|
"type": "string"
|
|
1224
1224
|
},
|
|
1225
1225
|
"exit": {
|
|
1226
|
-
"type": "integer"
|
|
1226
|
+
"type": "integer",
|
|
1227
|
+
"description": "The process exit code, or 1 when the command never produced one. Not a discriminator on its own — read `error` to tell a crash from a failure."
|
|
1227
1228
|
},
|
|
1228
1229
|
"pass": {
|
|
1229
1230
|
"type": "boolean"
|
|
1231
|
+
},
|
|
1232
|
+
"error": {
|
|
1233
|
+
"type": "string",
|
|
1234
|
+
"description": "Present ONLY when the command did not run to completion — a spawn failure, a maxBuffer overflow, or the 10-minute timeout. Its presence is the fact the ratchet grades as `crash` (tree restored, never counted as a reverted attempt); its absence means the command ran and the exit code is its own."
|
|
1235
|
+
},
|
|
1236
|
+
"stdout_tail": {
|
|
1237
|
+
"type": "string",
|
|
1238
|
+
"description": "The last 4000 characters of stdout, omitted when nothing was printed. Kept for passing commands too: a fixture that exits 0 having run zero tests is the false green this layer exists to catch. A truncated tail carries a leading `[truncated: kept the last N of M characters]` line, so a partial stream never reads as a complete one."
|
|
1239
|
+
},
|
|
1240
|
+
"stderr_tail": {
|
|
1241
|
+
"type": "string",
|
|
1242
|
+
"description": "The last 4000 characters of stderr, same bound and same truncation marker as `stdout_tail`."
|
|
1230
1243
|
}
|
|
1231
1244
|
}
|
|
1232
1245
|
},
|
package/kernel/verify/spec.mjs
CHANGED
|
@@ -235,11 +235,56 @@ export function lintScopes(scopes, repoFiles) {
|
|
|
235
235
|
}
|
|
236
236
|
|
|
237
237
|
/** Text forms worth scanning; anything else in a committed tree is not a reference carrier. */
|
|
238
|
-
const SCANNED = /\.(md|markdown|yml|yaml|json|txt)$/i;
|
|
238
|
+
export const SCANNED = /\.(md|markdown|yml|yaml|json|txt)$/i;
|
|
239
239
|
|
|
240
240
|
/** A machine-local board id. Strict on purpose: a committed tree has no reason to carry one at all. */
|
|
241
241
|
const TASK_ID = /\bTASK-[A-Za-z0-9][\w.-]*/;
|
|
242
242
|
|
|
243
|
+
/**
|
|
244
|
+
* The tier-direction violations one piece of text carries — the whole rule, on a string.
|
|
245
|
+
*
|
|
246
|
+
* EXPORTED BECAUSE IT HAS TWO ENFORCEMENT POINTS NOW, and they must not be allowed to drift.
|
|
247
|
+
* `lintCommittedTier` below walks files at GATE L1b; `hooks/tier-guard.mjs` refuses the same text
|
|
248
|
+
* at the moment a tool writes it, minutes earlier, while the writer still holds the context needed
|
|
249
|
+
* to rephrase. Four producers wrote committed files this rule reds and none of them learned from
|
|
250
|
+
* the lint, because by the time it speaks the dispatch that wrote the line is over. A second
|
|
251
|
+
* enforcement point is only worth having if it enforces the SAME predicate, so both call this.
|
|
252
|
+
*
|
|
253
|
+
* The detail strings are the message the writer reads, so they carry the remedy, not just the
|
|
254
|
+
* verdict: a board id resolves on the machine that wrote it and nowhere else, and a path into the
|
|
255
|
+
* gitignored tier dangles on every clone.
|
|
256
|
+
*
|
|
257
|
+
* @param {string} text - The file body, or the fragment a tool is about to write.
|
|
258
|
+
* @returns {Array<{kind:("board-id"|"local-path"), token:string, line:number, detail:string}>}
|
|
259
|
+
* One entry per offending line and form, in file order; [] when clean.
|
|
260
|
+
*/
|
|
261
|
+
export function tierLeaks(text) {
|
|
262
|
+
// Built from the LOCAL constant, never a literal — the storage roots have exactly one home.
|
|
263
|
+
const esc = LOCAL.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
|
|
264
|
+
// `\S+` where the walk's own rule is `\S`: the same LINES match either way (a path with one
|
|
265
|
+
// non-space character after the slash has at least one), and the longer form yields the token to
|
|
266
|
+
// quote back at the writer. Trailing punctuation the prose wrapped it in is trimmed off the
|
|
267
|
+
// quote only — never off the test.
|
|
268
|
+
const localPath = new RegExp(`${esc}/\\S+`);
|
|
269
|
+
const leaks = [];
|
|
270
|
+
String(text ?? "").split(/\r?\n/).forEach((line, i) => {
|
|
271
|
+
const task = line.match(TASK_ID);
|
|
272
|
+
if (task) {
|
|
273
|
+
leaks.push({ kind: "board-id", token: task[0], line: i + 1, detail:
|
|
274
|
+
`names ${task[0]} — a committed file cannot carry a board id. Boards live in ${LOCAL}/ ` +
|
|
275
|
+
"(gitignored) and renumber on every regeneration, so this resolves on the machine that wrote it " +
|
|
276
|
+
"and nowhere else. Cite the use case or the scope_id, which are stable." });
|
|
277
|
+
}
|
|
278
|
+
const path = line.match(localPath);
|
|
279
|
+
if (path) {
|
|
280
|
+
leaks.push({ kind: "local-path", token: path[0].replace(/[`)\]},.;:'"]+$/, ""), line: i + 1, detail:
|
|
281
|
+
`points into ${LOCAL}/ — a committed file cannot reference the gitignored tier; the path ` +
|
|
282
|
+
"dangles on every other clone. Name the committed artifact, or describe the tier without a path." });
|
|
283
|
+
}
|
|
284
|
+
});
|
|
285
|
+
return leaks;
|
|
286
|
+
}
|
|
287
|
+
|
|
243
288
|
/**
|
|
244
289
|
* Lint the WHOLE committed tree for references into the gitignored tier.
|
|
245
290
|
*
|
|
@@ -265,28 +310,15 @@ const TASK_ID = /\bTASK-[A-Za-z0-9][\w.-]*/;
|
|
|
265
310
|
export function lintCommittedTier({ cwd, slug }) {
|
|
266
311
|
const root = sharedRoot(cwd, slug);
|
|
267
312
|
if (!existsSync(root)) return [];
|
|
268
|
-
// Built from the LOCAL constant, never a literal — the storage roots have exactly one home.
|
|
269
|
-
const localPath = new RegExp(`${LOCAL.replace(/[.*+?^${}()|[\]\\]/g, "\\$&")}/\\S`);
|
|
270
313
|
const findings = [];
|
|
271
314
|
for (const rel of walkFiles(root)) {
|
|
272
315
|
if (!SCANNED.test(rel)) continue;
|
|
273
|
-
let
|
|
274
|
-
try {
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
findings.push({ rule: "TIER-DIRECTION", level: "red", detail:
|
|
280
|
-
`${at} names ${task[0]} — a committed file cannot carry a board id. Boards live in ${LOCAL}/ ` +
|
|
281
|
-
"(gitignored) and renumber on every regeneration, so this resolves on the machine that wrote it " +
|
|
282
|
-
"and nowhere else. Cite the use case or the scope_id, which are stable." });
|
|
283
|
-
}
|
|
284
|
-
if (localPath.test(line)) {
|
|
285
|
-
findings.push({ rule: "TIER-DIRECTION", level: "red", detail:
|
|
286
|
-
`${at} points into ${LOCAL}/ — a committed file cannot reference the gitignored tier; the path ` +
|
|
287
|
-
"dangles on every other clone. Name the committed artifact, or describe the tier without a path." });
|
|
288
|
-
}
|
|
289
|
-
});
|
|
316
|
+
let text;
|
|
317
|
+
try { text = readFileSync(join(root, rel), "utf8"); } catch { continue; }
|
|
318
|
+
for (const leak of tierLeaks(text)) {
|
|
319
|
+
findings.push({ rule: "TIER-DIRECTION", level: "red",
|
|
320
|
+
detail: `${relative(cwd, join(root, rel))}:${leak.line} ${leak.detail}` });
|
|
321
|
+
}
|
|
290
322
|
}
|
|
291
323
|
return findings;
|
|
292
324
|
}
|
package/kernel/verify/t0.mjs
CHANGED
|
@@ -79,6 +79,69 @@ function runCommand(cmd, cwd) {
|
|
|
79
79
|
return { cmd, exit: r.status ?? 1, pass: r.status === 0, stdout, stderr, ...(error ? { error } : {}) };
|
|
80
80
|
}
|
|
81
81
|
|
|
82
|
+
/**
|
|
83
|
+
* How much of each stream the persisted record keeps, per command.
|
|
84
|
+
*
|
|
85
|
+
* A bound rather than the whole stream, because one chatty fixture would otherwise make every
|
|
86
|
+
* reader of the run trace pay for it — and a bound at the END rather than the start, because that
|
|
87
|
+
* is where a stack trace, an assertion diff and a test summary all land.
|
|
88
|
+
*/
|
|
89
|
+
export const EVIDENCE_TAIL_CHARS = 4000;
|
|
90
|
+
|
|
91
|
+
/**
|
|
92
|
+
* The last `limit` characters of a stream, marked when anything was dropped.
|
|
93
|
+
*
|
|
94
|
+
* THE MARKER IS NOT DECORATION. A truncated tail that does not say so is a partial stream that
|
|
95
|
+
* reads as a complete one, which is the same class of defect as the one this whole change fixes:
|
|
96
|
+
* a record that overstates what it actually holds.
|
|
97
|
+
*
|
|
98
|
+
* @param {(string|null|undefined)} text - The captured stream.
|
|
99
|
+
* @param {number} [limit] - Characters to keep.
|
|
100
|
+
* @returns {(string|null)} The kept tail, or null when the stream was empty (the field is then
|
|
101
|
+
* omitted rather than stored as "", so "nothing was printed" stays visibly different from
|
|
102
|
+
* "nothing was kept").
|
|
103
|
+
*/
|
|
104
|
+
export function boundedTail(text, limit = EVIDENCE_TAIL_CHARS) {
|
|
105
|
+
const s = String(text ?? "");
|
|
106
|
+
if (!s) return null;
|
|
107
|
+
if (s.length <= limit) return s;
|
|
108
|
+
return `[truncated: kept the last ${limit} of ${s.length} characters]\n${s.slice(-limit)}`;
|
|
109
|
+
}
|
|
110
|
+
|
|
111
|
+
/**
|
|
112
|
+
* One command's outcome as the verdict artifact stores it — the evidence, not just the score.
|
|
113
|
+
*
|
|
114
|
+
* WHAT THIS FIXES. The artifact used to keep `{cmd, exit, pass}`, and `runCommand` maps a command
|
|
115
|
+
* that never started onto `exit: 1` — the same number a genuine failure returns. So a refused or
|
|
116
|
+
* timed-out command and a broken build were the SAME RECORD everywhere downstream: the digest, the
|
|
117
|
+
* hill, the report, the evaluator's citation and anyone reading the trace back afterwards. The
|
|
118
|
+
* kernel computed the difference (`error`, which the ratchet grades as `crash`) and discarded it
|
|
119
|
+
* one line later.
|
|
120
|
+
*
|
|
121
|
+
* `exit` is deliberately left as it is. Mapping a crash to some other number would change what the
|
|
122
|
+
* ratchet compares and what every existing reader parses; the distinction travels in `error`, which
|
|
123
|
+
* is the field that actually means "this never ran", and which the crash branch already reads.
|
|
124
|
+
*
|
|
125
|
+
* Output is kept for PASSING commands too, and that direction is not an afterthought: a fixture
|
|
126
|
+
* that exits 0 having run zero tests is the false green this evidence layer exists to catch, and
|
|
127
|
+
* its stdout is the only place that shows.
|
|
128
|
+
*
|
|
129
|
+
* @param {({cmd:string, exit:number, pass:boolean, stdout?:string, stderr?:string, error?:string}|null)} r
|
|
130
|
+
* A `runCommand` result, or null when no command was declared.
|
|
131
|
+
* @returns {(object|null)} The record to persist; null passes through unchanged.
|
|
132
|
+
*/
|
|
133
|
+
export function commandEvidence(r) {
|
|
134
|
+
if (!r) return null;
|
|
135
|
+
const stdout = boundedTail(r.stdout);
|
|
136
|
+
const stderr = boundedTail(r.stderr);
|
|
137
|
+
return {
|
|
138
|
+
cmd: r.cmd, exit: r.exit, pass: r.pass,
|
|
139
|
+
...(r.error ? { error: r.error } : {}),
|
|
140
|
+
...(stdout ? { stdout_tail: stdout } : {}),
|
|
141
|
+
...(stderr ? { stderr_tail: stderr } : {}),
|
|
142
|
+
};
|
|
143
|
+
}
|
|
144
|
+
|
|
82
145
|
/**
|
|
83
146
|
* Run every e2e fixture command for a scope.
|
|
84
147
|
* @param {string[]} fixtures - Fixture command lines (null/empty → no commands).
|
|
@@ -513,8 +576,10 @@ export async function cli(rawArgv) {
|
|
|
513
576
|
const { path, sha256: hash, trial } = writeArtifact(outDir, round, attempt, {
|
|
514
577
|
...(runId ? { run_id: runId } : {}),
|
|
515
578
|
scope_id: contract.scope_id,
|
|
516
|
-
|
|
517
|
-
|
|
579
|
+
// The evidence, not just the score — see `commandEvidence` for what the three-field record
|
|
580
|
+
// could not tell apart, and why `exit` still reads the way it always did.
|
|
581
|
+
fixtures: fixtures.results.map((r) => commandEvidence(r)),
|
|
582
|
+
db_probe: commandEvidence(dbProbe),
|
|
518
583
|
seesaw,
|
|
519
584
|
...verdict,
|
|
520
585
|
score: s,
|
package/package.json
CHANGED
|
@@ -281,6 +281,10 @@ Always use wikilinks (double brackets), never relative paths like `../domain-mod
|
|
|
281
281
|
(`"extracted from .shapeup/<slug>/intake.md"`) reds for the same reason a task
|
|
282
282
|
wikilink does: the path dangles on every other clone. Cite the committed pitch or
|
|
283
283
|
shaping doc instead, or describe the run tier without a path.
|
|
284
|
+
- **You will be stopped at the write, not at the gate.** The same rule is enforced as a
|
|
285
|
+
refusal on the Write/Edit itself, quoting the offending token back. There is nothing to
|
|
286
|
+
appeal: rephrase the line and write it again. Waiting for spec-lint to tell you at Board
|
|
287
|
+
Review costs the whole phase, which is what it used to cost every time.
|
|
284
288
|
- `[[tasks/...]]` wikilinks are valid only inside LOCAL documents (task files, the board,
|
|
285
289
|
EVAL reports), where they resolve against the LOCAL root (`.shapeup/<slug>/`);
|
|
286
290
|
every wikilink in a SHARED doc stays `spec_folder`-relative.
|
|
@@ -26,7 +26,8 @@ scales down, its *verification floor* does not.
|
|
|
26
26
|
ingest-result dispatch. Tiny never means "just edit the file inline".
|
|
27
27
|
- **T0 verification.** A tiny change still proves itself by running — never by claim. If there
|
|
28
28
|
is no runnable check at all, that is a fit-check failure, not a reason to skip T0.
|
|
29
|
-
- **The safety
|
|
29
|
+
- **The machine guards — safety spine, substrate sandbox, tier guard.** They do not scale down,
|
|
30
|
+
and a tiny lane is where a committed file is most likely to be written by hand.
|
|
30
31
|
- **The discovery ledger.** `lane: tiny` is recorded, so a later reader knows exactly what was
|
|
31
32
|
NOT checked (no EVAL verdict, no QA charter, no wiring assertion).
|
|
32
33
|
|