shapeup-sdlc 3.6.0 → 3.7.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +1 -1
- package/AGENTS.md +6 -4
- package/README.md +11 -4
- package/SECURITY.md +3 -2
- package/hooks/hooks.json +10 -0
- package/hooks/tier-guard.mjs +162 -0
- package/kernel/compile.mjs +52 -2
- package/kernel/gate.mjs +67 -2
- package/kernel/harness.mjs +8 -2
- package/kernel/probe/attempts.mjs +152 -0
- package/kernel/probe/resume.mjs +183 -21
- package/kernel/reduce/ship.mjs +39 -3
- package/kernel/schemas/domain.schema.json +15 -2
- package/kernel/verify/spec.mjs +52 -20
- package/kernel/verify/t0.mjs +67 -2
- package/package.json +1 -1
- package/skills/ba-pitch-analyzer/SKILL.md +1 -1
- package/skills/ba-pitch-analyzer/references/doc-schemas.md +10 -0
- package/skills/scope-hammer/SKILL.md +10 -2
- package/skills/tech-lead/references/tiny-lane.md +2 -1
- package/skills/tech-lead/workflows/shapeup-run.js +38 -10
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "shapeup-sdlc-plugin",
|
|
3
3
|
"displayName": "ShapeUp SDLC Plugin",
|
|
4
|
-
"version": "3.
|
|
4
|
+
"version": "3.7.1",
|
|
5
5
|
"description": "Shape Up SDLC harness for Claude Code: shaping, intake, orient, scope-mapping, building (T0-verified, sandboxed, scope-contracted), evaluation and QA skills orchestrated by a tech-lead.",
|
|
6
6
|
"author": {
|
|
7
7
|
"name": "Liberty Nguyen",
|
package/AGENTS.md
CHANGED
|
@@ -7,12 +7,14 @@ A three-phase Shape Up loop orchestrated by `/tech-lead`. Invariants live in the
|
|
|
7
7
|
|
|
8
8
|
Skills and commands are named short throughout this file; every one of them resolves only under the plugin's namespace — `Skill(shapeup-sdlc-plugin:orient)`, `/shapeup-sdlc-plugin:ship`. A bare name is not a typo you get a warning for: an unknown skill name is rejected upstream of the hook layer, so the dispatch fires no hook, leaves no decision row, and is answered by improvisation.
|
|
9
9
|
|
|
10
|
-
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, stopping a run with no receipt (the run's first act writes one).
|
|
10
|
+
- Hook-denied: dispatching a worker without a schema-valid WorkOrder, writing outside the scope's substrate, writing a committed file whose text names a run-trace path or a board id, stopping a run with no receipt (the run's first act writes one).
|
|
11
|
+
- **The tier-direction rule is refused at the write, not argued at L1b.** A committed file cannot carry a `.shapeup/` path or a `TASK-NNN` id — boards renumber per machine and the run trace is gitignored, so either one resolves only where it was written. Spec-lint has always red'd this at Board Review; the same rule now also refuses the write itself, naming the token, because four producers hit it and one of them hit it twice in one session — a worker carries no lesson across a dispatch, and by L1b the phase that wrote the line is over. The knowledge base is outside the rule: those files are instructions, not references anyone resolves.
|
|
11
12
|
- Attested, not assumed: a dispatch leaves a receipt naming the skill that ran, and ingest refuses an orchestrated result that has none. A dispatch that fails is answered by the sub-agent improvising the craft, which every other check accepts — so "the artifact exists" is not evidence that the shipped skill produced it. `--no-receipt-check` is the way through when the receipt channel itself fails.
|
|
12
13
|
- GATE L2 is advisory — warns when EVAL runs over unfinished tasks, permits the call (per-machine board, operator asked; ADR-0001) — a signal, not a bug.
|
|
13
14
|
- Sign-off is a file: each gate resolves from the answer set (`ci`/`guarded`/`interactive`) — cross, stop for the PO, or abort; the decision's source is ledgered.
|
|
15
|
+
- **Every ending a run does not come back from records a close** — its status, its cause and its timestamp — derived from the ending itself rather than from a hand-kept pair, so an ending nobody enumerated cannot silently record nothing. A breaker trip that routes to GATE H closes the run as `escalated`, carrying the breaker in its cause; a pause records no close **by design**, because a relaunch resumes it.
|
|
14
16
|
- **Your scope cut decides your concurrency, not the dial.** Two scopes that both declare a write to one path never build at the same time: `shared` substrate is the sanctioned escape from disjointness, and an edit is read-modify-write, so concurrent writers silently drop each other's work. An entry point listed as writable by five scopes therefore builds five scopes one at a time whatever `--parallel-scopes` says. The run states the ceiling its contracts actually permit before dispatching anything, as a BUILD-order line — a peak below that ceiling is a dispatch problem, a ceiling below the dial is a scope-cut problem, and only the second is fixable by re-cutting. The fix is a cut where exactly one scope owns each entry point. Note the ceiling is narrated, not gated: it does not travel in the ⏸ L1b block, so an unattended run carries it only in its log.
|
|
15
|
-
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round.
|
|
17
|
+
- The build+eval loop breaks only three ways ✦: EVAL PASS → QA → Ship; outer `round_budget` exhausted; opt-in `wall_clock_budget_s` tripped (the wall-clock axis event counters miss). Budget trips route to GATE H — ship what's green, never kill the run from outside. A scope exhausting its per-scope `attempt_budget` (T0 attempts) queues a GATE H proposal, never blocks the round. **An attempt is spent only when the attested channels say so** — a dispatch receipt, plus a leg row or a WorkResult. A compiled order and a T0 verdict are both writable by the scope being judged, so neither counts on its own, and an attempt still in flight holds the breaker open rather than tripping it. Opening the next attempt over an unanswered one is refused outright, because grading a tree the previous attempt may still be writing spends a budget on work nobody did.
|
|
16
18
|
|
|
17
19
|
### Phase 1 — Shaping (`/shapeup`)
|
|
18
20
|
1. Set Boundaries → `/shapeup shaping`
|
|
@@ -79,8 +81,8 @@ Everything discovered funnels into `.shapeup/<slug>/discovery/ledger.md` (Orient
|
|
|
79
81
|
through the chained launch, until the run can resume unattended again.
|
|
80
82
|
- Two storage tiers (ADR-0001): COMMITTED `shapeup/<slug>/` (shaping, spec, scopes, wiring-map, project-profile, requirements, hill, `REPORT.md` frozen at L4) vs GITIGNORED `.shapeup/` (board, orders/results, T0/eval/QA artifacts, ledgers, metrics, gate answers, and the run scripts staged for launch).
|
|
81
83
|
- **The run launches from a copy inside your project, and it has to.** The Workflow tool loads a script only from a directory the session may already read; the plugin installs outside your project, so naming the shipped path is refused before the run begins and no permission rule repairs it — the grant authorises the tool, not what it may read. Opening a run therefore re-copies the run scripts to `.shapeup/workflows/` and reports the path the launch names. A run in flight keeps the copy it started with: an upgrade reaches the next run, not the current round.
|
|
82
|
-
- **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright —
|
|
83
|
-
- Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded
|
|
84
|
+
- **A write the substrate does not cover is denied for as long as the dispatch is in flight, and no longer.** A dispatch is live from the moment its order is compiled until a result for *that* dispatch lands, and liveness is read off the run's order set — not off the pointer that names the run, which outlives it. A run closed any other way than shipping — escalated, aborted, killed outright — leaves an order still open by construction, and the pointer is what decides whether that still fences: the fence is enforced only while the pointer exists on disk, so a close that removes it releases the fence whatever the order set says, and restoring the pointer flips it straight back to denying. **Every terminal close retires it** (3.7.1 onward), not only a ship — an escalated or aborted run stops fencing the checkout the moment it closes, and a close that cannot remove a pointer says so in its own return rather than failing. Retiring is not answering, though, and the difference is the whole reason there are still two levers: the pointer says a run is in flight, the missing result says nobody came back, and only the first stops being true at a close. An order the run left unanswered stays genuinely unresolved — un-fenced, but not done — until `init run --force` writes a synthetic result for it, which is also what lets the attempt census keep telling work nobody did apart from work that failed. What the fence actually gates is narrower than "the project," too — it is this assistant's own edit path (a direct file write or edit call) outside the live order's substrate; a shell command, `git`, or any other editor still writes straight through it. Re-dispatch is fenced again, and the committed tier is never a worker's to write either way: those files belong to the orchestrator, whose window is a phase boundary rather than the middle of somebody else's order.
|
|
85
|
+
- Every run has a `run_id` — the receipt mints it, and orders, T0 artifacts, trial rows, agent-call journal rows and hook decisions all carry it. It is the only key that separates two runs of the same feature: everything else (`order_id`, round/attempt) repeats. It is **not** a time boundary — a relaunch resumes the same run and reuses the key, so one `run_id` legitimately spans every launch after a paused gate or a kill, with hours of wall clock between them, and `orders/<id>.json` is rewritten by each. Anything measuring elapsed time reads the append-only records, never the span of a key. SHIP S.7 exports the run's records as fact tables under `.shapeup/exports/<run_id>/` before the run trace is superseded — and so does **every other terminal ending**: a run that aborts, escalates or stops at GATE H exports on its close, because the runs whose records are worth most are the ones that never reach the Ship phase. The export is advisory at every one of them: a failed export degrades the trace and is reported in the run's state warnings, but never turns a close into a non-close. A WorkResult carries no `run_id` and reaches it through `order_id`.
|
|
84
86
|
- Every run projects a **run graph** — `.shapeup/<slug>/graph.jsonl`, append-only, written only by
|
|
85
87
|
`reduce graph`. Two families kept separate: work lineage (Run, Order, Result, Verdict, Trial,
|
|
86
88
|
GateDecision) and domain (Scope, UseCase, Requirement, Seam). It is derived from the artifacts,
|
package/README.md
CHANGED
|
@@ -210,7 +210,7 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
210
210
|
| **Runtime** — the kernel and the run script | When the run goes through the harness. A schema rejection or a non-zero exit stops the step. | A lane that never calls the kernel is never checked — which is why the hooks below cover the doors, not the steps. |
|
|
211
211
|
| **Advisory** — a report section | When somebody reads the artifact. | Silently, if nobody does. It is a cleanup list, never a verdict. |
|
|
212
212
|
|
|
213
|
-
**
|
|
213
|
+
**Five walls.** These are hooks because nothing in the runtime can substitute for them:
|
|
214
214
|
|
|
215
215
|
- `PreToolUse` (`Skill`) — **`hooks/gate-intake.mjs` denies a `tech-lead` dispatch that carries no
|
|
216
216
|
pitch, no spec folder, and no requirement text.** Observed, not theorized: when the requirement
|
|
@@ -227,6 +227,13 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
227
227
|
checked first and outranks everything, across every live contract — including the carve-out that
|
|
228
228
|
otherwise keeps the active feature's own run trace writable, so a path a live order froze stays
|
|
229
229
|
frozen wherever it lives.
|
|
230
|
+
- `PreToolUse` (`Edit|Write|MultiEdit`) — **`hooks/tier-guard.mjs` refuses a committed-tier write
|
|
231
|
+
whose content names a path into the local run trace or a board id.** Same rule spec-lint reds at
|
|
232
|
+
GATE L1b, asked at the moment of writing: four different producers wrote a committed file their
|
|
233
|
+
own lint then rejected, and one of them did it twice in one session, an hour apart, after fixing
|
|
234
|
+
the first occurrence itself. A worker carries no lesson across a dispatch, so the remedy is a
|
|
235
|
+
refusal the writer can act on while it still knows what it meant to say. The knowledge base is
|
|
236
|
+
outside it — those files are instructions, not references a reader has to resolve.
|
|
230
237
|
- `PreToolUse` (`Bash|Read|Write|Edit|MultiEdit`) — **`hooks/safety-spine.mjs` denies destructive
|
|
231
238
|
commands** (`rm -rf` on unrecoverable targets, force-push/push-to-main, `git reset --hard`,
|
|
232
239
|
`DROP TABLE`) and secret-file reads. A machine guard, not a pipeline guard; the escape hatch is
|
|
@@ -265,7 +272,7 @@ the layer that carries it, and the three layers here fail differently:
|
|
|
265
272
|
| `anti-rationalization` (claims the facts contradict) | The ship report's census, derived from the board and the T0 artifacts | The facts are in an artifact a teammate finds on `git pull`, not in a transcript nobody re-reads. |
|
|
266
273
|
| `slop-cleaner` (TODO/`console.log` leftovers) | The ship report's **Leftovers** section | Same scan, same added-lines-only rule; it lands somewhere checkable. |
|
|
267
274
|
|
|
268
|
-
**Nothing load-bearing depends on permission mode.** The
|
|
275
|
+
**Nothing load-bearing depends on permission mode.** The five walls plus the zero-work gate run
|
|
269
276
|
under every mode. The kernel needs a grant to be *invoked* — two Bash lines `npx shapeup-sdlc init`
|
|
270
277
|
writes — but a session that never gets that grant is a session that cannot run the pipeline at all,
|
|
271
278
|
not one that runs it unguarded.
|
|
@@ -338,8 +345,8 @@ kernel/{verify,reduce,probe,init,report}/ # its subcommands, plus compile and
|
|
|
338
345
|
kernel/lib/ # argv (the typed CLI boundary), paths (+ the run key), contract (shape)
|
|
339
346
|
kernel/schemas/ # the envelope port: WorkOrder, WorkResult, domain registry
|
|
340
347
|
commands/*.md # slash commands (/ship + the 9 phase commands)
|
|
341
|
-
hooks/ # hooks.json + the
|
|
342
|
-
# (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
|
|
348
|
+
hooks/ # hooks.json + the five walls: safety-spine, gate-intake, sandbox-guard,
|
|
349
|
+
# tier-guard (PreToolUse) + gate-zerowork (Stop, the one blocking hook)
|
|
343
350
|
# + dispatch-receipt (PostToolUse, denies nothing, attests which skill ran)
|
|
344
351
|
# + lib/decision.mjs (every hook records allow / deny / error)
|
|
345
352
|
oracles/ # the evaluation-contract oracle registry (test · snapshot · http · process)
|
package/SECURITY.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Security
|
|
2
2
|
|
|
3
|
-
This plugin installs **
|
|
4
|
-
`PreToolUse` position, **all
|
|
3
|
+
This plugin installs **eight hook entries (seven Node scripts + one `echo`)**: five in a
|
|
4
|
+
`PreToolUse` position, **all five of which can deny a tool call**, one `PostToolUse` hook that has no
|
|
5
5
|
deny path at all, and one `Stop`-position hook that can block a session from ending. That is the
|
|
6
6
|
product — and it is also exactly the kind of surface a careful reviewer should want spelled out
|
|
7
7
|
before installing. This page is that spelling-out.
|
|
@@ -73,6 +73,7 @@ sitting, and reading them is the recommended review.
|
|
|
73
73
|
| [`gate-intake.mjs`](hooks/gate-intake.mjs) | PreToolUse (`Skill`) | The `tech-lead` dispatch's own arguments | Yes — an orchestrator dispatch carrying no resolvable intake (no pitch, spec, resume or requirement text) | Fails open on `--order` and on any ambiguous arg shape |
|
|
74
74
|
| [`harness verify envelope`](kernel/verify/envelope.mjs) | PreToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the JSON schemas | Yes — a worker dispatch whose order file is missing or schema-invalid | Never gates a dispatch that carries no `--order` (standalone skill use stays free) |
|
|
75
75
|
| [`sandbox-guard.mjs`](hooks/sandbox-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the `substrate` block of every LIVE order — compiled, with no result at least as new as the order's own `compiled_at`, and, for a run-level phase dispatch, not yet superseded by a later phase's order | Yes — any write no live order permits: inside any `frozen`, outside every `allowed`/`shared`, or a `Write` to an `append_only` path. `frozen` is checked FIRST, so it also outranks the run-trace carve-out | No-op unless an order is live, which a finished run no longer is: the pointer names the run, never a dispatch. The active feature's own `.shapeup/<slug>/` run trace is writable EXCEPT where a live order freezes a path inside it. Appends denials to the local pathology log |
|
|
76
|
+
| [`tier-guard.mjs`](hooks/tier-guard.mjs) | PreToolUse (`Edit\|Write\|MultiEdit`) | The target path; the text the call is about to write (`content`, `new_string`, or any edit in a `MultiEdit` batch) | Yes — a write into `shapeup/<slug>/` whose content names a `.shapeup/` path or a `TASK-NNN` board id: the same tier-direction rule spec-lint reds at GATE L1b, refused at the moment of writing, with the offending token quoted back | No-op outside `shapeup/<slug>/`, on file forms the lint does not scan, and on `shapeup/knowledge-base/` (instructions to a worker, not references). Reads nothing off disk and writes only its decision row; it never inspects a file it is not being asked to write |
|
|
76
77
|
| [`dispatch-receipt.mjs`](hooks/dispatch-receipt.mjs) | PostToolUse (`Skill\|Agent`) | The `--order` file named in the dispatch; the tool result's own report of which skill ran | **No — it has no deny path at all.** It records that the shipped skill ran, so `harness reduce ingest` can refuse a result no dispatch produced | Never writes an attestation for a result that does not name a resolved skill; never fails the call it observes (every write is inside `try`/`catch`) |
|
|
77
78
|
| [`gate-zerowork.mjs`](hooks/gate-zerowork.mjs) | Stop | Run receipts on disk; the session transcript; the decision ledger | **Yes — the one blocking hook.** Returns `decision:"block"` when the session dispatched the orchestrator and produced no run receipt | Defers the moment any receipt exists; `stop_hook_active` caps it at one block per stop chain |
|
|
78
79
|
|
package/hooks/hooks.json
CHANGED
|
@@ -50,6 +50,16 @@
|
|
|
50
50
|
"timeout": 10
|
|
51
51
|
}
|
|
52
52
|
]
|
|
53
|
+
},
|
|
54
|
+
{
|
|
55
|
+
"matcher": "Edit|Write|MultiEdit",
|
|
56
|
+
"hooks": [
|
|
57
|
+
{
|
|
58
|
+
"type": "command",
|
|
59
|
+
"command": "node \"${CLAUDE_PLUGIN_ROOT}/hooks/tier-guard.mjs\"",
|
|
60
|
+
"timeout": 10
|
|
61
|
+
}
|
|
62
|
+
]
|
|
53
63
|
}
|
|
54
64
|
],
|
|
55
65
|
"PostToolUse": [
|
|
@@ -0,0 +1,162 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// Tier guard — PreToolUse hook: the committed tier's write boundary.
|
|
3
|
+
//
|
|
4
|
+
// Refuses an Edit/Write/MultiEdit into `shapeup/<slug>/` whose CONTENT carries a reference the
|
|
5
|
+
// committed tier cannot hold — a path into the gitignored run trace, or a machine-local board id.
|
|
6
|
+
// It is the same rule spec-lint's TIER-DIRECTION already enforces, asked at a different moment.
|
|
7
|
+
//
|
|
8
|
+
// WHY A SECOND ENFORCEMENT POINT FOR A RULE THAT WAS NEVER IN DOUBT. Four different producers wrote
|
|
9
|
+
// committed files this harness's own lint then reds: the requirements registry, the ship report, the
|
|
10
|
+
// project profile, the coverage clauses. The lint was right every time and caught every one of them
|
|
11
|
+
// — at GATE L1b, a whole phase after the sentence was written, addressed to a worker that no longer
|
|
12
|
+
// holds the context needed to rephrase it. The run pays for the round trip, and the fix lands in
|
|
13
|
+
// whatever wording the next reader guesses at.
|
|
14
|
+
//
|
|
15
|
+
// TEACHING DOES NOT CLOSE IT, and that is measured rather than assumed. One of the four recurred
|
|
16
|
+
// INSIDE A SINGLE SESSION, an hour apart, on a different line of the same file, after that same
|
|
17
|
+
// author had fixed the first occurrence and watched its own lint go green. A worker carries no
|
|
18
|
+
// lesson across a dispatch. So the remedy has to be a precondition the model cannot talk past, and
|
|
19
|
+
// it has to speak AT THE WRITE — the one moment the writer still knows what it meant to say.
|
|
20
|
+
//
|
|
21
|
+
// THE PREDICATE IS IMPORTED, NEVER RESTATED. `tierLeaks` is the lint's own scanner
|
|
22
|
+
// (kernel/verify/spec.mjs), so this hook cannot become narrower than the rule it fronts — which
|
|
23
|
+
// would let the defect through to L1b exactly as before — nor wider, which is how a guard earns
|
|
24
|
+
// being switched off. The structural suite executes both over one corpus and requires agreement.
|
|
25
|
+
//
|
|
26
|
+
// Design (deliberately conservative, the shape `sandbox-guard.mjs` argues for):
|
|
27
|
+
// • Fail-OPEN on everything it cannot positively prove: not a write tool, no resolvable path, a
|
|
28
|
+
// path outside `shapeup/<slug>/`, a file form the lint does not scan, or a payload carrying no
|
|
29
|
+
// content to read. A guard that blocks legitimate work gets disabled, and a disabled guard
|
|
30
|
+
// enforces nothing.
|
|
31
|
+
// • `shapeup/knowledge-base/` is OUTSIDE by construction, exactly as it is outside the lint's
|
|
32
|
+
// walk: those files are instructions read by a worker at runtime, not references a reader is
|
|
33
|
+
// expected to resolve, and the defect register they hold has every reason to quote a local path.
|
|
34
|
+
// • Fail-CLOSED only on a leak the lint would red, with the file, the line, the offending token,
|
|
35
|
+
// why it is wrong and what to write instead — all in the denial, because a denial the writer
|
|
36
|
+
// cannot act on immediately just becomes the L1b round trip with extra steps.
|
|
37
|
+
//
|
|
38
|
+
// WHAT IT DOES NOT COVER, stated because the gap is the same one the substrate fence has: this is a
|
|
39
|
+
// PreToolUse hook, so it sees this assistant's own edit path and nothing else. A kernel subcommand
|
|
40
|
+
// writing a committed file, a shell heredoc, or any other editor writes straight through it. Those
|
|
41
|
+
// producers are answered where they are written; this closes the channel the four measured
|
|
42
|
+
// recurrences actually came through.
|
|
43
|
+
//
|
|
44
|
+
// Contract: PreToolUse stdin JSON { tool_name, tool_input:{file_path, content | new_string |
|
|
45
|
+
// edits[]}, cwd }. Deny via { hookSpecificOutput: { hookEventName, permissionDecision:"deny",
|
|
46
|
+
// permissionDecisionReason } }.
|
|
47
|
+
|
|
48
|
+
import { resolve, relative, sep } from "node:path";
|
|
49
|
+
import { isMain } from "../kernel/lib/argv.mjs";
|
|
50
|
+
import { SHARED } from "../kernel/lib/paths.mjs";
|
|
51
|
+
import { SCANNED, tierLeaks } from "../kernel/verify/spec.mjs";
|
|
52
|
+
import { runHook, readStdin, settle, projectRoot } from "./lib/decision.mjs";
|
|
53
|
+
|
|
54
|
+
/** Committed-tier trees that hold instructions rather than references — outside the lint's walk. */
|
|
55
|
+
const NOT_A_SLUG = new Set(["knowledge-base"]);
|
|
56
|
+
|
|
57
|
+
/**
|
|
58
|
+
* The (path, content) pairs one write tool call is about to commit to disk.
|
|
59
|
+
*
|
|
60
|
+
* THE THREE WRITE TOOLS SPELL CONTENT THREE WAYS, and a guard that reads only `content` is blind to
|
|
61
|
+
* the Edit that adds the same line — which is the form the measured recurrence took, since the file
|
|
62
|
+
* already existed by then.
|
|
63
|
+
*
|
|
64
|
+
* @param {object} toolInput - The PreToolUse `tool_input` block.
|
|
65
|
+
* @returns {Array<{path:string, content:string}>} Pairs with a usable path and string content.
|
|
66
|
+
*/
|
|
67
|
+
export function extractWrites(toolInput) {
|
|
68
|
+
const out = [];
|
|
69
|
+
const base = toolInput?.file_path;
|
|
70
|
+
if (typeof toolInput?.content === "string" && base) out.push({ path: base, content: toolInput.content });
|
|
71
|
+
if (typeof toolInput?.new_string === "string" && base) out.push({ path: base, content: toolInput.new_string });
|
|
72
|
+
if (Array.isArray(toolInput?.edits)) {
|
|
73
|
+
for (const e of toolInput.edits) {
|
|
74
|
+
const p = e?.file_path || base;
|
|
75
|
+
if (p && typeof e?.new_string === "string") out.push({ path: p, content: e.new_string });
|
|
76
|
+
}
|
|
77
|
+
}
|
|
78
|
+
return out;
|
|
79
|
+
}
|
|
80
|
+
|
|
81
|
+
/**
|
|
82
|
+
* Is this path inside the committed tier the lint walks — `shapeup/<slug>/…`, scanned form?
|
|
83
|
+
*
|
|
84
|
+
* @param {string} rel - Path relative to the project root, as the lint would name it.
|
|
85
|
+
* @returns {boolean} True when a leak here is a leak the lint would red.
|
|
86
|
+
*/
|
|
87
|
+
export function inCommittedTier(rel) {
|
|
88
|
+
if (!rel || rel.startsWith("..") || rel.startsWith(sep)) return false;
|
|
89
|
+
const parts = rel.split(/[\\/]/);
|
|
90
|
+
// `shapeup/<slug>/<file>` — the lint walks a slug's tree, so a file at the tier root belongs to
|
|
91
|
+
// no run and is left alone, and `knowledge-base` is a sibling of the slugs, not one of them.
|
|
92
|
+
if (parts[0] !== SHARED || parts.length < 3 || NOT_A_SLUG.has(parts[1])) return false;
|
|
93
|
+
return SCANNED.test(rel);
|
|
94
|
+
}
|
|
95
|
+
|
|
96
|
+
async function main() {
|
|
97
|
+
await runHook("tier-guard", async () => {
|
|
98
|
+
const raw = await readStdin();
|
|
99
|
+
let p;
|
|
100
|
+
/** Fail-open, with the reason on the record (hooks/lib/decision.mjs). */
|
|
101
|
+
const defer = (reason, rule) => settle({
|
|
102
|
+
verdict: "allow", event: "PreToolUse", tool: p?.tool_name ?? null, cwd: p?.cwd, reason, rule,
|
|
103
|
+
});
|
|
104
|
+
try { p = JSON.parse(raw || "{}"); }
|
|
105
|
+
catch (e) { settle({ verdict: "error", event: "PreToolUse", reason: `unparseable payload: ${e.message}` }); }
|
|
106
|
+
|
|
107
|
+
if (!["Edit", "Write", "MultiEdit"].includes(p.tool_name)) {
|
|
108
|
+
defer(`${p.tool_name ?? "no tool_name"} is not a write tool — out of scope`, "not-write-tool");
|
|
109
|
+
}
|
|
110
|
+
|
|
111
|
+
// The shell's cwd is where the call fired; the tier lives at the project root. Same split the
|
|
112
|
+
// substrate fence has to make, and for the same reason: a worker that `cd`s into a sub-folder
|
|
113
|
+
// must not thereby leave the tier unguarded. Raw tool paths still resolve against the shell.
|
|
114
|
+
const cwd = p.cwd || process.cwd();
|
|
115
|
+
const root = projectRoot(cwd);
|
|
116
|
+
|
|
117
|
+
const writes = extractWrites(p.tool_input);
|
|
118
|
+
if (writes.length === 0) defer("no readable path+content pair in the tool input", "no-content");
|
|
119
|
+
|
|
120
|
+
const inTier = writes
|
|
121
|
+
.map((w) => ({ ...w, rel: relative(root, resolve(cwd, w.path)) }))
|
|
122
|
+
.filter((w) => inCommittedTier(w.rel));
|
|
123
|
+
if (inTier.length === 0) defer(`${writes.length} write(s), none inside ${SHARED}/<slug>/ in a scanned form`, "not-committed-tier");
|
|
124
|
+
|
|
125
|
+
const blocked = [];
|
|
126
|
+
for (const w of inTier) {
|
|
127
|
+
// THE TOKEN IS QUOTED BACK, which the lint's own message does not do and does not need to —
|
|
128
|
+
// it reports against a file on disk the reader can open at that line. Here the line does not
|
|
129
|
+
// exist yet, so "line 1 points into the local tier" leaves the writer hunting through a
|
|
130
|
+
// fragment it is holding in its head. Naming the exact string is the difference between a
|
|
131
|
+
// denial that is acted on and one that is retried verbatim.
|
|
132
|
+
for (const leak of tierLeaks(w.content)) blocked.push(`${w.rel}:${leak.line} — \`${leak.token}\` ${leak.detail}`);
|
|
133
|
+
}
|
|
134
|
+
|
|
135
|
+
if (blocked.length === 0) {
|
|
136
|
+
defer(`${inTier.length} committed-tier write(s) carry no tier-direction leak — permitted`, "tier-clean");
|
|
137
|
+
}
|
|
138
|
+
|
|
139
|
+
return {
|
|
140
|
+
verdict: "deny", event: "PreToolUse", tool: p.tool_name, subject: inTier[0].rel, cwd: root,
|
|
141
|
+
rule: "tier-direction",
|
|
142
|
+
reason: `${blocked.length} tier-direction leak(s) refused at the write boundary: ${blocked.join("; ")}`,
|
|
143
|
+
payload: {
|
|
144
|
+
hookSpecificOutput: {
|
|
145
|
+
hookEventName: "PreToolUse",
|
|
146
|
+
permissionDecision: "deny",
|
|
147
|
+
permissionDecisionReason:
|
|
148
|
+
"Tier guard (TIER-DIRECTION) — this write would put a reference into the committed tier that " +
|
|
149
|
+
"cannot survive the trip to another machine:\n" +
|
|
150
|
+
`${blocked.join("\n")}\n` +
|
|
151
|
+
`Refused here rather than at GATE L1b, where spec-lint reds the same file after the phase is over. ` +
|
|
152
|
+
"Rephrase the line now: cite the committed artifact, the use case or the scope_id, or describe the " +
|
|
153
|
+
"tier without naming a path.",
|
|
154
|
+
},
|
|
155
|
+
},
|
|
156
|
+
};
|
|
157
|
+
});
|
|
158
|
+
}
|
|
159
|
+
|
|
160
|
+
if (isMain(import.meta.url)) {
|
|
161
|
+
main();
|
|
162
|
+
}
|
package/kernel/compile.mjs
CHANGED
|
@@ -31,7 +31,7 @@ import { fileURLToPath } from "node:url";
|
|
|
31
31
|
import { validate } from "./verify/envelope.mjs";
|
|
32
32
|
import { readTrials } from "./verify/t0.mjs";
|
|
33
33
|
import { runArgs } from "./lib/argv.mjs";
|
|
34
|
-
import { readRunId } from "./lib/paths.mjs";
|
|
34
|
+
import { readRunId, dispatchReceipts, legLedger } from "./lib/paths.mjs";
|
|
35
35
|
// `specDir` is aliased: this module has a local `let specDir` holding the resolved, possibly
|
|
36
36
|
// --spec-overridden directory, and the import is the convention-derived default.
|
|
37
37
|
import {
|
|
@@ -41,6 +41,8 @@ import {
|
|
|
41
41
|
import { readContract, readAllContracts, tasksForScope, SCOPE_CONTRACT } from "./lib/contract.mjs";
|
|
42
42
|
import { writeActiveOrder } from "./probe/resume.mjs";
|
|
43
43
|
import { greenVerdict } from "./probe/t0.mjs";
|
|
44
|
+
import { attemptEvidence, readReceipts } from "./probe/attempts.mjs";
|
|
45
|
+
import { readLegs } from "./probe/leg.mjs";
|
|
44
46
|
import { latestRoundBuild } from "./verify/build.mjs";
|
|
45
47
|
// The SAME matcher the sandbox hook enforces with. "Is this cited file inside this scope's
|
|
46
48
|
// substrate" has to mean exactly what the guard means, or a bug is addressed to a scope that is
|
|
@@ -816,6 +818,40 @@ export async function cli(rawArgv) {
|
|
|
816
818
|
const round = flag("round");
|
|
817
819
|
const attempt = flag("attempt");
|
|
818
820
|
|
|
821
|
+
// AN ATTEMPT MAY NOT OPEN OVER AN UNANSWERED ONE — the write half of the attested-channel rule.
|
|
822
|
+
//
|
|
823
|
+
// `--attempt` is a flag: the CALLER picks the number, and nothing here used to check that the
|
|
824
|
+
// previous one had come back. Measured on a consumer run: attempt 2 was compiled and T0-verified
|
|
825
|
+
// while attempt 1 was still in flight, so the round graded the tree attempt 1 was still writing,
|
|
826
|
+
// counted the attempt, and stopped at GATE H with most of its budget and clock unspent. An order
|
|
827
|
+
// and a T0 verdict are both writable by the very scope being judged; only a dispatch receipt, a
|
|
828
|
+
// leg row or a WorkResult attests that work actually happened.
|
|
829
|
+
//
|
|
830
|
+
// FAILS OPEN, NEVER CLOSED, unless the bad state is positively proven. The proof required is the
|
|
831
|
+
// receipts ledger EXISTING while carrying no row for the previous attempt: a lane that does not
|
|
832
|
+
// attest dispatches at all (no ledger on disk) cannot be judged by this rule and is waved
|
|
833
|
+
// through, so `--tiny`, a prose round loop and a standalone build are untouched.
|
|
834
|
+
if (scope?.scope_id && round && attempt > 1) {
|
|
835
|
+
const receiptsPath = dispatchReceipts(cwd, slug);
|
|
836
|
+
if (existsSync(receiptsPath)) {
|
|
837
|
+
const prev = attemptEvidence(
|
|
838
|
+
cwd, slug, scope.scope_id, round, attempt - 1,
|
|
839
|
+
readReceipts(receiptsPath), readLegs(legLedger(cwd, slug)),
|
|
840
|
+
);
|
|
841
|
+
if (prev.state !== "spent") {
|
|
842
|
+
const why = prev.state === "unattested"
|
|
843
|
+
? "no dispatch receipt was ever written for it"
|
|
844
|
+
: "it was dispatched but has neither a leg-completion row nor a WorkResult";
|
|
845
|
+
console.error(
|
|
846
|
+
`compile-order: refusing to open attempt ${attempt} for "${scope.scope_id}" in round ${round} — ` +
|
|
847
|
+
`attempt ${attempt - 1} (${prev.orderId}) is unanswered: ${why}. Grading a tree the previous ` +
|
|
848
|
+
`attempt may still be writing counts an attempt that never ran, and spends a budget on work ` +
|
|
849
|
+
`nobody did. Wait for it to return, or record its outcome, before opening the next one.`);
|
|
850
|
+
process.exit(3);
|
|
851
|
+
}
|
|
852
|
+
}
|
|
853
|
+
}
|
|
854
|
+
|
|
819
855
|
// The fix round's inbound evidence. Derived here, from the ledgered verdict, for every lane —
|
|
820
856
|
// the workflow, `--tiny`, the prose round loop and a standalone `/build` all compile through
|
|
821
857
|
// this line, and none of them can pass a payload to a build order (see the banner above).
|
|
@@ -971,7 +1007,21 @@ export async function cli(rawArgv) {
|
|
|
971
1007
|
// than the flailing it detects. It advises; the orchestrator queues the GATE H proposal.
|
|
972
1008
|
if (scope?.scope_id) {
|
|
973
1009
|
const k = Number(scope.no_progress_k ?? payloadExtra.no_progress_k ?? 2);
|
|
974
|
-
|
|
1010
|
+
// SCOPED TO THIS RUN, for the same reason the attempt census is. `t0/trials.jsonl` is per-slug
|
|
1011
|
+
// and append-only, so a streak survives the run that produced it. Measured on a consumer: two
|
|
1012
|
+
// non-kept trials from earlier runs — one of them graded against an order no worker was ever
|
|
1013
|
+
// dispatched for — read as a stagnation streak that every later run of that scope tripped on,
|
|
1014
|
+
// seven hours after the tree they graded had been fixed. The breaker escalated on evidence that
|
|
1015
|
+
// predated its own fix, and no attempt of the tripping run had been dispatched at all.
|
|
1016
|
+
//
|
|
1017
|
+
// A trial with no run key counts for no run; an unresolvable current run matches nothing. Both
|
|
1018
|
+
// leave the breaker un-tripped, which is the fail-open direction this repo's guards take when
|
|
1019
|
+
// the bad state cannot be positively proven.
|
|
1020
|
+
const myRunId = readRunId(cwd, slug);
|
|
1021
|
+
const st = stagnation(
|
|
1022
|
+
allTrials.filter((t) => t.scope_id === scope.scope_id && myRunId != null && t.run_id === myRunId),
|
|
1023
|
+
k,
|
|
1024
|
+
);
|
|
975
1025
|
if (st.stagnant) {
|
|
976
1026
|
console.error(JSON.stringify({
|
|
977
1027
|
breaker: "stagnation", scope_id: scope.scope_id, streak: st.streak, no_progress_k: st.k,
|
package/kernel/gate.mjs
CHANGED
|
@@ -53,7 +53,7 @@
|
|
|
53
53
|
import { readFileSync, writeFileSync, appendFileSync, existsSync, mkdirSync } from "node:fs";
|
|
54
54
|
import { join, dirname } from "node:path";
|
|
55
55
|
import { runArgs } from "./lib/argv.mjs";
|
|
56
|
-
import { gateAnswerCandidates, gates as gatesPath, LOCAL } from "./lib/paths.mjs";
|
|
56
|
+
import { gateAnswerCandidates, gates as gatesPath, LOCAL, resultsDir } from "./lib/paths.mjs";
|
|
57
57
|
|
|
58
58
|
export const GATE_IDS = ["L0", "L1a", "L1a.5", "L1b", "L2", "L3", "QA", "H", "L4", "COACH-1"];
|
|
59
59
|
|
|
@@ -168,6 +168,71 @@ export function validate(set) {
|
|
|
168
168
|
* Returns { gate, decision, source, note, status } where status ∈ ok | ask | abort.
|
|
169
169
|
* Pure: the CLI turns `status` into the exit code, nothing here exits.
|
|
170
170
|
*/
|
|
171
|
+
/** The three verdicts `scope-hammer` may return (shapeup-run.js's HAMMER schema). */
|
|
172
|
+
export const HAMMER_VERDICTS = ["ship-now", "ship-after-fixes", "cannot-ship"];
|
|
173
|
+
|
|
174
|
+
/**
|
|
175
|
+
* The census verdict on disk for one run, or null when no census has run.
|
|
176
|
+
*
|
|
177
|
+
* Derived from the hammer's own WorkResult rather than taken from a caller: the verdict is the
|
|
178
|
+
* product of a dispatch, and a gate that accepts it as an argument accepts whatever the caller
|
|
179
|
+
* believes. `null` is a real answer here — "scope-hammer has not run" — and is the state a run that
|
|
180
|
+
* returned `gate_h` from a circuit breaker is in, because the breaker returns long before the
|
|
181
|
+
* hammer is dispatched.
|
|
182
|
+
*
|
|
183
|
+
* @param {string} cwd - Project root.
|
|
184
|
+
* @param {string} slug - Feature slug.
|
|
185
|
+
* @returns {(string|null)} The verdict, or null when there is no readable census.
|
|
186
|
+
*/
|
|
187
|
+
export function censusVerdict(cwd, slug) {
|
|
188
|
+
if (!slug) return null;
|
|
189
|
+
try {
|
|
190
|
+
const p = join(resultsDir(cwd, slug), "hammer.json");
|
|
191
|
+
if (!existsSync(p)) return null;
|
|
192
|
+
const r = JSON.parse(readFileSync(p, "utf8"));
|
|
193
|
+
const v = r?.payload?.verdict ?? r?.verdict ?? null;
|
|
194
|
+
return HAMMER_VERDICTS.includes(v) ? v : null;
|
|
195
|
+
} catch { return null; }
|
|
196
|
+
}
|
|
197
|
+
|
|
198
|
+
/**
|
|
199
|
+
* Narrow a resolved gate answer to what the run's own evidence supports.
|
|
200
|
+
*
|
|
201
|
+
* THE RULE, and it is about answers rather than callers. An answer set chooses among the answers a
|
|
202
|
+
* gate ALLOWS; it can never supply the evidence that makes one allowed. `ship` at L4 asserts the
|
|
203
|
+
* census cleared the run, so `ship` is only on the menu when a census exists and did clear it.
|
|
204
|
+
*
|
|
205
|
+
* Measured on a consumer: a run whose dispatch receipts carried `orient` and `task-executor` and
|
|
206
|
+
* nothing else — `scope-hammer` never dispatched — recorded `H → accept-cut-list` and `L4 → ship`
|
|
207
|
+
* from the `ci` preset, while its own ledger read `status: escalated`, `final_verdict: ~`, every
|
|
208
|
+
* requirement at `no evidence`. `shapeup-run.js` does guard this, but only after the hammer
|
|
209
|
+
* dispatch, so a run that returns `gate_h` from a breaker reaches L4 by a route the guard does not
|
|
210
|
+
* cover. A check on one call site is not a property.
|
|
211
|
+
*
|
|
212
|
+
* Absence of a census disqualifies `ship` on its own: "no one looked" and "someone looked and it
|
|
213
|
+
* was fine" are different facts, and only the second warrants a ship.
|
|
214
|
+
*
|
|
215
|
+
* @param {object} r - A resolved gate result from {@link resolve}.
|
|
216
|
+
* @param {string} cwd - Project root.
|
|
217
|
+
* @param {(string|null)} slug - Feature slug, when the gate is being resolved inside a run.
|
|
218
|
+
* @returns {object} The result, or an `ask` carrying why `ship` was not available.
|
|
219
|
+
*/
|
|
220
|
+
export function narrowToEvidence(r, cwd, slug) {
|
|
221
|
+
if (r?.gate !== "L4" || r?.status !== "ok" || r?.decision !== "ship") return r;
|
|
222
|
+
const verdict = censusVerdict(cwd, slug);
|
|
223
|
+
if (verdict === "ship-now" || verdict === "ship-after-fixes") return r;
|
|
224
|
+
const why = verdict === "cannot-ship"
|
|
225
|
+
? "scope-hammer's census returned CANNOT SHIP"
|
|
226
|
+
: "scope-hammer has not run, so no census exists";
|
|
227
|
+
return {
|
|
228
|
+
gate: r.gate, status: "ask", source: r.source, decision: "ask", note: r.note,
|
|
229
|
+
refused: "ship", census: verdict,
|
|
230
|
+
reason: `GATE L4 cannot be answered "ship": ${why}. An answer set chooses among the answers a ` +
|
|
231
|
+
`gate allows; it cannot supply the evidence that makes one allowed. Put the block to the ` +
|
|
232
|
+
`PO, or run the census first.`,
|
|
233
|
+
};
|
|
234
|
+
}
|
|
235
|
+
|
|
171
236
|
export function resolve(set, gate, source) {
|
|
172
237
|
if (!GATE_IDS.includes(gate)) {
|
|
173
238
|
return { gate, status: "error", reason: `unknown gate "${gate}" — known: ${GATE_IDS.join(", ")}` };
|
|
@@ -362,7 +427,7 @@ export function cli(rawArgv) {
|
|
|
362
427
|
const gate = args.resolve ?? null;
|
|
363
428
|
if (!gate) die("nothing to do — pass --init, --list, --verify, or --resolve <gate-id>");
|
|
364
429
|
|
|
365
|
-
const r = resolve(found.set, gate, found.source);
|
|
430
|
+
const r = narrowToEvidence(resolve(found.set, gate, found.source), cwd, args.slug ?? null);
|
|
366
431
|
if (r.status === "error") die(r.reason);
|
|
367
432
|
// A gate with no `--slug` (e.g. `--file` used ad hoc, outside any run) has nowhere to file a
|
|
368
433
|
// per-run ledger row — the same reasoning `resolveRunId` uses for "no run is active": absence is
|
package/kernel/harness.mjs
CHANGED
|
@@ -32,7 +32,7 @@
|
|
|
32
32
|
// gate An answer file with a source, not a vibe.
|
|
33
33
|
// probe resume · t0 · stats · digest · Read-only queries over run state. `concurrency`
|
|
34
34
|
// concurrency · leg · eval · answers how many legs ran at once and what the
|
|
35
|
-
// owner · requirements
|
|
35
|
+
// owner · requirements · attempts
|
|
36
36
|
// fan-out bought, and refuses a figure the record set
|
|
37
37
|
// cannot support rather than printing a plausible one.
|
|
38
38
|
// `leg` answers whether a scope's work reached the
|
|
@@ -50,6 +50,12 @@
|
|
|
50
50
|
// answers which pitch clause a verdict reached, joined
|
|
51
51
|
// through the plan's own covers: edge — the L4 line and
|
|
52
52
|
// GATE H's census cite it for the same reason.
|
|
53
|
+
// `attempts` answers how many of a scope's attempts are
|
|
54
|
+
// ATTESTED (a dispatch receipt AND a leg row or a
|
|
55
|
+
// WorkResult), never the order set or the T0 verdict
|
|
56
|
+
// set alone — the round loop's inner breaker and
|
|
57
|
+
// scope-hammer's census both cite it, so they cannot
|
|
58
|
+
// disagree about the same exhaustion again.
|
|
53
59
|
// init run · fit · run-args Opens a run, or refuses it (exit 3). `run-args`
|
|
54
60
|
// writes GATE L0.9b's launch record and echoes it.
|
|
55
61
|
// report export Projects the run's records as fact tables.
|
|
@@ -86,7 +92,7 @@ export const ROUTES = {
|
|
|
86
92
|
resume: "./probe/resume.mjs", t0: "./probe/t0.mjs", stats: "./probe/stats.mjs",
|
|
87
93
|
digest: "./probe/digest.mjs", concurrency: "./probe/concurrency.mjs",
|
|
88
94
|
leg: "./probe/leg.mjs", eval: "./probe/eval.mjs", owner: "./probe/owner.mjs",
|
|
89
|
-
requirements: "./probe/requirements.mjs",
|
|
95
|
+
requirements: "./probe/requirements.mjs", attempts: "./probe/attempts.mjs",
|
|
90
96
|
},
|
|
91
97
|
init: { run: "./init/run.mjs", fit: "./init/fit.mjs", "run-args": "./init/run-args.mjs" },
|
|
92
98
|
report: { export: "./report/export.mjs", _default: "export" },
|