@hecer/yoke 1.6.0 → 1.6.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.claude-plugin/plugin.json +13 -13
- package/.codex-plugin/plugin.json +7 -7
- package/CHANGELOG.md +294 -288
- package/README.md +874 -874
- package/TODOS.md +5 -5
- package/agents/docs.toml +6 -6
- package/agents/implementer.toml +6 -6
- package/agents/reviewer.toml +6 -6
- package/agents/security.toml +6 -6
- package/bench/README.md +86 -86
- package/bench/RESULTS.md +35 -35
- package/bench/output-compaction.mjs +65 -65
- package/bench/result-schema.mjs +12 -12
- package/bench/results/claude-2026-07-27T18-03-26.json +50 -50
- package/bench/results/codex-unavailable-1785175418318.json +15 -15
- package/bench/results/gemini-2026-07-27T18-03-44.json +46 -46
- package/bench/run-matrix.mjs +26 -26
- package/bench/run.mjs +106 -106
- package/canon/AGENTS.md +30 -30
- package/canon/context/DECISIONS.md +4 -4
- package/canon/context/GLOSSARY.md +11 -11
- package/canon/context/KNOWLEDGE.md +4 -4
- package/canon/context/PROJECT.md +15 -15
- package/canon/loop/loop-spec.md +65 -65
- package/canon/loop/prd.schema.md +43 -43
- package/canon/manifest.yaml +59 -59
- package/canon/policy/gates.md +7 -7
- package/canon/policy/roles.md +9 -9
- package/canon/skills/ATTRIBUTION.md +99 -99
- package/canon/skills/authoring-prd/SKILL.md +58 -58
- package/canon/skills/brainstorming/SKILL.md +164 -164
- package/canon/skills/codebase-design/DEEPENING.md +15 -15
- package/canon/skills/codebase-design/DESIGN-IT-TWICE.md +12 -12
- package/canon/skills/codebase-design/SKILL.md +39 -39
- package/canon/skills/dispatching-parallel-agents/SKILL.md +182 -182
- package/canon/skills/document-release/SKILL.md +302 -302
- package/canon/skills/domain-modeling/ADR-FORMAT.md +19 -19
- package/canon/skills/domain-modeling/CONTEXT-FORMAT.md +39 -39
- package/canon/skills/domain-modeling/SKILL.md +35 -35
- package/canon/skills/executing-plans/SKILL.md +70 -70
- package/canon/skills/finishing-a-development-branch/SKILL.md +200 -200
- package/canon/skills/health/SKILL.md +177 -177
- package/canon/skills/maintaining-context/SKILL.md +34 -34
- package/canon/skills/minimal-code/SKILL.md +21 -21
- package/canon/skills/no-ai-slop/SKILL.md +103 -103
- package/canon/skills/no-ai-slop/eval.md +43 -43
- package/canon/skills/plan-ceo-review/SKILL.md +541 -541
- package/canon/skills/plan-eng-review/SKILL.md +362 -362
- package/canon/skills/receiving-code-review/SKILL.md +213 -213
- package/canon/skills/requesting-code-review/SKILL.md +105 -105
- package/canon/skills/resolving-merge-conflicts/SKILL.md +18 -18
- package/canon/skills/retro/SKILL.md +397 -397
- package/canon/skills/review/SKILL.md +246 -246
- package/canon/skills/ship/SKILL.md +691 -691
- package/canon/skills/subagent-driven-development/SKILL.md +277 -277
- package/canon/skills/systematic-debugging/SKILL.md +296 -296
- package/canon/skills/tdd/SKILL.md +371 -371
- package/canon/skills/unslop-ui/SKILL.md +34 -34
- package/canon/skills/using-git-worktrees/SKILL.md +218 -218
- package/canon/skills/verification-before-completion/SKILL.md +139 -139
- package/canon/skills/visual-verification/SKILL.md +54 -54
- package/canon/skills/workflow/SKILL.md +22 -22
- package/canon/skills/writing-for-agents/SKILL-MECHANICS.md +27 -27
- package/canon/skills/writing-for-agents/SKILL.md +42 -42
- package/canon/skills/writing-plans/SKILL.md +152 -152
- package/canon/skills/writing-skills/SKILL.md +655 -655
- package/canon/skills/yoke-retrofit/SKILL.md +26 -26
- package/canon/skills/yoke-workflow/SKILL.md +20 -20
- package/canon/tools/codex-rtk-hook.mjs +35 -35
- package/canon/tools/graphify.md +3 -3
- package/canon/tools/playwright-mcp.md +3 -3
- package/canon/tools/rtk.md +7 -7
- package/canon/tools/serena.md +6 -6
- package/dist/agents/process.js +3 -0
- package/dist/loop/watchdog.js +1 -1
- package/dist/prd/command.js +17 -17
- package/dist/retrofit/planners/claude.js +14 -14
- package/dist/retrofit/preserve.js +2 -2
- package/docs/MIGRATING-TO-1.0.md +33 -33
- package/docs/MIGRATING-TO-1.1.md +27 -27
- package/docs/MIGRATING-TO-1.4.md +70 -70
- package/docs/PUBLISHING.md +91 -91
- package/docs/superpowers/plans/2026-06-28-baustein-e-context-layer.md +981 -981
- package/docs/superpowers/plans/2026-06-29-baustein-f-routing.md +258 -258
- package/docs/superpowers/plans/2026-06-29-baustein-g-loop-observability.md +1006 -1006
- package/docs/superpowers/plans/2026-06-29-baustein-h-loop-robustness.md +374 -374
- package/docs/superpowers/plans/2026-06-30-baustein-i-visual-design-verification.md +450 -450
- package/docs/superpowers/plans/2026-07-02-baustein-k-zero-to-100-bootstrap.md +1024 -1024
- package/docs/superpowers/plans/2026-07-02-baustein-m-flow-smoke-proofs.md +574 -574
- package/docs/superpowers/plans/2026-08-13-gauntlet-quality-loop.md +537 -537
- package/docs/superpowers/plans/2026-08-16-artifact-backed-output-compaction.md +329 -329
- package/docs/superpowers/specs/2026-06-28-baustein-e-context-layer-design.md +146 -146
- package/docs/superpowers/specs/2026-06-29-baustein-f-routing-design.md +106 -106
- package/docs/superpowers/specs/2026-06-29-baustein-g-loop-observability-design.md +186 -186
- package/docs/superpowers/specs/2026-06-29-baustein-h-loop-robustness-design.md +113 -113
- package/docs/superpowers/specs/2026-06-30-baustein-i-visual-design-verification-design.md +98 -98
- package/docs/superpowers/specs/2026-07-02-baustein-k-zero-to-100-bootstrap-design.md +200 -200
- package/docs/superpowers/specs/2026-07-02-baustein-m-flow-smoke-proofs-design.md +155 -155
- package/docs/superpowers/specs/2026-08-13-gauntlet-quality-loop-design.md +422 -422
- package/docs/superpowers/specs/2026-08-16-artifact-backed-output-compaction-design.md +166 -166
- package/gemini-extension.json +6 -6
- package/hooks/hooks.json +19 -19
- package/package.json +87 -87
|
@@ -1,186 +1,186 @@
|
|
|
1
|
-
# Baustein G — Loop Observability (heartbeat, timeout, live feedback)
|
|
2
|
-
|
|
3
|
-
**Status:** Design approved 2026-06-29
|
|
4
|
-
**Component:** Yoke (🐂)
|
|
5
|
-
**Relates to:** [[harness-loop-technique]], [[harness-build-progress]]
|
|
6
|
-
|
|
7
|
-
## Problem & Goal
|
|
8
|
-
|
|
9
|
-
A real incident: the autonomous loop ran on a downstream project (NewMarket) and sat
|
|
10
|
-
**blocked for ~5 hours, completely invisibly**. A story failed its verify gate, the loop
|
|
11
|
-
correctly returned `blocked` and stopped — but that outcome was printed to the stdout of a
|
|
12
|
-
detached background command and never surfaced. The user saw "running forever," couldn't tell
|
|
13
|
-
progress from a hang, and `yoke loop status` only reported the PRD count (`18/45`), not *why*
|
|
14
|
-
nothing was advancing.
|
|
15
|
-
|
|
16
|
-
Two gaps caused this:
|
|
17
|
-
1. **No durable, queryable run state.** `blocked` (and its reason) lived only in transient
|
|
18
|
-
stdout. There was no artifact to glance at.
|
|
19
|
-
2. **No agent-run timeout.** A genuinely hung nested agent (`claude -p`) would wedge the loop
|
|
20
|
-
indefinitely — only `verify` had a timeout, not the implementation run.
|
|
21
|
-
|
|
22
|
-
**Timeout must not punish slow-but-working agents.** A fixed *total-runtime* timeout cannot tell
|
|
23
|
-
"hung" from "working, just slow" and would kill a legitimately long story. So G uses an
|
|
24
|
-
**inactivity (idle) timeout**, not a wall-clock one: it measures time since the agent's *last
|
|
25
|
-
sign of life* (its last output byte), not total duration. An agent that keeps emitting output
|
|
26
|
-
runs for an unbounded time; only true silence is treated as a hang. The same output stream that
|
|
27
|
-
gives the user live feedback IS the liveness signal — feedback and hang-detection are one.
|
|
28
|
-
|
|
29
|
-
**Goal:** make the loop observable and self-limiting. The loop continuously reports where it
|
|
30
|
-
is (which story, which phase, why blocked) through **token-free, harness-side channels**, and
|
|
31
|
-
a per-iteration timeout breaks true hangs.
|
|
32
|
-
|
|
33
|
-
**Design principle — maximize free feedback.** All feedback added here is produced by the loop
|
|
34
|
-
driver (Node: console + local files), never by an agent. It costs **zero agent tokens**. The
|
|
35
|
-
only token cost in the loop remains the nested agent runs themselves, which G does not change.
|
|
36
|
-
So we deliberately make local feedback dense and frequent.
|
|
37
|
-
|
|
38
|
-
## Key Decisions (locked)
|
|
39
|
-
|
|
40
|
-
| Decision | Choice |
|
|
41
|
-
|---|---|
|
|
42
|
-
| Scope | Heartbeat status file + `yoke loop status` upgrade + agent-run timeout + blocked-leftover reporting + live console narration + append log |
|
|
43
|
-
| Timeout kind | **Idle (no-output) timeout**, NOT total runtime. Resets on every output byte; only true silence counts. Total runtime is unbounded while the agent keeps producing output. |
|
|
44
|
-
| Timeout on expiry | Kill the child → treated as a normal failure → `blocked` → loop stops (consistent with runner-fail / verify-fail / review-reject). No silent skip. |
|
|
45
|
-
| Timeout default | 20 minutes of **silence**; `--timeout=<minutes>` flag > `config.loop.timeoutMinutes` > default. `0` disables. |
|
|
46
|
-
| Timeout mechanism | A small watchdog wrapper (`src/loop/watchdog.ts`) runs the agent as a child, forwards stdio (prompt in, output live out), and kills on idle expiry. Keeps `runLoop`/runner synchronous — no async ripple. |
|
|
47
|
-
| Feedback channels | Console narration + `.yoke/loop-status.json` (current state) + `.yoke/loop.log` (append-only timeline). All Node-side, 0 tokens. |
|
|
48
|
-
| Backwards-compat | No status file → `yoke loop status` behaves exactly as today. Reporter is injectable and defaults to a real impl; existing `runLoop` tests pass a no-op. |
|
|
49
|
-
| Out of scope (YAGNI) | No token metering of nested agents (technically not exposable), no log streaming server, no web dashboard. |
|
|
50
|
-
|
|
51
|
-
## Architecture
|
|
52
|
-
|
|
53
|
-
All new surface lives in the `loop/` subsystem.
|
|
54
|
-
|
|
55
|
-
### 1. `src/loop/reporter.ts` (new) — status types + reporter + reader
|
|
56
|
-
```ts
|
|
57
|
-
type LoopState = 'running' | 'blocked' | 'complete' | 'cap-reached'
|
|
58
|
-
type LoopPhase = 'implementing' | 'verifying' | 'reviewing' | 'committing'
|
|
59
|
-
|
|
60
|
-
interface LoopStatus {
|
|
61
|
-
state: LoopState
|
|
62
|
-
phase?: LoopPhase // only meaningful while state === 'running'
|
|
63
|
-
story?: string
|
|
64
|
-
storyTitle?: string
|
|
65
|
-
reason?: string // populated on 'blocked'
|
|
66
|
-
iteration: number
|
|
67
|
-
progress: { passed: number; total: number }
|
|
68
|
-
startedAt: string // ISO — when the current story/iteration began
|
|
69
|
-
updatedAt: string // ISO — last heartbeat write
|
|
70
|
-
}
|
|
71
|
-
|
|
72
|
-
interface LoopReporter {
|
|
73
|
-
storyStart(story, iteration, progress): void // → state:running, phase:implementing
|
|
74
|
-
phase(phase: LoopPhase): void // → updates phase + updatedAt
|
|
75
|
-
blocked(reason: string): void // → state:blocked
|
|
76
|
-
complete(progress): void // → state:complete
|
|
77
|
-
capReached(progress): void // → state:cap-reached
|
|
78
|
-
}
|
|
79
|
-
|
|
80
|
-
makeReporter(dir, opts?: { quiet?: boolean }, now?: () => Date): LoopReporter
|
|
81
|
-
readStatus(dir): LoopStatus | null
|
|
82
|
-
```
|
|
83
|
-
The default `makeReporter` writes three places on every event:
|
|
84
|
-
- **`.yoke/loop-status.json`** — the single current `LoopStatus` (overwritten atomically: temp file + rename).
|
|
85
|
-
- **`.yoke/loop.log`** — appends one line: `<ISO> <state/phase> <story> <detail>`.
|
|
86
|
-
- **console** — a human line (suppressed when `quiet`): e.g.
|
|
87
|
-
- `▶ S6-seed-rich (19/45) — implementing…`
|
|
88
|
-
- ` ✓ verified` / ` ✓ reviewed` / ` ✓ committed → 20/45`
|
|
89
|
-
- `■ blocked on S6-seed-rich: verify failed (working tree left dirty — clean before restart)`
|
|
90
|
-
|
|
91
|
-
`now` is injectable for deterministic tests. Timestamps use the Node runtime (`Date`),
|
|
92
|
-
which is available (the loop is not a Workflow script).
|
|
93
|
-
|
|
94
|
-
### 2. `src/loop/loop.ts` — drive the reporter
|
|
95
|
-
`runLoop` takes an optional `reporter: LoopReporter` in `LoopOptions` (default: a no-op so the
|
|
96
|
-
existing tests are untouched; `run-command` injects the real one). At each boundary it calls:
|
|
97
|
-
`storyStart` before dispatch → `phase('verifying')` before verify → `phase('reviewing')` before
|
|
98
|
-
review → `phase('committing')` before commit → `complete`/`capReached`/`blocked` at terminal
|
|
99
|
-
returns. The **blocked** call's reason is enriched: if the working tree is dirty
|
|
100
|
-
(`opts.git.isClean(targetDir) === false`) after a block, append
|
|
101
|
-
`" (working tree has uncommitted changes from the blocked story — review/clean before re-running)"`.
|
|
102
|
-
This is the S5 leftover case made explicit.
|
|
103
|
-
|
|
104
|
-
### 3. `src/loop/watchdog.ts` (new) — idle-timeout wrapper
|
|
105
|
-
A tiny standalone CLI: `node watchdog.js --idle-ms=<N> -- <command> [args...]`.
|
|
106
|
-
- Spawns `<command>` with piped stdio. Forwards its own stdin to the child (so the prompt still
|
|
107
|
-
reaches `claude -p`) and the child's stdout/stderr to its own (so live output still reaches
|
|
108
|
-
the user — the feedback channel).
|
|
109
|
-
- Maintains an idle timer reset on **every** stdout/stderr chunk. On expiry it kills the child
|
|
110
|
-
(SIGTERM, then SIGKILL after a short grace; on win32 `taskkill /pid /t /f` for the tree) and
|
|
111
|
-
exits **124** (timeout convention). Otherwise it exits with the child's own exit code.
|
|
112
|
-
- `--idle-ms=0` means no watchdog (spawn-through with no timer).
|
|
113
|
-
All the async/stream complexity lives here, isolated and unit-testable. Nothing else in the
|
|
114
|
-
loop becomes async.
|
|
115
|
-
|
|
116
|
-
### 4. `src/loop/runner.ts` — run the agent through the watchdog
|
|
117
|
-
`Invocation` gains an optional `idleTimeoutMs?: number`. When set (>0), `runCli` runs the agent
|
|
118
|
-
**through the watchdog**: `node <dist>/loop/watchdog.js --idle-ms=<N> -- <command> <args...>`,
|
|
119
|
-
with the prompt still piped via stdin. `execSync` blocks on the watchdog (runner stays
|
|
120
|
-
synchronous). A watchdog exit of 124 (or any non-zero) is caught as today →
|
|
121
|
-
`{ success: false, summary: 'no output for Nm — treated as hung' }` → the loop blocks. When
|
|
122
|
-
`idleTimeoutMs` is unset/0, the agent runs directly as before. `makeRunner`/`makeReviewRunner`
|
|
123
|
-
accept and thread `idleTimeoutMs`; the reviewer run uses the same budget.
|
|
124
|
-
|
|
125
|
-
### 5. `src/loop/run-command.ts` — wire + upgrade `loop status`
|
|
126
|
-
- `runLoopCommand` resolves the idle timeout (`--timeout` flag > `config.loop.timeoutMinutes` >
|
|
127
|
-
default 20 minutes; `0` disables), builds the real `makeReporter(targetDir)`, builds runners
|
|
128
|
-
with the resolved `idleTimeoutMs`, and passes the reporter into `runLoop`.
|
|
129
|
-
- `loopStatus(targetDir)` reads `readStatus(targetDir)`. If present, it renders state + phase +
|
|
130
|
-
story + title + reason + iteration + a relative `updatedAt`, and — when `state==='running'`
|
|
131
|
-
and `updatedAt` is older than the timeout — a `possibly stuck (no update in Nm)` hint. If
|
|
132
|
-
absent, it falls back to today's `enabled + PRD progress` output verbatim.
|
|
133
|
-
|
|
134
|
-
### 6. `src/retrofit/config.ts` — optional config field
|
|
135
|
-
`loop.timeoutMinutes?: number` (idle minutes) added to the schema (optional, backwards-compatible).
|
|
136
|
-
|
|
137
|
-
### 7. `src/cli.ts` — flag + gitignore note
|
|
138
|
-
Parse `--timeout=<minutes>` on `yoke loop run`. `.yoke/loop-status.json` and `.yoke/loop.log`
|
|
139
|
-
are runtime artifacts — ensure the retrofit gitignore covers them (or document that they are
|
|
140
|
-
local-only). They must NOT block the clean-tree gate: the loop writes them under `.yoke/`,
|
|
141
|
-
which retrofit already gitignores for `worktrees`/`backup`; extend that to these two files.
|
|
142
|
-
|
|
143
|
-
## Data flow (one iteration)
|
|
144
|
-
```
|
|
145
|
-
storyStart(S6) ─► [console ▶ / status running:implementing / log]
|
|
146
|
-
runner(idle-watchdog) ─► phase(verifying) ─► verify ─► phase(reviewing) ─► [review]
|
|
147
|
-
─► phase(committing) ─► appendDecision ─► savePrd(passes) ─► commitAll
|
|
148
|
-
─► [console ✓ committed → 20/45]
|
|
149
|
-
on any failure ─► blocked(reason + leftover-hint) ─► [console ■ / status blocked / log] ─► return
|
|
150
|
-
```
|
|
151
|
-
|
|
152
|
-
## Error handling
|
|
153
|
-
- Status writes are best-effort and atomic (temp + rename); a write failure never aborts the
|
|
154
|
-
loop (wrapped, logged to console, swallowed) — observability must not break execution.
|
|
155
|
-
- Timeout kills the child and surfaces as `blocked`; the clean-tree gate + leftover hint guard
|
|
156
|
-
the restart.
|
|
157
|
-
- `readStatus` on a missing/corrupt file returns `null` → `loop status` falls back gracefully.
|
|
158
|
-
|
|
159
|
-
## Testing (subagent-driven TDD, like A–F)
|
|
160
|
-
- **reporter.ts:** status round-trip (write→read); missing file → null; each event sets the
|
|
161
|
-
right state/phase; `loop.log` appends one line per event; `quiet` suppresses console;
|
|
162
|
-
injected `now` makes timestamps deterministic; atomic write leaves no temp file behind.
|
|
163
|
-
- **loop.ts:** reporter receives `storyStart`/`phase`/`complete` in order on success;
|
|
164
|
-
`blocked` with reason on a failed verify; leftover hint appended when `git.isClean` is false
|
|
165
|
-
on block; existing loop tests still pass with the default no-op reporter.
|
|
166
|
-
- **watchdog.ts:** wrapping a child that keeps emitting output past the idle window is **not**
|
|
167
|
-
killed (proves slow-but-working survives); wrapping a silent child that outputs nothing is
|
|
168
|
-
killed after the idle window and exits 124; a fast child passes its exit code through; stdin
|
|
169
|
-
is forwarded to the child; `--idle-ms=0` never kills. (Use tiny `node -e` children with short
|
|
170
|
-
idle windows for deterministic, fast tests.)
|
|
171
|
-
- **runner.ts:** when `idleTimeoutMs>0` the invocation runs through the watchdog wrapper (assert
|
|
172
|
-
the built command); when unset/0 it runs the agent directly; resolution order flag>config>20.
|
|
173
|
-
- **run-command.ts:** `loopStatus` renders state+phase+reason when a status file exists; falls
|
|
174
|
-
back to PRD-only when absent; `--timeout` parsed and forwarded.
|
|
175
|
-
- **config.ts:** `timeoutMinutes` accepted and optional.
|
|
176
|
-
|
|
177
|
-
## What this would have done for the incident
|
|
178
|
-
`yoke loop status .` →
|
|
179
|
-
```
|
|
180
|
-
Loop: BLOCKED on S5-segment-schemas "All 9 segment attribute schemas…"
|
|
181
|
-
phase: verifying · iteration 19 · 18/45 · last update 5h ago
|
|
182
|
-
reason: story did not verify (working tree has uncommitted changes — clean before restart)
|
|
183
|
-
```
|
|
184
|
-
…instead of `18/45`. And a true hang (no agent output at all) would self-terminate after 20
|
|
185
|
-
minutes of silence as `blocked` — while a genuinely slow story that keeps streaming progress
|
|
186
|
-
runs as long as it needs.
|
|
1
|
+
# Baustein G — Loop Observability (heartbeat, timeout, live feedback)
|
|
2
|
+
|
|
3
|
+
**Status:** Design approved 2026-06-29
|
|
4
|
+
**Component:** Yoke (🐂)
|
|
5
|
+
**Relates to:** [[harness-loop-technique]], [[harness-build-progress]]
|
|
6
|
+
|
|
7
|
+
## Problem & Goal
|
|
8
|
+
|
|
9
|
+
A real incident: the autonomous loop ran on a downstream project (NewMarket) and sat
|
|
10
|
+
**blocked for ~5 hours, completely invisibly**. A story failed its verify gate, the loop
|
|
11
|
+
correctly returned `blocked` and stopped — but that outcome was printed to the stdout of a
|
|
12
|
+
detached background command and never surfaced. The user saw "running forever," couldn't tell
|
|
13
|
+
progress from a hang, and `yoke loop status` only reported the PRD count (`18/45`), not *why*
|
|
14
|
+
nothing was advancing.
|
|
15
|
+
|
|
16
|
+
Two gaps caused this:
|
|
17
|
+
1. **No durable, queryable run state.** `blocked` (and its reason) lived only in transient
|
|
18
|
+
stdout. There was no artifact to glance at.
|
|
19
|
+
2. **No agent-run timeout.** A genuinely hung nested agent (`claude -p`) would wedge the loop
|
|
20
|
+
indefinitely — only `verify` had a timeout, not the implementation run.
|
|
21
|
+
|
|
22
|
+
**Timeout must not punish slow-but-working agents.** A fixed *total-runtime* timeout cannot tell
|
|
23
|
+
"hung" from "working, just slow" and would kill a legitimately long story. So G uses an
|
|
24
|
+
**inactivity (idle) timeout**, not a wall-clock one: it measures time since the agent's *last
|
|
25
|
+
sign of life* (its last output byte), not total duration. An agent that keeps emitting output
|
|
26
|
+
runs for an unbounded time; only true silence is treated as a hang. The same output stream that
|
|
27
|
+
gives the user live feedback IS the liveness signal — feedback and hang-detection are one.
|
|
28
|
+
|
|
29
|
+
**Goal:** make the loop observable and self-limiting. The loop continuously reports where it
|
|
30
|
+
is (which story, which phase, why blocked) through **token-free, harness-side channels**, and
|
|
31
|
+
a per-iteration timeout breaks true hangs.
|
|
32
|
+
|
|
33
|
+
**Design principle — maximize free feedback.** All feedback added here is produced by the loop
|
|
34
|
+
driver (Node: console + local files), never by an agent. It costs **zero agent tokens**. The
|
|
35
|
+
only token cost in the loop remains the nested agent runs themselves, which G does not change.
|
|
36
|
+
So we deliberately make local feedback dense and frequent.
|
|
37
|
+
|
|
38
|
+
## Key Decisions (locked)
|
|
39
|
+
|
|
40
|
+
| Decision | Choice |
|
|
41
|
+
|---|---|
|
|
42
|
+
| Scope | Heartbeat status file + `yoke loop status` upgrade + agent-run timeout + blocked-leftover reporting + live console narration + append log |
|
|
43
|
+
| Timeout kind | **Idle (no-output) timeout**, NOT total runtime. Resets on every output byte; only true silence counts. Total runtime is unbounded while the agent keeps producing output. |
|
|
44
|
+
| Timeout on expiry | Kill the child → treated as a normal failure → `blocked` → loop stops (consistent with runner-fail / verify-fail / review-reject). No silent skip. |
|
|
45
|
+
| Timeout default | 20 minutes of **silence**; `--timeout=<minutes>` flag > `config.loop.timeoutMinutes` > default. `0` disables. |
|
|
46
|
+
| Timeout mechanism | A small watchdog wrapper (`src/loop/watchdog.ts`) runs the agent as a child, forwards stdio (prompt in, output live out), and kills on idle expiry. Keeps `runLoop`/runner synchronous — no async ripple. |
|
|
47
|
+
| Feedback channels | Console narration + `.yoke/loop-status.json` (current state) + `.yoke/loop.log` (append-only timeline). All Node-side, 0 tokens. |
|
|
48
|
+
| Backwards-compat | No status file → `yoke loop status` behaves exactly as today. Reporter is injectable and defaults to a real impl; existing `runLoop` tests pass a no-op. |
|
|
49
|
+
| Out of scope (YAGNI) | No token metering of nested agents (technically not exposable), no log streaming server, no web dashboard. |
|
|
50
|
+
|
|
51
|
+
## Architecture
|
|
52
|
+
|
|
53
|
+
All new surface lives in the `loop/` subsystem.
|
|
54
|
+
|
|
55
|
+
### 1. `src/loop/reporter.ts` (new) — status types + reporter + reader
|
|
56
|
+
```ts
|
|
57
|
+
type LoopState = 'running' | 'blocked' | 'complete' | 'cap-reached'
|
|
58
|
+
type LoopPhase = 'implementing' | 'verifying' | 'reviewing' | 'committing'
|
|
59
|
+
|
|
60
|
+
interface LoopStatus {
|
|
61
|
+
state: LoopState
|
|
62
|
+
phase?: LoopPhase // only meaningful while state === 'running'
|
|
63
|
+
story?: string
|
|
64
|
+
storyTitle?: string
|
|
65
|
+
reason?: string // populated on 'blocked'
|
|
66
|
+
iteration: number
|
|
67
|
+
progress: { passed: number; total: number }
|
|
68
|
+
startedAt: string // ISO — when the current story/iteration began
|
|
69
|
+
updatedAt: string // ISO — last heartbeat write
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
interface LoopReporter {
|
|
73
|
+
storyStart(story, iteration, progress): void // → state:running, phase:implementing
|
|
74
|
+
phase(phase: LoopPhase): void // → updates phase + updatedAt
|
|
75
|
+
blocked(reason: string): void // → state:blocked
|
|
76
|
+
complete(progress): void // → state:complete
|
|
77
|
+
capReached(progress): void // → state:cap-reached
|
|
78
|
+
}
|
|
79
|
+
|
|
80
|
+
makeReporter(dir, opts?: { quiet?: boolean }, now?: () => Date): LoopReporter
|
|
81
|
+
readStatus(dir): LoopStatus | null
|
|
82
|
+
```
|
|
83
|
+
The default `makeReporter` writes three places on every event:
|
|
84
|
+
- **`.yoke/loop-status.json`** — the single current `LoopStatus` (overwritten atomically: temp file + rename).
|
|
85
|
+
- **`.yoke/loop.log`** — appends one line: `<ISO> <state/phase> <story> <detail>`.
|
|
86
|
+
- **console** — a human line (suppressed when `quiet`): e.g.
|
|
87
|
+
- `▶ S6-seed-rich (19/45) — implementing…`
|
|
88
|
+
- ` ✓ verified` / ` ✓ reviewed` / ` ✓ committed → 20/45`
|
|
89
|
+
- `■ blocked on S6-seed-rich: verify failed (working tree left dirty — clean before restart)`
|
|
90
|
+
|
|
91
|
+
`now` is injectable for deterministic tests. Timestamps use the Node runtime (`Date`),
|
|
92
|
+
which is available (the loop is not a Workflow script).
|
|
93
|
+
|
|
94
|
+
### 2. `src/loop/loop.ts` — drive the reporter
|
|
95
|
+
`runLoop` takes an optional `reporter: LoopReporter` in `LoopOptions` (default: a no-op so the
|
|
96
|
+
existing tests are untouched; `run-command` injects the real one). At each boundary it calls:
|
|
97
|
+
`storyStart` before dispatch → `phase('verifying')` before verify → `phase('reviewing')` before
|
|
98
|
+
review → `phase('committing')` before commit → `complete`/`capReached`/`blocked` at terminal
|
|
99
|
+
returns. The **blocked** call's reason is enriched: if the working tree is dirty
|
|
100
|
+
(`opts.git.isClean(targetDir) === false`) after a block, append
|
|
101
|
+
`" (working tree has uncommitted changes from the blocked story — review/clean before re-running)"`.
|
|
102
|
+
This is the S5 leftover case made explicit.
|
|
103
|
+
|
|
104
|
+
### 3. `src/loop/watchdog.ts` (new) — idle-timeout wrapper
|
|
105
|
+
A tiny standalone CLI: `node watchdog.js --idle-ms=<N> -- <command> [args...]`.
|
|
106
|
+
- Spawns `<command>` with piped stdio. Forwards its own stdin to the child (so the prompt still
|
|
107
|
+
reaches `claude -p`) and the child's stdout/stderr to its own (so live output still reaches
|
|
108
|
+
the user — the feedback channel).
|
|
109
|
+
- Maintains an idle timer reset on **every** stdout/stderr chunk. On expiry it kills the child
|
|
110
|
+
(SIGTERM, then SIGKILL after a short grace; on win32 `taskkill /pid /t /f` for the tree) and
|
|
111
|
+
exits **124** (timeout convention). Otherwise it exits with the child's own exit code.
|
|
112
|
+
- `--idle-ms=0` means no watchdog (spawn-through with no timer).
|
|
113
|
+
All the async/stream complexity lives here, isolated and unit-testable. Nothing else in the
|
|
114
|
+
loop becomes async.
|
|
115
|
+
|
|
116
|
+
### 4. `src/loop/runner.ts` — run the agent through the watchdog
|
|
117
|
+
`Invocation` gains an optional `idleTimeoutMs?: number`. When set (>0), `runCli` runs the agent
|
|
118
|
+
**through the watchdog**: `node <dist>/loop/watchdog.js --idle-ms=<N> -- <command> <args...>`,
|
|
119
|
+
with the prompt still piped via stdin. `execSync` blocks on the watchdog (runner stays
|
|
120
|
+
synchronous). A watchdog exit of 124 (or any non-zero) is caught as today →
|
|
121
|
+
`{ success: false, summary: 'no output for Nm — treated as hung' }` → the loop blocks. When
|
|
122
|
+
`idleTimeoutMs` is unset/0, the agent runs directly as before. `makeRunner`/`makeReviewRunner`
|
|
123
|
+
accept and thread `idleTimeoutMs`; the reviewer run uses the same budget.
|
|
124
|
+
|
|
125
|
+
### 5. `src/loop/run-command.ts` — wire + upgrade `loop status`
|
|
126
|
+
- `runLoopCommand` resolves the idle timeout (`--timeout` flag > `config.loop.timeoutMinutes` >
|
|
127
|
+
default 20 minutes; `0` disables), builds the real `makeReporter(targetDir)`, builds runners
|
|
128
|
+
with the resolved `idleTimeoutMs`, and passes the reporter into `runLoop`.
|
|
129
|
+
- `loopStatus(targetDir)` reads `readStatus(targetDir)`. If present, it renders state + phase +
|
|
130
|
+
story + title + reason + iteration + a relative `updatedAt`, and — when `state==='running'`
|
|
131
|
+
and `updatedAt` is older than the timeout — a `possibly stuck (no update in Nm)` hint. If
|
|
132
|
+
absent, it falls back to today's `enabled + PRD progress` output verbatim.
|
|
133
|
+
|
|
134
|
+
### 6. `src/retrofit/config.ts` — optional config field
|
|
135
|
+
`loop.timeoutMinutes?: number` (idle minutes) added to the schema (optional, backwards-compatible).
|
|
136
|
+
|
|
137
|
+
### 7. `src/cli.ts` — flag + gitignore note
|
|
138
|
+
Parse `--timeout=<minutes>` on `yoke loop run`. `.yoke/loop-status.json` and `.yoke/loop.log`
|
|
139
|
+
are runtime artifacts — ensure the retrofit gitignore covers them (or document that they are
|
|
140
|
+
local-only). They must NOT block the clean-tree gate: the loop writes them under `.yoke/`,
|
|
141
|
+
which retrofit already gitignores for `worktrees`/`backup`; extend that to these two files.
|
|
142
|
+
|
|
143
|
+
## Data flow (one iteration)
|
|
144
|
+
```
|
|
145
|
+
storyStart(S6) ─► [console ▶ / status running:implementing / log]
|
|
146
|
+
runner(idle-watchdog) ─► phase(verifying) ─► verify ─► phase(reviewing) ─► [review]
|
|
147
|
+
─► phase(committing) ─► appendDecision ─► savePrd(passes) ─► commitAll
|
|
148
|
+
─► [console ✓ committed → 20/45]
|
|
149
|
+
on any failure ─► blocked(reason + leftover-hint) ─► [console ■ / status blocked / log] ─► return
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
## Error handling
|
|
153
|
+
- Status writes are best-effort and atomic (temp + rename); a write failure never aborts the
|
|
154
|
+
loop (wrapped, logged to console, swallowed) — observability must not break execution.
|
|
155
|
+
- Timeout kills the child and surfaces as `blocked`; the clean-tree gate + leftover hint guard
|
|
156
|
+
the restart.
|
|
157
|
+
- `readStatus` on a missing/corrupt file returns `null` → `loop status` falls back gracefully.
|
|
158
|
+
|
|
159
|
+
## Testing (subagent-driven TDD, like A–F)
|
|
160
|
+
- **reporter.ts:** status round-trip (write→read); missing file → null; each event sets the
|
|
161
|
+
right state/phase; `loop.log` appends one line per event; `quiet` suppresses console;
|
|
162
|
+
injected `now` makes timestamps deterministic; atomic write leaves no temp file behind.
|
|
163
|
+
- **loop.ts:** reporter receives `storyStart`/`phase`/`complete` in order on success;
|
|
164
|
+
`blocked` with reason on a failed verify; leftover hint appended when `git.isClean` is false
|
|
165
|
+
on block; existing loop tests still pass with the default no-op reporter.
|
|
166
|
+
- **watchdog.ts:** wrapping a child that keeps emitting output past the idle window is **not**
|
|
167
|
+
killed (proves slow-but-working survives); wrapping a silent child that outputs nothing is
|
|
168
|
+
killed after the idle window and exits 124; a fast child passes its exit code through; stdin
|
|
169
|
+
is forwarded to the child; `--idle-ms=0` never kills. (Use tiny `node -e` children with short
|
|
170
|
+
idle windows for deterministic, fast tests.)
|
|
171
|
+
- **runner.ts:** when `idleTimeoutMs>0` the invocation runs through the watchdog wrapper (assert
|
|
172
|
+
the built command); when unset/0 it runs the agent directly; resolution order flag>config>20.
|
|
173
|
+
- **run-command.ts:** `loopStatus` renders state+phase+reason when a status file exists; falls
|
|
174
|
+
back to PRD-only when absent; `--timeout` parsed and forwarded.
|
|
175
|
+
- **config.ts:** `timeoutMinutes` accepted and optional.
|
|
176
|
+
|
|
177
|
+
## What this would have done for the incident
|
|
178
|
+
`yoke loop status .` →
|
|
179
|
+
```
|
|
180
|
+
Loop: BLOCKED on S5-segment-schemas "All 9 segment attribute schemas…"
|
|
181
|
+
phase: verifying · iteration 19 · 18/45 · last update 5h ago
|
|
182
|
+
reason: story did not verify (working tree has uncommitted changes — clean before restart)
|
|
183
|
+
```
|
|
184
|
+
…instead of `18/45`. And a true hang (no agent output at all) would self-terminate after 20
|
|
185
|
+
minutes of silence as `blocked` — while a genuinely slow story that keeps streaming progress
|
|
186
|
+
runs as long as it needs.
|
|
@@ -1,113 +1,113 @@
|
|
|
1
|
-
# Baustein H — Loop Robustness (verify-as-truth + flaky-tolerant verify)
|
|
2
|
-
|
|
3
|
-
**Status:** Design approved 2026-06-29
|
|
4
|
-
**Component:** Yoke (🐂)
|
|
5
|
-
**Relates to:** [[harness-build-progress]], [[harness-loop-technique]]
|
|
6
|
-
|
|
7
|
-
## Problem & Goal
|
|
8
|
-
|
|
9
|
-
Evidence from a real autonomous run (NewMarket, 52/57 stories): the loop blocked and needed
|
|
10
|
-
human intervention on a handful of stories. Two of those block causes are **harness defects**,
|
|
11
|
-
not project defects, and both are cheap to fix:
|
|
12
|
-
|
|
13
|
-
1. **Runner exit-code is treated as ground truth (SE5).** The story was implemented and the full
|
|
14
|
-
suite was green (769/769), but the `claude` `.cmd` wrapper on Windows exited **127**. The loop
|
|
15
|
-
checks `result.success` *before* running verify, so it **blocked a successful story** on a
|
|
16
|
-
spurious exit code and didn't auto-commit. This contradicts the loop's own stated philosophy
|
|
17
|
-
(Baustein C2): *"independent verification is the source of truth, not the agent's exit code."*
|
|
18
|
-
|
|
19
|
-
2. **No flaky-test tolerance (T4).** Under autonomous load, a heavy background task caused an
|
|
20
|
-
async test to time out — a **false** red, not a real defect. The single-shot verify gate
|
|
21
|
-
blocked the story. The user had to manually add `retry: 2` to the project's `vitest.config`.
|
|
22
|
-
That resilience belongs in the harness.
|
|
23
|
-
|
|
24
|
-
**Goal:** make the loop trust **verify**, not the agent's exit code, and tolerate transient
|
|
25
|
-
flakes — closing two real block causes without weakening the gate.
|
|
26
|
-
|
|
27
|
-
**Out of scope (deferred):** cross-story regression repair (S5/S6 — a later story breaking an
|
|
28
|
-
earlier story's tests). That needs its own design; this spec does not address it.
|
|
29
|
-
|
|
30
|
-
## Key Decisions (locked)
|
|
31
|
-
|
|
32
|
-
| Decision | Choice |
|
|
33
|
-
|---|---|
|
|
34
|
-
| Verify authority | The **implementer** runner's `success` flag is advisory; **verify decides**. Run verify regardless of runner exit, block only on verify failure. |
|
|
35
|
-
| Runner-fail + verify-pass | Proceed (commit) — the exit code was a ghost. Note it via the reporter / summary. |
|
|
36
|
-
| Runner-fail + verify-fail | Block, with a reason naming both signals. |
|
|
37
|
-
| Reviewer runner | **Unchanged** — its non-zero exit *is* the reject verdict (there is no separate verify for the review). |
|
|
38
|
-
| Flaky tolerance | `retryingVerifier(inner, retries)` re-runs a failing verify up to `retries` times; first pass wins. |
|
|
39
|
-
| Retry default | `config.verify.retries` default **1** (one retry). `0` = strict/no-retry. A real failure still fails (twice). |
|
|
40
|
-
| Out of scope | Cross-story regression repair; changing the reviewer's exit-code semantics. |
|
|
41
|
-
|
|
42
|
-
## Architecture
|
|
43
|
-
|
|
44
|
-
### 1. `src/loop/verify.ts` — `retryingVerifier`
|
|
45
|
-
```ts
|
|
46
|
-
export function retryingVerifier(inner: Verifier, retries: number): Verifier {
|
|
47
|
-
return (targetDir) => {
|
|
48
|
-
let last = inner(targetDir)
|
|
49
|
-
let attempt = 0
|
|
50
|
-
while (!last.passed && attempt < retries) {
|
|
51
|
-
attempt++
|
|
52
|
-
last = inner(targetDir)
|
|
53
|
-
}
|
|
54
|
-
if (last.passed && attempt > 0) {
|
|
55
|
-
return { passed: true, summary: `${last.summary} (passed on retry ${attempt})` }
|
|
56
|
-
}
|
|
57
|
-
return attempt > 0 && !last.passed
|
|
58
|
-
? { passed: false, summary: `${last.summary} (still failing after ${attempt} retr${attempt === 1 ? 'y' : 'ies'})` }
|
|
59
|
-
: last
|
|
60
|
-
}
|
|
61
|
-
}
|
|
62
|
-
```
|
|
63
|
-
Pure, injectable `inner` for tests (no real command needed).
|
|
64
|
-
|
|
65
|
-
### 2. `src/loop/loop.ts` — verify is the gate
|
|
66
|
-
In **both** the isolate and non-isolate paths, restructure the implementer step so verify always
|
|
67
|
-
runs:
|
|
68
|
-
|
|
69
|
-
- Run the implementer runner (keep `result` for its summary).
|
|
70
|
-
- `reporter.phase('verifying')`; run `opts.verify(dir)`.
|
|
71
|
-
- If verify **fails**: block. Reason = verify summary, and if the runner *also* reported failure,
|
|
72
|
-
prepend that (`runner reported failure (<summary>); verify also red: <verify summary>`).
|
|
73
|
-
- If verify **passes**: proceed to review/commit even if `result.success` was false. When the
|
|
74
|
-
runner had reported failure, the committed decision/summary notes
|
|
75
|
-
`(runner exited non-zero but verify is green)` so the ghost is auditable.
|
|
76
|
-
|
|
77
|
-
The reviewer step is untouched: `if (!reviewResult.success) → block` stays (exit = verdict).
|
|
78
|
-
|
|
79
|
-
The Baustein-E invariant (no `passes:true` without a commit; decision rollback on commit failure)
|
|
80
|
-
and the leftover-hint are preserved exactly — only the implementer-failure gate moves from
|
|
81
|
-
"before verify" to "verify decides".
|
|
82
|
-
|
|
83
|
-
### 3. `src/retrofit/config.ts` — `verify.retries`
|
|
84
|
-
Extend the `verify` config object: `verify: { command: string; retries?: number }`.
|
|
85
|
-
Backwards-compatible (optional).
|
|
86
|
-
|
|
87
|
-
### 4. `src/loop/run-command.ts` — wire the retry
|
|
88
|
-
When building the verifier from config, wrap it:
|
|
89
|
-
`verify = retryingVerifier(commandVerifier(command), config.verify?.retries ?? 1)`.
|
|
90
|
-
The injected-verifier test path (`opts.verify`) is unchanged.
|
|
91
|
-
|
|
92
|
-
### 5. README + loop-spec
|
|
93
|
-
Document: verify is the source of truth (a spurious agent exit code can't block a green story),
|
|
94
|
-
and `verify.retries` (default 1) for flaky suites. Update `canon/loop/loop-spec.md` step 5.
|
|
95
|
-
|
|
96
|
-
## Testing (subagent-driven TDD)
|
|
97
|
-
- **verify.ts:** `retryingVerifier` passes immediately when inner passes (no retry); passes on
|
|
98
|
-
retry N when inner fails then passes; fails after exhausting retries; `retries: 0` = single
|
|
99
|
-
shot; summary notes the retry/exhaustion. Inner is a stub Verifier.
|
|
100
|
-
- **loop.ts:** runner-fail + verify-pass → story is committed + `passes:true` (the SE5 case);
|
|
101
|
-
runner-fail + verify-fail → blocked with a combined reason; runner-pass + verify-fail →
|
|
102
|
-
blocked (unchanged); the reviewer still blocks on its own failure; existing invariant tests
|
|
103
|
-
(commit-failure revert, leftover hint) stay green; applies to both isolate and non-isolate.
|
|
104
|
-
- **config.ts:** `verify.retries` accepted + optional.
|
|
105
|
-
- **run-command.ts:** the built verifier is wrapped with the resolved retry count (assert via a
|
|
106
|
-
small seam or by config round-trip).
|
|
107
|
-
- Full suite green; `tsc` clean.
|
|
108
|
-
|
|
109
|
-
## What this would have done for the run
|
|
110
|
-
- **SE5:** runner exits 127 but suite is 769/769 green → verify passes → story auto-commits. No
|
|
111
|
-
manual finalize.
|
|
112
|
-
- **T4:** the load-flake fails once, the retry passes → story proceeds. No manual `vitest.config`
|
|
113
|
-
patch, no false block.
|
|
1
|
+
# Baustein H — Loop Robustness (verify-as-truth + flaky-tolerant verify)
|
|
2
|
+
|
|
3
|
+
**Status:** Design approved 2026-06-29
|
|
4
|
+
**Component:** Yoke (🐂)
|
|
5
|
+
**Relates to:** [[harness-build-progress]], [[harness-loop-technique]]
|
|
6
|
+
|
|
7
|
+
## Problem & Goal
|
|
8
|
+
|
|
9
|
+
Evidence from a real autonomous run (NewMarket, 52/57 stories): the loop blocked and needed
|
|
10
|
+
human intervention on a handful of stories. Two of those block causes are **harness defects**,
|
|
11
|
+
not project defects, and both are cheap to fix:
|
|
12
|
+
|
|
13
|
+
1. **Runner exit-code is treated as ground truth (SE5).** The story was implemented and the full
|
|
14
|
+
suite was green (769/769), but the `claude` `.cmd` wrapper on Windows exited **127**. The loop
|
|
15
|
+
checks `result.success` *before* running verify, so it **blocked a successful story** on a
|
|
16
|
+
spurious exit code and didn't auto-commit. This contradicts the loop's own stated philosophy
|
|
17
|
+
(Baustein C2): *"independent verification is the source of truth, not the agent's exit code."*
|
|
18
|
+
|
|
19
|
+
2. **No flaky-test tolerance (T4).** Under autonomous load, a heavy background task caused an
|
|
20
|
+
async test to time out — a **false** red, not a real defect. The single-shot verify gate
|
|
21
|
+
blocked the story. The user had to manually add `retry: 2` to the project's `vitest.config`.
|
|
22
|
+
That resilience belongs in the harness.
|
|
23
|
+
|
|
24
|
+
**Goal:** make the loop trust **verify**, not the agent's exit code, and tolerate transient
|
|
25
|
+
flakes — closing two real block causes without weakening the gate.
|
|
26
|
+
|
|
27
|
+
**Out of scope (deferred):** cross-story regression repair (S5/S6 — a later story breaking an
|
|
28
|
+
earlier story's tests). That needs its own design; this spec does not address it.
|
|
29
|
+
|
|
30
|
+
## Key Decisions (locked)
|
|
31
|
+
|
|
32
|
+
| Decision | Choice |
|
|
33
|
+
|---|---|
|
|
34
|
+
| Verify authority | The **implementer** runner's `success` flag is advisory; **verify decides**. Run verify regardless of runner exit, block only on verify failure. |
|
|
35
|
+
| Runner-fail + verify-pass | Proceed (commit) — the exit code was a ghost. Note it via the reporter / summary. |
|
|
36
|
+
| Runner-fail + verify-fail | Block, with a reason naming both signals. |
|
|
37
|
+
| Reviewer runner | **Unchanged** — its non-zero exit *is* the reject verdict (there is no separate verify for the review). |
|
|
38
|
+
| Flaky tolerance | `retryingVerifier(inner, retries)` re-runs a failing verify up to `retries` times; first pass wins. |
|
|
39
|
+
| Retry default | `config.verify.retries` default **1** (one retry). `0` = strict/no-retry. A real failure still fails (twice). |
|
|
40
|
+
| Out of scope | Cross-story regression repair; changing the reviewer's exit-code semantics. |
|
|
41
|
+
|
|
42
|
+
## Architecture
|
|
43
|
+
|
|
44
|
+
### 1. `src/loop/verify.ts` — `retryingVerifier`
|
|
45
|
+
```ts
|
|
46
|
+
export function retryingVerifier(inner: Verifier, retries: number): Verifier {
|
|
47
|
+
return (targetDir) => {
|
|
48
|
+
let last = inner(targetDir)
|
|
49
|
+
let attempt = 0
|
|
50
|
+
while (!last.passed && attempt < retries) {
|
|
51
|
+
attempt++
|
|
52
|
+
last = inner(targetDir)
|
|
53
|
+
}
|
|
54
|
+
if (last.passed && attempt > 0) {
|
|
55
|
+
return { passed: true, summary: `${last.summary} (passed on retry ${attempt})` }
|
|
56
|
+
}
|
|
57
|
+
return attempt > 0 && !last.passed
|
|
58
|
+
? { passed: false, summary: `${last.summary} (still failing after ${attempt} retr${attempt === 1 ? 'y' : 'ies'})` }
|
|
59
|
+
: last
|
|
60
|
+
}
|
|
61
|
+
}
|
|
62
|
+
```
|
|
63
|
+
Pure, injectable `inner` for tests (no real command needed).
|
|
64
|
+
|
|
65
|
+
### 2. `src/loop/loop.ts` — verify is the gate
|
|
66
|
+
In **both** the isolate and non-isolate paths, restructure the implementer step so verify always
|
|
67
|
+
runs:
|
|
68
|
+
|
|
69
|
+
- Run the implementer runner (keep `result` for its summary).
|
|
70
|
+
- `reporter.phase('verifying')`; run `opts.verify(dir)`.
|
|
71
|
+
- If verify **fails**: block. Reason = verify summary, and if the runner *also* reported failure,
|
|
72
|
+
prepend that (`runner reported failure (<summary>); verify also red: <verify summary>`).
|
|
73
|
+
- If verify **passes**: proceed to review/commit even if `result.success` was false. When the
|
|
74
|
+
runner had reported failure, the committed decision/summary notes
|
|
75
|
+
`(runner exited non-zero but verify is green)` so the ghost is auditable.
|
|
76
|
+
|
|
77
|
+
The reviewer step is untouched: `if (!reviewResult.success) → block` stays (exit = verdict).
|
|
78
|
+
|
|
79
|
+
The Baustein-E invariant (no `passes:true` without a commit; decision rollback on commit failure)
|
|
80
|
+
and the leftover-hint are preserved exactly — only the implementer-failure gate moves from
|
|
81
|
+
"before verify" to "verify decides".
|
|
82
|
+
|
|
83
|
+
### 3. `src/retrofit/config.ts` — `verify.retries`
|
|
84
|
+
Extend the `verify` config object: `verify: { command: string; retries?: number }`.
|
|
85
|
+
Backwards-compatible (optional).
|
|
86
|
+
|
|
87
|
+
### 4. `src/loop/run-command.ts` — wire the retry
|
|
88
|
+
When building the verifier from config, wrap it:
|
|
89
|
+
`verify = retryingVerifier(commandVerifier(command), config.verify?.retries ?? 1)`.
|
|
90
|
+
The injected-verifier test path (`opts.verify`) is unchanged.
|
|
91
|
+
|
|
92
|
+
### 5. README + loop-spec
|
|
93
|
+
Document: verify is the source of truth (a spurious agent exit code can't block a green story),
|
|
94
|
+
and `verify.retries` (default 1) for flaky suites. Update `canon/loop/loop-spec.md` step 5.
|
|
95
|
+
|
|
96
|
+
## Testing (subagent-driven TDD)
|
|
97
|
+
- **verify.ts:** `retryingVerifier` passes immediately when inner passes (no retry); passes on
|
|
98
|
+
retry N when inner fails then passes; fails after exhausting retries; `retries: 0` = single
|
|
99
|
+
shot; summary notes the retry/exhaustion. Inner is a stub Verifier.
|
|
100
|
+
- **loop.ts:** runner-fail + verify-pass → story is committed + `passes:true` (the SE5 case);
|
|
101
|
+
runner-fail + verify-fail → blocked with a combined reason; runner-pass + verify-fail →
|
|
102
|
+
blocked (unchanged); the reviewer still blocks on its own failure; existing invariant tests
|
|
103
|
+
(commit-failure revert, leftover hint) stay green; applies to both isolate and non-isolate.
|
|
104
|
+
- **config.ts:** `verify.retries` accepted + optional.
|
|
105
|
+
- **run-command.ts:** the built verifier is wrapped with the resolved retry count (assert via a
|
|
106
|
+
small seam or by config round-trip).
|
|
107
|
+
- Full suite green; `tsc` clean.
|
|
108
|
+
|
|
109
|
+
## What this would have done for the run
|
|
110
|
+
- **SE5:** runner exits 127 but suite is 769/769 green → verify passes → story auto-commits. No
|
|
111
|
+
manual finalize.
|
|
112
|
+
- **T4:** the load-flake fails once, the retry passes → story proceeds. No manual `vitest.config`
|
|
113
|
+
patch, no false block.
|