@ngockhoale/ukit 2.0.6 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -22,6 +22,66 @@ $ARGUMENTS
22
22
 
23
23
  ---
24
24
 
25
+ ## Autonomy Contract — read before anything else
26
+
27
+ This command is a **one-shot pipeline**. The human is not watching. Treat every phase below
28
+ as unattended: they may walk away after submitting `$ARGUMENTS` and come back to a finished
29
+ result.
30
+
31
+ **The single question window is P0, before any file is written.** If `$ARGUMENTS` leaves a
32
+ choice that would change what gets built — not how it gets built — batch every such question
33
+ into **one** `AskUserQuestion` call (max 4 questions, each with a `(Recommended)` first
34
+ option) and resolve them all at once. Record the answers in `PLAN.md` §1.
35
+
36
+ **After P0 closes, do not ask the human anything until the Final Report.** Every decision
37
+ point in the phases below has a defined automatic resolution. Take it. Specifically:
38
+
39
+ | Situation | Old behavior | Required behavior now |
40
+ |-----------|--------------|----------------------|
41
+ | Scope spans multiple subsystems | wait for confirmation | decompose into sequential cycles, run cycle 1 now, queue the rest in `INDEX.md` |
42
+ | Plan review returns `Issues Found` | up to 3 rounds, then stop | 1 revision round; if round 2 still has findings, apply them and proceed — log to `## Plan Review Log` |
43
+ | Two same-wave tasks share a file | STOP | move the later task to the next wave (add the dependency), continue |
44
+ | INDEX has `pending_review`/`blocked` tasks | STOP | **resume** — see Resume below |
45
+ | Reviewer returns `changes_requested`/`critical_block` | report and stop | run the Phase 4 auto-fix loop (max 2 rounds) |
46
+ | Verification command fails | stop | fix and re-verify inside the auto-fix loop |
47
+
48
+ Escalate to the human **only** for: a blocker outside the repo (missing credential, external
49
+ service down), or a task still failing after both auto-fix rounds. Even then, finish every
50
+ other task first and report what was left undone.
51
+
52
+ **Never pause to ask for `/compact`.** Context is managed structurally — subagents write full
53
+ logs to disk and return short summaries (see 3b), and the run cursor makes the pipeline
54
+ resumable. If compaction does happen, the run resumes from the cursor automatically.
55
+
56
+ ### Run cursor — write after every step
57
+
58
+ Maintain `docs/AI_HANDOFF/RUN.md`. Rewrite it (whole file, 6 lines) immediately after every
59
+ numbered step completes:
60
+
61
+ ```
62
+ Command: handoff-fullstack
63
+ Goal: <one sentence from $ARGUMENTS>
64
+ Base: <BASE branch>
65
+ Phase: <P1|P2|P2.5|P3|I1|I2|I3|I4|R1|R2|R3|R4|R5|done>
66
+ Cursor: wave <N> batch <M> — <what just finished>
67
+ Next: <the exact next step to run>
68
+ ```
69
+
70
+ This file is the resume contract. It costs one small write per step and is what turns an
71
+ interrupted run into a continuable one.
72
+
73
+ ### Resume — when the run is re-invoked mid-flight
74
+
75
+ If `docs/AI_HANDOFF/RUN.md` exists with `Phase:` not `done`, this is a **continuation, not a
76
+ new cycle**. Do NOT re-plan, do NOT overwrite `PLAN.md`, do NOT ask the human whether to
77
+ continue. Read the cursor, then jump straight to `Next:` and carry on. Tasks already `done`
78
+ or `pending_review` are skipped; only `ready`, `in_progress`, `changes_requested` and
79
+ `blocked` tasks are picked up.
80
+
81
+ Only when `RUN.md` is absent or `Phase: done` does `$ARGUMENTS` start a fresh cycle.
82
+
83
+ ---
84
+
25
85
  ## Phase 1+2 — Plan (strong model)
26
86
 
27
87
  ### P1 — Read context (lite model)
@@ -45,7 +105,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
45
105
 
46
106
  1. **Guard — check INDEX statuses:**
47
107
  - All tasks are `ready` (planning only) → re-run is allowed: overwrite `PLAN.md` and `TASK-xxx.md` freely (iterative refinement).
48
- - Any task has status `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **STOP**: warn the human — tasks are already in implement/review phase, cannot overwrite safely.
108
+ - Any task has status `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **do not overwrite, and do not stop either.** Mid-flight work exists. Skip the rest of P2 entirely, write the run cursor with `Phase: I1`, and hand control back to the orchestrator to resume from Phase 3 with the surviving tasks (see Resume in the Autonomy Contract). Overwriting is what is unsafe here — stopping is not required.
49
109
  - No tasks → fresh cycle, proceed.
50
110
 
51
111
  2. **Resolve base branch:**
@@ -57,7 +117,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
57
117
  - §1 Intent — problem + success definition
58
118
  - §2 Scope — in / out of scope; same-wave tasks must not modify the same file (prevents merge conflicts)
59
119
  - §3 Approach — solution, trade-offs, alternatives rejected
60
- - §4 Test Plan — happy path + ≥1 edge case + regression if bugfix (non-negotiable)
120
+ - §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
61
121
  - §5 Verification — exact shell commands executor will run
62
122
  - §6 Acceptance — done checklist (prefer verifiable criteria with commands)
63
123
 
@@ -71,7 +131,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
71
131
  Every task MUST have:
72
132
  - Target Files (exact paths — no two tasks in same wave share a file)
73
133
  - Dependencies (`TASK-xxx` or `none` — wave structure inferred from this)
74
- - Test Cases (Type | Name | Expected — ≥1 happy + ≥1 edge case)
134
+ - Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
75
135
  - Test Files (exact paths)
76
136
  - Verification Commands (runnable shell commands)
77
137
  - Acceptance Criteria (verifiable checklist)
@@ -89,6 +149,25 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
89
149
 
90
150
  7. **Report:** task IDs, dependency graph, any `needs_breakdown` tasks + reason.
91
151
 
152
+ ### P2.5 — Independent plan review (strong model, separate agent)
153
+
154
+ **Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet).
155
+
156
+ - count = 0 → run the review below.
157
+ - count = 1 and the round returned `Issues Found` → planner revises once, then run **one** more review round.
158
+ - count ≥ 2 → do NOT invoke the reviewer again. **Do not stop and do not ask the human.** Have the planner apply every outstanding finding directly to `PLAN.md` and the affected `TASK-xxx.md` files, append `### Round <N> — findings applied without re-review` to the Plan Review Log listing what was changed, then proceed to P3.
159
+
160
+ The gate still does its job — two independent opus passes shape the plan before a line of code
161
+ is written. What it no longer does is hand a stalled plan back to a human who isn't there.
162
+
163
+ **Claude Code — MANDATORY, do this before anything else in P2.5:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from P2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
164
+
165
+ 1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
166
+ 2. `Issues Found` → route back to P2: `handoff-planner` revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for one more P2.5 review round — subject to the loop cap above.
167
+ 3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed to P3.
168
+
169
+ > Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from P2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
170
+
92
171
  ### P3 — Commit the plan (lite model)
93
172
 
94
173
  **Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for this commit step (lite tier — haiku/unic-lite). Run:
@@ -117,7 +196,11 @@ Verify working tree is clean after the plan commit:
117
196
  git status # must be clean — plan commit already done in P3
118
197
  ```
119
198
 
120
- If working tree is dirty → stop. Ask human to resolve uncommitted changes first.
199
+ If the working tree is dirty, do not stop. Commit the stragglers as a checkpoint so they stay
200
+ recoverable and the wave copy-back starts from a known state:
201
+ ```bash
202
+ git add -A && git commit -m "handoff: checkpoint before implement"
203
+ ```
121
204
 
122
205
  ### I2 — Infer wave groups
123
206
 
@@ -127,8 +210,20 @@ Read each `docs/AI_HANDOFF/tasks/TASK-xxx.md` for `Dependencies` field:
127
210
  - Chain A→B→C = 3 waves of 1 task each (sequential)
128
211
  - Independent A, B, C = 1 wave of 3 tasks (parallel)
129
212
 
213
+ **Conflict check — mandatory, before spawning any wave.** For every pair of tasks landing in
214
+ the same wave, compare their `Target Files` lists. If any file path appears in both, **resolve
215
+ it automatically — do not stop.** Keep the lower-numbered task in the current wave and push
216
+ the other one into the next wave by adding `Dependencies: TASK-<lower>` to its task file. Two
217
+ tasks that touch the same file are safe as long as they never run concurrently, and
218
+ serializing them is the one resolution that is always correct. Note the demotion in the wave
219
+ plan and in `RUN.md`.
220
+
221
+ This is the code-level safety net for the planner's own §2 Scope constraint (same-wave tasks
222
+ must not share a file) in case it slipped through review. Only mark a task `needs_breakdown`
223
+ if it is missing required fields — never merely for sharing a file.
224
+
130
225
  **Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
131
- `.ukit/storage/config.json` (default **3**). A wave with more tasks than that is split
226
+ `.ukit/storage/config.json` (default **10**). A wave with more tasks than that is split
132
227
  into consecutive batches of at most that many; finish one batch completely (including 3c
133
228
  copy-back and worktree deletion) before starting the next.
134
229
 
@@ -139,7 +234,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
139
234
 
140
235
  ### I3 — Execute wave by wave (code model agents)
141
236
 
142
- **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 3), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
237
+ **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 10), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
143
238
 
144
239
  For each wave:
145
240
 
@@ -156,7 +251,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
156
251
  TDD — mandatory:
157
252
  1. Write tests from §Test Cases
158
253
  cd .worktrees/task-xxx && <test command>
159
- Confirm RED (immediately GREEN → test is wrong — flag this)
254
+ Confirm RED — paste the actual failing output, don't just assert it happened
255
+ (immediately GREEN → test is wrong — flag this)
160
256
  2. Implement → run → confirm GREEN
161
257
  3. Run §Verification Commands inside the worktree:
162
258
  cd .worktrees/task-xxx && <each verification command>
@@ -169,9 +265,25 @@ Executor Report (append to task file — do NOT touch INDEX.md):
169
265
  EXECUTOR_TOOL: <tool>
170
266
  EXECUTOR_MODEL: <exact model ID — mandatory>
171
267
  EXECUTOR_SUBAGENT: <name or "-">
268
+ RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
269
+ "confirmed" without pasted output is not acceptable>
172
270
  Verification Output: <paste full output>
173
271
  Status: PASS | FAIL
174
272
  Note: <issues or "none">
273
+
274
+ Then RETURN TO THE ORCHESTRATOR AT MOST 10 LINES, in exactly this shape:
275
+ TASK: TASK-xxx
276
+ STATUS: PASS | FAIL
277
+ EXECUTOR_MODEL: <exact model ID>
278
+ FILES: <comma-separated changed paths>
279
+ RED: confirmed | not-confirmed
280
+ VERIFY: <n> commands, all pass | <first failing command + one-line reason>
281
+ NOTE: <one line or "none">
282
+
283
+ Do NOT repeat RED_OUTPUT or verification logs in the returned message. They are already
284
+ written to the task file on disk, which is where the reviewer reads them from. Pasting
285
+ them back a second time is what blows up the orchestrator's context window and kills the
286
+ run mid-pipeline.
175
287
  ```
176
288
 
177
289
  **3c — Copy changes back + delete worktrees** (orchestrator, after each task reports):
@@ -212,6 +324,36 @@ git branch -D handoff/task-xxx
212
324
 
213
325
  Worktrees are **always deleted immediately** — no exceptions.
214
326
 
327
+ **3d — Wave boundary: commit + context checkpoint** (orchestrator, after every wave)
328
+
329
+ A wave boundary is the only safe place to shed context, because everything of value is
330
+ already on disk (task files hold the full logs, git holds the code). Do all four, in order:
331
+
332
+ 1. **Checkpoint the code** — one commit per wave, so every wave is independently revertible:
333
+ ```bash
334
+ git add -A && git commit -m "handoff: wave <N> — TASK-00x, TASK-00y"
335
+ ```
336
+ Do not push here; the single push happens at R5.
337
+
338
+ 2. **Update the run cursor** — rewrite `docs/AI_HANDOFF/RUN.md` with `Phase: I3`,
339
+ `Cursor: wave <N> done`, `Next: wave <N+1>` (or `I4` if that was the last wave).
340
+
341
+ 3. **Collapse the wave in working memory.** From this point on, refer to the finished wave
342
+ only by its one-line-per-task summary (`TASK-xxx PASS <files>`). Do not re-read the task
343
+ files, do not restate executor reports, do not quote diffs from earlier waves. Anything
344
+ the reviewer needs, the reviewer reads from disk itself.
345
+
346
+ 4. **Check the pressure before spawning the next wave.** If a UKit context warning has fired
347
+ this run, or the wave just finished involved more than ~5 agents, state one line —
348
+ `context checkpoint: wave <N> collapsed, <M> tasks summarized` — and continue anyway. Do
349
+ not ask the human to run `/compact`, and do not abandon the run. If the session is
350
+ compacted or restarted for any reason, `RUN.md` plus the `SessionStart` resume hook bring
351
+ the pipeline back exactly here.
352
+
353
+ With a ~200k window and per-task returns capped at 10 lines (3b), a wave of 10 tasks costs
354
+ the orchestrator roughly 100 lines instead of the thousands that pasted RED and verification
355
+ logs used to cost. That difference is what makes a multi-wave cycle finish in one run.
356
+
215
357
  ### I4 — Consolidate + update INDEX (lite model)
216
358
 
217
359
  **Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for I4. After all waves complete, ask it to:
@@ -244,13 +386,21 @@ git status # should show modified/new files, no handoff branches/worktrees
244
386
  git diff --stat # summary of all changes
245
387
  ```
246
388
 
247
- If `git diff` is empty and `git status` is clean → implement was not completed. Stop and re-run Phase 3.
389
+ Because 3d commits each wave, the changes under review are in **commits since the plan commit**, not only in the working tree. Resolve the review range once and use it everywhere below:
390
+ ```bash
391
+ PLAN_COMMIT=$(git log --format=%H --grep='^handoff: plan' -n 1)
392
+ git diff $PLAN_COMMIT # all handoff changes: committed waves + anything uncommitted
393
+ ```
394
+
395
+ If that diff is empty → implement was not completed. Do not stop: re-enter Phase 3 from the run cursor with the remaining `ready` tasks.
248
396
 
249
397
  > Orchestrator (this session) handles the R1 guard check directly.
250
398
 
251
399
  ### R2 — Model isolation check (strong model, always first)
252
400
 
253
- **Claude Code — MANDATORY, do this before anything else in R2–R4:** for each `pending_review` task, call the Agent tool with `subagent_type: "code-reviewer"`. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
401
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
402
+
403
+ **Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
254
404
 
255
405
  The spawned reviewer agent reads `EXECUTOR_MODEL` from each task file `## Executor Report`:
256
406
 
@@ -272,8 +422,8 @@ If any command fails → verdict `critical_block`. Stop. Do NOT proceed to R4.
272
422
  ### R4 — Review unified diff + append verdict (strong model)
273
423
 
274
424
  ```bash
275
- git diff # all uncommitted changes vs $BASE
276
- git status # overview of modified/new/deleted files
425
+ git diff $PLAN_COMMIT # every handoff change: committed waves + working tree
426
+ git status # overview of modified/new/deleted files
277
427
  ```
278
428
 
279
429
  Review as one unified diff — correctness, regression risk, security, edge cases, maintainability. Cross-reference each task's intent in `docs/AI_HANDOFF/tasks/TASK-xxx.md`.
@@ -292,21 +442,70 @@ FINDINGS:
292
442
  NEXT_STATUS_FOR_INDEX: <status>
293
443
  ```
294
444
 
295
- ### R5 — Commit + push or stop (orchestrator)
445
+ Then return to the orchestrator **at most 6 lines** — the full verdict is already on disk:
446
+ ```
447
+ TASK: TASK-xxx
448
+ VERDICT: approved | approved_minor | changes_requested | critical_block
449
+ REVIEWER_MODEL: <exact model ID>
450
+ VERIFICATION_RERUN: PASS | FAIL
451
+ BLOCKING: <one line per critical/important finding, or "none">
452
+ ```
453
+ Do not paste the diff, the findings prose, or verification logs back into the orchestrator.
454
+
455
+ ### R4.5 — Auto-fix loop (max 2 rounds)
456
+
457
+ Any task with `changes_requested` or `critical_block` enters this loop. **Do not report to
458
+ the human and stop** — that is the single biggest reason a cycle used to stretch across days.
459
+
460
+ For round in 1..2:
461
+
462
+ 1. Collect every task still not `approved`/`approved_minor`. If none, exit the loop.
463
+ 2. Group them into waves exactly as I2 does (dependencies + the same auto-serializing file
464
+ conflict rule). Spawn one `feature-implementer` per task **in parallel**, at most
465
+ `handoff.maxParallelAgents`, each in a fresh worktree per 3a. Each agent gets:
466
+ - the task file path (it reads its own `## Reviewer Verdict` from disk — do not paste
467
+ findings into the prompt),
468
+ - the instruction: fix every `critical` and `important` finding, leave `minor` alone
469
+ unless trivially safe, keep the existing tests green, add a regression test for each
470
+ critical finding, then re-run §Verification Commands and append a new
471
+ `## Executor Report (fix round <N>)` to the task file.
472
+ - the same ≤10-line return contract as 3b.
473
+ 3. Copy back and delete worktrees per 3c. Commit the round:
474
+ ```bash
475
+ git add -A && git commit -m "handoff: fix round <N> — TASK-00x"
476
+ ```
477
+ 4. Re-review **only the tasks touched this round**, in parallel, per R2–R4. A reviewer must
478
+ still differ from the executor model, and still re-runs verification itself.
479
+ 5. Update `INDEX.md` and `RUN.md`. If everything is now approved, exit the loop.
296
480
 
297
- **All tasks `approved` or `approved_minor`:**
481
+ After round 2, any task still not approved is genuinely stuck: leave it `blocked`, record why
482
+ in its `## Discussion` thread, and carry it into the Final Report. **Every other task still
483
+ proceeds to R5** — one stubborn task must never hold the whole cycle hostage.
298
484
 
299
- Update `docs/AI_HANDOFF/INDEX.md`: all approved tasks → `done`.
485
+ ### R5 — Commit + push (orchestrator)
486
+
487
+ R4.5 has already run, so by this point every task is either approved or genuinely stuck.
488
+
489
+ Update `docs/AI_HANDOFF/INDEX.md`: approved tasks → `done`; anything still failing after both
490
+ fix rounds → `blocked`.
491
+
492
+ Commit whatever remains uncommitted (the per-wave commits from 3d are already in), then push
493
+ once — this is the run's only push:
300
494
 
301
495
  ```bash
302
- git add . && git commit -m "handoff: implement <goal>" && git push origin $BASE
496
+ git add -A && git commit -m "handoff: implement <goal>" # skip if nothing left to commit
497
+ git push origin $BASE
303
498
  ```
304
499
 
305
- Replace `<goal>` with the one-sentence goal from ACTIVE.md.
500
+ Replace `<goal>` with the one-sentence goal from ACTIVE.md. Push without asking — `git push`
501
+ is in the settings `allow` list precisely so a one-shot run never stalls at the last step.
502
+ Force-push stays denied, so this can only ever fast-forward.
306
503
 
307
- **Any task `changes_requested` or `critical_block`:**
504
+ **If some tasks are `blocked`:** still push. The approved work is reviewed, tested, and
505
+ committed; withholding it helps no one, and the per-wave commits make any subset revertible.
506
+ Name the blocked tasks in the Final Report.
308
507
 
309
- Report which tasks need fixes. Do NOT push. Executor re-runs Phase 3 for those tasks.
508
+ Finally set `Phase: done` in `docs/AI_HANDOFF/RUN.md`.
310
509
 
311
510
  ---
312
511
 
@@ -337,7 +536,19 @@ Re-run assertions after removal to confirm clean state.
337
536
  handoff-fullstack complete:
338
537
  Cycle: <ID>
339
538
  Tasks: <N> approved, <M> blocked
340
- Git: committed + pushed to <branch> | pending human (<M> blocked)
539
+ Fix rounds used: <0|1|2>
540
+ Git: <K> wave commits + pushed to <branch>
341
541
  Worktrees: all cleaned
342
542
  Branches: all cleaned
543
+ Blocked (needs you): <TASK-xxx — one-line reason> | none
343
544
  ```
545
+
546
+ Then, as the last line of the run:
547
+
548
+ > Cycle finished. Run `/compact` before starting the next cycle — this session is carrying
549
+ > the whole pipeline's history and the next cycle deserves a clean window.
550
+
551
+ This is the **only** place in the pipeline that may ask for a compaction. Asking at a cycle
552
+ boundary costs nothing: all state is in git, `INDEX.md` and `RUN.md`, so a compacted or
553
+ brand-new session picks the next cycle up with no loss. Asking mid-cycle is forbidden — see
554
+ the Autonomy Contract.
@@ -8,7 +8,48 @@
8
8
  $ARGUMENTS
9
9
  _Empty = all `ready` tasks. Or: "TASK-001" for a specific task._
10
10
 
11
- > **NO git commit. NO git push. Ever.** AI only makes file changes. Human commits when ready.
11
+ > **Commits: yes, one per wave. Push: no.** Each wave is checkpointed to git so any step is
12
+ > revertible, but nothing leaves the machine — `/ukit:handoff-review` decides what gets pushed.
13
+
14
+ ---
15
+
16
+ ## Autonomy Contract — read before anything else
17
+
18
+ Run every `ready` task to completion in this one invocation. The human is not watching and
19
+ will not answer mid-run.
20
+
21
+ **Ask nothing after the run starts.** Every decision below has a defined automatic
22
+ resolution — take it, and note it in the wave summary:
23
+
24
+ | Situation | Required behavior |
25
+ |-----------|------------------|
26
+ | Working tree dirty at Step 1 | commit a checkpoint, continue |
27
+ | Two same-wave tasks share a file | move the later task to the next wave, continue |
28
+ | A task's agent reports FAIL | leave it `blocked`, finish every other task, report at the end |
29
+ | Context warning fires | collapse the finished wave to one line per task, continue |
30
+ | No `ready` tasks but `changes_requested` ones exist | treat those as the work — this is a fix pass |
31
+
32
+ Escalate only for a blocker outside the repo. Never pause mid-wave to ask for `/compact`:
33
+ context is managed structurally (short agent returns + wave-boundary collapse), and the run
34
+ cursor makes an interrupted run resumable. A compaction request belongs at the end of the
35
+ command, never inside it.
36
+
37
+ ### Run cursor — write after every step
38
+
39
+ Maintain `docs/AI_HANDOFF/RUN.md`, rewritten (whole file) after every numbered step:
40
+
41
+ ```
42
+ Command: handoff-implement
43
+ Goal: <one sentence>
44
+ Base: <BASE>
45
+ Phase: <1|2|3|4|done>
46
+ Cursor: wave <N> batch <M> — <what just finished>
47
+ Next: <the exact next step to run>
48
+ ```
49
+
50
+ If this file already exists with `Phase:` not `done` when the command starts, this is a
51
+ **continuation**: read the cursor, jump to `Next:`, skip tasks already `pending_review` or
52
+ `done`. Do not restart from Step 1 and do not ask whether to continue.
12
53
 
13
54
  ---
14
55
 
@@ -16,14 +57,19 @@ _Empty = all `ready` tasks. Or: "TASK-001" for a specific task._
16
57
 
17
58
  Read `docs/AI_HANDOFF/ACTIVE.md` → get `Base: <BASE>`.
18
59
  Read `docs/AI_HANDOFF/INDEX.md` → collect `ready` tasks (or specific task from $ARGUMENTS).
19
- If no ready tasks → report and stop.
60
+ If there are no `ready` tasks, check for `changes_requested`/`blocked` ones and run those
61
+ instead. Only if nothing is actionable → report and stop.
20
62
 
21
- Verify working tree is clean before starting:
63
+ Verify working tree state:
22
64
  ```bash
23
- git status # must be clean — no uncommitted changes
65
+ git status
24
66
  ```
25
67
 
26
- If working tree is dirty → stop. Ask human to commit or stash first.
68
+ If the working tree is dirty, do not stop — commit a checkpoint so the wave copy-back starts
69
+ from a known state and the pre-existing edits stay recoverable:
70
+ ```bash
71
+ git add -A && git commit -m "handoff: checkpoint before implement"
72
+ ```
27
73
 
28
74
  ## Step 2 — Build wave groups
29
75
 
@@ -33,9 +79,22 @@ Read each `tasks/TASK-xxx.md` for `Dependencies` field:
33
79
  - Chain A→B→C = 3 waves of 1 task each (sequential, no parallel)
34
80
  - Independent A, B, C = 1 wave of 3 tasks (parallel)
35
81
 
82
+ ### Conflict check — mandatory, before spawning any wave
83
+
84
+ For every pair of tasks landing in the same wave, compare their `Target Files` lists. If any
85
+ file path appears in both, **resolve it automatically — do not stop.** Keep the
86
+ lower-numbered task in the current wave and push the other into the next wave by adding
87
+ `Dependencies: TASK-<lower>` to its task file. Two tasks touching the same file are safe as
88
+ long as they never run concurrently, and serializing them is the one resolution that is
89
+ always correct.
90
+
91
+ This is the code-level safety net for the planner's own §2 Scope constraint (same-wave tasks
92
+ must not share a file) in case it slipped through review. Only mark a task `needs_breakdown`
93
+ when it is missing required fields — never merely for sharing a file.
94
+
36
95
  ### Batch each wave — mandatory
37
96
 
38
- Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **3**). A wave
97
+ Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). A wave
39
98
  with more tasks than that is split into consecutive batches of at most that many; finish
40
99
  one batch completely (including 3c copy-back and worktree deletion) before starting the
41
100
  next.
@@ -68,7 +127,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
68
127
  TDD — mandatory:
69
128
  1. Write tests from §Test Cases
70
129
  cd .worktrees/task-xxx && <test command>
71
- Confirm RED (immediately GREEN → test is wrong — flag this)
130
+ Confirm RED — paste the actual failing output, don't just assert it happened
131
+ (immediately GREEN → test is wrong — flag this)
72
132
  2. Implement → run → confirm GREEN
73
133
  3. Run §Verification Commands inside the worktree:
74
134
  cd .worktrees/task-xxx && <each verification command>
@@ -81,9 +141,24 @@ Executor Report (append to task file — do NOT touch INDEX.md):
81
141
  EXECUTOR_TOOL: <tool>
82
142
  EXECUTOR_MODEL: <exact model ID — mandatory>
83
143
  EXECUTOR_SUBAGENT: <name or "-">
144
+ RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
145
+ "confirmed" without pasted output is not acceptable>
84
146
  Verification Output: <paste>
85
147
  Status: PASS | FAIL
86
148
  Note: <issues or "none">
149
+
150
+ Then RETURN TO THE ORCHESTRATOR AT MOST 10 LINES, in exactly this shape:
151
+ TASK: TASK-xxx
152
+ STATUS: PASS | FAIL
153
+ EXECUTOR_MODEL: <exact model ID>
154
+ FILES: <comma-separated changed paths>
155
+ RED: confirmed | not-confirmed
156
+ VERIFY: <n> commands, all pass | <first failing command + one-line reason>
157
+ NOTE: <one line or "none">
158
+
159
+ Do NOT repeat RED_OUTPUT or verification logs in the returned message. They are already on
160
+ disk in the task file, which is where the reviewer reads them from. Pasting them back a
161
+ second time is what blows up the orchestrator's context window and kills the run mid-wave.
87
162
  ```
88
163
 
89
164
  > Other tools without subagent support: open each task in a separate session with the code model.
@@ -132,21 +207,44 @@ git branch -D handoff/task-xxx
132
207
  **Orchestrator writes INDEX.md** — agents never touch INDEX directly.
133
208
  **Worktrees are always deleted immediately** — no exceptions.
134
209
 
135
- ### 3d — Next wave
136
- Create new worktrees from `$BASE`. Repeat 3a–3c.
210
+ ### 3d — Wave boundary: commit + context checkpoint
211
+
212
+ A wave boundary is the only safe place to shed context, because everything of value is
213
+ already on disk (task files hold the full logs, git holds the code). Do all four, in order:
214
+
215
+ 1. **Checkpoint the code** — one commit per wave, so every wave is independently revertible:
216
+ ```bash
217
+ git add -A && git commit -m "handoff: wave <N> — TASK-00x, TASK-00y"
218
+ ```
219
+ Do not push. `/ukit:handoff-review` owns that decision.
220
+
221
+ 2. **Update the run cursor** — rewrite `docs/AI_HANDOFF/RUN.md` with `Cursor: wave <N> done`,
222
+ `Next: wave <N+1>` (or `Step 4` if that was the last wave).
223
+
224
+ 3. **Collapse the wave in working memory.** From here on, refer to the finished wave only by
225
+ its one-line-per-task summary. Do not re-read task files, restate executor reports, or
226
+ quote earlier diffs. The reviewer reads what it needs from disk itself.
227
+
228
+ 4. **Continue.** If a UKit context warning has fired, say one line —
229
+ `context checkpoint: wave <N> collapsed` — and start the next wave anyway. Do not ask for
230
+ `/compact` here. If the session ends regardless, `RUN.md` plus the `SessionStart` resume
231
+ hook bring the run back to exactly this point.
232
+
233
+ Then create the next wave's worktrees from `$BASE` and repeat 3a–3c.
137
234
 
138
235
  ## Step 4 — Finalize: verify cleanup
139
236
 
140
237
  All waves complete. Verify:
141
238
 
142
239
  ```bash
143
- git worktree list # must show only main worktree
240
+ git worktree list # must show only main worktree
144
241
  git branch | grep handoff # must be empty
145
- git status # shows uncommitted changes — this is expected
146
- git diff --stat # summary of everything that changed
242
+ git log --oneline -n 10 # one commit per wave
243
+ git status # clean, or only trailing edits from the last wave
147
244
  ```
148
245
 
149
- All changes are in the **main working tree as uncommitted files**. No branches or worktrees remain.
246
+ All changes are committed as **one commit per wave on `$BASE`**, nothing pushed. No branches
247
+ or worktrees remain. Set `Phase: done` in `docs/AI_HANDOFF/RUN.md`.
150
248
 
151
249
  ## Step 5 — Report
152
250
 
@@ -154,8 +252,13 @@ All changes are in the **main working tree as uncommitted files**. No branches o
154
252
  Wave summary:
155
253
  Wave 1: TASK-001 [pending_review], TASK-002 [blocked]
156
254
  Wave 2: TASK-003 [pending_review]
255
+ Commits: <K> wave commits on <BASE>, not pushed
157
256
  All worktrees and branches removed.
158
- Uncommitted changes in working tree — human commits when ready.
257
+ Blocked (needs you): <TASK-xxx — one-line reason> | none
159
258
  ```
160
259
 
161
260
  **Next:** switch to strong model → `/ukit:handoff-review`
261
+
262
+ If this session is now heavy, this is the right moment to `/compact` — the wave commits,
263
+ `INDEX.md` and `RUN.md` hold all the state the review phase needs. Suggest it here, at the
264
+ command boundary, never inside a wave.