@ngockhoale/ukit 2.0.6 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +41 -0
- package/manifests/platform.full.yaml +11 -0
- package/package.json +1 -1
- package/templates/.claude/agents/code-reviewer.md +28 -1
- package/templates/.claude/agents/feature-implementer.md +31 -3
- package/templates/.claude/agents/handoff-planner.md +113 -9
- package/templates/.claude/commands/ukit/handoff-create.md +47 -3
- package/templates/.claude/commands/ukit/handoff-fullstack.md +230 -19
- package/templates/.claude/commands/ukit/handoff-implement.md +117 -14
- package/templates/.claude/commands/ukit/handoff-review.md +104 -33
- package/templates/.claude/hooks/context-window-guard.sh +34 -1
- package/templates/.claude/hooks/handoff-model-guard.sh +12 -0
- package/templates/.claude/hooks/handoff-resume.sh +90 -0
- package/templates/.claude/settings.json +9 -1
- package/templates/docs/AI_HANDOFF/RULES.md +55 -7
- package/templates/ukit/storage/config.json +18 -4
|
@@ -22,6 +22,66 @@ $ARGUMENTS
|
|
|
22
22
|
|
|
23
23
|
---
|
|
24
24
|
|
|
25
|
+
## Autonomy Contract — read before anything else
|
|
26
|
+
|
|
27
|
+
This command is a **one-shot pipeline**. The human is not watching. Treat every phase below
|
|
28
|
+
as unattended: they may walk away after submitting `$ARGUMENTS` and come back to a finished
|
|
29
|
+
result.
|
|
30
|
+
|
|
31
|
+
**The single question window is P0, before any file is written.** If `$ARGUMENTS` leaves a
|
|
32
|
+
choice that would change what gets built — not how it gets built — batch every such question
|
|
33
|
+
into **one** `AskUserQuestion` call (max 4 questions, each with a `(Recommended)` first
|
|
34
|
+
option) and resolve them all at once. Record the answers in `PLAN.md` §1.
|
|
35
|
+
|
|
36
|
+
**After P0 closes, do not ask the human anything until the Final Report.** Every decision
|
|
37
|
+
point in the phases below has a defined automatic resolution. Take it. Specifically:
|
|
38
|
+
|
|
39
|
+
| Situation | Old behavior | Required behavior now |
|
|
40
|
+
|-----------|--------------|----------------------|
|
|
41
|
+
| Scope spans multiple subsystems | wait for confirmation | decompose into sequential cycles, run cycle 1 now, queue the rest in `INDEX.md` |
|
|
42
|
+
| Plan review returns `Issues Found` | up to 3 rounds, then stop | 1 revision round; if round 2 still has findings, apply them and proceed — log to `## Plan Review Log` |
|
|
43
|
+
| Two same-wave tasks share a file | STOP | move the later task to the next wave (add the dependency), continue |
|
|
44
|
+
| INDEX has `pending_review`/`blocked` tasks | STOP | **resume** — see Resume below |
|
|
45
|
+
| Reviewer returns `changes_requested`/`critical_block` | report and stop | run the Phase 4 auto-fix loop (max 2 rounds) |
|
|
46
|
+
| Verification command fails | stop | fix and re-verify inside the auto-fix loop |
|
|
47
|
+
|
|
48
|
+
Escalate to the human **only** for: a blocker outside the repo (missing credential, external
|
|
49
|
+
service down), or a task still failing after both auto-fix rounds. Even then, finish every
|
|
50
|
+
other task first and report what was left undone.
|
|
51
|
+
|
|
52
|
+
**Never pause to ask for `/compact`.** Context is managed structurally — subagents write full
|
|
53
|
+
logs to disk and return short summaries (see 3b), and the run cursor makes the pipeline
|
|
54
|
+
resumable. If compaction does happen, the run resumes from the cursor automatically.
|
|
55
|
+
|
|
56
|
+
### Run cursor — write after every step
|
|
57
|
+
|
|
58
|
+
Maintain `docs/AI_HANDOFF/RUN.md`. Rewrite it (whole file, 6 lines) immediately after every
|
|
59
|
+
numbered step completes:
|
|
60
|
+
|
|
61
|
+
```
|
|
62
|
+
Command: handoff-fullstack
|
|
63
|
+
Goal: <one sentence from $ARGUMENTS>
|
|
64
|
+
Base: <BASE branch>
|
|
65
|
+
Phase: <P1|P2|P2.5|P3|I1|I2|I3|I4|R1|R2|R3|R4|R5|done>
|
|
66
|
+
Cursor: wave <N> batch <M> — <what just finished>
|
|
67
|
+
Next: <the exact next step to run>
|
|
68
|
+
```
|
|
69
|
+
|
|
70
|
+
This file is the resume contract. It costs one small write per step and is what turns an
|
|
71
|
+
interrupted run into a continuable one.
|
|
72
|
+
|
|
73
|
+
### Resume — when the run is re-invoked mid-flight
|
|
74
|
+
|
|
75
|
+
If `docs/AI_HANDOFF/RUN.md` exists with `Phase:` not `done`, this is a **continuation, not a
|
|
76
|
+
new cycle**. Do NOT re-plan, do NOT overwrite `PLAN.md`, do NOT ask the human whether to
|
|
77
|
+
continue. Read the cursor, then jump straight to `Next:` and carry on. Tasks already `done`
|
|
78
|
+
or `pending_review` are skipped; only `ready`, `in_progress`, `changes_requested` and
|
|
79
|
+
`blocked` tasks are picked up.
|
|
80
|
+
|
|
81
|
+
Only when `RUN.md` is absent or `Phase: done` does `$ARGUMENTS` start a fresh cycle.
|
|
82
|
+
|
|
83
|
+
---
|
|
84
|
+
|
|
25
85
|
## Phase 1+2 — Plan (strong model)
|
|
26
86
|
|
|
27
87
|
### P1 — Read context (lite model)
|
|
@@ -45,7 +105,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
45
105
|
|
|
46
106
|
1. **Guard — check INDEX statuses:**
|
|
47
107
|
- All tasks are `ready` (planning only) → re-run is allowed: overwrite `PLAN.md` and `TASK-xxx.md` freely (iterative refinement).
|
|
48
|
-
- Any task has status `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **
|
|
108
|
+
- Any task has status `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **do not overwrite, and do not stop either.** Mid-flight work exists. Skip the rest of P2 entirely, write the run cursor with `Phase: I1`, and hand control back to the orchestrator to resume from Phase 3 with the surviving tasks (see Resume in the Autonomy Contract). Overwriting is what is unsafe here — stopping is not required.
|
|
49
109
|
- No tasks → fresh cycle, proceed.
|
|
50
110
|
|
|
51
111
|
2. **Resolve base branch:**
|
|
@@ -57,7 +117,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
57
117
|
- §1 Intent — problem + success definition
|
|
58
118
|
- §2 Scope — in / out of scope; same-wave tasks must not modify the same file (prevents merge conflicts)
|
|
59
119
|
- §3 Approach — solution, trade-offs, alternatives rejected
|
|
60
|
-
- §4 Test Plan — happy path + ≥
|
|
120
|
+
- §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
|
|
61
121
|
- §5 Verification — exact shell commands executor will run
|
|
62
122
|
- §6 Acceptance — done checklist (prefer verifiable criteria with commands)
|
|
63
123
|
|
|
@@ -71,7 +131,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
71
131
|
Every task MUST have:
|
|
72
132
|
- Target Files (exact paths — no two tasks in same wave share a file)
|
|
73
133
|
- Dependencies (`TASK-xxx` or `none` — wave structure inferred from this)
|
|
74
|
-
- Test Cases (Type | Name | Expected — ≥1 happy + ≥
|
|
134
|
+
- Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
|
|
75
135
|
- Test Files (exact paths)
|
|
76
136
|
- Verification Commands (runnable shell commands)
|
|
77
137
|
- Acceptance Criteria (verifiable checklist)
|
|
@@ -89,6 +149,25 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
89
149
|
|
|
90
150
|
7. **Report:** task IDs, dependency graph, any `needs_breakdown` tasks + reason.
|
|
91
151
|
|
|
152
|
+
### P2.5 — Independent plan review (strong model, separate agent)
|
|
153
|
+
|
|
154
|
+
**Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet).
|
|
155
|
+
|
|
156
|
+
- count = 0 → run the review below.
|
|
157
|
+
- count = 1 and the round returned `Issues Found` → planner revises once, then run **one** more review round.
|
|
158
|
+
- count ≥ 2 → do NOT invoke the reviewer again. **Do not stop and do not ask the human.** Have the planner apply every outstanding finding directly to `PLAN.md` and the affected `TASK-xxx.md` files, append `### Round <N> — findings applied without re-review` to the Plan Review Log listing what was changed, then proceed to P3.
|
|
159
|
+
|
|
160
|
+
The gate still does its job — two independent opus passes shape the plan before a line of code
|
|
161
|
+
is written. What it no longer does is hand a stalled plan back to a human who isn't there.
|
|
162
|
+
|
|
163
|
+
**Claude Code — MANDATORY, do this before anything else in P2.5:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from P2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
|
|
164
|
+
|
|
165
|
+
1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
|
|
166
|
+
2. `Issues Found` → route back to P2: `handoff-planner` revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for one more P2.5 review round — subject to the loop cap above.
|
|
167
|
+
3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed to P3.
|
|
168
|
+
|
|
169
|
+
> Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from P2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
|
|
170
|
+
|
|
92
171
|
### P3 — Commit the plan (lite model)
|
|
93
172
|
|
|
94
173
|
**Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for this commit step (lite tier — haiku/unic-lite). Run:
|
|
@@ -117,7 +196,11 @@ Verify working tree is clean after the plan commit:
|
|
|
117
196
|
git status # must be clean — plan commit already done in P3
|
|
118
197
|
```
|
|
119
198
|
|
|
120
|
-
If working tree is dirty
|
|
199
|
+
If the working tree is dirty, do not stop. Commit the stragglers as a checkpoint so they stay
|
|
200
|
+
recoverable and the wave copy-back starts from a known state:
|
|
201
|
+
```bash
|
|
202
|
+
git add -A && git commit -m "handoff: checkpoint before implement"
|
|
203
|
+
```
|
|
121
204
|
|
|
122
205
|
### I2 — Infer wave groups
|
|
123
206
|
|
|
@@ -127,8 +210,20 @@ Read each `docs/AI_HANDOFF/tasks/TASK-xxx.md` for `Dependencies` field:
|
|
|
127
210
|
- Chain A→B→C = 3 waves of 1 task each (sequential)
|
|
128
211
|
- Independent A, B, C = 1 wave of 3 tasks (parallel)
|
|
129
212
|
|
|
213
|
+
**Conflict check — mandatory, before spawning any wave.** For every pair of tasks landing in
|
|
214
|
+
the same wave, compare their `Target Files` lists. If any file path appears in both, **resolve
|
|
215
|
+
it automatically — do not stop.** Keep the lower-numbered task in the current wave and push
|
|
216
|
+
the other one into the next wave by adding `Dependencies: TASK-<lower>` to its task file. Two
|
|
217
|
+
tasks that touch the same file are safe as long as they never run concurrently, and
|
|
218
|
+
serializing them is the one resolution that is always correct. Note the demotion in the wave
|
|
219
|
+
plan and in `RUN.md`.
|
|
220
|
+
|
|
221
|
+
This is the code-level safety net for the planner's own §2 Scope constraint (same-wave tasks
|
|
222
|
+
must not share a file) in case it slipped through review. Only mark a task `needs_breakdown`
|
|
223
|
+
if it is missing required fields — never merely for sharing a file.
|
|
224
|
+
|
|
130
225
|
**Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
|
|
131
|
-
`.ukit/storage/config.json` (default **
|
|
226
|
+
`.ukit/storage/config.json` (default **10**). A wave with more tasks than that is split
|
|
132
227
|
into consecutive batches of at most that many; finish one batch completely (including 3c
|
|
133
228
|
copy-back and worktree deletion) before starting the next.
|
|
134
229
|
|
|
@@ -139,7 +234,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
|
|
|
139
234
|
|
|
140
235
|
### I3 — Execute wave by wave (code model agents)
|
|
141
236
|
|
|
142
|
-
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default
|
|
237
|
+
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 10), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
|
|
143
238
|
|
|
144
239
|
For each wave:
|
|
145
240
|
|
|
@@ -156,7 +251,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
|
|
|
156
251
|
TDD — mandatory:
|
|
157
252
|
1. Write tests from §Test Cases
|
|
158
253
|
cd .worktrees/task-xxx && <test command>
|
|
159
|
-
Confirm RED
|
|
254
|
+
Confirm RED — paste the actual failing output, don't just assert it happened
|
|
255
|
+
(immediately GREEN → test is wrong — flag this)
|
|
160
256
|
2. Implement → run → confirm GREEN
|
|
161
257
|
3. Run §Verification Commands inside the worktree:
|
|
162
258
|
cd .worktrees/task-xxx && <each verification command>
|
|
@@ -169,9 +265,25 @@ Executor Report (append to task file — do NOT touch INDEX.md):
|
|
|
169
265
|
EXECUTOR_TOOL: <tool>
|
|
170
266
|
EXECUTOR_MODEL: <exact model ID — mandatory>
|
|
171
267
|
EXECUTOR_SUBAGENT: <name or "-">
|
|
268
|
+
RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
|
|
269
|
+
"confirmed" without pasted output is not acceptable>
|
|
172
270
|
Verification Output: <paste full output>
|
|
173
271
|
Status: PASS | FAIL
|
|
174
272
|
Note: <issues or "none">
|
|
273
|
+
|
|
274
|
+
Then RETURN TO THE ORCHESTRATOR AT MOST 10 LINES, in exactly this shape:
|
|
275
|
+
TASK: TASK-xxx
|
|
276
|
+
STATUS: PASS | FAIL
|
|
277
|
+
EXECUTOR_MODEL: <exact model ID>
|
|
278
|
+
FILES: <comma-separated changed paths>
|
|
279
|
+
RED: confirmed | not-confirmed
|
|
280
|
+
VERIFY: <n> commands, all pass | <first failing command + one-line reason>
|
|
281
|
+
NOTE: <one line or "none">
|
|
282
|
+
|
|
283
|
+
Do NOT repeat RED_OUTPUT or verification logs in the returned message. They are already
|
|
284
|
+
written to the task file on disk, which is where the reviewer reads them from. Pasting
|
|
285
|
+
them back a second time is what blows up the orchestrator's context window and kills the
|
|
286
|
+
run mid-pipeline.
|
|
175
287
|
```
|
|
176
288
|
|
|
177
289
|
**3c — Copy changes back + delete worktrees** (orchestrator, after each task reports):
|
|
@@ -212,6 +324,36 @@ git branch -D handoff/task-xxx
|
|
|
212
324
|
|
|
213
325
|
Worktrees are **always deleted immediately** — no exceptions.
|
|
214
326
|
|
|
327
|
+
**3d — Wave boundary: commit + context checkpoint** (orchestrator, after every wave)
|
|
328
|
+
|
|
329
|
+
A wave boundary is the only safe place to shed context, because everything of value is
|
|
330
|
+
already on disk (task files hold the full logs, git holds the code). Do all four, in order:
|
|
331
|
+
|
|
332
|
+
1. **Checkpoint the code** — one commit per wave, so every wave is independently revertible:
|
|
333
|
+
```bash
|
|
334
|
+
git add -A && git commit -m "handoff: wave <N> — TASK-00x, TASK-00y"
|
|
335
|
+
```
|
|
336
|
+
Do not push here; the single push happens at R5.
|
|
337
|
+
|
|
338
|
+
2. **Update the run cursor** — rewrite `docs/AI_HANDOFF/RUN.md` with `Phase: I3`,
|
|
339
|
+
`Cursor: wave <N> done`, `Next: wave <N+1>` (or `I4` if that was the last wave).
|
|
340
|
+
|
|
341
|
+
3. **Collapse the wave in working memory.** From this point on, refer to the finished wave
|
|
342
|
+
only by its one-line-per-task summary (`TASK-xxx PASS <files>`). Do not re-read the task
|
|
343
|
+
files, do not restate executor reports, do not quote diffs from earlier waves. Anything
|
|
344
|
+
the reviewer needs, the reviewer reads from disk itself.
|
|
345
|
+
|
|
346
|
+
4. **Check the pressure before spawning the next wave.** If a UKit context warning has fired
|
|
347
|
+
this run, or the wave just finished involved more than ~5 agents, state one line —
|
|
348
|
+
`context checkpoint: wave <N> collapsed, <M> tasks summarized` — and continue anyway. Do
|
|
349
|
+
not ask the human to run `/compact`, and do not abandon the run. If the session is
|
|
350
|
+
compacted or restarted for any reason, `RUN.md` plus the `SessionStart` resume hook bring
|
|
351
|
+
the pipeline back exactly here.
|
|
352
|
+
|
|
353
|
+
With a ~200k window and per-task returns capped at 10 lines (3b), a wave of 10 tasks costs
|
|
354
|
+
the orchestrator roughly 100 lines instead of the thousands that pasted RED and verification
|
|
355
|
+
logs used to cost. That difference is what makes a multi-wave cycle finish in one run.
|
|
356
|
+
|
|
215
357
|
### I4 — Consolidate + update INDEX (lite model)
|
|
216
358
|
|
|
217
359
|
**Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for I4. After all waves complete, ask it to:
|
|
@@ -244,13 +386,21 @@ git status # should show modified/new files, no handoff branches/worktrees
|
|
|
244
386
|
git diff --stat # summary of all changes
|
|
245
387
|
```
|
|
246
388
|
|
|
247
|
-
|
|
389
|
+
Because 3d commits each wave, the changes under review are in **commits since the plan commit**, not only in the working tree. Resolve the review range once and use it everywhere below:
|
|
390
|
+
```bash
|
|
391
|
+
PLAN_COMMIT=$(git log --format=%H --grep='^handoff: plan' -n 1)
|
|
392
|
+
git diff $PLAN_COMMIT # all handoff changes: committed waves + anything uncommitted
|
|
393
|
+
```
|
|
394
|
+
|
|
395
|
+
If that diff is empty → implement was not completed. Do not stop: re-enter Phase 3 from the run cursor with the remaining `ready` tasks.
|
|
248
396
|
|
|
249
397
|
> Orchestrator (this session) handles the R1 guard check directly.
|
|
250
398
|
|
|
251
399
|
### R2 — Model isolation check (strong model, always first)
|
|
252
400
|
|
|
253
|
-
**
|
|
401
|
+
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
|
|
402
|
+
|
|
403
|
+
**Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
|
|
254
404
|
|
|
255
405
|
The spawned reviewer agent reads `EXECUTOR_MODEL` from each task file `## Executor Report`:
|
|
256
406
|
|
|
@@ -272,8 +422,8 @@ If any command fails → verdict `critical_block`. Stop. Do NOT proceed to R4.
|
|
|
272
422
|
### R4 — Review unified diff + append verdict (strong model)
|
|
273
423
|
|
|
274
424
|
```bash
|
|
275
|
-
git diff
|
|
276
|
-
git status
|
|
425
|
+
git diff $PLAN_COMMIT # every handoff change: committed waves + working tree
|
|
426
|
+
git status # overview of modified/new/deleted files
|
|
277
427
|
```
|
|
278
428
|
|
|
279
429
|
Review as one unified diff — correctness, regression risk, security, edge cases, maintainability. Cross-reference each task's intent in `docs/AI_HANDOFF/tasks/TASK-xxx.md`.
|
|
@@ -292,21 +442,70 @@ FINDINGS:
|
|
|
292
442
|
NEXT_STATUS_FOR_INDEX: <status>
|
|
293
443
|
```
|
|
294
444
|
|
|
295
|
-
|
|
445
|
+
Then return to the orchestrator **at most 6 lines** — the full verdict is already on disk:
|
|
446
|
+
```
|
|
447
|
+
TASK: TASK-xxx
|
|
448
|
+
VERDICT: approved | approved_minor | changes_requested | critical_block
|
|
449
|
+
REVIEWER_MODEL: <exact model ID>
|
|
450
|
+
VERIFICATION_RERUN: PASS | FAIL
|
|
451
|
+
BLOCKING: <one line per critical/important finding, or "none">
|
|
452
|
+
```
|
|
453
|
+
Do not paste the diff, the findings prose, or verification logs back into the orchestrator.
|
|
454
|
+
|
|
455
|
+
### R4.5 — Auto-fix loop (max 2 rounds)
|
|
456
|
+
|
|
457
|
+
Any task with `changes_requested` or `critical_block` enters this loop. **Do not report to
|
|
458
|
+
the human and stop** — that is the single biggest reason a cycle used to stretch across days.
|
|
459
|
+
|
|
460
|
+
For round in 1..2:
|
|
461
|
+
|
|
462
|
+
1. Collect every task still not `approved`/`approved_minor`. If none, exit the loop.
|
|
463
|
+
2. Group them into waves exactly as I2 does (dependencies + the same auto-serializing file
|
|
464
|
+
conflict rule). Spawn one `feature-implementer` per task **in parallel**, at most
|
|
465
|
+
`handoff.maxParallelAgents`, each in a fresh worktree per 3a. Each agent gets:
|
|
466
|
+
- the task file path (it reads its own `## Reviewer Verdict` from disk — do not paste
|
|
467
|
+
findings into the prompt),
|
|
468
|
+
- the instruction: fix every `critical` and `important` finding, leave `minor` alone
|
|
469
|
+
unless trivially safe, keep the existing tests green, add a regression test for each
|
|
470
|
+
critical finding, then re-run §Verification Commands and append a new
|
|
471
|
+
`## Executor Report (fix round <N>)` to the task file.
|
|
472
|
+
- the same ≤10-line return contract as 3b.
|
|
473
|
+
3. Copy back and delete worktrees per 3c. Commit the round:
|
|
474
|
+
```bash
|
|
475
|
+
git add -A && git commit -m "handoff: fix round <N> — TASK-00x"
|
|
476
|
+
```
|
|
477
|
+
4. Re-review **only the tasks touched this round**, in parallel, per R2–R4. A reviewer must
|
|
478
|
+
still differ from the executor model, and still re-runs verification itself.
|
|
479
|
+
5. Update `INDEX.md` and `RUN.md`. If everything is now approved, exit the loop.
|
|
296
480
|
|
|
297
|
-
|
|
481
|
+
After round 2, any task still not approved is genuinely stuck: leave it `blocked`, record why
|
|
482
|
+
in its `## Discussion` thread, and carry it into the Final Report. **Every other task still
|
|
483
|
+
proceeds to R5** — one stubborn task must never hold the whole cycle hostage.
|
|
298
484
|
|
|
299
|
-
|
|
485
|
+
### R5 — Commit + push (orchestrator)
|
|
486
|
+
|
|
487
|
+
R4.5 has already run, so by this point every task is either approved or genuinely stuck.
|
|
488
|
+
|
|
489
|
+
Update `docs/AI_HANDOFF/INDEX.md`: approved tasks → `done`; anything still failing after both
|
|
490
|
+
fix rounds → `blocked`.
|
|
491
|
+
|
|
492
|
+
Commit whatever remains uncommitted (the per-wave commits from 3d are already in), then push
|
|
493
|
+
once — this is the run's only push:
|
|
300
494
|
|
|
301
495
|
```bash
|
|
302
|
-
git add
|
|
496
|
+
git add -A && git commit -m "handoff: implement <goal>" # skip if nothing left to commit
|
|
497
|
+
git push origin $BASE
|
|
303
498
|
```
|
|
304
499
|
|
|
305
|
-
Replace `<goal>` with the one-sentence goal from ACTIVE.md.
|
|
500
|
+
Replace `<goal>` with the one-sentence goal from ACTIVE.md. Push without asking — `git push`
|
|
501
|
+
is in the settings `allow` list precisely so a one-shot run never stalls at the last step.
|
|
502
|
+
Force-push stays denied, so this can only ever fast-forward.
|
|
306
503
|
|
|
307
|
-
**
|
|
504
|
+
**If some tasks are `blocked`:** still push. The approved work is reviewed, tested, and
|
|
505
|
+
committed; withholding it helps no one, and the per-wave commits make any subset revertible.
|
|
506
|
+
Name the blocked tasks in the Final Report.
|
|
308
507
|
|
|
309
|
-
|
|
508
|
+
Finally set `Phase: done` in `docs/AI_HANDOFF/RUN.md`.
|
|
310
509
|
|
|
311
510
|
---
|
|
312
511
|
|
|
@@ -337,7 +536,19 @@ Re-run assertions after removal to confirm clean state.
|
|
|
337
536
|
handoff-fullstack complete:
|
|
338
537
|
Cycle: <ID>
|
|
339
538
|
Tasks: <N> approved, <M> blocked
|
|
340
|
-
|
|
539
|
+
Fix rounds used: <0|1|2>
|
|
540
|
+
Git: <K> wave commits + pushed to <branch>
|
|
341
541
|
Worktrees: all cleaned
|
|
342
542
|
Branches: all cleaned
|
|
543
|
+
Blocked (needs you): <TASK-xxx — one-line reason> | none
|
|
343
544
|
```
|
|
545
|
+
|
|
546
|
+
Then, as the last line of the run:
|
|
547
|
+
|
|
548
|
+
> Cycle finished. Run `/compact` before starting the next cycle — this session is carrying
|
|
549
|
+
> the whole pipeline's history and the next cycle deserves a clean window.
|
|
550
|
+
|
|
551
|
+
This is the **only** place in the pipeline that may ask for a compaction. Asking at a cycle
|
|
552
|
+
boundary costs nothing: all state is in git, `INDEX.md` and `RUN.md`, so a compacted or
|
|
553
|
+
brand-new session picks the next cycle up with no loss. Asking mid-cycle is forbidden — see
|
|
554
|
+
the Autonomy Contract.
|
|
@@ -8,7 +8,48 @@
|
|
|
8
8
|
$ARGUMENTS
|
|
9
9
|
_Empty = all `ready` tasks. Or: "TASK-001" for a specific task._
|
|
10
10
|
|
|
11
|
-
> **
|
|
11
|
+
> **Commits: yes, one per wave. Push: no.** Each wave is checkpointed to git so any step is
|
|
12
|
+
> revertible, but nothing leaves the machine — `/ukit:handoff-review` decides what gets pushed.
|
|
13
|
+
|
|
14
|
+
---
|
|
15
|
+
|
|
16
|
+
## Autonomy Contract — read before anything else
|
|
17
|
+
|
|
18
|
+
Run every `ready` task to completion in this one invocation. The human is not watching and
|
|
19
|
+
will not answer mid-run.
|
|
20
|
+
|
|
21
|
+
**Ask nothing after the run starts.** Every decision below has a defined automatic
|
|
22
|
+
resolution — take it, and note it in the wave summary:
|
|
23
|
+
|
|
24
|
+
| Situation | Required behavior |
|
|
25
|
+
|-----------|------------------|
|
|
26
|
+
| Working tree dirty at Step 1 | commit a checkpoint, continue |
|
|
27
|
+
| Two same-wave tasks share a file | move the later task to the next wave, continue |
|
|
28
|
+
| A task's agent reports FAIL | leave it `blocked`, finish every other task, report at the end |
|
|
29
|
+
| Context warning fires | collapse the finished wave to one line per task, continue |
|
|
30
|
+
| No `ready` tasks but `changes_requested` ones exist | treat those as the work — this is a fix pass |
|
|
31
|
+
|
|
32
|
+
Escalate only for a blocker outside the repo. Never pause mid-wave to ask for `/compact`:
|
|
33
|
+
context is managed structurally (short agent returns + wave-boundary collapse), and the run
|
|
34
|
+
cursor makes an interrupted run resumable. A compaction request belongs at the end of the
|
|
35
|
+
command, never inside it.
|
|
36
|
+
|
|
37
|
+
### Run cursor — write after every step
|
|
38
|
+
|
|
39
|
+
Maintain `docs/AI_HANDOFF/RUN.md`, rewritten (whole file) after every numbered step:
|
|
40
|
+
|
|
41
|
+
```
|
|
42
|
+
Command: handoff-implement
|
|
43
|
+
Goal: <one sentence>
|
|
44
|
+
Base: <BASE>
|
|
45
|
+
Phase: <1|2|3|4|done>
|
|
46
|
+
Cursor: wave <N> batch <M> — <what just finished>
|
|
47
|
+
Next: <the exact next step to run>
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
If this file already exists with `Phase:` not `done` when the command starts, this is a
|
|
51
|
+
**continuation**: read the cursor, jump to `Next:`, skip tasks already `pending_review` or
|
|
52
|
+
`done`. Do not restart from Step 1 and do not ask whether to continue.
|
|
12
53
|
|
|
13
54
|
---
|
|
14
55
|
|
|
@@ -16,14 +57,19 @@ _Empty = all `ready` tasks. Or: "TASK-001" for a specific task._
|
|
|
16
57
|
|
|
17
58
|
Read `docs/AI_HANDOFF/ACTIVE.md` → get `Base: <BASE>`.
|
|
18
59
|
Read `docs/AI_HANDOFF/INDEX.md` → collect `ready` tasks (or specific task from $ARGUMENTS).
|
|
19
|
-
If no ready tasks
|
|
60
|
+
If there are no `ready` tasks, check for `changes_requested`/`blocked` ones and run those
|
|
61
|
+
instead. Only if nothing is actionable → report and stop.
|
|
20
62
|
|
|
21
|
-
Verify working tree
|
|
63
|
+
Verify working tree state:
|
|
22
64
|
```bash
|
|
23
|
-
git status
|
|
65
|
+
git status
|
|
24
66
|
```
|
|
25
67
|
|
|
26
|
-
If working tree is dirty
|
|
68
|
+
If the working tree is dirty, do not stop — commit a checkpoint so the wave copy-back starts
|
|
69
|
+
from a known state and the pre-existing edits stay recoverable:
|
|
70
|
+
```bash
|
|
71
|
+
git add -A && git commit -m "handoff: checkpoint before implement"
|
|
72
|
+
```
|
|
27
73
|
|
|
28
74
|
## Step 2 — Build wave groups
|
|
29
75
|
|
|
@@ -33,9 +79,22 @@ Read each `tasks/TASK-xxx.md` for `Dependencies` field:
|
|
|
33
79
|
- Chain A→B→C = 3 waves of 1 task each (sequential, no parallel)
|
|
34
80
|
- Independent A, B, C = 1 wave of 3 tasks (parallel)
|
|
35
81
|
|
|
82
|
+
### Conflict check — mandatory, before spawning any wave
|
|
83
|
+
|
|
84
|
+
For every pair of tasks landing in the same wave, compare their `Target Files` lists. If any
|
|
85
|
+
file path appears in both, **resolve it automatically — do not stop.** Keep the
|
|
86
|
+
lower-numbered task in the current wave and push the other into the next wave by adding
|
|
87
|
+
`Dependencies: TASK-<lower>` to its task file. Two tasks touching the same file are safe as
|
|
88
|
+
long as they never run concurrently, and serializing them is the one resolution that is
|
|
89
|
+
always correct.
|
|
90
|
+
|
|
91
|
+
This is the code-level safety net for the planner's own §2 Scope constraint (same-wave tasks
|
|
92
|
+
must not share a file) in case it slipped through review. Only mark a task `needs_breakdown`
|
|
93
|
+
when it is missing required fields — never merely for sharing a file.
|
|
94
|
+
|
|
36
95
|
### Batch each wave — mandatory
|
|
37
96
|
|
|
38
|
-
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **
|
|
97
|
+
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). A wave
|
|
39
98
|
with more tasks than that is split into consecutive batches of at most that many; finish
|
|
40
99
|
one batch completely (including 3c copy-back and worktree deletion) before starting the
|
|
41
100
|
next.
|
|
@@ -68,7 +127,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
|
|
|
68
127
|
TDD — mandatory:
|
|
69
128
|
1. Write tests from §Test Cases
|
|
70
129
|
cd .worktrees/task-xxx && <test command>
|
|
71
|
-
Confirm RED
|
|
130
|
+
Confirm RED — paste the actual failing output, don't just assert it happened
|
|
131
|
+
(immediately GREEN → test is wrong — flag this)
|
|
72
132
|
2. Implement → run → confirm GREEN
|
|
73
133
|
3. Run §Verification Commands inside the worktree:
|
|
74
134
|
cd .worktrees/task-xxx && <each verification command>
|
|
@@ -81,9 +141,24 @@ Executor Report (append to task file — do NOT touch INDEX.md):
|
|
|
81
141
|
EXECUTOR_TOOL: <tool>
|
|
82
142
|
EXECUTOR_MODEL: <exact model ID — mandatory>
|
|
83
143
|
EXECUTOR_SUBAGENT: <name or "-">
|
|
144
|
+
RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
|
|
145
|
+
"confirmed" without pasted output is not acceptable>
|
|
84
146
|
Verification Output: <paste>
|
|
85
147
|
Status: PASS | FAIL
|
|
86
148
|
Note: <issues or "none">
|
|
149
|
+
|
|
150
|
+
Then RETURN TO THE ORCHESTRATOR AT MOST 10 LINES, in exactly this shape:
|
|
151
|
+
TASK: TASK-xxx
|
|
152
|
+
STATUS: PASS | FAIL
|
|
153
|
+
EXECUTOR_MODEL: <exact model ID>
|
|
154
|
+
FILES: <comma-separated changed paths>
|
|
155
|
+
RED: confirmed | not-confirmed
|
|
156
|
+
VERIFY: <n> commands, all pass | <first failing command + one-line reason>
|
|
157
|
+
NOTE: <one line or "none">
|
|
158
|
+
|
|
159
|
+
Do NOT repeat RED_OUTPUT or verification logs in the returned message. They are already on
|
|
160
|
+
disk in the task file, which is where the reviewer reads them from. Pasting them back a
|
|
161
|
+
second time is what blows up the orchestrator's context window and kills the run mid-wave.
|
|
87
162
|
```
|
|
88
163
|
|
|
89
164
|
> Other tools without subagent support: open each task in a separate session with the code model.
|
|
@@ -132,21 +207,44 @@ git branch -D handoff/task-xxx
|
|
|
132
207
|
**Orchestrator writes INDEX.md** — agents never touch INDEX directly.
|
|
133
208
|
**Worktrees are always deleted immediately** — no exceptions.
|
|
134
209
|
|
|
135
|
-
### 3d —
|
|
136
|
-
|
|
210
|
+
### 3d — Wave boundary: commit + context checkpoint
|
|
211
|
+
|
|
212
|
+
A wave boundary is the only safe place to shed context, because everything of value is
|
|
213
|
+
already on disk (task files hold the full logs, git holds the code). Do all four, in order:
|
|
214
|
+
|
|
215
|
+
1. **Checkpoint the code** — one commit per wave, so every wave is independently revertible:
|
|
216
|
+
```bash
|
|
217
|
+
git add -A && git commit -m "handoff: wave <N> — TASK-00x, TASK-00y"
|
|
218
|
+
```
|
|
219
|
+
Do not push. `/ukit:handoff-review` owns that decision.
|
|
220
|
+
|
|
221
|
+
2. **Update the run cursor** — rewrite `docs/AI_HANDOFF/RUN.md` with `Cursor: wave <N> done`,
|
|
222
|
+
`Next: wave <N+1>` (or `Step 4` if that was the last wave).
|
|
223
|
+
|
|
224
|
+
3. **Collapse the wave in working memory.** From here on, refer to the finished wave only by
|
|
225
|
+
its one-line-per-task summary. Do not re-read task files, restate executor reports, or
|
|
226
|
+
quote earlier diffs. The reviewer reads what it needs from disk itself.
|
|
227
|
+
|
|
228
|
+
4. **Continue.** If a UKit context warning has fired, say one line —
|
|
229
|
+
`context checkpoint: wave <N> collapsed` — and start the next wave anyway. Do not ask for
|
|
230
|
+
`/compact` here. If the session ends regardless, `RUN.md` plus the `SessionStart` resume
|
|
231
|
+
hook bring the run back to exactly this point.
|
|
232
|
+
|
|
233
|
+
Then create the next wave's worktrees from `$BASE` and repeat 3a–3c.
|
|
137
234
|
|
|
138
235
|
## Step 4 — Finalize: verify cleanup
|
|
139
236
|
|
|
140
237
|
All waves complete. Verify:
|
|
141
238
|
|
|
142
239
|
```bash
|
|
143
|
-
git worktree list
|
|
240
|
+
git worktree list # must show only main worktree
|
|
144
241
|
git branch | grep handoff # must be empty
|
|
145
|
-
git
|
|
146
|
-
git
|
|
242
|
+
git log --oneline -n 10 # one commit per wave
|
|
243
|
+
git status # clean, or only trailing edits from the last wave
|
|
147
244
|
```
|
|
148
245
|
|
|
149
|
-
All changes are
|
|
246
|
+
All changes are committed as **one commit per wave on `$BASE`**, nothing pushed. No branches
|
|
247
|
+
or worktrees remain. Set `Phase: done` in `docs/AI_HANDOFF/RUN.md`.
|
|
150
248
|
|
|
151
249
|
## Step 5 — Report
|
|
152
250
|
|
|
@@ -154,8 +252,13 @@ All changes are in the **main working tree as uncommitted files**. No branches o
|
|
|
154
252
|
Wave summary:
|
|
155
253
|
Wave 1: TASK-001 [pending_review], TASK-002 [blocked]
|
|
156
254
|
Wave 2: TASK-003 [pending_review]
|
|
255
|
+
Commits: <K> wave commits on <BASE>, not pushed
|
|
157
256
|
All worktrees and branches removed.
|
|
158
|
-
|
|
257
|
+
Blocked (needs you): <TASK-xxx — one-line reason> | none
|
|
159
258
|
```
|
|
160
259
|
|
|
161
260
|
**Next:** switch to strong model → `/ukit:handoff-review`
|
|
261
|
+
|
|
262
|
+
If this session is now heavy, this is the right moment to `/compact` — the wave commits,
|
|
263
|
+
`INDEX.md` and `RUN.md` hold all the state the review phase needs. Suggest it here, at the
|
|
264
|
+
command boundary, never inside a wave.
|