@ngockhoale/ukit 2.0.6 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,47 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.1.0 - 2026-08-16
6
+
7
+ Handoff autonomy wave: `/ukit:handoff-fullstack` runs a whole goal end-to-end without stopping to ask. Questions are collected once, up front, in the planning phase; everything that used to be a mid-run STOP now resolves automatically and keeps going.
8
+
9
+ ### Added
10
+
11
+ - **Autonomy Contract in all four handoff commands.** `handoff-create` owns the single question window — one batched `AskUserQuestion` (max 4, recommended option first) before planning starts, covering scope, conflicting interpretations, subsystem choice and acceptance criteria only. `handoff-fullstack`, `handoff-implement` and `handoff-review` ask nothing between that window and the Final Report; each command carries a table mapping every former STOP to its required automatic behavior.
12
+ - **Run cursor + resume protocol (`docs/AI_HANDOFF/RUN.md`).** Every phase transition writes `Phase`/`Cursor`/`Next`. A `Phase:` other than `done` means the next session is a continuation: no re-plan, no re-scope, no questions. New config `handoff.autonomy.resumeFromCursor` / `runCursorFile`.
13
+ - **`handoff-resume.sh` SessionStart hook** (manifest id `hook-handoff-resume`). Reads the run cursor and prints a resume banner with Goal/Base/Phase/Cursor/Next, plus two extra lines when the session started from `source: compact`. Advisory only, always exits 0, silent when there is no unfinished run.
14
+ - **Auto-fix loop after review (max 2 rounds).** Non-approved tasks are re-waved and re-implemented in parallel — agents read their own `## Reviewer Verdict` from disk — then committed and re-reviewed. After round 2 the stuck tasks stay `blocked` and the rest of the pipeline proceeds instead of halting. Config: `handoff.autonomy.autoFixRounds`.
15
+ - **Planner Grounding + 12-point Self-Audit.** `handoff-planner` must verify that every Target File exists (or is marked `(new)`), that every Verification Command is a real `package.json` script, and that test file paths match the neighbouring test layout; then run a 12-point self-audit (coverage / correctness / test quality, including "could every listed test pass against an empty implementation?") and report the result. Compensates for dropping plan review from 3 rounds to 2.
16
+ - **Wave-boundary commit + context checkpoint.** Each wave ends with `git add -A && git commit -m "handoff: wave <N> — TASK-00x, ..."`, a cursor update, and a collapse of that wave to one line per task in working memory. Config: `handoff.autonomy.commitPerWave`.
17
+
18
+ ### Changed
19
+
20
+ - **`git push` moved from the `ask` list to `allow`** in default settings, alongside `git add`, `git commit` and `git worktree`. Runs no longer stop at the final step for a confirmation. `git push --force` and `--force-with-lease` remain in `deny`, which still wins over `allow`.
21
+ - **Subagent return contracts are now capped.** Executors return ≤10 lines (`TASK`/`STATUS`/`EXECUTOR_MODEL`/`FILES`/`RED`/`VERIFY`/`NOTE`), reviewers ≤6. Verbose agent reports injected back into the orchestrator were the main cause of context exhaustion on long fullstack runs; this addresses it structurally rather than by raising caps.
22
+ - **`context-window-guard.sh` is now run-cursor aware.** With an unfinished run it emits shed-context-and-continue instructions instead of the stop-and-`/compact` directive. Asking the user to `/compact` is permitted only at a cycle boundary, in the Final Report — never mid-task. Config: `handoff.autonomy.compactAtCycleBoundaryOnly`.
23
+ - **Same-wave file conflicts auto-serialize** instead of marking both tasks `needs_breakdown`: the higher-numbered task gains a `Dependencies: TASK-<lower>` line and moves to the next wave. Config: `handoff.autonomy.autoSerializeFileConflicts`.
24
+ - **A dirty working tree no longer blocks Phase 3.** It is committed as `handoff: checkpoint before implement` and the run continues.
25
+ - **Plan review loop cap: 3 rounds → 2**, then findings are applied and the run proceeds (logged as `### Round <N> — findings applied without re-review`). Config: `handoff.autonomy.planReviewRounds`.
26
+ - **Review diff range is now anchored to the plan commit** (`git log --grep='^handoff: plan' -n 1`) rather than the working tree, since waves are committed as they finish.
27
+ - **`handoff-planner` maximizes wave width.** `Dependencies: none` unless task B genuinely cannot be written without A's output — ordering preference, same code area and review order are explicitly not dependencies.
28
+ - **`feature-implementer` and `code-reviewer` run unattended in handoff mode.** Ambiguity resolves through task file → PLAN.md → surrounding code → recommended choice, recorded in `## Discussion`; reviewer findings must be actionable defects (`file:line` + concrete failure + what correct looks like), never questions. Daily (non-handoff) behavior is unchanged.
29
+ - **`handoff-model-guard.sh` gained a bash fast path** that skips the node spawn when the payload contains neither `AI_HANDOFF` nor `git push`. Behavior-identical (verified against the previous version across a 6-payload sweep); measured 67ms → 38ms per ordinary Bash/Edit/Write call, paid on every call in every parallel subagent.
30
+
31
+ ## 2.0.7 - 2026-08-15
32
+
33
+ ### Added
34
+
35
+ - **`/ukit:handoff-create` and `/ukit:handoff-fullstack` gain a P2.5 independent plan review gate.** After the planner writes `PLAN.md` and before it's committed, a separate `code-reviewer` agent invocation (`REVIEW_TARGET_TYPE=plan`, fresh context — not the planner reviewing its own work) checks Completeness/Consistency/Clarity/Scope/YAGNI. `Issues Found` routes back to the planner to revise and resubmit; `Approved` unlocks the commit. Each round is logged to PLAN.md's new `## Plan Review Log` section, and the loop is capped at 3 `Issues Found` rounds — past that it escalates to the human instead of looping indefinitely.
36
+ - **Same-wave file-conflict precheck in Phase 3.** Before spawning any implementation wave, `handoff-implement.md`/`handoff-fullstack.md` now compare `Target Files` across every task in that wave; any overlap stops the wave and marks both tasks `needs_breakdown` instead of risking a mid-wave merge conflict. Code-level backstop for the planner's existing "no shared files in one wave" scope constraint.
37
+ - **TDD RED-state now requires real evidence.** The Executor Report template adds a mandatory `RED_OUTPUT` field — the actual failing-test output pasted before implementation, not a bare "confirmed" claim. `code-reviewer` checks this field during Test Plan adherence review and requests changes if it's missing or vague.
38
+ - **`handoff.plan.requireLintOrTypecheckInVerification`** (default `true`): when the project has a lint/typecheck script, task Verification Commands must run it, not just tests; the planner must state an explicit N/A if the project has none. Enforced by `code-reviewer` as review-order step 0.
39
+
40
+ ### Changed
41
+
42
+ - **`handoff.maxParallelAgents`**: 3 → 10, raising the default cap on concurrent background agents per wave (config comments still warn against exceeding ~10-15 — each agent's report gets injected back into the orchestrator's own context on completion).
43
+ - **Phase 4 (Review) is now parallelized**, batched by `maxParallelAgents` the same way Phase 3 (Implement) already was — reviewer agents only read a diff and append a verdict to their own task file, so parallel review carries none of Phase 3's shared-worktree conflict risk.
44
+ - **`handoff.plan.minTestsEdgeCase`**: 1 → 2, and the two edge cases must now be of different kinds (e.g. null/empty **and** boundary/concurrent) — two near-duplicate cases no longer satisfy the requirement. Wired through `handoff-planner`, `code-reviewer`, `feature-implementer`'s inline-test-plan fallback, `RULES.md`, and both handoff pipeline commands.
45
+
5
46
  ## 2.0.6 - 2026-08-12
6
47
 
7
48
  ### Changed
@@ -1121,6 +1121,17 @@ items:
1121
1121
  packs:
1122
1122
  - core
1123
1123
 
1124
+ - id: hook-handoff-resume
1125
+ type: hook
1126
+ sourceTemplate: .claude/hooks/handoff-resume.sh
1127
+ targetPath: .claude/hooks/handoff-resume.sh
1128
+ requires: []
1129
+ mergeStrategy: overwrite_with_backup
1130
+ variables: []
1131
+ enabledByDefault: true
1132
+ packs:
1133
+ - core
1134
+
1124
1135
  - id: hook-reset-compact-pressure
1125
1136
  type: hook
1126
1137
  sourceTemplate: .claude/hooks/reset-compact-pressure.sh
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.0.6",
3
+ "version": "2.1.0",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, Antigravity, OpenAI Codex, and OpenCode.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -27,7 +27,8 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
27
27
 
28
28
  ### Review order
29
29
 
30
- 1. **Test Plan adherence** — Were all tests in §4 actually implemented? Run them yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
30
+ 0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
31
+ 1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
31
32
  2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
32
33
  3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider test suite if shared code was touched.
33
34
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
@@ -66,6 +67,27 @@ NOTES: [1-2 sentences for human reviewer if needed]
66
67
 
67
68
  After writing the verdict, update `docs/AI_HANDOFF/INDEX.md` row for this task: set Status = NEXT_STATUS_FOR_INDEX, set Reviewer = your model name.
68
69
 
70
+ **Then return to the orchestrator at most 6 lines** — the full verdict is already on disk:
71
+
72
+ ```
73
+ TASK: TASK-xxx
74
+ VERDICT: approved | approved_minor | changes_requested | critical_block
75
+ REVIEWER_MODEL: <exact model ID>
76
+ VERIFICATION_RERUN: PASS | FAIL
77
+ BLOCKING: <one line per critical/important finding, or "none">
78
+ ```
79
+
80
+ Do not paste the diff, the findings prose, or verification output into the returned message.
81
+ The orchestrator is driving a whole pipeline in one context window and re-reads what it needs
82
+ from the task file; pasted reviewer logs are a common reason a run exhausts its context and
83
+ dies before the cycle finishes.
84
+
85
+ **You may be running unattended.** Findings do not end the run — the caller feeds them to an
86
+ auto-fix round and re-review. So write findings that a fresh executor can act on with no human
87
+ present: point at `file:line`, state the concrete failure, and say what correct looks like. A
88
+ finding phrased as a question ("should this handle null?") is not actionable; phrase it as the
89
+ defect ("`parse()` at foo.js:41 throws on null input; expected an empty result").
90
+
69
91
  ### Model isolation check (FIRST thing you do)
70
92
 
71
93
  UKit cannot force any tool to use a specific model. The contract is enforced HERE, by you, via self-report comparison.
@@ -112,7 +134,10 @@ Only flag issues that would cause real problems during implementation planning.
112
134
 
113
135
  ### Output
114
136
 
137
+ Append (do NOT overwrite) this block to the end of the reviewed document under a `## Plan Review Log` section — create the section if it doesn't exist yet, keep all prior round entries:
138
+
115
139
  ```
140
+ ### Round <N> — <YYYY-MM-DD> · <your model>
116
141
  Status: Approved | Issues Found
117
142
 
118
143
  COMPLETENESS:
@@ -128,3 +153,5 @@ YAGNI:
128
153
 
129
154
  NOTES: [1-2 sentences if needed]
130
155
  ```
156
+
157
+ `<N>` = 1 + however many `### Round` entries already exist in the log (1 if this is the first review).
@@ -12,7 +12,15 @@ Implement requested behavior with minimal scope drift.
12
12
  - **Daily/ad-hoc mode** (DEFAULT): task didn't come from `docs/AI_HANDOFF/` → use the original lightweight workflow. Tests only when touched code already has coverage. No reviewer trigger.
13
13
  - **Handoff mode**: task file is `docs/AI_HANDOFF/tasks/TASK-xxx.md` OR user explicitly invokes handoff (e.g. "execute task TASK-001") → activate full Quality Gate: test-first → green → reviewer.
14
14
 
15
- If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
15
+ If unsure which mode applies, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
16
+
17
+ **In Handoff mode you are running unattended — ask nothing.** You were spawned by an
18
+ orchestrator driving a pipeline; there is no human in your conversation to answer, and a
19
+ question there is silently dropped while the run stalls. Resolve ambiguity in this order:
20
+ the task file → `PLAN.md` → the surrounding code's existing patterns → the choice you would
21
+ recommend. Record what you chose and why in the task's `## Discussion` thread. Only a blocker
22
+ outside the repo (missing credential, unreachable service) justifies reporting `FAIL` early —
23
+ and even then, report it, don't ask about it.
16
24
 
17
25
  ## Workflow
18
26
 
@@ -25,12 +33,12 @@ If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
25
33
  - non-trivial: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
26
34
  - Identify target files and existing patterns.
27
35
  - If task came from handoff, read `tasks/TASK-xxx.md` and locate its **Test Plan** + **Verification Commands**.
28
- - If confidence is low or risk is high, ask one short clarifying question before deeper analysis.
36
+ - Daily mode: if confidence is low or risk is high, ask one short clarifying question before deeper analysis. Handoff mode: do not ask — decide and record the decision (see above).
29
37
 
30
38
  ### 2. Plan Approach (< 1 minute)
31
39
 
32
40
  - List files to create/modify (max diff).
33
- - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥1 edge case; regression test if fixing a bug). In daily mode, skip this step.
41
+ - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
34
42
 
35
43
  ### 3. Test First (RED) — Handoff mode
36
44
 
@@ -79,6 +87,26 @@ NEXT: [follow-up needed, or "ready for review"]
79
87
 
80
88
  > **Self-report rule:** UKit cannot force any tool/host to use a specific model. Your self-reported `EXECUTOR_MODEL` is how the reviewer (in another tool or subagent) knows what to compare against its own model. Misreporting → reviewer refuses and asks the human to confirm.
81
89
 
90
+ **Handoff mode — where that report goes.** Append the block above, in full, to the task file
91
+ as `## Executor Report` (or `## Executor Report (fix round <N>)` on a re-run). Then return to
92
+ the orchestrator **at most 10 lines**:
93
+
94
+ ```
95
+ TASK: TASK-xxx
96
+ STATUS: PASS | FAIL
97
+ EXECUTOR_MODEL: <exact model ID>
98
+ FILES: <comma-separated changed paths>
99
+ RED: confirmed | not-confirmed
100
+ VERIFY: <n> commands, all pass | <first failing command + one-line reason>
101
+ NOTE: <one line, or "none">
102
+ ```
103
+
104
+ Do not repeat RED output, verification logs, diffs, or file contents in the returned message.
105
+ The reviewer reads all of that from the task file on disk. The orchestrator is driving a whole
106
+ pipeline in one context window, and pasted subagent logs are the single biggest reason a run
107
+ runs out of context and dies half-finished. Being terse here is not a style preference — it is
108
+ what lets the cycle reach the end.
109
+
82
110
  ### 7. Trigger Reviewer — Handoff mode ONLY
83
111
 
84
112
  - Daily mode: skip this step entirely. Just report and stop.
@@ -17,15 +17,43 @@ Use the strongest model available — planning with a weak model produces weak t
17
17
 
18
18
  ## Scope Check (before Phase 1)
19
19
 
20
- Before refining the request into `PLAN.md`, check whether it describes multiple independent subsystems (e.g. "CRM + AI + billing + analytics + mobile app"). If yes, stop and output a decomposition table instead of planning the mega-spec:
20
+ Before refining the request into `PLAN.md`, check whether it describes multiple independent subsystems (e.g. "CRM + AI + billing + analytics + mobile app"). If yes, record the decomposition:
21
21
 
22
22
  ```
23
23
  Scope complexity: HIGH
24
24
  Detected systems: [...]
25
- Recommended decomposition: N modules — plan each as a separate PLAN.md
25
+ Decomposition: N modules — module 1 planned now, modules 2..N queued
26
26
  ```
27
27
 
28
- Wait for human confirmation before writing `PLAN.md`.
28
+ Then **plan module 1 in full and keep going** — do not wait for confirmation. A mega-spec is
29
+ what you must avoid, not the work itself. Write the remaining modules into `PLAN.md` §2 as
30
+ explicitly out-of-scope-for-this-cycle, and add one `queued` row per module to `INDEX.md` so
31
+ the next cycle picks them up. Order modules so the one others depend on is planned first.
32
+
33
+ The caller may be running unattended (`/ukit:handoff-fullstack`). Stopping to ask means the
34
+ run dies and the human returns hours later to a question, which is the failure mode this
35
+ whole pipeline exists to prevent.
36
+
37
+ ## Grounding — every path and command must be real
38
+
39
+ The most common way a plan fails is not bad reasoning, it is **confident invention**: a target
40
+ file that doesn't exist, a test command the project doesn't have, an import path that never
41
+ resolved. Executors then burn a whole round discovering it.
42
+
43
+ Before writing any path or command into `PLAN.md` or a task file, verify it:
44
+
45
+ - **Target Files** — for each path, confirm it exists (`ls`/Glob), or that its parent
46
+ directory exists and the file is genuinely new. Mark new files `(new)` explicitly.
47
+ - **Verification Commands** — read `package.json` `scripts` (or the stack's equivalent) and
48
+ use the script names that are actually defined. Do not write `npm test` when the repo uses
49
+ `yarn test`, and do not invent a `lint` script that isn't there. Run `--help` or a dry check
50
+ if unsure.
51
+ - **Test Files** — follow the existing test layout and naming; open one neighbouring test file
52
+ and match its structure, framework and import style.
53
+ - **Interfaces** — quote real signatures from the source, not plausible-looking ones.
54
+
55
+ If something cannot be verified, say so in the task's `## Discussion` rather than guessing.
56
+ A stated unknown costs the executor one read; a wrong path costs it a round.
29
57
 
30
58
  ## Phase 1 — Write PLAN.md
31
59
 
@@ -37,9 +65,15 @@ Write all 7 sections to `docs/AI_HANDOFF/PLAN.md`:
37
65
  CONSTRAINT: tasks in the same wave must not modify the same file.
38
66
  If two tasks need the same file → make one depend on the other.
39
67
  §3 Approach — technical solution, trade-offs, alternatives rejected
40
- §4 Test Plan — happy path × N + ≥1 edge case + regression (if bugfix)
68
+ §4 Test Plan — happy path × N + ≥2 edge cases of DIFFERENT kinds (e.g. null/empty AND
69
+ boundary/concurrent — two near-duplicate cases do not satisfy this) +
70
+ regression (if bugfix)
41
71
  table: | Type | Test Name | Expected |
42
- §5 Verification — exact shell commands executor will run
72
+ §5 Verification — exact shell commands executor will run. If the project has a lint or
73
+ typecheck script (check package.json `scripts`, or the equivalent for
74
+ the project's stack), it MUST be included here, not just the test
75
+ command. If the project genuinely has none, state that explicitly —
76
+ do not omit silently.
43
77
  §6 Acceptance — checklist of done criteria (prefer verifiable/command-based criteria)
44
78
  §7 Global Constraints — one line each: version floors, dependency limits, naming/copy
45
79
  rules, platform requirements. Every TASK-xxx.md inherits this section
@@ -55,6 +89,8 @@ Append this footer to `PLAN.md` — mandatory, checked by a hook before the writ
55
89
  PLANNER_MODEL: <your exact model ID — e.g. claude-opus-5>
56
90
  ```
57
91
 
92
+ Your output does not go straight to implementation: an independent `code-reviewer` pass (`REVIEW_TARGET_TYPE=plan`) reviews `PLAN.md` next. If it returns `Issues Found`, you'll be re-invoked to revise and resubmit — write §1-§7 tight enough to pass on the first pass.
93
+
58
94
  ## Phase 2 — Split into TASK-xxx.md
59
95
 
60
96
  **Right-sizing rule:** A task is the smallest unit that carries its own test cycle and is worth a fresh reviewer's gate. Split only where a reviewer could meaningfully approve one task while rejecting its neighbor. Each task ends with an independently testable deliverable.
@@ -67,9 +103,9 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
67
103
  |-------|------|
68
104
  | Target Files | Exact paths — no two tasks in same wave share a file |
69
105
  | Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
70
- | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥1 edge case |
106
+ | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
71
107
  | Test Files | Exact test file paths to create/modify |
72
- | Verification Commands | Runnable shell commands |
108
+ | Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
73
109
  | Acceptance Criteria | Verifiable checklist |
74
110
 
75
111
  Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
@@ -80,6 +116,30 @@ Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
80
116
  - Chain: A → B → C runs as 3 sequential waves (1 task each, no parallelism)
81
117
  - Independent: A, B, C (all `none`) runs as 1 wave, all parallel
82
118
 
119
+ ### Maximize wave width — dependencies are expensive
120
+
121
+ Wave width is the single biggest lever on how long a cycle takes: a wave of 6 finishes in
122
+ roughly the time of its slowest task, while a chain of 6 takes six times that. Executors run
123
+ up to `handoff.maxParallelAgents` (default 10) at once, so a plan that produces `none`
124
+ dependencies for most tasks is dramatically faster than one that produces a chain.
125
+
126
+ **Write `Dependencies: none` unless B genuinely cannot be written without A's output.** A real
127
+ dependency means B imports a symbol A creates, or B tests behavior A implements. These are
128
+ *not* dependencies:
129
+
130
+ - "B is logically later" or "B builds on the same feature" — ordering preference, not a
131
+ dependency.
132
+ - "Both touch the same area of the codebase" — irrelevant unless they touch the same *file*.
133
+ - "A should be reviewed before B starts" — that is what the review phase is for.
134
+
135
+ Same-file collisions are the one real constraint, and they are cheap to design around: split
136
+ along file boundaries so each task owns its files outright. If two pieces of work truly must
137
+ edit one file, prefer merging them into a single task over chaining two — one task with two
138
+ test groups beats two waves.
139
+
140
+ Before finalizing, count your waves. If the dependency graph is mostly a chain, re-examine it:
141
+ most chains are ordering preferences that a wide wave 1 would satisfy just as well.
142
+
83
143
  ## Phase 3 — Update state files
84
144
 
85
145
  **INDEX.md** row per task: `| TASK-001 | <name> | ready | none | - |`
@@ -93,15 +153,59 @@ Status: planning_done — ready for executor
93
153
  ```
94
154
  Wave structure is NOT stored here — inferred from task `Dependencies` fields at runtime.
95
155
 
156
+ ## Self-Audit — run before reporting, every time
157
+
158
+ The downstream pipeline is unattended: an incomplete or wrong plan is not caught by a human
159
+ skimming it, it is caught by an executor failing three hours later. Independent plan review
160
+ runs at most twice, so this audit is your own last line of defense.
161
+
162
+ Walk the checklist and **fix what fails before you report** — do not report and hope review
163
+ catches it:
164
+
165
+ **Coverage — is anything missing?**
166
+ 1. Does every acceptance criterion in §6 trace to at least one task? Name the task per criterion.
167
+ 2. Does every task trace back to something in §1/§6? A task nothing asks for is scope creep — cut it.
168
+ 3. Do the tasks together actually deliver §1's success definition, or only the easy part of it? State the gap if there is one.
169
+ 4. Is the *unhappy* path planned — errors, empty input, permissions, migration of existing data — or only the feature?
170
+
171
+ **Correctness — is anything wrong?**
172
+ 5. Every `Target Files` path verified per Grounding? Any `(new)` file marked as such?
173
+ 6. Every `Verification Command` a real, defined script in this repo?
174
+ 7. Any two same-wave tasks sharing a file? If yes, merge them or add the dependency now — the executor will otherwise serialize them for you and your wave plan was a lie.
175
+ 8. Does any task depend on a symbol or file that no earlier task creates? That's a missing task, not a dependency.
176
+
177
+ **Test quality — the part most often faked:**
178
+ 9. Each task: ≥1 happy path + ≥2 edge cases of genuinely *different kinds*. "empty string" and "null" are the same kind — boundary, concurrency, malformed input, permission denied, duplicate, ordering are different kinds.
179
+ 10. Does each test case state a concrete `Expected` value, not "works correctly" or "returns successfully"? A test whose expectation cannot fail is not a test.
180
+ 11. Bugfix task → is there a regression test that fails against today's code?
181
+ 12. Could every listed test pass against an empty implementation? If so, the test cases describe nothing and must be rewritten.
182
+
183
+ Append the result to `PLAN.md`:
184
+ ```
185
+ ## Planner Self-Audit
186
+ Checklist: 12/12 pass
187
+ Fixed during audit: <what you changed, or "nothing">
188
+ Known gaps: <what you deliberately left out and why, or "none">
189
+ ```
190
+
191
+ An honest `Known gaps` line is worth more than a clean sheet — it is the one thing the reviewer
192
+ and executor cannot recover on their own.
193
+
96
194
  ## Output
97
195
 
196
+ Keep the returned message under 25 lines — the caller may be an orchestrator whose context
197
+ budget is the constraint on the whole run. Detail belongs in `PLAN.md`, not in the reply.
198
+
98
199
  - Task count + IDs
99
200
  - Dependency graph (text form: TASK-001 → TASK-003, TASK-002 independent)
201
+ - Wave plan: `wave 1: N tasks | wave 2: M tasks` — flag it if the graph is mostly a chain
202
+ - Self-audit result + any `Known gaps`
100
203
  - Any `needs_breakdown` tasks + reason
101
204
  - Next step: "Switch to Sonnet/unic-code → run `/ukit:handoff-implement`"
102
205
 
103
206
  ## Rules
104
207
 
105
208
  - Do NOT implement. Job ends when all tasks are `ready` or `needs_breakdown`.
106
- - Undone tasks in INDEX.md → warn human: continue old cycle or start fresh?
107
- - Cannot determine test cases → `needs_breakdown` + Discussion thread note.
209
+ - Undone tasks in INDEX.md → continue that cycle; plan only what is still missing. Do not ask whether to start fresh, and do not overwrite task files that already carry an executor report or reviewer verdict.
210
+ - Cannot determine test cases → `needs_breakdown` + Discussion thread note. Never mark a task `ready` with vague tests just to keep the pipeline moving — a task with fake tests passes review and ships a bug, which costs far more than one blocked task.
211
+ - You may be invoked unattended. Ask nothing; where the caller allowed questions, that window closed before you were spawned. Resolve ambiguity by reading the codebase, choose the option you would recommend, and record the choice and its rationale in `PLAN.md §3`.
@@ -14,6 +14,29 @@ $ARGUMENTS
14
14
 
15
15
  ---
16
16
 
17
+ ## Question Policy — this is the one phase that may ask
18
+
19
+ Planning is the **only** place in the handoff pipeline where asking the human is allowed, and
20
+ it is where every ambiguity must be burned off. Whatever is left unresolved here becomes a
21
+ guess during implementation, where nobody is available to correct it.
22
+
23
+ **Ask once, up front, batched.** Before Step 1, if `$ARGUMENTS` leaves a choice that changes
24
+ *what gets built* — not how — collect every such question into a **single** `AskUserQuestion`
25
+ call (max 4 questions, each with a `(Recommended)` first option). Then close the window.
26
+
27
+ Ask about: ambiguous scope boundaries, conflicting interpretations of the request, which
28
+ existing subsystem to extend versus replace, and any acceptance criterion you would otherwise
29
+ invent. Do **not** ask about: file layout, naming, test framework choice, library versions
30
+ already used in the repo, or anything the codebase or `docs/` already answers — resolve those
31
+ by reading, and record the resolution in `PLAN.md §3`.
32
+
33
+ Record every answer verbatim in `PLAN.md §1 Intent`. Downstream phases treat that section as
34
+ the human's own words and never re-ask.
35
+
36
+ After the batched call, run Steps 1–2.5 to completion without further questions.
37
+
38
+ ---
39
+
17
40
  ## Step 1 — Read context (lite model)
18
41
 
19
42
  **Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"`. Do NOT read these files yourself in the current session — this step is contracted to the lite tier (haiku/unic-lite), which only the spawned agent's frontmatter model guarantees. Ask the agent to:
@@ -36,7 +59,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
36
59
 
37
60
  1. Check INDEX.md task statuses:
38
61
  - All tasks are `ready` (planning only) → **re-run allowed**: overwrite PLAN.md and TASK-xxx.md freely — this is iterative refinement.
39
- - Any task is `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **STOP**: warn human — tasks are already in implement/review phase, cannot overwrite safely.
62
+ - Any task is `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **do not overwrite**: that cycle is mid-flight and its task files carry executor reports and verdicts. Report which tasks are in flight and point the human at `/ukit:handoff-review` (to finish the cycle) or `/ukit:handoff-clear` (to abandon it). Planning is the one phase where stopping is correct — there is nothing safe to do automatically here.
40
63
  - No tasks → fresh cycle, proceed normally.
41
64
 
42
65
  2. Resolve base branch:
@@ -48,7 +71,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
48
71
  - §1 Intent — problem + success definition
49
72
  - §2 Scope — in / out of scope. **Add a constraint**: same-wave tasks must not modify the same file (prevents merge conflicts). If two tasks need the same file, make one depend on the other.
50
73
  - §3 Approach — solution, trade-offs, alternatives rejected
51
- - §4 Test Plan — happy path + ≥1 edge case + regression if bugfix (non-negotiable)
74
+ - §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
52
75
  - §5 Verification — exact shell commands executor will run
53
76
  - §6 Acceptance — done checklist (prefer verifiable criteria with commands)
54
77
 
@@ -62,7 +85,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
62
85
  Every task MUST have:
63
86
  - Target Files (exact paths — no two tasks in same wave share a file)
64
87
  - Dependencies (`TASK-xxx` or `none` — wave structure inferred from this, not stored separately)
65
- - Test Cases (Type | Name | Expected — ≥1 happy + ≥1 edge case)
88
+ - Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
66
89
  - Test Files (exact paths)
67
90
  - Verification Commands (runnable shell commands)
68
91
  - Acceptance Criteria (verifiable checklist)
@@ -85,4 +108,25 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
85
108
 
86
109
  ---
87
110
 
111
+ ## Step 2.5 — Independent plan review (strong model, separate agent)
112
+
113
+ **Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet).
114
+
115
+ - count = 0 → run the review below.
116
+ - count = 1 and the round returned `Issues Found` → planner revises once, then run **one** more review round.
117
+ - count ≥ 2 → do NOT invoke the reviewer again, and do not stop. Have the planner apply every outstanding finding directly to `PLAN.md` and the affected `TASK-xxx.md` files, append `### Round <N> — findings applied without re-review` listing what changed, then finish.
118
+
119
+ Two independent strong-model passes shape the plan before any code is written — that gate is
120
+ intact. What it no longer does is hand a stalled plan back and wait.
121
+
122
+ **Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
123
+
124
+ 1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
125
+ 2. `Issues Found` → route back to Step 2: planner revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another Step 2.5 review (this becomes the next round). Do NOT commit or hand off to executor on `Issues Found`.
126
+ 3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed.
127
+
128
+ > Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from Step 2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
129
+
130
+ ---
131
+
88
132
  **Next:** switch to code model → `/ukit:handoff-implement`