@ngockhoale/ukit 2.0.6 → 2.1.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +41 -0
- package/manifests/platform.full.yaml +11 -0
- package/package.json +1 -1
- package/templates/.claude/agents/code-reviewer.md +28 -1
- package/templates/.claude/agents/feature-implementer.md +31 -3
- package/templates/.claude/agents/handoff-planner.md +113 -9
- package/templates/.claude/commands/ukit/handoff-create.md +47 -3
- package/templates/.claude/commands/ukit/handoff-fullstack.md +230 -19
- package/templates/.claude/commands/ukit/handoff-implement.md +117 -14
- package/templates/.claude/commands/ukit/handoff-review.md +104 -33
- package/templates/.claude/hooks/context-window-guard.sh +34 -1
- package/templates/.claude/hooks/handoff-model-guard.sh +12 -0
- package/templates/.claude/hooks/handoff-resume.sh +90 -0
- package/templates/.claude/settings.json +9 -1
- package/templates/docs/AI_HANDOFF/RULES.md +55 -7
- package/templates/ukit/storage/config.json +18 -4
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,47 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.1.0 - 2026-08-16
|
|
6
|
+
|
|
7
|
+
Handoff autonomy wave: `/ukit:handoff-fullstack` runs a whole goal end-to-end without stopping to ask. Questions are collected once, up front, in the planning phase; everything that used to be a mid-run STOP now resolves automatically and keeps going.
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
|
|
11
|
+
- **Autonomy Contract in all four handoff commands.** `handoff-create` owns the single question window — one batched `AskUserQuestion` (max 4, recommended option first) before planning starts, covering scope, conflicting interpretations, subsystem choice and acceptance criteria only. `handoff-fullstack`, `handoff-implement` and `handoff-review` ask nothing between that window and the Final Report; each command carries a table mapping every former STOP to its required automatic behavior.
|
|
12
|
+
- **Run cursor + resume protocol (`docs/AI_HANDOFF/RUN.md`).** Every phase transition writes `Phase`/`Cursor`/`Next`. A `Phase:` other than `done` means the next session is a continuation: no re-plan, no re-scope, no questions. New config `handoff.autonomy.resumeFromCursor` / `runCursorFile`.
|
|
13
|
+
- **`handoff-resume.sh` SessionStart hook** (manifest id `hook-handoff-resume`). Reads the run cursor and prints a resume banner with Goal/Base/Phase/Cursor/Next, plus two extra lines when the session started from `source: compact`. Advisory only, always exits 0, silent when there is no unfinished run.
|
|
14
|
+
- **Auto-fix loop after review (max 2 rounds).** Non-approved tasks are re-waved and re-implemented in parallel — agents read their own `## Reviewer Verdict` from disk — then committed and re-reviewed. After round 2 the stuck tasks stay `blocked` and the rest of the pipeline proceeds instead of halting. Config: `handoff.autonomy.autoFixRounds`.
|
|
15
|
+
- **Planner Grounding + 12-point Self-Audit.** `handoff-planner` must verify that every Target File exists (or is marked `(new)`), that every Verification Command is a real `package.json` script, and that test file paths match the neighbouring test layout; then run a 12-point self-audit (coverage / correctness / test quality, including "could every listed test pass against an empty implementation?") and report the result. Compensates for dropping plan review from 3 rounds to 2.
|
|
16
|
+
- **Wave-boundary commit + context checkpoint.** Each wave ends with `git add -A && git commit -m "handoff: wave <N> — TASK-00x, ..."`, a cursor update, and a collapse of that wave to one line per task in working memory. Config: `handoff.autonomy.commitPerWave`.
|
|
17
|
+
|
|
18
|
+
### Changed
|
|
19
|
+
|
|
20
|
+
- **`git push` moved from the `ask` list to `allow`** in default settings, alongside `git add`, `git commit` and `git worktree`. Runs no longer stop at the final step for a confirmation. `git push --force` and `--force-with-lease` remain in `deny`, which still wins over `allow`.
|
|
21
|
+
- **Subagent return contracts are now capped.** Executors return ≤10 lines (`TASK`/`STATUS`/`EXECUTOR_MODEL`/`FILES`/`RED`/`VERIFY`/`NOTE`), reviewers ≤6. Verbose agent reports injected back into the orchestrator were the main cause of context exhaustion on long fullstack runs; this addresses it structurally rather than by raising caps.
|
|
22
|
+
- **`context-window-guard.sh` is now run-cursor aware.** With an unfinished run it emits shed-context-and-continue instructions instead of the stop-and-`/compact` directive. Asking the user to `/compact` is permitted only at a cycle boundary, in the Final Report — never mid-task. Config: `handoff.autonomy.compactAtCycleBoundaryOnly`.
|
|
23
|
+
- **Same-wave file conflicts auto-serialize** instead of marking both tasks `needs_breakdown`: the higher-numbered task gains a `Dependencies: TASK-<lower>` line and moves to the next wave. Config: `handoff.autonomy.autoSerializeFileConflicts`.
|
|
24
|
+
- **A dirty working tree no longer blocks Phase 3.** It is committed as `handoff: checkpoint before implement` and the run continues.
|
|
25
|
+
- **Plan review loop cap: 3 rounds → 2**, then findings are applied and the run proceeds (logged as `### Round <N> — findings applied without re-review`). Config: `handoff.autonomy.planReviewRounds`.
|
|
26
|
+
- **Review diff range is now anchored to the plan commit** (`git log --grep='^handoff: plan' -n 1`) rather than the working tree, since waves are committed as they finish.
|
|
27
|
+
- **`handoff-planner` maximizes wave width.** `Dependencies: none` unless task B genuinely cannot be written without A's output — ordering preference, same code area and review order are explicitly not dependencies.
|
|
28
|
+
- **`feature-implementer` and `code-reviewer` run unattended in handoff mode.** Ambiguity resolves through task file → PLAN.md → surrounding code → recommended choice, recorded in `## Discussion`; reviewer findings must be actionable defects (`file:line` + concrete failure + what correct looks like), never questions. Daily (non-handoff) behavior is unchanged.
|
|
29
|
+
- **`handoff-model-guard.sh` gained a bash fast path** that skips the node spawn when the payload contains neither `AI_HANDOFF` nor `git push`. Behavior-identical (verified against the previous version across a 6-payload sweep); measured 67ms → 38ms per ordinary Bash/Edit/Write call, paid on every call in every parallel subagent.
|
|
30
|
+
|
|
31
|
+
## 2.0.7 - 2026-08-15
|
|
32
|
+
|
|
33
|
+
### Added
|
|
34
|
+
|
|
35
|
+
- **`/ukit:handoff-create` and `/ukit:handoff-fullstack` gain a P2.5 independent plan review gate.** After the planner writes `PLAN.md` and before it's committed, a separate `code-reviewer` agent invocation (`REVIEW_TARGET_TYPE=plan`, fresh context — not the planner reviewing its own work) checks Completeness/Consistency/Clarity/Scope/YAGNI. `Issues Found` routes back to the planner to revise and resubmit; `Approved` unlocks the commit. Each round is logged to PLAN.md's new `## Plan Review Log` section, and the loop is capped at 3 `Issues Found` rounds — past that it escalates to the human instead of looping indefinitely.
|
|
36
|
+
- **Same-wave file-conflict precheck in Phase 3.** Before spawning any implementation wave, `handoff-implement.md`/`handoff-fullstack.md` now compare `Target Files` across every task in that wave; any overlap stops the wave and marks both tasks `needs_breakdown` instead of risking a mid-wave merge conflict. Code-level backstop for the planner's existing "no shared files in one wave" scope constraint.
|
|
37
|
+
- **TDD RED-state now requires real evidence.** The Executor Report template adds a mandatory `RED_OUTPUT` field — the actual failing-test output pasted before implementation, not a bare "confirmed" claim. `code-reviewer` checks this field during Test Plan adherence review and requests changes if it's missing or vague.
|
|
38
|
+
- **`handoff.plan.requireLintOrTypecheckInVerification`** (default `true`): when the project has a lint/typecheck script, task Verification Commands must run it, not just tests; the planner must state an explicit N/A if the project has none. Enforced by `code-reviewer` as review-order step 0.
|
|
39
|
+
|
|
40
|
+
### Changed
|
|
41
|
+
|
|
42
|
+
- **`handoff.maxParallelAgents`**: 3 → 10, raising the default cap on concurrent background agents per wave (config comments still warn against exceeding ~10-15 — each agent's report gets injected back into the orchestrator's own context on completion).
|
|
43
|
+
- **Phase 4 (Review) is now parallelized**, batched by `maxParallelAgents` the same way Phase 3 (Implement) already was — reviewer agents only read a diff and append a verdict to their own task file, so parallel review carries none of Phase 3's shared-worktree conflict risk.
|
|
44
|
+
- **`handoff.plan.minTestsEdgeCase`**: 1 → 2, and the two edge cases must now be of different kinds (e.g. null/empty **and** boundary/concurrent) — two near-duplicate cases no longer satisfy the requirement. Wired through `handoff-planner`, `code-reviewer`, `feature-implementer`'s inline-test-plan fallback, `RULES.md`, and both handoff pipeline commands.
|
|
45
|
+
|
|
5
46
|
## 2.0.6 - 2026-08-12
|
|
6
47
|
|
|
7
48
|
### Changed
|
|
@@ -1121,6 +1121,17 @@ items:
|
|
|
1121
1121
|
packs:
|
|
1122
1122
|
- core
|
|
1123
1123
|
|
|
1124
|
+
- id: hook-handoff-resume
|
|
1125
|
+
type: hook
|
|
1126
|
+
sourceTemplate: .claude/hooks/handoff-resume.sh
|
|
1127
|
+
targetPath: .claude/hooks/handoff-resume.sh
|
|
1128
|
+
requires: []
|
|
1129
|
+
mergeStrategy: overwrite_with_backup
|
|
1130
|
+
variables: []
|
|
1131
|
+
enabledByDefault: true
|
|
1132
|
+
packs:
|
|
1133
|
+
- core
|
|
1134
|
+
|
|
1124
1135
|
- id: hook-reset-compact-pressure
|
|
1125
1136
|
type: hook
|
|
1126
1137
|
sourceTemplate: .claude/hooks/reset-compact-pressure.sh
|
package/package.json
CHANGED
|
@@ -27,7 +27,8 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
27
27
|
|
|
28
28
|
### Review order
|
|
29
29
|
|
|
30
|
-
|
|
30
|
+
0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
|
|
31
|
+
1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
|
|
31
32
|
2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
|
|
32
33
|
3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider test suite if shared code was touched.
|
|
33
34
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
@@ -66,6 +67,27 @@ NOTES: [1-2 sentences for human reviewer if needed]
|
|
|
66
67
|
|
|
67
68
|
After writing the verdict, update `docs/AI_HANDOFF/INDEX.md` row for this task: set Status = NEXT_STATUS_FOR_INDEX, set Reviewer = your model name.
|
|
68
69
|
|
|
70
|
+
**Then return to the orchestrator at most 6 lines** — the full verdict is already on disk:
|
|
71
|
+
|
|
72
|
+
```
|
|
73
|
+
TASK: TASK-xxx
|
|
74
|
+
VERDICT: approved | approved_minor | changes_requested | critical_block
|
|
75
|
+
REVIEWER_MODEL: <exact model ID>
|
|
76
|
+
VERIFICATION_RERUN: PASS | FAIL
|
|
77
|
+
BLOCKING: <one line per critical/important finding, or "none">
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
Do not paste the diff, the findings prose, or verification output into the returned message.
|
|
81
|
+
The orchestrator is driving a whole pipeline in one context window and re-reads what it needs
|
|
82
|
+
from the task file; pasted reviewer logs are a common reason a run exhausts its context and
|
|
83
|
+
dies before the cycle finishes.
|
|
84
|
+
|
|
85
|
+
**You may be running unattended.** Findings do not end the run — the caller feeds them to an
|
|
86
|
+
auto-fix round and re-review. So write findings that a fresh executor can act on with no human
|
|
87
|
+
present: point at `file:line`, state the concrete failure, and say what correct looks like. A
|
|
88
|
+
finding phrased as a question ("should this handle null?") is not actionable; phrase it as the
|
|
89
|
+
defect ("`parse()` at foo.js:41 throws on null input; expected an empty result").
|
|
90
|
+
|
|
69
91
|
### Model isolation check (FIRST thing you do)
|
|
70
92
|
|
|
71
93
|
UKit cannot force any tool to use a specific model. The contract is enforced HERE, by you, via self-report comparison.
|
|
@@ -112,7 +134,10 @@ Only flag issues that would cause real problems during implementation planning.
|
|
|
112
134
|
|
|
113
135
|
### Output
|
|
114
136
|
|
|
137
|
+
Append (do NOT overwrite) this block to the end of the reviewed document under a `## Plan Review Log` section — create the section if it doesn't exist yet, keep all prior round entries:
|
|
138
|
+
|
|
115
139
|
```
|
|
140
|
+
### Round <N> — <YYYY-MM-DD> · <your model>
|
|
116
141
|
Status: Approved | Issues Found
|
|
117
142
|
|
|
118
143
|
COMPLETENESS:
|
|
@@ -128,3 +153,5 @@ YAGNI:
|
|
|
128
153
|
|
|
129
154
|
NOTES: [1-2 sentences if needed]
|
|
130
155
|
```
|
|
156
|
+
|
|
157
|
+
`<N>` = 1 + however many `### Round` entries already exist in the log (1 if this is the first review).
|
|
@@ -12,7 +12,15 @@ Implement requested behavior with minimal scope drift.
|
|
|
12
12
|
- **Daily/ad-hoc mode** (DEFAULT): task didn't come from `docs/AI_HANDOFF/` → use the original lightweight workflow. Tests only when touched code already has coverage. No reviewer trigger.
|
|
13
13
|
- **Handoff mode**: task file is `docs/AI_HANDOFF/tasks/TASK-xxx.md` OR user explicitly invokes handoff (e.g. "execute task TASK-001") → activate full Quality Gate: test-first → green → reviewer.
|
|
14
14
|
|
|
15
|
-
If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
|
|
15
|
+
If unsure which mode applies, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
|
|
16
|
+
|
|
17
|
+
**In Handoff mode you are running unattended — ask nothing.** You were spawned by an
|
|
18
|
+
orchestrator driving a pipeline; there is no human in your conversation to answer, and a
|
|
19
|
+
question there is silently dropped while the run stalls. Resolve ambiguity in this order:
|
|
20
|
+
the task file → `PLAN.md` → the surrounding code's existing patterns → the choice you would
|
|
21
|
+
recommend. Record what you chose and why in the task's `## Discussion` thread. Only a blocker
|
|
22
|
+
outside the repo (missing credential, unreachable service) justifies reporting `FAIL` early —
|
|
23
|
+
and even then, report it, don't ask about it.
|
|
16
24
|
|
|
17
25
|
## Workflow
|
|
18
26
|
|
|
@@ -25,12 +33,12 @@ If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
|
|
|
25
33
|
- non-trivial: `docs/MEMORY.md` + `docs/PROJECT.md` + `docs/CODE_MAP.md`
|
|
26
34
|
- Identify target files and existing patterns.
|
|
27
35
|
- If task came from handoff, read `tasks/TASK-xxx.md` and locate its **Test Plan** + **Verification Commands**.
|
|
28
|
-
-
|
|
36
|
+
- Daily mode: if confidence is low or risk is high, ask one short clarifying question before deeper analysis. Handoff mode: do not ask — decide and record the decision (see above).
|
|
29
37
|
|
|
30
38
|
### 2. Plan Approach (< 1 minute)
|
|
31
39
|
|
|
32
40
|
- List files to create/modify (max diff).
|
|
33
|
-
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥
|
|
41
|
+
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
34
42
|
|
|
35
43
|
### 3. Test First (RED) — Handoff mode
|
|
36
44
|
|
|
@@ -79,6 +87,26 @@ NEXT: [follow-up needed, or "ready for review"]
|
|
|
79
87
|
|
|
80
88
|
> **Self-report rule:** UKit cannot force any tool/host to use a specific model. Your self-reported `EXECUTOR_MODEL` is how the reviewer (in another tool or subagent) knows what to compare against its own model. Misreporting → reviewer refuses and asks the human to confirm.
|
|
81
89
|
|
|
90
|
+
**Handoff mode — where that report goes.** Append the block above, in full, to the task file
|
|
91
|
+
as `## Executor Report` (or `## Executor Report (fix round <N>)` on a re-run). Then return to
|
|
92
|
+
the orchestrator **at most 10 lines**:
|
|
93
|
+
|
|
94
|
+
```
|
|
95
|
+
TASK: TASK-xxx
|
|
96
|
+
STATUS: PASS | FAIL
|
|
97
|
+
EXECUTOR_MODEL: <exact model ID>
|
|
98
|
+
FILES: <comma-separated changed paths>
|
|
99
|
+
RED: confirmed | not-confirmed
|
|
100
|
+
VERIFY: <n> commands, all pass | <first failing command + one-line reason>
|
|
101
|
+
NOTE: <one line, or "none">
|
|
102
|
+
```
|
|
103
|
+
|
|
104
|
+
Do not repeat RED output, verification logs, diffs, or file contents in the returned message.
|
|
105
|
+
The reviewer reads all of that from the task file on disk. The orchestrator is driving a whole
|
|
106
|
+
pipeline in one context window, and pasted subagent logs are the single biggest reason a run
|
|
107
|
+
runs out of context and dies half-finished. Being terse here is not a style preference — it is
|
|
108
|
+
what lets the cycle reach the end.
|
|
109
|
+
|
|
82
110
|
### 7. Trigger Reviewer — Handoff mode ONLY
|
|
83
111
|
|
|
84
112
|
- Daily mode: skip this step entirely. Just report and stop.
|
|
@@ -17,15 +17,43 @@ Use the strongest model available — planning with a weak model produces weak t
|
|
|
17
17
|
|
|
18
18
|
## Scope Check (before Phase 1)
|
|
19
19
|
|
|
20
|
-
Before refining the request into `PLAN.md`, check whether it describes multiple independent subsystems (e.g. "CRM + AI + billing + analytics + mobile app"). If yes,
|
|
20
|
+
Before refining the request into `PLAN.md`, check whether it describes multiple independent subsystems (e.g. "CRM + AI + billing + analytics + mobile app"). If yes, record the decomposition:
|
|
21
21
|
|
|
22
22
|
```
|
|
23
23
|
Scope complexity: HIGH
|
|
24
24
|
Detected systems: [...]
|
|
25
|
-
|
|
25
|
+
Decomposition: N modules — module 1 planned now, modules 2..N queued
|
|
26
26
|
```
|
|
27
27
|
|
|
28
|
-
|
|
28
|
+
Then **plan module 1 in full and keep going** — do not wait for confirmation. A mega-spec is
|
|
29
|
+
what you must avoid, not the work itself. Write the remaining modules into `PLAN.md` §2 as
|
|
30
|
+
explicitly out-of-scope-for-this-cycle, and add one `queued` row per module to `INDEX.md` so
|
|
31
|
+
the next cycle picks them up. Order modules so the one others depend on is planned first.
|
|
32
|
+
|
|
33
|
+
The caller may be running unattended (`/ukit:handoff-fullstack`). Stopping to ask means the
|
|
34
|
+
run dies and the human returns hours later to a question, which is the failure mode this
|
|
35
|
+
whole pipeline exists to prevent.
|
|
36
|
+
|
|
37
|
+
## Grounding — every path and command must be real
|
|
38
|
+
|
|
39
|
+
The most common way a plan fails is not bad reasoning, it is **confident invention**: a target
|
|
40
|
+
file that doesn't exist, a test command the project doesn't have, an import path that never
|
|
41
|
+
resolved. Executors then burn a whole round discovering it.
|
|
42
|
+
|
|
43
|
+
Before writing any path or command into `PLAN.md` or a task file, verify it:
|
|
44
|
+
|
|
45
|
+
- **Target Files** — for each path, confirm it exists (`ls`/Glob), or that its parent
|
|
46
|
+
directory exists and the file is genuinely new. Mark new files `(new)` explicitly.
|
|
47
|
+
- **Verification Commands** — read `package.json` `scripts` (or the stack's equivalent) and
|
|
48
|
+
use the script names that are actually defined. Do not write `npm test` when the repo uses
|
|
49
|
+
`yarn test`, and do not invent a `lint` script that isn't there. Run `--help` or a dry check
|
|
50
|
+
if unsure.
|
|
51
|
+
- **Test Files** — follow the existing test layout and naming; open one neighbouring test file
|
|
52
|
+
and match its structure, framework and import style.
|
|
53
|
+
- **Interfaces** — quote real signatures from the source, not plausible-looking ones.
|
|
54
|
+
|
|
55
|
+
If something cannot be verified, say so in the task's `## Discussion` rather than guessing.
|
|
56
|
+
A stated unknown costs the executor one read; a wrong path costs it a round.
|
|
29
57
|
|
|
30
58
|
## Phase 1 — Write PLAN.md
|
|
31
59
|
|
|
@@ -37,9 +65,15 @@ Write all 7 sections to `docs/AI_HANDOFF/PLAN.md`:
|
|
|
37
65
|
CONSTRAINT: tasks in the same wave must not modify the same file.
|
|
38
66
|
If two tasks need the same file → make one depend on the other.
|
|
39
67
|
§3 Approach — technical solution, trade-offs, alternatives rejected
|
|
40
|
-
§4 Test Plan — happy path × N + ≥
|
|
68
|
+
§4 Test Plan — happy path × N + ≥2 edge cases of DIFFERENT kinds (e.g. null/empty AND
|
|
69
|
+
boundary/concurrent — two near-duplicate cases do not satisfy this) +
|
|
70
|
+
regression (if bugfix)
|
|
41
71
|
table: | Type | Test Name | Expected |
|
|
42
|
-
§5 Verification — exact shell commands executor will run
|
|
72
|
+
§5 Verification — exact shell commands executor will run. If the project has a lint or
|
|
73
|
+
typecheck script (check package.json `scripts`, or the equivalent for
|
|
74
|
+
the project's stack), it MUST be included here, not just the test
|
|
75
|
+
command. If the project genuinely has none, state that explicitly —
|
|
76
|
+
do not omit silently.
|
|
43
77
|
§6 Acceptance — checklist of done criteria (prefer verifiable/command-based criteria)
|
|
44
78
|
§7 Global Constraints — one line each: version floors, dependency limits, naming/copy
|
|
45
79
|
rules, platform requirements. Every TASK-xxx.md inherits this section
|
|
@@ -55,6 +89,8 @@ Append this footer to `PLAN.md` — mandatory, checked by a hook before the writ
|
|
|
55
89
|
PLANNER_MODEL: <your exact model ID — e.g. claude-opus-5>
|
|
56
90
|
```
|
|
57
91
|
|
|
92
|
+
Your output does not go straight to implementation: an independent `code-reviewer` pass (`REVIEW_TARGET_TYPE=plan`) reviews `PLAN.md` next. If it returns `Issues Found`, you'll be re-invoked to revise and resubmit — write §1-§7 tight enough to pass on the first pass.
|
|
93
|
+
|
|
58
94
|
## Phase 2 — Split into TASK-xxx.md
|
|
59
95
|
|
|
60
96
|
**Right-sizing rule:** A task is the smallest unit that carries its own test cycle and is worth a fresh reviewer's gate. Split only where a reviewer could meaningfully approve one task while rejecting its neighbor. Each task ends with an independently testable deliverable.
|
|
@@ -67,9 +103,9 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
|
|
|
67
103
|
|-------|------|
|
|
68
104
|
| Target Files | Exact paths — no two tasks in same wave share a file |
|
|
69
105
|
| Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
|
|
70
|
-
| Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥
|
|
106
|
+
| Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
|
|
71
107
|
| Test Files | Exact test file paths to create/modify |
|
|
72
|
-
| Verification Commands | Runnable shell commands |
|
|
108
|
+
| Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
|
|
73
109
|
| Acceptance Criteria | Verifiable checklist |
|
|
74
110
|
|
|
75
111
|
Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
|
|
@@ -80,6 +116,30 @@ Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
|
|
|
80
116
|
- Chain: A → B → C runs as 3 sequential waves (1 task each, no parallelism)
|
|
81
117
|
- Independent: A, B, C (all `none`) runs as 1 wave, all parallel
|
|
82
118
|
|
|
119
|
+
### Maximize wave width — dependencies are expensive
|
|
120
|
+
|
|
121
|
+
Wave width is the single biggest lever on how long a cycle takes: a wave of 6 finishes in
|
|
122
|
+
roughly the time of its slowest task, while a chain of 6 takes six times that. Executors run
|
|
123
|
+
up to `handoff.maxParallelAgents` (default 10) at once, so a plan that produces `none`
|
|
124
|
+
dependencies for most tasks is dramatically faster than one that produces a chain.
|
|
125
|
+
|
|
126
|
+
**Write `Dependencies: none` unless B genuinely cannot be written without A's output.** A real
|
|
127
|
+
dependency means B imports a symbol A creates, or B tests behavior A implements. These are
|
|
128
|
+
*not* dependencies:
|
|
129
|
+
|
|
130
|
+
- "B is logically later" or "B builds on the same feature" — ordering preference, not a
|
|
131
|
+
dependency.
|
|
132
|
+
- "Both touch the same area of the codebase" — irrelevant unless they touch the same *file*.
|
|
133
|
+
- "A should be reviewed before B starts" — that is what the review phase is for.
|
|
134
|
+
|
|
135
|
+
Same-file collisions are the one real constraint, and they are cheap to design around: split
|
|
136
|
+
along file boundaries so each task owns its files outright. If two pieces of work truly must
|
|
137
|
+
edit one file, prefer merging them into a single task over chaining two — one task with two
|
|
138
|
+
test groups beats two waves.
|
|
139
|
+
|
|
140
|
+
Before finalizing, count your waves. If the dependency graph is mostly a chain, re-examine it:
|
|
141
|
+
most chains are ordering preferences that a wide wave 1 would satisfy just as well.
|
|
142
|
+
|
|
83
143
|
## Phase 3 — Update state files
|
|
84
144
|
|
|
85
145
|
**INDEX.md** row per task: `| TASK-001 | <name> | ready | none | - |`
|
|
@@ -93,15 +153,59 @@ Status: planning_done — ready for executor
|
|
|
93
153
|
```
|
|
94
154
|
Wave structure is NOT stored here — inferred from task `Dependencies` fields at runtime.
|
|
95
155
|
|
|
156
|
+
## Self-Audit — run before reporting, every time
|
|
157
|
+
|
|
158
|
+
The downstream pipeline is unattended: an incomplete or wrong plan is not caught by a human
|
|
159
|
+
skimming it, it is caught by an executor failing three hours later. Independent plan review
|
|
160
|
+
runs at most twice, so this audit is your own last line of defense.
|
|
161
|
+
|
|
162
|
+
Walk the checklist and **fix what fails before you report** — do not report and hope review
|
|
163
|
+
catches it:
|
|
164
|
+
|
|
165
|
+
**Coverage — is anything missing?**
|
|
166
|
+
1. Does every acceptance criterion in §6 trace to at least one task? Name the task per criterion.
|
|
167
|
+
2. Does every task trace back to something in §1/§6? A task nothing asks for is scope creep — cut it.
|
|
168
|
+
3. Do the tasks together actually deliver §1's success definition, or only the easy part of it? State the gap if there is one.
|
|
169
|
+
4. Is the *unhappy* path planned — errors, empty input, permissions, migration of existing data — or only the feature?
|
|
170
|
+
|
|
171
|
+
**Correctness — is anything wrong?**
|
|
172
|
+
5. Every `Target Files` path verified per Grounding? Any `(new)` file marked as such?
|
|
173
|
+
6. Every `Verification Command` a real, defined script in this repo?
|
|
174
|
+
7. Any two same-wave tasks sharing a file? If yes, merge them or add the dependency now — the executor will otherwise serialize them for you and your wave plan was a lie.
|
|
175
|
+
8. Does any task depend on a symbol or file that no earlier task creates? That's a missing task, not a dependency.
|
|
176
|
+
|
|
177
|
+
**Test quality — the part most often faked:**
|
|
178
|
+
9. Each task: ≥1 happy path + ≥2 edge cases of genuinely *different kinds*. "empty string" and "null" are the same kind — boundary, concurrency, malformed input, permission denied, duplicate, ordering are different kinds.
|
|
179
|
+
10. Does each test case state a concrete `Expected` value, not "works correctly" or "returns successfully"? A test whose expectation cannot fail is not a test.
|
|
180
|
+
11. Bugfix task → is there a regression test that fails against today's code?
|
|
181
|
+
12. Could every listed test pass against an empty implementation? If so, the test cases describe nothing and must be rewritten.
|
|
182
|
+
|
|
183
|
+
Append the result to `PLAN.md`:
|
|
184
|
+
```
|
|
185
|
+
## Planner Self-Audit
|
|
186
|
+
Checklist: 12/12 pass
|
|
187
|
+
Fixed during audit: <what you changed, or "nothing">
|
|
188
|
+
Known gaps: <what you deliberately left out and why, or "none">
|
|
189
|
+
```
|
|
190
|
+
|
|
191
|
+
An honest `Known gaps` line is worth more than a clean sheet — it is the one thing the reviewer
|
|
192
|
+
and executor cannot recover on their own.
|
|
193
|
+
|
|
96
194
|
## Output
|
|
97
195
|
|
|
196
|
+
Keep the returned message under 25 lines — the caller may be an orchestrator whose context
|
|
197
|
+
budget is the constraint on the whole run. Detail belongs in `PLAN.md`, not in the reply.
|
|
198
|
+
|
|
98
199
|
- Task count + IDs
|
|
99
200
|
- Dependency graph (text form: TASK-001 → TASK-003, TASK-002 independent)
|
|
201
|
+
- Wave plan: `wave 1: N tasks | wave 2: M tasks` — flag it if the graph is mostly a chain
|
|
202
|
+
- Self-audit result + any `Known gaps`
|
|
100
203
|
- Any `needs_breakdown` tasks + reason
|
|
101
204
|
- Next step: "Switch to Sonnet/unic-code → run `/ukit:handoff-implement`"
|
|
102
205
|
|
|
103
206
|
## Rules
|
|
104
207
|
|
|
105
208
|
- Do NOT implement. Job ends when all tasks are `ready` or `needs_breakdown`.
|
|
106
|
-
- Undone tasks in INDEX.md →
|
|
107
|
-
- Cannot determine test cases → `needs_breakdown` + Discussion thread note.
|
|
209
|
+
- Undone tasks in INDEX.md → continue that cycle; plan only what is still missing. Do not ask whether to start fresh, and do not overwrite task files that already carry an executor report or reviewer verdict.
|
|
210
|
+
- Cannot determine test cases → `needs_breakdown` + Discussion thread note. Never mark a task `ready` with vague tests just to keep the pipeline moving — a task with fake tests passes review and ships a bug, which costs far more than one blocked task.
|
|
211
|
+
- You may be invoked unattended. Ask nothing; where the caller allowed questions, that window closed before you were spawned. Resolve ambiguity by reading the codebase, choose the option you would recommend, and record the choice and its rationale in `PLAN.md §3`.
|
|
@@ -14,6 +14,29 @@ $ARGUMENTS
|
|
|
14
14
|
|
|
15
15
|
---
|
|
16
16
|
|
|
17
|
+
## Question Policy — this is the one phase that may ask
|
|
18
|
+
|
|
19
|
+
Planning is the **only** place in the handoff pipeline where asking the human is allowed, and
|
|
20
|
+
it is where every ambiguity must be burned off. Whatever is left unresolved here becomes a
|
|
21
|
+
guess during implementation, where nobody is available to correct it.
|
|
22
|
+
|
|
23
|
+
**Ask once, up front, batched.** Before Step 1, if `$ARGUMENTS` leaves a choice that changes
|
|
24
|
+
*what gets built* — not how — collect every such question into a **single** `AskUserQuestion`
|
|
25
|
+
call (max 4 questions, each with a `(Recommended)` first option). Then close the window.
|
|
26
|
+
|
|
27
|
+
Ask about: ambiguous scope boundaries, conflicting interpretations of the request, which
|
|
28
|
+
existing subsystem to extend versus replace, and any acceptance criterion you would otherwise
|
|
29
|
+
invent. Do **not** ask about: file layout, naming, test framework choice, library versions
|
|
30
|
+
already used in the repo, or anything the codebase or `docs/` already answers — resolve those
|
|
31
|
+
by reading, and record the resolution in `PLAN.md §3`.
|
|
32
|
+
|
|
33
|
+
Record every answer verbatim in `PLAN.md §1 Intent`. Downstream phases treat that section as
|
|
34
|
+
the human's own words and never re-ask.
|
|
35
|
+
|
|
36
|
+
After the batched call, run Steps 1–2.5 to completion without further questions.
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
17
40
|
## Step 1 — Read context (lite model)
|
|
18
41
|
|
|
19
42
|
**Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"`. Do NOT read these files yourself in the current session — this step is contracted to the lite tier (haiku/unic-lite), which only the spawned agent's frontmatter model guarantees. Ask the agent to:
|
|
@@ -36,7 +59,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
36
59
|
|
|
37
60
|
1. Check INDEX.md task statuses:
|
|
38
61
|
- All tasks are `ready` (planning only) → **re-run allowed**: overwrite PLAN.md and TASK-xxx.md freely — this is iterative refinement.
|
|
39
|
-
- Any task is `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **
|
|
62
|
+
- Any task is `pending_review`, `changes_requested`, `merge_conflict`, or `blocked` → **do not overwrite**: that cycle is mid-flight and its task files carry executor reports and verdicts. Report which tasks are in flight and point the human at `/ukit:handoff-review` (to finish the cycle) or `/ukit:handoff-clear` (to abandon it). Planning is the one phase where stopping is correct — there is nothing safe to do automatically here.
|
|
40
63
|
- No tasks → fresh cycle, proceed normally.
|
|
41
64
|
|
|
42
65
|
2. Resolve base branch:
|
|
@@ -48,7 +71,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
48
71
|
- §1 Intent — problem + success definition
|
|
49
72
|
- §2 Scope — in / out of scope. **Add a constraint**: same-wave tasks must not modify the same file (prevents merge conflicts). If two tasks need the same file, make one depend on the other.
|
|
50
73
|
- §3 Approach — solution, trade-offs, alternatives rejected
|
|
51
|
-
- §4 Test Plan — happy path + ≥
|
|
74
|
+
- §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
|
|
52
75
|
- §5 Verification — exact shell commands executor will run
|
|
53
76
|
- §6 Acceptance — done checklist (prefer verifiable criteria with commands)
|
|
54
77
|
|
|
@@ -62,7 +85,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
62
85
|
Every task MUST have:
|
|
63
86
|
- Target Files (exact paths — no two tasks in same wave share a file)
|
|
64
87
|
- Dependencies (`TASK-xxx` or `none` — wave structure inferred from this, not stored separately)
|
|
65
|
-
- Test Cases (Type | Name | Expected — ≥1 happy + ≥
|
|
88
|
+
- Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
|
|
66
89
|
- Test Files (exact paths)
|
|
67
90
|
- Verification Commands (runnable shell commands)
|
|
68
91
|
- Acceptance Criteria (verifiable checklist)
|
|
@@ -85,4 +108,25 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
85
108
|
|
|
86
109
|
---
|
|
87
110
|
|
|
111
|
+
## Step 2.5 — Independent plan review (strong model, separate agent)
|
|
112
|
+
|
|
113
|
+
**Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet).
|
|
114
|
+
|
|
115
|
+
- count = 0 → run the review below.
|
|
116
|
+
- count = 1 and the round returned `Issues Found` → planner revises once, then run **one** more review round.
|
|
117
|
+
- count ≥ 2 → do NOT invoke the reviewer again, and do not stop. Have the planner apply every outstanding finding directly to `PLAN.md` and the affected `TASK-xxx.md` files, append `### Round <N> — findings applied without re-review` listing what changed, then finish.
|
|
118
|
+
|
|
119
|
+
Two independent strong-model passes shape the plan before any code is written — that gate is
|
|
120
|
+
intact. What it no longer does is hand a stalled plan back and wait.
|
|
121
|
+
|
|
122
|
+
**Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
|
|
123
|
+
|
|
124
|
+
1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
|
|
125
|
+
2. `Issues Found` → route back to Step 2: planner revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another Step 2.5 review (this becomes the next round). Do NOT commit or hand off to executor on `Issues Found`.
|
|
126
|
+
3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed.
|
|
127
|
+
|
|
128
|
+
> Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from Step 2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
|
|
129
|
+
|
|
130
|
+
---
|
|
131
|
+
|
|
88
132
|
**Next:** switch to code model → `/ukit:handoff-implement`
|