@ngockhoale/ukit 2.0.6 → 2.0.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +15 -0
- package/package.json +1 -1
- package/templates/.claude/agents/code-reviewer.md +7 -1
- package/templates/.claude/agents/feature-implementer.md +1 -1
- package/templates/.claude/agents/handoff-planner.md +12 -4
- package/templates/.claude/commands/ukit/handoff-create.md +16 -2
- package/templates/.claude/commands/ukit/handoff-fullstack.md +31 -6
- package/templates/.claude/commands/ukit/handoff-implement.md +14 -2
- package/templates/.claude/commands/ukit/handoff-review.md +4 -2
- package/templates/docs/AI_HANDOFF/RULES.md +5 -5
- package/templates/ukit/storage/config.json +6 -4
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,21 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.0.7 - 2026-08-15
|
|
6
|
+
|
|
7
|
+
### Added
|
|
8
|
+
|
|
9
|
+
- **`/ukit:handoff-create` and `/ukit:handoff-fullstack` gain a P2.5 independent plan review gate.** After the planner writes `PLAN.md` and before it's committed, a separate `code-reviewer` agent invocation (`REVIEW_TARGET_TYPE=plan`, fresh context — not the planner reviewing its own work) checks Completeness/Consistency/Clarity/Scope/YAGNI. `Issues Found` routes back to the planner to revise and resubmit; `Approved` unlocks the commit. Each round is logged to PLAN.md's new `## Plan Review Log` section, and the loop is capped at 3 `Issues Found` rounds — past that it escalates to the human instead of looping indefinitely.
|
|
10
|
+
- **Same-wave file-conflict precheck in Phase 3.** Before spawning any implementation wave, `handoff-implement.md`/`handoff-fullstack.md` now compare `Target Files` across every task in that wave; any overlap stops the wave and marks both tasks `needs_breakdown` instead of risking a mid-wave merge conflict. Code-level backstop for the planner's existing "no shared files in one wave" scope constraint.
|
|
11
|
+
- **TDD RED-state now requires real evidence.** The Executor Report template adds a mandatory `RED_OUTPUT` field — the actual failing-test output pasted before implementation, not a bare "confirmed" claim. `code-reviewer` checks this field during Test Plan adherence review and requests changes if it's missing or vague.
|
|
12
|
+
- **`handoff.plan.requireLintOrTypecheckInVerification`** (default `true`): when the project has a lint/typecheck script, task Verification Commands must run it, not just tests; the planner must state an explicit N/A if the project has none. Enforced by `code-reviewer` as review-order step 0.
|
|
13
|
+
|
|
14
|
+
### Changed
|
|
15
|
+
|
|
16
|
+
- **`handoff.maxParallelAgents`**: 3 → 10, raising the default cap on concurrent background agents per wave (config comments still warn against exceeding ~10-15 — each agent's report gets injected back into the orchestrator's own context on completion).
|
|
17
|
+
- **Phase 4 (Review) is now parallelized**, batched by `maxParallelAgents` the same way Phase 3 (Implement) already was — reviewer agents only read a diff and append a verdict to their own task file, so parallel review carries none of Phase 3's shared-worktree conflict risk.
|
|
18
|
+
- **`handoff.plan.minTestsEdgeCase`**: 1 → 2, and the two edge cases must now be of different kinds (e.g. null/empty **and** boundary/concurrent) — two near-duplicate cases no longer satisfy the requirement. Wired through `handoff-planner`, `code-reviewer`, `feature-implementer`'s inline-test-plan fallback, `RULES.md`, and both handoff pipeline commands.
|
|
19
|
+
|
|
5
20
|
## 2.0.6 - 2026-08-12
|
|
6
21
|
|
|
7
22
|
### Changed
|
package/package.json
CHANGED
|
@@ -27,7 +27,8 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
27
27
|
|
|
28
28
|
### Review order
|
|
29
29
|
|
|
30
|
-
|
|
30
|
+
0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
|
|
31
|
+
1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
|
|
31
32
|
2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
|
|
32
33
|
3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider test suite if shared code was touched.
|
|
33
34
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
@@ -112,7 +113,10 @@ Only flag issues that would cause real problems during implementation planning.
|
|
|
112
113
|
|
|
113
114
|
### Output
|
|
114
115
|
|
|
116
|
+
Append (do NOT overwrite) this block to the end of the reviewed document under a `## Plan Review Log` section — create the section if it doesn't exist yet, keep all prior round entries:
|
|
117
|
+
|
|
115
118
|
```
|
|
119
|
+
### Round <N> — <YYYY-MM-DD> · <your model>
|
|
116
120
|
Status: Approved | Issues Found
|
|
117
121
|
|
|
118
122
|
COMPLETENESS:
|
|
@@ -128,3 +132,5 @@ YAGNI:
|
|
|
128
132
|
|
|
129
133
|
NOTES: [1-2 sentences if needed]
|
|
130
134
|
```
|
|
135
|
+
|
|
136
|
+
`<N>` = 1 + however many `### Round` entries already exist in the log (1 if this is the first review).
|
|
@@ -30,7 +30,7 @@ If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
|
|
|
30
30
|
### 2. Plan Approach (< 1 minute)
|
|
31
31
|
|
|
32
32
|
- List files to create/modify (max diff).
|
|
33
|
-
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥
|
|
33
|
+
- **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
|
|
34
34
|
|
|
35
35
|
### 3. Test First (RED) — Handoff mode
|
|
36
36
|
|
|
@@ -37,9 +37,15 @@ Write all 7 sections to `docs/AI_HANDOFF/PLAN.md`:
|
|
|
37
37
|
CONSTRAINT: tasks in the same wave must not modify the same file.
|
|
38
38
|
If two tasks need the same file → make one depend on the other.
|
|
39
39
|
§3 Approach — technical solution, trade-offs, alternatives rejected
|
|
40
|
-
§4 Test Plan — happy path × N + ≥
|
|
40
|
+
§4 Test Plan — happy path × N + ≥2 edge cases of DIFFERENT kinds (e.g. null/empty AND
|
|
41
|
+
boundary/concurrent — two near-duplicate cases do not satisfy this) +
|
|
42
|
+
regression (if bugfix)
|
|
41
43
|
table: | Type | Test Name | Expected |
|
|
42
|
-
§5 Verification — exact shell commands executor will run
|
|
44
|
+
§5 Verification — exact shell commands executor will run. If the project has a lint or
|
|
45
|
+
typecheck script (check package.json `scripts`, or the equivalent for
|
|
46
|
+
the project's stack), it MUST be included here, not just the test
|
|
47
|
+
command. If the project genuinely has none, state that explicitly —
|
|
48
|
+
do not omit silently.
|
|
43
49
|
§6 Acceptance — checklist of done criteria (prefer verifiable/command-based criteria)
|
|
44
50
|
§7 Global Constraints — one line each: version floors, dependency limits, naming/copy
|
|
45
51
|
rules, platform requirements. Every TASK-xxx.md inherits this section
|
|
@@ -55,6 +61,8 @@ Append this footer to `PLAN.md` — mandatory, checked by a hook before the writ
|
|
|
55
61
|
PLANNER_MODEL: <your exact model ID — e.g. claude-opus-5>
|
|
56
62
|
```
|
|
57
63
|
|
|
64
|
+
Your output does not go straight to implementation: an independent `code-reviewer` pass (`REVIEW_TARGET_TYPE=plan`) reviews `PLAN.md` next. If it returns `Issues Found`, you'll be re-invoked to revise and resubmit — write §1-§7 tight enough to pass on the first pass.
|
|
65
|
+
|
|
58
66
|
## Phase 2 — Split into TASK-xxx.md
|
|
59
67
|
|
|
60
68
|
**Right-sizing rule:** A task is the smallest unit that carries its own test cycle and is worth a fresh reviewer's gate. Split only where a reviewer could meaningfully approve one task while rejecting its neighbor. Each task ends with an independently testable deliverable.
|
|
@@ -67,9 +75,9 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
|
|
|
67
75
|
|-------|------|
|
|
68
76
|
| Target Files | Exact paths — no two tasks in same wave share a file |
|
|
69
77
|
| Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
|
|
70
|
-
| Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥
|
|
78
|
+
| Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
|
|
71
79
|
| Test Files | Exact test file paths to create/modify |
|
|
72
|
-
| Verification Commands | Runnable shell commands |
|
|
80
|
+
| Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
|
|
73
81
|
| Acceptance Criteria | Verifiable checklist |
|
|
74
82
|
|
|
75
83
|
Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
|
|
@@ -48,7 +48,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
48
48
|
- §1 Intent — problem + success definition
|
|
49
49
|
- §2 Scope — in / out of scope. **Add a constraint**: same-wave tasks must not modify the same file (prevents merge conflicts). If two tasks need the same file, make one depend on the other.
|
|
50
50
|
- §3 Approach — solution, trade-offs, alternatives rejected
|
|
51
|
-
- §4 Test Plan — happy path + ≥
|
|
51
|
+
- §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
|
|
52
52
|
- §5 Verification — exact shell commands executor will run
|
|
53
53
|
- §6 Acceptance — done checklist (prefer verifiable criteria with commands)
|
|
54
54
|
|
|
@@ -62,7 +62,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
62
62
|
Every task MUST have:
|
|
63
63
|
- Target Files (exact paths — no two tasks in same wave share a file)
|
|
64
64
|
- Dependencies (`TASK-xxx` or `none` — wave structure inferred from this, not stored separately)
|
|
65
|
-
- Test Cases (Type | Name | Expected — ≥1 happy + ≥
|
|
65
|
+
- Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
|
|
66
66
|
- Test Files (exact paths)
|
|
67
67
|
- Verification Commands (runnable shell commands)
|
|
68
68
|
- Acceptance Criteria (verifiable checklist)
|
|
@@ -85,4 +85,18 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
|
|
|
85
85
|
|
|
86
86
|
---
|
|
87
87
|
|
|
88
|
+
## Step 2.5 — Independent plan review (strong model, separate agent)
|
|
89
|
+
|
|
90
|
+
**Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet). If count ≥ 3 (3 prior rounds already returned `Issues Found`), do NOT invoke the reviewer again — STOP and escalate to the human: show the accumulated findings from all 3 rounds, ask them to manually revise PLAN.md, narrow scope, or explicitly approve an override. Otherwise, proceed below.
|
|
91
|
+
|
|
92
|
+
**Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
|
|
93
|
+
|
|
94
|
+
1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
|
|
95
|
+
2. `Issues Found` → route back to Step 2: planner revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another Step 2.5 review (this becomes the next round). Do NOT commit or hand off to executor on `Issues Found`.
|
|
96
|
+
3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed.
|
|
97
|
+
|
|
98
|
+
> Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from Step 2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
|
|
99
|
+
|
|
100
|
+
---
|
|
101
|
+
|
|
88
102
|
**Next:** switch to code model → `/ukit:handoff-implement`
|
|
@@ -57,7 +57,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
57
57
|
- §1 Intent — problem + success definition
|
|
58
58
|
- §2 Scope — in / out of scope; same-wave tasks must not modify the same file (prevents merge conflicts)
|
|
59
59
|
- §3 Approach — solution, trade-offs, alternatives rejected
|
|
60
|
-
- §4 Test Plan — happy path + ≥
|
|
60
|
+
- §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
|
|
61
61
|
- §5 Verification — exact shell commands executor will run
|
|
62
62
|
- §6 Acceptance — done checklist (prefer verifiable criteria with commands)
|
|
63
63
|
|
|
@@ -71,7 +71,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
71
71
|
Every task MUST have:
|
|
72
72
|
- Target Files (exact paths — no two tasks in same wave share a file)
|
|
73
73
|
- Dependencies (`TASK-xxx` or `none` — wave structure inferred from this)
|
|
74
|
-
- Test Cases (Type | Name | Expected — ≥1 happy + ≥
|
|
74
|
+
- Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
|
|
75
75
|
- Test Files (exact paths)
|
|
76
76
|
- Verification Commands (runnable shell commands)
|
|
77
77
|
- Acceptance Criteria (verifiable checklist)
|
|
@@ -89,6 +89,18 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
|
|
|
89
89
|
|
|
90
90
|
7. **Report:** task IDs, dependency graph, any `needs_breakdown` tasks + reason.
|
|
91
91
|
|
|
92
|
+
### P2.5 — Independent plan review (strong model, separate agent)
|
|
93
|
+
|
|
94
|
+
**Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet). If count ≥ 3 (3 prior rounds already returned `Issues Found`), do NOT invoke the reviewer again — STOP and escalate to the human: show the accumulated findings from all 3 rounds, ask them to manually revise PLAN.md, narrow scope, or explicitly approve an override. Otherwise, proceed below.
|
|
95
|
+
|
|
96
|
+
**Claude Code — MANDATORY, do this before anything else in P2.5:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from P2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
|
|
97
|
+
|
|
98
|
+
1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
|
|
99
|
+
2. `Issues Found` → route back to P2: `handoff-planner` revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another P2.5 review (this becomes the next round). Do NOT proceed to P3 on `Issues Found`.
|
|
100
|
+
3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed to P3.
|
|
101
|
+
|
|
102
|
+
> Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from P2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
|
|
103
|
+
|
|
92
104
|
### P3 — Commit the plan (lite model)
|
|
93
105
|
|
|
94
106
|
**Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for this commit step (lite tier — haiku/unic-lite). Run:
|
|
@@ -127,8 +139,16 @@ Read each `docs/AI_HANDOFF/tasks/TASK-xxx.md` for `Dependencies` field:
|
|
|
127
139
|
- Chain A→B→C = 3 waves of 1 task each (sequential)
|
|
128
140
|
- Independent A, B, C = 1 wave of 3 tasks (parallel)
|
|
129
141
|
|
|
142
|
+
**Conflict check — mandatory, before spawning any wave.** For every pair of tasks landing
|
|
143
|
+
in the same wave, compare their `Target Files` lists. If any file path appears in both,
|
|
144
|
+
STOP — do not spawn that wave. This is the code-level safety net for the planner's own §2
|
|
145
|
+
Scope constraint (same-wave tasks must not share a file) in case it slipped through
|
|
146
|
+
review. Mark both conflicting tasks `needs_breakdown` in `INDEX.md` and report the
|
|
147
|
+
conflicting file(s) to the human; do not resolve it yourself by reordering or guessing a
|
|
148
|
+
dependency.
|
|
149
|
+
|
|
130
150
|
**Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
|
|
131
|
-
`.ukit/storage/config.json` (default **
|
|
151
|
+
`.ukit/storage/config.json` (default **10**). A wave with more tasks than that is split
|
|
132
152
|
into consecutive batches of at most that many; finish one batch completely (including 3c
|
|
133
153
|
copy-back and worktree deletion) before starting the next.
|
|
134
154
|
|
|
@@ -139,7 +159,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
|
|
|
139
159
|
|
|
140
160
|
### I3 — Execute wave by wave (code model agents)
|
|
141
161
|
|
|
142
|
-
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default
|
|
162
|
+
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 10), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
|
|
143
163
|
|
|
144
164
|
For each wave:
|
|
145
165
|
|
|
@@ -156,7 +176,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
|
|
|
156
176
|
TDD — mandatory:
|
|
157
177
|
1. Write tests from §Test Cases
|
|
158
178
|
cd .worktrees/task-xxx && <test command>
|
|
159
|
-
Confirm RED
|
|
179
|
+
Confirm RED — paste the actual failing output, don't just assert it happened
|
|
180
|
+
(immediately GREEN → test is wrong — flag this)
|
|
160
181
|
2. Implement → run → confirm GREEN
|
|
161
182
|
3. Run §Verification Commands inside the worktree:
|
|
162
183
|
cd .worktrees/task-xxx && <each verification command>
|
|
@@ -169,6 +190,8 @@ Executor Report (append to task file — do NOT touch INDEX.md):
|
|
|
169
190
|
EXECUTOR_TOOL: <tool>
|
|
170
191
|
EXECUTOR_MODEL: <exact model ID — mandatory>
|
|
171
192
|
EXECUTOR_SUBAGENT: <name or "-">
|
|
193
|
+
RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
|
|
194
|
+
"confirmed" without pasted output is not acceptable>
|
|
172
195
|
Verification Output: <paste full output>
|
|
173
196
|
Status: PASS | FAIL
|
|
174
197
|
Note: <issues or "none">
|
|
@@ -250,7 +273,9 @@ If `git diff` is empty and `git status` is clean → implement was not completed
|
|
|
250
273
|
|
|
251
274
|
### R2 — Model isolation check (strong model, always first)
|
|
252
275
|
|
|
253
|
-
**
|
|
276
|
+
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
|
|
277
|
+
|
|
278
|
+
**Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
|
|
254
279
|
|
|
255
280
|
The spawned reviewer agent reads `EXECUTOR_MODEL` from each task file `## Executor Report`:
|
|
256
281
|
|
|
@@ -33,9 +33,18 @@ Read each `tasks/TASK-xxx.md` for `Dependencies` field:
|
|
|
33
33
|
- Chain A→B→C = 3 waves of 1 task each (sequential, no parallel)
|
|
34
34
|
- Independent A, B, C = 1 wave of 3 tasks (parallel)
|
|
35
35
|
|
|
36
|
+
### Conflict check — mandatory, before spawning any wave
|
|
37
|
+
|
|
38
|
+
For every pair of tasks landing in the same wave, compare their `Target Files` lists. If
|
|
39
|
+
any file path appears in both, STOP — do not spawn that wave. This is the code-level
|
|
40
|
+
safety net for the planner's own §2 Scope constraint (same-wave tasks must not share a
|
|
41
|
+
file) in case it slipped through review. Mark both conflicting tasks `needs_breakdown` in
|
|
42
|
+
`INDEX.md` and report the conflicting file(s) to the human; do not resolve it yourself by
|
|
43
|
+
reordering or guessing a dependency.
|
|
44
|
+
|
|
36
45
|
### Batch each wave — mandatory
|
|
37
46
|
|
|
38
|
-
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **
|
|
47
|
+
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). A wave
|
|
39
48
|
with more tasks than that is split into consecutive batches of at most that many; finish
|
|
40
49
|
one batch completely (including 3c copy-back and worktree deletion) before starting the
|
|
41
50
|
next.
|
|
@@ -68,7 +77,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
|
|
|
68
77
|
TDD — mandatory:
|
|
69
78
|
1. Write tests from §Test Cases
|
|
70
79
|
cd .worktrees/task-xxx && <test command>
|
|
71
|
-
Confirm RED
|
|
80
|
+
Confirm RED — paste the actual failing output, don't just assert it happened
|
|
81
|
+
(immediately GREEN → test is wrong — flag this)
|
|
72
82
|
2. Implement → run → confirm GREEN
|
|
73
83
|
3. Run §Verification Commands inside the worktree:
|
|
74
84
|
cd .worktrees/task-xxx && <each verification command>
|
|
@@ -81,6 +91,8 @@ Executor Report (append to task file — do NOT touch INDEX.md):
|
|
|
81
91
|
EXECUTOR_TOOL: <tool>
|
|
82
92
|
EXECUTOR_MODEL: <exact model ID — mandatory>
|
|
83
93
|
EXECUTOR_SUBAGENT: <name or "-">
|
|
94
|
+
RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
|
|
95
|
+
"confirmed" without pasted output is not acceptable>
|
|
84
96
|
Verification Output: <paste>
|
|
85
97
|
Status: PASS | FAIL
|
|
86
98
|
Note: <issues or "none">
|
|
@@ -28,7 +28,9 @@ If `git diff` is empty and `git status` is clean → handoff-implement was not c
|
|
|
28
28
|
|
|
29
29
|
## Step 2 — Review the diff
|
|
30
30
|
|
|
31
|
-
**
|
|
31
|
+
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
|
|
32
|
+
|
|
33
|
+
**Claude Code — MANDATORY, do this before anything else:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
|
|
32
34
|
|
|
33
35
|
The spawned reviewer agent performs 2a–2d below per task:
|
|
34
36
|
|
|
@@ -115,4 +117,4 @@ Review summary:
|
|
|
115
117
|
- Has fixes → executor re-runs `/ukit:handoff-implement TASK-xxx`
|
|
116
118
|
|
|
117
119
|
> Orchestrator (this session) handles Step 1, 2e, and 3 directly — those are not delegated.
|
|
118
|
-
> Other tools without subagent support: manually switch to the strong model, run 2a–2d
|
|
120
|
+
> Other tools without subagent support: manually switch to the strong model, run 2a–2d per task (one session at a time).
|
|
@@ -77,9 +77,9 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
|
|
|
77
77
|
**Phase 2 — Create Tasks (TDD-embedded, MANDATORY)** (smart/reasoning model, thường cùng phase 1)
|
|
78
78
|
- Human approve plan → AI split `PLAN.md §7` sang nhiều `tasks/TASK-xxx.md`.
|
|
79
79
|
- **Mỗi TASK file BẮT BUỘC có Test Plan của riêng nó**, không chỉ trỏ về PLAN.md. Cụ thể:
|
|
80
|
-
- `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥
|
|
80
|
+
- `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥2 edge case KHÁC loại nhau (vd null/empty + boundary/concurrent, không tính 2 case gần giống nhau) + regression (nếu fix bug).
|
|
81
81
|
- `§ Test Files`: đường dẫn cụ thể file test sẽ tạo/sửa (ví dụ `tests/auth/login.test.js`).
|
|
82
|
-
- `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS.
|
|
82
|
+
- `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS. Nếu project có sẵn lint/typecheck script → BẮT BUỘC liệt kê ở đây, không chỉ lệnh test. Project không có thì ghi rõ N/A, không được bỏ qua im lặng.
|
|
83
83
|
- `§ Acceptance Criteria`: checklist.
|
|
84
84
|
- Nếu split mà task nào không kèm được Test Cases + Test Files cụ thể → task đó chưa đủ `ready`, đánh `needs_breakdown`.
|
|
85
85
|
- Update `INDEX.md`: thêm row mỗi task với status `ready`.
|
|
@@ -88,7 +88,7 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
|
|
|
88
88
|
|
|
89
89
|
**Phase 3 — Implement + Test** (cheap-smart/code model)
|
|
90
90
|
- User: "execute next task" / "làm TASK-001" / "implement task 1".
|
|
91
|
-
- Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT` + verification output) vào cuối task file → đổi status `pending_review`.
|
|
91
|
+
- Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED (paste output failing thật, không chỉ khai đã confirm) → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT`/`RED_OUTPUT` + verification output) vào cuối task file → đổi status `pending_review`.
|
|
92
92
|
- KHÔNG được claim DONE nếu chưa có PASS fresh.
|
|
93
93
|
|
|
94
94
|
**Phase 4 — Review + Test** (reviewer model — KHÁC model executor)
|
|
@@ -108,8 +108,8 @@ A task is `ready` only when it has:
|
|
|
108
108
|
- Dependencies stated
|
|
109
109
|
- **Interfaces** — Consumes/Produces với chữ ký thật (function/endpoint/type), không placeholder;
|
|
110
110
|
`(none)` hợp lệ nếu task không có input/output liên task
|
|
111
|
-
- **Test Plan** (PLAN.md §4) — happy path + ≥
|
|
112
|
-
- Verification command (lệnh executor sẽ chạy)
|
|
111
|
+
- **Test Plan** (PLAN.md §4) — happy path + ≥2 edge case khác loại (+ regression test nếu fix bug); hoặc `N/A` kèm lý do
|
|
112
|
+
- Verification command (lệnh executor sẽ chạy) — PHẢI gồm lint/typecheck nếu project có sẵn
|
|
113
113
|
- Acceptance criteria
|
|
114
114
|
|
|
115
115
|
Missing any → `needs_breakdown`, `blocked`, or `needs_human`.
|
|
@@ -186,11 +186,12 @@
|
|
|
186
186
|
"handoff": {
|
|
187
187
|
"enabled": true,
|
|
188
188
|
"crossTool": true,
|
|
189
|
-
"maxParallelAgents":
|
|
189
|
+
"maxParallelAgents": 10,
|
|
190
190
|
"plan": {
|
|
191
191
|
"requireTestPlan": true,
|
|
192
192
|
"minTestsHappyPath": 1,
|
|
193
|
-
"minTestsEdgeCase":
|
|
193
|
+
"minTestsEdgeCase": 2,
|
|
194
|
+
"requireLintOrTypecheckInVerification": true,
|
|
194
195
|
"regressionTestRequiredForBugfix": true,
|
|
195
196
|
"smartModelHint": "claude-opus-5"
|
|
196
197
|
},
|
|
@@ -479,11 +480,12 @@
|
|
|
479
480
|
"handoff": {
|
|
480
481
|
"enabled": "Bật Quality Gate cho handoff: plan có Test Plan, executor test-first, reviewer model khác. Tắt = quay về flow cũ (dễ lọt lỗi vặt).",
|
|
481
482
|
"crossTool": "true nghĩa là handoff truyền qua file (PLAN/INDEX/tasks) chứ không qua in-process subagent — cho phép plan ở Claude Code, execute ở Kilo Code, review ở Claude Code khác model.",
|
|
482
|
-
"maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 3). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt. Lý do: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc
|
|
483
|
+
"maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 10, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
|
|
483
484
|
"plan": {
|
|
484
485
|
"requireTestPlan": "Bắt buộc PLAN.md §4 phải có Test Plan trước khi task chuyển ready.",
|
|
485
486
|
"minTestsHappyPath": "Tối thiểu test cho happy path.",
|
|
486
|
-
"minTestsEdgeCase": "Tối thiểu test cho edge case (null/empty
|
|
487
|
+
"minTestsEdgeCase": "Tối thiểu test cho edge case (mặc định 2, tăng từ 1 — phải khác loại nhau, vd null/empty + boundary/concurrent, không tính 2 test gần giống nhau là đủ).",
|
|
488
|
+
"requireLintOrTypecheckInVerification": "Nếu project có sẵn lint/typecheck script (package.json hoặc tương đương), Verification Commands của mỗi task BẮT BUỘC phải gồm lệnh đó — không chỉ chạy test. Project không có lint/typecheck thì planner phải ghi rõ lý do N/A thay vì bỏ qua im lặng.",
|
|
487
489
|
"regressionTestRequiredForBugfix": "Bug fix phải có regression test fail-trước-fix.",
|
|
488
490
|
"smartModelHint": "Gợi ý model mạnh nhất cho phase plan (ví dụ claude-opus-5). UKit không tự ép, chỉ ghi hint vào task."
|
|
489
491
|
},
|