@ngockhoale/ukit 2.0.6 → 2.0.7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,21 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.0.7 - 2026-08-15
6
+
7
+ ### Added
8
+
9
+ - **`/ukit:handoff-create` and `/ukit:handoff-fullstack` gain a P2.5 independent plan review gate.** After the planner writes `PLAN.md` and before it's committed, a separate `code-reviewer` agent invocation (`REVIEW_TARGET_TYPE=plan`, fresh context — not the planner reviewing its own work) checks Completeness/Consistency/Clarity/Scope/YAGNI. `Issues Found` routes back to the planner to revise and resubmit; `Approved` unlocks the commit. Each round is logged to PLAN.md's new `## Plan Review Log` section, and the loop is capped at 3 `Issues Found` rounds — past that it escalates to the human instead of looping indefinitely.
10
+ - **Same-wave file-conflict precheck in Phase 3.** Before spawning any implementation wave, `handoff-implement.md`/`handoff-fullstack.md` now compare `Target Files` across every task in that wave; any overlap stops the wave and marks both tasks `needs_breakdown` instead of risking a mid-wave merge conflict. Code-level backstop for the planner's existing "no shared files in one wave" scope constraint.
11
+ - **TDD RED-state now requires real evidence.** The Executor Report template adds a mandatory `RED_OUTPUT` field — the actual failing-test output pasted before implementation, not a bare "confirmed" claim. `code-reviewer` checks this field during Test Plan adherence review and requests changes if it's missing or vague.
12
+ - **`handoff.plan.requireLintOrTypecheckInVerification`** (default `true`): when the project has a lint/typecheck script, task Verification Commands must run it, not just tests; the planner must state an explicit N/A if the project has none. Enforced by `code-reviewer` as review-order step 0.
13
+
14
+ ### Changed
15
+
16
+ - **`handoff.maxParallelAgents`**: 3 → 10, raising the default cap on concurrent background agents per wave (config comments still warn against exceeding ~10-15 — each agent's report gets injected back into the orchestrator's own context on completion).
17
+ - **Phase 4 (Review) is now parallelized**, batched by `maxParallelAgents` the same way Phase 3 (Implement) already was — reviewer agents only read a diff and append a verdict to their own task file, so parallel review carries none of Phase 3's shared-worktree conflict risk.
18
+ - **`handoff.plan.minTestsEdgeCase`**: 1 → 2, and the two edge cases must now be of different kinds (e.g. null/empty **and** boundary/concurrent) — two near-duplicate cases no longer satisfy the requirement. Wired through `handoff-planner`, `code-reviewer`, `feature-implementer`'s inline-test-plan fallback, `RULES.md`, and both handoff pipeline commands.
19
+
5
20
  ## 2.0.6 - 2026-08-12
6
21
 
7
22
  ### Changed
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.0.6",
3
+ "version": "2.0.7",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, Antigravity, OpenAI Codex, and OpenCode.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -27,7 +27,8 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
27
27
 
28
28
  ### Review order
29
29
 
30
- 1. **Test Plan adherence** — Were all tests in §4 actually implemented? Run them yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
30
+ 0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
31
+ 1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
31
32
  2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
32
33
  3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider test suite if shared code was touched.
33
34
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
@@ -112,7 +113,10 @@ Only flag issues that would cause real problems during implementation planning.
112
113
 
113
114
  ### Output
114
115
 
116
+ Append (do NOT overwrite) this block to the end of the reviewed document under a `## Plan Review Log` section — create the section if it doesn't exist yet, keep all prior round entries:
117
+
115
118
  ```
119
+ ### Round <N> — <YYYY-MM-DD> · <your model>
116
120
  Status: Approved | Issues Found
117
121
 
118
122
  COMPLETENESS:
@@ -128,3 +132,5 @@ YAGNI:
128
132
 
129
133
  NOTES: [1-2 sentences if needed]
130
134
  ```
135
+
136
+ `<N>` = 1 + however many `### Round` entries already exist in the log (1 if this is the first review).
@@ -30,7 +30,7 @@ If unsure, ask the user. Don't apply Handoff mode rules to a quick one-off fix.
30
30
  ### 2. Plan Approach (< 1 minute)
31
31
 
32
32
  - List files to create/modify (max diff).
33
- - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥1 edge case; regression test if fixing a bug). In daily mode, skip this step.
33
+ - **Handoff mode only:** if no Test Plan exists in the task file and task is not `trivial`, write one inline before implementing (happy + ≥2 edge cases of different kinds; regression test if fixing a bug). In daily mode, skip this step.
34
34
 
35
35
  ### 3. Test First (RED) — Handoff mode
36
36
 
@@ -37,9 +37,15 @@ Write all 7 sections to `docs/AI_HANDOFF/PLAN.md`:
37
37
  CONSTRAINT: tasks in the same wave must not modify the same file.
38
38
  If two tasks need the same file → make one depend on the other.
39
39
  §3 Approach — technical solution, trade-offs, alternatives rejected
40
- §4 Test Plan — happy path × N + ≥1 edge case + regression (if bugfix)
40
+ §4 Test Plan — happy path × N + ≥2 edge cases of DIFFERENT kinds (e.g. null/empty AND
41
+ boundary/concurrent — two near-duplicate cases do not satisfy this) +
42
+ regression (if bugfix)
41
43
  table: | Type | Test Name | Expected |
42
- §5 Verification — exact shell commands executor will run
44
+ §5 Verification — exact shell commands executor will run. If the project has a lint or
45
+ typecheck script (check package.json `scripts`, or the equivalent for
46
+ the project's stack), it MUST be included here, not just the test
47
+ command. If the project genuinely has none, state that explicitly —
48
+ do not omit silently.
43
49
  §6 Acceptance — checklist of done criteria (prefer verifiable/command-based criteria)
44
50
  §7 Global Constraints — one line each: version floors, dependency limits, naming/copy
45
51
  rules, platform requirements. Every TASK-xxx.md inherits this section
@@ -55,6 +61,8 @@ Append this footer to `PLAN.md` — mandatory, checked by a hook before the writ
55
61
  PLANNER_MODEL: <your exact model ID — e.g. claude-opus-5>
56
62
  ```
57
63
 
64
+ Your output does not go straight to implementation: an independent `code-reviewer` pass (`REVIEW_TARGET_TYPE=plan`) reviews `PLAN.md` next. If it returns `Issues Found`, you'll be re-invoked to revise and resubmit — write §1-§7 tight enough to pass on the first pass.
65
+
58
66
  ## Phase 2 — Split into TASK-xxx.md
59
67
 
60
68
  **Right-sizing rule:** A task is the smallest unit that carries its own test cycle and is worth a fresh reviewer's gate. Split only where a reviewer could meaningfully approve one task while rejecting its neighbor. Each task ends with an independently testable deliverable.
@@ -67,9 +75,9 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
67
75
  |-------|------|
68
76
  | Target Files | Exact paths — no two tasks in same wave share a file |
69
77
  | Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
70
- | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥1 edge case |
78
+ | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
71
79
  | Test Files | Exact test file paths to create/modify |
72
- | Verification Commands | Runnable shell commands |
80
+ | Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
73
81
  | Acceptance Criteria | Verifiable checklist |
74
82
 
75
83
  Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
@@ -48,7 +48,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
48
48
  - §1 Intent — problem + success definition
49
49
  - §2 Scope — in / out of scope. **Add a constraint**: same-wave tasks must not modify the same file (prevents merge conflicts). If two tasks need the same file, make one depend on the other.
50
50
  - §3 Approach — solution, trade-offs, alternatives rejected
51
- - §4 Test Plan — happy path + ≥1 edge case + regression if bugfix (non-negotiable)
51
+ - §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
52
52
  - §5 Verification — exact shell commands executor will run
53
53
  - §6 Acceptance — done checklist (prefer verifiable criteria with commands)
54
54
 
@@ -62,7 +62,7 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
62
62
  Every task MUST have:
63
63
  - Target Files (exact paths — no two tasks in same wave share a file)
64
64
  - Dependencies (`TASK-xxx` or `none` — wave structure inferred from this, not stored separately)
65
- - Test Cases (Type | Name | Expected — ≥1 happy + ≥1 edge case)
65
+ - Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
66
66
  - Test Files (exact paths)
67
67
  - Verification Commands (runnable shell commands)
68
68
  - Acceptance Criteria (verifiable checklist)
@@ -85,4 +85,18 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
85
85
 
86
86
  ---
87
87
 
88
+ ## Step 2.5 — Independent plan review (strong model, separate agent)
89
+
90
+ **Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet). If count ≥ 3 (3 prior rounds already returned `Issues Found`), do NOT invoke the reviewer again — STOP and escalate to the human: show the accumulated findings from all 3 rounds, ask them to manually revise PLAN.md, narrow scope, or explicitly approve an override. Otherwise, proceed below.
91
+
92
+ **Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
93
+
94
+ 1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
95
+ 2. `Issues Found` → route back to Step 2: planner revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another Step 2.5 review (this becomes the next round). Do NOT commit or hand off to executor on `Issues Found`.
96
+ 3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed.
97
+
98
+ > Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from Step 2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
99
+
100
+ ---
101
+
88
102
  **Next:** switch to code model → `/ukit:handoff-implement`
@@ -57,7 +57,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
57
57
  - §1 Intent — problem + success definition
58
58
  - §2 Scope — in / out of scope; same-wave tasks must not modify the same file (prevents merge conflicts)
59
59
  - §3 Approach — solution, trade-offs, alternatives rejected
60
- - §4 Test Plan — happy path + ≥1 edge case + regression if bugfix (non-negotiable)
60
+ - §4 Test Plan — happy path + ≥2 edge cases of different kinds + regression if bugfix (non-negotiable)
61
61
  - §5 Verification — exact shell commands executor will run
62
62
  - §6 Acceptance — done checklist (prefer verifiable criteria with commands)
63
63
 
@@ -71,7 +71,7 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
71
71
  Every task MUST have:
72
72
  - Target Files (exact paths — no two tasks in same wave share a file)
73
73
  - Dependencies (`TASK-xxx` or `none` — wave structure inferred from this)
74
- - Test Cases (Type | Name | Expected — ≥1 happy + ≥1 edge case)
74
+ - Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
75
75
  - Test Files (exact paths)
76
76
  - Verification Commands (runnable shell commands)
77
77
  - Acceptance Criteria (verifiable checklist)
@@ -89,6 +89,18 @@ The planner agent does the following (use P1 summary — do NOT re-read files):
89
89
 
90
90
  7. **Report:** task IDs, dependency graph, any `needs_breakdown` tasks + reason.
91
91
 
92
+ ### P2.5 — Independent plan review (strong model, separate agent)
93
+
94
+ **Loop cap — check first:** count `### Round` entries in PLAN.md's `## Plan Review Log` (0 if the section doesn't exist yet). If count ≥ 3 (3 prior rounds already returned `Issues Found`), do NOT invoke the reviewer again — STOP and escalate to the human: show the accumulated findings from all 3 rounds, ask them to manually revise PLAN.md, narrow scope, or explicitly approve an override. Otherwise, proceed below.
95
+
96
+ **Claude Code — MANDATORY, do this before anything else in P2.5:** call the Agent tool with `subagent_type: "code-reviewer"`, passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from P2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
97
+
98
+ 1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
99
+ 2. `Issues Found` → route back to P2: `handoff-planner` revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another P2.5 review (this becomes the next round). Do NOT proceed to P3 on `Issues Found`.
100
+ 3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed to P3.
101
+
102
+ > Other tools without subagent support: manually switch to the strong model in a **separate** chat/session from P2, paste PLAN.md, review using the Spec/Plan Review checklist in `.claude/agents/code-reviewer.md`.
103
+
92
104
  ### P3 — Commit the plan (lite model)
93
105
 
94
106
  **Claude Code — MANDATORY:** call the Agent tool with `subagent_type: "ukit-small-task-maintainer"` for this commit step (lite tier — haiku/unic-lite). Run:
@@ -127,8 +139,16 @@ Read each `docs/AI_HANDOFF/tasks/TASK-xxx.md` for `Dependencies` field:
127
139
  - Chain A→B→C = 3 waves of 1 task each (sequential)
128
140
  - Independent A, B, C = 1 wave of 3 tasks (parallel)
129
141
 
142
+ **Conflict check — mandatory, before spawning any wave.** For every pair of tasks landing
143
+ in the same wave, compare their `Target Files` lists. If any file path appears in both,
144
+ STOP — do not spawn that wave. This is the code-level safety net for the planner's own §2
145
+ Scope constraint (same-wave tasks must not share a file) in case it slipped through
146
+ review. Mark both conflicting tasks `needs_breakdown` in `INDEX.md` and report the
147
+ conflicting file(s) to the human; do not resolve it yourself by reordering or guessing a
148
+ dependency.
149
+
130
150
  **Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
131
- `.ukit/storage/config.json` (default **3**). A wave with more tasks than that is split
151
+ `.ukit/storage/config.json` (default **10**). A wave with more tasks than that is split
132
152
  into consecutive batches of at most that many; finish one batch completely (including 3c
133
153
  copy-back and worktree deletion) before starting the next.
134
154
 
@@ -139,7 +159,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
139
159
 
140
160
  ### I3 — Execute wave by wave (code model agents)
141
161
 
142
- **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 3), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
162
+ **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 10), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
143
163
 
144
164
  For each wave:
145
165
 
@@ -156,7 +176,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
156
176
  TDD — mandatory:
157
177
  1. Write tests from §Test Cases
158
178
  cd .worktrees/task-xxx && <test command>
159
- Confirm RED (immediately GREEN → test is wrong — flag this)
179
+ Confirm RED — paste the actual failing output, don't just assert it happened
180
+ (immediately GREEN → test is wrong — flag this)
160
181
  2. Implement → run → confirm GREEN
161
182
  3. Run §Verification Commands inside the worktree:
162
183
  cd .worktrees/task-xxx && <each verification command>
@@ -169,6 +190,8 @@ Executor Report (append to task file — do NOT touch INDEX.md):
169
190
  EXECUTOR_TOOL: <tool>
170
191
  EXECUTOR_MODEL: <exact model ID — mandatory>
171
192
  EXECUTOR_SUBAGENT: <name or "-">
193
+ RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
194
+ "confirmed" without pasted output is not acceptable>
172
195
  Verification Output: <paste full output>
173
196
  Status: PASS | FAIL
174
197
  Note: <issues or "none">
@@ -250,7 +273,9 @@ If `git diff` is empty and `git status` is clean → implement was not completed
250
273
 
251
274
  ### R2 — Model isolation check (strong model, always first)
252
275
 
253
- **Claude Code — MANDATORY, do this before anything else in R2–R4:** for each `pending_review` task, call the Agent tool with `subagent_type: "code-reviewer"`. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
276
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
277
+
278
+ **Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
254
279
 
255
280
  The spawned reviewer agent reads `EXECUTOR_MODEL` from each task file `## Executor Report`:
256
281
 
@@ -33,9 +33,18 @@ Read each `tasks/TASK-xxx.md` for `Dependencies` field:
33
33
  - Chain A→B→C = 3 waves of 1 task each (sequential, no parallel)
34
34
  - Independent A, B, C = 1 wave of 3 tasks (parallel)
35
35
 
36
+ ### Conflict check — mandatory, before spawning any wave
37
+
38
+ For every pair of tasks landing in the same wave, compare their `Target Files` lists. If
39
+ any file path appears in both, STOP — do not spawn that wave. This is the code-level
40
+ safety net for the planner's own §2 Scope constraint (same-wave tasks must not share a
41
+ file) in case it slipped through review. Mark both conflicting tasks `needs_breakdown` in
42
+ `INDEX.md` and report the conflicting file(s) to the human; do not resolve it yourself by
43
+ reordering or guessing a dependency.
44
+
36
45
  ### Batch each wave — mandatory
37
46
 
38
- Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **3**). A wave
47
+ Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). A wave
39
48
  with more tasks than that is split into consecutive batches of at most that many; finish
40
49
  one batch completely (including 3c copy-back and worktree deletion) before starting the
41
50
  next.
@@ -68,7 +77,8 @@ Read docs/AI_HANDOFF/tasks/TASK-xxx.md
68
77
  TDD — mandatory:
69
78
  1. Write tests from §Test Cases
70
79
  cd .worktrees/task-xxx && <test command>
71
- Confirm RED (immediately GREEN → test is wrong — flag this)
80
+ Confirm RED — paste the actual failing output, don't just assert it happened
81
+ (immediately GREEN → test is wrong — flag this)
72
82
  2. Implement → run → confirm GREEN
73
83
  3. Run §Verification Commands inside the worktree:
74
84
  cd .worktrees/task-xxx && <each verification command>
@@ -81,6 +91,8 @@ Executor Report (append to task file — do NOT touch INDEX.md):
81
91
  EXECUTOR_TOOL: <tool>
82
92
  EXECUTOR_MODEL: <exact model ID — mandatory>
83
93
  EXECUTOR_SUBAGENT: <name or "-">
94
+ RED_OUTPUT: <paste the actual failing-test output from step 1 — a claim like
95
+ "confirmed" without pasted output is not acceptable>
84
96
  Verification Output: <paste>
85
97
  Status: PASS | FAIL
86
98
  Note: <issues or "none">
@@ -28,7 +28,9 @@ If `git diff` is empty and `git status` is clean → handoff-implement was not c
28
28
 
29
29
  ## Step 2 — Review the diff
30
30
 
31
- **Claude Code — MANDATORY, do this before anything else:** for each `pending_review` task, call the Agent tool with `subagent_type: "code-reviewer"`. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
31
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
32
+
33
+ **Claude Code — MANDATORY, do this before anything else:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
32
34
 
33
35
  The spawned reviewer agent performs 2a–2d below per task:
34
36
 
@@ -115,4 +117,4 @@ Review summary:
115
117
  - Has fixes → executor re-runs `/ukit:handoff-implement TASK-xxx`
116
118
 
117
119
  > Orchestrator (this session) handles Step 1, 2e, and 3 directly — those are not delegated.
118
- > Other tools without subagent support: manually switch to the strong model, run 2a–2d sequentially per task.
120
+ > Other tools without subagent support: manually switch to the strong model, run 2a–2d per task (one session at a time).
@@ -77,9 +77,9 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
77
77
  **Phase 2 — Create Tasks (TDD-embedded, MANDATORY)** (smart/reasoning model, thường cùng phase 1)
78
78
  - Human approve plan → AI split `PLAN.md §7` sang nhiều `tasks/TASK-xxx.md`.
79
79
  - **Mỗi TASK file BẮT BUỘC có Test Plan của riêng nó**, không chỉ trỏ về PLAN.md. Cụ thể:
80
- - `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥1 edge case + regression (nếu fix bug).
80
+ - `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥2 edge case KHÁC loại nhau (vd null/empty + boundary/concurrent, không tính 2 case gần giống nhau) + regression (nếu fix bug).
81
81
  - `§ Test Files`: đường dẫn cụ thể file test sẽ tạo/sửa (ví dụ `tests/auth/login.test.js`).
82
- - `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS.
82
+ - `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS. Nếu project có sẵn lint/typecheck script → BẮT BUỘC liệt kê ở đây, không chỉ lệnh test. Project không có thì ghi rõ N/A, không được bỏ qua im lặng.
83
83
  - `§ Acceptance Criteria`: checklist.
84
84
  - Nếu split mà task nào không kèm được Test Cases + Test Files cụ thể → task đó chưa đủ `ready`, đánh `needs_breakdown`.
85
85
  - Update `INDEX.md`: thêm row mỗi task với status `ready`.
@@ -88,7 +88,7 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
88
88
 
89
89
  **Phase 3 — Implement + Test** (cheap-smart/code model)
90
90
  - User: "execute next task" / "làm TASK-001" / "implement task 1".
91
- - Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT` + verification output) vào cuối task file → đổi status `pending_review`.
91
+ - Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED (paste output failing thật, không chỉ khai đã confirm) → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT`/`RED_OUTPUT` + verification output) vào cuối task file → đổi status `pending_review`.
92
92
  - KHÔNG được claim DONE nếu chưa có PASS fresh.
93
93
 
94
94
  **Phase 4 — Review + Test** (reviewer model — KHÁC model executor)
@@ -108,8 +108,8 @@ A task is `ready` only when it has:
108
108
  - Dependencies stated
109
109
  - **Interfaces** — Consumes/Produces với chữ ký thật (function/endpoint/type), không placeholder;
110
110
  `(none)` hợp lệ nếu task không có input/output liên task
111
- - **Test Plan** (PLAN.md §4) — happy path + ≥1 edge case (+ regression test nếu fix bug); hoặc `N/A` kèm lý do
112
- - Verification command (lệnh executor sẽ chạy)
111
+ - **Test Plan** (PLAN.md §4) — happy path + ≥2 edge case khác loại (+ regression test nếu fix bug); hoặc `N/A` kèm lý do
112
+ - Verification command (lệnh executor sẽ chạy) — PHẢI gồm lint/typecheck nếu project có sẵn
113
113
  - Acceptance criteria
114
114
 
115
115
  Missing any → `needs_breakdown`, `blocked`, or `needs_human`.
@@ -186,11 +186,12 @@
186
186
  "handoff": {
187
187
  "enabled": true,
188
188
  "crossTool": true,
189
- "maxParallelAgents": 3,
189
+ "maxParallelAgents": 10,
190
190
  "plan": {
191
191
  "requireTestPlan": true,
192
192
  "minTestsHappyPath": 1,
193
- "minTestsEdgeCase": 1,
193
+ "minTestsEdgeCase": 2,
194
+ "requireLintOrTypecheckInVerification": true,
194
195
  "regressionTestRequiredForBugfix": true,
195
196
  "smartModelHint": "claude-opus-5"
196
197
  },
@@ -479,11 +480,12 @@
479
480
  "handoff": {
480
481
  "enabled": "Bật Quality Gate cho handoff: plan có Test Plan, executor test-first, reviewer model khác. Tắt = quay về flow cũ (dễ lọt lỗi vặt).",
481
482
  "crossTool": "true nghĩa là handoff truyền qua file (PLAN/INDEX/tasks) chứ không qua in-process subagent — cho phép plan ở Claude Code, execute ở Kilo Code, review ở Claude Code khác model.",
482
- "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 3). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt. Lý do: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc là cách nhanh nhất làm session chính vượt context window. Hạ xuống 2 nếu task nặng; không nên vượt 3.",
483
+ "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 10, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
483
484
  "plan": {
484
485
  "requireTestPlan": "Bắt buộc PLAN.md §4 phải có Test Plan trước khi task chuyển ready.",
485
486
  "minTestsHappyPath": "Tối thiểu test cho happy path.",
486
- "minTestsEdgeCase": "Tối thiểu test cho edge case (null/empty/boundary/concurrent…).",
487
+ "minTestsEdgeCase": "Tối thiểu test cho edge case (mặc định 2, tăng từ 1 — phải khác loại nhau, vd null/empty + boundary/concurrent, không tính 2 test gần giống nhau là đủ).",
488
+ "requireLintOrTypecheckInVerification": "Nếu project có sẵn lint/typecheck script (package.json hoặc tương đương), Verification Commands của mỗi task BẮT BUỘC phải gồm lệnh đó — không chỉ chạy test. Project không có lint/typecheck thì planner phải ghi rõ lý do N/A thay vì bỏ qua im lặng.",
487
489
  "regressionTestRequiredForBugfix": "Bug fix phải có regression test fail-trước-fix.",
488
490
  "smartModelHint": "Gợi ý model mạnh nhất cho phase plan (ví dụ claude-opus-5). UKit không tự ép, chỉ ghi hint vào task."
489
491
  },