@ngockhoale/ukit 2.0.6 → 2.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -8,7 +8,26 @@
8
8
  $ARGUMENTS
9
9
  _Empty = all `pending_review` tasks. Or: "TASK-001" for a specific task._
10
10
 
11
- > **NO git commit. NO git push.** Reviewer only reads and reports. Human decides when to commit.
11
+ > **The reviewer agents never commit.** The orchestrator does: it checkpoints each fix round
12
+ > and, once the set is approved, pushes. Force-push stays denied.
13
+
14
+ ---
15
+
16
+ ## Autonomy Contract — read before anything else
17
+
18
+ Review every `pending_review` task and drive the set to a decision in this one invocation.
19
+ The human is not watching.
20
+
21
+ **Ask nothing after the run starts.** Findings do not end the run — they start the auto-fix
22
+ loop (Step 3). Escalate only for a blocker outside the repo, or a task still failing after
23
+ both fix rounds, and even then finish everything else first.
24
+
25
+ Never pause mid-review to ask for `/compact`. Reviewer agents return ≤6 lines and write the
26
+ full verdict to disk; a compaction request belongs at the end of the command, never inside it.
27
+
28
+ Maintain `docs/AI_HANDOFF/RUN.md` (`Command: handoff-review`, `Phase:`, `Cursor:`, `Next:`)
29
+ after every step. If it already exists with `Phase:` not `done`, this is a **continuation**:
30
+ jump to `Next:` instead of restarting, and do not ask whether to continue.
12
31
 
13
32
  ---
14
33
 
@@ -18,17 +37,22 @@ Read `docs/AI_HANDOFF/ACTIVE.md` → get `Base: <BASE>`.
18
37
  Read `docs/AI_HANDOFF/INDEX.md` → collect `pending_review` tasks (or specific task).
19
38
  If none → report and stop.
20
39
 
21
- Verify there are uncommitted changes from handoff-implement:
40
+ `/ukit:handoff-implement` commits one commit per wave, so the changes under review are in
41
+ commits since the plan commit, not only in the working tree. Resolve the review range once:
22
42
  ```bash
23
- git status # should show modified/new files, no handoff branches/worktrees
24
- git diff --stat # summary of all changes
43
+ PLAN_COMMIT=$(git log --format=%H --grep='^handoff: plan' -n 1)
44
+ git diff --stat $PLAN_COMMIT # every handoff change: wave commits + working tree
45
+ git status # no handoff branches/worktrees should remain
25
46
  ```
26
47
 
27
- If `git diff` is empty and `git status` is clean → handoff-implement was not completed. Stop and re-run it.
48
+ If that diff is empty → handoff-implement was not completed. Report which tasks are still
49
+ `ready` and stop; there is nothing to review yet.
28
50
 
29
51
  ## Step 2 — Review the diff
30
52
 
31
- **Claude Code — MANDATORY, do this before anything else:** for each `pending_review` task, call the Agent tool with `subagent_type: "code-reviewer"`. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
53
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
54
+
55
+ **Claude Code — MANDATORY, do this before anything else:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
32
56
 
33
57
  The spawned reviewer agent performs 2a–2d below per task:
34
58
 
@@ -49,13 +73,14 @@ Read `EXECUTOR_MODEL` from each task file `## Executor Report`.
49
73
  <each command from §Verification Commands across all tasks>
50
74
  ```
51
75
 
52
- If any command fails → verdict: `critical_block`. Stop.
76
+ If any command fails → verdict: `critical_block` for this task. That ends **this task's**
77
+ review, not the run — the orchestrator picks it up in Step 3's auto-fix loop.
53
78
 
54
79
  ### 2c — Review unified diff
55
80
 
56
81
  ```bash
57
- git diff # all uncommitted changes vs $BASE
58
- git status # overview of modified/new/deleted files
82
+ git diff $PLAN_COMMIT # every handoff change: wave commits + working tree
83
+ git status # overview of modified/new/deleted files
59
84
  ```
60
85
 
61
86
  Review as one unified diff — correctness, regression risk, security, edge cases, maintainability.
@@ -76,43 +101,89 @@ FINDINGS:
76
101
  NEXT_STATUS_FOR_INDEX: <status>
77
102
  ```
78
103
 
104
+ Then return to the orchestrator **at most 6 lines** — the full verdict is already on disk:
105
+ ```
106
+ TASK: TASK-xxx
107
+ VERDICT: approved | approved_minor | changes_requested | critical_block
108
+ REVIEWER_MODEL: <exact model ID>
109
+ VERIFICATION_RERUN: PASS | FAIL
110
+ BLOCKING: <one line per critical/important finding, or "none">
111
+ ```
112
+ Do not paste the diff, the findings prose, or verification logs back into the orchestrator.
113
+
79
114
  ### 2e — Orchestrator updates INDEX.md
80
115
 
81
116
  Set `status = NEXT_STATUS_FOR_INDEX`. Orchestrator writes INDEX, not the reviewer.
82
117
 
83
- ## Step 3 — After all tasks reviewed
84
-
85
- **All `approved` or `approved_minor`:**
86
-
87
- Update INDEX.md: all approved tasks → `done`.
88
-
89
- **DO NOT commit. DO NOT push.** Leave the working tree as-is for human review.
90
-
91
- Report:
92
- ```
93
- ✅ All tasks approved.
94
- Uncommitted changes are in the working tree — human commits when ready.
95
- git diff --stat ← review summary
96
- git add . && git commit -m "..." ← when ready
118
+ ## Step 3 — Auto-fix loop (max 2 rounds)
119
+
120
+ Any task with `changes_requested` or `critical_block` enters this loop. **Do not report to
121
+ the human and stop** — an unattended review that ends in a to-do list is what stretches a
122
+ cycle across days.
123
+
124
+ For round in 1..2:
125
+
126
+ 1. Collect every task not yet `approved`/`approved_minor`. If none, exit the loop.
127
+ 2. Group them into waves by `Dependencies`, auto-serializing any pair that shares a target
128
+ file. Spawn one `feature-implementer` per task **in parallel**, at most
129
+ `handoff.maxParallelAgents`, each in its own worktree. Each agent gets:
130
+ - the task file path — it reads its own `## Reviewer Verdict` from disk, so do not paste
131
+ findings into the prompt;
132
+ - the instruction: fix every `critical` and `important` finding, leave `minor` alone
133
+ unless trivially safe, keep existing tests green, add a regression test per critical
134
+ finding, re-run §Verification Commands, and append `## Executor Report (fix round <N>)`
135
+ to the task file;
136
+ - a ≤10-line return contract (`TASK` / `STATUS` / `EXECUTOR_MODEL` / `FILES` / `VERIFY` /
137
+ `NOTE`) — no pasted logs.
138
+ 3. Copy changes back, delete worktrees, then checkpoint:
139
+ ```bash
140
+ git add -A && git commit -m "handoff: fix round <N> — TASK-00x"
141
+ ```
142
+ 4. Re-review **only the tasks touched this round**, in parallel, per 2a–2e. The reviewer must
143
+ still differ from the executor model and still re-runs verification itself.
144
+ 5. Update `INDEX.md` and `RUN.md`. Exit when everything is approved.
145
+
146
+ After round 2, a task still failing is genuinely stuck: leave it `blocked`, record why in its
147
+ `## Discussion` thread, and name it in the report. **Every other task still proceeds to Step
148
+ 4** — one stubborn task must not hold the rest hostage.
149
+
150
+ ## Step 4 — Commit + push
151
+
152
+ Update `INDEX.md`: approved tasks → `done`; still-failing tasks → `blocked`.
153
+
154
+ Commit anything left uncommitted, then push once:
155
+ ```bash
156
+ git add -A && git commit -m "handoff: review <goal>" # skip if nothing left to commit
157
+ git push origin $BASE
97
158
  ```
98
159
 
99
- **Any `changes_requested` or `critical_block`:**
160
+ Push without asking — `git push` is in the settings `allow` list precisely so an unattended
161
+ run never stalls at the last step. Force-push stays denied, so this can only fast-forward.
162
+
163
+ If some tasks are `blocked`, still push: the approved work is reviewed, tested and
164
+ committed, and the per-wave commits make any subset revertible.
100
165
 
101
- Report which files need fixes. Executor re-runs `/ukit:handoff-implement TASK-xxx`.
102
- Changes from the failed task remain in the working tree — executor fixes them in place (no new worktree needed for simple fixes).
166
+ Set `Phase: done` in `docs/AI_HANDOFF/RUN.md`.
103
167
 
104
- ## Step 4 — Report
168
+ ## Step 5 — Report
105
169
 
106
170
  ```
107
171
  Review summary:
108
172
  TASK-001: approved
109
- TASK-002: changes_requested — see §Discussion
110
- Working tree: unchanged — human commits when satisfied
173
+ TASK-002: approved (fix round 1)
174
+ TASK-003: blocked — <one-line reason>
175
+ Fix rounds used: <0|1|2>
176
+ Git: <K> commits pushed to <BASE>
177
+ Blocked (needs you): <TASK-xxx> | none
111
178
  ```
112
179
 
113
- **Next:**
114
- - All approved → human reviews `git diff` and commits manually.
115
- - Has fixes → executor re-runs `/ukit:handoff-implement TASK-xxx`
180
+ Then, as the last line:
181
+
182
+ > Cycle finished. Run `/compact` before starting the next cycle — this session carries the
183
+ > whole review history and the next cycle deserves a clean window.
184
+
185
+ This is the only place this command may ask for a compaction. All state lives in git,
186
+ `INDEX.md` and `RUN.md`, so a compacted or fresh session resumes with nothing lost.
116
187
 
117
188
  > Orchestrator (this session) handles Step 1, 2e, and 3 directly — those are not delegated.
118
- > Other tools without subagent support: manually switch to the strong model, run 2a–2d sequentially per task.
189
+ > Other tools without subagent support: manually switch to the strong model, run 2a–2d per task (one session at a time).
@@ -164,6 +164,28 @@ const now = Date.now();
164
164
  const withinCooldown = guardState.lastPhase === phase
165
165
  && (now - Number(guardState.lastActionAt || 0)) < COOLDOWN_MS;
166
166
 
167
+ // A one-shot pipeline (handoff-fullstack) is mid-flight when RUN.md exists with a phase
168
+ // other than `done`. Telling the agent to stop and ask for /compact there is actively
169
+ // harmful: the human is deliberately away, the pipeline has a disk cursor that makes it
170
+ // resumable, and a mid-cycle halt is what used to strand a cycle for days. The right
171
+ // directive during a run is "shed context at the next wave boundary and keep going" —
172
+ // compaction is requested once, at the cycle boundary, by the command's Final Report.
173
+ function readRunCursor() {
174
+ try {
175
+ const text = fs.readFileSync(path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md'), 'utf8');
176
+ const runPhase = (text.match(/^Phase:\s*(.+)$/m)?.[1] || '').trim();
177
+ if (!runPhase || /^done$/i.test(runPhase)) return null;
178
+ return {
179
+ phase: runPhase,
180
+ cursor: (text.match(/^Cursor:\s*(.+)$/m)?.[1] || '').trim(),
181
+ };
182
+ } catch {
183
+ return null;
184
+ }
185
+ }
186
+
187
+ const run = readRunCursor();
188
+
167
189
  const lines_out = [];
168
190
  if (ratio >= 1) {
169
191
  lines_out.push(`UKIT CONTEXT ALERT — live context ~${estimatedTokens.toLocaleString()} tokens, at/over the ${hardCap.toLocaleString()} cap (${pct}%).`);
@@ -174,10 +196,21 @@ if (ratio >= 1) {
174
196
 
175
197
  if (withinCooldown) {
176
198
  lines_out.push('(Directive already issued this phase — act on it now if not already done; this is just a reminder.)');
199
+ } else if (run) {
200
+ lines_out.push(`A handoff run is mid-flight (RUN.md — Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}). DO NOT stop, DO NOT ask the user to /compact, DO NOT re-scope the remaining work into a new cycle.`);
201
+ lines_out.push('ACTION THIS TURN — shed context, then keep going:');
202
+ lines_out.push('1) Finish the step in flight, then update docs/AI_HANDOFF/RUN.md so the cursor is current.');
203
+ lines_out.push('2) At the next wave boundary, collapse the finished wave to one line per task and stop re-reading task files, executor reports, and earlier diffs — that detail is already on disk.');
204
+ lines_out.push('3) Make sure spawned agents are returning their short summary form (<=10 lines), not pasted RED/verification logs. Pasted subagent logs are the usual cause of this warning.');
205
+ lines_out.push('4) Continue to the next wave/phase. The user is asked to /compact once, at the end of the cycle — not now.');
206
+ if (sidechainEntries > 0) {
207
+ lines_out.push(`Note: ${sidechainEntries} subagent entries in this stretch. Keep concurrency at or below handoff.maxParallelAgents and keep returns short; do not widen the batch while this warning stands.`);
208
+ }
209
+ writeGuardState({ lastPhase: phase, lastActionAt: now });
177
210
  } else {
178
211
  lines_out.push('ACTION THIS TURN, before starting new investigation/subagents/pipeline phases:');
179
212
  lines_out.push('1) Persist current progress now (update docs/STATUS.md, e.g. via the update-status skill) so nothing is lost.');
180
- lines_out.push('2) If the remaining work still has multiple steps or pipeline phases left (e.g. a handoff-fullstack run with implement/review/deploy still ahead), split what remains into bounded docs/AI_HANDOFF tasks instead of continuing in this one thread — smaller phases keep each turn under the cap.');
213
+ lines_out.push('2) If the remaining work still has multiple steps or pipeline phases left, split what remains into bounded docs/AI_HANDOFF tasks instead of continuing in this one thread — smaller phases keep each turn under the cap.');
181
214
  lines_out.push('3) Otherwise, tell the user plainly: run /compact now, or wrap up and start a fresh session — work already committed/persisted to disk is not lost.');
182
215
  if (sidechainEntries > 0) {
183
216
  lines_out.push(`Note: ${sidechainEntries} subagent entries in this stretch — each teammate carries its own context window, and every finished report is injected back here, so running many at once is the fastest way to overflow this session. Avoid spawning more until context drops back under the cap.`);
@@ -14,6 +14,18 @@
14
14
  INPUT=$(cat)
15
15
  PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
16
16
 
17
+ # Fast path — skip the node spawn (~70ms) when this hook provably has nothing to say.
18
+ # The two branches below only ever act on (a) Edit/Write under docs/AI_HANDOFF/, or
19
+ # (b) a Bash `git push`. If neither marker appears anywhere in the raw payload, the node
20
+ # program is guaranteed to exit 0, so running it is pure latency — paid on EVERY Bash,
21
+ # Edit and Write in every session and every parallel subagent, which is where a
22
+ # many-agent pipeline quietly loses minutes.
23
+ # Conservative by construction: a payload that merely mentions these strings still falls
24
+ # through to the real check below. This can only skip work, never skip a block.
25
+ if ! printf '%s' "$INPUT" | grep -qE 'AI_HANDOFF|git[[:space:]]+push'; then
26
+ exit 0
27
+ fi
28
+
17
29
  INPUT="$INPUT" PROJECT_ROOT="$PROJECT_ROOT" node <<'NODE'
18
30
  const fs = require('fs');
19
31
  const path = require('path');
@@ -0,0 +1,90 @@
1
+ #!/bin/bash
2
+ # SessionStart hook: resume an interrupted handoff-fullstack run.
3
+ #
4
+ # Why this exists:
5
+ # A one-shot pipeline is only one-shot if an interruption is recoverable without the human
6
+ # retyping anything. Compaction, a crash, or closing the terminal all end the session while
7
+ # the cycle is still mid-flight. docs/AI_HANDOFF/RUN.md is the cursor the command writes
8
+ # after every step; this hook reads it back at session start and injects the continuation
9
+ # instruction, so the very next turn picks up where the run stopped.
10
+ #
11
+ # Pairs with context-window-guard.sh, which stops telling the agent to halt for /compact
12
+ # while a run is active. Between them: never stop mid-cycle, and if the session ends anyway,
13
+ # resume automatically.
14
+ #
15
+ # Fires on every SessionStart source (startup / resume / clear / compact). `compact` is the
16
+ # important one — that is the path the command's own Final Report tells the user to take
17
+ # between cycles, and the path an over-full session takes mid-cycle.
18
+ #
19
+ # ADVISORY ONLY — always exit 0. A resume hint must never be able to block a session from
20
+ # starting.
21
+
22
+ INPUT=$(cat)
23
+ PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
24
+
25
+ INPUT="$INPUT" PROJECT_ROOT="$PROJECT_ROOT" node <<'NODE' || true
26
+ const fs = require('fs');
27
+ const path = require('path');
28
+
29
+ const payload = (() => {
30
+ try {
31
+ const parsed = JSON.parse(process.env.INPUT || '');
32
+ return parsed && typeof parsed === 'object' ? parsed : {};
33
+ } catch {
34
+ return {};
35
+ }
36
+ })();
37
+
38
+ const projectRoot = process.env.PROJECT_ROOT;
39
+ const runPath = path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md');
40
+
41
+ let text;
42
+ try {
43
+ text = fs.readFileSync(runPath, 'utf8');
44
+ } catch {
45
+ // No cursor file: nothing was running, or the last cycle was cleared. Stay silent —
46
+ // a resume banner on every ordinary session start would be pure noise.
47
+ process.exit(0);
48
+ }
49
+
50
+ const field = (name) => (text.match(new RegExp(`^${name}:\\s*(.+)$`, 'm'))?.[1] || '').trim();
51
+
52
+ const phase = field('Phase');
53
+ // `done` means the last cycle finished cleanly. Anything else means a step was in flight.
54
+ if (!phase || /^done$/i.test(phase)) process.exit(0);
55
+
56
+ const goal = field('Goal');
57
+ const base = field('Base');
58
+ const cursor = field('Cursor');
59
+ const next = field('Next');
60
+ const source = typeof payload.source === 'string' ? payload.source : 'startup';
61
+
62
+ const out = [
63
+ 'UKIT HANDOFF RESUME — an unfinished handoff-fullstack run was found on disk.',
64
+ ` Goal: ${goal || '(not recorded)'}`,
65
+ ` Base: ${base || '(not recorded)'}`,
66
+ ` Phase: ${phase}`,
67
+ ` Cursor: ${cursor || '(not recorded)'}`,
68
+ ` Next: ${next || '(not recorded)'}`,
69
+ '',
70
+ 'This is a CONTINUATION, not a new cycle. Do not re-plan, do not overwrite PLAN.md or the',
71
+ 'task files, and do not ask the user whether to continue — they already asked for a',
72
+ 'one-shot run and the pipeline was interrupted, not cancelled.',
73
+ '',
74
+ 'Read docs/AI_HANDOFF/RUN.md and docs/AI_HANDOFF/INDEX.md, then execute the `Next:` step',
75
+ 'above using .claude/commands/ukit/handoff-fullstack.md as the procedure. Skip tasks',
76
+ 'already `done` or `pending_review`; pick up `ready`, `in_progress`, `changes_requested`',
77
+ 'and `blocked` ones. Keep writing the cursor after every step.',
78
+ ];
79
+
80
+ if (source === 'compact') {
81
+ out.push('');
82
+ out.push('The session was just compacted, so the window is clean — this is the ideal moment');
83
+ out.push('to continue. Resume immediately rather than reporting status and waiting.');
84
+ }
85
+
86
+ process.stdout.write(`${out.join('\n')}\n`);
87
+ process.exit(0);
88
+ NODE
89
+
90
+ exit 0
@@ -26,10 +26,13 @@
26
26
  "Bash(git status:*)",
27
27
  "Bash(git diff:*)",
28
28
  "Bash(git log:*)",
29
+ "Bash(git add:*)",
30
+ "Bash(git commit:*)",
31
+ "Bash(git push:*)",
32
+ "Bash(git worktree:*)",
29
33
  "Bash(jq:*)"
30
34
  ],
31
35
  "ask": [
32
- "Bash(git push:*)",
33
36
  "Bash(gh pr create:*)",
34
37
  "Bash(gh pr comment:*)",
35
38
  "Bash(gh issue comment:*)"
@@ -197,6 +200,11 @@
197
200
  "type": "command",
198
201
  "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/reset-compact-pressure.sh\"",
199
202
  "timeout": 8
203
+ },
204
+ {
205
+ "type": "command",
206
+ "command": "\"$CLAUDE_PROJECT_DIR/.claude/hooks/handoff-resume.sh\"",
207
+ "timeout": 8
200
208
  }
201
209
  ]
202
210
  }
@@ -38,6 +38,53 @@ Khi cần hỏi-lại / push-back / gợi ý cho phase khác, AI ghi vào `## Di
38
38
 
39
39
  Phase kế tiếp PHẢI đọc Discussion trước khi tiếp tục — coi như inbox.
40
40
 
41
+ ## Autonomy Model — hỏi một lần, chạy tới hết
42
+
43
+ Handoff được thiết kế để chạy **không cần người ngồi canh**. Config: `.ukit/storage/config.json` → `handoff.autonomy`.
44
+
45
+ **Cửa sổ hỏi duy nhất là lúc plan.** `/ukit:handoff-create` (và bước P0 của `/ukit:handoff-fullstack`) gom mọi câu hỏi vào **một** lần `AskUserQuestion` rồi đóng cửa sổ. Câu trả lời ghi nguyên văn vào `PLAN.md §1` — các phase sau coi đó là lời của người dùng và không hỏi lại.
46
+
47
+ Từ sau đó, mọi ngã ba đều có cách giải quyết tự động:
48
+
49
+ | Tình huống | Xử lý tự động |
50
+ |-----------|---------------|
51
+ | Scope nhiều subsystem | tách module, plan module 1, queue phần còn lại vào `INDEX.md` |
52
+ | Plan review còn `Issues Found` sau 2 vòng | planner áp findings rồi đi tiếp, ghi vào Plan Review Log |
53
+ | 2 task cùng wave đụng 1 file | đẩy task sau xuống wave kế (thêm dependency) |
54
+ | Working tree bẩn | commit checkpoint rồi chạy tiếp |
55
+ | Reviewer trả `changes_requested`/`critical_block` | vào vòng auto-fix, tối đa 2 vòng |
56
+ | Task vẫn fail sau 2 vòng fix | để `blocked`, **các task khác vẫn đi tiếp tới push** |
57
+
58
+ Chỉ escalate cho người khi: blocker nằm ngoài repo (thiếu credential, service chết), hoặc task còn fail sau cả 2 vòng fix. Kể cả vậy vẫn phải làm xong mọi task khác trước rồi mới báo.
59
+
60
+ **Quality gate không bị nới.** Bỏ chỗ *hỏi người*, không bỏ chỗ *kiểm tra*: TDD RED→GREEN vẫn bắt buộc, reviewer vẫn phải khác model executor, vẫn re-run Verification Commands, vẫn cấm claim DONE khi chưa có PASS tươi.
61
+
62
+ ### Run cursor — `RUN.md`
63
+
64
+ Mỗi command ghi lại `docs/AI_HANDOFF/RUN.md` sau **mỗi bước**:
65
+
66
+ ```
67
+ Command: <handoff-fullstack|handoff-implement|handoff-review>
68
+ Goal: <1 câu>
69
+ Base: <branch>
70
+ Phase: <phase hiện tại | done>
71
+ Cursor: wave <N> batch <M> — <vừa xong cái gì>
72
+ Next: <bước kế tiếp chính xác>
73
+ ```
74
+
75
+ `Phase:` khác `done` = có run đang dở. Command được gọi lại sẽ **chạy tiếp từ cursor**, không plan lại, không hỏi. Hook `SessionStart` (`handoff-resume.sh`) đọc file này và tự inject lệnh chạy tiếp — nên compact hay mất session giữa chừng đều không làm mất run.
76
+
77
+ ### Context: không bao giờ dừng giữa nhiệm vụ
78
+
79
+ - Subagent ghi **full log vào task file trên đĩa**, chỉ trả về orchestrator ≤10 dòng (executor) / ≤6 dòng (reviewer). Paste log ngược lại orchestrator là nguyên nhân số 1 làm run chết vì hết context.
80
+ - Hết mỗi wave: commit, ghi cursor, **collapse** wave đó còn 1 dòng/task trong bộ nhớ làm việc, rồi chạy tiếp.
81
+ - Yêu cầu `/compact` **chỉ** được đặt ở cuối command, giữa 2 cycle. Giữa cycle thì tuyệt đối không — state đã nằm hết ở git + `INDEX.md` + `RUN.md` nên compact ở ranh giới cycle không mất gì.
82
+
83
+ ### Git
84
+
85
+ - `handoff-implement`: **1 commit / wave**, không push. Mỗi wave revert độc lập được.
86
+ - `handoff-review` / `handoff-fullstack`: commit thêm mỗi vòng auto-fix, rồi **push 1 lần** ở cuối. Push không hỏi (`Bash(git push:*)` nằm trong `allow`); force-push vẫn bị `deny`.
87
+
41
88
  ## Handoff Flow (tool-agnostic, file-based state machine)
42
89
 
43
90
  UKit handoff hoạt động qua **file state**. Anh tự chọn tool nào cho từng phase — Claude Code / Kilo Code / Codex / OpenCode / tool mới sau này — đều được. UKit chỉ care về **role của model**, không care tool.
@@ -72,23 +119,24 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
72
119
  **Phase 1 — Idea + Plan** (smart/reasoning model)
73
120
  - Human submit ideas (natural language).
74
121
  - AI ghi vào `PLAN.md`: §1 Intent, §2 Scope, §3 Approach, **§4 Test Plan (bắt buộc TDD-style)**, §5 Verification Commands, §6 Acceptance Criteria.
75
- - Output: PLAN.md đầy đủ, chờ human approve.
122
+ - Đây là **cửa sổ hỏi duy nhất** của cả pipeline: gom mọi câu hỏi vào 1 lần `AskUserQuestion` trước khi viết, ghi câu trả lời vào §1.
123
+ - Output: PLAN.md đầy đủ + Planner Self-Audit. Chạy standalone (`/ukit:handoff-create`) thì dừng ở đây chờ human xem; chạy trong `/ukit:handoff-fullstack` thì đi thẳng tiếp sang Phase 2 — plan review độc lập là gate thay cho human.
76
124
 
77
125
  **Phase 2 — Create Tasks (TDD-embedded, MANDATORY)** (smart/reasoning model, thường cùng phase 1)
78
126
  - Human approve plan → AI split `PLAN.md §7` sang nhiều `tasks/TASK-xxx.md`.
79
127
  - **Mỗi TASK file BẮT BUỘC có Test Plan của riêng nó**, không chỉ trỏ về PLAN.md. Cụ thể:
80
- - `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥1 edge case + regression (nếu fix bug).
128
+ - `§ Test Cases`: bảng test (loại, tên test, expected) cho phần task này — happy + ≥2 edge case KHÁC loại nhau (vd null/empty + boundary/concurrent, không tính 2 case gần giống nhau) + regression (nếu fix bug).
81
129
  - `§ Test Files`: đường dẫn cụ thể file test sẽ tạo/sửa (ví dụ `tests/auth/login.test.js`).
82
- - `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS.
130
+ - `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS. Nếu project có sẵn lint/typecheck script → BẮT BUỘC liệt kê ở đây, không chỉ lệnh test. Project không có thì ghi rõ N/A, không được bỏ qua im lặng.
83
131
  - `§ Acceptance Criteria`: checklist.
84
132
  - Nếu split mà task nào không kèm được Test Cases + Test Files cụ thể → task đó chưa đủ `ready`, đánh `needs_breakdown`.
85
133
  - Update `INDEX.md`: thêm row mỗi task với status `ready`.
86
- - Đây là **điểm cắt human-approval**: phase này xong, executor được phép pick.
134
+ - Đây là **điểm cắt cuối trước khi code chạy**: phase này xong, executor được phép pick. Trong `/ukit:handoff-fullstack`, gate ở đây là plan review độc lập (model mạnh, context riêng) chứ không phải human — vì người dùng đã chủ động chọn chạy one-shot.
87
135
  - Mục tiêu: executor (cheap-smart model) đọc task file là biết NGAY test gì cần viết trước, KHÔNG phải tự suy diễn.
88
136
 
89
137
  **Phase 3 — Implement + Test** (cheap-smart/code model)
90
138
  - User: "execute next task" / "làm TASK-001" / "implement task 1".
91
- - Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT` + verification output) vào cuối task file → đổi status `pending_review`.
139
+ - Executor đọc `INDEX.md` → pick `ready` task → đổi `in_progress` → **viết test trước → RED (paste output failing thật, không chỉ khai đã confirm) → implement → GREEN** → chạy Verification Commands fresh trong turn → append `## Executor Report` (gồm `EXECUTOR_TOOL`/`EXECUTOR_MODEL`/`EXECUTOR_SUBAGENT`/`RED_OUTPUT` + verification output) vào cuối task file → đổi status `pending_review`.
92
140
  - KHÔNG được claim DONE nếu chưa có PASS fresh.
93
141
 
94
142
  **Phase 4 — Review + Test** (reviewer model — KHÁC model executor)
@@ -108,8 +156,8 @@ A task is `ready` only when it has:
108
156
  - Dependencies stated
109
157
  - **Interfaces** — Consumes/Produces với chữ ký thật (function/endpoint/type), không placeholder;
110
158
  `(none)` hợp lệ nếu task không có input/output liên task
111
- - **Test Plan** (PLAN.md §4) — happy path + ≥1 edge case (+ regression test nếu fix bug); hoặc `N/A` kèm lý do
112
- - Verification command (lệnh executor sẽ chạy)
159
+ - **Test Plan** (PLAN.md §4) — happy path + ≥2 edge case khác loại (+ regression test nếu fix bug); hoặc `N/A` kèm lý do
160
+ - Verification command (lệnh executor sẽ chạy) — PHẢI gồm lint/typecheck nếu project có sẵn
113
161
  - Acceptance criteria
114
162
 
115
163
  Missing any → `needs_breakdown`, `blocked`, or `needs_human`.
@@ -186,11 +186,24 @@
186
186
  "handoff": {
187
187
  "enabled": true,
188
188
  "crossTool": true,
189
- "maxParallelAgents": 3,
189
+ "maxParallelAgents": 10,
190
+ "autonomy": {
191
+ "askWindow": "plan-only",
192
+ "planReviewRounds": 2,
193
+ "autoFixRounds": 2,
194
+ "autoSerializeFileConflicts": true,
195
+ "resumeFromCursor": true,
196
+ "runCursorFile": "docs/AI_HANDOFF/RUN.md",
197
+ "commitPerWave": true,
198
+ "pushOnComplete": true,
199
+ "compactAtCycleBoundaryOnly": true,
200
+ "agentReturnMaxLines": 10
201
+ },
190
202
  "plan": {
191
203
  "requireTestPlan": true,
192
204
  "minTestsHappyPath": 1,
193
- "minTestsEdgeCase": 1,
205
+ "minTestsEdgeCase": 2,
206
+ "requireLintOrTypecheckInVerification": true,
194
207
  "regressionTestRequiredForBugfix": true,
195
208
  "smartModelHint": "claude-opus-5"
196
209
  },
@@ -479,11 +492,12 @@
479
492
  "handoff": {
480
493
  "enabled": "Bật Quality Gate cho handoff: plan có Test Plan, executor test-first, reviewer model khác. Tắt = quay về flow cũ (dễ lọt lỗi vặt).",
481
494
  "crossTool": "true nghĩa là handoff truyền qua file (PLAN/INDEX/tasks) chứ không qua in-process subagent — cho phép plan ở Claude Code, execute ở Kilo Code, review ở Claude Code khác model.",
482
- "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 3). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt. Lý do: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc là cách nhanh nhất làm session chính vượt context window. Hạ xuống 2 nếu task nặng; không nên vượt 3.",
495
+ "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 10, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
483
496
  "plan": {
484
497
  "requireTestPlan": "Bắt buộc PLAN.md §4 phải có Test Plan trước khi task chuyển ready.",
485
498
  "minTestsHappyPath": "Tối thiểu test cho happy path.",
486
- "minTestsEdgeCase": "Tối thiểu test cho edge case (null/empty/boundary/concurrent…).",
499
+ "minTestsEdgeCase": "Tối thiểu test cho edge case (mặc định 2, tăng từ 1 — phải khác loại nhau, vd null/empty + boundary/concurrent, không tính 2 test gần giống nhau là đủ).",
500
+ "requireLintOrTypecheckInVerification": "Nếu project có sẵn lint/typecheck script (package.json hoặc tương đương), Verification Commands của mỗi task BẮT BUỘC phải gồm lệnh đó — không chỉ chạy test. Project không có lint/typecheck thì planner phải ghi rõ lý do N/A thay vì bỏ qua im lặng.",
487
501
  "regressionTestRequiredForBugfix": "Bug fix phải có regression test fail-trước-fix.",
488
502
  "smartModelHint": "Gợi ý model mạnh nhất cho phase plan (ví dụ claude-opus-5). UKit không tự ép, chỉ ghi hint vào task."
489
503
  },