@ngockhoale/ukit 2.5.1 → 2.6.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (53) hide show
  1. package/CHANGELOG.md +80 -0
  2. package/manifests/platform.full.yaml +16 -0
  3. package/package.json +1 -1
  4. package/src/cli/commands/install.js +49 -2
  5. package/src/cli/commands/update.js +5 -0
  6. package/src/core/output/index.js +77 -0
  7. package/src/core/update.js +36 -5
  8. package/templates/.claude/agents/code-reviewer.md +12 -1
  9. package/templates/.claude/agents/feature-implementer.md +4 -2
  10. package/templates/.claude/agents/handoff-planner.md +36 -1
  11. package/templates/.claude/commands/ukit/handoff-clear.md +4 -0
  12. package/templates/.claude/commands/ukit/handoff-create.md +23 -7
  13. package/templates/.claude/commands/ukit/handoff-fullstack.md +181 -17
  14. package/templates/.claude/commands/ukit/handoff-implement.md +9 -2
  15. package/templates/.claude/commands/ukit/handoff-review.md +4 -1
  16. package/templates/.claude/commands/ukit/handoff-status.md +6 -2
  17. package/templates/.claude/hooks/auto-allow-bash.sh +17 -5
  18. package/templates/.claude/hooks/block-dangerous.sh +18 -5
  19. package/templates/.claude/hooks/completion-gate.sh +17 -5
  20. package/templates/.claude/hooks/compress-output.sh +15 -6
  21. package/templates/.claude/hooks/context-hardcap-gate.sh +77 -46
  22. package/templates/.claude/hooks/context-window-guard.sh +41 -27
  23. package/templates/.claude/hooks/handoff-model-guard.sh +62 -14
  24. package/templates/.claude/hooks/handoff-resume.sh +20 -7
  25. package/templates/.claude/hooks/post-edit-verify.sh +17 -5
  26. package/templates/.claude/hooks/pre-edit-backup.sh +17 -5
  27. package/templates/.claude/hooks/project-important.sh +18 -1
  28. package/templates/.claude/hooks/protect-files.sh +18 -5
  29. package/templates/.claude/hooks/record-execution.sh +17 -5
  30. package/templates/.claude/hooks/sensitive-data-guard.sh +48 -9
  31. package/templates/.claude/hooks/skill-router.sh +124 -86
  32. package/templates/.claude/hooks/stale-spec-guard.sh +22 -6
  33. package/templates/.claude/hooks/task-watchdog.sh +22 -7
  34. package/templates/.claude/hooks/verification-guard.sh +17 -5
  35. package/templates/.claude/hooks/vision-router.sh +17 -5
  36. package/templates/.claude/settings.json +15 -10
  37. package/templates/.claude/ukit/index/provision-worktree.mjs +30 -2
  38. package/templates/.claude/ukit/runtime/execution-ledger.mjs +99 -1
  39. package/templates/.claude/ukit/runtime/hook-input.sh +25 -0
  40. package/templates/.claude/ukit/runtime/hook-telemetry.mjs +86 -2
  41. package/templates/.claude/ukit/runtime/output-compression.mjs +87 -0
  42. package/templates/.claude/ukit/runtime/stop-coordinator.mjs +201 -11
  43. package/templates/.omp/RULES.md +9 -1
  44. package/templates/.omp/agents/code-reviewer.md +12 -1
  45. package/templates/.omp/agents/feature-implementer.md +4 -2
  46. package/templates/.omp/agents/handoff-planner.md +36 -1
  47. package/templates/.omp/hooks/pre/ukit-bridge.js +110 -3
  48. package/templates/AGENTS.md +14 -0
  49. package/templates/CLAUDE.md +14 -0
  50. package/templates/docs/AI_HANDOFF/RULES.md +37 -4
  51. package/templates/docs/AI_HANDOFF/SPEC.md +98 -0
  52. package/templates/docs/AI_HANDOFF/tasks/_TEMPLATE.md +4 -1
  53. package/templates/ukit/storage/config.json +48 -9
package/CHANGELOG.md CHANGED
@@ -2,6 +2,86 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.6.2 - 2026-09-18
6
+
7
+ Stability release — 17 confirmed hang / silent-idle / silent-hook defects fixed with
8
+ root causes (full register: `docs/AI_HANDOFF/BUGS-2.6.2.md`), focused on the "system
9
+ goes quiet and does nothing" failure class.
10
+
11
+ - **No more silent degraded input (§8 contract, BUG-C21-01/02/03)**: every hook that
12
+ stages stdin now *announces* when staging fails, stalls, or refuses — fail-closed
13
+ hooks (`block-dangerous`, `protect-files`, `context-hardcap-gate`,
14
+ `sensitive-data-guard`) emit an `ask`/`systemMessage` JSON + stderr `BLOCKED` +
15
+ `exit 2`; advisory hooks emit a `systemMessage` + `exit 0`. The old bounded-cut
16
+ silently truncated and proceeded. `sensitive-data-guard` now also fails closed
17
+ (exit 2) on any non-refusal staging failure and mirrors its deadline block to
18
+ stderr so the reason survives `process.exit` truncation.
19
+ - **Sync-fs defeats inside self-deadlines removed (BUG-C21-04)**: the
20
+ deadline-guarded node blocks in `context-window-guard.sh` and
21
+ `context-hardcap-gate.sh` no longer call synchronous `fs.*Sync` inside the kill
22
+ window — a stalled `.ukit` mount can no longer freeze the gate so the unref'd
23
+ deadline can never fire. (Residual call-frame past the grant boundary recorded
24
+ as BUG-C21-R1/R2 for the next cycle.)
25
+ - **Host timeouts now cover internal budgets (BUG-C21-05)**: every hook registered
26
+ in `settings.json` gets `timeout ≥ stage bound + node deadline + margin`; the
27
+ regression test iterates *all* registered hooks instead of a hand-picked list,
28
+ and four under-budgeted hooks (`auto-allow-bash`, `vision-router`,
29
+ `auto-prune-bash`, `reset-compact-pressure`) were raised to 12s.
30
+ - **Missing `await` in skill-router (review major)**: an un-awaited async context
31
+ snapshot could leave the router writing a stale/partial decision silently.
32
+ - **CLI hangs eliminated (BUG-C21-12/13)**: `ukit update`/`install` `spawnSync`
33
+ calls now carry timeouts with `killSignal: 'SIGKILL'` (SIGTERM alone keeps
34
+ waiting on signal-resistant children); a wedged `npm root -g` now reports
35
+ `global-root-query-timeout` instead of hanging forever. `promptYesNo` resolves
36
+ on stdin EOF / `ERR_USE_AFTER_CLOSE` instead of hanging or crashing.
37
+ - **Unbounded growth bounded (BUG-C21-07/09/10/11)**: `tee/` cache, hook-errors
38
+ jsonl, hook-latency and exec-ledger session files now have count caps; the
39
+ ledger protects receipt files (`.journal`, `.quarantine`, gate crash counter)
40
+ from count-cap eviction, and the overflow basis uses `entries.length`
41
+ (post-filter) everywhere.
42
+ - **Handoff machinery (BUG-C21-15)**: all four `RUN.md` `Phase:` consumers treat
43
+ `done`/`blocked` as terminal — dead runs no longer nag forever.
44
+ - **Wiring/parity (BUG-C21-06/08/14/16/17)**: stale-spec-guard deadline exits
45
+ distinguish "stall" from "match"; `provision-worktree` spawnSync calls carry
46
+ timeouts; `task-budget-validator.mjs` is manifest-covered so the documented
47
+ command works on fresh installs; `verification-guard.sh` is actually registered
48
+ on PreToolUse Bash; shipped `config.json` regained `codex-context-budget` and a
49
+ parity test now guards every array field under `subagents`.
50
+
51
+
52
+ ## 2.6.0 - 2026-09-18
53
+
54
+ handoff-fullstack v2 — autonomous create→implement→review loop that runs to 100%
55
+ completion of all unfinished work, with a mandatory detailed SPEC gate.
56
+
57
+ - **Never-stall Stop gate**: `stop-coordinator.mjs` gained a `handoff-cursor` lane that
58
+ reads `docs/AI_HANDOFF/RUN.md` and blocks stop while `Phase:` is not `done`/`blocked`,
59
+ carrying the cursor's `Next:` step. Liveness breaker releases after
60
+ `handoff.fullstack.stopGateMaxStalledBlocks` (default 12) un-advancing blocks; the lane
61
+ is advisory-on-failure and `stopGateEnabled` disables it. This fixes the observed
62
+ stall where runs went idle after a recap.
63
+ - **Mandatory SPEC.md**: `handoff-create` now writes `docs/AI_HANDOFF/SPEC.md` — a
64
+ 15-section template (FRs with Given/When/Then, fullstack scope, data model, API
65
+ contract, test matrix, acceptance criteria, migration, chosen defaults) — before task
66
+ files. Tasks carry `Spec references`; `handoff-model-guard.sh` hard-blocks fresh task
67
+ creation without a filled SPEC (`handoff.fullstack.specRequired=false` bypasses).
68
+ - **Autonomous loop contract**: only `HANDOFF FULLSTACK COMPLETE` / `HANDOFF FULLSTACK
69
+ BLOCKED` may end a run; `CHECKPOINT — WORK CONTINUING` replaces bare recaps. Phase 0
70
+ sweeps INDEX, task files, HISTORY/archive, `docs/TASKS.md` Ready-for-AI, and
71
+ uncommitted WIP. Stuck tasks recover via `cancelled_superseded` + `TASK-xxx-R<n>`
72
+ replacement tasks. `idleWatchdogMin` (default 5) scheduled wakeups keep the loop
73
+ alive off-Claude-Code; `quietScansRequired` (default 2) guards completion; Phase F
74
+ runs docs sync → `archive/cycle-NN/` → `Phase: done` → marker report.
75
+ - **Agent + command wiring**: `handoff-planner` gained legacy-sweep and SPEC-authoring
76
+ phases; `code-reviewer` reviews SPEC+PLAN against a 7-point spec quality gate;
77
+ `feature-implementer` treats `Spec references` as outranking task-file wording;
78
+ `handoff-implement`/`handoff-review` codify spec-is-the-contract and `-R<n>`
79
+ recovery. `handoff-clear` archives SPEC and must close RUN.md; `handoff-status`
80
+ reports run phase + spec presence.
81
+ - **Config**: `handoff.fullstack.*` keys (`stopGateEnabled`, `stopGateMaxStalledBlocks`,
82
+ `quietScansRequired`, `idleWatchdogMin`, `autoArchive`, `specRequired`) with `_help`
83
+ notes, in both config trees.
84
+
5
85
  ## 2.5.1 - 2026-09-18
6
86
 
7
87
  Post-release fixes from C20 reviewer advisories.
@@ -1548,6 +1548,22 @@ items:
1548
1548
  packs:
1549
1549
  - core
1550
1550
 
1551
+ # BUG-C21-14: this file shipped in templates/ but was never manifest-listed, so
1552
+ # autoDiscoverTemplates:false silently skipped it and the agent-documented
1553
+ # `node .claude/ukit/index/task-budget-validator.mjs` command 404'd on installs.
1554
+ # tests/manifest/settingsHookCoverage.test.js now asserts every file under
1555
+ # templates/.claude/ukit/index/ is covered, which immunizes the next added module.
1556
+ - id: ukit-index-task-budget-validator-script
1557
+ type: config
1558
+ sourceTemplate: .claude/ukit/index/task-budget-validator.mjs
1559
+ targetPath: .claude/ukit/index/task-budget-validator.mjs
1560
+ requires: []
1561
+ mergeStrategy: overwrite_with_backup
1562
+ variables: []
1563
+ enabledByDefault: true
1564
+ packs:
1565
+ - core
1566
+
1551
1567
  - id: gitignore-root
1552
1568
  type: config
1553
1569
  sourceTemplate: .gitignore
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.5.1",
3
+ "version": "2.6.2",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -81,9 +81,46 @@ async function loadTrackedManagedPathSet(projectRoot) {
81
81
  );
82
82
  }
83
83
 
84
+ // BUG-C21-13: `readline/promises` `question()` does not settle when the input
85
+ // stream hits EOF without a pending newline — the promise stays pending after
86
+ // the interface `close` event, so an unattended Ctrl+D / piped stdin froze
87
+ // `runInstall` forever. Race every question against the interface's `close`
88
+ // event; on EOF treat it as the default answer (decline → retain) and stop.
89
+ // A post-EOF `question()` also throws ERR_USE_AFTER_CLOSE — same treatment.
90
+ function onceInterfaceClosed(line) {
91
+ return new Promise((resolve) => {
92
+ if (typeof line.once === 'function') {
93
+ line.once('close', resolve);
94
+ }
95
+ });
96
+ }
97
+
98
+ async function askWithEofGuard(line, question) {
99
+ if (line.closed === true) {
100
+ return null;
101
+ }
102
+ try {
103
+ const answer = await Promise.race([line.question(question), onceInterfaceClosed(line)]);
104
+ return typeof answer === 'string' ? answer : null;
105
+ } catch (error) {
106
+ if (error?.code === 'ERR_USE_AFTER_CLOSE') {
107
+ return null;
108
+ }
109
+ throw error;
110
+ }
111
+ }
112
+
84
113
  async function promptYesNo(line, question) {
85
114
  while (true) {
86
- const answer = (await line.question(question)).trim().toLowerCase();
115
+ const raw = await askWithEofGuard(line, question);
116
+
117
+ if (raw === null) {
118
+ // EOF / interface closed: default answer (No) — never re-prompt a dead stream.
119
+ console.log('[UKit] Input ended (EOF) — keeping existing files.');
120
+ return false;
121
+ }
122
+
123
+ const answer = raw.trim().toLowerCase();
87
124
 
88
125
  if (answer === '') {
89
126
  return false;
@@ -97,6 +134,12 @@ async function promptYesNo(line, question) {
97
134
  return false;
98
135
  }
99
136
 
137
+ // A real terminal stays open after a bad answer, so re-prompting is safe —
138
+ // but if the stream just closed mid-loop, bail out as EOF.
139
+ if (line.closed === true) {
140
+ console.log('[UKit] Input ended (EOF) — keeping existing files.');
141
+ return false;
142
+ }
100
143
  console.log('[UKit] Please answer y/yes or n/no.');
101
144
  }
102
145
  }
@@ -191,7 +234,11 @@ export async function pruneDeselectedAdapters({
191
234
  console.log(`[UKit] Removed ${removedCount} ${adapter.label} path(s).`);
192
235
  }
193
236
  } finally {
194
- line.close();
237
+ try {
238
+ line.close();
239
+ } catch {
240
+ // close() is best-effort cleanup — never let it mask the real result.
241
+ }
195
242
  }
196
243
 
197
244
  return { retainedManagedPaths: [...new Set(retainedManagedPaths)] };
@@ -44,6 +44,11 @@ export async function runUpdate({ packageVersion, argv = [], cwd = process.cwd()
44
44
  console.log('[UKit] Run `ukit install` inside a project to set it up.');
45
45
  return;
46
46
  }
47
+ if (reason === 'global-root-query-timeout') {
48
+ console.log('[UKit] Update complete, but locating the upgraded CLI timed out.');
49
+ console.log('[UKit] `npm root -g` did not answer within 15s — check for a wedged npm/network, then run `ukit install` to refresh this project.');
50
+ return;
51
+ }
47
52
  if (reason === 'global-bin-not-found') {
48
53
  console.log('[UKit] Update complete, but the upgraded CLI could not be located.');
49
54
  console.log('[UKit] Run `ukit install` to refresh this project.');
@@ -353,6 +353,81 @@ function buildRecoveryHintSummary(summary, rawPath, { tokensBefore = 0 } = {}) {
353
353
  return null;
354
354
  }
355
355
 
356
+ // --- tee/ cache bound (BUG-C21-07, FR-006) -----------------------------------
357
+ // Mirror of the bound in templates/.claude/ukit/runtime/output-compression.mjs —
358
+ // keep the two implementations identical. persistRawOutput wrote one preserved
359
+ // file per compression and nothing ever deleted from tee/. Sweeps are SAMPLED
360
+ // (~1/16 writes, UKIT_TEE_SWEEP_PROBABILITY to force/disable) and BOUNDED
361
+ // (<= maxEntries stats, <= maxRemovals unlinks per sweep).
362
+ const TEE_SWEEP_PROBABILITY_DEFAULT = 1 / 16;
363
+ const TEE_SWEEP_MAX_ENTRIES = 128;
364
+ const TEE_SWEEP_MAX_REMOVALS = 64;
365
+ const TEE_MAX_FILES = 500;
366
+ const TEE_MAX_AGE_MS = 7 * 24 * 60 * 60 * 1000;
367
+
368
+ function teeSweepProbabilityFromEnv() {
369
+ const raw = Number(process.env.UKIT_TEE_SWEEP_PROBABILITY);
370
+ if (!Number.isFinite(raw)) return TEE_SWEEP_PROBABILITY_DEFAULT;
371
+ return Math.min(1, Math.max(0, raw));
372
+ }
373
+
374
+ // See the runtime mirror for the contract: removes entries older than maxAgeMs,
375
+ // then the oldest scanned entries while the dir count exceeds maxFiles — all
376
+ // bounded by maxEntries/maxRemovals so an oversized dir is amortized down.
377
+ export async function sweepTeeCache(dir, {
378
+ now = Date.now,
379
+ maxAgeMs = TEE_MAX_AGE_MS,
380
+ maxFiles = TEE_MAX_FILES,
381
+ maxEntries = TEE_SWEEP_MAX_ENTRIES,
382
+ maxRemovals = TEE_SWEEP_MAX_REMOVALS,
383
+ } = {}) {
384
+ let names;
385
+ try {
386
+ names = await fs.readdir(dir);
387
+ } catch {
388
+ return { sampled: true, scanned: 0, removed: 0 };
389
+ }
390
+ const cutoff = now() - maxAgeMs;
391
+ const entries = [];
392
+ let scanned = 0;
393
+ for (const name of names) {
394
+ if (scanned >= maxEntries) break;
395
+ scanned += 1;
396
+ try {
397
+ const stats = await fs.stat(path.join(dir, name));
398
+ if (stats.isFile()) entries.push({ name, mtimeMs: stats.mtimeMs });
399
+ } catch { /* raced away — fine */ }
400
+ }
401
+ entries.sort((a, b) => a.mtimeMs - b.mtimeMs);
402
+ // Overflow counts eligible FILES only — foreign entries (non-matching names,
403
+ // dirs) inflate `names.length` and would evict real entries while the dir is
404
+ // actually under the cap.
405
+ const overflow = Math.max(0, entries.length - maxFiles);
406
+ let removed = 0;
407
+ for (const entry of entries) {
408
+ if (removed >= maxRemovals) break;
409
+ if (removed < overflow || entry.mtimeMs < cutoff) {
410
+ try {
411
+ await fs.rm(path.join(dir, entry.name), { force: true });
412
+ removed += 1;
413
+ } catch { /* raced away — fine */ }
414
+ }
415
+ }
416
+ return { sampled: true, scanned, removed };
417
+ }
418
+
419
+ export async function maybeSweepTeeCache(dir, {
420
+ probability,
421
+ random = Math.random,
422
+ ...options
423
+ } = {}) {
424
+ const p = Number.isFinite(probability)
425
+ ? Math.min(1, Math.max(0, probability))
426
+ : teeSweepProbabilityFromEnv();
427
+ if (random() >= p) return { sampled: false, scanned: 0, removed: 0 };
428
+ return sweepTeeCache(dir, options);
429
+ }
430
+
356
431
  async function persistRawOutput(projectRoot, {
357
432
  command = '',
358
433
  summary = '',
@@ -379,6 +454,8 @@ async function persistRawOutput(projectRoot, {
379
454
 
380
455
  await fs.mkdir(teeCacheDir, { recursive: true });
381
456
  await fs.writeFile(absolutePath, rawOutputText, 'utf8');
457
+ // Bounded, sampled tee/ prune (BUG-C21-07): advisory only.
458
+ await maybeSweepTeeCache(teeCacheDir).catch(() => {});
382
459
 
383
460
  return {
384
461
  rawSaved: true,
@@ -10,6 +10,22 @@ export const UKIT_PACKAGE_NAME = '@ngockhoale/ukit';
10
10
  // CLAUDE.md and edits .gitignore — an unwanted surprise in, say, a home directory).
11
11
  const INSTALLED_MARKERS = ['.ukit', '.claude', 'CLAUDE.md', 'AGENTS.md', '.codex'];
12
12
 
13
+ // BUG-C21-12 (SPEC FR-005): spawnSync defaults to NO timeout, so a wedged npm
14
+ // (network partition, hung credential helper) would block the CLI forever.
15
+ // Per-operation CLI budgets: a pure `npm root -g` resolution is cheap (~15s);
16
+ // `npm install -g` and the post-update reinstall child get a generous 10-minute
17
+ // bound so slow-but-alive installs still finish.
18
+ // killSignal SIGKILL, not the default SIGTERM: a wedged child that ignores
19
+ // SIGTERM (npm's signal-exit cleanup handler can itself wedge) would keep
20
+ // spawnSync waiting past the timeout — verified on Node v22: SIGTERM on a
21
+ // signal-ignoring child blocks indefinitely, SIGKILL returns ~immediately with
22
+ // ETIMEDOUT. Direct-child kill only (no process-group tree kill): these are
23
+ // short-lived CLI one-shots; the detached-group TERM→KILL posture stays in
24
+ // templates/.claude/ukit/runtime/hook-process.mjs where hooks need it.
25
+ export const NPM_QUERY_TIMEOUT_MS = 15_000;
26
+ export const NPM_INSTALL_TIMEOUT_MS = 10 * 60 * 1000;
27
+ const KILL_ON_TIMEOUT = 'SIGKILL';
28
+
13
29
  export function hasUkitInstalled(projectRoot) {
14
30
  return INSTALLED_MARKERS.some((marker) => fs.existsSync(path.join(projectRoot, marker)));
15
31
  }
@@ -23,12 +39,23 @@ export function hasUkitInstalled(projectRoot) {
23
39
  * Only a fresh child process picks up the newly written package.
24
40
  */
25
41
  export function resolveGlobalUkitBin({ spawnSync = defaultSpawnSync } = {}) {
26
- const result = spawnSync('npm', ['root', '-g'], { encoding: 'utf8' });
42
+ return resolveGlobalUkitBinDetailed({ spawnSync }).binPath;
43
+ }
44
+
45
+ // Detailed variant: also reports WHY resolution failed, so callers can surface
46
+ // a timeout as actionable ("npm query timed out — is the network/proxy wedged?")
47
+ // instead of misattributing it to a missing file.
48
+ function resolveGlobalUkitBinDetailed({ spawnSync = defaultSpawnSync } = {}) {
49
+ const result = spawnSync('npm', ['root', '-g'], {
50
+ encoding: 'utf8',
51
+ timeout: NPM_QUERY_TIMEOUT_MS,
52
+ killSignal: KILL_ON_TIMEOUT,
53
+ });
27
54
  if (result.error || result.status !== 0 || typeof result.stdout !== 'string') {
28
- return null;
55
+ return { binPath: null, timedOut: result.error?.code === 'ETIMEDOUT' };
29
56
  }
30
57
  const binPath = path.join(result.stdout.trim(), ...UKIT_PACKAGE_NAME.split('/'), 'bin', 'ukit');
31
- return fs.existsSync(binPath) ? binPath : null;
58
+ return { binPath: fs.existsSync(binPath) ? binPath : null, timedOut: false };
32
59
  }
33
60
 
34
61
  /**
@@ -45,14 +72,16 @@ export function runInstallAfterUpdate({
45
72
  return { ran: false, reason: 'not-a-ukit-project' };
46
73
  }
47
74
 
48
- const binPath = resolveGlobalUkitBin({ spawnSync });
75
+ const { binPath, timedOut } = resolveGlobalUkitBinDetailed({ spawnSync });
49
76
  if (!binPath) {
50
- return { ran: false, reason: 'global-bin-not-found' };
77
+ return { ran: false, reason: timedOut ? 'global-root-query-timeout' : 'global-bin-not-found' };
51
78
  }
52
79
 
53
80
  const result = spawnSync(execPath, [binPath, 'install'], {
54
81
  cwd: projectRoot,
55
82
  stdio: 'inherit',
83
+ timeout: NPM_INSTALL_TIMEOUT_MS,
84
+ killSignal: KILL_ON_TIMEOUT,
56
85
  });
57
86
  if (result.error) {
58
87
  return { ran: false, reason: `install-failed: ${result.error.message}` };
@@ -66,6 +95,8 @@ export function runInstallAfterUpdate({
66
95
  export function updateUkit({ spawnSync = defaultSpawnSync } = {}) {
67
96
  const result = spawnSync('npm', ['install', '-g', UKIT_PACKAGE_NAME], {
68
97
  stdio: 'inherit',
98
+ timeout: NPM_INSTALL_TIMEOUT_MS,
99
+ killSignal: KILL_ON_TIMEOUT,
69
100
  });
70
101
 
71
102
  if (result.error) {
@@ -127,17 +127,28 @@ Same model is the most common silent failure. Do not skip this check.
127
127
  ### Inputs you expect
128
128
 
129
129
  - Path to the spec/plan document (e.g. `docs/plans/*.md`). No diff, no task file, no executor report — review the document itself.
130
+ - When invoked from the handoff pipeline you get BOTH `docs/AI_HANDOFF/SPEC.md` and `docs/AI_HANDOFF/PLAN.md`. Review them as one unit: the spec is the contract, the plan is the decomposition. Verdicts still append to PLAN.md's `## Plan Review Log`.
130
131
 
131
132
  ### Review order
132
133
 
133
134
  | Category | What to look for |
134
135
  |---|---|
135
136
  | Completeness | TODO/TBD/placeholders, incomplete sections |
136
- | Consistency | internal contradictions, conflicting requirements |
137
+ | Consistency | internal contradictions, conflicting requirements; SPEC and PLAN contradicting each other |
137
138
  | Clarity | requirements ambiguous enough to cause a wrong build |
138
139
  | Scope | focused enough for one plan, not silently covering multiple subsystems |
139
140
  | YAGNI | unrequested features, over-engineering |
140
141
 
142
+ **Handoff spec quality gate** (only when reviewing `docs/AI_HANDOFF/SPEC.md`):
143
+
144
+ 1. Every functional requirement is testable — Given/When/Then or a command, never "should work".
145
+ 2. Every applicable fullstack layer is covered or explicitly `N/A` with a reason.
146
+ 3. Every FR traces forward to plan scope; nothing in the plan is unbacked by the spec.
147
+ 4. Dependencies between parts are explicit.
148
+ 5. Legacy/unfinished work discovered by the Phase 0 sweep is either planned or recorded out-of-scope.
149
+ 6. Rollback/migration impact is addressed when data or schema changes.
150
+ 7. No vague instruction survives — "improve UI", "faster", "better UX" without defined behavior fails Clarity.
151
+
141
152
  Only flag issues that would cause real problems during implementation planning. Approve unless there are serious gaps that would lead to a flawed plan.
142
153
 
143
154
  ### Output
@@ -19,8 +19,10 @@ reasoning to the parent agent so it can decide whether to re-route.
19
19
  **In Handoff mode you are running unattended — ask nothing.** You were spawned by an
20
20
  orchestrator driving a pipeline; there is no human in your conversation to answer, and a
21
21
  question there is silently dropped while the run stalls. Resolve ambiguity in this order:
22
- the task file → `PLAN.md` → the surrounding code's existing patterns → the choice you would
23
- recommend. Record what you chose and why in the task's `## Discussion` thread. Only a blocker
22
+ the task file → the `Spec references` sections of `SPEC.md` → `PLAN.md` → the surrounding
23
+ code's existing patterns → the choice you would recommend. The spec is the contract; if the
24
+ task file and spec disagree, implement the spec and note it in the task's `## Discussion`
25
+ thread. Record what you chose and why in that thread. Only a blocker
24
26
  outside the repo (missing credential, unreachable service) justifies reporting `FAIL` early —
25
27
  and even then, report it, don't ask about it.
26
28
 
@@ -55,6 +55,35 @@ Before writing any path or command into `PLAN.md` or a task file, verify it:
55
55
  If something cannot be verified, say so in the task's `## Discussion` rather than guessing.
56
56
  A stated unknown costs the executor one read; a wrong path costs it a round.
57
57
 
58
+ ## Phase 0 — Legacy sweep (before writing anything)
59
+
60
+ The plan owns ALL unfinished work, not only the new request. Scan and fold in:
61
+
62
+ - `INDEX.md` rows that are not `done`/`cancelled_superseded`.
63
+ - Task files in stale `in_progress` / `blocked` / `changes_requested` /
64
+ `needs_executor_report` / `needs_breakdown` from dead sessions.
65
+ - `docs/AI_HANDOFF/HISTORY.md` + `archive/` — cycles closed with leftovers.
66
+ - `docs/TASKS.md` — `Ready for AI` items are newly-assigned work.
67
+ - `git status` — uncommitted work-in-progress (finish or checkpoint, never drop silently).
68
+
69
+ Every discovered item becomes either a task row in the new plan or an explicitly recorded
70
+ out-of-scope line in PLAN.md §2. Silent omission is a plan defect.
71
+
72
+ **Recovery:** a stuck task record that cannot be cleanly resumed (orphaned worktree,
73
+ contradicting reports, invalid state) is marked `cancelled_superseded` and replaced by
74
+ `TASK-xxx-R1` (`-R2`, …) carrying the same spec references, acceptance criteria and
75
+ verification — link both files' `## Discussion` threads.
76
+
77
+ ## Phase 0.5 — Write SPEC.md
78
+
79
+ Write `docs/AI_HANDOFF/SPEC.md` from the template at `templates/docs/AI_HANDOFF/SPEC.md`
80
+ (15 sections). The spec is the contract executors implement against — concrete enough that
81
+ nothing is guessed: exact paths, module and API names, schemas, validation rules,
82
+ permissions, empty/error states, migration behavior, test expectations. Every section is
83
+ filled or marked `N/A — <reason>`; every open question is resolved to a chosen default
84
+ recorded in §14. A vague line ("improve UI", "make it faster" with no number) is a spec
85
+ defect — fix it before writing tasks.
86
+
58
87
  ## Phase 1 — Write PLAN.md
59
88
 
60
89
  Write all 7 sections to `docs/AI_HANDOFF/PLAN.md`:
@@ -101,6 +130,7 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
101
130
 
102
131
  | Field | Rule |
103
132
  |-------|------|
133
+ | Spec references | SPEC.md section/FR IDs this task implements — every task traces to the spec |
104
134
  | Target Files | Exact paths — no two tasks in same wave share a file |
105
135
  | Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
106
136
  | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
@@ -160,6 +190,7 @@ most chains are ordering preferences that a wide wave 1 would satisfy just as we
160
190
  ```
161
191
  Cycle: <ID> Date: <YYYY-MM-DD> Base: <current HEAD branch>
162
192
  Goal: <1 sentence>
193
+ Spec: docs/AI_HANDOFF/SPEC.md
163
194
  Tasks: <N> total
164
195
  Status: planning_done — ready for executor
165
196
  ```
@@ -179,6 +210,8 @@ catches it:
179
210
  2. Does every task trace back to something in §1/§6? A task nothing asks for is scope creep — cut it.
180
211
  3. Do the tasks together actually deliver §1's success definition, or only the easy part of it? State the gap if there is one.
181
212
  4. Is the *unhappy* path planned — errors, empty input, permissions, migration of existing data — or only the feature?
213
+ 4b. Does every task carry `Spec references` into SPEC.md, and does every SPEC.md FR trace to at least one task? A spec section no task implements is a silent hole.
214
+ 4c. Did the Phase 0 sweep leave anything unplanned — stale tasks, legacy leftovers, Ready-for-AI items, uncommitted WIP — without a §2 out-of-scope line?
182
215
 
183
216
  **Correctness — is anything wrong?**
184
217
  5. Every `Target Files` path verified per Grounding? Any `(new)` file marked as such?
@@ -195,7 +228,7 @@ catches it:
195
228
  Append the result to `PLAN.md`:
196
229
  ```
197
230
  ## Planner Self-Audit
198
- Checklist: 12/12 pass
231
+ Checklist: 14/14 pass
199
232
  Fixed during audit: <what you changed, or "nothing">
200
233
  Known gaps: <what you deliberately left out and why, or "none">
201
234
  ```
@@ -209,6 +242,8 @@ Keep the returned message under 25 lines — the caller may be an orchestrator w
209
242
  budget is the constraint on the whole run. Detail belongs in `PLAN.md`, not in the reply.
210
243
 
211
244
  - Task count + IDs
245
+ - Spec path + one-line coverage statement (`SPEC.md §5 FR-001→TASK-003`, …)
246
+ - Recovery/superseded pairs (`TASK-007 → TASK-007-R1`) | none
212
247
  - Dependency graph (text form: TASK-001 → TASK-003, TASK-002 independent)
213
248
  - Wave plan: `wave 1: N tasks | wave 2: M tasks` — flag it if the graph is mostly a chain
214
249
  - Self-audit result + any `Known gaps`
@@ -37,6 +37,7 @@ Write `docs/AI_HANDOFF/archive/cycle-NNN.md` (if there is anything worth archivi
37
37
  ```
38
38
  # Cycle NNN — <YYYY-MM-DD> — ABORTED
39
39
  ## Summary: cycle was cleared before completion
40
+ ## Spec: <copy SPEC.md — or note its path if archived separately>
40
41
  ## Tasks: <copy INDEX.md table as-is>
41
42
  ```
42
43
  If `archive/` has > 3 files → delete oldest, append 1-line summary to `HISTORY.md`.
@@ -45,8 +46,11 @@ If `archive/` has > 3 files → delete oldest, append 1-line summary to `HISTORY
45
46
 
46
47
  ```
47
48
  PLAN.md → "# PLAN\n_(empty)_"
49
+ SPEC.md → restore the untouched template (keep section headings, clear content)
48
50
  INDEX.md → empty table header only
49
51
  ACTIVE.md → "# ACTIVE\n_(no active cycle)_"
52
+ RUN.md → set `Phase: done` (or delete) — a live cursor that is neither `done` nor `blocked`
53
+ makes the Stop gate refuse the next session's stops; clearing a cycle MUST close the cursor.
50
54
  tasks/TASK-*.md → delete all (keep _TEMPLATE.md)
51
55
  ```
52
56
 
@@ -69,7 +69,21 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
69
69
  BASE=$(git symbolic-ref --short HEAD)
70
70
  ```
71
71
 
72
- 3. Write `docs/AI_HANDOFF/PLAN.md` — all 6 sections mandatory:
72
+ 3. Write `docs/AI_HANDOFF/SPEC.md` — the detailed implementation spec (BẮT BUỘC before tasks).
73
+ Follow `docs/AI_HANDOFF/SPEC.md`'s template sections (or `templates/docs/AI_HANDOFF/SPEC.md`
74
+ on a fresh tree): problem/context, goals, non-goals, user journeys, functional requirements
75
+ with Given/When/Then + error cases, fullstack scope (backend, schema/migrations, API
76
+ contract, UI+state, integration, security, performance, observability, deploy/rollback),
77
+ data model, edge cases, test matrix, acceptance criteria, migration steps, open questions
78
+ with chosen defaults, review checklist.
79
+
80
+ The spec must be concrete enough that an executor implements it WITHOUT guessing: exact
81
+ file paths, module names, API methods, statuses, schemas, validation rules, permissions,
82
+ empty/error states, migration behavior, and test expectations. Open questions are
83
+ resolved to a chosen default recorded inline — the plan phase is the only question
84
+ window, so anything left "TBD" becomes a guess downstream.
85
+
86
+ 4. Write `docs/AI_HANDOFF/PLAN.md` — all 6 sections mandatory:
73
87
  - §1 Intent — problem + success definition
74
88
  - §2 Scope — in / out of scope. **Add a constraint**: same-wave tasks must not modify the same file (prevents merge conflicts). If two tasks need the same file, make one depend on the other.
75
89
  - §3 Approach — solution, trade-offs, alternatives rejected
@@ -83,8 +97,9 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
83
97
  PLANNER_MODEL: <your exact model ID>
84
98
  ```
85
99
 
86
- 4. Create `docs/AI_HANDOFF/tasks/TASK-001.md`, `TASK-002.md`... from `_TEMPLATE.md`
100
+ 5. Create `docs/AI_HANDOFF/tasks/TASK-001.md`, `TASK-002.md`... from `_TEMPLATE.md`
87
101
  Every task MUST have:
102
+ - Spec references (SPEC.md section IDs this task implements)
88
103
  - Target Files (exact paths — no two tasks in same wave share a file)
89
104
  - Dependencies (`TASK-xxx` or `none` — wave structure inferred from this, not stored separately)
90
105
  - Test Cases (Type | Name | Expected — ≥1 happy + ≥2 edge cases of different kinds)
@@ -93,18 +108,19 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
93
108
  - Acceptance Criteria (verifiable checklist)
94
109
  Missing any field → status: `needs_breakdown`, never `ready`
95
110
 
96
- 5. Update `INDEX.md` — one row per task, `status=ready`
111
+ 6. Update `INDEX.md` — one row per task, `status=ready`
97
112
 
98
- 6. Update `ACTIVE.md`:
113
+ 7. Update `ACTIVE.md`:
99
114
  ```
100
115
  Cycle: <ID> Date: <YYYY-MM-DD> Base: <BASE>
101
116
  Goal: <1 sentence>
117
+ Spec: docs/AI_HANDOFF/SPEC.md
102
118
  Tasks: <N> total
103
119
  Status: planning_done — ready for executor
104
120
  ```
105
121
  Note: wave structure is inferred from task Dependencies fields — not stored here.
106
122
 
107
- 7. Report: task IDs, dependency graph, any `needs_breakdown` + reason
123
+ 8. Report: task IDs, dependency graph, any `needs_breakdown` + reason
108
124
 
109
125
  > **For the human operator, on a tool with no agent support (Codex, OpenCode) — not an instruction to the model:** manually switch to the strong model, execute steps 1–7 above yourself.
110
126
 
@@ -121,9 +137,9 @@ The planner agent does the following (use Step 1 summary — do NOT re-read file
121
137
  Two independent strong-model passes shape the plan before any code is written — that gate is
122
138
  intact. What it no longer does is hand a stalled plan back and wait.
123
139
 
124
- **Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"` (omp: the `task` tool with `agent: "code-reviewer"`), passing `REVIEW_TARGET_TYPE=plan` and the path to `docs/AI_HANDOFF/PLAN.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
140
+ **Claude Code — MANDATORY, do this before anything else:** call the Agent tool with `subagent_type: "code-reviewer"` (omp: the `task` tool with `agent: "code-reviewer"`), passing `REVIEW_TARGET_TYPE=plan` and the paths to `docs/AI_HANDOFF/PLAN.md` AND `docs/AI_HANDOFF/SPEC.md`. This MUST be a separate agent invocation from Step 2's `handoff-planner` call (fresh context) — same-session self-review defeats the purpose of an independent gate.
125
141
 
126
- 1. Reviewer reads `PLAN.md` only (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
142
+ 1. Reviewer reads `SPEC.md` + `PLAN.md` (no diff, no task files, no executor report), checks Completeness / Consistency / Clarity / Scope / YAGNI plus the spec quality gate — every requirement testable, every fullstack layer covered, dependencies explicit, no vague instruction left — see `.claude/agents/code-reviewer.md` → Spec/Plan Review — and appends its verdict to PLAN.md's `## Plan Review Log` (new round entry, prior rounds kept).
127
143
  2. `Issues Found` → route back to Step 2: planner revises `PLAN.md` and the affected `TASK-xxx.md` files to address every finding, then re-submit for another Step 2.5 review (this becomes the next round). Do NOT commit or hand off to executor on `Issues Found`.
128
144
  3. `Approved` → append `PLAN_REVIEW: Approved by <reviewer model>` to PLAN.md's `## Planner Report` footer, then proceed.
129
145