@ngockhoale/ukit 2.1.3 → 2.1.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,27 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.1.4 - 2026-08-19
6
+
7
+ More room before compaction, the handoff hooks stop taxing every call they were meant to speed up, and the shipped handoff instructions stop contradicting the config they tell agents to read.
8
+
9
+ ### Changed
10
+
11
+ - **`compact.hardCapTokens` 160,000 → 220,000 and `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW` 150,000 → 180,000.** 2.1.3 fixed the deadlock but left the pair sized for a 200k window, which meant compacting far more often than necessary on a 256k model. The new pair leaves roughly 40k of working room after auto-compact fires, and the ordering invariant `autoCompactWindow < hardCapTokens` still holds — `tests/core/autoCompactWindow.test.js` fails if either number moves in isolation. **On a 200k model, lower `hardCapTokens` back to 160,000 or below**; the `_help` text in `config.json` says so, and the gate cannot protect you if the cap sits above your real context window.
12
+ - **Config-missing fallbacks follow the default**: `compact-threshold.mjs` and `context-window-guard.sh` now fall back to `220_000` rather than `160_000`, so a corrupt or absent `config.json` behaves like a fresh install instead of silently reverting to the old cap.
13
+
14
+ ### Fixed
15
+
16
+ - **The shipped handoff instructions no longer restate a config value that had gone stale.** Eight lines across `handoff-fullstack.md`, `handoff-implement.md`, `handoff-review.md` and `handoff-planner.md` told agents `handoff.maxParallelAgents` defaults to **10**, while `config.json` in the same release ships **12** — so an installed project got instructions that contradicted its own runtime config. Every one of those lines already said "read from `.ukit/storage/config.json`", so the restated number was redundant as well as wrong; it is now gone and the value has exactly one home. This was the second occurrence (the first was the `3 → 10` bump), and the cause both times was hand-copying a config value into prose. `tests/consistency/configDocsSync.test.js` now fails the suite if any doc restates a numeric `handoff.*` default that contradicts config, in English or Vietnamese, and if a `templates/.claude` file without `{{ }}` variables drifts from its active copy — the "keep the two mirrors identical" rule was written in two places but had never been enforced by anything except a maintainer running `diff` by hand.
17
+ - **The worktree early-exit guard no longer costs more than it saves.** `skill-router.sh` and `stale-spec-guard.sh` skip routing for tool calls scoped to a disposable `.worktrees/task-*` tree — but the first cut spawned `node -e` on every call to decide that, measured at 53ms → 124ms per call in the main tree, where the skip never applies. A pure-bash `case "$INPUT" in *".worktrees/"*)` pre-check now gates the node spawn; it is a strict superset of what the node guard can match, so the main tree pays nothing. Both hooks still fail **open** — any parse error, missing field or node failure falls through to the normal path.
18
+
19
+ ### Added
20
+
21
+ - **`.claude/ukit/index/provision-worktree.mjs`** — deterministic worktree provisioning for handoff waves, packaged via `manifests/platform.full.yaml`. It hides the worktree's `node_modules` symlink through the worktree's own `.git/info/exclude` (a symlink never matched `.gitignore`'s `node_modules/`, so an orchestrator `git add -A` would have committed a machine-local absolute path), and skips git-tracked files when copying `.claude/` so it cannot clobber `.claude/commands/ukit/handoff-fullstack.md`. **Shipped but not yet wired into `handoff-implement`** — the wiring is queued for a later release.
22
+ - **Test-selection policy in `docs/AI_HANDOFF/RULES.md`** — per-task related tests are the floor, and a full suite at the wave/cycle boundary is mandatory, not optional. This is the compensating net that makes narrowed per-task scope safe; the agent definitions already cited it before it existed.
23
+ - **`scripts/bench/parallel-agents.mjs`** — harness for measuring concurrent test-process contention at several parallelism levels, recording `slowdownFactor` / `perRunCost` / a `recommended` value plus the load averages the run happened under. It measures test-process contention, not LLM agents.
24
+ - **`handoff.maxParallelAgents`: 10 → 12.** Chosen, not measured — the benchmark above was never run on a quiet enough machine, so this is the middle of the ~10-15 safe band the config's own `_help` has always documented. Lower it to 3-5 if your tasks produce long verification output or you see compaction firing repeatedly.
25
+
5
26
  ## 2.1.3 - 2026-08-16
6
27
 
7
28
  The reason auto-compact never fired in the terminal was UKit itself. This makes it fire.
@@ -1238,6 +1238,19 @@ items:
1238
1238
  packs:
1239
1239
  - core
1240
1240
 
1241
+ - id: ukit-index-provision-worktree-script
1242
+ type: config
1243
+ sourceTemplate: .claude/ukit/index/provision-worktree.mjs
1244
+ targetPath: .claude/ukit/index/provision-worktree.mjs
1245
+ requires:
1246
+ - ukit-runtime-text-profile-script
1247
+ - ukit-runtime-safe-patch-core-script
1248
+ mergeStrategy: overwrite_with_backup
1249
+ variables: []
1250
+ enabledByDefault: true
1251
+ packs:
1252
+ - core
1253
+
1241
1254
  - id: ukit-index-stale-spec-check-script
1242
1255
  type: config
1243
1256
  sourceTemplate: .claude/ukit/index/stale-spec-check.mjs
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.1.3",
3
+ "version": "2.1.4",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, Antigravity, OpenAI Codex, and OpenCode.",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -0,0 +1,153 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * parallel-agents.mjs — TASK-004 harness.
4
+ *
5
+ * Measures the *machine-contention* component of N parallel agents by spawning N concurrent
6
+ * child processes (default `yarn test:release-core`) per level and recording wall-clock.
7
+ * This deliberately does NOT claim to benchmark LLM agent wall-clock (token counts, server
8
+ * load and retries make that irreproducible) — see TASK-004 planner note #1.
9
+ *
10
+ * Usage:
11
+ * node scripts/bench/parallel-agents.mjs [--levels 1,3,5,10] [--cmd "<test command>"] [--out <path>]
12
+ */
13
+
14
+ import fs from 'node:fs';
15
+ import os from 'node:os';
16
+ import path from 'node:path';
17
+ import { spawn } from 'node:child_process';
18
+
19
+ const SLOWDOWN_LIMIT = 2.0;
20
+
21
+ function parseArgs(argv) {
22
+ const opts = { levels: '1,3,5,10', cmd: 'yarn test:release-core', out: '.cache/bench/parallel-agents.json' };
23
+ for (let i = 0; i < argv.length; i += 2) {
24
+ const key = argv[i];
25
+ const val = argv[i + 1];
26
+ if ((key === '--levels' || key === '--cmd' || key === '--out') && val !== undefined) {
27
+ opts[key.slice(2)] = val;
28
+ }
29
+ }
30
+ return opts;
31
+ }
32
+
33
+ function parseLevels(raw) {
34
+ if (!/^\d+(,\d+)*$/.test(raw.trim())) {
35
+ throw new Error(`unparseable --levels value: "${raw}" (expected comma-separated positive integers, e.g. 1,3,5,10)`);
36
+ }
37
+ const levels = raw.split(',').map((s) => parseInt(s, 10));
38
+ if (levels.some((n) => n < 1)) throw new Error(`--levels values must be >= 1, got: ${raw}`);
39
+ return levels;
40
+ }
41
+
42
+ function runLevel(n, cmd) {
43
+ return new Promise((resolve) => {
44
+ const t0 = Date.now();
45
+ let live = 0;
46
+ let highWater = 0;
47
+ let failures = 0;
48
+ let settled = 0;
49
+ const children = [];
50
+ for (let i = 0; i < n; i += 1) {
51
+ const child = spawn(cmd, { shell: true, stdio: 'ignore' });
52
+ live += 1;
53
+ highWater = Math.max(highWater, live);
54
+ children.push(child);
55
+ child.on('error', () => {
56
+ live -= 1;
57
+ settled += 1;
58
+ failures += 1;
59
+ maybeFinish();
60
+ });
61
+ child.on('exit', (code) => {
62
+ live -= 1;
63
+ settled += 1;
64
+ if (code !== 0) failures += 1;
65
+ maybeFinish();
66
+ });
67
+ }
68
+ function maybeFinish() {
69
+ if (settled === n) {
70
+ resolve({ n, wallClockMs: Date.now() - t0, failures, concurrencyHighWaterMark: highWater });
71
+ }
72
+ }
73
+ });
74
+ }
75
+
76
+ function fail(msg) {
77
+ process.stderr.write(`${msg}\n`);
78
+ process.exit(1);
79
+ }
80
+
81
+ async function main() {
82
+ const opts = parseArgs(process.argv.slice(2));
83
+ let levels;
84
+ try {
85
+ levels = parseLevels(opts.levels);
86
+ } catch (err) {
87
+ fail(err.message);
88
+ }
89
+
90
+ const cpuCount = os.cpus().length;
91
+ const loadAvgStart = os.loadavg();
92
+ const startedAt = new Date().toISOString();
93
+
94
+ const rows = [];
95
+ for (const n of levels) {
96
+ // eslint-disable-next-line no-await-in-loop -- levels are measured sequentially on purpose
97
+ rows.push(await runLevel(n, opts.cmd));
98
+ }
99
+ const base = rows[0].wallClockMs;
100
+ for (const row of rows) {
101
+ row.slowdownFactor = row.wallClockMs / base;
102
+ row.perRunCost = row.wallClockMs / row.n;
103
+ }
104
+
105
+ const eligible = rows.filter((r) => r.slowdownFactor <= SLOWDOWN_LIMIT);
106
+ const recommended = eligible.reduce((best, r) => (r.perRunCost < best.perRunCost ? r : best), eligible[0]).n;
107
+
108
+ const loadAvgEnd = os.loadavg();
109
+ const finishedAt = new Date().toISOString();
110
+
111
+ const result = {
112
+ cpuCount,
113
+ levels: rows.map(({ n, wallClockMs, slowdownFactor, perRunCost, failures, concurrencyHighWaterMark }) => ({
114
+ n, wallClockMs, slowdownFactor, perRunCost, failures, concurrencyHighWaterMark,
115
+ })),
116
+ recommended,
117
+ measurementConditions: { loadAvgStart, loadAvgEnd, startedAt, finishedAt, cpuCount },
118
+ };
119
+
120
+ // Print the human table first so a failed --out write still shows the numbers.
121
+ const header = 'n'.padStart(4) + ' ' + 'wallClockMs'.padStart(12) + ' ' + 'slowdownFactor'.padStart(15) + ' ' + 'perRunCost'.padStart(11) + ' ' + 'failures'.padStart(8) + ' ' + 'concurrencyHighWaterMark'.padStart(23);
122
+ console.log(header);
123
+ for (const r of result.levels) {
124
+ console.log(
125
+ String(r.n).padStart(4)
126
+ + ' ' + String(r.wallClockMs).padStart(12)
127
+ + ' ' + r.slowdownFactor.toFixed(2).padStart(15)
128
+ + ' ' + r.perRunCost.toFixed(1).padStart(11)
129
+ + ' ' + String(r.failures).padStart(8)
130
+ + ' ' + String(r.concurrencyHighWaterMark).padStart(23),
131
+ );
132
+ }
133
+ console.log(`recommended: ${recommended} (best perRunCost with slowdownFactor <= ${SLOWDOWN_LIMIT}) on ${cpuCount} CPUs`);
134
+ console.log(`loadAvgStart: [${loadAvgStart.map((v) => v.toFixed(2)).join(', ')}] loadAvgEnd: [${loadAvgEnd.map((v) => v.toFixed(2)).join(', ')}]`);
135
+
136
+ // Overwrite (never append), and never leave a partial file behind.
137
+ const outPath = path.resolve(opts.out);
138
+ try {
139
+ fs.mkdirSync(path.dirname(outPath), { recursive: true });
140
+ const body = JSON.stringify(result, null, 2) + '\n';
141
+ fs.writeFileSync(outPath, body);
142
+ } catch (err) {
143
+ fail(`cannot write --out ${outPath}: ${err.message}`);
144
+ }
145
+
146
+ const totalFailures = result.levels.reduce((s, r) => s + r.failures, 0);
147
+ if (totalFailures > 0) {
148
+ const perLevel = result.levels.filter((r) => r.failures > 0).map((r) => `n=${r.n}: ${r.failures}/${r.n} failed`).join('; ');
149
+ fail(`bench command failed: ${perLevel}`);
150
+ }
151
+ }
152
+
153
+ main().catch((err) => fail(err && err.message ? err.message : String(err)));
@@ -30,7 +30,7 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
30
30
  0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
31
31
  1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
32
32
  2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
33
- 3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider test suite if shared code was touched.
33
+ 3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider suite only when the diff touched shared code; otherwise the task's own targeted commands are the gate and the wave-boundary full `yarn test` (see docs/AI_HANDOFF/RULES.md "Test selection") is the regression net.
34
34
  4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
35
35
  5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
36
36
  6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
@@ -105,7 +105,7 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
105
105
  | Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
106
106
  | Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
107
107
  | Test Files | Exact test file paths to create/modify |
108
- | Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
108
+ | Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above). Apply the "Test selection" resolution order from docs/AI_HANDOFF/RULES.md: `src/`/`scripts/` targets → the `tests` array in `.cache/index/tests-map.json`; `templates/.claude/**`/`.claude/**` targets → path convention (hooks → `tests/hooks/` + `tests/handoff/cycle*/`; manifest/settings → `tests/manifest/`; runtime `.mjs` mirrors → `tests/core/*Parity*` + `tests/index/`); if both resolve to fewer than one test file the task MUST fall back to `yarn test:release-core` — never the full suite by default, never an empty selection |
109
109
  | Acceptance Criteria | Verifiable checklist |
110
110
 
111
111
  Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
@@ -120,7 +120,7 @@ Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
120
120
 
121
121
  Wave width is the single biggest lever on how long a cycle takes: a wave of 6 finishes in
122
122
  roughly the time of its slowest task, while a chain of 6 takes six times that. Executors run
123
- up to `handoff.maxParallelAgents` (default 10) at once, so a plan that produces `none`
123
+ up to `handoff.maxParallelAgents` at once, so a plan that produces `none`
124
124
  dependencies for most tasks is dramatically faster than one that produces a chain.
125
125
 
126
126
  **Write `Dependencies: none` unless B genuinely cannot be written without A's output.** A real
@@ -223,7 +223,7 @@ must not share a file) in case it slipped through review. Only mark a task `need
223
223
  if it is missing required fields — never merely for sharing a file.
224
224
 
225
225
  **Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
226
- `.ukit/storage/config.json` (default **10**). A wave with more tasks than that is split
226
+ `.ukit/storage/config.json`. A wave with more tasks than that is split
227
227
  into consecutive batches of at most that many; finish one batch completely (including 3c
228
228
  copy-back and worktree deletion) before starting the next.
229
229
 
@@ -234,7 +234,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
234
234
 
235
235
  ### I3 — Execute wave by wave (code model agents)
236
236
 
237
- **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents` — default 10), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
237
+ **Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents`), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
238
238
 
239
239
  For each wave:
240
240
 
@@ -398,7 +398,7 @@ If that diff is empty → implement was not completed. Do not stop: re-enter Pha
398
398
 
399
399
  ### R2 — Model isolation check (strong model, always first)
400
400
 
401
- **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
401
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
402
402
 
403
403
  **Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
404
404
 
@@ -94,7 +94,7 @@ when it is missing required fields — never merely for sharing a file.
94
94
 
95
95
  ### Batch each wave — mandatory
96
96
 
97
- Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). A wave
97
+ Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. A wave
98
98
  with more tasks than that is split into consecutive batches of at most that many; finish
99
99
  one batch completely (including 3c copy-back and worktree deletion) before starting the
100
100
  next.
@@ -50,7 +50,7 @@ If that diff is empty → handoff-implement was not completed. Report which task
50
50
 
51
51
  ## Step 2 — Review the diff
52
52
 
53
- **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json` (default **10**). If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
53
+ **Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
54
54
 
55
55
  **Claude Code — MANDATORY, do this before anything else:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
56
56
 
@@ -1,6 +1,7 @@
1
1
  #!/bin/bash
2
2
  # PreToolUse hook: hard-enforce an absolute context token cap (compact.hardCapTokens,
3
- # default 160000 — must stay below the model's real context window), separate from the
3
+ # default 220000, sized for a 256k window — must stay below the model's real context
4
+ # window, so lower it on a 200k model), separate from the
4
5
  # soft/hard advisory pressure phases in
5
6
  # compact-threshold.mjs (default soft=50000/hard=80000, which only print a suggestion).
6
7
  #
@@ -75,7 +75,7 @@ function loadHardCap() {
75
75
  const value = JSON.parse(raw)?.compact?.hardCapTokens;
76
76
  if (Number.isFinite(value) && value > 0) return value;
77
77
  } catch { /* fall through to the default */ }
78
- return 160_000;
78
+ return 220_000;
79
79
  }
80
80
 
81
81
  let text;
@@ -4,6 +4,39 @@
4
4
 
5
5
  INPUT=$(cat)
6
6
  PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
7
+
8
+ # Worktree early-exit: when the tool call is scoped to a disposable .worktrees/task-*
9
+ # tree, there is nothing to route — exit before spawning the router. Reads
10
+ # tool_input.file_path / tool_input.path / tool_input.command (never tool_input.pattern).
11
+ # Skips only when ".worktrees/" starts a path component. Fails OPEN: any parse
12
+ # problem, missing field, or node error falls through to the normal hook path below.
13
+ # Pure-bash pre-check first: the node guard can only ever match when the literal
14
+ # ".worktrees/" substring is present, so this case is a strict superset and avoids
15
+ # paying the node spawn cost on the far more common non-worktree call.
16
+ case "$INPUT" in
17
+ *".worktrees/"*)
18
+ if printf '%s' "$INPUT" | node -e '
19
+ const chunks = [];
20
+ process.stdin.on("data", (c) => chunks.push(c));
21
+ process.stdin.on("end", () => {
22
+ try {
23
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
24
+ const toolInput = (payload && typeof payload === "object" && payload.tool_input
25
+ && typeof payload.tool_input === "object") ? payload.tool_input : {};
26
+ const candidates = [toolInput.file_path, toolInput.path, toolInput.command];
27
+ const QMARKS = String.fromCharCode(34, 39, 96);
28
+ const worktreeComponent = new RegExp("(^|[\\/\\s=" + QMARKS + "])\\.worktrees/");
29
+ process.exit(candidates.some((v) => typeof v === "string" && worktreeComponent.test(v)) ? 0 : 1);
30
+ } catch {
31
+ process.exit(1);
32
+ }
33
+ });
34
+ ' >/dev/null 2>&1; then
35
+ exit 0
36
+ fi
37
+ ;;
38
+ esac
39
+
7
40
  STATE_FILE="$PROJECT_ROOT/.claude/ukit/skill-router-state.json"
8
41
  HOOK_DIR="$(cd "$(dirname "$0")" && pwd)"
9
42
  THRESHOLD_SCRIPT="$HOOK_DIR/../ukit/runtime/compact-threshold.mjs"
@@ -10,4 +10,36 @@ if [ ! -f "$SCRIPT" ]; then
10
10
  exit 0
11
11
  fi
12
12
 
13
+ # Worktree early-exit: when the tool call is scoped to a disposable .worktrees/task-*
14
+ # tree, skip the stale-spec check entirely. Reads tool_input.file_path /
15
+ # tool_input.path / tool_input.command (never tool_input.pattern). Skips only when
16
+ # ".worktrees/" starts a path component. Fails OPEN: any parse problem, missing field,
17
+ # or node error falls through to the normal check below.
18
+ # Pure-bash pre-check first: the node guard can only ever match when the literal
19
+ # ".worktrees/" substring is present, so this case is a strict superset and avoids
20
+ # paying the node spawn cost on the far more common non-worktree call.
21
+ case "$INPUT" in
22
+ *".worktrees/"*)
23
+ if printf '%s' "$INPUT" | node -e '
24
+ const chunks = [];
25
+ process.stdin.on("data", (c) => chunks.push(c));
26
+ process.stdin.on("end", () => {
27
+ try {
28
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
29
+ const toolInput = (payload && typeof payload === "object" && payload.tool_input
30
+ && typeof payload.tool_input === "object") ? payload.tool_input : {};
31
+ const candidates = [toolInput.file_path, toolInput.path, toolInput.command];
32
+ const QMARKS = String.fromCharCode(34, 39, 96);
33
+ const worktreeComponent = new RegExp("(^|[\\/\\s=" + QMARKS + "])\\.worktrees/");
34
+ process.exit(candidates.some((v) => typeof v === "string" && worktreeComponent.test(v)) ? 0 : 1);
35
+ } catch {
36
+ process.exit(1);
37
+ }
38
+ });
39
+ ' >/dev/null 2>&1; then
40
+ exit 0
41
+ fi
42
+ ;;
43
+ esac
44
+
13
45
  printf '%s' "$INPUT" | node "$SCRIPT"
@@ -11,7 +11,7 @@
11
11
  "defaultView": "chat",
12
12
  "autoCompactEnabled": true,
13
13
  "env": {
14
- "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "150000"
14
+ "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "180000"
15
15
  },
16
16
  "permissions": {
17
17
  "defaultMode": "bypassPermissions",
@@ -0,0 +1,200 @@
1
+ #!/usr/bin/env node
2
+ /**
3
+ * provision-worktree.mjs — deterministic worktree provisioner (TASK-002).
4
+ *
5
+ * Provisions what `git worktree add` cannot bring across (all four are gitignored):
6
+ * - node_modules -> SYMLINK to the main tree (28M; avoids N concurrent installs)
7
+ * - .claude/ -> REAL COPY preserving file modes (never a symlink: a worktree
8
+ * edit must not write through into the live main-tree mirror).
9
+ * Mode preservation matters: every .claude/hooks/*.sh is 755 and
10
+ * tests/handoff/cycle4/vision-gate.test.mjs asserts the exec bit.
11
+ * - .ukit/storage/config.json -> copy
12
+ * - .cache/index/ -> copy
13
+ *
14
+ * CLI: node .claude/ukit/index/provision-worktree.mjs <worktree-path> [--main-root <path>]
15
+ * exit 0 on success (including a no-op second run); non-zero + one-line stderr otherwise.
16
+ */
17
+
18
+ import fs from 'node:fs';
19
+ import path from 'node:path';
20
+ import { spawnSync } from 'node:child_process';
21
+
22
+ function fail(msg) {
23
+ process.stderr.write(`provision-worktree: ${msg}\n`);
24
+ process.exit(1);
25
+ }
26
+
27
+ const args = process.argv.slice(2);
28
+ let wtPath = null;
29
+ let mainRoot = null;
30
+ for (let i = 0; i < args.length; i += 1) {
31
+ if (args[i] === '--main-root') {
32
+ if (i + 1 >= args.length) fail('--main-root requires a value');
33
+ mainRoot = args[(i += 1)];
34
+ } else if (wtPath === null) {
35
+ wtPath = args[i];
36
+ } else {
37
+ fail(`unexpected argument: ${args[i]}`);
38
+ }
39
+ }
40
+ if (!wtPath) fail('usage: provision-worktree.mjs <worktree-path> [--main-root <path>]');
41
+ wtPath = path.resolve(wtPath);
42
+ mainRoot = path.resolve(mainRoot ?? process.cwd());
43
+
44
+ // --- validate before mutating anything -------------------------------------
45
+
46
+ if (!fs.existsSync(wtPath) || !fs.statSync(wtPath).isDirectory()) {
47
+ fail(`worktree path is not an existing directory: ${wtPath}`);
48
+ }
49
+ try {
50
+ fs.accessSync(wtPath, fs.constants.W_OK);
51
+ } catch {
52
+ fail(`worktree dir is not writable: ${wtPath}`);
53
+ }
54
+
55
+ const mainModules = path.join(mainRoot, 'node_modules');
56
+ if (!fs.existsSync(mainModules)) {
57
+ fail(`main tree has no node_modules at ${mainModules}; run the install in the main tree first`);
58
+ }
59
+
60
+ function readPkg(p) {
61
+ if (!fs.existsSync(p)) return null;
62
+ try {
63
+ return JSON.parse(fs.readFileSync(p, 'utf8'));
64
+ } catch (err) {
65
+ fail(`cannot parse ${p}: ${err.message}`);
66
+ return null; // unreachable — fail() exits the process
67
+ }
68
+ }
69
+ const mainPkg = readPkg(path.join(mainRoot, 'package.json'));
70
+ const wtPkg = readPkg(path.join(wtPath, 'package.json'));
71
+ if (mainPkg && wtPkg) {
72
+ const depsOf = (o) => JSON.stringify({ d: o.dependencies ?? {}, dd: o.devDependencies ?? {} });
73
+ if (depsOf(mainPkg) !== depsOf(wtPkg)) {
74
+ fail('package.json dependencies/devDependencies differ from the main tree; run a real install in this worktree instead of symlinking node_modules');
75
+ }
76
+ }
77
+
78
+ // --- helpers -----------------------------------------------------------------
79
+
80
+ // The symlinked node_modules must stay invisible to `git status`/`git add -A`
81
+ // in this worktree: .gitignore's `node_modules/` (trailing slash) matches a
82
+ // real directory but NOT a symlink, so without this a handoff pipeline that
83
+ // runs `git add -A` would stage a symlink whose target is an absolute,
84
+ // machine-local path. `info/exclude` resolves to the shared git common dir
85
+ // (verified: worktrees do not get a private copy), so this is a one-time,
86
+ // idempotent, untracked local-metadata change — never part of any commit.
87
+ function excludeNodeModulesFromGit(root) {
88
+ try {
89
+ const common = spawnSync('git', ['-C', root, 'rev-parse', '--git-common-dir'], { encoding: 'utf8' });
90
+ if (common.status !== 0) return false; // not a git repo (e.g. unit-test fixture) — best effort
91
+ const gitCommonDir = path.resolve(root, common.stdout.trim());
92
+ const excludePath = path.join(gitCommonDir, 'info', 'exclude');
93
+ fs.mkdirSync(path.dirname(excludePath), { recursive: true });
94
+ const existing = fs.existsSync(excludePath) ? fs.readFileSync(excludePath, 'utf8') : '';
95
+ if (existing.split('\n').some((l) => l.trim() === 'node_modules')) return false; // already present
96
+ const sep = existing.length > 0 && !existing.endsWith('\n') ? '\n' : '';
97
+ fs.appendFileSync(excludePath, `${sep}node_modules\n`);
98
+ return true;
99
+ } catch {
100
+ return false; // best effort — a missing/unwritable git dir must not fail provisioning
101
+ }
102
+ }
103
+
104
+ // The `.claude` copy must never clobber a git-TRACKED file (today exactly
105
+ // `.claude/commands/ukit/handoff-fullstack.md` — owned by TASK-005). A fresh
106
+ // worktree already has that file checked out by `git worktree add`; if the
107
+ // gitignored dev mirror in mainRoot has locally drifted from the committed
108
+ // copy, an unconditional recursive copy would silently overwrite the
109
+ // worktree's tracked file with main's unreviewed local state.
110
+ function trackedClaudeFiles(root) {
111
+ try {
112
+ const res = spawnSync('git', ['-C', root, 'ls-files', '--', '.claude'], { encoding: 'utf8' });
113
+ if (res.status !== 0 || !res.stdout) return new Set();
114
+ return new Set(res.stdout.split('\n').map((s) => s.trim()).filter(Boolean));
115
+ } catch {
116
+ return new Set();
117
+ }
118
+ }
119
+
120
+ // --- provision -------------------------------------------------------------
121
+
122
+ // 1. node_modules symlink
123
+ const wtModules = path.join(wtPath, 'node_modules');
124
+ let modulesSt = null;
125
+ try {
126
+ modulesSt = fs.lstatSync(wtModules);
127
+ } catch {
128
+ /* absent */
129
+ }
130
+ if (modulesSt === null) {
131
+ try {
132
+ fs.symlinkSync(mainModules, wtModules, 'dir');
133
+ } catch (err) {
134
+ fail(`cannot create node_modules symlink: ${err.message}`);
135
+ }
136
+ console.log(`provisioned node_modules -> symlink ${mainModules}`);
137
+ } else if (modulesSt.isSymbolicLink()) {
138
+ const target = fs.readlinkSync(wtModules);
139
+ if (path.resolve(wtPath, target) !== path.resolve(mainModules)) {
140
+ fail(`node_modules symlink already exists but points at ${target}, not ${mainModules}`);
141
+ }
142
+ console.log(`node_modules symlink already correct -> ${target}`);
143
+ } else {
144
+ fail(`${wtModules} already exists and is not a symlink`);
145
+ }
146
+ if (excludeNodeModulesFromGit(wtPath)) {
147
+ console.log('node_modules excluded from git via .git/info/exclude');
148
+ }
149
+
150
+ // 2. .claude real copy, preserving modes
151
+ const srcClaude = path.join(mainRoot, '.claude');
152
+ const dstClaude = path.join(wtPath, '.claude');
153
+ if (!fs.existsSync(srcClaude)) {
154
+ fail(`main tree has no .claude directory to copy: ${srcClaude}`);
155
+ } else {
156
+ // Always (re)copy: a fresh worktree may already carry a partial, git-tracked
157
+ // .claude/ (e.g. commands), which must not make us skip the hooks/scripts.
158
+ // Never overwrite a git-tracked file under .claude/ (e.g.
159
+ // handoff-fullstack.md) — leave whatever `git worktree add` already
160
+ // checked out for it alone.
161
+ const trackedInClaude = trackedClaudeFiles(mainRoot);
162
+ fs.cpSync(srcClaude, dstClaude, {
163
+ recursive: true,
164
+ filter: (src) => {
165
+ const rel = path.relative(mainRoot, src).split(path.sep).join('/');
166
+ return !trackedInClaude.has(rel);
167
+ },
168
+ });
169
+ console.log(
170
+ `provisioned .claude -> copy of ${srcClaude}` +
171
+ (trackedInClaude.size > 0 ? ` (skipped ${trackedInClaude.size} git-tracked file(s))` : ''),
172
+ );
173
+ }
174
+
175
+ // 3. .ukit/storage/config.json
176
+ const srcConfig = path.join(mainRoot, '.ukit', 'storage', 'config.json');
177
+ const dstConfig = path.join(wtPath, '.ukit', 'storage', 'config.json');
178
+ if (fs.existsSync(dstConfig)) {
179
+ console.log('.ukit/storage/config.json already present (skipped copy)');
180
+ } else if (!fs.existsSync(srcConfig)) {
181
+ fail(`main tree has no .ukit/storage/config.json to copy: ${srcConfig}`);
182
+ } else {
183
+ fs.mkdirSync(path.dirname(dstConfig), { recursive: true });
184
+ fs.copyFileSync(srcConfig, dstConfig);
185
+ console.log('provisioned .ukit/storage/config.json');
186
+ }
187
+
188
+ // 4. .cache/index
189
+ const srcCache = path.join(mainRoot, '.cache', 'index');
190
+ const dstCache = path.join(wtPath, '.cache', 'index');
191
+ if (fs.existsSync(dstCache)) {
192
+ console.log('.cache/index already present in worktree (skipped copy)');
193
+ } else if (!fs.existsSync(srcCache)) {
194
+ fail(`main tree has no .cache/index directory to copy: ${srcCache}`);
195
+ } else {
196
+ fs.cpSync(srcCache, dstCache, { recursive: true });
197
+ console.log('provisioned .cache/index');
198
+ }
199
+
200
+ process.exit(0);
@@ -442,11 +442,13 @@ export function buildCompactThresholds(config = {}) {
442
442
  //
443
443
  // Must stay BELOW the model's real context window or the cap is unreachable: the API
444
444
  // rejects the request with "input exceeds the context window" long before an estimate
445
- // climbing toward a higher number ever trips this. The previous 220_000 default sat
446
- // above a 200k window, so the gate could never fire on a standard model — it was dead
447
- // code in exactly the situation it existed to prevent. 160_000 leaves real headroom on
448
- // a 200k window for the response plus the estimator's own undercount.
449
- const hardCapTokens = Math.max(1, finiteNumber(config?.compact?.hardCapTokens, 160_000));
445
+ // climbing toward a higher number ever trips this. 220_000 is sized for a 256k window
446
+ // and leaves headroom for the response plus the estimator's own undercount. On a 200k
447
+ // model this default sits ABOVE the window, so the gate can never fire and is dead code
448
+ // in exactly the situation it exists to prevent — set compact.hardCapTokens to 160_000
449
+ // or lower for those. Pair it with env.CLAUDE_CODE_AUTO_COMPACT_WINDOW, which must stay
450
+ // strictly below this cap so the client auto-compacts before the gate blocks tools.
451
+ const hardCapTokens = Math.max(1, finiteNumber(config?.compact?.hardCapTokens, 220_000));
450
452
 
451
453
  return {
452
454
  softThreshold,
@@ -79,8 +79,8 @@ Next: <bước kế tiếp chính xác>
79
79
  - Subagent ghi **full log vào task file trên đĩa**, chỉ trả về orchestrator ≤10 dòng (executor) / ≤6 dòng (reviewer). Paste log ngược lại orchestrator là nguyên nhân số 1 làm run chết vì hết context.
80
80
  - Hết mỗi wave: commit, ghi cursor, **collapse** wave đó còn 1 dòng/task trong bộ nhớ làm việc, rồi chạy tiếp.
81
81
  - Yêu cầu `/compact` **chỉ** được đặt ở cuối command, giữa 2 cycle. Giữa cycle thì tuyệt đối không — state đã nằm hết ở git + `INDEX.md` + `RUN.md` nên compact ở ranh giới cycle không mất gì.
82
- - Vượt `compact.hardCapTokens` (mặc định 160k) mà `RUN.md` còn run dở: `context-hardcap-gate` cho thêm `compact.hardCapGraceCalls` (mặc định 10) tool call rồi mới chặn cứng. **Grace đó chỉ để hạ cánh** — hoàn tất edit đang dở, commit, ghi cursor, push. Không mở task mới, không đọc thêm file, không spawn agent. Hết grace là chặn thật; budget chỉ reset khi ước lượng token thực sự giảm (có compact thật), không reset theo wave.
83
- - Không hook nào gọi được `/compact` — đó là lệnh client-only. Nhưng từ 2.1.3, settings mặc định đặt `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW = 150000` < `hardCapTokens` (160k), nên **client tự auto-compact trước khi gate chặn**. Đường thường: auto-compact chạy → `handoff-resume.sh` replay cursor → chạy tiếp, không cần người gõ gì. Grace window ở trên chỉ còn là lưới an toàn.
82
+ - Vượt `compact.hardCapTokens` (mặc định 220k, cỡ cho context window 256k) mà `RUN.md` còn run dở: `context-hardcap-gate` cho thêm `compact.hardCapGraceCalls` (mặc định 10) tool call rồi mới chặn cứng. **Grace đó chỉ để hạ cánh** — hoàn tất edit đang dở, commit, ghi cursor, push. Không mở task mới, không đọc thêm file, không spawn agent. Hết grace là chặn thật; budget chỉ reset khi ước lượng token thực sự giảm (có compact thật), không reset theo wave.
83
+ - Không hook nào gọi được `/compact` — đó là lệnh client-only. Nhưng từ 2.1.3, settings mặc định đặt `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW = 180000` < `hardCapTokens` (220k), nên **client tự auto-compact trước khi gate chặn**. Đường thường: auto-compact chạy → `handoff-resume.sh` replay cursor → chạy tiếp, không cần người gõ gì. Grace window ở trên chỉ còn là lưới an toàn.
84
84
  - Sửa một trong hai số đó thì phải giữ `autoCompactWindow < hardCapTokens`. Đảo thứ tự là deadlock: gate chặn tool trước → transcript ngừng lớn → ngưỡng auto-compact không bao giờ tới. `tests/core/autoCompactWindow.test.js` khóa bất biến này.
85
85
 
86
86
  ### Git
@@ -132,6 +132,17 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
132
132
  - `§ Test Files`: đường dẫn cụ thể file test sẽ tạo/sửa (ví dụ `tests/auth/login.test.js`).
133
133
  - `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS. Nếu project có sẵn lint/typecheck script → BẮT BUỘC liệt kê ở đây, không chỉ lệnh test. Project không có thì ghi rõ N/A, không được bỏ qua im lặng.
134
134
  - `§ Acceptance Criteria`: checklist.
135
+
136
+ #### Test selection (which tests the Verification Commands run)
137
+
138
+ Resolution order — exactly three steps, in this order. A task's Verification Commands MUST NOT default to the full suite:
139
+
140
+ 1. Target File under `src/` or `scripts/` → read `.cache/index/tests-map.json` and take the `tests` array for that `sourceFile`.
141
+ 2. Target File under `templates/.claude/**` or `.claude/**` → `tests-map.json` has no coverage of these paths, so use the path convention: hooks → `tests/hooks/` + `tests/handoff/cycle*/`; manifest/settings → `tests/manifest/`; runtime `.mjs` mirrors → `tests/core/*Parity*` + `tests/index/`.
142
+ 3. **Mandatory non-empty floor** — if steps 1–2 resolve to fewer than one test file, the task MUST fall back to `yarn test:release-core`. An empty selection is never permitted. This floor is a fallback for a single task's narrowed selection, not a default — most tasks resolve via steps 1–2 and never reach it.
143
+
144
+ **Wave/cycle boundary regression net** — the three steps above narrow one task's Verification Commands only; they are not a substitute for full-suite coverage. A full `yarn test` run at each wave/cycle boundary MUST happen and is the regression net for every per-task narrowed selection made under this policy. `code-reviewer.md`'s "wave-boundary full `yarn test` ... is the regression net" sentence refers to this paragraph.
145
+
135
146
  - Nếu split mà task nào không kèm được Test Cases + Test Files cụ thể → task đó chưa đủ `ready`, đánh `needs_breakdown`.
136
147
  - Update `INDEX.md`: thêm row mỗi task với status `ready`.
137
148
  - Đây là **điểm cắt cuối trước khi code chạy**: phase này xong, executor được phép pick. Trong `/ukit:handoff-fullstack`, gate ở đây là plan review độc lập (model mạnh, context riêng) chứ không phải human — vì người dùng đã chủ động chọn chạy one-shot.
@@ -9,7 +9,7 @@
9
9
  "compact": {
10
10
  "enabled": true,
11
11
  "tokenThreshold": 50000,
12
- "hardCapTokens": 160000,
12
+ "hardCapTokens": 220000,
13
13
  "hardCapBlock": true,
14
14
  "hardCapGraceCalls": 10,
15
15
  "contextRotDetection": true,
@@ -187,7 +187,7 @@
187
187
  "handoff": {
188
188
  "enabled": true,
189
189
  "crossTool": true,
190
- "maxParallelAgents": 10,
190
+ "maxParallelAgents": 12,
191
191
  "autonomy": {
192
192
  "askWindow": "plan-only",
193
193
  "planReviewRounds": 2,
@@ -384,7 +384,7 @@
384
384
  "compact": {
385
385
  "enabled": "Bật/tắt toàn bộ helper compact của UKit.",
386
386
  "tokenThreshold": "Ngưỡng token chung cho runtime compact dùng chung.",
387
- "hardCapTokens": "Ngưỡng cứng tuyệt đối (mặc định 160000 token ước lượng). Chạm/vượt ngưỡng này thì context coi như quá dài — không phải gợi ý nữa, là bắt buộc. PHẢI thấp hơn context window thật của model (200k), nếu không API sẽ báo lỗi vượt context trước khi gate kịp chặn.",
387
+ "hardCapTokens": "Ngưỡng cứng tuyệt đối (mặc định 220000 token ước lượng, cỡ cho context window 256k). Chạm/vượt ngưỡng này thì context coi như quá dài — không phải gợi ý nữa, là bắt buộc. PHẢI thấp hơn context window thật của model, nếu không API sẽ báo lỗi vượt context trước khi gate kịp chặn — model 200k thì hạ xuống 160000 hoặc thấp hơn. Đi kèm env.CLAUDE_CODE_AUTO_COMPACT_WINDOW trong .claude/settings.json (mặc định 180000), số đó phải nhỏ hơn ngưỡng này để client tự auto-compact trước khi gate chặn tool.",
388
388
  "hardCapBlock": "Nếu true, hook context-hardcap-gate chặn cứng Edit/Write/Bash (exit 2) khi vượt hardCapTokens, cho tới khi có compact thật (PreCompact) reset lại bộ đếm.",
389
389
  "hardCapGraceCalls": "Số tool call được phép chạy tiếp sau khi vượt hardCapTokens KHI docs/AI_HANDOFF/RUN.md còn run dở (mặc định 10). Dùng để run kịp commit + ghi cursor + push rồi mới bị chặn, thay vì chết giữa lúc đang Edit. Hết grace là chặn cứng như cũ. Budget tính theo mỗi đợt vượt cap, chỉ reset khi ước lượng token thật sự giảm (có compact thật) — không reset theo wave.",
390
390
  "contextRotDetection": "Phát hiện context quá dài/dễ mục để giữ lại state quan trọng trước khi AI nhớ sai.",
@@ -494,7 +494,7 @@
494
494
  "handoff": {
495
495
  "enabled": "Bật Quality Gate cho handoff: plan có Test Plan, executor test-first, reviewer model khác. Tắt = quay về flow cũ (dễ lọt lỗi vặt).",
496
496
  "crossTool": "true nghĩa là handoff truyền qua file (PLAN/INDEX/tasks) chứ không qua in-process subagent — cho phép plan ở Claude Code, execute ở Kilo Code, review ở Claude Code khác model.",
497
- "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 10, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
497
+ "maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 12, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập; 12 là số chọn theo dải an toàn ~10-15 dưới đây, không phải số đo được). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
498
498
  "plan": {
499
499
  "requireTestPlan": "Bắt buộc PLAN.md §4 phải có Test Plan trước khi task chuyển ready.",
500
500
  "minTestsHappyPath": "Tối thiểu test cho happy path.",