@ngockhoale/ukit 2.1.3 → 2.1.5
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +34 -0
- package/manifests/platform.full.yaml +13 -0
- package/package.json +1 -1
- package/scripts/bench/parallel-agents.mjs +153 -0
- package/src/core/runtimeConfig.js +6 -1
- package/src/index/taskRouting.js +9 -2
- package/templates/.claude/agents/code-reviewer.md +43 -2
- package/templates/.claude/agents/handoff-planner.md +2 -2
- package/templates/.claude/commands/ukit/handoff-fullstack.md +3 -3
- package/templates/.claude/commands/ukit/handoff-implement.md +1 -1
- package/templates/.claude/commands/ukit/handoff-review.md +1 -1
- package/templates/.claude/hooks/context-hardcap-gate.sh +2 -1
- package/templates/.claude/hooks/context-window-guard.sh +1 -1
- package/templates/.claude/hooks/skill-router.sh +33 -0
- package/templates/.claude/hooks/stale-spec-guard.sh +32 -0
- package/templates/.claude/settings.json +1 -1
- package/templates/.claude/ukit/index/provision-worktree.mjs +200 -0
- package/templates/.claude/ukit/runtime/compact-threshold.mjs +7 -5
- package/templates/CLAUDE.md +8 -0
- package/templates/docs/AI_HANDOFF/RULES.md +13 -2
- package/templates/ukit/storage/config.json +12 -7
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,40 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.1.5 - 2026-08-19
|
|
6
|
+
|
|
7
|
+
`local-build` was the one execution contract in the daily (non-handoff) flow that let the executor claim done off write evidence alone — and it is also the busiest lane a weaker `unic-code` backend hits. This closes that gap and adds a non-blocking second opinion, without adding wait time to the main task.
|
|
8
|
+
|
|
9
|
+
### Changed
|
|
10
|
+
|
|
11
|
+
- **`local-build` now requires verification evidence, not just write evidence, before claiming done.** `orchestration.contracts['local-build'].completionRule` was `require-write`; every other multi-step contract (`shared-edit`, `find-cause`, `map-impact`) already required verification too, so `local-build` was the odd one out — the exact lane where a weaker executor model producing "silly bugs" that pass a glance but fail a real check would go unnoticed. `completionEvidence` now includes `verification-evidence` alongside `write-evidence` in `src/index/taskRouting.js`, mirrored in `src/core/runtimeConfig.js` and both `config.json` copies. `tests/index/taskRouting.test.js` and `tests/core/runtimeConfig.test.js` lock the new requirement.
|
|
12
|
+
|
|
13
|
+
### Added
|
|
14
|
+
|
|
15
|
+
- **Non-blocking sidecar diff review for `local-build` and `shared-edit`.** After write + verification evidence exist for one of these tasks, routed state now carries a `review=code-reviewer(diff)` segment (`routeSummary.postEditReview`) instructing the agent to launch `code-reviewer` in the background (`run_in_background: true`, `REVIEW_TARGET_TYPE=diff`, model `unic-smart` per the new `subagents.diffReviewModel`) — reusing the reviewer-model-differs-from-executor precedent from `handoff.reviewer`, since a background pass can afford a stronger model without slowing the main task down. Findings are advisory only; the main task's completion is never blocked or reopened on their account. New `code-reviewer.md` mode: `REVIEW_TARGET_TYPE=diff`, documented in `templates/.claude/agents/code-reviewer.md` (mirrored byte-identically to `.claude/agents/code-reviewer.md`) and wired into `templates/CLAUDE.md` / this repo's own `CLAUDE.md` under "Post-Edit Sidecar Review (internal)".
|
|
16
|
+
- **`subagents.diffReviewEnabled` / `diffReviewAgent` / `diffReviewModel`** — new config fields in both `.ukit/storage/config.json` and `templates/ukit/storage/config.json`, defaulting to `true` / `code-reviewer` / `unic-smart`.
|
|
17
|
+
|
|
18
|
+
## 2.1.4 - 2026-08-19
|
|
19
|
+
|
|
20
|
+
More room before compaction, the handoff hooks stop taxing every call they were meant to speed up, and the shipped handoff instructions stop contradicting the config they tell agents to read.
|
|
21
|
+
|
|
22
|
+
### Changed
|
|
23
|
+
|
|
24
|
+
- **`compact.hardCapTokens` 160,000 → 220,000 and `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW` 150,000 → 180,000.** 2.1.3 fixed the deadlock but left the pair sized for a 200k window, which meant compacting far more often than necessary on a 256k model. The new pair leaves roughly 40k of working room after auto-compact fires, and the ordering invariant `autoCompactWindow < hardCapTokens` still holds — `tests/core/autoCompactWindow.test.js` fails if either number moves in isolation. **On a 200k model, lower `hardCapTokens` back to 160,000 or below**; the `_help` text in `config.json` says so, and the gate cannot protect you if the cap sits above your real context window.
|
|
25
|
+
- **Config-missing fallbacks follow the default**: `compact-threshold.mjs` and `context-window-guard.sh` now fall back to `220_000` rather than `160_000`, so a corrupt or absent `config.json` behaves like a fresh install instead of silently reverting to the old cap.
|
|
26
|
+
|
|
27
|
+
### Fixed
|
|
28
|
+
|
|
29
|
+
- **The shipped handoff instructions no longer restate a config value that had gone stale.** Eight lines across `handoff-fullstack.md`, `handoff-implement.md`, `handoff-review.md` and `handoff-planner.md` told agents `handoff.maxParallelAgents` defaults to **10**, while `config.json` in the same release ships **12** — so an installed project got instructions that contradicted its own runtime config. Every one of those lines already said "read from `.ukit/storage/config.json`", so the restated number was redundant as well as wrong; it is now gone and the value has exactly one home. This was the second occurrence (the first was the `3 → 10` bump), and the cause both times was hand-copying a config value into prose. `tests/consistency/configDocsSync.test.js` now fails the suite if any doc restates a numeric `handoff.*` default that contradicts config, in English or Vietnamese, and if a `templates/.claude` file without `{{ }}` variables drifts from its active copy — the "keep the two mirrors identical" rule was written in two places but had never been enforced by anything except a maintainer running `diff` by hand.
|
|
30
|
+
- **The worktree early-exit guard no longer costs more than it saves.** `skill-router.sh` and `stale-spec-guard.sh` skip routing for tool calls scoped to a disposable `.worktrees/task-*` tree — but the first cut spawned `node -e` on every call to decide that, measured at 53ms → 124ms per call in the main tree, where the skip never applies. A pure-bash `case "$INPUT" in *".worktrees/"*)` pre-check now gates the node spawn; it is a strict superset of what the node guard can match, so the main tree pays nothing. Both hooks still fail **open** — any parse error, missing field or node failure falls through to the normal path.
|
|
31
|
+
|
|
32
|
+
### Added
|
|
33
|
+
|
|
34
|
+
- **`.claude/ukit/index/provision-worktree.mjs`** — deterministic worktree provisioning for handoff waves, packaged via `manifests/platform.full.yaml`. It hides the worktree's `node_modules` symlink through the worktree's own `.git/info/exclude` (a symlink never matched `.gitignore`'s `node_modules/`, so an orchestrator `git add -A` would have committed a machine-local absolute path), and skips git-tracked files when copying `.claude/` so it cannot clobber `.claude/commands/ukit/handoff-fullstack.md`. **Shipped but not yet wired into `handoff-implement`** — the wiring is queued for a later release.
|
|
35
|
+
- **Test-selection policy in `docs/AI_HANDOFF/RULES.md`** — per-task related tests are the floor, and a full suite at the wave/cycle boundary is mandatory, not optional. This is the compensating net that makes narrowed per-task scope safe; the agent definitions already cited it before it existed.
|
|
36
|
+
- **`scripts/bench/parallel-agents.mjs`** — harness for measuring concurrent test-process contention at several parallelism levels, recording `slowdownFactor` / `perRunCost` / a `recommended` value plus the load averages the run happened under. It measures test-process contention, not LLM agents.
|
|
37
|
+
- **`handoff.maxParallelAgents`: 10 → 12.** Chosen, not measured — the benchmark above was never run on a quiet enough machine, so this is the middle of the ~10-15 safe band the config's own `_help` has always documented. Lower it to 3-5 if your tasks produce long verification output or you see compaction firing repeatedly.
|
|
38
|
+
|
|
5
39
|
## 2.1.3 - 2026-08-16
|
|
6
40
|
|
|
7
41
|
The reason auto-compact never fired in the terminal was UKit itself. This makes it fire.
|
|
@@ -1238,6 +1238,19 @@ items:
|
|
|
1238
1238
|
packs:
|
|
1239
1239
|
- core
|
|
1240
1240
|
|
|
1241
|
+
- id: ukit-index-provision-worktree-script
|
|
1242
|
+
type: config
|
|
1243
|
+
sourceTemplate: .claude/ukit/index/provision-worktree.mjs
|
|
1244
|
+
targetPath: .claude/ukit/index/provision-worktree.mjs
|
|
1245
|
+
requires:
|
|
1246
|
+
- ukit-runtime-text-profile-script
|
|
1247
|
+
- ukit-runtime-safe-patch-core-script
|
|
1248
|
+
mergeStrategy: overwrite_with_backup
|
|
1249
|
+
variables: []
|
|
1250
|
+
enabledByDefault: true
|
|
1251
|
+
packs:
|
|
1252
|
+
- core
|
|
1253
|
+
|
|
1241
1254
|
- id: ukit-index-stale-spec-check-script
|
|
1242
1255
|
type: config
|
|
1243
1256
|
sourceTemplate: .claude/ukit/index/stale-spec-check.mjs
|
package/package.json
CHANGED
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* parallel-agents.mjs — TASK-004 harness.
|
|
4
|
+
*
|
|
5
|
+
* Measures the *machine-contention* component of N parallel agents by spawning N concurrent
|
|
6
|
+
* child processes (default `yarn test:release-core`) per level and recording wall-clock.
|
|
7
|
+
* This deliberately does NOT claim to benchmark LLM agent wall-clock (token counts, server
|
|
8
|
+
* load and retries make that irreproducible) — see TASK-004 planner note #1.
|
|
9
|
+
*
|
|
10
|
+
* Usage:
|
|
11
|
+
* node scripts/bench/parallel-agents.mjs [--levels 1,3,5,10] [--cmd "<test command>"] [--out <path>]
|
|
12
|
+
*/
|
|
13
|
+
|
|
14
|
+
import fs from 'node:fs';
|
|
15
|
+
import os from 'node:os';
|
|
16
|
+
import path from 'node:path';
|
|
17
|
+
import { spawn } from 'node:child_process';
|
|
18
|
+
|
|
19
|
+
const SLOWDOWN_LIMIT = 2.0;
|
|
20
|
+
|
|
21
|
+
function parseArgs(argv) {
|
|
22
|
+
const opts = { levels: '1,3,5,10', cmd: 'yarn test:release-core', out: '.cache/bench/parallel-agents.json' };
|
|
23
|
+
for (let i = 0; i < argv.length; i += 2) {
|
|
24
|
+
const key = argv[i];
|
|
25
|
+
const val = argv[i + 1];
|
|
26
|
+
if ((key === '--levels' || key === '--cmd' || key === '--out') && val !== undefined) {
|
|
27
|
+
opts[key.slice(2)] = val;
|
|
28
|
+
}
|
|
29
|
+
}
|
|
30
|
+
return opts;
|
|
31
|
+
}
|
|
32
|
+
|
|
33
|
+
function parseLevels(raw) {
|
|
34
|
+
if (!/^\d+(,\d+)*$/.test(raw.trim())) {
|
|
35
|
+
throw new Error(`unparseable --levels value: "${raw}" (expected comma-separated positive integers, e.g. 1,3,5,10)`);
|
|
36
|
+
}
|
|
37
|
+
const levels = raw.split(',').map((s) => parseInt(s, 10));
|
|
38
|
+
if (levels.some((n) => n < 1)) throw new Error(`--levels values must be >= 1, got: ${raw}`);
|
|
39
|
+
return levels;
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
function runLevel(n, cmd) {
|
|
43
|
+
return new Promise((resolve) => {
|
|
44
|
+
const t0 = Date.now();
|
|
45
|
+
let live = 0;
|
|
46
|
+
let highWater = 0;
|
|
47
|
+
let failures = 0;
|
|
48
|
+
let settled = 0;
|
|
49
|
+
const children = [];
|
|
50
|
+
for (let i = 0; i < n; i += 1) {
|
|
51
|
+
const child = spawn(cmd, { shell: true, stdio: 'ignore' });
|
|
52
|
+
live += 1;
|
|
53
|
+
highWater = Math.max(highWater, live);
|
|
54
|
+
children.push(child);
|
|
55
|
+
child.on('error', () => {
|
|
56
|
+
live -= 1;
|
|
57
|
+
settled += 1;
|
|
58
|
+
failures += 1;
|
|
59
|
+
maybeFinish();
|
|
60
|
+
});
|
|
61
|
+
child.on('exit', (code) => {
|
|
62
|
+
live -= 1;
|
|
63
|
+
settled += 1;
|
|
64
|
+
if (code !== 0) failures += 1;
|
|
65
|
+
maybeFinish();
|
|
66
|
+
});
|
|
67
|
+
}
|
|
68
|
+
function maybeFinish() {
|
|
69
|
+
if (settled === n) {
|
|
70
|
+
resolve({ n, wallClockMs: Date.now() - t0, failures, concurrencyHighWaterMark: highWater });
|
|
71
|
+
}
|
|
72
|
+
}
|
|
73
|
+
});
|
|
74
|
+
}
|
|
75
|
+
|
|
76
|
+
function fail(msg) {
|
|
77
|
+
process.stderr.write(`${msg}\n`);
|
|
78
|
+
process.exit(1);
|
|
79
|
+
}
|
|
80
|
+
|
|
81
|
+
async function main() {
|
|
82
|
+
const opts = parseArgs(process.argv.slice(2));
|
|
83
|
+
let levels;
|
|
84
|
+
try {
|
|
85
|
+
levels = parseLevels(opts.levels);
|
|
86
|
+
} catch (err) {
|
|
87
|
+
fail(err.message);
|
|
88
|
+
}
|
|
89
|
+
|
|
90
|
+
const cpuCount = os.cpus().length;
|
|
91
|
+
const loadAvgStart = os.loadavg();
|
|
92
|
+
const startedAt = new Date().toISOString();
|
|
93
|
+
|
|
94
|
+
const rows = [];
|
|
95
|
+
for (const n of levels) {
|
|
96
|
+
// eslint-disable-next-line no-await-in-loop -- levels are measured sequentially on purpose
|
|
97
|
+
rows.push(await runLevel(n, opts.cmd));
|
|
98
|
+
}
|
|
99
|
+
const base = rows[0].wallClockMs;
|
|
100
|
+
for (const row of rows) {
|
|
101
|
+
row.slowdownFactor = row.wallClockMs / base;
|
|
102
|
+
row.perRunCost = row.wallClockMs / row.n;
|
|
103
|
+
}
|
|
104
|
+
|
|
105
|
+
const eligible = rows.filter((r) => r.slowdownFactor <= SLOWDOWN_LIMIT);
|
|
106
|
+
const recommended = eligible.reduce((best, r) => (r.perRunCost < best.perRunCost ? r : best), eligible[0]).n;
|
|
107
|
+
|
|
108
|
+
const loadAvgEnd = os.loadavg();
|
|
109
|
+
const finishedAt = new Date().toISOString();
|
|
110
|
+
|
|
111
|
+
const result = {
|
|
112
|
+
cpuCount,
|
|
113
|
+
levels: rows.map(({ n, wallClockMs, slowdownFactor, perRunCost, failures, concurrencyHighWaterMark }) => ({
|
|
114
|
+
n, wallClockMs, slowdownFactor, perRunCost, failures, concurrencyHighWaterMark,
|
|
115
|
+
})),
|
|
116
|
+
recommended,
|
|
117
|
+
measurementConditions: { loadAvgStart, loadAvgEnd, startedAt, finishedAt, cpuCount },
|
|
118
|
+
};
|
|
119
|
+
|
|
120
|
+
// Print the human table first so a failed --out write still shows the numbers.
|
|
121
|
+
const header = 'n'.padStart(4) + ' ' + 'wallClockMs'.padStart(12) + ' ' + 'slowdownFactor'.padStart(15) + ' ' + 'perRunCost'.padStart(11) + ' ' + 'failures'.padStart(8) + ' ' + 'concurrencyHighWaterMark'.padStart(23);
|
|
122
|
+
console.log(header);
|
|
123
|
+
for (const r of result.levels) {
|
|
124
|
+
console.log(
|
|
125
|
+
String(r.n).padStart(4)
|
|
126
|
+
+ ' ' + String(r.wallClockMs).padStart(12)
|
|
127
|
+
+ ' ' + r.slowdownFactor.toFixed(2).padStart(15)
|
|
128
|
+
+ ' ' + r.perRunCost.toFixed(1).padStart(11)
|
|
129
|
+
+ ' ' + String(r.failures).padStart(8)
|
|
130
|
+
+ ' ' + String(r.concurrencyHighWaterMark).padStart(23),
|
|
131
|
+
);
|
|
132
|
+
}
|
|
133
|
+
console.log(`recommended: ${recommended} (best perRunCost with slowdownFactor <= ${SLOWDOWN_LIMIT}) on ${cpuCount} CPUs`);
|
|
134
|
+
console.log(`loadAvgStart: [${loadAvgStart.map((v) => v.toFixed(2)).join(', ')}] loadAvgEnd: [${loadAvgEnd.map((v) => v.toFixed(2)).join(', ')}]`);
|
|
135
|
+
|
|
136
|
+
// Overwrite (never append), and never leave a partial file behind.
|
|
137
|
+
const outPath = path.resolve(opts.out);
|
|
138
|
+
try {
|
|
139
|
+
fs.mkdirSync(path.dirname(outPath), { recursive: true });
|
|
140
|
+
const body = JSON.stringify(result, null, 2) + '\n';
|
|
141
|
+
fs.writeFileSync(outPath, body);
|
|
142
|
+
} catch (err) {
|
|
143
|
+
fail(`cannot write --out ${outPath}: ${err.message}`);
|
|
144
|
+
}
|
|
145
|
+
|
|
146
|
+
const totalFailures = result.levels.reduce((s, r) => s + r.failures, 0);
|
|
147
|
+
if (totalFailures > 0) {
|
|
148
|
+
const perLevel = result.levels.filter((r) => r.failures > 0).map((r) => `n=${r.n}: ${r.failures}/${r.n} failed`).join('; ');
|
|
149
|
+
fail(`bench command failed: ${perLevel}`);
|
|
150
|
+
}
|
|
151
|
+
}
|
|
152
|
+
|
|
153
|
+
main().catch((err) => fail(err && err.message ? err.message : String(err)));
|
|
@@ -145,8 +145,9 @@ export function buildDefaultRuntimeConfig(overrides = {}) {
|
|
|
145
145
|
maxReadPasses: 2,
|
|
146
146
|
maxContextPulls: 1,
|
|
147
147
|
verificationPolicy: 'targeted-if-covered',
|
|
148
|
-
completionRule: 'require-write',
|
|
148
|
+
completionRule: 'require-write-and-verification',
|
|
149
149
|
delegationPolicy: 'disallow-by-default',
|
|
150
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
150
151
|
},
|
|
151
152
|
'find-cause': {
|
|
152
153
|
maxReadPassesBeforeReassess: 3,
|
|
@@ -160,6 +161,7 @@ export function buildDefaultRuntimeConfig(overrides = {}) {
|
|
|
160
161
|
verificationPolicy: 'targeted-then-widen-on-risk',
|
|
161
162
|
completionRule: 'require-write-and-verification',
|
|
162
163
|
delegationPolicy: 'allow-qualified-sidecar',
|
|
164
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
163
165
|
},
|
|
164
166
|
'map-impact': {
|
|
165
167
|
maxReadPasses: 3,
|
|
@@ -224,6 +226,9 @@ export function buildDefaultRuntimeConfig(overrides = {}) {
|
|
|
224
226
|
enabled: true,
|
|
225
227
|
smallTaskModel: 'unic-lite',
|
|
226
228
|
smallTaskAgent: 'ukit-small-task-maintainer',
|
|
229
|
+
diffReviewEnabled: true,
|
|
230
|
+
diffReviewAgent: 'code-reviewer',
|
|
231
|
+
diffReviewModel: 'unic-smart',
|
|
227
232
|
visionEnabled: true,
|
|
228
233
|
visionModel: 'unic-vision',
|
|
229
234
|
visionAgent: 'ukit-vision-analyst',
|
package/src/index/taskRouting.js
CHANGED
|
@@ -235,6 +235,9 @@ export function buildRouteSummary({
|
|
|
235
235
|
executionCandidates,
|
|
236
236
|
});
|
|
237
237
|
const executionContract = buildExecutionContract(executionMode);
|
|
238
|
+
const postEditReview = executionContract?.postEditReviewPolicy
|
|
239
|
+
? { policy: executionContract.postEditReviewPolicy, agent: 'code-reviewer', reviewTargetType: 'diff' }
|
|
240
|
+
: null;
|
|
238
241
|
const completionState = buildCompletionState({
|
|
239
242
|
executionMode,
|
|
240
243
|
verificationRecommendation,
|
|
@@ -263,6 +266,7 @@ export function buildRouteSummary({
|
|
|
263
266
|
editGuardHint ? `editGuard=${editGuardHint}` : null,
|
|
264
267
|
delegationRecommendation?.hint ? `delegate=${delegationRecommendation.hint}` : null,
|
|
265
268
|
policyMode ? `policy=${policyMode}` : null,
|
|
269
|
+
postEditReview ? `review=${postEditReview.agent}(${postEditReview.reviewTargetType})` : null,
|
|
266
270
|
handoffBudget?.warning ? `budget=${handoffBudget.warning}` : null,
|
|
267
271
|
worklogBudget?.warning ? `budget=${worklogBudget.warning}` : null,
|
|
268
272
|
].filter(Boolean).join(' | ');
|
|
@@ -279,6 +283,7 @@ export function buildRouteSummary({
|
|
|
279
283
|
approachSelector,
|
|
280
284
|
executionContract,
|
|
281
285
|
completionState,
|
|
286
|
+
postEditReview,
|
|
282
287
|
continuationState,
|
|
283
288
|
intentMode: routingContext.intentMode ?? null,
|
|
284
289
|
handoffFile,
|
|
@@ -580,9 +585,10 @@ function buildExecutionContract(executionMode = null) {
|
|
|
580
585
|
maxReadPasses: 2,
|
|
581
586
|
maxContextPulls: 1,
|
|
582
587
|
verificationPolicy: 'targeted-if-covered',
|
|
583
|
-
completionRule: 'require-write',
|
|
588
|
+
completionRule: 'require-write-and-verification',
|
|
584
589
|
delegationPolicy: 'disallow-by-default',
|
|
585
|
-
completionEvidence: ['write-evidence'],
|
|
590
|
+
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
591
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
586
592
|
},
|
|
587
593
|
'find-cause': {
|
|
588
594
|
maxReadPassesBeforeReassess: 3,
|
|
@@ -599,6 +605,7 @@ function buildExecutionContract(executionMode = null) {
|
|
|
599
605
|
delegationPolicy: 'allow-qualified-sidecar',
|
|
600
606
|
completionEvidence: ['write-evidence', 'verification-evidence'],
|
|
601
607
|
mirrorConsistencyRequired: true,
|
|
608
|
+
postEditReviewPolicy: 'sidecar-non-blocking',
|
|
602
609
|
},
|
|
603
610
|
'map-impact': {
|
|
604
611
|
maxReadPasses: 3,
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: code-reviewer
|
|
3
|
-
description: "Independent reviewer for handoff Phase 3,
|
|
3
|
+
description: "Independent reviewer for handoff Phase 3, for spec/plan documents, and for non-blocking sidecar diff review of daily-flow edits. For code (default): use after executor reports STATUS: DONE on a handoff task, MUST run with a model different from the executor (configured in .ukit/storage/config.json → handoff.reviewer.model, default unic-smart), produces a verdict: APPROVED | APPROVED-WITH-MINOR | CHANGES-REQUESTED | CRITICAL. For spec/plan documents (set REVIEW_TARGET_TYPE=spec or plan): reviews a docs/plans/*.md file for completeness/consistency/clarity/scope/YAGNI, produces Status: Approved | Issues Found. For sidecar diff review (set REVIEW_TARGET_TYPE=diff): reviews the current uncommitted git diff after a local-build/shared-edit task, runs in the background and never blocks the main task, produces STATUS: clean | issues-found."
|
|
4
4
|
model: opus # unic-smart
|
|
5
5
|
color: yellow
|
|
6
6
|
tools: ["Read", "Grep", "Glob", "Bash"]
|
|
@@ -14,6 +14,7 @@ You are the independent reviewer for UKit's handoff Quality Gate. Your model is
|
|
|
14
14
|
|
|
15
15
|
- `code` (default, if not specified) — reviewing a handoff task diff. Follow **Code Review** below, unchanged.
|
|
16
16
|
- `spec` | `plan` — reviewing a document (e.g. `docs/plans/*.md`), no diff/task file/executor report involved. Skip straight to **Spec/Plan Review** at the end of this file instead.
|
|
17
|
+
- `diff` — non-blocking sidecar review of the current uncommitted diff in the daily (non-handoff) flow. No task file/executor report/model-isolation check involved. Skip straight to **Sidecar Diff Review** at the end of this file instead.
|
|
17
18
|
|
|
18
19
|
## Code Review (REVIEW_TARGET_TYPE=code)
|
|
19
20
|
|
|
@@ -30,7 +31,7 @@ If any input is missing, return `CHANGES-REQUESTED` with reason "incomplete hand
|
|
|
30
31
|
0. **Verification package completeness** — Check whether the project has a lint or typecheck script (`package.json` scripts, or the stack's equivalent). If it does and the task's Verification Commands don't run it, that is `CHANGES-REQUESTED`: "verification commands missing lint/typecheck — re-run planner or add the command and re-verify" — do this before anything else below.
|
|
31
32
|
1. **Test Plan adherence** — Were all tests in §4 actually implemented, including the ≥2 edge cases required by `handoff.plan.minTestsEdgeCase`? Check the Executor Report's `RED_OUTPUT` field: it must contain actual failing-test output (assertion failure, stack trace, non-zero exit), not a bare claim like "confirmed" or "yes". Missing or vague `RED_OUTPUT` → `CHANGES-REQUESTED`: "no evidence tests were RED before implementation — re-run TDD cycle and paste real output". Then run the tests yourself: `<task Verification Commands>`. Fresh PASS required, no trusting executor's output blindly.
|
|
32
33
|
2. **Correctness** — Does the diff implement the requested behavior? Any obvious wrong assumptions, stale refs, missing cases?
|
|
33
|
-
3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider
|
|
34
|
+
3. **Regression risk** — What existing behavior could this break? Are shared paths/tests/contracts still aligned? Run the wider suite only when the diff touched shared code; otherwise the task's own targeted commands are the gate and the wave-boundary full `yarn test` (see docs/AI_HANDOFF/RULES.md "Test selection") is the regression net.
|
|
34
35
|
4. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
35
36
|
5. **Performance / scale** — Accidental N+1, repeated I/O, large scans inside hot paths.
|
|
36
37
|
6. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
@@ -155,3 +156,43 @@ NOTES: [1-2 sentences if needed]
|
|
|
155
156
|
```
|
|
156
157
|
|
|
157
158
|
`<N>` = 1 + however many `### Round` entries already exist in the log (1 if this is the first review).
|
|
159
|
+
|
|
160
|
+
## Sidecar Diff Review (REVIEW_TARGET_TYPE=diff)
|
|
161
|
+
|
|
162
|
+
This mode exists so a weaker daily-flow executor model still gets a second pair of eyes,
|
|
163
|
+
without adding wait time to the main task. You are launched in the background right after
|
|
164
|
+
the main task already has write + verification evidence; the caller is not waiting on you.
|
|
165
|
+
|
|
166
|
+
### Inputs you expect
|
|
167
|
+
|
|
168
|
+
- No task file, no executor report, no model-isolation check. Just read the current uncommitted
|
|
169
|
+
diff yourself: `git diff` (and `git diff --stat` for an overview). If there is no diff, report
|
|
170
|
+
`STATUS: clean` with `FINDINGS: none` and stop.
|
|
171
|
+
|
|
172
|
+
### Review order
|
|
173
|
+
|
|
174
|
+
Apply the same lenses as Code Review's steps 2-6, scoped to what the diff actually touches:
|
|
175
|
+
|
|
176
|
+
1. **Correctness** — Does the diff do what it looks like it's trying to do? Wrong assumptions, stale refs, missing cases?
|
|
177
|
+
2. **Regression risk** — Any existing behavior/tests/contracts this plausibly breaks?
|
|
178
|
+
3. **Safety / security / data loss** — Destructive actions, auth/permission, path handling, unsafe shell/DB/file ops.
|
|
179
|
+
4. **Performance / scale** — Accidental N+1, repeated I/O, large scans in hot paths.
|
|
180
|
+
5. **Maintainability** — Duplicated logic, dead branches, misleading naming, drift between docs/tests/source.
|
|
181
|
+
|
|
182
|
+
Do not re-run the project's full verification suite here — this is an advisory pass, not a gate.
|
|
183
|
+
You may read files for context but this mode never edits anything.
|
|
184
|
+
|
|
185
|
+
### Output
|
|
186
|
+
|
|
187
|
+
Keep it short — this is a quick advisory pass, not a full verdict:
|
|
188
|
+
|
|
189
|
+
```
|
|
190
|
+
STATUS: clean | issues-found
|
|
191
|
+
FINDINGS:
|
|
192
|
+
- file:line — what's wrong, why it matters
|
|
193
|
+
NOTES: [advisory only, non-blocking — 1 sentence if needed]
|
|
194
|
+
```
|
|
195
|
+
|
|
196
|
+
There is no task file or INDEX.md to update in this mode. Findings are advisory only: the main
|
|
197
|
+
task is not blocked on this review and may already be reported done by the time you finish.
|
|
198
|
+
Report back to the caller in a few lines; do not paste the full diff.
|
|
@@ -105,7 +105,7 @@ Use `_TEMPLATE.md` structure (from pre-read context or file).
|
|
|
105
105
|
| Dependencies | `TASK-xxx` or `none` — wave order is inferred from this |
|
|
106
106
|
| Test Cases | Type \| Test Name \| Expected — ≥1 happy + ≥2 edge cases of different kinds |
|
|
107
107
|
| Test Files | Exact test file paths to create/modify |
|
|
108
|
-
| Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above) |
|
|
108
|
+
| Verification Commands | Runnable shell commands — MUST include the project's lint/typecheck command if one exists (see §5 rule above). Apply the "Test selection" resolution order from docs/AI_HANDOFF/RULES.md: `src/`/`scripts/` targets → the `tests` array in `.cache/index/tests-map.json`; `templates/.claude/**`/`.claude/**` targets → path convention (hooks → `tests/hooks/` + `tests/handoff/cycle*/`; manifest/settings → `tests/manifest/`; runtime `.mjs` mirrors → `tests/core/*Parity*` + `tests/index/`); if both resolve to fewer than one test file the task MUST fall back to `yarn test:release-core` — never the full suite by default, never an empty selection |
|
|
109
109
|
| Acceptance Criteria | Verifiable checklist |
|
|
110
110
|
|
|
111
111
|
Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
|
|
@@ -120,7 +120,7 @@ Missing any field → `needs_breakdown`. Never mark incomplete tasks `ready`.
|
|
|
120
120
|
|
|
121
121
|
Wave width is the single biggest lever on how long a cycle takes: a wave of 6 finishes in
|
|
122
122
|
roughly the time of its slowest task, while a chain of 6 takes six times that. Executors run
|
|
123
|
-
up to `handoff.maxParallelAgents`
|
|
123
|
+
up to `handoff.maxParallelAgents` at once, so a plan that produces `none`
|
|
124
124
|
dependencies for most tasks is dramatically faster than one that produces a chain.
|
|
125
125
|
|
|
126
126
|
**Write `Dependencies: none` unless B genuinely cannot be written without A's output.** A real
|
|
@@ -223,7 +223,7 @@ must not share a file) in case it slipped through review. Only mark a task `need
|
|
|
223
223
|
if it is missing required fields — never merely for sharing a file.
|
|
224
224
|
|
|
225
225
|
**Batch each wave — mandatory.** Read `handoff.maxParallelAgents` from
|
|
226
|
-
`.ukit/storage/config.json
|
|
226
|
+
`.ukit/storage/config.json`. A wave with more tasks than that is split
|
|
227
227
|
into consecutive batches of at most that many; finish one batch completely (including 3c
|
|
228
228
|
copy-back and worktree deletion) before starting the next.
|
|
229
229
|
|
|
@@ -234,7 +234,7 @@ leaves worktrees behind. Batching only ever narrows a wave, never reorders acros
|
|
|
234
234
|
|
|
235
235
|
### I3 — Execute wave by wave (code model agents)
|
|
236
236
|
|
|
237
|
-
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents`
|
|
237
|
+
**Claude Code — MANDATORY, do this before anything else in I3:** for each wave, call the Agent tool once per task **in the current batch** (in parallel, at most `handoff.maxParallelAgents`), each with `subagent_type: "feature-implementer"`. Do NOT implement the tasks yourself in the current session — this step is contracted to the code tier (sonnet/unic-code), which only the spawned agent's frontmatter model guarantees.
|
|
238
238
|
|
|
239
239
|
For each wave:
|
|
240
240
|
|
|
@@ -398,7 +398,7 @@ If that diff is empty → implement was not completed. Do not stop: re-enter Pha
|
|
|
398
398
|
|
|
399
399
|
### R2 — Model isolation check (strong model, always first)
|
|
400
400
|
|
|
401
|
-
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json
|
|
401
|
+
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (R2–R4, appended to each task file) before starting the next.
|
|
402
402
|
|
|
403
403
|
**Claude Code — MANDATORY, do this before anything else in R2–R4:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees.
|
|
404
404
|
|
|
@@ -94,7 +94,7 @@ when it is missing required fields — never merely for sharing a file.
|
|
|
94
94
|
|
|
95
95
|
### Batch each wave — mandatory
|
|
96
96
|
|
|
97
|
-
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json
|
|
97
|
+
Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. A wave
|
|
98
98
|
with more tasks than that is split into consecutive batches of at most that many; finish
|
|
99
99
|
one batch completely (including 3c copy-back and worktree deletion) before starting the
|
|
100
100
|
next.
|
|
@@ -50,7 +50,7 @@ If that diff is empty → handoff-implement was not completed. Report which task
|
|
|
50
50
|
|
|
51
51
|
## Step 2 — Review the diff
|
|
52
52
|
|
|
53
|
-
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json
|
|
53
|
+
**Batch the review set — mandatory.** Read `handoff.maxParallelAgents` from `.ukit/storage/config.json`. If more `pending_review` tasks exist than that, split into consecutive batches of at most that many; finish one batch's verdicts (2a–2d, appended to each task file) before starting the next.
|
|
54
54
|
|
|
55
55
|
**Claude Code — MANDATORY, do this before anything else:** for each batch, call the Agent tool once per `pending_review` task **in parallel** (at most `handoff.maxParallelAgents`), each with `subagent_type: "code-reviewer"`. Reviewer agents only read the diff and append a verdict to their own task file — no worktree, no shared write target — so running them in parallel carries none of Phase 3's file-conflict risk. Do NOT review the diff yourself in the current session — this step is contracted to the strong tier (opus/unic-smart) and MUST differ from the executor's model, which only the spawned agent's frontmatter model guarantees. Pass each agent: the task file path, the executor's report, and the diff.
|
|
56
56
|
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
#!/bin/bash
|
|
2
2
|
# PreToolUse hook: hard-enforce an absolute context token cap (compact.hardCapTokens,
|
|
3
|
-
# default
|
|
3
|
+
# default 220000, sized for a 256k window — must stay below the model's real context
|
|
4
|
+
# window, so lower it on a 200k model), separate from the
|
|
4
5
|
# soft/hard advisory pressure phases in
|
|
5
6
|
# compact-threshold.mjs (default soft=50000/hard=80000, which only print a suggestion).
|
|
6
7
|
#
|
|
@@ -4,6 +4,39 @@
|
|
|
4
4
|
|
|
5
5
|
INPUT=$(cat)
|
|
6
6
|
PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
|
|
7
|
+
|
|
8
|
+
# Worktree early-exit: when the tool call is scoped to a disposable .worktrees/task-*
|
|
9
|
+
# tree, there is nothing to route — exit before spawning the router. Reads
|
|
10
|
+
# tool_input.file_path / tool_input.path / tool_input.command (never tool_input.pattern).
|
|
11
|
+
# Skips only when ".worktrees/" starts a path component. Fails OPEN: any parse
|
|
12
|
+
# problem, missing field, or node error falls through to the normal hook path below.
|
|
13
|
+
# Pure-bash pre-check first: the node guard can only ever match when the literal
|
|
14
|
+
# ".worktrees/" substring is present, so this case is a strict superset and avoids
|
|
15
|
+
# paying the node spawn cost on the far more common non-worktree call.
|
|
16
|
+
case "$INPUT" in
|
|
17
|
+
*".worktrees/"*)
|
|
18
|
+
if printf '%s' "$INPUT" | node -e '
|
|
19
|
+
const chunks = [];
|
|
20
|
+
process.stdin.on("data", (c) => chunks.push(c));
|
|
21
|
+
process.stdin.on("end", () => {
|
|
22
|
+
try {
|
|
23
|
+
const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
|
|
24
|
+
const toolInput = (payload && typeof payload === "object" && payload.tool_input
|
|
25
|
+
&& typeof payload.tool_input === "object") ? payload.tool_input : {};
|
|
26
|
+
const candidates = [toolInput.file_path, toolInput.path, toolInput.command];
|
|
27
|
+
const QMARKS = String.fromCharCode(34, 39, 96);
|
|
28
|
+
const worktreeComponent = new RegExp("(^|[\\/\\s=" + QMARKS + "])\\.worktrees/");
|
|
29
|
+
process.exit(candidates.some((v) => typeof v === "string" && worktreeComponent.test(v)) ? 0 : 1);
|
|
30
|
+
} catch {
|
|
31
|
+
process.exit(1);
|
|
32
|
+
}
|
|
33
|
+
});
|
|
34
|
+
' >/dev/null 2>&1; then
|
|
35
|
+
exit 0
|
|
36
|
+
fi
|
|
37
|
+
;;
|
|
38
|
+
esac
|
|
39
|
+
|
|
7
40
|
STATE_FILE="$PROJECT_ROOT/.claude/ukit/skill-router-state.json"
|
|
8
41
|
HOOK_DIR="$(cd "$(dirname "$0")" && pwd)"
|
|
9
42
|
THRESHOLD_SCRIPT="$HOOK_DIR/../ukit/runtime/compact-threshold.mjs"
|
|
@@ -10,4 +10,36 @@ if [ ! -f "$SCRIPT" ]; then
|
|
|
10
10
|
exit 0
|
|
11
11
|
fi
|
|
12
12
|
|
|
13
|
+
# Worktree early-exit: when the tool call is scoped to a disposable .worktrees/task-*
|
|
14
|
+
# tree, skip the stale-spec check entirely. Reads tool_input.file_path /
|
|
15
|
+
# tool_input.path / tool_input.command (never tool_input.pattern). Skips only when
|
|
16
|
+
# ".worktrees/" starts a path component. Fails OPEN: any parse problem, missing field,
|
|
17
|
+
# or node error falls through to the normal check below.
|
|
18
|
+
# Pure-bash pre-check first: the node guard can only ever match when the literal
|
|
19
|
+
# ".worktrees/" substring is present, so this case is a strict superset and avoids
|
|
20
|
+
# paying the node spawn cost on the far more common non-worktree call.
|
|
21
|
+
case "$INPUT" in
|
|
22
|
+
*".worktrees/"*)
|
|
23
|
+
if printf '%s' "$INPUT" | node -e '
|
|
24
|
+
const chunks = [];
|
|
25
|
+
process.stdin.on("data", (c) => chunks.push(c));
|
|
26
|
+
process.stdin.on("end", () => {
|
|
27
|
+
try {
|
|
28
|
+
const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
|
|
29
|
+
const toolInput = (payload && typeof payload === "object" && payload.tool_input
|
|
30
|
+
&& typeof payload.tool_input === "object") ? payload.tool_input : {};
|
|
31
|
+
const candidates = [toolInput.file_path, toolInput.path, toolInput.command];
|
|
32
|
+
const QMARKS = String.fromCharCode(34, 39, 96);
|
|
33
|
+
const worktreeComponent = new RegExp("(^|[\\/\\s=" + QMARKS + "])\\.worktrees/");
|
|
34
|
+
process.exit(candidates.some((v) => typeof v === "string" && worktreeComponent.test(v)) ? 0 : 1);
|
|
35
|
+
} catch {
|
|
36
|
+
process.exit(1);
|
|
37
|
+
}
|
|
38
|
+
});
|
|
39
|
+
' >/dev/null 2>&1; then
|
|
40
|
+
exit 0
|
|
41
|
+
fi
|
|
42
|
+
;;
|
|
43
|
+
esac
|
|
44
|
+
|
|
13
45
|
printf '%s' "$INPUT" | node "$SCRIPT"
|
|
@@ -0,0 +1,200 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
/**
|
|
3
|
+
* provision-worktree.mjs — deterministic worktree provisioner (TASK-002).
|
|
4
|
+
*
|
|
5
|
+
* Provisions what `git worktree add` cannot bring across (all four are gitignored):
|
|
6
|
+
* - node_modules -> SYMLINK to the main tree (28M; avoids N concurrent installs)
|
|
7
|
+
* - .claude/ -> REAL COPY preserving file modes (never a symlink: a worktree
|
|
8
|
+
* edit must not write through into the live main-tree mirror).
|
|
9
|
+
* Mode preservation matters: every .claude/hooks/*.sh is 755 and
|
|
10
|
+
* tests/handoff/cycle4/vision-gate.test.mjs asserts the exec bit.
|
|
11
|
+
* - .ukit/storage/config.json -> copy
|
|
12
|
+
* - .cache/index/ -> copy
|
|
13
|
+
*
|
|
14
|
+
* CLI: node .claude/ukit/index/provision-worktree.mjs <worktree-path> [--main-root <path>]
|
|
15
|
+
* exit 0 on success (including a no-op second run); non-zero + one-line stderr otherwise.
|
|
16
|
+
*/
|
|
17
|
+
|
|
18
|
+
import fs from 'node:fs';
|
|
19
|
+
import path from 'node:path';
|
|
20
|
+
import { spawnSync } from 'node:child_process';
|
|
21
|
+
|
|
22
|
+
function fail(msg) {
|
|
23
|
+
process.stderr.write(`provision-worktree: ${msg}\n`);
|
|
24
|
+
process.exit(1);
|
|
25
|
+
}
|
|
26
|
+
|
|
27
|
+
const args = process.argv.slice(2);
|
|
28
|
+
let wtPath = null;
|
|
29
|
+
let mainRoot = null;
|
|
30
|
+
for (let i = 0; i < args.length; i += 1) {
|
|
31
|
+
if (args[i] === '--main-root') {
|
|
32
|
+
if (i + 1 >= args.length) fail('--main-root requires a value');
|
|
33
|
+
mainRoot = args[(i += 1)];
|
|
34
|
+
} else if (wtPath === null) {
|
|
35
|
+
wtPath = args[i];
|
|
36
|
+
} else {
|
|
37
|
+
fail(`unexpected argument: ${args[i]}`);
|
|
38
|
+
}
|
|
39
|
+
}
|
|
40
|
+
if (!wtPath) fail('usage: provision-worktree.mjs <worktree-path> [--main-root <path>]');
|
|
41
|
+
wtPath = path.resolve(wtPath);
|
|
42
|
+
mainRoot = path.resolve(mainRoot ?? process.cwd());
|
|
43
|
+
|
|
44
|
+
// --- validate before mutating anything -------------------------------------
|
|
45
|
+
|
|
46
|
+
if (!fs.existsSync(wtPath) || !fs.statSync(wtPath).isDirectory()) {
|
|
47
|
+
fail(`worktree path is not an existing directory: ${wtPath}`);
|
|
48
|
+
}
|
|
49
|
+
try {
|
|
50
|
+
fs.accessSync(wtPath, fs.constants.W_OK);
|
|
51
|
+
} catch {
|
|
52
|
+
fail(`worktree dir is not writable: ${wtPath}`);
|
|
53
|
+
}
|
|
54
|
+
|
|
55
|
+
const mainModules = path.join(mainRoot, 'node_modules');
|
|
56
|
+
if (!fs.existsSync(mainModules)) {
|
|
57
|
+
fail(`main tree has no node_modules at ${mainModules}; run the install in the main tree first`);
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
function readPkg(p) {
|
|
61
|
+
if (!fs.existsSync(p)) return null;
|
|
62
|
+
try {
|
|
63
|
+
return JSON.parse(fs.readFileSync(p, 'utf8'));
|
|
64
|
+
} catch (err) {
|
|
65
|
+
fail(`cannot parse ${p}: ${err.message}`);
|
|
66
|
+
return null; // unreachable — fail() exits the process
|
|
67
|
+
}
|
|
68
|
+
}
|
|
69
|
+
const mainPkg = readPkg(path.join(mainRoot, 'package.json'));
|
|
70
|
+
const wtPkg = readPkg(path.join(wtPath, 'package.json'));
|
|
71
|
+
if (mainPkg && wtPkg) {
|
|
72
|
+
const depsOf = (o) => JSON.stringify({ d: o.dependencies ?? {}, dd: o.devDependencies ?? {} });
|
|
73
|
+
if (depsOf(mainPkg) !== depsOf(wtPkg)) {
|
|
74
|
+
fail('package.json dependencies/devDependencies differ from the main tree; run a real install in this worktree instead of symlinking node_modules');
|
|
75
|
+
}
|
|
76
|
+
}
|
|
77
|
+
|
|
78
|
+
// --- helpers -----------------------------------------------------------------
|
|
79
|
+
|
|
80
|
+
// The symlinked node_modules must stay invisible to `git status`/`git add -A`
|
|
81
|
+
// in this worktree: .gitignore's `node_modules/` (trailing slash) matches a
|
|
82
|
+
// real directory but NOT a symlink, so without this a handoff pipeline that
|
|
83
|
+
// runs `git add -A` would stage a symlink whose target is an absolute,
|
|
84
|
+
// machine-local path. `info/exclude` resolves to the shared git common dir
|
|
85
|
+
// (verified: worktrees do not get a private copy), so this is a one-time,
|
|
86
|
+
// idempotent, untracked local-metadata change — never part of any commit.
|
|
87
|
+
function excludeNodeModulesFromGit(root) {
|
|
88
|
+
try {
|
|
89
|
+
const common = spawnSync('git', ['-C', root, 'rev-parse', '--git-common-dir'], { encoding: 'utf8' });
|
|
90
|
+
if (common.status !== 0) return false; // not a git repo (e.g. unit-test fixture) — best effort
|
|
91
|
+
const gitCommonDir = path.resolve(root, common.stdout.trim());
|
|
92
|
+
const excludePath = path.join(gitCommonDir, 'info', 'exclude');
|
|
93
|
+
fs.mkdirSync(path.dirname(excludePath), { recursive: true });
|
|
94
|
+
const existing = fs.existsSync(excludePath) ? fs.readFileSync(excludePath, 'utf8') : '';
|
|
95
|
+
if (existing.split('\n').some((l) => l.trim() === 'node_modules')) return false; // already present
|
|
96
|
+
const sep = existing.length > 0 && !existing.endsWith('\n') ? '\n' : '';
|
|
97
|
+
fs.appendFileSync(excludePath, `${sep}node_modules\n`);
|
|
98
|
+
return true;
|
|
99
|
+
} catch {
|
|
100
|
+
return false; // best effort — a missing/unwritable git dir must not fail provisioning
|
|
101
|
+
}
|
|
102
|
+
}
|
|
103
|
+
|
|
104
|
+
// The `.claude` copy must never clobber a git-TRACKED file (today exactly
|
|
105
|
+
// `.claude/commands/ukit/handoff-fullstack.md` — owned by TASK-005). A fresh
|
|
106
|
+
// worktree already has that file checked out by `git worktree add`; if the
|
|
107
|
+
// gitignored dev mirror in mainRoot has locally drifted from the committed
|
|
108
|
+
// copy, an unconditional recursive copy would silently overwrite the
|
|
109
|
+
// worktree's tracked file with main's unreviewed local state.
|
|
110
|
+
function trackedClaudeFiles(root) {
|
|
111
|
+
try {
|
|
112
|
+
const res = spawnSync('git', ['-C', root, 'ls-files', '--', '.claude'], { encoding: 'utf8' });
|
|
113
|
+
if (res.status !== 0 || !res.stdout) return new Set();
|
|
114
|
+
return new Set(res.stdout.split('\n').map((s) => s.trim()).filter(Boolean));
|
|
115
|
+
} catch {
|
|
116
|
+
return new Set();
|
|
117
|
+
}
|
|
118
|
+
}
|
|
119
|
+
|
|
120
|
+
// --- provision -------------------------------------------------------------
|
|
121
|
+
|
|
122
|
+
// 1. node_modules symlink
|
|
123
|
+
const wtModules = path.join(wtPath, 'node_modules');
|
|
124
|
+
let modulesSt = null;
|
|
125
|
+
try {
|
|
126
|
+
modulesSt = fs.lstatSync(wtModules);
|
|
127
|
+
} catch {
|
|
128
|
+
/* absent */
|
|
129
|
+
}
|
|
130
|
+
if (modulesSt === null) {
|
|
131
|
+
try {
|
|
132
|
+
fs.symlinkSync(mainModules, wtModules, 'dir');
|
|
133
|
+
} catch (err) {
|
|
134
|
+
fail(`cannot create node_modules symlink: ${err.message}`);
|
|
135
|
+
}
|
|
136
|
+
console.log(`provisioned node_modules -> symlink ${mainModules}`);
|
|
137
|
+
} else if (modulesSt.isSymbolicLink()) {
|
|
138
|
+
const target = fs.readlinkSync(wtModules);
|
|
139
|
+
if (path.resolve(wtPath, target) !== path.resolve(mainModules)) {
|
|
140
|
+
fail(`node_modules symlink already exists but points at ${target}, not ${mainModules}`);
|
|
141
|
+
}
|
|
142
|
+
console.log(`node_modules symlink already correct -> ${target}`);
|
|
143
|
+
} else {
|
|
144
|
+
fail(`${wtModules} already exists and is not a symlink`);
|
|
145
|
+
}
|
|
146
|
+
if (excludeNodeModulesFromGit(wtPath)) {
|
|
147
|
+
console.log('node_modules excluded from git via .git/info/exclude');
|
|
148
|
+
}
|
|
149
|
+
|
|
150
|
+
// 2. .claude real copy, preserving modes
|
|
151
|
+
const srcClaude = path.join(mainRoot, '.claude');
|
|
152
|
+
const dstClaude = path.join(wtPath, '.claude');
|
|
153
|
+
if (!fs.existsSync(srcClaude)) {
|
|
154
|
+
fail(`main tree has no .claude directory to copy: ${srcClaude}`);
|
|
155
|
+
} else {
|
|
156
|
+
// Always (re)copy: a fresh worktree may already carry a partial, git-tracked
|
|
157
|
+
// .claude/ (e.g. commands), which must not make us skip the hooks/scripts.
|
|
158
|
+
// Never overwrite a git-tracked file under .claude/ (e.g.
|
|
159
|
+
// handoff-fullstack.md) — leave whatever `git worktree add` already
|
|
160
|
+
// checked out for it alone.
|
|
161
|
+
const trackedInClaude = trackedClaudeFiles(mainRoot);
|
|
162
|
+
fs.cpSync(srcClaude, dstClaude, {
|
|
163
|
+
recursive: true,
|
|
164
|
+
filter: (src) => {
|
|
165
|
+
const rel = path.relative(mainRoot, src).split(path.sep).join('/');
|
|
166
|
+
return !trackedInClaude.has(rel);
|
|
167
|
+
},
|
|
168
|
+
});
|
|
169
|
+
console.log(
|
|
170
|
+
`provisioned .claude -> copy of ${srcClaude}` +
|
|
171
|
+
(trackedInClaude.size > 0 ? ` (skipped ${trackedInClaude.size} git-tracked file(s))` : ''),
|
|
172
|
+
);
|
|
173
|
+
}
|
|
174
|
+
|
|
175
|
+
// 3. .ukit/storage/config.json
|
|
176
|
+
const srcConfig = path.join(mainRoot, '.ukit', 'storage', 'config.json');
|
|
177
|
+
const dstConfig = path.join(wtPath, '.ukit', 'storage', 'config.json');
|
|
178
|
+
if (fs.existsSync(dstConfig)) {
|
|
179
|
+
console.log('.ukit/storage/config.json already present (skipped copy)');
|
|
180
|
+
} else if (!fs.existsSync(srcConfig)) {
|
|
181
|
+
fail(`main tree has no .ukit/storage/config.json to copy: ${srcConfig}`);
|
|
182
|
+
} else {
|
|
183
|
+
fs.mkdirSync(path.dirname(dstConfig), { recursive: true });
|
|
184
|
+
fs.copyFileSync(srcConfig, dstConfig);
|
|
185
|
+
console.log('provisioned .ukit/storage/config.json');
|
|
186
|
+
}
|
|
187
|
+
|
|
188
|
+
// 4. .cache/index
|
|
189
|
+
const srcCache = path.join(mainRoot, '.cache', 'index');
|
|
190
|
+
const dstCache = path.join(wtPath, '.cache', 'index');
|
|
191
|
+
if (fs.existsSync(dstCache)) {
|
|
192
|
+
console.log('.cache/index already present in worktree (skipped copy)');
|
|
193
|
+
} else if (!fs.existsSync(srcCache)) {
|
|
194
|
+
fail(`main tree has no .cache/index directory to copy: ${srcCache}`);
|
|
195
|
+
} else {
|
|
196
|
+
fs.cpSync(srcCache, dstCache, { recursive: true });
|
|
197
|
+
console.log('provisioned .cache/index');
|
|
198
|
+
}
|
|
199
|
+
|
|
200
|
+
process.exit(0);
|
|
@@ -442,11 +442,13 @@ export function buildCompactThresholds(config = {}) {
|
|
|
442
442
|
//
|
|
443
443
|
// Must stay BELOW the model's real context window or the cap is unreachable: the API
|
|
444
444
|
// rejects the request with "input exceeds the context window" long before an estimate
|
|
445
|
-
// climbing toward a higher number ever trips this.
|
|
446
|
-
//
|
|
447
|
-
//
|
|
448
|
-
//
|
|
449
|
-
|
|
445
|
+
// climbing toward a higher number ever trips this. 220_000 is sized for a 256k window
|
|
446
|
+
// and leaves headroom for the response plus the estimator's own undercount. On a 200k
|
|
447
|
+
// model this default sits ABOVE the window, so the gate can never fire and is dead code
|
|
448
|
+
// in exactly the situation it exists to prevent — set compact.hardCapTokens to 160_000
|
|
449
|
+
// or lower for those. Pair it with env.CLAUDE_CODE_AUTO_COMPACT_WINDOW, which must stay
|
|
450
|
+
// strictly below this cap so the client auto-compacts before the gate blocks tools.
|
|
451
|
+
const hardCapTokens = Math.max(1, finiteNumber(config?.compact?.hardCapTokens, 220_000));
|
|
450
452
|
|
|
451
453
|
return {
|
|
452
454
|
softThreshold,
|
package/templates/CLAUDE.md
CHANGED
|
@@ -128,6 +128,14 @@ Khi Handoff mode: đọc `docs/AI_HANDOFF/RULES.md` để biết 4 phase (Idea+P
|
|
|
128
128
|
- This is optional internal orchestration config from `.ukit/storage/config.json`; never turn it into an end-user workflow.
|
|
129
129
|
- Always preserve the CoDev priority: quality > safety > speed > token discipline.
|
|
130
130
|
|
|
131
|
+
## Post-Edit Sidecar Review (internal)
|
|
132
|
+
|
|
133
|
+
- If routed state's `routeSummary.line` includes a `review=code-reviewer(diff)` segment, a `local-build` or `shared-edit` task qualifies for a non-blocking second opinion — this exists because the daily-flow executor model can miss edge cases.
|
|
134
|
+
- Only launch it once write evidence AND verification evidence already exist for the task (never before; never as a substitute for either).
|
|
135
|
+
- Launch the `code-reviewer` agent via the Agent tool with `REVIEW_TARGET_TYPE=diff` and `run_in_background: true` (model `unic-smart` per `subagents.diffReviewModel`). Do not wait for it — continue and report the task as done using the normal completion rules.
|
|
136
|
+
- Its findings are advisory only: never re-open, block, or delay the already-reported completion on their account. Surface them to the user as a follow-up note if/when they arrive.
|
|
137
|
+
- This is internal orchestration — end users never invoke it directly; `ukit install` plus natural language remains the whole surface. No new commands.
|
|
138
|
+
|
|
131
139
|
## Selective Subagent Policy (internal only)
|
|
132
140
|
|
|
133
141
|
- Keep direct execution as the default for trivial/simple work.
|
|
@@ -79,8 +79,8 @@ Next: <bước kế tiếp chính xác>
|
|
|
79
79
|
- Subagent ghi **full log vào task file trên đĩa**, chỉ trả về orchestrator ≤10 dòng (executor) / ≤6 dòng (reviewer). Paste log ngược lại orchestrator là nguyên nhân số 1 làm run chết vì hết context.
|
|
80
80
|
- Hết mỗi wave: commit, ghi cursor, **collapse** wave đó còn 1 dòng/task trong bộ nhớ làm việc, rồi chạy tiếp.
|
|
81
81
|
- Yêu cầu `/compact` **chỉ** được đặt ở cuối command, giữa 2 cycle. Giữa cycle thì tuyệt đối không — state đã nằm hết ở git + `INDEX.md` + `RUN.md` nên compact ở ranh giới cycle không mất gì.
|
|
82
|
-
- Vượt `compact.hardCapTokens` (mặc định
|
|
83
|
-
- Không hook nào gọi được `/compact` — đó là lệnh client-only. Nhưng từ 2.1.3, settings mặc định đặt `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW =
|
|
82
|
+
- Vượt `compact.hardCapTokens` (mặc định 220k, cỡ cho context window 256k) mà `RUN.md` còn run dở: `context-hardcap-gate` cho thêm `compact.hardCapGraceCalls` (mặc định 10) tool call rồi mới chặn cứng. **Grace đó chỉ để hạ cánh** — hoàn tất edit đang dở, commit, ghi cursor, push. Không mở task mới, không đọc thêm file, không spawn agent. Hết grace là chặn thật; budget chỉ reset khi ước lượng token thực sự giảm (có compact thật), không reset theo wave.
|
|
83
|
+
- Không hook nào gọi được `/compact` — đó là lệnh client-only. Nhưng từ 2.1.3, settings mặc định đặt `env.CLAUDE_CODE_AUTO_COMPACT_WINDOW = 180000` < `hardCapTokens` (220k), nên **client tự auto-compact trước khi gate chặn**. Đường thường: auto-compact chạy → `handoff-resume.sh` replay cursor → chạy tiếp, không cần người gõ gì. Grace window ở trên chỉ còn là lưới an toàn.
|
|
84
84
|
- Sửa một trong hai số đó thì phải giữ `autoCompactWindow < hardCapTokens`. Đảo thứ tự là deadlock: gate chặn tool trước → transcript ngừng lớn → ngưỡng auto-compact không bao giờ tới. `tests/core/autoCompactWindow.test.js` khóa bất biến này.
|
|
85
85
|
|
|
86
86
|
### Git
|
|
@@ -132,6 +132,17 @@ pending_review ──[reviewer]──▶ approved | approved_minor ──▶ don
|
|
|
132
132
|
- `§ Test Files`: đường dẫn cụ thể file test sẽ tạo/sửa (ví dụ `tests/auth/login.test.js`).
|
|
133
133
|
- `§ Verification Commands`: lệnh executor sẽ chạy để xác nhận PASS. Nếu project có sẵn lint/typecheck script → BẮT BUỘC liệt kê ở đây, không chỉ lệnh test. Project không có thì ghi rõ N/A, không được bỏ qua im lặng.
|
|
134
134
|
- `§ Acceptance Criteria`: checklist.
|
|
135
|
+
|
|
136
|
+
#### Test selection (which tests the Verification Commands run)
|
|
137
|
+
|
|
138
|
+
Resolution order — exactly three steps, in this order. A task's Verification Commands MUST NOT default to the full suite:
|
|
139
|
+
|
|
140
|
+
1. Target File under `src/` or `scripts/` → read `.cache/index/tests-map.json` and take the `tests` array for that `sourceFile`.
|
|
141
|
+
2. Target File under `templates/.claude/**` or `.claude/**` → `tests-map.json` has no coverage of these paths, so use the path convention: hooks → `tests/hooks/` + `tests/handoff/cycle*/`; manifest/settings → `tests/manifest/`; runtime `.mjs` mirrors → `tests/core/*Parity*` + `tests/index/`.
|
|
142
|
+
3. **Mandatory non-empty floor** — if steps 1–2 resolve to fewer than one test file, the task MUST fall back to `yarn test:release-core`. An empty selection is never permitted. This floor is a fallback for a single task's narrowed selection, not a default — most tasks resolve via steps 1–2 and never reach it.
|
|
143
|
+
|
|
144
|
+
**Wave/cycle boundary regression net** — the three steps above narrow one task's Verification Commands only; they are not a substitute for full-suite coverage. A full `yarn test` run at each wave/cycle boundary MUST happen and is the regression net for every per-task narrowed selection made under this policy. `code-reviewer.md`'s "wave-boundary full `yarn test` ... is the regression net" sentence refers to this paragraph.
|
|
145
|
+
|
|
135
146
|
- Nếu split mà task nào không kèm được Test Cases + Test Files cụ thể → task đó chưa đủ `ready`, đánh `needs_breakdown`.
|
|
136
147
|
- Update `INDEX.md`: thêm row mỗi task với status `ready`.
|
|
137
148
|
- Đây là **điểm cắt cuối trước khi code chạy**: phase này xong, executor được phép pick. Trong `/ukit:handoff-fullstack`, gate ở đây là plan review độc lập (model mạnh, context riêng) chứ không phải human — vì người dùng đã chủ động chọn chạy one-shot.
|
|
@@ -9,7 +9,7 @@
|
|
|
9
9
|
"compact": {
|
|
10
10
|
"enabled": true,
|
|
11
11
|
"tokenThreshold": 50000,
|
|
12
|
-
"hardCapTokens":
|
|
12
|
+
"hardCapTokens": 220000,
|
|
13
13
|
"hardCapBlock": true,
|
|
14
14
|
"hardCapGraceCalls": 10,
|
|
15
15
|
"contextRotDetection": true,
|
|
@@ -105,8 +105,9 @@
|
|
|
105
105
|
"maxReadPasses": 2,
|
|
106
106
|
"maxContextPulls": 1,
|
|
107
107
|
"verificationPolicy": "targeted-if-covered",
|
|
108
|
-
"completionRule": "require-write",
|
|
109
|
-
"delegationPolicy": "disallow-by-default"
|
|
108
|
+
"completionRule": "require-write-and-verification",
|
|
109
|
+
"delegationPolicy": "disallow-by-default",
|
|
110
|
+
"postEditReviewPolicy": "sidecar-non-blocking"
|
|
110
111
|
},
|
|
111
112
|
"find-cause": {
|
|
112
113
|
"modelTier": "code",
|
|
@@ -121,7 +122,8 @@
|
|
|
121
122
|
"maxContextPulls": 2,
|
|
122
123
|
"verificationPolicy": "targeted-then-widen-on-risk",
|
|
123
124
|
"completionRule": "require-write-and-verification",
|
|
124
|
-
"delegationPolicy": "allow-qualified-sidecar"
|
|
125
|
+
"delegationPolicy": "allow-qualified-sidecar",
|
|
126
|
+
"postEditReviewPolicy": "sidecar-non-blocking"
|
|
125
127
|
},
|
|
126
128
|
"map-impact": {
|
|
127
129
|
"modelTier": "code",
|
|
@@ -187,7 +189,7 @@
|
|
|
187
189
|
"handoff": {
|
|
188
190
|
"enabled": true,
|
|
189
191
|
"crossTool": true,
|
|
190
|
-
"maxParallelAgents":
|
|
192
|
+
"maxParallelAgents": 12,
|
|
191
193
|
"autonomy": {
|
|
192
194
|
"askWindow": "plan-only",
|
|
193
195
|
"planReviewRounds": 2,
|
|
@@ -241,6 +243,9 @@
|
|
|
241
243
|
"enabled": true,
|
|
242
244
|
"smallTaskModel": "unic-lite",
|
|
243
245
|
"smallTaskAgent": "ukit-small-task-maintainer",
|
|
246
|
+
"diffReviewEnabled": true,
|
|
247
|
+
"diffReviewAgent": "code-reviewer",
|
|
248
|
+
"diffReviewModel": "unic-smart",
|
|
244
249
|
"visionEnabled": true,
|
|
245
250
|
"visionModel": "unic-vision",
|
|
246
251
|
"visionAgent": "ukit-vision-analyst",
|
|
@@ -384,7 +389,7 @@
|
|
|
384
389
|
"compact": {
|
|
385
390
|
"enabled": "Bật/tắt toàn bộ helper compact của UKit.",
|
|
386
391
|
"tokenThreshold": "Ngưỡng token chung cho runtime compact dùng chung.",
|
|
387
|
-
"hardCapTokens": "Ngưỡng cứng tuyệt đối (mặc định
|
|
392
|
+
"hardCapTokens": "Ngưỡng cứng tuyệt đối (mặc định 220000 token ước lượng, cỡ cho context window 256k). Chạm/vượt ngưỡng này thì context coi như quá dài — không phải gợi ý nữa, là bắt buộc. PHẢI thấp hơn context window thật của model, nếu không API sẽ báo lỗi vượt context trước khi gate kịp chặn — model 200k thì hạ xuống 160000 hoặc thấp hơn. Đi kèm env.CLAUDE_CODE_AUTO_COMPACT_WINDOW trong .claude/settings.json (mặc định 180000), số đó phải nhỏ hơn ngưỡng này để client tự auto-compact trước khi gate chặn tool.",
|
|
388
393
|
"hardCapBlock": "Nếu true, hook context-hardcap-gate chặn cứng Edit/Write/Bash (exit 2) khi vượt hardCapTokens, cho tới khi có compact thật (PreCompact) reset lại bộ đếm.",
|
|
389
394
|
"hardCapGraceCalls": "Số tool call được phép chạy tiếp sau khi vượt hardCapTokens KHI docs/AI_HANDOFF/RUN.md còn run dở (mặc định 10). Dùng để run kịp commit + ghi cursor + push rồi mới bị chặn, thay vì chết giữa lúc đang Edit. Hết grace là chặn cứng như cũ. Budget tính theo mỗi đợt vượt cap, chỉ reset khi ước lượng token thật sự giảm (có compact thật) — không reset theo wave.",
|
|
390
395
|
"contextRotDetection": "Phát hiện context quá dài/dễ mục để giữ lại state quan trọng trước khi AI nhớ sai.",
|
|
@@ -494,7 +499,7 @@
|
|
|
494
499
|
"handoff": {
|
|
495
500
|
"enabled": "Bật Quality Gate cho handoff: plan có Test Plan, executor test-first, reviewer model khác. Tắt = quay về flow cũ (dễ lọt lỗi vặt).",
|
|
496
501
|
"crossTool": "true nghĩa là handoff truyền qua file (PLAN/INDEX/tasks) chứ không qua in-process subagent — cho phép plan ở Claude Code, execute ở Kilo Code, review ở Claude Code khác model.",
|
|
497
|
-
"maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định
|
|
502
|
+
"maxParallelAgents": "Số agent chạy song song TỐI ĐA trong một wave (mặc định 12, tăng từ 3 để giảm thời gian chờ khi có nhiều task độc lập; 12 là số chọn theo dải an toàn ~10-15 dưới đây, không phải số đo được). Một wave có nhiều task hơn số này sẽ được chia thành nhiều batch chạy lần lượt — áp dụng cho cả Phase 3 Implement và Phase 4 Review. Lý do giới hạn vẫn còn: mỗi agent nền có context window riêng, và report của agent khi xong sẽ được inject ngược vào session chính — chạy quá nhiều cùng lúc (vd 20+) vẫn có thể làm session chính vượt context window và bỏ lại worktree rác. Hạ xuống 3-5 nếu task nặng (verification output dài) hoặc thấy compact bị trigger liên tục; tránh vượt quá ~10-15.",
|
|
498
503
|
"plan": {
|
|
499
504
|
"requireTestPlan": "Bắt buộc PLAN.md §4 phải có Test Plan trước khi task chuyển ready.",
|
|
500
505
|
"minTestsHappyPath": "Tối thiểu test cho happy path.",
|