@ngockhoale/ukit 2.3.3 → 2.3.5

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,86 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.5 - 2026-09-10
6
+
7
+ Source-audit hardening release: the code index cannot corrupt itself, and a corrupt runtime
8
+ cache can no longer take down `ukit status` or `ukit memory`.
9
+
10
+ Index artifacts (`files`, `symbols`, `imports`, `calls`, `tests-map`, `hotspots`,
11
+ `archetypes`, `relations`, `analogs`, `meta`) are now written through same-directory
12
+ `.tmp-<timestamp>-<random>` files followed by an atomic `rename`, in both the source builder
13
+ (`src/index/buildIndex.js`) and the shipped `templates/.claude/ukit/index/lib/index-core.mjs`
14
+ runtime mirror. A build that crashes mid-write — or two builds racing in one repo — can no
15
+ longer leave a half-written artifact that poisons every subsequent index read; failed writes
16
+ clean their temp file up before rethrowing.
17
+
18
+ `ukit status` and `ukit memory` now read their cache/memory JSON through tolerant readers:
19
+ a corrupt `compact-history.json`, `compact-pressure.json`, `prompt-cache.json`,
20
+ `output-history.json`, or memory file prints a warning naming the exact file path and falls
21
+ back to defaults instead of crashing the whole command mid-report. `readJsonIfExists` itself
22
+ stays strict everywhere else, so genuinely bad callers still fail loudly.
23
+
24
+ Nested test layouts (`packages/widget/tests/unit/...`) are now recognized as tests by the
25
+ source path classifier and both runtime-mirror predicates — those files previously slipped
26
+ past test detection, broke archetype classification, and never appeared in the test map.
27
+
28
+ Each fix ships with regression coverage: a rename-spy test asserting all ten artifacts move
29
+ through temp files, corrupt-fixture tests asserting the warned paths and surviving output,
30
+ and a nested-`tests/` mapping/archetype test. Full suite: 77 files, 1,348 tests green.
31
+
32
+ ## 2.3.4 - 2026-09-10
33
+
34
+ Silent-stall audit and compact-recovery release. Two themes: every block reports a
35
+ user-visible reason, and compaction must actually buy back context.
36
+
37
+ Security hooks can no longer be bypassed or silently disabled. `block-dangerous.sh` catches
38
+ split recursive/force flags (`rm -r -f`, `rm --force -r`) and evaluates every delete target
39
+ per shell segment, so one allowlisted word can no longer legitimise a sibling target
40
+ (`rm -rf dist /`); a bare `/` stays classified unsafe, wrapped forms keep the whole-command
41
+ fallback, and detections surface as structured `ask` decisions that never echo the raw
42
+ command. `block-dangerous.sh`, `protect-files.sh`, and `auto-allow-bash.sh` fall back to
43
+ Node when `jq` is absent (stock macOS), so a missing optional binary cannot disable a gate.
44
+ `protect-files.sh` matches exact sensitive basenames and directory components instead of
45
+ substrings — `.env.d.ts` and `app.envrc` are source files, not secrets — and emits a
46
+ structured human-decision response on real blocks. `sensitive-data-guard.sh` only treats
47
+ `-u/--user` as credentials on tools that authenticate with it, so `docker run --user
48
+ 1000:1000` no longer false-blocks while `curl -u user:pass` still does.
49
+
50
+ Stops always carry a reason. `handoff-model-guard.sh` pairs its stderr refusal with a
51
+ structured `systemMessage`; the completion gate reports a lost route state after real
52
+ session activity instead of releasing silently; hard-cap grace and sensitive-data blocks
53
+ already speak `systemMessage` from 2.3.3. `Phase: done …` with trailing status text now
54
+ counts as finished (`^done\b`, not `^done$`) in `handoff-resume.sh`,
55
+ `context-hardcap-gate.sh`, and `context-window-guard.sh`, so annotated done-phases stop
56
+ being resumed forever. Multi-file resume intents survive compaction: identity is session +
57
+ prompt key, because the router re-keys `requestKey` on every tool call and that churn is
58
+ what made post-compact resume a silent no-op.
59
+
60
+ The hard-cap grace budget is race-proof. Parallel subagents hitting the cap in one episode
61
+ used to race a single read-modify-write counter and multiply the budget; each spent call
62
+ now claims an exclusive-create sentinel file, so the sentinels present are the true count
63
+ and one claim can never be counted twice (verified by an 8-process parallel test: exactly
64
+ the configured grants, the next call blocked).
65
+
66
+ omp payloads moved off argv. PostToolUse Bash payloads embed whole tool outputs and exceed
67
+ macOS's ~256KB-per-argument limit, failing exec with E2BIG before any hook ran. The bridge
68
+ now writes the payload to a temp file (`@path` marker, hourly sweep, inline fallback for
69
+ read-only roots) and `hook-chain-runner.mjs` reads the marker; raw JSON never starts with
70
+ `@`, so both forms stay unambiguous.
71
+
72
+ Compact recovery, directed. Near-cap advisories now order the work: land the smallest
73
+ finishable item end-to-end, defer the rest as one-line notes to disk, delegate broad
74
+ reads/searches to subagents, then ask the user to compact. A new post-compact check fires
75
+ once per compact boundary when the live context after the boundary is still ≥60% of the
76
+ cap: stop re-reading old files (the fastest way to re-inflate a just-compacted session),
77
+ delegate, or finish in a fresh session — UKit resume state carries the goal across with a
78
+ near-empty window. The reinjected PROJECT CONTEXT block carries the same discipline.
79
+
80
+ Also fixed: `ukit install|diff --with-codegraph` crashed before the flag was consumed;
81
+ `route-task.mjs` helper commands use POSIX single-quote escaping instead of `JSON.stringify`
82
+ (printed commands were open to `$(...)`/backtick expansion) and reject absolute outside-root
83
+ targets instead of routing foreign-file context.
84
+
5
85
  ## 2.3.3 - 2026-09-10
6
86
 
7
87
  False-stop and long-context release. Four fixes that share one theme: UKit must run a prompt
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.3",
3
+ "version": "2.3.5",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -15,7 +15,7 @@ export async function runDiff({ packageRoot, projectRoot, packageVersion, argv =
15
15
  return;
16
16
  }
17
17
 
18
- const toolsArg = parseToolsArg(argv);
18
+ const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
19
19
  const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
20
20
  const selectedAdapterItemIds = toSelectedAdapterItemIds(selectedOptionalTools);
21
21
  const withCodegraph = argv.includes('--with-codegraph');
@@ -197,7 +197,10 @@ export async function pruneDeselectedAdapters({
197
197
  }
198
198
 
199
199
  export async function runInstall({ packageRoot, projectRoot, packageVersion, argv = [] }) {
200
- const toolsArg = parseToolsArg(argv);
200
+ // parseToolsArg owns only --tools and throws on anything else, so hand it the argv with
201
+ // the flags consumed here (--with-codegraph) already removed — order-dependent parsing
202
+ // used to make the documented `ukit install --with-codegraph` fail 100% of the time.
203
+ const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
201
204
  const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
202
205
  const withCodegraph = argv.includes('--with-codegraph');
203
206
 
@@ -14,6 +14,15 @@ function defaultUserMemory() {
14
14
  };
15
15
  }
16
16
 
17
+ async function readMemoryJson(filePath) {
18
+ try {
19
+ return await readJsonIfExists(filePath);
20
+ } catch (error) {
21
+ console.warn(`[UKit] Ignoring corrupt memory JSON file ${filePath}: ${error?.message ?? String(error)}`);
22
+ return null;
23
+ }
24
+ }
25
+
17
26
  async function readDirectoryJsonItems(dirPath) {
18
27
  let entries = [];
19
28
  try {
@@ -29,7 +38,7 @@ async function readDirectoryJsonItems(dirPath) {
29
38
  }
30
39
 
31
40
  const fullPath = path.join(dirPath, entry.name);
32
- const content = await readJsonIfExists(fullPath);
41
+ const content = await readMemoryJson(fullPath);
33
42
  if (content) {
34
43
  items.push({
35
44
  fileName: entry.name,
@@ -91,7 +100,7 @@ function createSearchableText(type, content) {
91
100
 
92
101
  export async function exportMemory(projectRoot) {
93
102
  const runtimePaths = buildRuntimePaths(projectRoot);
94
- const user = (await readJsonIfExists(runtimePaths.userMemoryPath)) ?? defaultUserMemory();
103
+ const user = (await readMemoryJson(runtimePaths.userMemoryPath)) ?? defaultUserMemory();
95
104
  const projects = (await readDirectoryJsonItems(runtimePaths.projectsDir)).map((item) => item.content);
96
105
  const sessions = (await readDirectoryJsonItems(runtimePaths.sessionsDir)).map((item) => item.content);
97
106
 
@@ -100,7 +109,7 @@ export async function exportMemory(projectRoot) {
100
109
 
101
110
  export async function listMemoryItems(projectRoot) {
102
111
  const runtimePaths = buildRuntimePaths(projectRoot);
103
- const userMemory = (await readJsonIfExists(runtimePaths.userMemoryPath)) ?? defaultUserMemory();
112
+ const userMemory = (await readMemoryJson(runtimePaths.userMemoryPath)) ?? defaultUserMemory();
104
113
  const projectMemories = await readDirectoryJsonItems(runtimePaths.projectsDir);
105
114
  const sessionMemories = await readDirectoryJsonItems(runtimePaths.sessionsDir);
106
115
 
@@ -188,7 +197,7 @@ function projectMemoryPath(runtimePaths, projectId) {
188
197
 
189
198
  async function readProjectMemoryForId(runtimePaths, projectId) {
190
199
  const filePath = projectMemoryPath(runtimePaths, projectId);
191
- const existing = await readJsonIfExists(filePath);
200
+ const existing = await readMemoryJson(filePath);
192
201
  const memory = {
193
202
  id: projectId,
194
203
  conventions: [],
@@ -204,7 +213,7 @@ async function appendSessionArchive(runtimePaths, projectId, archivedSessions) {
204
213
  }
205
214
 
206
215
  const archivePath = path.join(runtimePaths.projectsDir, `${sanitizeProjectId(projectId)}.archive.json`);
207
- const existing = (await readJsonIfExists(archivePath)) ?? { sessions: [] };
216
+ const existing = (await readMemoryJson(archivePath)) ?? { sessions: [] };
208
217
  const sessions = Array.isArray(existing.sessions) ? existing.sessions : [];
209
218
  await writeJson(archivePath, { sessions: [...sessions, ...archivedSessions] });
210
219
  }
@@ -5,7 +5,7 @@ import { loadRuntimeConfig } from './runtimeConfig.js';
5
5
  import { summarizeOutputHistory } from './output/index.js';
6
6
  import { buildPromptCacheStats } from './token/index.js';
7
7
  import { countMemoryItems } from './memory/store.js';
8
- import { readCompactPressureState } from './compact/threshold.js';
8
+ import { buildCompactPressureState } from './compact/threshold.js';
9
9
  import { detectProjectContext } from '../context/detectProjectContext.js';
10
10
 
11
11
  function formatPrimaryAgent(agentKey) {
@@ -122,6 +122,15 @@ function formatRecoveryLabel(count) {
122
122
  return ` / ${count} recover${count === 1 ? 'y' : 'ies'}`;
123
123
  }
124
124
 
125
+ async function readStatusJson(filePath) {
126
+ try {
127
+ return await readJsonIfExists(filePath);
128
+ } catch (error) {
129
+ console.warn(`[UKit] Ignoring corrupt status JSON file ${filePath}: ${error?.message ?? String(error)}`);
130
+ return null;
131
+ }
132
+ }
133
+
125
134
  export async function buildStatusReport(projectRoot) {
126
135
  const runtimePaths = buildRuntimePaths(projectRoot);
127
136
  const runtimeExists = await pathExists(runtimePaths.runtimeRoot);
@@ -131,11 +140,14 @@ export async function buildStatusReport(projectRoot) {
131
140
 
132
141
  const projectContext = await detectProjectContext(projectRoot);
133
142
  const config = await loadRuntimeConfig(projectRoot);
134
- const compactHistory = normalizeCompactEntries(await readJsonIfExists(runtimePaths.compactHistoryPath));
143
+ const compactHistory = normalizeCompactEntries(await readStatusJson(runtimePaths.compactHistoryPath));
135
144
  const compactSummary = buildCompactSummary(compactHistory);
136
- const compactPressure = await readCompactPressureState(projectRoot, config);
137
- const promptCacheStats = buildPromptCacheStats(await readJsonIfExists(runtimePaths.promptCachePath));
138
- const outputCompressionStats = summarizeOutputHistory(await readJsonIfExists(runtimePaths.outputHistoryPath));
145
+ const compactPressure = buildCompactPressureState(
146
+ await readStatusJson(runtimePaths.compactPressurePath),
147
+ config,
148
+ );
149
+ const promptCacheStats = buildPromptCacheStats(await readStatusJson(runtimePaths.promptCachePath));
150
+ const outputCompressionStats = summarizeOutputHistory(await readStatusJson(runtimePaths.outputHistoryPath));
139
151
  const memoryCounts = await countMemoryItems(projectRoot);
140
152
  const adapters = await detectAdapterLabels(projectRoot, config.agent);
141
153
 
@@ -1197,7 +1197,20 @@ function areSourceFingerprintsEqual(previousFingerprint, nextFingerprint) {
1197
1197
 
1198
1198
  async function writeArtifact(rootDir, artifactName, payload) {
1199
1199
  const artifactPath = getArtifactPath(rootDir, artifactName);
1200
- await fs.writeFile(artifactPath, JSON.stringify(payload, null, 2));
1200
+ const temporaryPath = `${artifactPath}.tmp-${Date.now()}-${Math.random().toString(16).slice(2)}`;
1201
+
1202
+ try {
1203
+ await fs.writeFile(temporaryPath, JSON.stringify(payload, null, 2));
1204
+ await fs.rename(temporaryPath, artifactPath);
1205
+ } catch (error) {
1206
+ try {
1207
+ await fs.rm(temporaryPath, { force: true });
1208
+ } catch {
1209
+ // ignore cleanup errors
1210
+ }
1211
+
1212
+ throw error;
1213
+ }
1201
1214
  }
1202
1215
 
1203
1216
  async function readArtifactIfExists(rootDir, artifactName) {
@@ -37,6 +37,7 @@ export function isLikelyTestFilePath(filePath) {
37
37
  || lower.startsWith('specs/')
38
38
  || lower.includes('/__tests__/')
39
39
  || lower.includes('/test/')
40
+ || lower.includes('/tests/')
40
41
  || lower.includes('/spec/')
41
42
  || lower.includes('/specs/')
42
43
  || lower.endsWith('.test.js')
@@ -3,7 +3,22 @@
3
3
  # Optimized: fast-path grep to skip Node.js when rule already exists.
4
4
 
5
5
  INPUT=$(cat)
6
- COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty' 2>/dev/null)
6
+ # jq is absent on stock macOS. Keep jq as the low-latency normal path, but fall back to
7
+ # UKit's required Node runtime so safe commands still receive managed metadata.
8
+ if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
9
+ COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
10
+ else
11
+ COMMAND=$(printf '%s' "$INPUT" | node -e '
12
+ const chunks = [];
13
+ process.stdin.on("data", (chunk) => chunks.push(chunk));
14
+ process.stdin.on("end", () => {
15
+ try {
16
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
17
+ process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
18
+ } catch {}
19
+ });
20
+ ' 2>/dev/null)
21
+ fi
7
22
 
8
23
  if [ -z "$COMMAND" ]; then
9
24
  exit 0
@@ -3,7 +3,22 @@
3
3
  # Matched on: Bash
4
4
 
5
5
  INPUT=$(cat)
6
- COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty')
6
+ # jq is not installed on stock macOS. Keep jq as the low-latency normal path, but use
7
+ # UKit's required Node runtime as a fallback so its absence cannot silently disable the gate.
8
+ if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
9
+ COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
10
+ else
11
+ COMMAND=$(printf '%s' "$INPUT" | node -e '
12
+ const chunks = [];
13
+ process.stdin.on("data", (chunk) => chunks.push(chunk));
14
+ process.stdin.on("end", () => {
15
+ try {
16
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
17
+ process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
18
+ } catch {}
19
+ });
20
+ ' 2>/dev/null)
21
+ fi
7
22
 
8
23
  if [ -z "$COMMAND" ]; then
9
24
  exit 0
@@ -59,22 +74,80 @@ for pattern in "${DANGEROUS_PATTERNS[@]}"; do
59
74
  fi
60
75
  done
61
76
 
62
- # Handle rm commands: allow safe cleanup targets, block risky recursive force-deletes
63
- if echo "$SCAN_COMMAND" | grep -qE "(^|[;&|[:space:]])rm([[:space:]]|$)"; then
64
- if echo "$SCAN_COMMAND" | grep -qE "rm\s+(-[a-zA-Z]*r[a-zA-Z]*f|-[a-zA-Z]*f[a-zA-Z]*r|--recursive\s+--force|--force\s+--recursive|--recursive\s+-f|-f\s+--recursive)"; then
65
- SAFE_DELETE_REGEX='(^|[[:space:]])(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/|[[:space:]]|$)'
66
- UNSAFE_TARGET_REGEX='(^|[[:space:]])(/|~|\.\.?($|/)|\.($|/))'
77
+ # Handle rm commands: allow safe cleanup targets only, defer everything else to a human.
78
+ # Direct rm invocations are parsed per shell segment (`;`, `&`, `|` split) with per-target
79
+ # allowlist evaluation: split flags (`rm -r -f`, `rm --force -r`) register, and one
80
+ # allowlisted word can never legitimise a sibling target (`rm -rf dist /`). A whole-command
81
+ # fallback still catches rm inside wrapped/quoted contexts (`bash -c "rm -r -f /x"`) and
82
+ # flags placed after targets (`rm src -rf`), where segment parsing cannot see rm as a head.
83
+ SAFE_DELETE_REGEX='(^|[[:space:]])(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/|[[:space:]]|$)'
84
+ UNSAFE_TARGET_REGEX='(^|[[:space:]])(/|~|\.\.?($|/)|\.($|/))'
85
+ SAFE_ONE_TARGET_REGEX='^(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/.*)?$'
86
+ UNSAFE_ONE_TARGET_REGEX='^(/.*|~|~/.*|\.\.|\.\./.*|\.)$'
87
+ RM_WORD_REGEX=$'(^|[;&|[:space:]"\x27])rm([[:space:]"\x27]|$)'
67
88
 
68
- # Check the allowlist first: `./dist` legitimately contains a leading "./" that would
69
- # otherwise be mistaken for a bare current-directory target by UNSAFE_TARGET_REGEX below.
70
- if echo "$SCAN_COMMAND" | grep -qE "$SAFE_DELETE_REGEX"; then
89
+ RM_VERDICT_UNSAFE=0
90
+ RM_VERDICT_GENERIC=0
91
+ RM_DIRECT_SEEN=0
92
+ set -f
93
+ while IFS= read -r segment; do
94
+ [ -n "$segment" ] || continue
95
+ seg_head=$(printf '%s' "$segment" | sed -E 's/^[[:space:]]*//' | awk '{print $1}' | xargs -I{} basename {} 2>/dev/null)
96
+ [ "$seg_head" = "rm" ] || continue
97
+ seg_flags=$(printf '%s' "$segment" | awk '{for(i=2;i<=NF;i++){if($i=="--")break;if($i ~ /^-/)printf "%s\n",$i;else break}}')
98
+ seg_has_r=0
99
+ seg_has_f=0
100
+ for flag in $seg_flags; do
101
+ case "$flag" in
102
+ --recursive) seg_has_r=1 ;;
103
+ --force) seg_has_f=1 ;;
104
+ --*) ;;
105
+ -*) case "$flag" in *[rR]*) seg_has_r=1 ;; esac
106
+ case "$flag" in *f*) seg_has_f=1 ;; esac ;;
107
+ esac
108
+ done
109
+ [ "$seg_has_r" -eq 1 ] && [ "$seg_has_f" -eq 1 ] || continue
110
+ RM_DIRECT_SEEN=1
111
+ seg_targets=$(printf '%s' "$segment" | awk '{saw=0;for(i=2;i<=NF;i++){if(!saw){if($i=="--"){saw=1;continue}if($i ~ /^-/)continue;saw=1;printf "%s\n",$i}else if($i !~ /^-/)printf "%s\n",$i}}')
112
+ seg_all_safe=1
113
+ seg_any_unsafe=0
114
+ for target in $seg_targets; do
115
+ # Strip ONE trailing slash so `dist/` matches the allowlist — but never shrink a
116
+ # bare `/` (or `~`) to an empty string, which would silently dodge the unsafe check.
117
+ if [ "${#target}" -gt 1 ]; then target="${target%/}"; fi
118
+ if printf '%s' "$target" | grep -qE "$UNSAFE_ONE_TARGET_REGEX"; then
119
+ seg_any_unsafe=1
120
+ seg_all_safe=0
121
+ elif printf '%s' "$target" | grep -qE "$SAFE_ONE_TARGET_REGEX"; then
71
122
  :
72
- elif echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
73
- emit_dangerous_decision "Unsafe delete target detected; recursive force-delete could cause irreversible damage. UKit defers this to a human decision."
74
123
  else
75
- emit_dangerous_decision "Recursive force-delete outside safe cleanup targets (dist/build/coverage/.next/.nuxt/.turbo/tmp/.cache/node_modules) is destructive. UKit defers this to a human decision."
124
+ seg_all_safe=0
76
125
  fi
126
+ done
127
+ if [ "$seg_any_unsafe" -eq 1 ]; then RM_VERDICT_UNSAFE=1
128
+ elif [ "$seg_all_safe" -eq 0 ]; then RM_VERDICT_GENERIC=1
129
+ fi
130
+ done <<SEGMENTS
131
+ $(printf '%s' "$SCAN_COMMAND" | tr ';&|' '\n\n\n')
132
+ SEGMENTS
133
+ set +f
134
+
135
+ if [ "$RM_DIRECT_SEEN" -eq 0 ] && echo "$SCAN_COMMAND" | grep -qE "$RM_WORD_REGEX" \
136
+ && echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--recursive|(^|[[:space:]])-[a-zA-Z]*[rR]" \
137
+ && echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--force|(^|[[:space:]])-[a-zA-Z]*f"; then
138
+ # Unsafe wins over the allowlist here: a safe word elsewhere in the line must never
139
+ # legitimise an absolute/home/parent target we can also see.
140
+ if echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
141
+ RM_VERDICT_UNSAFE=1
142
+ elif ! echo "$SCAN_COMMAND" | grep -qE "$SAFE_DELETE_REGEX"; then
143
+ RM_VERDICT_GENERIC=1
77
144
  fi
78
145
  fi
79
146
 
147
+ if [ "$RM_VERDICT_UNSAFE" -eq 1 ]; then
148
+ emit_dangerous_decision "Unsafe delete target detected; recursive force-delete could cause irreversible damage. UKit defers this to a human decision."
149
+ elif [ "$RM_VERDICT_GENERIC" -eq 1 ]; then
150
+ emit_dangerous_decision "Recursive force-delete outside safe cleanup targets (dist/build/coverage/.next/.nuxt/.turbo/tmp/.cache/node_modules) is destructive. UKit defers this to a human decision."
151
+ fi
152
+
80
153
  exit 0
@@ -75,7 +75,7 @@ function readRunCursor() {
75
75
  try {
76
76
  const text = fs.readFileSync(path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md'), 'utf8');
77
77
  const runPhase = (text.match(/^Phase:\s*(.+)$/m)?.[1] || '').trim();
78
- if (!runPhase || /^done$/i.test(runPhase)) return null;
78
+ if (!runPhase || /^done\b/i.test(runPhase)) return null;
79
79
  return {
80
80
  phase: runPhase,
81
81
  cursor: (text.match(/^Cursor:\s*(.+)$/m)?.[1] || '').trim(),
@@ -122,11 +122,35 @@ function readRunCursor() {
122
122
  }
123
123
  const thresholds = mod.buildCompactThresholds(config);
124
124
 
125
+ // The grace budget must survive parallel subagents hitting the cap in the same episode.
126
+ // A read-modify-write of one JSON file under-counts (two processes both read used=3 and
127
+ // both write 4, silently multiplying the budget), so each spent call is claimed with an
128
+ // exclusive-create sentinel file (openSync 'wx'): the sentinels present ARE the number
129
+ // of grace calls spent, and one claim can never be counted twice.
125
130
  const gracePath = path.join(projectRoot, '.ukit', 'storage', 'cache', 'hardcap-grace.json');
131
+ const graceDir = path.dirname(gracePath);
132
+ const slotPath = (n) => path.join(graceDir, `hardcap-grace.json.${n}`);
133
+ const readSlots = () => {
134
+ try {
135
+ return fs.readdirSync(graceDir)
136
+ .map((name) => /^hardcap-grace\.json\.(\d+)$/.exec(name))
137
+ .filter(Boolean)
138
+ .map((m) => Number(m[1]))
139
+ .sort((a, b) => a - b);
140
+ } catch {
141
+ return [];
142
+ }
143
+ };
144
+ const clearGrace = () => {
145
+ try { fs.rmSync(gracePath, { force: true }); } catch { /* advisory only */ }
146
+ for (const n of readSlots()) {
147
+ try { fs.rmSync(slotPath(n), { force: true }); } catch { /* advisory only */ }
148
+ }
149
+ };
126
150
 
127
151
  if (state.estimatedTotalTokens < thresholds.hardCapTokens) {
128
152
  // Back under the cap — the episode is over, so the next one starts with a full budget.
129
- fs.rmSync(gracePath, { force: true });
153
+ clearGrace();
130
154
  process.exit(0);
131
155
  return;
132
156
  }
@@ -145,43 +169,87 @@ function readRunCursor() {
145
169
  const graceCalls = Number.isFinite(config?.compact?.hardCapGraceCalls)
146
170
  ? config.compact.hardCapGraceCalls
147
171
  : 10;
148
- const prior = readJsonSafe(gracePath, null);
172
+
173
+ let slots = readSlots();
174
+ let startedAtTokens = null;
175
+ for (const n of slots) {
176
+ const data = readJsonSafe(slotPath(n), null);
177
+ if (data && Number.isFinite(data.startedAtTokens)) {
178
+ startedAtTokens = Math.max(startedAtTokens ?? -Infinity, data.startedAtTokens);
179
+ }
180
+ }
149
181
  // Reset only when the estimate actually fell since grace started: that is the signal a
150
182
  // real compaction/shed happened. Advancing the cursor alone must NOT top the budget up,
151
183
  // otherwise a long run gets unlimited grace and the ceiling stops meaning anything.
152
- const carried =
153
- prior && Number.isFinite(prior.startedAtTokens) && state.estimatedTotalTokens >= prior.startedAtTokens
154
- ? prior
155
- : { startedAtTokens: state.estimatedTotalTokens, used: 0 };
184
+ if (startedAtTokens !== null && state.estimatedTotalTokens < startedAtTokens) {
185
+ clearGrace();
186
+ slots = [];
187
+ startedAtTokens = null;
188
+ }
189
+ const spent = slots.length > 0 ? slots[slots.length - 1] : 0;
190
+ if (startedAtTokens === null) startedAtTokens = state.estimatedTotalTokens;
156
191
 
157
- if (carried.used < graceCalls) {
158
- const used = carried.used + 1;
192
+ if (spent < graceCalls) {
193
+ // Claim the next slot atomically: EEXIST means a parallel worker already claimed it,
194
+ // step to the next. A filesystem error keeps the old fail-open behavior — losing the
195
+ // counter must never block a landing — at the cost of one extra granted call.
196
+ let used = 0;
197
+ let counterFailed = false;
159
198
  try {
160
- fs.mkdirSync(path.dirname(gracePath), { recursive: true });
161
- fs.writeFileSync(gracePath, JSON.stringify({ ...carried, used }));
162
- if (ordinaryTask && ledgerMod && used === 1) {
163
- await ledgerMod.writeResumeIntent(projectRoot, payload);
199
+ fs.mkdirSync(graceDir, { recursive: true });
200
+ for (let n = spent + 1; n <= graceCalls; n += 1) {
201
+ try {
202
+ const fd = fs.openSync(slotPath(n), 'wx');
203
+ fs.writeSync(fd, JSON.stringify({ startedAtTokens, used: n, ts: Date.now() }));
204
+ fs.closeSync(fd);
205
+ used = n;
206
+ break;
207
+ } catch (err) {
208
+ if (err?.code === 'EEXIST') continue;
209
+ throw err;
210
+ }
164
211
  }
165
212
  } catch {
166
- // Losing the counter or advisory resume intent must not block the run.
213
+ counterFailed = true;
214
+ }
215
+ if (used === 0 && counterFailed) used = spent + 1;
216
+ if (used > 0 && used <= graceCalls) {
217
+ try {
218
+ fs.writeFileSync(gracePath, JSON.stringify({ startedAtTokens, used }));
219
+ if (ordinaryTask && ledgerMod && used === 1) {
220
+ await ledgerMod.writeResumeIntent(projectRoot, payload);
221
+ }
222
+ } catch {
223
+ // Losing the counter or advisory resume intent must not block the run.
224
+ }
225
+ const landing = run
226
+ ? `An unfinished handoff run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`
227
+ : 'An unfinished ordinary routed task is in flight, so this call is allowed instead of stranding it before compaction.';
228
+ process.stderr.write(
229
+ [
230
+ `CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
231
+ landing,
232
+ run
233
+ ? 'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.'
234
+ : 'Spend the remaining grace on LANDING, not on new work: finish the current mutation or verification, then compact as soon as the host allows it.',
235
+ 'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
236
+ 'Tell the user in your reply that context is over the cap and they should run /compact as soon as this landing step is complete.',
237
+ 'After compaction, the SessionStart resume hook replays the safe cursor and the task continues automatically.',
238
+ ].join('\n') + '\n',
239
+ );
240
+ // The stderr contract above only reaches the model — if the turn dies before the
241
+ // model relays it, the user sees nothing. Mirror the first grace call as a
242
+ // structured systemMessage so the /compact advice is user-visible regardless.
243
+ if (used === 1) {
244
+ process.stdout.write(`${JSON.stringify({
245
+ systemMessage: `UKit: context is over the hard cap (~${state.estimatedTotalTokens} tokens). A short landing window (${graceCalls} calls) is active — run /compact when it ends; the task resumes automatically after compaction.`,
246
+ })}\n`);
247
+ }
248
+ process.exit(0);
249
+ return;
167
250
  }
168
- const landing = run
169
- ? `An unfinished handoff run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`
170
- : 'An unfinished ordinary routed task is in flight, so this call is allowed instead of stranding it before compaction.';
171
- process.stderr.write(
172
- [
173
- `CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
174
- landing,
175
- run
176
- ? 'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.'
177
- : 'Spend the remaining grace on LANDING, not on new work: finish the current mutation or verification, then compact as soon as the host allows it.',
178
- 'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
179
- 'Tell the user in your reply that context is over the cap and they should run /compact as soon as this landing step is complete.',
180
- 'After compaction, the SessionStart resume hook replays the safe cursor and the task continues automatically.',
181
- ].join('\n') + '\n',
182
- );
183
- process.exit(0);
184
- return;
251
+ // Every slot was claimed while we raced — the budget is genuinely spent; fall
252
+ // through to the hard block below.
185
253
  }
186
254
  }
187
255
 
@@ -197,6 +265,11 @@ function readRunCursor() {
197
265
  'Do not work around this by summarizing inline and continuing, and do not reach for a non-gated write tool.',
198
266
  ];
199
267
  process.stderr.write(`${lines.join('\n')}\n`);
268
+ // stderr reaches the model, not reliably the user. Pair the hard block with a
269
+ // structured systemMessage so "why did it stop" always has a user-visible answer.
270
+ process.stdout.write(`${JSON.stringify({
271
+ systemMessage: `UKit blocked ${toolName}: context is over the hard cap (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}) and the landing window is spent. Run /compact now; after compaction the task continues automatically.`,
272
+ })}\n`);
200
273
  process.exit(2);
201
274
  })().catch((err) => {
202
275
  // A logic error here fails OPEN: this is a backstop on top of advisory nudges,
@@ -94,12 +94,14 @@ if (truncated) lines.shift(); // a tail read almost certainly split the first li
94
94
  // Everything before the last compact boundary is already summarised away and is NOT in
95
95
  // the live context. Counting it is what makes naive file-size estimates useless.
96
96
  let start = 0;
97
+ let boundaryTs = null;
97
98
  for (let i = lines.length - 1; i >= 0; i -= 1) {
98
99
  if (!lines[i].includes('compact_boundary')) continue;
99
100
  try {
100
101
  const entry = JSON.parse(lines[i]);
101
102
  if (entry?.type === 'system' && entry?.subtype === 'compact_boundary') {
102
103
  start = i + 1;
104
+ boundaryTs = Date.parse(entry.timestamp) || null;
103
105
  break;
104
106
  }
105
107
  } catch { /* not a usable boundary line */ }
@@ -200,6 +202,31 @@ if (swapDetected && persistedState.lastSwapModel !== currentModel) {
200
202
  }
201
203
  }
202
204
 
205
+ // ── Post-compact shrink check ──
206
+ // Compaction only helps if the live context AFTER the boundary is actually small. When the
207
+ // prompts right after a fresh compact already measure >= 60% of the cap, the summary did
208
+ // not shrink the session enough — and re-compacting the same session recovers very little.
209
+ // The honest recovery moves are: stop re-inflating it with rereads, delegate broad work to
210
+ // subagents (separate windows), or finish in a FRESH session (UKit resume state carries
211
+ // the goal across). Fires once per compact boundary.
212
+ const POST_COMPACT_RATIO = 0.6;
213
+ const POST_COMPACT_WINDOW_MS = 10 * 60 * 1000;
214
+ const boundaryAgeMs = boundaryTs !== null ? Date.now() - boundaryTs : null;
215
+ const justCompacted = boundaryAgeMs !== null && boundaryAgeMs >= 0 && boundaryAgeMs < POST_COMPACT_WINDOW_MS;
216
+ if (justCompacted && ratio >= POST_COMPACT_RATIO && persistedState.lastPostCompactBoundary !== String(boundaryTs)) {
217
+ persistedState.lastPostCompactBoundary = String(boundaryTs);
218
+ writeGuardState(persistedState);
219
+ process.stdout.write(
220
+ [
221
+ `UKIT POST-COMPACT CHECK — compaction ~${Math.max(1, Math.round(boundaryAgeMs / 60000))} min ago, but live context is still ~${estimatedTokens.toLocaleString()} tokens (${Math.round(ratio * 100)}% of the ${hardCap.toLocaleString()} cap).`,
222
+ 'The compact did not shrink this session enough. Recover context now, in this order:',
223
+ '1) STOP re-reading old files, logs and tool outputs — durable state is in the PROJECT CONTEXT block or on disk. Rereads are what re-inflate a just-compacted session.',
224
+ '2) DELEGATE remaining broad work (searches, big reads, multi-file edits) to subagents: their tool output lives in their own windows; keep this context for decisions and short summaries only.',
225
+ '3) If substantial work still remains: land the current step, persist docs/STATUS.md, then ask the user to start a FRESH session — UKit resume state carries the goal forward with a near-empty window. Another /compact here recovers far less.',
226
+ ].join('\n') + '\n',
227
+ );
228
+ }
229
+
203
230
  if (ratio < 0.8) process.exit(0);
204
231
 
205
232
  const pct = Math.round(ratio * 100);
@@ -226,7 +253,7 @@ function readRunCursor() {
226
253
  try {
227
254
  const text = fs.readFileSync(path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md'), 'utf8');
228
255
  const runPhase = (text.match(/^Phase:\s*(.+)$/m)?.[1] || '').trim();
229
- if (!runPhase || /^done$/i.test(runPhase)) return null;
256
+ if (!runPhase || /^done\b/i.test(runPhase)) return null;
230
257
  return {
231
258
  phase: runPhase,
232
259
  cursor: (text.match(/^Cursor:\s*(.+)$/m)?.[1] || '').trim(),
@@ -261,9 +288,10 @@ if (withinCooldown) {
261
288
  writeGuardState({ ...persistedState, lastPhase: phase, lastActionAt: now });
262
289
  } else {
263
290
  lines_out.push('ACTION THIS TURN, before starting new investigation/subagents/pipeline phases:');
264
- lines_out.push('1) Persist current progress now (update docs/STATUS.md, e.g. via the update-status skill) so nothing is lost.');
265
- lines_out.push('2) If the remaining work still has multiple steps or pipeline phases left, split what remains into bounded docs/AI_HANDOFF tasks instead of continuing in this one thread — smaller phases keep each turn under the cap.');
266
- lines_out.push('3) Otherwise, make the user\'s next action the FIRST sentence of your reply so it cannot be missed: "Context sắp đầy — gõ /compact ngay (run /compact now)", or wrap up and start a fresh session — work already committed/persisted to disk is not lost. Do not keep working silently until tools start being refused.');
291
+ lines_out.push('1) LAND ONE THING: pick the smallest in-flight item you can finish end-to-end in ≤3 tool calls (edit + verify), complete exactly that one, and report it done. Do not open anything new.');
292
+ lines_out.push('2) DEFER THE REST: write each remaining step as one line into docs/STATUS.md (update-status skill), or split what remains into bounded docs/AI_HANDOFF tasks — a fresh context then continues from disk without re-reading this thread.');
293
+ lines_out.push('3) DELEGATE anything broad that must still run (searches, big reads, multi-file edits) to subagents instead of doing it inline — their tool output stays in their own windows, not this one.');
294
+ lines_out.push('4) Then make the user\'s next action the FIRST sentence of your reply: "Context sắp đầy — gõ /compact ngay (run /compact now)". After /compact UKit auto-resumes from disk; if the compacted session is still large, ask for a fresh session instead. Do not keep working silently until tools start being refused.');
267
295
  if (sidechainEntries > 0) {
268
296
  lines_out.push(`Note: ${sidechainEntries} subagent entries in this stretch — each teammate carries its own context window, and every finished report is injected back here, so running many at once is the fastest way to overflow this session. Avoid spawning more until context drops back under the cap.`);
269
297
  }
@@ -39,6 +39,13 @@ const input = payload?.tool_input || {};
39
39
 
40
40
  function block(message) {
41
41
  process.stderr.write(`BLOCKED (handoff model-tier guard): ${message}\n`);
42
+ // stderr reaches the model, not reliably the user. Pair the refusal with a structured
43
+ // systemMessage (first line only — the full problem list stays in stderr) so a
44
+ // handoff-contract block never looks like a silent stall in the UI.
45
+ const firstLine = String(message).split('\n').find((l) => l.trim()) || 'model-tier contract violated';
46
+ process.stdout.write(`${JSON.stringify({
47
+ systemMessage: `UKit handoff model-tier guard refused this action: ${firstLine.trim()}`,
48
+ })}\n`);
42
49
  process.exit(2);
43
50
  }
44
51
 
@@ -79,7 +79,7 @@ async function emitOrdinaryResume() {
79
79
 
80
80
  const phase = field('Phase');
81
81
  // `done` means the last cycle finished cleanly. Anything else means a step was in flight.
82
- if (!phase || /^done$/i.test(phase)) {
82
+ if (!phase || /^done\b/i.test(phase)) {
83
83
  await emitOrdinaryResume();
84
84
  process.exit(0);
85
85
  }
@@ -3,7 +3,22 @@
3
3
  # Matched on: Edit|Write
4
4
 
5
5
  INPUT=$(cat)
6
- FILE_PATH=$(echo "$INPUT" | jq -r '.tool_input.file_path // empty')
6
+ # jq is absent on stock macOS. Keep jq as the low-latency normal path, but fall back to
7
+ # UKit's required Node runtime so a missing optional binary cannot disable this gate.
8
+ if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
9
+ FILE_PATH=$(printf '%s' "$INPUT" | jq -r '.tool_input.file_path // empty')
10
+ else
11
+ FILE_PATH=$(printf '%s' "$INPUT" | node -e '
12
+ const chunks = [];
13
+ process.stdin.on("data", (chunk) => chunks.push(chunk));
14
+ process.stdin.on("end", () => {
15
+ try {
16
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
17
+ process.stdout.write(typeof payload?.tool_input?.file_path === "string" ? payload.tool_input.file_path : "");
18
+ } catch {}
19
+ });
20
+ ' 2>/dev/null)
21
+ fi
7
22
 
8
23
  if [ -z "$FILE_PATH" ]; then
9
24
  exit 0
@@ -16,23 +31,32 @@ if [[ "$BASENAME" == ".env.example" ]]; then
16
31
  exit 0
17
32
  fi
18
33
 
19
- PROTECTED_PATTERNS=(
20
- ".env"
21
- ".env.local"
22
- ".env.production"
23
- "package-lock.json"
24
- "yarn.lock"
25
- ".git/"
26
- "node_modules/"
27
- ".claude/settings.local.json"
28
- ".ukit/storage/security/"
29
- )
34
+ # Match exact sensitive filenames and directory components, never arbitrary substrings:
35
+ # `.env.d.ts` and `app.envrc` are source files, not secrets. The old `*".env"*`
36
+ # check blocked both, which could appear as a mysterious stopped edit.
37
+ PATH_SLASHED="/${FILE_PATH#/}"
38
+ PROTECTED_MATCH=""
39
+ case "$BASENAME" in
40
+ .env|.env.local|.env.production|package-lock.json|yarn.lock)
41
+ PROTECTED_MATCH="$BASENAME"
42
+ ;;
43
+ esac
44
+ if [ -z "$PROTECTED_MATCH" ]; then
45
+ case "$PATH_SLASHED" in
46
+ */.git/*) PROTECTED_MATCH=".git/" ;;
47
+ */node_modules/*) PROTECTED_MATCH="node_modules/" ;;
48
+ */.claude/settings.local.json) PROTECTED_MATCH=".claude/settings.local.json" ;;
49
+ */.ukit/storage/security/*) PROTECTED_MATCH=".ukit/storage/security/" ;;
50
+ esac
51
+ fi
30
52
 
31
- for pattern in "${PROTECTED_PATTERNS[@]}"; do
32
- if [[ "$FILE_PATH" == *"$pattern"* ]]; then
33
- echo "BLOCKED: Cannot modify '$FILE_PATH' — matches protected pattern '$pattern'. Ask the user to modify this file manually." >&2
34
- exit 2
35
- fi
36
- done
53
+ if [ -n "$PROTECTED_MATCH" ]; then
54
+ # A stderr-only refusal can look like a silent stall when the host does not render
55
+ # hook stderr. Do not expose the path/pattern on stdout: the generic explanation is
56
+ # enough for the user and avoids leaking a potentially sensitive filename.
57
+ printf '%s\n' '{"hookSpecificOutput":{"hookEventName":"PreToolUse","permissionDecision":"ask","permissionDecisionReason":"UKit protected-file guard blocked this edit. Ask the user to modify the protected file manually."}}'
58
+ echo "BLOCKED: Cannot modify '$FILE_PATH' — matches protected pattern '$PROTECTED_MATCH'. Ask the user to modify this file manually." >&2
59
+ exit 2
60
+ fi
37
61
 
38
62
  exit 0
@@ -163,8 +163,18 @@ function scanBashShapes(command) {
163
163
  }
164
164
  }
165
165
 
166
- if (/(^|\s)(-u|--user)\s+["']?[^\s"':]+:[^\s"']+/.test(segment)) {
167
- found.push({ label: 'curl/wget inline credentials (-u user:password)', value: segment });
166
+ // `-u/--user user:pass` is only credentials on tools that authenticate with it.
167
+ // The same flag shape on docker/runuser/chown (`--user 1000:1000`, a uid:gid pair)
168
+ // is NOT a secret — matching it anyway false-positives every containerised test run.
169
+ const headBase = head.replace(/^.*\//, '');
170
+ const CRED_USER_TOOLS = new Set(['curl', 'wget', 'ftp', 'lftp', 'aria2c', 'http', 'https']);
171
+ const inlineUser = /(^|\s)(-u|--user)\s+["']?([^\s"':]+):([^\s"']+)/.exec(segment);
172
+ if (inlineUser && CRED_USER_TOOLS.has(headBase)) {
173
+ // Even on credential tools, a purely numeric pair is a uid:gid, not a password.
174
+ const numericIds = /^\d+$/.test(inlineUser[3]) && /^\d+$/.test(inlineUser[4]);
175
+ if (!numericIds) {
176
+ found.push({ label: 'curl/wget inline credentials (-u user:password)', value: segment });
177
+ }
168
178
  }
169
179
  if (/https?:\/\/[^\s/@'"]+:[^\s/@'"]+@/.test(segment)) {
170
180
  found.push({ label: 'URL with embedded credentials (user:pass@host)', value: segment });
@@ -263,6 +273,12 @@ const lines = [
263
273
  'Never repeat or guess the detected value in your reply. Ask the user instead.',
264
274
  ];
265
275
  process.stderr.write(`${lines.join('\n')}\n`);
276
+ // stderr reaches the model; the user needs a stop reason they can actually see. Emit a
277
+ // structured systemMessage (redacted — labels only, never the matched value) so a
278
+ // sensitive-data block never looks like a silent stall in the UI.
279
+ process.stdout.write(`${JSON.stringify({
280
+ systemMessage: `UKit blocked this call: ${findings.length} potential secret(s) detected (${findings.map((f) => f.label).slice(0, 3).join('; ')}${findings.length > 3 ? '; …' : ''}). Redact the value, allowlist its sha256, or ask the user to decide.`,
281
+ })}\n`);
266
282
  process.exit(2);
267
283
  NODE
268
284
 
@@ -2580,6 +2580,34 @@ const { pathToFileURL } = require('url');
2580
2580
  routingContext,
2581
2581
  previousContextFingerprint: previousContext?.fingerprint || null,
2582
2582
  });
2583
+ // Persist the completion contract BEFORE any cache import or indexed-context resolution.
2584
+ // Those later steps can take longer than the host's PreToolUse hook timeout. Previously a
2585
+ // timeout killed this process before its only state write, and completion-gate.sh then had
2586
+ // no route to evaluate — so it released Stop with no message. This provisional state has
2587
+ // the full execution/evidence contract; the final richer state below atomically replaces it.
2588
+ const provisionalRouteSummary = buildRouteSummary({
2589
+ activeSkills: selected,
2590
+ routingContext,
2591
+ });
2592
+ provisionalRouteSummary.continuationState = advanceContinuationState(
2593
+ provisionalRouteSummary.continuationState,
2594
+ previous?.routeSummary?.continuationState || null,
2595
+ );
2596
+ applyRescueBiasToRouteSummary(provisionalRouteSummary);
2597
+ const provisionalState = {
2598
+ ...(sessionId ? { sessionId } : {}),
2599
+ requestKey,
2600
+ ts: now,
2601
+ source: 'skill-router-provisional',
2602
+ activeSkills: compactActiveSkills(selected),
2603
+ routingContext: compactRoutingContext(routingContext),
2604
+ ...(previousContext ? { previousContext: compactPreviousContext(previousContext) } : {}),
2605
+ routeSummary: compactRouteSummary(provisionalRouteSummary),
2606
+ };
2607
+ provisionalState.fingerprint = buildRouteStateFingerprint(provisionalState);
2608
+ ensureDir(statePath);
2609
+ fs.writeFileSync(statePath, JSON.stringify(provisionalState));
2610
+
2583
2611
  let cacheUtils = null;
2584
2612
  if (fs.existsSync(cacheUtilsPath)) {
2585
2613
  try {
@@ -2028,7 +2028,21 @@ function areBugIndexSnapshotsEqual(previousSnapshot, nextSnapshot) {
2028
2028
  }
2029
2029
 
2030
2030
  async function writeArtifact(rootDir, artifactName, payload) {
2031
- await fs.writeFile(getArtifactPath(rootDir, artifactName), JSON.stringify(payload, null, 2));
2031
+ const artifactPath = getArtifactPath(rootDir, artifactName);
2032
+ const temporaryPath = `${artifactPath}.tmp-${Date.now()}-${Math.random().toString(16).slice(2)}`;
2033
+
2034
+ try {
2035
+ await fs.writeFile(temporaryPath, JSON.stringify(payload, null, 2));
2036
+ await fs.rename(temporaryPath, artifactPath);
2037
+ } catch (error) {
2038
+ try {
2039
+ await fs.rm(temporaryPath, { force: true });
2040
+ } catch {
2041
+ // ignore cleanup errors
2042
+ }
2043
+
2044
+ throw error;
2045
+ }
2032
2046
  }
2033
2047
 
2034
2048
  async function readArtifact(rootDir, artifactName) {
@@ -2557,6 +2571,7 @@ function isLikelyTestFilePath(filePath) {
2557
2571
  || lower.startsWith('specs/')
2558
2572
  || lower.includes('/__tests__/')
2559
2573
  || lower.includes('/test/')
2574
+ || lower.includes('/tests/')
2560
2575
  || lower.includes('/spec/')
2561
2576
  || lower.includes('/specs/')
2562
2577
  || lower.endsWith('.test.js')
@@ -2988,6 +3003,7 @@ function isLikelyTestFile(filePath) {
2988
3003
  || lower.startsWith('specs/')
2989
3004
  || lower.includes('/__tests__/')
2990
3005
  || lower.includes('/test/')
3006
+ || lower.includes('/tests/')
2991
3007
  || lower.includes('/spec/')
2992
3008
  || lower.includes('/specs/')
2993
3009
  || lower.endsWith('.test.js')
@@ -2455,10 +2455,18 @@ function inferTaskType({ promptText, commandText, selectedIds, targetFile = null
2455
2455
  return inferred;
2456
2456
  }
2457
2457
 
2458
+ // Must mirror src/index/taskRouting.js: the printed helper command is meant to be pasted
2459
+ // into a shell, so arguments need POSIX single-quote escaping — JSON.stringify's double
2460
+ // quotes leave $(...), backticks, and $VAR open to shell expansion.
2461
+ function shellEscape(str) {
2462
+ if (typeof str !== 'string') return '';
2463
+ return "'" + str.replace(/'/g, "'\\''") + "'";
2464
+ }
2465
+
2458
2466
  function buildHelperCommand({ commandNamespace = '.claude', scriptName, intent = '', targetFile = null, taskType = null }) {
2459
2467
  const parts = ['node', `${commandNamespace}/ukit/index/${scriptName}`];
2460
- if (intent) parts.push(JSON.stringify(intent));
2461
- if (targetFile) parts.push('--target', JSON.stringify(targetFile));
2468
+ if (intent) parts.push(shellEscape(intent));
2469
+ if (targetFile) parts.push('--target', shellEscape(targetFile));
2462
2470
  if (taskType) parts.push('--type', taskType);
2463
2471
  return parts.join(' ');
2464
2472
  }
@@ -2666,7 +2674,9 @@ function normalizeRelativeFile(rootDir, rawFilePath) {
2666
2674
  if (relative && !relative.startsWith('..') && !path.isAbsolute(relative)) {
2667
2675
  return relative.replaceAll('\\', '/');
2668
2676
  }
2669
- return trimmed.replaceAll('\\', '/');
2677
+ // Absolute targets outside the project root are not routable repo files (src twin
2678
+ // returns null here) — emitting them pulled outlines of foreign files into context.
2679
+ return null;
2670
2680
  }
2671
2681
 
2672
2682
  return trimmed.replace(/^\.\/+/, '').replaceAll('\\', '/');
@@ -154,10 +154,15 @@ export async function readResumeIntent(projectRoot, payload = {}, { consume = fa
154
154
  const state = await readRouteState(projectRoot, payload);
155
155
  const ledger = await readExecutionLedger(projectRoot, payload);
156
156
  const promptKey = explicitPromptKey(state || {});
157
+ // Identity is session + promptKey, NOT requestKey. The router re-keys requestKey on
158
+ // every tool call (commandText/targetFile are part of the key), so requiring requestKey
159
+ // equality invalidated the intent after the FIRST tool call between the gate block and
160
+ // the compaction — the post-compact session then had no continuation instruction and
161
+ // sat idle silently. hasUnfinishedCompletion still re-validates against the CURRENT
162
+ // ledger evidence, so a request that finished meanwhile consumes nothing and never
163
+ // resumes; a genuinely different prompt produces a different promptKey and stays dead.
157
164
  if (!state || !ledger || ledger.sessionId !== sessionId
158
165
  || ledger.sessionKey !== safeSegment(sessionId)
159
- || state.requestKey !== intent.requestKey
160
- || ledger.requestKey !== intent.requestKey
161
166
  || promptKey !== intent.promptKey
162
167
  || ledger.promptKey !== intent.promptKey
163
168
  || !hasUnfinishedCompletion(state, ledger)) {
@@ -610,6 +615,20 @@ async function main() {
610
615
  const ledger = await readExecutionLedger(projectRoot, payload) || {};
611
616
  const result = evaluateCompletion({ state, ledger });
612
617
 
618
+ // A session ledger with activity proves this session was routed and accumulating
619
+ // evidence. If the route state is now missing or unreadable (router crashed/timed out
620
+ // before persisting it, or the session stamp was lost), the gate cannot evaluate
621
+ // completion — and ending silently there is indistinguishable from a mid-run stall.
622
+ // Surface the loss so the user sees WHY the run released. Sessions that were never
623
+ // routed (no ledger activity at all) stay silent: there is nothing to report.
624
+ if (!state && (ledger.sourceSucceeded || ledger.writeAttempted
625
+ || ledger.verificationAttempted || (ledger.receipts || []).length > 0)) {
626
+ process.stdout.write(`${JSON.stringify({
627
+ systemMessage: 'UKit completion gate: route state for this session was lost before the stop was evaluated (router did not persist it), so unfinished work could not be verified. Send the task again in a new message to re-route and finish it.',
628
+ })}\n`);
629
+ return;
630
+ }
631
+
613
632
  // Claude Code invokes Stop again after a Stop hook blocks the first stop. Re-blocking
614
633
  // that recovery turn creates a self-sustaining loop, so let it end normally instead.
615
634
  // If work still lacks evidence, surface the recovery reason to the user rather than
@@ -106,7 +106,18 @@ function run(payloadText, scriptPaths) {
106
106
  }
107
107
 
108
108
  try {
109
- const [, , payloadText = '{}', ...scriptPaths] = process.argv;
109
+ const [, , payloadArg = '{}', ...scriptPaths] = process.argv;
110
+ // The bridge passes the payload as a temp file (`@path`) when it can: argv is capped
111
+ // (~256KB per arg on macOS) and PostToolUse Bash payloads embed whole tool outputs.
112
+ // A leading '@' cannot occur in raw JSON, so the two forms are unambiguous.
113
+ let payloadText = payloadArg;
114
+ if (payloadArg.startsWith('@')) {
115
+ try {
116
+ payloadText = fs.readFileSync(payloadArg.slice(1), 'utf8');
117
+ } catch {
118
+ payloadText = '{}';
119
+ }
120
+ }
110
121
  process.stdout.write(JSON.stringify(run(payloadText, scriptPaths)));
111
122
  } catch (error) {
112
123
  process.stdout.write(JSON.stringify({
@@ -58,6 +58,10 @@ async function main() {
58
58
  ...await buildRecentOutputLines(projectRoot, config),
59
59
  ...await buildResolverHintLines(projectRoot, state),
60
60
  ...await buildDuraOneHintLines(projectRoot),
61
+ // Post-compact discipline: the host summary keeps whatever it kept — the fastest way to
62
+ // re-inflate a just-compacted session is re-reading old files/logs to "recover" detail
63
+ // that is already on disk. Say it once, right where the fresh context starts.
64
+ '- Post-compact discipline: do NOT re-read old files/logs/tool outputs to rebuild context — durable state is on disk and in this block. Delegate broad reads/searches to subagents (separate windows), keep replies short, and let the main window rebuild headroom before any wide investigation.',
61
65
  '- If Claude hooks block broad verification, follow the routed targeted order or ask for explicit full-suite confirmation.',
62
66
  ].filter(Boolean);
63
67
 
@@ -248,6 +248,39 @@ function resolveNodeExecutable() {
248
248
  return cachedNodeExecutable;
249
249
  }
250
250
 
251
+ // argv is bounded: macOS allows ~256KB per argument (E2BIG), and PostToolUse Bash payloads
252
+ // embed the whole tool output — routinely past that limit — so the exec fails before ANY
253
+ // hook runs and the chain verdict is lost. Pass the payload through a temp file instead
254
+ // (runner reads `@path` markers); inline argv stays as the fallback for read-only project
255
+ // roots, and remains unambiguous because raw JSON never starts with '@'.
256
+ function sweepStalePayloadFiles(dir) {
257
+ try {
258
+ const cutoff = Date.now() - 60 * 60 * 1000;
259
+ for (const name of fs.readdirSync(dir)) {
260
+ const filePath = path.join(dir, name);
261
+ try {
262
+ if (fs.statSync(filePath).mtimeMs < cutoff) fs.rmSync(filePath, { force: true });
263
+ } catch { /* raced away — fine */ }
264
+ }
265
+ } catch { /* sweeping is best effort */ }
266
+ }
267
+
268
+ function writePayloadFile(projectRoot, payload) {
269
+ try {
270
+ const dir = path.join(projectRoot, '.ukit', 'storage', 'cache', 'hook-payloads');
271
+ fs.mkdirSync(dir, { recursive: true });
272
+ sweepStalePayloadFiles(dir);
273
+ const filePath = path.join(
274
+ dir,
275
+ `${Date.now()}-${process.pid}-${Math.random().toString(36).slice(2, 8)}.json`,
276
+ );
277
+ fs.writeFileSync(filePath, JSON.stringify(payload), 'utf8');
278
+ return filePath;
279
+ } catch {
280
+ return null;
281
+ }
282
+ }
283
+
251
284
  export async function runScriptChain(
252
285
  pi,
253
286
  scripts,
@@ -264,15 +297,21 @@ export async function runScriptChain(
264
297
  const scriptPaths = scripts.map((scriptName) => path.join(projectRoot, '.claude', 'hooks', scriptName));
265
298
  const nodeExecutable = resolveNodeExecutable();
266
299
  const startedAt = Date.now();
300
+ const payloadFile = writePayloadFile(projectRoot, payload);
301
+ const payloadArg = payloadFile ? `@${payloadFile}` : JSON.stringify(payload);
267
302
  let execResult;
268
303
  try {
269
304
  execResult = await pi.exec(
270
305
  nodeExecutable,
271
- [runnerPath, JSON.stringify(payload), ...scriptPaths],
306
+ [runnerPath, payloadArg, ...scriptPaths],
272
307
  { cwd: projectRoot, timeout: chainExecTimeoutMs(scripts.length) },
273
308
  );
274
309
  } catch (error) {
275
310
  execResult = { code: 1, stdout: '', stderr: error?.message ?? String(error), killed: false };
311
+ } finally {
312
+ if (payloadFile) {
313
+ try { fs.rmSync(payloadFile, { force: true }); } catch { /* best effort cleanup */ }
314
+ }
276
315
  }
277
316
  const elapsedMs = Date.now() - startedAt;
278
317