@ngockhoale/ukit 2.3.2 → 2.3.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,106 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.4 - 2026-09-10
6
+
7
+ Silent-stall audit and compact-recovery release. Two themes: every block reports a
8
+ user-visible reason, and compaction must actually buy back context.
9
+
10
+ Security hooks can no longer be bypassed or silently disabled. `block-dangerous.sh` catches
11
+ split recursive/force flags (`rm -r -f`, `rm --force -r`) and evaluates every delete target
12
+ per shell segment, so one allowlisted word can no longer legitimise a sibling target
13
+ (`rm -rf dist /`); a bare `/` stays classified unsafe, wrapped forms keep the whole-command
14
+ fallback, and detections surface as structured `ask` decisions that never echo the raw
15
+ command. `block-dangerous.sh`, `protect-files.sh`, and `auto-allow-bash.sh` fall back to
16
+ Node when `jq` is absent (stock macOS), so a missing optional binary cannot disable a gate.
17
+ `protect-files.sh` matches exact sensitive basenames and directory components instead of
18
+ substrings — `.env.d.ts` and `app.envrc` are source files, not secrets — and emits a
19
+ structured human-decision response on real blocks. `sensitive-data-guard.sh` only treats
20
+ `-u/--user` as credentials on tools that authenticate with it, so `docker run --user
21
+ 1000:1000` no longer false-blocks while `curl -u user:pass` still does.
22
+
23
+ Stops always carry a reason. `handoff-model-guard.sh` pairs its stderr refusal with a
24
+ structured `systemMessage`; the completion gate reports a lost route state after real
25
+ session activity instead of releasing silently; hard-cap grace and sensitive-data blocks
26
+ already speak `systemMessage` from 2.3.3. `Phase: done …` with trailing status text now
27
+ counts as finished (`^done\b`, not `^done$`) in `handoff-resume.sh`,
28
+ `context-hardcap-gate.sh`, and `context-window-guard.sh`, so annotated done-phases stop
29
+ being resumed forever. Multi-file resume intents survive compaction: identity is session +
30
+ prompt key, because the router re-keys `requestKey` on every tool call and that churn is
31
+ what made post-compact resume a silent no-op.
32
+
33
+ The hard-cap grace budget is race-proof. Parallel subagents hitting the cap in one episode
34
+ used to race a single read-modify-write counter and multiply the budget; each spent call
35
+ now claims an exclusive-create sentinel file, so the sentinels present are the true count
36
+ and one claim can never be counted twice (verified by an 8-process parallel test: exactly
37
+ the configured grants, the next call blocked).
38
+
39
+ omp payloads moved off argv. PostToolUse Bash payloads embed whole tool outputs and exceed
40
+ macOS's ~256KB-per-argument limit, failing exec with E2BIG before any hook ran. The bridge
41
+ now writes the payload to a temp file (`@path` marker, hourly sweep, inline fallback for
42
+ read-only roots) and `hook-chain-runner.mjs` reads the marker; raw JSON never starts with
43
+ `@`, so both forms stay unambiguous.
44
+
45
+ Compact recovery, directed. Near-cap advisories now order the work: land the smallest
46
+ finishable item end-to-end, defer the rest as one-line notes to disk, delegate broad
47
+ reads/searches to subagents, then ask the user to compact. A new post-compact check fires
48
+ once per compact boundary when the live context after the boundary is still ≥60% of the
49
+ cap: stop re-reading old files (the fastest way to re-inflate a just-compacted session),
50
+ delegate, or finish in a fresh session — UKit resume state carries the goal across with a
51
+ near-empty window. The reinjected PROJECT CONTEXT block carries the same discipline.
52
+
53
+ Also fixed: `ukit install|diff --with-codegraph` crashed before the flag was consumed;
54
+ `route-task.mjs` helper commands use POSIX single-quote escaping instead of `JSON.stringify`
55
+ (printed commands were open to `$(...)`/backtick expansion) and reject absolute outside-root
56
+ targets instead of routing foreign-file context.
57
+
58
+ ## 2.3.3 - 2026-09-10
59
+
60
+ False-stop and long-context release. Four fixes that share one theme: UKit must run a prompt
61
+ to a finished result, and any stop must say why.
62
+
63
+ Session-bound completion gate. Project-scoped route state could outlive the session-scoped
64
+ execution ledger, so a new session could inherit an older session's completion contract and
65
+ stop with a false `missing write-evidence` message. `skill-router.sh` now binds fresh, cached,
66
+ and route-less state to `sessionId`; `execution-ledger.mjs` accepts route state only when its
67
+ session matches the hook payload and ignores unbound legacy state for session-aware payloads.
68
+ Same-session gating and prompt-scoped evidence carry remain intact. Regression coverage now
69
+ covers mismatched sessions, unbound legacy state, and prior-session route isolation.
70
+
71
+ Orchestration envelopes no longer corrupt the active route. `<task-notification>`,
72
+ `<tool-result>`, `<system-reminder>`, and `<local-command-caveat>…</local-command-stdout>`
73
+ payloads delivered through the prompt boundary were treated as explicit user prompts,
74
+ rewriting route/cache/audit state mid-run. `skill-router.sh` now ignores complete envelopes
75
+ before routing (fail-open, opening-tag attributes accepted); router exceptions surface on
76
+ stderr instead of disappearing. Related ownership fixes: Stop evaluation loads route state
77
+ with the current session payload, execution-ledger/route-task reject incompatible explicit
78
+ session ownership without inventing identity, subagent tool calls cannot rewrite the main
79
+ route, and cache entries are not touched before ownership is established.
80
+
81
+ Ordinary long-context tasks preserve and resume. A natural-language routed task with no
82
+ `docs/AI_HANDOFF/RUN.md` previously hit the hard-cap gate with no landing window and no
83
+ resume cursor after compaction. Ordinary ledger state now gets the existing finite
84
+ `hardCapGraceCalls` landing budget (only with explicit session ownership, matching
85
+ session/prompt/request identity, unfinished evidence, and no blocker); the first grace call
86
+ writes a 30-minute hash-only resume intent that `handoff-resume.sh` consumes once and
87
+ emits as a continuation instruction with no raw prompt/command data. The
88
+ `context-window-guard.sh` fallback moved from 220k to the shipped 500k default.
89
+
90
+ New `vibecode` autonomy level. One typed prompt must run to a finished result: with
91
+ `autonomy.level: vibecode`, the completion gate opts out of the continuation cap entirely and
92
+ keeps blocking (including the reentrant `stop_hook_active` Stop) until completion evidence
93
+ arrives or a genuine blocker is recorded — every block carries its reason. `deriveTaskRoute`
94
+ now propagates `autonomyLevel` and a `continuousExecution` policy
95
+ (`{ enabled, stopOn, maxContinuations }`) into `routeSummary` for every level, and
96
+ `validateRuntimeConfig` accepts the fourth level.
97
+
98
+ Dangerous commands are a human decision, not a silent block. `block-dangerous.sh` now emits a
99
+ structured `permissionDecision: "ask"` (PreToolUse JSON) with the raw command scrubbed from
100
+ all output — arguments and comments can carry secrets — while exit 2 still refuses the call
101
+ for harnesses that ignore structured output. The omp bridge chain layer translates the same
102
+ structured decision into an `ask` verdict; because omp has no native ask, its host boundary
103
+ surfaces it as a block carrying the decision reason.
104
+
5
105
  ## 2.3.2 - 2026-09-09
6
106
 
7
107
  Docs-contract release. "Every stop says why — no silent idle" is now a shipping Execution
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.2",
3
+ "version": "2.3.4",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -15,7 +15,7 @@ export async function runDiff({ packageRoot, projectRoot, packageVersion, argv =
15
15
  return;
16
16
  }
17
17
 
18
- const toolsArg = parseToolsArg(argv);
18
+ const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
19
19
  const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
20
20
  const selectedAdapterItemIds = toSelectedAdapterItemIds(selectedOptionalTools);
21
21
  const withCodegraph = argv.includes('--with-codegraph');
@@ -197,7 +197,10 @@ export async function pruneDeselectedAdapters({
197
197
  }
198
198
 
199
199
  export async function runInstall({ packageRoot, projectRoot, packageVersion, argv = [] }) {
200
- const toolsArg = parseToolsArg(argv);
200
+ // parseToolsArg owns only --tools and throws on anything else, so hand it the argv with
201
+ // the flags consumed here (--with-codegraph) already removed — order-dependent parsing
202
+ // used to make the documented `ukit install --with-codegraph` fail 100% of the time.
203
+ const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
201
204
  const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
202
205
  const withCodegraph = argv.includes('--with-codegraph');
203
206
 
@@ -257,7 +257,7 @@ export function validateRuntimeConfig(config) {
257
257
  errors.push('version must be a non-empty string.');
258
258
  }
259
259
 
260
- const VALID_AUTONOMY_LEVELS = new Set(['conservative', 'balanced', 'free-run']);
260
+ const VALID_AUTONOMY_LEVELS = new Set(['conservative', 'balanced', 'free-run', 'vibecode']);
261
261
  if (!isPlainObject(config.autonomy)) {
262
262
  errors.push('autonomy must be an object.');
263
263
  } else {
@@ -240,6 +240,7 @@ export function buildRouteSummary({
240
240
  executionCandidates,
241
241
  });
242
242
  const executionContract = buildExecutionContract(executionMode);
243
+ const continuousExecution = buildContinuousExecutionPolicy(autonomyLevel);
243
244
  const postEditReview = executionContract?.postEditReviewPolicy
244
245
  ? { policy: executionContract.postEditReviewPolicy, agent: 'code-reviewer', reviewTargetType: 'diff' }
245
246
  : null;
@@ -290,6 +291,8 @@ export function buildRouteSummary({
290
291
  completionState,
291
292
  postEditReview,
292
293
  continuationState,
294
+ autonomyLevel,
295
+ continuousExecution,
293
296
  intentMode: routingContext.intentMode ?? null,
294
297
  handoffFile,
295
298
  handoffBudget,
@@ -303,6 +306,19 @@ export function buildRouteSummary({
303
306
  };
304
307
  }
305
308
 
309
+ // Execution-ledger.mjs mirrors this policy for its Stop gate; the two must stay aligned.
310
+ // MAX_CONTINUATIONS (6) below is the same constant the ledger enforces for bounded modes.
311
+ function buildContinuousExecutionPolicy(autonomyLevel) {
312
+ if (autonomyLevel === 'vibecode') {
313
+ return {
314
+ enabled: true,
315
+ stopOn: ['completion-evidence', 'genuine-blocker', 'dangerous-command-decision'],
316
+ maxContinuations: null,
317
+ };
318
+ }
319
+ return { enabled: false, maxContinuations: 6 };
320
+ }
321
+
306
322
  function deriveContextMode(taskType) {
307
323
  if (taskType === 'trivial' || taskType === 'simple') return 'LITE';
308
324
  if (taskType === 'non-trivial' || taskType === 'shared-simple') return 'FULL';
@@ -3,7 +3,22 @@
3
3
  # Optimized: fast-path grep to skip Node.js when rule already exists.
4
4
 
5
5
  INPUT=$(cat)
6
- COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty' 2>/dev/null)
6
+ # jq is absent on stock macOS. Keep jq as the low-latency normal path, but fall back to
7
+ # UKit's required Node runtime so safe commands still receive managed metadata.
8
+ if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
9
+ COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
10
+ else
11
+ COMMAND=$(printf '%s' "$INPUT" | node -e '
12
+ const chunks = [];
13
+ process.stdin.on("data", (chunk) => chunks.push(chunk));
14
+ process.stdin.on("end", () => {
15
+ try {
16
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
17
+ process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
18
+ } catch {}
19
+ });
20
+ ' 2>/dev/null)
21
+ fi
7
22
 
8
23
  if [ -z "$COMMAND" ]; then
9
24
  exit 0
@@ -3,7 +3,22 @@
3
3
  # Matched on: Bash
4
4
 
5
5
  INPUT=$(cat)
6
- COMMAND=$(echo "$INPUT" | jq -r '.tool_input.command // empty')
6
+ # jq is not installed on stock macOS. Keep jq as the low-latency normal path, but use
7
+ # UKit's required Node runtime as a fallback so its absence cannot silently disable the gate.
8
+ if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
9
+ COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
10
+ else
11
+ COMMAND=$(printf '%s' "$INPUT" | node -e '
12
+ const chunks = [];
13
+ process.stdin.on("data", (chunk) => chunks.push(chunk));
14
+ process.stdin.on("end", () => {
15
+ try {
16
+ const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
17
+ process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
18
+ } catch {}
19
+ });
20
+ ' 2>/dev/null)
21
+ fi
7
22
 
8
23
  if [ -z "$COMMAND" ]; then
9
24
  exit 0
@@ -42,31 +57,97 @@ DANGEROUS_PATTERNS=(
42
57
  "dd if=/dev/"
43
58
  )
44
59
 
60
+ # Dangerous detections surface as a structured `ask` decision (stdout JSON) so a human
61
+ # decides; exit 2 still refuses the call for harnesses that ignore structured output.
62
+ # The raw command and the matched pattern are NEVER echoed — arguments and comments can
63
+ # carry secrets, and the pattern text itself restates the dangerous command.
64
+ emit_dangerous_decision() {
65
+ REASON="$1"
66
+ printf '%s\n' "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"ask\",\"permissionDecisionReason\":\"$REASON\"}}"
67
+ echo "BLOCKED: $REASON" >&2
68
+ exit 2
69
+ }
70
+
45
71
  for pattern in "${DANGEROUS_PATTERNS[@]}"; do
46
72
  if echo "$SCAN_COMMAND" | grep -qE "$pattern"; then
47
- echo "BLOCKED: Dangerous command detected matching pattern '$pattern'. This command could cause irreversible damage. Ask the user to run it manually if truly needed." >&2
48
- exit 2
73
+ emit_dangerous_decision "Dangerous command pattern detected; it could cause irreversible damage. UKit defers this to a human decision."
49
74
  fi
50
75
  done
51
76
 
52
- # Handle rm commands: allow safe cleanup targets, block risky recursive force-deletes
53
- if echo "$SCAN_COMMAND" | grep -qE "(^|[;&|[:space:]])rm([[:space:]]|$)"; then
54
- if echo "$SCAN_COMMAND" | grep -qE "rm\s+(-[a-zA-Z]*r[a-zA-Z]*f|-[a-zA-Z]*f[a-zA-Z]*r|--recursive\s+--force|--force\s+--recursive|--recursive\s+-f|-f\s+--recursive)"; then
55
- SAFE_DELETE_REGEX='(^|[[:space:]])(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/|[[:space:]]|$)'
56
- UNSAFE_TARGET_REGEX='(^|[[:space:]])(/|~|\.\.?($|/)|\.($|/))'
77
+ # Handle rm commands: allow safe cleanup targets only, defer everything else to a human.
78
+ # Direct rm invocations are parsed per shell segment (`;`, `&`, `|` split) with per-target
79
+ # allowlist evaluation: split flags (`rm -r -f`, `rm --force -r`) register, and one
80
+ # allowlisted word can never legitimise a sibling target (`rm -rf dist /`). A whole-command
81
+ # fallback still catches rm inside wrapped/quoted contexts (`bash -c "rm -r -f /x"`) and
82
+ # flags placed after targets (`rm src -rf`), where segment parsing cannot see rm as a head.
83
+ SAFE_DELETE_REGEX='(^|[[:space:]])(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/|[[:space:]]|$)'
84
+ UNSAFE_TARGET_REGEX='(^|[[:space:]])(/|~|\.\.?($|/)|\.($|/))'
85
+ SAFE_ONE_TARGET_REGEX='^(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/.*)?$'
86
+ UNSAFE_ONE_TARGET_REGEX='^(/.*|~|~/.*|\.\.|\.\./.*|\.)$'
87
+ RM_WORD_REGEX=$'(^|[;&|[:space:]"\x27])rm([[:space:]"\x27]|$)'
57
88
 
58
- # Check the allowlist first: `./dist` legitimately contains a leading "./" that would
59
- # otherwise be mistaken for a bare current-directory target by UNSAFE_TARGET_REGEX below.
60
- if echo "$SCAN_COMMAND" | grep -qE "$SAFE_DELETE_REGEX"; then
89
+ RM_VERDICT_UNSAFE=0
90
+ RM_VERDICT_GENERIC=0
91
+ RM_DIRECT_SEEN=0
92
+ set -f
93
+ while IFS= read -r segment; do
94
+ [ -n "$segment" ] || continue
95
+ seg_head=$(printf '%s' "$segment" | sed -E 's/^[[:space:]]*//' | awk '{print $1}' | xargs -I{} basename {} 2>/dev/null)
96
+ [ "$seg_head" = "rm" ] || continue
97
+ seg_flags=$(printf '%s' "$segment" | awk '{for(i=2;i<=NF;i++){if($i=="--")break;if($i ~ /^-/)printf "%s\n",$i;else break}}')
98
+ seg_has_r=0
99
+ seg_has_f=0
100
+ for flag in $seg_flags; do
101
+ case "$flag" in
102
+ --recursive) seg_has_r=1 ;;
103
+ --force) seg_has_f=1 ;;
104
+ --*) ;;
105
+ -*) case "$flag" in *[rR]*) seg_has_r=1 ;; esac
106
+ case "$flag" in *f*) seg_has_f=1 ;; esac ;;
107
+ esac
108
+ done
109
+ [ "$seg_has_r" -eq 1 ] && [ "$seg_has_f" -eq 1 ] || continue
110
+ RM_DIRECT_SEEN=1
111
+ seg_targets=$(printf '%s' "$segment" | awk '{saw=0;for(i=2;i<=NF;i++){if(!saw){if($i=="--"){saw=1;continue}if($i ~ /^-/)continue;saw=1;printf "%s\n",$i}else if($i !~ /^-/)printf "%s\n",$i}}')
112
+ seg_all_safe=1
113
+ seg_any_unsafe=0
114
+ for target in $seg_targets; do
115
+ # Strip ONE trailing slash so `dist/` matches the allowlist — but never shrink a
116
+ # bare `/` (or `~`) to an empty string, which would silently dodge the unsafe check.
117
+ if [ "${#target}" -gt 1 ]; then target="${target%/}"; fi
118
+ if printf '%s' "$target" | grep -qE "$UNSAFE_ONE_TARGET_REGEX"; then
119
+ seg_any_unsafe=1
120
+ seg_all_safe=0
121
+ elif printf '%s' "$target" | grep -qE "$SAFE_ONE_TARGET_REGEX"; then
61
122
  :
62
- elif echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
63
- echo "BLOCKED: Unsafe delete target detected. Refusing recursive force-delete." >&2
64
- exit 2
65
123
  else
66
- echo "BLOCKED: 'rm -rf' is only auto-allowed for safe cleanup targets (dist/build/coverage/.next/.nuxt/.turbo/tmp/.cache/node_modules)." >&2
67
- exit 2
124
+ seg_all_safe=0
68
125
  fi
126
+ done
127
+ if [ "$seg_any_unsafe" -eq 1 ]; then RM_VERDICT_UNSAFE=1
128
+ elif [ "$seg_all_safe" -eq 0 ]; then RM_VERDICT_GENERIC=1
129
+ fi
130
+ done <<SEGMENTS
131
+ $(printf '%s' "$SCAN_COMMAND" | tr ';&|' '\n\n\n')
132
+ SEGMENTS
133
+ set +f
134
+
135
+ if [ "$RM_DIRECT_SEEN" -eq 0 ] && echo "$SCAN_COMMAND" | grep -qE "$RM_WORD_REGEX" \
136
+ && echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--recursive|(^|[[:space:]])-[a-zA-Z]*[rR]" \
137
+ && echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--force|(^|[[:space:]])-[a-zA-Z]*f"; then
138
+ # Unsafe wins over the allowlist here: a safe word elsewhere in the line must never
139
+ # legitimise an absolute/home/parent target we can also see.
140
+ if echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
141
+ RM_VERDICT_UNSAFE=1
142
+ elif ! echo "$SCAN_COMMAND" | grep -qE "$SAFE_DELETE_REGEX"; then
143
+ RM_VERDICT_GENERIC=1
69
144
  fi
70
145
  fi
71
146
 
147
+ if [ "$RM_VERDICT_UNSAFE" -eq 1 ]; then
148
+ emit_dangerous_decision "Unsafe delete target detected; recursive force-delete could cause irreversible damage. UKit defers this to a human decision."
149
+ elif [ "$RM_VERDICT_GENERIC" -eq 1 ]; then
150
+ emit_dangerous_decision "Recursive force-delete outside safe cleanup targets (dist/build/coverage/.next/.nuxt/.turbo/tmp/.cache/node_modules) is destructive. UKit defers this to a human decision."
151
+ fi
152
+
72
153
  exit 0
@@ -1,9 +1,9 @@
1
1
  #!/bin/bash
2
2
  # PreToolUse hook: hard-enforce an absolute context token cap (compact.hardCapTokens,
3
- # default 220000, sized for a 256k window — must stay below the model's real context
4
- # window, so lower it on a 200k model), separate from the
3
+ # default 500000, sized for a 1M window — must stay below the model's real context
4
+ # window, so lower it on a 200k/256k model), separate from the
5
5
  # soft/hard advisory pressure phases in
6
- # compact-threshold.mjs (default soft=50000/hard=80000, which only print a suggestion).
6
+ # compact-threshold.mjs (default soft=150000/hard=240000, which only print a suggestion).
7
7
  #
8
8
  # Those advisory phases are just injected text — nothing stops the agent from ignoring
9
9
  # them and letting a session run to hundreds of thousands of tokens with no compaction.
@@ -18,12 +18,13 @@
18
18
  # real compaction.
19
19
  #
20
20
  # Grace window (compact.hardCapGraceCalls, default 10): blocking the very first tool call
21
- # past the cap strands a handoff run mid-edit — files half-written, nothing committed, no
22
- # cursor — which is strictly worse than letting it land. When docs/AI_HANDOFF/RUN.md shows
23
- # an unfinished run, the gate therefore allows a BOUNDED number of further calls so the run
24
- # can commit, write its cursor and push; after that it blocks exactly as before. The budget
25
- # is per over-cap episode, not per wave: it only resets when the estimate actually drops
26
- # (i.e. a real compaction happened), so a run cannot mint itself fresh grace forever.
21
+ # past the cap strands a handoff run or ordinary routed task mid-edit — files half-written,
22
+ # no durable completion state — which is strictly worse than letting it land. When RUN.md
23
+ # shows an unfinished handoff run, or the shared ledger proves an unfinished session-bound
24
+ # ordinary task, the gate allows a BOUNDED number of further calls so it can land the current
25
+ # edit/verification and persist resume state; after that it blocks exactly as before. The
26
+ # budget is per over-cap episode, not per wave: it only resets when the estimate actually
27
+ # drops (i.e. a real compaction happened), so a task cannot mint itself fresh grace forever.
27
28
  #
28
29
  # Config toggle: compact.hardCapBlock (default true). Set to false only to debug this
29
30
  # gate itself; it must not become a normal escape hatch.
@@ -74,7 +75,7 @@ function readRunCursor() {
74
75
  try {
75
76
  const text = fs.readFileSync(path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md'), 'utf8');
76
77
  const runPhase = (text.match(/^Phase:\s*(.+)$/m)?.[1] || '').trim();
77
- if (!runPhase || /^done$/i.test(runPhase)) return null;
78
+ if (!runPhase || /^done\b/i.test(runPhase)) return null;
78
79
  return {
79
80
  phase: runPhase,
80
81
  cursor: (text.match(/^Cursor:\s*(.+)$/m)?.[1] || '').trim(),
@@ -98,6 +99,10 @@ function readRunCursor() {
98
99
  }
99
100
 
100
101
  const mod = await import(pathToFileURL(thresholdModulePath).href);
102
+ const ledgerModulePath = path.join(hookDir, '..', 'ukit', 'runtime', 'execution-ledger.mjs');
103
+ const ledgerMod = fs.existsSync(ledgerModulePath)
104
+ ? await import(pathToFileURL(ledgerModulePath).href)
105
+ : null;
101
106
  const pressurePath = path.join(projectRoot, '.ukit', 'storage', 'cache', 'compact-pressure.json');
102
107
  const rawState = readJsonSafe(pressurePath, null);
103
108
 
@@ -117,49 +122,134 @@ function readRunCursor() {
117
122
  }
118
123
  const thresholds = mod.buildCompactThresholds(config);
119
124
 
125
+ // The grace budget must survive parallel subagents hitting the cap in the same episode.
126
+ // A read-modify-write of one JSON file under-counts (two processes both read used=3 and
127
+ // both write 4, silently multiplying the budget), so each spent call is claimed with an
128
+ // exclusive-create sentinel file (openSync 'wx'): the sentinels present ARE the number
129
+ // of grace calls spent, and one claim can never be counted twice.
120
130
  const gracePath = path.join(projectRoot, '.ukit', 'storage', 'cache', 'hardcap-grace.json');
131
+ const graceDir = path.dirname(gracePath);
132
+ const slotPath = (n) => path.join(graceDir, `hardcap-grace.json.${n}`);
133
+ const readSlots = () => {
134
+ try {
135
+ return fs.readdirSync(graceDir)
136
+ .map((name) => /^hardcap-grace\.json\.(\d+)$/.exec(name))
137
+ .filter(Boolean)
138
+ .map((m) => Number(m[1]))
139
+ .sort((a, b) => a - b);
140
+ } catch {
141
+ return [];
142
+ }
143
+ };
144
+ const clearGrace = () => {
145
+ try { fs.rmSync(gracePath, { force: true }); } catch { /* advisory only */ }
146
+ for (const n of readSlots()) {
147
+ try { fs.rmSync(slotPath(n), { force: true }); } catch { /* advisory only */ }
148
+ }
149
+ };
121
150
 
122
151
  if (state.estimatedTotalTokens < thresholds.hardCapTokens) {
123
152
  // Back under the cap — the episode is over, so the next one starts with a full budget.
124
- fs.rmSync(gracePath, { force: true });
153
+ clearGrace();
125
154
  process.exit(0);
126
155
  return;
127
156
  }
128
157
 
129
158
  const run = readRunCursor();
130
- if (run) {
159
+ let ordinaryTask = null;
160
+ if (!run && ledgerMod) {
161
+ try {
162
+ ordinaryTask = await ledgerMod.readResumableExecution(projectRoot, payload);
163
+ } catch {
164
+ ordinaryTask = null;
165
+ }
166
+ }
167
+ const resumable = run || ordinaryTask;
168
+ if (resumable) {
131
169
  const graceCalls = Number.isFinite(config?.compact?.hardCapGraceCalls)
132
170
  ? config.compact.hardCapGraceCalls
133
171
  : 10;
134
- const prior = readJsonSafe(gracePath, null);
172
+
173
+ let slots = readSlots();
174
+ let startedAtTokens = null;
175
+ for (const n of slots) {
176
+ const data = readJsonSafe(slotPath(n), null);
177
+ if (data && Number.isFinite(data.startedAtTokens)) {
178
+ startedAtTokens = Math.max(startedAtTokens ?? -Infinity, data.startedAtTokens);
179
+ }
180
+ }
135
181
  // Reset only when the estimate actually fell since grace started: that is the signal a
136
182
  // real compaction/shed happened. Advancing the cursor alone must NOT top the budget up,
137
183
  // otherwise a long run gets unlimited grace and the ceiling stops meaning anything.
138
- const carried =
139
- prior && Number.isFinite(prior.startedAtTokens) && state.estimatedTotalTokens >= prior.startedAtTokens
140
- ? prior
141
- : { startedAtTokens: state.estimatedTotalTokens, used: 0 };
184
+ if (startedAtTokens !== null && state.estimatedTotalTokens < startedAtTokens) {
185
+ clearGrace();
186
+ slots = [];
187
+ startedAtTokens = null;
188
+ }
189
+ const spent = slots.length > 0 ? slots[slots.length - 1] : 0;
190
+ if (startedAtTokens === null) startedAtTokens = state.estimatedTotalTokens;
142
191
 
143
- if (carried.used < graceCalls) {
144
- const used = carried.used + 1;
192
+ if (spent < graceCalls) {
193
+ // Claim the next slot atomically: EEXIST means a parallel worker already claimed it,
194
+ // step to the next. A filesystem error keeps the old fail-open behavior — losing the
195
+ // counter must never block a landing — at the cost of one extra granted call.
196
+ let used = 0;
197
+ let counterFailed = false;
145
198
  try {
146
- fs.mkdirSync(path.dirname(gracePath), { recursive: true });
147
- fs.writeFileSync(gracePath, JSON.stringify({ ...carried, used }));
199
+ fs.mkdirSync(graceDir, { recursive: true });
200
+ for (let n = spent + 1; n <= graceCalls; n += 1) {
201
+ try {
202
+ const fd = fs.openSync(slotPath(n), 'wx');
203
+ fs.writeSync(fd, JSON.stringify({ startedAtTokens, used: n, ts: Date.now() }));
204
+ fs.closeSync(fd);
205
+ used = n;
206
+ break;
207
+ } catch (err) {
208
+ if (err?.code === 'EEXIST') continue;
209
+ throw err;
210
+ }
211
+ }
148
212
  } catch {
149
- // Losing the counter must not block the run; worst case grace restarts.
213
+ counterFailed = true;
214
+ }
215
+ if (used === 0 && counterFailed) used = spent + 1;
216
+ if (used > 0 && used <= graceCalls) {
217
+ try {
218
+ fs.writeFileSync(gracePath, JSON.stringify({ startedAtTokens, used }));
219
+ if (ordinaryTask && ledgerMod && used === 1) {
220
+ await ledgerMod.writeResumeIntent(projectRoot, payload);
221
+ }
222
+ } catch {
223
+ // Losing the counter or advisory resume intent must not block the run.
224
+ }
225
+ const landing = run
226
+ ? `An unfinished handoff run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`
227
+ : 'An unfinished ordinary routed task is in flight, so this call is allowed instead of stranding it before compaction.';
228
+ process.stderr.write(
229
+ [
230
+ `CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
231
+ landing,
232
+ run
233
+ ? 'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.'
234
+ : 'Spend the remaining grace on LANDING, not on new work: finish the current mutation or verification, then compact as soon as the host allows it.',
235
+ 'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
236
+ 'Tell the user in your reply that context is over the cap and they should run /compact as soon as this landing step is complete.',
237
+ 'After compaction, the SessionStart resume hook replays the safe cursor and the task continues automatically.',
238
+ ].join('\n') + '\n',
239
+ );
240
+ // The stderr contract above only reaches the model — if the turn dies before the
241
+ // model relays it, the user sees nothing. Mirror the first grace call as a
242
+ // structured systemMessage so the /compact advice is user-visible regardless.
243
+ if (used === 1) {
244
+ process.stdout.write(`${JSON.stringify({
245
+ systemMessage: `UKit: context is over the hard cap (~${state.estimatedTotalTokens} tokens). A short landing window (${graceCalls} calls) is active — run /compact when it ends; the task resumes automatically after compaction.`,
246
+ })}\n`);
247
+ }
248
+ process.exit(0);
249
+ return;
150
250
  }
151
- process.stderr.write(
152
- [
153
- `CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
154
- `An unfinished run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`,
155
- 'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.',
156
- 'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
157
- 'Tell the user in your reply that context is over the cap and they should run /compact as soon as this run lands.',
158
- 'After the user runs /compact, the SessionStart resume hook replays the cursor and the run continues automatically.',
159
- ].join('\n') + '\n',
160
- );
161
- process.exit(0);
162
- return;
251
+ // Every slot was claimed while we raced — the budget is genuinely spent; fall
252
+ // through to the hard block below.
163
253
  }
164
254
  }
165
255
 
@@ -169,12 +259,17 @@ function readRunCursor() {
169
259
  'Typing "continue" alone will hit this same block; only /compact (or a new session) resets the counter. After compacting, resume the interrupted task.',
170
260
  'This is an absolute ceiling (compact.hardCapTokens), separate from the soft/hard advisory phases — those were apparently not followed.',
171
261
  `Edit/Write/Bash refused (tool_name=${toolName}) until real compaction happens.`,
172
- run
262
+ resumable
173
263
  ? 'The unfinished-run grace window (compact.hardCapGraceCalls) is already exhausted — progress should be committed and the cursor written by now.'
174
- : 'No unfinished run in docs/AI_HANDOFF/RUN.md, so there is no grace window to spend.',
264
+ : 'No resumable unfinished routed task was found, so there is no grace window to spend.',
175
265
  'Do not work around this by summarizing inline and continuing, and do not reach for a non-gated write tool.',
176
266
  ];
177
267
  process.stderr.write(`${lines.join('\n')}\n`);
268
+ // stderr reaches the model, not reliably the user. Pair the hard block with a
269
+ // structured systemMessage so "why did it stop" always has a user-visible answer.
270
+ process.stdout.write(`${JSON.stringify({
271
+ systemMessage: `UKit blocked ${toolName}: context is over the hard cap (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}) and the landing window is spent. Run /compact now; after compaction the task continues automatically.`,
272
+ })}\n`);
178
273
  process.exit(2);
179
274
  })().catch((err) => {
180
275
  // A logic error here fails OPEN: this is a backstop on top of advisory nudges,