@ngockhoale/ukit 2.3.2 → 2.3.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +100 -0
- package/package.json +1 -1
- package/src/cli/commands/diff.js +1 -1
- package/src/cli/commands/install.js +4 -1
- package/src/core/runtimeConfig.js +1 -1
- package/src/index/taskRouting.js +16 -0
- package/templates/.claude/hooks/auto-allow-bash.sh +16 -1
- package/templates/.claude/hooks/block-dangerous.sh +97 -16
- package/templates/.claude/hooks/context-hardcap-gate.sh +131 -36
- package/templates/.claude/hooks/context-window-guard.sh +33 -5
- package/templates/.claude/hooks/handoff-model-guard.sh +7 -0
- package/templates/.claude/hooks/handoff-resume.sh +71 -40
- package/templates/.claude/hooks/protect-files.sh +42 -18
- package/templates/.claude/hooks/sensitive-data-guard.sh +18 -2
- package/templates/.claude/hooks/skill-router.sh +81 -4
- package/templates/.claude/ukit/index/route-task.mjs +90 -23
- package/templates/.claude/ukit/runtime/execution-ledger.mjs +149 -11
- package/templates/.claude/ukit/runtime/hook-chain-runner.mjs +12 -1
- package/templates/.claude/ukit/runtime/reinject-context.mjs +4 -0
- package/templates/.omp/hooks/pre/ukit-bridge.js +83 -5
- package/templates/AGENTS.md +1 -1
- package/templates/CLAUDE.md +1 -1
- package/templates/docs/AI_HANDOFF/RULES.md +2 -2
- package/templates/ukit/storage/config.json +2 -2
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,106 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.3.4 - 2026-09-10
|
|
6
|
+
|
|
7
|
+
Silent-stall audit and compact-recovery release. Two themes: every block reports a
|
|
8
|
+
user-visible reason, and compaction must actually buy back context.
|
|
9
|
+
|
|
10
|
+
Security hooks can no longer be bypassed or silently disabled. `block-dangerous.sh` catches
|
|
11
|
+
split recursive/force flags (`rm -r -f`, `rm --force -r`) and evaluates every delete target
|
|
12
|
+
per shell segment, so one allowlisted word can no longer legitimise a sibling target
|
|
13
|
+
(`rm -rf dist /`); a bare `/` stays classified unsafe, wrapped forms keep the whole-command
|
|
14
|
+
fallback, and detections surface as structured `ask` decisions that never echo the raw
|
|
15
|
+
command. `block-dangerous.sh`, `protect-files.sh`, and `auto-allow-bash.sh` fall back to
|
|
16
|
+
Node when `jq` is absent (stock macOS), so a missing optional binary cannot disable a gate.
|
|
17
|
+
`protect-files.sh` matches exact sensitive basenames and directory components instead of
|
|
18
|
+
substrings — `.env.d.ts` and `app.envrc` are source files, not secrets — and emits a
|
|
19
|
+
structured human-decision response on real blocks. `sensitive-data-guard.sh` only treats
|
|
20
|
+
`-u/--user` as credentials on tools that authenticate with it, so `docker run --user
|
|
21
|
+
1000:1000` no longer false-blocks while `curl -u user:pass` still does.
|
|
22
|
+
|
|
23
|
+
Stops always carry a reason. `handoff-model-guard.sh` pairs its stderr refusal with a
|
|
24
|
+
structured `systemMessage`; the completion gate reports a lost route state after real
|
|
25
|
+
session activity instead of releasing silently; hard-cap grace and sensitive-data blocks
|
|
26
|
+
already speak `systemMessage` from 2.3.3. `Phase: done …` with trailing status text now
|
|
27
|
+
counts as finished (`^done\b`, not `^done$`) in `handoff-resume.sh`,
|
|
28
|
+
`context-hardcap-gate.sh`, and `context-window-guard.sh`, so annotated done-phases stop
|
|
29
|
+
being resumed forever. Multi-file resume intents survive compaction: identity is session +
|
|
30
|
+
prompt key, because the router re-keys `requestKey` on every tool call and that churn is
|
|
31
|
+
what made post-compact resume a silent no-op.
|
|
32
|
+
|
|
33
|
+
The hard-cap grace budget is race-proof. Parallel subagents hitting the cap in one episode
|
|
34
|
+
used to race a single read-modify-write counter and multiply the budget; each spent call
|
|
35
|
+
now claims an exclusive-create sentinel file, so the sentinels present are the true count
|
|
36
|
+
and one claim can never be counted twice (verified by an 8-process parallel test: exactly
|
|
37
|
+
the configured grants, the next call blocked).
|
|
38
|
+
|
|
39
|
+
omp payloads moved off argv. PostToolUse Bash payloads embed whole tool outputs and exceed
|
|
40
|
+
macOS's ~256KB-per-argument limit, failing exec with E2BIG before any hook ran. The bridge
|
|
41
|
+
now writes the payload to a temp file (`@path` marker, hourly sweep, inline fallback for
|
|
42
|
+
read-only roots) and `hook-chain-runner.mjs` reads the marker; raw JSON never starts with
|
|
43
|
+
`@`, so both forms stay unambiguous.
|
|
44
|
+
|
|
45
|
+
Compact recovery, directed. Near-cap advisories now order the work: land the smallest
|
|
46
|
+
finishable item end-to-end, defer the rest as one-line notes to disk, delegate broad
|
|
47
|
+
reads/searches to subagents, then ask the user to compact. A new post-compact check fires
|
|
48
|
+
once per compact boundary when the live context after the boundary is still ≥60% of the
|
|
49
|
+
cap: stop re-reading old files (the fastest way to re-inflate a just-compacted session),
|
|
50
|
+
delegate, or finish in a fresh session — UKit resume state carries the goal across with a
|
|
51
|
+
near-empty window. The reinjected PROJECT CONTEXT block carries the same discipline.
|
|
52
|
+
|
|
53
|
+
Also fixed: `ukit install|diff --with-codegraph` crashed before the flag was consumed;
|
|
54
|
+
`route-task.mjs` helper commands use POSIX single-quote escaping instead of `JSON.stringify`
|
|
55
|
+
(printed commands were open to `$(...)`/backtick expansion) and reject absolute outside-root
|
|
56
|
+
targets instead of routing foreign-file context.
|
|
57
|
+
|
|
58
|
+
## 2.3.3 - 2026-09-10
|
|
59
|
+
|
|
60
|
+
False-stop and long-context release. Four fixes that share one theme: UKit must run a prompt
|
|
61
|
+
to a finished result, and any stop must say why.
|
|
62
|
+
|
|
63
|
+
Session-bound completion gate. Project-scoped route state could outlive the session-scoped
|
|
64
|
+
execution ledger, so a new session could inherit an older session's completion contract and
|
|
65
|
+
stop with a false `missing write-evidence` message. `skill-router.sh` now binds fresh, cached,
|
|
66
|
+
and route-less state to `sessionId`; `execution-ledger.mjs` accepts route state only when its
|
|
67
|
+
session matches the hook payload and ignores unbound legacy state for session-aware payloads.
|
|
68
|
+
Same-session gating and prompt-scoped evidence carry remain intact. Regression coverage now
|
|
69
|
+
covers mismatched sessions, unbound legacy state, and prior-session route isolation.
|
|
70
|
+
|
|
71
|
+
Orchestration envelopes no longer corrupt the active route. `<task-notification>`,
|
|
72
|
+
`<tool-result>`, `<system-reminder>`, and `<local-command-caveat>…</local-command-stdout>`
|
|
73
|
+
payloads delivered through the prompt boundary were treated as explicit user prompts,
|
|
74
|
+
rewriting route/cache/audit state mid-run. `skill-router.sh` now ignores complete envelopes
|
|
75
|
+
before routing (fail-open, opening-tag attributes accepted); router exceptions surface on
|
|
76
|
+
stderr instead of disappearing. Related ownership fixes: Stop evaluation loads route state
|
|
77
|
+
with the current session payload, execution-ledger/route-task reject incompatible explicit
|
|
78
|
+
session ownership without inventing identity, subagent tool calls cannot rewrite the main
|
|
79
|
+
route, and cache entries are not touched before ownership is established.
|
|
80
|
+
|
|
81
|
+
Ordinary long-context tasks preserve and resume. A natural-language routed task with no
|
|
82
|
+
`docs/AI_HANDOFF/RUN.md` previously hit the hard-cap gate with no landing window and no
|
|
83
|
+
resume cursor after compaction. Ordinary ledger state now gets the existing finite
|
|
84
|
+
`hardCapGraceCalls` landing budget (only with explicit session ownership, matching
|
|
85
|
+
session/prompt/request identity, unfinished evidence, and no blocker); the first grace call
|
|
86
|
+
writes a 30-minute hash-only resume intent that `handoff-resume.sh` consumes once and
|
|
87
|
+
emits as a continuation instruction with no raw prompt/command data. The
|
|
88
|
+
`context-window-guard.sh` fallback moved from 220k to the shipped 500k default.
|
|
89
|
+
|
|
90
|
+
New `vibecode` autonomy level. One typed prompt must run to a finished result: with
|
|
91
|
+
`autonomy.level: vibecode`, the completion gate opts out of the continuation cap entirely and
|
|
92
|
+
keeps blocking (including the reentrant `stop_hook_active` Stop) until completion evidence
|
|
93
|
+
arrives or a genuine blocker is recorded — every block carries its reason. `deriveTaskRoute`
|
|
94
|
+
now propagates `autonomyLevel` and a `continuousExecution` policy
|
|
95
|
+
(`{ enabled, stopOn, maxContinuations }`) into `routeSummary` for every level, and
|
|
96
|
+
`validateRuntimeConfig` accepts the fourth level.
|
|
97
|
+
|
|
98
|
+
Dangerous commands are a human decision, not a silent block. `block-dangerous.sh` now emits a
|
|
99
|
+
structured `permissionDecision: "ask"` (PreToolUse JSON) with the raw command scrubbed from
|
|
100
|
+
all output — arguments and comments can carry secrets — while exit 2 still refuses the call
|
|
101
|
+
for harnesses that ignore structured output. The omp bridge chain layer translates the same
|
|
102
|
+
structured decision into an `ask` verdict; because omp has no native ask, its host boundary
|
|
103
|
+
surfaces it as a block carrying the decision reason.
|
|
104
|
+
|
|
5
105
|
## 2.3.2 - 2026-09-09
|
|
6
106
|
|
|
7
107
|
Docs-contract release. "Every stop says why — no silent idle" is now a shipping Execution
|
package/package.json
CHANGED
package/src/cli/commands/diff.js
CHANGED
|
@@ -15,7 +15,7 @@ export async function runDiff({ packageRoot, projectRoot, packageVersion, argv =
|
|
|
15
15
|
return;
|
|
16
16
|
}
|
|
17
17
|
|
|
18
|
-
const toolsArg = parseToolsArg(argv);
|
|
18
|
+
const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
|
|
19
19
|
const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
|
|
20
20
|
const selectedAdapterItemIds = toSelectedAdapterItemIds(selectedOptionalTools);
|
|
21
21
|
const withCodegraph = argv.includes('--with-codegraph');
|
|
@@ -197,7 +197,10 @@ export async function pruneDeselectedAdapters({
|
|
|
197
197
|
}
|
|
198
198
|
|
|
199
199
|
export async function runInstall({ packageRoot, projectRoot, packageVersion, argv = [] }) {
|
|
200
|
-
|
|
200
|
+
// parseToolsArg owns only --tools and throws on anything else, so hand it the argv with
|
|
201
|
+
// the flags consumed here (--with-codegraph) already removed — order-dependent parsing
|
|
202
|
+
// used to make the documented `ukit install --with-codegraph` fail 100% of the time.
|
|
203
|
+
const toolsArg = parseToolsArg(argv.filter((arg) => arg !== '--with-codegraph'));
|
|
201
204
|
const selectedOptionalTools = resolveOptionalToolKeys(toolsArg);
|
|
202
205
|
const withCodegraph = argv.includes('--with-codegraph');
|
|
203
206
|
|
|
@@ -257,7 +257,7 @@ export function validateRuntimeConfig(config) {
|
|
|
257
257
|
errors.push('version must be a non-empty string.');
|
|
258
258
|
}
|
|
259
259
|
|
|
260
|
-
const VALID_AUTONOMY_LEVELS = new Set(['conservative', 'balanced', 'free-run']);
|
|
260
|
+
const VALID_AUTONOMY_LEVELS = new Set(['conservative', 'balanced', 'free-run', 'vibecode']);
|
|
261
261
|
if (!isPlainObject(config.autonomy)) {
|
|
262
262
|
errors.push('autonomy must be an object.');
|
|
263
263
|
} else {
|
package/src/index/taskRouting.js
CHANGED
|
@@ -240,6 +240,7 @@ export function buildRouteSummary({
|
|
|
240
240
|
executionCandidates,
|
|
241
241
|
});
|
|
242
242
|
const executionContract = buildExecutionContract(executionMode);
|
|
243
|
+
const continuousExecution = buildContinuousExecutionPolicy(autonomyLevel);
|
|
243
244
|
const postEditReview = executionContract?.postEditReviewPolicy
|
|
244
245
|
? { policy: executionContract.postEditReviewPolicy, agent: 'code-reviewer', reviewTargetType: 'diff' }
|
|
245
246
|
: null;
|
|
@@ -290,6 +291,8 @@ export function buildRouteSummary({
|
|
|
290
291
|
completionState,
|
|
291
292
|
postEditReview,
|
|
292
293
|
continuationState,
|
|
294
|
+
autonomyLevel,
|
|
295
|
+
continuousExecution,
|
|
293
296
|
intentMode: routingContext.intentMode ?? null,
|
|
294
297
|
handoffFile,
|
|
295
298
|
handoffBudget,
|
|
@@ -303,6 +306,19 @@ export function buildRouteSummary({
|
|
|
303
306
|
};
|
|
304
307
|
}
|
|
305
308
|
|
|
309
|
+
// Execution-ledger.mjs mirrors this policy for its Stop gate; the two must stay aligned.
|
|
310
|
+
// MAX_CONTINUATIONS (6) below is the same constant the ledger enforces for bounded modes.
|
|
311
|
+
function buildContinuousExecutionPolicy(autonomyLevel) {
|
|
312
|
+
if (autonomyLevel === 'vibecode') {
|
|
313
|
+
return {
|
|
314
|
+
enabled: true,
|
|
315
|
+
stopOn: ['completion-evidence', 'genuine-blocker', 'dangerous-command-decision'],
|
|
316
|
+
maxContinuations: null,
|
|
317
|
+
};
|
|
318
|
+
}
|
|
319
|
+
return { enabled: false, maxContinuations: 6 };
|
|
320
|
+
}
|
|
321
|
+
|
|
306
322
|
function deriveContextMode(taskType) {
|
|
307
323
|
if (taskType === 'trivial' || taskType === 'simple') return 'LITE';
|
|
308
324
|
if (taskType === 'non-trivial' || taskType === 'shared-simple') return 'FULL';
|
|
@@ -3,7 +3,22 @@
|
|
|
3
3
|
# Optimized: fast-path grep to skip Node.js when rule already exists.
|
|
4
4
|
|
|
5
5
|
INPUT=$(cat)
|
|
6
|
-
|
|
6
|
+
# jq is absent on stock macOS. Keep jq as the low-latency normal path, but fall back to
|
|
7
|
+
# UKit's required Node runtime so safe commands still receive managed metadata.
|
|
8
|
+
if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
|
|
9
|
+
COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
|
|
10
|
+
else
|
|
11
|
+
COMMAND=$(printf '%s' "$INPUT" | node -e '
|
|
12
|
+
const chunks = [];
|
|
13
|
+
process.stdin.on("data", (chunk) => chunks.push(chunk));
|
|
14
|
+
process.stdin.on("end", () => {
|
|
15
|
+
try {
|
|
16
|
+
const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
|
|
17
|
+
process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
|
|
18
|
+
} catch {}
|
|
19
|
+
});
|
|
20
|
+
' 2>/dev/null)
|
|
21
|
+
fi
|
|
7
22
|
|
|
8
23
|
if [ -z "$COMMAND" ]; then
|
|
9
24
|
exit 0
|
|
@@ -3,7 +3,22 @@
|
|
|
3
3
|
# Matched on: Bash
|
|
4
4
|
|
|
5
5
|
INPUT=$(cat)
|
|
6
|
-
|
|
6
|
+
# jq is not installed on stock macOS. Keep jq as the low-latency normal path, but use
|
|
7
|
+
# UKit's required Node runtime as a fallback so its absence cannot silently disable the gate.
|
|
8
|
+
if command -v jq >/dev/null 2>&1 && jq --version >/dev/null 2>&1; then
|
|
9
|
+
COMMAND=$(printf '%s' "$INPUT" | jq -r '.tool_input.command // empty')
|
|
10
|
+
else
|
|
11
|
+
COMMAND=$(printf '%s' "$INPUT" | node -e '
|
|
12
|
+
const chunks = [];
|
|
13
|
+
process.stdin.on("data", (chunk) => chunks.push(chunk));
|
|
14
|
+
process.stdin.on("end", () => {
|
|
15
|
+
try {
|
|
16
|
+
const payload = JSON.parse(Buffer.concat(chunks).toString("utf8") || "{}");
|
|
17
|
+
process.stdout.write(typeof payload?.tool_input?.command === "string" ? payload.tool_input.command : "");
|
|
18
|
+
} catch {}
|
|
19
|
+
});
|
|
20
|
+
' 2>/dev/null)
|
|
21
|
+
fi
|
|
7
22
|
|
|
8
23
|
if [ -z "$COMMAND" ]; then
|
|
9
24
|
exit 0
|
|
@@ -42,31 +57,97 @@ DANGEROUS_PATTERNS=(
|
|
|
42
57
|
"dd if=/dev/"
|
|
43
58
|
)
|
|
44
59
|
|
|
60
|
+
# Dangerous detections surface as a structured `ask` decision (stdout JSON) so a human
|
|
61
|
+
# decides; exit 2 still refuses the call for harnesses that ignore structured output.
|
|
62
|
+
# The raw command and the matched pattern are NEVER echoed — arguments and comments can
|
|
63
|
+
# carry secrets, and the pattern text itself restates the dangerous command.
|
|
64
|
+
emit_dangerous_decision() {
|
|
65
|
+
REASON="$1"
|
|
66
|
+
printf '%s\n' "{\"hookSpecificOutput\":{\"hookEventName\":\"PreToolUse\",\"permissionDecision\":\"ask\",\"permissionDecisionReason\":\"$REASON\"}}"
|
|
67
|
+
echo "BLOCKED: $REASON" >&2
|
|
68
|
+
exit 2
|
|
69
|
+
}
|
|
70
|
+
|
|
45
71
|
for pattern in "${DANGEROUS_PATTERNS[@]}"; do
|
|
46
72
|
if echo "$SCAN_COMMAND" | grep -qE "$pattern"; then
|
|
47
|
-
|
|
48
|
-
exit 2
|
|
73
|
+
emit_dangerous_decision "Dangerous command pattern detected; it could cause irreversible damage. UKit defers this to a human decision."
|
|
49
74
|
fi
|
|
50
75
|
done
|
|
51
76
|
|
|
52
|
-
# Handle rm commands: allow safe cleanup targets,
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
77
|
+
# Handle rm commands: allow safe cleanup targets only, defer everything else to a human.
|
|
78
|
+
# Direct rm invocations are parsed per shell segment (`;`, `&`, `|` split) with per-target
|
|
79
|
+
# allowlist evaluation: split flags (`rm -r -f`, `rm --force -r`) register, and one
|
|
80
|
+
# allowlisted word can never legitimise a sibling target (`rm -rf dist /`). A whole-command
|
|
81
|
+
# fallback still catches rm inside wrapped/quoted contexts (`bash -c "rm -r -f /x"`) and
|
|
82
|
+
# flags placed after targets (`rm src -rf`), where segment parsing cannot see rm as a head.
|
|
83
|
+
SAFE_DELETE_REGEX='(^|[[:space:]])(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/|[[:space:]]|$)'
|
|
84
|
+
UNSAFE_TARGET_REGEX='(^|[[:space:]])(/|~|\.\.?($|/)|\.($|/))'
|
|
85
|
+
SAFE_ONE_TARGET_REGEX='^(\./)?(dist|build|coverage|\.next|\.nuxt|\.turbo|tmp|temp|\.cache|node_modules)(/.*)?$'
|
|
86
|
+
UNSAFE_ONE_TARGET_REGEX='^(/.*|~|~/.*|\.\.|\.\./.*|\.)$'
|
|
87
|
+
RM_WORD_REGEX=$'(^|[;&|[:space:]"\x27])rm([[:space:]"\x27]|$)'
|
|
57
88
|
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
89
|
+
RM_VERDICT_UNSAFE=0
|
|
90
|
+
RM_VERDICT_GENERIC=0
|
|
91
|
+
RM_DIRECT_SEEN=0
|
|
92
|
+
set -f
|
|
93
|
+
while IFS= read -r segment; do
|
|
94
|
+
[ -n "$segment" ] || continue
|
|
95
|
+
seg_head=$(printf '%s' "$segment" | sed -E 's/^[[:space:]]*//' | awk '{print $1}' | xargs -I{} basename {} 2>/dev/null)
|
|
96
|
+
[ "$seg_head" = "rm" ] || continue
|
|
97
|
+
seg_flags=$(printf '%s' "$segment" | awk '{for(i=2;i<=NF;i++){if($i=="--")break;if($i ~ /^-/)printf "%s\n",$i;else break}}')
|
|
98
|
+
seg_has_r=0
|
|
99
|
+
seg_has_f=0
|
|
100
|
+
for flag in $seg_flags; do
|
|
101
|
+
case "$flag" in
|
|
102
|
+
--recursive) seg_has_r=1 ;;
|
|
103
|
+
--force) seg_has_f=1 ;;
|
|
104
|
+
--*) ;;
|
|
105
|
+
-*) case "$flag" in *[rR]*) seg_has_r=1 ;; esac
|
|
106
|
+
case "$flag" in *f*) seg_has_f=1 ;; esac ;;
|
|
107
|
+
esac
|
|
108
|
+
done
|
|
109
|
+
[ "$seg_has_r" -eq 1 ] && [ "$seg_has_f" -eq 1 ] || continue
|
|
110
|
+
RM_DIRECT_SEEN=1
|
|
111
|
+
seg_targets=$(printf '%s' "$segment" | awk '{saw=0;for(i=2;i<=NF;i++){if(!saw){if($i=="--"){saw=1;continue}if($i ~ /^-/)continue;saw=1;printf "%s\n",$i}else if($i !~ /^-/)printf "%s\n",$i}}')
|
|
112
|
+
seg_all_safe=1
|
|
113
|
+
seg_any_unsafe=0
|
|
114
|
+
for target in $seg_targets; do
|
|
115
|
+
# Strip ONE trailing slash so `dist/` matches the allowlist — but never shrink a
|
|
116
|
+
# bare `/` (or `~`) to an empty string, which would silently dodge the unsafe check.
|
|
117
|
+
if [ "${#target}" -gt 1 ]; then target="${target%/}"; fi
|
|
118
|
+
if printf '%s' "$target" | grep -qE "$UNSAFE_ONE_TARGET_REGEX"; then
|
|
119
|
+
seg_any_unsafe=1
|
|
120
|
+
seg_all_safe=0
|
|
121
|
+
elif printf '%s' "$target" | grep -qE "$SAFE_ONE_TARGET_REGEX"; then
|
|
61
122
|
:
|
|
62
|
-
elif echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
|
|
63
|
-
echo "BLOCKED: Unsafe delete target detected. Refusing recursive force-delete." >&2
|
|
64
|
-
exit 2
|
|
65
123
|
else
|
|
66
|
-
|
|
67
|
-
exit 2
|
|
124
|
+
seg_all_safe=0
|
|
68
125
|
fi
|
|
126
|
+
done
|
|
127
|
+
if [ "$seg_any_unsafe" -eq 1 ]; then RM_VERDICT_UNSAFE=1
|
|
128
|
+
elif [ "$seg_all_safe" -eq 0 ]; then RM_VERDICT_GENERIC=1
|
|
129
|
+
fi
|
|
130
|
+
done <<SEGMENTS
|
|
131
|
+
$(printf '%s' "$SCAN_COMMAND" | tr ';&|' '\n\n\n')
|
|
132
|
+
SEGMENTS
|
|
133
|
+
set +f
|
|
134
|
+
|
|
135
|
+
if [ "$RM_DIRECT_SEEN" -eq 0 ] && echo "$SCAN_COMMAND" | grep -qE "$RM_WORD_REGEX" \
|
|
136
|
+
&& echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--recursive|(^|[[:space:]])-[a-zA-Z]*[rR]" \
|
|
137
|
+
&& echo "$SCAN_COMMAND" | grep -qE "(^|[[:space:]])--force|(^|[[:space:]])-[a-zA-Z]*f"; then
|
|
138
|
+
# Unsafe wins over the allowlist here: a safe word elsewhere in the line must never
|
|
139
|
+
# legitimise an absolute/home/parent target we can also see.
|
|
140
|
+
if echo "$SCAN_COMMAND" | grep -qE "$UNSAFE_TARGET_REGEX"; then
|
|
141
|
+
RM_VERDICT_UNSAFE=1
|
|
142
|
+
elif ! echo "$SCAN_COMMAND" | grep -qE "$SAFE_DELETE_REGEX"; then
|
|
143
|
+
RM_VERDICT_GENERIC=1
|
|
69
144
|
fi
|
|
70
145
|
fi
|
|
71
146
|
|
|
147
|
+
if [ "$RM_VERDICT_UNSAFE" -eq 1 ]; then
|
|
148
|
+
emit_dangerous_decision "Unsafe delete target detected; recursive force-delete could cause irreversible damage. UKit defers this to a human decision."
|
|
149
|
+
elif [ "$RM_VERDICT_GENERIC" -eq 1 ]; then
|
|
150
|
+
emit_dangerous_decision "Recursive force-delete outside safe cleanup targets (dist/build/coverage/.next/.nuxt/.turbo/tmp/.cache/node_modules) is destructive. UKit defers this to a human decision."
|
|
151
|
+
fi
|
|
152
|
+
|
|
72
153
|
exit 0
|
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
#!/bin/bash
|
|
2
2
|
# PreToolUse hook: hard-enforce an absolute context token cap (compact.hardCapTokens,
|
|
3
|
-
# default
|
|
4
|
-
# window, so lower it on a 200k model), separate from the
|
|
3
|
+
# default 500000, sized for a 1M window — must stay below the model's real context
|
|
4
|
+
# window, so lower it on a 200k/256k model), separate from the
|
|
5
5
|
# soft/hard advisory pressure phases in
|
|
6
|
-
# compact-threshold.mjs (default soft=
|
|
6
|
+
# compact-threshold.mjs (default soft=150000/hard=240000, which only print a suggestion).
|
|
7
7
|
#
|
|
8
8
|
# Those advisory phases are just injected text — nothing stops the agent from ignoring
|
|
9
9
|
# them and letting a session run to hundreds of thousands of tokens with no compaction.
|
|
@@ -18,12 +18,13 @@
|
|
|
18
18
|
# real compaction.
|
|
19
19
|
#
|
|
20
20
|
# Grace window (compact.hardCapGraceCalls, default 10): blocking the very first tool call
|
|
21
|
-
# past the cap strands a handoff run mid-edit — files half-written,
|
|
22
|
-
#
|
|
23
|
-
# an unfinished run, the
|
|
24
|
-
#
|
|
25
|
-
#
|
|
26
|
-
#
|
|
21
|
+
# past the cap strands a handoff run or ordinary routed task mid-edit — files half-written,
|
|
22
|
+
# no durable completion state — which is strictly worse than letting it land. When RUN.md
|
|
23
|
+
# shows an unfinished handoff run, or the shared ledger proves an unfinished session-bound
|
|
24
|
+
# ordinary task, the gate allows a BOUNDED number of further calls so it can land the current
|
|
25
|
+
# edit/verification and persist resume state; after that it blocks exactly as before. The
|
|
26
|
+
# budget is per over-cap episode, not per wave: it only resets when the estimate actually
|
|
27
|
+
# drops (i.e. a real compaction happened), so a task cannot mint itself fresh grace forever.
|
|
27
28
|
#
|
|
28
29
|
# Config toggle: compact.hardCapBlock (default true). Set to false only to debug this
|
|
29
30
|
# gate itself; it must not become a normal escape hatch.
|
|
@@ -74,7 +75,7 @@ function readRunCursor() {
|
|
|
74
75
|
try {
|
|
75
76
|
const text = fs.readFileSync(path.join(projectRoot, 'docs', 'AI_HANDOFF', 'RUN.md'), 'utf8');
|
|
76
77
|
const runPhase = (text.match(/^Phase:\s*(.+)$/m)?.[1] || '').trim();
|
|
77
|
-
if (!runPhase || /^done
|
|
78
|
+
if (!runPhase || /^done\b/i.test(runPhase)) return null;
|
|
78
79
|
return {
|
|
79
80
|
phase: runPhase,
|
|
80
81
|
cursor: (text.match(/^Cursor:\s*(.+)$/m)?.[1] || '').trim(),
|
|
@@ -98,6 +99,10 @@ function readRunCursor() {
|
|
|
98
99
|
}
|
|
99
100
|
|
|
100
101
|
const mod = await import(pathToFileURL(thresholdModulePath).href);
|
|
102
|
+
const ledgerModulePath = path.join(hookDir, '..', 'ukit', 'runtime', 'execution-ledger.mjs');
|
|
103
|
+
const ledgerMod = fs.existsSync(ledgerModulePath)
|
|
104
|
+
? await import(pathToFileURL(ledgerModulePath).href)
|
|
105
|
+
: null;
|
|
101
106
|
const pressurePath = path.join(projectRoot, '.ukit', 'storage', 'cache', 'compact-pressure.json');
|
|
102
107
|
const rawState = readJsonSafe(pressurePath, null);
|
|
103
108
|
|
|
@@ -117,49 +122,134 @@ function readRunCursor() {
|
|
|
117
122
|
}
|
|
118
123
|
const thresholds = mod.buildCompactThresholds(config);
|
|
119
124
|
|
|
125
|
+
// The grace budget must survive parallel subagents hitting the cap in the same episode.
|
|
126
|
+
// A read-modify-write of one JSON file under-counts (two processes both read used=3 and
|
|
127
|
+
// both write 4, silently multiplying the budget), so each spent call is claimed with an
|
|
128
|
+
// exclusive-create sentinel file (openSync 'wx'): the sentinels present ARE the number
|
|
129
|
+
// of grace calls spent, and one claim can never be counted twice.
|
|
120
130
|
const gracePath = path.join(projectRoot, '.ukit', 'storage', 'cache', 'hardcap-grace.json');
|
|
131
|
+
const graceDir = path.dirname(gracePath);
|
|
132
|
+
const slotPath = (n) => path.join(graceDir, `hardcap-grace.json.${n}`);
|
|
133
|
+
const readSlots = () => {
|
|
134
|
+
try {
|
|
135
|
+
return fs.readdirSync(graceDir)
|
|
136
|
+
.map((name) => /^hardcap-grace\.json\.(\d+)$/.exec(name))
|
|
137
|
+
.filter(Boolean)
|
|
138
|
+
.map((m) => Number(m[1]))
|
|
139
|
+
.sort((a, b) => a - b);
|
|
140
|
+
} catch {
|
|
141
|
+
return [];
|
|
142
|
+
}
|
|
143
|
+
};
|
|
144
|
+
const clearGrace = () => {
|
|
145
|
+
try { fs.rmSync(gracePath, { force: true }); } catch { /* advisory only */ }
|
|
146
|
+
for (const n of readSlots()) {
|
|
147
|
+
try { fs.rmSync(slotPath(n), { force: true }); } catch { /* advisory only */ }
|
|
148
|
+
}
|
|
149
|
+
};
|
|
121
150
|
|
|
122
151
|
if (state.estimatedTotalTokens < thresholds.hardCapTokens) {
|
|
123
152
|
// Back under the cap — the episode is over, so the next one starts with a full budget.
|
|
124
|
-
|
|
153
|
+
clearGrace();
|
|
125
154
|
process.exit(0);
|
|
126
155
|
return;
|
|
127
156
|
}
|
|
128
157
|
|
|
129
158
|
const run = readRunCursor();
|
|
130
|
-
|
|
159
|
+
let ordinaryTask = null;
|
|
160
|
+
if (!run && ledgerMod) {
|
|
161
|
+
try {
|
|
162
|
+
ordinaryTask = await ledgerMod.readResumableExecution(projectRoot, payload);
|
|
163
|
+
} catch {
|
|
164
|
+
ordinaryTask = null;
|
|
165
|
+
}
|
|
166
|
+
}
|
|
167
|
+
const resumable = run || ordinaryTask;
|
|
168
|
+
if (resumable) {
|
|
131
169
|
const graceCalls = Number.isFinite(config?.compact?.hardCapGraceCalls)
|
|
132
170
|
? config.compact.hardCapGraceCalls
|
|
133
171
|
: 10;
|
|
134
|
-
|
|
172
|
+
|
|
173
|
+
let slots = readSlots();
|
|
174
|
+
let startedAtTokens = null;
|
|
175
|
+
for (const n of slots) {
|
|
176
|
+
const data = readJsonSafe(slotPath(n), null);
|
|
177
|
+
if (data && Number.isFinite(data.startedAtTokens)) {
|
|
178
|
+
startedAtTokens = Math.max(startedAtTokens ?? -Infinity, data.startedAtTokens);
|
|
179
|
+
}
|
|
180
|
+
}
|
|
135
181
|
// Reset only when the estimate actually fell since grace started: that is the signal a
|
|
136
182
|
// real compaction/shed happened. Advancing the cursor alone must NOT top the budget up,
|
|
137
183
|
// otherwise a long run gets unlimited grace and the ceiling stops meaning anything.
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
|
|
184
|
+
if (startedAtTokens !== null && state.estimatedTotalTokens < startedAtTokens) {
|
|
185
|
+
clearGrace();
|
|
186
|
+
slots = [];
|
|
187
|
+
startedAtTokens = null;
|
|
188
|
+
}
|
|
189
|
+
const spent = slots.length > 0 ? slots[slots.length - 1] : 0;
|
|
190
|
+
if (startedAtTokens === null) startedAtTokens = state.estimatedTotalTokens;
|
|
142
191
|
|
|
143
|
-
if (
|
|
144
|
-
|
|
192
|
+
if (spent < graceCalls) {
|
|
193
|
+
// Claim the next slot atomically: EEXIST means a parallel worker already claimed it,
|
|
194
|
+
// step to the next. A filesystem error keeps the old fail-open behavior — losing the
|
|
195
|
+
// counter must never block a landing — at the cost of one extra granted call.
|
|
196
|
+
let used = 0;
|
|
197
|
+
let counterFailed = false;
|
|
145
198
|
try {
|
|
146
|
-
fs.mkdirSync(
|
|
147
|
-
|
|
199
|
+
fs.mkdirSync(graceDir, { recursive: true });
|
|
200
|
+
for (let n = spent + 1; n <= graceCalls; n += 1) {
|
|
201
|
+
try {
|
|
202
|
+
const fd = fs.openSync(slotPath(n), 'wx');
|
|
203
|
+
fs.writeSync(fd, JSON.stringify({ startedAtTokens, used: n, ts: Date.now() }));
|
|
204
|
+
fs.closeSync(fd);
|
|
205
|
+
used = n;
|
|
206
|
+
break;
|
|
207
|
+
} catch (err) {
|
|
208
|
+
if (err?.code === 'EEXIST') continue;
|
|
209
|
+
throw err;
|
|
210
|
+
}
|
|
211
|
+
}
|
|
148
212
|
} catch {
|
|
149
|
-
|
|
213
|
+
counterFailed = true;
|
|
214
|
+
}
|
|
215
|
+
if (used === 0 && counterFailed) used = spent + 1;
|
|
216
|
+
if (used > 0 && used <= graceCalls) {
|
|
217
|
+
try {
|
|
218
|
+
fs.writeFileSync(gracePath, JSON.stringify({ startedAtTokens, used }));
|
|
219
|
+
if (ordinaryTask && ledgerMod && used === 1) {
|
|
220
|
+
await ledgerMod.writeResumeIntent(projectRoot, payload);
|
|
221
|
+
}
|
|
222
|
+
} catch {
|
|
223
|
+
// Losing the counter or advisory resume intent must not block the run.
|
|
224
|
+
}
|
|
225
|
+
const landing = run
|
|
226
|
+
? `An unfinished handoff run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`
|
|
227
|
+
: 'An unfinished ordinary routed task is in flight, so this call is allowed instead of stranding it before compaction.';
|
|
228
|
+
process.stderr.write(
|
|
229
|
+
[
|
|
230
|
+
`CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
|
|
231
|
+
landing,
|
|
232
|
+
run
|
|
233
|
+
? 'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.'
|
|
234
|
+
: 'Spend the remaining grace on LANDING, not on new work: finish the current mutation or verification, then compact as soon as the host allows it.',
|
|
235
|
+
'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
|
|
236
|
+
'Tell the user in your reply that context is over the cap and they should run /compact as soon as this landing step is complete.',
|
|
237
|
+
'After compaction, the SessionStart resume hook replays the safe cursor and the task continues automatically.',
|
|
238
|
+
].join('\n') + '\n',
|
|
239
|
+
);
|
|
240
|
+
// The stderr contract above only reaches the model — if the turn dies before the
|
|
241
|
+
// model relays it, the user sees nothing. Mirror the first grace call as a
|
|
242
|
+
// structured systemMessage so the /compact advice is user-visible regardless.
|
|
243
|
+
if (used === 1) {
|
|
244
|
+
process.stdout.write(`${JSON.stringify({
|
|
245
|
+
systemMessage: `UKit: context is over the hard cap (~${state.estimatedTotalTokens} tokens). A short landing window (${graceCalls} calls) is active — run /compact when it ends; the task resumes automatically after compaction.`,
|
|
246
|
+
})}\n`);
|
|
247
|
+
}
|
|
248
|
+
process.exit(0);
|
|
249
|
+
return;
|
|
150
250
|
}
|
|
151
|
-
|
|
152
|
-
|
|
153
|
-
`CONTEXT OVER CAP — grace ${used}/${graceCalls} (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}).`,
|
|
154
|
-
`An unfinished run is in flight (Phase: ${run.phase}${run.cursor ? `, Cursor: ${run.cursor}` : ''}), so this call is allowed instead of stranding it mid-edit.`,
|
|
155
|
-
'Spend the remaining grace on LANDING, not on new work: finish the current edit, commit, update docs/AI_HANDOFF/RUN.md, push.',
|
|
156
|
-
'Do NOT start a new task, open new files, or spawn agents. When grace runs out the gate blocks hard.',
|
|
157
|
-
'Tell the user in your reply that context is over the cap and they should run /compact as soon as this run lands.',
|
|
158
|
-
'After the user runs /compact, the SessionStart resume hook replays the cursor and the run continues automatically.',
|
|
159
|
-
].join('\n') + '\n',
|
|
160
|
-
);
|
|
161
|
-
process.exit(0);
|
|
162
|
-
return;
|
|
251
|
+
// Every slot was claimed while we raced — the budget is genuinely spent; fall
|
|
252
|
+
// through to the hard block below.
|
|
163
253
|
}
|
|
164
254
|
}
|
|
165
255
|
|
|
@@ -169,12 +259,17 @@ function readRunCursor() {
|
|
|
169
259
|
'Typing "continue" alone will hit this same block; only /compact (or a new session) resets the counter. After compacting, resume the interrupted task.',
|
|
170
260
|
'This is an absolute ceiling (compact.hardCapTokens), separate from the soft/hard advisory phases — those were apparently not followed.',
|
|
171
261
|
`Edit/Write/Bash refused (tool_name=${toolName}) until real compaction happens.`,
|
|
172
|
-
|
|
262
|
+
resumable
|
|
173
263
|
? 'The unfinished-run grace window (compact.hardCapGraceCalls) is already exhausted — progress should be committed and the cursor written by now.'
|
|
174
|
-
: 'No unfinished
|
|
264
|
+
: 'No resumable unfinished routed task was found, so there is no grace window to spend.',
|
|
175
265
|
'Do not work around this by summarizing inline and continuing, and do not reach for a non-gated write tool.',
|
|
176
266
|
];
|
|
177
267
|
process.stderr.write(`${lines.join('\n')}\n`);
|
|
268
|
+
// stderr reaches the model, not reliably the user. Pair the hard block with a
|
|
269
|
+
// structured systemMessage so "why did it stop" always has a user-visible answer.
|
|
270
|
+
process.stdout.write(`${JSON.stringify({
|
|
271
|
+
systemMessage: `UKit blocked ${toolName}: context is over the hard cap (~${state.estimatedTotalTokens} tokens >= ${thresholds.hardCapTokens}) and the landing window is spent. Run /compact now; after compaction the task continues automatically.`,
|
|
272
|
+
})}\n`);
|
|
178
273
|
process.exit(2);
|
|
179
274
|
})().catch((err) => {
|
|
180
275
|
// A logic error here fails OPEN: this is a backstop on top of advisory nudges,
|