@ngockhoale/ukit 2.3.19 → 2.3.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +49 -0
- package/manifests/platform.full.yaml +14 -0
- package/package.json +1 -1
- package/scripts/release/verify-release.mjs +45 -0
- package/templates/.claude/hooks/completion-gate.sh +13 -1
- package/templates/.claude/hooks/record-execution.sh +13 -1
- package/templates/.claude/hooks/skill-router.sh +10 -3
- package/templates/.claude/ukit/runtime/execution-ledger.mjs +158 -40
- package/templates/.omp/hooks/pre/ukit-bridge.js +0 -18
- package/templates/AGENTS.md +8 -0
- package/templates/CLAUDE.md +8 -0
- package/templates/docs/PROMPT_CACHING.md +127 -0
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,55 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.3.21 - 2026-09-13
|
|
6
|
+
|
|
7
|
+
C16 release — the prompt-caching ruleset turns from research into shipped guidance, and the
|
|
8
|
+
prompt-assembly surfaces that inject into model context are made deterministic. No new commands
|
|
9
|
+
and no default runtime behavior changes: the tool-call policy ships as guidance only, pending the
|
|
10
|
+
A/B required before any behavior change.
|
|
11
|
+
|
|
12
|
+
- **Canonical prompt-caching ruleset (`docs/PROMPT_CACHING.md`, repo-local).** Distills the UNIC
|
|
13
|
+
caching roadmap into CTX-01..10 (MUST), the SHOULD set, and the never-do list, with
|
|
14
|
+
evidence-labeled vendor sections (Anthropic, OpenAI, DeepSeek, GLM, MiniMax) and an
|
|
15
|
+
adopt/adapt/reject matrix. Upstream vendor docs are reference only — never UNIC guarantees
|
|
16
|
+
(UNIC behavior is labeled `unknown`/`inferred`). Repo-local by design; not shipped.
|
|
17
|
+
- **Shipped caching guidance (`templates/docs/PROMPT_CACHING.md`).** The distilled CTX rules,
|
|
18
|
+
never-do list, and short tool-call policy now ship with UKit installs through a new manifest
|
|
19
|
+
item `docs-prompt-caching` (`mergeStrategy: overwrite_with_backup`, so `ukit update` refreshes
|
|
20
|
+
it). `templates/CLAUDE.md` and `templates/AGENTS.md` gain a matching `## Prompt Caching`
|
|
21
|
+
pointer section, and the repo-local `CLAUDE.md`/`AGENTS.md` carry a minimal pointer too.
|
|
22
|
+
- **Deterministic prompt assembly.** `templates/.claude/hooks/skill-router.sh` ordered
|
|
23
|
+
model-visible memory-recall output by volatile `updatedAt` timestamps, so the same state could
|
|
24
|
+
assemble different prompt bytes across runs and defeat prompt caching (CTX-01/CTX-04). Ordering
|
|
25
|
+
is now a stable sort with a deterministic tiebreak and no volatile timestamps in model-visible
|
|
26
|
+
output; the live `.claude/hooks/` twin is byte-identical. Locked by the new
|
|
27
|
+
`tests/hooks/promptAssemblyDeterminism.test.js` plus an ordering regression in
|
|
28
|
+
`tests/hooks/skillRouterHook.test.js`.
|
|
29
|
+
|
|
30
|
+
## 2.3.20 - 2026-09-13
|
|
31
|
+
|
|
32
|
+
C15 bug-fix release record — this release completes the silent-stop class on top of the
|
|
33
|
+
already-published 2.3.19 (published to npm 2026-09-13, registry verified, carrying the C14
|
|
34
|
+
valve removal); it adds no new features. npm is the distribution channel and `ukit update`
|
|
35
|
+
pulls from npm, so an install on 2.3.19 (or older) needs an update to receive these fixes.
|
|
36
|
+
|
|
37
|
+
- **`--evaluate-stop` never releases a routed stop silently.**
|
|
38
|
+
`templates/.claude/ukit/runtime/execution-ledger.mjs` had a final dispatch with no `else`, so any
|
|
39
|
+
evaluation result of `{continue:false, notify:false}` that was not `capped` wrote nothing at all
|
|
40
|
+
— the Stop was allowed and the session idled with no visible reason. That default path is now a
|
|
41
|
+
loud release: an invisible blocker and a route with empty/short completion evidence both emit a
|
|
42
|
+
`systemMessage` naming the cause and how to resume, instead of silence. Malformed stdin no longer
|
|
43
|
+
crashes the hook into a silent stderr exit — it blocks with the crash detail, bounded by a crash
|
|
44
|
+
counter that loudly releases after the threshold and resets on a clean evaluate.
|
|
45
|
+
- **Fail-loud completion-gate wrappers.**
|
|
46
|
+
`templates/.claude/hooks/completion-gate.sh` and `templates/.claude/hooks/record-execution.sh`
|
|
47
|
+
exited 0 silently when the runtime script was missing or failed; both now emit a
|
|
48
|
+
`systemMessage` naming the missing/failing runtime and the `ukit install` remedy.
|
|
49
|
+
- **omp reentrant stop gets the same never-silent contract as the CLI.**
|
|
50
|
+
`templates/.omp/hooks/pre/ukit-bridge.js` released a reentrant stop with a non-blocking notice
|
|
51
|
+
and no continuation; it now blocks with the actionable reason like any other stop, stays bounded
|
|
52
|
+
by the ledger's continuation cap, and releases loudly at the cap.
|
|
53
|
+
|
|
5
54
|
## 2.3.19 - 2026-09-12
|
|
6
55
|
|
|
7
56
|
C14 bug-fix release record — one correctness fix is shipped; this release adds no new features.
|
|
@@ -218,6 +218,20 @@ items:
|
|
|
218
218
|
packs:
|
|
219
219
|
- core
|
|
220
220
|
|
|
221
|
+
# Shipped prompt-caching guidance (CTX-01..10 + never-do + tool-call policy + vendor cheat
|
|
222
|
+
# sheet). `overwrite_with_backup` so `ukit update` refreshes the ruleset — deliberately NOT
|
|
223
|
+
# `docs/PROJECT.md`, which is `mergeStrategy: skip` and must never be auto-overwritten.
|
|
224
|
+
- id: docs-prompt-caching
|
|
225
|
+
type: config
|
|
226
|
+
sourceTemplate: docs/PROMPT_CACHING.md
|
|
227
|
+
targetPath: docs/PROMPT_CACHING.md
|
|
228
|
+
requires: []
|
|
229
|
+
mergeStrategy: overwrite_with_backup
|
|
230
|
+
variables: []
|
|
231
|
+
enabledByDefault: true
|
|
232
|
+
packs:
|
|
233
|
+
- core
|
|
234
|
+
|
|
221
235
|
- id: core-skill-delivery
|
|
222
236
|
type: skill
|
|
223
237
|
sourceTemplate: .claude/skills/delivery/SKILL.md
|
package/package.json
CHANGED
|
@@ -1,10 +1,35 @@
|
|
|
1
1
|
import { spawn } from 'node:child_process';
|
|
2
|
+
import { readFileSync } from 'node:fs';
|
|
2
3
|
import os from 'node:os';
|
|
3
4
|
import path from 'node:path';
|
|
4
5
|
|
|
5
6
|
const rootDir = process.cwd();
|
|
6
7
|
const npmCacheDir = path.join(os.tmpdir(), 'ukit-npm-cache');
|
|
7
8
|
|
|
9
|
+
// Opt-in post-publish registry-parity check (Release Policy invariant: npm and git always hold
|
|
10
|
+
// the same latest version). Only meaningful AFTER `npm publish` — before publish the registry is
|
|
11
|
+
// behind by definition — so it is gated behind --post-publish and never runs by default.
|
|
12
|
+
const postPublish = process.argv.includes('--post-publish');
|
|
13
|
+
|
|
14
|
+
if (postPublish) {
|
|
15
|
+
const localVersion = JSON.parse(readFileSync(path.join(rootDir, 'package.json'), 'utf8')).version;
|
|
16
|
+
const registryVersion = await readRegistryVersion();
|
|
17
|
+
|
|
18
|
+
console.log(`[release:verify] Registry parity (post-publish)`);
|
|
19
|
+
console.log(`[release:verify] package.json = ${localVersion}`);
|
|
20
|
+
console.log(`[release:verify] npm registry = ${registryVersion}`);
|
|
21
|
+
|
|
22
|
+
if (registryVersion !== localVersion) {
|
|
23
|
+
console.error(
|
|
24
|
+
`[release:verify] FAILED: registry version ${registryVersion} !== package.json ${localVersion}`,
|
|
25
|
+
);
|
|
26
|
+
process.exit(1);
|
|
27
|
+
}
|
|
28
|
+
|
|
29
|
+
console.log('\n[release:verify] All release checks passed.');
|
|
30
|
+
process.exit(0);
|
|
31
|
+
}
|
|
32
|
+
|
|
8
33
|
const steps = [
|
|
9
34
|
{
|
|
10
35
|
label: 'Artifact smoke',
|
|
@@ -54,3 +79,23 @@ function runStep({ command, args, env }) {
|
|
|
54
79
|
child.on('error', () => resolve(1));
|
|
55
80
|
});
|
|
56
81
|
}
|
|
82
|
+
|
|
83
|
+
function readRegistryVersion() {
|
|
84
|
+
return new Promise((resolve) => {
|
|
85
|
+
const child = spawn('npm', ['view', '@ngockhoale/ukit', 'version'], {
|
|
86
|
+
cwd: rootDir,
|
|
87
|
+
env: { ...process.env, npm_config_cache: npmCacheDir },
|
|
88
|
+
stdio: ['ignore', 'pipe', 'inherit'],
|
|
89
|
+
});
|
|
90
|
+
|
|
91
|
+
let stdout = '';
|
|
92
|
+
child.stdout.on('data', (chunk) => {
|
|
93
|
+
stdout += chunk;
|
|
94
|
+
});
|
|
95
|
+
child.on('close', () => {
|
|
96
|
+
const version = stdout.trim();
|
|
97
|
+
resolve(version.length > 0 ? version : '<unavailable>');
|
|
98
|
+
});
|
|
99
|
+
child.on('error', () => resolve('<unavailable>'));
|
|
100
|
+
});
|
|
101
|
+
}
|
|
@@ -1,12 +1,24 @@
|
|
|
1
1
|
#!/bin/bash
|
|
2
2
|
# Stop hook: block premature terminal stops while routed completion evidence is missing.
|
|
3
|
+
# ADVISORY ONLY — always exit 0. A missing or failing runtime must be loud, not a silent pass.
|
|
3
4
|
|
|
4
5
|
INPUT=$(cat)
|
|
5
6
|
PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
|
|
6
7
|
SCRIPT="$PROJECT_ROOT/.claude/ukit/runtime/execution-ledger.mjs"
|
|
7
8
|
|
|
8
9
|
if [ ! -f "$SCRIPT" ]; then
|
|
10
|
+
printf '%s\n' '{"systemMessage":"UKit completion gate: runtime script missing — run: ukit install"}'
|
|
9
11
|
exit 0
|
|
10
12
|
fi
|
|
11
13
|
|
|
12
|
-
printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --evaluate-stop
|
|
14
|
+
OUTPUT=$(printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --evaluate-stop)
|
|
15
|
+
STATUS=$?
|
|
16
|
+
|
|
17
|
+
if [ "$STATUS" -ne 0 ]; then
|
|
18
|
+
printf '%s\n' '{"systemMessage":"UKit completion gate: runtime script failed — run: ukit install"}'
|
|
19
|
+
exit 0
|
|
20
|
+
fi
|
|
21
|
+
|
|
22
|
+
if [ -n "$OUTPUT" ]; then
|
|
23
|
+
printf '%s\n' "$OUTPUT"
|
|
24
|
+
fi
|
|
@@ -1,12 +1,24 @@
|
|
|
1
1
|
#!/bin/bash
|
|
2
2
|
# PostToolUse hook: persist session-scoped source/write/verification receipts.
|
|
3
|
+
# ADVISORY ONLY — always exit 0. A missing or failing runtime must be loud, not a silent pass.
|
|
3
4
|
|
|
4
5
|
INPUT=$(cat)
|
|
5
6
|
PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
|
|
6
7
|
SCRIPT="$PROJECT_ROOT/.claude/ukit/runtime/execution-ledger.mjs"
|
|
7
8
|
|
|
8
9
|
if [ ! -f "$SCRIPT" ]; then
|
|
10
|
+
printf '%s\n' '{"systemMessage":"UKit record execution: runtime script missing — run: ukit install"}'
|
|
9
11
|
exit 0
|
|
10
12
|
fi
|
|
11
13
|
|
|
12
|
-
printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --record
|
|
14
|
+
OUTPUT=$(printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --record)
|
|
15
|
+
STATUS=$?
|
|
16
|
+
|
|
17
|
+
if [ "$STATUS" -ne 0 ]; then
|
|
18
|
+
printf '%s\n' '{"systemMessage":"UKit record execution: runtime script failed — run: ukit install"}'
|
|
19
|
+
exit 0
|
|
20
|
+
fi
|
|
21
|
+
|
|
22
|
+
if [ -n "$OUTPUT" ]; then
|
|
23
|
+
printf '%s\n' "$OUTPUT"
|
|
24
|
+
fi
|
|
@@ -961,8 +961,12 @@ const { pathToFileURL } = require('url');
|
|
|
961
961
|
return 0;
|
|
962
962
|
}
|
|
963
963
|
|
|
964
|
-
|
|
965
|
-
|
|
964
|
+
// TASK-027 CTX-01/CTX-04: the ranking must be a pure function of the logical memory
|
|
965
|
+
// content. A recency bonus scaled by Date.now() made the score — and therefore the
|
|
966
|
+
// injected previous-context block, its persisted order, and the PreCompact reinjection —
|
|
967
|
+
// shift on every wall-clock tick. Eligibility is unchanged (still score > 0); only the
|
|
968
|
+
// volatile ordering signal is dropped.
|
|
969
|
+
return score;
|
|
966
970
|
}
|
|
967
971
|
|
|
968
972
|
function buildPreviousContextSnippet(item) {
|
|
@@ -1017,9 +1021,12 @@ const { pathToFileURL } = require('url');
|
|
|
1017
1021
|
score: scoreMemoryItem(item, queryTokens),
|
|
1018
1022
|
}))
|
|
1019
1023
|
.filter((entry) => entry.score > 0)
|
|
1024
|
+
// CTX-01: ties break on the stable logical id (never on a volatile timestamp), so the
|
|
1025
|
+
// same memory set always yields the same injected order regardless of when the
|
|
1026
|
+
// entries were last written/ended.
|
|
1020
1027
|
.sort((left, right) => (
|
|
1021
1028
|
right.score - left.score
|
|
1022
|
-
||
|
|
1029
|
+
|| left.item.id.localeCompare(right.item.id)
|
|
1023
1030
|
))
|
|
1024
1031
|
.slice(0, 2)
|
|
1025
1032
|
.map((entry) => entry.item);
|
|
@@ -13,6 +13,12 @@ const RESUME_INTENT_TTL_MS = 30 * 60 * 1000;
|
|
|
13
13
|
const MAX_RECEIPTS = 24;
|
|
14
14
|
const MAX_SOURCE_FILES = 16;
|
|
15
15
|
const MAX_CONTINUATIONS = 6;
|
|
16
|
+
// A malformed stdin payload crashes `--evaluate-stop` before a route/session can be read, so
|
|
17
|
+
// the only identity available is the project root. Count consecutive crashes there and, once
|
|
18
|
+
// this bounded budget is exhausted, release LOUD (naming the crash + `ukit install`) instead of
|
|
19
|
+
// blocking forever on a corrupt hook input. A successful evaluate resets the count. Mirrors
|
|
20
|
+
// MAX_CONTINUATIONS style — a hard-coded module const, deliberately not a config knob.
|
|
21
|
+
const MAX_GATE_CRASHES = 3;
|
|
16
22
|
// Two adjacent identical failed verifications (no edit attempt between them) are a loop;
|
|
17
23
|
// vibecode routes — which bypass the cap and every reentrant valve — additionally get a
|
|
18
24
|
// cumulative escape after the same verification failed this many times across edits.
|
|
@@ -82,6 +88,17 @@ function ledgerPath(projectRoot, payload = {}) {
|
|
|
82
88
|
);
|
|
83
89
|
}
|
|
84
90
|
|
|
91
|
+
function crashCounterPath(projectRoot) {
|
|
92
|
+
return path.join(
|
|
93
|
+
projectRoot,
|
|
94
|
+
'.ukit',
|
|
95
|
+
'storage',
|
|
96
|
+
'cache',
|
|
97
|
+
'exec-ledger',
|
|
98
|
+
'gate-crash-counter.json',
|
|
99
|
+
);
|
|
100
|
+
}
|
|
101
|
+
|
|
85
102
|
function resumeIntentPath(projectRoot, sessionId) {
|
|
86
103
|
return path.join(
|
|
87
104
|
projectRoot,
|
|
@@ -727,10 +744,32 @@ export function evaluateCompletion({ state = {}, ledger = {} } = {}) {
|
|
|
727
744
|
reason: `UKit stopped automatic recovery: ${ledger.blocker.detail || 'a recorded blocker ended this run.'}`,
|
|
728
745
|
};
|
|
729
746
|
}
|
|
747
|
+
// An invisible blocker (e.g. a permission decision the model already named in its own
|
|
748
|
+
// reply) must not loop — but it must not end silent either. Release loud so the user can
|
|
749
|
+
// tell a deliberate stop from a stall.
|
|
750
|
+
return {
|
|
751
|
+
continue: false,
|
|
752
|
+
notify: true,
|
|
753
|
+
missingEvidence: [],
|
|
754
|
+
reason: `UKit stopped automatic recovery: ${ledger.blocker.detail || 'a recorded blocker ended this run.'}`,
|
|
755
|
+
};
|
|
756
|
+
}
|
|
757
|
+
if (!state || !state.routeSummary) {
|
|
758
|
+
// No routable state at all: a never-routed session (or a foreign-owned one) is silent
|
|
759
|
+
// success — there is nothing to report. The CLI's lost-route path handles the case where
|
|
760
|
+
// THIS session has ledger activity but the state vanished; callers that pass no state
|
|
761
|
+
// (e.g. the omp bridge for a session it does not own) must not be force-notified.
|
|
730
762
|
return { continue: false, notify: false, missingEvidence: [] };
|
|
731
763
|
}
|
|
732
764
|
if (evidence.length === 0) {
|
|
733
|
-
|
|
765
|
+
// A route with no completion evidence has no actionable block reason — blocking would
|
|
766
|
+
// loop unboundedly. Release loud instead of ending indistinguishable from a stall.
|
|
767
|
+
return {
|
|
768
|
+
continue: false,
|
|
769
|
+
notify: true,
|
|
770
|
+
missingEvidence: [],
|
|
771
|
+
reason: 'UKit completion gate: this route carries no completion evidence to verify, so the stop released without checking unfinished work. Send the task again in a new message to re-route it with a completion contract.',
|
|
772
|
+
};
|
|
734
773
|
}
|
|
735
774
|
|
|
736
775
|
// Request identity is the prompt (see evidencePromptKey): the router re-keys requestKey on
|
|
@@ -742,7 +781,10 @@ export function evaluateCompletion({ state = {}, ledger = {} } = {}) {
|
|
|
742
781
|
const effectiveLedger = sameRequest ? ledger : {};
|
|
743
782
|
const missingEvidence = evidence.filter((item) => !evidenceSatisfied(item, effectiveLedger, routeSummary));
|
|
744
783
|
if (missingEvidence.length === 0) {
|
|
745
|
-
|
|
784
|
+
// Silent success: the route is present and every required evidence is satisfied. Marked
|
|
785
|
+
// `complete` so the CLI dispatch recognizes it BEFORE the loud final else — otherwise a
|
|
786
|
+
// bare loud else would turn a clean stop into user-visible noise.
|
|
787
|
+
return { continue: false, notify: false, complete: true, missingEvidence: [] };
|
|
746
788
|
}
|
|
747
789
|
|
|
748
790
|
// `find-cause` can validly end clean: an investigation may establish that no actionable
|
|
@@ -907,7 +949,121 @@ async function readStdin() {
|
|
|
907
949
|
return chunks.join('');
|
|
908
950
|
}
|
|
909
951
|
|
|
952
|
+
// Malformed hook input crashes the evaluate BEFORE any route/session can be read, so the only
|
|
953
|
+
// stable identity is the project root. Count consecutive crashes there; past MAX_GATE_CRASHES,
|
|
954
|
+
// release LOUD naming the crash and the remedy instead of blocking a corrupt hook forever. Any
|
|
955
|
+
// counter I/O failure fails SAFE to block (still loud, never silent). A successful evaluate
|
|
956
|
+
// resets the count.
|
|
957
|
+
async function handleEvaluateCrash(projectRoot, error) {
|
|
958
|
+
const detail = `malformed hook input crashed the completion gate (${error?.message || error})`;
|
|
959
|
+
const blockReason = () => `${detail}. This is the UKit completion gate, not the task — re-send the task in a new message so the hook payload is regenerated.`;
|
|
960
|
+
let count = 0;
|
|
961
|
+
try {
|
|
962
|
+
const current = await readJson(crashCounterPath(projectRoot), null);
|
|
963
|
+
count = Number(current?.count) || 0;
|
|
964
|
+
} catch {
|
|
965
|
+
count = 0;
|
|
966
|
+
}
|
|
967
|
+
const next = count + 1;
|
|
968
|
+
try {
|
|
969
|
+
await writeJsonAtomic(crashCounterPath(projectRoot), { count: next, updatedAt: Date.now() });
|
|
970
|
+
} catch {
|
|
971
|
+
// Counter unwritable: we cannot bound the crash loop, but silence is never an option.
|
|
972
|
+
process.stdout.write(`${JSON.stringify({ decision: 'block', reason: blockReason() })}\n`);
|
|
973
|
+
return;
|
|
974
|
+
}
|
|
975
|
+
if (next >= MAX_GATE_CRASHES) {
|
|
976
|
+
const message = `UKit completion gate: ${detail}. This recurred ${next} times and cannot be bounded, so the stop is releasing loudly instead of blocking again. The hook payload is corrupt — run: ukit install to refresh the runtime, then re-send the task in a new message.`;
|
|
977
|
+
process.stderr.write(`[ukit-completion] ${message}\n`);
|
|
978
|
+
process.stdout.write(`${JSON.stringify({ systemMessage: message })}\n`);
|
|
979
|
+
return;
|
|
980
|
+
}
|
|
981
|
+
process.stdout.write(`${JSON.stringify({ decision: 'block', reason: blockReason() })}\n`);
|
|
982
|
+
}
|
|
983
|
+
|
|
984
|
+
async function resetGateCrashCounter(projectRoot) {
|
|
985
|
+
// Only touch disk when a crash was actually counted — a clean stop on a healthy project
|
|
986
|
+
// must not create/rewrite a crash-counter file.
|
|
987
|
+
try {
|
|
988
|
+
await fs.access(crashCounterPath(projectRoot));
|
|
989
|
+
} catch {
|
|
990
|
+
return;
|
|
991
|
+
}
|
|
992
|
+
try {
|
|
993
|
+
await writeJsonAtomic(crashCounterPath(projectRoot), { count: 0, updatedAt: Date.now() });
|
|
994
|
+
} catch {
|
|
995
|
+
// Best-effort: an unwritable counter only means the next crash blocks again, which is safe.
|
|
996
|
+
}
|
|
997
|
+
}
|
|
998
|
+
|
|
999
|
+
async function runEvaluateStop() {
|
|
1000
|
+
const projectRoot = process.env.CLAUDE_PROJECT_DIR || process.cwd();
|
|
1001
|
+
let payload;
|
|
1002
|
+
try {
|
|
1003
|
+
payload = JSON.parse((await readStdin()) || '{}');
|
|
1004
|
+
} catch (error) {
|
|
1005
|
+
await handleEvaluateCrash(projectRoot, error);
|
|
1006
|
+
return;
|
|
1007
|
+
}
|
|
1008
|
+
// A parse that reached here is a functioning evaluate — reset the crash budget so the next
|
|
1009
|
+
// malformed payload starts blocking again.
|
|
1010
|
+
await resetGateCrashCounter(projectRoot);
|
|
1011
|
+
|
|
1012
|
+
const state = await readRouteState(projectRoot, payload);
|
|
1013
|
+
const ledger = await readExecutionLedger(projectRoot, payload) || {};
|
|
1014
|
+
|
|
1015
|
+
// A session ledger with activity proves this session was routed and accumulating evidence.
|
|
1016
|
+
// If the route state is now missing or unreadable (router crashed/timed out before
|
|
1017
|
+
// persisting it, or the session stamp was lost), the gate cannot evaluate completion — and
|
|
1018
|
+
// ending silently there is indistinguishable from a mid-run stall. Surface the loss so the
|
|
1019
|
+
// user sees WHY the run released. Sessions that were never routed (no ledger activity at
|
|
1020
|
+
// all) stay silent: there is nothing to report (silent-success carve-out).
|
|
1021
|
+
if (!state) {
|
|
1022
|
+
const hasActivity = ledger.sourceSucceeded || ledger.writeAttempted
|
|
1023
|
+
|| ledger.verificationAttempted || (ledger.receipts || []).length > 0;
|
|
1024
|
+
if (!hasActivity) return;
|
|
1025
|
+
process.stdout.write(`${JSON.stringify({
|
|
1026
|
+
systemMessage: 'UKit completion gate: route state for this session was lost before the stop was evaluated (router did not persist it), so unfinished work could not be verified. Send the task again in a new message to re-route and finish it.',
|
|
1027
|
+
})}\n`);
|
|
1028
|
+
return;
|
|
1029
|
+
}
|
|
1030
|
+
|
|
1031
|
+
const result = evaluateCompletion({ state, ledger });
|
|
1032
|
+
|
|
1033
|
+
// A reentrant Stop (Claude Code re-fires Stop after this hook already blocked once) is
|
|
1034
|
+
// gated exactly like any other Stop: while evidence is still missing it blocks again with
|
|
1035
|
+
// the actionable reason instead of silently releasing the recovery turn. Loop termination
|
|
1036
|
+
// is guaranteed elsewhere — non-vibecode routes end at the continuation cap (final-notice
|
|
1037
|
+
// block -> capped visible release), and vibecode ends on the visible verification-loop
|
|
1038
|
+
// blocker.
|
|
1039
|
+
if (result.continue) {
|
|
1040
|
+
if (result.finalNotice) await markNotified(projectRoot, payload, ledger);
|
|
1041
|
+
else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
|
|
1042
|
+
process.stdout.write(`${JSON.stringify({ decision: 'block', reason: result.reason })}\n`);
|
|
1043
|
+
return;
|
|
1044
|
+
}
|
|
1045
|
+
if (result.capped || result.notify) {
|
|
1046
|
+
// Non-blocking endings (non-gated modes, an invisible blocker, an evidence-less route, or
|
|
1047
|
+
// cap reached after the final notice) must still tell the user what is unfinished — a
|
|
1048
|
+
// silent end is indistinguishable from a stall.
|
|
1049
|
+
process.stderr.write(`[ukit-completion] ${result.reason}\n`);
|
|
1050
|
+
process.stdout.write(`${JSON.stringify({ systemMessage: result.reason })}\n`);
|
|
1051
|
+
return;
|
|
1052
|
+
}
|
|
1053
|
+
if (result.complete) return; // silent success: route present, all evidence satisfied
|
|
1054
|
+
// Defensive final else: no evaluation result may ever end a routed stop silently.
|
|
1055
|
+
const fallback = `UKit completion gate: the stop was evaluated without a decision (missing evidence: ${(result.missingEvidence || []).join(', ') || 'none'}). If work is unfinished, re-send the task in a new message.`;
|
|
1056
|
+
process.stderr.write(`[ukit-completion] ${fallback}\n`);
|
|
1057
|
+
process.stdout.write(`${JSON.stringify({ systemMessage: fallback })}\n`);
|
|
1058
|
+
}
|
|
1059
|
+
|
|
910
1060
|
async function main() {
|
|
1061
|
+
// Dispatch the stop flag BEFORE parsing stdin: a malformed payload used to throw out of
|
|
1062
|
+
// main(), hit the .catch (stderr + exit 1) with no decision, and release the stop silently.
|
|
1063
|
+
if (process.argv.includes('--evaluate-stop')) {
|
|
1064
|
+
await runEvaluateStop();
|
|
1065
|
+
return;
|
|
1066
|
+
}
|
|
911
1067
|
const payload = JSON.parse((await readStdin()) || '{}');
|
|
912
1068
|
const projectRoot = process.env.CLAUDE_PROJECT_DIR || payload.cwd || process.cwd();
|
|
913
1069
|
if (process.argv.includes('--record')) {
|
|
@@ -917,44 +1073,6 @@ async function main() {
|
|
|
917
1073
|
toolName: payload.tool_name,
|
|
918
1074
|
harness: process.env.UKIT_HARNESS || 'claude-code',
|
|
919
1075
|
});
|
|
920
|
-
return;
|
|
921
|
-
}
|
|
922
|
-
if (process.argv.includes('--evaluate-stop')) {
|
|
923
|
-
const state = await readRouteState(projectRoot, payload);
|
|
924
|
-
const ledger = await readExecutionLedger(projectRoot, payload) || {};
|
|
925
|
-
const result = evaluateCompletion({ state, ledger });
|
|
926
|
-
|
|
927
|
-
// A session ledger with activity proves this session was routed and accumulating
|
|
928
|
-
// evidence. If the route state is now missing or unreadable (router crashed/timed out
|
|
929
|
-
// before persisting it, or the session stamp was lost), the gate cannot evaluate
|
|
930
|
-
// completion — and ending silently there is indistinguishable from a mid-run stall.
|
|
931
|
-
// Surface the loss so the user sees WHY the run released. Sessions that were never
|
|
932
|
-
// routed (no ledger activity at all) stay silent: there is nothing to report.
|
|
933
|
-
if (!state && (ledger.sourceSucceeded || ledger.writeAttempted
|
|
934
|
-
|| ledger.verificationAttempted || (ledger.receipts || []).length > 0)) {
|
|
935
|
-
process.stdout.write(`${JSON.stringify({
|
|
936
|
-
systemMessage: 'UKit completion gate: route state for this session was lost before the stop was evaluated (router did not persist it), so unfinished work could not be verified. Send the task again in a new message to re-route and finish it.',
|
|
937
|
-
})}\n`);
|
|
938
|
-
return;
|
|
939
|
-
}
|
|
940
|
-
|
|
941
|
-
// A reentrant Stop (Claude Code re-fires Stop after this hook already blocked once) is
|
|
942
|
-
// gated exactly like any other Stop: while evidence is still missing it blocks again
|
|
943
|
-
// with the actionable reason instead of silently releasing the recovery turn. Loop
|
|
944
|
-
// termination is guaranteed elsewhere — non-vibecode routes end at the continuation cap
|
|
945
|
-
// (final-notice block -> capped visible release), and vibecode ends on the visible
|
|
946
|
-
// verification-loop blocker.
|
|
947
|
-
|
|
948
|
-
if (result.continue) {
|
|
949
|
-
if (result.finalNotice) await markNotified(projectRoot, payload, ledger);
|
|
950
|
-
else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
|
|
951
|
-
process.stdout.write(`${JSON.stringify({ decision: 'block', reason: result.reason })}\n`);
|
|
952
|
-
} else if (result.capped || result.notify) {
|
|
953
|
-
// Non-blocking endings (non-gated modes, or cap reached after the final notice) must
|
|
954
|
-
// still tell the user what is unfinished — a silent end is indistinguishable from a stall.
|
|
955
|
-
process.stderr.write(`[ukit-completion] ${result.reason}\n`);
|
|
956
|
-
process.stdout.write(`${JSON.stringify({ systemMessage: result.reason })}\n`);
|
|
957
|
-
}
|
|
958
1076
|
}
|
|
959
1077
|
}
|
|
960
1078
|
|
|
@@ -14,7 +14,6 @@ import {
|
|
|
14
14
|
evidencePromptKey,
|
|
15
15
|
incrementContinuation,
|
|
16
16
|
markNotified,
|
|
17
|
-
noteStopProgress,
|
|
18
17
|
readExecutionLedger,
|
|
19
18
|
readRouteState,
|
|
20
19
|
recordExecutionReceipt,
|
|
@@ -634,23 +633,6 @@ export async function runSessionStop(
|
|
|
634
633
|
}
|
|
635
634
|
|
|
636
635
|
if (suppliedLedger === undefined) {
|
|
637
|
-
// Claude Code flags reentrant stops with stop_hook_active and the ledger CLI releases on
|
|
638
|
-
// them; omp's session_stop carries no such field. Detect the equivalent from the ledger:
|
|
639
|
-
// when the recovery turn produced no new receipts since the stop that blocked it,
|
|
640
|
-
// blocking again would only burn another turn in a self-sustaining loop — release with a
|
|
641
|
-
// visible reason instead. Vibecode autonomy keeps pushing (same as the CLI valve).
|
|
642
|
-
let reentrantStop = false;
|
|
643
|
-
try {
|
|
644
|
-
reentrantStop = (await noteStopProgress(projectRoot, payload)).reentrant === true;
|
|
645
|
-
} catch {
|
|
646
|
-
reentrantStop = false;
|
|
647
|
-
}
|
|
648
|
-
if (reentrantStop && state?.routeSummary?.autonomyLevel !== 'vibecode') {
|
|
649
|
-
const notice = `UKit stopped automatic recovery after one continuation: ${evaluation.reason}`;
|
|
650
|
-
pi.logger?.warn?.(`[UKit] ${notice}`);
|
|
651
|
-
sendContext(pi, [notice], 'nextTurn', { display: true });
|
|
652
|
-
return undefined;
|
|
653
|
-
}
|
|
654
636
|
try {
|
|
655
637
|
if (evaluation.finalNotice) await markNotified(projectRoot, payload, ledger);
|
|
656
638
|
else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
|
package/templates/AGENTS.md
CHANGED
|
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
|
|
|
106
106
|
- Threshold-based compact pressure is internal orchestration; do not expose it to users.
|
|
107
107
|
- For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
|
|
108
108
|
|
|
109
|
+
## Prompt Caching
|
|
110
|
+
|
|
111
|
+
- Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
|
|
112
|
+
- Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
|
|
113
|
+
- CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
|
|
114
|
+
- Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
|
|
115
|
+
- Upstream vendor docs are reference only — never a guarantee about the route you actually use.
|
|
116
|
+
|
|
109
117
|
## Safe Patch Protocol
|
|
110
118
|
|
|
111
119
|
- Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
|
package/templates/CLAUDE.md
CHANGED
|
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
|
|
|
106
106
|
- Threshold-based compact pressure is internal orchestration; do not expose it to users.
|
|
107
107
|
- For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
|
|
108
108
|
|
|
109
|
+
## Prompt Caching
|
|
110
|
+
|
|
111
|
+
- Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
|
|
112
|
+
- Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
|
|
113
|
+
- CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
|
|
114
|
+
- Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
|
|
115
|
+
- Upstream vendor docs are reference only — never a guarantee about the route you actually use.
|
|
116
|
+
|
|
109
117
|
## Safe Patch Protocol
|
|
110
118
|
|
|
111
119
|
- Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
|
|
@@ -0,0 +1,127 @@
|
|
|
1
|
+
# Prompt Caching — guidance for UKit projects
|
|
2
|
+
|
|
3
|
+
Prompt caching is a **prefix match**: a provider can reuse the computation of an input prefix
|
|
4
|
+
when a later request reproduces that prefix byte-for-byte. UKit cannot control the transport or
|
|
5
|
+
the provider cache engine, but it *does* control the instruction, skill, tool and hook-injected
|
|
6
|
+
content it renders. This file ships with UKit so every installed project gets the same rules for
|
|
7
|
+
keeping that content deterministic and stable. Read it on demand — it is deliberately not loaded
|
|
8
|
+
into every session.
|
|
9
|
+
|
|
10
|
+
## Why stable context matters
|
|
11
|
+
|
|
12
|
+
- Cache reuse is **reported**, not controllable. A `cache_read` counter (or a vendor equivalent)
|
|
13
|
+
proves the provider *reported* a reuse; it never proves which layer served it.
|
|
14
|
+
- Three mechanisms must never be conflated: the **provider prompt cache** (reuses input
|
|
15
|
+
computation), a **local tool-result cache** (the host reuses a still-valid result), and a
|
|
16
|
+
**response cache** (the application replays a stored answer).
|
|
17
|
+
- Three counts must never be conflated: a **logical model turn**, a **client HTTP attempt**, and a
|
|
18
|
+
**tool execution**.
|
|
19
|
+
- The rules below hold even if no cache capability exists behind your provider, because they
|
|
20
|
+
govern content UKit itself renders.
|
|
21
|
+
|
|
22
|
+
## CTX rules — MUST
|
|
23
|
+
|
|
24
|
+
| ID | Rule | What it means in practice |
|
|
25
|
+
|---|---|---|
|
|
26
|
+
| CTX-01 | Same logical input and config produce the same segment bytes | Render owned blocks deterministically; no ordering that varies run-to-run |
|
|
27
|
+
| CTX-02 | Preserve instruction/user/tool roles and conversation order | Never move a user ask into system or reorder history to lengthen a prefix |
|
|
28
|
+
| CTX-03 | Keep all tool IDs and required native continuation state intact | Never strip tool-call IDs or reasoning/continuation fields to shrink a request |
|
|
29
|
+
| CTX-04 | Do not inject clock/random IDs into static instructions | Static blocks carry no date, UUID or counter; volatile values go in the tail |
|
|
30
|
+
| CTX-05 | Every summary/compaction creates a new context epoch | Compaction is a decision with a cost; when it happens it starts a new epoch |
|
|
31
|
+
| CTX-06 | Never change data or code to match the cache | Cache optimization never edits content semantics — correctness over prefix |
|
|
32
|
+
| CTX-07 | Do not self-send a field the adapter has not confirmed as supported | Optional cache params only when capability is verified; HTTP 200 is not proof |
|
|
33
|
+
| CTX-08 | Tool-result cache reuse only when freshness/dependency is valid | Local reuse needs a freshness predicate and dependency fingerprint, not just a query match |
|
|
34
|
+
| CTX-09 | Never treat missing usage as zero | Missing cache counters are unknown plus lowered coverage, never a reported miss |
|
|
35
|
+
| CTX-10 | Never exceed existing instructions/permissions to cut calls | Reducing tool calls never means skipping a required check or test |
|
|
36
|
+
|
|
37
|
+
## Never do
|
|
38
|
+
|
|
39
|
+
- Never sort messages or reasoning blocks alphabetically.
|
|
40
|
+
- Never trim code literals, signed content, or data where whitespace is meaningful.
|
|
41
|
+
- Never rewrite native reasoning/signature/encrypted continuation fields.
|
|
42
|
+
- Never move a user request into system to lengthen the stable prefix.
|
|
43
|
+
- Never promote untrusted documents or tool output to the developer/system role.
|
|
44
|
+
- Never hash one text and assume every occurrence shares the provider cache.
|
|
45
|
+
- Never add a timestamp just to log — keep log metadata outside model-visible content.
|
|
46
|
+
|
|
47
|
+
## Runtime tool-call policy (condensed, guidance only)
|
|
48
|
+
|
|
49
|
+
This is guidance for how an assistant should decide when to call tools. It is **not** a rigid
|
|
50
|
+
"always at least N tools" or "max N tools" rule, and UKit does not change any default behavior
|
|
51
|
+
on its basis.
|
|
52
|
+
|
|
53
|
+
1. Before calling a tool, decide which data is still missing to finish the request.
|
|
54
|
+
2. Reuse evidence you already hold if it is still valid; check the version before reuse.
|
|
55
|
+
3. Batch independent reads when the interface supports it.
|
|
56
|
+
4. For dependent actions, wait for the needed result before deciding the next step.
|
|
57
|
+
5. Prefer scoped queries with enough output to verify.
|
|
58
|
+
6. If a result was truncated or insufficient, widen deliberately.
|
|
59
|
+
7. On tool error, distinguish parameter error, transient error, and unknown state.
|
|
60
|
+
8. After a change, run the check appropriate to the risk and the repo's requirements.
|
|
61
|
+
9. When the completion condition is met, return the result — check further only for specific
|
|
62
|
+
remaining risk.
|
|
63
|
+
|
|
64
|
+
## Vendor cheat sheet
|
|
65
|
+
|
|
66
|
+
The sections below describe how each upstream vendor documents its own caching. They are
|
|
67
|
+
**reference material only** — see "Upstream docs are not provider guarantees" below. Provider
|
|
68
|
+
minimums and rates change; treat the shapes here as orientation and **verify against the current
|
|
69
|
+
official docs** (links in each section) before relying on a number.
|
|
70
|
+
|
|
71
|
+
### Anthropic
|
|
72
|
+
|
|
73
|
+
- **Mechanism:** explicit cache breakpoints (`cache_control: {"type": "ephemeral"}`) on cacheable
|
|
74
|
+
content blocks; a single top-level marker enables automatic caching on the last eligible block.
|
|
75
|
+
- **Minimum cacheable size:** model-dependent, roughly 512 to 4096 input tokens; shorter prompts
|
|
76
|
+
are silently not cached.
|
|
77
|
+
- **Discount shape:** cache reads are billed at a small fraction of the normal input rate; cache
|
|
78
|
+
writes carry a premium; an optional longer TTL costs more to write.
|
|
79
|
+
- **Docs:** https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ·
|
|
80
|
+
https://docs.anthropic.com/en/docs/about-claude/pricing
|
|
81
|
+
|
|
82
|
+
### OpenAI
|
|
83
|
+
|
|
84
|
+
- **Mechanism:** implicit (automatic) caching; recent models also accept explicit cache markers,
|
|
85
|
+
and a prompt cache key can steer routing or cache accounting.
|
|
86
|
+
- **Minimum cacheable size:** roughly 1024 visible input tokens on recent models; varies with
|
|
87
|
+
request settings on older ones.
|
|
88
|
+
- **Discount shape:** cached reads are heavily discounted relative to input; cache writes are
|
|
89
|
+
billed at a modest premium on recent models and carry no extra write charge on older ones.
|
|
90
|
+
- **Docs:** https://developers.openai.com/api/docs/guides/prompt-caching
|
|
91
|
+
|
|
92
|
+
### DeepSeek
|
|
93
|
+
|
|
94
|
+
- **Mechanism:** automatic, best-effort prefix caching; a cached prefix is an indivisible unit, so
|
|
95
|
+
partial overlap does not hit.
|
|
96
|
+
- **Minimum cacheable size:** not documented.
|
|
97
|
+
- **Discount shape:** a separate lower per-model cache-hit rate versus the cache-miss rate.
|
|
98
|
+
- **Note:** with tools present, reasoning content must be passed back on every later request.
|
|
99
|
+
- **Docs:** https://api-docs.deepseek.com/guides/kv_cache
|
|
100
|
+
|
|
101
|
+
### GLM
|
|
102
|
+
|
|
103
|
+
- **Mechanism:** implicit caching triggered by content similarity; no explicit create or
|
|
104
|
+
invalidate API is documented.
|
|
105
|
+
- **Minimum cacheable size:** not documented.
|
|
106
|
+
- **Discount shape:** a separate per-model cached-input rate, not a universal ratio — do not
|
|
107
|
+
assume a fixed percentage.
|
|
108
|
+
- **Docs:** https://docs.z.ai/guides/capabilities/cache
|
|
109
|
+
|
|
110
|
+
### MiniMax
|
|
111
|
+
|
|
112
|
+
- **Mechanism:** passive automatic prefix caching, plus explicit Anthropic-compatible
|
|
113
|
+
`cache_control` breakpoints on cacheable blocks.
|
|
114
|
+
- **Minimum cacheable size:** caching applies from roughly 512 tokens upward.
|
|
115
|
+
- **Discount shape:** explicit cache writes are billed at a premium and reads at a fraction of
|
|
116
|
+
input; passive cache writes carry no additional charge.
|
|
117
|
+
- **Docs:** https://platform.minimax.io/docs/api-reference/text-prompt-caching.md
|
|
118
|
+
|
|
119
|
+
## Upstream docs are not provider guarantees
|
|
120
|
+
|
|
121
|
+
A vendor documenting a cache feature does **not** mean the gateway or route you actually use
|
|
122
|
+
forwards it, returns the same usage fields, or bills it the same way. Treat every statement about
|
|
123
|
+
cache behavior behind a gateway as a hypothesis, not a guarantee: verify cache parameters and
|
|
124
|
+
usage counters against the raw response of the exact endpoint you call, and against the current
|
|
125
|
+
official docs linked above. If a usage counter is absent, it is **unknown** — never a reported
|
|
126
|
+
zero (CTX-09). When project information changes mid-session, prefer correctness over a preserved
|
|
127
|
+
prefix: send the new version and accept whatever cache invalidation follows.
|