@ngockhoale/ukit 2.3.19 → 2.3.21

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,55 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.21 - 2026-09-13
6
+
7
+ C16 release — the prompt-caching ruleset turns from research into shipped guidance, and the
8
+ prompt-assembly surfaces that inject into model context are made deterministic. No new commands
9
+ and no default runtime behavior changes: the tool-call policy ships as guidance only, pending the
10
+ A/B required before any behavior change.
11
+
12
+ - **Canonical prompt-caching ruleset (`docs/PROMPT_CACHING.md`, repo-local).** Distills the UNIC
13
+ caching roadmap into CTX-01..10 (MUST), the SHOULD set, and the never-do list, with
14
+ evidence-labeled vendor sections (Anthropic, OpenAI, DeepSeek, GLM, MiniMax) and an
15
+ adopt/adapt/reject matrix. Upstream vendor docs are reference only — never UNIC guarantees
16
+ (UNIC behavior is labeled `unknown`/`inferred`). Repo-local by design; not shipped.
17
+ - **Shipped caching guidance (`templates/docs/PROMPT_CACHING.md`).** The distilled CTX rules,
18
+ never-do list, and short tool-call policy now ship with UKit installs through a new manifest
19
+ item `docs-prompt-caching` (`mergeStrategy: overwrite_with_backup`, so `ukit update` refreshes
20
+ it). `templates/CLAUDE.md` and `templates/AGENTS.md` gain a matching `## Prompt Caching`
21
+ pointer section, and the repo-local `CLAUDE.md`/`AGENTS.md` carry a minimal pointer too.
22
+ - **Deterministic prompt assembly.** `templates/.claude/hooks/skill-router.sh` ordered
23
+ model-visible memory-recall output by volatile `updatedAt` timestamps, so the same state could
24
+ assemble different prompt bytes across runs and defeat prompt caching (CTX-01/CTX-04). Ordering
25
+ is now a stable sort with a deterministic tiebreak and no volatile timestamps in model-visible
26
+ output; the live `.claude/hooks/` twin is byte-identical. Locked by the new
27
+ `tests/hooks/promptAssemblyDeterminism.test.js` plus an ordering regression in
28
+ `tests/hooks/skillRouterHook.test.js`.
29
+
30
+ ## 2.3.20 - 2026-09-13
31
+
32
+ C15 bug-fix release record — this release completes the silent-stop class on top of the
33
+ already-published 2.3.19 (published to npm 2026-09-13, registry verified, carrying the C14
34
+ valve removal); it adds no new features. npm is the distribution channel and `ukit update`
35
+ pulls from npm, so an install on 2.3.19 (or older) needs an update to receive these fixes.
36
+
37
+ - **`--evaluate-stop` never releases a routed stop silently.**
38
+ `templates/.claude/ukit/runtime/execution-ledger.mjs` had a final dispatch with no `else`, so any
39
+ evaluation result of `{continue:false, notify:false}` that was not `capped` wrote nothing at all
40
+ — the Stop was allowed and the session idled with no visible reason. That default path is now a
41
+ loud release: an invisible blocker and a route with empty/short completion evidence both emit a
42
+ `systemMessage` naming the cause and how to resume, instead of silence. Malformed stdin no longer
43
+ crashes the hook into a silent stderr exit — it blocks with the crash detail, bounded by a crash
44
+ counter that loudly releases after the threshold and resets on a clean evaluate.
45
+ - **Fail-loud completion-gate wrappers.**
46
+ `templates/.claude/hooks/completion-gate.sh` and `templates/.claude/hooks/record-execution.sh`
47
+ exited 0 silently when the runtime script was missing or failed; both now emit a
48
+ `systemMessage` naming the missing/failing runtime and the `ukit install` remedy.
49
+ - **omp reentrant stop gets the same never-silent contract as the CLI.**
50
+ `templates/.omp/hooks/pre/ukit-bridge.js` released a reentrant stop with a non-blocking notice
51
+ and no continuation; it now blocks with the actionable reason like any other stop, stays bounded
52
+ by the ledger's continuation cap, and releases loudly at the cap.
53
+
5
54
  ## 2.3.19 - 2026-09-12
6
55
 
7
56
  C14 bug-fix release record — one correctness fix is shipped; this release adds no new features.
@@ -218,6 +218,20 @@ items:
218
218
  packs:
219
219
  - core
220
220
 
221
+ # Shipped prompt-caching guidance (CTX-01..10 + never-do + tool-call policy + vendor cheat
222
+ # sheet). `overwrite_with_backup` so `ukit update` refreshes the ruleset — deliberately NOT
223
+ # `docs/PROJECT.md`, which is `mergeStrategy: skip` and must never be auto-overwritten.
224
+ - id: docs-prompt-caching
225
+ type: config
226
+ sourceTemplate: docs/PROMPT_CACHING.md
227
+ targetPath: docs/PROMPT_CACHING.md
228
+ requires: []
229
+ mergeStrategy: overwrite_with_backup
230
+ variables: []
231
+ enabledByDefault: true
232
+ packs:
233
+ - core
234
+
221
235
  - id: core-skill-delivery
222
236
  type: skill
223
237
  sourceTemplate: .claude/skills/delivery/SKILL.md
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.19",
3
+ "version": "2.3.21",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -1,10 +1,35 @@
1
1
  import { spawn } from 'node:child_process';
2
+ import { readFileSync } from 'node:fs';
2
3
  import os from 'node:os';
3
4
  import path from 'node:path';
4
5
 
5
6
  const rootDir = process.cwd();
6
7
  const npmCacheDir = path.join(os.tmpdir(), 'ukit-npm-cache');
7
8
 
9
+ // Opt-in post-publish registry-parity check (Release Policy invariant: npm and git always hold
10
+ // the same latest version). Only meaningful AFTER `npm publish` — before publish the registry is
11
+ // behind by definition — so it is gated behind --post-publish and never runs by default.
12
+ const postPublish = process.argv.includes('--post-publish');
13
+
14
+ if (postPublish) {
15
+ const localVersion = JSON.parse(readFileSync(path.join(rootDir, 'package.json'), 'utf8')).version;
16
+ const registryVersion = await readRegistryVersion();
17
+
18
+ console.log(`[release:verify] Registry parity (post-publish)`);
19
+ console.log(`[release:verify] package.json = ${localVersion}`);
20
+ console.log(`[release:verify] npm registry = ${registryVersion}`);
21
+
22
+ if (registryVersion !== localVersion) {
23
+ console.error(
24
+ `[release:verify] FAILED: registry version ${registryVersion} !== package.json ${localVersion}`,
25
+ );
26
+ process.exit(1);
27
+ }
28
+
29
+ console.log('\n[release:verify] All release checks passed.');
30
+ process.exit(0);
31
+ }
32
+
8
33
  const steps = [
9
34
  {
10
35
  label: 'Artifact smoke',
@@ -54,3 +79,23 @@ function runStep({ command, args, env }) {
54
79
  child.on('error', () => resolve(1));
55
80
  });
56
81
  }
82
+
83
+ function readRegistryVersion() {
84
+ return new Promise((resolve) => {
85
+ const child = spawn('npm', ['view', '@ngockhoale/ukit', 'version'], {
86
+ cwd: rootDir,
87
+ env: { ...process.env, npm_config_cache: npmCacheDir },
88
+ stdio: ['ignore', 'pipe', 'inherit'],
89
+ });
90
+
91
+ let stdout = '';
92
+ child.stdout.on('data', (chunk) => {
93
+ stdout += chunk;
94
+ });
95
+ child.on('close', () => {
96
+ const version = stdout.trim();
97
+ resolve(version.length > 0 ? version : '<unavailable>');
98
+ });
99
+ child.on('error', () => resolve('<unavailable>'));
100
+ });
101
+ }
@@ -1,12 +1,24 @@
1
1
  #!/bin/bash
2
2
  # Stop hook: block premature terminal stops while routed completion evidence is missing.
3
+ # ADVISORY ONLY — always exit 0. A missing or failing runtime must be loud, not a silent pass.
3
4
 
4
5
  INPUT=$(cat)
5
6
  PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
6
7
  SCRIPT="$PROJECT_ROOT/.claude/ukit/runtime/execution-ledger.mjs"
7
8
 
8
9
  if [ ! -f "$SCRIPT" ]; then
10
+ printf '%s\n' '{"systemMessage":"UKit completion gate: runtime script missing — run: ukit install"}'
9
11
  exit 0
10
12
  fi
11
13
 
12
- printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --evaluate-stop
14
+ OUTPUT=$(printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --evaluate-stop)
15
+ STATUS=$?
16
+
17
+ if [ "$STATUS" -ne 0 ]; then
18
+ printf '%s\n' '{"systemMessage":"UKit completion gate: runtime script failed — run: ukit install"}'
19
+ exit 0
20
+ fi
21
+
22
+ if [ -n "$OUTPUT" ]; then
23
+ printf '%s\n' "$OUTPUT"
24
+ fi
@@ -1,12 +1,24 @@
1
1
  #!/bin/bash
2
2
  # PostToolUse hook: persist session-scoped source/write/verification receipts.
3
+ # ADVISORY ONLY — always exit 0. A missing or failing runtime must be loud, not a silent pass.
3
4
 
4
5
  INPUT=$(cat)
5
6
  PROJECT_ROOT="${CLAUDE_PROJECT_DIR:-$PWD}"
6
7
  SCRIPT="$PROJECT_ROOT/.claude/ukit/runtime/execution-ledger.mjs"
7
8
 
8
9
  if [ ! -f "$SCRIPT" ]; then
10
+ printf '%s\n' '{"systemMessage":"UKit record execution: runtime script missing — run: ukit install"}'
9
11
  exit 0
10
12
  fi
11
13
 
12
- printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --record
14
+ OUTPUT=$(printf '%s' "$INPUT" | UKIT_HARNESS=claude-code node "$SCRIPT" --record)
15
+ STATUS=$?
16
+
17
+ if [ "$STATUS" -ne 0 ]; then
18
+ printf '%s\n' '{"systemMessage":"UKit record execution: runtime script failed — run: ukit install"}'
19
+ exit 0
20
+ fi
21
+
22
+ if [ -n "$OUTPUT" ]; then
23
+ printf '%s\n' "$OUTPUT"
24
+ fi
@@ -961,8 +961,12 @@ const { pathToFileURL } = require('url');
961
961
  return 0;
962
962
  }
963
963
 
964
- const recencyBonus = getMemoryTimestamp(item) > 0 ? Math.min(1, getMemoryTimestamp(item) / Date.now()) : 0;
965
- return score + recencyBonus;
964
+ // TASK-027 CTX-01/CTX-04: the ranking must be a pure function of the logical memory
965
+ // content. A recency bonus scaled by Date.now() made the score — and therefore the
966
+ // injected previous-context block, its persisted order, and the PreCompact reinjection —
967
+ // shift on every wall-clock tick. Eligibility is unchanged (still score > 0); only the
968
+ // volatile ordering signal is dropped.
969
+ return score;
966
970
  }
967
971
 
968
972
  function buildPreviousContextSnippet(item) {
@@ -1017,9 +1021,12 @@ const { pathToFileURL } = require('url');
1017
1021
  score: scoreMemoryItem(item, queryTokens),
1018
1022
  }))
1019
1023
  .filter((entry) => entry.score > 0)
1024
+ // CTX-01: ties break on the stable logical id (never on a volatile timestamp), so the
1025
+ // same memory set always yields the same injected order regardless of when the
1026
+ // entries were last written/ended.
1020
1027
  .sort((left, right) => (
1021
1028
  right.score - left.score
1022
- || getMemoryTimestamp(right.item) - getMemoryTimestamp(left.item)
1029
+ || left.item.id.localeCompare(right.item.id)
1023
1030
  ))
1024
1031
  .slice(0, 2)
1025
1032
  .map((entry) => entry.item);
@@ -13,6 +13,12 @@ const RESUME_INTENT_TTL_MS = 30 * 60 * 1000;
13
13
  const MAX_RECEIPTS = 24;
14
14
  const MAX_SOURCE_FILES = 16;
15
15
  const MAX_CONTINUATIONS = 6;
16
+ // A malformed stdin payload crashes `--evaluate-stop` before a route/session can be read, so
17
+ // the only identity available is the project root. Count consecutive crashes there and, once
18
+ // this bounded budget is exhausted, release LOUD (naming the crash + `ukit install`) instead of
19
+ // blocking forever on a corrupt hook input. A successful evaluate resets the count. Mirrors
20
+ // MAX_CONTINUATIONS style — a hard-coded module const, deliberately not a config knob.
21
+ const MAX_GATE_CRASHES = 3;
16
22
  // Two adjacent identical failed verifications (no edit attempt between them) are a loop;
17
23
  // vibecode routes — which bypass the cap and every reentrant valve — additionally get a
18
24
  // cumulative escape after the same verification failed this many times across edits.
@@ -82,6 +88,17 @@ function ledgerPath(projectRoot, payload = {}) {
82
88
  );
83
89
  }
84
90
 
91
+ function crashCounterPath(projectRoot) {
92
+ return path.join(
93
+ projectRoot,
94
+ '.ukit',
95
+ 'storage',
96
+ 'cache',
97
+ 'exec-ledger',
98
+ 'gate-crash-counter.json',
99
+ );
100
+ }
101
+
85
102
  function resumeIntentPath(projectRoot, sessionId) {
86
103
  return path.join(
87
104
  projectRoot,
@@ -727,10 +744,32 @@ export function evaluateCompletion({ state = {}, ledger = {} } = {}) {
727
744
  reason: `UKit stopped automatic recovery: ${ledger.blocker.detail || 'a recorded blocker ended this run.'}`,
728
745
  };
729
746
  }
747
+ // An invisible blocker (e.g. a permission decision the model already named in its own
748
+ // reply) must not loop — but it must not end silent either. Release loud so the user can
749
+ // tell a deliberate stop from a stall.
750
+ return {
751
+ continue: false,
752
+ notify: true,
753
+ missingEvidence: [],
754
+ reason: `UKit stopped automatic recovery: ${ledger.blocker.detail || 'a recorded blocker ended this run.'}`,
755
+ };
756
+ }
757
+ if (!state || !state.routeSummary) {
758
+ // No routable state at all: a never-routed session (or a foreign-owned one) is silent
759
+ // success — there is nothing to report. The CLI's lost-route path handles the case where
760
+ // THIS session has ledger activity but the state vanished; callers that pass no state
761
+ // (e.g. the omp bridge for a session it does not own) must not be force-notified.
730
762
  return { continue: false, notify: false, missingEvidence: [] };
731
763
  }
732
764
  if (evidence.length === 0) {
733
- return { continue: false, notify: false, missingEvidence: [] };
765
+ // A route with no completion evidence has no actionable block reason — blocking would
766
+ // loop unboundedly. Release loud instead of ending indistinguishable from a stall.
767
+ return {
768
+ continue: false,
769
+ notify: true,
770
+ missingEvidence: [],
771
+ reason: 'UKit completion gate: this route carries no completion evidence to verify, so the stop released without checking unfinished work. Send the task again in a new message to re-route it with a completion contract.',
772
+ };
734
773
  }
735
774
 
736
775
  // Request identity is the prompt (see evidencePromptKey): the router re-keys requestKey on
@@ -742,7 +781,10 @@ export function evaluateCompletion({ state = {}, ledger = {} } = {}) {
742
781
  const effectiveLedger = sameRequest ? ledger : {};
743
782
  const missingEvidence = evidence.filter((item) => !evidenceSatisfied(item, effectiveLedger, routeSummary));
744
783
  if (missingEvidence.length === 0) {
745
- return { continue: false, notify: false, missingEvidence: [] };
784
+ // Silent success: the route is present and every required evidence is satisfied. Marked
785
+ // `complete` so the CLI dispatch recognizes it BEFORE the loud final else — otherwise a
786
+ // bare loud else would turn a clean stop into user-visible noise.
787
+ return { continue: false, notify: false, complete: true, missingEvidence: [] };
746
788
  }
747
789
 
748
790
  // `find-cause` can validly end clean: an investigation may establish that no actionable
@@ -907,7 +949,121 @@ async function readStdin() {
907
949
  return chunks.join('');
908
950
  }
909
951
 
952
+ // Malformed hook input crashes the evaluate BEFORE any route/session can be read, so the only
953
+ // stable identity is the project root. Count consecutive crashes there; past MAX_GATE_CRASHES,
954
+ // release LOUD naming the crash and the remedy instead of blocking a corrupt hook forever. Any
955
+ // counter I/O failure fails SAFE to block (still loud, never silent). A successful evaluate
956
+ // resets the count.
957
+ async function handleEvaluateCrash(projectRoot, error) {
958
+ const detail = `malformed hook input crashed the completion gate (${error?.message || error})`;
959
+ const blockReason = () => `${detail}. This is the UKit completion gate, not the task — re-send the task in a new message so the hook payload is regenerated.`;
960
+ let count = 0;
961
+ try {
962
+ const current = await readJson(crashCounterPath(projectRoot), null);
963
+ count = Number(current?.count) || 0;
964
+ } catch {
965
+ count = 0;
966
+ }
967
+ const next = count + 1;
968
+ try {
969
+ await writeJsonAtomic(crashCounterPath(projectRoot), { count: next, updatedAt: Date.now() });
970
+ } catch {
971
+ // Counter unwritable: we cannot bound the crash loop, but silence is never an option.
972
+ process.stdout.write(`${JSON.stringify({ decision: 'block', reason: blockReason() })}\n`);
973
+ return;
974
+ }
975
+ if (next >= MAX_GATE_CRASHES) {
976
+ const message = `UKit completion gate: ${detail}. This recurred ${next} times and cannot be bounded, so the stop is releasing loudly instead of blocking again. The hook payload is corrupt — run: ukit install to refresh the runtime, then re-send the task in a new message.`;
977
+ process.stderr.write(`[ukit-completion] ${message}\n`);
978
+ process.stdout.write(`${JSON.stringify({ systemMessage: message })}\n`);
979
+ return;
980
+ }
981
+ process.stdout.write(`${JSON.stringify({ decision: 'block', reason: blockReason() })}\n`);
982
+ }
983
+
984
+ async function resetGateCrashCounter(projectRoot) {
985
+ // Only touch disk when a crash was actually counted — a clean stop on a healthy project
986
+ // must not create/rewrite a crash-counter file.
987
+ try {
988
+ await fs.access(crashCounterPath(projectRoot));
989
+ } catch {
990
+ return;
991
+ }
992
+ try {
993
+ await writeJsonAtomic(crashCounterPath(projectRoot), { count: 0, updatedAt: Date.now() });
994
+ } catch {
995
+ // Best-effort: an unwritable counter only means the next crash blocks again, which is safe.
996
+ }
997
+ }
998
+
999
+ async function runEvaluateStop() {
1000
+ const projectRoot = process.env.CLAUDE_PROJECT_DIR || process.cwd();
1001
+ let payload;
1002
+ try {
1003
+ payload = JSON.parse((await readStdin()) || '{}');
1004
+ } catch (error) {
1005
+ await handleEvaluateCrash(projectRoot, error);
1006
+ return;
1007
+ }
1008
+ // A parse that reached here is a functioning evaluate — reset the crash budget so the next
1009
+ // malformed payload starts blocking again.
1010
+ await resetGateCrashCounter(projectRoot);
1011
+
1012
+ const state = await readRouteState(projectRoot, payload);
1013
+ const ledger = await readExecutionLedger(projectRoot, payload) || {};
1014
+
1015
+ // A session ledger with activity proves this session was routed and accumulating evidence.
1016
+ // If the route state is now missing or unreadable (router crashed/timed out before
1017
+ // persisting it, or the session stamp was lost), the gate cannot evaluate completion — and
1018
+ // ending silently there is indistinguishable from a mid-run stall. Surface the loss so the
1019
+ // user sees WHY the run released. Sessions that were never routed (no ledger activity at
1020
+ // all) stay silent: there is nothing to report (silent-success carve-out).
1021
+ if (!state) {
1022
+ const hasActivity = ledger.sourceSucceeded || ledger.writeAttempted
1023
+ || ledger.verificationAttempted || (ledger.receipts || []).length > 0;
1024
+ if (!hasActivity) return;
1025
+ process.stdout.write(`${JSON.stringify({
1026
+ systemMessage: 'UKit completion gate: route state for this session was lost before the stop was evaluated (router did not persist it), so unfinished work could not be verified. Send the task again in a new message to re-route and finish it.',
1027
+ })}\n`);
1028
+ return;
1029
+ }
1030
+
1031
+ const result = evaluateCompletion({ state, ledger });
1032
+
1033
+ // A reentrant Stop (Claude Code re-fires Stop after this hook already blocked once) is
1034
+ // gated exactly like any other Stop: while evidence is still missing it blocks again with
1035
+ // the actionable reason instead of silently releasing the recovery turn. Loop termination
1036
+ // is guaranteed elsewhere — non-vibecode routes end at the continuation cap (final-notice
1037
+ // block -> capped visible release), and vibecode ends on the visible verification-loop
1038
+ // blocker.
1039
+ if (result.continue) {
1040
+ if (result.finalNotice) await markNotified(projectRoot, payload, ledger);
1041
+ else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
1042
+ process.stdout.write(`${JSON.stringify({ decision: 'block', reason: result.reason })}\n`);
1043
+ return;
1044
+ }
1045
+ if (result.capped || result.notify) {
1046
+ // Non-blocking endings (non-gated modes, an invisible blocker, an evidence-less route, or
1047
+ // cap reached after the final notice) must still tell the user what is unfinished — a
1048
+ // silent end is indistinguishable from a stall.
1049
+ process.stderr.write(`[ukit-completion] ${result.reason}\n`);
1050
+ process.stdout.write(`${JSON.stringify({ systemMessage: result.reason })}\n`);
1051
+ return;
1052
+ }
1053
+ if (result.complete) return; // silent success: route present, all evidence satisfied
1054
+ // Defensive final else: no evaluation result may ever end a routed stop silently.
1055
+ const fallback = `UKit completion gate: the stop was evaluated without a decision (missing evidence: ${(result.missingEvidence || []).join(', ') || 'none'}). If work is unfinished, re-send the task in a new message.`;
1056
+ process.stderr.write(`[ukit-completion] ${fallback}\n`);
1057
+ process.stdout.write(`${JSON.stringify({ systemMessage: fallback })}\n`);
1058
+ }
1059
+
910
1060
  async function main() {
1061
+ // Dispatch the stop flag BEFORE parsing stdin: a malformed payload used to throw out of
1062
+ // main(), hit the .catch (stderr + exit 1) with no decision, and release the stop silently.
1063
+ if (process.argv.includes('--evaluate-stop')) {
1064
+ await runEvaluateStop();
1065
+ return;
1066
+ }
911
1067
  const payload = JSON.parse((await readStdin()) || '{}');
912
1068
  const projectRoot = process.env.CLAUDE_PROJECT_DIR || payload.cwd || process.cwd();
913
1069
  if (process.argv.includes('--record')) {
@@ -917,44 +1073,6 @@ async function main() {
917
1073
  toolName: payload.tool_name,
918
1074
  harness: process.env.UKIT_HARNESS || 'claude-code',
919
1075
  });
920
- return;
921
- }
922
- if (process.argv.includes('--evaluate-stop')) {
923
- const state = await readRouteState(projectRoot, payload);
924
- const ledger = await readExecutionLedger(projectRoot, payload) || {};
925
- const result = evaluateCompletion({ state, ledger });
926
-
927
- // A session ledger with activity proves this session was routed and accumulating
928
- // evidence. If the route state is now missing or unreadable (router crashed/timed out
929
- // before persisting it, or the session stamp was lost), the gate cannot evaluate
930
- // completion — and ending silently there is indistinguishable from a mid-run stall.
931
- // Surface the loss so the user sees WHY the run released. Sessions that were never
932
- // routed (no ledger activity at all) stay silent: there is nothing to report.
933
- if (!state && (ledger.sourceSucceeded || ledger.writeAttempted
934
- || ledger.verificationAttempted || (ledger.receipts || []).length > 0)) {
935
- process.stdout.write(`${JSON.stringify({
936
- systemMessage: 'UKit completion gate: route state for this session was lost before the stop was evaluated (router did not persist it), so unfinished work could not be verified. Send the task again in a new message to re-route and finish it.',
937
- })}\n`);
938
- return;
939
- }
940
-
941
- // A reentrant Stop (Claude Code re-fires Stop after this hook already blocked once) is
942
- // gated exactly like any other Stop: while evidence is still missing it blocks again
943
- // with the actionable reason instead of silently releasing the recovery turn. Loop
944
- // termination is guaranteed elsewhere — non-vibecode routes end at the continuation cap
945
- // (final-notice block -> capped visible release), and vibecode ends on the visible
946
- // verification-loop blocker.
947
-
948
- if (result.continue) {
949
- if (result.finalNotice) await markNotified(projectRoot, payload, ledger);
950
- else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
951
- process.stdout.write(`${JSON.stringify({ decision: 'block', reason: result.reason })}\n`);
952
- } else if (result.capped || result.notify) {
953
- // Non-blocking endings (non-gated modes, or cap reached after the final notice) must
954
- // still tell the user what is unfinished — a silent end is indistinguishable from a stall.
955
- process.stderr.write(`[ukit-completion] ${result.reason}\n`);
956
- process.stdout.write(`${JSON.stringify({ systemMessage: result.reason })}\n`);
957
- }
958
1076
  }
959
1077
  }
960
1078
 
@@ -14,7 +14,6 @@ import {
14
14
  evidencePromptKey,
15
15
  incrementContinuation,
16
16
  markNotified,
17
- noteStopProgress,
18
17
  readExecutionLedger,
19
18
  readRouteState,
20
19
  recordExecutionReceipt,
@@ -634,23 +633,6 @@ export async function runSessionStop(
634
633
  }
635
634
 
636
635
  if (suppliedLedger === undefined) {
637
- // Claude Code flags reentrant stops with stop_hook_active and the ledger CLI releases on
638
- // them; omp's session_stop carries no such field. Detect the equivalent from the ledger:
639
- // when the recovery turn produced no new receipts since the stop that blocked it,
640
- // blocking again would only burn another turn in a self-sustaining loop — release with a
641
- // visible reason instead. Vibecode autonomy keeps pushing (same as the CLI valve).
642
- let reentrantStop = false;
643
- try {
644
- reentrantStop = (await noteStopProgress(projectRoot, payload)).reentrant === true;
645
- } catch {
646
- reentrantStop = false;
647
- }
648
- if (reentrantStop && state?.routeSummary?.autonomyLevel !== 'vibecode') {
649
- const notice = `UKit stopped automatic recovery after one continuation: ${evaluation.reason}`;
650
- pi.logger?.warn?.(`[UKit] ${notice}`);
651
- sendContext(pi, [notice], 'nextTurn', { display: true });
652
- return undefined;
653
- }
654
636
  try {
655
637
  if (evaluation.finalNotice) await markNotified(projectRoot, payload, ledger);
656
638
  else await incrementContinuation(projectRoot, payload, ledger, state?.requestKey || null, evidencePromptKey(state));
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -0,0 +1,127 @@
1
+ # Prompt Caching — guidance for UKit projects
2
+
3
+ Prompt caching is a **prefix match**: a provider can reuse the computation of an input prefix
4
+ when a later request reproduces that prefix byte-for-byte. UKit cannot control the transport or
5
+ the provider cache engine, but it *does* control the instruction, skill, tool and hook-injected
6
+ content it renders. This file ships with UKit so every installed project gets the same rules for
7
+ keeping that content deterministic and stable. Read it on demand — it is deliberately not loaded
8
+ into every session.
9
+
10
+ ## Why stable context matters
11
+
12
+ - Cache reuse is **reported**, not controllable. A `cache_read` counter (or a vendor equivalent)
13
+ proves the provider *reported* a reuse; it never proves which layer served it.
14
+ - Three mechanisms must never be conflated: the **provider prompt cache** (reuses input
15
+ computation), a **local tool-result cache** (the host reuses a still-valid result), and a
16
+ **response cache** (the application replays a stored answer).
17
+ - Three counts must never be conflated: a **logical model turn**, a **client HTTP attempt**, and a
18
+ **tool execution**.
19
+ - The rules below hold even if no cache capability exists behind your provider, because they
20
+ govern content UKit itself renders.
21
+
22
+ ## CTX rules — MUST
23
+
24
+ | ID | Rule | What it means in practice |
25
+ |---|---|---|
26
+ | CTX-01 | Same logical input and config produce the same segment bytes | Render owned blocks deterministically; no ordering that varies run-to-run |
27
+ | CTX-02 | Preserve instruction/user/tool roles and conversation order | Never move a user ask into system or reorder history to lengthen a prefix |
28
+ | CTX-03 | Keep all tool IDs and required native continuation state intact | Never strip tool-call IDs or reasoning/continuation fields to shrink a request |
29
+ | CTX-04 | Do not inject clock/random IDs into static instructions | Static blocks carry no date, UUID or counter; volatile values go in the tail |
30
+ | CTX-05 | Every summary/compaction creates a new context epoch | Compaction is a decision with a cost; when it happens it starts a new epoch |
31
+ | CTX-06 | Never change data or code to match the cache | Cache optimization never edits content semantics — correctness over prefix |
32
+ | CTX-07 | Do not self-send a field the adapter has not confirmed as supported | Optional cache params only when capability is verified; HTTP 200 is not proof |
33
+ | CTX-08 | Tool-result cache reuse only when freshness/dependency is valid | Local reuse needs a freshness predicate and dependency fingerprint, not just a query match |
34
+ | CTX-09 | Never treat missing usage as zero | Missing cache counters are unknown plus lowered coverage, never a reported miss |
35
+ | CTX-10 | Never exceed existing instructions/permissions to cut calls | Reducing tool calls never means skipping a required check or test |
36
+
37
+ ## Never do
38
+
39
+ - Never sort messages or reasoning blocks alphabetically.
40
+ - Never trim code literals, signed content, or data where whitespace is meaningful.
41
+ - Never rewrite native reasoning/signature/encrypted continuation fields.
42
+ - Never move a user request into system to lengthen the stable prefix.
43
+ - Never promote untrusted documents or tool output to the developer/system role.
44
+ - Never hash one text and assume every occurrence shares the provider cache.
45
+ - Never add a timestamp just to log — keep log metadata outside model-visible content.
46
+
47
+ ## Runtime tool-call policy (condensed, guidance only)
48
+
49
+ This is guidance for how an assistant should decide when to call tools. It is **not** a rigid
50
+ "always at least N tools" or "max N tools" rule, and UKit does not change any default behavior
51
+ on its basis.
52
+
53
+ 1. Before calling a tool, decide which data is still missing to finish the request.
54
+ 2. Reuse evidence you already hold if it is still valid; check the version before reuse.
55
+ 3. Batch independent reads when the interface supports it.
56
+ 4. For dependent actions, wait for the needed result before deciding the next step.
57
+ 5. Prefer scoped queries with enough output to verify.
58
+ 6. If a result was truncated or insufficient, widen deliberately.
59
+ 7. On tool error, distinguish parameter error, transient error, and unknown state.
60
+ 8. After a change, run the check appropriate to the risk and the repo's requirements.
61
+ 9. When the completion condition is met, return the result — check further only for specific
62
+ remaining risk.
63
+
64
+ ## Vendor cheat sheet
65
+
66
+ The sections below describe how each upstream vendor documents its own caching. They are
67
+ **reference material only** — see "Upstream docs are not provider guarantees" below. Provider
68
+ minimums and rates change; treat the shapes here as orientation and **verify against the current
69
+ official docs** (links in each section) before relying on a number.
70
+
71
+ ### Anthropic
72
+
73
+ - **Mechanism:** explicit cache breakpoints (`cache_control: {"type": "ephemeral"}`) on cacheable
74
+ content blocks; a single top-level marker enables automatic caching on the last eligible block.
75
+ - **Minimum cacheable size:** model-dependent, roughly 512 to 4096 input tokens; shorter prompts
76
+ are silently not cached.
77
+ - **Discount shape:** cache reads are billed at a small fraction of the normal input rate; cache
78
+ writes carry a premium; an optional longer TTL costs more to write.
79
+ - **Docs:** https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ·
80
+ https://docs.anthropic.com/en/docs/about-claude/pricing
81
+
82
+ ### OpenAI
83
+
84
+ - **Mechanism:** implicit (automatic) caching; recent models also accept explicit cache markers,
85
+ and a prompt cache key can steer routing or cache accounting.
86
+ - **Minimum cacheable size:** roughly 1024 visible input tokens on recent models; varies with
87
+ request settings on older ones.
88
+ - **Discount shape:** cached reads are heavily discounted relative to input; cache writes are
89
+ billed at a modest premium on recent models and carry no extra write charge on older ones.
90
+ - **Docs:** https://developers.openai.com/api/docs/guides/prompt-caching
91
+
92
+ ### DeepSeek
93
+
94
+ - **Mechanism:** automatic, best-effort prefix caching; a cached prefix is an indivisible unit, so
95
+ partial overlap does not hit.
96
+ - **Minimum cacheable size:** not documented.
97
+ - **Discount shape:** a separate lower per-model cache-hit rate versus the cache-miss rate.
98
+ - **Note:** with tools present, reasoning content must be passed back on every later request.
99
+ - **Docs:** https://api-docs.deepseek.com/guides/kv_cache
100
+
101
+ ### GLM
102
+
103
+ - **Mechanism:** implicit caching triggered by content similarity; no explicit create or
104
+ invalidate API is documented.
105
+ - **Minimum cacheable size:** not documented.
106
+ - **Discount shape:** a separate per-model cached-input rate, not a universal ratio — do not
107
+ assume a fixed percentage.
108
+ - **Docs:** https://docs.z.ai/guides/capabilities/cache
109
+
110
+ ### MiniMax
111
+
112
+ - **Mechanism:** passive automatic prefix caching, plus explicit Anthropic-compatible
113
+ `cache_control` breakpoints on cacheable blocks.
114
+ - **Minimum cacheable size:** caching applies from roughly 512 tokens upward.
115
+ - **Discount shape:** explicit cache writes are billed at a premium and reads at a fraction of
116
+ input; passive cache writes carry no additional charge.
117
+ - **Docs:** https://platform.minimax.io/docs/api-reference/text-prompt-caching.md
118
+
119
+ ## Upstream docs are not provider guarantees
120
+
121
+ A vendor documenting a cache feature does **not** mean the gateway or route you actually use
122
+ forwards it, returns the same usage fields, or bills it the same way. Treat every statement about
123
+ cache behavior behind a gateway as a hypothesis, not a guarantee: verify cache parameters and
124
+ usage counters against the raw response of the exact endpoint you call, and against the current
125
+ official docs linked above. If a usage counter is absent, it is **unknown** — never a reported
126
+ zero (CTX-09). When project information changes mid-session, prefer correctness over a preserved
127
+ prefix: send the new version and accept whatever cache invalidation follows.