@ngockhoale/ukit 2.3.20 → 2.3.21

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,31 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.21 - 2026-09-13
6
+
7
+ C16 release — the prompt-caching ruleset turns from research into shipped guidance, and the
8
+ prompt-assembly surfaces that inject into model context are made deterministic. No new commands
9
+ and no default runtime behavior changes: the tool-call policy ships as guidance only, pending the
10
+ A/B required before any behavior change.
11
+
12
+ - **Canonical prompt-caching ruleset (`docs/PROMPT_CACHING.md`, repo-local).** Distills the UNIC
13
+ caching roadmap into CTX-01..10 (MUST), the SHOULD set, and the never-do list, with
14
+ evidence-labeled vendor sections (Anthropic, OpenAI, DeepSeek, GLM, MiniMax) and an
15
+ adopt/adapt/reject matrix. Upstream vendor docs are reference only — never UNIC guarantees
16
+ (UNIC behavior is labeled `unknown`/`inferred`). Repo-local by design; not shipped.
17
+ - **Shipped caching guidance (`templates/docs/PROMPT_CACHING.md`).** The distilled CTX rules,
18
+ never-do list, and short tool-call policy now ship with UKit installs through a new manifest
19
+ item `docs-prompt-caching` (`mergeStrategy: overwrite_with_backup`, so `ukit update` refreshes
20
+ it). `templates/CLAUDE.md` and `templates/AGENTS.md` gain a matching `## Prompt Caching`
21
+ pointer section, and the repo-local `CLAUDE.md`/`AGENTS.md` carry a minimal pointer too.
22
+ - **Deterministic prompt assembly.** `templates/.claude/hooks/skill-router.sh` ordered
23
+ model-visible memory-recall output by volatile `updatedAt` timestamps, so the same state could
24
+ assemble different prompt bytes across runs and defeat prompt caching (CTX-01/CTX-04). Ordering
25
+ is now a stable sort with a deterministic tiebreak and no volatile timestamps in model-visible
26
+ output; the live `.claude/hooks/` twin is byte-identical. Locked by the new
27
+ `tests/hooks/promptAssemblyDeterminism.test.js` plus an ordering regression in
28
+ `tests/hooks/skillRouterHook.test.js`.
29
+
5
30
  ## 2.3.20 - 2026-09-13
6
31
 
7
32
  C15 bug-fix release record — this release completes the silent-stop class on top of the
@@ -218,6 +218,20 @@ items:
218
218
  packs:
219
219
  - core
220
220
 
221
+ # Shipped prompt-caching guidance (CTX-01..10 + never-do + tool-call policy + vendor cheat
222
+ # sheet). `overwrite_with_backup` so `ukit update` refreshes the ruleset — deliberately NOT
223
+ # `docs/PROJECT.md`, which is `mergeStrategy: skip` and must never be auto-overwritten.
224
+ - id: docs-prompt-caching
225
+ type: config
226
+ sourceTemplate: docs/PROMPT_CACHING.md
227
+ targetPath: docs/PROMPT_CACHING.md
228
+ requires: []
229
+ mergeStrategy: overwrite_with_backup
230
+ variables: []
231
+ enabledByDefault: true
232
+ packs:
233
+ - core
234
+
221
235
  - id: core-skill-delivery
222
236
  type: skill
223
237
  sourceTemplate: .claude/skills/delivery/SKILL.md
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.20",
3
+ "version": "2.3.21",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -961,8 +961,12 @@ const { pathToFileURL } = require('url');
961
961
  return 0;
962
962
  }
963
963
 
964
- const recencyBonus = getMemoryTimestamp(item) > 0 ? Math.min(1, getMemoryTimestamp(item) / Date.now()) : 0;
965
- return score + recencyBonus;
964
+ // TASK-027 CTX-01/CTX-04: the ranking must be a pure function of the logical memory
965
+ // content. A recency bonus scaled by Date.now() made the score — and therefore the
966
+ // injected previous-context block, its persisted order, and the PreCompact reinjection —
967
+ // shift on every wall-clock tick. Eligibility is unchanged (still score > 0); only the
968
+ // volatile ordering signal is dropped.
969
+ return score;
966
970
  }
967
971
 
968
972
  function buildPreviousContextSnippet(item) {
@@ -1017,9 +1021,12 @@ const { pathToFileURL } = require('url');
1017
1021
  score: scoreMemoryItem(item, queryTokens),
1018
1022
  }))
1019
1023
  .filter((entry) => entry.score > 0)
1024
+ // CTX-01: ties break on the stable logical id (never on a volatile timestamp), so the
1025
+ // same memory set always yields the same injected order regardless of when the
1026
+ // entries were last written/ended.
1020
1027
  .sort((left, right) => (
1021
1028
  right.score - left.score
1022
- || getMemoryTimestamp(right.item) - getMemoryTimestamp(left.item)
1029
+ || left.item.id.localeCompare(right.item.id)
1023
1030
  ))
1024
1031
  .slice(0, 2)
1025
1032
  .map((entry) => entry.item);
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -0,0 +1,127 @@
1
+ # Prompt Caching — guidance for UKit projects
2
+
3
+ Prompt caching is a **prefix match**: a provider can reuse the computation of an input prefix
4
+ when a later request reproduces that prefix byte-for-byte. UKit cannot control the transport or
5
+ the provider cache engine, but it *does* control the instruction, skill, tool and hook-injected
6
+ content it renders. This file ships with UKit so every installed project gets the same rules for
7
+ keeping that content deterministic and stable. Read it on demand — it is deliberately not loaded
8
+ into every session.
9
+
10
+ ## Why stable context matters
11
+
12
+ - Cache reuse is **reported**, not controllable. A `cache_read` counter (or a vendor equivalent)
13
+ proves the provider *reported* a reuse; it never proves which layer served it.
14
+ - Three mechanisms must never be conflated: the **provider prompt cache** (reuses input
15
+ computation), a **local tool-result cache** (the host reuses a still-valid result), and a
16
+ **response cache** (the application replays a stored answer).
17
+ - Three counts must never be conflated: a **logical model turn**, a **client HTTP attempt**, and a
18
+ **tool execution**.
19
+ - The rules below hold even if no cache capability exists behind your provider, because they
20
+ govern content UKit itself renders.
21
+
22
+ ## CTX rules — MUST
23
+
24
+ | ID | Rule | What it means in practice |
25
+ |---|---|---|
26
+ | CTX-01 | Same logical input and config produce the same segment bytes | Render owned blocks deterministically; no ordering that varies run-to-run |
27
+ | CTX-02 | Preserve instruction/user/tool roles and conversation order | Never move a user ask into system or reorder history to lengthen a prefix |
28
+ | CTX-03 | Keep all tool IDs and required native continuation state intact | Never strip tool-call IDs or reasoning/continuation fields to shrink a request |
29
+ | CTX-04 | Do not inject clock/random IDs into static instructions | Static blocks carry no date, UUID or counter; volatile values go in the tail |
30
+ | CTX-05 | Every summary/compaction creates a new context epoch | Compaction is a decision with a cost; when it happens it starts a new epoch |
31
+ | CTX-06 | Never change data or code to match the cache | Cache optimization never edits content semantics — correctness over prefix |
32
+ | CTX-07 | Do not self-send a field the adapter has not confirmed as supported | Optional cache params only when capability is verified; HTTP 200 is not proof |
33
+ | CTX-08 | Tool-result cache reuse only when freshness/dependency is valid | Local reuse needs a freshness predicate and dependency fingerprint, not just a query match |
34
+ | CTX-09 | Never treat missing usage as zero | Missing cache counters are unknown plus lowered coverage, never a reported miss |
35
+ | CTX-10 | Never exceed existing instructions/permissions to cut calls | Reducing tool calls never means skipping a required check or test |
36
+
37
+ ## Never do
38
+
39
+ - Never sort messages or reasoning blocks alphabetically.
40
+ - Never trim code literals, signed content, or data where whitespace is meaningful.
41
+ - Never rewrite native reasoning/signature/encrypted continuation fields.
42
+ - Never move a user request into system to lengthen the stable prefix.
43
+ - Never promote untrusted documents or tool output to the developer/system role.
44
+ - Never hash one text and assume every occurrence shares the provider cache.
45
+ - Never add a timestamp just to log — keep log metadata outside model-visible content.
46
+
47
+ ## Runtime tool-call policy (condensed, guidance only)
48
+
49
+ This is guidance for how an assistant should decide when to call tools. It is **not** a rigid
50
+ "always at least N tools" or "max N tools" rule, and UKit does not change any default behavior
51
+ on its basis.
52
+
53
+ 1. Before calling a tool, decide which data is still missing to finish the request.
54
+ 2. Reuse evidence you already hold if it is still valid; check the version before reuse.
55
+ 3. Batch independent reads when the interface supports it.
56
+ 4. For dependent actions, wait for the needed result before deciding the next step.
57
+ 5. Prefer scoped queries with enough output to verify.
58
+ 6. If a result was truncated or insufficient, widen deliberately.
59
+ 7. On tool error, distinguish parameter error, transient error, and unknown state.
60
+ 8. After a change, run the check appropriate to the risk and the repo's requirements.
61
+ 9. When the completion condition is met, return the result — check further only for specific
62
+ remaining risk.
63
+
64
+ ## Vendor cheat sheet
65
+
66
+ The sections below describe how each upstream vendor documents its own caching. They are
67
+ **reference material only** — see "Upstream docs are not provider guarantees" below. Provider
68
+ minimums and rates change; treat the shapes here as orientation and **verify against the current
69
+ official docs** (links in each section) before relying on a number.
70
+
71
+ ### Anthropic
72
+
73
+ - **Mechanism:** explicit cache breakpoints (`cache_control: {"type": "ephemeral"}`) on cacheable
74
+ content blocks; a single top-level marker enables automatic caching on the last eligible block.
75
+ - **Minimum cacheable size:** model-dependent, roughly 512 to 4096 input tokens; shorter prompts
76
+ are silently not cached.
77
+ - **Discount shape:** cache reads are billed at a small fraction of the normal input rate; cache
78
+ writes carry a premium; an optional longer TTL costs more to write.
79
+ - **Docs:** https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ·
80
+ https://docs.anthropic.com/en/docs/about-claude/pricing
81
+
82
+ ### OpenAI
83
+
84
+ - **Mechanism:** implicit (automatic) caching; recent models also accept explicit cache markers,
85
+ and a prompt cache key can steer routing or cache accounting.
86
+ - **Minimum cacheable size:** roughly 1024 visible input tokens on recent models; varies with
87
+ request settings on older ones.
88
+ - **Discount shape:** cached reads are heavily discounted relative to input; cache writes are
89
+ billed at a modest premium on recent models and carry no extra write charge on older ones.
90
+ - **Docs:** https://developers.openai.com/api/docs/guides/prompt-caching
91
+
92
+ ### DeepSeek
93
+
94
+ - **Mechanism:** automatic, best-effort prefix caching; a cached prefix is an indivisible unit, so
95
+ partial overlap does not hit.
96
+ - **Minimum cacheable size:** not documented.
97
+ - **Discount shape:** a separate lower per-model cache-hit rate versus the cache-miss rate.
98
+ - **Note:** with tools present, reasoning content must be passed back on every later request.
99
+ - **Docs:** https://api-docs.deepseek.com/guides/kv_cache
100
+
101
+ ### GLM
102
+
103
+ - **Mechanism:** implicit caching triggered by content similarity; no explicit create or
104
+ invalidate API is documented.
105
+ - **Minimum cacheable size:** not documented.
106
+ - **Discount shape:** a separate per-model cached-input rate, not a universal ratio — do not
107
+ assume a fixed percentage.
108
+ - **Docs:** https://docs.z.ai/guides/capabilities/cache
109
+
110
+ ### MiniMax
111
+
112
+ - **Mechanism:** passive automatic prefix caching, plus explicit Anthropic-compatible
113
+ `cache_control` breakpoints on cacheable blocks.
114
+ - **Minimum cacheable size:** caching applies from roughly 512 tokens upward.
115
+ - **Discount shape:** explicit cache writes are billed at a premium and reads at a fraction of
116
+ input; passive cache writes carry no additional charge.
117
+ - **Docs:** https://platform.minimax.io/docs/api-reference/text-prompt-caching.md
118
+
119
+ ## Upstream docs are not provider guarantees
120
+
121
+ A vendor documenting a cache feature does **not** mean the gateway or route you actually use
122
+ forwards it, returns the same usage fields, or bills it the same way. Treat every statement about
123
+ cache behavior behind a gateway as a hypothesis, not a guarantee: verify cache parameters and
124
+ usage counters against the raw response of the exact endpoint you call, and against the current
125
+ official docs linked above. If a usage counter is absent, it is **unknown** — never a reported
126
+ zero (CTX-09). When project information changes mid-session, prefer correctness over a preserved
127
+ prefix: send the new version and accept whatever cache invalidation follows.