@ngockhoale/ukit 2.3.20 → 2.3.21
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md
CHANGED
|
@@ -2,6 +2,31 @@
|
|
|
2
2
|
|
|
3
3
|
All notable changes to UKit are documented here.
|
|
4
4
|
|
|
5
|
+
## 2.3.21 - 2026-09-13
|
|
6
|
+
|
|
7
|
+
C16 release — the prompt-caching ruleset turns from research into shipped guidance, and the
|
|
8
|
+
prompt-assembly surfaces that inject into model context are made deterministic. No new commands
|
|
9
|
+
and no default runtime behavior changes: the tool-call policy ships as guidance only, pending the
|
|
10
|
+
A/B required before any behavior change.
|
|
11
|
+
|
|
12
|
+
- **Canonical prompt-caching ruleset (`docs/PROMPT_CACHING.md`, repo-local).** Distills the UNIC
|
|
13
|
+
caching roadmap into CTX-01..10 (MUST), the SHOULD set, and the never-do list, with
|
|
14
|
+
evidence-labeled vendor sections (Anthropic, OpenAI, DeepSeek, GLM, MiniMax) and an
|
|
15
|
+
adopt/adapt/reject matrix. Upstream vendor docs are reference only — never UNIC guarantees
|
|
16
|
+
(UNIC behavior is labeled `unknown`/`inferred`). Repo-local by design; not shipped.
|
|
17
|
+
- **Shipped caching guidance (`templates/docs/PROMPT_CACHING.md`).** The distilled CTX rules,
|
|
18
|
+
never-do list, and short tool-call policy now ship with UKit installs through a new manifest
|
|
19
|
+
item `docs-prompt-caching` (`mergeStrategy: overwrite_with_backup`, so `ukit update` refreshes
|
|
20
|
+
it). `templates/CLAUDE.md` and `templates/AGENTS.md` gain a matching `## Prompt Caching`
|
|
21
|
+
pointer section, and the repo-local `CLAUDE.md`/`AGENTS.md` carry a minimal pointer too.
|
|
22
|
+
- **Deterministic prompt assembly.** `templates/.claude/hooks/skill-router.sh` ordered
|
|
23
|
+
model-visible memory-recall output by volatile `updatedAt` timestamps, so the same state could
|
|
24
|
+
assemble different prompt bytes across runs and defeat prompt caching (CTX-01/CTX-04). Ordering
|
|
25
|
+
is now a stable sort with a deterministic tiebreak and no volatile timestamps in model-visible
|
|
26
|
+
output; the live `.claude/hooks/` twin is byte-identical. Locked by the new
|
|
27
|
+
`tests/hooks/promptAssemblyDeterminism.test.js` plus an ordering regression in
|
|
28
|
+
`tests/hooks/skillRouterHook.test.js`.
|
|
29
|
+
|
|
5
30
|
## 2.3.20 - 2026-09-13
|
|
6
31
|
|
|
7
32
|
C15 bug-fix release record — this release completes the silent-stop class on top of the
|
|
@@ -218,6 +218,20 @@ items:
|
|
|
218
218
|
packs:
|
|
219
219
|
- core
|
|
220
220
|
|
|
221
|
+
# Shipped prompt-caching guidance (CTX-01..10 + never-do + tool-call policy + vendor cheat
|
|
222
|
+
# sheet). `overwrite_with_backup` so `ukit update` refreshes the ruleset — deliberately NOT
|
|
223
|
+
# `docs/PROJECT.md`, which is `mergeStrategy: skip` and must never be auto-overwritten.
|
|
224
|
+
- id: docs-prompt-caching
|
|
225
|
+
type: config
|
|
226
|
+
sourceTemplate: docs/PROMPT_CACHING.md
|
|
227
|
+
targetPath: docs/PROMPT_CACHING.md
|
|
228
|
+
requires: []
|
|
229
|
+
mergeStrategy: overwrite_with_backup
|
|
230
|
+
variables: []
|
|
231
|
+
enabledByDefault: true
|
|
232
|
+
packs:
|
|
233
|
+
- core
|
|
234
|
+
|
|
221
235
|
- id: core-skill-delivery
|
|
222
236
|
type: skill
|
|
223
237
|
sourceTemplate: .claude/skills/delivery/SKILL.md
|
package/package.json
CHANGED
|
@@ -961,8 +961,12 @@ const { pathToFileURL } = require('url');
|
|
|
961
961
|
return 0;
|
|
962
962
|
}
|
|
963
963
|
|
|
964
|
-
|
|
965
|
-
|
|
964
|
+
// TASK-027 CTX-01/CTX-04: the ranking must be a pure function of the logical memory
|
|
965
|
+
// content. A recency bonus scaled by Date.now() made the score — and therefore the
|
|
966
|
+
// injected previous-context block, its persisted order, and the PreCompact reinjection —
|
|
967
|
+
// shift on every wall-clock tick. Eligibility is unchanged (still score > 0); only the
|
|
968
|
+
// volatile ordering signal is dropped.
|
|
969
|
+
return score;
|
|
966
970
|
}
|
|
967
971
|
|
|
968
972
|
function buildPreviousContextSnippet(item) {
|
|
@@ -1017,9 +1021,12 @@ const { pathToFileURL } = require('url');
|
|
|
1017
1021
|
score: scoreMemoryItem(item, queryTokens),
|
|
1018
1022
|
}))
|
|
1019
1023
|
.filter((entry) => entry.score > 0)
|
|
1024
|
+
// CTX-01: ties break on the stable logical id (never on a volatile timestamp), so the
|
|
1025
|
+
// same memory set always yields the same injected order regardless of when the
|
|
1026
|
+
// entries were last written/ended.
|
|
1020
1027
|
.sort((left, right) => (
|
|
1021
1028
|
right.score - left.score
|
|
1022
|
-
||
|
|
1029
|
+
|| left.item.id.localeCompare(right.item.id)
|
|
1023
1030
|
))
|
|
1024
1031
|
.slice(0, 2)
|
|
1025
1032
|
.map((entry) => entry.item);
|
package/templates/AGENTS.md
CHANGED
|
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
|
|
|
106
106
|
- Threshold-based compact pressure is internal orchestration; do not expose it to users.
|
|
107
107
|
- For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
|
|
108
108
|
|
|
109
|
+
## Prompt Caching
|
|
110
|
+
|
|
111
|
+
- Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
|
|
112
|
+
- Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
|
|
113
|
+
- CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
|
|
114
|
+
- Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
|
|
115
|
+
- Upstream vendor docs are reference only — never a guarantee about the route you actually use.
|
|
116
|
+
|
|
109
117
|
## Safe Patch Protocol
|
|
110
118
|
|
|
111
119
|
- Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
|
package/templates/CLAUDE.md
CHANGED
|
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
|
|
|
106
106
|
- Threshold-based compact pressure is internal orchestration; do not expose it to users.
|
|
107
107
|
- For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
|
|
108
108
|
|
|
109
|
+
## Prompt Caching
|
|
110
|
+
|
|
111
|
+
- Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
|
|
112
|
+
- Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
|
|
113
|
+
- CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
|
|
114
|
+
- Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
|
|
115
|
+
- Upstream vendor docs are reference only — never a guarantee about the route you actually use.
|
|
116
|
+
|
|
109
117
|
## Safe Patch Protocol
|
|
110
118
|
|
|
111
119
|
- Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
|
|
@@ -0,0 +1,127 @@
|
|
|
1
|
+
# Prompt Caching — guidance for UKit projects
|
|
2
|
+
|
|
3
|
+
Prompt caching is a **prefix match**: a provider can reuse the computation of an input prefix
|
|
4
|
+
when a later request reproduces that prefix byte-for-byte. UKit cannot control the transport or
|
|
5
|
+
the provider cache engine, but it *does* control the instruction, skill, tool and hook-injected
|
|
6
|
+
content it renders. This file ships with UKit so every installed project gets the same rules for
|
|
7
|
+
keeping that content deterministic and stable. Read it on demand — it is deliberately not loaded
|
|
8
|
+
into every session.
|
|
9
|
+
|
|
10
|
+
## Why stable context matters
|
|
11
|
+
|
|
12
|
+
- Cache reuse is **reported**, not controllable. A `cache_read` counter (or a vendor equivalent)
|
|
13
|
+
proves the provider *reported* a reuse; it never proves which layer served it.
|
|
14
|
+
- Three mechanisms must never be conflated: the **provider prompt cache** (reuses input
|
|
15
|
+
computation), a **local tool-result cache** (the host reuses a still-valid result), and a
|
|
16
|
+
**response cache** (the application replays a stored answer).
|
|
17
|
+
- Three counts must never be conflated: a **logical model turn**, a **client HTTP attempt**, and a
|
|
18
|
+
**tool execution**.
|
|
19
|
+
- The rules below hold even if no cache capability exists behind your provider, because they
|
|
20
|
+
govern content UKit itself renders.
|
|
21
|
+
|
|
22
|
+
## CTX rules — MUST
|
|
23
|
+
|
|
24
|
+
| ID | Rule | What it means in practice |
|
|
25
|
+
|---|---|---|
|
|
26
|
+
| CTX-01 | Same logical input and config produce the same segment bytes | Render owned blocks deterministically; no ordering that varies run-to-run |
|
|
27
|
+
| CTX-02 | Preserve instruction/user/tool roles and conversation order | Never move a user ask into system or reorder history to lengthen a prefix |
|
|
28
|
+
| CTX-03 | Keep all tool IDs and required native continuation state intact | Never strip tool-call IDs or reasoning/continuation fields to shrink a request |
|
|
29
|
+
| CTX-04 | Do not inject clock/random IDs into static instructions | Static blocks carry no date, UUID or counter; volatile values go in the tail |
|
|
30
|
+
| CTX-05 | Every summary/compaction creates a new context epoch | Compaction is a decision with a cost; when it happens it starts a new epoch |
|
|
31
|
+
| CTX-06 | Never change data or code to match the cache | Cache optimization never edits content semantics — correctness over prefix |
|
|
32
|
+
| CTX-07 | Do not self-send a field the adapter has not confirmed as supported | Optional cache params only when capability is verified; HTTP 200 is not proof |
|
|
33
|
+
| CTX-08 | Tool-result cache reuse only when freshness/dependency is valid | Local reuse needs a freshness predicate and dependency fingerprint, not just a query match |
|
|
34
|
+
| CTX-09 | Never treat missing usage as zero | Missing cache counters are unknown plus lowered coverage, never a reported miss |
|
|
35
|
+
| CTX-10 | Never exceed existing instructions/permissions to cut calls | Reducing tool calls never means skipping a required check or test |
|
|
36
|
+
|
|
37
|
+
## Never do
|
|
38
|
+
|
|
39
|
+
- Never sort messages or reasoning blocks alphabetically.
|
|
40
|
+
- Never trim code literals, signed content, or data where whitespace is meaningful.
|
|
41
|
+
- Never rewrite native reasoning/signature/encrypted continuation fields.
|
|
42
|
+
- Never move a user request into system to lengthen the stable prefix.
|
|
43
|
+
- Never promote untrusted documents or tool output to the developer/system role.
|
|
44
|
+
- Never hash one text and assume every occurrence shares the provider cache.
|
|
45
|
+
- Never add a timestamp just to log — keep log metadata outside model-visible content.
|
|
46
|
+
|
|
47
|
+
## Runtime tool-call policy (condensed, guidance only)
|
|
48
|
+
|
|
49
|
+
This is guidance for how an assistant should decide when to call tools. It is **not** a rigid
|
|
50
|
+
"always at least N tools" or "max N tools" rule, and UKit does not change any default behavior
|
|
51
|
+
on its basis.
|
|
52
|
+
|
|
53
|
+
1. Before calling a tool, decide which data is still missing to finish the request.
|
|
54
|
+
2. Reuse evidence you already hold if it is still valid; check the version before reuse.
|
|
55
|
+
3. Batch independent reads when the interface supports it.
|
|
56
|
+
4. For dependent actions, wait for the needed result before deciding the next step.
|
|
57
|
+
5. Prefer scoped queries with enough output to verify.
|
|
58
|
+
6. If a result was truncated or insufficient, widen deliberately.
|
|
59
|
+
7. On tool error, distinguish parameter error, transient error, and unknown state.
|
|
60
|
+
8. After a change, run the check appropriate to the risk and the repo's requirements.
|
|
61
|
+
9. When the completion condition is met, return the result — check further only for specific
|
|
62
|
+
remaining risk.
|
|
63
|
+
|
|
64
|
+
## Vendor cheat sheet
|
|
65
|
+
|
|
66
|
+
The sections below describe how each upstream vendor documents its own caching. They are
|
|
67
|
+
**reference material only** — see "Upstream docs are not provider guarantees" below. Provider
|
|
68
|
+
minimums and rates change; treat the shapes here as orientation and **verify against the current
|
|
69
|
+
official docs** (links in each section) before relying on a number.
|
|
70
|
+
|
|
71
|
+
### Anthropic
|
|
72
|
+
|
|
73
|
+
- **Mechanism:** explicit cache breakpoints (`cache_control: {"type": "ephemeral"}`) on cacheable
|
|
74
|
+
content blocks; a single top-level marker enables automatic caching on the last eligible block.
|
|
75
|
+
- **Minimum cacheable size:** model-dependent, roughly 512 to 4096 input tokens; shorter prompts
|
|
76
|
+
are silently not cached.
|
|
77
|
+
- **Discount shape:** cache reads are billed at a small fraction of the normal input rate; cache
|
|
78
|
+
writes carry a premium; an optional longer TTL costs more to write.
|
|
79
|
+
- **Docs:** https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ·
|
|
80
|
+
https://docs.anthropic.com/en/docs/about-claude/pricing
|
|
81
|
+
|
|
82
|
+
### OpenAI
|
|
83
|
+
|
|
84
|
+
- **Mechanism:** implicit (automatic) caching; recent models also accept explicit cache markers,
|
|
85
|
+
and a prompt cache key can steer routing or cache accounting.
|
|
86
|
+
- **Minimum cacheable size:** roughly 1024 visible input tokens on recent models; varies with
|
|
87
|
+
request settings on older ones.
|
|
88
|
+
- **Discount shape:** cached reads are heavily discounted relative to input; cache writes are
|
|
89
|
+
billed at a modest premium on recent models and carry no extra write charge on older ones.
|
|
90
|
+
- **Docs:** https://developers.openai.com/api/docs/guides/prompt-caching
|
|
91
|
+
|
|
92
|
+
### DeepSeek
|
|
93
|
+
|
|
94
|
+
- **Mechanism:** automatic, best-effort prefix caching; a cached prefix is an indivisible unit, so
|
|
95
|
+
partial overlap does not hit.
|
|
96
|
+
- **Minimum cacheable size:** not documented.
|
|
97
|
+
- **Discount shape:** a separate lower per-model cache-hit rate versus the cache-miss rate.
|
|
98
|
+
- **Note:** with tools present, reasoning content must be passed back on every later request.
|
|
99
|
+
- **Docs:** https://api-docs.deepseek.com/guides/kv_cache
|
|
100
|
+
|
|
101
|
+
### GLM
|
|
102
|
+
|
|
103
|
+
- **Mechanism:** implicit caching triggered by content similarity; no explicit create or
|
|
104
|
+
invalidate API is documented.
|
|
105
|
+
- **Minimum cacheable size:** not documented.
|
|
106
|
+
- **Discount shape:** a separate per-model cached-input rate, not a universal ratio — do not
|
|
107
|
+
assume a fixed percentage.
|
|
108
|
+
- **Docs:** https://docs.z.ai/guides/capabilities/cache
|
|
109
|
+
|
|
110
|
+
### MiniMax
|
|
111
|
+
|
|
112
|
+
- **Mechanism:** passive automatic prefix caching, plus explicit Anthropic-compatible
|
|
113
|
+
`cache_control` breakpoints on cacheable blocks.
|
|
114
|
+
- **Minimum cacheable size:** caching applies from roughly 512 tokens upward.
|
|
115
|
+
- **Discount shape:** explicit cache writes are billed at a premium and reads at a fraction of
|
|
116
|
+
input; passive cache writes carry no additional charge.
|
|
117
|
+
- **Docs:** https://platform.minimax.io/docs/api-reference/text-prompt-caching.md
|
|
118
|
+
|
|
119
|
+
## Upstream docs are not provider guarantees
|
|
120
|
+
|
|
121
|
+
A vendor documenting a cache feature does **not** mean the gateway or route you actually use
|
|
122
|
+
forwards it, returns the same usage fields, or bills it the same way. Treat every statement about
|
|
123
|
+
cache behavior behind a gateway as a hypothesis, not a guarantee: verify cache parameters and
|
|
124
|
+
usage counters against the raw response of the exact endpoint you call, and against the current
|
|
125
|
+
official docs linked above. If a usage counter is absent, it is **unknown** — never a reported
|
|
126
|
+
zero (CTX-09). When project information changes mid-session, prefer correctness over a preserved
|
|
127
|
+
prefix: send the new version and accept whatever cache invalidation follows.
|