@ngockhoale/ukit 2.3.20 → 2.3.22

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -2,6 +2,56 @@
2
2
 
3
3
  All notable changes to UKit are documented here.
4
4
 
5
+ ## 2.3.22 - 2026-09-13
6
+
7
+ Consultation-question routing fix — pure questions no longer fabricate mutation debt that the
8
+ completion gate then enforces against an answer turn. Observed live the same day: two Vietnamese
9
+ consultation questions (no English signal word, no target file) scored every candidate mode ≤ 0,
10
+ the mode ladder's upward bias resolved the find-cause/shared-edit tie to `shared-edit`, and the
11
+ gate blocked the answer turn six times ("missing write-evidence / verification-evidence") until
12
+ the continuation cap released it.
13
+
14
+ - **Consultation questions carry no completion contract.** A question-shaped prompt (trailing `?`,
15
+ a leading English or ASCII-folded Vietnamese interrogative, or a `là gì / được không` tail) with
16
+ no implement/review/debug/failure/impact signal and no target file routes `informational` — the
17
+ same no-mutation-debt lane delivery-only requests already use. Applied to
18
+ `src/index/taskRouting.js`, the standalone router `route-task.mjs`, and the hook-embedded
19
+ routing copy in `skill-router.sh` (live twins byte-identical).
20
+ - **Leading interrogatives override action words.** The router's signal regexes are English-only,
21
+ so Vietnamese how-questions mentioning an action ("Làm sao mà đo được khi tôi cài hệ thống
22
+ này…?") kept a mutation contract. A leading interrogative (làm sao / thế nào / tại sao / phải
23
+ làm gì / …) now marks the prompt consultative even when it contains implement verbs
24
+ (sửa / thêm / cài / cài đặt / …) — asking HOW something is done is not ordering the action
25
+ performed. Question-phrased edit orders ("bạn sửa giúp tôi… không?") keep their write contract.
26
+ - **Verification.** New routing regressions lock the two incident prompts to `informational` with
27
+ empty contract evidence, and lock both guard directions (English signal-word questions and
28
+ question-phrased Vietnamese edit orders stay in their lanes); full suite green.
29
+
30
+ ## 2.3.21 - 2026-09-13
31
+
32
+ C16 release — the prompt-caching ruleset turns from research into shipped guidance, and the
33
+ prompt-assembly surfaces that inject into model context are made deterministic. No new commands
34
+ and no default runtime behavior changes: the tool-call policy ships as guidance only, pending the
35
+ A/B required before any behavior change.
36
+
37
+ - **Canonical prompt-caching ruleset (`docs/PROMPT_CACHING.md`, repo-local).** Distills the UNIC
38
+ caching roadmap into CTX-01..10 (MUST), the SHOULD set, and the never-do list, with
39
+ evidence-labeled vendor sections (Anthropic, OpenAI, DeepSeek, GLM, MiniMax) and an
40
+ adopt/adapt/reject matrix. Upstream vendor docs are reference only — never UNIC guarantees
41
+ (UNIC behavior is labeled `unknown`/`inferred`). Repo-local by design; not shipped.
42
+ - **Shipped caching guidance (`templates/docs/PROMPT_CACHING.md`).** The distilled CTX rules,
43
+ never-do list, and short tool-call policy now ship with UKit installs through a new manifest
44
+ item `docs-prompt-caching` (`mergeStrategy: overwrite_with_backup`, so `ukit update` refreshes
45
+ it). `templates/CLAUDE.md` and `templates/AGENTS.md` gain a matching `## Prompt Caching`
46
+ pointer section, and the repo-local `CLAUDE.md`/`AGENTS.md` carry a minimal pointer too.
47
+ - **Deterministic prompt assembly.** `templates/.claude/hooks/skill-router.sh` ordered
48
+ model-visible memory-recall output by volatile `updatedAt` timestamps, so the same state could
49
+ assemble different prompt bytes across runs and defeat prompt caching (CTX-01/CTX-04). Ordering
50
+ is now a stable sort with a deterministic tiebreak and no volatile timestamps in model-visible
51
+ output; the live `.claude/hooks/` twin is byte-identical. Locked by the new
52
+ `tests/hooks/promptAssemblyDeterminism.test.js` plus an ordering regression in
53
+ `tests/hooks/skillRouterHook.test.js`.
54
+
5
55
  ## 2.3.20 - 2026-09-13
6
56
 
7
57
  C15 bug-fix release record — this release completes the silent-stop class on top of the
@@ -218,6 +218,20 @@ items:
218
218
  packs:
219
219
  - core
220
220
 
221
+ # Shipped prompt-caching guidance (CTX-01..10 + never-do + tool-call policy + vendor cheat
222
+ # sheet). `overwrite_with_backup` so `ukit update` refreshes the ruleset — deliberately NOT
223
+ # `docs/PROJECT.md`, which is `mergeStrategy: skip` and must never be auto-overwritten.
224
+ - id: docs-prompt-caching
225
+ type: config
226
+ sourceTemplate: docs/PROMPT_CACHING.md
227
+ targetPath: docs/PROMPT_CACHING.md
228
+ requires: []
229
+ mergeStrategy: overwrite_with_backup
230
+ variables: []
231
+ enabledByDefault: true
232
+ packs:
233
+ - core
234
+
221
235
  - id: core-skill-delivery
222
236
  type: skill
223
237
  sourceTemplate: .claude/skills/delivery/SKILL.md
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@ngockhoale/ukit",
3
- "version": "2.3.20",
3
+ "version": "2.3.22",
4
4
  "description": "Install/update an index-first AI workspace for Claude Code, OpenAI Codex, OpenCode, and omp (Oh My Pi).",
5
5
  "license": "MIT",
6
6
  "type": "module",
@@ -452,6 +452,47 @@ function deriveExecutionMode({
452
452
  && !scores.smallFixSignal
453
453
  && !scores.sharedRisk
454
454
  && !targetFile;
455
+ // A pure consultation question ("Rồi việc kế tiếp của tôi phải làm gì…?", "how do I
456
+ // measure this on another machine?") asks for an answer, not a repository change. The
457
+ // signal regexes are English-only, so a Vietnamese (or signal-free English) question
458
+ // with no target file scores every candidate mode ≤ 0 and the ladder's upward bias
459
+ // resolves the 0-score find-cause/shared-edit tie to shared-edit — fabricating
460
+ // write+verification debt that a text answer can never satisfy, after which the
461
+ // completion gate blocks the answer turn until the continuation cap releases it.
462
+ // Consultation questions carry no completion contract instead. A question that also
463
+ // carries an implement/review/debug/failure signal keeps its lane; targeted questions
464
+ // (explicit file target) keep their contract so a Vietnamese edit order against a
465
+ // named file stays enforced.
466
+ const trimmedPromptText = String(promptText || '').trim();
467
+ const foldedPromptText = trimmedPromptText
468
+ .toLowerCase()
469
+ .normalize('NFD')
470
+ .replace(/[\u0300-\u036f]/g, '')
471
+ .replace(/\u0111/g, 'd');
472
+ const leadingInterrogative = /^(?:what|how|which|when|where|who)\b/.test(foldedPromptText)
473
+ || /^(?:lam sao|the nao|nhu the nao|nhu nao|vi sao|tai sao|khi nao|bao gio|bao nhieu|o dau|phai lam gi|lam gi|lam cach nao)\b/.test(foldedPromptText);
474
+ const questionShapedPrompt = /\?\s*$/.test(trimmedPromptText)
475
+ || leadingInterrogative
476
+ || /\b(?:la gi|duoc khong)\s*$/.test(foldedPromptText);
477
+ // Implement verbs (Vietnamese included — the English-only scores above cannot see
478
+ // them) mark an action order even when phrased as a question ("ban sua giup toi…?" is
479
+ // an edit order, not a consultation). A leading interrogative overrides them:
480
+ // "Lam sao ma do duoc khi toi cai…?" asks HOW something is done, it does not order
481
+ // the action performed.
482
+ const consultationImplementWords = /(?<![A-Za-z0-9_])(?:implement|apply|update|modify|add|create|ship|deliver|fix|refactor|remove|delete|rename|change|write|build|make|install|run|deploy|execute|sửa|thêm|tạo|xóa|đổi|thay thế|cập nhật|viết|chạy|cài)(?![A-Za-z0-9_])/.test(trimmedPromptText);
483
+ const consultationOnlySignal = questionShapedPrompt
484
+ && (leadingInterrogative || !consultationImplementWords)
485
+ && scores.editCertainty === 0
486
+ && !scores.implementSignal
487
+ && !scores.reviewSignal
488
+ && !scores.debugSignal
489
+ && !scores.failureSignal
490
+ && !scores.impactSignal
491
+ && !scores.buildSignal
492
+ && !scores.directTransformSignal
493
+ && !scores.smallFixSignal
494
+ && !scores.sharedRisk
495
+ && !targetFile;
455
496
 
456
497
  if (
457
498
  releaseVerificationContinuation
@@ -460,7 +501,7 @@ function deriveExecutionMode({
460
501
  return 'review-release';
461
502
  }
462
503
 
463
- if (deliveryOnlySignal) {
504
+ if (deliveryOnlySignal || consultationOnlySignal) {
464
505
  return 'informational';
465
506
  }
466
507
 
@@ -961,8 +961,12 @@ const { pathToFileURL } = require('url');
961
961
  return 0;
962
962
  }
963
963
 
964
- const recencyBonus = getMemoryTimestamp(item) > 0 ? Math.min(1, getMemoryTimestamp(item) / Date.now()) : 0;
965
- return score + recencyBonus;
964
+ // TASK-027 CTX-01/CTX-04: the ranking must be a pure function of the logical memory
965
+ // content. A recency bonus scaled by Date.now() made the score — and therefore the
966
+ // injected previous-context block, its persisted order, and the PreCompact reinjection —
967
+ // shift on every wall-clock tick. Eligibility is unchanged (still score > 0); only the
968
+ // volatile ordering signal is dropped.
969
+ return score;
966
970
  }
967
971
 
968
972
  function buildPreviousContextSnippet(item) {
@@ -1017,9 +1021,12 @@ const { pathToFileURL } = require('url');
1017
1021
  score: scoreMemoryItem(item, queryTokens),
1018
1022
  }))
1019
1023
  .filter((entry) => entry.score > 0)
1024
+ // CTX-01: ties break on the stable logical id (never on a volatile timestamp), so the
1025
+ // same memory set always yields the same injected order regardless of when the
1026
+ // entries were last written/ended.
1020
1027
  .sort((left, right) => (
1021
1028
  right.score - left.score
1022
- || getMemoryTimestamp(right.item) - getMemoryTimestamp(left.item)
1029
+ || left.item.id.localeCompare(right.item.id)
1023
1030
  ))
1024
1031
  .slice(0, 2)
1025
1032
  .map((entry) => entry.item);
@@ -1645,7 +1652,38 @@ const { pathToFileURL } = require('url');
1645
1652
  return 'review-release';
1646
1653
  }
1647
1654
 
1648
- if (deliveryOnlySignal) {
1655
+ const trimmedPromptText = String(promptText || '').trim();
1656
+ const foldedPromptText = trimmedPromptText
1657
+ .toLowerCase()
1658
+ .normalize('NFD')
1659
+ .replace(/[\u0300-\u036f]/g, '')
1660
+ .replace(/\u0111/g, 'd');
1661
+ const leadingInterrogative = /^(?:what|how|which|when|where|who)\b/.test(foldedPromptText)
1662
+ || /^(?:lam sao|the nao|nhu the nao|nhu nao|vi sao|tai sao|khi nao|bao gio|bao nhieu|o dau|phai lam gi|lam gi|lam cach nao)\b/.test(foldedPromptText);
1663
+ const questionShapedPrompt = /\?\s*$/.test(trimmedPromptText)
1664
+ || leadingInterrogative
1665
+ || /\b(?:la gi|duoc khong)\s*$/.test(foldedPromptText);
1666
+ // Implement verbs (Vietnamese included — the English-only scores above cannot see
1667
+ // them) mark an action order even when phrased as a question ("ban sua giup toi…?" is
1668
+ // an edit order, not a consultation). A leading interrogative overrides them:
1669
+ // "Lam sao ma do duoc khi toi cai…?" asks HOW something is done, it does not order
1670
+ // the action performed.
1671
+ const consultationImplementWords = /(?<![A-Za-z0-9_])(?:implement|apply|update|modify|add|create|ship|deliver|fix|refactor|remove|delete|rename|change|write|build|make|install|run|deploy|execute|sửa|thêm|tạo|xóa|đổi|thay thế|cập nhật|viết|chạy|cài)(?![A-Za-z0-9_])/.test(trimmedPromptText);
1672
+ const consultationOnlySignal = questionShapedPrompt
1673
+ && (leadingInterrogative || !consultationImplementWords)
1674
+ && scores.editCertainty === 0
1675
+ && !scores.implementSignal
1676
+ && !scores.reviewSignal
1677
+ && !scores.debugSignal
1678
+ && !scores.failureSignal
1679
+ && !scores.impactSignal
1680
+ && !scores.buildSignal
1681
+ && !scores.directTransformSignal
1682
+ && !scores.smallFixSignal
1683
+ && !scores.sharedRisk
1684
+ && !targetFile;
1685
+
1686
+ if (deliveryOnlySignal || consultationOnlySignal) {
1649
1687
  return 'informational';
1650
1688
  }
1651
1689
 
@@ -1832,7 +1832,38 @@ function deriveExecutionMode({
1832
1832
  return 'review-release';
1833
1833
  }
1834
1834
 
1835
- if (deliveryOnlySignal) {
1835
+ const trimmedPromptText = String(promptText || '').trim();
1836
+ const foldedPromptText = trimmedPromptText
1837
+ .toLowerCase()
1838
+ .normalize('NFD')
1839
+ .replace(/[\u0300-\u036f]/g, '')
1840
+ .replace(/\u0111/g, 'd');
1841
+ const leadingInterrogative = /^(?:what|how|which|when|where|who)\b/.test(foldedPromptText)
1842
+ || /^(?:lam sao|the nao|nhu the nao|nhu nao|vi sao|tai sao|khi nao|bao gio|bao nhieu|o dau|phai lam gi|lam gi|lam cach nao)\b/.test(foldedPromptText);
1843
+ const questionShapedPrompt = /\?\s*$/.test(trimmedPromptText)
1844
+ || leadingInterrogative
1845
+ || /\b(?:la gi|duoc khong)\s*$/.test(foldedPromptText);
1846
+ // Implement verbs (Vietnamese included — the English-only scores above cannot see
1847
+ // them) mark an action order even when phrased as a question ("ban sua giup toi…?" is
1848
+ // an edit order, not a consultation). A leading interrogative overrides them:
1849
+ // "Lam sao ma do duoc khi toi cai…?" asks HOW something is done, it does not order
1850
+ // the action performed.
1851
+ const consultationImplementWords = /(?<![A-Za-z0-9_])(?:implement|apply|update|modify|add|create|ship|deliver|fix|refactor|remove|delete|rename|change|write|build|make|install|run|deploy|execute|sửa|thêm|tạo|xóa|đổi|thay thế|cập nhật|viết|chạy|cài)(?![A-Za-z0-9_])/.test(trimmedPromptText);
1852
+ const consultationOnlySignal = questionShapedPrompt
1853
+ && (leadingInterrogative || !consultationImplementWords)
1854
+ && scores.editCertainty === 0
1855
+ && !scores.implementSignal
1856
+ && !scores.reviewSignal
1857
+ && !scores.debugSignal
1858
+ && !scores.failureSignal
1859
+ && !scores.impactSignal
1860
+ && !scores.buildSignal
1861
+ && !scores.directTransformSignal
1862
+ && !scores.smallFixSignal
1863
+ && !scores.sharedRisk
1864
+ && !targetFile;
1865
+
1866
+ if (deliveryOnlySignal || consultationOnlySignal) {
1836
1867
  return 'informational';
1837
1868
  }
1838
1869
 
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -106,6 +106,14 @@ For clearly non-code specialist lanes (docs-only, status, task queue), skip the
106
106
  - Threshold-based compact pressure is internal orchestration; do not expose it to users.
107
107
  - For Codex Desktop long sessions, UKit can use soft auto-compact handoffs. Default `compact.codexContext.compactTarget=150` means about 150 compact handoff lines (120-150 preferred, hard max 170), not 150 tokens.
108
108
 
109
+ ## Prompt Caching
110
+
111
+ - Deterministic, stable context lets a provider reuse a prompt prefix — and it is worth doing even when no caching is guaranteed.
112
+ - Full ruleset: `docs/PROMPT_CACHING.md` (read on demand; it is not loaded into every session).
113
+ - CTX-01 deterministic segment bytes · CTX-02 keep roles and order · CTX-03 keep tool IDs and continuation state · CTX-04 no clock/random IDs in static blocks · CTX-05 compaction starts a new epoch · CTX-06 never change data to match a cache · CTX-07 no unconfirmed cache fields · CTX-08 tool-result reuse needs valid freshness · CTX-09 missing usage is unknown, not zero · CTX-10 never cut a required check to reduce calls.
114
+ - Never sort messages, trim meaningful whitespace, rewrite reasoning fields, or move a user request into system context.
115
+ - Upstream vendor docs are reference only — never a guarantee about the route you actually use.
116
+
109
117
  ## Safe Patch Protocol
110
118
 
111
119
  - Safe Patch is internal orchestration: normal users still only need `ukit install` and natural language.
@@ -0,0 +1,127 @@
1
+ # Prompt Caching — guidance for UKit projects
2
+
3
+ Prompt caching is a **prefix match**: a provider can reuse the computation of an input prefix
4
+ when a later request reproduces that prefix byte-for-byte. UKit cannot control the transport or
5
+ the provider cache engine, but it *does* control the instruction, skill, tool and hook-injected
6
+ content it renders. This file ships with UKit so every installed project gets the same rules for
7
+ keeping that content deterministic and stable. Read it on demand — it is deliberately not loaded
8
+ into every session.
9
+
10
+ ## Why stable context matters
11
+
12
+ - Cache reuse is **reported**, not controllable. A `cache_read` counter (or a vendor equivalent)
13
+ proves the provider *reported* a reuse; it never proves which layer served it.
14
+ - Three mechanisms must never be conflated: the **provider prompt cache** (reuses input
15
+ computation), a **local tool-result cache** (the host reuses a still-valid result), and a
16
+ **response cache** (the application replays a stored answer).
17
+ - Three counts must never be conflated: a **logical model turn**, a **client HTTP attempt**, and a
18
+ **tool execution**.
19
+ - The rules below hold even if no cache capability exists behind your provider, because they
20
+ govern content UKit itself renders.
21
+
22
+ ## CTX rules — MUST
23
+
24
+ | ID | Rule | What it means in practice |
25
+ |---|---|---|
26
+ | CTX-01 | Same logical input and config produce the same segment bytes | Render owned blocks deterministically; no ordering that varies run-to-run |
27
+ | CTX-02 | Preserve instruction/user/tool roles and conversation order | Never move a user ask into system or reorder history to lengthen a prefix |
28
+ | CTX-03 | Keep all tool IDs and required native continuation state intact | Never strip tool-call IDs or reasoning/continuation fields to shrink a request |
29
+ | CTX-04 | Do not inject clock/random IDs into static instructions | Static blocks carry no date, UUID or counter; volatile values go in the tail |
30
+ | CTX-05 | Every summary/compaction creates a new context epoch | Compaction is a decision with a cost; when it happens it starts a new epoch |
31
+ | CTX-06 | Never change data or code to match the cache | Cache optimization never edits content semantics — correctness over prefix |
32
+ | CTX-07 | Do not self-send a field the adapter has not confirmed as supported | Optional cache params only when capability is verified; HTTP 200 is not proof |
33
+ | CTX-08 | Tool-result cache reuse only when freshness/dependency is valid | Local reuse needs a freshness predicate and dependency fingerprint, not just a query match |
34
+ | CTX-09 | Never treat missing usage as zero | Missing cache counters are unknown plus lowered coverage, never a reported miss |
35
+ | CTX-10 | Never exceed existing instructions/permissions to cut calls | Reducing tool calls never means skipping a required check or test |
36
+
37
+ ## Never do
38
+
39
+ - Never sort messages or reasoning blocks alphabetically.
40
+ - Never trim code literals, signed content, or data where whitespace is meaningful.
41
+ - Never rewrite native reasoning/signature/encrypted continuation fields.
42
+ - Never move a user request into system to lengthen the stable prefix.
43
+ - Never promote untrusted documents or tool output to the developer/system role.
44
+ - Never hash one text and assume every occurrence shares the provider cache.
45
+ - Never add a timestamp just to log — keep log metadata outside model-visible content.
46
+
47
+ ## Runtime tool-call policy (condensed, guidance only)
48
+
49
+ This is guidance for how an assistant should decide when to call tools. It is **not** a rigid
50
+ "always at least N tools" or "max N tools" rule, and UKit does not change any default behavior
51
+ on its basis.
52
+
53
+ 1. Before calling a tool, decide which data is still missing to finish the request.
54
+ 2. Reuse evidence you already hold if it is still valid; check the version before reuse.
55
+ 3. Batch independent reads when the interface supports it.
56
+ 4. For dependent actions, wait for the needed result before deciding the next step.
57
+ 5. Prefer scoped queries with enough output to verify.
58
+ 6. If a result was truncated or insufficient, widen deliberately.
59
+ 7. On tool error, distinguish parameter error, transient error, and unknown state.
60
+ 8. After a change, run the check appropriate to the risk and the repo's requirements.
61
+ 9. When the completion condition is met, return the result — check further only for specific
62
+ remaining risk.
63
+
64
+ ## Vendor cheat sheet
65
+
66
+ The sections below describe how each upstream vendor documents its own caching. They are
67
+ **reference material only** — see "Upstream docs are not provider guarantees" below. Provider
68
+ minimums and rates change; treat the shapes here as orientation and **verify against the current
69
+ official docs** (links in each section) before relying on a number.
70
+
71
+ ### Anthropic
72
+
73
+ - **Mechanism:** explicit cache breakpoints (`cache_control: {"type": "ephemeral"}`) on cacheable
74
+ content blocks; a single top-level marker enables automatic caching on the last eligible block.
75
+ - **Minimum cacheable size:** model-dependent, roughly 512 to 4096 input tokens; shorter prompts
76
+ are silently not cached.
77
+ - **Discount shape:** cache reads are billed at a small fraction of the normal input rate; cache
78
+ writes carry a premium; an optional longer TTL costs more to write.
79
+ - **Docs:** https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching ·
80
+ https://docs.anthropic.com/en/docs/about-claude/pricing
81
+
82
+ ### OpenAI
83
+
84
+ - **Mechanism:** implicit (automatic) caching; recent models also accept explicit cache markers,
85
+ and a prompt cache key can steer routing or cache accounting.
86
+ - **Minimum cacheable size:** roughly 1024 visible input tokens on recent models; varies with
87
+ request settings on older ones.
88
+ - **Discount shape:** cached reads are heavily discounted relative to input; cache writes are
89
+ billed at a modest premium on recent models and carry no extra write charge on older ones.
90
+ - **Docs:** https://developers.openai.com/api/docs/guides/prompt-caching
91
+
92
+ ### DeepSeek
93
+
94
+ - **Mechanism:** automatic, best-effort prefix caching; a cached prefix is an indivisible unit, so
95
+ partial overlap does not hit.
96
+ - **Minimum cacheable size:** not documented.
97
+ - **Discount shape:** a separate lower per-model cache-hit rate versus the cache-miss rate.
98
+ - **Note:** with tools present, reasoning content must be passed back on every later request.
99
+ - **Docs:** https://api-docs.deepseek.com/guides/kv_cache
100
+
101
+ ### GLM
102
+
103
+ - **Mechanism:** implicit caching triggered by content similarity; no explicit create or
104
+ invalidate API is documented.
105
+ - **Minimum cacheable size:** not documented.
106
+ - **Discount shape:** a separate per-model cached-input rate, not a universal ratio — do not
107
+ assume a fixed percentage.
108
+ - **Docs:** https://docs.z.ai/guides/capabilities/cache
109
+
110
+ ### MiniMax
111
+
112
+ - **Mechanism:** passive automatic prefix caching, plus explicit Anthropic-compatible
113
+ `cache_control` breakpoints on cacheable blocks.
114
+ - **Minimum cacheable size:** caching applies from roughly 512 tokens upward.
115
+ - **Discount shape:** explicit cache writes are billed at a premium and reads at a fraction of
116
+ input; passive cache writes carry no additional charge.
117
+ - **Docs:** https://platform.minimax.io/docs/api-reference/text-prompt-caching.md
118
+
119
+ ## Upstream docs are not provider guarantees
120
+
121
+ A vendor documenting a cache feature does **not** mean the gateway or route you actually use
122
+ forwards it, returns the same usage fields, or bills it the same way. Treat every statement about
123
+ cache behavior behind a gateway as a hypothesis, not a guarantee: verify cache parameters and
124
+ usage counters against the raw response of the exact endpoint you call, and against the current
125
+ official docs linked above. If a usage counter is absent, it is **unknown** — never a reported
126
+ zero (CTX-09). When project information changes mid-session, prefer correctness over a preserved
127
+ prefix: send the new version and accept whatever cache invalidation follows.