dsh-local-models 0.3.0 → 0.3.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/lib/index.js CHANGED
@@ -89,11 +89,22 @@ const MMPROJ_CPU = process.env.LOCAL_MODELS_MMPROJ_CPU !== "0";
89
89
  * the advertised maxTokens becomes the adapter's defaultMaxTokens, which
90
90
  * dsh-compaction-basic uses as its reserved output `O` in `W - O - headroom`.
91
91
  * Advertising the whole window (the old 131K) leaves no message budget at
92
- * `W = O` — proactive compaction warns once and never fires. Half-window
93
- * cap keeps at least half the context for history so headroom tuning can
94
- * always make compaction viable. Heavy xhigh thinking that needs more must
95
- * raise per-request maxTokens explicitly (which honestly moves the
96
- * compaction threshold earlier instead of silently disabling it).
92
+ * `W = O` — proactive compaction warns once and never fires.
93
+ *
94
+ * NOTE — the 32K cap alone does NOT buy a late threshold under the core
95
+ * defaults (headroomTokens 65536, thresholdRatio 0.8, retainRatio 0.16; see
96
+ * compactionBudgetFor() below). For a 131072 window with O = 32768 the
97
+ * effective threshold is min(0.8 * W, W - O - headroom) = 32768, i.e. the
98
+ * session starts compacting at ~32-38K pressure tokens (the meter prices
99
+ * tools + system on top of the surface, so the visible trigger lands a few K
100
+ * above the raw threshold) — and with xhigh + preserveThinking every turn
101
+ * adds 10-20K thinking tokens, so compaction fires again on the next step:
102
+ * the "compaction loop". The honest fix for mid-size windows is a per-model
103
+ * headroom override in the dsh profile (headroom ~6-16K moves the 131K
104
+ * threshold to ~82-92K), not a smaller O — shrinking O truncates the xhigh
105
+ * thinking blocks this cap exists to protect. Heavy xhigh thinking that
106
+ * needs more must raise per-request maxTokens explicitly (which honestly
107
+ * moves the compaction threshold earlier instead of silently disabling it).
97
108
  */
98
109
  const MAX_OUTPUT_TOKENS = 32768;
99
110
  /** Default output share of the window: never offer more than half as max
@@ -101,6 +112,71 @@ const MAX_OUTPUT_TOKENS = 32768;
101
112
  function defaultMaxOutput(contextWindow) {
102
113
  return Math.max(1, Math.min(MAX_OUTPUT_TOKENS, Math.floor(contextWindow / 2)));
103
114
  }
115
+ /**
116
+ * dsh-compaction-basic defaults this plugin's advertisement is priced
117
+ * against. Duplicated here (not imported) because the core package is not a
118
+ * dependency of this plugin — keep in sync with upstream's resolveConfig().
119
+ */
120
+ export const COMPACTION_DEFAULTS = {
121
+ thresholdRatio: 0.8,
122
+ headroomTokens: 65536,
123
+ retainRatio: 0.16,
124
+ };
125
+ /**
126
+ * Price one advertised window the way dsh-compaction-basic's
127
+ * resolveCompactSpec() does, so the tab/skill can tell the user WHEN
128
+ * proactive compaction will actually fire — before they hit the loop.
129
+ *
130
+ * messageBudget = W - O (history available to messages)
131
+ * pressure = messageBudget - headroom
132
+ * threshold = min(floor(W * ratio), pressure) (fire at/above this)
133
+ * retain = floor(messageBudget * retainRatio) (tail kept per compact)
134
+ *
135
+ * `viable` is false when the pressure budget is <= 0: proactive compaction
136
+ * is then disabled entirely (core warns once and only recovers on overflow).
137
+ * That is the 96K-window cliff: W = 98304, O = 32768 leaves exactly 0.
138
+ *
139
+ * `recommendedHeadroom` is the headroom override that lands the threshold at
140
+ * ~70% of the window (clamped to [4096, 65536] so a summary + one retry
141
+ * always fit): the value to put in a `modelPolicies` entry for this route.
142
+ * Pure; `maxTokens` defaults to what buildProviderProfile would advertise.
143
+ */
144
+ export function compactionBudgetFor(contextWindow, opts = {}) {
145
+ const W = Number.isInteger(contextWindow) && contextWindow > 0 ? contextWindow : 8192;
146
+ const O = Number.isInteger(opts.maxTokens) && opts.maxTokens > 0
147
+ ? opts.maxTokens
148
+ : defaultMaxOutput(W);
149
+ const ratio = typeof opts.thresholdRatio === "number" && opts.thresholdRatio > 0 && opts.thresholdRatio <= 1
150
+ ? opts.thresholdRatio
151
+ : COMPACTION_DEFAULTS.thresholdRatio;
152
+ const headroom = Number.isInteger(opts.headroomTokens) && opts.headroomTokens >= 0
153
+ ? opts.headroomTokens
154
+ : COMPACTION_DEFAULTS.headroomTokens;
155
+ const retainRatio = typeof opts.retainRatio === "number" && opts.retainRatio > 0 && opts.retainRatio < ratio
156
+ ? opts.retainRatio
157
+ : COMPACTION_DEFAULTS.retainRatio;
158
+ const messageBudget = W - O;
159
+ const pressureBudget = messageBudget - headroom;
160
+ const thresholdTokens = Math.floor(Math.min(W * ratio, pressureBudget));
161
+ const retainTokens = Math.floor(messageBudget * retainRatio);
162
+ // Headroom that would put the threshold at ~70% of the window: the
163
+ // pressure budget must cover 0.7 * W, i.e. headroom <= message - 0.7W.
164
+ const recommendedHeadroom = Math.max(
165
+ 4096,
166
+ Math.min(COMPACTION_DEFAULTS.headroomTokens, messageBudget - Math.floor(W * 0.7)),
167
+ );
168
+ return {
169
+ contextWindow: W,
170
+ maxTokens: O,
171
+ headroomTokens: headroom,
172
+ messageBudget,
173
+ pressureBudget,
174
+ thresholdTokens,
175
+ retainTokens,
176
+ viable: pressureBudget > 0 && retainTokens < thresholdTokens,
177
+ recommendedHeadroom,
178
+ };
179
+ }
104
180
  /** Fixed-MTP ceiling. The tab offers 0-7 and the API clamps here: upstream
105
181
  * accepts any `--spec-draft-n-max` and clamps the effective depth to the
106
182
  * model's own nextn depth at load. Depth 3 is the measured 16 GiB sweet spot
@@ -1634,7 +1710,10 @@ export function buildProviderProfile(st, requestedRoute) {
1634
1710
  // series is a heavy thinker and at xhigh its thinking block alone
1635
1711
  // blows past small ceilings, so 32K is the default — but never more
1636
1712
  // than half the window, or compaction's `W - O - headroom` budget
1637
- // collapses and proactive compaction never fires. Need more room
1713
+ // collapses and proactive compaction never fires. NOTE: under the
1714
+ // core defaults (headroom 65536) a 131072 window still thresholds
1715
+ // at ~32K — see compactionBudgetFor(); mid-size windows need a
1716
+ // per-model headroom override, not a smaller O. Need more room
1638
1717
  // for a monster thinking block? Raise per-request maxTokens
1639
1718
  // explicitly instead.
1640
1719
  maxTokens: defaultMaxOutput(contextWindow),
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-local-models",
3
- "version": "0.3.0",
3
+ "version": "0.3.1",
4
4
  "description": "dsh addon: a Local Models settings tab that starts and stops a llama-server child process living with the dsh host process.",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
@@ -105,6 +105,30 @@ The registered model advertises `maxTokens: min(32768, floor(ctx/2))` — a
105
105
  message budget and proactive compaction never fires. Heavy-thinking models
106
106
  at xhigh that need more than 32K must raise per-request maxTokens
107
107
  explicitly (which honestly moves the compaction threshold earlier).
108
+ Compaction threshold under the core defaults (headroomTokens 65536,
109
+ thresholdRatio 0.8, retainRatio 0.16) is `min(0.8*W, W - O - headroom)` —
110
+ so a 131072 window with O = 32768 thresholds at **32768**, firing at
111
+ ~32-38K pressure tokens and looping with xhigh + preserveThinking (every
112
+ turn re-adds 10-20K thinking tokens). A 98304 window is worse: pressure
113
+ exactly 0, proactive compaction disabled entirely. The fix is a per-model
114
+ headroom override in the dsh profile (needs a `dsh web` restart — loader
115
+ normalization happens at boot), NOT a smaller O (that truncates xhigh
116
+ thinking). Compute it with `compactionBudgetFor(ctx).recommendedHeadroom`
117
+ (~6.5K for 131K → ~92K threshold; ~4K floor for 96K → ~61K). Example for
118
+ the Swift 131K route (also add siblings as needed):
119
+
120
+ ```yaml
121
+ - id: compaction-basic
122
+ name: '@deepseek-ai/dsh-compaction-basic'
123
+ config:
124
+ modelPolicies:
125
+ - provider: local-router
126
+ model: swift-1-5-qwen3-8-27b
127
+ headroomTokens: 6554
128
+ ```
129
+
130
+ Shortcut with zero config edits: run the bigger window instead — the
131
+ 200K/250K profiles threshold at ~106K/158K under defaults.
108
132
  For opencode against the spawned server:
109
133
  - the opencode config resolves per directory (project `opencode.json` beats
110
134
  nothing; the global `~/.config/opencode/opencode.json(c)` is authoritative