dsh-local-models 0.3.1 → 0.3.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/lib/index.js CHANGED
@@ -100,11 +100,14 @@ const MMPROJ_CPU = process.env.LOCAL_MODELS_MMPROJ_CPU !== "0";
100
100
  * above the raw threshold) — and with xhigh + preserveThinking every turn
101
101
  * adds 10-20K thinking tokens, so compaction fires again on the next step:
102
102
  * the "compaction loop". The honest fix for mid-size windows is a per-model
103
- * headroom override in the dsh profile (headroom ~6-16K moves the 131K
104
- * threshold to ~82-92K), not a smaller O — shrinking O truncates the xhigh
105
- * thinking blocks this cap exists to protect. Heavy xhigh thinking that
106
- * needs more must raise per-request maxTokens explicitly (which honestly
107
- * moves the compaction threshold earlier instead of silently disabling it).
103
+ * headroom override on the agent-preset row that mounts the running engine
104
+ * (`preset-standard` → `plugins` → `compaction` group → `compaction-basic`;
105
+ * headroom ~6-16K moves the 131K threshold to ~82-92K) — a patch on the
106
+ * top-level `compaction-basic` row is dead config (that row is disabled),
107
+ * and a smaller O would truncate the xhigh thinking blocks this cap
108
+ * exists to protect. Heavy xhigh thinking that needs more must raise
109
+ * per-request maxTokens explicitly (which honestly moves the compaction
110
+ * threshold earlier instead of silently disabling it).
108
111
  */
109
112
  const MAX_OUTPUT_TOKENS = 32768;
110
113
  /** Default output share of the window: never offer more than half as max
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "dsh-local-models",
3
- "version": "0.3.1",
3
+ "version": "0.3.2",
4
4
  "description": "dsh addon: a Local Models settings tab that starts and stops a llama-server child process living with the dsh host process.",
5
5
  "type": "module",
6
6
  "main": "lib/index.js",
@@ -111,21 +111,21 @@ so a 131072 window with O = 32768 thresholds at **32768**, firing at
111
111
  ~32-38K pressure tokens and looping with xhigh + preserveThinking (every
112
112
  turn re-adds 10-20K thinking tokens). A 98304 window is worse: pressure
113
113
  exactly 0, proactive compaction disabled entirely. The fix is a per-model
114
- headroom override in the dsh profile (needs a `dsh web` restart — loader
115
- normalization happens at boot), NOT a smaller O (that truncates xhigh
116
- thinking). Compute it with `compactionBudgetFor(ctx).recommendedHeadroom`
117
- (~6.5K for 131K → ~92K threshold; ~4K floor for 96K → ~61K). Example for
118
- the Swift 131K route (also add siblings as needed):
119
-
120
- ```yaml
121
- - id: compaction-basic
122
- name: '@deepseek-ai/dsh-compaction-basic'
123
- config:
124
- modelPolicies:
125
- - provider: local-router
126
- model: swift-1-5-qwen3-8-27b
127
- headroomTokens: 6554
128
- ```
114
+ headroom override — but placement matters: the engines that run live
115
+ inside the agent-preset scopes (`preset-standard` → `plugins` →
116
+ `compaction` group → `compaction-basic`), NOT the top-level
117
+ `compaction-basic` row (disabled there; a patch on it is dead config that
118
+ dump-config will happily show). So the override must ride a wholesale
119
+ `config` on the preset row (row config is replaced whole, not merged):
120
+ copy the installed preset file
121
+ (`dsh-web-app/presets/standard.patch.yml` → `insert[0].config`), add
122
+ `modelPolicies` to the nested engine, paste as `- id: preset-standard`.
123
+ Compute values with `compactionBudgetFor(ctx).recommendedHeadroom`
124
+ (~6.5K for 131K → ~92K threshold; ~4K floor for 96K → ~61K). Because it
125
+ is wholesale, the row freezes: re-generate after any dsh upgrade that
126
+ touches the preset file, or upstream changes to that preset are silently
127
+ dropped. Only patch presets actually used (`standard` here); `ptc` /
128
+ `cordis` / `minimal` need the same treatment if adopted.
129
129
 
130
130
  Shortcut with zero config edits: run the bigger window instead — the
131
131
  200K/250K profiles threshold at ~106K/158K under defaults.