dsh-local-models 0.3.1 → 0.3.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/lib/index.js +8 -5
- package/package.json +1 -1
- package/skills/dsh-local-models-ops/SKILL.md +15 -15
package/lib/index.js
CHANGED
|
@@ -100,11 +100,14 @@ const MMPROJ_CPU = process.env.LOCAL_MODELS_MMPROJ_CPU !== "0";
|
|
|
100
100
|
* above the raw threshold) — and with xhigh + preserveThinking every turn
|
|
101
101
|
* adds 10-20K thinking tokens, so compaction fires again on the next step:
|
|
102
102
|
* the "compaction loop". The honest fix for mid-size windows is a per-model
|
|
103
|
-
* headroom override
|
|
104
|
-
*
|
|
105
|
-
*
|
|
106
|
-
*
|
|
107
|
-
*
|
|
103
|
+
* headroom override on the agent-preset row that mounts the running engine
|
|
104
|
+
* (`preset-standard` → `plugins` → `compaction` group → `compaction-basic`;
|
|
105
|
+
* headroom ~6-16K moves the 131K threshold to ~82-92K) — a patch on the
|
|
106
|
+
* top-level `compaction-basic` row is dead config (that row is disabled),
|
|
107
|
+
* and a smaller O would truncate the xhigh thinking blocks this cap
|
|
108
|
+
* exists to protect. Heavy xhigh thinking that needs more must raise
|
|
109
|
+
* per-request maxTokens explicitly (which honestly moves the compaction
|
|
110
|
+
* threshold earlier instead of silently disabling it).
|
|
108
111
|
*/
|
|
109
112
|
const MAX_OUTPUT_TOKENS = 32768;
|
|
110
113
|
/** Default output share of the window: never offer more than half as max
|
package/package.json
CHANGED
|
@@ -111,21 +111,21 @@ so a 131072 window with O = 32768 thresholds at **32768**, firing at
|
|
|
111
111
|
~32-38K pressure tokens and looping with xhigh + preserveThinking (every
|
|
112
112
|
turn re-adds 10-20K thinking tokens). A 98304 window is worse: pressure
|
|
113
113
|
exactly 0, proactive compaction disabled entirely. The fix is a per-model
|
|
114
|
-
headroom override
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
114
|
+
headroom override — but placement matters: the engines that run live
|
|
115
|
+
inside the agent-preset scopes (`preset-standard` → `plugins` →
|
|
116
|
+
`compaction` group → `compaction-basic`), NOT the top-level
|
|
117
|
+
`compaction-basic` row (disabled there; a patch on it is dead config that
|
|
118
|
+
dump-config will happily show). So the override must ride a wholesale
|
|
119
|
+
`config` on the preset row (row config is replaced whole, not merged):
|
|
120
|
+
copy the installed preset file
|
|
121
|
+
(`dsh-web-app/presets/standard.patch.yml` → `insert[0].config`), add
|
|
122
|
+
`modelPolicies` to the nested engine, paste as `- id: preset-standard`.
|
|
123
|
+
Compute values with `compactionBudgetFor(ctx).recommendedHeadroom`
|
|
124
|
+
(~6.5K for 131K → ~92K threshold; ~4K floor for 96K → ~61K). Because it
|
|
125
|
+
is wholesale, the row freezes: re-generate after any dsh upgrade that
|
|
126
|
+
touches the preset file, or upstream changes to that preset are silently
|
|
127
|
+
dropped. Only patch presets actually used (`standard` here); `ptc` /
|
|
128
|
+
`cordis` / `minimal` need the same treatment if adopted.
|
|
129
129
|
|
|
130
130
|
Shortcut with zero config edits: run the bigger window instead — the
|
|
131
131
|
200K/250K profiles threshold at ~106K/158K under defaults.
|