dsh-local-models 0.2.0 → 0.3.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README-es.md +1 -1
- package/README-hi.md +1 -1
- package/README-pt.md +1 -1
- package/README-zh.md +1 -1
- package/README.md +5 -1
- package/lib/index.js +104 -15
- package/package.json +1 -1
- package/skills/dsh-local-models-ops/SKILL.md +31 -4
package/README-es.md
CHANGED
|
@@ -11,7 +11,7 @@ Construido sobre el `llama.cpp` upstream sin modificaciones (`llama-server`). Si
|
|
|
11
11
|
- **Estimación de VRAM en vivo** — pesos + los tipos de caché K/V seleccionados + estado recurrente + cómputo/grafo + overhead frente al total de GPU detectado (nvidia-smi / sysfs de amdgpu, sumados entre todas las GPUs, 16 GB asumidos cuando se desconoce), con filas de entra / margen de seguridad / ctx máximo que entra (ver [Problemas conocidos](./KNOWN_ISSUES.md) para la precisión en la familia Gemma)
|
|
12
12
|
- **Profiles** — guarda configuraciones de lanzamiento con nombre y recárgalas con un clic
|
|
13
13
|
- **Modo router** — sirve todos los perfiles guardados desde un único endpoint compatible con OpenAI (`--models-preset`); los modelos se cargan bajo demanda, uno residente a la vez por defecto. Arrancar el router registra (o vuelve a registrar) automáticamente sus modelos en dsh — sin pulsar Register a mano.
|
|
14
|
-
- **Register in dsh** — escribe el servidor listo como ruta de proveedor `llm-pi-ai` (con modalidad de visión + niveles de razonamiento, y salida máxima anunciada en
|
|
14
|
+
- **Register in dsh** — escribe el servidor listo como ruta de proveedor `llm-pi-ai` (con modalidad de visión + niveles de razonamiento, y salida máxima anunciada en 32K tokens (limitada a la mitad de la ventana para que la compactación conserve presupuesto de presión; sube maxTokens por request explícitamente para bloques largos de razonamiento xhigh))
|
|
15
15
|
- **Overlay de terminal** — tail en vivo del log de `llama-server` desde la pestaña
|
|
16
16
|
|
|
17
17
|
## Requirements
|
package/README-hi.md
CHANGED
|
@@ -11,7 +11,7 @@ stock upstream `llama.cpp` (`llama-server`) के विरुद्ध बन
|
|
|
11
11
|
- **Live VRAM estimate** — weights + चुने गए K/V cache types + recurrent state + compute/graph + overhead, detected GPU total के विरुद्ध (nvidia-smi / amdgpu sysfs, सभी GPUs पर summed, अज्ञात होने पर 16 GB माना गया), fits / safe-margin / max-ctx-that-fits पंक्तियों के साथ (Gemma-family सटीकता के लिए [Known issues](./KNOWN_ISSUES.md) देखें)
|
|
12
12
|
- **Profiles** — नामित launch configurations सहेजें, एक क्लिक में फिर लोड करें
|
|
13
13
|
- **Router mode** — सभी सहेजे गए profiles को एक OpenAI-compatible endpoint (`--models-preset`) से serve करें; models माँग पर लोड होते हैं, default रूप से एक बार में एक ही resident रहता है। router शुरू करने पर dsh में उसके models अपने आप (फिर से) register हो जाते हैं — manual Register दबाने की ज़रूरत नहीं।
|
|
14
|
-
- **Register in dsh** — तैयार server को `llm-pi-ai` provider route के रूप में लिखता है (vision modality + thinking levels शामिल, max output
|
|
14
|
+
- **Register in dsh** — तैयार server को `llm-pi-ai` provider route के रूप में लिखता है (vision modality + thinking levels शामिल, max output 32K tokens बताया गया (window के आधे तक सीमित ताकि compaction का pressure budget बचा रहे; लंबे xhigh thinking blocks के लिए per-request maxTokens स्पष्ट रूप से बढ़ाएं))
|
|
15
15
|
- **Terminal overlay** — tab से ही `llama-server` log का live tail
|
|
16
16
|
|
|
17
17
|
## Requirements
|
package/README-pt.md
CHANGED
|
@@ -11,7 +11,7 @@ Construído sobre o `llama.cpp` upstream sem modificações (`llama-server`). Se
|
|
|
11
11
|
- **Estimativa de VRAM ao vivo** — pesos + os tipos de cache K/V selecionados + estado recorrente + compute/grafo + overhead contra o total detectado de GPU (nvidia-smi / sysfs do amdgpu, somados entre as GPUs, 16 GB presumidos quando desconhecido), com linhas de cabe / margem de segurança / ctx máximo que cabe (veja [Known issues](./KNOWN_ISSUES.md) para a precisão na família Gemma)
|
|
12
12
|
- **Profiles** — salve configurações de inicialização nomeadas e recarregue com um clique
|
|
13
13
|
- **Modo Router** — serve todos os profiles salvos a partir de um único endpoint compatível com OpenAI (`--models-preset`); os modelos carregam sob demanda, um residente por vez por padrão. Iniciar o router registra (ou re-registra) automaticamente seus modelos no dsh — sem precisar clicar em Register manualmente.
|
|
14
|
-
- **Register in dsh** — grava o servidor pronto como uma rota de provedor `llm-pi-ai` (modalidade de visão + thinking levels incluídos, saída máxima anunciada em
|
|
14
|
+
- **Register in dsh** — grava o servidor pronto como uma rota de provedor `llm-pi-ai` (modalidade de visão + thinking levels incluídos, saída máxima anunciada em 32K tokens (limitada à metade da janela para que a compactação mantenha budget de pressão; suba o maxTokens por request explicitamente para blocos longos de thinking xhigh))
|
|
15
15
|
- **Overlay de terminal** — tail ao vivo do log do `llama-server` direto da aba
|
|
16
16
|
|
|
17
17
|
## Requirements
|
package/README-zh.md
CHANGED
|
@@ -11,7 +11,7 @@
|
|
|
11
11
|
- **实时显存估算** —— 权重 + 所选的 K/V cache 类型 + 循环状态 + 计算/图 + 额外开销,与检测到的 GPU 总显存对比(nvidia-smi / amdgpu sysfs,多卡求和,未知时按 16 GB 计),并给出 fits / safe-margin / max-ctx-that-fits 三行结果(Gemma 系列的准确性见 [Known issues](./KNOWN_ISSUES.md))
|
|
12
12
|
- **Profiles** —— 保存具名的启动配置,一键重新加载
|
|
13
13
|
- **路由模式** —— 用一个 OpenAI 兼容端点(`--models-preset`)服务所有已保存的 profile;模型按需加载,默认同一时刻只驻留一个。启动路由会自动在 dsh 中(重新)注册它的模型 —— 无需手动点 Register。
|
|
14
|
-
- **Register in dsh** —— 把已就绪的服务器写成一个 `llm-pi-ai` 提供方路由(包含视觉模态 + thinking levels,最大输出声明为
|
|
14
|
+
- **Register in dsh** —— 把已就绪的服务器写成一个 `llm-pi-ai` 提供方路由(包含视觉模态 + thinking levels,最大输出声明为 32K tokens(上限为窗口的一半,以便 compact 保留压力预算;较长的 xhigh thinking 块请显式提高单次请求的 maxTokens))
|
|
15
15
|
- **终端浮层** —— 在标签页里实时跟踪 `llama-server` 日志
|
|
16
16
|
|
|
17
17
|
## Requirements
|
package/README.md
CHANGED
|
@@ -1,5 +1,7 @@
|
|
|
1
1
|
# dsh-local-models
|
|
2
2
|
|
|
3
|
+
[](https://dsh-plugin.org/plugins/vmarcelo49/dsh-local-models)
|
|
4
|
+
|
|
3
5
|
A `dsh` addon that adds a **Local Models** tab to the dsh Web GUI: pick a `.gguf` file, tune context and speculative decoding, watch a live VRAM estimate, and load it through `llama-server` — then register the running server as an LLM provider in dsh with one click.
|
|
4
6
|
|
|
5
7
|
Built against stock upstream `llama.cpp` (`llama-server`). No fork, no patches, no build step: the client bundle is hand-written `React.createElement` (no JSX toolchain) and the node half is dependency-free.
|
|
@@ -11,7 +13,7 @@ Built against stock upstream `llama.cpp` (`llama-server`). No fork, no patches,
|
|
|
11
13
|
- **Live VRAM estimate** — weights + the selected K/V cache types + recurrent state + compute/graph + overhead against the detected GPU total (nvidia-smi / amdgpu sysfs, summed across GPUs, 16 GB assumed when unknown), with fits / safe-margin / max-ctx-that-fits rows (see [Known issues](./KNOWN_ISSUES.md) for Gemma-family accuracy)
|
|
12
14
|
- **Profiles** — save named launch configurations, reload in one click
|
|
13
15
|
- **Router mode** — serve all saved profiles from one OpenAI-compatible endpoint (`--models-preset`); models load on demand, one resident at a time by default. Starting the router automatically (re-)registers its models in dsh — no manual Register press.
|
|
14
|
-
- **Register in dsh** — writes the ready server as an `llm-pi-ai` provider route (vision modality + thinking levels included, max output advertised at
|
|
16
|
+
- **Register in dsh** — writes the ready server as an `llm-pi-ai` provider route (vision modality + thinking levels included, max output advertised at 32K tokens (capped at half the window so compaction keeps a pressure budget; raise per-request maxTokens explicitly for long xhigh thinking blocks))
|
|
15
17
|
- **Terminal overlay** — live tail of the `llama-server` log from the tab
|
|
16
18
|
|
|
17
19
|
## Requirements
|
|
@@ -136,3 +138,5 @@ See [KNOWN_ISSUES.md](./KNOWN_ISSUES.md) — most notably, the VRAM estimate is
|
|
|
136
138
|
## License
|
|
137
139
|
|
|
138
140
|
MIT — see [LICENSE](./LICENSE).
|
|
141
|
+
|
|
142
|
+
|
package/lib/index.js
CHANGED
|
@@ -84,15 +84,99 @@ const LOG_PATH = join(dataDir(), "llama-server.log");
|
|
|
84
84
|
* ~15-22 s per image. Set LOCAL_MODELS_MMPROJ_CPU=0 to keep vision on the GPU. */
|
|
85
85
|
const MMPROJ_CPU = process.env.LOCAL_MODELS_MMPROJ_CPU !== "0";
|
|
86
86
|
/**
|
|
87
|
-
* Max output advertised when a server is registered in dsh:
|
|
88
|
-
* the
|
|
89
|
-
* the
|
|
90
|
-
*
|
|
91
|
-
*
|
|
92
|
-
*
|
|
93
|
-
*
|
|
87
|
+
* Max output advertised when a server is registered in dsh: 32K tokens,
|
|
88
|
+
* matching the llm-pi-ai adapter default. This matters beyond truncation:
|
|
89
|
+
* the advertised maxTokens becomes the adapter's defaultMaxTokens, which
|
|
90
|
+
* dsh-compaction-basic uses as its reserved output `O` in `W - O - headroom`.
|
|
91
|
+
* Advertising the whole window (the old 131K) leaves no message budget at
|
|
92
|
+
* `W = O` — proactive compaction warns once and never fires.
|
|
93
|
+
*
|
|
94
|
+
* NOTE — the 32K cap alone does NOT buy a late threshold under the core
|
|
95
|
+
* defaults (headroomTokens 65536, thresholdRatio 0.8, retainRatio 0.16; see
|
|
96
|
+
* compactionBudgetFor() below). For a 131072 window with O = 32768 the
|
|
97
|
+
* effective threshold is min(0.8 * W, W - O - headroom) = 32768, i.e. the
|
|
98
|
+
* session starts compacting at ~32-38K pressure tokens (the meter prices
|
|
99
|
+
* tools + system on top of the surface, so the visible trigger lands a few K
|
|
100
|
+
* above the raw threshold) — and with xhigh + preserveThinking every turn
|
|
101
|
+
* adds 10-20K thinking tokens, so compaction fires again on the next step:
|
|
102
|
+
* the "compaction loop". The honest fix for mid-size windows is a per-model
|
|
103
|
+
* headroom override in the dsh profile (headroom ~6-16K moves the 131K
|
|
104
|
+
* threshold to ~82-92K), not a smaller O — shrinking O truncates the xhigh
|
|
105
|
+
* thinking blocks this cap exists to protect. Heavy xhigh thinking that
|
|
106
|
+
* needs more must raise per-request maxTokens explicitly (which honestly
|
|
107
|
+
* moves the compaction threshold earlier instead of silently disabling it).
|
|
108
|
+
*/
|
|
109
|
+
const MAX_OUTPUT_TOKENS = 32768;
|
|
110
|
+
/** Default output share of the window: never offer more than half as max
|
|
111
|
+
* output, so a message budget always survives (`W - O > 0`). */
|
|
112
|
+
function defaultMaxOutput(contextWindow) {
|
|
113
|
+
return Math.max(1, Math.min(MAX_OUTPUT_TOKENS, Math.floor(contextWindow / 2)));
|
|
114
|
+
}
|
|
115
|
+
/**
|
|
116
|
+
* dsh-compaction-basic defaults this plugin's advertisement is priced
|
|
117
|
+
* against. Duplicated here (not imported) because the core package is not a
|
|
118
|
+
* dependency of this plugin — keep in sync with upstream's resolveConfig().
|
|
94
119
|
*/
|
|
95
|
-
const
|
|
120
|
+
export const COMPACTION_DEFAULTS = {
|
|
121
|
+
thresholdRatio: 0.8,
|
|
122
|
+
headroomTokens: 65536,
|
|
123
|
+
retainRatio: 0.16,
|
|
124
|
+
};
|
|
125
|
+
/**
|
|
126
|
+
* Price one advertised window the way dsh-compaction-basic's
|
|
127
|
+
* resolveCompactSpec() does, so the tab/skill can tell the user WHEN
|
|
128
|
+
* proactive compaction will actually fire — before they hit the loop.
|
|
129
|
+
*
|
|
130
|
+
* messageBudget = W - O (history available to messages)
|
|
131
|
+
* pressure = messageBudget - headroom
|
|
132
|
+
* threshold = min(floor(W * ratio), pressure) (fire at/above this)
|
|
133
|
+
* retain = floor(messageBudget * retainRatio) (tail kept per compact)
|
|
134
|
+
*
|
|
135
|
+
* `viable` is false when the pressure budget is <= 0: proactive compaction
|
|
136
|
+
* is then disabled entirely (core warns once and only recovers on overflow).
|
|
137
|
+
* That is the 96K-window cliff: W = 98304, O = 32768 leaves exactly 0.
|
|
138
|
+
*
|
|
139
|
+
* `recommendedHeadroom` is the headroom override that lands the threshold at
|
|
140
|
+
* ~70% of the window (clamped to [4096, 65536] so a summary + one retry
|
|
141
|
+
* always fit): the value to put in a `modelPolicies` entry for this route.
|
|
142
|
+
* Pure; `maxTokens` defaults to what buildProviderProfile would advertise.
|
|
143
|
+
*/
|
|
144
|
+
export function compactionBudgetFor(contextWindow, opts = {}) {
|
|
145
|
+
const W = Number.isInteger(contextWindow) && contextWindow > 0 ? contextWindow : 8192;
|
|
146
|
+
const O = Number.isInteger(opts.maxTokens) && opts.maxTokens > 0
|
|
147
|
+
? opts.maxTokens
|
|
148
|
+
: defaultMaxOutput(W);
|
|
149
|
+
const ratio = typeof opts.thresholdRatio === "number" && opts.thresholdRatio > 0 && opts.thresholdRatio <= 1
|
|
150
|
+
? opts.thresholdRatio
|
|
151
|
+
: COMPACTION_DEFAULTS.thresholdRatio;
|
|
152
|
+
const headroom = Number.isInteger(opts.headroomTokens) && opts.headroomTokens >= 0
|
|
153
|
+
? opts.headroomTokens
|
|
154
|
+
: COMPACTION_DEFAULTS.headroomTokens;
|
|
155
|
+
const retainRatio = typeof opts.retainRatio === "number" && opts.retainRatio > 0 && opts.retainRatio < ratio
|
|
156
|
+
? opts.retainRatio
|
|
157
|
+
: COMPACTION_DEFAULTS.retainRatio;
|
|
158
|
+
const messageBudget = W - O;
|
|
159
|
+
const pressureBudget = messageBudget - headroom;
|
|
160
|
+
const thresholdTokens = Math.floor(Math.min(W * ratio, pressureBudget));
|
|
161
|
+
const retainTokens = Math.floor(messageBudget * retainRatio);
|
|
162
|
+
// Headroom that would put the threshold at ~70% of the window: the
|
|
163
|
+
// pressure budget must cover 0.7 * W, i.e. headroom <= message - 0.7W.
|
|
164
|
+
const recommendedHeadroom = Math.max(
|
|
165
|
+
4096,
|
|
166
|
+
Math.min(COMPACTION_DEFAULTS.headroomTokens, messageBudget - Math.floor(W * 0.7)),
|
|
167
|
+
);
|
|
168
|
+
return {
|
|
169
|
+
contextWindow: W,
|
|
170
|
+
maxTokens: O,
|
|
171
|
+
headroomTokens: headroom,
|
|
172
|
+
messageBudget,
|
|
173
|
+
pressureBudget,
|
|
174
|
+
thresholdTokens,
|
|
175
|
+
retainTokens,
|
|
176
|
+
viable: pressureBudget > 0 && retainTokens < thresholdTokens,
|
|
177
|
+
recommendedHeadroom,
|
|
178
|
+
};
|
|
179
|
+
}
|
|
96
180
|
/** Fixed-MTP ceiling. The tab offers 0-7 and the API clamps here: upstream
|
|
97
181
|
* accepts any `--spec-draft-n-max` and clamps the effective depth to the
|
|
98
182
|
* model's own nextn depth at load. Depth 3 is the measured 16 GiB sweet spot
|
|
@@ -1622,12 +1706,17 @@ export function buildProviderProfile(st, requestedRoute) {
|
|
|
1622
1706
|
id: modelId,
|
|
1623
1707
|
name: modelId,
|
|
1624
1708
|
contextWindow,
|
|
1625
|
-
//
|
|
1626
|
-
// is a heavy thinker and at xhigh its thinking block alone
|
|
1627
|
-
// past
|
|
1628
|
-
//
|
|
1629
|
-
//
|
|
1630
|
-
|
|
1709
|
+
// Capped at half the window (see MAX_OUTPUT_TOKENS): the qwen3.x
|
|
1710
|
+
// series is a heavy thinker and at xhigh its thinking block alone
|
|
1711
|
+
// blows past small ceilings, so 32K is the default — but never more
|
|
1712
|
+
// than half the window, or compaction's `W - O - headroom` budget
|
|
1713
|
+
// collapses and proactive compaction never fires. NOTE: under the
|
|
1714
|
+
// core defaults (headroom 65536) a 131072 window still thresholds
|
|
1715
|
+
// at ~32K — see compactionBudgetFor(); mid-size windows need a
|
|
1716
|
+
// per-model headroom override, not a smaller O. Need more room
|
|
1717
|
+
// for a monster thinking block? Raise per-request maxTokens
|
|
1718
|
+
// explicitly instead.
|
|
1719
|
+
maxTokens: defaultMaxOutput(contextWindow),
|
|
1631
1720
|
// Declaring image input is what makes a hand-declared vision model
|
|
1632
1721
|
// usable: without it dsh rejects image attachments ("model does not
|
|
1633
1722
|
// declare image input"). mmprojPath is set when a --mmproj file was
|
|
@@ -1940,7 +2029,7 @@ export function buildRouterProfile(profiles, modalitiesByModel) {
|
|
|
1940
2029
|
id,
|
|
1941
2030
|
name: id,
|
|
1942
2031
|
contextWindow: p.ctx ?? 8192,
|
|
1943
|
-
maxTokens:
|
|
2032
|
+
maxTokens: defaultMaxOutput(p.ctx ?? 8192),
|
|
1944
2033
|
input,
|
|
1945
2034
|
reasoningEfforts: { off: "none", low: "low", medium: "medium", xhigh: "xhigh" },
|
|
1946
2035
|
compat: { supportsReasoningEffort: true },
|
package/package.json
CHANGED
|
@@ -98,10 +98,37 @@ provider route on the dsh webserver. Single-model mode needs the manual
|
|
|
98
98
|
**Register in dsh** press; router mode has no Register button — starting
|
|
99
99
|
the router auto-registers the `local-router` route once the server is
|
|
100
100
|
ready (a refresh poll fires it; stopping first cancels a pending one).
|
|
101
|
-
The registered model advertises `maxTokens: min(
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
101
|
+
The registered model advertises `maxTokens: min(32768, floor(ctx/2))` — a
|
|
102
|
+
32K cap at half the window. The advertised value becomes the adapter's
|
|
103
|
+
`defaultMaxTokens`, which compaction uses as its reserved output `O` in
|
|
104
|
+
`W - O - headroom`: advertising the whole window (`O = W`) leaves no
|
|
105
|
+
message budget and proactive compaction never fires. Heavy-thinking models
|
|
106
|
+
at xhigh that need more than 32K must raise per-request maxTokens
|
|
107
|
+
explicitly (which honestly moves the compaction threshold earlier).
|
|
108
|
+
Compaction threshold under the core defaults (headroomTokens 65536,
|
|
109
|
+
thresholdRatio 0.8, retainRatio 0.16) is `min(0.8*W, W - O - headroom)` —
|
|
110
|
+
so a 131072 window with O = 32768 thresholds at **32768**, firing at
|
|
111
|
+
~32-38K pressure tokens and looping with xhigh + preserveThinking (every
|
|
112
|
+
turn re-adds 10-20K thinking tokens). A 98304 window is worse: pressure
|
|
113
|
+
exactly 0, proactive compaction disabled entirely. The fix is a per-model
|
|
114
|
+
headroom override in the dsh profile (needs a `dsh web` restart — loader
|
|
115
|
+
normalization happens at boot), NOT a smaller O (that truncates xhigh
|
|
116
|
+
thinking). Compute it with `compactionBudgetFor(ctx).recommendedHeadroom`
|
|
117
|
+
(~6.5K for 131K → ~92K threshold; ~4K floor for 96K → ~61K). Example for
|
|
118
|
+
the Swift 131K route (also add siblings as needed):
|
|
119
|
+
|
|
120
|
+
```yaml
|
|
121
|
+
- id: compaction-basic
|
|
122
|
+
name: '@deepseek-ai/dsh-compaction-basic'
|
|
123
|
+
config:
|
|
124
|
+
modelPolicies:
|
|
125
|
+
- provider: local-router
|
|
126
|
+
model: swift-1-5-qwen3-8-27b
|
|
127
|
+
headroomTokens: 6554
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
Shortcut with zero config edits: run the bigger window instead — the
|
|
131
|
+
200K/250K profiles threshold at ~106K/158K under defaults.
|
|
105
132
|
For opencode against the spawned server:
|
|
106
133
|
- the opencode config resolves per directory (project `opencode.json` beats
|
|
107
134
|
nothing; the global `~/.config/opencode/opencode.json(c)` is authoritative
|