dsh-local-models 0.0.0-stage → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,3 +1,138 @@
1
- # Temporary Holding Version
1
+ # dsh-local-models
2
2
 
3
- This version is a temporary placeholder for this package. An operational version to replace this has been submitted for review and is awaiting a staged release.
3
+ A `dsh` addon that adds a **Local Models** tab to the dsh Web GUI: pick a `.gguf` file, tune context and speculative decoding, watch a live VRAM estimate, and load it through `llama-server` — then register the running server as an LLM provider in dsh with one click.
4
+
5
+ Built against stock upstream `llama.cpp` (`llama-server`). No fork, no patches, no build step: the client bundle is hand-written `React.createElement` (no JSX toolchain) and the node half is dependency-free.
6
+
7
+ ## Features
8
+
9
+ - **Model picker** — in-app file browser (directories + `.gguf` only) with a header-only GGUF parse (architecture, quant, layers, context length, MoE detection) behind `POST /local-models/gguf-meta`
10
+ - **Launch options** — context slider (8K steps, capped at the model's trained context) + fine-tune input, KV cache quantization selectors (one for K, one for V — every type `llama-server` accepts, with bytes-per-element shown), fixed MTP draft depth (0–7, upstream clamps to the model's nextn depth), thinking level (`off`/`low`/`medium`/`xhigh`) + preserve-thinking toggle (`--reasoning-preserve` vs `--no-reasoning-preserve`, default off), optional vision `mmproj` (GPU or CPU offload), MoE expert placement (`--cpu-moe` / `--n-cpu-moe` / top-k override) with a fit-to-VRAM helper
11
+ - **Live VRAM estimate** — weights + the selected K/V cache types + recurrent state + compute/graph + overhead against the detected GPU total (nvidia-smi / amdgpu sysfs, summed across GPUs, 16 GB assumed when unknown), with fits / safe-margin / max-ctx-that-fits rows (see [Known issues](./KNOWN_ISSUES.md) for Gemma-family accuracy)
12
+ - **Profiles** — save named launch configurations, reload in one click
13
+ - **Router mode** — serve all saved profiles from one OpenAI-compatible endpoint (`--models-preset`); models load on demand, one resident at a time by default. Starting the router automatically (re-)registers its models in dsh — no manual Register press.
14
+ - **Register in dsh** — writes the ready server as an `llm-pi-ai` provider route (vision modality + thinking levels included, max output advertised at 131K tokens so long xhigh thinking blocks aren't truncated)
15
+ - **Terminal overlay** — live tail of the `llama-server` log from the tab
16
+
17
+ ## Requirements
18
+
19
+ - `dsh` with the `web` profile (the plugin composes into it)
20
+ - A `llama-server` binary (upstream `llama.cpp`, Vulkan/CUDA/CPU — whatever your machine uses)
21
+ - The VRAM budget is detected (`nvidia-smi` for NVIDIA, amdgpu sysfs for AMD, all visible GPUs summed) and can be pinned in the tab's Runtime card; the 16 GB fallback and the safety margin live at the top of `lib/client.js` (`TOTAL_VRAM_BYTES`, `SAFE_MARGIN_BYTES`)
22
+
23
+ ## Install
24
+
25
+ A plugin lives inside a dsh **profile**, which is a pnpm project under
26
+ `$DSH_HOME/profiles/<name>`; `dsh plugin` forwards its arguments to pnpm in
27
+ that directory.
28
+
29
+ ```bash
30
+ # from the npm registry
31
+ dsh plugin --profile web add dsh-local-models
32
+
33
+ # straight from git (plain ESM, no build step)
34
+ dsh plugin --profile web add github:Vmarcelo49/dsh-local-models
35
+
36
+ # from a local clone, for development (symlinked: edits apply on reload)
37
+ dsh plugin --profile web add link:/path/to/dsh-local-models
38
+ ```
39
+
40
+ `dsh plugin add` writes the dependency **and** appends the package to
41
+ `dsh.profile.bundles` in `$DSH_HOME/profiles/web/package.json` — that array is
42
+ what mounts it, so there is nothing to edit by hand. Restart the dsh web
43
+ process (bundle composition happens at boot), refresh the browser and open
44
+ Settings → **Local Models**.
45
+
46
+ Check the composition without booting, and remove it again with:
47
+
48
+ ```bash
49
+ dsh --profile web --dump-config | grep -A 2 dsh-local-models
50
+ dsh plugin --profile web remove dsh-local-models
51
+ ```
52
+
53
+ - **pnpm must be on `PATH`.** npm or bun can install the `dsh` CLI itself, but
54
+ plugin management inside a profile is pnpm's (`dsh plugin` shells out to it
55
+ and prints `pnpm was not found` otherwise).
56
+ - **No version gate, no exemption.** The package declares no `@deepseek-ai/*`
57
+ peer dependencies — it only uses injected services (`settings`,
58
+ `credentials`, `webServer`) and client slots — so `dsh plugin` never refuses
59
+ it for a dsh mismatch and no `dsh plugin allow-version` is needed.
60
+ - **No build step.** There is no `prepare` script, so the pnpm `allowBuilds`
61
+ gate that git-hosted plugins hit never triggers.
62
+ - **Manifest check.** [`dsh-plugin-dev check`](https://www.npmjs.com/package/dsh-plugin-guide)
63
+ (from `dsh-plugin-guide`) validates the bundle manifest: `cordis.patch.yml`,
64
+ the `dsh.bundle.patch` pointer, `engines` and the `files` whitelist.
65
+
66
+ > Node-half changes (routes, inject list) need a dsh restart; client-half changes only need a page refresh.
67
+
68
+ ## Usage
69
+
70
+ 1. **Choose GGUF…** — pick a model file (Home / Models shortcuts, Up navigation).
71
+ 2. Tune **context**, **KV cache K / V**, **Max MTP head** (fixed draft, 0-7; 3 is the tuned sweet spot — deeper collapses at large ctx), **thinking level** + **preserve thinking** checkbox, optional **mmproj** and **MoE** settings.
72
+ 3. **Load model**, watch the status card, inspect output via **Open terminal**.
73
+ 4. **Register in dsh** — the route (default `local-<alias>`) appears in the Models picker.
74
+ 5. Alternatively, save **profiles** and **Start router (from profiles)** for a multi-model endpoint.
75
+ 6. Tick **"Start the router automatically when dsh starts"** (Router card) to launch the router at boot and register its `local-router` route once healthy — models stay usable without opening the tab. Needs at least one saved profile; progress lands in `llama-server.log` (`[autostart]` lines, visible via Open terminal).
76
+ 7. **Idle eviction** (Router card, "Unload models after …", default 30 min idle) frees VRAM via upstream `--sleep-idle-seconds` on both single loads and the router; the sleeping server keeps answering `/health` and reloads automatically on the next request (one slow request). `0` disables it. Takes effect on the next start — the tab warns when the running server uses a different timer.
77
+
78
+ ## Configuration
79
+
80
+ | Variable | Default | Meaning |
81
+ |---|---|---|
82
+ | `LOCAL_MODELS_PORT` | `8080` | `llama-server` port |
83
+ | `LOCAL_MODELS_BIN` | — (auto-detect) | server binary or the dir holding it; the Runtime card's setting wins over it |
84
+ | `LOCAL_MODELS_SHORTCUTS` | — (none) | colon-separated file-browser shortcut dirs (`name=path` for custom labels); the Runtime card's folder list takes over once saved |
85
+ | `LOCAL_MODELS_MMPROJ_CPU` | `1` | vision projector weights in RAM (`0` = offload to GPU) |
86
+ | `LOCAL_MODELS_ROUTER_MAX` | `1` | max simultaneously resident router models |
87
+ | `LOCAL_MODELS_MAX_IMAGE_BYTES` | `10485760` | vision image guard |
88
+ | `LOCAL_MODELS_IMAGE_PIXEL_BUDGET` | `4194304` | vision pixel budget |
89
+ | `DSH_HOME` | `~/.dsh` | data dir (`local-models/profiles.json`, `local-models/settings.json`, `llama-server.log`) |
90
+
91
+ The tab's VRAM budget is detected, not hardcoded: NVIDIA through `nvidia-smi`,
92
+ AMD through sysfs (`mem_info_vram_total`, with the product name resolved from
93
+ `pci.ids` when present), all visible GPUs summed, and `CUDA_VISIBLE_DEVICES` /
94
+ `HIP_VISIBLE_DEVICES` honored. Hardware that cannot be read falls back to the
95
+ historic 16 GiB, and the Runtime card's **VRAM budget** field pins the number
96
+ by hand (`settings.json` → `vramGb`, 0 = auto).
97
+
98
+ Launch flags are fixed to the validated daily config: full offload, `-b 2048 -ub 512 -t 4 -np 1`, `--flash-attn on --kv-unified`, reasoning `--reasoning auto --reasoning-format deepseek --reasoning-effort <level>` plus `--reasoning-preserve` when the preserve toggle (profile `preserveThinking`) is on else `--no-reasoning-preserve`, MTP `--spec-type draft-mtp --spec-draft-n-max N --spec-draft-p-min 0` (ungated — upstream's own default; the confidence gate only pays on bandwidth-starved cards, on this 16 GB card it costs ~32% decode at n-max 3 while *raising* acceptance 63.5% → 91.1%, see [bench/mtp_tuning.md](./bench/mtp_tuning.md); the tab offers depths 0-7, upstream clamps the effective depth to the model's nextn depth, and the draft is unconditional at any ctx — the old “ignore the MTP ctx softcap” checkbox is gone, so a deep draft at large ctx can still OOM or collapse decode), multi-GPU placement `--split-mode` / `--tensor-split` when a profile sets them (default: llama.cpp's own layer split, no flags — the control only appears when more than one GPU is detected), and the KV cache pair from the tab's K/V selectors (`--cache-type-k` / `--cache-type-v`, profile fields `kvTypeK` / `kvTypeV`). Every type this `llama-server` accepts is offered (`f32 f16 bf16 q8_0 q5_1 q5_0 q4_1 iq4_nl q4_0`, labeled with its bytes/element); the default `q5_0` K / `q4_1` V is the measured 16 GB sweet spot, and legacy profiles without the fields launch with exactly that pair. Quantized V needs flash-attn (always on here) and the MTP draft KV stays pinned to `q4_0`. MLA models (DeepSeek-style latent KV) reject mixed K/V types in llama.cpp, so the tab warns and keeps Load disabled until both match, and the `/run` route refuses such a launch with a clear error. Router presets carry the same per-profile KV pair and `reasoning-preserve = 1/0` choice.
99
+
100
+ ## HTTP API (mounted under `/local-models`)
101
+
102
+ | Route | Meaning |
103
+ |---|---|
104
+ | `GET /local-models/browse?dir=` | dirs + `.gguf` files |
105
+ | `POST /local-models/gguf-meta` | `{path}` → parsed GGUF header (cached) |
106
+ | `GET /local-models/status` | state + fresh `/health` probe |
107
+ | `GET /local-models/logs?offset=&max=` | incremental tail of `llama-server.log` |
108
+ | `POST /local-models/run` | spawn the server |
109
+ | `POST /local-models/stop` | stop the child (or reap the port) |
110
+ | `POST /local-models/profiles` / `GET` | save (upsert) / list profiles |
111
+ | `POST /local-models/profiles/remove` | delete a profile |
112
+ | `GET /local-models/settings` / `POST` | read / update plugin settings (`autostartRouter`, `autoUnloadMins`, `binPath`, `shortcuts`, `vramGb`) |
113
+ | `POST /local-models/runtime/check` | `{binPath}` → resolve + `<bin> --version` (the Runtime card's Check) |
114
+ | `POST /local-models/router/start` | build presets from profiles + start router |
115
+ | `POST /local-models/router/unload` | unload one router model |
116
+ | `POST /local-models/router/unload-all` | unload all router models |
117
+ | `POST /local-models/register` | add the ready server as an `llm-pi-ai` route |
118
+
119
+ ## Project layout
120
+
121
+ ```
122
+ lib/index.js node half: process manager, GGUF parser, routes, presets
123
+ lib/client.js browser half: settings tab (single build-free bundle)
124
+ skills/ operator skill: spawn-parity checklist, profile audits
125
+ docs/ UI mockup
126
+ ```
127
+
128
+ Pure, exported helpers (`normalizeEffort`, `moeArgsFor`, `generateRouterPresets`, `buildProviderProfile`, profiles store) are covered by `npm test` (node's built-in runner, `test/`); `node lib/index.js /path/to/model.gguf` dumps a parsed header as a self-test.
129
+
130
+ Host-provided modules: `@deepseek-ai/dsh-client-runtime` and `@deepseek-ai/dsh-client-ui-settings` are injected by the dsh host at bundle time (see the `dsh.client.inject` list in `package.json`) and are deliberately **not** in `dependencies` — they don't exist on npm and must not be installed.
131
+
132
+ ## Known issues
133
+
134
+ See [KNOWN_ISSUES.md](./KNOWN_ISSUES.md) — most notably, the VRAM estimate is approximate for Gemma-family layouts.
135
+
136
+ ## License
137
+
138
+ MIT — see [LICENSE](./LICENSE).
@@ -0,0 +1,7 @@
1
+ # dsh-local-models bundle patch: mount the addon as one composition row.
2
+ # Bundle layers insert rows over the shared root, so the row goes under an
3
+ # `insert` key (a plain row would only patch an existing entry).
4
+
5
+ - insert:
6
+ - id: dsh-local-models
7
+ name: dsh-local-models