dsh-llama-cpp-sampling-params 0.1.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/LICENSE ADDED
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 dsh-llama-cpp-sampling-params contributors
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
package/README.md ADDED
@@ -0,0 +1,116 @@
1
+ # dsh-llama-cpp-sampling-params
2
+
3
+ A [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`) plugin that injects **per-model sampling parameters** into every chat-completions request sent to a local **llama.cpp / llama-server** gateway.
4
+
5
+ ## Why
6
+
7
+ dsh's `LlmCallConfig` only carries `temperature` / `maxTokens` / `stop` — it never sends `top_p`, `top_k`, `min_p`, `repeat_penalty`, `presence_penalty`, or `frequency_penalty`. llama.cpp accepts all of these **per request**. This plugin stamps the matching model's sampling set onto the wire body of chat-completions calls, so you can switch sampling behavior (e.g. a high-temperature creative role vs. a low-temperature coding role) **by switching the model alias** — no model reload, no extra VRAM.
8
+
9
+ ## How it works
10
+
11
+ The plugin wraps `globalThis.fetch` at the transport layer:
12
+
13
+ 1. Every `fetch()` call in the dsh process flows through the wrapper.
14
+ 2. For requests whose URL contains `/v1/chat/completions`, the wrapper parses the JSON body.
15
+ 3. It reads `body.model` (the alias), looks it up in the `sampling-params` `models` table.
16
+ 4. If found, it stamps the sampling set onto the parsed body and re-serializes `init.body`.
17
+ 5. The original fetch sends the modified body.
18
+
19
+ **Why a fetch wrapper:** dsh's agent-loop requests arrive **deep-frozen** (mutation throws), and the pi-ai adapter does not forward `onPayload` from `GenerateOptions`. Wrapping `fetch` operates below all of those abstractions — it edits the already-built JSON body before it leaves the process, so no dsh internal is touched.
20
+
21
+ ## How roles work
22
+
23
+ - llama.cpp can serve **one loaded GGUF under several aliases** (e.g. `model-i`, `model-p`, `model-t`), all pointing at the same underlying weights — so switching alias costs **no extra VRAM**.
24
+ - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so switching the model in the UI picks a sampling role.
25
+ - This plugin reads the `model` field of each request (the alias) and looks it up in the `sampling-params` `models` table. **Unconfigured aliases pass through untouched.**
26
+
27
+ ## Zero conflict
28
+
29
+ - The fields dsh already sends (`temperature` / `maxTokens`) are left to dsh unless a model explicitly sets them.
30
+ - Every other sampling field is one dsh never sends, so there is no competing source — the plugin simply overrides the llama.cpp server default.
31
+
32
+ ## Install
33
+
34
+ ```sh
35
+ # From npm (once published)
36
+ dsh plugin --profile web add dsh-llama-cpp-sampling-params
37
+
38
+ # From a local directory
39
+ dsh plugin --profile web add link:C:/path/to/dsh-llama-cpp-sampling-params
40
+ ```
41
+
42
+ Then restart `dsh web` (or refresh the GUI page).
43
+
44
+ ## Configure models
45
+
46
+ Set the per-model sampling table in `$DSH_HOME/settings.yaml`. The section applies **live** — edits take effect without a restart.
47
+
48
+ ```yaml
49
+ sampling-params:
50
+ models:
51
+ # Key = the exact model id (alias) sent on the wire.
52
+ model-t: # Think
53
+ temperature: 0.6
54
+ top_p: 0.95
55
+ top_k: 20
56
+ min_p: 0.05
57
+ repeat_penalty: 1.0
58
+ presence_penalty: 0.0
59
+ frequency_penalty: 0.0
60
+ model-i: # Instruct
61
+ temperature: 0.2
62
+ top_p: 0.8
63
+ top_k: 20
64
+ min_p: 0.05
65
+ repeat_penalty: 1.0
66
+ presence_penalty: 0.0
67
+ frequency_penalty: 0.0
68
+ model-p: # Planner
69
+ temperature: 1.0
70
+ top_p: 0.95
71
+ top_k: 20
72
+ min_p: 0.0
73
+ repeat_penalty: 1.0
74
+ presence_penalty: 0.0
75
+ frequency_penalty: 0.0
76
+ ```
77
+
78
+ > Each model is a **complete** sampling set — every field is required, so configure each alias with all seven numbers.
79
+
80
+ ## Ports
81
+
82
+ This plugin **never hardcodes an LLM port**. It matches on the `/v1/chat/completions` path in the request URL — the host and port come from your dsh provider configuration (e.g. `llm-pi-ai.providers.llama.baseURL`). The examples below use `<host>:<port>` as a placeholder.
83
+
84
+ ## Wire field names
85
+
86
+ llama.cpp uses snake_case wire names. **Note:** repetition penalty is `repeat_penalty`, **not** `repetition_penalty`.
87
+
88
+ | Model key | llama.cpp wire field |
89
+ |-----------|----------------------|
90
+ | `temperature` | `temperature` |
91
+ | `top_p` | `top_p` |
92
+ | `top_k` | `top_k` |
93
+ | `min_p` | `min_p` |
94
+ | `repeat_penalty` | `repeat_penalty` |
95
+ | `presence_penalty` | `presence_penalty` |
96
+ | `frequency_penalty` | `frequency_penalty` |
97
+
98
+ ## Verifying the sampling was applied
99
+
100
+ llama-server's `/slots` endpoint echoes the sampling parameters the server actually adopted. Query it for a model to confirm a request's sampling took effect:
101
+
102
+ ```sh
103
+ curl "http://<host>:<port>/slots?model=<model-id>"
104
+ ```
105
+
106
+ The `params` object of the returned slot reflects `temperature`, `top_p`, `top_k`, `min_p`, `repeat_penalty`, `presence_penalty`, and `frequency_penalty` — a match against the model's configured values confirms the injection reached the server.
107
+
108
+ ## Compatibility
109
+
110
+ - Requires dsh with `@deepseek-ai/dsh-settings` and `@deepseek-ai/schemastery` (both ship with dsh).
111
+ - Target must be an OpenAI-compatible gateway that honors these wire fields (llama.cpp `llama-server` does).
112
+ - The wrapper is guarded by a `Symbol.for` flag so hot-reload never double-wraps.
113
+
114
+ ## License
115
+
116
+ MIT
package/README.zh.md ADDED
@@ -0,0 +1,116 @@
1
+ # dsh-llama-cpp-sampling-params
2
+
3
+ 一個 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`dsh`)插件,將**每個模型的採樣參數**注入發送到本地 **llama.cpp / llama-server** 閘道的每個 chat-completions 請求。
4
+
5
+ ## 為什麼
6
+
7
+ dsh 的 `LlmCallConfig` 只帶 `temperature` / `maxTokens` / `stop`——從不發送 `top_p`、`top_k`、`min_p`、`repeat_penalty`、`presence_penalty` 或 `frequency_penalty`。llama.cpp 全部支援這些**按請求**參數。本插件將匹配模型的採樣組蓋印到 chat-completions 呼叫的 wire body 上,讓你**切換模型別名**就能切換採樣行為(例如高溫創意角色 vs. 低溫編碼角色)——無需重新載入模型、不佔額外 VRAM。
8
+
9
+ ## 運作原理
10
+
11
+ 插件在 transport 層 wrap `globalThis.fetch`:
12
+
13
+ 1. dsh process 內每個 `fetch()` 呼叫都經過 wrapper。
14
+ 2. URL 包含 `/v1/chat/completions` 的請求,wrapper 解析 JSON body。
15
+ 3. 讀取 `body.model`(別名),在 `sampling-params` 的 `models` 表查詢。
16
+ 4. 找到就將採樣組蓋印到解析後的 body,重新序列化 `init.body`。
17
+ 5. 原始 fetch 發送修改後的 body。
18
+
19
+ **為什麼用 fetch wrapper:** dsh 的 agent-loop 請求到達時是 **deep-frozen**(mutate 會 throw),且 pi-ai adapter 不會從 `GenerateOptions` 轉發 `onPayload`。Wrap `fetch` 在所有這些抽象之下操作——它直接編輯已組好的 JSON body,不碰任何 dsh 內部。
20
+
21
+ ## 角色如何運作
22
+
23
+ - llama.cpp 可將**一個已載入的 GGUF 以多個別名**(如 `model-i`、`model-p`、`model-t`)提供,全部指向相同的底層權重——切換別名**不佔額外 VRAM**。
24
+ - dsh 的 `llm-pi-ai` 將每個別名列為可選模型,切換 UI 裡的模型就選定採樣角色。
25
+ - 本插件讀取每個請求的 `model` 欄位(別名),在 `sampling-params` 的 `models` 表查詢。**未配置的別名原樣通過。**
26
+
27
+ ## 零衝突
28
+
29
+ - dsh 已發送的欄位(`temperature` / `maxTokens`)留給 dsh,除非模型明確設定。
30
+ - 其他採樣欄位是 dsh 從不發送的,因此沒有競爭來源——插件僅覆蓋 llama.cpp server 預設值。
31
+
32
+ ## 安裝
33
+
34
+ ```sh
35
+ # 從 npm(發布後)
36
+ dsh plugin --profile web add dsh-llama-cpp-sampling-params
37
+
38
+ # 從本地目錄
39
+ dsh plugin --profile web add link:C:/path/to/dsh-llama-cpp-sampling-params
40
+ ```
41
+
42
+ 然後重啟 `dsh web`(或刷新 GUI 頁面)。
43
+
44
+ ## 配置模型
45
+
46
+ 在 `$DSH_HOME/settings.yaml` 設定每個模型的採樣表。該區段**即時生效**——修改無需重啟。
47
+
48
+ ```yaml
49
+ sampling-params:
50
+ models:
51
+ # 鍵 = 發送到 wire 的精確 model id(別名)。
52
+ model-t: # Think
53
+ temperature: 0.6
54
+ top_p: 0.95
55
+ top_k: 20
56
+ min_p: 0.05
57
+ repeat_penalty: 1.0
58
+ presence_penalty: 0.0
59
+ frequency_penalty: 0.0
60
+ model-i: # Instruct
61
+ temperature: 0.2
62
+ top_p: 0.8
63
+ top_k: 20
64
+ min_p: 0.05
65
+ repeat_penalty: 1.0
66
+ presence_penalty: 0.0
67
+ frequency_penalty: 0.0
68
+ model-p: # Planner
69
+ temperature: 1.0
70
+ top_p: 0.95
71
+ top_k: 20
72
+ min_p: 0.0
73
+ repeat_penalty: 1.0
74
+ presence_penalty: 0.0
75
+ frequency_penalty: 0.0
76
+ ```
77
+
78
+ > 每個模型是**完整**採樣組——每個欄位必填,因此為每個別名配置全部七個數值。
79
+
80
+ ## 連接埠
81
+
82
+ 本插件**從不硬編碼 LLM port**。它匹配請求 URL 中的 `/v1/chat/completions` 路徑——host 和 port 來自你的 dsh provider 配置(如 `llm-pi-ai.providers.llama.baseURL`)。以下範例以 `<host>:<port>` 作為佔位符。
83
+
84
+ ## Wire 欄位名稱
85
+
86
+ llama.cpp 使用 snake_case wire 名稱。**注意:** 重複懲罰是 `repeat_penalty`,**不是** `repetition_penalty`。
87
+
88
+ | 模型鍵 | llama.cpp wire 欄位 |
89
+ |--------|---------------------|
90
+ | `temperature` | `temperature` |
91
+ | `top_p` | `top_p` |
92
+ | `top_k` | `top_k` |
93
+ | `min_p` | `min_p` |
94
+ | `repeat_penalty` | `repeat_penalty` |
95
+ | `presence_penalty` | `presence_penalty` |
96
+ | `frequency_penalty` | `frequency_penalty` |
97
+
98
+ ## 驗證採樣已生效
99
+
100
+ llama-server 的 `/slots` 端點回顯 server 實際採用的採樣參數。查詢模型以確認請求的採樣生效:
101
+
102
+ ```sh
103
+ curl "http://<host>:<port>/slots?model=<model-id>"
104
+ ```
105
+
106
+ 返回 slot 的 `params` 物件反映 `temperature`、`top_p`、`top_k`、`min_p`、`repeat_penalty`、`presence_penalty` 和 `frequency_penalty`——與模型配置值相符即確認注入已達 server。
107
+
108
+ ## 相容性
109
+
110
+ - 需要 dsh 內含 `@deepseek-ai/dsh-settings` 與 `@deepseek-ai/schemastery`(兩者隨 dsh 提供)。
111
+ - 目標必須是遵循這些 wire 欄位的 OpenAI 相容閘道(llama.cpp `llama-server` 符合)。
112
+ - Wrapper 以 `Symbol.for` 標記保護,hot-reload 不會重複 wrap。
113
+
114
+ ## 授權
115
+
116
+ MIT
@@ -0,0 +1,49 @@
1
+ # dsh-llama-cpp-sampling-params — bundle patch.
2
+ #
3
+ # Registers the plugin on the host. The plugin wraps globalThis.fetch to
4
+ # intercept chat-completions requests and stamp the matching sampling set
5
+ # onto the wire body before it leaves the process.
6
+ # The config block below is the BASE layer: a `sampling-params:` section in
7
+ # $DSH_HOME/settings.yaml overrides it without a restart.
8
+ - insert:
9
+ - id: sampling-params
10
+ name: dsh-llama-cpp-sampling-params
11
+ config:
12
+ # One entry per model. Key = the exact model id (alias) sent on the
13
+ # wire (the `model` field in the chat-completions request body).
14
+ #
15
+ # The plugin never hardcodes an LLM port — it matches on the
16
+ # /v1/chat/completions path in the request URL. Host and port come
17
+ # from dsh's provider baseURL (e.g. llm-pi-ai.providers.llama.baseURL).
18
+ #
19
+ # Wire field names (llama.cpp): temperature, top_p, top_k, min_p,
20
+ # repeat_penalty (NOT repetition_penalty), presence_penalty,
21
+ # frequency_penalty. Each model is a complete set — all fields required.
22
+ #
23
+ # Example (keys are placeholders):
24
+ # models:
25
+ # model-t: # Think
26
+ # temperature: 0.6
27
+ # top_p: 0.95
28
+ # top_k: 20
29
+ # min_p: 0.05
30
+ # repeat_penalty: 1.0
31
+ # presence_penalty: 0.0
32
+ # frequency_penalty: 0.0
33
+ # model-i: # Instruct
34
+ # temperature: 0.2
35
+ # top_p: 0.8
36
+ # top_k: 20
37
+ # min_p: 0.05
38
+ # repeat_penalty: 1.0
39
+ # presence_penalty: 0.0
40
+ # frequency_penalty: 0.0
41
+ # model-p: # Planner
42
+ # temperature: 1.0
43
+ # top_p: 0.95
44
+ # top_k: 20
45
+ # min_p: 0.0
46
+ # repeat_penalty: 1.0
47
+ # presence_penalty: 0.0
48
+ # frequency_penalty: 0.0
49
+ models: {}
package/index.js ADDED
@@ -0,0 +1,138 @@
1
+ /**
2
+ * dsh-llama-cpp-sampling-params
3
+ *
4
+ * Injects per-model-alias sampling parameters into every chat-completions
5
+ * request sent to an OpenAI-compatible local gateway (llama.cpp).
6
+ *
7
+ * How models work:
8
+ * - llama.cpp serves ONE loaded model under several aliases (e.g.
9
+ * `model-i`, `model-p`, `model-t`) that all point at the same underlying
10
+ * GGUF — so there is NO extra VRAM cost.
11
+ * - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
12
+ * switching model in the UI picks a sampling role.
13
+ * - This plugin reads the `model` field of each request (the alias) and
14
+ * looks it up in the `sampling-params` `models` table. If the alias is
15
+ * configured, it stamps that model's sampling set onto the wire body.
16
+ * Unconfigured aliases pass through byte-identical.
17
+ *
18
+ * Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
19
+ * maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
20
+ * presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
21
+ * deep-frozen (mutation throws) and the pi-ai adapter does not forward
22
+ * `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
23
+ * the transport layer, below all of those abstractions, and simply rewrites
24
+ * the already-built JSON body before it leaves the process.
25
+ *
26
+ * Zero conflict: the fields dsh already sends (temperature / maxTokens) are
27
+ * left to dsh unless a model explicitly sets them; every other sampling field
28
+ * is one dsh never sends, so there is no competing source.
29
+ */
30
+ import z from "@deepseek-ai/schemastery";
31
+ import { settingsNamespace } from "@deepseek-ai/dsh-settings";
32
+
33
+ const name = "sampling-params";
34
+ const inject = ["settings"];
35
+
36
+ // llama.cpp wire field names (snake_case). NOTE: repetition penalty is
37
+ // `repeat_penalty` on llama.cpp, NOT `repetition_penalty`.
38
+ const WIRE_KEYS = [
39
+ "temperature",
40
+ "top_p",
41
+ "top_k",
42
+ "min_p",
43
+ "repeat_penalty",
44
+ "presence_penalty",
45
+ "frequency_penalty",
46
+ ];
47
+
48
+ // One model's full sampling set. Each field is required (a model sets every
49
+ // number explicitly); a model that omits a field is rejected by the schema, so
50
+ // configure each alias with a complete set.
51
+ const Model = z.object({
52
+ temperature: z.number(),
53
+ top_p: z.number(),
54
+ top_k: z.number(),
55
+ min_p: z.number(),
56
+ repeat_penalty: z.number(),
57
+ presence_penalty: z.number(),
58
+ frequency_penalty: z.number(),
59
+ });
60
+
61
+ // Plugin config: a table keyed by the exact model id (alias) sent on the
62
+ // wire. Each key's value is the sampling set stamped onto requests that name
63
+ // that alias. Example:
64
+ // models:
65
+ // model-i: { temperature: 0.2, top_p: 0.8, ... }
66
+ // model-p: { temperature: 1.0, ... }
67
+ // model-t: { temperature: 0.6, ... }
68
+ const Config = z.object({
69
+ models: z.dict(Model).default({}),
70
+ });
71
+
72
+ // Mark so we never double-wrap globalThis.fetch (hot-reload safe).
73
+ const WRAPPED = Symbol.for("dsh-llama-cpp-sampling-params.fetch-wrapped");
74
+
75
+ // The OpenAI-compatible chat-completions path we inject into. Only requests
76
+ // whose URL carries this path are touched; everything else passes through
77
+ // byte-identical.
78
+ const CHAT_COMPLETIONS_PATH = "/v1/chat/completions";
79
+
80
+ function apply(ctx, config = {}) {
81
+ const log = ctx.logger("sampling-params");
82
+
83
+ // Register the settings section live so role values can be edited without
84
+ // a restart.
85
+ const scope = ctx.settings.register(settingsNamespace("sampling-params"), Config, {
86
+ base: config,
87
+ applies: "live",
88
+ });
89
+
90
+ // Read the sampling set for one exact model id/alias.
91
+ const readModel = (modelId) => {
92
+ try {
93
+ const section = scope.get();
94
+ return section?.models?.[modelId];
95
+ } catch {
96
+ return undefined;
97
+ }
98
+ };
99
+
100
+ const registeredModels = (() => {
101
+ try { return Object.keys(scope.get()?.models ?? {}); } catch { return []; }
102
+ })();
103
+ log.info(`sampling-params applied (fetch wrapper); models=${registeredModels.length}`);
104
+
105
+ if (typeof globalThis.fetch === "function" && !globalThis.fetch[WRAPPED]) {
106
+ const originalFetch = globalThis.fetch;
107
+ globalThis.fetch = async function (input, init) {
108
+ try {
109
+ const url = typeof input === "string" ? input : input?.url;
110
+ const bodyText = typeof init?.body === "string" ? init.body : null;
111
+ if (
112
+ typeof url === "string" &&
113
+ url.includes(CHAT_COMPLETIONS_PATH) &&
114
+ bodyText !== null
115
+ ) {
116
+ const body = JSON.parse(bodyText);
117
+ if (body && typeof body.model === "string") {
118
+ const model = readModel(body.model);
119
+ if (model) {
120
+ for (const key of WIRE_KEYS) {
121
+ if (typeof model[key] === "number") body[key] = model[key];
122
+ }
123
+ init.body = JSON.stringify(body);
124
+ }
125
+ }
126
+ }
127
+ } catch (error) {
128
+ // Never break LLM traffic: on any parse/matching error, send the
129
+ // request through untouched.
130
+ log.warn(`sampling-params inject skipped: ${error?.message ?? error}`);
131
+ }
132
+ return originalFetch.call(this, input, init);
133
+ };
134
+ globalThis.fetch[WRAPPED] = true;
135
+ }
136
+ }
137
+
138
+ export { apply, name, inject, Config };
package/package.json ADDED
@@ -0,0 +1,47 @@
1
+ {
2
+ "name": "dsh-llama-cpp-sampling-params",
3
+ "version": "0.1.0",
4
+ "description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to a local llama.cpp / llama-server gateway.",
5
+ "author": "devhang",
6
+ "repository": {
7
+ "type": "git",
8
+ "url": "git+https://github.com/devhang/dsh-llama-cpp-sampling-params.git"
9
+ },
10
+ "homepage": "https://github.com/devhang/dsh-llama-cpp-sampling-params#readme",
11
+ "bugs": {
12
+ "url": "https://github.com/devhang/dsh-llama-cpp-sampling-params/issues"
13
+ },
14
+ "type": "module",
15
+ "main": "index.js",
16
+ "files": [
17
+ "index.js",
18
+ "cordis.patch.yml",
19
+ "README.md",
20
+ "README.zh.md",
21
+ "LICENSE"
22
+ ],
23
+ "keywords": [
24
+ "deepseek-harness",
25
+ "dsh",
26
+ "dsh-plugin",
27
+ "llama.cpp",
28
+ "llama-server",
29
+ "sampling",
30
+ "top_p",
31
+ "min_p",
32
+ "repeat_penalty"
33
+ ],
34
+ "license": "MIT",
35
+ "dsh": {
36
+ "bundle": {
37
+ "patch": "./cordis.patch.yml"
38
+ }
39
+ },
40
+ "peerDependencies": {
41
+ "@deepseek-ai/cordis": "^4.0.1",
42
+ "@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.2.0-0"
43
+ },
44
+ "dependencies": {
45
+ "@deepseek-ai/schemastery": "^3.18.2"
46
+ }
47
+ }