dsh-llm-sampling-params 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +139 -0
- package/README.zh.md +139 -0
- package/cordis.patch.yml +50 -0
- package/index.js +216 -0
- package/package.json +51 -0
package/LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 dsh-llm-sampling-params contributors
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
package/README.md
ADDED
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# dsh-llm-sampling-params
|
|
2
|
+
|
|
3
|
+
A [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`dsh`) plugin that injects **per-model sampling parameters** into every chat-completions request sent to an OpenAI-compatible LLM gateway (llama.cpp, SGLang, vLLM).
|
|
4
|
+
|
|
5
|
+
## Why
|
|
6
|
+
|
|
7
|
+
dsh's `LlmCallConfig` only carries `temperature` / `maxTokens` / `stop` — it never sends `top_p`, `top_k`, `min_p`, `repeat_penalty`, `presence_penalty`, or `frequency_penalty`. OpenAI-compatible gateways (llama.cpp, SGLang, vLLM) accept all of these **per request**. This plugin stamps the matching model's sampling set onto the wire body of chat-completions calls, so you can switch sampling behavior (e.g. a high-temperature creative role vs. a low-temperature coding role) **by switching the model alias** — no model reload, no extra VRAM.
|
|
8
|
+
|
|
9
|
+
## How it works
|
|
10
|
+
|
|
11
|
+
The plugin wraps `globalThis.fetch` at the transport layer:
|
|
12
|
+
|
|
13
|
+
1. Every `fetch()` call in the dsh process flows through the wrapper.
|
|
14
|
+
2. For requests whose URL contains `/v1/chat/completions`, the wrapper parses the JSON body.
|
|
15
|
+
3. It reads `body.model` (the alias), looks it up in the `sampling-params` `models` table.
|
|
16
|
+
4. If found, it stamps the sampling set onto the parsed body and re-serializes `init.body`.
|
|
17
|
+
5. The original fetch sends the modified body.
|
|
18
|
+
|
|
19
|
+
**Why a fetch wrapper:** dsh's agent-loop requests arrive **deep-frozen** (mutation throws), and the pi-ai adapter does not forward `onPayload` from `GenerateOptions`. Wrapping `fetch` operates below all of those abstractions — it edits the already-built JSON body before it leaves the process, so no dsh internal is touched.
|
|
20
|
+
|
|
21
|
+
## How roles work
|
|
22
|
+
|
|
23
|
+
- A gateway can serve **one loaded model under several aliases** (e.g. `model-i`, `model-p`, `model-t`), all pointing at the same underlying weights — so switching alias costs **no extra VRAM**. (llama.cpp: `a =` alias list; vLLM: multiple `--served-model-name`; SGLang: single served name with lenient matching.)
|
|
24
|
+
- dsh's `llm-pi-ai` lists each alias as a separate selectable model, so switching the model in the UI picks a sampling role.
|
|
25
|
+
- This plugin reads the `model` field of each request (the alias) and looks it up in the `sampling-params` `models` table. **Unconfigured aliases pass through untouched.**
|
|
26
|
+
|
|
27
|
+
## Zero conflict
|
|
28
|
+
|
|
29
|
+
- The fields dsh already sends (`temperature` / `maxTokens`) are left to dsh unless a model explicitly sets them.
|
|
30
|
+
- Every other sampling field is one dsh never sends, so there is no competing source — the plugin simply overrides the server default.
|
|
31
|
+
|
|
32
|
+
## Install
|
|
33
|
+
|
|
34
|
+
```sh
|
|
35
|
+
# From npm (once published)
|
|
36
|
+
dsh plugin --profile web add dsh-llm-sampling-params
|
|
37
|
+
|
|
38
|
+
# From a local directory
|
|
39
|
+
dsh plugin --profile web add link:C:/path/to/dsh-llm-sampling-params
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
Then restart `dsh web` (or refresh the GUI page).
|
|
43
|
+
|
|
44
|
+
## Configure models
|
|
45
|
+
|
|
46
|
+
The `models` table lives in your dsh settings. **Where it goes depends on your dsh version:**
|
|
47
|
+
|
|
48
|
+
| dsh | where the table goes |
|
|
49
|
+
|-----|----------------------|
|
|
50
|
+
| **0.2.x** | the profile's `cordis.patch.yml`, as a top-level `sampling-params` entry |
|
|
51
|
+
| **0.1.x** | `$DSH_HOME/settings.yaml`, under a `sampling-params:` section (applies live) |
|
|
52
|
+
|
|
53
|
+
### DSH 0.2.x
|
|
54
|
+
|
|
55
|
+
Add a top-level entry to your profile's `cordis.patch.yml` (edit `cordis.patch.yml`, not `cordis.yml`). The `id: sampling-params` is the binding key and must match the plugin's namespace:
|
|
56
|
+
|
|
57
|
+
```yaml
|
|
58
|
+
- id: sampling-params
|
|
59
|
+
name: dsh-llm-sampling-params
|
|
60
|
+
config:
|
|
61
|
+
models:
|
|
62
|
+
# Key = the exact model id (alias) sent on the wire.
|
|
63
|
+
model-t: # Think
|
|
64
|
+
temperature: 0.6
|
|
65
|
+
top_p: 0.95
|
|
66
|
+
top_k: 20
|
|
67
|
+
min_p: 0.05
|
|
68
|
+
repeat_penalty: 1.0
|
|
69
|
+
presence_penalty: 0.0
|
|
70
|
+
frequency_penalty: 0.0
|
|
71
|
+
model-i: # Instruct
|
|
72
|
+
temperature: 0.2
|
|
73
|
+
top_p: 0.8
|
|
74
|
+
top_k: 20
|
|
75
|
+
min_p: 0.05
|
|
76
|
+
repeat_penalty: 1.0
|
|
77
|
+
presence_penalty: 0.0
|
|
78
|
+
frequency_penalty: 0.0
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
### Migrating from 0.1.x to 0.2.x
|
|
82
|
+
|
|
83
|
+
Upgrading dsh to 0.2.x renames `$DSH_HOME/settings.yaml` to `settings.yaml.imported` (a frozen backup, no longer read). Your `sampling-params.models` table is stranded there. To migrate, copy the `models:` block out of `settings.yaml.imported` and add it as the `sampling-params` entry in your profile's `cordis.patch.yml`, re-indenting the block by **+2 spaces** to nest under `config:`. The one-time auto-import does not re-run, so do this by hand (or via the settings UI).
|
|
84
|
+
|
|
85
|
+
> Each model is a **complete** sampling set — every field is required, so configure each alias with all seven numbers.
|
|
86
|
+
|
|
87
|
+
## Ports
|
|
88
|
+
|
|
89
|
+
This plugin **never hardcodes an LLM port**. It matches on the `/v1/chat/completions` path in the request URL — the host and port come from your dsh provider configuration (e.g. `llm-pi-ai.providers.llama.baseURL`). The examples below use `<host>:<port>` as a placeholder.
|
|
90
|
+
|
|
91
|
+
## Wire field names
|
|
92
|
+
|
|
93
|
+
The wire body uses snake_case names. The config's `repeat_penalty` is sent as **both** `repeat_penalty` (llama.cpp) and `repetition_penalty` (SGLang / vLLM), so the right one is honored regardless of backend.
|
|
94
|
+
|
|
95
|
+
| Model key | wire field(s) sent |
|
|
96
|
+
|-----------|----------------------|
|
|
97
|
+
| `temperature` | `temperature` |
|
|
98
|
+
| `top_p` | `top_p` |
|
|
99
|
+
| `top_k` | `top_k` |
|
|
100
|
+
| `min_p` | `min_p` |
|
|
101
|
+
| `repeat_penalty` | `repeat_penalty` + `repetition_penalty` |
|
|
102
|
+
| `presence_penalty` | `presence_penalty` |
|
|
103
|
+
| `frequency_penalty` | `frequency_penalty` |
|
|
104
|
+
|
|
105
|
+
## Verifying the sampling was applied
|
|
106
|
+
|
|
107
|
+
llama.cpp's `/slots` endpoint echoes the sampling parameters the server actually adopted (llama.cpp-specific; for SGLang / vLLM use the debug log below). Query it for a model to confirm a request's sampling took effect:
|
|
108
|
+
|
|
109
|
+
```sh
|
|
110
|
+
curl "http://<host>:<port>/slots?model=<model-id>"
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
The `params` object of the returned slot reflects `temperature`, `top_p`, `top_k`, `min_p`, `repeat_penalty`, `presence_penalty`, and `frequency_penalty` — a match against the model's configured values confirms the injection reached the server.
|
|
114
|
+
|
|
115
|
+
## Debugging the injection (any provider)
|
|
116
|
+
|
|
117
|
+
To confirm the outgoing request actually carries the sampling params — works with **any** OpenAI-compatible provider, not just llama-server — enable the built-in debug log:
|
|
118
|
+
|
|
119
|
+
- Start dsh with env `DSH_SAMPLING_DEBUG=1`, **or** create the marker file `~/.dsh/_sampling-debug-on`.
|
|
120
|
+
- Restart dsh so `apply()` re-reads the switch.
|
|
121
|
+
- Have a conversation in the GUI (select a configured alias).
|
|
122
|
+
- Each injected request appends one JSON line to `~/.dsh/_sampling-debug.log`:
|
|
123
|
+
|
|
124
|
+
```json
|
|
125
|
+
{"ts":"2026-01-01T00:00:00.000Z","model":"model-i","injected":{"temperature":0.2,"top_p":0.8,"top_k":20,"min_p":0.05,"repeat_penalty":1.0,"presence_penalty":0.0,"frequency_penalty":0.0}}
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
- Disable by unsetting the env / deleting the marker file (takes effect on the next restart). The log is best-effort and never affects the request.
|
|
129
|
+
|
|
130
|
+
## Compatibility
|
|
131
|
+
|
|
132
|
+
- Supports dsh **0.1.x and 0.2.x** — the settings read path is version-adaptive (`settings.register` on 0.1.x, the profile-entry config on 0.2.x).
|
|
133
|
+
- Requires dsh with `@deepseek-ai/dsh-settings` and `@deepseek-ai/schemastery` (both ship with dsh).
|
|
134
|
+
- Target must be an OpenAI-compatible gateway that honors these wire fields (llama.cpp, SGLang, and vLLM do).
|
|
135
|
+
- The wrapper is guarded by a `Symbol.for` flag so hot-reload never double-wraps.
|
|
136
|
+
|
|
137
|
+
## License
|
|
138
|
+
|
|
139
|
+
MIT
|
package/README.zh.md
ADDED
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
# dsh-llm-sampling-params
|
|
2
|
+
|
|
3
|
+
一個 [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)(`dsh`)插件,將**每個模型的採樣參數**注入發送到 OpenAI 相容 LLM 閘道(llama.cpp、SGLang、vLLM)的每個 chat-completions 請求。
|
|
4
|
+
|
|
5
|
+
## 為什麼
|
|
6
|
+
|
|
7
|
+
dsh 的 `LlmCallConfig` 只帶 `temperature` / `maxTokens` / `stop`——從不發送 `top_p`、`top_k`、`min_p`、`repeat_penalty`、`presence_penalty` 或 `frequency_penalty`。OpenAI 相容閘道(llama.cpp、SGLang、vLLM)全部支援這些**按請求**參數。本插件將匹配模型的採樣組蓋印到 chat-completions 呼叫的 wire body 上,讓你**切換模型別名**就能切換採樣行為(例如高溫創意角色 vs. 低溫編碼角色)——無需重新載入模型、不佔額外 VRAM。
|
|
8
|
+
|
|
9
|
+
## 運作原理
|
|
10
|
+
|
|
11
|
+
插件在 transport 層 wrap `globalThis.fetch`:
|
|
12
|
+
|
|
13
|
+
1. dsh process 內每個 `fetch()` 呼叫都經過 wrapper。
|
|
14
|
+
2. URL 包含 `/v1/chat/completions` 的請求,wrapper 解析 JSON body。
|
|
15
|
+
3. 讀取 `body.model`(別名),在 `sampling-params` 的 `models` 表查詢。
|
|
16
|
+
4. 找到就將採樣組蓋印到解析後的 body,重新序列化 `init.body`。
|
|
17
|
+
5. 原始 fetch 發送修改後的 body。
|
|
18
|
+
|
|
19
|
+
**為什麼用 fetch wrapper:** dsh 的 agent-loop 請求到達時是 **deep-frozen**(mutate 會 throw),且 pi-ai adapter 不會從 `GenerateOptions` 轉發 `onPayload`。Wrap `fetch` 在所有這些抽象之下操作——它直接編輯已組好的 JSON body,不碰任何 dsh 內部。
|
|
20
|
+
|
|
21
|
+
## 角色如何運作
|
|
22
|
+
|
|
23
|
+
- 閘道可將**一個已載入的模型以多個別名**(如 `model-i`、`model-p`、`model-t`)提供,全部指向相同的底層權重——切換別名**不佔額外 VRAM**。(llama.cpp:`a =` alias 清單;vLLM:多個 `--served-model-name`;SGLang:單一 served name、寬鬆匹配。)
|
|
24
|
+
- dsh 的 `llm-pi-ai` 將每個別名列為可選模型,切換 UI 裡的模型就選定採樣角色。
|
|
25
|
+
- 本插件讀取每個請求的 `model` 欄位(別名),在 `sampling-params` 的 `models` 表查詢。**未配置的別名原樣通過。**
|
|
26
|
+
|
|
27
|
+
## 零衝突
|
|
28
|
+
|
|
29
|
+
- dsh 已發送的欄位(`temperature` / `maxTokens`)留給 dsh,除非模型明確設定。
|
|
30
|
+
- 其他採樣欄位是 dsh 從不發送的,因此沒有競爭來源——插件僅覆蓋 server 預設值。
|
|
31
|
+
|
|
32
|
+
## 安裝
|
|
33
|
+
|
|
34
|
+
```sh
|
|
35
|
+
# 從 npm(發布後)
|
|
36
|
+
dsh plugin --profile web add dsh-llm-sampling-params
|
|
37
|
+
|
|
38
|
+
# 從本地目錄
|
|
39
|
+
dsh plugin --profile web add link:C:/path/to/dsh-llm-sampling-params
|
|
40
|
+
```
|
|
41
|
+
|
|
42
|
+
然後重啟 `dsh web`(或刷新 GUI 頁面)。
|
|
43
|
+
|
|
44
|
+
## 配置模型
|
|
45
|
+
|
|
46
|
+
`models` 表放在你的 dsh 設定裡。**放哪裡取決於你的 dsh 版本:**
|
|
47
|
+
|
|
48
|
+
| dsh | 表放哪裡 |
|
|
49
|
+
|-----|---------|
|
|
50
|
+
| **0.2.x** | profile 的 `cordis.patch.yml`,作為頂層 `sampling-params` entry |
|
|
51
|
+
| **0.1.x** | `$DSH_HOME/settings.yaml` 的 `sampling-params:` 區段(即時生效) |
|
|
52
|
+
|
|
53
|
+
### DSH 0.2.x
|
|
54
|
+
|
|
55
|
+
在 profile 的 `cordis.patch.yml` 加一個頂層 entry(改 `cordis.patch.yml`,不是 `cordis.yml`)。`id: sampling-params` 是 binding key,必須與插件 namespace 相符:
|
|
56
|
+
|
|
57
|
+
```yaml
|
|
58
|
+
- id: sampling-params
|
|
59
|
+
name: dsh-llm-sampling-params
|
|
60
|
+
config:
|
|
61
|
+
models:
|
|
62
|
+
# 鍵 = 發送到 wire 的精確 model id(別名)。
|
|
63
|
+
model-t: # Think
|
|
64
|
+
temperature: 0.6
|
|
65
|
+
top_p: 0.95
|
|
66
|
+
top_k: 20
|
|
67
|
+
min_p: 0.05
|
|
68
|
+
repeat_penalty: 1.0
|
|
69
|
+
presence_penalty: 0.0
|
|
70
|
+
frequency_penalty: 0.0
|
|
71
|
+
model-i: # Instruct
|
|
72
|
+
temperature: 0.2
|
|
73
|
+
top_p: 0.8
|
|
74
|
+
top_k: 20
|
|
75
|
+
min_p: 0.05
|
|
76
|
+
repeat_penalty: 1.0
|
|
77
|
+
presence_penalty: 0.0
|
|
78
|
+
frequency_penalty: 0.0
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
### 從 0.1.x 遷移到 0.2.x
|
|
82
|
+
|
|
83
|
+
升級 dsh 到 0.2.x 會把 `$DSH_HOME/settings.yaml` 改名成 `settings.yaml.imported`(凍結備份,不再讀取),你的 `sampling-params.models` 表就擱淺在那裡。要遷移:把 `models:` 區塊從 `settings.yaml.imported` 抽出,加進 profile 的 `cordis.patch.yml` 作為 `sampling-params` entry,並將該區塊**+2 空格**縮排以嵌在 `config:` 下。一次性自動 import 不會重跑,所以要手動搬(或經設定 UI)。
|
|
84
|
+
|
|
85
|
+
> 每個模型是**完整**採樣組——每個欄位必填,因此為每個別名配置全部七個數值。
|
|
86
|
+
|
|
87
|
+
## 連接埠
|
|
88
|
+
|
|
89
|
+
本插件**從不硬編碼 LLM port**。它匹配請求 URL 中的 `/v1/chat/completions` 路徑——host 和 port 來自你的 dsh provider 配置(如 `llm-pi-ai.providers.llama.baseURL`)。以下範例以 `<host>:<port>` 作為佔位符。
|
|
90
|
+
|
|
91
|
+
## Wire 欄位名稱
|
|
92
|
+
|
|
93
|
+
wire body 使用 snake_case 名稱。配置的 `repeat_penalty` 會**同時**以 `repeat_penalty`(llama.cpp)和 `repetition_penalty`(SGLang / vLLM)兩個名稱送出,所以無論後端是哪家,正確的那個都會被採用。
|
|
94
|
+
|
|
95
|
+
| 模型鍵 | wire 欄位(送出) |
|
|
96
|
+
|--------|---------------------|
|
|
97
|
+
| `temperature` | `temperature` |
|
|
98
|
+
| `top_p` | `top_p` |
|
|
99
|
+
| `top_k` | `top_k` |
|
|
100
|
+
| `min_p` | `min_p` |
|
|
101
|
+
| `repeat_penalty` | `repeat_penalty` + `repetition_penalty` |
|
|
102
|
+
| `presence_penalty` | `presence_penalty` |
|
|
103
|
+
| `frequency_penalty` | `frequency_penalty` |
|
|
104
|
+
|
|
105
|
+
## 驗證採樣已生效
|
|
106
|
+
|
|
107
|
+
llama.cpp 的 `/slots` 端點回顯 server 實際採用的採樣參數(llama.cpp 專屬;SGLang / vLLM 用下方的 debug log)。查詢模型以確認請求的採樣生效:
|
|
108
|
+
|
|
109
|
+
```sh
|
|
110
|
+
curl "http://<host>:<port>/slots?model=<model-id>"
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
返回 slot 的 `params` 物件反映 `temperature`、`top_p`、`top_k`、`min_p`、`repeat_penalty`、`presence_penalty` 和 `frequency_penalty`——與模型配置值相符即確認注入已達 server。
|
|
114
|
+
|
|
115
|
+
## 除錯注入(任何 provider)
|
|
116
|
+
|
|
117
|
+
要確認出向請求**真的**帶上 sampling 參數——對**任何** OpenAI 相容 provider 都適用,不只 llama-server——開啟內建 debug log:
|
|
118
|
+
|
|
119
|
+
- 用 env `DSH_SAMPLING_DEBUG=1` 啟動 dsh,**或**建立 marker 檔 `~/.dsh/_sampling-debug-on`。
|
|
120
|
+
- 重啟 dsh 讓 `apply()` 重讀開關。
|
|
121
|
+
- 在 GUI 對話(選一個已配置的 alias)。
|
|
122
|
+
- 每個注入的請求會 append 一行 JSON 到 `~/.dsh/_sampling-debug.log`:
|
|
123
|
+
|
|
124
|
+
```json
|
|
125
|
+
{"ts":"2026-01-01T00:00:00.000Z","model":"model-i","injected":{"temperature":0.2,"top_p":0.8,"top_k":20,"min_p":0.05,"repeat_penalty":1.0,"presence_penalty":0.0,"frequency_penalty":0.0}}
|
|
126
|
+
```
|
|
127
|
+
|
|
128
|
+
- 取消:unset env / 刪 marker 檔(下次重啟生效)。log 是 best-effort,不影響請求。
|
|
129
|
+
|
|
130
|
+
## 相容性
|
|
131
|
+
|
|
132
|
+
- 支援 dsh **0.1.x 與 0.2.x**——設定讀取路徑是 version-adaptive(0.1.x 用 `settings.register`、0.2.x 用 profile-entry config)。
|
|
133
|
+
- 需要 dsh 內含 `@deepseek-ai/dsh-settings` 與 `@deepseek-ai/schemastery`(兩者隨 dsh 提供)。
|
|
134
|
+
- 目標必須是遵循這些 wire 欄位的 OpenAI 相容閘道(llama.cpp、SGLang、vLLM 都符合)。
|
|
135
|
+
- Wrapper 以 `Symbol.for` 標記保護,hot-reload 不會重複 wrap。
|
|
136
|
+
|
|
137
|
+
## 授權
|
|
138
|
+
|
|
139
|
+
MIT
|
package/cordis.patch.yml
ADDED
|
@@ -0,0 +1,50 @@
|
|
|
1
|
+
# dsh-llm-sampling-params — bundle patch.
|
|
2
|
+
#
|
|
3
|
+
# Registers the plugin on the host. The plugin wraps globalThis.fetch to
|
|
4
|
+
# intercept chat-completions requests and stamp the matching sampling set
|
|
5
|
+
# onto the wire body before it leaves the process.
|
|
6
|
+
# The config block below is the BASE layer: a `sampling-params:` section in
|
|
7
|
+
# $DSH_HOME/settings.yaml overrides it without a restart.
|
|
8
|
+
- insert:
|
|
9
|
+
- id: sampling-params
|
|
10
|
+
name: dsh-llm-sampling-params
|
|
11
|
+
config:
|
|
12
|
+
# One entry per model. Key = the exact model id (alias) sent on the
|
|
13
|
+
# wire (the `model` field in the chat-completions request body).
|
|
14
|
+
#
|
|
15
|
+
# The plugin never hardcodes an LLM port — it matches on the
|
|
16
|
+
# /v1/chat/completions path in the request URL. Host and port come
|
|
17
|
+
# from dsh's provider baseURL (e.g. llm-pi-ai.providers.llama.baseURL).
|
|
18
|
+
#
|
|
19
|
+
# Wire field names: temperature, top_p, top_k, min_p, presence_penalty,
|
|
20
|
+
# frequency_penalty. The repetition penalty is sent as BOTH repeat_penalty
|
|
21
|
+
# (llama.cpp) and repetition_penalty (SGLang / vLLM) so the right one is
|
|
22
|
+
# honored regardless of backend. Each model is a complete set — all fields required.
|
|
23
|
+
#
|
|
24
|
+
# Example (keys are placeholders):
|
|
25
|
+
# models:
|
|
26
|
+
# model-t: # Think
|
|
27
|
+
# temperature: 0.6
|
|
28
|
+
# top_p: 0.95
|
|
29
|
+
# top_k: 20
|
|
30
|
+
# min_p: 0.05
|
|
31
|
+
# repeat_penalty: 1.0
|
|
32
|
+
# presence_penalty: 0.0
|
|
33
|
+
# frequency_penalty: 0.0
|
|
34
|
+
# model-i: # Instruct
|
|
35
|
+
# temperature: 0.2
|
|
36
|
+
# top_p: 0.8
|
|
37
|
+
# top_k: 20
|
|
38
|
+
# min_p: 0.05
|
|
39
|
+
# repeat_penalty: 1.0
|
|
40
|
+
# presence_penalty: 0.0
|
|
41
|
+
# frequency_penalty: 0.0
|
|
42
|
+
# model-p: # Planner
|
|
43
|
+
# temperature: 1.0
|
|
44
|
+
# top_p: 0.95
|
|
45
|
+
# top_k: 20
|
|
46
|
+
# min_p: 0.0
|
|
47
|
+
# repeat_penalty: 1.0
|
|
48
|
+
# presence_penalty: 0.0
|
|
49
|
+
# frequency_penalty: 0.0
|
|
50
|
+
models: {}
|
package/index.js
ADDED
|
@@ -0,0 +1,216 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* dsh-llm-sampling-params
|
|
3
|
+
*
|
|
4
|
+
* Injects per-model-alias sampling parameters into every chat-completions
|
|
5
|
+
* request sent to an OpenAI-compatible LLM gateway (llama.cpp, SGLang, vLLM).
|
|
6
|
+
*
|
|
7
|
+
* How models work:
|
|
8
|
+
* - An OpenAI-compatible gateway serves ONE loaded model under several
|
|
9
|
+
* aliases (e.g. `model-i`, `model-p`, `model-t`) that all point at the
|
|
10
|
+
* same underlying weights — so there is NO extra VRAM cost. (llama.cpp:
|
|
11
|
+
* `a =` alias list; vLLM: multiple `--served-model-name`.)
|
|
12
|
+
* - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
|
|
13
|
+
* switching model in the UI picks a sampling role.
|
|
14
|
+
* - This plugin reads the `model` field of each request (the alias) and
|
|
15
|
+
* looks it up in the `sampling-params` `models` table. If the alias is
|
|
16
|
+
* configured, it stamps that model's sampling set onto the wire body.
|
|
17
|
+
* Unconfigured aliases pass through byte-identical.
|
|
18
|
+
*
|
|
19
|
+
* Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
|
|
20
|
+
* maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
|
|
21
|
+
* presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
|
|
22
|
+
* deep-frozen (mutation throws) and the pi-ai adapter does not forward
|
|
23
|
+
* `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
|
|
24
|
+
* the transport layer, below all of those abstractions, and simply rewrites
|
|
25
|
+
* the already-built JSON body before it leaves the process.
|
|
26
|
+
*
|
|
27
|
+
* Config source across DSH versions: the `models` table is the plugin's
|
|
28
|
+
* settings. On DSH 0.1.x the plugin registers that section with
|
|
29
|
+
* `ctx.settings.register(...)` and reads it live via the returned scope. On
|
|
30
|
+
* 0.2.x the settings service was rebuilt around a profile-entry model
|
|
31
|
+
* (`describe` / `mutate` / `configure`) and no longer exposes `register`; a
|
|
32
|
+
* settings edit restarts the plugin fiber and re-runs `apply(ctx, config)`
|
|
33
|
+
* with the new config, so the plugin caches that config in a module-level
|
|
34
|
+
* `liveConfig` and reads it there. `apply` detects which path is available at
|
|
35
|
+
* runtime, so the same code works on both.
|
|
36
|
+
*
|
|
37
|
+
* Zero conflict: the fields dsh already sends (temperature / maxTokens) are
|
|
38
|
+
* left to dsh unless a model explicitly sets them; every other sampling field
|
|
39
|
+
* is one dsh never sends, so there is no competing source.
|
|
40
|
+
*/
|
|
41
|
+
import z from "@deepseek-ai/schemastery";
|
|
42
|
+
import fs from "node:fs";
|
|
43
|
+
import os from "node:os";
|
|
44
|
+
import path from "node:path";
|
|
45
|
+
|
|
46
|
+
const name = "sampling-params";
|
|
47
|
+
const inject = ["settings"];
|
|
48
|
+
|
|
49
|
+
// Wire field names (snake_case). The config uses `repeat_penalty`; the wire
|
|
50
|
+
// body sends BOTH `repeat_penalty` (llama.cpp) and `repetition_penalty`
|
|
51
|
+
// (SGLang / vLLM) so the right one is honored regardless of backend.
|
|
52
|
+
const WIRE_KEYS = [
|
|
53
|
+
"temperature",
|
|
54
|
+
"top_p",
|
|
55
|
+
"top_k",
|
|
56
|
+
"min_p",
|
|
57
|
+
"repeat_penalty",
|
|
58
|
+
"presence_penalty",
|
|
59
|
+
"frequency_penalty",
|
|
60
|
+
];
|
|
61
|
+
|
|
62
|
+
// One model's full sampling set. Each field is required (a model sets every
|
|
63
|
+
// number explicitly); a model that omits a field is rejected by the schema, so
|
|
64
|
+
// configure each alias with a complete set.
|
|
65
|
+
const Model = z.object({
|
|
66
|
+
temperature: z.number(),
|
|
67
|
+
top_p: z.number(),
|
|
68
|
+
top_k: z.number(),
|
|
69
|
+
min_p: z.number(),
|
|
70
|
+
repeat_penalty: z.number(),
|
|
71
|
+
presence_penalty: z.number(),
|
|
72
|
+
frequency_penalty: z.number(),
|
|
73
|
+
});
|
|
74
|
+
|
|
75
|
+
// Plugin config: a table keyed by the exact model id (alias) sent on the
|
|
76
|
+
// wire. Each key's value is the sampling set stamped onto requests that name
|
|
77
|
+
// that alias. Example:
|
|
78
|
+
// models:
|
|
79
|
+
// model-i: { temperature: 0.2, top_p: 0.8, ... }
|
|
80
|
+
// model-p: { temperature: 1.0, ... }
|
|
81
|
+
// model-t: { temperature: 0.6, ... }
|
|
82
|
+
const Config = z.object({
|
|
83
|
+
models: z.dict(Model).default({}),
|
|
84
|
+
});
|
|
85
|
+
|
|
86
|
+
// Mark so we never double-wrap globalThis.fetch (hot-reload safe).
|
|
87
|
+
const WRAPPED = Symbol.for("dsh-llm-sampling-params.fetch-wrapped");
|
|
88
|
+
|
|
89
|
+
// The OpenAI-compatible chat-completions path we inject into. Only requests
|
|
90
|
+
// whose URL carries this path are touched; everything else passes through
|
|
91
|
+
// byte-identical.
|
|
92
|
+
const CHAT_COMPLETIONS_PATH = "/v1/chat/completions";
|
|
93
|
+
|
|
94
|
+
// Module-level live config, refreshed on every apply() call. Under DSH 0.2.x a
|
|
95
|
+
// settings edit restarts the plugin fiber and re-runs apply() with the new
|
|
96
|
+
// config, so this always reflects the latest `models` table. Under 0.1.x the
|
|
97
|
+
// registered settings scope (created inside apply below) is the live source
|
|
98
|
+
// instead and this only holds the initial value.
|
|
99
|
+
let liveConfig = { models: {} };
|
|
100
|
+
|
|
101
|
+
// Debug switch (best-effort): when enabled, every injection appends one JSON
|
|
102
|
+
// line to ~/.dsh/_sampling-debug.log so you can confirm the outgoing body
|
|
103
|
+
// actually carries the sampling params. Enable with env DSH_SAMPLING_DEBUG=1
|
|
104
|
+
// or by creating the marker file ~/.dsh/_sampling-debug-on; disable by
|
|
105
|
+
// unsetting the env / deleting the file (takes effect on the next apply()).
|
|
106
|
+
let debugEnabled = false;
|
|
107
|
+
|
|
108
|
+
function apply(ctx, config = {}) {
|
|
109
|
+
const log = ctx.logger("sampling-params");
|
|
110
|
+
liveConfig = config;
|
|
111
|
+
try {
|
|
112
|
+
debugEnabled =
|
|
113
|
+
Boolean(process.env.DSH_SAMPLING_DEBUG) ||
|
|
114
|
+
fs.existsSync(path.join(os.homedir(), ".dsh", "_sampling-debug-on"));
|
|
115
|
+
} catch {
|
|
116
|
+
debugEnabled = false;
|
|
117
|
+
}
|
|
118
|
+
|
|
119
|
+
// Register a live settings scope on DSH 0.1.x, where `ctx.settings.register`
|
|
120
|
+
// exists. On 0.2.x that method is gone (the settings service moved to a
|
|
121
|
+
// profile-entry model: describe / mutate / configure), so scope stays null
|
|
122
|
+
// and readModel below falls back to the config that 0.2.x re-passes to apply
|
|
123
|
+
// on every restart. The namespace is passed as a plain string (not via
|
|
124
|
+
// `settingsNamespace()`) so importing it never breaks plugin load.
|
|
125
|
+
let scope = null;
|
|
126
|
+
const settings = ctx.settings;
|
|
127
|
+
if (settings && typeof settings.register === "function") {
|
|
128
|
+
try {
|
|
129
|
+
scope = settings.register("sampling-params", Config, { base: config, applies: "live" });
|
|
130
|
+
} catch (error) {
|
|
131
|
+
log.warn(`sampling-params settings.register failed; using config param: ${error?.message ?? error}`);
|
|
132
|
+
scope = null;
|
|
133
|
+
}
|
|
134
|
+
}
|
|
135
|
+
|
|
136
|
+
// Read the sampling set for one exact model id/alias. Prefer the live scope
|
|
137
|
+
// (0.1.x); otherwise read the latest config passed to apply (0.2.x restart).
|
|
138
|
+
const readModel = (modelId) => {
|
|
139
|
+
if (scope) {
|
|
140
|
+
try {
|
|
141
|
+
const model = scope.get()?.models?.[modelId];
|
|
142
|
+
if (model) return model;
|
|
143
|
+
} catch {
|
|
144
|
+
// fall through to the config-param path
|
|
145
|
+
}
|
|
146
|
+
}
|
|
147
|
+
try {
|
|
148
|
+
return liveConfig?.models?.[modelId];
|
|
149
|
+
} catch {
|
|
150
|
+
return undefined;
|
|
151
|
+
}
|
|
152
|
+
};
|
|
153
|
+
|
|
154
|
+
const registeredModels = (() => {
|
|
155
|
+
try {
|
|
156
|
+
const models = scope ? (scope.get()?.models ?? {}) : (liveConfig?.models ?? {});
|
|
157
|
+
return Object.keys(models);
|
|
158
|
+
} catch {
|
|
159
|
+
return [];
|
|
160
|
+
}
|
|
161
|
+
})();
|
|
162
|
+
log.info(`sampling-params applied (fetch wrapper); models=${registeredModels.length}`);
|
|
163
|
+
|
|
164
|
+
if (typeof globalThis.fetch === "function" && !globalThis.fetch[WRAPPED]) {
|
|
165
|
+
const originalFetch = globalThis.fetch;
|
|
166
|
+
globalThis.fetch = async function (input, init) {
|
|
167
|
+
try {
|
|
168
|
+
const url = typeof input === "string" ? input : input?.url;
|
|
169
|
+
const bodyText = typeof init?.body === "string" ? init.body : null;
|
|
170
|
+
if (
|
|
171
|
+
typeof url === "string" &&
|
|
172
|
+
url.includes(CHAT_COMPLETIONS_PATH) &&
|
|
173
|
+
bodyText !== null
|
|
174
|
+
) {
|
|
175
|
+
const body = JSON.parse(bodyText);
|
|
176
|
+
if (body && typeof body.model === "string") {
|
|
177
|
+
const model = readModel(body.model);
|
|
178
|
+
if (model) {
|
|
179
|
+
for (const key of WIRE_KEYS) {
|
|
180
|
+
if (typeof model[key] === "number") body[key] = model[key];
|
|
181
|
+
}
|
|
182
|
+
// SGLang names the repetition penalty `repetition_penalty` (llama.cpp
|
|
183
|
+
// uses `repeat_penalty`). Send both so the right one is honored regardless
|
|
184
|
+
// of backend; each server ignores the name it does not know.
|
|
185
|
+
if (typeof model.repeat_penalty === "number") {
|
|
186
|
+
body.repetition_penalty = model.repeat_penalty;
|
|
187
|
+
}
|
|
188
|
+
init.body = JSON.stringify(body);
|
|
189
|
+
if (debugEnabled) {
|
|
190
|
+
try {
|
|
191
|
+
const injected = {};
|
|
192
|
+
for (const key of WIRE_KEYS) if (typeof model[key] === "number") injected[key] = model[key];
|
|
193
|
+
if (typeof model.repeat_penalty === "number") injected.repetition_penalty = model.repeat_penalty;
|
|
194
|
+
fs.appendFileSync(
|
|
195
|
+
path.join(os.homedir(), ".dsh", "_sampling-debug.log"),
|
|
196
|
+
JSON.stringify({ ts: new Date().toISOString(), model: body.model, injected }) + "\n"
|
|
197
|
+
);
|
|
198
|
+
} catch {
|
|
199
|
+
// debug logging is best-effort; never break traffic
|
|
200
|
+
}
|
|
201
|
+
}
|
|
202
|
+
}
|
|
203
|
+
}
|
|
204
|
+
}
|
|
205
|
+
} catch (error) {
|
|
206
|
+
// Never break LLM traffic: on any parse/matching error, send the
|
|
207
|
+
// request through untouched.
|
|
208
|
+
log.warn(`sampling-params inject skipped: ${error?.message ?? error}`);
|
|
209
|
+
}
|
|
210
|
+
return originalFetch.call(this, input, init);
|
|
211
|
+
};
|
|
212
|
+
globalThis.fetch[WRAPPED] = true;
|
|
213
|
+
}
|
|
214
|
+
}
|
|
215
|
+
|
|
216
|
+
export { apply, name, inject, Config };
|
package/package.json
ADDED
|
@@ -0,0 +1,51 @@
|
|
|
1
|
+
{
|
|
2
|
+
"name": "dsh-llm-sampling-params",
|
|
3
|
+
"version": "0.2.1",
|
|
4
|
+
"description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty/repetition_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to an OpenAI-compatible gateway (llama.cpp, SGLang, vLLM).",
|
|
5
|
+
"author": "devhang",
|
|
6
|
+
"repository": {
|
|
7
|
+
"type": "git",
|
|
8
|
+
"url": "git+https://github.com/devhang/dsh-llm-sampling-params.git"
|
|
9
|
+
},
|
|
10
|
+
"homepage": "https://github.com/devhang/dsh-llm-sampling-params#readme",
|
|
11
|
+
"bugs": {
|
|
12
|
+
"url": "https://github.com/devhang/dsh-llm-sampling-params/issues"
|
|
13
|
+
},
|
|
14
|
+
"type": "module",
|
|
15
|
+
"main": "index.js",
|
|
16
|
+
"files": [
|
|
17
|
+
"index.js",
|
|
18
|
+
"cordis.patch.yml",
|
|
19
|
+
"README.md",
|
|
20
|
+
"README.zh.md",
|
|
21
|
+
"LICENSE"
|
|
22
|
+
],
|
|
23
|
+
"keywords": [
|
|
24
|
+
"deepseek-harness",
|
|
25
|
+
"dsh",
|
|
26
|
+
"dsh-plugin",
|
|
27
|
+
"llama.cpp",
|
|
28
|
+
"llama-server",
|
|
29
|
+
"sglang",
|
|
30
|
+
"vllm",
|
|
31
|
+
"openai-compatible",
|
|
32
|
+
"sampling",
|
|
33
|
+
"top_p",
|
|
34
|
+
"min_p",
|
|
35
|
+
"repeat_penalty",
|
|
36
|
+
"repetition_penalty"
|
|
37
|
+
],
|
|
38
|
+
"license": "MIT",
|
|
39
|
+
"dsh": {
|
|
40
|
+
"bundle": {
|
|
41
|
+
"patch": "./cordis.patch.yml"
|
|
42
|
+
}
|
|
43
|
+
},
|
|
44
|
+
"peerDependencies": {
|
|
45
|
+
"@deepseek-ai/cordis": "^4.0.1",
|
|
46
|
+
"@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.3.0-0"
|
|
47
|
+
},
|
|
48
|
+
"dependencies": {
|
|
49
|
+
"@deepseek-ai/schemastery": "^3.18.2"
|
|
50
|
+
}
|
|
51
|
+
}
|