dsh-llama-cpp-sampling-params 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +36 -28
- package/README.zh.md +36 -28
- package/index.js +178 -140
- package/package.json +47 -47
package/README.md
CHANGED
|
@@ -43,38 +43,45 @@ Then restart `dsh web` (or refresh the GUI page).
|
|
|
43
43
|
|
|
44
44
|
## Configure models
|
|
45
45
|
|
|
46
|
-
|
|
46
|
+
The `models` table lives in your dsh settings. **Where it goes depends on your dsh version:**
|
|
47
|
+
|
|
48
|
+
| dsh | where the table goes |
|
|
49
|
+
|-----|----------------------|
|
|
50
|
+
| **0.2.x** | the profile's `cordis.patch.yml`, as a top-level `sampling-params` entry |
|
|
51
|
+
| **0.1.x** | `$DSH_HOME/settings.yaml`, under a `sampling-params:` section (applies live) |
|
|
52
|
+
|
|
53
|
+
### DSH 0.2.x
|
|
54
|
+
|
|
55
|
+
Add a top-level entry to your profile's `cordis.patch.yml` (edit `cordis.patch.yml`, not `cordis.yml`). The `id: sampling-params` is the binding key and must match the plugin's namespace:
|
|
47
56
|
|
|
48
57
|
```yaml
|
|
49
|
-
sampling-params
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
top_p: 0.95
|
|
71
|
-
top_k: 20
|
|
72
|
-
min_p: 0.0
|
|
73
|
-
repeat_penalty: 1.0
|
|
74
|
-
presence_penalty: 0.0
|
|
75
|
-
frequency_penalty: 0.0
|
|
58
|
+
- id: sampling-params
|
|
59
|
+
name: dsh-llama-cpp-sampling-params
|
|
60
|
+
config:
|
|
61
|
+
models:
|
|
62
|
+
# Key = the exact model id (alias) sent on the wire.
|
|
63
|
+
model-t: # Think
|
|
64
|
+
temperature: 0.6
|
|
65
|
+
top_p: 0.95
|
|
66
|
+
top_k: 20
|
|
67
|
+
min_p: 0.05
|
|
68
|
+
repeat_penalty: 1.0
|
|
69
|
+
presence_penalty: 0.0
|
|
70
|
+
frequency_penalty: 0.0
|
|
71
|
+
model-i: # Instruct
|
|
72
|
+
temperature: 0.2
|
|
73
|
+
top_p: 0.8
|
|
74
|
+
top_k: 20
|
|
75
|
+
min_p: 0.05
|
|
76
|
+
repeat_penalty: 1.0
|
|
77
|
+
presence_penalty: 0.0
|
|
78
|
+
frequency_penalty: 0.0
|
|
76
79
|
```
|
|
77
80
|
|
|
81
|
+
### Migrating from 0.1.x to 0.2.x
|
|
82
|
+
|
|
83
|
+
Upgrading dsh to 0.2.x renames `$DSH_HOME/settings.yaml` to `settings.yaml.imported` (a frozen backup, no longer read). Your `sampling-params.models` table is stranded there. To migrate, copy the `models:` block out of `settings.yaml.imported` and add it as the `sampling-params` entry in your profile's `cordis.patch.yml`, re-indenting the block by **+2 spaces** to nest under `config:`. The one-time auto-import does not re-run, so do this by hand (or via the settings UI).
|
|
84
|
+
|
|
78
85
|
> Each model is a **complete** sampling set — every field is required, so configure each alias with all seven numbers.
|
|
79
86
|
|
|
80
87
|
## Ports
|
|
@@ -107,6 +114,7 @@ The `params` object of the returned slot reflects `temperature`, `top_p`, `top_k
|
|
|
107
114
|
|
|
108
115
|
## Compatibility
|
|
109
116
|
|
|
117
|
+
- Supports dsh **0.1.x and 0.2.x** — the settings read path is version-adaptive (`settings.register` on 0.1.x, the profile-entry config on 0.2.x).
|
|
110
118
|
- Requires dsh with `@deepseek-ai/dsh-settings` and `@deepseek-ai/schemastery` (both ship with dsh).
|
|
111
119
|
- Target must be an OpenAI-compatible gateway that honors these wire fields (llama.cpp `llama-server` does).
|
|
112
120
|
- The wrapper is guarded by a `Symbol.for` flag so hot-reload never double-wraps.
|
package/README.zh.md
CHANGED
|
@@ -43,38 +43,45 @@ dsh plugin --profile web add link:C:/path/to/dsh-llama-cpp-sampling-params
|
|
|
43
43
|
|
|
44
44
|
## 配置模型
|
|
45
45
|
|
|
46
|
-
|
|
46
|
+
`models` 表放在你的 dsh 設定裡。**放哪裡取決於你的 dsh 版本:**
|
|
47
|
+
|
|
48
|
+
| dsh | 表放哪裡 |
|
|
49
|
+
|-----|---------|
|
|
50
|
+
| **0.2.x** | profile 的 `cordis.patch.yml`,作為頂層 `sampling-params` entry |
|
|
51
|
+
| **0.1.x** | `$DSH_HOME/settings.yaml` 的 `sampling-params:` 區段(即時生效) |
|
|
52
|
+
|
|
53
|
+
### DSH 0.2.x
|
|
54
|
+
|
|
55
|
+
在 profile 的 `cordis.patch.yml` 加一個頂層 entry(改 `cordis.patch.yml`,不是 `cordis.yml`)。`id: sampling-params` 是 binding key,必須與插件 namespace 相符:
|
|
47
56
|
|
|
48
57
|
```yaml
|
|
49
|
-
sampling-params
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
top_p: 0.95
|
|
71
|
-
top_k: 20
|
|
72
|
-
min_p: 0.0
|
|
73
|
-
repeat_penalty: 1.0
|
|
74
|
-
presence_penalty: 0.0
|
|
75
|
-
frequency_penalty: 0.0
|
|
58
|
+
- id: sampling-params
|
|
59
|
+
name: dsh-llama-cpp-sampling-params
|
|
60
|
+
config:
|
|
61
|
+
models:
|
|
62
|
+
# 鍵 = 發送到 wire 的精確 model id(別名)。
|
|
63
|
+
model-t: # Think
|
|
64
|
+
temperature: 0.6
|
|
65
|
+
top_p: 0.95
|
|
66
|
+
top_k: 20
|
|
67
|
+
min_p: 0.05
|
|
68
|
+
repeat_penalty: 1.0
|
|
69
|
+
presence_penalty: 0.0
|
|
70
|
+
frequency_penalty: 0.0
|
|
71
|
+
model-i: # Instruct
|
|
72
|
+
temperature: 0.2
|
|
73
|
+
top_p: 0.8
|
|
74
|
+
top_k: 20
|
|
75
|
+
min_p: 0.05
|
|
76
|
+
repeat_penalty: 1.0
|
|
77
|
+
presence_penalty: 0.0
|
|
78
|
+
frequency_penalty: 0.0
|
|
76
79
|
```
|
|
77
80
|
|
|
81
|
+
### 從 0.1.x 遷移到 0.2.x
|
|
82
|
+
|
|
83
|
+
升級 dsh 到 0.2.x 會把 `$DSH_HOME/settings.yaml` 改名成 `settings.yaml.imported`(凍結備份,不再讀取),你的 `sampling-params.models` 表就擱淺在那裡。要遷移:把 `models:` 區塊從 `settings.yaml.imported` 抽出,加進 profile 的 `cordis.patch.yml` 作為 `sampling-params` entry,並將該區塊**+2 空格**縮排以嵌在 `config:` 下。一次性自動 import 不會重跑,所以要手動搬(或經設定 UI)。
|
|
84
|
+
|
|
78
85
|
> 每個模型是**完整**採樣組——每個欄位必填,因此為每個別名配置全部七個數值。
|
|
79
86
|
|
|
80
87
|
## 連接埠
|
|
@@ -107,6 +114,7 @@ curl "http://<host>:<port>/slots?model=<model-id>"
|
|
|
107
114
|
|
|
108
115
|
## 相容性
|
|
109
116
|
|
|
117
|
+
- 支援 dsh **0.1.x 與 0.2.x**——設定讀取路徑是 version-adaptive(0.1.x 用 `settings.register`、0.2.x 用 profile-entry config)。
|
|
110
118
|
- 需要 dsh 內含 `@deepseek-ai/dsh-settings` 與 `@deepseek-ai/schemastery`(兩者隨 dsh 提供)。
|
|
111
119
|
- 目標必須是遵循這些 wire 欄位的 OpenAI 相容閘道(llama.cpp `llama-server` 符合)。
|
|
112
120
|
- Wrapper 以 `Symbol.for` 標記保護,hot-reload 不會重複 wrap。
|
package/index.js
CHANGED
|
@@ -1,140 +1,178 @@
|
|
|
1
|
-
/**
|
|
2
|
-
* dsh-llama-cpp-sampling-params
|
|
3
|
-
*
|
|
4
|
-
* Injects per-model-alias sampling parameters into every chat-completions
|
|
5
|
-
* request sent to an OpenAI-compatible local gateway (llama.cpp).
|
|
6
|
-
*
|
|
7
|
-
* How models work:
|
|
8
|
-
* - llama.cpp serves ONE loaded model under several aliases (e.g.
|
|
9
|
-
* `model-i`, `model-p`, `model-t`) that all point at the same underlying
|
|
10
|
-
* GGUF — so there is NO extra VRAM cost.
|
|
11
|
-
* - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
|
|
12
|
-
* switching model in the UI picks a sampling role.
|
|
13
|
-
* - This plugin reads the `model` field of each request (the alias) and
|
|
14
|
-
* looks it up in the `sampling-params` `models` table. If the alias is
|
|
15
|
-
* configured, it stamps that model's sampling set onto the wire body.
|
|
16
|
-
* Unconfigured aliases pass through byte-identical.
|
|
17
|
-
*
|
|
18
|
-
* Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
|
|
19
|
-
* maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
|
|
20
|
-
* presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
|
|
21
|
-
* deep-frozen (mutation throws) and the pi-ai adapter does not forward
|
|
22
|
-
* `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
|
|
23
|
-
* the transport layer, below all of those abstractions, and simply rewrites
|
|
24
|
-
* the already-built JSON body before it leaves the process.
|
|
25
|
-
*
|
|
26
|
-
*
|
|
27
|
-
*
|
|
28
|
-
*
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
//
|
|
72
|
-
|
|
73
|
-
|
|
74
|
-
//
|
|
75
|
-
//
|
|
76
|
-
//
|
|
77
|
-
const
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
|
|
114
|
-
|
|
115
|
-
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
|
|
121
|
-
|
|
122
|
-
|
|
123
|
-
|
|
124
|
-
|
|
125
|
-
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
130
|
-
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
1
|
+
/**
|
|
2
|
+
* dsh-llama-cpp-sampling-params
|
|
3
|
+
*
|
|
4
|
+
* Injects per-model-alias sampling parameters into every chat-completions
|
|
5
|
+
* request sent to an OpenAI-compatible local gateway (llama.cpp).
|
|
6
|
+
*
|
|
7
|
+
* How models work:
|
|
8
|
+
* - llama.cpp serves ONE loaded model under several aliases (e.g.
|
|
9
|
+
* `model-i`, `model-p`, `model-t`) that all point at the same underlying
|
|
10
|
+
* GGUF — so there is NO extra VRAM cost.
|
|
11
|
+
* - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
|
|
12
|
+
* switching model in the UI picks a sampling role.
|
|
13
|
+
* - This plugin reads the `model` field of each request (the alias) and
|
|
14
|
+
* looks it up in the `sampling-params` `models` table. If the alias is
|
|
15
|
+
* configured, it stamps that model's sampling set onto the wire body.
|
|
16
|
+
* Unconfigured aliases pass through byte-identical.
|
|
17
|
+
*
|
|
18
|
+
* Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
|
|
19
|
+
* maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
|
|
20
|
+
* presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
|
|
21
|
+
* deep-frozen (mutation throws) and the pi-ai adapter does not forward
|
|
22
|
+
* `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
|
|
23
|
+
* the transport layer, below all of those abstractions, and simply rewrites
|
|
24
|
+
* the already-built JSON body before it leaves the process.
|
|
25
|
+
*
|
|
26
|
+
* Config source across DSH versions: the `models` table is the plugin's
|
|
27
|
+
* settings. On DSH 0.1.x the plugin registers that section with
|
|
28
|
+
* `ctx.settings.register(...)` and reads it live via the returned scope. On
|
|
29
|
+
* 0.2.x the settings service was rebuilt around a profile-entry model
|
|
30
|
+
* (`describe` / `mutate` / `configure`) and no longer exposes `register`; a
|
|
31
|
+
* settings edit restarts the plugin fiber and re-runs `apply(ctx, config)`
|
|
32
|
+
* with the new config, so the plugin caches that config in a module-level
|
|
33
|
+
* `liveConfig` and reads it there. `apply` detects which path is available at
|
|
34
|
+
* runtime, so the same code works on both.
|
|
35
|
+
*
|
|
36
|
+
* Zero conflict: the fields dsh already sends (temperature / maxTokens) are
|
|
37
|
+
* left to dsh unless a model explicitly sets them; every other sampling field
|
|
38
|
+
* is one dsh never sends, so there is no competing source.
|
|
39
|
+
*/
|
|
40
|
+
import z from "@deepseek-ai/schemastery";
|
|
41
|
+
|
|
42
|
+
const name = "sampling-params";
|
|
43
|
+
const inject = ["settings"];
|
|
44
|
+
|
|
45
|
+
// llama.cpp wire field names (snake_case). NOTE: repetition penalty is
|
|
46
|
+
// `repeat_penalty` on llama.cpp, NOT `repetition_penalty`.
|
|
47
|
+
const WIRE_KEYS = [
|
|
48
|
+
"temperature",
|
|
49
|
+
"top_p",
|
|
50
|
+
"top_k",
|
|
51
|
+
"min_p",
|
|
52
|
+
"repeat_penalty",
|
|
53
|
+
"presence_penalty",
|
|
54
|
+
"frequency_penalty",
|
|
55
|
+
];
|
|
56
|
+
|
|
57
|
+
// One model's full sampling set. Each field is required (a model sets every
|
|
58
|
+
// number explicitly); a model that omits a field is rejected by the schema, so
|
|
59
|
+
// configure each alias with a complete set.
|
|
60
|
+
const Model = z.object({
|
|
61
|
+
temperature: z.number(),
|
|
62
|
+
top_p: z.number(),
|
|
63
|
+
top_k: z.number(),
|
|
64
|
+
min_p: z.number(),
|
|
65
|
+
repeat_penalty: z.number(),
|
|
66
|
+
presence_penalty: z.number(),
|
|
67
|
+
frequency_penalty: z.number(),
|
|
68
|
+
});
|
|
69
|
+
|
|
70
|
+
// Plugin config: a table keyed by the exact model id (alias) sent on the
|
|
71
|
+
// wire. Each key's value is the sampling set stamped onto requests that name
|
|
72
|
+
// that alias. Example:
|
|
73
|
+
// models:
|
|
74
|
+
// model-i: { temperature: 0.2, top_p: 0.8, ... }
|
|
75
|
+
// model-p: { temperature: 1.0, ... }
|
|
76
|
+
// model-t: { temperature: 0.6, ... }
|
|
77
|
+
const Config = z.object({
|
|
78
|
+
models: z.dict(Model).default({}),
|
|
79
|
+
});
|
|
80
|
+
|
|
81
|
+
// Mark so we never double-wrap globalThis.fetch (hot-reload safe).
|
|
82
|
+
const WRAPPED = Symbol.for("dsh-llama-cpp-sampling-params.fetch-wrapped");
|
|
83
|
+
|
|
84
|
+
// The OpenAI-compatible chat-completions path we inject into. Only requests
|
|
85
|
+
// whose URL carries this path are touched; everything else passes through
|
|
86
|
+
// byte-identical.
|
|
87
|
+
const CHAT_COMPLETIONS_PATH = "/v1/chat/completions";
|
|
88
|
+
|
|
89
|
+
// Module-level live config, refreshed on every apply() call. Under DSH 0.2.x a
|
|
90
|
+
// settings edit restarts the plugin fiber and re-runs apply() with the new
|
|
91
|
+
// config, so this always reflects the latest `models` table. Under 0.1.x the
|
|
92
|
+
// registered settings scope (created inside apply below) is the live source
|
|
93
|
+
// instead and this only holds the initial value.
|
|
94
|
+
let liveConfig = { models: {} };
|
|
95
|
+
|
|
96
|
+
function apply(ctx, config = {}) {
|
|
97
|
+
const log = ctx.logger("sampling-params");
|
|
98
|
+
liveConfig = config;
|
|
99
|
+
|
|
100
|
+
// Register a live settings scope on DSH 0.1.x, where `ctx.settings.register`
|
|
101
|
+
// exists. On 0.2.x that method is gone (the settings service moved to a
|
|
102
|
+
// profile-entry model: describe / mutate / configure), so scope stays null
|
|
103
|
+
// and readModel below falls back to the config that 0.2.x re-passes to apply
|
|
104
|
+
// on every restart. The namespace is passed as a plain string (not via
|
|
105
|
+
// `settingsNamespace()`) so importing it never breaks plugin load.
|
|
106
|
+
let scope = null;
|
|
107
|
+
const settings = ctx.settings;
|
|
108
|
+
if (settings && typeof settings.register === "function") {
|
|
109
|
+
try {
|
|
110
|
+
scope = settings.register("sampling-params", Config, { base: config, applies: "live" });
|
|
111
|
+
} catch (error) {
|
|
112
|
+
log.warn(`sampling-params settings.register failed; using config param: ${error?.message ?? error}`);
|
|
113
|
+
scope = null;
|
|
114
|
+
}
|
|
115
|
+
}
|
|
116
|
+
|
|
117
|
+
// Read the sampling set for one exact model id/alias. Prefer the live scope
|
|
118
|
+
// (0.1.x); otherwise read the latest config passed to apply (0.2.x restart).
|
|
119
|
+
const readModel = (modelId) => {
|
|
120
|
+
if (scope) {
|
|
121
|
+
try {
|
|
122
|
+
const model = scope.get()?.models?.[modelId];
|
|
123
|
+
if (model) return model;
|
|
124
|
+
} catch {
|
|
125
|
+
// fall through to the config-param path
|
|
126
|
+
}
|
|
127
|
+
}
|
|
128
|
+
try {
|
|
129
|
+
return liveConfig?.models?.[modelId];
|
|
130
|
+
} catch {
|
|
131
|
+
return undefined;
|
|
132
|
+
}
|
|
133
|
+
};
|
|
134
|
+
|
|
135
|
+
const registeredModels = (() => {
|
|
136
|
+
try {
|
|
137
|
+
const models = scope ? (scope.get()?.models ?? {}) : (liveConfig?.models ?? {});
|
|
138
|
+
return Object.keys(models);
|
|
139
|
+
} catch {
|
|
140
|
+
return [];
|
|
141
|
+
}
|
|
142
|
+
})();
|
|
143
|
+
log.info(`sampling-params applied (fetch wrapper); models=${registeredModels.length}`);
|
|
144
|
+
|
|
145
|
+
if (typeof globalThis.fetch === "function" && !globalThis.fetch[WRAPPED]) {
|
|
146
|
+
const originalFetch = globalThis.fetch;
|
|
147
|
+
globalThis.fetch = async function (input, init) {
|
|
148
|
+
try {
|
|
149
|
+
const url = typeof input === "string" ? input : input?.url;
|
|
150
|
+
const bodyText = typeof init?.body === "string" ? init.body : null;
|
|
151
|
+
if (
|
|
152
|
+
typeof url === "string" &&
|
|
153
|
+
url.includes(CHAT_COMPLETIONS_PATH) &&
|
|
154
|
+
bodyText !== null
|
|
155
|
+
) {
|
|
156
|
+
const body = JSON.parse(bodyText);
|
|
157
|
+
if (body && typeof body.model === "string") {
|
|
158
|
+
const model = readModel(body.model);
|
|
159
|
+
if (model) {
|
|
160
|
+
for (const key of WIRE_KEYS) {
|
|
161
|
+
if (typeof model[key] === "number") body[key] = model[key];
|
|
162
|
+
}
|
|
163
|
+
init.body = JSON.stringify(body);
|
|
164
|
+
}
|
|
165
|
+
}
|
|
166
|
+
}
|
|
167
|
+
} catch (error) {
|
|
168
|
+
// Never break LLM traffic: on any parse/matching error, send the
|
|
169
|
+
// request through untouched.
|
|
170
|
+
log.warn(`sampling-params inject skipped: ${error?.message ?? error}`);
|
|
171
|
+
}
|
|
172
|
+
return originalFetch.call(this, input, init);
|
|
173
|
+
};
|
|
174
|
+
globalThis.fetch[WRAPPED] = true;
|
|
175
|
+
}
|
|
176
|
+
}
|
|
177
|
+
|
|
178
|
+
export { apply, name, inject, Config };
|
package/package.json
CHANGED
|
@@ -1,47 +1,47 @@
|
|
|
1
|
-
{
|
|
2
|
-
"name": "dsh-llama-cpp-sampling-params",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to a local llama.cpp / llama-server gateway.",
|
|
5
|
-
"author": "devhang",
|
|
6
|
-
"repository": {
|
|
7
|
-
"type": "git",
|
|
8
|
-
"url": "git+https://github.com/devhang/dsh-llama-cpp-sampling-params.git"
|
|
9
|
-
},
|
|
10
|
-
"homepage": "https://github.com/devhang/dsh-llama-cpp-sampling-params#readme",
|
|
11
|
-
"bugs": {
|
|
12
|
-
"url": "https://github.com/devhang/dsh-llama-cpp-sampling-params/issues"
|
|
13
|
-
},
|
|
14
|
-
"type": "module",
|
|
15
|
-
"main": "index.js",
|
|
16
|
-
"files": [
|
|
17
|
-
"index.js",
|
|
18
|
-
"cordis.patch.yml",
|
|
19
|
-
"README.md",
|
|
20
|
-
"README.zh.md",
|
|
21
|
-
"LICENSE"
|
|
22
|
-
],
|
|
23
|
-
"keywords": [
|
|
24
|
-
"deepseek-harness",
|
|
25
|
-
"dsh",
|
|
26
|
-
"dsh-plugin",
|
|
27
|
-
"llama.cpp",
|
|
28
|
-
"llama-server",
|
|
29
|
-
"sampling",
|
|
30
|
-
"top_p",
|
|
31
|
-
"min_p",
|
|
32
|
-
"repeat_penalty"
|
|
33
|
-
],
|
|
34
|
-
"license": "MIT",
|
|
35
|
-
"dsh": {
|
|
36
|
-
"bundle": {
|
|
37
|
-
"patch": "./cordis.patch.yml"
|
|
38
|
-
}
|
|
39
|
-
},
|
|
40
|
-
"peerDependencies": {
|
|
41
|
-
"@deepseek-ai/cordis": "^4.0.1",
|
|
42
|
-
"@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.
|
|
43
|
-
},
|
|
44
|
-
"dependencies": {
|
|
45
|
-
"@deepseek-ai/schemastery": "^3.18.2"
|
|
46
|
-
}
|
|
47
|
-
}
|
|
1
|
+
{
|
|
2
|
+
"name": "dsh-llama-cpp-sampling-params",
|
|
3
|
+
"version": "0.2.0",
|
|
4
|
+
"description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to a local llama.cpp / llama-server gateway.",
|
|
5
|
+
"author": "devhang",
|
|
6
|
+
"repository": {
|
|
7
|
+
"type": "git",
|
|
8
|
+
"url": "git+https://github.com/devhang/dsh-llama-cpp-sampling-params.git"
|
|
9
|
+
},
|
|
10
|
+
"homepage": "https://github.com/devhang/dsh-llama-cpp-sampling-params#readme",
|
|
11
|
+
"bugs": {
|
|
12
|
+
"url": "https://github.com/devhang/dsh-llama-cpp-sampling-params/issues"
|
|
13
|
+
},
|
|
14
|
+
"type": "module",
|
|
15
|
+
"main": "index.js",
|
|
16
|
+
"files": [
|
|
17
|
+
"index.js",
|
|
18
|
+
"cordis.patch.yml",
|
|
19
|
+
"README.md",
|
|
20
|
+
"README.zh.md",
|
|
21
|
+
"LICENSE"
|
|
22
|
+
],
|
|
23
|
+
"keywords": [
|
|
24
|
+
"deepseek-harness",
|
|
25
|
+
"dsh",
|
|
26
|
+
"dsh-plugin",
|
|
27
|
+
"llama.cpp",
|
|
28
|
+
"llama-server",
|
|
29
|
+
"sampling",
|
|
30
|
+
"top_p",
|
|
31
|
+
"min_p",
|
|
32
|
+
"repeat_penalty"
|
|
33
|
+
],
|
|
34
|
+
"license": "MIT",
|
|
35
|
+
"dsh": {
|
|
36
|
+
"bundle": {
|
|
37
|
+
"patch": "./cordis.patch.yml"
|
|
38
|
+
}
|
|
39
|
+
},
|
|
40
|
+
"peerDependencies": {
|
|
41
|
+
"@deepseek-ai/cordis": "^4.0.1",
|
|
42
|
+
"@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.3.0-0"
|
|
43
|
+
},
|
|
44
|
+
"dependencies": {
|
|
45
|
+
"@deepseek-ai/schemastery": "^3.18.2"
|
|
46
|
+
}
|
|
47
|
+
}
|