dsh-llama-cpp-sampling-params 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (4) hide show
  1. package/README.md +36 -28
  2. package/README.zh.md +36 -28
  3. package/index.js +178 -140
  4. package/package.json +47 -47
package/README.md CHANGED
@@ -43,38 +43,45 @@ Then restart `dsh web` (or refresh the GUI page).
43
43
 
44
44
  ## Configure models
45
45
 
46
- Set the per-model sampling table in `$DSH_HOME/settings.yaml`. The section applies **live** — edits take effect without a restart.
46
+ The `models` table lives in your dsh settings. **Where it goes depends on your dsh version:**
47
+
48
+ | dsh | where the table goes |
49
+ |-----|----------------------|
50
+ | **0.2.x** | the profile's `cordis.patch.yml`, as a top-level `sampling-params` entry |
51
+ | **0.1.x** | `$DSH_HOME/settings.yaml`, under a `sampling-params:` section (applies live) |
52
+
53
+ ### DSH 0.2.x
54
+
55
+ Add a top-level entry to your profile's `cordis.patch.yml` (edit `cordis.patch.yml`, not `cordis.yml`). The `id: sampling-params` is the binding key and must match the plugin's namespace:
47
56
 
48
57
  ```yaml
49
- sampling-params:
50
- models:
51
- # Key = the exact model id (alias) sent on the wire.
52
- model-t: # Think
53
- temperature: 0.6
54
- top_p: 0.95
55
- top_k: 20
56
- min_p: 0.05
57
- repeat_penalty: 1.0
58
- presence_penalty: 0.0
59
- frequency_penalty: 0.0
60
- model-i: # Instruct
61
- temperature: 0.2
62
- top_p: 0.8
63
- top_k: 20
64
- min_p: 0.05
65
- repeat_penalty: 1.0
66
- presence_penalty: 0.0
67
- frequency_penalty: 0.0
68
- model-p: # Planner
69
- temperature: 1.0
70
- top_p: 0.95
71
- top_k: 20
72
- min_p: 0.0
73
- repeat_penalty: 1.0
74
- presence_penalty: 0.0
75
- frequency_penalty: 0.0
58
+ - id: sampling-params
59
+ name: dsh-llama-cpp-sampling-params
60
+ config:
61
+ models:
62
+ # Key = the exact model id (alias) sent on the wire.
63
+ model-t: # Think
64
+ temperature: 0.6
65
+ top_p: 0.95
66
+ top_k: 20
67
+ min_p: 0.05
68
+ repeat_penalty: 1.0
69
+ presence_penalty: 0.0
70
+ frequency_penalty: 0.0
71
+ model-i: # Instruct
72
+ temperature: 0.2
73
+ top_p: 0.8
74
+ top_k: 20
75
+ min_p: 0.05
76
+ repeat_penalty: 1.0
77
+ presence_penalty: 0.0
78
+ frequency_penalty: 0.0
76
79
  ```
77
80
 
81
+ ### Migrating from 0.1.x to 0.2.x
82
+
83
+ Upgrading dsh to 0.2.x renames `$DSH_HOME/settings.yaml` to `settings.yaml.imported` (a frozen backup, no longer read). Your `sampling-params.models` table is stranded there. To migrate, copy the `models:` block out of `settings.yaml.imported` and add it as the `sampling-params` entry in your profile's `cordis.patch.yml`, re-indenting the block by **+2 spaces** to nest under `config:`. The one-time auto-import does not re-run, so do this by hand (or via the settings UI).
84
+
78
85
  > Each model is a **complete** sampling set — every field is required, so configure each alias with all seven numbers.
79
86
 
80
87
  ## Ports
@@ -107,6 +114,7 @@ The `params` object of the returned slot reflects `temperature`, `top_p`, `top_k
107
114
 
108
115
  ## Compatibility
109
116
 
117
+ - Supports dsh **0.1.x and 0.2.x** — the settings read path is version-adaptive (`settings.register` on 0.1.x, the profile-entry config on 0.2.x).
110
118
  - Requires dsh with `@deepseek-ai/dsh-settings` and `@deepseek-ai/schemastery` (both ship with dsh).
111
119
  - Target must be an OpenAI-compatible gateway that honors these wire fields (llama.cpp `llama-server` does).
112
120
  - The wrapper is guarded by a `Symbol.for` flag so hot-reload never double-wraps.
package/README.zh.md CHANGED
@@ -43,38 +43,45 @@ dsh plugin --profile web add link:C:/path/to/dsh-llama-cpp-sampling-params
43
43
 
44
44
  ## 配置模型
45
45
 
46
- 在 `$DSH_HOME/settings.yaml` 設定每個模型的採樣表。該區段**即時生效**——修改無需重啟。
46
+ `models` 表放在你的 dsh 設定裡。**放哪裡取決於你的 dsh 版本:**
47
+
48
+ | dsh | 表放哪裡 |
49
+ |-----|---------|
50
+ | **0.2.x** | profile 的 `cordis.patch.yml`,作為頂層 `sampling-params` entry |
51
+ | **0.1.x** | `$DSH_HOME/settings.yaml` 的 `sampling-params:` 區段(即時生效) |
52
+
53
+ ### DSH 0.2.x
54
+
55
+ 在 profile 的 `cordis.patch.yml` 加一個頂層 entry(改 `cordis.patch.yml`,不是 `cordis.yml`)。`id: sampling-params` 是 binding key,必須與插件 namespace 相符:
47
56
 
48
57
  ```yaml
49
- sampling-params:
50
- models:
51
- # 鍵 = 發送到 wire 的精確 model id(別名)。
52
- model-t: # Think
53
- temperature: 0.6
54
- top_p: 0.95
55
- top_k: 20
56
- min_p: 0.05
57
- repeat_penalty: 1.0
58
- presence_penalty: 0.0
59
- frequency_penalty: 0.0
60
- model-i: # Instruct
61
- temperature: 0.2
62
- top_p: 0.8
63
- top_k: 20
64
- min_p: 0.05
65
- repeat_penalty: 1.0
66
- presence_penalty: 0.0
67
- frequency_penalty: 0.0
68
- model-p: # Planner
69
- temperature: 1.0
70
- top_p: 0.95
71
- top_k: 20
72
- min_p: 0.0
73
- repeat_penalty: 1.0
74
- presence_penalty: 0.0
75
- frequency_penalty: 0.0
58
+ - id: sampling-params
59
+ name: dsh-llama-cpp-sampling-params
60
+ config:
61
+ models:
62
+ # 鍵 = 發送到 wire 的精確 model id(別名)。
63
+ model-t: # Think
64
+ temperature: 0.6
65
+ top_p: 0.95
66
+ top_k: 20
67
+ min_p: 0.05
68
+ repeat_penalty: 1.0
69
+ presence_penalty: 0.0
70
+ frequency_penalty: 0.0
71
+ model-i: # Instruct
72
+ temperature: 0.2
73
+ top_p: 0.8
74
+ top_k: 20
75
+ min_p: 0.05
76
+ repeat_penalty: 1.0
77
+ presence_penalty: 0.0
78
+ frequency_penalty: 0.0
76
79
  ```
77
80
 
81
+ ### 從 0.1.x 遷移到 0.2.x
82
+
83
+ 升級 dsh 到 0.2.x 會把 `$DSH_HOME/settings.yaml` 改名成 `settings.yaml.imported`(凍結備份,不再讀取),你的 `sampling-params.models` 表就擱淺在那裡。要遷移:把 `models:` 區塊從 `settings.yaml.imported` 抽出,加進 profile 的 `cordis.patch.yml` 作為 `sampling-params` entry,並將該區塊**+2 空格**縮排以嵌在 `config:` 下。一次性自動 import 不會重跑,所以要手動搬(或經設定 UI)。
84
+
78
85
  > 每個模型是**完整**採樣組——每個欄位必填,因此為每個別名配置全部七個數值。
79
86
 
80
87
  ## 連接埠
@@ -107,6 +114,7 @@ curl "http://<host>:<port>/slots?model=<model-id>"
107
114
 
108
115
  ## 相容性
109
116
 
117
+ - 支援 dsh **0.1.x 與 0.2.x**——設定讀取路徑是 version-adaptive(0.1.x 用 `settings.register`、0.2.x 用 profile-entry config)。
110
118
  - 需要 dsh 內含 `@deepseek-ai/dsh-settings` 與 `@deepseek-ai/schemastery`(兩者隨 dsh 提供)。
111
119
  - 目標必須是遵循這些 wire 欄位的 OpenAI 相容閘道(llama.cpp `llama-server` 符合)。
112
120
  - Wrapper 以 `Symbol.for` 標記保護,hot-reload 不會重複 wrap。
package/index.js CHANGED
@@ -1,140 +1,178 @@
1
- /**
2
- * dsh-llama-cpp-sampling-params
3
- *
4
- * Injects per-model-alias sampling parameters into every chat-completions
5
- * request sent to an OpenAI-compatible local gateway (llama.cpp).
6
- *
7
- * How models work:
8
- * - llama.cpp serves ONE loaded model under several aliases (e.g.
9
- * `model-i`, `model-p`, `model-t`) that all point at the same underlying
10
- * GGUF — so there is NO extra VRAM cost.
11
- * - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
12
- * switching model in the UI picks a sampling role.
13
- * - This plugin reads the `model` field of each request (the alias) and
14
- * looks it up in the `sampling-params` `models` table. If the alias is
15
- * configured, it stamps that model's sampling set onto the wire body.
16
- * Unconfigured aliases pass through byte-identical.
17
- *
18
- * Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
19
- * maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
20
- * presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
21
- * deep-frozen (mutation throws) and the pi-ai adapter does not forward
22
- * `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
23
- * the transport layer, below all of those abstractions, and simply rewrites
24
- * the already-built JSON body before it leaves the process.
25
- *
26
- * Zero conflict: the fields dsh already sends (temperature / maxTokens) are
27
- * left to dsh unless a model explicitly sets them; every other sampling field
28
- * is one dsh never sends, so there is no competing source.
29
- */
30
- import z from "@deepseek-ai/schemastery";
31
-
32
- const name = "sampling-params";
33
- const inject = ["settings"];
34
-
35
- // llama.cpp wire field names (snake_case). NOTE: repetition penalty is
36
- // `repeat_penalty` on llama.cpp, NOT `repetition_penalty`.
37
- const WIRE_KEYS = [
38
- "temperature",
39
- "top_p",
40
- "top_k",
41
- "min_p",
42
- "repeat_penalty",
43
- "presence_penalty",
44
- "frequency_penalty",
45
- ];
46
-
47
- // One model's full sampling set. Each field is required (a model sets every
48
- // number explicitly); a model that omits a field is rejected by the schema, so
49
- // configure each alias with a complete set.
50
- const Model = z.object({
51
- temperature: z.number(),
52
- top_p: z.number(),
53
- top_k: z.number(),
54
- min_p: z.number(),
55
- repeat_penalty: z.number(),
56
- presence_penalty: z.number(),
57
- frequency_penalty: z.number(),
58
- });
59
-
60
- // Plugin config: a table keyed by the exact model id (alias) sent on the
61
- // wire. Each key's value is the sampling set stamped onto requests that name
62
- // that alias. Example:
63
- // models:
64
- // model-i: { temperature: 0.2, top_p: 0.8, ... }
65
- // model-p: { temperature: 1.0, ... }
66
- // model-t: { temperature: 0.6, ... }
67
- const Config = z.object({
68
- models: z.dict(Model).default({}),
69
- });
70
-
71
- // Mark so we never double-wrap globalThis.fetch (hot-reload safe).
72
- const WRAPPED = Symbol.for("dsh-llama-cpp-sampling-params.fetch-wrapped");
73
-
74
- // The OpenAI-compatible chat-completions path we inject into. Only requests
75
- // whose URL carries this path are touched; everything else passes through
76
- // byte-identical.
77
- const CHAT_COMPLETIONS_PATH = "/v1/chat/completions";
78
-
79
- function apply(ctx, config = {}) {
80
- const log = ctx.logger("sampling-params");
81
-
82
- // Register the settings section live so role values can be edited without
83
- // a restart. NOTE: we pass the namespace as a plain string, not via
84
- // `settingsNamespace()` from @deepseek-ai/dsh-settings — the export was
85
- // added in 0.1.1 and is absent in some older/newer versions, so importing
86
- // it directly breaks plugin load on those dsh builds.
87
- const scope = ctx.settings.register("sampling-params", Config, {
88
- base: config,
89
- applies: "live",
90
- });
91
-
92
- // Read the sampling set for one exact model id/alias.
93
- const readModel = (modelId) => {
94
- try {
95
- const section = scope.get();
96
- return section?.models?.[modelId];
97
- } catch {
98
- return undefined;
99
- }
100
- };
101
-
102
- const registeredModels = (() => {
103
- try { return Object.keys(scope.get()?.models ?? {}); } catch { return []; }
104
- })();
105
- log.info(`sampling-params applied (fetch wrapper); models=${registeredModels.length}`);
106
-
107
- if (typeof globalThis.fetch === "function" && !globalThis.fetch[WRAPPED]) {
108
- const originalFetch = globalThis.fetch;
109
- globalThis.fetch = async function (input, init) {
110
- try {
111
- const url = typeof input === "string" ? input : input?.url;
112
- const bodyText = typeof init?.body === "string" ? init.body : null;
113
- if (
114
- typeof url === "string" &&
115
- url.includes(CHAT_COMPLETIONS_PATH) &&
116
- bodyText !== null
117
- ) {
118
- const body = JSON.parse(bodyText);
119
- if (body && typeof body.model === "string") {
120
- const model = readModel(body.model);
121
- if (model) {
122
- for (const key of WIRE_KEYS) {
123
- if (typeof model[key] === "number") body[key] = model[key];
124
- }
125
- init.body = JSON.stringify(body);
126
- }
127
- }
128
- }
129
- } catch (error) {
130
- // Never break LLM traffic: on any parse/matching error, send the
131
- // request through untouched.
132
- log.warn(`sampling-params inject skipped: ${error?.message ?? error}`);
133
- }
134
- return originalFetch.call(this, input, init);
135
- };
136
- globalThis.fetch[WRAPPED] = true;
137
- }
138
- }
139
-
140
- export { apply, name, inject, Config };
1
+ /**
2
+ * dsh-llama-cpp-sampling-params
3
+ *
4
+ * Injects per-model-alias sampling parameters into every chat-completions
5
+ * request sent to an OpenAI-compatible local gateway (llama.cpp).
6
+ *
7
+ * How models work:
8
+ * - llama.cpp serves ONE loaded model under several aliases (e.g.
9
+ * `model-i`, `model-p`, `model-t`) that all point at the same underlying
10
+ * GGUF — so there is NO extra VRAM cost.
11
+ * - dsh's `llm-pi-ai` lists each alias as a separate selectable model, so
12
+ * switching model in the UI picks a sampling role.
13
+ * - This plugin reads the `model` field of each request (the alias) and
14
+ * looks it up in the `sampling-params` `models` table. If the alias is
15
+ * configured, it stamps that model's sampling set onto the wire body.
16
+ * Unconfigured aliases pass through byte-identical.
17
+ *
18
+ * Why a fetch wrapper: dsh's `GenerateOptions` only carries temperature /
19
+ * maxTokens / stop — it never sends top_p, top_k, min_p, repeat_penalty,
20
+ * presence_penalty, or frequency_penalty. Moreover, agent-loop requests arrive
21
+ * deep-frozen (mutation throws) and the pi-ai adapter does not forward
22
+ * `onPayload` from GenerateOptions. Wrapping `globalThis.fetch` operates at
23
+ * the transport layer, below all of those abstractions, and simply rewrites
24
+ * the already-built JSON body before it leaves the process.
25
+ *
26
+ * Config source across DSH versions: the `models` table is the plugin's
27
+ * settings. On DSH 0.1.x the plugin registers that section with
28
+ * `ctx.settings.register(...)` and reads it live via the returned scope. On
29
+ * 0.2.x the settings service was rebuilt around a profile-entry model
30
+ * (`describe` / `mutate` / `configure`) and no longer exposes `register`; a
31
+ * settings edit restarts the plugin fiber and re-runs `apply(ctx, config)`
32
+ * with the new config, so the plugin caches that config in a module-level
33
+ * `liveConfig` and reads it there. `apply` detects which path is available at
34
+ * runtime, so the same code works on both.
35
+ *
36
+ * Zero conflict: the fields dsh already sends (temperature / maxTokens) are
37
+ * left to dsh unless a model explicitly sets them; every other sampling field
38
+ * is one dsh never sends, so there is no competing source.
39
+ */
40
+ import z from "@deepseek-ai/schemastery";
41
+
42
+ const name = "sampling-params";
43
+ const inject = ["settings"];
44
+
45
+ // llama.cpp wire field names (snake_case). NOTE: repetition penalty is
46
+ // `repeat_penalty` on llama.cpp, NOT `repetition_penalty`.
47
+ const WIRE_KEYS = [
48
+ "temperature",
49
+ "top_p",
50
+ "top_k",
51
+ "min_p",
52
+ "repeat_penalty",
53
+ "presence_penalty",
54
+ "frequency_penalty",
55
+ ];
56
+
57
+ // One model's full sampling set. Each field is required (a model sets every
58
+ // number explicitly); a model that omits a field is rejected by the schema, so
59
+ // configure each alias with a complete set.
60
+ const Model = z.object({
61
+ temperature: z.number(),
62
+ top_p: z.number(),
63
+ top_k: z.number(),
64
+ min_p: z.number(),
65
+ repeat_penalty: z.number(),
66
+ presence_penalty: z.number(),
67
+ frequency_penalty: z.number(),
68
+ });
69
+
70
+ // Plugin config: a table keyed by the exact model id (alias) sent on the
71
+ // wire. Each key's value is the sampling set stamped onto requests that name
72
+ // that alias. Example:
73
+ // models:
74
+ // model-i: { temperature: 0.2, top_p: 0.8, ... }
75
+ // model-p: { temperature: 1.0, ... }
76
+ // model-t: { temperature: 0.6, ... }
77
+ const Config = z.object({
78
+ models: z.dict(Model).default({}),
79
+ });
80
+
81
+ // Mark so we never double-wrap globalThis.fetch (hot-reload safe).
82
+ const WRAPPED = Symbol.for("dsh-llama-cpp-sampling-params.fetch-wrapped");
83
+
84
+ // The OpenAI-compatible chat-completions path we inject into. Only requests
85
+ // whose URL carries this path are touched; everything else passes through
86
+ // byte-identical.
87
+ const CHAT_COMPLETIONS_PATH = "/v1/chat/completions";
88
+
89
+ // Module-level live config, refreshed on every apply() call. Under DSH 0.2.x a
90
+ // settings edit restarts the plugin fiber and re-runs apply() with the new
91
+ // config, so this always reflects the latest `models` table. Under 0.1.x the
92
+ // registered settings scope (created inside apply below) is the live source
93
+ // instead and this only holds the initial value.
94
+ let liveConfig = { models: {} };
95
+
96
+ function apply(ctx, config = {}) {
97
+ const log = ctx.logger("sampling-params");
98
+ liveConfig = config;
99
+
100
+ // Register a live settings scope on DSH 0.1.x, where `ctx.settings.register`
101
+ // exists. On 0.2.x that method is gone (the settings service moved to a
102
+ // profile-entry model: describe / mutate / configure), so scope stays null
103
+ // and readModel below falls back to the config that 0.2.x re-passes to apply
104
+ // on every restart. The namespace is passed as a plain string (not via
105
+ // `settingsNamespace()`) so importing it never breaks plugin load.
106
+ let scope = null;
107
+ const settings = ctx.settings;
108
+ if (settings && typeof settings.register === "function") {
109
+ try {
110
+ scope = settings.register("sampling-params", Config, { base: config, applies: "live" });
111
+ } catch (error) {
112
+ log.warn(`sampling-params settings.register failed; using config param: ${error?.message ?? error}`);
113
+ scope = null;
114
+ }
115
+ }
116
+
117
+ // Read the sampling set for one exact model id/alias. Prefer the live scope
118
+ // (0.1.x); otherwise read the latest config passed to apply (0.2.x restart).
119
+ const readModel = (modelId) => {
120
+ if (scope) {
121
+ try {
122
+ const model = scope.get()?.models?.[modelId];
123
+ if (model) return model;
124
+ } catch {
125
+ // fall through to the config-param path
126
+ }
127
+ }
128
+ try {
129
+ return liveConfig?.models?.[modelId];
130
+ } catch {
131
+ return undefined;
132
+ }
133
+ };
134
+
135
+ const registeredModels = (() => {
136
+ try {
137
+ const models = scope ? (scope.get()?.models ?? {}) : (liveConfig?.models ?? {});
138
+ return Object.keys(models);
139
+ } catch {
140
+ return [];
141
+ }
142
+ })();
143
+ log.info(`sampling-params applied (fetch wrapper); models=${registeredModels.length}`);
144
+
145
+ if (typeof globalThis.fetch === "function" && !globalThis.fetch[WRAPPED]) {
146
+ const originalFetch = globalThis.fetch;
147
+ globalThis.fetch = async function (input, init) {
148
+ try {
149
+ const url = typeof input === "string" ? input : input?.url;
150
+ const bodyText = typeof init?.body === "string" ? init.body : null;
151
+ if (
152
+ typeof url === "string" &&
153
+ url.includes(CHAT_COMPLETIONS_PATH) &&
154
+ bodyText !== null
155
+ ) {
156
+ const body = JSON.parse(bodyText);
157
+ if (body && typeof body.model === "string") {
158
+ const model = readModel(body.model);
159
+ if (model) {
160
+ for (const key of WIRE_KEYS) {
161
+ if (typeof model[key] === "number") body[key] = model[key];
162
+ }
163
+ init.body = JSON.stringify(body);
164
+ }
165
+ }
166
+ }
167
+ } catch (error) {
168
+ // Never break LLM traffic: on any parse/matching error, send the
169
+ // request through untouched.
170
+ log.warn(`sampling-params inject skipped: ${error?.message ?? error}`);
171
+ }
172
+ return originalFetch.call(this, input, init);
173
+ };
174
+ globalThis.fetch[WRAPPED] = true;
175
+ }
176
+ }
177
+
178
+ export { apply, name, inject, Config };
package/package.json CHANGED
@@ -1,47 +1,47 @@
1
- {
2
- "name": "dsh-llama-cpp-sampling-params",
3
- "version": "0.1.1",
4
- "description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to a local llama.cpp / llama-server gateway.",
5
- "author": "devhang",
6
- "repository": {
7
- "type": "git",
8
- "url": "git+https://github.com/devhang/dsh-llama-cpp-sampling-params.git"
9
- },
10
- "homepage": "https://github.com/devhang/dsh-llama-cpp-sampling-params#readme",
11
- "bugs": {
12
- "url": "https://github.com/devhang/dsh-llama-cpp-sampling-params/issues"
13
- },
14
- "type": "module",
15
- "main": "index.js",
16
- "files": [
17
- "index.js",
18
- "cordis.patch.yml",
19
- "README.md",
20
- "README.zh.md",
21
- "LICENSE"
22
- ],
23
- "keywords": [
24
- "deepseek-harness",
25
- "dsh",
26
- "dsh-plugin",
27
- "llama.cpp",
28
- "llama-server",
29
- "sampling",
30
- "top_p",
31
- "min_p",
32
- "repeat_penalty"
33
- ],
34
- "license": "MIT",
35
- "dsh": {
36
- "bundle": {
37
- "patch": "./cordis.patch.yml"
38
- }
39
- },
40
- "peerDependencies": {
41
- "@deepseek-ai/cordis": "^4.0.1",
42
- "@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.2.0-0"
43
- },
44
- "dependencies": {
45
- "@deepseek-ai/schemastery": "^3.18.2"
46
- }
47
- }
1
+ {
2
+ "name": "dsh-llama-cpp-sampling-params",
3
+ "version": "0.2.0",
4
+ "description": "DSH plugin: inject per-model sampling parameters (top_p, top_k, min_p, repeat_penalty, presence_penalty, frequency_penalty) into every chat-completions request sent to a local llama.cpp / llama-server gateway.",
5
+ "author": "devhang",
6
+ "repository": {
7
+ "type": "git",
8
+ "url": "git+https://github.com/devhang/dsh-llama-cpp-sampling-params.git"
9
+ },
10
+ "homepage": "https://github.com/devhang/dsh-llama-cpp-sampling-params#readme",
11
+ "bugs": {
12
+ "url": "https://github.com/devhang/dsh-llama-cpp-sampling-params/issues"
13
+ },
14
+ "type": "module",
15
+ "main": "index.js",
16
+ "files": [
17
+ "index.js",
18
+ "cordis.patch.yml",
19
+ "README.md",
20
+ "README.zh.md",
21
+ "LICENSE"
22
+ ],
23
+ "keywords": [
24
+ "deepseek-harness",
25
+ "dsh",
26
+ "dsh-plugin",
27
+ "llama.cpp",
28
+ "llama-server",
29
+ "sampling",
30
+ "top_p",
31
+ "min_p",
32
+ "repeat_penalty"
33
+ ],
34
+ "license": "MIT",
35
+ "dsh": {
36
+ "bundle": {
37
+ "patch": "./cordis.patch.yml"
38
+ }
39
+ },
40
+ "peerDependencies": {
41
+ "@deepseek-ai/cordis": "^4.0.1",
42
+ "@deepseek-ai/dsh-settings": ">=0.1.1-rc.1 <0.3.0-0"
43
+ },
44
+ "dependencies": {
45
+ "@deepseek-ai/schemastery": "^3.18.2"
46
+ }
47
+ }