pi-llama-cpp 0.11.0 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -32,6 +32,16 @@ A [Pi Coding Agent](https://pi.dev/) extension that integrates with running [lla
32
32
 
33
33
  > **Note:** You can run your server with API authentication with `llama-server --api-key <your key> ...`.
34
34
 
35
+ ### Server Health Indicators
36
+
37
+ When browsing servers via `/models servers`, each server URL is prefixed with a health indicator:
38
+
39
+ | Icon | Status | Description |
40
+ | ---- | ----------- | --------------------------------------------- |
41
+ | 🟢 | Healthy | Server responded successfully to health check |
42
+ | 🟡 | Timeout | Server health check timed out |
43
+ | 🔴 | Unreachable | Server could not be reached |
44
+
35
45
  ## Installation
36
46
 
37
47
  This package is a Pi extension. Install it with
@@ -61,6 +71,18 @@ The recommended way to configure the extension is using the `llamaSettings` key.
61
71
 
62
72
  Add this to your `.pi/settings.json` (project) or `~/.pi/agent/settings.json` (global):
63
73
 
74
+ #### Minimal configuration
75
+
76
+ ```json
77
+ {
78
+ "llamaSettings": {
79
+ "servers": [{ "url": "http://127.0.0.1:8080" }]
80
+ }
81
+ }
82
+ ```
83
+
84
+ #### Full configuration
85
+
64
86
  ```json
65
87
  {
66
88
  "llamaSettings": {
@@ -68,7 +90,8 @@ Add this to your `.pi/settings.json` (project) or `~/.pi/agent/settings.json` (g
68
90
  {
69
91
  "url": "http://127.0.0.1:8080",
70
92
  "id": "local",
71
- "name": "Local Server"
93
+ "name": "Local Server",
94
+ "overrides": {}
72
95
  },
73
96
  {
74
97
  "url": "http://10.0.0.5:8080",
@@ -92,7 +115,7 @@ With this config, the servers will appear in Pi as **Llama.cpp (Local Server)**
92
115
  | ------ | ------ | -------- | ---------------------------------------------------------------------------- |
93
116
  | `url` | string | Yes | The URL of the llama.cpp server |
94
117
  | `id` | string | No | Custom provider ID (used for API key auth). Defaults to `llama-server=<url>` |
95
- | `name` | string | No | Display name for the server in the UI (shown as `Llama.cpp — <name>`) |
118
+ | `name` | string | No | Display name for the server in the UI (shown as `Llama.cpp (<name>)`) |
96
119
 
97
120
  > **Note:** If you set a custom `id`, you can use it in `~/.pi/agent/auth.json`. The extension will also fall back to the URL-based ID if no key is found for the custom `id`.
98
121
 
@@ -110,48 +133,25 @@ With this config, the servers will appear in Pi as **Llama.cpp (Local Server)**
110
133
 
111
134
  #### In-session settings menu
112
135
 
113
- Run `/models settings` to edit the scalar settings above without hand-editing JSON:
114
-
115
- - **Enter/Space** cycles the value under the cursor; **Esc** closes the menu.
116
- - Booleans toggle `on`/`off`, `sortBy` cycles through the sort orders, and the
117
- timeouts cycle through presets (`pollingTimeout`: 15s/30s/60s/120s/300s,
118
- `serverTimeout`: 500ms/1s/2s/5s/10s).
119
- - Changes are written to the **global** `~/.pi/agent/settings.json` only. If a
120
- project `.pi/settings.json` defines the same key, its value keeps winning in
121
- the merged view until you remove it there.
122
- - Boolean and sort changes apply immediately; timeout changes apply on the next
123
- model load.
124
- - The `servers` list is edited with `/models servers` (see below).
136
+ Run `/models settings` to edit the scalar settings above without hand-editing JSON. Changes are written to the **project** `.pi/settings.json` if it exists, otherwise to **global** `~/.pi/agent/settings.json`. Boolean and sort changes apply immediately; timeout changes apply on the next model load. The `servers` list is edited with `/models servers` (see below), and per-server model overrides with `/models overrides` (see [Model Overrides](#model-overrides)).
125
137
 
126
138
  #### Server list editor
127
139
 
128
140
  Run `/models servers` to add, edit or remove entries of `llamaSettings.servers`
129
- without hand-editing JSON:
130
-
131
- - **↑/↓** moves the cursor, **Enter/e** edits the selected URL, **i** edits
132
- its `id`, **n** its `name`, **a** adds a new entry, **d** deletes it
133
- (after an "Are you sure?" confirmation — only **y** confirms;
134
- **Enter** is ignored, **Esc/n** cancels), **Esc** closes the editor.
135
- - One URL per entry (`http://host:port`). Trailing slashes are stripped on
136
- save; `;`-separated values are rejected — use separate entries instead.
137
- - Each change is written immediately to the **global**
138
- `~/.pi/agent/settings.json`. If a project `.pi/settings.json` defines
139
- `servers`, its list keeps winning in the merged view until you remove it
140
- there.
141
- - Changes apply the next time providers are scanned — run `/models` to see
142
- them. Additions, removals, and URL/`id`/`name` edits all take effect on
143
- the next `/models`: new servers register their providers, removed ones
144
- leave pi's registry immediately, and edited ones are re-registered with
145
- the fresh config — no restart needed.
146
- - Limitation: a model already loading in the background on a removed or
147
- edited server finishes loading, but its progress notifications stop;
148
- re-select it from the (new) provider afterwards.
149
- - The editor shows a warning when the `LLAMA_SERVER_URL` environment variable
150
- is set, since it overrides the configured servers.
151
- - Per-server `id`/`name` overrides can be edited with **i**/**n**; saving an
152
- empty value clears the override. The list shows them as a
153
- `(<id> - <name>)` suffix, falling back to the auto-detected
154
- `llama-server=<url>` id when no custom `id` is set.
141
+ without hand-editing JSON. Each server URL is prefixed with a health indicator
142
+ (🟢 healthy, 🟡 timeout, 🔴 unreachable) that reflects the result of a
143
+ health check against the server. Each change is written immediately to the **project**
144
+ `.pi/settings.json` if it exists, otherwise to **global**
145
+ `~/.pi/agent/settings.json`.
146
+
147
+ Changes take effect immediately after closing the editor: new servers
148
+ register their providers, removed ones leave pi's registry right away,
149
+ and edited ones are re-registered with the fresh config — no restart or
150
+ `/models` needed.
151
+
152
+ Limitation: a model already loading in the background on a removed or
153
+ edited server finishes loading, but its progress notifications stop;
154
+ re-select it from the (new) provider afterwards.
155
155
 
156
156
  #### Environment variable
157
157
 
@@ -247,6 +247,7 @@ llama-server --model path/to/model.gguf ...
247
247
 
248
248
  The extension determines the context size as follows:
249
249
 
250
+ - A per-model `contextSize` override (see [Model Overrides](#model-overrides)) takes precedence over everything below
250
251
  - **Router mode**
251
252
  - When loaded, reads `meta.n_ctx` from the `/v1/models` endpoint
252
253
  - When not loaded, reads `--ctx-size` and/or `--fit-ctx` from the server arguments (which can also originate from the **presets.ini** file the llama.cpp server uses to load its models).
@@ -256,13 +257,14 @@ The extension determines the context size as follows:
256
257
 
257
258
  ### Commands
258
259
 
259
- | Command | Description |
260
- | ------------------ | ---------------------------------------------------------------------------------- |
261
- | `/models` | Browse your models with live status. Select a model to load, switch, or unload it. |
262
- | `/models info` | Show detailed information for all available models at once. |
263
- | `/models unload` | Unload all loaded models at once. |
264
- | `/models servers` | Add, edit or remove llama.cpp server URLs via a TUI editor. |
265
- | `/models settings` | Open a menu to edit the scalar `llamaSettings` fields. |
260
+ | Command | Description |
261
+ | ------------------- | --------------------------------------------------------------------------------------- |
262
+ | `/models` | Browse your models with live status. Select a model to load, switch, or unload it. |
263
+ | `/models info` | Show detailed information for all available models at once. |
264
+ | `/models unload` | Unload all loaded models at once. |
265
+ | `/models settings` | Open a menu to edit the scalar `llamaSettings` fields. |
266
+ | `/models servers` | Add, edit or remove llama.cpp server URLs via a TUI editor. |
267
+ | `/models overrides` | Edit per-server model overrides (`llamaSettings.servers[].overrides`) via a TUI editor. |
266
268
 
267
269
  > **Note:** When a llama.cpp server is slow to respond, it will be skipped at startup with a warning. Run `/models` to retry without timeout and see all models.
268
270
 
@@ -272,15 +274,16 @@ The extension determines the context size as follows:
272
274
 
273
275
  #### Model sorting
274
276
 
275
- The order of models in the `/models` menu is controlled by the `sortBy` setting:
277
+ The order of models in the `/models` menu is controlled by the `sortBy` setting.
278
+ Servers maintain their order from `llamaSettings`; sorting applies **within each server**:
276
279
 
277
- | Value | Description |
278
- | ------------- | --------------------------------------------------------------------------------------- |
279
- | `"asc"` | Sort by model ID ascending (default) |
280
- | `"desc"` | Sort by model ID descending |
281
- | `"asc-name"` | Sort by model name ascending (ties broken by ID) |
282
- | `"desc-name"` | Sort by model name descending (ties broken by ID) |
283
- | `"api"` | No sorting — models appear in the order returned by each server's `/v1/models` endpoint |
280
+ | Value | Description |
281
+ | ------------- | ------------------------------------------------------------------------------------------------------------------------- |
282
+ | `"asc"` | Sort by model ID ascending (default) |
283
+ | `"desc"` | Sort by model ID descending |
284
+ | `"asc-name"` | Sort by model name ascending (ties broken by ID) |
285
+ | `"desc-name"` | Sort by model name descending (ties broken by ID) |
286
+ | `"api"` | No sorting — models appear in the order returned by each server's `/v1/models` endpoint, servers in `llamaSettings` order |
284
287
 
285
288
  ### Model Actions
286
289
 
@@ -327,6 +330,129 @@ User-defined budgets can override the defaults by adding a `thinkingBudgets` obj
327
330
  Only `minimal`, `low`, `medium`, `high` and `xhigh` are configurable — `off` (0) and `max` (-1, unlimited) are fixed.
328
331
  The extension automatically injects the appropriate `thinking_budget_tokens` into each request payload based on the selected level.
329
332
 
333
+ ### Model Overrides
334
+
335
+ A locally-run `llama.cpp` server is free, but you can simulate costs for budgeting, experimentation, or comparison purposes — and fine-tune what the extension reports about each model.
336
+
337
+ This extension supports **per-model, per-server configuration** via the `overrides` key inside each server entry of `llamaSettings.servers`. Each entry can override the model's `cost`, `capabilities`, `reasoning`, `contextSize`, `maxTokens`, and `compat`, regardless of what the server reports.
338
+
339
+ Add overrides to your server configuration:
340
+
341
+ ```json
342
+ {
343
+ "llamaSettings": {
344
+ "servers": [
345
+ {
346
+ "url": "http://127.0.0.1:8080",
347
+ "overrides": {
348
+ "qwen-3.8-27b": {
349
+ "cost": { "input": 0.42, "output": 3.0, "cacheRead": 0.085 }
350
+ },
351
+ "glm-5.3-flash": {
352
+ "cost": { "input": 0.15, "output": 0.5, "cacheRead": 0.03 },
353
+ "capabilities": ["text"],
354
+ "reasoning": false,
355
+ "contextSize": 32768,
356
+ "maxTokens": 4096
357
+ }
358
+ }
359
+ }
360
+ ]
361
+ }
362
+ }
363
+ ```
364
+
365
+ Every field of an override is optional — absent fields fall back to what the extension detects (`capabilities`) or to its defaults (`reasoning: true`, zeroed cost).
366
+
367
+ #### Override editor
368
+
369
+ Run `/models overrides` to edit a server's override entries without hand-editing
370
+ JSON. It opens a settings menu (same UX as `/models settings`).
371
+
372
+ Each change is written immediately to the **project**
373
+ `.pi/settings.json` if it exists, otherwise to **global**
374
+ `~/.pi/agent/settings.json`.
375
+
376
+ Overrides take effect on the next provider request after closing the editor —
377
+ no `/reload` needed.
378
+
379
+ #### Cost Fields
380
+
381
+ Inside an override, the `cost` object accepts:
382
+
383
+ | Field | Type | Description |
384
+ | ------------ | ------ | ----------------------------------- |
385
+ | `input` | number | Cost per million input tokens |
386
+ | `output` | number | Cost per million output tokens |
387
+ | `cacheRead` | number | Cost per million cache read tokens |
388
+ | `cacheWrite` | number | Cost per million cache write tokens |
389
+
390
+ All four fields are optional — unspecified fields default to zero.
391
+ In the override editor, entering `0` (or leaving a field empty) removes the field from the settings — and the `cost` object itself once no fields remain — with the same effect as leaving it unset.
392
+
393
+ #### Other Fields
394
+
395
+ | Field | Type | Description |
396
+ | -------------- | ---------------- | ---------------------------------------------------------------------------------------------- |
397
+ | `capabilities` | array of strings | Pi capabilities for the model (`"text"`, `"image"`). Fully replaces the detected capabilities. |
398
+ | `reasoning` | boolean | Whether the model is a reasoning model. Defaults to `true` when absent. |
399
+ | `contextSize` | number | Override the model's context size in tokens. Falls back to autodetection when absent or `0`. |
400
+ | `maxTokens` | number | Override max generation tokens. Falls back to context size when absent or `0`. |
401
+ | `compat` | object | OpenAI-compatible provider compatibility settings (see below). |
402
+
403
+ #### Compatibility (`compat`)
404
+
405
+ The `compat` field accepts any subset of [OpenAI-compatible provider compatibility settings](https://github.com/earendil-works/pi/blob/main/packages/ai/src/types.ts) used by the `openai-completions` API. These control how the extension talks to your server — for example, disabling `developer` role support, choosing the thinking format, enabling Anthropic-style cache control, or setting thinking token budgets.
406
+
407
+ Example:
408
+
409
+ ```json
410
+ {
411
+ "llamaSettings": {
412
+ "servers": [
413
+ {
414
+ "url": "http://127.0.0.1:8080",
415
+ "overrides": {
416
+ "llama-3": {
417
+ "compat": {
418
+ "supportsDeveloperRole": false,
419
+ "thinkingFormat": "openai",
420
+ "thinkingTokenBudgetField": "thinking_budget_tokens"
421
+ }
422
+ }
423
+ }
424
+ }
425
+ ]
426
+ }
427
+ }
428
+ ```
429
+
430
+ ### Prefix matching
431
+
432
+ Override keys are treated as **prefix filters** — a model ID matches if it starts with the key. When multiple patterns match, the **longest (most specific) match wins**. This lets you define broad patterns at the top of your overrides and override them with more specific ones below.
433
+
434
+ Example:
435
+
436
+ ```json
437
+ {
438
+ "llama": { "cost": { "input": 0.01, "output": 0.02 } },
439
+ "llama-3": { "reasoning": false },
440
+ "llama-3-8b": { "cost": { "input": 0.2, "output": 0.6 } }
441
+ }
442
+ ```
443
+
444
+ | Model ID | Matching keys | Winner (longest) | Effective override |
445
+ | ------------- | -------------------------------- | ---------------- | ------------------------------------------------------ |
446
+ | `llama-3-8b` | `llama`, `llama-3`, `llama-3-8b` | `llama-3-8b` | `{ cost: { input: 0.2, output: 0.6 } }` |
447
+ | `llama-3-70b` | `llama`, `llama-3` | `llama-3` | `{ reasoning: false }` |
448
+ | `mistral-7b` | `llama` (no) | none | defaults (zero cost, detected caps, `reasoning: true`) |
449
+
450
+ > **Note:** Exact model IDs still work — they are simply the longest possible prefix for themselves. Empty keys are silently ignored.
451
+
452
+ Model matching uses this prefix system — the model ID must start with the override key for a match.
453
+
454
+ > **Note:** Overrides are resolved through the same settings merge logic (project overrides global), so they follow the same precedence chain as other server settings. If the same URL appears multiple times with different `overrides`, only the first one's overrides will be used (consistent with existing dedup behavior).
455
+
330
456
  ### Model Selection Event
331
457
 
332
458
  When you switch models via Pi's model picker (instead of using the `/models` command), the extension listens for the `model_select` event, which also loads the requested model before the conversation begins.
@@ -349,9 +475,11 @@ If loading takes longer than **60 seconds** (configurable via `pollingTimeout`),
349
475
 
350
476
  Each model exposed to Pi includes the following defaults:
351
477
 
352
- - **`maxTokens`** — dynamically set to the model's context window (detected from llama-server)
353
- - **`reasoning`** — `true` (assumed, as llama.cpp's `/v1/models` endpoint does not expose it)
354
- - **`cost`** — all zero (local models)
478
+ - **`contextWindow`** — detected from llama-server (see how the extension determines the context size above); can be overridden per-model via `llamaSettings.servers[].overrides` (see [Model Overrides](#model-overrides))
479
+ - **`maxTokens`** — dynamically set to the model's context window (detected from llama-server); can be overridden per-model via `llamaSettings.servers[].overrides` (see [Model Overrides](#model-overrides))
480
+ - **`reasoning`** — `true` by default (llama.cpp's `/v1/models` endpoint does not expose it); can be overridden per-model via `llamaSettings.servers[].overrides` (see [Model Overrides](#model-overrides))
481
+ - **`cost`** — all zero by default; can be customized per-model via `llamaSettings.servers[].overrides` (see [Model Overrides](#model-overrides))
482
+ - **`compat`** — OpenAI-compatible provider compatibility settings; can be set per-model via `llamaSettings.servers[].overrides` (see [Model Overrides](#model-overrides))
355
483
 
356
484
  ## Dependencies
357
485
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "pi-llama-cpp",
3
- "version": "0.11.0",
3
+ "version": "0.13.0",
4
4
  "description": "Pi extension for llama.cpp integration. Supports router, single and legacy models. Supports multiple servers.",
5
5
  "keywords": [
6
6
  "pi",
@@ -32,14 +32,14 @@
32
32
  ]
33
33
  },
34
34
  "peerDependencies": {
35
- "@earendil-works/pi-ai": ">=0.84.0",
36
- "@earendil-works/pi-coding-agent": ">=0.84.0",
37
- "@earendil-works/pi-tui": ">=0.84.0"
35
+ "@earendil-works/pi-ai": ">=0.85.1",
36
+ "@earendil-works/pi-coding-agent": ">=0.85.1",
37
+ "@earendil-works/pi-tui": ">=0.85.1"
38
38
  },
39
39
  "type": "module",
40
40
  "devDependencies": {
41
41
  "@types/node": "^26.4.1",
42
42
  "prettier-plugin-organize-imports": "^4.3.0",
43
- "vitest": "^5.0.0"
43
+ "vitest": "^5.0.1"
44
44
  }
45
45
  }
package/src/constants.ts CHANGED
@@ -28,6 +28,14 @@ export const API_KEY_PLACEHOLDER = "sk-placeholder";
28
28
  */
29
29
  export const LLAMA_SERVER_URL = "http://127.0.0.1:8080";
30
30
 
31
+ /**
32
+ * Endpoint prefix of the OpenAI-compatible API exposed by llama-server.
33
+ * Native routes (/health, /props, /models/sse, ...) live directly under the
34
+ * server root; only OpenAI-dialect consumers (the Pi provider registration
35
+ * and the /models listing) address the API under this prefix.
36
+ */
37
+ export const ENDPOINT_PREFIX = "/v1";
38
+
31
39
  /**
32
40
  * The default context if the server didn't expose it
33
41
  */
@@ -1,3 +1,5 @@
1
+ import type { ModelOverride } from "./settings";
2
+
1
3
  /**
2
4
  * Identity of a llama.cpp server endpoint.
3
5
  *
@@ -17,4 +19,9 @@ export interface ServerOptions {
17
19
  * Custom provider name suffix; falls back to the base URL.
18
20
  */
19
21
  customName?: string;
22
+ /**
23
+ * Per-model overrides resolved from `llamaSettings.servers`. See
24
+ * {@link ModelOverride} for the fallback semantics of each field.
25
+ */
26
+ overrides?: Record<string, ModelOverride>;
20
27
  }
@@ -1,11 +1,50 @@
1
+ import type { ModelCost, OpenAICompletionsCompat } from "@earendil-works/pi-ai";
1
2
  import type { SortBy } from "../constants";
2
3
 
4
+ /**
5
+ * Per-model overrides applied on top of what llama-server reports.
6
+ * Every field is optional — absent fields fall back to detection/defaults.
7
+ */
8
+ export interface ModelOverride {
9
+ /**
10
+ * Per-model token pricing. All four cost fields are optional —
11
+ * unspecified fields default to zero.
12
+ */
13
+ cost?: Partial<ModelCost>;
14
+ /**
15
+ * Pi capabilities for the model. When present, **fully replaces** the
16
+ * capabilities detected from the server (no merging).
17
+ */
18
+ capabilities?: ("text" | "image")[];
19
+ /**
20
+ * Whether the model is a reasoning model. When absent, defaults to `true`.
21
+ */
22
+ reasoning?: boolean;
23
+ /**
24
+ * Override the model's context size (in tokens), replacing the value
25
+ * autodetected from the server. When absent or `0`, falls back to
26
+ * detection (then `FALLBACK_CTX`).
27
+ */
28
+ contextSize?: number;
29
+ /**
30
+ * Override the maximum number of tokens the model can generate. When
31
+ * absent, falls back to the context size detected from the server.
32
+ */
33
+ maxTokens?: number;
34
+ /**
35
+ * OpenAI-compatible provider compatibility settings. Merged with any
36
+ * provider-level compat when the model is registered with Pi.
37
+ */
38
+ compat?: Partial<OpenAICompletionsCompat>;
39
+ }
40
+
3
41
  /**
4
42
  * A description of a server in the "llamaSettings" key
5
43
  */
6
44
  export interface LlamaServer {
7
45
  /**
8
- * The URL of the llama.cpp server.
46
+ * The URL of the llama.cpp server. Must be the bare origin — the
47
+ * OpenAI-compatible API is assumed under `/v1` (see {@link ENDPOINT_PREFIX}).
9
48
  */
10
49
  url: string;
11
50
  /**
@@ -16,6 +55,33 @@ export interface LlamaServer {
16
55
  * Custom display name for this server.
17
56
  */
18
57
  name?: string;
58
+ /**
59
+ * Per-model overrides for this server. Keys are **prefix filters** —
60
+ * a model ID matches if it starts with the key. When multiple patterns
61
+ * match, the **longest (most specific) match wins**.
62
+ *
63
+ * All fields of an override are optional — absent fields fall back to
64
+ * detection (`capabilities`) or defaults (`reasoning: true`, zero costs).
65
+ *
66
+ * Example:
67
+ * ```json
68
+ * {
69
+ * "llama": { "cost": { "input": 0.01, "output": 0.02 } },
70
+ * "llama-3": { "reasoning": false },
71
+ * "llama-3-8b": {
72
+ * "cost": { "input": 0.2, "output": 0.6, "cacheRead": 0.01 },
73
+ * "capabilities": ["text", "image"]
74
+ * }
75
+ * }
76
+ * ```
77
+ *
78
+ * For model `"llama-3-8b"`:
79
+ * - `"llama"` matches → cost `{ input: 0.01, output: 0.02 }`
80
+ * - `"llama-3"` matches → reasoning `false`
81
+ * - `"llama-3-8b"` matches → cost + capabilities fully replaced
82
+ * - **Winner**: `"llama-3-8b"` (longest match)
83
+ */
84
+ overrides?: Record<string, ModelOverride>;
19
85
  }
20
86
 
21
87
  /**