@knightcodeai/cli-linux-x64 0.10.0 → 0.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/bin/CHANGELOG.md +24 -0
- package/bin/docs/docs.json +4 -0
- package/bin/docs/extensions.md +5 -1
- package/bin/docs/images/doom-extension.png +0 -0
- package/bin/docs/images/interactive-mode.png +0 -0
- package/bin/docs/images/tree-view.png +0 -0
- package/bin/docs/keybindings.md +1 -1
- package/bin/docs/llama-cpp.md +11 -0
- package/bin/docs/session-format.md +5 -1
- package/bin/docs/sessions.md +2 -0
- package/bin/docs/virtual-models.md +114 -0
- package/bin/knightcode +2 -2
- package/bin/package.json +6 -6
- package/package.json +1 -1
package/bin/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,29 @@
|
|
|
1
1
|
# @knightcodeai/cli
|
|
2
2
|
|
|
3
|
+
## 0.11.0
|
|
4
|
+
|
|
5
|
+
### Added
|
|
6
|
+
|
|
7
|
+
- Added virtual models: an extension can register a selectable model, with `registerVirtualModel`, that routes each request to a physical model and thinking level. Virtual models appear in `/model`, `--model` and scoped models, the footer shows the routed model, and assistant messages record the thinking level the agent loop requested. See `docs/virtual-models.md`.
|
|
8
|
+
|
|
9
|
+
- Added Claude Sonnet 5.5, with managed effort levels, mid-conversation system messages and tool changes, and no temperature control.
|
|
10
|
+
|
|
11
|
+
- Added a classifier for every connected llama.cpp chat model, so choice, bool, and score questions are answered from next-token label probabilities, with an optional temperature that softens the distribution.
|
|
12
|
+
|
|
13
|
+
### Fixed
|
|
14
|
+
|
|
15
|
+
- Fixed OpenCode and OpenCode Go Qwen 3.8 Flash rejecting thinking blocks that have an empty signature.
|
|
16
|
+
|
|
17
|
+
- Fixed Mistral reasoning requests using a hardcoded model list. A model with a thinking-level map now sends `reasoning_effort` for the requested level, including max on GLM 5.2, and sends the model's off value when thinking is off. GLM 5.3 no longer uses `prompt_mode`.
|
|
18
|
+
|
|
19
|
+
- Fixed the OpenCode Go default model pointing at Kimi K2.6; it now defaults to Kimi K3.
|
|
20
|
+
|
|
21
|
+
- Fixed pasting files copied in the macOS Finder inserting their icon instead of their paths. Copied files now paste as their original paths, and a failed clipboard paste shows an error instead of failing silently.
|
|
22
|
+
|
|
23
|
+
- Fixed OpenAI Responses streams running tool calls that never finished. A server that omits `output_index`, such as llama.cpp, could turn two parallel calls into three with cut-off or mixed-up arguments; the stream now fails instead of running them.
|
|
24
|
+
|
|
25
|
+
- Fixed the Together default model pointing at Kimi K2.6, which Together no longer lists; it now defaults to Kimi K3.
|
|
26
|
+
|
|
3
27
|
## 0.10.0
|
|
4
28
|
|
|
5
29
|
### Added
|
package/bin/docs/docs.json
CHANGED
package/bin/docs/extensions.md
CHANGED
|
@@ -80,6 +80,7 @@ Automatic retries, recovery, compaction, or queued work can continue afterward.
|
|
|
80
80
|
| Persist non-context session data | `knightcode.appendEntry()` |
|
|
81
81
|
| Change active tools, model, or thinking level | Session control methods on `knightcode` |
|
|
82
82
|
| Add a model provider | `knightcode.registerProvider()` |
|
|
83
|
+
| Route each request to a model | [`knightcode.registerVirtualModel()`](virtual-models.md) |
|
|
83
84
|
| Add terminal rendering | Renderer registration and `ctx.ui` |
|
|
84
85
|
| Communicate with another extension | `knightcode.events` |
|
|
85
86
|
|
|
@@ -194,7 +195,10 @@ Use `ctx.ui.custom()` only when the interaction needs its own rendering and inpu
|
|
|
194
195
|
See [Terminal UI](tui.md) for component, focus, overlay, theme, and performance guidance.
|
|
195
196
|
|
|
196
197
|
Extensions load in interactive, RPC, JSON, and print modes.
|
|
197
|
-
Interactive mode provides the complete terminal UI.
|
|
198
|
+
Interactive mode provides the complete terminal UI. The [DOOM overlay example](../examples/extensions/doom-overlay/) uses a custom overlay component to render a game frame by frame:
|
|
199
|
+
|
|
200
|
+
<p align="center"><img src="images/doom-extension.png" alt="The DOOM overlay example running over a KnightCode session" width="750"></p>
|
|
201
|
+
|
|
198
202
|
RPC can forward supported dialogs and notifications through the [RPC Extension UI protocol](rpc-extension-ui.md), but not custom terminal components; JSON and print modes have no UI.
|
|
199
203
|
Guard terminal-only behavior with `ctx.mode === "tui"` and use `ctx.hasUI` for interactions supported by interactive and RPC clients.
|
|
200
204
|
|
|
Binary file
|
|
Binary file
|
|
Binary file
|
package/bin/docs/keybindings.md
CHANGED
|
@@ -125,7 +125,7 @@ In fullscreen mode, these actions control the transcript and take precedence ove
|
|
|
125
125
|
| `app.exit` | `ctrl+d` | Exit (when editor empty) |
|
|
126
126
|
| `app.suspend` | `ctrl+z` (None on Windows) | Suspend to background |
|
|
127
127
|
| `app.editor.external` | `ctrl+g` | Open in external editor (`externalEditor`, `$VISUAL`, `$EDITOR`, Notepad on Windows, or `nano` elsewhere) |
|
|
128
|
-
| `app.clipboard.pasteImage` | `ctrl+v` (`alt+v` on Windows and WSL) | Paste
|
|
128
|
+
| `app.clipboard.pasteImage` | `ctrl+v` (`alt+v` on Windows and WSL) | Paste files on macOS, images, or text from clipboard |
|
|
129
129
|
|
|
130
130
|
On native Windows, `app.suspend` has no default because Windows terminals do not support Unix job control. If you assign it manually, KnightCode shows a status message instead of suspending. WSL uses the normal `ctrl+z` and `fg` behavior.
|
|
131
131
|
|
package/bin/docs/llama-cpp.md
CHANGED
|
@@ -86,6 +86,17 @@ Loaded and sleeping models appear in `/model`. Sleeping models wake automaticall
|
|
|
86
86
|
|
|
87
87
|
If the router disconnects, `/llama` shows **Retry** and **Close**. Retry reconnects and refreshes model state without replaying the interrupted operation.
|
|
88
88
|
|
|
89
|
+
## Classification
|
|
90
|
+
|
|
91
|
+
Every model listed for chat is also listed as a classifier model with the same ID and the `llama-cpp-classify` API. Classifier models answer typed `choice`, `bool`, and `score` questions about JSON state, like TypeSafe's Jev models.
|
|
92
|
+
|
|
93
|
+
The model does not generate an answer. Each question becomes one chat prompt: the state, every question of the request, the state again, and then the question with its answers under single-token labels. Labels are letters for a choice (up to 62 options), `Yes`/`No` for a bool, and digits for a score (up to 10 levels). The second copy of the state is read with the questions in view, which improved accuracy on JevBench with small models. KnightCode reads the probabilities of the labels as the next token and normalizes them. A choice returns every option's probability and a confidence of `(n * peak - 1) / (n - 1)`; a score returns the expected level.
|
|
94
|
+
|
|
95
|
+
- Raw label probabilities are usually overconfident. The per-request `temperature` option divides the label logits before normalizing; values above 1 soften the distribution. It changes no answer.
|
|
96
|
+
- Questions run one after another. Everything before the final question is the same for all questions of a request, so the server's prompt cache evaluates it once. The state appears twice, so it needs twice its size in context.
|
|
97
|
+
- Small models may follow instructions written inside the state. The prompt tells the model to judge the state as data, but that is not a guarantee.
|
|
98
|
+
- Hybrid models such as Qwen3.5 cannot rewind a partially cached prompt without context checkpoints. If each question reprocesses the whole state, start the router with `--ctx-checkpoints 32 --checkpoint-min-step 0`.
|
|
99
|
+
|
|
89
100
|
## Troubleshooting
|
|
90
101
|
|
|
91
102
|
Check that the router is reachable:
|
|
@@ -92,9 +92,11 @@ Sessions created before system messages existed have no leading system message;
|
|
|
92
92
|
{"type":"message","id":"c3d4e5f6","parentId":"b2c3d4e5","timestamp":"2024-12-03T14:00:03.000Z","message":{"role":"toolResult","toolCallId":"call_123","toolName":"bash","content":[{"type":"text","text":"output"}],"isError":false,"timestamp":1733234403000}}
|
|
93
93
|
```
|
|
94
94
|
|
|
95
|
+
Assistant messages name the model that produced them. Newer messages also record `thinkingLevel`, the KnightCode thinking level requested for that response.
|
|
96
|
+
|
|
95
97
|
### ModelChangeEntry
|
|
96
98
|
|
|
97
|
-
Emitted when the user switches models mid-session.
|
|
99
|
+
Emitted when the user switches models mid-session. The latest entry is the selected model, which may be a [virtual model](virtual-models.md); assistant messages then name the physical model that answered.
|
|
98
100
|
|
|
99
101
|
```json
|
|
100
102
|
{"type":"model_change","id":"d4e5f6g7","parentId":"c3d4e5f6","timestamp":"2024-12-03T14:05:00.000Z","provider":"openai","modelId":"gpt-4o"}
|
|
@@ -169,6 +171,8 @@ Extension state persistence. Does NOT participate in LLM context.
|
|
|
169
171
|
|
|
170
172
|
Use `customType` to identify your extension's entries on reload. Interactive mode can render custom entries via `knightcode.registerEntryRenderer(customType, renderer)`, but they still do not participate in LLM context.
|
|
171
173
|
|
|
174
|
+
KnightCode stores [virtual model](virtual-models.md) router state as custom entries with `customType` `knightcode.virtual-model-state` and `data` `{ provider, modelId, state }`.
|
|
175
|
+
|
|
172
176
|
### CustomMessageEntry
|
|
173
177
|
|
|
174
178
|
Extension-injected messages that DO participate in LLM context.
|
package/bin/docs/sessions.md
CHANGED
|
@@ -30,6 +30,8 @@ KnightCode stores entries as a tree, so returning to an earlier point does not e
|
|
|
30
30
|
|
|
31
31
|
In `/tree`, select a user message to put its text back in the editor. Edit and submit it to create another branch. Selecting an assistant response or another entry continues after that entry with an empty editor.
|
|
32
32
|
|
|
33
|
+
<p align="center"><img src="images/tree-view.png" alt="The /tree view showing a session with two branches, user messages, assistant replies, and tool calls" width="750"></p>
|
|
34
|
+
|
|
33
35
|
Selecting a point while the model is responding cancels that response. Navigation cannot proceed while compaction or another tree navigation is still running; wait for it to finish and retry.
|
|
34
36
|
|
|
35
37
|
When you leave a branch, KnightCode can summarize it and attach that summary to the branch you enter. This preserves relevant work from the abandoned path without including every message from it.
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
# Virtual Models
|
|
2
|
+
|
|
3
|
+
A virtual model is a selectable model that picks a physical model for each request. Use one to route by task, cost, or conversation state. For example, a router can send quick questions to a small model and hard problems to a large one, while the user selects a single model.
|
|
4
|
+
|
|
5
|
+
Register virtual models from an [extension](extensions.md). They appear in `/model`, `--model`, scoped models, and settings like any other model. A virtual model can be listed under any provider, including one with physical models, such as `openai-codex/auto`.
|
|
6
|
+
|
|
7
|
+
## Selection and dispatch
|
|
8
|
+
|
|
9
|
+
A virtual model selects a model and a thinking level. A router maps that pair to a physical pair for each request:
|
|
10
|
+
|
|
11
|
+
```
|
|
12
|
+
selected (virtual model, virtual level) -> dispatched (physical model, physical level)
|
|
13
|
+
jev/auto:low -> anthropic/claude-sonnet-4-5:high
|
|
14
|
+
```
|
|
15
|
+
|
|
16
|
+
The virtual thinking level is an input to the router. Its meaning is up to the router; it need not correspond to a reasoning budget.
|
|
17
|
+
|
|
18
|
+
KnightCode keeps the two pairs apart:
|
|
19
|
+
|
|
20
|
+
| | Selection | Dispatch |
|
|
21
|
+
|---|---|---|
|
|
22
|
+
| Recorded in | `model_change` and `thinking_level_change` entries | Each assistant message: `provider`, `api`, `model`, `thinkingLevel` |
|
|
23
|
+
| Visible as | `ctx.model`, `ctx.thinkingLevel`, `KNIGHTCODE_MODEL`, `KNIGHTCODE_REASONING_LEVEL`, `/model` | The assistant message of each response |
|
|
24
|
+
|
|
25
|
+
Providers only receive physical models. Assistant messages name the physical model, so replaying a conversation across different physical models works the same as after a manual model switch. Resuming a session restores the virtual selection from its latest `model_change` entry. If the virtual model is no longer registered, KnightCode falls back to the physical model that answered last.
|
|
26
|
+
|
|
27
|
+
In interactive mode, the footer shows the routed model next to the selection, for example `auto • high → gpt-5.6-luna • medium`. `/session` lists the cost for each physical model.
|
|
28
|
+
|
|
29
|
+
Context usage uses the limits of the physical model that produced the latest response, even if that response came before switching to the virtual model. Without such a response, it uses the limits declared on the virtual model, if any. Compaction checks the same limits, and again the limits of the model each request is routed to. If that model's context window is too small for the conversation, KnightCode compacts before sending the request; the route stays as the router chose it.
|
|
30
|
+
|
|
31
|
+
## Register a virtual model
|
|
32
|
+
|
|
33
|
+
```typescript
|
|
34
|
+
import type { ExtensionAPI } from "@knightcodeai/cli";
|
|
35
|
+
|
|
36
|
+
export default function (knightcode: ExtensionAPI) {
|
|
37
|
+
knightcode.registerVirtualModel({
|
|
38
|
+
provider: "router",
|
|
39
|
+
id: "auto",
|
|
40
|
+
name: "Auto",
|
|
41
|
+
thinkingLevels: ["low", "high"],
|
|
42
|
+
route(request, ctx) {
|
|
43
|
+
// Tool follow-ups and retries stay on the model that handled the turn.
|
|
44
|
+
const sticky = request.failed ?? request.previous;
|
|
45
|
+
if (request.reason !== "user" && sticky) {
|
|
46
|
+
return { model: sticky.model, thinkingLevel: sticky.thinkingLevel ?? "medium" };
|
|
47
|
+
}
|
|
48
|
+
const id = request.thinkingLevel === "high" ? "claude-sonnet-4-5" : "claude-haiku-4-5";
|
|
49
|
+
return { model: ctx.modelRegistry.find("anthropic", id)!, thinkingLevel: "medium" };
|
|
50
|
+
},
|
|
51
|
+
});
|
|
52
|
+
}
|
|
53
|
+
```
|
|
54
|
+
|
|
55
|
+
- `provider` is the provider the model is listed under. It can be any provider ID. A provider can list several virtual models next to its physical ones. On a physical provider, the virtual model is available when that provider has credentials. Under an ID that no provider uses, it is always available.
|
|
56
|
+
- `id` must not be the ID of a physical model of that provider. If a catalog refresh later adds a physical model with the same ID, the virtual model hides it.
|
|
57
|
+
- `thinkingLevels` lists the levels offered for selection. It defaults to `["off"]`.
|
|
58
|
+
- `contextWindow` and `maxTokens` are shown before the first response. Unset limits are unknown.
|
|
59
|
+
- `input` lists the input types offered for selection. It defaults to text and images; physical models without image support receive placeholders.
|
|
60
|
+
|
|
61
|
+
Registration follows the same queuing and reload rules as `knightcode.registerProvider()`. Registering the same provider and ID again replaces the virtual model. `knightcode.unregisterVirtualModel(provider, id)` removes it; `knightcode.unregisterProvider()` does not. SDK code can register one without an extension: `modelRuntime.registerVirtualModel(definition)`.
|
|
62
|
+
|
|
63
|
+
## Route requests
|
|
64
|
+
|
|
65
|
+
`route(request, ctx)` runs before every request made with the virtual model and returns `{ model, thinkingLevel }`. The model can be any physical model in the catalog whose provider has credentials; look it up with `ctx.modelRegistry`. A virtual model cannot route to another virtual model. KnightCode clamps the thinking level to the returned model.
|
|
66
|
+
|
|
67
|
+
| Field | Meaning |
|
|
68
|
+
|---|---|
|
|
69
|
+
| `model`, `thinkingLevel` | The selected virtual model and level |
|
|
70
|
+
| `reason` | Why the request is made, see below |
|
|
71
|
+
| `previous` | Physical model and thinking level of the latest successful response in `messages` |
|
|
72
|
+
| `failed` | For `retry`: physical model, thinking level, and assistant `message` of the failed request, which `messages` no longer contains. The message carries `stopReason` and `errorMessage`. Absent when routing itself failed |
|
|
73
|
+
| `state` | Router state last returned on this session branch, see below |
|
|
74
|
+
| `messages` | The conversation for this request, including system messages |
|
|
75
|
+
| `signal` | Abort signal of the request |
|
|
76
|
+
|
|
77
|
+
| `reason` | Request |
|
|
78
|
+
|---|---|
|
|
79
|
+
| `user` | First request after a message the user wrote, including steering and follow-up messages |
|
|
80
|
+
| `continuation` | Any other request in the agent loop, such as after tool results or extension messages |
|
|
81
|
+
| `retry` | Automatic retry after a failed request, including after compaction for a context overflow |
|
|
82
|
+
| `direct` | Request made outside the agent loop, such as a compaction summary or an extension calling `ctx.modelRegistry.streamSimple()` |
|
|
83
|
+
|
|
84
|
+
Returning `previous` for `continuation` and `failed` for `retry` keeps prompt caches and thinking signatures valid. Switching models between turns is allowed but loses the prompt cache. A retry can also switch to another model, for example when `failed.message.errorMessage` reports that a provider is overloaded or the context overflowed.
|
|
85
|
+
|
|
86
|
+
If `route()` throws, or returns a virtual model or a model without credentials, the request ends with an error response.
|
|
87
|
+
|
|
88
|
+
## Keep routing state
|
|
89
|
+
|
|
90
|
+
`route()` can return `state` next to the model. KnightCode stores it on the session branch and passes it back as `request.state` on later requests. Use it for decisions the transcript does not record, such as classifier results or a routing phase:
|
|
91
|
+
|
|
92
|
+
```typescript
|
|
93
|
+
knightcode.registerVirtualModel<{ phase: "plan" | "build" }>({
|
|
94
|
+
provider: "router",
|
|
95
|
+
id: "phased",
|
|
96
|
+
name: "Phased",
|
|
97
|
+
route(request, ctx) {
|
|
98
|
+
const state = request.state ?? { phase: "plan" };
|
|
99
|
+
const id = state.phase === "plan" ? "claude-opus-4-5" : "claude-haiku-4-5";
|
|
100
|
+
return { model: ctx.modelRegistry.find("anthropic", id)!, thinkingLevel: "medium", state };
|
|
101
|
+
},
|
|
102
|
+
});
|
|
103
|
+
```
|
|
104
|
+
|
|
105
|
+
- State must be JSON-serializable. Returning `undefined` or `request.state` itself keeps the current state.
|
|
106
|
+
- KnightCode stores any other returned object as new state, before the request is sent, even when it equals the current state. Return a new object only when the state changes. The state stays stored if the request later fails.
|
|
107
|
+
- State follows the session tree, so forks and `/tree` navigation see the state of their branch. It survives compaction.
|
|
108
|
+
- `direct` requests have no state, and KnightCode ignores state they return.
|
|
109
|
+
|
|
110
|
+
The transcript already records the selection and every dispatched model, and `ctx.sessionManager.getBranch()` exposes both.
|
|
111
|
+
|
|
112
|
+
Routers can call other models through `ctx.modelRegistry`, for example `ctx.modelRegistry.classify()` with a classifier model from `ctx.modelRegistry.findOfType("classifier", provider, id)`. The call adds latency before the first token of the turn.
|
|
113
|
+
|
|
114
|
+
See [`jev-router.ts`](../examples/extensions/jev-router.ts) for a complete router. It plans on a strong OpenAI Codex model chosen by the Jev classifier, lets that model make the first edit, and then switches once to a cheaper model, accepting a single prompt-cache miss. It keeps the phase as router state.
|
package/bin/knightcode
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
1
|
[diffend] Oversized file quarantined before diffing.
|
|
2
2
|
name: package/bin/knightcode
|
|
3
|
-
size:
|
|
4
|
-
sha256:
|
|
3
|
+
size: 118354601 bytes
|
|
4
|
+
sha256: a671810feb8a88b437a08df5505351b5a01b9e3b74ad834825f90c7e15a75c2b
|
package/bin/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@knightcodeai/cli",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.11.0",
|
|
4
4
|
"description": "KnightCode — a local, BYOK terminal coding agent powered by OpenRouter.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"repository": {
|
|
@@ -37,11 +37,11 @@
|
|
|
37
37
|
"test": "vitest --run"
|
|
38
38
|
},
|
|
39
39
|
"optionalDependencies": {
|
|
40
|
-
"@knightcodeai/cli-linux-x64": "0.
|
|
41
|
-
"@knightcodeai/cli-linux-arm64": "0.
|
|
42
|
-
"@knightcodeai/cli-darwin-x64": "0.
|
|
43
|
-
"@knightcodeai/cli-darwin-arm64": "0.
|
|
44
|
-
"@knightcodeai/cli-win32-x64": "0.
|
|
40
|
+
"@knightcodeai/cli-linux-x64": "0.11.0",
|
|
41
|
+
"@knightcodeai/cli-linux-arm64": "0.11.0",
|
|
42
|
+
"@knightcodeai/cli-darwin-x64": "0.11.0",
|
|
43
|
+
"@knightcodeai/cli-darwin-arm64": "0.11.0",
|
|
44
|
+
"@knightcodeai/cli-win32-x64": "0.11.0"
|
|
45
45
|
},
|
|
46
46
|
"devDependencies": {
|
|
47
47
|
"@agentclientprotocol/sdk": "1.4.0",
|