@knightcodeai/cli-darwin-x64 0.10.0 → 0.11.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/bin/CHANGELOG.md CHANGED
@@ -1,5 +1,67 @@
1
1
  # @knightcodeai/cli
2
2
 
3
+ ## 0.11.1
4
+
5
+ ### Added
6
+
7
+ - Added the `codemode` tool, which runs JavaScript in a sandbox whose only capability is calling tools, including MCP tools that are not declared to the model.
8
+
9
+ - Added the `tool_search` tool so the model can find tools that are not declared yet and declare the matches on its next call.
10
+
11
+ - Added Sign in with ChatGPT for the OpenAI provider. Provider sign-in flows share one local callback server, and the browser page still shows the KnightCode knight.
12
+
13
+ - Added MCP servers, configured in `mcp.json` and managed with `knightcode mcp` and `/mcp`, including OAuth sign-in for HTTP servers.
14
+
15
+ - Added GPT-6.1 Sol (`gpt-6.1-sol`) for OpenAI, Azure OpenAI Responses, and OpenAI Codex, and made it the OpenAI Codex default model.
16
+
17
+ - Added Jev classifier models on Vercel AI Gateway (`typesafe-ai/jev`) and OpenCode (`jev-1.13` and `jev-1.13-free`).
18
+
19
+ - Added the `fullscreenWheelScrollLines` setting so fullscreen mode can scroll one line per wheel event, a fixed number of lines, or speed up fast spins. Alt+wheel moves five times as far.
20
+
21
+ ### Changed
22
+
23
+ - Changed `defaultTools` so a list of `+name` and `-name` entries adds or removes tools instead of replacing the whole selection.
24
+
25
+ - Changed built-in extensions so `knightcode config` lists them as `builtin:<name>` and settings can disable one with `-builtin:<name>`.
26
+
27
+ - Changed tool calls that have no custom renderer to show their arguments, collapsed on one line until the call is expanded.
28
+
29
+ - Changed session cost to include token usage from codemode classifier calls and from tools those scripts call.
30
+
31
+ - Changed bash results returned to scripts so they keep up to 1 MiB of output, with the full output in `full_output_path` when that is still not enough.
32
+
33
+ - Changed System One classifier results to include token usage and its catalog cost when the service reports token counts.
34
+
35
+ - Changed fullscreen redraws to keep parsed markdown and unpadded child lines across frames, and shell output now drops only control characters and interlinear annotations instead of every format character.
36
+
37
+ ### Fixed
38
+
39
+ - Fixed llama.cpp context windows so a reload keeps the last known size when the server has not reported one, and a configured `--ctx-size` is used before the training context.
40
+
41
+ ## 0.11.0
42
+
43
+ ### Added
44
+
45
+ - Added virtual models: an extension can register a selectable model, with `registerVirtualModel`, that routes each request to a physical model and thinking level. Virtual models appear in `/model`, `--model` and scoped models, the footer shows the routed model, and assistant messages record the thinking level the agent loop requested. See `docs/virtual-models.md`.
46
+
47
+ - Added Claude Sonnet 5.5, with managed effort levels, mid-conversation system messages and tool changes, and no temperature control.
48
+
49
+ - Added a classifier for every connected llama.cpp chat model, so choice, bool, and score questions are answered from next-token label probabilities, with an optional temperature that softens the distribution.
50
+
51
+ ### Fixed
52
+
53
+ - Fixed OpenCode and OpenCode Go Qwen 3.8 Flash rejecting thinking blocks that have an empty signature.
54
+
55
+ - Fixed Mistral reasoning requests using a hardcoded model list. A model with a thinking-level map now sends `reasoning_effort` for the requested level, including max on GLM 5.2, and sends the model's off value when thinking is off. GLM 5.3 no longer uses `prompt_mode`.
56
+
57
+ - Fixed the OpenCode Go default model pointing at Kimi K2.6; it now defaults to Kimi K3.
58
+
59
+ - Fixed pasting files copied in the macOS Finder inserting their icon instead of their paths. Copied files now paste as their original paths, and a failed clipboard paste shows an error instead of failing silently.
60
+
61
+ - Fixed OpenAI Responses streams running tool calls that never finished. A server that omits `output_index`, such as llama.cpp, could turn two parallel calls into three with cut-off or mixed-up arguments; the stream now fails instead of running them.
62
+
63
+ - Fixed the Together default model pointing at Kimi K2.6, which Together no longer lists; it now defaults to Kimi K3.
64
+
3
65
  ## 0.10.0
4
66
 
5
67
  ### Added
package/bin/docs/cli.md CHANGED
@@ -13,6 +13,7 @@ knightcode update [target] [options]
13
13
  knightcode list
14
14
  knightcode config [options]
15
15
  knightcode auth <check|print-api-key|print-bearer-token> [options]
16
+ knightcode mcp <list|login|logout> [options]
16
17
  ```
17
18
 
18
19
  <a id="modes"></a>
@@ -124,7 +125,7 @@ See [Settings](settings.md#tools) for configuring the default tool selection.
124
125
  - `-nt`, `--no-tools`<br>
125
126
  Starts with all built-in, extension, and custom tools disabled.
126
127
 
127
- Default enabled tools are `read`, `bash`, `edit`, and `write`, unless `defaultTools` changes them.
128
+ Default enabled tools are `read`, `bash`, `edit`, and `write`, unless `defaultTools` changes them. `--tools` replaces the whole selection, so name every tool you want; `defaultTools` also accepts `+name` and `-name` to change the defaults instead.
128
129
 
129
130
  | Built-in | Purpose |
130
131
  |---|---|
@@ -139,6 +140,49 @@ Default enabled tools are `read`, `bash`, `edit`, and `write`, unless `defaultTo
139
140
 
140
141
  The [web tools](usage.md#web-tools) `webfetch` and `websearch` are off until enabled with `/tools`. A tool excluded here stays off whatever `/tools` says.
141
142
 
143
+ Built-in extensions add two more tools. They are off by default; the MCP extension turns them on when an MCP server needs them (see [MCP](mcp.md#exposure)). To enable them yourself, name them in `--tools` or `defaultTools`.
144
+
145
+ | Built-in extension | Purpose |
146
+ |---|---|
147
+ | `codemode` | Run JavaScript that calls the other tools, for example in parallel with `Promise.allSettled`; only the script's output reaches the model |
148
+ | `tool_search` | Search tools that are not declared to the model (`codemode` and `deferred` exposure, such as MCP tools) and declare the matches for the next call |
149
+
150
+ ### Enable codemode
151
+
152
+ To turn on `codemode` for every session, add it to the default tools in `~/.knightcode/agent/settings.json` or a project's `.knightcode/settings.json`:
153
+
154
+ ```json
155
+ {
156
+ "defaultTools": ["+codemode"]
157
+ }
158
+ ```
159
+
160
+ This keeps `read`, `bash`, `edit`, and `write` and adds `codemode`. For one invocation, list every tool, since `--tools` replaces the selection:
161
+
162
+ ```sh
163
+ knightcode --tools read,bash,edit,write,codemode
164
+ ```
165
+
166
+ Codemode is useful without MCP: scripts can run several tool calls in parallel, filter large output before it reaches the model, and call classifier models such as TypeSafe's Jev through `models.classify()` (see [Classifier models](models.md#use-classifier-models)).
167
+
168
+ ### How codemode works
169
+
170
+ Codemode scripts run in a QuickJS sandbox that can only reach the other tools, through `tools.<name>(args)`; `ALL_TOOLS` lists them. Output comes from `text(value)`, `image(dataUrlOrImageContent)`, `console.*`, and a top-level `return value`; `exit()` ends the script early. The result starts with `Script completed` or `Script failed`, the wall time, and the output; a failed script keeps its partial output, followed by `Script error:` and the error.
171
+
172
+ A script may start with an options line such as `// @options: {"max_output_tokens": 2000, "timeout_ms": 60000}`. `max_output_tokens` (default 10000) limits the output: longer output keeps its start and end, and the full text is written to a temp file whose path is included in the result. `timeout_ms` is a hard deadline, unset by default.
173
+
174
+ While `codemode` is active, `codemode.mode` in [settings](settings.md#tools) decides how the other tools are presented. With `on` (default) declared tools keep being declared and their descriptions show how to call them from scripts. With `only` they are hidden from the model and listed in the `codemode` description instead, so the model calls them through scripts.
175
+
176
+ The `codemode` description lists the callable tools with their TypeScript declarations, grouped by namespace (for example one MCP server). Declarations share a budget of 3000 estimated tokens (`codemode.inlineBudget` in [settings](settings.md#tools)); every namespace is still listed with its tool count, and the description says whether the list is complete. Scripts find the rest with `await searchTools(query, { limit, namespace })`, which ranks tools with BM25, and `await describeTool(name)`, or by filtering `ALL_TOOLS`.
177
+
178
+ Tools with an output schema resolve to structured values: `bash` to `{ output, truncated, full_output_path?, exit_code, wall_time_seconds }`, also for non-zero exit codes, and MCP tools to their `CallToolResult`. Other tools resolve to their text output. The `output` of `bash` is not limited to the 2000 lines or 50KB the model sees: it holds up to 1 MiB, and longer output keeps its first and last 512 KiB around an omission marker, with `truncated` set and the full output in `full_output_path`.
179
+
180
+ `store(key, value)` and `load(key)` keep JSON values across `codemode` calls: each successful script that stores values appends a `codemode-store` custom entry to the session, so resumed sessions keep the values and each branch sees only the values written on its path. Scripts can also use `models`: `getModelsOfType`, `getAvailableOfType`, and `getModelOfType` list the model catalog, and `classify(model, context)` runs a classifier model with the session's credentials, at most four at a time per script.
181
+
182
+ ### Tool search
183
+
184
+ `tool_search` is off by default; enable it with `"defaultTools": ["+tool_search"]` or `--tools`. It uses the same ranking as `searchTools()` over tools that are not declared yet and declares the matches for the next model call. Loaded tools are recorded in the session like other tool changes, so they stay declared on that branch.
185
+
142
186
  <a id="resource-options"></a>
143
187
 
144
188
  ## Resources
@@ -150,9 +194,9 @@ knightcode --extension ./review.ts
150
194
  See [Configuration](configuration.md) for conventional directories and project trust, [Settings](settings.md#resources) for configured paths, and [KnightCode Packages](packages.md) for package sources.
151
195
 
152
196
  - `-e`, `--extension <path>`<br>
153
- Loads an extension file or directory and is repeatable.
197
+ Loads an extension file or directory, or a built-in extension such as `builtin:mcp`, and is repeatable.
154
198
  - `-ne`, `--no-extensions`<br>
155
- Disables discovered and configured extensions. Explicit `-e` paths still load.
199
+ Disables discovered, configured, and built-in extensions. Explicit `-e` paths still load, so `knightcode -ne -e builtin:mcp` keeps only the built-in MCP support.
156
200
  - `--skill <path>`<br>
157
201
  Loads a skill file or directory and is repeatable.
158
202
  - `-ns`, `--no-skills`<br>
@@ -268,3 +312,20 @@ Authentication commands require `--provider <provider>` or `--model <model>`. Se
268
312
  | `--min-expiry <duration>` | `print-bearer-token` | Require remaining token lifetime using `ms`, `s`, `m`, or `h`, such as `30m` |
269
313
 
270
314
  Credential-printing commands write secrets to stdout.
315
+
316
+ ## MCP commands
317
+
318
+ These commands work outside a session, so agents can run them through `bash`. See [MCP Servers](mcp.md).
319
+
320
+ | Command | Description |
321
+ |---|---|
322
+ | `knightcode mcp add <server> [options] -- <command> [args...]` | Add or replace a stdio server in `mcp.json`; `--env KEY=VALUE` (repeatable) and `--cwd <dir>` set its environment and working directory. Arguments after the command are passed to it |
323
+ | `knightcode mcp add <server> [options] --url <url>` | Add or replace a streamable HTTP server; `--header KEY=VALUE` (repeatable), `--bearer-token-env-var <NAME>` (sends `Authorization: Bearer ${NAME}`), `--oauth-client-id`, `--oauth-client-secret`, and `--oauth-callback-port` configure authentication |
324
+ | `knightcode mcp remove <server>` | Remove a server from `mcp.json`; stored OAuth credentials are kept |
325
+ | `knightcode mcp list [--json]` | Connect to every enabled server and print its state, tools, and errors; exit with `1` when a config entry is invalid or an enabled server is not connected |
326
+ | `knightcode mcp login <server> [--timeout <seconds>]` | Sign in to an OAuth server: open the authorization page and wait for the browser (default 300 seconds); a terminal also accepts the pasted redirect URL |
327
+ | `knightcode mcp logout <server>` | Delete the stored OAuth credentials of a server |
328
+
329
+ `add` and `remove` change `~/.knightcode/agent/mcp.json`, or `.knightcode/mcp.json` in the current directory with `--local` (`-l`). `add` also takes `--exposure <mode>` (see [Exposure](mcp.md#exposure)) and does not connect; run `knightcode mcp list` to check the server.
330
+
331
+ Project `.knightcode/mcp.json` files are only read for projects that are already trusted.
@@ -12,6 +12,7 @@ The agent directory is shown as `<agent-dir>` below. Set its location with the `
12
12
  |---|---|
13
13
  | `<agent-dir>/settings.json` | User-level [settings](settings.md), including preferences, defaults, resource paths, and KnightCode package declarations. |
14
14
  | `<agent-dir>/keybindings.json` | Custom terminal UI and application [keybindings](keybindings.md). |
15
+ | `<agent-dir>/mcp.json` | [MCP servers](mcp.md) available in every project. |
15
16
  | `<agent-dir>/models.json` | [Compatible endpoints, models, and model overrides](models.md#configure-a-compatible-endpoint). |
16
17
  | `<agent-dir>/auth.json` | Saved API keys and OAuth credentials. |
17
18
  | `<agent-dir>/tools.json` | [Web tool](usage.md#web-tools), scratchpad, and classifier gate settings written by `/tools`, including the Brave Search key. |
@@ -28,6 +29,7 @@ The agent directory is shown as `<agent-dir>` below. Set its location with the `
28
29
  | Path | Responsibility |
29
30
  |---|---|
30
31
  | `.knightcode/settings.json` | Project-level [settings](settings.md), resource paths, and KnightCode package declarations. |
32
+ | `.knightcode/mcp.json` | Project [MCP servers](mcp.md). |
31
33
  | `.knightcode/SYSTEM.md` | Replaces the system prompt for the project. |
32
34
  | `.knightcode/APPEND_SYSTEM.md` | Adds project-specific instructions to the system prompt. |
33
35
  | `.knightcode/extensions/` | Project extensions. |
@@ -95,6 +95,10 @@
95
95
  {
96
96
  "title": "Use KnightCode Packages",
97
97
  "path": "packages.md"
98
+ },
99
+ {
100
+ "title": "Connect MCP Servers",
101
+ "path": "mcp.md"
98
102
  }
99
103
  ]
100
104
  },
@@ -109,6 +113,10 @@
109
113
  "title": "Add Custom Providers",
110
114
  "path": "custom-provider.md"
111
115
  },
116
+ {
117
+ "title": "Route with Virtual Models",
118
+ "path": "virtual-models.md"
119
+ },
112
120
  {
113
121
  "title": "Build Terminal UI Components",
114
122
  "path": "tui.md"
@@ -80,6 +80,8 @@ Automatic retries, recovery, compaction, or queued work can continue afterward.
80
80
  | Persist non-context session data | `knightcode.appendEntry()` |
81
81
  | Change active tools, model, or thinking level | Session control methods on `knightcode` |
82
82
  | Add a model provider | `knightcode.registerProvider()` |
83
+ | Add an MCP server | `knightcode.registerMcpServer()` |
84
+ | Route each request to a model | [`knightcode.registerVirtualModel()`](virtual-models.md) |
83
85
  | Add terminal rendering | Renderer registration and `ctx.ui` |
84
86
  | Communicate with another extension | `knightcode.events` |
85
87
 
@@ -141,14 +143,61 @@ Use sequential execution when tools share mutable in-memory state.
141
143
  File-mutating tools should wrap the complete read-modify-write operation with `withFileMutationQueue()`.
142
144
  Truncate large model-facing results and tell the model where to read the complete output.
143
145
 
146
+ Declare `outputSchema` and return a matching `structuredContent` when the result is data. The model still receives `content`; programmatic callers such as codemode scripts receive `structuredContent` instead of the text. Tools without `outputSchema` are passed to scripts as their text content. To report a failure that still carries data, return the result with `isError: true` instead of throwing: the model sees an error, and scripts still receive `structuredContent`.
147
+
148
+ A tool can run other tools with `ctx.executeTool(name, args, { signal, onUpdate })`. Nested calls go through argument validation and the `tool_call` and `tool_result` handlers like model-issued calls, and emit `tool_execution_start`, `tool_execution_update`, and `tool_execution_end`; all of these events carry `parentToolCallId`, and their `toolCallId` is assigned by knightcode as `<parent id>/<n>`. These ids do not appear as tool calls or tool results in the transcript. Nested calls do not add transcript entries: their results only reach the calling tool, which reports them itself, for example through `onUpdate` and `details`. The session keeps a bounded record of them (name, arguments, status, duration, error; never results) as `nestedCalls` on the calling tool's result message. It is used for compaction file lists and shown in HTML exports. Arguments over 8 KiB per call or 32 KiB per tool result are omitted, at most 256 calls are kept, and `complete: false` marks a record that lost anything. The `usage` of nested results, at every depth, is added to the calling tool's result `usage`, so a tool reports only its own usage, not that of the tools it called. `ctx.tools` lists the tools `ctx.executeTool()` can call. `tool_result` handlers that redact `content` should also replace `structuredContent`; replacing only `content` drops it.
149
+
144
150
  See [`hello.ts`](../examples/extensions/hello.ts), [`todo.ts`](../examples/extensions/todo.ts), [`dynamic-tools.ts`](../examples/extensions/dynamic-tools.ts), and [`truncated-tool.ts`](../examples/extensions/truncated-tool.ts).
145
151
 
152
+ ### Tool exposure
153
+
154
+ `exposure` controls how the model reaches a tool. "Callable" means callable from other tools through `ctx.executeTool()` (`ctx.tools`), as the `codemode` tool's scripts do:
155
+
156
+ - `direct` (default): declared to the model while active, and callable while active.
157
+ - `model-only`: declared to the model while active, never callable. Use it for tools that orchestrate other tools or ask the user.
158
+ - `codemode`: callable whenever registered, and listed by the `codemode` tool. Not declared to the model unless activated explicitly.
159
+ - `deferred`: like `codemode`, but codemode tools do not list it; `tool_search` can find and activate it.
160
+ - `hidden`: registered but unreachable. Re-register a tool with `exposure: "hidden"` to withdraw it, since tools cannot be unregistered.
161
+
162
+ `namespace: { name, description }` groups related tools, as MCP servers do. Codemode tools list a namespace under one heading.
163
+
164
+ Registering a `direct` or `model-only` tool activates it; the other exposures are not activated on registration. The active set (`knightcode.getActiveTools()`, `knightcode.setActiveTools()`) is the set of tools declared to the model. `knightcode.getAllTools()` reports each tool's `exposure`, `namespace`, and `annotations`.
165
+
166
+ `annotations` are hints about what a tool does, with the meaning of MCP tool annotations: `readOnlyHint`, `destructiveHint`, `idempotentHint`, and `openWorldHint`. MCP tools carry the hints their server declares. Missing hints take the MCP defaults: a tool is not read-only, and may be destructive and reach an open world. The hints are not verified, but a permission extension can use them to decide which calls to confirm. This confirms the calls Codex asks approval for:
167
+
168
+ ```typescript
169
+ knightcode.on("tool_call", async (event, ctx) => {
170
+ const hints = knightcode.getAllTools().find((tool) => tool.name === event.toolName)?.annotations;
171
+ const needsApproval =
172
+ hints?.destructiveHint === true ||
173
+ (!hints?.readOnlyHint && ((hints?.destructiveHint ?? true) || (hints?.openWorldHint ?? true)));
174
+ if (needsApproval && !(await ctx.ui.confirm("Allow tool call?", event.toolName))) {
175
+ return { block: true, reason: `${event.toolName} was not approved` };
176
+ }
177
+ });
178
+ ```
179
+
180
+ A tool that orchestrates other tools can adjust what the model sees while it is active with `prepareLoadout(loadout)`. It runs whenever the active tools change and receives the declared tools, the callable tools, and every registered tool with its exposure and namespace. It returns replacement `descriptions` for declared tools (including its own) and `hiddenDeclarations`: active tools whose declarations requests leave out while they stay active and callable. `codemode` and `tool_search` use only this hook, `exposure`, and `ctx.executeTool()`, so another tool can implement the same behavior under a different name.
181
+
146
182
  ### Activate tools dynamically
147
183
 
148
184
  Register every tool first, keep optional tools inactive, and use `knightcode.setActiveTools()` from a loader tool to select the desired active tools. Names must already be registered; unknown names are ignored.
149
185
 
150
186
  KnightCode records the initial prompt and tool set in the transcript's first system message, then appends tool and prompt changes before the next model request. Providers that cannot represent the transition receive a complete transcript checkpoint, which can invalidate the cached prefix.
151
187
 
188
+ ### MCP servers
189
+
190
+ `knightcode.registerMcpServer(name, config)` adds an MCP server for the current session. `config` has the shape of an `mcpServers` entry in [`mcp.json`](mcp.md): `command`, `args`, `env`, and `cwd` for stdio servers, `url`, `headers`, and `oauth` for HTTP servers, plus `exposure`, `toolExposure`, `enabled`, and `timeout`.
191
+
192
+ ```typescript
193
+ knightcode.registerMcpServer("jira", { url: "https://mcp.example.com/jira", exposure: "codemode" });
194
+ knightcode.unregisterMcpServer("jira");
195
+ ```
196
+
197
+ Servers registered while the extension loads connect when the session starts, together with the `mcp.json` servers; servers registered later connect right away, and `knightcode.unregisterMcpServer()` closes the connection and makes the server's tools unreachable. Registrations are not saved: register again on every load, for example based on the extension's own settings. A server in `mcp.json` with the same name takes precedence, and `/mcp` shows the override. Registering the same name again replaces the extension's earlier registration; names registered by another extension, invalid names, and invalid configs throw.
198
+
199
+ The built-in MCP support connects registered servers. When nothing does, because another extension replaced it (see [MCP](mcp.md#other-mcp-extensions)), each registration is reported as an extension error. Other MCP extensions can connect registered servers too: read them with `knightcode.getMcpServers()` on `session_start` and handle the `mcp_servers_change` event for later changes.
200
+
152
201
  <a id="extensioncontext"></a>
153
202
  <a id="extensioncommandcontext"></a>
154
203
  <a id="use-extension-context"></a>
@@ -194,7 +243,10 @@ Use `ctx.ui.custom()` only when the interaction needs its own rendering and inpu
194
243
  See [Terminal UI](tui.md) for component, focus, overlay, theme, and performance guidance.
195
244
 
196
245
  Extensions load in interactive, RPC, JSON, and print modes.
197
- Interactive mode provides the complete terminal UI.
246
+ Interactive mode provides the complete terminal UI. The [DOOM overlay example](../examples/extensions/doom-overlay/) uses a custom overlay component to render a game frame by frame:
247
+
248
+ <p align="center"><img src="images/doom-extension.png" alt="The DOOM overlay example running over a KnightCode session" width="750"></p>
249
+
198
250
  RPC can forward supported dialogs and notifications through the [RPC Extension UI protocol](rpc-extension-ui.md), but not custom terminal components; JSON and print modes have no UI.
199
251
  Guard terminal-only behavior with `ctx.mode === "tui"` and use `ctx.hasUI` for interactions supported by interactive and RPC clients.
200
252
 
Binary file
Binary file
@@ -125,7 +125,7 @@ In fullscreen mode, these actions control the transcript and take precedence ove
125
125
  | `app.exit` | `ctrl+d` | Exit (when editor empty) |
126
126
  | `app.suspend` | `ctrl+z` (None on Windows) | Suspend to background |
127
127
  | `app.editor.external` | `ctrl+g` | Open in external editor (`externalEditor`, `$VISUAL`, `$EDITOR`, Notepad on Windows, or `nano` elsewhere) |
128
- | `app.clipboard.pasteImage` | `ctrl+v` (`alt+v` on Windows and WSL) | Paste image or text from clipboard |
128
+ | `app.clipboard.pasteImage` | `ctrl+v` (`alt+v` on Windows and WSL) | Paste files on macOS, images, or text from clipboard |
129
129
 
130
130
  On native Windows, `app.suspend` has no default because Windows terminals do not support Unix job control. If you assign it manually, KnightCode shows a status message instead of suspending. WSL uses the normal `ctrl+z` and `fg` behavior.
131
131
 
@@ -86,6 +86,17 @@ Loaded and sleeping models appear in `/model`. Sleeping models wake automaticall
86
86
 
87
87
  If the router disconnects, `/llama` shows **Retry** and **Close**. Retry reconnects and refreshes model state without replaying the interrupted operation.
88
88
 
89
+ ## Classification
90
+
91
+ Every model listed for chat is also listed as a classifier model with the same ID and the `llama-cpp-classify` API. Classifier models answer typed `choice`, `bool`, and `score` questions about JSON state, like TypeSafe's Jev models. The model reaches them from [`codemode`](cli.md#enable-codemode) scripts, and extensions through `ctx.modelRegistry.classify()`; see [Classifier models](models.md#use-classifier-models).
92
+
93
+ The model does not generate an answer. Each question becomes one chat prompt: the state, every question of the request, the state again, and then the question with its answers under single-token labels. Labels are letters for a choice (up to 62 options), `Yes`/`No` for a bool, and digits for a score (up to 10 levels). The second copy of the state is read with the questions in view, which improved accuracy on JevBench with small models. KnightCode reads the probabilities of the labels as the next token and normalizes them. A choice returns every option's probability and a confidence of `(n * peak - 1) / (n - 1)`; a score returns the expected level.
94
+
95
+ - Raw label probabilities are usually overconfident. The per-request `temperature` option divides the label logits before normalizing; values above 1 soften the distribution. It changes no answer.
96
+ - Questions run one after another. Everything before the final question is the same for all questions of a request, so the server's prompt cache evaluates it once. The state appears twice, so it needs twice its size in context.
97
+ - Small models may follow instructions written inside the state. The prompt tells the model to judge the state as data, but that is not a guarantee.
98
+ - Hybrid models such as Qwen3.5 cannot rewind a partially cached prompt without context checkpoints. If each question reprocesses the whole state, start the router with `--ctx-checkpoints 32 --checkpoint-min-step 0`.
99
+
89
100
  ## Troubleshooting
90
101
 
91
102
  Check that the router is reachable:
@@ -99,3 +110,5 @@ curl http://127.0.0.1:8080/models
99
110
  - **Model missing from `/model` with `--no-models-autoload`:** Load it with `/llama` first.
100
111
  - **Load fails or uses too much memory:** Lower `-c` or unload another model.
101
112
  - **Server is not in router mode:** Start it without `--model`, `-m`, or `-hf`.
113
+
114
+ To remove the `llama.cpp` provider and `/llama`, disable `llama.cpp` under Built-in in `knightcode config`, or set `"extensions": ["-builtin:llama.cpp"]` in [settings](settings.md#resources).
@@ -0,0 +1,188 @@
1
+ # MCP Servers
2
+
3
+ KnightCode connects to [Model Context Protocol](https://modelcontextprotocol.io) servers over stdio or streamable HTTP and makes their tools available to the model.
4
+
5
+ ## Configure servers
6
+
7
+ Add servers to `~/.knightcode/agent/mcp.json`, or to `.knightcode/mcp.json` in a project. The format matches other MCP clients, so existing `mcpServers` entries can be copied over:
8
+
9
+ ```json
10
+ {
11
+ "mcpServers": {
12
+ "filesystem": {
13
+ "command": "npx",
14
+ "args": ["-y", "@modelcontextprotocol/server-filesystem", "."]
15
+ },
16
+ "docs": {
17
+ "url": "https://example.com/mcp",
18
+ "headers": { "Authorization": "Bearer ${DOCS_TOKEN}" },
19
+ "exposure": "direct"
20
+ }
21
+ }
22
+ }
23
+ ```
24
+
25
+ - stdio servers take `command`, `args`, `env`, and `cwd`. Relative `cwd` resolves against the session directory. A leading `~/` in `command`, an argument, or `cwd` names the home directory.
26
+ - HTTP servers take `url`, `headers`, and `oauth` (see [Sign in with OAuth](#sign-in-with-oauth)). The legacy SSE transport is not supported.
27
+ - `env` and `headers` values can reference environment variables (`${NAME}`) or commands (`!command`), like provider API keys.
28
+ - `timeout` sets the per-request timeout in seconds (default 60). Progress notifications from the server reset it.
29
+ - `enabled: false` keeps an entry without connecting to it.
30
+
31
+ Project entries replace global entries with the same name. A project `mcp.json` is only read after the project is trusted, because stdio servers run commands.
32
+
33
+ `knightcode mcp add` and `knightcode mcp remove` edit the file from a shell (see [MCP commands](cli.md#mcp-commands)):
34
+
35
+ ```bash
36
+ knightcode mcp add filesystem -- npx -y @modelcontextprotocol/server-filesystem .
37
+ knightcode mcp add docs --url https://example.com/mcp --bearer-token-env-var DOCS_TOKEN --exposure direct
38
+ knightcode mcp add -l tools --env API_KEY='${TOOLS_KEY}' -- uvx tools-mcp
39
+ knightcode mcp remove docs
40
+ ```
41
+
42
+ Rules that are easy to get wrong:
43
+
44
+ - Server names may only contain letters, digits, `_`, and `-`. Tools are named `mcp__<server>__<tool>`.
45
+ - `type` is optional: a `command` makes a stdio server and a `url` a streamable HTTP server. When present, it must be `stdio`, `http`, or `streamable-http`. `sse` is rejected; most servers that document an SSE endpoint also serve streamable HTTP, often at `/mcp` instead of `/sse`.
46
+ - `command` is a single executable and `args` its arguments, not one shell string.
47
+ - Keep secrets out of the file: use `${NAME}` for environment variables, as in `"Authorization": "Bearer ${GITHUB_TOKEN}"`, or `!command` to run a command. A command must make up the whole value, so it has to print the header value itself: `"Authorization": "!echo Bearer $(gh auth token)"`.
48
+ - Invalid entries are skipped and reported; the other servers still connect.
49
+
50
+ ## Set up servers
51
+
52
+ When asked to add an MCP server, the agent should:
53
+
54
+ 1. Add simple servers with `knightcode mcp add` (add `-l` for the project file), or edit `mcp.json` directly for settings the command does not cover. Put personal servers and servers with credentials in `~/.knightcode/agent/mcp.json`. Use the project `.knightcode/mcp.json` only for servers the project itself needs, and only in trusted projects.
55
+ 2. Convert entries written for other clients:
56
+ - Claude Desktop, Claude Code, and Cursor use the same `mcpServers` shape; copy the entry.
57
+ - VS Code uses a top-level `servers` object and `inputs` prompts; move the entry under `mcpServers` and replace `${input:...}` with `${NAME}` environment variables.
58
+ - Codex uses TOML (`[mcp_servers.<name>]` with `command`, `args`, `env`, or `url`); write the same fields as JSON.
59
+ - opencode uses `"type": "local"` with `command` as an array (split it into `command` and `args`), `"type": "remote"` for URLs, `environment` for `env`, and `{env:NAME}` for `${NAME}`.
60
+ 3. Run `knightcode mcp list` to check the entry. It connects to every enabled server and prints the state, the tools, and errors such as the stderr of a stdio server that failed to start. It exits with 1 while anything is wrong.
61
+ 4. For a server that needs a sign-in, run `knightcode mcp login <server>`. It opens the authorization page in the user's browser and waits until the user approves access; tell the user to approve it. A running session uses the new credentials on its next turn.
62
+ 5. Tell the user to run `/reload` (or start a new session) so the running session connects to added or changed servers.
63
+
64
+ KnightCode connects when a session starts. The first prompt waits up to 10 seconds for startup connections; the tools of servers that take longer become available once they connect. HTTP connections that fail with a network error or a transient status (408, 429, 5xx) are retried twice. A server that drops its connection shows as disconnected and is reconnected on the next call. When a server announces that its tool list changed, new tools are added and withdrawn tools become unreachable until the server offers them again.
65
+
66
+ Config errors, servers that failed to connect, and servers that need a sign-in are reported once after startup.
67
+
68
+ Log messages servers send with MCP logging notifications are appended to `~/.knightcode/agent/mcp.log` as `<time> [<server>] <level> <logger>: <message>`. The file is moved to `mcp.log.1` when it grows past 5 MB.
69
+
70
+ ## Manage servers
71
+
72
+ `/mcp` opens the server manager. It lists every configured server with its state, tool count, exposure, and whether it comes from the global or the project `mcp.json`; servers that need attention come first. Select a server to:
73
+
74
+ - sign in, for OAuth servers that need it (see [Sign in with OAuth](#sign-in-with-oauth))
75
+ - see its tools, its command or URL, and the full connection error, including the tail of a stdio server's stderr
76
+ - reconnect
77
+ - sign out, which deletes the stored OAuth credentials
78
+ - change its exposure (see [Exposure](#exposure))
79
+ - disable or enable it
80
+
81
+ Exposure changes and enabling or disabling are saved to the `mcp.json` that defines the server; other content of the file is kept. Disabled servers stay listed so they can be enabled again.
82
+
83
+ Outside the interactive TUI, `/mcp` prints the server status. `/mcp login <server>`, `/mcp logout <server>`, and `/mcp reconnect <server>` run those actions directly.
84
+
85
+ From a shell, `knightcode mcp add`, `knightcode mcp remove`, `knightcode mcp list`, `knightcode mcp login <server>`, and `knightcode mcp logout <server>` manage servers without a session (see [MCP commands](cli.md#mcp-commands)).
86
+
87
+ Stopping a stdio server closes its stdin, then sends SIGTERM and finally SIGKILL to its whole process group, so servers started through wrappers such as `npx` or `uvx` do not linger.
88
+
89
+ ## Sign in with OAuth
90
+
91
+ Remote servers that use OAuth, such as Sentry, need no credentials in `mcp.json`:
92
+
93
+ ```json
94
+ {
95
+ "mcpServers": {
96
+ "sentry": { "url": "https://mcp.sentry.dev/mcp" }
97
+ }
98
+ }
99
+ ```
100
+
101
+ When such a server rejects the connection, `/mcp` shows it as needing sign-in. Select it and choose "Sign in" (or run `/mcp login sentry`, or `knightcode mcp login sentry` in a shell) to open the authorization page in your browser. After you approve access, the browser redirects to a temporary server on `127.0.0.1` and knightcode connects. If the browser runs on another machine, for example over SSH, paste the URL it was redirected to into the sign-in screen instead.
102
+
103
+ KnightCode registers itself with the authorization server (dynamic client registration), stores tokens in `~/.knightcode/agent/mcp-auth.json`, and refreshes access tokens automatically when they expire or the server rejects them. If the server later asks for more scope than was granted, it shows as needing sign-in again, and signing in requests the new scope. "Sign out" in `/mcp` (or `/mcp logout sentry`) deletes the stored credentials.
104
+
105
+ OAuth applies to HTTP servers without an `Authorization` header. For authorization servers that do not support dynamic client registration, configure a pre-registered client:
106
+
107
+ ```json
108
+ {
109
+ "mcpServers": {
110
+ "example": {
111
+ "url": "https://mcp.example.com/mcp",
112
+ "oauth": { "clientId": "my-client", "clientSecret": "${EXAMPLE_SECRET}", "callbackPort": 8765 }
113
+ }
114
+ }
115
+ }
116
+ ```
117
+
118
+ The redirect URI must match the one registered for the client. `callbackPort` fixes it to `http://127.0.0.1:<port>/callback`. For another redirect URI, set `callbackUrl`, for example `"callbackUrl": "http://localhost:8080/oauth/callback"`. It must be an `http` URI on `localhost`, `127.0.0.1`, or `[::1]`, and is sent exactly as written. Without a port in `callbackUrl`, knightcode listens on `callbackPort`, or on a free port, and adds it to the URI; authorization servers accept any port for loopback redirects (RFC 8252). `clientSecret` is optional and can reference environment variables or commands.
119
+
120
+ `scope` sets the scopes to request, separated by spaces, for servers that do not advertise the ones they need. Without it, knightcode requests the scopes the server advertises. When a server later asks for more scope, knightcode requests those on top of `scope`.
121
+
122
+ ## Exposure
123
+
124
+ Each server's tools are registered as `mcp__<server>__<tool>`. The `exposure` setting controls how the model reaches them:
125
+
126
+ - `codemode` (default): the tools are callable from [`codemode`](cli.md#tools) scripts and listed in the `codemode` tool's description, but are not declared to the model. Large MCP tool lists stay out of the model's tool declarations, and scripts can call several MCP tools, in parallel if needed, while returning only the part of the result the model needs. KnightCode activates the `codemode` tool when such a server connects. Large servers do not fill the description: declarations share a token budget, and scripts find the remaining tools with `searchTools()` (see [`codemode`](cli.md#tools)).
127
+ - `codemode-deferred`: like `codemode`, but the tools are not listed in the `codemode` tool's description either; it only names the server and its tool count. Scripts call them by name and find them with `searchTools()` or in `ALL_TOOLS`. Use it for large servers that codemode scripts use rarely.
128
+ - `deferred`: the tools are not declared to the model until the [`tool_search`](cli.md#tools) tool loads them. The model searches, and the matches are declared from its next call on and called directly, without codemode. KnightCode activates the `tool_search` tool when such a server connects. Use it for large servers without codemode.
129
+ - `direct`: the tools are declared to the model like built-in tools, and are also callable from codemode.
130
+ - `hidden`: the tools are registered but cannot be called.
131
+
132
+ `toolExposure` sets the exposure of single tools and overrides `exposure` for them. Keys are tool names as the server offers them, or patterns where `*` matches any characters. An exact name wins over patterns; among patterns, the first match in the object wins. With `hidden` as the server's exposure, only the listed tools are reachable:
133
+
134
+ ```json
135
+ {
136
+ "mcpServers": {
137
+ "github": {
138
+ "url": "https://api.githubcopilot.com/mcp/",
139
+ "exposure": "deferred",
140
+ "toolExposure": {
141
+ "search_code": "direct",
142
+ "get_*": "codemode",
143
+ "delete_*": "hidden"
144
+ }
145
+ }
146
+ }
147
+ }
148
+ ```
149
+
150
+ `knightcode mcp list` marks tools whose exposure differs from the server's, and the Tools view in `/mcp` shows it too.
151
+
152
+ Tools that are not declared (`codemode`, `codemode-deferred`, and `deferred` exposure) are reachable through either tool: codemode scripts can call all of them, and `tool_search` can load any of them. For example, with `codemode` active, scripts can call the tools of a `deferred` server, and with `tool_search` active, the model can load the tools of a `codemode` server.
153
+
154
+ Tools called from codemode scripts do not depend on the active tool set, so they stay callable after `/tree`, resume, and fork. Tools loaded by `tool_search` are recorded in the transcript like any other tool change and stay declared on that branch. To keep `codemode` active without MCP servers too, add `"defaultTools": ["+codemode"]` to [settings](settings.md#tools). To keep knightcode from activating the `codemode` tool, set `"autoEnableCodemode": false` at the top level of `mcp.json`, next to `mcpServers`. A project `mcp.json` value overrides the global one. KnightCode warns once when neither `codemode` nor `tool_search` is active, since the tools then cannot be called.
155
+
156
+ Text results over 20KB reach the model with the middle cut out, in the format Codex uses: the start and end of the text around a `…N chars truncated…` marker. The full text is saved to a temp file whose path the result names. Codemode scripts always receive the whole result, so a script can filter a large result down to what the model needs.
157
+
158
+ Codemode scripts receive an MCP tool's whole `CallToolResult` (`content` blocks as sent by the server, `structuredContent`, and `isError`), and the `codemode` description declares it as `CallToolResult<T>`. A result with `isError` resolves in scripts and is reported to the model as an error for direct calls. `image(result.content[0])` forwards an image block to the model. The server's `instructions` describe its tools in the `codemode` description.
159
+
160
+ ## Resources
161
+
162
+ When a connected server offers [resources](https://modelcontextprotocol.io/specification/2025-11-25/server/resources), knightcode adds the resource tools Codex and opencode use:
163
+
164
+ - `list_mcp_resources` lists resources as JSON: `{ server?, resources: [{ server, uri, name, ... }], nextCursor? }`. With `server`, it lists one page of that server, and `cursor` continues with the next one. Without, it lists every resource of every server.
165
+ - `list_mcp_resource_templates` lists URI templates for resources the servers do not list, in the same way.
166
+ - `read_mcp_resource` reads a resource given `server` and `uri`. Text resources reach the model as text and images as images; other binary resources are saved to temp files, and the model sees the file path. Scripts receive `{ server, uri, contents }`.
167
+
168
+ The tools reach every enabled server with resources whose exposure is not `hidden`, and take the widest exposure among them: `direct` if one of the servers is direct, else `codemode`, else `codemode-deferred`, else `deferred`. Resource links in tool results name `read_mcp_resource` and the server.
169
+
170
+ Resources for MCP Apps (`ui://` URIs or `text/html;profile=mcp-app`) are left out of the listings, since knightcode does not render them, and so are resource icons.
171
+
172
+ Reading and listing resources is retried once after a transient HTTP error (408, 429, 5xx). Tool calls are not retried, since the server may have run them.
173
+
174
+ ## Permissions
175
+
176
+ Every MCP call goes through knightcode's tool pipeline, so `tool_call` and `tool_result` extension handlers, including permission gates, apply to MCP tools. Calls made from codemode scripts carry the `codemode` call's id as `parentToolCallId`. `knightcode.getAllTools()` reports the tool annotations servers declare (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`), so a permission extension can confirm only calls that change something (see [Extensions](extensions.md#tool-exposure)). The resource tools are marked read-only.
177
+
178
+ ## Servers from extensions
179
+
180
+ Extensions can add servers for the current session with `knightcode.registerMcpServer(name, config)`, using the same config shape as `mcp.json` (see [Extensions](extensions.md#mcp-servers)). They connect like configured servers and appear in `/mcp` with the extension as their source. Enabling, disabling, and exposure changes for them apply to the current session only. A server in `mcp.json` with the same name takes precedence; `/mcp` lists the overridden registration. `knightcode mcp` shell commands do not load extensions and only see `mcp.json` servers.
181
+
182
+ ## Other MCP extensions
183
+
184
+ An installed extension that registers the `/mcp` command, such as `knightcode-mcp-adapter`, replaces the built-in MCP support: knightcode then neither reads `mcp.json` in sessions nor connects servers, and `/mcp` belongs to that extension. Remove the extension to use the built-in support. To turn off the built-in support without installing another extension, disable `mcp` under Built-in in `knightcode config`, or set `"extensions": ["-builtin:mcp"]` in [settings](settings.md#resources); `knightcode mcp` shell commands still work. Likewise, an extension that registers a tool named `codemode` or `tool_search` replaces the built-in tool of that name. `knightcode mcp` shell commands always use the built-in support.
185
+
186
+ ## SDK
187
+
188
+ SDK sessions do not load the built-in extensions. Add the MCP extension, the codemode extension for `codemode` and `codemode-deferred` servers, and the tool search extension for `deferred` servers to the resource loader. See [SDK](sdk.md#codemode-mcp).
@@ -100,6 +100,41 @@ Choose the conservative end of any published range. A model without a lifetime f
100
100
 
101
101
  Compatibility settings should describe verified differences in the endpoint's request or response behavior. Do not enable them based only on an endpoint advertising OpenAI or Anthropic compatibility.
102
102
 
103
+ ## Use classifier models
104
+
105
+ Classifier models do not chat. They answer typed questions about JSON state: pick one of several choices, answer yes or no, or give a score, each with probabilities. KnightCode includes TypeSafe's Jev model from these providers:
106
+
107
+ | Provider | Model IDs | Authentication |
108
+ |---|---|---|
109
+ | `typesafe` | `jev-latest` | `TYPESAFE_API_KEY` |
110
+ | `openrouter` | `typesafe/jev-1.13`, `~typesafe/jev-latest` | `OPENROUTER_API_KEY` or `/login` |
111
+ | `cloudflare-workers-ai` | `typesafe/jev` | `CLOUDFLARE_API_KEY` and `CLOUDFLARE_ACCOUNT_ID` |
112
+ | `vercel-ai-gateway` | `typesafe-ai/jev` | `AI_GATEWAY_API_KEY` |
113
+ | `opencode` | `jev-1.13`, `jev-1.13-free` | `OPENCODE_API_KEY` |
114
+
115
+ Chat models on a [llama.cpp router](llama-cpp.md#classification) are also listed as classifier models.
116
+
117
+ Classifier models do not appear in `/model`. The [classifier gate](usage.md#classifier-gate) uses one to screen risky tool calls. The model reaches them through the [`codemode`](cli.md#enable-codemode) tool, which is off unless an MCP server turned it on. Enable it with `"defaultTools": ["+codemode"]` in [settings](settings.md#tools). Scripts then list classifier models with `models.getAvailableOfType("classifier")` and call `models.classify(model, { state, questions })`:
118
+
119
+ ```js
120
+ const jev = await models.getModelOfType("classifier", "typesafe", "jev-latest");
121
+ const result = await models.classify(jev, {
122
+ state: { message: "The change works, thanks." },
123
+ questions: {
124
+ approved: {
125
+ type: "bool",
126
+ instructions: "Does the user approve of the result?",
127
+ criteria: { true: "Approval", false: "No approval" },
128
+ },
129
+ },
130
+ });
131
+ return result.answers;
132
+ ```
133
+
134
+ When the service reports token counts, as all System One services do, `result.usage` carries them with their cost. KnightCode adds the usage of a script's classifier calls to the `codemode` tool result, so it counts toward the session cost in the footer and `/session`. The cost uses the model's catalog price; models without one, such as TypeSafe's direct `jev-latest`, report tokens at no cost.
135
+
136
+ Extensions call classifiers through `ctx.modelRegistry.classify()`, without codemode. [Virtual models](virtual-models.md#route-requests) can use them to route requests; see the `jev-router.ts` example.
137
+
103
138
  ## Add a custom provider
104
139
 
105
140
  Use an extension when the provider needs custom streaming, model discovery, or authentication behavior. See [Custom Providers](custom-provider.md) for the extension workflow.
@@ -118,7 +118,7 @@ For each resource type:
118
118
 
119
119
  Filters narrow the package manifest. They do not expose resources that the package itself did not declare.
120
120
 
121
- Run `knightcode config` to enable or disable discovered resources. It starts with personal configuration; press Tab to switch scope, or run `knightcode config --local` to start with project overrides.
121
+ Run `knightcode config` to enable or disable discovered resources and knightcode's built-in extensions. It starts with personal configuration; press Tab to switch scope, or run `knightcode config --local` to start with project overrides.
122
122
 
123
123
  ## Understand scope and identity
124
124
 
@@ -50,6 +50,7 @@ This table covers providers with a single primary API-key variable. Providers th
50
50
  | ZAI Coding Plan (China) | `ZAI_CODING_CN_API_KEY` |
51
51
  | OpenCode Zen and Go | `OPENCODE_API_KEY` |
52
52
  | Radius | `RADIUS_API_KEY` |
53
+ | TypeSafe ([classifier models](models.md#use-classifier-models)) | `TYPESAFE_API_KEY` |
53
54
  | Hugging Face | `HF_TOKEN` |
54
55
  | Fireworks | `FIREWORKS_API_KEY` |
55
56
  | Together AI | `TOGETHER_API_KEY` |
package/bin/docs/sdk.md CHANGED
@@ -109,7 +109,11 @@ Use `DefaultResourceLoader` when you want standard discovery with selected overr
109
109
 
110
110
  <a id="inlineextension"></a>
111
111
 
112
- Inline extension factories can be supplied through `DefaultResourceLoader`. Give one an `InlineExtension` name only when it needs a stable name in diagnostics and startup output.
112
+ Inline extension factories can be supplied through `DefaultResourceLoader`. Give one an `InlineExtension` name only when it needs a stable name in diagnostics and startup output. A named inline extension with `replaceable: true` is left out when another extension registers a tool, command, or flag with a name it registers during loading, instead of both loading with a conflict. The CLI's built-in codemode, tool search, and MCP extensions are replaceable. A named entry with `builtin: true` is not an inline extension: it supplies the code of the `builtin:<name>` extension, which loads like a configured extension file. It loads by default, is listed in `knightcode config`, and is disabled by `-builtin:<name>` in the `extensions` setting or by `noExtensions`; `additionalExtensionPaths: ["builtin:<name>"]` loads it explicitly. It loads after project trust is resolved, so it cannot handle `project_trust`. The CLI's built-in extensions use it.
113
+
114
+ <a id="codemode-mcp"></a>
115
+
116
+ The CLI loads `codemode`, `tool_search`, and MCP as built-in extensions. SDK sessions do not; add `createCodemodeExtension()`, `createToolSearchExtension()`, and `createMcpExtension()` to the `extensionFactories` of `DefaultResourceLoader`. `codemode` and `tool_search` are registered inactive: enable them through the `defaultTools` setting (`["+codemode", "+tool_search"]` keeps the other default tools), or let the MCP extension activate them: `codemode` for servers with `codemode` or `codemode-deferred` exposure, `tool_search` for servers with `deferred` exposure. The MCP extension connects its servers on `session_start`, so call `session.bindExtensions()`. See [Codemode and MCP](../examples/sdk/14-codemode-mcp.ts).
113
117
 
114
118
  See the focused examples for [models](../examples/sdk/02-custom-model.ts), [tools](../examples/sdk/05-tools.ts), [extensions](../examples/sdk/06-extensions.ts), and [full control](../examples/sdk/12-full-control.ts).
115
119
 
@@ -130,6 +134,7 @@ See the focused examples for [models](../examples/sdk/02-custom-model.ts), [tool
130
134
  | [Sessions](../examples/sdk/11-sessions.ts) | Control session persistence and restoration |
131
135
  | [Full control](../examples/sdk/12-full-control.ts) | Replace default discovery and state services |
132
136
  | [Session runtime](../examples/sdk/13-session-runtime.ts) | Replace the active session safely |
137
+ | [Codemode and MCP](../examples/sdk/14-codemode-mcp.ts) | Add the `codemode`, `tool_search`, and MCP extensions |
133
138
 
134
139
  <a id="exports"></a>
135
140
 
@@ -37,6 +37,7 @@ Project trust does not limit what tool calls can access or affect. After KnightC
37
37
  KnightCode requires a project-trust decision when it finds any of these resources from the current working directory:
38
38
 
39
39
  - `.knightcode/settings.json`
40
+ - `.knightcode/mcp.json`
40
41
  - `.knightcode/extensions`, `.knightcode/skills`, `.knightcode/prompts`, or `.knightcode/themes`
41
42
  - `.knightcode/SYSTEM.md` or `.knightcode/APPEND_SYSTEM.md`
42
43
  - project `.agents/skills` in the current directory or an ancestor directory
@@ -46,6 +47,7 @@ A bare `.knightcode` directory does not require project trust.
46
47
  Granting project trust allows KnightCode to load:
47
48
 
48
49
  - project settings
50
+ - project MCP servers from `.knightcode/mcp.json`
49
51
  - extensions, skills, prompt templates, themes, and system-prompt files under `.knightcode`
50
52
  - missing packages configured through project settings
51
53
  - project-local and project-package extensions
@@ -92,9 +92,11 @@ Sessions created before system messages existed have no leading system message;
92
92
  {"type":"message","id":"c3d4e5f6","parentId":"b2c3d4e5","timestamp":"2024-12-03T14:00:03.000Z","message":{"role":"toolResult","toolCallId":"call_123","toolName":"bash","content":[{"type":"text","text":"output"}],"isError":false,"timestamp":1733234403000}}
93
93
  ```
94
94
 
95
+ Assistant messages name the model that produced them. Newer messages also record `thinkingLevel`, the KnightCode thinking level requested for that response.
96
+
95
97
  ### ModelChangeEntry
96
98
 
97
- Emitted when the user switches models mid-session.
99
+ Emitted when the user switches models mid-session. The latest entry is the selected model, which may be a [virtual model](virtual-models.md); assistant messages then name the physical model that answered.
98
100
 
99
101
  ```json
100
102
  {"type":"model_change","id":"d4e5f6g7","parentId":"c3d4e5f6","timestamp":"2024-12-03T14:05:00.000Z","provider":"openai","modelId":"gpt-4o"}
@@ -169,6 +171,8 @@ Extension state persistence. Does NOT participate in LLM context.
169
171
 
170
172
  Use `customType` to identify your extension's entries on reload. Interactive mode can render custom entries via `knightcode.registerEntryRenderer(customType, renderer)`, but they still do not participate in LLM context.
171
173
 
174
+ KnightCode stores [virtual model](virtual-models.md) router state as custom entries with `customType` `knightcode.virtual-model-state` and `data` `{ provider, modelId, state }`.
175
+
172
176
  ### CustomMessageEntry
173
177
 
174
178
  Extension-injected messages that DO participate in LLM context.
@@ -30,6 +30,8 @@ KnightCode stores entries as a tree, so returning to an earlier point does not e
30
30
 
31
31
  In `/tree`, select a user message to put its text back in the editor. Edit and submit it to create another branch. Selecting an assistant response or another entry continues after that entry with an empty editor.
32
32
 
33
+ <p align="center"><img src="images/tree-view.png" alt="The /tree view showing a session with two branches, user messages, assistant replies, and tool calls" width="750"></p>
34
+
33
35
  Selecting a point while the model is responding cancels that response. Navigation cannot proceed while compaction or another tree navigation is still running; wait for it to finish and retry.
34
36
 
35
37
  When you leave a branch, KnightCode can summarize it and attach that summary to the branch you enter. This preserves relevant work from the abandoned path without including every message from it.
@@ -37,9 +37,23 @@ See [Choose a Model](models.md) for model selection and thinking controls.
37
37
 
38
38
  | Setting | Type | Default | Description |
39
39
  |---|---|---|---|
40
- | `defaultTools` | `string[]` | `read`, `bash`, `edit`, `write` | Built-in tools enabled at startup. An empty array disables all built-in tools but not extension or SDK tools. |
40
+ | `defaultTools` | `string[]` | `read`, `bash`, `edit`, `write` | Tools enabled at startup. Plain names replace the defaults; `+name` adds a tool and `-name` removes one. An empty array disables all built-in tools but not extension or SDK tools. |
41
+ | `codemode.mode` | `"on"` \| `"only"` | `"on"` | How the `codemode` tool presents tools while it is active. `on`: declared tools get their `codemode` declaration appended to their description, and `codemode` lists only tools that are not declared (MCP `codemode` exposure). `only`: `codemode` lists every tool scripts can call, and active built-in and extension tools are hidden from the model, so it reaches them through `codemode`. |
42
+ | `codemode.inlineBudget` | number | `3000` | Estimated tokens (characters / 4) the `codemode` tool's description may spend on tool declarations. Tools that do not fit are left out and found with `searchTools()`. `0` lists only namespaces. |
41
43
 
42
- Available built-in tools are `read`, `bash`, `powershell`, `edit`, `write`, `grep`, `find`, and `ls`. CLI tool options override this setting for one invocation. See [Command Line](cli.md#tools).
44
+ Available built-in tools are `read`, `bash`, `powershell`, `edit`, `write`, `grep`, `find`, and `ls`. `defaultTools` can also name `codemode` and `tool_search`, which built-in extensions register inactive, and other extension tools registered inactive.
45
+
46
+ A list of only `+name` and `-name` entries changes the inherited selection instead of replacing it. For example, this enables `codemode` next to the default tools:
47
+
48
+ ```json
49
+ {
50
+ "defaultTools": ["+codemode"]
51
+ }
52
+ ```
53
+
54
+ This replaces `bash` with `powershell` and enables `grep`: `["-bash", "+powershell", "+grep"]`. Project settings apply on top of user settings: a project list with only `+name` and `-name` entries changes the user's selection, and a project list with a plain name replaces it. In one list, plain names form the selection, and `+name` and `-name` then apply in order.
55
+
56
+ CLI tool options override this setting for one invocation; `--tools` does not accept `+name` or `-name`. See [Command Line](cli.md#tools).
43
57
 
44
58
  The [web tools](usage.md#web-tools) `webfetch` and `websearch` are not part of `defaultTools`. `/tools` turns them on and stores their settings, including the search provider and Brave key, in `~/.knightcode/agent/tools.json`.
45
59
 
@@ -81,6 +95,7 @@ See [Compaction Reference](compaction.md) for trigger, summarization, and valida
81
95
  | `fullscreenExitOutput` | `"transcript" \| "resume-hint"` | `"transcript"` | Output printed when fullscreen mode exits. |
82
96
  | `fullscreenScrollbar` | `"auto" \| "always" \| "hidden"` | `"auto"` | Fullscreen transcript scrollbar behavior. |
83
97
  | `fullscreenCopyOnSelect` | boolean | `true` | Copy selected text automatically in fullscreen mode. When disabled, selections stay highlighted and `Ctrl+X` copies the active selection. Has no effect in regular mode. |
98
+ | `fullscreenWheelScrollLines` | `"auto"` \| number | `"auto"` | Lines per mouse-wheel event in fullscreen mode, from 1 to 100. `"auto"` moves one line per event in local macOS terminals, which already accelerate wheel and trackpad input; elsewhere, and over SSH, it speeds up fast wheel spins to at most 6 lines per event. Alt+wheel moves five times as far. |
84
99
  | `editorPaddingX` | number | `0` | Horizontal editor padding from 0 to 3 cells. |
85
100
  | `outputPad` | `0 \| 1` | `1` | Horizontal transcript padding. |
86
101
  | `autocompleteMaxVisible` | number | `5` | Visible autocomplete entries, from 3 to 20. |
@@ -142,6 +157,8 @@ Resource paths in user settings resolve from the agent directory. Paths in proje
142
157
 
143
158
  Resource arrays support glob exclusions with `!pattern`, exact inclusion with `+path`, and exact exclusion with `-path`. KnightCode loads resources listed in both user-level and project settings.
144
159
 
160
+ The built-in extensions are named `builtin:mcp`, `builtin:llama.cpp`, `builtin:codemode`, and `builtin:tool-search` in `extensions`. They load by default; `-builtin:mcp` disables one. A `+builtin:<name>` or `-builtin:<name>` entry in project settings overrides the user setting. `knightcode config` lists them under Built-in. `--no-extensions` disables them too, and `-e builtin:<name>` loads one explicitly.
161
+
145
162
  ## Updates, telemetry, and warnings
146
163
 
147
164
  | Setting | Type | Default | Description |
@@ -0,0 +1,114 @@
1
+ # Virtual Models
2
+
3
+ A virtual model is a selectable model that picks a physical model for each request. Use one to route by task, cost, or conversation state. For example, a router can send quick questions to a small model and hard problems to a large one, while the user selects a single model.
4
+
5
+ Register virtual models from an [extension](extensions.md). They appear in `/model`, `--model`, scoped models, and settings like any other model. A virtual model can be listed under any provider, including one with physical models, such as `openai-codex/auto`.
6
+
7
+ ## Selection and dispatch
8
+
9
+ A virtual model selects a model and a thinking level. A router maps that pair to a physical pair for each request:
10
+
11
+ ```
12
+ selected (virtual model, virtual level) -> dispatched (physical model, physical level)
13
+ jev/auto:low -> anthropic/claude-sonnet-4-5:high
14
+ ```
15
+
16
+ The virtual thinking level is an input to the router. Its meaning is up to the router; it need not correspond to a reasoning budget.
17
+
18
+ KnightCode keeps the two pairs apart:
19
+
20
+ | | Selection | Dispatch |
21
+ |---|---|---|
22
+ | Recorded in | `model_change` and `thinking_level_change` entries | Each assistant message: `provider`, `api`, `model`, `thinkingLevel` |
23
+ | Visible as | `ctx.model`, `ctx.thinkingLevel`, `KNIGHTCODE_MODEL`, `KNIGHTCODE_REASONING_LEVEL`, `/model` | The assistant message of each response |
24
+
25
+ Providers only receive physical models. Assistant messages name the physical model, so replaying a conversation across different physical models works the same as after a manual model switch. Resuming a session restores the virtual selection from its latest `model_change` entry. If the virtual model is no longer registered, KnightCode falls back to the physical model that answered last.
26
+
27
+ In interactive mode, the footer shows the routed model next to the selection, for example `auto • high → gpt-5.6-luna • medium`. `/session` lists the cost for each physical model.
28
+
29
+ Context usage uses the limits of the physical model that produced the latest response, even if that response came before switching to the virtual model. Without such a response, it uses the limits declared on the virtual model, if any. Compaction checks the same limits, and again the limits of the model each request is routed to. If that model's context window is too small for the conversation, KnightCode compacts before sending the request; the route stays as the router chose it.
30
+
31
+ ## Register a virtual model
32
+
33
+ ```typescript
34
+ import type { ExtensionAPI } from "@knightcodeai/cli";
35
+
36
+ export default function (knightcode: ExtensionAPI) {
37
+ knightcode.registerVirtualModel({
38
+ provider: "router",
39
+ id: "auto",
40
+ name: "Auto",
41
+ thinkingLevels: ["low", "high"],
42
+ route(request, ctx) {
43
+ // Tool follow-ups and retries stay on the model that handled the turn.
44
+ const sticky = request.failed ?? request.previous;
45
+ if (request.reason !== "user" && sticky) {
46
+ return { model: sticky.model, thinkingLevel: sticky.thinkingLevel ?? "medium" };
47
+ }
48
+ const id = request.thinkingLevel === "high" ? "claude-sonnet-4-5" : "claude-haiku-4-5";
49
+ return { model: ctx.modelRegistry.find("anthropic", id)!, thinkingLevel: "medium" };
50
+ },
51
+ });
52
+ }
53
+ ```
54
+
55
+ - `provider` is the provider the model is listed under. It can be any provider ID. A provider can list several virtual models next to its physical ones. On a physical provider, the virtual model is available when that provider has credentials. Under an ID that no provider uses, it is always available.
56
+ - `id` must not be the ID of a physical model of that provider. If a catalog refresh later adds a physical model with the same ID, the virtual model hides it.
57
+ - `thinkingLevels` lists the levels offered for selection. It defaults to `["off"]`.
58
+ - `contextWindow` and `maxTokens` are shown before the first response. Unset limits are unknown.
59
+ - `input` lists the input types offered for selection. It defaults to text and images; physical models without image support receive placeholders.
60
+
61
+ Registration follows the same queuing and reload rules as `knightcode.registerProvider()`. Registering the same provider and ID again replaces the virtual model. `knightcode.unregisterVirtualModel(provider, id)` removes it; `knightcode.unregisterProvider()` does not. SDK code can register one without an extension: `modelRuntime.registerVirtualModel(definition)`.
62
+
63
+ ## Route requests
64
+
65
+ `route(request, ctx)` runs before every request made with the virtual model and returns `{ model, thinkingLevel }`. The model can be any physical model in the catalog whose provider has credentials; look it up with `ctx.modelRegistry`. A virtual model cannot route to another virtual model. KnightCode clamps the thinking level to the returned model.
66
+
67
+ | Field | Meaning |
68
+ |---|---|
69
+ | `model`, `thinkingLevel` | The selected virtual model and level |
70
+ | `reason` | Why the request is made, see below |
71
+ | `previous` | Physical model and thinking level of the latest successful response in `messages` |
72
+ | `failed` | For `retry`: physical model, thinking level, and assistant `message` of the failed request, which `messages` no longer contains. The message carries `stopReason` and `errorMessage`. Absent when routing itself failed |
73
+ | `state` | Router state last returned on this session branch, see below |
74
+ | `messages` | The conversation for this request, including system messages |
75
+ | `signal` | Abort signal of the request |
76
+
77
+ | `reason` | Request |
78
+ |---|---|
79
+ | `user` | First request after a message the user wrote, including steering and follow-up messages |
80
+ | `continuation` | Any other request in the agent loop, such as after tool results or extension messages |
81
+ | `retry` | Automatic retry after a failed request, including after compaction for a context overflow |
82
+ | `direct` | Request made outside the agent loop, such as a compaction summary or an extension calling `ctx.modelRegistry.streamSimple()` |
83
+
84
+ Returning `previous` for `continuation` and `failed` for `retry` keeps prompt caches and thinking signatures valid. Switching models between turns is allowed but loses the prompt cache. A retry can also switch to another model, for example when `failed.message.errorMessage` reports that a provider is overloaded or the context overflowed.
85
+
86
+ If `route()` throws, or returns a virtual model or a model without credentials, the request ends with an error response.
87
+
88
+ ## Keep routing state
89
+
90
+ `route()` can return `state` next to the model. KnightCode stores it on the session branch and passes it back as `request.state` on later requests. Use it for decisions the transcript does not record, such as classifier results or a routing phase:
91
+
92
+ ```typescript
93
+ knightcode.registerVirtualModel<{ phase: "plan" | "build" }>({
94
+ provider: "router",
95
+ id: "phased",
96
+ name: "Phased",
97
+ route(request, ctx) {
98
+ const state = request.state ?? { phase: "plan" };
99
+ const id = state.phase === "plan" ? "claude-opus-4-5" : "claude-haiku-4-5";
100
+ return { model: ctx.modelRegistry.find("anthropic", id)!, thinkingLevel: "medium", state };
101
+ },
102
+ });
103
+ ```
104
+
105
+ - State must be JSON-serializable. Returning `undefined` or `request.state` itself keeps the current state.
106
+ - KnightCode stores any other returned object as new state, before the request is sent, even when it equals the current state. Return a new object only when the state changes. The state stays stored if the request later fails.
107
+ - State follows the session tree, so forks and `/tree` navigation see the state of their branch. It survives compaction.
108
+ - `direct` requests have no state, and KnightCode ignores state they return.
109
+
110
+ The transcript already records the selection and every dispatched model, and `ctx.sessionManager.getBranch()` exposes both.
111
+
112
+ Routers can call other models through `ctx.modelRegistry`, for example `ctx.modelRegistry.classify()` with a classifier model from `ctx.modelRegistry.findOfType("classifier", provider, id)`. The call adds latency before the first token of the turn.
113
+
114
+ See [`jev-router.ts`](../examples/extensions/jev-router.ts) for a complete router. It plans on a strong OpenAI Codex model chosen by the Jev classifier, lets that model make the first edit, and then switches once to a cheaper model, accepting a single prompt-cache miss. It keeps the phase as router state.
@@ -42,6 +42,8 @@ To replace the model-facing `bash` tool with `powershell`, add this to `~/.knigh
42
42
  }
43
43
  ```
44
44
 
45
+ `["-bash", "+powershell"]` does the same while keeping any other default tools you configured.
46
+
45
47
  Restart KnightCode, then ask it to run a harmless PowerShell command. The `!` and `!!` editor commands continue to use Bash. The `powershell` tool is available only when KnightCode runs as a native Windows process.
46
48
 
47
49
  See [Settings](settings.md#tools) for other tool combinations.
@@ -930,6 +930,21 @@
930
930
  '</div>';
931
931
  };
932
932
 
933
+ // Calls this tool made to other tools (for example from a codemode script), recorded without results.
934
+ const renderNestedCalls = () => {
935
+ const nested = result?.nestedCalls;
936
+ if (!nested || !Array.isArray(nested.calls) || nested.calls.length === 0) return '';
937
+ const icons = { ok: '✓', error: '✗', unfinished: '…' };
938
+ const lines = nested.calls.map(c => {
939
+ const args = c.arguments ? JSON.stringify(c.arguments) : `[arguments omitted, ${c.argumentsBytes} bytes]`;
940
+ const duration = c.durationMs !== undefined ? ` ${c.durationMs}ms` : '';
941
+ const error = c.error ? `\n ${c.error.split('\n').join('\n ')}` : '';
942
+ return `${icons[c.status] || '?'} ${c.name} ${args}${duration}${error}`;
943
+ });
944
+ const title = `Nested calls: ${nested.calls.length}${nested.complete ? '' : ' (incomplete record)'}`;
945
+ return formatExpandableOutput([title, ...lines].join('\n'), 1);
946
+ };
947
+
933
948
  const toolDomId = `tool-call-${escapeHtml(call.id)}`;
934
949
  let html = `<div class="tool-execution ${statusClass}" id="${toolDomId}">`;
935
950
  const args = call.arguments || {};
@@ -1063,6 +1078,7 @@
1063
1078
  }
1064
1079
  }
1065
1080
 
1081
+ html += renderNestedCalls();
1066
1082
  html += '</div>';
1067
1083
  return html;
1068
1084
  }
package/bin/knightcode CHANGED
@@ -1,4 +1,4 @@
1
1
  [diffend] Oversized file quarantined before diffing.
2
2
  name: package/bin/knightcode
3
- size: 80044368 bytes
4
- sha256: 0adea2003ebba3c8d3fb42f75c4506c8a143fa9fb5c4c8c240982b40138690c5
3
+ size: 81125712 bytes
4
+ sha256: 9811e8dab12686f35d348c62ee99354269b169757bc41bf7ab52b6e42c672613
package/bin/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@knightcodeai/cli",
3
- "version": "0.10.0",
3
+ "version": "0.11.1",
4
4
  "description": "KnightCode — a local, BYOK terminal coding agent powered by OpenRouter.",
5
5
  "type": "module",
6
6
  "repository": {
@@ -37,17 +37,19 @@
37
37
  "test": "vitest --run"
38
38
  },
39
39
  "optionalDependencies": {
40
- "@knightcodeai/cli-linux-x64": "0.10.0",
41
- "@knightcodeai/cli-linux-arm64": "0.10.0",
42
- "@knightcodeai/cli-darwin-x64": "0.10.0",
43
- "@knightcodeai/cli-darwin-arm64": "0.10.0",
44
- "@knightcodeai/cli-win32-x64": "0.10.0"
40
+ "@knightcodeai/cli-linux-x64": "0.11.1",
41
+ "@knightcodeai/cli-linux-arm64": "0.11.1",
42
+ "@knightcodeai/cli-darwin-x64": "0.11.1",
43
+ "@knightcodeai/cli-darwin-arm64": "0.11.1",
44
+ "@knightcodeai/cli-win32-x64": "0.11.1"
45
45
  },
46
46
  "devDependencies": {
47
47
  "@agentclientprotocol/sdk": "1.4.0",
48
48
  "@knightcode/agent": "workspace:*",
49
49
  "@knightcode/ai": "workspace:*",
50
50
  "@knightcode/client": "workspace:*",
51
+ "@knightcode/codemode": "workspace:*",
52
+ "@knightcode/mcp": "workspace:*",
51
53
  "@knightcode/protocol": "workspace:*",
52
54
  "@knightcode/remote": "workspace:*",
53
55
  "@knightcode/tools": "workspace:*",
@@ -70,6 +72,7 @@
70
72
  "marked": "18.0.11",
71
73
  "minimatch": "10.2.6",
72
74
  "proper-lockfile": "4.1.2",
75
+ "quickjs-wasi": "3.6.2",
73
76
  "semver": "7.8.5",
74
77
  "typebox": "1.3.27",
75
78
  "undici": "8.10.2",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@knightcodeai/cli-darwin-x64",
3
- "version": "0.10.0",
3
+ "version": "0.11.1",
4
4
  "license": "MIT",
5
5
  "repository": {
6
6
  "type": "git",