@knightcodeai/cli-linux-arm64 0.7.0 → 0.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/bin/CHANGELOG.md CHANGED
@@ -1,5 +1,91 @@
1
1
  # @knightcodeai/cli
2
2
 
3
+ ## 0.9.0
4
+
5
+ ### Added
6
+
7
+ - Added unsubscribe to extension events: `knightcode.on(event, handler)` now returns a function that removes that one registration, so an extension can listen once or stop listening without reloading. Removing a handler never affects a dispatch already in progress.
8
+
9
+ - Added what the desktop IDE's first run needs to `knightcode-engine`:
10
+
11
+ - `PUT /v1/models/default` records the model you pick in the IDE as your default, the same setting `/model` writes in the CLI, so both use one choice.
12
+ - `GET` and `PUT /v1/settings/telemetry` read and change the install-telemetry setting the CLI already has, so the IDE asks once and both follow the answer. The switch refuses to move while `KNIGHTCODE_TELEMETRY` is set in the environment.
13
+ - When the IDE starts the engine, it sends the same anonymous install ping as the CLI, once per IDE version, and not at all if you turned install telemetry off.
14
+
15
+ - Added `/undo`: pick an earlier user message and go back to it, with the option to restore every file the abandoned turns edited. Files are backed up before each `edit` and `write`, so the restore needs no git; `/tree` and double-Escape offer the same restore. `KNIGHTCODE_DISABLE_FILE_CHECKPOINTS=1` turns the backups off.
16
+
17
+ - Added click-to-expand for compaction, branch-summary and skill entries in the transcript. A left click toggles them the same way the expand shortcut does.
18
+
19
+ - Added a `terminal` option to `InteractiveMode`, so an embedder or test can drive the interactive UI against its own terminal implementation instead of the process's.
20
+
21
+ ### Changed
22
+
23
+ - Changed long shell command timings to read as minutes and hours (`26m 32s`, `2h 3m 4s`) instead of a raw seconds count; anything under a minute still shows tenths of a second.
24
+
25
+ - Changed the copy shortcut's description in the hotkeys list and keybinding settings to say what it does: it copies the current selection, or the last assistant message when nothing is selected.
26
+
27
+ - Changed the transcript types so tool-call arguments and tool-result `details` are declared as JSON values instead of `any`. A tool that puts a `Date`, a function or `undefined` into `details` is now a type error rather than something that silently fails to round-trip through the session file.
28
+
29
+ - Changed the Kimi For Coding catalog to read from the provider's current listing, which reports the 1M-token context window and low / high / max thinking levels for `kimi-for-coding`.
30
+
31
+ ### Fixed
32
+
33
+ - Fixed thinking blocks being dropped on the next turn when an Anthropic-compatible endpoint reports a different model name than the one requested. The requested model stays on the message, and the reported one is kept separately for cost attribution.
34
+
35
+ - Fixed the "Anthropic dropped thinking" notice repeating on every turn and flooding the transcript with per-block reasons. It now shows a short count, only when a response drops more blocks than the previous one, and not again when a session is reloaded; the details stay in the session file.
36
+
37
+ - Fixed a shell command killed by a signal (an OOM kill, `kill -9`, a `SIGTERM`) being reported to the model as a success. It now fails with the conventional exit code (137 for `SIGKILL`, 143 for `SIGTERM`) and keeps the output it produced before dying.
38
+
39
+ - Fixed a bare `400` or `413` with no body from any provider being treated as a context overflow and triggering compaction. Only Cerebras reports overflow that way, so the rule now applies to Cerebras alone.
40
+
41
+ - Fixed automatic compaction doing nothing when the newest tool result alone exceeds the retained-token budget. The cut now falls back to the assistant message that made the tool call, so older history is summarized before the next request instead of the session overflowing.
42
+
43
+ - Fixed DeepSeek V4.1 Flash offering the wrong thinking levels on OpenRouter and OpenCode Go. Both now expose the low / high / max efforts the model actually accepts, and OpenCode Go lists it under its current ID, `deepseek-v4.1-flash`.
44
+
45
+ - Fixed `knightcode-engine` answering a login request whose body is not JSON, or is too large, with a 500 instead of a 400 or 413.
46
+
47
+ - Fixed a `before_agent_start` handler that returns `systemPrompt` (or sets `forceSystemPrompt`) not actually replacing the prompt on models that accept mid-conversation system messages: they kept the original prompt at the head and received the forced text as a later update. The forced prompt is now sent as the leading system prompt for the run, and the session transcript keeps recording the structured sections instead of the forced text.
48
+
49
+ - Fixed fuzzy search in the model, session and file pickers lagging on long lists; matching now skips ahead with a native substring search and returns the same results in the same order.
50
+
51
+ - Fixed Gemini requests asking for thinking levels a model does not support. Turning thinking off, or picking a level the model lacks, now falls back to the lowest level that model advertises instead of a hard-coded Gemini 3 Pro / Flash guess.
52
+
53
+ - Fixed llama.cpp models whose chat template supports thinking (Qwen-style `enable_thinking`) always running as non-reasoning models. Loaded models are now checked through the server's `/props`, and those templates get an on/off thinking toggle.
54
+
55
+ - Fixed two transient provider failures ending the turn instead of retrying: Cloudflare `520` responses and Azure's "currently experiencing high demand" peak-load rejections are now retried like other overload errors.
56
+
57
+ - Fixed Vercel AI Gateway conversations losing their thinking on the next turn. The gateway returns thinking without a signature for models it translates, and those blocks are now replayed instead of being stripped.
58
+
59
+ ## 0.8.0
60
+
61
+ ### Added
62
+
63
+ - Added `allowedFallbackModels` to Anthropic model overrides in `models.json`, so you can choose which models the server may fall back to — or set an empty array to turn server-side fallback off.
64
+
65
+ - Added webfetch and websearch tools — a page as pageable, greppable markdown, and results from DuckDuckGo or Brave Search — off until enabled with the new `/tools` command, which turns each built-in tool off, on for the session, or on by default, and for websearch also picks the provider and stores the Brave API key.
66
+
67
+ ### Changed
68
+
69
+ - Changed the system prompt and tool set to live in the transcript instead of being rewritten behind it. A session records when its instructions changed or tools became available, resuming or moving between branches restores that state, and providers that support it keep their cached prompt prefix across the change. Extensions can replace individual prompt sections through `systemPromptOptions.sections`, and the deferred-tool loading path is replaced by mid-conversation tool additions on the models that accept them.
70
+
71
+ - Changed a failed copy to say why: instead of a bare "Copy failed" flash, the message now names the missing clipboard backend — `wl-clipboard`, `xclip`/`xsel` or the Termux API package — and stays on screen for five seconds.
72
+
73
+ ### Fixed
74
+
75
+ - Fixed a `user_bash` handler that throws or returns a malformed result silently falling back to the local shell — the command is now reported as failed instead of running somewhere the extension meant to prevent. A handler must return `undefined`, exactly one of `{ operations }` or `{ result }`, and nothing else.
76
+
77
+ - Fixed Baseten requests to carry session-affinity headers so a conversation keeps hitting the same replica and benefits from automatic prompt caching.
78
+
79
+ - Fixed Bedrock cost reporting to bill one-hour cache writes at their higher rate instead of charging every cache write at the five-minute rate.
80
+
81
+ - Fixed a local copy reporting success when no clipboard backend actually took the text: the terminal-escape fallback now only counts in a remote session, where it is the terminal that owns the clipboard.
82
+
83
+ - Fixed the package's public types so extension authors can import the event and result types their hooks receive, such as `ModelSelectEvent`, `ThinkingLevelSelectEvent` and the `*Result` types, instead of redeclaring them.
84
+
85
+ - Fixed the thinking levels offered for Gemini models: they now follow the reasoning efforts each model actually advertises instead of a version-number guess, so newer Flash and Pro models expose their real low/medium/high range.
86
+
87
+ - Fixed resuming a session by its exact ID reading every transcript in the session directory first, which made startup slow in directories with long histories.
88
+
3
89
  ## 0.7.0
4
90
 
5
91
  ### Added
@@ -151,7 +151,7 @@ See [`prepareCompaction()`](https://github.com/KnightCodeAI/knightcode/blob/main
151
151
 
152
152
  ### When It Triggers
153
153
 
154
- When you use `/tree` to navigate to a different branch, KnightCode offers to summarize the work you're leaving. This injects context from the left branch into the new branch.
154
+ When you use `/tree` to navigate to a different branch, KnightCode offers to summarize the work you're leaving. `/undo` never summarizes; it only rewinds. This injects context from the left branch into the new branch.
155
155
 
156
156
  ### How It Works
157
157
 
@@ -406,25 +406,31 @@ For providers with non-standard APIs, implement `streamSimple`. Study the existi
406
406
 
407
407
  ### Stream Pattern
408
408
 
409
- All providers follow the same pattern:
409
+ All providers follow the same pattern. The context is a normalized transcript: the system prompt and tool declarations live in its system messages, so read them with `getCurrentSystemPrompt(context.messages)` and `getCurrentTools(context.messages)` rather than expecting `context.systemPrompt` or `context.tools`. Models that accept system messages mid-conversation can send them in place; otherwise call `collapseSystemMessages(context)` first to fold later system messages into the leading one.
410
410
 
411
411
  ```typescript
412
412
  import {
413
413
  type AssistantMessage,
414
414
  type AssistantMessageEventStream,
415
- type Context,
416
415
  type Model,
417
416
  type SimpleStreamOptions,
417
+ type TranscriptContext,
418
418
  calculateCost,
419
+ collapseSystemMessages,
419
420
  createAssistantMessageEventStream,
421
+ getCurrentSystemPrompt,
422
+ getCurrentTools,
420
423
  } from "@knightcode/ai";
421
424
 
422
425
  function streamMyProvider(
423
426
  model: Model<any>,
424
- context: Context,
427
+ context: TranscriptContext,
425
428
  options?: SimpleStreamOptions
426
429
  ): AssistantMessageEventStream {
427
430
  const stream = createAssistantMessageEventStream();
431
+ const transcript = collapseSystemMessages(context);
432
+ const systemPrompt = getCurrentSystemPrompt(transcript.messages);
433
+ const tools = getCurrentTools(transcript.messages);
428
434
 
429
435
  (async () => {
430
436
  // Initialize output message
@@ -668,10 +674,10 @@ interface ProviderConfig {
668
674
  /** API type for streaming. Required at provider or model level when defining models. */
669
675
  api?: Api;
670
676
 
671
- /** Custom streaming implementation for non-standard APIs. */
677
+ /** Custom streaming implementation for non-standard APIs. Receives a normalized transcript. */
672
678
  streamSimple?: (
673
679
  model: Model<Api>,
674
- context: Context,
680
+ context: TranscriptContext,
675
681
  options?: SimpleStreamOptions
676
682
  ) => AssistantMessageEventStream;
677
683
 
@@ -84,6 +84,7 @@ These variables are read by KnightCode itself:
84
84
  | `KNIGHTCODE_SERVER_DIR` | Override the experimental server profile and socket directory; default is `~/.knightcode/server` |
85
85
  | `KNIGHTCODE_SERVER_ID` | Select the logical experimental server ID when `--server-id` is omitted |
86
86
  | `KNIGHTCODE_OFFLINE` | Disable startup network operations, including update checks, package updates, and install/update telemetry |
87
+ | `KNIGHTCODE_DISABLE_FILE_CHECKPOINTS` | Set to `1`, `true`, or `yes` to stop backing up files before edits; `/undo` then rewinds the conversation only. See [Sessions](sessions.md#file-restore) |
87
88
  | `KNIGHTCODE_SKIP_VERSION_CHECK` | Disable the `knightcode.dev` latest-version request |
88
89
  | `KNIGHTCODE_TELEMETRY` | Override install/update telemetry and provider attribution headers: `1`/`true`/`yes` or `0`/`false`/`no` |
89
90
  | `KNIGHTCODE_CACHE_RETENTION` | Set to `long` for extended provider prompt caching where supported |
@@ -93,6 +94,7 @@ These variables are read by KnightCode itself:
93
94
  | `KNIGHTCODE_IMAGE_PROTOCOL` | Override inline image detection with `kitty`, `iterm2`, `none`, or `auto` |
94
95
  | `KNIGHTCODE_TRUE_COLOR` | Override truecolor detection with `1`, `0`, or `auto` |
95
96
  | `KNIGHTCODE_TUI_ESC_TIMEOUT` | How long to wait after a lone ESC before treating it as Escape, in milliseconds; defaults to `100` over SSH and `10` otherwise. Increase if Alt-key input is misread as Escape |
97
+ | `BRAVE_API_KEY` | Brave Search key for `websearch` when none is stored with `/tools`; see [Web tools](usage.md#web-tools) |
96
98
  | `VISUAL`, `EDITOR` | External editor fallback when `externalEditor` is unset |
97
99
  | `HTTP_PROXY`, `HTTPS_PROXY` | Proxy outbound HTTP requests |
98
100
 
@@ -333,7 +333,7 @@ user sends another prompt ◄─────────────────
333
333
  ├─► session_compact (success)
334
334
  └─► session_compact_failed (failure or abort)
335
335
 
336
- /tree navigation
336
+ /tree, /undo, or double-Escape navigation
337
337
  ├─► session_before_tree (can cancel or customize)
338
338
  └─► session_tree
339
339
 
@@ -492,7 +492,7 @@ knightcode.on("session_compact_failed", async (event, ctx) => {
492
492
 
493
493
  #### session_before_tree / session_tree
494
494
 
495
- Fired on `/tree` navigation. See [Sessions](sessions.md) for tree navigation concepts.
495
+ Fired on `/tree`, `/undo`, and double-Escape navigation. See [Sessions](sessions.md) for tree navigation concepts.
496
496
 
497
497
  ```typescript
498
498
  knightcode.on("session_before_tree", async (event, ctx) => {
@@ -538,10 +538,13 @@ knightcode.on("before_agent_start", async (event, ctx) => {
538
538
  // event.systemPrompt - current chained system prompt for this handler
539
539
  // (includes changes from earlier before_agent_start handlers)
540
540
  // event.systemPromptOptions - structured options used to build the system prompt
541
- // .customPrompt - any custom system prompt (from --system-prompt, SYSTEM.md, or custom templates)
541
+ // .customPrompt - exact prompt prefix from --system-prompt, SYSTEM.md, or custom templates
542
+ // .forceSystemPrompt - optional exact replacement for the complete prompt
542
543
  // .selectedTools - tools currently active in the prompt
543
544
  // .toolSnippets - one-line descriptions for each tool
544
- // .promptGuidelines - custom guideline bullets
545
+ // .toolGuidelines - guideline bullets keyed by tool name
546
+ // .promptGuidelines - additional custom guideline bullets
547
+ // .sections - custom XML-wrapped sections keyed by tag name
545
548
  // .appendSystemPrompt - text from --append-system-prompt flags
546
549
  // .cwd - working directory
547
550
  // .contextFiles - AGENTS.md files and other loaded context files
@@ -560,7 +563,7 @@ knightcode.on("before_agent_start", async (event, ctx) => {
560
563
  });
561
564
  ```
562
565
 
563
- The `systemPromptOptions` field gives extensions access to the same structured data KnightCode uses to build the system prompt. This lets you inspect what KnightCode has loaded — custom prompts, guidelines, tool snippets, context files, skills — without re-discovering resources or re-parsing flags. Use it when your extension needs to make deep, informed changes to the system prompt while respecting user-provided configuration.
566
+ The `systemPromptOptions` field gives extensions access to the same structured data KnightCode uses to build the system prompt. Collections are mutable. Prefer changing `sections`, `selectedTools`, or `promptGuidelines`: KnightCode diffs the resulting prompt sections against what the model already has and appends one system message patching only the changed sections. Returning `systemPrompt`, or setting `forceSystemPrompt`, replaces the whole prompt for the run: every provider receives the forced text as its leading system prompt (a cache miss when it changes), and the session transcript keeps recording the structured sections. Tool selection changes update both the prompt contributions and executable provider tools; calling `knightcode.setActiveTools()` inside the handler has the same effect as editing `selectedTools`. Models that accept system messages mid-conversation receive the patch in place and keep their cached prefix; other models get the replayed prompt as their system prompt, which is a cache miss once per change.
564
567
 
565
568
  Inside `before_agent_start`, `event.systemPrompt` and `ctx.getSystemPrompt()` both reflect the chained system prompt as of the current handler. Later `before_agent_start` handlers can still modify it again.
566
569
 
@@ -906,6 +909,8 @@ knightcode.on("user_bash", (event, ctx) => {
906
909
  });
907
910
  ```
908
911
 
912
+ Returning `undefined` continues to the next handler, then local execution if none handles the event. A valid result stops propagation: `operations` executes the command through the supplied backend, while `result` records the completed command without executing it.
913
+
909
914
  ### Input Events
910
915
 
911
916
  #### input
@@ -1125,7 +1130,7 @@ const options = ctx.getSystemPromptOptions();
1125
1130
  const contextPaths = options.contextFiles?.map((file) => file.path) ?? [];
1126
1131
  ```
1127
1132
 
1128
- This has the same shape and mutability as `before_agent_start` `event.systemPromptOptions`: custom prompt, active tools, tool snippets, prompt guidelines, appended system prompt text, cwd, loaded context files, and loaded skills. It may include full context file contents, so treat it as sensitive extension-local data and avoid exposing it through command lists, logs, or autocomplete metadata.
1133
+ This has the same shape and mutability as `before_agent_start` `event.systemPromptOptions`: custom or forced prompt, active tools, tool snippets, per-tool and custom rules, custom sections, appended prompt text, cwd, loaded context files, and loaded skills. It may include full context file contents, so treat it as sensitive extension-local data and avoid exposing it through command lists, logs, or autocomplete metadata.
1129
1134
 
1130
1135
  This reports the current base prompt inputs. It does not include per-turn `before_agent_start` chained system-prompt changes, later `context` event message mutations, or `before_provider_request` payload rewrites.
1131
1136
 
@@ -1366,7 +1371,16 @@ export default function (knightcode: ExtensionAPI) {
1366
1371
 
1367
1372
  ### knightcode.on(event, handler)
1368
1373
 
1369
- Subscribe to events. See [Events](#events) for event types and return values.
1374
+ Subscribe to events. Returns an unsubscribe function that removes only that registration. See [Events](#events) for event types and return values.
1375
+
1376
+ ```typescript
1377
+ const unsubscribe = knightcode.on("agent_end", async (event) => {
1378
+ unsubscribe();
1379
+ await updateIntegration(event.messages);
1380
+ });
1381
+ ```
1382
+
1383
+ Handlers run in extension load order, then registration order within each extension. Adding or removing a handler does not affect a dispatch already in progress.
1370
1384
 
1371
1385
  ### knightcode.registerTool(definition)
1372
1386
 
@@ -2370,42 +2384,13 @@ If a slot renderer is not defined or throws:
2370
2384
 
2371
2385
  ### Dynamic Tool Loading
2372
2386
 
2373
- Extensions can register many tools while keeping only a small initial set active. A tool can then add more tools with `knightcode.setActiveTools()` during execution. KnightCode detects purely additive changes, records the newly available tool names on that tool result, and applies the updated active set before the next model request.
2374
-
2375
- This works with every model. Models with native deferred-loading support preserve the stable prompt prefix and load the new definitions at the tool-result position. Other models use the fallback described below.
2387
+ Extensions can register many tools while keeping only a small initial set active. A tool can then change the active set with `knightcode.setActiveTools()` during execution. KnightCode stores the initial prompt and tool loadout in the transcript's first system message, then appends tool and prompt deltas before the next model request. Providers that cannot represent a transition receive a complete transcript checkpoint, which may invalidate the cached prefix.
2376
2388
 
2377
2389
  The lifecycle is:
2378
2390
 
2379
2391
  1. Register every tool with `knightcode.registerTool()` so it appears in `knightcode.getAllTools()`.
2380
2392
  2. Keep loader tools, such as `search_tools`, active and leave searchable tools inactive.
2381
- 3. During loader execution, call `knightcode.setActiveTools([...currentTools, ...matchingTools])`. The change must be additive: do not remove currently active tools in the same call.
2382
- 4. KnightCode records which tools were added on the loader's tool result.
2383
- 5. Before the next model response, KnightCode exposes the added definitions using native deferred loading when supported, or the normal active tool list otherwise.
2384
-
2385
- You do not need to return provider-specific tool references or mark the loader as a special search tool. The active-tool change is the signal. Names passed to `knightcode.setActiveTools()` must already be registered; unknown names are ignored.
2386
-
2387
- #### Models with native deferred loading
2388
-
2389
- - **Anthropic**
2390
- - **Models:** Sonnet, Opus, Fable version 4.5 or newer (without Haiku)
2391
- - **Native representation:** Deferred definitions use `defer_loading`; the load point uses `tool_reference` content.
2392
- - **Fireworks Messages API**
2393
- - **Native representation:** Deferred definitions use `defer_loading`; the load point uses `tool_reference` content.
2394
- - **Loader names:** Use `ToolSearch` or `tool_search` for prefix deferral. Other loader names still work, but Fireworks includes the loaded schemas in the initial tool prefix, losing the cache benefit.
2395
- - This does not change API routing: Fireworks GLM models and Kimi K3 use Chat Completions, not Messages.
2396
- - **OpenAI**
2397
- - **Models:** `gpt-5.4` and newer family
2398
- - **Native representation:** KnightCode adds completed client `tool_search_call` and `tool_search_output` items at the load point.
2399
-
2400
- For a verified custom model or proxy, native handling can be enabled with `compat.supportsToolReferences: true` for `anthropic-messages`, or `compat.supportsToolSearch: true` for `openai-responses` and `openai-codex-responses`. Leave these disabled unless the endpoint and model accept the corresponding native protocol.
2401
-
2402
- #### Fallback behavior
2403
-
2404
- For all other models and providers, dynamic activation still works: KnightCode sends the complete current active tool list normally on the next request. The model can call the newly activated tools, but adding their definitions may invalidate the provider's cached prompt prefix.
2405
-
2406
- KnightCode also uses this safe fallback when the active set is not purely additive, such as replacing one group of tools with another. Tool removals therefore work, but they do not use deferred loading.
2407
-
2408
- For the best cache behavior, keep the loader tool active for the whole session and add tools instead of replacing the active set. Also note that activating a tool with `promptSnippet` or `promptGuidelines` rebuilds the system prompt; that system-prompt change can invalidate the prefix even when the provider supports deferred schemas. Lazily loaded tools should usually rely on their tool `description` and omit active-only prompt metadata.
2393
+ 3. During loader execution, call `knightcode.setActiveTools()` with the desired active tool names. Names must already be registered; unknown names are ignored.
2409
2394
 
2410
2395
  #### Search tool example
2411
2396
 
@@ -2507,7 +2492,7 @@ export default function (knightcode: ExtensionAPI) {
2507
2492
  }
2508
2493
  ```
2509
2494
 
2510
- When `search_tools` adds a match, the model receives that definition on the immediately following request. On a native-capable model the definition is anchored after the search result without changing the initial tool-schema prefix. On other models it appears in the normal tool list on that same following request.
2495
+ When `search_tools` adds a match, the model receives the complete updated tool list on the immediately following request.
2511
2496
 
2512
2497
  ## Custom UI
2513
2498
 
@@ -436,6 +436,7 @@ Built-in Anthropic models enable `supportsStrictTools` in their model metadata.
436
436
  | `supportsMidConvoEffort` | Whether the exact Claude model transport supports per-turn effort system messages and thinking binding controls. KnightCode persists native effort levels and always sends `drop_block` when enabled. Default: `false`. |
437
437
  | `allowEmptySignature` | Whether to replay empty thinking signatures as `signature: ""` instead of converting thinking to text. Default: `false`. |
438
438
  | `supportsStrictTools` | Whether the provider accepts strict JSON-schema tool definitions. Default: `false`; built-in Anthropic models enable it in generated metadata. |
439
+ | `allowedFallbackModels` | Up to three server-side fallback models, each with `provider`, `model`, and complete `cost` metadata. An empty array disables fallback. |
439
440
 
440
441
  ## OpenAI Compatibility
441
442
 
@@ -482,7 +483,6 @@ For providers with partial OpenAI compatibility, use the `compat` field.
482
483
  | `sessionAffinityFormat` | For `openai-completions` and `openai-responses`, the session-affinity header format: `openai` sends `session_id`/`x-client-request-id` (completions also `x-session-affinity`), `openai-nosession` omits the underscore-containing `session_id` header, `openrouter` sends `x-session-id`. Does not affect the `prompt_cache_key` body param. Default: auto-detected. |
483
484
  | `supportsStrictMode` | Whether the provider accepts strict JSON-schema function tool definitions. Defaults depend on the API; built-in OpenAI models carry explicit capability metadata. |
484
485
  | `supportsOpenAIGrammarTools` | Whether OpenAI-compatible APIs emit custom Lark/regex grammar tools. When `false`, grammar-constrained tools fall back to normal function tools. Default: `false`; the built-in model catalog enables it for GPT-5+ models on OpenAI, OpenAI Codex, Azure OpenAI, GitHub Copilot, opencode, and Cloudflare AI Gateway. |
485
- | `deferredToolsMode` | Use provider-specific deferred tool serialization. Currently only `"kimi"` is supported for Kimi's OpenAI-compatible Chat Completions format. |
486
486
  | `supportsLongCacheRetention` | Whether the provider accepts long cache retention when cache retention is `long`: `prompt_cache_options.ttl: "30m"` for GPT-5.6+ Responses models, `prompt_cache_retention: "24h"` for earlier OpenAI models, or `cache_control.ttl: "1h"` when `cacheControlFormat` is `anthropic`. Default: `true`. |
487
487
  | `openRouterRouting` | OpenRouter provider routing preferences. This object is sent as-is in the `provider` field of the [OpenRouter API request](https://openrouter.ai/docs/guides/routing/provider-selection). |
488
488
  | `vercelGatewayRouting` | Vercel AI Gateway routing config for provider selection (`only`, `order`) |
@@ -142,7 +142,7 @@ knightcode --name "my task" # Set session display name at startup
142
142
  knightcode --session <path|id> # Open a specific session
143
143
  ```
144
144
 
145
- Inside knightcode, use `/resume`, `/new`, `/tree`, `/fork`, and `/clone` to manage sessions.
145
+ Inside knightcode, use `/resume`, `/new`, `/undo`, `/tree`, `/fork`, and `/clone` to manage sessions.
146
146
 
147
147
  ### Non-interactive mode
148
148
 
package/bin/docs/sdk.md CHANGED
@@ -246,8 +246,8 @@ const state = session.agent.state;
246
246
  // state.messages: AgentMessage[] - conversation history
247
247
  // state.model: Model - current model
248
248
  // state.thinkingLevel: ThinkingLevel - current thinking level
249
- // state.systemPrompt: string - system prompt
250
- // state.tools: AgentTool[] - available tools
249
+ // state.systemPrompt: string - read-only, replayed from the transcript's system messages
250
+ // state.tools: AgentTool[] - executable tools; changes are declared to the model before the next request
251
251
  // state.streamingMessage?: AgentMessage - current partial assistant message
252
252
  // state.errorMessage?: string - latest assistant error
253
253
 
@@ -77,6 +77,14 @@ interface ToolCall {
77
77
  ### Base Message Types (from @knightcode/ai)
78
78
 
79
79
  ```typescript
80
+ interface SystemMessage {
81
+ role: "system";
82
+ content: string | TextContent[];
83
+ toolsAdded?: Tool[];
84
+ toolsRemoved?: Array<{ name: string }>;
85
+ timestamp: number; // Unix ms
86
+ }
87
+
80
88
  interface UserMessage {
81
89
  role: "user";
82
90
  content: string | (TextContent | ImageContent)[];
@@ -109,7 +117,6 @@ interface ToolResultMessage {
109
117
  content: (TextContent | ImageContent)[];
110
118
  details?: any; // Tool-specific metadata
111
119
  usage?: Usage; // Nested LLM work performed by the tool
112
- addedToolNames?: string[];
113
120
  isError: boolean;
114
121
  timestamp: number;
115
122
  }
@@ -177,6 +184,7 @@ interface CompactionSummaryMessage {
177
184
 
178
185
  ```typescript
179
186
  type AgentMessage =
187
+ | SystemMessage
180
188
  | UserMessage
181
189
  | AssistantMessage
182
190
  | ToolResultMessage
@@ -217,7 +225,14 @@ For sessions with a parent (created via `/fork`, `/clone`, or `newSession({ pare
217
225
 
218
226
  ### SessionMessageEntry
219
227
 
220
- A message in the conversation. The `message` field contains an `AgentMessage`.
228
+ A message in the conversation. The `message` field contains an `AgentMessage`. System messages carry the prompt and tool loadout: the first request of a session persists one with every prompt section and tool declaration, and later changes persist as system messages that patch `sections` by name (`null` removes one) and list `toolsAdded`/`toolsRemoved`. Replaying them in order yields the current prompt and tools; there is no separate prompt state entry.
229
+
230
+ ```json
231
+ {"type":"message","id":"a0b1c2d3","parentId":null,"timestamp":"2024-12-03T14:00:00.000Z","message":{"role":"system","content":"","sections":{"preamble":"You are an expert coding assistant...","tools":"<tools>\n- read: ...\n</tools>","cwd":"/project"},"toolsAdded":[{"name":"read","description":"...","parameters":{}}],"timestamp":1733234400000}}
232
+ {"type":"message","id":"d4e5f6g7","parentId":"c3d4e5f6","timestamp":"2024-12-03T14:04:00.000Z","message":{"role":"system","content":"","sections":{"skills":"<skills>...</skills>"},"toolsRemoved":[{"name":"write"}],"timestamp":1733234640000}}
233
+ ```
234
+
235
+ Sessions created before system messages existed have no leading system message; the first request declares the current prompt as a later system message, which replays the same way.
221
236
 
222
237
  ```json
223
238
  {"type":"message","id":"a1b2c3d4","parentId":"prev1234","timestamp":"2024-12-03T14:00:01.000Z","message":{"role":"user","content":"Hello","timestamp":1733234401000}}
@@ -243,15 +258,16 @@ Emitted when the user changes the thinking/reasoning level.
243
258
 
244
259
  ### CompactionEntry
245
260
 
246
- Created when context is compacted. Stores a summary of earlier messages.
261
+ Created when context is compacted. Stores a summary of earlier messages and a complete system prompt/tool checkpoint.
247
262
 
248
263
  ```json
249
- {"type":"compaction","id":"f6g7h8i9","parentId":"e5f6g7h8","timestamp":"2024-12-03T14:10:00.000Z","summary":"User discussed X, Y, Z...","firstKeptEntryId":"c3d4e5f6","tokensBefore":50000}
264
+ {"type":"compaction","id":"f6g7h8i9","parentId":"e5f6g7h8","timestamp":"2024-12-03T14:10:00.000Z","summary":"User discussed X, Y, Z...","firstKeptEntryId":"c3d4e5f6","tokensBefore":50000,"systemMessage":{"role":"system","content":"You are a coding assistant.","toolsAdded":[],"timestamp":1733235000000}}
250
265
  ```
251
266
 
252
267
  `firstKeptEntryId` is required. It identifies the first entry retained from before the compaction entry. When rebuilding context, KnightCode replaces older summarized entries with the compaction summary and keeps the range beginning at this entry.
253
268
 
254
269
  Optional fields:
270
+ - `systemMessage`: The replayed prompt sections and tool declarations at the compaction boundary; it becomes the leading system message of the compacted context, and system messages among the kept entries are dropped in its favor. It is absent on older session entries.
255
271
  - `usage`: LLM usage from generating the summary; included in session token and cost totals
256
272
  - `details`: Implementation-specific data (e.g., `{ readFiles: string[], modifiedFiles: string[] }` for default, or custom data for extensions)
257
273
  - `fromHook`: `true` if generated by an extension, `false`/`undefined` if knightcode-generated (legacy field name)
@@ -336,7 +352,7 @@ Entries normally form one tree, but navigation APIs can create multiple roots:
336
352
  1. Collects all entries on the path
337
353
  2. If one or more `CompactionEntry` values are on the path, uses the latest one:
338
354
  - Includes the compaction entry first
339
- - Includes entries from `firstKeptEntryId` up to, but not including, the compaction entry
355
+ - Includes non-system entries from `firstKeptEntryId` up to, but not including, the compaction entry
340
356
  - Includes entries after the compaction entry
341
357
  3. Preserves non-message entries in the selected range so interactive mode can render them
342
358
 
@@ -345,12 +361,12 @@ Entries normally form one tree, but navigation APIs can create multiple roots:
345
361
  1. Extracts current model and thinking level settings from the full path
346
362
  2. Converts selected entries to messages:
347
363
  - `message` -> stored `AgentMessage`
348
- - `compaction` -> `compactionSummary`
364
+ - `compaction` -> complete system checkpoint followed by `compactionSummary`
349
365
  - `branch_summary` -> `branchSummary`
350
366
  - `custom_message` -> `CustomMessage`
351
367
  - `custom` -> no context message
352
368
 
353
- The compaction summary replaces entries before `firstKeptEntryId`. The retained entries and all entries after the compaction remain available to the LLM.
369
+ The compaction summary replaces entries before `firstKeptEntryId`. Pre-compaction system messages are folded into the complete checkpoint rather than replayed from the retained range. Retained non-system entries and all entries after the compaction remain available to the LLM.
354
370
 
355
371
  ## Parsing Example
356
372
 
@@ -28,6 +28,7 @@ For the JSONL file format and SessionManager API, see [Session Format](session-f
28
28
  | `/name <name>` | Set the current session display name |
29
29
  | `/session` | Show session info |
30
30
  | `/tree` | Navigate the current session tree |
31
+ | `/undo` | Go back to an earlier user message, optionally restoring files |
31
32
  | `/fork` | Create a new session from a previous user message |
32
33
  | `/clone` | Duplicate the current active branch into a new session |
33
34
  | `/compact [prompt]` | Summarize older context; see [Compaction](compaction.md) |
@@ -117,6 +118,40 @@ Selecting an assistant, tool, compaction, or other non-user entry:
117
118
 
118
119
  Selecting the root user message resets the leaf to an empty conversation and places the original prompt in the editor.
119
120
 
121
+ ## Undoing with `/undo`
122
+
123
+ `/undo` is the quick form of `/tree` for the common case: pick a user message on the current
124
+ branch and go back to it. It lists only user messages, newest at the bottom, with the time
125
+ and how many files an undo to that point would restore. Enter goes back; the turns after the
126
+ chosen message are left as an abandoned branch (still reachable through `/tree`) and the
127
+ message text returns to the editor for editing and resubmitting.
128
+
129
+ ### File Restore
130
+
131
+ Before `edit` or `write` changes a file, KnightCode copies the original into
132
+ `~/.knightcode/agent/file-history/<session id>/`, once per file per user turn. When you go
133
+ back past a turn that changed files, through `/undo`, `/tree`, or double-Escape, you are asked:
134
+
135
+ - **Conversation only**: files stay as they are.
136
+ - **Conversation and N files**: every file the abandoned turns touched is put back to its
137
+ content from before those turns; files they created are deleted.
138
+
139
+ Escape at that prompt cancels the navigation entirely.
140
+
141
+ Limits:
142
+
143
+ - Changes made by shell commands (`bash`, `powershell`) are not tracked. The prompt says so
144
+ when a shell command ran in the abandoned turns.
145
+ - Edits made by subagents or outside KnightCode are not tracked either; a tracked file is
146
+ restored to its backup regardless of who changed it afterwards.
147
+ - If you go back without restoring files and later go back further, only the edits on the
148
+ branch you are on are restored.
149
+ - Jumping sideways to another branch with `/tree` restores to the fork point; edits on the
150
+ target branch are not re-applied.
151
+
152
+ Backups older than 30 days are removed at startup. Set `KNIGHTCODE_DISABLE_FILE_CHECKPOINTS=1`
153
+ to turn file tracking off; `/undo` then rewinds the conversation only.
154
+
120
155
  ## `/tree`, `/fork`, and `/clone`
121
156
 
122
157
  | Feature | `/tree` | `/fork` | `/clone` |
@@ -27,8 +27,8 @@ Use `/trust` in interactive mode to save a project trust decision for future ses
27
27
 
28
28
  | Setting | Type | Default | Description |
29
29
  |---------|------|---------|-------------|
30
- | `defaultProvider` | string | - | Startup provider (e.g., `"anthropic"`, `"openai"`; saved with Ctrl+S in `/model`, or edited manually) |
31
- | `defaultModel` | string | - | Startup model ID (saved with Ctrl+S in `/model`, or edited manually) |
30
+ | `defaultProvider` | string | - | Startup provider (e.g., `"anthropic"`, `"openai"`; saved with Ctrl+S in `/model`, by the IDE's model picker, or edited manually) |
31
+ | `defaultModel` | string | - | Startup model ID (saved with Ctrl+S in `/model`, by the IDE's model picker, or edited manually) |
32
32
  | `defaultThinkingLevel` | string | - | Startup thinking level (saved with Ctrl+S in `/thinking`, or edited manually): `"off"`, `"minimal"`, `"low"`, `"medium"`, `"high"`, `"xhigh"`, `"max"` |
33
33
  | `modelThinkingLevels` | object | - | Per-model startup thinking levels keyed by `"provider/modelId"`; configure from `/settings` → Default thinking level per model or edit manually |
34
34
  | `hideThinkingBlock` | boolean | `false` | Hide thinking blocks in output |
@@ -60,7 +60,7 @@ Use `/trust` in interactive mode to save a project trust decision for future ses
60
60
  | `enableInstallTelemetry` | boolean | `true` | Send the anonymous install/update ping and selected provider attribution headers. This does not control update checks |
61
61
  | `enableAnalytics` | boolean | `false` | Opt-in analytics data sharing. Currently only asked for during the experimental first-time setup (`KNIGHTCODE_EXPERIMENTAL=1`) |
62
62
  | `trackingId` | string | - | Analytics tracking identifier, generated when `enableAnalytics` is turned on |
63
- | `doubleEscapeAction` | string | `"tree"` | Action for double-escape: `"tree"`, `"fork"`, or `"none"` |
63
+ | `doubleEscapeAction` | string | `"tree"` | Action for double-escape: `"tree"`, `"fork"`, or `"none"`. `/undo` is the quick way to a user message |
64
64
  | `treeFilterMode` | string | `"default"` | Default filter for `/tree`: `"default"`, `"no-tools"`, `"user-only"`, `"labeled-only"`, `"all"` |
65
65
  | `editorPaddingX` | number | `0` | Horizontal padding for input editor (0-3) |
66
66
  | `outputPad` | number | `1` | Horizontal padding for user messages, assistant messages, and thinking (0 or 1) |
@@ -83,6 +83,8 @@ For VS Code, include `--wait` so knightcode resumes after the editor exits:
83
83
 
84
84
  `enableInstallTelemetry` controls the anonymous install/update ping to `https://knightcode.dev/api/report-install` and KnightCode attribution headers for OpenRouter, NVIDIA NIM, and Cloudflare provider requests. Opting out disables both. It does not disable update checks; KnightCode can still fetch `https://knightcode.dev/api/latest-version` to look for the latest version.
85
85
 
86
+ The KnightCode IDE shares this setting rather than keeping its own. Its first run and its settings write `enableInstallTelemetry` here, except while `KNIGHTCODE_TELEMETRY` is set, because the variable outranks the setting. The engine the IDE starts sends the same ping once per IDE version, with a `knightcode-ide` user agent, and records the version it was delivered for as `lastIdeVersion`. A ping that fails is retried on the next start.
87
+
86
88
  Set `KNIGHTCODE_SKIP_VERSION_CHECK=1` to disable the KnightCode version update check. Use `--offline` or `KNIGHTCODE_OFFLINE=1` to disable all startup network operations described here, including update checks, package update checks, and install/update telemetry.
87
89
 
88
90
  ### Network
@@ -279,6 +281,8 @@ On Windows, select `powershell` instead of `bash`, or include both:
279
281
 
280
282
  An empty array starts with no built-in tools while preserving extension and SDK custom tools. `--tools` replaces this behavior with a strict allowlist for all tools, `--no-tools` disables all tools, and `--no-builtin-tools` disables the built-in defaults. `--exclude-tools` filters the resulting list. A project `defaultTools` array replaces the global array.
281
283
 
284
+ The [web tools](usage.md#web-tools) `webfetch` and `websearch` are not part of `defaultTools`. `/tools` turns them on and stores their settings, including the search provider and Brave key, in `~/.knightcode/agent/tools.json`.
285
+
282
286
  ### Sessions
283
287
 
284
288
  | Setting | Type | Default | Description |
package/bin/docs/usage.md CHANGED
@@ -42,11 +42,13 @@ Type `/` in the editor to open command completion. Extensions can register custo
42
42
  | `/thinking` | Switch thinking level; Ctrl+S or Ctrl+D in the picker saves the startup default |
43
43
  | `/scoped-models` | Enable/disable models for Ctrl+P cycling |
44
44
  | `/settings` | Theme, message delivery, transport, and other preferences |
45
+ | [`/tools`](#web-tools) | Turn `webfetch` and `websearch` off, on for this session, or on by default; pick the search provider and store its key |
45
46
  | `/resume` | Pick from previous sessions |
46
47
  | `/new` | Start a new session |
47
48
  | `/name <name>` | Set session display name |
48
49
  | `/session` | Show session file, ID, messages, tokens, and cost |
49
50
  | `/tree` | Jump to any point in the session and continue from there |
51
+ | `/undo` | Go back to an earlier user message; optionally restore the files edited after it |
50
52
  | `/trust` | Save project trust decision for future sessions |
51
53
  | `/fork` | Create a new session from a previous user message |
52
54
  | `/clone` | Duplicate the current active branch into a new session |
@@ -60,6 +62,37 @@ Type `/` in the editor to open command completion. Extensions can register custo
60
62
  | `/changelog` | Display version history |
61
63
  | `/quit` | Quit knightcode |
62
64
 
65
+ ## Web Tools
66
+
67
+ Two tools give the agent read access to the web. Both ship disabled; turn them on with `/tools`.
68
+
69
+ | Tool | What it does |
70
+ |------|--------------|
71
+ | `webfetch` | Fetches a URL and returns the page as markdown, 400 lines at a time. The agent pages with `offset`/`limit` like `read`, or passes `grep` to get only matching lines. Responses are cached for 15 minutes, capped at 5 MB, and time out after 30 seconds. Private, loopback, and link-local addresses are refused. |
72
+ | `websearch` | Searches the web and returns up to 10 results (default 5) as title, URL, and snippet. DuckDuckGo needs no key but may rate-limit; Brave Search needs an API key. |
73
+
74
+ `/tools` lists each tool with its current state; pick one to open its settings panel. **Status** cycles through three states:
75
+
76
+ | State | Effect |
77
+ |-------|--------|
78
+ | Disabled | The tool is not offered to the model |
79
+ | Enabled for this session | On until knightcode exits, including across `/new`, `/resume`, and `/fork`; nothing is written to disk |
80
+ | Enabled by default | On in every session |
81
+
82
+ The `websearch` panel adds two rows: **Provider** (`duckduckgo` or `brave`) and **Brave API key**. Choosing Brave without a stored key opens the key prompt at once. Get a key at https://brave.com/search/api/; the free plan is enough. The key is also read from `BRAVE_API_KEY` when none is stored, but the provider only switches to Brave when you pick it. Keys show masked in the panel.
83
+
84
+ The same changes work without the panel:
85
+
86
+ ```bash
87
+ /tools websearch on # this session
88
+ /tools webfetch always # every session
89
+ /tools webfetch off
90
+ ```
91
+
92
+ Settings are stored in `~/.knightcode/agent/tools.json`, owner-readable only because it can hold the Brave key. `--tools` and `--exclude-tools` still apply: a tool excluded on the command line stays off whatever `/tools` says.
93
+
94
+ Fetched pages and search results are marked as untrusted in the tool output so the model treats instructions inside them as data, but treat the tools like any other network access: a fetched page can still influence what the agent does next.
95
+
63
96
  ## Message Queue
64
97
 
65
98
  You can submit messages while the agent is still working:
@@ -90,6 +123,7 @@ Useful session commands:
90
123
 
91
124
  - `/session` shows the current session file and ID.
92
125
  - `/tree` navigates the in-file session tree and can summarize abandoned branches.
126
+ - `/undo` goes back to an earlier user message and can restore the files edited after it.
93
127
  - `/fork` creates a new session from an earlier user message.
94
128
  - `/clone` duplicates the current active branch into a new session file.
95
129
  - `/compact` summarizes older messages to free context.
@@ -212,7 +246,7 @@ cat README.md | knightcode -p "Summarize this text"
212
246
  | `--no-builtin-tools`, `-nbt` | Disable built-in tools but keep extension/custom tools enabled |
213
247
  | `--no-tools`, `-nt` | Disable all tools |
214
248
 
215
- Built-in tools: `read`, `bash`, `powershell` (Windows), `edit`, `write`, `grep`, `find`, `ls`.
249
+ Built-in tools: `read`, `bash`, `powershell` (Windows), `edit`, `write`, `grep`, `find`, `ls`. The [web tools](#web-tools) `webfetch` and `websearch` are off until enabled with `/tools`.
216
250
 
217
251
  ### Resource Options
218
252
 
package/bin/knightcode CHANGED
@@ -1,4 +1,4 @@
1
1
  [diffend] Oversized file quarantined before diffing.
2
2
  name: package/bin/knightcode
3
- size: 109154089 bytes
4
- sha256: 96eb9e025122d5cfe89f26d8bb900f09ac75d1412a50d8f6b0e422cf7670d652
3
+ size: 109982795 bytes
4
+ sha256: 353976a412cf82079c26f715c54c7a4ea79af505b52be2d9774de018e390d21b
package/bin/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@knightcodeai/cli",
3
- "version": "0.7.0",
3
+ "version": "0.9.0",
4
4
  "description": "KnightCode — a local, BYOK terminal coding agent powered by OpenRouter.",
5
5
  "type": "module",
6
6
  "repository": {
@@ -37,11 +37,11 @@
37
37
  "test": "vitest --run"
38
38
  },
39
39
  "optionalDependencies": {
40
- "@knightcodeai/cli-linux-x64": "0.7.0",
41
- "@knightcodeai/cli-linux-arm64": "0.7.0",
42
- "@knightcodeai/cli-darwin-x64": "0.7.0",
43
- "@knightcodeai/cli-darwin-arm64": "0.7.0",
44
- "@knightcodeai/cli-win32-x64": "0.7.0"
40
+ "@knightcodeai/cli-linux-x64": "0.9.0",
41
+ "@knightcodeai/cli-linux-arm64": "0.9.0",
42
+ "@knightcodeai/cli-darwin-x64": "0.9.0",
43
+ "@knightcodeai/cli-darwin-arm64": "0.9.0",
44
+ "@knightcodeai/cli-win32-x64": "0.9.0"
45
45
  },
46
46
  "devDependencies": {
47
47
  "@agentclientprotocol/sdk": "1.4.0",
@@ -50,6 +50,7 @@
50
50
  "@knightcode/client": "workspace:*",
51
51
  "@knightcode/protocol": "workspace:*",
52
52
  "@knightcode/remote": "workspace:*",
53
+ "@knightcode/tools": "workspace:*",
53
54
  "@knightcode/server": "workspace:*",
54
55
  "@knightcode/session-backend-sqlite": "workspace:*",
55
56
  "@knightcode/tui": "workspace:*",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@knightcodeai/cli-linux-arm64",
3
- "version": "0.7.0",
3
+ "version": "0.9.0",
4
4
  "license": "MIT",
5
5
  "repository": {
6
6
  "type": "git",