samagotchi 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +162 -1
- data/README.md +29 -2
- data/bin/chi +60 -69
- data/docs/cli.md +211 -77
- data/docs/configuration.md +118 -21
- data/docs/desktop.md +39 -4
- data/docs/guardrails.md +11 -0
- data/docs/hooks.md +89 -7
- data/docs/memory.md +40 -0
- data/docs/plugins.md +50 -0
- data/docs/releasing.md +15 -12
- data/docs/sessions.md +20 -18
- data/lib/samagotchi/bootstrap/config_writer.rb +1 -2
- data/lib/samagotchi/bridge/sse_writer.rb +0 -3
- data/lib/samagotchi/bridge/turn_accumulator.rb +2 -0
- data/lib/samagotchi/bridge.rb +20 -12
- data/lib/samagotchi/bundles/skills/manifest.yml +10 -0
- data/lib/samagotchi/bundles/skills/plugin.rb +419 -0
- data/lib/samagotchi/bundles/source-links/hooks/source_links.rb +178 -5
- data/lib/samagotchi/bundles/source-links/manifest.yml +3 -3
- data/lib/samagotchi/bundles/source-links/source_links.md +1 -1
- data/lib/samagotchi/bundles/system/config_modification_protocol.md +9 -10
- data/lib/samagotchi/bundles/system/delegated.md +6 -7
- data/lib/samagotchi/bundles/system/identity.md +5 -0
- data/lib/samagotchi/bundles/system/manifest.yml +6 -6
- data/lib/samagotchi/bundles/system/memory_guide.md +26 -0
- data/lib/samagotchi/bundles/system/self_map.md +2 -1
- data/lib/samagotchi/client.rb +25 -26
- data/lib/samagotchi/commands/registry.rb +8 -0
- data/lib/samagotchi/config.rb +97 -113
- data/lib/samagotchi/desktop/macos/App.swift +12 -8
- data/lib/samagotchi/desktop/macos/ChiRunner.swift +4 -2
- data/lib/samagotchi/desktop/macos/Images.swift +113 -0
- data/lib/samagotchi/desktop/macos/Info.plist.erb +6 -0
- data/lib/samagotchi/desktop/macos/Panel.swift +112 -9
- data/lib/samagotchi/desktop/macos.rb +59 -8
- data/lib/samagotchi/desktop_command.rb +6 -3
- data/lib/samagotchi/edit_preview.rb +82 -0
- data/lib/samagotchi/engine.rb +236 -443
- data/lib/samagotchi/gem_update.rb +89 -0
- data/lib/samagotchi/guardrails/approval.rb +26 -4
- data/lib/samagotchi/guardrails/load_failures.rb +9 -3
- data/lib/samagotchi/host_registry.rb +8 -12
- data/lib/samagotchi/idle_client.rb +24 -15
- data/lib/samagotchi/idle_reminders.rb +2 -2
- data/lib/samagotchi/image_store.rb +10 -6
- data/lib/samagotchi/kernel_loop.rb +59 -123
- data/lib/samagotchi/live_versions.rb +65 -0
- data/lib/samagotchi/llm/api_key.rb +41 -0
- data/lib/samagotchi/llm/chat_loop.rb +77 -13
- data/lib/samagotchi/llm/errors.rb +38 -7
- data/lib/samagotchi/llm/http.rb +19 -22
- data/lib/samagotchi/llm/openai_chat.rb +22 -26
- data/lib/samagotchi/memory_bundle/installer.rb +65 -63
- data/lib/samagotchi/memory_bundle/provenance.rb +51 -12
- data/lib/samagotchi/memory_bundle/shipped_update.rb +157 -0
- data/lib/samagotchi/memory_bundle/status.rb +4 -1
- data/lib/samagotchi/memory_bundle/system_bundle.rb +81 -53
- data/lib/samagotchi/model_profile.rb +27 -10
- data/lib/samagotchi/note_command.rb +2 -1
- data/lib/samagotchi/prompt.rb +4 -2
- data/lib/samagotchi/reminder_store.rb +1 -9
- data/lib/samagotchi/reply_wait.rb +48 -4
- data/lib/samagotchi/self_report.rb +37 -5
- data/lib/samagotchi/send_command.rb +190 -17
- data/lib/samagotchi/session.rb +4 -2
- data/lib/samagotchi/session_commands.rb +38 -8
- data/lib/samagotchi/session_manager.rb +19 -53
- data/lib/samagotchi/system_prompt.rb +403 -0
- data/lib/samagotchi/terminal_ui/attach_launcher.rb +5 -3
- data/lib/samagotchi/terminal_ui/attached_loop.rb +141 -108
- data/lib/samagotchi/terminal_ui/attached_view.rb +27 -12
- data/lib/samagotchi/terminal_ui/event_renderer.rb +29 -8
- data/lib/samagotchi/terminal_ui/formatting.rb +41 -22
- data/lib/samagotchi/terminal_ui/input_support.rb +7 -23
- data/lib/samagotchi/terminal_ui/plain_surface.rb +13 -7
- data/lib/samagotchi/terminal_ui/question_prompt.rb +35 -0
- data/lib/samagotchi/terminal_ui/status_row.rb +81 -0
- data/lib/samagotchi/terminal_ui/surface.rb +1 -1
- data/lib/samagotchi/terminal_ui.rb +142 -923
- data/lib/samagotchi/text_diff.rb +181 -0
- data/lib/samagotchi/thinking.rb +126 -0
- data/lib/samagotchi/tool_activity.rb +52 -2
- data/lib/samagotchi/tool_runner.rb +37 -1
- data/lib/samagotchi/tools/ask_user_question.rb +41 -33
- data/lib/samagotchi/tools/edit.rb +23 -9
- data/lib/samagotchi/tools/execute.rb +3 -3
- data/lib/samagotchi/tools/output_guardrails.rb +8 -7
- data/lib/samagotchi/tools/read.rb +4 -4
- data/lib/samagotchi/tools/write.rb +4 -0
- data/lib/samagotchi/turn_flow.rb +12 -2
- data/lib/samagotchi/update_command.rb +309 -0
- data/lib/samagotchi/update_hint.rb +59 -0
- data/lib/samagotchi/version.rb +1 -1
- data/lib/samagotchi/vision_support.rb +6 -4
- data/lib/samagotchi/web/app.rb +173 -38
- data/lib/samagotchi/web/lan.rb +99 -0
- data/lib/samagotchi/web/message_parts.rb +19 -10
- data/lib/samagotchi/web/public/activity.js +10 -0
- data/lib/samagotchi/web/public/app.js +135 -78
- data/lib/samagotchi/web/public/chat_view.js +8 -1
- data/lib/samagotchi/web/public/data.js +2 -0
- data/lib/samagotchi/web/public/diff_view.js +58 -0
- data/lib/samagotchi/web/public/index.html +185 -18
- data/lib/samagotchi/web/public/model_pick.js +136 -0
- data/lib/samagotchi/web/public/model_picker.js +224 -0
- data/lib/samagotchi/web/public/notify.js +10 -0
- data/lib/samagotchi/web/public/question_card.js +3 -1
- data/lib/samagotchi/web/public/stage_model.js +110 -0
- data/lib/samagotchi/web/public/stage_view.js +580 -0
- data/lib/samagotchi/web/public/timing.js +6 -2
- data/lib/samagotchi/web/public/turn_events.js +38 -10
- data/lib/samagotchi/web/public/turn_model.js +11 -3
- data/lib/samagotchi/web/public/turn_view.js +76 -20
- data/lib/samagotchi/web/qr.rb +40 -0
- data/lib/samagotchi/web/server.rb +101 -11
- data/lib/samagotchi/web/token.rb +97 -0
- data/lib/samagotchi/worker.rb +5 -4
- metadata +38 -3
- data/lib/samagotchi/terminal_ui/legacy_surface.rb +0 -111
data/docs/cli.md
CHANGED
|
@@ -12,14 +12,16 @@
|
|
|
12
12
|
- `chi --attach <session-id>` — attach the terminal to a session's worker (e.g. one started from the Web UI), waking one if it has exited
|
|
13
13
|
- A session id can be shortened to any unique prefix (like git): `chi --attach 2ea8`. `--resume`, `--attach`, `sessions stop`, `sessions archive` and `sessions delete` take one; an ambiguous prefix lists the sessions it matches.
|
|
14
14
|
- `chi web [--port 4567] [--open] [--scope=all]` — start the Web UI (single localhost port session control plane) on this git project's sessions (`--scope=all`, or a folder in no repo: every session); if a chi web already runs on the port, print (with `--open`, open) its page for this folder and exit. Something else on the port (an older chi web too) exits 1 with "port N is in use"
|
|
15
|
+
- `chi web --web-host lan` — the Web UI on your home network too, for your phone: a link with an access token and its QR code (see [chi web on your phone](#chi-web-on-your-phone)); `chi web --new-token` replaces the token
|
|
15
16
|
- `chi web --web-markdown` — opt in to sanitized Markdown rendering for completed assistant messages
|
|
16
|
-
- `chi web --
|
|
17
|
+
- `chi web --web-view stage|chat` — draw turns with the stage view (the running turn pinned above the composer) or as the classic row of bubbles instead of the default turn view (each turn as one block of steps, the running one at the bottom); `?view=turn|stage|chat` on the page URL overrides it (see [Web views](#web-views))
|
|
17
18
|
- `chi sessions list|stop|archive|unarchive|delete|prune|clean` — manage persisted sessions; `list` shows this git project's, `list --scope=all` every one, a delegated session with `↳ <parent>`, `list --archived` the archived ones too (see [Sessions](sessions.md))
|
|
18
19
|
- `chi note [--source NAME] [-m TEXT] (ID|PREFIX)... | --all` — add a context note (TEXT or stdin) to sessions: background the model sees on its next turn; it starts no turn (see [Sessions: Context notes](sessions.md#context-notes))
|
|
19
|
-
- `chi send [-m TEXT] (ID|PREFIX)...` — send a message to sessions as if typed there: a turn starts (or a running one picks it up); piped stdin goes above `-m` as quoted context (see [Sessions: Sending a message](sessions.md#sending-a-message)); `--new` starts a session with it instead, and `--wait` prints the answer (`--wait ID` with no message waits for the next reply without sending; see [Starting a session](sessions.md#starting-a-session))
|
|
20
|
+
- `chi send [-m TEXT] [--image PATH]... (ID|PREFIX)...` — send a message to sessions as if typed there: a turn starts (or a running one picks it up); piped stdin goes above `-m` as quoted context, and `--image` attaches images (see [Sessions: Sending a message](sessions.md#sending-a-message)); `--new` starts a session with it instead, and `--wait` prints the answer (`--wait ID` with no message waits for the next reply without sending; see [Starting a session](sessions.md#starting-a-session))
|
|
20
21
|
- `chi desktop install|upgrade|uninstall|status` — the macOS "Send to chi" helper: a Service and a ⌃⌥⌘N hotkey that send text to live sessions as context notes (see [Desktop helper](desktop.md))
|
|
21
22
|
- `chi self` — print version, source dir (checkout or installed gem), config/memory/session paths, model/host and bundles
|
|
22
|
-
- `chi
|
|
23
|
+
- `chi update [--dry-run] [--no-gem] [--no-bundles] [--no-desktop]` — update an installed chi: the gem, the system bundle, the shipped bundles you installed and the desktop helper, in one table (see [Updating](#updating))
|
|
24
|
+
- `chi bundle install|upgrade|uninstall|status|diff|list|build` — manage memory bundles (see [Bundle hooks](hooks.md#bundle-hooks-unified-workflow-bundle)); `list` shows the installed ones and the ones shipped with chi, which `install <name>` installs (see [Guardrails](guardrails.md), [Plugins](plugins.md#the-btw-bundle), [the mcp bundle](plugins.md#the-mcp-bundle) [the loop-guard bundle](plugins.md#the-loop-guard-bundle), [the check-in bundle](plugins.md#the-check-in-bundle) and [the skills bundle](plugins.md#the-skills-bundle))
|
|
23
25
|
|
|
24
26
|
### First setup
|
|
25
27
|
|
|
@@ -57,6 +59,65 @@ chi bootstrap # try localhost 8080, 11434, 1234, 8000
|
|
|
57
59
|
or with anchors gets the lines printed to paste instead. `--dry-run` shows
|
|
58
60
|
what it would write.
|
|
59
61
|
|
|
62
|
+
### Updating
|
|
63
|
+
|
|
64
|
+
`chi update` brings an installed chi up to date and prints one table:
|
|
65
|
+
|
|
66
|
+
```
|
|
67
|
+
component from to status
|
|
68
|
+
chi (gem) 0.2.0 0.3.0 updated
|
|
69
|
+
system bundle 0.2.0 0.3.0 updated (kept your edits in identity.md: chi bundle diff samagotchi-system identity.md)
|
|
70
|
+
btw 0.1.1 up to date
|
|
71
|
+
known-names 0.1.0 0.1.1 updated
|
|
72
|
+
infra_tools 1.0.0 skipped (not from chi)
|
|
73
|
+
Chi Helper 0.2.0 up to date (launch file refreshed)
|
|
74
|
+
workers 2 live on 0.2.0: they move to 0.3.0 at idle exit (30 min) or chi sessions stop 2ea8c1f0 91b0d2aa
|
|
75
|
+
Also shipped, not installed: check-in, skills, source-links (chi bundle install NAME)
|
|
76
|
+
done
|
|
77
|
+
```
|
|
78
|
+
|
|
79
|
+
- **The gem.** It asks rubygems.org for the newest samagotchi (5 s timeout)
|
|
80
|
+
and, when that's newer, runs `gem install samagotchi` with the gem command
|
|
81
|
+
of the Ruby chi runs on (the real one, not a mise/rbenv/asdf shim). Then it
|
|
82
|
+
hands over to the new chi, which does the rest and prints the table. Old
|
|
83
|
+
versions stay installed: running workers and an old `chi web` still use
|
|
84
|
+
them (so don't `gem cleanup` while they run). Offline, the row says
|
|
85
|
+
"couldn't check" and the rest still runs; a failed install fails the row
|
|
86
|
+
and the rest runs on the current version. Under Bundler (`bundle exec`)
|
|
87
|
+
the row says `bundle update samagotchi` instead.
|
|
88
|
+
- **The system bundle** normally updated itself when the new chi started;
|
|
89
|
+
the row says what it did.
|
|
90
|
+
- **Shipped bundles**: each one you installed from chi (`chi bundle install
|
|
91
|
+
NAME`) is upgraded when chi ships a newer version. Memory files get the
|
|
92
|
+
3-way merge of `chi bundle upgrade`: an unedited file is updated, an edited
|
|
93
|
+
one that the new version also changes is kept, and the row says so (`chi
|
|
94
|
+
bundle diff NAME FILE` shows it; `chi bundle upgrade NAME --force` takes the
|
|
95
|
+
bundle's). Hooks, rules and the plugin are replaced; an edited one didn't
|
|
96
|
+
load anyway (its sha no longer matched) and the row says it was replaced.
|
|
97
|
+
A bundle of the same name from elsewhere (a zip, git) is skipped ("not from
|
|
98
|
+
chi"), a newer installed one is left, one whose new version needs a newer
|
|
99
|
+
chi is skipped, and bundles you didn't install stay uninstalled.
|
|
100
|
+
- **The desktop helper** (macOS) is rebuilt and restarted only when its Swift
|
|
101
|
+
sources changed (or the Ruby it runs moved); otherwise only its launch file
|
|
102
|
+
is refreshed. See [Desktop helper](desktop.md).
|
|
103
|
+
- **Running processes** are reported, never stopped: live workers on another
|
|
104
|
+
version, and a `chi web` on `web.port` running an older chi (sessions it
|
|
105
|
+
starts run that version too: restart it).
|
|
106
|
+
|
|
107
|
+
`--dry-run` shows the table with "would update" and changes nothing. It is
|
|
108
|
+
this version's view: bundles that only a newer gem ships newer show up once
|
|
109
|
+
that gem is installed (the real run installs it first and hands over).
|
|
110
|
+
`--no-gem`, `--no-bundles` and `--no-desktop` leave a part alone for one run;
|
|
111
|
+
`update.gem`, `update.bundles` and `update.desktop: false` in config.yml turn
|
|
112
|
+
one off for good. It exits 0 when nothing failed (kept edits and skips are
|
|
113
|
+
fine), 1 when a part failed, 2 on a usage error. A second run changes nothing
|
|
114
|
+
and ends with "everything is up to date".
|
|
115
|
+
|
|
116
|
+
From a checkout it refuses (`git pull`, or `chi bundle upgrade NAME` for one
|
|
117
|
+
bundle). After a gem update, the first interactive start of the new version
|
|
118
|
+
(`chi`, `chi web`; not `-p` or `--non-interactive`) says in one line when
|
|
119
|
+
bundles or the helper can be updated.
|
|
120
|
+
|
|
60
121
|
## Flags
|
|
61
122
|
|
|
62
123
|
Samagotchi exposes one flag that feeds a prompt (`-p`, `--prompt`) and one that
|
|
@@ -71,6 +132,7 @@ controls exit behavior (`--non-interactive`); `--resume` composes with both.
|
|
|
71
132
|
| `--no-shared` | Run the plain in-process REPL for this run. |
|
|
72
133
|
| `--attach SESSION_ID` | Attach to a session's worker, waking one if it has exited. |
|
|
73
134
|
| `--model NAME` | Use this model for the run (overrides the configured default and a resumed session's model). |
|
|
135
|
+
| `--thinking LEVEL` | How much the model thinks this run: `off`, `low`, `medium`, `high` or `default` (env `SAMAGOTCHI_THINKING_LEVEL`), over the config's levels. A session already running keeps its own. See "Thinking" in configuration.md. |
|
|
74
136
|
| `--profile NAME` | Prompt profile (`qwen36` or `gemma4`) for every model in this run, over config and the server's template (same as `--model-profile`, env `SAMAGOTCHI_MODEL_PROFILE`). See "Prompt profile" in configuration.md. |
|
|
75
137
|
| `--memory NAME` | Preload a memory entry into the system prompt (repeatable; a comma list too). Merged under the config.yml `memories:` baseline. Works attached: the list is stored on the session, so its worker builds the same prompt on every respawn. |
|
|
76
138
|
| `--mute NAME` | Hide a memory from this session (repeatable; a comma list too): its index line is not in the prompt, `memory_read` refuses it, the identity auto-load skips it, and it is dropped from the preloads (config baseline or `--memory`). A name matches in both scopes (`gh-helper`, `project/gh-helper` and `gh-helper.md` all hide `gh-helper`). Nothing on disk changes. See [Muting a memory](#muting-a-memory). |
|
|
@@ -87,8 +149,8 @@ flag also works as `--kebab-case VALUE`, e.g. `--server-host`, `--server-port`,
|
|
|
87
149
|
`api: openai` in config.yml is driven through the OpenAI chat API (streamed; a remote
|
|
88
150
|
provider via `url:` and `api_key_env:`); every other host gets chi's own raw-prompt loop. `/model` and `--model host:model` switch hosts, and
|
|
89
151
|
the loop with them. See [Configuration](configuration.md) (`hosts:` and `api:`).
|
|
90
|
-
`--backend`, `SAMAGOTCHI_BACKEND` and a `backend:` key were removed
|
|
91
|
-
|
|
152
|
+
`--backend`, `SAMAGOTCHI_BACKEND` and a `backend:` key were removed (`backend:` warns
|
|
153
|
+
as an unknown key).
|
|
92
154
|
|
|
93
155
|
### Entrypoint scenarios
|
|
94
156
|
|
|
@@ -279,7 +341,7 @@ REPL alike:
|
|
|
279
341
|
typed comes back once it closes. The choices then go, and one line stays:
|
|
280
342
|
`? Pick a fruit → Banana`.
|
|
281
343
|
- Ctrl-C cancels the turn and keeps what you typed.
|
|
282
|
-
- In the plain REPL, Ctrl-D on an empty prompt (or `exit`, `/exit`) mid-turn
|
|
344
|
+
- In the plain REPL, Ctrl-D on an empty prompt (or `exit`, `/exit`, `/quit`) mid-turn
|
|
283
345
|
exits once the turn ends: `(exits after this turn; Ctrl-C cancels it)`
|
|
284
346
|
(`/exit --delete` deletes the session then too). In an
|
|
285
347
|
attached terminal it detaches at once and the turn goes on in the worker
|
|
@@ -290,7 +352,7 @@ turns.
|
|
|
290
352
|
|
|
291
353
|
### Images
|
|
292
354
|
|
|
293
|
-
A model that can see images gets them
|
|
355
|
+
A model that can see images gets them these ways:
|
|
294
356
|
|
|
295
357
|
- **`@path` in a prompt** (REPL, attached terminal, `-p`): `what's wrong in
|
|
296
358
|
@shot.png?`, `@~/Desktop/a.jpg`, `@"my shot.png"`. Each `@` token that names an
|
|
@@ -305,6 +367,12 @@ A model that can see images gets them three ways:
|
|
|
305
367
|
- **The Web UI**: paste or drop images into the composer. Each shows as a chip
|
|
306
368
|
(× removes it) and is sent with the message; an image alone is sent as
|
|
307
369
|
`[image: name]`. Messages show thumbnails; a click opens one full size.
|
|
370
|
+
- **`chi send --image PATH`** (repeatable, up to 20) with a message, from a
|
|
371
|
+
script or another terminal: `chi send --image shot.png -m "why is this red?"
|
|
372
|
+
3fa2`. The attached terminal and the web show it like an image typed there.
|
|
373
|
+
- **The desktop panel** (macOS, [Desktop](desktop.md#images)): a screenshot on
|
|
374
|
+
the clipboard, an image selected in Finder, or one dropped on the panel goes
|
|
375
|
+
as an attachment with the message.
|
|
308
376
|
|
|
309
377
|
Images are downscaled to a 1568 px long side (with `sips` on macOS or
|
|
310
378
|
ImageMagick; without either, a larger image is refused with a hint) and stored
|
|
@@ -350,16 +418,54 @@ touch screen), and so does each code block of a rendered answer. An answer
|
|
|
350
418
|
copies its Markdown source, not the rendered text; a code block copies just
|
|
351
419
|
its code; a prompt copies the text as you typed it.
|
|
352
420
|
|
|
353
|
-
### Web
|
|
421
|
+
### Web views
|
|
422
|
+
|
|
423
|
+
`web.view` picks how the page draws a turn: `turn` (the default, below),
|
|
424
|
+
`stage` (the running turn pinned above the composer, below) or `chat` (the
|
|
425
|
+
classic row of bubbles: one thinking block, one activity panel and one
|
|
426
|
+
answer bubble per generation).
|
|
427
|
+
|
|
428
|
+
```sh
|
|
429
|
+
chi web --web-view stage # or chat; turn is the default
|
|
430
|
+
```
|
|
431
|
+
|
|
432
|
+
The setting also supports `SAMAGOTCHI_WEB_VIEW=stage` or the global config:
|
|
433
|
+
|
|
434
|
+
```yaml
|
|
435
|
+
web:
|
|
436
|
+
view: stage
|
|
437
|
+
```
|
|
438
|
+
|
|
439
|
+
`?view=turn`, `?view=stage` or `?view=chat` on the page URL picks the view
|
|
440
|
+
for that page load, whatever the config says; the parameter is dropped
|
|
441
|
+
when you switch between the project and all-sessions views. The terminal
|
|
442
|
+
UIs are not affected.
|
|
354
443
|
|
|
355
|
-
The
|
|
356
|
-
|
|
357
|
-
|
|
358
|
-
|
|
444
|
+
**The stage view** pins the running turn above the composer, in its own
|
|
445
|
+
card, so it stays on screen without scrolling: a status row (what it is
|
|
446
|
+
doing, the step, the elapsed time), your prompt on one line, the newest
|
|
447
|
+
narration sentence (else the newest thinking one, in italics; click it for
|
|
448
|
+
the step's reasoning), the running tool with what it does, and the last
|
|
449
|
+
three calls. A card that needs you (a question, an approval, check-in)
|
|
450
|
+
sits in the stage too. The chip under it, `N steps · M tool calls` with one
|
|
451
|
+
tick per call, opens the turn view's block in place (newest step first
|
|
452
|
+
while it runs; a tick opens its step). `▾` folds the stage to its status
|
|
453
|
+
row, remembered in this browser. When the turn ends the answer shows in
|
|
454
|
+
the stage, and the whole turn moves up into the history once you are not
|
|
455
|
+
using the stage (the pointer over it, a touch or scroll in the last 4 s,
|
|
456
|
+
keyboard focus or a selection keep it) for 1.5 s; sending the next message
|
|
457
|
+
moves it at once. The history is never scrolled while a turn runs, and only
|
|
458
|
+
follows the hand-off if you were at its end. The `/` command list opens
|
|
459
|
+
over the stage's lower edge.
|
|
460
|
+
|
|
461
|
+
**The turn view** shows a turn as *one block* where the work
|
|
462
|
+
happens. The running
|
|
359
463
|
generation is the live part at the bottom (its thinking, its narration, its
|
|
360
464
|
tool rows), the earlier ones stack above it collapsed to one line each
|
|
361
|
-
(their narration's first line, else
|
|
362
|
-
count), expandable for inspection.
|
|
465
|
+
(their narration's first line, else their first call's title such as
|
|
466
|
+
`edit lib/a.rb`, and a call count), expandable for inspection. A tool row
|
|
467
|
+
says what the call did: a file's path relative to the session's folder, a
|
|
468
|
+
command without its leading `cd … &&` (the full parameters on hover). The live thinking is one line: the
|
|
363
469
|
newest complete sentence, changing at most once per 1.5 s. Click it for the
|
|
364
470
|
full text; a peek is per step (the next step's thinking starts closed
|
|
365
471
|
again). When a step ends its thinking closes to a plain `thinking` line
|
|
@@ -375,23 +481,10 @@ from the saved messages (each step's thinking, narration, tool parameters
|
|
|
375
481
|
and output, the output capped at 2000 characters) and the timing records
|
|
376
482
|
(status and duration per row). On an `api: openai` host the model's
|
|
377
483
|
reasoning is saved with each step for this (never sent back to the model);
|
|
378
|
-
steps saved before that have none, so they show no thinking.
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
```
|
|
383
|
-
|
|
384
|
-
The setting also supports `SAMAGOTCHI_WEB_TURN_VIEW=false` or the global config:
|
|
385
|
-
|
|
386
|
-
```yaml
|
|
387
|
-
web:
|
|
388
|
-
turn_view: false
|
|
389
|
-
```
|
|
390
|
-
|
|
391
|
-
`?view=chat` on the page URL forces the classic chat view for that page load
|
|
392
|
-
and `?view=turn` the turn view, whatever the config says; the parameter is dropped
|
|
393
|
-
when you switch between the project and all-sessions views. The terminal
|
|
394
|
-
UIs are not affected.
|
|
484
|
+
steps saved before that have none, so they show no thinking. An `edit` or
|
|
485
|
+
`write` row has a closed `diff +3 −1` under it that opens to the change it
|
|
486
|
+
made (up to 120 lines or 8 KB), live and after a reload (see
|
|
487
|
+
[Guardrails](guardrails.md#ask) for the diff an approval shows first).
|
|
395
488
|
|
|
396
489
|
### Web annotate presets
|
|
397
490
|
|
|
@@ -419,6 +512,51 @@ web:
|
|
|
419
512
|
there means the default, not "none": use `""` in the file or on the command
|
|
420
513
|
line. A `chi web` that already runs keeps its list; restart it.
|
|
421
514
|
|
|
515
|
+
### chi web on your phone
|
|
516
|
+
|
|
517
|
+
`chi web` listens on 127.0.0.1 only. `--web-host lan` (or `web.host: lan`
|
|
518
|
+
in config.yml) also opens it on this machine's private IPv4 address, for a
|
|
519
|
+
phone on the same Wi-Fi:
|
|
520
|
+
|
|
521
|
+
```sh
|
|
522
|
+
chi web --web-host lan
|
|
523
|
+
```
|
|
524
|
+
|
|
525
|
+
```
|
|
526
|
+
Chi Web on http://127.0.0.1:4567/?dir=/Users/me/projects/app (public: …)
|
|
527
|
+
LAN: http://192.168.1.55:4567/?token=… ← anyone with this link can run commands as you
|
|
528
|
+
Plain http: the link and your traffic can be read by anyone on this Wi-Fi.
|
|
529
|
+
<the link's QR code>
|
|
530
|
+
```
|
|
531
|
+
|
|
532
|
+
Scan the QR code with the phone's camera. The page trades the token in the
|
|
533
|
+
link for a cookie (kept 400 days) and drops it from the address, so a
|
|
534
|
+
bookmark or a home-screen icon keeps working across restarts. A home-screen
|
|
535
|
+
web app on iOS has cookies of its own: if it opens on "needs chi web's
|
|
536
|
+
access token", paste the token there (the part of the link after
|
|
537
|
+
`token=`), or open the link once in it.
|
|
538
|
+
|
|
539
|
+
- Every request from another machine needs the token (the cookie, or
|
|
540
|
+
`Authorization: Bearer <token>` for curl); this Mac (127.0.0.1) needs none.
|
|
541
|
+
Without it the page says how to get in and the API answers 401.
|
|
542
|
+
- The token lives in `$XDG_STATE_HOME/samagotchi/web-token` (0600).
|
|
543
|
+
`chi web --new-token` replaces it: a running chi web takes the new one at
|
|
544
|
+
once, and every phone has to scan the new QR code.
|
|
545
|
+
- A second `chi web` prints the LAN link and QR code again. A plain
|
|
546
|
+
`chi web --web-host lan` while a chi web without LAN access runs asks you
|
|
547
|
+
to stop that one first. `chi self` says whether chi web runs on the LAN.
|
|
548
|
+
- `lan` picks the first private address (10.x, 172.16–31.x, 192.168.x) on an
|
|
549
|
+
interface that is up, not a VPN tunnel, bridge, VM or container, and names
|
|
550
|
+
the others; `web.host: 10.0.0.3` picks one yourself. An address outside
|
|
551
|
+
those ranges (a Tailscale 100.x one, a public one) works, with a warning.
|
|
552
|
+
After the address changes (a new Wi-Fi), restart chi web. IPv6 isn't
|
|
553
|
+
offered.
|
|
554
|
+
- It is plain http: the token and everything you do travel unencrypted on
|
|
555
|
+
the Wi-Fi. Use it on your home network, never on a shared one (a café, an
|
|
556
|
+
office guest network), and run `chi web --new-token` if a link leaks.
|
|
557
|
+
- On http the browser has no notifications: the bell is hidden on the phone,
|
|
558
|
+
and the tab title still counts what needs you.
|
|
559
|
+
|
|
422
560
|
## Runtime Model Switch (Assist Mode)
|
|
423
561
|
|
|
424
562
|
In interactive assist mode, you can switch the request model without restarting:
|
|
@@ -506,38 +644,23 @@ Behavior details:
|
|
|
506
644
|
|
|
507
645
|
## Status Line
|
|
508
646
|
|
|
509
|
-
|
|
510
|
-
|
|
511
|
-
|
|
512
|
-
Behavior:
|
|
513
|
-
|
|
514
|
-
- A static status line is printed before the next `>` prompt in assist mode.
|
|
515
|
-
- During spinner rendering, status details are rendered in the spinner block.
|
|
516
|
-
- When llama.cpp streaming payload includes usage fields, status prefers server-derived token telemetry (`p`, `c`, `t`) and context percent.
|
|
517
|
-
- If server usage fields are absent, status falls back to the `:context_status` estimate telemetry.
|
|
518
|
-
- When a memory is loaded between tool rounds, the spinner line includes a `loaded: <memory>` notification immediately after the spinner frame.
|
|
519
|
-
- After responses, memory details are shown via the same unified `status>` line.
|
|
520
|
-
- The legacy standalone `memories>` summary line is no longer emitted.
|
|
521
|
-
- With `--mute`, the sticky and idle rows add `muted: <names>` after `mem:` (the
|
|
522
|
-
spinner row doesn't). Attached, `mem:` shows the used memories and the
|
|
523
|
-
session's `--memory` list before the first turn records them.
|
|
524
|
-
|
|
525
|
-
Configuration:
|
|
526
|
-
|
|
527
|
-
- `SAMAGOTCHI_STATUS_LINE` (default `on`): set to `off`, `false`, or `0` to disable status-line rendering.
|
|
528
|
-
- `SAMAGOTCHI_STATUS_WIDTH_MODE` (default `terminal_cap`): one of `terminal_cap`, `fixed`.
|
|
529
|
-
- `SAMAGOTCHI_STATUS_MAX_WIDTH` (default `160`): maximum width used by `terminal_cap`.
|
|
530
|
-
- `SAMAGOTCHI_STATUS_FIXED_WIDTH` (default `120`): fixed width used by `fixed` mode.
|
|
647
|
+
The REPL and attached mode show one status row under the prompt, drawn when what it says
|
|
648
|
+
changes:
|
|
531
649
|
|
|
532
|
-
|
|
533
|
-
|
|
534
|
-
|
|
535
|
-
- `fixed`: use `SAMAGOTCHI_STATUS_FIXED_WIDTH`, single-line with `+N` overflow indicator.
|
|
536
|
-
|
|
537
|
-
Notes:
|
|
650
|
+
```
|
|
651
|
+
status> model=Qwen3.6-35B | ↳ 3f2a1c9e | ctx=12.3% (under20) | mem: notes, cli_usage | muted: gh-helper
|
|
652
|
+
```
|
|
538
653
|
|
|
539
|
-
-
|
|
540
|
-
|
|
654
|
+
- `model=`: the model in use, `(default: …)` beside it when it isn't the config's default, and
|
|
655
|
+
`model=<served> (served; asked <name>)` when the server said it served another model.
|
|
656
|
+
- `↳ <id>`: the session that delegated this one.
|
|
657
|
+
- `ctx=`: the kernel's context estimate and its bucket, updated during a turn (on hosts that
|
|
658
|
+
report none, `api: openai`, at the turn's end).
|
|
659
|
+
- `mem:`: the memories the session read, with its `--memory` list; `muted:` its `--mute` list.
|
|
660
|
+
Up to 8 names each, then `+N`.
|
|
661
|
+
- The row is cut to the terminal's width. Without a live region (output or input not a terminal,
|
|
662
|
+
`TERM=dumb`) it prints as a line when it changes.
|
|
663
|
+
- `SAMAGOTCHI_STATUS_LINE` / `status.line` (default `on`): `off`, `false` or `0` hides it.
|
|
541
664
|
|
|
542
665
|
## Tool Tally
|
|
543
666
|
|
|
@@ -552,37 +675,48 @@ It lists the top 3 tools by count (ties go to the tool used first), the failed c
|
|
|
552
675
|
(a call a guardrail or an approval blocked counts as a call, not as failed) and the last
|
|
553
676
|
call with its parameters, cut to the terminal width. It starts over with each turn.
|
|
554
677
|
|
|
555
|
-
-
|
|
678
|
+
- The REPL and attached mode: the second row of the activity slot, shown while the slot is (the
|
|
556
679
|
model generating or a tool running). Joining a turn mid-way seeds it from the turn so far.
|
|
557
|
-
- The REPL (`--no-shared`): a row under the spinner row. The spinner stops while tools
|
|
558
|
-
run (the `tool>` lines show them), so the tally shows while the model generates
|
|
559
|
-
between tool rounds; the spinner block is one row taller from then on.
|
|
560
680
|
- The web: the activity panel's summary reads `activity · 12 tool calls (2 failed) · execute ×7 · …`
|
|
561
681
|
(without `last:`: the rows show it).
|
|
562
682
|
|
|
563
|
-
##
|
|
683
|
+
## Activity Row
|
|
684
|
+
|
|
685
|
+
While a turn runs, one row above the prompt says what it is doing, the spinner frame first
|
|
686
|
+
(the REPL and attached mode alike):
|
|
564
687
|
|
|
565
|
-
|
|
566
|
-
|
|
567
|
-
|
|
568
|
-
|
|
569
|
-
|
|
570
|
-
|
|
688
|
+
- `| thinking…` while the model starts, then `| thinking · <sentence>`: the newest complete
|
|
689
|
+
sentence of its thinking, or of its answer (`writing ·`), like the web's thinking ticker: the
|
|
690
|
+
same sentence rules (a list number such as `118.` is no sentence end; a newline is one), and the
|
|
691
|
+
row changes at most once every 1.5 s so it doesn't flicker. A long sentence is cut with `…`; a
|
|
692
|
+
Qwen `TURN:` prefix and inline markdown are left out.
|
|
693
|
+
- `| waiting for the first token… 5s` after 2 s with nothing streamed.
|
|
694
|
+
- `| running execute…` while a tool runs (its `tool>` line prints when it ends).
|
|
695
|
+
- `| retrying (1/4 in 0.5s): Errno::ECONNREFUSED` while a network error is retried.
|
|
696
|
+
- `| mcp: starting servers…` while a plugin's slow setup (an init task) runs, between turns too;
|
|
697
|
+
`mcp> ✓ …` prints when it is done.
|
|
571
698
|
|
|
572
|
-
|
|
573
|
-
|
|
574
|
-
|
|
699
|
+
The spinner turns with time, so a turn that gets no chunks still looks alive. Without a live
|
|
700
|
+
region (output or input not a terminal, `TERM=dumb`) there is no activity row; the turn's lines
|
|
701
|
+
still print as they end.
|
|
575
702
|
|
|
576
703
|
## Thinking-Phase Cancellation
|
|
577
704
|
|
|
578
|
-
|
|
705
|
+
While a turn runs, you can cancel it without exiting the process:
|
|
579
706
|
|
|
580
707
|
- Press `Ctrl-C` to cancel the active request.
|
|
581
708
|
|
|
582
709
|
Behavior notes:
|
|
583
710
|
|
|
584
711
|
- Cancellation returns control to the prompt immediately; what you typed there stays.
|
|
585
|
-
-
|
|
712
|
+
- The turn ends with one line, `✕ turn canceled (Ctrl-C) · 3.1s`, and for a prompt turn a dim
|
|
713
|
+
`partial progress kept; !rollback restores the pre-turn state` under it (a canceled continue is back where it
|
|
714
|
+
started). A failed turn ends with `✕ turn failed: <summary> · 2.0s` and a dim `prompt restored for retry`. The
|
|
715
|
+
REPL and attached mode say the same; the web says `✕ canceled (Ctrl-C)` (`stopped` for its Stop button,
|
|
716
|
+
`by a hook` for a hook's).
|
|
717
|
+
- Visible text the canceled request had streamed stays in the conversation, marked `[interrupted]`, so the next
|
|
718
|
+
message (or a continue) picks up from the half-finished reply; the canceled request's thinking and any unfinished
|
|
719
|
+
tool call are dropped.
|
|
586
720
|
|
|
587
721
|
## Iteration Limit Behavior
|
|
588
722
|
|
data/docs/configuration.md
CHANGED
|
@@ -28,7 +28,7 @@ server: # the model server when there is no hosts: map below
|
|
|
28
28
|
host: 192.0.2.10
|
|
29
29
|
port: 8081
|
|
30
30
|
thinking:
|
|
31
|
-
|
|
31
|
+
level: default # off | low | medium | high | default; see "Thinking"
|
|
32
32
|
|
|
33
33
|
# Multi-host (optional): aggregated /models and per-model routing.
|
|
34
34
|
# A bare default.model uses the default host; host:model pins to a host.
|
|
@@ -87,6 +87,11 @@ Behavior:
|
|
|
87
87
|
`config: unknown key 'default.modle' (did you mean 'default.model'?)`.
|
|
88
88
|
Names you choose under the maps below (host names, model ids) don't warn.
|
|
89
89
|
- Environment variables and CLI flags win over config-file values.
|
|
90
|
+
- An edit to the file needs no restart of `chi web`: a new session's worker
|
|
91
|
+
reads the file when it starts, and a running chi takes a changed value the
|
|
92
|
+
next time it reads that setting (a worker keeps `hosts:` and what it set up
|
|
93
|
+
at its start until it is stopped). A worker gets the CLI flags of the chi
|
|
94
|
+
that started it through its environment.
|
|
90
95
|
- Workers inherit hosts via `SAMAGOTCHI_HOSTS_JSON` propagated through `SessionManager.spawn_options`.
|
|
91
96
|
|
|
92
97
|
This lets you run `chi` without repeating common defaults such as model
|
|
@@ -114,15 +119,9 @@ one shell (`SAMAGOTCHI_LOG_LEVEL=debug chi`); keep lasting choices in the file.
|
|
|
114
119
|
Most settings also have a CLI flag: the dotted name in kebab case
|
|
115
120
|
(`--server-read-timeout 900`); `chi --help` lists them.
|
|
116
121
|
|
|
117
|
-
|
|
118
|
-
|
|
119
|
-
|
|
120
|
-
(`SAMAGOTCHI_DEFAULT_MODEL: my-model`). They are still read, but every run
|
|
121
|
-
warns (`config key 'SAMAGOTCHI_DEFAULT_MODEL' is legacy UPPER — use
|
|
122
|
-
'default.model'`). When a file has both, the nested key wins and the warning
|
|
123
|
-
names both (`both 'SAMAGOTCHI_DEFAULT_MODEL' and 'default.model' are set;
|
|
124
|
-
using 'default.model', remove the flat key`). Move each to its nested form (`default: {model: my-model}`) and delete
|
|
125
|
-
the flat line. `/model --default` already writes the nested form.
|
|
122
|
+
An environment name used as a top-level key (`SAMAGOTCHI_DEFAULT_MODEL: my-model`,
|
|
123
|
+
the old flat form) is not read: it warns as an unknown key with the nested one
|
|
124
|
+
to use (`did you mean 'default.model'?`).
|
|
126
125
|
|
|
127
126
|
## Model Server Transport
|
|
128
127
|
|
|
@@ -214,6 +213,20 @@ hosts:
|
|
|
214
213
|
|
|
215
214
|
`chi self` shows the variable and whether it is set (`api key FIREWORKS_API_KEY (set)`).
|
|
216
215
|
|
|
216
|
+
Every host with `api_key_env:` sends `Authorization: Bearer <key>` on each request,
|
|
217
|
+
chat or raw-prompt alike, so a llama.cpp started with `--api-key` works as a native
|
|
218
|
+
host too (`chi bootstrap --key-env VAR` writes such an entry). An unset variable
|
|
219
|
+
fails the turn before any request (`set VAR`); a 401/403 names the variable to
|
|
220
|
+
check, or, on a host without `api_key_env:`, suggests adding it:
|
|
221
|
+
|
|
222
|
+
```yaml
|
|
223
|
+
hosts:
|
|
224
|
+
box:
|
|
225
|
+
host: 192.0.2.20
|
|
226
|
+
port: 8080
|
|
227
|
+
api_key_env: BOX_LLAMA_KEY
|
|
228
|
+
```
|
|
229
|
+
|
|
217
230
|
For models on that host, chi uses the chat loop (its own OpenAI chat adapter): it
|
|
218
231
|
takes the OpenAI base (`url:`, else `http://HOST:PORT/v1`) and streams messages plus
|
|
219
232
|
function schemas from `/v1/chat/completions`; the model's reasoning (`reasoning_content`)
|
|
@@ -355,6 +368,72 @@ models:
|
|
|
355
368
|
`SAMAGOTCHI_HOSTS_JSON`, and reads `models:` from the config file each turn: after changing a host's `sampling:`,
|
|
356
369
|
stop the session's worker (`chi sessions stop`) for it to take effect.
|
|
357
370
|
|
|
371
|
+
## Thinking
|
|
372
|
+
|
|
373
|
+
How much a model thinks before it answers. One level, `off`, `low`, `medium`, `high` or `default`, set per model,
|
|
374
|
+
per host or for everything:
|
|
375
|
+
|
|
376
|
+
```yaml
|
|
377
|
+
thinking:
|
|
378
|
+
level: default # every model without its own level
|
|
379
|
+
hosts:
|
|
380
|
+
openrouter:
|
|
381
|
+
url: https://openrouter.ai/api/v1
|
|
382
|
+
api: openai
|
|
383
|
+
api_key_env: OPENROUTER_API_KEY
|
|
384
|
+
thinking: low
|
|
385
|
+
models:
|
|
386
|
+
qwen3.6-35b-a3b:
|
|
387
|
+
thinking: off # unquoted off works (YAML reads it as false)
|
|
388
|
+
```
|
|
389
|
+
|
|
390
|
+
- `default` sends nothing: the provider's or the chat template's own default, which is chi's behaviour without the
|
|
391
|
+
setting. For some hybrid models that default is *no* thinking (DeepSeek V3.1 on OpenRouter); `medium` turns it on.
|
|
392
|
+
`on` isn't a level.
|
|
393
|
+
- Order, first set wins: `--thinking LEVEL` or `SAMAGOTCHI_THINKING_LEVEL`, then the `models:` entry (found the way
|
|
394
|
+
`profile:` is), then the `hosts:` entry, then `thinking.level` in the file, then `default`. Anything else than a
|
|
395
|
+
level warns once and counts as unset.
|
|
396
|
+
- The flag reaches the sessions that start with it; a session already running keeps its level. `models:` levels
|
|
397
|
+
and `thinking.level` in the file are read every turn; a host's `thinking:` reaches a worker when it starts, as
|
|
398
|
+
its `sampling:` does (`chi sessions stop` to change it).
|
|
399
|
+
- `/model` shows the level and where it came from (`thinking: off (models: qwen3.6-35b-a3b)`), `chi self` too.
|
|
400
|
+
- The idle recap and plugins' side questions always ask with thinking off, whatever the level.
|
|
401
|
+
|
|
402
|
+
What each backend gets:
|
|
403
|
+
|
|
404
|
+
| Backend | `off` | `low` / `medium` / `high` |
|
|
405
|
+
|---|---|---|
|
|
406
|
+
| native (`/completion`), `qwen36` | an empty thought after the assistant cue, and no turn preamble | no knob: thinking stays as the model has it, one notice |
|
|
407
|
+
| native, `gemma4` | no `<\|think\|>` token at the start of the system prompt | no knob, one notice |
|
|
408
|
+
| chat host (`api: openai`) | `chat_template_kwargs: {enable_thinking: false}` and `reasoning_effort: "none"` | `reasoning_effort: <level>` |
|
|
409
|
+
|
|
410
|
+
On chat hosts: llama.cpp honours both off switches but ignores the effort (when its `/props` says
|
|
411
|
+
`chat_template_caps.supports_reasoning_effort: false`, chi says so once per session and host, from the `/props`
|
|
412
|
+
answer the turn already fetched for the window); Splash takes `reasoning_effort` (off only
|
|
413
|
+
through `none`) and scales with it; OpenRouter translates `reasoning_effort` per model (some can't turn thinking off:
|
|
414
|
+
Qwen3-30B-A3B thinks anyway, gpt-oss refuses).
|
|
415
|
+
|
|
416
|
+
When the model thinks although the level is `off`, chi says so once per session and host
|
|
417
|
+
(`thinking> warning: off wasn't honoured by … (N chars of thinking)`) and logs `thinking_not_honoured` each time.
|
|
418
|
+
When a host answers the thinking fields with an HTTP 400 about reasoning (gpt-oss: "Reasoning is mandatory"), chi
|
|
419
|
+
sends the request again without them, leaves them out for that model from then on, and says so once.
|
|
420
|
+
|
|
421
|
+
The fields go under the `sampling:` map ("Sampling"): a `sampling:` key wins over the level's, and
|
|
422
|
+
`chat_template_kwargs` merges per sub-key. A `null` there drops a field the level would send, at any depth, for a
|
|
423
|
+
host that refuses one of them:
|
|
424
|
+
|
|
425
|
+
```yaml
|
|
426
|
+
hosts:
|
|
427
|
+
strict:
|
|
428
|
+
url: https://llm.example.com/v1
|
|
429
|
+
api: openai
|
|
430
|
+
thinking: off
|
|
431
|
+
sampling: { reasoning_effort: null } # sends only enable_thinking: false
|
|
432
|
+
```
|
|
433
|
+
|
|
434
|
+
A different level changes a native model's system prompt (Gemma's token, Qwen's turn preamble), so the next turn
|
|
435
|
+
reads the whole context again once; on a chat host only the end of the prompt changes.
|
|
436
|
+
|
|
358
437
|
## Llama HTTP Timeouts
|
|
359
438
|
|
|
360
439
|
Long-running llama.cpp completions can exceed Ruby's default HTTP read timeout.
|
|
@@ -362,6 +441,7 @@ Raise these to avoid premature request failures:
|
|
|
362
441
|
|
|
363
442
|
- `server.open_timeout` (default: `10`, env `SAMAGOTCHI_SERVER_OPEN_TIMEOUT`) connection timeout in seconds.
|
|
364
443
|
- `server.read_timeout` (default: `600`, env `SAMAGOTCHI_SERVER_READ_TIMEOUT`) response read timeout in seconds.
|
|
444
|
+
Either timeout at `0` (or anything not a positive number) is its default, on every host.
|
|
365
445
|
|
|
366
446
|
```yaml
|
|
367
447
|
server:
|
|
@@ -409,15 +489,32 @@ the running server (llama.cpp's `/props`), else the window the host's model list
|
|
|
409
489
|
gives (`context_length`, `context_window`, `max_model_len` or llama.cpp's
|
|
410
490
|
`meta.n_ctx`), else `context.window_tokens`.
|
|
411
491
|
|
|
492
|
+
A `:` in a model name is often part of the id (`qwen3:8b`, `mistral:7b`,
|
|
493
|
+
`unsloth/Qwen3-8B-GGUF:Q4_K_M`), so the part before the first `:` picks a host
|
|
494
|
+
only when it is a configured host's name. An unknown prefix is refused with an
|
|
495
|
+
error naming it and the configured hosts (with a "did you mean" for a near
|
|
496
|
+
miss) when either
|
|
497
|
+
- the rest is an `org/model` id (`nosuch:anthropic/claude-sonnet-4`), or
|
|
498
|
+
- the prefix is a hosted provider's name: `openrouter`, `openai`, `anthropic`,
|
|
499
|
+
`google`, `gemini`, `groq`, `xai`, `together` or `fireworks`
|
|
500
|
+
(`openai:gpt-4o` with no `openai` host).
|
|
501
|
+
|
|
502
|
+
The check applies wherever the model comes in: `--model`, `default.model`, an
|
|
503
|
+
alias, `/model`, `chi send --new --model`, the web's new-session model and a
|
|
504
|
+
delegate's model. Any other unknown prefix (`nosuch:x`) is sent to the default
|
|
505
|
+
host as the model id.
|
|
506
|
+
|
|
412
507
|
## Llama Network Retry Behavior
|
|
413
508
|
|
|
414
509
|
Transient network failures are retried automatically with exponential backoff.
|
|
415
510
|
|
|
416
511
|
- Default retries: `5` (up to `6` total attempts including the first call).
|
|
417
512
|
- Default backoff: `0.5s`, `1s`, `2s`, `4s`, `8s`.
|
|
418
|
-
- Retry scope: transient network errors (timeouts,
|
|
513
|
+
- Retry scope: transient network errors (timeouts, reset connections, EOF/socket reachability failures),
|
|
419
514
|
HTTP 429 and HTTP 500/502/503/504/529. A `Retry-After` header replaces the backoff delay; one longer than
|
|
420
515
|
60s is not waited out and the error is reported instead.
|
|
516
|
+
- A refused connection (nothing listening) is not retried: the turn fails at once with `can't reach host <name> at
|
|
517
|
+
<address> (connection refused) — is the server running?`.
|
|
421
518
|
- A stream that has already produced output is never retried (the retry would repeat it); it fails the turn.
|
|
422
519
|
- Cancellation (`Ctrl-C`) is never retried.
|
|
423
520
|
|
|
@@ -455,7 +552,7 @@ message (before, a failed llama.cpp `/completion` ended the turn as
|
|
|
455
552
|
|
|
456
553
|
| Kind | When | Retried |
|
|
457
554
|
|---|---|---|
|
|
458
|
-
| connection |
|
|
555
|
+
| connection | reset, timed out, dropped mid-stream; refused | yes (network retry), not mid-stream; refused: no |
|
|
459
556
|
| rate limited | HTTP 429 | yes, honouring `Retry-After` |
|
|
460
557
|
| server | HTTP 5xx, llama.cpp's mid-stream `error:` event | 500/502/503/504/529 only |
|
|
461
558
|
| auth | HTTP 401/403 | no |
|
|
@@ -630,7 +727,7 @@ described in their own sections.
|
|
|
630
727
|
| `session.max_children` | `4` | | Running delegated sessions one session may have. |
|
|
631
728
|
| `session.retention_days` | `14` | yes | Delete sessions not updated for N days; `0` = forever. See [Sessions](sessions.md). |
|
|
632
729
|
| `session.max_count` | `500` | yes | Keep the newest N; `0` = uncapped. |
|
|
633
|
-
| `session.keep_status` |
|
|
730
|
+
| `session.keep_status` | none | yes | Comma list of statuses never pruned (a session a worker or `chi` has open is never pruned anyway). |
|
|
634
731
|
| `session.sweep_interval_hours` | `24` | yes | How often the retention sweep runs. |
|
|
635
732
|
| `image.max_side` | `1568` | | See "Images". |
|
|
636
733
|
| `image.max_bytes` | `3750000` | | See "Images". |
|
|
@@ -639,23 +736,23 @@ described in their own sections.
|
|
|
639
736
|
| `log.file` | state dir | yes | See "Debug Log File". |
|
|
640
737
|
| `log.disable` | `false` | yes | No file logging. |
|
|
641
738
|
| `log.level` | `info` | yes | `debug`, `info`, `warn`, `error`. |
|
|
642
|
-
| `status.line` | `on` | yes | The REPL
|
|
643
|
-
| `status.width_mode` | `terminal_cap` | yes | `terminal_cap` (terminal width up to `max_width`) or `fixed`. |
|
|
644
|
-
| `status.max_width` | `160` | yes | Cap for `terminal_cap`. |
|
|
645
|
-
| `status.fixed_width` | `120` | yes | Width for `fixed`. |
|
|
739
|
+
| `status.line` | `on` | yes | The status row under the prompt (the REPL's and attached mode's), `on` or `off`. |
|
|
646
740
|
| `context.status` | `true` | yes | Context-usage telemetry for the model. See [context telemetry](internals/context-telemetry.md). |
|
|
647
741
|
| `context.window_tokens` | server's, else 256000 | yes | Context window when the server doesn't report one. |
|
|
648
742
|
| `context.chars_per_token` | `4.0` | yes | Estimate ratio when the server reports no usage. |
|
|
649
743
|
| `context.status_thresholds` | `20,40,60,80` | yes | Percentages that trigger a status. |
|
|
650
744
|
| `context.status_cadence` | `0` | yes | Also every N rounds; `0` = thresholds only. |
|
|
651
|
-
| `thinking.ui` | `spinner` | yes | `spinner` or `off`. |
|
|
652
|
-
| `thinking.render_interval` | `0.08` | yes | Seconds between thinking redraws. |
|
|
653
745
|
| `thinking.turn_preamble` | `true` | yes | Ask a `qwen36` model to open its thinking with a short `TURN:` line (the step label). |
|
|
746
|
+
| `thinking.level` | `default` | `--thinking` | `off`, `low`, `medium`, `high` or `default` for every model; the flag and env outrank the `models:`/`hosts:` entries, the file's value doesn't. See "Thinking". |
|
|
747
|
+
| `models.<key>.thinking`, `hosts.<name>.thinking` | none | | A model's or host's level. See "Thinking". |
|
|
654
748
|
| `max_tool_output_chars` | `10000` | yes | Tool output kept in the conversation; a top-level key (see below). |
|
|
655
749
|
| `retry.max` | `5` | yes | See "Llama Network Retry Behavior". |
|
|
656
750
|
| `retry.base_delay` | `0.5` | yes | |
|
|
657
751
|
| `retry.max_delay` | `8.0` | yes | |
|
|
658
752
|
| `retry.empty_answer` | `1` | | Times a turn asks again after an empty answer (at most 3, `0` = off). See "Llama Network Retry Behavior". |
|
|
753
|
+
| `update.gem` | `true` | | `false`: `chi update` never installs a newer gem (`--no-gem` for one run). See [CLI: Updating](cli.md#updating). |
|
|
754
|
+
| `update.bundles` | `true` | | `false`: `chi update` leaves the shipped bundles to `chi bundle upgrade` (`--no-bundles`). |
|
|
755
|
+
| `update.desktop` | `true` | | `false`: `chi update` leaves the desktop helper alone (`--no-desktop`). |
|
|
659
756
|
| `read.truncate_at_bytes` | `65536` | yes | A `read` result larger than this is cut to a preview. |
|
|
660
757
|
| `read.preview_bytes` | `12288` | yes | Size of that preview. |
|
|
661
758
|
| `read.hard_max_bytes` | `2097152` | yes | Largest file `read` opens. |
|
|
@@ -664,9 +761,9 @@ described in their own sections.
|
|
|
664
761
|
| `execute.preview_bytes` | `12288` | yes | |
|
|
665
762
|
| `execute.telemetry_threshold_pct` | `80` | yes | |
|
|
666
763
|
| `web.port` | `4567` | `--port` | `chi web`'s port. See [CLI](cli.md). |
|
|
667
|
-
| `web.host` | `127.0.0.1` | yes | `127.0.0.1`, `::1` or `localhost
|
|
764
|
+
| `web.host` | `127.0.0.1` | yes | `127.0.0.1`, `::1` or `localhost`; `lan` (this machine's private IPv4 address) or one of its IPv4 addresses also opens `chi web` to the network, with an access token (`chi web --new-token` replaces it). Anything else binds `127.0.0.1` with a warning. See [CLI: chi web on your phone](cli.md#chi-web-on-your-phone). |
|
|
668
765
|
| `web.markdown` | `false` | yes | Render answers as Markdown in `chi web`. |
|
|
669
|
-
| `web.
|
|
766
|
+
| `web.view` | `turn` | yes | How `chi web` draws a turn: `turn` (one block per turn), `stage` (the running turn pinned above the composer) or `chat` (the row of bubbles). See [CLI](cli.md#web-views). |
|
|
670
767
|
| `web.annotate_presets` | `Agreed\|Could you please elaborate?` | yes | Quick replies next to Annotate in `chi web`, `\|`-separated (a YAML list works too); `""` in the file or on the CLI leaves only Annotate (an empty env value means the default). See [CLI](cli.md#web-annotate-presets). |
|
|
671
768
|
| `history.file` | state dir | | Prompt history path. |
|
|
672
769
|
| `no_interrupt` | `false` | `--no-interrupt` | Raise the tool-call limit of a turn to 1000; a top-level key. |
|