samagotchi 0.2.0 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (131) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +198 -1
  3. data/README.md +56 -4
  4. data/bin/chi +118 -50
  5. data/docs/cli.md +184 -9
  6. data/docs/configuration.md +333 -47
  7. data/docs/desktop.md +45 -4
  8. data/docs/guardrails.md +11 -0
  9. data/docs/hooks.md +208 -5
  10. data/docs/plugins.md +68 -2
  11. data/docs/releasing.md +23 -13
  12. data/docs/sessions.md +45 -17
  13. data/lib/samagotchi/answer_display.rb +95 -0
  14. data/lib/samagotchi/archive_store.rb +90 -0
  15. data/lib/samagotchi/bootstrap/config_writer.rb +342 -0
  16. data/lib/samagotchi/bootstrap/probe.rb +262 -0
  17. data/lib/samagotchi/bootstrap_command.rb +347 -0
  18. data/lib/samagotchi/bridge/pending_card.rb +89 -0
  19. data/lib/samagotchi/bridge/turn_accumulator.rb +15 -3
  20. data/lib/samagotchi/bridge.rb +13 -1
  21. data/lib/samagotchi/bridge_client.rb +6 -2
  22. data/lib/samagotchi/bundles/check-in/manifest.yml +10 -0
  23. data/lib/samagotchi/bundles/check-in/plugin.rb +244 -0
  24. data/lib/samagotchi/bundles/source-links/hooks/source_links.rb +531 -0
  25. data/lib/samagotchi/bundles/source-links/manifest.yml +14 -0
  26. data/lib/samagotchi/bundles/source-links/source_links.md +5 -0
  27. data/lib/samagotchi/bundles/system/config_modification_protocol.md +10 -6
  28. data/lib/samagotchi/bundles/system/delegated.md +6 -7
  29. data/lib/samagotchi/bundles/system/manifest.yml +4 -4
  30. data/lib/samagotchi/bundles/system/self_map.md +8 -2
  31. data/lib/samagotchi/client.rb +81 -19
  32. data/lib/samagotchi/commands/registry.rb +8 -0
  33. data/lib/samagotchi/config.rb +252 -48
  34. data/lib/samagotchi/desktop/macos/App.swift +12 -8
  35. data/lib/samagotchi/desktop/macos/ChiRunner.swift +17 -9
  36. data/lib/samagotchi/desktop/macos/Images.swift +113 -0
  37. data/lib/samagotchi/desktop/macos/Info.plist.erb +6 -0
  38. data/lib/samagotchi/desktop/macos/Panel.swift +180 -25
  39. data/lib/samagotchi/desktop/macos.rb +59 -8
  40. data/lib/samagotchi/desktop_command.rb +6 -3
  41. data/lib/samagotchi/edit_preview.rb +82 -0
  42. data/lib/samagotchi/empty_answer_retry.rb +43 -0
  43. data/lib/samagotchi/engine.rb +434 -140
  44. data/lib/samagotchi/gem_update.rb +89 -0
  45. data/lib/samagotchi/guardrails/approval.rb +35 -4
  46. data/lib/samagotchi/guardrails/load_failures.rb +9 -3
  47. data/lib/samagotchi/guardrails/scratch_writes.rb +40 -0
  48. data/lib/samagotchi/guardrails.rb +1 -0
  49. data/lib/samagotchi/hooks/registry.rb +24 -5
  50. data/lib/samagotchi/host_registry.rb +9 -12
  51. data/lib/samagotchi/idle_client.rb +24 -15
  52. data/lib/samagotchi/idle_recap.rb +5 -1
  53. data/lib/samagotchi/idle_reminders.rb +2 -2
  54. data/lib/samagotchi/image_store.rb +10 -6
  55. data/lib/samagotchi/kernel_loop.rb +73 -94
  56. data/lib/samagotchi/live_versions.rb +59 -0
  57. data/lib/samagotchi/llm/api_key.rb +41 -0
  58. data/lib/samagotchi/llm/chat_loop.rb +132 -29
  59. data/lib/samagotchi/llm/errors.rb +41 -9
  60. data/lib/samagotchi/llm/http.rb +57 -17
  61. data/lib/samagotchi/llm/openai_chat.rb +17 -30
  62. data/lib/samagotchi/log_subscriber.rb +18 -3
  63. data/lib/samagotchi/memory_bundle/installer.rb +65 -63
  64. data/lib/samagotchi/memory_bundle/provenance.rb +51 -12
  65. data/lib/samagotchi/memory_bundle/shipped_update.rb +157 -0
  66. data/lib/samagotchi/memory_bundle/status.rb +4 -1
  67. data/lib/samagotchi/memory_bundle/system_bundle.rb +81 -53
  68. data/lib/samagotchi/model_profile.rb +24 -1
  69. data/lib/samagotchi/plugin/context.rb +22 -1
  70. data/lib/samagotchi/plugin/sessions.rb +3 -1
  71. data/lib/samagotchi/prompt.rb +4 -2
  72. data/lib/samagotchi/reminder_store.rb +1 -9
  73. data/lib/samagotchi/reply_wait.rb +126 -0
  74. data/lib/samagotchi/sampling_settings.rb +58 -0
  75. data/lib/samagotchi/self_report.rb +18 -3
  76. data/lib/samagotchi/send_command.rb +252 -11
  77. data/lib/samagotchi/session.rb +52 -11
  78. data/lib/samagotchi/session_archive_command.rb +107 -0
  79. data/lib/samagotchi/session_commands.rb +46 -7
  80. data/lib/samagotchi/session_manager.rb +115 -25
  81. data/lib/samagotchi/session_metrics.rb +222 -106
  82. data/lib/samagotchi/steer.rb +72 -0
  83. data/lib/samagotchi/terminal_ui/attached_loop.rb +57 -28
  84. data/lib/samagotchi/terminal_ui/event_renderer.rb +21 -11
  85. data/lib/samagotchi/terminal_ui/formatting.rb +40 -8
  86. data/lib/samagotchi/terminal_ui/input_support.rb +7 -19
  87. data/lib/samagotchi/terminal_ui/question_prompt.rb +35 -0
  88. data/lib/samagotchi/terminal_ui.rb +134 -247
  89. data/lib/samagotchi/text_diff.rb +181 -0
  90. data/lib/samagotchi/thinking.rb +115 -0
  91. data/lib/samagotchi/tool_activity.rb +3 -1
  92. data/lib/samagotchi/tool_runner.rb +34 -1
  93. data/lib/samagotchi/tools/ask_user_question.rb +41 -33
  94. data/lib/samagotchi/tools/builtins.rb +15 -4
  95. data/lib/samagotchi/tools/delegate_wait.rb +26 -69
  96. data/lib/samagotchi/tools/edit.rb +23 -9
  97. data/lib/samagotchi/tools/execute.rb +52 -14
  98. data/lib/samagotchi/tools/task_runtime.rb +19 -0
  99. data/lib/samagotchi/tools/task_wait.rb +27 -3
  100. data/lib/samagotchi/tools/write.rb +4 -0
  101. data/lib/samagotchi/turn_flow.rb +12 -2
  102. data/lib/samagotchi/turn_note.rb +60 -6
  103. data/lib/samagotchi/update_command.rb +308 -0
  104. data/lib/samagotchi/update_hint.rb +59 -0
  105. data/lib/samagotchi/version.rb +1 -1
  106. data/lib/samagotchi/vision_support.rb +7 -9
  107. data/lib/samagotchi/web/app.rb +91 -7
  108. data/lib/samagotchi/web/message_parts.rb +8 -3
  109. data/lib/samagotchi/web/public/activity.js +13 -1
  110. data/lib/samagotchi/web/public/annotate_presets.js +26 -0
  111. data/lib/samagotchi/web/public/annotations.js +13 -0
  112. data/lib/samagotchi/web/public/app.js +472 -111
  113. data/lib/samagotchi/web/public/card.js +5 -3
  114. data/lib/samagotchi/web/public/chat_view.js +13 -1
  115. data/lib/samagotchi/web/public/copy.js +20 -4
  116. data/lib/samagotchi/web/public/ctx.js +15 -0
  117. data/lib/samagotchi/web/public/data.js +23 -6
  118. data/lib/samagotchi/web/public/diff_view.js +58 -0
  119. data/lib/samagotchi/web/public/format.js +9 -0
  120. data/lib/samagotchi/web/public/index.html +60 -3
  121. data/lib/samagotchi/web/public/notify.js +175 -0
  122. data/lib/samagotchi/web/public/question_card.js +5 -2
  123. data/lib/samagotchi/web/public/sessions_list.js +7 -0
  124. data/lib/samagotchi/web/public/timing.js +39 -14
  125. data/lib/samagotchi/web/public/turn_events.js +75 -5
  126. data/lib/samagotchi/web/public/turn_view.js +49 -8
  127. data/lib/samagotchi/web/server.rb +8 -4
  128. data/lib/samagotchi/web/session_hub.rb +2 -1
  129. data/lib/samagotchi/web/session_summary.rb +24 -1
  130. data/lib/samagotchi/worker.rb +16 -4
  131. metadata +31 -1
@@ -2,26 +2,38 @@
2
2
 
3
3
  ## Global Config File
4
4
 
5
- Chi can preload a global config file and expose those entries as environment
6
- variables before the app boots.
5
+ Chi reads its settings from three places; the first that sets a value wins:
6
+
7
+ 1. a CLI flag (`--server-port 8081`),
8
+ 2. an environment variable (`SAMAGOTCHI_SERVER_PORT=8081`),
9
+ 3. the global config file (`server: {port: 8081}`),
10
+
11
+ then the built-in default. [All settings](#all-settings) lists them.
12
+
13
+ `chi bootstrap HOST[:PORT]` writes a first config for a model server, or adds
14
+ it to this file as a `hosts:` entry (see [First setup](cli.md#first-setup)).
7
15
 
8
16
  Default path:
9
17
 
10
18
  - `$XDG_CONFIG_HOME/samagotchi/config.yml`
11
19
  - Fallback when `XDG_CONFIG_HOME` is unset: `~/.config/samagotchi/config.yml`
12
20
 
13
- Example:
21
+ Example (every section is optional except `default.model`; README.md has a
22
+ minimal one):
14
23
 
15
24
  ```yaml
16
- SAMAGOTCHI_DEFAULT_MODEL: Qwen3-14B-Instruct
17
- server:
25
+ default:
26
+ model: Qwen3-14B-Instruct
27
+ server: # the model server when there is no hosts: map below
18
28
  host: 192.0.2.10
19
29
  port: 8081
20
- SAMAGOTCHI_THINKING_UI: spinner
30
+ thinking:
31
+ ui: spinner
32
+ level: default # off | low | medium | high | default; see "Thinking"
21
33
 
22
34
  # Multi-host (optional): aggregated /models and per-model routing.
23
- # Bare SAMAGOTCHI_DEFAULT_MODEL uses the default host; host:model pins to a host.
24
- # Transport per host overrides SAMAGOTCHI_SERVER_TRANSPORT; api: openai makes a
35
+ # A bare default.model uses the default host; host:model pins to a host.
36
+ # transport: on a host overrides server.transport; api: openai makes a
25
37
  # host use the OpenAI chat API instead of chi's raw prompt.
26
38
  hosts:
27
39
  main:
@@ -58,36 +70,61 @@ session:
58
70
  # keep_empty: false # true keeps sessions nothing happened in (default: deleted when left)
59
71
  # max_children: 4 # running sessions one session may have delegated at a time (the delegate tool)
60
72
 
61
- # Baseline memories preloaded into the system prompt (same shape as --memory).
73
+ # Baseline memories preloaded into the system prompt (same shape as --memory);
74
+ # name entries you have, or chi warns at each start.
62
75
  # CLI --memory entries are appended after these, deduped.
63
76
  memories:
64
- - system/user_preferences
65
- - project/feature-env-template
77
+ # - system/user_preferences
78
+ # - project/feature-env-template
66
79
  ```
67
80
 
68
81
  Behavior:
69
82
 
70
83
  - The file is optional.
71
- - Top level is a YAML mapping of scalar env overrides plus nested sections.
72
- - Real environment variables still win over config-file values.
84
+ - Top level is a YAML mapping of sections (`default:`, `server:`, `recap:`, …)
85
+ and the maps described below. Keys are lower `snake_case`, one level per
86
+ dot: `server.read_timeout` is `server: {read_timeout: 600}`.
87
+ - A key chi doesn't read warns at start, with the closest known key:
88
+ `config: unknown key 'default.modle' (did you mean 'default.model'?)`.
89
+ Names you choose under the maps below (host names, model ids) don't warn.
90
+ - Environment variables and CLI flags win over config-file values.
73
91
  - Workers inherit hosts via `SAMAGOTCHI_HOSTS_JSON` propagated through `SessionManager.spawn_options`.
74
92
 
75
93
  This lets you run `chi` without repeating common defaults such as model
76
94
  and llama host/port on every invocation.
77
95
 
78
- Note: The global config file supports both flat scalar entries (for env vars)
79
- and nested sections like `hosts:`, `recap:`, `hooks:`, `guardrails:` (see
80
- [Guardrails](guardrails.md)), `bundles:` (a bundle's settings for its hooks,
81
- see [Hooks: Settings](hooks.md#settings)), `model_aliases:`,
82
- `memories:`. Scalar entries are loaded as environment
83
- variables; non-scalar sections are skipped by the env-loader and parsed by
84
- their respective subsystems (e.g. the hooks system, `HostRegistry`). The
85
- `memories:` list is the persistent baseline for preloaded memory entries —
96
+ Besides the settings sections, the file holds maps that are read by their
97
+ own subsystems: `hosts:` (below), `models:` (see "Prompt profile" and
98
+ "Images"), `model_aliases:`, `hooks:` (see [Hooks](hooks.md)),
99
+ `guardrails:` (see [Guardrails](guardrails.md)), `bundles:` (a bundle's
100
+ settings for its hooks, see [Hooks: Settings](hooks.md#settings)) and
101
+ `memories:`. Their entry names (host names, model ids, aliases) are yours to
102
+ choose. The `memories:` list is the persistent baseline for preloaded memory entries —
86
103
  the same name shape as `--memory` (bare name or `scope/name`), merged under
87
104
  any per-run `--memory` values (config baseline first, deduped). A per-run
88
105
  `--mute NAME` removes an entry from the merged list for that session (see
89
106
  "Muting a memory" in cli.md).
90
107
 
108
+ ### Environment variables
109
+
110
+ Every setting has an environment variable: `SAMAGOTCHI_` plus its dotted
111
+ name in upper case, with `_` for each dot. `default.model` is
112
+ `SAMAGOTCHI_DEFAULT_MODEL`, `server.read_timeout` is
113
+ `SAMAGOTCHI_SERVER_READ_TIMEOUT`. Use them to override the file for one run or
114
+ one shell (`SAMAGOTCHI_LOG_LEVEL=debug chi`); keep lasting choices in the file.
115
+ Most settings also have a CLI flag: the dotted name in kebab case
116
+ (`--server-read-timeout 900`); `chi --help` lists them.
117
+
118
+ ### Legacy flat keys
119
+
120
+ Older configs used the environment names as top-level keys
121
+ (`SAMAGOTCHI_DEFAULT_MODEL: my-model`). They are still read, but every run
122
+ warns (`config key 'SAMAGOTCHI_DEFAULT_MODEL' is legacy UPPER — use
123
+ 'default.model'`). When a file has both, the nested key wins and the warning
124
+ names both (`both 'SAMAGOTCHI_DEFAULT_MODEL' and 'default.model' are set;
125
+ using 'default.model', remove the flat key`). Move each to its nested form (`default: {model: my-model}`) and delete
126
+ the flat line. `/model --default` already writes the nested form.
127
+
91
128
  ## Model Server Transport
92
129
 
93
130
  Chi talks to a model server over HTTP and supports three transports:
@@ -99,11 +136,13 @@ Chi talks to a model server over HTTP and supports three transports:
99
136
  continuous batching + tiered SSD KV cache) using the same `/v1/completions`
100
137
  and `/v1/models` endpoints as `mlx`.
101
138
 
102
- Select the transport with `SAMAGOTCHI_SERVER_TRANSPORT` (`llama_cpp`, `mlx`, or
103
- `omlx`). `SAMAGOTCHI_SERVER_HOST`/`SAMAGOTCHI_SERVER_PORT` are reused for all three —
139
+ > **mlx and oMLX** are not verified with recent chi versions; llama.cpp and
140
+ > OpenAI-compatible hosts (`api: openai`) are the tested paths.
141
+
142
+ Select the transport with `server.transport` (`llama_cpp`, `mlx`, or `omlx`;
143
+ env `SAMAGOTCHI_SERVER_TRANSPORT`). `server.host`/`server.port` are reused for all three —
104
144
  only the request/response shape differs. oMLX's default server port is `8000` (not
105
- `8080`), so point `SAMAGOTCHI_SERVER_PORT` at it, e.g. `SAMAGOTCHI_SERVER_PORT=8000`.
106
- With `hosts:` each entry may set `transport: llama_cpp|mlx|omlx` to override the
145
+ `8080`), so point `server.port` at it. With `hosts:` each entry may set `transport: llama_cpp|mlx|omlx` to override the
107
146
  global transport per host (`lib/samagotchi/host_registry.rb`).
108
147
 
109
148
  Each host may also set `api:`, which says how chi talks to it:
@@ -116,11 +155,11 @@ Without `api:` a host uses the raw-prompt loop, as before. The loop follows the
116
155
  model's host, so `/model other-host:model` can move a session between the two.
117
156
  Workers started by plain `chi`, `chi web` or `--attach` get the same hosts, `api:` included.
118
157
 
119
- Example for mlx-lm:
158
+ Example for mlx-lm (not verified with recent chi versions, like oMLX below):
120
159
 
121
160
  ```yaml
122
- SAMAGOTCHI_SERVER_TRANSPORT: mlx
123
161
  server:
162
+ transport: mlx
124
163
  host: 127.0.0.1
125
164
  port: 8080
126
165
  ```
@@ -132,8 +171,8 @@ mlx_lm.server --model mlx-community/Qwen3-14B-Instruct-4bit
132
171
  Example for oMLX:
133
172
 
134
173
  ```yaml
135
- SAMAGOTCHI_SERVER_TRANSPORT: omlx
136
174
  server:
175
+ transport: omlx
137
176
  host: 192.0.2.10
138
177
  port: 8000
139
178
  ```
@@ -176,6 +215,20 @@ hosts:
176
215
 
177
216
  `chi self` shows the variable and whether it is set (`api key FIREWORKS_API_KEY (set)`).
178
217
 
218
+ Every host with `api_key_env:` sends `Authorization: Bearer <key>` on each request,
219
+ chat or raw-prompt alike, so a llama.cpp started with `--api-key` works as a native
220
+ host too (`chi bootstrap --key-env VAR` writes such an entry). An unset variable
221
+ fails the turn before any request (`set VAR`); a 401/403 names the variable to
222
+ check, or, on a host without `api_key_env:`, suggests adding it:
223
+
224
+ ```yaml
225
+ hosts:
226
+ box:
227
+ host: 192.0.2.20
228
+ port: 8080
229
+ api_key_env: BOX_LLAMA_KEY
230
+ ```
231
+
179
232
  For models on that host, chi uses the chat loop (its own OpenAI chat adapter): it
180
233
  takes the OpenAI base (`url:`, else `http://HOST:PORT/v1`) and streams messages plus
181
234
  function schemas from `/v1/chat/completions`; the model's reasoning (`reasoning_content`)
@@ -205,9 +258,9 @@ oMLX's known tool-call limitation (a stream filter that strips markup) only
205
258
  affects its `/v1/chat/completions` endpoint, not the `/v1/completions` endpoint
206
259
  chi uses, so raw `[[…]]`/`<|tool_call>` markers stream through untouched.
207
260
 
208
- `SAMAGOTCHI_DEFAULT_MODEL` (config default) and `/model` (runtime effective) pick the model; the status line and `/model`
261
+ `default.model` (config default) and `/model` (runtime effective) pick the model; the status line and `/model`
209
262
  output always render the runtime effective model (showing default when diverged). Which prompt format it gets is the
210
- prompt profile (see "Prompt profile" below). How the selector reaches the request differs by transport:
263
+ prompt profile (see "Prompt profile" below). How the selector reaches the request differs by transport (mlx and oMLX: not verified with recent chi versions):
211
264
 
212
265
  - **mlx** (`mlx_lm.server`): the `model` field is omitted entirely — the server
213
266
  uses whatever was loaded via its own `--model` CLI flag.
@@ -217,7 +270,7 @@ prompt profile (see "Prompt profile" below). How the selector reaches the reques
217
270
  (case-insensitive) first, then substring, then passed through unchanged. That
218
271
  resolved id is usually prefixed (e.g. `mlx-community--gemma-3-4b-it-4bit`), so a
219
272
  short selector such as `gemma-3-4b-it-4bit` is what you set in
220
- `SAMAGOTCHI_DEFAULT_MODEL`. An unknown selector passes through raw and oMLX 404s,
273
+ `default.model`. An unknown selector passes through raw and oMLX 404s,
221
274
  listing its available models; if `/v1/models` is unreachable, samagotchi falls
222
275
  back to the raw selector and lets the server decide (its own 400/404). Either
223
276
  error fails the turn with the server's message (see "Server errors" below). Runtime
@@ -231,8 +284,8 @@ syntax, thought tags and stop sequences. That is the prompt profile, `qwen36` or
231
284
  worse output: a ChatML model under `gemma4` never hits a stop sequence, generates until its limit and then runs the
232
285
  tool calls it made up on the way. The first of these that says something wins:
233
286
 
234
- 1. `--profile NAME` (or `--model-profile NAME`), then `SAMAGOTCHI_MODEL_PROFILE`: for every model in the process,
235
- including one picked later with `/model`.
287
+ 1. `--profile NAME` (or `--model-profile NAME`), then `SAMAGOTCHI_MODEL_PROFILE` (`model.profile` has no config-file
288
+ form): for every model in the process, including one picked later with `/model`.
236
289
  2. `models:` in `config.yml`, keyed by model id or alias (case-insensitive; the name as typed, alias-resolved or
237
290
  without its host prefix):
238
291
 
@@ -274,13 +327,127 @@ Workers get `hosts:` (with `profile:`) through `SAMAGOTCHI_HOSTS_JSON` and read
274
327
  `--profile` reaches the worker a chi starts, but a worker that another process wakes later (`chi web`, `--attach`
275
328
  after an idle exit) gets that process's environment, so put a lasting choice in config.
276
329
 
330
+ ## Sampling
331
+
332
+ chi asks chat hosts (`api: openai`) for greedy decoding (`temperature: 0.0`) and sends native hosts no sampling
333
+ fields, so their server's defaults apply (llama.cpp: temperature 0.8). `sampling:` on a `hosts:` entry or a `models:`
334
+ entry sets request fields for the model's turns:
335
+
336
+ ```yaml
337
+ hosts:
338
+ work:
339
+ url: https://llm.example.com/v1
340
+ api: openai
341
+ api_key_env: WORK_API_KEY
342
+ sampling: { temperature: 0.6, presence_penalty: 1.5 }
343
+ models:
344
+ qwen3.6-35b-a3b:
345
+ sampling: { temperature: 0.6, top_p: 0.95, repeat_penalty: 1.1 }
346
+ ```
347
+
348
+ - The fields go into the request as written; chi doesn't check the names, since providers differ (`repeat_penalty`,
349
+ `min_p`, `dry_multiplier` are llama.cpp's; `presence_penalty` is OpenAI-style). A provider that refuses one fails
350
+ the turn with its error (`host work rejected the request: HTTP 400: …`; one server answered a `presence_penalty`
351
+ with "the requested logits or output transformation is not supported"): remove the last field you added from that
352
+ host's or model's `sampling:`, or set it to `null` in the model's entry. A nested map passes through too, e.g.
353
+ `chat_template_kwargs: { enable_thinking: false }`.
354
+ - A `models:` entry's fields win over its host's, field by field (the host can set a penalty and the model move only
355
+ the temperature). The entry is found the way `profile:` is (the name as typed, alias-resolved or without its host
356
+ prefix).
357
+ - `temperature: null` (or `~`) sends no temperature, so the provider's default applies (vLLM takes it from the
358
+ model's `generation_config.json`).
359
+ - Fields chi sets itself are refused with a warning: `model`, `messages`, `prompt`, `stream`, `stream_options`,
360
+ `tools`, `tool_choice`, `stop`, `n_predict`, `max_tokens`, `n`, `parallel_tool_calls`, `response_format`,
361
+ `cache_prompt`. A `sampling:` that isn't a map warns and is skipped.
362
+ - Greedy decoding can make a heavily quantized thinking model loop in its reasoning ("Let me write the reply…"
363
+ for minutes) or end with an empty answer. Qwen's own advice for its thinking models is `temperature: 0.6,
364
+ top_p: 0.95` (not greedy), with `presence_penalty` between 0 and 2 against endless repetition.
365
+ - The idle recap and side questions keep their own short, deterministic settings.
366
+
367
+ `/model` shows what applies (`sampling: temperature=0.6 presence_penalty=1.5 (hosts.work)`), and each request's
368
+ `stream` line in the debug log carries a `sampling=` field with what was sent. The fields are read every turn, so a
369
+ `/model` switch takes the new model's. A worker gets `hosts:` (with `sampling:`) when it starts, through
370
+ `SAMAGOTCHI_HOSTS_JSON`, and reads `models:` from the config file each turn: after changing a host's `sampling:`,
371
+ stop the session's worker (`chi sessions stop`) for it to take effect.
372
+
373
+ ## Thinking
374
+
375
+ How much a model thinks before it answers. One level, `off`, `low`, `medium`, `high` or `default`, set per model,
376
+ per host or for everything:
377
+
378
+ ```yaml
379
+ thinking:
380
+ level: default # every model without its own level
381
+ hosts:
382
+ openrouter:
383
+ url: https://openrouter.ai/api/v1
384
+ api: openai
385
+ api_key_env: OPENROUTER_API_KEY
386
+ thinking: low
387
+ models:
388
+ qwen3.6-35b-a3b:
389
+ thinking: off # unquoted off works (YAML reads it as false)
390
+ ```
391
+
392
+ - `default` sends nothing: the provider's or the chat template's own default, which is chi's behaviour without the
393
+ setting. For some hybrid models that default is *no* thinking (DeepSeek V3.1 on OpenRouter); `medium` turns it on.
394
+ `on` isn't a level.
395
+ - Order, first set wins: `--thinking LEVEL` or `SAMAGOTCHI_THINKING_LEVEL`, then the `models:` entry (found the way
396
+ `profile:` is), then the `hosts:` entry, then `thinking.level` in the file, then `default`. Anything else than a
397
+ level warns once and counts as unset.
398
+ - The flag reaches the sessions that start with it; a session already running keeps its level. `models:` levels
399
+ and `thinking.level` in the file are read every turn; a host's `thinking:` reaches a worker when it starts, as
400
+ its `sampling:` does (`chi sessions stop` to change it).
401
+ - `/model` shows the level and where it came from (`thinking: off (models: qwen3.6-35b-a3b)`), `chi self` too.
402
+ - The idle recap and plugins' side questions always ask with thinking off, whatever the level.
403
+
404
+ What each backend gets:
405
+
406
+ | Backend | `off` | `low` / `medium` / `high` |
407
+ |---|---|---|
408
+ | native (`/completion`), `qwen36` | an empty thought after the assistant cue, and no turn preamble | no knob: thinking stays as the model has it, one notice |
409
+ | native, `gemma4` | no `<\|think\|>` token at the start of the system prompt | no knob, one notice |
410
+ | chat host (`api: openai`) | `chat_template_kwargs: {enable_thinking: false}` and `reasoning_effort: "none"` | `reasoning_effort: <level>` |
411
+
412
+ On chat hosts: llama.cpp honours both off switches but ignores the effort; Splash takes `reasoning_effort` (off only
413
+ through `none`) and scales with it; OpenRouter translates `reasoning_effort` per model (some can't turn thinking off:
414
+ Qwen3-30B-A3B thinks anyway, gpt-oss refuses).
415
+
416
+ When the model thinks although the level is `off`, chi says so once per session and host
417
+ (`thinking> warning: off wasn't honoured by … (N chars of thinking)`) and logs `thinking_not_honoured` each time.
418
+ When a host answers the thinking fields with an HTTP 400 about reasoning (gpt-oss: "Reasoning is mandatory"), chi
419
+ sends the request again without them, leaves them out for that model from then on, and says so once.
420
+
421
+ The fields go under the `sampling:` map ("Sampling"): a `sampling:` key wins over the level's, and
422
+ `chat_template_kwargs` merges per sub-key. A `null` there drops a field the level would send, at any depth, for a
423
+ host that refuses one of them:
424
+
425
+ ```yaml
426
+ hosts:
427
+ strict:
428
+ url: https://llm.example.com/v1
429
+ api: openai
430
+ thinking: off
431
+ sampling: { reasoning_effort: null } # sends only enable_thinking: false
432
+ ```
433
+
434
+ A different level changes a native model's system prompt (Gemma's token, Qwen's turn preamble), so the next turn
435
+ reads the whole context again once; on a chat host only the end of the prompt changes.
436
+
277
437
  ## Llama HTTP Timeouts
278
438
 
279
439
  Long-running llama.cpp completions can exceed Ruby's default HTTP read timeout.
280
- Configure these environment variables to avoid premature request failures:
440
+ Raise these to avoid premature request failures:
281
441
 
282
- - `SAMAGOTCHI_SERVER_OPEN_TIMEOUT` (default: `10`) connection timeout in seconds.
283
- - `SAMAGOTCHI_SERVER_READ_TIMEOUT` (default: `600`) response read timeout in seconds.
442
+ - `server.open_timeout` (default: `10`, env `SAMAGOTCHI_SERVER_OPEN_TIMEOUT`) connection timeout in seconds.
443
+ - `server.read_timeout` (default: `600`, env `SAMAGOTCHI_SERVER_READ_TIMEOUT`) response read timeout in seconds.
444
+ Either timeout at `0` (or anything not a positive number) is its default, on every host.
445
+
446
+ ```yaml
447
+ server:
448
+ open_timeout: 10
449
+ read_timeout: 900
450
+ ```
284
451
 
285
452
  A streamed answer also has a **first-token limit**: the seconds it may take to show its first text, reasoning or
286
453
  tool call. A remote provider can keep a queued request open for minutes with SSE keep-alive comments
@@ -308,9 +475,10 @@ prompt caches hit and every turn is answered by the same model. Servers that don
308
475
 
309
476
  To explicitly route requests to a named model in llama.cpp, set:
310
477
 
311
- - `SAMAGOTCHI_DEFAULT_MODEL` (required): model name/id sent as the `model` field on `/completion` requests.
478
+ - `default.model` (required; env `SAMAGOTCHI_DEFAULT_MODEL`, `--model` per run): model name/id sent as the `model`
479
+ field on `/completion` requests.
312
480
 
313
- When `SAMAGOTCHI_DEFAULT_MODEL` is unset or blank, Samagotchi fails fast with a clear startup/configuration error.
481
+ When `default.model` is unset or blank, Samagotchi fails fast with a clear startup/configuration error.
314
482
 
315
483
  With several `hosts:`, an unqualified model name goes to the host whose `/models`
316
484
  list has it (after `/models` ran), by exact id first, then by substring. A
@@ -321,23 +489,55 @@ the running server (llama.cpp's `/props`), else the window the host's model list
321
489
  gives (`context_length`, `context_window`, `max_model_len` or llama.cpp's
322
490
  `meta.n_ctx`), else `context.window_tokens`.
323
491
 
492
+ A `:` in a model name is often part of the id (`qwen3:8b`, `mistral:7b`,
493
+ `unsloth/Qwen3-8B-GGUF:Q4_K_M`), so the part before the first `:` picks a host
494
+ only when it is a configured host's name. An unknown prefix is refused with an
495
+ error naming it and the configured hosts (with a "did you mean" for a near
496
+ miss) when either
497
+ - the rest is an `org/model` id (`nosuch:anthropic/claude-sonnet-4`), or
498
+ - the prefix is a hosted provider's name: `openrouter`, `openai`, `anthropic`,
499
+ `google`, `gemini`, `groq`, `xai`, `together` or `fireworks`
500
+ (`openai:gpt-4o` with no `openai` host).
501
+
502
+ The check applies wherever the model comes in: `--model`, `default.model`, an
503
+ alias, `/model`, `chi send --new --model`, the web's new-session model and a
504
+ delegate's model. Any other unknown prefix (`nosuch:x`) is sent to the default
505
+ host as the model id.
506
+
324
507
  ## Llama Network Retry Behavior
325
508
 
326
509
  Transient network failures are retried automatically with exponential backoff.
327
510
 
328
511
  - Default retries: `5` (up to `6` total attempts including the first call).
329
512
  - Default backoff: `0.5s`, `1s`, `2s`, `4s`, `8s`.
330
- - Retry scope: transient network errors (timeouts, refused/reset connections, EOF/socket reachability failures),
513
+ - Retry scope: transient network errors (timeouts, reset connections, EOF/socket reachability failures),
331
514
  HTTP 429 and HTTP 500/502/503/504/529. A `Retry-After` header replaces the backoff delay; one longer than
332
515
  60s is not waited out and the error is reported instead.
516
+ - A refused connection (nothing listening) is not retried: the turn fails at once with `can't reach host <name> at
517
+ <address> (connection refused) — is the server running?`.
333
518
  - A stream that has already produced output is never retried (the retry would repeat it); it fails the turn.
334
519
  - Cancellation (`Ctrl-C`) is never retried.
335
520
 
336
521
  Configuration:
337
522
 
338
- - `SAMAGOTCHI_RETRY_MAX` (default `5`): number of retries after the first failed attempt.
339
- - `SAMAGOTCHI_RETRY_BASE_DELAY` (default `0.5`): backoff base delay in seconds.
340
- - `SAMAGOTCHI_RETRY_MAX_DELAY` (default `8.0`): cap for backoff delay in seconds.
523
+ - `retry.max` (default `5`, env `SAMAGOTCHI_RETRY_MAX`): number of retries after the first failed attempt.
524
+ - `retry.base_delay` (default `0.5`, env `SAMAGOTCHI_RETRY_BASE_DELAY`): backoff base delay in seconds.
525
+ - `retry.max_delay` (default `8.0`, env `SAMAGOTCHI_RETRY_MAX_DELAY`): cap for backoff delay in seconds.
526
+
527
+ ```yaml
528
+ retry:
529
+ max: 5
530
+ base_delay: 0.5
531
+ max_delay: 8.0
532
+ ```
533
+
534
+ `retry.empty_answer` is a different thing: the request worked, but the model's answer had no visible text and no
535
+ tool calls (thinking only, or nothing; a thinking loop cut by the provider's output cap looks like this). chi then
536
+ asks again in the same turn with a hidden note ("your last reply had no visible answer…"), at `temperature: 0.6`
537
+ unless `sampling:` sets one; the REPL prints `↻ empty answer, asking again (1/1)` and the web shows it as a row of
538
+ the step. `retry.empty_answer` (default `1`, at most `3`, env `SAMAGOTCHI_RETRY_EMPTY_ANSWER`, no CLI flag) is how
539
+ many times per turn; `0` ends the turn at the empty answer as before. An answer cut because the context is full
540
+ (90 % or more) is not retried. When the retries run out the turn ends with "(the model returned an empty answer)".
341
541
 
342
542
  Assist-mode UX:
343
543
 
@@ -352,7 +552,7 @@ message (before, a failed llama.cpp `/completion` ended the turn as
352
552
 
353
553
  | Kind | When | Retried |
354
554
  |---|---|---|
355
- | connection | refused, reset, timed out, dropped mid-stream | yes (network retry), not mid-stream |
555
+ | connection | reset, timed out, dropped mid-stream; refused | yes (network retry), not mid-stream; refused: no |
356
556
  | rate limited | HTTP 429 | yes, honouring `Retry-After` |
357
557
  | server | HTTP 5xx, llama.cpp's mid-stream `error:` event | 500/502/503/504/529 only |
358
558
  | auth | HTTP 401/403 | no |
@@ -467,7 +667,7 @@ Whether a model can see images is found out before a turn with images is sent:
467
667
  - a native llama.cpp host: `/props` must report `modalities.vision` (the server
468
668
  runs with `--mmproj`) and a media marker, and the prompt profile must know the
469
669
  chat template's image wrapping (qwen36 does; gemma4 not yet);
470
- - mlx and oMLX hosts: no;
670
+ - mlx and oMLX hosts: no (not verified with recent chi versions);
471
671
  - an OpenAI-API host: a local llama.cpp's `/props`, else the host's model list
472
672
  (OpenRouter's `architecture.input_modalities`); when it doesn't say, the image
473
673
  is sent and a refusal is reported.
@@ -489,6 +689,92 @@ modalities check: without a media marker the prompt can't carry an image.
489
689
  If an AGENT.md file is present in the project root, samagotchi injects its
490
690
  contents into the system prompt under a "Project specific description:" section.
491
691
 
492
- To skip loading AGENT.md, set:
493
-
494
- `SAMAGOTCHI_SKIP_AGENT_MD=true`
692
+ To skip loading AGENT.md, set `skip_agent_md: true` at the top level of
693
+ `config.yml` (env `SAMAGOTCHI_SKIP_AGENT_MD=true`).
694
+
695
+ ## All settings
696
+
697
+ Every setting below takes the three forms described in
698
+ [Environment variables](#environment-variables): a nested key in
699
+ `config.yml`, `SAMAGOTCHI_<DOTTED_NAME>` in the environment and, where the CLI
700
+ column says so, a `--kebab-name` flag. The maps (`hosts:`, `models:`,
701
+ `model_aliases:`, `hooks:`, `guardrails:` rules, `bundles:`, `memories:`) are
702
+ described in their own sections.
703
+
704
+ | Setting | Default | CLI | What it does |
705
+ |---|---|---|---|
706
+ | `default.model` | (required) | `--model` | The model a new session starts with; `host:model` pins a host. |
707
+ | `default.input` | none | | Text pre-filled at the first prompt (a trailing space is kept); `--no-default-input` skips it. See [CLI](cli.md). |
708
+ | `default.n_predict` | server's | yes | Most tokens one generation may produce (native hosts). |
709
+ | `model.profile` | none | `--profile` | Prompt profile for every model (`qwen36`, `gemma4`); env and CLI only. See "Prompt profile". |
710
+ | `server.transport` | `llama_cpp` | yes | `llama_cpp`, `mlx` or `omlx`; see "Model Server Transport". |
711
+ | `server.host` | `localhost` | yes | The model server when there is no `hosts:` map. |
712
+ | `server.port` | `8080` | yes | Its port. |
713
+ | `server.open_timeout` | `10` | yes | Connection timeout, seconds. |
714
+ | `server.read_timeout` | `600` | yes | Read timeout, seconds. |
715
+ | `server.first_token_timeout` | 120 remote, off local | | Seconds to the first token; `0` = off. `hosts.<name>.first_token_timeout` wins. |
716
+ | `recap.enabled` | on | | `false` (or `recap: false`) turns the idle recap off. |
717
+ | `recap.model` | session's | yes | Model that writes the recap. |
718
+ | `recap.host_ref` | session's | yes | A `hosts:` name to ask (`host:` is accepted too). |
719
+ | `recap.base_url` | none | yes | An OpenAI API base to ask instead (`http://h:8081/v1`). |
720
+ | `recap.inactivity` | `180` | yes | Idle seconds before a recap. |
721
+ | `recap.timeout` | `30` | yes | Seconds a recap request may take. |
722
+ | `recap.min_user_turns` | `2` | yes | Prompts a session needs before it gets a recap. |
723
+ | `recap.sentences` | `2-4` | yes | Recap length, `N` or `N-M` (1–10). |
724
+ | `session.shared` | `true` | | Plain `chi` runs its session in a worker and attaches; `--no-shared` per run. |
725
+ | `session.idle_exit_minutes` | `30` | yes | An unused worker exits after this; `0` = never. |
726
+ | `session.keep_empty` | `false` | | Keep sessions nothing happened in. |
727
+ | `session.max_children` | `4` | | Running delegated sessions one session may have. |
728
+ | `session.retention_days` | `14` | yes | Delete sessions not updated for N days; `0` = forever. See [Sessions](sessions.md). |
729
+ | `session.max_count` | `500` | yes | Keep the newest N; `0` = uncapped. |
730
+ | `session.keep_status` | `running` | yes | Comma list of statuses never pruned. |
731
+ | `session.sweep_interval_hours` | `24` | yes | How often the retention sweep runs. |
732
+ | `image.max_side` | `1568` | | See "Images". |
733
+ | `image.max_bytes` | `3750000` | | See "Images". |
734
+ | `image.max_per_request` | `20` | | See "Images". |
735
+ | `guardrails.enabled` | `true` | | See [Guardrails](guardrails.md). |
736
+ | `log.file` | state dir | yes | See "Debug Log File". |
737
+ | `log.disable` | `false` | yes | No file logging. |
738
+ | `log.level` | `info` | yes | `debug`, `info`, `warn`, `error`. |
739
+ | `status.line` | `on` | yes | The REPL status line, `on` or `off`. |
740
+ | `status.width_mode` | `terminal_cap` | yes | `terminal_cap` (terminal width up to `max_width`) or `fixed`. |
741
+ | `status.max_width` | `160` | yes | Cap for `terminal_cap`. |
742
+ | `status.fixed_width` | `120` | yes | Width for `fixed`. |
743
+ | `context.status` | `true` | yes | Context-usage telemetry for the model. See [context telemetry](internals/context-telemetry.md). |
744
+ | `context.window_tokens` | server's, else 256000 | yes | Context window when the server doesn't report one. |
745
+ | `context.chars_per_token` | `4.0` | yes | Estimate ratio when the server reports no usage. |
746
+ | `context.status_thresholds` | `20,40,60,80` | yes | Percentages that trigger a status. |
747
+ | `context.status_cadence` | `0` | yes | Also every N rounds; `0` = thresholds only. |
748
+ | `thinking.ui` | `spinner` | yes | `spinner` or `off`. |
749
+ | `thinking.render_interval` | `0.08` | yes | Seconds between thinking redraws. |
750
+ | `thinking.turn_preamble` | `true` | yes | Ask a `qwen36` model to open its thinking with a short `TURN:` line (the step label). |
751
+ | `thinking.level` | `default` | `--thinking` | `off`, `low`, `medium`, `high` or `default` for every model; the flag and env outrank the `models:`/`hosts:` entries, the file's value doesn't. See "Thinking". |
752
+ | `models.<key>.thinking`, `hosts.<name>.thinking` | none | | A model's or host's level. See "Thinking". |
753
+ | `max_tool_output_chars` | `10000` | yes | Tool output kept in the conversation; a top-level key (see below). |
754
+ | `retry.max` | `5` | yes | See "Llama Network Retry Behavior". |
755
+ | `retry.base_delay` | `0.5` | yes | |
756
+ | `retry.max_delay` | `8.0` | yes | |
757
+ | `retry.empty_answer` | `1` | | Times a turn asks again after an empty answer (at most 3, `0` = off). See "Llama Network Retry Behavior". |
758
+ | `update.gem` | `true` | | `false`: `chi update` never installs a newer gem (`--no-gem` for one run). See [CLI: Updating](cli.md#updating). |
759
+ | `update.bundles` | `true` | | `false`: `chi update` leaves the shipped bundles to `chi bundle upgrade` (`--no-bundles`). |
760
+ | `update.desktop` | `true` | | `false`: `chi update` leaves the desktop helper alone (`--no-desktop`). |
761
+ | `read.truncate_at_bytes` | `65536` | yes | A `read` result larger than this is cut to a preview. |
762
+ | `read.preview_bytes` | `12288` | yes | Size of that preview. |
763
+ | `read.hard_max_bytes` | `2097152` | yes | Largest file `read` opens. |
764
+ | `read.telemetry_threshold_pct` | `80` | yes | A `read` result that alone fills this % of the context window carries a token estimate. |
765
+ | `execute.truncate_at_bytes` | `65536` | yes | The same for `execute` output. |
766
+ | `execute.preview_bytes` | `12288` | yes | |
767
+ | `execute.telemetry_threshold_pct` | `80` | yes | |
768
+ | `web.port` | `4567` | `--port` | `chi web`'s port. See [CLI](cli.md). |
769
+ | `web.host` | `127.0.0.1` | yes | `127.0.0.1`, `::1` or `localhost`. |
770
+ | `web.markdown` | `false` | yes | Render answers as Markdown in `chi web`. |
771
+ | `web.turn_view` | `true` | yes | One block per turn; `false` = the row of bubbles. |
772
+ | `web.annotate_presets` | `Agreed\|Could you please elaborate?` | yes | Quick replies next to Annotate in `chi web`, `\|`-separated (a YAML list works too); `""` in the file or on the CLI leaves only Annotate (an empty env value means the default). See [CLI](cli.md#web-annotate-presets). |
773
+ | `history.file` | state dir | | Prompt history path. |
774
+ | `no_interrupt` | `false` | `--no-interrupt` | Raise the tool-call limit of a turn to 1000; a top-level key. |
775
+ | `no_default_input` | `false` | `--no-default-input` | Don't pre-fill `default.input`; a top-level key. |
776
+ | `skip_agent_md` | `false` | | Don't load AGENT.md; a top-level key (see below). |
777
+
778
+ `max_tool_output_chars`, `skip_agent_md`, `no_interrupt` and
779
+ `no_default_input` have no section: in `config.yml` they are top-level keys as
780
+ written (`max_tool_output_chars: 20000`).
data/docs/desktop.md CHANGED
@@ -1,6 +1,7 @@
1
1
  # Desktop helper (macOS)
2
2
 
3
- `chi desktop` installs **Chi Helper**, a small native app. It sends text you selected in any app to a chi session,
3
+ `chi desktop` installs **Chi Helper**, a small native app. It sends text you selected in any app (or a screenshot,
4
+ an image: see [Images](#images)) to a chi session,
4
5
  either with a question as [your message](sessions.md#sending-a-message) (a turn runs, and the answer shows in the
5
6
  attached terminal or web page), or as a [context note](sessions.md#context-notes) (the model sees it on its next
6
7
  turn, and no turn starts).
@@ -27,8 +28,40 @@ it's still live; if there's only one live session, that one is (a recent one nev
27
28
  session starts its worker; a note to one waits for its next start, and the panel shows chi's line saying so. After a
28
29
  send the panel shows chi's line and closes. On an error it stays open and shows the error.
29
30
 
31
+ The first row, **New session in <folder>** (⌘0), starts a session with the message instead: `chi send --new --dir
32
+ <folder>`, so it shows in `chi web` at once (see [Starting a session](sessions.md#starting-a-session)). The folder is
33
+ that of the most recently updated live session, else of the newest recent one, else your home folder. It is ticked
34
+ alone (ticking it clears the sessions and the reverse) and is preselected when no session is live. ⌘⏎ on it beeps: a
35
+ note needs a session. After the send the panel shows `started <id>…` for 3 s.
36
+
30
37
  A session open in a `chi --no-shared` REPL isn't listed: it takes no notes or messages.
31
38
 
39
+ ## Images
40
+
41
+ The panel also sends images, as attachments of the message (`chi send --image`, see
42
+ [Sending a message](sessions.md#sending-a-message)):
43
+
44
+ - **A screenshot:** ⌃⇧⌘4 (a region to the clipboard), then ⌃⌥⌘N. The panel shows it as a thumbnail, named
45
+ `clipboard.png`; the context box stays empty.
46
+ - **Finder:** right-click image files → Services → **Send to chi** (the same Service as for text, so its shortcut
47
+ works here too), or ⌘C on them and ⌃⌥⌘N. Up to 20 at once.
48
+ - **Other apps:** a picture selected in Preview, Safari's Copy Image, anything that puts image data on the clipboard.
49
+ - **Drop** image files or the floating screenshot thumbnail (⇧⌘4 to a file) anywhere on the open panel: they are
50
+ added to the ones there.
51
+
52
+ When the clipboard holds both text and a picture (cells copied in Numbers or Excel), the text wins, as before. The
53
+ thumbnails sit under the message line, 48 px high, the file name as a tooltip; hover one for its ✕.
54
+
55
+ - A message is required with images: the line says "Say something about the image…", and ⏎ does nothing until
56
+ there is text (the context box counts).
57
+ - Notes are text only: ⌘⏎ with images beeps and says "Notes are text only: ⏎ sends the image as a message".
58
+ - The session's model must see images. A text-only one fails the turn in the session (the attached terminal or the
59
+ web shows why); the panel has already said it was sent.
60
+ - A session busy with a turn runs the image message as its next turn: the line says `(runs after the current turn)`.
61
+ - Clipboard and dropped image data goes to temp files under `$TMPDIR/chi-helper`, deleted after the send or when
62
+ the panel closes; files from Finder are sent as they are, never touched. A send with images may take up to 30 s
63
+ (converting, a worker starting) before the panel gives up.
64
+
32
65
  ## Install
33
66
 
34
67
  ```sh
@@ -47,16 +80,22 @@ don't break it. From a checkout it runs **that** checkout's `bin/chi`; installin
47
80
  warning, because the helper stops working once that worktree is removed. After switching between a checkout and a
48
81
  gem install, run `chi desktop upgrade` from the one you now use.
49
82
 
83
+ `chi update` keeps it current: it rebuilds and restarts the helper only when its Swift sources changed since the
84
+ build (`launch.json` records their digest) or the Ruby it runs moved. A new chi that left the sources alone only
85
+ rewrites `launch.json`, which the app reads at each send, so the app keeps its older version number and that's fine.
86
+
50
87
  ## Commands
51
88
 
52
89
  | Command | Does |
53
90
  |---|---|
54
91
  | `chi desktop install [--force] [--login]` | builds, installs and starts it; `--force` replaces an existing copy |
55
- | `chi desktop upgrade` | rebuilds it for this chi and restarts it, keeping the login setting |
92
+ | `chi desktop upgrade` | rebuilds it for this chi and restarts it, keeping the login setting (always; `chi update` does it only when needed) |
56
93
  | `chi desktop uninstall` | quits it and removes the app, its login item, launch file and settings |
57
94
  | `chi desktop status` | version against chi's, how it runs chi, state dirs, Service, hotkey, process, login item |
58
95
 
59
- `chi self` has a `desktop` line: `0.1.x (matches)`, `0.1.w (chi is 0.1.x: chi desktop upgrade)` or `not installed`.
96
+ `chi self` has a `desktop` line: `0.1.x (matches)`, `0.1.w (up to date for chi 0.1.x)` (an older build whose sources
97
+ haven't changed), `0.1.w (chi is 0.1.x: chi update)` (a rebuild is due) or `not installed`. `chi desktop status`
98
+ says the same.
60
99
 
61
100
  ## How it runs chi
62
101
 
@@ -86,7 +125,7 @@ Each call is stopped after 10 s. A stopped `chi note` says the note may be partl
86
125
  ## Troubleshooting
87
126
 
88
127
  - **"chi not found at …, run `chi desktop upgrade`"**: the Ruby or checkout in `launch.json` moved (a Ruby upgrade,
89
- a removed worktree). Run `chi desktop upgrade` from the chi you use now.
128
+ a removed worktree). Run `chi desktop upgrade` (or `chi update`) from the chi you use now.
90
129
  - **No "Send to chi" in the Services menu:** check `chi desktop status` (service). Try
91
130
  `/System/Library/CoreServices/pbs -update`, start the app again, or log out and back in. It must be ticked in
92
131
  System Settings → Keyboard → Keyboard Shortcuts… → Services → Text.
@@ -94,4 +133,6 @@ Each call is stopped after 10 s. A stopped `chi note` says the note may be partl
94
133
  restarts or you log out: macOS caches Services.
95
134
  - **⌃⌥⌘N does nothing:** `status` says whether another app holds it. macOS doesn't report clashes with its own
96
135
  shortcuts.
136
+ - **"Send to chi" missing on images in Finder or Preview** after an upgrade: the Services cache still has the old
137
+ (text-only) entry; `/System/Library/CoreServices/pbs -update`, restart the helper, or log out and back in.
97
138
  - **"No live sessions":** start one with `chi` in a terminal; `chi sessions list --live --scope=all` shows the same list.
data/docs/guardrails.md CHANGED
@@ -38,6 +38,17 @@ once it closes. On a short terminal the list shrinks (the hint row, then the `in
38
38
  lines, then the header go, then the options fold onto fewer rows). Once answered, one
39
39
  line stays in the scrollback: `! execute: git push origin main → Allow once`.
40
40
 
41
+ An `edit` or `write` also shows the change it would make, computed without
42
+ touching the file: a `change: +3 −1` line in the question (`new file, 12 lines`;
43
+ `would fail: old text not found in …` when the edit can't apply), and the
44
+ unified diff itself. The web card shows the diff under the path (20 lines,
45
+ then "show all"); the terminals print it above the question, green and red,
46
+ 40 lines at most (the rest is on the web). Binary files and files over 1 MB
47
+ say so instead of a diff. After the call runs, its row shows what really
48
+ changed: `diff +3 −1` under the row on the web (closed, it survives a
49
+ reload) and ` +3 −1` at the end of the terminal's tool line. The model never
50
+ sees these diffs.
51
+
41
52
  Who answers:
42
53
 
43
54
  - REPL (`chi --no-shared`, `-p` without `--non-interactive`): at the `? ` prompt.