samagotchi 0.2.0 → 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (96) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +116 -1
  3. data/README.md +40 -4
  4. data/bin/chi +86 -19
  5. data/docs/cli.md +108 -5
  6. data/docs/configuration.md +229 -45
  7. data/docs/desktop.md +6 -0
  8. data/docs/hooks.md +126 -5
  9. data/docs/plugins.md +68 -2
  10. data/docs/releasing.md +18 -8
  11. data/docs/sessions.md +30 -4
  12. data/lib/samagotchi/answer_display.rb +95 -0
  13. data/lib/samagotchi/archive_store.rb +90 -0
  14. data/lib/samagotchi/bootstrap/config_writer.rb +342 -0
  15. data/lib/samagotchi/bootstrap/probe.rb +262 -0
  16. data/lib/samagotchi/bootstrap_command.rb +347 -0
  17. data/lib/samagotchi/bridge/pending_card.rb +89 -0
  18. data/lib/samagotchi/bridge/turn_accumulator.rb +14 -3
  19. data/lib/samagotchi/bridge.rb +9 -0
  20. data/lib/samagotchi/bridge_client.rb +6 -2
  21. data/lib/samagotchi/bundles/check-in/manifest.yml +10 -0
  22. data/lib/samagotchi/bundles/check-in/plugin.rb +244 -0
  23. data/lib/samagotchi/bundles/source-links/hooks/source_links.rb +358 -0
  24. data/lib/samagotchi/bundles/source-links/manifest.yml +14 -0
  25. data/lib/samagotchi/bundles/source-links/source_links.md +5 -0
  26. data/lib/samagotchi/bundles/system/config_modification_protocol.md +10 -6
  27. data/lib/samagotchi/bundles/system/manifest.yml +3 -3
  28. data/lib/samagotchi/bundles/system/self_map.md +8 -2
  29. data/lib/samagotchi/client.rb +72 -13
  30. data/lib/samagotchi/config.rb +196 -36
  31. data/lib/samagotchi/desktop/macos/ChiRunner.swift +13 -7
  32. data/lib/samagotchi/desktop/macos/Panel.swift +71 -19
  33. data/lib/samagotchi/empty_answer_retry.rb +43 -0
  34. data/lib/samagotchi/engine.rb +233 -36
  35. data/lib/samagotchi/guardrails/approval.rb +9 -0
  36. data/lib/samagotchi/guardrails/scratch_writes.rb +40 -0
  37. data/lib/samagotchi/guardrails.rb +1 -0
  38. data/lib/samagotchi/hooks/registry.rb +24 -5
  39. data/lib/samagotchi/host_registry.rb +4 -3
  40. data/lib/samagotchi/idle_recap.rb +5 -1
  41. data/lib/samagotchi/kernel_loop.rb +47 -15
  42. data/lib/samagotchi/llm/chat_loop.rb +59 -20
  43. data/lib/samagotchi/llm/errors.rb +21 -3
  44. data/lib/samagotchi/llm/http.rb +42 -13
  45. data/lib/samagotchi/llm/openai_chat.rb +12 -4
  46. data/lib/samagotchi/log_subscriber.rb +18 -3
  47. data/lib/samagotchi/model_profile.rb +1 -1
  48. data/lib/samagotchi/plugin/context.rb +22 -1
  49. data/lib/samagotchi/plugin/sessions.rb +3 -1
  50. data/lib/samagotchi/reply_wait.rb +126 -0
  51. data/lib/samagotchi/sampling_settings.rb +58 -0
  52. data/lib/samagotchi/self_report.rb +1 -0
  53. data/lib/samagotchi/send_command.rb +153 -7
  54. data/lib/samagotchi/session.rb +52 -11
  55. data/lib/samagotchi/session_archive_command.rb +107 -0
  56. data/lib/samagotchi/session_commands.rb +11 -2
  57. data/lib/samagotchi/session_manager.rb +114 -9
  58. data/lib/samagotchi/session_metrics.rb +222 -106
  59. data/lib/samagotchi/steer.rb +72 -0
  60. data/lib/samagotchi/terminal_ui/attached_loop.rb +39 -6
  61. data/lib/samagotchi/terminal_ui/event_renderer.rb +13 -8
  62. data/lib/samagotchi/terminal_ui/formatting.rb +31 -8
  63. data/lib/samagotchi/terminal_ui/input_support.rb +3 -0
  64. data/lib/samagotchi/terminal_ui.rb +77 -4
  65. data/lib/samagotchi/tool_activity.rb +3 -1
  66. data/lib/samagotchi/tools/builtins.rb +15 -4
  67. data/lib/samagotchi/tools/delegate_wait.rb +26 -69
  68. data/lib/samagotchi/tools/execute.rb +52 -14
  69. data/lib/samagotchi/tools/task_runtime.rb +19 -0
  70. data/lib/samagotchi/tools/task_wait.rb +27 -3
  71. data/lib/samagotchi/turn_note.rb +60 -6
  72. data/lib/samagotchi/version.rb +1 -1
  73. data/lib/samagotchi/vision_support.rb +2 -6
  74. data/lib/samagotchi/web/app.rb +88 -4
  75. data/lib/samagotchi/web/public/activity.js +10 -1
  76. data/lib/samagotchi/web/public/annotate_presets.js +26 -0
  77. data/lib/samagotchi/web/public/annotations.js +13 -0
  78. data/lib/samagotchi/web/public/app.js +437 -88
  79. data/lib/samagotchi/web/public/card.js +5 -3
  80. data/lib/samagotchi/web/public/chat_view.js +10 -1
  81. data/lib/samagotchi/web/public/copy.js +20 -4
  82. data/lib/samagotchi/web/public/ctx.js +15 -0
  83. data/lib/samagotchi/web/public/data.js +21 -6
  84. data/lib/samagotchi/web/public/format.js +9 -0
  85. data/lib/samagotchi/web/public/index.html +38 -2
  86. data/lib/samagotchi/web/public/notify.js +175 -0
  87. data/lib/samagotchi/web/public/question_card.js +2 -1
  88. data/lib/samagotchi/web/public/sessions_list.js +7 -0
  89. data/lib/samagotchi/web/public/timing.js +39 -14
  90. data/lib/samagotchi/web/public/turn_events.js +46 -0
  91. data/lib/samagotchi/web/public/turn_view.js +47 -7
  92. data/lib/samagotchi/web/server.rb +8 -4
  93. data/lib/samagotchi/web/session_hub.rb +2 -1
  94. data/lib/samagotchi/web/session_summary.rb +24 -1
  95. data/lib/samagotchi/worker.rb +11 -0
  96. metadata +20 -1
@@ -2,26 +2,37 @@
2
2
 
3
3
  ## Global Config File
4
4
 
5
- Chi can preload a global config file and expose those entries as environment
6
- variables before the app boots.
5
+ Chi reads its settings from three places; the first that sets a value wins:
6
+
7
+ 1. a CLI flag (`--server-port 8081`),
8
+ 2. an environment variable (`SAMAGOTCHI_SERVER_PORT=8081`),
9
+ 3. the global config file (`server: {port: 8081}`),
10
+
11
+ then the built-in default. [All settings](#all-settings) lists them.
12
+
13
+ `chi bootstrap HOST[:PORT]` writes a first config for a model server, or adds
14
+ it to this file as a `hosts:` entry (see [First setup](cli.md#first-setup)).
7
15
 
8
16
  Default path:
9
17
 
10
18
  - `$XDG_CONFIG_HOME/samagotchi/config.yml`
11
19
  - Fallback when `XDG_CONFIG_HOME` is unset: `~/.config/samagotchi/config.yml`
12
20
 
13
- Example:
21
+ Example (every section is optional except `default.model`; README.md has a
22
+ minimal one):
14
23
 
15
24
  ```yaml
16
- SAMAGOTCHI_DEFAULT_MODEL: Qwen3-14B-Instruct
17
- server:
25
+ default:
26
+ model: Qwen3-14B-Instruct
27
+ server: # the model server when there is no hosts: map below
18
28
  host: 192.0.2.10
19
29
  port: 8081
20
- SAMAGOTCHI_THINKING_UI: spinner
30
+ thinking:
31
+ ui: spinner
21
32
 
22
33
  # Multi-host (optional): aggregated /models and per-model routing.
23
- # Bare SAMAGOTCHI_DEFAULT_MODEL uses the default host; host:model pins to a host.
24
- # Transport per host overrides SAMAGOTCHI_SERVER_TRANSPORT; api: openai makes a
34
+ # A bare default.model uses the default host; host:model pins to a host.
35
+ # transport: on a host overrides server.transport; api: openai makes a
25
36
  # host use the OpenAI chat API instead of chi's raw prompt.
26
37
  hosts:
27
38
  main:
@@ -58,36 +69,61 @@ session:
58
69
  # keep_empty: false # true keeps sessions nothing happened in (default: deleted when left)
59
70
  # max_children: 4 # running sessions one session may have delegated at a time (the delegate tool)
60
71
 
61
- # Baseline memories preloaded into the system prompt (same shape as --memory).
72
+ # Baseline memories preloaded into the system prompt (same shape as --memory);
73
+ # name entries you have, or chi warns at each start.
62
74
  # CLI --memory entries are appended after these, deduped.
63
75
  memories:
64
- - system/user_preferences
65
- - project/feature-env-template
76
+ # - system/user_preferences
77
+ # - project/feature-env-template
66
78
  ```
67
79
 
68
80
  Behavior:
69
81
 
70
82
  - The file is optional.
71
- - Top level is a YAML mapping of scalar env overrides plus nested sections.
72
- - Real environment variables still win over config-file values.
83
+ - Top level is a YAML mapping of sections (`default:`, `server:`, `recap:`, …)
84
+ and the maps described below. Keys are lower `snake_case`, one level per
85
+ dot: `server.read_timeout` is `server: {read_timeout: 600}`.
86
+ - A key chi doesn't read warns at start, with the closest known key:
87
+ `config: unknown key 'default.modle' (did you mean 'default.model'?)`.
88
+ Names you choose under the maps below (host names, model ids) don't warn.
89
+ - Environment variables and CLI flags win over config-file values.
73
90
  - Workers inherit hosts via `SAMAGOTCHI_HOSTS_JSON` propagated through `SessionManager.spawn_options`.
74
91
 
75
92
  This lets you run `chi` without repeating common defaults such as model
76
93
  and llama host/port on every invocation.
77
94
 
78
- Note: The global config file supports both flat scalar entries (for env vars)
79
- and nested sections like `hosts:`, `recap:`, `hooks:`, `guardrails:` (see
80
- [Guardrails](guardrails.md)), `bundles:` (a bundle's settings for its hooks,
81
- see [Hooks: Settings](hooks.md#settings)), `model_aliases:`,
82
- `memories:`. Scalar entries are loaded as environment
83
- variables; non-scalar sections are skipped by the env-loader and parsed by
84
- their respective subsystems (e.g. the hooks system, `HostRegistry`). The
85
- `memories:` list is the persistent baseline for preloaded memory entries —
95
+ Besides the settings sections, the file holds maps that are read by their
96
+ own subsystems: `hosts:` (below), `models:` (see "Prompt profile" and
97
+ "Images"), `model_aliases:`, `hooks:` (see [Hooks](hooks.md)),
98
+ `guardrails:` (see [Guardrails](guardrails.md)), `bundles:` (a bundle's
99
+ settings for its hooks, see [Hooks: Settings](hooks.md#settings)) and
100
+ `memories:`. Their entry names (host names, model ids, aliases) are yours to
101
+ choose. The `memories:` list is the persistent baseline for preloaded memory entries —
86
102
  the same name shape as `--memory` (bare name or `scope/name`), merged under
87
103
  any per-run `--memory` values (config baseline first, deduped). A per-run
88
104
  `--mute NAME` removes an entry from the merged list for that session (see
89
105
  "Muting a memory" in cli.md).
90
106
 
107
+ ### Environment variables
108
+
109
+ Every setting has an environment variable: `SAMAGOTCHI_` plus its dotted
110
+ name in upper case, with `_` for each dot. `default.model` is
111
+ `SAMAGOTCHI_DEFAULT_MODEL`, `server.read_timeout` is
112
+ `SAMAGOTCHI_SERVER_READ_TIMEOUT`. Use them to override the file for one run or
113
+ one shell (`SAMAGOTCHI_LOG_LEVEL=debug chi`); keep lasting choices in the file.
114
+ Most settings also have a CLI flag: the dotted name in kebab case
115
+ (`--server-read-timeout 900`); `chi --help` lists them.
116
+
117
+ ### Legacy flat keys
118
+
119
+ Older configs used the environment names as top-level keys
120
+ (`SAMAGOTCHI_DEFAULT_MODEL: my-model`). They are still read, but every run
121
+ warns (`config key 'SAMAGOTCHI_DEFAULT_MODEL' is legacy UPPER — use
122
+ 'default.model'`). When a file has both, the nested key wins and the warning
123
+ names both (`both 'SAMAGOTCHI_DEFAULT_MODEL' and 'default.model' are set;
124
+ using 'default.model', remove the flat key`). Move each to its nested form (`default: {model: my-model}`) and delete
125
+ the flat line. `/model --default` already writes the nested form.
126
+
91
127
  ## Model Server Transport
92
128
 
93
129
  Chi talks to a model server over HTTP and supports three transports:
@@ -99,11 +135,13 @@ Chi talks to a model server over HTTP and supports three transports:
99
135
  continuous batching + tiered SSD KV cache) using the same `/v1/completions`
100
136
  and `/v1/models` endpoints as `mlx`.
101
137
 
102
- Select the transport with `SAMAGOTCHI_SERVER_TRANSPORT` (`llama_cpp`, `mlx`, or
103
- `omlx`). `SAMAGOTCHI_SERVER_HOST`/`SAMAGOTCHI_SERVER_PORT` are reused for all three —
138
+ > **mlx and oMLX** are not verified with recent chi versions; llama.cpp and
139
+ > OpenAI-compatible hosts (`api: openai`) are the tested paths.
140
+
141
+ Select the transport with `server.transport` (`llama_cpp`, `mlx`, or `omlx`;
142
+ env `SAMAGOTCHI_SERVER_TRANSPORT`). `server.host`/`server.port` are reused for all three —
104
143
  only the request/response shape differs. oMLX's default server port is `8000` (not
105
- `8080`), so point `SAMAGOTCHI_SERVER_PORT` at it, e.g. `SAMAGOTCHI_SERVER_PORT=8000`.
106
- With `hosts:` each entry may set `transport: llama_cpp|mlx|omlx` to override the
144
+ `8080`), so point `server.port` at it. With `hosts:` each entry may set `transport: llama_cpp|mlx|omlx` to override the
107
145
  global transport per host (`lib/samagotchi/host_registry.rb`).
108
146
 
109
147
  Each host may also set `api:`, which says how chi talks to it:
@@ -116,11 +154,11 @@ Without `api:` a host uses the raw-prompt loop, as before. The loop follows the
116
154
  model's host, so `/model other-host:model` can move a session between the two.
117
155
  Workers started by plain `chi`, `chi web` or `--attach` get the same hosts, `api:` included.
118
156
 
119
- Example for mlx-lm:
157
+ Example for mlx-lm (not verified with recent chi versions, like oMLX below):
120
158
 
121
159
  ```yaml
122
- SAMAGOTCHI_SERVER_TRANSPORT: mlx
123
160
  server:
161
+ transport: mlx
124
162
  host: 127.0.0.1
125
163
  port: 8080
126
164
  ```
@@ -132,8 +170,8 @@ mlx_lm.server --model mlx-community/Qwen3-14B-Instruct-4bit
132
170
  Example for oMLX:
133
171
 
134
172
  ```yaml
135
- SAMAGOTCHI_SERVER_TRANSPORT: omlx
136
173
  server:
174
+ transport: omlx
137
175
  host: 192.0.2.10
138
176
  port: 8000
139
177
  ```
@@ -205,9 +243,9 @@ oMLX's known tool-call limitation (a stream filter that strips markup) only
205
243
  affects its `/v1/chat/completions` endpoint, not the `/v1/completions` endpoint
206
244
  chi uses, so raw `[[…]]`/`<|tool_call>` markers stream through untouched.
207
245
 
208
- `SAMAGOTCHI_DEFAULT_MODEL` (config default) and `/model` (runtime effective) pick the model; the status line and `/model`
246
+ `default.model` (config default) and `/model` (runtime effective) pick the model; the status line and `/model`
209
247
  output always render the runtime effective model (showing default when diverged). Which prompt format it gets is the
210
- prompt profile (see "Prompt profile" below). How the selector reaches the request differs by transport:
248
+ prompt profile (see "Prompt profile" below). How the selector reaches the request differs by transport (mlx and oMLX: not verified with recent chi versions):
211
249
 
212
250
  - **mlx** (`mlx_lm.server`): the `model` field is omitted entirely — the server
213
251
  uses whatever was loaded via its own `--model` CLI flag.
@@ -217,7 +255,7 @@ prompt profile (see "Prompt profile" below). How the selector reaches the reques
217
255
  (case-insensitive) first, then substring, then passed through unchanged. That
218
256
  resolved id is usually prefixed (e.g. `mlx-community--gemma-3-4b-it-4bit`), so a
219
257
  short selector such as `gemma-3-4b-it-4bit` is what you set in
220
- `SAMAGOTCHI_DEFAULT_MODEL`. An unknown selector passes through raw and oMLX 404s,
258
+ `default.model`. An unknown selector passes through raw and oMLX 404s,
221
259
  listing its available models; if `/v1/models` is unreachable, samagotchi falls
222
260
  back to the raw selector and lets the server decide (its own 400/404). Either
223
261
  error fails the turn with the server's message (see "Server errors" below). Runtime
@@ -231,8 +269,8 @@ syntax, thought tags and stop sequences. That is the prompt profile, `qwen36` or
231
269
  worse output: a ChatML model under `gemma4` never hits a stop sequence, generates until its limit and then runs the
232
270
  tool calls it made up on the way. The first of these that says something wins:
233
271
 
234
- 1. `--profile NAME` (or `--model-profile NAME`), then `SAMAGOTCHI_MODEL_PROFILE`: for every model in the process,
235
- including one picked later with `/model`.
272
+ 1. `--profile NAME` (or `--model-profile NAME`), then `SAMAGOTCHI_MODEL_PROFILE` (`model.profile` has no config-file
273
+ form): for every model in the process, including one picked later with `/model`.
236
274
  2. `models:` in `config.yml`, keyed by model id or alias (case-insensitive; the name as typed, alias-resolved or
237
275
  without its host prefix):
238
276
 
@@ -274,13 +312,62 @@ Workers get `hosts:` (with `profile:`) through `SAMAGOTCHI_HOSTS_JSON` and read
274
312
  `--profile` reaches the worker a chi starts, but a worker that another process wakes later (`chi web`, `--attach`
275
313
  after an idle exit) gets that process's environment, so put a lasting choice in config.
276
314
 
315
+ ## Sampling
316
+
317
+ chi asks chat hosts (`api: openai`) for greedy decoding (`temperature: 0.0`) and sends native hosts no sampling
318
+ fields, so their server's defaults apply (llama.cpp: temperature 0.8). `sampling:` on a `hosts:` entry or a `models:`
319
+ entry sets request fields for the model's turns:
320
+
321
+ ```yaml
322
+ hosts:
323
+ work:
324
+ url: https://llm.example.com/v1
325
+ api: openai
326
+ api_key_env: WORK_API_KEY
327
+ sampling: { temperature: 0.6, presence_penalty: 1.5 }
328
+ models:
329
+ qwen3.6-35b-a3b:
330
+ sampling: { temperature: 0.6, top_p: 0.95, repeat_penalty: 1.1 }
331
+ ```
332
+
333
+ - The fields go into the request as written; chi doesn't check the names, since providers differ (`repeat_penalty`,
334
+ `min_p`, `dry_multiplier` are llama.cpp's; `presence_penalty` is OpenAI-style). A provider that refuses one fails
335
+ the turn with its error (`host work rejected the request: HTTP 400: …`; one server answered a `presence_penalty`
336
+ with "the requested logits or output transformation is not supported"): remove the last field you added from that
337
+ host's or model's `sampling:`, or set it to `null` in the model's entry. A nested map passes through too, e.g.
338
+ `chat_template_kwargs: { enable_thinking: false }`.
339
+ - A `models:` entry's fields win over its host's, field by field (the host can set a penalty and the model move only
340
+ the temperature). The entry is found the way `profile:` is (the name as typed, alias-resolved or without its host
341
+ prefix).
342
+ - `temperature: null` (or `~`) sends no temperature, so the provider's default applies (vLLM takes it from the
343
+ model's `generation_config.json`).
344
+ - Fields chi sets itself are refused with a warning: `model`, `messages`, `prompt`, `stream`, `stream_options`,
345
+ `tools`, `tool_choice`, `stop`, `n_predict`, `max_tokens`, `n`, `parallel_tool_calls`, `response_format`,
346
+ `cache_prompt`. A `sampling:` that isn't a map warns and is skipped.
347
+ - Greedy decoding can make a heavily quantized thinking model loop in its reasoning ("Let me write the reply…"
348
+ for minutes) or end with an empty answer. Qwen's own advice for its thinking models is `temperature: 0.6,
349
+ top_p: 0.95` (not greedy), with `presence_penalty` between 0 and 2 against endless repetition.
350
+ - The idle recap and side questions keep their own short, deterministic settings.
351
+
352
+ `/model` shows what applies (`sampling: temperature=0.6 presence_penalty=1.5 (hosts.work)`), and each request's
353
+ `stream` line in the debug log carries a `sampling=` field with what was sent. The fields are read every turn, so a
354
+ `/model` switch takes the new model's. A worker gets `hosts:` (with `sampling:`) when it starts, through
355
+ `SAMAGOTCHI_HOSTS_JSON`, and reads `models:` from the config file each turn: after changing a host's `sampling:`,
356
+ stop the session's worker (`chi sessions stop`) for it to take effect.
357
+
277
358
  ## Llama HTTP Timeouts
278
359
 
279
360
  Long-running llama.cpp completions can exceed Ruby's default HTTP read timeout.
280
- Configure these environment variables to avoid premature request failures:
361
+ Raise these to avoid premature request failures:
362
+
363
+ - `server.open_timeout` (default: `10`, env `SAMAGOTCHI_SERVER_OPEN_TIMEOUT`) connection timeout in seconds.
364
+ - `server.read_timeout` (default: `600`, env `SAMAGOTCHI_SERVER_READ_TIMEOUT`) response read timeout in seconds.
281
365
 
282
- - `SAMAGOTCHI_SERVER_OPEN_TIMEOUT` (default: `10`) connection timeout in seconds.
283
- - `SAMAGOTCHI_SERVER_READ_TIMEOUT` (default: `600`) response read timeout in seconds.
366
+ ```yaml
367
+ server:
368
+ open_timeout: 10
369
+ read_timeout: 900
370
+ ```
284
371
 
285
372
  A streamed answer also has a **first-token limit**: the seconds it may take to show its first text, reasoning or
286
373
  tool call. A remote provider can keep a queued request open for minutes with SSE keep-alive comments
@@ -308,9 +395,10 @@ prompt caches hit and every turn is answered by the same model. Servers that don
308
395
 
309
396
  To explicitly route requests to a named model in llama.cpp, set:
310
397
 
311
- - `SAMAGOTCHI_DEFAULT_MODEL` (required): model name/id sent as the `model` field on `/completion` requests.
398
+ - `default.model` (required; env `SAMAGOTCHI_DEFAULT_MODEL`, `--model` per run): model name/id sent as the `model`
399
+ field on `/completion` requests.
312
400
 
313
- When `SAMAGOTCHI_DEFAULT_MODEL` is unset or blank, Samagotchi fails fast with a clear startup/configuration error.
401
+ When `default.model` is unset or blank, Samagotchi fails fast with a clear startup/configuration error.
314
402
 
315
403
  With several `hosts:`, an unqualified model name goes to the host whose `/models`
316
404
  list has it (after `/models` ran), by exact id first, then by substring. A
@@ -335,9 +423,24 @@ Transient network failures are retried automatically with exponential backoff.
335
423
 
336
424
  Configuration:
337
425
 
338
- - `SAMAGOTCHI_RETRY_MAX` (default `5`): number of retries after the first failed attempt.
339
- - `SAMAGOTCHI_RETRY_BASE_DELAY` (default `0.5`): backoff base delay in seconds.
340
- - `SAMAGOTCHI_RETRY_MAX_DELAY` (default `8.0`): cap for backoff delay in seconds.
426
+ - `retry.max` (default `5`, env `SAMAGOTCHI_RETRY_MAX`): number of retries after the first failed attempt.
427
+ - `retry.base_delay` (default `0.5`, env `SAMAGOTCHI_RETRY_BASE_DELAY`): backoff base delay in seconds.
428
+ - `retry.max_delay` (default `8.0`, env `SAMAGOTCHI_RETRY_MAX_DELAY`): cap for backoff delay in seconds.
429
+
430
+ ```yaml
431
+ retry:
432
+ max: 5
433
+ base_delay: 0.5
434
+ max_delay: 8.0
435
+ ```
436
+
437
+ `retry.empty_answer` is a different thing: the request worked, but the model's answer had no visible text and no
438
+ tool calls (thinking only, or nothing; a thinking loop cut by the provider's output cap looks like this). chi then
439
+ asks again in the same turn with a hidden note ("your last reply had no visible answer…"), at `temperature: 0.6`
440
+ unless `sampling:` sets one; the REPL prints `↻ empty answer, asking again (1/1)` and the web shows it as a row of
441
+ the step. `retry.empty_answer` (default `1`, at most `3`, env `SAMAGOTCHI_RETRY_EMPTY_ANSWER`, no CLI flag) is how
442
+ many times per turn; `0` ends the turn at the empty answer as before. An answer cut because the context is full
443
+ (90 % or more) is not retried. When the retries run out the turn ends with "(the model returned an empty answer)".
341
444
 
342
445
  Assist-mode UX:
343
446
 
@@ -467,7 +570,7 @@ Whether a model can see images is found out before a turn with images is sent:
467
570
  - a native llama.cpp host: `/props` must report `modalities.vision` (the server
468
571
  runs with `--mmproj`) and a media marker, and the prompt profile must know the
469
572
  chat template's image wrapping (qwen36 does; gemma4 not yet);
470
- - mlx and oMLX hosts: no;
573
+ - mlx and oMLX hosts: no (not verified with recent chi versions);
471
574
  - an OpenAI-API host: a local llama.cpp's `/props`, else the host's model list
472
575
  (OpenRouter's `architecture.input_modalities`); when it doesn't say, the image
473
576
  is sent and a refusal is reported.
@@ -489,6 +592,87 @@ modalities check: without a media marker the prompt can't carry an image.
489
592
  If an AGENT.md file is present in the project root, samagotchi injects its
490
593
  contents into the system prompt under a "Project specific description:" section.
491
594
 
492
- To skip loading AGENT.md, set:
493
-
494
- `SAMAGOTCHI_SKIP_AGENT_MD=true`
595
+ To skip loading AGENT.md, set `skip_agent_md: true` at the top level of
596
+ `config.yml` (env `SAMAGOTCHI_SKIP_AGENT_MD=true`).
597
+
598
+ ## All settings
599
+
600
+ Every setting below takes the three forms described in
601
+ [Environment variables](#environment-variables): a nested key in
602
+ `config.yml`, `SAMAGOTCHI_<DOTTED_NAME>` in the environment and, where the CLI
603
+ column says so, a `--kebab-name` flag. The maps (`hosts:`, `models:`,
604
+ `model_aliases:`, `hooks:`, `guardrails:` rules, `bundles:`, `memories:`) are
605
+ described in their own sections.
606
+
607
+ | Setting | Default | CLI | What it does |
608
+ |---|---|---|---|
609
+ | `default.model` | (required) | `--model` | The model a new session starts with; `host:model` pins a host. |
610
+ | `default.input` | none | | Text pre-filled at the first prompt (a trailing space is kept); `--no-default-input` skips it. See [CLI](cli.md). |
611
+ | `default.n_predict` | server's | yes | Most tokens one generation may produce (native hosts). |
612
+ | `model.profile` | none | `--profile` | Prompt profile for every model (`qwen36`, `gemma4`); env and CLI only. See "Prompt profile". |
613
+ | `server.transport` | `llama_cpp` | yes | `llama_cpp`, `mlx` or `omlx`; see "Model Server Transport". |
614
+ | `server.host` | `localhost` | yes | The model server when there is no `hosts:` map. |
615
+ | `server.port` | `8080` | yes | Its port. |
616
+ | `server.open_timeout` | `10` | yes | Connection timeout, seconds. |
617
+ | `server.read_timeout` | `600` | yes | Read timeout, seconds. |
618
+ | `server.first_token_timeout` | 120 remote, off local | | Seconds to the first token; `0` = off. `hosts.<name>.first_token_timeout` wins. |
619
+ | `recap.enabled` | on | | `false` (or `recap: false`) turns the idle recap off. |
620
+ | `recap.model` | session's | yes | Model that writes the recap. |
621
+ | `recap.host_ref` | session's | yes | A `hosts:` name to ask (`host:` is accepted too). |
622
+ | `recap.base_url` | none | yes | An OpenAI API base to ask instead (`http://h:8081/v1`). |
623
+ | `recap.inactivity` | `180` | yes | Idle seconds before a recap. |
624
+ | `recap.timeout` | `30` | yes | Seconds a recap request may take. |
625
+ | `recap.min_user_turns` | `2` | yes | Prompts a session needs before it gets a recap. |
626
+ | `recap.sentences` | `2-4` | yes | Recap length, `N` or `N-M` (1–10). |
627
+ | `session.shared` | `true` | | Plain `chi` runs its session in a worker and attaches; `--no-shared` per run. |
628
+ | `session.idle_exit_minutes` | `30` | yes | An unused worker exits after this; `0` = never. |
629
+ | `session.keep_empty` | `false` | | Keep sessions nothing happened in. |
630
+ | `session.max_children` | `4` | | Running delegated sessions one session may have. |
631
+ | `session.retention_days` | `14` | yes | Delete sessions not updated for N days; `0` = forever. See [Sessions](sessions.md). |
632
+ | `session.max_count` | `500` | yes | Keep the newest N; `0` = uncapped. |
633
+ | `session.keep_status` | `running` | yes | Comma list of statuses never pruned. |
634
+ | `session.sweep_interval_hours` | `24` | yes | How often the retention sweep runs. |
635
+ | `image.max_side` | `1568` | | See "Images". |
636
+ | `image.max_bytes` | `3750000` | | See "Images". |
637
+ | `image.max_per_request` | `20` | | See "Images". |
638
+ | `guardrails.enabled` | `true` | | See [Guardrails](guardrails.md). |
639
+ | `log.file` | state dir | yes | See "Debug Log File". |
640
+ | `log.disable` | `false` | yes | No file logging. |
641
+ | `log.level` | `info` | yes | `debug`, `info`, `warn`, `error`. |
642
+ | `status.line` | `on` | yes | The REPL status line, `on` or `off`. |
643
+ | `status.width_mode` | `terminal_cap` | yes | `terminal_cap` (terminal width up to `max_width`) or `fixed`. |
644
+ | `status.max_width` | `160` | yes | Cap for `terminal_cap`. |
645
+ | `status.fixed_width` | `120` | yes | Width for `fixed`. |
646
+ | `context.status` | `true` | yes | Context-usage telemetry for the model. See [context telemetry](internals/context-telemetry.md). |
647
+ | `context.window_tokens` | server's, else 256000 | yes | Context window when the server doesn't report one. |
648
+ | `context.chars_per_token` | `4.0` | yes | Estimate ratio when the server reports no usage. |
649
+ | `context.status_thresholds` | `20,40,60,80` | yes | Percentages that trigger a status. |
650
+ | `context.status_cadence` | `0` | yes | Also every N rounds; `0` = thresholds only. |
651
+ | `thinking.ui` | `spinner` | yes | `spinner` or `off`. |
652
+ | `thinking.render_interval` | `0.08` | yes | Seconds between thinking redraws. |
653
+ | `thinking.turn_preamble` | `true` | yes | Ask a `qwen36` model to open its thinking with a short `TURN:` line (the step label). |
654
+ | `max_tool_output_chars` | `10000` | yes | Tool output kept in the conversation; a top-level key (see below). |
655
+ | `retry.max` | `5` | yes | See "Llama Network Retry Behavior". |
656
+ | `retry.base_delay` | `0.5` | yes | |
657
+ | `retry.max_delay` | `8.0` | yes | |
658
+ | `retry.empty_answer` | `1` | | Times a turn asks again after an empty answer (at most 3, `0` = off). See "Llama Network Retry Behavior". |
659
+ | `read.truncate_at_bytes` | `65536` | yes | A `read` result larger than this is cut to a preview. |
660
+ | `read.preview_bytes` | `12288` | yes | Size of that preview. |
661
+ | `read.hard_max_bytes` | `2097152` | yes | Largest file `read` opens. |
662
+ | `read.telemetry_threshold_pct` | `80` | yes | A `read` result that alone fills this % of the context window carries a token estimate. |
663
+ | `execute.truncate_at_bytes` | `65536` | yes | The same for `execute` output. |
664
+ | `execute.preview_bytes` | `12288` | yes | |
665
+ | `execute.telemetry_threshold_pct` | `80` | yes | |
666
+ | `web.port` | `4567` | `--port` | `chi web`'s port. See [CLI](cli.md). |
667
+ | `web.host` | `127.0.0.1` | yes | `127.0.0.1`, `::1` or `localhost`. |
668
+ | `web.markdown` | `false` | yes | Render answers as Markdown in `chi web`. |
669
+ | `web.turn_view` | `true` | yes | One block per turn; `false` = the row of bubbles. |
670
+ | `web.annotate_presets` | `Agreed\|Could you please elaborate?` | yes | Quick replies next to Annotate in `chi web`, `\|`-separated (a YAML list works too); `""` in the file or on the CLI leaves only Annotate (an empty env value means the default). See [CLI](cli.md#web-annotate-presets). |
671
+ | `history.file` | state dir | | Prompt history path. |
672
+ | `no_interrupt` | `false` | `--no-interrupt` | Raise the tool-call limit of a turn to 1000; a top-level key. |
673
+ | `no_default_input` | `false` | `--no-default-input` | Don't pre-fill `default.input`; a top-level key. |
674
+ | `skip_agent_md` | `false` | | Don't load AGENT.md; a top-level key (see below). |
675
+
676
+ `max_tool_output_chars`, `skip_agent_md`, `no_interrupt` and
677
+ `no_default_input` have no section: in `config.yml` they are top-level keys as
678
+ written (`max_tool_output_chars: 20000`).
data/docs/desktop.md CHANGED
@@ -27,6 +27,12 @@ it's still live; if there's only one live session, that one is (a recent one nev
27
27
  session starts its worker; a note to one waits for its next start, and the panel shows chi's line saying so. After a
28
28
  send the panel shows chi's line and closes. On an error it stays open and shows the error.
29
29
 
30
+ The first row, **New session in <folder>** (⌘0), starts a session with the message instead: `chi send --new --dir
31
+ <folder>`, so it shows in `chi web` at once (see [Starting a session](sessions.md#starting-a-session)). The folder is
32
+ that of the most recently updated live session, else of the newest recent one, else your home folder. It is ticked
33
+ alone (ticking it clears the sessions and the reverse) and is preselected when no session is live. ⌘⏎ on it beeps: a
34
+ note needs a session. After the send the panel shows `started <id>…` for 3 s.
35
+
30
36
  A session open in a `chi --no-shared` REPL isn't listed: it takes no notes or messages.
31
37
 
32
38
  ## Install
data/docs/hooks.md CHANGED
@@ -50,7 +50,7 @@ The plugin class must respond to `#call(event)` — duck-typed, no base class re
50
50
  |-------|--------------|---------------|
51
51
  | `:session_start` | First turn of the session | `{ type: :session_start, session_id: "..." }` |
52
52
  | `:before_turn` | Before each turn starts | `{ type: :before_turn, session_id: "...", prompt: "..." (nil on a continue), messages: [...] (the history before this turn) }` |
53
- | `:after_turn` | After a turn completed or was cancelled (not after one that failed) | `{ type: :after_turn, status: "completed" \| "canceled", messages: [...] (the conversation the turn stored; a cancelled or empty turn ends it with a `kind: turn_note` system message, and a context line is `kind: context`, see [sessions.md](sessions.md#notes-a-turn-leaves-for-the-model)) }` |
53
+ | `:after_turn` | After a turn completed or was cancelled (not after one that failed) | `{ type: :after_turn, status: "completed" \| "canceled", present: (see [Presenting the answer](#presenting-the-answer-display-only)), messages: [...] (the conversation the turn stored; a cancelled or empty turn ends it with a `kind: turn_note` system message, and a context line is `kind: context`, see [sessions.md](sessions.md#notes-a-turn-leaves-for-the-model)) }` |
54
54
  | `:before_generation` | Before each LLM API call (both loops) | `{ type: :before_generation, iteration: N }` |
55
55
  | `:after_generation` | After LLM returns (both loops) | `{ type: :after_generation, iteration: N, response: "...", messages: [...] (the conversation as sent) }` |
56
56
  | `:before_tool_call` | Before tool dispatch (and before `tool_call_started`) | `{ type: :before_tool_call, iteration: N, call: {...}, params: "...", guardrail: Verdict, context: {...}, targets: {...}, blocked: false, block_reason: nil }` |
@@ -59,7 +59,7 @@ The plugin class must respond to `#call(event)` — duck-typed, no base class re
59
59
 
60
60
  Every event also carries the hook runtime (next section): `hook:` (the label
61
61
  of the hook about to run) and the callables `notify:`, `ask_user:`,
62
- `stop_turn:`.
62
+ `stop_turn:`, `steer:`.
63
63
 
64
64
  `messages:` is a **read-only copy**: a frozen array of copied message hashes
65
65
  (`{role:, content:, …}`). A hook that mutates it, or its strings, gets
@@ -69,8 +69,8 @@ cheap).
69
69
  ## What a hook can do: the runtime
70
70
 
71
71
  Besides reading (and, on `:before_tool_call`, voting on) its event, a hook
72
- can talk to the user through three callables the registry puts on every
73
- event:
72
+ can talk to the user, and to the running turn, through four callables the
73
+ registry puts on every event:
74
74
 
75
75
  ```ruby
76
76
  class Watchful
@@ -94,6 +94,14 @@ class Watchful
94
94
  # that call, and the rest of the batch is denied; from :after_turn or
95
95
  # :session_end it does nothing (false).
96
96
  event[:stop_turn].call("too many iterations without progress") if event[:iteration] > 20
97
+ when :after_tool_call
98
+ # Put text into the running turn, as the user's steering does: at the
99
+ # loop's next boundary it joins the conversation as its own user
100
+ # message, and every UI shows a nudge line ("<bundle> nudged: …").
101
+ # True when queued; false with no turn, and from :after_turn or
102
+ # :session_end. A steer that arrives after the model's final answer is
103
+ # dropped (logged), not merged: it never keeps a finished turn going.
104
+ event[:steer].call("Say briefly what you have found so far.") if event[:iteration] == 30
97
105
  end
98
106
  end
99
107
  end
@@ -103,11 +111,55 @@ end
103
111
  known-names)` for a bundle hook, `audit.rb (config)` for a config hook,
104
112
  `turn hook` for one registered at runtime.
105
113
 
114
+ A steer is saved in the session as `{role: "user", kind: "steer", source:
115
+ "<bundle>", content: "…"}`; the model reads only its text, as a user turn.
116
+ Its `source` is the hook's bundle (a config or turn hook's label otherwise).
117
+
106
118
  Timing: a notice from `:after_turn` or `:session_end` shows after the turn's
107
119
  end line. A question from `:before_tool_call` shows **before** the tool
108
120
  line (the gate runs first), so its text should name the call. The notices
109
121
  are also logged (`turn` tag, `hook_notice`).
110
122
 
123
+ ## Presenting the answer (display only)
124
+
125
+ `:after_turn` carries one more callable, `present:`. It changes how the
126
+ turn's answer is **shown**, never what the model said: the block gets the
127
+ current display text (the answer's content until a hook changed it) and
128
+ returns the new one.
129
+
130
+ ```ruby
131
+ class Shout
132
+ def call(event)
133
+ return unless event[:type] == :after_turn
134
+
135
+ event[:present].call { |text| text.gsub(/\bTODO\b/, "**TODO**") }
136
+ end
137
+ end
138
+ ```
139
+
140
+ - The result is kept as `display` on the answer's model message in the
141
+ session file. The model never sees it: the prompts and chat requests take
142
+ the fields they send, and the copies of the conversation given to hooks
143
+ (`messages:`), plugins (`ctx.messages`) and the recap leave it out.
144
+ - Calls chain in hook order (bundle hooks by priority, then config hooks,
145
+ then turn hooks): each block gets what the one before returned. The call
146
+ returns the display text after it.
147
+ - A block that raises, returns something other than a String, or returns
148
+ more than 200 000 characters leaves the display as it was (logged as
149
+ `present_rejected` with the hook's label).
150
+ - It works on the stored conversation's last message only when that is the
151
+ model's answer: after a cancelled, failed or empty turn there is none, and
152
+ the call returns nil without running the block.
153
+ - **The web** renders `display` instead of the answer (markdown, sanitised
154
+ like every answer: raw HTML is escaped, only http(s)/mailto links are
155
+ kept), on a live turn and after a reload. Its copy button copies the
156
+ display text. The page learns about it from an `answer_display` event
157
+ that comes after `turn_completed` (the hooks run after the turn ended).
158
+ - **Terminals** (the REPL, the attached TUI) have printed the answer by then
159
+ and do not change it; use `event[:notify]` for something they should show.
160
+
161
+ A plugin gets the same from `chi.on(:after_turn) { |event, ctx| event[:present].call { … } }`.
162
+
111
163
  ## Settings
112
164
 
113
165
  A hook class whose `initialize` takes an argument gets its settings: **one
@@ -299,7 +351,7 @@ Notes:
299
351
  - Ordering: bundle hooks fire by `(priority, bundle_name, hook_name)` (lower priority first), then plain `config.yml` hooks in registration order.
300
352
  - Settings: a hook class with `initialize(settings = {})` gets the bundle's section of `config.yml` `bundles:` (see [Settings](#settings)).
301
353
  - A bundle can also ship a `plugin.rb` whose `chi.on(event)` blocks are bundle hooks too, next to commands and tools; see [Plugins](plugins.md).
302
- - Shipped bundles: `chi bundle install guardrails` (rules, see [Guardrails](guardrails.md#the-guardrails-bundle)) and `chi bundle install known-names` (a hook, see [Guardrails](guardrails.md#the-known-names-bundle)), `chi bundle install btw` (a plugin: `/btw`, see [Plugins](plugins.md#the-btw-bundle)), `chi bundle install mcp` (a plugin: tools from MCP servers, see [Plugins](plugins.md#the-mcp-bundle)) and `chi bundle install loop-guard` (a plugin: breaks tool-call loops, see [Plugins](plugins.md#the-loop-guard-bundle)).
354
+ - Shipped bundles: `chi bundle install guardrails` (rules, see [Guardrails](guardrails.md#the-guardrails-bundle)) and `chi bundle install known-names` (a hook, see [Guardrails](guardrails.md#the-known-names-bundle)), `chi bundle install source-links` (a hook: announces source refs, see [The source-links bundle](#the-source-links-bundle)), `chi bundle install btw` (a plugin: `/btw`, see [Plugins](plugins.md#the-btw-bundle)), `chi bundle install mcp` (a plugin: tools from MCP servers, see [Plugins](plugins.md#the-mcp-bundle)) `chi bundle install loop-guard` (a plugin: breaks tool-call loops, see [Plugins](plugins.md#the-loop-guard-bundle)) and `chi bundle install check-in` (a plugin: checks on a long turn, see [Plugins](plugins.md#the-check-in-bundle)).
303
355
  - Installing a bundle executes its hook code at `Engine` startup. Only install bundles you trust, as you would a gem. Hooks are **not** executed at install time (copy-only); they are `module_eval`'d at `Engine.new` inside per-bundle `Samagotchi::Bundles::<name>` namespaces (no top-level `require` collisions). Keep hook files side-effect-free at load time; do work in `#call` — top-level side effects (require, IO, `at_exit`, global assignment) run once per `Engine.new` (class redefinition is idempotent).
304
356
 
305
357
  Lifecycle:
@@ -307,3 +359,72 @@ Lifecycle:
307
359
  - `chi bundle install <source>` copies `hooks/*.rb` to `~/.config/samagotchi/memories/.bundles/<name>/hooks/` and persists metadata + `trust_level` + `source_commit` (git HEAD) to provenance.
308
360
  - `Engine.new` loads `config.yml` hooks first, then bundle hooks via `Provenance.each_installed_holding_hooks` → `Hooks::BundleLoader.load`. Bundle hooks are process-scoped (they survive the per-turn `clear_hooks`; only plain hooks are cleared). Experimental bundles emit a one-line startup warning.
309
361
  - `chi bundle status`, `diff`, `uninstall`, `build` are hook-aware (counts, metadata, removal).
362
+
363
+ ## The source-links bundle
364
+
365
+ ```sh
366
+ chi bundle install source-links
367
+ ```
368
+
369
+ installs one `after_turn` hook and a short memory. When the model's answer
370
+ mentions a known source ref — a JIRA ticket, a GitHub issue, an internal
371
+ wiki page — the web links it in the answer (below), and every UI shows one
372
+ line right after the message:
373
+
374
+ ```
375
+ sources: JIRA JIRA-123 → https://myjira.com/browse/JIRA-123, JIRA JIRA-10 → https://myjira.com/browse/JIRA-10
376
+ ```
377
+
378
+ The note is **not part of the conversation**: it is an event shown to the
379
+ user, never stored in the session file. A UI replays it while the session's
380
+ worker lives (a page reload keeps it; a stopped worker loses it). Only the
381
+ model's final answer is scanned (the last `role: "model"` message), and only
382
+ the first 20 000 characters of it. A ref that is already a link is skipped —
383
+ inside a bare URL (`https://x.com/JIRA-123`), in a markdown link's target, or
384
+ in a markdown link's label when the target names the same ref
385
+ (`[JIRA-123](https://x.com/JIRA-123)`) — while `see https://x.com JIRA-123`
386
+ and `[fix for JIRA-123](https://github.com/o/r/pull/9)` still link. A ref
387
+ glued to URL punctuation (`/browse/JIRA-1`, `?key=JIRA-1`, `JIRA-1/foo`) is
388
+ skipped too; `Ticket:JIRA-5` and `#JIRA-123` are ordinary plain text and do
389
+ link. Refs are deduped within the turn (case-insensitively) and listed in
390
+ first-occurrence order, whatever order the sources are configured in. With no
391
+ sources configured the hook is a silent no-op.
392
+
393
+ ```yaml
394
+ bundles:
395
+ source-links:
396
+ sources:
397
+ - name: JIRA
398
+ prefix: JIRA # simple form: \bJIRA-(\d+)\b
399
+ base_url: https://myjira.com/browse/
400
+ - name: GitHub
401
+ pattern: '\bGH-(\d+)\b' # full form: a regex
402
+ url: 'https://github.com/org/repo/issues/{match}'
403
+ case_insensitive: false # optional, default false
404
+ max: 10 # optional: refs per line, default 10
405
+ note: false # optional: no sources line, default true
406
+ ```
407
+
408
+ **In the web** the refs are also links in the answer itself: each
409
+ occurrence becomes `[JIRA-123](https://myjira.com/browse/JIRA-123)` through
410
+ [`event[:present]`](#presenting-the-answer-display-only), so the model's
411
+ text stays as it was, and the links survive a reload and a stopped worker
412
+ (they are the answer's `display` in the session file). The same skip rules
413
+ apply, and a ref in code (a `` `span` `` or a fenced block) or anywhere in a
414
+ markdown link is left as it is; past the first 20 000 characters the answer
415
+ is unchanged. The terminals see only the line; `note: false` drops it and
416
+ keeps the web links.
417
+
418
+ The `prefix:` form compiles to `\b<prefix>-(\d+)\b` and the URL is
419
+ `base_url` + the full ref text (`JIRA-123`). The `pattern:` form takes a
420
+ regex; `{match}` in `url` is replaced with the first capture group (or the
421
+ full match when the pattern has none). `case_insensitive: true` adds the
422
+ `/i` flag. Past `max` refs the line ends with `… +N more`.
423
+
424
+ Each regex is compiled with a per-regex timeout (0.5 s, per match attempt),
425
+ so a catastrophic pattern is abandoned instead of hanging the turn: that
426
+ source is skipped whole (its partial matches are discarded) with a warning,
427
+ and the others still report. An entry with neither `prefix:` nor `pattern:`,
428
+ or a pattern that does not compile, is skipped with a warning at load. The
429
+ hook is `on_error: log`: a bug in it warns and the turn is unaffected. As with
430
+ every bundle hook, a running worker picks it up after its next start.