@iowarp/clio-coder 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (226) hide show
  1. package/CHANGELOG.md +407 -0
  2. package/CODE_OF_CONDUCT.md +21 -0
  3. package/CONTRIBUTING.md +224 -0
  4. package/LICENSE +202 -0
  5. package/NOTICE +9 -0
  6. package/README.md +798 -0
  7. package/SECURITY.md +72 -0
  8. package/assets/clio-coder-logo-128.webp +0 -0
  9. package/damage-control-rules.yaml +419 -0
  10. package/dist/acp-UMLFVA3F.js +92 -0
  11. package/dist/agents-Q4MYPMUW.js +91 -0
  12. package/dist/auth-O6HYIJ6J.js +521 -0
  13. package/dist/chunk-262G75JS.js +35 -0
  14. package/dist/chunk-26BZQOAD.js +1281 -0
  15. package/dist/chunk-2J63S4SF.js +508 -0
  16. package/dist/chunk-3DANZDGR.js +717 -0
  17. package/dist/chunk-4UQA7NCT.js +29 -0
  18. package/dist/chunk-527KG6XR.js +497 -0
  19. package/dist/chunk-5LDRNKX2.js +1063 -0
  20. package/dist/chunk-5N2FG33Q.js +25 -0
  21. package/dist/chunk-67MTHP2E.js +135 -0
  22. package/dist/chunk-6CWDTGUC.js +20 -0
  23. package/dist/chunk-7BHLZB3A.js +2115 -0
  24. package/dist/chunk-7RBKDI66.js +348 -0
  25. package/dist/chunk-AMFR5YA3.js +541 -0
  26. package/dist/chunk-BBUH4VAA.js +1224 -0
  27. package/dist/chunk-BYEU76JP.js +899 -0
  28. package/dist/chunk-CLJ5HLUD.js +458 -0
  29. package/dist/chunk-D5YD55AR.js +116 -0
  30. package/dist/chunk-DXQNI4PC.js +61 -0
  31. package/dist/chunk-E3NYWENM.js +1004 -0
  32. package/dist/chunk-GNGDQYDU.js +34688 -0
  33. package/dist/chunk-GOTUR54M.js +9 -0
  34. package/dist/chunk-HBU5MTAM.js +41 -0
  35. package/dist/chunk-HMYNFFY4.js +28 -0
  36. package/dist/chunk-JPOWPFCU.js +1010 -0
  37. package/dist/chunk-JWHCJDCI.js +1215 -0
  38. package/dist/chunk-KBR4MZZR.js +41 -0
  39. package/dist/chunk-KKKPTZLM.js +93 -0
  40. package/dist/chunk-ME6DNWIU.js +66 -0
  41. package/dist/chunk-NI4DEJMC.js +88 -0
  42. package/dist/chunk-O4EJEDHO.js +659 -0
  43. package/dist/chunk-PIDUD6M2.js +31 -0
  44. package/dist/chunk-PS4PFJQP.js +29459 -0
  45. package/dist/chunk-QV47YRF4.js +48 -0
  46. package/dist/chunk-RQDWMVRB.js +279 -0
  47. package/dist/chunk-TFSSEXL6.js +136 -0
  48. package/dist/chunk-TKHQ4DGZ.js +8290 -0
  49. package/dist/chunk-TPOCL34A.js +2876 -0
  50. package/dist/chunk-UGYAX5YI.js +565 -0
  51. package/dist/chunk-UHTSULZS.js +461 -0
  52. package/dist/chunk-UU3R62TT.js +128 -0
  53. package/dist/chunk-UWIJNAOB.js +3906 -0
  54. package/dist/chunk-VOO7NYPP.js +914 -0
  55. package/dist/chunk-VPAWTYLY.js +117 -0
  56. package/dist/chunk-WD6AJM35.js +1216 -0
  57. package/dist/chunk-X3BR7HWV.js +115 -0
  58. package/dist/chunk-X3NE4WVW.js +120 -0
  59. package/dist/chunk-XNISANGE.js +1395 -0
  60. package/dist/chunk-XV4ZJ6ZM.js +3177 -0
  61. package/dist/cli/index.js +236 -0
  62. package/dist/clio-KIQ5SNDS.js +53 -0
  63. package/dist/components-JVHMUBEB.js +653 -0
  64. package/dist/config-ZFCDBMDC.js +372 -0
  65. package/dist/configure-G4E3A2PG.js +27 -0
  66. package/dist/context-CDXTP2MP.js +293 -0
  67. package/dist/context-E3KIFVXI.js +185 -0
  68. package/dist/context-clear-3F4PLXOS.js +102 -0
  69. package/dist/context-index-Q7YSYTR3.js +106 -0
  70. package/dist/docs-YIETIWZI.js +280 -0
  71. package/dist/doctor-M5HJJZOL.js +61 -0
  72. package/dist/domains/agents/builtins/architect.md +33 -0
  73. package/dist/domains/agents/builtins/coder.md +31 -0
  74. package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
  75. package/dist/domains/agents/builtins/debugger.md +30 -0
  76. package/dist/domains/agents/builtins/documenter.md +31 -0
  77. package/dist/domains/agents/builtins/git-master.md +30 -0
  78. package/dist/domains/agents/builtins/provenance.md +30 -0
  79. package/dist/domains/agents/builtins/researcher.md +71 -0
  80. package/dist/domains/agents/builtins/scout.md +42 -0
  81. package/dist/domains/agents/builtins/tester.md +31 -0
  82. package/dist/domains/agents/builtins/verifier.md +30 -0
  83. package/dist/domains/agents/builtins/wiki-writer.md +41 -0
  84. package/dist/eval-B3KZZESM.js +2674 -0
  85. package/dist/evidence-V67CHM35.js +233 -0
  86. package/dist/evolve-YDZSUQYA.js +518 -0
  87. package/dist/extensions-SRG7XCAH.js +207 -0
  88. package/dist/fleet-CA2CRTVG.js +760 -0
  89. package/dist/fleet-preflight-CLIAX7YR.js +21 -0
  90. package/dist/init-2OZDJE2D.js +227 -0
  91. package/dist/memory-3PIQQAKX.js +207 -0
  92. package/dist/models-DY35XI7Y.js +237 -0
  93. package/dist/paths-5OMXW7Z4.js +57 -0
  94. package/dist/preload-KZVHET2B.js +11 -0
  95. package/dist/reset-PIFYNOS3.js +216 -0
  96. package/dist/run-3VSPP24F.js +735 -0
  97. package/dist/share-D36RQCXM.js +241 -0
  98. package/dist/skills-F2MRLELY.js +445 -0
  99. package/dist/skills-eval-E2ZTW4PL.js +932 -0
  100. package/dist/targets-DZMEZAH4.js +977 -0
  101. package/dist/trace-7NYCUI2J.js +250 -0
  102. package/dist/uninstall-AD3JWHBB.js +322 -0
  103. package/dist/upgrade-WYYBKGDY.js +301 -0
  104. package/dist/usage-ULIDAGFF.js +755 -0
  105. package/dist/version-ROZ6CZKH.js +16 -0
  106. package/dist/wiki-generate-PKFIX6OB.js +377 -0
  107. package/dist/worker/entry.js +1739 -0
  108. package/docs/README.md +93 -0
  109. package/docs/acp.md +120 -0
  110. package/docs/alcf-provider.md +72 -0
  111. package/docs/architecture.md +172 -0
  112. package/docs/artifact-versions.md +54 -0
  113. package/docs/built-in-agents.md +265 -0
  114. package/docs/capacity-and-scheduling.md +97 -0
  115. package/docs/commands-and-modes.md +554 -0
  116. package/docs/config-knobs-audit.md +115 -0
  117. package/docs/configuration-and-targets.md +812 -0
  118. package/docs/context-engine.md +236 -0
  119. package/docs/dispatch-architecture-rationale.md +126 -0
  120. package/docs/documentation-coverage.md +46 -0
  121. package/docs/documentation-guide.md +166 -0
  122. package/docs/environment-variables.md +105 -0
  123. package/docs/eval-runner.md +205 -0
  124. package/docs/evals-internal.md +298 -0
  125. package/docs/evidence-and-memory.md +243 -0
  126. package/docs/evolution.md +143 -0
  127. package/docs/exit-codes-and-output.md +74 -0
  128. package/docs/extensions-and-sharing.md +306 -0
  129. package/docs/fleet-demo-runbook.md +179 -0
  130. package/docs/fleet-dispatch.md +591 -0
  131. package/docs/glossary.md +75 -0
  132. package/docs/html/agents_blueprint.html +936 -0
  133. package/docs/html/alcf_blueprint.html +324 -0
  134. package/docs/html/architecture_blueprint.html +850 -0
  135. package/docs/html/commands_blueprint.html +794 -0
  136. package/docs/html/config_knobs_audit_blueprint.html +178 -0
  137. package/docs/html/configuration_blueprint.html +1080 -0
  138. package/docs/html/context_blueprint.html +603 -0
  139. package/docs/html/documentation_blueprint.html +832 -0
  140. package/docs/html/environment_blueprint.html +404 -0
  141. package/docs/html/eval_blueprint.html +743 -0
  142. package/docs/html/evals_internal_blueprint.html +190 -0
  143. package/docs/html/evolution_blueprint.html +674 -0
  144. package/docs/html/extensions_blueprint.html +2065 -0
  145. package/docs/html/fleet_dispatch_blueprint.html +286 -0
  146. package/docs/html/index.html +919 -0
  147. package/docs/html/lifecycle_blueprint.html +723 -0
  148. package/docs/html/memory_blueprint.html +699 -0
  149. package/docs/html/middleware_blueprint.html +664 -0
  150. package/docs/html/models_blueprint.html +2366 -0
  151. package/docs/html/observability_blueprint.html +683 -0
  152. package/docs/html/provider_adapter_blueprint.html +245 -0
  153. package/docs/html/safety_blueprint.html +1386 -0
  154. package/docs/html/shared.css +571 -0
  155. package/docs/html/shared.js +143 -0
  156. package/docs/html/skills_blueprint.html +671 -0
  157. package/docs/html/soak_blueprint.html +182 -0
  158. package/docs/html/tool_usage_blueprint.html +350 -0
  159. package/docs/html/tools_blueprint.html +2249 -0
  160. package/docs/html/trace_blueprint.html +235 -0
  161. package/docs/html/tui_design_blueprint.html +314 -0
  162. package/docs/html/validation_blueprint.html +961 -0
  163. package/docs/html/worker_dispatch_blueprint.html +231 -0
  164. package/docs/installation-and-lifecycle.md +308 -0
  165. package/docs/middleware-and-components.md +148 -0
  166. package/docs/model-catalog.md +189 -0
  167. package/docs/observability.md +233 -0
  168. package/docs/proactive-memory.md +452 -0
  169. package/docs/prompt-envelope-and-tools.md +142 -0
  170. package/docs/provider-adapter-cookbook.md +148 -0
  171. package/docs/release-cut-checklist.md +138 -0
  172. package/docs/safety-model.md +357 -0
  173. package/docs/scientific-validation.md +105 -0
  174. package/docs/session-lifecycle.md +156 -0
  175. package/docs/skills-marketplace.md +46 -0
  176. package/docs/tool-usage.md +527 -0
  177. package/docs/trace-store.md +132 -0
  178. package/docs/troubleshooting.md +33 -0
  179. package/docs/tui-design.md +239 -0
  180. package/docs/worker-dispatch-mechanics.md +242 -0
  181. package/package.json +132 -0
  182. package/skills/README.md +408 -0
  183. package/skills/git/commit-crafting/SKILL.md +79 -0
  184. package/skills/git/commit-crafting/evals.md +92 -0
  185. package/skills/git/create-pr/SKILL.md +116 -0
  186. package/skills/git/create-pr/evals.md +114 -0
  187. package/skills/git/investigate-issue/SKILL.md +139 -0
  188. package/skills/git/investigate-issue/evals.md +94 -0
  189. package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
  190. package/skills/git/resolve-merge-conflicts/evals.md +58 -0
  191. package/skills/git/review-changes/SKILL.md +103 -0
  192. package/skills/git/review-changes/evals.md +85 -0
  193. package/skills/git/worktree-create/SKILL.md +92 -0
  194. package/skills/git/worktree-create/evals.md +97 -0
  195. package/skills/git/worktree-create/references/worktree-setup.md +66 -0
  196. package/skills/git/worktree-merge/SKILL.md +95 -0
  197. package/skills/git/worktree-merge/evals.md +114 -0
  198. package/skills/skill-marketplace.json +261 -0
  199. package/skills/workflow/cut-it/SKILL.md +86 -0
  200. package/skills/workflow/cut-it/evals.md +42 -0
  201. package/src/domains/agents/builtins/architect.md +33 -0
  202. package/src/domains/agents/builtins/coder.md +31 -0
  203. package/src/domains/agents/builtins/context-bootstrap.md +38 -0
  204. package/src/domains/agents/builtins/debugger.md +30 -0
  205. package/src/domains/agents/builtins/documenter.md +31 -0
  206. package/src/domains/agents/builtins/git-master.md +30 -0
  207. package/src/domains/agents/builtins/provenance.md +30 -0
  208. package/src/domains/agents/builtins/researcher.md +71 -0
  209. package/src/domains/agents/builtins/scout.md +42 -0
  210. package/src/domains/agents/builtins/tester.md +31 -0
  211. package/src/domains/agents/builtins/verifier.md +30 -0
  212. package/src/domains/agents/builtins/wiki-writer.md +41 -0
  213. package/src/domains/agents/fleets/build-review.md +34 -0
  214. package/src/domains/agents/fleets/build-test.md +35 -0
  215. package/src/domains/agents/fleets/sdlc.md +86 -0
  216. package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
  217. package/src/domains/prompts/fragments/identity/clio.md +26 -0
  218. package/src/domains/prompts/fragments/operating/contract.md +64 -0
  219. package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
  220. package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
  221. package/src/domains/prompts/fragments/safety/read-only.md +13 -0
  222. package/src/domains/prompts/fragments/safety/suggest.md +13 -0
  223. package/src/domains/prompts/fragments/wiki/page.md +75 -0
  224. package/src/domains/prompts/fragments/wiki/plan.md +48 -0
  225. package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
  226. package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
@@ -0,0 +1,452 @@
1
+ # Proactive task memory
2
+
3
+ > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.0).
4
+
5
+ Clio's proactive task memory protects long-running work from behavioral state
6
+ decay: a requirement, environment fact, failed attempt, or diagnosis can still
7
+ exist in the transcript while no longer influencing the next action. The design
8
+ follows Wu et al., *Remember When It Matters: Proactive Memory Agent for
9
+ Long-Horizon Agents* (2026), adapted to Clio's visible middleware and local-model
10
+ routing.
11
+
12
+ The rules-only tier is enabled by default and makes no model calls. An LLM memory
13
+ tier is opt-in through the independent `background` route. The action agent's
14
+ system prompt and tool surface do not change, and disabling
15
+ `memory.intervention.enabled` removes observation, bank writes, model resolution,
16
+ reminders, handoff offers, and handoff seeding.
17
+
18
+ ## Architecture
19
+
20
+ ```mermaid
21
+ flowchart LR
22
+ T[tool and lifecycle hooks] --> R[memory intervention registration]
23
+ R --> B[session task bank]
24
+ R --> D{trigger boundary}
25
+ D -->|rules only| S[deterministic policy]
26
+ D -->|background configured| L[two-phase local model policy]
27
+ S --> V[visible advisory reminder]
28
+ L --> V
29
+ V --> U[next user turn and session ledger]
30
+ R -. counts and outcomes only .-> J[bounded state JSONL]
31
+ B -. explicit context-handoff .-> H[redacted handoff snapshot]
32
+ ```
33
+
34
+ The task bank is in `src/domains/memory/` and belongs to one live session. It is
35
+ separate from both durable approved lessons and the regenerable repository
36
+ context engine.
37
+
38
+ - **Private status** is the memory policy's short progress model. It can be
39
+ inspected with `/memory`, but is never rendered to either model, injected into
40
+ the action turn, or exported to a handoff.
41
+ - **Knowledge** contains stable task facts such as requirements, paths,
42
+ environment facts, and constraints.
43
+ - **Procedural memory** contains attempts and outcomes such as failed commands,
44
+ ruled-out hypotheses, diagnoses, and fixes that worked.
45
+
46
+ Knowledge and procedural entries have stable short IDs. A visible reminder is
47
+ one `Memory:` advisory block and records the cited entry IDs' injection counts.
48
+ The existing middleware path places that block in the next submitted user turn
49
+ and persists its attribution in the session ledger; there is no hidden
50
+ `transformContext` injection.
51
+
52
+ ## Paper mapping and Clio constraints
53
+
54
+ | Paper mechanism | Clio implementation |
55
+ | --- | --- |
56
+ | Separate memory agent | One stateful middleware registration beside the unmodified action agent |
57
+ | Status, knowledge, and procedural bank | Bounded in-memory `TaskMemoryBank`; status stays out of every reminder except the post-compaction restore |
58
+ | Phase 1 bank maintenance | Strict `update_status`, `save_knowledge`, `save_procedural`, and `delete` operations |
59
+ | Phase 2 intervene or stay silent | One advisory `inject_reminder` effect or explicit silence |
60
+ | Fixed memory cadence | Deterministic decay signals plus a coarse interval floor |
61
+ | Learned intervention calibration | Structural authority gate: spontaneous reminders must cite a bank entry; deterministic triggers may be uncited |
62
+ | Passive and always-on ablations | A/B harness compares baseline, rules, and LLM tiers and flags always-noisy ties as regressions |
63
+
64
+ Model output uses a strict two-line grammar parsed by `src/domains/memory/task-memory-policy.ts`:
65
+
66
+ ```text
67
+ <operations>[{"op":"update_status","content":"Tracking the current requirement."}]</operations>
68
+ <no_intervention/>
69
+ ```
70
+
71
+ When a visible reminder cites an existing bank entry, the second line uses:
72
+ ```text
73
+ <operations>[]</operations>
74
+ <context_for_action>Restored requirement [tm-k-1]</context_for_action>
75
+ ```
76
+
77
+ Knowledge and procedural saves use `{"op":"save_knowledge","content":"..."}` or
78
+ `{"op":"save_procedural","content":"..."}`; an `id` is only valid when updating
79
+ an existing entry.
80
+
81
+ The parser locates that envelope rather than matching the response byte for byte,
82
+ because a small local model routinely delivers a correct decision inside
83
+ imperfect packaging. A markdown fence, a `<think>` block, a leading "Here is my
84
+ step:", a closing pleasantry, and a pretty-printed multi-line operations array
85
+ are all accepted. Two shapes are read conservatively rather than generously:
86
+
87
+ - A response with no phase-two line at all is silence, since silence is the
88
+ prompt's documented default. Its phase-one writes still apply. An envelope
89
+ truncated mid-reasoning yields nothing at all.
90
+ - Tag shapes are stripped from the reminder before it is emitted, so the memory
91
+ model cannot close the `<system-reminder>` block it rides inside. Ordinary
92
+ comparisons and arrows survive.
93
+
94
+ Operations are validated structurally as a batch: a malformed entry or more than
95
+ eight operations rejects the whole list and changes nothing, because both say
96
+ the model did not produce an operation list at all. Two narrower mistakes cost
97
+ one operation instead of the step, because a small model makes both routinely
98
+ and the notes beside them are the point of the step.
99
+
100
+ Identity is repaired rather than rejected, because a small model invents a
101
+ descriptive id for content it is recording for the first time. A
102
+ `save_knowledge` or `save_procedural` whose id names no entry of that class
103
+ becomes a new entry, and a `delete` of an unknown id is dropped.
104
+
105
+ An unrecognized `op` is dropped the same way. Handed a JSON tool trajectory, a
106
+ small model borrows that trajectory's shape for an entry or two and answers
107
+ `{"op":"read","path":"..."}` beside otherwise valid saves; on the reference
108
+ route that happened in a third of sampled steps, and three of four such batches
109
+ carried a valid operation that the old whole-batch rejection discarded. A step
110
+ whose every operation was invented still records `malformed` rather than
111
+ passing as silence, since recovering nothing is not a decision to stay quiet.
112
+
113
+ Phase 1 writes remain valid when Phase 2 is gated or yields to a deterministic
114
+ reminder; an over-budget reminder is recorded as `gated` and suppressed rather
115
+ than discarding the writes that came with it. A timeout, provider failure,
116
+ malformed response, or telemetry failure is silent and never blocks a tool.
117
+
118
+ ### Intervention Defaults & Cadence Knobs
119
+ - `memory.intervention.enabled` (default `true`): Enables observation, task bank writes, and reminder injection.
120
+ - `memory.intervention.everyNTools` (default `10`): Minimum completed-tool interval between background interventions.
121
+ - `memory.intervention.windowSteps` (default `8`): Completed tool-trajectory window analyzed during background evaluation.
122
+ - `memory.intervention.maxTokens` (default `400`): Bounds the rendered memory-bank and reminder context budget; the policy model output cap is a separate fixed `4,000`-token contract in `task-memory-policy.ts`, sized so that a model which reasons anyway still reaches its envelope.
123
+ - `memory.intervention.timeoutMs` (default `180000`): Wall-clock limit for one background memory-policy request. The step is detached, so this deadline never delays a turn; set it above the observed step time for your route or finished work is discarded as a timeout. Step latency on a small local route is long-tailed rather than tightly clustered, so size this off a high percentile and not off a median.
124
+
125
+ ## Trigger semantics
126
+
127
+ Memory does not call a model after every tool. Signals accumulate and coalesce at
128
+ a turn-end boundary; at most one prompted step is started for that boundary, and
129
+ it runs detached from it.
130
+
131
+ | Trigger | Behavior |
132
+ | --- | --- |
133
+ | Interval | After `memory.intervention.everyNTools` completed tools since the last prompted step; default 10. This is the nondeterministic/citation-gated path. |
134
+ | Tool-error streak | Two consecutive error outcomes. A successful tool resets the streak. |
135
+ | Loop signal | Reuses the orchestrator loop guard's verdict; it does not infer a second competing loop detector. |
136
+ | Repeated failure | The rules tier records failed tool fingerprints and annotates the failing tool result once the same failure appears twice in the bounded trajectory. |
137
+ | Post-compaction | The first turn start after compaction restores status and knowledge once, without a model call, because compaction is precisely where execution facts leave the active window. |
138
+
139
+ ### Two delivery channels
140
+
141
+ A turn boundary is the wrong place to warn about a failure that happened forty
142
+ tool calls earlier in the same turn, because the reminder cannot reach the model
143
+ until the operator submits again. Memory therefore has two channels, and each
144
+ repeated failure uses exactly one of them:
145
+
146
+ - **Mid-turn annotation.** The second identical failure appends one cited
147
+ `Memory:` advisory to that tool's own result, through the existing
148
+ `annotate_tool_result` effect the loop guard already uses. The advisory digest
149
+ takes the first line of the tool error that names a problem, falling back to
150
+ the first line when no line names one. The model reads it on its very next round.
151
+ This is spent once per fingerprint per turn and re-earned in a later turn, because
152
+ the same command failing again after an operator turn is news again.
153
+ - **Next-turn reminder.** Post-compaction reactivation and any background-model
154
+ reminder ride the `inject_reminder` buffer into the next submitted turn, inside
155
+ the visible `<system-reminder>` block, and persist in the session ledger.
156
+
157
+ A boundary that already spoke through the annotation stays silent at turn end and
158
+ records one telemetry row, not two.
159
+
160
+ ## Background steps never hold a turn open
161
+
162
+ The pi agent does not become idle until every `agent_end` listener settles, so a
163
+ memory step awaited at that boundary would add its full latency to the visible
164
+ end of every triggered turn. Measured on the reference route below, step latency
165
+ has a median of 18.6 seconds and ranges up to 220.8 seconds. This makes an awaited
166
+ step intolerable as an end-of-turn pause.
167
+
168
+ The prompted step is therefore detached. `evaluateAsync` starts it and returns
169
+ immediately; the turn ends on schedule. When the step resolves, its reminder is
170
+ delivered through the deferred-reminder path into the next submitted turn, which
171
+ is exactly where an awaited turn_end reminder would have been buffered anyway.
172
+ Two consequences follow, both deliberate:
173
+
174
+ - At most one background step is alive per session. A boundary that arrives while
175
+ a step is still running is dropped rather than queued, so a model slower than
176
+ the turns that trigger it can never build a backlog. The drop is recorded as
177
+ its own telemetry row, because a cadence starved by a slow route and one that
178
+ simply never triggered are otherwise identical in the step log. Its triggers
179
+ stay pending, so the next free boundary still runs for them.
180
+ - A reminder can arrive one turn later than the trajectory that earned it. The
181
+ rules tier is unaffected and stays synchronous, so deterministic protection
182
+ keeps its original timing.
183
+
184
+ `/memory` shows whether a step is in flight, and the footer's memory row shows
185
+ `working` while one is running.
186
+
187
+ The error-streak, loop, repeated-failure, and post-compaction paths are
188
+ deterministic. A prompted reminder from one of those paths may be uncited. An
189
+ interval-only prompted reminder must cite at least one current knowledge or
190
+ procedural ID or it is recorded as `gated` and remains invisible.
191
+
192
+ ### Outcome semantics
193
+
194
+ The `/memory` overlay displays `last <decision>` where `<decision>` is the
195
+ combined outcome of the most recent actual memory boundary. A **memory boundary**
196
+ is a turn-end evaluation that includes newly completed tools or an explicit
197
+ deterministic trigger (interval, error-streak, loop signal). A no-tool
198
+ middleware continuation is another turn-end with no new tools since the previous
199
+ boundary. It is not a new memory boundary and does not replace the prior outcome.
200
+ Thus `last` remains `injected` across such continuations until a later
201
+ tool-bearing or explicitly triggered memory step produces a new outcome (e.g.,
202
+ a healthy tool leading to `silent`).
203
+
204
+ ## Choosing a background model
205
+
206
+ Memory reads a trajectory and writes a fixed envelope. It does not plan, and it
207
+ does not need to be clever. A small non-reasoning model is the right choice, and
208
+ Clio always requests the background route with thinking off regardless of
209
+ `background.thinkingLevel`.
210
+
211
+ That request reaches the wire wherever the runtime carries a thinking control:
212
+ llama.cpp reads `chat_template_kwargs.enable_thinking`, and LM Studio reads
213
+ `reasoning_effort`, where `none` is the off value.
214
+
215
+ A model that reasons anyway still works. Some genuinely cannot be silenced, and
216
+ the catalog records those as always-on so the level reads `forced` rather than
217
+ `off`; the shipped background model `qwopus3.5-9b-v3` is one of them. Reasoning
218
+ blocks are discarded and only the envelope is kept, and the output budget is
219
+ sized to let a reasoning preamble run its course first. The cost is latency,
220
+ which the detached step absorbs.
221
+
222
+ One configuration is refused rather than degraded. If the background role names
223
+ the same target and model as the orchestrator, and that model reasons, the LLM
224
+ memory tier stays off and memory runs on its free deterministic tier. A single
225
+ reasoning model already driving chat, workers, and shadow agents cannot also
226
+ deliberate over memory steps without contending with the work the operator
227
+ actually asked for.
228
+
229
+ This is a mix-and-match plane, not a local-only one. The background role resolves
230
+ through the same target machinery as every other role, so the useful shapes are:
231
+
232
+ - A frontier model for chat and a small efficient model for memory, whether that
233
+ small model is co-hosted, on another node, or a cheap cloud tier.
234
+ - A local workhorse for chat and a co-resident small local model for memory, with
235
+ the co-residency caveat below.
236
+ - Everything cloud: pick the provider's small fast model for memory and spend the
237
+ budget on the agent and the fleet.
238
+
239
+ Local co-residency still matters. The background model, the action model, their
240
+ KV caches, and parallel slots must all fit the target's available memory.
241
+
242
+ ## Operator setup
243
+
244
+ The shipped defaults are:
245
+
246
+ ```yaml
247
+ background:
248
+ target: null
249
+ model: null
250
+ thinkingLevel: off
251
+
252
+ memory:
253
+ intervention:
254
+ enabled: true
255
+ everyNTools: 10
256
+ windowSteps: 8
257
+ maxTokens: 400
258
+ timeoutMs: 180000
259
+ ```
260
+
261
+ With `background.target` and `background.model` unset, Clio stays in the
262
+ zero-cost rules tier. `/memory` shows the current tier, last decision, approved
263
+ durable lessons, the live bank, and a bounded history of the last twenty memory
264
+ steps with their trigger, decision, write count, cited-entry count, tier, and
265
+ latency. That history is the only place a capture, a gate, or a timeout becomes
266
+ visible, since those outcomes produce no transcript entry by design; only an
267
+ actual injection reaches the transcript. It carries counts and outcomes only,
268
+ never bank or trajectory text. `/settings` exposes every key above. In
269
+ `/targets`, select an eligible local target and press `b` to make it the saved
270
+ background-memory default; this is independent of `u` for chat and `f` for the
271
+ fleet default. A running session owns its routing snapshot, while the saved
272
+ selection becomes the default for new sessions.
273
+
274
+ The reference live configuration is an LM Studio server on the `zbook` node with
275
+ the wire model `qwopus3.5-9b-v3`:
276
+
277
+ ```yaml
278
+ background:
279
+ target: zbook
280
+ model: qwopus3.5-9b-v3
281
+ thinkingLevel: off
282
+ ```
283
+
284
+ A small model is the intended shape for this role. Across 60 measured steps on
285
+ that route, latency ran 4.4 to 220.8 seconds with a median of 18.6, a 90th
286
+ percentile of 79.9, and a 95th of 131.6. Capability is not the constraint;
287
+ latency is, its spread is wide, and the detached step above is what makes the
288
+ tier usable anyway.
289
+
290
+ Size `timeoutMs` off that tail rather than off the median. The shipped 180000
291
+ captures roughly the whole distribution on this route. A 20000 setting looks
292
+ generous against an 18.6-second median and in practice discarded about half of
293
+ all steps, since the step still runs to completion and only its result is thrown
294
+ away. A route whose steps mostly record `timeout` is a misconfigured deadline
295
+ before it is a slow model.
296
+
297
+ The target ID is not hard-coded. Any configured orchestrator-eligible local
298
+ target and wire model can fill the background role. Before enabling it, use the
299
+ real target surfaces to verify the route:
300
+
301
+ ```bash
302
+ clio-coder targets --probe
303
+ clio-coder models --target zbook
304
+ clio-coder
305
+ ```
306
+
307
+ Then inspect `/targets`, `/settings`, and `/memory`. Local co-residency still
308
+ matters: the background model, action model, their KV caches, and parallel slots
309
+ must fit the target's available memory. Increase `timeoutMs` for a deliberately
310
+ slow local route; lowering `maxTokens` bounds the visible reminder but does not
311
+ change the background model's strict output grammar.
312
+
313
+ For an immediate kill switch, set `memory.intervention.enabled` to `false` in
314
+ `/settings`. Removing the background target instead returns to rules-only
315
+ operation while leaving deterministic protection active.
316
+
317
+ ## What the LLM tier actually writes
318
+
319
+ Measured on the shipped prompt against `google/gemma-4-26b-a4b-qat`, across ten
320
+ live steps and forty controlled runs on the same route.
321
+
322
+ The tier writes `update_status` reliably and `save_knowledge` rarely, and that is
323
+ correct rather than broken. A trajectory step carries the tool name, a bounded
324
+ call description, an outcome, and a result digest. On success the digest is an
325
+ opaque result fingerprint, so a window of successful reads tells the model which
326
+ files were touched and nothing about what is in them. There is no durable fact in
327
+ that input, and a status line is the only faithful thing to write about it.
328
+
329
+ Three candidate causes were ruled out by controlled runs that changed one
330
+ variable at a time:
331
+
332
+ - rewriting the prompt's second worked example to carry a `save_knowledge` moved
333
+ nothing, and made the model emit no operations at all in four of five runs;
334
+ - seeding the bank with existing knowledge entries so the model could learn the
335
+ shape by example moved nothing;
336
+ - giving successful steps a content-bearing digest moved nothing on its own.
337
+
338
+ What does elicit knowledge is a durable fact in the input. With a task stating two
339
+ explicit constraints the model wrote both as knowledge; with that same task plus
340
+ content-bearing digests it wrote six. Error digests already carry a real
341
+ diagnostic line, which is why `save_procedural` fires on failing windows.
342
+
343
+ The consequence for reactivation is why the post-compaction block restores status.
344
+ Restoring knowledge alone restored nothing in the common case, because the one
345
+ class it read was usually the one class the model had not written.
346
+
347
+ Two further numbers from the same route. Roughly a quarter of live steps returned a
348
+ malformed envelope, usually `<operations>` with no list followed by
349
+ `<no_intervention/>`, which is recorded as `malformed`/`unparseable` and is
350
+ model behavior rather than a route fault. The boundary drop rate remains 0% at
351
+ shipped settings.
352
+
353
+ ## Handoff continuity
354
+
355
+ The bank normally dies with the session. When `context-handoff` is explicitly
356
+ requested, Clio supplies the skill a redacted `clio-task-memory` fenced snapshot
357
+ containing knowledge and procedural entries only. Ordinary turns receive no
358
+ snapshot. The handoff artifact remains under ignored `.clio-coder/handoffs/`; private
359
+ status and secret-shaped values do not cross the export boundary.
360
+
361
+ After `/resume`, Clio checks only the newest handoff and offers `/memory seed` if
362
+ it contains a valid snapshot. Seeding is explicit and deduplicated. It resets
363
+ injection attribution for the new session, and the master kill switch disables
364
+ both the offer and writes. `/new`, `/fork`, `/resume`, and ACP session changes
365
+ clear the prior heap bank before the new session can observe it.
366
+
367
+ ## Telemetry
368
+
369
+ Each completed memory step appends one content-free record to:
370
+
371
+ ```text
372
+ <stateDir>/memory/steps.jsonl
373
+ ```
374
+
375
+ Use `clio-coder paths --json` to resolve the state directory (the `"state"` property). The log rotates after 1 MiB and
376
+ keeps one previous generation as `steps.jsonl.1`. Every exact-schema record has:
377
+
378
+ - timestamp and schema version;
379
+ - one to three coalesced trigger reasons;
380
+ - `rules` or `llm` tier;
381
+ - per-class added, updated, and deleted entry counts;
382
+ - `silent`, `injected`, `gated`, `timeout`, `malformed`, or `dropped` decision;
383
+ - count of cited entries, input/output/total memory-model tokens, and latency.
384
+
385
+ `dropped` is the one outcome that ran no step: the boundary triggered while an
386
+ earlier step still held the single in-flight slot. It costs no tokens and no
387
+ latency, its triggers survive to the next free boundary, and it does not replace
388
+ the operator-visible last decision. Counting `dropped` rows against `llm` rows
389
+ over a session is how a starved cadence becomes visible.
390
+
391
+ The log contains no task, trajectory, bank, error, or reminder text. File creation,
392
+ rotation, serialization, and injected sinks are all best effort; a read-only
393
+ state directory or full disk cannot alter intervention behavior.
394
+
395
+ Note that routine no-tool continuation checks do not emit telemetry rows, as they
396
+ are not considered new memory boundaries. Only actual tool-bearing or explicitly
397
+ triggered memory steps produce rows.
398
+
399
+ ## Evaluation and promotion bar
400
+
401
+ `src/domains/eval/proactive-memory.ts` exports a fixed three-task, matched A/B
402
+ harness. It executes `baseline`, `rules`, and `llm` variants in stable order and
403
+ accepts any `{ id, model }` target. A runner adapter owns isolated task execution
404
+ and returns action tokens/latency plus the exact telemetry rows emitted for that
405
+ trial. The report provides:
406
+
407
+ - pass rate;
408
+ - injected and cited reminder counts;
409
+ - reminders per task and citation rate;
410
+ - total and baseline-relative added tokens and latency;
411
+ - an `alwaysNoisyRegression` verdict.
412
+
413
+ The deterministic end-to-end harness contract can be run directly:
414
+
415
+ ```bash
416
+ npm run test:file -- tests/contracts/proactive-memory-eval.test.ts
417
+ ```
418
+
419
+ For a live local comparison, an adapter should route only the `llm` variant
420
+ through the request's target/model (the reference is `zbook` /
421
+ `qwopus3.5-9b-v3`), keep baseline memory telemetry empty, and run all
422
+ nine trials in equivalent isolated workspaces. Do not promote the LLM tier from
423
+ one anecdotal task. The evidence bar is a pass-rate gain from a small number of
424
+ specific, usually cited reminders at acceptable added token and latency cost.
425
+ Injecting at least once per task while merely tying or losing to baseline is
426
+ always a regression, even when every reminder is cited.
427
+
428
+ ## Worker growth path
429
+
430
+ Worker-side intervention is intentionally not implemented in this sprint. The
431
+ bank, policy client, telemetry, and registration interfaces carry no interactive
432
+ chat-loop types, so they can be instantiated per worker later without moving the
433
+ policy into the action agent.
434
+
435
+ The existing transport already exposes the required seams:
436
+
437
+ 1. `src/domains/dispatch/worker-spawn.ts` receives worker NDJSON events and its
438
+ `SpawnedWorker.send` path can write bounded control messages while the worker
439
+ is alive.
440
+ 2. Worker steering already drains between tool batches, which is the safe point
441
+ for a visible memory advisory; it must not interrupt a tool in flight.
442
+ 3. Workers already maintain per-worker loop detectors and tool-call caps. A
443
+ future registration should consume those verdicts instead of re-deriving
444
+ them.
445
+ 4. Each dispatched run needs its own bank, cadence, spend guard, telemetry
446
+ attribution, and teardown. Parent session memory must not leak into sibling
447
+ workers implicitly.
448
+
449
+ The future sequence is therefore worker events → worker-local registration → one
450
+ bounded steering advisory between batches. It must preserve the current receipt,
451
+ safety, timeout, and permission semantics, and it should ship only after a
452
+ Terminal-Bench-style long-run evaluation shows a selective benefit.
@@ -0,0 +1,142 @@
1
+ # Prompt Envelope and Tools
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/tools_blueprint.html](html/tools_blueprint.html) (Version: 0.3.0).
5
+
6
+ Clio Coder keeps the model-facing envelope stable and moves enforcement into the runtime registry and safety policy.
7
+
8
+ Source of truth: `src/core/tool-names.ts`, `src/tools/agent-tools.ts`, `src/tools/bootstrap.ts`, `src/tools/policy.ts`, `src/tools/observation.ts`, `src/tools/ignore-policy.ts`, and the per-tool modules under `src/tools/**`.
9
+
10
+ ## One system prompt per session
11
+
12
+ The chat loop compiles one provider-facing system prompt for a session. The compile key is `target|model|autonomy|sessionId|workingContextPaths`, with the working-context paths sorted before hashing into the key.
13
+
14
+ The compiled prompt is reused byte-for-byte on ordinary submits. It recompiles only when that key changes or when config hot-reload invalidates the prompt cache. Path-scoped project rules can therefore recompile the prompt when a matching file enters working context. When recompilation changes the text, the session ledger records a `promptRecompiled` entry with the previous hash, new hash, and token estimate.
15
+
16
+ Prompt extensions can add dynamic fragments for project rules, the operator profile, and Clio source-tree awareness. Pending skill requests and middleware reminders are visible text in the user message, not hidden prompt machinery.
17
+
18
+ The Tool Contract section of the prompt renders a fixed set of base lines plus one optional guidance sentence per tool, sourced from the tool registry (`ToolMetadata.promptHint` in `src/tools/registry.ts`, assigned in `src/tools/bootstrap.ts`). The base lines cover the complete-surface rule, tool-free answering, orientation preferences, a deterministic routing order (structured observation before bash, task board for multi-step work, bounded dispatch with receipt synthesis, validation before final claims), failure recovery through `context(scope="docs")` instead of blind retries, and the skill-listing gate (skill-shaped tasks or explicit operator skill requests only). The chat loop derives the hint list once from the session's frozen tool surface at compile time, and the compiler renders the hints sorted by tool name, so the compiled text depends only on which hinted tools are on the surface. Today five tools carry hints: `ask_user`, `code_nav`, `context`, `dispatch`, and `tasks`. Removing a tool from the surface removes its hint with no compiler change; adding a hint to a tool is a deliberate prompt-text change that must land with updated prompt contract tests and a CHANGELOG note.
19
+
20
+ ## One tool surface per session
21
+
22
+ For tool-capable providers, Clio sends the full registry as the session tool surface. The list is deterministic and sorted through the worker-tool resolver (`resolveAgentTools` in `src/tools/agent-tools.ts`), so the serialized schemas stay byte-identical on every submit. `src/tools/agent-tools.ts` is the single agent-tool adapter across the codebase. Both the orchestrator session and worker subprocesses resolve their tool set through the same `effectiveToolNames` narrowing function, ensuring that the attested signature and runtime surface cannot diverge.
23
+
24
+ Tools are keyed strictly by the canonical `ToolName` union defined in `src/core/tool-names.ts` with no alias table. Pure and idempotent `prepareArguments` normalizers defined on `ToolSpec` serve as the sole leniency layer for coercing legacy or weak-model parameter formats.
25
+
26
+ Tool visibility is not a per-turn hinting system. Pending-skill policy, ask-user policy, Bash policy, path policy, protected artifacts, dispatch admission, middleware, and the autonomy mapping are enforced when a tool is invoked. The `autonomy` level is applied at registry admission after the safety net passes a call; the safety prompt fragment mirrors that enforced matrix as guidance to the model. Prompt text and provider schemas do not bypass the registry.
27
+
28
+ Providers that cannot call tools receive no schemas, and the prompt tells the model to proceed without tool calls.
29
+
30
+ ## Canonical worker harness
31
+
32
+ Native and mediated dispatch workers use a separate prompts-domain compiler over the same loaded fragment table. Its stable system prompt has exactly five sections: identity-lite, the shared operating contract plus assigned-task rules, a tool contract sliced to the final canonical toolkit, safety for the single effective autonomy, and one final persona. A request persona override replaces only the recipe body; eligible bound-skill instructions are composed inside that same final persona and never widen tools.
33
+
34
+ The compiler runs after target capability and tool-profile admission. Its canonical tool names are the same names transported in `WorkerSpec.allowedTools` and attached as schemas; routine non-Scout work removes `code_nav`, narrow profiles remove their excluded schemas and guidance, tool-incapable targets get an explicit no-tools contract, and Claude SDK aliases are filtered from the same canonical set. ACP's external inventory is unknown, so ACP bounded-role admission continues to validate the unchanged raw persona rather than fabricating a complete native schema list.
35
+
36
+ Project context, memory, bounded dispatch briefing, pipeline input, the assigned task, and the per-run safety-posture reminder remain dynamic user messages. A briefing is a separately delimited message labeled as untrusted task context/data; it is never concatenated into the task or stable system prompt. Dynamic ordering is project, safety, memory, briefing, then pipeline input, with pipeline input last. These messages do not affect the stable composition hash. Persona, effective autonomy, target tool capability, or final toolkit changes do affect it.
37
+
38
+ ## Seven planes, twenty tools
39
+
40
+ The builtin surface is 20 registered tools organized in seven planes. Each plane is one policy unit: its tools share an action class, a size posture, a details schema, and a concurrency rule. `src/tools/policy.ts` asserts these invariants at bootstrap, so drift between the plane design, the safety classifier, and the registered specs fails loudly instead of shipping a surface that behaves differently from what the policy engine assumes.
41
+
42
+ | Plane | Tools | Action class | Concurrency |
43
+ | --- | --- | --- | --- |
44
+ | OBSERVE | `read`, `grep`, `find`, `ls`, `code_nav`, `context`, `credential_present` | read | parallel |
45
+ | MUTATE | `write`, `edit` | write | sequential |
46
+ | EXECUTE | `bash`, `verify` | execute | sequential |
47
+ | EXECUTE | `git` | read | parallel |
48
+ | ORCHESTRATE | `dispatch`, `steer` | dispatch | sequential |
49
+ | ORCHESTRATE | `monitor` | read | parallel |
50
+ | ORCHESTRATE | `tasks` | read | sequential |
51
+ | ORCHESTRATE | `ledger` | read | sequential |
52
+ | RETRIEVE | `web_fetch` | read | parallel |
53
+ | INTERACT | `ask_user` | read | sequential |
54
+ | ARTIFACT | `artifact` | write | sequential |
55
+
56
+ Three tools sit in a plane for containment rather than class. `git` is read-only inspection (op=status/diff/log) that runs on the safe-exec spine, so it lives in the EXECUTE plane with read-class safety disposition. `monitor` never mutates a run, so it stays read class and parallel inside the ORCHESTRATE plane. `tasks` orchestrates the agent's own work rather than workers: it mutates only the session's task ledger, never the workspace, so it keeps read class (never gated behind a confirmation) but runs sequential so two board mutations in one batch cannot interleave. `ledger` is the agent ledger, the coordination board concurrent dispatch workers share: a post reaches a one-way control lane and a read answers from a local mirror, so it touches no workspace and stays read class, and reviewers and judges are pinned to read-only autonomy where a write class would block the peer review the board exists for.
57
+
58
+ Registration is conditional on wiring: `context` gains its workspace scope only when a session contract is bound, `dispatch`/`monitor`/`steer` register only with a dispatch contract, and `ask_user` registers only when an interactive handler exists. Dispatch tool profiles narrow the surface for workers: `minimal-local` is `read`, `grep`, `find`, `ls`, `git`, `context`, `code_nav`; `science-local` adds `verify`; `full-agent` keeps everything.
59
+
60
+ ### Consolidated call shapes
61
+
62
+ Several tools absorb what used to be separate tools:
63
+
64
+ - `find(pattern, path?, order?, limit?, include_ignored?)` locates paths by glob pattern (`*`, `**`, `?`, `[abc]`), default limit 500. `order="path"` (default) returns fd's native order; `order="mtime"` returns newest first from a bounded candidate set instead of statting the whole tree, and reports `details.candidates` when the candidate cap made the ordering approximate.
65
+ - `grep(pattern, path?, mode?, glob?, ignore_case?, literal?, context?, limit?, include_ignored?)` searches file contents with ripgrep, degrading to a bounded pure-Node search when rg is absent. `mode=content` (default) returns line-referenced matches, `mode=files` returns matching paths, `mode=count` returns per-file counts. Context lines are consumed from rg's `--json` stream.
66
+ - `context(scope="workspace"|"docs"|"skills")` is the one OBSERVE entry point for material about the working environment: the session workspace snapshot, retrieval over Clio's bundled documentation (`query` required), and skill listing or loading (`name` optional, `include_tree` for the skill's resource files).
67
+ - `verify(check?, path?, args?, browser?, cwd?, timeout_ms?)` runs declared verification. `verify()` with no arguments lists declared checks grouped by source (package.json verification scripts today), `verify(check="<script>")` runs one through the safe-exec spine with no shell, and `verify(check="frontend", path=...)` validates an HTML/CSS/JS artifact without granting shell access.
68
+ - `artifact(kind="plan"|"review"|"report", content, ...)` writes named artifacts behind one surface: Markdown documents (default `PLAN.md`/`REVIEW.md`/`REPORT.md` at the project root; `path` may override inside the workspace) that terminate the turn, because writing the artifact is the answer. Skills are not artifacts; a `SKILL.md` is written with the ordinary write tool and validated by the skills loader.
69
+ - `dispatch(task?, tasks?, mode?, ...)` supports a first-class singular assignment (`task`) and a batch (`tasks`), never both. `task` is worker instructions; `briefing` is optional bounded parent context/data and cannot replace it. Briefing stays a separate dynamic message and receipt provenance, never part of the receipt task. A shared top-level briefing applies to strings and objects without an override; an object-level briefing wins. Blank values are omitted, the cap is 12,000 UTF-8 bytes, and approval pins the exact canonical value. Ordinary handles enter one registered event consumer immediately. Synchronous calls auto-wait for stream-and-receipt completion; `detach:true` returns ids after durable batch registration while the same consumer continues. Review and compete retain gate-sensitive direct drains. Task objects may include `persona` and `tool_profile`. Pipeline output is threaded as bounded data. A successful native or ACP run requires a nonempty receipt-sealed final output; exit zero without one fails as `worker_final_output_missing`, with unfinished text retained only as partial diagnostics. `dispatch(list=true)` renders the catalog.
70
+ - `monitor(run_id?, mode?)` is read-only visibility into known synchronous and detached runs: `list` enumerates, `status` reports one, `peek` returns the in-process event tail, `receipt` exposes the stored evidence, and `wait` observes one run without collecting or canceling it. `collect` is the authoritative terminal batch operation over a detached batch or run-id list; collect before final synthesis. Completed output reports receipt integrity, evidence verification, briefing provenance, and bounded project-context provenance as different fields.
71
+ - `steer(run_id, action, message?)` controls a running worker: `guide` writes a canonical trimmed steering message to an HTTP or SDK worker and `cancel` terminates it. Successfully written steers gain ordered byte/hash/timestamp provenance; after the runtime accepts the guidance, `clio_steer_received` acknowledges the exact matching sequence, and prose is never stored in ledger or receipt. Single-shot subprocess runtimes and ACP remain non-steerable. Interactive operators can steer synchronous live-input runs; parent-model steering requires detached ids because model tools are sequential.
72
+
73
+ ### One ignore policy for path walkers
74
+
75
+ `grep`, `find`, and their pure-Node fallbacks answer "which parts of the tree are visible" from one shared policy in `src/tools/ignore-policy.ts`. Three layers apply: `.clio-coder`, `.fallow`, and `.git` are always excluded; `.gitignore` is honored natively by rg/fd; and one generated-dirs list (`node_modules`, `dist`, `build`, `coverage`, `.venv`, and similar) is force-excluded even when a project forgot to gitignore it. `include_ignored: true` lifts the gitignore and generated-dirs layers together. The clio-internal layer always stands, except that pointing a tool directly at one of those directories means the caller wants those paths.
76
+
77
+ ## The observation envelope
78
+
79
+ The six content-returning OBSERVE tools (`read`, `grep`, `find`, `ls`, `code_nav`, `context`) close every result through one shared envelope in `src/tools/observation.ts`. `credential_present` sits in the OBSERVE plane but returns a typed boolean and carries no envelope cap. The envelope owns four guarantees.
80
+
81
+ **One notice line, one format.** A truncated text result appends exactly one notice:
82
+
83
+ ```text
84
+ [<tool>: <shown>/<total> <unit> shown (<shownSize> of <totalSize>) | full: <offloadPath> | next: <exact-call>]
85
+ ```
86
+
87
+ Unknown segments are omitted. `<total>` renders as `N+` when the search was killed early at its limit, meaning matches beyond it exist but were never counted. `next` is always an exact continuation call fragment such as `limit=200` or `offset=451`, never prose. Untruncated results get no notice. Empty results are standardized: `grep` returns `No matches found`, `find` returns `No files found matching pattern`, `ls` returns `(empty directory)`, and the JSON-format tools return valid JSON with empty arrays and `next` populated.
88
+
89
+ **Offload on truncation.** When a byte cap cuts collected content, the tool spills its full rendering to the per-session scratch file (`<stateDir>/scratch/<sessionId>/<toolCallId>.txt`) and reports the path in the notice, so no collected match, path, or line is ever unrecoverable. Two deliberate exceptions exist: `read` never offloads because the source file is directly re-addressable via `next: offset=N`, and a bare item-limit truncation without a byte cut continues via `next` alone, since an offload would only duplicate the body.
90
+
91
+ **Always-valid JSON.** `code_nav` and the JSON scopes of `context` declare `format: "json"`. A JSON payload must parse or be replaced whole; it is never cut mid-document. An oversize payload is offloaded and the body is replaced by the parseable stub:
92
+
93
+ ```json
94
+ {"error":"result exceeded <cap>","offloadPath":"...","next":"..."}
95
+ ```
96
+
97
+ **One turn budget.** All six envelope tools draw from a single per-turn pool keyed `sessionId:turnId`, default 192KB, overridable with `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES`. Each call reserves the minimum of its self cap and the remaining budget before doing the work. An exhausted pool short-circuits with an `[observation budget exhausted ...]` notice naming the tool, the subject, and the used/limit sizes, instead of paying for a search whose output could not be returned. A call whose cap was reduced by the pool appends a budget note telling the model to narrow its arguments or continue in a follow-up turn.
98
+
99
+ Per-call self caps: `read` 50KB (`CLIO_CODER_READ_MAX_BYTES`), `grep` 16KB for `mode=content` and 8KB for `files`/`count`, `find` 8KB, `ls` 8KB, `code_nav` 16KB, `context` 16KB for docs and 50KB for skills/workspace. The registry backstop cap for each envelope tool is its self cap plus 2KB slack, so a tool's own notice with its exact continuation call survives shaping instead of being cut again and replaced by a generic hint; the bootstrap policy assertion fails loudly if a cap ever drops below that.
100
+
101
+ Every envelope result carries `details.observation` (`{tool, unit, shownCount, totalCount, shownBytes, totalBytes, truncated, format, next?, offloadPath?, budget?}`) for the TUI ledger, session turns, and observers.
102
+
103
+ ## Description tiering
104
+
105
+ Tool descriptions are tiered by how much a wrong call costs. The hot tools the model calls constantly (`read`, `grep`, `find`, `dispatch`) embed their operational contract in the description: caps, modes, ignore semantics, and how truncated results continue. Every other tool carries a one-to-two-sentence statement of what it does, and deep usage guidance lives in the bundled docs corpus ([tool-usage.md](tool-usage.md)) rather than the prompt prefix, retrievable on demand through `context(scope="docs")`. This keeps the serialized schema block small and byte-stable while still giving the model a path to depth when it needs one.
106
+
107
+ ## The gateway reservation
108
+
109
+ `gateway` is a design-reserved name in `src/core/tool-names.ts`, not an implemented tool. The reserved contract sketch is `gateway(op: "find" | "describe" | "call", capability?, args?)`: an MCP/database proxy with one fixed schema, where external capabilities surface through find/describe/call results rather than as per-capability schemas in the prompt prefix. It would carry the network action class and run sequentially. Reserving the name keeps classifiers and profiles from ever assigning `gateway` to a dynamic tool.
110
+
111
+ ## Context protection
112
+
113
+ Clio uses two context-protection mechanisms.
114
+
115
+ 1. Tool results are capped at the source and again at the registry boundary. OBSERVE tools use the envelope caps above. Exact mutation tools (`write`, `edit`, `artifact`) use 8KB; `steer` and `credential_present` use 4KB; `ask_user` has a 20KB policy. Summary-kind tools (`bash`, `git`, `verify`, `dispatch`, `monitor`) use 16KB at the registry boundary. `web_fetch` is bounded at 16KB after shaping and may read more before it: its `max_bytes` argument defaults to 600KB and is hard-capped at 5MB. Tools without an explicit result-size policy use an approximately 18KB generic backstop. Over-cap generic results are shown briefly and, when possible, saved under `<stateDir>/scratch/<sessionId>/<toolCallId>.txt` with an `offloadPath` detail and a 10MB scratch-file cap.
116
+ 2. Auto-compaction uses one pressure threshold. The default threshold is 0.8. When pressure crosses the threshold, Clio first masks stale tool observations and stale thinking older than `excludeLastTurns`. If pressure remains above the threshold, it runs the LLM summary compaction path and replays from the compacted session view.
117
+
118
+ Manual `/context compact`, `CLIO_CODER_FORCE_COMPACT=1`, and overflow recovery force the LLM summary path directly.
119
+
120
+ Compaction rewrites history, so the next turn on a local single-slot backend is expected to lose prefix-cache alignment. Clio records `expectedColdReasons` and shows one dim notice for that turn.
121
+
122
+ ## Inspecting a session
123
+
124
+ Timing and cache behavior are persisted per API call, so a finished session can be inspected from its stored artifacts alone. Each assistant entry in the session ledger (`current.jsonl`, under the directory reported by `clio-coder paths`) carries `timing { ttftMs, apiMs }` and `promptCache { input, cacheRead, cacheWrite, backendVerdict }`, and the run's first persisted call also carries `expectedColdReasons`. Cache verdicts are `hot`, `partial`, `cold`, or `small`.
125
+
126
+ For aggregate cost and token facts across sessions, use `clio-coder usage report --days <n>`. Inside the TUI, `/cost` shows session totals and `/context` opens the context-window ledger overlay.
127
+
128
+ ## Self-documentation retrieval
129
+
130
+ `context(scope="docs")` is the model-facing companion to the human `clio-coder docs` server. The server serves bundled `docs/html/**` blueprints for people; the docs scope indexes the bundled markdown corpus for agents. It is deterministic and offline: no embeddings service, network call, or filesystem write is needed.
131
+
132
+ The search index splits markdown into heading-delimited sections, records heading breadcrumbs and line ranges, and ranks results with light stemming, controlled Clio vocabulary aliases, phrase boosts, and BM25-style body scoring. The tool returns compact JSON containing corpus metadata, normalized and expanded query terms, and ranked hits with `file`, `heading`, `breadcrumb`, `anchor`, section `lines`, `snippetLines`, a bounded `snippet`, `matchedTerms`, `signals`, `coverage`, and `score`. `limit` defaults to 5 sections and caps at 12. The per-file filter the pre-consolidation docs tool accepted was dropped; narrow with more specific query terms instead. Even an empty result is valid JSON with empty arrays and a populated `next` continuation.
133
+
134
+ ## Edit matching safety
135
+
136
+ The `edit` tool first attempts exact matching. If the model's old text differs
137
+ only by normalized quote, dash, whitespace, or indentation details, Clio maps
138
+ the normalized match back to the original line span and splices only the
139
+ intended replacement. Unchanged spans keep their original bytes, including
140
+ smart punctuation and CRLF line endings. Ambiguous duplicate matches,
141
+ overlapping hunks, empty changes, and no-op edits are rejected instead of
142
+ guessing.