@iowarp/clio-coder 0.3.4 → 0.3.6

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (265) hide show
  1. package/CHANGELOG.md +37 -2
  2. package/CONTRIBUTING.md +6 -6
  3. package/README.md +2 -2
  4. package/dist/{acp-S5R4RR5B.js → acp-2BEHC4DL.js} +4 -4
  5. package/dist/{agents-P6DMMVZY.js → agents-LNNFTM53.js} +13 -11
  6. package/dist/assets/codewiki.json +1 -1
  7. package/dist/{auth-2XCZLPKS.js → auth-KXXFI2VS.js} +6 -6
  8. package/dist/{chunk-YCWGATWI.js → chunk-24I7BN55.js} +2 -2
  9. package/dist/{chunk-EKMEHE4H.js → chunk-33YXPOE3.js} +2 -3
  10. package/dist/chunk-3BPUFZDL.js +37 -0
  11. package/dist/{chunk-WPQLXFOZ.js → chunk-43AOLP7E.js} +2 -2
  12. package/dist/{chunk-N4CZJQRK.js → chunk-5JGRAMKL.js} +4 -4
  13. package/dist/{chunk-BRXQQJFP.js → chunk-6US73PDB.js} +568 -47
  14. package/dist/{chunk-K6WL7QZT.js → chunk-6XXKFVSN.js} +2 -2
  15. package/dist/{chunk-QQK64KLB.js → chunk-CJUB2JJ2.js} +138 -20
  16. package/dist/{chunk-HV5X7OR2.js → chunk-CKXWIANG.js} +12 -12
  17. package/dist/{chunk-UZHIZC5S.js → chunk-CYQKWTG3.js} +61 -76
  18. package/dist/{chunk-QWU7ZBO7.js → chunk-DJVECN66.js} +204 -45
  19. package/dist/{chunk-ZWMF7253.js → chunk-E2ER4LJF.js} +304 -9
  20. package/dist/{chunk-7RXG6QRZ.js → chunk-EKY57CSP.js} +2 -75
  21. package/dist/{chunk-EDRHSCIE.js → chunk-EYPA3EGJ.js} +10 -2
  22. package/dist/{chunk-TTNYS3EA.js → chunk-G7MUEIGA.js} +1 -1
  23. package/dist/{chunk-BPGS2WCQ.js → chunk-GEYXPTRF.js} +2 -1
  24. package/dist/{chunk-BEY543CS.js → chunk-GOXNB3AO.js} +5 -2
  25. package/dist/{chunk-G4BMMOKF.js → chunk-HVDIIIQW.js} +2 -2
  26. package/dist/chunk-HWUFFB6L.js +83 -0
  27. package/dist/{chunk-35MKKU5R.js → chunk-K7T3E2SR.js} +15 -8
  28. package/dist/{chunk-VAWWTKDP.js → chunk-KHSFENX2.js} +2 -2
  29. package/dist/chunk-LCGCVYZ4.js +57 -0
  30. package/dist/{chunk-X6COSD2O.js → chunk-LYF7OHWH.js} +41 -14
  31. package/dist/{chunk-POHLU5DW.js → chunk-M6L6IDJG.js} +3 -3
  32. package/dist/{chunk-X4RCMKVQ.js → chunk-NDINPTJ4.js} +2 -2
  33. package/dist/{chunk-5M54SPOL.js → chunk-ODFEOB4F.js} +161 -5
  34. package/dist/{chunk-3JLKSKD7.js → chunk-OH3TOQTB.js} +5 -1
  35. package/dist/{chunk-MEQ45TQ4.js → chunk-PBTHKCPN.js} +18 -4
  36. package/dist/{chunk-ED4KHGC3.js → chunk-PPAMZ32Z.js} +9 -2
  37. package/dist/{chunk-QQL5RT5M.js → chunk-QM3F2GKX.js} +94 -36
  38. package/dist/{chunk-A2GZF7DC.js → chunk-QNQHSOLF.js} +4 -4
  39. package/dist/{chunk-KRPY7NTG.js → chunk-R46L2BIR.js} +3 -3
  40. package/dist/{chunk-BP4OYD6A.js → chunk-RY3LY4J5.js} +20 -2
  41. package/dist/{chunk-34475P3I.js → chunk-TSHXZTOQ.js} +5 -4
  42. package/dist/{chunk-VJWL6YS5.js → chunk-UUVG37B4.js} +2 -2
  43. package/dist/{chunk-2TZWSW76.js → chunk-WHGPSPT5.js} +2 -2
  44. package/dist/{chunk-TW3WDMVS.js → chunk-WHJYKASB.js} +2 -2
  45. package/dist/{chunk-YHZX5GEU.js → chunk-XAKHZX5N.js} +2 -2
  46. package/dist/{chunk-HXG4IURW.js → chunk-XE2VEJHX.js} +2 -2
  47. package/dist/{chunk-3HZ5RWN2.js → chunk-XF5N4U5A.js} +7 -6
  48. package/dist/{chunk-ZYKPLLNQ.js → chunk-XXQNGV4M.js} +590 -32
  49. package/dist/{chunk-4JUF2NNX.js → chunk-XYDYPRZI.js} +4 -4
  50. package/dist/{chunk-VMNQ6OZA.js → chunk-ZRGEBJ4T.js} +971 -794
  51. package/dist/{chunk-2LZI5CAG.js → chunk-ZXF4XRKW.js} +75 -33
  52. package/dist/{chunk-VSNATDE6.js → chunk-ZZMN5OM4.js} +2 -2
  53. package/dist/cli/index.js +31 -31
  54. package/dist/{clio-J5JIOIDS.js → clio-M2KGYUFZ.js} +2 -2
  55. package/dist/{code-nav-AXCXSBHX.js → code-nav-GQNL7XA6.js} +5 -5
  56. package/dist/codewiki/build-worker.js +4 -4
  57. package/dist/{components-KELWS457.js → components-5TTYYX6G.js} +3 -3
  58. package/dist/{config-OEBMIN2U.js → config-XUUYQIWO.js} +27 -25
  59. package/dist/{configure-PUQOSIXQ.js → configure-IHJ7YOMV.js} +7 -7
  60. package/dist/{context-URSXPBCK.js → context-74JLXAWD.js} +12 -12
  61. package/dist/{context-MGSE4Z2T.js → context-75MIWW3U.js} +24 -22
  62. package/dist/{context-EKDCKUUZ.js → context-ZQ7SIFJV.js} +8 -7
  63. package/dist/{context-clear-KDAJRNUK.js → context-clear-GYKWNUML.js} +24 -22
  64. package/dist/{context-index-BZ4UYMTC.js → context-index-SSR5ECNE.js} +3 -3
  65. package/dist/{context-working-set-SBKMPPI2.js → context-working-set-UX5KEP4J.js} +11 -10
  66. package/dist/{dispatch-runner-MSWN72NK.js → dispatch-runner-GIJBHNFL.js} +21 -20
  67. package/dist/{docs-2C2LTVT2.js → docs-6FZSCG5B.js} +3 -3
  68. package/dist/{doctor-7BSE27PJ.js → doctor-SVJ5BZCW.js} +4 -4
  69. package/dist/{eval-IZGDOO4H.js → eval-CG6LLBLD.js} +47 -232
  70. package/dist/{evidence-SR7WXB5B.js → evidence-ZYFIEN42.js} +19 -18
  71. package/dist/{evolve-K7VE2CBX.js → evolve-QGEXEMDW.js} +19 -18
  72. package/dist/{extensions-QVDOHDGJ.js → extensions-ADGNCJJD.js} +3 -3
  73. package/dist/{fleet-7XMJNQNF.js → fleet-S5R4ZOQY.js} +49 -30
  74. package/dist/{fleet-preflight-AQNAH644.js → fleet-preflight-BHSNPBMH.js} +2 -2
  75. package/dist/{init-JGNPAYXT.js → init-5DRU55YR.js} +31 -29
  76. package/dist/memory-7YKKR6UC.js +467 -0
  77. package/dist/{models-ZMMLFJNN.js → models-ZPOLRU2C.js} +10 -10
  78. package/dist/{monitor-2F3T5KHP.js → monitor-US5F5YGZ.js} +33 -18
  79. package/dist/{orchestrator-ORHT43JB.js → orchestrator-E2AL4T5N.js} +1092 -659
  80. package/dist/{paths-UXLN5YYZ.js → paths-E7KYAQWE.js} +3 -3
  81. package/dist/{reset-NXGTYNUO.js → reset-KZ652EK6.js} +3 -3
  82. package/dist/{run-RF4WJGMT.js → run-SRNBKDWD.js} +52 -40
  83. package/dist/{share-UT3W6E4M.js → share-CGZE33UP.js} +3 -3
  84. package/dist/{skills-PSACKC5Q.js → skills-S2X4DLY5.js} +4 -4
  85. package/dist/{skills-eval-WJSI55RZ.js → skills-eval-W2GGIC4R.js} +19 -18
  86. package/dist/{targets-PIIRAOYS.js → targets-54SWINWB.js} +14 -12
  87. package/dist/{terminal-lease-ULWXWNVY.js → terminal-lease-SAIF2OGY.js} +5 -4
  88. package/dist/{uninstall-FZCQCDKC.js → uninstall-BVLWXKBT.js} +3 -3
  89. package/dist/{upgrade-346TZ6AV.js → upgrade-JKAR27XC.js} +8 -8
  90. package/dist/{usage-6KKXR32N.js → usage-MSAWCLX4.js} +60 -27
  91. package/dist/{verifiers-4UUM6TEE.js → verifiers-NCBTHHN2.js} +60 -54
  92. package/dist/{wiki-generate-7STOCIFZ.js → wiki-generate-GUSOQ6ZP.js} +30 -28
  93. package/dist/worker/entry.js +69 -58
  94. package/dist/{workspace-G4ZWUIPR.js → workspace-ZJ6BFM3Q.js} +4 -4
  95. package/docs/README.md +3 -3
  96. package/docs/acp.md +1 -1
  97. package/docs/alcf-provider.md +1 -1
  98. package/docs/architecture.md +2 -2
  99. package/docs/artifact-placement.md +1 -2
  100. package/docs/artifact-versions.md +1 -1
  101. package/docs/built-in-agents.md +1 -1
  102. package/docs/capacity-and-scheduling.md +1 -1
  103. package/docs/commands-and-modes.md +9 -7
  104. package/docs/configuration-and-targets.md +12 -1
  105. package/docs/context-engine.md +4 -2
  106. package/docs/context-working-set.md +4 -4
  107. package/docs/development-pipeline.md +1 -1
  108. package/docs/documentation-coverage.md +3 -3
  109. package/docs/documentation-guide.md +2 -2
  110. package/docs/eval-runner.md +1 -1
  111. package/docs/evals-internal.md +4 -45
  112. package/docs/evidence-and-memory.md +67 -7
  113. package/docs/evolution.md +1 -1
  114. package/docs/exit-codes-and-output.md +1 -1
  115. package/docs/extensions-and-sharing.md +2 -2
  116. package/docs/fleet-dispatch.md +28 -2
  117. package/docs/installation-and-lifecycle.md +2 -2
  118. package/docs/middleware-and-components.md +19 -2
  119. package/docs/model-catalog.md +1 -1
  120. package/docs/observability.md +3 -3
  121. package/docs/proactive-memory.md +26 -16
  122. package/docs/prompt-envelope-and-tools.md +4 -2
  123. package/docs/provider-adapter-cookbook.md +1 -1
  124. package/docs/release-cut-checklist.md +38 -35
  125. package/docs/safety-model.md +29 -7
  126. package/docs/scientific-validation.md +3 -3
  127. package/docs/session-lifecycle.md +1 -1
  128. package/docs/skills-marketplace.md +1 -1
  129. package/docs/tool-usage.md +2 -2
  130. package/docs/trace-store.md +1 -1
  131. package/docs/troubleshooting.md +1 -1
  132. package/docs/tui-design.md +38 -4
  133. package/docs/worker-dispatch-mechanics.md +1 -1
  134. package/package.json +7 -4
  135. package/src/cli/agents.ts +2 -3
  136. package/src/cli/argv.ts +14 -1
  137. package/src/cli/fleet.ts +15 -0
  138. package/src/cli/index.ts +1 -1
  139. package/src/cli/memory.ts +272 -10
  140. package/src/cli/modes/json-stream.ts +2 -2
  141. package/src/cli/modes/print.ts +12 -1
  142. package/src/cli/run.ts +22 -2
  143. package/src/cli/targets.ts +12 -3
  144. package/src/cli/usage.ts +55 -7
  145. package/src/core/bus-events.ts +3 -0
  146. package/src/core/response-model-id.ts +134 -0
  147. package/src/core/toml.ts +62 -0
  148. package/src/core/workspace-files.ts +0 -1
  149. package/src/domains/agents/builtins/architect.md +1 -1
  150. package/src/domains/agents/catalog.ts +5 -4
  151. package/src/domains/agents/recipe.ts +54 -14
  152. package/src/domains/agents/result-contract.ts +7 -4
  153. package/src/domains/context/bootstrap.ts +36 -27
  154. package/src/domains/context/project-metadata.ts +19 -63
  155. package/src/domains/context/prompt-context.ts +8 -0
  156. package/src/domains/context/working-set/policies/index.ts +3 -4
  157. package/src/domains/dispatch/budget-envelope.ts +396 -0
  158. package/src/domains/dispatch/contract.ts +2 -0
  159. package/src/domains/dispatch/extension.ts +81 -27
  160. package/src/domains/dispatch/orphan-recovery.ts +1 -0
  161. package/src/domains/dispatch/receipt-integrity.ts +4 -0
  162. package/src/domains/dispatch/state.ts +1 -0
  163. package/src/domains/dispatch/types.ts +10 -3
  164. package/src/domains/dispatch/validation.ts +14 -0
  165. package/src/domains/dispatch/worker-spawn.ts +14 -3
  166. package/src/domains/eval/metrics/evidence.ts +0 -116
  167. package/src/domains/eval/metrics/invariants.ts +1 -1
  168. package/src/domains/eval/runners/clio-run.ts +1 -10
  169. package/src/domains/eval/runners/external-command.ts +2 -29
  170. package/src/domains/eval/schema/suite.ts +0 -7
  171. package/src/domains/eval/suites/run.ts +1 -7
  172. package/src/domains/memory/index.ts +22 -0
  173. package/src/domains/memory/operations.ts +58 -1
  174. package/src/domains/memory/promotion.ts +281 -0
  175. package/src/domains/memory/prompt-section.ts +25 -5
  176. package/src/domains/memory/proposal.ts +51 -7
  177. package/src/domains/memory/task-bank.ts +3 -2
  178. package/src/domains/memory/task-memory-handoff.ts +181 -24
  179. package/src/domains/memory/task-memory-policy.ts +3 -1
  180. package/src/domains/memory/types.ts +37 -0
  181. package/src/domains/memory/validate.ts +178 -0
  182. package/src/domains/middleware/memory-intervention.ts +35 -25
  183. package/src/domains/middleware/runtime.ts +6 -0
  184. package/src/domains/middleware/skills-reminder.ts +19 -4
  185. package/src/domains/middleware/stalled-turn.ts +43 -1
  186. package/src/domains/middleware/types.ts +10 -0
  187. package/src/domains/observability/contract.ts +6 -1
  188. package/src/domains/observability/cost.ts +20 -4
  189. package/src/domains/observability/extension.ts +2 -2
  190. package/src/domains/providers/index.ts +3 -0
  191. package/src/domains/providers/model-discovery.ts +9 -0
  192. package/src/domains/providers/runtime-resolution.ts +38 -1
  193. package/src/domains/providers/runtimes/common/probe-helpers.ts +97 -16
  194. package/src/domains/providers/types/context-window-slots.ts +18 -0
  195. package/src/domains/providers/types/runtime-descriptor.ts +3 -1
  196. package/src/domains/safety/call-target.ts +211 -14
  197. package/src/domains/safety/decision-presentation.ts +268 -0
  198. package/src/domains/safety/redaction.ts +73 -0
  199. package/src/domains/session/context-ledger.ts +10 -1
  200. package/src/domains/session/decision-board.ts +4 -0
  201. package/src/domains/session/entries.ts +3 -0
  202. package/src/domains/session/history.ts +68 -19
  203. package/src/domains/session/usage.ts +24 -7
  204. package/src/engine/acp/event-mapper.ts +7 -0
  205. package/src/engine/acp/server.ts +29 -2
  206. package/src/engine/apis/lmstudio.ts +25 -4
  207. package/src/engine/apis/openai-completions.ts +147 -22
  208. package/src/engine/claude/sdk-runtime.ts +8 -2
  209. package/src/engine/claude/tool-safety.ts +13 -0
  210. package/src/engine/loop-guard.ts +27 -3
  211. package/src/engine/worker-events.ts +4 -3
  212. package/src/engine/worker-runtime.ts +59 -54
  213. package/src/entry/orchestrator.ts +18 -1
  214. package/src/interactive/chat-loop-messages.ts +22 -0
  215. package/src/interactive/chat-loop.ts +13 -0
  216. package/src/interactive/chat-renderer.ts +19 -3
  217. package/src/interactive/clio-editor.ts +44 -7
  218. package/src/interactive/context-overlay.ts +43 -5
  219. package/src/interactive/cost-overlay.ts +39 -8
  220. package/src/interactive/dispatch-board.ts +212 -35
  221. package/src/interactive/footer/widgets.ts +13 -0
  222. package/src/interactive/interactive-application.ts +6 -1
  223. package/src/interactive/interactive-input-runtime.ts +11 -1
  224. package/src/interactive/interactive-presentation.ts +11 -1
  225. package/src/interactive/memory-overlay.ts +89 -4
  226. package/src/interactive/overlay-ask-user-lifecycle.ts +1 -1
  227. package/src/interactive/overlay-frame.ts +5 -2
  228. package/src/interactive/overlay-general-openers.ts +40 -1
  229. package/src/interactive/overlay-key-routing.ts +41 -1
  230. package/src/interactive/overlay-lifecycle.ts +11 -4
  231. package/src/interactive/overlay-permission-lifecycle.ts +23 -8
  232. package/src/interactive/overlay-transitions.ts +11 -0
  233. package/src/interactive/overlays/ask-user.ts +74 -30
  234. package/src/interactive/overlays/decisions.ts +3 -1
  235. package/src/interactive/permission-hint.ts +35 -0
  236. package/src/interactive/permission-overlay.ts +95 -45
  237. package/src/interactive/renderers/tool-execution.ts +19 -49
  238. package/src/interactive/session-last-turn.ts +8 -1
  239. package/src/interactive/session-usage-reseed.ts +36 -10
  240. package/src/interactive/slash-commands.ts +2 -2
  241. package/src/interactive/status/summary.ts +5 -0
  242. package/src/interactive/status/types.ts +5 -0
  243. package/src/interactive/terminal-lease.ts +1 -0
  244. package/src/interactive/turn-context.ts +96 -23
  245. package/src/interactive/turn-middleware.ts +1 -0
  246. package/src/interactive/turn-runtime.ts +37 -8
  247. package/src/interactive/turn-state.ts +3 -0
  248. package/src/interactive/worker-progress.ts +440 -0
  249. package/src/interactive/worker-stream.ts +51 -110
  250. package/src/tools/agent-tools.ts +28 -3
  251. package/src/tools/ask-user.ts +21 -1
  252. package/src/tools/context/index.ts +2 -2
  253. package/src/tools/dispatch-arguments.ts +8 -0
  254. package/src/tools/dispatch-event-text.ts +19 -0
  255. package/src/tools/dispatch.ts +24 -1
  256. package/src/tools/monitor.ts +15 -0
  257. package/src/tools/registry.ts +15 -5
  258. package/src/tools/result-disposition.ts +156 -0
  259. package/src/tools/result-shaping.ts +59 -1
  260. package/src/tools/verify/authoring.ts +55 -54
  261. package/src/tools/worker-evidence.ts +19 -0
  262. package/src/worker/spec-contract.ts +43 -3
  263. package/dist/chunk-EFADSJET.js +0 -18
  264. package/dist/memory-4ALKDJ4Q.js +0 -246
  265. package/src/domains/eval/metrics/chaos-stream.ts +0 -93
@@ -1,7 +1,7 @@
1
1
  # Internal Eval Suites
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Private suites should live outside this repository. Keep datasets, prompts,
7
7
  live fleet coordinates, calibration outputs, and raw run artifacts in a private
@@ -15,9 +15,9 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
15
15
  ```
16
16
 
17
17
  Use `--out <dir>` when the artifact should be written outside the default Clio
18
- data directory. Public summaries can be copied into
19
- `benchmarks/results/<suite>/<run-id>/` only after they have been sanitized down
20
- to `manifest.json` and `summary.json`.
18
+ data directory. Product eval artifacts and external benchmark campaigns are
19
+ separate: public benchmark adapters live under `benchmarks/community/` and do
20
+ not use the eval runner.
21
21
 
22
22
  ## Context Regression Seed
23
23
 
@@ -268,44 +268,3 @@ thresholds:
268
268
  value: 0
269
269
  ```
270
270
 
271
- ---
272
-
273
- ## Soak Benchmark Suite
274
-
275
- The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
276
-
277
- Every suite runs through the product's own eval runner against a configured
278
- target; there is no separate soak runner:
279
-
280
- ```bash
281
- npm run build
282
- clio-coder eval run --suite benchmarks/soak/clio-soak.yaml \
283
- --target <id> --model <wireId> --clio-coder-entry dist/cli/index.js
284
- ```
285
-
286
- `tests/contracts/eval-soak-suite.test.ts` loads all four files in CI and drives
287
- `clio-soak.yaml` against a stub that seals receipts on purpose, so the gate is
288
- known to fail when sealing fails; the model runs themselves are operator-run.
289
-
290
- The soak suite comprises four specialized suite files:
291
-
292
- ### 1. Machinery Under Load (`clio-soak.yaml`)
293
- Evaluates the same task workload across two execution surfaces: the headless main-agent surface (`clio-run`) and a dispatched worker surface (`agent: coder`). It tests single-file bugs, multi-file bugs, and compaction continuity across restarts.
294
- - **Surface Differences**: Main-agent tasks verify session ledger continuity (`ledger.formatVersion`, `ledger.toolPairsUnmatched`, `ledger.assistantBetweenCallAndResult`), while dispatch worker tasks verify process group cleanup (`process.orphanedChildren == 0`).
295
- - **Compaction Continuity**: Verifies that compaction summaries are present (`continuity.compactionSummaryPresent`) and that pre-compaction facts are preserved (`continuity.answeredFromPreCompaction`).
296
- - **Suite-Wide Gates**: Gates on `receipt.sealed`, `receipt.integrityValid`, `receipt.outcomeMatchesExit`, `tokens.measured`, `stream.cumulativeSnapshots == 0`, `stream.usageDoubleCounted == false`, and `stream.segmentUsageMatchesMessages == true`.
297
-
298
- ### 2. Per-Step Write Boundaries (`clio-soak-boundary.yaml`)
299
- Validates write boundary enforcement across steps without model participation. Enforcement is strictly detect-and-rollback and is never sandboxing.
300
- - `write-boundary.rolled-back`: Verifies clean detection of allowlist violations (`writes_boundary_violation`), git-level file restoration, and sealed verdict generation (`boundary.violationsRolledBack == 1`, `boundary.rollbackIncomplete == 0`).
301
- - `write-boundary.rollback-incomplete`: Tests honest failure reporting when a path was dirty prior to snapshot taking. prior bytes exist only in the overwritten tree, so rollback leaves the tree unchanged and records incomplete rollback (`boundary.rollbackIncomplete == 1`, `boundary.violationsRolledBack == 0`).
302
-
303
- ### 3. Fault Injection Chaos (`clio-soak-chaos.yaml`)
304
- Evaluates system resilience against process signals.
305
- - `chaos.sigint-mid-tool`: Prompts Clio for a long-running bash tool call and injects `SIGINT` once the subprocess initializes. Asserts exit code `130`, confirms no orphaned children remain (`process.orphanedChildren == 0`), and verifies receipt sealing, receipt integrity, and provider token reporting.
306
-
307
- ### 4. Bounded Loops (`clio-soak-loop.yaml`)
308
- Validates iteration bounds and receipt accounting for fleet loops (`bounded-loop.fleet`).
309
- - **Loop Bounds**: Asserts that verification attempts do not exceed declared limits (`loop.attemptsSpent <= 3`), recovery attempts seal individual receipts (`loop.receiptsMatchRepairs == true`), and unneeded nodes report as `unneeded` rather than skipped or failed (`loop.skippedNodes == 0`).
310
- - **Two Token Accountings**: Distinguishes `tokens.*` (folded live off wire stdout by `createStreamInvariantFold`) from `receiptUsage.*` (journal receipts sealed and authenticated against ledger envelopes). On fleet runs, wire streaming is absent (`tokens.measured == false`), while journal receipts provide authenticated usage (`receiptUsage.measured == true`).
311
-
@@ -1,9 +1,9 @@
1
1
  # Evidence Corpus and Long-Term Memory
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.4). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
4
+ > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.6). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
5
5
 
6
- Clio Coder treats run claims and agent lessons as structured artifacts to support reproducibility and scientific provenance. In evaluations such as [SWE-bench](https://www.swebench.com), capturing granular execution evidence is essential for validating agent claims. Evidence corpora are deterministic directories built from run ledgers, receipts, sessions, audits, and eval artifacts. In v0.3.4, forensic evidence auto-builds on dispatch run completion: when a run finalizes, the observability domain automatically compiles the evidence bundle under `<dataDir>/evidence/run-<id>/` and updates a compact sidecar index row in `<stateDir>/evidence-index.json`. Long-term memory records are local, evidence-linked, and only injected after explicit approval. Use the TUI [`/view`](observability.md) command for interactive inspection of receipts, dispatch output, durable tool output, compaction summaries, and session accountability before building or citing evidence.
6
+ Clio Coder treats run claims and agent lessons as structured artifacts to support reproducibility and scientific provenance. In evaluations such as [SWE-bench](https://www.swebench.com), capturing granular execution evidence is essential for validating agent claims. Evidence corpora are deterministic directories built from run ledgers, receipts, sessions, audits, and eval artifacts. In v0.3.6, forensic evidence auto-builds on dispatch run completion: when a run finalizes, the observability domain automatically compiles the evidence bundle under `<dataDir>/evidence/run-<id>/` and updates a compact sidecar index row in `<stateDir>/evidence-index.json`. Long-term memory records are local, evidence-linked, and only injected after explicit approval. Use the TUI [`/view`](observability.md) command for interactive inspection of receipts, dispatch output, durable tool output, compaction summaries, and session accountability before building or citing evidence.
7
7
 
8
8
  Source of truth: `src/domains/evidence/**`, `src/domains/memory/**`, `src/cli/evidence.ts`, and `src/cli/memory.ts`.
9
9
 
@@ -246,7 +246,8 @@ Mutation-report receipts are grounded directly against observed tool events reco
246
246
 
247
247
  ```bash
248
248
  clio-coder memory list
249
- clio-coder memory propose --from-evidence <evidenceId>
249
+ clio-coder memory propose --from-evidence <evidenceId> [scope options]
250
+ clio-coder memory promote --from-handoff <path> [--entry <id>...] --scope <scope> [scope options]
250
251
  clio-coder memory approve <memoryId>
251
252
  clio-coder memory reject <memoryId>
252
253
  clio-coder memory prune --stale
@@ -267,6 +268,8 @@ The store is capped at `500` records and is sorted by scope, key, creation time,
267
268
  ```mermaid
268
269
  stateDiagram-v2
269
270
  evidence --> proposed: propose --from-evidence
271
+ taskBank --> proposed: /memory selected-entry action
272
+ redactedHandoff --> proposed: promote --from-handoff
270
273
  proposed --> approved: approve <id>
271
274
  proposed --> rejected: reject <id>
272
275
  approved --> rejected: reject <id>
@@ -277,6 +280,45 @@ stateDiagram-v2
277
280
 
278
281
  Records must cite at least one evidence ID to be considered for prompt injection. Rejected records remain in the store until stale pruning so the same bad lesson is not immediately re-proposed from the same evidence.
279
282
 
283
+ Task-bank promotion is a reviewed export from transient execution memory. The
284
+ `/memory` overlay offers repo and global proposal actions only on selected
285
+ knowledge and procedural rows. Status remains private and cannot enter the
286
+ promotion service. The first global action arms a warning, and the second
287
+ action acknowledges the broader applicability. A successful action writes an
288
+ unapproved record and names the separate `memory approve` command required to
289
+ make it injectable.
290
+
291
+ The CLI consumes a version 2 `clio-task-memory` handoff snapshot. Omitting
292
+ `--entry` proposes every knowledge and procedural entry; repeating `--entry`
293
+ selects exact entry IDs. Version 2 snapshots carry source session, evidence,
294
+ runtime, agent, timestamps, and export-redaction facts. Version 1 snapshots
295
+ remain seedable but cannot be promoted because they do not carry source
296
+ session or evidence provenance.
297
+
298
+ Every promotion redacts secret-shaped values before `records.json` is written.
299
+ The durable provenance block records the source kind, session, selected entry,
300
+ entry class and timestamps, plus the replacement count and source field paths.
301
+ Promotion never approves its own output.
302
+
303
+ ### Explicit scope selection
304
+
305
+ Reviewed scope options are closed to four choices:
306
+
307
+ | Scope | Required selection | Validation |
308
+ | --- | --- | --- |
309
+ | `repo` | `--repository <canonical-absolute-path>` | The path must exist and already equal its canonical absolute identity. Symlink aliases and paths containing unresolved segments are rejected. |
310
+ | `global` | `--acknowledge-global` | The acknowledgement is separate from `--scope global`. |
311
+ | `runtime` | `--runtime <id>` | The ID must be valid and must occur in the source provenance. |
312
+ | `agent` | `--agent <id>` | The ID must be valid and must occur in the source provenance. |
313
+
314
+ The same options may be added to `memory propose --from-evidence`. With no
315
+ scope option, evidence proposals keep the existing inference order. An
316
+ explicit repository may differ from the repository that produced the
317
+ evidence, which supports a reviewed lesson about repository A learned while
318
+ working in repository B. Runtime and agent overrides may only select an exact
319
+ identity already recorded by the evidence. Global scope always requires its
320
+ own acknowledgement. No inference path widens an explicit choice.
321
+
280
322
  ---
281
323
 
282
324
  ## Prompt injection rules
@@ -287,7 +329,7 @@ Defaults:
287
329
 
288
330
  | Constraint | Default |
289
331
  | --- | --- |
290
- | Scopes | `global`, `repo` |
332
+ | Base scopes | `global`, `repo` |
291
333
  | Token budget | `400` estimated tokens |
292
334
  | Max records | `5` |
293
335
  | Required status | `approved: true` |
@@ -296,6 +338,12 @@ Defaults:
296
338
 
297
339
  Rendered memory lines always cite record ID, scope, lesson, and evidence IDs. The prompt tells the model not to extrapolate beyond cited findings.
298
340
 
341
+ Interactive main-agent sessions additionally admit records for the exact
342
+ active runtime. `clio-coder run --agent` admits records for the exact resolved
343
+ runtime and selected agent. Runtime and agent records use structured identity
344
+ fields; `appliesWhen` text cannot grant either applicability. Missing,
345
+ malformed, or different active identities exclude those records.
346
+
299
347
  ### Repository-scoped identity
300
348
 
301
349
  Repository memory is selected by an exact canonical absolute-path identity. The interactive orchestrator and `clio-coder run --agent` compute that identity from the active working directory; symlink aliases collapse to the same key. A repository move, a different Git worktree path, a subdirectory launch, a malformed identity, or a missing identity does not inherit another repository's memory. Global records are unaffected.
@@ -308,15 +356,27 @@ Every `scope: "repo"` record must carry:
308
356
 
309
357
  The structured `repository` field is the only applicability mechanism: store validation rejects repo records without it, and `appliesWhen` tokens never grant repository applicability. There is intentionally no automatic path rewrite for moved repositories or worktrees: a filesystem move produces a different identity and the record simply stops applying until it is re-scoped with new evidence.
310
358
 
359
+ Runtime and agent records follow the same fail-closed shape:
360
+
361
+ ```json
362
+ { "runtime": { "kind": "runtime", "key": "openai" } }
363
+ ```
364
+
365
+ ```json
366
+ { "agent": { "kind": "agent", "key": "coder" } }
367
+ ```
368
+
369
+ Only the field matching the record scope is present.
370
+
311
371
  ---
312
372
 
313
373
  ## Recommended workflow
314
374
 
315
375
  1. Build evidence from the run/session/eval that taught the lesson.
316
376
  2. Inspect the evidence and findings.
317
- 3. Propose memory from the evidence.
318
- 4. Review the proposed lesson for correctness and scope.
319
- 5. Approve only if it is durable and useful.
377
+ 3. Propose memory from the evidence, or promote selected public task memory from `/memory` or a redacted handoff.
378
+ 4. Review the proposed lesson, source provenance, redaction facts, and exact scope.
379
+ 5. Approve only if it is durable and useful under that scope.
320
380
  6. Reject incorrect or overbroad records.
321
381
  7. Prune stale records periodically.
322
382
 
package/docs/evolution.md CHANGED
@@ -1,7 +1,7 @@
1
1
  # Evolution and Change Manifests
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
7
7
 
@@ -1,6 +1,6 @@
1
1
  # Exit Codes & Machine-Readable Output Contracts
2
2
 
3
- This document specifies the process exit codes, machine-readable JSON streaming formats, standard I/O separation rules, and `--help` conventions across all Clio Coder CLI commands in `v0.3.4`.
3
+ This document specifies the process exit codes, machine-readable JSON streaming formats, standard I/O separation rules, and `--help` conventions across all Clio Coder CLI commands in `v0.3.6`.
4
4
 
5
5
  Source implementations: `src/cli/` and `src/entry/`.
6
6
 
@@ -1,7 +1,7 @@
1
1
  # Extensions, Prompt Templates, Skills, and Share Archives
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Clio Coder has lightweight community-oriented resource packaging. Extensions are filesystem bundles that contribute prompts and skills. Share archives are portable JSON files for moving project/user Clio resources between machines or collaborators. Themes are built into the engine and are no longer loaded from extensions.
7
7
 
@@ -251,7 +251,7 @@ Share archives are single JSON files:
251
251
  "formatVersion": 1,
252
252
  "manifest": {
253
253
  "format": "clio.share.v1",
254
- "clioVersion": "0.3.4",
254
+ "clioVersion": "0.3.6",
255
255
  "createdAt": "...",
256
256
  "files": []
257
257
  },
@@ -1,6 +1,6 @@
1
1
  # Fleet Dispatch
2
2
 
3
- > **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.4).
3
+ > **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.6).
4
4
 
5
5
  Clio Coder dispatches bounded worker agents. With a fleet configured, those
6
6
  workers run on remote machines over SSH while the orchestrator keeps every
@@ -77,7 +77,11 @@ briefing wins. Supplying both `task` and `tasks` fails instead of choosing one.
77
77
  After approval, execution consumes only the registry-owned resolved plan, so
78
78
  later mutation of raw arguments cannot change either field.
79
79
 
80
- Recipes may declare `budget: {toolCalls, readReserve, synthesis}`. `toolCalls` is the admitted-call phase boundary; the final `readReserve` slots accept canonical `read` plus the agent's granted mutation tools, so a writer can still deliver inside its own reserve; `synthesis: true` forces a text-only final round, while `false` stops after the admitted phase. `guardrails.workerToolCallCap` is transported separately as the ceiling on executed calls and always wins when lower. Native workers and Claude SDK enforce this policy. Claude Code and Antigravity reject explicit-budget recipes because their black-box loops cannot provide equivalent per-call mediation. Before launch, every admitted WorkerSpec v3 contains one concrete effective budget and a settings fingerprint, including custom recipes whose source omitted a budget.
80
+ Recipes declare a default with `budget: {toolCalls, readReserve, synthesis}`. They may also declare `maximum: {toolCalls, readReserve}` inside that object. A recipe without `maximum` is an exact pin, which preserves the fixed behavior of existing recipes. A ranged recipe admits the optional dispatch request `budget: {toolCalls, readReserve, retryRevision?}` only when the request is inside its maximum. `retryRevision` has the same two integer fields and preauthorizes the ceiling that a later automatic retry, bounded result-contract revision, or review revision may select. The loop guard raises a result-contract revision boundary only in this case. A phase without that ceiling cannot grow and retains the existing text-only repair behavior.
81
+
82
+ `toolCalls` is the admitted-call phase boundary. The final `readReserve` slots accept canonical `read` plus the agent's granted mutation tools, so a writer can still deliver inside its own reserve. Admission requires integers and `0 <= readReserve < toolCalls` for every declared phase. `synthesis: true` forces a text-only final round, while `false` stops after the admitted phase. `guardrails.workerToolCallCap` remains the operator-controlled lifetime ceiling and always wins when lower. A default may be clamped by a lower operator cap so default callers retain their prior behavior; an explicit request outside the operator cap is denied.
83
+
84
+ Admission computes one immutable envelope with the recipe policy, invocation request, effective worker budget, and every clamp or escalation reason. Native workers and Claude SDK enforce the effective budget. Claude Code, Antigravity, and ACP delegation reject invocation envelopes because their black-box loops cannot provide equivalent per-call mediation. Before launch, every admitted WorkerSpec v3 still contains one concrete effective budget and a settings fingerprint. The envelope provenance is sealed in the run ledger and receipt and appears in monitor, fleet status, and the live fleet card.
81
85
 
82
86
  ## Node setup
83
87
 
@@ -658,6 +662,27 @@ hard block.
658
662
  renders `local`), gate badges (`gate reviewer c2`), reroute badges, live
659
663
  tool activity (names only; arguments never cross the worker stdout seam),
660
664
  and a per-worker context meter.
665
+ - `Enter` on the selected Fleet Runs row opens its worker detail: the phase,
666
+ the running call with a redacted action descriptor (`bash running npm
667
+ test`), and the bounded tail of the worker's own prose. The default list
668
+ stays compact, so a fan-out of scouts costs one card each until an operator
669
+ opens one. Detail follows the cursor rather than pinning to a run.
670
+ - The board and the transcript worker block read one projection
671
+ (`src/interactive/worker-progress.ts`), so they cannot disagree about what a
672
+ worker is saying or touching. It keeps 40 lines and 4096 bytes of tail, 8
673
+ distinct tool names, 4 recent actions, and accepts 16 KB of delta bytes per
674
+ 250 ms; what the bounds refuse is counted and named on the card beside the
675
+ `/view dispatch:<runId>` deep link.
676
+ - Action descriptors are composed where the arguments are trusted: the tool
677
+ registry's admission path, the Claude tool mapper, and the ACP update
678
+ mapper. Each reads a fixed verb vocabulary and a fixed argument-field
679
+ allowlist, scrubs credentials, strips escape sequences, and bounds the
680
+ result to 64 characters before it crosses the worker stdout seam. Raw
681
+ argument objects never cross at all.
682
+ - Reasoning content is never displayed. The detail may name a `thinking`
683
+ phase and the usage facts the card already carries, never the text.
684
+ - Settlement replaces the provisional tail with the sealed receipt's answer;
685
+ a run whose receipt cannot be read keeps its own last durable message.
661
686
  - The context meter renders the worker's last-message context occupancy
662
687
  against the model's context window: healthy below 80 percent, warn from 80,
663
688
  critical from 95.
@@ -667,6 +692,7 @@ hard block.
667
692
  - The monitor tool reports the node and reroute lineage on `status`, `list`,
668
693
  and `collect`.
669
694
  - `clio-coder fleet status [--json]` shows the durable ledger view cross-process.
695
+ - A worker permission escalation uses the `Worker escalation` consequence tier in operator presentation. The tier names the worker agent and run and describes where the one-shot answer returns. It does not approve the request, change the worker's inherited autonomy, or weaken the safety net; the existing worker escalation protocol remains the only resolution path.
670
696
 
671
697
  ## Speculation observer
672
698
 
@@ -3,7 +3,7 @@
3
3
  Clio Coder is designed to be self-contained and platform-compliant. This document outlines the default directory paths, file purposes, permission levels, and lifecycle commands (`install`, `reset`, `upgrade`, and `uninstall`). Clio Coder installs from npm as `@iowarp/clio-coder` (`npm install -g @iowarp/clio-coder`, published since v0.3.0) or from a source checkout with a deterministic local symlink; the CLI classifies both install kinds and `clio-coder upgrade` handles each.
4
4
 
5
5
  > [!TIP]
6
- > **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.4). You can open it directly in any web browser to view details dynamically.
6
+ > **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.6). You can open it directly in any web browser to view details dynamically.
7
7
 
8
8
  ---
9
9
 
@@ -229,7 +229,7 @@ Upgrading from 0.3.1 to 0.3.3 is automated:
229
229
  clio-coder upgrade
230
230
  ```
231
231
 
232
- Key lifecycle and operational updates in v0.3.4:
232
+ Key lifecycle and operational updates in v0.3.6:
233
233
  - Upgraded the underlying engine SDK libraries to 0.84.0 with signal-aware OAuth cancellation.
234
234
  - Hardened migration resilience: damaged `credentials.yaml` files no longer block upgrades when no renames are needed (#121); `--skip-migrations` is available as a recovery override.
235
235
  - Fullscreen TUI mode (`terminal.tuiMode`, `terminal.fullscreenScrollbar`) is available via Settings → Terminal (restart required). Adaptive presentation pacing is the live `terminal.smoothStreaming` setting; 0.3.3 defaults it to `off`, with conservative `auto` and explicit `on` available from the same section.
@@ -1,7 +1,7 @@
1
1
  # Middleware and Component Registry
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Clio Coder has two related but separate surfaces:
7
7
 
@@ -114,7 +114,24 @@ Middleware hook budgets are phase-aware through `DEFAULT_MIDDLEWARE_HOOK_BUDGETS
114
114
 
115
115
  Per-phase budgets can be overridden via `CLIO_CODER_HOOK_BUDGET_<PHASE>_MS` or global `CLIO_CODER_HOOK_BUDGET_MS`. Warmup grace exempts initial calls (`DEFAULT_HOOK_BUDGET_WARMUP_CALLS = 1`), and steady-state warnings trigger when at least 3 of the last 5 post-warmup calls exceed budget (`DEFAULT_HOOK_BUDGET_WINDOW = 5`, `DEFAULT_HOOK_BUDGET_THRESHOLD = 3`). Overruns are reported but do not abort the turn. The orchestrator and workers share the middleware contract, but worker guard state is process-local.
116
116
 
117
- Middleware reminders are visible request text, not hidden prompt state. `turn_start` reminders flush into the same accepted request; `turn_end` reminders flush once on the next request. The built-in stalled-turn rule can request one automatic continuation for a user prompt, then stops rather than looping forever.
117
+ Middleware reminders are visible request text, not hidden prompt state. `turn_start` reminders flush into the same accepted request; `turn_end` reminders flush once on the next request. A `request_continuation` from any producer is capped at one automatic continuation per user prompt; a second producer in the same prompt gets a footer notice that the nudge is spent, and the turn is handed back to the operator rather than looped.
118
+
119
+ ### Built-in registrations
120
+
121
+ These ship in every interactive session. Each is one bounded behavior with a visible reminder; none changes a tool policy or a safety verdict.
122
+
123
+ | Id | Hooks | What it does |
124
+ | --- | --- | --- |
125
+ | `nudge.stalled-turn` | `turn_end` | The one declarative rule. A turn that called no tools and ended on an announced action ("Next I will inspect `src/cli/index.ts`") is continued once with a reminder to perform it or say plainly that it is finished. Questions, "let me know", conditional offers ("if you want me to"), and completion statements are not announcements. |
126
+ | `observer.skills-reminder` | `turn_start`, `turn_end` | Once per session, on the first substantive turn, when installed or installable skills exist, injects one line teaching the suggestion protocol: list with `context(scope="skills")`, open the reply with `Suggested skill: /skill <name>` when one matches, then continue the task in the same turn. Only the operator loads a skill. At `turn_end`, a reply that made the suggestion and stopped with only listing calls behind it is continued once (#184): the suggestion is not the task. Greetings do not spend the session's one reminder; a resumed or forked session never gets one. |
127
+ | `observer.task-board-reminder` | `turn_start` | Once per session, when the operator's text literally enumerates three or more steps (`1)`, `2.`, `step 3:`, or three bulleted lines), injects one line asking for `tasks action="plan"` before the first edit. Prose that merely mentions numbers never counts. |
128
+ | `nudge.open-tasks` | `turn_end` | A settled work turn (one that called tools) that ends while the session task board still has pending or active tasks is continued once with the open list. Pure conversation turns, aborted or errored turns, and boards where every remaining task is blocked do not trigger. |
129
+ | `nudge.detached-dispatch` | `turn_end` | A settled turn that ends while a detached dispatch batch has every run terminal and uncollected is continued once, naming the ready batches; `monitor mode="collect"` clears it, including across resume. Batches with runs still in flight, and surfaces without `monitor`, do not trigger. |
130
+ | `nudge.read-only-exploration` | `after_tool`, `turn_end` | After nine or more read-only calls (`read`, `grep`, `find`, `ls`, `code_nav`, read-only shell) in one user turn without a successful Scout dispatch, injects one advisory to delegate broad reconnaissance to Scout. One advisory per user turn, and only on surfaces that have `dispatch`. |
131
+ | `rail.unbacked-worker-claim` | `after_tool`, `turn_end` | A reply that reports worker or Scout results in a turn with no `dispatch` call gets one warning that the claim is not backed by a receipt. A `[worker result]` note the operator shared is receipt-backed and exempt. No continuation: the operator decides. |
132
+ | `observer.memory-intervention` | `after_tool` | Every `memory.intervention.everyNTools` tool calls, asks a background model for a bounded reflection over the recent window and injects it as a reminder when it arrives. Governed by the `memory.intervention` settings block. |
133
+
134
+ Two coded controls sit beside the registrations rather than among them. `tool-choice-control` turns `require_tool` and `lock_tools` effects into the provider's tool-choice field for the next round: a required tool clears when that tool starts, a lock lasts until the next submitted turn and outranks later requirements. `hook-receipts` is the durable ring (200 entries, throttled to one write per two seconds) of user-defined hook executions that `clio-coder config inspect` reads.
118
135
 
119
136
  User-defined hook declarations load from three places: `<extensionRoot>/hooks.yaml`, `.clio-coder/hooks.yaml`, and `.clio-coder/hooks.local.yaml`. A hook can be `prompt`, `effect`, or `command`. Command hooks run an argv array without a shell, under the workspace with a timeout and bounded output, and every hook execution emits a receipt.
120
137
 
@@ -1,7 +1,7 @@
1
1
  # Model Catalog, Runtime Refresh, and Field Notes
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Clio Coder treats a selectable model as the intersection of three sources:
7
7
 
@@ -1,7 +1,7 @@
1
1
  # Observability Viewer
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/observability_blueprint.html](html/observability_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/observability_blueprint.html](html/observability_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  `/view` is the interactive artifact viewer for a Clio session. It keeps the live transcript compact while preserving a full inspection path for durable artifacts, task ledgers, and successful workspace outputs.
7
7
 
@@ -159,8 +159,8 @@ The base provenance sets, steering, routing, quality, worker identity, and resul
159
159
  | `safety.toolTelemetry.ingestionErrors` | `number` | Current dispatch receipts | Malformed or lost frames, event-fold/source errors, and drain timeouts that make otherwise mediated telemetry incomplete | experimental |
160
160
  | `safety.toolTelemetry.unfinished` | `{ tool, count }[]` | Current dispatch receipts | Tool starts that had no matching finish when the receipt sealed | experimental |
161
161
  | `safety.toolTelemetry.workspaceMutationPossible` | `boolean` | Current dispatch receipts | Whether incomplete or unavailable telemetry could conceal a shared-workspace mutation; retry admission fails closed when true | experimental |
162
- | `autonomyEnforcement.grade` | `string` | Always in v0.3.4 | The autonomy grade level enforced for the run | experimental |
163
- | `autonomyEnforcement.autonomy` | `string` | Always in v0.3.4 | The effective autonomy level name (e.g. auto-edit, suggest, read-only, full-auto) | experimental |
162
+ | `autonomyEnforcement.grade` | `string` | Always in v0.3.6 | The autonomy grade level enforced for the run | experimental |
163
+ | `autonomyEnforcement.autonomy` | `string` | Always in v0.3.6 | The effective autonomy level name (e.g. auto-edit, suggest, read-only, full-auto) | experimental |
164
164
  | `autonomyEnforcement.externalMode` | `string` | When running external worker | The execution mode of the external worker runtime | experimental |
165
165
  | `autonomyEnforcement.dangerousBypass` | `boolean` | When running external worker | Whether a safety bypass was explicitly activated | experimental |
166
166
  | `validationGrounding.claimed` | `number` | Validation grounding evaluated | Count of validations claimed by worker | experimental |
@@ -1,6 +1,6 @@
1
1
  # Proactive task memory
2
2
 
3
- > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.4).
3
+ > **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.6).
4
4
 
5
5
  Clio's proactive task memory protects long-running work from behavioral state
6
6
  decay: a requirement, environment fact, failed attempt, or diagnosis can still
@@ -133,7 +133,7 @@ it runs detached from it.
133
133
  | Interval | After `memory.intervention.everyNTools` completed tools since the last prompted step; default 10. This is the nondeterministic/citation-gated path. |
134
134
  | Tool-error streak | Two consecutive error outcomes. A successful tool resets the streak. |
135
135
  | Loop signal | Reuses the orchestrator loop guard's verdict; it does not infer a second competing loop detector. |
136
- | Repeated failure | The rules tier records failed tool fingerprints and annotates the failing tool result once the same failure appears twice in the bounded trajectory. |
136
+ | Repeated failure | The rules tier records failed operation fingerprints and annotates the failing tool result once the same failure appears twice in the bounded trajectory. |
137
137
  | Post-compaction | The first turn start after compaction restores status and knowledge once, without a model call, because compaction is precisely where execution facts leave the active window. |
138
138
 
139
139
  ### Two delivery channels
@@ -145,11 +145,13 @@ repeated failure uses exactly one of them:
145
145
 
146
146
  - **Mid-turn annotation.** The second identical failure appends one cited
147
147
  `Memory:` advisory to that tool's own result, through the existing
148
- `annotate_tool_result` effect the loop guard already uses. The advisory digest
149
- takes the first line of the tool error that names a problem, falling back to
150
- the first line when no line names one. The model reads it on its very next round.
151
- This is spent once per fingerprint per turn and re-earned in a later turn, because
152
- the same command failing again after an operator turn is news again.
148
+ `annotate_tool_result` effect the loop guard already uses. The advisory uses
149
+ the canonical result-disposition digest when one is available. Older hook
150
+ producers fall back to the first tool-error line that names a problem. Every
151
+ digest is redacted and byte-capped before it reaches the task bank. The model
152
+ reads the advisory on its very next round. This is spent once per operation
153
+ fingerprint per turn and re-earned in a later turn, because the same command
154
+ failing again after an operator turn is news again.
153
155
  - **Next-turn reminder.** Post-compaction reactivation and any background-model
154
156
  reminder ride the `inject_reminder` buffer into the next submitted turn, inside
155
157
  the visible `<system-reminder>` block, and persist in the session ledger.
@@ -319,15 +321,23 @@ operation while leaving deterministic protection active.
319
321
  Measured on the shipped prompt against `google/gemma-4-26b-a4b-qat`, across ten
320
322
  live steps and forty controlled runs on the same route.
321
323
 
322
- The tier writes `update_status` reliably and `save_knowledge` rarely, and that is
323
- correct rather than broken. A trajectory step carries the tool name, a bounded
324
- call description, an outcome, and a result digest. On success the digest is an
325
- opaque result fingerprint, so a window of successful reads tells the model which
326
- files were touched and nothing about what is in them. There is no durable fact in
327
- that input, and a status line is the only faithful thing to write about it.
328
-
329
- Three candidate causes were ruled out by controlled runs that changed one
330
- variable at a time:
324
+ Earlier measurements found that the tier wrote `update_status` reliably and
325
+ `save_knowledge` rarely. At that time a successful trajectory step carried an
326
+ opaque result fingerprint, so a window of successful reads told the model which
327
+ files were touched and nothing about what was in them.
328
+
329
+ A current trajectory step keeps two fields with different jobs. The operation
330
+ fingerprint identifies repeated calls and remains derived only from the tool name
331
+ and arguments. The result digest is human-readable diagnostic content from the
332
+ canonical result-disposition projection, with explicit source provenance. Secret
333
+ redaction and a 240-byte cap apply before the digest reaches the task bank or the
334
+ background request. A metadata-only disposition contributes outcome facts and no
335
+ captured body. Results without a canonical disposition use a redacted deterministic
336
+ fallback, so older tool producers remain useful without gaining a second model
337
+ summarizer.
338
+
339
+ Three candidate causes were ruled out in the earlier implementation by
340
+ controlled runs that changed one variable at a time:
331
341
 
332
342
  - rewriting the prompt's second worked example to carry a `save_knowledge` moved
333
343
  nothing, and made the model emit no operations at all in four of five runs;
@@ -1,7 +1,7 @@
1
1
  # Prompt Envelope and Tools
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/tools_blueprint.html](html/tools_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/tools_blueprint.html](html/tools_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  Clio Coder keeps the model-facing envelope stable and moves enforcement into the runtime registry and safety policy.
7
7
 
@@ -78,6 +78,8 @@ Three tools sit in a plane for containment rather than class. `git` is read-only
78
78
 
79
79
  Registration is conditional on wiring: `context` gains its workspace scope only when a session contract is bound, `dispatch`/`monitor`/`steer` register only with a dispatch contract, and `ask_user` registers only when an interactive handler exists. Dispatch tool profiles narrow the surface for workers: `minimal-local` is `read`, `grep`, `find`, `ls`, `git`, `context`, `code_nav`; `science-local` adds `verify`; `full-agent` keeps everything.
80
80
 
81
+ `ask_user` keeps its typed `exposure: local | outward` admission fact separate from caller prose. The registry uses exposure only in the enforced autonomy mapping. After admission, the host carries the normalized fact into the shared decision-presentation classifier; question text, headers, options, summaries, and requested color or severity words cannot select a consequence tier. The resulting presentation object contains no admission disposition and cannot grant authority.
82
+
81
83
  ### Consolidated call shapes
82
84
 
83
85
  Several tools absorb what used to be separate tools:
@@ -87,7 +89,7 @@ Several tools absorb what used to be separate tools:
87
89
  - `context(scope="workspace"|"docs"|"skills")` is the one OBSERVE entry point for material about the working environment: the session workspace snapshot, retrieval over Clio's bundled documentation (`query` required), and skill listing or loading (`name` optional, `include_tree` for the skill's resource files).
88
90
  - `verify(check?, path?, args?, browser?, cwd?, timeout_ms?)` runs declared verification. `verify()` lists package.json verification scripts and strict version-1 `.clio-coder/verifiers.yaml` entries through the same `{id, description, command, cwd, timeoutMs, tags, source}` projection. `verify(check="<id>")` runs a package script or the catalog's exact argv/cwd/timeout through safe-exec with no shell. Model `args`, cwd, timeout, output-cap, and environment fields cannot mutate a project entry. `verify(check="frontend", path=...)` validates an HTML/CSS/JS artifact without granting shell access.
89
91
  - `artifact(kind="plan"|"review"|"report", content, ...)` writes named artifacts behind one surface: Markdown documents (default `.clio-coder/artifacts/PLAN.md`/`REVIEW.md`/`REPORT.md`; `path` may override inside the workspace) that terminate the turn, because writing the artifact is the answer. Skills are not artifacts; a `SKILL.md` is written with the ordinary write tool and validated by the skills loader.
90
- - `dispatch(task?, tasks?, mode?, ...)` supports a first-class singular assignment (`task`) and a batch (`tasks`), never both. `task` is worker instructions; `briefing` is optional bounded parent context/data and cannot replace it. Briefing stays a separate dynamic message and receipt provenance, never part of the receipt task. A shared top-level briefing applies to strings and objects without an override; an object-level briefing wins. Blank values are omitted, the cap is 12,000 UTF-8 bytes, and approval pins the exact canonical value. Ordinary handles enter one registered event consumer immediately. Synchronous calls auto-wait for stream-and-receipt completion; `detach:true` returns ids after durable batch registration while the same consumer continues. Review and compete retain gate-sensitive direct drains. Task objects may include `persona` and `tool_profile`. Pipeline output is threaded as bounded data. A successful native or ACP run requires a nonempty receipt-sealed final output; exit zero without one fails as `worker_final_output_missing`, with unfinished text retained only as partial diagnostics. `dispatch(list=true)` renders the catalog.
92
+ - `dispatch(task?, tasks?, mode?, ...)` supports a first-class singular assignment (`task`) and a batch (`tasks`), never both. `task` is worker instructions; `briefing` is optional bounded parent context/data and cannot replace it. Briefing stays a separate dynamic message and receipt provenance, never part of the receipt task. A shared top-level briefing applies to strings and objects without an override; an object-level briefing wins. Blank values are omitted, the cap is 12,000 UTF-8 bytes, and approval pins the exact canonical value. Ordinary handles enter one registered event consumer immediately. Synchronous calls auto-wait for stream-and-receipt completion; `detach:true` returns ids after durable batch registration while the same consumer continues. Review and compete retain gate-sensitive direct drains. Task objects may include `persona`, `tool_profile`, and a typed `budget: {toolCalls, readReserve, retryRevision?}`. The budget must fit the recipe's authored range and the operator lifetime cap. `retryRevision` is the only authority for a later retry, result-contract revision, or review revision to grow its phase. Pipeline output is threaded as bounded data. A successful native or ACP run requires a nonempty receipt-sealed final output; exit zero without one fails as `worker_final_output_missing`, with unfinished text retained only as partial diagnostics. `dispatch(list=true)` renders the catalog.
91
93
  - `monitor(run_id?, mode?)` is read-only visibility into known synchronous and detached runs: `list` enumerates, `status` reports one, `peek` returns the in-process event tail, `receipt` exposes the stored evidence, and `wait` observes one run without collecting or canceling it. `collect` is the authoritative terminal batch operation over a detached batch or run-id list; collect before final synthesis. Completed output reports receipt integrity, evidence verification, briefing provenance, and bounded project-context provenance as different fields.
92
94
  - `steer(run_id, action, message?)` controls a running worker: `guide` writes a canonical trimmed steering message to an HTTP or SDK worker and `cancel` terminates it. Successfully written steers gain ordered byte/hash/timestamp provenance; after the runtime accepts the guidance, `clio_steer_received` acknowledges the exact matching sequence, and prose is never stored in ledger or receipt. Single-shot subprocess runtimes and ACP remain non-steerable. Interactive operators can steer synchronous live-input runs; parent-model steering requires detached ids because model tools are sequential.
93
95
 
@@ -1,7 +1,7 @@
1
1
  # Provider Adapter Cookbook
2
2
 
3
3
  > [!TIP]
4
- > **Interactive Spec Available:** An interactive runtime adapter descriptor builder and probe sequence capability checklist is located at [docs/html/provider_adapter_blueprint.html](html/provider_adapter_blueprint.html) (Version: 0.3.4).
4
+ > **Interactive Spec Available:** An interactive runtime adapter descriptor builder and probe sequence capability checklist is located at [docs/html/provider_adapter_blueprint.html](html/provider_adapter_blueprint.html) (Version: 0.3.6).
5
5
 
6
6
  This cookbook guides developers through implementing custom model runtimes and inference server integrations within Clio Coder. It explains the runtime descriptor interfaces, probing protocols, model synthesis, and how to configure reasoning and thinking behaviors.
7
7