@iowarp/clio-coder 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (226) hide show
  1. package/CHANGELOG.md +407 -0
  2. package/CODE_OF_CONDUCT.md +21 -0
  3. package/CONTRIBUTING.md +224 -0
  4. package/LICENSE +202 -0
  5. package/NOTICE +9 -0
  6. package/README.md +798 -0
  7. package/SECURITY.md +72 -0
  8. package/assets/clio-coder-logo-128.webp +0 -0
  9. package/damage-control-rules.yaml +419 -0
  10. package/dist/acp-UMLFVA3F.js +92 -0
  11. package/dist/agents-Q4MYPMUW.js +91 -0
  12. package/dist/auth-O6HYIJ6J.js +521 -0
  13. package/dist/chunk-262G75JS.js +35 -0
  14. package/dist/chunk-26BZQOAD.js +1281 -0
  15. package/dist/chunk-2J63S4SF.js +508 -0
  16. package/dist/chunk-3DANZDGR.js +717 -0
  17. package/dist/chunk-4UQA7NCT.js +29 -0
  18. package/dist/chunk-527KG6XR.js +497 -0
  19. package/dist/chunk-5LDRNKX2.js +1063 -0
  20. package/dist/chunk-5N2FG33Q.js +25 -0
  21. package/dist/chunk-67MTHP2E.js +135 -0
  22. package/dist/chunk-6CWDTGUC.js +20 -0
  23. package/dist/chunk-7BHLZB3A.js +2115 -0
  24. package/dist/chunk-7RBKDI66.js +348 -0
  25. package/dist/chunk-AMFR5YA3.js +541 -0
  26. package/dist/chunk-BBUH4VAA.js +1224 -0
  27. package/dist/chunk-BYEU76JP.js +899 -0
  28. package/dist/chunk-CLJ5HLUD.js +458 -0
  29. package/dist/chunk-D5YD55AR.js +116 -0
  30. package/dist/chunk-DXQNI4PC.js +61 -0
  31. package/dist/chunk-E3NYWENM.js +1004 -0
  32. package/dist/chunk-GNGDQYDU.js +34688 -0
  33. package/dist/chunk-GOTUR54M.js +9 -0
  34. package/dist/chunk-HBU5MTAM.js +41 -0
  35. package/dist/chunk-HMYNFFY4.js +28 -0
  36. package/dist/chunk-JPOWPFCU.js +1010 -0
  37. package/dist/chunk-JWHCJDCI.js +1215 -0
  38. package/dist/chunk-KBR4MZZR.js +41 -0
  39. package/dist/chunk-KKKPTZLM.js +93 -0
  40. package/dist/chunk-ME6DNWIU.js +66 -0
  41. package/dist/chunk-NI4DEJMC.js +88 -0
  42. package/dist/chunk-O4EJEDHO.js +659 -0
  43. package/dist/chunk-PIDUD6M2.js +31 -0
  44. package/dist/chunk-PS4PFJQP.js +29459 -0
  45. package/dist/chunk-QV47YRF4.js +48 -0
  46. package/dist/chunk-RQDWMVRB.js +279 -0
  47. package/dist/chunk-TFSSEXL6.js +136 -0
  48. package/dist/chunk-TKHQ4DGZ.js +8290 -0
  49. package/dist/chunk-TPOCL34A.js +2876 -0
  50. package/dist/chunk-UGYAX5YI.js +565 -0
  51. package/dist/chunk-UHTSULZS.js +461 -0
  52. package/dist/chunk-UU3R62TT.js +128 -0
  53. package/dist/chunk-UWIJNAOB.js +3906 -0
  54. package/dist/chunk-VOO7NYPP.js +914 -0
  55. package/dist/chunk-VPAWTYLY.js +117 -0
  56. package/dist/chunk-WD6AJM35.js +1216 -0
  57. package/dist/chunk-X3BR7HWV.js +115 -0
  58. package/dist/chunk-X3NE4WVW.js +120 -0
  59. package/dist/chunk-XNISANGE.js +1395 -0
  60. package/dist/chunk-XV4ZJ6ZM.js +3177 -0
  61. package/dist/cli/index.js +236 -0
  62. package/dist/clio-KIQ5SNDS.js +53 -0
  63. package/dist/components-JVHMUBEB.js +653 -0
  64. package/dist/config-ZFCDBMDC.js +372 -0
  65. package/dist/configure-G4E3A2PG.js +27 -0
  66. package/dist/context-CDXTP2MP.js +293 -0
  67. package/dist/context-E3KIFVXI.js +185 -0
  68. package/dist/context-clear-3F4PLXOS.js +102 -0
  69. package/dist/context-index-Q7YSYTR3.js +106 -0
  70. package/dist/docs-YIETIWZI.js +280 -0
  71. package/dist/doctor-M5HJJZOL.js +61 -0
  72. package/dist/domains/agents/builtins/architect.md +33 -0
  73. package/dist/domains/agents/builtins/coder.md +31 -0
  74. package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
  75. package/dist/domains/agents/builtins/debugger.md +30 -0
  76. package/dist/domains/agents/builtins/documenter.md +31 -0
  77. package/dist/domains/agents/builtins/git-master.md +30 -0
  78. package/dist/domains/agents/builtins/provenance.md +30 -0
  79. package/dist/domains/agents/builtins/researcher.md +71 -0
  80. package/dist/domains/agents/builtins/scout.md +42 -0
  81. package/dist/domains/agents/builtins/tester.md +31 -0
  82. package/dist/domains/agents/builtins/verifier.md +30 -0
  83. package/dist/domains/agents/builtins/wiki-writer.md +41 -0
  84. package/dist/eval-B3KZZESM.js +2674 -0
  85. package/dist/evidence-V67CHM35.js +233 -0
  86. package/dist/evolve-YDZSUQYA.js +518 -0
  87. package/dist/extensions-SRG7XCAH.js +207 -0
  88. package/dist/fleet-CA2CRTVG.js +760 -0
  89. package/dist/fleet-preflight-CLIAX7YR.js +21 -0
  90. package/dist/init-2OZDJE2D.js +227 -0
  91. package/dist/memory-3PIQQAKX.js +207 -0
  92. package/dist/models-DY35XI7Y.js +237 -0
  93. package/dist/paths-5OMXW7Z4.js +57 -0
  94. package/dist/preload-KZVHET2B.js +11 -0
  95. package/dist/reset-PIFYNOS3.js +216 -0
  96. package/dist/run-3VSPP24F.js +735 -0
  97. package/dist/share-D36RQCXM.js +241 -0
  98. package/dist/skills-F2MRLELY.js +445 -0
  99. package/dist/skills-eval-E2ZTW4PL.js +932 -0
  100. package/dist/targets-DZMEZAH4.js +977 -0
  101. package/dist/trace-7NYCUI2J.js +250 -0
  102. package/dist/uninstall-AD3JWHBB.js +322 -0
  103. package/dist/upgrade-WYYBKGDY.js +301 -0
  104. package/dist/usage-ULIDAGFF.js +755 -0
  105. package/dist/version-ROZ6CZKH.js +16 -0
  106. package/dist/wiki-generate-PKFIX6OB.js +377 -0
  107. package/dist/worker/entry.js +1739 -0
  108. package/docs/README.md +93 -0
  109. package/docs/acp.md +120 -0
  110. package/docs/alcf-provider.md +72 -0
  111. package/docs/architecture.md +172 -0
  112. package/docs/artifact-versions.md +54 -0
  113. package/docs/built-in-agents.md +265 -0
  114. package/docs/capacity-and-scheduling.md +97 -0
  115. package/docs/commands-and-modes.md +554 -0
  116. package/docs/config-knobs-audit.md +115 -0
  117. package/docs/configuration-and-targets.md +812 -0
  118. package/docs/context-engine.md +236 -0
  119. package/docs/dispatch-architecture-rationale.md +126 -0
  120. package/docs/documentation-coverage.md +46 -0
  121. package/docs/documentation-guide.md +166 -0
  122. package/docs/environment-variables.md +105 -0
  123. package/docs/eval-runner.md +205 -0
  124. package/docs/evals-internal.md +298 -0
  125. package/docs/evidence-and-memory.md +243 -0
  126. package/docs/evolution.md +143 -0
  127. package/docs/exit-codes-and-output.md +74 -0
  128. package/docs/extensions-and-sharing.md +306 -0
  129. package/docs/fleet-demo-runbook.md +179 -0
  130. package/docs/fleet-dispatch.md +591 -0
  131. package/docs/glossary.md +75 -0
  132. package/docs/html/agents_blueprint.html +936 -0
  133. package/docs/html/alcf_blueprint.html +324 -0
  134. package/docs/html/architecture_blueprint.html +850 -0
  135. package/docs/html/commands_blueprint.html +794 -0
  136. package/docs/html/config_knobs_audit_blueprint.html +178 -0
  137. package/docs/html/configuration_blueprint.html +1080 -0
  138. package/docs/html/context_blueprint.html +603 -0
  139. package/docs/html/documentation_blueprint.html +832 -0
  140. package/docs/html/environment_blueprint.html +404 -0
  141. package/docs/html/eval_blueprint.html +743 -0
  142. package/docs/html/evals_internal_blueprint.html +190 -0
  143. package/docs/html/evolution_blueprint.html +674 -0
  144. package/docs/html/extensions_blueprint.html +2065 -0
  145. package/docs/html/fleet_dispatch_blueprint.html +286 -0
  146. package/docs/html/index.html +919 -0
  147. package/docs/html/lifecycle_blueprint.html +723 -0
  148. package/docs/html/memory_blueprint.html +699 -0
  149. package/docs/html/middleware_blueprint.html +664 -0
  150. package/docs/html/models_blueprint.html +2366 -0
  151. package/docs/html/observability_blueprint.html +683 -0
  152. package/docs/html/provider_adapter_blueprint.html +245 -0
  153. package/docs/html/safety_blueprint.html +1386 -0
  154. package/docs/html/shared.css +571 -0
  155. package/docs/html/shared.js +143 -0
  156. package/docs/html/skills_blueprint.html +671 -0
  157. package/docs/html/soak_blueprint.html +182 -0
  158. package/docs/html/tool_usage_blueprint.html +350 -0
  159. package/docs/html/tools_blueprint.html +2249 -0
  160. package/docs/html/trace_blueprint.html +235 -0
  161. package/docs/html/tui_design_blueprint.html +314 -0
  162. package/docs/html/validation_blueprint.html +961 -0
  163. package/docs/html/worker_dispatch_blueprint.html +231 -0
  164. package/docs/installation-and-lifecycle.md +308 -0
  165. package/docs/middleware-and-components.md +148 -0
  166. package/docs/model-catalog.md +189 -0
  167. package/docs/observability.md +233 -0
  168. package/docs/proactive-memory.md +452 -0
  169. package/docs/prompt-envelope-and-tools.md +142 -0
  170. package/docs/provider-adapter-cookbook.md +148 -0
  171. package/docs/release-cut-checklist.md +138 -0
  172. package/docs/safety-model.md +357 -0
  173. package/docs/scientific-validation.md +105 -0
  174. package/docs/session-lifecycle.md +156 -0
  175. package/docs/skills-marketplace.md +46 -0
  176. package/docs/tool-usage.md +527 -0
  177. package/docs/trace-store.md +132 -0
  178. package/docs/troubleshooting.md +33 -0
  179. package/docs/tui-design.md +239 -0
  180. package/docs/worker-dispatch-mechanics.md +242 -0
  181. package/package.json +132 -0
  182. package/skills/README.md +408 -0
  183. package/skills/git/commit-crafting/SKILL.md +79 -0
  184. package/skills/git/commit-crafting/evals.md +92 -0
  185. package/skills/git/create-pr/SKILL.md +116 -0
  186. package/skills/git/create-pr/evals.md +114 -0
  187. package/skills/git/investigate-issue/SKILL.md +139 -0
  188. package/skills/git/investigate-issue/evals.md +94 -0
  189. package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
  190. package/skills/git/resolve-merge-conflicts/evals.md +58 -0
  191. package/skills/git/review-changes/SKILL.md +103 -0
  192. package/skills/git/review-changes/evals.md +85 -0
  193. package/skills/git/worktree-create/SKILL.md +92 -0
  194. package/skills/git/worktree-create/evals.md +97 -0
  195. package/skills/git/worktree-create/references/worktree-setup.md +66 -0
  196. package/skills/git/worktree-merge/SKILL.md +95 -0
  197. package/skills/git/worktree-merge/evals.md +114 -0
  198. package/skills/skill-marketplace.json +261 -0
  199. package/skills/workflow/cut-it/SKILL.md +86 -0
  200. package/skills/workflow/cut-it/evals.md +42 -0
  201. package/src/domains/agents/builtins/architect.md +33 -0
  202. package/src/domains/agents/builtins/coder.md +31 -0
  203. package/src/domains/agents/builtins/context-bootstrap.md +38 -0
  204. package/src/domains/agents/builtins/debugger.md +30 -0
  205. package/src/domains/agents/builtins/documenter.md +31 -0
  206. package/src/domains/agents/builtins/git-master.md +30 -0
  207. package/src/domains/agents/builtins/provenance.md +30 -0
  208. package/src/domains/agents/builtins/researcher.md +71 -0
  209. package/src/domains/agents/builtins/scout.md +42 -0
  210. package/src/domains/agents/builtins/tester.md +31 -0
  211. package/src/domains/agents/builtins/verifier.md +30 -0
  212. package/src/domains/agents/builtins/wiki-writer.md +41 -0
  213. package/src/domains/agents/fleets/build-review.md +34 -0
  214. package/src/domains/agents/fleets/build-test.md +35 -0
  215. package/src/domains/agents/fleets/sdlc.md +86 -0
  216. package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
  217. package/src/domains/prompts/fragments/identity/clio.md +26 -0
  218. package/src/domains/prompts/fragments/operating/contract.md +64 -0
  219. package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
  220. package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
  221. package/src/domains/prompts/fragments/safety/read-only.md +13 -0
  222. package/src/domains/prompts/fragments/safety/suggest.md +13 -0
  223. package/src/domains/prompts/fragments/wiki/page.md +75 -0
  224. package/src/domains/prompts/fragments/wiki/plan.md +48 -0
  225. package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
  226. package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
@@ -0,0 +1,148 @@
1
+ # Middleware and Component Registry
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.0).
5
+
6
+ Clio Coder has two related but separate surfaces:
7
+
8
+ 1. **Components**: deterministic inventory of files that can affect harness behavior.
9
+ 2. **Middleware**: an experimental hook/effect contract around tool, turn, and compaction lifecycle points.
10
+
11
+ The components surface is active and user-facing through `clio-coder components`. The middleware runtime is intentionally conservative in the current alpha: the hook/effect types, validation helpers, declarative rule engine, built-in registrations, and a receipted hook-file surface exist, but arbitrary repository or user middleware packages are not a shipped public extension point. Enforcing guard registrations ride the same hook runtime at the composition root: the loop guard, protected-artifacts guard, dispatch dedup, file and skill observers, tool-prose checks, and finish-contract assessor form the middleware tier of the safety net (see [safety-model.md](safety-model.md)).
12
+
13
+ ---
14
+
15
+ ## Component scanner
16
+
17
+ Source: `src/domains/components/scan.ts` and `src/domains/components/types.ts`.
18
+
19
+ The scanner reads files, computes SHA-256 hashes, and emits a stable `ComponentSnapshot`:
20
+
21
+ ```ts
22
+ interface ComponentSnapshot {
23
+ version: 1;
24
+ generatedAt: string;
25
+ root: string;
26
+ components: HarnessComponent[];
27
+ }
28
+ ```
29
+
30
+ It does not execute scanned files.
31
+
32
+ ### Component kinds
33
+
34
+ `COMPONENT_KINDS` currently contains:
35
+
36
+ | Kind | Typical source | Authority |
37
+ | --- | --- | --- |
38
+ | `prompt-fragment` | `src/domains/prompts/fragments/**/*.md` | advisory |
39
+ | `agent-recipe` | `src/domains/agents/builtins/*.md` | advisory |
40
+ | `tool-implementation` | `src/tools/*.ts` | enforcing |
41
+ | `tool-helper` | selected helper files such as `src/tools/registry.ts` | enforcing |
42
+ | `runtime-descriptor` | `src/domains/providers/runtimes/**/*.ts` | runtime-critical |
43
+ | `safety-rule-pack` | `damage-control-rules.yaml` | enforcing |
44
+ | `config-schema` | `src/core/defaults.ts`, `src/core/config.ts` (scanner list in `src/domains/components/scan.ts`) | runtime-critical |
45
+ | `session-schema` | session entry/contract files | runtime-critical |
46
+ | `receipt-schema` | dispatch receipt/integrity files | runtime-critical |
47
+ | `context-file` | `CLIO-CODER.md`, `CONTRIBUTING.md`, `SECURITY.md` | advisory |
48
+ | `doc-spec` | currently `docs/specs/**/*.md` if present | descriptive |
49
+ | `middleware` | reserved kind | enforcing |
50
+ | `memory` | reserved kind | advisory |
51
+ | `eval-suite` | reserved kind | descriptive |
52
+
53
+ > [!WARNING]
54
+ > The current scanner still looks for `doc-spec` files under `docs/specs/`. Most public docs now live flat under `docs/*.md`, so public docs may not appear as `doc-spec` components until the scanner is updated.
55
+
56
+ ### Reload classes
57
+
58
+ | Reload class | Meaning |
59
+ | --- | --- |
60
+ | `hot` | Can be reread during an active process where supported. |
61
+ | `next-dispatch` | Affects the next fleet worker dispatch. |
62
+ | `restart-required` | Low-level schemas/rules/runtimes should be treated as restart-bound. |
63
+ | `static` | Descriptive specs and suites. |
64
+
65
+ ---
66
+
67
+ ## Component CLI
68
+
69
+ ```bash
70
+ clio-coder components
71
+ clio-coder components --json
72
+ clio-coder components snapshot --out before.json
73
+ clio-coder components diff --from before.json --to after.json
74
+ ```
75
+
76
+ Snapshots are useful in reviews because they show behavior-affecting changes even when the raw diff is broad.
77
+
78
+ ---
79
+
80
+ ## Middleware contract
81
+
82
+ Source: `src/domains/middleware/types.ts`, `validate.ts`, `budget.ts`, and `runtime.ts`.
83
+
84
+ Supported hooks:
85
+
86
+ | Hook ID | Current use |
87
+ | --- | --- |
88
+ | `before_tool` | Guard and annotate a tool call before execution. Rejected or parked attempts still reach loop detection. |
89
+ | `after_tool` | Observe or annotate a completed tool result. File mutation and skill activation observers listen here and cannot change the result. |
90
+ | `turn_start` | Inject visible `<system-reminder>` text into the accepted request. |
91
+ | `turn_end` | Buffer reminders for the next request, including stalled-turn, tool-prose, and finish-contract advisories. |
92
+ | `on_compaction` | Observe compaction events. Effects from this hook are discarded by design. |
93
+
94
+ Supported effect kinds:
95
+
96
+ | Effect | Current meaning |
97
+ | --- | --- |
98
+ | `inject_reminder` | Structured reminder payload. |
99
+ | `annotate_tool_result` | Append deterministic annotation to a tool result. |
100
+ | `block_tool` | Hard-block a tool before execution. |
101
+ | `protect_path` | Register a protected artifact path in session state. |
102
+ | `request_continuation` | Ask the chat loop for one bounded automatic continuation. |
103
+ | `require_tool` | Require a specific tool for the next turn. |
104
+ | `lock_tools` | Lock available tools to the current subset. |
105
+
106
+ Declarative rules run before coded registrations. Scoped registrations match by hook and, for tool hooks, by tool name. Hook failures emit diagnostics and later hooks still run.
107
+
108
+ Middleware hook budgets are phase-aware through `DEFAULT_MIDDLEWARE_HOOK_BUDGETS_MS`:
109
+ - `before_tool`: 25 ms
110
+ - `after_tool`: 25 ms
111
+ - `turn_start`: 50 ms
112
+ - `turn_end`: 75 ms
113
+ - `on_compaction`: 150 ms
114
+
115
+ Per-phase budgets can be overridden via `CLIO_CODER_HOOK_BUDGET_<PHASE>_MS` or global `CLIO_CODER_HOOK_BUDGET_MS`. Warmup grace exempts initial calls (`DEFAULT_HOOK_BUDGET_WARMUP_CALLS = 1`), and steady-state warnings trigger when at least 3 of the last 5 post-warmup calls exceed budget (`DEFAULT_HOOK_BUDGET_WINDOW = 5`, `DEFAULT_HOOK_BUDGET_THRESHOLD = 3`). Overruns are reported but do not abort the turn. The orchestrator and workers share the middleware contract, but worker guard state is process-local.
116
+
117
+ Middleware reminders are visible request text, not hidden prompt state. `turn_start` reminders flush into the same accepted request; `turn_end` reminders flush once on the next request. The built-in stalled-turn rule can request one automatic continuation for a user prompt, then stops rather than looping forever.
118
+
119
+ User-defined hook declarations load from three places: `<extensionRoot>/hooks.yaml`, `.clio-coder/hooks.yaml`, and `.clio-coder/hooks.local.yaml`. A hook can be `prompt`, `effect`, or `command`. Command hooks run an argv array without a shell, under the workspace with a timeout and bounded output, and every hook execution emits a receipt.
120
+
121
+ ---
122
+
123
+ ## Validation helpers
124
+
125
+ Middleware validators enforce closed fields and known enum values. Minimal valid rule object:
126
+
127
+ ```json
128
+ {
129
+ "id": "lab.require-validation",
130
+ "source": "builtin",
131
+ "description": "Require validation after generated artifact writes.",
132
+ "enabled": true,
133
+ "hooks": ["turn_end"],
134
+ "effectKinds": ["request_continuation"]
135
+ }
136
+ ```
137
+
138
+ Minimal valid effect object examples:
139
+
140
+ ```json
141
+ { "kind": "block_tool", "reason": "protected path", "severity": "hard-block" }
142
+ ```
143
+
144
+ ```json
145
+ { "kind": "protect_path", "path": "out/checkpoint.nc", "reason": "validated output" }
146
+ ```
147
+
148
+ The current `MiddlewareRuleSource` is only `builtin`. Hook files compile into coded registrations on the same runtime, but they are not custom declarative rule sources and do not grant new tool authority.
@@ -0,0 +1,189 @@
1
+ # Model Catalog, Runtime Refresh, and Field Notes
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.0).
5
+
6
+ Clio Coder treats a selectable model as the intersection of three sources:
7
+
8
+ 1. **Configured targets** in `settings.yaml` (`targets[]`, `defaultModel`, and optional `wireModels`).
9
+ 2. **Live runtime probes** (`probe()` / `probeModels()`), which discover models that appeared after Clio started.
10
+ 3. **Catalog knowledge** from pi-ai provider catalogs, Clio's bundled local YAML knowledge base under `src/domains/providers/models/**`, and user/project model-catalog overlays.
11
+
12
+ ## Runtime refresh controls
13
+
14
+ - `/targets`: `r` probes the selected target; `R` probes all targets.
15
+ - `/model` or `/models`: `r` refreshes the selected row's target; `R` refreshes all targets.
16
+ - `clio-coder models`: probes live targets by default before printing the CLI model list. Use `--offline` to skip live probing. Former `--probe` and `--no-probe` flags are gone.
17
+
18
+ Configured `wireModels` and a target `defaultModel` remain selectable before a
19
+ live catalog is known; Clio labels those rows as `configured` or `default`.
20
+ Once a target returns a live catalog, that catalog is authoritative and models
21
+ the runtime no longer reports stop resolving. Live probe discoveries are labeled
22
+ `live` and carry load-state metadata when the runtime exposes it. This preserves
23
+ operator-curated defaults while still letting runtime discovery take over after
24
+ newly installed local models or newly entitled cloud models appear. Catalog YAML
25
+ entries are loaded when the provider domain is built, so bundled or overlay
26
+ edits require a restart or rebuild before they affect capability and quirk
27
+ matching.
28
+
29
+ Live provider probes are the preferred source for loaded context and per-model metadata. Clio keeps a 128k local-coding context recommendation, but it no longer treats that recommendation as provider truth for unknown local models: effective context comes from live probe data, an explicit target override, a model catalog/KB entry, or the runtime descriptor default. If the live target is below the recommendation, Clio reports a warning rather than silently inflating the displayed window.
30
+
31
+ `probeCapabilitiesForModel` is the one exact-id selector. When a router serves several models, capability resolution queries `probeCapabilitiesForModel` to ensure probe data is extracted only from the `/v1/models` row keyed to its own exact wire model ID.
32
+
33
+ Transient probe failures preserve the last-good catalog, load states,
34
+ capabilities, and notes for the same target identity, but the target health is
35
+ reported as down or unavailable with the probe error as the reason. Worker
36
+ dispatch canonicalizes requested model ids against the live catalog when one is
37
+ available, so a short alias can resolve to the canonical live id before the
38
+ worker spec and receipt are written.
39
+
40
+ ## Benchmarking Models
41
+
42
+ Model and config benchmark adapters ship under [benchmarks/community/](../benchmarks/community/). These adapters (such as `bench:swe`, `bench:scicode`, and the fleet benchmark `bench:tb`) drive Clio through the CLI or `clio-coder eval`.
43
+
44
+ For example, to run the fleet benchmark:
45
+ ```sh
46
+ npm run bench:tb -- --limit 3
47
+ ```
48
+
49
+ The benchmarks record context-window, thinking, sampling, weight quantization, and KV-cache settings so sweeps can be compared consistently.
50
+
51
+ ## What "sanctioned" means
52
+
53
+ A model family is "sanctioned" only when we can say what was tested and under which runtime. It is not a blanket endorsement. For each family, capture:
54
+
55
+ - exact model id / artifact / quantization;
56
+ - provider or runtime surface (`lmstudio-native`, `ollama-native`, `llamacpp`, `openrouter`, `openai-codex`, etc.);
57
+ - hardware and serving configuration;
58
+ - context window and max output actually exercised;
59
+ - tool-use, reasoning, vision, embeddings/rerank/FIM behavior where relevant;
60
+ - quirks needed by the engine (thinking mechanism, sampling, KV cache);
61
+ - failures and "do not use this route yet" notes.
62
+
63
+ Engine-visible quirks belong in catalog YAML entries under `quirks.kvCache`, `quirks.sampling`, and `quirks.thinking`. Bundled entries under `src/domains/providers/models/**/*.yaml` are for curated Clio-supported families. User/lab/project experiments should start as overlays before they are promoted into source. Free-form notes can live alongside catalog entries and in this docs area for later cookbooks/blog posts.
64
+
65
+ ## Local catalog overlays
66
+
67
+ Use model-catalog overlays when a local endpoint needs per-model facts that the
68
+ runtime cannot reliably probe, but the model is not yet ready to become bundled
69
+ Clio catalog knowledge.
70
+
71
+ Overlay roots are loaded in this order, with later roots winning equally
72
+ specific `matchPatterns`:
73
+
74
+ 1. Bundled Clio catalog: `src/domains/providers/models/**` or packaged `dist/providers-models`.
75
+ 2. User overlay: `$CLIO_CODER_CONFIG_DIR/model-catalog.d` or the platform config equivalent.
76
+ 3. Project overlay: `.clio-coder/model-catalog.d` under the current working directory.
77
+ 4. Extra overlay roots from `CLIO_CODER_MODEL_CATALOG_DIRS`, separated by the platform path delimiter.
78
+
79
+ Longest matching pattern still wins across all roots. This means a broad project
80
+ overlay such as `qwen` will not replace a more specific bundled entry such as
81
+ `qwen3.6-35b-a3b`; an equal match does replace it. Missing overlay directories
82
+ are ignored, so operators can create them only when needed.
83
+
84
+ Overlay files are ordinary YAML lists using the same schema as the bundled
85
+ catalog:
86
+
87
+ ```yaml
88
+ - family: ornith-1.0-35b-local
89
+ matchPatterns:
90
+ - ornith-1.0-35b
91
+ - Ornith-1.0-35B-Q4_K_M-262K
92
+ capabilities:
93
+ chat: true
94
+ tools: true
95
+ toolCallFormat: qwen
96
+ reasoning: true
97
+ thinkingFormat: qwen-chat-template
98
+ structuredOutputs: json-schema
99
+ vision: false
100
+ audio: false
101
+ embeddings: false
102
+ rerank: false
103
+ fim: false
104
+ contextWindow: 262144
105
+ maxTokens: 65536
106
+ quirks:
107
+ sampling:
108
+ thinking:
109
+ temperature: 0.6
110
+ topP: 0.95
111
+ topK: 20
112
+ thinking:
113
+ mechanism: always-on
114
+ guidance: |
115
+ The serving endpoint separates Qwen-style thinking into
116
+ reasoning_content while content contains the final answer.
117
+ ```
118
+
119
+ Use `settings.yaml` `wireModels` for target inventory. Use overlays for
120
+ per-model semantics and quirks. Promote an overlay into the bundled catalog only
121
+ after the model behavior is verified and useful beyond one operator's target.
122
+
123
+ ## Field note template
124
+
125
+ Use this shape when testing a subscription model, homelab GPU target, research-lab allocation, or new local runtime:
126
+
127
+ ```md
128
+ ## <model family or exact model> on <runtime>
129
+
130
+ - Date:
131
+ - Operator / lab:
132
+ - Runtime target:
133
+ - Provider / endpoint:
134
+ - Hardware:
135
+ - Model id / artifact:
136
+ - Quantization / precision:
137
+ - Context / max output tested:
138
+ - Auth / subscription tier:
139
+
140
+ ### Serving config
141
+ - Command or UI settings:
142
+ - GPU layers / tensor parallel / KV cache:
143
+ - Sampler defaults:
144
+
145
+ ### Smoke tests
146
+ - Tool calling:
147
+ - Reasoning control:
148
+ - Long-context behavior:
149
+ - Vision / embeddings / rerank / FIM:
150
+ - Latency / throughput notes:
151
+
152
+ ### Outcome
153
+ - Status: candidate | verified | limited | avoid
154
+ - Recommended Clio runtime:
155
+ - Required catalog quirks:
156
+ - Known failures:
157
+ - Follow-up benchmarks:
158
+ ```
159
+
160
+ ## Reasoning Controls and Thinking Replay Semantics
161
+
162
+ The Context Engine evaluates thinking mechanisms per model target and manages live reasoning streams. Depending on the runtime capabilities, Clio Coder employs specific thinking replay semantics to ensure chain-of-thought data is preserved or replayed correctly in the conversation history:
163
+
164
+ - **Ollama Native (`ollama-native`):** Ollama utilizes the native `thinking` field in the request and response payloads. The engine handles Ollama-specific effort levels and streams reasoning increments cleanly through the native thinking channel.
165
+ - **LM Studio Native (`lmstudio-native`):** Because LM Studio does not expose a native reasoning field, the engine replays prior thinking blocks by prepending them to assistant message payloads. These are formatted as a text-prepended part wrapped in `<think>` and `</think>` tags.
166
+ - **OpenAI Completions (`openai-completions`):** The OpenAI-compatible completions provider preserves reasoning blocks within assistant messages. It replays thinking blocks via the `reasoning_content` parameter in the message history, ensuring that the model maintains its chain-of-thought across conversational turns without stripping the data.
167
+ - **Anthropic OAuth / API (`anthropic-max`):** Uses the `anthropic-extended` thinking format. The engine supports Anthropic's native extended thinking block protocol, streaming thinking increments and outputting them wrapped appropriately or natively depending on target capabilities.
168
+ - **Reasoning-Never Models (`thinking.mechanism: none`):** When a model is configured or cataloged with `thinking.mechanism: none`, it is treated as a reasoning-never model. For these models, Clio must not send any thinking fields or parameters in requests, must not replay thinking blocks, must not surface thinking events to the TUI, and must not preserve or log reasoning token usage in metrics.
169
+
170
+ ---
171
+
172
+ ## Subscription Catalog Models
173
+
174
+ Subscription models are registered and managed as standard HTTP/cloud targets:
175
+
176
+ - **`openai-codex` (ChatGPT Plus/Pro OAuth):** Maps to catalog-backed Codex model ids surfaced by `clio-coder configure --list` and `clio-coder models` via a browser-minted subscription OAuth token, supporting complete chat, vision, and tool-use capabilities.
177
+ - **`anthropic-max` (Claude Pro/Max OAuth):** Powers chat and workers using catalog-backed Claude model ids surfaced by `clio-coder configure --list` and `clio-coder models`. It relies on pi-ai's `anthropic` OAuth provider. During auth initialization, it alerts the operator to usage-terms caveat via:
178
+ `Connects with your Claude Pro/Max subscription via OAuth (the same path Claude Code uses). Using subscription credentials outside Anthropic's first-party apps may not align with their terms of service; enable at your own discretion.`
179
+
180
+ ---
181
+
182
+ ## Promotion path
183
+
184
+ 1. Capture raw field notes in docs or a lab notebook.
185
+ 2. Add or update a user/project catalog overlay with capabilities and quirks.
186
+ 3. Add focused unit/integration coverage when behavior changes engine routing.
187
+ 4. Refresh `/models` with `R` and verify the selected row reports the expected source/caps.
188
+ 5. Promote the cleaned overlay into the bundled catalog only when the model family is ready to bless for Clio users.
189
+ 6. Promote the cleaned field note into a cookbook, guideline, or community blog post.
@@ -0,0 +1,233 @@
1
+ # Observability Viewer
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** An interactive dashboard is located at [docs/html/observability_blueprint.html](html/observability_blueprint.html) (Version: 0.3.0).
5
+
6
+ `/view` is the interactive artifact viewer for a Clio session. It keeps the live transcript compact while preserving a full inspection path for durable artifacts.
7
+
8
+ ```text
9
+ /view
10
+ /view <id-or-filter>
11
+ /view verify <runId>
12
+ ```
13
+
14
+ `/view` opens a full-screen split viewer. The left pane groups artifacts by category and supports type-to-filter. The right pane renders the selected artifact with pager controls. `Tab` or `Shift+Tab` switches between the artifact list and details. `Left` and `Right` jump to the previous or next non-empty category from either pane; `Up` and `Down` select artifacts in the list or scroll details in the content pane. Category jumps honor the active filter and wrap at the ends. `v` verifies a selected receipt. `o` shows the absolute backing path through the notice channel when the selected artifact has one; pathless artifacts produce a warning notice instead. In the list pane, `Esc` clears a non-empty filter before a second `Esc` closes the viewer.
15
+
16
+ ---
17
+
18
+ ## The Evidence Spine End-to-End
19
+
20
+ Clio Coder operates on a single, unified accountability spine that connects in-loop execution diagnostics with durable forensic evidence. The spine operates at two distinct layers:
21
+
22
+ 1. **Live Layer (In-Loop Assessment)**: During a session, the safety domain executes a cheap in-memory scan over the last 80 entries on every `turn_end` hook. It detects validation commands, dispatch receipts, protected artifacts, and requested inspections to check if completion claims are backed by evidence. The live kinds are mapped to the canonical evidence taxonomy.
23
+ 2. **Forensic Layer (Post-Completion Aggregator)**: When a run completes, the observability domain aggregates the session ledger, receipts, transcripts, and audit logs into a rich forensic bundle.
24
+
25
+ ```mermaid
26
+ graph TD
27
+ dispatch[Dispatch Completed/Failed Event] --> obs[Observability Bus Subscriber]
28
+ obs --> build[Asynchronous buildEvidence]
29
+ build --> bundle[Write Forensic Bundle to dataDir]
30
+ build --> index[Append row to evidence-index.json in stateDir]
31
+ index --> view[Surfaced in /view Accountability Panel]
32
+ ```
33
+
34
+ ### Auto-Build on Dispatch Completion
35
+
36
+ The observability extension subscribes to the `dispatch.completed` and `dispatch.failed` channels via the `SafeEventBus` (in `src/domains/observability/extension.ts`). When a run terminates, the domain initiates the forensic builder without blocking the event bus or TUI rendering.
37
+
38
+ The builder runs `buildEvidence({ dataDir, stateDir, runId })` to read the state files and compile a detailed bundle under `<dataDir>/evidence/run-<runId>/`. If a headless run is executing, the observability stop hook (`stop()`) flushes all in-flight build promises before the process exits, ensuring no data is lost. Any build failures are swallowed and logged to stderr to prevent compiler or file-lock issues from crashing the main run.
39
+
40
+ ### The Sidecar Index
41
+
42
+ After building the forensic bundle, the domain appends a metadata row to the sidecar index file located at `<stateDir>/evidence-index.json`. The file is kept as a JSON array acting as a bounded ring (capped at 1000 rows).
43
+
44
+ To prevent concurrent Clio processes from corrupting the index, writes are queued within the process and serialize across processes using the shared state-file locking mechanism.
45
+
46
+ An `EvidenceIndexRow` has the following schema:
47
+ ```json
48
+ {
49
+ "runId": "4f89d2a9c12",
50
+ "evidenceId": "run-4f89d2a9c12",
51
+ "tags": ["test-failure", "session-linked"],
52
+ "firstPassSuccess": false,
53
+ "findingCount": 2,
54
+ "generatedAt": "2026-06-25T14:30:00.000Z"
55
+ }
56
+ ```
57
+
58
+ ---
59
+
60
+ ## Artifact Categories and Path Layouts
61
+
62
+ Clio resolves directories under platform-specific XDG defaults (on Linux, these default to `~/.config/clio-coder/`, `~/.local/share/clio-coder/`, and `~/.local/state/clio-coder/`).
63
+
64
+ | Category | Description | Backing Path |
65
+ | --- | --- | --- |
66
+ | **Accountability** | Rolling first-pass-success rate and failure-cause histogram. | `<stateDir>/evidence-index.json` |
67
+ | **Evidence bundles** | Deterministic run or session overviews, findings, totals, and linked files. | `<dataDir>/evidence/<evidenceId>/` |
68
+ | **Receipts** | Durable run receipts verified by SHA-256 integrity digests. | `<stateDir>/receipts/<runId>.json` |
69
+ | **Dispatch outputs** | Logs and ledger records detailing worker execution. | `<stateDir>/runs.json` and `<stateDir>/receipts/<runId>.json` |
70
+ | **Task ledgers** | Per-turn task-board goals, active runs, and required validation evidence. | `<stateDir>/sessions/<cwdHash>/<sessionId>/current.jsonl` |
71
+ | **Tool outputs** | Offloaded large outputs or execution logs. | `<stateDir>/scratch/<sessionId>/<toolCallId>.txt` |
72
+ | **Protected artifacts** | Validation-protected artifact metadata and its absolute artifact path when available. | Session ledger record plus the protected workspace path |
73
+ | **Compaction** | Summaries of compacted history sessions. | `<stateDir>/sessions/<cwdHash>/<sessionId>/current.jsonl` |
74
+ | **Prompt manifests** | One validated record per prompt compile: `systemPromptHash`, previous hash, token estimate, thinking dial at compile time, per-section token estimates, and per-fragment content hashes. Identifies the exact compiled prompt and supports hash diffs without storing prompt text. Malformed records appear as an explicit read-error artifact. | `<stateDir>/sessions/<cwdHash>/<sessionId>/prompt-manifest.jsonl` |
75
+ | **Safety audit rows** | Current-session safety and permission decisions, with malformed ledger lines surfaced separately. | `<stateDir>/audit/<date>.jsonl` |
76
+
77
+ ---
78
+
79
+ ## The Accountability Panel
80
+
81
+ The first category in the TUI split viewer is **Accountability**. It reads the sidecar index directly to present a live summary without loading heavy forensic logs.
82
+
83
+ ### First-Pass Success Rate
84
+ A run is marked as a first-pass success when:
85
+ - The terminal dispatch outcome succeeded.
86
+ - The run had zero dispatch retries (attempt 0).
87
+ - The built bundle contains validation evidence, meaning the `no-validation` tag is absent.
88
+
89
+ The TUI displays this rate as:
90
+ `first-pass success: <succeeded-attempts>/<total-attempts> (<pct>%)`
91
+
92
+ ### Failure-Cause Histogram
93
+ The TUI lists the top failure causes sorted by frequency (descending), then by tag name (ascending). The histogram filters out provenance and quality tags (such as `audit-linked`, `session-linked`, and `no-validation`) and displays only real failure causes:
94
+ - `timeout`
95
+ - `auth-failure`
96
+ - `missing-dependency`
97
+ - `build-failure`
98
+ - `test-failure`
99
+ - `blocked-tool`
100
+
101
+ ---
102
+
103
+ ## Receipt Integrity Verification
104
+
105
+ Pressing `v` on a selected receipt or running `/view verify <runId>` performs cryptographic integrity checks:
106
+
107
+ 1. **Read Receipt**: Reads the receipt JSON from `<stateDir>/receipts/<runId>.json`.
108
+ 2. **Resolve Ledger**: Looks up the run envelope inside `<stateDir>/runs.json`.
109
+ 3. **Verify Integrity**: Recomputes the SHA-256 digest over the strict v15 receipt and reconstructible ledger fields. The digest covers every current field, including steering, routing intent and decision, route quality, worker identity, execution role, and result-contract conformance. Every version other than 15 fails verification; there is no historical receipt reader.
110
+ 4. **Report Result**: The viewer reports `ok` or the verification failure reason. It does not rename or delete the receipt. Startup orphan recovery may quarantine corrupt orphan receipt files as `<name>.json.corrupt`, but `/view verify` is read-only.
111
+
112
+ ---
113
+
114
+ ## Receipt Fields for Dispatch Provenance
115
+
116
+ A receipt carries optional provenance and context blocks that answer "what happened" for a chained (pipeline), composed (persona override), escalated, briefed, steered, or external run. Those optional blocks remain absent when unused. Current receipts carry strict integrity v15 and an explicit `outcomeCode: null` when no classified deterministic failure occurred. Automation consumers must treat the optional blocks below as absent by default and `outcomeCode` as nullable; older receipt versions are invalid.
117
+
118
+ Receipt integrity verification and evidence verification are independent.
119
+ `receipt_integrity=verified/v15/sha256` means Clio called the receipt verifier
120
+ against the ledger envelope; merely finding an embedded digest is not enough.
121
+ `evidence_verification=<verified|unverified|not_applicable|unknown>/<basis>`
122
+ describes validation evidence inside that verified receipt. Likewise,
123
+ `briefing` authenticates parent-supplied dispatch data, while
124
+ `project_context` authenticates the separately rendered bounded project
125
+ message. Model-facing dispatch and collect output name all four concepts
126
+ separately and never substitute one hash for another.
127
+
128
+ The evidence bundle renders these sets in `transcript.md` (human sentences) and `trace.cleaned.jsonl` (structured run rows), `clio-coder evidence inspect` prints them as a `provenance <runId>:` block, and the `dispatch` tool appends a compact suffix to each run line plus additive keys on `details.runs[]`. A timed-out or denied escalation also raises an `escalation` finding in the bundle.
129
+
130
+ The base provenance sets, steering, routing, quality, worker identity, and result-conformance coverage all enter in v0.2.9. These fields are labeled `experimental`: their strict v15 shape is frozen for the release, but the labels stay experimental until the schema is promoted post-1.0. For the complete version registry and migration contract across all artifacts, see [artifact-versions.md](artifact-versions.md).
131
+
132
+ | Field path | Type | When present | Meaning | Status |
133
+ | --- | --- | --- | --- | --- |
134
+ | `pipeline.fromRunId` | `string \| null` | Pipeline step after the first | Run whose final output was threaded in as input data; `null` when the upstream run id is unknown | experimental |
135
+ | `pipeline.position` | `number` | Pipeline step after the first | 1-based index of this step in the chain | experimental |
136
+ | `pipeline.inputBytes` | `number` | Pipeline step after the first | UTF-8 byte length of the threaded upstream text before the 12000-char cap | experimental |
137
+ | `pipeline.inputTruncated` | `boolean` | Pipeline step after the first | `true` when the 12000-char cap clipped the threaded input | experimental |
138
+ | `briefing.bytes` | `number` | A bounded parent briefing was sent | UTF-8 byte count of the exact canonical briefing content | experimental |
139
+ | `briefing.contentHash` | `string` | A bounded parent briefing was sent | SHA-256 of exact canonical briefing content; prose is not copied into the receipt | experimental |
140
+ | `projectContext.tier` | `"none" \| "bounded"` | Current receipts | Effective project-context policy | experimental |
141
+ | `projectContext.chars` | `number` | Bounded project context | Character count of the rendered project-context message | experimental |
142
+ | `projectContext.contentHash` | `string` | Nonempty bounded project context | SHA-256 of the rendered project-context message | experimental |
143
+ | `steering[].sequence` | `number` | A steer was successfully written | Stable 1-based order within the run | experimental |
144
+ | `steering[].bytes` | `number` | A steer was successfully written | UTF-8 bytes of the exact canonical trimmed steer | experimental |
145
+ | `steering[].contentHash` | `string` | A steer was successfully written | SHA-256 of the canonical steer; prose is not persisted | experimental |
146
+ | `steering[].sentAt` | `string` | A steer was successfully written | Write timestamp | experimental |
147
+ | `steering[].acknowledged` | `boolean` | A steer was successfully written | Whether a worker acknowledgement was actually observed | experimental |
148
+ | `steering[].acknowledgedAt` | `string` | Acknowledgement was observed | Acknowledgement timestamp | experimental |
149
+ | `outcomeCode` | five-value stable string union or `null` | Every v15 terminal receipt | Non-null for `vram_capacity_fit_failure`, `worker_tool_call_cap_exhausted`, `loop_guard_tools_disabled_exhausted`, `result_contract_exhausted`, or `worker_final_output_missing`; otherwise `null`. Each non-null code denotes terminal deterministic failure and is incompatible with `outcome: "succeeded"`. Dispatch retry policy consumes this code only, never diagnostic prose. | experimental |
150
+ | `personaOverride.promptHash` | `string` | Ad-hoc specialist whose persona replaced the recipe body | Hash of the composed static prompt; equals `staticCompositionHash` for the run | experimental |
151
+ | `safety.decisions.escalationRequested` | `number` | Run saw at least one permission escalation | Parked permission asks handed to the operator | experimental |
152
+ | `safety.decisions.escalationApproved` | `number` | Run saw at least one permission escalation | Escalations the operator approved | experimental |
153
+ | `safety.decisions.escalationDenied` | `number` | Run saw at least one permission escalation | Escalations the operator denied | experimental |
154
+ | `safety.decisions.escalationTimedOut` | `number` | Run saw at least one permission escalation | Escalations resolved by the timeout fallback (no operator decision) | experimental |
155
+ | `safety.toolTelemetry.coverage` | `"complete" \| "partial" \| "unavailable"` | Current dispatch receipts | Whether Clio can account for the runtime's complete tool start/finish stream | experimental |
156
+ | `safety.toolTelemetry.ingestionErrors` | `number` | Current dispatch receipts | Malformed or lost frames, event-fold/source errors, and drain timeouts that make otherwise mediated telemetry incomplete | experimental |
157
+ | `safety.toolTelemetry.unfinished` | `{ tool, count }[]` | Current dispatch receipts | Tool starts that had no matching finish when the receipt sealed | experimental |
158
+ | `safety.toolTelemetry.workspaceMutationPossible` | `boolean` | Current dispatch receipts | Whether incomplete or unavailable telemetry could conceal a shared-workspace mutation; retry admission fails closed when true | experimental |
159
+ | `autonomyEnforcement.grade` | `string` | Always in v0.3.0 | The autonomy grade level enforced for the run | experimental |
160
+ | `autonomyEnforcement.autonomy` | `string` | Always in v0.3.0 | The effective autonomy level name (e.g. auto-edit, suggest, read-only, full-auto) | experimental |
161
+ | `autonomyEnforcement.externalMode` | `string` | When running external worker | The execution mode of the external worker runtime | experimental |
162
+ | `autonomyEnforcement.dangerousBypass` | `boolean` | When running external worker | Whether a safety bypass was explicitly activated | experimental |
163
+ | `validationGrounding.claimed` | `number` | Validation grounding evaluated | Count of validations claimed by worker | experimental |
164
+ | `validationGrounding.grounded` | `number` | Validation grounding evaluated | Count of claimed validations matched against executed commands | experimental |
165
+ | `validationGrounding.ungrounded` | `string[]` | Validation grounding evaluated | Claim names with no matching execution, stably ordered and bounded | experimental |
166
+ | `validationGrounding.basis` | `"no-command-executed" \| "unmatched-command"` | Validation grounding evaluated | Why the unmatched claims are unmatched. Only `no-command-executed` takes a quality label away | experimental |
167
+ | `capabilityMismatch.verdict` | `"refuse" \| "flag"` | Capability assessment evaluated | Mismatch verdict for agent capability class versus task shape | experimental |
168
+ | `capabilityMismatch.agentId` | `string` | Capability assessment evaluated | Dispatched agent ID | experimental |
169
+ | `capabilityMismatch.capabilityClass` | `string` | Capability assessment evaluated | Admitted agent capability class | experimental |
170
+ | `capabilityMismatch.taskType` | `string` | Capability assessment evaluated | Classified task shape | experimental |
171
+ | `capabilityMismatch.suggestedAgentId` | `string \| null` | Capability assessment evaluated | Installed recipe that can do this work, or null when none is installed | experimental |
172
+ | `capabilityMismatch.detail` | `string` | Capability assessment evaluated | Human-readable explanation of the mismatch | experimental |
173
+
174
+ Only Clio-owned native/SSH worker wrappers and the Claude SDK path may transport worker-authored outcome events. ACP and black-box subprocess output cannot self-assert an outcome code; Clio may still assign `worker_final_output_missing` at its trusted finalization seam.
175
+
176
+ The escalation counters appear together and only when `escalationRequested` is present (at least one escalation occurred), so a deny-all or non-escalating run keeps its `safety.decisions` block unchanged.
177
+
178
+ ---
179
+
180
+ ## Worked Example: End-to-End Spine Flow
181
+
182
+ Here is a step-by-step trace of how a run passes through the spine.
183
+
184
+ ### 1. Dispatch Completion
185
+ A dispatched task to execute tests finishes. The dispatch domain persists the run envelope and the receipt, then emits the completion event:
186
+ ```json
187
+ {
188
+ "runId": "abc1234",
189
+ "status": "completed",
190
+ "exitCode": 0,
191
+ "lineage": { "attempt": 0 }
192
+ }
193
+ ```
194
+
195
+ ### 2. Forensic Build
196
+ The observability domain catches the event and triggers `buildAndIndexEvidence`. It generates the overview under `<dataDir>/evidence/run-abc1234/overview.json`:
197
+ ```json
198
+ {
199
+ "version": 1,
200
+ "evidenceId": "run-abc1234",
201
+ "source": { "kind": "run", "runId": "abc1234" },
202
+ "tags": ["session-linked"],
203
+ "totals": {
204
+ "toolCalls": 5,
205
+ "toolErrors": 0
206
+ }
207
+ }
208
+ ```
209
+
210
+ ### 3. Sidecar Append
211
+ Observability maps the result into an index row and writes it to `<stateDir>/evidence-index.json`. Because there was a successful validation tool call, `firstPassSuccess` is `true`:
212
+ ```json
213
+ {
214
+ "runId": "abc1234",
215
+ "evidenceId": "run-abc1234",
216
+ "tags": ["session-linked"],
217
+ "firstPassSuccess": true,
218
+ "findingCount": 0,
219
+ "generatedAt": "2026-06-25T14:31:00.000Z"
220
+ }
221
+ ```
222
+
223
+ ### 4. Surfacing in /view
224
+ When the operator opens the TUI split viewer `/view`, the Accountability panel reads the index and displays:
225
+ ```text
226
+ # Accountability
227
+
228
+ first-pass success: 1/1 (100%)
229
+
230
+ ## Top failure causes
231
+
232
+ none
233
+ ```