@iowarp/clio-coder 0.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (226) hide show
  1. package/CHANGELOG.md +407 -0
  2. package/CODE_OF_CONDUCT.md +21 -0
  3. package/CONTRIBUTING.md +224 -0
  4. package/LICENSE +202 -0
  5. package/NOTICE +9 -0
  6. package/README.md +798 -0
  7. package/SECURITY.md +72 -0
  8. package/assets/clio-coder-logo-128.webp +0 -0
  9. package/damage-control-rules.yaml +419 -0
  10. package/dist/acp-UMLFVA3F.js +92 -0
  11. package/dist/agents-Q4MYPMUW.js +91 -0
  12. package/dist/auth-O6HYIJ6J.js +521 -0
  13. package/dist/chunk-262G75JS.js +35 -0
  14. package/dist/chunk-26BZQOAD.js +1281 -0
  15. package/dist/chunk-2J63S4SF.js +508 -0
  16. package/dist/chunk-3DANZDGR.js +717 -0
  17. package/dist/chunk-4UQA7NCT.js +29 -0
  18. package/dist/chunk-527KG6XR.js +497 -0
  19. package/dist/chunk-5LDRNKX2.js +1063 -0
  20. package/dist/chunk-5N2FG33Q.js +25 -0
  21. package/dist/chunk-67MTHP2E.js +135 -0
  22. package/dist/chunk-6CWDTGUC.js +20 -0
  23. package/dist/chunk-7BHLZB3A.js +2115 -0
  24. package/dist/chunk-7RBKDI66.js +348 -0
  25. package/dist/chunk-AMFR5YA3.js +541 -0
  26. package/dist/chunk-BBUH4VAA.js +1224 -0
  27. package/dist/chunk-BYEU76JP.js +899 -0
  28. package/dist/chunk-CLJ5HLUD.js +458 -0
  29. package/dist/chunk-D5YD55AR.js +116 -0
  30. package/dist/chunk-DXQNI4PC.js +61 -0
  31. package/dist/chunk-E3NYWENM.js +1004 -0
  32. package/dist/chunk-GNGDQYDU.js +34688 -0
  33. package/dist/chunk-GOTUR54M.js +9 -0
  34. package/dist/chunk-HBU5MTAM.js +41 -0
  35. package/dist/chunk-HMYNFFY4.js +28 -0
  36. package/dist/chunk-JPOWPFCU.js +1010 -0
  37. package/dist/chunk-JWHCJDCI.js +1215 -0
  38. package/dist/chunk-KBR4MZZR.js +41 -0
  39. package/dist/chunk-KKKPTZLM.js +93 -0
  40. package/dist/chunk-ME6DNWIU.js +66 -0
  41. package/dist/chunk-NI4DEJMC.js +88 -0
  42. package/dist/chunk-O4EJEDHO.js +659 -0
  43. package/dist/chunk-PIDUD6M2.js +31 -0
  44. package/dist/chunk-PS4PFJQP.js +29459 -0
  45. package/dist/chunk-QV47YRF4.js +48 -0
  46. package/dist/chunk-RQDWMVRB.js +279 -0
  47. package/dist/chunk-TFSSEXL6.js +136 -0
  48. package/dist/chunk-TKHQ4DGZ.js +8290 -0
  49. package/dist/chunk-TPOCL34A.js +2876 -0
  50. package/dist/chunk-UGYAX5YI.js +565 -0
  51. package/dist/chunk-UHTSULZS.js +461 -0
  52. package/dist/chunk-UU3R62TT.js +128 -0
  53. package/dist/chunk-UWIJNAOB.js +3906 -0
  54. package/dist/chunk-VOO7NYPP.js +914 -0
  55. package/dist/chunk-VPAWTYLY.js +117 -0
  56. package/dist/chunk-WD6AJM35.js +1216 -0
  57. package/dist/chunk-X3BR7HWV.js +115 -0
  58. package/dist/chunk-X3NE4WVW.js +120 -0
  59. package/dist/chunk-XNISANGE.js +1395 -0
  60. package/dist/chunk-XV4ZJ6ZM.js +3177 -0
  61. package/dist/cli/index.js +236 -0
  62. package/dist/clio-KIQ5SNDS.js +53 -0
  63. package/dist/components-JVHMUBEB.js +653 -0
  64. package/dist/config-ZFCDBMDC.js +372 -0
  65. package/dist/configure-G4E3A2PG.js +27 -0
  66. package/dist/context-CDXTP2MP.js +293 -0
  67. package/dist/context-E3KIFVXI.js +185 -0
  68. package/dist/context-clear-3F4PLXOS.js +102 -0
  69. package/dist/context-index-Q7YSYTR3.js +106 -0
  70. package/dist/docs-YIETIWZI.js +280 -0
  71. package/dist/doctor-M5HJJZOL.js +61 -0
  72. package/dist/domains/agents/builtins/architect.md +33 -0
  73. package/dist/domains/agents/builtins/coder.md +31 -0
  74. package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
  75. package/dist/domains/agents/builtins/debugger.md +30 -0
  76. package/dist/domains/agents/builtins/documenter.md +31 -0
  77. package/dist/domains/agents/builtins/git-master.md +30 -0
  78. package/dist/domains/agents/builtins/provenance.md +30 -0
  79. package/dist/domains/agents/builtins/researcher.md +71 -0
  80. package/dist/domains/agents/builtins/scout.md +42 -0
  81. package/dist/domains/agents/builtins/tester.md +31 -0
  82. package/dist/domains/agents/builtins/verifier.md +30 -0
  83. package/dist/domains/agents/builtins/wiki-writer.md +41 -0
  84. package/dist/eval-B3KZZESM.js +2674 -0
  85. package/dist/evidence-V67CHM35.js +233 -0
  86. package/dist/evolve-YDZSUQYA.js +518 -0
  87. package/dist/extensions-SRG7XCAH.js +207 -0
  88. package/dist/fleet-CA2CRTVG.js +760 -0
  89. package/dist/fleet-preflight-CLIAX7YR.js +21 -0
  90. package/dist/init-2OZDJE2D.js +227 -0
  91. package/dist/memory-3PIQQAKX.js +207 -0
  92. package/dist/models-DY35XI7Y.js +237 -0
  93. package/dist/paths-5OMXW7Z4.js +57 -0
  94. package/dist/preload-KZVHET2B.js +11 -0
  95. package/dist/reset-PIFYNOS3.js +216 -0
  96. package/dist/run-3VSPP24F.js +735 -0
  97. package/dist/share-D36RQCXM.js +241 -0
  98. package/dist/skills-F2MRLELY.js +445 -0
  99. package/dist/skills-eval-E2ZTW4PL.js +932 -0
  100. package/dist/targets-DZMEZAH4.js +977 -0
  101. package/dist/trace-7NYCUI2J.js +250 -0
  102. package/dist/uninstall-AD3JWHBB.js +322 -0
  103. package/dist/upgrade-WYYBKGDY.js +301 -0
  104. package/dist/usage-ULIDAGFF.js +755 -0
  105. package/dist/version-ROZ6CZKH.js +16 -0
  106. package/dist/wiki-generate-PKFIX6OB.js +377 -0
  107. package/dist/worker/entry.js +1739 -0
  108. package/docs/README.md +93 -0
  109. package/docs/acp.md +120 -0
  110. package/docs/alcf-provider.md +72 -0
  111. package/docs/architecture.md +172 -0
  112. package/docs/artifact-versions.md +54 -0
  113. package/docs/built-in-agents.md +265 -0
  114. package/docs/capacity-and-scheduling.md +97 -0
  115. package/docs/commands-and-modes.md +554 -0
  116. package/docs/config-knobs-audit.md +115 -0
  117. package/docs/configuration-and-targets.md +812 -0
  118. package/docs/context-engine.md +236 -0
  119. package/docs/dispatch-architecture-rationale.md +126 -0
  120. package/docs/documentation-coverage.md +46 -0
  121. package/docs/documentation-guide.md +166 -0
  122. package/docs/environment-variables.md +105 -0
  123. package/docs/eval-runner.md +205 -0
  124. package/docs/evals-internal.md +298 -0
  125. package/docs/evidence-and-memory.md +243 -0
  126. package/docs/evolution.md +143 -0
  127. package/docs/exit-codes-and-output.md +74 -0
  128. package/docs/extensions-and-sharing.md +306 -0
  129. package/docs/fleet-demo-runbook.md +179 -0
  130. package/docs/fleet-dispatch.md +591 -0
  131. package/docs/glossary.md +75 -0
  132. package/docs/html/agents_blueprint.html +936 -0
  133. package/docs/html/alcf_blueprint.html +324 -0
  134. package/docs/html/architecture_blueprint.html +850 -0
  135. package/docs/html/commands_blueprint.html +794 -0
  136. package/docs/html/config_knobs_audit_blueprint.html +178 -0
  137. package/docs/html/configuration_blueprint.html +1080 -0
  138. package/docs/html/context_blueprint.html +603 -0
  139. package/docs/html/documentation_blueprint.html +832 -0
  140. package/docs/html/environment_blueprint.html +404 -0
  141. package/docs/html/eval_blueprint.html +743 -0
  142. package/docs/html/evals_internal_blueprint.html +190 -0
  143. package/docs/html/evolution_blueprint.html +674 -0
  144. package/docs/html/extensions_blueprint.html +2065 -0
  145. package/docs/html/fleet_dispatch_blueprint.html +286 -0
  146. package/docs/html/index.html +919 -0
  147. package/docs/html/lifecycle_blueprint.html +723 -0
  148. package/docs/html/memory_blueprint.html +699 -0
  149. package/docs/html/middleware_blueprint.html +664 -0
  150. package/docs/html/models_blueprint.html +2366 -0
  151. package/docs/html/observability_blueprint.html +683 -0
  152. package/docs/html/provider_adapter_blueprint.html +245 -0
  153. package/docs/html/safety_blueprint.html +1386 -0
  154. package/docs/html/shared.css +571 -0
  155. package/docs/html/shared.js +143 -0
  156. package/docs/html/skills_blueprint.html +671 -0
  157. package/docs/html/soak_blueprint.html +182 -0
  158. package/docs/html/tool_usage_blueprint.html +350 -0
  159. package/docs/html/tools_blueprint.html +2249 -0
  160. package/docs/html/trace_blueprint.html +235 -0
  161. package/docs/html/tui_design_blueprint.html +314 -0
  162. package/docs/html/validation_blueprint.html +961 -0
  163. package/docs/html/worker_dispatch_blueprint.html +231 -0
  164. package/docs/installation-and-lifecycle.md +308 -0
  165. package/docs/middleware-and-components.md +148 -0
  166. package/docs/model-catalog.md +189 -0
  167. package/docs/observability.md +233 -0
  168. package/docs/proactive-memory.md +452 -0
  169. package/docs/prompt-envelope-and-tools.md +142 -0
  170. package/docs/provider-adapter-cookbook.md +148 -0
  171. package/docs/release-cut-checklist.md +138 -0
  172. package/docs/safety-model.md +357 -0
  173. package/docs/scientific-validation.md +105 -0
  174. package/docs/session-lifecycle.md +156 -0
  175. package/docs/skills-marketplace.md +46 -0
  176. package/docs/tool-usage.md +527 -0
  177. package/docs/trace-store.md +132 -0
  178. package/docs/troubleshooting.md +33 -0
  179. package/docs/tui-design.md +239 -0
  180. package/docs/worker-dispatch-mechanics.md +242 -0
  181. package/package.json +132 -0
  182. package/skills/README.md +408 -0
  183. package/skills/git/commit-crafting/SKILL.md +79 -0
  184. package/skills/git/commit-crafting/evals.md +92 -0
  185. package/skills/git/create-pr/SKILL.md +116 -0
  186. package/skills/git/create-pr/evals.md +114 -0
  187. package/skills/git/investigate-issue/SKILL.md +139 -0
  188. package/skills/git/investigate-issue/evals.md +94 -0
  189. package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
  190. package/skills/git/resolve-merge-conflicts/evals.md +58 -0
  191. package/skills/git/review-changes/SKILL.md +103 -0
  192. package/skills/git/review-changes/evals.md +85 -0
  193. package/skills/git/worktree-create/SKILL.md +92 -0
  194. package/skills/git/worktree-create/evals.md +97 -0
  195. package/skills/git/worktree-create/references/worktree-setup.md +66 -0
  196. package/skills/git/worktree-merge/SKILL.md +95 -0
  197. package/skills/git/worktree-merge/evals.md +114 -0
  198. package/skills/skill-marketplace.json +261 -0
  199. package/skills/workflow/cut-it/SKILL.md +86 -0
  200. package/skills/workflow/cut-it/evals.md +42 -0
  201. package/src/domains/agents/builtins/architect.md +33 -0
  202. package/src/domains/agents/builtins/coder.md +31 -0
  203. package/src/domains/agents/builtins/context-bootstrap.md +38 -0
  204. package/src/domains/agents/builtins/debugger.md +30 -0
  205. package/src/domains/agents/builtins/documenter.md +31 -0
  206. package/src/domains/agents/builtins/git-master.md +30 -0
  207. package/src/domains/agents/builtins/provenance.md +30 -0
  208. package/src/domains/agents/builtins/researcher.md +71 -0
  209. package/src/domains/agents/builtins/scout.md +42 -0
  210. package/src/domains/agents/builtins/tester.md +31 -0
  211. package/src/domains/agents/builtins/verifier.md +30 -0
  212. package/src/domains/agents/builtins/wiki-writer.md +41 -0
  213. package/src/domains/agents/fleets/build-review.md +34 -0
  214. package/src/domains/agents/fleets/build-test.md +35 -0
  215. package/src/domains/agents/fleets/sdlc.md +86 -0
  216. package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
  217. package/src/domains/prompts/fragments/identity/clio.md +26 -0
  218. package/src/domains/prompts/fragments/operating/contract.md +64 -0
  219. package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
  220. package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
  221. package/src/domains/prompts/fragments/safety/read-only.md +13 -0
  222. package/src/domains/prompts/fragments/safety/suggest.md +13 -0
  223. package/src/domains/prompts/fragments/wiki/page.md +75 -0
  224. package/src/domains/prompts/fragments/wiki/plan.md +48 -0
  225. package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
  226. package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
@@ -0,0 +1,105 @@
1
+ # Environment Variables
2
+
3
+ Every environment variable the shipped `src/` tree reads, grouped by role. Settings.yaml is the durable home for operator policy; env vars exist for per-process overrides (CI, one-off experiments), directory layout, debugging, and internal plumbing. When prose and source disagree, prefer the source; the table cites the read site.
4
+
5
+ This page is the complete inventory, and `tests/contracts/environment-variable-inventory.test.ts` fails if `src/` reads a variable that has no row here.
6
+
7
+ > [!TIP]
8
+ > [docs/html/environment_blueprint.html](html/environment_blueprint.html) is a browsable walkthrough of the most commonly set variables with an effective-path resolver. It covers a curated subset, so use the tables below when you need the full list.
9
+
10
+ ## Guardrail overrides
11
+
12
+ Durable values live in the `guardrails:` section of settings.yaml (see [configuration-and-targets.md](configuration-and-targets.md)). These env vars override them for one process; resolution is env > settings > built-in default, and every value is a positive integer. Resolution lives in `src/core/guardrails.ts`.
13
+
14
+ | Variable | Settings key | Default | Controls |
15
+ | --- | --- | --- | --- |
16
+ | `CLIO_CODER_TURN_TOOL_CALL_BUDGET` | `guardrails.turnToolCallBudget` | 60 | Orchestrator per-turn soft tool-call budget; the hard interrupt ceiling sits 15 above it (`src/engine/loop-guard.ts`). |
17
+ | `CLIO_CODER_WORKER_TOOL_CALL_CAP` | `guardrails.workerToolCallCap` | 150 | Lifetime ceiling on tool calls one dispatched worker may execute. Calls the harness refused (reserve steering, synthesis-lockout denials) never spend it. Agent recipe budgets may narrow but never widen it (`src/engine/loop-guard.ts`). |
18
+ | `CLIO_CODER_MAX_RUNS` | `guardrails.maxDispatchRuns` | 1000 | Dispatch run-ledger retention cap (`src/domains/dispatch/state.ts`). |
19
+ | `CLIO_CODER_READ_MAX_BYTES` | `guardrails.readMaxBytes` | 51200 | Per-call byte cap for the read tool, floored at 1024 (`src/tools/read.ts`). |
20
+ | `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES` | `guardrails.observationTurnBudgetBytes` | 196608 | Shared per-turn byte pool across observation tools (`src/tools/observation.ts`). |
21
+ | `CLIO_CODER_INTERNAL_DISPATCH_TIMEOUT_MS` | `guardrails.internalDispatchTimeoutMs` | 900000 | Wall-clock cap for one internal generator dispatch: the wiki documenter and the bootstrap scout (`src/cli/internal-dispatch.ts`). |
22
+
23
+ ## Behavior knobs without a settings key
24
+
25
+ | Variable | Default | Controls |
26
+ | --- | --- | --- |
27
+ | `NO_COLOR` | unset | Set to any non-empty value to drop every foreground and background color. Bold, dim, italic, and underline stay, because they are what is left to read the interface by (`src/interactive/theme/tokens.ts`). |
28
+ | `CLIO_CODER_RIGOR` | repo-derived | Finish-contract evidence bar, `normal` or `high`, layered over the repo-derived default (`src/domains/safety/rigor.ts`). |
29
+ | `CLIO_CODER_RESIDENCY` | managed | `observe`/`off` stops Clio managing model residency on every local runtime path, llama.cpp routers included; per-target opt-out via `lifecycle: user-managed` (`src/engine/apis/residency.ts`). |
30
+ | `CLIO_CODER_TRUST_PROJECT_SKILLS` | off | `1` trusts project-local skills for execution (`src/domains/resources/skills/loader.ts`). |
31
+ | `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS` | off | `1` lets full-auto pass through to external CLI runtimes with their own full access (`src/engine/claude/subprocess-runtime.ts`, `src/engine/antigravity/subprocess-runtime.ts`). |
32
+ | `CLIO_CODER_FORCE_COMPACT` | off | `1` forces compaction on the next interactive turn (`src/interactive/chat-loop.ts`). |
33
+ | `CLIO_CODER_STATUS_STUCK_MS` | 180000 | Stuck-turn watchdog threshold (`src/interactive/status/watchdog.ts`). |
34
+ | `CLIO_CODER_SHUTDOWN_HOOK_MS` | 500 | Wall-clock budget per shutdown hook (`src/core/termination.ts`). |
35
+ | `CLIO_CODER_HOOK_BUDGET_MS` | per-phase built-ins | Global middleware hook wall-clock budget (`src/domains/middleware/budget.ts`). |
36
+ | `CLIO_CODER_HOOK_BUDGET_<PHASE>_MS` | per-phase built-ins | Per-phase hook budget, e.g. `CLIO_CODER_HOOK_BUDGET_TURN_END_MS`; beats the global var. |
37
+ | `CLIO_CODER_HOOK_BUDGET_WARMUP_CALLS` | 1 | Hook calls exempted from budget accounting at startup. |
38
+ | `CLIO_CODER_HOOK_BUDGET_WINDOW` | 5 | Sliding-window size for steady-state hook-budget warnings. |
39
+ | `CLIO_CODER_HOOK_BUDGET_THRESHOLD` | 3 | Overruns within the window before a steady-state warning. |
40
+ | `CLIO_CODER_LMSTUDIO_SDK_PREDICT` | off | `1` sends LM Studio predictions over the SDK again instead of its OpenAI-compatible port. Predictions moved to HTTP because the SDK surface ignores the thinking control, so this is an escape hatch back to the older transport and not a debug toggle. Listing, loading, and unloading always use the SDK (`src/engine/apis/lmstudio-native.ts`). |
41
+ | `CLIO_CODER_SKILL_CATALOG_DIR` | unset | Local skill-catalog directory override (`src/domains/resources/skills/marketplace.ts`). |
42
+ | `CLIO_CODER_SKILL_MARKETPLACE_INDEX` | unset | Skill-marketplace index path override (`src/domains/resources/skills/marketplace.ts`). |
43
+ | `CLIO_CODER_MODEL_CATALOG_DIRS` | unset | Extra model-catalog directories (`src/domains/providers/knowledge-base-path.ts`). |
44
+ | `CLIO_CODER_NO_NETWORK_TOOLS` | off | `1` strips network tools from every registry in the process; the skills-eval harness sets it for hermetic arms; `--allow-network` clears it (`src/tools/network-policy.ts`). |
45
+
46
+ ## Directory and install layout
47
+
48
+ | Variable | Default | Controls |
49
+ | --- | --- | --- |
50
+ | `CLIO_CODER_HOME` | unset | Single-tree install root; the per-role vars below beat it (`src/core/xdg.ts`). |
51
+ | `CLIO_CODER_CONFIG_DIR`, `CLIO_CODER_DATA_DIR`, `CLIO_CODER_STATE_DIR`, `CLIO_CODER_CACHE_DIR` | XDG platform defaults | Per-role directory overrides (`src/core/xdg.ts`). |
52
+ | `CLIO_CODER_BIN_DIR` | `~/.local/bin` | Launcher symlink location (`src/cli/uninstall.ts`). |
53
+ | `CLIO_CODER_PACKAGE_ROOT` | auto-detected | Package root for bundled-asset resolution (`src/core/package-root.ts`). |
54
+
55
+ ## Debug and trace toggles
56
+
57
+ All default off; enable with `1`.
58
+
59
+ | Variable | Controls |
60
+ | --- | --- |
61
+ | `CLIO_CODER_BUS_TRACE` | Event-bus channel tracing to stderr (`src/core/bus-trace.ts`). |
62
+ | `CLIO_CODER_TRACE_BOOT` | Boot-phase timing trace (`src/core/boot-trace.ts`). |
63
+ | `CLIO_CODER_TIMING` | Startup timing report (`src/entry/orchestrator.ts`). |
64
+ | `CLIO_CODER_DEBUG_SHUTDOWN` | Shutdown-path diagnostics (`src/core/termination.ts`). |
65
+ | `CLIO_CODER_DEBUG_LMSTUDIO` | LM Studio wire logging (`src/domains/providers/runtimes/common/lmstudio-logger.ts`). |
66
+ | `CLIO_CODER_RUNTIME_VERBOSE` | Verbose runtime logging (`src/engine/apis/lmstudio-native.ts`). |
67
+ | `CLIO_CODER_HOOK_BUDGET_DEBUG` | Per-overrun hook-budget diagnostics (`src/domains/middleware/runtime.ts`). |
68
+
69
+ ### File-writing traces
70
+
71
+ These two take a path, not `1`. Setting either to `1` writes a file named `1` in the working directory. Both are off when unset or empty, and both create parent directories.
72
+
73
+ | Variable | Contents | Controls |
74
+ | --- | --- | --- |
75
+ | `CLIO_CODER_RENDER_TRACE` | timing only | Per-frame render timing for the interactive TUI, truncated on open so one file is one session. Records frame durations and counts and no conversation text, which makes it the instrument for reproducing a frame-cost claim at a given terminal size (`src/interactive/render-trace.ts`). |
76
+ | `CLIO_CODER_MEMORY_TRACE` | conversation text | Proactive task-memory step envelopes, including up to 8000 characters of the text each step saw. This is content-bearing by construction, so the file carries whatever the session carried. Do not enable it on work you would not paste, and do not attach the file to a bug report without reading it first (`src/domains/memory/task-memory-trace.ts`). |
77
+
78
+ Example:
79
+
80
+ ```bash
81
+ CLIO_CODER_RENDER_TRACE=/tmp/clio-render.jsonl clio-coder
82
+ ```
83
+
84
+ ## Internal plumbing
85
+
86
+ Set by Clio for its own processes; not operator knobs.
87
+
88
+ | Variable | Purpose |
89
+ | --- | --- |
90
+ | `CLIO_CODER_INTERACTIVE` | Marks the interactive TUI process; scrubbed from bash-tool children so nested invocations do not inherit it (`src/cli/clio.ts`, `src/core/bash-exec.ts`). |
91
+ | `CLIO_CODER_RUN_OVERRIDES` | JSON envelope for run-scoped CLI options (`--max-context-tokens`, `--kv-cache-mode`, sampling flags). One typed variable instead of one env var per option; worker subprocesses inherit it (`src/core/run-overrides.ts`). |
92
+ | `CLIO_CODER_RESUME_SESSION_ID` | Session id handed across a self-restart; consumed and deleted at boot (`src/entry/orchestrator.ts`). |
93
+ | `CLIO_CODER_BOOTSTRAP_GENERATE_CHILD` | Marks the CLIO-CODER.md-generation child so it skips recursion (`src/domains/context/extension.ts`). |
94
+ | `CLIO_CODER_WORKER_LABELS` | Comma-separated labels a dispatched worker reports as its own (`src/domains/dispatch/transport.ts`, `src/worker/entry.ts`). |
95
+ | `CLIO_CODER_WORKER_PGID` | Process-group id the transport assigns a worker so its whole tree can be signalled (`src/domains/dispatch/transport.ts`, `src/worker/entry.ts`). |
96
+
97
+ ## Test-only
98
+
99
+ | Variable | Purpose |
100
+ | --- | --- |
101
+ | `CLIO_CODER_WORKER_FAUX` (+ `_MODEL`, `_TEXT`, `_STOP_REASON`, `_ERROR_MESSAGE`) | Fake worker model for tests (`src/engine/ai.ts`). |
102
+ | `CLIO_CODER_TEST_UPGRADE_NO_NETWORK` | Skips npm install during upgrade tests (`src/cli/upgrade.ts`). |
103
+ | `CLIO_CODER_REQUIRE_HOME_PREFIX` | Test guardrail: abort if resolved directories escape `CLIO_CODER_HOME` (`src/core/init.ts`). |
104
+
105
+ Variables used only by `scripts/` and `benchmarks/` harnesses (the `CLIO_CODER_LIVE_*` smoke-test family, benchmark fleet configuration, install-script inputs) are not part of the shipped runtime and are documented inline where they are consumed.
@@ -0,0 +1,205 @@
1
+ # Clio Coder Local Evaluation Runner
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.0).
5
+
6
+ The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
7
+
8
+ Source of truth: [src/domains/eval/](../src/domains/eval/) and [src/cli/eval.ts](../src/cli/eval.ts).
9
+
10
+ ---
11
+
12
+ ## CLI Commands
13
+
14
+ The CLI commands under `clio-coder eval` support running, validating, reporting, comparing, and gating evaluation suites.
15
+
16
+ ```bash
17
+ clio-coder eval validate --suite <suite.yaml>
18
+ clio-coder eval run --suite <suite.yaml> [--target <id>] [--model <id>] [--out <path>] [--clio-entry <path>]
19
+ clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--clio-entry <path>]
20
+ clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
21
+ clio-coder eval compare <baselineEvalId> <candidateEvalId>
22
+ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
23
+ ```
24
+
25
+ ### Command Roles
26
+ * **`validate`**: Validates the structure and constraints of a Suite v2 YAML file without executing it.
27
+ * **`run`**: Runs a Suite v2 (via `--suite`) or a compatibility v1 task file (via `--task-file`). Outputs a text summary and writes an eval artifact under `<dataDir>/evals/` and evidence under `<dataDir>/evidence/eval-<evalId>/`.
28
+ * **`report`**: Formats and prints a report from a saved `evalId`. Supports multiple `--format` outputs:
29
+ * `text` (default): Human-readable stdout summary.
30
+ * `json`: Raw JSON structure of the artifact.
31
+ * `md`: Markdown document with tables and summaries.
32
+ * `swe-jsonl`: Standardized JSONL format representing task runs (e.g. for SWE-bench comparisons).
33
+ * `junit`: XML report for CI/CD integration.
34
+ * **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
35
+ * **`gate`**: Compares candidate metrics against baseline or absolute thresholds, exiting non-zero if assertions fail (useful for PR gating).
36
+
37
+ Exit codes:
38
+
39
+ | Command | Success | Failure |
40
+ | --- | --- | --- |
41
+ | `eval validate` | `0` when validation passes | `2` for validation issues |
42
+ | `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
43
+ | `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
44
+ | `eval compare` | `0` when both artifacts load and compare succeeds | `1` if artifacts cannot be read, `2` for invalid ID |
45
+ | `eval gate` | `0` when all threshold assertions pass | `1` if assertions fail, `2` for config/invalid ID errors |
46
+
47
+ ---
48
+
49
+ ## Suite v2 Schema
50
+
51
+ Suite v2 files define matrix targets, workspaces, runner parameters, validation metrics, assertions, and path blocklists.
52
+
53
+ ```yaml
54
+ version: 2
55
+ suite:
56
+ id: "science-suite"
57
+ title: "Scientific Software Evaluation"
58
+ visibility: "local"
59
+ description: "Suite for verifying HPC integrations."
60
+ matrix:
61
+ targets:
62
+ - id: "local-gemini"
63
+ model: "gemini-3.5-flash"
64
+ - id: "local-claude"
65
+ model: "claude-sonnet-5"
66
+ repeats: 3
67
+ tasks:
68
+ - id: "fft-tolerance-check"
69
+ tags:
70
+ - numeric
71
+ - fast
72
+ workspace:
73
+ kind: "temp-copy" # local | git | temp-copy
74
+ path: "fixtures/fft-src"
75
+ excludes:
76
+ - "**/node_modules/**"
77
+ runner:
78
+ kind: "clio-run" # clio-run | context-index | context-init | external-command
79
+ prompt: "Optimize the FFT tolerance bounds in solver.ts"
80
+ timeoutMs: 60000
81
+ verify:
82
+ commands:
83
+ - "npm run test"
84
+ assertions:
85
+ - metric: "result.pass"
86
+ op: "eq"
87
+ value: true
88
+ - metric: "tokens.total"
89
+ op: "lt"
90
+ value: 15000
91
+ forbidPaths:
92
+ - "**/credentials.yaml"
93
+ metrics:
94
+ collect:
95
+ - "tokens.total"
96
+ - "latency.wallMs"
97
+ timeoutMs: 90000
98
+ ```
99
+
100
+ ### Schema Field Reference
101
+
102
+ | Field / Section | Sub-fields | Description |
103
+ | --- | --- | --- |
104
+ | `version` | - | Must equal `2`. |
105
+ | `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
106
+ | `matrix` | `targets[]`, `repeats` | Matrix of execution targets (specifying model and thinking flags) and the repetition count. |
107
+ | `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
108
+ | `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
109
+ | `verify` | `commands`, `assertions`, `forbidPaths` | Validation steps: shell commands, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
110
+ | `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
111
+
112
+ ---
113
+
114
+ ## Workspace Kinds
115
+ * **`local`**: Executes the task directly in the specified local path.
116
+ * **`git`**: Clones the repository from `url`, checks out the specified `commit` or `checkout` ref, and runs there.
117
+ * **`temp-copy`**: Copies the directory at `path` to a temporary workspace location before running. This prevents side-effects from polluting other task runs.
118
+
119
+ ---
120
+
121
+ ## Runner Kinds
122
+ * **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
123
+ * **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
124
+ * **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
125
+ * **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
126
+
127
+ ---
128
+
129
+ ## Metric Assertions
130
+
131
+ Metrics collected during runs can be validated automatically using the `verify.assertions` list. Supported operator fields (`op`) are:
132
+ * `lt` (less than)
133
+ * `lte` (less than or equal)
134
+ * `gt` (greater than)
135
+ * `gte` (greater than or equal)
136
+ * `eq` (equal)
137
+ * `neq` (not equal)
138
+
139
+ Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `result.pass`.
140
+
141
+ ---
142
+
143
+ ## Failure Classes
144
+
145
+ Evaluation tasks may fail with one of the following classes:
146
+
147
+ | Failure Class | Meaning |
148
+ | --- | --- |
149
+ | `setup_failed` | A setup command exited non-zero. |
150
+ | `verifier_failed` | A verifier command exited non-zero. |
151
+ | `timeout` | A setup or verifier command timed out. |
152
+ | `cwd_missing` | Resolved task cwd does not exist. |
153
+ | `command_error` | Reserved class for command spawn/system errors. |
154
+
155
+ ---
156
+
157
+ ## v1 Compatibility
158
+
159
+ Version 1 task files can still be run directly via:
160
+ ```bash
161
+ clio-coder eval run --task-file tasks.yaml [--repeat <n>]
162
+ ```
163
+ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
164
+ * Workspace kind: `local` (using task `cwd` as workspace path)
165
+ * Runner kind: `external-command` (executing task `setup` commands)
166
+ * Verify commands: Task `verifier` commands
167
+ * Timeout and tags mapped directly
168
+
169
+ ---
170
+
171
+ ## Token Accounting & Provenance
172
+
173
+ Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
174
+
175
+ 1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events watched by `token-stream.ts` / `createStreamInvariantFold`. This represents usage reported by the provider for assistant messages watched over the wire. On surfaces without stdout streaming (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
176
+ 2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
177
+
178
+ ### Fail-Closed Reporting
179
+ Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
180
+
181
+ ---
182
+
183
+ ## Eval Artifact Format (v4)
184
+
185
+ Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
186
+
187
+ ```typescript
188
+ export interface EvalArtifactV4 {
189
+ version: 4;
190
+ evalId: string;
191
+ suite: { id: string; hash: string };
192
+ clio: EvalClioProvenance;
193
+ environment: EvalEnvironmentProvenance;
194
+ matrix: { target: string; model: string | null; thinking: string | null };
195
+ summary: EvalArtifactSummaryV4;
196
+ results: EvalArtifactResultV4[];
197
+ }
198
+ ```
199
+
200
+ ---
201
+
202
+ ## Task Outcome Measurement (`verify.measure`)
203
+
204
+ Task outcome commands declared under `verify.measure` evaluate whether the model solved the workload and record metrics (`task.solved`, `task.exitCode`). A non-zero exit from `verify.measure` is recorded as data and **never fails the evaluation item**. Task solution outcome is a measurement, while only machinery invariant behavior operates as a gate.
205
+
@@ -0,0 +1,298 @@
1
+ # Internal Eval Suites
2
+
3
+ > [!TIP]
4
+ > **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.0).
5
+
6
+ Private suites should live outside this repository. Keep datasets, prompts,
7
+ live fleet coordinates, calibration outputs, and raw run artifacts in a private
8
+ checkout or object store.
9
+
10
+ Run a private suite from this source checkout with:
11
+
12
+ ```sh
13
+ npm run build
14
+ clio-coder eval run --suite <external-path> --clio-entry dist/cli/index.js
15
+ ```
16
+
17
+ Use `--out <dir>` when the artifact should be written outside the default Clio
18
+ data directory. Public summaries can be copied into
19
+ `benchmarks/results/<suite>/<run-id>/` only after they have been sanitized down
20
+ to `manifest.json` and `summary.json`.
21
+
22
+ ## Context Regression Seed
23
+
24
+ ```yaml
25
+ version: 2
26
+ suite:
27
+ id: internal-context-regression
28
+ title: Internal context regression
29
+ visibility: private
30
+ description: Private context index determinism and coverage regression.
31
+ provenance:
32
+ owner: internal-eval
33
+ source: private
34
+ matrix:
35
+ targets:
36
+ - id: local
37
+ repeats: 3
38
+ tasks:
39
+ - id: context-index-private-checkout
40
+ tags:
41
+ - internal
42
+ - context
43
+ - offline
44
+ workspace:
45
+ kind: local
46
+ path: .
47
+ excludes:
48
+ - .git
49
+ - node_modules
50
+ - dist
51
+ - coverage
52
+ runner:
53
+ kind: context-index
54
+ verify:
55
+ assertions:
56
+ - metric: context.indexedFiles
57
+ op: gt
58
+ value: 0
59
+ - metric: context.coverage
60
+ op: gt
61
+ value: 0
62
+ forbidPaths:
63
+ - data/private-dump
64
+ - runs/raw
65
+ metrics:
66
+ collect:
67
+ - context.indexedFiles
68
+ - context.coverage
69
+ - context.structuralHash
70
+ - context.digestTokens
71
+ - latency.wallMs
72
+ timeoutMs: 120000
73
+ thresholds:
74
+ fail:
75
+ - metric: result.pass
76
+ op: eq
77
+ value: false
78
+ - metric: context.coverage
79
+ op: lte
80
+ value: 0
81
+ ```
82
+
83
+ ## Model Matrix Seed
84
+
85
+ ```yaml
86
+ version: 2
87
+ suite:
88
+ id: internal-model-matrix
89
+ title: Internal model matrix
90
+ visibility: private
91
+ description: Private model target smoke matrix for headless Clio runs.
92
+ provenance:
93
+ owner: internal-eval
94
+ source: private
95
+ matrix:
96
+ targets:
97
+ - id: local-fast
98
+ model: example-fast-model
99
+ thinking: low
100
+ - id: local-deep
101
+ model: example-deep-model
102
+ thinking: high
103
+ repeats: 2
104
+ tasks:
105
+ - id: version-smoke
106
+ tags:
107
+ - internal
108
+ - matrix
109
+ - smoke
110
+ workspace:
111
+ kind: temp-copy
112
+ path: fixtures/version-smoke
113
+ excludes:
114
+ - node_modules
115
+ - dist
116
+ runner:
117
+ kind: external-command
118
+ commands:
119
+ - node -e "console.log('matrix smoke')"
120
+ verify:
121
+ assertions:
122
+ - metric: verifier.exitCode
123
+ op: eq
124
+ value: 0
125
+ metrics:
126
+ collect:
127
+ - latency.wallMs
128
+ - verifier.exitCode
129
+ timeoutMs: 30000
130
+ - id: tiny-clio-run
131
+ tags:
132
+ - internal
133
+ - matrix
134
+ - model-required
135
+ workspace:
136
+ kind: temp-copy
137
+ path: fixtures/tiny-repair
138
+ excludes:
139
+ - node_modules
140
+ - dist
141
+ runner:
142
+ kind: clio-run
143
+ prompt: Fix the intentionally broken function so the local verifier passes.
144
+ verify:
145
+ commands:
146
+ - npm test
147
+ assertions:
148
+ - metric: result.pass
149
+ op: eq
150
+ value: true
151
+ metrics:
152
+ collect:
153
+ - tokens.input
154
+ - tokens.output
155
+ - tokens.total
156
+ - latency.wallMs
157
+ - patch.filesChanged
158
+ - verifier.exitCode
159
+ timeoutMs: 600000
160
+ thresholds:
161
+ fail:
162
+ - metric: result.pass
163
+ op: eq
164
+ value: false
165
+ - metric: tools.blocked
166
+ op: gt
167
+ value: 0
168
+ ```
169
+
170
+ ## Live Turns Seed
171
+
172
+ ```yaml
173
+ version: 2
174
+ suite:
175
+ id: internal-live-turns
176
+ title: Internal live turns
177
+ visibility: private
178
+ description: Private live-fleet prompts for end-to-end Clio turn checks.
179
+ provenance:
180
+ owner: internal-eval
181
+ source: private
182
+ matrix:
183
+ targets:
184
+ - id: live-orchestrator
185
+ model: example-live-model
186
+ thinking: low
187
+ repeats: 1
188
+ tasks:
189
+ - id: readonly-repo-summary
190
+ tags:
191
+ - internal
192
+ - live
193
+ - readonly
194
+ - model-required
195
+ workspace:
196
+ kind: temp-copy
197
+ path: fixtures/live-readonly
198
+ excludes:
199
+ - node_modules
200
+ - dist
201
+ - .clio-coder
202
+ runner:
203
+ kind: clio-run
204
+ prompt: Summarize the repository purpose and make no file changes.
205
+ verify:
206
+ forbidPaths:
207
+ - unexpected-output.txt
208
+ assertions:
209
+ - metric: result.pass
210
+ op: eq
211
+ value: true
212
+ - metric: patch.filesChanged
213
+ op: eq
214
+ value: 0
215
+ metrics:
216
+ collect:
217
+ - tokens.total
218
+ - latency.wallMs
219
+ - tools.totalCalls
220
+ - tools.failed
221
+ - patch.filesChanged
222
+ timeoutMs: 300000
223
+ - id: live-small-edit
224
+ tags:
225
+ - internal
226
+ - live
227
+ - edit
228
+ - model-required
229
+ workspace:
230
+ kind: temp-copy
231
+ path: fixtures/live-small-edit
232
+ excludes:
233
+ - node_modules
234
+ - dist
235
+ - .clio-coder
236
+ runner:
237
+ kind: clio-run
238
+ prompt: Fix the failing unit test with the smallest source change.
239
+ verify:
240
+ commands:
241
+ - npm test
242
+ assertions:
243
+ - metric: result.pass
244
+ op: eq
245
+ value: true
246
+ - metric: patch.testFilesModified
247
+ op: eq
248
+ value: 0
249
+ metrics:
250
+ collect:
251
+ - tokens.input
252
+ - tokens.output
253
+ - tokens.total
254
+ - latency.wallMs
255
+ - tools.totalCalls
256
+ - tools.failed
257
+ - patch.filesChanged
258
+ - patch.testFilesModified
259
+ - verifier.exitCode
260
+ timeoutMs: 600000
261
+ thresholds:
262
+ fail:
263
+ - metric: result.pass
264
+ op: eq
265
+ value: false
266
+ - metric: tools.failed
267
+ op: gt
268
+ value: 0
269
+ ```
270
+
271
+ ---
272
+
273
+ ## Soak Benchmark Suite
274
+
275
+ The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
276
+
277
+ The soak suite comprises four specialized suite files:
278
+
279
+ ### 1. Machinery Under Load (`clio-soak.yaml`)
280
+ Evaluates the same task workload across two execution surfaces: the headless main-agent surface (`clio-run`) and a dispatched worker surface (`agent: coder`). It tests single-file bugs, multi-file bugs, and compaction continuity across restarts.
281
+ - **Surface Differences**: Main-agent tasks verify session ledger continuity (`ledger.formatVersion`, `ledger.toolPairsUnmatched`, `ledger.assistantBetweenCallAndResult`), while dispatch worker tasks verify process group cleanup (`process.orphanedChildren == 0`).
282
+ - **Compaction Continuity**: Verifies that compaction summaries are present (`continuity.compactionSummaryPresent`) and that pre-compaction facts are preserved (`continuity.answeredFromPreCompaction`).
283
+ - **Suite-Wide Gates**: Gates on `receipt.sealed`, `receipt.integrityValid`, `receipt.outcomeMatchesExit`, `tokens.measured`, `stream.cumulativeSnapshots == 0`, `stream.usageDoubleCounted == false`, and `stream.segmentUsageMatchesMessages == true`.
284
+
285
+ ### 2. Per-Step Write Boundaries (`clio-soak-boundary.yaml`)
286
+ Validates write boundary enforcement across steps without model participation. Enforcement is strictly detect-and-rollback and is never sandboxing.
287
+ - `write-boundary.rolled-back`: Verifies clean detection of allowlist violations (`writes_boundary_violation`), git-level file restoration, and sealed verdict generation (`boundary.violationsRolledBack == 1`, `boundary.rollbackIncomplete == 0`).
288
+ - `write-boundary.rollback-incomplete`: Tests honest failure reporting when a path was dirty prior to snapshot taking. prior bytes exist only in the overwritten tree, so rollback leaves the tree unchanged and records incomplete rollback (`boundary.rollbackIncomplete == 1`, `boundary.violationsRolledBack == 0`).
289
+
290
+ ### 3. Fault Injection Chaos (`clio-soak-chaos.yaml`)
291
+ Evaluates system resilience against process signals.
292
+ - `chaos.sigint-mid-tool`: Prompts Clio for a long-running bash tool call and injects `SIGINT` once the subprocess initializes. Asserts exit code `130`, confirms no orphaned children remain (`process.orphanedChildren == 0`), and verifies receipt sealing, receipt integrity, and provider token reporting.
293
+
294
+ ### 4. Bounded Loops (`clio-soak-loop.yaml`)
295
+ Validates iteration bounds and receipt accounting for fleet loops (`bounded-loop.fleet`).
296
+ - **Loop Bounds**: Asserts that verification attempts do not exceed declared limits (`loop.attemptsSpent <= 3`), recovery attempts seal individual receipts (`loop.receiptsMatchRepairs == true`), and unneeded nodes report as `unneeded` rather than skipped or failed (`loop.skippedNodes == 0`).
297
+ - **Two Token Accountings**: Distinguishes `tokens.*` (folded live off wire stdout by `createStreamInvariantFold`) from `receiptUsage.*` (journal receipts sealed and authenticated against ledger envelopes). On fleet runs, wire streaming is absent (`tokens.measured == false`), while journal receipts provide authenticated usage (`receiptUsage.measured == true`).
298
+