@iowarp/clio-coder 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +407 -0
- package/CODE_OF_CONDUCT.md +21 -0
- package/CONTRIBUTING.md +224 -0
- package/LICENSE +202 -0
- package/NOTICE +9 -0
- package/README.md +798 -0
- package/SECURITY.md +72 -0
- package/assets/clio-coder-logo-128.webp +0 -0
- package/damage-control-rules.yaml +419 -0
- package/dist/acp-UMLFVA3F.js +92 -0
- package/dist/agents-Q4MYPMUW.js +91 -0
- package/dist/auth-O6HYIJ6J.js +521 -0
- package/dist/chunk-262G75JS.js +35 -0
- package/dist/chunk-26BZQOAD.js +1281 -0
- package/dist/chunk-2J63S4SF.js +508 -0
- package/dist/chunk-3DANZDGR.js +717 -0
- package/dist/chunk-4UQA7NCT.js +29 -0
- package/dist/chunk-527KG6XR.js +497 -0
- package/dist/chunk-5LDRNKX2.js +1063 -0
- package/dist/chunk-5N2FG33Q.js +25 -0
- package/dist/chunk-67MTHP2E.js +135 -0
- package/dist/chunk-6CWDTGUC.js +20 -0
- package/dist/chunk-7BHLZB3A.js +2115 -0
- package/dist/chunk-7RBKDI66.js +348 -0
- package/dist/chunk-AMFR5YA3.js +541 -0
- package/dist/chunk-BBUH4VAA.js +1224 -0
- package/dist/chunk-BYEU76JP.js +899 -0
- package/dist/chunk-CLJ5HLUD.js +458 -0
- package/dist/chunk-D5YD55AR.js +116 -0
- package/dist/chunk-DXQNI4PC.js +61 -0
- package/dist/chunk-E3NYWENM.js +1004 -0
- package/dist/chunk-GNGDQYDU.js +34688 -0
- package/dist/chunk-GOTUR54M.js +9 -0
- package/dist/chunk-HBU5MTAM.js +41 -0
- package/dist/chunk-HMYNFFY4.js +28 -0
- package/dist/chunk-JPOWPFCU.js +1010 -0
- package/dist/chunk-JWHCJDCI.js +1215 -0
- package/dist/chunk-KBR4MZZR.js +41 -0
- package/dist/chunk-KKKPTZLM.js +93 -0
- package/dist/chunk-ME6DNWIU.js +66 -0
- package/dist/chunk-NI4DEJMC.js +88 -0
- package/dist/chunk-O4EJEDHO.js +659 -0
- package/dist/chunk-PIDUD6M2.js +31 -0
- package/dist/chunk-PS4PFJQP.js +29459 -0
- package/dist/chunk-QV47YRF4.js +48 -0
- package/dist/chunk-RQDWMVRB.js +279 -0
- package/dist/chunk-TFSSEXL6.js +136 -0
- package/dist/chunk-TKHQ4DGZ.js +8290 -0
- package/dist/chunk-TPOCL34A.js +2876 -0
- package/dist/chunk-UGYAX5YI.js +565 -0
- package/dist/chunk-UHTSULZS.js +461 -0
- package/dist/chunk-UU3R62TT.js +128 -0
- package/dist/chunk-UWIJNAOB.js +3906 -0
- package/dist/chunk-VOO7NYPP.js +914 -0
- package/dist/chunk-VPAWTYLY.js +117 -0
- package/dist/chunk-WD6AJM35.js +1216 -0
- package/dist/chunk-X3BR7HWV.js +115 -0
- package/dist/chunk-X3NE4WVW.js +120 -0
- package/dist/chunk-XNISANGE.js +1395 -0
- package/dist/chunk-XV4ZJ6ZM.js +3177 -0
- package/dist/cli/index.js +236 -0
- package/dist/clio-KIQ5SNDS.js +53 -0
- package/dist/components-JVHMUBEB.js +653 -0
- package/dist/config-ZFCDBMDC.js +372 -0
- package/dist/configure-G4E3A2PG.js +27 -0
- package/dist/context-CDXTP2MP.js +293 -0
- package/dist/context-E3KIFVXI.js +185 -0
- package/dist/context-clear-3F4PLXOS.js +102 -0
- package/dist/context-index-Q7YSYTR3.js +106 -0
- package/dist/docs-YIETIWZI.js +280 -0
- package/dist/doctor-M5HJJZOL.js +61 -0
- package/dist/domains/agents/builtins/architect.md +33 -0
- package/dist/domains/agents/builtins/coder.md +31 -0
- package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
- package/dist/domains/agents/builtins/debugger.md +30 -0
- package/dist/domains/agents/builtins/documenter.md +31 -0
- package/dist/domains/agents/builtins/git-master.md +30 -0
- package/dist/domains/agents/builtins/provenance.md +30 -0
- package/dist/domains/agents/builtins/researcher.md +71 -0
- package/dist/domains/agents/builtins/scout.md +42 -0
- package/dist/domains/agents/builtins/tester.md +31 -0
- package/dist/domains/agents/builtins/verifier.md +30 -0
- package/dist/domains/agents/builtins/wiki-writer.md +41 -0
- package/dist/eval-B3KZZESM.js +2674 -0
- package/dist/evidence-V67CHM35.js +233 -0
- package/dist/evolve-YDZSUQYA.js +518 -0
- package/dist/extensions-SRG7XCAH.js +207 -0
- package/dist/fleet-CA2CRTVG.js +760 -0
- package/dist/fleet-preflight-CLIAX7YR.js +21 -0
- package/dist/init-2OZDJE2D.js +227 -0
- package/dist/memory-3PIQQAKX.js +207 -0
- package/dist/models-DY35XI7Y.js +237 -0
- package/dist/paths-5OMXW7Z4.js +57 -0
- package/dist/preload-KZVHET2B.js +11 -0
- package/dist/reset-PIFYNOS3.js +216 -0
- package/dist/run-3VSPP24F.js +735 -0
- package/dist/share-D36RQCXM.js +241 -0
- package/dist/skills-F2MRLELY.js +445 -0
- package/dist/skills-eval-E2ZTW4PL.js +932 -0
- package/dist/targets-DZMEZAH4.js +977 -0
- package/dist/trace-7NYCUI2J.js +250 -0
- package/dist/uninstall-AD3JWHBB.js +322 -0
- package/dist/upgrade-WYYBKGDY.js +301 -0
- package/dist/usage-ULIDAGFF.js +755 -0
- package/dist/version-ROZ6CZKH.js +16 -0
- package/dist/wiki-generate-PKFIX6OB.js +377 -0
- package/dist/worker/entry.js +1739 -0
- package/docs/README.md +93 -0
- package/docs/acp.md +120 -0
- package/docs/alcf-provider.md +72 -0
- package/docs/architecture.md +172 -0
- package/docs/artifact-versions.md +54 -0
- package/docs/built-in-agents.md +265 -0
- package/docs/capacity-and-scheduling.md +97 -0
- package/docs/commands-and-modes.md +554 -0
- package/docs/config-knobs-audit.md +115 -0
- package/docs/configuration-and-targets.md +812 -0
- package/docs/context-engine.md +236 -0
- package/docs/dispatch-architecture-rationale.md +126 -0
- package/docs/documentation-coverage.md +46 -0
- package/docs/documentation-guide.md +166 -0
- package/docs/environment-variables.md +105 -0
- package/docs/eval-runner.md +205 -0
- package/docs/evals-internal.md +298 -0
- package/docs/evidence-and-memory.md +243 -0
- package/docs/evolution.md +143 -0
- package/docs/exit-codes-and-output.md +74 -0
- package/docs/extensions-and-sharing.md +306 -0
- package/docs/fleet-demo-runbook.md +179 -0
- package/docs/fleet-dispatch.md +591 -0
- package/docs/glossary.md +75 -0
- package/docs/html/agents_blueprint.html +936 -0
- package/docs/html/alcf_blueprint.html +324 -0
- package/docs/html/architecture_blueprint.html +850 -0
- package/docs/html/commands_blueprint.html +794 -0
- package/docs/html/config_knobs_audit_blueprint.html +178 -0
- package/docs/html/configuration_blueprint.html +1080 -0
- package/docs/html/context_blueprint.html +603 -0
- package/docs/html/documentation_blueprint.html +832 -0
- package/docs/html/environment_blueprint.html +404 -0
- package/docs/html/eval_blueprint.html +743 -0
- package/docs/html/evals_internal_blueprint.html +190 -0
- package/docs/html/evolution_blueprint.html +674 -0
- package/docs/html/extensions_blueprint.html +2065 -0
- package/docs/html/fleet_dispatch_blueprint.html +286 -0
- package/docs/html/index.html +919 -0
- package/docs/html/lifecycle_blueprint.html +723 -0
- package/docs/html/memory_blueprint.html +699 -0
- package/docs/html/middleware_blueprint.html +664 -0
- package/docs/html/models_blueprint.html +2366 -0
- package/docs/html/observability_blueprint.html +683 -0
- package/docs/html/provider_adapter_blueprint.html +245 -0
- package/docs/html/safety_blueprint.html +1386 -0
- package/docs/html/shared.css +571 -0
- package/docs/html/shared.js +143 -0
- package/docs/html/skills_blueprint.html +671 -0
- package/docs/html/soak_blueprint.html +182 -0
- package/docs/html/tool_usage_blueprint.html +350 -0
- package/docs/html/tools_blueprint.html +2249 -0
- package/docs/html/trace_blueprint.html +235 -0
- package/docs/html/tui_design_blueprint.html +314 -0
- package/docs/html/validation_blueprint.html +961 -0
- package/docs/html/worker_dispatch_blueprint.html +231 -0
- package/docs/installation-and-lifecycle.md +308 -0
- package/docs/middleware-and-components.md +148 -0
- package/docs/model-catalog.md +189 -0
- package/docs/observability.md +233 -0
- package/docs/proactive-memory.md +452 -0
- package/docs/prompt-envelope-and-tools.md +142 -0
- package/docs/provider-adapter-cookbook.md +148 -0
- package/docs/release-cut-checklist.md +138 -0
- package/docs/safety-model.md +357 -0
- package/docs/scientific-validation.md +105 -0
- package/docs/session-lifecycle.md +156 -0
- package/docs/skills-marketplace.md +46 -0
- package/docs/tool-usage.md +527 -0
- package/docs/trace-store.md +132 -0
- package/docs/troubleshooting.md +33 -0
- package/docs/tui-design.md +239 -0
- package/docs/worker-dispatch-mechanics.md +242 -0
- package/package.json +132 -0
- package/skills/README.md +408 -0
- package/skills/git/commit-crafting/SKILL.md +79 -0
- package/skills/git/commit-crafting/evals.md +92 -0
- package/skills/git/create-pr/SKILL.md +116 -0
- package/skills/git/create-pr/evals.md +114 -0
- package/skills/git/investigate-issue/SKILL.md +139 -0
- package/skills/git/investigate-issue/evals.md +94 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
- package/skills/git/resolve-merge-conflicts/evals.md +58 -0
- package/skills/git/review-changes/SKILL.md +103 -0
- package/skills/git/review-changes/evals.md +85 -0
- package/skills/git/worktree-create/SKILL.md +92 -0
- package/skills/git/worktree-create/evals.md +97 -0
- package/skills/git/worktree-create/references/worktree-setup.md +66 -0
- package/skills/git/worktree-merge/SKILL.md +95 -0
- package/skills/git/worktree-merge/evals.md +114 -0
- package/skills/skill-marketplace.json +261 -0
- package/skills/workflow/cut-it/SKILL.md +86 -0
- package/skills/workflow/cut-it/evals.md +42 -0
- package/src/domains/agents/builtins/architect.md +33 -0
- package/src/domains/agents/builtins/coder.md +31 -0
- package/src/domains/agents/builtins/context-bootstrap.md +38 -0
- package/src/domains/agents/builtins/debugger.md +30 -0
- package/src/domains/agents/builtins/documenter.md +31 -0
- package/src/domains/agents/builtins/git-master.md +30 -0
- package/src/domains/agents/builtins/provenance.md +30 -0
- package/src/domains/agents/builtins/researcher.md +71 -0
- package/src/domains/agents/builtins/scout.md +42 -0
- package/src/domains/agents/builtins/tester.md +31 -0
- package/src/domains/agents/builtins/verifier.md +30 -0
- package/src/domains/agents/builtins/wiki-writer.md +41 -0
- package/src/domains/agents/fleets/build-review.md +34 -0
- package/src/domains/agents/fleets/build-test.md +35 -0
- package/src/domains/agents/fleets/sdlc.md +86 -0
- package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
- package/src/domains/prompts/fragments/identity/clio.md +26 -0
- package/src/domains/prompts/fragments/operating/contract.md +64 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
- package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
- package/src/domains/prompts/fragments/safety/read-only.md +13 -0
- package/src/domains/prompts/fragments/safety/suggest.md +13 -0
- package/src/domains/prompts/fragments/wiki/page.md +75 -0
- package/src/domains/prompts/fragments/wiki/plan.md +48 -0
- package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
|
@@ -0,0 +1,105 @@
|
|
|
1
|
+
# Environment Variables
|
|
2
|
+
|
|
3
|
+
Every environment variable the shipped `src/` tree reads, grouped by role. Settings.yaml is the durable home for operator policy; env vars exist for per-process overrides (CI, one-off experiments), directory layout, debugging, and internal plumbing. When prose and source disagree, prefer the source; the table cites the read site.
|
|
4
|
+
|
|
5
|
+
This page is the complete inventory, and `tests/contracts/environment-variable-inventory.test.ts` fails if `src/` reads a variable that has no row here.
|
|
6
|
+
|
|
7
|
+
> [!TIP]
|
|
8
|
+
> [docs/html/environment_blueprint.html](html/environment_blueprint.html) is a browsable walkthrough of the most commonly set variables with an effective-path resolver. It covers a curated subset, so use the tables below when you need the full list.
|
|
9
|
+
|
|
10
|
+
## Guardrail overrides
|
|
11
|
+
|
|
12
|
+
Durable values live in the `guardrails:` section of settings.yaml (see [configuration-and-targets.md](configuration-and-targets.md)). These env vars override them for one process; resolution is env > settings > built-in default, and every value is a positive integer. Resolution lives in `src/core/guardrails.ts`.
|
|
13
|
+
|
|
14
|
+
| Variable | Settings key | Default | Controls |
|
|
15
|
+
| --- | --- | --- | --- |
|
|
16
|
+
| `CLIO_CODER_TURN_TOOL_CALL_BUDGET` | `guardrails.turnToolCallBudget` | 60 | Orchestrator per-turn soft tool-call budget; the hard interrupt ceiling sits 15 above it (`src/engine/loop-guard.ts`). |
|
|
17
|
+
| `CLIO_CODER_WORKER_TOOL_CALL_CAP` | `guardrails.workerToolCallCap` | 150 | Lifetime ceiling on tool calls one dispatched worker may execute. Calls the harness refused (reserve steering, synthesis-lockout denials) never spend it. Agent recipe budgets may narrow but never widen it (`src/engine/loop-guard.ts`). |
|
|
18
|
+
| `CLIO_CODER_MAX_RUNS` | `guardrails.maxDispatchRuns` | 1000 | Dispatch run-ledger retention cap (`src/domains/dispatch/state.ts`). |
|
|
19
|
+
| `CLIO_CODER_READ_MAX_BYTES` | `guardrails.readMaxBytes` | 51200 | Per-call byte cap for the read tool, floored at 1024 (`src/tools/read.ts`). |
|
|
20
|
+
| `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES` | `guardrails.observationTurnBudgetBytes` | 196608 | Shared per-turn byte pool across observation tools (`src/tools/observation.ts`). |
|
|
21
|
+
| `CLIO_CODER_INTERNAL_DISPATCH_TIMEOUT_MS` | `guardrails.internalDispatchTimeoutMs` | 900000 | Wall-clock cap for one internal generator dispatch: the wiki documenter and the bootstrap scout (`src/cli/internal-dispatch.ts`). |
|
|
22
|
+
|
|
23
|
+
## Behavior knobs without a settings key
|
|
24
|
+
|
|
25
|
+
| Variable | Default | Controls |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `NO_COLOR` | unset | Set to any non-empty value to drop every foreground and background color. Bold, dim, italic, and underline stay, because they are what is left to read the interface by (`src/interactive/theme/tokens.ts`). |
|
|
28
|
+
| `CLIO_CODER_RIGOR` | repo-derived | Finish-contract evidence bar, `normal` or `high`, layered over the repo-derived default (`src/domains/safety/rigor.ts`). |
|
|
29
|
+
| `CLIO_CODER_RESIDENCY` | managed | `observe`/`off` stops Clio managing model residency on every local runtime path, llama.cpp routers included; per-target opt-out via `lifecycle: user-managed` (`src/engine/apis/residency.ts`). |
|
|
30
|
+
| `CLIO_CODER_TRUST_PROJECT_SKILLS` | off | `1` trusts project-local skills for execution (`src/domains/resources/skills/loader.ts`). |
|
|
31
|
+
| `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS` | off | `1` lets full-auto pass through to external CLI runtimes with their own full access (`src/engine/claude/subprocess-runtime.ts`, `src/engine/antigravity/subprocess-runtime.ts`). |
|
|
32
|
+
| `CLIO_CODER_FORCE_COMPACT` | off | `1` forces compaction on the next interactive turn (`src/interactive/chat-loop.ts`). |
|
|
33
|
+
| `CLIO_CODER_STATUS_STUCK_MS` | 180000 | Stuck-turn watchdog threshold (`src/interactive/status/watchdog.ts`). |
|
|
34
|
+
| `CLIO_CODER_SHUTDOWN_HOOK_MS` | 500 | Wall-clock budget per shutdown hook (`src/core/termination.ts`). |
|
|
35
|
+
| `CLIO_CODER_HOOK_BUDGET_MS` | per-phase built-ins | Global middleware hook wall-clock budget (`src/domains/middleware/budget.ts`). |
|
|
36
|
+
| `CLIO_CODER_HOOK_BUDGET_<PHASE>_MS` | per-phase built-ins | Per-phase hook budget, e.g. `CLIO_CODER_HOOK_BUDGET_TURN_END_MS`; beats the global var. |
|
|
37
|
+
| `CLIO_CODER_HOOK_BUDGET_WARMUP_CALLS` | 1 | Hook calls exempted from budget accounting at startup. |
|
|
38
|
+
| `CLIO_CODER_HOOK_BUDGET_WINDOW` | 5 | Sliding-window size for steady-state hook-budget warnings. |
|
|
39
|
+
| `CLIO_CODER_HOOK_BUDGET_THRESHOLD` | 3 | Overruns within the window before a steady-state warning. |
|
|
40
|
+
| `CLIO_CODER_LMSTUDIO_SDK_PREDICT` | off | `1` sends LM Studio predictions over the SDK again instead of its OpenAI-compatible port. Predictions moved to HTTP because the SDK surface ignores the thinking control, so this is an escape hatch back to the older transport and not a debug toggle. Listing, loading, and unloading always use the SDK (`src/engine/apis/lmstudio-native.ts`). |
|
|
41
|
+
| `CLIO_CODER_SKILL_CATALOG_DIR` | unset | Local skill-catalog directory override (`src/domains/resources/skills/marketplace.ts`). |
|
|
42
|
+
| `CLIO_CODER_SKILL_MARKETPLACE_INDEX` | unset | Skill-marketplace index path override (`src/domains/resources/skills/marketplace.ts`). |
|
|
43
|
+
| `CLIO_CODER_MODEL_CATALOG_DIRS` | unset | Extra model-catalog directories (`src/domains/providers/knowledge-base-path.ts`). |
|
|
44
|
+
| `CLIO_CODER_NO_NETWORK_TOOLS` | off | `1` strips network tools from every registry in the process; the skills-eval harness sets it for hermetic arms; `--allow-network` clears it (`src/tools/network-policy.ts`). |
|
|
45
|
+
|
|
46
|
+
## Directory and install layout
|
|
47
|
+
|
|
48
|
+
| Variable | Default | Controls |
|
|
49
|
+
| --- | --- | --- |
|
|
50
|
+
| `CLIO_CODER_HOME` | unset | Single-tree install root; the per-role vars below beat it (`src/core/xdg.ts`). |
|
|
51
|
+
| `CLIO_CODER_CONFIG_DIR`, `CLIO_CODER_DATA_DIR`, `CLIO_CODER_STATE_DIR`, `CLIO_CODER_CACHE_DIR` | XDG platform defaults | Per-role directory overrides (`src/core/xdg.ts`). |
|
|
52
|
+
| `CLIO_CODER_BIN_DIR` | `~/.local/bin` | Launcher symlink location (`src/cli/uninstall.ts`). |
|
|
53
|
+
| `CLIO_CODER_PACKAGE_ROOT` | auto-detected | Package root for bundled-asset resolution (`src/core/package-root.ts`). |
|
|
54
|
+
|
|
55
|
+
## Debug and trace toggles
|
|
56
|
+
|
|
57
|
+
All default off; enable with `1`.
|
|
58
|
+
|
|
59
|
+
| Variable | Controls |
|
|
60
|
+
| --- | --- |
|
|
61
|
+
| `CLIO_CODER_BUS_TRACE` | Event-bus channel tracing to stderr (`src/core/bus-trace.ts`). |
|
|
62
|
+
| `CLIO_CODER_TRACE_BOOT` | Boot-phase timing trace (`src/core/boot-trace.ts`). |
|
|
63
|
+
| `CLIO_CODER_TIMING` | Startup timing report (`src/entry/orchestrator.ts`). |
|
|
64
|
+
| `CLIO_CODER_DEBUG_SHUTDOWN` | Shutdown-path diagnostics (`src/core/termination.ts`). |
|
|
65
|
+
| `CLIO_CODER_DEBUG_LMSTUDIO` | LM Studio wire logging (`src/domains/providers/runtimes/common/lmstudio-logger.ts`). |
|
|
66
|
+
| `CLIO_CODER_RUNTIME_VERBOSE` | Verbose runtime logging (`src/engine/apis/lmstudio-native.ts`). |
|
|
67
|
+
| `CLIO_CODER_HOOK_BUDGET_DEBUG` | Per-overrun hook-budget diagnostics (`src/domains/middleware/runtime.ts`). |
|
|
68
|
+
|
|
69
|
+
### File-writing traces
|
|
70
|
+
|
|
71
|
+
These two take a path, not `1`. Setting either to `1` writes a file named `1` in the working directory. Both are off when unset or empty, and both create parent directories.
|
|
72
|
+
|
|
73
|
+
| Variable | Contents | Controls |
|
|
74
|
+
| --- | --- | --- |
|
|
75
|
+
| `CLIO_CODER_RENDER_TRACE` | timing only | Per-frame render timing for the interactive TUI, truncated on open so one file is one session. Records frame durations and counts and no conversation text, which makes it the instrument for reproducing a frame-cost claim at a given terminal size (`src/interactive/render-trace.ts`). |
|
|
76
|
+
| `CLIO_CODER_MEMORY_TRACE` | conversation text | Proactive task-memory step envelopes, including up to 8000 characters of the text each step saw. This is content-bearing by construction, so the file carries whatever the session carried. Do not enable it on work you would not paste, and do not attach the file to a bug report without reading it first (`src/domains/memory/task-memory-trace.ts`). |
|
|
77
|
+
|
|
78
|
+
Example:
|
|
79
|
+
|
|
80
|
+
```bash
|
|
81
|
+
CLIO_CODER_RENDER_TRACE=/tmp/clio-render.jsonl clio-coder
|
|
82
|
+
```
|
|
83
|
+
|
|
84
|
+
## Internal plumbing
|
|
85
|
+
|
|
86
|
+
Set by Clio for its own processes; not operator knobs.
|
|
87
|
+
|
|
88
|
+
| Variable | Purpose |
|
|
89
|
+
| --- | --- |
|
|
90
|
+
| `CLIO_CODER_INTERACTIVE` | Marks the interactive TUI process; scrubbed from bash-tool children so nested invocations do not inherit it (`src/cli/clio.ts`, `src/core/bash-exec.ts`). |
|
|
91
|
+
| `CLIO_CODER_RUN_OVERRIDES` | JSON envelope for run-scoped CLI options (`--max-context-tokens`, `--kv-cache-mode`, sampling flags). One typed variable instead of one env var per option; worker subprocesses inherit it (`src/core/run-overrides.ts`). |
|
|
92
|
+
| `CLIO_CODER_RESUME_SESSION_ID` | Session id handed across a self-restart; consumed and deleted at boot (`src/entry/orchestrator.ts`). |
|
|
93
|
+
| `CLIO_CODER_BOOTSTRAP_GENERATE_CHILD` | Marks the CLIO-CODER.md-generation child so it skips recursion (`src/domains/context/extension.ts`). |
|
|
94
|
+
| `CLIO_CODER_WORKER_LABELS` | Comma-separated labels a dispatched worker reports as its own (`src/domains/dispatch/transport.ts`, `src/worker/entry.ts`). |
|
|
95
|
+
| `CLIO_CODER_WORKER_PGID` | Process-group id the transport assigns a worker so its whole tree can be signalled (`src/domains/dispatch/transport.ts`, `src/worker/entry.ts`). |
|
|
96
|
+
|
|
97
|
+
## Test-only
|
|
98
|
+
|
|
99
|
+
| Variable | Purpose |
|
|
100
|
+
| --- | --- |
|
|
101
|
+
| `CLIO_CODER_WORKER_FAUX` (+ `_MODEL`, `_TEXT`, `_STOP_REASON`, `_ERROR_MESSAGE`) | Fake worker model for tests (`src/engine/ai.ts`). |
|
|
102
|
+
| `CLIO_CODER_TEST_UPGRADE_NO_NETWORK` | Skips npm install during upgrade tests (`src/cli/upgrade.ts`). |
|
|
103
|
+
| `CLIO_CODER_REQUIRE_HOME_PREFIX` | Test guardrail: abort if resolved directories escape `CLIO_CODER_HOME` (`src/core/init.ts`). |
|
|
104
|
+
|
|
105
|
+
Variables used only by `scripts/` and `benchmarks/` harnesses (the `CLIO_CODER_LIVE_*` smoke-test family, benchmark fleet configuration, install-script inputs) are not part of the shipped runtime and are documented inline where they are consumed.
|
|
@@ -0,0 +1,205 @@
|
|
|
1
|
+
# Clio Coder Local Evaluation Runner
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
|
+
|
|
8
|
+
Source of truth: [src/domains/eval/](../src/domains/eval/) and [src/cli/eval.ts](../src/cli/eval.ts).
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## CLI Commands
|
|
13
|
+
|
|
14
|
+
The CLI commands under `clio-coder eval` support running, validating, reporting, comparing, and gating evaluation suites.
|
|
15
|
+
|
|
16
|
+
```bash
|
|
17
|
+
clio-coder eval validate --suite <suite.yaml>
|
|
18
|
+
clio-coder eval run --suite <suite.yaml> [--target <id>] [--model <id>] [--out <path>] [--clio-entry <path>]
|
|
19
|
+
clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--clio-entry <path>]
|
|
20
|
+
clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
|
|
21
|
+
clio-coder eval compare <baselineEvalId> <candidateEvalId>
|
|
22
|
+
clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
|
|
23
|
+
```
|
|
24
|
+
|
|
25
|
+
### Command Roles
|
|
26
|
+
* **`validate`**: Validates the structure and constraints of a Suite v2 YAML file without executing it.
|
|
27
|
+
* **`run`**: Runs a Suite v2 (via `--suite`) or a compatibility v1 task file (via `--task-file`). Outputs a text summary and writes an eval artifact under `<dataDir>/evals/` and evidence under `<dataDir>/evidence/eval-<evalId>/`.
|
|
28
|
+
* **`report`**: Formats and prints a report from a saved `evalId`. Supports multiple `--format` outputs:
|
|
29
|
+
* `text` (default): Human-readable stdout summary.
|
|
30
|
+
* `json`: Raw JSON structure of the artifact.
|
|
31
|
+
* `md`: Markdown document with tables and summaries.
|
|
32
|
+
* `swe-jsonl`: Standardized JSONL format representing task runs (e.g. for SWE-bench comparisons).
|
|
33
|
+
* `junit`: XML report for CI/CD integration.
|
|
34
|
+
* **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
|
|
35
|
+
* **`gate`**: Compares candidate metrics against baseline or absolute thresholds, exiting non-zero if assertions fail (useful for PR gating).
|
|
36
|
+
|
|
37
|
+
Exit codes:
|
|
38
|
+
|
|
39
|
+
| Command | Success | Failure |
|
|
40
|
+
| --- | --- | --- |
|
|
41
|
+
| `eval validate` | `0` when validation passes | `2` for validation issues |
|
|
42
|
+
| `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
|
|
43
|
+
| `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
|
|
44
|
+
| `eval compare` | `0` when both artifacts load and compare succeeds | `1` if artifacts cannot be read, `2` for invalid ID |
|
|
45
|
+
| `eval gate` | `0` when all threshold assertions pass | `1` if assertions fail, `2` for config/invalid ID errors |
|
|
46
|
+
|
|
47
|
+
---
|
|
48
|
+
|
|
49
|
+
## Suite v2 Schema
|
|
50
|
+
|
|
51
|
+
Suite v2 files define matrix targets, workspaces, runner parameters, validation metrics, assertions, and path blocklists.
|
|
52
|
+
|
|
53
|
+
```yaml
|
|
54
|
+
version: 2
|
|
55
|
+
suite:
|
|
56
|
+
id: "science-suite"
|
|
57
|
+
title: "Scientific Software Evaluation"
|
|
58
|
+
visibility: "local"
|
|
59
|
+
description: "Suite for verifying HPC integrations."
|
|
60
|
+
matrix:
|
|
61
|
+
targets:
|
|
62
|
+
- id: "local-gemini"
|
|
63
|
+
model: "gemini-3.5-flash"
|
|
64
|
+
- id: "local-claude"
|
|
65
|
+
model: "claude-sonnet-5"
|
|
66
|
+
repeats: 3
|
|
67
|
+
tasks:
|
|
68
|
+
- id: "fft-tolerance-check"
|
|
69
|
+
tags:
|
|
70
|
+
- numeric
|
|
71
|
+
- fast
|
|
72
|
+
workspace:
|
|
73
|
+
kind: "temp-copy" # local | git | temp-copy
|
|
74
|
+
path: "fixtures/fft-src"
|
|
75
|
+
excludes:
|
|
76
|
+
- "**/node_modules/**"
|
|
77
|
+
runner:
|
|
78
|
+
kind: "clio-run" # clio-run | context-index | context-init | external-command
|
|
79
|
+
prompt: "Optimize the FFT tolerance bounds in solver.ts"
|
|
80
|
+
timeoutMs: 60000
|
|
81
|
+
verify:
|
|
82
|
+
commands:
|
|
83
|
+
- "npm run test"
|
|
84
|
+
assertions:
|
|
85
|
+
- metric: "result.pass"
|
|
86
|
+
op: "eq"
|
|
87
|
+
value: true
|
|
88
|
+
- metric: "tokens.total"
|
|
89
|
+
op: "lt"
|
|
90
|
+
value: 15000
|
|
91
|
+
forbidPaths:
|
|
92
|
+
- "**/credentials.yaml"
|
|
93
|
+
metrics:
|
|
94
|
+
collect:
|
|
95
|
+
- "tokens.total"
|
|
96
|
+
- "latency.wallMs"
|
|
97
|
+
timeoutMs: 90000
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
### Schema Field Reference
|
|
101
|
+
|
|
102
|
+
| Field / Section | Sub-fields | Description |
|
|
103
|
+
| --- | --- | --- |
|
|
104
|
+
| `version` | - | Must equal `2`. |
|
|
105
|
+
| `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
|
|
106
|
+
| `matrix` | `targets[]`, `repeats` | Matrix of execution targets (specifying model and thinking flags) and the repetition count. |
|
|
107
|
+
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
|
|
108
|
+
| `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
|
|
109
|
+
| `verify` | `commands`, `assertions`, `forbidPaths` | Validation steps: shell commands, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
|
|
110
|
+
| `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
|
|
111
|
+
|
|
112
|
+
---
|
|
113
|
+
|
|
114
|
+
## Workspace Kinds
|
|
115
|
+
* **`local`**: Executes the task directly in the specified local path.
|
|
116
|
+
* **`git`**: Clones the repository from `url`, checks out the specified `commit` or `checkout` ref, and runs there.
|
|
117
|
+
* **`temp-copy`**: Copies the directory at `path` to a temporary workspace location before running. This prevents side-effects from polluting other task runs.
|
|
118
|
+
|
|
119
|
+
---
|
|
120
|
+
|
|
121
|
+
## Runner Kinds
|
|
122
|
+
* **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
|
|
123
|
+
* **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
|
|
124
|
+
* **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
|
|
125
|
+
* **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
|
|
126
|
+
|
|
127
|
+
---
|
|
128
|
+
|
|
129
|
+
## Metric Assertions
|
|
130
|
+
|
|
131
|
+
Metrics collected during runs can be validated automatically using the `verify.assertions` list. Supported operator fields (`op`) are:
|
|
132
|
+
* `lt` (less than)
|
|
133
|
+
* `lte` (less than or equal)
|
|
134
|
+
* `gt` (greater than)
|
|
135
|
+
* `gte` (greater than or equal)
|
|
136
|
+
* `eq` (equal)
|
|
137
|
+
* `neq` (not equal)
|
|
138
|
+
|
|
139
|
+
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `result.pass`.
|
|
140
|
+
|
|
141
|
+
---
|
|
142
|
+
|
|
143
|
+
## Failure Classes
|
|
144
|
+
|
|
145
|
+
Evaluation tasks may fail with one of the following classes:
|
|
146
|
+
|
|
147
|
+
| Failure Class | Meaning |
|
|
148
|
+
| --- | --- |
|
|
149
|
+
| `setup_failed` | A setup command exited non-zero. |
|
|
150
|
+
| `verifier_failed` | A verifier command exited non-zero. |
|
|
151
|
+
| `timeout` | A setup or verifier command timed out. |
|
|
152
|
+
| `cwd_missing` | Resolved task cwd does not exist. |
|
|
153
|
+
| `command_error` | Reserved class for command spawn/system errors. |
|
|
154
|
+
|
|
155
|
+
---
|
|
156
|
+
|
|
157
|
+
## v1 Compatibility
|
|
158
|
+
|
|
159
|
+
Version 1 task files can still be run directly via:
|
|
160
|
+
```bash
|
|
161
|
+
clio-coder eval run --task-file tasks.yaml [--repeat <n>]
|
|
162
|
+
```
|
|
163
|
+
Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
|
|
164
|
+
* Workspace kind: `local` (using task `cwd` as workspace path)
|
|
165
|
+
* Runner kind: `external-command` (executing task `setup` commands)
|
|
166
|
+
* Verify commands: Task `verifier` commands
|
|
167
|
+
* Timeout and tags mapped directly
|
|
168
|
+
|
|
169
|
+
---
|
|
170
|
+
|
|
171
|
+
## Token Accounting & Provenance
|
|
172
|
+
|
|
173
|
+
Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
|
|
174
|
+
|
|
175
|
+
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events watched by `token-stream.ts` / `createStreamInvariantFold`. This represents usage reported by the provider for assistant messages watched over the wire. On surfaces without stdout streaming (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
|
|
176
|
+
2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
|
|
177
|
+
|
|
178
|
+
### Fail-Closed Reporting
|
|
179
|
+
Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
|
|
180
|
+
|
|
181
|
+
---
|
|
182
|
+
|
|
183
|
+
## Eval Artifact Format (v4)
|
|
184
|
+
|
|
185
|
+
Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
|
|
186
|
+
|
|
187
|
+
```typescript
|
|
188
|
+
export interface EvalArtifactV4 {
|
|
189
|
+
version: 4;
|
|
190
|
+
evalId: string;
|
|
191
|
+
suite: { id: string; hash: string };
|
|
192
|
+
clio: EvalClioProvenance;
|
|
193
|
+
environment: EvalEnvironmentProvenance;
|
|
194
|
+
matrix: { target: string; model: string | null; thinking: string | null };
|
|
195
|
+
summary: EvalArtifactSummaryV4;
|
|
196
|
+
results: EvalArtifactResultV4[];
|
|
197
|
+
}
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
---
|
|
201
|
+
|
|
202
|
+
## Task Outcome Measurement (`verify.measure`)
|
|
203
|
+
|
|
204
|
+
Task outcome commands declared under `verify.measure` evaluate whether the model solved the workload and record metrics (`task.solved`, `task.exitCode`). A non-zero exit from `verify.measure` is recorded as data and **never fails the evaluation item**. Task solution outcome is a measurement, while only machinery invariant behavior operates as a gate.
|
|
205
|
+
|
|
@@ -0,0 +1,298 @@
|
|
|
1
|
+
# Internal Eval Suites
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
|
+
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
8
|
+
checkout or object store.
|
|
9
|
+
|
|
10
|
+
Run a private suite from this source checkout with:
|
|
11
|
+
|
|
12
|
+
```sh
|
|
13
|
+
npm run build
|
|
14
|
+
clio-coder eval run --suite <external-path> --clio-entry dist/cli/index.js
|
|
15
|
+
```
|
|
16
|
+
|
|
17
|
+
Use `--out <dir>` when the artifact should be written outside the default Clio
|
|
18
|
+
data directory. Public summaries can be copied into
|
|
19
|
+
`benchmarks/results/<suite>/<run-id>/` only after they have been sanitized down
|
|
20
|
+
to `manifest.json` and `summary.json`.
|
|
21
|
+
|
|
22
|
+
## Context Regression Seed
|
|
23
|
+
|
|
24
|
+
```yaml
|
|
25
|
+
version: 2
|
|
26
|
+
suite:
|
|
27
|
+
id: internal-context-regression
|
|
28
|
+
title: Internal context regression
|
|
29
|
+
visibility: private
|
|
30
|
+
description: Private context index determinism and coverage regression.
|
|
31
|
+
provenance:
|
|
32
|
+
owner: internal-eval
|
|
33
|
+
source: private
|
|
34
|
+
matrix:
|
|
35
|
+
targets:
|
|
36
|
+
- id: local
|
|
37
|
+
repeats: 3
|
|
38
|
+
tasks:
|
|
39
|
+
- id: context-index-private-checkout
|
|
40
|
+
tags:
|
|
41
|
+
- internal
|
|
42
|
+
- context
|
|
43
|
+
- offline
|
|
44
|
+
workspace:
|
|
45
|
+
kind: local
|
|
46
|
+
path: .
|
|
47
|
+
excludes:
|
|
48
|
+
- .git
|
|
49
|
+
- node_modules
|
|
50
|
+
- dist
|
|
51
|
+
- coverage
|
|
52
|
+
runner:
|
|
53
|
+
kind: context-index
|
|
54
|
+
verify:
|
|
55
|
+
assertions:
|
|
56
|
+
- metric: context.indexedFiles
|
|
57
|
+
op: gt
|
|
58
|
+
value: 0
|
|
59
|
+
- metric: context.coverage
|
|
60
|
+
op: gt
|
|
61
|
+
value: 0
|
|
62
|
+
forbidPaths:
|
|
63
|
+
- data/private-dump
|
|
64
|
+
- runs/raw
|
|
65
|
+
metrics:
|
|
66
|
+
collect:
|
|
67
|
+
- context.indexedFiles
|
|
68
|
+
- context.coverage
|
|
69
|
+
- context.structuralHash
|
|
70
|
+
- context.digestTokens
|
|
71
|
+
- latency.wallMs
|
|
72
|
+
timeoutMs: 120000
|
|
73
|
+
thresholds:
|
|
74
|
+
fail:
|
|
75
|
+
- metric: result.pass
|
|
76
|
+
op: eq
|
|
77
|
+
value: false
|
|
78
|
+
- metric: context.coverage
|
|
79
|
+
op: lte
|
|
80
|
+
value: 0
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
## Model Matrix Seed
|
|
84
|
+
|
|
85
|
+
```yaml
|
|
86
|
+
version: 2
|
|
87
|
+
suite:
|
|
88
|
+
id: internal-model-matrix
|
|
89
|
+
title: Internal model matrix
|
|
90
|
+
visibility: private
|
|
91
|
+
description: Private model target smoke matrix for headless Clio runs.
|
|
92
|
+
provenance:
|
|
93
|
+
owner: internal-eval
|
|
94
|
+
source: private
|
|
95
|
+
matrix:
|
|
96
|
+
targets:
|
|
97
|
+
- id: local-fast
|
|
98
|
+
model: example-fast-model
|
|
99
|
+
thinking: low
|
|
100
|
+
- id: local-deep
|
|
101
|
+
model: example-deep-model
|
|
102
|
+
thinking: high
|
|
103
|
+
repeats: 2
|
|
104
|
+
tasks:
|
|
105
|
+
- id: version-smoke
|
|
106
|
+
tags:
|
|
107
|
+
- internal
|
|
108
|
+
- matrix
|
|
109
|
+
- smoke
|
|
110
|
+
workspace:
|
|
111
|
+
kind: temp-copy
|
|
112
|
+
path: fixtures/version-smoke
|
|
113
|
+
excludes:
|
|
114
|
+
- node_modules
|
|
115
|
+
- dist
|
|
116
|
+
runner:
|
|
117
|
+
kind: external-command
|
|
118
|
+
commands:
|
|
119
|
+
- node -e "console.log('matrix smoke')"
|
|
120
|
+
verify:
|
|
121
|
+
assertions:
|
|
122
|
+
- metric: verifier.exitCode
|
|
123
|
+
op: eq
|
|
124
|
+
value: 0
|
|
125
|
+
metrics:
|
|
126
|
+
collect:
|
|
127
|
+
- latency.wallMs
|
|
128
|
+
- verifier.exitCode
|
|
129
|
+
timeoutMs: 30000
|
|
130
|
+
- id: tiny-clio-run
|
|
131
|
+
tags:
|
|
132
|
+
- internal
|
|
133
|
+
- matrix
|
|
134
|
+
- model-required
|
|
135
|
+
workspace:
|
|
136
|
+
kind: temp-copy
|
|
137
|
+
path: fixtures/tiny-repair
|
|
138
|
+
excludes:
|
|
139
|
+
- node_modules
|
|
140
|
+
- dist
|
|
141
|
+
runner:
|
|
142
|
+
kind: clio-run
|
|
143
|
+
prompt: Fix the intentionally broken function so the local verifier passes.
|
|
144
|
+
verify:
|
|
145
|
+
commands:
|
|
146
|
+
- npm test
|
|
147
|
+
assertions:
|
|
148
|
+
- metric: result.pass
|
|
149
|
+
op: eq
|
|
150
|
+
value: true
|
|
151
|
+
metrics:
|
|
152
|
+
collect:
|
|
153
|
+
- tokens.input
|
|
154
|
+
- tokens.output
|
|
155
|
+
- tokens.total
|
|
156
|
+
- latency.wallMs
|
|
157
|
+
- patch.filesChanged
|
|
158
|
+
- verifier.exitCode
|
|
159
|
+
timeoutMs: 600000
|
|
160
|
+
thresholds:
|
|
161
|
+
fail:
|
|
162
|
+
- metric: result.pass
|
|
163
|
+
op: eq
|
|
164
|
+
value: false
|
|
165
|
+
- metric: tools.blocked
|
|
166
|
+
op: gt
|
|
167
|
+
value: 0
|
|
168
|
+
```
|
|
169
|
+
|
|
170
|
+
## Live Turns Seed
|
|
171
|
+
|
|
172
|
+
```yaml
|
|
173
|
+
version: 2
|
|
174
|
+
suite:
|
|
175
|
+
id: internal-live-turns
|
|
176
|
+
title: Internal live turns
|
|
177
|
+
visibility: private
|
|
178
|
+
description: Private live-fleet prompts for end-to-end Clio turn checks.
|
|
179
|
+
provenance:
|
|
180
|
+
owner: internal-eval
|
|
181
|
+
source: private
|
|
182
|
+
matrix:
|
|
183
|
+
targets:
|
|
184
|
+
- id: live-orchestrator
|
|
185
|
+
model: example-live-model
|
|
186
|
+
thinking: low
|
|
187
|
+
repeats: 1
|
|
188
|
+
tasks:
|
|
189
|
+
- id: readonly-repo-summary
|
|
190
|
+
tags:
|
|
191
|
+
- internal
|
|
192
|
+
- live
|
|
193
|
+
- readonly
|
|
194
|
+
- model-required
|
|
195
|
+
workspace:
|
|
196
|
+
kind: temp-copy
|
|
197
|
+
path: fixtures/live-readonly
|
|
198
|
+
excludes:
|
|
199
|
+
- node_modules
|
|
200
|
+
- dist
|
|
201
|
+
- .clio-coder
|
|
202
|
+
runner:
|
|
203
|
+
kind: clio-run
|
|
204
|
+
prompt: Summarize the repository purpose and make no file changes.
|
|
205
|
+
verify:
|
|
206
|
+
forbidPaths:
|
|
207
|
+
- unexpected-output.txt
|
|
208
|
+
assertions:
|
|
209
|
+
- metric: result.pass
|
|
210
|
+
op: eq
|
|
211
|
+
value: true
|
|
212
|
+
- metric: patch.filesChanged
|
|
213
|
+
op: eq
|
|
214
|
+
value: 0
|
|
215
|
+
metrics:
|
|
216
|
+
collect:
|
|
217
|
+
- tokens.total
|
|
218
|
+
- latency.wallMs
|
|
219
|
+
- tools.totalCalls
|
|
220
|
+
- tools.failed
|
|
221
|
+
- patch.filesChanged
|
|
222
|
+
timeoutMs: 300000
|
|
223
|
+
- id: live-small-edit
|
|
224
|
+
tags:
|
|
225
|
+
- internal
|
|
226
|
+
- live
|
|
227
|
+
- edit
|
|
228
|
+
- model-required
|
|
229
|
+
workspace:
|
|
230
|
+
kind: temp-copy
|
|
231
|
+
path: fixtures/live-small-edit
|
|
232
|
+
excludes:
|
|
233
|
+
- node_modules
|
|
234
|
+
- dist
|
|
235
|
+
- .clio-coder
|
|
236
|
+
runner:
|
|
237
|
+
kind: clio-run
|
|
238
|
+
prompt: Fix the failing unit test with the smallest source change.
|
|
239
|
+
verify:
|
|
240
|
+
commands:
|
|
241
|
+
- npm test
|
|
242
|
+
assertions:
|
|
243
|
+
- metric: result.pass
|
|
244
|
+
op: eq
|
|
245
|
+
value: true
|
|
246
|
+
- metric: patch.testFilesModified
|
|
247
|
+
op: eq
|
|
248
|
+
value: 0
|
|
249
|
+
metrics:
|
|
250
|
+
collect:
|
|
251
|
+
- tokens.input
|
|
252
|
+
- tokens.output
|
|
253
|
+
- tokens.total
|
|
254
|
+
- latency.wallMs
|
|
255
|
+
- tools.totalCalls
|
|
256
|
+
- tools.failed
|
|
257
|
+
- patch.filesChanged
|
|
258
|
+
- patch.testFilesModified
|
|
259
|
+
- verifier.exitCode
|
|
260
|
+
timeoutMs: 600000
|
|
261
|
+
thresholds:
|
|
262
|
+
fail:
|
|
263
|
+
- metric: result.pass
|
|
264
|
+
op: eq
|
|
265
|
+
value: false
|
|
266
|
+
- metric: tools.failed
|
|
267
|
+
op: gt
|
|
268
|
+
value: 0
|
|
269
|
+
```
|
|
270
|
+
|
|
271
|
+
---
|
|
272
|
+
|
|
273
|
+
## Soak Benchmark Suite
|
|
274
|
+
|
|
275
|
+
The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
|
|
276
|
+
|
|
277
|
+
The soak suite comprises four specialized suite files:
|
|
278
|
+
|
|
279
|
+
### 1. Machinery Under Load (`clio-soak.yaml`)
|
|
280
|
+
Evaluates the same task workload across two execution surfaces: the headless main-agent surface (`clio-run`) and a dispatched worker surface (`agent: coder`). It tests single-file bugs, multi-file bugs, and compaction continuity across restarts.
|
|
281
|
+
- **Surface Differences**: Main-agent tasks verify session ledger continuity (`ledger.formatVersion`, `ledger.toolPairsUnmatched`, `ledger.assistantBetweenCallAndResult`), while dispatch worker tasks verify process group cleanup (`process.orphanedChildren == 0`).
|
|
282
|
+
- **Compaction Continuity**: Verifies that compaction summaries are present (`continuity.compactionSummaryPresent`) and that pre-compaction facts are preserved (`continuity.answeredFromPreCompaction`).
|
|
283
|
+
- **Suite-Wide Gates**: Gates on `receipt.sealed`, `receipt.integrityValid`, `receipt.outcomeMatchesExit`, `tokens.measured`, `stream.cumulativeSnapshots == 0`, `stream.usageDoubleCounted == false`, and `stream.segmentUsageMatchesMessages == true`.
|
|
284
|
+
|
|
285
|
+
### 2. Per-Step Write Boundaries (`clio-soak-boundary.yaml`)
|
|
286
|
+
Validates write boundary enforcement across steps without model participation. Enforcement is strictly detect-and-rollback and is never sandboxing.
|
|
287
|
+
- `write-boundary.rolled-back`: Verifies clean detection of allowlist violations (`writes_boundary_violation`), git-level file restoration, and sealed verdict generation (`boundary.violationsRolledBack == 1`, `boundary.rollbackIncomplete == 0`).
|
|
288
|
+
- `write-boundary.rollback-incomplete`: Tests honest failure reporting when a path was dirty prior to snapshot taking. prior bytes exist only in the overwritten tree, so rollback leaves the tree unchanged and records incomplete rollback (`boundary.rollbackIncomplete == 1`, `boundary.violationsRolledBack == 0`).
|
|
289
|
+
|
|
290
|
+
### 3. Fault Injection Chaos (`clio-soak-chaos.yaml`)
|
|
291
|
+
Evaluates system resilience against process signals.
|
|
292
|
+
- `chaos.sigint-mid-tool`: Prompts Clio for a long-running bash tool call and injects `SIGINT` once the subprocess initializes. Asserts exit code `130`, confirms no orphaned children remain (`process.orphanedChildren == 0`), and verifies receipt sealing, receipt integrity, and provider token reporting.
|
|
293
|
+
|
|
294
|
+
### 4. Bounded Loops (`clio-soak-loop.yaml`)
|
|
295
|
+
Validates iteration bounds and receipt accounting for fleet loops (`bounded-loop.fleet`).
|
|
296
|
+
- **Loop Bounds**: Asserts that verification attempts do not exceed declared limits (`loop.attemptsSpent <= 3`), recovery attempts seal individual receipts (`loop.receiptsMatchRepairs == true`), and unneeded nodes report as `unneeded` rather than skipped or failed (`loop.skippedNodes == 0`).
|
|
297
|
+
- **Two Token Accountings**: Distinguishes `tokens.*` (folded live off wire stdout by `createStreamInvariantFold`) from `receiptUsage.*` (journal receipts sealed and authenticated against ledger envelopes). On fleet runs, wire streaming is absent (`tokens.measured == false`), while journal receipts provide authenticated usage (`receiptUsage.measured == true`).
|
|
298
|
+
|