@iowarp/clio-coder 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +407 -0
- package/CODE_OF_CONDUCT.md +21 -0
- package/CONTRIBUTING.md +224 -0
- package/LICENSE +202 -0
- package/NOTICE +9 -0
- package/README.md +798 -0
- package/SECURITY.md +72 -0
- package/assets/clio-coder-logo-128.webp +0 -0
- package/damage-control-rules.yaml +419 -0
- package/dist/acp-UMLFVA3F.js +92 -0
- package/dist/agents-Q4MYPMUW.js +91 -0
- package/dist/auth-O6HYIJ6J.js +521 -0
- package/dist/chunk-262G75JS.js +35 -0
- package/dist/chunk-26BZQOAD.js +1281 -0
- package/dist/chunk-2J63S4SF.js +508 -0
- package/dist/chunk-3DANZDGR.js +717 -0
- package/dist/chunk-4UQA7NCT.js +29 -0
- package/dist/chunk-527KG6XR.js +497 -0
- package/dist/chunk-5LDRNKX2.js +1063 -0
- package/dist/chunk-5N2FG33Q.js +25 -0
- package/dist/chunk-67MTHP2E.js +135 -0
- package/dist/chunk-6CWDTGUC.js +20 -0
- package/dist/chunk-7BHLZB3A.js +2115 -0
- package/dist/chunk-7RBKDI66.js +348 -0
- package/dist/chunk-AMFR5YA3.js +541 -0
- package/dist/chunk-BBUH4VAA.js +1224 -0
- package/dist/chunk-BYEU76JP.js +899 -0
- package/dist/chunk-CLJ5HLUD.js +458 -0
- package/dist/chunk-D5YD55AR.js +116 -0
- package/dist/chunk-DXQNI4PC.js +61 -0
- package/dist/chunk-E3NYWENM.js +1004 -0
- package/dist/chunk-GNGDQYDU.js +34688 -0
- package/dist/chunk-GOTUR54M.js +9 -0
- package/dist/chunk-HBU5MTAM.js +41 -0
- package/dist/chunk-HMYNFFY4.js +28 -0
- package/dist/chunk-JPOWPFCU.js +1010 -0
- package/dist/chunk-JWHCJDCI.js +1215 -0
- package/dist/chunk-KBR4MZZR.js +41 -0
- package/dist/chunk-KKKPTZLM.js +93 -0
- package/dist/chunk-ME6DNWIU.js +66 -0
- package/dist/chunk-NI4DEJMC.js +88 -0
- package/dist/chunk-O4EJEDHO.js +659 -0
- package/dist/chunk-PIDUD6M2.js +31 -0
- package/dist/chunk-PS4PFJQP.js +29459 -0
- package/dist/chunk-QV47YRF4.js +48 -0
- package/dist/chunk-RQDWMVRB.js +279 -0
- package/dist/chunk-TFSSEXL6.js +136 -0
- package/dist/chunk-TKHQ4DGZ.js +8290 -0
- package/dist/chunk-TPOCL34A.js +2876 -0
- package/dist/chunk-UGYAX5YI.js +565 -0
- package/dist/chunk-UHTSULZS.js +461 -0
- package/dist/chunk-UU3R62TT.js +128 -0
- package/dist/chunk-UWIJNAOB.js +3906 -0
- package/dist/chunk-VOO7NYPP.js +914 -0
- package/dist/chunk-VPAWTYLY.js +117 -0
- package/dist/chunk-WD6AJM35.js +1216 -0
- package/dist/chunk-X3BR7HWV.js +115 -0
- package/dist/chunk-X3NE4WVW.js +120 -0
- package/dist/chunk-XNISANGE.js +1395 -0
- package/dist/chunk-XV4ZJ6ZM.js +3177 -0
- package/dist/cli/index.js +236 -0
- package/dist/clio-KIQ5SNDS.js +53 -0
- package/dist/components-JVHMUBEB.js +653 -0
- package/dist/config-ZFCDBMDC.js +372 -0
- package/dist/configure-G4E3A2PG.js +27 -0
- package/dist/context-CDXTP2MP.js +293 -0
- package/dist/context-E3KIFVXI.js +185 -0
- package/dist/context-clear-3F4PLXOS.js +102 -0
- package/dist/context-index-Q7YSYTR3.js +106 -0
- package/dist/docs-YIETIWZI.js +280 -0
- package/dist/doctor-M5HJJZOL.js +61 -0
- package/dist/domains/agents/builtins/architect.md +33 -0
- package/dist/domains/agents/builtins/coder.md +31 -0
- package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
- package/dist/domains/agents/builtins/debugger.md +30 -0
- package/dist/domains/agents/builtins/documenter.md +31 -0
- package/dist/domains/agents/builtins/git-master.md +30 -0
- package/dist/domains/agents/builtins/provenance.md +30 -0
- package/dist/domains/agents/builtins/researcher.md +71 -0
- package/dist/domains/agents/builtins/scout.md +42 -0
- package/dist/domains/agents/builtins/tester.md +31 -0
- package/dist/domains/agents/builtins/verifier.md +30 -0
- package/dist/domains/agents/builtins/wiki-writer.md +41 -0
- package/dist/eval-B3KZZESM.js +2674 -0
- package/dist/evidence-V67CHM35.js +233 -0
- package/dist/evolve-YDZSUQYA.js +518 -0
- package/dist/extensions-SRG7XCAH.js +207 -0
- package/dist/fleet-CA2CRTVG.js +760 -0
- package/dist/fleet-preflight-CLIAX7YR.js +21 -0
- package/dist/init-2OZDJE2D.js +227 -0
- package/dist/memory-3PIQQAKX.js +207 -0
- package/dist/models-DY35XI7Y.js +237 -0
- package/dist/paths-5OMXW7Z4.js +57 -0
- package/dist/preload-KZVHET2B.js +11 -0
- package/dist/reset-PIFYNOS3.js +216 -0
- package/dist/run-3VSPP24F.js +735 -0
- package/dist/share-D36RQCXM.js +241 -0
- package/dist/skills-F2MRLELY.js +445 -0
- package/dist/skills-eval-E2ZTW4PL.js +932 -0
- package/dist/targets-DZMEZAH4.js +977 -0
- package/dist/trace-7NYCUI2J.js +250 -0
- package/dist/uninstall-AD3JWHBB.js +322 -0
- package/dist/upgrade-WYYBKGDY.js +301 -0
- package/dist/usage-ULIDAGFF.js +755 -0
- package/dist/version-ROZ6CZKH.js +16 -0
- package/dist/wiki-generate-PKFIX6OB.js +377 -0
- package/dist/worker/entry.js +1739 -0
- package/docs/README.md +93 -0
- package/docs/acp.md +120 -0
- package/docs/alcf-provider.md +72 -0
- package/docs/architecture.md +172 -0
- package/docs/artifact-versions.md +54 -0
- package/docs/built-in-agents.md +265 -0
- package/docs/capacity-and-scheduling.md +97 -0
- package/docs/commands-and-modes.md +554 -0
- package/docs/config-knobs-audit.md +115 -0
- package/docs/configuration-and-targets.md +812 -0
- package/docs/context-engine.md +236 -0
- package/docs/dispatch-architecture-rationale.md +126 -0
- package/docs/documentation-coverage.md +46 -0
- package/docs/documentation-guide.md +166 -0
- package/docs/environment-variables.md +105 -0
- package/docs/eval-runner.md +205 -0
- package/docs/evals-internal.md +298 -0
- package/docs/evidence-and-memory.md +243 -0
- package/docs/evolution.md +143 -0
- package/docs/exit-codes-and-output.md +74 -0
- package/docs/extensions-and-sharing.md +306 -0
- package/docs/fleet-demo-runbook.md +179 -0
- package/docs/fleet-dispatch.md +591 -0
- package/docs/glossary.md +75 -0
- package/docs/html/agents_blueprint.html +936 -0
- package/docs/html/alcf_blueprint.html +324 -0
- package/docs/html/architecture_blueprint.html +850 -0
- package/docs/html/commands_blueprint.html +794 -0
- package/docs/html/config_knobs_audit_blueprint.html +178 -0
- package/docs/html/configuration_blueprint.html +1080 -0
- package/docs/html/context_blueprint.html +603 -0
- package/docs/html/documentation_blueprint.html +832 -0
- package/docs/html/environment_blueprint.html +404 -0
- package/docs/html/eval_blueprint.html +743 -0
- package/docs/html/evals_internal_blueprint.html +190 -0
- package/docs/html/evolution_blueprint.html +674 -0
- package/docs/html/extensions_blueprint.html +2065 -0
- package/docs/html/fleet_dispatch_blueprint.html +286 -0
- package/docs/html/index.html +919 -0
- package/docs/html/lifecycle_blueprint.html +723 -0
- package/docs/html/memory_blueprint.html +699 -0
- package/docs/html/middleware_blueprint.html +664 -0
- package/docs/html/models_blueprint.html +2366 -0
- package/docs/html/observability_blueprint.html +683 -0
- package/docs/html/provider_adapter_blueprint.html +245 -0
- package/docs/html/safety_blueprint.html +1386 -0
- package/docs/html/shared.css +571 -0
- package/docs/html/shared.js +143 -0
- package/docs/html/skills_blueprint.html +671 -0
- package/docs/html/soak_blueprint.html +182 -0
- package/docs/html/tool_usage_blueprint.html +350 -0
- package/docs/html/tools_blueprint.html +2249 -0
- package/docs/html/trace_blueprint.html +235 -0
- package/docs/html/tui_design_blueprint.html +314 -0
- package/docs/html/validation_blueprint.html +961 -0
- package/docs/html/worker_dispatch_blueprint.html +231 -0
- package/docs/installation-and-lifecycle.md +308 -0
- package/docs/middleware-and-components.md +148 -0
- package/docs/model-catalog.md +189 -0
- package/docs/observability.md +233 -0
- package/docs/proactive-memory.md +452 -0
- package/docs/prompt-envelope-and-tools.md +142 -0
- package/docs/provider-adapter-cookbook.md +148 -0
- package/docs/release-cut-checklist.md +138 -0
- package/docs/safety-model.md +357 -0
- package/docs/scientific-validation.md +105 -0
- package/docs/session-lifecycle.md +156 -0
- package/docs/skills-marketplace.md +46 -0
- package/docs/tool-usage.md +527 -0
- package/docs/trace-store.md +132 -0
- package/docs/troubleshooting.md +33 -0
- package/docs/tui-design.md +239 -0
- package/docs/worker-dispatch-mechanics.md +242 -0
- package/package.json +132 -0
- package/skills/README.md +408 -0
- package/skills/git/commit-crafting/SKILL.md +79 -0
- package/skills/git/commit-crafting/evals.md +92 -0
- package/skills/git/create-pr/SKILL.md +116 -0
- package/skills/git/create-pr/evals.md +114 -0
- package/skills/git/investigate-issue/SKILL.md +139 -0
- package/skills/git/investigate-issue/evals.md +94 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
- package/skills/git/resolve-merge-conflicts/evals.md +58 -0
- package/skills/git/review-changes/SKILL.md +103 -0
- package/skills/git/review-changes/evals.md +85 -0
- package/skills/git/worktree-create/SKILL.md +92 -0
- package/skills/git/worktree-create/evals.md +97 -0
- package/skills/git/worktree-create/references/worktree-setup.md +66 -0
- package/skills/git/worktree-merge/SKILL.md +95 -0
- package/skills/git/worktree-merge/evals.md +114 -0
- package/skills/skill-marketplace.json +261 -0
- package/skills/workflow/cut-it/SKILL.md +86 -0
- package/skills/workflow/cut-it/evals.md +42 -0
- package/src/domains/agents/builtins/architect.md +33 -0
- package/src/domains/agents/builtins/coder.md +31 -0
- package/src/domains/agents/builtins/context-bootstrap.md +38 -0
- package/src/domains/agents/builtins/debugger.md +30 -0
- package/src/domains/agents/builtins/documenter.md +31 -0
- package/src/domains/agents/builtins/git-master.md +30 -0
- package/src/domains/agents/builtins/provenance.md +30 -0
- package/src/domains/agents/builtins/researcher.md +71 -0
- package/src/domains/agents/builtins/scout.md +42 -0
- package/src/domains/agents/builtins/tester.md +31 -0
- package/src/domains/agents/builtins/verifier.md +30 -0
- package/src/domains/agents/builtins/wiki-writer.md +41 -0
- package/src/domains/agents/fleets/build-review.md +34 -0
- package/src/domains/agents/fleets/build-test.md +35 -0
- package/src/domains/agents/fleets/sdlc.md +86 -0
- package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
- package/src/domains/prompts/fragments/identity/clio.md +26 -0
- package/src/domains/prompts/fragments/operating/contract.md +64 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
- package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
- package/src/domains/prompts/fragments/safety/read-only.md +13 -0
- package/src/domains/prompts/fragments/safety/suggest.md +13 -0
- package/src/domains/prompts/fragments/wiki/page.md +75 -0
- package/src/domains/prompts/fragments/wiki/plan.md +48 -0
- package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
package/README.md
ADDED
|
@@ -0,0 +1,798 @@
|
|
|
1
|
+
<p align="center">
|
|
2
|
+
<picture>
|
|
3
|
+
<source srcset="https://raw.githubusercontent.com/iowarp/clio-coder/main/assets/banner.webp" type="image/webp" />
|
|
4
|
+
<img src="https://raw.githubusercontent.com/iowarp/clio-coder/main/assets/banner.png" alt="Clio Coder, the coding agent in IOWarp's CLIO ecosystem of agentic science" width="100%" />
|
|
5
|
+
</picture>
|
|
6
|
+
</p>
|
|
7
|
+
|
|
8
|
+
<h1 align="center">Clio Coder</h1>
|
|
9
|
+
|
|
10
|
+
<p align="center"><strong>A supervised coding agent for research software, built to run on your models, on your machines, with a receipt for everything it did.</strong></p>
|
|
11
|
+
|
|
12
|
+
<p align="center">
|
|
13
|
+
<a href="https://github.com/iowarp/clio-coder/releases/latest"><img alt="Latest release" src="https://img.shields.io/github/v/tag/iowarp/clio-coder?sort=semver&label=release&color=00d4db&style=flat-square" /></a>
|
|
14
|
+
<a href="https://www.npmjs.com/package/@iowarp/clio-coder"><img alt="npm" src="https://img.shields.io/npm/v/%40iowarp%2Fclio-coder?label=npm&color=cb3837&style=flat-square" /></a>
|
|
15
|
+
<a href="https://github.com/iowarp/clio-coder/actions/workflows/ci.yml"><img alt="CI status" src="https://img.shields.io/github/actions/workflow/status/iowarp/clio-coder/ci.yml?branch=main&label=ci&style=flat-square" /></a>
|
|
16
|
+
<a href="#requirements"><img alt="Node >=22.19" src="https://img.shields.io/badge/node-%3E%3D22.19-147366?style=flat-square" /></a>
|
|
17
|
+
<a href="LICENSE"><img alt="License Apache-2.0" src="https://img.shields.io/badge/license-Apache--2.0-241131?style=flat-square" /></a>
|
|
18
|
+
<a href="https://iowarp.ai"><img alt="IOWarp CLIO" src="https://img.shields.io/badge/IOWarp-CLIO-00d4db?style=flat-square" /></a>
|
|
19
|
+
<a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=2411318"><img alt="NSF #2411318" src="https://img.shields.io/badge/NSF-%232411318-241131?style=flat-square" /></a>
|
|
20
|
+
</p>
|
|
21
|
+
|
|
22
|
+
> [!WARNING]
|
|
23
|
+
> **Experimental v0.3.0 release.** Clio Coder's behavior and interfaces may
|
|
24
|
+
> break or change without notice. Use version control, review proposed changes,
|
|
25
|
+
> and keep backups when operating on important repositories.
|
|
26
|
+
|
|
27
|
+
---
|
|
28
|
+
|
|
29
|
+
Clio Coder is a terminal coding agent for people who work on real scientific
|
|
30
|
+
and HPC codebases: simulation kernels, data pipelines, numerical libraries,
|
|
31
|
+
build systems that take twenty minutes and break in ways no cloud model has
|
|
32
|
+
ever seen.
|
|
33
|
+
|
|
34
|
+
You bring the model. A local llama.cpp, Ollama, LM Studio, vLLM, or SGLang
|
|
35
|
+
server; a cloud API; your ChatGPT or Claude subscription; or an Argonne
|
|
36
|
+
Leadership Computing Facility inference gateway. Clio brings the harness
|
|
37
|
+
around it: a terminal UI, twenty typed tools instead of an unrestricted
|
|
38
|
+
shell, a fleet of bounded worker agents that can run across your whole
|
|
39
|
+
cluster over SSH, durable sessions, and an integrity-sealed receipt for every run.
|
|
40
|
+
|
|
41
|
+
CLIO stands for Context Layer for Input/Output. Clio Coder is the interactive
|
|
42
|
+
coding agent in IOWarp's ecosystem of agentic science, named for the Greek
|
|
43
|
+
muse of history and built by the Gnosis Research Center at Illinois Tech.
|
|
44
|
+
|
|
45
|
+
### Pick your path
|
|
46
|
+
|
|
47
|
+
| | You are | Start here |
|
|
48
|
+
| --- | --- | --- |
|
|
49
|
+
| 🔬 | A researcher or developer who wants to use it | [Install](#install) → [Five-minute start](#five-minute-start) → [Bring your own model](#bring-your-own-model) |
|
|
50
|
+
| 🤖 | An AI agent that just landed in this repository | [For agents](#for-agents) |
|
|
51
|
+
| 🛠️ | A developer who wants to contribute | [For contributors](#for-contributors) |
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
## Why Clio is different
|
|
56
|
+
|
|
57
|
+
Most coding agents ask you to trust a remote model with a shell. Clio makes a
|
|
58
|
+
different bet: the harness should be strong enough that a 20B model running on
|
|
59
|
+
your own GPU is useful, and honest enough that you can reconstruct every
|
|
60
|
+
decision afterward.
|
|
61
|
+
|
|
62
|
+
**The model never gets a shell by default.** The tool surface is twenty
|
|
63
|
+
typed tools organized into seven policy planes. Bash is default-deny, filtered
|
|
64
|
+
through [damage-control rules](damage-control-rules.yaml) and per-project
|
|
65
|
+
policy. Reads are bounded, writes are queued and reviewable, and every
|
|
66
|
+
privileged call passes through one admission path that cannot be widened by
|
|
67
|
+
the model asking nicely.
|
|
68
|
+
|
|
69
|
+
**Local models are the design target, not a fallback.** llama.cpp and similar
|
|
70
|
+
servers expose a single prefix-cache slot. Clio keeps the compiled prompt and
|
|
71
|
+
provider tool schemas byte-stable so that slot stays hot across turns and
|
|
72
|
+
sessions, bounds every tool result so one `grep` cannot blow the window, and
|
|
73
|
+
records a per-call cache verdict (`hot`, `partial`, `cold`, `small`) in the
|
|
74
|
+
session ledger so you can see when and why the cache went cold.
|
|
75
|
+
|
|
76
|
+
**Work is delegated to bounded workers, not to one long context.** The
|
|
77
|
+
orchestrator dispatches focused agents with explicit tool profiles, call
|
|
78
|
+
budgets, cost ceilings, and typed result contracts. A worker that cannot
|
|
79
|
+
produce a conforming answer fails loudly instead of returning confident prose.
|
|
80
|
+
|
|
81
|
+
**Your cluster is the runtime.** Declare your nodes and the same worker
|
|
82
|
+
protocol tunnels over SSH. A remote worker gets the same prompts, the same
|
|
83
|
+
safety matrix, the same receipts. Placement is deterministic and pinnable, and
|
|
84
|
+
capacity is governed by durable expiring leases that survive process death.
|
|
85
|
+
|
|
86
|
+
**Everything is auditable.** Each run seals a receipt covering token usage,
|
|
87
|
+
priced cost, tool activity, safety decisions, routing intent, the resolved
|
|
88
|
+
route, worker attestation, and result-contract conformance. `clio-coder evidence`
|
|
89
|
+
and `/view verify` check them; nothing in the audit trail is reconstructed
|
|
90
|
+
from prose.
|
|
91
|
+
|
|
92
|
+
**Science is a first-class domain.** [clio-kit](https://github.com/iowarp/clio-kit)
|
|
93
|
+
contributes MCP servers for HDF5, Slurm, ParaView, Pandas, NetCDF, FITS, Zarr,
|
|
94
|
+
and ArXiv, and the shipped skills catalog includes scientific debugging and
|
|
95
|
+
experiment-protocol guides.
|
|
96
|
+
|
|
97
|
+
---
|
|
98
|
+
|
|
99
|
+
# For humans
|
|
100
|
+
|
|
101
|
+
## Requirements
|
|
102
|
+
|
|
103
|
+
- Node.js `>=22.19.0` and npm
|
|
104
|
+
- Linux or macOS. Windows is best effort until a stable release.
|
|
105
|
+
- At least one model target: a local OpenAI-compatible server, Ollama, LM
|
|
106
|
+
Studio, llama.cpp, vLLM, SGLang, a cloud API key, a ChatGPT or Claude
|
|
107
|
+
subscription login, an ALCF Globus account, or an installed `claude` command
|
|
108
|
+
|
|
109
|
+
## Install
|
|
110
|
+
|
|
111
|
+
From npm:
|
|
112
|
+
|
|
113
|
+
```bash
|
|
114
|
+
npm install -g @iowarp/clio-coder
|
|
115
|
+
clio-coder --version
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
From source, pinned to this release:
|
|
119
|
+
|
|
120
|
+
```bash
|
|
121
|
+
git clone --branch v0.3.0 https://github.com/iowarp/clio-coder.git
|
|
122
|
+
cd clio-coder
|
|
123
|
+
npm run install:local
|
|
124
|
+
export PATH="$HOME/.local/bin:$PATH"
|
|
125
|
+
hash -r
|
|
126
|
+
"$HOME/.local/bin/clio-coder" --version
|
|
127
|
+
```
|
|
128
|
+
|
|
129
|
+
The clone is pinned to `v0.3.0`, the release these instructions describe.
|
|
130
|
+
Without `--branch` you get the default branch, which is ahead of the release and
|
|
131
|
+
is not what the rest of this page documents.
|
|
132
|
+
|
|
133
|
+
`npm run install:local` verifies dependencies, builds the CLI, installs a
|
|
134
|
+
symlink at `${CLIO_CODER_BIN_DIR:-$HOME/.local/bin}/clio-coder`, and runs the installed
|
|
135
|
+
CLI's structure repair so a fresh install passes plain `clio-coder doctor` with no
|
|
136
|
+
manual steps. It warns if the bin directory is not on your `PATH` and prints the
|
|
137
|
+
`export` line above; the line is a no-op for the shell that already has it, and
|
|
138
|
+
it is what makes a bare `clio-coder` resolve in the shell that does not. Put it in
|
|
139
|
+
your shell profile to keep it across sessions. The symlink executes
|
|
140
|
+
`dist/cli/index.js`, so re-run `npm run build` after editing TypeScript sources.
|
|
141
|
+
|
|
142
|
+
The last line runs the launcher by its full path on purpose. A bare `clio-coder` may
|
|
143
|
+
resolve to an older install earlier on your `PATH`, so it verifies whichever one
|
|
144
|
+
that is rather than the one you just installed; the installer warns when it
|
|
145
|
+
finds that shadowing, and names the other path.
|
|
146
|
+
|
|
147
|
+
Before you switch to the bare name, ask which file it reaches:
|
|
148
|
+
|
|
149
|
+
```bash
|
|
150
|
+
command -v clio-coder # expect $HOME/.local/bin/clio-coder
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
Comparing `clio-coder --version` against `"$HOME/.local/bin/clio-coder" --version` does not
|
|
154
|
+
answer that. Two installs of the same release print the same version, so the
|
|
155
|
+
versions agree while the name still resolves to the other one. The path is the
|
|
156
|
+
question.
|
|
157
|
+
|
|
158
|
+
To remove it, preview first:
|
|
159
|
+
|
|
160
|
+
```bash
|
|
161
|
+
clio-coder uninstall --dry-run
|
|
162
|
+
clio-coder uninstall --remove-binary --force
|
|
163
|
+
hash -r
|
|
164
|
+
```
|
|
165
|
+
|
|
166
|
+
For selective wipes that keep settings or credentials, use `clio-coder reset`. Full
|
|
167
|
+
details live in [docs/installation-and-lifecycle.md](docs/installation-and-lifecycle.md).
|
|
168
|
+
|
|
169
|
+
## Five-minute start
|
|
170
|
+
|
|
171
|
+
Run Clio from the repository you want to work on and point one target at a
|
|
172
|
+
running model server. This example uses LM Studio; other local runtime ids
|
|
173
|
+
include `ollama-native`, `llamacpp`, `vllm`, and `sglang`.
|
|
174
|
+
|
|
175
|
+
```bash
|
|
176
|
+
cd /path/to/your/repo
|
|
177
|
+
|
|
178
|
+
clio-coder configure \
|
|
179
|
+
--id local-lmstudio \
|
|
180
|
+
--runtime lmstudio-native \
|
|
181
|
+
--url http://localhost:1234 \
|
|
182
|
+
--model your-model-id \
|
|
183
|
+
--set-orchestrator \
|
|
184
|
+
--set-fleet-default
|
|
185
|
+
|
|
186
|
+
clio-coder targets use local-lmstudio
|
|
187
|
+
clio-coder targets --probe
|
|
188
|
+
```
|
|
189
|
+
|
|
190
|
+
Once the target probes healthy, teach Clio about your project, try a headless
|
|
191
|
+
turn, then open the TUI:
|
|
192
|
+
|
|
193
|
+
```bash
|
|
194
|
+
clio-coder context init # bootstraps local generated CLIO-CODER.md context from your real source tree
|
|
195
|
+
clio-coder run "Summarize this repository layout and identify the main entry points."
|
|
196
|
+
clio-coder # interactive terminal UI
|
|
197
|
+
```
|
|
198
|
+
|
|
199
|
+
Inside the TUI, `/targets`, `/agents`, `/fleet`, and `/skill` confirm what the
|
|
200
|
+
session can see. `/help` opens the interactive help center.
|
|
201
|
+
|
|
202
|
+
## Bring your own model
|
|
203
|
+
|
|
204
|
+
Clio treats models as named **targets**. A target is a runtime plus an
|
|
205
|
+
endpoint plus a model plus credentials, and you can route interactive chat and
|
|
206
|
+
fleet dispatch through different targets independently.
|
|
207
|
+
|
|
208
|
+
### Local runtimes
|
|
209
|
+
|
|
210
|
+
| Runtime id | Server |
|
|
211
|
+
| --- | --- |
|
|
212
|
+
| `llamacpp`, `llamacpp-anthropic`, `llamacpp-completion` | llama.cpp and llama-swap routers |
|
|
213
|
+
| `lmstudio-native` | LM Studio |
|
|
214
|
+
| `ollama-native` | Ollama |
|
|
215
|
+
| `vllm`, `sglang` | vLLM and SGLang |
|
|
216
|
+
| `lemonade`, `lemonade-anthropic` | Lemonade |
|
|
217
|
+
| `openai-compat`, `anthropic-compat` | Any OpenAI- or Anthropic-shaped endpoint |
|
|
218
|
+
|
|
219
|
+
### Cloud APIs
|
|
220
|
+
|
|
221
|
+
`openai`, `anthropic`, `google`, `groq`, `mistral`, `deepseek`, `openrouter`,
|
|
222
|
+
`bedrock`, and `alcf` for Argonne's Sophia and Metis gateways over Globus
|
|
223
|
+
OAuth. See [docs/alcf-provider.md](docs/alcf-provider.md) for the HPC path.
|
|
224
|
+
|
|
225
|
+
### Subscriptions
|
|
226
|
+
|
|
227
|
+
You can drive Clio from a ChatGPT Plus/Pro or Claude Pro/Max subscription
|
|
228
|
+
instead of an API key.
|
|
229
|
+
|
|
230
|
+
```bash
|
|
231
|
+
clio-coder auth login anthropic-max # Claude Pro/Max OAuth
|
|
232
|
+
clio-coder auth login openai-codex # ChatGPT Plus/Pro OAuth
|
|
233
|
+
|
|
234
|
+
clio-coder configure --id claude-sub --runtime anthropic-max --model claude-sonnet-5 --set-orchestrator
|
|
235
|
+
clio-coder configure --id chatgpt-sub --runtime openai-codex --model gpt-5.4 --set-orchestrator
|
|
236
|
+
```
|
|
237
|
+
|
|
238
|
+
Pick model ids from `clio-coder models --target <id>` after login.
|
|
239
|
+
|
|
240
|
+
> [!NOTE]
|
|
241
|
+
> Connecting a Claude Pro/Max subscription over OAuth uses the same path as
|
|
242
|
+
> Claude Code. Using subscription credentials outside a vendor's first-party
|
|
243
|
+
> apps may not align with their terms of service. Enable at your own
|
|
244
|
+
> discretion.
|
|
245
|
+
|
|
246
|
+
### Subscription-backed workers
|
|
247
|
+
|
|
248
|
+
Clio can also drive other coding agents as workers while keeping its own
|
|
249
|
+
permission gating in front of them.
|
|
250
|
+
|
|
251
|
+
```bash
|
|
252
|
+
claude auth login # authenticate the official Claude CLI first
|
|
253
|
+
|
|
254
|
+
# Claude Code SDK worker, with enforced per-tool safety
|
|
255
|
+
clio-coder configure --id claude-sdk-worker --runtime claude-sdk --model sonnet --set-fleet-default
|
|
256
|
+
|
|
257
|
+
# claude -p subprocess worker, advisory permission-mode gating only
|
|
258
|
+
clio-coder configure --id claude-code-worker --runtime claude-code --model sonnet
|
|
259
|
+
|
|
260
|
+
# Google Antigravity subprocess worker, under your existing agy login
|
|
261
|
+
clio-coder configure --id agy-worker --runtime antigravity-code --model "Gemini 3.5 Flash (High)"
|
|
262
|
+
```
|
|
263
|
+
|
|
264
|
+
### Mixing them
|
|
265
|
+
|
|
266
|
+
The interesting configuration is a strong orchestrator with cheap local
|
|
267
|
+
muscle, or the reverse.
|
|
268
|
+
|
|
269
|
+
```bash
|
|
270
|
+
clio-coder configure --id chatgpt-orch --runtime openai-codex --model gpt-5.4 --set-orchestrator
|
|
271
|
+
clio-coder configure --id claude-worker --runtime claude-sdk --model sonnet
|
|
272
|
+
clio-coder configure --id local-fleet --runtime lmstudio-native --url http://localhost:1234 \
|
|
273
|
+
--model qwen-7b --set-fleet-default
|
|
274
|
+
|
|
275
|
+
clio-coder targets profile claude-sdk claude-worker --model sonnet
|
|
276
|
+
clio-coder run --agent coder "Refactor src/engine/parser.ts"
|
|
277
|
+
```
|
|
278
|
+
|
|
279
|
+
Full reference: [docs/configuration-and-targets.md](docs/configuration-and-targets.md).
|
|
280
|
+
|
|
281
|
+
### Keeping a scout model resident
|
|
282
|
+
|
|
283
|
+
Fast scout agents work best when a small scout model is already loaded beside
|
|
284
|
+
your main coding model on a local router. This is only safe when the combined
|
|
285
|
+
weights, KV caches, context windows, and parallel slots fit in GPU memory. If
|
|
286
|
+
the router spills into CPU RAM, both scout calls and main turns get slow.
|
|
287
|
+
|
|
288
|
+
Load both models manually on the target host, then point the orchestrator and
|
|
289
|
+
the scout worker profile at them. On llama.cpp routers, keep `max_instances`
|
|
290
|
+
at least as high as the number of models you want resident. Clio can see which
|
|
291
|
+
router instances are loaded and the router's instance limit, but current
|
|
292
|
+
llama.cpp router responses do not expose free VRAM, so confirming the loaded
|
|
293
|
+
set fits remains the operator's job. Workers on other nodes are unaffected.
|
|
294
|
+
|
|
295
|
+
## Safety and autonomy
|
|
296
|
+
|
|
297
|
+
There is one tool surface and one admission path. What changes is the autonomy
|
|
298
|
+
level, set in `/settings` or overridden for a single run with `--autonomy`.
|
|
299
|
+
|
|
300
|
+
| Level | Behavior |
|
|
301
|
+
| --- | --- |
|
|
302
|
+
| `read-only` | Inspection only. Every mutation and execution is denied. |
|
|
303
|
+
| `suggest` | Mutations are proposed and parked for your approval. |
|
|
304
|
+
| `auto-edit` | File edits proceed; execution and dispatch still gate. |
|
|
305
|
+
| `full-auto` | Approved classes proceed unattended, still inside damage-control rules. |
|
|
306
|
+
|
|
307
|
+
Notices name their mechanism so you always know who stopped a call:
|
|
308
|
+
`[safety-net]` for level-independent blocks, `[approval]` for parked calls,
|
|
309
|
+
`[autonomy]` for read-only denials, and `[middleware]` for hook diagnostics.
|
|
310
|
+
|
|
311
|
+
Workers can never exceed the orchestrator's authority. A dispatch request can
|
|
312
|
+
only narrow it, and reviewers and judges always run read-only. A `&&` chain is
|
|
313
|
+
judged at its most restrictive recognized member and refused whole if any
|
|
314
|
+
member is unrecognized, and `/tmp`, `/var/tmp`, and `/var/folders` are scratch
|
|
315
|
+
while the rest of the system roots stay protected. Details:
|
|
316
|
+
[docs/safety-model.md](docs/safety-model.md).
|
|
317
|
+
|
|
318
|
+
## The fleet: one machine or your whole cluster
|
|
319
|
+
|
|
320
|
+
Clio's orchestrator delegates work to bounded workers. With a fleet declared,
|
|
321
|
+
those workers run on other machines over SSH while every guarantee holds: one
|
|
322
|
+
admission path, one autonomy matrix, one receipt chain.
|
|
323
|
+
|
|
324
|
+
```mermaid
|
|
325
|
+
flowchart LR
|
|
326
|
+
U["you"] --> O["orchestrator TUI"]
|
|
327
|
+
O --> P["execution plan<br/>hashed DAG, capacity waves"]
|
|
328
|
+
P --> A["admission<br/>leases, queue, cost ceiling"]
|
|
329
|
+
A --> L["local worker"]
|
|
330
|
+
A --> S1["ssh node: blade"]
|
|
331
|
+
A --> S2["ssh node: dragon"]
|
|
332
|
+
L --> R["receipts and evidence"]
|
|
333
|
+
S1 --> R
|
|
334
|
+
S2 --> R
|
|
335
|
+
R --> O
|
|
336
|
+
```
|
|
337
|
+
|
|
338
|
+
Declare nodes in `settings.yaml` and the implicit `local` node is always
|
|
339
|
+
present:
|
|
340
|
+
|
|
341
|
+
```yaml
|
|
342
|
+
fleet:
|
|
343
|
+
nodes:
|
|
344
|
+
- id: blade
|
|
345
|
+
host: blade.example.net
|
|
346
|
+
maxWorkers: 2
|
|
347
|
+
residency: observe
|
|
348
|
+
- id: dragon
|
|
349
|
+
host: dragon.example.net
|
|
350
|
+
maxWorkers: 1
|
|
351
|
+
```
|
|
352
|
+
|
|
353
|
+
Then `clio-coder doctor` runs a per-node preflight, `clio-coder fleet list|run|status`
|
|
354
|
+
drives and observes contracts, and `clio-coder fleet drain|resume` closes or reopens
|
|
355
|
+
durable dispatch admission. A drain preserves running work, rejects every new
|
|
356
|
+
execution start, and expires after one hour unless renewed. The `/fleet`
|
|
357
|
+
overlay shows nodes, profiles, bindings, and live runs. Nodes must share the
|
|
358
|
+
project filesystem at the same absolute path; hosts that do not fail admission
|
|
359
|
+
with a clear reason. Target URLs resolve on the node the worker runs on, so
|
|
360
|
+
`localhost` means that node's own inference server and there is no central
|
|
361
|
+
proxy.
|
|
362
|
+
|
|
363
|
+
Everything you need to reproduce it end to end, including a recorded
|
|
364
|
+
multi-node demo script, is in [docs/fleet-dispatch.md](docs/fleet-dispatch.md)
|
|
365
|
+
and [docs/fleet-demo-runbook.md](docs/fleet-demo-runbook.md).
|
|
366
|
+
|
|
367
|
+
Routing is measured but conservative. Every dispatch records a joint decision
|
|
368
|
+
over agent, target, model, runtime, and node, while shadow mode leaves the
|
|
369
|
+
explicit route unchanged. Operators can activate only named read-only and
|
|
370
|
+
quality roles, and only after the exact tuple has enough integrity-valid
|
|
371
|
+
quality, reliability, cost, freshness, and decision-latency evidence:
|
|
372
|
+
|
|
373
|
+
```yaml
|
|
374
|
+
routing:
|
|
375
|
+
activeRoles: [researcher, verifier, reviewer, judge]
|
|
376
|
+
activePostures: [quality, balanced]
|
|
377
|
+
agentAutomation:
|
|
378
|
+
activeAgentRoles: [] # stays advisory until exact agent/role pairs are named
|
|
379
|
+
```
|
|
380
|
+
|
|
381
|
+
Manual pins and `failover: none` remain exact. Active mode fails closed when
|
|
382
|
+
no route is ready. `agent: auto` is separately bounded by recipe audience,
|
|
383
|
+
authority, tools, skills, result contract, locality, and approved governance;
|
|
384
|
+
changing from a read-only Scout phase to workspace editing requires an
|
|
385
|
+
authenticated plan approval or authority already granted by full-auto policy.
|
|
386
|
+
|
|
387
|
+
## Project context: CLIO-CODER.md
|
|
388
|
+
|
|
389
|
+
Clio loads a local `CLIO-CODER.md` as generated project context on every session.
|
|
390
|
+
`clio-coder context init` grounds a draft in the actual source tree, preserves an
|
|
391
|
+
existing handbook until an explicit replacement action, and can adopt existing
|
|
392
|
+
`CLAUDE.md`, `AGENTS.md`, `GEMINI.md`, Cursor, and Copilot context with
|
|
393
|
+
provenance and conflict reporting. `CLIO-CODER.md` is a gitignored runtime artifact,
|
|
394
|
+
not canonical repository documentation.
|
|
395
|
+
|
|
396
|
+
Alongside it, `clio-coder context index` builds a structural codewiki that the
|
|
397
|
+
`code_nav` tool navigates, so a model can find a symbol without reading half
|
|
398
|
+
the repository into its window.
|
|
399
|
+
|
|
400
|
+
## Skills
|
|
401
|
+
|
|
402
|
+
Skills are reusable `SKILL.md` guides the model loads on demand. Clio
|
|
403
|
+
discovers them from per-user and per-project roots, including `.clio-coder/skills`
|
|
404
|
+
and cross-harness layouts such as `.claude/skills` and `.codex/skills`. A
|
|
405
|
+
skill's `allowed-tools` declaration is enforced at tool admission, and a skill
|
|
406
|
+
can ship executable RED-GREEN evals that `clio-coder skills eval <name>` runs
|
|
407
|
+
instead of trusting the prose.
|
|
408
|
+
|
|
409
|
+
This repository ships a curated catalog under [skills/](skills/README.md) with
|
|
410
|
+
provenance frontmatter, evals, and content hashes pinned in
|
|
411
|
+
`skills/registry.yaml`, so an installed copy verifies against its audited
|
|
412
|
+
source at activation. Nothing auto-loads.
|
|
413
|
+
|
|
414
|
+
```bash
|
|
415
|
+
clio-coder skills install context-handoff # copy into .clio-coder/skills
|
|
416
|
+
clio-coder skills list # confirm Clio sees it
|
|
417
|
+
```
|
|
418
|
+
|
|
419
|
+
The catalog includes [`find-skills`](skills/meta/find-skills/), which routes
|
|
420
|
+
discovery through `clio-coder skills search` and `clio-coder skills install`. Install it
|
|
421
|
+
with `clio-coder skills install find-skills --user` so it outranks the community
|
|
422
|
+
skill of the same name that other installers drop into compat roots.
|
|
423
|
+
|
|
424
|
+
## Memory that survives long tasks
|
|
425
|
+
|
|
426
|
+
Long agentic runs decay: a requirement or a failed attempt is still in the
|
|
427
|
+
transcript but no longer influences the next action. Clio's proactive task
|
|
428
|
+
memory watches tool and lifecycle hooks, keeps a session task bank, and
|
|
429
|
+
surfaces visible advisory reminders at trigger boundaries.
|
|
430
|
+
|
|
431
|
+
The default tier is rules-only and makes no model calls. An LLM memory tier is
|
|
432
|
+
opt-in through an independent background route. The action agent's prompt and
|
|
433
|
+
tool surface never change, `/memory` inspects the bank, and disabling
|
|
434
|
+
`memory.intervention.enabled` removes the whole mechanism. Durable lessons are
|
|
435
|
+
separate, scoped, evidence-linked, and managed through `clio-coder memory
|
|
436
|
+
list|propose|approve|reject|prune`. Design notes:
|
|
437
|
+
[docs/proactive-memory.md](docs/proactive-memory.md).
|
|
438
|
+
|
|
439
|
+
## Status
|
|
440
|
+
|
|
441
|
+
Clio Coder is experimental software in a soft beta. The current release is
|
|
442
|
+
**v0.3.0**, installable from npm as
|
|
443
|
+
[`@iowarp/clio-coder`](https://www.npmjs.com/package/@iowarp/clio-coder) or
|
|
444
|
+
from source. Interfaces may still move between minor versions, and
|
|
445
|
+
model-specific behavior varies by target.
|
|
446
|
+
|
|
447
|
+
Release notes live in the [CHANGELOG](CHANGELOG.md), the implementation detail
|
|
448
|
+
behind each entry lives in the commit history, and every release is gated by
|
|
449
|
+
the deterministic `npm run ci:release` suite.
|
|
450
|
+
|
|
451
|
+
## Troubleshooting
|
|
452
|
+
|
|
453
|
+
| Problem | Try this |
|
|
454
|
+
| --- | --- |
|
|
455
|
+
| `clio-coder: command not found` | Run `npm run install:local`, then `hash -r`; confirm `${CLIO_CODER_BIN_DIR:-$HOME/.local/bin}` is on `PATH`. |
|
|
456
|
+
| No model target is available | Run `clio-coder configure`, then `clio-coder targets --probe`. |
|
|
457
|
+
| Local model does not respond | Confirm the local runtime is running and the target URL is correct. |
|
|
458
|
+
| Cloud model auth fails | Check `clio-coder auth status <target>` and verify the API key or login flow. |
|
|
459
|
+
| A fleet node never gets work | Run `clio-coder doctor`; per-node preflight reports filesystem parity and target facts. |
|
|
460
|
+
| Source changes do not appear | Re-run `npm run build`; the linked CLI points at `dist/`. |
|
|
461
|
+
| State appears corrupted | Run `clio-coder doctor`, then `clio-coder doctor --fix`. |
|
|
462
|
+
|
|
463
|
+
When filing an issue, include the output of `clio-coder --version`, `node
|
|
464
|
+
--version`, `clio-coder doctor`, and `clio-coder targets`. Redact secrets, private
|
|
465
|
+
prompts, logs, and proprietary code.
|
|
466
|
+
|
|
467
|
+
---
|
|
468
|
+
|
|
469
|
+
# For agents
|
|
470
|
+
|
|
471
|
+
If you are an AI agent operating inside this repository or driving Clio as a
|
|
472
|
+
tool, this section is the orientation you need.
|
|
473
|
+
|
|
474
|
+
## Orienting in this repository
|
|
475
|
+
|
|
476
|
+
Do not read broadly. Start from the codewiki, which indexes 927 source files.
|
|
477
|
+
Use `code_nav` in `entries`, `path`, or `symbol` mode before any wide read.
|
|
478
|
+
The indexed entry points are `src/cli/index.ts`, `src/domains/agents/index.ts`,
|
|
479
|
+
`src/domains/components/index.ts`, `src/domains/config/index.ts`,
|
|
480
|
+
`src/domains/context/bootstrap.ts`, `src/domains/context/index.ts`,
|
|
481
|
+
`src/domains/dispatch/index.ts`, and `src/domains/eval/index.ts`. When present,
|
|
482
|
+
read the local generated `CLIO-CODER.md` after that index-led orientation; it carries
|
|
483
|
+
the project-specific invariants and traps that are not obvious from the source.
|
|
484
|
+
|
|
485
|
+
## The tool surface
|
|
486
|
+
|
|
487
|
+
Twenty tools in seven planes. Each plane is one policy unit covering action
|
|
488
|
+
class, size posture, result schema, and concurrency rule.
|
|
489
|
+
|
|
490
|
+
| Plane | Tools | Posture |
|
|
491
|
+
| --- | --- | --- |
|
|
492
|
+
| OBSERVE | `read`, `grep`, `find`, `ls`, `code_nav`, `context`, `credential_present` | Read class, parallel, bounded by a truncation envelope |
|
|
493
|
+
| MUTATE | `write`, `edit` | Write class, sequential, queued through the file-mutation queue |
|
|
494
|
+
| EXECUTE | `bash`, `git`, `verify` | Containment posture; `bash` is default-deny, `git` is read-only inspection |
|
|
495
|
+
| ORCHESTRATE | `dispatch`, `monitor`, `steer`, `tasks`, `ledger` | Dispatch class, sequential except read-only `monitor` |
|
|
496
|
+
| RETRIEVE | `web_fetch` | Network read, parallel |
|
|
497
|
+
| INTERACT | `ask_user` | Host-owned operator interview |
|
|
498
|
+
| ARTIFACT | `artifact` | Plans, reviews, and reports as durable artifacts |
|
|
499
|
+
|
|
500
|
+
Every observation carries a truncation envelope with offload paths and next
|
|
501
|
+
hints, so a large result is bounded rather than silently cut. Full parameter
|
|
502
|
+
and payload reference: [docs/tool-usage.md](docs/tool-usage.md).
|
|
503
|
+
|
|
504
|
+
## Dispatch topologies
|
|
505
|
+
|
|
506
|
+
All topologies go through the same tool, admission chain, and autonomy matrix.
|
|
507
|
+
|
|
508
|
+
| Topology | Invocation | Semantics |
|
|
509
|
+
| --- | --- | --- |
|
|
510
|
+
| Singular | `task: "..."` | One assignment, with an optional separate `briefing`. |
|
|
511
|
+
| Parallel | `tasks: [...]` | Fan out, wait for all, one summary. |
|
|
512
|
+
| Sequential | `mode: "sequential"` | One at a time, stop reporting on timeout or abort. |
|
|
513
|
+
| Pipeline | `mode: "pipeline"` | Each step receives the previous step's output as data. |
|
|
514
|
+
| Detached | `detach: true` | Return assignment ids immediately and collect later. |
|
|
515
|
+
| Review gate | `review: {reviewer?, max_cycles?}` | Builder, read-only verifier verdict, bounded revise loop. |
|
|
516
|
+
| Compete | `mode: "compete", candidates: 2..4` | N candidates in scratch worktrees, read-only judge, winner applied. |
|
|
517
|
+
|
|
518
|
+
Detached batches are durable, so collection survives session exit. Use
|
|
519
|
+
`monitor` with `mode="collect"` as the authoritative terminal barrier over a
|
|
520
|
+
batch, and collect every detached batch before final synthesis; `mode="tools"`
|
|
521
|
+
reports what a run actually executed. `Alt+S` sends a running attached dispatch
|
|
522
|
+
to the background as a detached batch, which review gates, compete, pipelines,
|
|
523
|
+
and time-boxed calls refuse with a reason.
|
|
524
|
+
|
|
525
|
+
Admission refuses a pairing it can prove impossible, such as a pinned read-only
|
|
526
|
+
recipe aimed at mutation work, and flags the receipt where it is unsure rather
|
|
527
|
+
than blocking. Claimed file changes and validation commands are checked against
|
|
528
|
+
the run's own tool events, so a path the run never wrote cannot seal as done.
|
|
529
|
+
|
|
530
|
+
## Execution roles and typed results
|
|
531
|
+
|
|
532
|
+
Every dispatch carries an `ExecutionRole`: `builder`, `reviewer`, `judge`,
|
|
533
|
+
`researcher`, `verifier`, or `recovery`. The role is typed on every request,
|
|
534
|
+
ledger envelope, receipt, route candidate, plan task, and route decision.
|
|
535
|
+
Route statistics never mix roles, and any attempt after the first is
|
|
536
|
+
`recovery`.
|
|
537
|
+
|
|
538
|
+
Workers answer typed terminal contracts, not trailing prose. A `scout-report`
|
|
539
|
+
carries findings as `{claim, path, line}`, and grounding is structural rather
|
|
540
|
+
than a regex over prose: a cited line must fall inside a span this run
|
|
541
|
+
actually read, so an estimated line number cannot pass as an observation. A
|
|
542
|
+
worker validates its own result and spends a bounded number of repair rounds
|
|
543
|
+
before failing the run; the orchestrator's sealed validation is the authority.
|
|
544
|
+
|
|
545
|
+
Review and compete gates default to the builtin `verifier` and never fall back
|
|
546
|
+
to the builder agent. A gate decider's postcondition is the gate result
|
|
547
|
+
contract, not its own recipe contract.
|
|
548
|
+
|
|
549
|
+
## The built-in fleet
|
|
550
|
+
|
|
551
|
+
`architect`, `coder`, `tester`, `verifier`, `debugger`, `documenter`, `scout`,
|
|
552
|
+
`researcher`, `provenance`, and `git-master`. Each is a versioned frontmatter
|
|
553
|
+
recipe with an explicit tool profile, call and cost budget, and result
|
|
554
|
+
contract. Malformed custom recipes are quarantined with a diagnostic; malformed
|
|
555
|
+
builtins fail startup. Reference:
|
|
556
|
+
[docs/built-in-agents.md](docs/built-in-agents.md).
|
|
557
|
+
|
|
558
|
+
This roster is compiled into the session prompt whenever the dispatch tool is
|
|
559
|
+
available, so pin an id from it. `agent: "auto"` baselines from task shape and
|
|
560
|
+
is a fallback, not a router.
|
|
561
|
+
|
|
562
|
+
## Programmatic interfaces
|
|
563
|
+
|
|
564
|
+
```bash
|
|
565
|
+
clio-coder run "<task>" --json # one headless turn, JSONL events
|
|
566
|
+
clio-coder run "<task>" --agent coder # one explicit fleet agent, writes a receipt
|
|
567
|
+
clio-coder acp # serve ACP v1 over stdio for ACP frontends
|
|
568
|
+
clio-coder fleet run <contract> # run a fleet DAG contract
|
|
569
|
+
clio-coder fleet drain # pause new execution starts for up to one hour
|
|
570
|
+
clio-coder fleet resume # reopen durable dispatch admission
|
|
571
|
+
clio-coder evidence build|inspect|list # deterministic evidence artifacts
|
|
572
|
+
clio-coder eval validate|run|report|compare|gate
|
|
573
|
+
```
|
|
574
|
+
|
|
575
|
+
Dispatch can also delegate to external ACP agents while Clio mediates
|
|
576
|
+
permissions. Clio implements the
|
|
577
|
+
[Agent Client Protocol](https://agentclientprotocol.com) so the engine stays
|
|
578
|
+
decoupled from IDE frontends.
|
|
579
|
+
|
|
580
|
+
## What gets recorded
|
|
581
|
+
|
|
582
|
+
Every run seals a receipt. Receipt integrity is at v15 and covers normalized
|
|
583
|
+
routing intent, the resolved route, worker attestation, priced cost, phase
|
|
584
|
+
timing, tool activity, safety decisions, and result-contract conformance.
|
|
585
|
+
Gate decisions are v2 artifacts that seal route correlation across agent,
|
|
586
|
+
target, model family, runtime, and node, and they cross a staged durable
|
|
587
|
+
boundary rather than being written directly.
|
|
588
|
+
|
|
589
|
+
A worker attests its protocol version, pid, process-group id, host, settings
|
|
590
|
+
fingerprint, WorkerSpec digest, runtime, target, endpoint identity hash, wire
|
|
591
|
+
model, effective tool signature, and bounded resource facts before any model
|
|
592
|
+
call. Any drift from the approved identity kills the worker.
|
|
593
|
+
|
|
594
|
+
`clio-coder trace` reads the same store and now records interactive turns beside
|
|
595
|
+
dispatched runs, one event per tool call with its verdict.
|
|
596
|
+
|
|
597
|
+
Verify from the TUI with `/view verify <runId>`, or from the shell with `clio-coder evidence inspect`. See [docs/observability.md](docs/observability.md).
|
|
598
|
+
|
|
599
|
+
---
|
|
600
|
+
|
|
601
|
+
# For contributors
|
|
602
|
+
|
|
603
|
+
Contributions are welcome, and the fastest way in is to fix something you hit
|
|
604
|
+
while using it on your own research code.
|
|
605
|
+
|
|
606
|
+
## Architecture at a glance
|
|
607
|
+
|
|
608
|
+
```mermaid
|
|
609
|
+
flowchart TB
|
|
610
|
+
CLI["src/cli"] --> ENG["src/engine"]
|
|
611
|
+
TUI["src/interactive"] --> ENG
|
|
612
|
+
ENG --> TOOLS["src/tools<br/>20 typed tools, 7 planes"]
|
|
613
|
+
ENG --> DOM["src/domains"]
|
|
614
|
+
DOM --> DISP["dispatch<br/>plans, leases, routing, receipts"]
|
|
615
|
+
DOM --> CTX["context<br/>codewiki, compaction, CLIO-CODER.md"]
|
|
616
|
+
DOM --> PROV["providers<br/>runtimes, auth, catalog"]
|
|
617
|
+
DOM --> SAFE["safety<br/>damage control, policy"]
|
|
618
|
+
DISP --> WORK["src/worker<br/>bounded worker runtime"]
|
|
619
|
+
```
|
|
620
|
+
|
|
621
|
+
The largest indexed areas are `src/domains` (392 files), `tests/contracts`
|
|
622
|
+
(236), `src/interactive` (83), `src/cli` (48), `src/tools` (42), `src/engine`
|
|
623
|
+
(40), and `src/core` (35). Compile-time boundaries between domains are
|
|
624
|
+
enforced by a test suite, not by convention. Read
|
|
625
|
+
[docs/architecture.md](docs/architecture.md) before adding a cross-domain
|
|
626
|
+
import.
|
|
627
|
+
|
|
628
|
+
## Local development
|
|
629
|
+
|
|
630
|
+
```bash
|
|
631
|
+
npm ci
|
|
632
|
+
npm run dev # tsup watch build
|
|
633
|
+
npm run ci # the full local gate
|
|
634
|
+
```
|
|
635
|
+
|
|
636
|
+
Targeted checks when the risk is narrower:
|
|
637
|
+
|
|
638
|
+
| Check | Command |
|
|
639
|
+
| --- | --- |
|
|
640
|
+
| Types | `npm run typecheck` |
|
|
641
|
+
| Style | `npm run lint` |
|
|
642
|
+
| Contracts | `npm run test:contracts` |
|
|
643
|
+
| Smoke flows | `npm run test:smoke` |
|
|
644
|
+
| Domain boundaries | `npm run check:boundaries` |
|
|
645
|
+
| Everything | `npm run test` |
|
|
646
|
+
|
|
647
|
+
Conventions worth knowing before your first PR: local imports end in `.js`,
|
|
648
|
+
tests use `node:test`, and `any` needs a tracking issue.
|
|
649
|
+
|
|
650
|
+
## Release verification
|
|
651
|
+
|
|
652
|
+
```bash
|
|
653
|
+
npm run ci:release
|
|
654
|
+
```
|
|
655
|
+
|
|
656
|
+
That runs typecheck, Biome, the skills pin check, the production build, the
|
|
657
|
+
contract, smoke, and boundary suites, and the `check-release` dist and package
|
|
658
|
+
audit. Live model validation is separate, manual, and opt-in, because no
|
|
659
|
+
deterministic suite can promise that every local model behaves identically:
|
|
660
|
+
|
|
661
|
+
```bash
|
|
662
|
+
CLIO_CODER_LIVE_SMOKE=1 \
|
|
663
|
+
CLIO_CODER_LIVE_TARGET=openai-compat \
|
|
664
|
+
CLIO_CODER_LIVE_RUNTIME=openai-compat \
|
|
665
|
+
CLIO_CODER_LIVE_MODEL=your-model \
|
|
666
|
+
CLIO_CODER_LIVE_BASE_URL=http://localhost:8080/v1 \
|
|
667
|
+
npm run test:live
|
|
668
|
+
|
|
669
|
+
CLIO_CODER_LIVE_SMOKE=1 npm run test:live -- --delegation # needs local opencode and copilot
|
|
670
|
+
npm run test:live-eval:fleet-dispatch # multi-node dispatch regression
|
|
671
|
+
```
|
|
672
|
+
|
|
673
|
+
Benchmarks against public suites live under `benchmarks/`:
|
|
674
|
+
|
|
675
|
+
```bash
|
|
676
|
+
npm run bench:swe # SWE-bench Lite
|
|
677
|
+
npm run bench:scicode # SciCode
|
|
678
|
+
```
|
|
679
|
+
|
|
680
|
+
## Where to start
|
|
681
|
+
|
|
682
|
+
Read [CONTRIBUTING.md](CONTRIBUTING.md) for setup, architecture invariants,
|
|
683
|
+
branch and commit conventions, and the review rubric. Good first areas:
|
|
684
|
+
provider adapters for a runtime you use
|
|
685
|
+
([cookbook](docs/provider-adapter-cookbook.md)), skills for a scientific
|
|
686
|
+
domain you know ([catalog](skills/README.md)), and documentation gaps you hit
|
|
687
|
+
during onboarding.
|
|
688
|
+
|
|
689
|
+
Security reports go through [SECURITY.md](SECURITY.md), not public issues.
|
|
690
|
+
|
|
691
|
+
---
|
|
692
|
+
|
|
693
|
+
## Documentation
|
|
694
|
+
|
|
695
|
+
The full set lives under [docs/](docs/README.md), and `clio-coder docs` serves it
|
|
696
|
+
locally with interactive blueprints.
|
|
697
|
+
|
|
698
|
+
| Topic | Guide |
|
|
699
|
+
| --- | --- |
|
|
700
|
+
| Commands, slash commands, operating posture, keybindings, dispatch, verification, troubleshooting | [commands-and-modes.md](docs/commands-and-modes.md) |
|
|
701
|
+
| Multi-node fleet dispatch: SSH transport, doctor preflight, placement, topologies, receipts | [fleet-dispatch.md](docs/fleet-dispatch.md) |
|
|
702
|
+
| Executable multi-node demo with a reviewer gate and receipt provenance walkthrough | [fleet-demo-runbook.md](docs/fleet-demo-runbook.md) |
|
|
703
|
+
| NDJSON parent-child protocols, watchdog timers, and exit status mapping | [worker-dispatch-mechanics.md](docs/worker-dispatch-mechanics.md) |
|
|
704
|
+
| Built-in agent recipes, discovery roots, frontmatter schema, dispatch admission | [built-in-agents.md](docs/built-in-agents.md) |
|
|
705
|
+
| Context window resolution, probe capabilities, token accounting, compaction, priming | [context-engine.md](docs/context-engine.md) |
|
|
706
|
+
| Proactive task memory, session task bank, intervention rules, handoff carrying | [proactive-memory.md](docs/proactive-memory.md) |
|
|
707
|
+
| Runtime targets, local model configuration, fleet profiles, auth | [configuration-and-targets.md](docs/configuration-and-targets.md) |
|
|
708
|
+
| Argonne ALCF Sophia and Metis inference targets over Globus OAuth | [alcf-provider.md](docs/alcf-provider.md) |
|
|
709
|
+
| Safety posture, default-deny Bash, project policy, damage-control rules, typed validation | [safety-model.md](docs/safety-model.md) |
|
|
710
|
+
| Source layout, compile-time boundaries, domain loading, runtime data flow | [architecture.md](docs/architecture.md) |
|
|
711
|
+
| Reference for all 19 worker tools: parameters, payloads, error examples | [tool-usage.md](docs/tool-usage.md) |
|
|
712
|
+
| Prompt envelope reuse, provider tool delivery, bounded tool results | [prompt-envelope-and-tools.md](docs/prompt-envelope-and-tools.md) |
|
|
713
|
+
| Implementing custom model runtimes and inference server integrations | [provider-adapter-cookbook.md](docs/provider-adapter-cookbook.md) |
|
|
714
|
+
| Artifact browsing, receipt verification, dispatch diagnostics, observability routing | [observability.md](docs/observability.md) |
|
|
715
|
+
| Evidence directory structures, findings, operator-approved memory retrieval | [evidence-and-memory.md](docs/evidence-and-memory.md) |
|
|
716
|
+
| Local YAML eval suites, reports, comparisons, command evidence | [eval-runner.md](docs/eval-runner.md) |
|
|
717
|
+
| Installation, upgrade, reset, uninstallation, configuration folders, permissions | [installation-and-lifecycle.md](docs/installation-and-lifecycle.md) |
|
|
718
|
+
| Every environment variable the runtime reads | [environment-variables.md](docs/environment-variables.md) |
|
|
719
|
+
| Prompt and skill resources, extension manifests, portable share archives | [extensions-and-sharing.md](docs/extensions-and-sharing.md) |
|
|
720
|
+
| Skills Hub marketplace discovery, install actions, publishing | [skills-marketplace.md](docs/skills-marketplace.md) |
|
|
721
|
+
| Runtime model refresh, catalog sources, local and cloud model quirks | [model-catalog.md](docs/model-catalog.md) |
|
|
722
|
+
| Active component snapshots and the experimental middleware hook contract | [middleware-and-components.md](docs/middleware-and-components.md) |
|
|
723
|
+
| Advisory validation-contract patterns for scientific artifacts and HPC assumptions | [scientific-validation.md](docs/scientific-validation.md) |
|
|
724
|
+
| Falsifiable Change Manifest templates, auditability, and `clio-coder evolve` | [evolution.md](docs/evolution.md) |
|
|
725
|
+
| Interface layout, palette, Unicode vocabulary, drawing choreography | [tui-design.md](docs/tui-design.md) |
|
|
726
|
+
| Source-first docs workflow, mapping matrix, alpha wording guidance | [documentation-guide.md](docs/documentation-guide.md) |
|
|
727
|
+
| Private context index determinism and target smoke matrices (internal) | [evals-internal.md](docs/evals-internal.md) |
|
|
728
|
+
| Point-in-time inventory of legacy environment variables (historical) | [config-knobs-audit.md](docs/config-knobs-audit.md) |
|
|
729
|
+
|
|
730
|
+
## Measuring local model performance
|
|
731
|
+
|
|
732
|
+
llama.cpp and similar backends often expose a single prefix-cache slot. When
|
|
733
|
+
dispatch traffic or compaction invalidates it, the next turn records the
|
|
734
|
+
expected-cold reasons and shows one dim notice. Per-call cache verdicts
|
|
735
|
+
(`hot`, `partial`, `cold`, `small`) are persisted with timing and prompt-cache
|
|
736
|
+
counters in each session's `context-snapshots.jsonl`, so a slow session can be
|
|
737
|
+
diagnosed from the ledger alone.
|
|
738
|
+
|
|
739
|
+
```bash
|
|
740
|
+
clio-coder usage report --days 7 # cost and token facts with cited run ids
|
|
741
|
+
```
|
|
742
|
+
|
|
743
|
+
Inside the TUI, `/cost` shows session totals and `/context` opens the
|
|
744
|
+
context-window ledger. See [docs/context-engine.md](docs/context-engine.md)
|
|
745
|
+
for how the context engine measures and protects the prompt prefix.
|
|
746
|
+
|
|
747
|
+
---
|
|
748
|
+
|
|
749
|
+
## Heritage, lineage, and funding
|
|
750
|
+
|
|
751
|
+
Clio Coder is developed under the [IOWarp](https://iowarp.ai) project by the
|
|
752
|
+
[Gnosis Research Center](https://grc.iit.edu) at the
|
|
753
|
+
[Illinois Institute of Technology](https://www.iit.edu) in collaboration with
|
|
754
|
+
the University of Utah.
|
|
755
|
+
|
|
756
|
+
IOWarp and the CLIO (Context Layer for Input/Output) architecture are funded by
|
|
757
|
+
the National Science Foundation under
|
|
758
|
+
[Award #2411318](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2411318) for
|
|
759
|
+
2024 through 2029. Principal Investigator: Dr. Xian-He Sun. Co-Principal
|
|
760
|
+
Investigators: Dr. Anthony Kougkas, Dr. Jake Hochhalter, and Dr. Vivek
|
|
761
|
+
Srikumar.
|
|
762
|
+
|
|
763
|
+
Clio Coder is the interactive coding orchestrator in a larger ecosystem:
|
|
764
|
+
|
|
765
|
+
- [clio-core](https://github.com/iowarp/clio-core) is the foundational storage
|
|
766
|
+
layer using Chimaera-based tiered data and context storage.
|
|
767
|
+
- [clio-kit](https://github.com/iowarp/clio-kit) is a suite of 15+
|
|
768
|
+
[Model Context Protocol](https://modelcontextprotocol.io) servers exposing
|
|
769
|
+
150+ tools for scientific computing domains including HDF5, Slurm, ParaView,
|
|
770
|
+
Pandas, ArXiv, NetCDF, FITS, and Zarr.
|
|
771
|
+
|
|
772
|
+
### Built on
|
|
773
|
+
|
|
774
|
+
- **Pi Agent Framework** from [Earendil Works](https://github.com/earendil-works):
|
|
775
|
+
the [@earendil-works/pi-ai](https://www.npmjs.com/package/@earendil-works/pi-ai)
|
|
776
|
+
execution engine, [@earendil-works/pi-tui](https://www.npmjs.com/package/@earendil-works/pi-tui)
|
|
777
|
+
terminal rendering, and [@earendil-works/pi-agent-core](https://www.npmjs.com/package/@earendil-works/pi-agent-core)
|
|
778
|
+
subagent orchestration.
|
|
779
|
+
- **Anthropic Claude Agent SDK** through
|
|
780
|
+
[@anthropic-ai/claude-agent-sdk](https://www.npmjs.com/package/@anthropic-ai/claude-agent-sdk)
|
|
781
|
+
for Claude Code worker runs under Pro/Max subscriptions.
|
|
782
|
+
- **Agent Client Protocol** for decoupling the engine from IDE frontends.
|
|
783
|
+
- **Globus Auth** for authenticating against ALCF's Sophia and Metis inference
|
|
784
|
+
gateways.
|
|
785
|
+
|
|
786
|
+
### Evaluation
|
|
787
|
+
|
|
788
|
+
Subagents and prompt techniques are evaluated against
|
|
789
|
+
[SWE-bench](https://www.swebench.com) and SciCode. Every subagent run produces
|
|
790
|
+
structured execution evidence, matched against baseline and candidate
|
|
791
|
+
evaluations to catch silent regressions.
|
|
792
|
+
|
|
793
|
+
---
|
|
794
|
+
|
|
795
|
+
<p align="center">
|
|
796
|
+
Licensed under Apache-2.0. See <a href="LICENSE">LICENSE</a> and <a href="NOTICE">NOTICE</a>.<br />
|
|
797
|
+
<sub>Built for the people who maintain the code that science runs on.</sub>
|
|
798
|
+
</p>
|