@iowarp/clio-coder 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +407 -0
- package/CODE_OF_CONDUCT.md +21 -0
- package/CONTRIBUTING.md +224 -0
- package/LICENSE +202 -0
- package/NOTICE +9 -0
- package/README.md +798 -0
- package/SECURITY.md +72 -0
- package/assets/clio-coder-logo-128.webp +0 -0
- package/damage-control-rules.yaml +419 -0
- package/dist/acp-UMLFVA3F.js +92 -0
- package/dist/agents-Q4MYPMUW.js +91 -0
- package/dist/auth-O6HYIJ6J.js +521 -0
- package/dist/chunk-262G75JS.js +35 -0
- package/dist/chunk-26BZQOAD.js +1281 -0
- package/dist/chunk-2J63S4SF.js +508 -0
- package/dist/chunk-3DANZDGR.js +717 -0
- package/dist/chunk-4UQA7NCT.js +29 -0
- package/dist/chunk-527KG6XR.js +497 -0
- package/dist/chunk-5LDRNKX2.js +1063 -0
- package/dist/chunk-5N2FG33Q.js +25 -0
- package/dist/chunk-67MTHP2E.js +135 -0
- package/dist/chunk-6CWDTGUC.js +20 -0
- package/dist/chunk-7BHLZB3A.js +2115 -0
- package/dist/chunk-7RBKDI66.js +348 -0
- package/dist/chunk-AMFR5YA3.js +541 -0
- package/dist/chunk-BBUH4VAA.js +1224 -0
- package/dist/chunk-BYEU76JP.js +899 -0
- package/dist/chunk-CLJ5HLUD.js +458 -0
- package/dist/chunk-D5YD55AR.js +116 -0
- package/dist/chunk-DXQNI4PC.js +61 -0
- package/dist/chunk-E3NYWENM.js +1004 -0
- package/dist/chunk-GNGDQYDU.js +34688 -0
- package/dist/chunk-GOTUR54M.js +9 -0
- package/dist/chunk-HBU5MTAM.js +41 -0
- package/dist/chunk-HMYNFFY4.js +28 -0
- package/dist/chunk-JPOWPFCU.js +1010 -0
- package/dist/chunk-JWHCJDCI.js +1215 -0
- package/dist/chunk-KBR4MZZR.js +41 -0
- package/dist/chunk-KKKPTZLM.js +93 -0
- package/dist/chunk-ME6DNWIU.js +66 -0
- package/dist/chunk-NI4DEJMC.js +88 -0
- package/dist/chunk-O4EJEDHO.js +659 -0
- package/dist/chunk-PIDUD6M2.js +31 -0
- package/dist/chunk-PS4PFJQP.js +29459 -0
- package/dist/chunk-QV47YRF4.js +48 -0
- package/dist/chunk-RQDWMVRB.js +279 -0
- package/dist/chunk-TFSSEXL6.js +136 -0
- package/dist/chunk-TKHQ4DGZ.js +8290 -0
- package/dist/chunk-TPOCL34A.js +2876 -0
- package/dist/chunk-UGYAX5YI.js +565 -0
- package/dist/chunk-UHTSULZS.js +461 -0
- package/dist/chunk-UU3R62TT.js +128 -0
- package/dist/chunk-UWIJNAOB.js +3906 -0
- package/dist/chunk-VOO7NYPP.js +914 -0
- package/dist/chunk-VPAWTYLY.js +117 -0
- package/dist/chunk-WD6AJM35.js +1216 -0
- package/dist/chunk-X3BR7HWV.js +115 -0
- package/dist/chunk-X3NE4WVW.js +120 -0
- package/dist/chunk-XNISANGE.js +1395 -0
- package/dist/chunk-XV4ZJ6ZM.js +3177 -0
- package/dist/cli/index.js +236 -0
- package/dist/clio-KIQ5SNDS.js +53 -0
- package/dist/components-JVHMUBEB.js +653 -0
- package/dist/config-ZFCDBMDC.js +372 -0
- package/dist/configure-G4E3A2PG.js +27 -0
- package/dist/context-CDXTP2MP.js +293 -0
- package/dist/context-E3KIFVXI.js +185 -0
- package/dist/context-clear-3F4PLXOS.js +102 -0
- package/dist/context-index-Q7YSYTR3.js +106 -0
- package/dist/docs-YIETIWZI.js +280 -0
- package/dist/doctor-M5HJJZOL.js +61 -0
- package/dist/domains/agents/builtins/architect.md +33 -0
- package/dist/domains/agents/builtins/coder.md +31 -0
- package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
- package/dist/domains/agents/builtins/debugger.md +30 -0
- package/dist/domains/agents/builtins/documenter.md +31 -0
- package/dist/domains/agents/builtins/git-master.md +30 -0
- package/dist/domains/agents/builtins/provenance.md +30 -0
- package/dist/domains/agents/builtins/researcher.md +71 -0
- package/dist/domains/agents/builtins/scout.md +42 -0
- package/dist/domains/agents/builtins/tester.md +31 -0
- package/dist/domains/agents/builtins/verifier.md +30 -0
- package/dist/domains/agents/builtins/wiki-writer.md +41 -0
- package/dist/eval-B3KZZESM.js +2674 -0
- package/dist/evidence-V67CHM35.js +233 -0
- package/dist/evolve-YDZSUQYA.js +518 -0
- package/dist/extensions-SRG7XCAH.js +207 -0
- package/dist/fleet-CA2CRTVG.js +760 -0
- package/dist/fleet-preflight-CLIAX7YR.js +21 -0
- package/dist/init-2OZDJE2D.js +227 -0
- package/dist/memory-3PIQQAKX.js +207 -0
- package/dist/models-DY35XI7Y.js +237 -0
- package/dist/paths-5OMXW7Z4.js +57 -0
- package/dist/preload-KZVHET2B.js +11 -0
- package/dist/reset-PIFYNOS3.js +216 -0
- package/dist/run-3VSPP24F.js +735 -0
- package/dist/share-D36RQCXM.js +241 -0
- package/dist/skills-F2MRLELY.js +445 -0
- package/dist/skills-eval-E2ZTW4PL.js +932 -0
- package/dist/targets-DZMEZAH4.js +977 -0
- package/dist/trace-7NYCUI2J.js +250 -0
- package/dist/uninstall-AD3JWHBB.js +322 -0
- package/dist/upgrade-WYYBKGDY.js +301 -0
- package/dist/usage-ULIDAGFF.js +755 -0
- package/dist/version-ROZ6CZKH.js +16 -0
- package/dist/wiki-generate-PKFIX6OB.js +377 -0
- package/dist/worker/entry.js +1739 -0
- package/docs/README.md +93 -0
- package/docs/acp.md +120 -0
- package/docs/alcf-provider.md +72 -0
- package/docs/architecture.md +172 -0
- package/docs/artifact-versions.md +54 -0
- package/docs/built-in-agents.md +265 -0
- package/docs/capacity-and-scheduling.md +97 -0
- package/docs/commands-and-modes.md +554 -0
- package/docs/config-knobs-audit.md +115 -0
- package/docs/configuration-and-targets.md +812 -0
- package/docs/context-engine.md +236 -0
- package/docs/dispatch-architecture-rationale.md +126 -0
- package/docs/documentation-coverage.md +46 -0
- package/docs/documentation-guide.md +166 -0
- package/docs/environment-variables.md +105 -0
- package/docs/eval-runner.md +205 -0
- package/docs/evals-internal.md +298 -0
- package/docs/evidence-and-memory.md +243 -0
- package/docs/evolution.md +143 -0
- package/docs/exit-codes-and-output.md +74 -0
- package/docs/extensions-and-sharing.md +306 -0
- package/docs/fleet-demo-runbook.md +179 -0
- package/docs/fleet-dispatch.md +591 -0
- package/docs/glossary.md +75 -0
- package/docs/html/agents_blueprint.html +936 -0
- package/docs/html/alcf_blueprint.html +324 -0
- package/docs/html/architecture_blueprint.html +850 -0
- package/docs/html/commands_blueprint.html +794 -0
- package/docs/html/config_knobs_audit_blueprint.html +178 -0
- package/docs/html/configuration_blueprint.html +1080 -0
- package/docs/html/context_blueprint.html +603 -0
- package/docs/html/documentation_blueprint.html +832 -0
- package/docs/html/environment_blueprint.html +404 -0
- package/docs/html/eval_blueprint.html +743 -0
- package/docs/html/evals_internal_blueprint.html +190 -0
- package/docs/html/evolution_blueprint.html +674 -0
- package/docs/html/extensions_blueprint.html +2065 -0
- package/docs/html/fleet_dispatch_blueprint.html +286 -0
- package/docs/html/index.html +919 -0
- package/docs/html/lifecycle_blueprint.html +723 -0
- package/docs/html/memory_blueprint.html +699 -0
- package/docs/html/middleware_blueprint.html +664 -0
- package/docs/html/models_blueprint.html +2366 -0
- package/docs/html/observability_blueprint.html +683 -0
- package/docs/html/provider_adapter_blueprint.html +245 -0
- package/docs/html/safety_blueprint.html +1386 -0
- package/docs/html/shared.css +571 -0
- package/docs/html/shared.js +143 -0
- package/docs/html/skills_blueprint.html +671 -0
- package/docs/html/soak_blueprint.html +182 -0
- package/docs/html/tool_usage_blueprint.html +350 -0
- package/docs/html/tools_blueprint.html +2249 -0
- package/docs/html/trace_blueprint.html +235 -0
- package/docs/html/tui_design_blueprint.html +314 -0
- package/docs/html/validation_blueprint.html +961 -0
- package/docs/html/worker_dispatch_blueprint.html +231 -0
- package/docs/installation-and-lifecycle.md +308 -0
- package/docs/middleware-and-components.md +148 -0
- package/docs/model-catalog.md +189 -0
- package/docs/observability.md +233 -0
- package/docs/proactive-memory.md +452 -0
- package/docs/prompt-envelope-and-tools.md +142 -0
- package/docs/provider-adapter-cookbook.md +148 -0
- package/docs/release-cut-checklist.md +138 -0
- package/docs/safety-model.md +357 -0
- package/docs/scientific-validation.md +105 -0
- package/docs/session-lifecycle.md +156 -0
- package/docs/skills-marketplace.md +46 -0
- package/docs/tool-usage.md +527 -0
- package/docs/trace-store.md +132 -0
- package/docs/troubleshooting.md +33 -0
- package/docs/tui-design.md +239 -0
- package/docs/worker-dispatch-mechanics.md +242 -0
- package/package.json +132 -0
- package/skills/README.md +408 -0
- package/skills/git/commit-crafting/SKILL.md +79 -0
- package/skills/git/commit-crafting/evals.md +92 -0
- package/skills/git/create-pr/SKILL.md +116 -0
- package/skills/git/create-pr/evals.md +114 -0
- package/skills/git/investigate-issue/SKILL.md +139 -0
- package/skills/git/investigate-issue/evals.md +94 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
- package/skills/git/resolve-merge-conflicts/evals.md +58 -0
- package/skills/git/review-changes/SKILL.md +103 -0
- package/skills/git/review-changes/evals.md +85 -0
- package/skills/git/worktree-create/SKILL.md +92 -0
- package/skills/git/worktree-create/evals.md +97 -0
- package/skills/git/worktree-create/references/worktree-setup.md +66 -0
- package/skills/git/worktree-merge/SKILL.md +95 -0
- package/skills/git/worktree-merge/evals.md +114 -0
- package/skills/skill-marketplace.json +261 -0
- package/skills/workflow/cut-it/SKILL.md +86 -0
- package/skills/workflow/cut-it/evals.md +42 -0
- package/src/domains/agents/builtins/architect.md +33 -0
- package/src/domains/agents/builtins/coder.md +31 -0
- package/src/domains/agents/builtins/context-bootstrap.md +38 -0
- package/src/domains/agents/builtins/debugger.md +30 -0
- package/src/domains/agents/builtins/documenter.md +31 -0
- package/src/domains/agents/builtins/git-master.md +30 -0
- package/src/domains/agents/builtins/provenance.md +30 -0
- package/src/domains/agents/builtins/researcher.md +71 -0
- package/src/domains/agents/builtins/scout.md +42 -0
- package/src/domains/agents/builtins/tester.md +31 -0
- package/src/domains/agents/builtins/verifier.md +30 -0
- package/src/domains/agents/builtins/wiki-writer.md +41 -0
- package/src/domains/agents/fleets/build-review.md +34 -0
- package/src/domains/agents/fleets/build-test.md +35 -0
- package/src/domains/agents/fleets/sdlc.md +86 -0
- package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
- package/src/domains/prompts/fragments/identity/clio.md +26 -0
- package/src/domains/prompts/fragments/operating/contract.md +64 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
- package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
- package/src/domains/prompts/fragments/safety/read-only.md +13 -0
- package/src/domains/prompts/fragments/safety/suggest.md +13 -0
- package/src/domains/prompts/fragments/wiki/page.md +75 -0
- package/src/domains/prompts/fragments/wiki/plan.md +48 -0
- package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
|
@@ -0,0 +1,148 @@
|
|
|
1
|
+
# Provider Adapter Cookbook
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive runtime adapter descriptor builder and probe sequence capability checklist is located at [docs/html/provider_adapter_blueprint.html](html/provider_adapter_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
This cookbook guides developers through implementing custom model runtimes and inference server integrations within Clio Coder. It explains the runtime descriptor interfaces, probing protocols, model synthesis, and how to configure reasoning and thinking behaviors.
|
|
7
|
+
|
|
8
|
+
Source of truth:
|
|
9
|
+
- Runtime descriptor types: [src/domains/providers/types/runtime-descriptor.ts](../src/domains/providers/types/runtime-descriptor.ts)
|
|
10
|
+
- Registry loader: [src/domains/providers/registry.ts](../src/domains/providers/registry.ts)
|
|
11
|
+
- Probe reasoning helpers: [src/domains/providers/probe/reasoning.ts](../src/domains/providers/probe/reasoning.ts)
|
|
12
|
+
- Model capabilities resolver: [src/domains/providers/model-capabilities.ts](../src/domains/providers/model-capabilities.ts)
|
|
13
|
+
- Inference capability flags: [src/domains/providers/types/capability-flags.ts](../src/domains/providers/types/capability-flags.ts)
|
|
14
|
+
- Model target resolution: [src/domains/providers/runtime-resolution.ts](../src/domains/providers/runtime-resolution.ts)
|
|
15
|
+
|
|
16
|
+
---
|
|
17
|
+
|
|
18
|
+
## 1. Anatomy of a Runtime Descriptor
|
|
19
|
+
|
|
20
|
+
Every model runtime (e.g., Local Native, Cloud HTTP, Subprocess) implements the `RuntimeDescriptor` interface defined in `src/domains/providers/types/runtime-descriptor.ts`.
|
|
21
|
+
|
|
22
|
+
Here is a template for a new runtime plugin:
|
|
23
|
+
|
|
24
|
+
```typescript
|
|
25
|
+
import { Type } from "typebox";
|
|
26
|
+
import type { Api, Model } from "@earendil-works/pi-ai";
|
|
27
|
+
import type {
|
|
28
|
+
RuntimeDescriptor,
|
|
29
|
+
ProbeContext,
|
|
30
|
+
ProbeResult,
|
|
31
|
+
ReasoningProbeResult,
|
|
32
|
+
} from "../../types/runtime-descriptor.js";
|
|
33
|
+
import { type CapabilityFlags } from "../../types/capability-flags.js";
|
|
34
|
+
|
|
35
|
+
export const myCustomRuntime: RuntimeDescriptor = {
|
|
36
|
+
id: "my-custom-service",
|
|
37
|
+
displayName: "My Custom Service Native Client",
|
|
38
|
+
kind: "http", // "http" | "sdk" | "subprocess"
|
|
39
|
+
tier: "local-native", // "protocol" | "cloud" | "local-native" | "subscription"
|
|
40
|
+
apiFamily: "openai-responses", // api model mapping class
|
|
41
|
+
auth: "api-key", // "api-key" | "oauth" | "aws-sdk" | "vertex-adc" | "claude-cli" | "none"
|
|
42
|
+
credentialsEnvVar: "CUSTOM_SERVICE_API_KEY",
|
|
43
|
+
|
|
44
|
+
defaultCapabilities: {
|
|
45
|
+
chat: true,
|
|
46
|
+
tools: true,
|
|
47
|
+
toolCallFormat: "openai",
|
|
48
|
+
reasoning: false,
|
|
49
|
+
structuredOutputs: "json-schema",
|
|
50
|
+
vision: false,
|
|
51
|
+
audio: false,
|
|
52
|
+
embeddings: false,
|
|
53
|
+
rerank: false,
|
|
54
|
+
fim: false,
|
|
55
|
+
contextWindow: 8192,
|
|
56
|
+
maxTokens: 4096,
|
|
57
|
+
},
|
|
58
|
+
|
|
59
|
+
// Probes target endpoint health and loaded models.
|
|
60
|
+
async probe(target, ctx): Promise<ProbeResult> {
|
|
61
|
+
// Implementation here (see Section 2)
|
|
62
|
+
},
|
|
63
|
+
|
|
64
|
+
// Synthesizes the model client for execution.
|
|
65
|
+
synthesizeModel(target, wireModelId, kb): Model<Api> {
|
|
66
|
+
// Implementation here (see Section 3)
|
|
67
|
+
}
|
|
68
|
+
};
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## 2. Probing Mechanisms
|
|
74
|
+
|
|
75
|
+
Probes discover the current state of a target inference server when Clio starts or when `/targets` / `/models` are refreshed.
|
|
76
|
+
|
|
77
|
+
### 2.1 Endpoint Probing (`probe`)
|
|
78
|
+
The `probe` method validates endpoint reachability and collects loaded models:
|
|
79
|
+
|
|
80
|
+
* **Inputs:** `TargetDescriptor` (which holds target `url`, optional `apiKey`, and connection metadata) and `ProbeContext` (which provides timeout signals and credentials).
|
|
81
|
+
* **Return Value:** A `ProbeResult` indicating:
|
|
82
|
+
* `ok`: True if reachable.
|
|
83
|
+
* `serverVersion`: String identifier of the backend (e.g. `"Ollama/0.1.48"`).
|
|
84
|
+
* `models`: A list of strings representing the currently loaded/selectable models.
|
|
85
|
+
* `modelStates` (optional): Footprint mappings detailing VRAM/RAM loading stats.
|
|
86
|
+
|
|
87
|
+
### 2.2 Reasoning Probing (`probeReasoning`)
|
|
88
|
+
For local endpoints where models are loaded dynamically, the runtime can supply a `probeReasoning` method. It sends a short mock completion request to inspect whether the model outputs reasoning/thinking tags (such as `reasoning_content` in OpenAI completions or `<think>` tags in raw text streams).
|
|
89
|
+
|
|
90
|
+
Clio caches this result under the session's provider cache, preventing redundant network requests.
|
|
91
|
+
|
|
92
|
+
### 2.3 Exact-ID Capability Selection (`probeCapabilitiesForModel`)
|
|
93
|
+
`probeCapabilitiesForModel` is the one exact-id selector during capability resolution. When a router target serves several models, `probeCapabilitiesForModel` matches `probeModelCapabilities` keyed strictly to the requested wire model ID. A router serving multiple models thus answers only from the `/v1/models` row keyed to its own wire model, preventing capability flags or token limits from bleeding across different models on the same target.
|
|
94
|
+
|
|
95
|
+
---
|
|
96
|
+
|
|
97
|
+
## 3. Model Synthesis
|
|
98
|
+
|
|
99
|
+
The `synthesizeModel` method acts as the factory that creates the `pi-ai` compatible client interface for model turns.
|
|
100
|
+
|
|
101
|
+
* **Signature:**
|
|
102
|
+
```typescript
|
|
103
|
+
synthesizeModel(
|
|
104
|
+
target: TargetDescriptor,
|
|
105
|
+
wireModelId: string,
|
|
106
|
+
kb: KnowledgeBaseHit | null
|
|
107
|
+
): Model<Api>
|
|
108
|
+
```
|
|
109
|
+
* **Tasks:**
|
|
110
|
+
1. Retrieve configured API credentials using `providers.auth` persisted through `openAuthStorage()`.
|
|
111
|
+
2. Instantiate the adapter client (e.g., building a `pi-ai` OpenAI or Anthropic provider instance).
|
|
112
|
+
3. Bind custom prompt templates and FIM (Fill-in-the-Middle) properties where supported.
|
|
113
|
+
|
|
114
|
+
---
|
|
115
|
+
|
|
116
|
+
## 4. Configuring Reasoning & Thinking Formats
|
|
117
|
+
|
|
118
|
+
Clio supports diverse thinking mechanisms. If your model family uses a custom format, it must be mapped to one of the following mechanisms in the model's catalog YAML entry (`quirks.thinking.mechanism`):
|
|
119
|
+
|
|
120
|
+
| Mechanism | Behavior |
|
|
121
|
+
| --- | --- |
|
|
122
|
+
| `none` | **Reasoning-Never:** Clio strips thinking request fields (e.g., effort levels), avoids replaying thinking blocks in history, emits no TUI thinking events, and records no reasoning token usage metrics. |
|
|
123
|
+
| `ollama-native` | Standard Ollama native thinking streams. |
|
|
124
|
+
| `lmstudio-native` | Assistant thinking is prepended to output payloads wrapped in `<think>` and `</think>` tags. |
|
|
125
|
+
| `openai-completions` | Replays thinking blocks via `reasoning_content` message parameters. |
|
|
126
|
+
| `anthropic-max` | Anthropic extended thinking block protocol. |
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## 5. Adding the Adapter to Clio
|
|
131
|
+
|
|
132
|
+
Once your runtime adapter descriptor is implemented:
|
|
133
|
+
|
|
134
|
+
### 5.1 Static Built-in Registration
|
|
135
|
+
Add your descriptor to the static array export in [src/domains/providers/runtimes/builtins.ts](../src/domains/providers/runtimes/builtins.ts):
|
|
136
|
+
```typescript
|
|
137
|
+
import { myCustomRuntime } from "./custom/my-custom-runtime.js";
|
|
138
|
+
|
|
139
|
+
export const BUILTIN_RUNTIMES = [
|
|
140
|
+
// ...
|
|
141
|
+
myCustomRuntime,
|
|
142
|
+
];
|
|
143
|
+
```
|
|
144
|
+
|
|
145
|
+
### 5.2 Dynamic Plugin Loading
|
|
146
|
+
Clio's `RuntimeRegistry` can load custom runtimes dynamically at startup:
|
|
147
|
+
* **Directories:** Place compiled Javascript descriptors (`.js`) inside `$CLIO_CODER_CONFIG_DIR/runtimes/` (defaulting to `~/.config/clio-coder/runtimes/`).
|
|
148
|
+
* **Package exports:** Publish an npm package that exports a `clioRuntimes` array containing your runtime descriptors, then list the package name under `runtimePlugins` in your configuration settings.
|
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# v0.3.0 Release-Cut Checklist
|
|
2
|
+
|
|
3
|
+
The ordered steps that turn the prepared `v0.3.0` branch into a published
|
|
4
|
+
release. Everything above the line marked **AUTHORIZATION BOUNDARY** is
|
|
5
|
+
repeatable and reversible and was run during the hardening sessions. Everything
|
|
6
|
+
below it is external or destructive, was deliberately **not run**, and needs an
|
|
7
|
+
explicit decision from the operator.
|
|
8
|
+
|
|
9
|
+
Nothing in this checklist has been performed against `main`, a remote, a tag,
|
|
10
|
+
or the npm registry.
|
|
11
|
+
|
|
12
|
+
## Status of the prepared tree
|
|
13
|
+
|
|
14
|
+
| Item | State |
|
|
15
|
+
| --- | --- |
|
|
16
|
+
| Branch | `v0.3.0`, local only |
|
|
17
|
+
| `package.json` version | `0.3.0`, **not bumped by the hardening sessions** |
|
|
18
|
+
| `main` | untouched |
|
|
19
|
+
| Remotes | not contacted |
|
|
20
|
+
| Tags | none created |
|
|
21
|
+
| npm registry | not contacted |
|
|
22
|
+
|
|
23
|
+
---
|
|
24
|
+
|
|
25
|
+
## Part 1: verification (repeatable, already run)
|
|
26
|
+
|
|
27
|
+
1. `npm run typecheck`
|
|
28
|
+
2. `npm run lint`
|
|
29
|
+
3. `npm run check:boundaries`
|
|
30
|
+
4. `npm run build`
|
|
31
|
+
5. `npm run test`
|
|
32
|
+
6. `npm run test:trace-viewer`
|
|
33
|
+
7. `npm run ci` (runs 1, 2, `skills:check`, 4, 5, 6)
|
|
34
|
+
8. `npm run ci` again under the other supported Node major. Both Node 22 and
|
|
35
|
+
Node 24 must be green; the repo is developed against 22.22.3 and 24.9.0.
|
|
36
|
+
9. `npm run test:lifecycle` for the twenty-case lifecycle matrix against a real
|
|
37
|
+
`npm pack` installed into a temporary prefix. Case 9 needs `--live` plus
|
|
38
|
+
`CLIO_CODER_LIFECYCLE_URL` and `CLIO_CODER_LIFECYCLE_MODEL` naming a target whose model
|
|
39
|
+
is already resident.
|
|
40
|
+
10. `npm run ci:release`, which adds `scripts/check-release.mjs`: dist shebang
|
|
41
|
+
integrity, the forbidden-file list, the required runtime resources, and the
|
|
42
|
+
tarball and unpacked size budgets.
|
|
43
|
+
|
|
44
|
+
## Part 2: version and notes (repeatable, NOT run)
|
|
45
|
+
|
|
46
|
+
These edit the working tree only. They are reversible with `git checkout` and
|
|
47
|
+
are listed here because the hardening sessions were explicitly scoped out of
|
|
48
|
+
performing them.
|
|
49
|
+
|
|
50
|
+
11. Decide the released version. The tree currently reads `0.3.0` in
|
|
51
|
+
`package.json`. If that is the number to publish, no bump is needed; confirm
|
|
52
|
+
it deliberately rather than by default.
|
|
53
|
+
12. Files carrying a version reference, to update together if the number
|
|
54
|
+
changes:
|
|
55
|
+
- `package.json` (`version`)
|
|
56
|
+
- `CHANGELOG.md` (the `## 0.3.0 - <date>` heading and its date)
|
|
57
|
+
- `docs/environment-variables.md` and `docs/tui-design.md` (the
|
|
58
|
+
`(Version: 0.3.0)` markers on the interactive-blueprint tips)
|
|
59
|
+
- `docs/html/*.html` (the `Blueprint (v0.3.0)` titles)
|
|
60
|
+
- `scripts/check-release.mjs` (the measured-at figures in the budget
|
|
61
|
+
comment, if the package size moved materially)
|
|
62
|
+
13. Confirm the `## 0.3.0` section of `CHANGELOG.md` describes every
|
|
63
|
+
user-visible behavior change in the release, including the ones that alter
|
|
64
|
+
existing behavior:
|
|
65
|
+
- unknown slash commands now fail instead of reaching the model as chat
|
|
66
|
+
- `--remove-binary` launcher ownership is identity, not a path shape
|
|
67
|
+
- `reset` and `uninstall` exit 1 on partial failure instead of reporting
|
|
68
|
+
success
|
|
69
|
+
14. Re-run `npm run ci:release` after any version edit.
|
|
70
|
+
15. Commit the version and notes as one commit on `v0.3.0`.
|
|
71
|
+
|
|
72
|
+
---
|
|
73
|
+
|
|
74
|
+
## AUTHORIZATION BOUNDARY
|
|
75
|
+
|
|
76
|
+
Every step below leaves the local checkout, is externally visible, or cannot be
|
|
77
|
+
undone by a local `git` command. **None of them has been run.**
|
|
78
|
+
|
|
79
|
+
## Part 3: clean-install verification (external, NOT run)
|
|
80
|
+
|
|
81
|
+
16. **NOT RUN** — `npm pack` and install the resulting tarball into a fresh
|
|
82
|
+
temporary prefix on a machine that has never had Clio installed, with empty
|
|
83
|
+
XDG roots. `npm run test:lifecycle` covers this on the development machine;
|
|
84
|
+
a second machine is what proves no developer-local state is load-bearing.
|
|
85
|
+
17. **NOT RUN** — From that install, verify: `clio-coder --version`, `clio-coder --help`,
|
|
86
|
+
an empty-state non-TTY launch, `clio-coder configure` to a real target,
|
|
87
|
+
`clio-coder doctor`, one real turn, and `clio-coder uninstall --dry-run`.
|
|
88
|
+
18. **NOT RUN** — Inspect the artifact by hand: `tar -tzf` the tarball, confirm
|
|
89
|
+
no source maps, no `scripts/`, no `tests/`, no `benchmarks/`, no
|
|
90
|
+
`apps/trace-viewer`, and that `skills/`, `docs/*.md`, `docs/html/`, the
|
|
91
|
+
builtin agents, the model catalogs, and `damage-control-rules.yaml` are all
|
|
92
|
+
present.
|
|
93
|
+
|
|
94
|
+
## Part 4: branch integration (destructive to history, NOT run)
|
|
95
|
+
|
|
96
|
+
19. **NOT RUN** — Decide how `v0.3.0` reaches `main`. The hardening sessions
|
|
97
|
+
were forbidden to merge, rebase, or modify `main`, so no integration
|
|
98
|
+
strategy has been chosen or attempted.
|
|
99
|
+
20. **NOT RUN** — Integrate, then re-run `npm run ci:release` on the integrated
|
|
100
|
+
result. A gate that passed on the branch has not passed on the merge.
|
|
101
|
+
|
|
102
|
+
## Part 5: tag and push (external, NOT run)
|
|
103
|
+
|
|
104
|
+
21. **NOT RUN** — `git tag -a v0.3.0 -m "..."`.
|
|
105
|
+
22. **NOT RUN** — `git push origin <branch>`.
|
|
106
|
+
23. **NOT RUN** — `git push origin v0.3.0`.
|
|
107
|
+
|
|
108
|
+
## Part 6: publication (external and irreversible, NOT run)
|
|
109
|
+
|
|
110
|
+
24. **NOT RUN** — `npm publish`. Note that `prepublishOnly` runs
|
|
111
|
+
`npm run ci:release`, so publication re-gates the tree; that is a safety
|
|
112
|
+
net and not a substitute for step 20.
|
|
113
|
+
25. **NOT RUN** — Decide the dist-tag. Publishing to `latest` makes this the
|
|
114
|
+
default install for every user. An experimental release may warrant
|
|
115
|
+
`--tag next` instead; the CLI and README both describe v0.3.0 as
|
|
116
|
+
experimental, which argues for it.
|
|
117
|
+
26. **NOT RUN** — A published version cannot be replaced. `npm unpublish` is
|
|
118
|
+
restricted and time-limited, and a mistake is corrected by publishing a
|
|
119
|
+
higher version, not by removing the wrong one.
|
|
120
|
+
|
|
121
|
+
## Part 7: post-publish verification (external, NOT run)
|
|
122
|
+
|
|
123
|
+
27. **NOT RUN** — On a clean machine, `npm install -g @iowarp/clio-coder` from
|
|
124
|
+
the registry rather than from a local tarball, then repeat step 17 against
|
|
125
|
+
it. This is the only step that tests what users actually receive.
|
|
126
|
+
28. **NOT RUN** — Verify `clio-coder upgrade` finds and applies the published
|
|
127
|
+
version from an installation of the previous release.
|
|
128
|
+
29. **NOT RUN** — Publish the GitHub release with the `CHANGELOG.md` section
|
|
129
|
+
for this version.
|
|
130
|
+
|
|
131
|
+
---
|
|
132
|
+
|
|
133
|
+
## Rollback
|
|
134
|
+
|
|
135
|
+
There is no rollback for step 24. If a defect is found after publication, the
|
|
136
|
+
correction is a patch release. Before step 24, every step is reversible:
|
|
137
|
+
steps 21 through 23 by deleting the local and remote tag and force-updating the
|
|
138
|
+
branch, and steps 11 through 15 by `git reset`.
|
|
@@ -0,0 +1,357 @@
|
|
|
1
|
+
# Clio Coder Safety Model
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/safety_blueprint.html](html/safety_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
Clio Coder's safety posture is code-enforced, not prompt-only. As the orchestrator coding agent in the [IOWarp](https://iowarp.ai) ecosystem developed by the [Gnosis Research Center](https://grc.iit.edu) at Illinois Tech under NSF Award [#2411318](https://www.nsf.gov/awardsearch/showAward?AWD_ID=2411318), Clio gates execution by target capabilities, the tool registry, the safety policy engine, project policies, protected-artifact checks, and audit receipts.
|
|
7
|
+
|
|
8
|
+
Source of truth: `src/domains/safety/**`, `src/tools/registry.ts`, `src/tools/bootstrap.ts`, `src/tools/policy.ts`, `src/entry/orchestrator.ts`, `src/domains/dispatch/write-boundary.ts`, and `damage-control-rules.yaml`.
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Two axes: autonomy and the safety net
|
|
13
|
+
|
|
14
|
+
The `autonomy` setting (`read-only` | `suggest` | `auto-edit` | `full-auto`) is an enforced dial. It controls exactly one thing: which action classes run immediately, which park for operator approval, and which are auto-denied. The safety net (damage-control rules, path policy, protected artifacts, loop guard, dispatch scope admission) is independent of the dial and identical at every level. When a `[safety-net]` notice appears at full-auto, that is the always-on net working as designed, not a contradiction of the level.
|
|
15
|
+
|
|
16
|
+
In Clio Coder v0.3.0, effective autonomy resolution is strictly centralized in `src/entry/orchestrator.ts` through `resolveEffectiveAutonomy` and `resolveBaselineAutonomy`. Every admission surface (tool registry admission, dispatch plan provenance, and ACP session snapshots) delegates to this pair of functions so that fallback paths cannot diverge across execution contexts. `resolveBaselineAutonomy` evaluates dispatch settings overrides, headless CLI options, and configuration settings before applying the default `auto-edit` level. `resolveEffectiveAutonomy` combines any active ACP session autonomy level with the baseline resolution.
|
|
17
|
+
|
|
18
|
+
### Autonomy levels
|
|
19
|
+
|
|
20
|
+
| Action class | `read-only` | `suggest` | `auto-edit` (default) | `full-auto` |
|
|
21
|
+
|---|---|---|---|---|
|
|
22
|
+
| `read` | allow | allow | allow | allow |
|
|
23
|
+
| `write` | deny | ask | allow | allow |
|
|
24
|
+
| `execute`: builtin no-prompt set + project commands | deny | ask | allow | allow |
|
|
25
|
+
| `execute`: any other bash | deny | ask | ask | allow |
|
|
26
|
+
| `dispatch` | deny | ask | allow | allow |
|
|
27
|
+
| `system_modify` | deny | ask | ask | ask |
|
|
28
|
+
| `git_destructive` | net block | net block | net block | net block |
|
|
29
|
+
| `unknown` | deny | ask | ask | ask |
|
|
30
|
+
|
|
31
|
+
- **`read-only`**: Clio inspects and answers. Mutating calls are auto-denied with a rejection telling the model to propose the change instead; approvals are never invoked. Denials render as `[autonomy]` notices.
|
|
32
|
+
- **`suggest`**: every non-read action parks for one-shot approval. The operator drives.
|
|
33
|
+
- **`auto-edit`**: workspace edits and recognized commands run; unrecognized bash asks instead of blocking.
|
|
34
|
+
- **`full-auto`**: bash runs without prompting; the net is the protection, not the prompt. `system_modify` still asks because it reaches outside the workspace; `unknown` still asks because the net cannot reason about calls it cannot classify.
|
|
35
|
+
|
|
36
|
+
The `system_modify` confirm is level-invariant, so it is enforced and attributed as a safety-net confirm rail: the overlay, notices, and audit ledger name the net (reason code `system-modify-confirm`, policy source `builtin-classifier`), not the autonomy level. The matrix row above is unchanged in outcome at every level; only `read-only` converts the ask to a denial. `unknown` remains in the autonomy mapping because the registry substitutes a registered tool's base action class after the net evaluates.
|
|
37
|
+
|
|
38
|
+
The level is persisted as `autonomy` in `settings.yaml`, hot-reloads, and is edited in the `/settings` Autonomy & Safety section.
|
|
39
|
+
|
|
40
|
+
---
|
|
41
|
+
|
|
42
|
+
## Enforcement path
|
|
43
|
+
|
|
44
|
+
Every tool call, orchestrator or worker, evaluates in this order:
|
|
45
|
+
|
|
46
|
+
1. **Safety net** (policy engine + middleware guards): `block` is final at every level; `ask` is a confirm rail (damage-control `ask` rules, project `requireConfirmation`, `system_modify`) that parks at every level; `pass` hands off to step 2. Blocks precede asks: a damage-control `ask` rule never bypasses a hard block, so confirming an ask-rule command that targets a zero-access path still blocks. The built-in path protection (which includes zero-access blocklists for critical files like `.git/config` and `credentials.yaml`, resolved with symlink canonicalization to prevent bypasses) is evaluated even when `.clio-coder/safety.yaml` is malformed, invalid, or attempts to override it. A malformed project policy cannot disable built-in default path protection, so credential protection never fails open.
|
|
47
|
+
2. **Autonomy mapping**: the action class plus the level produce allow, ask, or deny per the matrix above.
|
|
48
|
+
3. **Approvals**: whatever asked in step 1 or 2 parks interactively, denies deterministically headless, resolves per `workers.onPermission` in workers, and non-stall denies in delegations.
|
|
49
|
+
|
|
50
|
+
```mermaid
|
|
51
|
+
graph TD
|
|
52
|
+
user[User request] --> surface[Provider tool capability]
|
|
53
|
+
surface --> registry[Tool registry admission]
|
|
54
|
+
registry --> net[Safety net verdict: block / confirm / pass]
|
|
55
|
+
net --> autonomy[Autonomy mapping: allow / ask / deny]
|
|
56
|
+
autonomy --> approvals[Approvals: park, deny, or resolve per context]
|
|
57
|
+
approvals --> middleware[Middleware + protected artifacts]
|
|
58
|
+
middleware --> run[Tool execution]
|
|
59
|
+
run --> shape[Result shaping]
|
|
60
|
+
shape --> receipt[Receipts, audit, evidence]
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
Net `confirm` is never auto-allowed by autonomy, including full-auto. Net `block` is never downgraded by anything. A tool hidden by target capability or explicit suppression is not shown to the model; a tool that is shown still must pass this path before it can run.
|
|
64
|
+
|
|
65
|
+
### Worker permission escalation
|
|
66
|
+
|
|
67
|
+
Dispatched workers run non-interactively, so step 3 resolves per `workers.onPermission`: `deny` turns the parked call into a structured denial, `fail` ends the run, and `escalate` hands the ask up to the interactive operator. Under `escalate` the worker parks the call, emits a `clio_permission_escalated` event, and waits; dispatch republishes the ask on the bus tagged with the run id; the operator resolves it in the same permission overlay used for the main agent; and the decision returns down the worker's stdin. Resolution is human-only by construction: no model-facing tool can approve a worker permission, and the dispatch `resolveWorkerPermission` method is reachable only from the interactive layer. This preserves the receipt's honesty, since a model approving its own fleet's asks would collapse the audit trail.
|
|
68
|
+
|
|
69
|
+
Escalation can never hang a run. Every escalated ask resolves by an operator decision or by the `workers.escalation` timeout fallback (`{ timeoutMs, fallback }`, defaults 120000 ms and `deny`); a headless session has no subscriber, so the timeout fallback always governs there. The worker keeps emitting heartbeats while parked, so the reconciler does not reap it, and every escalation and its resolution source (operator or timeout) is recorded on the receipt.
|
|
70
|
+
|
|
71
|
+
---
|
|
72
|
+
|
|
73
|
+
## Operating Posture and Visible Tools
|
|
74
|
+
|
|
75
|
+
Clio operates under a single operating posture with a standard, unified visible toolset. The 19 built-in tools are organized in seven planes; each plane is one policy unit for action class, size posture, and concurrency, asserted at bootstrap by `src/tools/policy.ts` so the classifier and the registered specs can never drift apart silently.
|
|
76
|
+
|
|
77
|
+
| Plane | Tools | Action class |
|
|
78
|
+
| --- | --- | --- |
|
|
79
|
+
| OBSERVE | `read`, `grep`, `find`, `ls`, `code_nav`, `context`, `credential_present` | `read` |
|
|
80
|
+
| MUTATE | `write`, `edit` | `write` |
|
|
81
|
+
| EXECUTE | `bash`, `verify` | `execute` |
|
|
82
|
+
| EXECUTE | `git` | `read` |
|
|
83
|
+
| ORCHESTRATE | `dispatch`, `steer` | `dispatch` |
|
|
84
|
+
| ORCHESTRATE | `monitor`, `tasks` | `read` |
|
|
85
|
+
| RETRIEVE | `web_fetch` | `read` |
|
|
86
|
+
| INTERACT | `ask_user` | `read` |
|
|
87
|
+
| ARTIFACT | `artifact` | `write` |
|
|
88
|
+
|
|
89
|
+
`git` is read-only inspection on the safe-exec spine, so it carries the read class despite living in the EXECUTE plane; `monitor` and `tasks` never mutate a run or the workspace, so they stay read class inside the ORCHESTRATE plane. `gateway` is a design-reserved name only (see `src/core/tool-names.ts`), not a registered tool.
|
|
90
|
+
|
|
91
|
+
Target capability, dispatch tool profiles, and recipe constraints can further narrow the tools available to a run. That narrowing is convenience and budget control; safety still lives in code gates.
|
|
92
|
+
|
|
93
|
+
---
|
|
94
|
+
|
|
95
|
+
## Skill tool surface narrowing
|
|
96
|
+
|
|
97
|
+
A `SKILL.md` may declare `allowed-tools` and `disallowed-tools`. The declaration is enforced at tool admission, between the safety net and the autonomy mapping, on every surface that activates skills (interactive turns, headless `clio-coder run` turns, and dispatched workers whose recipes declare skills).
|
|
98
|
+
|
|
99
|
+
- **Window.** Narrowing arms when `context` (scope="skills") successfully loads the skill and lasts for the lifetime of the pending-skill policy: to the end of the current turn for the main agent, and to the end of the run for a worker. A later turn is unrestricted until a skill is requested and loaded again.
|
|
100
|
+
- **Merge.** Denials win: a tool named in any loaded skill's `disallowed-tools` is blocked. Allow-narrowing applies only while every loaded skill declares `allowed-tools`; the merged surface is the union of those lists. A loaded skill that declares no `allowed-tools` keeps the full surface for its own workflow, which lifts the allow-narrowing (never the denials) for that window.
|
|
101
|
+
- **Exemptions.** `context` (the remaining requested skills of the turn must still load) and `ask_user` (the escape hatch the block message points at) are always admitted.
|
|
102
|
+
- **Direction.** Narrowing only blocks. It never grants a tool the safety net, damage-control rules, or autonomy mapping would refuse, and an out-of-surface call blocks terminally instead of parking for confirmation.
|
|
103
|
+
- **Block message.** The rejection names the tool, the active skill(s), and the merged surface, and states the remediation: work within the declared surface, or use `ask_user` (when available) to hand the step to the operator. The audit row carries reason code `skill_surface`.
|
|
104
|
+
|
|
105
|
+
---
|
|
106
|
+
|
|
107
|
+
## Damage-control rules
|
|
108
|
+
|
|
109
|
+
`damage-control-rules.yaml` is compiled into rule packs. Base rules apply broadly. Rules with `ask: true` park for one-shot confirmation (for example `git stash drop` and remote-branch deletion); hard-block rules and classifier-pattern `git_destructive` hits are always blocked. Both behaviors apply at every autonomy level.
|
|
110
|
+
|
|
111
|
+
Examples of patterns the rules target include destructive filesystem operations, dangerous device writes, fork bombs, pipe-to-shell installers, and destructive git operations.
|
|
112
|
+
|
|
113
|
+
Safety policy metadata records active rule IDs and hashes so receipts/evidence can explain which rule pack was active.
|
|
114
|
+
|
|
115
|
+
---
|
|
116
|
+
|
|
117
|
+
## Bash recognition
|
|
118
|
+
|
|
119
|
+
The policy engine tags every bash command as recognized or unrecognized; the autonomy mapping decides what happens next:
|
|
120
|
+
|
|
121
|
+
1. **Recognized**: a valid `.clio-coder/safety.yaml` command entry, or the narrow built-in no-prompt set such as `pwd`, simple `ls`, `git status`, bounded `git diff/log`, common test/lint/build commands, `pytest`, `cargo test`, `go test`, or `make test`. In addition, compound `&&` chains where every segment is a recognized safe command (up to `CHAIN_MAX_SEGMENTS = 6`) are recognized. Commands wrapped inside `sh -c` are expanded up to depth 3 (`INNER_SHELL_MAX_DEPTH = 3`) and evaluated segment by segment. Recognized commands run without prompting at `auto-edit` and `full-auto`.
|
|
122
|
+
2. **Unrecognized**: anything else. Unrecognized bash asks for one-shot approval at `auto-edit` and runs at `full-auto`; at `suggest` it asks like every mutation, and at `read-only` it is denied.
|
|
123
|
+
|
|
124
|
+
Shell operators split two ways. Unrecognized sequencing and redirection (`||`, `;`, pipes, redirects, newlines, or `&&` chains with unapproved commands or more than 6 segments) make a command unrecognized: it asks at `auto-edit` and runs at `full-auto`. Command substitution (`$(...)`, backticks) is a net confirm rail at every level, full-auto included, because the net cannot scan what it executes until runtime. The rule pack scans the full command string before either check, so a destructive verb behind an operator is caught regardless of level. Project policy entries reject operator kinds unless the entry sets `shellOperators: allow`.
|
|
125
|
+
|
|
126
|
+
Bash `cwd` is resolved under the workspace root. Escaping the workspace is blocked unless a reviewed project policy permits the exact command/cwd combination.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## Policy Engine Evaluation Order
|
|
131
|
+
|
|
132
|
+
Source: `src/domains/safety/policy-engine.ts`.
|
|
133
|
+
|
|
134
|
+
Every tool call entering the safety engine passes through a strict 10-step evaluation sequence. Safety-net blocks are final; autonomy mapping applies only after the safety net passes:
|
|
135
|
+
|
|
136
|
+
1. **Damage-Control Scan**: Evaluates compiled rule packs (`damage-control-rules.yaml`) against command strings, paths, and tool arguments.
|
|
137
|
+
2. **Write-Root Containment**: Verifies that mutations remain within configured write boundaries (`evaluateWriteRoots`).
|
|
138
|
+
3. **Hard Blocks**: Enforces unconditional blocks against destructive actions (such as `git_destructive` operations and zero-access credentials).
|
|
139
|
+
4. **Invalid Project Policy Fail-Closed**: If `.clio-coder/safety.yaml` contains parsing or schema errors, execution tools fail closed.
|
|
140
|
+
5. **Path Policy Enforcement**: Evaluates `zeroAccessPaths`, `readOnlyPaths`, and `noDeletePaths`.
|
|
141
|
+
6. **Bash Zero-Access Protocol**: Rejects shell commands attempting to read, exfiltrate, or redirect from protected credential files. A safe presence-check exception (`grep -sq "^NAME=" <file>`) is permitted for environment probing without exposing secrets.
|
|
142
|
+
7. **Ask Rails**: Evaluates rules requiring confirmation (such as project `requireConfirmation` or unanalyzable command substitutions `$(...)`).
|
|
143
|
+
8. **System Modify Checks**: Assesses operating-system level modification commands targeting system roots (`/etc`, `/usr`, `/var`, `/bin`, `/sbin`, `/run`), exempting temporary write paths (`/var/tmp`, `/var/folders`).
|
|
144
|
+
9. **Bash Allowlist Recognition**: Checks whether the command matches the known safe command allowlist or compound `&&` chain.
|
|
145
|
+
10. **Default Allow**: If no prior rule intervened, the action proceeds to autonomy-level evaluation.
|
|
146
|
+
|
|
147
|
+
---
|
|
148
|
+
|
|
149
|
+
## Project safety policy
|
|
150
|
+
|
|
151
|
+
Clio searches upward from the current working directory for `.clio-coder/safety.yaml`. The file is parsed once into a loaded policy. Invalid policy files fail closed for execution tools.
|
|
152
|
+
|
|
153
|
+
Minimal schema v1:
|
|
154
|
+
|
|
155
|
+
```yaml
|
|
156
|
+
version: 1
|
|
157
|
+
zeroAccessPaths:
|
|
158
|
+
- secrets/
|
|
159
|
+
- .env
|
|
160
|
+
readOnlyPaths:
|
|
161
|
+
- vendor/
|
|
162
|
+
noDeletePaths:
|
|
163
|
+
- out/validated/
|
|
164
|
+
commands:
|
|
165
|
+
- id: local-test
|
|
166
|
+
command: npm test
|
|
167
|
+
cwd: .
|
|
168
|
+
timeoutMs: 120000
|
|
169
|
+
maxOutputBytes: 600000
|
|
170
|
+
actionClass: execute
|
|
171
|
+
shellOperators: deny
|
|
172
|
+
env:
|
|
173
|
+
mode: none
|
|
174
|
+
allow: []
|
|
175
|
+
requireConfirmation: false
|
|
176
|
+
rationale: Standard local test command.
|
|
177
|
+
owner: maintainers
|
|
178
|
+
comment: Keep exact and reviewed.
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
Accepted root keys:
|
|
182
|
+
|
|
183
|
+
```text
|
|
184
|
+
version | commands | tasks | disableDefaultPathPolicy | zeroAccessPaths | readOnlyPaths | noDeletePaths
|
|
185
|
+
```
|
|
186
|
+
|
|
187
|
+
`tasks` is an alias for command policy entries. Unknown keys, wrong types, duplicate command IDs, absolute `cwd`, `..`-escaping `cwd`, and invalid path-policy entries make the policy invalid.
|
|
188
|
+
|
|
189
|
+
Path-policy behavior:
|
|
190
|
+
|
|
191
|
+
| Key | Effect |
|
|
192
|
+
| --- | --- |
|
|
193
|
+
| `zeroAccessPaths` | Blocks read, write, and delete. |
|
|
194
|
+
| `readOnlyPaths` | Allows read, blocks write/delete. |
|
|
195
|
+
| `noDeletePaths` | Blocks delete. |
|
|
196
|
+
| `disableDefaultPathPolicy` | Uses only project path policy rather than merging default damage-control paths. Note: This cannot disable built-in zero-access protection for critical system paths like `.git/config` and `credentials.yaml`. |
|
|
197
|
+
|
|
198
|
+
Command entry notes:
|
|
199
|
+
|
|
200
|
+
| Field | Meaning |
|
|
201
|
+
| --- | --- |
|
|
202
|
+
| `shellOperators` | `deny` by default; `allow` only for exact reviewed commands. |
|
|
203
|
+
| `env` | `mode: none` by default; `mode: allowlist` permits named environment variables. |
|
|
204
|
+
| `requireConfirmation` | Parks the call for super confirmation instead of immediate allow. |
|
|
205
|
+
|
|
206
|
+
---
|
|
207
|
+
|
|
208
|
+
## Typed validation tools
|
|
209
|
+
|
|
210
|
+
Prefer typed tools over Bash:
|
|
211
|
+
|
|
212
|
+
- `git` (op=status/diff/log) uses fixed command vectors.
|
|
213
|
+
- `verify(check="<script>")` runs a declared package.json verification script (the `test*/lint*/build*/typecheck*/check*/format*/ci*` family) through bounded execution helpers with no shell; `verify()` with no arguments lists the declared checks.
|
|
214
|
+
- `verify(check="frontend", path=...)` validates frontend artifacts without granting arbitrary shell access.
|
|
215
|
+
|
|
216
|
+
The frontend check accepts `.html`, `.htm`, `.css`, `.js`, `.mjs`, and `.cjs` under the workspace root. It checks HTML tag balance, local script/style references, JavaScript syntax, CSS brace/comment/string balance, and optionally loads HTML with an available headless Chromium/Chrome/Edge executable (`browser: auto|required|off`).
|
|
217
|
+
|
|
218
|
+
The `edit` tool also carries conservative matching rules. It preserves
|
|
219
|
+
unchanged bytes, handles common quote, dash, whitespace, and indentation drift
|
|
220
|
+
for matching only, and rejects ambiguous or no-op edits rather than applying a
|
|
221
|
+
guess.
|
|
222
|
+
|
|
223
|
+
### Write boundaries: detect-and-rollback
|
|
224
|
+
|
|
225
|
+
Post-run enforcement of declared write boundaries is strictly detect-and-rollback, never sandboxing. Nothing prevents a write during step execution: no container, no seccomp filter, no read-only mount, and no filesystem interception of the worker's tools. A step executes with whatever permissions its environment provides.
|
|
226
|
+
|
|
227
|
+
After step execution, the orchestrator compares the working checkout against the git snapshot recorded before the step (`src/domains/dispatch/write-boundary.ts`). It identifies modifications outside declared write boundaries, rolls back unauthorized file changes, and fails the step with a typed failure reason (`WriteBoundaryViolation`). Operators requiring steps to be physically incapable of writing outside designated paths must employ OS-level isolation, which Clio Coder deliberately does not claim to provide.
|
|
228
|
+
|
|
229
|
+
---
|
|
230
|
+
|
|
231
|
+
## Dispatch runtimes
|
|
232
|
+
|
|
233
|
+
Fleet dispatch is admitted only when the requested worker scope is a subset of the orchestrator scope and requested actions fit the worker scope.
|
|
234
|
+
|
|
235
|
+
Dispatch workers can run the same HTTP, native, or pi-ai-backed runtimes as the orchestrator, driven through the [pi SDK family](https://www.npmjs.com/package/@earendil-works/pi-agent-core) (including `@earendil-works/pi-agent-core` and its TUI and AI wrappers). Clio observes and governs those tool calls directly, so every worker run is subject to the same safety mapping and receipt accounting as an interactive turn.
|
|
236
|
+
|
|
237
|
+
Three integration paths exist for driving Claude Code, ranging from fully enforced to advisory gating:
|
|
238
|
+
|
|
239
|
+
- **`claude-sdk` (Enforced Safety):** Drives [@anthropic-ai/claude-agent-sdk](https://www.npmjs.com/package/@anthropic-ai/claude-agent-sdk) directly. This is the **strong safety path** because Clio enforces tool gating before execution. Clio registers a `PreToolUse` hook (which fires for all tool uses, including auto-allowed reads) and wraps `canUseTool` for permission paths. Every tool request is mapped into a Clio tool/action class, evaluated by the safety net, and passed through the active autonomy matrix. Because a dispatched worker is noninteractive, any `ask` decision is resolved as a non-stall denial (`workers.onPermission=deny` returns denial; `workers.onPermission=fail` terminates the run with a permission-required code).
|
|
240
|
+
- **`claude-code` (Subprocess Gating):** Drives `claude -p` as a subprocess. Because the CLI lacks a direct callback hook, Clio cannot evaluate each tool invocation. Instead, Clio maps the active autonomy level to the binary's command-line parameters (such as `--permission-mode` and tool allowlists). Unrecognized tools are gated by the subprocess runtime itself. Dispatch at autonomy `suggest` is refused outright (the same applies to `antigravity-code`): a subprocess cannot park a tool call for approval, so `suggest` has no honest mapping and the runner fails closed before launching the external CLI. A dangerous bypass (`--allow-dangerously-skip-permissions`) is only sent when autonomy is `full-auto` and `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1`, and it is never silent: the run's receipt records it (see the enforcement grades below) and evidence raises an external-bypass finding.
|
|
241
|
+
- **Claude Code over ACP (Advisory Gating):** Drives Zed's `@zed-industries/claude-code-acp` (or `@agentclientprotocol/claude-agent-acp`) bridge as an [Agent Client Protocol (ACP)](https://agentclientprotocol.com) delegation agent. Clio's ACP mediator intercepts tool calls and filters them against the safety net, but gating is ultimately **advisory** as Claude governs its own runtime execution. For strict, code-enforced per-tool safety, `claude-sdk` is preferred over ACP.
|
|
242
|
+
|
|
243
|
+
All Claude Code runtimes rely on the user's existing CLI authentication and store no credentials in Clio.
|
|
244
|
+
|
|
245
|
+
### Autonomy enforcement grades
|
|
246
|
+
|
|
247
|
+
How faithfully a runtime can honor the autonomy model is a recorded fact, not an assumption. Worker receipts carry an optional `autonomyEnforcement` block sealed into the integrity digest:
|
|
248
|
+
|
|
249
|
+
- **`mediated`**: per-call evaluation through Clio's net and autonomy mapping (native workers, `claude-sdk`).
|
|
250
|
+
- **`approximated`**: the posture is mapped to external harness flags with no per-call mediation (`claude-code`, `antigravity-code` subprocesses).
|
|
251
|
+
- **`bypassed`**: a dangerous external mode ran under `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS=1`, where even net blocks do not exist.
|
|
252
|
+
|
|
253
|
+
Evidence raises a warn-level external-bypass finding for bypassed runs and an info-level approximation note otherwise, and the provenance projection surfaces the block in `transcript.md`, `trace.cleaned.jsonl`, and compact dispatch output. The product statement is: safety-net blocks are final on mediated runtimes; on approximated runtimes Clio maps your autonomy level to the closest external posture and records that it did; bypassed runs are explicitly marked in receipts and evidence.
|
|
254
|
+
|
|
255
|
+
---
|
|
256
|
+
|
|
257
|
+
## Approvals
|
|
258
|
+
|
|
259
|
+
An `ask` can come from either axis: a safety-net confirm rail (damage-control `ask` rule, project `requireConfirmation`, `system_modify`) or the autonomy mapping. The permission overlay names the asking axis on its `Asked by:` line, and the transcript carries an `[approval]` notice for every parked call.
|
|
260
|
+
|
|
261
|
+
Every approvable ask has one canonical identity: a `requestId` minted at the approvals plane. The `PermissionRequested` and `PermissionResolved` bus payloads and the audit permission rows all carry it, along with `origin` (who asked), `axis` (which rail or level), and `decidedBy` (who or what answered), so a request joins its resolution on one key across the bus, the ledger, and receipts, and every request resolves exactly once. Worker escalations forward their full decision provenance (reasons, reason code, rule id, policy source), so the overlay names the real asking rail for a worker exactly as it does for the main agent.
|
|
262
|
+
|
|
263
|
+
How an ask resolves depends on the context:
|
|
264
|
+
|
|
265
|
+
### Interactive TUI Behavior
|
|
266
|
+
|
|
267
|
+
In interactive mode, a permission request opens a queued overlay prompt immediately in the TUI.
|
|
268
|
+
- **Queued Overlays:** If multiple tools or worker dispatches require permission during a single turn, the TUI queues the requests. Closing one overlay automatically pops the next permission overlay in the queue.
|
|
269
|
+
- **Operator Options:** The operator can grant permission once, which resumes only the parked tool call without changing the overall operating posture; the one-shot grant is scoped to the presented request's `requestId`. Denying rejects only the presented request and advances the queue; the next parked call re-presents. Cancel-all is reserved for shutdown, an aborted turn, headless runs, and transport failure, where no operator can answer.
|
|
270
|
+
|
|
271
|
+
### Deterministic Headless Behavior
|
|
272
|
+
|
|
273
|
+
When executing tasks in headless mode through `clio-coder run`, there is no terminal operator to prompt.
|
|
274
|
+
- **Deterministic Denials:** Any action that requires permission resolves as a deterministic tool denial. The model receives the rejection and may adapt within the same run.
|
|
275
|
+
- **Rejection Message:** The engine assigns a standard rejection reason to the denied action: `"clio-coder run cannot confirm permission requests; rerun interactively to approve this action."`
|
|
276
|
+
- **Run Exit:** The process exit code reflects the overall run outcome, not the permission denial by itself. A run can still exit 0 if the model completes successfully after receiving the rejected tool result.
|
|
277
|
+
- **Receipts and Audit:** Headless permission denials increment `permissionRequested` receipt/audit counts, so scripts can detect that approval would have been needed without treating every denial as a failed process.
|
|
278
|
+
|
|
279
|
+
### Workers and delegations
|
|
280
|
+
|
|
281
|
+
- **Workers** inherit the session's autonomy level, capped by dispatch scope admission. A worker ask resolves per `workers.onPermission`: `deny` continues the run with a rejection; `fail` ends it; `escalate` forwards it to the interactive operator (see the escalation section above). All three values are editable in the `/settings` center.
|
|
282
|
+
- **Delegations (ACP)** under `clio-policy` governance evaluate through the same net and autonomy mapping; an ask resolves as a non-stall deny so the external agent never hangs waiting for an operator.
|
|
283
|
+
- **ACP server sessions** (a remote client driving Clio) snapshot the autonomy level at `session/new`, so a mid-session settings change on the host cannot alter an in-flight remote session's admission decisions.
|
|
284
|
+
|
|
285
|
+
---
|
|
286
|
+
|
|
287
|
+
## Rigor Gate and Finish Contract
|
|
288
|
+
|
|
289
|
+
In addition to tool-level safety gates, Clio Coder enforces a completion boundary via the finish-contract assessor. This gate is governed by a single `rigor` setting (either `normal` or `high`), which is completely orthogonal to the autonomy permission levels.
|
|
290
|
+
|
|
291
|
+
### Autonomy versus Rigor
|
|
292
|
+
|
|
293
|
+
It is critical to distinguish these two control axes:
|
|
294
|
+
- **Autonomy (Permission)**: Determines what files, directories, tools, or shell commands the agent is permitted to touch, and whether it must ask the operator for permission before executing them.
|
|
295
|
+
- **Rigor (Evidence Bar)**: Determines the verification standard required to accept a task as "done". Autonomy gates what the agent *can do*; rigor gates what the agent must *verify* before it is allowed to finish.
|
|
296
|
+
|
|
297
|
+
| Setting | Axis | Governed By | Handled In |
|
|
298
|
+
| --- | --- | --- | --- |
|
|
299
|
+
| **Autonomy** | Authority | `autonomy` settings dial, `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS` | `src/tools/registry.ts`, `src/domains/safety/` |
|
|
300
|
+
| **Rigor** | Validation | `CLIO_CODER_RIGOR` override, workspace validation contracts | `src/domains/safety/rigor.ts`, `src/domains/safety/finish-contract-registration.ts` |
|
|
301
|
+
|
|
302
|
+
---
|
|
303
|
+
|
|
304
|
+
### Rigor Resolution
|
|
305
|
+
|
|
306
|
+
The effective rigor level for a session or dispatch run is resolved at boot time using the following prioritization:
|
|
307
|
+
|
|
308
|
+
1. **Explicit Override**: Checked via the `CLIO_CODER_RIGOR` environment variable. It is trimmed and parsed case-insensitively. A value of `"high"` or `"normal"` overrides any other setting.
|
|
309
|
+
2. **Repository-Derived Default**: If no override is present, Clio checks the workspace root for the presence of any of the following validation contract files:
|
|
310
|
+
- `.clio-coder/validation.yaml`
|
|
311
|
+
- `.clio-coder/validation.yml`
|
|
312
|
+
- `validation.yaml`
|
|
313
|
+
- `validation.yml`
|
|
314
|
+
- `VALIDATION.md`
|
|
315
|
+
|
|
316
|
+
If any of these files are present, the default rigor level is raised to `high`. Otherwise, the default is `normal`.
|
|
317
|
+
|
|
318
|
+
---
|
|
319
|
+
|
|
320
|
+
### The Finish Gate and Re-Prompt Behavior
|
|
321
|
+
|
|
322
|
+
On every settled `turn_end`, the finish-contract assessor scans entries since the last user message, capped at 80 entries. The trigger is action-scoped: the gate engages only when that window contains successful workspace mutation evidence and no validation evidence or explicit limitation. The model does not have to type a phrase such as `done` or `fixed`; the settled turn after mutation is the completion signal.
|
|
323
|
+
|
|
324
|
+
The assessor decision order is:
|
|
325
|
+
|
|
326
|
+
1. If the window has no successful mutating receipt or settled mutating `!` bash execution, the contract passes with `no_mutation`.
|
|
327
|
+
2. If the window has validation evidence, the contract passes with `validation_evidence`. Evidence includes successful validation commands, `verify` checks (declared verification scripts and the frontend check), passed dispatch receipts, and protected-artifact validation records.
|
|
328
|
+
3. If the assistant explicitly states what could not be verified and why, the contract passes with `explicit_limitation`.
|
|
329
|
+
4. Otherwise, the contract engages with `unvalidated_mutation`.
|
|
330
|
+
|
|
331
|
+
- **Normal Rigor**: Clio issues a soft advisory warning (`FINISH_CONTRACT_ADVISORY_MESSAGE`) injected as a reminder for the next turn, but permits the turn to settle.
|
|
332
|
+
- **High Rigor**: Clio withholds completion. The assessor emits `request_continuation` and a warning `inject_reminder` carrying `HIGH_RIGOR_REVALIDATION_MESSAGE`, instructing the model to run a verification-family command (e.g. `npm test`, `npm run build`) or explicitly declare a limitation before ending.
|
|
333
|
+
|
|
334
|
+
#### Exemptions and Safety Precautions
|
|
335
|
+
- **No-Mutation Turns**: Read-only status, alignment, and inspection turns are exempt because there is no successful workspace mutation in the recent window.
|
|
336
|
+
- **Limitation Claims**: If the model explicitly states what could not be verified and why, the assessor accepts the statement as an explicit limitation and allows the turn to settle cleanly.
|
|
337
|
+
- **Dynamic Injection**: All gate directives are injected dynamically through middleware effects. This ensures that the static system prompt prefix remains byte-stable, preserving prompt caches.
|
|
338
|
+
- **Prior Hard-Block Preservation**: If a prior middleware hook has already emitted a hard block (e.g. tool-prose violation), the high-rigor continuation is suppressed so that critical error guidance is not overwritten.
|
|
339
|
+
|
|
340
|
+
|
|
341
|
+
---
|
|
342
|
+
|
|
343
|
+
## Receipts and evidence
|
|
344
|
+
|
|
345
|
+
Safety decisions feed receipts, audit rows, and evidence artifacts.
|
|
346
|
+
|
|
347
|
+
- **Audit Ledger Durability (`src/domains/safety/audit.ts`)**: Audit entries are appended to `<stateDir>/audit/audit-<date>.jsonl` synchronously, backed by a debounced 5-second `fsyncSync` flush to guarantee on-disk durability without stalling interactive turns. Write errors are logged to stderr without throwing, ensuring auditing never crashes the main execution loop.
|
|
348
|
+
- **Timestamp Ordering**: Audit rows flush asynchronously from concurrent producers and are not strictly time-ordered within a raw `.jsonl` file; consumers must sort rows by `ts` before reasoning about sequence (the evidence builder does this automatically).
|
|
349
|
+
- **Inspection**: The interactive [`/view`](observability.md) surface can inspect and verify receipt artifacts without mutating them. When reporting a problem, include redacted receipts or evidence IDs when possible so maintainers can see:
|
|
350
|
+
|
|
351
|
+
- mode and requested action class;
|
|
352
|
+
- policy source and rule IDs;
|
|
353
|
+
- project policy hash/path;
|
|
354
|
+
- blocked/asked/allowed decision counts;
|
|
355
|
+
- permission request and resolution rows, joinable on `requestId`;
|
|
356
|
+
- the worker `autonomyEnforcement` grade where present;
|
|
357
|
+
- tool statistics and failure messages.
|