@iowarp/clio-coder 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +407 -0
- package/CODE_OF_CONDUCT.md +21 -0
- package/CONTRIBUTING.md +224 -0
- package/LICENSE +202 -0
- package/NOTICE +9 -0
- package/README.md +798 -0
- package/SECURITY.md +72 -0
- package/assets/clio-coder-logo-128.webp +0 -0
- package/damage-control-rules.yaml +419 -0
- package/dist/acp-UMLFVA3F.js +92 -0
- package/dist/agents-Q4MYPMUW.js +91 -0
- package/dist/auth-O6HYIJ6J.js +521 -0
- package/dist/chunk-262G75JS.js +35 -0
- package/dist/chunk-26BZQOAD.js +1281 -0
- package/dist/chunk-2J63S4SF.js +508 -0
- package/dist/chunk-3DANZDGR.js +717 -0
- package/dist/chunk-4UQA7NCT.js +29 -0
- package/dist/chunk-527KG6XR.js +497 -0
- package/dist/chunk-5LDRNKX2.js +1063 -0
- package/dist/chunk-5N2FG33Q.js +25 -0
- package/dist/chunk-67MTHP2E.js +135 -0
- package/dist/chunk-6CWDTGUC.js +20 -0
- package/dist/chunk-7BHLZB3A.js +2115 -0
- package/dist/chunk-7RBKDI66.js +348 -0
- package/dist/chunk-AMFR5YA3.js +541 -0
- package/dist/chunk-BBUH4VAA.js +1224 -0
- package/dist/chunk-BYEU76JP.js +899 -0
- package/dist/chunk-CLJ5HLUD.js +458 -0
- package/dist/chunk-D5YD55AR.js +116 -0
- package/dist/chunk-DXQNI4PC.js +61 -0
- package/dist/chunk-E3NYWENM.js +1004 -0
- package/dist/chunk-GNGDQYDU.js +34688 -0
- package/dist/chunk-GOTUR54M.js +9 -0
- package/dist/chunk-HBU5MTAM.js +41 -0
- package/dist/chunk-HMYNFFY4.js +28 -0
- package/dist/chunk-JPOWPFCU.js +1010 -0
- package/dist/chunk-JWHCJDCI.js +1215 -0
- package/dist/chunk-KBR4MZZR.js +41 -0
- package/dist/chunk-KKKPTZLM.js +93 -0
- package/dist/chunk-ME6DNWIU.js +66 -0
- package/dist/chunk-NI4DEJMC.js +88 -0
- package/dist/chunk-O4EJEDHO.js +659 -0
- package/dist/chunk-PIDUD6M2.js +31 -0
- package/dist/chunk-PS4PFJQP.js +29459 -0
- package/dist/chunk-QV47YRF4.js +48 -0
- package/dist/chunk-RQDWMVRB.js +279 -0
- package/dist/chunk-TFSSEXL6.js +136 -0
- package/dist/chunk-TKHQ4DGZ.js +8290 -0
- package/dist/chunk-TPOCL34A.js +2876 -0
- package/dist/chunk-UGYAX5YI.js +565 -0
- package/dist/chunk-UHTSULZS.js +461 -0
- package/dist/chunk-UU3R62TT.js +128 -0
- package/dist/chunk-UWIJNAOB.js +3906 -0
- package/dist/chunk-VOO7NYPP.js +914 -0
- package/dist/chunk-VPAWTYLY.js +117 -0
- package/dist/chunk-WD6AJM35.js +1216 -0
- package/dist/chunk-X3BR7HWV.js +115 -0
- package/dist/chunk-X3NE4WVW.js +120 -0
- package/dist/chunk-XNISANGE.js +1395 -0
- package/dist/chunk-XV4ZJ6ZM.js +3177 -0
- package/dist/cli/index.js +236 -0
- package/dist/clio-KIQ5SNDS.js +53 -0
- package/dist/components-JVHMUBEB.js +653 -0
- package/dist/config-ZFCDBMDC.js +372 -0
- package/dist/configure-G4E3A2PG.js +27 -0
- package/dist/context-CDXTP2MP.js +293 -0
- package/dist/context-E3KIFVXI.js +185 -0
- package/dist/context-clear-3F4PLXOS.js +102 -0
- package/dist/context-index-Q7YSYTR3.js +106 -0
- package/dist/docs-YIETIWZI.js +280 -0
- package/dist/doctor-M5HJJZOL.js +61 -0
- package/dist/domains/agents/builtins/architect.md +33 -0
- package/dist/domains/agents/builtins/coder.md +31 -0
- package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
- package/dist/domains/agents/builtins/debugger.md +30 -0
- package/dist/domains/agents/builtins/documenter.md +31 -0
- package/dist/domains/agents/builtins/git-master.md +30 -0
- package/dist/domains/agents/builtins/provenance.md +30 -0
- package/dist/domains/agents/builtins/researcher.md +71 -0
- package/dist/domains/agents/builtins/scout.md +42 -0
- package/dist/domains/agents/builtins/tester.md +31 -0
- package/dist/domains/agents/builtins/verifier.md +30 -0
- package/dist/domains/agents/builtins/wiki-writer.md +41 -0
- package/dist/eval-B3KZZESM.js +2674 -0
- package/dist/evidence-V67CHM35.js +233 -0
- package/dist/evolve-YDZSUQYA.js +518 -0
- package/dist/extensions-SRG7XCAH.js +207 -0
- package/dist/fleet-CA2CRTVG.js +760 -0
- package/dist/fleet-preflight-CLIAX7YR.js +21 -0
- package/dist/init-2OZDJE2D.js +227 -0
- package/dist/memory-3PIQQAKX.js +207 -0
- package/dist/models-DY35XI7Y.js +237 -0
- package/dist/paths-5OMXW7Z4.js +57 -0
- package/dist/preload-KZVHET2B.js +11 -0
- package/dist/reset-PIFYNOS3.js +216 -0
- package/dist/run-3VSPP24F.js +735 -0
- package/dist/share-D36RQCXM.js +241 -0
- package/dist/skills-F2MRLELY.js +445 -0
- package/dist/skills-eval-E2ZTW4PL.js +932 -0
- package/dist/targets-DZMEZAH4.js +977 -0
- package/dist/trace-7NYCUI2J.js +250 -0
- package/dist/uninstall-AD3JWHBB.js +322 -0
- package/dist/upgrade-WYYBKGDY.js +301 -0
- package/dist/usage-ULIDAGFF.js +755 -0
- package/dist/version-ROZ6CZKH.js +16 -0
- package/dist/wiki-generate-PKFIX6OB.js +377 -0
- package/dist/worker/entry.js +1739 -0
- package/docs/README.md +93 -0
- package/docs/acp.md +120 -0
- package/docs/alcf-provider.md +72 -0
- package/docs/architecture.md +172 -0
- package/docs/artifact-versions.md +54 -0
- package/docs/built-in-agents.md +265 -0
- package/docs/capacity-and-scheduling.md +97 -0
- package/docs/commands-and-modes.md +554 -0
- package/docs/config-knobs-audit.md +115 -0
- package/docs/configuration-and-targets.md +812 -0
- package/docs/context-engine.md +236 -0
- package/docs/dispatch-architecture-rationale.md +126 -0
- package/docs/documentation-coverage.md +46 -0
- package/docs/documentation-guide.md +166 -0
- package/docs/environment-variables.md +105 -0
- package/docs/eval-runner.md +205 -0
- package/docs/evals-internal.md +298 -0
- package/docs/evidence-and-memory.md +243 -0
- package/docs/evolution.md +143 -0
- package/docs/exit-codes-and-output.md +74 -0
- package/docs/extensions-and-sharing.md +306 -0
- package/docs/fleet-demo-runbook.md +179 -0
- package/docs/fleet-dispatch.md +591 -0
- package/docs/glossary.md +75 -0
- package/docs/html/agents_blueprint.html +936 -0
- package/docs/html/alcf_blueprint.html +324 -0
- package/docs/html/architecture_blueprint.html +850 -0
- package/docs/html/commands_blueprint.html +794 -0
- package/docs/html/config_knobs_audit_blueprint.html +178 -0
- package/docs/html/configuration_blueprint.html +1080 -0
- package/docs/html/context_blueprint.html +603 -0
- package/docs/html/documentation_blueprint.html +832 -0
- package/docs/html/environment_blueprint.html +404 -0
- package/docs/html/eval_blueprint.html +743 -0
- package/docs/html/evals_internal_blueprint.html +190 -0
- package/docs/html/evolution_blueprint.html +674 -0
- package/docs/html/extensions_blueprint.html +2065 -0
- package/docs/html/fleet_dispatch_blueprint.html +286 -0
- package/docs/html/index.html +919 -0
- package/docs/html/lifecycle_blueprint.html +723 -0
- package/docs/html/memory_blueprint.html +699 -0
- package/docs/html/middleware_blueprint.html +664 -0
- package/docs/html/models_blueprint.html +2366 -0
- package/docs/html/observability_blueprint.html +683 -0
- package/docs/html/provider_adapter_blueprint.html +245 -0
- package/docs/html/safety_blueprint.html +1386 -0
- package/docs/html/shared.css +571 -0
- package/docs/html/shared.js +143 -0
- package/docs/html/skills_blueprint.html +671 -0
- package/docs/html/soak_blueprint.html +182 -0
- package/docs/html/tool_usage_blueprint.html +350 -0
- package/docs/html/tools_blueprint.html +2249 -0
- package/docs/html/trace_blueprint.html +235 -0
- package/docs/html/tui_design_blueprint.html +314 -0
- package/docs/html/validation_blueprint.html +961 -0
- package/docs/html/worker_dispatch_blueprint.html +231 -0
- package/docs/installation-and-lifecycle.md +308 -0
- package/docs/middleware-and-components.md +148 -0
- package/docs/model-catalog.md +189 -0
- package/docs/observability.md +233 -0
- package/docs/proactive-memory.md +452 -0
- package/docs/prompt-envelope-and-tools.md +142 -0
- package/docs/provider-adapter-cookbook.md +148 -0
- package/docs/release-cut-checklist.md +138 -0
- package/docs/safety-model.md +357 -0
- package/docs/scientific-validation.md +105 -0
- package/docs/session-lifecycle.md +156 -0
- package/docs/skills-marketplace.md +46 -0
- package/docs/tool-usage.md +527 -0
- package/docs/trace-store.md +132 -0
- package/docs/troubleshooting.md +33 -0
- package/docs/tui-design.md +239 -0
- package/docs/worker-dispatch-mechanics.md +242 -0
- package/package.json +132 -0
- package/skills/README.md +408 -0
- package/skills/git/commit-crafting/SKILL.md +79 -0
- package/skills/git/commit-crafting/evals.md +92 -0
- package/skills/git/create-pr/SKILL.md +116 -0
- package/skills/git/create-pr/evals.md +114 -0
- package/skills/git/investigate-issue/SKILL.md +139 -0
- package/skills/git/investigate-issue/evals.md +94 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
- package/skills/git/resolve-merge-conflicts/evals.md +58 -0
- package/skills/git/review-changes/SKILL.md +103 -0
- package/skills/git/review-changes/evals.md +85 -0
- package/skills/git/worktree-create/SKILL.md +92 -0
- package/skills/git/worktree-create/evals.md +97 -0
- package/skills/git/worktree-create/references/worktree-setup.md +66 -0
- package/skills/git/worktree-merge/SKILL.md +95 -0
- package/skills/git/worktree-merge/evals.md +114 -0
- package/skills/skill-marketplace.json +261 -0
- package/skills/workflow/cut-it/SKILL.md +86 -0
- package/skills/workflow/cut-it/evals.md +42 -0
- package/src/domains/agents/builtins/architect.md +33 -0
- package/src/domains/agents/builtins/coder.md +31 -0
- package/src/domains/agents/builtins/context-bootstrap.md +38 -0
- package/src/domains/agents/builtins/debugger.md +30 -0
- package/src/domains/agents/builtins/documenter.md +31 -0
- package/src/domains/agents/builtins/git-master.md +30 -0
- package/src/domains/agents/builtins/provenance.md +30 -0
- package/src/domains/agents/builtins/researcher.md +71 -0
- package/src/domains/agents/builtins/scout.md +42 -0
- package/src/domains/agents/builtins/tester.md +31 -0
- package/src/domains/agents/builtins/verifier.md +30 -0
- package/src/domains/agents/builtins/wiki-writer.md +41 -0
- package/src/domains/agents/fleets/build-review.md +34 -0
- package/src/domains/agents/fleets/build-test.md +35 -0
- package/src/domains/agents/fleets/sdlc.md +86 -0
- package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
- package/src/domains/prompts/fragments/identity/clio.md +26 -0
- package/src/domains/prompts/fragments/operating/contract.md +64 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
- package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
- package/src/domains/prompts/fragments/safety/read-only.md +13 -0
- package/src/domains/prompts/fragments/safety/suggest.md +13 -0
- package/src/domains/prompts/fragments/wiki/page.md +75 -0
- package/src/domains/prompts/fragments/wiki/plan.md +48 -0
- package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
|
@@ -0,0 +1,243 @@
|
|
|
1
|
+
# Evidence Corpus and Long-Term Memory
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.0). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
|
|
5
|
+
|
|
6
|
+
Clio Coder treats run claims and agent lessons as structured artifacts to support reproducibility and scientific provenance. In evaluations such as [SWE-bench](https://www.swebench.com), capturing granular execution evidence is essential for validating agent claims. Evidence corpora are deterministic directories built from run ledgers, receipts, sessions, audits, and eval artifacts. In v0.3.0, forensic evidence auto-builds on dispatch run completion: when a run finalizes, the observability domain automatically compiles the evidence bundle under `<dataDir>/evidence/run-<id>/` and updates a compact sidecar index row in `<stateDir>/evidence-index.json`. Long-term memory records are local, evidence-linked, and only injected after explicit approval. Use the TUI [`/view`](observability.md) command for interactive inspection of receipts, dispatch output, durable tool output, compaction summaries, and session accountability before building or citing evidence.
|
|
7
|
+
|
|
8
|
+
Source of truth: `src/domains/evidence/**`, `src/domains/memory/**`, `src/cli/evidence.ts`, and `src/cli/memory.ts`.
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## Evidence CLI
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
clio-coder evidence build --run <runId>
|
|
16
|
+
clio-coder evidence build --session <sessionId>
|
|
17
|
+
clio-coder evidence build --eval <evalId>
|
|
18
|
+
clio-coder evidence inspect <evidenceId>
|
|
19
|
+
clio-coder evidence list
|
|
20
|
+
```
|
|
21
|
+
|
|
22
|
+
`clio-coder evidence inspect <id>` requires a valid evidence artifact ID. If the requested artifact does not exist on disk, it outputs `error: evidence artifact not found: <id> (see clio-coder evidence list)` and exits with code 1.
|
|
23
|
+
|
|
24
|
+
Evidence IDs are deterministic:
|
|
25
|
+
|
|
26
|
+
| Source | ID shape |
|
|
27
|
+
| --- | --- |
|
|
28
|
+
| Run | `run-<runId>` |
|
|
29
|
+
| Session | `session-<sessionId>` |
|
|
30
|
+
| Eval | `eval-<evalId>` |
|
|
31
|
+
|
|
32
|
+
Rebuilding the same evidence ID rewrites the same directory under `<dataDir>/evidence/`.
|
|
33
|
+
|
|
34
|
+
|
|
35
|
+
---
|
|
36
|
+
|
|
37
|
+
## Evidence directory layout
|
|
38
|
+
|
|
39
|
+
Run/session evidence files:
|
|
40
|
+
|
|
41
|
+
```text
|
|
42
|
+
<dataDir>/evidence/<evidenceId>/
|
|
43
|
+
├── overview.json
|
|
44
|
+
├── transcript.md
|
|
45
|
+
├── trace.raw.jsonl
|
|
46
|
+
├── trace.cleaned.jsonl
|
|
47
|
+
├── tool-events.jsonl
|
|
48
|
+
├── audit-linked.jsonl
|
|
49
|
+
├── receipt.json
|
|
50
|
+
├── gate-decisions.json
|
|
51
|
+
├── protected-artifacts.json
|
|
52
|
+
├── findings.json
|
|
53
|
+
└── findings.md
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
Eval evidence adds `eval-result.json` and uses empty receipt/protected-artifact placeholders when no linked receipts exist.
|
|
57
|
+
|
|
58
|
+
### Core files
|
|
59
|
+
|
|
60
|
+
| File | Purpose |
|
|
61
|
+
| --- | --- |
|
|
62
|
+
| `overview.json` | Stable summary: source, runs, sessions, statuses, tasks, models, totals, tags, and file list. |
|
|
63
|
+
| `transcript.md` | Human-readable run/session/eval transcript. |
|
|
64
|
+
| `trace.raw.jsonl` | Raw run ledger/receipt/eval rows. |
|
|
65
|
+
| `trace.cleaned.jsonl` | Compact normalized rows plus findings. |
|
|
66
|
+
| `tool-events.jsonl` | Tool summaries from session entries, audit rows, receipts, or eval commands. |
|
|
67
|
+
| `audit-linked.jsonl` | Audit rows linked to run/session context when available. |
|
|
68
|
+
| `receipt.json` | Receipt bundle (`{ version: 1, receipts: [...] }`); only receipts that pass integrity verification contribute verified fields. |
|
|
69
|
+
| `gate-decisions.json` | Integrity-verified review verdicts, compete winner selections, and winner confirmations discovered from linked receipt ids. |
|
|
70
|
+
| `protected-artifacts.json` | Protected artifact state/events. |
|
|
71
|
+
| `findings.json` / `findings.md` | Structured and readable findings. |
|
|
72
|
+
|
|
73
|
+
When a run was chained (pipeline), composed with a persona override, or escalated for a permission, `transcript.md` and `trace.cleaned.jsonl` surface the receipt's provenance field sets, and `clio-coder evidence inspect` prints them as a `provenance <runId>:` block. The field paths, types, and stability labels are documented in the [receipt provenance schema](./observability.md#receipt-fields-for-dispatch-provenance).
|
|
74
|
+
|
|
75
|
+
---
|
|
76
|
+
|
|
77
|
+
## Evidence Tag Taxonomy and Failure Causes
|
|
78
|
+
|
|
79
|
+
Clio Coder classifies every run, session, and eval record using a closed set of 25 canonical tags. These tags distinguish general execution characteristics (such as lineage linkages) from actual failure causes.
|
|
80
|
+
|
|
81
|
+
### Complete Taxonomy
|
|
82
|
+
|
|
83
|
+
| Tag | Category | Trigger / Meaning |
|
|
84
|
+
| --- | --- | --- |
|
|
85
|
+
| `audit-linked` | Provenance | Audit logs successfully linked to this run or session. |
|
|
86
|
+
| `audit-missing` | Provenance | No matching audit logs were found. |
|
|
87
|
+
| `best-effort-link` | Provenance | Inspection commands or logs linked via heuristics. |
|
|
88
|
+
| `timeout` | Failure | Execution exceeded the maximum duration limit. |
|
|
89
|
+
| `context-overflow` | Constraint | Model context limit was exceeded. |
|
|
90
|
+
| `provider-transient` | Transient | Temporary API or model gateway connection error. |
|
|
91
|
+
| `missing-dependency`| Failure | Python, Node, or system package dependency was missing. |
|
|
92
|
+
| `wrong-runtime` | Configuration | Execution failed due to incorrect compiler or runtime environment. |
|
|
93
|
+
| `proxy-validation` | Validation | Weak validation (e.g. only file-presence check rather than execution). |
|
|
94
|
+
| `no-validation` | Validation | Succeeded turn or run did not execute any verification commands. |
|
|
95
|
+
| `destructive-cleanup`| Precaution | Clean-up rules triggered to prevent workspace pollution or damage. |
|
|
96
|
+
| `blocked-tool` | Failure | The safety net blocked a tool call requested by the model. |
|
|
97
|
+
| `escalation` | Precaution | A worker permission escalation timed out or was denied; see the receipt provenance schema below. |
|
|
98
|
+
| `receipt-integrity` | Security | Forensic verification detected receipt modification or checksum mismatch. |
|
|
99
|
+
| `protected-artifact`| Precaution | Mutating a path protected by project or system safety policies. |
|
|
100
|
+
| `tool-loop` | Constraint | The model repeatedly called the same tool with identical arguments. |
|
|
101
|
+
| `test-failure` | Failure | A verification command containing test/lint keywords exited non-zero. |
|
|
102
|
+
| `build-failure` | Failure | A verification command containing build keywords exited non-zero. |
|
|
103
|
+
| `cwd-missing` | Configuration | The directory target specified for execution did not exist. |
|
|
104
|
+
| `session-linked` | Provenance | The run is linked back to its originating parent session. |
|
|
105
|
+
| `session-missing` | Provenance | No parent session could be resolved for this run. |
|
|
106
|
+
| `auth-failure` | Failure | Missing or invalid credentials/API keys. |
|
|
107
|
+
| `external-bypass` | Security | An external runner bypassed standard safety gates. |
|
|
108
|
+
| `external-approximation`| Validation | An external runner approximated results rather than fully executing. |
|
|
109
|
+
| `unknown` | Undefined | Unclassified execution failure. |
|
|
110
|
+
|
|
111
|
+
---
|
|
112
|
+
|
|
113
|
+
### Failure-Cause Tag Subset
|
|
114
|
+
|
|
115
|
+
A subset of the taxonomy represents actual failure causes (governed by the `FAILURE_CAUSE_TAG_ORDER` array). These are the only tags included in the receipt summaries and TUI observability histograms:
|
|
116
|
+
|
|
117
|
+
1. **`timeout`**: Triggered if the run outcome is `"timed_out"` or `"stalled"`, or if the error/failure text contains `"timed out"` or `"timeout"`.
|
|
118
|
+
2. **`auth-failure`**: Triggered if failure text contains keywords like `"auth"`, `"api key"`, `"credential"`, or `"unauthorized"`.
|
|
119
|
+
3. **`missing-dependency`**: Triggered if failure logs contain `"module not found"`, `"missing package"`, or `"missing dependency"`.
|
|
120
|
+
4. **`build-failure`**: Triggered in receipt summaries when a non-zero receipt exit is paired with build tool names in `toolStats` (e.g. `build`, `compile`, `make`, `cmake`, `cargo`, `gradle`, `ninja`, `tsc`). Forensic evidence can also classify a non-zero run from build language in the recorded task text.
|
|
121
|
+
5. **`test-failure`**: Triggered in receipt summaries when a non-zero receipt exit is paired with test or lint tool names in `toolStats` (e.g. `pytest`, `ctest`, `jest`, `vitest`, `test`, `lint`, `typecheck`). Forensic evidence can also classify a non-zero run from validation language in the recorded task text.
|
|
122
|
+
6. **`blocked-tool`**: Triggered if tool execution statistics show a blocked count greater than `0`.
|
|
123
|
+
|
|
124
|
+
---
|
|
125
|
+
|
|
126
|
+
## Receipt findingsSummary
|
|
127
|
+
|
|
128
|
+
Each run receipt (persisted under `<stateDir>/receipts/<runId>.json`) carries an optional `findingsSummary` block. This block provides a cheap, integrity-covered summary of the run's findings:
|
|
129
|
+
|
|
130
|
+
```json
|
|
131
|
+
"findingsSummary": {
|
|
132
|
+
"tags": ["test-failure"],
|
|
133
|
+
"firstPassSuccess": false,
|
|
134
|
+
"findingCount": 1
|
|
135
|
+
}
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
### Computation and Lifecycle
|
|
139
|
+
- **Circular Dependency Prevention**: To prevent circular dependencies, `findingsSummary` is calculated **cheaply in-memory** at receipt-record time using the draft envelope and tool statistics (in `src/domains/dispatch/receipt-findings.ts`). It never reads from disk or calls `buildEvidence`.
|
|
140
|
+
- **First-Pass Success**: Calculated as `true` only if the terminal outcome was `"succeeded"`, the lineage attempt was `0` (no dispatch retries), the tool stats confirm at least one successful validation tool was executed, and no failure-cause tags were detected.
|
|
141
|
+
- **Cryptographic Coverage**: Current receipts use strict v15 and authenticate every current receipt field, including briefing and steering provenance, routing intent and decision, route quality, worker identity, execution role, and result-contract conformance, against the reconstructed ledger. Every version other than v15 is rejected; there is no historical receipt reader.
|
|
142
|
+
|
|
143
|
+
| Version | Verification policy | Compatibility policy |
|
|
144
|
+
|---|---|---|
|
|
145
|
+
| v15 | Current canonical projection; every current receipt and reconstructible ledger field is authenticated | Accepted |
|
|
146
|
+
| Any other version | No reader | Rejected; remove or archive the incompatible state rather than expecting migration |
|
|
147
|
+
|
|
148
|
+
Receipt integrity and evidence verification answer different questions. The
|
|
149
|
+
former proves that a receipt matches its ledger envelope; the latter records
|
|
150
|
+
whether applicable validation evidence was observed. Briefing provenance is
|
|
151
|
+
also distinct from bounded project-context provenance: both can be absent or
|
|
152
|
+
present independently, and neither hash is evidence for the other.
|
|
153
|
+
|
|
154
|
+
### Mutation-Report Grounding
|
|
155
|
+
|
|
156
|
+
Mutation-report receipts are grounded directly against observed tool events recorded in the run ledger:
|
|
157
|
+
- When a worker run concludes, claimed modified files are validated against the actual write set observed from `edit` and `write` tool invocations.
|
|
158
|
+
- If a target file was untouched during the run but already existed on disk, the result contract seals route quality as `unmeasured` rather than `fail`.
|
|
159
|
+
- If a write attempt was refused or denied, mutation validation marks the outcome as `unmeasured`.
|
|
160
|
+
- If an unattempted mutation target does not exist on disk, postcondition validation fails.
|
|
161
|
+
|
|
162
|
+
---
|
|
163
|
+
|
|
164
|
+
|
|
165
|
+
## Memory CLI
|
|
166
|
+
|
|
167
|
+
```bash
|
|
168
|
+
clio-coder memory list
|
|
169
|
+
clio-coder memory propose --from-evidence <evidenceId>
|
|
170
|
+
clio-coder memory approve <memoryId>
|
|
171
|
+
clio-coder memory reject <memoryId>
|
|
172
|
+
clio-coder memory prune --stale
|
|
173
|
+
```
|
|
174
|
+
|
|
175
|
+
Memory records live in:
|
|
176
|
+
|
|
177
|
+
```text
|
|
178
|
+
<dataDir>/memory/records.json
|
|
179
|
+
```
|
|
180
|
+
|
|
181
|
+
The store is capped at `500` records and is sorted by scope, key, creation time, and id for stable writes.
|
|
182
|
+
|
|
183
|
+
---
|
|
184
|
+
|
|
185
|
+
## Memory record lifecycle
|
|
186
|
+
|
|
187
|
+
```mermaid
|
|
188
|
+
stateDiagram-v2
|
|
189
|
+
evidence --> proposed: propose --from-evidence
|
|
190
|
+
proposed --> approved: approve <id>
|
|
191
|
+
proposed --> rejected: reject <id>
|
|
192
|
+
approved --> rejected: reject <id>
|
|
193
|
+
proposed --> pruned: prune --stale after 30 days
|
|
194
|
+
rejected --> pruned: prune --stale after 30 days
|
|
195
|
+
approved --> pruned: prune --stale after 180 days since lastVerifiedAt/createdAt
|
|
196
|
+
```
|
|
197
|
+
|
|
198
|
+
Records must cite at least one evidence ID to be considered for prompt injection. Rejected records remain in the store until stale pruning so the same bad lesson is not immediately re-proposed from the same evidence.
|
|
199
|
+
|
|
200
|
+
---
|
|
201
|
+
|
|
202
|
+
## Prompt injection rules
|
|
203
|
+
|
|
204
|
+
The chat loop loads memory synchronously from the bounded local store and calls `buildMemoryPromptSection()`.
|
|
205
|
+
|
|
206
|
+
Defaults:
|
|
207
|
+
|
|
208
|
+
| Constraint | Default |
|
|
209
|
+
| --- | --- |
|
|
210
|
+
| Scopes | `global`, `repo` |
|
|
211
|
+
| Token budget | `400` estimated tokens |
|
|
212
|
+
| Max records | `5` |
|
|
213
|
+
| Required status | `approved: true` |
|
|
214
|
+
| Required provenance | At least one `evidenceRefs[]` entry |
|
|
215
|
+
| Suppression | Records with active `regressions[]` entries are skipped |
|
|
216
|
+
|
|
217
|
+
Rendered memory lines always cite record ID, scope, lesson, and evidence IDs. The prompt tells the model not to extrapolate beyond cited findings.
|
|
218
|
+
|
|
219
|
+
### Repository-scoped identity
|
|
220
|
+
|
|
221
|
+
Repository memory is selected by an exact canonical absolute-path identity. The interactive orchestrator and `clio-coder run --agent` compute that identity from the active working directory; symlink aliases collapse to the same key. A repository move, a different Git worktree path, a subdirectory launch, a malformed identity, or a missing identity does not inherit another repository's memory. Global records are unaffected.
|
|
222
|
+
|
|
223
|
+
Every `scope: "repo"` record must carry:
|
|
224
|
+
|
|
225
|
+
```json
|
|
226
|
+
"repository": { "kind": "canonical-path", "key": "/canonical/absolute/repository/path" }
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
The structured `repository` field is the only applicability mechanism: store validation rejects repo records without it, and `appliesWhen` tokens never grant repository applicability. There is intentionally no automatic path rewrite for moved repositories or worktrees: a filesystem move produces a different identity and the record simply stops applying until it is re-scoped with new evidence.
|
|
230
|
+
|
|
231
|
+
---
|
|
232
|
+
|
|
233
|
+
## Recommended workflow
|
|
234
|
+
|
|
235
|
+
1. Build evidence from the run/session/eval that taught the lesson.
|
|
236
|
+
2. Inspect the evidence and findings.
|
|
237
|
+
3. Propose memory from the evidence.
|
|
238
|
+
4. Review the proposed lesson for correctness and scope.
|
|
239
|
+
5. Approve only if it is durable and useful.
|
|
240
|
+
6. Reject incorrect or overbroad records.
|
|
241
|
+
7. Prune stale records periodically.
|
|
242
|
+
|
|
243
|
+
Memory is meant to reduce repeated mistakes, not to become an unreviewed second instruction system.
|
|
@@ -0,0 +1,143 @@
|
|
|
1
|
+
# Evolution and Change Manifests
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
|
|
7
|
+
|
|
8
|
+
Source of truth: `src/domains/evolution/manifest.ts`, `src/domains/evolution/validate.ts`, and `src/cli/evolve.ts`.
|
|
9
|
+
|
|
10
|
+
---
|
|
11
|
+
|
|
12
|
+
## CLI
|
|
13
|
+
|
|
14
|
+
```bash
|
|
15
|
+
clio-coder evolve manifest init > change-manifest.json
|
|
16
|
+
clio-coder evolve manifest validate change-manifest.json
|
|
17
|
+
clio-coder evolve manifest summarize change-manifest.json
|
|
18
|
+
```
|
|
19
|
+
|
|
20
|
+
`init` prints a template. `validate` exits non-zero and reports JSON paths when fields are missing or invalid. In `v0.2.9`, it additionally resolves each non-empty `evidenceRefs` entry (which must match `run-<id>` or `session-<id>` formatting) against the local evidence store, failing validation if a referenced bundle is not found. `summarize` prints the iteration id, base SHA, authority levels, components, files, predicted regressions, validation-step count, and reports evidence refs resolution.
|
|
21
|
+
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
## Minimal manifest
|
|
25
|
+
|
|
26
|
+
```json
|
|
27
|
+
{
|
|
28
|
+
"version": 1,
|
|
29
|
+
"iterationId": "exploratory-1",
|
|
30
|
+
"baseGitSha": "0000000000000000000000000000000000000000",
|
|
31
|
+
"createdAt": "2026-04-29T00:00:00.000Z",
|
|
32
|
+
"changes": [
|
|
33
|
+
{
|
|
34
|
+
"id": "change-1",
|
|
35
|
+
"componentIds": ["context-file:CLIO-CODER.md"],
|
|
36
|
+
"filesChanged": ["CLIO-CODER.md"],
|
|
37
|
+
"authorityLevel": "prompt",
|
|
38
|
+
"evidenceRefs": [],
|
|
39
|
+
"rootCause": "First exploratory iteration; no evidence corpus exists yet.",
|
|
40
|
+
"targetedFix": "Describe the smallest proposed harness change.",
|
|
41
|
+
"predictedFixes": ["One expected improvement."],
|
|
42
|
+
"predictedRegressions": [],
|
|
43
|
+
"validationPlan": ["npm run test"],
|
|
44
|
+
"rollbackPlan": "Revert the filesChanged entries for this change.",
|
|
45
|
+
"expectedBudgetImpact": {
|
|
46
|
+
"risk": "same"
|
|
47
|
+
}
|
|
48
|
+
}
|
|
49
|
+
]
|
|
50
|
+
}
|
|
51
|
+
```
|
|
52
|
+
|
|
53
|
+
Only the first exploratory iteration (`iterationId: "exploratory-1"`) is permitted to use an empty `evidenceRefs` array. For all other iterations, the manifest must cite valid evidence of verification.
|
|
54
|
+
|
|
55
|
+
### Evidence-Linked Validation (Slice 5a)
|
|
56
|
+
|
|
57
|
+
During `clio-coder evolve manifest validate` and `summarize` commands, Clio Coder validates the referenced evidence bundles:
|
|
58
|
+
- **Format Verification**: Every reference in the `evidenceRefs` array must follow the format `run-<id>` or `session-<id>`.
|
|
59
|
+
- **Durable Store Resolution**: Each reference must correspond to a folder that actually exists under `<dataDir>/evidence/`. If any referenced bundle is missing, validation fails and reports a dangling reference issue.
|
|
60
|
+
- **Engine Boundaries**: To maintain domain boundaries (`check:boundaries`), the validation function `validateChangeManifest` is completely decoupled. It accepts a `resolveEvidenceRef` predicate option. The CLI passes a resolver connected to the evidence store, keeping the evolution domain from directly importing the evidence domain.
|
|
61
|
+
|
|
62
|
+
### Self-Edit Gate Deferral (Slice 5b)
|
|
63
|
+
|
|
64
|
+
Enforcement of high-authority self-edits (Slice 5b) is deferred per the status and design in `src/domains/evolution/SELF_EDIT_GATE.md`.
|
|
65
|
+
|
|
66
|
+
#### Why it is Deferred
|
|
67
|
+
To enforce that Clio Coder cannot modify its own codebase (specifically paths mapped to `HIGH_AUTHORITY_LEVELS` such as `src/engine/**` or `src/domains/safety/**`) without a validated change manifest, the harness must distinguish between:
|
|
68
|
+
1. **Agent-Initiated Self-Edits**: Clio editing its own harness paths autonomously.
|
|
69
|
+
2. **Operator-Requested Edits**: The developer instructing Clio to modify files in the repository.
|
|
70
|
+
|
|
71
|
+
Currently, the middleware lacks a reliable signal to separate these two scenarios. Implementing a naive path-glob filter would block developers from making ordinary source changes, which is forbidden by the release specification.
|
|
72
|
+
|
|
73
|
+
#### Concrete Gate Design and Re-entry
|
|
74
|
+
When the required signal (`selfEditOrigin: "operator-requested" | "agent-initiated"`) is implemented, Slice 5b will run as a `before_tool` middleware registration. For any high-authority tool call classified as `agent-initiated`, it will require a change manifest whose `evidenceRefs` validate against the local evidence store, returning a `block_tool` effect if missing.
|
|
75
|
+
|
|
76
|
+
|
|
77
|
+
---
|
|
78
|
+
|
|
79
|
+
## Schema notes
|
|
80
|
+
|
|
81
|
+
| Field | Type | Requirement |
|
|
82
|
+
| --- | --- | --- |
|
|
83
|
+
| `version` | literal `1` | Required. |
|
|
84
|
+
| `iterationId` | non-empty string | Required. `exploratory-1` has special empty-evidence handling. |
|
|
85
|
+
| `baseGitSha` | non-empty string | Required. Use the commit the change was based on. |
|
|
86
|
+
| `createdAt` | non-empty string | Required; ISO timestamp recommended. |
|
|
87
|
+
| `changes` | array | Required. |
|
|
88
|
+
| `changes[].id` | non-empty string | Required. |
|
|
89
|
+
| `changes[].componentIds` | string array | Required; can be empty only if `filesChanged` is non-empty. |
|
|
90
|
+
| `changes[].filesChanged` | string array | Required; can be empty only if `componentIds` is non-empty. |
|
|
91
|
+
| `changes[].authorityLevel` | enum | Required; see below. |
|
|
92
|
+
| `changes[].evidenceRefs` | string array | Required; empty only for `exploratory-1`. |
|
|
93
|
+
| `changes[].rootCause` | non-empty string | Required. |
|
|
94
|
+
| `changes[].targetedFix` | non-empty string | Required. |
|
|
95
|
+
| `changes[].predictedFixes` | string array | Required. |
|
|
96
|
+
| `changes[].predictedRegressions` | string array | Required; high-authority changes require at least one entry. |
|
|
97
|
+
| `changes[].validationPlan` | string array | Required. |
|
|
98
|
+
| `changes[].rollbackPlan` | non-empty string | Required. |
|
|
99
|
+
| `changes[].expectedBudgetImpact` | object | Optional; when present, `risk` is required. |
|
|
100
|
+
|
|
101
|
+
Authority levels:
|
|
102
|
+
|
|
103
|
+
```text
|
|
104
|
+
prompt | tool-description | tool-implementation | middleware | memory | runtime | safety | schema | cli
|
|
105
|
+
```
|
|
106
|
+
|
|
107
|
+
High-authority levels are:
|
|
108
|
+
|
|
109
|
+
```text
|
|
110
|
+
tool-implementation | middleware | runtime | safety | schema | cli
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
High-authority changes must list predicted regressions. This is deliberate: the point of the manifest is not to promise safety, but to make risk reviewable.
|
|
114
|
+
|
|
115
|
+
Budget impact risk values:
|
|
116
|
+
|
|
117
|
+
```text
|
|
118
|
+
lower | same | higher
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Optional budget deltas:
|
|
122
|
+
|
|
123
|
+
```json
|
|
124
|
+
"expectedBudgetImpact": {
|
|
125
|
+
"tokenDelta": 500,
|
|
126
|
+
"wallTimeDeltaMs": 2000,
|
|
127
|
+
"risk": "higher"
|
|
128
|
+
}
|
|
129
|
+
```
|
|
130
|
+
|
|
131
|
+
---
|
|
132
|
+
|
|
133
|
+
## Recommended workflow
|
|
134
|
+
|
|
135
|
+
1. Capture the current commit: `git rev-parse HEAD`.
|
|
136
|
+
2. Run `clio-coder components snapshot --out before.json` if the change affects prompts, tools, runtimes, safety, schemas, or recipes.
|
|
137
|
+
3. Draft the change manifest.
|
|
138
|
+
4. Implement the smallest change set.
|
|
139
|
+
5. Run the validation plan and update `evidenceRefs` with resolved evidence bundle IDs such as `run-<id>` or `session-<id>` where available.
|
|
140
|
+
6. Run `clio-coder evolve manifest validate` and include the manifest in review notes.
|
|
141
|
+
7. If the change fails, use `rollbackPlan` rather than ad hoc cleanup.
|
|
142
|
+
|
|
143
|
+
Change manifests are especially useful for experimental CLIO work because they separate evidence-backed claims from plans, hypotheses, and future milestones.
|
|
@@ -0,0 +1,74 @@
|
|
|
1
|
+
# Exit Codes & Machine-Readable Output Contracts
|
|
2
|
+
|
|
3
|
+
This document specifies the process exit codes, machine-readable JSON streaming formats, standard I/O separation rules, and `--help` conventions across all Clio Coder CLI commands in `v0.3.0`.
|
|
4
|
+
|
|
5
|
+
Source implementations: `src/cli/` and `src/entry/`.
|
|
6
|
+
|
|
7
|
+
---
|
|
8
|
+
|
|
9
|
+
## 1. Global Exit Codes
|
|
10
|
+
|
|
11
|
+
Clio Coder follows a deterministic exit code taxonomy across all commands:
|
|
12
|
+
|
|
13
|
+
| Exit Code | Meaning | Typical Causes & Conditions |
|
|
14
|
+
| :--- | :--- | :--- |
|
|
15
|
+
| **`0`** | **Success** | Successful command execution, clean run settlement, `--version`, `--help` invocation, or missing trace database notice without an explicit `--db` flag. |
|
|
16
|
+
| **`1`** | **Operational Failure** | Execution error, model target unreachable, doctor diagnosis with unresolved issues, explicit `--db` path not found, or evaluation rubric failure (`fail` or `error` verdict). |
|
|
17
|
+
| **`2`** | **Syntax / Usage Error** | Unknown subcommand, invalid flag, missing required positional arguments, global flag placed after subcommand, or data mutation SQL keyword passed to `clio-coder trace sql`. |
|
|
18
|
+
| **`3`** | **Unmeasured / Harness State** | Specific to `clio-coder skills eval`: rubric could not be evaluated due to unparseable judge output, timeout, or evidence archive write failure (distinct from a failure/regression). |
|
|
19
|
+
|
|
20
|
+
---
|
|
21
|
+
|
|
22
|
+
## 2. The `--help` Standard
|
|
23
|
+
|
|
24
|
+
Every subcommand in Clio Coder adheres to the strict `--help` convention:
|
|
25
|
+
|
|
26
|
+
1. **Standard Output**: Usage instructions and options are printed exclusively to `stdout`.
|
|
27
|
+
2. **Zero Exit**: The process exits with code `0`.
|
|
28
|
+
3. **Zero Side Effects**: Running `clio-coder <subcommand> --help` executes no runtime setup, initiates no network probes, and mutates no state files.
|
|
29
|
+
|
|
30
|
+
### Global vs Subcommand Flag Positioning
|
|
31
|
+
|
|
32
|
+
Global options (such as `--cwd`, `--config-dir`, `--state-dir`, `--profile`, and `--debug`) must precede the subcommand. If a global flag is placed after the subcommand name, Clio prints a remediation guide to `stderr` and exits with code `2`:
|
|
33
|
+
|
|
34
|
+
```text
|
|
35
|
+
--config-dir is a global option and must come before the subcommand: clio-coder --config-dir <path> <command> ...
|
|
36
|
+
```
|
|
37
|
+
|
|
38
|
+
---
|
|
39
|
+
|
|
40
|
+
## 3. Standard I/O Separation & Headless Execution
|
|
41
|
+
|
|
42
|
+
In headless execution (`clio-coder run`):
|
|
43
|
+
|
|
44
|
+
1. **Standard Output (`stdout`)**: Reserved strictly for the final answer, deliverable artifact content, or machine-readable JSON streams.
|
|
45
|
+
2. **Standard Error (`stderr`)**: Reserved for progress notifications, permission denial advisories, telemetry warnings, and error diagnostics.
|
|
46
|
+
3. **Headless Permission Denials**: When a tool requires permission that cannot be granted in headless mode, Clio emits `HEADLESS_PERMISSION_DENIED_REASON`:
|
|
47
|
+
```text
|
|
48
|
+
clio-coder run cannot confirm permission requests; rerun interactively to approve this action.
|
|
49
|
+
```
|
|
50
|
+
The denial is delivered to the LLM as a tool result so the agent can adapt or report the limitation. If the run completes after the denial, the process exits `0` with the answer on `stdout`.
|
|
51
|
+
|
|
52
|
+
---
|
|
53
|
+
|
|
54
|
+
## 4. Machine-Readable Output Formats (`--json` & `--json-events`)
|
|
55
|
+
|
|
56
|
+
Many Clio CLI subcommands provide structured JSON output for integration with scripts, CI pipelines, and external orchestrators.
|
|
57
|
+
|
|
58
|
+
### Subcommand JSON Summary
|
|
59
|
+
|
|
60
|
+
| Subcommand | Flag | Output Structure |
|
|
61
|
+
| :--- | :--- | :--- |
|
|
62
|
+
| `clio-coder run` | `--json` | Stream of incremental NDJSON event frames (`message`, `tool_call`, `tool_result`, `terminal`). |
|
|
63
|
+
| `clio-coder run` | `--json-events terminal` | Filters event stream to emit only the final terminal run receipt. |
|
|
64
|
+
| `clio-coder agents` | `--json` | JSON array of registered agent recipe metadata objects. |
|
|
65
|
+
| `clio-coder targets` | `--json` | JSON object containing the configured `targets` array. |
|
|
66
|
+
| `clio-coder models` | `--json` | JSON array of catalog models with capability flags. |
|
|
67
|
+
| `clio-coder fleet status` | `--json` | JSON snapshot of cluster nodes, active leases, and drain status. |
|
|
68
|
+
| `clio-coder trace runs` | `--json` | JSON array of trace run records. |
|
|
69
|
+
| `clio-coder trace sql` | Positional query | JSON array of rows returned by the read-only SQLite `SELECT` query. |
|
|
70
|
+
| `clio-coder paths` | `--json` | JSON object mapping platform directory names to absolute paths. |
|
|
71
|
+
|
|
72
|
+
### Incremental Streaming Invariant
|
|
73
|
+
|
|
74
|
+
The `--json` stream from `clio-coder run` emits **deltas and increments**, never repeated whole-message snapshots. This guarantees that consumers receive stream tokens linearly without duplicating memory or bandwidth.
|