@iowarp/clio-coder 0.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +407 -0
- package/CODE_OF_CONDUCT.md +21 -0
- package/CONTRIBUTING.md +224 -0
- package/LICENSE +202 -0
- package/NOTICE +9 -0
- package/README.md +798 -0
- package/SECURITY.md +72 -0
- package/assets/clio-coder-logo-128.webp +0 -0
- package/damage-control-rules.yaml +419 -0
- package/dist/acp-UMLFVA3F.js +92 -0
- package/dist/agents-Q4MYPMUW.js +91 -0
- package/dist/auth-O6HYIJ6J.js +521 -0
- package/dist/chunk-262G75JS.js +35 -0
- package/dist/chunk-26BZQOAD.js +1281 -0
- package/dist/chunk-2J63S4SF.js +508 -0
- package/dist/chunk-3DANZDGR.js +717 -0
- package/dist/chunk-4UQA7NCT.js +29 -0
- package/dist/chunk-527KG6XR.js +497 -0
- package/dist/chunk-5LDRNKX2.js +1063 -0
- package/dist/chunk-5N2FG33Q.js +25 -0
- package/dist/chunk-67MTHP2E.js +135 -0
- package/dist/chunk-6CWDTGUC.js +20 -0
- package/dist/chunk-7BHLZB3A.js +2115 -0
- package/dist/chunk-7RBKDI66.js +348 -0
- package/dist/chunk-AMFR5YA3.js +541 -0
- package/dist/chunk-BBUH4VAA.js +1224 -0
- package/dist/chunk-BYEU76JP.js +899 -0
- package/dist/chunk-CLJ5HLUD.js +458 -0
- package/dist/chunk-D5YD55AR.js +116 -0
- package/dist/chunk-DXQNI4PC.js +61 -0
- package/dist/chunk-E3NYWENM.js +1004 -0
- package/dist/chunk-GNGDQYDU.js +34688 -0
- package/dist/chunk-GOTUR54M.js +9 -0
- package/dist/chunk-HBU5MTAM.js +41 -0
- package/dist/chunk-HMYNFFY4.js +28 -0
- package/dist/chunk-JPOWPFCU.js +1010 -0
- package/dist/chunk-JWHCJDCI.js +1215 -0
- package/dist/chunk-KBR4MZZR.js +41 -0
- package/dist/chunk-KKKPTZLM.js +93 -0
- package/dist/chunk-ME6DNWIU.js +66 -0
- package/dist/chunk-NI4DEJMC.js +88 -0
- package/dist/chunk-O4EJEDHO.js +659 -0
- package/dist/chunk-PIDUD6M2.js +31 -0
- package/dist/chunk-PS4PFJQP.js +29459 -0
- package/dist/chunk-QV47YRF4.js +48 -0
- package/dist/chunk-RQDWMVRB.js +279 -0
- package/dist/chunk-TFSSEXL6.js +136 -0
- package/dist/chunk-TKHQ4DGZ.js +8290 -0
- package/dist/chunk-TPOCL34A.js +2876 -0
- package/dist/chunk-UGYAX5YI.js +565 -0
- package/dist/chunk-UHTSULZS.js +461 -0
- package/dist/chunk-UU3R62TT.js +128 -0
- package/dist/chunk-UWIJNAOB.js +3906 -0
- package/dist/chunk-VOO7NYPP.js +914 -0
- package/dist/chunk-VPAWTYLY.js +117 -0
- package/dist/chunk-WD6AJM35.js +1216 -0
- package/dist/chunk-X3BR7HWV.js +115 -0
- package/dist/chunk-X3NE4WVW.js +120 -0
- package/dist/chunk-XNISANGE.js +1395 -0
- package/dist/chunk-XV4ZJ6ZM.js +3177 -0
- package/dist/cli/index.js +236 -0
- package/dist/clio-KIQ5SNDS.js +53 -0
- package/dist/components-JVHMUBEB.js +653 -0
- package/dist/config-ZFCDBMDC.js +372 -0
- package/dist/configure-G4E3A2PG.js +27 -0
- package/dist/context-CDXTP2MP.js +293 -0
- package/dist/context-E3KIFVXI.js +185 -0
- package/dist/context-clear-3F4PLXOS.js +102 -0
- package/dist/context-index-Q7YSYTR3.js +106 -0
- package/dist/docs-YIETIWZI.js +280 -0
- package/dist/doctor-M5HJJZOL.js +61 -0
- package/dist/domains/agents/builtins/architect.md +33 -0
- package/dist/domains/agents/builtins/coder.md +31 -0
- package/dist/domains/agents/builtins/context-bootstrap.md +38 -0
- package/dist/domains/agents/builtins/debugger.md +30 -0
- package/dist/domains/agents/builtins/documenter.md +31 -0
- package/dist/domains/agents/builtins/git-master.md +30 -0
- package/dist/domains/agents/builtins/provenance.md +30 -0
- package/dist/domains/agents/builtins/researcher.md +71 -0
- package/dist/domains/agents/builtins/scout.md +42 -0
- package/dist/domains/agents/builtins/tester.md +31 -0
- package/dist/domains/agents/builtins/verifier.md +30 -0
- package/dist/domains/agents/builtins/wiki-writer.md +41 -0
- package/dist/eval-B3KZZESM.js +2674 -0
- package/dist/evidence-V67CHM35.js +233 -0
- package/dist/evolve-YDZSUQYA.js +518 -0
- package/dist/extensions-SRG7XCAH.js +207 -0
- package/dist/fleet-CA2CRTVG.js +760 -0
- package/dist/fleet-preflight-CLIAX7YR.js +21 -0
- package/dist/init-2OZDJE2D.js +227 -0
- package/dist/memory-3PIQQAKX.js +207 -0
- package/dist/models-DY35XI7Y.js +237 -0
- package/dist/paths-5OMXW7Z4.js +57 -0
- package/dist/preload-KZVHET2B.js +11 -0
- package/dist/reset-PIFYNOS3.js +216 -0
- package/dist/run-3VSPP24F.js +735 -0
- package/dist/share-D36RQCXM.js +241 -0
- package/dist/skills-F2MRLELY.js +445 -0
- package/dist/skills-eval-E2ZTW4PL.js +932 -0
- package/dist/targets-DZMEZAH4.js +977 -0
- package/dist/trace-7NYCUI2J.js +250 -0
- package/dist/uninstall-AD3JWHBB.js +322 -0
- package/dist/upgrade-WYYBKGDY.js +301 -0
- package/dist/usage-ULIDAGFF.js +755 -0
- package/dist/version-ROZ6CZKH.js +16 -0
- package/dist/wiki-generate-PKFIX6OB.js +377 -0
- package/dist/worker/entry.js +1739 -0
- package/docs/README.md +93 -0
- package/docs/acp.md +120 -0
- package/docs/alcf-provider.md +72 -0
- package/docs/architecture.md +172 -0
- package/docs/artifact-versions.md +54 -0
- package/docs/built-in-agents.md +265 -0
- package/docs/capacity-and-scheduling.md +97 -0
- package/docs/commands-and-modes.md +554 -0
- package/docs/config-knobs-audit.md +115 -0
- package/docs/configuration-and-targets.md +812 -0
- package/docs/context-engine.md +236 -0
- package/docs/dispatch-architecture-rationale.md +126 -0
- package/docs/documentation-coverage.md +46 -0
- package/docs/documentation-guide.md +166 -0
- package/docs/environment-variables.md +105 -0
- package/docs/eval-runner.md +205 -0
- package/docs/evals-internal.md +298 -0
- package/docs/evidence-and-memory.md +243 -0
- package/docs/evolution.md +143 -0
- package/docs/exit-codes-and-output.md +74 -0
- package/docs/extensions-and-sharing.md +306 -0
- package/docs/fleet-demo-runbook.md +179 -0
- package/docs/fleet-dispatch.md +591 -0
- package/docs/glossary.md +75 -0
- package/docs/html/agents_blueprint.html +936 -0
- package/docs/html/alcf_blueprint.html +324 -0
- package/docs/html/architecture_blueprint.html +850 -0
- package/docs/html/commands_blueprint.html +794 -0
- package/docs/html/config_knobs_audit_blueprint.html +178 -0
- package/docs/html/configuration_blueprint.html +1080 -0
- package/docs/html/context_blueprint.html +603 -0
- package/docs/html/documentation_blueprint.html +832 -0
- package/docs/html/environment_blueprint.html +404 -0
- package/docs/html/eval_blueprint.html +743 -0
- package/docs/html/evals_internal_blueprint.html +190 -0
- package/docs/html/evolution_blueprint.html +674 -0
- package/docs/html/extensions_blueprint.html +2065 -0
- package/docs/html/fleet_dispatch_blueprint.html +286 -0
- package/docs/html/index.html +919 -0
- package/docs/html/lifecycle_blueprint.html +723 -0
- package/docs/html/memory_blueprint.html +699 -0
- package/docs/html/middleware_blueprint.html +664 -0
- package/docs/html/models_blueprint.html +2366 -0
- package/docs/html/observability_blueprint.html +683 -0
- package/docs/html/provider_adapter_blueprint.html +245 -0
- package/docs/html/safety_blueprint.html +1386 -0
- package/docs/html/shared.css +571 -0
- package/docs/html/shared.js +143 -0
- package/docs/html/skills_blueprint.html +671 -0
- package/docs/html/soak_blueprint.html +182 -0
- package/docs/html/tool_usage_blueprint.html +350 -0
- package/docs/html/tools_blueprint.html +2249 -0
- package/docs/html/trace_blueprint.html +235 -0
- package/docs/html/tui_design_blueprint.html +314 -0
- package/docs/html/validation_blueprint.html +961 -0
- package/docs/html/worker_dispatch_blueprint.html +231 -0
- package/docs/installation-and-lifecycle.md +308 -0
- package/docs/middleware-and-components.md +148 -0
- package/docs/model-catalog.md +189 -0
- package/docs/observability.md +233 -0
- package/docs/proactive-memory.md +452 -0
- package/docs/prompt-envelope-and-tools.md +142 -0
- package/docs/provider-adapter-cookbook.md +148 -0
- package/docs/release-cut-checklist.md +138 -0
- package/docs/safety-model.md +357 -0
- package/docs/scientific-validation.md +105 -0
- package/docs/session-lifecycle.md +156 -0
- package/docs/skills-marketplace.md +46 -0
- package/docs/tool-usage.md +527 -0
- package/docs/trace-store.md +132 -0
- package/docs/troubleshooting.md +33 -0
- package/docs/tui-design.md +239 -0
- package/docs/worker-dispatch-mechanics.md +242 -0
- package/package.json +132 -0
- package/skills/README.md +408 -0
- package/skills/git/commit-crafting/SKILL.md +79 -0
- package/skills/git/commit-crafting/evals.md +92 -0
- package/skills/git/create-pr/SKILL.md +116 -0
- package/skills/git/create-pr/evals.md +114 -0
- package/skills/git/investigate-issue/SKILL.md +139 -0
- package/skills/git/investigate-issue/evals.md +94 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +96 -0
- package/skills/git/resolve-merge-conflicts/evals.md +58 -0
- package/skills/git/review-changes/SKILL.md +103 -0
- package/skills/git/review-changes/evals.md +85 -0
- package/skills/git/worktree-create/SKILL.md +92 -0
- package/skills/git/worktree-create/evals.md +97 -0
- package/skills/git/worktree-create/references/worktree-setup.md +66 -0
- package/skills/git/worktree-merge/SKILL.md +95 -0
- package/skills/git/worktree-merge/evals.md +114 -0
- package/skills/skill-marketplace.json +261 -0
- package/skills/workflow/cut-it/SKILL.md +86 -0
- package/skills/workflow/cut-it/evals.md +42 -0
- package/src/domains/agents/builtins/architect.md +33 -0
- package/src/domains/agents/builtins/coder.md +31 -0
- package/src/domains/agents/builtins/context-bootstrap.md +38 -0
- package/src/domains/agents/builtins/debugger.md +30 -0
- package/src/domains/agents/builtins/documenter.md +31 -0
- package/src/domains/agents/builtins/git-master.md +30 -0
- package/src/domains/agents/builtins/provenance.md +30 -0
- package/src/domains/agents/builtins/researcher.md +71 -0
- package/src/domains/agents/builtins/scout.md +42 -0
- package/src/domains/agents/builtins/tester.md +31 -0
- package/src/domains/agents/builtins/verifier.md +30 -0
- package/src/domains/agents/builtins/wiki-writer.md +41 -0
- package/src/domains/agents/fleets/build-review.md +34 -0
- package/src/domains/agents/fleets/build-test.md +35 -0
- package/src/domains/agents/fleets/sdlc.md +86 -0
- package/src/domains/prompts/fragments/identity/clio-worker.md +11 -0
- package/src/domains/prompts/fragments/identity/clio.md +26 -0
- package/src/domains/prompts/fragments/operating/contract.md +64 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +14 -0
- package/src/domains/prompts/fragments/safety/full-auto.md +14 -0
- package/src/domains/prompts/fragments/safety/read-only.md +13 -0
- package/src/domains/prompts/fragments/safety/suggest.md +13 -0
- package/src/domains/prompts/fragments/wiki/page.md +75 -0
- package/src/domains/prompts/fragments/wiki/plan.md +48 -0
- package/src/domains/providers/models/cloud-models/alcf.yaml +40 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +993 -0
|
@@ -0,0 +1,452 @@
|
|
|
1
|
+
# Proactive task memory
|
|
2
|
+
|
|
3
|
+
> **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.0).
|
|
4
|
+
|
|
5
|
+
Clio's proactive task memory protects long-running work from behavioral state
|
|
6
|
+
decay: a requirement, environment fact, failed attempt, or diagnosis can still
|
|
7
|
+
exist in the transcript while no longer influencing the next action. The design
|
|
8
|
+
follows Wu et al., *Remember When It Matters: Proactive Memory Agent for
|
|
9
|
+
Long-Horizon Agents* (2026), adapted to Clio's visible middleware and local-model
|
|
10
|
+
routing.
|
|
11
|
+
|
|
12
|
+
The rules-only tier is enabled by default and makes no model calls. An LLM memory
|
|
13
|
+
tier is opt-in through the independent `background` route. The action agent's
|
|
14
|
+
system prompt and tool surface do not change, and disabling
|
|
15
|
+
`memory.intervention.enabled` removes observation, bank writes, model resolution,
|
|
16
|
+
reminders, handoff offers, and handoff seeding.
|
|
17
|
+
|
|
18
|
+
## Architecture
|
|
19
|
+
|
|
20
|
+
```mermaid
|
|
21
|
+
flowchart LR
|
|
22
|
+
T[tool and lifecycle hooks] --> R[memory intervention registration]
|
|
23
|
+
R --> B[session task bank]
|
|
24
|
+
R --> D{trigger boundary}
|
|
25
|
+
D -->|rules only| S[deterministic policy]
|
|
26
|
+
D -->|background configured| L[two-phase local model policy]
|
|
27
|
+
S --> V[visible advisory reminder]
|
|
28
|
+
L --> V
|
|
29
|
+
V --> U[next user turn and session ledger]
|
|
30
|
+
R -. counts and outcomes only .-> J[bounded state JSONL]
|
|
31
|
+
B -. explicit context-handoff .-> H[redacted handoff snapshot]
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
The task bank is in `src/domains/memory/` and belongs to one live session. It is
|
|
35
|
+
separate from both durable approved lessons and the regenerable repository
|
|
36
|
+
context engine.
|
|
37
|
+
|
|
38
|
+
- **Private status** is the memory policy's short progress model. It can be
|
|
39
|
+
inspected with `/memory`, but is never rendered to either model, injected into
|
|
40
|
+
the action turn, or exported to a handoff.
|
|
41
|
+
- **Knowledge** contains stable task facts such as requirements, paths,
|
|
42
|
+
environment facts, and constraints.
|
|
43
|
+
- **Procedural memory** contains attempts and outcomes such as failed commands,
|
|
44
|
+
ruled-out hypotheses, diagnoses, and fixes that worked.
|
|
45
|
+
|
|
46
|
+
Knowledge and procedural entries have stable short IDs. A visible reminder is
|
|
47
|
+
one `Memory:` advisory block and records the cited entry IDs' injection counts.
|
|
48
|
+
The existing middleware path places that block in the next submitted user turn
|
|
49
|
+
and persists its attribution in the session ledger; there is no hidden
|
|
50
|
+
`transformContext` injection.
|
|
51
|
+
|
|
52
|
+
## Paper mapping and Clio constraints
|
|
53
|
+
|
|
54
|
+
| Paper mechanism | Clio implementation |
|
|
55
|
+
| --- | --- |
|
|
56
|
+
| Separate memory agent | One stateful middleware registration beside the unmodified action agent |
|
|
57
|
+
| Status, knowledge, and procedural bank | Bounded in-memory `TaskMemoryBank`; status stays out of every reminder except the post-compaction restore |
|
|
58
|
+
| Phase 1 bank maintenance | Strict `update_status`, `save_knowledge`, `save_procedural`, and `delete` operations |
|
|
59
|
+
| Phase 2 intervene or stay silent | One advisory `inject_reminder` effect or explicit silence |
|
|
60
|
+
| Fixed memory cadence | Deterministic decay signals plus a coarse interval floor |
|
|
61
|
+
| Learned intervention calibration | Structural authority gate: spontaneous reminders must cite a bank entry; deterministic triggers may be uncited |
|
|
62
|
+
| Passive and always-on ablations | A/B harness compares baseline, rules, and LLM tiers and flags always-noisy ties as regressions |
|
|
63
|
+
|
|
64
|
+
Model output uses a strict two-line grammar parsed by `src/domains/memory/task-memory-policy.ts`:
|
|
65
|
+
|
|
66
|
+
```text
|
|
67
|
+
<operations>[{"op":"update_status","content":"Tracking the current requirement."}]</operations>
|
|
68
|
+
<no_intervention/>
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
When a visible reminder cites an existing bank entry, the second line uses:
|
|
72
|
+
```text
|
|
73
|
+
<operations>[]</operations>
|
|
74
|
+
<context_for_action>Restored requirement [tm-k-1]</context_for_action>
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Knowledge and procedural saves use `{"op":"save_knowledge","content":"..."}` or
|
|
78
|
+
`{"op":"save_procedural","content":"..."}`; an `id` is only valid when updating
|
|
79
|
+
an existing entry.
|
|
80
|
+
|
|
81
|
+
The parser locates that envelope rather than matching the response byte for byte,
|
|
82
|
+
because a small local model routinely delivers a correct decision inside
|
|
83
|
+
imperfect packaging. A markdown fence, a `<think>` block, a leading "Here is my
|
|
84
|
+
step:", a closing pleasantry, and a pretty-printed multi-line operations array
|
|
85
|
+
are all accepted. Two shapes are read conservatively rather than generously:
|
|
86
|
+
|
|
87
|
+
- A response with no phase-two line at all is silence, since silence is the
|
|
88
|
+
prompt's documented default. Its phase-one writes still apply. An envelope
|
|
89
|
+
truncated mid-reasoning yields nothing at all.
|
|
90
|
+
- Tag shapes are stripped from the reminder before it is emitted, so the memory
|
|
91
|
+
model cannot close the `<system-reminder>` block it rides inside. Ordinary
|
|
92
|
+
comparisons and arrows survive.
|
|
93
|
+
|
|
94
|
+
Operations are validated structurally as a batch: a malformed entry or more than
|
|
95
|
+
eight operations rejects the whole list and changes nothing, because both say
|
|
96
|
+
the model did not produce an operation list at all. Two narrower mistakes cost
|
|
97
|
+
one operation instead of the step, because a small model makes both routinely
|
|
98
|
+
and the notes beside them are the point of the step.
|
|
99
|
+
|
|
100
|
+
Identity is repaired rather than rejected, because a small model invents a
|
|
101
|
+
descriptive id for content it is recording for the first time. A
|
|
102
|
+
`save_knowledge` or `save_procedural` whose id names no entry of that class
|
|
103
|
+
becomes a new entry, and a `delete` of an unknown id is dropped.
|
|
104
|
+
|
|
105
|
+
An unrecognized `op` is dropped the same way. Handed a JSON tool trajectory, a
|
|
106
|
+
small model borrows that trajectory's shape for an entry or two and answers
|
|
107
|
+
`{"op":"read","path":"..."}` beside otherwise valid saves; on the reference
|
|
108
|
+
route that happened in a third of sampled steps, and three of four such batches
|
|
109
|
+
carried a valid operation that the old whole-batch rejection discarded. A step
|
|
110
|
+
whose every operation was invented still records `malformed` rather than
|
|
111
|
+
passing as silence, since recovering nothing is not a decision to stay quiet.
|
|
112
|
+
|
|
113
|
+
Phase 1 writes remain valid when Phase 2 is gated or yields to a deterministic
|
|
114
|
+
reminder; an over-budget reminder is recorded as `gated` and suppressed rather
|
|
115
|
+
than discarding the writes that came with it. A timeout, provider failure,
|
|
116
|
+
malformed response, or telemetry failure is silent and never blocks a tool.
|
|
117
|
+
|
|
118
|
+
### Intervention Defaults & Cadence Knobs
|
|
119
|
+
- `memory.intervention.enabled` (default `true`): Enables observation, task bank writes, and reminder injection.
|
|
120
|
+
- `memory.intervention.everyNTools` (default `10`): Minimum completed-tool interval between background interventions.
|
|
121
|
+
- `memory.intervention.windowSteps` (default `8`): Completed tool-trajectory window analyzed during background evaluation.
|
|
122
|
+
- `memory.intervention.maxTokens` (default `400`): Bounds the rendered memory-bank and reminder context budget; the policy model output cap is a separate fixed `4,000`-token contract in `task-memory-policy.ts`, sized so that a model which reasons anyway still reaches its envelope.
|
|
123
|
+
- `memory.intervention.timeoutMs` (default `180000`): Wall-clock limit for one background memory-policy request. The step is detached, so this deadline never delays a turn; set it above the observed step time for your route or finished work is discarded as a timeout. Step latency on a small local route is long-tailed rather than tightly clustered, so size this off a high percentile and not off a median.
|
|
124
|
+
|
|
125
|
+
## Trigger semantics
|
|
126
|
+
|
|
127
|
+
Memory does not call a model after every tool. Signals accumulate and coalesce at
|
|
128
|
+
a turn-end boundary; at most one prompted step is started for that boundary, and
|
|
129
|
+
it runs detached from it.
|
|
130
|
+
|
|
131
|
+
| Trigger | Behavior |
|
|
132
|
+
| --- | --- |
|
|
133
|
+
| Interval | After `memory.intervention.everyNTools` completed tools since the last prompted step; default 10. This is the nondeterministic/citation-gated path. |
|
|
134
|
+
| Tool-error streak | Two consecutive error outcomes. A successful tool resets the streak. |
|
|
135
|
+
| Loop signal | Reuses the orchestrator loop guard's verdict; it does not infer a second competing loop detector. |
|
|
136
|
+
| Repeated failure | The rules tier records failed tool fingerprints and annotates the failing tool result once the same failure appears twice in the bounded trajectory. |
|
|
137
|
+
| Post-compaction | The first turn start after compaction restores status and knowledge once, without a model call, because compaction is precisely where execution facts leave the active window. |
|
|
138
|
+
|
|
139
|
+
### Two delivery channels
|
|
140
|
+
|
|
141
|
+
A turn boundary is the wrong place to warn about a failure that happened forty
|
|
142
|
+
tool calls earlier in the same turn, because the reminder cannot reach the model
|
|
143
|
+
until the operator submits again. Memory therefore has two channels, and each
|
|
144
|
+
repeated failure uses exactly one of them:
|
|
145
|
+
|
|
146
|
+
- **Mid-turn annotation.** The second identical failure appends one cited
|
|
147
|
+
`Memory:` advisory to that tool's own result, through the existing
|
|
148
|
+
`annotate_tool_result` effect the loop guard already uses. The advisory digest
|
|
149
|
+
takes the first line of the tool error that names a problem, falling back to
|
|
150
|
+
the first line when no line names one. The model reads it on its very next round.
|
|
151
|
+
This is spent once per fingerprint per turn and re-earned in a later turn, because
|
|
152
|
+
the same command failing again after an operator turn is news again.
|
|
153
|
+
- **Next-turn reminder.** Post-compaction reactivation and any background-model
|
|
154
|
+
reminder ride the `inject_reminder` buffer into the next submitted turn, inside
|
|
155
|
+
the visible `<system-reminder>` block, and persist in the session ledger.
|
|
156
|
+
|
|
157
|
+
A boundary that already spoke through the annotation stays silent at turn end and
|
|
158
|
+
records one telemetry row, not two.
|
|
159
|
+
|
|
160
|
+
## Background steps never hold a turn open
|
|
161
|
+
|
|
162
|
+
The pi agent does not become idle until every `agent_end` listener settles, so a
|
|
163
|
+
memory step awaited at that boundary would add its full latency to the visible
|
|
164
|
+
end of every triggered turn. Measured on the reference route below, step latency
|
|
165
|
+
has a median of 18.6 seconds and ranges up to 220.8 seconds. This makes an awaited
|
|
166
|
+
step intolerable as an end-of-turn pause.
|
|
167
|
+
|
|
168
|
+
The prompted step is therefore detached. `evaluateAsync` starts it and returns
|
|
169
|
+
immediately; the turn ends on schedule. When the step resolves, its reminder is
|
|
170
|
+
delivered through the deferred-reminder path into the next submitted turn, which
|
|
171
|
+
is exactly where an awaited turn_end reminder would have been buffered anyway.
|
|
172
|
+
Two consequences follow, both deliberate:
|
|
173
|
+
|
|
174
|
+
- At most one background step is alive per session. A boundary that arrives while
|
|
175
|
+
a step is still running is dropped rather than queued, so a model slower than
|
|
176
|
+
the turns that trigger it can never build a backlog. The drop is recorded as
|
|
177
|
+
its own telemetry row, because a cadence starved by a slow route and one that
|
|
178
|
+
simply never triggered are otherwise identical in the step log. Its triggers
|
|
179
|
+
stay pending, so the next free boundary still runs for them.
|
|
180
|
+
- A reminder can arrive one turn later than the trajectory that earned it. The
|
|
181
|
+
rules tier is unaffected and stays synchronous, so deterministic protection
|
|
182
|
+
keeps its original timing.
|
|
183
|
+
|
|
184
|
+
`/memory` shows whether a step is in flight, and the footer's memory row shows
|
|
185
|
+
`working` while one is running.
|
|
186
|
+
|
|
187
|
+
The error-streak, loop, repeated-failure, and post-compaction paths are
|
|
188
|
+
deterministic. A prompted reminder from one of those paths may be uncited. An
|
|
189
|
+
interval-only prompted reminder must cite at least one current knowledge or
|
|
190
|
+
procedural ID or it is recorded as `gated` and remains invisible.
|
|
191
|
+
|
|
192
|
+
### Outcome semantics
|
|
193
|
+
|
|
194
|
+
The `/memory` overlay displays `last <decision>` where `<decision>` is the
|
|
195
|
+
combined outcome of the most recent actual memory boundary. A **memory boundary**
|
|
196
|
+
is a turn-end evaluation that includes newly completed tools or an explicit
|
|
197
|
+
deterministic trigger (interval, error-streak, loop signal). A no-tool
|
|
198
|
+
middleware continuation is another turn-end with no new tools since the previous
|
|
199
|
+
boundary. It is not a new memory boundary and does not replace the prior outcome.
|
|
200
|
+
Thus `last` remains `injected` across such continuations until a later
|
|
201
|
+
tool-bearing or explicitly triggered memory step produces a new outcome (e.g.,
|
|
202
|
+
a healthy tool leading to `silent`).
|
|
203
|
+
|
|
204
|
+
## Choosing a background model
|
|
205
|
+
|
|
206
|
+
Memory reads a trajectory and writes a fixed envelope. It does not plan, and it
|
|
207
|
+
does not need to be clever. A small non-reasoning model is the right choice, and
|
|
208
|
+
Clio always requests the background route with thinking off regardless of
|
|
209
|
+
`background.thinkingLevel`.
|
|
210
|
+
|
|
211
|
+
That request reaches the wire wherever the runtime carries a thinking control:
|
|
212
|
+
llama.cpp reads `chat_template_kwargs.enable_thinking`, and LM Studio reads
|
|
213
|
+
`reasoning_effort`, where `none` is the off value.
|
|
214
|
+
|
|
215
|
+
A model that reasons anyway still works. Some genuinely cannot be silenced, and
|
|
216
|
+
the catalog records those as always-on so the level reads `forced` rather than
|
|
217
|
+
`off`; the shipped background model `qwopus3.5-9b-v3` is one of them. Reasoning
|
|
218
|
+
blocks are discarded and only the envelope is kept, and the output budget is
|
|
219
|
+
sized to let a reasoning preamble run its course first. The cost is latency,
|
|
220
|
+
which the detached step absorbs.
|
|
221
|
+
|
|
222
|
+
One configuration is refused rather than degraded. If the background role names
|
|
223
|
+
the same target and model as the orchestrator, and that model reasons, the LLM
|
|
224
|
+
memory tier stays off and memory runs on its free deterministic tier. A single
|
|
225
|
+
reasoning model already driving chat, workers, and shadow agents cannot also
|
|
226
|
+
deliberate over memory steps without contending with the work the operator
|
|
227
|
+
actually asked for.
|
|
228
|
+
|
|
229
|
+
This is a mix-and-match plane, not a local-only one. The background role resolves
|
|
230
|
+
through the same target machinery as every other role, so the useful shapes are:
|
|
231
|
+
|
|
232
|
+
- A frontier model for chat and a small efficient model for memory, whether that
|
|
233
|
+
small model is co-hosted, on another node, or a cheap cloud tier.
|
|
234
|
+
- A local workhorse for chat and a co-resident small local model for memory, with
|
|
235
|
+
the co-residency caveat below.
|
|
236
|
+
- Everything cloud: pick the provider's small fast model for memory and spend the
|
|
237
|
+
budget on the agent and the fleet.
|
|
238
|
+
|
|
239
|
+
Local co-residency still matters. The background model, the action model, their
|
|
240
|
+
KV caches, and parallel slots must all fit the target's available memory.
|
|
241
|
+
|
|
242
|
+
## Operator setup
|
|
243
|
+
|
|
244
|
+
The shipped defaults are:
|
|
245
|
+
|
|
246
|
+
```yaml
|
|
247
|
+
background:
|
|
248
|
+
target: null
|
|
249
|
+
model: null
|
|
250
|
+
thinkingLevel: off
|
|
251
|
+
|
|
252
|
+
memory:
|
|
253
|
+
intervention:
|
|
254
|
+
enabled: true
|
|
255
|
+
everyNTools: 10
|
|
256
|
+
windowSteps: 8
|
|
257
|
+
maxTokens: 400
|
|
258
|
+
timeoutMs: 180000
|
|
259
|
+
```
|
|
260
|
+
|
|
261
|
+
With `background.target` and `background.model` unset, Clio stays in the
|
|
262
|
+
zero-cost rules tier. `/memory` shows the current tier, last decision, approved
|
|
263
|
+
durable lessons, the live bank, and a bounded history of the last twenty memory
|
|
264
|
+
steps with their trigger, decision, write count, cited-entry count, tier, and
|
|
265
|
+
latency. That history is the only place a capture, a gate, or a timeout becomes
|
|
266
|
+
visible, since those outcomes produce no transcript entry by design; only an
|
|
267
|
+
actual injection reaches the transcript. It carries counts and outcomes only,
|
|
268
|
+
never bank or trajectory text. `/settings` exposes every key above. In
|
|
269
|
+
`/targets`, select an eligible local target and press `b` to make it the saved
|
|
270
|
+
background-memory default; this is independent of `u` for chat and `f` for the
|
|
271
|
+
fleet default. A running session owns its routing snapshot, while the saved
|
|
272
|
+
selection becomes the default for new sessions.
|
|
273
|
+
|
|
274
|
+
The reference live configuration is an LM Studio server on the `zbook` node with
|
|
275
|
+
the wire model `qwopus3.5-9b-v3`:
|
|
276
|
+
|
|
277
|
+
```yaml
|
|
278
|
+
background:
|
|
279
|
+
target: zbook
|
|
280
|
+
model: qwopus3.5-9b-v3
|
|
281
|
+
thinkingLevel: off
|
|
282
|
+
```
|
|
283
|
+
|
|
284
|
+
A small model is the intended shape for this role. Across 60 measured steps on
|
|
285
|
+
that route, latency ran 4.4 to 220.8 seconds with a median of 18.6, a 90th
|
|
286
|
+
percentile of 79.9, and a 95th of 131.6. Capability is not the constraint;
|
|
287
|
+
latency is, its spread is wide, and the detached step above is what makes the
|
|
288
|
+
tier usable anyway.
|
|
289
|
+
|
|
290
|
+
Size `timeoutMs` off that tail rather than off the median. The shipped 180000
|
|
291
|
+
captures roughly the whole distribution on this route. A 20000 setting looks
|
|
292
|
+
generous against an 18.6-second median and in practice discarded about half of
|
|
293
|
+
all steps, since the step still runs to completion and only its result is thrown
|
|
294
|
+
away. A route whose steps mostly record `timeout` is a misconfigured deadline
|
|
295
|
+
before it is a slow model.
|
|
296
|
+
|
|
297
|
+
The target ID is not hard-coded. Any configured orchestrator-eligible local
|
|
298
|
+
target and wire model can fill the background role. Before enabling it, use the
|
|
299
|
+
real target surfaces to verify the route:
|
|
300
|
+
|
|
301
|
+
```bash
|
|
302
|
+
clio-coder targets --probe
|
|
303
|
+
clio-coder models --target zbook
|
|
304
|
+
clio-coder
|
|
305
|
+
```
|
|
306
|
+
|
|
307
|
+
Then inspect `/targets`, `/settings`, and `/memory`. Local co-residency still
|
|
308
|
+
matters: the background model, action model, their KV caches, and parallel slots
|
|
309
|
+
must fit the target's available memory. Increase `timeoutMs` for a deliberately
|
|
310
|
+
slow local route; lowering `maxTokens` bounds the visible reminder but does not
|
|
311
|
+
change the background model's strict output grammar.
|
|
312
|
+
|
|
313
|
+
For an immediate kill switch, set `memory.intervention.enabled` to `false` in
|
|
314
|
+
`/settings`. Removing the background target instead returns to rules-only
|
|
315
|
+
operation while leaving deterministic protection active.
|
|
316
|
+
|
|
317
|
+
## What the LLM tier actually writes
|
|
318
|
+
|
|
319
|
+
Measured on the shipped prompt against `google/gemma-4-26b-a4b-qat`, across ten
|
|
320
|
+
live steps and forty controlled runs on the same route.
|
|
321
|
+
|
|
322
|
+
The tier writes `update_status` reliably and `save_knowledge` rarely, and that is
|
|
323
|
+
correct rather than broken. A trajectory step carries the tool name, a bounded
|
|
324
|
+
call description, an outcome, and a result digest. On success the digest is an
|
|
325
|
+
opaque result fingerprint, so a window of successful reads tells the model which
|
|
326
|
+
files were touched and nothing about what is in them. There is no durable fact in
|
|
327
|
+
that input, and a status line is the only faithful thing to write about it.
|
|
328
|
+
|
|
329
|
+
Three candidate causes were ruled out by controlled runs that changed one
|
|
330
|
+
variable at a time:
|
|
331
|
+
|
|
332
|
+
- rewriting the prompt's second worked example to carry a `save_knowledge` moved
|
|
333
|
+
nothing, and made the model emit no operations at all in four of five runs;
|
|
334
|
+
- seeding the bank with existing knowledge entries so the model could learn the
|
|
335
|
+
shape by example moved nothing;
|
|
336
|
+
- giving successful steps a content-bearing digest moved nothing on its own.
|
|
337
|
+
|
|
338
|
+
What does elicit knowledge is a durable fact in the input. With a task stating two
|
|
339
|
+
explicit constraints the model wrote both as knowledge; with that same task plus
|
|
340
|
+
content-bearing digests it wrote six. Error digests already carry a real
|
|
341
|
+
diagnostic line, which is why `save_procedural` fires on failing windows.
|
|
342
|
+
|
|
343
|
+
The consequence for reactivation is why the post-compaction block restores status.
|
|
344
|
+
Restoring knowledge alone restored nothing in the common case, because the one
|
|
345
|
+
class it read was usually the one class the model had not written.
|
|
346
|
+
|
|
347
|
+
Two further numbers from the same route. Roughly a quarter of live steps returned a
|
|
348
|
+
malformed envelope, usually `<operations>` with no list followed by
|
|
349
|
+
`<no_intervention/>`, which is recorded as `malformed`/`unparseable` and is
|
|
350
|
+
model behavior rather than a route fault. The boundary drop rate remains 0% at
|
|
351
|
+
shipped settings.
|
|
352
|
+
|
|
353
|
+
## Handoff continuity
|
|
354
|
+
|
|
355
|
+
The bank normally dies with the session. When `context-handoff` is explicitly
|
|
356
|
+
requested, Clio supplies the skill a redacted `clio-task-memory` fenced snapshot
|
|
357
|
+
containing knowledge and procedural entries only. Ordinary turns receive no
|
|
358
|
+
snapshot. The handoff artifact remains under ignored `.clio-coder/handoffs/`; private
|
|
359
|
+
status and secret-shaped values do not cross the export boundary.
|
|
360
|
+
|
|
361
|
+
After `/resume`, Clio checks only the newest handoff and offers `/memory seed` if
|
|
362
|
+
it contains a valid snapshot. Seeding is explicit and deduplicated. It resets
|
|
363
|
+
injection attribution for the new session, and the master kill switch disables
|
|
364
|
+
both the offer and writes. `/new`, `/fork`, `/resume`, and ACP session changes
|
|
365
|
+
clear the prior heap bank before the new session can observe it.
|
|
366
|
+
|
|
367
|
+
## Telemetry
|
|
368
|
+
|
|
369
|
+
Each completed memory step appends one content-free record to:
|
|
370
|
+
|
|
371
|
+
```text
|
|
372
|
+
<stateDir>/memory/steps.jsonl
|
|
373
|
+
```
|
|
374
|
+
|
|
375
|
+
Use `clio-coder paths --json` to resolve the state directory (the `"state"` property). The log rotates after 1 MiB and
|
|
376
|
+
keeps one previous generation as `steps.jsonl.1`. Every exact-schema record has:
|
|
377
|
+
|
|
378
|
+
- timestamp and schema version;
|
|
379
|
+
- one to three coalesced trigger reasons;
|
|
380
|
+
- `rules` or `llm` tier;
|
|
381
|
+
- per-class added, updated, and deleted entry counts;
|
|
382
|
+
- `silent`, `injected`, `gated`, `timeout`, `malformed`, or `dropped` decision;
|
|
383
|
+
- count of cited entries, input/output/total memory-model tokens, and latency.
|
|
384
|
+
|
|
385
|
+
`dropped` is the one outcome that ran no step: the boundary triggered while an
|
|
386
|
+
earlier step still held the single in-flight slot. It costs no tokens and no
|
|
387
|
+
latency, its triggers survive to the next free boundary, and it does not replace
|
|
388
|
+
the operator-visible last decision. Counting `dropped` rows against `llm` rows
|
|
389
|
+
over a session is how a starved cadence becomes visible.
|
|
390
|
+
|
|
391
|
+
The log contains no task, trajectory, bank, error, or reminder text. File creation,
|
|
392
|
+
rotation, serialization, and injected sinks are all best effort; a read-only
|
|
393
|
+
state directory or full disk cannot alter intervention behavior.
|
|
394
|
+
|
|
395
|
+
Note that routine no-tool continuation checks do not emit telemetry rows, as they
|
|
396
|
+
are not considered new memory boundaries. Only actual tool-bearing or explicitly
|
|
397
|
+
triggered memory steps produce rows.
|
|
398
|
+
|
|
399
|
+
## Evaluation and promotion bar
|
|
400
|
+
|
|
401
|
+
`src/domains/eval/proactive-memory.ts` exports a fixed three-task, matched A/B
|
|
402
|
+
harness. It executes `baseline`, `rules`, and `llm` variants in stable order and
|
|
403
|
+
accepts any `{ id, model }` target. A runner adapter owns isolated task execution
|
|
404
|
+
and returns action tokens/latency plus the exact telemetry rows emitted for that
|
|
405
|
+
trial. The report provides:
|
|
406
|
+
|
|
407
|
+
- pass rate;
|
|
408
|
+
- injected and cited reminder counts;
|
|
409
|
+
- reminders per task and citation rate;
|
|
410
|
+
- total and baseline-relative added tokens and latency;
|
|
411
|
+
- an `alwaysNoisyRegression` verdict.
|
|
412
|
+
|
|
413
|
+
The deterministic end-to-end harness contract can be run directly:
|
|
414
|
+
|
|
415
|
+
```bash
|
|
416
|
+
npm run test:file -- tests/contracts/proactive-memory-eval.test.ts
|
|
417
|
+
```
|
|
418
|
+
|
|
419
|
+
For a live local comparison, an adapter should route only the `llm` variant
|
|
420
|
+
through the request's target/model (the reference is `zbook` /
|
|
421
|
+
`qwopus3.5-9b-v3`), keep baseline memory telemetry empty, and run all
|
|
422
|
+
nine trials in equivalent isolated workspaces. Do not promote the LLM tier from
|
|
423
|
+
one anecdotal task. The evidence bar is a pass-rate gain from a small number of
|
|
424
|
+
specific, usually cited reminders at acceptable added token and latency cost.
|
|
425
|
+
Injecting at least once per task while merely tying or losing to baseline is
|
|
426
|
+
always a regression, even when every reminder is cited.
|
|
427
|
+
|
|
428
|
+
## Worker growth path
|
|
429
|
+
|
|
430
|
+
Worker-side intervention is intentionally not implemented in this sprint. The
|
|
431
|
+
bank, policy client, telemetry, and registration interfaces carry no interactive
|
|
432
|
+
chat-loop types, so they can be instantiated per worker later without moving the
|
|
433
|
+
policy into the action agent.
|
|
434
|
+
|
|
435
|
+
The existing transport already exposes the required seams:
|
|
436
|
+
|
|
437
|
+
1. `src/domains/dispatch/worker-spawn.ts` receives worker NDJSON events and its
|
|
438
|
+
`SpawnedWorker.send` path can write bounded control messages while the worker
|
|
439
|
+
is alive.
|
|
440
|
+
2. Worker steering already drains between tool batches, which is the safe point
|
|
441
|
+
for a visible memory advisory; it must not interrupt a tool in flight.
|
|
442
|
+
3. Workers already maintain per-worker loop detectors and tool-call caps. A
|
|
443
|
+
future registration should consume those verdicts instead of re-deriving
|
|
444
|
+
them.
|
|
445
|
+
4. Each dispatched run needs its own bank, cadence, spend guard, telemetry
|
|
446
|
+
attribution, and teardown. Parent session memory must not leak into sibling
|
|
447
|
+
workers implicitly.
|
|
448
|
+
|
|
449
|
+
The future sequence is therefore worker events → worker-local registration → one
|
|
450
|
+
bounded steering advisory between batches. It must preserve the current receipt,
|
|
451
|
+
safety, timeout, and permission semantics, and it should ship only after a
|
|
452
|
+
Terminal-Bench-style long-run evaluation shows a selective benefit.
|
|
@@ -0,0 +1,142 @@
|
|
|
1
|
+
# Prompt Envelope and Tools
|
|
2
|
+
|
|
3
|
+
> [!TIP]
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/tools_blueprint.html](html/tools_blueprint.html) (Version: 0.3.0).
|
|
5
|
+
|
|
6
|
+
Clio Coder keeps the model-facing envelope stable and moves enforcement into the runtime registry and safety policy.
|
|
7
|
+
|
|
8
|
+
Source of truth: `src/core/tool-names.ts`, `src/tools/agent-tools.ts`, `src/tools/bootstrap.ts`, `src/tools/policy.ts`, `src/tools/observation.ts`, `src/tools/ignore-policy.ts`, and the per-tool modules under `src/tools/**`.
|
|
9
|
+
|
|
10
|
+
## One system prompt per session
|
|
11
|
+
|
|
12
|
+
The chat loop compiles one provider-facing system prompt for a session. The compile key is `target|model|autonomy|sessionId|workingContextPaths`, with the working-context paths sorted before hashing into the key.
|
|
13
|
+
|
|
14
|
+
The compiled prompt is reused byte-for-byte on ordinary submits. It recompiles only when that key changes or when config hot-reload invalidates the prompt cache. Path-scoped project rules can therefore recompile the prompt when a matching file enters working context. When recompilation changes the text, the session ledger records a `promptRecompiled` entry with the previous hash, new hash, and token estimate.
|
|
15
|
+
|
|
16
|
+
Prompt extensions can add dynamic fragments for project rules, the operator profile, and Clio source-tree awareness. Pending skill requests and middleware reminders are visible text in the user message, not hidden prompt machinery.
|
|
17
|
+
|
|
18
|
+
The Tool Contract section of the prompt renders a fixed set of base lines plus one optional guidance sentence per tool, sourced from the tool registry (`ToolMetadata.promptHint` in `src/tools/registry.ts`, assigned in `src/tools/bootstrap.ts`). The base lines cover the complete-surface rule, tool-free answering, orientation preferences, a deterministic routing order (structured observation before bash, task board for multi-step work, bounded dispatch with receipt synthesis, validation before final claims), failure recovery through `context(scope="docs")` instead of blind retries, and the skill-listing gate (skill-shaped tasks or explicit operator skill requests only). The chat loop derives the hint list once from the session's frozen tool surface at compile time, and the compiler renders the hints sorted by tool name, so the compiled text depends only on which hinted tools are on the surface. Today five tools carry hints: `ask_user`, `code_nav`, `context`, `dispatch`, and `tasks`. Removing a tool from the surface removes its hint with no compiler change; adding a hint to a tool is a deliberate prompt-text change that must land with updated prompt contract tests and a CHANGELOG note.
|
|
19
|
+
|
|
20
|
+
## One tool surface per session
|
|
21
|
+
|
|
22
|
+
For tool-capable providers, Clio sends the full registry as the session tool surface. The list is deterministic and sorted through the worker-tool resolver (`resolveAgentTools` in `src/tools/agent-tools.ts`), so the serialized schemas stay byte-identical on every submit. `src/tools/agent-tools.ts` is the single agent-tool adapter across the codebase. Both the orchestrator session and worker subprocesses resolve their tool set through the same `effectiveToolNames` narrowing function, ensuring that the attested signature and runtime surface cannot diverge.
|
|
23
|
+
|
|
24
|
+
Tools are keyed strictly by the canonical `ToolName` union defined in `src/core/tool-names.ts` with no alias table. Pure and idempotent `prepareArguments` normalizers defined on `ToolSpec` serve as the sole leniency layer for coercing legacy or weak-model parameter formats.
|
|
25
|
+
|
|
26
|
+
Tool visibility is not a per-turn hinting system. Pending-skill policy, ask-user policy, Bash policy, path policy, protected artifacts, dispatch admission, middleware, and the autonomy mapping are enforced when a tool is invoked. The `autonomy` level is applied at registry admission after the safety net passes a call; the safety prompt fragment mirrors that enforced matrix as guidance to the model. Prompt text and provider schemas do not bypass the registry.
|
|
27
|
+
|
|
28
|
+
Providers that cannot call tools receive no schemas, and the prompt tells the model to proceed without tool calls.
|
|
29
|
+
|
|
30
|
+
## Canonical worker harness
|
|
31
|
+
|
|
32
|
+
Native and mediated dispatch workers use a separate prompts-domain compiler over the same loaded fragment table. Its stable system prompt has exactly five sections: identity-lite, the shared operating contract plus assigned-task rules, a tool contract sliced to the final canonical toolkit, safety for the single effective autonomy, and one final persona. A request persona override replaces only the recipe body; eligible bound-skill instructions are composed inside that same final persona and never widen tools.
|
|
33
|
+
|
|
34
|
+
The compiler runs after target capability and tool-profile admission. Its canonical tool names are the same names transported in `WorkerSpec.allowedTools` and attached as schemas; routine non-Scout work removes `code_nav`, narrow profiles remove their excluded schemas and guidance, tool-incapable targets get an explicit no-tools contract, and Claude SDK aliases are filtered from the same canonical set. ACP's external inventory is unknown, so ACP bounded-role admission continues to validate the unchanged raw persona rather than fabricating a complete native schema list.
|
|
35
|
+
|
|
36
|
+
Project context, memory, bounded dispatch briefing, pipeline input, the assigned task, and the per-run safety-posture reminder remain dynamic user messages. A briefing is a separately delimited message labeled as untrusted task context/data; it is never concatenated into the task or stable system prompt. Dynamic ordering is project, safety, memory, briefing, then pipeline input, with pipeline input last. These messages do not affect the stable composition hash. Persona, effective autonomy, target tool capability, or final toolkit changes do affect it.
|
|
37
|
+
|
|
38
|
+
## Seven planes, twenty tools
|
|
39
|
+
|
|
40
|
+
The builtin surface is 20 registered tools organized in seven planes. Each plane is one policy unit: its tools share an action class, a size posture, a details schema, and a concurrency rule. `src/tools/policy.ts` asserts these invariants at bootstrap, so drift between the plane design, the safety classifier, and the registered specs fails loudly instead of shipping a surface that behaves differently from what the policy engine assumes.
|
|
41
|
+
|
|
42
|
+
| Plane | Tools | Action class | Concurrency |
|
|
43
|
+
| --- | --- | --- | --- |
|
|
44
|
+
| OBSERVE | `read`, `grep`, `find`, `ls`, `code_nav`, `context`, `credential_present` | read | parallel |
|
|
45
|
+
| MUTATE | `write`, `edit` | write | sequential |
|
|
46
|
+
| EXECUTE | `bash`, `verify` | execute | sequential |
|
|
47
|
+
| EXECUTE | `git` | read | parallel |
|
|
48
|
+
| ORCHESTRATE | `dispatch`, `steer` | dispatch | sequential |
|
|
49
|
+
| ORCHESTRATE | `monitor` | read | parallel |
|
|
50
|
+
| ORCHESTRATE | `tasks` | read | sequential |
|
|
51
|
+
| ORCHESTRATE | `ledger` | read | sequential |
|
|
52
|
+
| RETRIEVE | `web_fetch` | read | parallel |
|
|
53
|
+
| INTERACT | `ask_user` | read | sequential |
|
|
54
|
+
| ARTIFACT | `artifact` | write | sequential |
|
|
55
|
+
|
|
56
|
+
Three tools sit in a plane for containment rather than class. `git` is read-only inspection (op=status/diff/log) that runs on the safe-exec spine, so it lives in the EXECUTE plane with read-class safety disposition. `monitor` never mutates a run, so it stays read class and parallel inside the ORCHESTRATE plane. `tasks` orchestrates the agent's own work rather than workers: it mutates only the session's task ledger, never the workspace, so it keeps read class (never gated behind a confirmation) but runs sequential so two board mutations in one batch cannot interleave. `ledger` is the agent ledger, the coordination board concurrent dispatch workers share: a post reaches a one-way control lane and a read answers from a local mirror, so it touches no workspace and stays read class, and reviewers and judges are pinned to read-only autonomy where a write class would block the peer review the board exists for.
|
|
57
|
+
|
|
58
|
+
Registration is conditional on wiring: `context` gains its workspace scope only when a session contract is bound, `dispatch`/`monitor`/`steer` register only with a dispatch contract, and `ask_user` registers only when an interactive handler exists. Dispatch tool profiles narrow the surface for workers: `minimal-local` is `read`, `grep`, `find`, `ls`, `git`, `context`, `code_nav`; `science-local` adds `verify`; `full-agent` keeps everything.
|
|
59
|
+
|
|
60
|
+
### Consolidated call shapes
|
|
61
|
+
|
|
62
|
+
Several tools absorb what used to be separate tools:
|
|
63
|
+
|
|
64
|
+
- `find(pattern, path?, order?, limit?, include_ignored?)` locates paths by glob pattern (`*`, `**`, `?`, `[abc]`), default limit 500. `order="path"` (default) returns fd's native order; `order="mtime"` returns newest first from a bounded candidate set instead of statting the whole tree, and reports `details.candidates` when the candidate cap made the ordering approximate.
|
|
65
|
+
- `grep(pattern, path?, mode?, glob?, ignore_case?, literal?, context?, limit?, include_ignored?)` searches file contents with ripgrep, degrading to a bounded pure-Node search when rg is absent. `mode=content` (default) returns line-referenced matches, `mode=files` returns matching paths, `mode=count` returns per-file counts. Context lines are consumed from rg's `--json` stream.
|
|
66
|
+
- `context(scope="workspace"|"docs"|"skills")` is the one OBSERVE entry point for material about the working environment: the session workspace snapshot, retrieval over Clio's bundled documentation (`query` required), and skill listing or loading (`name` optional, `include_tree` for the skill's resource files).
|
|
67
|
+
- `verify(check?, path?, args?, browser?, cwd?, timeout_ms?)` runs declared verification. `verify()` with no arguments lists declared checks grouped by source (package.json verification scripts today), `verify(check="<script>")` runs one through the safe-exec spine with no shell, and `verify(check="frontend", path=...)` validates an HTML/CSS/JS artifact without granting shell access.
|
|
68
|
+
- `artifact(kind="plan"|"review"|"report", content, ...)` writes named artifacts behind one surface: Markdown documents (default `PLAN.md`/`REVIEW.md`/`REPORT.md` at the project root; `path` may override inside the workspace) that terminate the turn, because writing the artifact is the answer. Skills are not artifacts; a `SKILL.md` is written with the ordinary write tool and validated by the skills loader.
|
|
69
|
+
- `dispatch(task?, tasks?, mode?, ...)` supports a first-class singular assignment (`task`) and a batch (`tasks`), never both. `task` is worker instructions; `briefing` is optional bounded parent context/data and cannot replace it. Briefing stays a separate dynamic message and receipt provenance, never part of the receipt task. A shared top-level briefing applies to strings and objects without an override; an object-level briefing wins. Blank values are omitted, the cap is 12,000 UTF-8 bytes, and approval pins the exact canonical value. Ordinary handles enter one registered event consumer immediately. Synchronous calls auto-wait for stream-and-receipt completion; `detach:true` returns ids after durable batch registration while the same consumer continues. Review and compete retain gate-sensitive direct drains. Task objects may include `persona` and `tool_profile`. Pipeline output is threaded as bounded data. A successful native or ACP run requires a nonempty receipt-sealed final output; exit zero without one fails as `worker_final_output_missing`, with unfinished text retained only as partial diagnostics. `dispatch(list=true)` renders the catalog.
|
|
70
|
+
- `monitor(run_id?, mode?)` is read-only visibility into known synchronous and detached runs: `list` enumerates, `status` reports one, `peek` returns the in-process event tail, `receipt` exposes the stored evidence, and `wait` observes one run without collecting or canceling it. `collect` is the authoritative terminal batch operation over a detached batch or run-id list; collect before final synthesis. Completed output reports receipt integrity, evidence verification, briefing provenance, and bounded project-context provenance as different fields.
|
|
71
|
+
- `steer(run_id, action, message?)` controls a running worker: `guide` writes a canonical trimmed steering message to an HTTP or SDK worker and `cancel` terminates it. Successfully written steers gain ordered byte/hash/timestamp provenance; after the runtime accepts the guidance, `clio_steer_received` acknowledges the exact matching sequence, and prose is never stored in ledger or receipt. Single-shot subprocess runtimes and ACP remain non-steerable. Interactive operators can steer synchronous live-input runs; parent-model steering requires detached ids because model tools are sequential.
|
|
72
|
+
|
|
73
|
+
### One ignore policy for path walkers
|
|
74
|
+
|
|
75
|
+
`grep`, `find`, and their pure-Node fallbacks answer "which parts of the tree are visible" from one shared policy in `src/tools/ignore-policy.ts`. Three layers apply: `.clio-coder`, `.fallow`, and `.git` are always excluded; `.gitignore` is honored natively by rg/fd; and one generated-dirs list (`node_modules`, `dist`, `build`, `coverage`, `.venv`, and similar) is force-excluded even when a project forgot to gitignore it. `include_ignored: true` lifts the gitignore and generated-dirs layers together. The clio-internal layer always stands, except that pointing a tool directly at one of those directories means the caller wants those paths.
|
|
76
|
+
|
|
77
|
+
## The observation envelope
|
|
78
|
+
|
|
79
|
+
The six content-returning OBSERVE tools (`read`, `grep`, `find`, `ls`, `code_nav`, `context`) close every result through one shared envelope in `src/tools/observation.ts`. `credential_present` sits in the OBSERVE plane but returns a typed boolean and carries no envelope cap. The envelope owns four guarantees.
|
|
80
|
+
|
|
81
|
+
**One notice line, one format.** A truncated text result appends exactly one notice:
|
|
82
|
+
|
|
83
|
+
```text
|
|
84
|
+
[<tool>: <shown>/<total> <unit> shown (<shownSize> of <totalSize>) | full: <offloadPath> | next: <exact-call>]
|
|
85
|
+
```
|
|
86
|
+
|
|
87
|
+
Unknown segments are omitted. `<total>` renders as `N+` when the search was killed early at its limit, meaning matches beyond it exist but were never counted. `next` is always an exact continuation call fragment such as `limit=200` or `offset=451`, never prose. Untruncated results get no notice. Empty results are standardized: `grep` returns `No matches found`, `find` returns `No files found matching pattern`, `ls` returns `(empty directory)`, and the JSON-format tools return valid JSON with empty arrays and `next` populated.
|
|
88
|
+
|
|
89
|
+
**Offload on truncation.** When a byte cap cuts collected content, the tool spills its full rendering to the per-session scratch file (`<stateDir>/scratch/<sessionId>/<toolCallId>.txt`) and reports the path in the notice, so no collected match, path, or line is ever unrecoverable. Two deliberate exceptions exist: `read` never offloads because the source file is directly re-addressable via `next: offset=N`, and a bare item-limit truncation without a byte cut continues via `next` alone, since an offload would only duplicate the body.
|
|
90
|
+
|
|
91
|
+
**Always-valid JSON.** `code_nav` and the JSON scopes of `context` declare `format: "json"`. A JSON payload must parse or be replaced whole; it is never cut mid-document. An oversize payload is offloaded and the body is replaced by the parseable stub:
|
|
92
|
+
|
|
93
|
+
```json
|
|
94
|
+
{"error":"result exceeded <cap>","offloadPath":"...","next":"..."}
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
**One turn budget.** All six envelope tools draw from a single per-turn pool keyed `sessionId:turnId`, default 192KB, overridable with `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES`. Each call reserves the minimum of its self cap and the remaining budget before doing the work. An exhausted pool short-circuits with an `[observation budget exhausted ...]` notice naming the tool, the subject, and the used/limit sizes, instead of paying for a search whose output could not be returned. A call whose cap was reduced by the pool appends a budget note telling the model to narrow its arguments or continue in a follow-up turn.
|
|
98
|
+
|
|
99
|
+
Per-call self caps: `read` 50KB (`CLIO_CODER_READ_MAX_BYTES`), `grep` 16KB for `mode=content` and 8KB for `files`/`count`, `find` 8KB, `ls` 8KB, `code_nav` 16KB, `context` 16KB for docs and 50KB for skills/workspace. The registry backstop cap for each envelope tool is its self cap plus 2KB slack, so a tool's own notice with its exact continuation call survives shaping instead of being cut again and replaced by a generic hint; the bootstrap policy assertion fails loudly if a cap ever drops below that.
|
|
100
|
+
|
|
101
|
+
Every envelope result carries `details.observation` (`{tool, unit, shownCount, totalCount, shownBytes, totalBytes, truncated, format, next?, offloadPath?, budget?}`) for the TUI ledger, session turns, and observers.
|
|
102
|
+
|
|
103
|
+
## Description tiering
|
|
104
|
+
|
|
105
|
+
Tool descriptions are tiered by how much a wrong call costs. The hot tools the model calls constantly (`read`, `grep`, `find`, `dispatch`) embed their operational contract in the description: caps, modes, ignore semantics, and how truncated results continue. Every other tool carries a one-to-two-sentence statement of what it does, and deep usage guidance lives in the bundled docs corpus ([tool-usage.md](tool-usage.md)) rather than the prompt prefix, retrievable on demand through `context(scope="docs")`. This keeps the serialized schema block small and byte-stable while still giving the model a path to depth when it needs one.
|
|
106
|
+
|
|
107
|
+
## The gateway reservation
|
|
108
|
+
|
|
109
|
+
`gateway` is a design-reserved name in `src/core/tool-names.ts`, not an implemented tool. The reserved contract sketch is `gateway(op: "find" | "describe" | "call", capability?, args?)`: an MCP/database proxy with one fixed schema, where external capabilities surface through find/describe/call results rather than as per-capability schemas in the prompt prefix. It would carry the network action class and run sequentially. Reserving the name keeps classifiers and profiles from ever assigning `gateway` to a dynamic tool.
|
|
110
|
+
|
|
111
|
+
## Context protection
|
|
112
|
+
|
|
113
|
+
Clio uses two context-protection mechanisms.
|
|
114
|
+
|
|
115
|
+
1. Tool results are capped at the source and again at the registry boundary. OBSERVE tools use the envelope caps above. Exact mutation tools (`write`, `edit`, `artifact`) use 8KB; `steer` and `credential_present` use 4KB; `ask_user` has a 20KB policy. Summary-kind tools (`bash`, `git`, `verify`, `dispatch`, `monitor`) use 16KB at the registry boundary. `web_fetch` is bounded at 16KB after shaping and may read more before it: its `max_bytes` argument defaults to 600KB and is hard-capped at 5MB. Tools without an explicit result-size policy use an approximately 18KB generic backstop. Over-cap generic results are shown briefly and, when possible, saved under `<stateDir>/scratch/<sessionId>/<toolCallId>.txt` with an `offloadPath` detail and a 10MB scratch-file cap.
|
|
116
|
+
2. Auto-compaction uses one pressure threshold. The default threshold is 0.8. When pressure crosses the threshold, Clio first masks stale tool observations and stale thinking older than `excludeLastTurns`. If pressure remains above the threshold, it runs the LLM summary compaction path and replays from the compacted session view.
|
|
117
|
+
|
|
118
|
+
Manual `/context compact`, `CLIO_CODER_FORCE_COMPACT=1`, and overflow recovery force the LLM summary path directly.
|
|
119
|
+
|
|
120
|
+
Compaction rewrites history, so the next turn on a local single-slot backend is expected to lose prefix-cache alignment. Clio records `expectedColdReasons` and shows one dim notice for that turn.
|
|
121
|
+
|
|
122
|
+
## Inspecting a session
|
|
123
|
+
|
|
124
|
+
Timing and cache behavior are persisted per API call, so a finished session can be inspected from its stored artifacts alone. Each assistant entry in the session ledger (`current.jsonl`, under the directory reported by `clio-coder paths`) carries `timing { ttftMs, apiMs }` and `promptCache { input, cacheRead, cacheWrite, backendVerdict }`, and the run's first persisted call also carries `expectedColdReasons`. Cache verdicts are `hot`, `partial`, `cold`, or `small`.
|
|
125
|
+
|
|
126
|
+
For aggregate cost and token facts across sessions, use `clio-coder usage report --days <n>`. Inside the TUI, `/cost` shows session totals and `/context` opens the context-window ledger overlay.
|
|
127
|
+
|
|
128
|
+
## Self-documentation retrieval
|
|
129
|
+
|
|
130
|
+
`context(scope="docs")` is the model-facing companion to the human `clio-coder docs` server. The server serves bundled `docs/html/**` blueprints for people; the docs scope indexes the bundled markdown corpus for agents. It is deterministic and offline: no embeddings service, network call, or filesystem write is needed.
|
|
131
|
+
|
|
132
|
+
The search index splits markdown into heading-delimited sections, records heading breadcrumbs and line ranges, and ranks results with light stemming, controlled Clio vocabulary aliases, phrase boosts, and BM25-style body scoring. The tool returns compact JSON containing corpus metadata, normalized and expanded query terms, and ranked hits with `file`, `heading`, `breadcrumb`, `anchor`, section `lines`, `snippetLines`, a bounded `snippet`, `matchedTerms`, `signals`, `coverage`, and `score`. `limit` defaults to 5 sections and caps at 12. The per-file filter the pre-consolidation docs tool accepted was dropped; narrow with more specific query terms instead. Even an empty result is valid JSON with empty arrays and a populated `next` continuation.
|
|
133
|
+
|
|
134
|
+
## Edit matching safety
|
|
135
|
+
|
|
136
|
+
The `edit` tool first attempts exact matching. If the model's old text differs
|
|
137
|
+
only by normalized quote, dash, whitespace, or indentation details, Clio maps
|
|
138
|
+
the normalized match back to the original line span and splices only the
|
|
139
|
+
intended replacement. Unchanged spans keep their original bytes, including
|
|
140
|
+
smart punctuation and CRLF line endings. Ambiguous duplicate matches,
|
|
141
|
+
overlapping hunks, empty changes, and no-op edits are rejected instead of
|
|
142
|
+
guessing.
|