@iowarp/clio-coder 0.3.3 → 0.3.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +39 -0
- package/CONTRIBUTING.md +1 -1
- package/README.md +3 -3
- package/dist/{acp-P2AQILE2.js → acp-S5R4RR5B.js} +7 -6
- package/dist/{agents-72W3BI7I.js → agents-P6DMMVZY.js} +24 -21
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-5TWEIYDN.js → auth-2XCZLPKS.js} +12 -8
- package/dist/{chunk-GGXXDWE4.js → chunk-22NAGB7X.js} +2 -2
- package/dist/{chunk-2DJ2KNFG.js → chunk-2LZI5CAG.js} +133 -13
- package/dist/{chunk-OAO4GE4M.js → chunk-2TZWSW76.js} +2 -2
- package/dist/{chunk-OOJYHWRB.js → chunk-34475P3I.js} +2 -2
- package/dist/{chunk-4XUGQOHA.js → chunk-35MKKU5R.js} +4 -4
- package/dist/{chunk-LM5TQCJZ.js → chunk-3HZ5RWN2.js} +5 -5
- package/dist/{chunk-AGYYIBLL.js → chunk-3JLKSKD7.js} +2 -2
- package/dist/{chunk-KZWTDYJF.js → chunk-4JUF2NNX.js} +7 -7
- package/dist/{chunk-ZDOOVTXZ.js → chunk-4OC57DA6.js} +27 -4
- package/dist/chunk-5M54SPOL.js +926 -0
- package/dist/{chunk-STBPMHSX.js → chunk-7RXG6QRZ.js} +51 -11
- package/dist/{chunk-X6IAEBZR.js → chunk-A2GZF7DC.js} +5 -5
- package/dist/{chunk-A3CYT5EX.js → chunk-AD2SYQYC.js} +55 -2
- package/dist/chunk-AOCYTWAV.js +449 -0
- package/dist/chunk-BEY543CS.js +258 -0
- package/dist/{chunk-6N5PTWMY.js → chunk-BP4OYD6A.js} +32 -13
- package/dist/chunk-BPGS2WCQ.js +612 -0
- package/dist/{chunk-V6RTAOC2.js → chunk-BRXQQJFP.js} +8 -8
- package/dist/chunk-CFGTUFWB.js +67 -0
- package/dist/{chunk-CBCAPZAA.js → chunk-E25LMLRW.js} +2 -2
- package/dist/{chunk-DUYJ5IO6.js → chunk-EDRHSCIE.js} +4 -4
- package/dist/{chunk-XBXAASKX.js → chunk-EFADSJET.js} +2 -2
- package/dist/{chunk-M6SHUN7Q.js → chunk-FO5ZOVUY.js} +2 -2
- package/dist/chunk-FYYLNIL5.js +313 -0
- package/dist/{chunk-FNTMWMX5.js → chunk-HV5X7OR2.js} +14 -12
- package/dist/{chunk-TZTZS7QK.js → chunk-HXG4IURW.js} +5 -3
- package/dist/{chunk-5UFT4SUX.js → chunk-K6WL7QZT.js} +3 -3
- package/dist/chunk-K7VKOLQQ.js +15 -0
- package/dist/{chunk-BMEMKKIT.js → chunk-KOHPCX4K.js} +2 -2
- package/dist/{chunk-OC7FIQPC.js → chunk-KRPY7NTG.js} +10 -7
- package/dist/chunk-LL4KHSZI.js +22 -0
- package/dist/{chunk-DSELYM6W.js → chunk-MEQ45TQ4.js} +15 -9
- package/dist/{chunk-5UUP6MWO.js → chunk-MV3K5QF2.js} +5 -436
- package/dist/{chunk-6SGHMWE3.js → chunk-N4CZJQRK.js} +5 -5
- package/dist/{chunk-TZK7PACC.js → chunk-NILBFAPG.js} +14 -8
- package/dist/chunk-OZNBF4L3.js +23 -0
- package/dist/{verify-375KUB3Y.js → chunk-PCZJO5TI.js} +127 -42
- package/dist/chunk-QQK64KLB.js +1360 -0
- package/dist/{chunk-SRDMMSEP.js → chunk-QQL5RT5M.js} +979 -1619
- package/dist/{chunk-LZSJBIVT.js → chunk-QWU7ZBO7.js} +70 -720
- package/dist/{chunk-2TLUCQVG.js → chunk-RD5U66HV.js} +3 -3
- package/dist/{chunk-OKGUZO2U.js → chunk-SPULKLCF.js} +4 -3
- package/dist/{chunk-OQ33BKR3.js → chunk-TTNYS3EA.js} +3 -60
- package/dist/chunk-TW3WDMVS.js +677 -0
- package/dist/chunk-TZSKNMZG.js +434 -0
- package/dist/{chunk-7MNJORFF.js → chunk-UL3WSD3F.js} +6 -1
- package/dist/{chunk-PIWWS5BL.js → chunk-UZHIZC5S.js} +7 -7
- package/dist/{chunk-UFIIWP2H.js → chunk-VAWWTKDP.js} +8 -8
- package/dist/{chunk-COU2UHX6.js → chunk-VEZEGCGW.js} +170 -2
- package/dist/{chunk-LW6DSM3M.js → chunk-VMNQ6OZA.js} +98 -202
- package/dist/chunk-VSNATDE6.js +122 -0
- package/dist/chunk-W6GROXXM.js +69 -0
- package/dist/chunk-WPQLXFOZ.js +375 -0
- package/dist/{chunk-ORBHGJC5.js → chunk-WR67VIZY.js} +3 -3
- package/dist/{chunk-PAJK6MAQ.js → chunk-X6COSD2O.js} +5 -5
- package/dist/chunk-ZGVHUX3M.js +66 -0
- package/dist/{chunk-LWLEKMDQ.js → chunk-ZYKPLLNQ.js} +510 -547
- package/dist/cli/index.js +27 -23
- package/dist/{clio-JOU4FXVA.js → clio-J5JIOIDS.js} +7 -6
- package/dist/{code-nav-7AX6FYE6.js → code-nav-AXCXSBHX.js} +5 -3
- package/dist/{config-XCDVKR23.js → config-OEBMIN2U.js} +37 -27
- package/dist/{configure-4GAP54ZW.js → configure-PUQOSIXQ.js} +16 -13
- package/dist/{context-5VKGUVJJ.js → context-EKDCKUUZ.js} +82 -7
- package/dist/{context-4UOGGLQ5.js → context-MGSE4Z2T.js} +33 -23
- package/dist/{context-77FM5DV5.js → context-URSXPBCK.js} +17 -9
- package/dist/{context-clear-XXJRLCJJ.js → context-clear-KDAJRNUK.js} +33 -23
- package/dist/context-working-set-SBKMPPI2.js +1552 -0
- package/dist/{dispatch-runner-QPRDDBDX.js → dispatch-runner-MSWN72NK.js} +43 -29
- package/dist/{doctor-HR46URBJ.js → doctor-7BSE27PJ.js} +10 -10
- package/dist/{eval-XSSNATB4.js → eval-IZGDOO4H.js} +9 -8
- package/dist/{evidence-6HG2PY2B.js → evidence-SR7WXB5B.js} +51 -23
- package/dist/{evolve-K7YU3NCY.js → evolve-K7VE2CBX.js} +30 -20
- package/dist/{fleet-VY3HHKN6.js → fleet-7XMJNQNF.js} +48 -38
- package/dist/{fleet-preflight-DDN536IT.js → fleet-preflight-AQNAH644.js} +3 -3
- package/dist/{init-JYGXI3FK.js → init-JGNPAYXT.js} +41 -31
- package/dist/{memory-WFZMGYHX.js → memory-4ALKDJ4Q.js} +32 -22
- package/dist/{models-I5QWSEOM.js → models-ZMMLFJNN.js} +22 -19
- package/dist/{monitor-GE4ID3IA.js → monitor-2F3T5KHP.js} +55 -43
- package/dist/{orchestrator-EM5MC3HM.js → orchestrator-ORHT43JB.js} +572 -388
- package/dist/{reset-L2FQEE3E.js → reset-NXGTYNUO.js} +4 -3
- package/dist/{run-ZU3QMZPZ.js → run-RF4WJGMT.js} +51 -41
- package/dist/{share-S5BZQC5I.js → share-UT3W6E4M.js} +5 -4
- package/dist/{skills-X5VXCRNQ.js → skills-PSACKC5Q.js} +2 -2
- package/dist/{skills-eval-WKIHWTHR.js → skills-eval-WJSI55RZ.js} +34 -24
- package/dist/{targets-SNCPI2NR.js → targets-PIIRAOYS.js} +23 -20
- package/dist/{terminal-lease-BNAHVHBS.js → terminal-lease-ULWXWNVY.js} +4 -3
- package/dist/{upgrade-JQHHPQ4K.js → upgrade-346TZ6AV.js} +18 -17
- package/dist/{usage-OR4O5SMZ.js → usage-6KKXR32N.js} +34 -24
- package/dist/verifiers-4UUM6TEE.js +1214 -0
- package/dist/verify-X5HDROLA.js +25 -0
- package/dist/{wiki-generate-UEXP2ARI.js → wiki-generate-7STOCIFZ.js} +42 -31
- package/dist/worker/entry.js +33 -24
- package/docs/README.md +7 -6
- package/docs/acp.md +1 -1
- package/docs/alcf-provider.md +1 -1
- package/docs/architecture.md +2 -2
- package/docs/artifact-versions.md +1 -1
- package/docs/built-in-agents.md +1 -1
- package/docs/capacity-and-scheduling.md +1 -1
- package/docs/commands-and-modes.md +53 -21
- package/docs/config-knobs-audit.md +1 -2
- package/docs/configuration-and-targets.md +15 -1
- package/docs/context-engine.md +64 -12
- package/docs/context-working-set.md +194 -0
- package/docs/development-pipeline.md +1 -1
- package/docs/documentation-coverage.md +5 -5
- package/docs/documentation-guide.md +6 -5
- package/docs/environment-variables.md +2 -1
- package/docs/eval-runner.md +1 -1
- package/docs/evals-internal.md +14 -1
- package/docs/evidence-and-memory.md +74 -2
- package/docs/evolution.md +1 -1
- package/docs/exit-codes-and-output.md +1 -1
- package/docs/extensions-and-sharing.md +2 -2
- package/docs/fleet-dispatch.md +22 -7
- package/docs/glossary.md +21 -1
- package/docs/installation-and-lifecycle.md +2 -2
- package/docs/middleware-and-components.md +1 -1
- package/docs/model-catalog.md +7 -9
- package/docs/observability.md +4 -4
- package/docs/proactive-memory.md +1 -1
- package/docs/prompt-envelope-and-tools.md +4 -4
- package/docs/provider-adapter-cookbook.md +1 -1
- package/docs/release-cut-checklist.md +31 -31
- package/docs/safety-model.md +23 -4
- package/docs/scientific-validation.md +21 -3
- package/docs/session-lifecycle.md +3 -3
- package/docs/skills-marketplace.md +1 -1
- package/docs/tool-usage.md +79 -12
- package/docs/trace-store.md +1 -1
- package/docs/troubleshooting.md +1 -1
- package/docs/tui-design.md +1 -1
- package/docs/worker-dispatch-mechanics.md +11 -1
- package/package.json +8 -11
- package/skills/meta/clio-test/SKILL.md +20 -17
- package/skills/meta/clio-test/evals.md +3 -3
- package/skills/meta/clio-test/references/harness.md +35 -6
- package/skills/meta/clio-test/references/test-map.md +20 -10
- package/skills/registry.yaml +2 -2
- package/skills/skill-marketplace.json +1 -1
- package/src/cli/context-working-set.ts +513 -0
- package/src/cli/context.ts +8 -0
- package/src/cli/evidence.ts +20 -2
- package/src/cli/index.ts +4 -0
- package/src/cli/verifiers.ts +325 -0
- package/src/core/bash-exec.ts +39 -14
- package/src/core/bus-events.ts +19 -4
- package/src/core/config.ts +54 -0
- package/src/core/defaults.ts +50 -3
- package/src/core/verification-scripts.ts +6 -0
- package/src/domains/agents/builtins/verifier.md +3 -0
- package/src/domains/config/classify.ts +1 -0
- package/src/domains/context/working-set/contract.ts +161 -0
- package/src/domains/context/working-set/defaults.ts +28 -0
- package/src/domains/context/working-set/engine.ts +203 -0
- package/src/domains/context/working-set/fold.ts +62 -0
- package/src/domains/context/working-set/horizon.ts +38 -0
- package/src/domains/context/working-set/marker.ts +103 -0
- package/src/domains/context/working-set/path-index.ts +436 -0
- package/src/domains/context/working-set/payload.ts +152 -0
- package/src/domains/context/working-set/policies/age-horizon.ts +55 -0
- package/src/domains/context/working-set/policies/index.ts +21 -0
- package/src/domains/context/working-set/policies/structural.ts +160 -0
- package/src/domains/context/working-set/project.ts +132 -0
- package/src/domains/context/working-set/protect.ts +109 -0
- package/src/domains/context/working-set/recall.ts +177 -0
- package/src/domains/context/working-set/replay/controls.ts +112 -0
- package/src/domains/context/working-set/replay/load-clio.ts +199 -0
- package/src/domains/context/working-set/replay/metrics.ts +185 -0
- package/src/domains/context/working-set/replay/reference-graph.ts +79 -0
- package/src/domains/context/working-set/replay/report.ts +139 -0
- package/src/domains/context/working-set/replay/runner.ts +325 -0
- package/src/domains/context/working-set/replay/synthetic.ts +422 -0
- package/src/domains/context/working-set/replay/trace.ts +21 -0
- package/src/domains/context/working-set/visible.ts +54 -0
- package/src/domains/evidence/build.ts +112 -45
- package/src/domains/evidence/eval.ts +24 -7
- package/src/domains/evidence/index.ts +53 -0
- package/src/domains/evidence/ordering.ts +12 -0
- package/src/domains/evidence/run-trust.ts +221 -0
- package/src/domains/evidence/store.ts +46 -6
- package/src/domains/evidence/trust-status.ts +854 -0
- package/src/domains/evidence/types.ts +26 -0
- package/src/domains/middleware/memory-intervention.ts +3 -0
- package/src/domains/middleware/stalled-turn.ts +165 -4
- package/src/domains/safety/autonomy.ts +1 -1
- package/src/domains/safety/default-path-policy.ts +8 -0
- package/src/domains/safety/finish-contract.ts +4 -3
- package/src/domains/safety/policy-engine.ts +48 -6
- package/src/domains/session/compaction/compact.ts +23 -1
- package/src/domains/session/compaction/cut-point.ts +2 -0
- package/src/domains/session/compaction/tokens.ts +16 -1
- package/src/domains/session/context-ledger.ts +2 -0
- package/src/domains/session/entries.ts +107 -1
- package/src/domains/session/manager.ts +9 -2
- package/src/domains/session/migrations/index.ts +22 -3
- package/src/engine/acp/server.ts +3 -0
- package/src/engine/agent.ts +18 -1
- package/src/engine/session.ts +9 -3
- package/src/entry/orchestrator.ts +16 -4
- package/src/interactive/chat-loop-messages.ts +18 -6
- package/src/interactive/chat-panel.ts +17 -1
- package/src/interactive/chat-renderer.ts +30 -21
- package/src/interactive/context-meter.ts +10 -0
- package/src/interactive/context-overlay.ts +81 -6
- package/src/interactive/context-recall-command.ts +110 -0
- package/src/interactive/interactive-slash-runtime.ts +37 -1
- package/src/interactive/model-session-replay.ts +21 -0
- package/src/interactive/overlay-general-openers.ts +6 -0
- package/src/interactive/overlay-session-lifecycle.ts +8 -4
- package/src/interactive/renderers/tool-execution.ts +18 -2
- package/src/interactive/session-transcript.ts +2 -2
- package/src/interactive/slash-commands.ts +29 -2
- package/src/interactive/turn-context.ts +238 -88
- package/src/interactive/turn-middleware.ts +6 -6
- package/src/tools/agent-tools.ts +11 -4
- package/src/tools/bash.ts +144 -82
- package/src/tools/builtin-tool-catalog.ts +11 -5
- package/src/tools/context/index.ts +105 -3
- package/src/tools/context/surface.ts +3 -2
- package/src/tools/core-bootstrap.ts +21 -0
- package/src/tools/dispatch-runner.ts +9 -7
- package/src/tools/monitor.ts +28 -20
- package/src/tools/registry.ts +59 -7
- package/src/tools/result-disposition.ts +550 -0
- package/src/tools/result-shaping.ts +262 -19
- package/src/tools/safe-exec.ts +2 -0
- package/src/tools/verify/authoring.ts +1119 -0
- package/src/tools/verify/catalog.ts +346 -0
- package/src/tools/verify/index.ts +13 -3
- package/src/tools/verify/scripts.ts +135 -37
- package/src/tools/verify/surface.ts +9 -5
- package/src/tools/worker-evidence.ts +35 -12
- package/dist/chunk-J7CWMCQD.js +0 -255
- package/dist/chunk-T6YILFSB.js +0 -80
- package/dist/chunk-VAKQQHWR.js +0 -434
- package/dist/chunk-VPAYEGVX.js +0 -184
package/docs/context-engine.md
CHANGED
|
@@ -1,11 +1,13 @@
|
|
|
1
1
|
# Context Engine
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.4).
|
|
5
5
|
|
|
6
6
|
Clio Coder tracks context pressure, records per-turn snapshots, and protects the provider context with bounded tool results plus single-threshold compaction.
|
|
7
7
|
|
|
8
|
-
Source of truth lives in `src/domains/session/context-accounting.ts`, `src/domains/session/context-ledger.ts`, `src/domains/session/compaction/`, `src/domains/session/migrations/index.ts`, and the chat-loop integration in `src/interactive/chat-loop.ts`.
|
|
8
|
+
Source of truth lives in `src/domains/session/context-accounting.ts`, `src/domains/session/context-ledger.ts`, `src/domains/session/compaction/`, `src/domains/context/working-set/`, `src/domains/session/migrations/index.ts`, and the chat-loop integration in `src/interactive/chat-loop.ts`.
|
|
9
|
+
|
|
10
|
+
The non-destructive eviction layer has its own guide: [context-working-set.md](context-working-set.md).
|
|
9
11
|
|
|
10
12
|
## Context window resolution
|
|
11
13
|
|
|
@@ -23,7 +25,7 @@ The estimator in `context-accounting.ts` uses a four-characters-per-token family
|
|
|
23
25
|
|
|
24
26
|
At submit time, Clio captures a context snapshot and persists a slim JSONL record under the session directory as `context-snapshots.jsonl`. The slim record keeps token counts, segment metadata, signatures, and hashes, not the heavy prompt or transcript text. When provider usage arrives, `reconcileSnapshot` folds actual input and output counts back into the ledger.
|
|
25
27
|
|
|
26
|
-
Session metadata enforces session format version
|
|
28
|
+
Session metadata enforces session format version 4 (`CURRENT_SESSION_FORMAT_VERSION = 4`). Version 4 is additive: it adds the `contextEviction` and `contextRecall` records and changes no existing entry. A version 3 session therefore migrates to 4 in place when Clio opens it, and no entry is rewritten. Only a session written by a newer build is refused, with an error naming the version it read and pointing at upgrading. The bump is one-way for the operator: a 0.3.3 binary cannot open a session this release wrote.
|
|
27
29
|
|
|
28
30
|
The `/context` overlay and footer meter read the same ledger categories: `system`, `tools`, `agents`, `skills`, `memory`, `project`, `messages`, `pending`, `reserve`, `free`, and `streaming`.
|
|
29
31
|
|
|
@@ -31,31 +33,63 @@ The `/context` overlay and footer meter read the same ledger categories: `system
|
|
|
31
33
|
|
|
32
34
|
Auto-compaction is controlled by one pressure threshold. Pressure is `estimated_tokens / context_window`. The default threshold is `0.8`.
|
|
33
35
|
|
|
34
|
-
|
|
36
|
+
Crossing that threshold engages three mechanisms in a fixed order. The first two are cheap, reversible, and call no model. Only the third rewrites what the session says about itself.
|
|
37
|
+
|
|
38
|
+
### 1. Working-set eviction
|
|
39
|
+
|
|
40
|
+
When `compaction.auto` is enabled and pressure crosses the threshold before a request, Clio applies the configured working-set policy first. The policy selects tool-result bodies and closed-turn thinking blocks, `runAutoCompact` appends one `contextEviction` ledger entry, and `refreshAgentMessagesFromSession` projects those units out of model replay behind a one-line marker. Nothing is deleted: the ledger keeps the original bodies, the transcript keeps showing them, and `/resume`, `/tree`, `/fork`, and the HTML export are unaffected.
|
|
41
|
+
|
|
42
|
+
Already-evicted units are never selected again. Recent turns keep their full observations and thinking, governed by `context.workingSet.protectLastTurns`. Results whose estimated body is below `context.workingSet.minEvictableTokens` (200 tokens by default) are kept whatever their age as a low-yield churn guard. The engine separately refuses any candidate whose marker would save no tokens. The `age-horizon` policy is therefore the selection the old destructive mask made minus those small results, not a byte-identical reproduction of it; the default `structural-v1` policy applies its structural rules before any age rule.
|
|
43
|
+
|
|
44
|
+
If the projection drops pressure below the threshold, Clio sends the request and no summary runs. The policies, the protection predicates, the marker format, and the ledger records are documented in [context-working-set.md](context-working-set.md).
|
|
45
|
+
|
|
46
|
+
### 2. Recall
|
|
47
|
+
|
|
48
|
+
An evicted body comes back on demand and only on demand. The marker names the exact call: `context(scope="recall", ref="<turnId>")` returns the original body byte-exact through the observation envelope and appends a `contextRecall` entry. Operators use `/context recall <ref>`, which prints the body to the transcript without putting it into model context.
|
|
49
|
+
|
|
50
|
+
A recall does not un-evict. The marker stays byte-identical where it was, so the provider prefix cache is untouched, and repeated recalls of the same ref are the churn signal the `/context` overlay reports.
|
|
51
|
+
|
|
52
|
+
Offline replay does not infer those explicit decisions from a later read of the same path. A reread already returns current content, while recall returns a selected historical ref. The replay tables keep the token-weighted `recallTokens` demand bound and reserve recall count, churn, and tail-growth simulation for ledgers or corpora that record which refs were actually recalled.
|
|
53
|
+
|
|
54
|
+
### 3. LLM summary, as a last resort
|
|
35
55
|
|
|
36
|
-
|
|
56
|
+
If pressure remains above the threshold after eviction, Clio runs the summary compaction path: it calls the summarization model, appends a `compactionSummary` entry, refreshes projected replay messages from the session, and continues. This is the only mechanism that spends tokens and the only one whose output is a lossy paraphrase, which is why it runs last.
|
|
57
|
+
|
|
58
|
+
Iterative compaction has one raw-history boundary. The first pass searches from the start of the active path. A later pass searches strictly after the previous `compactionSummary`, while the prior checkpoint and its retained suffix are fed to the summarizer as canonical context for one cumulative replacement. The replay benchmark uses that same boundary. It must not run `findCutPoint` over the visible retained suffix again: doing so re-prices history already captured by the previous checkpoint and overstates repeated summary churn.
|
|
59
|
+
|
|
60
|
+
Manual `/context compact`, `CLIO_CODER_FORCE_COMPACT=1`, and overflow recovery force the summary path directly and skip every pre-stage. The overflow guard runs before the user turn is committed, so a blocked oversized request does not leave an unanswered user entry in the ledger.
|
|
61
|
+
|
|
62
|
+
### The legacy mask escape hatch
|
|
63
|
+
|
|
64
|
+
`CLIO_CODER_LEGACY_MASK=1` restores the destructive pre-stage working-set eviction replaced. It calls `session.replaceEntries` and rewrites the persisted bodies, so masked content is gone for the operator as well as the model. It uses the old marker format:
|
|
37
65
|
|
|
38
66
|
```text
|
|
39
67
|
[Observation masked: <tool> output was <lines> lines, <chars> chars - contents masked to save context. Re-run the tool for current content.] Preview: <preview>
|
|
40
68
|
```
|
|
41
69
|
|
|
42
|
-
|
|
70
|
+
It exists for one release as a compatibility diagnosis path and is removed in the next.
|
|
43
71
|
|
|
44
|
-
|
|
72
|
+
### Replay text
|
|
45
73
|
|
|
46
|
-
|
|
74
|
+
When the ledger is replayed to the model, compaction summaries, branch summaries, and bash executions become standardized user-role message text. Clio imports `COMPACTION_SUMMARY_PREFIX`, `BRANCH_SUMMARY_PREFIX`, their suffixes, and `bashExecutionToText` through `src/engine/messages.ts`; `src/interactive/chat-renderer.ts` maps Clio's entry shapes onto them and applies replay truncation. The working-set projection runs before that builder, so markers are what the replay text is built from.
|
|
47
75
|
|
|
48
76
|
## Cache-divergence honesty
|
|
49
77
|
|
|
50
|
-
|
|
78
|
+
Every provider Clio targets caches by exact prefix. Anthropic hashes the cumulative prefix up to a `cache_control` breakpoint and looks back at most 20 blocks for an earlier write; the minimum cacheable prefix is 512 to 4,096 tokens by model, reads cost 0.1x input and writes 1.25x. OpenAI caches automatically from 1,024 tokens in 128-token increments on exact prefix matches at 0.1x. vLLM hashes each KV block from its parent block's hash, so a change in one block invalidates every later block. llama.cpp (and LM Studio on top of it) picks the slot with the longest common prefix and re-evaluates only the suffix, and `--cache-reuse` can shift later KV chunks back into place after a mid-prompt removal. The consequence is the same everywhere except on llama.cpp with cache reuse: whatever bytes change, everything after the earliest changed position is re-prefilled. That is why a marker is byte-stable, why a recall rides the tail instead of restoring the body in place, why `structural-v1` batches evictions down to `target` instead of trimming on every turn, and why the replay tables report cold prefix tokens per event next to tokens evicted: at a 32k budget one event re-prefills most of the window whichever policy chose the items, so the lever that protects a cloud cache is the number of events, not their contents. A local backend with cache reuse pays less for the same removal, which is where finer-grained eviction and recall earn their keep.
|
|
79
|
+
|
|
80
|
+
The procedural replay target sweep measured 0.4, 0.5, 0.6, and an exhaustive rung-6 stop over 24 traces. Target 0.4 and exhaustive selection converged because un-evictable residue exhausted the candidate pool. Against 0.6, target 0.4 cut cold-prefix tokens by 2.8% at 64k and 7.3% at 128k, with no summary reduction and a 0.00072 reduction in retention covered at 128k. That is below the 10% cache-saving threshold set for changing a cross-tier default, so the default remains 0.6. The complete sweep and reopening rule are in the replay README.
|
|
51
81
|
|
|
52
|
-
|
|
82
|
+
Compaction and eviction both change the replayed history. On a local backend with a single prefix-cache slot, the next turn after either one is expected to be cold because the byte prefix moved. Dispatch traffic can disturb the same slot.
|
|
83
|
+
|
|
84
|
+
Clio records these disturbances once on the next assistant entry as `promptCache.expectedColdReasons`. The recorded reasons are `working_set_evict` for an applied eviction event, `compaction` for the summary path, and `dispatch` for interleaved worker traffic. `compaction` and `dispatch` are stamped only on `local-native` targets, because a single-slot local cache is the one an interleaved run actually disturbs. `working_set_evict` is stamped on every tier: the eviction moved the byte prefix itself, so the cloud prefix cache is cold for the same reason. The user sees one dim notice, and the same reasons persist on that entry in the session ledger next to the per-call cache data.
|
|
53
85
|
|
|
54
86
|
Per-call cache verdicts are `hot`, `partial`, `cold`, and `small`. They are derived from provider usage and persisted with `timing { ttftMs, apiMs }` and `promptCache { input, cacheRead, cacheWrite, backendVerdict }` when available.
|
|
55
87
|
|
|
88
|
+
The `/context` overlay closes the loop. When the last settled run came back `cold` and Clio had recorded a reason for it, the overlay adds a line naming that reason, for example `last cold turn: working-set eviction (expected)`, and reports the cache line without the warning token. A reused prompt shell with a cold backend and no recorded reason stays a warning: Clio kept the bytes stable and the provider re-prefilled anyway, which is a disagreement worth surfacing.
|
|
89
|
+
|
|
56
90
|
## Settings
|
|
57
91
|
|
|
58
|
-
The public settings
|
|
92
|
+
The public settings use one compaction threshold plus a non-destructive working-set stage:
|
|
59
93
|
|
|
60
94
|
```yaml
|
|
61
95
|
compaction:
|
|
@@ -64,9 +98,27 @@ compaction:
|
|
|
64
98
|
excludeLastTurns: 6
|
|
65
99
|
# model: provider/summary-model-id
|
|
66
100
|
# systemPrompt: ~/.config/clio-coder/prompts/compaction.md
|
|
101
|
+
|
|
102
|
+
context:
|
|
103
|
+
workingSet:
|
|
104
|
+
enabled: true
|
|
105
|
+
policy: structural-v1
|
|
106
|
+
target: 0.6
|
|
107
|
+
protectLastTurns: 6
|
|
108
|
+
minEvictableTokens: 200
|
|
67
109
|
```
|
|
68
110
|
|
|
69
|
-
`auto` controls the pre-request trigger. Manual `/context compact` still runs when `auto` is false. `model` optionally selects a dedicated summarization model
|
|
111
|
+
`compaction.auto` controls the pre-request trigger. Manual `/context compact` still runs when `auto` is false. `compaction.model` optionally selects a dedicated summarization model, and `compaction.systemPrompt` optionally points at a prompt override file. `compaction.excludeLastTurns` only governs the temporary legacy mask path; working-set protection uses `context.workingSet.protectLastTurns`.
|
|
112
|
+
|
|
113
|
+
| Key | Default | Accepted | Meaning |
|
|
114
|
+
| --- | --- | --- | --- |
|
|
115
|
+
| `context.workingSet.enabled` | `true` | boolean | `false` skips eviction and goes directly to summary compaction. It does not restore the destructive mask. |
|
|
116
|
+
| `context.workingSet.policy` | `structural-v1` | `age-horizon`, `structural-v1` | Candidate selection rule set. `age-horizon` is the pre-layer age selection. |
|
|
117
|
+
| `context.workingSet.target` | `0.6` | number greater than 0 and less than 1 | Used-over-window ratio an applied eviction event batches down to. |
|
|
118
|
+
| `context.workingSet.protectLastTurns` | `6` | integer ≥ 1 | Recent turns whose observations and thinking are never evicted. |
|
|
119
|
+
| `context.workingSet.minEvictableTokens` | `200` | integer ≥ 0 | Results below this body estimate are never evicted. The default is a measured low-yield churn guard; marker break-even is enforced separately. |
|
|
120
|
+
|
|
121
|
+
Set `CLIO_CODER_LEGACY_MASK=1` only as a temporary compatibility escape hatch for the old destructive mask stage. See [context-working-set.md](context-working-set.md) for what each policy selects and why.
|
|
70
122
|
|
|
71
123
|
Settings validation is strict: an older file still carrying the removed `compaction.thresholds` block fails to load with the exact key path during normal startup. Edit removed or unknown keys deliberately; `clio-coder doctor --fix` does not transform settings into the current schema.
|
|
72
124
|
|
|
@@ -0,0 +1,194 @@
|
|
|
1
|
+
# Working Set
|
|
2
|
+
|
|
3
|
+
The working set is the part of the session ledger the model actually receives on the next request. When context pressure crosses `compaction.threshold`, Clio narrows that view before it considers summarizing anything: selected tool-result bodies and closed-turn thinking blocks stop being replayed, and a one-line marker takes each body's place. Nothing is deleted. The ledger keeps every byte the tools produced, the transcript keeps showing them, and the model can ask for any evicted body back by ref.
|
|
4
|
+
|
|
5
|
+
Source of truth is `src/domains/context/working-set/` (`contract.ts`, `fold.ts`, `project.ts`, `marker.ts`, `protect.ts`, `engine.ts`, `recall.ts`, `policies/`), the ledger records in `src/domains/session/entries.ts`, and the compaction stage in `src/interactive/turn-context.ts` (`runAutoCompact`).
|
|
6
|
+
|
|
7
|
+
> [!WARNING]
|
|
8
|
+
> This is an experimental community alpha surface. The default policy is `structural-v1`, chosen from the replay tables under `benchmarks/results/context-replay/`. `age-horizon` reproduces the selection Clio made before this layer existed and stays available.
|
|
9
|
+
|
|
10
|
+
## Vocabulary
|
|
11
|
+
|
|
12
|
+
| Term | Definition |
|
|
13
|
+
| --- | --- |
|
|
14
|
+
| Working set | What the model sees on the next request: the ledger with the current projection applied. It is never a file. |
|
|
15
|
+
| Ledger | The durable append-only session record (`current.jsonl`). The working-set layer appends to it and never rewrites it. |
|
|
16
|
+
| Evicted | A unit whose body the projection replaces with a marker. The ledger entry that holds the original body is untouched. |
|
|
17
|
+
| Offloaded | A result the observation envelope already wrote to a file because it exceeded the per-call cap. Its marker carries the pointer instead of a preview, and recall returns the pointer rather than inlining the file. |
|
|
18
|
+
| Recall | Readmitting an evicted body by ref, through `context(scope="recall", ref=...)` for the model or `/context recall <ref>` for the operator. |
|
|
19
|
+
| Marker | The byte-stable one-line stub the projection renders in place of an evicted body. It names the ref, the reason, the size, and the exact call that brings the body back. |
|
|
20
|
+
| Projection | A pure, in-memory transform from ledger entries to the entries the replay builder hands the model. `projectWorkingSet(entries, view)` is that function. |
|
|
21
|
+
|
|
22
|
+
## Eviction is a projection, not a rewrite
|
|
23
|
+
|
|
24
|
+
The stage this layer replaces rewrote history. `maskStaleObservations` walked the entries, replaced observation bodies with a masked-out string, and called `session.replaceEntries`. That destroyed the only copy: after a mask, `/resume`, `/tree`, `/fork`, and the HTML export all showed the placeholder, and the content was gone for the operator as well as the model.
|
|
25
|
+
|
|
26
|
+
The working-set layer separates the two audiences. What leaves is recorded as a `contextEviction` entry, appended like any other. `refreshAgentMessagesFromSession` folds those entries into a `WorkingSetView` and applies `projectWorkingSet` before `buildReplayAgentMessagesFromTurns` runs, so only the messages bound for the provider carry markers. Every reader that shows the session to a human reads the raw ledger and sees the full bodies.
|
|
27
|
+
|
|
28
|
+
Three properties follow from that shape:
|
|
29
|
+
|
|
30
|
+
- **Idempotence.** Projecting an already-projected slice reproduces it byte for byte, because the marker comes from the ledger entry rather than from the body being replaced.
|
|
31
|
+
- **Branch safety.** The fold runs through `filterEntriesToActivePath` (issue #94), so an eviction recorded on a branch `/tree` later abandoned cannot project onto the live one, and a fork inherits the view of its shared prefix.
|
|
32
|
+
- **Determinism.** A policy is a pure function of `PolicyInput`. The same ledger and the same settings select the same units in a live session and in an offline replay of that session.
|
|
33
|
+
|
|
34
|
+
Usage anchors recorded before an eviction described a longer prompt than the model will now receive, so the projection stamps `contextUsageInvalidated` on assistant entries that precede the newest eviction event. Without that, `calculateContextTokens` would keep reporting the pre-eviction size and the pressure estimator would never see the space the event freed. The stamp is replay bookkeeping, not provider-visible message content, and the prompt estimator deliberately excludes it. That keeps plan-time `tokensAfter` and the post-event cold-prefix metric on the same byte definition.
|
|
35
|
+
|
|
36
|
+
## Ledger records and format v4
|
|
37
|
+
|
|
38
|
+
Two entry kinds carry the layer, both defined in `src/domains/session/entries.ts`:
|
|
39
|
+
|
|
40
|
+
| Kind | Fields | Meaning |
|
|
41
|
+
| --- | --- | --- |
|
|
42
|
+
| `contextEviction` | `policyId`, `trigger` (`pressure` or `operator`), `evicted[]`, `tokensBefore`, `tokensAfter`, `pressureBefore`, `snapshotIdBefore` | One applied event. Each `evicted[]` item is `{ ref, reason, tokensFreed, marker, by? }`. |
|
|
43
|
+
| `contextRecall` | `ref`, `trigger` (`tool` or `operator`), `tokensReadmitted`, `toolCallId?` | One readmission of one ref. It is a churn record, not an un-eviction. |
|
|
44
|
+
|
|
45
|
+
`reason` is one of `superseded_read`, `stale_after_mutation`, `listing_consumed`, `failure_resolved`, `thinking_turn_closed`, `age_horizon`, `operator`. A `ref` is the `turnId` of the ledger entry that holds the unit: for a `tool_result` message the unit is the result body, and for an `assistant` message it is every thinking block the message carries. Per-block eviction is deliberately not modelled.
|
|
46
|
+
|
|
47
|
+
Adding those kinds bumps the session format to version 4 (`CURRENT_SESSION_FORMAT_VERSION = 4` in `src/engine/session.ts`). The bump is additive: no existing entry kind changes shape, so a version 3 session migrates to 4 in place when Clio opens it and no entry is rewritten. `runMigrations` refuses only what it cannot read, a session written by a newer build, with "upgrade clio-coder to resume this session". The bump is still one-way for the operator: Clio 0.3.3 does not know these kinds and cannot open a session this release wrote.
|
|
48
|
+
|
|
49
|
+
## The marker contract
|
|
50
|
+
|
|
51
|
+
A marker is one line, its fields are in fixed order, and it carries no timestamp and no counter. That is not cosmetic. The marker is persisted inside the `contextEviction` entry and replayed on every subsequent request, so a marker whose bytes drifted between renders would cold-start the provider prefix cache on a turn that evicted nothing new. It would also make two replays of the same recorded ledger disagree.
|
|
52
|
+
|
|
53
|
+
Field order is `ref`, `reason`, `by`, `tool`, `path`, `size`, `offload`, `recall`, then the body tail. Undefined fields are omitted rather than rendered empty. `path` is the one file the result was about: `details.paths` when the tool recorded exactly one (`edit`, `write`, `artifact`), otherwise the `path` argument of the call as the model wrote it, which is how a `read` marker names its file. Real output from `renderMarker` in `src/domains/context/working-set/marker.ts`:
|
|
54
|
+
|
|
55
|
+
```text
|
|
56
|
+
[evicted ref=0198f3c2-7a10-7c31-9d44-2b0c5f1e88a3 reason=stale_after_mutation by=0198f3c2-9b02-7f55-8e10-6d21ac9e4471 tool=read path=src/domains/context/working-set/engine.ts size=41 lines/3.8KB recall=context(scope="recall", ref="0198f3c2-7a10-7c31-9d44-2b0c5f1e88a3") preview="export function planEviction(policy: WorkingSetPolicy, input: PolicyInput): EvictionPlan | null { export function planEv"]
|
|
57
|
+
```
|
|
58
|
+
|
|
59
|
+
```text
|
|
60
|
+
[evicted ref=0198f3c2-1d44-7a90-b201-77c0e1a2f5de reason=failure_resolved by=0198f3c3-0002-7ab1-9c33-14ff90bb2c07 tool=bash size=4 lines/152B recall=context(scope="recall", ref="0198f3c2-1d44-7a90-b201-77c0e1a2f5de") first_line="src/interactive/turn-context.ts(466,15): error TS2345: Argument of type 'PolicyInput' is not assignable to parameter of "]
|
|
61
|
+
```
|
|
62
|
+
|
|
63
|
+
```text
|
|
64
|
+
[evicted ref=0198f3c4-55aa-7be2-8f01-9a3d6c2b1e77 reason=listing_consumed tool=grep size=1 lines/234.4KB offload=/home/dev/.local/state/clio-coder/offload/0198f3c4-grep.txt recall=context(scope="recall", ref="0198f3c4-55aa-7be2-8f01-9a3d6c2b1e77")]
|
|
65
|
+
```
|
|
66
|
+
|
|
67
|
+
Three rules govern the tail. Most reasons render `preview`: the first 120 characters of the body, whitespace collapsed and double quotes escaped, so the preview cannot break the quoted field or spill onto a second line. A `failure_resolved` eviction renders `first_line` instead, because the line that says what failed is worth the marker's tokens where a preview of a stack trace is not. An offloaded body renders neither, because the `offload=` pointer already promises the full artifact at a stable path and a preview would spend tokens repeating it.
|
|
68
|
+
|
|
69
|
+
Thinking eviction renders no marker at all. The reasoning simply stops being replayed. A marker there would spend tokens announcing that something the model cannot act on is gone.
|
|
70
|
+
|
|
71
|
+
## Policies
|
|
72
|
+
|
|
73
|
+
A policy answers one question: which units should leave. It never writes, never reads a clock, and never calls a model. `planEviction` then materializes the selection into `EvictedItem`s with markers rendered and tokens measured, and prices the result against the projection the model will actually receive.
|
|
74
|
+
|
|
75
|
+
### Protection predicates
|
|
76
|
+
|
|
77
|
+
`protect.ts` runs before every rule in `structural-v1` and is absolute. A policy is allowed to be wrong about relevance; it is not allowed to drop these:
|
|
78
|
+
|
|
79
|
+
1. Anything that is not a `tool_result` or `assistant` message. Operator words, compaction and branch summaries, skill activations, task ledgers, worker runs, and bash executions are the session's record of itself.
|
|
80
|
+
2. Anything inside the recent window, which starts at `protectionCutoffIndex(entries, protectLastTurns)`. A turn starts at a user message, a `bashExecution`, or a `branchSummary`.
|
|
81
|
+
3. A result whose estimated body is below `minEvictableTokens`. This protects low-yield bodies from churn; the engine independently rejects a marker that would free no tokens.
|
|
82
|
+
4. A body the legacy destructive stage already replaced, which has nothing left to evict.
|
|
83
|
+
5. A call the safety rails blocked. A refused call is a decision the session made, not an observation it can re-fetch.
|
|
84
|
+
6. A write or edit the turn in flight is still standing on.
|
|
85
|
+
7. A failure nothing later resolved, and any unindexed failure, because without an observation there is no way to ask whether it was resolved.
|
|
86
|
+
|
|
87
|
+
### `age-horizon`
|
|
88
|
+
|
|
89
|
+
The rule `maskStaleObservations` applied, recorded instead of destroyed. Every `tool_result` body older than the protection horizon leaves the working set, and every `assistant` message older than the horizon loses its thinking blocks. Same turn-start definition, same cutoff, and a body carrying a legacy compaction marker is skipped the same way.
|
|
90
|
+
|
|
91
|
+
One skip condition is new, so this is today's selection minus small results rather than a byte-identical reproduction of it: a result whose estimated body is below `minEvictableTokens` (200 tokens by default) stays, whatever its age. The engine already rejects markers that save no tokens; the higher default is a measured low-yield churn guard. The old mask had no such floor and masked those results too. Thinking has no size floor either way, because dropping it renders no marker.
|
|
92
|
+
|
|
93
|
+
`age-horizon` has no target stop. It evicts everything beyond the horizon in one event, exactly as the mask did, and ignores `context.workingSet.target`; the replay tables show this as `saturated events = 1.000` on every row. That is deliberate: the policy exists to reproduce the old selection through the ledger, and an operator who wants batching to a target wants `structural-v1`. Candidates arrive newest-safe-first, so a caller that stops early has evicted the newest safe unit rather than the oldest one.
|
|
94
|
+
|
|
95
|
+
Age is not a quality signal. A file read twenty turns ago and never touched since is more useful than a directory listing from two turns ago, which is the whole reason `structural-v1` exists and is the default.
|
|
96
|
+
|
|
97
|
+
### `structural-v1` (default)
|
|
98
|
+
|
|
99
|
+
Rule order is the policy. Each rung emits candidates newest-first, every candidate passes `isProtected`, and no unit is claimed twice, so a read that is both stale and superseded is evicted for the reason that came first and carries the `by` ref that explains it. The rungs, in order:
|
|
100
|
+
|
|
101
|
+
| # | Reason | Fires when |
|
|
102
|
+
| --- | --- | --- |
|
|
103
|
+
| 1 | `stale_after_mutation` | A read-class observation is followed by a write or edit of the same file. Whatever the body said is now a claim about a file that no longer exists in that form. |
|
|
104
|
+
| 2 | `superseded_read` | A later successful read of the same file covers this one's lines. A full read covers everything; any other read covers only an identical or containing range, and an unknown range covers nothing. |
|
|
105
|
+
| 3 | `failure_resolved` | A later call succeeded with byte-identical arguments, or, for `read`, `grep`, and `find`, reached the same file by any route. |
|
|
106
|
+
| 4 | `listing_consumed` | Every path the listing surfaced went on to be read. One surfaced path still unread and the listing stays, because that is the path the agent comes back to. |
|
|
107
|
+
| 5 | `thinking_turn_closed` | An assistant message beyond the protection horizon carries thinking blocks. |
|
|
108
|
+
| 6 | `age_horizon` | Only under pressure, and only until the projection reaches `target`. |
|
|
109
|
+
|
|
110
|
+
Rungs 1 through 5 are unconditional: redundant content is free to drop, whatever the pressure. Rung 6 is the only one that looks at token counts, and it stops the moment the projected size reaches `context.workingSet.target × contextWindow`. Newest-first within a rung is a cost decision: evicting the youngest safe unit keeps the cold region after the eviction point small, so the turn that pays for the event pays least.
|
|
111
|
+
|
|
112
|
+
The long-trace sweep found that targets 0.4 and an exhaustive rung 6 produced identical results because the usable candidate pool ran out first. Relative to the 0.6 default, 0.4 reduced cold-prefix tokens by 2.8% at 64k and 7.3% at 128k, did not reduce summaries, and lowered retention covered by 0.00072 at 128k. The default therefore remains 0.6. The replay README records the full grid and the numeric reopening rule.
|
|
113
|
+
|
|
114
|
+
The facts the rungs read come from `path-index.ts`, one deterministic pass over the active-path entries producing one observation per tool result that names a path: which file, which line range, which paths a listing surfaced, whether the call failed, and where in the turn sequence it sits. Tools that observe no path (dispatch, web fetch, tasks, ask user, context) produce no observation. There are no content fingerprints.
|
|
115
|
+
|
|
116
|
+
## Recall
|
|
117
|
+
|
|
118
|
+
Recall is explicit and by ref. There is no auto-readmission: the marker tells the model exactly which call brings the body back, and the model decides.
|
|
119
|
+
|
|
120
|
+
`resolveRecall(entries, view, ref, activeLeafTurnId)` resolves a ref against the fold at the live leaf and returns the original body byte-exact, read with the same field precedence the projection would have used. It fails in three typed ways, and each message names the nearest valid ref when one exists:
|
|
121
|
+
|
|
122
|
+
- `invalid_ref` when the ref is empty or carries whitespace.
|
|
123
|
+
- `not_on_active_path` when the session has no such turn on this branch, which includes a ref from a branch `/tree` abandoned.
|
|
124
|
+
- `not_evicted` when the unit is still in context. An assistant turn reports separately that thinking is not recallable.
|
|
125
|
+
|
|
126
|
+
Both messages end with the refs that can be recalled on the active path (tool results only, up to eight, then a count), because a failed recall is usually a mistyped ref and the listing is what the next call needs.
|
|
127
|
+
|
|
128
|
+
An LLM summary also preserves recall discovery across its cut. When an evicted tool result falls before `firstKeptTurnId`, the generated checkpoint carries a `<recallable-refs>` block with the same `ref (tool path)` rows used by recall failures, bounded to eight rows plus a remaining count. Results that stay after the cut keep their ordinary markers and are not repeated in the block.
|
|
129
|
+
|
|
130
|
+
**A recall does not un-evict.** The key stays in `view.evicted`, the marker stays byte-identical at its original position, and the recalled body arrives at the tail of the working set inside the recall result. Readmitting it in place would duplicate the bytes and invalidate the provider prefix cache for everything after that point, which costs more than the recall saved.
|
|
131
|
+
|
|
132
|
+
That also makes recall the churn signal. `churn = recalls / itemsEvicted` over the active path. A high churn number means the policy keeps evicting content the session still needs, which is a reason to change the policy rather than to raise the threshold.
|
|
133
|
+
|
|
134
|
+
The procedural replay does not synthesize churn from path reuse. Its reference graph maps each earlier observation to every later reread or discovery of the same path, while a real `contextRecall` is an explicit model choice of one ref. A later reread already returns current content at the tail, so also injecting the old body would duplicate data and misread stale or superseded observations as recall demand. Replay reports `recallTokens` as a one-time demand bound per evicted item and waits for explicit `contextRecall` records before reporting recall count, churn, or tail growth. The graph-density measurements and reopening condition are in the replay README.
|
|
135
|
+
|
|
136
|
+
An offloaded result returns its pointer, never the file. The model gets the same `full: <path>` promise the original tool result ended with and reads it with `read` when it wants it.
|
|
137
|
+
|
|
138
|
+
The two entry points differ in where the body lands:
|
|
139
|
+
|
|
140
|
+
| Caller | Entry point | Where the body goes | Ledger record |
|
|
141
|
+
| --- | --- | --- | --- |
|
|
142
|
+
| Model | `context(scope="recall", ref=...)` | Back into the working set through the normal observation envelope, so the per-turn pool and the self cap still apply | `contextRecall` with `trigger: "tool"` and the `toolCallId` |
|
|
143
|
+
| Operator | `/context recall <ref>` | The transcript only. It is never submitted as a turn and never counted against the context window | `contextRecall` with `trigger: "operator"` |
|
|
144
|
+
|
|
145
|
+
Both publish `BusChannels.ContextRecalled`, and both route through the middleware `on_compaction` hook as stage `working_set_recall`.
|
|
146
|
+
|
|
147
|
+
## Settings
|
|
148
|
+
|
|
149
|
+
```yaml
|
|
150
|
+
context:
|
|
151
|
+
workingSet:
|
|
152
|
+
enabled: true
|
|
153
|
+
policy: structural-v1
|
|
154
|
+
target: 0.6
|
|
155
|
+
protectLastTurns: 6
|
|
156
|
+
minEvictableTokens: 200
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
| Key | Default | Accepted | Meaning |
|
|
160
|
+
| --- | --- | --- | --- |
|
|
161
|
+
| `context.workingSet.enabled` | `true` | boolean | Master switch. `false` skips eviction and goes straight to summary compaction. It does not restore the destructive mask. |
|
|
162
|
+
| `context.workingSet.policy` | `structural-v1` | `age-horizon`, `structural-v1` | Candidate selection rule set. |
|
|
163
|
+
| `context.workingSet.target` | `0.6` | number greater than 0 and less than 1 | Used-over-window ratio an applied `structural-v1` event batches down to. `age-horizon` ignores it. |
|
|
164
|
+
| `context.workingSet.protectLastTurns` | `6` | integer ≥ 1 | Recent turns whose observations and thinking are never evicted. |
|
|
165
|
+
| `context.workingSet.minEvictableTokens` | `200` | integer ≥ 0 | Results below this body estimate are never evicted. The default protects low-yield bodies; marker break-even is enforced separately. |
|
|
166
|
+
|
|
167
|
+
`compaction.excludeLastTurns` governs only the temporary legacy mask path; working-set protection uses `protectLastTurns`. Settings validation is strict, so an unknown key under this block fails startup with its exact path.
|
|
168
|
+
|
|
169
|
+
`CLIO_CODER_LEGACY_MASK=1` restores the destructive stale-observation stage for one release as a compatibility escape hatch. It rewrites the ledger, and it is removed in the next release.
|
|
170
|
+
|
|
171
|
+
## What the operator sees
|
|
172
|
+
|
|
173
|
+
- **`/context` overlay.** A working-set section under the category legend: the policy that produced the most recent event, evicted item count, evicted tokens, event count, recall count, and churn. Evicted tokens render as one line after the legend rather than as a meter category, because they are outside the window rather than a slice of it.
|
|
174
|
+
- **Transcript.** An evicted tool row keeps its full body and gains a dim `evicted · <reason>` tag. The transcript shows the ledger, never the projection, so `/resume`, `/tree`, `/fork`, and the HTML export are unaffected by eviction.
|
|
175
|
+
- **`/context recall <ref>`.** Prints the ref, why it was evicted, the token count, and the offload pointer when there is one, followed by the original body. Transcript only.
|
|
176
|
+
- **Prompt cache line.** Every applied event stamps `working_set_evict` on the next assistant entry's `promptCache.expectedColdReasons`. When the last settled run came back cold for that reason, the overlay adds `last cold turn: working-set eviction (expected)` and drops the shell-reused-but-backend-cold warning, because the cold turn is explained rather than surprising.
|
|
177
|
+
- **Notice.** One line per applied event: `[context engine] working set: N items evicted by <policy>; ~X -> ~Y tokens, recall by ref with context(scope="recall")`. The numbers are the plan's, priced over the visible ledger slice, and they are the same numbers the `contextEviction` entry, the `[Compaction] Reclaimed context` toast, and the overlay's `last compaction` line carry. The footer meter is a separate live estimate over the agent message list and can differ from them by the tool schemas and replay text it includes.
|
|
178
|
+
|
|
179
|
+
## Not in this release
|
|
180
|
+
|
|
181
|
+
These are tracked follow-ups, not available behavior:
|
|
182
|
+
|
|
183
|
+
- **Auto-readmission.** Nothing brings an evicted body back on its own. There are no path fingerprints and no registry of what the model is likely to need next.
|
|
184
|
+
- **Cost model and deferred scheduling.** Pressure is the only trigger. There is no break-even horizon, no deferred eviction plan, and no piggybacking beyond the fact that the working-set stage already runs first inside `runAutoCompact`.
|
|
185
|
+
- **Intra-turn eviction.** Eviction runs before a request is sent. A single turn whose tool results overflow the window is handled by the observation envelope's caps and by summary compaction, not by this layer.
|
|
186
|
+
- **Worker runtimes.** Dispatched workers replay their own ledgers without the working-set stage.
|
|
187
|
+
- **Digests.** A marker carries tool, size, and a first-line preview. The generated summaries from #165 are not embedded in it.
|
|
188
|
+
|
|
189
|
+
## See also
|
|
190
|
+
|
|
191
|
+
- `clio-coder context replay --sessions <path>...` replays Clio ledgers, and `--synthetic <ids>` replays the seeded procedural corpora, through the same fold, projection, and policy code with `none`, `random`, and `oracle` controls; `clio-coder context working-set --session <id|path>` prints one session's fold and path index. Both are described under [Working-set replay](commands-and-modes.md#working-set-replay), and the committed tables with the default-policy rule are under `benchmarks/results/context-replay/`.
|
|
192
|
+
- [context-engine.md](context-engine.md) for context window resolution, token accounting, and how this stage sits ahead of summary compaction.
|
|
193
|
+
- [session-lifecycle.md](session-lifecycle.md) for the ledger format, active-path lineage, and branching.
|
|
194
|
+
- [glossary.md](glossary.md) for the one-line definitions of these terms.
|
|
@@ -80,7 +80,7 @@ New `area:*` labels are proposed in an issue, not created ad hoc.
|
|
|
80
80
|
|
|
81
81
|
## Milestones are releases
|
|
82
82
|
|
|
83
|
-
Each open milestone is the next version (`v0.3.
|
|
83
|
+
Each open milestone is the next version (`v0.3.4`, `v0.4.0`). Triage means
|
|
84
84
|
assigning an issue to a milestone or explicitly leaving it in the backlog.
|
|
85
85
|
A release cut requires every issue in its milestone to be closed
|
|
86
86
|
or bumped; the milestone closes when the tag is published.
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Clio Coder Documentation Coverage Matrix
|
|
2
2
|
|
|
3
|
-
This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.
|
|
3
|
+
This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.4`.
|
|
4
4
|
|
|
5
5
|
## Coverage Matrix
|
|
6
6
|
|
|
@@ -18,7 +18,7 @@ This matrix maps every top-level directory in `src/` and every domain directory
|
|
|
18
18
|
| `src/domains/agents/` | 12 built-in recipes, agent catalog, recipe schema, fleet commands, fleet contract v4 | [built-in-agents.md](built-in-agents.md), [fleet-dispatch.md](fleet-dispatch.md) | `documented` | Documented in built-in agent recipes guide and fleet dispatch architecture. |
|
|
19
19
|
| `src/domains/components/` | Component scanning, snapshots, hashing, diffing | [middleware-and-components.md](middleware-and-components.md) | `documented` | Documented in active component snapshot and middleware guide. |
|
|
20
20
|
| `src/domains/config/` | Configuration contracts, file watcher, keybinding definitions, setting classifiers | [configuration-and-targets.md](configuration-and-targets.md), [commands-and-modes.md](commands-and-modes.md) | `documented` | Documented in configuration targets and command/keybinding reference. |
|
|
21
|
-
| `src/domains/context/` | `CLIO-CODER.md` bootstrap, codewiki generation, prompt context assembly, project rules | [context-engine.md](context-engine.md) | `documented` |
|
|
21
|
+
| `src/domains/context/` | `CLIO-CODER.md` bootstrap, codewiki generation, prompt context assembly, project rules, non-destructive working-set eviction (`age-horizon` and `structural-v1` policies, protection predicates, path index, byte-stable markers, recall by ref) | [context-engine.md](context-engine.md), [context-working-set.md](context-working-set.md) | `documented` | Context window, token accounting, and the three compaction mechanisms in the engine reference; the working-set layer has its own guide covering the vocabulary, both ledger record kinds and format v4, the marker contract, both policies with their rule order, recall semantics, and the operator surfaces. |
|
|
22
22
|
| `src/domains/dispatch/` | Fleet orchestration, assignment store, batch tracker, admission, route planner, receipt integrity v15 | [fleet-dispatch.md](fleet-dispatch.md), [dispatch-architecture-rationale.md](dispatch-architecture-rationale.md), [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `documented` | Multi-node fleet dispatch, admission invariants, and receipt verification fully documented. |
|
|
23
23
|
| `src/domains/eval/` | Suite v2 YAML schema, eval runner, metrics, reporters, workspace sandboxing | [eval-runner.md](eval-runner.md), [evals-internal.md](evals-internal.md) | `documented` | Documented in eval runner and soak benchmark guides. |
|
|
24
24
|
| `src/domains/evidence/` | Evidence bundles, findings taxonomy, provenance store, failure attribution | [evidence-and-memory.md](evidence-and-memory.md) | `documented` | Documented in evidence directory structures and memory retrieval guide. |
|
|
@@ -33,14 +33,14 @@ This matrix maps every top-level directory in `src/` and every domain directory
|
|
|
33
33
|
| `src/domains/resources/` | Skill package discovery, marketplace index resolution, prompt resources | [skills-marketplace.md](skills-marketplace.md), [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Skills marketplace, publishing flows, and resource managers documented. |
|
|
34
34
|
| `src/domains/safety/` | Policy engine, action classifiers, damage-control rules, path policies, finish contract, audit log | [safety-model.md](safety-model.md), [scientific-validation.md](scientific-validation.md) | `documented` | Policy evaluation order, 10-step sequence, write containment, and finish contract documented. |
|
|
35
35
|
| `src/domains/scheduling/` | Capacity lease acquisition, heartbeats, expiry, cross-process locks, cluster scheduling | [capacity-and-scheduling.md](capacity-and-scheduling.md), [fleet-dispatch.md](fleet-dispatch.md) | `documented` | Dedicated capacity leasing, heartbeat TTL, and cross-process lock reference. |
|
|
36
|
-
| `src/domains/session/` |
|
|
36
|
+
| `src/domains/session/` | Session ledger format v4, tree branching (`/tree`), `/fork`, `/resume`, checkpoints, protected-artifact journal | [session-lifecycle.md](session-lifecycle.md), [context-working-set.md](context-working-set.md) | `documented` | Dedicated session lifecycle guide covering branching, journal, and recovery; the `contextEviction` and `contextRecall` records added at format v4 are specified in the working-set guide. |
|
|
37
37
|
| `src/domains/share/` | Portable share archive bundles, manifest verification, import/export flows | [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Share archives and portable bundle formats documented in extensions guide. |
|
|
38
|
-
| `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.
|
|
38
|
+
| `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.4. |
|
|
39
39
|
|
|
40
40
|
## Cross-Cutting Reference Guides
|
|
41
41
|
|
|
42
42
|
In addition to source subsystem mappings, the documentation set includes cross-cutting contracts:
|
|
43
43
|
|
|
44
44
|
1. [artifact-versions.md](artifact-versions.md): Canonical version registry and migration contract for all 9 serialized artifact schemas across Clio Coder.
|
|
45
|
-
2. [glossary.md](glossary.md): Formal definitions of
|
|
45
|
+
2. [glossary.md](glossary.md): Formal definitions of 45 core architectural concepts mapped to their TypeScript types in `src/`.
|
|
46
46
|
3. [troubleshooting.md](troubleshooting.md): Comprehensive diagnostic and remediation guide keyed by exact user-facing error strings.
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Documentation Standards and Codebase Alignment
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.4).
|
|
5
5
|
|
|
6
6
|
Clio Coder is an experimental community alpha. Documentation should help contributors and early users work from the source of truth without overstating maturity. When docs drift, prefer the current source and tests over older prose or aspirational roadmap notes.
|
|
7
7
|
|
|
@@ -34,7 +34,8 @@ Classify claims clearly:
|
|
|
34
34
|
| [README.md](../README.md) | `CHANGELOG.md`, package metadata, release receipts | Product overview, install, first run, alpha framing, and release status. |
|
|
35
35
|
| [docs/README.md](README.md) | This docs directory | Documentation hub. |
|
|
36
36
|
| [commands-and-modes.md](commands-and-modes.md) | `src/cli/index.ts`, `src/cli/args.ts`, `src/interactive/slash-commands.ts`, `src/domains/dispatch/**` | CLI commands, headless run flags (`--session`, `--continue`, `--json-events`), session continuity, `--json` wire projection promise, slash commands, keybindings, live steering. |
|
|
37
|
-
| [context-engine.md](context-engine.md) | `src/domains/context/**`, `src/domains/session/context-accounting.ts`, `src/domains/session/context-ledger.ts`, `src/domains/session/compaction/` | Context window resolution, per-model probe capabilities, token accounting, snapshots,
|
|
37
|
+
| [context-engine.md](context-engine.md) | `src/domains/context/**`, `src/domains/session/context-accounting.ts`, `src/domains/session/context-ledger.ts`, `src/domains/session/compaction/` | Context window resolution, per-model probe capabilities, token accounting, snapshots, the three compaction mechanisms, model-driven `clio-coder context init`, format v4 session enforcement. |
|
|
38
|
+
| [context-working-set.md](context-working-set.md) | `src/domains/context/working-set/**`, `src/domains/session/entries.ts`, `src/interactive/turn-context.ts` | Working-set vocabulary, eviction as a projection, the `contextEviction` / `contextRecall` records, the marker contract, the `age-horizon` and `structural-v1` policies, recall semantics, and the operator surfaces. |
|
|
38
39
|
| [architecture.md](architecture.md) | `tests/boundaries/check-boundaries.ts`, `src/core/domain-loader.ts`, `src/engine/**`, `src/worker/**` | Source layout, 5 enforced boundary rules (dependency direction vs import form), runtime flow mermaid diagram, event/audit model, detect-and-rollback write boundaries. |
|
|
39
40
|
| [dispatch-architecture-rationale.md](dispatch-architecture-rationale.md) | `src/domains/dispatch/**`, `tests/boundaries/check-boundaries.ts` | Design rationale, not behavior: invariants that cross the seams a dispatch split would use, what any future split must preserve, the one dispatch→eval import, and the closed barrel-import decision. |
|
|
40
41
|
| [configuration-and-targets.md](configuration-and-targets.md) | `src/core/defaults.ts`, `src/core/config.ts`, `src/domains/providers/**`, `src/cli/configure.ts`, `src/cli/targets.ts`, `src/cli/models.ts`, `src/cli/auth.ts` | TargetDescriptor, contextWindowProvenance (`configured`, `discovered`, `catalog`, `runtime-default`), settings.yaml, strict validation, saved defaults vs live routing. |
|
|
@@ -49,16 +50,16 @@ Classify claims clearly:
|
|
|
49
50
|
| [capacity-and-scheduling.md](capacity-and-scheduling.md) | `src/domains/scheduling/**`, `src/domains/dispatch/capacity-lease.ts`, `src/domains/dispatch/reservation-store.ts` | Multi-process capacity leases (`dispatch-admission.json`), heartbeat TTLs, cross-process transaction locks (`dispatch-admission.json.lock`), and cluster drain controls. |
|
|
50
51
|
| [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `src/worker/**` | NDJSON parent-child socket protocols, control/bulk lane demuxing, watchdog timers, worker attestation (13 protocol fields), permission parking, exit codes. |
|
|
51
52
|
| [fleet-demo-runbook.md](fleet-demo-runbook.md) | `src/domains/dispatch/**` | Multi-node fleet demo: SSH setup, C++ build/repair workflow, reviewer gates, receipt verification v15. |
|
|
52
|
-
| [session-lifecycle.md](session-lifecycle.md) | `src/engine/session.ts`, `src/domains/session/**` | Session lifecycle, on-disk ledger format
|
|
53
|
+
| [session-lifecycle.md](session-lifecycle.md) | `src/engine/session.ts`, `src/domains/session/**` | Session lifecycle, on-disk ledger format v4 (`current.jsonl`), tree branching (`tree.json`), active-path lineage selection, `/fork`, `/resume`, checkpoints, and write-ahead protected-artifact journal. |
|
|
53
54
|
| [acp.md](acp.md) | `src/engine/acp/**`, `src/cli/acp.ts` | Agent Client Protocol (ACP) server over stdio, tool mediation, non-stall permission handling, timeout bounds, and error taxonomy. |
|
|
54
55
|
| [artifact-versions.md](artifact-versions.md) | `src/domains/dispatch/receipt-integrity.ts`, `src/engine/session.ts`, `src/worker/spec-contract.ts`, `src/domains/agents/fleet-contract.ts`, `src/domains/eval/schema/`, `src/domains/observability/trace-store.ts` | Version registry and migration policies for all 9 serialized artifact schemas across Clio Coder. |
|
|
55
56
|
| [exit-codes-and-output.md](exit-codes-and-output.md) | `src/cli/**`, `src/entry/**` | Global process exit codes (0, 1, 2, 3), `--help` standard on stdout, machine-readable JSON streaming (`--json`, `--json-events`), and headless stdout deliverable contracts. |
|
|
56
57
|
| [troubleshooting.md](troubleshooting.md) | `src/core/**`, `src/cli/**`, `src/domains/**` | Actionable error remediation and diagnostics keyed by exact user-facing messages. |
|
|
57
|
-
| [glossary.md](glossary.md) | `src/domains/dispatch/types.ts`, `src/tools/**`, `src/domains/agents/**`, `src/core/**` | Canonical definitions of
|
|
58
|
+
| [glossary.md](glossary.md) | `src/domains/dispatch/types.ts`, `src/tools/**`, `src/domains/agents/**`, `src/core/**` | Canonical definitions of 45 core architectural concepts mapped to `src/` types. |
|
|
58
59
|
| [documentation-coverage.md](documentation-coverage.md) | `src/**` | Complete source-to-documentation mapping matrix and subsystem coverage status. |
|
|
59
60
|
| [tui-design.md](tui-design.md) | `src/interactive/theme/tokens.ts`, `src/interactive/theme/glyphs.ts` | TUI color system, glyph vocabulary (`contextReserve`), structural layouts, state choreography, code ink. |
|
|
60
61
|
| [installation-and-lifecycle.md](installation-and-lifecycle.md) | `src/cli/paths.ts`, `src/cli/doctor.ts`, `src/cli/uninstall.ts`, `src/cli/removal.ts` | Installation, upgrade, reset, uninstallation, launcher ownership and what `--remove-binary` will and will not remove, partial-failure behavior, configuration folders (`credentials.yaml` `0o600`), and permissions. |
|
|
61
|
-
| [release-cut-checklist.md](release-cut-checklist.md) | `scripts/check-release.mjs`, `
|
|
62
|
+
| [release-cut-checklist.md](release-cut-checklist.md) | `scripts/check-release.mjs`, `tests/smoke/pack-install.test.ts`, `benchmarks/internal/`, `package.json` | Ordered release-cut steps with an explicit authorization boundary: everything external or irreversible is marked not run and needs an operator decision. |
|
|
62
63
|
| [observability.md](observability.md) | `src/domains/observability/**`, `src/interactive/view/**`, `src/domains/dispatch/**`, `src/core/bus-events.ts` | `/view` artifact browsing, receipt verification, worker diagnostics, event routing, and cost snapshots. |
|
|
63
64
|
| [evidence-and-memory.md](evidence-and-memory.md) | `src/domains/evidence/**`, `src/domains/memory/**`, `src/cli/evidence.ts`, `src/cli/memory.ts` | Evidence corpus layout, findings, memory lifecycle and prompt injection. |
|
|
64
65
|
| [proactive-memory.md](proactive-memory.md) | `src/domains/memory/**` | Proactive task memory architecture, session task bank, intervention rules, and handoff carrying. |
|
|
@@ -30,6 +30,7 @@ Durable values live in the `guardrails:` section of settings.yaml (see [configur
|
|
|
30
30
|
| `CLIO_CODER_TRUST_PROJECT_SKILLS` | off | `1` trusts project-local skills for execution (`src/domains/resources/skills/loader.ts`). |
|
|
31
31
|
| `CLIO_CODER_ALLOW_EXTERNAL_FULL_ACCESS` | off | `1` lets full-auto pass through to external CLI runtimes with their own full access (`src/engine/claude/subprocess-runtime.ts`, `src/engine/antigravity/subprocess-runtime.ts`). |
|
|
32
32
|
| `CLIO_CODER_FORCE_COMPACT` | off | `1` forces compaction on the next interactive turn (`src/interactive/chat-loop.ts`). |
|
|
33
|
+
| `CLIO_CODER_LEGACY_MASK` | off | `1` temporarily restores the destructive stale-observation mask before summary compaction; remove it after compatibility diagnosis. |
|
|
33
34
|
| `CLIO_CODER_STATUS_STUCK_MS` | 180000 | Stuck-turn watchdog threshold (`src/interactive/status/watchdog.ts`). |
|
|
34
35
|
| `CLIO_CODER_SHUTDOWN_HOOK_MS` | 500 | Wall-clock budget per shutdown hook (`src/core/termination.ts`). |
|
|
35
36
|
| `CLIO_CODER_HOOK_BUDGET_MS` | per-phase built-ins | Global middleware hook wall-clock budget (`src/domains/middleware/budget.ts`). |
|
|
@@ -111,4 +112,4 @@ Set by Clio for its own processes; not operator knobs.
|
|
|
111
112
|
| `CLIO_CODER_TEST_STAGE1_DELAY_MS`, `CLIO_CODER_TEST_STAGE1_FAIL` | `NODE_ENV=test`-only, bounded instant-shell interleaving and injected hydration failure seams for the built PTY acceptance suite (`src/cli/clio.ts`). |
|
|
112
113
|
| `CLIO_CODER_REQUIRE_HOME_PREFIX` | Test guardrail: abort if resolved directories escape `CLIO_CODER_HOME` (`src/core/init.ts`). |
|
|
113
114
|
|
|
114
|
-
Variables used only by
|
|
115
|
+
Variables used only by the Terminal-Bench agent under `benchmarks/community/` and by the install script are not part of the shipped runtime and are documented inline where they are consumed. The live drivers under `benchmarks/internal/` take no environment of their own: the target comes from `--target <id>`.
|
package/docs/eval-runner.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Clio Coder Local Evaluation Runner
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.4).
|
|
5
5
|
|
|
6
6
|
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
7
|
|
package/docs/evals-internal.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Internal Eval Suites
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** Interactive blueprints are available for internal evaluation suites at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) and soak benchmark suites at [docs/html/soak_blueprint.html](html/soak_blueprint.html) (Version: 0.3.4).
|
|
5
5
|
|
|
6
6
|
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
7
|
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
@@ -274,6 +274,19 @@ thresholds:
|
|
|
274
274
|
|
|
275
275
|
The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
|
|
276
276
|
|
|
277
|
+
Every suite runs through the product's own eval runner against a configured
|
|
278
|
+
target; there is no separate soak runner:
|
|
279
|
+
|
|
280
|
+
```bash
|
|
281
|
+
npm run build
|
|
282
|
+
clio-coder eval run --suite benchmarks/soak/clio-soak.yaml \
|
|
283
|
+
--target <id> --model <wireId> --clio-coder-entry dist/cli/index.js
|
|
284
|
+
```
|
|
285
|
+
|
|
286
|
+
`tests/contracts/eval-soak-suite.test.ts` loads all four files in CI and drives
|
|
287
|
+
`clio-soak.yaml` against a stub that seals receipts on purpose, so the gate is
|
|
288
|
+
known to fail when sealing fails; the model runs themselves are operator-run.
|
|
289
|
+
|
|
277
290
|
The soak suite comprises four specialized suite files:
|
|
278
291
|
|
|
279
292
|
### 1. Machinery Under Load (`clio-soak.yaml`)
|