@iowarp/clio-coder 0.3.4 → 0.3.7
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +65 -2
- package/CONTRIBUTING.md +6 -6
- package/README.md +16 -5
- package/dist/{acp-S5R4RR5B.js → acp-SK4MD6MM.js} +11 -11
- package/dist/{agents-P6DMMVZY.js → agents-2FN2K6ME.js} +33 -26
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-2XCZLPKS.js → auth-QIYZWM5I.js} +15 -15
- package/dist/{chunk-EKMEHE4H.js → chunk-33YXPOE3.js} +2 -3
- package/dist/{chunk-VAWWTKDP.js → chunk-3HAPLH5M.js} +11 -11
- package/dist/{chunk-YCWGATWI.js → chunk-465YSENW.js} +3 -3
- package/dist/{chunk-4OC57DA6.js → chunk-4DGYLA73.js} +53 -2
- package/dist/{chunk-UZHIZC5S.js → chunk-4DWFMQDR.js} +61 -76
- package/dist/{chunk-ZWMF7253.js → chunk-5C3AQNDW.js} +328 -9
- package/dist/{chunk-22NAGB7X.js → chunk-5C77SEEY.js} +5 -94
- package/dist/{chunk-WPQLXFOZ.js → chunk-5FR74PWO.js} +3 -2
- package/dist/{chunk-35MKKU5R.js → chunk-5UJ6ECTS.js} +18 -10
- package/dist/{chunk-BRXQQJFP.js → chunk-6M7VS3J3.js} +571 -50
- package/dist/{chunk-QQK64KLB.js → chunk-6TUKSZVF.js} +141 -23
- package/dist/{chunk-N4CZJQRK.js → chunk-AB4XIIVB.js} +8 -6
- package/dist/{chunk-KRPY7NTG.js → chunk-BMWK7ZIZ.js} +14 -20
- package/dist/{chunk-4BPJXDWC.js → chunk-C4JBQ5SR.js} +30 -14
- package/dist/{chunk-ZYKPLLNQ.js → chunk-CEYBNUGC.js} +821 -83
- package/dist/{chunk-VEZEGCGW.js → chunk-D4MDIG46.js} +20 -18
- package/dist/chunk-DJNLUABN.js +843 -0
- package/dist/{chunk-BP4OYD6A.js → chunk-DMD2AGVS.js} +21 -2
- package/dist/{chunk-KOHPCX4K.js → chunk-DOOEX22V.js} +2 -2
- package/dist/chunk-DQA7QLMD.js +123 -0
- package/dist/chunk-DR52UMZW.js +21 -0
- package/dist/{chunk-3HZ5RWN2.js → chunk-EBEFWSGL.js} +9 -7
- package/dist/{chunk-EDRHSCIE.js → chunk-EELBMBT6.js} +128 -13
- package/dist/{chunk-HV5X7OR2.js → chunk-EOOQZZDE.js} +16 -14
- package/dist/{chunk-WR67VIZY.js → chunk-FOT2FX5J.js} +63 -5
- package/dist/{chunk-BPGS2WCQ.js → chunk-GEYXPTRF.js} +2 -1
- package/dist/{chunk-FYYLNIL5.js → chunk-GH5622CP.js} +2 -2
- package/dist/{chunk-BEY543CS.js → chunk-GOXNB3AO.js} +5 -2
- package/dist/chunk-GWS3VEIW.js +195 -0
- package/dist/{chunk-G4BMMOKF.js → chunk-HVDIIIQW.js} +2 -2
- package/dist/chunk-HWUFFB6L.js +83 -0
- package/dist/{chunk-X6COSD2O.js → chunk-J7PIKKWC.js} +8 -436
- package/dist/{chunk-NILBFAPG.js → chunk-JNXPYBB4.js} +2 -2
- package/dist/{chunk-4VP4KH3K.js → chunk-JRIO5UD2.js} +4 -4
- package/dist/{chunk-K6WL7QZT.js → chunk-JTSEDYVQ.js} +7 -7
- package/dist/{chunk-QKMUKYO7.js → chunk-KCMKRQX4.js} +236 -84
- package/dist/chunk-KZ2H5X4G.js +1026 -0
- package/dist/{chunk-A2GZF7DC.js → chunk-LADCF22A.js} +13 -13
- package/dist/chunk-LCGCVYZ4.js +57 -0
- package/dist/chunk-M4AKACEO.js +382 -0
- package/dist/{chunk-POHLU5DW.js → chunk-M6L6IDJG.js} +3 -3
- package/dist/{chunk-4JUF2NNX.js → chunk-MXI6J5JF.js} +7 -7
- package/dist/{chunk-X4RCMKVQ.js → chunk-NDINPTJ4.js} +2 -2
- package/dist/{chunk-TTNYS3EA.js → chunk-OB5HIGJY.js} +1 -1
- package/dist/{chunk-7RXG6QRZ.js → chunk-OBMAI2DP.js} +61 -840
- package/dist/{chunk-5M54SPOL.js → chunk-ODFEOB4F.js} +161 -5
- package/dist/chunk-PD3MESLB.js +242 -0
- package/dist/{chunk-ED4KHGC3.js → chunk-PPAMZ32Z.js} +9 -2
- package/dist/{chunk-VMNQ6OZA.js → chunk-QCTRSGHQ.js} +963 -786
- package/dist/chunk-RVG5JXAL.js +41 -0
- package/dist/{chunk-RD5U66HV.js → chunk-SROCI7ZU.js} +7 -7
- package/dist/{chunk-MFFY33HR.js → chunk-THKY7CD7.js} +466 -205
- package/dist/{chunk-34475P3I.js → chunk-TSHXZTOQ.js} +5 -4
- package/dist/{chunk-PCZJO5TI.js → chunk-UFQ3F4FW.js} +13 -178
- package/dist/{chunk-AD2SYQYC.js → chunk-UHXRNZ2J.js} +121 -3
- package/dist/chunk-UND3GU2L.js +103 -0
- package/dist/{chunk-QQL5RT5M.js → chunk-UUANF5CR.js} +2323 -2114
- package/dist/{chunk-VJWL6YS5.js → chunk-UUVG37B4.js} +2 -2
- package/dist/chunk-UVDSQ6LW.js +472 -0
- package/dist/{chunk-QWU7ZBO7.js → chunk-VQNODYQ4.js} +215 -56
- package/dist/chunk-VREKEFLL.js +37 -0
- package/dist/{chunk-2TZWSW76.js → chunk-WHGPSPT5.js} +2 -2
- package/dist/{chunk-TW3WDMVS.js → chunk-WHJYKASB.js} +2 -2
- package/dist/{chunk-MEQ45TQ4.js → chunk-WJHBC77E.js} +21 -7
- package/dist/{chunk-HXG4IURW.js → chunk-X2KV5FXT.js} +2 -2
- package/dist/{chunk-YHZX5GEU.js → chunk-XAKHZX5N.js} +2 -2
- package/dist/{chunk-2LZI5CAG.js → chunk-XEGB6BCN.js} +228 -36
- package/dist/{chunk-E25LMLRW.js → chunk-YD734TPH.js} +2 -2
- package/dist/{verifiers-4UUM6TEE.js → chunk-YTYFXUI3.js} +121 -372
- package/dist/{chunk-3JLKSKD7.js → chunk-ZGH7FGS5.js} +17 -7
- package/dist/{chunk-VSNATDE6.js → chunk-ZZMN5OM4.js} +2 -2
- package/dist/cli/index.js +34 -32
- package/dist/{clio-J5JIOIDS.js → clio-WBVQEBKO.js} +7 -7
- package/dist/{code-nav-AXCXSBHX.js → code-nav-FGGFIE7L.js} +7 -7
- package/dist/codewiki/build-worker.js +4 -4
- package/dist/{components-KELWS457.js → components-F7OEATSO.js} +5 -5
- package/dist/{config-OEBMIN2U.js → config-TRBL3RCF.js} +48 -41
- package/dist/{configure-PUQOSIXQ.js → configure-OLCVPHNM.js} +17 -17
- package/dist/{context-URSXPBCK.js → context-MJIJ6GOX.js} +12 -12
- package/dist/{context-EKDCKUUZ.js → context-WFPKQSM6.js} +26 -9
- package/dist/{context-MGSE4Z2T.js → context-XEWE3MOJ.js} +44 -37
- package/dist/{context-clear-KDAJRNUK.js → context-clear-KNOS2JPB.js} +44 -37
- package/dist/{context-index-BZ4UYMTC.js → context-index-SSR5ECNE.js} +3 -3
- package/dist/{context-working-set-SBKMPPI2.js → context-working-set-EUXAZI6N.js} +14 -13
- package/dist/{dispatch-runner-MSWN72NK.js → dispatch-runner-B7MTOVKL.js} +321 -60
- package/dist/{docs-2C2LTVT2.js → docs-FLJTIDSE.js} +5 -5
- package/dist/{doctor-7BSE27PJ.js → doctor-RN4YKO2X.js} +15 -15
- package/dist/{eval-IZGDOO4H.js → eval-RUBJVSNQ.js} +52 -236
- package/dist/{evidence-SR7WXB5B.js → evidence-JZNBUOQZ.js} +39 -33
- package/dist/{evolve-K7VE2CBX.js → evolve-FJVC4KKI.js} +39 -33
- package/dist/{extensions-QVDOHDGJ.js → extensions-IQL36S7K.js} +5 -5
- package/dist/{fleet-7XMJNQNF.js → fleet-BDKYJFCP.js} +243 -370
- package/dist/fleet-commands-ZFIWZSB3.js +70 -0
- package/dist/fleet-graph-Y6HPXIVF.js +125 -0
- package/dist/fleet-new-RDVJLHHH.js +48 -0
- package/dist/{fleet-preflight-AQNAH644.js → fleet-preflight-BHSNPBMH.js} +2 -2
- package/dist/fleet-validate-BIYREGIK.js +79 -0
- package/dist/{init-JGNPAYXT.js → init-LQUB5COQ.js} +57 -48
- package/dist/library-NJAHIGG4.js +217 -0
- package/dist/memory-OG6HOYKM.js +472 -0
- package/dist/{models-ZMMLFJNN.js → models-5ZG5XY7J.js} +23 -22
- package/dist/{monitor-2F3T5KHP.js → monitor-TJ7AMTGB.js} +69 -35
- package/dist/{orchestrator-ORHT43JB.js → orchestrator-WZYB54DM.js} +4868 -1189
- package/dist/{paths-UXLN5YYZ.js → paths-XUC7GS6E.js} +5 -5
- package/dist/{reset-NXGTYNUO.js → reset-PXQT45IY.js} +8 -8
- package/dist/{run-RF4WJGMT.js → run-FQ74YF62.js} +82 -62
- package/dist/{share-UT3W6E4M.js → share-FW7SVCL3.js} +34 -10
- package/dist/{skills-PSACKC5Q.js → skills-7E7IRB3R.js} +25 -9
- package/dist/{skills-eval-WJSI55RZ.js → skills-eval-LI75W6OK.js} +43 -35
- package/dist/{targets-PIIRAOYS.js → targets-4CIFKCTW.js} +27 -24
- package/dist/{terminal-lease-ULWXWNVY.js → terminal-lease-WUZY7ZV5.js} +5 -4
- package/dist/{uninstall-FZCQCDKC.js → uninstall-7FV7IP4E.js} +5 -5
- package/dist/{upgrade-346TZ6AV.js → upgrade-K2HVIVMQ.js} +21 -20
- package/dist/{usage-6KKXR32N.js → usage-GTZELZQX.js} +159 -59
- package/dist/verifiers-RLAHT27O.js +336 -0
- package/dist/{verify-X5HDROLA.js → verify-BX3BRKH5.js} +7 -6
- package/dist/{wiki-generate-7STOCIFZ.js → wiki-generate-ASIFASCN.js} +58 -48
- package/dist/worker/entry.js +98 -84
- package/dist/{workspace-G4ZWUIPR.js → workspace-ZJ6BFM3Q.js} +4 -4
- package/docs/README.md +4 -3
- package/docs/acp.md +1 -1
- package/docs/alcf-provider.md +1 -1
- package/docs/architecture.md +2 -2
- package/docs/artifact-placement.md +1 -2
- package/docs/artifact-versions.md +10 -6
- package/docs/built-in-agents.md +26 -2
- package/docs/capacity-and-scheduling.md +1 -1
- package/docs/commands-and-modes.md +90 -8
- package/docs/configuration-and-targets.md +90 -2
- package/docs/context-engine.md +4 -2
- package/docs/context-working-set.md +4 -4
- package/docs/development-pipeline.md +1 -1
- package/docs/dispatch-architecture-rationale.md +1 -1
- package/docs/documentation-coverage.md +4 -4
- package/docs/documentation-guide.md +4 -4
- package/docs/eval-runner.md +1 -1
- package/docs/evals-internal.md +4 -45
- package/docs/evidence-and-memory.md +70 -10
- package/docs/evolution.md +1 -1
- package/docs/exit-codes-and-output.md +4 -1
- package/docs/extensions-and-sharing.md +6 -2
- package/docs/fleet-demo-runbook.md +2 -2
- package/docs/fleet-dispatch.md +224 -11
- package/docs/git-commit-provenance.md +2 -2
- package/docs/glossary.md +1 -1
- package/docs/installation-and-lifecycle.md +2 -2
- package/docs/middleware-and-components.md +20 -2
- package/docs/model-catalog.md +1 -1
- package/docs/observability.md +55 -8
- package/docs/proactive-memory.md +26 -16
- package/docs/prompt-envelope-and-tools.md +4 -2
- package/docs/provider-adapter-cookbook.md +1 -1
- package/docs/release-cut-checklist.md +83 -65
- package/docs/resource-library.md +59 -0
- package/docs/safety-model.md +29 -7
- package/docs/scientific-validation.md +3 -3
- package/docs/session-lifecycle.md +37 -1
- package/docs/skills-marketplace.md +16 -3
- package/docs/tool-usage.md +14 -7
- package/docs/trace-store.md +1 -1
- package/docs/troubleshooting.md +1 -1
- package/docs/tui-design.md +38 -4
- package/docs/worker-dispatch-mechanics.md +3 -3
- package/package.json +7 -4
- package/src/cli/agents.ts +2 -3
- package/src/cli/argv.ts +14 -1
- package/src/cli/fleet-commands.ts +37 -0
- package/src/cli/fleet-graph.ts +102 -0
- package/src/cli/fleet-new.ts +36 -0
- package/src/cli/fleet-preflight.ts +121 -0
- package/src/cli/fleet-validate.ts +30 -0
- package/src/cli/fleet.ts +188 -335
- package/src/cli/index.ts +4 -2
- package/src/cli/library.ts +190 -0
- package/src/cli/memory.ts +272 -10
- package/src/cli/modes/json-stream.ts +2 -2
- package/src/cli/modes/print.ts +12 -1
- package/src/cli/run.ts +22 -2
- package/src/cli/share.ts +13 -1
- package/src/cli/targets.ts +12 -3
- package/src/cli/usage.ts +160 -20
- package/src/core/bus-events.ts +7 -0
- package/src/core/commit-attribution.ts +4 -4
- package/src/core/config.ts +130 -0
- package/src/core/defaults.ts +81 -0
- package/src/core/response-model-id.ts +134 -0
- package/src/core/toml.ts +62 -0
- package/src/core/workspace-files.ts +0 -1
- package/src/domains/agents/builtins/architect.md +2 -1
- package/src/domains/agents/builtins/oracle.md +33 -0
- package/src/domains/agents/catalog.ts +18 -5
- package/src/domains/agents/fleet-contract.ts +278 -16
- package/src/domains/agents/index.ts +14 -0
- package/src/domains/agents/recipe.ts +54 -14
- package/src/domains/agents/result-contract.ts +242 -5
- package/src/domains/config/classify.ts +4 -0
- package/src/domains/context/bootstrap.ts +36 -27
- package/src/domains/context/project-metadata.ts +19 -63
- package/src/domains/context/prompt-context.ts +8 -0
- package/src/domains/context/working-set/policies/index.ts +3 -4
- package/src/domains/dispatch/active-route-planner.ts +14 -0
- package/src/domains/dispatch/backoff.ts +2 -1
- package/src/domains/dispatch/budget-envelope.ts +396 -0
- package/src/domains/dispatch/capability-match.ts +1 -0
- package/src/domains/dispatch/checkout-writer-lease.ts +175 -0
- package/src/domains/dispatch/contract.ts +36 -0
- package/src/domains/dispatch/delegation-plan.ts +167 -0
- package/src/domains/dispatch/execution-plan.ts +76 -5
- package/src/domains/dispatch/execution-role.ts +3 -1
- package/src/domains/dispatch/execution-scheduler.ts +183 -67
- package/src/domains/dispatch/extension.ts +339 -36
- package/src/domains/dispatch/fleet-gate.ts +14 -0
- package/src/domains/dispatch/fleet-plan.ts +63 -3
- package/src/domains/dispatch/fleet-run.ts +737 -0
- package/src/domains/dispatch/gate-role-prompts.ts +9 -0
- package/src/domains/dispatch/host-verification.ts +178 -0
- package/src/domains/dispatch/index.ts +38 -0
- package/src/domains/dispatch/intent.ts +159 -0
- package/src/domains/dispatch/orphan-recovery.ts +1 -0
- package/src/domains/dispatch/receipt-integrity.ts +12 -4
- package/src/domains/dispatch/state.ts +37 -3
- package/src/domains/dispatch/types.ts +61 -9
- package/src/domains/dispatch/validation.ts +80 -6
- package/src/domains/dispatch/worker-spawn.ts +14 -3
- package/src/domains/eval/metrics/evidence.ts +0 -116
- package/src/domains/eval/metrics/invariants.ts +1 -1
- package/src/domains/eval/runners/clio-run.ts +1 -10
- package/src/domains/eval/runners/external-command.ts +2 -29
- package/src/domains/eval/schema/suite.ts +0 -7
- package/src/domains/eval/suites/run.ts +1 -7
- package/src/domains/evidence/trust-status.ts +10 -1
- package/src/domains/memory/index.ts +22 -0
- package/src/domains/memory/operations.ts +58 -1
- package/src/domains/memory/promotion.ts +281 -0
- package/src/domains/memory/prompt-section.ts +25 -5
- package/src/domains/memory/proposal.ts +51 -7
- package/src/domains/memory/task-bank.ts +3 -2
- package/src/domains/memory/task-memory-handoff.ts +181 -24
- package/src/domains/memory/task-memory-policy.ts +3 -1
- package/src/domains/memory/types.ts +37 -0
- package/src/domains/memory/validate.ts +178 -0
- package/src/domains/middleware/index.ts +15 -0
- package/src/domains/middleware/memory-intervention.ts +35 -25
- package/src/domains/middleware/runtime.ts +6 -0
- package/src/domains/middleware/skills-reminder.ts +19 -4
- package/src/domains/middleware/stalled-turn.ts +43 -1
- package/src/domains/middleware/types.ts +10 -0
- package/src/domains/middleware/watchdog.ts +281 -0
- package/src/domains/observability/contract.ts +9 -2
- package/src/domains/observability/cost.ts +31 -4
- package/src/domains/observability/extension.ts +2 -2
- package/src/domains/observability/index.ts +10 -0
- package/src/domains/observability/out-of-turn-usage.ts +223 -0
- package/src/domains/providers/index.ts +3 -0
- package/src/domains/providers/model-discovery.ts +9 -0
- package/src/domains/providers/runtime-resolution.ts +38 -1
- package/src/domains/providers/runtimes/common/probe-helpers.ts +97 -16
- package/src/domains/providers/types/context-window-slots.ts +18 -0
- package/src/domains/providers/types/runtime-descriptor.ts +3 -1
- package/src/domains/resources/index.ts +20 -0
- package/src/domains/resources/library.ts +326 -0
- package/src/domains/resources/skills/marketplace.ts +37 -12
- package/src/domains/safety/call-target.ts +211 -14
- package/src/domains/safety/decision-presentation.ts +268 -0
- package/src/domains/safety/redaction.ts +73 -0
- package/src/domains/session/context-ledger.ts +10 -1
- package/src/domains/session/decision-board.ts +4 -0
- package/src/domains/session/entries.ts +3 -0
- package/src/domains/session/handoff.ts +629 -0
- package/src/domains/session/history.ts +68 -19
- package/src/domains/session/usage.ts +24 -7
- package/src/domains/share/archive.ts +67 -2
- package/src/engine/acp/event-mapper.ts +7 -0
- package/src/engine/acp/server.ts +29 -2
- package/src/engine/apis/lmstudio.ts +25 -4
- package/src/engine/apis/openai-completions.ts +147 -22
- package/src/engine/claude/sdk-runtime.ts +8 -2
- package/src/engine/claude/tool-safety.ts +13 -0
- package/src/engine/loop-guard.ts +27 -3
- package/src/engine/worker-events.ts +4 -3
- package/src/engine/worker-runtime.ts +59 -54
- package/src/entry/orchestrator.ts +55 -1
- package/src/interactive/bus-notices.ts +26 -0
- package/src/interactive/chat-loop-messages.ts +22 -0
- package/src/interactive/chat-loop.ts +248 -1
- package/src/interactive/chat-renderer.ts +41 -3
- package/src/interactive/clio-editor.ts +44 -7
- package/src/interactive/context-overlay.ts +43 -5
- package/src/interactive/cost-overlay.ts +70 -11
- package/src/interactive/council-dispatch.ts +30 -0
- package/src/interactive/council-grid.ts +213 -0
- package/src/interactive/council.ts +99 -0
- package/src/interactive/dispatch-board.ts +471 -50
- package/src/interactive/fleet-run-preview.ts +307 -0
- package/src/interactive/footer/notifications.ts +219 -0
- package/src/interactive/footer/widgets.ts +13 -0
- package/src/interactive/handoff-round.ts +56 -0
- package/src/interactive/interactive-application.ts +49 -2
- package/src/interactive/interactive-event-projection.ts +9 -1
- package/src/interactive/interactive-input-runtime.ts +11 -1
- package/src/interactive/interactive-presentation.ts +11 -1
- package/src/interactive/interactive-slash-runtime.ts +52 -2
- package/src/interactive/interactive-subscriptions.ts +14 -2
- package/src/interactive/memory-overlay.ts +89 -4
- package/src/interactive/oracle.ts +179 -0
- package/src/interactive/overlay-ask-user-lifecycle.ts +7 -1
- package/src/interactive/overlay-frame.ts +5 -2
- package/src/interactive/overlay-general-openers.ts +230 -2
- package/src/interactive/overlay-key-routing.ts +58 -2
- package/src/interactive/overlay-lifecycle.ts +52 -5
- package/src/interactive/overlay-permission-lifecycle.ts +33 -8
- package/src/interactive/overlay-resource-openers.ts +11 -3
- package/src/interactive/overlay-session-lifecycle.ts +234 -2
- package/src/interactive/overlay-transitions.ts +11 -0
- package/src/interactive/overlays/ask-user.ts +74 -30
- package/src/interactive/overlays/decisions.ts +3 -1
- package/src/interactive/overlays/fleet-run-approval.ts +208 -0
- package/src/interactive/overlays/handoff-review.ts +185 -0
- package/src/interactive/overlays/library-install-confirm.ts +151 -0
- package/src/interactive/overlays/list-overlay.ts +168 -2
- package/src/interactive/overlays/settings.ts +101 -4
- package/src/interactive/overlays/side-question.ts +139 -0
- package/src/interactive/overlays/skills-hub.ts +401 -15
- package/src/interactive/permission-hint.ts +35 -0
- package/src/interactive/permission-overlay.ts +95 -45
- package/src/interactive/renderers/tool-execution.ts +19 -49
- package/src/interactive/session-last-turn.ts +8 -1
- package/src/interactive/session-usage-reseed.ts +36 -10
- package/src/interactive/side-question.ts +171 -0
- package/src/interactive/slash-commands.ts +434 -7
- package/src/interactive/slash-spec.ts +19 -6
- package/src/interactive/status/summary.ts +5 -0
- package/src/interactive/status/types.ts +5 -0
- package/src/interactive/terminal-lease.ts +1 -0
- package/src/interactive/theme/tokens.ts +30 -0
- package/src/interactive/turn-context.ts +96 -23
- package/src/interactive/turn-middleware.ts +16 -1
- package/src/interactive/turn-runtime.ts +37 -8
- package/src/interactive/turn-state.ts +3 -0
- package/src/interactive/watchdog-run.ts +75 -0
- package/src/interactive/worker-progress.ts +440 -0
- package/src/interactive/worker-share.ts +56 -1
- package/src/interactive/worker-stream.ts +58 -110
- package/src/tools/agent-tools.ts +28 -3
- package/src/tools/ask-user.ts +21 -1
- package/src/tools/bootstrap.ts +3 -0
- package/src/tools/compete-worktrees.ts +13 -79
- package/src/tools/context/index.ts +2 -2
- package/src/tools/dispatch-admission.ts +242 -8
- package/src/tools/dispatch-arguments.ts +65 -1
- package/src/tools/dispatch-event-text.ts +19 -0
- package/src/tools/dispatch-plan.ts +136 -6
- package/src/tools/dispatch-runner.ts +319 -13
- package/src/tools/dispatch-types.ts +20 -1
- package/src/tools/dispatch.ts +96 -3
- package/src/tools/monitor.ts +31 -0
- package/src/tools/profiles.ts +18 -4
- package/src/tools/registry.ts +15 -5
- package/src/tools/result-disposition.ts +156 -0
- package/src/tools/result-shaping.ts +59 -1
- package/src/tools/task-worktree.ts +238 -0
- package/src/tools/verify/authoring.ts +116 -55
- package/src/tools/verify/scripts.ts +62 -0
- package/src/tools/worker-evidence.ts +21 -1
- package/src/worker/spec-contract.ts +44 -3
- package/dist/chunk-EFADSJET.js +0 -18
- package/dist/chunk-HC4CLZ2Y.js +0 -68
- package/dist/memory-4ALKDJ4Q.js +0 -246
- package/src/domains/eval/metrics/chaos-stream.ts +0 -93
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Commands and Modes
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/commands_blueprint.html](html/commands_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/commands_blueprint.html](html/commands_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
|
|
7
7
|
Clio Coder is a terminal-first alpha harness. This page keeps the command
|
|
@@ -56,7 +56,7 @@ For process exit codes, stdout deliverable guarantees, and machine-readable JSON
|
|
|
56
56
|
| `clio-coder dev components diff --from <a> --to <b> [--json]` | Compare component snapshots. |
|
|
57
57
|
| `clio-coder evidence build\|inspect\|list` | Build and inspect deterministic evidence artifacts. |
|
|
58
58
|
| `clio-coder eval validate\|run\|report\|compare\|gate` | Validate, run, report, compare, and gate local evaluation suites (Suite v2). |
|
|
59
|
-
| `clio-coder memory list\|propose\|approve\|reject\|prune` | Manage scoped, evidence-linked memory records. |
|
|
59
|
+
| `clio-coder memory list\|propose\|promote\|approve\|reject\|prune` | Manage scoped, evidence-linked memory records. |
|
|
60
60
|
| `clio-coder trace runs [--db PATH] [--limit N] [--json]` | List runs recorded in the durable trace mirror beside the ledger. |
|
|
61
61
|
| `clio-coder trace phases <runId> [--db PATH]` | Show one run's recorded phases. |
|
|
62
62
|
| `clio-coder trace tail <runId> [--follow] [--db PATH]` | Tail one run's recorded events; `--follow` streams as they land. |
|
|
@@ -145,20 +145,24 @@ The registry table below lists the available interactive slash commands. On a ba
|
|
|
145
145
|
| `/quit` | `/quit` | Exit Clio Coder |
|
|
146
146
|
| `/help` | `/help [query]` | Open the interactive help center showing commands and keys |
|
|
147
147
|
| `/skill` | `/skill [name] [task]` | Open the Skills Hub or invoke a skill |
|
|
148
|
+
| `/library` | `/library [kind]` | Open the Skills Hub on a resource library tab |
|
|
148
149
|
| `/prompts` | `/prompts` | List prompt templates |
|
|
149
150
|
| `/extensions` | `/extensions` | List installed extensions |
|
|
150
151
|
| `/interop` | `/interop` | Review other coding agents detected on this machine |
|
|
151
152
|
| `/share` | `/share [runId] \| /share export <path> \| /share import [--dry-run] [--force] <path>` | Share a worker result with the main agent, or export and import Clio archives |
|
|
152
153
|
| `/run` | `/run [--agent-profile <profile>] [--runtime <runtimeId>] [--target <id>] [--model <id>] [--thinking <level>] [--tool-profile <minimal-local\|science-local\|full-agent>] [--require <cap>] [--share] <agent> <task>` | Run a fleet agent |
|
|
153
154
|
| `/delegate` | `/delegate [--share] <agent-id> <task>` | Run an ACP delegation agent |
|
|
155
|
+
| `/btw` | `/btw <question>` | Ask a side question that never enters the session transcript |
|
|
156
|
+
| `/oracle` | `/oracle <question>` | Ask a read-only advisor to challenge a question against this session's settled decisions |
|
|
157
|
+
| `/council` | `/council [--roster <name>] [--rounds <n>] [--synthesis <judge\|vote\|none>] <task>` | Ask a roster of read-only members the same task, with an optional vote or judge synthesis |
|
|
154
158
|
| `/agents` | `/agents` | List Clio agents and ACP delegation agents |
|
|
155
159
|
| `/targets` | `/targets` | Open Settings → Targets: health, use, connect, probe, remove |
|
|
156
160
|
| `/cost` | `/cost` | Show session token and cost totals |
|
|
157
161
|
| `/context` | `/context compact [instructions] \| /context recall <ref> \| /context init \| /context refresh \| /context reset` | Context hub: window overlay plus compact, recall, init, refresh, and reset |
|
|
158
|
-
| `/fleet` | `/fleet
|
|
162
|
+
| `/fleet` | `/fleet run [--var <key=value>] <name>` | Open Settings → Fleet, or run a fleet contract with an approval preview |
|
|
159
163
|
| `/decisions` | `/decisions` | Show settled interview decisions and operator revisions |
|
|
160
164
|
| `/tasks` | `/tasks add <text> \| /tasks hand <id> \| /tasks done <id> \| /tasks drop <id>` | Show the session board or manage project operator tasks |
|
|
161
|
-
| `/memory` | `/memory seed` | Inspect
|
|
165
|
+
| `/memory` | `/memory seed` | Inspect, promote, or seed task memory |
|
|
162
166
|
| `/view` | `/view [filter] \| /view verify <runId>` | Browse session artifacts and verify receipts |
|
|
163
167
|
| `/thinking` | `/thinking [level]` | Set the chat thinking level, or open Settings → Orchestrator |
|
|
164
168
|
| `/output` | `/output [verbosity]` | Set transcript detail (minimal, default, verbose), or open Settings → Terminal |
|
|
@@ -167,6 +171,7 @@ The registry table below lists the available interactive slash commands. On a ba
|
|
|
167
171
|
| `/settings` | `/settings [section]` | Open interactive settings |
|
|
168
172
|
| `/resume` | `/resume` | Resume a past session |
|
|
169
173
|
| `/new` | `/new` | Start a fresh session |
|
|
174
|
+
| `/handoff` | `/handoff <goal>` | Hand this session's working state to a fresh session for a stated goal |
|
|
170
175
|
| `/tree` | `/tree` | Open session tree navigator |
|
|
171
176
|
| `/fork` | `/fork` | Fork from an assistant turn |
|
|
172
177
|
| `/export` | `/export [path]` | Export a self-contained HTML transcript by default; a `.md` path writes Markdown |
|
|
@@ -190,6 +195,72 @@ There are no slash-command aliases. `/context compact`, `/quit`, `/model`,
|
|
|
190
195
|
spellings stay errors that name `/help` instead of guessing which operation the
|
|
191
196
|
operator intended.
|
|
192
197
|
|
|
198
|
+
`/btw <question>` runs one model round beside the session and renders the answer
|
|
199
|
+
in an overlay. It sends the same compiled message history the next turn would
|
|
200
|
+
send, as read-only input, under a short system instruction saying this is a side
|
|
201
|
+
question, with no tools. Nothing about the round is appended: not the session
|
|
202
|
+
JSONL, not the transcript panel, not the context ledger, not the task board. That
|
|
203
|
+
is the point of it. A fleet run briefs its workers from the transcript, so a
|
|
204
|
+
question the operator asks to orient themselves mid-run would otherwise become
|
|
205
|
+
context every worker inherits. Esc closes the overlay, and cancels the round if it
|
|
206
|
+
is still streaming. `/btw` during an in-flight turn is refused with a notice
|
|
207
|
+
rather than queued, because a side question answered after the run it was asked
|
|
208
|
+
during has already missed its moment. The round's token usage still shows in
|
|
209
|
+
`/cost`, labeled as a side question, because it was a real call and cost real
|
|
210
|
+
money; it is deliberately not counted as a turn.
|
|
211
|
+
|
|
212
|
+
`/council [--roster <name>] [--rounds <n>] [--synthesis judge|vote|none] <task>`
|
|
213
|
+
asks a roster of two to five read-only members the same task and puts the group on
|
|
214
|
+
the Fleet Runs board as one card. It owns no dispatch path of its own: the command
|
|
215
|
+
builds dispatch-tool arguments and admits them through the tool registry, so a
|
|
216
|
+
supervised autonomy level parks the call and the approval overlay names every
|
|
217
|
+
member's label, target, model, node, round count, and synthesis mode before
|
|
218
|
+
anything runs. Members are pinned to read-only autonomy and the council tool
|
|
219
|
+
surface by admission, exactly as they are for a council the model asks for.
|
|
220
|
+
|
|
221
|
+
`--roster` names a `workers.rosters` entry. Without it the command takes
|
|
222
|
+
`workers.rosters.default` when that roster exists, and with neither it refuses
|
|
223
|
+
and names the setting to declare. A roster that is the only one configured is
|
|
224
|
+
still not the default: seating a council from whichever roster happens to be
|
|
225
|
+
present would run models the operator never chose. `--rounds` accepts one to
|
|
226
|
+
three and `--synthesis` accepts `judge`, `vote`, or `none`, which are the tool's
|
|
227
|
+
own bounds, enforced where the operator typed them so a council is never refused
|
|
228
|
+
after its plan has already been shown. `/council` during an in-flight turn is
|
|
229
|
+
refused with a notice rather than queued, for the same reason `/fleet run` is: an
|
|
230
|
+
approved plan describes the workspace as it stands. Nothing the members produce
|
|
231
|
+
enters the main agent's context until an operator runs `/share`.
|
|
232
|
+
|
|
233
|
+
`/handoff <goal>` carries this session's working state into a fresh session for a
|
|
234
|
+
goal the operator states. The goal is required and gated: a goal shorter than 12
|
|
235
|
+
characters is refused, and so is one of a small stoplist of non-goals such as
|
|
236
|
+
"continue", "next", or "resume". Both refusals name the rule they enforce, because
|
|
237
|
+
"keep going" is exactly the instruction a handoff exists to replace.
|
|
238
|
+
|
|
239
|
+
One model round then runs on the same out-of-turn seam `/btw` uses. It reads the
|
|
240
|
+
compiled message history the next turn would send, sends no tools, and answers
|
|
241
|
+
with JSON validated against a fixed response schema of decisions, facts, files,
|
|
242
|
+
commands, and open questions. Every list and every string is bounded; output over
|
|
243
|
+
a bound is truncated with a visible marker and the document names each bound that
|
|
244
|
+
fired, so nothing is cut silently and an over-eager answer is never a refusal.
|
|
245
|
+
|
|
246
|
+
Every file path the model names is checked against this session's read ledger and
|
|
247
|
+
never against the filesystem. Paths the session did not touch are dropped and
|
|
248
|
+
listed in the document under their own heading so the operator can see what the
|
|
249
|
+
model invented. Extracted decisions are merged with the session's settled decision
|
|
250
|
+
board, and the board wins. The result is one Markdown document opened for review:
|
|
251
|
+
Enter accepts it, `e` hands it to `$EDITOR`, and Esc cancels the whole handoff with
|
|
252
|
+
nothing written anywhere.
|
|
253
|
+
|
|
254
|
+
On accept, Clio mints a new session, writes the reviewed document into it as
|
|
255
|
+
bounded data labelled as a handoff from the old session id, and replays the old
|
|
256
|
+
session's skill activations so loaded skills carry forward. The document is never
|
|
257
|
+
written as a fabricated user turn. The old session is left untouched apart from one
|
|
258
|
+
terminal note recording the target session id. A handoff is a session operation
|
|
259
|
+
throughout: it writes no memory promotion candidate and never calls the task-memory
|
|
260
|
+
bank. `/handoff` during an in-flight turn is refused with a notice rather than
|
|
261
|
+
queued, because a document summarizing a session that is still moving would be
|
|
262
|
+
wrong by the time it was read.
|
|
263
|
+
|
|
193
264
|
The `/resume` picker accepts Page Up and Page Down to move by its 12 visible rows. Arrow keys continue to move one session at a time, and typing continues to filter the list.
|
|
194
265
|
|
|
195
266
|
Only active commands run. Typing anything command-shaped that the registry does
|
|
@@ -224,7 +295,7 @@ Configuration lives in one place: the `/settings` overlay. `/settings <section>`
|
|
|
224
295
|
|
|
225
296
|
Settings → Targets presents an operational console table (`HEALTH`, `ID`, `ROLES`, `RUNTIME`, `LATENCY`) with an in-place action/detail drawer for URL, default model, last probe error, and reachability. `Enter` opens actions for `Use` (switches active chat target and rebases model), `Connect` (runs the API-key or OAuth flow then probes), `Probe`, and `Remove` (with preflight analysis of affected routes/profiles). Probing runs live when the overlay opens or when explicitly requested. Target creation is initiated via `clio-coder targets add`.
|
|
226
297
|
|
|
227
|
-
Settings → Fleet is an entity workbench organized with dim group headers (`Defaults`, `Profiles`, `Agent routes`, `Placement`). Dispatched worker defaults and profile rows render as compact summaries (`fast-local node-a/example-coder-model high auto`), drilling into fields (`target`, `model`, `thinkingLevel`, `node`) on `Enter`. Profile removal is a named destructive action with affected-route preflight. Running and retrying dispatches live in the `Alt+W` Fleet Runs board, which also steers and cancels them.
|
|
298
|
+
Settings → Fleet is an entity workbench organized with dim group headers (`Defaults`, `Profiles`, `Agent routes`, `Placement`). Dispatched worker defaults and profile rows render as compact summaries (`fast-local node-a/example-coder-model high auto`), drilling into fields (`target`, `model`, `thinkingLevel`, `node`) on `Enter`. Profile removal is a named destructive action with affected-route preflight. Running and retrying dispatches live in the `Alt+W` Fleet Runs board, which also steers and cancels them. `Enter` opens the selected run's worker detail: the phase, the running call with its redacted action descriptor, and the bounded tail of the worker's own prose.
|
|
228
299
|
|
|
229
300
|
`/run` and `/delegate` put the worker's answer on screen. Both echo the typed
|
|
230
301
|
line dim above the block, then stream the run into the transcript as an attributed
|
|
@@ -256,6 +327,15 @@ operator steering whose run id names a receipt it can read, so a model that
|
|
|
256
327
|
never dispatched the run does not discard it as unattributed output. A turn
|
|
257
328
|
that only relays a shared note does not trip the unbacked-worker-claim
|
|
258
329
|
advisory.
|
|
330
|
+
A council run shares as a council. `/share <synthesis runId>` brings the whole
|
|
331
|
+
`council-report` in as one bounded block: every final-round member's answer under
|
|
332
|
+
its roster label, each with its verdict when it declared one, then the synthesis
|
|
333
|
+
line naming the mode, the verdict, the tally, and the judge run when there was
|
|
334
|
+
one. `/share <member runId>` brings that one member's answer in under its roster
|
|
335
|
+
label, so a single voice never reaches the main agent as an unattributed one. A
|
|
336
|
+
synthesis run whose sealed text does not parse as a report is shared verbatim
|
|
337
|
+
rather than dropped, because the operator named that run.
|
|
338
|
+
|
|
259
339
|
`/new` resets the transcript and the pool bare `/share` draws from, so a run
|
|
260
340
|
from the previous session cannot be shared into the new one. Worker tool
|
|
261
341
|
arguments never cross at all: the transcript carries tool names only, the same
|
|
@@ -350,7 +430,7 @@ editor reserves and can be rebound through `settings.yaml.keybindings`.
|
|
|
350
430
|
| `Alt+U` | Toggle the footer dashboard between compact (quiet 2-zone) and expanded (4-zone urgency) layouts. |
|
|
351
431
|
| `Alt+L` | Open the model and targets selector. |
|
|
352
432
|
| `Alt+J` / `Alt+K` | Cycle forward / backward through the scoped model set (when empty, displays a notice directing the operator to `/scoped-models`). |
|
|
353
|
-
| `Alt+W` | Toggle the Fleet Runs board (task, run ID, live telemetry, retry, and terminal history). |
|
|
433
|
+
| `Alt+W` | Toggle the Fleet Runs board (task, run ID, live telemetry, retry, and terminal history). Inside it, `Enter` opens the selected run's live worker detail, `s` steers, and `x` cancels. |
|
|
354
434
|
| `Alt+B` | Open the composite session and operator task board (`/tasks`). Approved application-boundary override of editor word-back. |
|
|
355
435
|
| `Alt+D` | Open the settled interview decision board (`/decisions`). Approved application-boundary override of editor word-delete. |
|
|
356
436
|
| `Alt+S` / `Ctrl+Alt+B` | Convert an active attached dispatch to a detached background batch. |
|
|
@@ -415,7 +495,9 @@ Tool and command execution is governed by:
|
|
|
415
495
|
- **Safety Net:** Granular rule packs loaded from `damage-control-rules.yaml`, project policies, and protected artifact paths; always on, identical at every autonomy level.
|
|
416
496
|
- **Autonomy Mapping:** Once the net passes a call, the level decides whether it runs, asks, or is denied. See [safety-model.md](safety-model.md) for the full matrix.
|
|
417
497
|
|
|
418
|
-
When an action asks for confirmation, whether from a safety-net rail or from the autonomy level, the call parks and the
|
|
498
|
+
When an action asks for confirmation, whether from a safety-net rail or from the autonomy level, the call parks and three surfaces say so at once. The transcript row reads `⏸ awaiting approval` with `action ·`, `axis ·`, and `target ·` lines under it; the footer phase pill reads `⏸ confirm`; and a consequence-tier dialog opens with the tool, target, action, authenticated requester, one-shot authority, reversibility, and deny and stop effects. Titles distinguish workspace authority, outward consequences, safety-net confirmation, system changes, and worker escalations. The dialog sits at bottom center with five rows reserved for the composer and footer, and it re-anchors on resize. The composer rail switches to `CONFIRM` and repeats the keys while the prompt owns the keyboard.
|
|
499
|
+
|
|
500
|
+
The keys are the same on both surfaces: `Enter` allows this one call, `Esc` denies it, and `s` denies it and stops the turn so nothing asks again. `Enter` allows only from an empty composer. While the composer holds a draft, the habitual send key does nothing, the rail and the dialog footer read `[Backspace] clear draft` instead of `[Enter] allow`, and only the deletion keys (`Backspace`, `Delete`, `Ctrl+U`, `Ctrl+W`, `Ctrl+K`) reach the editor until the draft is gone. Every other key is swallowed. A call that parks while another overlay holds the screen is announced with an `[approval]` notice and the dialog opens as soon as that overlay closes; the dialog lays itself out for any terminal width, so no width is too narrow for it. Approving or denying never changes the level.
|
|
419
501
|
|
|
420
502
|
Notice vocabulary, one prefix per mechanism: `[safety-net]` for level-independent blocks, `[approval]` for parked calls, `[autonomy]` for read-only denials, and `[middleware]` for hook diagnostics.
|
|
421
503
|
|
|
@@ -465,7 +547,7 @@ to execute through the existing engine worker path, the sanctioned Claude Code w
|
|
|
465
547
|
| --- | --- |
|
|
466
548
|
| `npm run ci` | Local and GitHub PR gate: typecheck, lint, skills pin check, build, the deterministic test suite, and the trace-viewer suite. |
|
|
467
549
|
| `npm run ci:release` | Maintainer release gate: `npm run ci`, then the `check-release` dist and packaging audit. |
|
|
468
|
-
| `npm run live:smoke -- --target <id>` | One real headless turn against a configured target. Add `--delegation` for the `opencode` and `copilot` ACP agents. The other operator-run drivers (`live:
|
|
550
|
+
| `npm run live:smoke -- --target <id>` | One real headless turn against a configured target. Add `--delegation` for the `opencode` and `copilot` ACP agents. The other operator-run drivers (`live:fleet-dispatch`, `live:tui`, `live:home`) are listed in `benchmarks/internal/README.md`. |
|
|
469
551
|
| `npm run typecheck` | Strict TypeScript pass. |
|
|
470
552
|
| `npm run lint` | Biome checks plus `scripts/check-hygiene.ts`, which runs the boundary invariants, the skills pin check, and the README and docs drift rules. |
|
|
471
553
|
| `npm run test` | Contract and smoke tests through the sharded runner. |
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Configuration, Targets, Runtimes, and Auth
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive configuration validator, target resolver, and CLI command generator is located at [docs/html/configuration_blueprint.html](html/configuration_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive configuration validator, target resolver, and CLI command generator is located at [docs/html/configuration_blueprint.html](html/configuration_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
Clio Coder is target-first: chat and fleet dispatch resolve through configured targets in `settings.yaml`, not through provider-specific ad hoc flags. Chat and print targets are HTTP and native engine-backed runtimes. Fleet dispatch can also target the sanctioned Claude Code subscription runtimes described below.
|
|
7
7
|
|
|
@@ -31,6 +31,8 @@ Default config file:
|
|
|
31
31
|
|
|
32
32
|
Role contents: config holds user-authored files (settings, credentials, agents, skills, prompts, extensions, runtimes); data holds durable artifacts (memory, evidence, evals); state holds machine-produced session state (sessions, audit, receipts, runs.json, recent-models.json, install.json, interop.json, interviews, scratch); cache holds disposable derived files.
|
|
33
33
|
|
|
34
|
+
The `library` settings block configures the private resource catalog. `library.catalog` is an optional path and defaults to `<configDir>/library.yaml`. `library.remote` is an optional git remote URL, and the catalog repository must name that git remote `library`. `library.sync` defaults to `false`, which makes sync and push refuse before spawning git. `library.confirmedRemote` is written by `clio-coder library remote confirm <url>` and must exactly match `library.remote` before sync or push can run. Confirmation sets both values when `library.remote` is unset and refuses a differing configured URL with `library_remote_mismatch`. See [resource-library.md](resource-library.md).
|
|
35
|
+
|
|
34
36
|
`clio-coder paths --json` prints the resolved directories and is the single source of truth for scripts.
|
|
35
37
|
|
|
36
38
|
---
|
|
@@ -174,6 +176,17 @@ workers:
|
|
|
174
176
|
model: your-model-id
|
|
175
177
|
thinkingLevel: off
|
|
176
178
|
profiles: {}
|
|
179
|
+
rosters:
|
|
180
|
+
design:
|
|
181
|
+
members:
|
|
182
|
+
- label: local-a
|
|
183
|
+
target: local-lmstudio
|
|
184
|
+
model: your-model-id
|
|
185
|
+
thinking: medium
|
|
186
|
+
color: accent
|
|
187
|
+
- label: local-b
|
|
188
|
+
target: local-vllm
|
|
189
|
+
color: "#5ba8ff"
|
|
177
190
|
agentBindings: {}
|
|
178
191
|
maxRetries: 2
|
|
179
192
|
onPermission: deny
|
|
@@ -208,6 +221,11 @@ terminal:
|
|
|
208
221
|
tuiMode: regular # regular terminal scrollback or fullscreen sticky layout
|
|
209
222
|
fullscreenScrollbar: auto # hidden, auto, or always in fullscreen mode
|
|
210
223
|
smoothStreaming: off # off, conservative auto, or explicit on
|
|
224
|
+
notify: false # content-free desktop notification, interactive TTY only
|
|
225
|
+
watchdog:
|
|
226
|
+
enabled: false # opt-in read-only review of every mutating turn
|
|
227
|
+
# target: local-lmstudio # route the review at a cheap model
|
|
228
|
+
# cadenceToolCalls: 20 # also review every N tool calls inside a turn
|
|
211
229
|
skills:
|
|
212
230
|
trustProjectCompatRoots: false
|
|
213
231
|
delegation:
|
|
@@ -306,6 +324,17 @@ LM Studio can require bearer authentication for its HTTP APIs
|
|
|
306
324
|
|
|
307
325
|
A model id on an LM Studio target is resolved against that host's loaded instances. A key with a loaded instance is never sent bare (which would JIT-load a second copy). An instance id reported loaded by two configured LM Studio targets on different hosts is an LM Link peer projection. When a bare model key is requested and multiple instances of it are loaded, Clio selects an instance in this order: the target's configured `defaultModel`, then an instance not cross-listed by another configured LM Studio target, and finally the first loaded instance. This behavior tracks issue #113.
|
|
308
326
|
|
|
327
|
+
When the selected instance is also loaded on a peer, a request may be answered by that peer (#185). Clio separates the requested model id, the response observation, and the model id used for accounting. Every new assistant ledger entry carries `responseModelIdObservation` in one of these explicit shapes:
|
|
328
|
+
|
|
329
|
+
| State | Meaning | Accounting attribution |
|
|
330
|
+
| --- | --- | --- |
|
|
331
|
+
| `{ "state": "reported", "reportedModelId": "<id>" }` | Clio observed an OpenAI-compatible event stream and the provider reported a model id. | The reported id. |
|
|
332
|
+
| `{ "state": "not-reported" }` | Clio observed the event stream and it contained no model id. | `unknown`, because the provider did not identify the responding model. |
|
|
333
|
+
| `{ "state": "not-observed" }` | This provider path did not expose response model-id presence to the stream tap. | A differing `responseModel` when available, otherwise the requested model id. |
|
|
334
|
+
| `{ "state": "legacy-difference-only", "differingModelId": "<id>" }` or the same shape with `null` | The ledger predates #193 and recorded only whether the response `model` differed from the request. This state is produced while reading historical rows; new rows do not write it. | The historical differing id when available, otherwise the requested model id. |
|
|
335
|
+
|
|
336
|
+
The adapter retains `responseModel` as the differing response id because providers outside the stream tap still supply that fact. `clio-coder usage report` emits `attributedModelId`, `requestedModelIds`, and `responseModelIdObservationCounts`. Its text table and the `/cost` overlay use the labels `attributed model`, `requested model ids`, and `response model id observation`; requested ids are printed as ids rather than as `same`. The footer's last-turn line uses `response model id observation <state>`, with the id after `reported` or a historical `legacy difference-only` state. Dispatch receipt `upstreamResponses` entries carry `requestedModelId`, `responseModelIdObservation`, `differingResponseModelId`, and `providerResponseId`. The peer warning is said once per process per distinct fact (target, requested id, resolved instance, peer set), not once per turn.
|
|
337
|
+
|
|
309
338
|
|
|
310
339
|
Prompt-template overrides, system prompts, GPU-offload ratios, KV-cache quantization, parallel slots,
|
|
311
340
|
context checkpoints, and speculative-decoding variants are not writable through this Clio settings
|
|
@@ -456,7 +485,8 @@ The Settings Center organizes all configuration under four non-selectable group
|
|
|
456
485
|
| **RUNTIME** | Budget (`budget`) | `budget.sessionCeilingUsd`, `defaults.maxTokens`, and `budget.concurrency` (restart required). |
|
|
457
486
|
| **RUNTIME** | Compaction (`compaction`) | `compaction.auto`, `compaction.threshold`, and `compaction.excludeLastTurns`. |
|
|
458
487
|
| **RUNTIME** | Retry (`retry`) | `retry.enabled`, `retry.maxRetries`, `retry.baseDelayMs`, and `retry.maxDelayMs`. |
|
|
459
|
-
| **EXPERIENCE** | Terminal (`terminal`) | `terminal.showTerminalProgress`, `terminal.outputVerbosity` (`minimal`, `default`, `verbose`), `terminal.tuiMode` (`regular`, `fullscreen`), `terminal.fullscreenScrollbar` (`hidden`, `auto`, `always`), `terminal.smoothStreaming` (`off`, `auto`, `on`), and `theme`. |
|
|
488
|
+
| **EXPERIENCE** | Terminal (`terminal`) | `terminal.showTerminalProgress`, `terminal.outputVerbosity` (`minimal`, `default`, `verbose`), `terminal.tuiMode` (`regular`, `fullscreen`), `terminal.fullscreenScrollbar` (`hidden`, `auto`, `always`), `terminal.smoothStreaming` (`off`, `auto`, `on`), `terminal.notify`, and `theme`. |
|
|
489
|
+
| **EXPERIENCE** | Watchdog (`watchdog`) | `watchdog.enabled`, `watchdog.target`, and `watchdog.cadenceToolCalls`. The two optional keys are editable text rows that render their absence as `(session target)` and `(turn end only)`; submitting an empty value removes the key from `settings.yaml` rather than storing a blank. |
|
|
460
490
|
| **EXPERIENCE** | Advanced (`advanced`) | `runtimePlugins`, `attribution.gitCommits`, `compaction.model`, `compaction.systemPrompt`, `delegation.defaults.connectTimeoutMs`, `delegation.defaults.turnTimeoutMs`, `delegation.defaults.permissionTimeoutMs`, `keybindings`, and `delegation.agents`. |
|
|
461
491
|
|
|
462
492
|
`retry.streamStallMs` has no Settings Center row; edit it in `settings.yaml`.
|
|
@@ -506,6 +536,10 @@ Label to config path mapping:
|
|
|
506
536
|
| TUI mode | `terminal.tuiMode` (`regular` or `fullscreen`, restart required) |
|
|
507
537
|
| Fullscreen scrollbar | `terminal.fullscreenScrollbar` (`hidden`, `auto`, or `always`, restart required) |
|
|
508
538
|
| Smooth streaming | `terminal.smoothStreaming` (`off`, `auto`, or `on`, live) |
|
|
539
|
+
| Desktop notifications | `terminal.notify` |
|
|
540
|
+
| Turn-end watchdog | `watchdog.enabled` |
|
|
541
|
+
| Watchdog target | `watchdog.target` (blank clears the key) |
|
|
542
|
+
| Watchdog cadence (tools) | `watchdog.cadenceToolCalls` (integer ≥ 1; blank clears the key) |
|
|
509
543
|
| Theme | `theme` |
|
|
510
544
|
| Runtime plugins | `runtimePlugins` |
|
|
511
545
|
| Clio commit provenance | `attribution.gitCommits` (`enabled` or `disabled`, live) |
|
|
@@ -544,6 +578,15 @@ These are saved defaults, not a live control surface. See [Live routing vs saved
|
|
|
544
578
|
|
|
545
579
|
### Safety and worker policy
|
|
546
580
|
|
|
581
|
+
`workers.rosters.<name>.members` defines council membership beside
|
|
582
|
+
`workers.profiles`. Every member accepts `label`, `target`, and the optional
|
|
583
|
+
keys `model`, `thinking`, and `color`. Labels must match
|
|
584
|
+
`[a-z][a-z0-9_-]{0,31}` and must be unique inside the roster. A roster contains
|
|
585
|
+
two to five members. Colors accept a theme token such as `accent`, `success`,
|
|
586
|
+
or `reason`, or a six-digit hexadecimal value such as `#5ba8ff`. Unknown roster
|
|
587
|
+
and member keys are rejected during configuration load. The existing settings
|
|
588
|
+
watcher validates and publishes roster changes with every other hot reload.
|
|
589
|
+
|
|
547
590
|
| Key | Default | Validation | When it applies |
|
|
548
591
|
| --- | --- | --- | --- |
|
|
549
592
|
| `autonomy` | `auto-edit` | `read-only`, `suggest`, `auto-edit`, `full-auto` | immediately |
|
|
@@ -553,8 +596,13 @@ These are saved defaults, not a live control surface. See [Live routing vs saved
|
|
|
553
596
|
| `workers.maxRetries` | `2` | integer ≥ 0 | next dispatch |
|
|
554
597
|
| `workers.resilienceCooldownMs` | `15000` | integer ≥ 0 | next dispatch |
|
|
555
598
|
| `workers.profiles` | `{}` | map of profile name to a target/model/thinking choice | next dispatch |
|
|
599
|
+
| `workers.rosters` | `{}` | map of roster name to 2 to 5 council members | next dispatch |
|
|
556
600
|
| `workers.agentBindings` | `{}` | map of agent id to a key present in `workers.profiles` | next dispatch |
|
|
557
601
|
| `skills.trustProjectCompatRoots` | `false` | boolean | restart |
|
|
602
|
+
| `library.catalog` | `null` | string or null | immediately |
|
|
603
|
+
| `library.remote` | `null` | string or null | immediately |
|
|
604
|
+
| `library.confirmedRemote` | `null` | string or null | immediately |
|
|
605
|
+
| `library.sync` | `false` | boolean | immediately |
|
|
558
606
|
|
|
559
607
|
### Git commit provenance
|
|
560
608
|
|
|
@@ -612,6 +660,34 @@ Generic provider and transport errors are classified by transient retry rules, i
|
|
|
612
660
|
| `memory.intervention.maxTokens` | `400` | integer ≥ 1 | next turn |
|
|
613
661
|
| `memory.intervention.timeoutMs` | `180000` | integer ≥ 1 | next turn |
|
|
614
662
|
|
|
663
|
+
### Turn-end watchdog
|
|
664
|
+
|
|
665
|
+
| Key | Default | Validation | When it applies |
|
|
666
|
+
| --- | --- | --- | --- |
|
|
667
|
+
| `watchdog.enabled` | `false` | boolean | immediately |
|
|
668
|
+
| `watchdog.target` | unset | non-empty target id | immediately |
|
|
669
|
+
| `watchdog.cadenceToolCalls` | unset | integer ≥ 1 | immediately |
|
|
670
|
+
|
|
671
|
+
The watchdog is off by default because it spends one worker run per mutating
|
|
672
|
+
turn. With `enabled: true`, a turn that changed the tree is handed to one
|
|
673
|
+
read-only `verifier` run briefed with the turn's coalesced diff and the task
|
|
674
|
+
board's current scope. Its blockers become one transcript notice naming the
|
|
675
|
+
count and the first three failed checks, and nothing else: it never follows up,
|
|
676
|
+
never queues a turn, and never mutates. A passing report emits nothing at all. A
|
|
677
|
+
turn with no file mutations never fires it.
|
|
678
|
+
|
|
679
|
+
`watchdog.target` routes the run at a named target, which is how a cheap local
|
|
680
|
+
model reviews turns run on a subscription route; unset, the run takes the
|
|
681
|
+
session's active target. `watchdog.cadenceToolCalls: N` additionally fires the
|
|
682
|
+
watchdog after every N tool calls inside a turn, with the same diff-and-scope
|
|
683
|
+
briefing, so mid-turn scope drift is visible before the turn ends. At most one
|
|
684
|
+
watchdog run is in flight at a time; a trigger that arrives while one is running
|
|
685
|
+
is dropped and counted rather than queued. Headless and ACP runs never fire the
|
|
686
|
+
watchdog regardless of the setting, because neither has an operator reading a
|
|
687
|
+
transcript. The block has its own Settings Center section under EXPERIENCE ›
|
|
688
|
+
Watchdog; clearing the target or the cadence row removes that key from
|
|
689
|
+
`settings.yaml` rather than writing an empty value.
|
|
690
|
+
|
|
615
691
|
### Delegation
|
|
616
692
|
|
|
617
693
|
| Key | Default | Validation | When it applies |
|
|
@@ -634,10 +710,22 @@ Generic provider and transport errors are classified by transient retry rules, i
|
|
|
634
710
|
| `terminal.tuiMode` | `regular` | `regular`, `fullscreen` | restart |
|
|
635
711
|
| `terminal.fullscreenScrollbar` | `auto` | `hidden`, `auto`, `always` | restart |
|
|
636
712
|
| `terminal.smoothStreaming` | `off` | `off`, `auto`, `on` | immediately |
|
|
713
|
+
| `terminal.notify` | `false` | boolean | immediately |
|
|
637
714
|
| `modelSelector.favorites` | `[]` | list of strings | immediately |
|
|
638
715
|
| `modelSelector.recentLimit` | `12` | integer ≥ 1 | immediately |
|
|
639
716
|
| `keybindings` | `{}` | map of binding id to a key string or list of them | restart |
|
|
640
717
|
|
|
718
|
+
`terminal.notify` turns on a content-free desktop notification for the three
|
|
719
|
+
moments an operator is waiting: a turn ends, a detached fleet batch settles, and
|
|
720
|
+
a worker permission or `ask_user` request parks. The payload is fixed. The title
|
|
721
|
+
is always `clio-coder` and the body comes from a closed vocabulary (`turn
|
|
722
|
+
finished`, `batch <shortId> settled`, `approval needed`), so no prompt text, file
|
|
723
|
+
path, or model output ever leaves the process in a notification. Clio emits OSC
|
|
724
|
+
777 by default and OSC 9 on iTerm2, Windows Terminal, and ConEmu, never both for
|
|
725
|
+
one event. Headless, ACP, and non-TTY runs never emit one regardless of the
|
|
726
|
+
setting. The knob has a Settings Center row under EXPERIENCE › Terminal,
|
|
727
|
+
labeled `Desktop notifications`.
|
|
728
|
+
|
|
641
729
|
Recently selected models are runtime state and live in `recent-models.json` under the state directory, not here. A `state.recentModels` key in `settings.yaml` is an unknown-key error.
|
|
642
730
|
|
|
643
731
|
### Structural and catalog keys
|
package/docs/context-engine.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Context Engine
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/context_blueprint.html](html/context_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
Clio Coder tracks context pressure, records per-turn snapshots, and protects the provider context with bounded tool results plus single-threshold compaction.
|
|
7
7
|
|
|
@@ -19,6 +19,8 @@ Local-native runtimes use a recommended minimum desired window of 128,000 tokens
|
|
|
19
19
|
|
|
20
20
|
The `/context` overlay states which layer answered, next to the token total: `loaded`, `probed`, `configured`, `declared`, or `assumed`.
|
|
21
21
|
|
|
22
|
+
A probed llama.cpp window is the share one request gets, not the server's total. llama.cpp splits `--ctx-size` evenly across `--parallel` slots unless `--kv-unified` is set, so a server started with `--ctx-size 786432 --parallel 4 --no-kv-unified` admits 196,608 tokens per request, and that is the figure autocompact and the meter plan against. The probe reads the flags (long and short forms, `-c`, `-np`, `-kvu`, and the last of `--kv-unified` or `--no-kv-unified` given) off the router's per-model status, keeps the split on the model's discovery state, and `/context` prints the derivation next to the share: `196,608 (786,432 / 4 slots)`. `clio-coder targets` does the same in its `ctx` note for the target's default model and adds a probe note naming the flags.
|
|
23
|
+
|
|
22
24
|
## Token accounting and snapshots
|
|
23
25
|
|
|
24
26
|
The estimator in `context-accounting.ts` uses a four-characters-per-token family for hot-path accounting. It estimates system prompt, tools, messages, pending input, and runtime categories without calling a model tokenizer on every TUI refresh.
|
|
@@ -190,7 +192,7 @@ In Git workspaces, the indexer uses the same visible file set across full builds
|
|
|
190
192
|
incremental updates, fingerprints, and project profiles: tracked files plus
|
|
191
193
|
untracked, unignored work in progress. It excludes symlinks, submodule gitlinks,
|
|
192
194
|
generated output, scratch space, and local-state directories such as `.git`,
|
|
193
|
-
`.clio-coder`, `.superpowers`, `.codex`, `.claude`,
|
|
195
|
+
`.clio-coder`, `.superpowers`, `.codex`, `.claude`, `node_modules`,
|
|
194
196
|
`dist`, `build`, `coverage`, virtualenvs, `target`, and `vendor`. Non-Git
|
|
195
197
|
workspaces use a bounded filesystem walk with the same directory exclusions.
|
|
196
198
|
Source coverage spans TypeScript, JavaScript, Python, Rust, Go, C, C++, CUDA
|
|
@@ -5,7 +5,7 @@ The working set is the part of the session ledger the model actually receives on
|
|
|
5
5
|
Source of truth is `src/domains/context/working-set/` (`contract.ts`, `fold.ts`, `project.ts`, `marker.ts`, `protect.ts`, `engine.ts`, `recall.ts`, `policies/`), the ledger records in `src/domains/session/entries.ts`, and the compaction stage in `src/interactive/turn-context.ts` (`runAutoCompact`).
|
|
6
6
|
|
|
7
7
|
> [!WARNING]
|
|
8
|
-
> This is an experimental community alpha surface. The default policy is `structural-v1
|
|
8
|
+
> This is an experimental community alpha surface. The default policy is `structural-v1`; `age-horizon` reproduces the selection Clio made before this layer existed and stays available.
|
|
9
9
|
|
|
10
10
|
## Vocabulary
|
|
11
11
|
|
|
@@ -170,7 +170,7 @@ context:
|
|
|
170
170
|
|
|
171
171
|
## What the operator sees
|
|
172
172
|
|
|
173
|
-
- **`/context` overlay.** A working-set section under the category legend: the policy
|
|
173
|
+
- **`/context` overlay.** A working-set section under the category legend: the configured policy with its state (`policy structural-v1 · no events yet` until the first event, `disabled` when `context.workingSet.enabled` is off, and `(last event by <policy>)` when the setting changed after an event), evicted item count, evicted tokens, event count, recall count, and churn. Evicted tokens render as one line after the legend rather than as a meter category, because they are outside the window rather than a slice of it.
|
|
174
174
|
- **Transcript.** An evicted tool row keeps its full body and gains a dim `evicted · <reason>` tag. The transcript shows the ledger, never the projection, so `/resume`, `/tree`, `/fork`, and the HTML export are unaffected by eviction.
|
|
175
175
|
- **`/context recall <ref>`.** Prints the ref, why it was evicted, the token count, and the offload pointer when there is one, followed by the original body. Transcript only.
|
|
176
176
|
- **Prompt cache line.** Every applied event stamps `working_set_evict` on the next assistant entry's `promptCache.expectedColdReasons`. When the last settled run came back cold for that reason, the overlay adds `last cold turn: working-set eviction (expected)` and drops the shell-reused-but-backend-cold warning, because the cold turn is explained rather than surprising.
|
|
@@ -181,14 +181,14 @@ context:
|
|
|
181
181
|
These are tracked follow-ups, not available behavior:
|
|
182
182
|
|
|
183
183
|
- **Auto-readmission.** Nothing brings an evicted body back on its own. There are no path fingerprints and no registry of what the model is likely to need next.
|
|
184
|
-
- **Cost model and deferred scheduling.** Pressure is the only trigger. There is no break-even horizon, no deferred eviction plan, and no piggybacking beyond the fact that the working-set stage already runs first inside `runAutoCompact`.
|
|
184
|
+
- **Cost model and deferred scheduling.** Pressure is the only trigger, and it is `compaction.threshold`, not `target`. The replay tables price every applied event by the cold prefix it re-prefills (about 29k tokens per event at a 64k budget), and batching from the threshold down to the target is what keeps one event per cycle; a trigger at the target would make every turn above 60% with one newly redundant read an event of its own, and no row in the sweep shows fewer summaries in return. There is no break-even horizon, no deferred eviction plan, and no piggybacking beyond the fact that the working-set stage already runs first inside `runAutoCompact`.
|
|
185
185
|
- **Intra-turn eviction.** Eviction runs before a request is sent. A single turn whose tool results overflow the window is handled by the observation envelope's caps and by summary compaction, not by this layer.
|
|
186
186
|
- **Worker runtimes.** Dispatched workers replay their own ledgers without the working-set stage.
|
|
187
187
|
- **Digests.** A marker carries tool, size, and a first-line preview. The generated summaries from #165 are not embedded in it.
|
|
188
188
|
|
|
189
189
|
## See also
|
|
190
190
|
|
|
191
|
-
- `clio-coder context replay --sessions <path>...` replays Clio ledgers, and `--synthetic <ids>` replays the seeded procedural corpora, through the same fold, projection, and policy code with `none`, `random`, and `oracle` controls; `clio-coder context working-set --session <id|path>` prints one session's fold and path index. Both are described under [Working-set replay](commands-and-modes.md#working-set-replay)
|
|
191
|
+
- `clio-coder context replay --sessions <path>...` replays Clio ledgers, and `--synthetic <ids>` replays the seeded procedural corpora, through the same fold, projection, and policy code with `none`, `random`, and `oracle` controls; `clio-coder context working-set --session <id|path>` prints one session's fold and path index. Both are described under [Working-set replay](commands-and-modes.md#working-set-replay). Generated replay tables are local artifacts rather than versioned benchmark results.
|
|
192
192
|
- [context-engine.md](context-engine.md) for context window resolution, token accounting, and how this stage sits ahead of summary compaction.
|
|
193
193
|
- [session-lifecycle.md](session-lifecycle.md) for the ledger format, active-path lineage, and branching.
|
|
194
194
|
- [glossary.md](glossary.md) for the one-line definitions of these terms.
|
|
@@ -80,7 +80,7 @@ New `area:*` labels are proposed in an issue, not created ad hoc.
|
|
|
80
80
|
|
|
81
81
|
## Milestones are releases
|
|
82
82
|
|
|
83
|
-
Each open milestone is the next version (`v0.3.
|
|
83
|
+
Each open milestone is the next version (`v0.3.7`, `v0.4.0`). Triage means
|
|
84
84
|
assigning an issue to a milestone or explicitly leaving it in the backlog.
|
|
85
85
|
A release cut requires every issue in its milestone to be closed
|
|
86
86
|
or bumped; the milestone closes when the tag is published.
|
|
@@ -28,7 +28,7 @@ split would use. They cross them.
|
|
|
28
28
|
| Write-boundary attribution is per scheduling *window*, so the compiler refuses a wave with two writers | scheduling, write boundaries, plan compilation | `execution-plan.ts`, `write-boundary.ts` |
|
|
29
29
|
| A loop's later nodes are `unneeded`, decided by the scheduler, not the plan | plan compilation, scheduling, receipts | `fleet-plan.ts`, `execution-scheduler.ts` |
|
|
30
30
|
| Staleness revalidation re-runs a verification a later workspace step invalidated | scheduling, plan compilation, code steps | `execution-scheduler.ts` |
|
|
31
|
-
| Receipt integrity
|
|
31
|
+
| Receipt integrity v16 seals normalized routing intent | routing, receipts | `receipt-integrity.ts`, `routing-intent.ts` |
|
|
32
32
|
|
|
33
33
|
The write-boundary and loop rows are the sharpest. Both are properties of a
|
|
34
34
|
*wave*, which is a scheduling concept computed by the plan compiler and enforced
|
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Clio Coder Documentation Coverage Matrix
|
|
2
2
|
|
|
3
|
-
This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.
|
|
3
|
+
This matrix maps every top-level directory in `src/` and every domain directory under `src/domains/` to its authoritative documentation page. It records coverage status (`documented`, `partial`, `undocumented`), missing concepts, and key source contracts for `v0.3.7`.
|
|
4
4
|
|
|
5
5
|
## Coverage Matrix
|
|
6
6
|
|
|
@@ -19,8 +19,8 @@ This matrix maps every top-level directory in `src/` and every domain directory
|
|
|
19
19
|
| `src/domains/components/` | Component scanning, snapshots, hashing, diffing | [middleware-and-components.md](middleware-and-components.md) | `documented` | Documented in active component snapshot and middleware guide. |
|
|
20
20
|
| `src/domains/config/` | Configuration contracts, file watcher, keybinding definitions, setting classifiers | [configuration-and-targets.md](configuration-and-targets.md), [commands-and-modes.md](commands-and-modes.md) | `documented` | Documented in configuration targets and command/keybinding reference. |
|
|
21
21
|
| `src/domains/context/` | `CLIO-CODER.md` bootstrap, codewiki generation, prompt context assembly, project rules, non-destructive working-set eviction (`age-horizon` and `structural-v1` policies, protection predicates, path index, byte-stable markers, recall by ref) | [context-engine.md](context-engine.md), [context-working-set.md](context-working-set.md) | `documented` | Context window, token accounting, and the three compaction mechanisms in the engine reference; the working-set layer has its own guide covering the vocabulary, both ledger record kinds and format v4, the marker contract, both policies with their rule order, recall semantics, and the operator surfaces. |
|
|
22
|
-
| `src/domains/dispatch/` | Fleet orchestration, assignment store, batch tracker, admission, route planner, receipt integrity
|
|
23
|
-
| `src/domains/eval/` | Suite v2 YAML schema, eval runner, metrics, reporters, workspace sandboxing | [eval-runner.md](eval-runner.md), [evals-internal.md](evals-internal.md) | `documented` |
|
|
22
|
+
| `src/domains/dispatch/` | Fleet orchestration, assignment store, batch tracker, admission, route planner, receipt integrity v16 | [fleet-dispatch.md](fleet-dispatch.md), [dispatch-architecture-rationale.md](dispatch-architecture-rationale.md), [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `documented` | Multi-node fleet dispatch, admission invariants, and receipt verification fully documented. |
|
|
23
|
+
| `src/domains/eval/` | Suite v2 YAML schema, eval runner, metrics, reporters, workspace sandboxing | [eval-runner.md](eval-runner.md), [evals-internal.md](evals-internal.md) | `documented` | Product evals are documented independently from external benchmarks. |
|
|
24
24
|
| `src/domains/evidence/` | Evidence bundles, findings taxonomy, provenance store, failure attribution | [evidence-and-memory.md](evidence-and-memory.md) | `documented` | Documented in evidence directory structures and memory retrieval guide. |
|
|
25
25
|
| `src/domains/evolution/` | Falsifiable Change Manifest JSON templates and `clio-coder evolve` self-edit gates | [evolution.md](evolution.md) | `documented` | Documented in evolution manifest reference and mutation validation rules. |
|
|
26
26
|
| `src/domains/extensions/` | Extension manifest schemas, resource roots, portable share archives | [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Documented in extensions and sharing guide. |
|
|
@@ -35,7 +35,7 @@ This matrix maps every top-level directory in `src/` and every domain directory
|
|
|
35
35
|
| `src/domains/scheduling/` | Capacity lease acquisition, heartbeats, expiry, cross-process locks, cluster scheduling | [capacity-and-scheduling.md](capacity-and-scheduling.md), [fleet-dispatch.md](fleet-dispatch.md) | `documented` | Dedicated capacity leasing, heartbeat TTL, and cross-process lock reference. |
|
|
36
36
|
| `src/domains/session/` | Session ledger format v4, tree branching (`/tree`), `/fork`, `/resume`, checkpoints, protected-artifact journal | [session-lifecycle.md](session-lifecycle.md), [context-working-set.md](context-working-set.md) | `documented` | Dedicated session lifecycle guide covering branching, journal, and recovery; the `contextEviction` and `contextRecall` records added at format v4 are specified in the working-set guide. |
|
|
37
37
|
| `src/domains/share/` | Portable share archive bundles, manifest verification, import/export flows | [extensions-and-sharing.md](extensions-and-sharing.md) | `documented` | Share archives and portable bundle formats documented in extensions guide. |
|
|
38
|
-
| `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.
|
|
38
|
+
| `src/domains/webhook/` | Empty directory | None (Inert) | `inert` | Directory contains no active modules or exports in v0.3.7. |
|
|
39
39
|
|
|
40
40
|
## Cross-Cutting Reference Guides
|
|
41
41
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Documentation Standards and Codebase Alignment
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive documentation link linter, phrasing/claim evaluator, and alignment portal is located at [docs/html/documentation_blueprint.html](html/documentation_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
Clio Coder is an experimental community alpha. Documentation should help contributors and early users work from the source of truth without overstating maturity. When docs drift, prefer the current source and tests over older prose or aspirational roadmap notes.
|
|
7
7
|
|
|
@@ -46,10 +46,10 @@ Classify claims clearly:
|
|
|
46
46
|
| [alcf-provider.md](alcf-provider.md) | `src/domains/providers/runtimes/cloud/alcf.ts`, `src/engine/alcf-oauth.ts` | Globus PKCE OAuth, openAuthStorage(), Sophia vLLM, Metis API, chatTemplateKwargsUnsupported. |
|
|
47
47
|
| [environment-variables.md](environment-variables.md) | `src/core/guardrails.ts`, `src/core/xdg.ts`, `src/domains/providers/knowledge-base-path.ts` | Comprehensive env var matrix: guardrail overrides, directory layout (CLIO_CODER_HOME), debug toggles, and internal plumbing. |
|
|
48
48
|
| [built-in-agents.md](built-in-agents.md) | `src/domains/agents/**`, `src/domains/agents/builtins/*.md`, `src/domains/dispatch/**` | Builtin agent recipes, discovery roots, frontmatter schema, fleet contract shadowing (`.clio-coder/fleets/<name>.md`), active route automation. |
|
|
49
|
-
| [fleet-dispatch.md](fleet-dispatch.md) | `src/domains/dispatch/**` | Multi-node SSH dispatch: process-safe admission, capacity leases, Contract v4 write boundaries (detect-and-rollback), bounded check/repair loops (`loop_bound_exhausted`), deterministic code steps, attestation, receipts
|
|
49
|
+
| [fleet-dispatch.md](fleet-dispatch.md) | `src/domains/dispatch/**` | Multi-node SSH dispatch: process-safe admission, capacity leases, Contract v4 write boundaries (detect-and-rollback), bounded check/repair loops (`loop_bound_exhausted`), deterministic code steps, attestation, receipts v16. |
|
|
50
50
|
| [capacity-and-scheduling.md](capacity-and-scheduling.md) | `src/domains/scheduling/**`, `src/domains/dispatch/capacity-lease.ts`, `src/domains/dispatch/reservation-store.ts` | Multi-process capacity leases (`dispatch-admission.json`), heartbeat TTLs, cross-process transaction locks (`dispatch-admission.json.lock`), and cluster drain controls. |
|
|
51
51
|
| [worker-dispatch-mechanics.md](worker-dispatch-mechanics.md) | `src/worker/**` | NDJSON parent-child socket protocols, control/bulk lane demuxing, watchdog timers, worker attestation (13 protocol fields), permission parking, exit codes. |
|
|
52
|
-
| [fleet-demo-runbook.md](fleet-demo-runbook.md) | `src/domains/dispatch/**` | Multi-node fleet demo: SSH setup, C++ build/repair workflow, reviewer gates, receipt verification
|
|
52
|
+
| [fleet-demo-runbook.md](fleet-demo-runbook.md) | `src/domains/dispatch/**` | Multi-node fleet demo: SSH setup, C++ build/repair workflow, reviewer gates, receipt verification v16. |
|
|
53
53
|
| [session-lifecycle.md](session-lifecycle.md) | `src/engine/session.ts`, `src/domains/session/**` | Session lifecycle, on-disk ledger format v4 (`current.jsonl`), tree branching (`tree.json`), active-path lineage selection, `/fork`, `/resume`, checkpoints, and write-ahead protected-artifact journal. |
|
|
54
54
|
| [acp.md](acp.md) | `src/engine/acp/**`, `src/cli/acp.ts` | Agent Client Protocol (ACP) server over stdio, tool mediation, non-stall permission handling, timeout bounds, and error taxonomy. |
|
|
55
55
|
| [artifact-versions.md](artifact-versions.md) | `src/domains/dispatch/receipt-integrity.ts`, `src/engine/session.ts`, `src/worker/spec-contract.ts`, `src/domains/agents/fleet-contract.ts`, `src/domains/eval/schema/`, `src/domains/observability/trace-store.ts` | Version registry and migration policies for all 9 serialized artifact schemas across Clio Coder. |
|
|
@@ -65,7 +65,7 @@ Classify claims clearly:
|
|
|
65
65
|
| [proactive-memory.md](proactive-memory.md) | `src/domains/memory/**` | Proactive task memory architecture, session task bank, intervention rules, and handoff carrying. |
|
|
66
66
|
| [trace-store.md](trace-store.md) | `src/cli/trace.ts`, `src/domains/observability/trace-store.ts` | WAL SQLite trace mirror database schema, rowid cursor queries, rebuildability, 6 `clio-coder trace` subcommands (`runs`, `phases`, `tail`, `procs`, read-only `sql` SELECT, `ui`). |
|
|
67
67
|
| [eval-runner.md](eval-runner.md) | `src/domains/eval/**`, `src/cli/eval.ts` | Local YAML eval tasks, dual token accountings (`tokens.*` wire vs `receiptUsage.*` journal), fail-closed null totals, EvalArtifactV4 format, `verify.measure` task outcome recording. |
|
|
68
|
-
| [evals-internal.md](evals-internal.md) | `src/domains/eval
|
|
68
|
+
| [evals-internal.md](evals-internal.md) | `src/domains/eval/**` | Private context index determinism and target smoke matrices. External model benchmarks are documented under `benchmarks/`. |
|
|
69
69
|
| [extensions-and-sharing.md](extensions-and-sharing.md) | `src/domains/extensions/**`, `src/domains/resources/**`, `src/domains/share/**`, `src/cli/extensions.ts`, `src/cli/share.ts` | Prompt and skill resources, extension manifests, portable share archives. |
|
|
70
70
|
| [skills-marketplace.md](skills-marketplace.md) | `src/interactive/overlays/skills-hub.ts`, `src/domains/resources/skills/marketplace.ts` | Skills Hub marketplace discovery through the install resolver, empty state, install actions, publishing flow. |
|
|
71
71
|
| [model-catalog.md](model-catalog.md) | `src/domains/providers/catalog.ts`, `src/domains/providers/models/**`, `src/domains/providers/probe/**`, `src/domains/providers/model-capabilities.ts` | Model catalog, live probes (`--offline` toggle), exact-id selector `probeCapabilitiesForModel`, field-note promotion. |
|
package/docs/eval-runner.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Clio Coder Local Evaluation Runner
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
7
|
|
package/docs/evals-internal.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Internal Eval Suites
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:**
|
|
4
|
+
> **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.7).
|
|
5
5
|
|
|
6
6
|
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
7
|
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
@@ -15,9 +15,9 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
|
|
|
15
15
|
```
|
|
16
16
|
|
|
17
17
|
Use `--out <dir>` when the artifact should be written outside the default Clio
|
|
18
|
-
data directory.
|
|
19
|
-
|
|
20
|
-
|
|
18
|
+
data directory. Product eval artifacts and external benchmark campaigns are
|
|
19
|
+
separate: public benchmark adapters live under `benchmarks/community/` and do
|
|
20
|
+
not use the eval runner.
|
|
21
21
|
|
|
22
22
|
## Context Regression Seed
|
|
23
23
|
|
|
@@ -268,44 +268,3 @@ thresholds:
|
|
|
268
268
|
value: 0
|
|
269
269
|
```
|
|
270
270
|
|
|
271
|
-
---
|
|
272
|
-
|
|
273
|
-
## Soak Benchmark Suite
|
|
274
|
-
|
|
275
|
-
The soak benchmark suite located under [`benchmarks/soak/`](../benchmarks/soak/) measures Clio's own machinery performance, integrity, and structural invariant promises under load. Unlike standard evaluation suites, the soak suite evaluates the reliability of Clio rather than model capability. A weak model that fails to solve the workload still passes the suite if Clio's machinery behaves correctly; a strong model fails the suite if Clio fails to seal a receipt, cannot authenticate a receipt, or violates a system invariant.
|
|
276
|
-
|
|
277
|
-
Every suite runs through the product's own eval runner against a configured
|
|
278
|
-
target; there is no separate soak runner:
|
|
279
|
-
|
|
280
|
-
```bash
|
|
281
|
-
npm run build
|
|
282
|
-
clio-coder eval run --suite benchmarks/soak/clio-soak.yaml \
|
|
283
|
-
--target <id> --model <wireId> --clio-coder-entry dist/cli/index.js
|
|
284
|
-
```
|
|
285
|
-
|
|
286
|
-
`tests/contracts/eval-soak-suite.test.ts` loads all four files in CI and drives
|
|
287
|
-
`clio-soak.yaml` against a stub that seals receipts on purpose, so the gate is
|
|
288
|
-
known to fail when sealing fails; the model runs themselves are operator-run.
|
|
289
|
-
|
|
290
|
-
The soak suite comprises four specialized suite files:
|
|
291
|
-
|
|
292
|
-
### 1. Machinery Under Load (`clio-soak.yaml`)
|
|
293
|
-
Evaluates the same task workload across two execution surfaces: the headless main-agent surface (`clio-run`) and a dispatched worker surface (`agent: coder`). It tests single-file bugs, multi-file bugs, and compaction continuity across restarts.
|
|
294
|
-
- **Surface Differences**: Main-agent tasks verify session ledger continuity (`ledger.formatVersion`, `ledger.toolPairsUnmatched`, `ledger.assistantBetweenCallAndResult`), while dispatch worker tasks verify process group cleanup (`process.orphanedChildren == 0`).
|
|
295
|
-
- **Compaction Continuity**: Verifies that compaction summaries are present (`continuity.compactionSummaryPresent`) and that pre-compaction facts are preserved (`continuity.answeredFromPreCompaction`).
|
|
296
|
-
- **Suite-Wide Gates**: Gates on `receipt.sealed`, `receipt.integrityValid`, `receipt.outcomeMatchesExit`, `tokens.measured`, `stream.cumulativeSnapshots == 0`, `stream.usageDoubleCounted == false`, and `stream.segmentUsageMatchesMessages == true`.
|
|
297
|
-
|
|
298
|
-
### 2. Per-Step Write Boundaries (`clio-soak-boundary.yaml`)
|
|
299
|
-
Validates write boundary enforcement across steps without model participation. Enforcement is strictly detect-and-rollback and is never sandboxing.
|
|
300
|
-
- `write-boundary.rolled-back`: Verifies clean detection of allowlist violations (`writes_boundary_violation`), git-level file restoration, and sealed verdict generation (`boundary.violationsRolledBack == 1`, `boundary.rollbackIncomplete == 0`).
|
|
301
|
-
- `write-boundary.rollback-incomplete`: Tests honest failure reporting when a path was dirty prior to snapshot taking. prior bytes exist only in the overwritten tree, so rollback leaves the tree unchanged and records incomplete rollback (`boundary.rollbackIncomplete == 1`, `boundary.violationsRolledBack == 0`).
|
|
302
|
-
|
|
303
|
-
### 3. Fault Injection Chaos (`clio-soak-chaos.yaml`)
|
|
304
|
-
Evaluates system resilience against process signals.
|
|
305
|
-
- `chaos.sigint-mid-tool`: Prompts Clio for a long-running bash tool call and injects `SIGINT` once the subprocess initializes. Asserts exit code `130`, confirms no orphaned children remain (`process.orphanedChildren == 0`), and verifies receipt sealing, receipt integrity, and provider token reporting.
|
|
306
|
-
|
|
307
|
-
### 4. Bounded Loops (`clio-soak-loop.yaml`)
|
|
308
|
-
Validates iteration bounds and receipt accounting for fleet loops (`bounded-loop.fleet`).
|
|
309
|
-
- **Loop Bounds**: Asserts that verification attempts do not exceed declared limits (`loop.attemptsSpent <= 3`), recovery attempts seal individual receipts (`loop.receiptsMatchRepairs == true`), and unneeded nodes report as `unneeded` rather than skipped or failed (`loop.skippedNodes == 0`).
|
|
310
|
-
- **Two Token Accountings**: Distinguishes `tokens.*` (folded live off wire stdout by `createStreamInvariantFold`) from `receiptUsage.*` (journal receipts sealed and authenticated against ledger envelopes). On fleet runs, wire streaming is absent (`tokens.measured == false`), while journal receipts provide authenticated usage (`receiptUsage.measured == true`).
|
|
311
|
-
|