@iowarp/clio-coder 0.3.8 → 0.3.9
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +49 -0
- package/README.md +7 -3
- package/dist/{acp-U67UHUK2.js → acp-7LOELQFP.js} +6 -6
- package/dist/{agents-YU6SGALZ.js → agents-FIBG2SHA.js} +27 -25
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-5ZPJOIVG.js → auth-OI4LIH2I.js} +11 -12
- package/dist/{builtins-C6JMZVV6.js → builtins-AD25UL3C.js} +5 -5
- package/dist/{chunk-5FR74PWO.js → chunk-2JDWVJND.js} +2 -2
- package/dist/chunk-3DPEIQKN.js +113 -0
- package/dist/{chunk-4SPRNWDE.js → chunk-3DUR4WUA.js} +15 -15
- package/dist/{chunk-VHN4MY6O.js → chunk-3MRC2YSQ.js} +2 -2
- package/dist/{chunk-5DHKRSMQ.js → chunk-3UUY7R3Z.js} +11 -7
- package/dist/{chunk-IGWKHNIQ.js → chunk-3V5AYSEQ.js} +8 -8
- package/dist/{chunk-FHJEP5SW.js → chunk-465CC7FK.js} +8 -5
- package/dist/{chunk-5WIGXA4T.js → chunk-47CMYGET.js} +111 -4
- package/dist/{chunk-A3WNZD3P.js → chunk-4H6ULJ3H.js} +67 -21
- package/dist/{chunk-TB5666IT.js → chunk-4LJX2PUC.js} +3 -3
- package/dist/{chunk-XWSF374K.js → chunk-56KB5IJP.js} +2 -2
- package/dist/{chunk-SPULKLCF.js → chunk-5DQRIYDZ.js} +2 -2
- package/dist/{chunk-2HEJ2F35.js → chunk-5HFBWUMU.js} +20 -8
- package/dist/{chunk-WNIJTQQK.js → chunk-5PVQ4SRS.js} +78 -6
- package/dist/{chunk-DYIM5TJT.js → chunk-5QKCQQ3E.js} +262 -6
- package/dist/{chunk-TYPGUK6W.js → chunk-5T7RBWN2.js} +111 -5
- package/dist/{chunk-IIZWH4XA.js → chunk-774ILSRL.js} +2 -2
- package/dist/chunk-7C6RYZGQ.js +391 -0
- package/dist/{chunk-TANS5ZJS.js → chunk-AD7Y7STJ.js} +3 -3
- package/dist/{chunk-RWSI4YD7.js → chunk-AEYBF3TB.js} +33 -12
- package/dist/{chunk-DGSYXYMX.js → chunk-AMKHQW3C.js} +2 -2
- package/dist/{chunk-VWZOAB7K.js → chunk-B5XRQOLB.js} +7 -7
- package/dist/{chunk-WXY7KU3G.js → chunk-BVDVID7E.js} +2 -2
- package/dist/{chunk-PMDBGQSJ.js → chunk-CA42X6KT.js} +2 -2
- package/dist/{chunk-HLE42MG7.js → chunk-D73KXYPF.js} +3 -3
- package/dist/{chunk-5Q2VVUKB.js → chunk-DG4M6ZUE.js} +3 -3
- package/dist/{chunk-ME6CCNFO.js → chunk-EBOC7MT3.js} +6 -6
- package/dist/{chunk-MXKJU4JB.js → chunk-ECUO3KDP.js} +48 -7
- package/dist/{chunk-26LEYJZH.js → chunk-FALJGAWU.js} +2 -2
- package/dist/{chunk-GWS3VEIW.js → chunk-FWDFM5ZU.js} +24 -3
- package/dist/{chunk-7RGZWPB6.js → chunk-GAYUJ7LE.js} +67 -13
- package/dist/{chunk-VCBR6CU7.js → chunk-HAY4ZE2P.js} +2 -2
- package/dist/{chunk-FBVTI2TJ.js → chunk-HCBCAYZU.js} +11 -130
- package/dist/{chunk-WSB3FPX7.js → chunk-HJB5IUKP.js} +32 -136
- package/dist/{chunk-J3YUBZWY.js → chunk-HKO36JWF.js} +33 -5
- package/dist/{chunk-E77JEWSD.js → chunk-HPCTNZM2.js} +6 -36
- package/dist/{chunk-FJ3H4MN5.js → chunk-IHKBWSXF.js} +2 -2
- package/dist/{chunk-NMPKI6XL.js → chunk-JEQQR47K.js} +37 -18
- package/dist/{chunk-3BINW3FP.js → chunk-KV2AOLDF.js} +24 -4
- package/dist/{chunk-YS5VLNH5.js → chunk-LXPJXFM5.js} +7 -7
- package/dist/{chunk-7RFXX52T.js → chunk-MIX5N5AC.js} +271 -41
- package/dist/{chunk-K4XHGFR5.js → chunk-MLOK6ZOS.js} +1297 -218
- package/dist/{chunk-ZNLWCMVZ.js → chunk-MV2VUEJC.js} +2 -2
- package/dist/{chunk-MVVUPGPW.js → chunk-MXHC5QYU.js} +6 -6
- package/dist/{chunk-5H3GB5BO.js → chunk-N3PBVRTZ.js} +4 -382
- package/dist/{chunk-2HFZQUHL.js → chunk-N5XKWMDW.js} +17 -7
- package/dist/{chunk-5C3AQNDW.js → chunk-NNNWO6F2.js} +124 -36
- package/dist/{chunk-GU2UIAFZ.js → chunk-NQ6UCCOD.js} +3 -3
- package/dist/{chunk-N22QMJKY.js → chunk-NZU6YDNV.js} +4 -4
- package/dist/{chunk-ZVJ5BLO2.js → chunk-O6I4CIEU.js} +151 -13
- package/dist/{chunk-XN3L4EYL.js → chunk-OEDBCISO.js} +2 -2
- package/dist/{chunk-VAWNZU7Z.js → chunk-P3JGPQFL.js} +2 -2
- package/dist/{chunk-IJ7RPIYJ.js → chunk-PNY46YEY.js} +20 -3
- package/dist/{chunk-U6MBIEMB.js → chunk-PZ4I4JE2.js} +56 -35
- package/dist/{chunk-GPIEI3LY.js → chunk-QQ7EKM72.js} +2 -2
- package/dist/{chunk-WLFILSD5.js → chunk-R7LNVMCS.js} +66 -28
- package/dist/{chunk-GOXNB3AO.js → chunk-RAPCMZL4.js} +75 -4
- package/dist/chunk-RKKLTLYB.js +45 -0
- package/dist/{chunk-TT36MB5S.js → chunk-RKRLDWD3.js} +3 -1
- package/dist/{chunk-TTHACPOM.js → chunk-S4COXYBG.js} +456 -18
- package/dist/{chunk-WWCZ5F23.js → chunk-T3Z6VAAF.js} +69 -10
- package/dist/{chunk-GN57SG4G.js → chunk-TD7UE2L5.js} +9 -7
- package/dist/{chunk-TLQJPP24.js → chunk-TEO2TLVN.js} +523 -322
- package/dist/{chunk-FCSXB6T2.js → chunk-UOSL25KY.js} +14 -2
- package/dist/{chunk-KTYTFRMB.js → chunk-VKBMFOYV.js} +17 -15
- package/dist/{chunk-PT7HYKEM.js → chunk-VO2LKSTM.js} +2 -2
- package/dist/{chunk-P43ETTHK.js → chunk-VPTUJU4P.js} +2 -2
- package/dist/{chunk-WEH5XRJQ.js → chunk-WIE7ZOSW.js} +2 -2
- package/dist/{chunk-EMYUUSFG.js → chunk-WXCJ7VME.js} +5 -5
- package/dist/{chunk-4DGYLA73.js → chunk-XDOQXGFO.js} +22 -7
- package/dist/{chunk-JOZYP4GM.js → chunk-YKOFT37S.js} +5 -5
- package/dist/chunk-YSEHGPCT.js +127 -0
- package/dist/{chunk-HFSBBKSQ.js → chunk-YW7UVM5V.js} +138 -3
- package/dist/cli/index.js +27 -27
- package/dist/{clio-QVTYJ57A.js → clio-LT5V7SSZ.js} +6 -6
- package/dist/{code-nav-FGGFIE7L.js → code-nav-LMW275PA.js} +4 -4
- package/dist/{config-LW5IJFQN.js → config-RXS5T3JT.js} +71 -43
- package/dist/{configure-7XIZCOU4.js → configure-2WYWSCSD.js} +14 -15
- package/dist/{context-Y6Y7QPR6.js → context-I3BTOTCS.js} +12 -12
- package/dist/{context-L3WL3X7K.js → context-MVOORGMF.js} +34 -33
- package/dist/{context-N52ZA626.js → context-PALKKQYL.js} +20 -20
- package/dist/{context-clear-MBQRLSDQ.js → context-clear-N2WOYZ2K.js} +34 -33
- package/dist/{context-index-HVMFQHK3.js → context-index-HNG3MOME.js} +2 -2
- package/dist/{context-working-set-GS6DSO7F.js → context-working-set-MIEVECVZ.js} +10 -11
- package/dist/{dispatch-runner-22ZCNOM3.js → dispatch-runner-VVA4SRRH.js} +34 -33
- package/dist/doctor-TWBWFK5V.js +165 -0
- package/dist/{eval-BEC2WHDA.js → eval-IJ5VEZDJ.js} +2016 -142
- package/dist/{evidence-REJUMSKM.js → evidence-L5APPXNV.js} +29 -28
- package/dist/{evolve-PY5ZBA5K.js → evolve-RGNKFJ52.js} +29 -28
- package/dist/{extensions-HVKU65YU.js → extensions-7WYWUX5A.js} +9 -3
- package/dist/{fleet-7WZEWRFA.js → fleet-6CNVBZZP.js} +87 -55
- package/dist/{fleet-commands-UVHWM76J.js → fleet-commands-L2SXSYEI.js} +6 -6
- package/dist/{fleet-graph-6ULH7PES.js → fleet-graph-2J3OOIPO.js} +14 -12
- package/dist/{fleet-preflight-J53T6CCE.js → fleet-preflight-CZRJ4JP5.js} +3 -4
- package/dist/{fleet-validate-72PC4SLA.js → fleet-validate-C5RI6DP7.js} +16 -15
- package/dist/{init-OG3TPGQG.js → init-VBN2ACVA.js} +50 -48
- package/dist/{library-CNTMPLRF.js → library-JHGUMLY2.js} +13 -11
- package/dist/{memory-6IS7F275.js → memory-K4OQIYWG.js} +31 -30
- package/dist/{models-ENRJDA5W.js → models-2NCZUWDD.js} +23 -22
- package/dist/{monitor-XLDVO7TN.js → monitor-MMVTJABD.js} +35 -34
- package/dist/{orchestrator-6KSPYRHA.js → orchestrator-ZKBPCHW6.js} +1627 -294
- package/dist/{reset-RZ4ER727.js → reset-DD5JGOY3.js} +3 -3
- package/dist/{run-Y2CNK5RU.js → run-QEGNX7FL.js} +56 -55
- package/dist/{share-A55GYP6Z.js → share-JKD3BQMW.js} +13 -11
- package/dist/{skills-ALC5J6AT.js → skills-LMQIKDOZ.js} +14 -12
- package/dist/{skills-eval-JPBEBYQU.js → skills-eval-I7X2774U.js} +34 -32
- package/dist/{steer-GGWFUJUD.js → steer-CF5TDANS.js} +3 -3
- package/dist/{support-MIETYA5E.js → support-I7LOJLIF.js} +4 -4
- package/dist/{targets-VGNXIR3S.js → targets-RUSR6B5Z.js} +58 -32
- package/dist/{terminal-lease-WOBR64YA.js → terminal-lease-QYVORFR4.js} +6 -4
- package/dist/{trace-PNCASAXC.js → trace-ODOQIVIW.js} +61 -6
- package/dist/{upgrade-FUSUAGHR.js → upgrade-XANW3FXB.js} +17 -16
- package/dist/{usage-N4MKVHKD.js → usage-4H7ZRXQT.js} +88 -44
- package/dist/{verifiers-YAWOJ3H2.js → verifiers-UZXNBZEB.js} +6 -6
- package/dist/{verify-LTDHYBGY.js → verify-BVKWTNDL.js} +5 -5
- package/dist/{wiki-generate-6M7GHTBJ.js → wiki-generate-MY7WV2QI.js} +49 -47
- package/dist/worker/entry.js +29 -33
- package/docs/alcf-provider.md +1 -1
- package/docs/architecture.md +1 -1
- package/docs/artifact-versions.md +6 -1
- package/docs/built-in-agents.md +1 -1
- package/docs/capacity-and-scheduling.md +23 -2
- package/docs/commands-and-modes.md +1 -1
- package/docs/configuration-and-targets.md +30 -3
- package/docs/context-engine.md +63 -4
- package/docs/documentation-coverage.md +3 -3
- package/docs/documentation-guide.md +1 -1
- package/docs/environment-variables.md +2 -0
- package/docs/eval-runner.md +262 -11
- package/docs/evals-internal.md +72 -2
- package/docs/evidence-and-memory.md +11 -10
- package/docs/evolution.md +1 -1
- package/docs/extensions-and-sharing.md +3 -1
- package/docs/fleet-dispatch.md +4 -4
- package/docs/installation-and-lifecycle.md +1 -1
- package/docs/middleware-and-components.md +1 -1
- package/docs/model-catalog.md +1 -1
- package/docs/observability.md +53 -2
- package/docs/proactive-memory.md +127 -14
- package/docs/prompt-envelope-and-tools.md +19 -1
- package/docs/provider-adapter-cookbook.md +1 -1
- package/docs/release-cut-checklist.md +19 -3
- package/docs/safety-model.md +1 -1
- package/docs/scientific-validation.md +1 -1
- package/docs/skills-marketplace.md +1 -1
- package/docs/tool-usage.md +1 -1
- package/docs/trace-store.md +1 -1
- package/docs/troubleshooting.md +87 -0
- package/docs/tui-design.md +1 -1
- package/docs/worker-dispatch-mechanics.md +1 -1
- package/package.json +2 -1
- package/src/cli/agents.ts +1 -1
- package/src/cli/config-inspect.ts +33 -6
- package/src/cli/config.ts +1 -1
- package/src/cli/doctor-state-size.ts +82 -0
- package/src/cli/doctor.ts +3 -1
- package/src/cli/eval.ts +80 -16
- package/src/cli/extensions.ts +5 -1
- package/src/cli/fleet.ts +32 -3
- package/src/cli/targets.ts +44 -13
- package/src/cli/trace.ts +63 -4
- package/src/cli/usage.ts +63 -14
- package/src/core/bus-events.ts +29 -1
- package/src/core/cache-telemetry.ts +42 -0
- package/src/core/config.ts +18 -0
- package/src/core/defaults.ts +36 -6
- package/src/core/endpoint-key.ts +27 -0
- package/src/core/residency-target-key.ts +25 -0
- package/src/core/response-schema.ts +36 -2
- package/src/domains/config/classify.ts +3 -0
- package/src/domains/context/codewiki/coordinator.ts +12 -4
- package/src/domains/dispatch/admission.ts +40 -3
- package/src/domains/dispatch/capacity-lease.ts +98 -9
- package/src/domains/dispatch/contract.ts +11 -0
- package/src/domains/dispatch/execution-plan.ts +44 -4
- package/src/domains/dispatch/extension.ts +166 -42
- package/src/domains/dispatch/fleet-run.ts +23 -3
- package/src/domains/dispatch/heartbeat.ts +32 -8
- package/src/domains/dispatch/index.ts +3 -0
- package/src/domains/dispatch/orphan-recovery.ts +5 -0
- package/src/domains/dispatch/reservation-store.ts +116 -8
- package/src/domains/dispatch/state.ts +4 -0
- package/src/domains/dispatch/worker-spawn.ts +25 -11
- package/src/domains/dispatch/write-boundary-enforcer.ts +20 -3
- package/src/domains/dispatch/write-boundary.ts +62 -1
- package/src/domains/eval/artifacts/store.ts +62 -0
- package/src/domains/eval/compare/behavioral.ts +224 -0
- package/src/domains/eval/compare/compare.ts +355 -2
- package/src/domains/eval/compare/envelope.ts +128 -0
- package/src/domains/eval/compare/gates.ts +24 -6
- package/src/domains/eval/compare/thresholds.ts +30 -3
- package/src/domains/eval/execution-provenance.ts +240 -0
- package/src/domains/eval/metrics/aggregate.ts +136 -0
- package/src/domains/eval/metrics/call-ledger-stream.ts +112 -0
- package/src/domains/eval/metrics/tracked.ts +413 -0
- package/src/domains/eval/provenance.ts +117 -0
- package/src/domains/eval/reports/comparison.ts +128 -0
- package/src/domains/eval/reports/junit.ts +17 -3
- package/src/domains/eval/reports/markdown.ts +3 -3
- package/src/domains/eval/reports/text.ts +14 -0
- package/src/domains/eval/run-compare.ts +20 -0
- package/src/domains/eval/runners/clio-run.ts +127 -0
- package/src/domains/eval/runners/external-command.ts +28 -3
- package/src/domains/eval/schema/adapter.ts +111 -0
- package/src/domains/eval/schema/artifact.ts +20 -0
- package/src/domains/eval/schema/behavioral-metrics.ts +204 -0
- package/src/domains/eval/schema/behavioral.ts +520 -0
- package/src/domains/eval/schema/execution-envelope.ts +194 -0
- package/src/domains/eval/schema/serving.ts +74 -0
- package/src/domains/eval/schema/suite.ts +38 -8
- package/src/domains/eval/schema/validate.ts +58 -3
- package/src/domains/eval/schema/verdict.ts +237 -0
- package/src/domains/eval/suites/resolve.ts +2 -0
- package/src/domains/eval/suites/run.ts +264 -33
- package/src/domains/eval/verifiers/command.ts +2 -1
- package/src/domains/eval/workspaces/temp-copy.ts +145 -13
- package/src/domains/evidence/build.ts +2 -13
- package/src/domains/evidence/eval.ts +2 -12
- package/src/domains/evidence/findings-markdown.ts +33 -0
- package/src/domains/evidence/run-trust.ts +7 -113
- package/src/domains/evidence/trust-projection.ts +2 -2
- package/src/domains/extensions/compatibility.ts +285 -0
- package/src/domains/extensions/discovery.ts +38 -3
- package/src/domains/extensions/resources.ts +1 -1
- package/src/domains/extensions/state.ts +12 -3
- package/src/domains/extensions/types.ts +2 -0
- package/src/domains/lifecycle/doctor.ts +69 -1
- package/src/domains/memory/index.ts +14 -0
- package/src/domains/memory/task-bank-promotion.ts +64 -0
- package/src/domains/memory/task-memory-policy.ts +77 -8
- package/src/domains/memory/task-memory-spend.ts +131 -0
- package/src/domains/memory/task-memory-status.ts +7 -0
- package/src/domains/memory/task-memory-telemetry.ts +2 -0
- package/src/domains/middleware/index.ts +1 -0
- package/src/domains/middleware/memory-intervention.ts +69 -5
- package/src/domains/middleware/memory-step-endpoint.ts +71 -0
- package/src/domains/observability/background-memory-usage.ts +140 -0
- package/src/domains/observability/cost.ts +1 -1
- package/src/domains/observability/index.ts +7 -0
- package/src/domains/observability/out-of-turn-usage.ts +51 -2
- package/src/domains/observability/trace-store.ts +192 -2
- package/src/domains/prompts/compiler.ts +100 -13
- package/src/domains/providers/endpoint-capacity.ts +96 -0
- package/src/domains/providers/index.ts +10 -0
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +243 -1
- package/src/domains/providers/runtime-resolution.ts +8 -1
- package/src/domains/providers/runtimes/common/probe-helpers.ts +31 -9
- package/src/domains/providers/runtimes/local-native/llamacpp-anthropic.ts +1 -1
- package/src/domains/providers/runtimes/local-native/llamacpp-completion.ts +1 -1
- package/src/domains/providers/runtimes/local-native/llamacpp-embed.ts +1 -1
- package/src/domains/providers/runtimes/local-native/llamacpp-rerank.ts +1 -1
- package/src/domains/providers/runtimes/local-native/llamacpp.ts +4 -1
- package/src/domains/providers/runtimes/local-native/lmstudio.ts +4 -1
- package/src/domains/providers/runtimes/local-native/ollama-native.ts +6 -1
- package/src/domains/providers/types/capability-flags.ts +2 -0
- package/src/domains/providers/types/target-descriptor.ts +2 -0
- package/src/domains/resources/prompts/loader.ts +95 -33
- package/src/domains/safety/call-target.ts +52 -0
- package/src/domains/safety/run-effects.ts +35 -4
- package/src/domains/session/context-accounting.ts +52 -1
- package/src/domains/session/context-ledger.ts +37 -13
- package/src/domains/session/index.ts +6 -0
- package/src/domains/session/prompt-cache.ts +140 -0
- package/src/domains/session/prompt-manifest.ts +42 -0
- package/src/engine/acp/adapter.ts +18 -3
- package/src/engine/ai.ts +35 -0
- package/src/engine/apis/llamacpp-residency.ts +55 -3
- package/src/engine/apis/lmstudio.ts +25 -5
- package/src/engine/apis/ollama-native.ts +2 -1
- package/src/engine/apis/openai-completions.ts +80 -17
- package/src/engine/apis/residency-lock.ts +3 -1
- package/src/engine/apis/residency.ts +34 -1
- package/src/engine/provider-payload.ts +29 -1
- package/src/entry/orchestrator.ts +176 -30
- package/src/interactive/chat-loop-messages.ts +26 -7
- package/src/interactive/chat-loop.ts +318 -41
- package/src/interactive/chat-panel.ts +62 -8
- package/src/interactive/clio-editor.ts +45 -8
- package/src/interactive/context-activity.ts +5 -1
- package/src/interactive/context-meter.ts +1 -1
- package/src/interactive/context-overlay.ts +40 -10
- package/src/interactive/cost-overlay.ts +64 -6
- package/src/interactive/dispatch-board.ts +84 -12
- package/src/interactive/fleet-run-preview.ts +41 -15
- package/src/interactive/handoff-round.ts +41 -2
- package/src/interactive/interactive-application.ts +24 -1
- package/src/interactive/interactive-input-runtime.ts +8 -0
- package/src/interactive/interactive-presentation.ts +4 -0
- package/src/interactive/interactive-shell.ts +20 -17
- package/src/interactive/interactive-slash-runtime.ts +27 -4
- package/src/interactive/memory-overlay.ts +8 -0
- package/src/interactive/mutation-preview.ts +295 -0
- package/src/interactive/overlay-general-openers.ts +16 -0
- package/src/interactive/overlay-key-routing.ts +38 -0
- package/src/interactive/overlay-lifecycle.ts +38 -5
- package/src/interactive/overlay-permission-lifecycle.ts +22 -2
- package/src/interactive/overlay-session-lifecycle.ts +73 -9
- package/src/interactive/overlays/ask-user.ts +91 -19
- package/src/interactive/overlays/help-reference.ts +4 -0
- package/src/interactive/overlays/prompts.ts +11 -1
- package/src/interactive/overlays/settings.ts +35 -1
- package/src/interactive/permission-hint.ts +34 -2
- package/src/interactive/permission-overlay.ts +159 -9
- package/src/interactive/prewarm.ts +197 -0
- package/src/interactive/render-trace.ts +162 -15
- package/src/interactive/renderers/tool-execution.ts +4 -0
- package/src/interactive/side-question.ts +58 -1
- package/src/interactive/status/controller.ts +11 -0
- package/src/interactive/status/state-machine.ts +54 -2
- package/src/interactive/status/types.ts +7 -0
- package/src/interactive/terminal-lease.ts +2 -0
- package/src/interactive/turn-context.ts +299 -31
- package/src/interactive/turn-persistence.ts +14 -4
- package/src/interactive/turn-prewarm.ts +364 -0
- package/src/interactive/turn-queues.ts +7 -4
- package/src/interactive/turn-runtime.ts +8 -1
- package/src/interactive/turn-state.ts +23 -0
- package/src/interactive/view/view-overlay.ts +28 -3
- package/src/tools/ask-user.ts +43 -2
- package/src/tools/dispatch-plan.ts +17 -9
- package/src/tools/dispatch-scout.ts +1 -1
- package/src/tools/registry.ts +16 -0
- package/dist/chunk-AOCYTWAV.js +0 -449
- package/dist/chunk-HWUFFB6L.js +0 -83
- package/dist/chunk-R346GLFC.js +0 -31
- package/dist/doctor-M7YEDGAE.js +0 -91
package/docs/eval-runner.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Clio Coder Local Evaluation Runner
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive task suite validator, subprocess execution simulator, and compare calculator is located at [docs/html/eval_blueprint.html](html/eval_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
7
|
|
|
@@ -15,10 +15,10 @@ The CLI commands under `clio-coder eval` support running, validating, reporting,
|
|
|
15
15
|
|
|
16
16
|
```bash
|
|
17
17
|
clio-coder eval validate --suite <suite.yaml>
|
|
18
|
-
clio-coder eval run --suite <suite.yaml> [--target <id>] [--model <id>] [--out <path>] [--clio-coder-entry <path>]
|
|
18
|
+
clio-coder eval run --suite <suite.yaml> [--trials <n>] [--target <id>] [--model <id>] [--out <path>] [--clio-coder-entry <path>]
|
|
19
19
|
clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--clio-coder-entry <path>]
|
|
20
20
|
clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
|
|
21
|
-
clio-coder eval compare <baselineEvalId> <candidateEvalId>
|
|
21
|
+
clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
|
|
22
22
|
clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
|
|
23
23
|
```
|
|
24
24
|
|
|
@@ -32,7 +32,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
|
|
|
32
32
|
* `swe-jsonl`: Standardized JSONL format representing task runs (e.g. for SWE-bench comparisons).
|
|
33
33
|
* `junit`: XML report for CI/CD integration.
|
|
34
34
|
* **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
|
|
35
|
-
* **`gate`**: Compares candidate metrics against baseline
|
|
35
|
+
* **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
|
|
36
36
|
|
|
37
37
|
Exit codes:
|
|
38
38
|
|
|
@@ -41,8 +41,8 @@ Exit codes:
|
|
|
41
41
|
| `eval validate` | `0` when validation passes | `2` for validation issues |
|
|
42
42
|
| `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
|
|
43
43
|
| `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
|
|
44
|
-
| `eval compare` | `0` when both artifacts
|
|
45
|
-
| `eval gate` | `0` when
|
|
44
|
+
| `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
|
|
45
|
+
| `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for any hard failure, `2` for config/invalid ID errors |
|
|
46
46
|
|
|
47
47
|
---
|
|
48
48
|
|
|
@@ -103,10 +103,11 @@ tasks:
|
|
|
103
103
|
| --- | --- | --- |
|
|
104
104
|
| `version` | - | Must equal `2`. |
|
|
105
105
|
| `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
|
|
106
|
-
| `matrix` | `targets[]`, `repeats` | Matrix of execution targets
|
|
106
|
+
| `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count, and the execution-envelope fields intentionally varied by the suite. |
|
|
107
107
|
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
|
|
108
108
|
| `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
|
|
109
|
-
| `
|
|
109
|
+
| `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
|
|
110
|
+
| `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
|
|
110
111
|
| `metrics` | `collect` | List of metric names to compile for the evaluation runs. |
|
|
111
112
|
|
|
112
113
|
---
|
|
@@ -114,7 +115,7 @@ tasks:
|
|
|
114
115
|
## Workspace Kinds
|
|
115
116
|
* **`local`**: Executes the task directly in the specified local path.
|
|
116
117
|
* **`git`**: Clones the repository from `url`, checks out the specified `commit` or `checkout` ref, and runs there.
|
|
117
|
-
* **`temp-copy`**: Copies the directory at `path` to a temporary workspace location before
|
|
118
|
+
* **`temp-copy`**: Copies the directory at `path` to a temporary workspace location immediately before the matrix item runs and removes it afterward. In a Git checkout, the copy contains exactly tracked files plus untracked files not excluded by Git ignore rules (`git ls-files --cached --others --exclude-standard`), with `excludes` applied afterward. Outside Git it retains the recursive directory copy. This prevents side-effects from polluting other task runs without copying ignored datasets or build trees.
|
|
118
119
|
|
|
119
120
|
---
|
|
120
121
|
|
|
@@ -194,12 +195,262 @@ export interface EvalArtifactV4 {
|
|
|
194
195
|
matrix: { target: string; model: string | null; thinking: string | null };
|
|
195
196
|
summary: EvalArtifactSummaryV4;
|
|
196
197
|
results: EvalArtifactResultV4[];
|
|
198
|
+
servingConfiguration?: EvalServingConfigurationV1;
|
|
199
|
+
aggregates?: EvalScenarioAggregateV1[];
|
|
197
200
|
}
|
|
198
201
|
```
|
|
199
202
|
|
|
203
|
+
`servingConfiguration` and `aggregates` are additive. The v4 reader still accepts an artifact that omits them, and each result's `verdict` is optional for the same reason, so an artifact written before this release loads unchanged.
|
|
204
|
+
|
|
200
205
|
---
|
|
201
206
|
|
|
202
|
-
##
|
|
207
|
+
## The verdict envelope
|
|
208
|
+
|
|
209
|
+
Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
|
|
210
|
+
|
|
211
|
+
```json
|
|
212
|
+
{
|
|
213
|
+
"schema": "clio.eval.verdict.v1",
|
|
214
|
+
"scenarioId": "latency-nonnegative",
|
|
215
|
+
"trialIndex": 0,
|
|
216
|
+
"outcome": "pass",
|
|
217
|
+
"machinery": "ok",
|
|
218
|
+
"reason": null,
|
|
219
|
+
"trackedMetrics": { "...": "see below" },
|
|
220
|
+
"behavioral": null,
|
|
221
|
+
"evidence": {
|
|
222
|
+
"assignmentId": "qcy5rfopdrfw",
|
|
223
|
+
"terminalReceiptDigest": "d85a3ad4f8ae...",
|
|
224
|
+
"graderExitCode": 0
|
|
225
|
+
}
|
|
226
|
+
}
|
|
227
|
+
```
|
|
228
|
+
|
|
229
|
+
The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`, or `unmeasured`; `machinery` is `ok` or `infrastructure_failure`; `reason` is null for a pass or unmeasured outcome and names the rule or failure class for every failure; the original `behavioral` reservation remains exactly `null`; and an envelope claiming both `infrastructure_failure` and `pass` is rejected at parse rather than recorded. A run whose harness broke therefore cannot be read as a model that succeeded. Behavioral results use the separately versioned sibling document below rather than changing this persisted schema.
|
|
230
|
+
|
|
231
|
+
### Behavioral scenario and verdict documents
|
|
232
|
+
|
|
233
|
+
Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
|
|
234
|
+
|
|
235
|
+
The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
|
|
236
|
+
|
|
237
|
+
Expected and forbidden rules contain typed predicates over facts sourced from `transcript`, `tool`, `receipt`, or `grader`. Facts cite a locator, SHA-256 digest, and optional bounded excerpt. The parser caps rules, facts, evidence per category, ids, and explanations. Before judging, facts and unavailable sources are sorted into a canonical representation and hashed as `judgeInputDigest`, so input order cannot change the judge result. Duplicate or conflicting facts, missing categories, malformed evidence, contradictory outcomes, and a behavioral document that references a different result are refused.
|
|
238
|
+
|
|
239
|
+
Suite execution adapts scalar run metrics into these observable facts at the Suite v2 to Artifact v4 boundary. A declared no-tool target leaves tool-dependent rules `unmeasured`, while an available evidence source that omits a required fact produces `unknown`. Categories a role-specific scenario does not claim to measure remain `unmeasured`; they are not numeric zero and do not silently satisfy a rule.
|
|
240
|
+
|
|
241
|
+
### Public built-in behavioral corpus
|
|
242
|
+
|
|
243
|
+
The repository ships corpus `public-built-in-behavior` version `1.0.0` under
|
|
244
|
+
`benchmarks/eval/`. It contains no private prompts, endpoints, credentials, or
|
|
245
|
+
mutable external dataset:
|
|
246
|
+
|
|
247
|
+
- `behavioral-machinery.yaml` provides one positive and one adversarial
|
|
248
|
+
machinery-only check for each of the 13 shipped built-in worker recipes. Its
|
|
249
|
+
deterministic driver loads the production recipe catalog, admits a real
|
|
250
|
+
dispatch through the production gate, runs a scripted worker, and verifies
|
|
251
|
+
the sealed receipt and result-contract outcome. The 26 scenarios require no
|
|
252
|
+
model; they do not infer behavior by grepping recipe frontmatter.
|
|
253
|
+
- `behavioral-model.yaml` provides four isolated main-agent scenarios on the
|
|
254
|
+
`mini` target: a focused edit, adversarial scope control, required
|
|
255
|
+
delegation, and recovery after Bash is denied. Together they cover all eight
|
|
256
|
+
behavioral categories with per-tool call and blocked-call counts, distinct
|
|
257
|
+
and allowlisted read-path counts, declared decoy hits, and grader-emitted
|
|
258
|
+
claim-support and completion facts.
|
|
259
|
+
- `behavioral-model-negative-control.yaml` intentionally reads a declared
|
|
260
|
+
decoy. A healthy run solves its literal task while recording
|
|
261
|
+
`behavioral_failure` with violated exploration and safety labels, proving
|
|
262
|
+
that the rules can reject observed model behavior rather than merely restate
|
|
263
|
+
aggregate success counters.
|
|
264
|
+
|
|
265
|
+
Build once, then run either focused suite from the repository root:
|
|
266
|
+
|
|
267
|
+
```sh
|
|
268
|
+
node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
|
|
269
|
+
node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
270
|
+
node dist/cli/index.js eval run --suite benchmarks/eval/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
271
|
+
```
|
|
272
|
+
|
|
273
|
+
The machinery tasks use the repository read-only and create only private
|
|
274
|
+
scratch state under `TMPDIR`; model tasks use a fresh `temp-copy` workspace and
|
|
275
|
+
remove it after the matrix item settles. The machinery suite is the fast
|
|
276
|
+
admission, worker, and receipt contract. The model suite is the live behavioral
|
|
277
|
+
measurement: keep its Artifact v4 output as evidence for the exact target and
|
|
278
|
+
serving configuration that ran, rather than treating one observed model result
|
|
279
|
+
as a universal guarantee. Behavioral facts and their evidence store only
|
|
280
|
+
bounded read counters, not path strings. As with other eval runs, the artifact's
|
|
281
|
+
bounded diagnostic stdout may retain the underlying tool event stream.
|
|
282
|
+
|
|
283
|
+
### `trackedMetrics`
|
|
284
|
+
|
|
285
|
+
Eleven numbers plus a reason histogram, each carrying the source it came from. `source` is `ledger` (the per-call ledger folded from the worker's own JSON stream), `receipt` (the sealed run receipt), or `estimated`, and `estimated` is what a missing observation is marked as rather than being silently counted as measured.
|
|
286
|
+
|
|
287
|
+
| Metric | Usual source |
|
|
288
|
+
| --- | --- |
|
|
289
|
+
| `modelCalls` | ledger |
|
|
290
|
+
| `uncachedPrefillTokens` | ledger, from `promptCache.backend` |
|
|
291
|
+
| `cacheReadTokens` | ledger, from `promptCache.backend`, falling back to pi-ai cache reads |
|
|
292
|
+
| `generatedTokens` | ledger |
|
|
293
|
+
| `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
|
|
294
|
+
| `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
|
|
295
|
+
| `ttftMsFirstCall` | ledger |
|
|
296
|
+
| `wallClockMs` | receipt |
|
|
297
|
+
| `contextTokensAtEnd` | ledger |
|
|
298
|
+
| `compactions` | ledger |
|
|
299
|
+
| `expectedColdReasons` | ledger, one sourced count per reason |
|
|
300
|
+
|
|
301
|
+
A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
|
|
302
|
+
|
|
303
|
+
### Scenario aggregates
|
|
304
|
+
|
|
305
|
+
`aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
|
|
306
|
+
|
|
307
|
+
### Behavioral multi-metric results
|
|
308
|
+
|
|
309
|
+
A result with a `clio.eval.behavior.v1` verdict also carries the additive
|
|
310
|
+
`clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
|
|
311
|
+
to its role and target/model envelope and records one `number | null`
|
|
312
|
+
observation for each closed metric. The source travels beside every value:
|
|
313
|
+
|
|
314
|
+
| Family | Metric | Direction | Gate | Source |
|
|
315
|
+
|---|---|---|---|---|
|
|
316
|
+
| correctness | `correctness.taskSolved` | higher | hard | grader |
|
|
317
|
+
| safety | `safety.violations` | lower | hard | behavioral label |
|
|
318
|
+
| behavior | `behavior.labelViolations` | lower | informational | behavioral labels |
|
|
319
|
+
| efficiency | `efficiency.toolCalls` | lower | informational | terminal tool events |
|
|
320
|
+
| exploration | `exploration.unnecessaryReads` | lower | informational | read observation counters |
|
|
321
|
+
| delegation | `delegation.quality` | higher | informational | behavioral label |
|
|
322
|
+
| claims | `claims.unsupported` | lower | informational | grader |
|
|
323
|
+
| tokens | `tokens.total` | lower | informational | runner usage stream |
|
|
324
|
+
| latency | `latency.wallMs` | lower | informational | monotonic runner clock |
|
|
325
|
+
| cost | `cost.usd` | lower | informational | sealed receipt |
|
|
326
|
+
|
|
327
|
+
Label metrics are numeric projections only when the category is `satisfied` or
|
|
328
|
+
`violated`; `unknown` and `unmeasured` remain null. A missing grader, token
|
|
329
|
+
stream, receipt, read observation, or category label likewise remains null.
|
|
330
|
+
The projection therefore records observation coverage without claiming that
|
|
331
|
+
silence was success, safety, or zero cost.
|
|
332
|
+
|
|
333
|
+
### Behavioral comparisons and variance
|
|
334
|
+
|
|
335
|
+
`eval compare` reduces the behavioral projection independently for every
|
|
336
|
+
scenario, role, target id, and model id. Each row contains the baseline and
|
|
337
|
+
candidate distributions, coverage, mean delta, variance delta, and two closed
|
|
338
|
+
classifications: `improved`, `regressed`, `unchanged`, or `incomparable` for the
|
|
339
|
+
mean and for variability. Lower variance is the improvement direction for the
|
|
340
|
+
variability classification.
|
|
341
|
+
|
|
342
|
+
Correctness and safety rows are hard. A measured regression fails the hard
|
|
343
|
+
gate even when pass rate, tokens, latency, or cost improved. Losing a
|
|
344
|
+
correctness or safety measurement that existed in the baseline is also a hard
|
|
345
|
+
failure; a category explicitly unmeasured on both sides stays incomparable but
|
|
346
|
+
does not invent a regression. Other families remain visible informational
|
|
347
|
+
tradeoffs. `--metric` accepts either a behavioral metric or family as well as a
|
|
348
|
+
tracked metric, but filtering displayed rows never filters the hard-gate
|
|
349
|
+
decision.
|
|
350
|
+
|
|
351
|
+
Comparison output supports `text`, `json`, `md`, and `junit`. All four carry
|
|
352
|
+
the same hard-gate result and closed classifications. JUnit failures represent
|
|
353
|
+
only hard behavioral failures; an informational efficiency or cost regression
|
|
354
|
+
is emitted as testcase output rather than a failed testcase.
|
|
355
|
+
|
|
356
|
+
### Execution-envelope provenance and comparability
|
|
357
|
+
|
|
358
|
+
Every newly written behavioral result carries an additive
|
|
359
|
+
`clio.eval.execution-envelope.v1` sibling. Artifact v4,
|
|
360
|
+
`clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
|
|
361
|
+
existing identities. The envelope records the selected prompt fragment ids,
|
|
362
|
+
authored versions or `unversioned` marker, fragment content hashes, prompt
|
|
363
|
+
composition hash, recipe id/version/fingerprint when a worker recipe applies,
|
|
364
|
+
target, wire model, runtime, thinking level, tool signature, effective
|
|
365
|
+
autonomy, rule-pack and project-policy hashes, bounded project-context
|
|
366
|
+
provenance, and corpus id/version. A machinery-only scenario uses explicit
|
|
367
|
+
nulls for model concepts that did not apply; null is not substituted for a
|
|
368
|
+
fact that was observed.
|
|
369
|
+
|
|
370
|
+
Suite v2 may declare `matrix.dimensions` from `prompt`, `recipe`, `target`,
|
|
371
|
+
`wireModel`, `runtime`, `thinkingLevel`, `toolSignature`, `autonomy`, `policy`,
|
|
372
|
+
`projectContext`, and `corpus`. Comparison ignores only dimensions declared by
|
|
373
|
+
both artifacts. Any other envelope difference marks every metric row for that
|
|
374
|
+
scenario/role/target incomparable and fails the behavioral gate. A missing
|
|
375
|
+
envelope on only one side is also incomparable. Two older artifacts that both
|
|
376
|
+
predate the sibling remain readable and compare under their existing data.
|
|
377
|
+
|
|
378
|
+
Text, JSON, Markdown, and JUnit comparison reports carry the same envelope
|
|
379
|
+
mismatch. Text and Markdown also include independent per-scenario and per-role
|
|
380
|
+
baseline/candidate counts for improved, regressed, unchanged, and incomparable
|
|
381
|
+
metric means and variances. When the prompt or recipe identity changes, the
|
|
382
|
+
generated evidence names each affected corpus scenario and role instead of
|
|
383
|
+
hiding it behind an aggregate score.
|
|
384
|
+
|
|
385
|
+
### Checked behavioral release baseline
|
|
386
|
+
|
|
387
|
+
The checked deterministic baseline is
|
|
388
|
+
`benchmarks/eval/behavioral-machinery-baseline.json`. The release gate runs all
|
|
389
|
+
26 machinery-only scenarios through the built CLI and compares a stable
|
|
390
|
+
projection of their labels, metrics, and execution envelopes with that file.
|
|
391
|
+
It requires no model, private endpoint, credential, or mutable dataset.
|
|
392
|
+
|
|
393
|
+
When an intentional prompt, recipe, policy, or expected-behavior change moves
|
|
394
|
+
the evidence, run the same machinery suite first, inspect the failing diff and
|
|
395
|
+
the named affected corpus results, then update explicitly:
|
|
396
|
+
|
|
397
|
+
```sh
|
|
398
|
+
npm run build
|
|
399
|
+
node benchmarks/eval/check-behavioral-release.mjs --update
|
|
400
|
+
git diff -- benchmarks/eval/behavioral-machinery-baseline.json
|
|
401
|
+
```
|
|
402
|
+
|
|
403
|
+
The baseline update belongs in the reviewed change that caused it. Do not use
|
|
404
|
+
the update command merely to make a red gate green. The model-required and
|
|
405
|
+
negative-control suites remain manual release evidence because their outputs
|
|
406
|
+
depend on a live target; they are never folded into the deterministic baseline.
|
|
407
|
+
The projection excludes `latency.wallMs` because scheduler timing is not stable
|
|
408
|
+
evidence. Behavioral labels, deterministic metrics, and the execution envelope
|
|
409
|
+
remain checked byte for byte.
|
|
410
|
+
|
|
411
|
+
### Hard thresholds and informational budgets
|
|
412
|
+
|
|
413
|
+
Suite and external threshold files keep two separate assertion lists:
|
|
414
|
+
|
|
415
|
+
```yaml
|
|
416
|
+
thresholds:
|
|
417
|
+
fail:
|
|
418
|
+
- metric: task.solved
|
|
419
|
+
op: eq
|
|
420
|
+
value: false
|
|
421
|
+
informational:
|
|
422
|
+
- metric: cost.usd
|
|
423
|
+
op: gt
|
|
424
|
+
value: 0.25
|
|
425
|
+
```
|
|
426
|
+
|
|
427
|
+
`fail` is the backwards-compatible hard list. A firing or unresolved hard
|
|
428
|
+
assertion makes `eval run` or `eval gate` exit nonzero. `informational` uses the
|
|
429
|
+
same typed predicates and reports every firing budget or missing measurement,
|
|
430
|
+
but never changes the exit status. `eval gate` additionally evaluates the
|
|
431
|
+
baseline-to-candidate correctness and safety hard gate, so a cheaper candidate
|
|
432
|
+
cannot offset a task or safety regression.
|
|
433
|
+
|
|
434
|
+
### `--trials N`
|
|
203
435
|
|
|
204
|
-
|
|
436
|
+
`--trials N` overrides the suite's `matrix.repeats` and asks for an isolated workspace per matrix item. A `local` workspace is converted to a temporary copy immediately before that item runs, so an explicit trial run never mutates the directory it was pointed at; `git` and `temp-copy` workspaces already produce a distinct preparation directory per item. Workspace and state directories are removed on the item's `finally` path, including runner, setup, and copy failures. The trial index rides through to each verdict's `trialIndex`.
|
|
437
|
+
|
|
438
|
+
### Serving-configuration provenance and drift refusal
|
|
439
|
+
|
|
440
|
+
`servingConfiguration` records what the numbers were measured against: `targetId`, `runtimeId`, `modelId`, `serverBuild`, `total_slots`, `thinkingLevel`, and `compiledPromptHash`. The build string and slot count are read from the server after the matrix has run while it is still awake, by fetching `/props` and falling back to the model-qualified slots query when `/props` exposes no `total_slots`. The prompt hash is the receipt's static composition hash, so a prompt change is visible as a configuration change rather than as a mysterious metric shift.
|
|
441
|
+
|
|
442
|
+
`eval compare` prints both configurations and refuses outright when they differ:
|
|
443
|
+
|
|
444
|
+
```text
|
|
445
|
+
serving configuration drift; pass --allow-config-drift to compare these runs
|
|
446
|
+
baseline serving: target=mini runtime=llamacpp model=... server_build=b226-2115b73d8 total_slots=1 thinking=off compiled_prompt_hash=...
|
|
447
|
+
candidate serving: ...
|
|
448
|
+
```
|
|
449
|
+
|
|
450
|
+
`--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
|
|
451
|
+
|
|
452
|
+
---
|
|
453
|
+
|
|
454
|
+
## Task Outcome Measurement (`verify.measure`)
|
|
205
455
|
|
|
456
|
+
Task outcome commands declared under `verify.measure` are the code grader for whether the model solved the workload and record metrics (`task.solved`, `task.exitCode`). A non-zero exit fails the final result and is named on its verdict as `reason: grader_failed`, while `machinery` remains `ok` when the runner and machinery verifiers succeeded. This keeps the artifact's `pass`, verdict outcome, scenario aggregates, and summary on one pass decision without misreporting a grader failure as broken machinery.
|
package/docs/evals-internal.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Internal Eval Suites
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive blueprint is available at [docs/html/evals_internal_blueprint.html](html/evals_internal_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
7
|
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
@@ -19,6 +19,77 @@ data directory. Product eval artifacts and external benchmark campaigns are
|
|
|
19
19
|
separate: public benchmark adapters live under `benchmarks/community/` and do
|
|
20
20
|
not use the eval runner.
|
|
21
21
|
|
|
22
|
+
The public behavioral corpus is the deliberate exception to the otherwise
|
|
23
|
+
private Suite v2 data policy. Its reviewable, synthetic suites live under
|
|
24
|
+
`benchmarks/eval/`: a model-free positive/adversarial authority pair for every
|
|
25
|
+
built-in worker recipe, four tiny main-agent model scenarios covering all eight
|
|
26
|
+
behavioral categories with event- and grader-derived facts, and an intentional
|
|
27
|
+
decoy negative control. The model-free driver uses the shipped recipe catalog,
|
|
28
|
+
real dispatch admission, scripted workers, and sealed receipts rather than
|
|
29
|
+
frontmatter inspection. See
|
|
30
|
+
[eval-runner.md](eval-runner.md#public-built-in-behavioral-corpus) for the
|
|
31
|
+
focused commands. Private prompts, calibration cases, fleet coordinates, and
|
|
32
|
+
campaign artifacts still belong outside this repository and must not be copied
|
|
33
|
+
into the public corpus.
|
|
34
|
+
|
|
35
|
+
## Running a private suite as a measurement
|
|
36
|
+
|
|
37
|
+
A private suite is usually run to answer whether a harness change moved
|
|
38
|
+
something, which makes it a measurement rather than a pass or fail. Three
|
|
39
|
+
mechanics matter for that, all documented in full in
|
|
40
|
+
[eval-runner.md](eval-runner.md#the-verdict-envelope).
|
|
41
|
+
|
|
42
|
+
Run repeated trials with `--trials N` rather than by editing `matrix.repeats`.
|
|
43
|
+
The flag overrides the suite's repeat count and asks for an isolated workspace
|
|
44
|
+
per matrix item, so a `local` workspace is copied for the run instead of being
|
|
45
|
+
mutated across trials. Each result's verdict carries its `trialIndex`, and the
|
|
46
|
+
artifact's `aggregates` reduce them per scenario: `k`, `passAtK` (any trial
|
|
47
|
+
passed), `passPowK` (every trial passed), and a mean and nearest-rank p90 for
|
|
48
|
+
every tracked metric. A single trial produces a `k: 1` aggregate whose mean and
|
|
49
|
+
p90 are the same observed value, which is a fact to state in a report rather
|
|
50
|
+
than a distribution to reason about.
|
|
51
|
+
|
|
52
|
+
Read the tracked metrics with their sources attached. A private suite on a
|
|
53
|
+
local target is measuring prefill economics as much as correctness, so
|
|
54
|
+
`uncachedPrefillTokens`, `cacheReadTokens`, `ttftMsFirstCall`, and the
|
|
55
|
+
`expectedColdReasons` histogram are the interesting columns, and each one says
|
|
56
|
+
whether it came from the ledger, from the receipt, or was `estimated`. A metric
|
|
57
|
+
marked `estimated` on one side of a comparison and measured on the other is
|
|
58
|
+
refused rather than differenced.
|
|
59
|
+
|
|
60
|
+
Behavioral suites add a second projection beside those tracked performance
|
|
61
|
+
metrics. Compare it per scenario, role, and target/model envelope rather than
|
|
62
|
+
reducing unlike roles into one pass rate. Correctness and safety are hard
|
|
63
|
+
regression gates; tool efficiency, unnecessary exploration, delegation
|
|
64
|
+
quality, unsupported claims, tokens, latency, and receipt cost remain separate
|
|
65
|
+
families with their own measured coverage and repeat variance. A missing value
|
|
66
|
+
is null and makes that row incomparable, never zero.
|
|
67
|
+
|
|
68
|
+
Put release-blocking assertions under `thresholds.fail` and non-blocking spend
|
|
69
|
+
or latency budgets under `thresholds.informational`. Informational findings are
|
|
70
|
+
printed in every gate run but do not change its exit status. Do not put a cost
|
|
71
|
+
budget in the hard list to compensate for weak correctness, and do not turn a
|
|
72
|
+
correctness rule into an informational budget; the comparison gate evaluates
|
|
73
|
+
correctness and safety before either kind of operator-authored threshold.
|
|
74
|
+
|
|
75
|
+
Record the serving configuration or the comparison is not one. The artifact
|
|
76
|
+
captures `targetId`, `runtimeId`, `modelId`, `serverBuild`, `total_slots`,
|
|
77
|
+
`thinkingLevel`, and `compiledPromptHash`, read from the server after the matrix
|
|
78
|
+
has run while it is still awake. `eval compare` refuses two artifacts whose
|
|
79
|
+
configurations differ unless `--allow-config-drift` is passed, and prints both
|
|
80
|
+
either way. Treat that refusal as the useful behavior it is: a private suite
|
|
81
|
+
compared across a server restart that changed a flag, a quantization, or the
|
|
82
|
+
thinking level is measuring the server, not the change under test.
|
|
83
|
+
|
|
84
|
+
The verdict envelope keeps its original `behavioral: null` field for compatibility.
|
|
85
|
+
A suite that declares a versioned behavioral scenario records the result as a
|
|
86
|
+
separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
|
|
87
|
+
to the unchanged verdict identity. Its labels come only from bounded transcript,
|
|
88
|
+
tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
|
|
89
|
+
`machinery: "infrastructure_failure"`, which the parser refuses to pair with a
|
|
90
|
+
`pass`, so a private suite cannot report a passing rate that includes runs
|
|
91
|
+
nothing measured.
|
|
92
|
+
|
|
22
93
|
## Context Regression Seed
|
|
23
94
|
|
|
24
95
|
```yaml
|
|
@@ -267,4 +338,3 @@ thresholds:
|
|
|
267
338
|
op: gt
|
|
268
339
|
value: 0
|
|
269
340
|
```
|
|
270
|
-
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Evidence Corpus and Long-Term Memory
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive memory lifecycle dashboard and simulator is located at [docs/html/memory_blueprint.html](html/memory_blueprint.html) (Version: 0.3.9). Use it to design, validate, and simulate memory proposals, approval loops, pruning rules, and token budgets.
|
|
5
5
|
|
|
6
6
|
Clio Coder treats run claims and agent lessons as structured artifacts to support reproducibility and scientific provenance. In evaluations such as [SWE-bench](https://www.swebench.com), capturing granular execution evidence is essential for validating agent claims. Evidence corpora are deterministic directories built from run ledgers, receipts, sessions, audits, and eval artifacts. In v0.3.7, forensic evidence auto-builds on dispatch run completion: when a run finalizes, the observability domain automatically compiles the evidence bundle under `<dataDir>/evidence/run-<id>/` and updates a compact sidecar index row in `<stateDir>/evidence-index.json`. Long-term memory records are local, evidence-linked, and only injected after explicit approval. Use the TUI [`/view`](observability.md) command for interactive inspection of receipts, dispatch output, durable tool output, compaction summaries, and session accountability before building or citing evidence.
|
|
7
7
|
|
|
@@ -68,9 +68,9 @@ Eval evidence adds `eval-result.json` and uses empty receipt/protected-artifact
|
|
|
68
68
|
| `audit-linked.jsonl` | Audit rows linked to run/session context when available. |
|
|
69
69
|
| `receipt.json` | Receipt bundle (`{ version: 1, receipts: [...] }`); only receipts that pass integrity verification contribute verified fields. |
|
|
70
70
|
| `gate-decisions.json` | Integrity-verified review verdicts, compete winner selections, and winner confirmations discovered from linked receipt ids. |
|
|
71
|
-
| `trust-status.json` | Canonical per-run six-axis trust projections derived from authenticated receipts, gate decisions, grounded validation artifacts
|
|
71
|
+
| `trust-status.json` | Canonical per-run six-axis trust projections derived from authenticated receipts, gate decisions, and grounded validation artifacts. |
|
|
72
72
|
| `protected-artifacts.json` | Protected artifact state/events. |
|
|
73
|
-
| `findings.json` / `findings.md` | Structured
|
|
73
|
+
| `findings.json` / `findings.md` | Structured findings plus a readable report that begins with each linked run's canonical tier, fixed-order summary, and six axes. |
|
|
74
74
|
|
|
75
75
|
### Run attribution under concurrency
|
|
76
76
|
|
|
@@ -219,19 +219,20 @@ They do not mutate receipt, gate-decision, evidence-bundle, or session formats.
|
|
|
219
219
|
| Valid bounded project context, valid none-tier workspace-root record, or valid briefing hash | Context provenance is `recorded`. A `none`-tier run still receives the workspace-root message, so a none-tier block naming exactly `workspace-root` with a well-formed count and hash is `recorded`. Explicit project-context tier `none` with no content and no briefing is `not_applicable`; a missing historical field is `unknown`; a contradictory block (a handbook section under a none policy, a hash with no section, a malformed count) is `invalid`. |
|
|
220
220
|
| Gate decision | An authenticated independent pass or fail maps to `passed` or `failed`. Correlated review maps to `not_independent`. Unauthenticated artifacts map to `unknown`; operator or full-auto confirmation alone is `not_applicable` to independent review. |
|
|
221
221
|
| Receipt autonomy grade | `mediated`, `approximated`, and `bypassed` map to `enforced`, `approximated`, and `bypassed`. A dangerous-bypass flag always normalizes to `bypassed`; a missing historical block is `unknown`. |
|
|
222
|
-
| Finish-contract assessment |
|
|
223
|
-
| Malformed audit row identifier | A blank or whitespace-only optional identifier
|
|
222
|
+
| Finish-contract assessment | The assessment remains linked in `audit-linked.jsonl` and contributes its domain findings and tags. It does not override the receipt-derived `completionEvidence` axis on the evidence surface alone. |
|
|
223
|
+
| Malformed audit row identifier | A blank or whitespace-only optional identifier remains linked as audit input and never aborts the bundle. It cannot affect the receipt-derived trust projection. |
|
|
224
224
|
| Bundle without `trust-status.json` | Inspection reports `projection: historical_format` with no canonical run projections. It never reconstructs positive states from older summary tags. |
|
|
225
225
|
|
|
226
226
|
Receipt inspection, worker output, monitor details, and evidence rebuilding all
|
|
227
227
|
use the same authenticated receipt projection boundary. Evidence rebuilding
|
|
228
|
-
then composes independently authenticated gate decisions
|
|
229
|
-
|
|
228
|
+
then composes independently authenticated gate decisions without changing
|
|
229
|
+
receipt-owned axes. Findings such as
|
|
230
230
|
`no-validation`, `proxy-validation`, `external-approximation`,
|
|
231
231
|
`external-bypass`, `independent-review`, `context-provenance`, and
|
|
232
|
-
`completion-evidence` are selected from the canonical states
|
|
233
|
-
|
|
234
|
-
receipt, gate, audit, and trace
|
|
232
|
+
`completion-evidence` are selected from the canonical states. `findings.md`
|
|
233
|
+
prints the tier, summary, and every axis before those diagnostic records, while
|
|
234
|
+
their detailed domain artifacts remain in the receipt, gate, audit, and trace
|
|
235
|
+
files.
|
|
235
236
|
|
|
236
237
|
The canonical aggregate is an additive projection for downstream work. Receipt
|
|
237
238
|
integrity remains version 18, evidence bundles remain version 1, gate decisions
|
package/docs/evolution.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Evolution and Change Manifests
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive change manifest editor, authority risk assessor, and checklist workspace is located at [docs/html/evolution_blueprint.html](html/evolution_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
|
|
7
7
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Extensions, Prompt Templates, Skills, and Share Archives
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard is located at [docs/html/extensions_blueprint.html](html/extensions_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
Clio Coder has lightweight community-oriented resource packaging. Extensions are filesystem bundles that contribute prompts and skills. Share archives are portable JSON files for moving project/user Clio resources between machines or collaborators. Themes are built into the engine and are no longer loaded from extensions.
|
|
7
7
|
|
|
@@ -197,6 +197,8 @@ Required fields are `manifestVersion: 1`, `id`, `version`, and `description`. `n
|
|
|
197
197
|
|
|
198
198
|
IDs must be lowercase and may include numbers, dots, underscores, and hyphens; they must start/end alphanumeric.
|
|
199
199
|
|
|
200
|
+
`compatibility.clio` is optional. When present, it must be a valid SemVer range such as `>=0.3.8`, `^0.3.8`, or `0.3.x`. Installation refuses a package whose range excludes the running Clio version and names the extension, its declared range, and that running version. Clio repeats the check whenever it loads installed extensions, so a package that becomes incompatible after a Clio version change stays visible in `extensions list` with its diagnostic but contributes no resources. An incompatible project package does not hide a compatible user package with the same ID. A manifest without `compatibility.clio` keeps the existing unrestricted behavior.
|
|
201
|
+
|
|
200
202
|
---
|
|
201
203
|
|
|
202
204
|
## Extension CLI
|
package/docs/fleet-dispatch.md
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
# Fleet Dispatch
|
|
2
2
|
|
|
3
|
-
> **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.
|
|
3
|
+
> **Interactive Spec Available:** An interactive fleet node topology planner, scout router, receipt verifier, and failure taxonomy simulator is located at [docs/html/fleet_dispatch_blueprint.html](html/fleet_dispatch_blueprint.html) (Version: 0.3.9).
|
|
4
4
|
|
|
5
5
|
Clio Coder dispatches bounded worker agents. With a fleet configured, those
|
|
6
6
|
workers run on remote machines over SSH while the orchestrator keeps every
|
|
@@ -531,14 +531,14 @@ The grammar for declared write boundary entries requires repository-relative POS
|
|
|
531
531
|
Write boundary enforcement is detect-and-rollback, never OS or filesystem sandboxing. A step runs with whatever filesystem permissions its underlying execution environment possesses. Upon step completion, the orchestrator inspects the working tree to verify compliance:
|
|
532
532
|
1. Snapshot baseline: Before a step executes, the orchestrator captures a snapshot (`captureWorkspaceSnapshot`) recording the baseline git HEAD commit and existing dirty path content tokens.
|
|
533
533
|
2. Workspace diffing: After step completion, the orchestrator runs git status inspection (`diffWorkspace`) to identify changed paths relative to the snapshot baseline commit.
|
|
534
|
-
3. Authorship attribution: A changed path outside the allowlist is blamed on the window only when it intersects what the window's own runs recorded writing. That record is the run's tool-call stream, folded by the same recorder that grounds a sealed mutation report and read back through `DispatchContract.
|
|
535
|
-
4. Open records: A run's recorded write set is read as a closed list only when the run could not have written outside it. A window whose steps do not all offer a closed record falls back to blaming every change outside the allowlist, which is the behavior that predates attribution. Three things open a record: a `kind: code` step, whose registered command publishes no tool events; a run whose tool telemetry coverage is not `complete`, such as one on a subprocess runtime; and a run that made a successful call to a tool able to mutate a path its own arguments do not name. That last set is derived from the tool surface rather than authored, as every registered tool outside the `read` and `write` action classes, which today is `bash`, `verify`, `dispatch`, and `steer`, plus any dynamic or MCP tool whose schema this process cannot read. The `git` tool is a closed status, diff, and log surface and stays enumerable. The verdict records `attributionComplete: false` and
|
|
534
|
+
3. Authorship attribution: A changed path outside the allowlist is blamed on the window only when it intersects what the window's own runs recorded writing. That record is the run's tool-call stream, folded by the same recorder that grounds a sealed mutation report and read back through `DispatchContract.observedRunWriteAttribution`. A change that no run in the window recorded is an unattributed concurrent change: it is listed under `unattributed` in the verdict and reported to the operator, and it is never rolled back. This is what keeps a file an operator edited while a fleet ran from being overwritten with its committed version.
|
|
535
|
+
4. Open records: A run's recorded write set is read as a closed list only when the run could not have written outside it. A window whose steps do not all offer a closed record falls back to blaming every change outside the allowlist, which is the behavior that predates attribution. Three things open a record: a `kind: code` step, whose registered command publishes no tool events; a run whose tool telemetry coverage is not `complete`, such as one on a subprocess runtime; and a run that made a successful call to a tool able to mutate a path its own arguments do not name. That last set is derived from the tool surface rather than authored, as every registered tool outside the `read` and `write` action classes, which today is `bash`, `verify`, `dispatch`, and `steer`, plus any dynamic or MCP tool whose schema this process cannot read. The `git` tool is a closed status, diff, and log surface and stays enumerable. When a successful opaque call opens the record, the live Alt+W card immediately names the tool and says that its arguments cannot enumerate every path it may write. The durable verdict records `attributionComplete: false` and an `attributionDowngrades` entry with the reason, tool name, tool call id, run id, and step id. The after-the-fact write-boundary message repeats that cause before explaining why the whole outside diff was blamed on the window.
|
|
536
536
|
5. Rollback execution: Attributed unauthorized changes are automatically rolled back (`rollbackPath`).
|
|
537
537
|
6. Content source: Rollback restores content strictly from what git already has in the pinned baseline commit (`snapshot.head`). If a path was already dirty when the step snapshot was captured, its prior content is not stored in git, so in-place restoration cannot be guaranteed. The working tree is left as the step made it, and the status settles as `rollback-incomplete`.
|
|
538
538
|
7. Violation handling: Any attributed unauthorized change fails the step with the typed reason `writes_boundary_violation`.
|
|
539
539
|
8. Window attribution: Enforcement evaluates scheduling windows (`wave-<n>` or `revalidate-<stepId>-<n>`). A wave window cannot combine steps with overlapping declared boundaries or multiple concurrent step writers, ensuring single-step attribution.
|
|
540
540
|
9. Ignored paths and state subtraction: Enforcement evaluates paths reported by git status, which never lists a git-ignored path. A declared `writes` entry the repository ignores is therefore refused before anything runs, by `fleet validate`, by `fleet run` preflight, and by the `/fleet run` preview, with a diagnostic naming the entry and the ignoring rule (for example `'work/' is ignored by .gitignore:1:work/`). Silently certifying such a window as clean is not an option, because nothing about it was observed. The Clio state directory (`.clio-coder/` or `clioStateDir()`) is subtracted from status checks so orchestrator receipts, code step log artifacts, and boundary verdicts do not trigger false violations.
|
|
541
|
-
10. Durable records: Verdicts are serialized as JSON records at `write-boundaries/<rootId>/<window>.json` under the Clio state directory, carrying the baseline HEAD commit, checked paths, violations, unattributed concurrent changes, the attribution completeness flag, rollback actions, status, and SHA-256 digest.
|
|
541
|
+
10. Durable records: Verdicts are serialized as JSON records at `write-boundaries/<rootId>/<window>.json` under the Clio state directory, carrying the baseline HEAD commit, checked paths, violations, unattributed concurrent changes, the attribution completeness flag and downgrade causes, rollback actions, status, and SHA-256 digest.
|
|
542
542
|
|
|
543
543
|
### Bounded check/repair loops
|
|
544
544
|
|
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
Clio Coder is designed to be self-contained and platform-compliant. This document outlines the default directory paths, file purposes, permission levels, and lifecycle commands (`install`, `reset`, `upgrade`, and `uninstall`). Clio Coder installs from npm as `@iowarp/clio-coder` (`npm install -g @iowarp/clio-coder`, published since v0.3.0) or from a source checkout with a deterministic local symlink; the CLI classifies both install kinds and `clio-coder upgrade` handles each.
|
|
4
4
|
|
|
5
5
|
> [!TIP]
|
|
6
|
-
> **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.
|
|
6
|
+
> **Interactive Spec Available:** An interactive dashboard with a path simulator and visual flowcharts is located at [docs/html/lifecycle_blueprint.html](html/lifecycle_blueprint.html) (Version: 0.3.9). You can open it directly in any web browser to view details dynamically.
|
|
7
7
|
|
|
8
8
|
---
|
|
9
9
|
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Middleware and Component Registry
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard with an interactive component scanner and a dynamic hook-and-effect pipeline is located at [docs/html/middleware_blueprint.html](html/middleware_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
Clio Coder has two related but separate surfaces:
|
|
7
7
|
|
package/docs/model-catalog.md
CHANGED
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Model Catalog, Runtime Refresh, and Field Notes
|
|
2
2
|
|
|
3
3
|
> [!TIP]
|
|
4
|
-
> **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.
|
|
4
|
+
> **Interactive Spec Available:** An interactive dashboard mapping capabilities, probe discovery, and target resolution is located at [docs/html/models_blueprint.html](html/models_blueprint.html) (Version: 0.3.9).
|
|
5
5
|
|
|
6
6
|
Clio Coder treats a selectable model as the intersection of three sources:
|
|
7
7
|
|