@iowarp/clio-coder 0.4.1 → 0.4.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +92 -0
- package/CONTRIBUTING.md +59 -36
- package/README.md +404 -472
- package/SECURITY.md +2 -1
- package/dist/{acp-ZILU3AUO.js → acp-TMDQZDIG.js} +7 -7
- package/dist/{agents-HYWGBGQR.js → agents-5N5NG3XG.js} +28 -28
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-N3QT7CBO.js → auth-Z5CCBXKQ.js} +8 -9
- package/dist/{builtins-UJLMOVOV.js → builtins-K6TNDT24.js} +4 -4
- package/dist/{chunk-GVQJ5CCZ.js → chunk-2HFQNRV3.js} +7 -7
- package/dist/{chunk-QMXC4JB7.js → chunk-2NHR3NAY.js} +163 -1401
- package/dist/chunk-2X4RYJTJ.js +39 -0
- package/dist/{chunk-Y45G3AXC.js → chunk-2Z2IKEXI.js} +6 -10
- package/dist/{chunk-EIMVLWB3.js → chunk-34BHNEE3.js} +7 -3
- package/dist/{chunk-GIZNH63R.js → chunk-35MSIRKH.js} +9 -4
- package/dist/chunk-3EBYEESD.js +314 -0
- package/dist/{chunk-CTJ4RNAA.js → chunk-3F7VUY77.js} +2 -2
- package/dist/{chunk-AP73CFDC.js → chunk-3KIPBMUA.js} +2 -2
- package/dist/{chunk-J5LZHVIT.js → chunk-3M6DQK6S.js} +113 -35
- package/dist/{chunk-VEGN6WIQ.js → chunk-462T4EGZ.js} +2 -2
- package/dist/{chunk-AFKWHWXF.js → chunk-4JDLP6ZS.js} +33 -16
- package/dist/{chunk-6FN3E6KX.js → chunk-4O6MANBS.js} +2 -2
- package/dist/{chunk-AKB4GYDL.js → chunk-54ODD65L.js} +5 -5
- package/dist/{chunk-BBTJOK6Y.js → chunk-5KW52TEP.js} +3 -3
- package/dist/{chunk-6CCS4G3W.js → chunk-5PFYMY2V.js} +2 -2
- package/dist/chunk-77QIVUZB.js +1334 -0
- package/dist/{chunk-7OBGU7UB.js → chunk-7BHIY2MW.js} +7 -13
- package/dist/{chunk-3QSOM6PA.js → chunk-AZ4WMN4W.js} +2 -2
- package/dist/{chunk-6NJQITNH.js → chunk-B74PXLU7.js} +6 -3
- package/dist/{chunk-R23Z6K6I.js → chunk-B7HM5Z7T.js} +15 -15
- package/dist/{chunk-R32CLGZ6.js → chunk-BO7Y52RY.js} +81 -20
- package/dist/{chunk-UEDMSP56.js → chunk-BYMNWQ7O.js} +123 -148
- package/dist/{chunk-ZJLUDYFY.js → chunk-CRFOIAX3.js} +4 -4
- package/dist/{chunk-2NM363SV.js → chunk-CYZW7JHJ.js} +7 -7
- package/dist/{chunk-6HMJX2VU.js → chunk-DYHAXKHD.js} +38 -10
- package/dist/{chunk-THYWACCR.js → chunk-DZAW46HP.js} +3 -3
- package/dist/{chunk-FYUN5KZ3.js → chunk-DZEK6CJN.js} +17 -17
- package/dist/{chunk-3I5NY75V.js → chunk-E7GT7O5N.js} +5 -5
- package/dist/{chunk-VKFQTNDV.js → chunk-F2I26BDK.js} +4 -4
- package/dist/{chunk-HLW2MRKE.js → chunk-F4EKGO4N.js} +3 -1
- package/dist/{chunk-IXJT6DCX.js → chunk-FVDGR2ZL.js} +3 -3
- package/dist/{chunk-TZSKNMZG.js → chunk-GTUD2WMY.js} +2 -1
- package/dist/{chunk-7EPLI7VL.js → chunk-HIICAHCJ.js} +2 -2
- package/dist/{chunk-E67WX76H.js → chunk-HKMD33FO.js} +29 -80
- package/dist/chunk-HLAFFSEK.js +360 -0
- package/dist/{chunk-UAPGZHYC.js → chunk-I64IFBLB.js} +9 -2
- package/dist/{chunk-XKA2ICR3.js → chunk-I66ZTYNP.js} +440 -175
- package/dist/{chunk-7PWAODYW.js → chunk-I7XBWTYH.js} +2 -2
- package/dist/{chunk-PVAMAVBB.js → chunk-IDNA72AH.js} +102 -2
- package/dist/{chunk-GCSMB2KY.js → chunk-IKOZFYBN.js} +1 -1
- package/dist/{chunk-2VG7KLYV.js → chunk-IKSLQ4XV.js} +5460 -3241
- package/dist/{chunk-QKIFBZKT.js → chunk-IMXMHHMQ.js} +166 -25
- package/dist/{chunk-74YWRRU5.js → chunk-JBCS7CRR.js} +2 -2
- package/dist/{chunk-BDPT6GTK.js → chunk-JWJGP5DQ.js} +2 -2
- package/dist/{chunk-K6BF4U2H.js → chunk-KKOJXO6R.js} +62 -14
- package/dist/chunk-KPXDY6QF.js +47 -0
- package/dist/{chunk-ABLSQ6JX.js → chunk-LJID3DYZ.js} +7 -1
- package/dist/{chunk-VKRH2TCS.js → chunk-M2DAX4F6.js} +2 -2
- package/dist/{chunk-6I5ILFOF.js → chunk-M2WXEHER.js} +2 -2
- package/dist/{chunk-YPI3QQCF.js → chunk-MCEPRMZW.js} +2 -4
- package/dist/{chunk-N5UK64DP.js → chunk-MCMZMDAC.js} +2 -2
- package/dist/{chunk-Y4CAGMM6.js → chunk-MNJGS2IN.js} +5 -6
- package/dist/{chunk-TVHHYFHE.js → chunk-NEDJ26B5.js} +2 -2
- package/dist/{chunk-U2WB7TZS.js → chunk-NMJXSHBJ.js} +97 -85
- package/dist/{chunk-HUAS7ITX.js → chunk-O3YUNJZ2.js} +13 -21
- package/dist/{chunk-MA3H6DM5.js → chunk-P75RZCJW.js} +25 -3
- package/dist/{chunk-IG7BCQBA.js → chunk-PGF63K6I.js} +2 -2
- package/dist/chunk-PJX3WQUQ.js +42 -0
- package/dist/{chunk-6DWBAZ5U.js → chunk-Q4XWMHX6.js} +4 -6
- package/dist/{chunk-OJTRZGR3.js → chunk-QQLGQY2A.js} +8 -8
- package/dist/{chunk-J4HBWF6Y.js → chunk-RLYRBIYQ.js} +115 -20
- package/dist/{chunk-NLFAQR7Z.js → chunk-S66XZJOF.js} +3 -23
- package/dist/{chunk-C537JADH.js → chunk-SSEYRH53.js} +6 -7
- package/dist/chunk-SZAA6XDG.js +30 -0
- package/dist/{chunk-MOPSG2X7.js → chunk-TPEQIQIE.js} +6 -6
- package/dist/{chunk-JA5QWE4Z.js → chunk-UBRFI4HS.js} +1879 -1650
- package/dist/{chunk-BTGG6BG2.js → chunk-UH347SHR.js} +154 -15
- package/dist/{chunk-5YHDIDBP.js → chunk-UH632ZYL.js} +2 -2
- package/dist/{chunk-BWW4HLO4.js → chunk-UXCU4E3T.js} +8 -6
- package/dist/{chunk-6VC4OV3Z.js → chunk-VIA6RFQZ.js} +3 -11
- package/dist/{chunk-ZAZB4JMW.js → chunk-VKPAQYEB.js} +27 -8
- package/dist/{chunk-UXN6JT4W.js → chunk-W4YEMFBX.js} +2 -2
- package/dist/{chunk-TD3PGPQA.js → chunk-W6NIE6OW.js} +2 -2
- package/dist/{chunk-TVH4ONAM.js → chunk-X7IARSHT.js} +3 -3
- package/dist/{chunk-PJJ6MY27.js → chunk-XE3PCIXH.js} +3 -3
- package/dist/{chunk-FEFIFZTL.js → chunk-XGDPUNND.js} +2 -2
- package/dist/{chunk-SCYB3HA4.js → chunk-XOXV5GKE.js} +51 -16
- package/dist/{chunk-QTFGO774.js → chunk-XQRY4DTA.js} +24 -11
- package/dist/{chunk-BJGUKIG4.js → chunk-YJISEZKC.js} +2 -2
- package/dist/{chunk-GPPB3JBE.js → chunk-ZGNYYXQ6.js} +2 -2
- package/dist/{chunk-SINK3QR6.js → chunk-ZNT2M6TG.js} +7 -7
- package/dist/{chunk-7RY5VZPH.js → chunk-ZW4HH5JJ.js} +6 -6
- package/dist/cli/index.js +33 -32
- package/dist/{clio-IT3G3VQH.js → clio-7VB377CC.js} +7 -7
- package/dist/{code-nav-RK6S7F6E.js → code-nav-YVLCYA7V.js} +85 -17
- package/dist/{config-3QZRWZJF.js → config-4HVOS65E.js} +88 -43
- package/dist/{configure-FL7Y3KJF.js → configure-PIWO7B24.js} +10 -10
- package/dist/{context-5HE7ODYK.js → context-IYEHL3WQ.js} +33 -31
- package/dist/{context-XNHL75JV.js → context-KQYIWPWT.js} +47 -34
- package/dist/{context-KYQFRVDC.js → context-N6ZE3LGJ.js} +11 -11
- package/dist/{context-clear-N545L53A.js → context-clear-G4OGZJDS.js} +33 -31
- package/dist/{context-working-set-QHKXSV2F.js → context-working-set-BWLF6LJP.js} +7 -7
- package/dist/{dispatch-runner-RGIE5PCT.js → dispatch-runner-2QQAITS3.js} +38 -38
- package/dist/{docs-5NAF6AU7.js → docs-PD3EXDKU.js} +21 -20
- package/dist/{doctor-ZGPEGHIP.js → doctor-LHBD36VU.js} +23 -22
- package/dist/{eval-GXLL44RD.js → eval-C45FYRJ6.js} +21 -20
- package/dist/{eval-inventory-HBWSWQOK.js → eval-inventory-6DEJPLBF.js} +2 -2
- package/dist/{evidence-HWLBRH3Q.js → evidence-6SHONYAF.js} +30 -28
- package/dist/{evolve-FTZBMNVW.js → evolve-KRKMV72X.js} +30 -28
- package/dist/{extensions-VHRBEID7.js → extensions-KPZ2UHBB.js} +5 -3
- package/dist/{fleet-CKZHJWZJ.js → fleet-IVTCKDHT.js} +62 -61
- package/dist/{fleet-commands-EXDXBMV6.js → fleet-commands-EDWL3IT7.js} +5 -5
- package/dist/{fleet-decisions-OTHB6KRL.js → fleet-decisions-YP3YEFGK.js} +4 -4
- package/dist/{fleet-graph-YTEZUCUT.js → fleet-graph-ZFWKHY2M.js} +16 -14
- package/dist/{fleet-inspect-SS6YMDCK.js → fleet-inspect-FVUNCBML.js} +31 -29
- package/dist/{fleet-preflight-PBY4VYOM.js → fleet-preflight-UN5XED4R.js} +2 -2
- package/dist/{fleet-validate-KMEM5L3S.js → fleet-validate-XOWC4HSX.js} +17 -15
- package/dist/{fleet-verify-QD5M7E7Q.js → fleet-verify-UN3SODEL.js} +30 -28
- package/dist/{fleet-view-WAMJYNDT.js → fleet-view-TWHJKCN6.js} +31 -29
- package/dist/{init-5XQRBOFV.js → init-T2QORQ3Y.js} +50 -49
- package/dist/{interop-34TVO25M.js → interop-IN5I2A66.js} +5 -5
- package/dist/{library-3QY6KF57.js → library-LSCATDLZ.js} +15 -13
- package/dist/{memory-L4UTIIIW.js → memory-HYOKAGGJ.js} +31 -29
- package/dist/{models-ZVX3QOWE.js → models-2GPMFYCM.js} +22 -21
- package/dist/{monitor-CEKVSYTS.js → monitor-E4ASVUJH.js} +34 -32
- package/dist/{orchestrator-77BAP6BC.js → orchestrator-DDMPR3PY.js} +984 -583
- package/dist/{panes-7STHOAUJ.js → panes-E3RUXOW5.js} +4 -4
- package/dist/{panes-SHAUIRXY.js → panes-IXKLOKA2.js} +23 -8
- package/dist/{reset-EOLM7GVE.js → reset-OAQP3W4O.js} +4 -4
- package/dist/{resources-74GKTLSF.js → resources-OTRSN34L.js} +15 -13
- package/dist/{run-HBAUJNNZ.js → run-5DEYH5QK.js} +60 -59
- package/dist/{share-G3APVLVP.js → share-IHWTLO3M.js} +19 -15
- package/dist/{skills-35HHUKCR.js → skills-IYMXMKW4.js} +17 -15
- package/dist/{skills-eval-QN4HSHDC.js → skills-eval-DROHSJAR.js} +36 -36
- package/dist/{skills-inventory-J357J34F.js → skills-inventory-D7X4L4ZX.js} +15 -13
- package/dist/{slash-commands-JZZCQA32.js → slash-commands-QBM7UZ3B.js} +21 -18
- package/dist/{steer-XAVHJM22.js → steer-Z5DO23FJ.js} +2 -2
- package/dist/{targets-DSM6CY3M.js → targets-P2FUC4IL.js} +25 -28
- package/dist/{terminal-lease-JOPFUVEM.js → terminal-lease-YREJ3JX2.js} +5 -5
- package/dist/{tools-MKNWVPBH.js → tools-5B7RO6MV.js} +4 -4
- package/dist/{trace-ECQ7TIYZ.js → trace-YMGMUM6A.js} +55 -7
- package/dist/{upgrade-H7TOM7YL.js → upgrade-PXK3S2YM.js} +11 -9
- package/dist/{usage-X52N3IDJ.js → usage-ME5MPXGX.js} +36 -34
- package/dist/{verifiers-EJTVVSMA.js → verifiers-BVZ7IWOO.js} +5 -5
- package/dist/{verify-YJL6XET2.js → verify-5K7ZKQFC.js} +4 -4
- package/dist/{web-fetch-MPIFL3LL.js → web-fetch-MPARV2K7.js} +2 -2
- package/dist/{wiki-generate-4NDZTQ4B.js → wiki-generate-F5W5QTYY.js} +48 -47
- package/dist/{with-panes-OBOBFIIR.js → with-panes-BYOJCLAM.js} +51 -255
- package/dist/worker/entry.js +45 -30
- package/docs/README.md +176 -81
- package/docs/{acp.md → architecture/acp.md} +36 -20
- package/docs/{alcf-provider.md → architecture/alcf-provider.md} +8 -5
- package/docs/{architecture.md → architecture/architecture.md} +43 -22
- package/docs/{artifact-placement.md → architecture/artifact-placement.md} +26 -23
- package/docs/architecture/artifact-versions.md +90 -0
- package/docs/{capacity-and-scheduling.md → architecture/capacity-and-scheduling.md} +26 -13
- package/docs/{context-engine.md → architecture/context-engine.md} +25 -25
- package/docs/{context-working-set.md → architecture/context-working-set.md} +13 -10
- package/docs/{dispatch-architecture-rationale.md → architecture/dispatch-architecture-rationale.md} +12 -9
- package/docs/{dispatch-typed-intent.md → architecture/dispatch-typed-intent.md} +68 -46
- package/docs/{evidence-and-memory.md → architecture/evidence-and-memory.md} +23 -16
- package/docs/{middleware-and-components.md → architecture/middleware-and-components.md} +11 -5
- package/docs/{model-catalog.md → architecture/model-catalog.md} +40 -17
- package/docs/{observability.md → architecture/observability.md} +26 -13
- package/docs/{pi-boundary.md → architecture/pi-boundary.md} +24 -11
- package/docs/{prompt-envelope-and-tools.md → architecture/prompt-envelope-and-tools.md} +55 -20
- package/docs/{provider-adapter-cookbook.md → architecture/provider-adapter-cookbook.md} +35 -24
- package/docs/{safety-model.md → architecture/safety-model.md} +20 -15
- package/docs/{session-lifecycle.md → architecture/session-lifecycle.md} +8 -5
- package/docs/architecture/time-conventions.md +125 -0
- package/docs/{trace-store.md → architecture/trace-store.md} +13 -5
- package/docs/{tui-design.md → architecture/tui-design.md} +13 -13
- package/docs/{worker-dispatch-mechanics.md → architecture/worker-dispatch-mechanics.md} +27 -30
- package/docs/{built-in-agents.md → guide/built-in-agents.md} +50 -34
- package/docs/{commands-and-modes.md → guide/commands-and-modes.md} +65 -60
- package/docs/{configuration-and-targets.md → guide/configuration-and-targets.md} +227 -289
- package/docs/guide/configuration-reference.md +1158 -0
- package/docs/{environment-variables.md → guide/environment-variables.md} +31 -28
- package/docs/{exit-codes-and-output.md → guide/exit-codes-and-output.md} +6 -3
- package/docs/{extensions-and-sharing.md → guide/extensions-and-sharing.md} +41 -14
- package/docs/{fleet-dispatch.md → guide/fleet-dispatch.md} +39 -43
- package/docs/{glossary.md → guide/glossary.md} +14 -11
- package/docs/{installation-and-lifecycle.md → guide/installation-and-lifecycle.md} +44 -13
- package/docs/guide/panes-and-files.md +290 -0
- package/docs/{proactive-memory.md → guide/proactive-memory.md} +79 -66
- package/docs/{resource-library.md → guide/resource-library.md} +13 -4
- package/docs/{skills-marketplace.md → guide/skills-marketplace.md} +7 -3
- package/docs/{tool-usage.md → guide/tool-usage.md} +87 -23
- package/docs/{troubleshooting.md → guide/troubleshooting.md} +9 -4
- package/docs/{config-knobs-audit.md → history/config-knobs-audit.md} +11 -11
- package/docs/{release-cut-checklist.md → history/release-cut-checklist.md} +29 -2
- package/docs/{development-pipeline.md → process/development-pipeline.md} +24 -26
- package/docs/process/documentation-coverage.md +100 -0
- package/docs/process/documentation-guide.md +187 -0
- package/docs/{eval-runner.md → process/eval-runner.md} +41 -50
- package/docs/{evals-internal.md → process/evals-internal.md} +10 -10
- package/docs/{evolution.md → process/evolution.md} +2 -2
- package/docs/{fleet-demo-runbook.md → process/fleet-demo-runbook.md} +11 -7
- package/docs/{git-commit-provenance.md → process/git-commit-provenance.md} +11 -4
- package/docs/{performance-methodology.md → process/performance-methodology.md} +87 -69
- package/docs/{scientific-validation.md → process/scientific-validation.md} +4 -4
- package/evals/README.md +2 -2
- package/package.json +9 -7
- package/skills/README.md +46 -37
- package/skills/coding/ast-grep/SKILL.md +2 -2
- package/skills/coding/coding-standards/SKILL.md +2 -2
- package/skills/coding/prototype/SKILL.md +2 -2
- package/skills/coding/tdd/SKILL.md +2 -2
- package/skills/context/context-handoff/SKILL.md +2 -2
- package/skills/context/context-prime/SKILL.md +2 -2
- package/skills/git/file-ticket/SKILL.md +2 -2
- package/skills/git/fix-issue/SKILL.md +3 -3
- package/skills/git/resolve-merge-conflicts/SKILL.md +2 -2
- package/skills/git/ship/SKILL.md +2 -2
- package/skills/git/worktree-create/SKILL.md +2 -2
- package/skills/git/worktree-merge/SKILL.md +2 -2
- package/skills/meta/clio-coder-dev/SKILL.md +9 -5
- package/skills/meta/clio-coder-dev/evals.md +3 -2
- package/skills/meta/clio-coder-test/SKILL.md +102 -95
- package/skills/meta/clio-coder-test/evals.md +9 -4
- package/skills/meta/clio-coder-test/references/harness.md +100 -124
- package/skills/meta/clio-coder-test/references/test-map.md +77 -50
- package/skills/meta/credentials/SKILL.md +2 -2
- package/skills/meta/find-skills/SKILL.md +2 -2
- package/skills/meta/herdr/SKILL.md +2 -2
- package/skills/meta/skill-craft/SKILL.md +22 -16
- package/skills/planning/architecture/SKILL.md +2 -2
- package/skills/planning/backlog/SKILL.md +2 -2
- package/skills/planning/prd/SKILL.md +2 -2
- package/skills/planning/product-intent/SKILL.md +2 -2
- package/skills/planning/tech-spec/SKILL.md +2 -2
- package/skills/registry.yaml +62 -62
- package/skills/research/arxiv-literature/SKILL.md +2 -2
- package/skills/research/experiment-protocol/SKILL.md +2 -2
- package/skills/research/scientific-debugging/SKILL.md +2 -2
- package/skills/research/scientific-modernization/SKILL.md +2 -2
- package/skills/skill-marketplace.json +62 -62
- package/skills/workflow/cut-it/SKILL.md +2 -2
- package/skills/workflow/design-council/SKILL.md +2 -2
- package/skills/workflow/grill-me/SKILL.md +2 -2
- package/skills/workflow/workflow-distiller/SKILL.md +2 -2
- package/src/cli/args.ts +2 -2
- package/src/cli/bootstrap-generate.ts +1 -1
- package/src/cli/config-inspect.ts +65 -12
- package/src/cli/configure.ts +0 -4
- package/src/cli/docs.ts +22 -14
- package/src/cli/doctor-naming.ts +5 -5
- package/src/cli/doctor-toolchain.ts +3 -3
- package/src/cli/eval.ts +1 -2
- package/src/cli/extensions.ts +2 -1
- package/src/cli/fleet.ts +1 -1
- package/src/cli/index.ts +2 -1
- package/src/cli/internal-dispatch.ts +3 -4
- package/src/cli/panes.ts +19 -5
- package/src/cli/run.ts +2 -2
- package/src/cli/share.ts +5 -1
- package/src/cli/skills-eval.ts +3 -3
- package/src/cli/targets.ts +2 -6
- package/src/cli/trace.ts +55 -4
- package/src/cli/wiki-generate.ts +1 -1
- package/src/core/artifact-paths.ts +1 -1
- package/src/core/bash-exec.ts +131 -86
- package/src/core/bus-events.ts +51 -6
- package/src/core/config.ts +5 -1
- package/src/core/defaults.ts +7 -4
- package/src/core/dispatch-outcome.ts +16 -0
- package/src/core/guardrails.ts +10 -49
- package/src/core/prompt-hint.ts +9 -0
- package/src/domains/agents/builtins/architect.md +2 -3
- package/src/domains/agents/builtins/coder.md +3 -2
- package/src/domains/agents/builtins/debugger.md +2 -2
- package/src/domains/agents/builtins/documenter.md +2 -2
- package/src/domains/agents/builtins/git-master.md +1 -1
- package/src/domains/agents/builtins/oracle.md +1 -1
- package/src/domains/agents/builtins/provenance.md +1 -1
- package/src/domains/agents/builtins/researcher.md +1 -1
- package/src/domains/agents/builtins/scout.md +1 -1
- package/src/domains/agents/builtins/tester.md +2 -2
- package/src/domains/agents/builtins/verifier.md +2 -2
- package/src/domains/agents/builtins/wiki-writer.md +1 -1
- package/src/domains/agents/catalog.ts +12 -14
- package/src/domains/agents/contract.ts +2 -0
- package/src/domains/agents/extension.ts +23 -1
- package/src/domains/config/keybindings.ts +8 -0
- package/src/domains/context/extension.ts +0 -3
- package/src/domains/context/working-set/path-index.ts +1 -0
- package/src/domains/dispatch/capability-match.ts +10 -0
- package/src/domains/dispatch/extension.ts +105 -22
- package/src/domains/dispatch/host-verification.ts +435 -39
- package/src/domains/dispatch/intent-requirements.ts +10 -0
- package/src/domains/dispatch/intent.ts +18 -1
- package/src/domains/dispatch/path-scope.ts +235 -24
- package/src/domains/dispatch/run-event-journal.ts +4 -15
- package/src/domains/dispatch/state.ts +2 -3
- package/src/domains/dispatch/transport.ts +45 -21
- package/src/domains/dispatch/types.ts +55 -3
- package/src/domains/eval/artifacts/store.ts +5 -0
- package/src/domains/eval/store.ts +8 -1
- package/src/domains/evidence/trust-status.ts +10 -1
- package/src/domains/extensions/contract.ts +15 -1
- package/src/domains/extensions/discovery.ts +238 -41
- package/src/domains/extensions/extension.ts +105 -6
- package/src/domains/extensions/index.ts +24 -0
- package/src/domains/extensions/integrity.ts +189 -0
- package/src/domains/extensions/manager.ts +17 -1
- package/src/domains/extensions/resource-path.ts +27 -0
- package/src/domains/extensions/resources.ts +18 -38
- package/src/domains/extensions/snapshot-store.ts +39 -0
- package/src/domains/extensions/snapshot.ts +180 -0
- package/src/domains/extensions/state.ts +385 -57
- package/src/domains/extensions/types.ts +118 -1
- package/src/domains/lifecycle/migrations/2026-09-01-extension-install-digests.ts +27 -0
- package/src/domains/lifecycle/migrations/index.ts +2 -0
- package/src/domains/lifecycle/naming-resources.ts +19 -4
- package/src/domains/lifecycle/naming-yazi.ts +10 -5
- package/src/domains/middleware/contract.ts +26 -0
- package/src/domains/middleware/extension.ts +24 -24
- package/src/domains/middleware/hook-receipts.ts +27 -4
- package/src/domains/middleware/hooks-io.ts +65 -32
- package/src/domains/middleware/hooks.ts +64 -0
- package/src/domains/middleware/index.ts +28 -4
- package/src/domains/middleware/registrations.ts +326 -0
- package/src/domains/middleware/runtime.ts +28 -0
- package/src/domains/middleware/snapshot.ts +20 -7
- package/src/domains/mux/contract.ts +38 -0
- package/src/domains/mux/detect.ts +6 -13
- package/src/domains/mux/index.ts +1 -1
- package/src/domains/mux/operations.ts +44 -5
- package/src/domains/mux/yazi/assets/yazi.toml +2 -2
- package/src/domains/mux/yazi/session.ts +53 -4
- package/src/domains/mux/yazi/theme.ts +117 -17
- package/src/domains/observability/contract.ts +10 -11
- package/src/domains/observability/extension.ts +11 -3
- package/src/domains/observability/projection.ts +14 -90
- package/src/domains/observability/trace-store.ts +43 -7
- package/src/domains/prompts/compiler.ts +73 -53
- package/src/domains/prompts/contract.ts +15 -3
- package/src/domains/prompts/extension.ts +97 -9
- package/src/domains/prompts/fragments/identity/clio-worker.md +1 -3
- package/src/domains/prompts/fragments/identity/clio.md +6 -12
- package/src/domains/prompts/fragments/identity/docs-routing.md +1 -2
- package/src/domains/prompts/fragments/identity/self-awareness.md +3 -11
- package/src/domains/prompts/fragments/operating/contract.md +7 -15
- package/src/domains/prompts/fragments/operating/delegation.md +32 -34
- package/src/domains/prompts/fragments/operating/skills.md +10 -24
- package/src/domains/prompts/fragments/operating/worker.md +1 -8
- package/src/domains/providers/index.ts +1 -1
- package/src/domains/providers/model-runtime-capabilities.ts +85 -21
- package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +669 -104
- package/src/domains/providers/runtime-resolution.ts +31 -0
- package/src/domains/providers/runtimes/common/probe-helpers.ts +7 -2
- package/src/domains/providers/runtimes/local-native/llamacpp.ts +9 -1
- package/src/domains/providers/types/cost-provenance.ts +19 -0
- package/src/domains/providers/types/local-model-quirks.ts +85 -37
- package/src/domains/resources/skills/loader.ts +16 -19
- package/src/domains/safety/call-target.ts +1 -1
- package/src/domains/safety/loop-detector.ts +7 -4
- package/src/domains/session/task-board.ts +10 -9
- package/src/domains/share/archive.ts +164 -7
- package/src/engine/acp/server.ts +62 -9
- package/src/engine/apis/llamacpp-residency.ts +3 -4
- package/src/engine/apis/lmstudio.ts +3 -3
- package/src/engine/apis/ollama-native.ts +6 -6
- package/src/engine/apis/openai-completions.ts +28 -25
- package/src/engine/apis/output-budget.ts +8 -18
- package/src/engine/apis/residency.ts +8 -27
- package/src/engine/gemma-channel-filter.ts +19 -0
- package/src/engine/loop-guard.ts +92 -12
- package/src/engine/worker-runtime.ts +40 -11
- package/src/engine/worker-tools.ts +3 -1
- package/src/entry/extension-hook-sources.ts +28 -0
- package/src/entry/extension-reload.ts +309 -0
- package/src/entry/orchestrator.ts +59 -35
- package/src/interactive/application-controller.ts +2 -1
- package/src/interactive/bus-notices.ts +8 -1
- package/src/interactive/chat-loop-messages.ts +3 -13
- package/src/interactive/chat-loop.ts +10 -1
- package/src/interactive/chat-panel.ts +36 -13
- package/src/interactive/chat-renderer.ts +71 -7
- package/src/interactive/dispatch-board.ts +6 -11
- package/src/interactive/footer/widgets.ts +13 -0
- package/src/interactive/interactive-application.ts +39 -4
- package/src/interactive/interactive-input-runtime.ts +4 -0
- package/src/interactive/interactive-presentation.ts +2 -2
- package/src/interactive/interactive-slash-runtime.ts +2 -0
- package/src/interactive/overlays/extensions.ts +9 -1
- package/src/interactive/overlays/help-reference.ts +13 -0
- package/src/interactive/overlays/settings.ts +27 -16
- package/src/interactive/panes-runtime.ts +111 -35
- package/src/interactive/prompt-cache-identity.ts +88 -0
- package/src/interactive/slash-commands.ts +129 -14
- package/src/interactive/stream-pacing-policy.ts +0 -23
- package/src/interactive/turn-context.ts +30 -15
- package/src/interactive/yazi-bridge.ts +60 -6
- package/src/tools/agent-tools.ts +30 -1
- package/src/tools/artifact.ts +2 -2
- package/src/tools/ask-user.ts +3 -3
- package/src/tools/bash.ts +1 -1
- package/src/tools/bootstrap.ts +4 -0
- package/src/tools/builtin-tool-catalog.ts +52 -22
- package/src/tools/codewiki/code-nav-surface.ts +6 -0
- package/src/tools/codewiki/code-nav.ts +99 -13
- package/src/tools/context/docs-engine.ts +20 -7
- package/src/tools/context/index.ts +29 -12
- package/src/tools/core-bootstrap.ts +28 -6
- package/src/tools/credential-present.ts +1 -2
- package/src/tools/dispatch-arguments.ts +5 -1
- package/src/tools/dispatch-plan.ts +48 -4
- package/src/tools/dispatch-run-events.ts +1 -1
- package/src/tools/dispatch-schema.ts +338 -0
- package/src/tools/dispatch-types.ts +3 -0
- package/src/tools/dispatch.ts +9 -254
- package/src/tools/ledger.ts +3 -5
- package/src/tools/monitor-surface.ts +5 -13
- package/src/tools/observation.ts +4 -5
- package/src/tools/panes-surface.ts +4 -11
- package/src/tools/panes.ts +4 -2
- package/src/tools/policy.ts +15 -2
- package/src/tools/read.ts +5 -6
- package/src/tools/registry.ts +30 -7
- package/src/tools/result-shaping.ts +18 -14
- package/src/tools/steer-surface.ts +1 -1
- package/src/tools/tasks.ts +1 -1
- package/src/tools/truncate.ts +6 -5
- package/src/tools/verify/surface.ts +6 -12
- package/src/tools/web-fetch-surface.ts +1 -3
- package/dist/chunk-5QIAJV2D.js +0 -48
- package/dist/chunk-JZWT5J3Y.js +0 -814
- package/dist/chunk-K7VKOLQQ.js +0 -15
- package/dist/chunk-PMZCIOCJ.js +0 -25
- package/dist/chunk-SUW5DORT.js +0 -819
- package/dist/chunk-UOV2BYIW.js +0 -107
- package/dist/chunk-WR6U3OVP.js +0 -45
- package/docs/artifact-versions.md +0 -67
- package/docs/documentation-coverage.md +0 -46
- package/docs/documentation-guide.md +0 -167
- package/docs/time-conventions.md +0 -101
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
# Clio Coder Local Evaluation Runner
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Clio Coder Local Evaluation Runner visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/eval_blueprint.html).
|
|
5
5
|
|
|
6
6
|
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
7
|
|
|
8
|
-
Source of truth: [src/domains/eval/](
|
|
8
|
+
Source of truth: [src/domains/eval/](../../src/domains/eval/) and [src/cli/eval.ts](../../src/cli/eval.ts).
|
|
9
9
|
|
|
10
10
|
---
|
|
11
11
|
|
|
@@ -20,6 +20,7 @@ clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--cl
|
|
|
20
20
|
clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
|
|
21
21
|
clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
|
|
22
22
|
clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
|
|
23
|
+
clio-coder eval inventory --json
|
|
23
24
|
```
|
|
24
25
|
|
|
25
26
|
### Command Roles
|
|
@@ -33,6 +34,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
|
|
|
33
34
|
* `junit`: XML report for CI/CD integration.
|
|
34
35
|
* **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
|
|
35
36
|
* **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
|
|
37
|
+
* **`inventory`**: Prints the fixed machine-readable inventory used by GUI hosts. It includes stored report identity, provenance, serving facts, accounting, and per-scenario outcomes without report attachments.
|
|
36
38
|
|
|
37
39
|
Exit codes:
|
|
38
40
|
|
|
@@ -42,7 +44,7 @@ Exit codes:
|
|
|
42
44
|
| `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
|
|
43
45
|
| `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
|
|
44
46
|
| `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
|
|
45
|
-
| `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for
|
|
47
|
+
| `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for hard failures, unreadable inputs, and malformed threshold files; `2` for an invalid eval ID or usage error |
|
|
46
48
|
|
|
47
49
|
---
|
|
48
50
|
|
|
@@ -75,7 +77,7 @@ tasks:
|
|
|
75
77
|
excludes:
|
|
76
78
|
- "**/node_modules/**"
|
|
77
79
|
runner:
|
|
78
|
-
kind: "clio-run" # clio-run | context-index | context-init | external-command
|
|
80
|
+
kind: "clio-coder-run" # clio-coder-run | context-index | context-init | external-command
|
|
79
81
|
prompt: "Optimize the FFT tolerance bounds in solver.ts"
|
|
80
82
|
timeoutMs: 60000
|
|
81
83
|
verify:
|
|
@@ -103,12 +105,12 @@ tasks:
|
|
|
103
105
|
| --- | --- | --- |
|
|
104
106
|
| `version` | - | Must equal `2`. |
|
|
105
107
|
| `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
|
|
106
|
-
| `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count,
|
|
107
|
-
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
|
|
108
|
-
| `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
|
|
109
|
-
| `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
|
|
108
|
+
| `matrix` | `targets[]`, `repeats`, `dimensions[]`, `maxCostUsd` | Matrix of execution targets, repetition count, execution-envelope fields intentionally varied by the suite, and an optional cumulative known-cost ceiling. |
|
|
109
|
+
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes`, `setup` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). Optional `setup` commands prepare the workspace before the runner starts. |
|
|
110
|
+
| `runner` | `kind`, `prompt`, `autonomy`, `agent`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-coder-run` (starts Clio's agent loop), `context-index` (runs the indexer), `context-init` (initializes context), or `external-command` (spawns a subprocess). `agent` selects a worker recipe and `autonomy` sets one-run headless authority. |
|
|
111
|
+
| `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio-coder.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
|
|
110
112
|
| `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
|
|
111
|
-
| `metrics` | `collect` |
|
|
113
|
+
| `metrics` | `collect`, `readObservation` | Metric names to compile plus optional public allowlisted and decoy paths reduced to bounded read counters. Raw path strings do not enter behavioral facts. |
|
|
112
114
|
|
|
113
115
|
---
|
|
114
116
|
|
|
@@ -120,7 +122,7 @@ tasks:
|
|
|
120
122
|
---
|
|
121
123
|
|
|
122
124
|
## Runner Kinds
|
|
123
|
-
* **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
|
|
125
|
+
* **`clio-coder-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools. The released `clio-run` spelling is accepted only as a legacy input alias and is normalized before validation; writers and new suites use `clio-coder-run`.
|
|
124
126
|
* **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
|
|
125
127
|
* **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
|
|
126
128
|
* **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
|
|
@@ -190,7 +192,7 @@ export interface EvalArtifactV4 {
|
|
|
190
192
|
version: 4;
|
|
191
193
|
evalId: string;
|
|
192
194
|
suite: { id: string; hash: string };
|
|
193
|
-
|
|
195
|
+
clioCoder: EvalClioProvenance;
|
|
194
196
|
environment: EvalEnvironmentProvenance;
|
|
195
197
|
matrix: { target: string; model: string | null; thinking: string | null };
|
|
196
198
|
summary: EvalArtifactSummaryV4;
|
|
@@ -206,11 +208,11 @@ export interface EvalArtifactV4 {
|
|
|
206
208
|
|
|
207
209
|
## The verdict envelope
|
|
208
210
|
|
|
209
|
-
Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
|
|
211
|
+
Every result carries a strictly parsed `clio-coder.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
|
|
210
212
|
|
|
211
213
|
```json
|
|
212
214
|
{
|
|
213
|
-
"schema": "clio.eval.verdict.v1",
|
|
215
|
+
"schema": "clio-coder.eval.verdict.v1",
|
|
214
216
|
"scenarioId": "latency-nonnegative",
|
|
215
217
|
"trialIndex": 0,
|
|
216
218
|
"outcome": "pass",
|
|
@@ -230,7 +232,7 @@ The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`,
|
|
|
230
232
|
|
|
231
233
|
### Behavioral scenario and verdict documents
|
|
232
234
|
|
|
233
|
-
Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
|
|
235
|
+
Behavioral evaluation is additive and does not change the persisted `clio-coder.eval.verdict.v1` reader. A Suite v2 task may declare a `clio-coder.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio-coder.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable. Readers normalize the released `clio.eval.*` identifiers for compatibility, but current writers emit only `clio-coder.eval.*` identifiers.
|
|
234
236
|
|
|
235
237
|
The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
|
|
236
238
|
|
|
@@ -240,8 +242,8 @@ Suite execution adapts scalar run metrics into these observable facts at the Sui
|
|
|
240
242
|
|
|
241
243
|
### Public built-in behavioral corpus
|
|
242
244
|
|
|
243
|
-
The repository
|
|
244
|
-
`
|
|
245
|
+
The source repository carries corpus `public-built-in-behavior` version `1.0.0`
|
|
246
|
+
under `evals/`. It contains no private prompts, endpoints, credentials, or
|
|
245
247
|
mutable external dataset:
|
|
246
248
|
|
|
247
249
|
- `behavioral-machinery.yaml` provides one positive and one adversarial
|
|
@@ -262,12 +264,15 @@ mutable external dataset:
|
|
|
262
264
|
that the rules can reject observed model behavior rather than merely restate
|
|
263
265
|
aggregate success counters.
|
|
264
266
|
|
|
265
|
-
|
|
267
|
+
These are source-checkout workflows: the npm archive keeps the inputs for
|
|
268
|
+
inspection and reproducibility, but the deterministic TypeScript driver uses
|
|
269
|
+
the repository development toolchain. Build once, then run either focused
|
|
270
|
+
suite from the repository root:
|
|
266
271
|
|
|
267
272
|
```sh
|
|
268
|
-
node dist/cli/index.js eval run --suite
|
|
269
|
-
node dist/cli/index.js eval run --suite
|
|
270
|
-
node dist/cli/index.js eval run --suite
|
|
273
|
+
node dist/cli/index.js eval run --suite evals/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
|
|
274
|
+
node dist/cli/index.js eval run --suite evals/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
275
|
+
node dist/cli/index.js eval run --suite evals/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
271
276
|
```
|
|
272
277
|
|
|
273
278
|
The machinery tasks use the repository read-only and create only private
|
|
@@ -306,8 +311,8 @@ A dispatched worker's receipt reports `sessionId: null` and writes no session ar
|
|
|
306
311
|
|
|
307
312
|
### Behavioral multi-metric results
|
|
308
313
|
|
|
309
|
-
A result with a `clio.eval.behavior.v1` verdict also carries the additive
|
|
310
|
-
`clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
|
|
314
|
+
A result with a `clio-coder.eval.behavior.v1` verdict also carries the additive
|
|
315
|
+
`clio-coder.eval.behavior.metrics.v1` projection. The projection binds the scenario
|
|
311
316
|
to its role and target/model envelope and records one `number | null`
|
|
312
317
|
observation for each closed metric. The source travels beside every value:
|
|
313
318
|
|
|
@@ -356,8 +361,8 @@ is emitted as testcase output rather than a failed testcase.
|
|
|
356
361
|
### Execution-envelope provenance and comparability
|
|
357
362
|
|
|
358
363
|
Every newly written behavioral result carries an additive
|
|
359
|
-
`clio.eval.execution-envelope.v1` sibling. Artifact v4,
|
|
360
|
-
`clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
|
|
364
|
+
`clio-coder.eval.execution-envelope.v1` sibling. Artifact v4,
|
|
365
|
+
`clio-coder.eval.verdict.v1`, and `clio-coder.eval.behavior.metrics.v1` retain their
|
|
361
366
|
existing identities. The envelope records the selected prompt fragment ids,
|
|
362
367
|
authored versions or `unversioned` marker, fragment content hashes, prompt
|
|
363
368
|
composition hash, recipe id/version/fingerprint when a worker recipe applies,
|
|
@@ -382,31 +387,15 @@ metric means and variances. When the prompt or recipe identity changes, the
|
|
|
382
387
|
generated evidence names each affected corpus scenario and role instead of
|
|
383
388
|
hiding it behind an aggregate score.
|
|
384
389
|
|
|
385
|
-
###
|
|
390
|
+
### Reference behavioral baseline
|
|
386
391
|
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
the evidence, run the same machinery suite first, inspect the failing diff and
|
|
395
|
-
the named affected corpus results, then update explicitly:
|
|
396
|
-
|
|
397
|
-
```sh
|
|
398
|
-
npm run build
|
|
399
|
-
node benchmarks/eval/check-behavioral-release.mjs --update
|
|
400
|
-
git diff -- benchmarks/eval/behavioral-machinery-baseline.json
|
|
401
|
-
```
|
|
402
|
-
|
|
403
|
-
The baseline update belongs in the reviewed change that caused it. Do not use
|
|
404
|
-
the update command merely to make a red gate green. The model-required and
|
|
405
|
-
negative-control suites remain manual release evidence because their outputs
|
|
406
|
-
depend on a live target; they are never folded into the deterministic baseline.
|
|
407
|
-
The projection excludes `latency.wallMs` because scheduler timing is not stable
|
|
408
|
-
evidence. Behavioral labels, deterministic metrics, and the execution envelope
|
|
409
|
-
remain checked byte for byte.
|
|
392
|
+
`evals/behavioral-machinery-baseline.json` is retained as reviewable reference
|
|
393
|
+
evidence for the machinery corpus. It is not a CI or release gate. Run the
|
|
394
|
+
current `evals/behavioral-machinery.yaml` through the built CLI when a prompt,
|
|
395
|
+
recipe, policy, or expected-behavior change needs a fresh measurement, inspect
|
|
396
|
+
the named scenario evidence, and update any retained baseline deliberately in
|
|
397
|
+
the reviewed change. Model-required and negative-control suites remain manual
|
|
398
|
+
measurements tied to their exact target and serving configuration.
|
|
410
399
|
|
|
411
400
|
### Hard thresholds and informational budgets
|
|
412
401
|
|
|
@@ -443,10 +432,12 @@ cannot offset a task or safety regression.
|
|
|
443
432
|
|
|
444
433
|
```text
|
|
445
434
|
serving configuration drift; pass --allow-config-drift to compare these runs
|
|
446
|
-
baseline serving: target=mini runtime=llamacpp model
|
|
435
|
+
baseline serving: target=mini runtime=llamacpp model=ornith1.5-35b-moe server_build=... total_slots=4 thinking=off compiled_prompt_hash=...
|
|
447
436
|
candidate serving: ...
|
|
448
437
|
```
|
|
449
438
|
|
|
439
|
+
The current reference `mini` endpoint is the llama.cpp router at `192.168.86.141:8080`. It serves `ornith1.5-35b-moe` with four parallel slots and 262144 context tokens per slot. These deployment facts are reference topology, not defaults imposed on another target; retain the artifact's observed serving configuration with every comparison.
|
|
440
|
+
|
|
450
441
|
`--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
|
|
451
442
|
|
|
452
443
|
---
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Internal Eval Suites
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Internal Eval Suites visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evals_internal_blueprint.html).
|
|
5
5
|
|
|
6
6
|
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
7
|
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
@@ -15,13 +15,13 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
|
|
|
15
15
|
```
|
|
16
16
|
|
|
17
17
|
Use `--out <dir>` when the artifact should be written outside the default Clio
|
|
18
|
-
data directory.
|
|
19
|
-
|
|
20
|
-
|
|
18
|
+
data directory. External benchmark campaigns should adapt their cases and
|
|
19
|
+
grader observations into the same eval engine while keeping private datasets,
|
|
20
|
+
credentials, endpoints, and raw artifacts outside this repository.
|
|
21
21
|
|
|
22
22
|
The public behavioral corpus is the deliberate exception to the otherwise
|
|
23
23
|
private Suite v2 data policy. Its reviewable, synthetic suites live under
|
|
24
|
-
`
|
|
24
|
+
`evals/`: a model-free positive/adversarial authority pair for every
|
|
25
25
|
built-in worker recipe, four tiny main-agent model scenarios covering all eight
|
|
26
26
|
behavioral categories with event- and grader-derived facts, and an intentional
|
|
27
27
|
decoy negative control. The model-free driver uses the shipped recipe catalog,
|
|
@@ -83,7 +83,7 @@ thinking level is measuring the server, not the change under test.
|
|
|
83
83
|
|
|
84
84
|
The verdict envelope keeps its original `behavioral: null` field for compatibility.
|
|
85
85
|
A suite that declares a versioned behavioral scenario records the result as a
|
|
86
|
-
separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
|
|
86
|
+
separate `clio-coder.eval.behavior.v1` document on the Artifact v4 result, cross-linked
|
|
87
87
|
to the unchanged verdict identity. Its labels come only from bounded transcript,
|
|
88
88
|
tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
|
|
89
89
|
`machinery: "infrastructure_failure"`, which the parser refuses to pair with a
|
|
@@ -210,7 +210,7 @@ tasks:
|
|
|
210
210
|
- node_modules
|
|
211
211
|
- dist
|
|
212
212
|
runner:
|
|
213
|
-
kind: clio-run
|
|
213
|
+
kind: clio-coder-run
|
|
214
214
|
prompt: Fix the intentionally broken function so the local verifier passes.
|
|
215
215
|
verify:
|
|
216
216
|
commands:
|
|
@@ -271,7 +271,7 @@ tasks:
|
|
|
271
271
|
- dist
|
|
272
272
|
- .clio-coder
|
|
273
273
|
runner:
|
|
274
|
-
kind: clio-run
|
|
274
|
+
kind: clio-coder-run
|
|
275
275
|
prompt: Summarize the repository purpose and make no file changes.
|
|
276
276
|
verify:
|
|
277
277
|
forbidPaths:
|
|
@@ -305,7 +305,7 @@ tasks:
|
|
|
305
305
|
- dist
|
|
306
306
|
- .clio-coder
|
|
307
307
|
runner:
|
|
308
|
-
kind: clio-run
|
|
308
|
+
kind: clio-coder-run
|
|
309
309
|
prompt: Fix the failing unit test with the smallest source change.
|
|
310
310
|
verify:
|
|
311
311
|
commands:
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Evolution and Change Manifests
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Evolution and Change Manifests visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evolution_blueprint.html).
|
|
5
5
|
|
|
6
6
|
Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
|
|
7
7
|
|
|
@@ -1,10 +1,13 @@
|
|
|
1
1
|
# Fleet Demo Runbook
|
|
2
2
|
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Fleet Demo Runbook visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/fleet_demo_blueprint.html).
|
|
5
|
+
|
|
3
6
|
A repeatable multi-node demonstration: one orchestrator drives a real
|
|
4
7
|
CMake/C++ fix through a reviewer-gated dispatch across SSH nodes, and every
|
|
5
8
|
worker's receipt (including the remote ones) verifies afterward. The steps
|
|
6
9
|
are executable in order; this document doubles as the recording script.
|
|
7
|
-
Background and reference: [fleet-dispatch.md](fleet-dispatch.md).
|
|
10
|
+
Background and reference: [fleet-dispatch.md](../guide/fleet-dispatch.md).
|
|
8
11
|
|
|
9
12
|
## Reference fabric
|
|
10
13
|
|
|
@@ -139,9 +142,9 @@ clio-coder evidence inspect <evidenceId>
|
|
|
139
142
|
run ledger; a tampered or mismatched receipt fails the build with the field
|
|
140
143
|
that diverged. The receipts of the remote runs verify on the orchestrator host because the
|
|
141
144
|
ledger and receipts live on the shared filesystem. Current receipts use strict
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
+
v20 and authenticate every current receipt and reconstructed-ledger field.
|
|
146
|
+
Lower versions are reported as retired and are never read as evidence or
|
|
147
|
+
migrated; malformed and future shapes fail verification.
|
|
145
148
|
|
|
146
149
|
## Provenance walkthrough: what a PI can verify from receipts alone
|
|
147
150
|
|
|
@@ -171,9 +174,10 @@ reconstruct:
|
|
|
171
174
|
complete receipt schema and its stable ledger row. `clio-coder evidence build
|
|
172
175
|
--run <id>` recomputes and cross-checks it; `verifyReceiptIntegrity` in
|
|
173
176
|
`src/domains/dispatch/receipt-integrity.ts` is the reference
|
|
174
|
-
implementation. Current receipts use
|
|
175
|
-
|
|
176
|
-
read as evidence through
|
|
177
|
+
implementation. Current receipts use v20. Lower versions are reported as
|
|
178
|
+
retired, while malformed and future shapes fail verification. Incompatible
|
|
179
|
+
state may be archived for inspection, but it is never read as evidence through
|
|
180
|
+
a compatibility verifier.
|
|
177
181
|
|
|
178
182
|
The walkthrough for an audience is three commands: `clio-coder evidence build
|
|
179
183
|
--run <id>` (it verifies), open the receipt JSON (read `node`, `gate`,
|
|
@@ -1,13 +1,20 @@
|
|
|
1
1
|
# Git Commit Provenance
|
|
2
2
|
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Git Commit Provenance visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/git_commit_provenance_blueprint.html).
|
|
5
|
+
|
|
3
6
|
Clio Coder adds evidence-aware role trailers to commits created through Clio.
|
|
4
7
|
The feature is enabled by default:
|
|
5
8
|
|
|
6
9
|
```yaml
|
|
7
|
-
|
|
8
|
-
|
|
10
|
+
integrations:
|
|
11
|
+
git:
|
|
12
|
+
commitAttribution: true
|
|
9
13
|
```
|
|
10
14
|
|
|
15
|
+
The released `attribution.gitCommits` path is a migration alias. Current
|
|
16
|
+
settings files and writers use `integrations.git.commitAttribution`.
|
|
17
|
+
|
|
11
18
|
Settings -> Advanced exposes the same switch as **Clio commit provenance**, with
|
|
12
19
|
`enabled` and `disabled` values. A change applies immediately to subsequent
|
|
13
20
|
commits in the session. When disabled, Clio leaves commit messages entirely
|
|
@@ -49,11 +56,11 @@ Co-authored-by: Clio Coder <clio-coder@iowarp.ai>
|
|
|
49
56
|
Existing human trailers stay in place. A Clio trailer already present in any
|
|
50
57
|
letter case is respected rather than repeated, line endings are normalized only
|
|
51
58
|
while attribution is enabled, and repeated processing is idempotent. When a directly relevant
|
|
52
|
-
receipt-
|
|
59
|
+
receipt-v20 digest passes integrity verification, Clio may additionally add the
|
|
53
60
|
full digest:
|
|
54
61
|
|
|
55
62
|
```text
|
|
56
|
-
Clio-Evidence: receipt-
|
|
63
|
+
Clio-Evidence: receipt-v20/sha256:<64-character digest>
|
|
57
64
|
```
|
|
58
65
|
|
|
59
66
|
Clio does not invent, shorten, or add an unrelated digest. The role trailers do
|