@iowarp/clio-coder 0.4.1 → 0.4.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +127 -0
- package/CONTRIBUTING.md +142 -52
- package/README.md +434 -473
- package/SECURITY.md +2 -1
- package/dist/{acp-ZILU3AUO.js → acp-H2NGRPWO.js} +12 -12
- package/dist/{agents-HYWGBGQR.js → agents-TL5LLUQP.js} +56 -55
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-N3QT7CBO.js → auth-E5SW4HMS.js} +23 -21
- package/dist/builtins-IA7V7FUC.js +22 -0
- package/dist/{chunk-7RY5VZPH.js → chunk-2APPQIER.js} +8 -8
- package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
- package/dist/{chunk-JA5QWE4Z.js → chunk-2UG5F4C5.js} +1973 -1664
- package/dist/{chunk-5YHDIDBP.js → chunk-2UH2KFUP.js} +2 -2
- package/dist/{chunk-CTJ4RNAA.js → chunk-2VIKGWFZ.js} +2 -2
- package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
- package/dist/{chunk-GIZNH63R.js → chunk-35MSIRKH.js} +9 -4
- package/dist/chunk-3EBYEESD.js +314 -0
- package/dist/{chunk-J5LZHVIT.js → chunk-3M6DQK6S.js} +113 -35
- package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
- package/dist/{chunk-6FN3E6KX.js → chunk-4O6MANBS.js} +2 -2
- package/dist/chunk-4UVU7BJ5.js +39 -0
- package/dist/{chunk-VKRH2TCS.js → chunk-4WR7VSYB.js} +2 -2
- package/dist/{chunk-BBTJOK6Y.js → chunk-54CBCGIR.js} +5 -5
- package/dist/{chunk-AP73CFDC.js → chunk-5ICU3EUH.js} +2 -2
- package/dist/chunk-5MEZN6CB.js +1334 -0
- package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
- package/dist/{chunk-ABLSQ6JX.js → chunk-64I3JVYM.js} +8 -2
- package/dist/{chunk-AFKWHWXF.js → chunk-6PTFB5VS.js} +39 -22
- package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
- package/dist/chunk-7DRAWPTZ.js +360 -0
- package/dist/chunk-7E7I3WLS.js +3762 -0
- package/dist/{chunk-BJGUKIG4.js → chunk-7ZYNNDKC.js} +7 -7
- package/dist/{chunk-XKA2ICR3.js → chunk-AF4YM7Z4.js} +652 -252
- package/dist/{chunk-GVQJ5CCZ.js → chunk-AX2THNSA.js} +12 -12
- package/dist/{chunk-IG7BCQBA.js → chunk-B4OAX3SI.js} +65 -3
- package/dist/{chunk-TD3PGPQA.js → chunk-B4VEBZKF.js} +3 -3
- package/dist/{chunk-74YWRRU5.js → chunk-BEPZRGGU.js} +10 -10
- package/dist/{chunk-FEFIFZTL.js → chunk-CE5AX47J.js} +2 -2
- package/dist/{chunk-UAPGZHYC.js → chunk-DWUOQKRU.js} +25 -11
- package/dist/{chunk-THYWACCR.js → chunk-E3TPLWFX.js} +3 -3
- package/dist/{chunk-7EPLI7VL.js → chunk-EKCHAPYA.js} +2 -2
- package/dist/{chunk-HLW2MRKE.js → chunk-F4EKGO4N.js} +3 -1
- package/dist/{chunk-PJJ6MY27.js → chunk-F5JHEYZM.js} +7 -7
- package/dist/{chunk-6CCS4G3W.js → chunk-FTMGRKEF.js} +3 -3
- package/dist/{chunk-SINK3QR6.js → chunk-G76U63X4.js} +17 -17
- package/dist/{chunk-EIMVLWB3.js → chunk-GHS5EBTQ.js} +64 -9
- package/dist/{chunk-QMXC4JB7.js → chunk-GI7YYQ3F.js} +187 -1419
- package/dist/{chunk-TZSKNMZG.js → chunk-GTUD2WMY.js} +2 -1
- package/dist/{chunk-6HMJX2VU.js → chunk-GWZNEVM2.js} +44 -12
- package/dist/chunk-GYV6VZOC.js +26 -0
- package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
- package/dist/{chunk-UXN6JT4W.js → chunk-HEQY7ZFI.js} +3 -3
- package/dist/{chunk-7PWAODYW.js → chunk-I7XBWTYH.js} +2 -2
- package/dist/{chunk-GCSMB2KY.js → chunk-I7ZPNEJM.js} +145 -102
- package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
- package/dist/{chunk-QTFGO774.js → chunk-IGLP3ODT.js} +29 -16
- package/dist/chunk-IJNZMHLA.js +101 -0
- package/dist/{chunk-BDPT6GTK.js → chunk-INY6HTFL.js} +7 -7
- package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
- package/dist/{chunk-6NJQITNH.js → chunk-IWT4SF4R.js} +6 -3
- package/dist/{chunk-R23Z6K6I.js → chunk-JDAY6FIL.js} +19 -19
- package/dist/chunk-JEQ3XTHC.js +42 -0
- package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
- package/dist/{chunk-TVH4ONAM.js → chunk-JKKCYP3C.js} +10 -10
- package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
- package/dist/{chunk-C537JADH.js → chunk-KK4JZPBQ.js} +19 -141
- package/dist/{chunk-K6BF4U2H.js → chunk-KKOJXO6R.js} +62 -14
- package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
- package/dist/{chunk-6DWBAZ5U.js → chunk-L47TF46W.js} +5 -7
- package/dist/{chunk-HUAS7ITX.js → chunk-LDJG7DW3.js} +91 -42
- package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
- package/dist/{chunk-YPI3QQCF.js → chunk-MCEPRMZW.js} +2 -4
- package/dist/{chunk-Y4CAGMM6.js → chunk-MNJGS2IN.js} +5 -6
- package/dist/{chunk-VKFQTNDV.js → chunk-MUW2BDDH.js} +4 -4
- package/dist/{chunk-E67WX76H.js → chunk-MWUZBSAQ.js} +104 -152
- package/dist/{chunk-OJTRZGR3.js → chunk-N2Z7HLVY.js} +21 -21
- package/dist/{chunk-TVHHYFHE.js → chunk-NEDJ26B5.js} +2 -2
- package/dist/{chunk-FYUN5KZ3.js → chunk-NIQJ66N4.js} +21 -21
- package/dist/{chunk-U2WB7TZS.js → chunk-NMJXSHBJ.js} +97 -85
- package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
- package/dist/{chunk-VEGN6WIQ.js → chunk-O5CVSAG5.js} +3 -3
- package/dist/{chunk-MOPSG2X7.js → chunk-OML5D5V5.js} +8 -8
- package/dist/{chunk-2VG7KLYV.js → chunk-PAJQJ7BS.js} +5816 -3255
- package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
- package/dist/{chunk-BTGG6BG2.js → chunk-QWGDJJYJ.js} +158 -19
- package/dist/chunk-R6Q67RJH.js +134 -0
- package/dist/{chunk-ZJLUDYFY.js → chunk-RRNP2ANY.js} +6 -6
- package/dist/{chunk-PVAMAVBB.js → chunk-RSJ25QSL.js} +102 -2
- package/dist/{chunk-NLFAQR7Z.js → chunk-S66XZJOF.js} +3 -23
- package/dist/chunk-SKHCAU7K.js +385 -0
- package/dist/chunk-SZAA6XDG.js +30 -0
- package/dist/{chunk-J4HBWF6Y.js → chunk-TM6LQDI3.js} +131 -28
- package/dist/chunk-UOIZ7DA4.js +41 -0
- package/dist/{chunk-MA3H6DM5.js → chunk-UPZU6GE4.js} +25 -3
- package/dist/{chunk-BWW4HLO4.js → chunk-UXCU4E3T.js} +8 -6
- package/dist/{chunk-N5UK64DP.js → chunk-V2ANDPVT.js} +4 -4
- package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
- package/dist/{chunk-6VC4OV3Z.js → chunk-VIA6RFQZ.js} +3 -11
- package/dist/{chunk-ZAZB4JMW.js → chunk-VKPAQYEB.js} +27 -8
- package/dist/{chunk-QKIFBZKT.js → chunk-VW6DOEDG.js} +497 -81
- package/dist/{chunk-SCYB3HA4.js → chunk-W6RRQCPQ.js} +63 -19
- package/dist/{chunk-2NM363SV.js → chunk-WBKFA554.js} +10 -10
- package/dist/{chunk-R32CLGZ6.js → chunk-WCXUNS7U.js} +82 -21
- package/dist/{chunk-GPPB3JBE.js → chunk-WRBAGUNF.js} +3 -3
- package/dist/{chunk-IXJT6DCX.js → chunk-XIVNBFZS.js} +85 -30
- package/dist/{chunk-UEDMSP56.js → chunk-XPWWI35G.js} +417 -201
- package/dist/chunk-XRZT5WY5.js +47 -0
- package/dist/{chunk-3QSOM6PA.js → chunk-Y3CBHOR6.js} +2 -2
- package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
- package/dist/{chunk-AKB4GYDL.js → chunk-YQWYVTMC.js} +5 -5
- package/dist/{chunk-6I5ILFOF.js → chunk-ZA4VCIGV.js} +3 -3
- package/dist/{chunk-7OBGU7UB.js → chunk-ZDN3Y73Y.js} +12 -18
- package/dist/{chunk-3I5NY75V.js → chunk-ZWPRK62N.js} +8 -5
- package/dist/cli/index.js +41 -39
- package/dist/{clio-IT3G3VQH.js → clio-CMMK4KRR.js} +9 -9
- package/dist/{code-nav-RK6S7F6E.js → code-nav-MDZNQS33.js} +89 -21
- package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
- package/dist/{config-3QZRWZJF.js → config-SVM5P5YI.js} +131 -84
- package/dist/{configure-FL7Y3KJF.js → configure-LE3IK2TJ.js} +28 -26
- package/dist/{context-5HE7ODYK.js → context-2OHRKS42.js} +69 -64
- package/dist/{context-KYQFRVDC.js → context-E3VC7RX5.js} +15 -11
- package/dist/{context-XNHL75JV.js → context-VNCR7KAG.js} +93 -65
- package/dist/{context-clear-N545L53A.js → context-clear-BW4O37TG.js} +64 -60
- package/dist/context-map-COB37XXN.js +505 -0
- package/dist/{context-working-set-QHKXSV2F.js → context-working-set-VDS25HXZ.js} +19 -18
- package/dist/{dispatch-runner-RGIE5PCT.js → dispatch-runner-5AHT53RF.js} +93 -82
- package/dist/{docs-5NAF6AU7.js → docs-PD3EXDKU.js} +21 -20
- package/dist/{doctor-ZGPEGHIP.js → doctor-WNNVO6FY.js} +48 -47
- package/dist/{eval-GXLL44RD.js → eval-7G7SGAYO.js} +287 -115
- package/dist/{eval-inventory-HBWSWQOK.js → eval-inventory-Y6QRFOH5.js} +4 -4
- package/dist/{evidence-HWLBRH3Q.js → evidence-VD6736FQ.js} +67 -64
- package/dist/{evolve-FTZBMNVW.js → evolve-AL3NGVRL.js} +65 -62
- package/dist/{extensions-VHRBEID7.js → extensions-MOVJ32NM.js} +9 -7
- package/dist/{fleet-CKZHJWZJ.js → fleet-QZHUMAGI.js} +114 -111
- package/dist/{fleet-commands-EXDXBMV6.js → fleet-commands-BAYT5FJZ.js} +10 -10
- package/dist/{fleet-decisions-OTHB6KRL.js → fleet-decisions-IREVMRU4.js} +7 -6
- package/dist/{fleet-graph-YTEZUCUT.js → fleet-graph-YCTT3HTI.js} +22 -19
- package/dist/{fleet-inspect-SS6YMDCK.js → fleet-inspect-QVJTDAVB.js} +58 -55
- package/dist/{fleet-preflight-PBY4VYOM.js → fleet-preflight-25QAFPK4.js} +4 -4
- package/dist/{fleet-validate-KMEM5L3S.js → fleet-validate-5O57AAJ7.js} +26 -23
- package/dist/{fleet-verify-QD5M7E7Q.js → fleet-verify-CPH2W2T6.js} +59 -56
- package/dist/{fleet-view-WAMJYNDT.js → fleet-view-SWBR3VGQ.js} +58 -55
- package/dist/{init-5XQRBOFV.js → init-J477LKZH.js} +82 -79
- package/dist/{interop-34TVO25M.js → interop-3FCM6XLG.js} +11 -11
- package/dist/{library-3QY6KF57.js → library-QUQEIUG6.js} +30 -27
- package/dist/{memory-L4UTIIIW.js → memory-SGGSEP65.js} +67 -64
- package/dist/{models-ZVX3QOWE.js → models-HEKUAXXK.js} +53 -46
- package/dist/{monitor-CEKVSYTS.js → monitor-HKU57TYQ.js} +63 -60
- package/dist/{orchestrator-77BAP6BC.js → orchestrator-VDFAEFAI.js} +1831 -1057
- package/dist/{panes-7STHOAUJ.js → panes-DN2SSFOH.js} +5 -5
- package/dist/{panes-SHAUIRXY.js → panes-TALGNPZT.js} +29 -14
- package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
- package/dist/reset-EAJFFJVB.js +344 -0
- package/dist/{resources-74GKTLSF.js → resources-OVKSEFVE.js} +29 -20
- package/dist/{run-HBAUJNNZ.js → run-7DP7ZF2J.js} +120 -115
- package/dist/{share-G3APVLVP.js → share-WML67FT3.js} +32 -27
- package/dist/{skills-35HHUKCR.js → skills-SG662R2K.js} +41 -31
- package/dist/{skills-eval-QN4HSHDC.js → skills-eval-VVZEUU46.js} +78 -77
- package/dist/{skills-inventory-J357J34F.js → skills-inventory-I2E23GET.js} +23 -20
- package/dist/{slash-commands-JZZCQA32.js → slash-commands-S7MBJDQK.js} +40 -36
- package/dist/{steer-XAVHJM22.js → steer-2LQOMCPB.js} +3 -3
- package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
- package/dist/{targets-DSM6CY3M.js → targets-4QC3HIEW.js} +54 -54
- package/dist/{terminal-lease-JOPFUVEM.js → terminal-lease-TUHIJ6Y2.js} +5 -5
- package/dist/{tools-MKNWVPBH.js → tools-TFGJICCU.js} +10 -10
- package/dist/{trace-ECQ7TIYZ.js → trace-FXMXUZUF.js} +55 -7
- package/dist/uninstall-5PEVOE5B.js +408 -0
- package/dist/upgrade-M4WXY6KN.js +303 -0
- package/dist/{usage-X52N3IDJ.js → usage-N7ZNVLEM.js} +151 -104
- package/dist/{verifiers-EJTVVSMA.js → verifiers-DJTP4XX6.js} +15 -15
- package/dist/{verify-YJL6XET2.js → verify-RWE4PPEK.js} +9 -9
- package/dist/{web-fetch-MPIFL3LL.js → web-fetch-MPARV2K7.js} +2 -2
- package/dist/{wiki-generate-4NDZTQ4B.js → wiki-generate-C7IQOXSP.js} +89 -86
- package/dist/{with-panes-OBOBFIIR.js → with-panes-4GCGSL7J.js} +53 -257
- package/dist/worker/entry.js +90 -74
- package/docs/README.md +176 -81
- package/docs/{acp.md → architecture/acp.md} +36 -20
- package/docs/{alcf-provider.md → architecture/alcf-provider.md} +8 -5
- package/docs/{architecture.md → architecture/architecture.md} +43 -22
- package/docs/{artifact-placement.md → architecture/artifact-placement.md} +27 -23
- package/docs/architecture/artifact-versions.md +90 -0
- package/docs/{capacity-and-scheduling.md → architecture/capacity-and-scheduling.md} +26 -13
- package/docs/{context-engine.md → architecture/context-engine.md} +29 -25
- package/docs/{context-working-set.md → architecture/context-working-set.md} +13 -10
- package/docs/{dispatch-architecture-rationale.md → architecture/dispatch-architecture-rationale.md} +12 -9
- package/docs/{dispatch-typed-intent.md → architecture/dispatch-typed-intent.md} +68 -46
- package/docs/{evidence-and-memory.md → architecture/evidence-and-memory.md} +23 -16
- package/docs/{middleware-and-components.md → architecture/middleware-and-components.md} +11 -5
- package/docs/{model-catalog.md → architecture/model-catalog.md} +61 -27
- package/docs/{observability.md → architecture/observability.md} +38 -14
- package/docs/{pi-boundary.md → architecture/pi-boundary.md} +24 -11
- package/docs/{prompt-envelope-and-tools.md → architecture/prompt-envelope-and-tools.md} +57 -20
- package/docs/{provider-adapter-cookbook.md → architecture/provider-adapter-cookbook.md} +99 -25
- package/docs/{safety-model.md → architecture/safety-model.md} +35 -20
- package/docs/{session-lifecycle.md → architecture/session-lifecycle.md} +8 -5
- package/docs/architecture/time-conventions.md +125 -0
- package/docs/{trace-store.md → architecture/trace-store.md} +13 -5
- package/docs/{tui-design.md → architecture/tui-design.md} +13 -13
- package/docs/{worker-dispatch-mechanics.md → architecture/worker-dispatch-mechanics.md} +27 -30
- package/docs/{built-in-agents.md → guide/built-in-agents.md} +65 -35
- package/docs/{commands-and-modes.md → guide/commands-and-modes.md} +66 -61
- package/docs/{configuration-and-targets.md → guide/configuration-and-targets.md} +323 -297
- package/docs/guide/configuration-reference.md +1163 -0
- package/docs/{environment-variables.md → guide/environment-variables.md} +33 -28
- package/docs/{exit-codes-and-output.md → guide/exit-codes-and-output.md} +6 -3
- package/docs/{extensions-and-sharing.md → guide/extensions-and-sharing.md} +41 -14
- package/docs/{fleet-dispatch.md → guide/fleet-dispatch.md} +39 -43
- package/docs/{glossary.md → guide/glossary.md} +14 -11
- package/docs/{installation-and-lifecycle.md → guide/installation-and-lifecycle.md} +81 -17
- package/docs/guide/panes-and-files.md +290 -0
- package/docs/{proactive-memory.md → guide/proactive-memory.md} +131 -107
- package/docs/{resource-library.md → guide/resource-library.md} +13 -4
- package/docs/{skills-marketplace.md → guide/skills-marketplace.md} +25 -3
- package/docs/{tool-usage.md → guide/tool-usage.md} +87 -23
- package/docs/{troubleshooting.md → guide/troubleshooting.md} +9 -4
- package/docs/{config-knobs-audit.md → history/config-knobs-audit.md} +11 -11
- package/docs/{release-cut-checklist.md → history/release-cut-checklist.md} +29 -2
- package/docs/process/development-pipeline.md +152 -0
- package/docs/process/documentation-coverage.md +100 -0
- package/docs/process/documentation-guide.md +187 -0
- package/docs/{eval-runner.md → process/eval-runner.md} +108 -53
- package/docs/{evals-internal.md → process/evals-internal.md} +10 -10
- package/docs/{evolution.md → process/evolution.md} +2 -2
- package/docs/{fleet-demo-runbook.md → process/fleet-demo-runbook.md} +11 -7
- package/docs/{git-commit-provenance.md → process/git-commit-provenance.md} +11 -4
- package/docs/{performance-methodology.md → process/performance-methodology.md} +87 -69
- package/docs/{scientific-validation.md → process/scientific-validation.md} +4 -4
- package/evals/README.md +2 -2
- package/evals/behavioral-model.yaml +3 -2
- package/package.json +10 -8
- package/skills/README.md +52 -41
- package/skills/coding/ast-grep/SKILL.md +102 -31
- package/skills/coding/ast-grep/evals.md +26 -0
- package/skills/coding/coding-standards/SKILL.md +41 -6
- package/skills/coding/coding-standards/evals.md +23 -0
- package/skills/coding/prototype/SKILL.md +88 -29
- package/skills/coding/prototype/evals.md +19 -0
- package/skills/coding/tdd/SKILL.md +81 -54
- package/skills/coding/tdd/evals.md +20 -0
- package/skills/context/context-handoff/SKILL.md +44 -3
- package/skills/context/context-handoff/evals.md +44 -0
- package/skills/context/context-prime/SKILL.md +46 -16
- package/skills/context/context-prime/evals.md +45 -0
- package/skills/git/branch-closeout/SKILL.md +132 -0
- package/skills/git/branch-closeout/evals.md +133 -0
- package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
- package/skills/git/file-ticket/SKILL.md +78 -64
- package/skills/git/file-ticket/assets/issue-template.md +22 -0
- package/skills/git/file-ticket/evals.md +31 -26
- package/skills/git/file-ticket/references/issue-discovery.md +49 -0
- package/skills/git/fix-issue/SKILL.md +88 -65
- package/skills/git/fix-issue/evals.md +35 -31
- package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +101 -52
- package/skills/git/resolve-merge-conflicts/evals.md +52 -25
- package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
- package/skills/git/ship/SKILL.md +103 -67
- package/skills/git/ship/assets/pr-template.md +21 -0
- package/skills/git/ship/evals.md +44 -28
- package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
- package/skills/git/worktree-create/SKILL.md +80 -50
- package/skills/git/worktree-create/evals.md +40 -33
- package/skills/git/worktree-create/references/worktree-setup.md +62 -66
- package/skills/git/worktree-merge/SKILL.md +112 -65
- package/skills/git/worktree-merge/evals.md +42 -34
- package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
- package/skills/meta/clio-coder-dev/SKILL.md +9 -5
- package/skills/meta/clio-coder-dev/evals.md +3 -2
- package/skills/meta/clio-coder-test/SKILL.md +102 -95
- package/skills/meta/clio-coder-test/evals.md +9 -4
- package/skills/meta/clio-coder-test/references/harness.md +100 -124
- package/skills/meta/clio-coder-test/references/test-map.md +77 -50
- package/skills/meta/credentials/SKILL.md +2 -2
- package/skills/meta/find-skills/SKILL.md +2 -2
- package/skills/meta/herdr/SKILL.md +2 -2
- package/skills/meta/skill-craft/SKILL.md +22 -16
- package/skills/planning/archify/SKILL.md +196 -0
- package/skills/planning/archify/evals.md +65 -0
- package/skills/planning/architecture/SKILL.md +62 -13
- package/skills/planning/architecture/evals.md +65 -0
- package/skills/planning/backlog/SKILL.md +131 -15
- package/skills/planning/backlog/evals.md +142 -0
- package/skills/planning/prd/SKILL.md +47 -7
- package/skills/planning/prd/evals.md +54 -0
- package/skills/planning/product-intent/SKILL.md +58 -3
- package/skills/planning/product-intent/evals.md +70 -0
- package/skills/planning/tech-spec/SKILL.md +54 -3
- package/skills/planning/tech-spec/evals.md +73 -0
- package/skills/registry.yaml +70 -62
- package/skills/remote.yaml +13 -0
- package/skills/research/arxiv-literature/SKILL.md +77 -19
- package/skills/research/arxiv-literature/evals.md +50 -0
- package/skills/research/experiment-protocol/SKILL.md +21 -2
- package/skills/research/experiment-protocol/evals.md +23 -0
- package/skills/research/scientific-debugging/SKILL.md +24 -2
- package/skills/research/scientific-debugging/evals.md +18 -0
- package/skills/research/scientific-modernization/SKILL.md +27 -2
- package/skills/research/scientific-modernization/evals.md +27 -0
- package/skills/skill-marketplace.json +97 -62
- package/skills/workflow/cut-it/SKILL.md +66 -6
- package/skills/workflow/cut-it/evals.md +101 -0
- package/skills/workflow/design-council/SKILL.md +118 -28
- package/skills/workflow/design-council/evals.md +161 -0
- package/skills/workflow/grill-me/SKILL.md +87 -11
- package/skills/workflow/grill-me/evals.md +153 -0
- package/skills/workflow/workflow-distiller/SKILL.md +77 -18
- package/skills/workflow/workflow-distiller/evals.md +118 -0
- package/src/cli/args.ts +2 -2
- package/src/cli/bootstrap-generate.ts +1 -1
- package/src/cli/config-inspect.ts +65 -12
- package/src/cli/configure-interop.ts +105 -13
- package/src/cli/configure-oauth.ts +57 -0
- package/src/cli/configure-onboarding.ts +980 -0
- package/src/cli/configure-target.ts +594 -0
- package/src/cli/configure.ts +1082 -532
- package/src/cli/context-map.ts +114 -0
- package/src/cli/context.ts +4 -0
- package/src/cli/docs.ts +22 -14
- package/src/cli/doctor-naming.ts +5 -5
- package/src/cli/doctor-toolchain.ts +3 -3
- package/src/cli/eval.ts +1 -2
- package/src/cli/extensions.ts +2 -1
- package/src/cli/fleet.ts +1 -1
- package/src/cli/index.ts +3 -1
- package/src/cli/internal-dispatch.ts +3 -4
- package/src/cli/lifecycle-presenter.ts +436 -0
- package/src/cli/models.ts +10 -2
- package/src/cli/modes/print.ts +5 -1
- package/src/cli/panes.ts +19 -5
- package/src/cli/reset.ts +228 -106
- package/src/cli/run.ts +9 -4
- package/src/cli/select.ts +664 -0
- package/src/cli/share.ts +5 -1
- package/src/cli/skills-eval.ts +3 -3
- package/src/cli/skills.ts +9 -2
- package/src/cli/targets.ts +5 -6
- package/src/cli/trace.ts +55 -4
- package/src/cli/uninstall.ts +233 -165
- package/src/cli/upgrade.ts +204 -149
- package/src/cli/usage.ts +86 -27
- package/src/cli/validate-model.ts +3 -3
- package/src/cli/wiki-generate.ts +1 -1
- package/src/core/artifact-paths.ts +1 -1
- package/src/core/bash-exec.ts +131 -86
- package/src/core/bus-events.ts +51 -6
- package/src/core/config.ts +61 -1
- package/src/core/defaults.ts +7 -4
- package/src/core/dispatch-outcome.ts +16 -0
- package/src/core/external-diagnostic.ts +44 -0
- package/src/core/gateway-routing.ts +157 -0
- package/src/core/guardrails.ts +10 -49
- package/src/core/prompt-hint.ts +9 -0
- package/src/core/safe-exec.ts +17 -2
- package/src/core/skill-activation.ts +89 -2
- package/src/domains/agents/builtins/architect.md +2 -3
- package/src/domains/agents/builtins/coder.md +3 -2
- package/src/domains/agents/builtins/debugger.md +2 -2
- package/src/domains/agents/builtins/documenter.md +2 -2
- package/src/domains/agents/builtins/git-master.md +1 -1
- package/src/domains/agents/builtins/oracle.md +1 -1
- package/src/domains/agents/builtins/provenance.md +1 -1
- package/src/domains/agents/builtins/researcher.md +1 -1
- package/src/domains/agents/builtins/scout.md +1 -1
- package/src/domains/agents/builtins/tester.md +2 -2
- package/src/domains/agents/builtins/verifier.md +2 -2
- package/src/domains/agents/builtins/wiki-writer.md +1 -1
- package/src/domains/agents/builtins/world-knowledge.md +31 -0
- package/src/domains/agents/catalog.ts +13 -15
- package/src/domains/agents/contract.ts +2 -0
- package/src/domains/agents/extension.ts +23 -1
- package/src/domains/agents/result-contract.ts +70 -0
- package/src/domains/config/keybindings.ts +8 -0
- package/src/domains/context/extension.ts +0 -3
- package/src/domains/context/wiki/map-seed.ts +589 -0
- package/src/domains/context/wiki/plan.ts +2 -2
- package/src/domains/context/working-set/path-index.ts +1 -0
- package/src/domains/dispatch/admission.ts +29 -0
- package/src/domains/dispatch/agent-candidates.ts +10 -0
- package/src/domains/dispatch/budget-envelope.ts +86 -1
- package/src/domains/dispatch/capability-match.ts +11 -0
- package/src/domains/dispatch/capacity-lease.ts +17 -0
- package/src/domains/dispatch/contract.ts +11 -1
- package/src/domains/dispatch/extension.ts +237 -49
- package/src/domains/dispatch/host-verification.ts +435 -39
- package/src/domains/dispatch/intent-requirements.ts +10 -0
- package/src/domains/dispatch/intent.ts +18 -1
- package/src/domains/dispatch/path-scope.ts +235 -24
- package/src/domains/dispatch/run-event-journal.ts +4 -15
- package/src/domains/dispatch/state.ts +2 -3
- package/src/domains/dispatch/transport.ts +45 -21
- package/src/domains/dispatch/types.ts +58 -3
- package/src/domains/dispatch/worker-model-metadata.ts +38 -0
- package/src/domains/eval/artifacts/store.ts +5 -0
- package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
- package/src/domains/eval/metrics/token-stream.ts +201 -31
- package/src/domains/eval/metrics/tracked.ts +40 -4
- package/src/domains/eval/runners/clio-run.ts +5 -2
- package/src/domains/eval/schema/suite.ts +28 -0
- package/src/domains/eval/schema/verdict.ts +2 -2
- package/src/domains/eval/store.ts +8 -1
- package/src/domains/eval/suites/resolve.ts +13 -1
- package/src/domains/eval/suites/run.ts +24 -3
- package/src/domains/evidence/trust-status.ts +10 -1
- package/src/domains/extensions/contract.ts +15 -1
- package/src/domains/extensions/discovery.ts +238 -41
- package/src/domains/extensions/extension.ts +105 -6
- package/src/domains/extensions/index.ts +24 -0
- package/src/domains/extensions/integrity.ts +189 -0
- package/src/domains/extensions/manager.ts +17 -1
- package/src/domains/extensions/resource-path.ts +27 -0
- package/src/domains/extensions/resources.ts +18 -38
- package/src/domains/extensions/snapshot-store.ts +39 -0
- package/src/domains/extensions/snapshot.ts +180 -0
- package/src/domains/extensions/state.ts +385 -57
- package/src/domains/extensions/types.ts +118 -1
- package/src/domains/interop/registry.ts +6 -2
- package/src/domains/interop/types.ts +4 -0
- package/src/domains/lifecycle/migrations/2026-09-01-extension-install-digests.ts +27 -0
- package/src/domains/lifecycle/migrations/index.ts +6 -0
- package/src/domains/lifecycle/naming-resources.ts +19 -4
- package/src/domains/lifecycle/naming-yazi.ts +10 -5
- package/src/domains/memory/task-memory-policy.ts +70 -26
- package/src/domains/memory/task-memory-telemetry.ts +1 -0
- package/src/domains/middleware/contract.ts +26 -0
- package/src/domains/middleware/extension.ts +24 -24
- package/src/domains/middleware/hook-receipts.ts +27 -4
- package/src/domains/middleware/hooks-io.ts +65 -32
- package/src/domains/middleware/hooks.ts +64 -0
- package/src/domains/middleware/index.ts +28 -5
- package/src/domains/middleware/marketplace-offer.ts +3 -35
- package/src/domains/middleware/memory-intervention.ts +127 -32
- package/src/domains/middleware/memory-step-endpoint.ts +3 -2
- package/src/domains/middleware/registrations.ts +326 -0
- package/src/domains/middleware/runtime.ts +28 -0
- package/src/domains/middleware/skills-reminder.ts +31 -2
- package/src/domains/middleware/snapshot.ts +20 -7
- package/src/domains/mux/contract.ts +38 -0
- package/src/domains/mux/detect.ts +6 -13
- package/src/domains/mux/index.ts +1 -1
- package/src/domains/mux/operations.ts +44 -5
- package/src/domains/mux/yazi/assets/yazi.toml +2 -2
- package/src/domains/mux/yazi/session.ts +53 -4
- package/src/domains/mux/yazi/theme.ts +117 -17
- package/src/domains/observability/compaction-usage.ts +118 -0
- package/src/domains/observability/contract.ts +10 -11
- package/src/domains/observability/cost.ts +1 -1
- package/src/domains/observability/extension.ts +17 -4
- package/src/domains/observability/out-of-turn-usage.ts +52 -21
- package/src/domains/observability/projection.ts +14 -90
- package/src/domains/observability/trace-store.ts +43 -7
- package/src/domains/prompts/compiler.ts +73 -53
- package/src/domains/prompts/contract.ts +15 -3
- package/src/domains/prompts/extension.ts +97 -9
- package/src/domains/prompts/fragments/identity/clio-worker.md +1 -3
- package/src/domains/prompts/fragments/identity/clio.md +6 -12
- package/src/domains/prompts/fragments/identity/docs-routing.md +1 -2
- package/src/domains/prompts/fragments/identity/self-awareness.md +3 -11
- package/src/domains/prompts/fragments/operating/contract.md +7 -15
- package/src/domains/prompts/fragments/operating/delegation.md +32 -34
- package/src/domains/prompts/fragments/operating/skills.md +10 -24
- package/src/domains/prompts/fragments/operating/worker.md +1 -8
- package/src/domains/providers/contract.ts +4 -1
- package/src/domains/providers/extension.ts +40 -9
- package/src/domains/providers/index.ts +1 -1
- package/src/domains/providers/model-capabilities.ts +9 -0
- package/src/domains/providers/model-discovery.ts +2 -0
- package/src/domains/providers/model-runtime-capabilities.ts +99 -25
- package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +699 -114
- package/src/domains/providers/runtime-resolution.ts +31 -0
- package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
- package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
- package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
- package/src/domains/providers/runtimes/common/probe-helpers.ts +7 -2
- package/src/domains/providers/runtimes/local-native/llamacpp.ts +9 -1
- package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
- package/src/domains/providers/support.ts +11 -5
- package/src/domains/providers/target-model-cache.ts +25 -2
- package/src/domains/providers/types/capability-flags.ts +2 -0
- package/src/domains/providers/types/cost-provenance.ts +19 -0
- package/src/domains/providers/types/local-model-quirks.ts +85 -37
- package/src/domains/providers/types/runtime-descriptor.ts +20 -1
- package/src/domains/providers/types/target-descriptor.ts +19 -0
- package/src/domains/resources/index.ts +3 -0
- package/src/domains/resources/skills/install.ts +72 -7
- package/src/domains/resources/skills/loader.ts +23 -19
- package/src/domains/resources/skills/marketplace.ts +63 -11
- package/src/domains/safety/autonomy.ts +15 -0
- package/src/domains/safety/call-target.ts +1 -1
- package/src/domains/safety/index.ts +1 -0
- package/src/domains/safety/loop-detector.ts +7 -4
- package/src/domains/safety/path-policy.ts +1 -1
- package/src/domains/safety/policy-engine.ts +34 -11
- package/src/domains/safety/protected-artifacts.ts +191 -88
- package/src/domains/safety/run-effects.ts +2 -22
- package/src/domains/safety/skill-authority.ts +55 -0
- package/src/domains/session/compaction/compact.ts +72 -22
- package/src/domains/session/entries.ts +6 -0
- package/src/domains/session/task-board.ts +10 -9
- package/src/domains/session/usage.ts +3 -3
- package/src/domains/share/archive.ts +164 -7
- package/src/engine/acp/server.ts +62 -9
- package/src/engine/agent.ts +13 -3
- package/src/engine/ai.ts +26 -8
- package/src/engine/antigravity/subprocess-runtime.ts +386 -120
- package/src/engine/api-registry.ts +3 -0
- package/src/engine/apis/llamacpp-residency.ts +3 -4
- package/src/engine/apis/lmstudio.ts +3 -3
- package/src/engine/apis/ollama-native.ts +6 -6
- package/src/engine/apis/openai-completions.ts +145 -39
- package/src/engine/apis/output-budget.ts +8 -18
- package/src/engine/apis/residency.ts +8 -27
- package/src/engine/external-subprocess.ts +114 -6
- package/src/engine/gemma-channel-filter.ts +19 -0
- package/src/engine/loop-guard.ts +92 -12
- package/src/engine/worker-runtime.ts +40 -11
- package/src/engine/worker-tools.ts +3 -1
- package/src/entry/background-model-metadata.ts +18 -0
- package/src/entry/compaction-prompt.ts +57 -0
- package/src/entry/extension-hook-sources.ts +28 -0
- package/src/entry/extension-reload.ts +309 -0
- package/src/entry/orchestrator.ts +464 -251
- package/src/entry/task-memory-lifecycle.ts +35 -0
- package/src/interactive/application-controller.ts +2 -1
- package/src/interactive/bus-notices.ts +8 -1
- package/src/interactive/chat-loop-messages.ts +16 -17
- package/src/interactive/chat-loop.ts +75 -3
- package/src/interactive/chat-panel.ts +36 -13
- package/src/interactive/chat-renderer.ts +72 -7
- package/src/interactive/cost-overlay.ts +26 -2
- package/src/interactive/dispatch-board.ts +6 -11
- package/src/interactive/footer/widgets.ts +13 -0
- package/src/interactive/interactive-application.ts +39 -4
- package/src/interactive/interactive-input-runtime.ts +4 -0
- package/src/interactive/interactive-presentation.ts +2 -2
- package/src/interactive/interactive-slash-runtime.ts +4 -1
- package/src/interactive/overlays/extensions.ts +9 -1
- package/src/interactive/overlays/help-reference.ts +13 -0
- package/src/interactive/overlays/settings.ts +27 -16
- package/src/interactive/panes-runtime.ts +111 -35
- package/src/interactive/prompt-cache-identity.ts +88 -0
- package/src/interactive/renderers/worker-entry.ts +32 -0
- package/src/interactive/slash-commands.ts +153 -20
- package/src/interactive/stream-pacing-policy.ts +0 -23
- package/src/interactive/theme/labels.ts +19 -13
- package/src/interactive/turn-context.ts +39 -20
- package/src/interactive/turn-recovery.ts +8 -0
- package/src/interactive/turn-runtime.ts +27 -11
- package/src/interactive/turn-state.ts +7 -0
- package/src/interactive/worker-receipts.ts +1 -0
- package/src/interactive/worker-stream.ts +6 -1
- package/src/interactive/yazi-bridge.ts +60 -6
- package/src/tools/agent-tools.ts +30 -1
- package/src/tools/artifact.ts +2 -2
- package/src/tools/ask-user.ts +3 -3
- package/src/tools/bash.ts +1 -1
- package/src/tools/bootstrap.ts +4 -0
- package/src/tools/builtin-tool-catalog.ts +52 -22
- package/src/tools/codewiki/code-nav-surface.ts +6 -0
- package/src/tools/codewiki/code-nav.ts +99 -13
- package/src/tools/context/docs-engine.ts +20 -7
- package/src/tools/context/index.ts +59 -21
- package/src/tools/core-bootstrap.ts +28 -6
- package/src/tools/credential-present.ts +1 -2
- package/src/tools/dispatch-arguments.ts +6 -1
- package/src/tools/dispatch-event-text.ts +10 -0
- package/src/tools/dispatch-plan.ts +49 -4
- package/src/tools/dispatch-run-events.ts +1 -1
- package/src/tools/dispatch-runner.ts +12 -0
- package/src/tools/dispatch-schema.ts +338 -0
- package/src/tools/dispatch-types.ts +3 -0
- package/src/tools/dispatch.ts +9 -254
- package/src/tools/ledger.ts +3 -5
- package/src/tools/monitor-surface.ts +5 -13
- package/src/tools/observation.ts +4 -5
- package/src/tools/panes-surface.ts +4 -11
- package/src/tools/panes.ts +4 -2
- package/src/tools/policy.ts +15 -2
- package/src/tools/read.ts +5 -6
- package/src/tools/registry.ts +41 -12
- package/src/tools/result-shaping.ts +18 -14
- package/src/tools/steer-surface.ts +1 -1
- package/src/tools/tasks.ts +1 -1
- package/src/tools/truncate.ts +6 -5
- package/src/tools/verify/surface.ts +6 -12
- package/src/tools/web-fetch-surface.ts +1 -3
- package/src/tools/worker-evidence.ts +3 -1
- package/src/worker/spec-contract.ts +4 -0
- package/dist/builtins-UJLMOVOV.js +0 -17
- package/dist/chunk-5QIAJV2D.js +0 -48
- package/dist/chunk-JZWT5J3Y.js +0 -814
- package/dist/chunk-K7VKOLQQ.js +0 -15
- package/dist/chunk-PMZCIOCJ.js +0 -25
- package/dist/chunk-SUW5DORT.js +0 -819
- package/dist/chunk-UOV2BYIW.js +0 -107
- package/dist/chunk-WR6U3OVP.js +0 -45
- package/dist/chunk-Y45G3AXC.js +0 -1558
- package/dist/reset-EOLM7GVE.js +0 -230
- package/dist/uninstall-N34PCTGJ.js +0 -331
- package/dist/upgrade-H7TOM7YL.js +0 -323
- package/docs/artifact-versions.md +0 -67
- package/docs/development-pipeline.md +0 -121
- package/docs/documentation-coverage.md +0 -46
- package/docs/documentation-guide.md +0 -167
- package/docs/time-conventions.md +0 -101
|
@@ -1,11 +1,11 @@
|
|
|
1
1
|
# Clio Coder Local Evaluation Runner
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Clio Coder Local Evaluation Runner visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/eval_blueprint.html).
|
|
5
5
|
|
|
6
6
|
The local evaluation runner executes repository-local YAML task suites as deterministic subprocess checks. It is useful for comparing harness changes, prompts, tools, or local workflows.
|
|
7
7
|
|
|
8
|
-
Source of truth: [src/domains/eval/](
|
|
8
|
+
Source of truth: [src/domains/eval/](../../src/domains/eval/) and [src/cli/eval.ts](../../src/cli/eval.ts).
|
|
9
9
|
|
|
10
10
|
---
|
|
11
11
|
|
|
@@ -20,6 +20,7 @@ clio-coder eval run --task-file <tasks.yaml> [--repeat <n>] [--out <path>] [--cl
|
|
|
20
20
|
clio-coder eval report <evalId> --format text|json|md|swe-jsonl|junit
|
|
21
21
|
clio-coder eval compare <baselineEvalId> <candidateEvalId> [--metric <name>] [--format text|json|md|junit] [--allow-config-drift]
|
|
22
22
|
clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds <file>]
|
|
23
|
+
clio-coder eval inventory --json
|
|
23
24
|
```
|
|
24
25
|
|
|
25
26
|
### Command Roles
|
|
@@ -33,6 +34,7 @@ clio-coder eval gate <candidateEvalId> --baseline <baselineEvalId> [--thresholds
|
|
|
33
34
|
* `junit`: XML report for CI/CD integration.
|
|
34
35
|
* **`compare`**: Compares two evaluation artifacts (baseline and candidate) by matching tasks.
|
|
35
36
|
* **`gate`**: Compares candidate metrics against baseline and absolute thresholds. Correctness and safety regressions fail independently of informational budgets.
|
|
37
|
+
* **`inventory`**: Prints the fixed machine-readable inventory used by GUI hosts. It includes stored report identity, provenance, serving facts, accounting, and per-scenario outcomes without report attachments.
|
|
36
38
|
|
|
37
39
|
Exit codes:
|
|
38
40
|
|
|
@@ -42,7 +44,7 @@ Exit codes:
|
|
|
42
44
|
| `eval run` | `0` when all task repetitions pass | `1` when any task fails, `2` for invalid configs |
|
|
43
45
|
| `eval report` | `0` when artifact loads | `1` if artifact cannot be read, `2` for invalid ID |
|
|
44
46
|
| `eval compare` | `0` when both artifacts compare and the behavioral hard gate passes | `1` for a hard regression or unreadable artifact, `2` for invalid ID |
|
|
45
|
-
| `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for
|
|
47
|
+
| `eval gate` | `0` when correctness, safety, and hard threshold assertions pass | `1` for hard failures, unreadable inputs, and malformed threshold files; `2` for an invalid eval ID or usage error |
|
|
46
48
|
|
|
47
49
|
---
|
|
48
50
|
|
|
@@ -75,7 +77,7 @@ tasks:
|
|
|
75
77
|
excludes:
|
|
76
78
|
- "**/node_modules/**"
|
|
77
79
|
runner:
|
|
78
|
-
kind: "clio-run" # clio-run | context-index | context-init | external-command
|
|
80
|
+
kind: "clio-coder-run" # clio-coder-run | context-index | context-init | external-command
|
|
79
81
|
prompt: "Optimize the FFT tolerance bounds in solver.ts"
|
|
80
82
|
timeoutMs: 60000
|
|
81
83
|
verify:
|
|
@@ -103,12 +105,12 @@ tasks:
|
|
|
103
105
|
| --- | --- | --- |
|
|
104
106
|
| `version` | - | Must equal `2`. |
|
|
105
107
|
| `suite` | `id`, `title`, `visibility`, `description` | Metadata identifying the evaluation suite. |
|
|
106
|
-
| `matrix` | `targets[]`, `repeats`, `dimensions[]` | Matrix of execution targets, repetition count,
|
|
107
|
-
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). |
|
|
108
|
-
| `runner` | `kind`, `prompt`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-run` (starts Clio agent loop), `context-index` (runs indexer), `context-init` (initializes context), `external-command` (spawns subprocess). |
|
|
109
|
-
| `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
|
|
108
|
+
| `matrix` | `targets[]`, `repeats`, `dimensions[]`, `maxCostUsd` | Matrix of execution targets, repetition count, execution-envelope fields intentionally varied by the suite, and an optional cumulative known-cost ceiling. |
|
|
109
|
+
| `workspace` | `kind`, `path`, `url`, `commit`, `checkout`, `excludes`, `setup` | Workspace strategy: `local` (run in-place), `git` (clone from URL), or `temp-copy` (isolated copy of a directory). Optional `setup` commands prepare the workspace before the runner starts. |
|
|
110
|
+
| `runner` | `kind`, `prompt`, `autonomy`, `agent`, `command`, `commands`, `args`, `timeoutMs` | Runner type: `clio-coder-run` (starts Clio's agent loop), `context-index` (runs the indexer), `context-init` (initializes context), or `external-command` (spawns a subprocess). `agent` selects a worker recipe and `autonomy` sets one-run headless authority. |
|
|
111
|
+
| `behavioral` | `schema`, `corpus`, `execution`, `expectedBehavior`, `forbiddenBehavior`, `judge` | Optional `clio-coder.eval.scenario.v1` behavioral contract. Rules name a closed category and a typed predicate over transcript, tool, receipt, or grader facts. |
|
|
110
112
|
| `verify` | `commands`, `measure`, `assertions`, `forbidPaths` | Validation steps: shell commands, a task-outcome grader, metric assertions (e.g. `op: lt` for max token counts), and files/directories that must not be created or modified (`forbidPaths`). |
|
|
111
|
-
| `metrics` | `collect` |
|
|
113
|
+
| `metrics` | `collect`, `readObservation` | Metric names to compile plus optional public allowlisted and decoy paths reduced to bounded read counters. Raw path strings do not enter behavioral facts. |
|
|
112
114
|
|
|
113
115
|
---
|
|
114
116
|
|
|
@@ -120,7 +122,7 @@ tasks:
|
|
|
120
122
|
---
|
|
121
123
|
|
|
122
124
|
## Runner Kinds
|
|
123
|
-
* **`clio-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools.
|
|
125
|
+
* **`clio-coder-run`**: Invokes the main Clio Coder agent loop with the task's prompt, tracing all tools. The released `clio-run` spelling is accepted only as a legacy input alias and is normalized before validation; writers and new suites use `clio-coder-run`.
|
|
124
126
|
* **`context-index`**: Triggers the context engine to build index structures (`codewiki`).
|
|
125
127
|
* **`context-init`**: Initializes workspace files (such as generating `CLIO-CODER.md`).
|
|
126
128
|
* **`external-command`**: Spawns an external command or sequence of commands in the task workspace.
|
|
@@ -137,7 +139,37 @@ Metrics collected during runs can be validated automatically using the `verify.a
|
|
|
137
139
|
* `eq` (equal)
|
|
138
140
|
* `neq` (not equal)
|
|
139
141
|
|
|
140
|
-
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `
|
|
142
|
+
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, `result.pass`, and the `provider.*` metrics below. Each `verify.assertions` condition must hold; an unavailable metric fails closed.
|
|
143
|
+
|
|
144
|
+
### Optional Provider-Health Gates
|
|
145
|
+
|
|
146
|
+
A task that recovers from a provider error can still pass its task checks by default. Provider health is a separate, opt-in requirement. To require observed provider events and no observed terminal errors, add these assertions to the task:
|
|
147
|
+
|
|
148
|
+
```yaml
|
|
149
|
+
verify:
|
|
150
|
+
assertions:
|
|
151
|
+
- metric: "provider.measured"
|
|
152
|
+
op: "eq"
|
|
153
|
+
value: true
|
|
154
|
+
- metric: "provider.stopReason.error"
|
|
155
|
+
op: "eq"
|
|
156
|
+
value: 0
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
Suite-level `thresholds.fail` uses the opposite condition: a matching condition is a failure. The corresponding hard gate checks each run as follows:
|
|
160
|
+
|
|
161
|
+
```yaml
|
|
162
|
+
thresholds:
|
|
163
|
+
fail:
|
|
164
|
+
- metric: "provider.measured"
|
|
165
|
+
op: "eq"
|
|
166
|
+
value: false
|
|
167
|
+
- metric: "provider.stopReason.error"
|
|
168
|
+
op: "gt"
|
|
169
|
+
value: 0
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
Unavailable metrics also fail closed in a hard threshold. `thresholds.informational` records findings without changing exit status. These examples reject an observed error followed by a successful recovery while leaving the default ungated task-pass policy unchanged. To reject any observed retry start or assistant abort as well, add conditions on `provider.retryStarted` or `provider.stopReason.aborted` with the same assertion-versus-failure polarity.
|
|
141
173
|
|
|
142
174
|
---
|
|
143
175
|
|
|
@@ -173,14 +205,44 @@ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
|
|
|
173
205
|
|
|
174
206
|
Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
|
|
175
207
|
|
|
176
|
-
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events
|
|
208
|
+
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events. These totals include all known usage on errored calls as well as successful calls; recovery never subtracts earlier spend. Only finite, nonnegative usage facts are admitted. On surfaces without the relevant stdout events (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
|
|
177
209
|
2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
|
|
178
210
|
|
|
179
211
|
### Fail-Closed Reporting
|
|
180
212
|
Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
|
|
181
213
|
|
|
214
|
+
An errored call must carry at least one positive reported token, reasoning, or cost fact before its usage is considered observed. Reasoning-only or cost-only observations do not establish ordinary token totals. A stream containing only errored calls with missing or synthetic all-zero usage remains `tokens.measured: false`. The current event shape cannot distinguish synthetic all-zero failures from genuinely reported zero usage, so it cannot establish measured zero spending in either case. Partial positive usage remains included as known spend. On failed calls, adapters can also initialize individual absent fields to zero; those zeros remain unattributed and make coverage incomplete. The existing inclusive numeric fields are known subtotals, so their zeros do not prove complete zero spending when failed usage is incomplete.
|
|
215
|
+
|
|
216
|
+
### Provider Observations and Failed-Call Share
|
|
217
|
+
|
|
218
|
+
The `provider.*` metrics describe events observed on live stdout, folded before diagnostic output is truncated. Native runs and multi-command external runners retain these observations from their executed commands. They do not enumerate SDK-internal retries or network attempts that were never emitted. Filtered or opaque output can leave provider health unobserved even when the process exits successfully or a receipt reports task success.
|
|
219
|
+
|
|
220
|
+
| Metric | Meaning |
|
|
221
|
+
| --- | --- |
|
|
222
|
+
| `provider.measured` | Whether an assistant terminal reason or a counted retry phase was observed. With no such observations this is `false`, and provider counters are absent. It does not certify complete provider coverage. |
|
|
223
|
+
| `provider.stopReason.stop`, `provider.stopReason.toolUse`, `provider.stopReason.length` | Counts of these terminal reasons on assistant `message_end` events. |
|
|
224
|
+
| `provider.stopReason.error`, `provider.stopReason.aborted`, `provider.stopReason.other` | Separate counts of errored, aborted, and other observed terminal reasons. An unrecognized terminal reason goes into `other`. Partial updates and repeated messages in `turn_end` or `agent_end` do not add counts. |
|
|
225
|
+
| `provider.retryScheduled` | Observed `scheduled` phases: planned retries, including ones cancelled before execution. |
|
|
226
|
+
| `provider.retryStarted` | Observed `retrying` phases: retry execution starts. Repeated `waiting` countdown frames do not count as attempts. |
|
|
227
|
+
| `provider.retryCancelled`, `provider.retryExhausted`, `provider.retryRecovered` | Counts of the corresponding observed phases. They describe retry-chain outcomes and do not fabricate additional assistant calls. Attempt numbers can restart for each chain. |
|
|
228
|
+
| `provider.errorUsageObservedCalls` | Errored calls with at least one positive reported token, reasoning, or cost fact. |
|
|
229
|
+
| `provider.errorUsageUnobservedCalls` | Errored calls with no positive reported token, reasoning, or cost fact, including absent or all-zero usage. |
|
|
230
|
+
| `provider.errorUsageIncompleteCalls` | Errored calls with unobserved usage, incomplete token fields, or ambiguous normalized zero fields. This can overlap `errorUsageObservedCalls` when only part of the usage is known. |
|
|
231
|
+
| `provider.errorCostUnobservedCalls` | Errored calls without a positive cost fact. Zero or absent cost does not prove that an error was free, including when some token usage is known. |
|
|
232
|
+
| `provider.errorTokens.input`, `provider.errorTokens.output`, `provider.errorTokens.total`, `provider.errorTokens.cacheRead`, `provider.errorTokens.cacheWrite` | Known positive token subtotals for errored calls. Absent or ambiguous zero fields remain unattributed; when total usage is absent, a total can still be summed from known token fields and remains incomplete. |
|
|
233
|
+
| `provider.errorCostUsd` | Known positive cost subtotal from errored calls' stream usage objects. Cost can come from adapter pricing; it is not independently certified provider billing. |
|
|
234
|
+
| `provider.errorReasoningTokens`, `provider.errorReasoningUnobservedCalls` | Known failed-call reasoning subtotal and calls without attributable reported reasoning. Reasoning can overlap output, so it is never added to ordinary token totals. |
|
|
235
|
+
|
|
236
|
+
The failed-call share covers `stopReason: error`; aborted calls remain separately labeled. Failed-share amounts and usage-coverage counters appear only after an errored terminal message is observed. These share metrics supplement the inclusive `tokens.*` totals without changing the summary token shape, receipt schema, or verdict schema. Missing or partial failed usage makes them known subtotals, not a complete amount to subtract from total spend. The native runner's `cost.usd` can use receipt evidence, so equality with the stream-based `provider.errorCostUsd` is not guaranteed.
|
|
237
|
+
|
|
238
|
+
Positive reported reasoning uses the normalized `usage.reasoning` field, `reasoning_tokens`, and supported nested provider-detail fields. The root `reasoningTokens` alias can be an adapter estimate without a provenance marker, so it is left unattributed rather than promoted to reported provider usage. Normalized zero reasoning is also unattributed because adapters can fill it when provider detail is absent. A reasoning-only failure is observed but has incomplete ordinary-token coverage; no output or total is inferred from it.
|
|
239
|
+
|
|
240
|
+
These observations also do not reconcile the separate `trackedMetrics` ledger selection (#276). Tracked metrics prefer durable assistant-call facts when available, retain durable compaction and tool records, and otherwise fall back to stream calls. Artifacts expose source counts and warnings, but partial or mixed ledgers can omit stream-only calls and fork-inherited history remains unreconciled. Neither those tracked values nor the new provider counters prove complete run accounting; failed-compaction usage is retained separately in the out-of-turn usage ledger and usage report, and is not included by this eval fold.
|
|
241
|
+
|
|
182
242
|
---
|
|
183
243
|
|
|
244
|
+
Full reconciliation across session, stdout, fork and out-of-turn evidence is deferred to v0.4.5 or later. Version 0.4.3 does not add an automatic rejection of cost or efficiency comparisons merely because those sources are partial or mixed. Matching source counts do not prove complete coverage or shared call identity. Existing missing-metric, serving-configuration and execution-envelope comparison gates still apply.
|
|
245
|
+
|
|
184
246
|
## Eval Artifact Format (v4)
|
|
185
247
|
|
|
186
248
|
Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
|
|
@@ -190,7 +252,7 @@ export interface EvalArtifactV4 {
|
|
|
190
252
|
version: 4;
|
|
191
253
|
evalId: string;
|
|
192
254
|
suite: { id: string; hash: string };
|
|
193
|
-
|
|
255
|
+
clioCoder: EvalClioProvenance;
|
|
194
256
|
environment: EvalEnvironmentProvenance;
|
|
195
257
|
matrix: { target: string; model: string | null; thinking: string | null };
|
|
196
258
|
summary: EvalArtifactSummaryV4;
|
|
@@ -206,11 +268,11 @@ export interface EvalArtifactV4 {
|
|
|
206
268
|
|
|
207
269
|
## The verdict envelope
|
|
208
270
|
|
|
209
|
-
Every result carries a strictly parsed `clio.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
|
|
271
|
+
Every result carries a strictly parsed `clio-coder.eval.verdict.v1` envelope (`src/domains/eval/schema/verdict.ts`). Suite v2 results are adapted into it at one explicit boundary (`src/domains/eval/schema/adapter.ts`) rather than by widening the artifact version, because the envelope carries no information a v4 artifact cannot hold.
|
|
210
272
|
|
|
211
273
|
```json
|
|
212
274
|
{
|
|
213
|
-
"schema": "clio.eval.verdict.v1",
|
|
275
|
+
"schema": "clio-coder.eval.verdict.v1",
|
|
214
276
|
"scenarioId": "latency-nonnegative",
|
|
215
277
|
"trialIndex": 0,
|
|
216
278
|
"outcome": "pass",
|
|
@@ -230,7 +292,7 @@ The envelope is fail-closed by construction. `outcome` is one of `pass`, `fail`,
|
|
|
230
292
|
|
|
231
293
|
### Behavioral scenario and verdict documents
|
|
232
294
|
|
|
233
|
-
Behavioral evaluation is additive and does not change the persisted `clio.eval.verdict.v1` reader. A Suite v2 task may declare a `clio.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable.
|
|
295
|
+
Behavioral evaluation is additive and does not change the persisted `clio-coder.eval.verdict.v1` reader. A Suite v2 task may declare a `clio-coder.eval.scenario.v1` block, and its Artifact v4 result then carries a sibling `clio-coder.eval.behavior.v1` document whose `verdictRef` names the verdict schema, scenario id, and trial index. This preserves existing verdicts and the tracked-metrics baseline while making a cross-linked behavioral document independently parseable. Readers normalize the released `clio.eval.*` identifiers for compatibility, but current writers emit only `clio-coder.eval.*` identifiers.
|
|
234
296
|
|
|
235
297
|
The closed categories are `tool_choice`, `exploration`, `delegation`, `safety_comprehension`, `claim_grounding`, `denied_tool_recovery`, `completion_behavior`, and `task_correctness`. Each category result is exactly one of `satisfied`, `violated`, `unknown`, or `unmeasured`. The document outcome is `pass`, `behavioral_failure`, `unknown`, `unmeasured`, or `infrastructure_failure`; missing facts are never invented as successes, and an infrastructure failure cannot become a behavioral pass.
|
|
236
298
|
|
|
@@ -240,8 +302,8 @@ Suite execution adapts scalar run metrics into these observable facts at the Sui
|
|
|
240
302
|
|
|
241
303
|
### Public built-in behavioral corpus
|
|
242
304
|
|
|
243
|
-
The repository
|
|
244
|
-
`
|
|
305
|
+
The source repository carries corpus `public-built-in-behavior` version `1.0.0`
|
|
306
|
+
under `evals/`. It contains no private prompts, endpoints, credentials, or
|
|
245
307
|
mutable external dataset:
|
|
246
308
|
|
|
247
309
|
- `behavioral-machinery.yaml` provides one positive and one adversarial
|
|
@@ -262,12 +324,15 @@ mutable external dataset:
|
|
|
262
324
|
that the rules can reject observed model behavior rather than merely restate
|
|
263
325
|
aggregate success counters.
|
|
264
326
|
|
|
265
|
-
|
|
327
|
+
These are source-checkout workflows: the npm archive keeps the inputs for
|
|
328
|
+
inspection and reproducibility, but the deterministic TypeScript driver uses
|
|
329
|
+
the repository development toolchain. Build once, then run either focused
|
|
330
|
+
suite from the repository root:
|
|
266
331
|
|
|
267
332
|
```sh
|
|
268
|
-
node dist/cli/index.js eval run --suite
|
|
269
|
-
node dist/cli/index.js eval run --suite
|
|
270
|
-
node dist/cli/index.js eval run --suite
|
|
333
|
+
node dist/cli/index.js eval run --suite evals/behavioral-machinery.yaml --clio-coder-entry dist/cli/index.js
|
|
334
|
+
node dist/cli/index.js eval run --suite evals/behavioral-model.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
335
|
+
node dist/cli/index.js eval run --suite evals/behavioral-model-negative-control.yaml --target mini --clio-coder-entry dist/cli/index.js
|
|
271
336
|
```
|
|
272
337
|
|
|
273
338
|
The machinery tasks use the repository read-only and create only private
|
|
@@ -292,7 +357,7 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
292
357
|
| `generatedTokens` | ledger |
|
|
293
358
|
| `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
|
|
294
359
|
| `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
|
|
295
|
-
| `ttftMsFirstCall` | ledger |
|
|
360
|
+
| `ttftMsFirstCall` | ledger; nullable when first-call timing is absent |
|
|
296
361
|
| `wallClockMs` | receipt |
|
|
297
362
|
| `contextTokensAtEnd` | ledger |
|
|
298
363
|
| `compactions` | ledger |
|
|
@@ -300,14 +365,18 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
300
365
|
|
|
301
366
|
A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
|
|
302
367
|
|
|
368
|
+
First-call TTFT uses the earliest recorded assistant-call timestamp across the selected ledgers; equal timestamps retain their observed order. Missing or invalid timing, or an invalid timestamp that prevents ordering the calls, yields `{ value: null, source: "estimated" }`. A measured zero remains `{ value: 0, source: "ledger" }`. Native session timing starts at each stream invocation and includes the provider's response-header wait. Stdout-only fallback timing starts at the provider's `message_start`, which can arrive after headers; it requires first output, but is not complete request latency and must not be compared as equivalent to native invocation timing. A completion alone supplies neither a zero duration nor a first-token measurement. Verdict v1 consumers must accept nullable TTFT. Historical numeric values, including estimated zeros and native spans that omitted the pre-header wait, remain readable and are not rewritten.
|
|
369
|
+
|
|
370
|
+
When a stream message lacks a valid timestamp, its ledger payload marks `timestampEstimated: true` beside the legacy ISO placeholder. This leaves first-call chronology unmeasured while preserving any observed per-call monotonic timing.
|
|
371
|
+
|
|
303
372
|
### Scenario aggregates
|
|
304
373
|
|
|
305
374
|
`aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
|
|
306
375
|
|
|
307
376
|
### Behavioral multi-metric results
|
|
308
377
|
|
|
309
|
-
A result with a `clio.eval.behavior.v1` verdict also carries the additive
|
|
310
|
-
`clio.eval.behavior.metrics.v1` projection. The projection binds the scenario
|
|
378
|
+
A result with a `clio-coder.eval.behavior.v1` verdict also carries the additive
|
|
379
|
+
`clio-coder.eval.behavior.metrics.v1` projection. The projection binds the scenario
|
|
311
380
|
to its role and target/model envelope and records one `number | null`
|
|
312
381
|
observation for each closed metric. The source travels beside every value:
|
|
313
382
|
|
|
@@ -356,8 +425,8 @@ is emitted as testcase output rather than a failed testcase.
|
|
|
356
425
|
### Execution-envelope provenance and comparability
|
|
357
426
|
|
|
358
427
|
Every newly written behavioral result carries an additive
|
|
359
|
-
`clio.eval.execution-envelope.v1` sibling. Artifact v4,
|
|
360
|
-
`clio.eval.verdict.v1`, and `clio.eval.behavior.metrics.v1` retain their
|
|
428
|
+
`clio-coder.eval.execution-envelope.v1` sibling. Artifact v4,
|
|
429
|
+
`clio-coder.eval.verdict.v1`, and `clio-coder.eval.behavior.metrics.v1` retain their
|
|
361
430
|
existing identities. The envelope records the selected prompt fragment ids,
|
|
362
431
|
authored versions or `unversioned` marker, fragment content hashes, prompt
|
|
363
432
|
composition hash, recipe id/version/fingerprint when a worker recipe applies,
|
|
@@ -382,31 +451,15 @@ metric means and variances. When the prompt or recipe identity changes, the
|
|
|
382
451
|
generated evidence names each affected corpus scenario and role instead of
|
|
383
452
|
hiding it behind an aggregate score.
|
|
384
453
|
|
|
385
|
-
###
|
|
454
|
+
### Reference behavioral baseline
|
|
386
455
|
|
|
387
|
-
|
|
388
|
-
|
|
389
|
-
|
|
390
|
-
|
|
391
|
-
|
|
392
|
-
|
|
393
|
-
|
|
394
|
-
the evidence, run the same machinery suite first, inspect the failing diff and
|
|
395
|
-
the named affected corpus results, then update explicitly:
|
|
396
|
-
|
|
397
|
-
```sh
|
|
398
|
-
npm run build
|
|
399
|
-
node benchmarks/eval/check-behavioral-release.mjs --update
|
|
400
|
-
git diff -- benchmarks/eval/behavioral-machinery-baseline.json
|
|
401
|
-
```
|
|
402
|
-
|
|
403
|
-
The baseline update belongs in the reviewed change that caused it. Do not use
|
|
404
|
-
the update command merely to make a red gate green. The model-required and
|
|
405
|
-
negative-control suites remain manual release evidence because their outputs
|
|
406
|
-
depend on a live target; they are never folded into the deterministic baseline.
|
|
407
|
-
The projection excludes `latency.wallMs` because scheduler timing is not stable
|
|
408
|
-
evidence. Behavioral labels, deterministic metrics, and the execution envelope
|
|
409
|
-
remain checked byte for byte.
|
|
456
|
+
`evals/behavioral-machinery-baseline.json` is retained as reviewable reference
|
|
457
|
+
evidence for the machinery corpus. It is not a CI or release gate. Run the
|
|
458
|
+
current `evals/behavioral-machinery.yaml` through the built CLI when a prompt,
|
|
459
|
+
recipe, policy, or expected-behavior change needs a fresh measurement, inspect
|
|
460
|
+
the named scenario evidence, and update any retained baseline deliberately in
|
|
461
|
+
the reviewed change. Model-required and negative-control suites remain manual
|
|
462
|
+
measurements tied to their exact target and serving configuration.
|
|
410
463
|
|
|
411
464
|
### Hard thresholds and informational budgets
|
|
412
465
|
|
|
@@ -443,10 +496,12 @@ cannot offset a task or safety regression.
|
|
|
443
496
|
|
|
444
497
|
```text
|
|
445
498
|
serving configuration drift; pass --allow-config-drift to compare these runs
|
|
446
|
-
baseline serving: target=mini runtime=llamacpp model
|
|
499
|
+
baseline serving: target=mini runtime=llamacpp model=ornith1.5-35b-moe server_build=... total_slots=4 thinking=off compiled_prompt_hash=...
|
|
447
500
|
candidate serving: ...
|
|
448
501
|
```
|
|
449
502
|
|
|
503
|
+
The current reference `mini` endpoint is the llama.cpp router at `192.168.86.141:8080`. It serves `ornith1.5-35b-moe` with four parallel slots and 262144 context tokens per slot. These deployment facts are reference topology, not defaults imposed on another target; retain the artifact's observed serving configuration with every comparison.
|
|
504
|
+
|
|
450
505
|
`--allow-config-drift` proceeds and labels the comparison `config drift: allowed`. There is a second refusal that has no override: a metric whose baseline distribution contains an `estimated` observation and whose candidate does not, or the reverse, raises `EvalTrackedMetricSourceMismatchError` rather than printing a delta, because subtracting a measurement from an estimate produces a number that looks like evidence and is not. `--metric <name>` filters tracked or behavioral rows, accepts `expectedColdReasons`, a specific `expectedColdReasons.<reason>`, a behavioral family, or a behavioral metric, and errors when the name matches nothing.
|
|
451
506
|
|
|
452
507
|
---
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Internal Eval Suites
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Internal Eval Suites visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evals_internal_blueprint.html).
|
|
5
5
|
|
|
6
6
|
Private suites should live outside this repository. Keep datasets, prompts,
|
|
7
7
|
live fleet coordinates, calibration outputs, and raw run artifacts in a private
|
|
@@ -15,13 +15,13 @@ clio-coder eval run --suite <external-path> --clio-coder-entry dist/cli/index.js
|
|
|
15
15
|
```
|
|
16
16
|
|
|
17
17
|
Use `--out <dir>` when the artifact should be written outside the default Clio
|
|
18
|
-
data directory.
|
|
19
|
-
|
|
20
|
-
|
|
18
|
+
data directory. External benchmark campaigns should adapt their cases and
|
|
19
|
+
grader observations into the same eval engine while keeping private datasets,
|
|
20
|
+
credentials, endpoints, and raw artifacts outside this repository.
|
|
21
21
|
|
|
22
22
|
The public behavioral corpus is the deliberate exception to the otherwise
|
|
23
23
|
private Suite v2 data policy. Its reviewable, synthetic suites live under
|
|
24
|
-
`
|
|
24
|
+
`evals/`: a model-free positive/adversarial authority pair for every
|
|
25
25
|
built-in worker recipe, four tiny main-agent model scenarios covering all eight
|
|
26
26
|
behavioral categories with event- and grader-derived facts, and an intentional
|
|
27
27
|
decoy negative control. The model-free driver uses the shipped recipe catalog,
|
|
@@ -83,7 +83,7 @@ thinking level is measuring the server, not the change under test.
|
|
|
83
83
|
|
|
84
84
|
The verdict envelope keeps its original `behavioral: null` field for compatibility.
|
|
85
85
|
A suite that declares a versioned behavioral scenario records the result as a
|
|
86
|
-
separate `clio.eval.behavior.v1` document on the Artifact v4 result, cross-linked
|
|
86
|
+
separate `clio-coder.eval.behavior.v1` document on the Artifact v4 result, cross-linked
|
|
87
87
|
to the unchanged verdict identity. Its labels come only from bounded transcript,
|
|
88
88
|
tool, receipt, or grader facts, never from an ungrounded judge paragraph. A run whose harness broke records
|
|
89
89
|
`machinery: "infrastructure_failure"`, which the parser refuses to pair with a
|
|
@@ -210,7 +210,7 @@ tasks:
|
|
|
210
210
|
- node_modules
|
|
211
211
|
- dist
|
|
212
212
|
runner:
|
|
213
|
-
kind: clio-run
|
|
213
|
+
kind: clio-coder-run
|
|
214
214
|
prompt: Fix the intentionally broken function so the local verifier passes.
|
|
215
215
|
verify:
|
|
216
216
|
commands:
|
|
@@ -271,7 +271,7 @@ tasks:
|
|
|
271
271
|
- dist
|
|
272
272
|
- .clio-coder
|
|
273
273
|
runner:
|
|
274
|
-
kind: clio-run
|
|
274
|
+
kind: clio-coder-run
|
|
275
275
|
prompt: Summarize the repository purpose and make no file changes.
|
|
276
276
|
verify:
|
|
277
277
|
forbidPaths:
|
|
@@ -305,7 +305,7 @@ tasks:
|
|
|
305
305
|
- dist
|
|
306
306
|
- .clio-coder
|
|
307
307
|
runner:
|
|
308
|
-
kind: clio-run
|
|
308
|
+
kind: clio-coder-run
|
|
309
309
|
prompt: Fix the failing unit test with the smallest source change.
|
|
310
310
|
verify:
|
|
311
311
|
commands:
|
|
@@ -1,7 +1,7 @@
|
|
|
1
1
|
# Evolution and Change Manifests
|
|
2
2
|
|
|
3
|
-
>
|
|
4
|
-
>
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Evolution and Change Manifests visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/evolution_blueprint.html).
|
|
5
5
|
|
|
6
6
|
Clio Coder uses change manifests to make harness changes reviewable, falsifiable, and rollback-friendly. CLIO stands for Context Layer for Input/Output, named for the Greek muse of history. A manifest is JSON, generated or checked with `clio-coder evolve manifest`, and should describe what changed, why, what evidence supports it, what could regress, how to validate it, and how to roll it back.
|
|
7
7
|
|
|
@@ -1,10 +1,13 @@
|
|
|
1
1
|
# Fleet Demo Runbook
|
|
2
2
|
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Fleet Demo Runbook visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/fleet_demo_blueprint.html).
|
|
5
|
+
|
|
3
6
|
A repeatable multi-node demonstration: one orchestrator drives a real
|
|
4
7
|
CMake/C++ fix through a reviewer-gated dispatch across SSH nodes, and every
|
|
5
8
|
worker's receipt (including the remote ones) verifies afterward. The steps
|
|
6
9
|
are executable in order; this document doubles as the recording script.
|
|
7
|
-
Background and reference: [fleet-dispatch.md](fleet-dispatch.md).
|
|
10
|
+
Background and reference: [fleet-dispatch.md](../guide/fleet-dispatch.md).
|
|
8
11
|
|
|
9
12
|
## Reference fabric
|
|
10
13
|
|
|
@@ -139,9 +142,9 @@ clio-coder evidence inspect <evidenceId>
|
|
|
139
142
|
run ledger; a tampered or mismatched receipt fails the build with the field
|
|
140
143
|
that diverged. The receipts of the remote runs verify on the orchestrator host because the
|
|
141
144
|
ledger and receipts live on the shared filesystem. Current receipts use strict
|
|
142
|
-
|
|
143
|
-
|
|
144
|
-
|
|
145
|
+
v20 and authenticate every current receipt and reconstructed-ledger field.
|
|
146
|
+
Lower versions are reported as retired and are never read as evidence or
|
|
147
|
+
migrated; malformed and future shapes fail verification.
|
|
145
148
|
|
|
146
149
|
## Provenance walkthrough: what a PI can verify from receipts alone
|
|
147
150
|
|
|
@@ -171,9 +174,10 @@ reconstruct:
|
|
|
171
174
|
complete receipt schema and its stable ledger row. `clio-coder evidence build
|
|
172
175
|
--run <id>` recomputes and cross-checks it; `verifyReceiptIntegrity` in
|
|
173
176
|
`src/domains/dispatch/receipt-integrity.ts` is the reference
|
|
174
|
-
implementation. Current receipts use
|
|
175
|
-
|
|
176
|
-
read as evidence through
|
|
177
|
+
implementation. Current receipts use v20. Lower versions are reported as
|
|
178
|
+
retired, while malformed and future shapes fail verification. Incompatible
|
|
179
|
+
state may be archived for inspection, but it is never read as evidence through
|
|
180
|
+
a compatibility verifier.
|
|
177
181
|
|
|
178
182
|
The walkthrough for an audience is three commands: `clio-coder evidence build
|
|
179
183
|
--run <id>` (it verifies), open the receipt JSON (read `node`, `gate`,
|
|
@@ -1,13 +1,20 @@
|
|
|
1
1
|
# Git Commit Provenance
|
|
2
2
|
|
|
3
|
+
> **Visual blueprint:** The source checkout includes the complete
|
|
4
|
+
> [Git Commit Provenance visual reference](https://github.com/iowarp/clio-coder/blob/main/docs/html/git_commit_provenance_blueprint.html).
|
|
5
|
+
|
|
3
6
|
Clio Coder adds evidence-aware role trailers to commits created through Clio.
|
|
4
7
|
The feature is enabled by default:
|
|
5
8
|
|
|
6
9
|
```yaml
|
|
7
|
-
|
|
8
|
-
|
|
10
|
+
integrations:
|
|
11
|
+
git:
|
|
12
|
+
commitAttribution: true
|
|
9
13
|
```
|
|
10
14
|
|
|
15
|
+
The released `attribution.gitCommits` path is a migration alias. Current
|
|
16
|
+
settings files and writers use `integrations.git.commitAttribution`.
|
|
17
|
+
|
|
11
18
|
Settings -> Advanced exposes the same switch as **Clio commit provenance**, with
|
|
12
19
|
`enabled` and `disabled` values. A change applies immediately to subsequent
|
|
13
20
|
commits in the session. When disabled, Clio leaves commit messages entirely
|
|
@@ -49,11 +56,11 @@ Co-authored-by: Clio Coder <clio-coder@iowarp.ai>
|
|
|
49
56
|
Existing human trailers stay in place. A Clio trailer already present in any
|
|
50
57
|
letter case is respected rather than repeated, line endings are normalized only
|
|
51
58
|
while attribution is enabled, and repeated processing is idempotent. When a directly relevant
|
|
52
|
-
receipt-
|
|
59
|
+
receipt-v20 digest passes integrity verification, Clio may additionally add the
|
|
53
60
|
full digest:
|
|
54
61
|
|
|
55
62
|
```text
|
|
56
|
-
Clio-Evidence: receipt-
|
|
63
|
+
Clio-Evidence: receipt-v20/sha256:<64-character digest>
|
|
57
64
|
```
|
|
58
65
|
|
|
59
66
|
Clio does not invent, shorten, or add an unrelated digest. The role trailers do
|