@iowarp/clio-coder 0.4.2 → 0.4.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +35 -0
- package/CONTRIBUTING.md +86 -19
- package/README.md +35 -6
- package/dist/{acp-TMDQZDIG.js → acp-H2NGRPWO.js} +11 -11
- package/dist/{agents-5N5NG3XG.js → agents-TL5LLUQP.js} +54 -53
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-Z5CCBXKQ.js → auth-E5SW4HMS.js} +19 -16
- package/dist/{builtins-K6TNDT24.js → builtins-IA7V7FUC.js} +9 -4
- package/dist/{chunk-ZW4HH5JJ.js → chunk-2APPQIER.js} +6 -6
- package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
- package/dist/{chunk-UBRFI4HS.js → chunk-2UG5F4C5.js} +127 -47
- package/dist/{chunk-UH632ZYL.js → chunk-2UH2KFUP.js} +2 -2
- package/dist/{chunk-3F7VUY77.js → chunk-2VIKGWFZ.js} +2 -2
- package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
- package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
- package/dist/{chunk-2X4RYJTJ.js → chunk-4UVU7BJ5.js} +2 -2
- package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
- package/dist/{chunk-5KW52TEP.js → chunk-54CBCGIR.js} +5 -5
- package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
- package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
- package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
- package/dist/{chunk-LJID3DYZ.js → chunk-64I3JVYM.js} +2 -2
- package/dist/{chunk-4JDLP6ZS.js → chunk-6PTFB5VS.js} +7 -7
- package/dist/{chunk-VN3SHNBN.js → chunk-7DICMOS6.js} +2 -2
- package/dist/{chunk-HLAFFSEK.js → chunk-7DRAWPTZ.js} +2 -2
- package/dist/chunk-7E7I3WLS.js +3762 -0
- package/dist/{chunk-YJISEZKC.js → chunk-7ZYNNDKC.js} +6 -6
- package/dist/{chunk-I66ZTYNP.js → chunk-AF4YM7Z4.js} +236 -101
- package/dist/{chunk-2HFQNRV3.js → chunk-AX2THNSA.js} +12 -12
- package/dist/{chunk-PGF63K6I.js → chunk-B4OAX3SI.js} +65 -3
- package/dist/{chunk-W6NIE6OW.js → chunk-B4VEBZKF.js} +3 -3
- package/dist/{chunk-JBCS7CRR.js → chunk-BEPZRGGU.js} +10 -10
- package/dist/{chunk-XGDPUNND.js → chunk-CE5AX47J.js} +2 -2
- package/dist/{chunk-I64IFBLB.js → chunk-DWUOQKRU.js} +17 -10
- package/dist/{chunk-DZAW46HP.js → chunk-E3TPLWFX.js} +3 -3
- package/dist/{chunk-HIICAHCJ.js → chunk-EKCHAPYA.js} +2 -2
- package/dist/{chunk-XE3PCIXH.js → chunk-F5JHEYZM.js} +7 -7
- package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
- package/dist/{chunk-ZNT2M6TG.js → chunk-G76U63X4.js} +17 -17
- package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
- package/dist/{chunk-2NHR3NAY.js → chunk-GI7YYQ3F.js} +40 -34
- package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
- package/dist/chunk-GYV6VZOC.js +26 -0
- package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
- package/dist/{chunk-W4YEMFBX.js → chunk-HEQY7ZFI.js} +2 -2
- package/dist/{chunk-IKOZFYBN.js → chunk-I7ZPNEJM.js} +145 -102
- package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
- package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
- package/dist/chunk-IJNZMHLA.js +101 -0
- package/dist/{chunk-JWJGP5DQ.js → chunk-INY6HTFL.js} +7 -7
- package/dist/{chunk-PBP4B7XR.js → chunk-IUE3Y34X.js} +2 -2
- package/dist/{chunk-B74PXLU7.js → chunk-IWT4SF4R.js} +3 -3
- package/dist/{chunk-B7HM5Z7T.js → chunk-JDAY6FIL.js} +5 -5
- package/dist/{chunk-PJX3WQUQ.js → chunk-JEQ3XTHC.js} +2 -2
- package/dist/{chunk-FSP7CMNU.js → chunk-JGRC33J2.js} +50 -4
- package/dist/{chunk-X7IARSHT.js → chunk-JKKCYP3C.js} +9 -9
- package/dist/{chunk-HJWWJ6IL.js → chunk-JSC3U7TI.js} +16 -4
- package/dist/{chunk-SSEYRH53.js → chunk-KK4JZPBQ.js} +19 -140
- package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
- package/dist/{chunk-Q4XWMHX6.js → chunk-L47TF46W.js} +2 -2
- package/dist/{chunk-O3YUNJZ2.js → chunk-LDJG7DW3.js} +81 -24
- package/dist/{chunk-CDNVLKUX.js → chunk-LLDJM5XK.js} +13 -7
- package/dist/{chunk-F2I26BDK.js → chunk-MUW2BDDH.js} +4 -4
- package/dist/{chunk-HKMD33FO.js → chunk-MWUZBSAQ.js} +79 -76
- package/dist/{chunk-QQLGQY2A.js → chunk-N2Z7HLVY.js} +20 -20
- package/dist/{chunk-DZEK6CJN.js → chunk-NIQJ66N4.js} +19 -19
- package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
- package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
- package/dist/{chunk-TPEQIQIE.js → chunk-OML5D5V5.js} +8 -8
- package/dist/{chunk-IKSLQ4XV.js → chunk-PAJQJ7BS.js} +558 -216
- package/dist/{chunk-ZW55JB7N.js → chunk-PUVDKJ2Y.js} +2 -2
- package/dist/{chunk-UH347SHR.js → chunk-QWGDJJYJ.js} +11 -11
- package/dist/chunk-R6Q67RJH.js +134 -0
- package/dist/{chunk-CRFOIAX3.js → chunk-RRNP2ANY.js} +6 -6
- package/dist/{chunk-IDNA72AH.js → chunk-RSJ25QSL.js} +2 -2
- package/dist/chunk-SKHCAU7K.js +385 -0
- package/dist/{chunk-RLYRBIYQ.js → chunk-TM6LQDI3.js} +20 -12
- package/dist/chunk-UOIZ7DA4.js +41 -0
- package/dist/{chunk-P75RZCJW.js → chunk-UPZU6GE4.js} +3 -3
- package/dist/{chunk-MCMZMDAC.js → chunk-V2ANDPVT.js} +4 -4
- package/dist/{chunk-AK5XEFVZ.js → chunk-VA5FNYMT.js} +26 -13
- package/dist/{chunk-IMXMHHMQ.js → chunk-VW6DOEDG.js} +332 -57
- package/dist/{chunk-XOXV5GKE.js → chunk-W6RRQCPQ.js} +16 -7
- package/dist/{chunk-CYZW7JHJ.js → chunk-WBKFA554.js} +8 -8
- package/dist/{chunk-BO7Y52RY.js → chunk-WCXUNS7U.js} +7 -7
- package/dist/{chunk-ZGNYYXQ6.js → chunk-WRBAGUNF.js} +3 -3
- package/dist/{chunk-FVDGR2ZL.js → chunk-XIVNBFZS.js} +85 -30
- package/dist/{chunk-BYMNWQ7O.js → chunk-XPWWI35G.js} +299 -58
- package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
- package/dist/{chunk-AZ4WMN4W.js → chunk-Y3CBHOR6.js} +2 -2
- package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
- package/dist/{chunk-54ODD65L.js → chunk-YQWYVTMC.js} +4 -4
- package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
- package/dist/{chunk-7BHIY2MW.js → chunk-ZDN3Y73Y.js} +6 -6
- package/dist/{chunk-E7GT7O5N.js → chunk-ZWPRK62N.js} +7 -4
- package/dist/cli/index.js +38 -37
- package/dist/{clio-7VB377CC.js → clio-CMMK4KRR.js} +7 -7
- package/dist/{code-nav-YVLCYA7V.js → code-nav-MDZNQS33.js} +7 -7
- package/dist/{components-UBWCQSRW.js → components-UCUQ4QXW.js} +4 -4
- package/dist/{config-4HVOS65E.js → config-SVM5P5YI.js} +76 -74
- package/dist/{configure-PIWO7B24.js → configure-LE3IK2TJ.js} +26 -24
- package/dist/{context-IYEHL3WQ.js → context-2OHRKS42.js} +66 -63
- package/dist/{context-N6ZE3LGJ.js → context-E3VC7RX5.js} +15 -11
- package/dist/{context-KQYIWPWT.js → context-VNCR7KAG.js} +60 -45
- package/dist/{context-clear-G4OGZJDS.js → context-clear-BW4O37TG.js} +61 -59
- package/dist/context-map-COB37XXN.js +505 -0
- package/dist/{context-working-set-BWLF6LJP.js → context-working-set-VDS25HXZ.js} +17 -16
- package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-5AHT53RF.js} +85 -74
- package/dist/{doctor-LHBD36VU.js → doctor-WNNVO6FY.js} +37 -37
- package/dist/{eval-C45FYRJ6.js → eval-7G7SGAYO.js} +285 -114
- package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
- package/dist/{evidence-6SHONYAF.js → evidence-VD6736FQ.js} +63 -62
- package/dist/{evolve-KRKMV72X.js → evolve-AL3NGVRL.js} +62 -61
- package/dist/{extensions-KPZ2UHBB.js → extensions-MOVJ32NM.js} +7 -7
- package/dist/{fleet-IVTCKDHT.js → fleet-QZHUMAGI.js} +110 -108
- package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-BAYT5FJZ.js} +10 -10
- package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-IREVMRU4.js} +7 -6
- package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-YCTT3HTI.js} +19 -18
- package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-QVJTDAVB.js} +55 -54
- package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-25QAFPK4.js} +4 -4
- package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-5O57AAJ7.js} +23 -22
- package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-CPH2W2T6.js} +56 -55
- package/dist/{fleet-view-TWHJKCN6.js → fleet-view-SWBR3VGQ.js} +55 -54
- package/dist/{init-T2QORQ3Y.js → init-J477LKZH.js} +78 -76
- package/dist/{interop-IN5I2A66.js → interop-3FCM6XLG.js} +11 -11
- package/dist/{library-LSCATDLZ.js → library-QUQEIUG6.js} +28 -27
- package/dist/{memory-HYOKAGGJ.js → memory-SGGSEP65.js} +64 -63
- package/dist/{models-2GPMFYCM.js → models-HEKUAXXK.js} +49 -43
- package/dist/{monitor-E4ASVUJH.js → monitor-HKU57TYQ.js} +61 -60
- package/dist/{orchestrator-DDMPR3PY.js → orchestrator-VDFAEFAI.js} +919 -546
- package/dist/{panes-E3RUXOW5.js → panes-DN2SSFOH.js} +3 -3
- package/dist/{panes-IXKLOKA2.js → panes-TALGNPZT.js} +8 -8
- package/dist/{paths-L7LGY6RN.js → paths-NBMFAIEZ.js} +5 -5
- package/dist/reset-EAJFFJVB.js +344 -0
- package/dist/{resources-OTRSN34L.js → resources-OVKSEFVE.js} +27 -20
- package/dist/{run-5DEYH5QK.js → run-7DP7ZF2J.js} +113 -109
- package/dist/{share-IHWTLO3M.js → share-WML67FT3.js} +26 -25
- package/dist/{skills-IYMXMKW4.js → skills-SG662R2K.js} +39 -31
- package/dist/{skills-eval-DROHSJAR.js → skills-eval-VVZEUU46.js} +74 -73
- package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-I2E23GET.js} +21 -20
- package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-S7MBJDQK.js} +35 -34
- package/dist/{steer-Z5DO23FJ.js → steer-2LQOMCPB.js} +3 -3
- package/dist/{support-U7QOWY26.js → support-CC2UJBJ6.js} +6 -6
- package/dist/{targets-P2FUC4IL.js → targets-4QC3HIEW.js} +48 -45
- package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-TUHIJ6Y2.js} +2 -2
- package/dist/{tools-5B7RO6MV.js → tools-TFGJICCU.js} +8 -8
- package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
- package/dist/uninstall-5PEVOE5B.js +408 -0
- package/dist/upgrade-M4WXY6KN.js +303 -0
- package/dist/{usage-ME5MPXGX.js → usage-N7ZNVLEM.js} +147 -102
- package/dist/{verifiers-BVZ7IWOO.js → verifiers-DJTP4XX6.js} +15 -15
- package/dist/{verify-5K7ZKQFC.js → verify-RWE4PPEK.js} +9 -9
- package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-C7IQOXSP.js} +84 -82
- package/dist/{with-panes-BYOJCLAM.js → with-panes-4GCGSL7J.js} +9 -9
- package/dist/worker/entry.js +61 -60
- package/docs/architecture/artifact-placement.md +1 -0
- package/docs/architecture/artifact-versions.md +1 -1
- package/docs/architecture/context-engine.md +4 -0
- package/docs/architecture/middleware-and-components.md +1 -1
- package/docs/architecture/model-catalog.md +21 -10
- package/docs/architecture/observability.md +12 -1
- package/docs/architecture/prompt-envelope-and-tools.md +2 -0
- package/docs/architecture/provider-adapter-cookbook.md +63 -0
- package/docs/architecture/safety-model.md +15 -5
- package/docs/guide/built-in-agents.md +17 -3
- package/docs/guide/commands-and-modes.md +1 -1
- package/docs/guide/configuration-and-targets.md +97 -9
- package/docs/guide/configuration-reference.md +7 -2
- package/docs/guide/environment-variables.md +2 -0
- package/docs/guide/installation-and-lifecycle.md +37 -4
- package/docs/guide/proactive-memory.md +66 -55
- package/docs/guide/skills-marketplace.md +18 -0
- package/docs/process/development-pipeline.md +34 -1
- package/docs/process/eval-runner.md +67 -3
- package/evals/behavioral-model.yaml +3 -2
- package/package.json +2 -2
- package/skills/README.md +7 -5
- package/skills/coding/ast-grep/SKILL.md +101 -30
- package/skills/coding/ast-grep/evals.md +26 -0
- package/skills/coding/coding-standards/SKILL.md +40 -5
- package/skills/coding/coding-standards/evals.md +23 -0
- package/skills/coding/prototype/SKILL.md +87 -28
- package/skills/coding/prototype/evals.md +19 -0
- package/skills/coding/tdd/SKILL.md +80 -53
- package/skills/coding/tdd/evals.md +20 -0
- package/skills/context/context-handoff/SKILL.md +43 -2
- package/skills/context/context-handoff/evals.md +44 -0
- package/skills/context/context-prime/SKILL.md +45 -15
- package/skills/context/context-prime/evals.md +45 -0
- package/skills/git/branch-closeout/SKILL.md +132 -0
- package/skills/git/branch-closeout/evals.md +133 -0
- package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
- package/skills/git/file-ticket/SKILL.md +77 -63
- package/skills/git/file-ticket/assets/issue-template.md +22 -0
- package/skills/git/file-ticket/evals.md +31 -26
- package/skills/git/file-ticket/references/issue-discovery.md +49 -0
- package/skills/git/fix-issue/SKILL.md +87 -64
- package/skills/git/fix-issue/evals.md +35 -31
- package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
- package/skills/git/resolve-merge-conflicts/evals.md +52 -25
- package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
- package/skills/git/ship/SKILL.md +103 -67
- package/skills/git/ship/assets/pr-template.md +21 -0
- package/skills/git/ship/evals.md +44 -28
- package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
- package/skills/git/worktree-create/SKILL.md +80 -50
- package/skills/git/worktree-create/evals.md +40 -33
- package/skills/git/worktree-create/references/worktree-setup.md +62 -66
- package/skills/git/worktree-merge/SKILL.md +112 -65
- package/skills/git/worktree-merge/evals.md +42 -34
- package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
- package/skills/planning/archify/SKILL.md +196 -0
- package/skills/planning/archify/evals.md +65 -0
- package/skills/planning/architecture/SKILL.md +61 -12
- package/skills/planning/architecture/evals.md +65 -0
- package/skills/planning/backlog/SKILL.md +130 -14
- package/skills/planning/backlog/evals.md +142 -0
- package/skills/planning/prd/SKILL.md +46 -6
- package/skills/planning/prd/evals.md +54 -0
- package/skills/planning/product-intent/SKILL.md +57 -2
- package/skills/planning/product-intent/evals.md +70 -0
- package/skills/planning/tech-spec/SKILL.md +53 -2
- package/skills/planning/tech-spec/evals.md +73 -0
- package/skills/registry.yaml +58 -50
- package/skills/remote.yaml +13 -0
- package/skills/research/arxiv-literature/SKILL.md +76 -18
- package/skills/research/arxiv-literature/evals.md +50 -0
- package/skills/research/experiment-protocol/SKILL.md +20 -1
- package/skills/research/experiment-protocol/evals.md +23 -0
- package/skills/research/scientific-debugging/SKILL.md +23 -1
- package/skills/research/scientific-debugging/evals.md +18 -0
- package/skills/research/scientific-modernization/SKILL.md +26 -1
- package/skills/research/scientific-modernization/evals.md +27 -0
- package/skills/skill-marketplace.json +63 -28
- package/skills/workflow/cut-it/SKILL.md +65 -5
- package/skills/workflow/cut-it/evals.md +101 -0
- package/skills/workflow/design-council/SKILL.md +117 -27
- package/skills/workflow/design-council/evals.md +161 -0
- package/skills/workflow/grill-me/SKILL.md +86 -10
- package/skills/workflow/grill-me/evals.md +153 -0
- package/skills/workflow/workflow-distiller/SKILL.md +76 -17
- package/skills/workflow/workflow-distiller/evals.md +118 -0
- package/src/cli/configure-interop.ts +105 -13
- package/src/cli/configure-oauth.ts +57 -0
- package/src/cli/configure-onboarding.ts +980 -0
- package/src/cli/configure-target.ts +594 -0
- package/src/cli/configure.ts +1082 -528
- package/src/cli/context-map.ts +114 -0
- package/src/cli/context.ts +4 -0
- package/src/cli/index.ts +1 -0
- package/src/cli/lifecycle-presenter.ts +436 -0
- package/src/cli/models.ts +10 -2
- package/src/cli/modes/print.ts +5 -1
- package/src/cli/reset.ts +228 -106
- package/src/cli/run.ts +7 -2
- package/src/cli/select.ts +664 -0
- package/src/cli/skills.ts +9 -2
- package/src/cli/targets.ts +3 -0
- package/src/cli/uninstall.ts +233 -165
- package/src/cli/upgrade.ts +204 -149
- package/src/cli/usage.ts +86 -27
- package/src/cli/validate-model.ts +3 -3
- package/src/core/config.ts +56 -0
- package/src/core/external-diagnostic.ts +44 -0
- package/src/core/gateway-routing.ts +157 -0
- package/src/core/safe-exec.ts +17 -2
- package/src/core/skill-activation.ts +89 -2
- package/src/domains/agents/builtins/world-knowledge.md +31 -0
- package/src/domains/agents/catalog.ts +1 -1
- package/src/domains/agents/result-contract.ts +70 -0
- package/src/domains/context/wiki/map-seed.ts +589 -0
- package/src/domains/context/wiki/plan.ts +2 -2
- package/src/domains/dispatch/admission.ts +29 -0
- package/src/domains/dispatch/agent-candidates.ts +10 -0
- package/src/domains/dispatch/budget-envelope.ts +86 -1
- package/src/domains/dispatch/capability-match.ts +1 -0
- package/src/domains/dispatch/capacity-lease.ts +17 -0
- package/src/domains/dispatch/contract.ts +11 -1
- package/src/domains/dispatch/extension.ts +134 -29
- package/src/domains/dispatch/types.ts +3 -0
- package/src/domains/dispatch/worker-model-metadata.ts +38 -0
- package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
- package/src/domains/eval/metrics/token-stream.ts +201 -31
- package/src/domains/eval/metrics/tracked.ts +40 -4
- package/src/domains/eval/runners/clio-run.ts +5 -2
- package/src/domains/eval/schema/suite.ts +28 -0
- package/src/domains/eval/schema/verdict.ts +2 -2
- package/src/domains/eval/suites/resolve.ts +13 -1
- package/src/domains/eval/suites/run.ts +24 -3
- package/src/domains/interop/registry.ts +6 -2
- package/src/domains/interop/types.ts +4 -0
- package/src/domains/lifecycle/migrations/index.ts +4 -0
- package/src/domains/memory/task-memory-policy.ts +70 -26
- package/src/domains/memory/task-memory-telemetry.ts +1 -0
- package/src/domains/middleware/index.ts +0 -1
- package/src/domains/middleware/marketplace-offer.ts +3 -35
- package/src/domains/middleware/memory-intervention.ts +127 -32
- package/src/domains/middleware/memory-step-endpoint.ts +3 -2
- package/src/domains/middleware/skills-reminder.ts +31 -2
- package/src/domains/observability/compaction-usage.ts +118 -0
- package/src/domains/observability/cost.ts +1 -1
- package/src/domains/observability/extension.ts +6 -1
- package/src/domains/observability/out-of-turn-usage.ts +52 -21
- package/src/domains/providers/contract.ts +4 -1
- package/src/domains/providers/extension.ts +40 -9
- package/src/domains/providers/model-capabilities.ts +9 -0
- package/src/domains/providers/model-discovery.ts +2 -0
- package/src/domains/providers/model-runtime-capabilities.ts +15 -5
- package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +32 -12
- package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
- package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
- package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
- package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
- package/src/domains/providers/support.ts +11 -5
- package/src/domains/providers/target-model-cache.ts +25 -2
- package/src/domains/providers/types/capability-flags.ts +2 -0
- package/src/domains/providers/types/runtime-descriptor.ts +20 -1
- package/src/domains/providers/types/target-descriptor.ts +19 -0
- package/src/domains/resources/index.ts +3 -0
- package/src/domains/resources/skills/install.ts +72 -7
- package/src/domains/resources/skills/loader.ts +7 -0
- package/src/domains/resources/skills/marketplace.ts +63 -11
- package/src/domains/safety/autonomy.ts +15 -0
- package/src/domains/safety/index.ts +1 -0
- package/src/domains/safety/path-policy.ts +1 -1
- package/src/domains/safety/policy-engine.ts +34 -11
- package/src/domains/safety/protected-artifacts.ts +191 -88
- package/src/domains/safety/run-effects.ts +2 -22
- package/src/domains/safety/skill-authority.ts +55 -0
- package/src/domains/session/compaction/compact.ts +72 -22
- package/src/domains/session/entries.ts +6 -0
- package/src/domains/session/usage.ts +3 -3
- package/src/engine/agent.ts +13 -3
- package/src/engine/ai.ts +26 -8
- package/src/engine/antigravity/subprocess-runtime.ts +386 -120
- package/src/engine/api-registry.ts +3 -0
- package/src/engine/apis/openai-completions.ts +117 -14
- package/src/engine/external-subprocess.ts +114 -6
- package/src/entry/background-model-metadata.ts +18 -0
- package/src/entry/compaction-prompt.ts +57 -0
- package/src/entry/orchestrator.ts +405 -216
- package/src/entry/task-memory-lifecycle.ts +35 -0
- package/src/interactive/chat-loop-messages.ts +13 -4
- package/src/interactive/chat-loop.ts +65 -2
- package/src/interactive/chat-renderer.ts +1 -0
- package/src/interactive/cost-overlay.ts +26 -2
- package/src/interactive/interactive-slash-runtime.ts +2 -1
- package/src/interactive/renderers/worker-entry.ts +32 -0
- package/src/interactive/slash-commands.ts +24 -6
- package/src/interactive/theme/labels.ts +19 -13
- package/src/interactive/turn-context.ts +9 -5
- package/src/interactive/turn-recovery.ts +8 -0
- package/src/interactive/turn-runtime.ts +27 -11
- package/src/interactive/turn-state.ts +7 -0
- package/src/interactive/worker-receipts.ts +1 -0
- package/src/interactive/worker-stream.ts +6 -1
- package/src/tools/context/index.ts +30 -9
- package/src/tools/dispatch-arguments.ts +1 -0
- package/src/tools/dispatch-event-text.ts +10 -0
- package/src/tools/dispatch-plan.ts +1 -0
- package/src/tools/dispatch-runner.ts +12 -0
- package/src/tools/registry.ts +11 -5
- package/src/tools/worker-evidence.ts +3 -1
- package/src/worker/spec-contract.ts +4 -0
- package/dist/chunk-2Z2IKEXI.js +0 -1554
- package/dist/reset-OAQP3W4O.js +0 -230
- package/dist/uninstall-N34PCTGJ.js +0 -331
- package/dist/upgrade-PXK3S2YM.js +0 -325
|
@@ -121,7 +121,7 @@ malformed response, or telemetry failure is silent and never blocks a tool.
|
|
|
121
121
|
- `context.memory.cadenceToolCalls` (default `10`): Minimum completed-tool interval between background interventions.
|
|
122
122
|
- `context.memory.trajectorySteps` (default `8`): Completed tool-trajectory window analyzed during background evaluation.
|
|
123
123
|
- `context.memory.maxOutputTokens` (default `2000`): Bounds the rendered memory-bank context and the ordinary policy-model completion. An always-on-thinking model receives additional reasoning headroom, at least `4,000` tokens when its model cap permits, so it can still reach the strict envelope.
|
|
124
|
-
- `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy
|
|
124
|
+
- `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy step, including its one permitted chat fallback attempt. The step is detached, so this deadline never delays a turn, but it does hold a request slot on a real inference endpoint that your own turns and dispatched workers queue against. Raise it only after inspecting the timeout and hit-rate evidence for the selected route.
|
|
125
125
|
|
|
126
126
|
## Trigger semantics
|
|
127
127
|
|
|
@@ -254,38 +254,48 @@ unexamined one:
|
|
|
254
254
|
would not have been spent. The current 60-second deadline is the source-backed
|
|
255
255
|
bound on what one optional call may hold a shared local server for, not a
|
|
256
256
|
prediction of when a route answers.
|
|
257
|
-
- A step
|
|
258
|
-
skipped with reason `endpoint_busy
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
257
|
+
- A step whose known endpoint occupancy exhausts the resolved request capacity is
|
|
258
|
+
skipped with reason `endpoint_busy`. The same gateway URL alone does not imply
|
|
259
|
+
a single request slot or a single physical model server.
|
|
260
|
+
|
|
261
|
+
### Dedicated routing, chat fallback and endpoint capacity
|
|
262
|
+
|
|
263
|
+
Clio prefers the explicitly configured memory target and model. If that route is
|
|
264
|
+
known unavailable (missing target/runtime, down target, absent model in a known
|
|
265
|
+
catalog, or a model reported unloaded/loading), Clio selects the active chat
|
|
266
|
+
route for that step. A runtime client error on the dedicated route permits one
|
|
267
|
+
chat attempt within the original remaining deadline. It does not retry after a
|
|
268
|
+
timeout, cancellation, session/branch switch, malformed envelope, or a model's
|
|
269
|
+
explicit silence. The fallback never falls back again. Both routes unavailable
|
|
270
|
+
produce a visible `client_error` outcome. An unset memory role remains rules-only;
|
|
271
|
+
chat fallback does not enable model-based memory by default.
|
|
272
|
+
|
|
273
|
+
A runtime notice names the selected chat fallback and its reason. Routing stays
|
|
274
|
+
session-local and does not edit saved settings. Each attempted call records its
|
|
275
|
+
own known usage under its actual target/model and emits its own telemetry outcome;
|
|
276
|
+
failed dedicated usage is retained alongside fallback usage. The final policy
|
|
277
|
+
result describes the final attempt, without merging different provider identities.
|
|
278
|
+
A generation change discards late content while preserving the original usage owner.
|
|
279
|
+
|
|
280
|
+
Capacity follows the same existing evidence as dispatch: explicit target
|
|
281
|
+
`maxConcurrentRequests`, current discovery, a valid discovery prior, then the
|
|
282
|
+
conservative one-slot default for a local-native runtime. Known occupancy includes
|
|
283
|
+
this process's foreground requests, active dispatch leases and held reservation
|
|
284
|
+
waves. A two-slot endpoint with one foreground request can admit memory; a full
|
|
285
|
+
one-slot endpoint skips it. Known dedicated saturation does not itself initiate
|
|
286
|
+
chat fallback. Larger declared capacities work without a new constant.
|
|
287
|
+
|
|
288
|
+
LiteLLM is a gateway protocol, so Clio does not invent a local one-slot limit for
|
|
289
|
+
its URL. Distinct model routes such as dynamo, mini and zbook can share that URL;
|
|
290
|
+
the gateway owns their physical routing and backend residency. An explicit or
|
|
291
|
+
observed endpoint-wide bound still applies when present. Clio's process-local
|
|
292
|
+
foreground holds and observed dispatch state are not a global scheduler for every
|
|
293
|
+
client using the gateway. Slot holds are released on success, failure and abort.
|
|
294
|
+
|
|
295
|
+
The existing `background_memory` expected-cold stamp remains keyed by endpoint
|
|
296
|
+
URL. It records a possible shared-endpoint cache disturbance, including after a
|
|
297
|
+
failed request, rather than proving that a different model behind the gateway
|
|
298
|
+
actually evicted the chat prefix.
|
|
289
299
|
|
|
290
300
|
## Choosing a background model
|
|
291
301
|
|
|
@@ -294,9 +304,11 @@ does not need to be clever. A small non-reasoning model is the right choice, and
|
|
|
294
304
|
Clio always requests the memory route with thinking off. Version 2 therefore
|
|
295
305
|
has no configurable memory thinking-level key.
|
|
296
306
|
|
|
297
|
-
|
|
307
|
+
The off request depends on the resolved runtime and model metadata:
|
|
298
308
|
llama.cpp reads `chat_template_kwargs.enable_thinking`, and LM Studio reads
|
|
299
|
-
`reasoning_effort`,
|
|
309
|
+
`reasoning_effort`, with the off value selected by the model family. A gateway
|
|
310
|
+
needs a recognized, unanimous upstream runtime declaration; unknown metadata
|
|
311
|
+
does not establish a dialect. Requesting off does not prove server compliance.
|
|
300
312
|
|
|
301
313
|
A model that reasons anyway still works. Some genuinely cannot be silenced, and
|
|
302
314
|
the catalog records those as always-on so the level reads `forced` rather than
|
|
@@ -306,12 +318,11 @@ blocks are discarded and only the envelope is kept, and the output budget is
|
|
|
306
318
|
sized to let a reasoning preamble run its course first. The cost is latency,
|
|
307
319
|
which the detached step absorbs.
|
|
308
320
|
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
actually asked for.
|
|
321
|
+
The active chat model may also serve memory, including a reasoning model, when
|
|
322
|
+
request capacity permits. Memory still requests thinking off and uses its own
|
|
323
|
+
bounded envelope and output budget. A dedicated small model is preferred for
|
|
324
|
+
latency and resource use; sharing a model is a capacity decision rather than an
|
|
325
|
+
unconditional refusal.
|
|
315
326
|
|
|
316
327
|
This is a mix-and-match plane, not a local-only one. The background role resolves
|
|
317
328
|
through the same target machinery as every other role, so the useful shapes are:
|
|
@@ -356,30 +367,30 @@ independent of the chat target and the fleet default. A running session owns
|
|
|
356
367
|
its routing snapshot, while the saved
|
|
357
368
|
selection becomes the default for new sessions.
|
|
358
369
|
|
|
359
|
-
|
|
370
|
+
An example gateway topology keeps the three backend routes distinct:
|
|
360
371
|
|
|
361
372
|
| Role | Target | Runtime and endpoint | Model | Capacity |
|
|
362
373
|
| --- | --- | --- | --- | --- |
|
|
363
|
-
| Chat | `dynamo` |
|
|
364
|
-
|
|
|
374
|
+
| Chat and memory fallback | `dynamo` | LiteLLM at `http://gateway:4000` | `dynamo/qwen3.8-27b` | Gateway-owned unless an explicit or observed endpoint bound is available |
|
|
375
|
+
| Preferred background memory | `zbook` | Same LiteLLM gateway | `zbook/ornith-1.5-35b-a3b` | Same capacity evidence rules |
|
|
376
|
+
| Alternative memory candidate | `mini` | Same LiteLLM gateway | `mini/ornith1.5-35b-moe` | Same capacity evidence rules |
|
|
365
377
|
|
|
366
|
-
The corresponding
|
|
378
|
+
The corresponding role selection is:
|
|
367
379
|
|
|
368
380
|
```yaml
|
|
369
381
|
chat:
|
|
370
382
|
target: dynamo
|
|
371
|
-
model: qwen3.8-27b
|
|
383
|
+
model: dynamo/qwen3.8-27b
|
|
372
384
|
context:
|
|
373
385
|
memory:
|
|
374
|
-
target:
|
|
375
|
-
model:
|
|
386
|
+
target: zbook
|
|
387
|
+
model: zbook/ornith-1.5-35b-a3b
|
|
376
388
|
```
|
|
377
389
|
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
effect on the operator's turn.
|
|
390
|
+
These are example configured target/model IDs, not built-in routes. Compare a
|
|
391
|
+
slow dedicated model with an alternative using actual step latency, valid bank
|
|
392
|
+
writes and known usage; the active chat route remains the fallback. Do not create
|
|
393
|
+
aliases or change endpoint URLs just to bypass capacity accounting.
|
|
383
394
|
|
|
384
395
|
The deadline is a bound on what an optional call may hold that server for, not a
|
|
385
396
|
figure sized to capture the tail. The shipped `60000` is a bounded compromise,
|
|
@@ -489,7 +500,7 @@ clear the prior heap bank before the new session can observe it.
|
|
|
489
500
|
|
|
490
501
|
## Telemetry
|
|
491
502
|
|
|
492
|
-
Each completed memory
|
|
503
|
+
Each completed memory-policy attempt appends one content-free record to:
|
|
493
504
|
|
|
494
505
|
```text
|
|
495
506
|
<stateDir>/memory/steps.jsonl
|
|
@@ -512,7 +523,7 @@ folds out of this file.
|
|
|
512
523
|
`dropped` is the one outcome that ran no step. It has two causes, separated by
|
|
513
524
|
the row's reason: `step_in_flight` means the boundary triggered while an earlier
|
|
514
525
|
step still held the single in-flight slot, and `endpoint_busy` means the step
|
|
515
|
-
|
|
526
|
+
found the known endpoint request capacity exhausted. Both cost
|
|
516
527
|
no tokens and no latency, both leave their triggers pending for the next free
|
|
517
528
|
boundary, and neither replaces the operator-visible last decision. Counting
|
|
518
529
|
`dropped` rows against `llm` rows over a session is how a starved cadence becomes
|
|
@@ -5,6 +5,14 @@
|
|
|
5
5
|
|
|
6
6
|
The Skills Hub (`/skill`) shows project skills, user skills, and the marketplace. Every marketplace row comes from the same local lookup that `clio-coder skills install <name>` and `/skill <name>` resolve through, so the hub lists nothing it cannot install.
|
|
7
7
|
|
|
8
|
+
## Operator ownership of installed skills
|
|
9
|
+
|
|
10
|
+
The active project `.clio-coder/skills/` tree and the resolved user `<configDir>/skills/` tree are operator-owned. Main-agent and worker tool admissions refuse writes, edits, artifacts and recognized shell mutations to either tree, including ancestor deletion and current symlink aliases. This boundary stays active at every autonomy level, even when the general default path policy is disabled. Reads and loading already installed skills retain their existing rules.
|
|
11
|
+
|
|
12
|
+
Draft a proposed skill outside these active trees, for example in `draft-skills/`. The operator installs or updates it through `clio-coder skills install`, `skills update`, `skills sync`, `library add --yes`, the Skills Hub, `/skill <name>`, or an explicitly accepted marketplace offer. Even `full-auto` must wait for that offer's bound operator answer; a task match alone no longer installs a skill. Model shell calls to recognizable Clio skill installation, update and sync commands are refused. Confirmed `library add --yes` commands are also reserved for operators, including additions of other resource kinds that may install skill dependencies. An unconfirmed `library add` only prints a plan and remains available, as do inventory, search, inspection and validation commands.
|
|
13
|
+
|
|
14
|
+
This is a tool-admission boundary, not an operating-system sandbox. Shell inspection covers literal paths and recognized command forms, respecting comments, quoted words and redirection operands when identifying Clio commands. Quoted search patterns remain read-only, and literal operator filenames do not hide later mutation operands. It cannot prove the effects of arbitrary scripts, dynamically constructed paths, aliases, or concurrent filesystem changes. Keep shell execution supervised when stronger confinement is required.
|
|
15
|
+
|
|
8
16
|
## Where marketplace rows come from
|
|
9
17
|
|
|
10
18
|
There is one source, `discoverMarketplaceSkills()` in `src/domains/resources/skills/marketplace.ts`, and it reads two kinds of real local data:
|
|
@@ -62,6 +70,16 @@ creation. Extension resource roots and share archives are documented in
|
|
|
62
70
|
[extensions-and-sharing.md](extensions-and-sharing.md); this page owns the TUI
|
|
63
71
|
Hub and marketplace behavior.
|
|
64
72
|
|
|
73
|
+
## Remote entries and overlays
|
|
74
|
+
|
|
75
|
+
Some skills are worth carrying in the catalog without vendoring their content. `skills/remote.yaml` lists them: each entry names a skill, its category, a `sourceUrl` that must be a GitHub tree URL at a pinned tag, an `overlay` package inside the catalog, and an optional `exclude` list of upstream top-level members. `npm run skills:pin` publishes such an entry into `skills/skill-marketplace.json` with `origin: "remote"`, the upstream URL as its `sourceUrl`, and the `overlay` and `exclude` fields attached. The overlay's `SKILL.md` is pinned in `skills/registry.yaml` like every other catalog skill, so `npm run skills:check` fails when it drifts.
|
|
76
|
+
|
|
77
|
+
`archify` is the worked example. Its entry points at `https://github.com/tt-a1i/archify/tree/v2.16.0/archify`, overlays `skills/planning/archify`, and excludes `test` and `package-lock.json`. Running `clio-coder skills install archify --project` clones that tag, drops the excluded members, copies the overlay over the clone so Clio's wrapper `SKILL.md` replaces the upstream one, validates the shaped tree, and swaps it into `.clio-coder/skills/archify/` with the usual provenance stamps. The renderer, its schemas, and its brand-mark notices come from upstream at install time and never enter the npm tarball. The wrapper omits upstream's update-awareness step on purpose: that step performs a network request during a chat turn, and no Clio chat turn depends on the network.
|
|
78
|
+
|
|
79
|
+
Discovery treats the overlay folder as part of the remote entry rather than as a skill of its own, so a bare-name install never lands the wrapper without the renderer it wraps. `clio-coder skills update` recovers the same overlay and exclude list from the marketplace, so an update refetches the pinned upstream and re-applies the wrapper instead of replacing it.
|
|
80
|
+
|
|
81
|
+
Two operational notes. A copy dropped into a project-scope `.claude/skills/archify` is a compat import and stays untrusted until `integrations.projectResources.trustProjectImports` is on; install through Clio to get a trusted, provenance-stamped copy. And since the skill's authoring loop is several `bash` calls (validate, deliver, verify), `auto-edit` is the sensible permission posture for a mapping session; full manual approval works but prompts on every command.
|
|
82
|
+
|
|
65
83
|
## Publishing a skill
|
|
66
84
|
|
|
67
85
|
Add a directory under `skills/<category>/<name>/` (or `skills/<name>/`) in the repo containing a `SKILL.md` with `name` and `description` frontmatter. The directory name must match `[A-Za-z0-9][A-Za-z0-9._-]*`. Run `npm run skills:pin` to republish `skills/skill-marketplace.json`, which is the index consumers point `CLIO_CODER_SKILL_MARKETPLACE_INDEX` at or copy to `<configDir>/skill-marketplace.json`. Scientific and niche coding domains are the marketplace's focus; see the existing `skills/` tree for the house format.
|
|
@@ -20,12 +20,13 @@ under pressure; git mechanics alone never justify a stage.
|
|
|
20
20
|
| --- | --- | --- |
|
|
21
21
|
| 1. File | [`file-ticket`](../../skills/git/file-ticket/) | A labeled GitHub issue with evidence and acceptance criteria |
|
|
22
22
|
| 2. Fix | [`fix-issue`](../../skills/git/fix-issue/) | An uncommitted, verified change where failing tests preceded the fix, self-reviewed against the issue's acceptance criteria |
|
|
23
|
-
| 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`)
|
|
23
|
+
| 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`); contributors push it to their fork and open a PR, while maintainer work stays local for gated integration; merge is a human decision |
|
|
24
24
|
|
|
25
25
|
Releases follow [release-cut-checklist.md](../history/release-cut-checklist.md) as a
|
|
26
26
|
human-gated checklist, not a skill. Worktrees
|
|
27
27
|
([`worktree-create`](../../skills/git/worktree-create/),
|
|
28
28
|
[`worktree-merge`](../../skills/git/worktree-merge/)),
|
|
29
|
+
[`branch-closeout`](../../skills/git/branch-closeout/),
|
|
29
30
|
[`resolve-merge-conflicts`](../../skills/git/resolve-merge-conflicts/), and
|
|
30
31
|
[`tdd`](../../skills/coding/tdd/) are à-la-carte tools reached for when the
|
|
31
32
|
situation calls for them, not stages every change passes through. An RCA
|
|
@@ -34,6 +35,38 @@ hard bugs, not a mandatory toll booth. Batch ticket creation from a PRD
|
|
|
34
35
|
bypasses stage 1 and uses [`backlog`](../../skills/planning/backlog/)
|
|
35
36
|
instead; everything downstream is identical.
|
|
36
37
|
|
|
38
|
+
## Closeout
|
|
39
|
+
|
|
40
|
+
A merged PR is not operationally finished until its local scaffolding is
|
|
41
|
+
closed. The reusable [`branch-closeout`](../../skills/git/branch-closeout/) skill automates this verification and teardown safely. After the human merge decision:
|
|
42
|
+
|
|
43
|
+
1. Fetch and prune, confirm the PR's merged state, and identify the resulting
|
|
44
|
+
commit on `origin/main`. Direct ancestry proves an ordinary merge; a squash
|
|
45
|
+
or cherry-pick needs the PR-to-result evidence because commit identity and
|
|
46
|
+
patch identity can both change during integration.
|
|
47
|
+
2. Inspect every associated worktree for tracked changes, untracked files, and
|
|
48
|
+
ignored state that carries evidence rather than rebuildable output. Remove
|
|
49
|
+
it through `git worktree remove`; forcing removal requires explicit approval
|
|
50
|
+
to discard what remains.
|
|
51
|
+
3. Delete the local source and integration branches. For a contributor PR,
|
|
52
|
+
delete the merged branch from the contributor's fork. The canonical
|
|
53
|
+
repository never hosts topic, integration, or release-candidate branches.
|
|
54
|
+
4. Turn unfinished experimental findings into an issue with evidence and a
|
|
55
|
+
next decision. Do not use indefinite `work/`, `wip/`, `keep/`, or temporary
|
|
56
|
+
tags as a substitute for backlog state.
|
|
57
|
+
5. Report the remaining worktrees, local branches, stashes, local-only tags,
|
|
58
|
+
and canonical remote heads. The expected canonical head set is exactly
|
|
59
|
+
`refs/heads/main`; every survivor needs an owner and purpose.
|
|
60
|
+
|
|
61
|
+
Maintainer release candidates are local-only and use a compact branch name
|
|
62
|
+
that cannot collide with their tag: branch `v043`, tag `v0.4.3`. Gate the exact
|
|
63
|
+
candidate, require fetched `origin/main` to be its ancestor, fast-forward local
|
|
64
|
+
`main`, fetch again, and push only `refs/heads/main:refs/heads/main` with
|
|
65
|
+
explicit authorization. After CI passes, push only the fully qualified
|
|
66
|
+
annotated tag. Once the tag's peeled commit equals the reviewed commit on
|
|
67
|
+
`main` and the release succeeds, delete the local candidate branch. Published
|
|
68
|
+
dotted release tags are immutable history and are never cleanup targets.
|
|
69
|
+
|
|
37
70
|
## Inheriting a Pi release
|
|
38
71
|
|
|
39
72
|
Pi dependency upgrades use a fixed five-step review so that upstream fixes
|
|
@@ -139,7 +139,37 @@ Metrics collected during runs can be validated automatically using the `verify.a
|
|
|
139
139
|
* `eq` (equal)
|
|
140
140
|
* `neq` (not equal)
|
|
141
141
|
|
|
142
|
-
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `
|
|
142
|
+
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, `result.pass`, and the `provider.*` metrics below. Each `verify.assertions` condition must hold; an unavailable metric fails closed.
|
|
143
|
+
|
|
144
|
+
### Optional Provider-Health Gates
|
|
145
|
+
|
|
146
|
+
A task that recovers from a provider error can still pass its task checks by default. Provider health is a separate, opt-in requirement. To require observed provider events and no observed terminal errors, add these assertions to the task:
|
|
147
|
+
|
|
148
|
+
```yaml
|
|
149
|
+
verify:
|
|
150
|
+
assertions:
|
|
151
|
+
- metric: "provider.measured"
|
|
152
|
+
op: "eq"
|
|
153
|
+
value: true
|
|
154
|
+
- metric: "provider.stopReason.error"
|
|
155
|
+
op: "eq"
|
|
156
|
+
value: 0
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
Suite-level `thresholds.fail` uses the opposite condition: a matching condition is a failure. The corresponding hard gate checks each run as follows:
|
|
160
|
+
|
|
161
|
+
```yaml
|
|
162
|
+
thresholds:
|
|
163
|
+
fail:
|
|
164
|
+
- metric: "provider.measured"
|
|
165
|
+
op: "eq"
|
|
166
|
+
value: false
|
|
167
|
+
- metric: "provider.stopReason.error"
|
|
168
|
+
op: "gt"
|
|
169
|
+
value: 0
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
Unavailable metrics also fail closed in a hard threshold. `thresholds.informational` records findings without changing exit status. These examples reject an observed error followed by a successful recovery while leaving the default ungated task-pass policy unchanged. To reject any observed retry start or assistant abort as well, add conditions on `provider.retryStarted` or `provider.stopReason.aborted` with the same assertion-versus-failure polarity.
|
|
143
173
|
|
|
144
174
|
---
|
|
145
175
|
|
|
@@ -175,14 +205,44 @@ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
|
|
|
175
205
|
|
|
176
206
|
Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
|
|
177
207
|
|
|
178
|
-
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events
|
|
208
|
+
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events. These totals include all known usage on errored calls as well as successful calls; recovery never subtracts earlier spend. Only finite, nonnegative usage facts are admitted. On surfaces without the relevant stdout events (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
|
|
179
209
|
2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
|
|
180
210
|
|
|
181
211
|
### Fail-Closed Reporting
|
|
182
212
|
Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
|
|
183
213
|
|
|
214
|
+
An errored call must carry at least one positive reported token, reasoning, or cost fact before its usage is considered observed. Reasoning-only or cost-only observations do not establish ordinary token totals. A stream containing only errored calls with missing or synthetic all-zero usage remains `tokens.measured: false`. The current event shape cannot distinguish synthetic all-zero failures from genuinely reported zero usage, so it cannot establish measured zero spending in either case. Partial positive usage remains included as known spend. On failed calls, adapters can also initialize individual absent fields to zero; those zeros remain unattributed and make coverage incomplete. The existing inclusive numeric fields are known subtotals, so their zeros do not prove complete zero spending when failed usage is incomplete.
|
|
215
|
+
|
|
216
|
+
### Provider Observations and Failed-Call Share
|
|
217
|
+
|
|
218
|
+
The `provider.*` metrics describe events observed on live stdout, folded before diagnostic output is truncated. Native runs and multi-command external runners retain these observations from their executed commands. They do not enumerate SDK-internal retries or network attempts that were never emitted. Filtered or opaque output can leave provider health unobserved even when the process exits successfully or a receipt reports task success.
|
|
219
|
+
|
|
220
|
+
| Metric | Meaning |
|
|
221
|
+
| --- | --- |
|
|
222
|
+
| `provider.measured` | Whether an assistant terminal reason or a counted retry phase was observed. With no such observations this is `false`, and provider counters are absent. It does not certify complete provider coverage. |
|
|
223
|
+
| `provider.stopReason.stop`, `provider.stopReason.toolUse`, `provider.stopReason.length` | Counts of these terminal reasons on assistant `message_end` events. |
|
|
224
|
+
| `provider.stopReason.error`, `provider.stopReason.aborted`, `provider.stopReason.other` | Separate counts of errored, aborted, and other observed terminal reasons. An unrecognized terminal reason goes into `other`. Partial updates and repeated messages in `turn_end` or `agent_end` do not add counts. |
|
|
225
|
+
| `provider.retryScheduled` | Observed `scheduled` phases: planned retries, including ones cancelled before execution. |
|
|
226
|
+
| `provider.retryStarted` | Observed `retrying` phases: retry execution starts. Repeated `waiting` countdown frames do not count as attempts. |
|
|
227
|
+
| `provider.retryCancelled`, `provider.retryExhausted`, `provider.retryRecovered` | Counts of the corresponding observed phases. They describe retry-chain outcomes and do not fabricate additional assistant calls. Attempt numbers can restart for each chain. |
|
|
228
|
+
| `provider.errorUsageObservedCalls` | Errored calls with at least one positive reported token, reasoning, or cost fact. |
|
|
229
|
+
| `provider.errorUsageUnobservedCalls` | Errored calls with no positive reported token, reasoning, or cost fact, including absent or all-zero usage. |
|
|
230
|
+
| `provider.errorUsageIncompleteCalls` | Errored calls with unobserved usage, incomplete token fields, or ambiguous normalized zero fields. This can overlap `errorUsageObservedCalls` when only part of the usage is known. |
|
|
231
|
+
| `provider.errorCostUnobservedCalls` | Errored calls without a positive cost fact. Zero or absent cost does not prove that an error was free, including when some token usage is known. |
|
|
232
|
+
| `provider.errorTokens.input`, `provider.errorTokens.output`, `provider.errorTokens.total`, `provider.errorTokens.cacheRead`, `provider.errorTokens.cacheWrite` | Known positive token subtotals for errored calls. Absent or ambiguous zero fields remain unattributed; when total usage is absent, a total can still be summed from known token fields and remains incomplete. |
|
|
233
|
+
| `provider.errorCostUsd` | Known positive cost subtotal from errored calls' stream usage objects. Cost can come from adapter pricing; it is not independently certified provider billing. |
|
|
234
|
+
| `provider.errorReasoningTokens`, `provider.errorReasoningUnobservedCalls` | Known failed-call reasoning subtotal and calls without attributable reported reasoning. Reasoning can overlap output, so it is never added to ordinary token totals. |
|
|
235
|
+
|
|
236
|
+
The failed-call share covers `stopReason: error`; aborted calls remain separately labeled. Failed-share amounts and usage-coverage counters appear only after an errored terminal message is observed. These share metrics supplement the inclusive `tokens.*` totals without changing the summary token shape, receipt schema, or verdict schema. Missing or partial failed usage makes them known subtotals, not a complete amount to subtract from total spend. The native runner's `cost.usd` can use receipt evidence, so equality with the stream-based `provider.errorCostUsd` is not guaranteed.
|
|
237
|
+
|
|
238
|
+
Positive reported reasoning uses the normalized `usage.reasoning` field, `reasoning_tokens`, and supported nested provider-detail fields. The root `reasoningTokens` alias can be an adapter estimate without a provenance marker, so it is left unattributed rather than promoted to reported provider usage. Normalized zero reasoning is also unattributed because adapters can fill it when provider detail is absent. A reasoning-only failure is observed but has incomplete ordinary-token coverage; no output or total is inferred from it.
|
|
239
|
+
|
|
240
|
+
These observations also do not reconcile the separate `trackedMetrics` ledger selection (#276). Tracked metrics prefer durable assistant-call facts when available, retain durable compaction and tool records, and otherwise fall back to stream calls. Artifacts expose source counts and warnings, but partial or mixed ledgers can omit stream-only calls and fork-inherited history remains unreconciled. Neither those tracked values nor the new provider counters prove complete run accounting; failed-compaction usage is retained separately in the out-of-turn usage ledger and usage report, and is not included by this eval fold.
|
|
241
|
+
|
|
184
242
|
---
|
|
185
243
|
|
|
244
|
+
Full reconciliation across session, stdout, fork and out-of-turn evidence is deferred to v0.4.5 or later. Version 0.4.3 does not add an automatic rejection of cost or efficiency comparisons merely because those sources are partial or mixed. Matching source counts do not prove complete coverage or shared call identity. Existing missing-metric, serving-configuration and execution-envelope comparison gates still apply.
|
|
245
|
+
|
|
186
246
|
## Eval Artifact Format (v4)
|
|
187
247
|
|
|
188
248
|
Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
|
|
@@ -297,7 +357,7 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
297
357
|
| `generatedTokens` | ledger |
|
|
298
358
|
| `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
|
|
299
359
|
| `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
|
|
300
|
-
| `ttftMsFirstCall` | ledger |
|
|
360
|
+
| `ttftMsFirstCall` | ledger; nullable when first-call timing is absent |
|
|
301
361
|
| `wallClockMs` | receipt |
|
|
302
362
|
| `contextTokensAtEnd` | ledger |
|
|
303
363
|
| `compactions` | ledger |
|
|
@@ -305,6 +365,10 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
305
365
|
|
|
306
366
|
A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
|
|
307
367
|
|
|
368
|
+
First-call TTFT uses the earliest recorded assistant-call timestamp across the selected ledgers; equal timestamps retain their observed order. Missing or invalid timing, or an invalid timestamp that prevents ordering the calls, yields `{ value: null, source: "estimated" }`. A measured zero remains `{ value: 0, source: "ledger" }`. Native session timing starts at each stream invocation and includes the provider's response-header wait. Stdout-only fallback timing starts at the provider's `message_start`, which can arrive after headers; it requires first output, but is not complete request latency and must not be compared as equivalent to native invocation timing. A completion alone supplies neither a zero duration nor a first-token measurement. Verdict v1 consumers must accept nullable TTFT. Historical numeric values, including estimated zeros and native spans that omitted the pre-header wait, remain readable and are not rewritten.
|
|
369
|
+
|
|
370
|
+
When a stream message lacks a valid timestamp, its ledger payload marks `timestampEstimated: true` beside the legacy ISO placeholder. This leaves first-call chronology unmeasured while preserving any observed per-call monotonic timing.
|
|
371
|
+
|
|
308
372
|
### Scenario aggregates
|
|
309
373
|
|
|
310
374
|
`aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
|
|
@@ -67,6 +67,7 @@ tasks:
|
|
|
67
67
|
metrics:
|
|
68
68
|
collect: &metrics
|
|
69
69
|
[task.solved, claims.unsupported, completion.reported, tools.calls.read, tools.calls.dispatch,
|
|
70
|
+
tools.succeeded.dispatch,
|
|
70
71
|
tools.blocked.bash, tools.read.distinctPaths, tools.read.outsideAllowed, tools.read.decoyHits]
|
|
71
72
|
readObservation:
|
|
72
73
|
allowedPaths: [evals/fixtures/behavioral-main.ts]
|
|
@@ -137,10 +138,10 @@ tasks:
|
|
|
137
138
|
expectedBehavior:
|
|
138
139
|
- id: delegation-chooses-dispatch
|
|
139
140
|
category: tool_choice
|
|
140
|
-
fact: {source: tool, key: tools.
|
|
141
|
+
fact: {source: tool, key: tools.succeeded.dispatch, op: gte, value: 1}
|
|
141
142
|
- id: delegation-occurs
|
|
142
143
|
category: delegation
|
|
143
|
-
fact: {source: tool, key: tools.
|
|
144
|
+
fact: {source: tool, key: tools.succeeded.dispatch, op: gte, value: 1}
|
|
144
145
|
- id: delegation-grounded
|
|
145
146
|
category: claim_grounding
|
|
146
147
|
fact: {source: grader, key: claims.unsupported, op: eq, value: 0}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@iowarp/clio-coder",
|
|
3
|
-
"version": "0.4.
|
|
3
|
+
"version": "0.4.3",
|
|
4
4
|
"description": "The terminal coding agent for people who maintain the code that science runs on, with user-chosen models and inspectable evidence.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"ai",
|
|
@@ -80,7 +80,7 @@
|
|
|
80
80
|
"test": "node --import tsx --import ./tests/harness/tmp-root.ts --test tests/contracts/*.test.ts tests/smoke/*.test.ts",
|
|
81
81
|
"test:trace-viewer": "npm --prefix apps/trace-viewer test",
|
|
82
82
|
"trace:ui": "node apps/trace-viewer/server.mjs",
|
|
83
|
-
"ci": "npm run typecheck && npm run lint && npm run
|
|
83
|
+
"ci": "npm run typecheck && npm run lint && npm run build && npm run test && npm run test:trace-viewer",
|
|
84
84
|
"ci:release": "npm run ci && node scripts/check-release.mjs",
|
|
85
85
|
"install:local": "bash scripts/install-local.sh",
|
|
86
86
|
"smoke:real-home": "bash scripts/smoke-real-home.sh",
|
package/skills/README.md
CHANGED
|
@@ -63,6 +63,7 @@ folder is presentation and provenance, not a namespace.
|
|
|
63
63
|
| [`worktree-create`](git/worktree-create/) | workflow | Stand up isolated worktrees for parallel branches: detected install/config/health-check, per-worktree verification. |
|
|
64
64
|
| [`worktree-merge`](git/worktree-merge/) | workflow | Integrate finished worktree branches through a throwaway integration branch with per-merge tests and a full final gate. |
|
|
65
65
|
| [`resolve-merge-conflicts`](git/resolve-merge-conflicts/) | workflow | A merge/rebase is stopped on conflicts. Resolves from both sides' reconstructed intent, validates, completes the operation. |
|
|
66
|
+
| [`branch-closeout`](git/branch-closeout/) | workflow | Proves merged work on the canonical base, inspects and removes associated worktrees through Git, deletes local branches safely, and audits surviving repository refs. |
|
|
66
67
|
|
|
67
68
|
### `research/` — scientific and literature work
|
|
68
69
|
|
|
@@ -101,11 +102,12 @@ folder is presentation and provenance, not a namespace.
|
|
|
101
102
|
| [`herdr`](meta/herdr/) | integration | The user asks to launch, drive, or inspect another agent or command in a Herdr pane — including a second Clio Coder instance. Requires `HERDR_ENV=1`. |
|
|
102
103
|
|
|
103
104
|
Each SKILL.md may declare `allowed-tools` / `disallowed-tools`. After a skill
|
|
104
|
-
loads, Clio enforces that declaration at tool admission
|
|
105
|
-
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
105
|
+
loads, Clio enforces that declaration at tool admission for the rest of the
|
|
106
|
+
session, across the operator's later turns, until another skill replaces it or
|
|
107
|
+
the operator runs `/skill off` (a worker run keeps the run-scoped lifetime):
|
|
108
|
+
calls outside the merged surface are blocked with reason code `skill_surface`,
|
|
109
|
+
with `context` and `ask_user` always admitted. A skill can narrow its tool
|
|
110
|
+
surface but never grant tools the host would not allow. See [Skill tool surface narrowing](../docs/architecture/safety-model.md#skill-tool-surface-narrowing)
|
|
109
111
|
for the full semantics.
|
|
110
112
|
|
|
111
113
|
## Install (activate a marketplace skill)
|