@iowarp/clio-coder 0.4.2 → 0.4.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +98 -0
- package/CONTRIBUTING.md +86 -19
- package/README.md +35 -6
- package/dist/{acp-TMDQZDIG.js → acp-WNAYYF4F.js} +12 -13
- package/dist/{agents-5N5NG3XG.js → agents-3OKXHLOI.js} +60 -57
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-Z5CCBXKQ.js → auth-VKNNMGPU.js} +21 -19
- package/dist/{builtins-K6TNDT24.js → builtins-WGALA46I.js} +9 -4
- package/dist/{chunk-XE3PCIXH.js → chunk-23L32XTI.js} +12 -9
- package/dist/{chunk-I64IFBLB.js → chunk-25QBEXRS.js} +18 -11
- package/dist/{chunk-CDNVLKUX.js → chunk-26QSH3EJ.js} +13 -7
- package/dist/{chunk-QQLGQY2A.js → chunk-2ASED4PZ.js} +22 -22
- package/dist/{chunk-MCEPRMZW.js → chunk-2CU2H6KE.js} +2 -2
- package/dist/chunk-2DSOYNFC.js +108 -0
- package/dist/{chunk-72GZI5EV.js → chunk-2JH2WHGE.js} +2 -2
- package/dist/{chunk-I66EAJFY.js → chunk-2WZ546HR.js} +267 -232
- package/dist/{chunk-O3YUNJZ2.js → chunk-2ZSONWVL.js} +82 -25
- package/dist/{chunk-2NHR3NAY.js → chunk-36CT5VVL.js} +331 -42
- package/dist/{chunk-2X4RYJTJ.js → chunk-3GY4F45V.js} +3 -3
- package/dist/{chunk-ZW55JB7N.js → chunk-3ODX73FK.js} +4 -6
- package/dist/{chunk-PBP4B7XR.js → chunk-3UNOLWNZ.js} +3 -3
- package/dist/{chunk-4JDLP6ZS.js → chunk-3UUXNFEX.js} +14 -10
- package/dist/{chunk-RKSR6VSF.js → chunk-4IUZQIJ3.js} +29 -1
- package/dist/{chunk-ZW4HH5JJ.js → chunk-4M6Z5QVF.js} +6 -6
- package/dist/{chunk-K6BSR66V.js → chunk-4NSRCOYP.js} +4 -1
- package/dist/{chunk-M2DAX4F6.js → chunk-4WR7VSYB.js} +2 -2
- package/dist/{chunk-FSP7CMNU.js → chunk-54X7T7DK.js} +61 -6
- package/dist/{chunk-54ODD65L.js → chunk-5636DCO5.js} +4 -4
- package/dist/chunk-57XXR6DR.js +3763 -0
- package/dist/{chunk-3KIPBMUA.js → chunk-5ICU3EUH.js} +2 -2
- package/dist/{chunk-77QIVUZB.js → chunk-5MEZN6CB.js} +4 -4
- package/dist/{chunk-O42A54GG.js → chunk-5OIVVPHF.js} +2 -2
- package/dist/{chunk-YJISEZKC.js → chunk-5TUB6SLS.js} +6 -6
- package/dist/{chunk-IMXMHHMQ.js → chunk-6OSVSQL5.js} +341 -57
- package/dist/{chunk-Q4XWMHX6.js → chunk-6PAZTBPA.js} +14 -2
- package/dist/{chunk-FVDGR2ZL.js → chunk-6Q3CYFD3.js} +112 -39
- package/dist/{chunk-IDNA72AH.js → chunk-6QOTUPRG.js} +155 -36
- package/dist/{chunk-X7IARSHT.js → chunk-6UINWWS6.js} +16 -10
- package/dist/{chunk-CYZW7JHJ.js → chunk-72YIHOZQ.js} +9 -9
- package/dist/{chunk-IKSLQ4XV.js → chunk-75W7L2E2.js} +752 -861
- package/dist/{chunk-CRFOIAX3.js → chunk-7UGL4MB5.js} +6 -6
- package/dist/{chunk-HIICAHCJ.js → chunk-AUPNRN7C.js} +2 -2
- package/dist/{chunk-7BHIY2MW.js → chunk-BJVFZO5U.js} +8 -50
- package/dist/{chunk-B74PXLU7.js → chunk-CUSRQKPU.js} +65 -3
- package/dist/chunk-DQOVN6KV.js +386 -0
- package/dist/{chunk-E7GT7O5N.js → chunk-DT3LWJOB.js} +7 -4
- package/dist/chunk-DXKJURES.js +671 -0
- package/dist/{chunk-JBCS7CRR.js → chunk-EL24TAU4.js} +10 -10
- package/dist/{chunk-TPEQIQIE.js → chunk-ELWDPP3Y.js} +8 -8
- package/dist/{chunk-NDINPTJ4.js → chunk-ELZVTCGV.js} +5 -4
- package/dist/chunk-EXLD33WO.js +381 -0
- package/dist/chunk-FEAXX7B6.js +101 -0
- package/dist/{chunk-RLYRBIYQ.js → chunk-FFUPXJC4.js} +90 -331
- package/dist/{chunk-5PFYMY2V.js → chunk-FTMGRKEF.js} +2 -2
- package/dist/{chunk-34BHNEE3.js → chunk-GHS5EBTQ.js} +58 -7
- package/dist/{chunk-DYHAXKHD.js → chunk-GWZNEVM2.js} +12 -8
- package/dist/chunk-GX5WYQO4.js +59 -0
- package/dist/chunk-GYV6VZOC.js +26 -0
- package/dist/{chunk-MQXIVJ35.js → chunk-HAXOFFRH.js} +5 -5
- package/dist/chunk-I2DWJ4GM.js +390 -0
- package/dist/{chunk-TXOTCRLG.js → chunk-I5FWO7L5.js} +5 -5
- package/dist/{chunk-WNP7O5WZ.js → chunk-ID64D7PE.js} +4 -4
- package/dist/{chunk-XQRY4DTA.js → chunk-IGLP3ODT.js} +10 -10
- package/dist/chunk-IRXAATOX.js +539 -0
- package/dist/chunk-IXIY2H4R.js +44 -0
- package/dist/{chunk-SSEYRH53.js → chunk-IZXGRF7P.js} +92 -147
- package/dist/{chunk-5TSRNF4G.js → chunk-JCI2ROMZ.js} +164 -6
- package/dist/{chunk-JWJGP5DQ.js → chunk-JEIYHLOR.js} +7 -7
- package/dist/{chunk-F2I26BDK.js → chunk-JQLNNIKT.js} +4 -4
- package/dist/{chunk-BYMNWQ7O.js → chunk-JSD46VO2.js} +315 -63
- package/dist/{chunk-AK5XEFVZ.js → chunk-JT2RFCC5.js} +64 -14
- package/dist/{chunk-MCMZMDAC.js → chunk-K6T2ZAMZ.js} +168 -6
- package/dist/{chunk-PGF63K6I.js → chunk-KFV5L5SK.js} +73 -4
- package/dist/{chunk-IHXBNWMM.js → chunk-KXDSS5WJ.js} +7 -3
- package/dist/{chunk-PJX3WQUQ.js → chunk-LLXSDWXS.js} +3 -3
- package/dist/{chunk-DZAW46HP.js → chunk-LTIKRKFL.js} +3 -3
- package/dist/{chunk-DZEK6CJN.js → chunk-N56KALIC.js} +21 -21
- package/dist/{chunk-B7HM5Z7T.js → chunk-NAI6ZFCY.js} +9 -5
- package/dist/{chunk-I66ZTYNP.js → chunk-NRO2BJRH.js} +2656 -2213
- package/dist/{chunk-ZGNYYXQ6.js → chunk-NXIMQY5W.js} +3 -3
- package/dist/{chunk-IKOZFYBN.js → chunk-NXYCB2VD.js} +149 -106
- package/dist/{chunk-CWVRRIEI.js → chunk-NZMNUPZZ.js} +2 -2
- package/dist/{chunk-462T4EGZ.js → chunk-O5CVSAG5.js} +2 -2
- package/dist/chunk-ODGTEFFI.js +50 -0
- package/dist/{chunk-3F7VUY77.js → chunk-OEJSLEPW.js} +2 -2
- package/dist/{chunk-KKOJXO6R.js → chunk-OMQNJVKW.js} +4 -2
- package/dist/{chunk-5KW52TEP.js → chunk-Q4WO54TA.js} +132 -77
- package/dist/{chunk-W6NIE6OW.js → chunk-QUFRYSWI.js} +13 -7
- package/dist/{chunk-42FMPA75.js → chunk-QZWQA4DE.js} +2 -2
- package/dist/chunk-R6Q67RJH.js +134 -0
- package/dist/{chunk-W4YEMFBX.js → chunk-RAY4OVGZ.js} +3 -3
- package/dist/{chunk-ZNT2M6TG.js → chunk-RQCKCSRL.js} +17 -17
- package/dist/{chunk-LJID3DYZ.js → chunk-RXTN6AKH.js} +3 -3
- package/dist/{chunk-P75RZCJW.js → chunk-RZDWV63N.js} +3 -3
- package/dist/{chunk-UH632ZYL.js → chunk-S6PYF2XF.js} +2 -2
- package/dist/{chunk-HJWWJ6IL.js → chunk-TOIVGRUX.js} +17 -5
- package/dist/{chunk-HLAFFSEK.js → chunk-TQAHXW6Y.js} +2 -2
- package/dist/{chunk-JIEGK6UF.js → chunk-U6TMQNSI.js} +48 -4
- package/dist/{chunk-2HFQNRV3.js → chunk-UEPWCCTY.js} +12 -12
- package/dist/chunk-UOIZ7DA4.js +41 -0
- package/dist/{chunk-UH347SHR.js → chunk-USR47QNF.js} +11 -11
- package/dist/{chunk-AZ4WMN4W.js → chunk-V6HJFQZE.js} +2 -2
- package/dist/chunk-V76WTFTW.js +318 -0
- package/dist/{chunk-NMJXSHBJ.js → chunk-W54I7H25.js} +2 -2
- package/dist/{chunk-KPXDY6QF.js → chunk-XRZT5WY5.js} +2 -2
- package/dist/{chunk-UBRFI4HS.js → chunk-XULDXHTN.js} +142 -50
- package/dist/chunk-XXYSBZIQ.js +283 -0
- package/dist/{chunk-HKMD33FO.js → chunk-Y55JBDO5.js} +405 -122
- package/dist/{chunk-XOXV5GKE.js → chunk-YD5GIKET.js} +17 -8
- package/dist/{chunk-XGDPUNND.js → chunk-YECAMM3D.js} +2 -2
- package/dist/{chunk-BO7Y52RY.js → chunk-YNFKXPEC.js} +7 -7
- package/dist/{chunk-VXMFAE2W.js → chunk-YPC6ZR5L.js} +19 -6
- package/dist/{chunk-M2WXEHER.js → chunk-ZA4VCIGV.js} +2 -2
- package/dist/cli/index.js +42 -40
- package/dist/{clio-7VB377CC.js → clio-QLICPCF5.js} +7 -7
- package/dist/{code-nav-YVLCYA7V.js → code-nav-IJR2DBPR.js} +9 -9
- package/dist/{components-UBWCQSRW.js → components-2TGAI2RC.js} +5 -6
- package/dist/{config-4HVOS65E.js → config-IUA6OYNS.js} +88 -81
- package/dist/{configure-PIWO7B24.js → configure-VEPX4NMX.js} +26 -25
- package/dist/{context-KQYIWPWT.js → context-2DKHWH2T.js} +60 -45
- package/dist/{context-IYEHL3WQ.js → context-4MPR7WKB.js} +78 -69
- package/dist/{context-N6ZE3LGJ.js → context-BOYF5EJM.js} +15 -11
- package/dist/{context-clear-G4OGZJDS.js → context-clear-S4ZJCQUX.js} +73 -65
- package/dist/context-map-COB37XXN.js +505 -0
- package/dist/{context-working-set-BWLF6LJP.js → context-working-set-3I3FYX6Y.js} +18 -17
- package/dist/detail-A7JAVSIG.js +98 -0
- package/dist/{dispatch-runner-2QQAITS3.js → dispatch-runner-RJ5I2F2O.js} +99 -75
- package/dist/{docs-PD3EXDKU.js → docs-SPOV3BAN.js} +3 -5
- package/dist/{doctor-LHBD36VU.js → doctor-DKICC2SN.js} +71 -48
- package/dist/{eval-C45FYRJ6.js → eval-OQOQUDHK.js} +308 -146
- package/dist/{eval-inventory-6DEJPLBF.js → eval-inventory-Y6QRFOH5.js} +4 -4
- package/dist/{evidence-6SHONYAF.js → evidence-4DQ25GUQ.js} +79 -175
- package/dist/evidence-4F5USFKH.js +208 -0
- package/dist/{evolve-KRKMV72X.js → evolve-GSS52E5J.js} +71 -67
- package/dist/{extensions-KPZ2UHBB.js → extensions-G7MFLYHT.js} +8 -9
- package/dist/{fleet-IVTCKDHT.js → fleet-Q37YHHAQ.js} +126 -118
- package/dist/{fleet-commands-EDWL3IT7.js → fleet-commands-G7E4N7SM.js} +16 -13
- package/dist/{fleet-decisions-YP3YEFGK.js → fleet-decisions-O7M6QBA2.js} +9 -8
- package/dist/{fleet-graph-ZFWKHY2M.js → fleet-graph-TOUBW6OW.js} +21 -20
- package/dist/{fleet-inspect-FVUNCBML.js → fleet-inspect-SW33JJNI.js} +65 -60
- package/dist/{fleet-preflight-UN5XED4R.js → fleet-preflight-CV2655TW.js} +4 -5
- package/dist/{fleet-validate-XOWC4HSX.js → fleet-validate-KESZX2YH.js} +25 -24
- package/dist/{fleet-verify-UN3SODEL.js → fleet-verify-UQPMTVE3.js} +66 -61
- package/dist/{fleet-view-TWHJKCN6.js → fleet-view-MN2VG4MR.js} +65 -60
- package/dist/{init-T2QORQ3Y.js → init-PXEXQSBF.js} +90 -82
- package/dist/{interop-IN5I2A66.js → interop-ZG5T62U3.js} +12 -13
- package/dist/inventory-C26CFDRR.js +101 -0
- package/dist/{library-LSCATDLZ.js → library-B2W4N74O.js} +29 -29
- package/dist/{memory-HYOKAGGJ.js → memory-YCANYS5A.js} +73 -69
- package/dist/{models-2GPMFYCM.js → models-GERTU3YI.js} +51 -47
- package/dist/{monitor-E4ASVUJH.js → monitor-CPNIUULB.js} +74 -67
- package/dist/{orchestrator-DDMPR3PY.js → orchestrator-J4BSH4WQ.js} +1288 -1606
- package/dist/{panes-E3RUXOW5.js → panes-BOHAEGYC.js} +4 -4
- package/dist/{panes-IXKLOKA2.js → panes-NXSLDQZ2.js} +10 -11
- package/dist/{paths-L7LGY6RN.js → paths-VSUWNC22.js} +6 -7
- package/dist/reset-TNWTB5LU.js +343 -0
- package/dist/{resources-OTRSN34L.js → resources-4PXNMD5G.js} +29 -22
- package/dist/{run-5DEYH5QK.js → run-D6XJ34CN.js} +132 -132
- package/dist/{share-IHWTLO3M.js → share-2NWMJJEE.js} +27 -27
- package/dist/{skills-IYMXMKW4.js → skills-KR7WON5G.js} +40 -33
- package/dist/{skills-eval-DROHSJAR.js → skills-eval-O2ZNOLDS.js} +81 -77
- package/dist/{skills-inventory-D7X4L4ZX.js → skills-inventory-ZZOUBK7O.js} +23 -22
- package/dist/{slash-commands-QBM7UZ3B.js → slash-commands-ZXPJD64J.js} +47 -37
- package/dist/{steer-Z5DO23FJ.js → steer-XA25PSCS.js} +4 -4
- package/dist/{support-U7QOWY26.js → support-7EMVWYG2.js} +6 -6
- package/dist/{targets-P2FUC4IL.js → targets-OMH2XCSN.js} +50 -49
- package/dist/tasks-IPAGMEIX.js +36 -0
- package/dist/{terminal-lease-YREJ3JX2.js → terminal-lease-C2J3JYRE.js} +4 -4
- package/dist/{tools-5B7RO6MV.js → tools-EFFEAIDP.js} +8 -9
- package/dist/{trace-YMGMUM6A.js → trace-FXMXUZUF.js} +7 -7
- package/dist/uninstall-HALS6BLF.js +407 -0
- package/dist/upgrade-MS72RJEP.js +306 -0
- package/dist/{usage-ME5MPXGX.js → usage-NHG6MCJM.js} +162 -108
- package/dist/{verifiers-BVZ7IWOO.js → verifiers-7AUNVXDY.js} +155 -22
- package/dist/{verify-5K7ZKQFC.js → verify-FWYGPKMR.js} +14 -12
- package/dist/{web-fetch-MPARV2K7.js → web-fetch-V4FKSDAV.js} +4 -4
- package/dist/{wiki-generate-F5W5QTYY.js → wiki-generate-743CIGJW.js} +99 -90
- package/dist/{with-panes-BYOJCLAM.js → with-panes-BDQEWBRT.js} +10 -10
- package/dist/worker/entry.js +72 -68
- package/docs/README.md +3 -2
- package/docs/architecture/acp.md +17 -0
- package/docs/architecture/artifact-placement.md +1 -0
- package/docs/architecture/artifact-versions.md +2 -2
- package/docs/architecture/context-engine.md +4 -0
- package/docs/architecture/dispatch-typed-intent.md +1 -1
- package/docs/architecture/evidence-and-memory.md +1 -1
- package/docs/architecture/middleware-and-components.md +1 -1
- package/docs/architecture/model-catalog.md +21 -10
- package/docs/architecture/observability.md +19 -2
- package/docs/architecture/prompt-envelope-and-tools.md +17 -5
- package/docs/architecture/provider-adapter-cookbook.md +63 -0
- package/docs/architecture/safety-model.md +25 -22
- package/docs/architecture/tui-design.md +1 -1
- package/docs/guide/built-in-agents.md +25 -11
- package/docs/guide/commands-and-modes.md +18 -3
- package/docs/guide/configuration-and-targets.md +100 -10
- package/docs/guide/configuration-reference.md +17 -7
- package/docs/guide/environment-variables.md +4 -2
- package/docs/guide/installation-and-lifecycle.md +37 -4
- package/docs/guide/proactive-memory.md +66 -55
- package/docs/guide/skills-marketplace.md +18 -0
- package/docs/guide/tool-usage.md +78 -3
- package/docs/history/config-knobs-audit.md +2 -2
- package/docs/process/development-pipeline.md +40 -2
- package/docs/process/eval-runner.md +67 -3
- package/docs/process/git-commit-provenance.md +15 -0
- package/docs/process/release-cut-checklist.md +207 -0
- package/docs/process/scientific-validation.md +18 -17
- package/evals/behavioral-machinery-support.ts +1 -0
- package/evals/behavioral-machinery.yaml +1 -1
- package/evals/behavioral-model.yaml +3 -2
- package/package.json +2 -2
- package/skills/README.md +7 -5
- package/skills/coding/ast-grep/SKILL.md +101 -30
- package/skills/coding/ast-grep/evals.md +26 -0
- package/skills/coding/coding-standards/SKILL.md +47 -2
- package/skills/coding/coding-standards/evals.md +23 -0
- package/skills/coding/prototype/SKILL.md +87 -28
- package/skills/coding/prototype/evals.md +19 -0
- package/skills/coding/tdd/SKILL.md +80 -53
- package/skills/coding/tdd/evals.md +20 -0
- package/skills/context/context-handoff/SKILL.md +43 -2
- package/skills/context/context-handoff/evals.md +44 -0
- package/skills/context/context-prime/SKILL.md +45 -15
- package/skills/context/context-prime/evals.md +45 -0
- package/skills/git/branch-closeout/SKILL.md +132 -0
- package/skills/git/branch-closeout/evals.md +133 -0
- package/skills/git/branch-closeout/references/closeout-checklist.md +81 -0
- package/skills/git/file-ticket/SKILL.md +77 -63
- package/skills/git/file-ticket/assets/issue-template.md +22 -0
- package/skills/git/file-ticket/evals.md +31 -26
- package/skills/git/file-ticket/references/issue-discovery.md +49 -0
- package/skills/git/fix-issue/SKILL.md +87 -64
- package/skills/git/fix-issue/evals.md +35 -31
- package/skills/git/fix-issue/references/diagnosis-and-rca.md +46 -0
- package/skills/git/resolve-merge-conflicts/SKILL.md +100 -51
- package/skills/git/resolve-merge-conflicts/evals.md +52 -25
- package/skills/git/resolve-merge-conflicts/references/conflict-matrix.md +126 -0
- package/skills/git/ship/SKILL.md +103 -67
- package/skills/git/ship/assets/pr-template.md +21 -0
- package/skills/git/ship/evals.md +44 -28
- package/skills/git/ship/references/remote-and-branch-policy.md +62 -0
- package/skills/git/worktree-create/SKILL.md +80 -50
- package/skills/git/worktree-create/evals.md +40 -33
- package/skills/git/worktree-create/references/worktree-setup.md +62 -66
- package/skills/git/worktree-merge/SKILL.md +112 -65
- package/skills/git/worktree-merge/evals.md +42 -34
- package/skills/git/worktree-merge/references/merge-strategies.md +52 -0
- package/skills/planning/archify/SKILL.md +196 -0
- package/skills/planning/archify/evals.md +65 -0
- package/skills/planning/architecture/SKILL.md +61 -12
- package/skills/planning/architecture/evals.md +65 -0
- package/skills/planning/backlog/SKILL.md +130 -14
- package/skills/planning/backlog/evals.md +142 -0
- package/skills/planning/prd/SKILL.md +47 -6
- package/skills/planning/prd/evals.md +54 -0
- package/skills/planning/product-intent/SKILL.md +57 -2
- package/skills/planning/product-intent/evals.md +70 -0
- package/skills/planning/tech-spec/SKILL.md +53 -2
- package/skills/planning/tech-spec/evals.md +73 -0
- package/skills/registry.yaml +58 -50
- package/skills/remote.yaml +13 -0
- package/skills/research/arxiv-literature/SKILL.md +76 -18
- package/skills/research/arxiv-literature/evals.md +50 -0
- package/skills/research/experiment-protocol/SKILL.md +20 -1
- package/skills/research/experiment-protocol/evals.md +23 -0
- package/skills/research/scientific-debugging/SKILL.md +24 -1
- package/skills/research/scientific-debugging/evals.md +18 -0
- package/skills/research/scientific-modernization/SKILL.md +26 -1
- package/skills/research/scientific-modernization/evals.md +27 -0
- package/skills/skill-marketplace.json +63 -28
- package/skills/workflow/cut-it/SKILL.md +64 -5
- package/skills/workflow/cut-it/evals.md +101 -0
- package/skills/workflow/design-council/SKILL.md +112 -27
- package/skills/workflow/design-council/evals.md +161 -0
- package/skills/workflow/grill-me/SKILL.md +85 -10
- package/skills/workflow/grill-me/evals.md +153 -0
- package/skills/workflow/workflow-distiller/SKILL.md +76 -17
- package/skills/workflow/workflow-distiller/evals.md +118 -0
- package/src/cli/args.ts +0 -8
- package/src/cli/configure-interop.ts +105 -13
- package/src/cli/configure-oauth.ts +57 -0
- package/src/cli/configure-onboarding.ts +980 -0
- package/src/cli/configure-target.ts +594 -0
- package/src/cli/configure.ts +1084 -529
- package/src/cli/context-map.ts +114 -0
- package/src/cli/context.ts +4 -0
- package/src/cli/doctor-state-size.ts +1 -12
- package/src/cli/doctor-validation-contract.ts +28 -0
- package/src/cli/doctor.ts +5 -0
- package/src/cli/evidence-detail.ts +1 -75
- package/src/cli/evidence-inventory.ts +1 -167
- package/src/cli/index.ts +3 -0
- package/src/cli/lifecycle-presenter.ts +436 -0
- package/src/cli/models.ts +10 -2
- package/src/cli/modes/print.ts +5 -1
- package/src/cli/reset.ts +228 -106
- package/src/cli/run.ts +7 -4
- package/src/cli/select.ts +664 -0
- package/src/cli/skills.ts +9 -2
- package/src/cli/targets.ts +3 -0
- package/src/cli/tasks.ts +84 -0
- package/src/cli/uninstall.ts +233 -165
- package/src/cli/upgrade.ts +210 -150
- package/src/cli/usage.ts +92 -27
- package/src/cli/validate-model.ts +3 -3
- package/src/cli/verifiers.ts +147 -1
- package/src/cli/wiki-generate.ts +1 -0
- package/src/core/commit-attribution.ts +41 -1
- package/src/core/config.ts +56 -0
- package/src/core/external-diagnostic.ts +44 -0
- package/src/core/gateway-routing.ts +157 -0
- package/src/core/git-commit-attribution.ts +46 -3
- package/src/core/run-overrides.ts +0 -5
- package/src/core/safe-exec.ts +17 -2
- package/src/core/skill-activation.ts +92 -2
- package/src/core/tool-names.ts +5 -2
- package/src/domains/agents/builtins/architect.md +1 -1
- package/src/domains/agents/builtins/coder.md +1 -1
- package/src/domains/agents/builtins/documenter.md +1 -1
- package/src/domains/agents/builtins/git-master.md +1 -1
- package/src/domains/agents/builtins/provenance.md +7 -7
- package/src/domains/agents/builtins/tester.md +1 -1
- package/src/domains/agents/builtins/verifier.md +2 -2
- package/src/domains/agents/builtins/wiki-writer.md +4 -3
- package/src/domains/agents/builtins/world-knowledge.md +31 -0
- package/src/domains/agents/catalog.ts +1 -1
- package/src/domains/agents/result-contract.ts +70 -0
- package/src/domains/context/extension.ts +31 -7
- package/src/domains/context/refresh.ts +3 -0
- package/src/domains/context/wiki/frontmatter.ts +5 -2
- package/src/domains/context/wiki/generate.ts +6 -0
- package/src/domains/context/wiki/map-seed.ts +589 -0
- package/src/domains/context/wiki/plan.ts +2 -2
- package/src/domains/context/wiki/prompts.ts +43 -0
- package/src/domains/dispatch/active-route-planner.ts +4 -0
- package/src/domains/dispatch/admission.ts +29 -0
- package/src/domains/dispatch/agent-candidates.ts +10 -0
- package/src/domains/dispatch/budget-envelope.ts +86 -1
- package/src/domains/dispatch/capability-match.ts +1 -0
- package/src/domains/dispatch/capacity-lease.ts +17 -0
- package/src/domains/dispatch/code-step.ts +11 -4
- package/src/domains/dispatch/contract.ts +42 -8
- package/src/domains/dispatch/execution-scheduler.ts +2 -0
- package/src/domains/dispatch/extension.ts +184 -56
- package/src/domains/dispatch/fleet-commit-attribution.ts +5 -0
- package/src/domains/dispatch/fleet-run.ts +1 -0
- package/src/domains/dispatch/host-verification.ts +114 -13
- package/src/domains/dispatch/intent.ts +28 -18
- package/src/domains/dispatch/orphan-recovery.ts +2 -0
- package/src/domains/dispatch/receipt-integrity.ts +4 -0
- package/src/domains/dispatch/reservation-store.ts +5 -3
- package/src/domains/dispatch/state.ts +15 -2
- package/src/domains/dispatch/types.ts +20 -2
- package/src/domains/dispatch/worker-model-metadata.ts +38 -0
- package/src/domains/eval/metrics/call-ledger-stream.ts +34 -11
- package/src/domains/eval/metrics/token-stream.ts +201 -31
- package/src/domains/eval/metrics/tracked.ts +40 -4
- package/src/domains/eval/runners/clio-run.ts +17 -11
- package/src/domains/eval/runners/context-index.ts +2 -7
- package/src/domains/eval/runners/context-init.ts +3 -6
- package/src/domains/eval/runners/external-command.ts +29 -11
- package/src/domains/eval/schema/suite.ts +28 -0
- package/src/domains/eval/schema/verdict.ts +2 -2
- package/src/domains/eval/suites/resolve.ts +13 -1
- package/src/domains/eval/suites/run.ts +24 -3
- package/src/domains/evidence/build.ts +102 -15
- package/src/domains/evidence/detail.ts +69 -0
- package/src/domains/evidence/eval.ts +13 -1
- package/src/domains/evidence/finish-contract-map.ts +5 -1
- package/src/domains/evidence/inventory.ts +167 -0
- package/src/domains/evidence/store.ts +16 -0
- package/src/domains/evidence/types.ts +12 -0
- package/src/domains/extensions/resources.ts +7 -0
- package/src/domains/interop/registry.ts +6 -2
- package/src/domains/interop/types.ts +4 -0
- package/src/domains/lifecycle/migrations/index.ts +4 -0
- package/src/domains/memory/task-memory-policy.ts +70 -26
- package/src/domains/memory/task-memory-telemetry.ts +1 -0
- package/src/domains/middleware/index.ts +0 -1
- package/src/domains/middleware/marketplace-offer.ts +22 -35
- package/src/domains/middleware/memory-intervention.ts +127 -32
- package/src/domains/middleware/memory-step-endpoint.ts +3 -2
- package/src/domains/middleware/runtime.ts +7 -3
- package/src/domains/middleware/skills-reminder.ts +31 -2
- package/src/domains/mux/detect.ts +3 -6
- package/src/domains/observability/accountability.ts +15 -1
- package/src/domains/observability/compaction-usage.ts +118 -0
- package/src/domains/observability/contract.ts +52 -7
- package/src/domains/observability/cost.ts +1 -1
- package/src/domains/observability/evidence-index.ts +10 -0
- package/src/domains/observability/extension.ts +15 -6
- package/src/domains/observability/out-of-turn-usage.ts +52 -21
- package/src/domains/observability/projection.ts +394 -45
- package/src/{interactive → domains/observability}/worker-progress.ts +3 -3
- package/src/domains/prompts/fragments/operating/contract.md +2 -0
- package/src/domains/prompts/fragments/wiki/page.md +8 -0
- package/src/domains/providers/contract.ts +4 -1
- package/src/domains/providers/extension.ts +40 -9
- package/src/domains/providers/model-capabilities.ts +9 -0
- package/src/domains/providers/model-discovery.ts +3 -4
- package/src/domains/providers/model-runtime-capabilities.ts +15 -5
- package/src/domains/providers/models/local-models/clio-coder-local-coding-targets.yaml +48 -26
- package/src/domains/providers/runtimes/antigravity/antigravity-code.ts +225 -45
- package/src/domains/providers/runtimes/claude/claude-code.ts +9 -0
- package/src/domains/providers/runtimes/common/lmstudio-http.ts +6 -2
- package/src/domains/providers/runtimes/common/local-synth.ts +2 -0
- package/src/domains/providers/runtimes/protocol/litellm.ts +119 -29
- package/src/domains/providers/support.ts +11 -5
- package/src/domains/providers/target-model-cache.ts +25 -2
- package/src/domains/providers/types/capability-flags.ts +2 -0
- package/src/domains/providers/types/runtime-descriptor.ts +20 -1
- package/src/domains/providers/types/target-descriptor.ts +19 -0
- package/src/domains/resources/index.ts +3 -0
- package/src/domains/resources/skills/install.ts +72 -7
- package/src/domains/resources/skills/loader.ts +7 -0
- package/src/domains/resources/skills/marketplace.ts +63 -11
- package/src/domains/safety/action-classifier.ts +7 -0
- package/src/domains/safety/autonomy.ts +15 -0
- package/src/domains/safety/default-path-policy.ts +2 -0
- package/src/domains/safety/finish-contract-registration.ts +29 -14
- package/src/domains/safety/finish-contract.ts +252 -40
- package/src/domains/safety/index.ts +21 -1
- package/src/domains/safety/path-policy.ts +1 -1
- package/src/domains/safety/policy-engine.ts +60 -17
- package/src/domains/safety/protected-artifacts.ts +191 -88
- package/src/domains/safety/rigor.ts +53 -39
- package/src/domains/safety/run-effects.ts +2 -22
- package/src/domains/safety/skill-authority.ts +55 -0
- package/src/domains/safety/validation-contract.ts +388 -0
- package/src/domains/session/archive-readers.ts +10 -1
- package/src/domains/session/compaction/compact.ts +72 -22
- package/src/domains/session/decision-board.ts +101 -2
- package/src/domains/session/entries.ts +50 -7
- package/src/domains/session/extension.ts +4 -4
- package/src/domains/session/handoff.ts +2 -1
- package/src/domains/session/manager.ts +2 -3
- package/src/domains/session/task-board.ts +14 -1
- package/src/domains/session/tree/fork.ts +1 -2
- package/src/domains/session/tree/navigator.ts +1 -1
- package/src/domains/session/usage.ts +3 -3
- package/src/domains/user-tasks/acceptance.ts +56 -0
- package/src/domains/user-tasks/active-acceptance.ts +40 -0
- package/src/domains/user-tasks/store.ts +34 -3
- package/src/engine/acp/adapter.ts +24 -6
- package/src/engine/acp/server.ts +21 -4
- package/src/engine/acp/transport.ts +53 -8
- package/src/engine/acp/types.ts +4 -0
- package/src/engine/agent.ts +13 -3
- package/src/engine/ai.ts +26 -8
- package/src/engine/antigravity/subprocess-runtime.ts +386 -120
- package/src/engine/api-registry.ts +3 -0
- package/src/engine/apis/ollama-native.ts +15 -0
- package/src/engine/apis/openai-completions.ts +117 -14
- package/src/engine/claude/subprocess-runtime.ts +107 -60
- package/src/engine/external-subprocess.ts +122 -6
- package/src/entry/background-model-metadata.ts +18 -0
- package/src/entry/compaction-prompt.ts +57 -0
- package/src/entry/orchestrator.ts +416 -218
- package/src/entry/task-memory-lifecycle.ts +35 -0
- package/src/interactive/chat-loop-messages.ts +13 -4
- package/src/interactive/chat-loop.ts +65 -2
- package/src/interactive/chat-renderer.ts +1 -0
- package/src/interactive/cost-overlay.ts +26 -2
- package/src/interactive/dispatch-board.ts +46 -717
- package/src/interactive/fleet-run-preview.ts +2 -1
- package/src/interactive/interactive-application.ts +3 -2
- package/src/interactive/interactive-presentation.ts +55 -12
- package/src/interactive/interactive-slash-runtime.ts +4 -2
- package/src/interactive/oracle.ts +5 -2
- package/src/interactive/overlays/fleet-run-approval.ts +3 -2
- package/src/interactive/overlays/message-picker.ts +2 -2
- package/src/interactive/overlays/settings.ts +2 -2
- package/src/interactive/overlays/tree-selector.ts +2 -2
- package/src/interactive/renderers/branch-summary.ts +1 -1
- package/src/interactive/renderers/worker-entry.ts +32 -0
- package/src/interactive/slash-autocomplete.ts +4 -6
- package/src/interactive/slash-commands.ts +49 -45
- package/src/interactive/slash-spec.ts +28 -0
- package/src/interactive/theme/labels.ts +19 -13
- package/src/interactive/turn-context.ts +9 -5
- package/src/interactive/turn-recovery.ts +8 -0
- package/src/interactive/turn-runtime.ts +27 -11
- package/src/interactive/turn-state.ts +7 -0
- package/src/interactive/view/artifacts.ts +2 -0
- package/src/interactive/worker-receipts.ts +1 -0
- package/src/interactive/worker-stream.ts +13 -4
- package/src/tools/bootstrap.ts +4 -0
- package/src/tools/builtin-tool-catalog.ts +31 -0
- package/src/tools/compete-worktrees.ts +7 -1
- package/src/tools/context/index.ts +30 -9
- package/src/tools/core-bootstrap.ts +16 -0
- package/src/tools/decide.ts +136 -0
- package/src/tools/dispatch-admission.ts +21 -0
- package/src/tools/dispatch-arguments.ts +1 -0
- package/src/tools/dispatch-event-text.ts +10 -0
- package/src/tools/dispatch-plan.ts +6 -2
- package/src/tools/dispatch-runner.ts +27 -1
- package/src/tools/dispatch-types.ts +6 -0
- package/src/tools/evidence.ts +96 -0
- package/src/tools/limitation.ts +76 -0
- package/src/tools/policy.ts +9 -0
- package/src/tools/presentation.ts +3 -0
- package/src/tools/registry.ts +11 -5
- package/src/tools/result-shaping.ts +17 -5
- package/src/tools/task-worktree.ts +13 -3
- package/src/tools/tasks.ts +10 -1
- package/src/tools/verify/authoring.ts +170 -83
- package/src/tools/verify/catalog.ts +122 -5
- package/src/tools/verify/index.ts +2 -1
- package/src/tools/verify/numeric.ts +298 -0
- package/src/tools/verify/perf.ts +143 -0
- package/src/tools/verify/scripts.ts +229 -2
- package/src/tools/worker-evidence.ts +3 -1
- package/src/worker/spec-contract.ts +4 -0
- package/dist/chunk-2Z2IKEXI.js +0 -1554
- package/dist/chunk-RVG5JXAL.js +0 -41
- package/dist/chunk-T56WDKA5.js +0 -183
- package/dist/chunk-VN3SHNBN.js +0 -313
- package/dist/chunk-VPKWYKEY.js +0 -169
- package/dist/reset-OAQP3W4O.js +0 -230
- package/dist/uninstall-N34PCTGJ.js +0 -331
- package/dist/upgrade-PXK3S2YM.js +0 -325
|
@@ -121,7 +121,7 @@ malformed response, or telemetry failure is silent and never blocks a tool.
|
|
|
121
121
|
- `context.memory.cadenceToolCalls` (default `10`): Minimum completed-tool interval between background interventions.
|
|
122
122
|
- `context.memory.trajectorySteps` (default `8`): Completed tool-trajectory window analyzed during background evaluation.
|
|
123
123
|
- `context.memory.maxOutputTokens` (default `2000`): Bounds the rendered memory-bank context and the ordinary policy-model completion. An always-on-thinking model receives additional reasoning headroom, at least `4,000` tokens when its model cap permits, so it can still reach the strict envelope.
|
|
124
|
-
- `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy
|
|
124
|
+
- `context.memory.timeoutMs` (default `60000`): Wall-clock limit for one background memory-policy step, including its one permitted chat fallback attempt. The step is detached, so this deadline never delays a turn, but it does hold a request slot on a real inference endpoint that your own turns and dispatched workers queue against. Raise it only after inspecting the timeout and hit-rate evidence for the selected route.
|
|
125
125
|
|
|
126
126
|
## Trigger semantics
|
|
127
127
|
|
|
@@ -254,38 +254,48 @@ unexamined one:
|
|
|
254
254
|
would not have been spent. The current 60-second deadline is the source-backed
|
|
255
255
|
bound on what one optional call may hold a shared local server for, not a
|
|
256
256
|
prediction of when a route answers.
|
|
257
|
-
- A step
|
|
258
|
-
skipped with reason `endpoint_busy
|
|
259
|
-
|
|
260
|
-
|
|
261
|
-
|
|
262
|
-
|
|
263
|
-
|
|
264
|
-
|
|
265
|
-
|
|
266
|
-
|
|
267
|
-
|
|
268
|
-
|
|
269
|
-
|
|
270
|
-
|
|
271
|
-
|
|
272
|
-
|
|
273
|
-
|
|
274
|
-
|
|
275
|
-
|
|
276
|
-
|
|
277
|
-
|
|
278
|
-
|
|
279
|
-
|
|
280
|
-
|
|
281
|
-
|
|
282
|
-
|
|
283
|
-
|
|
284
|
-
|
|
285
|
-
|
|
286
|
-
|
|
287
|
-
|
|
288
|
-
|
|
257
|
+
- A step whose known endpoint occupancy exhausts the resolved request capacity is
|
|
258
|
+
skipped with reason `endpoint_busy`. The same gateway URL alone does not imply
|
|
259
|
+
a single request slot or a single physical model server.
|
|
260
|
+
|
|
261
|
+
### Dedicated routing, chat fallback and endpoint capacity
|
|
262
|
+
|
|
263
|
+
Clio prefers the explicitly configured memory target and model. If that route is
|
|
264
|
+
known unavailable (missing target/runtime, down target, absent model in a known
|
|
265
|
+
catalog, or a model reported unloaded/loading), Clio selects the active chat
|
|
266
|
+
route for that step. A runtime client error on the dedicated route permits one
|
|
267
|
+
chat attempt within the original remaining deadline. It does not retry after a
|
|
268
|
+
timeout, cancellation, session/branch switch, malformed envelope, or a model's
|
|
269
|
+
explicit silence. The fallback never falls back again. Both routes unavailable
|
|
270
|
+
produce a visible `client_error` outcome. An unset memory role remains rules-only;
|
|
271
|
+
chat fallback does not enable model-based memory by default.
|
|
272
|
+
|
|
273
|
+
A runtime notice names the selected chat fallback and its reason. Routing stays
|
|
274
|
+
session-local and does not edit saved settings. Each attempted call records its
|
|
275
|
+
own known usage under its actual target/model and emits its own telemetry outcome;
|
|
276
|
+
failed dedicated usage is retained alongside fallback usage. The final policy
|
|
277
|
+
result describes the final attempt, without merging different provider identities.
|
|
278
|
+
A generation change discards late content while preserving the original usage owner.
|
|
279
|
+
|
|
280
|
+
Capacity follows the same existing evidence as dispatch: explicit target
|
|
281
|
+
`maxConcurrentRequests`, current discovery, a valid discovery prior, then the
|
|
282
|
+
conservative one-slot default for a local-native runtime. Known occupancy includes
|
|
283
|
+
this process's foreground requests, active dispatch leases and held reservation
|
|
284
|
+
waves. A two-slot endpoint with one foreground request can admit memory; a full
|
|
285
|
+
one-slot endpoint skips it. Known dedicated saturation does not itself initiate
|
|
286
|
+
chat fallback. Larger declared capacities work without a new constant.
|
|
287
|
+
|
|
288
|
+
LiteLLM is a gateway protocol, so Clio does not invent a local one-slot limit for
|
|
289
|
+
its URL. Distinct model routes such as dynamo, mini and zbook can share that URL;
|
|
290
|
+
the gateway owns their physical routing and backend residency. An explicit or
|
|
291
|
+
observed endpoint-wide bound still applies when present. Clio's process-local
|
|
292
|
+
foreground holds and observed dispatch state are not a global scheduler for every
|
|
293
|
+
client using the gateway. Slot holds are released on success, failure and abort.
|
|
294
|
+
|
|
295
|
+
The existing `background_memory` expected-cold stamp remains keyed by endpoint
|
|
296
|
+
URL. It records a possible shared-endpoint cache disturbance, including after a
|
|
297
|
+
failed request, rather than proving that a different model behind the gateway
|
|
298
|
+
actually evicted the chat prefix.
|
|
289
299
|
|
|
290
300
|
## Choosing a background model
|
|
291
301
|
|
|
@@ -294,9 +304,11 @@ does not need to be clever. A small non-reasoning model is the right choice, and
|
|
|
294
304
|
Clio always requests the memory route with thinking off. Version 2 therefore
|
|
295
305
|
has no configurable memory thinking-level key.
|
|
296
306
|
|
|
297
|
-
|
|
307
|
+
The off request depends on the resolved runtime and model metadata:
|
|
298
308
|
llama.cpp reads `chat_template_kwargs.enable_thinking`, and LM Studio reads
|
|
299
|
-
`reasoning_effort`,
|
|
309
|
+
`reasoning_effort`, with the off value selected by the model family. A gateway
|
|
310
|
+
needs a recognized, unanimous upstream runtime declaration; unknown metadata
|
|
311
|
+
does not establish a dialect. Requesting off does not prove server compliance.
|
|
300
312
|
|
|
301
313
|
A model that reasons anyway still works. Some genuinely cannot be silenced, and
|
|
302
314
|
the catalog records those as always-on so the level reads `forced` rather than
|
|
@@ -306,12 +318,11 @@ blocks are discarded and only the envelope is kept, and the output budget is
|
|
|
306
318
|
sized to let a reasoning preamble run its course first. The cost is latency,
|
|
307
319
|
which the detached step absorbs.
|
|
308
320
|
|
|
309
|
-
|
|
310
|
-
|
|
311
|
-
|
|
312
|
-
|
|
313
|
-
|
|
314
|
-
actually asked for.
|
|
321
|
+
The active chat model may also serve memory, including a reasoning model, when
|
|
322
|
+
request capacity permits. Memory still requests thinking off and uses its own
|
|
323
|
+
bounded envelope and output budget. A dedicated small model is preferred for
|
|
324
|
+
latency and resource use; sharing a model is a capacity decision rather than an
|
|
325
|
+
unconditional refusal.
|
|
315
326
|
|
|
316
327
|
This is a mix-and-match plane, not a local-only one. The background role resolves
|
|
317
328
|
through the same target machinery as every other role, so the useful shapes are:
|
|
@@ -356,30 +367,30 @@ independent of the chat target and the fleet default. A running session owns
|
|
|
356
367
|
its routing snapshot, while the saved
|
|
357
368
|
selection becomes the default for new sessions.
|
|
358
369
|
|
|
359
|
-
|
|
370
|
+
An example gateway topology keeps the three backend routes distinct:
|
|
360
371
|
|
|
361
372
|
| Role | Target | Runtime and endpoint | Model | Capacity |
|
|
362
373
|
| --- | --- | --- | --- | --- |
|
|
363
|
-
| Chat | `dynamo` |
|
|
364
|
-
|
|
|
374
|
+
| Chat and memory fallback | `dynamo` | LiteLLM at `http://gateway:4000` | `dynamo/qwen3.8-27b` | Gateway-owned unless an explicit or observed endpoint bound is available |
|
|
375
|
+
| Preferred background memory | `zbook` | Same LiteLLM gateway | `zbook/ornith-1.5-35b-a3b` | Same capacity evidence rules |
|
|
376
|
+
| Alternative memory candidate | `mini` | Same LiteLLM gateway | `mini/ornith1.5-35b-moe` | Same capacity evidence rules |
|
|
365
377
|
|
|
366
|
-
The corresponding
|
|
378
|
+
The corresponding role selection is:
|
|
367
379
|
|
|
368
380
|
```yaml
|
|
369
381
|
chat:
|
|
370
382
|
target: dynamo
|
|
371
|
-
model: qwen3.8-27b
|
|
383
|
+
model: dynamo/qwen3.8-27b
|
|
372
384
|
context:
|
|
373
385
|
memory:
|
|
374
|
-
target:
|
|
375
|
-
model:
|
|
386
|
+
target: zbook
|
|
387
|
+
model: zbook/ornith-1.5-35b-a3b
|
|
376
388
|
```
|
|
377
389
|
|
|
378
|
-
|
|
379
|
-
|
|
380
|
-
|
|
381
|
-
|
|
382
|
-
effect on the operator's turn.
|
|
390
|
+
These are example configured target/model IDs, not built-in routes. Compare a
|
|
391
|
+
slow dedicated model with an alternative using actual step latency, valid bank
|
|
392
|
+
writes and known usage; the active chat route remains the fallback. Do not create
|
|
393
|
+
aliases or change endpoint URLs just to bypass capacity accounting.
|
|
383
394
|
|
|
384
395
|
The deadline is a bound on what an optional call may hold that server for, not a
|
|
385
396
|
figure sized to capture the tail. The shipped `60000` is a bounded compromise,
|
|
@@ -489,7 +500,7 @@ clear the prior heap bank before the new session can observe it.
|
|
|
489
500
|
|
|
490
501
|
## Telemetry
|
|
491
502
|
|
|
492
|
-
Each completed memory
|
|
503
|
+
Each completed memory-policy attempt appends one content-free record to:
|
|
493
504
|
|
|
494
505
|
```text
|
|
495
506
|
<stateDir>/memory/steps.jsonl
|
|
@@ -512,7 +523,7 @@ folds out of this file.
|
|
|
512
523
|
`dropped` is the one outcome that ran no step. It has two causes, separated by
|
|
513
524
|
the row's reason: `step_in_flight` means the boundary triggered while an earlier
|
|
514
525
|
step still held the single in-flight slot, and `endpoint_busy` means the step
|
|
515
|
-
|
|
526
|
+
found the known endpoint request capacity exhausted. Both cost
|
|
516
527
|
no tokens and no latency, both leave their triggers pending for the next free
|
|
517
528
|
boundary, and neither replaces the operator-visible last decision. Counting
|
|
518
529
|
`dropped` rows against `llm` rows over a session is how a starved cadence becomes
|
|
@@ -5,6 +5,14 @@
|
|
|
5
5
|
|
|
6
6
|
The Skills Hub (`/skill`) shows project skills, user skills, and the marketplace. Every marketplace row comes from the same local lookup that `clio-coder skills install <name>` and `/skill <name>` resolve through, so the hub lists nothing it cannot install.
|
|
7
7
|
|
|
8
|
+
## Operator ownership of installed skills
|
|
9
|
+
|
|
10
|
+
The active project `.clio-coder/skills/` tree and the resolved user `<configDir>/skills/` tree are operator-owned. Main-agent and worker tool admissions refuse writes, edits, artifacts and recognized shell mutations to either tree, including ancestor deletion and current symlink aliases. This boundary stays active at every autonomy level, even when the general default path policy is disabled. Reads and loading already installed skills retain their existing rules.
|
|
11
|
+
|
|
12
|
+
Draft a proposed skill outside these active trees, for example in `draft-skills/`. The operator installs or updates it through `clio-coder skills install`, `skills update`, `skills sync`, `library add --yes`, the Skills Hub, `/skill <name>`, or an explicitly accepted marketplace offer. Even `full-auto` must wait for that offer's bound operator answer; a task match alone no longer installs a skill. Model shell calls to recognizable Clio skill installation, update and sync commands are refused. Confirmed `library add --yes` commands are also reserved for operators, including additions of other resource kinds that may install skill dependencies. An unconfirmed `library add` only prints a plan and remains available, as do inventory, search, inspection and validation commands.
|
|
13
|
+
|
|
14
|
+
This is a tool-admission boundary, not an operating-system sandbox. Shell inspection covers literal paths and recognized command forms, respecting comments, quoted words and redirection operands when identifying Clio commands. Quoted search patterns remain read-only, and literal operator filenames do not hide later mutation operands. It cannot prove the effects of arbitrary scripts, dynamically constructed paths, aliases, or concurrent filesystem changes. Keep shell execution supervised when stronger confinement is required.
|
|
15
|
+
|
|
8
16
|
## Where marketplace rows come from
|
|
9
17
|
|
|
10
18
|
There is one source, `discoverMarketplaceSkills()` in `src/domains/resources/skills/marketplace.ts`, and it reads two kinds of real local data:
|
|
@@ -62,6 +70,16 @@ creation. Extension resource roots and share archives are documented in
|
|
|
62
70
|
[extensions-and-sharing.md](extensions-and-sharing.md); this page owns the TUI
|
|
63
71
|
Hub and marketplace behavior.
|
|
64
72
|
|
|
73
|
+
## Remote entries and overlays
|
|
74
|
+
|
|
75
|
+
Some skills are worth carrying in the catalog without vendoring their content. `skills/remote.yaml` lists them: each entry names a skill, its category, a `sourceUrl` that must be a GitHub tree URL at a pinned tag, an `overlay` package inside the catalog, and an optional `exclude` list of upstream top-level members. `npm run skills:pin` publishes such an entry into `skills/skill-marketplace.json` with `origin: "remote"`, the upstream URL as its `sourceUrl`, and the `overlay` and `exclude` fields attached. The overlay's `SKILL.md` is pinned in `skills/registry.yaml` like every other catalog skill, so `npm run skills:check` fails when it drifts.
|
|
76
|
+
|
|
77
|
+
`archify` is the worked example. Its entry points at `https://github.com/tt-a1i/archify/tree/v2.16.0/archify`, overlays `skills/planning/archify`, and excludes `test` and `package-lock.json`. Running `clio-coder skills install archify --project` clones that tag, drops the excluded members, copies the overlay over the clone so Clio's wrapper `SKILL.md` replaces the upstream one, validates the shaped tree, and swaps it into `.clio-coder/skills/archify/` with the usual provenance stamps. The renderer, its schemas, and its brand-mark notices come from upstream at install time and never enter the npm tarball. The wrapper omits upstream's update-awareness step on purpose: that step performs a network request during a chat turn, and no Clio chat turn depends on the network.
|
|
78
|
+
|
|
79
|
+
Discovery treats the overlay folder as part of the remote entry rather than as a skill of its own, so a bare-name install never lands the wrapper without the renderer it wraps. `clio-coder skills update` recovers the same overlay and exclude list from the marketplace, so an update refetches the pinned upstream and re-applies the wrapper instead of replacing it.
|
|
80
|
+
|
|
81
|
+
Two operational notes. A copy dropped into a project-scope `.claude/skills/archify` is a compat import and stays untrusted until `integrations.projectResources.trustProjectImports` is on; install through Clio to get a trusted, provenance-stamped copy. And since the skill's authoring loop is several `bash` calls (validate, deliver, verify), `auto-edit` is the sensible permission posture for a mapping session; full manual approval works but prompts on every command.
|
|
82
|
+
|
|
65
83
|
## Publishing a skill
|
|
66
84
|
|
|
67
85
|
Add a directory under `skills/<category>/<name>/` (or `skills/<name>/`) in the repo containing a `SKILL.md` with `name` and `description` frontmatter. The directory name must match `[A-Za-z0-9][A-Za-z0-9._-]*`. Run `npm run skills:pin` to republish `skills/skill-marketplace.json`, which is the index consumers point `CLIO_CODER_SKILL_MARKETPLACE_INDEX` at or copy to `<configDir>/skill-marketplace.json`. Scientific and niche coding domains are the marketplace's focus; see the existing `skills/` tree for the house format.
|
package/docs/guide/tool-usage.md
CHANGED
|
@@ -311,7 +311,7 @@ Arguments:
|
|
|
311
311
|
Projects may commit a versioned executable catalog at `.clio-coder/verifiers.yaml`:
|
|
312
312
|
|
|
313
313
|
```yaml
|
|
314
|
-
version:
|
|
314
|
+
version: 2
|
|
315
315
|
checks:
|
|
316
316
|
- id: rust-workspace
|
|
317
317
|
description: Run the Rust workspace tests
|
|
@@ -319,9 +319,29 @@ checks:
|
|
|
319
319
|
cwd: .
|
|
320
320
|
timeoutMs: 600000
|
|
321
321
|
tags: [rust, test]
|
|
322
|
+
- id: grid-metadata
|
|
323
|
+
description: Compare the regional grid statistics against the reference
|
|
324
|
+
kind: numeric-compare
|
|
325
|
+
command: [python, tools/grid_stats.py, out/region_west.nc]
|
|
326
|
+
reference: tests/reference/region_west.json
|
|
327
|
+
tolerance: { relative: 1.0e-6, ulp: 4 }
|
|
328
|
+
cwd: .
|
|
329
|
+
timeoutMs: 120000
|
|
330
|
+
tags: [scientific, netcdf]
|
|
331
|
+
- id: solver-time
|
|
332
|
+
description: Keep the solver inside its wall-time budget
|
|
333
|
+
kind: perf-budget
|
|
334
|
+
command: [python, tools/solve.py, --small]
|
|
335
|
+
baseline: .clio-coder/baselines/solver-time.json
|
|
336
|
+
tolerance: { relative: 0.25 }
|
|
337
|
+
cwd: .
|
|
338
|
+
timeoutMs: 600000
|
|
339
|
+
tags: [scientific, performance]
|
|
322
340
|
```
|
|
323
341
|
|
|
324
|
-
|
|
342
|
+
Every check has a `kind`, absent or `command` by default. A version 1 file still loads and every check there is `kind: command`; the kind fields require `version: 2`. `kind: command` reads the exit code. `kind: numeric-compare` runs the command, parses its stdout as a JSON object of `string -> number | number[]`, and judges it against `reference` (a repository-relative JSON file of the same shape) under `tolerance`, which names at least one of `relative`, `absolute`, or `ulp`; a value passes only when every named tolerance holds, a key missing on either side fails with the key named, arrays compare elementwise and fail on length mismatch, and `NaN` or infinity fails. `kind: perf-budget` runs the command and judges the wall time the harness measured against either `budget: {wallTimeMs, tolerance?: {relative}}` or `baseline`, a repository-relative JSON `{wallTimeMs}` that `clio-coder verifiers baseline <id>` records from one clean run, with an optional `tolerance: {relative}` of headroom over it. Exactly one of `budget` and `baseline` is present. A command that exits non-zero, times out, or is aborted fails before any judgement. Both kinds record a structured `report` on the `verify` result details and on the host-verification check of a dispatch receipt (per-key worst deviation and the failed tolerance, or measured time, effective budget, and ratio); a failing judgement is a check failure, not a new evidence category.
|
|
343
|
+
|
|
344
|
+
Version 2 keeps version 1's strictness. Every root and check field shown above is required, unknown fields fail, and duplicate IDs fail. A project ID uses lowercase letters, digits, `.`, `_`, `:`, or `-`, begins with a letter or digit, and is at most 64 UTF-8 bytes. `frontend` is reserved. Descriptions are trimmed single-line text capped at 512 bytes. `command` is a nonempty argv array with at most 64 entries and 4096 bytes per entry. A shell command string is invalid, and explicit shell executables such as `sh`, `bash`, `pwsh`, and `cmd` are rejected. `cwd` is a repository-relative existing directory capped at 512 bytes; absolute paths, `..` escapes, and symbolic-link escapes fail. `timeoutMs` is a positive integer capped at 900000. A check may carry at most 16 distinct lowercase tags of at most 32 bytes each. The whole file is capped at 262144 bytes and may contain at most 128 checks. YAML aliases are disabled.
|
|
325
345
|
|
|
326
346
|
Provider IDs share one namespace. If a catalog ID collides with a discovered package script, listing and execution fail and identify both source files. Catalog parsing also fails closed before any package or project check runs.
|
|
327
347
|
|
|
@@ -348,8 +368,11 @@ clio-coder verifiers author
|
|
|
348
368
|
clio-coder verifiers author --exclude cmake-build-debug --rename go-test=go-suite
|
|
349
369
|
clio-coder verifiers author --dry-run go-suite --yes
|
|
350
370
|
clio-coder verifiers validate
|
|
371
|
+
clio-coder verifiers baseline solver-time
|
|
351
372
|
```
|
|
352
373
|
|
|
374
|
+
`author` also lists one incomplete `numeric-compare` check for every validation-contract artifact that declares `numerical_tolerances`, with the reference and tolerance filled in and the exact `verifiers add` line to complete; the command is the operator's to supply, so the proposal never enters the catalog on its own. `baseline <id>` runs a `perf-budget` check once and writes its wall time to the check's `baseline` path; a failing or timed-out command records nothing.
|
|
375
|
+
|
|
353
376
|
`validate` reads the committed file with the same parser used by `verify()`. `dry-run <id>` is an explicit request to execute one admitted check through the production `verify` path. `author --dry-run <id> --yes` writes only after confirmation and starts the selected dry run only after the write is accepted by production discovery.
|
|
354
377
|
|
|
355
378
|
Later changes use the same preview and confirmation boundary. `edit` preserves the ID unless `rename` is requested. Renames and additions reject collisions with catalog IDs and active package-script IDs. Removals state that the deleted command will no longer be executable through catalog authority. Generated IDs are stable for a stable ordered signal set; a collision receives the first available deterministic `-2`, `-3`, and later suffix.
|
|
@@ -362,7 +385,7 @@ clio-coder verifiers rename validate-grid validate-regional-grid --yes
|
|
|
362
385
|
clio-coder verifiers remove validate-regional-grid --yes
|
|
363
386
|
```
|
|
364
387
|
|
|
365
|
-
The `add` command is the explicit path for an unsupported or ambiguous project. `--command` must be a JSON argv array, so manual entry still cannot turn a shell command string into executable catalog authority.
|
|
388
|
+
The `add` command is the explicit path for an unsupported or ambiguous project. `--command` must be a JSON argv array, so manual entry still cannot turn a shell command string into executable catalog authority. `--kind numeric-compare` takes `--reference <path>` and `--tolerance '<json>'`; `--kind perf-budget` takes either `--budget-ms <n>` with optional `--budget-relative <r>` or `--baseline <path>` with optional `--tolerance '{"relative": r}'`. `edit` accepts the same options to change a check's kind.
|
|
366
389
|
|
|
367
390
|
`verify(check="frontend", path=<file>)` validates an HTML, CSS, or JavaScript artifact without shell access. The path must stay inside the workspace root and end in `.html`, `.htm`, `.css`, `.js`, `.mjs`, or `.cjs`. Checks per type: HTML tag balance (comment-aware, HTML5 optional end tags honored), inline and referenced script syntax (classic scripts parsed in-process, modules via `node --check`), inline and linked CSS brace/string/comment balance, local script and stylesheet references resolved and existence-checked (external and root-relative references are skipped), and an optional headless browser load. `browser="auto"` warns when no chromium/chrome/edge executable is on PATH, `"required"` fails, `"off"` skips. Each check reports pass, warn, fail, or skip; any fail makes the whole result an error. `details = {action: "verify", check: "frontend", path, browserMode, status, checks}`.
|
|
368
391
|
|
|
@@ -627,6 +650,58 @@ panes(action="open", preset="logs")
|
|
|
627
650
|
panes(action="close", target="all")
|
|
628
651
|
```
|
|
629
652
|
|
|
653
|
+
## evidence: inspect canonical evidence and trust status
|
|
654
|
+
|
|
655
|
+
Reads evidence bundles as JSON. Source: `src/tools/evidence.ts`. Read class; sequential, because `run` mode may materialize a bundle under Clio's data directory. It shares the inventory and trust projections behind `clio-coder evidence inventory` and `clio-coder evidence inspect`, so the model and the operator read the same record.
|
|
656
|
+
|
|
657
|
+
Arguments:
|
|
658
|
+
|
|
659
|
+
- `mode` (required). `list`, `inspect`, or `run`.
|
|
660
|
+
- `id` (required for `inspect`). An evidence bundle id.
|
|
661
|
+
- `runId` (required for `run`). A dispatch run id; the bundle is built first when none exists.
|
|
662
|
+
|
|
663
|
+
`list` returns the bounded newest-first inventory: provenance, tags, totals, and a worst-run trust verdict per bundle. `inspect` returns the bundle overview, the per-run trust axes and verdict, the gate decisions, and the findings. `run` resolves `run-<runId>` and builds the bundle when it is absent; a run with no ledger row is reported absent with `artifactAbsent: true` in the details. Results are capped at 16KB, and a truncated result stays valid JSON with a `preview`. Provenance requires this tool and Verifier may use it.
|
|
664
|
+
|
|
665
|
+
```text
|
|
666
|
+
evidence(mode="list")
|
|
667
|
+
evidence(mode="inspect", id="run-r-42")
|
|
668
|
+
evidence(mode="run", runId="r-42")
|
|
669
|
+
```
|
|
670
|
+
|
|
671
|
+
## limitation: record what a turn could not verify
|
|
672
|
+
|
|
673
|
+
Records a typed limitation receipt for the finish contract. Source: `src/tools/limitation.ts`. Read class; parallel. The tool is pure: it touches no filesystem and runs no shell, so the successful receipt in the session ledger is its whole effect.
|
|
674
|
+
|
|
675
|
+
Arguments:
|
|
676
|
+
|
|
677
|
+
- `scope` (required). What could not be verified, in one sentence.
|
|
678
|
+
- `reason` (required). `no-runner`, `blocked`, `out-of-scope`, `environment`, or `other`.
|
|
679
|
+
- `paths` (optional). Repository-relative paths left unverified.
|
|
680
|
+
|
|
681
|
+
Call it once, before the final reply, when files changed and validation could not run. The finish contract accepts a successful `limitation` receipt inside the same window as the mutation scan in place of validation evidence. A rejected call (empty scope, unknown reason) leaves no receipt and does not count, and the assistant's prose never does. The six mutating recipes carry the tool and the operating contract tells the model to call it; see [the finish gate](../architecture/safety-model.md#the-finish-gate-and-re-prompt-behavior).
|
|
682
|
+
|
|
683
|
+
```text
|
|
684
|
+
limitation(scope="CUDA kernels changed but no GPU is available here", reason="environment", paths=["src/kernels/solve.cu"])
|
|
685
|
+
```
|
|
686
|
+
|
|
687
|
+
## decide: record a design decision
|
|
688
|
+
|
|
689
|
+
Appends the model's own design choice to the session decision board beside operator `ask_user` answers. Source: `src/tools/decide.ts`. Read class; sequential, so two decisions in one batch cannot race the supersede lookup. The call succeeds only in a session with a decision board; a worker's call is refused.
|
|
690
|
+
|
|
691
|
+
Arguments:
|
|
692
|
+
|
|
693
|
+
- `key` (required). Stable kebab-case name, at most 64 bytes.
|
|
694
|
+
- `value` (required). The option chosen, at most 512 bytes.
|
|
695
|
+
- `alternatives` (required). One to six rejected options, at most 256 bytes each.
|
|
696
|
+
- `rationale` (required). Why the choice won, at most 1024 bytes.
|
|
697
|
+
- `label` (optional). Short title, at most 128 bytes.
|
|
698
|
+
|
|
699
|
+
The call appends one `decisionLedger` entry with `origin: "agent"` and returns the decision ref `<interviewId>/<key>`. A repeat key supersedes the earlier agent decision with the new rationale as its correction; an operator decision with the same key is never overwritten and the call fails. Dispatch seals every active ref onto the run request, envelope, and receipt, and Clio-controlled commits carry one `Clio-Decision:` trailer per ref; see [commit provenance](../process/git-commit-provenance.md).
|
|
700
|
+
|
|
701
|
+
```text
|
|
702
|
+
decide(key="cache-key-shape", value="capability tuple", alternatives=["node id"], rationale="matches the existing buckets and survives fleet changes", label="Cache key")
|
|
703
|
+
```
|
|
704
|
+
|
|
630
705
|
## ask_user: host-owned operator interviews
|
|
631
706
|
|
|
632
707
|
Runs a host-owned interactive interview or single-question prompt with the operator, recording decisions and/or free-form answers. Source: `src/tools/ask-user.ts`. Read class; sequential.
|
|
@@ -34,7 +34,7 @@ Every runtime-tunable value needs both halves; the pair is one knob, not two. Th
|
|
|
34
34
|
| `CLIO_CODER_MAX_TOOL_CALLS` | 50 | `src/engine/loop-guard.ts` → `src/engine/worker-runtime.ts` | Worker lifetime tool-call cap for a dispatched run. Different axis than the orchestrator budget despite the near-identical name. |
|
|
35
35
|
| `CLIO_CODER_MAX_DISPATCH_RUNS` | 1000 | `src/domains/dispatch/state.ts` | Dispatch run-ledger retention cap. |
|
|
36
36
|
| `CLIO_CODER_MAX_CONTEXT_TOKENS` | unset | `src/domains/providers/runtime-resolution.ts` | Context-window override for local runtimes. Also set internally by `clio-coder run --max-context-tokens` (see §6). |
|
|
37
|
-
| `CLIO_CODER_KV_CACHE_MODE` | unset | retired | KV-cache quantization mode. Also set internally by the former `clio-coder run
|
|
37
|
+
| `CLIO_CODER_KV_CACHE_MODE` | unset | retired | KV-cache quantization mode. Also set internally by the former `clio-coder run` path. |
|
|
38
38
|
| `CLIO_CODER_SAMPLING_OVERRIDES` | unset | `src/engine/apis/sampling-overrides.ts` | JSON sampling-parameter override. Set internally by print-mode sampling flags. |
|
|
39
39
|
| `CLIO_CODER_READ_MAX_BYTES` | 51200 (50 KB) | `src/tools/read.ts` | Per-call byte cap for the read tool. |
|
|
40
40
|
| `CLIO_CODER_OBSERVATION_TURN_BUDGET_BYTES` | 196608 (192 KB) | `src/tools/observation.ts` | Shared per-turn byte pool across all observation tools. |
|
|
@@ -93,7 +93,7 @@ All default off; all enabled with `1`.
|
|
|
93
93
|
|
|
94
94
|
## 6. CLI flags that bridge through env vars (Pre-consolidated State)
|
|
95
95
|
|
|
96
|
-
`clio-coder run --max-context-tokens`
|
|
96
|
+
`clio-coder run --max-context-tokens` (`src/cli/run.ts:143-234`) and the print-mode sampling flags (`src/cli/modes/print.ts:306-333`) do not plumb their values through function arguments. They mutate `process.env` (`CLIO_CODER_MAX_CONTEXT_TOKENS`, `CLIO_CODER_KV_CACHE_MODE`, `CLIO_CODER_SAMPLING_OVERRIDES`), run the command, then restore the previous value in a `finally`. The env var is the transport between the CLI layer and deep engine code.
|
|
97
97
|
|
|
98
98
|
## 7. Script- and benchmark-only vars (Pre-consolidated State)
|
|
99
99
|
|
|
@@ -20,12 +20,13 @@ under pressure; git mechanics alone never justify a stage.
|
|
|
20
20
|
| --- | --- | --- |
|
|
21
21
|
| 1. File | [`file-ticket`](../../skills/git/file-ticket/) | A labeled GitHub issue with evidence and acceptance criteria |
|
|
22
22
|
| 2. Fix | [`fix-issue`](../../skills/git/fix-issue/) | An uncommitted, verified change where failing tests preceded the fix, self-reviewed against the issue's acceptance criteria |
|
|
23
|
-
| 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`)
|
|
23
|
+
| 3. Ship | [`ship`](../../skills/git/ship/) | An atomic conventional commit referencing the issue (`fixes #N`); contributors push it to their fork and open a PR, while maintainer work stays local for gated integration; merge is a human decision |
|
|
24
24
|
|
|
25
|
-
Releases follow [release-cut-checklist.md](
|
|
25
|
+
Releases follow [release-cut-checklist.md](release-cut-checklist.md) as a
|
|
26
26
|
human-gated checklist, not a skill. Worktrees
|
|
27
27
|
([`worktree-create`](../../skills/git/worktree-create/),
|
|
28
28
|
[`worktree-merge`](../../skills/git/worktree-merge/)),
|
|
29
|
+
[`branch-closeout`](../../skills/git/branch-closeout/),
|
|
29
30
|
[`resolve-merge-conflicts`](../../skills/git/resolve-merge-conflicts/), and
|
|
30
31
|
[`tdd`](../../skills/coding/tdd/) are à-la-carte tools reached for when the
|
|
31
32
|
situation calls for them, not stages every change passes through. An RCA
|
|
@@ -34,6 +35,38 @@ hard bugs, not a mandatory toll booth. Batch ticket creation from a PRD
|
|
|
34
35
|
bypasses stage 1 and uses [`backlog`](../../skills/planning/backlog/)
|
|
35
36
|
instead; everything downstream is identical.
|
|
36
37
|
|
|
38
|
+
## Closeout
|
|
39
|
+
|
|
40
|
+
A merged PR is not operationally finished until its local scaffolding is
|
|
41
|
+
closed. The reusable [`branch-closeout`](../../skills/git/branch-closeout/) skill automates this verification and teardown safely. After the human merge decision:
|
|
42
|
+
|
|
43
|
+
1. Fetch and prune, confirm the PR's merged state, and identify the resulting
|
|
44
|
+
commit on `origin/main`. Direct ancestry proves an ordinary merge; a squash
|
|
45
|
+
or cherry-pick needs the PR-to-result evidence because commit identity and
|
|
46
|
+
patch identity can both change during integration.
|
|
47
|
+
2. Inspect every associated worktree for tracked changes, untracked files, and
|
|
48
|
+
ignored state that carries evidence rather than rebuildable output. Remove
|
|
49
|
+
it through `git worktree remove`; forcing removal requires explicit approval
|
|
50
|
+
to discard what remains.
|
|
51
|
+
3. Delete the local source and integration branches. For a contributor PR,
|
|
52
|
+
delete the merged branch from the contributor's fork. The canonical
|
|
53
|
+
repository never hosts topic, integration, or release-candidate branches.
|
|
54
|
+
4. Turn unfinished experimental findings into an issue with evidence and a
|
|
55
|
+
next decision. Do not use indefinite `work/`, `wip/`, `keep/`, or temporary
|
|
56
|
+
tags as a substitute for backlog state.
|
|
57
|
+
5. Report the remaining worktrees, local branches, stashes, local-only tags,
|
|
58
|
+
and canonical remote heads. The expected canonical head set is exactly
|
|
59
|
+
`refs/heads/main`; every survivor needs an owner and purpose.
|
|
60
|
+
|
|
61
|
+
Maintainer release candidates are local-only and use a compact branch name
|
|
62
|
+
that cannot collide with their tag: branch `v043`, tag `v0.4.3`. Gate the exact
|
|
63
|
+
candidate, require fetched `origin/main` to be its ancestor, fast-forward local
|
|
64
|
+
`main`, fetch again, and push only `refs/heads/main:refs/heads/main` with
|
|
65
|
+
explicit authorization. After CI passes, push only the fully qualified
|
|
66
|
+
annotated tag. Once the tag's peeled commit equals the reviewed commit on
|
|
67
|
+
`main` and the release succeeds, delete the local candidate branch. Published
|
|
68
|
+
dotted release tags are immutable history and are never cleanup targets.
|
|
69
|
+
|
|
37
70
|
## Inheriting a Pi release
|
|
38
71
|
|
|
39
72
|
Pi dependency upgrades use a fixed five-step review so that upstream fixes
|
|
@@ -73,6 +106,11 @@ There is no committed weighted-shard or special serial-lane runner. Keep timing
|
|
|
73
106
|
claims within the focused contract that owns them, and use the full `npm run ci`
|
|
74
107
|
gate before handoff.
|
|
75
108
|
|
|
109
|
+
The release smoke script `scripts/smoke-real-home.sh` (invoked via
|
|
110
|
+
`npm run smoke:real-home`) tests booting the built CLI binary against a copy of
|
|
111
|
+
the operator settings in a scratch home. An optional `--strict` flag makes doctor
|
|
112
|
+
exit 1 fail the smoke run on failing rows instead of tolerating fleet state.
|
|
113
|
+
|
|
76
114
|
## Issue conventions
|
|
77
115
|
|
|
78
116
|
- **Title**: conventional tag plus imperative summary (`fix: memory overlay
|
|
@@ -139,7 +139,37 @@ Metrics collected during runs can be validated automatically using the `verify.a
|
|
|
139
139
|
* `eq` (equal)
|
|
140
140
|
* `neq` (not equal)
|
|
141
141
|
|
|
142
|
-
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, and `
|
|
142
|
+
Metrics that can be validated include `tokens.input`, `tokens.output`, `tokens.total`, `latency.wallMs`, `tools.totalCalls`, `tools.failed`, `tools.blocked`, `verifier.exitCode`, `result.pass`, and the `provider.*` metrics below. Each `verify.assertions` condition must hold; an unavailable metric fails closed.
|
|
143
|
+
|
|
144
|
+
### Optional Provider-Health Gates
|
|
145
|
+
|
|
146
|
+
A task that recovers from a provider error can still pass its task checks by default. Provider health is a separate, opt-in requirement. To require observed provider events and no observed terminal errors, add these assertions to the task:
|
|
147
|
+
|
|
148
|
+
```yaml
|
|
149
|
+
verify:
|
|
150
|
+
assertions:
|
|
151
|
+
- metric: "provider.measured"
|
|
152
|
+
op: "eq"
|
|
153
|
+
value: true
|
|
154
|
+
- metric: "provider.stopReason.error"
|
|
155
|
+
op: "eq"
|
|
156
|
+
value: 0
|
|
157
|
+
```
|
|
158
|
+
|
|
159
|
+
Suite-level `thresholds.fail` uses the opposite condition: a matching condition is a failure. The corresponding hard gate checks each run as follows:
|
|
160
|
+
|
|
161
|
+
```yaml
|
|
162
|
+
thresholds:
|
|
163
|
+
fail:
|
|
164
|
+
- metric: "provider.measured"
|
|
165
|
+
op: "eq"
|
|
166
|
+
value: false
|
|
167
|
+
- metric: "provider.stopReason.error"
|
|
168
|
+
op: "gt"
|
|
169
|
+
value: 0
|
|
170
|
+
```
|
|
171
|
+
|
|
172
|
+
Unavailable metrics also fail closed in a hard threshold. `thresholds.informational` records findings without changing exit status. These examples reject an observed error followed by a successful recovery while leaving the default ungated task-pass policy unchanged. To reject any observed retry start or assistant abort as well, add conditions on `provider.retryStarted` or `provider.stopReason.aborted` with the same assertion-versus-failure polarity.
|
|
143
173
|
|
|
144
174
|
---
|
|
145
175
|
|
|
@@ -175,14 +205,44 @@ Under the hood, these are parsed and wrapped into a Suite v2 adapter with:
|
|
|
175
205
|
|
|
176
206
|
Clio maintains two distinct token accounting streams with different provenances. These accounts are never merged, reconciled, or treated as interchangeable:
|
|
177
207
|
|
|
178
|
-
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events
|
|
208
|
+
1. **`tokens.*` (Wire Streaming)**: Folded live off stdout from assistant `message_end` events. These totals include all known usage on errored calls as well as successful calls; recovery never subtracts earlier spend. Only finite, nonnegative usage facts are admitted. On surfaces without the relevant stdout events (such as `clio-coder fleet run --json`), `tokens.measured` is `false`.
|
|
179
209
|
2. **`receiptUsage.*` (Journal Receipts)**: Summed from an evaluation item's run journal. Every attempt writes a receipt carrying token counts and USD cost authenticated against its own ledger envelope.
|
|
180
210
|
|
|
181
211
|
### Fail-Closed Reporting
|
|
182
212
|
Both accounting streams report unmeasured state with no counts at all rather than a numeric zero. Reporting zero for an unmeasured run would falsely claim the run cost nothing. On an unmeasured run, `tokens.total` resolves to `null` and fails closed on metric threshold comparisons.
|
|
183
213
|
|
|
214
|
+
An errored call must carry at least one positive reported token, reasoning, or cost fact before its usage is considered observed. Reasoning-only or cost-only observations do not establish ordinary token totals. A stream containing only errored calls with missing or synthetic all-zero usage remains `tokens.measured: false`. The current event shape cannot distinguish synthetic all-zero failures from genuinely reported zero usage, so it cannot establish measured zero spending in either case. Partial positive usage remains included as known spend. On failed calls, adapters can also initialize individual absent fields to zero; those zeros remain unattributed and make coverage incomplete. The existing inclusive numeric fields are known subtotals, so their zeros do not prove complete zero spending when failed usage is incomplete.
|
|
215
|
+
|
|
216
|
+
### Provider Observations and Failed-Call Share
|
|
217
|
+
|
|
218
|
+
The `provider.*` metrics describe events observed on live stdout, folded before diagnostic output is truncated. Native runs and multi-command external runners retain these observations from their executed commands. They do not enumerate SDK-internal retries or network attempts that were never emitted. Filtered or opaque output can leave provider health unobserved even when the process exits successfully or a receipt reports task success.
|
|
219
|
+
|
|
220
|
+
| Metric | Meaning |
|
|
221
|
+
| --- | --- |
|
|
222
|
+
| `provider.measured` | Whether an assistant terminal reason or a counted retry phase was observed. With no such observations this is `false`, and provider counters are absent. It does not certify complete provider coverage. |
|
|
223
|
+
| `provider.stopReason.stop`, `provider.stopReason.toolUse`, `provider.stopReason.length` | Counts of these terminal reasons on assistant `message_end` events. |
|
|
224
|
+
| `provider.stopReason.error`, `provider.stopReason.aborted`, `provider.stopReason.other` | Separate counts of errored, aborted, and other observed terminal reasons. An unrecognized terminal reason goes into `other`. Partial updates and repeated messages in `turn_end` or `agent_end` do not add counts. |
|
|
225
|
+
| `provider.retryScheduled` | Observed `scheduled` phases: planned retries, including ones cancelled before execution. |
|
|
226
|
+
| `provider.retryStarted` | Observed `retrying` phases: retry execution starts. Repeated `waiting` countdown frames do not count as attempts. |
|
|
227
|
+
| `provider.retryCancelled`, `provider.retryExhausted`, `provider.retryRecovered` | Counts of the corresponding observed phases. They describe retry-chain outcomes and do not fabricate additional assistant calls. Attempt numbers can restart for each chain. |
|
|
228
|
+
| `provider.errorUsageObservedCalls` | Errored calls with at least one positive reported token, reasoning, or cost fact. |
|
|
229
|
+
| `provider.errorUsageUnobservedCalls` | Errored calls with no positive reported token, reasoning, or cost fact, including absent or all-zero usage. |
|
|
230
|
+
| `provider.errorUsageIncompleteCalls` | Errored calls with unobserved usage, incomplete token fields, or ambiguous normalized zero fields. This can overlap `errorUsageObservedCalls` when only part of the usage is known. |
|
|
231
|
+
| `provider.errorCostUnobservedCalls` | Errored calls without a positive cost fact. Zero or absent cost does not prove that an error was free, including when some token usage is known. |
|
|
232
|
+
| `provider.errorTokens.input`, `provider.errorTokens.output`, `provider.errorTokens.total`, `provider.errorTokens.cacheRead`, `provider.errorTokens.cacheWrite` | Known positive token subtotals for errored calls. Absent or ambiguous zero fields remain unattributed; when total usage is absent, a total can still be summed from known token fields and remains incomplete. |
|
|
233
|
+
| `provider.errorCostUsd` | Known positive cost subtotal from errored calls' stream usage objects. Cost can come from adapter pricing; it is not independently certified provider billing. |
|
|
234
|
+
| `provider.errorReasoningTokens`, `provider.errorReasoningUnobservedCalls` | Known failed-call reasoning subtotal and calls without attributable reported reasoning. Reasoning can overlap output, so it is never added to ordinary token totals. |
|
|
235
|
+
|
|
236
|
+
The failed-call share covers `stopReason: error`; aborted calls remain separately labeled. Failed-share amounts and usage-coverage counters appear only after an errored terminal message is observed. These share metrics supplement the inclusive `tokens.*` totals without changing the summary token shape, receipt schema, or verdict schema. Missing or partial failed usage makes them known subtotals, not a complete amount to subtract from total spend. The native runner's `cost.usd` can use receipt evidence, so equality with the stream-based `provider.errorCostUsd` is not guaranteed.
|
|
237
|
+
|
|
238
|
+
Positive reported reasoning uses the normalized `usage.reasoning` field, `reasoning_tokens`, and supported nested provider-detail fields. The root `reasoningTokens` alias can be an adapter estimate without a provenance marker, so it is left unattributed rather than promoted to reported provider usage. Normalized zero reasoning is also unattributed because adapters can fill it when provider detail is absent. A reasoning-only failure is observed but has incomplete ordinary-token coverage; no output or total is inferred from it.
|
|
239
|
+
|
|
240
|
+
These observations also do not reconcile the separate `trackedMetrics` ledger selection (#276). Tracked metrics prefer durable assistant-call facts when available, retain durable compaction and tool records, and otherwise fall back to stream calls. Artifacts expose source counts and warnings, but partial or mixed ledgers can omit stream-only calls and fork-inherited history remains unreconciled. Neither those tracked values nor the new provider counters prove complete run accounting; failed-compaction usage is retained separately in the out-of-turn usage ledger and usage report, and is not included by this eval fold.
|
|
241
|
+
|
|
184
242
|
---
|
|
185
243
|
|
|
244
|
+
Full reconciliation across session, stdout, fork and out-of-turn evidence is deferred to v0.4.5 or later. Version 0.4.3 does not add an automatic rejection of cost or efficiency comparisons merely because those sources are partial or mixed. Matching source counts do not prove complete coverage or shared call identity. Existing missing-metric, serving-configuration and execution-envelope comparison gates still apply.
|
|
245
|
+
|
|
186
246
|
## Eval Artifact Format (v4)
|
|
187
247
|
|
|
188
248
|
Evaluation artifacts use format version 4 (`EvalArtifactV4`). Summary token metrics report `measuredRuns` out of total `runs`:
|
|
@@ -297,7 +357,7 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
297
357
|
| `generatedTokens` | ledger |
|
|
298
358
|
| `reasoningTokens` | receipt; nullable, because absent and zero are different claims |
|
|
299
359
|
| `toolCalls`, `toolErrors` | ledger when present, otherwise receipt |
|
|
300
|
-
| `ttftMsFirstCall` | ledger |
|
|
360
|
+
| `ttftMsFirstCall` | ledger; nullable when first-call timing is absent |
|
|
301
361
|
| `wallClockMs` | receipt |
|
|
302
362
|
| `contextTokensAtEnd` | ledger |
|
|
303
363
|
| `compactions` | ledger |
|
|
@@ -305,6 +365,10 @@ Eleven numbers plus a reason histogram, each carrying the source it came from. `
|
|
|
305
365
|
|
|
306
366
|
A dispatched worker's receipt reports `sessionId: null` and writes no session archive, which is why the ledger source exists at all: the runner folds structured usage, backend timing, cache, and monotonic TTFT facts out of the worker's `message_end` events. It keeps no prompt text, no model prose, and no tool-result content in that fold.
|
|
307
367
|
|
|
368
|
+
First-call TTFT uses the earliest recorded assistant-call timestamp across the selected ledgers; equal timestamps retain their observed order. Missing or invalid timing, or an invalid timestamp that prevents ordering the calls, yields `{ value: null, source: "estimated" }`. A measured zero remains `{ value: 0, source: "ledger" }`. Native session timing starts at each stream invocation and includes the provider's response-header wait. Stdout-only fallback timing starts at the provider's `message_start`, which can arrive after headers; it requires first output, but is not complete request latency and must not be compared as equivalent to native invocation timing. A completion alone supplies neither a zero duration nor a first-token measurement. Verdict v1 consumers must accept nullable TTFT. Historical numeric values, including estimated zeros and native spans that omitted the pre-header wait, remain readable and are not rewritten.
|
|
369
|
+
|
|
370
|
+
When a stream message lacks a valid timestamp, its ledger payload marks `timestampEstimated: true` beside the legacy ISO placeholder. This leaves first-call chronology unmeasured while preserving any observed per-call monotonic timing.
|
|
371
|
+
|
|
308
372
|
### Scenario aggregates
|
|
309
373
|
|
|
310
374
|
`aggregates` groups verdicts by `scenarioId`, sets `k` to the trial count, and records `passAtK` (any trial passed) and `passPowK` (every trial passed). Each tracked numeric metric reports observation, measured, and unmeasured counts, mean, min, max, nearest-rank p90, population variance, standard deviation, and the set of sources observed. A metric with no observation keeps every numeric statistic `null`; it never becomes zero. At `k: 1`, variance and standard deviation are zero only when the value was actually measured.
|
|
@@ -66,6 +66,21 @@ Clio-Evidence: receipt-v20/sha256:<64-character digest>
|
|
|
66
66
|
Clio does not invent, shorten, or add an unrelated digest. The role trailers do
|
|
67
67
|
not depend on this optional line.
|
|
68
68
|
|
|
69
|
+
A commit also names the decisions it was made under, one trailer per active
|
|
70
|
+
decision on the session decision board (an `ask_user` answer or a design
|
|
71
|
+
choice the agent recorded with `decide`):
|
|
72
|
+
|
|
73
|
+
```text
|
|
74
|
+
Clio-Decision: <interviewId>/<key>
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
The refs come from the sealed receipt's `decisionRefs` at the fleet seam and
|
|
78
|
+
from the live board at the session seam. They are sorted, capped at 32, added
|
|
79
|
+
once, and only refs of the `<id>/<key>` shape are written, where the key is an
|
|
80
|
+
operator `snake_case` key or an agent `kebab-case` key. A decision trailer
|
|
81
|
+
records rationale provenance; it is not evidence that the decision was correct
|
|
82
|
+
or that its work was validated.
|
|
83
|
+
|
|
69
84
|
## Commit paths and hooks
|
|
70
85
|
|
|
71
86
|
The deterministic SDLC fleet attributes its plan, code, and documentation
|