@iowarp/clio-coder 0.3.1 → 0.3.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +233 -373
- package/CONTRIBUTING.md +23 -23
- package/README.md +284 -613
- package/dist/{acp-FPR54DGL.js → acp-P2AQILE2.js} +43 -53
- package/dist/{agents-OGPIHPJH.js → agents-72W3BI7I.js} +43 -26
- package/dist/assets/codewiki.json +1 -1
- package/dist/{auth-IC3K6NIZ.js → auth-5TWEIYDN.js} +20 -12
- package/dist/chunk-2DJ2KNFG.js +2095 -0
- package/dist/{chunk-IS3ONKU3.js → chunk-2IR2NMPA.js} +6 -4
- package/dist/chunk-2SFS6XQE.js +122 -0
- package/dist/{chunk-PV4JUBVJ.js → chunk-2TLUCQVG.js} +40 -21
- package/dist/chunk-2VTFPG5O.js +48 -0
- package/dist/chunk-4BJ5BYCE.js +61 -0
- package/dist/{chunk-474KN5II.js → chunk-4BPJXDWC.js} +111 -181
- package/dist/chunk-4VP4KH3K.js +962 -0
- package/dist/chunk-4XUGQOHA.js +797 -0
- package/dist/chunk-4ZG3XFUR.js +77 -0
- package/dist/chunk-5B2AEOW5.js +5407 -0
- package/dist/{chunk-K2ITRMHZ.js → chunk-5TSRNF4G.js} +6 -138
- package/dist/{chunk-4QKXUHSR.js → chunk-5UFT4SUX.js} +70 -20
- package/dist/{chunk-OLBBMFRD.js → chunk-5UUP6MWO.js} +24 -62
- package/dist/chunk-65DEGPJ6.js +52 -0
- package/dist/chunk-6EJMN2Y3.js +17 -0
- package/dist/chunk-6N5PTWMY.js +136 -0
- package/dist/chunk-6SGHMWE3.js +277 -0
- package/dist/chunk-6XLNIQDB.js +27 -0
- package/dist/chunk-7CR24IG7.js +242 -0
- package/dist/chunk-7MNJORFF.js +22 -0
- package/dist/{chunk-KY56HMHH.js → chunk-A3CYT5EX.js} +125 -31
- package/dist/chunk-AGYYIBLL.js +1069 -0
- package/dist/{chunk-GB6QRBXN.js → chunk-APJ265NV.js} +54 -1187
- package/dist/{chunk-673JJUWJ.js → chunk-BMEMKKIT.js} +2 -2
- package/dist/chunk-CBCAPZAA.js +229 -0
- package/dist/chunk-CMZWFGD2.js +352 -0
- package/dist/chunk-COU2UHX6.js +400 -0
- package/dist/chunk-DSELYM6W.js +1077 -0
- package/dist/chunk-DUYJ5IO6.js +644 -0
- package/dist/chunk-ECH6PKUQ.js +39 -0
- package/dist/chunk-ED4KHGC3.js +143 -0
- package/dist/chunk-EKMEHE4H.js +340 -0
- package/dist/chunk-FCSXB6T2.js +338 -0
- package/dist/chunk-FJ3H4MN5.js +48 -0
- package/dist/{chunk-RPTR2H26.js → chunk-FNTMWMX5.js} +21 -15
- package/dist/chunk-FQ4SKYE4.js +29 -0
- package/dist/chunk-G4BMMOKF.js +182 -0
- package/dist/{chunk-ZPY3JZ5E.js → chunk-GGXXDWE4.js} +183 -1233
- package/dist/chunk-HC4CLZ2Y.js +68 -0
- package/dist/{chunk-LU4TK2PR.js → chunk-HFSBBKSQ.js} +5 -56
- package/dist/{chunk-PIUMUEMV.js → chunk-HKIYEGME.js} +10 -6
- package/dist/chunk-I4HZDVNP.js +73 -0
- package/dist/chunk-IKCO5N3L.js +162 -0
- package/dist/chunk-IR4CFBFN.js +56 -0
- package/dist/{chunk-R5KLMSBV.js → chunk-J5Q24KAG.js} +2 -2
- package/dist/chunk-J7CWMCQD.js +255 -0
- package/dist/{chunk-K5XEMXTI.js → chunk-JVCV3ICN.js} +1 -1
- package/dist/chunk-KZWTDYJF.js +217 -0
- package/dist/chunk-LBMZMYH2.js +285 -0
- package/dist/{chunk-G34LV2PF.js → chunk-LM5TQCJZ.js} +84 -170
- package/dist/chunk-LW6DSM3M.js +5135 -0
- package/dist/chunk-LWLEKMDQ.js +3482 -0
- package/dist/{chunk-H6F6BYOH.js → chunk-LZSJBIVT.js} +7003 -7434
- package/dist/{chunk-HQQID6OA.js → chunk-M6SHUN7Q.js} +5 -5
- package/dist/{chunk-FST4FYJB.js → chunk-MFFY33HR.js} +99 -140
- package/dist/{chunk-BSU2YIWB.js → chunk-MVVUPGPW.js} +131 -136
- package/dist/chunk-OAO4GE4M.js +619 -0
- package/dist/{chunk-Q3RUPKEJ.js → chunk-OC7FIQPC.js} +58 -189
- package/dist/chunk-OKGUZO2U.js +34 -0
- package/dist/{chunk-GAEBEQVI.js → chunk-OOJYHWRB.js} +32 -346
- package/dist/{chunk-Q5WJOSJ7.js → chunk-OQ33BKR3.js} +2 -1
- package/dist/chunk-OQE5J4C6.js +73 -0
- package/dist/{chunk-KKNLWXI6.js → chunk-ORBHGJC5.js} +8 -8
- package/dist/{chunk-MAR7Y6HW.js → chunk-PAJK6MAQ.js} +23 -16
- package/dist/{chunk-M5T5VO65.js → chunk-PIWWS5BL.js} +837 -635
- package/dist/chunk-POHLU5DW.js +1186 -0
- package/dist/chunk-QKMUKYO7.js +4961 -0
- package/dist/chunk-SRDMMSEP.js +16405 -0
- package/dist/chunk-SST6Z5JA.js +80 -0
- package/dist/chunk-STBPMHSX.js +2456 -0
- package/dist/chunk-T6YILFSB.js +80 -0
- package/dist/chunk-TZK7PACC.js +174 -0
- package/dist/chunk-TZTZS7QK.js +227 -0
- package/dist/{chunk-ASND7OZK.js → chunk-UFIIWP2H.js} +13 -13
- package/dist/chunk-UOV2BYIW.js +107 -0
- package/dist/{chunk-PFEFKVGL.js → chunk-V6RTAOC2.js} +13 -11
- package/dist/chunk-VAKQQHWR.js +434 -0
- package/dist/chunk-VG7TBQIY.js +128 -0
- package/dist/chunk-VJWL6YS5.js +244 -0
- package/dist/{chunk-EYOKLTMF.js → chunk-VPAYEGVX.js} +17 -3
- package/dist/chunk-WEH5XRJQ.js +32 -0
- package/dist/chunk-X4RCMKVQ.js +641 -0
- package/dist/{chunk-TEKV33Q5.js → chunk-X6IAEBZR.js} +65 -33
- package/dist/chunk-XBXAASKX.js +18 -0
- package/dist/chunk-XN3L4EYL.js +46 -0
- package/dist/{chunk-RDLVBZEO.js → chunk-YCWGATWI.js} +6 -4
- package/dist/chunk-YHZX5GEU.js +193 -0
- package/dist/chunk-YXLYO42X.js +91 -0
- package/dist/{chunk-NMOX6HFD.js → chunk-ZDOOVTXZ.js} +29 -77
- package/dist/chunk-ZI647VB5.js +37 -0
- package/dist/{chunk-C4PTHK7P.js → chunk-ZWLZP4ZT.js} +5 -5
- package/dist/chunk-ZWMF7253.js +1882 -0
- package/dist/cli/index.js +62 -54
- package/dist/clio-JOU4FXVA.js +25 -0
- package/dist/code-nav-7AX6FYE6.js +600 -0
- package/dist/codewiki/build-worker.js +66 -0
- package/dist/compile-cache-CVJMMODC.js +18 -0
- package/dist/{components-DMAOEKFB.js → components-KELWS457.js} +11 -6
- package/dist/{config-IRUQ7SE4.js → config-XCDVKR23.js} +92 -55
- package/dist/configure-4GAP54ZW.js +42 -0
- package/dist/{context-5RADCKTR.js → context-4UOGGLQ5.js} +71 -35
- package/dist/context-5VKGUVJJ.js +866 -0
- package/dist/{context-3KWFLHJG.js → context-77FM5DV5.js} +15 -13
- package/dist/{context-clear-7TSNPAAI.js → context-clear-XXJRLCJJ.js} +54 -28
- package/dist/{context-index-W4RLWOQH.js → context-index-BZ4UYMTC.js} +30 -24
- package/dist/dispatch-runner-QPRDDBDX.js +1997 -0
- package/dist/{docs-5AWSPS37.js → docs-2C2LTVT2.js} +23 -10
- package/dist/{doctor-UC5NAJYQ.js → doctor-HR46URBJ.js} +27 -17
- package/dist/{eval-U6TJHRLX.js → eval-XSSNATB4.js} +29 -16
- package/dist/{evidence-YEGUW4L3.js → evidence-6HG2PY2B.js} +46 -26
- package/dist/{evolve-TXARCTPG.js → evolve-K7YU3NCY.js} +45 -25
- package/dist/{extensions-OZFJ3A3G.js → extensions-QVDOHDGJ.js} +16 -7
- package/dist/{fleet-6G3DHNYE.js → fleet-VY3HHKN6.js} +163 -54
- package/dist/{fleet-preflight-DSNT37JK.js → fleet-preflight-DDN536IT.js} +7 -4
- package/dist/{init-KZ5QTF6M.js → init-JYGXI3FK.js} +69 -32
- package/dist/{memory-73ESV5YC.js → memory-WFZMGYHX.js} +48 -27
- package/dist/{models-A4PVNWJK.js → models-I5QWSEOM.js} +39 -25
- package/dist/monitor-GE4ID3IA.js +661 -0
- package/dist/{chunk-FCIH3BIZ.js → orchestrator-EM5MC3HM.js} +15979 -12407
- package/dist/{paths-C4H6IV77.js → paths-UXLN5YYZ.js} +10 -5
- package/dist/{preload-6WVMHX3A.js → preload-P6DGH2PZ.js} +2 -2
- package/dist/{reset-BGW6OGMV.js → reset-L2FQEE3E.js} +16 -10
- package/dist/{run-YTPEYQOH.js → run-ZU3QMZPZ.js} +101 -61
- package/dist/{share-YIFFV4NQ.js → share-S5BZQC5I.js} +15 -8
- package/dist/{skills-2V6RA3OQ.js → skills-X5VXCRNQ.js} +34 -14
- package/dist/{skills-eval-S2TVJO4F.js → skills-eval-WKIHWTHR.js} +70 -34
- package/dist/steer-GGWFUJUD.js +77 -0
- package/dist/{targets-TYXLPB23.js → targets-SNCPI2NR.js} +43 -27
- package/dist/terminal-lease-BNAHVHBS.js +395 -0
- package/dist/{trace-GGOJ6Q6Z.js → trace-PNCASAXC.js} +41 -16
- package/dist/{chunk-N6F52NLF.js → tree-sitter-HGKH6LG4.js} +28 -2306
- package/dist/{uninstall-LLLT4F4W.js → uninstall-FZCQCDKC.js} +10 -5
- package/dist/{upgrade-33G2LMM5.js → upgrade-JQHHPQ4K.js} +45 -25
- package/dist/{usage-ZAFSXKKG.js → usage-OR4O5SMZ.js} +62 -31
- package/dist/verify-375KUB3Y.js +716 -0
- package/dist/web-fetch-2YHJ3KTG.js +638 -0
- package/dist/{wiki-generate-NUQCVOQ3.js → wiki-generate-UEXP2ARI.js} +74 -34
- package/dist/worker/entry.js +221 -36
- package/dist/workspace-G4ZWUIPR.js +22 -0
- package/docs/README.md +22 -17
- package/docs/acp.md +168 -16
- package/docs/alcf-provider.md +1 -1
- package/docs/architecture.md +136 -7
- package/docs/artifact-versions.md +1 -1
- package/docs/built-in-agents.md +1 -1
- package/docs/capacity-and-scheduling.md +1 -1
- package/docs/commands-and-modes.md +114 -71
- package/docs/config-knobs-audit.md +1 -3
- package/docs/configuration-and-targets.md +174 -46
- package/docs/context-engine.md +29 -6
- package/docs/development-pipeline.md +26 -1
- package/docs/dispatch-architecture-rationale.md +1 -1
- package/docs/documentation-coverage.md +2 -2
- package/docs/documentation-guide.md +1 -1
- package/docs/environment-variables.md +13 -5
- package/docs/eval-runner.md +1 -1
- package/docs/evals-internal.md +1 -1
- package/docs/evidence-and-memory.md +6 -2
- package/docs/evolution.md +2 -2
- package/docs/exit-codes-and-output.md +15 -9
- package/docs/extensions-and-sharing.md +9 -9
- package/docs/fleet-dispatch.md +7 -5
- package/docs/git-commit-provenance.md +120 -0
- package/docs/glossary.md +1 -1
- package/docs/installation-and-lifecycle.md +34 -27
- package/docs/middleware-and-components.md +1 -1
- package/docs/model-catalog.md +45 -14
- package/docs/observability.md +8 -5
- package/docs/performance-methodology.md +491 -0
- package/docs/pi-boundary.md +72 -0
- package/docs/proactive-memory.md +3 -3
- package/docs/prompt-envelope-and-tools.md +24 -3
- package/docs/provider-adapter-cookbook.md +57 -4
- package/docs/release-cut-checklist.md +129 -115
- package/docs/safety-model.md +9 -5
- package/docs/scientific-validation.md +3 -3
- package/docs/session-lifecycle.md +55 -12
- package/docs/skills-marketplace.md +12 -8
- package/docs/time-conventions.md +1 -1
- package/docs/tool-usage.md +3 -3
- package/docs/trace-store.md +1 -1
- package/docs/troubleshooting.md +10 -7
- package/docs/tui-design.md +47 -10
- package/docs/worker-dispatch-mechanics.md +1 -1
- package/package.json +19 -22
- package/skills/coding/ast-grep/SKILL.md +136 -0
- package/skills/coding/ast-grep/evals.md +56 -0
- package/skills/coding/ast-grep/references/rule_reference.md +297 -0
- package/skills/coding/coding-standards/SKILL.md +113 -0
- package/skills/coding/coding-standards/evals.md +34 -0
- package/skills/coding/prototype/SKILL.md +86 -0
- package/skills/coding/prototype/evals.md +42 -0
- package/skills/coding/prototype/references/LOGIC.md +67 -0
- package/skills/coding/prototype/references/UI.md +112 -0
- package/skills/coding/tdd/SKILL.md +101 -0
- package/skills/coding/tdd/evals.md +41 -0
- package/skills/coding/tdd/references/mocking.md +59 -0
- package/skills/coding/tdd/references/tests.md +77 -0
- package/skills/context/context-handoff/SKILL.md +126 -0
- package/skills/context/context-handoff/evals.md +57 -0
- package/skills/context/context-handoff/scripts/new-handoff.sh +26 -0
- package/skills/context/context-prime/SKILL.md +95 -0
- package/skills/context/context-prime/evals.md +54 -0
- package/skills/meta/clio-dev/SKILL.md +91 -0
- package/skills/meta/clio-dev/evals.md +45 -0
- package/skills/meta/clio-test/SKILL.md +130 -0
- package/skills/meta/clio-test/evals.md +43 -0
- package/skills/meta/clio-test/references/harness.md +97 -0
- package/skills/meta/clio-test/references/test-map.md +59 -0
- package/skills/meta/credentials/SKILL.md +125 -0
- package/skills/meta/credentials/evals.md +104 -0
- package/skills/meta/find-skills/SKILL.md +72 -0
- package/skills/meta/find-skills/evals.md +47 -0
- package/skills/meta/herdr/SKILL.md +127 -0
- package/skills/meta/herdr/evals.md +38 -0
- package/skills/meta/skill-craft/SKILL.md +102 -0
- package/skills/meta/skill-craft/evals.md +41 -0
- package/skills/planning/architecture/SKILL.md +129 -0
- package/skills/planning/architecture/evals.md +36 -0
- package/skills/planning/backlog/SKILL.md +90 -0
- package/skills/planning/backlog/evals.md +43 -0
- package/skills/planning/prd/SKILL.md +82 -0
- package/skills/planning/prd/evals.md +49 -0
- package/skills/planning/product-intent/SKILL.md +112 -0
- package/skills/planning/product-intent/evals.md +36 -0
- package/skills/planning/tech-spec/SKILL.md +115 -0
- package/skills/planning/tech-spec/evals.md +47 -0
- package/skills/registry.yaml +136 -0
- package/skills/research/arxiv-literature/SKILL.md +104 -0
- package/skills/research/arxiv-literature/evals.md +58 -0
- package/skills/research/experiment-protocol/SKILL.md +122 -0
- package/skills/research/experiment-protocol/evals.md +91 -0
- package/skills/research/scientific-debugging/SKILL.md +119 -0
- package/skills/research/scientific-debugging/evals.md +138 -0
- package/skills/research/scientific-modernization/SKILL.md +138 -0
- package/skills/research/scientific-modernization/evals.md +84 -0
- package/skills/workflow/design-council/SKILL.md +139 -0
- package/skills/workflow/design-council/evals.md +97 -0
- package/skills/workflow/grill-me/SKILL.md +186 -0
- package/skills/workflow/grill-me/evals.md +78 -0
- package/skills/workflow/workflow-distiller/SKILL.md +136 -0
- package/skills/workflow/workflow-distiller/evals.md +107 -0
- package/src/cli/acp.ts +31 -4
- package/src/cli/clio.ts +68 -6
- package/src/cli/config-inspect.ts +28 -22
- package/src/cli/configure.ts +47 -9
- package/src/cli/context-clear.ts +2 -2
- package/src/cli/context-index.ts +21 -23
- package/src/cli/context.ts +13 -8
- package/src/cli/default-target.ts +9 -17
- package/src/cli/docs.ts +11 -5
- package/src/cli/evidence.ts +4 -1
- package/src/cli/extensions.ts +10 -1
- package/src/cli/fleet.ts +47 -6
- package/src/cli/index.ts +55 -26
- package/src/cli/memory.ts +3 -1
- package/src/cli/models.ts +1 -1
- package/src/cli/modes/json-stream.ts +37 -1
- package/src/cli/modes/print.ts +24 -9
- package/src/cli/run.ts +2 -2
- package/src/cli/skills-eval.ts +23 -8
- package/src/cli/skills.ts +19 -4
- package/src/cli/targets.ts +4 -0
- package/src/cli/text-layout.ts +15 -5
- package/src/cli/trace.ts +62 -14
- package/src/cli/upgrade.ts +18 -2
- package/src/cli/usage.ts +10 -3
- package/src/cli/wiki-generate.ts +2 -1
- package/src/core/agent-environment.ts +7 -0
- package/src/core/bash-exec.ts +72 -1
- package/src/core/boot-trace.ts +9 -4
- package/src/core/bus-events.ts +20 -4
- package/src/core/commit-attribution.ts +157 -0
- package/src/core/compile-cache.ts +159 -0
- package/src/core/config.ts +131 -2
- package/src/core/defaults.ts +39 -5
- package/src/core/domain-loader.ts +12 -5
- package/src/core/git-commit-attribution.ts +387 -0
- package/src/core/incomplete-installation.ts +45 -0
- package/src/core/response-schema.ts +1 -1
- package/src/core/safe-exec.ts +13 -1
- package/src/core/settings-layers.ts +155 -21
- package/src/core/skill-activation.ts +1 -1
- package/src/core/startup-timer.ts +3 -3
- package/src/core/state-file-lock.ts +13 -1
- package/src/core/termination.ts +78 -5
- package/src/domains/config/classify.ts +15 -3
- package/src/domains/config/extension.ts +19 -13
- package/src/domains/config/index.ts +10 -0
- package/src/domains/config/keybindings.ts +45 -9
- package/src/domains/context/bootstrap-prompt.ts +1 -1
- package/src/domains/context/bootstrap.ts +111 -18
- package/src/domains/context/clear.ts +16 -11
- package/src/domains/context/clio-md.ts +111 -9
- package/src/domains/context/codewiki/artifact.ts +400 -0
- package/src/domains/context/codewiki/build-worker-protocol.ts +24 -0
- package/src/domains/context/codewiki/build-worker.ts +54 -0
- package/src/domains/context/codewiki/coordinator.ts +182 -0
- package/src/domains/context/codewiki/indexer.ts +59 -144
- package/src/domains/context/codewiki/paths.ts +67 -0
- package/src/domains/context/codewiki/schema.ts +80 -0
- package/src/domains/context/codewiki/tree-sitter.ts +1 -1
- package/src/domains/context/contract.ts +11 -5
- package/src/domains/context/extension.ts +94 -143
- package/src/domains/context/fingerprint.ts +3 -1
- package/src/domains/context/index.ts +12 -22
- package/src/domains/context/project-metadata.ts +19 -0
- package/src/domains/context/prompt-context.ts +9 -10
- package/src/domains/context/refresh.ts +29 -21
- package/src/domains/context/runtime.ts +17 -0
- package/src/domains/context/wiki/generate.ts +39 -34
- package/src/domains/context/wiki/plan.ts +1 -1
- package/src/domains/context/wiki/prompts.ts +21 -8
- package/src/domains/dispatch/code-step.ts +20 -1
- package/src/domains/dispatch/extension.ts +158 -21
- package/src/domains/dispatch/failure-classification.ts +6 -0
- package/src/domains/dispatch/fleet-commit-attribution.ts +56 -0
- package/src/domains/dispatch/orphan-recovery.ts +50 -8
- package/src/domains/dispatch/receipt-integrity.ts +5 -0
- package/src/domains/dispatch/state.ts +31 -6
- package/src/domains/dispatch/transport.ts +2 -1
- package/src/domains/dispatch/types.ts +14 -0
- package/src/domains/dispatch/worker-spawn.ts +21 -2
- package/src/domains/eval/metrics/context.ts +1 -1
- package/src/domains/eval/types.ts +0 -1
- package/src/domains/evidence/build.ts +41 -1
- package/src/domains/lifecycle/migrations/2026-08-18-lmstudio-runtime-id.ts +52 -0
- package/src/domains/lifecycle/migrations/index.ts +24 -4
- package/src/domains/middleware/hooks-io.ts +12 -0
- package/src/domains/middleware/skills-reminder.ts +30 -15
- package/src/domains/prompts/compiler.ts +142 -84
- package/src/domains/prompts/contract.ts +18 -2
- package/src/domains/prompts/extension.ts +39 -7
- package/src/domains/prompts/fragment-loader.ts +0 -1
- package/src/domains/prompts/fragments/identity/clio.md +2 -4
- package/src/domains/prompts/fragments/identity/docs-routing.md +10 -0
- package/src/domains/prompts/fragments/identity/self-awareness.md +1 -45
- package/src/domains/prompts/fragments/operating/contract.md +4 -50
- package/src/domains/prompts/fragments/operating/delegation.md +42 -0
- package/src/domains/prompts/fragments/operating/skills.md +26 -0
- package/src/domains/prompts/fragments/operating/worker.md +16 -0
- package/src/domains/prompts/fragments/safety/auto-edit.md +5 -5
- package/src/domains/prompts/fragments/safety/full-auto.md +3 -3
- package/src/domains/prompts/fragments/safety/read-only.md +4 -4
- package/src/domains/prompts/fragments/safety/suggest.md +2 -2
- package/src/domains/prompts/fragments/wiki/page.md +10 -0
- package/src/domains/prompts/fragments/wiki/plan.md +10 -0
- package/src/domains/prompts/preload.ts +3 -3
- package/src/domains/providers/auth/api-key.ts +1 -1
- package/src/domains/providers/auth/backend-file.ts +20 -10
- package/src/domains/providers/auth/backend-memory.ts +59 -4
- package/src/domains/providers/auth/boot-status.ts +65 -0
- package/src/domains/providers/auth/oauth.ts +2 -1
- package/src/domains/providers/auth/storage.ts +97 -38
- package/src/domains/providers/capabilities.ts +12 -4
- package/src/domains/providers/contract.ts +15 -4
- package/src/domains/providers/extension.ts +18 -6
- package/src/domains/providers/model-runtime-capabilities.ts +15 -4
- package/src/domains/providers/models/local-models/clio-local-coding-targets.yaml +118 -35
- package/src/domains/providers/plugins.ts +5 -3
- package/src/domains/providers/probe/fingerprint.ts +25 -5
- package/src/domains/providers/registry.ts +31 -10
- package/src/domains/providers/runtimes/boot-manifest.ts +55 -0
- package/src/domains/providers/runtimes/builtins.ts +2 -2
- package/src/domains/providers/runtimes/common/lmstudio-http.ts +423 -0
- package/src/domains/providers/runtimes/common/local-synth.ts +6 -7
- package/src/domains/providers/runtimes/local-native/lmstudio.ts +241 -0
- package/src/domains/providers/support.ts +6 -3
- package/src/domains/providers/types/local-model-quirks.ts +7 -9
- package/src/domains/providers/types/runtime-descriptor.ts +12 -1
- package/src/domains/providers/types/target-descriptor.ts +22 -0
- package/src/domains/resources/contract.ts +0 -1
- package/src/domains/resources/extension.ts +1 -3
- package/src/domains/resources/loader.ts +3 -4
- package/src/domains/resources/prompts/loader.ts +16 -2
- package/src/domains/resources/prompts/substitute.ts +1 -65
- package/src/domains/resources/skills/content-hash.ts +2 -0
- package/src/domains/resources/skills/install.ts +17 -0
- package/src/domains/resources/skills/loader.ts +17 -10
- package/src/domains/resources/skills/marketplace.ts +55 -9
- package/src/domains/safety/action-classifier.ts +4 -2
- package/src/domains/safety/audit.ts +8 -2
- package/src/domains/safety/extension.ts +1 -1
- package/src/domains/session/compaction/branch-summary.ts +3 -2
- package/src/domains/session/compaction/cut-point.ts +2 -1
- package/src/domains/session/compaction/tokens.ts +2 -1
- package/src/domains/session/context-ledger.ts +14 -0
- package/src/domains/session/contract.ts +15 -0
- package/src/domains/session/decision-board.ts +190 -0
- package/src/domains/session/entries.ts +66 -3
- package/src/domains/session/extension.ts +93 -12
- package/src/domains/session/retry.ts +10 -18
- package/src/domains/session/session-artifacts.ts +107 -0
- package/src/domains/session/task-board.ts +207 -13
- package/src/domains/session/tree/active-path.ts +44 -5
- package/src/domains/session/tree/fork.ts +26 -27
- package/src/domains/session/tree/preview.ts +2 -2
- package/src/domains/session/workspace/git-probe.ts +17 -11
- package/src/domains/user-tasks/store.ts +297 -0
- package/src/engine/acp/errors.ts +96 -0
- package/src/engine/acp/server.ts +1728 -146
- package/src/engine/acp/transport.ts +135 -14
- package/src/engine/acp/types.ts +26 -0
- package/src/engine/agent.ts +3 -3
- package/src/engine/ai.ts +32 -27
- package/src/engine/alcf-oauth.ts +26 -19
- package/src/engine/api-registry.ts +223 -0
- package/src/engine/apis/index.ts +3 -7
- package/src/engine/apis/llamacpp-residency.ts +49 -9
- package/src/engine/apis/lmstudio-residency.ts +5 -21
- package/src/engine/apis/lmstudio.ts +243 -0
- package/src/engine/apis/ollama-native.ts +24 -3
- package/src/engine/apis/openai-completions.ts +170 -91
- package/src/engine/apis/residency.ts +139 -3
- package/src/engine/apis/types.ts +16 -0
- package/src/engine/env-api-keys.ts +98 -0
- package/src/engine/gemma-channel-filter.ts +223 -0
- package/src/engine/instrumented-tui.ts +192 -0
- package/src/engine/messages.ts +14 -0
- package/src/engine/models.ts +42 -0
- package/src/engine/oauth.ts +16 -12
- package/src/engine/prompt-templates.ts +1 -0
- package/src/engine/provider-payload.ts +16 -59
- package/src/engine/strip-tokenizer-sentinels.ts +1 -1
- package/src/engine/truncate.ts +9 -0
- package/src/engine/tui.ts +17 -9
- package/src/engine/types.ts +3 -6
- package/src/engine/worker-runtime-capabilities.ts +5 -0
- package/src/engine/worker-runtime.ts +1 -1
- package/src/engine/worker-tools.ts +9 -4
- package/src/entry/boot-options.ts +50 -0
- package/src/entry/orchestrator.ts +288 -150
- package/src/interactive/application-controller.ts +89 -2
- package/src/interactive/chat-loop.ts +266 -41
- package/src/interactive/chat-panel.ts +715 -278
- package/src/interactive/chat-renderer.ts +306 -85
- package/src/interactive/clio-editor.ts +3 -8
- package/src/interactive/command-fallbacks.ts +2 -2
- package/src/interactive/context-overlay.ts +27 -1
- package/src/interactive/editor-submit.ts +253 -24
- package/src/interactive/export-html/ansi-to-html.ts +161 -0
- package/src/interactive/export-html/index.ts +51 -0
- package/src/interactive/export-html/template.ts +45 -0
- package/src/interactive/export-html/tool-renderer.ts +54 -0
- package/src/interactive/footer/dashboard.ts +4 -0
- package/src/interactive/footer/notifications.ts +1 -1
- package/src/interactive/footer/widgets.ts +42 -22
- package/src/interactive/footer-panel.ts +8 -3
- package/src/interactive/format-time.ts +14 -2
- package/src/interactive/interactive-application.ts +203 -17
- package/src/interactive/interactive-event-projection.ts +18 -1
- package/src/interactive/interactive-input-runtime.ts +50 -4
- package/src/interactive/interactive-presentation.ts +151 -19
- package/src/interactive/interactive-shell.ts +268 -14
- package/src/interactive/interactive-slash-runtime.ts +184 -117
- package/src/interactive/interactive-tickers.ts +38 -7
- package/src/interactive/keybinding-manager.ts +1 -1
- package/src/interactive/layout.ts +40 -3
- package/src/interactive/overlay-frame.ts +1 -1
- package/src/interactive/overlay-general-openers.ts +58 -1
- package/src/interactive/overlay-key-routing.ts +3 -0
- package/src/interactive/overlay-lifecycle.ts +13 -0
- package/src/interactive/overlay-permission-lifecycle.ts +2 -1
- package/src/interactive/overlay-session-lifecycle.ts +69 -12
- package/src/interactive/overlays/ask-user.ts +146 -24
- package/src/interactive/overlays/decisions.ts +300 -0
- package/src/interactive/overlays/help-reference.ts +15 -10
- package/src/interactive/overlays/model-selector.ts +34 -16
- package/src/interactive/overlays/session-selector.ts +18 -0
- package/src/interactive/overlays/settings.ts +105 -17
- package/src/interactive/overlays/skills-hub.ts +4 -4
- package/src/interactive/overlays/tree-selector.ts +41 -6
- package/src/interactive/render-trace.ts +499 -90
- package/src/interactive/renderers/compaction-summary.ts +2 -2
- package/src/interactive/renderers/diff.ts +115 -104
- package/src/interactive/renderers/mermaid.ts +53 -0
- package/src/interactive/renderers/tool-execution.ts +516 -168
- package/src/interactive/renderers/worker-entry.ts +20 -4
- package/src/interactive/session-switch-settlement.ts +10 -0
- package/src/interactive/slash-autocomplete.ts +6 -114
- package/src/interactive/slash-commands.ts +135 -47
- package/src/interactive/slash-spec.ts +9 -38
- package/src/interactive/status/controller.ts +5 -1
- package/src/interactive/status/index.ts +12 -1
- package/src/interactive/status/reasoning.ts +87 -0
- package/src/interactive/status/summary.ts +13 -2
- package/src/interactive/stdout-backpressure.ts +99 -0
- package/src/interactive/stream-pacer.ts +530 -0
- package/src/interactive/stream-pacing-policy.ts +66 -0
- package/src/interactive/tasks-overlay.ts +368 -14
- package/src/interactive/terminal-lease.ts +485 -0
- package/src/interactive/theme/tokens.ts +1 -1
- package/src/interactive/transcript-detail.ts +120 -0
- package/src/interactive/turn-context.ts +4 -3
- package/src/interactive/turn-persistence.ts +30 -13
- package/src/interactive/turn-queues.ts +12 -0
- package/src/interactive/turn-recovery.ts +25 -8
- package/src/interactive/turn-runtime.ts +79 -12
- package/src/interactive/turn-state.ts +10 -0
- package/src/interactive/view/artifacts.ts +114 -4
- package/src/interactive/view/view-overlay.ts +3 -0
- package/src/interactive/welcome-dashboard.ts +17 -16
- package/src/interactive/worker-receipts.ts +52 -3
- package/src/interactive/worker-stream.ts +5 -1
- package/src/tools/agent-tools.ts +23 -3
- package/src/tools/artifact.ts +2 -2
- package/src/tools/ask-user.ts +23 -13
- package/src/tools/bash.ts +30 -2
- package/src/tools/bootstrap.ts +34 -431
- package/src/tools/builtin-tool-catalog.ts +271 -0
- package/src/tools/codewiki/code-nav-surface.ts +29 -0
- package/src/tools/codewiki/code-nav.ts +8 -22
- package/src/tools/codewiki/shared.ts +41 -38
- package/src/tools/context/docs-engine.ts +14 -3
- package/src/tools/context/index.ts +107 -28
- package/src/tools/context/surface.ts +19 -0
- package/src/tools/core-bootstrap.ts +168 -0
- package/src/tools/credential-present.ts +5 -5
- package/src/tools/dispatch-admission.ts +533 -0
- package/src/tools/dispatch-background.ts +54 -0
- package/src/tools/dispatch-event-text.ts +6 -0
- package/src/tools/dispatch-plan.ts +9 -4
- package/src/tools/dispatch-run-events.ts +238 -0
- package/src/tools/dispatch-runner.ts +2370 -0
- package/src/tools/dispatch-scout-admission.ts +295 -0
- package/src/tools/dispatch-types.ts +77 -0
- package/src/tools/dispatch.ts +67 -3161
- package/src/tools/find.ts +4 -2
- package/src/tools/grep.ts +2 -2
- package/src/tools/lazy-tool.ts +60 -0
- package/src/tools/ledger.ts +3 -3
- package/src/tools/monitor-surface.ts +36 -0
- package/src/tools/monitor.ts +2 -32
- package/src/tools/observers.ts +2 -2
- package/src/tools/presentation.ts +107 -0
- package/src/tools/registry.ts +45 -27
- package/src/tools/safe-exec.ts +2 -2
- package/src/tools/steer-surface.ts +17 -0
- package/src/tools/steer.ts +2 -13
- package/src/tools/tasks.ts +108 -11
- package/src/tools/truncate.ts +25 -184
- package/src/tools/verify/frontend.ts +3 -1
- package/src/tools/verify/index.ts +3 -38
- package/src/tools/verify/surface.ts +46 -0
- package/src/tools/web-fetch-surface.ts +23 -0
- package/src/tools/web-fetch.ts +2 -20
- package/src/tools/write.ts +7 -2
- package/src/worker/entry.ts +39 -2
- package/src/worker/spec-contract.ts +26 -5
- package/dist/chunk-7SS2CTV2.js +0 -61361
- package/dist/chunk-DKGKUHFA.js +0 -924
- package/dist/chunk-GEP36Y4X.js +0 -12796
- package/dist/chunk-XYWBQRDM.js +0 -137
- package/dist/clio-BZVGEUFJ.js +0 -58
- package/dist/configure-S7S6F6CL.js +0 -32
- package/docs/html/agents_blueprint.html +0 -936
- package/docs/html/alcf_blueprint.html +0 -324
- package/docs/html/architecture_blueprint.html +0 -850
- package/docs/html/commands_blueprint.html +0 -939
- package/docs/html/config_knobs_audit_blueprint.html +0 -178
- package/docs/html/configuration_blueprint.html +0 -1080
- package/docs/html/context_blueprint.html +0 -603
- package/docs/html/documentation_blueprint.html +0 -832
- package/docs/html/environment_blueprint.html +0 -404
- package/docs/html/eval_blueprint.html +0 -743
- package/docs/html/evals_internal_blueprint.html +0 -190
- package/docs/html/evolution_blueprint.html +0 -674
- package/docs/html/extensions_blueprint.html +0 -2065
- package/docs/html/fleet_dispatch_blueprint.html +0 -286
- package/docs/html/index.html +0 -919
- package/docs/html/lifecycle_blueprint.html +0 -723
- package/docs/html/memory_blueprint.html +0 -699
- package/docs/html/middleware_blueprint.html +0 -664
- package/docs/html/models_blueprint.html +0 -2366
- package/docs/html/observability_blueprint.html +0 -683
- package/docs/html/provider_adapter_blueprint.html +0 -245
- package/docs/html/safety_blueprint.html +0 -1386
- package/docs/html/shared.css +0 -571
- package/docs/html/shared.js +0 -143
- package/docs/html/skills_blueprint.html +0 -671
- package/docs/html/soak_blueprint.html +0 -182
- package/docs/html/tool_usage_blueprint.html +0 -350
- package/docs/html/tools_blueprint.html +0 -2249
- package/docs/html/trace_blueprint.html +0 -235
- package/docs/html/tui_design_blueprint.html +0 -374
- package/docs/html/validation_blueprint.html +0 -961
- package/docs/html/worker_dispatch_blueprint.html +0 -231
- package/src/core/release.ts +0 -2
- package/src/domains/providers/runtimes/common/lmstudio-logger.ts +0 -32
- package/src/domains/providers/runtimes/local-native/lmstudio-native.ts +0 -491
- package/src/engine/apis/lmstudio-native.ts +0 -1438
- package/src/engine/apis/thinking-replay.ts +0 -11
- package/src/tools/string-enum.ts +0 -15
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
# Evals - scientific-modernization
|
|
2
|
+
|
|
3
|
+
Baseline scenarios (run a subagent WITHOUT the skill to capture the gap, then
|
|
4
|
+
WITH the skill to confirm it closes). Rubric is pass/fail per bullet.
|
|
5
|
+
|
|
6
|
+
## S1 - rewrite an established genomics tool in Rust
|
|
7
|
+
|
|
8
|
+
Setup: a maintained genomics parser has a large user base, a reference CLI,
|
|
9
|
+
and representative public datasets. The new repository has no validation
|
|
10
|
+
contract. Prompt: "Rewrite this parser in Rust and make it the new default."
|
|
11
|
+
|
|
12
|
+
Expected:
|
|
13
|
+
|
|
14
|
+
- Checks the upstream contribution path, maintainers, license, and downstream
|
|
15
|
+
compatibility obligations before choosing a fork or successor.
|
|
16
|
+
- Writes an acceptance contract before implementation, using the released
|
|
17
|
+
parser and fixed datasets as an independent parity oracle.
|
|
18
|
+
- Specifies file, API, CLI, error, ordering, missing-value, and numerical
|
|
19
|
+
compatibility rather than comparing only happy-path output.
|
|
20
|
+
- Splits the migration into independently validated stages with retained raw
|
|
21
|
+
comparison artifacts and rollback points.
|
|
22
|
+
- Calls the result a prototype unless ownership, release, migration, and
|
|
23
|
+
maintenance are settled.
|
|
24
|
+
|
|
25
|
+
## S2 - port a numerical solver to GPU
|
|
26
|
+
|
|
27
|
+
Setup: a CPU solver is trusted, but floating-point reduction order will change
|
|
28
|
+
on the GPU. Prompt: "Port the solver to CUDA and prove it is correct and
|
|
29
|
+
faster."
|
|
30
|
+
|
|
31
|
+
Expected:
|
|
32
|
+
|
|
33
|
+
- Uses the CPU implementation, analytical cases, conserved quantities, or
|
|
34
|
+
fixed simulated truth as an oracle independent of the CUDA code.
|
|
35
|
+
- Defines units, shapes, absolute and relative tolerances, nondeterminism, and
|
|
36
|
+
platform variation before the first GPU result.
|
|
37
|
+
- Separates correctness parity from a pre-registered performance experiment
|
|
38
|
+
and invokes experiment-protocol for the latter.
|
|
39
|
+
- Revalidates the full oracle matrix after each optimization rather than
|
|
40
|
+
inheriting correctness from an earlier version.
|
|
41
|
+
- Includes large inputs, fresh environments, failures, and numerical edge
|
|
42
|
+
cases in the last-mile plan.
|
|
43
|
+
|
|
44
|
+
## S3 - modernize packaging in place
|
|
45
|
+
|
|
46
|
+
Setup: a scientific Python package has a legacy build, multiple supported
|
|
47
|
+
platforms, and an active upstream. Prompt: "Replace the packaging system and
|
|
48
|
+
release it without changing scientific behavior."
|
|
49
|
+
|
|
50
|
+
Expected:
|
|
51
|
+
|
|
52
|
+
- Prefers an upstreamable bounded change over creating a replacement package.
|
|
53
|
+
- Captures fresh-install, upgrade, uninstall, import, CLI, and documented
|
|
54
|
+
workflow behavior from released artifacts before editing the build.
|
|
55
|
+
- Uses existing scientific outputs as a no-regression oracle even though the
|
|
56
|
+
requested change appears packaging-only.
|
|
57
|
+
- Delivers the migration in reversible stages and records platform-specific
|
|
58
|
+
evidence.
|
|
59
|
+
- Names release ownership, deprecation impact, and rollback instructions.
|
|
60
|
+
|
|
61
|
+
## S4 - anti-triggers
|
|
62
|
+
|
|
63
|
+
Setup: one request asks why a solver emits NaNs; another asks to benchmark two
|
|
64
|
+
MPI collectives without migrating software.
|
|
65
|
+
|
|
66
|
+
Expected:
|
|
67
|
+
|
|
68
|
+
- Routes the NaN diagnosis to scientific-debugging.
|
|
69
|
+
- Routes the isolated performance comparison to experiment-protocol.
|
|
70
|
+
- Does not manufacture a modernization or stewardship project for either.
|
|
71
|
+
|
|
72
|
+
## Baseline failure modes to watch for (RED)
|
|
73
|
+
|
|
74
|
+
- Starts a big-bang rewrite immediately and treats self-authored tests as proof.
|
|
75
|
+
- Declares parity from a few hand-picked examples with no tolerance semantics.
|
|
76
|
+
- Mixes correctness and performance so a speedup excuses changed answers.
|
|
77
|
+
- Discovers edge cases only in one final comparison after the rewrite.
|
|
78
|
+
- Replaces an active project without early maintainer coordination.
|
|
79
|
+
- Calls an unowned fork production-ready because installation succeeds.
|
|
80
|
+
|
|
81
|
+
## Smoke record (2026-08-13)
|
|
82
|
+
|
|
83
|
+
One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
|
|
84
|
+
(30B local, llamacpp on mini), full-auto sandbox. NOT COMPLETED: loop guard at 75 tool calls. The run also surfaced two harness containment findings (write-tool workspace escape; cross-arm workspace visibility) — harness issues, not skill issues.
|
|
@@ -0,0 +1,139 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: design-council
|
|
3
|
+
description: Use when a design decision has real tradeoffs and needs several expert perspectives that challenge each other before code is written, such as architecture choices, API shapes, storage formats, parallelization strategies, or dependency decisions. Quick mode runs a single round for a fast perspective check. Triggers on "council", "debate this", "multiple perspectives", "weigh the options", "what would experts say". Not for a one-question-at-a-time interrogation of a plan; use grill-me. Not for splitting implementation work across workers; use dispatch directly.
|
|
4
|
+
version: 0.3.0
|
|
5
|
+
license: Apache-2.0
|
|
6
|
+
allowed-tools:
|
|
7
|
+
- dispatch
|
|
8
|
+
- read
|
|
9
|
+
- grep
|
|
10
|
+
- find
|
|
11
|
+
- ls
|
|
12
|
+
- context
|
|
13
|
+
- code_nav
|
|
14
|
+
clio:
|
|
15
|
+
registry-id: iowarp/clio-coder
|
|
16
|
+
source-url: https://github.com/iowarp/clio-coder/tree/main/skills/workflow/design-council
|
|
17
|
+
audit: pass
|
|
18
|
+
provenance: designed
|
|
19
|
+
eval-status: scenarios-recorded
|
|
20
|
+
model-size: large
|
|
21
|
+
agents:
|
|
22
|
+
- scout
|
|
23
|
+
- researcher
|
|
24
|
+
- provenance
|
|
25
|
+
---
|
|
26
|
+
|
|
27
|
+
# Design Council
|
|
28
|
+
|
|
29
|
+
Run a bounded multi-perspective debate on a real design decision. The council
|
|
30
|
+
surfaces the crux of a disagreement before code commits to one side. It is not
|
|
31
|
+
a ritual: if experts would agree, do not convene it.
|
|
32
|
+
|
|
33
|
+
## Step 0 — Check the question is contested
|
|
34
|
+
|
|
35
|
+
Before composing anyone, ask: would credible experts actually disagree on the
|
|
36
|
+
answer? If every perspective you can imagine picks the same option and differs
|
|
37
|
+
only in caveats, stop here. Say the council is not needed, give the consensus
|
|
38
|
+
answer with the caveats attached, and end.
|
|
39
|
+
|
|
40
|
+
## Step 1 — Compose perspectives
|
|
41
|
+
|
|
42
|
+
Derive perspectives from the topic itself, never from a generic role menu.
|
|
43
|
+
Three is the default and the right number for almost every decision. Go to
|
|
44
|
+
four or five only when the decision genuinely has that many independent
|
|
45
|
+
stances, and never headless: each perspective is a worker run, and a model
|
|
46
|
+
that dispatches the round serially instead of in parallel turns five
|
|
47
|
+
perspectives into five sequential runs. If you are running without a user to
|
|
48
|
+
wait on you, use three perspectives and one round.
|
|
49
|
+
|
|
50
|
+
Each perspective gets:
|
|
51
|
+
|
|
52
|
+
- a name and a stance (what it argues for);
|
|
53
|
+
- the expertise it argues from;
|
|
54
|
+
- the specific thing it must attack in the other positions.
|
|
55
|
+
|
|
56
|
+
Example, "HDF5 vs Zarr for checkpoints": an HPC I/O veteran defending
|
|
57
|
+
single-file HDF5 on parallel filesystems; a cloud-native engineer arguing
|
|
58
|
+
object-store-first Zarr; an operator worried about tooling and recovery; a
|
|
59
|
+
numerics lead demanding bit-exact round-trips. Never "optimist, pessimist,
|
|
60
|
+
pragmatist".
|
|
61
|
+
|
|
62
|
+
## Step 2 — Dispatch each perspective as a read-only worker
|
|
63
|
+
|
|
64
|
+
The recipe fixes capability; your task prompt supplies the persona. Pick per
|
|
65
|
+
perspective from the read-only recipes in the live catalog:
|
|
66
|
+
|
|
67
|
+
- `scout`: stance grounded in this repository's code.
|
|
68
|
+
- `researcher`: stance leaning on external docs, standards, or papers.
|
|
69
|
+
- `provenance`: stance arguing from runtime evidence and receipts.
|
|
70
|
+
|
|
71
|
+
Run one round's perspectives in parallel: one `dispatch` call with the round's
|
|
72
|
+
task prompts in `tasks` and `mode="parallel"`. Rounds are sequential. Each
|
|
73
|
+
task prompt carries the persona block, the decision context, and the full
|
|
74
|
+
transcript so far. Workers never edit files; the debate is analysis only.
|
|
75
|
+
Dispatch receipts link every statement to a worker run.
|
|
76
|
+
|
|
77
|
+
## Step 3 — Run the rounds
|
|
78
|
+
|
|
79
|
+
1. **Positions.** Each perspective states its position, its strongest
|
|
80
|
+
argument, and what evidence would change its mind.
|
|
81
|
+
2. **Responses.** Each perspective receives the round 1 transcript and must
|
|
82
|
+
respond to named points from the others: concede, rebut, or sharpen.
|
|
83
|
+
3. **Convergence.** Each perspective states what it now agrees with, where it
|
|
84
|
+
still disagrees and why that crux is the crux, and its final
|
|
85
|
+
recommendation.
|
|
86
|
+
|
|
87
|
+
**Quick mode** (user asked for a light pass): round 1 plus synthesis. No
|
|
88
|
+
responses round.
|
|
89
|
+
|
|
90
|
+
**Early termination.** After round 1, judge disagreement on the decision
|
|
91
|
+
question itself, not on side conditions. If every position picks the same
|
|
92
|
+
option and differs only in caveats, toggles, or requests to measure later,
|
|
93
|
+
that is consensus: skip rounds 2 and 3, report that the council was not
|
|
94
|
+
needed, and return the consensus with caveats. Never manufacture friction.
|
|
95
|
+
|
|
96
|
+
## Step 4 — Synthesize
|
|
97
|
+
|
|
98
|
+
After the last round, you (the orchestrator) write:
|
|
99
|
+
|
|
100
|
+
```markdown
|
|
101
|
+
## Council Synthesis - <decision>
|
|
102
|
+
|
|
103
|
+
Agreements:
|
|
104
|
+
- <point> (all perspectives, round <n>)
|
|
105
|
+
|
|
106
|
+
Live disagreements:
|
|
107
|
+
- <point> - crux: <the fact or value judgment that would settle it>
|
|
108
|
+
|
|
109
|
+
Recommendation:
|
|
110
|
+
- <choice> because <reasoning grounded in the transcript>
|
|
111
|
+
|
|
112
|
+
Dissent preserved:
|
|
113
|
+
- <perspective>: <the objection that survives the recommendation>
|
|
114
|
+
```
|
|
115
|
+
|
|
116
|
+
Cite the transcript (perspective and round) for every claim. Done when every
|
|
117
|
+
synthesis line has a citation and the recommendation names its crux.
|
|
118
|
+
|
|
119
|
+
## Degraded mode
|
|
120
|
+
|
|
121
|
+
If dispatch is unavailable or admission-denied, run the same rounds inline:
|
|
122
|
+
write each perspective's contribution yourself, sequentially, same round
|
|
123
|
+
structure and synthesis format. Label the output as degraded (single-model
|
|
124
|
+
debate, no receipts).
|
|
125
|
+
|
|
126
|
+
## Boundaries
|
|
127
|
+
|
|
128
|
+
Stress-testing a plan by questioning its author one question at a time is
|
|
129
|
+
`grill-me`, not a council. Splitting implementation work across workers is
|
|
130
|
+
plain dispatch, not a council. Council workers analyze; they never build.
|
|
131
|
+
|
|
132
|
+
## Red Flags
|
|
133
|
+
|
|
134
|
+
- Perspectives named "optimist" and "pessimist" (role menu, not topic).
|
|
135
|
+
- More than five perspectives, or debate rounds beyond three.
|
|
136
|
+
- A synthesis that averages positions instead of naming the crux.
|
|
137
|
+
- Manufactured disagreement on a settled question.
|
|
138
|
+
- A worker asked to edit files as part of the debate.
|
|
139
|
+
- Synthesis statements that no transcript line supports.
|
|
@@ -0,0 +1,97 @@
|
|
|
1
|
+
# Evals - design-council
|
|
2
|
+
|
|
3
|
+
Baseline scenarios (run a subagent WITHOUT the skill to capture the gap, then
|
|
4
|
+
WITH the skill to confirm it closes). Rubric is pass/fail per bullet.
|
|
5
|
+
|
|
6
|
+
## S1 - storage-format decision with real tradeoffs
|
|
7
|
+
|
|
8
|
+
Setup: a project checkpointing large arrays; prompt: "council: should we use
|
|
9
|
+
HDF5 or Zarr for our checkpoint format?"
|
|
10
|
+
|
|
11
|
+
Expected:
|
|
12
|
+
|
|
13
|
+
- Composes 3 to 5 perspectives from the topic (each with a name, stance,
|
|
14
|
+
expertise, and a target to attack), not from a generic role menu.
|
|
15
|
+
- Dispatches perspectives as read-only workers (scout, researcher, or
|
|
16
|
+
provenance recipes), parallel within a round.
|
|
17
|
+
- Round 2 responses address named points from the round 1 transcript.
|
|
18
|
+
- Disagreement is surfaced with its crux stated, not averaged away.
|
|
19
|
+
- Synthesis lists agreements, live disagreements with the crux,
|
|
20
|
+
a recommendation, and preserved dissent, citing the transcript.
|
|
21
|
+
- Dispatch receipts exist linking statements to worker runs.
|
|
22
|
+
|
|
23
|
+
## S2 - consensus topic
|
|
24
|
+
|
|
25
|
+
Setup: prompt asks the council to debate a question with an obvious answer for
|
|
26
|
+
this repo (for example "should we hand-roll our own YAML parser or keep the
|
|
27
|
+
existing dependency?").
|
|
28
|
+
|
|
29
|
+
Expected:
|
|
30
|
+
|
|
31
|
+
- Round 1 comes back in agreement; the council stops there.
|
|
32
|
+
- Says explicitly that the council was not needed and returns the consensus.
|
|
33
|
+
- Does not run responses or convergence rounds, and does not manufacture
|
|
34
|
+
friction.
|
|
35
|
+
|
|
36
|
+
## S3 - dispatch unavailable
|
|
37
|
+
|
|
38
|
+
Setup: dispatch is admission-denied or absent in the environment.
|
|
39
|
+
|
|
40
|
+
Expected:
|
|
41
|
+
|
|
42
|
+
- Falls back to inline sequential perspectives with the same round structure
|
|
43
|
+
and synthesis format.
|
|
44
|
+
- Labels the result as degraded (single-model debate, no receipts).
|
|
45
|
+
|
|
46
|
+
## S4 - anti-trigger: plan stress-test
|
|
47
|
+
|
|
48
|
+
Setup: user says "poke holes in my plan" or wants their decisions interrogated
|
|
49
|
+
one at a time.
|
|
50
|
+
|
|
51
|
+
Expected:
|
|
52
|
+
|
|
53
|
+
- Refers to grill-me instead of convening a council.
|
|
54
|
+
|
|
55
|
+
## Baseline failure modes to watch for (RED)
|
|
56
|
+
|
|
57
|
+
- Generic personas (optimist, pessimist, devil's advocate) with no expertise
|
|
58
|
+
or attack target.
|
|
59
|
+
- One blended essay of pros and cons instead of independent positions that
|
|
60
|
+
respond to each other.
|
|
61
|
+
- No worker dispatch at all, or workers asked to edit files.
|
|
62
|
+
- A "balanced" synthesis that hides the crux of disagreement.
|
|
63
|
+
- Fabricated debate on a consensus topic to justify the ceremony.
|
|
64
|
+
|
|
65
|
+
## Observed live-smoke results
|
|
66
|
+
|
|
67
|
+
Run 2026-07-01/02, headless `clio-coder run --skill` against a scratch fixture (an
|
|
68
|
+
MPI checkpointing module writing one raw npy per rank per step).
|
|
69
|
+
|
|
70
|
+
- Full council (atomic rename vs direct write): three bounded rounds, four
|
|
71
|
+
topic-composed perspectives each dispatched as read-only `scout` workers
|
|
72
|
+
(parallel within rounds, sequential between), receipts on every statement
|
|
73
|
+
(exit 0, agentId scout), synthesis with agreements, a crux, and the
|
|
74
|
+
Performance Engineer's dissent preserved and answered. Note the topic
|
|
75
|
+
proved genuinely debatable, so full rounds were correct.
|
|
76
|
+
- Degraded fallback (HDF5 vs Zarr, run while the fleet target had model
|
|
77
|
+
residency failures): the skill attempted a parallel `dispatch` first,
|
|
78
|
+
observed worker failures in receipts, then ran the same rounds inline and
|
|
79
|
+
labeled the result degraded exactly as the clause requires.
|
|
80
|
+
- Early termination (S2, zero-padded vs unpadded checkpoint filenames): the
|
|
81
|
+
first draft of the skill ran full rounds on side-quibbles for two
|
|
82
|
+
debatable-in-hindsight "consensus" topics; the composition and termination
|
|
83
|
+
clauses were tightened to judge consensus on the decision question itself.
|
|
84
|
+
With the shipped wording the council declared itself not needed before any
|
|
85
|
+
dispatch, cited the rule, and returned the consensus with caveats. No
|
|
86
|
+
manufactured friction.
|
|
87
|
+
|
|
88
|
+
## Smoke record (2026-08-13)
|
|
89
|
+
|
|
90
|
+
One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
|
|
91
|
+
(30B local, llamacpp on mini), full-auto sandbox. NOT COMPLETED: treatment timed out at the 900s ceiling mid-council (serial perspectives). Needs a longer timeout or fewer personas for headless eval; skill was visibly working when cut.
|
|
92
|
+
|
|
93
|
+
Follow-up (2026-08-13, v0.3.0): Step 1 now makes three perspectives the
|
|
94
|
+
default and tells a headless run to use three and one round, which is the
|
|
95
|
+
blocker the timeout exposed. Not re-run, so `eval-status` stays
|
|
96
|
+
`scenarios-recorded`; the next campaign has to confirm the shortened council
|
|
97
|
+
fits the ceiling.
|
|
@@ -0,0 +1,186 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: grill-me
|
|
3
|
+
description: Use when the user wants a plan, design, or idea stress-tested through a phased one-question-at-a-time interview before any code is written, or when intent is too ambiguous to plan from. Scans available context first, reviews known facts, fills missing decisions, respects stop signals, and ends with a compact decision log. Triggers on "grill me", "interview me", "stress-test this plan", "poke holes in this".
|
|
4
|
+
version: 0.3.2
|
|
5
|
+
license: Apache-2.0
|
|
6
|
+
allowed-tools:
|
|
7
|
+
- read
|
|
8
|
+
- grep
|
|
9
|
+
- ls
|
|
10
|
+
- find
|
|
11
|
+
- git
|
|
12
|
+
- context
|
|
13
|
+
- code_nav
|
|
14
|
+
- ask_user
|
|
15
|
+
clio:
|
|
16
|
+
registry-id: iowarp/clio-coder
|
|
17
|
+
source-url: https://github.com/iowarp/clio-coder/tree/main/skills/workflow/grill-me
|
|
18
|
+
audit: pass
|
|
19
|
+
provenance: designed
|
|
20
|
+
eval-status: smoke-checked
|
|
21
|
+
model-size: large
|
|
22
|
+
---
|
|
23
|
+
|
|
24
|
+
# Grill Me
|
|
25
|
+
|
|
26
|
+
Run a rigorous, repo-aware interview that turns a vague plan into explicit
|
|
27
|
+
decisions. The point is not to interrogate for sport; it is to surface hidden
|
|
28
|
+
branches before anyone writes code.
|
|
29
|
+
|
|
30
|
+
## Operating Contract
|
|
31
|
+
|
|
32
|
+
- Use `ask_user` for the interview whenever it is active.
|
|
33
|
+
- For every interview round, call `ask_user` with `mode: "single_question"` and
|
|
34
|
+
exactly one question.
|
|
35
|
+
- On the first ask for a normal grill-me run, set `max_rounds` to a bounded
|
|
36
|
+
value, usually `12` and at most `16` unless the user explicitly asked for a
|
|
37
|
+
very deep interview.
|
|
38
|
+
- Put your recommended answer first when options are natural. Include 2-3 real
|
|
39
|
+
alternatives with short tradeoff descriptions.
|
|
40
|
+
- The user answers in natural language. You translate answers into compact
|
|
41
|
+
decision keys and rationale when you call `ask_user` with `action: "complete"`.
|
|
42
|
+
- If `ask_user` is unavailable, ask in plain text, still one question at a
|
|
43
|
+
time, and keep an internal decision log.
|
|
44
|
+
|
|
45
|
+
## Phase Map
|
|
46
|
+
|
|
47
|
+
Walk phases in order unless the user names a narrower target. Do not skip the
|
|
48
|
+
scan. A phase can be "review" when context already contains a plausible answer
|
|
49
|
+
or "fill" when the decision is genuinely missing.
|
|
50
|
+
|
|
51
|
+
| Phase | Scope | Default mode |
|
|
52
|
+
|---|---|---|
|
|
53
|
+
| 0 | Scan available context: user prompt, named files, repo structure, git state, existing plans/specs | No questions unless the subject is missing |
|
|
54
|
+
| 1 | Frame: user, problem, outcome, non-goals, success criteria | Fill or review |
|
|
55
|
+
| 2 | Current state: existing code, constraints, conventions, integration points, prior attempts | Review |
|
|
56
|
+
| 3 | Shape: data model, API/UX surface, ownership boundaries, naming, compatibility | Fill |
|
|
57
|
+
| 4 | Risk: failure modes, migrations, rollout, test strategy, observability, reversibility | Fill |
|
|
58
|
+
| 5 | Delivery: first slice, done-when checks, deferrals, handoff target (`prd`, `cut-it`, or direct implementation) | Review then complete |
|
|
59
|
+
|
|
60
|
+
## Workflow
|
|
61
|
+
|
|
62
|
+
### Step 1 - Scan
|
|
63
|
+
|
|
64
|
+
Read what the user already gave you. If the task references files, plans, code,
|
|
65
|
+
tests, or project conventions, inspect them before asking. Prefer
|
|
66
|
+
`context(scope="workspace")`, `grep`, `read`, and codewiki tools over guessing.
|
|
67
|
+
|
|
68
|
+
Privately build a phase map:
|
|
69
|
+
|
|
70
|
+
- known facts
|
|
71
|
+
- assumptions worth challenging
|
|
72
|
+
- missing decisions
|
|
73
|
+
- dependencies between decisions
|
|
74
|
+
- likely deferrals and why they might be safe
|
|
75
|
+
|
|
76
|
+
Never spend the user's attention on facts that the repo answers. If you found
|
|
77
|
+
the answer in code, summarize it briefly and ask only whether it should remain
|
|
78
|
+
true.
|
|
79
|
+
|
|
80
|
+
### Step 2 - Choose Review Or Fill
|
|
81
|
+
|
|
82
|
+
For each phase:
|
|
83
|
+
|
|
84
|
+
- **Review mode**: context already has an answer. Present the finding in one or
|
|
85
|
+
two sentences, then ask a targeted question such as "Is this still accurate?"
|
|
86
|
+
or "Should we keep this constraint?"
|
|
87
|
+
- **Fill mode**: context is sparse or ambiguous. Ask the highest-leverage
|
|
88
|
+
missing decision first.
|
|
89
|
+
|
|
90
|
+
Always resolve root decisions before leaves, in the Question Priority order
|
|
91
|
+
below.
|
|
92
|
+
|
|
93
|
+
### Step 3 - Ask One Question
|
|
94
|
+
|
|
95
|
+
Each `ask_user` round contains one question only:
|
|
96
|
+
|
|
97
|
+
- one stable `header`
|
|
98
|
+
- one concrete question
|
|
99
|
+
- recommended option first when options fit
|
|
100
|
+
- no multi-part wording hidden inside the question
|
|
101
|
+
|
|
102
|
+
Bad: "Who is this for, what should v1 include, and how should we test it?"
|
|
103
|
+
|
|
104
|
+
Good: "Which user should v1 optimize for first?"
|
|
105
|
+
|
|
106
|
+
If an answer is vague, ask a follow-up on the same branch. Do not jump to a new
|
|
107
|
+
branch while the current one is still unresolved.
|
|
108
|
+
|
|
109
|
+
### Step 4 - Respect Stop Signals
|
|
110
|
+
|
|
111
|
+
Stop immediately when the user says "stop", "enough", "later", "done", "next
|
|
112
|
+
time", or cancels the modal. Do not ask another question to confirm stopping.
|
|
113
|
+
|
|
114
|
+
If you have enough decisions to be useful, call:
|
|
115
|
+
|
|
116
|
+
```json
|
|
117
|
+
{
|
|
118
|
+
"action": "complete",
|
|
119
|
+
"summary": "Short interview closeout.",
|
|
120
|
+
"decisions": [
|
|
121
|
+
{
|
|
122
|
+
"key": "primary_outcome",
|
|
123
|
+
"value": "Smallest useful slice",
|
|
124
|
+
"rationale": "The user prioritized fast validation over broad architecture.",
|
|
125
|
+
"confidence": "high",
|
|
126
|
+
"source_question": "What should this plan optimize for first?"
|
|
127
|
+
}
|
|
128
|
+
]
|
|
129
|
+
}
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
Then provide the decision log. If the stop happened before enough context,
|
|
133
|
+
state the partial decisions and the next unresolved root question.
|
|
134
|
+
|
|
135
|
+
### Step 5 - Complete
|
|
136
|
+
|
|
137
|
+
Before final prose, call `ask_user` with `action: "complete"` and a compact
|
|
138
|
+
`decisions` array. Then write the final decision log:
|
|
139
|
+
|
|
140
|
+
```markdown
|
|
141
|
+
## Decision Log - <topic>
|
|
142
|
+
1. <decision> - chosen over <alternative> because <reason>
|
|
143
|
+
2. ...
|
|
144
|
+
|
|
145
|
+
Deferred:
|
|
146
|
+
- <item> - safe because <reason>
|
|
147
|
+
|
|
148
|
+
Open risks:
|
|
149
|
+
- <risk or unresolved branch>
|
|
150
|
+
|
|
151
|
+
Recommended next step:
|
|
152
|
+
- <prd | cut-it | direct implementation> - <why>
|
|
153
|
+
```
|
|
154
|
+
|
|
155
|
+
## Question Priority
|
|
156
|
+
|
|
157
|
+
Use this ordering when several questions are possible:
|
|
158
|
+
|
|
159
|
+
1. User and problem being solved.
|
|
160
|
+
2. Primary success measure.
|
|
161
|
+
3. Explicit non-goals.
|
|
162
|
+
4. Existing constraints from repo or environment.
|
|
163
|
+
5. Data/API/UX boundary.
|
|
164
|
+
6. Failure modes and recovery.
|
|
165
|
+
7. Tests and done-when checks.
|
|
166
|
+
8. First implementation slice.
|
|
167
|
+
9. Naming and polish.
|
|
168
|
+
|
|
169
|
+
## Decision Rules
|
|
170
|
+
|
|
171
|
+
- If the user says "whatever you think", record your recommendation as the
|
|
172
|
+
decision and say so.
|
|
173
|
+
- If two choices are both viable, choose the one that reduces irreversible
|
|
174
|
+
work unless the user explicitly values speed or breadth more.
|
|
175
|
+
- If a decision is safe to defer, record why and what later signal will force
|
|
176
|
+
it.
|
|
177
|
+
- If the plan is too vague to slice or implement, say that clearly and continue
|
|
178
|
+
interviewing instead of fabricating certainty.
|
|
179
|
+
|
|
180
|
+
## Red Flags
|
|
181
|
+
|
|
182
|
+
- Asking multiple questions in one `ask_user` round.
|
|
183
|
+
- Asking about facts discoverable from the repo.
|
|
184
|
+
- Letting `ask_user` hit the round limit without completing the interview.
|
|
185
|
+
- Ending with a summary paragraph instead of the decision log.
|
|
186
|
+
- Treating cancellation as permission to keep asking.
|
|
@@ -0,0 +1,78 @@
|
|
|
1
|
+
# Evals - grill-me
|
|
2
|
+
|
|
3
|
+
Baseline scenarios (run a subagent WITHOUT the skill to capture the gap, then
|
|
4
|
+
WITH the skill to confirm it closes). Rubric is pass/fail per bullet.
|
|
5
|
+
|
|
6
|
+
## S1 - vague feature request
|
|
7
|
+
|
|
8
|
+
Setup: repo with an existing config domain. Prompt: "grill me on adding plugin
|
|
9
|
+
support."
|
|
10
|
+
|
|
11
|
+
Expected:
|
|
12
|
+
|
|
13
|
+
- Scans the existing config/extension code before asking anything the repo can
|
|
14
|
+
answer.
|
|
15
|
+
- Starts with a root decision such as outcome, target user, or scope boundary.
|
|
16
|
+
- Uses `ask_user` with `mode: "single_question"` and exactly one question in
|
|
17
|
+
the round.
|
|
18
|
+
- Puts its recommended answer first when options are supplied.
|
|
19
|
+
- Ends by calling `ask_user` with `action: "complete"` and then writes a
|
|
20
|
+
decision log, not a summary paragraph.
|
|
21
|
+
|
|
22
|
+
## S2 - answerable-from-repo question
|
|
23
|
+
|
|
24
|
+
Setup: plan mentions "the test runner". The repo's package.json defines it.
|
|
25
|
+
|
|
26
|
+
Expected:
|
|
27
|
+
|
|
28
|
+
- Does NOT ask the user which test runner is used; reads package.json instead.
|
|
29
|
+
- States what it found and asks only whether that constraint should remain
|
|
30
|
+
true.
|
|
31
|
+
- Treats the area as review mode, not fill mode.
|
|
32
|
+
|
|
33
|
+
## S3 - user defers
|
|
34
|
+
|
|
35
|
+
Setup: mid-interview, user answers "whatever you think is best."
|
|
36
|
+
|
|
37
|
+
Expected:
|
|
38
|
+
|
|
39
|
+
- Records its own recommendation as the decision and says so explicitly.
|
|
40
|
+
- Does not silently skip the branch.
|
|
41
|
+
- Continues only if another root decision remains.
|
|
42
|
+
|
|
43
|
+
## S4 - long phased interview
|
|
44
|
+
|
|
45
|
+
Setup: prompt asks for a deep stress test of a large design.
|
|
46
|
+
|
|
47
|
+
Expected:
|
|
48
|
+
|
|
49
|
+
- First `ask_user` call sets a bounded `max_rounds` value such as 12 or 16.
|
|
50
|
+
- Still asks one question per round.
|
|
51
|
+
- Completes before the limit when decisions are sufficient.
|
|
52
|
+
- If the limit is near, closes with current decisions and open risks instead
|
|
53
|
+
of continuing to ask.
|
|
54
|
+
|
|
55
|
+
## S5 - stop signal
|
|
56
|
+
|
|
57
|
+
Setup: mid-interview, user says "stop", "enough", "later", or cancels the
|
|
58
|
+
modal.
|
|
59
|
+
|
|
60
|
+
Expected:
|
|
61
|
+
|
|
62
|
+
- Stops immediately and does not ask a confirmation question.
|
|
63
|
+
- Calls `ask_user` complete when possible with partial decisions.
|
|
64
|
+
- Final response includes partial decisions and the next unresolved root
|
|
65
|
+
question.
|
|
66
|
+
|
|
67
|
+
## Baseline failure modes to watch for (RED)
|
|
68
|
+
|
|
69
|
+
- Question batching: more than one question in an `ask_user` round.
|
|
70
|
+
- Interview starts cold without reading repo facts that are clearly relevant.
|
|
71
|
+
- Interview ends when the user gets tired, with no decision log.
|
|
72
|
+
- Asks about facts discoverable via grep/read.
|
|
73
|
+
- Hits the round limit without a useful closeout.
|
|
74
|
+
|
|
75
|
+
## Smoke record (2026-08-13)
|
|
76
|
+
|
|
77
|
+
One representative scenario via `clio-coder skills eval` against Nemo-3.5-Lightning
|
|
78
|
+
(30B local, llamacpp on mini), full-auto sandbox. PASS (smoke). Skill loaded and interviewed; judge emitted nothing (truncation).
|