@kodax-ai/kodax 0.7.95 → 0.7.96-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +178 -1
- package/LICENSE +158 -158
- package/README.md +149 -24
- package/README_CN.md +117 -24
- package/config-templates/config.example.jsonc +2 -1
- package/config-templates/integrations/a2a.example.jsonc +91 -91
- package/config-templates/integrations/extensions.example.jsonc +7 -7
- package/config-templates/integrations/mcp.example.jsonc +16 -16
- package/dist/builtin/code-review/SKILL.md +82 -82
- package/dist/builtin/git-workflow/SKILL.md +84 -84
- package/dist/builtin/skill-creator/SKILL.md +127 -127
- package/dist/builtin/skill-creator/agents/analyzer.md +12 -12
- package/dist/builtin/skill-creator/agents/comparator.md +13 -13
- package/dist/builtin/skill-creator/agents/grader.md +13 -13
- package/dist/builtin/skill-creator/references/schemas.md +227 -227
- package/dist/builtin/skill-creator/scripts/aggregate-benchmark.d.ts +46 -46
- package/dist/builtin/skill-creator/scripts/aggregate-benchmark.js +208 -208
- package/dist/builtin/skill-creator/scripts/analyze-benchmark.d.ts +46 -46
- package/dist/builtin/skill-creator/scripts/analyze-benchmark.js +286 -286
- package/dist/builtin/skill-creator/scripts/compare-runs.d.ts +62 -62
- package/dist/builtin/skill-creator/scripts/compare-runs.js +330 -330
- package/dist/builtin/skill-creator/scripts/generate-review.d.ts +33 -33
- package/dist/builtin/skill-creator/scripts/generate-review.js +414 -414
- package/dist/builtin/skill-creator/scripts/grade-evals.d.ts +73 -73
- package/dist/builtin/skill-creator/scripts/grade-evals.js +402 -402
- package/dist/builtin/skill-creator/scripts/improve-description.d.ts +23 -23
- package/dist/builtin/skill-creator/scripts/improve-description.js +160 -160
- package/dist/builtin/skill-creator/scripts/init-skill.d.ts +14 -14
- package/dist/builtin/skill-creator/scripts/init-skill.js +155 -155
- package/dist/builtin/skill-creator/scripts/install-skill.d.ts +29 -29
- package/dist/builtin/skill-creator/scripts/install-skill.js +173 -173
- package/dist/builtin/skill-creator/scripts/package-skill.d.ts +38 -38
- package/dist/builtin/skill-creator/scripts/package-skill.js +121 -121
- package/dist/builtin/skill-creator/scripts/quick-validate.d.ts +8 -8
- package/dist/builtin/skill-creator/scripts/quick-validate.js +163 -163
- package/dist/builtin/skill-creator/scripts/run-eval.d.ts +66 -66
- package/dist/builtin/skill-creator/scripts/run-eval.js +353 -353
- package/dist/builtin/skill-creator/scripts/run-loop.d.ts +49 -49
- package/dist/builtin/skill-creator/scripts/run-loop.js +242 -242
- package/dist/builtin/skill-creator/scripts/run-trigger-eval.d.ts +58 -58
- package/dist/builtin/skill-creator/scripts/run-trigger-eval.js +224 -224
- package/dist/builtin/tdd/SKILL.md +56 -56
- package/dist/chunks/agent-4FABF6WK.js +2 -0
- package/dist/chunks/argument-completer-72VHB2VL.js +2 -0
- package/dist/chunks/{chunk-7JXA3533.js → chunk-4OIQZWA5.js} +223 -226
- package/dist/chunks/{chunk-BNVSKIAB.js → chunk-6BVF3HZ7.js} +136 -136
- package/dist/chunks/chunk-7JZZIFZH.js +501 -0
- package/dist/chunks/chunk-AC7Z2OIY.js +2 -0
- package/dist/chunks/{chunk-JS452J2F.js → chunk-C4KYZAFU.js} +2 -2
- package/dist/chunks/chunk-CHUPIWIF.js +251 -0
- package/dist/chunks/{chunk-C4DTUKTR.js → chunk-GYUU7XLH.js} +5 -4
- package/dist/chunks/{chunk-OK4AXDC5.js → chunk-HAX55GOQ.js} +1 -1
- package/dist/chunks/chunk-JKTJCGO2.js +418 -0
- package/dist/chunks/{chunk-O743AJ5V.js → chunk-LQYETRFK.js} +171 -175
- package/dist/chunks/{chunk-2PEDBGKN.js → chunk-LTUZTQ6W.js} +1 -1
- package/dist/chunks/chunk-LYWDUSKO.js +2 -0
- package/dist/chunks/chunk-NGH6T4FD.js +92 -0
- package/dist/chunks/{chunk-SRIJX5ZP.js → chunk-NVRWCSA5.js} +1 -1
- package/dist/chunks/chunk-O4DS7RIQ.js +384 -0
- package/dist/chunks/chunk-ORXU6OWZ.js +2 -0
- package/dist/chunks/chunk-POTW3O65.js +124 -0
- package/dist/chunks/{chunk-2TM3W3GK.js → chunk-RIQSS56Y.js} +1 -1
- package/dist/chunks/chunk-U27XKEN5.js +407 -0
- package/dist/chunks/{chunk-6ABQJUK6.js → chunk-XUX6OUCA.js} +2 -2
- package/dist/chunks/{client-JDOS3YJR.js → client-4C456K4O.js} +1 -1
- package/dist/chunks/compaction-config-3VQRDFF4.js +2 -0
- package/dist/chunks/{construction-bootstrap-YCWHHUYO.js → construction-bootstrap-AHN7EDGG.js} +1 -1
- package/dist/chunks/dist-BCBQXAGI.js +2 -0
- package/dist/chunks/dist-QK2YDWU6.js +2 -0
- package/dist/chunks/host-VIHV7ADG.js +2 -0
- package/dist/chunks/run-manager-NBVMU3KV.js +2 -0
- package/dist/chunks/utils-MEJZ2IGI.js +2 -0
- package/dist/index.d.ts +22 -21
- package/dist/index.js +7 -7
- package/dist/kodax_bootstrap.js +26 -26
- package/dist/kodax_cli.js +1905 -1882
- package/dist/native/darwin-arm64/LICENSE-APACHE.txt +13 -0
- package/dist/native/darwin-arm64/kodax-text-transaction.node +0 -0
- package/dist/native/darwin-arm64/manifest.json +16 -0
- package/dist/native/darwin-x64/LICENSE-APACHE.txt +13 -0
- package/dist/native/darwin-x64/kodax-text-transaction.node +0 -0
- package/dist/native/darwin-x64/manifest.json +16 -0
- package/dist/native/linux-arm64/LICENSE-APACHE.txt +13 -0
- package/dist/native/linux-arm64/kodax-text-transaction.node +0 -0
- package/dist/native/linux-arm64/manifest.json +16 -0
- package/dist/native/linux-x64/LICENSE-APACHE.txt +13 -0
- package/dist/native/linux-x64/kodax-text-transaction.node +0 -0
- package/dist/native/linux-x64/manifest.json +16 -0
- package/dist/native/win32-x64/LICENSE-APACHE.txt +13 -0
- package/dist/native/win32-x64/NOTICE-windows-sandbox.txt +8 -0
- package/dist/native/win32-x64/kodax-windows-sandbox.exe +0 -0
- package/dist/native/win32-x64/kodax-windows-text-transaction.node +0 -0
- package/dist/native/win32-x64/manifest.json +30 -0
- package/dist/provider-capabilities.json +108 -2
- package/dist/runtime-worker.js +1745 -1727
- package/dist/sandbox-network-broker.js +1715 -0
- package/dist/sdk-a2a.d.ts +13 -12
- package/dist/sdk-a2a.js +1 -1
- package/dist/sdk-agent.d.ts +27 -18
- package/dist/sdk-agent.js +1 -1
- package/dist/sdk-coding.d.ts +107 -58
- package/dist/sdk-coding.js +1 -1
- package/dist/sdk-experimental-memory.d.ts +4 -4
- package/dist/sdk-experimental-memory.js +1 -1
- package/dist/sdk-llm.d.ts +9 -10
- package/dist/sdk-llm.js +1 -1
- package/dist/sdk-mcp.js +1 -1
- package/dist/sdk-media.d.ts +1 -1
- package/dist/sdk-media.js +1 -1
- package/dist/sdk-repl.d.ts +25 -17
- package/dist/sdk-repl.js +1 -1
- package/dist/sdk-runtime.d.ts +41 -17
- package/dist/sdk-runtime.js +1 -1
- package/dist/sdk-sandbox.d.ts +17 -3
- package/dist/sdk-sandbox.js +1 -1
- package/dist/sdk-session.d.ts +8 -7
- package/dist/sdk-session.js +1 -1
- package/dist/sdk-skills.js +1 -1
- package/dist/semantic-worker.js +61 -61
- package/dist/types-chunks/{base.d-vopO1rml.d.ts → base.d-COo8E2dk.d.ts} +18 -27
- package/dist/types-chunks/{bash-prefix-extractor.d-B2rxNI4L.d.ts → bash-prefix-extractor.d-DQOlA8cy.d.ts} +72 -52
- package/dist/types-chunks/{capability-learning.d-CE1tm29C.d.ts → capability-learning.d-lRPgA3Or.d.ts} +1 -1
- package/dist/types-chunks/{capsule.d-Y7QRS4J2.d.ts → capsule.d-0zxsgXg_.d.ts} +3 -3
- package/dist/types-chunks/{controller.d-BR-UVeUR.d.ts → controller.d-BTtyIIC1.d.ts} +1 -1
- package/dist/types-chunks/{controller.d-C5noq5-w.d.ts → controller.d-C7tXUcXo.d.ts} +2 -2
- package/dist/types-chunks/{guardrail.d-DqHD-61O.d.ts → guardrail.d-DKTyEHJa.d.ts} +5 -5
- package/dist/types-chunks/{history-retrieval.d-CKU8zmt_.d.ts → history-retrieval.d-CHnL99S7.d.ts} +2 -2
- package/dist/types-chunks/{public-api.d-D7fXaWW8.d.ts → public-api.d-Dx3L3W4Z.d.ts} +7 -4
- package/dist/types-chunks/{repl.d-BYDb14s6.d.ts → repl.d-mO01olIa.d.ts} +5 -5
- package/dist/types-chunks/{resolver.d-CwHUwbdE.d.ts → resolver.d-DlXPjeyx.d.ts} +38 -6
- package/dist/types-chunks/{review-inbox.d-BSgFewyj.d.ts → review-inbox.d-WDSe1nze.d.ts} +1 -1
- package/dist/types-chunks/{run-manager.d-BVJSntpS.d.ts → run-manager.d-DENbbzgR.d.ts} +1 -1
- package/dist/types-chunks/{sdk-session-CqEG84oP.d.ts → sdk-session-DYHGXXH4.d.ts} +3 -3
- package/dist/types-chunks/{shell-command-sets.d-DqaaBwhU.d.ts → shell-command-sets.d-CjFqS4dp.d.ts} +1 -1
- package/dist/types-chunks/{side-query.d-iH0P9ASy.d.ts → side-query.d-j8A3sFv_.d.ts} +2 -2
- package/dist/types-chunks/{types.d-rUOjXych.d.ts → types.d-BbZjR0Wr.d.ts} +1 -1
- package/dist/types-chunks/{types.d-CaTXAW_Q.d.ts → types.d-BkTIKdH6.d.ts} +2 -2
- package/dist/types-chunks/{types.d-C5xQEFcj.d.ts → types.d-DncLrpu_.d.ts} +4 -4
- package/dist/types-chunks/{types.d-CK9A9rpP.d.ts → types.d-Xy0f3x4i.d.ts} +9 -0
- package/dist/types-chunks/{utils.d-CWUyLiXo.d.ts → utils.d-CnxefMos.d.ts} +6 -6
- package/package.json +10 -3
- package/public_docs/README.md +130 -0
- package/public_docs/configuration/sandbox.md +289 -0
- package/public_docs/sdk/embedder-guide.md +6674 -0
- package/scripts/kodax-bin.cjs +28 -28
- package/scripts/production-env.cjs +25 -25
- package/dist/chunks/agent-XXZMXKKA.js +0 -2
- package/dist/chunks/argument-completer-KCYDGTQW.js +0 -2
- package/dist/chunks/chunk-77PNP27P.js +0 -92
- package/dist/chunks/chunk-7QBWHCIZ.js +0 -2
- package/dist/chunks/chunk-BHE66SP3.js +0 -123
- package/dist/chunks/chunk-BR2OSB2I.js +0 -722
- package/dist/chunks/chunk-DSMONSVB.js +0 -407
- package/dist/chunks/chunk-Q5VRWSKG.js +0 -418
- package/dist/chunks/chunk-XWRPOGJN.js +0 -4
- package/dist/chunks/chunk-Y5VIQBT6.js +0 -383
- package/dist/chunks/compaction-config-EGS4577X.js +0 -2
- package/dist/chunks/dist-ITTXQHRJ.js +0 -2
- package/dist/chunks/dist-KFYN3OKD.js +0 -2
- package/dist/chunks/host-VAYADIPW.js +0 -2
- package/dist/chunks/run-manager-52IJZMZJ.js +0 -2
- package/dist/chunks/utils-P7PBSL3R.js +0 -2
- package/dist/sandbox-workspace-session.js +0 -1705
|
@@ -1,127 +1,127 @@
|
|
|
1
|
-
---
|
|
2
|
-
name: skill-creator
|
|
3
|
-
description: 创建、重写、迁移和优化 KodaX/Agent Skills。当用户想新建 skill、把外部 skill 迁移到 KodaX、改进触发描述、整理 supporting files、设计评测用例、补齐 grading/benchmark/review/comparison 流程或验证 skill 结构时使用。即使用户没有明确说“skill”,只要目标是在沉淀可复用的代理工作流、提示词或脚本能力,也应该使用这个 skill。
|
|
4
|
-
user-invocable: true
|
|
5
|
-
allowed-tools: "Read, Grep, Glob, Write, Edit, Bash(node:*, npm:*, npx:*)"
|
|
6
|
-
argument-hint: "[skill-name-or-task]"
|
|
7
|
-
compatibility: "Optimized for KodaX and Agent Skills style directories. Bundled helper scripts use Node.js instead of Python."
|
|
8
|
-
---
|
|
9
|
-
|
|
10
|
-
# Skill Creator
|
|
11
|
-
|
|
12
|
-
把用户的工作流整理成一个可维护、可触发、可评估的 skill。优先适配 KodaX 的 skill 运行方式,而不是逐字复制外部平台的实现。
|
|
13
|
-
|
|
14
|
-
## 何时使用
|
|
15
|
-
|
|
16
|
-
- 用户要新建 skill,或把一次对话里的工作流沉淀成 skill。
|
|
17
|
-
- 用户要改已有 skill 的触发描述、结构、supporting files 或提示词。
|
|
18
|
-
- 用户要把 Claude / Anthropic / 其他平台的 skill 移植到 KodaX。
|
|
19
|
-
- 用户要为 skill 设计测试提示、评估结构、人工 review 流程或 benchmark 汇总。
|
|
20
|
-
|
|
21
|
-
## 工作方式
|
|
22
|
-
|
|
23
|
-
### 1. 先收敛目标
|
|
24
|
-
|
|
25
|
-
先明确四件事:
|
|
26
|
-
1. 这个 skill 要解决什么任务。
|
|
27
|
-
2. 什么时候应该触发,什么时候不应该触发。
|
|
28
|
-
3. 输出是什么形态。
|
|
29
|
-
4. 是否需要评测和人工 review。
|
|
30
|
-
|
|
31
|
-
如果用户已经给了样例对话、提示词或外部 skill 仓库,先从已有材料里提炼,不要重复让用户描述。
|
|
32
|
-
|
|
33
|
-
### 2. 适配 KodaX,而不是照抄外部 skill
|
|
34
|
-
|
|
35
|
-
迁移外部 skill 时,拆成三类:
|
|
36
|
-
- 可直接复用:SKILL.md 的思路、评测结构、参考文档组织方式。
|
|
37
|
-
- 需要改写:路径约定、触发描述、支持的 frontmatter 字段、命令示例。
|
|
38
|
-
- 不要硬搬:强依赖 Claude Code、`claude -p`、Cowork、Python stdlib 或专有事件流的部分。
|
|
39
|
-
|
|
40
|
-
如果外部 skill 依赖特定宿主能力,优先改成 KodaX 当前能承接的手工流程或 Node 工具,而不是留下名不副实的说明。
|
|
41
|
-
|
|
42
|
-
### 3. 写 KodaX 风格的 skill
|
|
43
|
-
|
|
44
|
-
- `description` 要写“做什么 + 什么时候用”,并稍微主动一点,避免 under-trigger。
|
|
45
|
-
- `SKILL.md` 负责主流程,不要把所有细节都塞进去。
|
|
46
|
-
- 重复性、机械性、易出错的步骤,放到 `scripts/`。
|
|
47
|
-
- 大块参考资料放到 `references/`。
|
|
48
|
-
- 模板或静态文件放到 `assets/`。
|
|
49
|
-
- 如果某个专家流程只服务这个 skill,可以放进 `agents/`,但先把它当私有 contract,不要自动上升成全局产品概念。
|
|
50
|
-
- `description` 同时决定模型发现。自然语言只会看到未设置 `disable-model-invocation: true` 的 name/description。
|
|
51
|
-
- 需要只在显式 slash 后加载时,写 `disable-model-invocation: true`。这只关闭模型 catalog 与模型 `skill` 工具。
|
|
52
|
-
- 每个已启用 Skill 都接受 `/<name>` 和 `/skill:<name>`。
|
|
53
|
-
- token 可在 query 头部或中间。后续文本是 Skill 参数。
|
|
54
|
-
- `user-invocable` 只保留解析兼容。不要把它当成执行权限:宿主忽略该字段后,作者会误以为 `false` 能禁止 slash/SDK 调用。
|
|
55
|
-
|
|
56
|
-
### 4. Bundled scripts 默认用 Node.js
|
|
57
|
-
|
|
58
|
-
KodaX 当前会把 builtin skill 目录直接复制到 `dist/`。因此:
|
|
59
|
-
|
|
60
|
-
- skill 内的可执行脚本默认使用 plain Node ESM `.js`。
|
|
61
|
-
- 只有在你同时修改了构建链、确保脚本会被编译时,才在 skill 内使用 `.ts`。
|
|
62
|
-
- 如果用户只是想“改成 node/typescript”,默认先落成 Node `.js`,这是最稳妥的内建交付方式。
|
|
63
|
-
|
|
64
|
-
### 5. 先验证,再评估
|
|
65
|
-
|
|
66
|
-
起草完成后,优先按下面的顺序推进:
|
|
67
|
-
|
|
68
|
-
1. 用 `node scripts/quick-validate.js <skill-dir>` 做结构检查。
|
|
69
|
-
2. 如果还没有 skill 骨架,用 `node scripts/init-skill.js <skill-name> --path <skills-dir>` 初始化。
|
|
70
|
-
3. 设计 2 到 3 个真实用户会说的测试提示。
|
|
71
|
-
4. 如果要跑端到端 skill eval,把提示整理到 `evals/evals.json`,再用 `node scripts/run-eval.js --skill-path <skill-dir> --evals <evals.json> --workspace <iteration-dir>` 生成 `with_skill` / `without_skill` workspace。
|
|
72
|
-
5. 如果要补第 3 阶段的专家评测流,先用 `node scripts/grade-evals.js <workspace>` 生成 `grading.json`,再用 `node scripts/aggregate-benchmark.js <workspace> --skill-name <name>` 聚合 benchmark,用 `node scripts/analyze-benchmark.js <workspace>` 产出分析结论,用 `node scripts/compare-runs.js <workspace>` 做 blind comparison。
|
|
73
|
-
6. 如果要评估 description 的触发效果,再用 `node scripts/run-trigger-eval.js --skill-path <skill-dir> --evals <evals.json>` 跑一轮触发评测。
|
|
74
|
-
7. 如果要迭代 description,可以用 `node scripts/improve-description.js --skill-path <skill-dir> --eval-results <results.json>` 生成候选描述,或用 `node scripts/run-loop.js --skill-path <skill-dir> --evals <evals.json> --workspace <workspace-dir>` 跑多轮优化。
|
|
75
|
-
8. 如果需要人工 review,把运行结果整理到 workspace,再用 `node scripts/generate-review.js <workspace> --static <html-file>` 或本地服务模式生成 review 页面。
|
|
76
|
-
9. 如果要分享给别的 KodaX/Agent Skills 风格环境,用 `node scripts/package-skill.js <skill-dir>` 打包,再用 `node scripts/install-skill.js <archive-or-dir>` 验证安装链路。
|
|
77
|
-
|
|
78
|
-
## 评估建议
|
|
79
|
-
|
|
80
|
-
- 客观任务:优先写断言、grading 结构和 benchmark。
|
|
81
|
-
- 主观任务:优先给人类 review 页面,再用 comparator 做 blind comparison,而不是强行只看单一分数。
|
|
82
|
-
- 描述优化:先整理误触发/漏触发样例,再跑 `run-trigger-eval`,需要时再用 `improve-description` 或 `run-loop` 迭代。
|
|
83
|
-
- 如果需要专家提示词,把 `agents/grader.md`、`agents/analyzer.md`、`agents/comparator.md` 当作私有专家 contract 使用,不要先把它们产品化成通用 swarm 概念。
|
|
84
|
-
|
|
85
|
-
## 输出要求
|
|
86
|
-
|
|
87
|
-
默认给出:
|
|
88
|
-
- 修改后的 `SKILL.md`
|
|
89
|
-
- 新增或更新的 supporting files
|
|
90
|
-
- 简短的 trigger/eval 样例
|
|
91
|
-
- 还没覆盖的风险或后续建议
|
|
92
|
-
|
|
93
|
-
如果用户是在移植外部 skill,还要额外说明:
|
|
94
|
-
- 哪些能力已经迁移
|
|
95
|
-
- 哪些能力因为宿主差异被删减或改写
|
|
96
|
-
- 哪些部分后续值得继续产品化
|
|
97
|
-
|
|
98
|
-
## 可用工具
|
|
99
|
-
|
|
100
|
-
- `agents/grader.md`:给 `grade-evals.js` 使用的专家评分契约。
|
|
101
|
-
- `agents/analyzer.md`:给 `analyze-benchmark.js` 使用的分析契约。
|
|
102
|
-
- `agents/comparator.md`:给 `compare-runs.js` 使用的盲比契约。
|
|
103
|
-
- `scripts/quick-validate.js`:校验 skill 结构和 frontmatter。
|
|
104
|
-
- `scripts/init-skill.js`:初始化一个新的 skill 骨架,并可一并创建 `evals/evals.json`。
|
|
105
|
-
- `scripts/run-eval.js`:运行端到端 skill eval,生成 `with_skill` / `without_skill` workspace 结果。
|
|
106
|
-
- `scripts/grade-evals.js`:消费 workspace,给每个 run 生成 `grading.json` 与 `grading-summary.json`。
|
|
107
|
-
- `scripts/aggregate-benchmark.js`:聚合 `grading.json` / `timing.json` 生成 `benchmark.json` 与 `benchmark.md`。
|
|
108
|
-
- `scripts/analyze-benchmark.js`:基于 benchmark 和 grading 产出 `analysis.json` 与 `analysis.md`。
|
|
109
|
-
- `scripts/compare-runs.js`:对两个 config 做 blind comparison,生成 `comparison.json` 与 `comparison.md`。
|
|
110
|
-
- `scripts/run-trigger-eval.js`:对 description 做 KodaX 原生触发评测,检查误触发和漏触发。
|
|
111
|
-
- `scripts/improve-description.js`:基于评测结果生成新的 description 候选。
|
|
112
|
-
- `scripts/run-loop.js`:把触发评测和 description 改写串成可重复的多轮优化流程。
|
|
113
|
-
- `scripts/generate-review.js`:把 workspace 结果生成静态或本地服务版 HTML review 页面。
|
|
114
|
-
- `scripts/package-skill.js`:把 skill 目录打成 `.skill` 归档,便于分享与分发。
|
|
115
|
-
- `scripts/install-skill.js`:把 `.skill` 归档或目录安装到目标 skills 目录。
|
|
116
|
-
- `references/schemas.md`:评测相关 JSON 结构参考。
|
|
117
|
-
|
|
118
|
-
这里的 description eval、loop、grading、analysis、comparison 和 packaging 都已经是 KodaX 原生实现,不再依赖 Anthropic 的 Python 脚本或 Claude Code 专有宿主能力。
|
|
119
|
-
|
|
120
|
-
## 使用示例
|
|
121
|
-
|
|
122
|
-
- `/skill:skill-creator 把这个 Claude skill 迁移成 KodaX builtin`
|
|
123
|
-
- `/skill:skill-creator 新建一个 release-notes skill`
|
|
124
|
-
- `/skill:skill-creator 优化现有 skill 的 description 和 evals`
|
|
125
|
-
- `/skill:skill-creator 给这个 skill 补 trigger eval、grading、benchmark、comparison 和 review 流程`
|
|
126
|
-
- `/skill:skill-creator 把这个 skill 打成可分享的 .skill 包`
|
|
127
|
-
- `/skill:skill-creator 初始化一个新 skill 骨架并生成 eval workspace`
|
|
1
|
+
---
|
|
2
|
+
name: skill-creator
|
|
3
|
+
description: 创建、重写、迁移和优化 KodaX/Agent Skills。当用户想新建 skill、把外部 skill 迁移到 KodaX、改进触发描述、整理 supporting files、设计评测用例、补齐 grading/benchmark/review/comparison 流程或验证 skill 结构时使用。即使用户没有明确说“skill”,只要目标是在沉淀可复用的代理工作流、提示词或脚本能力,也应该使用这个 skill。
|
|
4
|
+
user-invocable: true
|
|
5
|
+
allowed-tools: "Read, Grep, Glob, Write, Edit, Bash(node:*, npm:*, npx:*)"
|
|
6
|
+
argument-hint: "[skill-name-or-task]"
|
|
7
|
+
compatibility: "Optimized for KodaX and Agent Skills style directories. Bundled helper scripts use Node.js instead of Python."
|
|
8
|
+
---
|
|
9
|
+
|
|
10
|
+
# Skill Creator
|
|
11
|
+
|
|
12
|
+
把用户的工作流整理成一个可维护、可触发、可评估的 skill。优先适配 KodaX 的 skill 运行方式,而不是逐字复制外部平台的实现。
|
|
13
|
+
|
|
14
|
+
## 何时使用
|
|
15
|
+
|
|
16
|
+
- 用户要新建 skill,或把一次对话里的工作流沉淀成 skill。
|
|
17
|
+
- 用户要改已有 skill 的触发描述、结构、supporting files 或提示词。
|
|
18
|
+
- 用户要把 Claude / Anthropic / 其他平台的 skill 移植到 KodaX。
|
|
19
|
+
- 用户要为 skill 设计测试提示、评估结构、人工 review 流程或 benchmark 汇总。
|
|
20
|
+
|
|
21
|
+
## 工作方式
|
|
22
|
+
|
|
23
|
+
### 1. 先收敛目标
|
|
24
|
+
|
|
25
|
+
先明确四件事:
|
|
26
|
+
1. 这个 skill 要解决什么任务。
|
|
27
|
+
2. 什么时候应该触发,什么时候不应该触发。
|
|
28
|
+
3. 输出是什么形态。
|
|
29
|
+
4. 是否需要评测和人工 review。
|
|
30
|
+
|
|
31
|
+
如果用户已经给了样例对话、提示词或外部 skill 仓库,先从已有材料里提炼,不要重复让用户描述。
|
|
32
|
+
|
|
33
|
+
### 2. 适配 KodaX,而不是照抄外部 skill
|
|
34
|
+
|
|
35
|
+
迁移外部 skill 时,拆成三类:
|
|
36
|
+
- 可直接复用:SKILL.md 的思路、评测结构、参考文档组织方式。
|
|
37
|
+
- 需要改写:路径约定、触发描述、支持的 frontmatter 字段、命令示例。
|
|
38
|
+
- 不要硬搬:强依赖 Claude Code、`claude -p`、Cowork、Python stdlib 或专有事件流的部分。
|
|
39
|
+
|
|
40
|
+
如果外部 skill 依赖特定宿主能力,优先改成 KodaX 当前能承接的手工流程或 Node 工具,而不是留下名不副实的说明。
|
|
41
|
+
|
|
42
|
+
### 3. 写 KodaX 风格的 skill
|
|
43
|
+
|
|
44
|
+
- `description` 要写“做什么 + 什么时候用”,并稍微主动一点,避免 under-trigger。
|
|
45
|
+
- `SKILL.md` 负责主流程,不要把所有细节都塞进去。
|
|
46
|
+
- 重复性、机械性、易出错的步骤,放到 `scripts/`。
|
|
47
|
+
- 大块参考资料放到 `references/`。
|
|
48
|
+
- 模板或静态文件放到 `assets/`。
|
|
49
|
+
- 如果某个专家流程只服务这个 skill,可以放进 `agents/`,但先把它当私有 contract,不要自动上升成全局产品概念。
|
|
50
|
+
- `description` 同时决定模型发现。自然语言只会看到未设置 `disable-model-invocation: true` 的 name/description。
|
|
51
|
+
- 需要只在显式 slash 后加载时,写 `disable-model-invocation: true`。这只关闭模型 catalog 与模型 `skill` 工具。
|
|
52
|
+
- 每个已启用 Skill 都接受 `/<name>` 和 `/skill:<name>`。
|
|
53
|
+
- token 可在 query 头部或中间。后续文本是 Skill 参数。
|
|
54
|
+
- `user-invocable` 只保留解析兼容。不要把它当成执行权限:宿主忽略该字段后,作者会误以为 `false` 能禁止 slash/SDK 调用。
|
|
55
|
+
|
|
56
|
+
### 4. Bundled scripts 默认用 Node.js
|
|
57
|
+
|
|
58
|
+
KodaX 当前会把 builtin skill 目录直接复制到 `dist/`。因此:
|
|
59
|
+
|
|
60
|
+
- skill 内的可执行脚本默认使用 plain Node ESM `.js`。
|
|
61
|
+
- 只有在你同时修改了构建链、确保脚本会被编译时,才在 skill 内使用 `.ts`。
|
|
62
|
+
- 如果用户只是想“改成 node/typescript”,默认先落成 Node `.js`,这是最稳妥的内建交付方式。
|
|
63
|
+
|
|
64
|
+
### 5. 先验证,再评估
|
|
65
|
+
|
|
66
|
+
起草完成后,优先按下面的顺序推进:
|
|
67
|
+
|
|
68
|
+
1. 用 `node scripts/quick-validate.js <skill-dir>` 做结构检查。
|
|
69
|
+
2. 如果还没有 skill 骨架,用 `node scripts/init-skill.js <skill-name> --path <skills-dir>` 初始化。
|
|
70
|
+
3. 设计 2 到 3 个真实用户会说的测试提示。
|
|
71
|
+
4. 如果要跑端到端 skill eval,把提示整理到 `evals/evals.json`,再用 `node scripts/run-eval.js --skill-path <skill-dir> --evals <evals.json> --workspace <iteration-dir>` 生成 `with_skill` / `without_skill` workspace。
|
|
72
|
+
5. 如果要补第 3 阶段的专家评测流,先用 `node scripts/grade-evals.js <workspace>` 生成 `grading.json`,再用 `node scripts/aggregate-benchmark.js <workspace> --skill-name <name>` 聚合 benchmark,用 `node scripts/analyze-benchmark.js <workspace>` 产出分析结论,用 `node scripts/compare-runs.js <workspace>` 做 blind comparison。
|
|
73
|
+
6. 如果要评估 description 的触发效果,再用 `node scripts/run-trigger-eval.js --skill-path <skill-dir> --evals <evals.json>` 跑一轮触发评测。
|
|
74
|
+
7. 如果要迭代 description,可以用 `node scripts/improve-description.js --skill-path <skill-dir> --eval-results <results.json>` 生成候选描述,或用 `node scripts/run-loop.js --skill-path <skill-dir> --evals <evals.json> --workspace <workspace-dir>` 跑多轮优化。
|
|
75
|
+
8. 如果需要人工 review,把运行结果整理到 workspace,再用 `node scripts/generate-review.js <workspace> --static <html-file>` 或本地服务模式生成 review 页面。
|
|
76
|
+
9. 如果要分享给别的 KodaX/Agent Skills 风格环境,用 `node scripts/package-skill.js <skill-dir>` 打包,再用 `node scripts/install-skill.js <archive-or-dir>` 验证安装链路。
|
|
77
|
+
|
|
78
|
+
## 评估建议
|
|
79
|
+
|
|
80
|
+
- 客观任务:优先写断言、grading 结构和 benchmark。
|
|
81
|
+
- 主观任务:优先给人类 review 页面,再用 comparator 做 blind comparison,而不是强行只看单一分数。
|
|
82
|
+
- 描述优化:先整理误触发/漏触发样例,再跑 `run-trigger-eval`,需要时再用 `improve-description` 或 `run-loop` 迭代。
|
|
83
|
+
- 如果需要专家提示词,把 `agents/grader.md`、`agents/analyzer.md`、`agents/comparator.md` 当作私有专家 contract 使用,不要先把它们产品化成通用 swarm 概念。
|
|
84
|
+
|
|
85
|
+
## 输出要求
|
|
86
|
+
|
|
87
|
+
默认给出:
|
|
88
|
+
- 修改后的 `SKILL.md`
|
|
89
|
+
- 新增或更新的 supporting files
|
|
90
|
+
- 简短的 trigger/eval 样例
|
|
91
|
+
- 还没覆盖的风险或后续建议
|
|
92
|
+
|
|
93
|
+
如果用户是在移植外部 skill,还要额外说明:
|
|
94
|
+
- 哪些能力已经迁移
|
|
95
|
+
- 哪些能力因为宿主差异被删减或改写
|
|
96
|
+
- 哪些部分后续值得继续产品化
|
|
97
|
+
|
|
98
|
+
## 可用工具
|
|
99
|
+
|
|
100
|
+
- `agents/grader.md`:给 `grade-evals.js` 使用的专家评分契约。
|
|
101
|
+
- `agents/analyzer.md`:给 `analyze-benchmark.js` 使用的分析契约。
|
|
102
|
+
- `agents/comparator.md`:给 `compare-runs.js` 使用的盲比契约。
|
|
103
|
+
- `scripts/quick-validate.js`:校验 skill 结构和 frontmatter。
|
|
104
|
+
- `scripts/init-skill.js`:初始化一个新的 skill 骨架,并可一并创建 `evals/evals.json`。
|
|
105
|
+
- `scripts/run-eval.js`:运行端到端 skill eval,生成 `with_skill` / `without_skill` workspace 结果。
|
|
106
|
+
- `scripts/grade-evals.js`:消费 workspace,给每个 run 生成 `grading.json` 与 `grading-summary.json`。
|
|
107
|
+
- `scripts/aggregate-benchmark.js`:聚合 `grading.json` / `timing.json` 生成 `benchmark.json` 与 `benchmark.md`。
|
|
108
|
+
- `scripts/analyze-benchmark.js`:基于 benchmark 和 grading 产出 `analysis.json` 与 `analysis.md`。
|
|
109
|
+
- `scripts/compare-runs.js`:对两个 config 做 blind comparison,生成 `comparison.json` 与 `comparison.md`。
|
|
110
|
+
- `scripts/run-trigger-eval.js`:对 description 做 KodaX 原生触发评测,检查误触发和漏触发。
|
|
111
|
+
- `scripts/improve-description.js`:基于评测结果生成新的 description 候选。
|
|
112
|
+
- `scripts/run-loop.js`:把触发评测和 description 改写串成可重复的多轮优化流程。
|
|
113
|
+
- `scripts/generate-review.js`:把 workspace 结果生成静态或本地服务版 HTML review 页面。
|
|
114
|
+
- `scripts/package-skill.js`:把 skill 目录打成 `.skill` 归档,便于分享与分发。
|
|
115
|
+
- `scripts/install-skill.js`:把 `.skill` 归档或目录安装到目标 skills 目录。
|
|
116
|
+
- `references/schemas.md`:评测相关 JSON 结构参考。
|
|
117
|
+
|
|
118
|
+
这里的 description eval、loop、grading、analysis、comparison 和 packaging 都已经是 KodaX 原生实现,不再依赖 Anthropic 的 Python 脚本或 Claude Code 专有宿主能力。
|
|
119
|
+
|
|
120
|
+
## 使用示例
|
|
121
|
+
|
|
122
|
+
- `/skill:skill-creator 把这个 Claude skill 迁移成 KodaX builtin`
|
|
123
|
+
- `/skill:skill-creator 新建一个 release-notes skill`
|
|
124
|
+
- `/skill:skill-creator 优化现有 skill 的 description 和 evals`
|
|
125
|
+
- `/skill:skill-creator 给这个 skill 补 trigger eval、grading、benchmark、comparison 和 review 流程`
|
|
126
|
+
- `/skill:skill-creator 把这个 skill 打成可分享的 .skill 包`
|
|
127
|
+
- `/skill:skill-creator 初始化一个新 skill 骨架并生成 eval workspace`
|
|
@@ -1,12 +1,12 @@
|
|
|
1
|
-
# Analyzer
|
|
2
|
-
|
|
3
|
-
You are the benchmark analysis specialist for KodaX skill evaluation workspaces.
|
|
4
|
-
|
|
5
|
-
Your job is to look across benchmark and grading artifacts, identify signal vs noise, and recommend the next iteration.
|
|
6
|
-
|
|
7
|
-
Rules:
|
|
8
|
-
- Focus on stable deltas, variance, repeated failures, and likely weak assertions.
|
|
9
|
-
- Separate real improvement from measurement noise.
|
|
10
|
-
- Prefer specific, operational recommendations over generic advice.
|
|
11
|
-
- Call out when the data is inconclusive.
|
|
12
|
-
- Return JSON only.
|
|
1
|
+
# Analyzer
|
|
2
|
+
|
|
3
|
+
You are the benchmark analysis specialist for KodaX skill evaluation workspaces.
|
|
4
|
+
|
|
5
|
+
Your job is to look across benchmark and grading artifacts, identify signal vs noise, and recommend the next iteration.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
- Focus on stable deltas, variance, repeated failures, and likely weak assertions.
|
|
9
|
+
- Separate real improvement from measurement noise.
|
|
10
|
+
- Prefer specific, operational recommendations over generic advice.
|
|
11
|
+
- Call out when the data is inconclusive.
|
|
12
|
+
- Return JSON only.
|
|
@@ -1,13 +1,13 @@
|
|
|
1
|
-
# Comparator
|
|
2
|
-
|
|
3
|
-
You are the blind comparison specialist for KodaX skill eval outputs.
|
|
4
|
-
|
|
5
|
-
Your job is to compare candidate outputs without relying on config names or implementation details.
|
|
6
|
-
|
|
7
|
-
Rules:
|
|
8
|
-
- Judge against the eval prompt, expected outcome, and explicit assertions.
|
|
9
|
-
- Compare only the quality of the visible outputs.
|
|
10
|
-
- Prefer clearer, more complete, and less risky answers.
|
|
11
|
-
- Use `tie` when both outputs are similarly strong or similarly weak.
|
|
12
|
-
- Use `inconclusive` when the prompt does not provide enough evidence to decide.
|
|
13
|
-
- Return JSON only.
|
|
1
|
+
# Comparator
|
|
2
|
+
|
|
3
|
+
You are the blind comparison specialist for KodaX skill eval outputs.
|
|
4
|
+
|
|
5
|
+
Your job is to compare candidate outputs without relying on config names or implementation details.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
- Judge against the eval prompt, expected outcome, and explicit assertions.
|
|
9
|
+
- Compare only the quality of the visible outputs.
|
|
10
|
+
- Prefer clearer, more complete, and less risky answers.
|
|
11
|
+
- Use `tie` when both outputs are similarly strong or similarly weak.
|
|
12
|
+
- Use `inconclusive` when the prompt does not provide enough evidence to decide.
|
|
13
|
+
- Return JSON only.
|
|
@@ -1,13 +1,13 @@
|
|
|
1
|
-
# Grader
|
|
2
|
-
|
|
3
|
-
You are the grading specialist for KodaX skill eval runs.
|
|
4
|
-
|
|
5
|
-
Your job is to judge one run against the eval prompt, expected outcome, and explicit assertions.
|
|
6
|
-
|
|
7
|
-
Rules:
|
|
8
|
-
- Judge only what is visible in the provided artifacts.
|
|
9
|
-
- Do not assume hidden behavior or give credit for intentions.
|
|
10
|
-
- Prefer concrete evidence from the final output.
|
|
11
|
-
- If an expectation is only partially satisfied, mark it as failed and explain why.
|
|
12
|
-
- Keep uncertainty explicit instead of guessing.
|
|
13
|
-
- Return JSON only.
|
|
1
|
+
# Grader
|
|
2
|
+
|
|
3
|
+
You are the grading specialist for KodaX skill eval runs.
|
|
4
|
+
|
|
5
|
+
Your job is to judge one run against the eval prompt, expected outcome, and explicit assertions.
|
|
6
|
+
|
|
7
|
+
Rules:
|
|
8
|
+
- Judge only what is visible in the provided artifacts.
|
|
9
|
+
- Do not assume hidden behavior or give credit for intentions.
|
|
10
|
+
- Prefer concrete evidence from the final output.
|
|
11
|
+
- If an expectation is only partially satisfied, mark it as failed and explain why.
|
|
12
|
+
- Keep uncertainty explicit instead of guessing.
|
|
13
|
+
- Return JSON only.
|