oh-my-knowledge 0.27.0 → 0.28.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (122) hide show
  1. package/README.md +31 -34
  2. package/README.zh.md +31 -34
  3. package/dist/src/analysis/report-diagnostics.d.ts +2 -2
  4. package/dist/src/analysis/report-diagnostics.js +2 -2
  5. package/dist/src/analysis/sample-diagnostics.d.ts +3 -3
  6. package/dist/src/analysis/sample-diagnostics.js +3 -3
  7. package/dist/src/authoring/generator.js +1 -1
  8. package/dist/src/authoring/generator.js.map +1 -1
  9. package/dist/src/cli/commands/doctor.d.ts.map +1 -1
  10. package/dist/src/cli/commands/doctor.js +39 -2
  11. package/dist/src/cli/commands/doctor.js.map +1 -1
  12. package/dist/src/cli/commands/eval.d.ts.map +1 -1
  13. package/dist/src/cli/commands/eval.js +0 -5
  14. package/dist/src/cli/commands/eval.js.map +1 -1
  15. package/dist/src/cli/commands/{export.d.ts → evolve.d.ts} +1 -1
  16. package/dist/src/cli/commands/evolve.d.ts.map +1 -0
  17. package/dist/src/cli/commands/{improve-skill.js → evolve.js} +3 -3
  18. package/dist/src/cli/commands/evolve.js.map +1 -0
  19. package/dist/src/cli/commands/registry.d.ts.map +1 -1
  20. package/dist/src/cli/commands/registry.js +4 -22
  21. package/dist/src/cli/commands/registry.js.map +1 -1
  22. package/dist/src/cli/commands/{improve.d.ts → sample.d.ts} +1 -1
  23. package/dist/src/cli/commands/sample.d.ts.map +1 -0
  24. package/dist/src/cli/commands/{improve-samples.js → sample.js} +5 -4
  25. package/dist/src/cli/commands/sample.js.map +1 -0
  26. package/dist/src/cli/i18n-dict.d.ts +2 -4
  27. package/dist/src/cli/i18n-dict.d.ts.map +1 -1
  28. package/dist/src/cli/i18n-dict.js +165 -373
  29. package/dist/src/cli/i18n-dict.js.map +1 -1
  30. package/dist/src/cli/parse-run-config.d.ts +1 -1
  31. package/dist/src/cli/parse-run-config.js +1 -1
  32. package/dist/src/doctor/health/builtin-dimensions.d.ts +11 -0
  33. package/dist/src/doctor/health/builtin-dimensions.d.ts.map +1 -0
  34. package/dist/src/doctor/health/builtin-dimensions.js +94 -0
  35. package/dist/src/doctor/health/builtin-dimensions.js.map +1 -0
  36. package/dist/src/doctor/health/composer.d.ts +19 -0
  37. package/dist/src/doctor/health/composer.d.ts.map +1 -0
  38. package/dist/src/doctor/health/composer.js +289 -0
  39. package/dist/src/doctor/health/composer.js.map +1 -0
  40. package/dist/src/doctor/health/dimension-registry.d.ts +13 -0
  41. package/dist/src/doctor/health/dimension-registry.d.ts.map +1 -0
  42. package/dist/src/doctor/health/dimension-registry.js +28 -0
  43. package/dist/src/doctor/health/dimension-registry.js.map +1 -0
  44. package/dist/src/doctor/health/dimension-spec.d.ts +46 -0
  45. package/dist/src/doctor/health/dimension-spec.d.ts.map +1 -0
  46. package/dist/src/doctor/health/dimension-spec.js +12 -0
  47. package/dist/src/doctor/health/dimension-spec.js.map +1 -0
  48. package/dist/src/doctor/health/parser.d.ts +27 -0
  49. package/dist/src/doctor/health/parser.d.ts.map +1 -0
  50. package/dist/src/doctor/health/parser.js +190 -0
  51. package/dist/src/doctor/health/parser.js.map +1 -0
  52. package/dist/src/doctor/health/prompt-builder.d.ts +22 -0
  53. package/dist/src/doctor/health/prompt-builder.d.ts.map +1 -0
  54. package/dist/src/doctor/health/prompt-builder.js +162 -0
  55. package/dist/src/doctor/health/prompt-builder.js.map +1 -0
  56. package/dist/src/doctor/health/register.d.ts +13 -0
  57. package/dist/src/doctor/health/register.d.ts.map +1 -0
  58. package/dist/src/doctor/health/register.js +20 -0
  59. package/dist/src/doctor/health/register.js.map +1 -0
  60. package/dist/src/doctor/html-renderer.d.ts +20 -0
  61. package/dist/src/doctor/html-renderer.d.ts.map +1 -0
  62. package/dist/src/doctor/html-renderer.js +366 -0
  63. package/dist/src/doctor/html-renderer.js.map +1 -0
  64. package/dist/src/doctor/index.d.ts.map +1 -1
  65. package/dist/src/doctor/index.js +55 -6
  66. package/dist/src/doctor/index.js.map +1 -1
  67. package/dist/src/doctor/renderer.d.ts +4 -0
  68. package/dist/src/doctor/renderer.d.ts.map +1 -1
  69. package/dist/src/doctor/renderer.js +72 -18
  70. package/dist/src/doctor/renderer.js.map +1 -1
  71. package/dist/src/doctor/rules.d.ts +6 -5
  72. package/dist/src/doctor/rules.d.ts.map +1 -1
  73. package/dist/src/doctor/rules.js +4 -3
  74. package/dist/src/doctor/rules.js.map +1 -1
  75. package/dist/src/types/doctor.d.ts +41 -3
  76. package/dist/src/types/doctor.d.ts.map +1 -1
  77. package/dist/src/types/doctor.js +3 -0
  78. package/dist/src/types/doctor.js.map +1 -1
  79. package/dist/src/types/eval.d.ts +1 -1
  80. package/dist/src/types/report.d.ts +2 -2
  81. package/dist/src/types/report.d.ts.map +1 -1
  82. package/package.json +1 -1
  83. package/dist/src/cli/commands/eval-debias.d.ts +0 -2
  84. package/dist/src/cli/commands/eval-debias.d.ts.map +0 -1
  85. package/dist/src/cli/commands/eval-debias.js +0 -88
  86. package/dist/src/cli/commands/eval-debias.js.map +0 -1
  87. package/dist/src/cli/commands/export-diff.d.ts +0 -2
  88. package/dist/src/cli/commands/export-diff.d.ts.map +0 -1
  89. package/dist/src/cli/commands/export-diff.js +0 -177
  90. package/dist/src/cli/commands/export-diff.js.map +0 -1
  91. package/dist/src/cli/commands/export-saturation.d.ts +0 -2
  92. package/dist/src/cli/commands/export-saturation.d.ts.map +0 -1
  93. package/dist/src/cli/commands/export-saturation.js +0 -59
  94. package/dist/src/cli/commands/export-saturation.js.map +0 -1
  95. package/dist/src/cli/commands/export-verdict.d.ts +0 -2
  96. package/dist/src/cli/commands/export-verdict.d.ts.map +0 -1
  97. package/dist/src/cli/commands/export-verdict.js +0 -45
  98. package/dist/src/cli/commands/export-verdict.js.map +0 -1
  99. package/dist/src/cli/commands/export.d.ts.map +0 -1
  100. package/dist/src/cli/commands/export.js +0 -150
  101. package/dist/src/cli/commands/export.js.map +0 -1
  102. package/dist/src/cli/commands/improve-failures.d.ts +0 -2
  103. package/dist/src/cli/commands/improve-failures.d.ts.map +0 -1
  104. package/dist/src/cli/commands/improve-failures.js +0 -59
  105. package/dist/src/cli/commands/improve-failures.js.map +0 -1
  106. package/dist/src/cli/commands/improve-plan.d.ts +0 -2
  107. package/dist/src/cli/commands/improve-plan.d.ts.map +0 -1
  108. package/dist/src/cli/commands/improve-plan.js +0 -75
  109. package/dist/src/cli/commands/improve-plan.js.map +0 -1
  110. package/dist/src/cli/commands/improve-samples.d.ts +0 -2
  111. package/dist/src/cli/commands/improve-samples.d.ts.map +0 -1
  112. package/dist/src/cli/commands/improve-samples.js.map +0 -1
  113. package/dist/src/cli/commands/improve-skill.d.ts +0 -2
  114. package/dist/src/cli/commands/improve-skill.d.ts.map +0 -1
  115. package/dist/src/cli/commands/improve-skill.js.map +0 -1
  116. package/dist/src/cli/commands/improve.d.ts.map +0 -1
  117. package/dist/src/cli/commands/improve.js +0 -34
  118. package/dist/src/cli/commands/improve.js.map +0 -1
  119. package/dist/src/cli/coverage-renderer.d.ts +0 -15
  120. package/dist/src/cli/coverage-renderer.d.ts.map +0 -1
  121. package/dist/src/cli/coverage-renderer.js +0 -74
  122. package/dist/src/cli/coverage-renderer.js.map +0 -1
@@ -16,9 +16,7 @@
16
16
  * 2. **保留原文的白名单 (产品术语 / 命令 / 文件名)**
17
17
  * 以下 token 在两种语言里都保留原文, 不翻译:
18
18
  * - 产品名: omk, oh-my-knowledge, Claude, npm
19
- * - 子命令空间和命令名: init, doctor, eval, observe, improve, export,
20
- * studio, samples, skill, plan, failures, gold, debias, diff, verdict,
21
- * saturation
19
+ * - 命令名: init, doctor, eval, observe, evolve, sample, studio, gold
22
20
  * - omk 核心业务术语: skill, variant, sample, judge, executor (出现在产品
23
21
  * UI 里时首字母可大写如 "Skill 评测", 描述句中保持小写)
24
22
  * - 技术参数: --lang, --control, --treatment, --bootstrap, --judge-repeat,
@@ -213,8 +211,8 @@ export const CLI_DICT = {
213
211
  en: '\n💡 Non-interactive environment, skipping report server\n',
214
212
  },
215
213
  'cli.run.no_serve_view_hint': {
216
- zh: ' 导出报告:omk export {id} --reports-dir {dir}\n',
217
- en: ' Export report: omk export {id} --reports-dir {dir}\n',
214
+ zh: ' 查看报告:omk studio --reports-dir {dir}(报告 ID:{id})\n',
215
+ en: ' View report: omk studio --reports-dir {dir} (report id: {id})\n',
218
216
  },
219
217
  'cli.run.gold_load_failed': {
220
218
  zh: '\n⚠ gold dataset 加载失败 ({dir}):\n',
@@ -260,18 +258,6 @@ export const CLI_DICT = {
260
258
  zh: '⚠ 加载 samples 文件失败 ({path}): {message}\n',
261
259
  en: '⚠ Failed to load samples file ({path}): {message}\n',
262
260
  },
263
- 'cli.export.unsupported_format': {
264
- zh: '不支持的导出格式:{format}。可用格式:html / markdown / github-summary。',
265
- en: 'Unsupported export format: {format}. Available: html / markdown / github-summary.',
266
- },
267
- 'cli.export.html_done': {
268
- zh: '已导出 HTML:{path}',
269
- en: 'HTML exported to: {path}',
270
- },
271
- 'cli.export.done': {
272
- zh: '已导出:{path}',
273
- en: 'Exported to: {path}',
274
- },
275
261
  'cli.studio.started': {
276
262
  zh: 'studio 已启动:{url}',
277
263
  en: 'Studio running at {url}',
@@ -309,8 +295,8 @@ export const CLI_DICT = {
309
295
  en: '\nGenerated {n} eval-samples files. Review them, then run: omk eval --batch',
310
296
  },
311
297
  'cli.gen.specify_skill_path': {
312
- zh: '请指定 skill 文件路径, 例如: omk improve samples skills/my-skill.md',
313
- en: 'Please specify a skill file path, e.g.: omk improve samples skills/my-skill.md',
298
+ zh: '请指定 skill 文件路径, 例如: omk sample skills/my-skill.md',
299
+ en: 'Please specify a skill file path, e.g.: omk sample skills/my-skill.md',
314
300
  },
315
301
  'cli.gen.samples_already_exists': {
316
302
  zh: 'eval-samples.json 已存在。如需覆盖请先删除该文件。',
@@ -333,8 +319,8 @@ export const CLI_DICT = {
333
319
  en: 'Generation failed: {message}',
334
320
  },
335
321
  'cli.evolve.specify_skill_path': {
336
- zh: '请指定 skill 文件路径, 例如: omk improve skill skills/my-skill.md',
337
- en: 'Please specify a skill file path, e.g.: omk improve skill skills/my-skill.md',
322
+ zh: '请指定 skill 文件路径, 例如: omk evolve skills/my-skill.md',
323
+ en: 'Please specify a skill file path, e.g.: omk evolve skills/my-skill.md',
338
324
  },
339
325
  'cli.evolve.section_header': {
340
326
  zh: '\n=== Improve skill: {path} ===\n',
@@ -365,8 +351,8 @@ export const CLI_DICT = {
365
351
  en: 'All versions saved at: {dir}/\n',
366
352
  },
367
353
  'cli.evolve.report_link': {
368
- zh: '📊 评测报告:omk export {id} --format html\n',
369
- en: '📊 Report: omk export {id} --format html\n',
354
+ zh: '📊 查看报告:omk studio(报告 ID:{id})\n',
355
+ en: '📊 View report: omk studio (report id: {id})\n',
370
356
  },
371
357
  'cli.help.product_main': {
372
358
  zh: `
@@ -374,19 +360,18 @@ oh-my-knowledge — 知识载体工作台
374
360
 
375
361
  用法:
376
362
  omk init [dir] 初始化一个 skill 评测项目
377
- omk doctor [path] 静态健康检查:结构、依赖、样本、配置、污染风险
363
+ omk doctor [path] LLM 健康度审计(7 内置维度 + 可扩展);--static-only 切离线静态模式
378
364
  omk eval [options] 离线评测:比较版本,输出 verdict + report
379
365
  omk observe <sessions-dir> 线上观测:真实 session、gap、失败率、inbox
380
- omk improve <report-id> 改进建议:样本质量、失败聚类、skill patch 线索
381
- omk export <report-id> [options] 证据导出:PR / CI / audit
382
- omk studio 打开本地工作台
366
+ omk evolve <skill> 多轮自动迭代改进 skill
367
+ omk sample <skill> 生成或补齐 eval-samples 评测用例(或 --batch 批量模式)
368
+ omk studio 打开本地工作台浏览报告
383
369
 
384
370
  主路径:
385
371
  omk doctor
386
372
  omk eval --control code-review-v1 --treatment code-review-v2
387
373
  omk observe ~/.claude/projects/<project>
388
- omk improve <report-id>
389
- omk export <report-id> --format github-summary
374
+ omk evolve skills/code-review-v2/SKILL.md
390
375
  omk studio
391
376
 
392
377
  通用选项:
@@ -399,19 +384,18 @@ oh-my-knowledge — Knowledge Artifact Workbench
399
384
 
400
385
  Usage:
401
386
  omk init [dir] Scaffold a skill evaluation project
402
- omk doctor [path] Static health check: structure, deps, samples, config, contamination risk
387
+ omk doctor [path] LLM health audit (7 builtin dimensions, extensible); --static-only for offline static checks
403
388
  omk eval [options] Offline evaluation: compare versions, emit verdict + report
404
389
  omk observe <sessions-dir> Production observation: sessions, gaps, failure rate, inbox
405
- omk improve <report-id> Improvement advice: sample quality, failure clusters, skill patch hints
406
- omk export <report-id> [options] Evidence export for PR / CI / audit
407
- omk studio Open the local workbench
390
+ omk evolve <skill> Auto-iterate a skill through multi-round eval loops
391
+ omk sample <skill> Generate or fill eval-samples test cases (or --batch for all skills)
392
+ omk studio Open the local workbench to browse reports
408
393
 
409
394
  Main workflow:
410
395
  omk doctor
411
396
  omk eval --control code-review-v1 --treatment code-review-v2
412
397
  omk observe ~/.claude/projects/<project>
413
- omk improve <report-id>
414
- omk export <report-id> --format github-summary
398
+ omk evolve skills/code-review-v2/SKILL.md
415
399
  omk studio
416
400
 
417
401
  Common options:
@@ -461,7 +445,6 @@ omk eval — 离线评测 skill 版本,并给出 ship/no-ship verdict
461
445
  用法:
462
446
  omk eval --control <variant> --treatment <variant> [options]
463
447
  omk eval gold <init|validate|compare> ...
464
- omk eval debias length <report-id> ...
465
448
 
466
449
  常用选项:
467
450
  --samples <path> 用例文件(默认:eval-samples.json)
@@ -482,7 +465,8 @@ omk eval — 离线评测 skill 版本,并给出 ship/no-ship verdict
482
465
  --no-serve 评测后不自动启动报告 server
483
466
 
484
467
  示例:
485
- omk eval --control code-review-v1 --treatment code-review-v2
468
+ omk eval --control baseline --treatment my-skill # 单 skill 必要性测试(baseline 是保留 variant 名,代表「不注入 skill 的裸基线」)
469
+ omk eval --control code-review-v1 --treatment code-review-v2 # 多版本 A/B
486
470
  omk eval --config eval.yaml
487
471
  omk eval gold compare v1-vs-v2-20260505-1200 --gold-dir gold-dataset
488
472
  `,
@@ -492,7 +476,6 @@ omk eval — run offline skill evaluation and emit a ship/no-ship verdict
492
476
  Usage:
493
477
  omk eval --control <variant> --treatment <variant> [options]
494
478
  omk eval gold <init|validate|compare> ...
495
- omk eval debias length <report-id> ...
496
479
 
497
480
  Common options:
498
481
  --samples <path> Sample file (default: eval-samples.json)
@@ -513,7 +496,8 @@ Common options:
513
496
  --no-serve Do not auto-start report server after evaluation
514
497
 
515
498
  Examples:
516
- omk eval --control code-review-v1 --treatment code-review-v2
499
+ omk eval --control baseline --treatment my-skill # Single-skill necessity test (baseline is a reserved variant — "no skill injected")
500
+ omk eval --control code-review-v1 --treatment code-review-v2 # Multi-variant A/B
517
501
  omk eval --config eval.yaml
518
502
  omk eval gold compare v1-vs-v2-20260505-1200 --gold-dir gold-dataset
519
503
  `,
@@ -544,34 +528,6 @@ Options:
544
528
  --reports-dir <path> Reports directory for compare (default: ~/.oh-my-knowledge/reports)
545
529
  --variant <name> Variant in the report to compare
546
530
  --bootstrap-samples <n> Bootstrap resamples for compare
547
- `,
548
- },
549
- 'cli.help.eval_debias': {
550
- zh: `
551
- omk eval debias — 验证 length-debias 是否降低评委长度偏差
552
-
553
- 用法:
554
- omk eval debias length <reportId> [options]
555
-
556
- 选项:
557
- --samples <path> 用例文件;默认从 report.meta.request 读取
558
- --reports-dir <path> 报告目录(默认:~/.oh-my-knowledge/reports)
559
- --variant <name> 只验证指定 variant
560
- --judge-models <executor:model> 指定单评委
561
- --bootstrap-samples <n> bootstrap 重采样次数
562
- `,
563
- en: `
564
- omk eval debias — validate whether length-debias reduces judge length bias
565
-
566
- Usage:
567
- omk eval debias length <reportId> [options]
568
-
569
- Options:
570
- --samples <path> Sample file; defaults to report.meta.request
571
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports)
572
- --variant <name> Validate only one variant
573
- --judge-models <executor:model> Single judge to use
574
- --bootstrap-samples <n> Bootstrap resamples
575
531
  `,
576
532
  },
577
533
  'cli.help.observe': {
@@ -604,119 +560,51 @@ Options:
604
560
  --output-dir <path> Output directory (default: ~/.oh-my-knowledge/analyses)
605
561
  `,
606
562
  },
607
- 'cli.help.improve': {
563
+ 'cli.help.evolve': {
608
564
  zh: `
609
- omk improve — 从报告或 trace 中得到下一步改进建议
565
+ omk evolve — 多轮自动迭代改进 skill
610
566
 
611
567
  用法:
612
- omk improve <report-id> 输出样本质量诊断和改进计划
613
- omk improve plan <report-id> 同上,显式 plan 子命令
614
- omk improve failures <report-id> 聚类失败用例,生成根因和修复方向
615
- omk improve samples [skill] 为 skill 生成或补齐 eval samples
616
- omk improve skill <skill> 基于评测循环尝试改进 skill
568
+ omk evolve <skill-path> [options]
569
+
570
+ 选项:
571
+ --rounds <n> 迭代轮数(默认:5)
572
+ --target <score> 目标分数
573
+ --model <name> 任务执行模型,每轮跑 eval samples 的被测模型(默认:sonnet)
574
+ --improve-model <name> skill 改写模型,每轮根据反馈改写 skill 的模型(默认:sonnet)
575
+ --judge-models <executor:model> 单评委配置(默认:claude:haiku)
617
576
 
618
577
  示例:
619
- omk improve v1-vs-v2-20260505-1200
620
- omk improve samples skills/code-review/SKILL.md
621
- omk improve failures v1-vs-v2-20260505-1200
578
+ omk evolve skills/code-review/SKILL.md
579
+ omk evolve skills/code-review/SKILL.md --rounds 10 --target 4.5
580
+ omk evolve skills/code-review/SKILL.md --model sonnet --improve-model opus
622
581
  `,
623
582
  en: `
624
- omk improve — get next-step improvement advice from reports or traces
583
+ omk evolve — auto-iterate a skill through multi-round evaluation loops
625
584
 
626
585
  Usage:
627
- omk improve <report-id> Print sample diagnostics and repair plan
628
- omk improve plan <report-id> Same as above, explicit plan subcommand
629
- omk improve failures <report-id> Cluster failed cases into root causes and fixes
630
- omk improve samples [skill] Generate or fill eval samples for a skill
631
- omk improve skill <skill> Try to improve a skill through evaluation loops
586
+ omk evolve <skill-path> [options]
587
+
588
+ Options:
589
+ --rounds <n> Iteration rounds (default: 5)
590
+ --target <score> Target score
591
+ --model <name> Task executor model — runs eval samples each round (default: sonnet)
592
+ --improve-model <name> Skill rewriter model — rewrites the skill each round (default: sonnet)
593
+ --judge-models <executor:model> Single judge config (default: claude:haiku)
632
594
 
633
595
  Examples:
634
- omk improve v1-vs-v2-20260505-1200
635
- omk improve samples skills/code-review/SKILL.md
636
- omk improve failures v1-vs-v2-20260505-1200
596
+ omk evolve skills/code-review/SKILL.md
597
+ omk evolve skills/code-review/SKILL.md --rounds 10 --target 4.5
598
+ omk evolve skills/code-review/SKILL.md --model sonnet --improve-model opus
637
599
  `,
638
600
  },
639
- 'cli.help.improve_plan': {
640
- zh: [
641
- '',
642
- '用法: omk improve <reportId> [options]',
643
- ' omk improve plan <reportId> [options]',
644
- '',
645
- '诊断用例集本身的质量问题: 区分度低 / 重复 / 歧义 / 成本异常 / 全 fail。',
646
- '回答 "评测结论是否被坏用例污染"。',
647
- '',
648
- '选项:',
649
- ' --reports-dir <dir> 报告存储目录',
650
- ' --samples <path> 用例文件路径 (用于 near-duplicate 检测; 默认从 report.meta.request 读)',
651
- ' --top <n> 每类只显示前 N 个 (默认 10, 0=全部)',
652
- ' --duplicate-rouge <num> near-duplicate ROUGE-1 阈值 (默认 0.7)',
653
- ' --ambiguous-stddev <num> 歧义阈值, judge stddev (默认 1.0, 需要 --judge-repeat ≥ 2 数据)',
654
- ' --cost-k <num> 成本异常倍数 vs 中位数 (默认 3)',
655
- ' --latency-k <num> 耗时异常倍数 vs 中位数 (默认 3)',
656
- ' --flat <num> flat_scores 分差阈值 (默认 0.5)',
657
- '',
658
- ].join('\n'),
659
- en: [
660
- '',
661
- 'Usage: omk improve <reportId> [options]',
662
- ' omk improve plan <reportId> [options]',
663
- '',
664
- 'Diagnose quality issues in the sample set itself: low discrimination /',
665
- 'duplicates / ambiguity / cost anomalies / all-fail. Answers "is the verdict',
666
- 'tainted by bad samples?".',
667
- '',
668
- 'Options:',
669
- ' --reports-dir <dir> report store dir',
670
- ' --samples <path> sample file path (for near-duplicate detection; defaults to report.meta.request)',
671
- ' --top <n> top N per category (default 10, 0=all)',
672
- ' --duplicate-rouge <num> near-duplicate ROUGE-1 threshold (default 0.7)',
673
- ' --ambiguous-stddev <num> ambiguity threshold, judge stddev (default 1.0, requires --judge-repeat ≥ 2)',
674
- ' --cost-k <num> cost-outlier multiplier vs median (default 3)',
675
- ' --latency-k <num> latency-outlier multiplier vs median (default 3)',
676
- ' --flat <num> flat_scores spread threshold (default 0.5)',
677
- '',
678
- ].join('\n'),
679
- },
680
- 'cli.help.improve_failures': {
681
- zh: [
682
- '',
683
- '用法: omk improve failures <reportId> [options]',
684
- '',
685
- '把已有 report 的失败用例喂给一次 LLM 调用, 自动聚类并给出修复建议。',
686
- '失败定义: compositeScore < threshold 或 ok=false。',
687
- '',
688
- '选项:',
689
- ' --reports-dir <dir> 报告存储目录',
690
- ' --judge-models <executor:model> 评委 (默认: 沿用 report.meta.judgeModels[0]; failures 仅支持单评委)',
691
- ' --max-clusters <n> 最多聚成几类 (默认 5)',
692
- ' --threshold <num> compositeScore < threshold 算失败 (默认 3)',
693
- ' --max-feed <n> 最多喂给 LLM 多少条 (默认 50, 超出取最差)',
694
- '',
695
- ].join('\n'),
696
- en: [
697
- '',
698
- 'Usage: omk improve failures <reportId> [options]',
699
- '',
700
- 'Feed failing samples from an existing report to a single LLM call, auto-cluster',
701
- 'them, and produce per-cluster fix suggestions.',
702
- 'Failure definition: compositeScore < threshold or ok=false.',
703
- '',
704
- 'Options:',
705
- ' --reports-dir <dir> report store dir',
706
- ' --judge-models <executor:model> Judge (default: from report.meta.judgeModels[0]; failures is single-judge only)',
707
- ' --max-clusters <n> max number of clusters (default 5)',
708
- ' --threshold <num> compositeScore < threshold counts as failure (default 3)',
709
- ' --max-feed <n> max samples to feed the LLM (default 50, takes the worst)',
710
- '',
711
- ].join('\n'),
712
- },
713
- 'cli.help.improve_samples': {
601
+ 'cli.help.sample': {
714
602
  zh: `
715
- omk improve samples — 生成或补齐 eval-samples 评测用例
603
+ omk sample — 生成或补齐 eval-samples 评测用例
716
604
 
717
605
  用法:
718
- omk improve samples <skill-path> [options]
719
- omk improve samples --batch [--skill-dir <dir>] [options]
606
+ omk sample <skill-path> [options]
607
+ omk sample --batch [--skill-dir <dir>] [options]
720
608
 
721
609
  选项:
722
610
  --count <n> 生成用例数量(默认:5)
@@ -725,161 +613,17 @@ omk improve samples — 生成或补齐 eval-samples 评测用例
725
613
  --skill-dir <path> skill 目录(batch 使用,默认:skills)
726
614
  `,
727
615
  en: `
728
- omk improve samples — generate or fill eval-samples test cases
616
+ omk sample — generate or fill eval-samples test cases
729
617
 
730
618
  Usage:
731
- omk improve samples <skill-path> [options]
732
- omk improve samples --batch [--skill-dir <dir>] [options]
619
+ omk sample <skill-path> [options]
620
+ omk sample --batch [--skill-dir <dir>] [options]
733
621
 
734
622
  Options:
735
623
  --count <n> Number of test cases to generate (default: 5)
736
624
  --model <name> Generation model (default: sonnet)
737
625
  --batch Generate for skills that are missing eval-samples
738
626
  --skill-dir <path> Skill directory for batch mode (default: skills)
739
- `,
740
- },
741
- 'cli.help.improve_skill': {
742
- zh: `
743
- omk improve skill — 基于评测循环迭代改进 skill
744
-
745
- 用法:
746
- omk improve skill <skill-path> [options]
747
-
748
- 选项:
749
- --rounds <n> 迭代轮数(默认:3)
750
- --target <score> 目标分数
751
- --model <name> 改进模型
752
- --judge-models <executor:model> 单评委配置
753
- `,
754
- en: `
755
- omk improve skill — improve a skill through evaluation loops
756
-
757
- Usage:
758
- omk improve skill <skill-path> [options]
759
-
760
- Options:
761
- --rounds <n> Iteration rounds (default: 3)
762
- --target <score> Target score
763
- --model <name> Improvement model
764
- --judge-models <executor:model> Single judge config
765
- `,
766
- },
767
- 'cli.help.export': {
768
- zh: `
769
- omk export — 导出可贴到 PR / CI / audit 的证据包
770
-
771
- 用法:
772
- omk export <report-id> [options]
773
- omk export diff <report-id> [report-id] [options]
774
- omk export verdict <report-id> [options]
775
- omk export saturation <report-id> [options]
776
-
777
- 选项:
778
- --format <format> html / markdown / github-summary(默认:html)
779
- --out <path> 输出文件;markdown / github-summary 未指定时输出到 stdout
780
- --reports-dir <path> 报告目录(默认:~/.oh-my-knowledge/reports)
781
-
782
- 示例:
783
- omk export v1-vs-v2-20260505-1200 --format github-summary
784
- omk export v1-vs-v2-20260505-1200 --format markdown --out report.md
785
- omk export diff v1-vs-v2-20260505-1200 --regressions-only
786
- omk export verdict v1-vs-v2-20260505-1200
787
- `,
788
- en: `
789
- omk export — export evidence packs for PR / CI / audit
790
-
791
- Usage:
792
- omk export <report-id> [options]
793
- omk export diff <report-id> [report-id] [options]
794
- omk export verdict <report-id> [options]
795
- omk export saturation <report-id> [options]
796
-
797
- Options:
798
- --format <format> html / markdown / github-summary (default: html)
799
- --out <path> Output file; markdown / github-summary print to stdout by default
800
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports)
801
-
802
- Examples:
803
- omk export v1-vs-v2-20260505-1200 --format github-summary
804
- omk export v1-vs-v2-20260505-1200 --format markdown --out report.md
805
- omk export diff v1-vs-v2-20260505-1200 --regressions-only
806
- omk export verdict v1-vs-v2-20260505-1200
807
- `,
808
- },
809
- 'cli.help.export_diff': {
810
- zh: `
811
- omk export diff — 导出样本级或报告级差异
812
-
813
- 用法:
814
- omk export diff <report-id> [--variant <name>] [--regressions-only] [--top <n>]
815
- omk export diff <report-id-a> <report-id-b>
816
-
817
- 选项:
818
- --reports-dir <path> 报告目录(默认:~/.oh-my-knowledge/reports)
819
- --variant <name> 样本级 diff 的实验组 variant
820
- --regressions-only 只显示回退用例
821
- --top <n> 最多显示 N 条
822
- `,
823
- en: `
824
- omk export diff — export sample-level or cross-report differences
825
-
826
- Usage:
827
- omk export diff <report-id> [--variant <name>] [--regressions-only] [--top <n>]
828
- omk export diff <report-id-a> <report-id-b>
829
-
830
- Options:
831
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports)
832
- --variant <name> Treatment variant for sample-level diff
833
- --regressions-only Show regressions only
834
- --top <n> Show at most N rows
835
- `,
836
- },
837
- 'cli.help.export_verdict': {
838
- zh: `
839
- omk export verdict — 输出已有报告的一行 ship/no-ship verdict
840
-
841
- 用法:
842
- omk export verdict <report-id> [options]
843
-
844
- 选项:
845
- --reports-dir <path> 报告目录(默认:~/.oh-my-knowledge/reports)
846
- --threshold <number> 三层 gate 阈值
847
- --trivial-diff <number> 实际可忽略 diff
848
- --verbose 输出完整 verdict 解释
849
- `,
850
- en: `
851
- omk export verdict — print a one-line ship/no-ship verdict for an existing report
852
-
853
- Usage:
854
- omk export verdict <report-id> [options]
855
-
856
- Options:
857
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports)
858
- --threshold <number> Three-layer gate threshold
859
- --trivial-diff <number> Practically negligible diff
860
- --verbose Print the full verdict explanation
861
- `,
862
- },
863
- 'cli.help.export_saturation': {
864
- zh: `
865
- omk export saturation — 输出重复评测的饱和度证据
866
-
867
- 用法:
868
- omk export saturation <report-id> [--variant <name>]
869
-
870
- 选项:
871
- --reports-dir <path> 报告目录(默认:~/.oh-my-knowledge/reports)
872
- --variant <name> 只输出指定 variant
873
- `,
874
- en: `
875
- omk export saturation — print saturation evidence from repeated evaluations
876
-
877
- Usage:
878
- omk export saturation <report-id> [--variant <name>]
879
-
880
- Options:
881
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports)
882
- --variant <name> Print only one variant
883
627
  `,
884
628
  },
885
629
  'cli.help.studio': {
@@ -920,27 +664,6 @@ Examples:
920
664
  omk studio --no-open
921
665
  `,
922
666
  },
923
- // sample design coverage block strings
924
- 'cli.diagnose.coverage_header': {
925
- zh: '用例设计覆盖度 (Sample design coverage):',
926
- en: 'Sample design coverage:',
927
- },
928
- 'cli.diagnose.coverage_unspecified': {
929
- zh: '(未声明)',
930
- en: '(unspecified)',
931
- },
932
- 'cli.diagnose.coverage_chars': {
933
- zh: '字符',
934
- en: 'chars',
935
- },
936
- 'cli.diagnose.coverage_hint_empty': {
937
- zh: 'ℹ 该用例集未声明任何 capability / difficulty / construct / provenance 元数据。详见 docs/sample-design-spec.md',
938
- en: 'ℹ No samples in this set declare capability / difficulty / construct / provenance metadata. See docs/sample-design-spec.md',
939
- },
940
- 'cli.diagnose.coverage_declared': {
941
- zh: '声明',
942
- en: 'declared',
943
- },
944
667
  // ============ omk doctor 健康检查 ============
945
668
  'cli.doctor.rule.skill_readable': {
946
669
  zh: 'skill 文件可读',
@@ -958,6 +681,67 @@ Examples:
958
681
  zh: '用例 ↔ skill 输入约定',
959
682
  en: 'samples ↔ skill contract',
960
683
  },
684
+ 'cli.doctor.rule.skill_health_check': {
685
+ zh: '健康度体检',
686
+ en: 'Health check',
687
+ },
688
+ // ============ skill_health composer (CLI default; --static-only disables it) ============
689
+ 'cli.doctor.health.skipped': {
690
+ zh: '健康度体检已跳过(runHealthCheck=false)',
691
+ en: 'health check skipped (runHealthCheck=false)',
692
+ },
693
+ 'cli.doctor.health.no_dimensions': {
694
+ zh: '没有注册任何健康度维度,跳过',
695
+ en: 'no health dimensions registered, skipped',
696
+ },
697
+ 'cli.doctor.health.fail.executor': {
698
+ zh: 'LLM 调用失败: {error}',
699
+ en: 'LLM call failed: {error}',
700
+ },
701
+ 'cli.doctor.health.fail.parse': {
702
+ zh: 'LLM 输出解析失败: {error}',
703
+ en: 'failed to parse LLM output: {error}',
704
+ },
705
+ 'cli.doctor.health.fail.empty_output': {
706
+ zh: 'LLM 返回了空输出',
707
+ en: 'LLM returned empty output',
708
+ },
709
+ 'cli.doctor.health.hint.executor': {
710
+ zh: '检查 executor 配置(--executor / --model)与网络连通,或调大 --timeout',
711
+ en: 'Verify executor config (--executor / --model) and connectivity, or raise --timeout',
712
+ },
713
+ 'cli.doctor.health.hint.parse': {
714
+ zh: 'LLM 没返回合法 JSON;原文存在 detail.rawOutput 截断片段,可重跑或换 model',
715
+ en: 'LLM did not return valid JSON; raw snippet stored in detail.rawOutput. Re-run or switch model',
716
+ },
717
+ 'cli.doctor.health.dim.message': {
718
+ zh: '{level}: 错误 {err}/警告 {warn}/建议 {sug}',
719
+ en: '{level}: error {err}/warn {warn}/suggest {sug}',
720
+ },
721
+ 'cli.doctor.health.dim.missing': {
722
+ zh: 'LLM 未输出此维度({dim}),已置不适用',
723
+ en: 'LLM omitted dimension ({dim}); treated as N/A',
724
+ },
725
+ 'cli.doctor.health.summary.label': {
726
+ zh: '健康度总览',
727
+ en: 'Health summary',
728
+ },
729
+ 'cli.doctor.health.summary.message': {
730
+ zh: '{overall} | 维度: 健康 {h}/亚健康 {sh}/不健康 {bad}/不适用 {na} | finding: 错误 {err}/警告 {warn}/建议 {sug}',
731
+ en: '{overall} | dims: healthy {h}/sub {sh}/unhealthy {bad}/n-a {na} | findings: err {err}/warn {warn}/sug {sug}',
732
+ },
733
+ 'cli.doctor.health.summary.no_top': {
734
+ zh: '完整详情见 --json 输出或 --html 报告',
735
+ en: 'Full detail in --json output or --html report',
736
+ },
737
+ // 7 内置维度 labelKey (id-based)
738
+ 'cli.doctor.health.dim.trigger-boundary': { zh: '触发与边界', en: 'Trigger & boundary' },
739
+ 'cli.doctor.health.dim.doc-clarity': { zh: '文档清晰', en: 'Documentation clarity' },
740
+ 'cli.doctor.health.dim.instr-precision': { zh: '指令精确性', en: 'Instruction precision' },
741
+ 'cli.doctor.health.dim.dependency': { zh: '依赖检查', en: 'Dependency check' },
742
+ 'cli.doctor.health.dim.tool-conventions': { zh: '工具规范', en: 'Tool conventions' },
743
+ 'cli.doctor.health.dim.security': { zh: '安全与合规', en: 'Security & compliance' },
744
+ 'cli.doctor.health.dim.examples': { zh: '示例完备', en: 'Example completeness' },
961
745
  // pass
962
746
  'cli.doctor.skill_readable.pass': {
963
747
  zh: 'skill 内容长度 {length} 字符',
@@ -1072,10 +856,10 @@ Examples:
1072
856
  // ============ omk doctor CLI level ============
1073
857
  'cli.help.doctor_usage': {
1074
858
  zh: `
1075
- oh-my-knowledge — omk doctor 健康检查
859
+ oh-my-knowledge — omk doctor 健康度体检 (LLM-judge)
1076
860
 
1077
861
  用法:
1078
- omk doctor [path] 在 path 上跑评测前置健康检查
862
+ omk doctor [path] 在 path 上跑深度健康度体检
1079
863
  omk doctor 在当前目录(或 ./skills)批量跑
1080
864
 
1081
865
  参数:
@@ -1083,62 +867,70 @@ oh-my-knowledge — omk doctor 健康检查
1083
867
 
1084
868
  选项:
1085
869
  --json 把 DoctorReport 打到 stdout(CI 消费用)
1086
- --gate 静默模式: 通过 exit 0 / 不通过 exit 1, 仅 stderr 出问题摘要
1087
- --executor <name> executor 名(仅向后兼容, doctor 不直接打 LLM)
1088
- --model <name> model 名(同上)
870
+ --gate 静默模式: fatal 问题 exit 1; warnings_only 仍 exit 0, 仅 stderr 出摘要
871
+ --executor <name> LLM executor (默认 claude, 可换 anthropic-api/codex 等)
872
+ --model <name> 模型 (默认 sonnet)
1089
873
  --samples <path> 显式指定评测用例文件
1090
- --timeout <seconds> rule 执行超时(默认 8)
874
+ --timeout <seconds> 单次 LLM 会话超时 (默认 600)
875
+ --html <path> 产出可视化 HTML 报告到 <path> (可与 --json 同时用)
876
+ --static-only 离线模式: 只跑静态检查 (skill 可读性 / 元数据 / 依赖 / samples 契约), 不调 LLM
1091
877
  --lang <zh|en> 切换输出语言
1092
878
 
1093
879
  示例:
1094
- omk doctor examples/code-review/skills/v1.md
1095
- omk doctor examples/code-review/skills --json | jq .outcome # passed | warnings_only | failed
1096
- omk doctor --gate; echo $?
1097
-
1098
- doctor 检查项(纯静态 / 零 LLM 调用):
1099
- - skill 文件可读 + 内容有最小长度
1100
- - skill 元数据合法 (front-matter 若有)
1101
- - 前置依赖完整 (引用的 CLI 工具 / 文件 / 环境变量 / preflight 命令)
1102
- - 用例 ↔ skill 输入约定 (warn 级, 仅传 samples 时跑)
1103
-
1104
- executor / judge 连通性由 evaluation preflight 负责, 不在 doctor 范围内。
1105
- omk eval 内置 doctor 强制门禁, 不可 skip — 静态检查
1106
- 零成本无理由跳过。LLM 连通性可用 --skip-connectivity 跳过 (--resume 时自动)。
880
+ omk doctor my-skill --html /tmp/report.html # 深度体检 + HTML 报告 (默认)
881
+ omk doctor examples/code-review/skills --json > r.json # JSON 给 CI / 外部工具消费
882
+ omk doctor --gate; echo $? # CI 模式: fatal 问题 exit 1, 警告不阻断
883
+ omk doctor --static-only # 无 LLM 环境 (CI / 断网) 跑纯静态检查
884
+
885
+ doctor = LLM 健康度体检 (单次 LLM 会话):
886
+ - 7 个内置维度: 触发与边界 / 文档清晰 / 指令精确性 / 依赖检查 / 工具规范 / 安全与合规 / 示例完备
887
+ - 用户可扩展: 在自己代码里 registerHealthDimension(spec) 加自定义维度,
888
+ 会自动加入同一次 LLM 调用的 prompt + 报告 (顺序 = 注册顺序)
889
+ - 每维度独立给 健康/亚健康/不健康/不适用 + findings + 改进建议
890
+ - HTML 报告: 维度按 fail→warn→pass→skipped 排, 错误 finding 排前面
891
+
892
+ 注: omk eval 内部仍跑静态 skill-readability/metadata/dependency
893
+ gate 保护评测质量, 不走 omk doctor 这条 LLM 路径 (角色分离: doctor=审计, eval=评测)。
894
+ LLM 连通性可用 omk eval --skip-connectivity 跳过 (--resume 时自动)。
1107
895
  `.trim() + '\n',
1108
896
  en: `
1109
- oh-my-knowledge — omk doctor health check
897
+ oh-my-knowledge — omk doctor health audit (LLM-judge)
1110
898
 
1111
899
  Usage:
1112
- omk doctor [path] Run pre-evaluation health check on path
1113
- omk doctor Batch check current dir (or ./skills)
900
+ omk doctor [path] Run deep LLM-based health audit on path
901
+ omk doctor Batch audit current dir (or ./skills)
1114
902
 
1115
903
  Arguments:
1116
- path A .md file, directory, or omit (= cwd). Directory mode batches all skills.
904
+ path A .md file, directory, or omit (= cwd). Directory batches all skills.
1117
905
 
1118
906
  Options:
1119
907
  --json Print DoctorReport JSON to stdout (CI-friendly)
1120
- --gate Silent mode: exit 0 if pass, exit 1 if fail; brief stderr summary only
1121
- --executor <name> executor name (kept for compat; doctor does not call LLM)
1122
- --model <name> model name (same)
908
+ --gate Silent mode: exit 1 only on fatal failure; warnings_only exits 0
909
+ --executor <name> LLM executor (default 'claude'; switchable to anthropic-api/codex etc)
910
+ --model <name> model name (default 'sonnet')
1123
911
  --samples <path> Explicit eval samples file
1124
- --timeout <seconds> per-rule timeout (default 8)
912
+ --timeout <seconds> LLM session timeout (default 600)
913
+ --html <path> Also write a visual HTML report to <path> (combines with --json)
914
+ --static-only Offline mode: run only static checks (readability / metadata / deps / samples contract); no LLM call
1125
915
  --lang <zh|en> Output language
1126
916
 
1127
917
  Examples:
1128
- omk doctor examples/code-review/skills/v1.md
1129
- omk doctor examples/code-review/skills --json | jq .outcome # passed | warnings_only | failed
1130
- omk doctor --gate; echo $?
1131
-
1132
- Checks (pure static / zero LLM calls):
1133
- - skill file readable + minimum content length
1134
- - skill metadata valid (front-matter if present)
1135
- - dependencies present (referenced CLI tools / files / env vars / preflight commands)
1136
- - samples ↔ skill contract (warn-level, only when samples provided)
1137
-
1138
- executor / judge connectivity is handled by evaluation preflight, not doctor.
1139
- omk eval runs doctor as mandatory; no skip flag — static
1140
- checks cost nothing to run. LLM connectivity can be skipped with --skip-connectivity
1141
- (auto-skipped on --resume).
918
+ omk doctor my-skill --html /tmp/report.html # deep audit + HTML report (default)
919
+ omk doctor examples/code-review/skills --json > r.json # JSON for CI / external tools
920
+ omk doctor --gate; echo $? # CI mode: fatal failures exit 1; warnings do not block
921
+ omk doctor --static-only # offline (CI / no LLM) static checks only
922
+
923
+ doctor = LLM health audit (single LLM session):
924
+ - 7 builtin dimensions: trigger & boundary / doc clarity / instruction precision /
925
+ dependency / tool conventions / security & compliance / example completeness
926
+ - User-extensible: call registerHealthDimension(spec) in your code to add custom
927
+ dimensions; they join the same LLM call's prompt + report (order = registration order)
928
+ - Each dim graded healthy / sub-healthy / unhealthy / N-A + findings + suggestions
929
+ - HTML report: dims sorted fail→warn→pass→skipped; errors first within each dim
930
+
931
+ Note: omk eval still runs static skill-readability/metadata/dependency gates
932
+ internally (separate from this doctor command). Roles: doctor=audit, eval=evaluate.
933
+ LLM connectivity for omk eval can be skipped with --skip-connectivity (auto on --resume).
1142
934
  `.trim() + '\n',
1143
935
  },
1144
936
  'cli.doctor.no_skill_found': {