oh-my-knowledge 0.0.1 → 0.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (216) hide show
  1. package/LICENSE +21 -0
  2. package/README.md +600 -322
  3. package/README.zh.md +724 -0
  4. package/dist/src/analysis/coverage-analyzer.d.ts +57 -0
  5. package/dist/src/analysis/coverage-analyzer.d.ts.map +1 -0
  6. package/dist/src/analysis/coverage-analyzer.js +262 -0
  7. package/dist/src/analysis/coverage-analyzer.js.map +1 -0
  8. package/dist/src/analysis/gap-analyzer.d.ts +114 -0
  9. package/dist/src/analysis/gap-analyzer.d.ts.map +1 -0
  10. package/dist/src/analysis/gap-analyzer.js +439 -0
  11. package/dist/src/analysis/gap-analyzer.js.map +1 -0
  12. package/dist/src/analysis/hedging-classifier.d.ts +34 -0
  13. package/dist/src/analysis/hedging-classifier.d.ts.map +1 -0
  14. package/dist/src/analysis/hedging-classifier.js +144 -0
  15. package/dist/src/analysis/hedging-classifier.js.map +1 -0
  16. package/dist/src/analysis/report-diagnostics.d.ts +9 -0
  17. package/dist/src/analysis/report-diagnostics.d.ts.map +1 -0
  18. package/dist/src/analysis/report-diagnostics.js +631 -0
  19. package/dist/src/analysis/report-diagnostics.js.map +1 -0
  20. package/dist/src/authoring/evolver.d.ts +64 -0
  21. package/dist/src/authoring/evolver.d.ts.map +1 -0
  22. package/dist/src/authoring/evolver.js +275 -0
  23. package/dist/src/authoring/evolver.js.map +1 -0
  24. package/dist/src/authoring/generator.d.ts +13 -0
  25. package/dist/src/authoring/generator.d.ts.map +1 -0
  26. package/dist/src/authoring/generator.js +60 -0
  27. package/dist/src/authoring/generator.js.map +1 -0
  28. package/dist/src/cli.d.ts +3 -0
  29. package/dist/src/cli.d.ts.map +1 -0
  30. package/dist/src/cli.js +1012 -0
  31. package/dist/src/cli.js.map +1 -0
  32. package/dist/src/eval-core/cache.d.ts +11 -0
  33. package/dist/src/eval-core/cache.d.ts.map +1 -0
  34. package/dist/src/eval-core/cache.js +53 -0
  35. package/dist/src/eval-core/cache.js.map +1 -0
  36. package/dist/src/eval-core/ci-gates.d.ts +17 -0
  37. package/dist/src/eval-core/ci-gates.d.ts.map +1 -0
  38. package/dist/src/eval-core/ci-gates.js +42 -0
  39. package/dist/src/eval-core/ci-gates.js.map +1 -0
  40. package/dist/src/eval-core/dependency-checker.d.ts +37 -0
  41. package/dist/src/eval-core/dependency-checker.d.ts.map +1 -0
  42. package/dist/src/eval-core/dependency-checker.js +280 -0
  43. package/dist/src/eval-core/dependency-checker.js.map +1 -0
  44. package/dist/src/eval-core/evaluation-execution.d.ts +26 -0
  45. package/dist/src/eval-core/evaluation-execution.d.ts.map +1 -0
  46. package/dist/src/eval-core/evaluation-execution.js +196 -0
  47. package/dist/src/eval-core/evaluation-execution.js.map +1 -0
  48. package/dist/src/eval-core/evaluation-job.d.ts +50 -0
  49. package/dist/src/eval-core/evaluation-job.d.ts.map +1 -0
  50. package/dist/src/eval-core/evaluation-job.js +114 -0
  51. package/dist/src/eval-core/evaluation-job.js.map +1 -0
  52. package/dist/src/eval-core/evaluation-reporting.d.ts +28 -0
  53. package/dist/src/eval-core/evaluation-reporting.d.ts.map +1 -0
  54. package/dist/src/eval-core/evaluation-reporting.js +123 -0
  55. package/dist/src/eval-core/evaluation-reporting.js.map +1 -0
  56. package/dist/src/eval-core/execution-strategy.d.ts +11 -0
  57. package/dist/src/eval-core/execution-strategy.d.ts.map +1 -0
  58. package/dist/src/eval-core/execution-strategy.js +121 -0
  59. package/dist/src/eval-core/execution-strategy.js.map +1 -0
  60. package/dist/src/eval-core/fact-checker.d.ts +25 -0
  61. package/dist/src/eval-core/fact-checker.d.ts.map +1 -0
  62. package/dist/src/eval-core/fact-checker.js +66 -0
  63. package/dist/src/eval-core/fact-checker.js.map +1 -0
  64. package/dist/src/eval-core/schema.d.ts +30 -0
  65. package/dist/src/eval-core/schema.d.ts.map +1 -0
  66. package/dist/src/eval-core/schema.js +195 -0
  67. package/dist/src/eval-core/schema.js.map +1 -0
  68. package/dist/src/eval-core/statistics.d.ts +86 -0
  69. package/dist/src/eval-core/statistics.d.ts.map +1 -0
  70. package/dist/src/eval-core/statistics.js +210 -0
  71. package/dist/src/eval-core/statistics.js.map +1 -0
  72. package/dist/src/eval-core/task-planner.d.ts +4 -0
  73. package/dist/src/eval-core/task-planner.d.ts.map +1 -0
  74. package/dist/src/eval-core/task-planner.js +33 -0
  75. package/dist/src/eval-core/task-planner.js.map +1 -0
  76. package/dist/src/eval-workflows/each-evaluation-workflow.d.ts +137 -0
  77. package/dist/src/eval-workflows/each-evaluation-workflow.d.ts.map +1 -0
  78. package/dist/src/eval-workflows/each-evaluation-workflow.js +174 -0
  79. package/dist/src/eval-workflows/each-evaluation-workflow.js.map +1 -0
  80. package/dist/src/eval-workflows/evaluation-pipeline.d.ts +42 -0
  81. package/dist/src/eval-workflows/evaluation-pipeline.d.ts.map +1 -0
  82. package/dist/src/eval-workflows/evaluation-pipeline.js +218 -0
  83. package/dist/src/eval-workflows/evaluation-pipeline.js.map +1 -0
  84. package/dist/src/eval-workflows/evaluation-preparation.d.ts +64 -0
  85. package/dist/src/eval-workflows/evaluation-preparation.d.ts.map +1 -0
  86. package/dist/src/eval-workflows/evaluation-preparation.js +90 -0
  87. package/dist/src/eval-workflows/evaluation-preparation.js.map +1 -0
  88. package/dist/src/eval-workflows/run-evaluation.d.ts +111 -0
  89. package/dist/src/eval-workflows/run-evaluation.d.ts.map +1 -0
  90. package/dist/src/eval-workflows/run-evaluation.js +272 -0
  91. package/dist/src/eval-workflows/run-evaluation.js.map +1 -0
  92. package/dist/src/executors/anthropic-api.d.ts +3 -0
  93. package/dist/src/executors/anthropic-api.d.ts.map +1 -0
  94. package/dist/src/executors/anthropic-api.js +43 -0
  95. package/dist/src/executors/anthropic-api.js.map +1 -0
  96. package/dist/src/executors/claude-cli.d.ts +3 -0
  97. package/dist/src/executors/claude-cli.d.ts.map +1 -0
  98. package/dist/src/executors/claude-cli.js +103 -0
  99. package/dist/src/executors/claude-cli.js.map +1 -0
  100. package/dist/src/executors/claude-sdk-trace.d.ts +10 -0
  101. package/dist/src/executors/claude-sdk-trace.d.ts.map +1 -0
  102. package/dist/src/executors/claude-sdk-trace.js +100 -0
  103. package/dist/src/executors/claude-sdk-trace.js.map +1 -0
  104. package/dist/src/executors/claude-sdk.d.ts +3 -0
  105. package/dist/src/executors/claude-sdk.d.ts.map +1 -0
  106. package/dist/src/executors/claude-sdk.js +160 -0
  107. package/dist/src/executors/claude-sdk.js.map +1 -0
  108. package/dist/src/executors/gemini.d.ts +3 -0
  109. package/dist/src/executors/gemini.d.ts.map +1 -0
  110. package/dist/src/executors/gemini.js +78 -0
  111. package/dist/src/executors/gemini.js.map +1 -0
  112. package/dist/src/executors/index.d.ts +7 -0
  113. package/dist/src/executors/index.d.ts.map +1 -0
  114. package/dist/src/executors/index.js +22 -0
  115. package/dist/src/executors/index.js.map +1 -0
  116. package/dist/src/executors/openai-api.d.ts +3 -0
  117. package/dist/src/executors/openai-api.d.ts.map +1 -0
  118. package/dist/src/executors/openai-api.js +40 -0
  119. package/dist/src/executors/openai-api.js.map +1 -0
  120. package/dist/src/executors/openai-cli.d.ts +3 -0
  121. package/dist/src/executors/openai-cli.d.ts.map +1 -0
  122. package/dist/src/executors/openai-cli.js +60 -0
  123. package/dist/src/executors/openai-cli.js.map +1 -0
  124. package/dist/src/executors/script.d.ts +3 -0
  125. package/dist/src/executors/script.d.ts.map +1 -0
  126. package/dist/src/executors/script.js +63 -0
  127. package/dist/src/executors/script.js.map +1 -0
  128. package/dist/src/executors/shared.d.ts +117 -0
  129. package/dist/src/executors/shared.d.ts.map +1 -0
  130. package/dist/src/executors/shared.js +49 -0
  131. package/dist/src/executors/shared.js.map +1 -0
  132. package/dist/src/grading/assertions.d.ts +18 -0
  133. package/dist/src/grading/assertions.d.ts.map +1 -0
  134. package/dist/src/grading/assertions.js +239 -0
  135. package/dist/src/grading/assertions.js.map +1 -0
  136. package/dist/src/grading/index.d.ts +26 -0
  137. package/dist/src/grading/index.d.ts.map +1 -0
  138. package/dist/src/grading/index.js +98 -0
  139. package/dist/src/grading/index.js.map +1 -0
  140. package/dist/src/grading/judge.d.ts +13 -0
  141. package/dist/src/grading/judge.d.ts.map +1 -0
  142. package/dist/src/grading/judge.js +98 -0
  143. package/dist/src/grading/judge.js.map +1 -0
  144. package/dist/src/grading/layered-scores.d.ts +13 -0
  145. package/dist/src/grading/layered-scores.d.ts.map +1 -0
  146. package/dist/src/grading/layered-scores.js +62 -0
  147. package/dist/src/grading/layered-scores.js.map +1 -0
  148. package/dist/src/inputs/eval-config.d.ts +13 -0
  149. package/dist/src/inputs/eval-config.d.ts.map +1 -0
  150. package/dist/src/inputs/eval-config.js +136 -0
  151. package/dist/src/inputs/eval-config.js.map +1 -0
  152. package/dist/src/inputs/load-samples.d.ts +16 -0
  153. package/dist/src/inputs/load-samples.d.ts.map +1 -0
  154. package/dist/src/inputs/load-samples.js +51 -0
  155. package/dist/src/inputs/load-samples.js.map +1 -0
  156. package/dist/src/inputs/mcp-resolver.d.ts +50 -0
  157. package/dist/src/inputs/mcp-resolver.d.ts.map +1 -0
  158. package/dist/src/inputs/mcp-resolver.js +307 -0
  159. package/dist/src/inputs/mcp-resolver.js.map +1 -0
  160. package/dist/src/inputs/skill-loader.d.ts +18 -0
  161. package/dist/src/inputs/skill-loader.d.ts.map +1 -0
  162. package/dist/src/inputs/skill-loader.js +222 -0
  163. package/dist/src/inputs/skill-loader.js.map +1 -0
  164. package/dist/src/inputs/url-fetcher.d.ts +19 -0
  165. package/dist/src/inputs/url-fetcher.d.ts.map +1 -0
  166. package/dist/src/inputs/url-fetcher.js +241 -0
  167. package/dist/src/inputs/url-fetcher.js.map +1 -0
  168. package/dist/src/observability/skill-health-analyzer.d.ts +83 -0
  169. package/dist/src/observability/skill-health-analyzer.d.ts.map +1 -0
  170. package/dist/src/observability/skill-health-analyzer.js +238 -0
  171. package/dist/src/observability/skill-health-analyzer.js.map +1 -0
  172. package/dist/src/observability/trace-adapter.d.ts +75 -0
  173. package/dist/src/observability/trace-adapter.d.ts.map +1 -0
  174. package/dist/src/observability/trace-adapter.js +367 -0
  175. package/dist/src/observability/trace-adapter.js.map +1 -0
  176. package/dist/src/renderer/html-renderer.d.ts +9 -0
  177. package/dist/src/renderer/html-renderer.d.ts.map +1 -0
  178. package/dist/src/renderer/html-renderer.js +360 -0
  179. package/dist/src/renderer/html-renderer.js.map +1 -0
  180. package/dist/src/renderer/layout.d.ts +13 -0
  181. package/dist/src/renderer/layout.d.ts.map +1 -0
  182. package/dist/src/renderer/layout.js +582 -0
  183. package/dist/src/renderer/layout.js.map +1 -0
  184. package/dist/src/renderer/skill-health-renderer.d.ts +20 -0
  185. package/dist/src/renderer/skill-health-renderer.d.ts.map +1 -0
  186. package/dist/src/renderer/skill-health-renderer.js +253 -0
  187. package/dist/src/renderer/skill-health-renderer.js.map +1 -0
  188. package/dist/src/renderer/summary.d.ts +21 -0
  189. package/dist/src/renderer/summary.d.ts.map +1 -0
  190. package/dist/src/renderer/summary.js +1072 -0
  191. package/dist/src/renderer/summary.js.map +1 -0
  192. package/dist/src/renderer/table.d.ts +3 -0
  193. package/dist/src/renderer/table.d.ts.map +1 -0
  194. package/dist/src/renderer/table.js +134 -0
  195. package/dist/src/renderer/table.js.map +1 -0
  196. package/dist/src/renderer/trends.d.ts +6 -0
  197. package/dist/src/renderer/trends.d.ts.map +1 -0
  198. package/dist/src/renderer/trends.js +127 -0
  199. package/dist/src/renderer/trends.js.map +1 -0
  200. package/dist/src/server/job-store.d.ts +4 -0
  201. package/dist/src/server/job-store.d.ts.map +1 -0
  202. package/dist/src/server/job-store.js +68 -0
  203. package/dist/src/server/job-store.js.map +1 -0
  204. package/dist/src/server/report-server.d.ts +17 -0
  205. package/dist/src/server/report-server.d.ts.map +1 -0
  206. package/dist/src/server/report-server.js +591 -0
  207. package/dist/src/server/report-server.js.map +1 -0
  208. package/dist/src/server/report-store.d.ts +44 -0
  209. package/dist/src/server/report-store.d.ts.map +1 -0
  210. package/dist/src/server/report-store.js +190 -0
  211. package/dist/src/server/report-store.js.map +1 -0
  212. package/dist/src/types.d.ts +559 -0
  213. package/dist/src/types.d.ts.map +1 -0
  214. package/dist/src/types.js +2 -0
  215. package/dist/src/types.js.map +1 -0
  216. package/package.json +53 -5
package/README.md CHANGED
@@ -1,448 +1,726 @@
1
1
  # oh-my-knowledge
2
2
 
3
- Knowledge artifact evaluation toolkit — benchmark your skills with objective data.
3
+ [![npm version](https://img.shields.io/npm/v/oh-my-knowledge.svg)](https://www.npmjs.com/package/oh-my-knowledge)
4
+ [![CI](https://github.com/lizhiyao/oh-my-knowledge/actions/workflows/ci.yml/badge.svg)](https://github.com/lizhiyao/oh-my-knowledge/actions/workflows/ci.yml)
5
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](./LICENSE)
6
+ [![Node.js Version](https://img.shields.io/node/v/oh-my-knowledge.svg)](https://nodejs.org)
4
7
 
5
- English | [中文](./README.zh-CN.md)
8
+ **English** | [简体中文](./README.zh.md)
6
9
 
7
- **Fixed model, variable knowledge artifact, data speaks.**
10
+ Knowledge-artifact evaluation toolkit — measure your artifact's quality with objective data.
8
11
 
9
- ## Why
12
+ **Fix the model, vary the knowledge artifact, let the data speak.**
10
13
 
11
- Teams building AI skills (system prompts, knowledge packages, rule sets) need objective data to prove v2 is better than v1. `oh-my-knowledge` runs controlled experiments: same model, same test cases, only the knowledge artifact changes.
14
+ ## Why this tool
12
15
 
13
- ## Quick Start
16
+ Teams doing knowledge engineering produce lots of knowledge artifacts (skills today, but also prompts, agents, workflows…). When someone asks "why is v2 better than v1", you need objective data instead of gut feeling. `oh-my-knowledge` solves this with controlled experiments: **same model, same test samples, only the knowledge artifact changes.**
17
+
18
+ ## Key features
19
+
20
+ - **Controlled-variable offline bench** — fix the model and samples, vary only the artifact; works with Claude Code skills, CLAUDE.md prompts, RAG knowledge bases, or any markdown-based instruction
21
+ - **Six-dimension scoring** — separate signals for Fact / Behavior / LLM-judge / Cost / Efficiency / Stability, so a regression in one axis isn't hidden by gains in another
22
+ - **Production session observability** — parse Claude Code session JSONL traces, measure per-skill failure rate, latency, token cost, and knowledge-gap signals on real user sessions
23
+ - **Knowledge-gap detection** — severity-weighted signals (explicit markers / failed searches / hedging language / repeated failures) quantify risk exposure instead of claiming completeness
24
+ - **Pre-merge CI gate** — `omk bench ci` enforces three-layer all-pass (fact + behavior + llm-judge) semantics, catching single-layer regressions a composite score would hide
25
+
26
+ ## Quick start
14
27
 
15
28
  ```bash
16
- # Install globally
17
- npm i -g oh-my-knowledge
29
+ # install
30
+ npm i oh-my-knowledge -g
18
31
 
19
- # Scaffold a new eval project
32
+ # scaffold an eval project
20
33
  omk bench init my-eval
21
34
  cd my-eval
22
35
 
23
- # Preview the evaluation plan
36
+ # drop the artifacts you want to compare into skills/
37
+ # option 1: plain .md files (skills/v1.md, skills/v2.md)
38
+ # option 2: full artifact dirs (skills/my-skill-v1/SKILL.md, ...)
39
+ # a single artifact also works — baseline is auto-added as control
40
+
41
+ # preview the plan
24
42
  omk bench run --dry-run
25
43
 
26
- # Run the evaluation
27
- omk bench run --variants v1,v2
28
-
29
- # View the report
30
- omk bench report
31
- # Open http://127.0.0.1:7799
32
- ```
33
-
34
- ## How It Works
35
-
36
- ```
37
- eval-samples.json skills/v1.md skills/v2.md
38
- │ │ │
39
- └────────┬───────────┘ │
40
- │ │
41
- ┌──────▼──────┐ ┌──────▼──────┐
42
- │ sample + │ │ sample + │
43
- │ skill v1 │ │ skill v2 │
44
- └──────┬──────┘ └──────┬──────┘
45
- │ │
46
- ┌──────▼──────┐ ┌──────▼──────┐
47
- │ Executor │ │ Executor │
48
- │ claude │ │ claude │
49
- │ openai │ │ openai │
50
- │ gemini │ │ gemini │
51
- └──────┬──────┘ └──────┬──────┘
52
- │ │
53
- ┌──────▼──────────────────────▼──────┐
54
- │ Grading │
55
- │ ┌─────────────┐ ┌──────────────┐ │
56
- │ │ Assertions │ │ LLM Judge │ │
57
- │ │ (18 types) │ │ (rubric or │ │
58
- │ │ │ │ dimensions) │ │
59
- │ └─────────────┘ └──────────────┘ │
60
- └──────────────────┬─────────────────┘
61
- │
62
- ┌─────────▼─────────┐
63
- │ Report + Analysis │
64
- │ (JSON/HTML) │
65
- └───────────────────┘
66
- ```
67
-
68
- ## Eval Sample Format
69
-
70
- Supports both JSON and YAML (`eval-samples.json`, `eval-samples.yaml`, `eval-samples.yml`).
71
-
72
- The file contains an array of sample objects. Each sample represents one test case for evaluating a skill.
44
+ # run the evaluation (auto-discovers everything under skills/)
45
+ omk bench run
46
+ ```
47
+
48
+ ## Use inside Claude Code
49
+
50
+ After installing omk, talk to it in natural language from Claude Code:
51
+
52
+ ```
53
+ /omk eval # evaluate the artifact(s) in the current project
54
+ /omk evolve # auto-iterate to improve an artifact
55
+ /omk gen-samples # generate test cases
56
+ ```
57
+
58
+ You can also just say "compare v1 vs v2 for me" or "improve this artifact" — omk picks the right command.
59
+
60
+ ## Features
61
+
62
+ | Feature | What it does |
63
+ |---|---|
64
+ | **18 assertion types** | substring, regex, JSON Schema, semantic similarity, custom JS function, and more |
65
+ | **Six-dim evaluation** | Fact / Behavior / LLM-judge / Cost / Efficiency / Stability shown independently |
66
+ | **Multi-executor** | Claude CLI / Claude SDK / OpenAI / Gemini / any custom command |
67
+ | **MCP URL fetching** | pull content from private-doc URLs via an MCP server (SSO-protected knowledge bases, etc.) |
68
+ | **Blind A/B** | `--blind` hides variant names; HTML report has a reveal button |
69
+ | **Parallel execution** | `--concurrency N` runs N tasks at once |
70
+ | **Multi-run variance** | `--repeat N` repeats the eval and computes mean / SD / CI / t-test |
71
+ | **Auto analysis** | detects low-discrimination assertions, flat scores, all-pass / all-fail, expensive samples |
72
+ | **Traceability** | reports carry CLI version, Node version, artifact version fingerprint |
73
+ | **EN / ZH switch** | one-click language toggle in the HTML report |
74
+
75
+ ## How it works
76
+
77
+ Core idea: **fix the model and the samples, vary only the artifact and runtime context**, use interleaved scheduling to cancel time drift, score via assertions + LLM judge (dual channel), then layer on knowledge-gap signals to quantify risk exposure.
78
+
79
+ ```mermaid
80
+ flowchart TD
81
+ subgraph Input["① Input"]
82
+ S["eval-samples<br/>(JSON / YAML)"]
83
+ A["artifacts<br/>skills/*.md · SKILL.md<br/>baseline · git:name · @cwd"]
84
+ end
85
+
86
+ subgraph Prep["② Preprocess (resolve & fetch)"]
87
+ V["variant resolution<br/>variant → artifact + runtime context<br/>(cwd / project CLAUDE.md / local skills)"]
88
+ U["URL fetching<br/>URLs in prompt / context<br/>MCP Server(private docs) → HTTP"]
89
+ end
90
+
91
+ subgraph Schedule["③ Interleaved + concurrent scheduling"]
92
+ Q["s1-v1 → s1-v2 → s2-v1 → s2-v2 …<br/>--concurrency N · --repeat N"]
93
+ end
94
+
95
+ subgraph Exec["④ Executor (fixed model)"]
96
+ E["claude / claude-sdk / openai / gemini<br/>anthropic-api / openai-api / custom"]
97
+ T["claude-sdk extracts<br/>turns / toolCalls trace"]
98
+ E -.-> T
99
+ end
100
+
101
+ subgraph Score["⑤ Dual-channel scoring"]
102
+ AS["assertions (18 types)<br/>content / structure / cost / latency<br/>agent: tools_called · turns_min …"]
103
+ LS["LLM judge<br/>rubric · dimensions (independent per-dim scores)"]
104
+ CS["composite score<br/>mean of assertion & LLM when both present"]
105
+ AS --> CS
106
+ LS --> CS
107
+ end
108
+
109
+ subgraph Analyze["⑥ Auto analysis + knowledge gaps"]
110
+ D["low-discrimination / flat scores / all-pass or all-fail<br/>expensive samples · variance · t-test"]
111
+ G["knowledge-gap signals<br/>(quantify risk exposure, not completeness proof)"]
112
+ end
113
+
114
+ subgraph Report["⑦ Report"]
115
+ R["Six dims: Fact / Behavior / LLM-judge / Cost / Efficiency / Stability<br/>JSON + HTML · blind reveal<br/>CLI/Node/version fingerprint traceable"]
116
+ end
117
+
118
+ S --> U
119
+ A --> V
120
+ V --> Q
121
+ U --> Q
122
+ Q --> E
123
+ T --> AS
124
+ E --> AS
125
+ E --> LS
126
+ CS --> D
127
+ CS --> G
128
+ D --> R
129
+ G --> R
130
+ ```
131
+
132
+ **Key design choices:**
133
+
134
+ - **Interleaved scheduling** removes time drift: different variants of the same sample are dispatched alternately rather than "all of v1 then all of v2", so model load / network jitter can't be mis-attributed to the artifact.
135
+ - **variant = artifact + runtime context**: `name@cwd` lets control groups explicitly declare the "project directory" input, separating "project-level accumulated knowledge" from "explicit artifact injection".
136
+ - **Dual-channel scoring is complementary**: assertions catch deterministic defects (must call tool X, must contain field Y); the LLM judge catches subjective quality (readability, completeness). Mean is taken when both are present.
137
+ - **Knowledge-gap signals** are not part of the score — they are an independent tracking channel that tells you "how much risk exposure this evaluation covered", for convergence tracking, not as a completeness proof.
138
+
139
+ ## Eval sample format
140
+
141
+ Supports JSON and YAML (`eval-samples.json`, `eval-samples.yaml`, `eval-samples.yml`).
73
142
 
74
143
  ```json
75
144
  [
76
145
  {
77
146
  "sample_id": "s001",
78
- "prompt": "Review this code",
147
+ "prompt": "Review this code for security issues",
79
148
  "context": "function auth(u, p) { db.query('SELECT * FROM users WHERE name=' + u); }",
80
- "rubric": "Should identify SQL injection and suggest parameterized queries",
149
+ "rubric": "Should identify SQL injection risk and recommend parameterized queries",
81
150
  "assertions": [
82
151
  { "type": "contains", "value": "SQL injection", "weight": 1 },
83
152
  { "type": "contains", "value": "parameterized", "weight": 1 },
84
- { "type": "not_contains", "value": "looks good", "weight": 0.5 },
85
- { "type": "json_valid" },
86
- { "type": "cost_max", "value": 0.01 },
87
- { "type": "custom", "fn": "my-assertion.mjs", "weight": 1 }
153
+ { "type": "not_contains", "value": "looks fine", "weight": 0.5 }
88
154
  ],
89
155
  "dimensions": {
90
- "security": "Should identify injection vulnerability",
91
- "actionability": "Should provide concrete fix with code"
156
+ "security": "did it identify the injection vulnerability?",
157
+ "actionability": "did it give directly usable fix code?"
92
158
  }
93
159
  }
94
160
  ]
95
161
  ```
96
162
 
97
- ### Field Reference
163
+ ### Fields
98
164
 
99
165
  | Field | Type | Required | Description |
100
- |-------|------|----------|-------------|
101
- | `sample_id` | `string` | **Yes** | Unique identifier for the sample (e.g., `"s001"`). Used in reports and analysis to reference this test case. |
102
- | `prompt` | `string` | **Yes** | The user prompt sent to the model. This is the task or question the model should answer. |
103
- | `context` | `string` | No | Additional context appended to the prompt (e.g., code snippet, document text). If provided, it is wrapped in a code block and concatenated after `prompt`. |
104
- | `rubric` | `string` | No | Natural language scoring criteria for the LLM judge. The judge model reads this rubric and scores the output 1-5. Use when you need semantic/qualitative evaluation. |
105
- | `assertions` | `array` | No | List of deterministic and async checks applied to the model output. Each assertion is an object with a `type` field (see [Assertion Types](#assertion-types)). |
106
- | `assertions[].type` | `string` | **Yes** | The assertion type (e.g., `"contains"`, `"json_valid"`, `"custom"`). See full list below. |
107
- | `assertions[].value` | `string\|number` | Varies | The value to check against. Required for `contains`, `starts_with`, `equals`, `min_length`, `cost_max`, etc. |
108
- | `assertions[].values` | `array` | Varies | Array of strings. Required for `contains_all` and `contains_any`. |
109
- | `assertions[].pattern` | `string` | Varies | Regex pattern. Required for `regex` type. |
110
- | `assertions[].flags` | `string` | No | Regex flags (default: `"i"`). Only used with `regex` type. |
111
- | `assertions[].schema` | `object` | Varies | JSON Schema object. Required for `json_schema` type. Validated via [ajv](https://ajv.js.org/) (full JSON Schema spec). |
112
- | `assertions[].reference` | `string` | Varies | Reference text for semantic comparison. Required for `semantic_similarity` type. |
113
- | `assertions[].threshold` | `number` | No | Minimum score (1-5) to consider a semantic similarity match passing. Default: `3`. |
114
- | `assertions[].fn` | `string` | Varies | Path to a `.mjs` file exporting the check function. Required for `custom` type. Resolved relative to the samples file directory. |
115
- | `assertions[].weight` | `number` | No | Weight of this assertion in the composite score calculation. Default: `1`. Higher weight = more influence on the final assertion score. |
116
- | `dimensions` | `object` | No | Key-value map for multi-dimensional LLM scoring. Each key is a dimension name (e.g., `"security"`), and the value is the rubric text the LLM judge uses to score that dimension (1-5). Scores are averaged into a single LLM score. |
166
+ |---|---|---|---|
167
+ | `sample_id` | `string` | **yes** | Unique sample ID |
168
+ | `prompt` | `string` | **yes** | User prompt sent to the model |
169
+ | `context` | `string` | no | Extra context (e.g. code). Wrapped in a code block and appended to the prompt. URLs are auto-fetched at runtime. |
170
+ | `rubric` | `string` | no | Scoring guideline for the LLM judge (1-5 scale) |
171
+ | `assertions` | `array` | no | Assertion checks; see [assertion types](#assertion-types) |
172
+ | `assertions[].type` | `string` | **yes** | Assertion type |
173
+ | `assertions[].value` | `string\|number` | depends | Check value (required for `contains`, `min_length`, `cost_max`, etc.) |
174
+ | `assertions[].values` | `array` | depends | String array (required for `contains_all`, `contains_any`) |
175
+ | `assertions[].pattern` | `string` | depends | Regex pattern (required for `regex`) |
176
+ | `assertions[].flags` | `string` | no | Regex flags (default `"i"`) |
177
+ | `assertions[].schema` | `object` | depends | JSON Schema object (required for `json_schema`, via [ajv](https://ajv.js.org/)) |
178
+ | `assertions[].reference` | `string` | depends | Reference text (required for `semantic_similarity`) |
179
+ | `assertions[].threshold` | `number` | no | Pass threshold for semantic similarity (default 3) |
180
+ | `assertions[].fn` | `string` | depends | Path to a custom assertion JS file (required for `custom`) |
181
+ | `assertions[].weight` | `number` | no | Weight (default 1) |
182
+ | `dimensions` | `object` | no | Multi-dimension scoring; key = dimension name, value = scoring guideline |
183
+
184
+ ### URL auto-fetching
185
+
186
+ URLs in `prompt` and `context` are auto-fetched before evaluation and inlined into the text. Useful when referencing online docs, API references, etc.:
117
187
 
118
- **Scoring priority:** If both `assertions` and `rubric`/`dimensions` are present, the composite score is a 50/50 weighted average. If only one is present, that score is used directly. If none are present, the score is 0.
188
+ ```json
189
+ {
190
+ "sample_id": "s001",
191
+ "prompt": "Generate test cases from this PRD: https://wiki.example.com/prd/feature-x"
192
+ }
193
+ ```
194
+
195
+ At runtime, URLs are replaced with the actual content. Fetch order: MCP Server first for matching URLs (e.g. SSO-protected private docs), then plain HTTP for the rest. URLs already resolved by MCP are not re-fetched via HTTP.
196
+
197
+ **Private-doc URLs**: drop a `.mcp.json` config file into the project dir, or pass `--mcp-config <path>`:
198
+
199
+ ```json
200
+ {
201
+ "mcpServers": {
202
+ "docs": {
203
+ "command": "npx",
204
+ "args": ["@example/docs-mcp-server"],
205
+ "env": { "DOCS_API_TOKEN": "xxx" },
206
+ "urlPatterns": ["docs.example.com"],
207
+ "fetchTool": {
208
+ "name": "fetch_doc",
209
+ "urlTransform": {
210
+ "regex": "docs\\.example\\.com/([^/]+/[^/]+)/([^/?#]+)",
211
+ "params": { "namespace": "$1", "slug": "$2" }
212
+ },
213
+ "contentExtract": "data.body"
214
+ }
215
+ }
216
+ }
217
+ }
218
+ ```
119
219
 
120
- **Prompt construction:** The final prompt sent to the model is: `prompt` alone if no `context`, or `prompt + "\n\n```\n" + context + "\n```"` if `context` is provided.
220
+ **Public URLs**: fetched via plain HTTP. If they require auth, make sure the shell already has network access configured (VPN, proxy, etc.).
121
221
 
122
- ### Grading Strategy
222
+ ### Scoring strategy
123
223
 
124
- Each sample can use up to three grading methods. They can be used alone or combined.
224
+ #### 1. Assertion score
125
225
 
126
- #### 1. Assertions (deterministic scoring)
226
+ Rule-based local checks; each assertion yields pass/fail.
127
227
 
128
- Assertions are rule-based checks that run locally without any LLM calls (except `semantic_similarity` and `custom`). Each assertion produces a **pass/fail** result.
228
+ **Formula:**
129
229
 
130
- **How the assertion score is calculated:**
230
+ - Pass rate = sum of passed assertion weights / total weight (0–1)
231
+ - Score = 1 + pass_rate × 4 (mapped to 1–5)
232
+ - Example: 3 assertions (weight 1 each), 2 pass → pass rate 2/3 → score = 1 + 0.67 × 4 = **3.67**
131
233
 
132
- 1. Each assertion has a `weight` (default: 1)
133
- 2. Sum the weights of all passing assertions → `passedWeight`
134
- 3. Sum the weights of all assertions → `totalWeight`
135
- 4. Compute ratio: `passedWeight / totalWeight` (0.0 ~ 1.0)
136
- 5. Normalize to 1-5 scale: **`score = 1 + ratio × 4`**
234
+ #### 2. Rubric / Dimensions score
137
235
 
138
- Example: 3 assertions (weight 1 each), 2 pass → ratio = 2/3 → score = 1 + 2.67 = **3.67**
236
+ The judge model (default `haiku`) scores 1–5 against the rubric. In `dimensions` mode, each dimension is scored independently and then averaged.
139
237
 
140
- #### 2. Rubric (single LLM judge)
238
+ #### 3. Composite score
141
239
 
142
- A judge model (default: `haiku`, configurable via `--judge-model`) reads the model output and scores it against the rubric text. Returns an integer score from **1** (fail) to **5** (excellent) with a brief reason.
240
+ | Condition | Formula |
241
+ |---|---|
242
+ | Only assertions | `assertionScore` |
243
+ | Only LLM judge | `llmScore` |
244
+ | Both present | `(assertionScore + llmScore) / 2` |
245
+ | Neither | `0` |
143
246
 
144
- Only one of `rubric` or `dimensions` should be used per sample. If both are present, `dimensions` takes priority.
247
+ ### Assertion types
145
248
 
146
- #### 3. Dimensions (multi-dimensional LLM judge)
249
+ **Deterministic assertions (18 total):**
147
250
 
148
- Each dimension is scored independently by the judge model (1-5). The dimension scores are **averaged** to produce a single LLM score.
251
+ | Type | Description |
252
+ |---|---|
253
+ | `contains` / `not_contains` | substring must / must-not appear |
254
+ | `regex` | regex match |
255
+ | `min_length` / `max_length` | length bounds |
256
+ | `json_valid` / `json_schema` | JSON validation |
257
+ | `starts_with` / `ends_with` | prefix / suffix |
258
+ | `equals` / `not_equals` | exact match |
259
+ | `word_count_min` / `word_count_max` | word-count bounds |
260
+ | `contains_all` / `contains_any` | multi-value match |
261
+ | `cost_max` / `latency_max` | cost / latency caps |
262
+ | `semantic_similarity` | LLM-based semantic similarity |
263
+ | `custom` | custom JS function (30 s timeout) |
149
264
 
150
- Example: `security: 5`, `actionability: 3` → LLM score = **(5 + 3) / 2 = 4.0**
265
+ ### Custom assertion
151
266
 
152
- #### Composite Score
267
+ ```js
268
+ // my-assertion.mjs
269
+ export default function(output, { sample, assertion }) {
270
+ return { pass: output.includes('SQL'), message: 'checked for SQL keyword' };
271
+ }
272
+ ```
153
273
 
154
- | What's present | Composite score formula |
155
- |----------------|----------------------|
156
- | Assertions only | `assertionScore` |
157
- | LLM only (rubric or dimensions) | `llmScore` |
158
- | Both | `(assertionScore + llmScore) / 2` |
159
- | Neither | `0` |
274
+ ## Six-dim evaluation
160
275
 
161
- All scores are on a **1-5 scale**. A score of 0 means no grading criteria were defined.
276
+ Reports display results across six independent dimensions. The three scoring layers — Fact / Behavior / LLM-judge — are shown separately so you see **which layer regressed** instead of a single composite number:
162
277
 
163
- ### Assertion Types
278
+ | Dimension | Metric | Description |
279
+ |---|---|---|
280
+ | 📋 **Fact** | fact-assertion pass rate | rule-verifiable assertions like `contains` / `json_schema` / `fact_check`, mapped to 1-5 |
281
+ | 🛠️ **Behavior** | behavior-assertion pass rate | execution-compliance assertions like `tools_called` / `tool_output_contains` / `turns_max` |
282
+ | 💬 **LLM-judge** | rubric score | 1-5 scored by the judge model against a predefined rubric; subjective, catches what rules miss |
283
+ | 💰 **Cost** | total cost, input/output tokens | API cost based on token usage and model pricing |
284
+ | ⚡ **Efficiency** | average latency (ms) | end-to-end latency from request to full response |
285
+ | 🛡️ **Stability** | CV (coefficient of variation) | score consistency across repeated runs (`--repeat ≥ 2`); single-run shows `—`, **honestly acknowledging what can't be measured** |
164
286
 
165
- **Deterministic (sync, no LLM):**
287
+ ## CLI reference
166
288
 
167
- | Type | Fields | Description |
168
- |------|--------|-------------|
169
- | `contains` | `value`, `weight` | Output contains substring (case-insensitive) |
170
- | `not_contains` | `value`, `weight` | Output does NOT contain substring |
171
- | `regex` | `pattern`, `flags`, `weight` | Output matches regex |
172
- | `min_length` | `value`, `weight` | Output length >= value |
173
- | `max_length` | `value`, `weight` | Output length <= value |
174
- | `json_valid` | `weight` | Output is valid JSON |
175
- | `json_schema` | `schema`, `weight` | Output matches JSON Schema (full spec via ajv) |
176
- | `starts_with` | `value`, `weight` | Output starts with string (case-insensitive) |
177
- | `ends_with` | `value`, `weight` | Output ends with string (case-insensitive) |
178
- | `equals` | `value`, `weight` | Output exactly equals value (after trim) |
179
- | `not_equals` | `value`, `weight` | Output does not equal value (after trim) |
180
- | `word_count_min` | `value`, `weight` | Word count >= value |
181
- | `word_count_max` | `value`, `weight` | Word count <= value |
182
- | `contains_all` | `values`, `weight` | Output contains ALL substrings |
183
- | `contains_any` | `values`, `weight` | Output contains at least one substring |
184
- | `cost_max` | `value`, `weight` | Execution cost (USD) <= value |
185
- | `latency_max` | `value`, `weight` | Execution latency (ms) <= value |
289
+ ### `omk bench run`
186
290
 
187
- **Async (LLM-based):**
291
+ ```bash
292
+ omk bench run [options]
188
293
 
189
- | Type | Fields | Description |
190
- |------|--------|-------------|
191
- | `semantic_similarity` | `reference`, `threshold`, `weight` | LLM judges similarity to reference text (threshold default: 3) |
192
- | `custom` | `fn`, `weight` | Load external JS function (see below) |
294
+ options:
295
+ --samples <path> sample file (default: eval-samples.json, also detects .yaml/.yml)
296
+ --skill-dir <path> artifact dir (default: skills)
297
+ --control <expr> control-group variant expression (experiment role = control)
298
+ --treatment <v1,v2> treatment-group variant expressions, comma-separated
299
+ at least one of --control / --treatment is required
300
+ (unless you use --config or --each)
301
+ special values: baseline (empty artifact), git:name (git HEAD),
302
+ git:ref:name (specific commit), path with "/" (read file directly)
303
+ --config <path> YAML/JSON config file (evaluation-as-code); declares
304
+ samples + variants + model + executor in one file; CLI
305
+ flags override config fields when both are provided
306
+ --model <name> model under test (default: sonnet)
307
+ --judge-model <name> judge model (default: haiku)
308
+ --output-dir <path> output dir (default: ~/.oh-my-knowledge/reports/)
309
+ --no-judge skip the LLM judge
310
+ --no-cache disable result cache (on by default; identical inputs reuse)
311
+ --dry-run preview only
312
+ --blind blind mode
313
+ --concurrency <n> parallel tasks (default: 1)
314
+ --timeout <sec> per-task executor timeout (default: 120)
315
+ --repeat <n> repeat N times for variance analysis (default: 1)
316
+ --executor <name> executor (default: claude); supports custom commands
317
+ --skip-preflight skip pre-evaluation model reachability check
318
+ --mcp-config <path> MCP config for fetching private-doc URLs via MCP Server
319
+ (default: .mcp.json in cwd)
320
+ --no-serve don't auto-start the report server after the run
321
+ --verbose print per-sample details (duration, tokens, output preview)
322
+ --each batch mode: evaluate each artifact independently vs baseline
323
+ requires {name}.eval-samples.json paired with each artifact
324
+ ```
193
325
 
194
- ### Custom Assertions
326
+ ### `omk bench run --each` (batch mode)
195
327
 
196
- Create a `.mjs` file that exports a function:
328
+ When `skills/` contains several **independent** artifacts, use `--each` to evaluate each one against baseline and produce a merged report.
197
329
 
198
- ```js
199
- // my-assertion.mjs
200
- export default function(output, { sample, assertion }) {
201
- const hasKeyword = output.includes('SQL');
202
- return { pass: hasKeyword, message: 'Checked for SQL keyword' };
203
- }
330
+ ```
331
+ skills/
332
+ ├── asset.md ← artifact file
333
+ ├── asset.eval-samples.json ← paired samples
334
+ ├── home.md
335
+ ├── home.eval-samples.json
336
+ └── product/ ← directory format also supported
337
+ ├── SKILL.md
338
+ └── eval-samples.json
204
339
  ```
205
340
 
206
- Reference it in your sample: `{ "type": "custom", "fn": "my-assertion.mjs" }`. The `fn` path is resolved relative to the samples file directory.
341
+ Pairing rules:
207
342
 
208
- ## Report Example
343
+ - `{name}.md` → looks for `{name}.eval-samples.json` in the same dir
344
+ - `{name}/SKILL.md` → looks for `{name}/eval-samples.json`
345
+ - artifacts without paired samples are skipped with a warning
209
346
 
210
- After running `omk bench run --variants v1,v2`, the tool outputs a JSON report (also saved to `~/.oh-my-knowledge/reports/`):
347
+ ```bash
348
+ omk bench run --each
349
+ omk bench run --each --dry-run
350
+ ```
211
351
 
212
- ```json
213
- {
214
- "id": "2026-03-24T15-30-45-v1-v2",
215
- "meta": {
216
- "variants": ["v1", "v2"],
217
- "model": "sonnet",
218
- "judgeModel": "haiku",
219
- "executor": "claude",
220
- "sampleCount": 3,
221
- "taskCount": 6,
222
- "totalCostUSD": 0.0234,
223
- "timestamp": "2026-03-24T15:30:45.000Z",
224
- "cliVersion": "0.3.0",
225
- "nodeVersion": "v22.0.0",
226
- "skillHashes": { "v1": "a1b2c3d4e5f6", "v2": "f6e5d4c3b2a1" }
227
- },
228
- "summary": {
229
- "v1": {
230
- "totalSamples": 3,
231
- "successCount": 3,
232
- "errorCount": 0,
233
- "avgCompositeScore": 3.67,
234
- "avgAssertionScore": 3.0,
235
- "avgLlmScore": 4.33,
236
- "avgDurationMs": 2500,
237
- "avgTotalTokens": 1850,
238
- "totalCostUSD": 0.0112
239
- },
240
- "v2": {
241
- "totalSamples": 3,
242
- "successCount": 3,
243
- "errorCount": 0,
244
- "avgCompositeScore": 4.5,
245
- "avgAssertionScore": 5.0,
246
- "avgLlmScore": 4.0,
247
- "avgDurationMs": 2800,
248
- "avgTotalTokens": 2100,
249
- "totalCostUSD": 0.0122
250
- }
251
- },
252
- "results": [
253
- {
254
- "sample_id": "s001",
255
- "variants": {
256
- "v1": {
257
- "ok": true,
258
- "compositeScore": 3.5,
259
- "assertions": {
260
- "passed": 1,
261
- "total": 2,
262
- "score": 3.0,
263
- "details": [
264
- { "type": "contains", "value": "SQL injection", "weight": 1, "passed": true },
265
- { "type": "contains", "value": "parameterized", "weight": 1, "passed": false }
266
- ]
267
- },
268
- "llmScore": 4,
269
- "llmReason": "Identified the vulnerability but did not provide a complete fix",
270
- "durationMs": 2300,
271
- "inputTokens": 850,
272
- "outputTokens": 1200,
273
- "totalTokens": 2050,
274
- "costUSD": 0.0038,
275
- "outputPreview": "This code has a SQL injection vulnerability..."
276
- },
277
- "v2": {
278
- "ok": true,
279
- "compositeScore": 4.5,
280
- "assertions": {
281
- "passed": 2,
282
- "total": 2,
283
- "score": 5.0,
284
- "details": [
285
- { "type": "contains", "value": "SQL injection", "weight": 1, "passed": true },
286
- { "type": "contains", "value": "parameterized", "weight": 1, "passed": true }
287
- ]
288
- },
289
- "llmScore": 4,
290
- "llmReason": "Thorough analysis with actionable fix code",
291
- "durationMs": 2600,
292
- "inputTokens": 900,
293
- "outputTokens": 1400,
294
- "totalTokens": 2300,
295
- "costUSD": 0.0042,
296
- "outputPreview": "## Security Issue: SQL Injection\n\nThe code is vulnerable..."
297
- }
298
- }
299
- }
300
- ],
301
- "analysis": {
302
- "insights": [
303
- {
304
- "type": "uniform_scores",
305
- "severity": "info",
306
- "message": "1/3 samples show score difference < 0.5 between variants"
307
- }
308
- ],
309
- "suggestions": []
310
- }
311
- }
352
+ ### `omk bench gen-samples` (generate test cases)
353
+
354
+ Reads an artifact's content and uses an LLM to auto-generate eval-samples. Review and edit them before running eval.
355
+
356
+ ```bash
357
+ # generate for a specific artifact (writes eval-samples.json)
358
+ omk bench gen-samples skills/my-skill.md
359
+
360
+ # batch-generate for every artifact under skills/ that lacks samples
361
+ omk bench gen-samples --each
362
+
363
+ # specify sample count
364
+ omk bench gen-samples skills/my-skill.md --count 10
312
365
  ```
313
366
 
314
- **Reading the report:**
367
+ Options:
315
368
 
316
- - **`summary`** gives a quick comparison — in this example, v2 scores higher (4.5 vs 3.67) because it passes more assertions
317
- - **`results`** shows per-sample detail — you can see exactly which assertions passed/failed and why
318
- - **`analysis`** flags patterns — here it notes one sample has similar scores across variants
319
- - View the HTML version at `http://127.0.0.1:7799/run/{id}` after running `omk bench report`
369
+ ```
370
+ --each batch-generate for every artifact missing samples
371
+ --count <n> samples per artifact (default: 5)
372
+ --model <name> model used for generation (default: sonnet)
373
+ --skill-dir <path> artifact dir (default: skills), used with --each
374
+ ```
320
375
 
321
- ## CLI Reference
376
+ ### `omk bench evolve` (self-iterating improvement)
322
377
 
323
- ### `omk bench run`
378
+ Lets the AI iterate an artifact automatically: evaluate → analyze weak spots → LLM rewrites → evaluate again → keep if the score went up, drop otherwise → repeat.
324
379
 
325
380
  ```bash
326
- omk bench run [options]
381
+ # basic: iterate 5 rounds
382
+ omk bench evolve skills/my-skill.md
383
+
384
+ # set rounds and target score
385
+ omk bench evolve skills/my-skill.md --rounds 10 --target 4.5
386
+ ```
327
387
 
328
388
  Options:
329
- --samples <path> Sample file (default: eval-samples.json, auto-detects .yaml/.yml)
330
- --skill-dir <path> Skill directory (default: skills)
331
- --variants <v1,v2> Variant names (default: v1,v2)
332
- --model <name> Model under test (default: sonnet)
333
- --judge-model <name> Judge model (default: haiku)
334
- --output-dir <path> Output directory (default: ~/.oh-my-knowledge/reports/)
335
- --no-judge Skip LLM judging
336
- --dry-run Preview only
337
- --blind Blind A/B mode: hide variant names in report
338
- --concurrency <n> Number of parallel tasks (default: 1)
339
- --repeat <n> Run evaluation N times for variance analysis (default: 1)
340
- --executor <name> Executor (default: claude)
389
+
341
390
  ```
391
+ --rounds <n> max iteration rounds (default: 5)
392
+ --target <score> stop early when the score reaches this threshold
393
+ --samples <path> sample file (default: eval-samples.json)
394
+ --improve-model <name> model used for rewrites (default: sonnet)
395
+ ```
396
+
397
+ Each round's output is saved under `skills/evolve/` (`my-skill.r0.md`, `my-skill.r1.md`…), so you can `diff` to see what the AI changed. The best round is written back to the original file.
342
398
 
343
399
  ### `omk bench ci`
344
400
 
345
- Run evaluation in CI and exit with pass/fail code.
401
+ Run the evaluation inside CI. Exit code 0 on pass, 1 on fail — can be wired into gates directly.
402
+
403
+ The gate is **three-layer all-pass**: `avgFactScore >= threshold AND avgBehaviorScore >= threshold AND avgJudgeScore >= threshold`. Any layer below threshold is FAIL, and the output shows which layer broke. This stops cases like `fact 4.5→2.5 but judge 3→5` from passing via composite averaging — if one layer regresses, the gate catches it.
346
404
 
347
405
  ```bash
348
406
  omk bench ci [options]
349
-
350
- Options:
351
- (same as "bench run", plus:)
352
- --threshold <number> Minimum composite score to pass (default: 3.5)
407
+ --threshold <number> per-layer minimum score (default: 3.5); applied
408
+ independently to fact / behavior / judge
353
409
  ```
354
410
 
355
- Exit code 0 = all variants pass, 1 = at least one variant below threshold.
356
-
357
411
  ### `omk bench report`
358
412
 
413
+ Start the report server to browse historical reports, submit feedback, and delete reports.
414
+
359
415
  ```bash
360
416
  omk bench report [options]
361
-
362
- Options:
363
- --port <number> Server port (default: 7799)
364
- --reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports/)
417
+ --port <number> server port (default: 7799)
365
418
  ```
366
419
 
367
420
  ### `omk bench init`
368
421
 
369
422
  ```bash
370
- omk bench init [dir] # Scaffold a new eval project
423
+ omk bench init [dir] # scaffold an eval project
371
424
  ```
372
425
 
373
- ## Features
426
+ ## `omk analyze` — production observability
374
427
 
375
- ### Blind A/B Testing
428
+ `omk bench run` is **offline evaluation** (fixed controls, repeatable, scored). Production is different — no control group, no ground truth, no repetition, so scoring isn't valid there. `omk analyze` turns existing Claude Code session traces into **skill-health reports** (coverage, gap signals, execution stability, tokens/latency per skill). It gives you clues about **which skill is worth re-evaluating offline**, not a production score.
376
429
 
377
- Use `--blind` to hide variant names in reports. Variants are randomly labeled as "Variant A", "Variant B", etc. A reveal button in the HTML report shows the mapping.
430
+ ```bash
431
+ # analyze all cc sessions of the current project (auto-infers kb from the trace)
432
+ omk analyze ~/.claude/projects/-Users-you-Documents-my-project
378
433
 
379
- ### Parallel Execution
434
+ # restrict to the last 7 days / 24 hours / 30 minutes
435
+ omk analyze ~/.claude/projects/my-project --last 7d
380
436
 
381
- Use `--concurrency N` to run N tasks in parallel. Tasks maintain interleaved scheduling order to reduce time bias.
437
+ # absolute time window
438
+ omk analyze ~/.claude/projects/my-project --from 2026-04-01T00:00:00Z --to 2026-04-15T23:59:59Z
382
439
 
383
- ### Multi-run Variance Analysis
440
+ # whitelist specific skills
441
+ omk analyze ~/.claude/projects/my-project --skills audit,polish
384
442
 
385
- Use `--repeat N` to run the evaluation N times. The report includes:
386
- - Per-variant mean, standard deviation, 95% confidence interval
387
- - Pairwise Welch's t-test between variants (significance at p < 0.05)
443
+ # override the inferred knowledge-base root
444
+ omk analyze ~/.claude/projects/my-project --kb /path/to/project
445
+ ```
388
446
 
389
- ### Auto-analysis
447
+ The command writes `~/.oh-my-knowledge/analyses/<timestamp>-skill-health.json`. Browse results alongside bench reports with `omk bench report` — the homepage has a "📊 Skill Health Reports" link, and each skill card also has a "trend →" link to its time-series view. For two reports side-by-side, use the compare selector on `/analyses`.
390
448
 
391
- After each evaluation, the toolkit automatically detects:
392
- - **Low-discrimination assertions**: assertions with identical results across all variants
393
- - **Uniform scores**: samples where variants score within 0.5 of each other
394
- - **All-pass / all-fail**: assertions that may be too loose or too strict
395
- - **High-cost samples**: samples with disproportionately high cost
449
+ **What you get per skill:**
396
450
 
397
- Insights and suggestions are shown in the HTML report.
451
+ - **Knowledge usage** — which KB files this skill actually read (coverage %)
452
+ - **Knowledge gaps** — four weighted signals (failed search / model-flagged gap / hedging / repeated miss); hedging goes through an LLM-assisted classifier to filter out business-possibility hedging vs genuine knowledge uncertainty
453
+ - **Execution stability** — tool-failure rate; a skill with > 20% failures gets a warning that its gap signals may be environmental noise rather than real knowledge gaps
454
+ - **Usage cost** — billable tokens (input+output) separate from cached tokens, total duration
398
455
 
399
- ### Human Feedback
456
+ **What this is NOT:**
400
457
 
401
- The HTML report includes star rating (1-5) and comment forms for each sample-variant pair. Feedback is persisted to the report JSON via `POST /api/run/:id/feedback`.
458
+ - Not a general APM (request/response/latency tracing is Langfuse / Datadog territory)
459
+ - Not streaming / alerting (batch only — run on a cron if you want periodic snapshots)
460
+ - Not a production score (no control group, no ground truth — use `omk bench run` for scoring)
402
461
 
403
- ### Traceability
462
+ ## Executors
404
463
 
405
- Reports include `cliVersion`, `nodeVersion`, and `skillHashes` (SHA-256 of each skill file) in metadata for reproducibility.
464
+ ### Built-in executors
406
465
 
407
- ## Executors
466
+ | Executor | When to use | Description |
467
+ |---|---|---|
468
+ | `claude` | default | invokes `claude -p` via Claude CLI |
469
+ | `claude-sdk` | structured output | uses Claude Agent SDK — no stdout parsing, avoids buffer truncation |
470
+ | `openai` | cross-vendor comparison | invokes `openai api` CLI |
471
+ | `gemini` | cross-vendor comparison | invokes `gemini` CLI |
472
+ | `anthropic-api` | no CLI needed | calls Anthropic HTTP API directly (needs `ANTHROPIC_API_KEY`) |
473
+ | `openai-api` | no CLI needed | calls OpenAI HTTP API directly (needs `OPENAI_API_KEY`) |
474
+
475
+ API-direct executors support custom base URLs via env: `ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL`.
408
476
 
409
- Use `--executor` to select which model provider to use.
477
+ ### Custom executor
410
478
 
411
- | Executor | CLI Tool | Default Model | Auth |
412
- |----------|----------|---------------|------|
413
- | `claude` | `claude -p` | `sonnet` | Claude Max plan or API key |
414
- | `openai` | `openai api chat.completions.create` | `gpt-4o` | `OPENAI_API_KEY` env var |
415
- | `gemini` | `gemini` (stdin pipe) | Default Gemini model | Google account or `GOOGLE_API_KEY` |
479
+ Any shell command can serve as an executor, communicating via stdin/stdout JSON:
416
480
 
417
481
  ```bash
418
- # Use OpenAI
419
- omk bench run --executor openai --model gpt-4o --variants v1,v2
482
+ omk bench run --executor "python my_provider.py"
483
+ omk bench run --executor "./my-executor.sh"
484
+ ```
485
+
486
+ **Protocol:**
487
+
488
+ - **input** (stdin): JSON `{"model":"...","system":"...","prompt":"..."}`
489
+ - **output** (stdout): JSON `{"output":"model reply","inputTokens":0,"outputTokens":0,"costUSD":0}`
490
+ - stdout only needs to return the fields you care about; others default to 0. Plain-text output (no tokens/cost parsing) is also fine.
491
+ - non-zero exit code counts as failure
420
492
 
421
- # Use Gemini
422
- omk bench run --executor gemini --model gemini-2.5-pro --variants v1,v2
493
+ ### Artifact directory layout
423
494
 
424
- # Compare the same skill across providers (run separately, compare reports)
425
- omk bench run --executor claude --model sonnet --variants v1,v2
426
- omk bench run --executor openai --model gpt-4o --variants v1,v2
495
+ The built-in executors (claude / openai / gemini) support two artifact layouts, mixable in the same run:
496
+
497
+ ```
498
+ skills/
499
+ ├── v1.md # option 1: plain .md file
500
+ └── my-skill/ # option 2: full artifact dir
501
+ ├── SKILL.md # this file is auto-loaded as system prompt
502
+ ├── config.json # other files don't participate in eval, kept for completeness
503
+ └── scripts/
504
+ ```
505
+
506
+ **Variant resolution rules:**
507
+
508
+ `variant` is the experiment-group expression. After resolution, OMK produces an `artifact` plus an optional `runtime context` (currently mainly `cwd`).
509
+
510
+ | Format | Meaning |
511
+ |---|---|
512
+ | `name` | looks up `name.md` or `name/SKILL.md` in the artifact dir, resolves to one artifact |
513
+ | `baseline` | empty artifact, no system prompt — think "nothing at all" |
514
+ | `project-env@/path/to/project` | empty artifact, but run in the specified project dir — observe project-level runtime context alone |
515
+ | `git:name` | reads the last-committed version of an artifact from git HEAD |
516
+ | `git:ref:name` | reads an artifact from a specific commit |
517
+ | `./path/to/file.md` | path with `/`: read the file directly as an artifact |
518
+ | `variant@/path/to/project` | attach a run dir to any variant; supports `name@cwd`, `git:name@cwd`, `/file.md@cwd` |
519
+
520
+ When both `--control` and `--treatment` are omitted, use `--config eval.yaml` or `--each`. With `--each`, `baseline` is auto-added as control and every discovered artifact becomes a treatment.
521
+
522
+ ```bash
523
+ # explicit: one control, one or more treatments
524
+ omk bench run --control v1 --treatment v2
525
+ omk bench run --control baseline --treatment v1,v2,v3
526
+
527
+ # compare empty artifact vs explicit artifact
528
+ omk bench run --control baseline --treatment my-skill
529
+
530
+ # observe project-level runtime context in isolation (use a self-describing label)
531
+ omk bench run --control baseline --treatment project-env@/path/to/target-project
532
+
533
+ # compare "project-level runtime context" vs "explicit artifact injection"
534
+ omk bench run \
535
+ --control project-env@/path/to/target-project \
536
+ --treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
537
+
538
+ # before vs after (old version read from git history)
539
+ omk bench run --control git:my-skill --treatment my-skill
540
+
541
+ # direct file paths
542
+ omk bench run --control ./old-skill.md --treatment ./new-skill.md
543
+
544
+ # config-file driven (evaluation-as-code)
545
+ omk bench run --config eval.yaml
427
546
  ```
428
547
 
429
548
  **Prerequisites:**
430
- - **claude**: Install [Claude Code](https://claude.ai/code) and authenticate
549
+
550
+ - **claude**: install [Claude Code](https://claude.ai/code) and authenticate
551
+ - **claude-sdk**: install [Claude Code](https://claude.ai/code) and authenticate (uses Agent SDK, no CLI stdout parsing)
552
+ - **anthropic-api**: set the `ANTHROPIC_API_KEY` env var
431
553
  - **openai**: `pip install openai` and set `OPENAI_API_KEY`
432
- - **gemini**: `npm i -g @google/gemini-cli` and authenticate with Google
554
+ - **openai-api**: set the `OPENAI_API_KEY` env var
555
+ - **gemini**: `npm i -g @google/gemini-cli` and authenticate
556
+
557
+ ### Agent evaluation and project-level runtime context
558
+
559
+ When the executor is `claude-sdk`, OMK supports a first pass of agent-aware evaluation.
560
+
561
+ A few concepts worth keeping separate:
562
+
563
+ - `artifact`: the thing being evaluated — baseline, skill, prompt, agent
564
+ - `variant`: the CLI expression for an experiment group
565
+ - `runtime context`: the runtime environment; currently mainly `cwd`. In project-type agent scenarios it includes the project dir, its `CLAUDE.md`, local skills, and any other environmental factors that affect behavior
566
+
567
+ In OMK, `agent` is not a catch-all term and neither is `skill`. A cleaner phrasing: **you are comparing how different artifacts behave under different runtime contexts.**
568
+
569
+ - auto-extracts turns / toolCalls traces
570
+ - supports assertions on tool-call behavior
571
+ - supports running under a specified `cwd`, so Claude Code auto-loads the project's `CLAUDE.md`, skills, and local runtime context
572
+
573
+ #### Recommended executor
574
+
575
+ ```bash
576
+ omk bench run --executor claude-sdk
577
+ ```
578
+
579
+ #### Agent-related assertions
580
+
581
+ | Assertion | Meaning |
582
+ |---|---|
583
+ | `tools_called` | must call the specified tool(s) |
584
+ | `tools_not_called` | must not call the specified tool(s) |
585
+ | `tools_count_min` / `tools_count_max` | tool-call-count bounds |
586
+ | `tool_output_contains` | output of a specific tool must contain given content |
587
+ | `turns_min` / `turns_max` | turn-count bounds |
588
+
589
+ #### Three common control setups
590
+
591
+ **1. Bare-model baseline**
433
592
 
434
- ## Environment Variables
593
+ No system prompt and no knowledge-carrying project dir. Requires at least one treatment to compare against:
594
+
595
+ ```bash
596
+ omk bench run \
597
+ --executor claude-sdk \
598
+ --control baseline \
599
+ --treatment my-skill
600
+ ```
601
+
602
+ **2. Empty artifact + project-level runtime context**
603
+
604
+ No system prompt, but runs inside a project dir. This is **not** a strict "bare baseline" — it is "empty artifact + project-level runtime context".
605
+
606
+ ```bash
607
+ omk bench run \
608
+ --executor claude-sdk \
609
+ --control baseline \
610
+ --treatment project-env@/path/to/target-project
611
+ ```
612
+
613
+ **3. Explicit artifact injection**
614
+
615
+ Inject an external `SKILL.md` as the artifact while also keeping the project dir. Good for contrasting "project-level runtime context" vs "explicit single-artifact injection".
616
+
617
+ ```bash
618
+ omk bench run \
619
+ --executor claude-sdk \
620
+ --control project-env@/path/to/target-project \
621
+ --treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
622
+ ```
623
+
624
+ #### Recommended first-round design
625
+
626
+ For PRD / complex business-knowledge scenarios, start with:
627
+
628
+ ```bash
629
+ omk bench run \
630
+ --executor claude-sdk \
631
+ --samples skills/evaluate-review/eval-samples.yaml \
632
+ --control baseline \
633
+ --treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
634
+ ```
635
+
636
+ If you want to prove whether "the knowledge sitting inside the project directory" is effective on its own, add a second treatment:
637
+
638
+ ```bash
639
+ omk bench run \
640
+ --executor claude-sdk \
641
+ --samples skills/evaluate-review/eval-samples.yaml \
642
+ --control baseline \
643
+ --treatment project-env@/path/to/target-project,/path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
644
+ ```
645
+
646
+ #### Design tips
647
+
648
+ - **Always start with `--dry-run`** to confirm samples, variants, and `cwd` are parsed correctly
649
+ - **Project-level controls must differ in `cwd`**: the same prompt under different project dirs hits different runtime contexts
650
+ - **Try PRD scenarios first**: compared to pure coding, they make it easier to validate knowledge completeness, impact-area detection, and business correctness
651
+
652
+ ### Common model configurations
653
+
654
+ **Don't have Claude?** Most Chinese LLMs (GLM, Qwen, Moonshot, DeepSeek, etc.) are OpenAI-API compatible — use the `openai-api` executor directly:
655
+
656
+ ```bash
657
+ # GLM (Zhipu)
658
+ export OPENAI_API_KEY="your Zhipu API key"
659
+ export OPENAI_BASE_URL="https://open.bigmodel.cn/api/paas/v4"
660
+ omk bench run --executor openai-api --model glm-4-plus \
661
+ --judge-model glm-4-plus --no-cache
662
+
663
+ # Qwen (Alibaba)
664
+ export OPENAI_API_KEY="your Qwen API key"
665
+ export OPENAI_BASE_URL="https://dashscope.aliyuncs.com/compatible-mode/v1"
666
+ omk bench run --executor openai-api --model qwen-plus \
667
+ --judge-model qwen-plus
668
+
669
+ # DeepSeek
670
+ export OPENAI_API_KEY="your DeepSeek API key"
671
+ export OPENAI_BASE_URL="https://api.deepseek.com"
672
+ omk bench run --executor openai-api --model deepseek-chat \
673
+ --judge-model deepseek-chat
674
+
675
+ # Moonshot (Kimi)
676
+ export OPENAI_API_KEY="your Moonshot API key"
677
+ export OPENAI_BASE_URL="https://api.moonshot.cn/v1"
678
+ omk bench run --executor openai-api --model moonshot-v1-8k \
679
+ --judge-model moonshot-v1-8k
680
+ ```
681
+
682
+ **Ollama local model:**
683
+
684
+ ```bash
685
+ omk bench run --executor "python examples/custom-executor/ollama-executor.py" \
686
+ --model llama3 --no-judge
687
+ ```
688
+
689
+ **About the judge model:**
690
+
691
+ - `--judge-model` picks the model used by the LLM judge (default `haiku`)
692
+ - `--judge-executor` picks the executor the judge uses (defaults to `--executor`)
693
+ - If you don't have Claude, point `--judge-executor` and `--judge-model` at whatever model you have
694
+ - Add `--no-judge` to skip the LLM judge and rely on assertions alone
695
+
696
+ ## Environment variables
435
697
 
436
698
  | Variable | Description |
437
- |----------|-------------|
438
- | `CCV_PROXY_URL` | Route requests through cc-viewer proxy for real-time visualization |
439
- | `OMK_BENCH_PORT` | Report server port (default: 7799) |
699
+ |---|---|
700
+ | `CCV_PROXY_URL` | proxy requests through cc-viewer for live eval-traffic visualization |
701
+ | `OMK_BENCH_PORT` | report server port (default: 7799) |
440
702
 
441
703
  ## Requirements
442
704
 
443
705
  - Node.js >= 20
444
- - `claude` CLI installed and authenticated (Max plan works, no API key needed)
706
+ - `claude` CLI (for the default executor and LLM judge; see [Claude Code](https://claude.ai/code))
707
+ - not needed if you use other executors (openai / gemini) with `--no-judge`
708
+
709
+ ## Security notice
710
+
711
+ This tool is designed for **local trusted environments** (dev machines, CI pipelines). The following features execute local code — make sure inputs come from a trusted source:
712
+
713
+ | Feature | Risk | Scope |
714
+ |---|---|---|
715
+ | **Custom assertions** (`custom`) | dynamically loads and executes user-specified `.mjs` files | only use assertion files you authored or reviewed |
716
+ | **eval-samples.json** | assertion configs can reference external file paths | don't use sample files from untrusted sources |
717
+
718
+ **Recommendations:**
719
+
720
+ - Do not expose `omk bench report` on the public internet (no auth)
721
+ - Don't use third-party eval-samples you haven't vetted
722
+ - Custom assertions have a 30-second timeout but no sandbox isolation
445
723
 
446
- ## License
724
+ ---
447
725
 
448
- MIT
726
+ See [CHANGELOG](./CHANGELOG.md) for release notes. Contributions welcome — see [CONTRIBUTING](./CONTRIBUTING.md).