oh-my-knowledge 0.0.1 → 0.19.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +600 -322
- package/README.zh.md +724 -0
- package/dist/src/analysis/coverage-analyzer.d.ts +57 -0
- package/dist/src/analysis/coverage-analyzer.d.ts.map +1 -0
- package/dist/src/analysis/coverage-analyzer.js +262 -0
- package/dist/src/analysis/coverage-analyzer.js.map +1 -0
- package/dist/src/analysis/gap-analyzer.d.ts +114 -0
- package/dist/src/analysis/gap-analyzer.d.ts.map +1 -0
- package/dist/src/analysis/gap-analyzer.js +439 -0
- package/dist/src/analysis/gap-analyzer.js.map +1 -0
- package/dist/src/analysis/hedging-classifier.d.ts +34 -0
- package/dist/src/analysis/hedging-classifier.d.ts.map +1 -0
- package/dist/src/analysis/hedging-classifier.js +144 -0
- package/dist/src/analysis/hedging-classifier.js.map +1 -0
- package/dist/src/analysis/report-diagnostics.d.ts +9 -0
- package/dist/src/analysis/report-diagnostics.d.ts.map +1 -0
- package/dist/src/analysis/report-diagnostics.js +631 -0
- package/dist/src/analysis/report-diagnostics.js.map +1 -0
- package/dist/src/authoring/evolver.d.ts +64 -0
- package/dist/src/authoring/evolver.d.ts.map +1 -0
- package/dist/src/authoring/evolver.js +275 -0
- package/dist/src/authoring/evolver.js.map +1 -0
- package/dist/src/authoring/generator.d.ts +13 -0
- package/dist/src/authoring/generator.d.ts.map +1 -0
- package/dist/src/authoring/generator.js +60 -0
- package/dist/src/authoring/generator.js.map +1 -0
- package/dist/src/cli.d.ts +3 -0
- package/dist/src/cli.d.ts.map +1 -0
- package/dist/src/cli.js +1012 -0
- package/dist/src/cli.js.map +1 -0
- package/dist/src/eval-core/cache.d.ts +11 -0
- package/dist/src/eval-core/cache.d.ts.map +1 -0
- package/dist/src/eval-core/cache.js +53 -0
- package/dist/src/eval-core/cache.js.map +1 -0
- package/dist/src/eval-core/ci-gates.d.ts +17 -0
- package/dist/src/eval-core/ci-gates.d.ts.map +1 -0
- package/dist/src/eval-core/ci-gates.js +42 -0
- package/dist/src/eval-core/ci-gates.js.map +1 -0
- package/dist/src/eval-core/dependency-checker.d.ts +37 -0
- package/dist/src/eval-core/dependency-checker.d.ts.map +1 -0
- package/dist/src/eval-core/dependency-checker.js +280 -0
- package/dist/src/eval-core/dependency-checker.js.map +1 -0
- package/dist/src/eval-core/evaluation-execution.d.ts +26 -0
- package/dist/src/eval-core/evaluation-execution.d.ts.map +1 -0
- package/dist/src/eval-core/evaluation-execution.js +196 -0
- package/dist/src/eval-core/evaluation-execution.js.map +1 -0
- package/dist/src/eval-core/evaluation-job.d.ts +50 -0
- package/dist/src/eval-core/evaluation-job.d.ts.map +1 -0
- package/dist/src/eval-core/evaluation-job.js +114 -0
- package/dist/src/eval-core/evaluation-job.js.map +1 -0
- package/dist/src/eval-core/evaluation-reporting.d.ts +28 -0
- package/dist/src/eval-core/evaluation-reporting.d.ts.map +1 -0
- package/dist/src/eval-core/evaluation-reporting.js +123 -0
- package/dist/src/eval-core/evaluation-reporting.js.map +1 -0
- package/dist/src/eval-core/execution-strategy.d.ts +11 -0
- package/dist/src/eval-core/execution-strategy.d.ts.map +1 -0
- package/dist/src/eval-core/execution-strategy.js +121 -0
- package/dist/src/eval-core/execution-strategy.js.map +1 -0
- package/dist/src/eval-core/fact-checker.d.ts +25 -0
- package/dist/src/eval-core/fact-checker.d.ts.map +1 -0
- package/dist/src/eval-core/fact-checker.js +66 -0
- package/dist/src/eval-core/fact-checker.js.map +1 -0
- package/dist/src/eval-core/schema.d.ts +30 -0
- package/dist/src/eval-core/schema.d.ts.map +1 -0
- package/dist/src/eval-core/schema.js +195 -0
- package/dist/src/eval-core/schema.js.map +1 -0
- package/dist/src/eval-core/statistics.d.ts +86 -0
- package/dist/src/eval-core/statistics.d.ts.map +1 -0
- package/dist/src/eval-core/statistics.js +210 -0
- package/dist/src/eval-core/statistics.js.map +1 -0
- package/dist/src/eval-core/task-planner.d.ts +4 -0
- package/dist/src/eval-core/task-planner.d.ts.map +1 -0
- package/dist/src/eval-core/task-planner.js +33 -0
- package/dist/src/eval-core/task-planner.js.map +1 -0
- package/dist/src/eval-workflows/each-evaluation-workflow.d.ts +137 -0
- package/dist/src/eval-workflows/each-evaluation-workflow.d.ts.map +1 -0
- package/dist/src/eval-workflows/each-evaluation-workflow.js +174 -0
- package/dist/src/eval-workflows/each-evaluation-workflow.js.map +1 -0
- package/dist/src/eval-workflows/evaluation-pipeline.d.ts +42 -0
- package/dist/src/eval-workflows/evaluation-pipeline.d.ts.map +1 -0
- package/dist/src/eval-workflows/evaluation-pipeline.js +218 -0
- package/dist/src/eval-workflows/evaluation-pipeline.js.map +1 -0
- package/dist/src/eval-workflows/evaluation-preparation.d.ts +64 -0
- package/dist/src/eval-workflows/evaluation-preparation.d.ts.map +1 -0
- package/dist/src/eval-workflows/evaluation-preparation.js +90 -0
- package/dist/src/eval-workflows/evaluation-preparation.js.map +1 -0
- package/dist/src/eval-workflows/run-evaluation.d.ts +111 -0
- package/dist/src/eval-workflows/run-evaluation.d.ts.map +1 -0
- package/dist/src/eval-workflows/run-evaluation.js +272 -0
- package/dist/src/eval-workflows/run-evaluation.js.map +1 -0
- package/dist/src/executors/anthropic-api.d.ts +3 -0
- package/dist/src/executors/anthropic-api.d.ts.map +1 -0
- package/dist/src/executors/anthropic-api.js +43 -0
- package/dist/src/executors/anthropic-api.js.map +1 -0
- package/dist/src/executors/claude-cli.d.ts +3 -0
- package/dist/src/executors/claude-cli.d.ts.map +1 -0
- package/dist/src/executors/claude-cli.js +103 -0
- package/dist/src/executors/claude-cli.js.map +1 -0
- package/dist/src/executors/claude-sdk-trace.d.ts +10 -0
- package/dist/src/executors/claude-sdk-trace.d.ts.map +1 -0
- package/dist/src/executors/claude-sdk-trace.js +100 -0
- package/dist/src/executors/claude-sdk-trace.js.map +1 -0
- package/dist/src/executors/claude-sdk.d.ts +3 -0
- package/dist/src/executors/claude-sdk.d.ts.map +1 -0
- package/dist/src/executors/claude-sdk.js +160 -0
- package/dist/src/executors/claude-sdk.js.map +1 -0
- package/dist/src/executors/gemini.d.ts +3 -0
- package/dist/src/executors/gemini.d.ts.map +1 -0
- package/dist/src/executors/gemini.js +78 -0
- package/dist/src/executors/gemini.js.map +1 -0
- package/dist/src/executors/index.d.ts +7 -0
- package/dist/src/executors/index.d.ts.map +1 -0
- package/dist/src/executors/index.js +22 -0
- package/dist/src/executors/index.js.map +1 -0
- package/dist/src/executors/openai-api.d.ts +3 -0
- package/dist/src/executors/openai-api.d.ts.map +1 -0
- package/dist/src/executors/openai-api.js +40 -0
- package/dist/src/executors/openai-api.js.map +1 -0
- package/dist/src/executors/openai-cli.d.ts +3 -0
- package/dist/src/executors/openai-cli.d.ts.map +1 -0
- package/dist/src/executors/openai-cli.js +60 -0
- package/dist/src/executors/openai-cli.js.map +1 -0
- package/dist/src/executors/script.d.ts +3 -0
- package/dist/src/executors/script.d.ts.map +1 -0
- package/dist/src/executors/script.js +63 -0
- package/dist/src/executors/script.js.map +1 -0
- package/dist/src/executors/shared.d.ts +117 -0
- package/dist/src/executors/shared.d.ts.map +1 -0
- package/dist/src/executors/shared.js +49 -0
- package/dist/src/executors/shared.js.map +1 -0
- package/dist/src/grading/assertions.d.ts +18 -0
- package/dist/src/grading/assertions.d.ts.map +1 -0
- package/dist/src/grading/assertions.js +239 -0
- package/dist/src/grading/assertions.js.map +1 -0
- package/dist/src/grading/index.d.ts +26 -0
- package/dist/src/grading/index.d.ts.map +1 -0
- package/dist/src/grading/index.js +98 -0
- package/dist/src/grading/index.js.map +1 -0
- package/dist/src/grading/judge.d.ts +13 -0
- package/dist/src/grading/judge.d.ts.map +1 -0
- package/dist/src/grading/judge.js +98 -0
- package/dist/src/grading/judge.js.map +1 -0
- package/dist/src/grading/layered-scores.d.ts +13 -0
- package/dist/src/grading/layered-scores.d.ts.map +1 -0
- package/dist/src/grading/layered-scores.js +62 -0
- package/dist/src/grading/layered-scores.js.map +1 -0
- package/dist/src/inputs/eval-config.d.ts +13 -0
- package/dist/src/inputs/eval-config.d.ts.map +1 -0
- package/dist/src/inputs/eval-config.js +136 -0
- package/dist/src/inputs/eval-config.js.map +1 -0
- package/dist/src/inputs/load-samples.d.ts +16 -0
- package/dist/src/inputs/load-samples.d.ts.map +1 -0
- package/dist/src/inputs/load-samples.js +51 -0
- package/dist/src/inputs/load-samples.js.map +1 -0
- package/dist/src/inputs/mcp-resolver.d.ts +50 -0
- package/dist/src/inputs/mcp-resolver.d.ts.map +1 -0
- package/dist/src/inputs/mcp-resolver.js +307 -0
- package/dist/src/inputs/mcp-resolver.js.map +1 -0
- package/dist/src/inputs/skill-loader.d.ts +18 -0
- package/dist/src/inputs/skill-loader.d.ts.map +1 -0
- package/dist/src/inputs/skill-loader.js +222 -0
- package/dist/src/inputs/skill-loader.js.map +1 -0
- package/dist/src/inputs/url-fetcher.d.ts +19 -0
- package/dist/src/inputs/url-fetcher.d.ts.map +1 -0
- package/dist/src/inputs/url-fetcher.js +241 -0
- package/dist/src/inputs/url-fetcher.js.map +1 -0
- package/dist/src/observability/skill-health-analyzer.d.ts +83 -0
- package/dist/src/observability/skill-health-analyzer.d.ts.map +1 -0
- package/dist/src/observability/skill-health-analyzer.js +238 -0
- package/dist/src/observability/skill-health-analyzer.js.map +1 -0
- package/dist/src/observability/trace-adapter.d.ts +75 -0
- package/dist/src/observability/trace-adapter.d.ts.map +1 -0
- package/dist/src/observability/trace-adapter.js +367 -0
- package/dist/src/observability/trace-adapter.js.map +1 -0
- package/dist/src/renderer/html-renderer.d.ts +9 -0
- package/dist/src/renderer/html-renderer.d.ts.map +1 -0
- package/dist/src/renderer/html-renderer.js +360 -0
- package/dist/src/renderer/html-renderer.js.map +1 -0
- package/dist/src/renderer/layout.d.ts +13 -0
- package/dist/src/renderer/layout.d.ts.map +1 -0
- package/dist/src/renderer/layout.js +582 -0
- package/dist/src/renderer/layout.js.map +1 -0
- package/dist/src/renderer/skill-health-renderer.d.ts +20 -0
- package/dist/src/renderer/skill-health-renderer.d.ts.map +1 -0
- package/dist/src/renderer/skill-health-renderer.js +253 -0
- package/dist/src/renderer/skill-health-renderer.js.map +1 -0
- package/dist/src/renderer/summary.d.ts +21 -0
- package/dist/src/renderer/summary.d.ts.map +1 -0
- package/dist/src/renderer/summary.js +1072 -0
- package/dist/src/renderer/summary.js.map +1 -0
- package/dist/src/renderer/table.d.ts +3 -0
- package/dist/src/renderer/table.d.ts.map +1 -0
- package/dist/src/renderer/table.js +134 -0
- package/dist/src/renderer/table.js.map +1 -0
- package/dist/src/renderer/trends.d.ts +6 -0
- package/dist/src/renderer/trends.d.ts.map +1 -0
- package/dist/src/renderer/trends.js +127 -0
- package/dist/src/renderer/trends.js.map +1 -0
- package/dist/src/server/job-store.d.ts +4 -0
- package/dist/src/server/job-store.d.ts.map +1 -0
- package/dist/src/server/job-store.js +68 -0
- package/dist/src/server/job-store.js.map +1 -0
- package/dist/src/server/report-server.d.ts +17 -0
- package/dist/src/server/report-server.d.ts.map +1 -0
- package/dist/src/server/report-server.js +591 -0
- package/dist/src/server/report-server.js.map +1 -0
- package/dist/src/server/report-store.d.ts +44 -0
- package/dist/src/server/report-store.d.ts.map +1 -0
- package/dist/src/server/report-store.js +190 -0
- package/dist/src/server/report-store.js.map +1 -0
- package/dist/src/types.d.ts +559 -0
- package/dist/src/types.d.ts.map +1 -0
- package/dist/src/types.js +2 -0
- package/dist/src/types.js.map +1 -0
- package/package.json +53 -5
package/README.md
CHANGED
|
@@ -1,448 +1,726 @@
|
|
|
1
1
|
# oh-my-knowledge
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
[](https://www.npmjs.com/package/oh-my-knowledge)
|
|
4
|
+
[](https://github.com/lizhiyao/oh-my-knowledge/actions/workflows/ci.yml)
|
|
5
|
+
[](./LICENSE)
|
|
6
|
+
[](https://nodejs.org)
|
|
4
7
|
|
|
5
|
-
English | [
|
|
8
|
+
**English** | [简体中文](./README.zh.md)
|
|
6
9
|
|
|
7
|
-
|
|
10
|
+
Knowledge-artifact evaluation toolkit — measure your artifact's quality with objective data.
|
|
8
11
|
|
|
9
|
-
|
|
12
|
+
**Fix the model, vary the knowledge artifact, let the data speak.**
|
|
10
13
|
|
|
11
|
-
|
|
14
|
+
## Why this tool
|
|
12
15
|
|
|
13
|
-
|
|
16
|
+
Teams doing knowledge engineering produce lots of knowledge artifacts (skills today, but also prompts, agents, workflows…). When someone asks "why is v2 better than v1", you need objective data instead of gut feeling. `oh-my-knowledge` solves this with controlled experiments: **same model, same test samples, only the knowledge artifact changes.**
|
|
17
|
+
|
|
18
|
+
## Key features
|
|
19
|
+
|
|
20
|
+
- **Controlled-variable offline bench** — fix the model and samples, vary only the artifact; works with Claude Code skills, CLAUDE.md prompts, RAG knowledge bases, or any markdown-based instruction
|
|
21
|
+
- **Six-dimension scoring** — separate signals for Fact / Behavior / LLM-judge / Cost / Efficiency / Stability, so a regression in one axis isn't hidden by gains in another
|
|
22
|
+
- **Production session observability** — parse Claude Code session JSONL traces, measure per-skill failure rate, latency, token cost, and knowledge-gap signals on real user sessions
|
|
23
|
+
- **Knowledge-gap detection** — severity-weighted signals (explicit markers / failed searches / hedging language / repeated failures) quantify risk exposure instead of claiming completeness
|
|
24
|
+
- **Pre-merge CI gate** — `omk bench ci` enforces three-layer all-pass (fact + behavior + llm-judge) semantics, catching single-layer regressions a composite score would hide
|
|
25
|
+
|
|
26
|
+
## Quick start
|
|
14
27
|
|
|
15
28
|
```bash
|
|
16
|
-
#
|
|
17
|
-
npm i
|
|
29
|
+
# install
|
|
30
|
+
npm i oh-my-knowledge -g
|
|
18
31
|
|
|
19
|
-
#
|
|
32
|
+
# scaffold an eval project
|
|
20
33
|
omk bench init my-eval
|
|
21
34
|
cd my-eval
|
|
22
35
|
|
|
23
|
-
#
|
|
36
|
+
# drop the artifacts you want to compare into skills/
|
|
37
|
+
# option 1: plain .md files (skills/v1.md, skills/v2.md)
|
|
38
|
+
# option 2: full artifact dirs (skills/my-skill-v1/SKILL.md, ...)
|
|
39
|
+
# a single artifact also works — baseline is auto-added as control
|
|
40
|
+
|
|
41
|
+
# preview the plan
|
|
24
42
|
omk bench run --dry-run
|
|
25
43
|
|
|
26
|
-
#
|
|
27
|
-
omk bench run
|
|
28
|
-
|
|
29
|
-
|
|
30
|
-
|
|
31
|
-
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
52
|
-
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
|
|
64
|
-
|
|
65
|
-
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
|
|
71
|
-
|
|
72
|
-
|
|
44
|
+
# run the evaluation (auto-discovers everything under skills/)
|
|
45
|
+
omk bench run
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
## Use inside Claude Code
|
|
49
|
+
|
|
50
|
+
After installing omk, talk to it in natural language from Claude Code:
|
|
51
|
+
|
|
52
|
+
```
|
|
53
|
+
/omk eval # evaluate the artifact(s) in the current project
|
|
54
|
+
/omk evolve # auto-iterate to improve an artifact
|
|
55
|
+
/omk gen-samples # generate test cases
|
|
56
|
+
```
|
|
57
|
+
|
|
58
|
+
You can also just say "compare v1 vs v2 for me" or "improve this artifact" — omk picks the right command.
|
|
59
|
+
|
|
60
|
+
## Features
|
|
61
|
+
|
|
62
|
+
| Feature | What it does |
|
|
63
|
+
|---|---|
|
|
64
|
+
| **18 assertion types** | substring, regex, JSON Schema, semantic similarity, custom JS function, and more |
|
|
65
|
+
| **Six-dim evaluation** | Fact / Behavior / LLM-judge / Cost / Efficiency / Stability shown independently |
|
|
66
|
+
| **Multi-executor** | Claude CLI / Claude SDK / OpenAI / Gemini / any custom command |
|
|
67
|
+
| **MCP URL fetching** | pull content from private-doc URLs via an MCP server (SSO-protected knowledge bases, etc.) |
|
|
68
|
+
| **Blind A/B** | `--blind` hides variant names; HTML report has a reveal button |
|
|
69
|
+
| **Parallel execution** | `--concurrency N` runs N tasks at once |
|
|
70
|
+
| **Multi-run variance** | `--repeat N` repeats the eval and computes mean / SD / CI / t-test |
|
|
71
|
+
| **Auto analysis** | detects low-discrimination assertions, flat scores, all-pass / all-fail, expensive samples |
|
|
72
|
+
| **Traceability** | reports carry CLI version, Node version, artifact version fingerprint |
|
|
73
|
+
| **EN / ZH switch** | one-click language toggle in the HTML report |
|
|
74
|
+
|
|
75
|
+
## How it works
|
|
76
|
+
|
|
77
|
+
Core idea: **fix the model and the samples, vary only the artifact and runtime context**, use interleaved scheduling to cancel time drift, score via assertions + LLM judge (dual channel), then layer on knowledge-gap signals to quantify risk exposure.
|
|
78
|
+
|
|
79
|
+
```mermaid
|
|
80
|
+
flowchart TD
|
|
81
|
+
subgraph Input["① Input"]
|
|
82
|
+
S["eval-samples<br/>(JSON / YAML)"]
|
|
83
|
+
A["artifacts<br/>skills/*.md · SKILL.md<br/>baseline · git:name · @cwd"]
|
|
84
|
+
end
|
|
85
|
+
|
|
86
|
+
subgraph Prep["② Preprocess (resolve & fetch)"]
|
|
87
|
+
V["variant resolution<br/>variant → artifact + runtime context<br/>(cwd / project CLAUDE.md / local skills)"]
|
|
88
|
+
U["URL fetching<br/>URLs in prompt / context<br/>MCP Server(private docs) → HTTP"]
|
|
89
|
+
end
|
|
90
|
+
|
|
91
|
+
subgraph Schedule["③ Interleaved + concurrent scheduling"]
|
|
92
|
+
Q["s1-v1 → s1-v2 → s2-v1 → s2-v2 …<br/>--concurrency N · --repeat N"]
|
|
93
|
+
end
|
|
94
|
+
|
|
95
|
+
subgraph Exec["④ Executor (fixed model)"]
|
|
96
|
+
E["claude / claude-sdk / openai / gemini<br/>anthropic-api / openai-api / custom"]
|
|
97
|
+
T["claude-sdk extracts<br/>turns / toolCalls trace"]
|
|
98
|
+
E -.-> T
|
|
99
|
+
end
|
|
100
|
+
|
|
101
|
+
subgraph Score["⑤ Dual-channel scoring"]
|
|
102
|
+
AS["assertions (18 types)<br/>content / structure / cost / latency<br/>agent: tools_called · turns_min …"]
|
|
103
|
+
LS["LLM judge<br/>rubric · dimensions (independent per-dim scores)"]
|
|
104
|
+
CS["composite score<br/>mean of assertion & LLM when both present"]
|
|
105
|
+
AS --> CS
|
|
106
|
+
LS --> CS
|
|
107
|
+
end
|
|
108
|
+
|
|
109
|
+
subgraph Analyze["⑥ Auto analysis + knowledge gaps"]
|
|
110
|
+
D["low-discrimination / flat scores / all-pass or all-fail<br/>expensive samples · variance · t-test"]
|
|
111
|
+
G["knowledge-gap signals<br/>(quantify risk exposure, not completeness proof)"]
|
|
112
|
+
end
|
|
113
|
+
|
|
114
|
+
subgraph Report["⑦ Report"]
|
|
115
|
+
R["Six dims: Fact / Behavior / LLM-judge / Cost / Efficiency / Stability<br/>JSON + HTML · blind reveal<br/>CLI/Node/version fingerprint traceable"]
|
|
116
|
+
end
|
|
117
|
+
|
|
118
|
+
S --> U
|
|
119
|
+
A --> V
|
|
120
|
+
V --> Q
|
|
121
|
+
U --> Q
|
|
122
|
+
Q --> E
|
|
123
|
+
T --> AS
|
|
124
|
+
E --> AS
|
|
125
|
+
E --> LS
|
|
126
|
+
CS --> D
|
|
127
|
+
CS --> G
|
|
128
|
+
D --> R
|
|
129
|
+
G --> R
|
|
130
|
+
```
|
|
131
|
+
|
|
132
|
+
**Key design choices:**
|
|
133
|
+
|
|
134
|
+
- **Interleaved scheduling** removes time drift: different variants of the same sample are dispatched alternately rather than "all of v1 then all of v2", so model load / network jitter can't be mis-attributed to the artifact.
|
|
135
|
+
- **variant = artifact + runtime context**: `name@cwd` lets control groups explicitly declare the "project directory" input, separating "project-level accumulated knowledge" from "explicit artifact injection".
|
|
136
|
+
- **Dual-channel scoring is complementary**: assertions catch deterministic defects (must call tool X, must contain field Y); the LLM judge catches subjective quality (readability, completeness). Mean is taken when both are present.
|
|
137
|
+
- **Knowledge-gap signals** are not part of the score — they are an independent tracking channel that tells you "how much risk exposure this evaluation covered", for convergence tracking, not as a completeness proof.
|
|
138
|
+
|
|
139
|
+
## Eval sample format
|
|
140
|
+
|
|
141
|
+
Supports JSON and YAML (`eval-samples.json`, `eval-samples.yaml`, `eval-samples.yml`).
|
|
73
142
|
|
|
74
143
|
```json
|
|
75
144
|
[
|
|
76
145
|
{
|
|
77
146
|
"sample_id": "s001",
|
|
78
|
-
"prompt": "Review this code",
|
|
147
|
+
"prompt": "Review this code for security issues",
|
|
79
148
|
"context": "function auth(u, p) { db.query('SELECT * FROM users WHERE name=' + u); }",
|
|
80
|
-
"rubric": "Should identify SQL injection and
|
|
149
|
+
"rubric": "Should identify SQL injection risk and recommend parameterized queries",
|
|
81
150
|
"assertions": [
|
|
82
151
|
{ "type": "contains", "value": "SQL injection", "weight": 1 },
|
|
83
152
|
{ "type": "contains", "value": "parameterized", "weight": 1 },
|
|
84
|
-
{ "type": "not_contains", "value": "looks
|
|
85
|
-
{ "type": "json_valid" },
|
|
86
|
-
{ "type": "cost_max", "value": 0.01 },
|
|
87
|
-
{ "type": "custom", "fn": "my-assertion.mjs", "weight": 1 }
|
|
153
|
+
{ "type": "not_contains", "value": "looks fine", "weight": 0.5 }
|
|
88
154
|
],
|
|
89
155
|
"dimensions": {
|
|
90
|
-
"security": "
|
|
91
|
-
"actionability": "
|
|
156
|
+
"security": "did it identify the injection vulnerability?",
|
|
157
|
+
"actionability": "did it give directly usable fix code?"
|
|
92
158
|
}
|
|
93
159
|
}
|
|
94
160
|
]
|
|
95
161
|
```
|
|
96
162
|
|
|
97
|
-
###
|
|
163
|
+
### Fields
|
|
98
164
|
|
|
99
165
|
| Field | Type | Required | Description |
|
|
100
|
-
|
|
101
|
-
| `sample_id` | `string` | **
|
|
102
|
-
| `prompt` | `string` | **
|
|
103
|
-
| `context` | `string` |
|
|
104
|
-
| `rubric` | `string` |
|
|
105
|
-
| `assertions` | `array` |
|
|
106
|
-
| `assertions[].type` | `string` | **
|
|
107
|
-
| `assertions[].value` | `string\|number` |
|
|
108
|
-
| `assertions[].values` | `array` |
|
|
109
|
-
| `assertions[].pattern` | `string` |
|
|
110
|
-
| `assertions[].flags` | `string` |
|
|
111
|
-
| `assertions[].schema` | `object` |
|
|
112
|
-
| `assertions[].reference` | `string` |
|
|
113
|
-
| `assertions[].threshold` | `number` |
|
|
114
|
-
| `assertions[].fn` | `string` |
|
|
115
|
-
| `assertions[].weight` | `number` |
|
|
116
|
-
| `dimensions` | `object` |
|
|
166
|
+
|---|---|---|---|
|
|
167
|
+
| `sample_id` | `string` | **yes** | Unique sample ID |
|
|
168
|
+
| `prompt` | `string` | **yes** | User prompt sent to the model |
|
|
169
|
+
| `context` | `string` | no | Extra context (e.g. code). Wrapped in a code block and appended to the prompt. URLs are auto-fetched at runtime. |
|
|
170
|
+
| `rubric` | `string` | no | Scoring guideline for the LLM judge (1-5 scale) |
|
|
171
|
+
| `assertions` | `array` | no | Assertion checks; see [assertion types](#assertion-types) |
|
|
172
|
+
| `assertions[].type` | `string` | **yes** | Assertion type |
|
|
173
|
+
| `assertions[].value` | `string\|number` | depends | Check value (required for `contains`, `min_length`, `cost_max`, etc.) |
|
|
174
|
+
| `assertions[].values` | `array` | depends | String array (required for `contains_all`, `contains_any`) |
|
|
175
|
+
| `assertions[].pattern` | `string` | depends | Regex pattern (required for `regex`) |
|
|
176
|
+
| `assertions[].flags` | `string` | no | Regex flags (default `"i"`) |
|
|
177
|
+
| `assertions[].schema` | `object` | depends | JSON Schema object (required for `json_schema`, via [ajv](https://ajv.js.org/)) |
|
|
178
|
+
| `assertions[].reference` | `string` | depends | Reference text (required for `semantic_similarity`) |
|
|
179
|
+
| `assertions[].threshold` | `number` | no | Pass threshold for semantic similarity (default 3) |
|
|
180
|
+
| `assertions[].fn` | `string` | depends | Path to a custom assertion JS file (required for `custom`) |
|
|
181
|
+
| `assertions[].weight` | `number` | no | Weight (default 1) |
|
|
182
|
+
| `dimensions` | `object` | no | Multi-dimension scoring; key = dimension name, value = scoring guideline |
|
|
183
|
+
|
|
184
|
+
### URL auto-fetching
|
|
185
|
+
|
|
186
|
+
URLs in `prompt` and `context` are auto-fetched before evaluation and inlined into the text. Useful when referencing online docs, API references, etc.:
|
|
117
187
|
|
|
118
|
-
|
|
188
|
+
```json
|
|
189
|
+
{
|
|
190
|
+
"sample_id": "s001",
|
|
191
|
+
"prompt": "Generate test cases from this PRD: https://wiki.example.com/prd/feature-x"
|
|
192
|
+
}
|
|
193
|
+
```
|
|
194
|
+
|
|
195
|
+
At runtime, URLs are replaced with the actual content. Fetch order: MCP Server first for matching URLs (e.g. SSO-protected private docs), then plain HTTP for the rest. URLs already resolved by MCP are not re-fetched via HTTP.
|
|
196
|
+
|
|
197
|
+
**Private-doc URLs**: drop a `.mcp.json` config file into the project dir, or pass `--mcp-config <path>`:
|
|
198
|
+
|
|
199
|
+
```json
|
|
200
|
+
{
|
|
201
|
+
"mcpServers": {
|
|
202
|
+
"docs": {
|
|
203
|
+
"command": "npx",
|
|
204
|
+
"args": ["@example/docs-mcp-server"],
|
|
205
|
+
"env": { "DOCS_API_TOKEN": "xxx" },
|
|
206
|
+
"urlPatterns": ["docs.example.com"],
|
|
207
|
+
"fetchTool": {
|
|
208
|
+
"name": "fetch_doc",
|
|
209
|
+
"urlTransform": {
|
|
210
|
+
"regex": "docs\\.example\\.com/([^/]+/[^/]+)/([^/?#]+)",
|
|
211
|
+
"params": { "namespace": "$1", "slug": "$2" }
|
|
212
|
+
},
|
|
213
|
+
"contentExtract": "data.body"
|
|
214
|
+
}
|
|
215
|
+
}
|
|
216
|
+
}
|
|
217
|
+
}
|
|
218
|
+
```
|
|
119
219
|
|
|
120
|
-
**
|
|
220
|
+
**Public URLs**: fetched via plain HTTP. If they require auth, make sure the shell already has network access configured (VPN, proxy, etc.).
|
|
121
221
|
|
|
122
|
-
###
|
|
222
|
+
### Scoring strategy
|
|
123
223
|
|
|
124
|
-
|
|
224
|
+
#### 1. Assertion score
|
|
125
225
|
|
|
126
|
-
|
|
226
|
+
Rule-based local checks; each assertion yields pass/fail.
|
|
127
227
|
|
|
128
|
-
|
|
228
|
+
**Formula:**
|
|
129
229
|
|
|
130
|
-
|
|
230
|
+
- Pass rate = sum of passed assertion weights / total weight (0–1)
|
|
231
|
+
- Score = 1 + pass_rate × 4 (mapped to 1–5)
|
|
232
|
+
- Example: 3 assertions (weight 1 each), 2 pass → pass rate 2/3 → score = 1 + 0.67 × 4 = **3.67**
|
|
131
233
|
|
|
132
|
-
|
|
133
|
-
2. Sum the weights of all passing assertions → `passedWeight`
|
|
134
|
-
3. Sum the weights of all assertions → `totalWeight`
|
|
135
|
-
4. Compute ratio: `passedWeight / totalWeight` (0.0 ~ 1.0)
|
|
136
|
-
5. Normalize to 1-5 scale: **`score = 1 + ratio × 4`**
|
|
234
|
+
#### 2. Rubric / Dimensions score
|
|
137
235
|
|
|
138
|
-
|
|
236
|
+
The judge model (default `haiku`) scores 1–5 against the rubric. In `dimensions` mode, each dimension is scored independently and then averaged.
|
|
139
237
|
|
|
140
|
-
####
|
|
238
|
+
#### 3. Composite score
|
|
141
239
|
|
|
142
|
-
|
|
240
|
+
| Condition | Formula |
|
|
241
|
+
|---|---|
|
|
242
|
+
| Only assertions | `assertionScore` |
|
|
243
|
+
| Only LLM judge | `llmScore` |
|
|
244
|
+
| Both present | `(assertionScore + llmScore) / 2` |
|
|
245
|
+
| Neither | `0` |
|
|
143
246
|
|
|
144
|
-
|
|
247
|
+
### Assertion types
|
|
145
248
|
|
|
146
|
-
|
|
249
|
+
**Deterministic assertions (18 total):**
|
|
147
250
|
|
|
148
|
-
|
|
251
|
+
| Type | Description |
|
|
252
|
+
|---|---|
|
|
253
|
+
| `contains` / `not_contains` | substring must / must-not appear |
|
|
254
|
+
| `regex` | regex match |
|
|
255
|
+
| `min_length` / `max_length` | length bounds |
|
|
256
|
+
| `json_valid` / `json_schema` | JSON validation |
|
|
257
|
+
| `starts_with` / `ends_with` | prefix / suffix |
|
|
258
|
+
| `equals` / `not_equals` | exact match |
|
|
259
|
+
| `word_count_min` / `word_count_max` | word-count bounds |
|
|
260
|
+
| `contains_all` / `contains_any` | multi-value match |
|
|
261
|
+
| `cost_max` / `latency_max` | cost / latency caps |
|
|
262
|
+
| `semantic_similarity` | LLM-based semantic similarity |
|
|
263
|
+
| `custom` | custom JS function (30 s timeout) |
|
|
149
264
|
|
|
150
|
-
|
|
265
|
+
### Custom assertion
|
|
151
266
|
|
|
152
|
-
|
|
267
|
+
```js
|
|
268
|
+
// my-assertion.mjs
|
|
269
|
+
export default function(output, { sample, assertion }) {
|
|
270
|
+
return { pass: output.includes('SQL'), message: 'checked for SQL keyword' };
|
|
271
|
+
}
|
|
272
|
+
```
|
|
153
273
|
|
|
154
|
-
|
|
155
|
-
|----------------|----------------------|
|
|
156
|
-
| Assertions only | `assertionScore` |
|
|
157
|
-
| LLM only (rubric or dimensions) | `llmScore` |
|
|
158
|
-
| Both | `(assertionScore + llmScore) / 2` |
|
|
159
|
-
| Neither | `0` |
|
|
274
|
+
## Six-dim evaluation
|
|
160
275
|
|
|
161
|
-
|
|
276
|
+
Reports display results across six independent dimensions. The three scoring layers — Fact / Behavior / LLM-judge — are shown separately so you see **which layer regressed** instead of a single composite number:
|
|
162
277
|
|
|
163
|
-
|
|
278
|
+
| Dimension | Metric | Description |
|
|
279
|
+
|---|---|---|
|
|
280
|
+
| 📋 **Fact** | fact-assertion pass rate | rule-verifiable assertions like `contains` / `json_schema` / `fact_check`, mapped to 1-5 |
|
|
281
|
+
| 🛠️ **Behavior** | behavior-assertion pass rate | execution-compliance assertions like `tools_called` / `tool_output_contains` / `turns_max` |
|
|
282
|
+
| 💬 **LLM-judge** | rubric score | 1-5 scored by the judge model against a predefined rubric; subjective, catches what rules miss |
|
|
283
|
+
| 💰 **Cost** | total cost, input/output tokens | API cost based on token usage and model pricing |
|
|
284
|
+
| ⚡ **Efficiency** | average latency (ms) | end-to-end latency from request to full response |
|
|
285
|
+
| 🛡️ **Stability** | CV (coefficient of variation) | score consistency across repeated runs (`--repeat ≥ 2`); single-run shows `—`, **honestly acknowledging what can't be measured** |
|
|
164
286
|
|
|
165
|
-
|
|
287
|
+
## CLI reference
|
|
166
288
|
|
|
167
|
-
|
|
168
|
-
|------|--------|-------------|
|
|
169
|
-
| `contains` | `value`, `weight` | Output contains substring (case-insensitive) |
|
|
170
|
-
| `not_contains` | `value`, `weight` | Output does NOT contain substring |
|
|
171
|
-
| `regex` | `pattern`, `flags`, `weight` | Output matches regex |
|
|
172
|
-
| `min_length` | `value`, `weight` | Output length >= value |
|
|
173
|
-
| `max_length` | `value`, `weight` | Output length <= value |
|
|
174
|
-
| `json_valid` | `weight` | Output is valid JSON |
|
|
175
|
-
| `json_schema` | `schema`, `weight` | Output matches JSON Schema (full spec via ajv) |
|
|
176
|
-
| `starts_with` | `value`, `weight` | Output starts with string (case-insensitive) |
|
|
177
|
-
| `ends_with` | `value`, `weight` | Output ends with string (case-insensitive) |
|
|
178
|
-
| `equals` | `value`, `weight` | Output exactly equals value (after trim) |
|
|
179
|
-
| `not_equals` | `value`, `weight` | Output does not equal value (after trim) |
|
|
180
|
-
| `word_count_min` | `value`, `weight` | Word count >= value |
|
|
181
|
-
| `word_count_max` | `value`, `weight` | Word count <= value |
|
|
182
|
-
| `contains_all` | `values`, `weight` | Output contains ALL substrings |
|
|
183
|
-
| `contains_any` | `values`, `weight` | Output contains at least one substring |
|
|
184
|
-
| `cost_max` | `value`, `weight` | Execution cost (USD) <= value |
|
|
185
|
-
| `latency_max` | `value`, `weight` | Execution latency (ms) <= value |
|
|
289
|
+
### `omk bench run`
|
|
186
290
|
|
|
187
|
-
|
|
291
|
+
```bash
|
|
292
|
+
omk bench run [options]
|
|
188
293
|
|
|
189
|
-
|
|
190
|
-
|
|
191
|
-
|
|
192
|
-
|
|
294
|
+
options:
|
|
295
|
+
--samples <path> sample file (default: eval-samples.json, also detects .yaml/.yml)
|
|
296
|
+
--skill-dir <path> artifact dir (default: skills)
|
|
297
|
+
--control <expr> control-group variant expression (experiment role = control)
|
|
298
|
+
--treatment <v1,v2> treatment-group variant expressions, comma-separated
|
|
299
|
+
at least one of --control / --treatment is required
|
|
300
|
+
(unless you use --config or --each)
|
|
301
|
+
special values: baseline (empty artifact), git:name (git HEAD),
|
|
302
|
+
git:ref:name (specific commit), path with "/" (read file directly)
|
|
303
|
+
--config <path> YAML/JSON config file (evaluation-as-code); declares
|
|
304
|
+
samples + variants + model + executor in one file; CLI
|
|
305
|
+
flags override config fields when both are provided
|
|
306
|
+
--model <name> model under test (default: sonnet)
|
|
307
|
+
--judge-model <name> judge model (default: haiku)
|
|
308
|
+
--output-dir <path> output dir (default: ~/.oh-my-knowledge/reports/)
|
|
309
|
+
--no-judge skip the LLM judge
|
|
310
|
+
--no-cache disable result cache (on by default; identical inputs reuse)
|
|
311
|
+
--dry-run preview only
|
|
312
|
+
--blind blind mode
|
|
313
|
+
--concurrency <n> parallel tasks (default: 1)
|
|
314
|
+
--timeout <sec> per-task executor timeout (default: 120)
|
|
315
|
+
--repeat <n> repeat N times for variance analysis (default: 1)
|
|
316
|
+
--executor <name> executor (default: claude); supports custom commands
|
|
317
|
+
--skip-preflight skip pre-evaluation model reachability check
|
|
318
|
+
--mcp-config <path> MCP config for fetching private-doc URLs via MCP Server
|
|
319
|
+
(default: .mcp.json in cwd)
|
|
320
|
+
--no-serve don't auto-start the report server after the run
|
|
321
|
+
--verbose print per-sample details (duration, tokens, output preview)
|
|
322
|
+
--each batch mode: evaluate each artifact independently vs baseline
|
|
323
|
+
requires {name}.eval-samples.json paired with each artifact
|
|
324
|
+
```
|
|
193
325
|
|
|
194
|
-
###
|
|
326
|
+
### `omk bench run --each` (batch mode)
|
|
195
327
|
|
|
196
|
-
|
|
328
|
+
When `skills/` contains several **independent** artifacts, use `--each` to evaluate each one against baseline and produce a merged report.
|
|
197
329
|
|
|
198
|
-
```
|
|
199
|
-
|
|
200
|
-
|
|
201
|
-
|
|
202
|
-
|
|
203
|
-
|
|
330
|
+
```
|
|
331
|
+
skills/
|
|
332
|
+
├── asset.md ← artifact file
|
|
333
|
+
├── asset.eval-samples.json ← paired samples
|
|
334
|
+
├── home.md
|
|
335
|
+
├── home.eval-samples.json
|
|
336
|
+
└── product/ ← directory format also supported
|
|
337
|
+
├── SKILL.md
|
|
338
|
+
└── eval-samples.json
|
|
204
339
|
```
|
|
205
340
|
|
|
206
|
-
|
|
341
|
+
Pairing rules:
|
|
207
342
|
|
|
208
|
-
|
|
343
|
+
- `{name}.md` → looks for `{name}.eval-samples.json` in the same dir
|
|
344
|
+
- `{name}/SKILL.md` → looks for `{name}/eval-samples.json`
|
|
345
|
+
- artifacts without paired samples are skipped with a warning
|
|
209
346
|
|
|
210
|
-
|
|
347
|
+
```bash
|
|
348
|
+
omk bench run --each
|
|
349
|
+
omk bench run --each --dry-run
|
|
350
|
+
```
|
|
211
351
|
|
|
212
|
-
|
|
213
|
-
|
|
214
|
-
|
|
215
|
-
|
|
216
|
-
|
|
217
|
-
|
|
218
|
-
|
|
219
|
-
|
|
220
|
-
|
|
221
|
-
|
|
222
|
-
|
|
223
|
-
|
|
224
|
-
|
|
225
|
-
"nodeVersion": "v22.0.0",
|
|
226
|
-
"skillHashes": { "v1": "a1b2c3d4e5f6", "v2": "f6e5d4c3b2a1" }
|
|
227
|
-
},
|
|
228
|
-
"summary": {
|
|
229
|
-
"v1": {
|
|
230
|
-
"totalSamples": 3,
|
|
231
|
-
"successCount": 3,
|
|
232
|
-
"errorCount": 0,
|
|
233
|
-
"avgCompositeScore": 3.67,
|
|
234
|
-
"avgAssertionScore": 3.0,
|
|
235
|
-
"avgLlmScore": 4.33,
|
|
236
|
-
"avgDurationMs": 2500,
|
|
237
|
-
"avgTotalTokens": 1850,
|
|
238
|
-
"totalCostUSD": 0.0112
|
|
239
|
-
},
|
|
240
|
-
"v2": {
|
|
241
|
-
"totalSamples": 3,
|
|
242
|
-
"successCount": 3,
|
|
243
|
-
"errorCount": 0,
|
|
244
|
-
"avgCompositeScore": 4.5,
|
|
245
|
-
"avgAssertionScore": 5.0,
|
|
246
|
-
"avgLlmScore": 4.0,
|
|
247
|
-
"avgDurationMs": 2800,
|
|
248
|
-
"avgTotalTokens": 2100,
|
|
249
|
-
"totalCostUSD": 0.0122
|
|
250
|
-
}
|
|
251
|
-
},
|
|
252
|
-
"results": [
|
|
253
|
-
{
|
|
254
|
-
"sample_id": "s001",
|
|
255
|
-
"variants": {
|
|
256
|
-
"v1": {
|
|
257
|
-
"ok": true,
|
|
258
|
-
"compositeScore": 3.5,
|
|
259
|
-
"assertions": {
|
|
260
|
-
"passed": 1,
|
|
261
|
-
"total": 2,
|
|
262
|
-
"score": 3.0,
|
|
263
|
-
"details": [
|
|
264
|
-
{ "type": "contains", "value": "SQL injection", "weight": 1, "passed": true },
|
|
265
|
-
{ "type": "contains", "value": "parameterized", "weight": 1, "passed": false }
|
|
266
|
-
]
|
|
267
|
-
},
|
|
268
|
-
"llmScore": 4,
|
|
269
|
-
"llmReason": "Identified the vulnerability but did not provide a complete fix",
|
|
270
|
-
"durationMs": 2300,
|
|
271
|
-
"inputTokens": 850,
|
|
272
|
-
"outputTokens": 1200,
|
|
273
|
-
"totalTokens": 2050,
|
|
274
|
-
"costUSD": 0.0038,
|
|
275
|
-
"outputPreview": "This code has a SQL injection vulnerability..."
|
|
276
|
-
},
|
|
277
|
-
"v2": {
|
|
278
|
-
"ok": true,
|
|
279
|
-
"compositeScore": 4.5,
|
|
280
|
-
"assertions": {
|
|
281
|
-
"passed": 2,
|
|
282
|
-
"total": 2,
|
|
283
|
-
"score": 5.0,
|
|
284
|
-
"details": [
|
|
285
|
-
{ "type": "contains", "value": "SQL injection", "weight": 1, "passed": true },
|
|
286
|
-
{ "type": "contains", "value": "parameterized", "weight": 1, "passed": true }
|
|
287
|
-
]
|
|
288
|
-
},
|
|
289
|
-
"llmScore": 4,
|
|
290
|
-
"llmReason": "Thorough analysis with actionable fix code",
|
|
291
|
-
"durationMs": 2600,
|
|
292
|
-
"inputTokens": 900,
|
|
293
|
-
"outputTokens": 1400,
|
|
294
|
-
"totalTokens": 2300,
|
|
295
|
-
"costUSD": 0.0042,
|
|
296
|
-
"outputPreview": "## Security Issue: SQL Injection\n\nThe code is vulnerable..."
|
|
297
|
-
}
|
|
298
|
-
}
|
|
299
|
-
}
|
|
300
|
-
],
|
|
301
|
-
"analysis": {
|
|
302
|
-
"insights": [
|
|
303
|
-
{
|
|
304
|
-
"type": "uniform_scores",
|
|
305
|
-
"severity": "info",
|
|
306
|
-
"message": "1/3 samples show score difference < 0.5 between variants"
|
|
307
|
-
}
|
|
308
|
-
],
|
|
309
|
-
"suggestions": []
|
|
310
|
-
}
|
|
311
|
-
}
|
|
352
|
+
### `omk bench gen-samples` (generate test cases)
|
|
353
|
+
|
|
354
|
+
Reads an artifact's content and uses an LLM to auto-generate eval-samples. Review and edit them before running eval.
|
|
355
|
+
|
|
356
|
+
```bash
|
|
357
|
+
# generate for a specific artifact (writes eval-samples.json)
|
|
358
|
+
omk bench gen-samples skills/my-skill.md
|
|
359
|
+
|
|
360
|
+
# batch-generate for every artifact under skills/ that lacks samples
|
|
361
|
+
omk bench gen-samples --each
|
|
362
|
+
|
|
363
|
+
# specify sample count
|
|
364
|
+
omk bench gen-samples skills/my-skill.md --count 10
|
|
312
365
|
```
|
|
313
366
|
|
|
314
|
-
|
|
367
|
+
Options:
|
|
315
368
|
|
|
316
|
-
|
|
317
|
-
-
|
|
318
|
-
|
|
319
|
-
|
|
369
|
+
```
|
|
370
|
+
--each batch-generate for every artifact missing samples
|
|
371
|
+
--count <n> samples per artifact (default: 5)
|
|
372
|
+
--model <name> model used for generation (default: sonnet)
|
|
373
|
+
--skill-dir <path> artifact dir (default: skills), used with --each
|
|
374
|
+
```
|
|
320
375
|
|
|
321
|
-
|
|
376
|
+
### `omk bench evolve` (self-iterating improvement)
|
|
322
377
|
|
|
323
|
-
|
|
378
|
+
Lets the AI iterate an artifact automatically: evaluate → analyze weak spots → LLM rewrites → evaluate again → keep if the score went up, drop otherwise → repeat.
|
|
324
379
|
|
|
325
380
|
```bash
|
|
326
|
-
|
|
381
|
+
# basic: iterate 5 rounds
|
|
382
|
+
omk bench evolve skills/my-skill.md
|
|
383
|
+
|
|
384
|
+
# set rounds and target score
|
|
385
|
+
omk bench evolve skills/my-skill.md --rounds 10 --target 4.5
|
|
386
|
+
```
|
|
327
387
|
|
|
328
388
|
Options:
|
|
329
|
-
|
|
330
|
-
--skill-dir <path> Skill directory (default: skills)
|
|
331
|
-
--variants <v1,v2> Variant names (default: v1,v2)
|
|
332
|
-
--model <name> Model under test (default: sonnet)
|
|
333
|
-
--judge-model <name> Judge model (default: haiku)
|
|
334
|
-
--output-dir <path> Output directory (default: ~/.oh-my-knowledge/reports/)
|
|
335
|
-
--no-judge Skip LLM judging
|
|
336
|
-
--dry-run Preview only
|
|
337
|
-
--blind Blind A/B mode: hide variant names in report
|
|
338
|
-
--concurrency <n> Number of parallel tasks (default: 1)
|
|
339
|
-
--repeat <n> Run evaluation N times for variance analysis (default: 1)
|
|
340
|
-
--executor <name> Executor (default: claude)
|
|
389
|
+
|
|
341
390
|
```
|
|
391
|
+
--rounds <n> max iteration rounds (default: 5)
|
|
392
|
+
--target <score> stop early when the score reaches this threshold
|
|
393
|
+
--samples <path> sample file (default: eval-samples.json)
|
|
394
|
+
--improve-model <name> model used for rewrites (default: sonnet)
|
|
395
|
+
```
|
|
396
|
+
|
|
397
|
+
Each round's output is saved under `skills/evolve/` (`my-skill.r0.md`, `my-skill.r1.md`…), so you can `diff` to see what the AI changed. The best round is written back to the original file.
|
|
342
398
|
|
|
343
399
|
### `omk bench ci`
|
|
344
400
|
|
|
345
|
-
Run evaluation
|
|
401
|
+
Run the evaluation inside CI. Exit code 0 on pass, 1 on fail — can be wired into gates directly.
|
|
402
|
+
|
|
403
|
+
The gate is **three-layer all-pass**: `avgFactScore >= threshold AND avgBehaviorScore >= threshold AND avgJudgeScore >= threshold`. Any layer below threshold is FAIL, and the output shows which layer broke. This stops cases like `fact 4.5→2.5 but judge 3→5` from passing via composite averaging — if one layer regresses, the gate catches it.
|
|
346
404
|
|
|
347
405
|
```bash
|
|
348
406
|
omk bench ci [options]
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
(same as "bench run", plus:)
|
|
352
|
-
--threshold <number> Minimum composite score to pass (default: 3.5)
|
|
407
|
+
--threshold <number> per-layer minimum score (default: 3.5); applied
|
|
408
|
+
independently to fact / behavior / judge
|
|
353
409
|
```
|
|
354
410
|
|
|
355
|
-
Exit code 0 = all variants pass, 1 = at least one variant below threshold.
|
|
356
|
-
|
|
357
411
|
### `omk bench report`
|
|
358
412
|
|
|
413
|
+
Start the report server to browse historical reports, submit feedback, and delete reports.
|
|
414
|
+
|
|
359
415
|
```bash
|
|
360
416
|
omk bench report [options]
|
|
361
|
-
|
|
362
|
-
Options:
|
|
363
|
-
--port <number> Server port (default: 7799)
|
|
364
|
-
--reports-dir <path> Reports directory (default: ~/.oh-my-knowledge/reports/)
|
|
417
|
+
--port <number> server port (default: 7799)
|
|
365
418
|
```
|
|
366
419
|
|
|
367
420
|
### `omk bench init`
|
|
368
421
|
|
|
369
422
|
```bash
|
|
370
|
-
omk bench init [dir]
|
|
423
|
+
omk bench init [dir] # scaffold an eval project
|
|
371
424
|
```
|
|
372
425
|
|
|
373
|
-
##
|
|
426
|
+
## `omk analyze` — production observability
|
|
374
427
|
|
|
375
|
-
|
|
428
|
+
`omk bench run` is **offline evaluation** (fixed controls, repeatable, scored). Production is different — no control group, no ground truth, no repetition, so scoring isn't valid there. `omk analyze` turns existing Claude Code session traces into **skill-health reports** (coverage, gap signals, execution stability, tokens/latency per skill). It gives you clues about **which skill is worth re-evaluating offline**, not a production score.
|
|
376
429
|
|
|
377
|
-
|
|
430
|
+
```bash
|
|
431
|
+
# analyze all cc sessions of the current project (auto-infers kb from the trace)
|
|
432
|
+
omk analyze ~/.claude/projects/-Users-you-Documents-my-project
|
|
378
433
|
|
|
379
|
-
|
|
434
|
+
# restrict to the last 7 days / 24 hours / 30 minutes
|
|
435
|
+
omk analyze ~/.claude/projects/my-project --last 7d
|
|
380
436
|
|
|
381
|
-
|
|
437
|
+
# absolute time window
|
|
438
|
+
omk analyze ~/.claude/projects/my-project --from 2026-04-01T00:00:00Z --to 2026-04-15T23:59:59Z
|
|
382
439
|
|
|
383
|
-
|
|
440
|
+
# whitelist specific skills
|
|
441
|
+
omk analyze ~/.claude/projects/my-project --skills audit,polish
|
|
384
442
|
|
|
385
|
-
|
|
386
|
-
|
|
387
|
-
|
|
443
|
+
# override the inferred knowledge-base root
|
|
444
|
+
omk analyze ~/.claude/projects/my-project --kb /path/to/project
|
|
445
|
+
```
|
|
388
446
|
|
|
389
|
-
|
|
447
|
+
The command writes `~/.oh-my-knowledge/analyses/<timestamp>-skill-health.json`. Browse results alongside bench reports with `omk bench report` — the homepage has a "📊 Skill Health Reports" link, and each skill card also has a "trend →" link to its time-series view. For two reports side-by-side, use the compare selector on `/analyses`.
|
|
390
448
|
|
|
391
|
-
|
|
392
|
-
- **Low-discrimination assertions**: assertions with identical results across all variants
|
|
393
|
-
- **Uniform scores**: samples where variants score within 0.5 of each other
|
|
394
|
-
- **All-pass / all-fail**: assertions that may be too loose or too strict
|
|
395
|
-
- **High-cost samples**: samples with disproportionately high cost
|
|
449
|
+
**What you get per skill:**
|
|
396
450
|
|
|
397
|
-
|
|
451
|
+
- **Knowledge usage** — which KB files this skill actually read (coverage %)
|
|
452
|
+
- **Knowledge gaps** — four weighted signals (failed search / model-flagged gap / hedging / repeated miss); hedging goes through an LLM-assisted classifier to filter out business-possibility hedging vs genuine knowledge uncertainty
|
|
453
|
+
- **Execution stability** — tool-failure rate; a skill with > 20% failures gets a warning that its gap signals may be environmental noise rather than real knowledge gaps
|
|
454
|
+
- **Usage cost** — billable tokens (input+output) separate from cached tokens, total duration
|
|
398
455
|
|
|
399
|
-
|
|
456
|
+
**What this is NOT:**
|
|
400
457
|
|
|
401
|
-
|
|
458
|
+
- Not a general APM (request/response/latency tracing is Langfuse / Datadog territory)
|
|
459
|
+
- Not streaming / alerting (batch only — run on a cron if you want periodic snapshots)
|
|
460
|
+
- Not a production score (no control group, no ground truth — use `omk bench run` for scoring)
|
|
402
461
|
|
|
403
|
-
|
|
462
|
+
## Executors
|
|
404
463
|
|
|
405
|
-
|
|
464
|
+
### Built-in executors
|
|
406
465
|
|
|
407
|
-
|
|
466
|
+
| Executor | When to use | Description |
|
|
467
|
+
|---|---|---|
|
|
468
|
+
| `claude` | default | invokes `claude -p` via Claude CLI |
|
|
469
|
+
| `claude-sdk` | structured output | uses Claude Agent SDK — no stdout parsing, avoids buffer truncation |
|
|
470
|
+
| `openai` | cross-vendor comparison | invokes `openai api` CLI |
|
|
471
|
+
| `gemini` | cross-vendor comparison | invokes `gemini` CLI |
|
|
472
|
+
| `anthropic-api` | no CLI needed | calls Anthropic HTTP API directly (needs `ANTHROPIC_API_KEY`) |
|
|
473
|
+
| `openai-api` | no CLI needed | calls OpenAI HTTP API directly (needs `OPENAI_API_KEY`) |
|
|
474
|
+
|
|
475
|
+
API-direct executors support custom base URLs via env: `ANTHROPIC_BASE_URL`, `OPENAI_BASE_URL`.
|
|
408
476
|
|
|
409
|
-
|
|
477
|
+
### Custom executor
|
|
410
478
|
|
|
411
|
-
|
|
412
|
-
|----------|----------|---------------|------|
|
|
413
|
-
| `claude` | `claude -p` | `sonnet` | Claude Max plan or API key |
|
|
414
|
-
| `openai` | `openai api chat.completions.create` | `gpt-4o` | `OPENAI_API_KEY` env var |
|
|
415
|
-
| `gemini` | `gemini` (stdin pipe) | Default Gemini model | Google account or `GOOGLE_API_KEY` |
|
|
479
|
+
Any shell command can serve as an executor, communicating via stdin/stdout JSON:
|
|
416
480
|
|
|
417
481
|
```bash
|
|
418
|
-
|
|
419
|
-
omk bench run --executor
|
|
482
|
+
omk bench run --executor "python my_provider.py"
|
|
483
|
+
omk bench run --executor "./my-executor.sh"
|
|
484
|
+
```
|
|
485
|
+
|
|
486
|
+
**Protocol:**
|
|
487
|
+
|
|
488
|
+
- **input** (stdin): JSON `{"model":"...","system":"...","prompt":"..."}`
|
|
489
|
+
- **output** (stdout): JSON `{"output":"model reply","inputTokens":0,"outputTokens":0,"costUSD":0}`
|
|
490
|
+
- stdout only needs to return the fields you care about; others default to 0. Plain-text output (no tokens/cost parsing) is also fine.
|
|
491
|
+
- non-zero exit code counts as failure
|
|
420
492
|
|
|
421
|
-
|
|
422
|
-
omk bench run --executor gemini --model gemini-2.5-pro --variants v1,v2
|
|
493
|
+
### Artifact directory layout
|
|
423
494
|
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
|
|
495
|
+
The built-in executors (claude / openai / gemini) support two artifact layouts, mixable in the same run:
|
|
496
|
+
|
|
497
|
+
```
|
|
498
|
+
skills/
|
|
499
|
+
├── v1.md # option 1: plain .md file
|
|
500
|
+
└── my-skill/ # option 2: full artifact dir
|
|
501
|
+
├── SKILL.md # this file is auto-loaded as system prompt
|
|
502
|
+
├── config.json # other files don't participate in eval, kept for completeness
|
|
503
|
+
└── scripts/
|
|
504
|
+
```
|
|
505
|
+
|
|
506
|
+
**Variant resolution rules:**
|
|
507
|
+
|
|
508
|
+
`variant` is the experiment-group expression. After resolution, OMK produces an `artifact` plus an optional `runtime context` (currently mainly `cwd`).
|
|
509
|
+
|
|
510
|
+
| Format | Meaning |
|
|
511
|
+
|---|---|
|
|
512
|
+
| `name` | looks up `name.md` or `name/SKILL.md` in the artifact dir, resolves to one artifact |
|
|
513
|
+
| `baseline` | empty artifact, no system prompt — think "nothing at all" |
|
|
514
|
+
| `project-env@/path/to/project` | empty artifact, but run in the specified project dir — observe project-level runtime context alone |
|
|
515
|
+
| `git:name` | reads the last-committed version of an artifact from git HEAD |
|
|
516
|
+
| `git:ref:name` | reads an artifact from a specific commit |
|
|
517
|
+
| `./path/to/file.md` | path with `/`: read the file directly as an artifact |
|
|
518
|
+
| `variant@/path/to/project` | attach a run dir to any variant; supports `name@cwd`, `git:name@cwd`, `/file.md@cwd` |
|
|
519
|
+
|
|
520
|
+
When both `--control` and `--treatment` are omitted, use `--config eval.yaml` or `--each`. With `--each`, `baseline` is auto-added as control and every discovered artifact becomes a treatment.
|
|
521
|
+
|
|
522
|
+
```bash
|
|
523
|
+
# explicit: one control, one or more treatments
|
|
524
|
+
omk bench run --control v1 --treatment v2
|
|
525
|
+
omk bench run --control baseline --treatment v1,v2,v3
|
|
526
|
+
|
|
527
|
+
# compare empty artifact vs explicit artifact
|
|
528
|
+
omk bench run --control baseline --treatment my-skill
|
|
529
|
+
|
|
530
|
+
# observe project-level runtime context in isolation (use a self-describing label)
|
|
531
|
+
omk bench run --control baseline --treatment project-env@/path/to/target-project
|
|
532
|
+
|
|
533
|
+
# compare "project-level runtime context" vs "explicit artifact injection"
|
|
534
|
+
omk bench run \
|
|
535
|
+
--control project-env@/path/to/target-project \
|
|
536
|
+
--treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
|
|
537
|
+
|
|
538
|
+
# before vs after (old version read from git history)
|
|
539
|
+
omk bench run --control git:my-skill --treatment my-skill
|
|
540
|
+
|
|
541
|
+
# direct file paths
|
|
542
|
+
omk bench run --control ./old-skill.md --treatment ./new-skill.md
|
|
543
|
+
|
|
544
|
+
# config-file driven (evaluation-as-code)
|
|
545
|
+
omk bench run --config eval.yaml
|
|
427
546
|
```
|
|
428
547
|
|
|
429
548
|
**Prerequisites:**
|
|
430
|
-
|
|
549
|
+
|
|
550
|
+
- **claude**: install [Claude Code](https://claude.ai/code) and authenticate
|
|
551
|
+
- **claude-sdk**: install [Claude Code](https://claude.ai/code) and authenticate (uses Agent SDK, no CLI stdout parsing)
|
|
552
|
+
- **anthropic-api**: set the `ANTHROPIC_API_KEY` env var
|
|
431
553
|
- **openai**: `pip install openai` and set `OPENAI_API_KEY`
|
|
432
|
-
- **
|
|
554
|
+
- **openai-api**: set the `OPENAI_API_KEY` env var
|
|
555
|
+
- **gemini**: `npm i -g @google/gemini-cli` and authenticate
|
|
556
|
+
|
|
557
|
+
### Agent evaluation and project-level runtime context
|
|
558
|
+
|
|
559
|
+
When the executor is `claude-sdk`, OMK supports a first pass of agent-aware evaluation.
|
|
560
|
+
|
|
561
|
+
A few concepts worth keeping separate:
|
|
562
|
+
|
|
563
|
+
- `artifact`: the thing being evaluated — baseline, skill, prompt, agent
|
|
564
|
+
- `variant`: the CLI expression for an experiment group
|
|
565
|
+
- `runtime context`: the runtime environment; currently mainly `cwd`. In project-type agent scenarios it includes the project dir, its `CLAUDE.md`, local skills, and any other environmental factors that affect behavior
|
|
566
|
+
|
|
567
|
+
In OMK, `agent` is not a catch-all term and neither is `skill`. A cleaner phrasing: **you are comparing how different artifacts behave under different runtime contexts.**
|
|
568
|
+
|
|
569
|
+
- auto-extracts turns / toolCalls traces
|
|
570
|
+
- supports assertions on tool-call behavior
|
|
571
|
+
- supports running under a specified `cwd`, so Claude Code auto-loads the project's `CLAUDE.md`, skills, and local runtime context
|
|
572
|
+
|
|
573
|
+
#### Recommended executor
|
|
574
|
+
|
|
575
|
+
```bash
|
|
576
|
+
omk bench run --executor claude-sdk
|
|
577
|
+
```
|
|
578
|
+
|
|
579
|
+
#### Agent-related assertions
|
|
580
|
+
|
|
581
|
+
| Assertion | Meaning |
|
|
582
|
+
|---|---|
|
|
583
|
+
| `tools_called` | must call the specified tool(s) |
|
|
584
|
+
| `tools_not_called` | must not call the specified tool(s) |
|
|
585
|
+
| `tools_count_min` / `tools_count_max` | tool-call-count bounds |
|
|
586
|
+
| `tool_output_contains` | output of a specific tool must contain given content |
|
|
587
|
+
| `turns_min` / `turns_max` | turn-count bounds |
|
|
588
|
+
|
|
589
|
+
#### Three common control setups
|
|
590
|
+
|
|
591
|
+
**1. Bare-model baseline**
|
|
433
592
|
|
|
434
|
-
|
|
593
|
+
No system prompt and no knowledge-carrying project dir. Requires at least one treatment to compare against:
|
|
594
|
+
|
|
595
|
+
```bash
|
|
596
|
+
omk bench run \
|
|
597
|
+
--executor claude-sdk \
|
|
598
|
+
--control baseline \
|
|
599
|
+
--treatment my-skill
|
|
600
|
+
```
|
|
601
|
+
|
|
602
|
+
**2. Empty artifact + project-level runtime context**
|
|
603
|
+
|
|
604
|
+
No system prompt, but runs inside a project dir. This is **not** a strict "bare baseline" — it is "empty artifact + project-level runtime context".
|
|
605
|
+
|
|
606
|
+
```bash
|
|
607
|
+
omk bench run \
|
|
608
|
+
--executor claude-sdk \
|
|
609
|
+
--control baseline \
|
|
610
|
+
--treatment project-env@/path/to/target-project
|
|
611
|
+
```
|
|
612
|
+
|
|
613
|
+
**3. Explicit artifact injection**
|
|
614
|
+
|
|
615
|
+
Inject an external `SKILL.md` as the artifact while also keeping the project dir. Good for contrasting "project-level runtime context" vs "explicit single-artifact injection".
|
|
616
|
+
|
|
617
|
+
```bash
|
|
618
|
+
omk bench run \
|
|
619
|
+
--executor claude-sdk \
|
|
620
|
+
--control project-env@/path/to/target-project \
|
|
621
|
+
--treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
|
|
622
|
+
```
|
|
623
|
+
|
|
624
|
+
#### Recommended first-round design
|
|
625
|
+
|
|
626
|
+
For PRD / complex business-knowledge scenarios, start with:
|
|
627
|
+
|
|
628
|
+
```bash
|
|
629
|
+
omk bench run \
|
|
630
|
+
--executor claude-sdk \
|
|
631
|
+
--samples skills/evaluate-review/eval-samples.yaml \
|
|
632
|
+
--control baseline \
|
|
633
|
+
--treatment /path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
|
|
634
|
+
```
|
|
635
|
+
|
|
636
|
+
If you want to prove whether "the knowledge sitting inside the project directory" is effective on its own, add a second treatment:
|
|
637
|
+
|
|
638
|
+
```bash
|
|
639
|
+
omk bench run \
|
|
640
|
+
--executor claude-sdk \
|
|
641
|
+
--samples skills/evaluate-review/eval-samples.yaml \
|
|
642
|
+
--control baseline \
|
|
643
|
+
--treatment project-env@/path/to/target-project,/path/to/target-project/.claude/skills/prd/SKILL.md@/path/to/target-project
|
|
644
|
+
```
|
|
645
|
+
|
|
646
|
+
#### Design tips
|
|
647
|
+
|
|
648
|
+
- **Always start with `--dry-run`** to confirm samples, variants, and `cwd` are parsed correctly
|
|
649
|
+
- **Project-level controls must differ in `cwd`**: the same prompt under different project dirs hits different runtime contexts
|
|
650
|
+
- **Try PRD scenarios first**: compared to pure coding, they make it easier to validate knowledge completeness, impact-area detection, and business correctness
|
|
651
|
+
|
|
652
|
+
### Common model configurations
|
|
653
|
+
|
|
654
|
+
**Don't have Claude?** Most Chinese LLMs (GLM, Qwen, Moonshot, DeepSeek, etc.) are OpenAI-API compatible — use the `openai-api` executor directly:
|
|
655
|
+
|
|
656
|
+
```bash
|
|
657
|
+
# GLM (Zhipu)
|
|
658
|
+
export OPENAI_API_KEY="your Zhipu API key"
|
|
659
|
+
export OPENAI_BASE_URL="https://open.bigmodel.cn/api/paas/v4"
|
|
660
|
+
omk bench run --executor openai-api --model glm-4-plus \
|
|
661
|
+
--judge-model glm-4-plus --no-cache
|
|
662
|
+
|
|
663
|
+
# Qwen (Alibaba)
|
|
664
|
+
export OPENAI_API_KEY="your Qwen API key"
|
|
665
|
+
export OPENAI_BASE_URL="https://dashscope.aliyuncs.com/compatible-mode/v1"
|
|
666
|
+
omk bench run --executor openai-api --model qwen-plus \
|
|
667
|
+
--judge-model qwen-plus
|
|
668
|
+
|
|
669
|
+
# DeepSeek
|
|
670
|
+
export OPENAI_API_KEY="your DeepSeek API key"
|
|
671
|
+
export OPENAI_BASE_URL="https://api.deepseek.com"
|
|
672
|
+
omk bench run --executor openai-api --model deepseek-chat \
|
|
673
|
+
--judge-model deepseek-chat
|
|
674
|
+
|
|
675
|
+
# Moonshot (Kimi)
|
|
676
|
+
export OPENAI_API_KEY="your Moonshot API key"
|
|
677
|
+
export OPENAI_BASE_URL="https://api.moonshot.cn/v1"
|
|
678
|
+
omk bench run --executor openai-api --model moonshot-v1-8k \
|
|
679
|
+
--judge-model moonshot-v1-8k
|
|
680
|
+
```
|
|
681
|
+
|
|
682
|
+
**Ollama local model:**
|
|
683
|
+
|
|
684
|
+
```bash
|
|
685
|
+
omk bench run --executor "python examples/custom-executor/ollama-executor.py" \
|
|
686
|
+
--model llama3 --no-judge
|
|
687
|
+
```
|
|
688
|
+
|
|
689
|
+
**About the judge model:**
|
|
690
|
+
|
|
691
|
+
- `--judge-model` picks the model used by the LLM judge (default `haiku`)
|
|
692
|
+
- `--judge-executor` picks the executor the judge uses (defaults to `--executor`)
|
|
693
|
+
- If you don't have Claude, point `--judge-executor` and `--judge-model` at whatever model you have
|
|
694
|
+
- Add `--no-judge` to skip the LLM judge and rely on assertions alone
|
|
695
|
+
|
|
696
|
+
## Environment variables
|
|
435
697
|
|
|
436
698
|
| Variable | Description |
|
|
437
|
-
|
|
438
|
-
| `CCV_PROXY_URL` |
|
|
439
|
-
| `OMK_BENCH_PORT` |
|
|
699
|
+
|---|---|
|
|
700
|
+
| `CCV_PROXY_URL` | proxy requests through cc-viewer for live eval-traffic visualization |
|
|
701
|
+
| `OMK_BENCH_PORT` | report server port (default: 7799) |
|
|
440
702
|
|
|
441
703
|
## Requirements
|
|
442
704
|
|
|
443
705
|
- Node.js >= 20
|
|
444
|
-
- `claude` CLI
|
|
706
|
+
- `claude` CLI (for the default executor and LLM judge; see [Claude Code](https://claude.ai/code))
|
|
707
|
+
- not needed if you use other executors (openai / gemini) with `--no-judge`
|
|
708
|
+
|
|
709
|
+
## Security notice
|
|
710
|
+
|
|
711
|
+
This tool is designed for **local trusted environments** (dev machines, CI pipelines). The following features execute local code — make sure inputs come from a trusted source:
|
|
712
|
+
|
|
713
|
+
| Feature | Risk | Scope |
|
|
714
|
+
|---|---|---|
|
|
715
|
+
| **Custom assertions** (`custom`) | dynamically loads and executes user-specified `.mjs` files | only use assertion files you authored or reviewed |
|
|
716
|
+
| **eval-samples.json** | assertion configs can reference external file paths | don't use sample files from untrusted sources |
|
|
717
|
+
|
|
718
|
+
**Recommendations:**
|
|
719
|
+
|
|
720
|
+
- Do not expose `omk bench report` on the public internet (no auth)
|
|
721
|
+
- Don't use third-party eval-samples you haven't vetted
|
|
722
|
+
- Custom assertions have a 30-second timeout but no sandbox isolation
|
|
445
723
|
|
|
446
|
-
|
|
724
|
+
---
|
|
447
725
|
|
|
448
|
-
|
|
726
|
+
See [CHANGELOG](./CHANGELOG.md) for release notes. Contributions welcome — see [CONTRIBUTING](./CONTRIBUTING.md).
|