eval-builder 0.1.0__tar.gz → 0.1.2__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (127) hide show
  1. {eval_builder-0.1.0 → eval_builder-0.1.2}/.gitignore +2 -0
  2. {eval_builder-0.1.0 → eval_builder-0.1.2}/AGENTS.md +5 -1
  3. eval_builder-0.1.2/CHANGELOG.md +85 -0
  4. eval_builder-0.1.2/PKG-INFO +282 -0
  5. eval_builder-0.1.2/README.md +263 -0
  6. eval_builder-0.1.2/docs/demo.gif +0 -0
  7. eval_builder-0.1.2/docs/demo.tape +25 -0
  8. eval_builder-0.1.2/docs/label-sheet.png +0 -0
  9. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/README.md +5 -0
  10. eval_builder-0.1.2/examples/sample-labeling/README.md +79 -0
  11. eval_builder-0.1.2/examples/sample-labeling/decisions.json +98 -0
  12. eval_builder-0.1.2/examples/sample-labeling/drive_sheet.py +90 -0
  13. eval_builder-0.1.2/examples/sample-labeling/evalset/cases.yaml +2042 -0
  14. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/deepeval/dataset.json +1145 -0
  15. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/deepeval/judge.json +10 -0
  16. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/deepeval/test_eval_builder.py +165 -0
  17. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/inspect/dataset.jsonl +47 -0
  18. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/inspect/task.py +23 -0
  19. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/jsonl/cases.jsonl +47 -0
  20. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/manifest.json +90 -0
  21. eval_builder-0.1.2/examples/sample-labeling/evalset/exports/promptfoo/promptfooconfig.yaml +2314 -0
  22. eval_builder-0.1.2/examples/sample-labeling/evalset/ingest.json +33 -0
  23. eval_builder-0.1.2/examples/sample-labeling/evalset/judge_check.json +2188 -0
  24. eval_builder-0.1.2/examples/sample-labeling/evalset/judge_run_log.json +35 -0
  25. eval_builder-0.1.2/examples/sample-labeling/evalset/judgments.jsonl +1128 -0
  26. eval_builder-0.1.2/examples/sample-labeling/evalset/label_plan.json +673 -0
  27. eval_builder-0.1.2/examples/sample-labeling/evalset/label_sheet.csv +313 -0
  28. eval_builder-0.1.2/examples/sample-labeling/evalset/label_sheet.html +445 -0
  29. eval_builder-0.1.2/examples/sample-labeling/evalset/labels.jsonl +24 -0
  30. eval_builder-0.1.2/examples/sample-labeling/evalset/report.json +3593 -0
  31. eval_builder-0.1.2/examples/sample-labeling/evalset/report.md +148 -0
  32. eval_builder-0.1.2/examples/sample-labeling/evalset/rubric.yaml +100 -0
  33. eval_builder-0.1.2/examples/sample-labeling/evalset/selection.json +1227 -0
  34. eval_builder-0.1.2/examples/sample-labeling/evalset/traces.jsonl +48 -0
  35. eval_builder-0.1.2/examples/sample-labeling/fill.py +127 -0
  36. eval_builder-0.1.2/examples/sample-labeling/run.sh +35 -0
  37. eval_builder-0.1.2/examples/sample-labeling/sheet-run/drive_log.txt +12 -0
  38. eval_builder-0.1.2/examples/sample-labeling/sheet-run/labels.jsonl +24 -0
  39. eval_builder-0.1.2/examples/sample-labeling/sheet-run/sheet-export.png +0 -0
  40. eval_builder-0.1.2/examples/sample-labeling/sheet-run/sheet-first-case.png +0 -0
  41. eval_builder-0.1.2/examples/sample-labeling/sheet-run/sheet-phone.png +0 -0
  42. eval_builder-0.1.2/examples/sample_logs.jsonl +48 -0
  43. {eval_builder-0.1.0 → eval_builder-0.1.2}/pyproject.toml +5 -1
  44. {eval_builder-0.1.0 → eval_builder-0.1.2}/skills/eval-builder/SKILL.md +48 -17
  45. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/__init__.py +1 -1
  46. eval_builder-0.1.2/src/eval_builder/balance.py +55 -0
  47. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/cli.py +113 -9
  48. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/draft.py +3 -2
  49. eval_builder-0.1.2/src/eval_builder/export.py +610 -0
  50. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/ingest/__init__.py +5 -1
  51. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/ingest/formats.py +8 -2
  52. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/judge/check.py +139 -20
  53. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/judge/plan.py +52 -7
  54. eval_builder-0.1.2/src/eval_builder/label.py +549 -0
  55. eval_builder-0.1.2/src/eval_builder/label_sheet.py +481 -0
  56. eval_builder-0.1.2/src/eval_builder/mcp_server.py +253 -0
  57. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/report.py +14 -0
  58. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/select.py +34 -0
  59. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/setup_agents.py +14 -5
  60. eval_builder-0.1.2/src/eval_builder/status.py +60 -0
  61. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/workspace.py +12 -0
  62. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_cli_mcp.py +21 -0
  63. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_export.py +70 -2
  64. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_judge_check.py +15 -0
  65. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_judge_run.py +20 -0
  66. eval_builder-0.1.2/tests/test_label.py +349 -0
  67. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_setup.py +18 -0
  68. {eval_builder-0.1.0 → eval_builder-0.1.2}/uv.lock +1 -1
  69. eval_builder-0.1.0/CHANGELOG.md +0 -18
  70. eval_builder-0.1.0/PKG-INFO +0 -154
  71. eval_builder-0.1.0/README.md +0 -135
  72. eval_builder-0.1.0/src/eval_builder/export.py +0 -284
  73. eval_builder-0.1.0/src/eval_builder/mcp_server.py +0 -175
  74. eval_builder-0.1.0/src/eval_builder/status.py +0 -32
  75. {eval_builder-0.1.0 → eval_builder-0.1.2}/.github/workflows/ci.yml +0 -0
  76. {eval_builder-0.1.0 → eval_builder-0.1.2}/.github/workflows/release.yml +0 -0
  77. {eval_builder-0.1.0 → eval_builder-0.1.2}/CONTRIBUTING.md +0 -0
  78. {eval_builder-0.1.0 → eval_builder-0.1.2}/LICENSE +0 -0
  79. {eval_builder-0.1.0 → eval_builder-0.1.2}/SECURITY.md +0 -0
  80. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/judges/control_judges.py +0 -0
  81. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/judges/ollama_judge.py +0 -0
  82. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/expected_behaviors.yaml +0 -0
  83. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/fill_suite.py +0 -0
  84. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/cases.yaml +0 -0
  85. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/judge_check.json +0 -0
  86. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/judge_run_log.json +0 -0
  87. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/judgments.jsonl +0 -0
  88. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/labels.jsonl +0 -0
  89. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/report.json +0 -0
  90. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/report.md +0 -0
  91. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/judges/rubric.yaml +0 -0
  92. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/prepare.py +0 -0
  93. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/run.sh +0 -0
  94. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/cases.yaml +0 -0
  95. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/deepeval/dataset.json +0 -0
  96. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/deepeval/test_eval_builder.py +0 -0
  97. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/inspect/dataset.jsonl +0 -0
  98. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/inspect/task.py +0 -0
  99. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/jsonl/cases.jsonl +0 -0
  100. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/manifest.json +0 -0
  101. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/exports/promptfoo/promptfooconfig.yaml +0 -0
  102. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/ingest.json +0 -0
  103. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/judge_check.json +0 -0
  104. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/judge_run_log.json +0 -0
  105. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/judgments.jsonl +0 -0
  106. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/report.json +0 -0
  107. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/report.md +0 -0
  108. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/rubric.yaml +0 -0
  109. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/suite/selection.json +0 -0
  110. {eval_builder-0.1.0 → eval_builder-0.1.2}/examples/mt-bench/verify_exports.sh +0 -0
  111. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/io.py +0 -0
  112. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/judge/__init__.py +0 -0
  113. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/judge/run.py +0 -0
  114. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/judge/stats.py +0 -0
  115. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/redact.py +0 -0
  116. {eval_builder-0.1.0 → eval_builder-0.1.2}/src/eval_builder/schema.py +0 -0
  117. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/conftest.py +0 -0
  118. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/fixtures/anthropic_messages.json +0 -0
  119. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/fixtures/generic.jsonl +0 -0
  120. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/fixtures/langfuse_export.json +0 -0
  121. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/fixtures/openai_chat.jsonl +0 -0
  122. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/fixtures/otel_genai.json +0 -0
  123. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_draft.py +0 -0
  124. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_ingest.py +0 -0
  125. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_judge_stats.py +0 -0
  126. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_report.py +0 -0
  127. {eval_builder-0.1.0 → eval_builder-0.1.2}/tests/test_select.py +0 -0
@@ -9,9 +9,11 @@ build/
9
9
  *.egg-info/
10
10
  .DS_Store
11
11
  evalset/
12
+ !examples/sample-labeling/evalset/
12
13
  .deepeval/
13
14
  # Example inputs that prepare.py downloads or derives (reproducible from the pinned dataset)
14
15
  examples/mt-bench/data/
15
16
  # Large or derived files inside example workspaces
16
17
  examples/mt-bench/*/judge_requests.jsonl
17
18
  examples/mt-bench/*/traces.jsonl
19
+ examples/sample-labeling/evalset/judge_requests.jsonl
@@ -16,6 +16,9 @@ src/eval_builder/
16
16
  judge/run.py opt-in judge plugin runner (off by default, runs a user command)
17
17
  judge/check.py flip rate, kappa, accuracy, probes, verdicts
18
18
  judge/stats.py Wilson interval, Cohen's kappa, majority vote
19
+ label.py which cases a person should label, the CSV sheet, `label import`
20
+ label_sheet.py the offline HTML labeling sheet (one file, inline JS, no network)
21
+ balance.py the warning when cases or labels are mostly one outcome
19
22
  export.py promptfoo, DeepEval, Inspect AI, JSONL
20
23
  report.py report.md and report.json
21
24
  setup_agents.py `eval-builder setup` for Claude Code, Codex, Cursor
@@ -25,7 +28,8 @@ src/eval_builder/
25
28
 
26
29
  ## Rules
27
30
 
28
- - No model calls and no network access in the package. The judge runner only starts a
31
+ - No model calls and no network access in the package. (Exported files may call a
32
+ model when the user runs them, for example the DeepEval test calling a wired judge.) The judge runner only starts a
29
33
  command the user names, and only with an explicit flag.
30
34
  - Deterministic: same inputs and seed give the same selection and the same files.
31
35
  - Every number the tool reports must be traceable to a file in the workspace.
@@ -0,0 +1,85 @@
1
+ # Changelog
2
+
3
+ ## 0.1.2 (2026-10-08)
4
+
5
+ Human labels. No judge can be called trustworthy without them, and both real agent
6
+ sessions against 0.1.0 and 0.1.1 stopped at that step: they asked for labels and had
7
+ no way to collect them.
8
+
9
+ - `label` (CLI, and MCP `label`): picks the ready cases a person should label, 24 by
10
+ default (judge-check needs 20). The budget is split evenly across outcomes (the
11
+ judges' consensus verdict, or the logged failure flag before any judge has run), and
12
+ up to half of each share goes to cases where the judges disagree, flip across
13
+ repeats or move under padding. Every pick and its reason is in `label_plan.json`.
14
+ - `label` writes `label_sheet.html`: one self-contained file that works offline and
15
+ makes no network requests (its content security policy blocks them). One case per
16
+ screen with the earlier turns, the reply, the expected behavior and criteria;
17
+ pass/fail or A/B buttons (labels come from rubric.yaml), an optional note, keyboard
18
+ shortcuts, progress, and the judges' verdicts hidden. Progress survives a reload
19
+ through the browser's local storage. Export downloads `labels.jsonl` in exactly the
20
+ format judge-check reads and shows the same text to copy.
21
+ - `label` also writes `label_sheet.csv` for spreadsheet users (fill the label column).
22
+ - `label import <file>` (CLI, and MCP `label_import`): reads the sheet's labels.jsonl,
23
+ the filled-in CSV, a JSON list, or stdin (`-`); checks every case id and label
24
+ against cases.yaml and rubric.yaml; merges into `labels.jsonl` (a new label replaces
25
+ the same labeler's earlier one, other labelers are kept); lists rejected rows with
26
+ the reason and exits 1 when there are any.
27
+ - `select` and `judge-check` warn, with the real proportions, when at least 80% of the
28
+ selected cases or of the human labels share one outcome. judge-check also warns when
29
+ a judge gave the same verdict on every labeled case or its accuracy is no better
30
+ than always giving the most common label, and reports `majority_baseline` (that
31
+ always-the-common-label accuracy).
32
+ - DeepEval export: a pointwise judge that passed judge-check (or one forced with
33
+ `--judge`) now grades every case with its exact rubric.yaml prompt, through a custom
34
+ metric in `test_eval_builder.py` and `judge.json`. Ollama providers are called on the
35
+ local server; for other providers you fill in `call_judge`. Without a checked judge
36
+ the file keeps GEval. The export notes and README say that Inspect AI still uses its
37
+ default `model_graded_qa` grader.
38
+ - judge-check parses JSON verdicts that a token limit cut off mid-reason, and verdicts
39
+ wrapped in a json code fence.
40
+ - `status` and judge-check's `next` walk through the labeling step.
41
+ - `examples/sample-labeling/`: the whole flow on the bundled 48-conversation sample.
42
+
43
+ ## 0.1.1 (2026-10-08)
44
+
45
+ Fixes from a fresh-install audit and a real Claude Code session driving the MCP server.
46
+
47
+ - promptfoo export: the app now receives the whole conversation as chat messages
48
+ (`{{ messages | dump }}`). Before, only the last user turn was sent, so multi-turn cases
49
+ such as "Can you parallelize it?" reached the app without their context.
50
+ - promptfoo export wires a checked judge in as the llm-rubric grader: the first pointwise
51
+ judge that judge-check marked trustworthy and that has a new optional `provider` field in
52
+ rubric.yaml (a promptfoo provider id). `export --judge <id>` forces one, with a warning
53
+ when it did not pass. The judge's exact prompt becomes `rubricPrompt`.
54
+ - judge-check parses JSON verdicts such as `{"pass": false, "reason": "..."}`, so one judge
55
+ prompt works both in judge-check and as a promptfoo grader.
56
+ - judge-check summary rows (CLI `--json` and MCP) now carry 95% intervals, case counts and
57
+ label counts, and the result has a `next` step (for example how to write labels.jsonl).
58
+ - judge-plan documents the request row fields (`judge`, not `judge_id`), refuses to plan
59
+ when no case is ready or a judge is still TODO, and says how to write judgments.
60
+ - Clearer errors and next steps: judge-check without judgments, format detection failure
61
+ (shows the keys it saw and the shapes it accepts), `select` when fewer unique traces than
62
+ `-n` exist, `status` now walks through every step including judges.
63
+ - MCP `judge_check` takes `include_cases` to return per-case verdict counts and majorities.
64
+ - MCP tool descriptions say what each tool does, when to use it and what to call next.
65
+ - `setup`: installs the Codex skill in the same run when the codex CLI is on PATH but
66
+ `~/.codex` does not exist yet, and says to restart the agent after applying.
67
+ - `examples/sample_logs.jsonl`: 48 conversations from MT-Bench (CC BY 4.0) to try the
68
+ quickstart without your own logs.
69
+
70
+ ## 0.1.0 (2026-10-08)
71
+
72
+ First release.
73
+
74
+ - `ingest`: OpenAI chat JSONL (fine-tuning style and request/response logs), Anthropic
75
+ messages, Langfuse trace exports, OpenTelemetry GenAI spans (OTLP JSON and flat span
76
+ JSON), and generic input/output JSONL. Default-on redaction with a per-kind report.
77
+ - `select`: exact and near-duplicate removal, TF-IDF k-means clusters, stratum
78
+ coverage, failure oversampling, a recorded reason for every pick.
79
+ - `draft`, `validate`: case and rubric templates that the agent fills in with the user.
80
+ - `judge-plan`, `judge-run` (opt-in), `judge-check`: repeated trials, flip rate,
81
+ agreement with human labels (accuracy with Wilson intervals, Cohen's kappa), answer
82
+ order and padding probes, a verdict per judge.
83
+ - `export`: promptfoo, DeepEval, Inspect AI, JSONL.
84
+ - `report`: Markdown and JSON.
85
+ - `setup`: registers the MCP server with Claude Code, Codex and Cursor.
@@ -0,0 +1,282 @@
1
+ Metadata-Version: 2.5
2
+ Name: eval-builder
3
+ Version: 0.1.2
4
+ Summary: Turn real LLM app logs into an eval suite and measure which LLM judges you can trust.
5
+ Author: Abel Yagubyan
6
+ License-Expression: MIT
7
+ License-File: LICENSE
8
+ Keywords: deepeval,evals,inspect-ai,llm,llm-as-a-judge,mcp,promptfoo
9
+ Classifier: Operating System :: OS Independent
10
+ Classifier: Programming Language :: Python :: 3
11
+ Classifier: Topic :: Software Development :: Testing
12
+ Requires-Python: >=3.11
13
+ Requires-Dist: mcp>=1.2
14
+ Requires-Dist: numpy>=1.24
15
+ Requires-Dist: pyyaml>=6.0
16
+ Requires-Dist: scikit-learn>=1.3
17
+ Requires-Dist: scipy>=1.10
18
+ Description-Content-Type: text/markdown
19
+
20
+ # eval-builder
21
+
22
+ Your agent turns your app's real logs into an eval suite, and tells you which of its judges you can actually trust.
23
+
24
+ Try it on 48 sample conversations (no account, no API key, no model needed for these steps):
25
+
26
+ ```sh
27
+ curl -sLO https://raw.githubusercontent.com/Abelo9996/eval-builder/main/examples/sample_logs.jsonl
28
+ uvx eval-builder ingest sample_logs.jsonl # or your own logs: OpenAI, Anthropic, Langfuse, OpenTelemetry, JSONL
29
+ uvx eval-builder select -n 10 --stratify category && uvx eval-builder draft
30
+ ```
31
+
32
+ ![The quickstart running against eval-builder 0.1.1 from PyPI](docs/demo.gif)
33
+
34
+ What you get (real output of the commands above, eval-builder 0.1.1 from PyPI; the
35
+ curl and the three uvx calls took 35 s in total here; the very first uvx run also
36
+ downloads about 38 MiB, mostly numpy, scipy and scikit-learn, which took 6 s here):
37
+
38
+ ```
39
+ ingested 48 traces into evalset/traces.jsonl
40
+ sample_logs.jsonl: format=openai records=48 traces=48 skipped=0 sha256=2d5dd2295784988b
41
+ redactions: 0 {}
42
+ next: eval-builder select -n 30 (add --stratify <metadata keys> to cover them)
43
+ 48 traces -> 16 unique (32 exact dupes, 0 near dupes) -> selected 10 (8 failures) across 4 clusters
44
+ q121-llama-13b: failure (negative user feedback); 69% of unique traces are failures and at least 30% of picks are reserved for them
45
+ q81-alpaca-13b: failure (negative user feedback); 69% of unique traces are failures and at least 30% of picks are reserved for them
46
+ q102-alpaca-13b: covers category=reasoning (2 unique traces, 12%)
47
+ q111-alpaca-13b: adds variety within cluster 1 (8 traces, 50%; triangle, response, person); least similar to cases already picked there
48
+ ... (6 more picks)
49
+ next: run draft to turn the selection into cases.yaml
50
+ 10 case(s) added, 10 total in evalset/cases.yaml
51
+ next: read the cases (list_cases, or cases.yaml), define criteria and judges (set_rubric, or rubric.yaml), fill expected_behavior and criteria per case and set status: ready (update_case), then run validate
52
+ ```
53
+
54
+ `evalset/cases.yaml` now holds 10 real conversations with a TODO where the expected
55
+ behavior goes. Writing that, running judges and reading the statistics is the part your
56
+ coding agent does, through the MCP server (next section).
57
+
58
+ eval-builder is not another eval platform. It builds the suite and checks the judges,
59
+ then exports to the tools you already run: promptfoo, DeepEval, Inspect AI, or plain
60
+ JSONL. It never calls a model. Your coding agent (Claude Code, Codex, Cursor) does the
61
+ thinking; eval-builder does the selection, the bookkeeping and the statistics, and
62
+ writes down the evidence.
63
+
64
+ ## Use it with your agent
65
+
66
+ ```sh
67
+ uvx eval-builder setup --yes # registers the MCP server and the skill, then restart your agent
68
+ ```
69
+
70
+ Then type something like this into Claude Code, from the folder that has your logs:
71
+
72
+ > Our assistant's chat logs from last month are in ./logs. Build a small eval suite from
73
+ > them (about 10 cases) that I can run in promptfoo, and tell me whether I can trust a
74
+ > local judge for grading it. I have Ollama running with qwen2.5:7b-instruct and
75
+ > llama3.2:3b. Keep the judge runs small.
76
+
77
+ This is from a real headless Claude Code session (October 2026, Claude Code 2.1.293)
78
+ given exactly that prompt and 48 conversations from [`examples/sample_logs.jsonl`](examples/sample_logs.jsonl).
79
+ It ran ingest, select, draft, list_cases, set_rubric, ten update_case calls, validate,
80
+ judge_plan, 140 judge calls to Ollama through a short script it wrote, judge_check,
81
+ export and report in 32 turns and about 10 minutes. From its final answer:
82
+
83
+ > I built a 10-case promptfoo suite, but **you can't trust either local judge yet.**
84
+ > Neither one has been checked against human labels, and each confidently got an
85
+ > obvious case wrong.
86
+ >
87
+ > | judge | flip rate | padding helped | verdict |
88
+ > |---|---|---|---|
89
+ > | qwen2.5:7b-instruct | 10% [2%, 40%] (n=10) | 0% [0%, 28%] | not_enough_data |
90
+ > | llama3.2:3b | 20% [6%, 51%] (n=10) | 0% [0%, 28%] | not_enough_data |
91
+ >
92
+ > **qwen7b failed a clear case.** It passed the Hawaii rewrite 5 out of 5 times, though
93
+ > only the first sentence starts with "A".
94
+ >
95
+ > **No grader is wired in.** Because no judge passed, the `llm-rubric` checks fall back to
96
+ > promptfoo's default grader, which needs an OpenAI key. To grade locally anyway, export
97
+ > again with the judge set explicitly (the `export` tool's `judge` option, e.g. `qwen7b`).
98
+ > Its verdicts will be unverified.
99
+
100
+ It wrote every expected behavior itself because nobody was there to confirm them, and
101
+ said so on each case. It did not invent human labels; it asked for them. That is where
102
+ this session and the earlier one both stopped, so 0.1.2 adds the missing step: `label`
103
+ writes a sheet the person labels in their browser, and `label import` brings the
104
+ labels back for judge-check ([Human labels, end to end](#human-labels-end-to-end)).
105
+
106
+ An earlier session against 0.1.0 is where the promptfoo context bug fixed in 0.1.1 came
107
+ from. The agent read the export and told the user: "Right now it only sends the final
108
+ question (`{{input}}`), so your app won't see what "it" refers to in "Can you
109
+ parallelize it?"."
110
+
111
+ ## Example: real output
112
+
113
+ Everything below comes from runs on this machine (Apple M4, 16 GB) that are committed
114
+ in [`examples/mt-bench/`](examples/mt-bench/). The data is the LMSYS
115
+ [MT-Bench human judgments](https://huggingface.co/datasets/lmsys/mt_bench_human_judgments)
116
+ dataset (CC BY 4.0): 80 two-turn questions answered by 6 models, plus pairwise verdicts
117
+ from expert human judges.
118
+
119
+ **Logs to suite.** `ingest` read 480 conversations (OpenAI chat format, redaction on,
120
+ 0 matches). `select -n 24 --stratify category` removed 400 exact duplicates (each
121
+ question was asked to six models), clustered the 80 unique conversations into 9 topics
122
+ and picked 24 covering all 8 categories, 19 of them failures. From
123
+ [`suite/report.md`](examples/mt-bench/suite/report.md):
124
+
125
+ ```
126
+ | trace | cluster | represents | why it was picked |
127
+ | q81-alpaca-13b | 0 | 6 | failure (negative user feedback); 62% of unique traces are failures and at least 30% of picks are reserved for them |
128
+ | q82-alpaca-13b | 0 | 6 | adds variety within cluster 0 (11 traces, 14%; response, previous response, previous); least similar to cases already picked there |
129
+ | q156-alpaca-13b | 3 | 6 | covers category=humanities (10 unique traces, 12%) |
130
+ ```
131
+
132
+ The agent wrote expected behavior for the 24 cases, and `export` produced files that
133
+ Inspect AI 0.3.277, DeepEval 4.2.8 and promptfoo 0.124.0 all load (see
134
+ [`verify_exports.sh`](examples/mt-bench/verify_exports.sh)).
135
+
136
+ **Which judges can you trust?** Five local LLM judge setups (four models through
137
+ Ollama, one of them also at temperature 0) and two controls that are not language models judged 80 answer pairs where human experts
138
+ picked a winner. Each judge ran 5 times per pair, 5 more with the answers swapped, and
139
+ 5 more with an irrelevant paragraph appended to one answer: 8,400 calls, 0 errors.
140
+ `eval-builder judge-check`, from [`judges/report.md`](examples/mt-bench/judges/report.md):
141
+
142
+ | judge | verdict | flip rate | accuracy vs humans | kappa | survives answer swap | first-shown answer picked | padding helped |
143
+ |---|---|---|---|---|---|---|---|
144
+ | qwen2.5:7b-instruct, temp 0 | biased | 0% [0%, 5%] | 72% [62%, 81%] | 0.45 [0.26, 0.65] | 75% [65%, 83%] | 44% | 0% |
145
+ | qwen2.5:7b-instruct, temp 0.8 | biased | 12% [7%, 22%] | 70% [59%, 79%] | 0.40 [0.20, 0.60] | 75% [65%, 83%] | 42% | 1% |
146
+ | qwen2.5:3b-instruct | biased | 12% [7%, 22%] | 56% [45%, 67%] | 0.13 [-0.08, 0.35] | 44% [33%, 55%] | 35% | 9% |
147
+ | gemma2:2b | unstable | 59% [48%, 69%] | 56% [45%, 67%] | 0.14 [-0.07, 0.36] | 38% [28%, 48%] | 35% | 11% |
148
+ | llama3.2:3b | unstable | 56% [45%, 67%] | 51% [40%, 62%] | 0.04 [-0.17, 0.26] | 14% [8%, 23%] | 20% | 10% |
149
+ | control: longer answer wins | biased | 0% [0%, 5%] | 66% [55%, 76%] | 0.32 [0.12, 0.53] | 100% [95%, 100%] | 50% | 30% |
150
+ | control: coin flip | unstable | 95% [88%, 98%] | 51% [40%, 62%] | 0.03 [-0.19, 0.24] | 55% [44%, 65%] | 50% | 25% |
151
+
152
+ n = 80 pairs per judge; brackets are 95% intervals. No judge passed. The closest,
153
+ qwen2.5 7B, reaches kappa 0.40 to 0.45 with the human experts, but its verdict
154
+ flips on a quarter of the pairs when the answer order is swapped, and setting the
155
+ temperature to 0 removes the run-to-run flips without removing that order effect.
156
+ llama3.2 3B picked whichever answer it saw second 80% of the time. The longer-answer
157
+ rule never flips and agrees with humans 66% of the time, which is why stability and
158
+ accuracy alone are not enough: the padding probe catches it.
159
+
160
+ The same check on the suite's own pass/fail judge (qwen2.5 7B, 24 cases, no human
161
+ labels) gives `unstable`: its verdict changed across 5 identical calls on 7 of 24
162
+ cases (29%, interval 15% to 49%).
163
+
164
+ ### Human labels, end to end
165
+
166
+ ```sh
167
+ uvx eval-builder label # picks 24 cases, writes label_sheet.html
168
+ uvx eval-builder label import ~/Downloads/labels.jsonl # what the sheet's Export button saved
169
+ uvx eval-builder judge-check
170
+ ```
171
+
172
+ ![The labeling sheet: one case per screen, pass and fail buttons with keyboard shortcuts](docs/label-sheet.png)
173
+
174
+ Run on the 48 sample conversations with every model's answer kept as a case (47
175
+ cases), three local judges (1,128 calls, 0 errors), and 24 labels entered through the
176
+ sheet in headless Chrome. The labels were made by the developer (Claude Code reading
177
+ each case for him), not by an independent annotator, so read this as a demonstration
178
+ of the flow. Details and every file: [`examples/sample-labeling/`](examples/sample-labeling/).
179
+
180
+ - `label` split the 24 picks 12/12 between cases the judges called pass and fail; 15
181
+ are cases where the judges disagree, flip or move under padding.
182
+ - The sheet made no network requests, survived a reload mid-way, and its download was
183
+ byte-identical to the copy box. `label import` took 24 rows, rejected 0.
184
+ - `judge-check` against those labels (fail 18, pass 6):
185
+
186
+ | judge | verdict | flip rate | accuracy vs labels | kappa |
187
+ |---|---|---|---|---|
188
+ | qwen2.5:7b-instruct, temp 0 | trustworthy | 0% [0%, 8%] | 75% [55%, 88%] | 0.50 [0.15, 0.85] |
189
+ | qwen2.5:7b-instruct, temp 0.8 | trustworthy | 9% [3%, 20%] | 71% [51%, 85%] | 0.44 [0.09, 0.79] |
190
+ | llama3.2:3b | unstable | 47% [33%, 61%] | 54% [35%, 72%] | 0.12 [-0.26, 0.50] |
191
+
192
+ The qwen judges pass the default thresholds, but judge-check also warns that their
193
+ accuracy is no better than always answering "fail" (75% of the labels), and every
194
+ miss went the same way: both passed a reply that said "Here is an allegorical poem"
195
+ and then wrote no poem. With 24 labels the kappa interval runs from about 0.1 to 0.8.
196
+ The DeepEval export wired the 0.8-temperature qwen judge in with its exact prompt;
197
+ DeepEval 4.2.8 ran three logged cases through it and it made the same call on the
198
+ missing poem.
199
+
200
+ ## How it works
201
+
202
+ | step | what it does | what it uses |
203
+ |---|---|---|
204
+ | `ingest` | Reads OpenAI chat JSONL (fine-tuning style or request/response logs), Anthropic messages, Langfuse trace exports, OpenTelemetry GenAI spans (OTLP JSON, current `gen_ai.input.messages` and older `gen_ai.prompt.N` attributes), or generic input/output JSONL. Normalizes to one schema with input, output, earlier turns, tools called, error flag, user feedback, route and model. Redacts emails, API keys, Luhn-valid card numbers, phone numbers and SSN-shaped numbers by default and reports counts. | Python standard library |
205
+ | `select` | Removes exact duplicates (normalized hash of the user turns) and near-duplicates (character n-gram cosine), clusters the rest by topic, then picks: failures first (errors, negative feedback), at least one case per stratum (route, tool, feedback, any metadata key you name), one central case per cluster, and fills the rest by cluster size with the least similar remaining cases. Every pick records why it was picked. | scikit-learn (TF-IDF, k-means) |
206
+ | `draft`, `validate` | Writes `cases.yaml` and `rubric.yaml` with TODO markers. The agent fills in expected behavior and criteria with you. Validation refuses ready cases that still contain TODO or reference unknown criteria. | PyYAML |
207
+ | `judge-plan` | Lists every judge call to make: each case N times, plus probes that swap the answer order (pairwise judges) and pad an answer with an irrelevant paragraph. | |
208
+ | `judge-run` | Optional and off by default. Sends each request as a JSON line to a command you name (your script, your provider, your keys) and records the verdicts. eval-builder ships no API keys and no provider code. | your command |
209
+ | `label`, `label import` | Picks the ready cases a person should label (default 24): the budget is split across outcomes (the judges' consensus, or the logged failure flag before judges ran) and up to half of each share goes to cases where judges disagree, flip across repeats or move under padding. Writes `label_sheet.html`, one offline file (one case per screen, pass/fail or A/B buttons, a note, keyboard shortcuts, judge verdicts hidden, progress kept in the browser) whose Export button downloads `labels.jsonl` in the format judge-check reads, and `label_sheet.csv` for spreadsheet users. `label import` checks the file against `cases.yaml` and merges it into the workspace. | Python standard library; the sheet is plain HTML and JavaScript |
210
+ | `judge-check` | Per judge: flip rate across repeated calls (with a Wilson interval), self-agreement, majority-of-3 vote stability, accuracy and Cohen's kappa against your human labels (with intervals), position consistency and first-shown preference, and how often padding moved the verdict toward the padded answer. Verdict: `trustworthy`, `unstable`, `biased`, `misaligned`, or `not_enough_data`, with the numbers behind it. | |
211
+ | `export` | promptfoo `promptfooconfig.yaml` (the full conversation as chat messages, llm-rubric asserts, and a judge that passed judge-check wired in as the grader when its rubric entry names a promptfoo `provider`), DeepEval dataset plus a `deepeval test run` file (the same checked judge grades every case with its exact prompt; without one, GEval), Inspect AI dataset plus `task.py` (Inspect's default `model_graded_qa` grader), plain JSONL. The manifest lists file hashes and which judges passed. | |
212
+ | `report` | `report.md` and `report.json`: sources with sha256, counts, redactions, selection reasons, the judge table, and the limits. | |
213
+
214
+ The verdict thresholds are explicit flags with defaults: flip rate at most 20% of cases,
215
+ position consistency at least 80%, padding helps at most 10% of cases, kappa at least
216
+ 0.4 on at least 20 human-labeled cases. A judge without human labels is never called
217
+ trustworthy. When at least 80% of the selected cases or of the human labels share one
218
+ outcome, `select` and `judge-check` say so with the real proportions, because a judge
219
+ that always gives that answer would look accurate on them.
220
+
221
+ ## Setup for agents
222
+
223
+ ```sh
224
+ uvx eval-builder setup # shows what it would change
225
+ uvx eval-builder setup --yes # applies it; restart your agent afterwards
226
+ ```
227
+
228
+ `setup` registers the MCP server (`uvx eval-builder mcp`) with Claude Code (`claude mcp
229
+ add --scope user`), Codex (`[mcp_servers.eval-builder]` in `~/.codex/config.toml`) and
230
+ Cursor (`~/.cursor/mcp.json`), and copies the agent instructions to
231
+ `~/.claude/skills/eval-builder/` and `~/.codex/skills/eval-builder/`. It backs up any
232
+ file it edits and does nothing on a second run. The workflow the agent follows is in
233
+ [`skills/eval-builder/SKILL.md`](skills/eval-builder/SKILL.md).
234
+
235
+ MCP tools: `ingest`, `select`, `draft`, `list_cases`, `update_case`, `set_rubric`,
236
+ `validate`, `judge_plan`, `label`, `label_import`, `judge_check`, `export`, `report`,
237
+ `status`. The judge runner
238
+ is CLI only, because it executes a command.
239
+
240
+ ## What it can't do
241
+
242
+ - It does not write expected behavior or human labels. The agent drafts expected
243
+ behavior with you; labels must come from people. Without labels, judge-check can
244
+ tell you a judge is unstable or biased, but not that it is right. `label` makes the
245
+ labeling quick, but someone still has to read each case.
246
+ - Picking labels where judges disagree makes each label more informative, but those
247
+ cases are harder than average, so accuracy measured on them leans pessimistic.
248
+ `--uncertain-share 0` picks by outcome and topic only.
249
+ - The labeling sheet is a static page, so it cannot save files: the person has to
250
+ click Export (or copy the text) and import it. Unexported progress lives only in
251
+ that browser's local storage.
252
+ - Selection is lexical. Two requests that mean the same thing in different words can
253
+ land in different clusters, and near-duplicate detection only catches close textual
254
+ matches.
255
+ - Redaction is pattern-based. Names, street addresses and free-form secrets get
256
+ through. Look at `traces.jsonl` before sharing a workspace.
257
+ - The bias probes cover answer order and irrelevant length only. Self-preference,
258
+ style bias and rubric misreadings are not measured.
259
+ - Small samples give wide intervals. The report prints them; read them.
260
+ - It does not run your app, your judges or your eval. The agent (or a script you name
261
+ with `judge-run`) calls the judge model; the exported files run the eval in promptfoo,
262
+ DeepEval or Inspect AI.
263
+ - A checked judge is wired in as the grader in promptfoo and DeepEval, and only a
264
+ pointwise one. promptfoo needs its prompt to answer in JSON (`{"pass": ...,
265
+ "reason": ...}`), because llm-rubric cannot parse a bare "pass". The DeepEval test
266
+ calls Ollama judges itself; for any other provider you fill in `call_judge`. The
267
+ Inspect AI export still uses Inspect's default `model_graded_qa` grader, not your
268
+ checked judge.
269
+
270
+ ## Privacy and safety
271
+
272
+ Everything runs locally. eval-builder makes no network calls and sends nothing
273
+ anywhere. The only process it starts is the judge command you name with
274
+ `judge-run --enable-judge-plugin`. The labeling sheet is one HTML file with a content
275
+ security policy that blocks every network request; it keeps unexported labels in the
276
+ browser's local storage on that machine. Exported test files can call a model when
277
+ you run them (the DeepEval test calls the judge it was given). Redaction is on unless you pass `--no-redact`, and
278
+ the report lists what was replaced (counts and kinds, never the values).
279
+
280
+ ## License
281
+
282
+ MIT. See [LICENSE](LICENSE).