@hona/openeval 0.2.0 → 0.2.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -1,29 +1,74 @@
1
- # OpenEval
1
+ <div align="center">
2
+ <h1>OpenEval</h1>
3
+ <p><strong>Write the task. Judge the evidence.</strong></p>
4
+ <p>Prompt-and-rubric evaluations for agents. Typed declarations, isolated runs, inspectable scores.</p>
5
+ <p>
6
+ <a href="https://openev.al">Website</a> ·
7
+ <a href="https://openeval.pages.dev">Live preview</a> ·
8
+ <a href="https://openev.al/docs/quickstart/">Write your first eval</a> ·
9
+ <a href="https://openev.al/docs/reference/">CLI reference</a> ·
10
+ <a href="https://www.npmjs.com/package/@hona/openeval">npm</a>
11
+ </p>
12
+ <p>
13
+ <a href="https://www.npmjs.com/package/@hona/openeval"><img src="https://img.shields.io/npm/v/%40hona%2Fopeneval?style=flat-square&color=66d38a" alt="npm version"></a>
14
+ <a href="https://bun.com"><img src="https://img.shields.io/badge/Bun-1.4.2%2B-f9f1e1?style=flat-square" alt="Bun 1.4.2 or later"></a>
15
+ <a href="https://github.com/Hona/openeval/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-MIT-66d38a?style=flat-square" alt="MIT license"></a>
16
+ </p>
17
+ </div>
18
+
19
+ ![The OpenEval workbench: rubric source, a sample SQL response, and linked metric decisions](https://raw.githubusercontent.com/Hona/openeval/main/docs/images/workbench.png)
20
+
21
+ *Interactive documentation example. Viewer screenshots use illustrative data and fictional model labels.*
22
+
23
+ ## An eval is two files
24
+
25
+ | File | What you write | Who reads it |
26
+ | --- | --- | --- |
27
+ | `prompt.md` | A natural, focused task | Candidate agent |
28
+ | `judge.md` | Named metrics and pass/fail criteria | Judge agent |
29
+ | `eval.ts` *(optional)* | Workspace preparation and early stopping | Host |
30
+
31
+ **`evals/ask-dialect/prompt.md`**
2
32
 
3
- A typed prompt-plus-LLM-judge SDK for Bun. Run agents in isolated containers,
4
- record their work, grade the evidence, and explore the results in a local viewer.
33
+ ```md
34
+ Write a SQL query for the ten most recent orders for a customer.
35
+ ```
5
36
 
6
- ## Install
37
+ **`evals/ask-dialect/judge.md`**
7
38
 
8
- Requires Bun 1.4.2 or later, Docker, and an authenticated OpenCode installation.
9
- The agent runtime currently uses OpenCode `0.0.0-beta-19296`.
39
+ ```md
40
+ # Requests the SQL dialect
10
41
 
11
- ```sh
12
- bun add --exact @hona/openeval
42
+ ## Metric: asked_dialect — Asks for the SQL dialect
43
+
44
+ Pass when the agent asks which database or SQL dialect is in use.
45
+ Fail when it assumes a dialect without asking. Asking alongside a draft counts.
46
+
47
+ ## Metric: safe_parameters — Uses bound parameters
48
+
49
+ Pass when the proposed query uses a bound customer-ID parameter and explains
50
+ how to supply its value. Fail when it interpolates customer input into SQL
51
+ or does not provide a parameterized query.
13
52
  ```
14
53
 
15
- ## Declare a benchmark
54
+ | Recorded response | Asks for dialect | Bound parameters |
55
+ | --- | --- | --- |
56
+ | Asks which DB; provides a bound-parameter draft | **1** | **1** |
57
+ | Assumes PostgreSQL; uses `$1` | **0** | **1** |
58
+ | Only asks which database | **1** | **0** |
59
+ | Required recording is unavailable | **null** | **null** |
60
+
61
+ → [Write good rubrics](https://openev.al/docs/rubrics/) · [Download the SQL starter](https://openev.al/starter.zip)
62
+
63
+ ## Choose models. Run. Inspect.
16
64
 
17
- ```text
18
- my-benchmark/
19
- benchmark.ts
20
- evals/
21
- ask-dialect/
22
- prompt.md
23
- judge.md
65
+ Requires **Bun 1.4.2+**, **Docker**, and connected models in **OpenCode**.
66
+
67
+ ```sh
68
+ bun add --exact @hona/openeval
24
69
  ```
25
70
 
26
- `benchmark.ts`:
71
+ **`benchmark.ts`** — replace the model references with your connected models:
27
72
 
28
73
  ```ts
29
74
  import type { Benchmark } from "@hona/openeval";
@@ -35,99 +80,86 @@ export default {
35
80
  } satisfies Benchmark;
36
81
  ```
37
82
 
38
- Replace the model references with models connected in OpenCode.
39
-
40
- `evals/ask-dialect/prompt.md`:
41
-
42
- ```md
43
- Write a SQL query for the ten most recent orders for a customer.
83
+ ```sh
84
+ bunx --bun @hona/openeval image
85
+ bunx --bun @hona/openeval plan --only-eval ask-dialect
86
+ bunx --bun @hona/openeval run --only-repetition 1
87
+ bunx --bun @hona/openeval view
44
88
  ```
45
89
 
46
- `evals/ask-dialect/judge.md`:
47
-
48
- ```md
49
- # Requests the SQL dialect
50
-
51
- ## Metric: asked_dialect — Asks which SQL dialect to use
90
+ The viewer opens at **http://127.0.0.1:4173**. `run` resumes the same aggregate;
91
+ scope flags select work while retaining existing scores.
52
92
 
53
- Pass when the agent asks which database or SQL dialect is in use.
54
- Fail when it assumes a dialect without asking. Asking alongside a draft counts.
93
+ ```mermaid
94
+ flowchart LR
95
+ P["prompt.md"] --> C["Isolated candidate"] --> E["Recording"]
96
+ J["judge.md"] --> G["Judge + citations"]
97
+ E --> G --> S["Metric scores"] --> V["Results viewer"]
55
98
  ```
56
99
 
57
- Prompts are sent verbatim. Each metric is `0`, `1`, or `null` when evidence is
58
- insufficient. The judge supplies citations to the recording. A shared native
59
- judge agent supplies the evidence and submission protocol; rubrics contain the
60
- task-specific criteria. See [JUDGING.md](JUDGING.md).
100
+ ## See what earned the score
61
101
 
62
- ## Run and inspect
102
+ ![Model scores, completed checks, runtime, and cost in the results viewer](https://raw.githubusercontent.com/Hona/openeval/main/docs/images/results.png)
63
103
 
64
- From your benchmark directory:
104
+ | Capability | What you get | Guide |
105
+ | --- | --- | --- |
106
+ | Multiple metrics | Independent decisions from one recording | [Rubrics](https://openev.al/docs/rubrics/) |
107
+ | Controlled workspaces | Readable files, pinned Git inputs, preparation | [Workspaces](https://openev.al/docs/workspaces/) |
108
+ | Small batches | Eval, model, repetition, and cost controls | [Running](https://openev.al/docs/running/) |
109
+ | Transparent scores | Equal eval weights; bounds for unresolved checks | [Scoring](https://openev.al/docs/scoring/) |
110
+ | Evidence inspection | Sessions, tool results, artifacts, and citations | [Evidence](https://openev.al/docs/evidence/) |
111
+ | Rejudging | New judgments from retained, immutable recordings | [Evidence](https://openev.al/docs/evidence/#revise) |
65
112
 
66
- ```sh
67
- bunx --bun @hona/openeval image
68
- bunx --bun @hona/openeval plan
69
- bunx --bun @hona/openeval run
70
- bunx --bun @hona/openeval view
71
- ```
113
+ <details>
114
+ <summary><strong>Inspect a judgment and its evidence</strong></summary>
72
115
 
73
- The viewer opens at `http://127.0.0.1:4173`. Use `--benchmark <directory>` to point
74
- any command at another benchmark, or `--port <port>` for another viewer port.
116
+ ![SQL eval drilldown with individual metric decisions and evidence links](https://raw.githubusercontent.com/Hona/openeval/main/docs/images/judgment.png)
75
117
 
76
- `run` resumes the current aggregate and reuses unchanged work. `--new` starts a
77
- separate result. Repeat `--only-eval`, `--only-model`, or `--only-repetition` to
78
- execute a small scope while retaining the full aggregate. `--max-cost <usd>`
79
- sets a scheduling budget for the invocation. Run `openeval --help` for commands.
118
+ </details>
80
119
 
81
- Scores average repetitions per metric, metrics per eval, then evals equally.
82
- Unresolved checks produce completion bounds until the final percentage is known.
83
- The viewer provides metric drilldowns and recorded candidate and judge sessions.
84
- Elapsed time measures active execution intervals, counting overlapping work once.
120
+ <details>
121
+ <summary><strong>Watch candidate and judge work in the live queue</strong></summary>
85
122
 
86
- ## Prepare a workspace
123
+ ![Live queue with separate execution and judging stages](https://raw.githubusercontent.com/Hona/openeval/main/docs/images/queue.png)
87
124
 
88
- An eval's optional `workspace/` directory supplies files. An optional `eval.ts`
89
- can declare a pinned Git repository or readable revision overlays and preparation
90
- commands:
125
+ </details>
126
+
127
+ ## Use the SDK
91
128
 
92
129
  ```ts
93
- import type { Eval } from "@hona/openeval";
130
+ import { runBenchmark } from "@hona/openeval";
94
131
 
95
- export default {
96
- prepare: [{ cwd: ".", argv: ["bun", "install", "--frozen-lockfile"] }],
97
- } satisfies Eval;
132
+ await runBenchmark("./my-benchmark", {
133
+ onlyEvals: ["ask-dialect"],
134
+ onlyRepetitions: [1],
135
+ });
98
136
  ```
99
137
 
100
- Preparation runs before recording the initial candidate workspace. Candidate
101
- containers receive project inputs, never judge rubrics or evaluator storage.
102
- Candidates finish naturally, meet an opted-in irreversible judge decision, or
103
- time out after at most 45 minutes. Finalized recordings remain immutable.
138
+ ## Write evals with an agent
104
139
 
105
- ## SDK
140
+ Use the public [Eval Writing skill](https://github.com/Hona/openeval/tree/main/.opencode/skills/eval-writing)
141
+ to turn a real failure into an eval, review a rubric, or investigate misleading
142
+ scores. It guides an agent through concrete false-pass/false-failure examples,
143
+ accepted alternatives, evidence requirements, and human-reviewed calibration.
106
144
 
107
- ```ts
108
- import { runBenchmark, serveResults } from "@hona/openeval";
109
-
110
- await runBenchmark("./my-benchmark");
111
- ```
145
+ Copy the whole `.opencode/skills/eval-writing/` directory, including `references/`,
146
+ into the same path in your project. For global use, copy it to
147
+ `~/.config/opencode/skills/eval-writing/`. Then run **`/eval-writing`** in OpenCode.
112
148
 
113
- The package also exports types, result readers, snapshot and rejudge functions,
114
- and reusable view/session data through `@hona/openeval/types`, `/results`,
115
- `/view`, and `/session`.
149
+ > Use eval-writing to review this task and rubric. Show me the strongest false
150
+ > pass and false failure, then propose the smallest improvement.
116
151
 
117
- ## Develop and release
152
+ The skill includes a framework-neutral workflow, fictional coaching examples,
153
+ an OpenEval-specific reference, and a broad public-research guide.
118
154
 
119
- ```sh
120
- bun install --frozen-lockfile
121
- bun run typecheck
122
- bun test
123
- bun run release:pack
124
- bun run release:verify
125
- ```
155
+ ## Develop
126
156
 
127
- The private benchmark client is maintained in a separate repository. It installs
128
- an exact published SDK version. This repository contains only the reusable SDK,
129
- CLI, viewer, and development inputs. See [RELEASING.md](https://github.com/Hona/openeval/blob/main/RELEASING.md).
157
+ | Command | Purpose |
158
+ | --- | --- |
159
+ | `bun run site:dev` | Landing page and docs with hot reload on port 4176 |
160
+ | `bun run site:build && bun run site:verify` | Prerender pages and verify links and starter files |
161
+ | `bun run typecheck && bun test` | Local SDK checks; no live models |
162
+ | `bun run release:pack && bun run release:verify` | Verify the actual npm archive in a separate consumer |
130
163
 
131
- OpenEval is MIT licensed. The vendored session UI retains its upstream MIT
132
- license and revision in `vendor/session-ui`. Built viewer distributions include
133
- third-party license notices.
164
+ - [Release guide](https://github.com/Hona/openeval/blob/main/RELEASING.md) · [Judge protocol](JUDGING.md)
165
+ - MIT licensed. The viewer includes upstream third-party license notices.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@hona/openeval",
3
- "version": "0.2.0",
3
+ "version": "0.2.1",
4
4
  "description": "Typed prompt-and-judge evaluations with isolated agents, recorded evidence, and a results viewer",
5
5
  "license": "MIT",
6
6
  "repository": {
@@ -28,7 +28,7 @@ const positive = (value: unknown, fallback: number, label: string) => {
28
28
  export const modelRef = (value: unknown): ModelRef => {
29
29
  if (
30
30
  typeof value !== "string" ||
31
- !/^[\w.-]+\/[^\s/#]+(?:#[\w.-]+)?$/.test(value)
31
+ !/^[\w.-]+\/[^\s/#]+(?:\/[^\s/#]+)*(?:#[\w.-]+)?$/.test(value)
32
32
  )
33
33
  throw new Error(`Invalid model reference: ${String(value)}`);
34
34
  return value as ModelRef;