@hona/openeval 0.2.0 → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +120 -88
- package/package.json +1 -1
- package/src/app/load-benchmark.ts +1 -1
package/README.md
CHANGED
|
@@ -1,29 +1,74 @@
|
|
|
1
|
-
|
|
1
|
+
<div align="center">
|
|
2
|
+
<h1>OpenEval</h1>
|
|
3
|
+
<p><strong>Write the task. Judge the evidence.</strong></p>
|
|
4
|
+
<p>Prompt-and-rubric evaluations for agents. Typed declarations, isolated runs, inspectable scores.</p>
|
|
5
|
+
<p>
|
|
6
|
+
<a href="https://openev.al">Website</a> ·
|
|
7
|
+
<a href="https://openeval.pages.dev">Live preview</a> ·
|
|
8
|
+
<a href="https://openev.al/docs/quickstart/">Write your first eval</a> ·
|
|
9
|
+
<a href="https://openev.al/docs/reference/">CLI reference</a> ·
|
|
10
|
+
<a href="https://www.npmjs.com/package/@hona/openeval">npm</a>
|
|
11
|
+
</p>
|
|
12
|
+
<p>
|
|
13
|
+
<a href="https://www.npmjs.com/package/@hona/openeval"><img src="https://img.shields.io/npm/v/%40hona%2Fopeneval?style=flat-square&color=66d38a" alt="npm version"></a>
|
|
14
|
+
<a href="https://bun.com"><img src="https://img.shields.io/badge/Bun-1.4.2%2B-f9f1e1?style=flat-square" alt="Bun 1.4.2 or later"></a>
|
|
15
|
+
<a href="https://github.com/Hona/openeval/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-MIT-66d38a?style=flat-square" alt="MIT license"></a>
|
|
16
|
+
</p>
|
|
17
|
+
</div>
|
|
18
|
+
|
|
19
|
+

|
|
20
|
+
|
|
21
|
+
*Interactive documentation example. Viewer screenshots use illustrative data and fictional model labels.*
|
|
22
|
+
|
|
23
|
+
## An eval is two files
|
|
24
|
+
|
|
25
|
+
| File | What you write | Who reads it |
|
|
26
|
+
| --- | --- | --- |
|
|
27
|
+
| `prompt.md` | A natural, focused task | Candidate agent |
|
|
28
|
+
| `judge.md` | Named metrics and pass/fail criteria | Judge agent |
|
|
29
|
+
| `eval.ts` *(optional)* | Workspace preparation and early stopping | Host |
|
|
30
|
+
|
|
31
|
+
**`evals/ask-dialect/prompt.md`**
|
|
2
32
|
|
|
3
|
-
|
|
4
|
-
|
|
33
|
+
```md
|
|
34
|
+
Write a SQL query for the ten most recent orders for a customer.
|
|
35
|
+
```
|
|
5
36
|
|
|
6
|
-
|
|
37
|
+
**`evals/ask-dialect/judge.md`**
|
|
7
38
|
|
|
8
|
-
|
|
9
|
-
|
|
39
|
+
```md
|
|
40
|
+
# Requests the SQL dialect
|
|
10
41
|
|
|
11
|
-
|
|
12
|
-
|
|
42
|
+
## Metric: asked_dialect — Asks for the SQL dialect
|
|
43
|
+
|
|
44
|
+
Pass when the agent asks which database or SQL dialect is in use.
|
|
45
|
+
Fail when it assumes a dialect without asking. Asking alongside a draft counts.
|
|
46
|
+
|
|
47
|
+
## Metric: safe_parameters — Uses bound parameters
|
|
48
|
+
|
|
49
|
+
Pass when the proposed query uses a bound customer-ID parameter and explains
|
|
50
|
+
how to supply its value. Fail when it interpolates customer input into SQL
|
|
51
|
+
or does not provide a parameterized query.
|
|
13
52
|
```
|
|
14
53
|
|
|
15
|
-
|
|
54
|
+
| Recorded response | Asks for dialect | Bound parameters |
|
|
55
|
+
| --- | --- | --- |
|
|
56
|
+
| Asks which DB; provides a bound-parameter draft | **1** | **1** |
|
|
57
|
+
| Assumes PostgreSQL; uses `$1` | **0** | **1** |
|
|
58
|
+
| Only asks which database | **1** | **0** |
|
|
59
|
+
| Required recording is unavailable | **null** | **null** |
|
|
60
|
+
|
|
61
|
+
→ [Write good rubrics](https://openev.al/docs/rubrics/) · [Download the SQL starter](https://openev.al/starter.zip)
|
|
62
|
+
|
|
63
|
+
## Choose models. Run. Inspect.
|
|
16
64
|
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
ask-dialect/
|
|
22
|
-
prompt.md
|
|
23
|
-
judge.md
|
|
65
|
+
Requires **Bun 1.4.2+**, **Docker**, and connected models in **OpenCode**.
|
|
66
|
+
|
|
67
|
+
```sh
|
|
68
|
+
bun add --exact @hona/openeval
|
|
24
69
|
```
|
|
25
70
|
|
|
26
|
-
|
|
71
|
+
**`benchmark.ts`** — replace the model references with your connected models:
|
|
27
72
|
|
|
28
73
|
```ts
|
|
29
74
|
import type { Benchmark } from "@hona/openeval";
|
|
@@ -35,99 +80,86 @@ export default {
|
|
|
35
80
|
} satisfies Benchmark;
|
|
36
81
|
```
|
|
37
82
|
|
|
38
|
-
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
|
|
43
|
-
Write a SQL query for the ten most recent orders for a customer.
|
|
83
|
+
```sh
|
|
84
|
+
bunx --bun @hona/openeval image
|
|
85
|
+
bunx --bun @hona/openeval plan --only-eval ask-dialect
|
|
86
|
+
bunx --bun @hona/openeval run --only-repetition 1
|
|
87
|
+
bunx --bun @hona/openeval view
|
|
44
88
|
```
|
|
45
89
|
|
|
46
|
-
`
|
|
47
|
-
|
|
48
|
-
```md
|
|
49
|
-
# Requests the SQL dialect
|
|
50
|
-
|
|
51
|
-
## Metric: asked_dialect — Asks which SQL dialect to use
|
|
90
|
+
The viewer opens at **http://127.0.0.1:4173**. `run` resumes the same aggregate;
|
|
91
|
+
scope flags select work while retaining existing scores.
|
|
52
92
|
|
|
53
|
-
|
|
54
|
-
|
|
93
|
+
```mermaid
|
|
94
|
+
flowchart LR
|
|
95
|
+
P["prompt.md"] --> C["Isolated candidate"] --> E["Recording"]
|
|
96
|
+
J["judge.md"] --> G["Judge + citations"]
|
|
97
|
+
E --> G --> S["Metric scores"] --> V["Results viewer"]
|
|
55
98
|
```
|
|
56
99
|
|
|
57
|
-
|
|
58
|
-
insufficient. The judge supplies citations to the recording. A shared native
|
|
59
|
-
judge agent supplies the evidence and submission protocol; rubrics contain the
|
|
60
|
-
task-specific criteria. See [JUDGING.md](JUDGING.md).
|
|
100
|
+
## See what earned the score
|
|
61
101
|
|
|
62
|
-
|
|
102
|
+

|
|
63
103
|
|
|
64
|
-
|
|
104
|
+
| Capability | What you get | Guide |
|
|
105
|
+
| --- | --- | --- |
|
|
106
|
+
| Multiple metrics | Independent decisions from one recording | [Rubrics](https://openev.al/docs/rubrics/) |
|
|
107
|
+
| Controlled workspaces | Readable files, pinned Git inputs, preparation | [Workspaces](https://openev.al/docs/workspaces/) |
|
|
108
|
+
| Small batches | Eval, model, repetition, and cost controls | [Running](https://openev.al/docs/running/) |
|
|
109
|
+
| Transparent scores | Equal eval weights; bounds for unresolved checks | [Scoring](https://openev.al/docs/scoring/) |
|
|
110
|
+
| Evidence inspection | Sessions, tool results, artifacts, and citations | [Evidence](https://openev.al/docs/evidence/) |
|
|
111
|
+
| Rejudging | New judgments from retained, immutable recordings | [Evidence](https://openev.al/docs/evidence/#revise) |
|
|
65
112
|
|
|
66
|
-
|
|
67
|
-
|
|
68
|
-
bunx --bun @hona/openeval plan
|
|
69
|
-
bunx --bun @hona/openeval run
|
|
70
|
-
bunx --bun @hona/openeval view
|
|
71
|
-
```
|
|
113
|
+
<details>
|
|
114
|
+
<summary><strong>Inspect a judgment and its evidence</strong></summary>
|
|
72
115
|
|
|
73
|
-
|
|
74
|
-
any command at another benchmark, or `--port <port>` for another viewer port.
|
|
116
|
+

|
|
75
117
|
|
|
76
|
-
|
|
77
|
-
separate result. Repeat `--only-eval`, `--only-model`, or `--only-repetition` to
|
|
78
|
-
execute a small scope while retaining the full aggregate. `--max-cost <usd>`
|
|
79
|
-
sets a scheduling budget for the invocation. Run `openeval --help` for commands.
|
|
118
|
+
</details>
|
|
80
119
|
|
|
81
|
-
|
|
82
|
-
|
|
83
|
-
The viewer provides metric drilldowns and recorded candidate and judge sessions.
|
|
84
|
-
Elapsed time measures active execution intervals, counting overlapping work once.
|
|
120
|
+
<details>
|
|
121
|
+
<summary><strong>Watch candidate and judge work in the live queue</strong></summary>
|
|
85
122
|
|
|
86
|
-
|
|
123
|
+

|
|
87
124
|
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
|
|
125
|
+
</details>
|
|
126
|
+
|
|
127
|
+
## Use the SDK
|
|
91
128
|
|
|
92
129
|
```ts
|
|
93
|
-
import
|
|
130
|
+
import { runBenchmark } from "@hona/openeval";
|
|
94
131
|
|
|
95
|
-
|
|
96
|
-
|
|
97
|
-
|
|
132
|
+
await runBenchmark("./my-benchmark", {
|
|
133
|
+
onlyEvals: ["ask-dialect"],
|
|
134
|
+
onlyRepetitions: [1],
|
|
135
|
+
});
|
|
98
136
|
```
|
|
99
137
|
|
|
100
|
-
|
|
101
|
-
containers receive project inputs, never judge rubrics or evaluator storage.
|
|
102
|
-
Candidates finish naturally, meet an opted-in irreversible judge decision, or
|
|
103
|
-
time out after at most 45 minutes. Finalized recordings remain immutable.
|
|
138
|
+
## Write evals with an agent
|
|
104
139
|
|
|
105
|
-
|
|
140
|
+
Use the public [Eval Writing skill](https://github.com/Hona/openeval/tree/main/.opencode/skills/eval-writing)
|
|
141
|
+
to turn a real failure into an eval, review a rubric, or investigate misleading
|
|
142
|
+
scores. It guides an agent through concrete false-pass/false-failure examples,
|
|
143
|
+
accepted alternatives, evidence requirements, and human-reviewed calibration.
|
|
106
144
|
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
110
|
-
await runBenchmark("./my-benchmark");
|
|
111
|
-
```
|
|
145
|
+
Copy the whole `.opencode/skills/eval-writing/` directory, including `references/`,
|
|
146
|
+
into the same path in your project. For global use, copy it to
|
|
147
|
+
`~/.config/opencode/skills/eval-writing/`. Then run **`/eval-writing`** in OpenCode.
|
|
112
148
|
|
|
113
|
-
|
|
114
|
-
and
|
|
115
|
-
`/view`, and `/session`.
|
|
149
|
+
> Use eval-writing to review this task and rubric. Show me the strongest false
|
|
150
|
+
> pass and false failure, then propose the smallest improvement.
|
|
116
151
|
|
|
117
|
-
|
|
152
|
+
The skill includes a framework-neutral workflow, fictional coaching examples,
|
|
153
|
+
an OpenEval-specific reference, and a broad public-research guide.
|
|
118
154
|
|
|
119
|
-
|
|
120
|
-
bun install --frozen-lockfile
|
|
121
|
-
bun run typecheck
|
|
122
|
-
bun test
|
|
123
|
-
bun run release:pack
|
|
124
|
-
bun run release:verify
|
|
125
|
-
```
|
|
155
|
+
## Develop
|
|
126
156
|
|
|
127
|
-
|
|
128
|
-
|
|
129
|
-
|
|
157
|
+
| Command | Purpose |
|
|
158
|
+
| --- | --- |
|
|
159
|
+
| `bun run site:dev` | Landing page and docs with hot reload on port 4176 |
|
|
160
|
+
| `bun run site:build && bun run site:verify` | Prerender pages and verify links and starter files |
|
|
161
|
+
| `bun run typecheck && bun test` | Local SDK checks; no live models |
|
|
162
|
+
| `bun run release:pack && bun run release:verify` | Verify the actual npm archive in a separate consumer |
|
|
130
163
|
|
|
131
|
-
|
|
132
|
-
|
|
133
|
-
third-party license notices.
|
|
164
|
+
- [Release guide](https://github.com/Hona/openeval/blob/main/RELEASING.md) · [Judge protocol](JUDGING.md)
|
|
165
|
+
- MIT licensed. The viewer includes upstream third-party license notices.
|
package/package.json
CHANGED
|
@@ -28,7 +28,7 @@ const positive = (value: unknown, fallback: number, label: string) => {
|
|
|
28
28
|
export const modelRef = (value: unknown): ModelRef => {
|
|
29
29
|
if (
|
|
30
30
|
typeof value !== "string" ||
|
|
31
|
-
!/^[\w.-]+\/[^\s/#]+(?:#[\w.-]+)?$/.test(value)
|
|
31
|
+
!/^[\w.-]+\/[^\s/#]+(?:\/[^\s/#]+)*(?:#[\w.-]+)?$/.test(value)
|
|
32
32
|
)
|
|
33
33
|
throw new Error(`Invalid model reference: ${String(value)}`);
|
|
34
34
|
return value as ModelRef;
|