evals-lab 0.3.0 → 0.4.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,42 @@ Newest first. Read a release's **Upgrade notes** before installing it: an
5
5
  upgrade can rewrite what the lab keeps in your data directory, and an older
6
6
  version cannot always read it back.
7
7
 
8
+ ## 0.4.1
9
+
10
+ ### Changed
11
+
12
+ - The README is shorter, and describes the whole lab rather than one use
13
+ of it.
14
+
15
+ ## 0.4.0
16
+
17
+ ### Upgrade notes
18
+
19
+ - **Copy your data directory before upgrading** (see "Where your data is
20
+ kept" in the README). A dataset 0.4.0 saves is dataset version 7, which
21
+ 0.3.x cannot read. To go back, reinstall 0.3.0 and restore the copy.
22
+ - Stored datasets are not rewritten on upgrade: 0.4.0 reads a version-6
23
+ dataset as version 7, and writes version 7 the next time it is saved.
24
+ - Export JSON writes dataset version 7, which 0.3.x refuses to import.
25
+ Import reads versions 1 to 7.
26
+
27
+ ### Added
28
+
29
+ - Re-run, in a History row's menu and in Results' run menu: a new run of the
30
+ same pipeline over the same file revisions and the same dataset version.
31
+ It is refused while the run is still going, or when one of the files it
32
+ read is gone. The new run says which run it re-ran, and links back to it
33
+ in Results. Undo cancels it while it is still queued.
34
+ - A dataset carries four new sections, `scoring`, `grader`, `every` and
35
+ `run`: how it scores, which model grades it, and metrics for every item
36
+ and for the whole run. Edit raw… shows them. Runs do not read them yet.
37
+ A dataset from 0.3.x has them as scored All, by the lab's grader, with no
38
+ metrics of its own.
39
+
40
+ ### Changed
41
+
42
+ - Resume is in a History row's menu, beside Re-run.
43
+
8
44
  ## 0.3.0
9
45
 
10
46
  ### Upgrade notes
package/README.md CHANGED
@@ -1,10 +1,15 @@
1
1
  # evals-lab
2
2
 
3
- A browser bench for grading what a model answers: two wordings side by side
4
- over a set of files, scored against answers fixed in advance.
3
+ A lab in the browser for evaluating prompts and models. A pipeline takes items
4
+ from a Source — uploaded files, a Power Automate flow's run history, or a
5
+ prompt alone — and sends each to one or more targets, every target a model
6
+ with its own prompt. Its evals score the replies with metrics — text and JSON
7
+ checks, item counts, agreement with production's reply, or a grading model —
8
+ and a dataset holds each item's expected answers. Every run is kept in
9
+ History, and every prompt version in the Library.
5
10
 
6
- This is an early test release. It runs on your own machine and talks to the
7
- models you point it at.
11
+ An early test release. It runs on your machine and talks only to the models
12
+ and services you connect it to.
8
13
 
9
14
  ## Install and run
10
15
 
@@ -13,31 +18,23 @@ npm install -g evals-lab
13
18
  evals-lab
14
19
  ```
15
20
 
16
- Then open <http://localhost:8080>. A new lab opens on two demo datasets, each
17
- with its images and a starter pipeline. The first start takes a little longer
18
- than later ones while it installs them.
21
+ Open <http://localhost:8080>. A new lab starts with two demo datasets and
22
+ their pipelines; the first start takes longer while it installs them. Without
23
+ installing: `npx evals-lab`.
19
24
 
20
- Without installing: `npx evals-lab`.
21
-
22
- ## What it needs
23
-
24
- | | | Windows | Linux |
25
- |---|---|---|---|
26
- | **Node.js** 22.18 or newer | runs the lab's worker | you have it if you have `npm` | same |
27
- | **Python** 3.10 or newer | runs the lab's server | `winget install Python.Python.3.13` | your distribution's `python3` |
28
- | **ImageMagick** | sizes images before a model is sent them, and makes thumbnails | `winget install ImageMagick.ImageMagick` | your distribution's `imagemagick` |
29
-
30
- `evals-lab` says which of these it could not find. It starts without
31
- ImageMagick, but a run that sends images needs it. Open a new terminal after
32
- installing one, so it is on your `PATH`.
25
+ | Needs | Windows | Linux |
26
+ |---|---|---|
27
+ | **Node.js** 22.18+ | comes with `npm` | same |
28
+ | **Python** 3.10+ | `winget install Python.Python.3.13` | your distribution's `python3` |
29
+ | **ImageMagick**, for runs that send images | `winget install ImageMagick.ImageMagick` | your distribution's `imagemagick` |
33
30
 
34
- macOS should work the same way and has not been tried.
31
+ `evals-lab` names anything it cannot find. Open a new terminal after
32
+ installing one. macOS should work and is untried.
35
33
 
36
34
  ## Where your data is kept
37
35
 
38
- Datasets, pipelines, Setup profiles (API keys among them), prompts, uploaded
39
- files and runs are kept on your machine, in one directory, so the lab opens
40
- as you left it:
36
+ Datasets, pipelines, Setup profiles (with their API keys), prompts, uploaded
37
+ files and runs are kept in one directory, which upgrades leave alone:
41
38
 
42
39
  | | |
43
40
  |---|---|
@@ -45,8 +42,7 @@ as you left it:
45
42
  | Linux | `~/.local/share/evals-lab` (or `$XDG_DATA_HOME/evals-lab`) |
46
43
  | macOS | `~/Library/Application Support/evals-lab` |
47
44
 
48
- Upgrading or reinstalling the package leaves it alone. Delete the directory
49
- to start again from a new lab.
45
+ Delete it to start again from a new lab.
50
46
 
51
47
  ## Options
52
48
 
@@ -61,125 +57,60 @@ evals-lab [--port <n>] [--host <address>] [--data-dir <dir>]
61
57
  | `--data-dir` | `DATA_DIR` | the directory above |
62
58
  | | `LAB_PASSWORD` | none |
63
59
  | | `OLLAMA_URL` | `http://127.0.0.1:11434` |
64
- | | `M365_CLIENT_ID`, `M365_TENANT` | none; see Power Automate below |
65
60
  | | `PYTHON` | `py -3`, `python` or `python3`, whichever is found |
66
61
 
67
- `OLLAMA_URL` is only the connection a new lab starts with. Every other
68
- connection — an OpenAI-compatible endpoint, Anthropic, llama.cpp — is a Setup
69
- profile you add in the page.
62
+ `--host 0.0.0.0` makes the lab reachable from your network and is refused
63
+ without `LAB_PASSWORD`. Never put it on the open internet: the lab holds your
64
+ API keys and forwards requests to any address a profile names.
65
+
66
+ ## Models
70
67
 
71
- ## Power Automate (Microsoft 365)
68
+ A model is a Setup profile: **Setup › Add profile**, then pick its type —
69
+ Ollama (local or cloud), OpenAI-compatible, Anthropic, llama.cpp or an HTTP
70
+ endpoint — and give its address, model and key. **Test connection** checks
71
+ it. An Ollama profile left without an address uses `OLLAMA_URL`.
72
72
 
73
- The lab can read a Power Automate cloud flow, and its run history, from your
74
- organisation's Microsoft 365 tenant, then grade the flow's HTTP step against
75
- the same model and prompt. The sign-in stays in your browser tab: the lab's
76
- server never receives a Microsoft token.
73
+ ## Power Automate
77
74
 
78
- It needs a work or school account. Microsoft retired cloud flows for personal
79
- accounts (outlook.com, hotmail.com) in July 2025, so a personal account cannot
80
- be used. For a small business, a trial or a Microsoft 365 Developer Program
81
- tenant, follow the same steps as a company.
75
+ The lab can read a Power Automate cloud flow and its run history from your
76
+ organisation's Microsoft 365 tenant, then grade the flow's HTTP step. It needs
77
+ a work or school account; personal accounts (outlook.com, hotmail.com) have no
78
+ cloud flows. A small business, a trial or a Microsoft 365 Developer Program
79
+ tenant follows the same steps. The sign-in stays in your browser tab.
82
80
 
83
- ### Register an app in your tenant
81
+ ### Register an app
84
82
 
85
- An administrator of the tenant registers it once, at
86
- <https://entra.microsoft.com>:
83
+ A tenant administrator does this once, at <https://entra.microsoft.com>:
87
84
 
88
85
  1. **App registrations › New registration.**
89
86
  - Supported account types: *Accounts in this organizational directory only*.
90
- - Redirect URI: platform **Single-page application**, value `http://localhost/`.
91
- Microsoft ignores the port on a `localhost` redirect, so this covers
92
- `--port` too. A lab served on another address needs that address here
93
- as well, over `https`.
94
- 2. **API permissions › Add a permission › APIs my organization uses.** Search
95
- for *Microsoft Flow Service* and add the delegated permission
96
- **Flows.Read.All**. If the service isn't listed, open
97
- <https://make.powerautomate.com> once, then search again.
98
- 3. **Grant admin consent** for the tenant.
99
- 4. No client secret or certificate is needed: the lab signs in as a public
100
- client.
101
-
102
- ### Tell the lab about it
103
-
104
- In the lab, open **Setup › Connections**, then use the Microsoft 365 row's
105
- menu › **Configure…**. Enter:
106
-
107
- | | |
108
- |---|---|
109
- | Client ID | the app's *Application (client) ID* |
110
- | Tenant | the *Directory (tenant) ID*, or one of the tenant's domains |
111
-
112
- Fill in Tenant for an app registered as above. Left empty, it means any
113
- organisation's tenant, which only a multitenant app accepts: Microsoft
114
- refuses the sign-in of a single-tenant app with `AADSTS50194`.
115
-
116
- Both are saved as you type, and neither is secret. **Sign in** from the same
117
- menu, then add a Source in the Sources tab with **Add source › Power Automate
118
- workflow**.
119
-
120
- To set the app for the whole machine instead, start the lab with
121
- `M365_CLIENT_ID` and `M365_TENANT` set. Setup then shows those values and
122
- can't change them.
123
-
124
- Open the lab at `localhost`, the address `evals-lab` prints. A sign-in begun
125
- at `127.0.0.1` reloads the page at `localhost`, the address registered above,
126
- and carries on from there. Anything open in the tab, such as an Add source
127
- dialog, does not carry over.
128
-
129
- ## Google Drive
130
-
131
- **Not usable yet.** The lab can sign in to Google Drive, but nothing reads
132
- from Drive yet: a Source made from a Drive folder is still to come. Until it
133
- arrives, there is no reason to set up a Google app, and the steps below are
134
- provisional. They haven't been tried against a real Google project.
135
-
136
- When it arrives, the sign-in will be held by the lab's server, on your
137
- machine. It will reach only the files you pick: the lab asks for
138
- `drive.file` access, never your whole Drive.
139
-
140
- To sign in, open **Setup › Connections**, then use the Google Drive row's
141
- menu › **Sign in**. If the row says *Not Configured*, give the lab a Google
142
- app first, once:
143
-
144
- 1. At <https://console.cloud.google.com>, make a project. Enable the
145
- **Google Drive API** and the **Google Picker API**.
146
- 2. **OAuth consent screen:** choose *External*. Give the app a name and your
147
- email, add the scope `.../auth/drive.file`, then **Publish app**. It needs
148
- no review from Google, because that scope is non-sensitive. Left in
149
- *Testing*, a sign-in lapses after 7 days.
150
- 3. **Credentials › Create credentials › OAuth client ID**, of type
151
- **Desktop app**. Download its JSON file.
152
- 4. **Credentials › Create credentials › API key.** Restrict it to the Google
153
- Picker API.
154
- 5. In the lab, use the Google Drive row's menu › **Configure…**, then
155
- **Load client file…**, and choose the file from step 3. Paste the key
156
- from step 4 into **API key**.
157
-
158
- A Desktop app needs no redirect registered: Google sends the browser back to
159
- any port on this machine. A lab you reach at its own address, from another
160
- machine, needs a *Web application* client instead. Choose that type in
161
- Configure…, then register the **Redirect URI** it shows with the client.
162
-
163
- Everything saves as you type. The client's secret is kept by the lab and
164
- never shown again. The window says only that one is saved.
165
-
166
- ## Reaching it from another machine
167
-
168
- By default only your own machine can reach the lab. That is deliberate: the
169
- lab's server forwards requests to whatever address a Setup profile names, so
170
- anyone who can reach the lab can reach those addresses through it, and the
171
- lab holds your API keys.
172
-
173
- `--host 0.0.0.0` makes it reachable from your network, and is refused unless
174
- `LAB_PASSWORD` is set. Do not put it on the open internet.
87
+ - Redirect URI: **Single-page application**, `http://localhost/`. This
88
+ covers any `--port`. A lab served elsewhere needs its own address here
89
+ too, over `https`.
90
+ 2. **API permissions › Add a permission › APIs my organization uses ›
91
+ Microsoft Flow Service**, delegated **Flows.Read.All**. If it isn't listed,
92
+ open <https://make.powerautomate.com> once and search again.
93
+ 3. **Grant admin consent.** No client secret is needed.
94
+
95
+ ### Connect the lab
96
+
97
+ 1. **Setup › Connections**, the Microsoft 365 row's menu › **Configure…**:
98
+ - Client ID: the app's *Application (client) ID*.
99
+ - Tenant: the *Directory (tenant) ID* or one of its domains. Left empty, a
100
+ single-tenant app's sign-in fails with `AADSTS50194`.
101
+ 2. The same menu › **Sign in**, with the lab open at `localhost`. A sign-in
102
+ from `127.0.0.1` reloads the page there and loses any open dialog.
103
+ 3. **Library › Sources › Add source › Power Automate workflow.**
104
+
105
+ Or set `M365_CLIENT_ID` and `M365_TENANT` when starting the lab, for the
106
+ whole machine; Setup then shows them read-only.
175
107
 
176
108
  ## Upgrading and removing
177
109
 
178
- Read `CHANGELOG.md` in the package, or on
179
- <https://www.npmjs.com/package/evals-lab?activeTab=code>, before upgrading.
180
- Its **Upgrade notes** say when a release rewrites what is in your data
181
- directory in a form an older version cannot read; copy the directory first
182
- if you may want to go back.
110
+ Read `CHANGELOG.md` first, in the package or on
111
+ <https://www.npmjs.com/package/evals-lab?activeTab=code>. When its
112
+ **Upgrade notes** say a release rewrites your data, copy the data directory
113
+ before upgrading if you may want to go back.
183
114
 
184
115
  ```bash
185
116
  npm install -g evals-lab@latest # upgrade
@@ -188,13 +119,9 @@ npm uninstall -g evals-lab # remove; your data directory stays
188
119
 
189
120
  ## Licence
190
121
 
191
- Copyright 2026 Grainbox Limited. Licensed under the
122
+ Copyright 2026 Grainbox Limited, under the
192
123
  [PolyForm Internal Use License 1.0.0](https://polyformproject.org/licenses/internal-use/1.0.0)
193
- (`LICENSE.md` in the package): you and your company may use it, and change it,
194
- for your own internal business. You may not give it, or a changed copy, to
195
- anyone else, sell it, or offer it as a service.
196
-
197
- ## The demo images
198
-
199
- Every image in the demo datasets is in the public domain. Each has its line in
200
- `lab/demo/CREDITS.md` inside the package.
124
+ (`LICENSE.md`): you and your company may use and change it for your own
125
+ internal business, but not give it, or a changed copy, to anyone else, sell
126
+ it, or offer it as a service. The demo images are public domain; each is
127
+ credited in `lab/demo/CREDITS.md`.
package/lab/VERSION CHANGED
@@ -1 +1 @@
1
- 0.3.0 (2026.10.02-355)
1
+ 0.4.1 (2026.10.03-363)
@@ -983,22 +983,37 @@ const SLOTS = ["content", "target", "responses"];
983
983
 
984
984
 
985
985
 
986
-
986
+
987
987
 
988
-
988
+
989
+
990
+
991
+
992
+
993
+
994
+
995
+
996
+
997
+
989
998
 
990
999
 
991
1000
 
992
- /** The dataset body's version: 6 marks itself; 5 named its Source; 4 and
993
- earlier, neither. */
994
- const DATASET_BODY_VERSION = 6 ;
1001
+ /** An eval group's scoring (docs/pipeline-model.md §17). */
1002
+
1003
+
1004
+
1005
+
1006
+
1007
+ /** The dataset body's version: 7 is an eval group (§17); 6 marks itself; 5
1008
+ named its Source; 4 and earlier, neither. */
1009
+ const DATASET_BODY_VERSION = 7 ;
995
1010
 
996
1011
  /** The metrics whose Ignore case version 6 made mean what it says for a
997
1012
  reply read as a list, as well as one read as text. */
998
1013
  const CASE_FOLDING = ["contains", "contains-all", "contains-any"];
999
1014
 
1000
1015
  /**
1001
- * [body] as this version of a dataset (6), from any earlier one. Version 1
1016
+ * [body] as this version of a dataset (7), from any earlier one. Version 1
1002
1017
  * held `imageCases`, each with `minTags`/`maxTags`, and the parser's
1003
1018
  * `replays` and `conformance` (now fixtures/replays.json beside the checks).
1004
1019
  * Version 2 held `rules`, which clean a job's answer and so belong to the
@@ -1010,7 +1025,9 @@ const CASE_FOLDING = ["contains", "contains-all", "contains-any"];
1010
1025
  * `note`, `traits` go, and the body names no Source yet. Version 5 matched
1011
1026
  * a list's items ignoring case whatever a Contains metric's Ignore case
1012
1027
  * said, so each of its Contains metrics says Ignore case (`caseOfV5`), and
1013
- * the body says its version. Every reader of a
1028
+ * the body says its version. Version 6 is a group's cases alone: version 7
1029
+ * adds its scoring, grader, Every item and Whole run (`groupOfV6`), and a
1030
+ * stored row is read that way rather than rewritten. Every reader of a
1014
1031
  * body calls this: the runner, the page, a run's kept copy. A reader that
1015
1032
  * needs a version-2 body's rules -- to upgrade a pipeline graded against
1016
1033
  * it -- takes them first (`datasetRules`). Pure: the same body gives the
@@ -1021,14 +1038,26 @@ function upgradeDatasetBody (body ) {
1021
1038
  if (!isObj(body)) return body;
1022
1039
  if (body.version === DATASET_BODY_VERSION) return body;
1023
1040
  const b = body ;
1041
+ // A body saying any other version is one this lab does not read.
1042
+ if ("version" in b) return b.version === 6 ? groupOfV6(b) : body;
1024
1043
  // A body that names its Source, even as null, is version 5.
1025
1044
  if ("source" in b) {
1026
- return { version: DATASET_BODY_VERSION, ...b, ...(Array.isArray(b.cases) ? { cases: b.cases.map(caseOfV5) } : {}) };
1045
+ return groupOfV6({ ...b, ...(Array.isArray(b.cases) ? { cases: b.cases.map(caseOfV5) } : {}) });
1027
1046
  }
1028
1047
  const raw = Array.isArray(b.cases) ? b.cases : Array.isArray(b.imageCases) ? b.imageCases : null;
1029
1048
  // Not a body of any version -- a copy kept while a dataset was an overlay.
1030
1049
  if (!raw) return body;
1031
- return { version: DATASET_BODY_VERSION, source: null, cases: raw.map(caseOfV4).map(caseOfV5) };
1050
+ return groupOfV6({ source: null, cases: raw.map(caseOfV4).map(caseOfV5) });
1051
+ }
1052
+
1053
+ /** A version-6 body as a version-7 eval group: scored All, the lab's
1054
+ grader, and no metrics of its own for every item or the whole run --
1055
+ what a Metrics eval naming the dataset with none of its own graded, which
1056
+ is what every Graded set converted to. */
1057
+ function groupOfV6(b ) {
1058
+ const { version: _v, ...rest } = b;
1059
+ return { version: DATASET_BODY_VERSION, source: null, scoring: { mode: "all", threshold: null }, grader: null, every: [], run: [],
1060
+ ...rest } ;
1032
1061
  }
1033
1062
 
1034
1063
  /** A version-5 case as a version-6 one: each Contains metric says Ignore
@@ -3187,15 +3216,7 @@ EVAL_TYPES.metrics = {
3187
3216
  ? { ...t, grader: { id: ctx.grader.id, name: ctx.grader.name } } : t),
3188
3217
  wholeRun: t => t.over === "run",
3189
3218
  // Over the whole run: every metric over the replies together, or each alone.
3190
- verdict(t, ress, kind){
3191
- const run = runInput(ress, !(kind != null && OUTPUT_KINDS[kind]?.terms));
3192
- const metrics = readRun(t.metrics || [], run);
3193
- const s = scoreOf(metrics, t.mode, t.threshold);
3194
- const off = metrics.filter(r => !r.pass);
3195
- return { pass: s ? s.pass : true, detail: off.length ? off.map(r => `${r.label}: ${r.reason}`).join(" · ") : "all matched",
3196
- ran: run.replies.length,
3197
- checks: metrics.map((r, x) => ({ key: `m${x}`, label: r.label, want: "", pass: r.pass, ...(r.pass ? {} : { note: r.reason }) })) };
3198
- },
3219
+ verdict: (t, ress, kind) => wholeRunVerdict(t.metrics || [], ress, kind, t.mode, t.threshold),
3199
3220
  // Rule m{x} is the eval's own metric x, read in that place on every item.
3200
3221
  // A score stored before Metrics -- a Graded set's, read as its conversion --
3201
3222
  // has no readings of its own: its one metric's reading is the score's.
@@ -3296,6 +3317,38 @@ LEGACY_TESTS.single .toMetrics = t => {
3296
3317
  return { ...common(t), type: "metrics", mode: "all", threshold: null, grader: null, dataset: null, over: "run", metrics };
3297
3318
  };
3298
3319
 
3320
+ /** [list] read over a whole run's replies so far, scored the way [mode]
3321
+ says: a Metrics eval's verdict over the run, and an eval group's Whole run. */
3322
+ function wholeRunVerdict(list , ress , kind ,
3323
+ mode = "all", threshold = null) {
3324
+ const run = runInput(ress, !(kind != null && OUTPUT_KINDS[kind]?.terms));
3325
+ const metrics = readRun(list, run);
3326
+ const s = scoreOf(metrics, mode, threshold);
3327
+ const off = metrics.filter(r => !r.pass);
3328
+ return { pass: s ? s.pass : true, detail: off.length ? off.map(r => `${r.label}: ${r.reason}`).join(" · ") : "all matched",
3329
+ ran: run.replies.length,
3330
+ checks: metrics.map((r, x) => ({ key: `m${x}`, label: r.label, want: "", pass: r.pass, ...(r.pass ? {} : { note: r.reason }) })) };
3331
+ }
3332
+
3333
+ /**
3334
+ * An eval group's verdict on one item (docs/pipeline-model.md §17): its
3335
+ * Every item metrics and [kase]'s own -- the item's case in the group, or
3336
+ * null where it has none -- read of the reply under the group's scoring.
3337
+ * The reading a Metrics eval naming the group as its dataset gives, which
3338
+ * groups-check.js holds it to. Null where no metric had anything to read.
3339
+ */
3340
+ function readGroup(group , kase , res , more = {}) {
3341
+ return readMetrics([...group.every, ...(kase?.metrics ?? [])], metricInput(res, kase, more), graderCtx(group, more),
3342
+ group.scoring.mode, group.scoring.threshold);
3343
+ }
3344
+
3345
+ /** An eval group's Whole run verdict, from the replies so far: its `run`
3346
+ metrics under its scoring. Null for a group with no Whole run metrics,
3347
+ which has nothing to say of a run. */
3348
+ function readGroupRun(group , ress , kind = null) {
3349
+ return group.run.length ? wholeRunVerdict(group.run, ress, kind, group.scoring.mode, group.scoring.threshold) : null;
3350
+ }
3351
+
3299
3352
  /** The grader a Metrics eval names, reached through what the runner hands it. */
3300
3353
  function graderCtx(t , more ) {
3301
3354
  const ask = isRef(t.grader) ? more.grader?.(t.grader) : undefined;
@@ -4768,6 +4821,24 @@ function validateEvals(ev , files
4768
4821
  }
4769
4822
  if (ev.cases != null && !Array.isArray(ev.cases)) bad.push("cases has to be a list");
4770
4823
  if (ev.source != null && !isRef(ev.source)) bad.push("a dataset names its Source as { id, name }, or null");
4824
+ // An eval group's own sections (version 7): what a Metrics eval held.
4825
+ if (ev.scoring != null) {
4826
+ const sc = ev.scoring;
4827
+ if (!isObj(sc) || !SCORING_MODES.includes(sc.mode)) bad.push(`a group is scored ${SCORING_MODES.join(" or ")}`);
4828
+ else if (sc.mode === "weighted" && typeof sc.threshold !== "number") bad.push("a group scored in points needs Pass at: the points an item has to reach");
4829
+ else if (sc.threshold != null && typeof sc.threshold !== "number") bad.push("a group's Pass at has to be a number");
4830
+ }
4831
+ if (ev.grader != null && !isRef(ev.grader)) bad.push("a group names its grader as { id, name }, or null");
4832
+ if (ev.every != null) metricsProblems(ev.every, "Every item", bad);
4833
+ if (ev.run != null) {
4834
+ metricsProblems(ev.run, "Whole run", bad);
4835
+ // Read once over every reply: nothing that reads one item, or asks a grader.
4836
+ for (const m of Array.isArray(ev.run) ? ev.run : []) {
4837
+ const entry = isObj(m) ? METRICS[m.type] : undefined;
4838
+ if (entry?.graded) bad.push(`Whole run: ${entry.label} is model-graded, so it cannot read a whole run`);
4839
+ else if (entry?.perItem) bad.push(`Whole run: ${entry.label} reads one item at a time, so it cannot read a whole run`);
4840
+ }
4841
+ }
4771
4842
  if (bad.length) return bad;
4772
4843
 
4773
4844
  const graded = ev.cases || [];
@@ -4915,7 +4986,7 @@ registerMetrics({ registerKinds });
4915
4986
  export {
4916
4987
  TOKEN_DEFAULTS, TOKEN_TYPES, tokenMapping, tokenNames, resolvePrompt, tokenSet, textPrompt, words,
4917
4988
  loopReplyError, preparedSize,
4918
- termIn, forbiddenIn, readCase, gradedSetFrom, caseItem, caseMetrics, metricLines, upgradeDatasetBody, datasetRules, upgradeResults, emptyTally, addToTally, isHostedUrl,
4989
+ termIn, forbiddenIn, readCase, gradedSetFrom, caseItem, caseMetrics, metricLines, upgradeDatasetBody, readGroup, readGroupRun, datasetRules, upgradeResults, emptyTally, addToTally, isHostedUrl,
4919
4990
  SEED, REPLY_TOKENS_BEFORE, ANTHROPIC_MAX_TOKENS, asNumber, pinReplyTokens, mappingsFor,
4920
4991
  CONNECTION_TYPES, HTTP_APIS, httpApiFor, httpRequestOf, httpReplyOf, readFlowReply, WdlError, localAnswer, typeOf, profileType, convertProfile, splitOllama, DECODING_KEYS, apiBase,
4921
4992
  EDGE_448, budgetLabel, IMAGE_FORMATS, encoderQuality,
package/lab/run-evals.js CHANGED
@@ -296,8 +296,8 @@ function readDataset(file) {
296
296
  broken(`${file}: ${e.message}`);
297
297
  }
298
298
  if (isObject(doc) && doc.format === "evals-lab/dataset") {
299
- if (![1, 2, 3, 4, 5].includes(doc.version)) {
300
- broken(`${file} is version ${JSON.stringify(doc.version)}, and this reads versions 1 to 5`);
299
+ if (![1, 2, 3, 4, 5, 6, 7].includes(doc.version)) {
300
+ broken(`${file} is version ${JSON.stringify(doc.version)}, and this reads versions 1 to 7`);
301
301
  }
302
302
  doc = isObject(doc.dataset) ? doc.dataset.body : null;
303
303
  }