evals-lab 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,57 @@ Newest first. Read a release's **Upgrade notes** before installing it: an
5
5
  upgrade can rewrite what the lab keeps in your data directory, and an older
6
6
  version cannot always read it back.
7
7
 
8
+ ## 0.5.0
9
+
10
+ ### Upgrade notes
11
+
12
+ - **Copy your data directory before upgrading** (see "User data" in the
13
+ README). 0.5.0 keeps versions of each eval group and each run's verdict
14
+ beside what 0.4.x kept, and saves pipelines as pipeline version 13, which
15
+ 0.4.x cannot read. To go back, reinstall 0.4.1 and restore the copy.
16
+ - Pipelines are read as version 13 and saved as it at their next edit. An
17
+ eval that graded by a dataset becomes a link to that eval group, or a
18
+ group of the pipeline's own when it had metrics of its own.
19
+ - Export JSON writes an eval group file, which 0.4.x refuses to import.
20
+ Import still reads dataset files of versions 1 to 7.
21
+
22
+ ### Added
23
+
24
+ - `evals-lab run <bundle>` runs a pipeline's evals with no lab, for CI: the
25
+ pipeline's menu › Export for CI… downloads the bundle. It exits 0 when
26
+ every eval passes, 1 when one fails, 2 when the run is incomplete and 3
27
+ when it cannot run, and writes a JSON report (`--json`), JUnit XML
28
+ (`--junit`) and a Markdown summary (`--summary`). See "In CI" in the
29
+ README.
30
+ - `evals-lab run --lab <url> --pipeline <name> --wait` runs a lab's own
31
+ pipeline there, so the run is in its History, and exits with its verdict.
32
+ - Each Target profile has a slug in Setup. CI reads the profile's key from
33
+ `EVALSLAB_API_KEY_<SLUG>`.
34
+ - An eval group keeps its versions, and a run keeps the version it graded
35
+ with: Re-run grades with the same one.
36
+ - A run's row in History keeps its verdict.
37
+
38
+ ### Changed
39
+
40
+ - A Target profile's key is never sent back to the page: Setup shows it as
41
+ Set, and typing replaces it.
42
+ - Tables look the same on every tab, every named column sorts, and lists
43
+ can select many rows at once.
44
+ - A dialog's buttons sit at its bottom right, and one that edits in place
45
+ closes with Close, or Save & Close once something changed.
46
+
47
+ ### Fixed
48
+
49
+ - History skipped a run queued at the same moment as another when it loaded
50
+ older runs.
51
+
52
+ ## 0.4.1
53
+
54
+ ### Changed
55
+
56
+ - The README is shorter, and describes the whole lab rather than one use
57
+ of it.
58
+
8
59
  ## 0.4.0
9
60
 
10
61
  ### Upgrade notes
package/README.md CHANGED
@@ -1,10 +1,19 @@
1
1
  # evals-lab
2
2
 
3
- A browser bench for grading what a model answers: two wordings side by side
4
- over a set of files, scored against answers fixed in advance.
3
+ A lab for evaluating prompts and models in a friendly UI. A pipeline takes items from a Source (files, Power Automate runs, or a prompt) and sends each to one or more Targets, then runs evaluates Responses. Allows for quick iterations to compare changes in prompts and settings.
5
4
 
6
- This is an early test release. It runs on your own machine and talks to the
7
- models you point it at.
5
+ Model connections can be local (Ollama, llama.cpp) or online (Anthropic, OpenAI, any compatible HTTP endpoint).
6
+
7
+ The pipeline can be scored against metrics. The metrics include parsing checks, item counts, prod reply verification, grading models, eval groups (datasets), and more.
8
+
9
+ Runs are kept in history which can be exported.
10
+
11
+ Currently in early release. Runs locally and doesn't send your data anywhere else.
12
+
13
+ Example use-cases:
14
+ - Providing images to an LLM and comparing responses for accuracy
15
+ - Checking which model can meet evals in a Power Automate workflow
16
+ - Determining which prompt gets the highest score in metrics
8
17
 
9
18
  ## Install and run
10
19
 
@@ -13,31 +22,22 @@ npm install -g evals-lab
13
22
  evals-lab
14
23
  ```
15
24
 
16
- Then open <http://localhost:8080>. A new lab opens on two demo datasets, each
17
- with its images and a starter pipeline. The first start takes a little longer
18
- than later ones while it installs them.
19
-
20
- Without installing: `npx evals-lab`.
25
+ Open <http://localhost:8080>. A new lab starts with two demo datasets and
26
+ their pipelines. Note the first start takes longer while it installs them. Without
27
+ installing: `npx evals-lab`.
21
28
 
22
- ## What it needs
23
-
24
- | | | Windows | Linux |
25
- |---|---|---|---|
26
- | **Node.js** 22.18 or newer | runs the lab's worker | you have it if you have `npm` | same |
27
- | **Python** 3.10 or newer | runs the lab's server | `winget install Python.Python.3.13` | your distribution's `python3` |
28
- | **ImageMagick** | sizes images before a model is sent them, and makes thumbnails | `winget install ImageMagick.ImageMagick` | your distribution's `imagemagick` |
29
-
30
- `evals-lab` says which of these it could not find. It starts without
31
- ImageMagick, but a run that sends images needs it. Open a new terminal after
32
- installing one, so it is on your `PATH`.
29
+ | Needs | Windows | Linux |
30
+ |---|---|---|
31
+ | **Node.js** 22.18+ | comes with `npm` | same |
32
+ | **Python** 3.10+ | `winget install Python.Python.3.13` | your distribution's `python3` |
33
+ | **ImageMagick**, for runs that send images | `winget install ImageMagick.ImageMagick` | your distribution's `imagemagick` |
33
34
 
34
- macOS should work the same way and has not been tried.
35
+ `evals-lab` names anything it cannot find. Open a new terminal after
36
+ installing one. macOS should work and is untried.
35
37
 
36
- ## Where your data is kept
38
+ ## User data
37
39
 
38
- Datasets, pipelines, Setup profiles (API keys among them), prompts, uploaded
39
- files and runs are kept on your machine, in one directory, so the lab opens
40
- as you left it:
40
+ User data is kept in a single folder. Upgrades intend to keep user data. The folders are:
41
41
 
42
42
  | | |
43
43
  |---|---|
@@ -45,8 +45,7 @@ as you left it:
45
45
  | Linux | `~/.local/share/evals-lab` (or `$XDG_DATA_HOME/evals-lab`) |
46
46
  | macOS | `~/Library/Application Support/evals-lab` |
47
47
 
48
- Upgrading or reinstalling the package leaves it alone. Delete the directory
49
- to start again from a new lab.
48
+ Delete this folder to start again fresh.
50
49
 
51
50
  ## Options
52
51
 
@@ -61,125 +60,154 @@ evals-lab [--port <n>] [--host <address>] [--data-dir <dir>]
61
60
  | `--data-dir` | `DATA_DIR` | the directory above |
62
61
  | | `LAB_PASSWORD` | none |
63
62
  | | `OLLAMA_URL` | `http://127.0.0.1:11434` |
64
- | | `M365_CLIENT_ID`, `M365_TENANT` | none; see Power Automate below |
65
63
  | | `PYTHON` | `py -3`, `python` or `python3`, whichever is found |
66
64
 
67
- `OLLAMA_URL` is only the connection a new lab starts with. Every other
68
- connection — an OpenAI-compatible endpoint, Anthropic, llama.cpp — is a Setup
69
- profile you add in the page.
65
+ `--host 0.0.0.0` makes the lab reachable from your network and needs `LAB_PASSWORD` to be set (otherwise it refuses connections).
66
+
67
+ ## In CI
68
+
69
+ A pipeline can run with no lab, from a repository. In the lab, the pipeline's
70
+ menu › **Export for CI…** downloads it as a bundle: unzip it into the
71
+ repository, then:
70
72
 
71
- ## Power Automate (Microsoft 365)
73
+ ```bash
74
+ EVALSLAB_API_KEY_ANTHROPIC=... evals-lab run evals/
75
+ ```
72
76
 
73
- The lab can read a Power Automate cloud flow, and its run history, from your
74
- organisation's Microsoft 365 tenant, then grade the flow's HTTP step against
75
- the same model and prompt. The sign-in stays in your browser tab: the lab's
76
- server never receives a Microsoft token.
77
+ ```text
78
+ evals-lab run <bundle-dir | pipeline.yaml> [--items <dir>] [--json <file|->]
79
+ [--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
80
+ ```
77
81
 
78
- It needs a work or school account. Microsoft retired cloud flows for personal
79
- accounts (outlook.com, hotmail.com) in July 2025, so a personal account cannot
80
- be used. For a small business, a trial or a Microsoft 365 Developer Program
81
- tenant, follow the same steps as a company.
82
+ | Option | |
83
+ |---|---|
84
+ | `--items` | the items the pipeline's Source reads; default the bundle's `items/` |
85
+ | `--json` | the JSON report; `-` for stdout |
86
+ | `--junit` | JUnit XML, for a CI test panel |
87
+ | `--summary` | the results as Markdown, added to the file: for GitHub, `"$GITHUB_STEP_SUMMARY"` |
88
+ | `--min-pass` | an eval passes when this share of its items pass |
89
+ | `--progress` | a line per item as it finishes |
90
+
91
+ The bundle holds no key. Each Target profile's key comes from
92
+ `EVALSLAB_API_KEY_` and its slug in Setup, in capitals with hyphens as
93
+ underscores: a profile slugged `anthropic` reads `EVALSLAB_API_KEY_ANTHROPIC`,
94
+ and `gpt-4o-mini` reads `EVALSLAB_API_KEY_GPT_4O_MINI`. A table of results
95
+ goes to stderr.
96
+
97
+ | Exit code | |
98
+ |---|---|
99
+ | `0` | every eval passed |
100
+ | `1` | an eval failed |
101
+ | `2` | incomplete: an item is missing, or did not run |
102
+ | `3` | broken or refused |
103
+
104
+ An item a case names that is not in the items is reported as Missing.
105
+ Images need ImageMagick, as in the lab.
106
+
107
+ ### GitHub Actions
108
+
109
+ ```yaml
110
+ name: Evals
111
+ on: pull_request
112
+
113
+ jobs:
114
+ evals:
115
+ runs-on: ubuntu-latest
116
+ steps:
117
+ - uses: actions/checkout@v5
118
+ - uses: actions/setup-node@v4
119
+ with:
120
+ node-version: 24
121
+ # Only when the items are images:
122
+ # - run: sudo apt-get update && sudo apt-get install -y imagemagick
123
+ - run: >
124
+ npx --yes evals-lab@0.5 run evals/
125
+ --junit evals.xml --json evals.json --summary "$GITHUB_STEP_SUMMARY"
126
+ env:
127
+ EVALSLAB_API_KEY_ANTHROPIC: ${{ secrets.ANTHROPIC_API_KEY }}
128
+ - uses: actions/upload-artifact@v4
129
+ if: always()
130
+ with:
131
+ name: evals
132
+ path: evals.*
133
+ ```
82
134
 
83
- ### Register an app in your tenant
135
+ A failing eval fails the job, and the results show on the run's page. Other
136
+ CI systems run the same command and read `--junit` in their test panel.
84
137
 
85
- An administrator of the tenant registers it once, at
86
- <https://entra.microsoft.com>:
138
+ ### Against a running lab
87
139
 
88
- 1. **App registrations › New registration.**
89
- - Supported account types: *Accounts in this organizational directory only*.
90
- - Redirect URI: platform **Single-page application**, value `http://localhost/`.
91
- Microsoft ignores the port on a `localhost` redirect, so this covers
92
- `--port` too. A lab served on another address needs that address here
93
- as well, over `https`.
94
- 2. **API permissions › Add a permission › APIs my organization uses.** Search
95
- for *Microsoft Flow Service* and add the delegated permission
96
- **Flows.Read.All**. If the service isn't listed, open
97
- <https://make.powerautomate.com> once, then search again.
98
- 3. **Grant admin consent** for the tenant.
99
- 4. No client secret or certificate is needed: the lab signs in as a public
100
- client.
101
-
102
- ### Tell the lab about it
103
-
104
- In the lab, open **Setup › Connections**, then use the Microsoft 365 row's
105
- menu › **Configure…**. Enter:
140
+ CI can run a lab's own pipeline there instead, so the run shows in that lab's
141
+ History and its keys never leave it:
106
142
 
107
- | | |
143
+ ```bash
144
+ LAB_PASSWORD=... evals-lab run --lab https://lab.example --pipeline "Demo 1" --wait
145
+ ```
146
+
147
+ ```text
148
+ evals-lab run --lab <url> --pipeline <name> [--wait] [--json <file|->]
149
+ [--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
150
+ ```
151
+
152
+ | Option | |
108
153
  |---|---|
109
- | Client ID | the app's *Application (client) ID* |
110
- | Tenant | the *Directory (tenant) ID*, or one of the tenant's domains |
111
-
112
- Fill in Tenant for an app registered as above. Left empty, it means any
113
- organisation's tenant, which only a multitenant app accepts: Microsoft
114
- refuses the sign-in of a single-tenant app with `AADSTS50194`.
115
-
116
- Both are saved as you type, and neither is secret. **Sign in** from the same
117
- menu, then add a Source in the Sources tab with **Add source › Power Automate
118
- workflow**.
119
-
120
- To set the app for the whole machine instead, start the lab with
121
- `M365_CLIENT_ID` and `M365_TENANT` set. Setup then shows those values and
122
- can't change them.
123
-
124
- Open the lab at `localhost`, the address `evals-lab` prints. A sign-in begun
125
- at `127.0.0.1` reloads the page at `localhost`, the address registered above,
126
- and carries on from there. Anything open in the tab, such as an Add source
127
- dialog, does not carry over.
128
-
129
- ## Google Drive
130
-
131
- **Not usable yet.** The lab can sign in to Google Drive, but nothing reads
132
- from Drive yet: a Source made from a Drive folder is still to come. Until it
133
- arrives, there is no reason to set up a Google app, and the steps below are
134
- provisional. They haven't been tried against a real Google project.
135
-
136
- When it arrives, the sign-in will be held by the lab's server, on your
137
- machine. It will reach only the files you pick: the lab asks for
138
- `drive.file` access, never your whole Drive.
139
-
140
- To sign in, open **Setup › Connections**, then use the Google Drive row's
141
- menu › **Sign in**. If the row says *Not Configured*, give the lab a Google
142
- app first, once:
143
-
144
- 1. At <https://console.cloud.google.com>, make a project. Enable the
145
- **Google Drive API** and the **Google Picker API**.
146
- 2. **OAuth consent screen:** choose *External*. Give the app a name and your
147
- email, add the scope `.../auth/drive.file`, then **Publish app**. It needs
148
- no review from Google, because that scope is non-sensitive. Left in
149
- *Testing*, a sign-in lapses after 7 days.
150
- 3. **Credentials › Create credentials › OAuth client ID**, of type
151
- **Desktop app**. Download its JSON file.
152
- 4. **Credentials › Create credentials › API key.** Restrict it to the Google
153
- Picker API.
154
- 5. In the lab, use the Google Drive row's menu › **Configure…**, then
155
- **Load client file…**, and choose the file from step 3. Paste the key
156
- from step 4 into **API key**.
157
-
158
- A Desktop app needs no redirect registered: Google sends the browser back to
159
- any port on this machine. A lab you reach at its own address, from another
160
- machine, needs a *Web application* client instead. Choose that type in
161
- Configure…, then register the **Redirect URI** it shows with the client.
162
-
163
- Everything saves as you type. The client's secret is kept by the lab and
164
- never shown again. The window says only that one is saved.
165
-
166
- ## Reaching it from another machine
167
-
168
- By default only your own machine can reach the lab. That is deliberate: the
169
- lab's server forwards requests to whatever address a Setup profile names, so
170
- anyone who can reach the lab can reach those addresses through it, and the
171
- lab holds your API keys.
172
-
173
- `--host 0.0.0.0` makes it reachable from your network, and is refused unless
174
- `LAB_PASSWORD` is set. Do not put it on the open internet.
175
-
176
- ## Upgrading and removing
177
-
178
- Read `CHANGELOG.md` in the package, or on
179
- <https://www.npmjs.com/package/evals-lab?activeTab=code>, before upgrading.
180
- Its **Upgrade notes** say when a release rewrites what is in your data
181
- directory in a form an older version cannot read; copy the directory first
182
- if you may want to go back.
154
+ | `--lab` | the lab's address; `LAB_PASSWORD` is sent when it is set |
155
+ | `--pipeline` | the pipeline, by its name or its id |
156
+ | `--wait` | wait for the run, and exit with the lab's verdict on it |
157
+
158
+ Without `--wait` the run's id is printed once it is queued, and the exit code
159
+ is `0`; `--json`, `--junit`, `--summary`, `--min-pass` and `--progress` need
160
+ `--wait`. The exit codes are the ones above, and a cancelled run is `2`.
161
+
162
+ ## Getting started
163
+
164
+ First set up a Setup profile: **Setup › Add profile**, then pick its type —
165
+ Ollama (local or cloud), OpenAI-compatible, Anthropic, llama.cpp or an HTTP
166
+ endpoint — and give its address, model and key. NB: An Ollama profile without an address uses `OLLAMA_URL`.
167
+
168
+ Set up your data source(s) in Library.
169
+
170
+ Then create a Pipeline to Run. Changes are auto-saved. Use the "Restore" menu to go back and undo any changes you don't want to keep.
171
+
172
+ ## Power Automate
173
+
174
+ The lab can read a Power Automate cloud flow and its run history from your
175
+ org's Microsoft 365 tenant then grade the flow's HTTP step. It needs
176
+ a work or school account. Personal accounts (outlook.com, hotmail.com) have no cloud flows. The sign-in data is kept in your browser to avoid it being stored elsewhere, the downside is you need to re-sign in if you close/reopen your web browser. The evals can keep running however, the sign-in is just to get the workflow data, which does persist in the evals-lab user data locally.
177
+
178
+ ### Register an app
179
+
180
+ A tenant administrator does this once, at <https://entra.microsoft.com>:
181
+
182
+ 1. **App registrations › New registration.**
183
+ - Supported account types: *Accounts in this organizational directory only*.
184
+ - Redirect URI: **Single-page application**, `http://localhost/`. This
185
+ covers any `--port`. A lab served elsewhere needs its own address here
186
+ too, over `https`.
187
+ 2. **API permissions › Add a permission › APIs my organization uses ›
188
+ Microsoft Flow Service**, delegated **Flows.Read.All**. If it isn't listed,
189
+ open <https://make.powerautomate.com> once and search again.
190
+ 3. **Grant admin consent.** No client secret is needed.
191
+
192
+ ### Connect the lab
193
+
194
+ 1. **Setup › Connections**, the Microsoft 365 row's menu › **Configure…**:
195
+ - Client ID: the app's *Application (client) ID*.
196
+ - Tenant: the *Directory (tenant) ID* or one of its domains. Left empty, a
197
+ single-tenant app's sign-in fails with `AADSTS50194`.
198
+ 2. The same menu › **Sign in**, with the lab open at `localhost`. A sign-in
199
+ from `127.0.0.1` reloads the page there and loses any open dialog.
200
+ 3. **Library › Sources › Add source › Power Automate workflow.**
201
+
202
+ Or set `M365_CLIENT_ID` and `M365_TENANT` when starting the lab, for the
203
+ whole machine; Setup then shows them read-only.
204
+
205
+ ## Upgrading
206
+
207
+ Read `CHANGELOG.md` first, in the package or on
208
+ <https://www.npmjs.com/package/evals-lab?activeTab=code>. When its
209
+ **Upgrade notes** say a release rewrites your data, backup the data directory
210
+ before upgrading.
183
211
 
184
212
  ```bash
185
213
  npm install -g evals-lab@latest # upgrade
@@ -188,13 +216,9 @@ npm uninstall -g evals-lab # remove; your data directory stays
188
216
 
189
217
  ## Licence
190
218
 
191
- Copyright 2026 Grainbox Limited. Licensed under the
219
+ Copyright 2026 Grainbox Limited, under the
192
220
  [PolyForm Internal Use License 1.0.0](https://polyformproject.org/licenses/internal-use/1.0.0)
193
- (`LICENSE.md` in the package): you and your company may use it, and change it,
194
- for your own internal business. You may not give it, or a changed copy, to
195
- anyone else, sell it, or offer it as a service.
196
-
197
- ## The demo images
198
-
199
- Every image in the demo datasets is in the public domain. Each has its line in
200
- `lab/demo/CREDITS.md` inside the package.
221
+ (`LICENSE.md`): you and your company may use and change it for your own
222
+ internal business, but not give it, or a changed copy, to anyone else, sell
223
+ it, or offer it as a service. The demo images are public domain; each is
224
+ credited in `lab/demo/CREDITS.md`.
package/bin/evals-lab.js CHANGED
@@ -1,6 +1,8 @@
1
1
  #!/usr/bin/env node
2
2
  // The npm package's launcher: `evals-lab` starts the lab's server and says
3
- // where it is.
3
+ // where it is, and `evals-lab run <bundle>` runs an Export for CI bundle with
4
+ // no server at all, or `evals-lab run --lab <url>` a running lab's pipeline
5
+ // (run.js).
4
6
  //
5
7
  // The lab itself is server.py, a stdlib Python server, staged beside this file
6
8
  // in ../lab by build.js. What this adds is what a lab run by someone who only
@@ -26,6 +28,8 @@ const PYTHON_FLOOR = [3, 10];
26
28
  const LOCAL_OLLAMA = "http://127.0.0.1:11434";
27
29
 
28
30
  const USAGE = `Usage: evals-lab [options]
31
+ evals-lab run <bundle> [options] (evals-lab run --help)
32
+ evals-lab run --lab <url> --pipeline <name> [--wait] [options]
29
33
 
30
34
  --port <n> the port to listen on. Default: $PORT, else 8080.
31
35
  --host <address> the address to listen on. Default: $LAB_HOST, else
@@ -152,6 +156,10 @@ function fail(message) {
152
156
  }
153
157
 
154
158
  async function main() {
159
+ if (process.argv[2] === "run") {
160
+ process.exitCode = await require("./run.js").run(process.argv.slice(3), { lab: LAB });
161
+ return;
162
+ }
155
163
  let o;
156
164
  try { o = options(process.argv.slice(2)); } catch (e) { fail(`${e.message}\n\n${USAGE}`); }
157
165
  if (o.help) return console.log(USAGE);