evals-lab 0.4.1 → 0.6.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,87 @@ Newest first. Read a release's **Upgrade notes** before installing it: an
5
5
  upgrade can rewrite what the lab keeps in your data directory, and an older
6
6
  version cannot always read it back.
7
7
 
8
+ ## 0.6.0
9
+
10
+ ### Added
11
+
12
+ - Library › Evals (was Library › Datasets) lists every eval group: its
13
+ Source, cases, scoring, and the pipelines that link it. A group's page
14
+ holds its Grades, Every item, Whole run and cases.
15
+ - Runs › Evals links eval groups to a pipeline. Link eval group picks from
16
+ the Library and marks the groups made for this content. A link follows a
17
+ group's latest version, or is pinned to one. A pipeline's own group is
18
+ edited in place, and Move to Library shares it.
19
+ - A pipeline can link several eval groups. Its pass rule says whether every
20
+ one must pass, or at least a number. Results, History and `evals-lab run`
21
+ all grade by it.
22
+ - An Undo button in the header takes back the last 20 actions, newest
23
+ first, until the page reloads.
24
+ - Choose pipeline lists the 5 used most recently, then the rest, searchable
25
+ and sortable.
26
+ - A Profile, Source or Grader field opens a picker, most recent first.
27
+
28
+ ### Changed
29
+
30
+ - Results shows a row per eval group, with each Target's share and verdict,
31
+ then the overall verdict. History's Outcome is the overall verdict.
32
+ - A dialog is a draft: nothing reaches the lab until Save & Close. Cancel,
33
+ × and Esc discard it.
34
+ - An unnamed target is Target A, B, C.
35
+ - The run bar sits under the pipeline's header. It says "12 of 42" while a
36
+ run goes and "42 in 2m" when it is done.
37
+ - Verdicts show as a check or a cross, in Results and History.
38
+ - A metric's "How it counts" is Options, and Grades is Reads: Parsed reply
39
+ or Raw reply.
40
+
41
+ ### Fixed
42
+
43
+ - Add profile, then ×, added a profile.
44
+
45
+ ## 0.5.0
46
+
47
+ ### Upgrade notes
48
+
49
+ - **Copy your data directory before upgrading** (see "User data" in the
50
+ README). 0.5.0 keeps versions of each eval group and each run's verdict
51
+ beside what 0.4.x kept, and saves pipelines as pipeline version 13, which
52
+ 0.4.x cannot read. To go back, reinstall 0.4.1 and restore the copy.
53
+ - Pipelines are read as version 13 and saved as it at their next edit. An
54
+ eval that graded by a dataset becomes a link to that eval group, or a
55
+ group of the pipeline's own when it had metrics of its own.
56
+ - Export JSON writes an eval group file, which 0.4.x refuses to import.
57
+ Import still reads dataset files of versions 1 to 7.
58
+
59
+ ### Added
60
+
61
+ - `evals-lab run <bundle>` runs a pipeline's evals with no lab, for CI: the
62
+ pipeline's menu › Export for CI… downloads the bundle. It exits 0 when
63
+ every eval passes, 1 when one fails, 2 when the run is incomplete and 3
64
+ when it cannot run, and writes a JSON report (`--json`), JUnit XML
65
+ (`--junit`) and a Markdown summary (`--summary`). See "In CI" in the
66
+ README.
67
+ - `evals-lab run --lab <url> --pipeline <name> --wait` runs a lab's own
68
+ pipeline there, so the run is in its History, and exits with its verdict.
69
+ - Each Target profile has a slug in Setup. CI reads the profile's key from
70
+ `EVALSLAB_API_KEY_<SLUG>`.
71
+ - An eval group keeps its versions, and a run keeps the version it graded
72
+ with: Re-run grades with the same one.
73
+ - A run's row in History keeps its verdict.
74
+
75
+ ### Changed
76
+
77
+ - A Target profile's key is never sent back to the page: Setup shows it as
78
+ Set, and typing replaces it.
79
+ - Tables look the same on every tab, every named column sorts, and lists
80
+ can select many rows at once.
81
+ - A dialog's buttons sit at its bottom right, and one that edits in place
82
+ closes with Close, or Save & Close once something changed.
83
+
84
+ ### Fixed
85
+
86
+ - History skipped a run queued at the same moment as another when it loaded
87
+ older runs.
88
+
8
89
  ## 0.4.1
9
90
 
10
91
  ### Changed
package/README.md CHANGED
@@ -1,15 +1,19 @@
1
1
  # evals-lab
2
2
 
3
- A lab in the browser for evaluating prompts and models. A pipeline takes items
4
- from a Source — uploaded files, a Power Automate flow's run history, or a
5
- prompt alone — and sends each to one or more targets, every target a model
6
- with its own prompt. Its evals score the replies with metrics — text and JSON
7
- checks, item counts, agreement with production's reply, or a grading model —
8
- and a dataset holds each item's expected answers. Every run is kept in
9
- History, and every prompt version in the Library.
3
+ A lab for evaluating prompts and models in a friendly UI. A pipeline takes items from a Source (files, Power Automate runs, or a prompt) and sends each to one or more Targets, then runs evaluates Responses. Allows for quick iterations to compare changes in prompts and settings.
10
4
 
11
- An early test release. It runs on your machine and talks only to the models
12
- and services you connect it to.
5
+ Model connections can be local (Ollama, llama.cpp) or online (Anthropic, OpenAI, any compatible HTTP endpoint).
6
+
7
+ The pipeline can be scored against metrics. The metrics include parsing checks, item counts, prod reply verification, grading models, eval groups (datasets), and more.
8
+
9
+ Runs are kept in history which can be exported.
10
+
11
+ Currently in early release. Runs locally and doesn't send your data anywhere else.
12
+
13
+ Example use-cases:
14
+ - Providing images to an LLM and comparing responses for accuracy
15
+ - Checking which model can meet evals in a Power Automate workflow
16
+ - Determining which prompt gets the highest score in metrics
13
17
 
14
18
  ## Install and run
15
19
 
@@ -19,7 +23,7 @@ evals-lab
19
23
  ```
20
24
 
21
25
  Open <http://localhost:8080>. A new lab starts with two demo datasets and
22
- their pipelines; the first start takes longer while it installs them. Without
26
+ their pipelines. Note the first start takes longer while it installs them. Without
23
27
  installing: `npx evals-lab`.
24
28
 
25
29
  | Needs | Windows | Linux |
@@ -31,10 +35,9 @@ installing: `npx evals-lab`.
31
35
  `evals-lab` names anything it cannot find. Open a new terminal after
32
36
  installing one. macOS should work and is untried.
33
37
 
34
- ## Where your data is kept
38
+ ## User data
35
39
 
36
- Datasets, pipelines, Setup profiles (with their API keys), prompts, uploaded
37
- files and runs are kept in one directory, which upgrades leave alone:
40
+ User data is kept in a single folder. Upgrades intend to keep user data. The folders are:
38
41
 
39
42
  | | |
40
43
  |---|---|
@@ -42,7 +45,7 @@ files and runs are kept in one directory, which upgrades leave alone:
42
45
  | Linux | `~/.local/share/evals-lab` (or `$XDG_DATA_HOME/evals-lab`) |
43
46
  | macOS | `~/Library/Application Support/evals-lab` |
44
47
 
45
- Delete it to start again from a new lab.
48
+ Delete this folder to start again fresh.
46
49
 
47
50
  ## Options
48
51
 
@@ -59,24 +62,118 @@ evals-lab [--port <n>] [--host <address>] [--data-dir <dir>]
59
62
  | | `OLLAMA_URL` | `http://127.0.0.1:11434` |
60
63
  | | `PYTHON` | `py -3`, `python` or `python3`, whichever is found |
61
64
 
62
- `--host 0.0.0.0` makes the lab reachable from your network and is refused
63
- without `LAB_PASSWORD`. Never put it on the open internet: the lab holds your
64
- API keys and forwards requests to any address a profile names.
65
+ `--host 0.0.0.0` makes the lab reachable from your network and needs `LAB_PASSWORD` to be set (otherwise it refuses connections).
66
+
67
+ ## In CI
68
+
69
+ A pipeline can run with no lab, from a repository. In the lab, the pipeline's
70
+ menu › **Export for CI…** downloads it as a bundle: unzip it into the
71
+ repository, then:
72
+
73
+ ```bash
74
+ EVALSLAB_API_KEY_ANTHROPIC=... evals-lab run evals/
75
+ ```
76
+
77
+ ```text
78
+ evals-lab run <bundle-dir | pipeline.yaml> [--items <dir>] [--json <file|->]
79
+ [--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
80
+ ```
81
+
82
+ | Option | |
83
+ |---|---|
84
+ | `--items` | the items the pipeline's Source reads; default the bundle's `items/` |
85
+ | `--json` | the JSON report; `-` for stdout |
86
+ | `--junit` | JUnit XML, for a CI test panel |
87
+ | `--summary` | the results as Markdown, added to the file: for GitHub, `"$GITHUB_STEP_SUMMARY"` |
88
+ | `--min-pass` | an eval passes when this share of its items pass |
89
+ | `--progress` | a line per item as it finishes |
90
+
91
+ The bundle holds no key. Each Target profile's key comes from
92
+ `EVALSLAB_API_KEY_` and its slug in Setup, in capitals with hyphens as
93
+ underscores: a profile slugged `anthropic` reads `EVALSLAB_API_KEY_ANTHROPIC`,
94
+ and `gpt-4o-mini` reads `EVALSLAB_API_KEY_GPT_4O_MINI`. A table of results
95
+ goes to stderr.
96
+
97
+ | Exit code | |
98
+ |---|---|
99
+ | `0` | every eval passed |
100
+ | `1` | an eval failed |
101
+ | `2` | incomplete: an item is missing, or did not run |
102
+ | `3` | broken or refused |
103
+
104
+ An item a case names that is not in the items is reported as Missing.
105
+ Images need ImageMagick, as in the lab.
106
+
107
+ ### GitHub Actions
108
+
109
+ ```yaml
110
+ name: Evals
111
+ on: pull_request
112
+
113
+ jobs:
114
+ evals:
115
+ runs-on: ubuntu-latest
116
+ steps:
117
+ - uses: actions/checkout@v5
118
+ - uses: actions/setup-node@v4
119
+ with:
120
+ node-version: 24
121
+ # Only when the items are images:
122
+ # - run: sudo apt-get update && sudo apt-get install -y imagemagick
123
+ - run: >
124
+ npx --yes evals-lab@0.5 run evals/
125
+ --junit evals.xml --json evals.json --summary "$GITHUB_STEP_SUMMARY"
126
+ env:
127
+ EVALSLAB_API_KEY_ANTHROPIC: ${{ secrets.ANTHROPIC_API_KEY }}
128
+ - uses: actions/upload-artifact@v4
129
+ if: always()
130
+ with:
131
+ name: evals
132
+ path: evals.*
133
+ ```
134
+
135
+ A failing eval fails the job, and the results show on the run's page. Other
136
+ CI systems run the same command and read `--junit` in their test panel.
137
+
138
+ ### Against a running lab
139
+
140
+ CI can run a lab's own pipeline there instead, so the run shows in that lab's
141
+ History and its keys never leave it:
65
142
 
66
- ## Models
143
+ ```bash
144
+ LAB_PASSWORD=... evals-lab run --lab https://lab.example --pipeline "Demo 1" --wait
145
+ ```
67
146
 
68
- A model is a Setup profile: **Setup › Add profile**, then pick its type —
147
+ ```text
148
+ evals-lab run --lab <url> --pipeline <name> [--wait] [--json <file|->]
149
+ [--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
150
+ ```
151
+
152
+ | Option | |
153
+ |---|---|
154
+ | `--lab` | the lab's address; `LAB_PASSWORD` is sent when it is set |
155
+ | `--pipeline` | the pipeline, by its name or its id |
156
+ | `--wait` | wait for the run, and exit with the lab's verdict on it |
157
+
158
+ Without `--wait` the run's id is printed once it is queued, and the exit code
159
+ is `0`; `--json`, `--junit`, `--summary`, `--min-pass` and `--progress` need
160
+ `--wait`. The exit codes are the ones above, and a cancelled run is `2`.
161
+
162
+ ## Getting started
163
+
164
+ First set up a Setup profile: **Setup › Add profile**, then pick its type —
69
165
  Ollama (local or cloud), OpenAI-compatible, Anthropic, llama.cpp or an HTTP
70
- endpoint — and give its address, model and key. **Test connection** checks
71
- it. An Ollama profile left without an address uses `OLLAMA_URL`.
166
+ endpoint — and give its address, model and key. NB: An Ollama profile without an address uses `OLLAMA_URL`.
167
+
168
+ Set up your data source(s) in Library.
169
+
170
+ Then create a Pipeline to Run. Changes are auto-saved. Use the "Restore" menu to go back and undo any changes you don't want to keep.
72
171
 
73
172
  ## Power Automate
74
173
 
75
174
  The lab can read a Power Automate cloud flow and its run history from your
76
- organisation's Microsoft 365 tenant, then grade the flow's HTTP step. It needs
77
- a work or school account; personal accounts (outlook.com, hotmail.com) have no
78
- cloud flows. A small business, a trial or a Microsoft 365 Developer Program
79
- tenant follows the same steps. The sign-in stays in your browser tab.
175
+ org's Microsoft 365 tenant then grade the flow's HTTP step. It needs
176
+ a work or school account. Personal accounts (outlook.com, hotmail.com) have no cloud flows. The sign-in data is kept in your browser to avoid it being stored elsewhere, the downside is you need to re-sign in if you close/reopen your web browser. The evals can keep running however, the sign-in is just to get the workflow data, which does persist in the evals-lab user data locally.
80
177
 
81
178
  ### Register an app
82
179
 
@@ -105,12 +202,12 @@ A tenant administrator does this once, at <https://entra.microsoft.com>:
105
202
  Or set `M365_CLIENT_ID` and `M365_TENANT` when starting the lab, for the
106
203
  whole machine; Setup then shows them read-only.
107
204
 
108
- ## Upgrading and removing
205
+ ## Upgrading
109
206
 
110
207
  Read `CHANGELOG.md` first, in the package or on
111
208
  <https://www.npmjs.com/package/evals-lab?activeTab=code>. When its
112
- **Upgrade notes** say a release rewrites your data, copy the data directory
113
- before upgrading if you may want to go back.
209
+ **Upgrade notes** say a release rewrites your data, backup the data directory
210
+ before upgrading.
114
211
 
115
212
  ```bash
116
213
  npm install -g evals-lab@latest # upgrade
package/bin/evals-lab.js CHANGED
@@ -1,6 +1,8 @@
1
1
  #!/usr/bin/env node
2
2
  // The npm package's launcher: `evals-lab` starts the lab's server and says
3
- // where it is.
3
+ // where it is, and `evals-lab run <bundle>` runs an Export for CI bundle with
4
+ // no server at all, or `evals-lab run --lab <url>` a running lab's pipeline
5
+ // (run.js).
4
6
  //
5
7
  // The lab itself is server.py, a stdlib Python server, staged beside this file
6
8
  // in ../lab by build.js. What this adds is what a lab run by someone who only
@@ -26,6 +28,8 @@ const PYTHON_FLOOR = [3, 10];
26
28
  const LOCAL_OLLAMA = "http://127.0.0.1:11434";
27
29
 
28
30
  const USAGE = `Usage: evals-lab [options]
31
+ evals-lab run <bundle> [options] (evals-lab run --help)
32
+ evals-lab run --lab <url> --pipeline <name> [--wait] [options]
29
33
 
30
34
  --port <n> the port to listen on. Default: $PORT, else 8080.
31
35
  --host <address> the address to listen on. Default: $LAB_HOST, else
@@ -152,6 +156,10 @@ function fail(message) {
152
156
  }
153
157
 
154
158
  async function main() {
159
+ if (process.argv[2] === "run") {
160
+ process.exitCode = await require("./run.js").run(process.argv.slice(3), { lab: LAB });
161
+ return;
162
+ }
155
163
  let o;
156
164
  try { o = options(process.argv.slice(2)); } catch (e) { fail(`${e.message}\n\n${USAGE}`); }
157
165
  if (o.help) return console.log(USAGE);