evals-lab 0.4.1 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +44 -0
- package/README.md +125 -28
- package/bin/evals-lab.js +9 -1
- package/bin/run.js +531 -0
- package/lab/VERSION +1 -1
- package/lab/demo/pipelines/demo-1.json +8 -9
- package/lab/demo/pipelines/demo-2.json +8 -9
- package/lab/evals-core.mjs +672 -61
- package/lab/run-evals.js +195 -57
- package/lab/server.py +681 -127
- package/lab/web/dist/assets/gallery-BFf9vis6.js +3 -0
- package/lab/web/dist/assets/{gallery-DFeJkfUw.css → gallery-B_-TH0F-.css} +1 -1
- package/lab/web/dist/assets/main-DeeRLWnO.css +1 -0
- package/lab/web/dist/assets/main-LT0U2TYF.js +21 -0
- package/lab/web/dist/assets/tokens-CCEtCZtQ.js +59 -0
- package/lab/web/dist/assets/tokens-lq45aAPS.css +1 -0
- package/lab/web/dist/gallery.html +4 -4
- package/lab/web/dist/index.html +4 -4
- package/package.json +1 -1
- package/lab/web/dist/assets/gallery-B-7oyY37.js +0 -3
- package/lab/web/dist/assets/main-BkZTEix2.js +0 -21
- package/lab/web/dist/assets/main-C_b7QoTv.css +0 -1
- package/lab/web/dist/assets/tokens-BEn6_hVz.css +0 -1
- package/lab/web/dist/assets/tokens-DLRdTFGY.js +0 -55
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,50 @@ Newest first. Read a release's **Upgrade notes** before installing it: an
|
|
|
5
5
|
upgrade can rewrite what the lab keeps in your data directory, and an older
|
|
6
6
|
version cannot always read it back.
|
|
7
7
|
|
|
8
|
+
## 0.5.0
|
|
9
|
+
|
|
10
|
+
### Upgrade notes
|
|
11
|
+
|
|
12
|
+
- **Copy your data directory before upgrading** (see "User data" in the
|
|
13
|
+
README). 0.5.0 keeps versions of each eval group and each run's verdict
|
|
14
|
+
beside what 0.4.x kept, and saves pipelines as pipeline version 13, which
|
|
15
|
+
0.4.x cannot read. To go back, reinstall 0.4.1 and restore the copy.
|
|
16
|
+
- Pipelines are read as version 13 and saved as it at their next edit. An
|
|
17
|
+
eval that graded by a dataset becomes a link to that eval group, or a
|
|
18
|
+
group of the pipeline's own when it had metrics of its own.
|
|
19
|
+
- Export JSON writes an eval group file, which 0.4.x refuses to import.
|
|
20
|
+
Import still reads dataset files of versions 1 to 7.
|
|
21
|
+
|
|
22
|
+
### Added
|
|
23
|
+
|
|
24
|
+
- `evals-lab run <bundle>` runs a pipeline's evals with no lab, for CI: the
|
|
25
|
+
pipeline's menu › Export for CI… downloads the bundle. It exits 0 when
|
|
26
|
+
every eval passes, 1 when one fails, 2 when the run is incomplete and 3
|
|
27
|
+
when it cannot run, and writes a JSON report (`--json`), JUnit XML
|
|
28
|
+
(`--junit`) and a Markdown summary (`--summary`). See "In CI" in the
|
|
29
|
+
README.
|
|
30
|
+
- `evals-lab run --lab <url> --pipeline <name> --wait` runs a lab's own
|
|
31
|
+
pipeline there, so the run is in its History, and exits with its verdict.
|
|
32
|
+
- Each Target profile has a slug in Setup. CI reads the profile's key from
|
|
33
|
+
`EVALSLAB_API_KEY_<SLUG>`.
|
|
34
|
+
- An eval group keeps its versions, and a run keeps the version it graded
|
|
35
|
+
with: Re-run grades with the same one.
|
|
36
|
+
- A run's row in History keeps its verdict.
|
|
37
|
+
|
|
38
|
+
### Changed
|
|
39
|
+
|
|
40
|
+
- A Target profile's key is never sent back to the page: Setup shows it as
|
|
41
|
+
Set, and typing replaces it.
|
|
42
|
+
- Tables look the same on every tab, every named column sorts, and lists
|
|
43
|
+
can select many rows at once.
|
|
44
|
+
- A dialog's buttons sit at its bottom right, and one that edits in place
|
|
45
|
+
closes with Close, or Save & Close once something changed.
|
|
46
|
+
|
|
47
|
+
### Fixed
|
|
48
|
+
|
|
49
|
+
- History skipped a run queued at the same moment as another when it loaded
|
|
50
|
+
older runs.
|
|
51
|
+
|
|
8
52
|
## 0.4.1
|
|
9
53
|
|
|
10
54
|
### Changed
|
package/README.md
CHANGED
|
@@ -1,15 +1,19 @@
|
|
|
1
1
|
# evals-lab
|
|
2
2
|
|
|
3
|
-
A lab
|
|
4
|
-
from a Source — uploaded files, a Power Automate flow's run history, or a
|
|
5
|
-
prompt alone — and sends each to one or more targets, every target a model
|
|
6
|
-
with its own prompt. Its evals score the replies with metrics — text and JSON
|
|
7
|
-
checks, item counts, agreement with production's reply, or a grading model —
|
|
8
|
-
and a dataset holds each item's expected answers. Every run is kept in
|
|
9
|
-
History, and every prompt version in the Library.
|
|
3
|
+
A lab for evaluating prompts and models in a friendly UI. A pipeline takes items from a Source (files, Power Automate runs, or a prompt) and sends each to one or more Targets, then runs evaluates Responses. Allows for quick iterations to compare changes in prompts and settings.
|
|
10
4
|
|
|
11
|
-
|
|
12
|
-
|
|
5
|
+
Model connections can be local (Ollama, llama.cpp) or online (Anthropic, OpenAI, any compatible HTTP endpoint).
|
|
6
|
+
|
|
7
|
+
The pipeline can be scored against metrics. The metrics include parsing checks, item counts, prod reply verification, grading models, eval groups (datasets), and more.
|
|
8
|
+
|
|
9
|
+
Runs are kept in history which can be exported.
|
|
10
|
+
|
|
11
|
+
Currently in early release. Runs locally and doesn't send your data anywhere else.
|
|
12
|
+
|
|
13
|
+
Example use-cases:
|
|
14
|
+
- Providing images to an LLM and comparing responses for accuracy
|
|
15
|
+
- Checking which model can meet evals in a Power Automate workflow
|
|
16
|
+
- Determining which prompt gets the highest score in metrics
|
|
13
17
|
|
|
14
18
|
## Install and run
|
|
15
19
|
|
|
@@ -19,7 +23,7 @@ evals-lab
|
|
|
19
23
|
```
|
|
20
24
|
|
|
21
25
|
Open <http://localhost:8080>. A new lab starts with two demo datasets and
|
|
22
|
-
their pipelines
|
|
26
|
+
their pipelines. Note the first start takes longer while it installs them. Without
|
|
23
27
|
installing: `npx evals-lab`.
|
|
24
28
|
|
|
25
29
|
| Needs | Windows | Linux |
|
|
@@ -31,10 +35,9 @@ installing: `npx evals-lab`.
|
|
|
31
35
|
`evals-lab` names anything it cannot find. Open a new terminal after
|
|
32
36
|
installing one. macOS should work and is untried.
|
|
33
37
|
|
|
34
|
-
##
|
|
38
|
+
## User data
|
|
35
39
|
|
|
36
|
-
|
|
37
|
-
files and runs are kept in one directory, which upgrades leave alone:
|
|
40
|
+
User data is kept in a single folder. Upgrades intend to keep user data. The folders are:
|
|
38
41
|
|
|
39
42
|
| | |
|
|
40
43
|
|---|---|
|
|
@@ -42,7 +45,7 @@ files and runs are kept in one directory, which upgrades leave alone:
|
|
|
42
45
|
| Linux | `~/.local/share/evals-lab` (or `$XDG_DATA_HOME/evals-lab`) |
|
|
43
46
|
| macOS | `~/Library/Application Support/evals-lab` |
|
|
44
47
|
|
|
45
|
-
Delete
|
|
48
|
+
Delete this folder to start again fresh.
|
|
46
49
|
|
|
47
50
|
## Options
|
|
48
51
|
|
|
@@ -59,24 +62,118 @@ evals-lab [--port <n>] [--host <address>] [--data-dir <dir>]
|
|
|
59
62
|
| | `OLLAMA_URL` | `http://127.0.0.1:11434` |
|
|
60
63
|
| | `PYTHON` | `py -3`, `python` or `python3`, whichever is found |
|
|
61
64
|
|
|
62
|
-
`--host 0.0.0.0` makes the lab reachable from your network and
|
|
63
|
-
|
|
64
|
-
|
|
65
|
+
`--host 0.0.0.0` makes the lab reachable from your network and needs `LAB_PASSWORD` to be set (otherwise it refuses connections).
|
|
66
|
+
|
|
67
|
+
## In CI
|
|
68
|
+
|
|
69
|
+
A pipeline can run with no lab, from a repository. In the lab, the pipeline's
|
|
70
|
+
menu › **Export for CI…** downloads it as a bundle: unzip it into the
|
|
71
|
+
repository, then:
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
EVALSLAB_API_KEY_ANTHROPIC=... evals-lab run evals/
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
```text
|
|
78
|
+
evals-lab run <bundle-dir | pipeline.yaml> [--items <dir>] [--json <file|->]
|
|
79
|
+
[--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
|
|
80
|
+
```
|
|
81
|
+
|
|
82
|
+
| Option | |
|
|
83
|
+
|---|---|
|
|
84
|
+
| `--items` | the items the pipeline's Source reads; default the bundle's `items/` |
|
|
85
|
+
| `--json` | the JSON report; `-` for stdout |
|
|
86
|
+
| `--junit` | JUnit XML, for a CI test panel |
|
|
87
|
+
| `--summary` | the results as Markdown, added to the file: for GitHub, `"$GITHUB_STEP_SUMMARY"` |
|
|
88
|
+
| `--min-pass` | an eval passes when this share of its items pass |
|
|
89
|
+
| `--progress` | a line per item as it finishes |
|
|
90
|
+
|
|
91
|
+
The bundle holds no key. Each Target profile's key comes from
|
|
92
|
+
`EVALSLAB_API_KEY_` and its slug in Setup, in capitals with hyphens as
|
|
93
|
+
underscores: a profile slugged `anthropic` reads `EVALSLAB_API_KEY_ANTHROPIC`,
|
|
94
|
+
and `gpt-4o-mini` reads `EVALSLAB_API_KEY_GPT_4O_MINI`. A table of results
|
|
95
|
+
goes to stderr.
|
|
96
|
+
|
|
97
|
+
| Exit code | |
|
|
98
|
+
|---|---|
|
|
99
|
+
| `0` | every eval passed |
|
|
100
|
+
| `1` | an eval failed |
|
|
101
|
+
| `2` | incomplete: an item is missing, or did not run |
|
|
102
|
+
| `3` | broken or refused |
|
|
103
|
+
|
|
104
|
+
An item a case names that is not in the items is reported as Missing.
|
|
105
|
+
Images need ImageMagick, as in the lab.
|
|
106
|
+
|
|
107
|
+
### GitHub Actions
|
|
108
|
+
|
|
109
|
+
```yaml
|
|
110
|
+
name: Evals
|
|
111
|
+
on: pull_request
|
|
112
|
+
|
|
113
|
+
jobs:
|
|
114
|
+
evals:
|
|
115
|
+
runs-on: ubuntu-latest
|
|
116
|
+
steps:
|
|
117
|
+
- uses: actions/checkout@v5
|
|
118
|
+
- uses: actions/setup-node@v4
|
|
119
|
+
with:
|
|
120
|
+
node-version: 24
|
|
121
|
+
# Only when the items are images:
|
|
122
|
+
# - run: sudo apt-get update && sudo apt-get install -y imagemagick
|
|
123
|
+
- run: >
|
|
124
|
+
npx --yes evals-lab@0.5 run evals/
|
|
125
|
+
--junit evals.xml --json evals.json --summary "$GITHUB_STEP_SUMMARY"
|
|
126
|
+
env:
|
|
127
|
+
EVALSLAB_API_KEY_ANTHROPIC: ${{ secrets.ANTHROPIC_API_KEY }}
|
|
128
|
+
- uses: actions/upload-artifact@v4
|
|
129
|
+
if: always()
|
|
130
|
+
with:
|
|
131
|
+
name: evals
|
|
132
|
+
path: evals.*
|
|
133
|
+
```
|
|
134
|
+
|
|
135
|
+
A failing eval fails the job, and the results show on the run's page. Other
|
|
136
|
+
CI systems run the same command and read `--junit` in their test panel.
|
|
137
|
+
|
|
138
|
+
### Against a running lab
|
|
139
|
+
|
|
140
|
+
CI can run a lab's own pipeline there instead, so the run shows in that lab's
|
|
141
|
+
History and its keys never leave it:
|
|
65
142
|
|
|
66
|
-
|
|
143
|
+
```bash
|
|
144
|
+
LAB_PASSWORD=... evals-lab run --lab https://lab.example --pipeline "Demo 1" --wait
|
|
145
|
+
```
|
|
67
146
|
|
|
68
|
-
|
|
147
|
+
```text
|
|
148
|
+
evals-lab run --lab <url> --pipeline <name> [--wait] [--json <file|->]
|
|
149
|
+
[--junit <file>] [--summary <file>] [--min-pass <0..1>] [--progress]
|
|
150
|
+
```
|
|
151
|
+
|
|
152
|
+
| Option | |
|
|
153
|
+
|---|---|
|
|
154
|
+
| `--lab` | the lab's address; `LAB_PASSWORD` is sent when it is set |
|
|
155
|
+
| `--pipeline` | the pipeline, by its name or its id |
|
|
156
|
+
| `--wait` | wait for the run, and exit with the lab's verdict on it |
|
|
157
|
+
|
|
158
|
+
Without `--wait` the run's id is printed once it is queued, and the exit code
|
|
159
|
+
is `0`; `--json`, `--junit`, `--summary`, `--min-pass` and `--progress` need
|
|
160
|
+
`--wait`. The exit codes are the ones above, and a cancelled run is `2`.
|
|
161
|
+
|
|
162
|
+
## Getting started
|
|
163
|
+
|
|
164
|
+
First set up a Setup profile: **Setup › Add profile**, then pick its type —
|
|
69
165
|
Ollama (local or cloud), OpenAI-compatible, Anthropic, llama.cpp or an HTTP
|
|
70
|
-
endpoint — and give its address, model and key.
|
|
71
|
-
|
|
166
|
+
endpoint — and give its address, model and key. NB: An Ollama profile without an address uses `OLLAMA_URL`.
|
|
167
|
+
|
|
168
|
+
Set up your data source(s) in Library.
|
|
169
|
+
|
|
170
|
+
Then create a Pipeline to Run. Changes are auto-saved. Use the "Restore" menu to go back and undo any changes you don't want to keep.
|
|
72
171
|
|
|
73
172
|
## Power Automate
|
|
74
173
|
|
|
75
174
|
The lab can read a Power Automate cloud flow and its run history from your
|
|
76
|
-
|
|
77
|
-
a work or school account
|
|
78
|
-
cloud flows. A small business, a trial or a Microsoft 365 Developer Program
|
|
79
|
-
tenant follows the same steps. The sign-in stays in your browser tab.
|
|
175
|
+
org's Microsoft 365 tenant then grade the flow's HTTP step. It needs
|
|
176
|
+
a work or school account. Personal accounts (outlook.com, hotmail.com) have no cloud flows. The sign-in data is kept in your browser to avoid it being stored elsewhere, the downside is you need to re-sign in if you close/reopen your web browser. The evals can keep running however, the sign-in is just to get the workflow data, which does persist in the evals-lab user data locally.
|
|
80
177
|
|
|
81
178
|
### Register an app
|
|
82
179
|
|
|
@@ -105,12 +202,12 @@ A tenant administrator does this once, at <https://entra.microsoft.com>:
|
|
|
105
202
|
Or set `M365_CLIENT_ID` and `M365_TENANT` when starting the lab, for the
|
|
106
203
|
whole machine; Setup then shows them read-only.
|
|
107
204
|
|
|
108
|
-
## Upgrading
|
|
205
|
+
## Upgrading
|
|
109
206
|
|
|
110
207
|
Read `CHANGELOG.md` first, in the package or on
|
|
111
208
|
<https://www.npmjs.com/package/evals-lab?activeTab=code>. When its
|
|
112
|
-
**Upgrade notes** say a release rewrites your data,
|
|
113
|
-
before upgrading
|
|
209
|
+
**Upgrade notes** say a release rewrites your data, backup the data directory
|
|
210
|
+
before upgrading.
|
|
114
211
|
|
|
115
212
|
```bash
|
|
116
213
|
npm install -g evals-lab@latest # upgrade
|
package/bin/evals-lab.js
CHANGED
|
@@ -1,6 +1,8 @@
|
|
|
1
1
|
#!/usr/bin/env node
|
|
2
2
|
// The npm package's launcher: `evals-lab` starts the lab's server and says
|
|
3
|
-
// where it is
|
|
3
|
+
// where it is, and `evals-lab run <bundle>` runs an Export for CI bundle with
|
|
4
|
+
// no server at all, or `evals-lab run --lab <url>` a running lab's pipeline
|
|
5
|
+
// (run.js).
|
|
4
6
|
//
|
|
5
7
|
// The lab itself is server.py, a stdlib Python server, staged beside this file
|
|
6
8
|
// in ../lab by build.js. What this adds is what a lab run by someone who only
|
|
@@ -26,6 +28,8 @@ const PYTHON_FLOOR = [3, 10];
|
|
|
26
28
|
const LOCAL_OLLAMA = "http://127.0.0.1:11434";
|
|
27
29
|
|
|
28
30
|
const USAGE = `Usage: evals-lab [options]
|
|
31
|
+
evals-lab run <bundle> [options] (evals-lab run --help)
|
|
32
|
+
evals-lab run --lab <url> --pipeline <name> [--wait] [options]
|
|
29
33
|
|
|
30
34
|
--port <n> the port to listen on. Default: $PORT, else 8080.
|
|
31
35
|
--host <address> the address to listen on. Default: $LAB_HOST, else
|
|
@@ -152,6 +156,10 @@ function fail(message) {
|
|
|
152
156
|
}
|
|
153
157
|
|
|
154
158
|
async function main() {
|
|
159
|
+
if (process.argv[2] === "run") {
|
|
160
|
+
process.exitCode = await require("./run.js").run(process.argv.slice(3), { lab: LAB });
|
|
161
|
+
return;
|
|
162
|
+
}
|
|
155
163
|
let o;
|
|
156
164
|
try { o = options(process.argv.slice(2)); } catch (e) { fail(`${e.message}\n\n${USAGE}`); }
|
|
157
165
|
if (o.help) return console.log(USAGE);
|