lastlight-evals 0.1.2 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +98 -12
- package/dashboard/dist/assets/index-YwScAAxm.js +238 -0
- package/dashboard/dist/assets/index-uU9Sj3M4.css +1 -0
- package/dashboard/dist/index.html +19 -0
- package/dashboard/dist/logo.png +0 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/.github/workflows/ci.yml +16 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/README.md +7 -1
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/package.json +3 -1
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/scripts/lint.mjs +25 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/test/date-range.test.ts +14 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/tsconfig.json +13 -0
- package/datasets/code-fix/tests/codefix__date-range-off-by-one/date-range.test.ts +5 -5
- package/dist/config.js +100 -0
- package/dist/config.js.map +1 -0
- package/dist/mechanism.test.js +41 -0
- package/dist/mechanism.test.js.map +1 -1
- package/dist/metrics.js +83 -0
- package/dist/metrics.js.map +1 -1
- package/dist/paths.js +51 -1
- package/dist/paths.js.map +1 -1
- package/dist/report.js +112 -3
- package/dist/report.js.map +1 -1
- package/dist/run-instance.js +141 -14
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +342 -112
- package/dist/run.js.map +1 -1
- package/dist/serve.js +136 -0
- package/dist/serve.js.map +1 -0
- package/examples/overlay/README.md +30 -0
- package/examples/overlay/config.yaml +36 -0
- package/examples/overlay-anthropic/config.yaml +27 -0
- package/package.json +16 -6
- package/dist/html-report.js +0 -325
- package/dist/html-report.js.map +0 -1
package/README.md
CHANGED
|
@@ -2,9 +2,9 @@
|
|
|
2
2
|
|
|
3
3
|
### Which model should run your agent? Find out — with receipts.
|
|
4
4
|
|
|
5
|
-
[](https://
|
|
5
|
+
[](https://evals.lastlight.dev/)
|
|
6
6
|
|
|
7
|
-
> **[▶ Explore the live scorecard](https://
|
|
7
|
+
> **[▶ Explore the live scorecard](https://evals.lastlight.dev/)** — interactive, with per-instance detail. (Above: code-fix tier across 9 models.)
|
|
8
8
|
|
|
9
9
|
`lastlight-evals` takes [**Last Light**](https://lastlight.dev)'s *real*
|
|
10
10
|
production workflows — the actual prompts, skills, and agent loop that ship — and
|
|
@@ -45,10 +45,37 @@ instance (SWE-bench shape)
|
|
|
45
45
|
|
|
46
46
|
## Get started
|
|
47
47
|
|
|
48
|
-
Needs **Node 24+** and a provider API key.
|
|
49
|
-
|
|
50
|
-
|
|
51
|
-
|
|
48
|
+
Needs **Node 24+** and a provider API key.
|
|
49
|
+
|
|
50
|
+
### Easiest: let the Last Light agent skill set up *your own* workspace
|
|
51
|
+
|
|
52
|
+
Want to eval **your own** deployment — your workflows, your agent persona, your
|
|
53
|
+
config — not just the shipped samples? If you drive
|
|
54
|
+
[Last Light](https://lastlight.dev) from an agent (e.g. Claude Code), install its
|
|
55
|
+
skills once and then just *ask* — no flags to remember:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
lastlight skills install # installs the Last Light agent skills
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Then, in a **new empty folder**, tell your agent (point it at *your* instance
|
|
62
|
+
overlay repo):
|
|
63
|
+
|
|
64
|
+
> *Let's set up an evals workspace here, using my existing Last Light instance
|
|
65
|
+
> config in `cliftonc/lastlight-instance`.*
|
|
66
|
+
|
|
67
|
+
The `lastlight-evals` skill scaffolds the workspace, clones your overlay into
|
|
68
|
+
`instance/`, seeds the sample datasets, and wires it all up — under the hood it
|
|
69
|
+
runs `lastlight-evals init . --clone cliftonc/lastlight-instance`, after which a
|
|
70
|
+
bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
|
|
71
|
+
overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
|
|
72
|
+
models you care about. Prefer to drive it by hand? Keep reading.
|
|
73
|
+
|
|
74
|
+
### Manual: scaffold with `init`
|
|
75
|
+
|
|
76
|
+
The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
|
|
77
|
+
(your workflows + your datasets, seeded from the built-in samples) and optionally
|
|
78
|
+
creates a private GitHub repo for it:
|
|
52
79
|
|
|
53
80
|
```bash
|
|
54
81
|
npm install -g lastlight-evals
|
|
@@ -154,25 +181,64 @@ lastlight-evals run --overlay ~/work/lastlight-instance
|
|
|
154
181
|
# add your own datasets dir without an overlay
|
|
155
182
|
lastlight-evals run --datasets ~/my-evals/datasets
|
|
156
183
|
|
|
184
|
+
# CONFIG run type — eval a deployment's REAL per-step model config (different
|
|
185
|
+
# models per workflow phase, from the overlay's config.yaml) instead of forcing
|
|
186
|
+
# one model. This is the setup you actually ship. Try the bundled sample overlay:
|
|
187
|
+
lastlight-evals run code-fix --mode config --overlay examples/overlay
|
|
188
|
+
lastlight-evals run code-fix --mode config --overlay A --overlay B # 2 configs side-by-side
|
|
189
|
+
|
|
157
190
|
# ad-hoc model set / focus one instance / no browser
|
|
158
191
|
EVAL_MODELS="openai/gpt-5.5,anthropic/claude-sonnet-4-6" lastlight-evals run
|
|
159
192
|
EVAL_INSTANCE=off-by-one lastlight-evals run code-fix
|
|
160
193
|
lastlight-evals run triage --no-open
|
|
161
194
|
```
|
|
162
195
|
|
|
163
|
-
The
|
|
164
|
-
|
|
165
|
-
|
|
196
|
+
The report is a **JSON-driven dashboard**, not generated HTML — the harness only
|
|
197
|
+
ever writes `scorecard.json`, updating it (atomically) as the run proceeds. The
|
|
198
|
+
runner starts a tiny local server and opens `http://localhost:PORT` deep-linked
|
|
199
|
+
at the run, so you watch the scorecard fill in live (the SPA polls the JSON).
|
|
200
|
+
When the run finishes the server stays up so the dashboard keeps working — press
|
|
201
|
+
`Ctrl-C` to stop it. Each run lands in its **own** timestamped folder, so runs
|
|
202
|
+
accumulate instead of overwriting — `./eval-results/<tiers>/<runId>/` (override
|
|
203
|
+
the root with `LASTLIGHT_EVALS_OUT`), where `runId` is `<timestamp>-<git-sha>`:
|
|
166
204
|
|
|
167
|
-
- `
|
|
168
|
-
- `scorecard.json` — structured roll-up per model.
|
|
205
|
+
- `scorecard.json` — structured roll-up per model + per-instance results, carrying run `meta`.
|
|
169
206
|
- `predictions.jsonl` — SWE-bench predictions shape.
|
|
170
207
|
|
|
208
|
+
The dashboard's **overview** lists every run newest-first with a per-model trend
|
|
209
|
+
sparkline and links into each run's full scorecard; the **run view** is the
|
|
210
|
+
model-comparison table plus per-instance rows. To browse past runs anytime
|
|
211
|
+
without running models, start the server on its own:
|
|
212
|
+
|
|
213
|
+
```bash
|
|
214
|
+
lastlight-evals serve # opens the dashboard over ./eval-results
|
|
215
|
+
lastlight-evals serve --port 4319
|
|
216
|
+
```
|
|
217
|
+
|
|
171
218
|
Needs a provider key (`OPENAI_API_KEY` / `ANTHROPIC_API_KEY` /
|
|
172
219
|
`FIREWORKS_API_KEY` / `OPENROUTER_API_KEY`) in the environment or a cwd `.env`.
|
|
173
220
|
The runner exits non-zero **only** if the harness itself errors — a weak model
|
|
174
221
|
scoring poorly is the measurement, not a build failure.
|
|
175
222
|
|
|
223
|
+
### Two run types
|
|
224
|
+
|
|
225
|
+
A run compares N **arms** along one of two axes — pick with `--mode` (or, in a
|
|
226
|
+
TTY with no model flags, you're asked):
|
|
227
|
+
|
|
228
|
+
- **`models`** (default) — compare models, each **forced across every workflow
|
|
229
|
+
step**. `--model`/`--compare` select the set. Lands in `eval-results/<tier>/`
|
|
230
|
+
(or `<tier>-compare/`).
|
|
231
|
+
- **`config`** (`--mode config`) — run a deployment's **real per-step model
|
|
232
|
+
config**: the `models`/`variants` maps from an overlay's `config.yaml`, merged
|
|
233
|
+
over core's `config/default.yaml` exactly as production does, so each phase can
|
|
234
|
+
run on a different model. The arm is the config/overlay; pass `--overlay` more
|
|
235
|
+
than once to compare configs side-by-side, or re-run over time to compare as
|
|
236
|
+
you tweak prompts/skills/workflow/model-config. `--model` overrides a config's
|
|
237
|
+
`default` for quick what-ifs. Lands in `eval-results/<tier>-config/`, on its
|
|
238
|
+
own trend line. The run view shows a **Per-step models** panel with each
|
|
239
|
+
phase's resolved model. See [`examples/overlay`](examples/overlay) for a
|
|
240
|
+
ready-to-run sample.
|
|
241
|
+
|
|
176
242
|
## Your own workflows + datasets (overlays)
|
|
177
243
|
|
|
178
244
|
An **overlay** is a directory (often its own repo, like `lastlight-instance`)
|
|
@@ -194,7 +260,12 @@ lastlight-evals run --overlay ~/work/lastlight-instance # or LASTLIGHT_OVERL
|
|
|
194
260
|
- An overlay **`evals/models.json`** is picked up automatically (or pass
|
|
195
261
|
`--models-file`).
|
|
196
262
|
|
|
197
|
-
### `lastlight-evals init [dir]` — scaffold
|
|
263
|
+
### `lastlight-evals init [dir]` — scaffold an evals workspace
|
|
264
|
+
|
|
265
|
+
Two shapes, depending on whether you already have a deployment overlay repo:
|
|
266
|
+
|
|
267
|
+
**Plain** — a self-contained overlay+evals repo (its own `workflows/` `skills/`
|
|
268
|
+
`agent-context/` + `evals/`):
|
|
198
269
|
|
|
199
270
|
```bash
|
|
200
271
|
lastlight-evals init my-evals
|
|
@@ -206,6 +277,21 @@ Scaffolds `workflows/` `skills/` `agent-context/` (empty, to fill in),
|
|
|
206
277
|
`config.yaml`, and a `.gitignore`/`README`, then offers to `git init` + create a
|
|
207
278
|
private GitHub repo via `gh` (reusing core's `lastlight server setup` flow).
|
|
208
279
|
|
|
280
|
+
**Separate** (`--clone`) — the recommended shape when you already have a
|
|
281
|
+
deployment overlay (e.g. `lastlight-instance`) and want to eval **its** config.
|
|
282
|
+
The overlay is cloned into `<dir>/instance/` (its own git checkout, git-ignored)
|
|
283
|
+
with the evals at the workspace root; a bare run auto-detects both, no flags:
|
|
284
|
+
|
|
285
|
+
```bash
|
|
286
|
+
lastlight-evals init my-evals --clone cliftonc/lastlight-instance
|
|
287
|
+
cd my-evals && lastlight-evals run # auto: overlay ./instance + ./evals/datasets
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
This is exactly what the [`lastlight-evals` agent skill](#easiest-let-the-last-light-agent-skill-set-up-your-own-workspace)
|
|
291
|
+
does for you. Update the overlay later with `cd instance && git pull`; your evals
|
|
292
|
+
stay out of the deployment repo. Run `lastlight-evals init --help` for all flags
|
|
293
|
+
(`--yes`, `--no-git`, …).
|
|
294
|
+
|
|
209
295
|
## Datasets & tiers
|
|
210
296
|
|
|
211
297
|
A **tier** is a directory containing `instances.json` (+ an optional `tier.json`
|