lastlight-evals 0.1.1 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +198 -39
- package/dashboard/dist/assets/index-YwScAAxm.js +238 -0
- package/dashboard/dist/assets/index-uU9Sj3M4.css +1 -0
- package/dashboard/dist/index.html +19 -0
- package/dashboard/dist/logo.png +0 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/.github/workflows/ci.yml +16 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/README.md +7 -1
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/package.json +3 -1
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/scripts/lint.mjs +25 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/test/date-range.test.ts +14 -0
- package/datasets/code-fix/repos/codefix__date-range-off-by-one/tsconfig.json +13 -0
- package/datasets/code-fix/tests/codefix__date-range-off-by-one/date-range.test.ts +5 -5
- package/dist/config.js +100 -0
- package/dist/config.js.map +1 -0
- package/dist/init.js +229 -32
- package/dist/init.js.map +1 -1
- package/dist/mechanism.test.js +41 -0
- package/dist/mechanism.test.js.map +1 -1
- package/dist/metrics.js +91 -3
- package/dist/metrics.js.map +1 -1
- package/dist/paths.js +51 -1
- package/dist/paths.js.map +1 -1
- package/dist/report.js +128 -6
- package/dist/report.js.map +1 -1
- package/dist/run-instance.js +143 -14
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +401 -89
- package/dist/run.js.map +1 -1
- package/dist/serve.js +136 -0
- package/dist/serve.js.map +1 -0
- package/examples/overlay/README.md +30 -0
- package/examples/overlay/config.yaml +36 -0
- package/examples/overlay-anthropic/config.yaml +27 -0
- package/package.json +16 -6
- package/dist/html-report.js +0 -325
- package/dist/html-report.js.map +0 -1
package/README.md
CHANGED
|
@@ -1,16 +1,31 @@
|
|
|
1
1
|
# lastlight-evals
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
3
|
+
### Which model should run your agent? Find out — with receipts.
|
|
4
|
+
|
|
5
|
+
[](https://evals.lastlight.dev/)
|
|
6
|
+
|
|
7
|
+
> **[▶ Explore the live scorecard](https://evals.lastlight.dev/)** — interactive, with per-instance detail. (Above: code-fix tier across 9 models.)
|
|
8
|
+
|
|
9
|
+
`lastlight-evals` takes [**Last Light**](https://lastlight.dev)'s *real*
|
|
10
|
+
production workflows — the actual prompts, skills, and agent loop that ship — and
|
|
11
|
+
runs them end to end against a fully mocked GitHub, for whatever models you throw
|
|
12
|
+
at it. No toy benchmarks, no LLM-as-judge: every run is graded **deterministically**
|
|
13
|
+
(did the agent apply the right labels? did the held-out tests turn green?), then
|
|
14
|
+
ranked side by side on **pass rate, cost, and latency**.
|
|
15
|
+
|
|
16
|
+
The payoff is one scorecard that tells you, for *your* workflows, exactly what
|
|
17
|
+
each model delivers — and what it costs you per run. Swap a model, re-run, see the
|
|
18
|
+
difference. Drop in your own issues and repos and it evaluates *your* agent.
|
|
19
|
+
|
|
20
|
+
> 🛰️ Part of [**Last Light**](https://lastlight.dev) — the AI agent that triages,
|
|
21
|
+
> reviews, and fixes your GitHub repos.
|
|
22
|
+
> **[lastlight.dev](https://lastlight.dev)** · [Core repo](https://github.com/cliftonc/lastlight) · [Eval repo](https://github.com/cliftonc/lastlight-evals)
|
|
23
|
+
|
|
24
|
+
It's **SWE-bench-compatible**, and nothing here touches real GitHub: the agent's
|
|
25
|
+
`github_*` tool calls are served by an in-process fake (seeded + recording) and
|
|
26
|
+
`git push` goes to a local bare repo. The only deviations from production are the
|
|
27
|
+
two we can't do unattended — approval gates are disabled and outward side-effects
|
|
28
|
+
are mocked. Everything else is exactly what ships.
|
|
14
29
|
|
|
15
30
|
```
|
|
16
31
|
instance (SWE-bench shape)
|
|
@@ -28,40 +43,121 @@ instance (SWE-bench shape)
|
|
|
28
43
|
> (the base-URL mock, static-token mode, the no-clone seeding trick, the
|
|
29
44
|
> asset-bootstrap footgun, the metrics drain).
|
|
30
45
|
|
|
31
|
-
##
|
|
46
|
+
## Get started
|
|
47
|
+
|
|
48
|
+
Needs **Node 24+** and a provider API key.
|
|
49
|
+
|
|
50
|
+
### Easiest: let the Last Light agent skill set up *your own* workspace
|
|
51
|
+
|
|
52
|
+
Want to eval **your own** deployment — your workflows, your agent persona, your
|
|
53
|
+
config — not just the shipped samples? If you drive
|
|
54
|
+
[Last Light](https://lastlight.dev) from an agent (e.g. Claude Code), install its
|
|
55
|
+
skills once and then just *ask* — no flags to remember:
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
lastlight skills install # installs the Last Light agent skills
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
Then, in a **new empty folder**, tell your agent (point it at *your* instance
|
|
62
|
+
overlay repo):
|
|
63
|
+
|
|
64
|
+
> *Let's set up an evals workspace here, using my existing Last Light instance
|
|
65
|
+
> config in `cliftonc/lastlight-instance`.*
|
|
66
|
+
|
|
67
|
+
The `lastlight-evals` skill scaffolds the workspace, clones your overlay into
|
|
68
|
+
`instance/`, seeds the sample datasets, and wires it all up — under the hood it
|
|
69
|
+
runs `lastlight-evals init . --clone cliftonc/lastlight-instance`, after which a
|
|
70
|
+
bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
|
|
71
|
+
overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
|
|
72
|
+
models you care about. Prefer to drive it by hand? Keep reading.
|
|
32
73
|
|
|
33
|
-
|
|
34
|
-
|
|
35
|
-
|
|
36
|
-
|
|
37
|
-
|
|
38
|
-
core does.
|
|
74
|
+
### Manual: scaffold with `init`
|
|
75
|
+
|
|
76
|
+
The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
|
|
77
|
+
(your workflows + your datasets, seeded from the built-in samples) and optionally
|
|
78
|
+
creates a private GitHub repo for it:
|
|
39
79
|
|
|
40
80
|
```bash
|
|
41
|
-
npm install
|
|
81
|
+
npm install -g lastlight-evals
|
|
82
|
+
export OPENAI_API_KEY=... # or ANTHROPIC_ / FIREWORKS_ / OPENROUTER_
|
|
83
|
+
|
|
84
|
+
# 1. Scaffold your workspace (offers to `git init` + `gh repo create`).
|
|
85
|
+
lastlight-evals init my-evals
|
|
86
|
+
cd my-evals
|
|
87
|
+
|
|
88
|
+
# 2. Run it — drives the real workflows against your datasets, prints a scorecard.
|
|
89
|
+
lastlight-evals run --overlay .
|
|
42
90
|
```
|
|
43
91
|
|
|
44
|
-
|
|
45
|
-
`
|
|
46
|
-
|
|
92
|
+
That's the loop: edit `evals/datasets/` with your own issues/repos (and
|
|
93
|
+
`workflows/` with your own workflows), then re-run. `init` gives you a
|
|
94
|
+
self-contained, version-controllable repo that **shadows** the built-in
|
|
95
|
+
workflows/skills and datasets by name — see [overlays](#your-own-workflows--datasets-overlays)
|
|
96
|
+
and the [configuration docs](https://lastlight.dev/docs/configuration/).
|
|
97
|
+
|
|
98
|
+
**Just kicking the tires?** Skip `init` and run the shipped samples directly:
|
|
47
99
|
|
|
48
100
|
```bash
|
|
49
|
-
|
|
50
|
-
|
|
101
|
+
npm install -g lastlight-evals
|
|
102
|
+
lastlight-evals run triage # or: npx lastlight-evals run triage
|
|
51
103
|
```
|
|
52
104
|
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
105
|
+
> Installing pulls in `lastlight` (and `agentic-pi`). `lastlight-evals` is a thin
|
|
106
|
+
> CLI on the `lastlight` package — it runs core's published `workflows/`,
|
|
107
|
+
> `skills/`, and `agent-context/`, so the evals exercise the **exact same assets
|
|
108
|
+
> production does**.
|
|
109
|
+
|
|
110
|
+
### Configuration (`.env`)
|
|
111
|
+
|
|
112
|
+
The only thing you must provide is a **model provider key**. Set it in the
|
|
113
|
+
environment, or drop a `.env` file in the directory you run from (the runner
|
|
114
|
+
loads it automatically — KEY=VALUE lines, no quotes needed):
|
|
115
|
+
|
|
116
|
+
```bash
|
|
117
|
+
# .env — at least ONE of these. Set keys only for the providers you want to eval.
|
|
118
|
+
OPENAI_API_KEY=sk-...
|
|
119
|
+
ANTHROPIC_API_KEY=sk-ant-...
|
|
120
|
+
FIREWORKS_API_KEY=fw-... # GLM / DeepSeek / GPT-OSS (open models)
|
|
121
|
+
OPENROUTER_API_KEY=sk-or-...
|
|
122
|
+
```
|
|
123
|
+
|
|
124
|
+
- The default run uses one model (`default` in `models.json`); `--compare` fans
|
|
125
|
+
out across the `compare` set, **running only the models whose key is present**
|
|
126
|
+
— so set the keys for the providers you care about and the rest are skipped.
|
|
127
|
+
- **No GitHub credentials are needed** — GitHub is mocked end to end. The harness
|
|
128
|
+
sets a dummy `GITHUB_TOKEN` internally; don't put a real one in `.env`.
|
|
129
|
+
- An `init`-scaffolded repo already gitignores `.env`, so your keys never get
|
|
130
|
+
committed.
|
|
131
|
+
|
|
132
|
+
## What a run does
|
|
133
|
+
|
|
134
|
+
Each eval `instance` (an issue fixture, optionally with a code fixture + held-out
|
|
135
|
+
tests) is taken through the **real** production workflow end to end:
|
|
136
|
+
|
|
137
|
+
1. An **in-process fake GitHub** starts, seeded with the issue and recording
|
|
138
|
+
every mutating call the agent makes.
|
|
139
|
+
2. For `code-fix`, the **workspace is seeded** with the fixture repo at its base
|
|
140
|
+
commit plus a local bare `origin`, so `git push` works fully offline.
|
|
141
|
+
3. The **real workflow YAML** (`issue-triage`, `build`, …) is loaded from
|
|
142
|
+
`lastlight` and run with `sandbox:"none"`, the agent's `github_*` tools
|
|
143
|
+
pointed at the fake, and approval gates disabled (so it never pauses).
|
|
144
|
+
4. The result is **graded deterministically** — no LLM judge:
|
|
145
|
+
- **behavioral** — the recorded GitHub calls (labels, comments, PRs) vs the
|
|
146
|
+
instance's `expect_github` / `triage_gold`.
|
|
147
|
+
- **execution** (code-fix) — the held-out tests are applied and run; the case
|
|
148
|
+
is *resolved* only if every `FAIL_TO_PASS` passes and every `PASS_TO_PASS`
|
|
149
|
+
stays green (SWE-bench's criterion).
|
|
150
|
+
5. Token usage, cost, and latency are collected per run.
|
|
151
|
+
|
|
152
|
+
Run multiple models and you get a side-by-side **scorecard** (HTML + JSON)
|
|
153
|
+
ranking them on pass rate, cost, and latency.
|
|
58
154
|
|
|
59
155
|
## Run it
|
|
60
156
|
|
|
61
157
|
```bash
|
|
62
158
|
# no tier args → interactively pick which tiers to run (one or all).
|
|
63
159
|
# Non-interactive (CI / piped) falls back to the cheapest default.
|
|
64
|
-
lastlight-evals run
|
|
160
|
+
lastlight-evals run
|
|
65
161
|
|
|
66
162
|
# name tiers explicitly to skip the prompt
|
|
67
163
|
lastlight-evals run triage
|
|
@@ -85,43 +181,91 @@ lastlight-evals run --overlay ~/work/lastlight-instance
|
|
|
85
181
|
# add your own datasets dir without an overlay
|
|
86
182
|
lastlight-evals run --datasets ~/my-evals/datasets
|
|
87
183
|
|
|
184
|
+
# CONFIG run type — eval a deployment's REAL per-step model config (different
|
|
185
|
+
# models per workflow phase, from the overlay's config.yaml) instead of forcing
|
|
186
|
+
# one model. This is the setup you actually ship. Try the bundled sample overlay:
|
|
187
|
+
lastlight-evals run code-fix --mode config --overlay examples/overlay
|
|
188
|
+
lastlight-evals run code-fix --mode config --overlay A --overlay B # 2 configs side-by-side
|
|
189
|
+
|
|
88
190
|
# ad-hoc model set / focus one instance / no browser
|
|
89
191
|
EVAL_MODELS="openai/gpt-5.5,anthropic/claude-sonnet-4-6" lastlight-evals run
|
|
90
192
|
EVAL_INSTANCE=off-by-one lastlight-evals run code-fix
|
|
91
193
|
lastlight-evals run triage --no-open
|
|
92
194
|
```
|
|
93
195
|
|
|
94
|
-
The
|
|
95
|
-
|
|
96
|
-
|
|
196
|
+
The report is a **JSON-driven dashboard**, not generated HTML — the harness only
|
|
197
|
+
ever writes `scorecard.json`, updating it (atomically) as the run proceeds. The
|
|
198
|
+
runner starts a tiny local server and opens `http://localhost:PORT` deep-linked
|
|
199
|
+
at the run, so you watch the scorecard fill in live (the SPA polls the JSON).
|
|
200
|
+
When the run finishes the server stays up so the dashboard keeps working — press
|
|
201
|
+
`Ctrl-C` to stop it. Each run lands in its **own** timestamped folder, so runs
|
|
202
|
+
accumulate instead of overwriting — `./eval-results/<tiers>/<runId>/` (override
|
|
203
|
+
the root with `LASTLIGHT_EVALS_OUT`), where `runId` is `<timestamp>-<git-sha>`:
|
|
97
204
|
|
|
98
|
-
- `
|
|
99
|
-
- `scorecard.json` — structured roll-up per model.
|
|
205
|
+
- `scorecard.json` — structured roll-up per model + per-instance results, carrying run `meta`.
|
|
100
206
|
- `predictions.jsonl` — SWE-bench predictions shape.
|
|
101
207
|
|
|
208
|
+
The dashboard's **overview** lists every run newest-first with a per-model trend
|
|
209
|
+
sparkline and links into each run's full scorecard; the **run view** is the
|
|
210
|
+
model-comparison table plus per-instance rows. To browse past runs anytime
|
|
211
|
+
without running models, start the server on its own:
|
|
212
|
+
|
|
213
|
+
```bash
|
|
214
|
+
lastlight-evals serve # opens the dashboard over ./eval-results
|
|
215
|
+
lastlight-evals serve --port 4319
|
|
216
|
+
```
|
|
217
|
+
|
|
102
218
|
Needs a provider key (`OPENAI_API_KEY` / `ANTHROPIC_API_KEY` /
|
|
103
219
|
`FIREWORKS_API_KEY` / `OPENROUTER_API_KEY`) in the environment or a cwd `.env`.
|
|
104
220
|
The runner exits non-zero **only** if the harness itself errors — a weak model
|
|
105
221
|
scoring poorly is the measurement, not a build failure.
|
|
106
222
|
|
|
223
|
+
### Two run types
|
|
224
|
+
|
|
225
|
+
A run compares N **arms** along one of two axes — pick with `--mode` (or, in a
|
|
226
|
+
TTY with no model flags, you're asked):
|
|
227
|
+
|
|
228
|
+
- **`models`** (default) — compare models, each **forced across every workflow
|
|
229
|
+
step**. `--model`/`--compare` select the set. Lands in `eval-results/<tier>/`
|
|
230
|
+
(or `<tier>-compare/`).
|
|
231
|
+
- **`config`** (`--mode config`) — run a deployment's **real per-step model
|
|
232
|
+
config**: the `models`/`variants` maps from an overlay's `config.yaml`, merged
|
|
233
|
+
over core's `config/default.yaml` exactly as production does, so each phase can
|
|
234
|
+
run on a different model. The arm is the config/overlay; pass `--overlay` more
|
|
235
|
+
than once to compare configs side-by-side, or re-run over time to compare as
|
|
236
|
+
you tweak prompts/skills/workflow/model-config. `--model` overrides a config's
|
|
237
|
+
`default` for quick what-ifs. Lands in `eval-results/<tier>-config/`, on its
|
|
238
|
+
own trend line. The run view shows a **Per-step models** panel with each
|
|
239
|
+
phase's resolved model. See [`examples/overlay`](examples/overlay) for a
|
|
240
|
+
ready-to-run sample.
|
|
241
|
+
|
|
107
242
|
## Your own workflows + datasets (overlays)
|
|
108
243
|
|
|
109
244
|
An **overlay** is a directory (often its own repo, like `lastlight-instance`)
|
|
110
245
|
that carries its own `workflows/` / `skills/` / `agent-context/` (which shadow
|
|
111
|
-
the core built-ins by name) and its own `evals/datasets/`.
|
|
246
|
+
the core built-ins by name) and its own `evals/datasets/`. It's the same
|
|
247
|
+
deployment-overlay mechanism the production harness uses — see the [Last Light
|
|
248
|
+
configuration docs](https://lastlight.dev/docs/configuration/) for the full
|
|
249
|
+
story. One flag wires both:
|
|
112
250
|
|
|
113
251
|
```bash
|
|
114
252
|
lastlight-evals run --overlay ~/work/lastlight-instance # or LASTLIGHT_OVERLAY_DIR
|
|
115
253
|
```
|
|
116
254
|
|
|
117
|
-
- Overlay **workflows/skills** are layered over core via core's asset
|
|
118
|
-
(same mechanism the
|
|
255
|
+
- Overlay **workflows/skills** are layered over core via core's [asset
|
|
256
|
+
overlay](https://lastlight.dev/docs/configuration/) (same mechanism the
|
|
257
|
+
production harness uses).
|
|
119
258
|
- Overlay **datasets** are discovered at `<overlay>/evals/datasets/<tier>/`, and
|
|
120
259
|
shadow built-in tiers of the same name.
|
|
121
260
|
- An overlay **`evals/models.json`** is picked up automatically (or pass
|
|
122
261
|
`--models-file`).
|
|
123
262
|
|
|
124
|
-
### `lastlight-evals init [dir]` — scaffold
|
|
263
|
+
### `lastlight-evals init [dir]` — scaffold an evals workspace
|
|
264
|
+
|
|
265
|
+
Two shapes, depending on whether you already have a deployment overlay repo:
|
|
266
|
+
|
|
267
|
+
**Plain** — a self-contained overlay+evals repo (its own `workflows/` `skills/`
|
|
268
|
+
`agent-context/` + `evals/`):
|
|
125
269
|
|
|
126
270
|
```bash
|
|
127
271
|
lastlight-evals init my-evals
|
|
@@ -133,6 +277,21 @@ Scaffolds `workflows/` `skills/` `agent-context/` (empty, to fill in),
|
|
|
133
277
|
`config.yaml`, and a `.gitignore`/`README`, then offers to `git init` + create a
|
|
134
278
|
private GitHub repo via `gh` (reusing core's `lastlight server setup` flow).
|
|
135
279
|
|
|
280
|
+
**Separate** (`--clone`) — the recommended shape when you already have a
|
|
281
|
+
deployment overlay (e.g. `lastlight-instance`) and want to eval **its** config.
|
|
282
|
+
The overlay is cloned into `<dir>/instance/` (its own git checkout, git-ignored)
|
|
283
|
+
with the evals at the workspace root; a bare run auto-detects both, no flags:
|
|
284
|
+
|
|
285
|
+
```bash
|
|
286
|
+
lastlight-evals init my-evals --clone cliftonc/lastlight-instance
|
|
287
|
+
cd my-evals && lastlight-evals run # auto: overlay ./instance + ./evals/datasets
|
|
288
|
+
```
|
|
289
|
+
|
|
290
|
+
This is exactly what the [`lastlight-evals` agent skill](#easiest-let-the-last-light-agent-skill-set-up-your-own-workspace)
|
|
291
|
+
does for you. Update the overlay later with `cd instance && git pull`; your evals
|
|
292
|
+
stay out of the deployment repo. Run `lastlight-evals init --help` for all flags
|
|
293
|
+
(`--yes`, `--no-git`, …).
|
|
294
|
+
|
|
136
295
|
## Datasets & tiers
|
|
137
296
|
|
|
138
297
|
A **tier** is a directory containing `instances.json` (+ an optional `tier.json`
|