lastlight-evals 0.1.1 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (36) hide show
  1. package/README.md +198 -39
  2. package/dashboard/dist/assets/index-YwScAAxm.js +238 -0
  3. package/dashboard/dist/assets/index-uU9Sj3M4.css +1 -0
  4. package/dashboard/dist/index.html +19 -0
  5. package/dashboard/dist/logo.png +0 -0
  6. package/datasets/code-fix/repos/codefix__date-range-off-by-one/.github/workflows/ci.yml +16 -0
  7. package/datasets/code-fix/repos/codefix__date-range-off-by-one/README.md +7 -1
  8. package/datasets/code-fix/repos/codefix__date-range-off-by-one/package.json +3 -1
  9. package/datasets/code-fix/repos/codefix__date-range-off-by-one/scripts/lint.mjs +25 -0
  10. package/datasets/code-fix/repos/codefix__date-range-off-by-one/test/date-range.test.ts +14 -0
  11. package/datasets/code-fix/repos/codefix__date-range-off-by-one/tsconfig.json +13 -0
  12. package/datasets/code-fix/tests/codefix__date-range-off-by-one/date-range.test.ts +5 -5
  13. package/dist/config.js +100 -0
  14. package/dist/config.js.map +1 -0
  15. package/dist/init.js +229 -32
  16. package/dist/init.js.map +1 -1
  17. package/dist/mechanism.test.js +41 -0
  18. package/dist/mechanism.test.js.map +1 -1
  19. package/dist/metrics.js +91 -3
  20. package/dist/metrics.js.map +1 -1
  21. package/dist/paths.js +51 -1
  22. package/dist/paths.js.map +1 -1
  23. package/dist/report.js +128 -6
  24. package/dist/report.js.map +1 -1
  25. package/dist/run-instance.js +143 -14
  26. package/dist/run-instance.js.map +1 -1
  27. package/dist/run.js +401 -89
  28. package/dist/run.js.map +1 -1
  29. package/dist/serve.js +136 -0
  30. package/dist/serve.js.map +1 -0
  31. package/examples/overlay/README.md +30 -0
  32. package/examples/overlay/config.yaml +36 -0
  33. package/examples/overlay-anthropic/config.yaml +27 -0
  34. package/package.json +16 -6
  35. package/dist/html-report.js +0 -325
  36. package/dist/html-report.js.map +0 -1
package/README.md CHANGED
@@ -1,16 +1,31 @@
1
1
  # lastlight-evals
2
2
 
3
- A standalone, **SWE-bench-compatible** eval harness for [Last
4
- Light](https://github.com/cliftonc/lastlight) workflows. It drives the **real**
5
- production workflows (`issue-triage`, `build`, …) — their actual prompts and
6
- skills — against a **mocked GitHub**, grades the result deterministically, and
7
- prints a model-comparison scorecard. It answers "what do we expect from the
8
- agent, and which model does it best?"
9
-
10
- Nothing here talks to real GitHub. The agent's `github_*` tool calls are served
11
- by an in-process fake (seeded + recording), and `git push` goes to a local bare
12
- repo. The only deviations from production are the two we can't do unattended:
13
- approval gates are disabled and outward side-effects are mocked.
3
+ ### Which model should run your agent? Find out — with receipts.
4
+
5
+ [![Last Light eval scorecard — 9 models compared on pass rate, cost, and latency](docs/scorecard-v2.png)](https://evals.lastlight.dev/)
6
+
7
+ > **[▶ Explore the live scorecard](https://evals.lastlight.dev/)** — interactive, with per-instance detail. (Above: code-fix tier across 9 models.)
8
+
9
+ `lastlight-evals` takes [**Last Light**](https://lastlight.dev)'s *real*
10
+ production workflows — the actual prompts, skills, and agent loop that ship — and
11
+ runs them end to end against a fully mocked GitHub, for whatever models you throw
12
+ at it. No toy benchmarks, no LLM-as-judge: every run is graded **deterministically**
13
+ (did the agent apply the right labels? did the held-out tests turn green?), then
14
+ ranked side by side on **pass rate, cost, and latency**.
15
+
16
+ The payoff is one scorecard that tells you, for *your* workflows, exactly what
17
+ each model delivers — and what it costs you per run. Swap a model, re-run, see the
18
+ difference. Drop in your own issues and repos and it evaluates *your* agent.
19
+
20
+ > 🛰️ Part of [**Last Light**](https://lastlight.dev) — the AI agent that triages,
21
+ > reviews, and fixes your GitHub repos.
22
+ > **[lastlight.dev](https://lastlight.dev)** · [Core repo](https://github.com/cliftonc/lastlight) · [Eval repo](https://github.com/cliftonc/lastlight-evals)
23
+
24
+ It's **SWE-bench-compatible**, and nothing here touches real GitHub: the agent's
25
+ `github_*` tool calls are served by an in-process fake (seeded + recording) and
26
+ `git push` goes to a local bare repo. The only deviations from production are the
27
+ two we can't do unattended — approval gates are disabled and outward side-effects
28
+ are mocked. Everything else is exactly what ships.
14
29
 
15
30
  ```
16
31
  instance (SWE-bench shape)
@@ -28,40 +43,121 @@ instance (SWE-bench shape)
28
43
  > (the base-URL mock, static-token mode, the no-clone seeding trick, the
29
44
  > asset-bootstrap footgun, the metrics drain).
30
45
 
31
- ## How it depends on Last Light
46
+ ## Get started
47
+
48
+ Needs **Node 24+** and a provider API key.
49
+
50
+ ### Easiest: let the Last Light agent skill set up *your own* workspace
51
+
52
+ Want to eval **your own** deployment — your workflows, your agent persona, your
53
+ config — not just the shipped samples? If you drive
54
+ [Last Light](https://lastlight.dev) from an agent (e.g. Claude Code), install its
55
+ skills once and then just *ask* — no flags to remember:
56
+
57
+ ```bash
58
+ lastlight skills install # installs the Last Light agent skills
59
+ ```
60
+
61
+ Then, in a **new empty folder**, tell your agent (point it at *your* instance
62
+ overlay repo):
63
+
64
+ > *Let's set up an evals workspace here, using my existing Last Light instance
65
+ > config in `cliftonc/lastlight-instance`.*
66
+
67
+ The `lastlight-evals` skill scaffolds the workspace, clones your overlay into
68
+ `instance/`, seeds the sample datasets, and wires it all up — under the hood it
69
+ runs `lastlight-evals init . --clone cliftonc/lastlight-instance`, after which a
70
+ bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
71
+ overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
72
+ models you care about. Prefer to drive it by hand? Keep reading.
32
73
 
33
- `lastlight-evals` is a thin CLI on top of the `lastlight` npm package. It imports
34
- exactly four things from core's public `lastlight/evals` barrel —
35
- `getWorkflow`, `runWorkflow`, `ExecutorConfig`, `TemplateContext` — plus the
36
- `gh`-repo bootstrap helpers used by `init`. Core ships its `workflows/`,
37
- `skills/`, and `agent-context/` in the package, so the evals run the same assets
38
- core does.
74
+ ### Manual: scaffold with `init`
75
+
76
+ The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
77
+ (your workflows + your datasets, seeded from the built-in samples) and optionally
78
+ creates a private GitHub repo for it:
39
79
 
40
80
  ```bash
41
- npm install # installs `lastlight` (and agentic-pi as a peer)
81
+ npm install -g lastlight-evals
82
+ export OPENAI_API_KEY=... # or ANTHROPIC_ / FIREWORKS_ / OPENROUTER_
83
+
84
+ # 1. Scaffold your workspace (offers to `git init` + `gh repo create`).
85
+ lastlight-evals init my-evals
86
+ cd my-evals
87
+
88
+ # 2. Run it — drives the real workflows against your datasets, prints a scorecard.
89
+ lastlight-evals run --overlay .
42
90
  ```
43
91
 
44
- **Local development against an un-published core.** Until the matching
45
- `lastlight` version is on npm — or whenever you want to eval your working-tree
46
- core — link a checkout:
92
+ That's the loop: edit `evals/datasets/` with your own issues/repos (and
93
+ `workflows/` with your own workflows), then re-run. `init` gives you a
94
+ self-contained, version-controllable repo that **shadows** the built-in
95
+ workflows/skills and datasets by name — see [overlays](#your-own-workflows--datasets-overlays)
96
+ and the [configuration docs](https://lastlight.dev/docs/configuration/).
97
+
98
+ **Just kicking the tires?** Skip `init` and run the shipped samples directly:
47
99
 
48
100
  ```bash
49
- cd ../lastlight && npm run build && npm link
50
- cd ../lastlight-evals && npm link lastlight
101
+ npm install -g lastlight-evals
102
+ lastlight-evals run triage # or: npx lastlight-evals run triage
51
103
  ```
52
104
 
53
- Or set `LASTLIGHT_CORE_DIR=/path/to/lastlight` to point just the **asset roots**
54
- (workflows/skills/agent-context — the bulk of what `lastlight server update`
55
- ships) at a checkout without touching the npm dep. (The runner *code* still
56
- comes from `node_modules/lastlight`; use `npm link` to exercise working-tree
57
- engine code too.)
105
+ > Installing pulls in `lastlight` (and `agentic-pi`). `lastlight-evals` is a thin
106
+ > CLI on the `lastlight` package — it runs core's published `workflows/`,
107
+ > `skills/`, and `agent-context/`, so the evals exercise the **exact same assets
108
+ > production does**.
109
+
110
+ ### Configuration (`.env`)
111
+
112
+ The only thing you must provide is a **model provider key**. Set it in the
113
+ environment, or drop a `.env` file in the directory you run from (the runner
114
+ loads it automatically — KEY=VALUE lines, no quotes needed):
115
+
116
+ ```bash
117
+ # .env — at least ONE of these. Set keys only for the providers you want to eval.
118
+ OPENAI_API_KEY=sk-...
119
+ ANTHROPIC_API_KEY=sk-ant-...
120
+ FIREWORKS_API_KEY=fw-... # GLM / DeepSeek / GPT-OSS (open models)
121
+ OPENROUTER_API_KEY=sk-or-...
122
+ ```
123
+
124
+ - The default run uses one model (`default` in `models.json`); `--compare` fans
125
+ out across the `compare` set, **running only the models whose key is present**
126
+ — so set the keys for the providers you care about and the rest are skipped.
127
+ - **No GitHub credentials are needed** — GitHub is mocked end to end. The harness
128
+ sets a dummy `GITHUB_TOKEN` internally; don't put a real one in `.env`.
129
+ - An `init`-scaffolded repo already gitignores `.env`, so your keys never get
130
+ committed.
131
+
132
+ ## What a run does
133
+
134
+ Each eval `instance` (an issue fixture, optionally with a code fixture + held-out
135
+ tests) is taken through the **real** production workflow end to end:
136
+
137
+ 1. An **in-process fake GitHub** starts, seeded with the issue and recording
138
+ every mutating call the agent makes.
139
+ 2. For `code-fix`, the **workspace is seeded** with the fixture repo at its base
140
+ commit plus a local bare `origin`, so `git push` works fully offline.
141
+ 3. The **real workflow YAML** (`issue-triage`, `build`, …) is loaded from
142
+ `lastlight` and run with `sandbox:"none"`, the agent's `github_*` tools
143
+ pointed at the fake, and approval gates disabled (so it never pauses).
144
+ 4. The result is **graded deterministically** — no LLM judge:
145
+ - **behavioral** — the recorded GitHub calls (labels, comments, PRs) vs the
146
+ instance's `expect_github` / `triage_gold`.
147
+ - **execution** (code-fix) — the held-out tests are applied and run; the case
148
+ is *resolved* only if every `FAIL_TO_PASS` passes and every `PASS_TO_PASS`
149
+ stays green (SWE-bench's criterion).
150
+ 5. Token usage, cost, and latency are collected per run.
151
+
152
+ Run multiple models and you get a side-by-side **scorecard** (HTML + JSON)
153
+ ranking them on pass rate, cost, and latency.
58
154
 
59
155
  ## Run it
60
156
 
61
157
  ```bash
62
158
  # no tier args → interactively pick which tiers to run (one or all).
63
159
  # Non-interactive (CI / piped) falls back to the cheapest default.
64
- lastlight-evals run # (or: npm run eval)
160
+ lastlight-evals run
65
161
 
66
162
  # name tiers explicitly to skip the prompt
67
163
  lastlight-evals run triage
@@ -85,43 +181,91 @@ lastlight-evals run --overlay ~/work/lastlight-instance
85
181
  # add your own datasets dir without an overlay
86
182
  lastlight-evals run --datasets ~/my-evals/datasets
87
183
 
184
+ # CONFIG run type — eval a deployment's REAL per-step model config (different
185
+ # models per workflow phase, from the overlay's config.yaml) instead of forcing
186
+ # one model. This is the setup you actually ship. Try the bundled sample overlay:
187
+ lastlight-evals run code-fix --mode config --overlay examples/overlay
188
+ lastlight-evals run code-fix --mode config --overlay A --overlay B # 2 configs side-by-side
189
+
88
190
  # ad-hoc model set / focus one instance / no browser
89
191
  EVAL_MODELS="openai/gpt-5.5,anthropic/claude-sonnet-4-6" lastlight-evals run
90
192
  EVAL_INSTANCE=off-by-one lastlight-evals run code-fix
91
193
  lastlight-evals run triage --no-open
92
194
  ```
93
195
 
94
- The runner opens `index.html` and **rewrites it after every run** (auto-refresh,
95
- preserving the active tab + scroll), so you watch the scorecard fill in live.
96
- Output lands under `./eval-results/<tiers>/` (override with `LASTLIGHT_EVALS_OUT`):
196
+ The report is a **JSON-driven dashboard**, not generated HTML — the harness only
197
+ ever writes `scorecard.json`, updating it (atomically) as the run proceeds. The
198
+ runner starts a tiny local server and opens `http://localhost:PORT` deep-linked
199
+ at the run, so you watch the scorecard fill in live (the SPA polls the JSON).
200
+ When the run finishes the server stays up so the dashboard keeps working — press
201
+ `Ctrl-C` to stop it. Each run lands in its **own** timestamped folder, so runs
202
+ accumulate instead of overwriting — `./eval-results/<tiers>/<runId>/` (override
203
+ the root with `LASTLIGHT_EVALS_OUT`), where `runId` is `<timestamp>-<git-sha>`:
97
204
 
98
- - `index.html` — styled scorecard.
99
- - `scorecard.json` — structured roll-up per model.
205
+ - `scorecard.json` — structured roll-up per model + per-instance results, carrying run `meta`.
100
206
  - `predictions.jsonl` — SWE-bench predictions shape.
101
207
 
208
+ The dashboard's **overview** lists every run newest-first with a per-model trend
209
+ sparkline and links into each run's full scorecard; the **run view** is the
210
+ model-comparison table plus per-instance rows. To browse past runs anytime
211
+ without running models, start the server on its own:
212
+
213
+ ```bash
214
+ lastlight-evals serve # opens the dashboard over ./eval-results
215
+ lastlight-evals serve --port 4319
216
+ ```
217
+
102
218
  Needs a provider key (`OPENAI_API_KEY` / `ANTHROPIC_API_KEY` /
103
219
  `FIREWORKS_API_KEY` / `OPENROUTER_API_KEY`) in the environment or a cwd `.env`.
104
220
  The runner exits non-zero **only** if the harness itself errors — a weak model
105
221
  scoring poorly is the measurement, not a build failure.
106
222
 
223
+ ### Two run types
224
+
225
+ A run compares N **arms** along one of two axes — pick with `--mode` (or, in a
226
+ TTY with no model flags, you're asked):
227
+
228
+ - **`models`** (default) — compare models, each **forced across every workflow
229
+ step**. `--model`/`--compare` select the set. Lands in `eval-results/<tier>/`
230
+ (or `<tier>-compare/`).
231
+ - **`config`** (`--mode config`) — run a deployment's **real per-step model
232
+ config**: the `models`/`variants` maps from an overlay's `config.yaml`, merged
233
+ over core's `config/default.yaml` exactly as production does, so each phase can
234
+ run on a different model. The arm is the config/overlay; pass `--overlay` more
235
+ than once to compare configs side-by-side, or re-run over time to compare as
236
+ you tweak prompts/skills/workflow/model-config. `--model` overrides a config's
237
+ `default` for quick what-ifs. Lands in `eval-results/<tier>-config/`, on its
238
+ own trend line. The run view shows a **Per-step models** panel with each
239
+ phase's resolved model. See [`examples/overlay`](examples/overlay) for a
240
+ ready-to-run sample.
241
+
107
242
  ## Your own workflows + datasets (overlays)
108
243
 
109
244
  An **overlay** is a directory (often its own repo, like `lastlight-instance`)
110
245
  that carries its own `workflows/` / `skills/` / `agent-context/` (which shadow
111
- the core built-ins by name) and its own `evals/datasets/`. One flag wires both:
246
+ the core built-ins by name) and its own `evals/datasets/`. It's the same
247
+ deployment-overlay mechanism the production harness uses — see the [Last Light
248
+ configuration docs](https://lastlight.dev/docs/configuration/) for the full
249
+ story. One flag wires both:
112
250
 
113
251
  ```bash
114
252
  lastlight-evals run --overlay ~/work/lastlight-instance # or LASTLIGHT_OVERLAY_DIR
115
253
  ```
116
254
 
117
- - Overlay **workflows/skills** are layered over core via core's asset overlay
118
- (same mechanism the production harness uses).
255
+ - Overlay **workflows/skills** are layered over core via core's [asset
256
+ overlay](https://lastlight.dev/docs/configuration/) (same mechanism the
257
+ production harness uses).
119
258
  - Overlay **datasets** are discovered at `<overlay>/evals/datasets/<tier>/`, and
120
259
  shadow built-in tiers of the same name.
121
260
  - An overlay **`evals/models.json`** is picked up automatically (or pass
122
261
  `--models-file`).
123
262
 
124
- ### `lastlight-evals init [dir]` — scaffold a fresh overlay+evals repo
263
+ ### `lastlight-evals init [dir]` — scaffold an evals workspace
264
+
265
+ Two shapes, depending on whether you already have a deployment overlay repo:
266
+
267
+ **Plain** — a self-contained overlay+evals repo (its own `workflows/` `skills/`
268
+ `agent-context/` + `evals/`):
125
269
 
126
270
  ```bash
127
271
  lastlight-evals init my-evals
@@ -133,6 +277,21 @@ Scaffolds `workflows/` `skills/` `agent-context/` (empty, to fill in),
133
277
  `config.yaml`, and a `.gitignore`/`README`, then offers to `git init` + create a
134
278
  private GitHub repo via `gh` (reusing core's `lastlight server setup` flow).
135
279
 
280
+ **Separate** (`--clone`) — the recommended shape when you already have a
281
+ deployment overlay (e.g. `lastlight-instance`) and want to eval **its** config.
282
+ The overlay is cloned into `<dir>/instance/` (its own git checkout, git-ignored)
283
+ with the evals at the workspace root; a bare run auto-detects both, no flags:
284
+
285
+ ```bash
286
+ lastlight-evals init my-evals --clone cliftonc/lastlight-instance
287
+ cd my-evals && lastlight-evals run # auto: overlay ./instance + ./evals/datasets
288
+ ```
289
+
290
+ This is exactly what the [`lastlight-evals` agent skill](#easiest-let-the-last-light-agent-skill-set-up-your-own-workspace)
291
+ does for you. Update the overlay later with `cd instance && git pull`; your evals
292
+ stay out of the deployment repo. Run `lastlight-evals init --help` for all flags
293
+ (`--yes`, `--no-git`, …).
294
+
136
295
  ## Datasets & tiers
137
296
 
138
297
  A **tier** is a directory containing `instances.json` (+ an optional `tier.json`