lastlight-evals 0.1.2 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (34) hide show
  1. package/README.md +98 -12
  2. package/dashboard/dist/assets/index-YwScAAxm.js +238 -0
  3. package/dashboard/dist/assets/index-uU9Sj3M4.css +1 -0
  4. package/dashboard/dist/index.html +19 -0
  5. package/dashboard/dist/logo.png +0 -0
  6. package/datasets/code-fix/repos/codefix__date-range-off-by-one/.github/workflows/ci.yml +16 -0
  7. package/datasets/code-fix/repos/codefix__date-range-off-by-one/README.md +7 -1
  8. package/datasets/code-fix/repos/codefix__date-range-off-by-one/package.json +3 -1
  9. package/datasets/code-fix/repos/codefix__date-range-off-by-one/scripts/lint.mjs +25 -0
  10. package/datasets/code-fix/repos/codefix__date-range-off-by-one/test/date-range.test.ts +14 -0
  11. package/datasets/code-fix/repos/codefix__date-range-off-by-one/tsconfig.json +13 -0
  12. package/datasets/code-fix/tests/codefix__date-range-off-by-one/date-range.test.ts +5 -5
  13. package/dist/config.js +100 -0
  14. package/dist/config.js.map +1 -0
  15. package/dist/mechanism.test.js +41 -0
  16. package/dist/mechanism.test.js.map +1 -1
  17. package/dist/metrics.js +83 -0
  18. package/dist/metrics.js.map +1 -1
  19. package/dist/paths.js +51 -1
  20. package/dist/paths.js.map +1 -1
  21. package/dist/report.js +112 -3
  22. package/dist/report.js.map +1 -1
  23. package/dist/run-instance.js +141 -14
  24. package/dist/run-instance.js.map +1 -1
  25. package/dist/run.js +342 -112
  26. package/dist/run.js.map +1 -1
  27. package/dist/serve.js +136 -0
  28. package/dist/serve.js.map +1 -0
  29. package/examples/overlay/README.md +30 -0
  30. package/examples/overlay/config.yaml +36 -0
  31. package/examples/overlay-anthropic/config.yaml +27 -0
  32. package/package.json +16 -6
  33. package/dist/html-report.js +0 -325
  34. package/dist/html-report.js.map +0 -1
package/README.md CHANGED
@@ -2,9 +2,9 @@
2
2
 
3
3
  ### Which model should run your agent? Find out — with receipts.
4
4
 
5
- [![Last Light eval scorecard — 9 models compared on pass rate, cost, and latency](docs/scorecard.png)](https://cliftonc.github.io/lastlight-evals/)
5
+ [![Last Light eval scorecard — 9 models compared on pass rate, cost, and latency](docs/scorecard-v2.png)](https://evals.lastlight.dev/)
6
6
 
7
- > **[▶ Explore the live scorecard](https://cliftonc.github.io/lastlight-evals/)** — interactive, with per-instance detail. (Above: triage tier across 9 models.)
7
+ > **[▶ Explore the live scorecard](https://evals.lastlight.dev/)** — interactive, with per-instance detail. (Above: code-fix tier across 9 models.)
8
8
 
9
9
  `lastlight-evals` takes [**Last Light**](https://lastlight.dev)'s *real*
10
10
  production workflows — the actual prompts, skills, and agent loop that ship — and
@@ -45,10 +45,37 @@ instance (SWE-bench shape)
45
45
 
46
46
  ## Get started
47
47
 
48
- Needs **Node 24+** and a provider API key. The fastest path is
49
- **`init`** — it scaffolds *your own* evals workspace (your workflows + your
50
- datasets, seeded from the built-in samples) and optionally creates a private
51
- GitHub repo for it:
48
+ Needs **Node 24+** and a provider API key.
49
+
50
+ ### Easiest: let the Last Light agent skill set up *your own* workspace
51
+
52
+ Want to eval **your own** deployment — your workflows, your agent persona, your
53
+ config — not just the shipped samples? If you drive
54
+ [Last Light](https://lastlight.dev) from an agent (e.g. Claude Code), install its
55
+ skills once and then just *ask* — no flags to remember:
56
+
57
+ ```bash
58
+ lastlight skills install # installs the Last Light agent skills
59
+ ```
60
+
61
+ Then, in a **new empty folder**, tell your agent (point it at *your* instance
62
+ overlay repo):
63
+
64
+ > *Let's set up an evals workspace here, using my existing Last Light instance
65
+ > config in `cliftonc/lastlight-instance`.*
66
+
67
+ The `lastlight-evals` skill scaffolds the workspace, clones your overlay into
68
+ `instance/`, seeds the sample datasets, and wires it all up — under the hood it
69
+ runs `lastlight-evals init . --clone cliftonc/lastlight-instance`, after which a
70
+ bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
71
+ overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
72
+ models you care about. Prefer to drive it by hand? Keep reading.
73
+
74
+ ### Manual: scaffold with `init`
75
+
76
+ The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
77
+ (your workflows + your datasets, seeded from the built-in samples) and optionally
78
+ creates a private GitHub repo for it:
52
79
 
53
80
  ```bash
54
81
  npm install -g lastlight-evals
@@ -154,25 +181,64 @@ lastlight-evals run --overlay ~/work/lastlight-instance
154
181
  # add your own datasets dir without an overlay
155
182
  lastlight-evals run --datasets ~/my-evals/datasets
156
183
 
184
+ # CONFIG run type — eval a deployment's REAL per-step model config (different
185
+ # models per workflow phase, from the overlay's config.yaml) instead of forcing
186
+ # one model. This is the setup you actually ship. Try the bundled sample overlay:
187
+ lastlight-evals run code-fix --mode config --overlay examples/overlay
188
+ lastlight-evals run code-fix --mode config --overlay A --overlay B # 2 configs side-by-side
189
+
157
190
  # ad-hoc model set / focus one instance / no browser
158
191
  EVAL_MODELS="openai/gpt-5.5,anthropic/claude-sonnet-4-6" lastlight-evals run
159
192
  EVAL_INSTANCE=off-by-one lastlight-evals run code-fix
160
193
  lastlight-evals run triage --no-open
161
194
  ```
162
195
 
163
- The runner opens `index.html` and **rewrites it after every run** (auto-refresh,
164
- preserving the active tab + scroll), so you watch the scorecard fill in live.
165
- Output lands under `./eval-results/<tiers>/` (override with `LASTLIGHT_EVALS_OUT`):
196
+ The report is a **JSON-driven dashboard**, not generated HTML — the harness only
197
+ ever writes `scorecard.json`, updating it (atomically) as the run proceeds. The
198
+ runner starts a tiny local server and opens `http://localhost:PORT` deep-linked
199
+ at the run, so you watch the scorecard fill in live (the SPA polls the JSON).
200
+ When the run finishes the server stays up so the dashboard keeps working — press
201
+ `Ctrl-C` to stop it. Each run lands in its **own** timestamped folder, so runs
202
+ accumulate instead of overwriting — `./eval-results/<tiers>/<runId>/` (override
203
+ the root with `LASTLIGHT_EVALS_OUT`), where `runId` is `<timestamp>-<git-sha>`:
166
204
 
167
- - `index.html` — styled scorecard.
168
- - `scorecard.json` — structured roll-up per model.
205
+ - `scorecard.json` — structured roll-up per model + per-instance results, carrying run `meta`.
169
206
  - `predictions.jsonl` — SWE-bench predictions shape.
170
207
 
208
+ The dashboard's **overview** lists every run newest-first with a per-model trend
209
+ sparkline and links into each run's full scorecard; the **run view** is the
210
+ model-comparison table plus per-instance rows. To browse past runs anytime
211
+ without running models, start the server on its own:
212
+
213
+ ```bash
214
+ lastlight-evals serve # opens the dashboard over ./eval-results
215
+ lastlight-evals serve --port 4319
216
+ ```
217
+
171
218
  Needs a provider key (`OPENAI_API_KEY` / `ANTHROPIC_API_KEY` /
172
219
  `FIREWORKS_API_KEY` / `OPENROUTER_API_KEY`) in the environment or a cwd `.env`.
173
220
  The runner exits non-zero **only** if the harness itself errors — a weak model
174
221
  scoring poorly is the measurement, not a build failure.
175
222
 
223
+ ### Two run types
224
+
225
+ A run compares N **arms** along one of two axes — pick with `--mode` (or, in a
226
+ TTY with no model flags, you're asked):
227
+
228
+ - **`models`** (default) — compare models, each **forced across every workflow
229
+ step**. `--model`/`--compare` select the set. Lands in `eval-results/<tier>/`
230
+ (or `<tier>-compare/`).
231
+ - **`config`** (`--mode config`) — run a deployment's **real per-step model
232
+ config**: the `models`/`variants` maps from an overlay's `config.yaml`, merged
233
+ over core's `config/default.yaml` exactly as production does, so each phase can
234
+ run on a different model. The arm is the config/overlay; pass `--overlay` more
235
+ than once to compare configs side-by-side, or re-run over time to compare as
236
+ you tweak prompts/skills/workflow/model-config. `--model` overrides a config's
237
+ `default` for quick what-ifs. Lands in `eval-results/<tier>-config/`, on its
238
+ own trend line. The run view shows a **Per-step models** panel with each
239
+ phase's resolved model. See [`examples/overlay`](examples/overlay) for a
240
+ ready-to-run sample.
241
+
176
242
  ## Your own workflows + datasets (overlays)
177
243
 
178
244
  An **overlay** is a directory (often its own repo, like `lastlight-instance`)
@@ -194,7 +260,12 @@ lastlight-evals run --overlay ~/work/lastlight-instance # or LASTLIGHT_OVERL
194
260
  - An overlay **`evals/models.json`** is picked up automatically (or pass
195
261
  `--models-file`).
196
262
 
197
- ### `lastlight-evals init [dir]` — scaffold a fresh overlay+evals repo
263
+ ### `lastlight-evals init [dir]` — scaffold an evals workspace
264
+
265
+ Two shapes, depending on whether you already have a deployment overlay repo:
266
+
267
+ **Plain** — a self-contained overlay+evals repo (its own `workflows/` `skills/`
268
+ `agent-context/` + `evals/`):
198
269
 
199
270
  ```bash
200
271
  lastlight-evals init my-evals
@@ -206,6 +277,21 @@ Scaffolds `workflows/` `skills/` `agent-context/` (empty, to fill in),
206
277
  `config.yaml`, and a `.gitignore`/`README`, then offers to `git init` + create a
207
278
  private GitHub repo via `gh` (reusing core's `lastlight server setup` flow).
208
279
 
280
+ **Separate** (`--clone`) — the recommended shape when you already have a
281
+ deployment overlay (e.g. `lastlight-instance`) and want to eval **its** config.
282
+ The overlay is cloned into `<dir>/instance/` (its own git checkout, git-ignored)
283
+ with the evals at the workspace root; a bare run auto-detects both, no flags:
284
+
285
+ ```bash
286
+ lastlight-evals init my-evals --clone cliftonc/lastlight-instance
287
+ cd my-evals && lastlight-evals run # auto: overlay ./instance + ./evals/datasets
288
+ ```
289
+
290
+ This is exactly what the [`lastlight-evals` agent skill](#easiest-let-the-last-light-agent-skill-set-up-your-own-workspace)
291
+ does for you. Update the overlay later with `cd instance && git pull`; your evals
292
+ stay out of the deployment repo. Run `lastlight-evals init --help` for all flags
293
+ (`--yes`, `--no-git`, …).
294
+
209
295
  ## Datasets & tiers
210
296
 
211
297
  A **tier** is a directory containing `instances.json` (+ an optional `tier.json`