lastlight-evals 0.4.0 → 0.6.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +117 -8
- package/dashboard/dist/assets/index-D5AO7YF_.js +274 -0
- package/dashboard/dist/assets/index-DllAFtF8.css +1 -0
- package/dashboard/dist/index.html +20 -3
- package/datasets/pr-review/README.md +58 -0
- package/datasets/pr-review/martian-leaderboard.json +5785 -0
- package/datasets/pr-review/tier.json +5 -0
- package/dist/add-case.js +188 -12
- package/dist/add-case.js.map +1 -1
- package/dist/discovery.js +10 -4
- package/dist/discovery.js.map +1 -1
- package/dist/fake-github.js +187 -6
- package/dist/fake-github.js.map +1 -1
- package/dist/grade.js +252 -0
- package/dist/grade.js.map +1 -1
- package/dist/judge.js +145 -0
- package/dist/judge.js.map +1 -0
- package/dist/mechanism.test.js +332 -3
- package/dist/mechanism.test.js.map +1 -1
- package/dist/report.js +123 -1
- package/dist/report.js.map +1 -1
- package/dist/run-instance.js +152 -14
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +175 -36
- package/dist/run.js.map +1 -1
- package/dist/sandbox-preflight.js +107 -0
- package/dist/sandbox-preflight.js.map +1 -0
- package/dist/seed.js +223 -0
- package/dist/seed.js.map +1 -1
- package/models.json +1 -0
- package/package.json +2 -2
- package/dashboard/dist/assets/index-D6qxQ8uN.js +0 -258
- package/dashboard/dist/assets/index-DYwJHkLo.css +0 -1
package/README.md
CHANGED
|
@@ -9,9 +9,10 @@
|
|
|
9
9
|
`lastlight-evals` takes [**Last Light**](https://lastlight.dev)'s *real*
|
|
10
10
|
production workflows — the actual prompts, skills, and agent loop that ship — and
|
|
11
11
|
runs them end to end against a fully mocked GitHub, for whatever models you throw
|
|
12
|
-
at it. No toy benchmarks
|
|
13
|
-
|
|
14
|
-
|
|
12
|
+
at it. No toy benchmarks: grading is **deterministic** by default (did the agent
|
|
13
|
+
apply the right labels? did the held-out tests turn green?) — with one scoped
|
|
14
|
+
LLM-judge for the `pr-review` tier's precision/recall/F-beta (F1 by default) — then ranked side by
|
|
15
|
+
side on **pass/score rate, cost, and latency**.
|
|
15
16
|
|
|
16
17
|
The payoff is one scorecard that tells you, for *your* workflows, exactly what
|
|
17
18
|
each model delivers — and what it costs you per run. Swap a model, re-run, see the
|
|
@@ -144,19 +145,51 @@ tests) is taken through the **real** production workflow end to end:
|
|
|
144
145
|
2. For `code-fix`, the **workspace is seeded** with the fixture repo at its base
|
|
145
146
|
commit plus a local bare `origin`, so `git push` works fully offline.
|
|
146
147
|
3. The **real workflow YAML** (`issue-triage`, `build`, …) is loaded from
|
|
147
|
-
`lastlight` and run with `sandbox:"none"
|
|
148
|
-
pointed at the fake, and approval gates disabled (so it never
|
|
149
|
-
|
|
148
|
+
`lastlight` and run with `sandbox:"none"` by default (in-process), the agent's
|
|
149
|
+
`github_*` tools pointed at the fake, and approval gates disabled (so it never
|
|
150
|
+
pauses). Pass `--sandbox gondolin` to isolate the agent's bash/file tools in a
|
|
151
|
+
QEMU micro-VM (see [Isolation](#isolation---sandbox) below).
|
|
152
|
+
4. The result is **graded deterministically** (with one scoped judge for
|
|
153
|
+
pr-review):
|
|
150
154
|
- **behavioral** — the recorded GitHub calls (labels, comments, PRs) vs the
|
|
151
155
|
instance's `expect_github` / `triage_gold`.
|
|
152
156
|
- **execution** (code-fix) — the held-out tests are applied and run; the case
|
|
153
157
|
is *resolved* only if every `FAIL_TO_PASS` passes and every `PASS_TO_PASS`
|
|
154
158
|
stays green (SWE-bench's criterion).
|
|
159
|
+
- **review** (pr-review) — the posted review is matched to a human-verified
|
|
160
|
+
gold set by an **LLM judge** → precision / recall / **F-beta** (F1 by default,
|
|
161
|
+
Martian's leaderboard metric; `EVAL_F_BETA` reweights). The one,
|
|
162
|
+
deliberately-scoped exception; triage/code-fix stay judge-free.
|
|
155
163
|
5. Token usage, cost, and latency are collected per run.
|
|
156
164
|
|
|
157
165
|
Run multiple models and you get a side-by-side **scorecard** (HTML + JSON)
|
|
158
166
|
ranking them on pass rate, cost, and latency.
|
|
159
167
|
|
|
168
|
+
### Isolation (`--sandbox`)
|
|
169
|
+
|
|
170
|
+
By default the agent runs **in-process** (`sandbox:"none"`) — fast and CI-friendly,
|
|
171
|
+
but with **no filesystem restriction**: the agent process can read any absolute
|
|
172
|
+
path, including this repo's held-out gold data (`datasets/<tier>/tests/`,
|
|
173
|
+
`instances.json`, `.eval-cache/`). A capable model that explores the disk could
|
|
174
|
+
find and spoil the answer key.
|
|
175
|
+
|
|
176
|
+
`--sandbox gondolin` (or `EVAL_SANDBOX=gondolin`) closes that gap: the agent's
|
|
177
|
+
bash/file tools execute inside a **QEMU micro-VM** that only sees its own
|
|
178
|
+
workspace, so host gold paths are invisible. The agent runtime and `github_*`
|
|
179
|
+
tools stay in-process, so the fake-GitHub mock still works unchanged — this is
|
|
180
|
+
why gondolin, and not `docker`, is the supported isolation backend (`docker`/`smol`
|
|
181
|
+
run the *whole* agent in the container/VM, where the in-process fake GitHub isn't
|
|
182
|
+
reachable).
|
|
183
|
+
|
|
184
|
+
Gondolin needs QEMU with hardware acceleration and runs **natively** (macOS via
|
|
185
|
+
Apple's Hypervisor.framework, Linux via KVM) — install it with `brew install qemu`
|
|
186
|
+
(macOS) or your distro's `qemu-system` package. It does **not** work inside a
|
|
187
|
+
container on macOS (no `/dev/kvm`, and the failure is a silent hang), so the
|
|
188
|
+
harness runs a fail-fast preflight and aborts with guidance rather than wedging.
|
|
189
|
+
Expect a one-time ~13s VM cold start plus per-tool-call overhead, so keep
|
|
190
|
+
`--sandbox gondolin` for trustworthy/anti-spoil runs and leave the default `none`
|
|
191
|
+
for quick iteration.
|
|
192
|
+
|
|
160
193
|
## Run it
|
|
161
194
|
|
|
162
195
|
```bash
|
|
@@ -180,6 +213,10 @@ lastlight-evals run triage --model glm,deepseek # a comma-list also works
|
|
|
180
213
|
# repeat each case N times; verdicts WORST-case, cost/tokens/latency MEAN
|
|
181
214
|
lastlight-evals run triage --runs 3
|
|
182
215
|
|
|
216
|
+
# isolate the agent's tools in a QEMU micro-VM so it can't read host gold data
|
|
217
|
+
# (anti-spoil; needs QEMU natively — see "Isolation" above). Default is none.
|
|
218
|
+
lastlight-evals run pr-review --sandbox gondolin
|
|
219
|
+
|
|
183
220
|
# run against an overlay repo's OWN workflows + datasets (see below)
|
|
184
221
|
lastlight-evals run --overlay ~/work/lastlight-instance
|
|
185
222
|
|
|
@@ -303,7 +340,10 @@ A **tier** is a directory containing `instances.json` (+ an optional `tier.json`
|
|
|
303
340
|
declaring its `defaultWorkflow`). Tiers are discovered from three roots, merged
|
|
304
341
|
by name with **overlay > user (`--datasets`) > built-in** precedence:
|
|
305
342
|
|
|
306
|
-
- **built-in** (shipped here): `triage` → `issue-triage`, `code-fix` → `build
|
|
343
|
+
- **built-in** (shipped here): `triage` → `issue-triage`, `code-fix` → `build`,
|
|
344
|
+
`pr-review` → `pr-review` (ships empty — populate with
|
|
345
|
+
`scripts/import-martian.ts`; see [PR-review tier](#pr-review-tier-code-review-bench)
|
|
346
|
+
below and `datasets/pr-review/README.md`).
|
|
307
347
|
- **user**: `--datasets <dir>` / `LASTLIGHT_EVALS_DATASETS`.
|
|
308
348
|
- **overlay**: `<overlay>/evals/datasets/*`.
|
|
309
349
|
|
|
@@ -366,6 +406,73 @@ A new tier just needs a directory with an `instances.json` and a `tier.json`
|
|
|
366
406
|
(`{ "name", "defaultWorkflow", "description" }`); per-instance `workflow` wins
|
|
367
407
|
when present.
|
|
368
408
|
|
|
409
|
+
## PR-review tier (Code Review Bench)
|
|
410
|
+
|
|
411
|
+
The **`pr-review`** tier measures review *quality* against
|
|
412
|
+
[Martian's Code Review Bench](https://github.com/withmartian/code-review-benchmark):
|
|
413
|
+
the review the real `pr-review` workflow posts is scored against a human-verified
|
|
414
|
+
**gold set** of the issues a reviewer should have caught. It's the **one** tier
|
|
415
|
+
graded by an LLM judge — matching free-text findings to semantic gold comments
|
|
416
|
+
can't be done deterministically — so triage and code-fix stay judge-free.
|
|
417
|
+
|
|
418
|
+
**Cases** come from Martian's *offline* set — 50 real merged PRs across Sentry,
|
|
419
|
+
Grafana, Cal.com, Discourse, and Keycloak, each carrying inlined `golden_comments`.
|
|
420
|
+
They ship **empty** (`datasets/pr-review/instances.json` is `[]`) because they're
|
|
421
|
+
large real-repo PRs — *generated*, not vendored:
|
|
422
|
+
|
|
423
|
+
```bash
|
|
424
|
+
npx tsx scripts/import-martian.ts # resolve all 50 via gh (pins base/head SHAs)
|
|
425
|
+
npx tsx scripts/import-martian.ts --limit 3 # a quick subset first
|
|
426
|
+
```
|
|
427
|
+
|
|
428
|
+
**Seeding** clones the real repo into the gitignored `./.eval-cache/` and checks
|
|
429
|
+
out the PR **head** (mirroring production's pre-clone contract), so the skill's
|
|
430
|
+
`git diff origin/<base>...HEAD` works fully offline — no fixture is vendored.
|
|
431
|
+
|
|
432
|
+
**Grading** (`gradeReview`, `src/grade.ts`) is a two-step LLM judge:
|
|
433
|
+
|
|
434
|
+
1. **Extract** the review's distinct, concrete findings (drop praise/summaries).
|
|
435
|
+
2. **Match** each finding to a gold comment ("same underlying issue?").
|
|
436
|
+
|
|
437
|
+
From the matches: **precision** = matched ÷ posted, **recall** = matched ÷ gold,
|
|
438
|
+
combined as **F-beta**. The headline is **F1** (β=1 — precision and recall weighted
|
|
439
|
+
equally, Martian's leaderboard metric). Pass **`--f-beta 0.5`** (or `EVAL_F_BETA=0.5`)
|
|
440
|
+
to weight precision 2× (F0.5), mirroring Martian's adjustable F-beta; the dashboard
|
|
441
|
+
relabels itself `F{β}` to match.
|
|
442
|
+
|
|
443
|
+
> **Gold-set caveat.** Martian's own methodology documents the gold set as
|
|
444
|
+
> *incomplete* — it caps at human performance, so a real issue the annotators
|
|
445
|
+
> missed is scored as a false positive. That understates precision, which is why
|
|
446
|
+
> the default is F1, not the precision-weighted F0.5. Treat the score as a
|
|
447
|
+
> **relative** signal and inspect each match with the dashboard's **judge** button.
|
|
448
|
+
|
|
449
|
+
**The judge model is independent** of the models under test — a strong default per
|
|
450
|
+
your provider key (`EVAL_JUDGE_MODEL` overrides). A judge failure marks the case
|
|
451
|
+
*errored* (ungraded), never a silent zero. Alongside the judge score, a cheap
|
|
452
|
+
deterministic `review_submitted` proxy checks a review was actually posted.
|
|
453
|
+
|
|
454
|
+
**Diff-blind by default.** The judge sees only the posted review (body + inline
|
|
455
|
+
comments) matched against the gold set — *not* the PR diff — mirroring Martian's
|
|
456
|
+
offline judge. This can penalize terse, location-anchored comments (`off-by-one
|
|
457
|
+
here` on a line the judge can't see). Pass **`--judge-with-diff`** to feed the PR
|
|
458
|
+
diff into the judge for higher-fidelity matching (the judge is instructed never to
|
|
459
|
+
invent findings from the diff); this trades away leaderboard parity, and the
|
|
460
|
+
dashboard marks such grades **`diff-aware`**.
|
|
461
|
+
|
|
462
|
+
**Run it** (heavy — clones real repos + calls the judge):
|
|
463
|
+
|
|
464
|
+
```bash
|
|
465
|
+
lastlight-evals run pr-review --model <model> # full tier
|
|
466
|
+
lastlight-evals run pr-review --model <model> --limit 3 # first 3 cases (controlled)
|
|
467
|
+
lastlight-evals run pr-review --model <model> --f-beta 0.5 # weight precision 2×
|
|
468
|
+
lastlight-evals run pr-review --model <model> --judge-with-diff # give the judge the diff
|
|
469
|
+
```
|
|
470
|
+
|
|
471
|
+
In the dashboard, each row's **judge** button opens the judge's working — the
|
|
472
|
+
findings it extracted, the gold set, the finding↔gold pairing (matched / false
|
|
473
|
+
positive / missed), and its raw replies — so the F1 score is inspectable, not a
|
|
474
|
+
black box.
|
|
475
|
+
|
|
369
476
|
## Models (`models.json`)
|
|
370
477
|
|
|
371
478
|
- `default` — the single model `run` uses.
|
|
@@ -378,5 +485,7 @@ when present.
|
|
|
378
485
|
|
|
379
486
|
- **`lastlight-evals extract <owner>/<repo>#<n>`** — generate eval cases from
|
|
380
487
|
GitHub historical issues/PRs (issue → fixture, merged PR → held-out tests).
|
|
381
|
-
- Docker-backed runs
|
|
488
|
+
- Docker-backed sandboxed runs (needs the fake GitHub reachable from inside the
|
|
489
|
+
container — `--sandbox gondolin` already gives native isolation today); real
|
|
490
|
+
SWE-bench Lite ingestion; per-fixture test runners.
|
|
382
491
|
- LLM-as-judge stays out by design — grading is deterministic.
|