lastlight-evals 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -10,8 +10,8 @@
10
10
  rel="stylesheet"
11
11
  />
12
12
  <title>Last Light — Eval Dashboard</title>
13
- <script type="module" crossorigin src="/assets/index-D6qxQ8uN.js"></script>
14
- <link rel="stylesheet" crossorigin href="/assets/index-DYwJHkLo.css">
13
+ <script type="module" crossorigin src="/assets/index-GvQi4Peq.js"></script>
14
+ <link rel="stylesheet" crossorigin href="/assets/index-DbWEwzO3.css">
15
15
  </head>
16
16
  <body>
17
17
  <div id="root"></div>
@@ -0,0 +1,58 @@
1
+ # pr-review tier
2
+
3
+ Measures **PR-review quality** the way Martian's
4
+ [Code Review Bench](https://github.com/withmartian/code-review-benchmark) does: the
5
+ review the `pr-review` workflow posts is matched, by an LLM judge, against a
6
+ human-verified **gold set** of real issues, scoring **precision / recall / F-beta**.
7
+ The headline is **F1** (β=1, precision and recall weighted equally — Martian's
8
+ leaderboard metric); set `EVAL_F_BETA=0.5` to weight precision 2× (F0.5), mirroring
9
+ Martian's adjustable F-beta. Cases come from their **offline** set
10
+ (`offline/results/benchmark_data.json`).
11
+
12
+ `instances.json` is **gitignored** — the cases are *generated* from Martian's
13
+ benchmark, not vendored, so they don't live in this repo. Populate it **once**
14
+ locally and it persists across runs (no git noise, no re-import each time). It
15
+ holds 50 PRs across Sentry / Grafana / Cal.com / Discourse / Keycloak:
16
+
17
+ ```bash
18
+ # needs `gh` (authenticated) + network; pins base/head SHAs into instances.json
19
+ npx tsx scripts/import-martian.ts # full 50
20
+ npx tsx scripts/import-martian.ts --limit 3 # a quick subset first
21
+ npx tsx scripts/import-martian.ts --dry-run # preview without writing
22
+ ```
23
+
24
+ Then run the tier (heavy — clones the real repos, calls a judge model):
25
+
26
+ ```bash
27
+ # grade one model; the judge defaults to a strong model per your provider keys
28
+ # (override with EVAL_JUDGE_MODEL). See src/judge.ts.
29
+ npx tsx src/run.ts run pr-review --model <model> # full tier
30
+ npx tsx src/run.ts run pr-review --model <model> --limit 3 # first 3 cases (controlled/cheap)
31
+ ```
32
+
33
+ `--limit N` caps the tier to its first N instances (in file order) — the
34
+ lightest way to smoke-test the plumbing before cloning + grading all 50. Combine
35
+ with `--instance <id>` to pin exact cases.
36
+
37
+ Each case's shape (`src/schema.ts`):
38
+
39
+ - `pr` — the PR fixture served by the fake GitHub + checked out at its **head**
40
+ (base + head refs/commits, so `git diff origin/<base>...HEAD` works offline).
41
+ - `review_gold` — the gold comments (`severity` + `description`; file/line are
42
+ absent in the Martian set, so the judge matches on substance).
43
+ - `expect_github.review_submitted` — a cheap deterministic proxy (a review was
44
+ posted) alongside the judge grade.
45
+
46
+ > Comparability caveat: our F1 won't equal the public leaderboard (different judge
47
+ > model + harness). Treat it as a **relative** optimisation signal, and inspect the
48
+ > per-case match with the dashboard's **judge** button. Martian's gold set is known
49
+ > (by their own methodology) to be **incomplete** — it caps at human performance, so
50
+ > a real issue the annotators missed scores as a false positive. That understates
51
+ > precision, which is why the default is F1 rather than the precision-weighted F0.5.
52
+
53
+ ---
54
+
55
+ **Attribution.** Cases derive from Martian's
56
+ [Code Review Bench](https://github.com/withmartian/code-review-benchmark)
57
+ (© 2025 Martian, MIT). The importer pins the PRs' base/head SHAs and inlines the
58
+ gold comments locally; nothing from Martian's dataset is committed to this repo.