lastlight-evals 0.4.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +117 -8
- package/dashboard/dist/assets/{index-DYwJHkLo.css → index-DbWEwzO3.css} +1 -1
- package/dashboard/dist/assets/index-GvQi4Peq.js +264 -0
- package/dashboard/dist/index.html +2 -2
- package/datasets/pr-review/README.md +58 -0
- package/datasets/pr-review/martian-leaderboard.json +5785 -0
- package/datasets/pr-review/tier.json +5 -0
- package/dist/add-case.js +188 -12
- package/dist/add-case.js.map +1 -1
- package/dist/discovery.js +10 -4
- package/dist/discovery.js.map +1 -1
- package/dist/fake-github.js +187 -6
- package/dist/fake-github.js.map +1 -1
- package/dist/grade.js +252 -0
- package/dist/grade.js.map +1 -1
- package/dist/judge.js +145 -0
- package/dist/judge.js.map +1 -0
- package/dist/mechanism.test.js +279 -2
- package/dist/mechanism.test.js.map +1 -1
- package/dist/report.js +123 -1
- package/dist/report.js.map +1 -1
- package/dist/run-instance.js +97 -13
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +160 -36
- package/dist/run.js.map +1 -1
- package/dist/sandbox-preflight.js +107 -0
- package/dist/sandbox-preflight.js.map +1 -0
- package/dist/seed.js +167 -0
- package/dist/seed.js.map +1 -1
- package/models.json +1 -0
- package/package.json +2 -2
- package/dashboard/dist/assets/index-D6qxQ8uN.js +0 -258
|
@@ -10,8 +10,8 @@
|
|
|
10
10
|
rel="stylesheet"
|
|
11
11
|
/>
|
|
12
12
|
<title>Last Light — Eval Dashboard</title>
|
|
13
|
-
<script type="module" crossorigin src="/assets/index-
|
|
14
|
-
<link rel="stylesheet" crossorigin href="/assets/index-
|
|
13
|
+
<script type="module" crossorigin src="/assets/index-GvQi4Peq.js"></script>
|
|
14
|
+
<link rel="stylesheet" crossorigin href="/assets/index-DbWEwzO3.css">
|
|
15
15
|
</head>
|
|
16
16
|
<body>
|
|
17
17
|
<div id="root"></div>
|
|
@@ -0,0 +1,58 @@
|
|
|
1
|
+
# pr-review tier
|
|
2
|
+
|
|
3
|
+
Measures **PR-review quality** the way Martian's
|
|
4
|
+
[Code Review Bench](https://github.com/withmartian/code-review-benchmark) does: the
|
|
5
|
+
review the `pr-review` workflow posts is matched, by an LLM judge, against a
|
|
6
|
+
human-verified **gold set** of real issues, scoring **precision / recall / F-beta**.
|
|
7
|
+
The headline is **F1** (β=1, precision and recall weighted equally — Martian's
|
|
8
|
+
leaderboard metric); set `EVAL_F_BETA=0.5` to weight precision 2× (F0.5), mirroring
|
|
9
|
+
Martian's adjustable F-beta. Cases come from their **offline** set
|
|
10
|
+
(`offline/results/benchmark_data.json`).
|
|
11
|
+
|
|
12
|
+
`instances.json` is **gitignored** — the cases are *generated* from Martian's
|
|
13
|
+
benchmark, not vendored, so they don't live in this repo. Populate it **once**
|
|
14
|
+
locally and it persists across runs (no git noise, no re-import each time). It
|
|
15
|
+
holds 50 PRs across Sentry / Grafana / Cal.com / Discourse / Keycloak:
|
|
16
|
+
|
|
17
|
+
```bash
|
|
18
|
+
# needs `gh` (authenticated) + network; pins base/head SHAs into instances.json
|
|
19
|
+
npx tsx scripts/import-martian.ts # full 50
|
|
20
|
+
npx tsx scripts/import-martian.ts --limit 3 # a quick subset first
|
|
21
|
+
npx tsx scripts/import-martian.ts --dry-run # preview without writing
|
|
22
|
+
```
|
|
23
|
+
|
|
24
|
+
Then run the tier (heavy — clones the real repos, calls a judge model):
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
# grade one model; the judge defaults to a strong model per your provider keys
|
|
28
|
+
# (override with EVAL_JUDGE_MODEL). See src/judge.ts.
|
|
29
|
+
npx tsx src/run.ts run pr-review --model <model> # full tier
|
|
30
|
+
npx tsx src/run.ts run pr-review --model <model> --limit 3 # first 3 cases (controlled/cheap)
|
|
31
|
+
```
|
|
32
|
+
|
|
33
|
+
`--limit N` caps the tier to its first N instances (in file order) — the
|
|
34
|
+
lightest way to smoke-test the plumbing before cloning + grading all 50. Combine
|
|
35
|
+
with `--instance <id>` to pin exact cases.
|
|
36
|
+
|
|
37
|
+
Each case's shape (`src/schema.ts`):
|
|
38
|
+
|
|
39
|
+
- `pr` — the PR fixture served by the fake GitHub + checked out at its **head**
|
|
40
|
+
(base + head refs/commits, so `git diff origin/<base>...HEAD` works offline).
|
|
41
|
+
- `review_gold` — the gold comments (`severity` + `description`; file/line are
|
|
42
|
+
absent in the Martian set, so the judge matches on substance).
|
|
43
|
+
- `expect_github.review_submitted` — a cheap deterministic proxy (a review was
|
|
44
|
+
posted) alongside the judge grade.
|
|
45
|
+
|
|
46
|
+
> Comparability caveat: our F1 won't equal the public leaderboard (different judge
|
|
47
|
+
> model + harness). Treat it as a **relative** optimisation signal, and inspect the
|
|
48
|
+
> per-case match with the dashboard's **judge** button. Martian's gold set is known
|
|
49
|
+
> (by their own methodology) to be **incomplete** — it caps at human performance, so
|
|
50
|
+
> a real issue the annotators missed scores as a false positive. That understates
|
|
51
|
+
> precision, which is why the default is F1 rather than the precision-weighted F0.5.
|
|
52
|
+
|
|
53
|
+
---
|
|
54
|
+
|
|
55
|
+
**Attribution.** Cases derive from Martian's
|
|
56
|
+
[Code Review Bench](https://github.com/withmartian/code-review-benchmark)
|
|
57
|
+
(© 2025 Martian, MIT). The importer pins the PRs' base/head SHAs and inlines the
|
|
58
|
+
gold comments locally; nothing from Martian's dataset is committed to this repo.
|