lastlight-evals 0.10.0 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dashboard/dist/assets/index-B4TAaIUR.css +1 -0
- package/dashboard/dist/assets/index-LFdFc6Z9.js +109 -0
- package/dashboard/dist/index.html +2 -2
- package/datasets/aacr-bench/README.md +146 -0
- package/datasets/pr-review/README.md +41 -0
- package/datasets/pr-review/anchors.json +3709 -0
- package/dist/add-case.js.map +1 -1
- package/dist/arm.js +14 -5
- package/dist/arm.js.map +1 -1
- package/dist/bootstrap.js +14 -0
- package/dist/bootstrap.js.map +1 -1
- package/dist/clean.js.map +1 -1
- package/dist/config.js +59 -20
- package/dist/config.js.map +1 -1
- package/dist/fake-github.js +72 -0
- package/dist/fake-github.js.map +1 -1
- package/dist/grade.js +308 -64
- package/dist/grade.js.map +1 -1
- package/dist/init.js.map +1 -1
- package/dist/judge.js +19 -0
- package/dist/judge.js.map +1 -1
- package/dist/mechanism.test.js +303 -13
- package/dist/mechanism.test.js.map +1 -1
- package/dist/metrics.js +88 -9
- package/dist/metrics.js.map +1 -1
- package/dist/paths.js +135 -1
- package/dist/paths.js.map +1 -1
- package/dist/phase-models.js +56 -0
- package/dist/phase-models.js.map +1 -0
- package/dist/pool.js +32 -0
- package/dist/pool.js.map +1 -0
- package/dist/pool.test.js +180 -0
- package/dist/pool.test.js.map +1 -0
- package/dist/pr-context.js +77 -6
- package/dist/pr-context.js.map +1 -1
- package/dist/repeats.test.js +317 -0
- package/dist/repeats.test.js.map +1 -0
- package/dist/report.js +99 -5
- package/dist/report.js.map +1 -1
- package/dist/review-metrics.js +571 -0
- package/dist/review-metrics.js.map +1 -0
- package/dist/review-metrics.test.js +634 -0
- package/dist/review-metrics.test.js.map +1 -0
- package/dist/review-pipeline-stats.js +455 -0
- package/dist/review-pipeline-stats.js.map +1 -0
- package/dist/review-pipeline-stats.test.js +413 -0
- package/dist/review-pipeline-stats.test.js.map +1 -0
- package/dist/run-instance.js +176 -45
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +637 -122
- package/dist/run.js.map +1 -1
- package/dist/schema.js +22 -1
- package/dist/schema.js.map +1 -1
- package/package.json +9 -8
- package/dashboard/dist/assets/index-B1jNgf2P.js +0 -313
- package/dashboard/dist/assets/index-Sic0jI7y.css +0 -1
|
@@ -28,8 +28,8 @@
|
|
|
28
28
|
rel="stylesheet"
|
|
29
29
|
/>
|
|
30
30
|
<title>Last Light — Eval Dashboard</title>
|
|
31
|
-
<script type="module" crossorigin src="/assets/index-
|
|
32
|
-
<link rel="stylesheet" crossorigin href="/assets/index-
|
|
31
|
+
<script type="module" crossorigin src="/assets/index-LFdFc6Z9.js"></script>
|
|
32
|
+
<link rel="stylesheet" crossorigin href="/assets/index-B4TAaIUR.css">
|
|
33
33
|
</head>
|
|
34
34
|
<body>
|
|
35
35
|
<div id="root"></div>
|
|
@@ -0,0 +1,146 @@
|
|
|
1
|
+
# AACR-Bench — the external adjudication instrument
|
|
2
|
+
|
|
3
|
+
This directory holds **no data**. AACR-Bench is downloaded on first use into
|
|
4
|
+
`.eval-cache/aacr-bench/dataset.json` (gitignored) and pinned by sha256 in every
|
|
5
|
+
report. This file is the record of what the dataset is, what it is good for, and
|
|
6
|
+
the three things about it that will mislead you if nobody writes them down.
|
|
7
|
+
|
|
8
|
+
- **Source:** [`Alibaba-Aone/aacr-bench`](https://huggingface.co/datasets/Alibaba-Aone/aacr-bench) on Hugging Face
|
|
9
|
+
- **Licence:** Apache-2.0
|
|
10
|
+
- **Published alongside:** [`alibaba/open-code-review`](https://github.com/alibaba/open-code-review) (Apache-2.0, Go)
|
|
11
|
+
- **Single file:** `dataset.json`, ~2.1 MB, a JSON array of **2,145** objects
|
|
12
|
+
- **Fetched from:** `https://huggingface.co/datasets/Alibaba-Aone/aacr-bench/resolve/main/dataset.json`
|
|
13
|
+
- **sha256 (first fetch, 2026-08-22):** `0804505f0a474765ce2840c832cfeaa6c4f0250dd6ccb169fe73c6758b245a86`
|
|
14
|
+
|
|
15
|
+
## Why it is not committed
|
|
16
|
+
|
|
17
|
+
2.1 MB of somebody else's corpus does not belong in this tree, and the eval
|
|
18
|
+
package's `files` list would ship it to npm. The reproducibility story is the
|
|
19
|
+
**sha256 stamped in `report.json`** instead, exactly as `datasets/pr-review/instances.json`
|
|
20
|
+
is generated rather than vendored. `.eval-cache/` is already in `apps/evals/.gitignore`.
|
|
21
|
+
|
|
22
|
+
Override the cache location with `LASTLIGHT_EVALS_CACHE`, or point at a local copy
|
|
23
|
+
with `--dataset <path>`.
|
|
24
|
+
|
|
25
|
+
## The consumer
|
|
26
|
+
|
|
27
|
+
`apps/evals/scripts/aacr-adjudicate.ts` — see its header doc-block, which is the
|
|
28
|
+
long-form version of everything below.
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
cd apps/evals
|
|
32
|
+
|
|
33
|
+
# The two deterministic floors. No model calls, no keys, no spend.
|
|
34
|
+
npx tsx scripts/aacr-adjudicate.ts --arm keep-all --all
|
|
35
|
+
npx tsx scripts/aacr-adjudicate.ts --arm drop-all --all
|
|
36
|
+
|
|
37
|
+
# The whole pipeline including prompt construction, with a stub decision.
|
|
38
|
+
npx tsx scripts/aacr-adjudicate.ts --arm llm --dry-run --limit 50 --print-prompt
|
|
39
|
+
|
|
40
|
+
# A real run. Prints its cost first and REFUSES without --yes.
|
|
41
|
+
npx tsx scripts/aacr-adjudicate.ts --arm llm --limit 50 --yes --out
|
|
42
|
+
```
|
|
43
|
+
|
|
44
|
+
## Row shape
|
|
45
|
+
|
|
46
|
+
Every key is present on every one of the 2,145 rows (verified).
|
|
47
|
+
|
|
48
|
+
| Field | Type | Notes |
|
|
49
|
+
|---|---|---|
|
|
50
|
+
| `project_main_language` | str | C++ 508, TypeScript 422, Java 272, Go 245, C 206, Python 151, JavaScript 141, Rust 92, PHP 56, C# 52 |
|
|
51
|
+
| `pr_url` | str | `https://github.com/OWNER/REPO/pull/N` — 200 distinct PRs across 50 repos |
|
|
52
|
+
| `pr_source_commit` | str | **This is `base.sha`, not the head** (verified 12/12) |
|
|
53
|
+
| `pr_target_commit` | str | A historical head sha |
|
|
54
|
+
| `pr_change_line_count` | int | |
|
|
55
|
+
| `pr_category` | str | e.g. `Code Refactoring / Architectural Improvement` |
|
|
56
|
+
| `is_ai_comment` | bool | **1,597 true / 548 false** |
|
|
57
|
+
| `note` | str | **The review comment text — the whole input.** p50 283 chars, p95 867, max 2,223 |
|
|
58
|
+
| `path` | str | File the comment is on |
|
|
59
|
+
| `side` | str | `right` \| `left` |
|
|
60
|
+
| `source_model` | str | Which model wrote it; **empty string on the 548 human rows** |
|
|
61
|
+
| `from_line`, `to_line` | int | |
|
|
62
|
+
| `category` | str | Code Defect 1022, Maintainability and Readability 905, Performance 144, Security Vulnerability 74 |
|
|
63
|
+
| `context` | str | Diff Level 1017, File Level 744, Repo Level 384 |
|
|
64
|
+
| `label` | int | **1 = expert-verified CORRECT (1,505), 0 = INCORRECT (640)** |
|
|
65
|
+
|
|
66
|
+
Cross-tabs worth having:
|
|
67
|
+
|
|
68
|
+
| | label=1 | label=0 | base rate |
|
|
69
|
+
|---|---|---|---|
|
|
70
|
+
| AI-authored | 1,114 | 483 | 69.8% |
|
|
71
|
+
| human-authored | 391 | 157 | 71.4% |
|
|
72
|
+
|
|
73
|
+
`source_model`: GPT-5.2 575, (human) 548, Claude-Code/Claude-4.5-Sonnet 279,
|
|
74
|
+
Qwen-Coder-480B 254, GLM-4.7 229, Deepseek-V3.2 136, Gemini-3-Pro 124.
|
|
75
|
+
|
|
76
|
+
## No clone. No checkout. That is the point
|
|
77
|
+
|
|
78
|
+
Adjudication takes the **comment text plus its metadata** and nothing else. No
|
|
79
|
+
repo is cloned, no sha is resolved, `gh` is never invoked, and the two commit
|
|
80
|
+
fields are recorded for provenance only. That is why this instrument could be
|
|
81
|
+
built before any of the review-reconstruction work in
|
|
82
|
+
[`docs/plans/deterministic-pr-levers.md`](../../../../docs/plans/deterministic-pr-levers.md) —
|
|
83
|
+
it has no dependency on WP1–WP5.
|
|
84
|
+
|
|
85
|
+
## The three things that will mislead you
|
|
86
|
+
|
|
87
|
+
### 1. This is not review recall
|
|
88
|
+
|
|
89
|
+
The script measures whether an arm can tell a valid comment from an invalid one
|
|
90
|
+
**when handed the comment**. It says nothing about whether our reviewer would
|
|
91
|
+
have *generated* that comment — which is the actual bottleneck (1 of 25 gold
|
|
92
|
+
findings on `skillspro`). An arm can score perfectly here and move review recall
|
|
93
|
+
by zero.
|
|
94
|
+
|
|
95
|
+
### 2. It is not the AACR-Bench leaderboard's metric
|
|
96
|
+
|
|
97
|
+
The published leaderboard scores review **generation** against the 1,505 label=1
|
|
98
|
+
rows as a gold set to be found — Open Code Review v1.3.1 at 20.0% recall
|
|
99
|
+
(301/1505), Claude Code v2.1.169 at 28.9% (435/1505). Different task, different
|
|
100
|
+
denominator. Per the `01b` house rule and
|
|
101
|
+
[WP9 AC9](../../../../docs/plans/deterministic-pr-levers.md#external-validation-wp9),
|
|
102
|
+
our number and theirs are never averaged, pooled, or put in the same column.
|
|
103
|
+
|
|
104
|
+
The 640 label=0 rows are, as far as the leaderboard is concerned, not part of the
|
|
105
|
+
gold set at all. They exist for exactly the classification task this script runs,
|
|
106
|
+
and WP9's write-up of the dataset ("1,505 annotated ground-truth issues") does not
|
|
107
|
+
mention them.
|
|
108
|
+
|
|
109
|
+
### 3. 74% of it is machine-authored
|
|
110
|
+
|
|
111
|
+
1,597 comments were written by GPT-5.2, Claude-4.5-Sonnet, Qwen-Coder-480B,
|
|
112
|
+
GLM-4.7, Deepseek-V3.2 or Gemini-3-Pro and *then* expert-verified. Only 548 are
|
|
113
|
+
human-authored. So the corpus over-represents the findings AI reviewers already
|
|
114
|
+
produce, and an adjudicator tuned on it is tuned to police machine output. The
|
|
115
|
+
script always reports the two halves separately and never pools them.
|
|
116
|
+
|
|
117
|
+
Contamination applies as it does to the Martian set: these are public
|
|
118
|
+
repositories, and the PRs are historical. Treat it as a *contaminated, large,
|
|
119
|
+
public* set complementary to the *clean, tiny, private* `skillspro` one — never
|
|
120
|
+
pooled with it either.
|
|
121
|
+
|
|
122
|
+
## What a result here means for WP6
|
|
123
|
+
|
|
124
|
+
[WP6](../../../../docs/plans/deterministic-pr-levers.md#adjudication-and-the-attention-boundary-wp6)'s
|
|
125
|
+
adjudicator **may re-rank, re-tier and demote, but may not delete a finding
|
|
126
|
+
without a probe transcript refuting it.** Every time a filter has been measured
|
|
127
|
+
against a conservative reviewer it has cost recall — our own candidate v2
|
|
128
|
+
(micro-recall 1/25 → 2/25, F1 halved, reverted), BitsAI-CR's ReviewFilter
|
|
129
|
+
(precision 54.5 → 67.1, recall 45.5 → **39.8**), and Open Code Review's own
|
|
130
|
+
deterministic layer (discards ~5,100 findings and loses 134 real defects doing
|
|
131
|
+
it).
|
|
132
|
+
|
|
133
|
+
So the headline metric is **retention** (kept ÷ label=1), not interception. The
|
|
134
|
+
deterministic arms pin what those numbers are *without* the intervention, which
|
|
135
|
+
is the only thing that makes the intervention's numbers a guard rather than a
|
|
136
|
+
decoration:
|
|
137
|
+
|
|
138
|
+
| Arm | retention | interception | precision | F1 |
|
|
139
|
+
|---|---|---|---|---|
|
|
140
|
+
| `keep-all` — **what production does today** | 100.0% (1505/1505) | 0.0% (0/640) | 70.2% (1505/2145) | 0.825 |
|
|
141
|
+
| `drop-all` | 0.0% (0/1505) | 100.0% (640/640) | n/a (0/0) | n/a |
|
|
142
|
+
|
|
143
|
+
`keep-all`'s precision is the corpus base rate and its F1 of **0.825** is the bar.
|
|
144
|
+
An arm scoring below it has traded recall for nothing — which is the exact shape
|
|
145
|
+
of the reverted candidate v2. A WP6 adjudicator drops in as a fourth arm by
|
|
146
|
+
implementing the `Arm` interface in the script and registering in `ARMS`.
|
|
@@ -19,8 +19,15 @@ holds 50 PRs across Sentry / Grafana / Cal.com / Discourse / Keycloak:
|
|
|
19
19
|
npx tsx scripts/import-martian.ts # full 50
|
|
20
20
|
npx tsx scripts/import-martian.ts --limit 3 # a quick subset first
|
|
21
21
|
npx tsx scripts/import-martian.ts --dry-run # preview without writing
|
|
22
|
+
npx tsx scripts/import-martian.ts --tier pr-review-heldout # import into a sibling tier dir
|
|
22
23
|
```
|
|
23
24
|
|
|
25
|
+
`--tier <name>` (default `pr-review`) picks the tier directory the import writes
|
|
26
|
+
(`datasets/<name>/`) and the `name` stamped into a freshly-created `tier.json`,
|
|
27
|
+
so a held-out split can live beside this tier under its own name. The tier's
|
|
28
|
+
`defaultWorkflow` stays `pr-review` either way — every imported case runs the
|
|
29
|
+
pr-review workflow.
|
|
30
|
+
|
|
24
31
|
Then run the tier (heavy — clones the real repos, calls a judge model):
|
|
25
32
|
|
|
26
33
|
```bash
|
|
@@ -50,6 +57,40 @@ Each case's shape (`src/schema.ts`):
|
|
|
50
57
|
> a real issue the annotators missed scores as a false positive. That understates
|
|
51
58
|
> precision, which is why the default is F1 rather than the precision-weighted F0.5.
|
|
52
59
|
|
|
60
|
+
## `anchors.json` — the deterministic anchor labels
|
|
61
|
+
|
|
62
|
+
`anchors.json` **is** committed (unlike `instances.json`) and is the frozen input to the
|
|
63
|
+
code-facts *evidence-coverage* metric. Martian's gold set carries only
|
|
64
|
+
`{severity, description}` — no file, no line — so there is no deterministic join from a gold
|
|
65
|
+
finding to a place in the diff. `scripts/facts-anchors.ts` builds one: it pulls code-shaped
|
|
66
|
+
tokens out of the gold prose (tokenizer `v1`, documented rule-by-rule in the script) and marks a
|
|
67
|
+
finding **anchored** when one of them matches, on a word boundary, an added-or-removed line of the
|
|
68
|
+
three-dot (`base...head`, merge-base) diff. No model is involved — an LLM in the denominator would
|
|
69
|
+
make every coverage number downstream unfalsifiable.
|
|
70
|
+
|
|
71
|
+
```bash
|
|
72
|
+
npx tsx scripts/facts-anchors.ts # regenerate + print the report
|
|
73
|
+
npx tsx scripts/facts-anchors.ts --dry-run --audit # the hand-audit sheet
|
|
74
|
+
npx tsx scripts/facts-anchors.ts --dry-run --unanchored # what could not be anchored
|
|
75
|
+
```
|
|
76
|
+
|
|
77
|
+
Three things about it are load-bearing:
|
|
78
|
+
|
|
79
|
+
- **Freeze the labels, not the tokenizer.** The metric's denominator IS the tokenizer's output, so
|
|
80
|
+
the artifact is committed and stamps `tokenizer: "v1"`. A better tokenizer ships as `v2` in a new
|
|
81
|
+
file; it never rewrites what past numbers meant.
|
|
82
|
+
- **It carries its own error bar.** The `audit` block records a seeded random sample of 20 anchored
|
|
83
|
+
findings, each read by hand, with per-finding verdicts. Quote the anchor rate with that rate
|
|
84
|
+
attached, and don't tune the tokenizer to improve it.
|
|
85
|
+
- **No gold text.** Only derived labels (anchors, `path:line`, severity, Martian's `bug_type` /
|
|
86
|
+
`requires_context` / `language`). Join back to your local `instances.json` on
|
|
87
|
+
`instanceId` + `goldIndex` to read a description.
|
|
88
|
+
|
|
89
|
+
The headline: **99/137 anchored (72.3%)**. That is a property of the *gold text*, not of
|
|
90
|
+
code-facts — it is the ceiling on what any identifier-level evidence layer could ever be scored
|
|
91
|
+
against. The 38 unanchored are mostly prose with nothing code-shaped in it (20 of them name no
|
|
92
|
+
identifier at all), plus a stylesheet/i18n cluster where the finding is a value, not a name.
|
|
93
|
+
|
|
53
94
|
---
|
|
54
95
|
|
|
55
96
|
**Attribution.** Cases derive from Martian's
|