lastlight-evals 0.10.0 → 0.11.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (56) hide show
  1. package/dashboard/dist/assets/index-B4TAaIUR.css +1 -0
  2. package/dashboard/dist/assets/index-LFdFc6Z9.js +109 -0
  3. package/dashboard/dist/index.html +2 -2
  4. package/datasets/aacr-bench/README.md +146 -0
  5. package/datasets/pr-review/README.md +41 -0
  6. package/datasets/pr-review/anchors.json +3709 -0
  7. package/dist/add-case.js.map +1 -1
  8. package/dist/arm.js +14 -5
  9. package/dist/arm.js.map +1 -1
  10. package/dist/bootstrap.js +14 -0
  11. package/dist/bootstrap.js.map +1 -1
  12. package/dist/clean.js.map +1 -1
  13. package/dist/config.js +59 -20
  14. package/dist/config.js.map +1 -1
  15. package/dist/fake-github.js +72 -0
  16. package/dist/fake-github.js.map +1 -1
  17. package/dist/grade.js +308 -64
  18. package/dist/grade.js.map +1 -1
  19. package/dist/init.js.map +1 -1
  20. package/dist/judge.js +19 -0
  21. package/dist/judge.js.map +1 -1
  22. package/dist/mechanism.test.js +303 -13
  23. package/dist/mechanism.test.js.map +1 -1
  24. package/dist/metrics.js +88 -9
  25. package/dist/metrics.js.map +1 -1
  26. package/dist/paths.js +135 -1
  27. package/dist/paths.js.map +1 -1
  28. package/dist/phase-models.js +56 -0
  29. package/dist/phase-models.js.map +1 -0
  30. package/dist/pool.js +32 -0
  31. package/dist/pool.js.map +1 -0
  32. package/dist/pool.test.js +180 -0
  33. package/dist/pool.test.js.map +1 -0
  34. package/dist/pr-context.js +71 -6
  35. package/dist/pr-context.js.map +1 -1
  36. package/dist/repeats.test.js +317 -0
  37. package/dist/repeats.test.js.map +1 -0
  38. package/dist/report.js +99 -5
  39. package/dist/report.js.map +1 -1
  40. package/dist/review-metrics.js +571 -0
  41. package/dist/review-metrics.js.map +1 -0
  42. package/dist/review-metrics.test.js +634 -0
  43. package/dist/review-metrics.test.js.map +1 -0
  44. package/dist/review-pipeline-stats.js +414 -0
  45. package/dist/review-pipeline-stats.js.map +1 -0
  46. package/dist/review-pipeline-stats.test.js +413 -0
  47. package/dist/review-pipeline-stats.test.js.map +1 -0
  48. package/dist/run-instance.js +176 -45
  49. package/dist/run-instance.js.map +1 -1
  50. package/dist/run.js +637 -122
  51. package/dist/run.js.map +1 -1
  52. package/dist/schema.js +22 -1
  53. package/dist/schema.js.map +1 -1
  54. package/package.json +9 -8
  55. package/dashboard/dist/assets/index-B1jNgf2P.js +0 -313
  56. package/dashboard/dist/assets/index-Sic0jI7y.css +0 -1
@@ -28,8 +28,8 @@
28
28
  rel="stylesheet"
29
29
  />
30
30
  <title>Last Light — Eval Dashboard</title>
31
- <script type="module" crossorigin src="/assets/index-B1jNgf2P.js"></script>
32
- <link rel="stylesheet" crossorigin href="/assets/index-Sic0jI7y.css">
31
+ <script type="module" crossorigin src="/assets/index-LFdFc6Z9.js"></script>
32
+ <link rel="stylesheet" crossorigin href="/assets/index-B4TAaIUR.css">
33
33
  </head>
34
34
  <body>
35
35
  <div id="root"></div>
@@ -0,0 +1,146 @@
1
+ # AACR-Bench — the external adjudication instrument
2
+
3
+ This directory holds **no data**. AACR-Bench is downloaded on first use into
4
+ `.eval-cache/aacr-bench/dataset.json` (gitignored) and pinned by sha256 in every
5
+ report. This file is the record of what the dataset is, what it is good for, and
6
+ the three things about it that will mislead you if nobody writes them down.
7
+
8
+ - **Source:** [`Alibaba-Aone/aacr-bench`](https://huggingface.co/datasets/Alibaba-Aone/aacr-bench) on Hugging Face
9
+ - **Licence:** Apache-2.0
10
+ - **Published alongside:** [`alibaba/open-code-review`](https://github.com/alibaba/open-code-review) (Apache-2.0, Go)
11
+ - **Single file:** `dataset.json`, ~2.1 MB, a JSON array of **2,145** objects
12
+ - **Fetched from:** `https://huggingface.co/datasets/Alibaba-Aone/aacr-bench/resolve/main/dataset.json`
13
+ - **sha256 (first fetch, 2026-08-22):** `0804505f0a474765ce2840c832cfeaa6c4f0250dd6ccb169fe73c6758b245a86`
14
+
15
+ ## Why it is not committed
16
+
17
+ 2.1 MB of somebody else's corpus does not belong in this tree, and the eval
18
+ package's `files` list would ship it to npm. The reproducibility story is the
19
+ **sha256 stamped in `report.json`** instead, exactly as `datasets/pr-review/instances.json`
20
+ is generated rather than vendored. `.eval-cache/` is already in `apps/evals/.gitignore`.
21
+
22
+ Override the cache location with `LASTLIGHT_EVALS_CACHE`, or point at a local copy
23
+ with `--dataset <path>`.
24
+
25
+ ## The consumer
26
+
27
+ `apps/evals/scripts/aacr-adjudicate.ts` — see its header doc-block, which is the
28
+ long-form version of everything below.
29
+
30
+ ```bash
31
+ cd apps/evals
32
+
33
+ # The two deterministic floors. No model calls, no keys, no spend.
34
+ npx tsx scripts/aacr-adjudicate.ts --arm keep-all --all
35
+ npx tsx scripts/aacr-adjudicate.ts --arm drop-all --all
36
+
37
+ # The whole pipeline including prompt construction, with a stub decision.
38
+ npx tsx scripts/aacr-adjudicate.ts --arm llm --dry-run --limit 50 --print-prompt
39
+
40
+ # A real run. Prints its cost first and REFUSES without --yes.
41
+ npx tsx scripts/aacr-adjudicate.ts --arm llm --limit 50 --yes --out
42
+ ```
43
+
44
+ ## Row shape
45
+
46
+ Every key is present on every one of the 2,145 rows (verified).
47
+
48
+ | Field | Type | Notes |
49
+ |---|---|---|
50
+ | `project_main_language` | str | C++ 508, TypeScript 422, Java 272, Go 245, C 206, Python 151, JavaScript 141, Rust 92, PHP 56, C# 52 |
51
+ | `pr_url` | str | `https://github.com/OWNER/REPO/pull/N` — 200 distinct PRs across 50 repos |
52
+ | `pr_source_commit` | str | **This is `base.sha`, not the head** (verified 12/12) |
53
+ | `pr_target_commit` | str | A historical head sha |
54
+ | `pr_change_line_count` | int | |
55
+ | `pr_category` | str | e.g. `Code Refactoring / Architectural Improvement` |
56
+ | `is_ai_comment` | bool | **1,597 true / 548 false** |
57
+ | `note` | str | **The review comment text — the whole input.** p50 283 chars, p95 867, max 2,223 |
58
+ | `path` | str | File the comment is on |
59
+ | `side` | str | `right` \| `left` |
60
+ | `source_model` | str | Which model wrote it; **empty string on the 548 human rows** |
61
+ | `from_line`, `to_line` | int | |
62
+ | `category` | str | Code Defect 1022, Maintainability and Readability 905, Performance 144, Security Vulnerability 74 |
63
+ | `context` | str | Diff Level 1017, File Level 744, Repo Level 384 |
64
+ | `label` | int | **1 = expert-verified CORRECT (1,505), 0 = INCORRECT (640)** |
65
+
66
+ Cross-tabs worth having:
67
+
68
+ | | label=1 | label=0 | base rate |
69
+ |---|---|---|---|
70
+ | AI-authored | 1,114 | 483 | 69.8% |
71
+ | human-authored | 391 | 157 | 71.4% |
72
+
73
+ `source_model`: GPT-5.2 575, (human) 548, Claude-Code/Claude-4.5-Sonnet 279,
74
+ Qwen-Coder-480B 254, GLM-4.7 229, Deepseek-V3.2 136, Gemini-3-Pro 124.
75
+
76
+ ## No clone. No checkout. That is the point
77
+
78
+ Adjudication takes the **comment text plus its metadata** and nothing else. No
79
+ repo is cloned, no sha is resolved, `gh` is never invoked, and the two commit
80
+ fields are recorded for provenance only. That is why this instrument could be
81
+ built before any of the review-reconstruction work in
82
+ [`docs/plans/deterministic-pr-levers.md`](../../../../docs/plans/deterministic-pr-levers.md) —
83
+ it has no dependency on WP1–WP5.
84
+
85
+ ## The three things that will mislead you
86
+
87
+ ### 1. This is not review recall
88
+
89
+ The script measures whether an arm can tell a valid comment from an invalid one
90
+ **when handed the comment**. It says nothing about whether our reviewer would
91
+ have *generated* that comment — which is the actual bottleneck (1 of 25 gold
92
+ findings on `skillspro`). An arm can score perfectly here and move review recall
93
+ by zero.
94
+
95
+ ### 2. It is not the AACR-Bench leaderboard's metric
96
+
97
+ The published leaderboard scores review **generation** against the 1,505 label=1
98
+ rows as a gold set to be found — Open Code Review v1.3.1 at 20.0% recall
99
+ (301/1505), Claude Code v2.1.169 at 28.9% (435/1505). Different task, different
100
+ denominator. Per the `01b` house rule and
101
+ [WP9 AC9](../../../../docs/plans/deterministic-pr-levers.md#external-validation-wp9),
102
+ our number and theirs are never averaged, pooled, or put in the same column.
103
+
104
+ The 640 label=0 rows are, as far as the leaderboard is concerned, not part of the
105
+ gold set at all. They exist for exactly the classification task this script runs,
106
+ and WP9's write-up of the dataset ("1,505 annotated ground-truth issues") does not
107
+ mention them.
108
+
109
+ ### 3. 74% of it is machine-authored
110
+
111
+ 1,597 comments were written by GPT-5.2, Claude-4.5-Sonnet, Qwen-Coder-480B,
112
+ GLM-4.7, Deepseek-V3.2 or Gemini-3-Pro and *then* expert-verified. Only 548 are
113
+ human-authored. So the corpus over-represents the findings AI reviewers already
114
+ produce, and an adjudicator tuned on it is tuned to police machine output. The
115
+ script always reports the two halves separately and never pools them.
116
+
117
+ Contamination applies as it does to the Martian set: these are public
118
+ repositories, and the PRs are historical. Treat it as a *contaminated, large,
119
+ public* set complementary to the *clean, tiny, private* `skillspro` one — never
120
+ pooled with it either.
121
+
122
+ ## What a result here means for WP6
123
+
124
+ [WP6](../../../../docs/plans/deterministic-pr-levers.md#adjudication-and-the-attention-boundary-wp6)'s
125
+ adjudicator **may re-rank, re-tier and demote, but may not delete a finding
126
+ without a probe transcript refuting it.** Every time a filter has been measured
127
+ against a conservative reviewer it has cost recall — our own candidate v2
128
+ (micro-recall 1/25 → 2/25, F1 halved, reverted), BitsAI-CR's ReviewFilter
129
+ (precision 54.5 → 67.1, recall 45.5 → **39.8**), and Open Code Review's own
130
+ deterministic layer (discards ~5,100 findings and loses 134 real defects doing
131
+ it).
132
+
133
+ So the headline metric is **retention** (kept ÷ label=1), not interception. The
134
+ deterministic arms pin what those numbers are *without* the intervention, which
135
+ is the only thing that makes the intervention's numbers a guard rather than a
136
+ decoration:
137
+
138
+ | Arm | retention | interception | precision | F1 |
139
+ |---|---|---|---|---|
140
+ | `keep-all` — **what production does today** | 100.0% (1505/1505) | 0.0% (0/640) | 70.2% (1505/2145) | 0.825 |
141
+ | `drop-all` | 0.0% (0/1505) | 100.0% (640/640) | n/a (0/0) | n/a |
142
+
143
+ `keep-all`'s precision is the corpus base rate and its F1 of **0.825** is the bar.
144
+ An arm scoring below it has traded recall for nothing — which is the exact shape
145
+ of the reverted candidate v2. A WP6 adjudicator drops in as a fourth arm by
146
+ implementing the `Arm` interface in the script and registering in `ARMS`.
@@ -19,8 +19,15 @@ holds 50 PRs across Sentry / Grafana / Cal.com / Discourse / Keycloak:
19
19
  npx tsx scripts/import-martian.ts # full 50
20
20
  npx tsx scripts/import-martian.ts --limit 3 # a quick subset first
21
21
  npx tsx scripts/import-martian.ts --dry-run # preview without writing
22
+ npx tsx scripts/import-martian.ts --tier pr-review-heldout # import into a sibling tier dir
22
23
  ```
23
24
 
25
+ `--tier <name>` (default `pr-review`) picks the tier directory the import writes
26
+ (`datasets/<name>/`) and the `name` stamped into a freshly-created `tier.json`,
27
+ so a held-out split can live beside this tier under its own name. The tier's
28
+ `defaultWorkflow` stays `pr-review` either way — every imported case runs the
29
+ pr-review workflow.
30
+
24
31
  Then run the tier (heavy — clones the real repos, calls a judge model):
25
32
 
26
33
  ```bash
@@ -50,6 +57,40 @@ Each case's shape (`src/schema.ts`):
50
57
  > a real issue the annotators missed scores as a false positive. That understates
51
58
  > precision, which is why the default is F1 rather than the precision-weighted F0.5.
52
59
 
60
+ ## `anchors.json` — the deterministic anchor labels
61
+
62
+ `anchors.json` **is** committed (unlike `instances.json`) and is the frozen input to the
63
+ code-facts *evidence-coverage* metric. Martian's gold set carries only
64
+ `{severity, description}` — no file, no line — so there is no deterministic join from a gold
65
+ finding to a place in the diff. `scripts/facts-anchors.ts` builds one: it pulls code-shaped
66
+ tokens out of the gold prose (tokenizer `v1`, documented rule-by-rule in the script) and marks a
67
+ finding **anchored** when one of them matches, on a word boundary, an added-or-removed line of the
68
+ three-dot (`base...head`, merge-base) diff. No model is involved — an LLM in the denominator would
69
+ make every coverage number downstream unfalsifiable.
70
+
71
+ ```bash
72
+ npx tsx scripts/facts-anchors.ts # regenerate + print the report
73
+ npx tsx scripts/facts-anchors.ts --dry-run --audit # the hand-audit sheet
74
+ npx tsx scripts/facts-anchors.ts --dry-run --unanchored # what could not be anchored
75
+ ```
76
+
77
+ Three things about it are load-bearing:
78
+
79
+ - **Freeze the labels, not the tokenizer.** The metric's denominator IS the tokenizer's output, so
80
+ the artifact is committed and stamps `tokenizer: "v1"`. A better tokenizer ships as `v2` in a new
81
+ file; it never rewrites what past numbers meant.
82
+ - **It carries its own error bar.** The `audit` block records a seeded random sample of 20 anchored
83
+ findings, each read by hand, with per-finding verdicts. Quote the anchor rate with that rate
84
+ attached, and don't tune the tokenizer to improve it.
85
+ - **No gold text.** Only derived labels (anchors, `path:line`, severity, Martian's `bug_type` /
86
+ `requires_context` / `language`). Join back to your local `instances.json` on
87
+ `instanceId` + `goldIndex` to read a description.
88
+
89
+ The headline: **99/137 anchored (72.3%)**. That is a property of the *gold text*, not of
90
+ code-facts — it is the ceiling on what any identifier-level evidence layer could ever be scored
91
+ against. The 38 unanchored are mostly prose with nothing code-shaped in it (20 of them name no
92
+ identifier at all), plus a stylesheet/i18n cluster where the finding is a value, not a name.
93
+
53
94
  ---
54
95
 
55
96
  **Attribution.** Cases derive from Martian's