lastlight-evals 0.17.0 → 0.18.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +51 -0
- package/dashboard/dist/assets/{index-DSu_AxvO.js → index-BiY2Jj1T.js} +38 -38
- package/dashboard/dist/assets/index-Dzo-4g9I.css +1 -0
- package/dashboard/dist/index.html +2 -2
- package/dist/fake-github.js +111 -20
- package/dist/fake-github.js.map +1 -1
- package/dist/phase-replay-context.js +60 -0
- package/dist/phase-replay-context.js.map +1 -0
- package/dist/phase-replay-node.js +12 -3
- package/dist/phase-replay-node.js.map +1 -1
- package/dist/pr-context.js +29 -9
- package/dist/pr-context.js.map +1 -1
- package/dist/rereview-node.js +320 -0
- package/dist/rereview-node.js.map +1 -0
- package/dist/rereview.js +199 -0
- package/dist/rereview.js.map +1 -0
- package/dist/rereview.test.js +512 -0
- package/dist/rereview.test.js.map +1 -0
- package/dist/run-instance.js +543 -253
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +4 -1
- package/dist/run.js.map +1 -1
- package/dist/schema.js.map +1 -1
- package/dist/seed.js +57 -7
- package/dist/seed.js.map +1 -1
- package/package.json +4 -4
- package/dashboard/dist/assets/index-CcEkcmk_.css +0 -1
package/README.md
CHANGED
|
@@ -516,6 +516,57 @@ findings it extracted, the gold set, the finding↔gold pairing (matched / false
|
|
|
516
516
|
positive / missed), and its raw replies — so the F1 score is inspectable, not a
|
|
517
517
|
black box.
|
|
518
518
|
|
|
519
|
+
### Re-review chains (`rounds`)
|
|
520
|
+
|
|
521
|
+
A pr-review case can replay a PR's whole review history: list the heads it was
|
|
522
|
+
reviewed at, oldest first, and the harness runs the real `pr-review` workflow
|
|
523
|
+
once per head, in order, against one fake GitHub and one workspace — so round 2
|
|
524
|
+
sees round 1's review, threads and review ledger exactly as a production
|
|
525
|
+
re-review would.
|
|
526
|
+
|
|
527
|
+
```jsonc
|
|
528
|
+
{
|
|
529
|
+
"instance_id": "prreview__skillspro-1680",
|
|
530
|
+
"pr": { "number": 1680, "head_commit": "557eb5bc…", /* … */ },
|
|
531
|
+
"rounds": [
|
|
532
|
+
{ "head_commit": "1d60d667…", "label": "opened" },
|
|
533
|
+
{ "head_commit": "fcfc6296…", "label": "after the first fix push" },
|
|
534
|
+
{ "head_commit": "557eb5bc…" } // = pr.head_commit — the scored round
|
|
535
|
+
]
|
|
536
|
+
}
|
|
537
|
+
```
|
|
538
|
+
|
|
539
|
+
Full 40-hex SHAs; the last round must be `pr.head_commit`. The case's grade,
|
|
540
|
+
cost and artifacts stay the LAST round's, so it scores like the single-round
|
|
541
|
+
case of its final head; every round's own record lands on the result's
|
|
542
|
+
`rereview` — late discoveries (round ≥ 2 inline comments on lines unchanged
|
|
543
|
+
since the previous head), `converged` / `already-raised` withholds, coverage,
|
|
544
|
+
the ledger, gold matched and cost — shown under **Re-review rounds** and the
|
|
545
|
+
row's **rounds** button in the dashboard. A case with no `rounds` (or one)
|
|
546
|
+
runs exactly as before.
|
|
547
|
+
|
|
548
|
+
Two things keep an earlier round honest:
|
|
549
|
+
|
|
550
|
+
- **Its diff is its own.** Each round's changed files are taken from the merge
|
|
551
|
+
base of `pr.base_commit` and that round's head, as GitHub computes them — a
|
|
552
|
+
branch that later merged or rebased onto a newer main still shows round 1
|
|
553
|
+
only its own changes.
|
|
554
|
+
- **It cannot read the future.** Human discussion seeded on the case is served
|
|
555
|
+
from the start unless it carries `from_round` (on `reviews`,
|
|
556
|
+
`review_comments` or `issue_comments`): the first round (1-based) that may see
|
|
557
|
+
it — the round whose real bot review came after it. Without it, round 1 would
|
|
558
|
+
read the human review later rounds are graded on.
|
|
559
|
+
|
|
560
|
+
```jsonc
|
|
561
|
+
"reviews": [
|
|
562
|
+
{ "user": "SociableSteve", "state": "CHANGES_REQUESTED", "body": "…", "commit_id": "fcfc6296…", "from_round": 3 }
|
|
563
|
+
]
|
|
564
|
+
```
|
|
565
|
+
|
|
566
|
+
For the $0 version — no model, just which code each round had already seen —
|
|
567
|
+
run `npx tsx scripts/rereview-delta-replay.ts --repo <checkout> --base <ref>
|
|
568
|
+
--heads <sha1>,<sha2>,… [--comments comments.json]`.
|
|
569
|
+
|
|
519
570
|
## Improving an eval — the loop (`lastlight-evals-loop`)
|
|
520
571
|
|
|
521
572
|
Running an eval gives you a score; the **improvement loop** raises it *without
|