lastlight-evals 0.17.0 → 0.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -516,6 +516,57 @@ findings it extracted, the gold set, the finding↔gold pairing (matched / false
516
516
  positive / missed), and its raw replies — so the F1 score is inspectable, not a
517
517
  black box.
518
518
 
519
+ ### Re-review chains (`rounds`)
520
+
521
+ A pr-review case can replay a PR's whole review history: list the heads it was
522
+ reviewed at, oldest first, and the harness runs the real `pr-review` workflow
523
+ once per head, in order, against one fake GitHub and one workspace — so round 2
524
+ sees round 1's review, threads and review ledger exactly as a production
525
+ re-review would.
526
+
527
+ ```jsonc
528
+ {
529
+ "instance_id": "prreview__skillspro-1680",
530
+ "pr": { "number": 1680, "head_commit": "557eb5bc…", /* … */ },
531
+ "rounds": [
532
+ { "head_commit": "1d60d667…", "label": "opened" },
533
+ { "head_commit": "fcfc6296…", "label": "after the first fix push" },
534
+ { "head_commit": "557eb5bc…" } // = pr.head_commit — the scored round
535
+ ]
536
+ }
537
+ ```
538
+
539
+ Full 40-hex SHAs; the last round must be `pr.head_commit`. The case's grade,
540
+ cost and artifacts stay the LAST round's, so it scores like the single-round
541
+ case of its final head; every round's own record lands on the result's
542
+ `rereview` — late discoveries (round ≥ 2 inline comments on lines unchanged
543
+ since the previous head), `converged` / `already-raised` withholds, coverage,
544
+ the ledger, gold matched and cost — shown under **Re-review rounds** and the
545
+ row's **rounds** button in the dashboard. A case with no `rounds` (or one)
546
+ runs exactly as before.
547
+
548
+ Two things keep an earlier round honest:
549
+
550
+ - **Its diff is its own.** Each round's changed files are taken from the merge
551
+ base of `pr.base_commit` and that round's head, as GitHub computes them — a
552
+ branch that later merged or rebased onto a newer main still shows round 1
553
+ only its own changes.
554
+ - **It cannot read the future.** Human discussion seeded on the case is served
555
+ from the start unless it carries `from_round` (on `reviews`,
556
+ `review_comments` or `issue_comments`): the first round (1-based) that may see
557
+ it — the round whose real bot review came after it. Without it, round 1 would
558
+ read the human review later rounds are graded on.
559
+
560
+ ```jsonc
561
+ "reviews": [
562
+ { "user": "SociableSteve", "state": "CHANGES_REQUESTED", "body": "…", "commit_id": "fcfc6296…", "from_round": 3 }
563
+ ]
564
+ ```
565
+
566
+ For the $0 version — no model, just which code each round had already seen —
567
+ run `npx tsx scripts/rereview-delta-replay.ts --repo <checkout> --base <ref>
568
+ --heads <sha1>,<sha2>,… [--comments comments.json]`.
569
+
519
570
  ## Improving an eval — the loop (`lastlight-evals-loop`)
520
571
 
521
572
  Running an eval gives you a score; the **improvement loop** raises it *without