lastlight-evals 0.5.0 → 0.7.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -473,6 +473,53 @@ findings it extracted, the gold set, the finding↔gold pairing (matched / false
473
473
  positive / missed), and its raw replies — so the F1 score is inspectable, not a
474
474
  black box.
475
475
 
476
+ ## Improving an eval — the loop (`lastlight-evals-loop`)
477
+
478
+ Running an eval gives you a score; the **improvement loop** raises it *without
479
+ gaming it*. It's driven by the sibling **`lastlight-evals-loop`** skill (say
480
+ *"raise the pr-review F1"*) and two read-only helpers in `scripts/`. The method —
481
+ **mine the failures → propose a few minimal candidate fixes → keep the one that
482
+ survives a blind held-out gate** — follows
483
+ [*Self-Harness: Harnesses That Improve Themselves*](https://arxiv.org/abs/2606.09498),
484
+ adapted to keep the anti-gaming discipline below.
485
+
486
+ One round:
487
+
488
+ 1. **Diagnose (mine).** `scripts/mine-failures.ts` reads the **TRAIN** split of a
489
+ scorecard and clusters the judge's `falseNegatives` (recall loss, weighted by
490
+ severity) and `falsePositives` (precision loss) into a **ranked signature
491
+ bundle** — the systematic patterns, ordered by F1 headroom, instead of reading
492
+ traces by hand.
493
+
494
+ ```bash
495
+ npx tsx scripts/mine-failures.ts <train-scorecard>.json --train <train-ids> --keywords
496
+ ```
497
+
498
+ 2. **Propose.** Draft a few (K=2–4) **minimal, diverse** candidate edits for the
499
+ top pattern — lowest lever first: a generic overlay prompt/skill/persona edit,
500
+ or a synthetic `AGENTS.md` injected into the checkout. Never a core change.
501
+ 3. **Select on TRAIN, confirm on HELD-OUT once.** Rank candidates on the train
502
+ split, then give the single winner **one** blind held-out confirmation (gating
503
+ every candidate on held-out would inflate it). `scripts/diff-runs.ts` computes
504
+ the keep/revert verdict:
505
+
506
+ ```bash
507
+ npx tsx scripts/diff-runs.ts <baseline>.json <winner>.json \
508
+ --train <train-ids> --heldout <heldout-ids>
509
+ # VERDICT: KEEP — train ↑ and held-out held (or REVERT — OVERFIT: held-out regressed)
510
+ # --symmetric swaps in the paper's non-regressive gate (neither split may regress).
511
+ ```
512
+
513
+ 4. **Keep one, journal, repeat** until a target F1 or a plateau.
514
+
515
+ **What keeps it honest:** a fixed **train / blind held-out** split (the empirical
516
+ gate), **one change kept per round** (attribution), an adversarial **generality +
517
+ leak auditor** that rejects any edit naming a specific repo/file or encoding the
518
+ gold answer, and **generic-first** levers — core is never touched. The loop
519
+ produces two durable outputs: workflow improvements (better prompts/skills for
520
+ every repo) and per-repo recommendations (context a maintainer can commit), each
521
+ backed by a measured held-out lift.
522
+
476
523
  ## Models (`models.json`)
477
524
 
478
525
  - `default` — the single model `run` uses.