lastlight-evals 0.5.0 → 0.7.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +47 -0
- package/dashboard/dist/assets/index-D5AO7YF_.js +274 -0
- package/dashboard/dist/assets/index-DllAFtF8.css +1 -0
- package/dashboard/dist/index.html +20 -3
- package/dist/mechanism.test.js +54 -2
- package/dist/mechanism.test.js.map +1 -1
- package/dist/run-instance.js +56 -2
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +15 -0
- package/dist/run.js.map +1 -1
- package/dist/seed.js +56 -0
- package/dist/seed.js.map +1 -1
- package/package.json +2 -2
- package/dashboard/dist/assets/index-DbWEwzO3.css +0 -1
- package/dashboard/dist/assets/index-GvQi4Peq.js +0 -264
package/README.md
CHANGED
|
@@ -473,6 +473,53 @@ findings it extracted, the gold set, the finding↔gold pairing (matched / false
|
|
|
473
473
|
positive / missed), and its raw replies — so the F1 score is inspectable, not a
|
|
474
474
|
black box.
|
|
475
475
|
|
|
476
|
+
## Improving an eval — the loop (`lastlight-evals-loop`)
|
|
477
|
+
|
|
478
|
+
Running an eval gives you a score; the **improvement loop** raises it *without
|
|
479
|
+
gaming it*. It's driven by the sibling **`lastlight-evals-loop`** skill (say
|
|
480
|
+
*"raise the pr-review F1"*) and two read-only helpers in `scripts/`. The method —
|
|
481
|
+
**mine the failures → propose a few minimal candidate fixes → keep the one that
|
|
482
|
+
survives a blind held-out gate** — follows
|
|
483
|
+
[*Self-Harness: Harnesses That Improve Themselves*](https://arxiv.org/abs/2606.09498),
|
|
484
|
+
adapted to keep the anti-gaming discipline below.
|
|
485
|
+
|
|
486
|
+
One round:
|
|
487
|
+
|
|
488
|
+
1. **Diagnose (mine).** `scripts/mine-failures.ts` reads the **TRAIN** split of a
|
|
489
|
+
scorecard and clusters the judge's `falseNegatives` (recall loss, weighted by
|
|
490
|
+
severity) and `falsePositives` (precision loss) into a **ranked signature
|
|
491
|
+
bundle** — the systematic patterns, ordered by F1 headroom, instead of reading
|
|
492
|
+
traces by hand.
|
|
493
|
+
|
|
494
|
+
```bash
|
|
495
|
+
npx tsx scripts/mine-failures.ts <train-scorecard>.json --train <train-ids> --keywords
|
|
496
|
+
```
|
|
497
|
+
|
|
498
|
+
2. **Propose.** Draft a few (K=2–4) **minimal, diverse** candidate edits for the
|
|
499
|
+
top pattern — lowest lever first: a generic overlay prompt/skill/persona edit,
|
|
500
|
+
or a synthetic `AGENTS.md` injected into the checkout. Never a core change.
|
|
501
|
+
3. **Select on TRAIN, confirm on HELD-OUT once.** Rank candidates on the train
|
|
502
|
+
split, then give the single winner **one** blind held-out confirmation (gating
|
|
503
|
+
every candidate on held-out would inflate it). `scripts/diff-runs.ts` computes
|
|
504
|
+
the keep/revert verdict:
|
|
505
|
+
|
|
506
|
+
```bash
|
|
507
|
+
npx tsx scripts/diff-runs.ts <baseline>.json <winner>.json \
|
|
508
|
+
--train <train-ids> --heldout <heldout-ids>
|
|
509
|
+
# VERDICT: KEEP — train ↑ and held-out held (or REVERT — OVERFIT: held-out regressed)
|
|
510
|
+
# --symmetric swaps in the paper's non-regressive gate (neither split may regress).
|
|
511
|
+
```
|
|
512
|
+
|
|
513
|
+
4. **Keep one, journal, repeat** until a target F1 or a plateau.
|
|
514
|
+
|
|
515
|
+
**What keeps it honest:** a fixed **train / blind held-out** split (the empirical
|
|
516
|
+
gate), **one change kept per round** (attribution), an adversarial **generality +
|
|
517
|
+
leak auditor** that rejects any edit naming a specific repo/file or encoding the
|
|
518
|
+
gold answer, and **generic-first** levers — core is never touched. The loop
|
|
519
|
+
produces two durable outputs: workflow improvements (better prompts/skills for
|
|
520
|
+
every repo) and per-repo recommendations (context a maintainer can commit), each
|
|
521
|
+
backed by a measured held-out lift.
|
|
522
|
+
|
|
476
523
|
## Models (`models.json`)
|
|
477
524
|
|
|
478
525
|
- `default` — the single model `run` uses.
|