lastlight-evals 0.2.1 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -71,6 +71,11 @@ bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
71
71
  overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
72
72
  models you care about. Prefer to drive it by hand? Keep reading.
73
73
 
74
+ > The skill itself lives in a separate repo — it's bundled in the **`lastlight`
75
+ > plugin** ([`cliftonc/lastlight`](https://github.com/cliftonc/lastlight), under
76
+ > `plugins/lastlight/skills/lastlight-evals/`) and tracks this CLI's `init` /
77
+ > `run` surface, so the two are kept in sync.
78
+
74
79
  ### Manual: scaffold with `init`
75
80
 
76
81
  The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
@@ -318,7 +323,20 @@ by name with **overlay > user (`--datasets`) > built-in** precedence:
318
323
  }
319
324
  ```
320
325
 
321
- **Code-fix** — three things keyed by `instance_id`, all under the tier dir:
326
+ Or scaffold one from a **real, resolved issue** — its content, the labels that were
327
+ applied (with who applied them), and reviewer comments become the gold case:
328
+
329
+ ```bash
330
+ lastlight-evals add-case --issue https://github.com/owner/repo/issues/42 --dry-run
331
+ ```
332
+
333
+ It seeds the issue *without* its triage labels (so the agent triages fresh), sets
334
+ `expect_github.labels_added` to the applied labels (+ `issue_closed` if it was
335
+ closed), and prints the labels/comments as evidence; you then assign
336
+ `triage_gold` (category/state) per your deployment's taxonomy.
337
+
338
+ **Code-fix (vendored fixture)** — three things keyed by `instance_id`, all under
339
+ the tier dir:
322
340
 
323
341
  ```
324
342
  <tier>/instances.json # the SweBenchInstance (FAIL_TO_PASS / PASS_TO_PASS)
@@ -326,6 +344,24 @@ by name with **overlay > user (`--datasets`) > built-in** precedence:
326
344
  <tier>/tests/<id>/ # held-out test files, copied in at grade time
327
345
  ```
328
346
 
347
+ **Code-fix from a real PR (git-source)** — point the CLI at a merged PR instead
348
+ of hand-building a fixture:
349
+
350
+ ```bash
351
+ lastlight-evals add-case --pr https://github.com/owner/repo/pull/123 --dry-run
352
+ ```
353
+
354
+ It reads the PR with `gh`, computes `base_commit` (the merge-base of the base
355
+ branch and the PR head) + `head_commit`, captures the PR's **test** diff as the
356
+ held-out `test_patch`, and — unless `--no-validate` — runs the tests at base
357
+ (red) vs head (green) to fill `FAIL_TO_PASS` / `PASS_TO_PASS`. Drop `--dry-run`
358
+ to write it (to `--datasets <dir>` / `--overlay <dir>`, else `./datasets`). No
359
+ `repos/<id>/` is vendored: at run time the harness clones the repo into the
360
+ gitignored `./.eval-cache/` and checks out `base_commit`. Non-`node --test`
361
+ runners work via `--test-cmd "<cmd>"` (+ `--setup-cmd "<cmd>"`), graded on the
362
+ test command's exit code (suite mode) when it emits no TAP names. The repo's
363
+ tests run real code — only use trusted repos.
364
+
329
365
  A new tier just needs a directory with an `instances.json` and a `tier.json`
330
366
  (`{ "name", "defaultWorkflow", "description" }`); per-instance `workflow` wins
331
367
  when present.