lastlight-evals 0.2.1 → 0.4.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +37 -1
- package/dashboard/dist/assets/index-D6qxQ8uN.js +258 -0
- package/dashboard/dist/assets/index-DYwJHkLo.css +1 -0
- package/dashboard/dist/index.html +2 -2
- package/dist/add-case.js +532 -0
- package/dist/add-case.js.map +1 -0
- package/dist/clean.js +127 -0
- package/dist/clean.js.map +1 -0
- package/dist/grade.js +58 -8
- package/dist/grade.js.map +1 -1
- package/dist/init.js +6 -2
- package/dist/init.js.map +1 -1
- package/dist/mechanism.test.js +135 -2
- package/dist/mechanism.test.js.map +1 -1
- package/dist/report.js +34 -9
- package/dist/report.js.map +1 -1
- package/dist/run-instance.js +108 -17
- package/dist/run-instance.js.map +1 -1
- package/dist/run.js +126 -9
- package/dist/run.js.map +1 -1
- package/dist/seed.js +120 -16
- package/dist/seed.js.map +1 -1
- package/dist/serve.js +2 -0
- package/dist/serve.js.map +1 -1
- package/package.json +2 -2
- package/dashboard/dist/assets/index-YwScAAxm.js +0 -238
- package/dashboard/dist/assets/index-uU9Sj3M4.css +0 -1
package/README.md
CHANGED
|
@@ -71,6 +71,11 @@ bare `lastlight-evals run` "just works" (it auto-detects `./instance` as the
|
|
|
71
71
|
overlay and `./evals/datasets`). Now you're evaluating *your* agent against the
|
|
72
72
|
models you care about. Prefer to drive it by hand? Keep reading.
|
|
73
73
|
|
|
74
|
+
> The skill itself lives in a separate repo — it's bundled in the **`lastlight`
|
|
75
|
+
> plugin** ([`cliftonc/lastlight`](https://github.com/cliftonc/lastlight), under
|
|
76
|
+
> `plugins/lastlight/skills/lastlight-evals/`) and tracks this CLI's `init` /
|
|
77
|
+
> `run` surface, so the two are kept in sync.
|
|
78
|
+
|
|
74
79
|
### Manual: scaffold with `init`
|
|
75
80
|
|
|
76
81
|
The fastest CLI path is **`init`** — it scaffolds *your own* evals workspace
|
|
@@ -318,7 +323,20 @@ by name with **overlay > user (`--datasets`) > built-in** precedence:
|
|
|
318
323
|
}
|
|
319
324
|
```
|
|
320
325
|
|
|
321
|
-
|
|
326
|
+
Or scaffold one from a **real, resolved issue** — its content, the labels that were
|
|
327
|
+
applied (with who applied them), and reviewer comments become the gold case:
|
|
328
|
+
|
|
329
|
+
```bash
|
|
330
|
+
lastlight-evals add-case --issue https://github.com/owner/repo/issues/42 --dry-run
|
|
331
|
+
```
|
|
332
|
+
|
|
333
|
+
It seeds the issue *without* its triage labels (so the agent triages fresh), sets
|
|
334
|
+
`expect_github.labels_added` to the applied labels (+ `issue_closed` if it was
|
|
335
|
+
closed), and prints the labels/comments as evidence; you then assign
|
|
336
|
+
`triage_gold` (category/state) per your deployment's taxonomy.
|
|
337
|
+
|
|
338
|
+
**Code-fix (vendored fixture)** — three things keyed by `instance_id`, all under
|
|
339
|
+
the tier dir:
|
|
322
340
|
|
|
323
341
|
```
|
|
324
342
|
<tier>/instances.json # the SweBenchInstance (FAIL_TO_PASS / PASS_TO_PASS)
|
|
@@ -326,6 +344,24 @@ by name with **overlay > user (`--datasets`) > built-in** precedence:
|
|
|
326
344
|
<tier>/tests/<id>/ # held-out test files, copied in at grade time
|
|
327
345
|
```
|
|
328
346
|
|
|
347
|
+
**Code-fix from a real PR (git-source)** — point the CLI at a merged PR instead
|
|
348
|
+
of hand-building a fixture:
|
|
349
|
+
|
|
350
|
+
```bash
|
|
351
|
+
lastlight-evals add-case --pr https://github.com/owner/repo/pull/123 --dry-run
|
|
352
|
+
```
|
|
353
|
+
|
|
354
|
+
It reads the PR with `gh`, computes `base_commit` (the merge-base of the base
|
|
355
|
+
branch and the PR head) + `head_commit`, captures the PR's **test** diff as the
|
|
356
|
+
held-out `test_patch`, and — unless `--no-validate` — runs the tests at base
|
|
357
|
+
(red) vs head (green) to fill `FAIL_TO_PASS` / `PASS_TO_PASS`. Drop `--dry-run`
|
|
358
|
+
to write it (to `--datasets <dir>` / `--overlay <dir>`, else `./datasets`). No
|
|
359
|
+
`repos/<id>/` is vendored: at run time the harness clones the repo into the
|
|
360
|
+
gitignored `./.eval-cache/` and checks out `base_commit`. Non-`node --test`
|
|
361
|
+
runners work via `--test-cmd "<cmd>"` (+ `--setup-cmd "<cmd>"`), graded on the
|
|
362
|
+
test command's exit code (suite mode) when it emits no TAP names. The repo's
|
|
363
|
+
tests run real code — only use trusted repos.
|
|
364
|
+
|
|
329
365
|
A new tier just needs a directory with an `instances.json` and a `tier.json`
|
|
330
366
|
(`{ "name", "defaultWorkflow", "description" }`); per-instance `workflow` wins
|
|
331
367
|
when present.
|