@tangle-network/agent-eval 0.180.0 → 0.181.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +60 -0
- package/README.md +119 -159
- package/dist/adapters/http.d.ts +2 -2
- package/dist/{agent-profile-B7yErX0q.d.ts → agent-profile-CivaSsSy.d.ts} +4 -4
- package/dist/{agent-profile-B7yErX0q.d.ts.map → agent-profile-CivaSsSy.d.ts.map} +1 -1
- package/dist/{agent-profile-cell-0gSi5ffD.js → agent-profile-cell-Cv6UA-W_.js} +20 -57
- package/dist/agent-profile-cell-Cv6UA-W_.js.map +1 -0
- package/dist/{agent-profile-cell-CTOZJUuE.d.ts → agent-profile-cell-s__adRnK.d.ts} +3 -3
- package/dist/agent-profile-cell-s__adRnK.d.ts.map +1 -0
- package/dist/analyst/index.d.ts +10 -10
- package/dist/analyst/index.js +3 -3
- package/dist/ast-CP9ae9B0.js +557 -0
- package/dist/ast-CP9ae9B0.js.map +1 -0
- package/dist/ast-hI-vjW6J.d.ts +457 -0
- package/dist/ast-hI-vjW6J.d.ts.map +1 -0
- package/dist/{benchmark-command-D4vpnAdO.js → benchmark-command-B57n9vjz.js} +7 -6
- package/dist/{benchmark-command-D4vpnAdO.js.map → benchmark-command-B57n9vjz.js.map} +1 -1
- package/dist/benchmarks/index.d.ts +4 -4
- package/dist/benchmarks/index.js +3 -3
- package/dist/campaign/index.d.ts +6 -6
- package/dist/campaign/index.js +8 -8
- package/dist/{campaign-BGEurASO.js → campaign-4_ppJW5X.js} +12 -12
- package/dist/{campaign-BGEurASO.js.map → campaign-4_ppJW5X.js.map} +1 -1
- package/dist/{campaign-evidence-D8DBLqLI.js → campaign-evidence-B8oF9xQ6.js} +515 -471
- package/dist/campaign-evidence-B8oF9xQ6.js.map +1 -0
- package/dist/cli.js +4 -7
- package/dist/cli.js.map +1 -1
- package/dist/{client-BlLY6o2w.js → client-CXE-U1SA.js} +3 -1
- package/dist/client-CXE-U1SA.js.map +1 -0
- package/dist/{client-CuQgX33c.d.ts → client-kh2jOjTK.d.ts} +4 -4
- package/dist/{client-CuQgX33c.d.ts.map → client-kh2jOjTK.d.ts.map} +1 -1
- package/dist/contract/index.d.ts +13 -13
- package/dist/contract/index.js +10 -9
- package/dist/contract/index.js.map +1 -1
- package/dist/{default-registry-IGDE9XIC.d.ts → default-registry-BwDSWVzg.d.ts} +6 -6
- package/dist/{default-registry-IGDE9XIC.d.ts.map → default-registry-BwDSWVzg.d.ts.map} +1 -1
- package/dist/{define-agent-eval-Cx4Ls9ta.d.ts → define-agent-eval-CwOWWQt_.d.ts} +33 -12
- package/dist/define-agent-eval-CwOWWQt_.d.ts.map +1 -0
- package/dist/{define-agent-eval-Dzidv34q.js → define-agent-eval-Ddu33JH9.js} +134 -67
- package/dist/define-agent-eval-Ddu33JH9.js.map +1 -0
- package/dist/{dspy-rlm-engine-xKiWmj_G.js → dspy-rlm-engine-S53V0HhE.js} +2 -2
- package/dist/{dspy-rlm-engine-xKiWmj_G.js.map → dspy-rlm-engine-S53V0HhE.js.map} +1 -1
- package/dist/{engine-CX8ReXkn.d.ts → engine-DS1cysJy.d.ts} +10 -7
- package/dist/engine-DS1cysJy.d.ts.map +1 -0
- package/dist/{eval-campaign-Cs-7MiCs.js → eval-campaign-aYdtjtJR.js} +4 -4
- package/dist/{eval-campaign-Cs-7MiCs.js.map → eval-campaign-aYdtjtJR.js.map} +1 -1
- package/dist/{exact-types-B7LC1EyX.d.ts → exact-types-BZDe0W2D.d.ts} +2 -2
- package/dist/{exact-types-B7LC1EyX.d.ts.map → exact-types-BZDe0W2D.d.ts.map} +1 -1
- package/dist/experiment/index.d.ts +27 -477
- package/dist/experiment/index.d.ts.map +1 -1
- package/dist/experiment/index.js +95 -559
- package/dist/experiment/index.js.map +1 -1
- package/dist/{experiment-tracker-B3TiF5-u.d.ts → experiment-tracker-C7PfnF4b.d.ts} +2 -2
- package/dist/{experiment-tracker-B3TiF5-u.d.ts.map → experiment-tracker-C7PfnF4b.d.ts.map} +1 -1
- package/dist/{external-optimizer-process-Dlz8YxrT.js → external-optimizer-process-QDRURJAM.js} +3 -3
- package/dist/{external-optimizer-process-Dlz8YxrT.js.map → external-optimizer-process-QDRURJAM.js.map} +1 -1
- package/dist/{external-optimizer-subprocess-q3VzlGAO.js → external-optimizer-subprocess-D4dzUBZI.js} +3 -2
- package/dist/{external-optimizer-subprocess-q3VzlGAO.js.map → external-optimizer-subprocess-D4dzUBZI.js.map} +1 -1
- package/dist/{feedback-trajectory-eHWNv5Aj.d.ts → feedback-trajectory-CXmtITBo.d.ts} +3 -3
- package/dist/{feedback-trajectory-eHWNv5Aj.d.ts.map → feedback-trajectory-CXmtITBo.d.ts.map} +1 -1
- package/dist/hosted/index.d.ts +2 -2
- package/dist/hosted/index.d.ts.map +1 -1
- package/dist/hosted/index.js +1 -1
- package/dist/{index-BxWvILU8.d.ts → index-Bp_6sj3x.d.ts} +109 -56
- package/dist/index-Bp_6sj3x.d.ts.map +1 -0
- package/dist/{index-e7LXeRVa.d.ts → index-CJ3LhKIX.d.ts} +2 -2
- package/dist/{index-e7LXeRVa.d.ts.map → index-CJ3LhKIX.d.ts.map} +1 -1
- package/dist/{index-CiUjjEIa.d.ts → index-DNntP4ch.d.ts} +7 -7
- package/dist/{index-CiUjjEIa.d.ts.map → index-DNntP4ch.d.ts.map} +1 -1
- package/dist/{index-DxNYmx4a.d.ts → index-DoykkxW0.d.ts} +11 -11
- package/dist/{index-DxNYmx4a.d.ts.map → index-DoykkxW0.d.ts.map} +1 -1
- package/dist/index.d.ts +28 -28
- package/dist/index.js +24 -15
- package/dist/index.js.map +1 -1
- package/dist/{insight-report-DETqPc_A.d.ts → insight-report-D1qa0HWs.d.ts} +9 -5
- package/dist/{insight-report-DETqPc_A.d.ts.map → insight-report-D1qa0HWs.d.ts.map} +1 -1
- package/dist/{integrity-BKTcA-HP.d.ts → integrity-rGOfSUle.d.ts} +2 -2
- package/dist/{integrity-BKTcA-HP.d.ts.map → integrity-rGOfSUle.d.ts.map} +1 -1
- package/dist/{ledger-core-Cs9f7385.js → journal-Cs9f7385.js} +1 -1
- package/dist/journal-Cs9f7385.js.map +1 -0
- package/dist/{judge-calibration-C5CbMYce.d.ts → judge-calibration-DFtEMlde.d.ts} +31 -2
- package/dist/judge-calibration-DFtEMlde.d.ts.map +1 -0
- package/dist/{judge-calibration-BnpVKtnb.js → judge-calibration-DYmaBtJr.js} +48 -2
- package/dist/{judge-calibration-BnpVKtnb.js.map → judge-calibration-DYmaBtJr.js.map} +1 -1
- package/dist/ledger-core/index.d.ts +1 -1
- package/dist/ledger-core/index.js +1 -1
- package/dist/{llm-judge-v80Kmu9g.js → llm-judge-DEFZeSiu.js} +645 -456
- package/dist/llm-judge-DEFZeSiu.js.map +1 -0
- package/dist/{matrix-DeMmnWrP.d.ts → matrix-CyhW-vgJ.d.ts} +2 -2
- package/dist/{matrix-DeMmnWrP.d.ts.map → matrix-CyhW-vgJ.d.ts.map} +1 -1
- package/dist/meta-eval/index.d.ts +138 -7
- package/dist/meta-eval/index.d.ts.map +1 -1
- package/dist/meta-eval/index.js +245 -97
- package/dist/meta-eval/index.js.map +1 -1
- package/dist/{mint-Cc1_zwRQ.js → mint-ySIIkKlV.js} +2 -2
- package/dist/{mint-Cc1_zwRQ.js.map → mint-ySIIkKlV.js.map} +1 -1
- package/dist/multishot/golden/index.d.ts +1 -1
- package/dist/multishot/index.d.ts +2 -2
- package/dist/openapi.json +1 -1
- package/dist/outcome-store-BXlkwMPR.js +131 -0
- package/dist/outcome-store-BXlkwMPR.js.map +1 -0
- package/dist/{outcome-store-BYHIuO0e.d.ts → outcome-store-CNt4iZ67.d.ts} +18 -25
- package/dist/outcome-store-CNt4iZ67.d.ts.map +1 -0
- package/dist/{paired-promotion-decision-CGzg0cI_.d.ts → paired-promotion-decision-DPsMQm-0.d.ts} +13 -7
- package/dist/{paired-promotion-decision-CGzg0cI_.d.ts.map → paired-promotion-decision-DPsMQm-0.d.ts.map} +1 -1
- package/dist/pipelines/index.js +1 -1
- package/dist/{produced-state-Cv0kJJuP.js → produced-state-BHboMaab.js} +3 -3
- package/dist/{produced-state-Cv0kJJuP.js.map → produced-state-BHboMaab.js.map} +1 -1
- package/dist/profile-cell.d.ts +1 -1
- package/dist/profile-cell.js +1 -1
- package/dist/{promotion-policy-DWOm70gx.js → promotion-policy-CDMMxzb6.js} +28 -40
- package/dist/promotion-policy-CDMMxzb6.js.map +1 -0
- package/dist/{registry-ByVld1-5.d.ts → registry-BRbB6Y0v.d.ts} +4 -4
- package/dist/{registry-ByVld1-5.d.ts.map → registry-BRbB6Y0v.d.ts.map} +1 -1
- package/dist/{release-confidence-BAcNYOf1.d.ts → release-confidence-BcqeQTHW.d.ts} +3 -3
- package/dist/{release-confidence-BAcNYOf1.d.ts.map → release-confidence-BcqeQTHW.d.ts.map} +1 -1
- package/dist/{release-confidence-BcGCclTB.js → release-confidence-DMg8n18l.js} +2 -2
- package/dist/{release-confidence-BcGCclTB.js.map → release-confidence-DMg8n18l.js.map} +1 -1
- package/dist/reporting.d.ts +4 -4
- package/dist/reporting.js +3 -3
- package/dist/{researcher-jsW1X94L.d.ts → researcher-64T49THL.d.ts} +6 -6
- package/dist/{researcher-jsW1X94L.d.ts.map → researcher-64T49THL.d.ts.map} +1 -1
- package/dist/{reward-hacking-ZXEi9VCq.d.ts → reward-hacking-uzO_ihep.d.ts} +2 -2
- package/dist/{reward-hacking-ZXEi9VCq.d.ts.map → reward-hacking-uzO_ihep.d.ts.map} +1 -1
- package/dist/rl.d.ts +53 -99
- package/dist/rl.d.ts.map +1 -1
- package/dist/rl.js +182 -169
- package/dist/rl.js.map +1 -1
- package/dist/rollout/index.d.ts +1 -1
- package/dist/rollout/index.js +2 -2
- package/dist/{rollout-DmoJVqrF.js → rollout-B-UF5R6w.js} +2 -2
- package/dist/{rollout-DmoJVqrF.js.map → rollout-B-UF5R6w.js.map} +1 -1
- package/dist/rubric-predictive-validity-Bmj2_cll.d.ts +79 -0
- package/dist/rubric-predictive-validity-Bmj2_cll.d.ts.map +1 -0
- package/dist/rubric-predictive-validity-CCK-1B7w.js +178 -0
- package/dist/rubric-predictive-validity-CCK-1B7w.js.map +1 -0
- package/dist/{run-record-DTv1MdjK.d.ts → run-record-BiTWauyO.d.ts} +2 -2
- package/dist/{run-record-DTv1MdjK.d.ts.map → run-record-BiTWauyO.d.ts.map} +1 -1
- package/dist/run-record-Br-Yzt_k.js +464 -0
- package/dist/run-record-Br-Yzt_k.js.map +1 -0
- package/dist/{run-record-DQpSf7t-.js → run-record-DualPTn2.js} +2 -2
- package/dist/{run-record-DQpSf7t-.js.map → run-record-DualPTn2.js.map} +1 -1
- package/dist/{semantic-concept-judge-Bi6_iGqg.js → semantic-concept-judge-Bm5JDEKO.js} +3 -3
- package/dist/{semantic-concept-judge-Bi6_iGqg.js.map → semantic-concept-judge-Bm5JDEKO.js.map} +1 -1
- package/dist/{sequential-B5gXgcyp.js → sequential-DAsyV2T9.js} +42 -25
- package/dist/sequential-DAsyV2T9.js.map +1 -0
- package/dist/{series-convergence-DeG33RpC.d.ts → series-convergence-BnMs_uAr.d.ts} +3 -3
- package/dist/{series-convergence-DeG33RpC.d.ts.map → series-convergence-BnMs_uAr.d.ts.map} +1 -1
- package/dist/{skillopt-optimization-method-C3oYul8v.js → skillopt-optimization-method-CL_0aArC.js} +5 -5
- package/dist/{skillopt-optimization-method-C3oYul8v.js.map → skillopt-optimization-method-CL_0aArC.js.map} +1 -1
- package/dist/{statistical-heldout-0La5ZTlv.d.ts → statistical-heldout-CpVd6FmY.d.ts} +207 -144
- package/dist/statistical-heldout-CpVd6FmY.d.ts.map +1 -0
- package/dist/{store-tool-spans-4J1EDElP.d.ts → store-tool-spans-Dt-YdAuE.d.ts} +6 -6
- package/dist/{store-tool-spans-4J1EDElP.d.ts.map → store-tool-spans-Dt-YdAuE.d.ts.map} +1 -1
- package/dist/{summary-report-gMrbYawB.d.ts → summary-report-D1h4dlrK.d.ts} +3 -3
- package/dist/{summary-report-gMrbYawB.d.ts.map → summary-report-D1h4dlrK.d.ts.map} +1 -1
- package/dist/{summary-report-B16xy9Kd.js → summary-report-e-MaOAHV.js} +2 -2
- package/dist/{summary-report-B16xy9Kd.js.map → summary-report-e-MaOAHV.js.map} +1 -1
- package/dist/{tool-groups-2QA0S7dK.d.ts → tool-groups-B2bSNaJB.d.ts} +3 -3
- package/dist/tool-groups-B2bSNaJB.d.ts.map +1 -0
- package/dist/{tool-waste-B9tdWV6g.js → tool-waste-C7MU9u1e.js} +2 -2
- package/dist/{tool-waste-B9tdWV6g.js.map → tool-waste-C7MU9u1e.js.map} +1 -1
- package/dist/trace-repair/index.d.ts +2 -2
- package/dist/traces.d.ts +6 -6
- package/dist/traces.js +1 -1
- package/dist/{types-BmlkCrg0.d.ts → types-BvZoPTGa.d.ts} +3 -3
- package/dist/{types-BmlkCrg0.d.ts.map → types-BvZoPTGa.d.ts.map} +1 -1
- package/dist/{types-gvRsyJLh.d.ts → types-CBbLtr2J.d.ts} +38 -3
- package/dist/{types-gvRsyJLh.d.ts.map → types-CBbLtr2J.d.ts.map} +1 -1
- package/dist/{types-C34V4Vto.d.ts → types-CS0qk_Yp.d.ts} +4 -4
- package/dist/{types-C34V4Vto.d.ts.map → types-CS0qk_Yp.d.ts.map} +1 -1
- package/dist/{types-DzuaM493.d.ts → types-D7gEdPoQ.d.ts} +3 -3
- package/dist/{types-DzuaM493.d.ts.map → types-D7gEdPoQ.d.ts.map} +1 -1
- package/dist/wire/index.d.ts +2 -2
- package/docs/adapters-observability.md +14 -0
- package/docs/campaign-proposers.md +86 -128
- package/docs/charter.md +108 -112
- package/docs/concepts.md +157 -69
- package/docs/design/mlbenchmarks-book-review.md +440 -0
- package/docs/design/mlbenchmarks-review/observations.json +713 -0
- package/docs/design/mlbenchmarks-review/probes.mts +476 -0
- package/docs/design/mlbenchmarks-review/sources.json +200 -0
- package/docs/design/self-improvement-evidence-audit.md +263 -0
- package/docs/design.md +2 -1
- package/docs/eval-surface-map.md +95 -42
- package/docs/evaluation-integrity.md +220 -0
- package/docs/experiment.md +111 -55
- package/docs/feature-guide.md +5 -6
- package/docs/hosted-ingest-spec.md +4 -11
- package/docs/insight-report.md +187 -455
- package/docs/outcome-validity.md +182 -0
- package/docs/product-eval-adoption.md +1 -2
- package/docs/research-report-methodology.md +7 -7
- package/docs/search-history-receipts.md +8 -0
- package/docs/statistical-evidence.md +129 -0
- package/docs/verdicts.md +76 -49
- package/package.json +1 -1
- package/dist/agent-profile-cell-0gSi5ffD.js.map +0 -1
- package/dist/agent-profile-cell-CTOZJUuE.d.ts.map +0 -1
- package/dist/campaign-evidence-D8DBLqLI.js.map +0 -1
- package/dist/client-BlLY6o2w.js.map +0 -1
- package/dist/define-agent-eval-Cx4Ls9ta.d.ts.map +0 -1
- package/dist/define-agent-eval-Dzidv34q.js.map +0 -1
- package/dist/engine-CX8ReXkn.d.ts.map +0 -1
- package/dist/index-BxWvILU8.d.ts.map +0 -1
- package/dist/judge-calibration-C5CbMYce.d.ts.map +0 -1
- package/dist/ledger-core-Cs9f7385.js.map +0 -1
- package/dist/llm-judge-v80Kmu9g.js.map +0 -1
- package/dist/outcome-store-BYHIuO0e.d.ts.map +0 -1
- package/dist/outcome-store-ChBKlTd_.js +0 -75
- package/dist/outcome-store-ChBKlTd_.js.map +0 -1
- package/dist/promotion-policy-DWOm70gx.js.map +0 -1
- package/dist/rubric-predictive-validity-2D5Gw9z9.js +0 -131
- package/dist/rubric-predictive-validity-2D5Gw9z9.js.map +0 -1
- package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts +0 -75
- package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts.map +0 -1
- package/dist/run-record-CR63CpHK.js +0 -216
- package/dist/run-record-CR63CpHK.js.map +0 -1
- package/dist/sequential-B5gXgcyp.js.map +0 -1
- package/dist/statistical-heldout-0La5ZTlv.d.ts.map +0 -1
- package/dist/tool-groups-2QA0S7dK.d.ts.map +0 -1
|
@@ -0,0 +1,263 @@
|
|
|
1
|
+
# Evidence audit: selfImprove and optimizer benefit
|
|
2
|
+
|
|
3
|
+
Historical runs contain gains, nulls, and regressions.
|
|
4
|
+
They do not establish optimizer superiority over a direct edit or current-branch improvement across tasks.
|
|
5
|
+
A historical GEPA analyst comparison reported 0.4285 to 0.4809 pooled micro F1, with important validity limits described below.
|
|
6
|
+
A later GEPA challenger lost 0.0561 on fresh agent families.
|
|
7
|
+
These results justify preserving candidates, checking transfer, and reporting negative or inconclusive outcomes.
|
|
8
|
+
|
|
9
|
+
This was a read-only audit on 2026-09-13 of `feat/evaluation-integrity` over `dda9941437190c9c541b3f54946bfeeb153366fe`.
|
|
10
|
+
The implementation changes were uncommitted during inspection.
|
|
11
|
+
No paid calls or new model experiments ran for this audit.
|
|
12
|
+
Current-branch fixture tests are distinct from all historical model results below.
|
|
13
|
+
|
|
14
|
+
## Inspected evidence
|
|
15
|
+
|
|
16
|
+
Paths below are relative to the repository root unless stated otherwise.
|
|
17
|
+
|
|
18
|
+
- All 10 [registry records](../../evidence/INDEX.md): 2 `CERTIFIED`, 6 `MEASURED-ONCE`, 1 `RESOLVED-NULL`, and 1 `UNVERIFIED`.
|
|
19
|
+
These are registry labels; this audit did not recertify those records.
|
|
20
|
+
- The archived extraction comparison, its original implementation, and 20 optimization-related notebook records.
|
|
21
|
+
- Four committed CodeTraceBench result files, both certification reports, and both preregistrations.
|
|
22
|
+
- Current selfImprove, final-comparison, method-integrity, and final-evidence fixtures.
|
|
23
|
+
|
|
24
|
+
Searches covered `evidence/`, `examples/`, `benchmarks/`, `docs/`, and `.evolve/`.
|
|
25
|
+
There was no root `results/` directory and only one archived method-comparison JSON in examples.
|
|
26
|
+
Queries used `rg`, `git show` at the producing revision, and Python JSON parsing.
|
|
27
|
+
For each committed analyst result, model observation counts, positive-label micro F1, call totals, and cost sums were independently recomputed.
|
|
28
|
+
|
|
29
|
+
The following referenced raw artifacts were unavailable locally:
|
|
30
|
+
|
|
31
|
+
- `~/bench-cache/ctb-20260801/certification/`: incumbent and manual-width-adaptive arms.
|
|
32
|
+
- `~/bench-cache/ctb-20260801/cert2/`: rejected G2 arms.
|
|
33
|
+
- `.evolve/substrate-proof/appworld/d3-scaled-comparison.json`.
|
|
34
|
+
- `.evolve/compare-optimization-methods/` and `.evolve/compare-drivers/`.
|
|
35
|
+
- `examples/findings-ablation/index.ts` and the referenced session scratch artifacts for bridge and CAD proofs.
|
|
36
|
+
|
|
37
|
+
The registry's `.evolve/certification-2026-08-02-preregistration.md` reference is stale.
|
|
38
|
+
The committed replacement is [benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/preregistration.md](../../benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/preregistration.md).
|
|
39
|
+
|
|
40
|
+
## Model-backed evidence
|
|
41
|
+
|
|
42
|
+
| Date and task | Optimizer and comparison | Observed result | Units and costs | Evidence limit |
|
|
43
|
+
| --- | --- | --- | --- | --- |
|
|
44
|
+
| June 1, transaction extraction | Package-local GEPA-style reflection, GEPA-style Pareto, and SkillOpt-style patching versus an intentionally underspecified prompt | Baseline 0.625; reflection and SkillOpt 1.0; Pareto 0.958 | 8 search cases, 6 reported final cases; 182 captured worker calls; $0.013237 captured worker cost | SkillOpt selected on the reported final set; optimizer model costs omitted; no official Python optimizers |
|
|
45
|
+
| June 1, findings ablation | Local GEPA with versus without analyst findings | Both 1.0 from baseline 0.625; difference 0, archived CI [0, 0] | 6 reported cases, 130 calls, $0.009, 131 seconds | Zero findings were generated; the proposed mechanism never activated; notebook only |
|
|
46
|
+
| May 30, legal agents | Local gepaDriver candidates versus incumbent | Candidate fee scores 100 to 83/92; hallucination-free score 100 to 85; selected baseline | 4 scorable personas; 2 named final personas; gen1/pop2/reps2 | Model, cost, and complete paired rows missing; notebook only |
|
|
47
|
+
| June 1, AppWorld difficulty 3 | Local drivers versus a competent baseline, deepseek-v4-pro worker | Small baseline 0.794, lift 0; scaled baseline 0.885, both GEPA lifts 0; memory -0.047 | Small n=6; scaled notebook n=8 with unclear repeated-cell meaning; scaled cost $2.58 | Raw comparison absent; null does not isolate whether remaining errors are prompt-fixable |
|
|
48
|
+
| August 1, CodeTraceBench analyst | Real Python GEPA instructions versus stock and manual width-adaptive arm, glm-5.2 | Pooled micro F1: GEPA 0.4809, stock 0.4285, manual 0.4047 | 69 cases, 2 reps, 138 model observations per arm; search $4.46; six comparison runs documented $50.63 | Paired intervals cross zero; point-estimate promotion; split3 later retired; baseline raw files unavailable |
|
|
49
|
+
| August 2, CodeTraceBench transfer | Real Python GEPA G2 versus shipping G1, same analyst engine/model | Fresh pooled micro F1 0.1928 versus 0.2489; G2 rejected | 64 cases, 128 observations per arm; search $7.05; four comparison runs documented $31.07 | G2 raw files unavailable; intervals cross zero; Terminus2 has 30 source clusters for 32 cases |
|
|
50
|
+
| August 3, cross-family selection | Real GEPA prompt search with a macro objective | Macro 0.163 to 0.187; micro 0.3399 to 0.1897; TP 26 to 11; recall 0.325 to 0.138 | 24 train / 16 selection cases, 56 evaluations, $6.42 | Selection-only rejection; no fresh final experiment; notebook only |
|
|
51
|
+
| August 19–20, AIME | Official Python GEPA and live bridge; glm-5.3 worker, deepseek-v4-flash optimizer | Search completed in attempts 11–14; final comparison never completed | Train8/selection8/final10; 14 launches; about $4 reported | Machinery evidence, no lift verdict; substantial shared-host contention and task timeouts |
|
|
52
|
+
| August 20, CLI bridge toy | Official GEPA autoresearch engine and unmodified Claude CLI through a loopback route | Deterministic length objective about 0.286 to 1.0 | 12 optimizer evaluations, 2 submitted candidates, 8/8 wire calls, $0.029778336, 39.7 seconds | Real candidate-authoring calls; deterministic toy score; registry-only artifacts |
|
|
53
|
+
| June 8, OpenSCAD directive | GEPA versus handwritten directive | Reported +9.5 percentage points on compiled-CAD quality | One final split; task count and cost missing | Weak external-repository pointer; no pinned command or run directory |
|
|
54
|
+
|
|
55
|
+
The AppWorld notebook also reports GSM8K baseline 1.0 with both deepseek-v4-pro and deepseek-v4-flash.
|
|
56
|
+
Its sample counts and costs are absent.
|
|
57
|
+
Its aggregate phrase “five configs” does not reconcile with its enumerated task variants; no total run count is inferred here.
|
|
58
|
+
The findings and extraction records are related demonstrations and must not be counted as independent replication.
|
|
59
|
+
|
|
60
|
+
### June extraction: exact limitations
|
|
61
|
+
|
|
62
|
+
Artifact: [examples/compare-optimization-methods/results/deepseek-chat-20260601.json](../../examples/compare-optimization-methods/results/deepseek-chat-20260601.json).
|
|
63
|
+
It was created at `a648fae334c740d8e0f368e81f68ae933cdb1135` under `examples/compare-drivers-canonical/results/`.
|
|
64
|
+
The current directory name does not identify the historical optimizer implementation.
|
|
65
|
+
|
|
66
|
+
At that revision, inspect these producing sources with `git show`:
|
|
67
|
+
|
|
68
|
+
- `examples/compare-drivers-canonical/index.ts`.
|
|
69
|
+
- `examples/_shared/extraction-task.ts`.
|
|
70
|
+
- `src/campaign/presets/compare-drivers.ts`.
|
|
71
|
+
- `src/campaign/presets/run-skill-opt.ts`.
|
|
72
|
+
- `src/campaign/presets/run-improvement-loop.ts`.
|
|
73
|
+
- `src/campaign/drivers/gepa.ts` and `src/campaign/drivers/skill-opt.ts`.
|
|
74
|
+
|
|
75
|
+
The baseline was `Extract the transaction info from the message as JSON.`
|
|
76
|
+
The optimizer received the omitted schema, formatting rules, and suggested mutation primitives.
|
|
77
|
+
Winners mainly supplied merchant, amount, date, category, and formatting requirements.
|
|
78
|
+
The deterministic composite averages four normalized field matches; six transactions are the sample units, not 24 independent fields.
|
|
79
|
+
No direct edit control tested whether copying the supplied requirements achieved the same benefit.
|
|
80
|
+
|
|
81
|
+
The same `HOLDOUT` entered the inner optimizers and the outer comparison.
|
|
82
|
+
SkillOpt accepted patches using those six cases and fed rejection scores plus accepted-delta notes into subsequent proposals.
|
|
83
|
+
Its reported final score is therefore selection-set performance.
|
|
84
|
+
The GEPA-style entries selected on train, then exposed those cases to an inner gate before outer rescoring.
|
|
85
|
+
Their returned surface did not depend on that gate.
|
|
86
|
+
This establishes unequal final-data access; it does not prove that every numerical GEPA gain was false.
|
|
87
|
+
|
|
88
|
+
| Historical driver | Baseline | Candidate | Lift | Archived lift CI | Captured driver scoring cost |
|
|
89
|
+
| --- | ---: | ---: | ---: | --- | ---: |
|
|
90
|
+
| gepa-reflection | 0.625 | 1.000 | 0.375 | [0.167, 0.542] | $0.002921 |
|
|
91
|
+
| skill-opt | 0.625 | 1.000 | 0.375 | [0.167, 0.542] | $0.004005 |
|
|
92
|
+
| gepa-pareto | 0.625 | 0.958 | 0.333 | [0.167, 0.542] | $0.002929 |
|
|
93
|
+
|
|
94
|
+
Reflection versus SkillOpt had delta 0, CI [0, 0].
|
|
95
|
+
Reflection versus Pareto had delta 0.042, CI [0, 0.125].
|
|
96
|
+
Both archived comparisons used `favored: 'tie'`; that does not establish equivalence.
|
|
97
|
+
|
|
98
|
+
Captured worker usage was 18,226 input tokens and 7,560 output tokens across 182 records over 126 seconds.
|
|
99
|
+
The rate calculation `18226 * 0.27 / 1e6 + 7560 * 1.10 / 1e6 = 0.01323702` matches the rounded artifact cost.
|
|
100
|
+
The listed driver costs sum to $0.009855; $0.003382 of captured worker cost has no named phase breakdown in this artifact.
|
|
101
|
+
More importantly, the captured `records` array was populated only by the extraction worker.
|
|
102
|
+
The optimizer drivers called `callLlm` directly and discarded model usage and cost.
|
|
103
|
+
Therefore, end-to-end optimization cost and total model-call count are unknown.
|
|
104
|
+
The archived `honestVerdict: 'lift-proven'` is not valid current certification or proof of optimizer superiority.
|
|
105
|
+
|
|
106
|
+
### CodeTraceBench: useful gains and failed transfer
|
|
107
|
+
|
|
108
|
+
Full source reports and all measured fields:
|
|
109
|
+
|
|
110
|
+
- [benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/README.md](../../benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/README.md) and `preregistration.md`.
|
|
111
|
+
- [benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/README.md](../../benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/README.md) and `preregistration.md`.
|
|
112
|
+
- `.evolve/experiments.jsonl`, lines 26, 29, 33, 34, and 41.
|
|
113
|
+
|
|
114
|
+
Round 1 used real Python GEPA at historical script revision `0eb2e32`, with 40 evaluations on 10 train and 6 selection cases.
|
|
115
|
+
The output contract stayed fixed.
|
|
116
|
+
The selected prompt hash was `d3829fb855690a3a385f498049801c14bb990c6e49858a6739bd331c0ab324e1`.
|
|
117
|
+
Search selection composite rose from 0.281 to 0.331; this is separate from final micro F1.
|
|
118
|
+
|
|
119
|
+
The final experiment used glm-5.2 through z.ai, stock DSPy RLM execution, seed0, and two repetitions.
|
|
120
|
+
Arms ran serially; each run used concurrency6, maxOutput8192, timeout1,200,000ms, maxCost30, and maxArtifact8MiB.
|
|
121
|
+
The environment was Node24.16.0 on Linux x64.
|
|
122
|
+
The dataset revision was `aa213b84ffb6690fc37ca15766d6ca174ec36d4d`.
|
|
123
|
+
Model names were recorded; a provider-served immutable model snapshot was not demonstrated by this audit.
|
|
124
|
+
|
|
125
|
+
| Arm and final split | Micro F1 | Macro F1 | Recall | Precision | Failed model observations | Documented cost |
|
|
126
|
+
| --- | ---: | ---: | ---: | ---: | --- | ---: |
|
|
127
|
+
| Incumbent / holdout2 | 0.5641 | 0.5596 | 0.6436 | 0.5021 | 0/64 | $7.97 |
|
|
128
|
+
| Manual W / holdout2 | 0.5224 | 0.5125 | 0.5426 | 0.5037 | 1/64 | $7.67 |
|
|
129
|
+
| GEPA G / holdout2 | 0.6288 | 0.5789 | 0.6622 | 0.5986 | 0/64 | $7.96 |
|
|
130
|
+
| Incumbent / split3 | 0.1693 | 0.1830 | 0.3276 | 0.1141 | 1/74 | $9.24 |
|
|
131
|
+
| Manual W / split3 | 0.1805 | 0.1791 | 0.3190 | 0.1259 | 1/74 | $8.72 |
|
|
132
|
+
| GEPA G / split3 | 0.1799 | 0.1844 | 0.3017 | 0.1282 | 0/74 | $9.07 |
|
|
133
|
+
|
|
134
|
+
There were 32 holdout2 cases and 37 split3 cases: 69 cases and 138 model observations per arm.
|
|
135
|
+
Only 30 holdout2 cases had positive labels; the paired F1 table consequently used 30 + 37 = 67 positive cases.
|
|
136
|
+
The notebook's “67 fresh cases” must not replace the execution denominator.
|
|
137
|
+
Pooled macro F1 was G0.3611, incumbent0.3516, and W0.3284.
|
|
138
|
+
|
|
139
|
+
G versus incumbent paired F1 intervals were [-0.027, +0.067] on holdout2 and [-0.042, +0.046] on split3.
|
|
140
|
+
W intervals were [-0.096, +0.001] and [-0.064, +0.053], respectively.
|
|
141
|
+
Reported median paired delta was zero for all four comparisons.
|
|
142
|
+
Promotion followed a preregistered pooled point-estimate rule; it did not require statistical exclusion of zero.
|
|
143
|
+
The wide-cascade holdout2 gain of 0.0647 is a useful historical signal.
|
|
144
|
+
|
|
145
|
+
After certification, split3 was retired: 27/37 cases label the final submit step.
|
|
146
|
+
A constant last-step prediction scored micro F1 0.568 there, versus about 0.180 for the analyst.
|
|
147
|
+
Consequently, the pooled result cannot support a general analysis-quality claim.
|
|
148
|
+
The reports' stronger language about unbiased instruments and certification must be read with this disclosed correction.
|
|
149
|
+
|
|
150
|
+
G2 search later improved a weighted selection objective by 0.064 on 12 cases, with micro F1 0.400 to 0.509 and macro -0.032.
|
|
151
|
+
On fresh OpenHands and Terminus2, G2 pooled micro was 0.1928 versus stock0.2489, with 3/128 versus 0/128 failures.
|
|
152
|
+
Per-family G2 micro was OH0.2086/T20.1822, versus stock OH0.2896/T20.2162.
|
|
153
|
+
Paired intervals were OH[-0.126, +0.058] and T2[-0.143, +0.023].
|
|
154
|
+
The fixed promotion rule rejected G2; its selection gain did not transfer.
|
|
155
|
+
The later August3 macro-versus-micro divergence was found on selection data and is not another fresh final result.
|
|
156
|
+
|
|
157
|
+
### Raw recomputation and accounting
|
|
158
|
+
|
|
159
|
+
These four files contain the surviving model arm and an empty control; they do not contain the unavailable incumbent/W/G2 arms.
|
|
160
|
+
Historical comparison deltas and intervals therefore remain documented evidence rather than independently reconstructed comparisons.
|
|
161
|
+
|
|
162
|
+
| Label | Committed raw result | Run identity SHA256 |
|
|
163
|
+
| --- | --- | --- |
|
|
164
|
+
| G/h2 | [benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/result-holdout2.json](../../benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/result-holdout2.json) | `24883695f29e0b928f3a55d000e985682d18f810eb63c2c70b00e95418a34fac` |
|
|
165
|
+
| G/s3 | [benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/result-split3.json](../../benchmarks/trace-analysis/codetracebench-glm52-certified-20260801/result-split3.json) | `417161b661138fb03977e475ae76be22e54d2153b32bb1842c0a6dde49ecc200` |
|
|
166
|
+
| Stock/OH | [benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/result-stock-openhands.json](../../benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/result-stock-openhands.json) | `30d80958be3d71f56a780850c7791cb4d078a8d98bc696140181611f94f7be95` |
|
|
167
|
+
| Stock/T2 | [benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/result-stock-terminus2.json](../../benchmarks/trace-analysis/codetracebench-crossfamily-cert-20260802/result-stock-terminus2.json) | `abda3035ded04b3625981b2073d6b24a152fb4223836471481d4858f8a9589a3` |
|
|
168
|
+
|
|
169
|
+
| Label | Observations / cases / clusters | Positive / negative / unlabeled observations | Positive-label TP / FP / FN | Recomputed micro F1 | Model calls | Estimated model cost |
|
|
170
|
+
| --- | --- | --- | --- | ---: | ---: | ---: |
|
|
171
|
+
| G/h2 | 64 / 32 / 32 | 60 / 0 / 4 | 249 / 167 / 127 | 0.628787879 | 912 | $7.9586602 |
|
|
172
|
+
| G/s3 | 74 / 37 / 37 | 74 / 0 / 0 | 35 / 238 / 81 | 0.179948586 | 1045 | $9.0744110 |
|
|
173
|
+
| Stock/OH | 64 / 32 / 32 | 32 / 28 / 4 | 43 / 80 / 131 | 0.289562290 | 855 | $7.3335628 |
|
|
174
|
+
| Stock/T2 | 64 / 32 / 30 | 32 / 20 / 12 | 40 / 130 / 160 | 0.216216216 | 883 | $7.5553616 |
|
|
175
|
+
|
|
176
|
+
Counts, call totals, and estimated cost sums reconcile against each result summary.
|
|
177
|
+
All four model arms have zero failed observations, zero unknown-cost observations, and zero unknown token-usage observations.
|
|
178
|
+
Every model observation explicitly labels its cost `estimated`; the reports' wording “measured cost” must not imply provider-billed receipts.
|
|
179
|
+
The table does not include GEPA search cost or missing comparison arms.
|
|
180
|
+
|
|
181
|
+
| Label | Input tokens | Output tokens | Cached tokens | Reasoning tokens | Unknown cache-write usage observations |
|
|
182
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
183
|
+
| G/h2 | 2,055,464 | 578,167 | 9,089,024 | 0 | 64/64 |
|
|
184
|
+
| G/s3 | 2,553,890 | 659,891 | 10,150,528 | 0 | 74/74 |
|
|
185
|
+
| Stock/OH | 2,336,470 | 445,036 | 8,254,336 | 0 | 64/64 |
|
|
186
|
+
| Stock/T2 | 2,257,624 | 457,760 | 8,656,192 | 0 | 64/64 |
|
|
187
|
+
|
|
188
|
+
Cache-write usage is unknown for every observation; summary zero totals do not establish measured zeros.
|
|
189
|
+
Trusted-negative false-positive rates are OH0.5 on 28 observations and T20.6 on 20 observations.
|
|
190
|
+
They are null on h2/s3 because those model arms have no trusted-negative observations.
|
|
191
|
+
The normalized calibration F1 includes negative predictions, giving OH0.243626062 and T20.189573460.
|
|
192
|
+
Those are different quantities from the historical positive-label micro F1 above.
|
|
193
|
+
|
|
194
|
+
| Label | Latency minimum ms | Median ms | Mean ms | p95 ms | Maximum ms |
|
|
195
|
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
|
196
|
+
| G/h2 | 58,387.888 | 134,080.197 | 149,395.929 | 280,446.514 | 388,678.939 |
|
|
197
|
+
| G/s3 | 66,049.095 | 131,120.154 | 150,738.482 | 281,274.392 | 348,210.556 |
|
|
198
|
+
| Stock/OH | 50,270.252 | 101,125.396 | 112,446.304 | 203,914.120 | 233,030.922 |
|
|
199
|
+
| Stock/T2 | 40,764.354 | 108,713.527 | 118,436.742 | 206,919.628 | 223,641.819 |
|
|
200
|
+
|
|
201
|
+
Latency is per completed analyst observation under concurrency6; it is not serial campaign duration.
|
|
202
|
+
The exact protocol, implementation, dependency-lock, labels, trace, and candidate hashes remain in the linked raw results.
|
|
203
|
+
|
|
204
|
+
## Current selected-candidate behavior
|
|
205
|
+
|
|
206
|
+
The implementation preserves the search-selected candidate independently of its final gate decision.
|
|
207
|
+
`runFinalComparison()` compares that selected surface without selecting a replacement.
|
|
208
|
+
`src/contract/self-improve-method.ts` returns `selected.winnerSurface`.
|
|
209
|
+
`src/contract/self-improve.ts` returns `result.winnerSurface`.
|
|
210
|
+
|
|
211
|
+
Inspected fixtures establish these intended behaviors:
|
|
212
|
+
|
|
213
|
+
- `tests/campaign/final-evidence-integration.test.ts`: eight independent binary pairs produce lift1 and default-gate `ship`.
|
|
214
|
+
The declared-unit comparison runs twice without consuming evidence when no ledger policy is supplied.
|
|
215
|
+
- The same file: four fractional pairs preserve `WIN` and lift0.215 while the default gate returns `hold`.
|
|
216
|
+
- `tests/contract-self-improve-method-integrity.test.ts`: method-selected `WIN` remains selected despite losing on train.
|
|
217
|
+
Four final dispatches produce lift0.4; this test injects an always-ship gate.
|
|
218
|
+
- Deferred-final fixtures preserve the selected candidate, execute zero final calls, and leave final score/lift absent.
|
|
219
|
+
- No-op fixtures return the baseline and hold; these reflect unchanged selection rather than gate-driven candidate replacement.
|
|
220
|
+
- Ledger lifecycle fixtures cover method and proposer paths with an injected hold gate.
|
|
221
|
+
|
|
222
|
+
These are fixed, marker, or echo fixtures.
|
|
223
|
+
Some controlled receipt fixtures set backend `real`; that literal does not turn them into model-backed efficacy evidence.
|
|
224
|
+
This historical audit inspected fixture source independently of the implementation's test runs.
|
|
225
|
+
The implementation also adds a default-gate regression for a selected candidate that loses on final tasks.
|
|
226
|
+
It checks negative lift, a refusing gate, and preservation of the selected candidate.
|
|
227
|
+
|
|
228
|
+
The branch's `claim` option declares units separately from optional fresh-evidence accounting.
|
|
229
|
+
This separation supports repeated development comparisons without misrepresenting them as new confirmation.
|
|
230
|
+
The interval API also seals the cluster-bootstrap `value` selector in `IntervalSpec`; submitted row evidence no longer chooses it after sealing.
|
|
231
|
+
Neither API correction is itself evidence of improved model behavior.
|
|
232
|
+
|
|
233
|
+
## Smallest meaningful benefit experiment
|
|
234
|
+
|
|
235
|
+
Use the current public `selfImprove({ method, claim })` path with the maintained official optimizer and the production agent entrypoint.
|
|
236
|
+
Choose a real task panel with observed prompt-fixable errors under a reasonable current baseline.
|
|
237
|
+
Do not manufacture benefit solely by withholding known output requirements from that baseline.
|
|
238
|
+
|
|
239
|
+
1. Run a small execution-and-capture smoke before a full search.
|
|
240
|
+
Verify candidate identity, worker and optimizer receipts, missingness, and retained raw scores.
|
|
241
|
+
2. Use three arms: unchanged baseline, a direct edit from the same development evidence, and the official optimizer.
|
|
242
|
+
Fix authoring/search resource limits and preserve actual spending separately for each arm.
|
|
243
|
+
3. Partition by source task or incident into development, selection, and fresh final units.
|
|
244
|
+
Pin the population, practical effect, primary metric, failure policy, and stopping rule before final exposure.
|
|
245
|
+
4. Estimate final sample size from development variation and the desired practical effect using maintained power helpers.
|
|
246
|
+
Count independent source units and account for clustering; do not substitute repeated calls or a universal 20-unit floor.
|
|
247
|
+
5. Freeze the selected candidates and compare all arms on paired final units under a balanced execution schedule.
|
|
248
|
+
Report baseline-to-candidate lift and optimizer-to-direct-edit lift, including uncertainty, failures, costs, and latency.
|
|
249
|
+
6. Use the final-evidence ledger when making a fresh-confirmation claim.
|
|
250
|
+
Retain the candidate, its diff, and all observed scores even when the gate holds or evidence is inconclusive.
|
|
251
|
+
|
|
252
|
+
No paid experiment was launched by this audit.
|
|
253
|
+
A cost forecast should use the current configured model rates and observed smoke usage before the search starts.
|
|
254
|
+
A single successful search demonstrates that run's benefit; repeat searches are required to estimate optimizer reliability across seeds or tasks.
|
|
255
|
+
|
|
256
|
+
## Supported communication
|
|
257
|
+
|
|
258
|
+
The implementation improves how users declare claims, preserve evidence, inspect uncertainty, and compare selected candidates.
|
|
259
|
+
Historical useful-task optimization has produced gains, nulls, and regressions.
|
|
260
|
+
The evidence supports testing automated evaluation engineering through explicit outcome checks and independent confirmation.
|
|
261
|
+
It does not support claiming universal self-improvement, optimizer superiority over a direct edit, or current-branch model-quality lift.
|
|
262
|
+
|
|
263
|
+
See [evaluation integrity](../evaluation-integrity.md) for the public API and the book chapters motivating its methodology.
|
package/docs/design.md
CHANGED
|
@@ -31,7 +31,7 @@ The stack context matters only if you adopt more of it later.
|
|
|
31
31
|
## The dependency rule
|
|
32
32
|
|
|
33
33
|
This section is the public rationale.
|
|
34
|
-
The enforceable maintainer rule lives in [`CLAUDE.md`](../CLAUDE.md#
|
|
34
|
+
The enforceable maintainer rule lives in [`CLAUDE.md`](../CLAUDE.md#dependency-and-evidence-boundaries).
|
|
35
35
|
|
|
36
36
|
**`agent-eval` has zero upward dependencies on a consumer.**
|
|
37
37
|
This is what keeps the package reusable outside our own stack: nothing in here imports from `agent-runtime`, `agent-knowledge`, or `sandbox`, whether at runtime, in development dependencies, or as type-only imports.
|
|
@@ -66,5 +66,6 @@ They are not adoption reference:
|
|
|
66
66
|
- [`building-doctrine.md`](./building-doctrine.md): conventions our agents follow when consuming this package (reachable model defaults, probe-before-debug, experiment integrity checklist)
|
|
67
67
|
- [`design/loop-taxonomy.md`](./design/loop-taxonomy.md): the internal vocabulary for execution drivers, workers, measurements, and proposers
|
|
68
68
|
- [`design/statistics-decisions.md`](./design/statistics-decisions.md): per-statistic trust status, the no-runtime-dependency verdict, and the exact-versus-asymptotic policy at 3–10 repetitions
|
|
69
|
+
- [`design/mlbenchmarks-book-review.md`](./design/mlbenchmarks-book-review.md): book review, source-backed gaps, and proposals for automated evaluation engineering at a pinned repository revision
|
|
69
70
|
- [`research-report-methodology.md`](./research-report-methodology.md): the evidence standard our own research reports are held to
|
|
70
71
|
- [`.claude/skills/agent-eval/SKILL.md`](../.claude/skills/agent-eval/SKILL.md): directives for LLM agents writing integration code, encoding bug classes we have already shipped and fixed once
|
package/docs/eval-surface-map.md
CHANGED
|
@@ -1,20 +1,21 @@
|
|
|
1
1
|
# Eval surface map: which primitive, when
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
| `
|
|
13
|
-
| `
|
|
14
|
-
| `
|
|
15
|
-
| `
|
|
16
|
-
| `
|
|
17
|
-
| `runEvalCampaign` |
|
|
3
|
+
Choose the public entry point by the decision you need to make.
|
|
4
|
+
The [README](../README.md#choose-a-workflow) starts with the common product workflows.
|
|
5
|
+
This reference covers direct execution controls and specialist modules.
|
|
6
|
+
|
|
7
|
+
## The run primitives
|
|
8
|
+
|
|
9
|
+
| Primitive | Import | Use | Returns |
|
|
10
|
+
|---|---|---|---|
|
|
11
|
+
| [`runCampaign()`](../examples/plan-before-you-spend/) | `/campaign` | Execute and judge a scenarios × repetitions grid through caller-owned dispatch. | `CampaignResult` |
|
|
12
|
+
| [`runEval()`](../src/campaign/presets/run-eval.ts) | `/contract` or `/campaign` | Score one surface with campaign defaults. | `CampaignResult` |
|
|
13
|
+
| [`runProfileMatrix()`](../examples/profile-matrix/) | `/campaign` | Run the same cases across named agent profiles with provenance and backend checks. | `RunProfileMatrixResult`, including `.records`. |
|
|
14
|
+
| [`runOptimization()`](./campaign-proposers.md#write-a-custom-candidate-generator) | `/campaign` | Generate, measure, and select candidates on development cases. | Generations and a winner surface. |
|
|
15
|
+
| [`runImprovementLoop()`](./multi-shot-optimization.md) | `/contract` or `/campaign` | Search, compare on final cases, and apply a release gate. | Final comparison, winner, and gate decision. |
|
|
16
|
+
| [`compareOptimizationMethods()`](../examples/compare-optimization-methods/) | `/campaign` | Compare selected surfaces from several methods under declared budgets. | Final paired contrasts, uncertainty, and costs. |
|
|
17
|
+
| [`runEvalCampaign()`](../src/eval-campaign.ts) | Root | Run with a caller-supplied trace sink and emitter. | Campaign result and records. |
|
|
18
|
+
| [`runAgentMatrix()`](../src/matrix/runner.ts) | `/matrix` | Schedule a general Cartesian grid without campaign scoring semantics. | Cell results. |
|
|
18
19
|
|
|
19
20
|
When variants of the same task run inside one `runCampaign`, give those scenarios the same `seedGroup` so each repetition uses common randomness.
|
|
20
21
|
Use `runProfileMatrix` instead when profiles are separate campaign axes.
|
|
@@ -34,14 +35,64 @@ to retry failed cells from Eval's durable campaign cache.
|
|
|
34
35
|
Finalization refuses missing, overlapping, stale, corrupt, or duplicate rows,
|
|
35
36
|
then returns the ordinary `runProfileMatrix` result and its distributions.
|
|
36
37
|
Coverage reports missing, failed, and zero-score rows separately.
|
|
37
|
-
| `runAgentMatrix` | The bare N-axis cartesian scheduler with concurrency control. The layer beneath the eval surface: reach for it only when you need raw scheduling, not eval semantics. | cell results |
|
|
38
38
|
|
|
39
|
-
|
|
40
|
-
**generate** (`runOptimization`) → **gate** (`runImprovementLoop`). `runEvalCampaign`
|
|
41
|
-
is `runCampaign` with capture inverted; `runAgentMatrix` is the scheduler underneath.
|
|
39
|
+
## Claims, evaluator checks, and final evidence
|
|
42
40
|
|
|
43
|
-
|
|
44
|
-
|
|
41
|
+
| Concern | Import | Public API |
|
|
42
|
+
|---|---|---|
|
|
43
|
+
| Declare population and independent units | `/experiment` | `defineEvaluationClaim()`, `summarizeEvaluationUnits()` |
|
|
44
|
+
| Register an executable decision | `/experiment` | `defineExperiment()`, `sealExperiment()`, `openSealedExperiment()` |
|
|
45
|
+
| Check design adequacy at a practical effect | `/experiment` | `clusteredPower()`, `assertDesignAdequate()` |
|
|
46
|
+
| Track final-data reservation and exposure | `/experiment` | `openFinalEvidenceLedger()` |
|
|
47
|
+
| Group reusable comparisons by source unit | `/contract` or `/campaign` | The top-level `claim` option. |
|
|
48
|
+
| Reserve fresh evidence for confirmation | `/contract` or `/campaign` | Optional `finalEvidence: { ledger, requestId, evaluatorDigest }`. |
|
|
49
|
+
| Admit an evaluator against both error limits | `/meta-eval` | `auditEvaluator()` |
|
|
50
|
+
| Test a grader with known incorrect items | `/meta-eval` | [`definePlant()`, `seedPlants()`, `catchRate()`](./plants.md) |
|
|
51
|
+
| Measure agreement and known bias patterns | `/meta-eval` | `calibrateJudgeContinuous()`, `continuousAgreement()`, `positionalBias()`, `verbosityBias()`, `selfPreference()` |
|
|
52
|
+
| Relate scores to declared deployment outcomes | `/meta-eval` | `rubricPredictiveValidity()`, `correlationStudy()`, `calibrationFromPairs()`, `calibrationCurve()` |
|
|
53
|
+
|
|
54
|
+
The top-level `claim` declares the independent unit for ordinary reusable comparisons.
|
|
55
|
+
Include `minimumEffect` when the decision concerns a practical improvement.
|
|
56
|
+
Optional `finalEvidence` requires a comparison or certification claim.
|
|
57
|
+
It reserves units before search and records exposure before final dispatch.
|
|
58
|
+
Its ledger must be shared across related campaigns.
|
|
59
|
+
Keep source unit identifiers stable when scenarios or populations are renamed.
|
|
60
|
+
The host controls access to private evidence and must preserve author/auditor separation.
|
|
61
|
+
[Evaluation integrity](./evaluation-integrity.md) describes these boundaries and the public result shapes.
|
|
62
|
+
|
|
63
|
+
Outcome associations require an explicit desired direction for predictive validity.
|
|
64
|
+
They produce descriptive evidence and experiment hypotheses.
|
|
65
|
+
They do not establish a causal benefit from changing a rubric.
|
|
66
|
+
See [outcome validity](./outcome-validity.md).
|
|
67
|
+
|
|
68
|
+
Method comparisons retain `scenarioScores`, `unitScores`, the `units` summary, and `pairedCellN` separately.
|
|
69
|
+
`favored: null` means the paired decision did not establish a preferred method.
|
|
70
|
+
It does not establish equivalence.
|
|
71
|
+
Use the decision diagnostics and intervals to distinguish insufficient evidence from a supported improvement.
|
|
72
|
+
|
|
73
|
+
## Specialist imports
|
|
74
|
+
|
|
75
|
+
| Subpath | Use |
|
|
76
|
+
|---|---|
|
|
77
|
+
| `/traces`, `/trace-attributes` | Store trace evidence, [connect observability exporters](./adapters-observability.md), and use canonical measurement attribute names. |
|
|
78
|
+
| `/analyst` | Execute declared analysts against recorded evidence. |
|
|
79
|
+
| `/reporting`, `/pipelines` | Compare runs, render [research reports](./research-report-methodology.md), and extract recorded failure patterns. |
|
|
80
|
+
| `/supervisor-run` | Read recursive run directories and their evidence coverage. |
|
|
81
|
+
| `/trace-repair`, `/trajectory-replay` | Execute proposed repairs or replay recorded shell trajectories. |
|
|
82
|
+
| `/benchmarks`, `/fuzz` | Adapt benchmark data and explore a declared behavior space. |
|
|
83
|
+
| `/builder-eval`, `/multishot`, [`/multishot/golden`](./multishot-golden-records.md) | Evaluate generated applications and multi-turn conversations. |
|
|
84
|
+
| `/matrix` | Schedule Cartesian experiment grids. |
|
|
85
|
+
| `/rl` | Build reward, preference, and supervised datasets from eligible evidence. |
|
|
86
|
+
| `/profile-cell` | Create and validate portable agent-profile identities. |
|
|
87
|
+
| `/authenticity`, `/ledger-core` | Check evidence authenticity and maintain canonical hash-chained journals. |
|
|
88
|
+
| [`/rollout`](./rollout.md), `/storyboard` | Serialize training rows and render recorded work. |
|
|
89
|
+
| `/hosted`, `/wire`, `/adapters/http` | Connect hosted storage or expose evaluation through HTTP and RPC. |
|
|
90
|
+
|
|
91
|
+
For existing coding-agent transcripts, start with [session intake](./code-agent-intake.md).
|
|
92
|
+
|
|
93
|
+
Root `Scenario`, `JudgeScore`, and `GateDecision` match `/contract`.
|
|
94
|
+
Use root `ProductScenario` and `DimensionJudgeScore` for the product-judging functions.
|
|
95
|
+
`HeldOutGate.evaluate()` returns the separate root type `HeldOutGateDecision`.
|
|
45
96
|
|
|
46
97
|
## What a campaign result reports: the mean and the spread
|
|
47
98
|
|
|
@@ -50,6 +101,8 @@ release-gate). Keep them separate; pick by the table.
|
|
|
50
101
|
`byScenario` holds one `ScenarioAggregate` per scenario that produced at least one composite.
|
|
51
102
|
|
|
52
103
|
Each aggregate reports a mean, a seeded bootstrap `ci95` band, `n`, and a `distribution`.
|
|
104
|
+
Here, `n` counts observed scores.
|
|
105
|
+
Use registered gates for inference across independent source units; their `pairedCellN` retains the raw paired denominator.
|
|
53
106
|
`distribution` is the `SeriesDistribution` value `summarizeNumberSeries` returns: `n`, `min`, `p50`, `p90`, `max`, and `sum` over the exact scores the mean was taken over.
|
|
54
107
|
Quantiles use the nearest-rank definition, so every reported quantile is a score the campaign measured.
|
|
55
108
|
|
|
@@ -61,22 +114,28 @@ A judge that produced no score has no entry at all.
|
|
|
61
114
|
An absent aggregate is the honest record of an unmeasured judge, and a zero-filled distribution would read as a measured all-zero series.
|
|
62
115
|
|
|
63
116
|
`SeriesDistribution` is the one distribution summary in this package.
|
|
64
|
-
|
|
117
|
+
The [insight report](./insight-report.md) explains why its `ScalarDistribution` has a separate shape.
|
|
65
118
|
|
|
66
119
|
## Planning the cell grid without a run directory
|
|
67
120
|
|
|
68
121
|
`buildCellSchedule(scenarios, seed, reps)` returns the `(scenario × rep)` fan-out: one `CellScheduleSlot` per cell, with its `cellId` and its per-cell seed.
|
|
69
|
-
|
|
122
|
+
This function does not access the filesystem.
|
|
123
|
+
It can size a design and check cell counts and seeds before a run directory exists.
|
|
70
124
|
Scenarios that share a `seedGroup` receive the same per-replicate seeds, which is what makes a paired comparison see common randomness.
|
|
71
125
|
|
|
72
|
-
Use `planCampaignRun`
|
|
126
|
+
Use `planCampaignRun()` to classify cached, pending, and blocked cells.
|
|
127
|
+
That call reads the durable cache in a real run directory.
|
|
73
128
|
`cellDirectory` and `cellCachePath` name a cell's location once a run directory is chosen.
|
|
129
|
+
Use [eval fixtures](./eval-fixtures.md) to load and fingerprint cases from folders before planning their campaign.
|
|
74
130
|
|
|
75
131
|
## Evidence receipts: `attest`
|
|
76
132
|
|
|
77
|
-
`attest(report, provenance)`
|
|
133
|
+
`attest(report, provenance)` binds a serializable report to its provenance through content hashes.
|
|
134
|
+
Provenance records model versions, seeds, the price-table hash, code revision, and input digest.
|
|
78
135
|
`verifyAttestation(report, attested)` returns a typed outcome rather than throwing, so a pipeline records why a report failed to verify instead of dying.
|
|
79
136
|
`ATTESTATION_ALGORITHM` is the hash-scheme tag every attestation carries, and a verifier rejects an unknown algorithm instead of guessing.
|
|
137
|
+
Verification requires an `envelopeHash` that binds the report hash to its provenance.
|
|
138
|
+
An absent, malformed, or mismatched envelope makes the attestation invalid.
|
|
80
139
|
Signing stays with the consumer: an `AttestedReport` is a stable byte-identical payload to sign, and this package never holds keys.
|
|
81
140
|
|
|
82
141
|
## Failed cells: receipts and bounded retry
|
|
@@ -93,24 +152,19 @@ A retried attempt keeps its receipt at `<cell>/failure-receipt.attempt-<n>.json`
|
|
|
93
152
|
With `abortOnCellError`, the abort fires only when a cell's final attempt fails.
|
|
94
153
|
Without `cellRetry`, a failed cell is final: one transient 503 leaves campaign coverage incomplete, and `runImprovementLoop` then refuses the holdout comparison.
|
|
95
154
|
|
|
96
|
-
##
|
|
155
|
+
## Grade produced state through a judge
|
|
97
156
|
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
`verifyCompletion`**: not a dedicated runner. The pipeline:
|
|
157
|
+
Use `extractProducedState()` and `verifyCompletion()` inside a `JudgeConfig` to grade the work an agent produced.
|
|
158
|
+
The same judge can run through `runCampaign()` or `runProfileMatrix()`.
|
|
101
159
|
|
|
102
|
-
```
|
|
103
|
-
runtime
|
|
104
|
-
|
|
105
|
-
verifyCompletion(taskGold, state, correctnessChecker)
|
|
106
|
-
│
|
|
107
|
-
inject as a JudgeConfig into runProfileMatrix / runCampaign
|
|
160
|
+
```text
|
|
161
|
+
runtime events -> extractProducedState(events) -> ProducedState
|
|
162
|
+
-> verifyCompletion(taskGold, state, checker) -> JudgeConfig score
|
|
108
163
|
```
|
|
109
164
|
|
|
110
|
-
|
|
111
|
-
|
|
112
|
-
|
|
113
|
-
that is already one judge. (Archetype: `playback.ts` `scoreUserStory`.)
|
|
165
|
+
The host supplies the correctness checker and the events.
|
|
166
|
+
This composition shares campaign execution, capture, and reporting without another runner.
|
|
167
|
+
See [product patterns](./product-eval-adoption.md#product-patterns) for host adapters and [knowledge readiness](./knowledge-readiness.md) for required context checks.
|
|
114
168
|
|
|
115
169
|
### The in-band body contract
|
|
116
170
|
|
|
@@ -123,6 +177,5 @@ product database to recover it:
|
|
|
123
177
|
proposal is graded presence-only (and, by the completion oracle's rule, does
|
|
124
178
|
not count as a completed deliverable).
|
|
125
179
|
|
|
126
|
-
|
|
127
|
-
|
|
128
|
-
add an enrichment band-aid.
|
|
180
|
+
Carry deliverable content in the event when the host persists it.
|
|
181
|
+
This lets the grader inspect the recorded artifact without querying mutable product storage.
|