@tangle-network/agent-eval 0.179.0 → 0.181.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (236) hide show
  1. package/CHANGELOG.md +66 -0
  2. package/README.md +119 -146
  3. package/dist/adapters/http.d.ts +2 -2
  4. package/dist/{agent-profile-B7yErX0q.d.ts → agent-profile-CivaSsSy.d.ts} +4 -4
  5. package/dist/{agent-profile-B7yErX0q.d.ts.map → agent-profile-CivaSsSy.d.ts.map} +1 -1
  6. package/dist/{agent-profile-cell-0gSi5ffD.js → agent-profile-cell-Cv6UA-W_.js} +20 -57
  7. package/dist/agent-profile-cell-Cv6UA-W_.js.map +1 -0
  8. package/dist/{agent-profile-cell-CTOZJUuE.d.ts → agent-profile-cell-s__adRnK.d.ts} +3 -3
  9. package/dist/agent-profile-cell-s__adRnK.d.ts.map +1 -0
  10. package/dist/analyst/index.d.ts +10 -10
  11. package/dist/analyst/index.js +4 -4
  12. package/dist/ast-CP9ae9B0.js +557 -0
  13. package/dist/ast-CP9ae9B0.js.map +1 -0
  14. package/dist/ast-hI-vjW6J.d.ts +457 -0
  15. package/dist/ast-hI-vjW6J.d.ts.map +1 -0
  16. package/dist/{benchmark-command-CY6Dg5t5.js → benchmark-command-B57n9vjz.js} +7 -6
  17. package/dist/{benchmark-command-CY6Dg5t5.js.map → benchmark-command-B57n9vjz.js.map} +1 -1
  18. package/dist/benchmarks/index.d.ts +4 -4
  19. package/dist/benchmarks/index.js +3 -3
  20. package/dist/campaign/index.d.ts +6 -6
  21. package/dist/campaign/index.js +8 -8
  22. package/dist/{campaign-BGEurASO.js → campaign-4_ppJW5X.js} +12 -12
  23. package/dist/{campaign-BGEurASO.js.map → campaign-4_ppJW5X.js.map} +1 -1
  24. package/dist/{campaign-evidence-D8DBLqLI.js → campaign-evidence-B8oF9xQ6.js} +515 -471
  25. package/dist/campaign-evidence-B8oF9xQ6.js.map +1 -0
  26. package/dist/cli.js +5 -8
  27. package/dist/cli.js.map +1 -1
  28. package/dist/{client-BlLY6o2w.js → client-CXE-U1SA.js} +3 -1
  29. package/dist/client-CXE-U1SA.js.map +1 -0
  30. package/dist/{client-CuQgX33c.d.ts → client-kh2jOjTK.d.ts} +4 -4
  31. package/dist/{client-CuQgX33c.d.ts.map → client-kh2jOjTK.d.ts.map} +1 -1
  32. package/dist/contract/index.d.ts +13 -13
  33. package/dist/contract/index.js +11 -10
  34. package/dist/contract/index.js.map +1 -1
  35. package/dist/{default-registry-IGDE9XIC.d.ts → default-registry-BwDSWVzg.d.ts} +6 -6
  36. package/dist/{default-registry-IGDE9XIC.d.ts.map → default-registry-BwDSWVzg.d.ts.map} +1 -1
  37. package/dist/{default-registry-BryMEmr8.js → default-registry-aL7xUrUz.js} +2 -2
  38. package/dist/{default-registry-BryMEmr8.js.map → default-registry-aL7xUrUz.js.map} +1 -1
  39. package/dist/{define-agent-eval-Cx4Ls9ta.d.ts → define-agent-eval-CwOWWQt_.d.ts} +33 -12
  40. package/dist/define-agent-eval-CwOWWQt_.d.ts.map +1 -0
  41. package/dist/{define-agent-eval-Dzidv34q.js → define-agent-eval-Ddu33JH9.js} +134 -67
  42. package/dist/define-agent-eval-Ddu33JH9.js.map +1 -0
  43. package/dist/{dspy-rlm-engine-xKiWmj_G.js → dspy-rlm-engine-S53V0HhE.js} +2 -2
  44. package/dist/{dspy-rlm-engine-xKiWmj_G.js.map → dspy-rlm-engine-S53V0HhE.js.map} +1 -1
  45. package/dist/{engine-CX8ReXkn.d.ts → engine-DS1cysJy.d.ts} +10 -7
  46. package/dist/engine-DS1cysJy.d.ts.map +1 -0
  47. package/dist/{eval-campaign-Cs-7MiCs.js → eval-campaign-aYdtjtJR.js} +4 -4
  48. package/dist/{eval-campaign-Cs-7MiCs.js.map → eval-campaign-aYdtjtJR.js.map} +1 -1
  49. package/dist/{exact-types-B7LC1EyX.d.ts → exact-types-BZDe0W2D.d.ts} +2 -2
  50. package/dist/{exact-types-B7LC1EyX.d.ts.map → exact-types-BZDe0W2D.d.ts.map} +1 -1
  51. package/dist/experiment/index.d.ts +27 -477
  52. package/dist/experiment/index.d.ts.map +1 -1
  53. package/dist/experiment/index.js +95 -559
  54. package/dist/experiment/index.js.map +1 -1
  55. package/dist/{experiment-tracker-B3TiF5-u.d.ts → experiment-tracker-C7PfnF4b.d.ts} +2 -2
  56. package/dist/{experiment-tracker-B3TiF5-u.d.ts.map → experiment-tracker-C7PfnF4b.d.ts.map} +1 -1
  57. package/dist/{external-optimizer-process-Dlz8YxrT.js → external-optimizer-process-QDRURJAM.js} +3 -3
  58. package/dist/{external-optimizer-process-Dlz8YxrT.js.map → external-optimizer-process-QDRURJAM.js.map} +1 -1
  59. package/dist/{external-optimizer-subprocess-q3VzlGAO.js → external-optimizer-subprocess-D4dzUBZI.js} +3 -2
  60. package/dist/{external-optimizer-subprocess-q3VzlGAO.js.map → external-optimizer-subprocess-D4dzUBZI.js.map} +1 -1
  61. package/dist/{feedback-trajectory-eHWNv5Aj.d.ts → feedback-trajectory-CXmtITBo.d.ts} +3 -3
  62. package/dist/{feedback-trajectory-eHWNv5Aj.d.ts.map → feedback-trajectory-CXmtITBo.d.ts.map} +1 -1
  63. package/dist/hosted/index.d.ts +2 -2
  64. package/dist/hosted/index.d.ts.map +1 -1
  65. package/dist/hosted/index.js +1 -1
  66. package/dist/{index-BxWvILU8.d.ts → index-Bp_6sj3x.d.ts} +109 -56
  67. package/dist/index-Bp_6sj3x.d.ts.map +1 -0
  68. package/dist/{index-e7LXeRVa.d.ts → index-CJ3LhKIX.d.ts} +2 -2
  69. package/dist/{index-e7LXeRVa.d.ts.map → index-CJ3LhKIX.d.ts.map} +1 -1
  70. package/dist/{index-CbLmrWCa.d.ts → index-DNntP4ch.d.ts} +8 -8
  71. package/dist/{index-CbLmrWCa.d.ts.map → index-DNntP4ch.d.ts.map} +1 -1
  72. package/dist/{index-DxNYmx4a.d.ts → index-DoykkxW0.d.ts} +11 -11
  73. package/dist/{index-DxNYmx4a.d.ts.map → index-DoykkxW0.d.ts.map} +1 -1
  74. package/dist/index.d.ts +28 -28
  75. package/dist/index.js +25 -16
  76. package/dist/index.js.map +1 -1
  77. package/dist/{insight-report-DETqPc_A.d.ts → insight-report-D1qa0HWs.d.ts} +9 -5
  78. package/dist/{insight-report-DETqPc_A.d.ts.map → insight-report-D1qa0HWs.d.ts.map} +1 -1
  79. package/dist/{integrity-DsHWCebQ.js → integrity-DH5ng72x.js} +2 -2
  80. package/dist/{integrity-DsHWCebQ.js.map → integrity-DH5ng72x.js.map} +1 -1
  81. package/dist/{integrity-BKTcA-HP.d.ts → integrity-rGOfSUle.d.ts} +2 -2
  82. package/dist/{integrity-BKTcA-HP.d.ts.map → integrity-rGOfSUle.d.ts.map} +1 -1
  83. package/dist/{ledger-core-Cs9f7385.js → journal-Cs9f7385.js} +1 -1
  84. package/dist/journal-Cs9f7385.js.map +1 -0
  85. package/dist/{judge-calibration-C5CbMYce.d.ts → judge-calibration-DFtEMlde.d.ts} +31 -2
  86. package/dist/judge-calibration-DFtEMlde.d.ts.map +1 -0
  87. package/dist/{judge-calibration-BnpVKtnb.js → judge-calibration-DYmaBtJr.js} +48 -2
  88. package/dist/{judge-calibration-BnpVKtnb.js.map → judge-calibration-DYmaBtJr.js.map} +1 -1
  89. package/dist/ledger-core/index.d.ts +1 -1
  90. package/dist/ledger-core/index.js +1 -1
  91. package/dist/{llm-judge-v80Kmu9g.js → llm-judge-DEFZeSiu.js} +645 -456
  92. package/dist/llm-judge-DEFZeSiu.js.map +1 -0
  93. package/dist/{matrix-DeMmnWrP.d.ts → matrix-CyhW-vgJ.d.ts} +2 -2
  94. package/dist/{matrix-DeMmnWrP.d.ts.map → matrix-CyhW-vgJ.d.ts.map} +1 -1
  95. package/dist/meta-eval/index.d.ts +138 -7
  96. package/dist/meta-eval/index.d.ts.map +1 -1
  97. package/dist/meta-eval/index.js +245 -97
  98. package/dist/meta-eval/index.js.map +1 -1
  99. package/dist/{mint-Cc1_zwRQ.js → mint-ySIIkKlV.js} +2 -2
  100. package/dist/{mint-Cc1_zwRQ.js.map → mint-ySIIkKlV.js.map} +1 -1
  101. package/dist/multishot/golden/index.d.ts +1 -1
  102. package/dist/multishot/index.d.ts +2 -2
  103. package/dist/openapi.json +1 -1
  104. package/dist/outcome-store-BXlkwMPR.js +131 -0
  105. package/dist/outcome-store-BXlkwMPR.js.map +1 -0
  106. package/dist/{outcome-store-BYHIuO0e.d.ts → outcome-store-CNt4iZ67.d.ts} +18 -25
  107. package/dist/outcome-store-CNt4iZ67.d.ts.map +1 -0
  108. package/dist/{paired-promotion-decision-CGzg0cI_.d.ts → paired-promotion-decision-DPsMQm-0.d.ts} +13 -7
  109. package/dist/{paired-promotion-decision-CGzg0cI_.d.ts.map → paired-promotion-decision-DPsMQm-0.d.ts.map} +1 -1
  110. package/dist/pipelines/index.js +1 -1
  111. package/dist/{produced-state-Cv0kJJuP.js → produced-state-BHboMaab.js} +3 -3
  112. package/dist/{produced-state-Cv0kJJuP.js.map → produced-state-BHboMaab.js.map} +1 -1
  113. package/dist/profile-cell.d.ts +1 -1
  114. package/dist/profile-cell.js +1 -1
  115. package/dist/{promotion-policy-DWOm70gx.js → promotion-policy-CDMMxzb6.js} +28 -40
  116. package/dist/promotion-policy-CDMMxzb6.js.map +1 -0
  117. package/dist/{registry-ByVld1-5.d.ts → registry-BRbB6Y0v.d.ts} +4 -4
  118. package/dist/{registry-ByVld1-5.d.ts.map → registry-BRbB6Y0v.d.ts.map} +1 -1
  119. package/dist/{release-confidence-BAcNYOf1.d.ts → release-confidence-BcqeQTHW.d.ts} +3 -3
  120. package/dist/{release-confidence-BAcNYOf1.d.ts.map → release-confidence-BcqeQTHW.d.ts.map} +1 -1
  121. package/dist/{release-confidence-BcGCclTB.js → release-confidence-DMg8n18l.js} +2 -2
  122. package/dist/{release-confidence-BcGCclTB.js.map → release-confidence-DMg8n18l.js.map} +1 -1
  123. package/dist/{report-command-DKlXfU5r.js → report-command-V1ecVgAv.js} +27 -3
  124. package/dist/report-command-V1ecVgAv.js.map +1 -0
  125. package/dist/reporting.d.ts +4 -4
  126. package/dist/reporting.js +3 -3
  127. package/dist/{researcher-jsW1X94L.d.ts → researcher-64T49THL.d.ts} +6 -6
  128. package/dist/{researcher-jsW1X94L.d.ts.map → researcher-64T49THL.d.ts.map} +1 -1
  129. package/dist/{reward-hacking-ZXEi9VCq.d.ts → reward-hacking-uzO_ihep.d.ts} +2 -2
  130. package/dist/{reward-hacking-ZXEi9VCq.d.ts.map → reward-hacking-uzO_ihep.d.ts.map} +1 -1
  131. package/dist/rl.d.ts +53 -99
  132. package/dist/rl.d.ts.map +1 -1
  133. package/dist/rl.js +182 -169
  134. package/dist/rl.js.map +1 -1
  135. package/dist/rollout/index.d.ts +1 -1
  136. package/dist/rollout/index.js +2 -2
  137. package/dist/{rollout-DmoJVqrF.js → rollout-B-UF5R6w.js} +2 -2
  138. package/dist/{rollout-DmoJVqrF.js.map → rollout-B-UF5R6w.js.map} +1 -1
  139. package/dist/rubric-predictive-validity-Bmj2_cll.d.ts +79 -0
  140. package/dist/rubric-predictive-validity-Bmj2_cll.d.ts.map +1 -0
  141. package/dist/rubric-predictive-validity-CCK-1B7w.js +178 -0
  142. package/dist/rubric-predictive-validity-CCK-1B7w.js.map +1 -0
  143. package/dist/{run-record-DTv1MdjK.d.ts → run-record-BiTWauyO.d.ts} +2 -2
  144. package/dist/{run-record-DTv1MdjK.d.ts.map → run-record-BiTWauyO.d.ts.map} +1 -1
  145. package/dist/run-record-Br-Yzt_k.js +464 -0
  146. package/dist/run-record-Br-Yzt_k.js.map +1 -0
  147. package/dist/{run-record-DQpSf7t-.js → run-record-DualPTn2.js} +2 -2
  148. package/dist/{run-record-DQpSf7t-.js.map → run-record-DualPTn2.js.map} +1 -1
  149. package/dist/{semantic-concept-judge-Bi6_iGqg.js → semantic-concept-judge-Bm5JDEKO.js} +3 -3
  150. package/dist/{semantic-concept-judge-Bi6_iGqg.js.map → semantic-concept-judge-Bm5JDEKO.js.map} +1 -1
  151. package/dist/{sequential-B5gXgcyp.js → sequential-DAsyV2T9.js} +42 -25
  152. package/dist/sequential-DAsyV2T9.js.map +1 -0
  153. package/dist/{series-convergence-DeG33RpC.d.ts → series-convergence-BnMs_uAr.d.ts} +3 -3
  154. package/dist/{series-convergence-DeG33RpC.d.ts.map → series-convergence-BnMs_uAr.d.ts.map} +1 -1
  155. package/dist/{skillopt-optimization-method-C3oYul8v.js → skillopt-optimization-method-CL_0aArC.js} +5 -5
  156. package/dist/{skillopt-optimization-method-C3oYul8v.js.map → skillopt-optimization-method-CL_0aArC.js.map} +1 -1
  157. package/dist/{statistical-heldout-0La5ZTlv.d.ts → statistical-heldout-CpVd6FmY.d.ts} +207 -144
  158. package/dist/statistical-heldout-CpVd6FmY.d.ts.map +1 -0
  159. package/dist/{store-tool-spans-4J1EDElP.d.ts → store-tool-spans-Dt-YdAuE.d.ts} +6 -6
  160. package/dist/{store-tool-spans-4J1EDElP.d.ts.map → store-tool-spans-Dt-YdAuE.d.ts.map} +1 -1
  161. package/dist/{summary-report-gMrbYawB.d.ts → summary-report-D1h4dlrK.d.ts} +3 -3
  162. package/dist/{summary-report-gMrbYawB.d.ts.map → summary-report-D1h4dlrK.d.ts.map} +1 -1
  163. package/dist/{summary-report-B16xy9Kd.js → summary-report-e-MaOAHV.js} +2 -2
  164. package/dist/{summary-report-B16xy9Kd.js.map → summary-report-e-MaOAHV.js.map} +1 -1
  165. package/dist/supervisor-run/index.d.ts +4 -2
  166. package/dist/supervisor-run/index.d.ts.map +1 -1
  167. package/dist/supervisor-run/index.js +3 -3
  168. package/dist/{terminal-record-Ce9_UjRz.js → terminal-record-BtPwKTSr.js} +58 -26
  169. package/dist/terminal-record-BtPwKTSr.js.map +1 -0
  170. package/dist/{tool-groups-2QA0S7dK.d.ts → tool-groups-B2bSNaJB.d.ts} +3 -3
  171. package/dist/tool-groups-B2bSNaJB.d.ts.map +1 -0
  172. package/dist/{tool-waste-B9tdWV6g.js → tool-waste-C7MU9u1e.js} +2 -2
  173. package/dist/{tool-waste-B9tdWV6g.js.map → tool-waste-C7MU9u1e.js.map} +1 -1
  174. package/dist/trace-repair/index.d.ts +2 -2
  175. package/dist/traces.d.ts +6 -6
  176. package/dist/traces.js +1 -1
  177. package/dist/{types-BmlkCrg0.d.ts → types-BvZoPTGa.d.ts} +3 -3
  178. package/dist/{types-BmlkCrg0.d.ts.map → types-BvZoPTGa.d.ts.map} +1 -1
  179. package/dist/{types-gvRsyJLh.d.ts → types-CBbLtr2J.d.ts} +38 -3
  180. package/dist/{types-gvRsyJLh.d.ts.map → types-CBbLtr2J.d.ts.map} +1 -1
  181. package/dist/{types-C34V4Vto.d.ts → types-CS0qk_Yp.d.ts} +4 -4
  182. package/dist/{types-C34V4Vto.d.ts.map → types-CS0qk_Yp.d.ts.map} +1 -1
  183. package/dist/{types-DzuaM493.d.ts → types-D7gEdPoQ.d.ts} +3 -3
  184. package/dist/{types-DzuaM493.d.ts.map → types-D7gEdPoQ.d.ts.map} +1 -1
  185. package/dist/{types-vUdAx2Cj.d.ts → types-lPkDQNqJ.d.ts} +20 -2
  186. package/dist/{types-vUdAx2Cj.d.ts.map → types-lPkDQNqJ.d.ts.map} +1 -1
  187. package/dist/wire/index.d.ts +2 -2
  188. package/docs/adapters-observability.md +14 -0
  189. package/docs/campaign-proposers.md +86 -128
  190. package/docs/charter.md +108 -112
  191. package/docs/concepts.md +157 -69
  192. package/docs/design/mlbenchmarks-book-review.md +440 -0
  193. package/docs/design/mlbenchmarks-review/observations.json +713 -0
  194. package/docs/design/mlbenchmarks-review/probes.mts +476 -0
  195. package/docs/design/mlbenchmarks-review/sources.json +200 -0
  196. package/docs/design/self-improvement-evidence-audit.md +263 -0
  197. package/docs/design.md +2 -1
  198. package/docs/eval-surface-map.md +95 -42
  199. package/docs/evaluation-integrity.md +220 -0
  200. package/docs/experiment.md +111 -55
  201. package/docs/feature-guide.md +5 -6
  202. package/docs/hosted-ingest-spec.md +4 -11
  203. package/docs/insight-report.md +187 -455
  204. package/docs/outcome-validity.md +182 -0
  205. package/docs/product-eval-adoption.md +1 -2
  206. package/docs/research-report-methodology.md +7 -7
  207. package/docs/search-history-receipts.md +8 -0
  208. package/docs/statistical-evidence.md +129 -0
  209. package/docs/verdicts.md +76 -49
  210. package/package.json +1 -1
  211. package/dist/agent-profile-cell-0gSi5ffD.js.map +0 -1
  212. package/dist/agent-profile-cell-CTOZJUuE.d.ts.map +0 -1
  213. package/dist/campaign-evidence-D8DBLqLI.js.map +0 -1
  214. package/dist/client-BlLY6o2w.js.map +0 -1
  215. package/dist/define-agent-eval-Cx4Ls9ta.d.ts.map +0 -1
  216. package/dist/define-agent-eval-Dzidv34q.js.map +0 -1
  217. package/dist/engine-CX8ReXkn.d.ts.map +0 -1
  218. package/dist/index-BxWvILU8.d.ts.map +0 -1
  219. package/dist/judge-calibration-C5CbMYce.d.ts.map +0 -1
  220. package/dist/ledger-core-Cs9f7385.js.map +0 -1
  221. package/dist/llm-judge-v80Kmu9g.js.map +0 -1
  222. package/dist/outcome-store-BYHIuO0e.d.ts.map +0 -1
  223. package/dist/outcome-store-ChBKlTd_.js +0 -75
  224. package/dist/outcome-store-ChBKlTd_.js.map +0 -1
  225. package/dist/promotion-policy-DWOm70gx.js.map +0 -1
  226. package/dist/report-command-DKlXfU5r.js.map +0 -1
  227. package/dist/rubric-predictive-validity-2D5Gw9z9.js +0 -131
  228. package/dist/rubric-predictive-validity-2D5Gw9z9.js.map +0 -1
  229. package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts +0 -75
  230. package/dist/rubric-predictive-validity-Dl1dvKCv.d.ts.map +0 -1
  231. package/dist/run-record-CR63CpHK.js +0 -216
  232. package/dist/run-record-CR63CpHK.js.map +0 -1
  233. package/dist/sequential-B5gXgcyp.js.map +0 -1
  234. package/dist/statistical-heldout-0La5ZTlv.d.ts.map +0 -1
  235. package/dist/terminal-record-Ce9_UjRz.js.map +0 -1
  236. package/dist/tool-groups-2QA0S7dK.d.ts.map +0 -1
package/docs/concepts.md CHANGED
@@ -2,52 +2,50 @@
2
2
 
3
3
  `agent-eval` records agent runs, scores their outputs, compares variants, and applies caller-defined release rules.
4
4
 
5
- A model can say a task is complete while the build fails, a browser flow is broken, an integration is disconnected, or required sources are missing.
5
+ An agent can claim success while a build, browser flow, or integration fails.
6
+ Required source evidence can also be missing.
6
7
  This package lets code, model judges, and human feedback check those outcomes through the same run format.
7
8
 
8
9
  ## The top-level functions
9
10
 
10
- Start with `/contract` and `defineAgentEval()` for a new integration.
11
- Use the lower-level functions when you need direct control over execution, storage, or statistics.
11
+ Start with `defineAgentEval()` from `/contract` for one agent, judge, case set, and baseline surface.
12
+ Its `evaluate()` method returns campaign measurements.
13
+ Its `improve()` method searches and returns a final comparison with a release decision.
14
+ Use `selfImprove()` directly when you do not need shared configuration.
12
15
 
13
- | Function | When to call it | What you give it | What you get back |
14
- |---|---|---|---|
15
- | **`defineAgentEval()`** | You have scenarios, an agent, a judge, and a baseline surface, and you want one object you can score or improve. | scenarios, agent, judge, baseline surface | `{ evaluate(), improve() }` where `evaluate()` returns a campaign result and `improve()` returns a report |
16
- | **`selfImprove()`** | You want candidate generation, scoring, and a release decision in one call. | scenarios, agent, judge, baseline surface | report, winner surface, and a `gateDecision` (see below) |
17
- | **`loadEvalFixtureScenarios()`** | You want agents to add evals as folders with `PROMPT.md`, checks, and starter files. | `evals/<name>/PROMPT.md + EVAL.ts + package.json` | `Scenario[]` that runs through `runCampaign`; pair with `planEvalFixtureRun()` before spending tokens |
18
- | **`analyzeRuns()`** | You have existing runs and do not need to invoke an agent. | `RunRecord[]` and options | `InsightReport` |
19
- | **Intake adapters** (`fromFeedbackTable`, `fromOtelSpans`) | Your data isn't already in `RunRecord` shape: it's in Obsidian, Sheets, an OTel collector, etc. | source-specific input | `RunRecord[]` ready to pipe into `analyzeRuns()` |
20
- | **`sealExperiment()` / `openSealedExperiment()`** | The result must convince a reader who does not trust you, so the rules must be fixed before the data arrives. | arms, admission funnel, estimand, interval, decision table | a hashed rule tree plus executors that can run no other rule ([`experiment.md`](./experiment.md)) |
21
- | **`runEquivalenceCheck()`** | The work has no held-out test suite, so no answer key exists to grade against. | a claim, two blind arms, an injected checker | a certification naming who vouched and how it can fail ([`verification-strategies.md`](./verification-strategies.md)) |
22
- | **`AnalystRegistry.runExact()`** | A batch of runs failed and you need cited findings, with the caller owning every execution choice. | recorded evidence, a declared analyst list | findings with evidence references, an execution plan, and a receipt ([`trace-analysis.md`](./trace-analysis.md)) |
16
+ Use `analyzeRuns()` from `/contract` for existing `RunRecord[]` evidence.
17
+ [The README workflows](../README.md#choose-a-workflow) link to runnable examples.
18
+ [The surface map](./eval-surface-map.md) lists lower-level execution and analysis APIs.
23
19
 
24
- See [`customer-journeys.md`](./customer-journeys.md) for runnable paths from existing logs, human ratings, and a callable agent.
25
- The [README front-door table](../README.md#which-front-door) lists every callable entry point with a runnable example.
20
+ Root `Scenario`, `JudgeScore`, and `GateDecision` use the same definitions as `/contract` and `/campaign`.
21
+ The product-judging shapes have explicit root names: `ProductScenario` and `DimensionJudgeScore`.
22
+ The separate `HeldOutGate` class returns `HeldOutGateDecision` over `RunRecord` comparisons.
26
23
 
27
24
  ### The five release decisions
28
25
 
29
- `selfImprove()` and every gate return a `GateDecision`, not a two-way ship/hold flag.
30
- Folding the last three into `hold` throws away the action each one names.
26
+ `selfImprove()` returns a `gateDecision` from the campaign `GateDecision` union.
27
+ Keep its five values distinct because they require different actions.
31
28
 
32
29
  | Decision | What it means | What to do next |
33
30
  |---|---|---|
34
- | `ship` | Every gate passed on sufficient evidence. | Release the candidate. |
35
- | `hold` | A gate failed on sufficient evidence. | Reject this candidate. |
36
- | `need_more_work` | A gate could not decide: the evidence was missing, or the paired sample was too small to claim significance. | Gather more runs, then gate again. |
31
+ | `ship` | All required configured checks support release. | Review the evidence and release the candidate. |
32
+ | `hold` | The gate does not justify release. A required check can fail or lack sufficient evidence. | Inspect the contributions to distinguish regression from an unresolved comparison. |
33
+ | `need_more_work` | The gate reports that more work or evidence is required. | Address the reported gap before another decision. |
37
34
  | `model_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the model. | Handle it; no gate in this package emits it. |
38
35
  | `arch_ceiling` | Reserved for a caller-supplied gate that attributes the limit to the architecture. | Handle it; no gate in this package emits it. |
39
36
 
40
- The last two are part of the taxonomy and of the composition order, but no built-in gate returns them today.
41
- Handle all five anyway: a caller's own gate may return either, and the type will not let you ignore them.
37
+ Handle all five values when accepting caller-supplied gates.
42
38
 
43
- `need_more_work` is not a quiet `hold`.
44
- "Gather more evidence" and "reject this candidate" are different actions, and folding the first into the second abandons a real gain that was only underpowered.
39
+ Read the contributing checks before interpreting a refused release.
40
+ An unresolved comparison does not establish that the candidate is worse.
41
+ Absent optional checks remain `not_evaluated`, including when the required checks support `ship`.
45
42
 
46
43
  When gates are composed, `ship` requires every gate to ship.
47
44
  Otherwise the strongest hold wins, in this order: `arch_ceiling`, `model_ceiling`, `hold`, `need_more_work`.
48
45
 
49
- `analyzeRuns()` and the high-level contract return the same `InsightReport` shape.
50
- It contains score distributions, paired lift intervals, judge agreement, cost, failure clusters, contamination checks, outcome correlation, and recommendations.
46
+ `analyzeRuns()` returns an `InsightReport`; `selfImprove()` includes one in its result.
47
+ The report includes score distributions, cost, and recommendations.
48
+ Paired lift, failure clusters, contamination checks, and outcome associations require their corresponding inputs.
51
49
  [`insight-report.md`](./insight-report.md) defines every field.
52
50
 
53
51
  ## Package Boundary
@@ -68,7 +66,8 @@ Use the profile improvement functions from `/contract` when a host owns immutabl
68
66
 
69
67
  This API never activates a candidate or runs an agent itself.
70
68
  The host owns authorization, billing, task isolation, profile materialization, execution, and durable evidence.
71
- The first portable profile contract accepts prompt and skill changes only; a host must add its own exact-state adapter before measuring tools, MCP servers, hooks, subagents, or external knowledge.
69
+ The portable profile contract accepts prompt and skill changes.
70
+ A host needs an adapter for exact state before measuring tools, MCP servers, hooks, subagents, or external knowledge.
72
71
 
73
72
  ## Main Objects
74
73
 
@@ -85,7 +84,7 @@ Traces, datasets, optimization, statistics, and reports build on these objects.
85
84
 
86
85
  Every entry in `GateResult.contributingGates` has a `status` of `pass`, `fail`, or `not_evaluated`.
87
86
  `pass` and `fail` mean the check ran with sufficient input.
88
- `not_evaluated` means the check lacked enough evidence to run.
87
+ `not_evaluated` means the check was unconfigured or lacked required input or evidence.
89
88
  `defaultProductionGate` always requires held-out significance.
90
89
  Its other checks are optional until their input is configured or their name is included in `requiredChecks`.
91
90
  A required check with missing or insufficient evidence remains `not_evaluated` and holds the release decision.
@@ -93,6 +92,29 @@ An absent optional check records `not_evaluated` and never appears as a successf
93
92
  Run history is shared input only.
94
93
  Enable reward-hacking and canary monitoring independently with `rewardHacking` and `canary`.
95
94
 
95
+ `selfImprove()` uses `defaultProductionGate()` unless you pass a custom `gate`.
96
+ The optional `paretoSignificanceGate()` applies each objective's regression floor to its deciding confidence interval.
97
+ Its confidence level applies per objective; the gate does not adjust for multiple objectives or repeated comparisons.
98
+ It pairs execution cells directly; custom gates must apply any grouping required by a claim.
99
+ For detected binary outcomes, that interval accounts for uncertainty even when every observed pair agrees.
100
+ At 95% confidence, 20 matching all-positive binary pairs leave approximately 16 percentage points of uncertainty in either direction.
101
+ A declared five-point regression tolerance therefore holds that candidate; 100 matching pairs narrow the interval enough to clear that floor.
102
+ Another objective must still show a significant gain before promotion.
103
+ An axis labeled `regressed` has not cleared its configured floor.
104
+ Its interval can permit a regression without demonstrating one.
105
+ An interval with zero width or non-finite bounds produces an `indeterminate` axis verdict and a `not_evaluated` check.
106
+ If no other axis breaches its regression floor, the gate returns `need_more_work`.
107
+ This includes undeclared all-zero outcomes, whose binary scale cannot be inferred, and continuous observations with constant paired differences.
108
+ Additional identical observations do not resolve an unknown outcome scale or a collapsed bootstrap interval.
109
+ Consumers with custom policies must handle `indeterminate` as unresolved evidence.
110
+
111
+ Set an objective's `binaryScale: 1` for known `{0, 1}` observations, including all-zero error indicators.
112
+ Use `binaryScale: 100` for `{0, 100}` observations; its default regression tolerance is 5.
113
+ The shared `decidePairedPromotion()` function accepts the same declaration.
114
+ The scale must be positive and finite, and every paired cell score must be zero or that scale.
115
+ Declared binary outcomes use the risk-difference mean and reject `statistic: 'median'`.
116
+ The declaration identifies the outcome support; it does not reduce sample requirements or supply missing observations.
117
+
96
118
  When the thing being evaluated is an agent that should keep working, use
97
119
  [`runAgentControlLoop`](./control-runtime.md). It turns validators into a
98
120
  runtime loop: observe typed state, validate it, decide the next action, act,
@@ -115,19 +137,19 @@ that can seed memory, replay scenarios, and optimization.
115
137
  | **Layer** | One stage of a verifier pipeline (install, typecheck, build, semantic, …). |
116
138
  | **Finding** | A specific issue a judge found: file, line, severity, message. |
117
139
  | **Trace store** | The append-only log of every span/event during a run. Replay = read this back. |
118
- | **Composite score** | A 0..1 number combining all dimensions. The single number you gate on. |
119
- | **Rubric version** | A stable hash of the rubric. Scores from different rubric versions are not comparable. |
140
+ | **Composite score** | An aggregate on the judge's declared scale. Gates must use thresholds on that scale. |
141
+ | **Rubric version** | A stable hash of the rubric. Comparing revisions requires calibration against shared independent labels. |
120
142
 
121
143
  ### Running an evaluation
122
144
 
123
145
  | Term | Plain English |
124
146
  |---|---|
125
- | **Case** (`Scenario`) | One task the agent must do. The unit every score is per. |
147
+ | **Case** (`Scenario`) | One task the agent must do. Variants can share an independent source unit. |
126
148
  | **Surface** | The value being changed: a prompt, a skill, or a serialized configuration. |
127
149
  | **Dispatch** | The function that runs your agent on one case and returns the artifact. |
128
150
  | **Campaign** | One complete pass of every case, executed, scored, and cached under a run directory. |
129
- | **Cell** | One (case × replicate) of a campaign. Cells are cached, so a rerun skips the ones that finished. |
130
- | **Receipt** | The record of what one paid call actually cost, in dollars and tokens. Absent when nothing measured it. |
151
+ | **Cell** | One (case × replicate) of a campaign. With caching enabled, matching completed cells can be reused. |
152
+ | **Receipt** | A settled call record with cost, token usage, and flags for unknown measurements. |
131
153
  | **Cost ledger** | The spend account receipts are written to. A capped ledger refuses a call that would exceed the cap. |
132
154
  | **Provenance** | Where a number came from: the package version, the source revision, the run identity, the exact attempt. |
133
155
  | **`RunRecord`** | The analysis-time projection of one run: who ran, on what, with which seed, at what cost, and what it scored. |
@@ -145,20 +167,52 @@ that can seed memory, replay scenarios, and optimization.
145
167
  | **Selection cases** | Evidence the optimizer reads to choose among its candidates. |
146
168
  | **Final cases** | Held back from the optimizer entirely. They produce the reported lift. |
147
169
 
148
- The three-way split is the reason a reported lift means anything.
149
- An optimizer that saw the final cases can score well on them without the agent getting better.
170
+ Keep scenario identifiers disjoint across the three partitions.
171
+ For new-unit claims or fresh final evidence, also keep source units separate between development and final cases.
172
+ Fixed-roster development can share sources while retaining independent-unit counts in its reports.
173
+ Renamed variants from one incident can leak information across splits.
174
+ A final comparison supports only the declared population and measured conditions.
150
175
 
151
176
  ### Proving a result
152
177
 
153
- | Term | Plain English |
178
+ | Term | Meaning |
154
179
  |---|---|
155
- | **Experiment** | The rules — arms, funnel, estimand, interval, decision written as data before the data arrives. |
156
- | **Seal** | A hash of that whole rule tree. The execution surface accepts no rule outside it. |
157
- | **Estimand** | The exact quantity being measured, for example the paired difference in pass rate. |
158
- | **Funnel** | The denominator chain: how many rows entered, what each stage removed, and how many remain. |
159
- | **Verification strategy** | One of ten ways to certify a result, each with a documented way it can certify a wrong one. |
160
- | **Certification** | Who vouched for a verdict, with what checker version, and what the checker did not check. |
161
- | **Analyst** | A function that reads recorded evidence and returns findings that cite it. |
180
+ | **Claim** | The intended use, population, sampling frame, independent unit, generalization target, and optional minimum useful effect. |
181
+ | **Independent unit** | The source task, incident, or family that contributes one independent observation to an inference. |
182
+ | **Experiment** | Arms, admission, estimand, interval, and decision rules declared before results are inspected. |
183
+ | **Seal** | A digest binding the experiment's rules and claim to the executed specification. |
184
+ | **Estimand** | The quantity being estimated, such as the mean difference across independent task families. |
185
+ | **Funnel** | Counts of input rows, exclusions at each stage, and retained evidence. |
186
+ | **Final-evidence reservation** | A durable claim on source units before search; exposure records the evaluated candidates before dispatch. |
187
+ | **Verification strategy** | A method of checking a result, with documented assumptions and failure modes. |
188
+ | **Certification** | The checker identity, strategy, unverified assumptions, and evidence associated with a verdict. |
189
+ | **Analyst** | A function that reads recorded evidence and returns cited findings. |
190
+
191
+ Use `defineEvaluationClaim()` from `/experiment` to declare what a result can describe.
192
+ Pass it as the top-level `claim` when improving a surface or comparing optimization methods.
193
+ `fixed-roster` concerns the listed units; `new-units` attempts to generalize to further units from the declared population.
194
+ A declared sampling frame does not itself establish representative sampling.
195
+
196
+ Count repetitions separately from independent units.
197
+ For example, 100 retries of one incident produce 100 observations and one independent incident.
198
+ Campaign aggregate `n` describes its observed scores.
199
+ The default improvement gate and method comparisons retain their independent-unit and paired-cell counts.
200
+
201
+ Set `minimumEffect` when the decision concerns a practically useful change.
202
+ Development and absolute-rate claims can omit it.
203
+ Sealed power checks assess the declared effect.
204
+ A design that detects only much larger effects cannot pass that adequacy check.
205
+ Inference also needs the interval and clustering rule to match the claim.
206
+
207
+ Unit-aware comparison does not require a final-evidence ledger.
208
+ For fresh confirmation, add `finalEvidence: { ledger, requestId, evaluatorDigest }` alongside the top-level `claim`.
209
+ This reserves final units before search.
210
+ Use one durable ledger across related campaigns and stable source identities across renamed variants.
211
+ Exposure remains recorded if measurement fails or the process stops.
212
+ The host enforces access isolation; the ledger cannot inspect reads outside this workflow.
213
+
214
+ See [evaluation integrity](./evaluation-integrity.md) for claims, final evidence, and evaluator admission.
215
+ [Registered experiments](./experiment.md) describes seals, decision rules, and refusal artifacts.
162
216
 
163
217
  ## The feedback trajectory loop
164
218
 
@@ -178,7 +232,8 @@ rows, optimizer rows, and held-out examples for overfit checks.
178
232
 
179
233
  ## Code Generator Eval
180
234
 
181
- When the artifact is generated code, agent-eval scores it at three independent layers. Each layer fails differently, and you want to know which one broke:
235
+ Generated-code evaluations can score the agent session, the build, and the running application.
236
+ Each layer detects different failures:
182
237
 
183
238
  ```
184
239
  L0 builder Did the agent's session itself work?
@@ -190,21 +245,22 @@ L1 app-build Does the artifact build / typecheck / test?
190
245
 
191
246
 
192
247
  L2 app-runtime Does the artifact actually run end-to-end?
193
- (Dynamic signal: only worth checking if L1 passed.)
248
+ (Dynamic signal: requires a runnable application.)
194
249
  ```
195
250
 
196
- `BuilderSession` orchestrates this. It opens at `startChat`, runs the build at `ship`, runs the runtime check at `runAppScenario`. Each layer emits a trace span. Composite score aggregates them with `scoreProject`.
197
-
198
- Why three? Because each catches a different failure mode:
199
- - L0 misses: agent crashed mid-generation, you have a half-written file.
200
- - L1 misses: files exist but typecheck fails. LLM judges can't reliably catch this.
201
- - L2 misses: code compiles but does the wrong thing at runtime.
251
+ `BuilderSession` coordinates these checks.
252
+ It opens at `startChat`, runs the build at `ship`, and runs the application check at `runAppScenario`.
253
+ Each layer emits a trace span.
254
+ `scoreProject` reports each layer's score and whether the required measurements are complete.
202
255
 
203
- If you only check one layer, you ship the bugs that the other two layers would have caught.
256
+ - L0: The agent crashed during generation and left an incomplete artifact.
257
+ - L1: Files exist but do not typecheck or build.
258
+ - L2: Code compiles but behaves incorrectly when executed.
204
259
 
205
260
  ## How rubrics work
206
261
 
207
262
  A rubric describes:
263
+
208
264
  1. **Dimensions**: the axes you score on (e.g. `buyer_quality`, `voice`, `signal`).
209
265
  2. **Weights**: how to combine dimensions into a composite (`0.5 * buyer_quality + 0.3 * voice + 0.2 * signal`).
210
266
  3. **Failure modes**: named patterns the judge looks for ("ai-cadence", "vague-claim").
@@ -214,7 +270,10 @@ A rubric describes:
214
270
  Built-in rubrics ship in `src/wire/rubrics.ts`, including `anti-slop` for technical-buyer voice.
215
271
  You can also pass the same rubric shape inline at the call site.
216
272
 
217
- A rubric is plain data. The digest of that data, tagged with the scheme that produced it, is the `rubricVersion`. Two scores are only comparable if they used the same `rubricVersion`: change the rubric and you start a new comparison series.
273
+ A rubric is plain data.
274
+ Its digest and encoding scheme identify the `rubricVersion`.
275
+ Changing the rubric starts a new comparison series.
276
+ Evaluate rubric revisions against independent labels before combining their scores.
218
277
 
219
278
  ## How verifiers work
220
279
 
@@ -229,7 +288,7 @@ const verifier = new MultiLayerVerifier([
229
288
  ])
230
289
 
231
290
  const report = await verifier.run({ env })
232
- report.allPass // boolean: every layer passed
291
+ report.allPass // every layer passed and the task measurement is complete
233
292
  report.taskScore // complete task score, or undefined
234
293
  report.blendedScore // diagnostic weighted aggregate, possibly partial
235
294
  report.layers // per-layer status, findings, duration
@@ -238,21 +297,27 @@ report.layers // per-layer status, findings, duration
238
297
  `env` carries the sandbox driver, the working directory, and the harness commands each layer runs.
239
298
 
240
299
  Use `taskScore` when creating task labels or training data.
241
- An errored, timed-out, skipped, or incomplete scoring panel leaves `taskScore` undefined.
300
+ A complete scoring panel needs at least one valid score from a positive-weight layer.
301
+ Positive-weight layers that error, time out, or skip leave `taskScore` undefined.
302
+ Passing layers can omit a numeric score.
303
+ A failed layer contributes only when `failContributesToScore` is enabled and it supplies a valid score.
304
+ Zero-weight layers can still fail `allPass` without removing `taskScore`.
242
305
  Use `blendedScore` only to inspect the measurements that did complete.
243
306
 
244
307
  Two rules that will save you bugs:
245
308
 
246
- 1. **Run both gates.** Build gates catch code that doesn't compile; structural assertions catch missing files. Run both unconditionally: they catch orthogonal failures.
247
-
248
- 2. **Pair LLM judges with build outcomes.** An LLM judge will rate non-compiling code as "looks right" (0.8). Always short-circuit on `buildOutcome.passed === false` before any LLM judging.
309
+ 1. Run build checks and structural assertions.
310
+ They detect different failures.
311
+ 2. Preserve a failed build as a deterministic release failure.
312
+ A semantic score cannot override it.
249
313
 
250
314
  ## Judge calibration
251
315
 
252
- Two questions to answer before trusting any LLM judge:
316
+ Compare a judge with independent labels and other judges before using its scores for decisions:
253
317
 
254
- 1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and worst-N miscalibrations vs a human golden set.
255
- 2. **Does it agree with itself / other judges?** `continuousAgreement(scores)` and `calibrateJudgeContinuous(golden, candidate)` report κ_w + ICC(2,1) + Pearson + Spearman with bootstrap 95% CIs on the raw [0,1] scores.
318
+ 1. **Does it agree with humans?** `calibrateJudge(golden, candidate)` reports Pearson, MAE, integer-rounded κ, and the five largest errors on matched item IDs.
319
+ 2. **Does it agree with other judges?**
320
+ `continuousAgreement()` and `calibrateJudgeContinuous()` report agreement and bootstrap intervals on continuous scores.
256
321
 
257
322
  Each statistic answers a different question:
258
323
 
@@ -262,7 +327,7 @@ Each statistic answers a different question:
262
327
  | Spearman | Do they rank the same way? | The size of any gap |
263
328
  | MAE (mean absolute error) | How far apart are they, on average? | Whether the gap is systematic |
264
329
  | κ (Cohen's kappa) | Do they agree more than chance? | Everything below the rounding step |
265
- | ICC(2,1) | Do they agree in absolute value, not just in shape? | |
330
+ | ICC(2,1) | Do raters agree in absolute score under its variance model? | Shared errors against the intended outcome |
266
331
 
267
332
  Use two flavours of κ for one reason.
268
333
  `calibrateJudge` rounds each score to an integer first.
@@ -273,15 +338,36 @@ ICC(2,1) catches a bias Pearson cannot see.
273
338
  If judge B always scores twice judge A, the two move together perfectly and Pearson stays near 1, while ICC drops.
274
339
  That drop is the signal.
275
340
 
276
- Every reported interval is a bootstrap 95 % interval: the statistic is recomputed on many resamples of the data, and the middle 95 % of those values is the interval.
341
+ ICC and continuous weighted κ have percentile bootstrap intervals, with `ciLevel: 0.95` by default.
342
+ Pearson, Spearman, and MAE are point estimates in these reports.
343
+
344
+ Import calibration and bias functions from `/meta-eval`.
345
+
346
+ | Probe | Input | Observation |
347
+ |---|---|---|
348
+ | `positionalBias()` | The same items judged with their presentation order swapped. | Mean paired score difference by position. |
349
+ | `verbosityBias()` | Output lengths and judge scores. | Correlation between length and score. |
350
+ | `selfPreference()` | Scores grouped by whether judge and output share a model family. | Difference between the group means. |
351
+
352
+ These probes are descriptive diagnostics.
353
+ Length and family groups can also differ in task quality; an observed association alone does not isolate bias.
354
+ Inspect sample counts before interpreting a diagnostic, especially `n: 0`.
355
+
356
+ Use `auditEvaluator()` for admission against predeclared false-acceptance and false-rejection limits.
357
+ Its observation records distinguish fresh controls, development exposure, and unknown judgments.
358
+ It counts source families rather than repeated variants and reports simultaneous exact bounds for both error rates.
359
+ The host must enforce independent authorship and control access.
277
360
 
278
- `verbosityBias` is the one exported bias probe: it finds a judge that rewards length regardless of quality.
279
- The `JudgeInsight` report shape also carries optional `positionalBias` and `selfPreference` fields for caller-computed probes.
280
- No built-in computes those two fields.
361
+ Use `rubricPredictiveValidity()` to compare rubric scores with declared deployment outcomes.
362
+ Specify whether each outcome should increase or decrease.
363
+ The report preserves signed associations, direction-aligned associations, and exclusions.
364
+ An `inverse` association is a reason to investigate; it does not prove that reversing a rubric will improve behavior.
365
+ See [outcome validity](./outcome-validity.md).
281
366
 
282
367
  ## Trace Model
283
368
 
284
- Every operation emits structured spans into a `TraceStore`. A run is a tree:
369
+ Instrumented execution writes structured spans into a `TraceStore`.
370
+ A builder run can have this tree:
285
371
 
286
372
  ```
287
373
  builder-session [span]
@@ -294,7 +380,9 @@ builder-session [span]
294
380
  └── scenario.run [span]
295
381
  ```
296
382
 
297
- Spans are append-only and have stable ids: replay is reading the same store back. OTLP export ships them out for distributed tracing.
383
+ Recorded spans preserve their identifiers and relationships.
384
+ Trace inspection reads this evidence; executable replay separately reruns recorded operations.
385
+ OTLP export sends spans to distributed tracing systems.
298
386
 
299
387
  You usually should not build this tree by hand. Product runtimes,
300
388
  `runAgentControlLoop`, harnesses, and verifiers should emit it while they run.
@@ -317,4 +405,4 @@ release decision.
317
405
  - **Certifying a result with no answer key?** Read [verification-strategies.md](./verification-strategies.md) for the ten-member family and the blind two-arm protocol.
318
406
  - **Reading a verdict someone else produced?** Read [verdicts.md](./verdicts.md) for what `certification` carries and what an absent one means.
319
407
  - **Grading a finding by executing its repair?** Read [trace-repair-grader.md](./trace-repair-grader.md), and [trajectory-replay.md](./trajectory-replay.md) for re-executing a recorded failure.
320
- - **Wondering why this package exists at all?** Read [charter.md](./charter.md) for the four end-states it is built against.
408
+ - **Checking package ownership?** Read [charter.md](./charter.md) for the implemented foundations and host responsibilities.